Commit Graph
205 Commits
Author SHA1 Message Date
Brennan BensonandClaude 24edf0f64b fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm (#23467)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm

When the provider gave no verdict, the structured status projection now derives
one from the newest turn's lifecycle: interrupted -> interruption, unverifiable ->
unconfirmed. Nothing new is journaled, the completion feed stays provider-only, and
the legacy interrupted flag stays a user stop only. Every verdict reader handles
both arms explicitly.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* fix(native-chat): a folded turn a crash cut off reads Interrupted after N

The settled-turn timing now carries the turn's verdict, derived by the same
agentTurnVerdict the status row uses. A turn that ended interrupted with no
provider verdict heads its fold 'Interrupted after N'; a user's stop keeps
'Worked for N'.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* fix(native-chat): the user's close of a chat records the turn it cuts short as their cancellation

The expected-close settle writes outcome cancellation when the user aimed the
stop at this chat: a Stop while the agent starts, the chat's tab closed (the
agentSession.close RPC, or session.tabs.close with a user reason), or /clear.
A quit, an idle eviction, a worktree teardown or an orchestration stop leaves the
turn with no verdict, so it still reads Interrupted.

* test(native-chat): a Claude turn a newer send superseded reads Interrupted

The supersede fires for any send Orca dispatched, the user's or another agent's,
and nothing at that site records the sender, so the turn keeps no verdict and
folds as Interrupted after N.

* fix(native-chat): the user's close records cancellation on the turn the provider settled on its way out

The Codex and Claude adapters settle their open turn as interrupted, with no
verdict, while the host stops them. The expected-close settle then found no
running turn, so a user's close of a mid-turn chat read Interrupted. The stop
now reads the running turns before it reaches the provider and records the
user's cancellation on each one it cut short, unless the provider gave a
verdict of its own. The close-verdict test's adapter now settles its turn on
close the way the real adapters do.

* test(native-chat): update the close and settled-turn expectations for the host-observed verdict

agentSession.close now passes the user's word to the host, and a settled
interrupted turn with no provider verdict carries `interruption`. Also merge a
duplicate import the code-quality gate rejects.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): the user's close cancels a turn whose start landed as the provider stopped

The close read which turns were running before the stop. A turn whose start was
still in flight (a send echo not yet journaled) was absent from that read, so the
provider's verdict-less settle on the way out left it Interrupted. The close now
reads which turns were already over instead, and records the user's cancellation
on every other turn the stop left running or interrupted with no verdict. A turn
cut off earlier, or finished during the stop, keeps its end.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

* refactor(native-chat): the adapter settles the turn a stop cuts with the stop's typed cause

The host hands its stop's cause to the adapter's close. Each adapter settles its own
open turn on 'ended' through one mapping, turnVerdictForChildEnd, and the host's
dead-generation fallback uses the same mapping for any turn no adapter settled. A
user's close or stop of this chat is their cancellation; a quit, eviction, teardown
or an exit the adapter saw first is news.

Deletes the snapshot-and-diff reconstruction (endedTurnItemIds,
userStoppedTurnRevisions, settleUserStoppedTurns) and the requestedByUser flag.
host.close now takes a required cause.

* test(native-chat): a user's close drops the chat's status row like an eviction

* test(native-chat): expect the eviction cause on the host closes of idle release, worker stop and worker discard

The stop's typed cause now travels into host.close and the adapter's close, so these
three non-user closes assert the 'evict' they pass.

* refactor(native-chat): every stop names its cause, so none defaults to the user's cancellation

stopStructuredAgentSessionAgentUnderSerialize defaulted its ending to 'user-stop', which now
settles the cut turn as the user's cancellation. Every caller already passes a cause; the
parameter is now required, and a type-level test fails to compile if the default returns.

* test(native-chat): pin who a chat's session.tabs.close speaks for, older clients' reasonless close included

The mapping lived inline in a type-unchecked file, and only the explicit user reason had a test: an older client's reasonless close, or a lifecycle echo read as the user's, stayed green. It is now one exhaustive, type-checked function with a case per reason.

* style(mobile): draw the unconfirmed dot in the theme's status amber, not an inline hex

* docs(agent-status): the main agent's outcome also carries the host-observed end, interruption or unconfirmed

* fix(native-chat): a chat the user closed while its agent started is not a failed start

A still-starting child the user's close cut counted as a failed start, since only 'user-stop' was excluded: a start-failure row, and queued messages rejected as a provider failure. Whether an ending fails its start is now one exhaustive switch, shared by the failed-start read and the delivery loop's handover: a user's stop or close never does; an exit, a failed attach and the host's own stops still do.

* fix(native-chat): a chat the user closed closes its queued messages, and starts no agent for them

After 876b6989f1 a user's close of a still-starting chat went on like a Stop, so when the close did not complete the delivery loop started a new agent for the message queued behind it. A child's end now has three dispositions, not a failed-start boolean: a user's Stop lets the queue go on, the user's close closes what was queued before it, and any other end fails it. The close is the one a completed close does (the provider-closed rejection, no verdict), applied at the top of each delivery step and ordered against the close so a later send still goes on.

* fix(activity): a crash-cut turn draws the interrupted glyph; only a user's Stop keeps the done check

The Activity page drew every Interrupted row with the done check, which #2569 chose for a user's Stop. With a crash now reading Interrupted, that put a green check on a turn nobody asked to stop. The row's glyph is now an exhaustive switch over the verdict: a cancellation keeps the done check, and an interruption draws the existing interrupted dot. An unconfirmed end already drew its own glyph.

* fix(activity): the Interrupted group header draws the done check only when every row is a user's Stop

A user's Stop and a crash share the Interrupted group, and its header took its newest row's glyph, so a Stop newer than a crash put a green check over the crash. The header is now folded over the group's rows: the done check only when every row draws it, the interrupted dot otherwise.

* fix(native-chat): a failed close of what the user closed starts no agent for it

Rejecting the messages a user's close left queued swallowed a journal write failure, so the delivery step went on to start an agent for a message in a chat the user closed. The rejection now reports whether it landed, and a step whose rejection failed stops instead; the next wake re-derives and retries it. Also pins that the ordering against the close holds only within its epoch, since a later epoch's sequences restart.

* test(native-chat): the idle sweep's stop is an eviction, so its close carries that cause

Main's idle sweep now stops an idle agent through the conversation lifetime, which this branch gives the 'evict' cause; its expectations name it.

* fix(native-chat): a retried stop keeps the cause of the stop it finishes

A user's Stop or close whose wind-down failed after the child was proven gone was finished by the idle sweep as an eviction, so the turn it cut read Interrupted. The owed wind-down now carries its stop's cause, and a retry with no child settles with it.

* test(native-chat): the idle sweep's close of a retrying Claude chat carries the eviction cause

Main's new test expected the adapter close with the session id alone; every stop now names its cause, and the idle sweep's is 'evict'.

* fix(status): a user's Stop marks done on the tab and sidebar; red Interrupted is only a turn cut short by something else

The tab, the worktree card and the sidebar rows drew a Stop with the same red dot as a crash. The
verdict mark now maps a cancellation to done, still saying "Interrupted by user" in the row text,
and the mobile mirror follows. The Activity page keeps grouping a Stop under Interrupted with the
done check, as before.

* test(cross-version): a new phone reads a user's Stop as done; an old phone still draws it interrupted

* fix(native-chat): a user's Stop inside a live Claude chat reads as their cancellation

Stopping a running Claude turn interrupts it and keeps the session, so the turn's end comes from
the CLI's result frame. Claude CLIs before 2.1.91 send that frame with no terminal_reason, and later
ones may still omit it, so the user's own Stop was recorded as a failure with an error row.

Orca now records the stop on the open turn when it sends the interrupt. An error result for that
turn reads as the user's cancellation whatever reason the CLI gives. The stop belongs to that one
turn, so it cannot reach the next, and it is withdrawn when the CLI refuses the interrupt.

* docs(agent-status): a user's stop marks done; name the tab close cause by its type

The reference still said a stop marks a row interrupted and ranks between live work and an
unconfirmed end. A cancellation now marks done, and only a turn cut short by something else ranks
as interrupted. The runtime's tab close restated the close cause's union; it now uses the type.

* test(native-chat): a proven crash reads as an interruption on the status feed and in the chat

A crash the relaunch proves now settles its turn interrupted, and the status feed works the verdict
out from that record, so the restart test expects interruption for a proven crash and unconfirmed
for one it cannot prove, never a cancellation. A chat read before the proof lands reports
unconfirmed, then interruption and a folded "Interrupted after 27s" once the proof revises it.

* fix(status): a user's Stop reads Interrupted, and a turn anything else cut short reads Failed

The verdict mark now maps a cancellation, the user's own Stop, to interrupted, and an interruption,
a turn cut short by a crash or a killed agent, to failed, the same as a failure, which outranks live
subagent work. An unconfirmed end is unchanged. This applies to every agent, in a terminal or a chat,
on the tab, the sidebar rows and worktree card, the dashboard row, Cmd+J and the phone. A Stop is
not news, so the rollups rank it below an unconfirmed end, and notifications word an interruption
"failed". Recording is unchanged.

* fix(activity): group a user's Stop under Interrupted and a crash with failures

A user's Stop draws the interrupted glyph and sits alone in Interrupted, and a turn anything else cut
short sits in Failed, titled "Agent failed". Every row in a status group now draws the group's own
glyph, so the header is the group's status and the rule that folded a Stop's done check into the
header is gone. Interrupted ranks below an unconfirmed end, as in the sidebar.

* fix(status): draw a user's Stop in the muted tone, not the fault red

The interrupted dot, which now means only a user's Stop, draws in the muted foreground token on the
agent rows, the sidebar card and the phone. Red stays for a failure or a turn cut short by anything
else, and green for a finish.

* fix(native-chat): fold a stopped turn as "Interrupted after N" and a failed or crash-cut one as "Failed after N"

The settled turn header now follows the verdict mark: a user's Stop reads "Interrupted after N", and
a failure or a turn anything else cut short reads "Failed after N", under the new key
components.native-chat.status.failedAfter in all six catalogs and the boot catalog. Desktop and
phone share the one description, so they agree.

* docs(agent-status): describe the Interrupted and Failed marks

The reference and the phone's turn bar still described a user's stop as done and a crash as
interrupted. A fault now reads failed, a user's stop reads interrupted in the muted tone, and the
rollups rank an unconfirmed end above a stop.

* test(status): a crash the relaunch recovers marks failed

The recovery test still expected a recovered interruption to mark interrupted; it now marks failed,
as a failure does. Formatting only elsewhere.

* fix(native-chat): record a turn a newer request replaced as superseded, and show it Interrupted

A Claude turn that a newer send replaced before its result arrived was recorded as interrupted with
no verdict, which reads as a turn cut short by something else, now "Failed". It is now recorded with
its own outcome, `superseded`, where the replacement is detected. That outcome names no sender, so a
dispatch from another agent is never recorded as the user's Stop, and it sets no legacy flag.

Every reader handles it in an exhaustive switch: it draws the muted Interrupted mark with the plain
text "Interrupted", folds as "Interrupted after N", and attention demotes it with a Stop, through
the renamed agentTurnEndedOnRequest. Older builds read an arm they do not know as no verdict, which
is what this turn carried before, so their rows keep reading done; the cross-version suites pin an
older desktop's journal and status readers and an older phone.

* refactor(status): name the attention predicate for a turn ended on purpose

agentTurnEndedOnRequest becomes agentTurnEndedOnPurpose: a user's Stop or a newer request's
replacement, never a fault. The Claude turn-end comment no longer says a replaced turn carries no
verdict.

* test(native-chat): a turn cut off by a restart or by quitting Orca reads Failed after N

On main a restart-cut turn shows the done tick. Pin the chat's turn bar and
the tab's mark for both cuts, through the recovery settlement and the quit's
child-end mapping, and pin the quit's turn bar through the host's own quit.

* style(native-chat): format the superseded turn-bar expectations

* fix(claude): a Stop that names no turn is the user's stop of the open turn

The chat's Stop button names no turn. Claude's conversation Stop recorded the
user's stop only for a named turn, so an older CLI's error result after that
Stop read Failed. It now records it on the open turn through the same intent,
dropped when Claude refuses the interrupt and never carried to the next turn.

* fix(native-chat): keep the attach context's publishStatus required

The lifetime context type makes publishStatus optional, so the attach context
that spreads it no longer satisfied its own type once its duplicate
publishStatus went. The host's lifetime context is now inferred, and checked
with satisfies, so the spread carries the member it always sets.

* fix(activity): rank a user's Stop below live work in the status grouping

The Activity page's status grouping put the Interrupted group (a user's Stop, or a
turn a newer request replaced) above Working and Monitoring, so a Stop still sorted
like news there while the sidebar, worktree card and Cmd+J rank it below live work.
It now follows live work and stays above Done; Failed and Couldn't confirm keep
their places above live work.

* docs(agent-status): say which turn outcomes the journal records and which are derived

The journal now records superseded as well as the provider's verdict and a stop;
interruption and unconfirmed are derived on read. The resume row no longer claims
interrupted renders red.

* test(native-chat): the idle sweep's held-send rest closes with the evict cause, like its siblings

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-30 15:02:44 -07:00
Neil a781a602a8 test: retire duplicate cases that replay an owner across a re-export or provider shim (#24114)
Resolves 208 candidate pairs where the same case title appears verbatim in two or more
files, produced by a repo-wide scan calibrated against a known positive. 46 case
declarations removed across 32 files, 798 lines gone. No file deleted whole, no
production code touched.

The headline result is the measurement, not the deletions: across the three buckets that
reported in detail, the signal ran roughly 86% false-positive (3/42, 9/42, and the rest).
It has good recall and poor precision, and it reorders a reading queue rather than
replacing one. Calibrating a detector against a known positive proves recall, not
precision.

What the deletions were:

- Duplicate invocation through a re-export shim. `native-chat-tool-summary.ts` is a
  ten-line `export {...} from '../../../../shared/native-chat-tool-summary'`, and
  `agent-status.ts:161` is `export { isExplicitAgentStatusFresh } from
  './pane-agent-evidence'`. Cases on the shim side were byte-equivalent to the owner's
  with no rendering or transport hop.
- Provider-local replays of a shared helper: three `repository-ref` providers that are
  each `createRemoteRefProbeCache(parseXRef)` and contribute nothing to transient
  handling; two `local-pty` and `daemon/session` tables replaying
  `shell-startup-output-scanner`, whose owner additionally checks every split point.
- A reader-side replay of store policy. `runtime-worktree-agent-rows-structured.test.ts`
  asserted an attention-to-blocked mapping; the reader contains zero `attention` or
  `blocked` tokens and copies `state` through. The mapping lives in
  `structuredAgentSessionAgentStatus`. Consistent with
  `docs/reference/agent-status-store.md`: readers keep only presentation policy.
- Constructor-only subclass duplication: the shared capability-cache case is covered by
  `codex-app-server-capability-cache.test.ts`, whose ten cases include the identical
  title plus all four risks `docs/reference/git-compatibility.md` names — first fallback,
  later cached call, concurrent probes, per-host isolation.
- A private predicate duplicated at a real boundary, varying only a path passed straight
  into the shared predicate.

Why most pairs were KEPT, because the false positives are principled rather than noise:

- Two independent execution hosts. `src/relay/git-handler-*` and `src/main/git/*` are
  separate Git implementations that cannot import each other and hold separate capability
  caches, exactly as the compatibility doc requires; the repo already ships
  `status-branch-line-total-relay-parity.test.ts` to pin the duality deliberately. Neither
  side's argv, timeout or cache regression is visible to the other.
- Deliberately duplicated production siblings: Codex vs Claude (different account fields,
  different CLIs, different wire protocols), gitea vs bitbucket (`/pulls/42` vs
  `/pullrequests/42`), gl-utils vs gh-utils (separate in-flight maps). Same contract
  shape, different implementations — an identical title is the correct naming.
- Shared-predicate consumers: one side tests the predicate, the other tests a caller's
  wiring to it. A caller that forgot to call the predicate passes the shared test.

In a codebase with intentional provider and host symmetry, identical test titles are
expected, and the signal cannot distinguish "copied" from "parallel by design" because
both produce the same prose. Only reading both bodies separates them.

Verified: 6,968 desktop test files pass; the three modified mobile files pass (39 cases);
`check-reliability-gates.mjs` 140 gates; nothing under
`mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` touched.

62 local failures across 12 files were each accounted for and none is caused by this
change: `browser-manager-tab-identity`, `browser-manager-viewport-ownership`,
`session-scanner-codex-workers` and `managed-hook-script-refresh` all fail identically in
a pristine `origin/main` worktree; five `mobile-web-app-*-render` tests need Playwright
browsers this machine lacks; `structured-agent-session-restart-ownership` and
`ssh-remote-commands` pass in isolation and fail only under concurrent load.
2026-09-30 03:30:23 -07:00
Neil b99462ac1c test: retire mobile, cloud, config and e2e cases their input cannot reach (#24077)
Completes the first pass over every test area in the repository. Sweep over
`mobile/src`, `config/scripts`, `cloud/`, and `tests/` (1,494 files in scope, with
the 24 files under `mobile/src/test-support/rpc-recording/` deliberately excluded).
31 case declarations removed across 17 files, 2 test files deleted, 356 lines gone.

What went, by pattern:

- Cross-boundary replays of a shared helper. A whole mobile file re-ran
  `extractPendingAsk`/`parseAskFromStatus`/`formatAskAnswer`, all owned by
  `src/shared/native-chat-ask.test.ts`, `native-chat-ask-fifo.test.ts` and the
  renderer's interactive-prompt suite — one case title was verbatim identical to the
  owner's, and the owners' inputs are supersets. The mobile file imported the shared
  module directly and exercised no mobile transport, lifecycle or rendering.
- A case whose input cannot reach the behavior its title names: "arms it on Android
  while the drawer is open", where `use-back-claim.ts` has zero
  Platform/OS references, so flipping the mocked OS changes only shadow styles.
- Identity copiers, including one asserting `prSidebarRenderBranch(state) ===
  state.kind` against a production body that is `return state.kind`. The function
  stays; it has three live callers.
- A test of the runtime rather than the product: a case asserting Node's own
  `EventEmitter` crash contract on a bare emitter, with zero production code in the
  path. The guard it documents is exercised behaviourally by the case after it.
- Duplicate invocations, one of them provable rather than eyeballed: with
  `MODULE_SCOPE_ENV_WRITER_PIN = 0`, `files.size <= 0` is strictly implied by the
  sibling's `expect(offenders).toEqual([])`, since a non-empty `offenders` forces
  `files.size >= 1`. The pin's own doc says it may only ever be decreased from 0, so
  it could never become a meaningful bound either. Its policy guidance survives as a
  comment; the file's real ratchet and its regex self-test both stay.
- Expected values produced by the test's own arithmetic, and a p95 case strictly
  implied by a sibling that already pins exact p95 and exact max over a wider range.

One production line goes: the `export` keyword on `assignmentCleanupSteps` in
`cloud/apps/relay/src/assignment-cleanup-steps.ts`. The function itself stays and is
still called internally; only the test-only export was orphaned.

Kept deliberately: everything a gate cites, checked by case title and not only by
file path; a gate-cited case that does not deliver its claim (reported instead — see
below); a cross-version wire cell whose ledger is never invoked, left under the
raised bar for wire coverage; and every limit, bound, quota and provenance guard.

Nothing under `mobile/src/test-support/rpc-recording/` or
`mobile/rpc-foundation/goldens/` was touched — those bytes feed a `recorderSha256`
digest pinning 398 golden recordings.

Verified: `mobile` vitest over the modified mobile files (8 files, 50 cases);
`mobile/scripts/check-tests-typecheck-ratchet.mjs` OK (898 files in program, 125
grandfathered, none @ts-nocheck); relay suite 799 passed; `check-reliability-gates.mjs`
140 gates; both deleted files confirmed absent from the gate manifest,
`cloud/package.json` and `mobile/tests-typecheck-baseline.txt`.

Seven local failures were investigated and none is caused by this change: five
`mobile-web-app-*-render` tests drive `playwright-core` chromium/webkit and need
browsers this machine lacks, `release-checkout.unit.test.ts` needs cross-version git
refs, and `e2e-worker-env-isolation.unit.test.ts` fails identically with its HEAD
content restored — it recurses `tests/e2e` with symlink-following `statSync` and no
depth guard.
2026-09-30 02:05:30 -07:00
Jinwoo Hong 2d31941286 fix(mobile): the update-mobile wall opens the exact release (#23789)
* fix(mobile): the update-mobile wall opens the exact release

When the in-app checker knows the newest release, the update-mobile wall
offers "Get Orca <version>" and opens that release's page instead of the
generic releases list or a hand-built App Store link. Mounting the wall
runs the single-flight, bounded checker once. Dismissal is ignored here.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the update checker out of the page closure

The wall now takes the offered release as a prop and opens it through
the external-link seam. A shell-only hook runs the checker once per
update-mobile wall and supplies the release; MobileWebShellScreen, which
no page route reaches, wires it. HostProtocolGate is in every page
closure, so it stays unwired pending a ruling.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): platform-split the wall's release offer and throttle its check

The hook is the one module whose behaviour differs by platform: the native
file runs the checker, the .web.ts sibling offers nothing, so both
HostProtocolGate and MobileWebShellScreen wire it the same way and the
page closure never carries the checker. A remount checks only when no
release is known and the last check is absent or over the retry interval
old, since each check is an unauthenticated GitHub call.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): let the checker decide whether the wall's check is due

The hook's own throttle read lastCheckedAt, which moves only on success,
so after a failed lookup every wall remount looked up again, and before
preferences loaded a cold-start wall looked up despite a fresh stored
result. The checker already owns the cadence: checkIfDue waits for the
stored state and checks only past the retry or daily interval.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): the wall reads the known release and triggers nothing

The started checker's own timer, cold-start and foreground paths already
run every due check, so a check requested by the wall could never be due.
The hook now only returns the known release; the checker is back to main.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): the wall reads its release through useWallAppUpdate

The .web.ts sibling alone keeps the page closure clean, so the screen
calls the hook itself and the optional prop, the gate wiring and the
shell wiring go. HostProtocolGate and MobileWebShellScreen are back to
their base content.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-29 19:02:40 -04:00
Neil 6e1b7e7fa3 test: remove junk tests that assert source text instead of behavior (#23815)
Deletes 101 test files and trims 112 more, all matching documented junk
patterns: exact source/import/string greps, copied inventories and export
lists, duplicate invocations of a contract another test already owns,
typeof-shape checks TypeScript already enforces, and self-comparisons.

The largest group read a production `.ts` file and asserted on its text —
for example a TaskPage test that required the source to contain
`selectedRepos.find((r) => r.id === newIssueRepoId) ?? selectedRepos[0] ?? null`.
Any behavior-preserving rename broke it; no behavior change ever did.

Production-side follow-through: exports that only these tests imported are
de-exported or deleted, stale comments pointing at removed censuses are
dropped, and the reliability-gate registry, `cloud/package.json` test lists,
and orphaned source-reading helpers are updated so nothing references a
deleted file.

Two files kept their real coverage and lost only the census scaffolding:
`agent-status-producer-census.test.ts` now drives all five producers end to
end instead of grepping the source tree, and `config-toml-trust-stale-writes`
replaces an export-list parity check.
2026-09-29 01:21:53 -07:00
Neil ccdb324b63 Add CodeBuddy as a built-in coding agent (#23740)
* feat(agents): integrate CodeBuddy launch, status and session history

* docs: record CodeBuddy lifecycle verification

* fix(codebuddy): backfill scoped history and negotiate remote resume

* test(cli): include CodeBuddy in known search agents
2026-09-28 18:11:25 -07:00
Brennan Benson 29c7d5d983 fix(mobile): start AI-button agents through agent.launch, never a bare shell (#22762)
* fix(mobile): start AI-button agents through agent.launch, never a bare shell

"Fix checks with AI", "Resolve conflicts with AI", commit-failure recovery and
diff review's "New Agent Session" created a terminal with no agent and typed the
multi-line prompt into the shell, so each line ran as a shell command.

They now call agent.launchReplay into the existing workspace with the prompt;
the host picks chat or terminal from the user's default and delivers the prompt.
Hosts without the launch capabilities get the buttons disabled with update copy.

The agent comes from the desktop's own resolution (moved to src/shared). The
replay loop and capability read are shared with the workspace-create launch.

* test(mobile): repin bridged-parity tallies for the AI-button launch goldens

The corpus goes from 787 to 790 goldens: five shell-path goldens are removed and
eight agent.launch ones added; one lands in identical and two in
result-absent-settlement.

* test(mobile): re-record goldens for AI-button launches through agent.launch

Repinned baseline to 514ab7f868 and re-recorded all goldens. Against the branch
point: 781 header-only (baseline on all; adapterSha256 on the 48 goldens whose
adapter module changed; scenarioSha256 on 3), one body moved
(pr-triage-launch: createTerminal + terminal.send becomes agent.launchReplay),
eight added (the new launch outcomes and their reply matrices) and five deleted
(the shell-path scenarios and their matrices).

* fix(mobile): show review notes' agent launch progress and failures, once

"New Agent Session" left the sheet open with no progress for the whole launch
(up to a minute while a terminal agent readies), so a second tap started a
second agent, and a launch that never started or could not be confirmed
rejected an unobserved promise, showing nothing. The sheet now closes on tap,
the review screen says "Starting an agent...", one launch runs at a time, and
every outcome lands in the review screen's status line.

Marking notes sent now reads the screen state when the launch settles, so a
note written during the wait is not dropped by the whole-list save.

* test(mobile): re-record goldens for review notes' launch outcome on the review screen

Repins the corpus to 6f3018576d. One golden body moves:
review-create-agent-refused now fulfils with "Workspace not found" in the
review screen's status line and the sheet closed, where it previously
rejected an unobserved promise and left the status line empty. The other
789 goldens move only their baseline header.

* test(agent-status): drop the retired PR-triage terminal send from the identity inventory

The phone's AI buttons no longer create a terminal and send the prompt into it
(`createTerminalAndSendPrompt` is gone); the host's agent launch delivers it.
There is no terminal action consumer left in that file to pin.

* fix(runtime): publish saved source-control launch recipes to paired clients

settings.get is an allowlist and omitted sourceControlAi, so the phone never
saw an agent saved globally for "Fix checks", "Resolve conflicts" or commit
recovery and always fell back to the default agent. The host now publishes
the launch actions' recipes (agent, prompt template, agent args), normalized
so legacy saved defaults are already migrated. A new optional reply field:
older clients ignore it, and a client talking to an older host sees none and
keeps using the default agent.

* fix(mobile): ask to update Orca only when the host answered without agent launch

An unread or failed status read settles with no capabilities, which the AI
buttons read as an old host and showed "Update Orca on your computer". The
update copy now needs a status the host actually returned; an unread one keeps
the buttons disabled without blaming the desktop's version.

* fix(mobile): send an AI button's saved agent arguments with its launch

The desktop's direct launches for "Fix checks", "Resolve conflicts" and commit
recovery pass the action's saved agent arguments to agent.launch; the phone
honoured the saved agent but dropped its arguments. It now sends them the same
way: absent when none are saved, so the host keeps the user's configured
defaults. A host that predates the field ignores it.

* test(mobile): repin the RPC recording corpus after the launch recipe and availability fixes

Repins to fe85e0346f. All 790 goldens move only their baseline header: no
scenario saves agent arguments or reads an unreadable status, so no recorded
behaviour changes.

* fix(mobile): say the host status is unreadable instead of nothing when it is

With the update copy now reserved for a host that answered without agent
launch, an unread status left the AI buttons disabled with no explanation.
They now say "Could not read this host's status. Go back and reopen it.", the
words the mobile web shell already uses for the same failure; leaving the host
re-reads its status.

* fix(mobile): wrap an AI button's prompt in the action's saved template

The desktop renders every source-control launch's prompt through the action's
saved template (buildSourceControlRecoveryAgentCommandInput); the phone sent
its built-in prompt as is. Now that the host publishes the recipes, the phone
renders through the same shared function, refuses an empty result as the
desktop does, and offers the rendered text when it could not be delivered.
Review notes have no recipe and are unchanged.

* test(mobile): re-record goldens for the templated AI-button prompt

Repins to 7fd1555d20. One golden body moves: pr-triage-prompt-not-delivered
now carries the prompt as sent (rendered through the action's template) on its
prompt-not-sent result, which is what Copy prompt offers. The other 789
goldens move only their baseline header.

* fix(mobile): re-read a host status that failed while the connection stayed up

A status.get that timed out or was cut over settled the host's gates closed
and was never asked again until the connection state changed, so the phone's
AI buttons stayed disabled behind "Could not read this host's status" on a
link that was working. The gate still settles closed at once, so a failed
read never holds the host screen, but it now re-asks in the background with
the same backoff the runtime capability probe uses, and opens once a status
lands. A reply this app cannot decode is not re-asked.

* test(mobile): repin the RPC recording corpus after the host status re-read

The status gate change moves no recorded behavior: every golden's body is
unchanged and only its baseline header moves to the new pin.

* fix(mobile): show a launch's host warning as a note, not an error

A launch that went ahead can carry a host warning (a structured chat ignores saved agent
arguments, including the '' a template-only save writes). The AI buttons rendered it in the red
error line beside a success haptic. The notice now carries it separately as secondary text, and
review notes keep saying they were sent. Commit recovery also takes the synchronous in-flight lock
the PR triage buttons use, so two taps before a re-render start one agent.

* chore(mobile): record the host status re-read timer for React Doctor

The status re-read arms one retry timer from inside its read and clears it in the effect's
cleanup. React Doctor reports that self-rescheduling shape even in its minimal form, which failed
both changed-lines gates. Suppressed the same way as the session startup timers.

* fix(mobile): say review notes are waiting for the desktop instead of doing nothing

With no live connection, New Agent Session threw from a handler whose promise the sheet drops, so
the tap did nothing visible while the button stayed enabled (proven capabilities survive a drop).
It now closes the sheet and shows "Waiting for desktop..." as the other AI buttons do.

* fix(mobile): stop sending an AI button's saved agent arguments

Whether saved arguments apply depends on the route and shell the host settles after the request
(a chat ignores them and warns; malformed ones fail after admission), and the desktop sends them
only when they apply. The phone cannot know that, so it now leaves them out and the agent's default
arguments apply, as before this series. The saved agent and prompt template still apply.

* fix(mobile): mark review notes sent through the latest save

The sent marks after an agent launch went through the save callback captured at tap time, whose
rollback restores the screen from that moment, so a failed save could drop notes written during
the launch. It now uses the latest render's save, as it already did for the screen state.

* refactor(mobile): own the host status re-read outside the effect

The re-read loop lived inside the effect body, so React Doctor could not see its cleanup and
needed an inline suppression plus a config allowlist entry. The loop is now a plain function that
returns its stop handle, and the effect returns that handle, the same shape every caller of the
runtime capability probe uses. Both suppressions are removed; behaviour is unchanged.

* test(mobile): record the host descriptor from a background status re-read

Pins that the status read records the host descriptor when a re-read succeeds after a failed first
read, not only on the first answer.

* fix(mobile): show a PR AI launch notice only under the button that launched it

Fix checks and Resolve conflicts shared one error, warning and undelivered prompt, so a Fix checks
launch whose prompt was not sent also offered "Copy prompt" under Resolve conflicts, copying the
fix-checks prompt. Notices are now kept per button. The host availability notice stays under each
disabled button, since it explains why that button cannot be tapped.

* fix(mobile): say the host status is being retried instead of asking to reopen it

The host status gate now re-reads a failed status in the background, so "Go back and reopen it"
asked the user for a step that is no longer needed. The review sheet hint uses the same words.
The mobile web shell keeps its own copy.

* refactor(mobile): run the host status gate on the shared status probe

The gate had its own copy of the status probe's retry loop (same delays, same cutover and backoff
split, same stop on an undecodable status). The probe now takes an optional callback for each
failed attempt, which the gate uses to settle closed on the first failure, and the duplicate loop
and its now-unused reader are removed. Existing probe callers are unchanged.

* test(mobile): repin the RPC recording corpus after merging main

Re-records every golden against the merge commit and drops the three goldens whose
scenarios this branch removed, which the merge had restored from main.

* test(mobile): re-record the RPC goldens on the merge with main

Conflicted goldens were seeded from main and re-recorded against the merged
tree; every value either side recorded survives except main's terminal.send in
the PR triage launch, which this branch removes. Drops three goldens main still
had for scenarios this branch deleted.

* feat(mobile): confirm an AI button's agent started, naming the workspace

Fix checks, Resolve conflicts and commit recovery now show "Agent started in
<workspace>" under the button once the host started the agent with its prompt,
so a tap is no longer silent. The workspace label comes from the Source Control
panel and falls back to the branch.

* test(mobile): repin the recording baseline to the success-confirmation commit (header-only)

* fix(mobile): name the workspace in the diff review's AI-button confirmation

The diff review screen mounted the PR sidebar without a workspace label, so
"Agent started in ..." under Fix checks and Resolve conflicts named the branch
while the screen header named the workspace. The sidebar now requires the label
so no screen can drop it, and the diff review passes the one its header shows.

* test(mobile): re-record the RPC goldens on the merge with main

Repins baseline to the merge commit, the last commit to touch a fenced
path. Against this branch before the merge, only header fields move:
baseline on every golden, and adapterSha256 on the 14 review-action goldens
whose adapter main now drives through the review sheet state. No recorded
body changed.

* test(mobile): re-record the RPC goldens on the merge with main

Repins baseline to the merge commit, the last commit to touch a fenced
path. Against this branch before the merge only the baseline header moves,
on every golden; no recorded body changed.

* test(mobile): give the send-sheet stacking test the review controller's host status inputs

The merge with main brought in #22951's stacking test, which builds the review
controller without the host capability and status inputs this branch made
required, so the mobile test typecheck ratchet failed.

* test(mobile): re-record the RPC goldens on the merge with main

Repins baseline to the merge commit 03995ae29d, the last commit to touch a
fenced path. Against this branch before the merge only headers move: baseline
on every golden, and adapterSha256 on the 14 goldens recorded through the
terminal adapter main changed in #23080. No recorded body changed, and the
merged corpus differs from main exactly as this branch did before.
2026-09-28 14:28:09 -07:00
Jinwoo Hong 0034ede120 fix(mobile): left-align every line of the desktop host card (#23673)
* fix(mobile): left-align every line of the desktop host card

StatusDot carried its own marginRight on top of each row's spacing, so the
host card's status text sat 14 px in and the worktree line was hand-indented
to match. Move the dot spacing to the row gap in every consumer and drop the
worktree-line indent.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin tasks style parity hash for the title-row gap

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 15:31:40 -04:00
8b410b4893 feat: add first-class Qoder CLI support (#23581)
feat: add first-class Qoder CLI support

Integrate Qoder launch, identity, canonical hook status, trust and resume.
Verify with captured Qoder 1.1.64 transcripts and hidden Electron sidebar checks.

Builds on and cross-reviews #7502, #8611, #9655, #12910, #13311 and #15291.

Co-authored-by: dalveytech-vincent <vincent@dalveytech.com>
Co-authored-by: Eridanus117 <45489268+Eridanus117@users.noreply.github.com>
Co-authored-by: xingqingzzp-gif <xingqingzzp-gif@users.noreply.github.com>
Co-authored-by: jyang2004 <jyang2004@users.noreply.github.com>
Co-authored-by: yunqian <yunqian@alibaba-inc.com>
Co-authored-by: huzhening.hzn <huzhening.hzn@alibaba-inc.com>
2026-09-28 02:59:50 -07:00
400e4e7957 feat(agents): add Freebuff launch and sidebar status support (#23567)
Add Freebuff launch support and execution-host status reporting for the sidebar, including running, question, blocked, and settled states. Validate against captured CLI transcripts and real rendered sidebar evidence.

Cross-referenced community implementations #17065, #20839, and the Freebuff portion of #18790. Preserve their agent/catalog/mobile/documentation coverage and add canonical status publication and regression tests.

Co-authored-by: Harkaran Brar <18134082+harkaranbrar7@users.noreply.github.com>
Co-authored-by: Prarambha369 <98906077+Prarambha369@users.noreply.github.com>
Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com>
2026-09-28 02:32:41 -07:00
Jinwoo Hong ba858ee446 feat(mobile): the keyboard covers the page like a native screen, and the shell says its height (#23110)
The shell no longer shortens the WebView for the keyboard; it publishes the keyboard height like the safe-area insets, so native's keyboard lift, refit hold and dismiss key run on the page unchanged. Keyboard and inset arithmetic read the shell's OS through a host-os seam. One page-version floor (manifest pageVersion, shell floor 1) replaces per-feature accept negotiation; a page below the floor gets the existing update wall, a desktop with no bundle keeps native screens. iOS shell drops the form accessory bar and its own keyboard observers. Native session screens untouched.
2026-09-28 04:07:35 -04:00
Neil 45f3512a33 feat(agents): add first-class DeepSeek Harness (dsh) support (#22468)
* feat(agents): add first-class DeepSeek Harness (dsh) support

Register DSH as a supervised Orca agent: catalog entry and detection for its
dsh-tui profile, status/question hooks through DeepSeek's own Claude-Code hook
bridge, composer-ready prompt delivery, session resume, headless Source Control
AI, and title identity that no longer collides with Gemini's.

* fix(dsh): reach Orca through DSH's credential scrub and stop reading its title as Gemini

DSH runs command hooks through its own shell executor, which drops every env var whose
name contains KEY, TOKEN, SECRET or PASSWORD — taking ORCA_PANE_KEY and
ORCA_AGENT_LAUNCH_TOKEN with it, so every hook exited without posting. Mirror both onto
scrub-safe aliases at spawn and restore them at the top of the DSH hook script.

Its title collided too: DSH rests on the same glyph Gemini works on, so a resting DSH
pane was relabelled Gemini CLI and reported working forever. Defer both the Gemini
classifier and the title status detector on DSH's whale, in the base module both copies
of that classifier read.

* test(mobile): repin the session-route closure for the DSH agent icon

* fix(dsh): address review — never splice user rows, cover remote panes, keep the diff off argv

- findManagedDshPatchRegion paired an orphan start marker with a later block's end, so a
  truncated write made install/remove delete the user's own rows. Pair each end with the
  nearest preceding start; regression test fails without the fix.
- The relay PTY env builder never applied the scrub-safe aliases, so remote DSH status
  silently never appeared even with the remote hook installed.
- Source Control AI sent the whole diff on argv; send it over stdin with DSH's '-' marker.
- dsh-tui/dst already chose the interactive profile, so a workspace folder named 'web' or
  'plugin' no longer marks a live agent pane non-interactive.
- Isolate USERPROFILE as well as HOME so a Windows run cannot edit the real home.
- Drop the duplicate README badge and revert an incidental doc reformat.

* refactor(dsh): share the managed-hooks reader and tighten the new modules

Reuse before reimplementing: readManagedDshHookEvents was a near-verbatim copy of Muse's,
with byte-identical private helpers. Both now call one readManagedHookEventsFromJson.

Also: one readTextOrAbsent instead of two spellings of the same read (dropping an
existsSync TOCTOU), one status() builder instead of four inline literals, rmSync(force)
instead of exists-then-unlink, and a redundant empty-string guard before JSON.parse.
The patch-file transforms lose their index juggling for a predicate plus a filter.

* fix(dsh): refuse a flow-style patch file, keep its mode, and stop the relay inheriting a pane

- applyManagedDshPatch matched only an exact `[]`, so `[] # keep empty` or a non-empty
  flow sequence got a block entry appended after it — invalid YAML that would leave DSH
  unable to load the user's own patch layer either. It now strips the token from an empty
  sequence (keeping a trailing comment) and returns null for a non-empty one; install
  reports that and changes nothing.
- The patch rewrite dropped an owner-only file to the umask default (CWE-732); pass
  preserveMode.
- The relay PTY env never dropped inherited pane identity the way the local and daemon
  builders do, so a spawn that specified none could inherit the relay's own and every
  agent's hook would report against that pane.

* fix(dsh): keep the flow-style refusal in every status read, and scope the mode test to POSIX

A refused patch file carries no managed region, so getStatus() fell through to a bare
not_installed with detail null — the actionable 'rewrite it as a block sequence' message
only ever reached the one-shot install() return. Export the predicate and check it first,
behind one shared message constant.

The owner-only mode assertion cannot hold on Windows, where chmod only toggles the
read-only attribute and mode & 0o777 reads 0o666 for any writable file.

* docs(readme): restore the DeepSeek Harness badge lost in the rebase

* test(mobile): repin the session-route closure to the measured 4221

Measured, not derived: 4220 without the DSH icon entry, 4221 with it. Two of the three
modules above main's 4218 pin are not this change's — they arrived with the mobile work
after #22570 and were never repinned; the changelog records that split explicitly.

* fix(dsh): settle tui-idle on the agent's own hook, so supervised workers see it ready

Reported by a tester on the adhoc build: `terminal wait --for tui-idle` ran to its 90s
timeout against an already-ready DSH composer, so a supervised worker never sees the agent
as ready.

Every existing tier reads the title, and DSH deliberately carries no title status: its rest
prefix is Gemini's working glyph, so the detector reports none. A fresh first-party `done`
is better evidence than any title anyway — it is the agent's own account of its own turn,
and normalizeDshEvent drops subagent events, so it is the lead's. Scoped to DSH: for agents
whose hooks report child turns, a mid-turn `done` is the #6011 class this file prevents.

* test(daemon): record the DSH transcript's true-colour I2 divergences

Adding the dsh-tui capture to __fixtures__ enrolled it in the serialize replay sweep, where
it reports 10 I2 divergences and failed the unlisted-transcript default of 0.

Every one is the same shape — visible-grid row=0, a 24-bit background the round trip does
not restore to default — which is DSH's whale intro painting whole rows of true colour.
Verified as an upstream limitation rather than a regression by replaying against the
previous build (build-serialize-addon-at-ref.mjs --ref origin/main): I1 and I3 both hold.

* fix(dsh): return the new tui-idle verdict from the first-party done lane

Main refactored isTuiIdleSatisfied into evaluateTuiIdle, which returns a verdict rather
than a boolean. The DSH lane still returned `true`; it is tier-1 positive evidence, so it
returns READY_STRONG like the title/body lane above it. Re-verified the regression test
still fails without the lane.

* test(relay): pin the scrub-safe pane-identity aliases on the relay spawn path

The relay builds a remote pane's env itself, so the alias mirroring there had no
test: removing the call left every suite green while remote DSH status silently
vanished. Both cases fail without it.

* docs(dsh): point the hook service at the integration reference

The reference doc had no inbound link from anywhere in the repo.
2026-09-27 22:44:18 -07:00
Brennan BensonandClaude 85067494a1 fix(native-chat): a request that failed reads as failed (#22944)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 22:23:49 -07:00
Brennan Benson 78771646af fix(mobile): show one review sheet at a time so the review screen never freezes (#22951)
* fix(mobile): close the review sheet before opening Send Notes

On iPhone, Review Actions > Send Unsent Notes and Review Complete > Send Notes
opened the Send Notes sheet while their own sheet was still on screen. iOS
cannot present a second native sheet until the first has unmounted, so Send
Notes never appeared and every later tap on the review screen was swallowed
until the app restarted. Send Unsent Notes now uses the action sheet's
closeBeforePress, and the Review Complete drawer opens Send Notes from its
onAfterClose, the same sequencing the action sheet already uses.

* test(mobile): drive Send Unsent Notes through the real action sheet

The overflow test only checked the closeBeforePress flag, and nothing tested that the
action sheet actually defers such an action until it has closed. Press the real row and
assert Send Notes opens only from the sheet's after-close callback.

* fix(mobile): ignore a close request on a drawer that is already hiding

Android Back during a drawer's close animation restarted the hide, which
cancelled it, so the drawer never unmounted and its invisible Modal kept
swallowing every tap. Send Notes, which now opens after the review sheet
closes, never appeared either.

* refactor(mobile): give the review screen one sheet state so sheets cannot stack

The review screen kept five independent sheet flags that any caller could
set at any time. iOS cannot present a sheet while another is still on
screen (even mid-close), so any overlap froze every tap. Two openers could
still produce one: Review Complete appearing after the mark-reviewed save
landed on a sheet opened meanwhile, and the Send Notes list reopening a
sheet the user had already dismissed.

The five flags become one reducer that mounts at most one sheet. Switching
sheets closes the current one and shows the next only after its drawer
reports it has finished closing; Review Complete waits for the user's
sheet instead of closing it; a late send list only fills a Send Notes that
is still shown or queued. The two per-call-site sequencers this branch
added are removed in favour of it.

* test(mobile): repin the RPC goldens for the review sheet state

The review-actions recording adapter now drives the screen's sheet reducer
instead of the removed Send Notes setter, which moves adapterSha256 on the
14 goldens recorded through it. baseline is repinned to the refactor commit
so the recorder's product fence matches; no recorded body changed (only the
baseline and adapterSha256 header fields move).

* fix(mobile): settle a review sheet that closed before it was ever shown

A drawer mounts only once a commit shows it, so a sheet closed or displaced in
the same batch it opened in (e.g. Review Complete landing in the same frame as
a tap that opens another sheet) never sends onAfterClose. The sheet state then
waited on it forever and every later sheet on the screen stayed queued.

Track which sheet's drawer a commit actually showed and settle a closing sheet
that never reached the screen instead of waiting for a close it cannot send.

* fix(mobile): keep a queued Send Notes when a late Review Complete lands

A second Mark Reviewed save resolving while Review Complete was closing
toward Send Notes reopened Review Complete and dropped the Send Notes the
user had just asked for. A background opener now yields to any queued
user sheet, including behind a closing sheet of its own kind.

* refactor(mobile): present review sheets through one keyed drawer

iOS cannot present a native Modal while another is still presented, even
during its close animation. The review screen now renders all five sheets
through one KeyedBottomDrawer that alone decides what is presented: a
request for a different sheet hides the current one, and the next is
mounted only after its hide finished and a commit without any Modal has
landed. A request replaced before it was shown is never mounted.

The screen's sheet state now records only what the user asked for
(`requested`, plus a background Review Complete in `deferred`), so the
presented-sheet bookkeeping, the per-drawer close callbacks and the
screen-side mounted-sheet inference are gone.

BottomDrawer becomes a constant-key adapter over the same drawer, so the
app has one mount/close lifecycle. A hide that finished just before a
reopen is now ignored instead of latching, which used to swallow the next
close and leave an invisible Modal eating taps.

* test(mobile): repin the RPC goldens for the keyed review drawer

The review action adapter now drives the requested-sheet state, which
moves that family's adapterSha256 on its 14 goldens; the repin to the
keyed-drawer commit moves `baseline` on all 787. No golden body moved.
2026-09-27 21:38:29 -07:00
02a5594434 fix(mobile): give each host field one writer so relay routing can't revert edits (#22956)
* fix(mobile): give each host field one writer so relay routing can't revert edits

Relay learners (director re-resolution, credential rotation, direct-to-relay
upgrade) saved the whole HostProfile snapshot their connection opened with,
so a late relay write reverted an Edit Host endpoint and rewrote the device
token. The relay overlay also stored a copy of the paired address, which the
direct probe kept dialling after an edit.

Pairing is now the only full-profile writer (savePairedHost, fenced by a
census). Learners call setRelayRouting(hostId, relay), which never touches the
row or keychain, refuses a removed host, and skips unchanged routing. The
overlay is relay-only in memory; one serializer writes the v2 shape older
builds parse, without the direct entry. Saving a new endpoint rebuilds even a
Relay-active client from the saved row.

Co-authored-by: mmarabel <166927047+mmarabel@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>

* test(mobile): re-record RPC goldens for relay-only host routing

Baseline moves to the product commit. adapterSha256 moves for the pairing and
relay-credential adapters, which now read relay.relayHostId and inject
saveRelayRouting. Only the four direct-upgrade bodies change: the settled host
no longer carries the overlay's endpoints copy (direct-primary, relay-primary)
or the duplicate relayHostId; relay and every effect are unchanged.

Co-authored-by: mmarabel <166927047+mmarabel@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>

* refactor(mobile): relay learners own only relay routing

Follow-up to the field-ownership split. The supervisor keeps its host
read-only and holds the relay in a small owner that persists moves through
setRelayRouting; the direct upgrade returns { relay, bundle } and the
lifecycle composes the profile once. With one direct endpoint left, the probe
takes a bound openDirect and the lifecycle computes the direct path once.

One name per writer: setRelayRouting and savePairedHost are also the
dependency keys, and the removed-host error is RelayRoutingHostRemovedError.
The census now counts real value imports of savePairedHost, not mentions.

Co-authored-by: mmarabel <166927047+mmarabel@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>

* test(mobile): re-record RPC goldens for the { relay, bundle } upgrade result

Baseline moves to the refactor commit, and adapterSha256 moves for the
pairing and relay-credential adapters (dependency keys renamed to
savePairedHost and setRelayRouting). Only the four direct-upgrade bodies
change: the settled upgrade result is now { relay, bundle } instead of
{ host, bundle }, with identical relay, bundle, state and effects.

Co-authored-by: mmarabel <166927047+mmarabel@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>

* refactor(mobile): refresh edited hosts and let the establisher own relay

Edit Host now reuses refreshHostClient, which already rebuilds an owned
client (relay-active included) from the saved row after re-pairing, instead
of a savedAddressChanged flag on forceReconnect.

The session establisher, the only relay dialer, holds the relay and adopts
learned routing through setRelayRouting; the supervisor takes a host id and a
required relay, so the no-relay guards and the synthetic upgraded profile go.
enqueueHostListMutation returns its operation's value, and the overlay store
names its API after routing.

Co-authored-by: mmarabel <166927047+mmarabel@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>

---------

Co-authored-by: mmarabel <166927047+mmarabel@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>
2026-09-25 23:15:12 -04:00
8846987c99 feat(rate-limits): add Cursor usage tracking (#22633)
* feat(rate-limits): add Cursor usage tracking

## ELI5

If you use Cursor, Orca now shows how much of your monthly Cursor plan you
have used, next to the Claude, Codex and Grok meters, and in Settings →
Accounts. It reads the sign-in Cursor already saved on this computer and never
changes it.

## What changed

Cursor becomes a rate-limit provider like Grok: a status-bar meter (default-on,
with its own toggle), a row in the usage roster, and a Settings → Accounts
section naming the signed-in account.

The credential is read from whichever of three stores has it, first match wins,
all read-only:

- the macOS login keychain item `cursor-access-token` / `cursor-user`, which is
  where `cursor-agent` 2026.06+ keeps the session;
- `~/.cursor/auth.json` and its platform variants, used by older CLIs;
- the Cursor IDE's `state.vscdb` (`cursorAuth/accessToken`), for people who
  never run the CLI.

The keychain entry is the one current CLIs use, and reading only `auth.json`
finds nothing on an up-to-date macOS install. A locked keychain cannot mask a
readable `auth.json`, and a locked `state.vscdb` cannot mask either.
`~/.cursor/cli-config.json` supplies the account's email and display name; it
never holds a token.

Usage comes from the dashboard route the Cursor web dashboard itself reads,
because Cursor documents no individual-user usage API — every documented API is
team- or Enterprise-scoped. Per Cursor's pricing docs an individual plan has two
pools, Cursor Models and Other Models, both resetting with the billing cycle,
plus optional on-demand spend; each becomes a named bucket. The headline
percentage prefers `used / limit` over the sibling percentage fields, which are
pre-rounded for the dashboard's own copy. Because the route is undocumented the
mapping is defensive: an unrecognised payload resolves to `unavailable` and
hides the bar rather than publishing a zero that reads as "no usage".

Orca never runs `cursor-agent login` and never writes, refreshes or rotates a
Cursor credential. An expired token short-circuits to an actionable
"run cursor-agent login" instead of spending a request that can only 401 — not a
rare case, since `cursor-agent status` still reports `isAuthenticated: true`
against a token that expired months ago.

## Why this shape

Six open PRs implement this feature and none reads the keychain, so each finds
nothing for a large share of users; this takes the auth layer further and keeps
what those PRs verified live. The bar is not gated on `cursor-agent` being on
PATH, unlike other CLI providers, because an IDE-only session is real usage with
no CLI to detect.

`readKeychainPassword` moved out of the Claude keychain reader into
`src/main/macos-keychain/generic-password.ts` so both providers share one
`security(1)` wrapper. It is a byte-for-byte relocation, so Claude's credential
path is unchanged; the two child_process allowlists move the entry with it and
neither ratchet count changes.

Co-authored-by: Preschian Febryantara <preschian@users.noreply.github.com>
Co-authored-by: Qwesdy <qwezdi@proton.me>
Co-authored-by: ivo922 <github.concur614@passmail.net>
Co-authored-by: Mihail Vratchanski <mivrkiki@gmail.com>
Co-authored-by: Tauri-EPO <enrico.pin@gmail.com>
Co-authored-by: Raajik <44516546+Raajik@users.noreply.github.com>

* test(rate-limits): name the JWT helper's segment type in the Cursor tests

The anti-slop gate rejects a bare `object` parameter; the fixtures build a
claims record, so say that.

* fix(rate-limits): render Cursor's pools and keep its plan total visible

Review of the first commit found the meter effectively blank for a healthy
account, which the screenshots missed because the only Cursor session on hand
had expired and never reached the success path.

- The verbose status-bar segment filtered buckets through an allowlist written
  for Gemini's experimental models, so both Cursor pools were dropped and the
  fallback needed a `session` window Cursor never reports. A signed-in account
  rendered an icon and no number. The allowlist now admits Cursor's pools, and
  the fallback accepts a monthly window.
- `getWindowSections` dropped `monthly` whenever buckets existed. Cursor puts
  the plan total there and its sub-pools in buckets, so a plan at 92% showed as
  50% in the roster, the tooltip, and the tightest-usage pick.
- A plan reporting `enabled: false` still published its 0% pools, painting a
  healthy meter for a pool the account does not own and skipping the
  request-quota fallback.
- `redirect: 'error'` turned the dashboard's bounce to /login into a generic
  network failure, hiding the actionable sign-in message.
- A busy `state.vscdb` (the IDE holds it open) surfaced as a provider error,
  which would pin an alert bar on Cursor IDE users who never set Cursor up in
  Orca. It falls through to "no credential" instead.
- Refreshing the Accounts section read the keychain twice for one update.

* fix(rate-limits): pin the platform in the Cursor keychain tests

Review caught three cases that assumed macOS: the keychain source is behind an
explicit `process.platform` check, so on the Linux CI runner the mocked read was
never reached and the tests read the CLI file instead. They now set the platform
they mean, and two new cases assert the off-macOS fall-through.

Also track the credentials reference doc (docs/** is ignored by default, so a
new reference needs its own allowlist entry) and give the visibility fixtures
their own provider id instead of Grok's.

* fix(rate-limits): prefer a live Cursor session and report a failed refresh

Review round two, from CodeRabbit and Pullfrog.

- Credential precedence returned the first token that parsed, so an expired
  keychain token in front of a fresh Cursor IDE session reported "sign-in
  expired" on every poll while a usable session sat one source below. A live
  session now wins; the expired one is returned only when nothing live exists,
  so the actionable message still appears in that case.
- The usage schema took `.optional()` where the route sends `null` for an absent
  sub-object, so one null pool failed the parse for the whole body and threw
  away valid pools and the billing cycle with it.
- Cursor usage could survive an account switch: a failed refresh for account B
  kept account A's figures beside B's name in Accounts. The snapshot now carries
  a hashed account fingerprint, and a known-and-changed identity clears the
  previous reading. A refresh that names no account still keeps its own.
- The Accounts section rendered nothing at all when a signed-in account's fetch
  failed, and could repaint an older account when two status reads overlapped.
  It now states the failure — beside the numbers when a stale snapshot remains —
  and ignores superseded reads.
- A web client claimed "not signed in" for a host it cannot read, contradicting
  the meter beside it; it now says the detail is host-only.
- Signed-out copy named `cursor-agent login` as the only way in, though an IDE
  sign-in works just as well.
- The census comment ended at 4219 after the pacer squash without naming the two
  modules #22616 added; recorded them, re-measured on a clean origin/main.
- Narrowed the docs claim: Cursor documents all-plan APIs, but no individual
  usage endpoint.

* fix(i18n): localize the web client's Cursor host-only notice

It reaches the Accounts pane like any other string, so the coverage gate is
right to want it in the catalog rather than allowlisted.

* fix(rate-limits): name the Cursor account on failed refreshes, and ship the reworded copy

Review round three. Both findings say an earlier fix did not actually take.

- The account-switch guard reads `authProvenance` off the fresh result, but the
  fetcher stamped it only on success and network failures. The `stale-token`,
  429, 5xx and parse results omitted it, and so did the expired-session branch —
  so a switch whose first refresh failed, which is precisely the case the guard
  exists for, still rendered the previous account's figures under the new name.
  Every failure holding a readable session now names its account; a missing or
  unreadable credential still names none. The service test also fed a result
  shape the fetcher never produces, so it proved nothing; it now uses the real
  stale-token shape, and the fetcher test asserts provenance across 401/429/5xx
  and expiry.
- The reworded signed-out copy never rendered: a present catalog value beats the
  `translate()` fallback, and `sync:localization-catalog` only adds missing keys
  rather than updating changed defaults. Updated both strings in en.json, which
  also prunes them from the runtime-required catalog now that they match.

* docs: keep the Cursor credentials reference out of the tree

Its content lives in the PR description instead; docs/** stays ignored rather
than gaining an allowlist entry for this branch.

* test(mobile): drop the census note main no longer pins

main removed `SESSION_ROUTE_MODULES` and re-pinned this lane on a different
count, so the paragraph this branch added documents a number series that is
gone. The branch touches nothing in this file now.

---------

Co-authored-by: Preschian Febryantara <preschian@users.noreply.github.com>
Co-authored-by: Qwesdy <qwezdi@proton.me>
Co-authored-by: ivo922 <github.concur614@passmail.net>
Co-authored-by: Mihail Vratchanski <mivrkiki@gmail.com>
Co-authored-by: Tauri-EPO <enrico.pin@gmail.com>
Co-authored-by: Raajik <44516546+Raajik@users.noreply.github.com>
2026-09-25 18:54:11 -07:00
Neil 90801e2deb feat(agents): add first-class ZCode harness (#22464)
* feat(agents): add first-class ZCode harness

Add ZCode (Z.ai's `zcode` CLI) as a supervised Orca agent: managed lifecycle
hooks on local, SSH and Windows hosts; status, question and approval reporting;
synthetic status titles; session resume; orchestration worker launch options;
and desktop + mobile agent-picker registration.

Written against the newly open-sourced `zai-org/ZCode` (agent CLI 0.16.9), not
against a remembered screen:

- ZCode's hook runner writes a Claude-compatible stdin alias set, so it routes
  through the existing Claude-compatible vendor path while keeping its own
  identity in the sidebar.
- `PermissionRequest` fires only once the approval card is on screen and racing
  the user's answer, so it is proof the pane is blocked, not an auto-approval.
- ZCode's clarification tool is literally `AskUserQuestion` with Claude's
  questions/options shape, so Orca's question card renders it unchanged.
- ZCode's `hooks.enabled` defaults to false, which is why configured hooks were
  reported as never firing; the installer sets it.
- ZCode renames its own process to `zcode-cli`, so the expected foreground
  process cannot be the launch command or dispatch refuses the pane.
- ZCode emits no OSC title in any state and repaints its ASCII banner forever,
  so readiness comes from Orca's synthetic hook title and launch drafts wait on
  the composer box rather than on a quiet render window.

Three files crossed their max-lines limit, so each is split along a real seam:
command-line entrypoint parsing out of agent process recognition, skill
classification out of skill root discovery, and registry coverage out of the
remote hook installer tests.

Refs #10564

* fix(zcode): drop the session-option catalog and pin the orchestration contract

ZCode's CLI exposes no `--model` flag at all, and the session-option launch path
refuses to apply any option until a model id is chosen. A catalog therefore could
not deliver `--mode` per worker, and would have accepted `--model` only to drop
it silently. Take opencode's position instead: no catalog, so `worker-start
--model` is refused with a clear message and ZCode launches with the model from
its own config. `--mode` stays reachable through agent args, which is also how
the yolo default is applied.

Add a contract test covering the parts that make ZCode a usable worker:
dispatchable foreground process, stdin prompt delivery, the prompt staying out
of the launch command, and the composer-gated draft paste.

* refactor(zcode): reuse shared helpers and cut the harness down

No behaviour change; every ZCode test still passes.

- Use installer-utils' own `hookDefinitionHasManagedCommand` instead of
  re-walking a hook definition by hand, which also drops a local string reader.
- Share one `readZCodeEventMap` instead of keeping the same narrowing in both
  hook-settings and hook-config-json.
- Collapse five identical error returns into one `zcodeHookError` builder, and
  return early from the status branches instead of assigning through `let`.
- Split the event-to-status decision out of `normalizeZCodeEvent` into a pure
  `readZCodeTurn`, so the normalizer reads as decide-then-build and stops
  computing the tool name for events that never look at it.
- Take a script file name in `readManagedZCodeHookEvents` like its siblings,
  which removes a `Parameters<typeof …>` indirection at the call site.
- Drop the unused `ZCodeHookEvent` export and inline a single-use path helper.
- Correct a stale comment: ZCode's loader is a strict `JSON.parse`, so the
  in-place edit preserves key order and indentation, not comments.

* fix(zcode): address review — keep unmanaged event keys, correct comment, de-dupe README

- `removeZCodeManagedHooks` deleted any event key whose list ended up empty, so an
  unrelated `"Notification": []` the user wrote was removed as collateral whenever a
  managed hook elsewhere made the write happen. Only touch an event Orca actually
  owned something in; covered by a new regression test.
- The `isNewTurnEvent` comment claimed UserPromptSubmit was ZCode's only turn
  boundary while the expression below it also returned true for SessionStart. Say
  what the code does: SessionStart lands the idle boundary, UserPromptSubmit is the
  turn boundary (the Codex/Claude shape).
- ZCode appeared twice in the README's single agent-badge block; keep the
  local-icon entry the link checker validates and drop the favicon duplicate.

* docs(zcode): call out that the desktop bundle's CLI cannot open a session

From live testing on #22464: pointing `zcode` at the desktop app's bundled
`glm/zcode.cjs` installs Orca's hooks fine but then fails with
`Cannot find package '@zcode/tui'`, so the pane never opens a session. The
symptom reads as a broken harness when the CLI simply has no TUI. Say which
build to use and how to check before reporting a problem.

Reported-by: JWu527
2026-09-25 02:17:51 -07:00
Brennan Benson 8f7cbad07b feat(mobile): name the machine after pairing (#22104)
* feat(mobile): confirm host identity after pairing

* refactor(mobile): unify host descriptor state

* chore(i18n): translate the last-known host descriptor label

The remote-host row's "Last known" label shipped in English only. Every other locale now carries it, worded as each catalog already words "last known".

* fix(mobile): make the pairing naming step safe to abandon and show the machine live

Pairing:
- An unreadable status.get reply no longer strands the pairing race: the
  descriptor read ran inside the race's success handler and threw, so the
  candidate was never counted and a direct-only pairing sat on
  "Connecting..." until the timeout. The race now reads the status through
  a reader that returns null instead of throwing.
- A pending pairing is now a small state machine: a save in flight owns the
  outcome (Cancel, back, or unmount no longer clear the journal under it),
  a failed save stays pending and the naming screen shows the error with
  Save still available, and Cancel never rejects.
- When the desktop refuses relay provisioning, the journal is cleared before
  the naming step instead of at save, so an app kill on that screen no longer
  blocks every later scan with "recovery pending".
- A pairing that resolves after the screen went away is cancelled rather
  than left with its journal.
- The screens keep their root ref callbacks stable (the latest pending
  pairing is read from a ref) instead of re-creating them per pairing.
- The label field's placeholder shows the name that an empty label saves.
- "Is this an existing host" is derived from the id identity resolution
  hands back, so host-store and its tests go back to main's shape.

Machine descriptor:
- Mobile keeps the host-reported machine name and OS in memory only, filled
  by the status reads that already happen, and labels it "Last known" from the
  row's own connection state. This drops the per-host AsyncStorage copy, its
  web sibling and web-overrides entry, and the removal cleanup. It also fixes
  a latch: freshness used to stay true for the whole process once a host
  had answered.
- Desktop reads the descriptor through lastVerifiedRuntimeStatus and marks
  it "Last known" using the same reachability verdict as the row's dot.
- Drops the unused hostname/previousPlatform resolver inputs and moves the
  "OS · machine" formatting into the shared resolver.

Also restores main's page-only Reconnect gate in the host header (#22326),
undoes the no-op toStoredHostProfile reformat, and re-measures the web
session route at 4222 modules (one fewer: the dropped web persistence file).

* test(mobile): re-record RPC goldens for the deferred pairing save

Repins baseline to the pairing fix commit and re-records every golden.
Against main, 781 goldens move only in the header (baseline everywhere,
adapterSha256 for the pairing adapter family). Six pairing goldens move in
the body, and only in effect order: the pairing now resolves (closing its
candidate sockets) before the naming step saves the host, and a refused
relay provision clears its journal before the save rather than after. The
set of effects and every outcome match main.

This also removes the unhandled-rejection effects and the failed
result-absent/result-null cells that the previous recording captured from
the pairing race's throwing status read.

* fix(mobile): keep the machine name out of the host header title while the label loads

The host screen starts with an empty saved label and loads it asynchronously. With a descriptor
already in memory, the resolver fell back to the machine name as the title for that window, so
opening "Windows-Low Spec" briefly titled the header with the Mac's name. Read the descriptor
only once the label is known.

* test(mobile): build the pending-pairing status through its schema so the test typechecks

The mobile tests typecheck ratchet rejected a partial status literal: the reply type keeps every
optional field as a required key. Parsing the literal through the status schema yields that shape
without a type assertion.

* test(mobile-web): re-pin the session route closure after merging main

#22301 added two src/shared modules this route reaches; its own CI never ran this suite, so the
merged branch read 4224 against the 4222 pin. Measured on the merge.

* feat(mobile): name each host by what its desktop reports

Pairing saves the host immediately again and names it after the machine name the
desktop publishes; every connection refreshes that name and OS, so a rename on the
desktop reaches the phone. A name typed on the phone's Edit screen is kept as a
phone-local override that wins; clearing it returns to the desktop's name.

- Stored host profiles gain optional personalName, lastKnownMachineName and
  lastKnownHostPlatform; `name` stays the resolved value older builds read.
  Legacy records classify a generated "Host N" as desktop-named and any other
  name as a phone override.
- The connection layer runs one retrying status probe per connected host and
  records the descriptor; the capability probe becomes a projection of it.
- Rows and the header always show the OS, add the machine name under a phone
  override that hides it, and mark it "Last known" while the host is offline.
- Removes the deferred-save naming screen and the one-shot home status fetch.
- Name rules move to host-name-identity.ts and the host-list mutation queue to
  host-list-mutation-queue.ts; an unchanged mutation no longer rewrites storage.

* test(mobile): restore the RPC recordings to main's

Pairing saves the host back to back again and the recording adapter is main's, so
every recording matches main byte for byte; the earlier deferred-save re-record and
its baseline repin no longer apply.

* fix(mobile): keep stored host name identity across snapshot saves and on the web page

A connection re-saves its host profile snapshot on relay credential rotation or relay
re-resolution. The re-pair merge let that snapshot's name identity win, so a phone rename
cleared or changed since connect came back, and a newer desktop machine name rolled back.
The stored record now keeps the name identity on every save; the save supplies the rest.

The web page receives only the app's resolved host name, with no identity fields, so the
docked host header treated a phone rename as "no override" and titled the host with the
live machine name. The display hook now reads such a source the way storage reads a legacy
record: a non-generated name is the user's label.

* fix(mobile): hand the web page the host's stored name identity

The page received only the app's resolved host name, so it had to guess whether that
name was the phone's override or the desktop's name. It guessed "override" for any
non-generated name, which froze a desktop-adopted name as the title after the desktop
was renamed, and left an offline page header without the last-known OS and machine name.

The shell now puts the stored personalName and last-known descriptor on the init host as
optional fields. The page reads them exactly as the app does; a page handed its host by an
older shell still falls back to classifying the name, and an older page ignores the fields.

* fix(mobile): drop an unreadable host name field, not the whole host

The stored host record and the page's init host checked the platform against a closed list
and required non-empty names. A value this build does not know, such as a platform added in a
later build, failed the whole record: the host list dropped the paired host and the next write
persisted the list without it, and the page refused its init message. Those three optional
fields are now salvaged, so an unreadable value drops only that field.

Also states the one exception to the stored name rule (an OS reported without a machine name
keeps the adopted name), and brings four mobile test files in line with the branch: the edit
screen now saves `personalName`, and host opens now start a descriptor status probe.

* refactor(mobile): name the shared host name fields for their role

* fix(mobile): drop the "Last known" prefix from the host machine line

The OS and machine name line under a host's name reads the same whether or not the
host is connected; the connection status already says when it is offline, and a
prefix that users could read as applying to the name added nothing. Removes the
resolver's liveness input and the desktop row's translation key.
2026-09-24 22:35:22 -07:00
Jinjing b419b3183e test: remove redundant mobile and GitLab checks (#22748) 2026-09-24 20:45:15 -07:00
Brennan Benson 7a4f080086 revert: #18790 (orchestration incarnation reap fallback and bundled Freebuff agent) (#22601)
This reverts commit 0677271709.

#18790 was merged as one squash commit that carried two unrelated changes:
a process-incarnation fallback for reaping leaked orchestration worker
terminals, and an unannounced "Freebuff" third-party agent (catalog entry,
icon, locale strings, README rows). The Freebuff agent was never meant to
ship, so the whole PR is reverted; the reap fix should be re-submitted on
its own.

Until that re-land, a worker whose durable terminal handle goes stale is
again reported missing on release/stop instead of being re-found through
its process incarnation, so its terminal can leak on Remote Server.

The mobile session page closure pin moves 4218 -> 4219: the revert drops
the freebuff icon #22119 pinned (-1), and #22452 had already added two
src/shared modules without re-pinning (+2).
2026-09-23 21:39:28 -07:00
Jinwoo Hong 5b6a857e41 fix(mobile): route the bottom drawer's keyboard through the platform seam (OTA phase C follow-up) (#22556)
* fix(mobile): route the bottom drawer's keyboard through the platform seam

Fill-mode sheets called Keyboard.metrics() directly, which react-native-web
does not implement, so opening one on the page threw and the shell
re-downloaded the workspace. The drawer now reads useSoftKeyboard, whose
native half seeds from metrics() and carries the event duration, and whose
web half answers from the window (duration 0). The fill/content-sized seed
rule and resolveBottomDrawerKeyboardInset are unchanged. A census keeps
Keyboard.metrics/addListener inside the seam plus the tab-sheet hide wait.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(config): retire the drawer's exemption from the page keyboard census

The bottom drawer now reads the keyboard seam, so no module in the
source-control or review closures names react-native-web's Keyboard stub.
The census also flags Keyboard.metrics, which the stub lacks.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): give the drawer an imperative keyboard pair from the seam

The seam now exports subscribeSoftKeyboard and currentSoftKeyboardHeight
beside its hooks. The drawer's effect is back to its original shape with
only its Keyboard calls swapped for the pair, and useSoftKeyboard is back
to {height, visible} with no metrics() seed. Seeding every consumer opened
an iOS window between willHide and didHide where metrics() still reads
open. The web pair answers from visualViewport, so it stays silent inside
the shell and lifts sheets in a plain mobile browser.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): start the web keyboard subscription from the current strip

A keyboard already covering the page when subscribeSoftKeyboard attached
never produced onHide when it closed, so the occlusion hook and a seeded
fill sheet stayed lifted. Outside the shell only.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-23 19:46:52 -04:00
NeilandAdrien De oliveira ebed0964a2 feat(agents): add first-class Muse Code harness (#22216)
* feat(agents): add first-class Muse Code harness

Add Muse as a supervised Orca agent across desktop, mobile, session history, source control, local hooks, SSH, WSL, and native Windows. Preserve user settings, support Muse 1.3 hook environment allowlists, and recognize versioned foreground processes. Include question, waiting, completion, resume, and readiness coverage.

Co-authored-by: homesh-dev <300847526+homesh-dev@users.noreply.github.com>

Co-authored-by: jeffhuen <32542276+jeffhuen@users.noreply.github.com>

Co-authored-by: John Cusack <johncusackccm@gmail.com>

Co-authored-by: Adrien De oliveira <75085839+adriendeoliveira@users.noreply.github.com>

* test(agents): cover Muse remote hook registration

* test(agents): cover Muse hook and source-control contracts

* test(agents): exclude Muse hook metadata from script mode check

* test(agents): keep Muse skill picker coverage stable

* test(ai-vault): include Muse in every-agent fixture

* test(mobile): repin Muse agent icon closure

* fix(muse): detect questions and approvals from structured Muse signals

Muse 1.3 fires no hook for request_user_input, so a pending question left
the pane "working". Its internal reminder subagents also post hooks with
their own session ids (even after Stop), which surfaced "tool failed" rows
and flipped finished panes back to working.

- Read pending questions from Muse's session log
  (user_input_prompt_requested/settled) via the existing transcript poll,
  now generalized from Codex subagents to Muse on main and relay.
- Drop child-session hooks (SubagentStart ids, or turn_id === session_id).
- Treat Notification permission_prompt as the approval wait; PermissionRequest
  also fires for auto-approved calls, so it only caches the approval card.
- Ignore Notification copy as the prompt; poll replays are not new prompts
  or turn boundaries.
- Allowlist USERPROFILE so Windows cmd AutoRun doesn't fail every hook.

* perf(muse): parse only question events from the session log

Most Muse session-log lines are large model/tool records. Filter raw lines
by the user_input_prompt_ marker before JSON.parse via an optional
readJsonlCursor line filter.

* fix(muse): unwrap batched log records and scope questions to the live turn

Review follow-ups: question events inside retained_frame batches were
skipped, and a question left open by a crash or interrupt stayed pending
for the pane's life. Share the history scanner's retained_frame unwrapper,
and only report a pending question whose run_id matches the hook turn_id.

* refactor(muse): drop type assertion in retained_frame unwrap

* fix(agent-hooks): satisfy exhaustive-switch lint in transcript poll policy

---------

Co-authored-by: Adrien De oliveira <75085839+adriendeoliveira@users.noreply.github.com>
2026-09-22 19:13:11 -07:00
Jinwoo Hong 11083ac4d3 fix(mobile): the page's auth-failed banner offers Re-pair again (#22363)
#22283 dropped all three banner actions on the page, on the claim that
/pair-scan sits outside the page's route root so Re-pair cannot work there.
It does work: the host screen's router is the route handoff, which posts a
target the page does not serve to the shell (route-handoff.web.ts:207), and
on the emulator the tap opened the native scan screen and Back returned to
the same page document.

The page's sibling now renders Re-pair as native does, wired to the same
onRepair, plus a muted line for the two it still cannot honour: "Reconnect
or remove this host from the Orca app." forceReconnect stays null on the
page and removal keeps refusing; native renders its three actions as
before. The doc comments and the web-overrides reason are corrected, and
the reason's drifted citations re-resolved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-22 20:24:13 -04:00
Jinwoo Hong 0e6862cbcc fix(mobile): the page offers no control whose only effect is a re-dial it cannot make (#22326)
* fix(mobile): a Retry that can only re-dial is not offered where nothing dials

Six failed-load screens share one Retry shape: re-dial a host that is not
connected, otherwise re-read. On the page the re-dial is inert
(`client-context.web.tsx:55`) and each screen's load already re-runs when the
shell's client reconnects, so in the disconnected state that Retry did
nothing at all. `connectionRetryAction` makes the decision once and answers
null when a re-dial is needed and none exists; agent history, the file
explorer root, the file preview, git history, the source-control status gate
and the diff review render no Retry for null.

The explorer's per-folder Retry keeps its control: it queues the folder, and
the queue drains on the next `connected` whoever brought it back.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): the page offers no re-dial, so no header offers one

`forceReconnect` on the page was `() => Promise.resolve()`: the shell owns
the connection and nothing in the document can re-dial it. The host header's
Reconnect and the session header's "tap to retry" were wired to it and did
nothing there. The context member is now nullable and the page's provider
hands out null, so the compiler found every caller: both headers render no
reconnect affordance for null, and the session status keeps the verdict
label without promising a tap.

Native providers and the recording adapters still pass a function, so
nothing a phone renders changes. The session route's host-JSX parity hash
moves for the header's extra null check; the page test doubles that stubbed
the old inert re-dial now stub null.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): the auth-failed banner cites the page's re-dial as null

Three comments and the banner's override reason still said the page's
`forceReconnect` was an inert `() => Promise.resolve()`, and cited
`client-context.web.tsx` lines the previous commit moved. They now say null
and point at the lines that hold it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the session closure gains the page's Retry decision

`connection-retry-action.ts` is the one module the Retry fix adds to the
session route's page closure, reached through the explorer, source control
and git history it docks. Measured on this head with all five generators run
first, and diffed against the pre-change closure: one local module added,
none removed.

Session route closure 4207 -> 4208 modules, local 1021 -> 1022.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): the capability probe belongs on the page, and says why

The push fence excluded `runtime-capability-probe.ts` because the session
route and the host screen run it. The session half holds, and the probe
works there: `status.get` carries no client identity and makes no write, the
shell forwards it like any non-`native.` request, and the desktop's mobile
allowlist admits it. The host-screen half no longer does:
`codex-reset-credit-capability.ts` is reached only from `accounts.tsx`, which
the bundle carries and the page hands to the native screen.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): keep the session retry test's cast under its disable line

The formatter wrapped the cast onto the line after the disable comment,
which left it uncovered. The cast now sits on its own line directly below
the SAFETY note.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the agent-history Retry test mocks the pathname the handoff reads

Main's page route handoff now subscribes to `usePathname` (#22300), and the
Retry suite this branch added mounts that handoff with an `expo-router` mock
that lacked it. Same one-line addition main made to the back-handoff suite.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the reload each hidden page Retry relies on

Hiding a Retry on the page rests on the screen's load re-running when the
shell's client reconnects, because nothing on the page re-dials. Only the
explorer's folder drain pinned that. Each other site now has a case that
starts unreachable with no Retry and asserts the load goes out on the
client and state the reconnect delivers: agent history (status.get), file
preview (the preview read), diff review (the snapshot load), git history
(git.history) and source-control status (git.status, in the loaders suite
because the panel test mocks the state hook).

Each goes red when the `client`/`connState` dependencies it guards are
removed; for source control that is both `loadStatus` and the
`loadBranchCompare` it depends on.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): one import of the transport types in the source-control loaders test

CI's native code-quality audit denies the duplicate-import warning the reload pin added.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-22 19:41:38 -04:00
Jinwoo Hong 11db2b9a7d feat(mobile): the device Back key reaches the page (#22308)
* feat(mobile): the page can claim the device Back key

The shell's page had no way to hear Android Back: every sheet inside it
early-returned on web, so the key popped the whole session route. Adds the
first negotiated shell-to-page frame kind alongside it.

- `back-claim`, page to shell, declared in `init.accepts`: the document is
  holding the key, or has let it go.
- `back`, shell to page, declared in `ready.accepts`: one press, dispatched to
  the newest consumer that takes it. A press nothing takes is handed back as a
  `navigate-back` rather than dropped.

Both are optional fields on frames the other side already reads, so an old
shell never hears a claim and an old page is never sent a press; each pops as
it does today. No protocol bump and no stream opcode.

`bridge-host.ts` was at its line cap, so the notify forwarder moves to
`bridge-host-notify.ts` unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): Android Back closes the sheet on the page, not the screen

Inside the shell's page every sheet early-returned on web, so one press left
the session route with the sheet still open. The drawer, the right drawer and
the file-preview prompt now claim the key through one seam on both platforms:
`use-back-claim.ts` is the hardware key, `use-back-claim.web.ts` is a claim on
the shell's. All sixteen session sheets render through `MountedBottomDrawer`,
so the one claim there covers every one of them, and a census fails if a sheet
bypasses it.

`route-handoff.web.ts` claims while the page grew a stack of its own, and
hands the press back when it did not.

The shell takes the key off the navigator only while a claim is live: Android
gets a `hardwareBackPress` handler that returns the host's own answer, iOS
loses the stack's swipe-back. The claim is cleared on `document-started`, on a
remount, on a new `ready`, on anything that takes the generation off screen,
on the page's `close` and on dispose.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(config): a Back press closes a sheet on the real bundle

The unit suites reach both halves of the lane but never the two together on a
document a browser rendered. The render rig can now post a `back` frame, and
the drawer check opens the Filter sheet, reads the claim off the notify list,
sends one press and pins that the sheet closed with no `navigate-back` behind
it. Red without the drawer's claim: the claim never arrives.

Also fixes a fragility the rich-markdown rig caught. `MountedBottomDrawer` is
shared with the native app and mounts under no page provider in a bare tree,
where `usePageBridgeClient` threw; the seam now reads the bridge through
`usePageBridgeClientIfPresent` and claims nothing without one.

Session route closure 4207 -> 4209: `use-back-claim.web.ts` through the route
handoff, `bridge-page-back.ts` through the envelope.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): mirror the Back seam's latest values from an effect

Three `ref.current = …` writes sat in render, which React replays and
discards. Each moves into a dependency-list-free effect declared ahead of the
registration that reads it, the shape `use-mobile-web-shell-bridge.ts` already
uses for the same reason: the caller rebuilds the value every render, so there
is nothing to depend on, and `useRef` seeds the first mount. The registration
still keys on the claim alone, so a rebuilt handler re-registers nothing.

The web seam's test drops its two type assertions for a named fixture type.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): the page says its Back claim again on every init

The claim was edge-triggered and the shell forgets on purpose: it drops the
claim answering every `ready`, and a host rebuilt under a live page — a client
swap through forceReconnect, which leaves the WebView mounted — starts with
none at all. A document still holding a sheet was then unknown to the shell,
and the next press popped the screen out from under it.

`init` is the shell saying it is here now, so the page answers each one with
the state rather than with a transition. Posted after the session has taken
the frame, so the gate reads that `init`'s own `accepts` and a shell that
never named the claim still hears nothing.

Nothing is said while nothing is held. Every `init` answering a `ready` comes
from a host that dropped the claim first, so it already holds false; the only
other one carries a rewritten route, where a stale true needs a `false` the
page posted to have never left, and a port that refused that frame refuses
this one too.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): a rebuilt host keeps the session's Back claim

A host is rebuilt when the client under it changes, and the page document does
not move: the WebView stays mounted, the session id holds, and the page is
never told. The rebuilt host started with no claim and no `accepts`, so it
refused every press and the navigator popped the screen out from under an open
sheet. Two clients on the same generation leave the page nothing to refuse, so
nothing made it re-ask and re-assert.

What the page declared and what it is holding are facts about the session, the
way `sessionEstablished` already is. `createBridgeHostBack` takes them as a
seed, `readSessionBack()` hands them on, and the hook holds them stamped with
the session so a record left by one never seeds the next.

`dispose()` no longer reports the claim gone: a host retiring is not a
document ending, and that report was the thing taking the key off a live
sheet. Every reset path is unchanged and still has its own case — the page's
`ready`, its `close`, and the session's own store for `document-started`,
`remounted` and anything that takes the generation off screen.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-22 16:21:26 -04:00
Jinwoo Hong a8786c040d fix(mobile): the page never removes a host, and stops bundling push (#22283)
* fix(mobile): the page never removes a host

The page holds one host profile from `init.host` and no credential, so
`removeHost` on web resolved without doing anything and the screen reported
success for a host that was still paired. Its `.web` sibling refuses with a
typed error instead, and the auth-failed banner's Remove — the one surface
that opens the confirm — is absent on the page, because a control that can
only refuse should not be there.

Refusing is also the fence that keeps `push-registration.ts` out of the page
bundle: the native lifecycle file's import was that subsystem's only path
into a page route.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): drop the push families the page closure no longer reaches

`host-removal-lifecycle.web.ts` was `src/notifications`'s only path into a
page route, so the whole directory left the page bundle and the capability
probe left the C1 layout closure with it. The derived family set shrank by
two; the pin tables and their counts now match what the closure reaches.

The expo-notifications fence grows a second claim and loses a precondition
that had become false: `push-token.web.ts` and
`desktop-notification-channel.web.ts` are no longer in the bundle either, so
the fence is stated as the absence of the directory.

C1 20 families / 94 goldens, C2 68 / 257, C3 26 / 116, C5 25 / 125.
Session route closure 4211 -> 4207 modules.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): the page's auth-failed banner offers no control it can honour

The banner is reachable on the page — the shell forwards the native client's
state verbatim (`bridge-host.ts:381`) and `auth-failed` is in the wire enum
(`bridge/bridge-envelope.ts:43`) — and the page can honour none of its three
actions. `forceReconnect` is `() => Promise.resolve()` there
(`client-context.web.tsx:55`, read through `host-client-hooks.ts:87`),
`/pair-scan` sits outside the page's route root of `app/h`
(`mobile-web-app-route-manifest.mjs:6`), and removal refuses. The previous
commit hid only Remove and claimed the other two still worked; they do not.

The whole action row moves into `AuthFailedBannerActions`, whose `.web`
sibling renders no control and one line naming the app. The sentence above it
is unchanged: re-pairing from the desktop is still what to do.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-22 12:49:25 -04:00
Jinwoo Hong 3d76c22b57 fix(mobile): a refused update serves the cached generation while the host is up (#22237)
* fix(mobile): a refused update serves the cached generation while the host is up

A newer generation that fails to fetch or stage was refused with a named
reason and then painted a wall with "Try again" over an intact generation
already on disk — the same one the offline branch opens without being asked
the moment the host goes away.

`onDownloadFailed` now branches on what is cached rather than on which side
refused: with nothing on disk the refusal is still the screen, and with a
generation on disk it is opened through the offline branch, judged by its own
route list. The refused generation is never staged, committed or persisted, and
nothing about the refusal is written, so the next launch asks the host again.

The bundle-side refusal is named as a dismissible notice above the page, on the
existing host-route banner. It says what happened and promises no retry,
because the shell schedules none.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): judge the cached generation against the host it can reach before serving it

The fallback served a cached generation on the strength of the offline rule,
which skips the compat check because a host nobody can reach cannot have
changed. On this path the host has just answered, and an update usually exists
precisely because it moved — so bytes that were inside the protocol window when
they were written may be outside it now.

`CachedGeneration` now carries the three fields the compat verdict reads,
projected in `openCache` off the manifest stored beside the assets. That
manifest is never absent: `readActiveGeneration` answers null for a generation
whose manifest did not parse, and the read schema requires all three.

`cachedGenerationWall` lives beside `gateVerdict`, because only an `open` gate
is judged further. The other verdicts already have answers there: an absent
capability is the native-route rule, and an unreadable status leaves the same
empty list, so walling on either would be the `bundle-unavailable` wall that
file exists to keep off a host that simply did not reply.

A generation outside the window now earns the wall with its verdict, not the
download-failed screen, and nothing is deleted: the bytes are intact and a
newer host is not what makes them wrong.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): ask the cached generation's own routes before walling it

The compat wall ran before the route question, so a cached generation that does
not carry this route — or predates route listing entirely, `routes: undefined` —
earned a terminal `bundle-incompatible` wall where the answer is `native-route`.
`onManifestRead` has always taken the other order: a route that stays native has
nothing to wall about. `openByOwnRoutes` now asks the route first and applies
the wall only on the served branch, and the update notice moved to `served`, so
a native answer carries no notice about a screen it is not showing.

The other half is the verdict the gates hold at the moment of the refusal. A
`fetching` session does not await the gates, so a refusal can land under a
verdict the flow never started on. `gateState` is now the one mapping from a
gate verdict to a screen, shared by the entry into the flow and by the fallback,
so the two cannot answer the same verdict differently: a host that stopped
serving a bundle is `native-route`, a status that went unreadable says so and
re-arms, a dial in progress or a pending status waits in `checking`, an
unreachable host keeps the offline rule and serves the cache unjudged, and
`open` is the only answer that leaves a host to judge the generation against.

That inverts two round-2 assertions that expected the cached page to be served
when the capability list had gone empty. Both were wrong for the same reason:
an empty list is the gate's question, not a verdict about a bundle.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): announce the host route notice banner, as loudly as its tone

The banner is inserted into a screen that is already on screen, so a reader who
has moved past the top of the list never arrives at it. It carried no live
region and no role, so nothing carried it to them.

The urgency follows `tone` rather than being assertive for everything. The
failure tone is an action that did not happen — a refused worktree action, or
the shell's refused update — and interrupts with `alert` and an assertive
region. The notice tone is a bounced route, context for a list already being
read, and waits its turn politely; interrupting for that would train people to
ignore the first. No role on that arm: React Native has no `status` role, so the
polite region is the whole of the answer.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-22 08:39:21 -04:00
Jinwoo Hong 48bdbb8e24 refactor(mobile): the host-scoping rewrite moves out of the terminal's HTML (#22172)
`scopeDocumentStyleToHost` and `scopeStyleToHost` are one rewrite of a flat stylesheet, and two
page mounts read it: the terminal's and the rich Markdown editor's. The module lived under
`terminal/terminal-webview-html/`, so the editor reached across the terminal's directory for it.
It moves to `src/style-scoping/`, named for what it does rather than for its first caller, and
both mounts import it from there. No re-export shim: the old path is gone.

Its test does not follow it whole. Four of its five cases read the terminal's own sheets
(`TERMINAL_DOCUMENT_*`, `XTERM_ENGINE_CSS`) and the first asserts `document-style.ts`'s split
identity, which is not about the rewrite at all -- so that file stays in the terminal directory as
`document-style.test.ts`, beside the module it is about. What moves is the part that names no
caller: which selectors read as the document's own, and the sheet shapes both exports refuse.

The page-closure pin names the module by path and is updated in the same commit. The closure is
otherwise unmoved: 4,363 modules and 1,021 local before and after, one path swapped for another.
The five generated artifacts are byte-identical -- none of them bundles this module.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-22 03:44:35 -04:00
Jinwoo Hong e47ef8cc28 feat(mobile): the shell tells a page which optional capabilities it has (OTA phase D, C8.1) (#22141)
* fix(mobile): publish page-route pairs the strict host schema accepts (OTA phase D, C8.1)

`routeViewOf` handed the manifest's own route entries to the host as
`pageRouteGrants`. The phone reads a manifest route loosely, so an entry
arrives carrying whatever field the desktop that wrote it knew about, and
`BridgePageRouteGrantsSchema` is `.strict()`: one unread key refuses the
pairs, `createBridgeHost` refuses the route with them, and the page gets no
`init` at all rather than losing one field.

Fixed before any route carries an optional grant (ruling 37.4), so the
manifest field the next commits add costs an installed shell nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore: drop the closure and bundle probe scripts from the tree

Scratch measurements for C8.1 (which route closures reach the HTML preview,
and what the preview render rig costs to bundle with a client provider). They
belong outside the repository and were swept in by the previous commit's
`git add -A`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): a manifest route may declare optional grants (OTA phase D, C8.1)

Design B of design-ota-c8-1.md, ruling 37. `MobileWebBundleRouteSchema` grows
`optionalGrants` under the required lane's own grammar, with the 16-name
ceiling applied over the union of the two lists rather than to each. Serving a
route still reads `grants` alone, so a capability a screen cannot work without
stays required and takes the route native; a session's granted list is
`[...grants, ...optionalGrants]` narrowed to what this shell implements, from
one helper that both `grantsForRoute` and the `pageRouteGrants` publish read.

The ruling's compatibility rationale is corrected in place. `z.looseObject`
passes unknown members through rather than dropping them (measured, zod
4.4.3), so a shell older than the field still receives the key; what it lacks
is a policy that reads one. What makes the lane safe against such a shell is
therefore the previous commit's publish fix, not the reader.

BRIDGE_PROTOCOL_VERSION stays 1. No new notify, verb or frame field.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): name the shell's cancelled-navigation behaviour as a grant (OTA phase D, C8.1)

`externalNavigation` joins `MOBILE_WEB_SHELL_GRANTS` beside `screencastBinary`
and `haptics`, declared in `cancelled-navigation-target.ts` because that is
the module holding the rule which acts on it. A third token that is neither a
verb nor a notify: the page posts nothing to make a cancelled top-frame
navigation happen, so this list is the only thing that can tell a page whether
a tap inside the sealed HTML-preview frame escapes at all.

A constant and not a platform read (ruling 37.1): both engines dispatch the
event, `ios/MobileWebShellView.swift:481` and Android's
`MobileWebShellView.kt:382`, so an app build carries the behaviour on both or
on neither.

The policy census grows the half that was only pinned by the verb table: the
implemented set is that table plus exactly three non-verb tokens, each read
off the module that declares it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): the bundle builder carries a route's optional grants (OTA phase D, C8.1)

`resolveMobileWebPageRoutes` maps each declaration member by member, so a
field the declaration grows reaches a phone only once the map names it: until
now `optionalGrants` would have been dropped in silence and every route would
have declared nothing optional. Omitted when the route declares none, because
absent and empty are the same answer to a shell.

The declaration suite grows the rule rather than a row: the map carries the
lane through and writes no key without one, and the lane is held to the
manifest's own grammar and to the ceiling over the union of the two lists.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): the HTML preview hides its links on a shell that cannot open one (OTA phase D, C8.1)

The session route declares `externalNavigation` on the optional lane, and the
preview asks for it before it renders an artifact's links as links. Ruling
37.2's three readings are what "hide" means here, and removing `href` is what
delivers all three at once: `a:any-link` stops matching, so the UA stylesheet
stops underlining, the element leaves the tab order, and there is no dead
anchor a tap does nothing on. The text the author wrote stays where it was,
the artifact paints, and the Preview/Source toggle is untouched.

Done with the browser's own parser rather than over the source text: an `href`
inside a comment or a `<template>` is text to a browser, and a pass that
rewrote either would be editing the artifact instead of its links. The frame
also loses `allow-top-navigation-by-user-activation` on that path, so a link
the pass somehow missed is refused by the browsing context as well.

One route, measured rather than assumed: the design said two, and the file
preview route's closure does not reach the HTML preview at all - it renders
`MobileFilePreviewScreen`. The new closure census derives that list from the
hook's callers.

The render rig grows the case on both engines and the readings it needs, and
`mobile-web-app-preview-frame-readings.mjs` is split out of it at the
readings/arms boundary, because the two were over the 600-line cap together.
Two engine findings are recorded in the rig: an `<a>` with no `href` still
answers `tabIndex` 0 on both, so focusability is asked by focusing; and WebKit
computes `cursor: auto` for a real link, so that reading is pinned where it
discriminates and its blindness pinned where it does not.

The hop-coverage census now reads the effective set, because that is what the
running rule compares. Inert today: the session route is the only declarer and
an opener into every other route.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the preview's hidden link path where the unit suite can reach it (OTA phase D, C8.1)

The mobile suite runs in a `node` environment whose resolver has no `.web`
precedence, so `MobileHtmlPreview.web.tsx`'s import of the grant hook lands on
the native sibling, which answers yes unconditionally. That is why the
existing component suite still measured the granted frame without knowing a
grant exists, and it means the hidden path had no coverage in the sharded
`test` job, where the render rig is skipped for want of the bundler's
dependencies.

So the wiring gets its own file with the module replaced: that the component
asks, and that both the frame's sandbox and the document it is handed follow
the one answer. happy-dom rather than the suite default, because the inerting
pass parses with the browser's own `DOMParser`.

`String(node.type)` rather than a literal comparison: `node.type` is
`ElementType`, which overlaps a real intrinsic tag and not the host strings
these mocks render, so `=== 'Pressable'` is a no-overlap error under
`tsconfig.test.json` and the tests-typecheck ratchet reds on it.

Also replaces a `Reflect.get` the anti-slop gate refuses with an `in` check.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the session page closure at 4,362 for C8.1's three modules

Measured on both sides with `mobileWebAppRouteClosure(SESSION_ROUTE)` at base
`841d06a969` with all five postinstall generators run first, and the two
`local` lists diffed rather than the total inferred: 4,359 -> 4,362 modules,
1,017 -> 1,020 local.

All three are local source modules and none is vendored: the page's read of
`init.grants.native`, the pass that turns an artifact's links back into text
without the grant, and the module declaring the token beside the rule that
acts on it - reached both by that hook and by `page-route-policy.ts`. The
`bridge-caps.ts` it imports was already in this closure, and the hook's native
sibling is replaced rather than joined.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): allowlist the preview's grant sibling among the .web.* overrides

`mobile-web-app-web-overrides.test.mjs` pins the allowlist against the `.web.*`
files on disk, so a new web sibling reds it until the file says why the page
needs one. Red before: `expected [ …(36) ] to deeply equal [ …(37) ]`, naming
`src/components/use-html-preview-link-grant.web.ts`.

The preview's own entry is corrected with it: its reason said
`allow-top-navigation-by-user-activation` is granted, and that token is now
conditional on the shell answering that it can open such a navigation.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the hidden-link render case waits on the frame's own reading (round 1)

CI's chromium arm timed out at the full 240 s on this case alone while the
WebKit sibling passed in 1.5 s and it passed 26/26 locally. The cause is the
third arm: it tapped the granted link and waited through
`expectNavigation: 'main-frame'`, and `waitForRecordedNavigation` has no bound
but the case's own timeout. Under CI load the click missed its 2 s
actionability window, no navigation was ever recorded, and the arm sat in that
wait until vitest gave up - `recorded []`, with the frame attached only at
38.9 s. Three arms sharing one budget is what made this the case to find it.

The arm is dropped rather than its wait lengthened or retried. Every verdict
left is a reading the frame itself publishes: the anchors its document holds,
the style the engine computed for one, whether focus lands on it, and now
whether the tap this arm made landed at all - `actError` is asserted null, so
a click that never reached its target is no longer the same three zeros as a
tap that did nothing.

Nothing is lost. The tap's outcome on a granted shell is the next case, on
these same counters from this same rig and with a budget of its own, which is
the presence precondition this file already uses elsewhere for the same
reason. The WebKit sibling's discriminating reads are untouched.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): the inert-link pass changes nothing an engine renders but the links (round 2)

Round 2's ruling: the hidden-link path may change nothing about the artifact's
rendering except that links are not links. A parse and a reserialise is not
free of that by default, and all four findings reproduced on Chromium 147 and
WebKit 26.4.

A same-document fragment link is kept. It starts no navigation at all, so it
goes on working inside the sealed frame whatever the shell can do, and taking
it away would be degradation over a capability it never needed - an artifact's
own table of contents is the case. Its `target` still goes, because a fragment
aimed at another frame is a navigation rather than a scroll, and `href=""` is
not a fragment: it resolves to the frame's own URL.

Links inside `template.content` are reached, recursively. `<template
shadowrootmode>` is a declarative shadow root the frame's parser attaches and
renders, and `querySelectorAll` does not walk into template content, so those
links arrived live inside a sandbox that refuses their navigation - the dead
anchor ruling 37.2 forbids. Measured: `parseFromString` attaches no such root
on either engine or in happy-dom, so the pass can reach them.

The leading newline of a `pre`, `listing` or `textarea` is written back. A
parser drops one after the start tag and the serialiser is specified to put it
back; measured, neither engine's does, so a round trip lost a blank line from
every such block.

The doctype is carried whole, and the reason is corrected from the one the
finding gave. It cannot move this frame between layout modes: a `srcdoc`
document takes its mode from its embedder, and measured, a quirks doctype, the
bare name and no doctype at all all read `CSS1Compat` inside the frame. What
rewriting it does is change the document the author wrote for no reason, with
`document.doctype` observable beside a Source tab showing the original. The
render case pins `compatMode` as the blind reading it is and reads the frame's
own doctype identifiers as the one that discriminates.

Option B was not available: the frame has no `allow-scripts` and inherits
`script-src 'self'`, so nothing runs inside it and there is no injection to
carry the work.

Also drops a vacuous half of the affordance test. `renderSource()` is called
with no argument, so the markup a Source view shows is the caller's own
closure and asserting it equals the fixture passed whatever the component did.
What the component decides is whether the rewritten frame stays mounted
underneath, and that is what is read now.

`mobile-web-app-preview-arm-driver.mjs` is split out of the render rig at the
boundary the readings module already names - the rig holds what each case
claims, the driver how an arm is driven, the readings what it reports - since
the three were over the 600-line cap together. No max-lines disable or bump.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): a fragment link is a frame navigation in this preview, so it is inerted too (round 3)

pullfrog is right, reproduced on both engines before believing it. Round 2
kept `#`-prefixed hrefs on the theory that they are same-document scrolls. In
this frame they are not: the document's URL is `about:srcdoc` while its base
URL is inherited from the embedder, so `#section` resolves against the shell's
own URL and the destination differs from the document's by more than a
fragment - which makes activating it a frame navigation, and the shipped
`frame-src 'none'` refuses it.

Measured under the shipped policy, one tap, with something to scroll:

  Chromium 147   scrollY 0, frame becomes chrome-error://chromewebdata/,
                 artifact gone, embedder reports frame-src <origin>/preview
  WebKit 26.4    scrollY 0, frame stays about:srcdoc and intact, same report

So the destruction is Chromium-only but the absence of a scroll is not: there
was no working affordance to carve out for, and the carve-out left a live link
that destroys the preview - worse than the inert text it was meant to avoid.
Both sandbox values behave the same, so this is the base URL and the policy
rather than the sandbox.

The same tap does the same thing on the granted path, where this pass does not
run, so an artifact's internal links have never worked in the preview. That is
not this change's to fix; it is recorded in
`followup-html-preview-fragment-links.md`, and the render case reads the
granted arm's violation as its presence precondition so the behaviour is
pinned rather than merely known.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-22 02:11:40 -04:00
Jinwoo Hong 35897da0aa fix(mobile): serialize a list the engine nested inside a paragraph (#22145)
`insertUnorderedList` puts the `<ul>` inside the `<p>` it was given rather than replacing it —
measured on WebKit 26.4 and Chromium 147 both — and `blockMarkdown` read such a paragraph inline.
A bullet list the user typed came back as the paragraph's own text with no marker, so it did not
survive a markdown round trip, on the page and in the native WebView alike.

The serializer now reads structure wherever the list sits: text before it is a paragraph, the list
is a list, text after is a paragraph. The DOM is left as the engine made it and no branch asks
which engine is running. The parse side needs no mirror — it already renders `- x` as a top-level
`<ul>`, which is the shape the fixed serializer reports, and the flat control case pins that.

The unit fixture is built through the paragraph's own `innerHTML`: the HTML parser closes a `<p>`
before a `<ul>`, so a markup string on the editor gives two siblings and would measure the flat
shape. Each case asserts the nesting it got before it reads anything.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-22 01:11:13 -04:00
Jinwoo Hong 841d06a969 feat(mobile): the rich Markdown editor mounts on the page (OTA phase C, C7.10 C2) (#22099)
* feat(mobile): the editor document reads its surface from its host's root

The markup gives the editable surface an id, and inside the WebView that is
unambiguous because the document is the page. On the page it is not: a stack
transition keeps the outgoing session screen mounted while the incoming one
starts, so two hosts carry `#editor` at once and a page-wide `getElementById`
hands both documents the first one. The seventh seam is the root, exactly as it
is the terminal's ninth (ruling 24): the WebView names none of them and gets the
whole page, the page names the element its mount planted the markup in.

Red first, `vitest run src/components/rich-markdown/document-host-root.test.ts`
against the page-wide read: 4 failed, 1 passed — content written into the second
host landed in the first, both documents serialized the first surface, an edit in
the second reported through the first document's host, and stopping the first
took the listeners off the surface the second was still using. The one that
passed is the control: a document with no root still reads the whole page, which
is what the WebView gets.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): hold the editor's document rules under its host element

The editor's sheet says `:root`, `*`, `html` and `body` because inside the
WebView it owns the page. Appended to the head of a React Native Web application
all four restyle every screen the shell can show, so the page mount may inject
only what it owns — ruling 19's rule for `window.onerror`, applied to CSS.

The terminal's half of the scoper drops those rules and repaints through a seam,
because the colour `html, body` was setting belongs to the application. The
editor has no such seam and needs none: its host element *is* that editor's page,
so `scopeDocumentStyleToHost` moves the document's own rules onto the host — the
variables every other rule reads, the surface colour, the font, the box model —
and everything else hangs under it. A selector that merely starts at the document
(`body p`) throws rather than being rewritten into something it did not say.

Red first, `vitest run src/components/rich-markdown/page-stylesheet.test.ts`:
6 failed, 0 passed, all on the absent export.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): the fifteen toolbar commands are one row both surfaces render

The row of controls is not the WebView's: a press becomes an injected
`runCommand` there and a call on the page, and neither difference belongs in the
toolbar. Extracted so the page's editor does not declare fifteen rows of its own
that would drift from the phone's.

`MobileRichMarkdownToolbar.test.tsx` adds the fence a second copy would have
needed: the row names every command in the contract, exactly once. Verified red
by dropping `codeBlock` from the row — "names every command in the contract,
once" failed on the 14-member list before the case went back. The native
component's own test and the web fallbacks file stay green unchanged, which is
what says the extraction moved nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): keep the toolbar test inside the tests-typecheck ratchet

`check-tests-typecheck-ratchet.mjs` reported the new file as newly failing
`tsc -p tsconfig.test.json`: the `ScrollView` mock's spread did not match any
`createElement` overload, and comparing a node's `ElementType` against the string
`'Pressable'` is a no-overlap comparison. Host strings for the mock and
`String(node.type)` for the read, rather than a cast. Ratchet back to OK at 800
files.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): the rich Markdown editor mounts on the page

`react-native-webview` renders nothing in a browser, so C7.6 gave the page a
plain Markdown field and recorded the toolbar and the rendered view as a
degradation. Ruling 26 makes that debt rather than done: the page mounts the
document itself. `rich-markdown-web-document-mount.ts` is the editor's half of
what `terminal-web-document-mount.ts` does for the terminal — the sheet held
under the host's class, the markup planted in the host, one factory call, and a
dispose that gives the host back. `MobileRichMarkdownEditor.web.tsx` is the
component over it, with the same fifteen-command toolbar and the same controller
the phone uses, so `MarkdownReader` cannot tell which sibling it has.

Three seams are the page's rather than the window's. Messages reach
`handleMessage` directly and never `window.ReactNativeWebView`, which on the page
is the shell's bridge. The URL for Link and Image comes from `TextInputModal`:
`window.prompt` was measured to return null in both shells, so those two commands
silently did nothing. And no inset source is supplied, so `onKeyboardInsetChange`
is never called — the screen's `keyboard-occlusion.web.ts` measures the same
viewport with the same formula, and a report here would lift its bar twice.

Red first, two runs. `rich-markdown-web-document-mount.test.ts`: 9 failed on the
absent module, and its listener case is the one that holds ruling 21 — a second
mount reports its own edits and the first mount's detached surface reports
nothing, with an event dispatched on it to say so. The four new cases in
`mobile-webview-editor-web-fallbacks.test.tsx`, run against the plain field still
in the tree: 4 failed, 6 passed — no toolbar, no URL modal, and the `TextInput`
the page is meant to have lost.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): put the editor's surface on the 16 px floor, and grow a census that can see it

`#editor` computed to 14 px, measured in both engines. iOS zooms the page on
focus of any editable under 16 px and does not zoom back, and
`keyboard-occlusion.web.ts` answers 0 for the rest of the session at a scale
other than 1 — the exact failure the floor exists for, on the page's only
full-screen writing surface. The size now comes from the text-input seam, which
is also where the two hosts part: the phone keeps the app's body size because a
WebView has no page to zoom, the page gets the raise, and one binding moves both
if the floor ever does.

The `TextInput` census could not have caught it. `modulesDeclaringTextInput`
matches JSX tags and `style` props, and this is a `contenteditable` in a markup
string sized by a rule in a stylesheet. `mobile-web-app-editable-host-font-size.mjs`
starts from the markup instead: it finds every editable host a closure declares,
follows its id to the rule beside it, and reads the size the same way — a literal
at or above the floor, or the seam's own export imported from the seam's module.
An editable with no id, or one no sibling sheet styles, is reported unresolved
rather than passed.

Red first. The rule's own file reported
`src/components/rich-markdown/document-style.ts:36` as the offender before the
fix (4 failed, 3 passed on the first run, the other three being the brace scanner
and the line-start anchor the fixtures found). The closure case in
`mobile-web-app-session-terminal-closure.test.mjs` now names the editor as the
one editable in the session route's closure and its offender list is empty:
1 passed, 4 skipped under `-t editables`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): the caret survives the host's URL dialog, so Link and Image insert

Measured in the render check, on both engines: the Link and Image commands opened
the modal, took the URL, and inserted nothing. The dialog is what takes the
caret — the modal focuses its own field — and `execCommand` on a document that
does not hold the selection does nothing at all. So the page had swapped one
silent failure for another: `window.prompt` returning null on the phone, and a
command with no selection on the page.

Two halves. The document remembers its caret before it waits and puts it back
after (`restoreRememberedSelection`, unconditional where `restoreSelectionOrEnd`
needs a flag, because the wait itself is the blur); if the host replaced the
content while the dialog was open, the remembered range is gone from the document
and the caret goes to the end instead. And the component answers the promise from
the drawer's `onAfterClose` rather than from the submit, because WebKit would not
take the focus back while the field still held it — with the answer released on
submit, chromium inserted and WebKit did not.

`TextInputModal` forwards `onAfterClose` for that, which is the one thing it did
not already pass through to `BottomDrawer`.

Red first, `editor-selection.test.ts` against the previous `editor-commands.ts`:
2 failed, 7 passed — the caret was left in the dialog's field, and a replaced
document did not fall back to the end. The render check's Link/Image case went
from failing on both engines to inserting on both, with the inserted image's
`naturalWidth` above zero under the shipped policy.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the page's rich Markdown editor, in both engines under the shipped header

`config/scripts/mobile-web-app-rich-markdown-render.test.mjs`: the real component,
mounted by the real React, driven through its toolbar in chromium and webkit under
the policy read out of the shell's own Kotlin source. Sixteen cases, eight per
engine.

What it measures rather than asserts: all fifteen commands change the document,
each with the precondition that what it produces was not there first; the surface
computes to the 16 px floor and the document's own `--editor-surface` variable is
set on the host and nowhere on the root element; `ready` and `change` cross the
seam while `window.ReactNativeWebView` — defined by the rig so its absence is a
reading — is never touched; Link and Image are answered by the modal, and the
inserted image paints with a non-zero `naturalWidth`; one change per checkbox tap
and one per inline code; a link tap reaches the host instead of navigating; a
remount leaves the listener snapshot and the scheduler exactly where one whole
cycle left them (rulings 20 and 21); and two editors on one page hold their own
content and report their own edits.

Four harness facts the first runs found, each now in a comment: the entry needs
four of `MOBILE_WEB_APP_SHIMS` (`isFabric` threw `global is not defined` and every
case failed at `data-ready`); `.web.jsx` in `resolveExtensions` or
`react-native-svg` resolves its Fabric components; a `SafeAreaProvider`, which the
route's navigator supplies and a bare mount does not; and the document's markup,
not its text, as the oracle for a content reset — `### body text here` and `body
text here` read the same, so a text wait passed on the document it was replacing.

One finding, reported not fixed: WebKit's `insertUnorderedList` nests the `<ul>`
inside the `<p>` it was given and the serializer walks back out with the same
text, so a bullet list does not survive a round trip there. The phone's WebView is
the same engine, so this is not something the page introduces; the case names the
command's own element and the reset is numbered per command to work around it.

Run 5 of 5: 16 passed, 0 failed, 0 errors, exit 0, 6.58s.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the session route's closure for the editor on the page

Both sides measured with `mobileWebAppRouteClosure(SESSION_ROUTE)` at base
`9267423f22`, all five postinstall generators run first, the before side a scratch
worktree detached at that sha:

  modules        4333 -> 4360   (+27)
  local modules   991 -> 1018   (+27)

All 27 are local and none is vendored, which is the point: the editor is the app's
own code, not a library. The document's 24 modules under `src/components/rich-markdown/`
were reachable from nothing on the page while it rendered a plain field, and the
other three are the mount, the shared toolbar, and the controller with its
keyboard-inset module. Nothing leaves, because the web sibling replaces its own
native file and that file was never in this closure. Named by diffing the two
`local` lists, not inferred from the total.

`document-style-scoping.ts` is on both sides: the terminal's mount already brings
it, so the editor's second export costs no module.

The generation, measured the same way on both sides: 8,028,418 -> 8,056,166 bytes
(+27,748) across 109 assets against the 9 MiB ceiling, 85.1% -> 85.4%. The script
count does not move (67 against the 76 the chunk fence allows for 15 routes) and
neither does the entry's static closure (1,612,052 bytes against 3 MiB) — this is
code the route already reached for, not a new chunk boundary.

The grant census needs nothing: `openExternalLink` is the editor's only seam with a
grant behind it, and the session route already declares `externalLink` for six
other openers. `mobile-web-app-page-grant-call-sites.test.mjs` passes unchanged.

Closure, webview-consumer and grant censuses together: 28 passed, 0 failed, exit 0.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): drop an oxlint directive the changed-code gate reads as unused

`check-changed-code-quality.mjs` failed with one finding: the mount effect's
`react-hooks/exhaustive-deps` disable reports no problem under that config, so the
directive itself is the finding. The reason it carried is worth keeping and now
reads as a plain comment — the effect mounts once, with `promptForUrl` taken from
the closure, because re-running it would throw away a live document and the caret
in it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(config): the editable-host census counts every editable tag, not every id

CodeRabbit, `mobile-web-app-editable-host-font-size.mjs:50`, and right on the
code: the pattern started from `id="…"`, so it matched only hosts that carry one.
The no-id guard fired for a file with *no* named host at all, which means a file
holding a named host beside an anonymous one reported the named one as clean and
said nothing about the other. An editable is its tag; the id is read out of the
tag afterwards.

Also `:145`, also right: the sibling search was `startsWith(directory + '/')`,
which reaches the subtree, and the walk stops at the first file whose sheet opens
the host's selector. The closure's order is the bundler's rather than
alphabetical, so a sheet one directory down could answer for the sibling the host
actually gets. Now the immediate directory only.

Red first, both cases in the census's own file. The mixed fixture reported one
host where two were planted (1 failed, 7 passed); the nested fixture, with the
nested sheet first in the closure and a compliant 18 px rule in it, hid a 14 px
sibling and reported no offender (1 failed, 8 passed). 9 passed after.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(config): an editable with no declared size is unresolved, not a pass

CodeRabbit, `mobile-web-app-editable-host-font-size.mjs:110`, and right for the
CSS case: `readFontSize` answered `onSeam: true` for a rule that declares no
`font-size`, so the offender check accepted the host without being able to say
what size it gets. The inherited value comes from a rule this walk does not read —
the host element's own, or the page's root — and it can be 14 px.

So "no declaration" becomes "cannot say" and lands in
`unresolvedEditableHostStyles`, which the session closure census holds at empty.
Not an offender: an offender is a size this walk read and found under the floor.

The `TextInput` half of the seam still lets an absent `fontSize` through as
inheritance. That is main's policy and it is about a prop rather than a cascade, so
it is not touched here; the divergence is stated in the reader's own comment.

Red first: the inheritance fixture reported no unresolved host where the size is
unknowable (1 failed, 8 passed), 9 passed after. The real tree is unaffected —
the editor declares its size on the seam — and the closure census still reads an
empty unresolved list.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(config): the editable-host census reads the font-size the cascade uses

CodeRabbit, `mobile-web-app-editable-host-font-size.mjs:119`, and right: the walk
read the first `font-size` in a rule, and CSS takes the last of equal importance.
`font-size: 16px; font-size: 14px;` was therefore reported compliant for a surface
the browser renders at 14 px. `!important` outranks every declaration that is not,
whatever the order.

The flag is also stripped from the value, which the finding did not name but the
fixture caught: without that, a compliant size carrying `!important` was reported
as an offender, because it matched neither the literal nor the substitution shape.

The declarations are split on the separator rather than matched with a value
pattern. A pattern excluding `}` cut `${TEXT_INPUT_FONT_SIZE}px` at the brace of
its own interpolation and reported the real editor as an offender — caught on the
first run of the fix, and the reason the split is the shape here.

Red first: 2 failed, 9 passed — the repeated-declaration fixture reported no
offender, and the important-declaration pair reported the wrong one of the two.
11 passed after.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(config): the render check reads the floor from the seam instead of retyping it

pullfrog, `mobile-web-app-rich-markdown-render.test.mjs:37`, and right: the comment
said the floor was read from the seam and the constant was the literal `16`, which
is the shape the seam exists to prevent. It now comes from
`textInputFontSizeFloor(mobileDir)`, the same reader the closure census uses, which
throws rather than defaulting when the seam is gone.

The assertion becomes "at or above the floor" rather than equal to it. The seam is
`Math.max(bodySize, floor)`, so a theme raising the body size past the floor raises
what the page computes; equality against the floor would have been the same stale
literal one module further away.

Two controls, both run. Raising `TEXT_INPUT_FONT_SIZE_FLOOR` to 18 in the seam
keeps the case green on both engines, because the stylesheet reads the same module
and the page computed 18 — the two moving together is the point. Replacing the
stylesheet's `${TEXT_INPUT_FONT_SIZE}px` with a literal `14px` reds it on both
engines, `expected 14 to be greater than or equal to 16`, which is what says the
assertion carries weight. Both files were restored.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): the editable-host census reads every rule that sizes the host (round 3)

`ruleFor` returned the first exact `#id` rule and the walk stopped there, so a
later exact rule of equal specificity, or a higher-specificity subject rule that
still targets the host, could lower the rendered size unseen.

Every exact rule in the sheet is now collected in source order and read as one
cascade, and any other rule whose subject compound targets the host and declares
`font-size` makes the host unresolved rather than compliant. No specificity
arithmetic, and the sibling walk is unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 22:54:32 -04:00
Jinwoo Hong 226f4a0775 fix(mobile): two more table parsers hold a pipe in a cell (OTA phase C follow-up) (#22114)
* test(mobile): pin escaped pipes in mobile markdown table cells

The mobile preview parser splits a table row on every pipe, so a cell
that escaped one becomes two cells and keeps the backslash.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): read table rows through the shared row splitter

The editor's markdown-table-rows already splits on unescaped pipes only
and unescapes the cell; it has no imports of its own, so owning the rule
once costs the preview parser nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin escaped-pipe rows in PR comment tables

Its splitter strips the trailing pipe before walking escapes and reads
`\\|` as an escaped pipe, so a row ending in `\|` loses the pipe and a
cell holding a backslash swallows the separator after it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): split PR comment table rows on unescaped pipes only

Its own delimiter grammar stays local: a single dash still opens a table
here, which the editor's three-dash separator would reject.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(config): repin the session route closure at 4,332 modules

markdown-table-rows.ts joins through the PR comment renderer. Measured on
this head: 4,332 modules / 990 local, and it is the only file under
rich-markdown/ in the closure, so nothing came with it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 20:26:41 -04:00
0677271709 fix(orchestration): reap leaked worker terminals via process-incarnation fallback — stops an unbounded PTY/process leak on Remote Server (OOM / cgroup PID exhaustion) (#18790)
* fix(orchestration): remint live handle from process incarnation on worker release

When a durable terminal handle goes stale (rendererGraphEpoch fence),
inspectWorkerTerminal re-mints a live handle via
resolveTerminalHandleByProcessIncarnation + matchesProcessIncarnation so
release/stop/read act on the still-running PTY instead of reporting
missing and leaking the agent process tree.

- keep main shared host-scope re-exports; add matchesProcessIncarnation
- wire observation.terminalHandle through control/stop/release
- rebuild release-completion on main structured paths
- on missing/unattached + provably exited: settleDead fence first, then
  same-incarnation settleWorker fall back (archive may block settleDead
  mid-request); settle before recovery defer

* fix(orchestration): derive SSH host scope from the reminted handle; reuse fresh-request recovery guidance for structured workers

Addresses two open CodeRabbit review comments on PR #18790.

inspectWorkerTerminal read the dispatch authority with the stale durable
terminalHandle, so after a remint the lookup resolved nowhere and
currentHostScope was always undefined — an SSH worker with no liveness
verdict and no persisted host_scope got classified from terminal.connected
instead of unverifiable. It now reads the same effectiveHandle every other
observation in the function uses.

stopStructuredWorkerForRelease told the caller to repeat the release with
the same --retry-request, which only replays the stale release_unknown
receipt and made a structured-worker close failure permanently unretryable.
It now sources releaseUnknownRecovery from worker-release-completion so the
fresh-request-ID guidance lives in one place.

Pre-commit lint-staged (oxlint + oxfmt) run manually: clean.

* test(orchestration): exercise incarnation recovery through runtime paths

* test(orchestration): pin the incarnation read scenario to the reminted terminal

The read scenario only asserted that the call resolved, so it documented
nothing about which handle the read reached. Assert that the handle
readTerminal received resolves to the registered pane and incarnation, so
the scenario proves the read went through the reminted terminal instead of
passing on the incarnation fence's throw.

* refactor(orchestration): drop redundant incarnation prefix check; require liveTerminalHandle

* feat: add freebuff as a first-class TUI agent (#42)

<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every
commit. -->

| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 0 | 0 | 0 | 0 |
| Prod | 28 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$​37 | 0 |
$\color{#1a7f37}{\Huge{\mathbf{+}}}$​37 |

<!-- /orca-pr-loc -->

## ELI5

Add Freebuff (`freebuff`) as a recognized first-class TUI coding agent
in Orca alongside Codebuff and other supported agents.

## What Changed

- Registered `freebuff` across shared TUI agent definitions,
configuration catalogs, display names, and telemetry schemas.
- Added agent icons, favicons, status mappings, and mobile asset
references for Freebuff.
- Added localization strings across supported language packs (`en`,
`es`, `fr`, `ja`, `ko`, `zh`) and updated locale translation policy.
- Documented Freebuff CLI in README agent table (`npm i -g freebuff`).

## Why

Freebuff is a CLI coding agent twin of Codebuff (`npm i -g freebuff`).
Adding it to the catalog enables users to launch worktrees, run
automated sessions, and pick Freebuff directly within Orca.

## Linked Issue

N/A

## Visual Proof

`N/A` - Catalog registration and metadata definition for CLI agent
launch; UI rendering uses existing TUI agent picker and status
components.

## Testing

- Verified TypeScript contracts, schemas, and catalog configurations.
- Tested CLI detection / agent picker integration locally on Linux
(`worktree create --agent freebuff`).

## AI Disclosure

Assisted by AI coding tooling.

## Checklist

- [x] This PR is small and focused
- [x] I explained what changed and why (including ELI5)
- [x] Before/after screenshots or videos attached for UI changes, or
`N/A` with reason
- [x] Self-reviewed for correctness, security, and performance
- [x] Cross-platform, SSH/remote, and path/shortcut impact considered
(or N/A)

---------

Co-authored-by: Lesley Murfin <lesley@revivebusiness.ca>

* test(orchestration): erase method overloads in worker reap fixtures

* test: document worker fixture type boundaries

* test: simplify worker fixture typing

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: svc-orca[bot] <313947298+svc-orca[bot]@users.noreply.github.com>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-21 17:23:33 -07:00
Jinwoo Hong 3cfb070294 feat(mobile): register the session page route (OTA phase C, C7.7) (#21977)
* feat(mobile): switch the session route to the shell, still unregistered (OTA phase C, C7.7)

The review switch's shape, for its reasons. The session screen becomes
`MobileSessionRouteScreen` in `src/session` because `useMobileSessionController` is 32 hooks
deep and opens the terminal, chat and tab subscriptions: at the switch's top level it would
open every one of them behind the page as well as in front of it, since hooks cannot be
conditional. As an element passed for `fallback` it is built and not mounted.

Four query params carried rather than re-derived, each omitted when empty: `name` is a label
the screen otherwise derives from the workspace, `created` is the create flow's one-shot flag,
`warning` is the host's own text, and `paneKey` is a notification tap. `paneKey` is the one the
screen writes back — `use-notification-pane-navigation.ts` rewrites it to empty once it has
switched, through `setParams` on the handoff, which inside the page is the document's own
router — so it has to arrive in the page for that to happen at all.

Inert on its own. A switched route renders the shell only once `MOBILE_WEB_PAGE_ROUTES` lists
it; until then the flag is the only thing that changes and it is off.

Three censuses red without their rows, measured on this tree:
- `shell-screen-route-census.test.ts` `walks the route tree and finds them` named
  `session/[worktreeId].tsx` as a ninth switch the list did not have.
- `mobile-web-shell-flag-census.test.ts` `reaches the switched routes through that hook and no
  others` reds without `SESSION_ROUTE` in `SWITCHED_ROUTES`.
- `mobile-web-app-web-overrides.test.mjs` `lists exactly the .web.* files on disk` named the
  new sibling.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): root the session parity family at the screen the route mounts (OTA phase C, C7.7)

The extraction parity pin walks from a root function in `app/h/[hostId]/session/[worktreeId].tsx`,
which is now the flag switch: the walk found no `SessionScreen`, and the runtime-string count went
534 -> 542 on the switch's own param names and path literals.

Rooted at `MobileSessionRouteScreen` instead, which is the function that calls the controller. The
switch's business is which of the two screens renders, not what the session screen does, and its
literals have no place in a hash about the extraction.

Every pinned hash is unchanged, which is what says the body moved and nothing else did: 275 hooks,
77 callbacks, 24 effects, 534 runtime strings, 124 host and 61 leaf JSX facts, 172 style
references, all at the same SHA-256 they had before the move.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): give the page the session screen's stored preferences (OTA phase C, C7.7)

Ruling 7: nothing silently no-ops. The allowlist was one exact key and one prefix, so every
preference the session screen reads inside the page fell back to its default and kept working
outside it — a state the user cannot tell from a preference that does not exist.

The keys are derived from the route's own closure, not copied from design §6. Nine join the list:
`orca:terminal-accessory-layout`, `orca:custom-accessory-keys`, `orca:defaultSessionView`,
`orca:mobileStructuredSendOperations:v1`, the three terminal preferences ruling 7 names
(`orca:terminalTextScale`, `orca:terminalAutocompleteEnabled`, `orca:terminalLinkOpenMode`), and
two the design did not: `orca:hostDockWidth`, which `use-mobile-dock-resize.ts` drags on this
screen, and `orca:hostSidebarWidth`, which `app/h/_layout.tsx` reads above every page route and
which the manifest already names as the reason agent-history declares `storage` at all.

Two are per workspace, not per host. Design §6 has `orca:nativeChatTabs:<worktreeId>`; the module
builds `<prefix><enc(hostId)>:<enc(worktreeId)>`, and `orca:terminalLiveInputDisabled:` has the
same shape. So the narrowing goes one level in from C2.9's: `pageStorageKeysForRoute` and
`isPageStorageKeyForRoute` replace the host-scoped pair, and a session page opened on one workspace
can no more rewrite the tabs of the one beside it than it can another host's pins. Both sides read
the workspace off the route pathname, which is the one fact the shell and the page are each handed.

Every new key's writer notes the mirror before it persists, as `savePinnedIds` does: `init` is
built synchronously, so a write that only reached the store would be one `init` behind.

A refusal is a rejection, not a dropped write. The real AsyncStorage rejects when its store
refuses, and the caller that matters already catches: the durable send journal answers
"Message not sent" rather than putting a mutation on the wire with an operation id no store holds,
which after a crash would send the message twice. `PageStorageRefusedError` names the key and which
of the three refusals it was.

Measured on this tree, which is why the journal needed more than an allowlist entry: one journal
entry with no attachment serializes to 342 characters and 48 unsettled sends put the value past
`PAGE_STORAGE_MAX_VALUE_CHARS` (47 is under it), against a schema that admits 4,096. `init`'s own
`BridgeInitStorageSchema` refines on that bound, so handing the journal over whole refuses the
*frame* and the session screen never opens at all. `pageStorageEntriesForInit` drops such a value
and names it; the page reads a default, which is a degradation rather than a page that does not
start.

Red first, measured here:
- 21 cases across three files on the host-scoped helpers being gone.
- `leaves out a value the page would refuse the whole frame over` reds with the filter bypassed.
- The journal case reds without the rejection, with the operation claimed against a store that
  never took it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): register the session page route (OTA phase C, C7.7)

One entry in `MOBILE_WEB_PAGE_ROUTES` with ten grants, every one read off a call site in this
route's own closure rather than carried from design §1. Measured here:

  navigate                 7 handoff sites
  externalLink             6 openers
  haptics                 24 trigger sites
  native.clipboard.write   6 sites
  native.clipboard.read    3 sites
  native.media.*           2 sites, one seam (`useMediaPicker`)
  screencastBinary         1 site (`MobileBrowserPane.tsx`)
  storage                 10 exact keys and 2 workspace-scoped, the previous commit's

`pageRouteGrants` is derived from this list, so the row is a consequence of the entry and there is
no second table to edit. The design's list was exactly right; the counts are what say so.

The hop census goes 16 -> 23, measured. All seven new rows are `X -> /h/[hostId]/session/
[worktreeId]`, one from each other page route, and none goes the other way: the session's ten
grants are a strict superset of every other route's, so every hop into it is handed to the shell
and every one of its own targets stays in the document. That second half is asserted as grant
coverage rather than as the absence of seven rows — absent is also what an unregistered route
looks like, which is the shape C4 already had to correct once.

Two censuses gained the route and one is new:
- The haptics seam census, whose route-module map moves to
  `mobile-web-app-page-route-modules.mjs` so the new census below shares it rather than keeping a
  second copy that stops growing when the first one does.
- `page-served-back-control-a11y.test.ts`, which named two controls with no `accessibilityRole`:
  `MobileSessionHeader.tsx:64 role=none label=Back to worktrees` and
  `QuickCommandsSheet.tsx:160 role=none label=Back`. Both get the role. Inside the shell there is
  no native chrome behind them, so a bare Pressable is absent from the accessibility tree.
- `mobile-web-app-screencast-lane-grant.test.mjs` derives `screencastBinary` from the closures the
  way the haptics census derives its token. C6 could not write it: the pane is mounted by a route
  rather than registered as one, so there was no route to pin the grant against (C6 ruling 3).

The derivation census gains C6's half measured against this route rather than against a module
closure read on its own, which is the other half of C6 ruling 3. The composed row for the session
route's own families waits on C7.8's table, and on C4.5's split before it.

Numbers, both ends measured on this tree, never summed:
- Session route closure 4,328 -> 4,329 modules, 978 -> 979 local. The +1 is
  `MobileSessionRouteScreen.tsx`; the route file is one input either way, now the `.web.tsx`.
- Chunk count 65 before and 65 after, against the 72 the fence allows at 14 route keys. The fence
  is untouched: a `.web.tsx` sibling is not a new route key, and this route shared its split.
- Bundle 8,020,519 -> 8,022,202 bytes, 108 assets either side.

`mobile-web-app-route-chunk-closure.mjs` looked the route module up by its exact path, and
`resolveExtensions` puts `.web.tsx` first: the first route with a sibling to be asked for reached
"no output". It tries the sibling first now, which is what the build actually chunked.

Without the manifest entry these red on this tree: `pins every hop the handoff must take away from
the page`, `keeps every hop out of the session local`, `declares only routes the bundle has a
module for`, `reaches the built manifest`, `covers every page route and finds a control in each`,
both haptics-seam cases, and `declares the screencast lane on exactly the routes whose closure
asks for it`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): render-check the session route, and quiet the two things it found (OTA phase C, C7.7)

The render check mounts the registered route in a real browser on the built bundle, under the
header the shells send. It asserts the session screen paints rather than the Unmatched route, that
the Back control reaches the accessibility tree as a real `<button>` with its name, that the
route's own chunk arrives on a client-side navigation, that nothing it paints leaves the origin or
logs a policy violation, and that the three reads the screen makes carry the workspace the route
named — the precondition the rest needs, since a screen that mounted and asked for nothing would
paint the same chrome.

It also asserts, strictly, that the page and console errors are `[]`, which is what found both
fixes here. Measured on this tree before them: two console lines and one uncaught rejection on
every mount of the route, none of them visible natively.

- `use-mobile-session-markdown-actions.ts` registered `BackHandler.addEventListener` with no
  platform guard, and the effect re-registers whenever the dirty-draft list changes. React Native
  Web answers "BackHandler is not supported on web and should not be used." and hands back an inert
  subscription, so the guard was never armed on the page anyway. Gated on `Platform.OS`, as the
  right drawer, the bottom drawer and the file preview already are. There is no hardware back in a
  WebView; the shell owns the phone's, and the page's Back control is where the prompt lives.
- `use-mobile-session-diff-comments.ts` ran `void loadDiffComments()` in an effect with no catch.
  The loader returns on a *refused* `worktree.show` and nothing caught a *rejected* one, so a host
  that will not answer produced `Uncaught (in promise)` on every session mount. Caught at the
  effect rather than inside the loader, whose promise the golden recorder awaits; notes that did
  not arrive leave the ones on screen as they were, which is the module's own policy for a refusal.

**The terminal is not painted here and the file says so at both ends.** A terminal on screen needs
the host protocol handshake, a tab snapshot, a terminal inventory and a `terminal.subscribe`
stream — five hand-written fixtures against five Zod schemas inside a transport double, which is
what the harness's docstring refuses to become. Scripting `status.get` alone was measured here:
the protocol gate reads it and the page paints "Update Orca on your computer" instead of the
screen. What the terminal does under the shipped header is
`mobile-web-app-terminal-render.test.mjs`, on the same component and the same build options.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): refresh the session parity pins for the two seam edits (OTA phase C, C7.7)

The previous commit's two fixes are inside the parity family, so three pins moved. Both edits are
one token each and neither changes what a phone renders:

- `'web'`, the `Platform.OS` guard the Markdown actions' `BackHandler` registration gained.
- `"button"`, the accessibility role the session header's Back control gained.

Runtime strings 534 -> 536, with the effect hash and the host-JSX hash moving for the same two.
Everything else is unchanged: 275 hooks, 77 callbacks, 24 effects, 61 leaf JSX facts, 172 style
references, all at the SHA-256 they had before. A separate commit because a reported head does not
move by amend, and because the moved hashes are worth reading on their own.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): report the diff-notes rejection instead of catching it (OTA phase C, C7.7)

The `.catch` the previous commit added to `loadDiffComments` moved a golden, which is a finding
rather than something to record over: `matrix-session.diff-notes-worktree.show-1` certifies the
unhandled rejection as an effect of its loaded checkpoint, so the corpus says the app raises it
today and a fix is a re-record and a review event.

Reverted to `void loadDiffComments()`, with the defect written where a reader of that effect will
find it. `family-recordings.test.ts > session.diff-notes: reply partitions at worktree.show#1` is
green again; it was the one failure in an otherwise clean 8,699-test run.

The render check keeps the observation rather than losing it. Its error assertion is now the exact
list `['RenderCheckShellDouble: the render check answers no RPC']` instead of `[]`, so a second
error reds it and so does this one going away — which makes the file the place the fix is noticed
when someone lands it with the re-record.

The defect, for that PR: the loader returns on a *refused* `worktree.show` and nothing catches a
*rejected* one, so a host that will not answer raises an unhandled rejection on every session
mount. It is not a page fault — the shell's `fault` notify comes from the React boundary and
nothing reaches it — so the generation is not dropped and the screen works; the cost is a
document-level error on every mount.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the session effect hash after the diff-notes revert (OTA phase C, C7.7)

The effect pin was refreshed while `loadDiffComments` carried a `.catch`; reverting that (the fix
moves a golden, so it is a finding rather than a line) moves the same hash back off it. Repinned on
the uncaught `void` call, which is what the tree holds and what the corpus certifies.

Count unchanged at 24 effects; nothing else in the family moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): reject a page storage write for size only, and log the rest (OTA phase C, C7.7 round 1)

Ruling 33.4. `PageStorageRefusedError` was raised for all three refusals, and two of them have no
catcher: a page-closure writer of an unlisted key awaits `setItem` with nothing around it —
`notification-delivery-preferences.ts:39` plainly, `preferences.ts` in several places — so a key
the page was never allowed to keep became an unhandled rejection in the document. That is a worse
failure than the silent drop it replaced, and it is the one the page can least afford, because an
uncaught rejection there is a document-level error on a screen that is otherwise working.

Scope is now one refusal. `too-large` rejects, because the caller that needs it is written for it:
the durable send journal's composer catches it and answers "Message not sent" rather than sending a
mutation whose operation id was never written down (ruling 7). `not-allowed` and `not-delivered`
resolve and are logged as `[page-bridge] storage-write-dropped`, which is the old behaviour plus
the line a device log needs — a preference that did not stick looks identical to one nobody set.

A batch applies every pair it can, logs every drop, and rejects only if one of them was oversize.

Red first, measured here: seven cases in `page-async-storage.test.ts` red on the rejection, among
them a `notificationDeliveryPreferences` write resolving, another host's pins, another workspace's
chat tabs, and a write the shell would not take. The oversize case is unchanged and still asserts
`PageStorageRefusedError` with the key and the character bound in its message, so the narrowing is
visible as the difference between the two.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): make every count in the session route say the same number (OTA phase C, C7.7 round 1)

Ruling 33.5. Three numbers were stated more than once and two of them had drifted when the merge
took the grant list from ten to fourteen.

- Grants. `mobile-web-page-routes.mjs:100` and `mobile-web-page-route-hop-coverage.test.mjs:57`
  both still said ten. Fourteen in both, and the manifest comment now names the audio verbs beside
  the media ones as things only this route asks for.
- Keys. The manifest said `storage` covers "the ten exact keys and two workspace-scoped ones",
  which counts `orca:last-visited-worktree` — a key this route did not add. Nine exact plus the
  two workspace-scoped, which is what C7.7 put in `page-storage-keys.ts`.
- The journal entry. 342 and 343 are both real and answer different questions, which is exactly
  why one number had to win: an entry serializes to 342 characters on its own and costs 343 in the
  array, the difference being the comma that joins it. 343 is the one that drives the threshold,
  so it is the one stated, with the 342 kept beside it as its derivation. Re-measured here rather
  than carried: 47 entries are 16,140 characters and 48 are 16,483, against the 16,384 cap.

Comments only; no behaviour and no assertion moved. The threshold case in
`mobile-structured-send-page-storage-refusal.test.ts` already asserted the boundary both ways and
still passes unchanged, which is what says the arithmetic above is the code's and not the prose's.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): red-first for a pane request over a re-sent init (OTA phase C, C7.7 round 1)

Ruling 33.1's four cases plus the compatibility one, all red: `publishRoute`
is not a member of the host, `onRouteUpdate` is not a member of the page's
client, and `ready` carries no `accepts`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): deliver a pane request to the mounted page over a re-sent init (OTA phase C, C7.7 round 1)

Ruling 33.1. The session switch keyed on the whole route, so a notification
tap for another pane of the session on screen either remounted the shell (a
bridge teardown and a page reload for a tab switch) or, for the pane already
showing, moved nothing at all: the page cleared `paneKey` on its own router
and the native param kept it, so `SET_PARAMS` wrote the value already there.

`paneKey` leaves the key and travels as a route update. The page declares
`accepts: ['route-update']` on `ready`; the shell re-sends `init` for a
same-path param change only to a page that declared it, and treats a second
`init` for the session the page already holds as a route update rather than a
replacement -- in-flight requests, subscriptions, the storage snapshot (the
same object, asserted) and the generation all stay. The screen reports
delivery and the switch clears the native param, so no later `init` replays a
spent tap. `use-notification-pane-navigation.web.ts` reads the request off a
standing listener; the native file is unchanged.

Wire-compatible both ways without a version bump: `accepts` is optional, an
older page is never sent a second `init`, and an older shell never sends one.
Both degrade to today's lost repeat tap. `BRIDGE_PROTOCOL_VERSION` and every
released native RPC are untouched.

Two files were at their line cap, so two modules came out at their own
boundaries rather than a cap bump: `bridge-init-route.ts` (the route half of
`init`, wanted by the switches, the host and the page) and `bridge-host-route.ts`
(one host's held route and what it may publish).

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the session switch and its hardware-back gate get their own tests (OTA phase C, C7.7 round 1)

Ruling 33.2. `mobile-web-shell-session-route.test.tsx` mirrors the eight cases
the files switch has -- route built, native fallback while the flag settles,
repeated params, dot-segment refusal, segment encoding, flag off, remount on a
route change, remount on a param change -- plus the two pane cases: a repeat
tap for the same pane reaches the mounted page twice and a different pane
reaches it once, both with one mount in the lifecycle.

The `BackHandler` gate gets a unit test in the shape of its three siblings.
Reaching it meant the hook declaring the fourteen fields it reads instead of
taking all 268 of the session model, so a probe can render it without building
a session; `MobileSessionDiffCommentsModel` satisfies that by construction and
the one caller is unchanged. No pin in the session parity census moves.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(config): a call-site census for the six grants that had none (OTA phase C, C7.7 round 1)

Ruling 33.3. `navigate`, `storage`, `externalLink`, the two clipboard verbs
and the media three were pinned only by the list they were copied from, so
striking any of them out of a manifest entry reddened nothing. Each is now
derived from the route's own closure by parsing the call sites -- a call, not
a mention in a comment or a string, and not an import the module never calls
-- and each row has a named control case driven over the session entry with
that row's grants struck out.

It found one thing. `app/h/_layout.tsx` wraps every `/h` route in
`HostProtocolGate`, whose wall offers an Update Orca link through
`openExternalLink`, and two routes reach that without declaring
`externalLink`: on them the link posts a notify the shell refuses. Recorded
exactly as `KNOWN_UNDECLARED` rather than exempted, because widening two other
routes' grants is a capability decision and this is pre-existing on main.

`notificationPaneTab` moves to its own module. A `.web.ts` sibling cannot
import its native neighbour by the plain path: the bundler's
`resolveExtensions` answers with the `.web.ts` file, so that import was the
file itself and esbuild refused the page bundle with a cycle. The mobile suite
does not bundle, so only the closure walk saw it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(config): make a struck-out grant red a case named after it (OTA phase C, C7.7 round 1)

The first shape checked the whole manifest at once, so removing any one of
the eight reddened all seven cases and named none of them: the per-row control
read `session.grants` off the manifest the removal had just changed. Each row
now has its own manifest case, and each control is built from what the
session route's closure reaches rather than from what its entry declares, so
it stays green whatever the manifest says. Closures are memoised, which is
what pays for walking all eight once per row.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold the pane request in a ref, not in state (OTA phase C, C7.7 round 1)

The changed-code quality gate's React Doctor found it:
`no-adjust-state-on-prop-change`. A tap can arrive before the terminals have
loaded, so the request has to wait; holding it in state meant the effect that
consumed it set state on a prop change, and the stale selection renders first.
The request waits in a ref now and a counter wakes the effect, so the effect
reads and clears rather than adjusting anything.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(config): re-measure the session closure and list the new web sibling (OTA phase C, C7.7 round 1)

The full `config/scripts` suite found both. The closure reads 4,326 modules
and 984 local, two more than the merge, and the two are named rather than
counted: `notification-pane-tab.ts` and `bridge-init-route.ts`. The pane
hook's web sibling replaces the native file rather than joining it, so it
costs nothing -- but it is a `.web.ts`, so it needs its row in
`web-overrides.json` saying why the native one cannot run on the page.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the page off a journal init could not carry (OTA phase C, C7.7 round 1)

Ruling 33.6, from pullfrog on `beda1cc384`. Dropping an over-cap value from
`init` did not revoke the page's write access to it: the key stays in
`pageStorageKeysForRoute`, so the page read no journal, `parseJournal(null)`
gave it an empty one, and its first send wrote a one-entry value over the
device's -- every native entry lost and a fresh `operationId` for an operation
the native journal already held, which is the duplicate send ruling 7 exists
to prevent.

`pageStorageEntriesForInit` now reports `oversize` beside `dropped`: only the
value-cap drops, because an entry-cap drop is a key that fits and the page's
own write of it is the size the shell would have carried anyway. The shell
sends those names as `init.storageOversize`, and a page write to one of them
rejects with `PageStorageRefusedError` under the size contract of 33.4, which
the composer already shows as "Message not sent". The native journal is
untouched until the user is back on native or it drains.

`storageOversize` is optional in both directions: an older shell sends none
and an older page ignores it, which is exactly today's behaviour. No version
bump; omitted rather than sent empty, so no golden moves.

Red first, with the two states replaced by ones the shell produces. The 47/48
case drives `pageStorageEntriesForInit` rather than publishing a journal value
the shell strips before `publishPageStorage` ever sees it, and the case that
used to assert a successful write now asserts the native entries survive: it
was the clobber, recorded as success.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): deliver a route update only when a param moved, and only after init went out (OTA phase C, C7.7 round 2)

Round 2, findings 1 and 2, both red first.

`onRouteUpdate` fired on every re-sent `init` for the session the page holds,
not only on one whose route moved. The shell answers every `ready` with the
route it holds and the page re-asks on its own backoff and again after a
refused `state` frame, so one tap reached the pane hook as `['', 'pane-1']`.
Both ends now read one definition of moved, `bridgeRouteMoved`, which is the
page's own `shellScreenRouteKey`: the host will not send an `init` for a route
that did not move and the page will not publish one it was sent anyway. The
`.web.ts` hook keeps its empty-pane guard and its comment now says why it is
load-bearing rather than defensive -- the shell's own clear arrives as a move.

`onRouteDelivered` ran on the `ready` path without checking that an `init` had
gone out. A refused route answers the ask with nothing, so the caller would
clear a one-shot param the page never received. `sendInit` reports whether a
frame left and `onPageReady` carries it. Unreachable from the session switch,
which parses the route before it mounts the shell; the prop's contract says it
anyway, and the publish path already honoured it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): attach the init-storage doc to its type, keep the overrides escape (OTA phase C, C7.7 round 2)

Round 2, findings 4 and 5, neither a behaviour change.

The block describing `pageStorageEntriesForInit` had `PageStorageForInit` and
its own one-line doc between it and the function, so it documented neither.
The type moves above it and the block sits on the function it describes.

`web-overrides.json` had an escaped em dash re-encoded as a literal one when
this branch added its rows through a JSON round trip, on a line about the
keyboard stub that has nothing to do with C7.7. Main's `—` is restored;
`oxfmt --check` accepts the file either way, so this is main's spelling kept
rather than a formatter's demand.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): catch the custom-key save the page store refuses (OTA phase C, C7.7 round 2)

Round 2 addendum. `addKey` awaited `saveCustomKeys` with no catch and both of
its callers are `void addKey(...)`, so the rejection had nowhere to go.
`orca:custom-accessory-keys` is in the session route's page allowlist and a
page write over `PAGE_STORAGE_MAX_VALUE_CHARS` rejects rather than drops (the
size contract of 33.4, extended by 33.6 to a key `init` could not carry), so
past ~16 KB of accessory keys this surfaced as an unhandled rejection in the
page -- which the fault boundary reports and which drops the generation.
Every other allowlisted writer in this closure already catches: the two write
chains in `TerminalShortcutSettings`, the live-input save and the session-view
preference.

Caught at the boundary and logged, and the drawer neither announces the key
nor closes: a row on the accessory bar that no store holds, gone at the next
load, is the failure the allowlist exists to avoid. Red first -- the case saw
the refusal escape with the page's own message -- and the control reds again
when the catch rethrows.

Belongs in `8b4c559e90` by the brief; it is its own commit because that one
was already made and amending is forbidden.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): reject a batch of oversize writes once, not once per pair (OTA phase C, C7.7 round 2)

CodeRabbit and pullfrog, same site. `settleBatch` called `settle` per refusal
and kept the first rejected promise, so a `multiSet` or `multiRemove` with two
over-cap pairs built a second rejected promise nobody held -- an unhandled
rejection in the page, the outcome ruling 33.4's rejection scope exists to
avoid. Two oversize keys is all it takes, and `storageOversize` made a second
way to reach it.

A refusal is now an error or nothing, and only the caller's one rejection ever
becomes a promise. Red first under an `unhandledRejection` listener with two
over-cap pairs: one orphan before, none after, and the caller still hears
about the first key.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): record a route key only once a frame carried it (OTA phase C, C7.7 round 2)

CodeRabbit on `MobileWebShellScreen.tsx:271`. The effect recorded the route's
key and then published, so a publish the hook refused for having no host was
remembered as though it had gone out. `publishRoute` is now keyed on
everything the host is built from rather than on the session alone, so the
render that brings the host re-runs the effect, and the key is written only
after a frame has left.

Reported honestly: this does not repair a lost tap, and the case beside it
says so. The host is built from the route the render holds, so a route that
moved before it existed rides the first `init` either way and `publishRoute`
then answers "did not move". What the change removes is a key recorded for a
frame nobody sent -- the same contract finding 2 fixed on the `ready` path.
The case pins the delivery count across the gap: nothing reported while there
is no host, nothing reported once there is one and it has sent nothing, and
exactly one report when the `init` answering the page's ask carries the route.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): refuse an oversize-dropped key at the shell, not only at the page (OTA phase C, C7.7 round 2)

pullfrog's rollout gap on ruling 33.6. `storageOversize` is honoured by a page
built with it, and the page is served from the desktop: a document from an
older bundle ignores the field and writes the key whole, which for the send
journal replaces every entry the device holds. The shell is the half that
updates with the app, so the shell is where the refusal has to live.

The host now refuses a `storage` notify for a key it could not hand the page,
answering it as the drop it already answers an unlisted key with. The page's
own rejection stays as the fast path -- it reaches the composer as
"Message not sent" with no round trip -- and the schema comment says the field
is advisory and the shell enforces it.

Red first: a host holding the journal as oversize received a page write for it
and posted it to native storage; now it posts nothing and the entries survive,
while a key it did hand over is still writable.

The three refusals became one predicate in `page-storage-keys.ts`, where the
keys are, because inlining the third put `bridge-host.ts` over its line cap.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the page's pane hook gets its own test (OTA phase C, C7.7 round 2)

pullfrog: `use-notification-pane-navigation.web.ts` had no cover. The native
file's test mounts the native file, and `bridge-route-update.test.ts` stops at
the client, so the half that turns a route update into a tab switch was
untested.

Seven cases: the seed from the route the page was opened on, a request held
until the terminals load, a repeat tap on the pane already showing, a
different pane, the clear the shell posts after each delivery, a pane that has
since closed, and a page opened on no pane at all.

Two controls, so the cases are not all satisfied by one behaviour. Dropping
the seed reds the two that read the first `init`. Deduplicating by value
instead of counting deliveries reds the repeat tap, which is the case the
counter exists for.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): roll the custom-keys mirror back when the store refuses (OTA phase C, C7.7 round 2)

CodeRabbit. `saveCustomKeys` notes the write in the mirror before it persists,
because a reader is answered from the map rather than from the store and the
shell builds `init` synchronously from that map. On a refused write the note
stood: the page's next `init` carried the value native had rejected, and every
native reader of the key saw it too.

The previous mirrored value is captured and put back on the failure path, and
the error still goes to the caller so `addKey` keeps withholding the key.

Red first: with the store refusing, the mirror held the rejected value where
the pre-save value belonged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): report a route delivered only after its frame was posted (OTA phase C, C7.7 round 2)

CodeRabbit on the send path. `sendInit` answered "sent" the moment it handed
the JSON to `post`, and a rejected post was reported a turn later as a
diagnostic -- so the screen spent the one-shot `paneKey` on a frame the page
never received, cleared the native param, and the tap was gone. `publishRoute`
was fire-and-forget the same way.

Delivery is a promise now, settled after `options.post` resolves and false on
either throw or reject. Readiness stays separate: `onPageReady` fires on the
ask, as the shell's wait needs, and carries the delivery promise beside it.
The screen records the route key and calls `onRouteDelivered` only when that
promise answers true, and a refusal leaves nothing recorded so the next render
that can carry the route tries again.

Red first: with the view refusing what it was handed, the frame was built and
posted and the screen reported delivery anyway. Now it reports none while the
page's ask is still reported, and a host-level case pins the same split.

`bridge-host.ts` was at 299 of 300 lines, so the send half came out as
`bridge-host-frames.ts` rather than growing it; the file now measures 280.
The screen's own test harness never attached the view handle, so every post in
it rejected unobserved -- it attaches one now, which is what let the case see
the frame at all.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): fold two doc blocks back onto what they describe (OTA phase C, C7.7 round 2)

pullfrog's two nits, no behaviour change. `page-async-storage.ts` kept the old
`settle` block above `refusalError` when the function it described moved down
with a one-liner of its own; the orphan goes. `bridge-host.ts` had two stacked
blocks on `sendInit` after it grew a return value; they are one, and it now
says the frame is still built synchronously and only the post is awaited --
which is the property the golden recorder depends on.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): let the host own a pending route until its frame lands

A route the page has not received is now the host's, not the screen's. `publish` keeps it
pending until a post resolves true, marks it delivered only then, and reports that through a
callback registered once per host. Movement is measured against what a frame actually reached
the page with rather than against what the host holds, so a refused frame leaves the route owed
instead of reading as one that did not move.

Three things the old shape lost, each a case here: a frame the view refused was never retried,
because only another render could try and a mounted page has none coming; a render while a post
was in flight cancelled the report the switch spends to clear the param; and a repeat tap for
the same pane was held, because the host had already moved its held route on the attempt that
failed. The retries are the moments delivery becomes possible again — the next `ready`, and a
view handle the host regains — and one frame goes out at a time.

Also folds round 4's doc nits: the stale delivery comment the screen no longer has a ref for,
and a leftover `an` in the `ready` branch.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): roll the journal mirror back when persistence fails

`writeEntries` noted the mirror before the store took it, which is what keeps an `init` built in
the same turn current — but it kept the note when the store refused. The page then received a
journal the device never wrote and resumed operations nothing was holding.

Restored on the error path, the same shape as the custom-keys save, and on both halves: the
removal that empties the journal had the same gap as the write that fills it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep publishing until the held route is the delivered one

Two halves of the same gap, both found by the bots on `fba8cb3f3d`.

A route that moved while a frame was out was held for the turn and then had nothing to wake it:
the post settling only cleared the in-flight flag, and on a mounted page no `ready`, handle or
tap need ever come along. A landing is now itself a moment to publish again, while what the host
holds is not what the page has. Only on a landing — a refused post that re-attempted itself
would spin, and that one still waits for whatever makes delivery possible again.

And the report carries the route a frame reached the page with, which the session switch was
ignoring: the older pane landing wiped the `paneKey` naming the newer one, so the page stayed
where it was and the second tap was gone. The switch now spends the param only for the pane that
was delivered.

The delivery cases render through one helper rather than six copies of the same setup, which is
what keeps the file under its cap.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): drop the second block describing a parameter onPageReady no longer takes

The field is documented by the block above it; this one still described the `delivered` promise
the handler was handed before the host took ownership of the pending route.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): let the page erase the route param it was handed

Ruling 34, step one: one page-to-shell frame that asks the shell to clear a one-shot route param,
naming the value the page applied.

Closed at both ends. The param is an enum of what the shell hands over, so a page cannot edit a
route it was never given; the notify name is a member of the closed union, so it gets a row in
the grant table by compilation rather than by memory, and rides no grant because it can only
spend something this shell put there. The shell declares it in `init`, the mirror of
`ready.accepts`: no shipped shell serves a page, so nothing needs negotiating today and the
page's check exists from the first version that can post one.

The comparison belongs to whoever holds the param, which is the session switch: a tap that moved
on while the page was applying the one before it leaves a newer key, and a clear naming the older
one is not for it.

`bridge-envelope.ts` went over its cap, so the page-to-shell union moved to
`bridge-notify-envelope.ts` and the fields both halves spell to `bridge-frame-fields.ts`, which
the envelope re-exports. No cap was raised.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): re-send init on a route change and track nothing else

Ruling 34, step two: the tracked handoff is gone. Deleted, not patched — the pending route, the
delivered route, the in-flight flag, the landing callback, `retryPendingRoute`, the delivery
promise `onPageReady` used to carry, and the `onRouteDelivered` that ran from the host through
the hook and the screen to the switch.

What is left is the rule in one line: `publish` sends one `init` when the route moved and the
page said it takes one, and every `ready` is answered with the route the shell holds then. A
frame the view refused is repaired by the next ask, not by a retry; the request it carried is
spent by the page.

The cases that tested the deleted mechanism go with it. The outcomes they protected are pinned
where they now live: one frame per move and none for a render that moved nothing, a lost frame
repaired by the next ask, no second `init` to a page that never said it takes one, and the
repeat tap measured through the page's erase rather than through a delivery report.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): the page applies a pane and erases the request that carried it

Ruling 34, step three. The page hook applies the pane an `init` names and asks the shell to erase
the param it came on, naming what it applied.

Two rules, and both are the page's because the shell has none. The erase is asked for on every
`init` that carries a pane rather than only on the one that changed something: a clear that never
reached the shell leaves the param in place, and the next frame carrying it is the repair. The
switch happens once per value: a re-asked `ready` is answered with the route the shell still
holds, and applying that again would drag the page off a tab the user has since moved to.

A repeat tap for the same pane still arrives as a request, because the erase went through in
between and the tap wrote the param back. The client refuses to post the frame to a shell that
did not declare it takes one, which is every shell older than the field.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the page bundle's readers pointed at the module that declares each name

The envelope split left three source-text readers and one import pointed at a file that now
re-exports what they read.

`shell-screen-route.ts` read the route schema back through the envelope, which reaches that file
again through the page-to-shell union: a cycle esbuild resolves to `undefined`, so every page
route mounted onto a schema that was not there yet and the browser render suite failed on twelve
files with a TypeError rather than on a build error. It reads the declaring module now.

The render harness read `BRIDGE_PROTOCOL_VERSION` and `BRIDGE_FAULT_GRANT` out of the envelope by
regex; both moved, and a regex over a re-export answers for whichever file the last split left
them in. Both point at `bridge-frame-fields.ts`, and the throw names it.

The session route's page closure is re-measured on this tree at 4,330 / 988 and the four new
modules are named, not inferred: the two halves of the split envelope, and the route-update
module and route-key reader the page-to-shell union now reaches through it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): count posted inits with the page's own reader, not a cast

The changed-code quality gate refuses a type assertion, and it is right to here: a frame the
page's reader would refuse is not an `init` the page ever saw, so a case counting them must not
count one either.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): put every mirrored write on one path that notes what the store took

Ruling 35. Fourteen call sites in six files noted the shell's mirror before persisting, and on
the page a persist can be refused: twelve left the map holding a value no store had taken, and
the next `init` handed the page exactly that. Two undid it by hand.

`persistMirrored` is the one path now, and it seats the map from what the store holds after the
write rather than from what it was handed. That is what makes the note follow acceptance without
a second opinion about it: the page's adapter resolves a `not-allowed` write and logs it, so a
rejection is not the only refusal there is, and reading back is the only answer that covers both.
The cost is one store read per mirrored write on the device, where the store refuses nothing.

`writeMirroredStorage` keeps its note-then-persist order and loses every caller but one: the
shell taking a value the page has already applied, into the device store, which has no allowlist
and no frame cap to refuse against. It builds the next `init` synchronously in the same turn, so
noting on the store's reply there would hand the page back the value it just changed. The
last-visited key moved off it, because that module is in the page's own closure.

Both rollbacks are gone with the notes that needed them, and `noteMirroredWrite` is private. A
source-scanning census holds each writer to the path by name.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): answer a page batch write at the first pair it cannot take

Ruling 35's other half. `settleBatch` collected a refusal per pair, logged each, and rejected
with the first that could reject while the rest of the batch went in anyway — one promise
describing a call where some pairs landed and some did not, which is not something a caller can
act on.

A batch is one call with one answer now: every pair before the refusal is applied, the refusal is
the answer, and nothing after it is attempted. No page-closure writer calls `multiSet` or
`multiRemove` today, so this is the rule for whoever writes the first one rather than a change to
anyone's behaviour; both directions are pinned, including the refused first pair that stops the
rest.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(mobile): format the mirrored write path's census and journal writer

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): count the note-first callers the mirror module says are held to one

Two gaps pullfrog found in the census. The last-visited key moved onto `persistMirrored` with no
row naming it, so removing its write path reddened nothing; and `mirrored-storage-keys.ts` says
the census holds `writeMirroredStorage` to one caller while nothing counted them.

Counted now, over every module under `mobile/src` rather than over a list of files a new caller
could sit outside of: a second one is either a writer that wants note-then-persist without the
store that earns it, or a page-reachable module that would note a refusal as an accepted write.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): fold sendInit's doc onto the function it now describes

It still described an awaited post that answered whether the page received the frame, which
ruling 34 deleted: it fires the frame and answers nothing, a refused route sends nothing at all,
and a post the view would not take is one diagnostic and no further attempt.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): let the page own a frame it received and could not handle

Ruling 34's addendum. On iOS the host's post is `callAsyncJavaScript`, which rejects when the
page's synchronous `onmessage` throws — with the document still mounted. The shell reads that as
a frame that never arrived, and it tracks nothing about posts, so nothing would ever send it
again. It is not a lost frame either: the page had it, one of its own listeners failed, and a
retry would fail the same way.

`receive` catches it and reports `inbound-listener-threw`, so the delivery is the channel's and
the handling is the page's. Nothing is swallowed and nothing is retried.

Two cases pinned the throw escaping and now pin it being reported: the ack that a listener bug
must not wedge, and the bootstrap stamp a tree that throws still leaves behind.

With that path closed, a post is refused only when no document holds the view, and the comments
on both halves of the route seam say so instead of naming a backoff that is stopped by then. The
repair is pinned rather than described: a tap that arrives while the view is gone is carried to
the next document's `ready`, because the held route advances on `hold` as well as on `send`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): move the page's held session out of the client, which was at its cap

The listener catch put `bridge-rpc-client.ts` at 305 counted lines against a cap of 300, so this
splits rather than bumps.

The session is the one piece of the client with a lifecycle rather than a value: a second `init`
for the same session updates it in place, a different one replaces it and takes the requests and
streams of the session before it, and each case has its own listeners to fire in its own order.
The client keeps the frames and the ports; `bridge-client-shell-session.ts` keeps what they are
for, and the client's three members delegate to it.

The client measures 277 counted lines after the move. The session route's page closure is
unchanged at 4,330 / 988: the page reaches its client from the entry, not from the route module.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 16:07:14 -04:00
Jinwoo Hong 83128ea7ae fix(mobile): four pre-existing rich Markdown editor defects (OTA phase C, C1 follow-up) (#22054)
* fix(mobile): a checkbox tap in the rich editor reports one change

A tap raises click, input and change, all three bubble to `#editor`, and each
handler emitted: three identical changes under one generation. The tick and the
change now come from `change` alone.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): inline code over a selection reports one change

`wrapSelection` emitted and `runCommand` emits after every command, so the one
command that wraps rather than execs reported twice. The wrap helper now emits
nothing; `inlineCode` is its only caller.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): a rich-editor code block can hold a backtick fence

A fixed three-backtick fence ends at the first three-backtick run inside it, so
a block containing a fence rendered as paragraphs. The writer now measures the
longest run and opens one longer; the reader carries the run it opened with and
closes only on one at least as long.

The bundle's input count moves by the one module the two halves now share.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): a rich-editor table cell can hold a pipe

The reader split on every pipe and the writer joined without escaping, so a cell
containing a pipe became two columns and the backslash that hid it survived as
text. The reader now splits on unescaped pipes and undoes the escapes; the
writer escapes backslashes before pipes, which is the order that round-trips a
cell holding both.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the checkbox input's caret-flag clear

The checkbox fix skipped the whole input handler, which also stopped clearing
`selectionDroppedOnBlur`. That flag is not part of the duplicate-change defect,
so it is cleared as it always was and only the emit moves to `change`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 14:25:57 -04:00
Jinwoo Hong a3c6d4266a fix(mobile): admit https: images on the web shell's CSP (OTA phase C, ruling 27) (#21964)
* fix(mobile): admit https: images on the web shell's CSP (OTA phase C, ruling 27)

Native markdown and the native rich editor load images the author referenced
by URL, so the page has to as well or a remote image is a blank where native
paints a picture. `img-src` widens to `img-src 'self' data: https:` on both
platforms; `script-src`, `connect-src`, `object-src`, `frame-src` and
`child-src` do not move.

`http:` stays out, and the pins say so directly rather than by absence: the
Kotlin test's blanket `!contains("http")` could not survive `https:`, so both
native pins now check `http:` (not a substring of `https:`) and check that
`https:` appears in `img-src` and nowhere else, the same shape the `data:`
pin already had.

No behaviour change on released phones: the shell ships in no released tag
(mobile-v0.0.9 predates it), so this reaches devices with the Phase E native
build and not before.

Neither native module has a CI job, so both ran locally: swiftc over the
module plus MobileWebShellChecks, and
`:orca-mobile-web-shell:testDebugUnitTest`. Both were confirmed red against
the old directive first.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): correct what the sealed preview frame is stricter about

The doc comment said the page was deliberately stricter than the native
preview because it loads no remote image and runs no script. Since `img-src`
gained `https:` only the script half is true: the frame loads a remote image
exactly as the native WebView does.

Says instead what an artifact's image URL now is -- a channel that fires on
view and carries whatever its author encoded, with nothing dynamic behind it
because no script runs -- and names `referrerPolicy` as what keeps the
document's own origin out of the request.

Comment only; no behaviour and no test moves.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): measure both halves of the preview frame's image fence

"fetches nothing of the artifact that leaves the origin" stopped being what
the sealed arm proves once `img-src` gained `https:`. The fixture's foreign
origin is `http://127.0.0.1`, so its two images are refused on the scheme
alone and only the font is refused by `font-src 'none'`. Renamed to say
exactly that.

The half that was missing is an https arm. Playwright route interception
answers an `https://…invalid` origin in the page, so the arm needs no TLS
server and no new dependency, and a request only reaches the handler if the
policy let it out. Under the shipped header, on Chromium and WebKit, the
`<img>` and the CSS background are both requested -- `img-src` governs a
background too -- and the font still is not.

`artifact()` takes the subresource origin; the links stay on the cleartext
one so no existing navigation case changes.

Red-first: with `img-src 'self' data:` put back into the parsed Kotlin
policy, the new arm fails on both engines with `expected [] to deeply equal
[ '/css-bg.png', '/img.png' ]`. The directive was restored byte-identical
before this commit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(scripts): split the preview frame's settling out of the render check

The https arm pushed mobile-web-app-html-preview-render.test.mjs to 620
counted lines, over the 600 cap config/scripts carries. Split at a module
boundary rather than bumped: the four wait-and-settle functions are rig
mechanics with no assertion in them, and they now sit beside the diagnosis
module they already reported through.

`waitForLoadedFrame` and `settleAfterMount` are the two the render check
calls; `waitForRecordedNavigation` and `settleWithoutNavigation` stay
internal to the new module.

Move only. Same 20 tests pass on both engines.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): send Referrer-Policy: no-referrer on the shell document

`img-src https:` gave the page somewhere to send a request, and the document
origin is `orca-mobile-web://<sessionId>/`, so a request that carries a
referrer carries the session id to whatever host an artifact or a markdown
document named.

`referrerPolicy="no-referrer"` on the preview iframe does not cover it.
Measured in the render rig against a permissive control policy: WebKit puts
the embedder's URL on a srcdoc frame's image request despite the attribute,
and Chromium sends none. So the guarantee belongs on the document, where one
header covers every request the page makes, and it rides the document alone
with the policy -- the referrer of a request is decided by the document that
made it, so on a subresource response it would govern nothing.

WKWebView under the custom scheme is unverified: the rig is Playwright
WebKit over http, not WKWebView over `orca-mobile-web://`. The header is the
hedge, and it costs nothing if that host never leaked.

Pinned three ways, each confirmed red first:
- Swift, exit 133 with the header removed.
- Kotlin, MobileWebShellResponseHeadersTest "sends the policy on the
  document" FAILED at :17 with it removed.
- The rig, through a new `readShellDocumentHeaders` that parses the Kotlin
  source the way `readShellCsp` does and throws rather than returning an
  empty map. With the value flipped to `unsafe-url` the WebKit arm fails
  `expected [ …(2) ] to deeply equal [ null, null ]`; with the line deleted
  the parse throws "could not parse the shell document headers".

The rig's arm carries its own presence precondition: a third server serves
the shipped policy with `unsafe-url`, so the WebKit reading is the header
doing the work, and Chromium's null either way is pinned as the browser's
behaviour rather than sold as evidence the header arrived.

MobileHtmlPreview.web.tsx said the iframe attribute kept the origin out of
the request. Corrected to name the header, since the measurement above is
what disproved it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): quote the current directive where the old text was written down

Three comments still read `img-src 'self' data:`, so a grep for the old
directive found live prose that no longer matches the header. Each stays
about `data:`, which is what those paths rest on; only the quoted policy
changes.

The two remaining hits in the repo are src/main/browser/doc-preview-protocol,
which is the desktop preview's own policy and not this one.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): name the surfaces img-src https: actually unblocks today

The comment justified `https:` with markdown and the rich editor, and
neither renders a remote image on the page. Verified in the tree:
MobileMarkdown paints `![](...)` as a tappable link at both of its image
branches and never mounts an Image, and it has no `.web` sibling, so that is
what native does too; MobileRichMarkdownEditor.web.tsx is a 92-line
multiline TextInput, still C7.6's plain source field.

What the directive unblocks today is four surfaces, none of them overridden
on the page:
- MobileAgentIcon's favicon, a hardcoded `google.com/s2/favicons` URL, used
  by thirteen callers including the session header and the worktree rows;
- MobileRepoIcon's project icon, a host-named favicon, avatar or upload, on
  the worktree list and the host workspace list;
- PRCommentCard's author avatar, from the review reply schema;
- the sealed HTML preview frame, which inherits the policy.

Markdown and the editor are named as the anticipated surfaces ruling 26
points at, so a later reader does not take the loosening as already covering
them. Both native pins carried the same wrong claim and are corrected.

That comment is the only record of why the policy loosened, so it says what
is true now and what is coming, separately.

Comment only: the parsed header is unchanged, checked through the harness
reader the render suite uses.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): point the new source-control route pin at the current directive

Merge resolution, not a conflict git could see. #21957 landed the
source-control and review page routes on main while this branch was open,
and its render check pins the directive text twice: `cspHeader` by substring,
which survives the widening, and the Swift source by the quoted literal
`"img-src 'self' data:"`, which does not. Two PRs green alone, red on the
merge.

Both pins now read the current directive.

One comment goes with it. "Not one request left the origin, so there is
nothing for the policy to have refused" now needs saying why: `https:` is
admitted, so an empty host list is these two closures fetching nothing
rather than the policy refusing something. The avatar that would fetch needs
provider data this page never gets, which the file's own closing note
already explains.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): wait for the admitted images before reading their hits

CI's Chrome 152 recorded the CSS background and not the `<img>` by the time
the bounded settle returned, so both https arms failed on a count: "expected
[ '/css-bg.png' ] to deeply equal [ '/css-bg.png', '/img.png' ]" and
"expected 1 to be 2". The reads were absence-shaped -- two frames and 200 ms
-- and the claim they carry is a presence.

So the arms wait for their own evidence, the way the `'refusal'` arm already
does. `frameReady: 'images'` polls until both admitted paths are recorded,
bounded by nothing but the case's own `ctx.signal`. It sits after the marker
wait, because an image is requested by a document that has parsed, and the
arm hands its reader in rather than the settling module reaching for state
that belongs to an arm.

One reader now serves the wait and the reading. An arm that waits on one
list and asserts on another has proved nothing about the list it asserts on.

The `/probe.woff2` absence is untouched and is now an absence standing
behind two presences rather than beside them.

What the wait prints when it does not arrive, captured by making the paths
unsatisfiable against a 12 s case:

  [html-preview-render] the arm recorded ["/img.png","/css-bg.png"] of
  ["/css-bg.png","/img.png","/never-arrives.png"]; #remote
  {"complete":true,"naturalWidth":1,
  "currentSrc":"https://artifact-images.invalid/img.png?n=n1",
  "loading":null}: arm csp=shipped sandbox=product frameReady=images
  nonce=n1 | browser 147.0.7727.15 | ... | frames [...]

`complete` with a zero `naturalWidth` is a request that finished and
produced no image; `complete` false is one still in flight. So a Chrome that
never issues the request says which of those it was, instead of a bare count.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): say why an admitted image never arrived, and hand back the context

CI's Chrome 152 read the `<img>` as complete with a zero naturalWidth and a
resolved currentSrc while the route handler never saw the request, and the
CSS background from the same origin did reach it. The diagnosis could say
the image failed but not why, because nothing was watching the request.

Now four sources are, for the `.invalid` origin only, in a module of their
own so the rig file stays under its cap: `request` says whether the page
asked at all, `requestfailed` carries the browser's `errorText`, and CDP's
`Network.loadingFailed` adds `blockedReason` and `corsErrorStatus`, which is
the only place a refusal names itself once the request never reaches a route
handler. `Network.requestWillBeSent` records the resource type, the initiator
and the frame, which separates an image the parser found from one nothing
asked for. They fill arrays while an arm passes and are only read on abort.

Proved by forcing the abort rather than assuming: with the awaited paths made
unsatisfiable, the reading names the font's refusal in both vocabularies at
once, `failed [{"url":".../probe.woff2","errorText":"csp"}]` and `cdp
loadingFailed [{"errorText":"","blockedReason":"csp",...,"type":"Font"}]`,
beside `cdp sent` showing every request's type, initiator and frameId.

Teardown: `open()` now takes an explicit context and closes both the page and
the context in a `finally`. The close used to sit on the happy path, so an
arm whose wait aborted and whose result reads then raced vitest's teardown
left its page and its implicit context open on a browser every later case in
that engine still runs on.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): correct three rationales the widening left wrong

(a) A review comment's avatar is not a surface the widening unblocks.
PRCommentCard renders it only under `Platform.OS !== 'web'` and a component
test pins the skip, so on the page it never renders. Dropped from both native
rationales and moved to the anticipated list beside markdown and the editor,
with the reason each is anticipated rather than current.

(b) The Kotlin rationale quoted the iOS origin. Android serves from
`https://<sha256(sessionId) first 32 hex>.orca-mobile-web.invalid/`, so a
referrer there carries a stable per-session handle and not the id itself,
while iOS serves `orca-mobile-web://<sessionId>/` and carries it verbatim.
Both are something an image host can key on across requests, which is what
the header is for; each file now names its own origin.

(c) "Only the script half of that is stricter than native" overstated it.
`font-src 'none'` and `connect-src 'self'` are stricter too. Images are the
one of the four that stopped being stricter, and the comment now says which
three remain and why.

A fourth, found while checking (a): the skip's own comment justified itself
with `img-src` being `'self' data:`, so a provider avatar would be "one
refused request per card". That is no longer true -- the avatar would load
now -- so the skip is a page capability gap rather than a policy consequence.
Recorded as such at the guard. Whether to lift the guard is a ruling-26
question and not this PR's.

Comments only. The parsed policy and document headers are unchanged, checked
through the harness readers the render suite uses.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): probe why Chrome never asks for the artifact image

CI's read was decisive: on Chrome 152 only the CSS background was requested,
while the `<img>` reported complete with a zero naturalWidth and a resolved
currentSrc. A request that went out and failed cannot produce both readings,
so the next probe asks the frame rather than the network.

On abort it now reads, inside the artifact frame: readyState, the init
script's own moment, document.images.length, every
`performance.getEntriesByType('resource')` name, the navigation entry types,
and for #remote its src, isConnected, complete, naturalWidth, currentSrc and
the outcome of decode(). A resource entry for a URL the rig never saw would
mean the request left the frame and died before reaching it.

Then it issues a `new Image()` at a URL that has never existed and reports two
seconds later whether the rig saw it. That splits the two live explanations: if
the fresh request is seen and the artifact's was not, the frame can fetch and
the parser-inserted element is the cause; if neither is seen, requests from
this frame are not reaching the rig at all. Subframe document commits are
counted from mount, because a second parse is a new window and leaves nothing
behind to count, and a second parse could be meeting a failure the first
cached.

`cdp sent` was empty on CI even for a request Playwright did record, so the
page's own session is blind to the frame. Chromium isolates sandboxed iframes
into their own process, srcdoc included, so flattened Target.setAutoAttach now
puts each child target on the same connection with Network.enable on the
child, and the attached list reports whether the frame is a separate target
at all.

The navigation arm gets the same reading, since CI showed it fails on its own
rather than behind the aborted image arms.

Verified by forcing the abort rather than assumed. Locally the reading prints
one subframe parse, decode resolved, every resource the document fetched, and
`fresh ... issued true seen true`, with the attached list empty, which is
consistent with this Chrome not isolating the frame and its page session
seeing the requests.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): time the artifact image against the frame's attachment

CI's second read showed the frame did issue the request -- it has a
resource-timing entry and decode rejected with EncodingError -- while the rig
saw only the CSS background, and a fresh image created later from the same
frame was both issued and seen. The remaining question is whether the entry
starts before anything was listening to that frame.

So the entry is now reported in full for the element under test:
responseStatus, transferSize, encodedBodySize, nextHopProtocol, startTime and
duration. A zero status with a zero transferSize is a fetch that reached the
network stack and came back with nothing, which is what an unintercepted
request looks like once `.invalid` fails to resolve.

Both sides of the comparison get a wall clock: `Target.attachedToTarget` and
Playwright's own `frameattached` now carry the moment they fired, and every
recorded request carries the moment it was seen. An entry that starts before
the attachment is the race stated rather than inferred.

Abort path only; the passing run is unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): serve the artifact's https assets from a real TLS listener

Interception could not measure what the directive admits. Chrome 152 isolates
the sandboxed srcdoc frame into its own target and the parser-inserted `<img>`
is the document's first fetch, issued before interception attaches there: the
request escaped to the real network, `artifact-images.invalid` did not
resolve, and the rig recorded nothing while the frame's own resource timing
showed the fetch and a later fresh image was both issued and seen.

So the assets come from a listener that is already accepting before the page
exists. It cannot be raced: the request arrives or it does not, and either
answer is the measurement. Hits and referrers are recorded server-side, the
way this rig's cleartext origin already does it, and read per arm by nonce.
`img-src 'self' data: https:` matches on scheme, so `https://127.0.0.1:<port>`
exercises the same directive as any other https host.

Lifecycle: started in beforeAll before any browser, closed in afterAll beside
the other servers. Its certificate is generated per run by openssl into the
suite's own scratch directory under `mobile/.tmp`, which the root gitignore
already covers and into which the server writes a second `.gitignore` as well;
the key never leaves that directory and nothing trusts it, since the context
is created with `ignoreHTTPSErrors`. No arm shares state: one hit list keyed
by each arm's nonce, and the permissive-Referrer-Policy control stays what it
was, a second bundle server serving the page, because the control is the
document's header and not the image host's.

The navigation record moves off interception too. It is now `page.on('request')`,
one subscription over every frame, armed after the rig's own `goto` exactly
where the route used to be registered; the route stays only for what only a
route can do, refuse the navigation. That answers the top-nav arm's `recorded
[]`: its record depended on the same per-target interception.

And the arms stop swallowing their clicks. `click(...).catch(() => {})` made a
tap that never landed and a tap that produced no navigation the same empty
counter; `open()` now records the error and the two top-nav arms assert it is
null before reading any count.

One correction to the reading added in the previous commit. The resource-timing
fields came back zero for a request that had plainly succeeded: they are opaque
cross-origin. The listener now sends `Timing-Allow-Origin`, after which
transferSize, encodedBodySize and nextHopProtocol carry real values.
`responseStatus` still reads zero on a successful request, so the comment names
the three that discriminate rather than the four that are printed.

24/24 on both local engines.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): compare the artifact fetch and the attachment on one clock

The early-or-late comparison spanned two clocks and could not answer the
question it was written for. Every `at` in the request log is Node's
`performance.now()`, counting from process start; the resource entry's
`startTime` is the frame's own, counting from that document's navigation. A
frame entry reads as earlier than a Node attachment by roughly the process
uptime, so the comparison would have reported the race as confirmed on every
run, including runs where there was no race. A green CI would not have caught
it.

So the comparison is stated where both numbers actually live: `asked` against
`attached` in the request log, on the Node clock alone. `startTime` and
`duration` stay, labelled as the frame's own account and explicitly not
comparable to an attachment time. The module docstring says the same, so the
next reading added here starts from the rule rather than rediscovering it.

The commit message of b5e82065f3 carries the same overstatement and is left
as it stands; this is the correction.

Also the stale route-handler references, now that the asset listener records
the secure origin and the navigation record is a page subscription. Three were
in the review; two more were not, and both were stale for the same reason:
`waitForRecordedNavigation`'s docstring still credited the route with
recording a main-frame navigation, which stopped being true when the record
moved off interception, and the request log described a refusal as one the
request never reached a route handler with. The route now only refuses; it
counts nothing. The one remaining mention is the deliberate contrast in the
rig that says the record is the page's event and not the route's.

Comments only. 24/24 on both local engines.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 08:34:38 -04:00
Jinwoo Hong d9954000b3 refactor(mobile): the rich editor's document becomes scope-threaded modules and a bundled factory (OTA phase C, C7.10 C1) (#21969)
* refactor(mobile): split the rich editor document's stylesheet and markup apart

The body constant carried the tail of a `:root` block, every CSS rule and the
editable surface's markup in one string, which only the HTML builder could
splice. A page mounting the document needs the stylesheet and the markup
separately, so they become a function over the theme and a constant.

Byte-for-byte inert: `mobile-rich-markdown-editor-document.test.ts`'s digest of
the shipped document is unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): give the keyboard-inset normaliser its own module

It is the host's half of the inset, read by the controller, and it sat in the
module holding the document's in-page script. The script is about to become
ordinary TypeScript under `rich-markdown/`, where a native-side normaliser does
not belong.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): the rich editor's document becomes scope-threaded modules and a factory

The editor's ~600-line program lived in seven string constants a concatenator
glued into one `<script>`: unreadable, untypeable, and unreachable from a page,
which is where the OTA shell has to run it (ruling 26).

It is now ordinary TypeScript under `src/components/rich-markdown/`. Every
function that touches editor state takes `scope: RichMarkdownEditorScope` first,
`createRichMarkdownEditorDocument(host)` builds the scope, runs the start
sequence and returns `{ send, stop }`, and the six window reads the script did
are host seams with those reads as their defaults: `postToHost`, `promptForUrl`,
`keyboardInsetSource`, `clearTimer`, `getSelection`, `getDocument`.
`runCommand` is async because a host that answers the URL prompt with a modal
cannot answer synchronously; the thirteen commands that never wait stay one
synchronous act.

No module holds a `let` and none does work at parse time (rulings 20, 21), so a
second mount starts from its own state and `stop` takes back both the surface's
four listeners and the viewport's two.

The native document is an esbuild IIFE bundle of `native-document-entry.ts`,
written beside the terminal document's artifact by a fifth postinstall
generator. Nothing ships it yet: the HTML builder still splices the old strings,
which the next commit changes.

Red-first: `rich-markdown-document-parse-time.test.ts` and
`rich-markdown-host-seams.test.ts`. Their readers are the terminal census's,
extracted to `src/test-support/webview-document-census.ts` and pointed at both
documents rather than copied.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): ship the bundled document and retire the editor's script strings

`buildMobileRichMarkdownEditorHtml` splices the esbuild bundle of
`src/components/rich-markdown/`, and the seven string constants and their
concatenator go. `escapeInjectedJavaScriptString` stays: it is the escape for
`injectJavaScript`, which is still how the native host reaches the document.

Equivalence, since a byte golden over the script cannot survive a bundler:

- `rich-markdown/native-document-bundle.test.ts` evaluates the shipped artifact
  exactly as the WebView does — its markup, its bridge, its `execCommand`, its
  `prompt`, its `visualViewport` — and drives it through the injected handle:
  `keyboardInset` then `ready`, all five members, a markdown round trip through
  the real escape, an edit under the host's generation, every toolbar command's
  engine verb, the `javascript:` refusal, a tapped link, and the module list.
- `mobile-rich-markdown-editor-document.test.ts` keeps a byte pin, now over the
  page around the document. Measured on main's own document with its script
  region removed and on this one: 5,621 bytes, both
  `5054e1d5c87e4ce1805d4856ddc8bf36804e697675e6013d84da453d3e81af25`. The
  whole-document digest it replaces was `1ef29c88…`, 29,852 bytes.

Every assertion `mobile-rich-markdown-editor-html.test.ts` made by extracting
functions out of the emitted text is kept, aimed at the modules:

- nested/ordered/task list rendering and serialization, entities, explicit
  numbering, the parent-start fallback, read-only checkboxes →
  `markdown-round-trip.test.ts`, over real elements rather than shaped objects.
- the emitChange/setEditable guards and the generation carried through a
  replacement → `editor-content.test.ts`, behaviourally.
- dismissKeyboard, the tapped caret, the label tap, the restored caret, the
  end-of-document fallback, the detached caret → `editor-selection.test.ts`,
  with a blur that drops the ranges the way WebKit does.
- parseable script and the injection escape stay in the HTML test.

New with the factory: `document-lifecycle.test.ts` — stop takes the four surface
listeners and the viewport observer off, a second mount is its own document, two
documents do not share `editable`, and a start that throws unwinds.

`use-mobile-rich-markdown-editor-controller`, `MobileRichMarkdownEditor` and the
web fallback tests are untouched and green.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read the document's mutable bindings from the tree, not the line start

The census matched `/^(let|var) /gm`, so `export let`, a declaration indented
inside a top-level block and a `for (let …)` head were all invisible — three
shapes of the one binding two documents would share — and its single
precondition proved only the shape it could already see.

`moduleLevelMutableBindings` walks the program instead and stops at every
function body, because a binding one call owns is not module state. Its
preconditions are one per shape, with the kind each reports, and a negative case
over a `const` and a function-local `let`/`var` so the empty list is a
measurement rather than a reader that refuses everything.

Red-first: `export let pendingReport = 0` planted in `keyboard-inset.ts` reds it
with `keyboard-inset: let pendingReport`, which the old matcher passed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): pin both WebView document bundles to the mobile root

esbuild writes each module's path into a bundle as a comment relative to the
working directory, and neither generator set `absWorkingDir`. So the artifact's
bytes followed the cwd of whatever postinstall run wrote it: measured from the
repo root, `mobile/`, and `mobile/src`, three digests — and from outside the
repo the comments carried `/Users/<name>/…`, a machine path in the one file
every bundle test compares against a build it makes itself.

Both generators now pin the mobile root, so the four cwds measured agree, and
both bundle tests carry the pin: a digest built in a child process from the OS
temp directory equals the committed artifact's, and no comment in either
artifact is an absolute path or climbs out with `../`.

`build-terminal-document-script.mjs` had the defect verbatim on main; C1 copied
its shape, so both are fixed here rather than leaving the original to be found
again. Neither artifact's bytes move: both were generated from `mobile/`, which
is what `absWorkingDir` now names.

Red-first: deleting the `absWorkingDir` line from either generator reds that
generator's case.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): cover the getSelection seam's override, not just its default

Five of the six seams had both halves and this one had only its window default,
which is the half that cannot fail on the page: there the caret has to come from
the object the host hands over, because a document mounted inside a screen
shares `window` with every other field on it.

The case gives the document a selection of its own, blurs the surface the way
WebKit does — dropping the ranges, which is the whole reason a caret is saved —
and reads the restored caret back out of the host's object. The window's own
selection stays empty throughout, which is what says the default was never
consulted.

Red-first: `rememberSelection` reading `window.getSelection()` instead of the
field reds it; every other case in the file stays green.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make the editor document's stop cancel its pending timer

`stop` took the surface's four listeners and the viewport's observer off and
left the input timer, while the scope kept the handle and the `clearTimer` seam
kept the means to cancel it. A listener comes off with the element it was on; a
scheduled callback holds the scope and fires into a document the host has
already unmounted, posting a change under the generation of content it has
replaced.

`stopEditorContent` cancels it through the seam and clears the field, and the
sequence runs it last — after the listeners that could have scheduled another
one are gone.

Nothing schedules the handle today. The cancel is here because the seam and the
field exist for the day something does, and that is not the moment to discover
`stop` never reached it. The case plants the pending change rather than waiting
for a debounce, and carries its own control: the same timer posts while the
document is running, and posts nothing once it is stopped.

Red-first: dropping `stopEditorContent` from the sequence reds both that case
and the parse-time census's start/stop set comparison.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): correct the postinstall generator count in both censuses

Two comments said four generators and six generated files. There are five
generators writing six files, and the six are not the six either comment
described: `census-source-files.ts` still named the page's copy of the terminal
document, which ruling 25 retired and #21962 stopped ignoring, while C7.10 C1
added the rich Markdown editor's.

Both now name the lists of record — `mobile/package.json`'s postinstall for the
generators, `mobile/.gitignore` for the files — and say the count is a reading
that grows rather than a fence, which is what made the old numbers wrong twice
over.

Verified against both lists: 5 and 6.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say which digest is the document and which is the page around it

The docstring put main's whole-document digest and byte count in the sentence
introducing the shell pin, so it read as if `1ef29c88…` and 29,852 bytes were
what the constant below asserts. They are not: that digest is of main's whole
document, script included, and nothing in the file reproduces it. The constant
is of the document with its `<script>` region emptied, taken on main's document
and on this one.

Both are now named and separated, with what each covers and why the shell one
was read twice.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): name the parse-time fixture by its role

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): drop an editor command whose dialog answered after the host moved on

C7.10 C1 made `runCommand` async so a host can answer the URL prompt with a
modal. Inside the WebView that changes nothing — `window.prompt` resolves within
a microtask, and the host reaches the document through `injectJavaScript`, which
is a later task — but on the page the modal is a real task boundary, and while
it is open the host can replace the content, make the editor read-only or
unmount it entirely. The continuation ran anyway: `createLink` against markdown
nobody chose, and a change posted under the new generation carrying an edit made
against the old one.

`acceptsCommands` is the question both halves ask: not stopped, still editable,
still the same generation, still contenteditable. `insertUrl` asks it before
`execCommand` and `runCommand` asks it again before emitting, each against the
generation read before its own wait.

The scope gains `stopped`, which `stopRichMarkdownEditorDocument` sets.

Inert on native, where no state can change across a microtask, so the answer to
both questions is the one the old code assumed.

Red-first: with either check removed, the new case reports
`[ 'createLink', 'createLink' ]` against `[ 'createLink' ]`. The case carries its
own control — an answer that arrives while nothing has moved is still applied
and still reported.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make the editor's block reader always consume a line

`markdownToHtml` looped forever on `# `, `- ` and `1. `. `isBlockStart` admits a
marker followed by a space, and the list test admits the same, but the heading
reader requires text after the hashes and `parseListLine` requires text after
the marker — so on those lines the list branch consumed nothing and returned the
index it was given, and the paragraph loop gathered nothing and pushed an empty
paragraph without advancing. A one-line file the host handed to `setMarkdown`
froze the WebView.

Two guards, both by the same rule: a branch may only commit if it moved the
index. The list branch falls through when its run is empty, and the paragraph
falls back to the line itself when it gathered none.

Present on main verbatim, so this is inherited rather than introduced — but the
fix is observationally inert, because the only inputs it changes are the ones
that previously never returned. Every input that produced output produces the
same output.

Evidence, from a probe that bounds the loop from the inside rather than waiting
on it: before, `# ` and `- ` both UNBOUNDED; after, twenty marker and fence
shapes all return. The pinned cases carry their own control, `# ok` and `- ok`,
so the fallback is not swallowing the readers it falls back from.

A red-first case is not possible here: without the fix the case does not fail,
it hangs the worker. The probe above is the measurement.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): see every declaration that runs as a document module is evaluated

The parse-time reader inspected only variable declarations, while
`DECLARATION_KINDS` admits classes and default exports. So
`class A { static value = install() }`, a static block, and
`export default install()` all passed a census whose whole job is to refuse
exactly that — and a static field reading `document` passed too, which is the
remount defect the rule exists for, wearing a different shape.

Three shapes now, each reported by what it does rather than what it looks like:
a variable initialiser, a class's static members, and a default export that is
an expression. `DECLARES_WITHOUT_RUNNING` keeps the last one from walking into
the body of `export default function () {}`, whose calls run when something
calls it.

The preconditions are one per shape, with a negative case beside them: an
instance field runs per `new` and nothing in a document is ever constructed, and
a default-exported function declares a body rather than running one.

Inherited from the terminal's census, which had the same reader; both use this
one, and both are green.

Red-first: removing the class branch reds the new precondition case.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read lifecycle exports from the tree, not from one exact spelling

The reader was a regular expression needing `export function`, one line, the
scope parameter and no return type. `export async function startX(`, a return
type, or a parameter list the formatter wrapped made a real lifecycle export
vanish — and the comparison it feeds is a set against the names the sequence
calls, so a function missing from *both* lists makes them agree. A start nobody
runs would have read as a start nobody needs.

It now qualifies a function by what it is: exported, named for its lifecycle,
and taking the document's scope as its only parameter. That last clause is
ruling 20's own wording — a start takes nothing the scope does not already carry
— and the regex was enforcing it by accident, through the single parameter its
pattern happened to allow.

Surfaced by the change: the terminal's `startEdgeScroll(scope, dir)`, which the
regex never matched and the sequence never calls. It takes a direction, so it is
the overlay's act for a drag rather than a module's lifecycle, and the one-
parameter rule refuses it for the stated reason instead of by accident. Both
censuses are green.

Red-first: restoring the regex reds the new precondition case, which covers
`async`, a return type and wrapped parameters, with refusals beside them for a
two-parameter start, another document's scope type, and an unexported function.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 07:16:23 -04:00
Jinwoo Hong 23207bfde2 feat(mobile): register the source-control and review page routes (OTA phase C, C4.4) (#21957)
* refactor(mobile): move the review route body onto a component and the handoff seam (OTA phase C, C4.4)

The review route file called `useMobileDiffReviewController` at its top level. A switch cannot
keep it there: hooks are unconditional, so the whole controller — its client subscriptions
included — would run behind the shell's page whenever the shell renders. As an element passed for
`fallback` it is created and not mounted, which is how the explorer switch already behaves.

`useRouter` becomes `useRouteHandoff` in the same move. It was the one raw expo-router router left
in the review closure (measured: the only other value import of one is the seam's own web sibling),
and inside the page the session screen `openSession` replaces to is native, so that target has to
be handed back to the app rather than posted into a document that does not render it.

The params are read in the component rather than handed down, so this is the route body and the
route file above it is free to become a switch.

`session-router-seam-census.test.ts` gains the module by name. Kept with `useRouter` the census
reds twice — `imports nothing from expo-router that can navigate` names
`MobileDiffReviewRouteScreen.tsx (useRouter)`, and the completeness case gains `useRouter` — which
is what forces the swap rather than leaving it to a reviewer.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): switch the source-control and review routes to the shell, still unregistered (OTA phase C, C4.4)

Both take the files switch's shape: `firstParam`/`firstReviewParam` on every param, `shellScreenRoute`
as the one predicate, `MobileWebShellScreen` keyed on `shellScreenRouteKey`, the native screen built
as an element and passed for `fallback`. Both gain a `.web.tsx` sibling for `index.web.tsx`'s reason —
the native file reaches OrcaMobileWebShellView, whose module throws at import in a browser, and the
route manifest imports every route.

Inert on its own. A switched route renders the shell only once `MOBILE_WEB_PAGE_ROUTES` lists it,
which is the next commit; until then the flag is the only thing that changes and it is off.

Query params are omitted rather than sent empty, and the whole record is omitted when none was
named: `tab=` is a lens named nothing and lands on `changes` through a different branch than an
absent one, and the same holds for `name`, `origin`, `scope`, `file` and `area`.

`pr` and `history` are deliberately not switched. Both are `Redirect`s into `source-control`, and a
redirect inside the page would leave the session bound to a pathname the page has left; left native
they replace into this route and its switch mounts the shell.

Three censuses red without their rows, measured on this tree:
- `mobile-web-app-web-overrides.test.mjs` `lists exactly the .web.* files on disk` names the two new
  siblings; `states a reason for every override` reds on a placeholder under 20 characters.
- `mobile-web-shell-flag-census.test.ts` `reaches the switched routes through that hook and no
  others` reds without the two `SWITCHED_ROUTES` names.
- `shell-screen-route-census.test.ts` `walks the route tree and finds them` reds without the two
  switch names.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): register the source-control and review page routes (OTA phase C, C4.4)

Two entries in `MOBILE_WEB_PAGE_ROUTES`, five grants each, with the reason for each grant read off
the screen that needs it. The two lists are equal on purpose: the hub's changed-file rows push
review and review replaces back, and a target declaring no more than its opener is a hop the
handoff keeps inside the document. Registering either alone would have put a native frame and a
second bridge session between a changed-file row and its diff.

`pageRouteGrants` is derived from this list, so the two rows are a consequence of the entries and
there is no second table to edit. `pr` and `history` stay native redirects and are never listed; the
derived target list at this tree is [files, files/preview, source-control, [p], accounts,
agent-history, review, session, tasks, web], with no `pr` or `history` row, because the census reads
call sites and both redirects name `source-control`.

Measured on this tree, not carried from the draft:
- The hop census goes 8 -> 16. The eight new rows are exactly `{/h/[hostId], agent-history,
  files/[worktreeId], files/preview} -> {source-control, review}`, each handed off for
  `native.clipboard.write` and the first four also for `externalLink`. `source-control <-> review`
  is absent in both directions, which a new case now asserts as grant-list equality rather than as
  the absence of a row — absent is also what an unregistered route looks like.
- The Back census now walks six trees and finds 8 controls, both rules printing empty. The two new
  ones are `MobileSourceControlHeader.tsx:46 role=button label=Back to session` and
  `MobileDiffReviewHeader.tsx:48 role=button label=Back`, which is what C4.3 bought. The
  `ARRIVING_SCREENS` describe it wrote for this moment is removed: with the rows in
  `PAGE_SERVED_SCREENS` its trees are covered and its cases were a second reading of the same thing.
- Both closures reach the haptics seam, so `haptics` is declared by measurement: the seam census
  derives the reaching set and its two cases pass with the routes in its `ROUTE_MODULES` map.

Without the two manifest entries these red on this tree: `pins every hop the handoff must take away
from the page`, `keeps the hub and review local to each other`, `declares only routes the bundle has
a module for`, `reaches the built manifest`, `covers every page route and finds a control in each`,
and both haptics-seam cases.

`build-mobile-web-app-bundle.test.mjs` is split rather than fenced. The two pinned entries put it at
607 non-comment lines against the 600 cap, and the declaration block is a different concern from how
the bundle is built — it grows once per registered domain while that file does not. It moves whole
into `mobile-web-page-routes.test.mjs`, named for the module it is written against, so the next route
to register does not have to choose between a lint fence and a split it did not ask for.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): render-check the two page routes, and make the oversized stage-all readable (OTA phase C, C4.4)

The render check mounts both routes in a real browser, asserts each paints its own screen rather
than the Unmatched route with no console error and no page fault, and asserts each fetches its own
chunk on a client-side navigation. It also reads the shipped `img-src 'self' data:` out of the
Kotlin source it is served with, pins the Swift twin beside it, and asserts neither route leaves the
origin or logs a policy violation while it paints.

The avatar skip itself (ruling 3) is `PRCommentCard`: on web it renders its existing empty-avatar
`View` rather than letting one `<Image>` per comment attempt a fetch the policy refuses. Its branch
is pinned by a component test, which reds on the platform check being removed. The render check's
off-origin case is honest about being the negative half only — no comment card renders there,
because the PR chain behind it is not scripted, and the file says so.

The `useAnimatedScrollHandler` risk is answered by the two static facts rather than by a probe, and
they are recorded as assertions: the hook is deliberately outside the four `MAPPER_HOOKS` because it
is an event handler, and its updater's only effect is a write to `scrollOffsetY`, which
`RightDrawer.tsx` assigns in two places and reads in none. A later read reds that case the moment it
is added.

The `oversized` stage-all refusal (ruling 2, made testable by ruling 5) was a silent no-op, and this
is the fix as well as the case. Measured on this tree before it: `git.bulkStage` with 12,000 paths
posts one 1,033,012-byte frame, the shell's reader drops it with `{ kind: 'refused', refusal:
'oversized' }`, and the page's promise never settles — `busyAction` never cleared and
`setActionError` was never called. Both new cases red by timing out at 15s against that path.

Refused at the page's own send boundary instead, under the shell reader's own predicate rather than
a second spelling of it: `isBridgeFrameWithinCap` is extracted from `parseBridgeMessage` and used by
both sides. `sendFrame` answers `sent` / `oversized` / `port-failed`, so `sendRequest` rejects with a
`BridgeRequestOversizedError` whose message is a sentence the panel puts on its error surface, and
the members whose contract is a boolean keep it. No delivery-unknown mark: the frame never left, so
nothing ran on the desktop and the smaller retry is safe to offer.

The case runs the real chain — bridge port pair, `useMobileGitRequests`, `runGitWorkflow` — with
only react-native and the haptics seam mocked, and asserts the message that lands is a sentence and
that the busy flag is raised and then cleared.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): drop the type assertions from the two new C4.4 test files (OTA phase C, C4.4)

The changed-code quality gate named five, all in the files the previous commit added, and a fence is
not the answer to any of them. A separate commit because a reported head does not move by amend.

- The comment fixture is a real `PRComment` rather than a cast: the type's six required fields are
  all this case needs, and the SAFETY disable that stood in for them was inert anyway — oxfmt had
  wrapped it onto three lines, and a wrapped `oxlint-disable-next-line` matches nothing.
- The image lookup goes through `findAll` on the host tag rather than `findAllByType`, which takes a
  component. Through `String`, because `node.type` is `ElementType` and React Native declares no
  intrinsic elements, so the compiler reads a bare tag comparison as unreachable.
- The runners hook takes its router from `useRouteHandoff` with expo-router mocked under it, which
  is how a `RouteHandoff` is obtained rather than asserted into existence. No target is pressed.
- The rejection and the diagnostic are read through narrowings instead of casts, which also drops
  an `expect.any` that only type-checked because of one.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep a frame the page cannot serialize inside the send contract (OTA phase C, C4.4 round 1)

Round 1 finding 1. The oversized refusal moved `JSON.stringify` outside `sendFrame`'s `try`, so a
frame carrying a cycle, a `BigInt` or a throwing `toJSON` threw past the whole send path. Three
things followed, all measured here on a cyclic `params`:

- the caller was rejected with a bare `TypeError` from `JSON.stringify` instead of the
  `BridgeSendFailedError` every other undelivered frame raises;
- no `send-failed` diagnostic was raised, so nothing recorded that a frame had been lost;
- `sendRequest` opens the id before it posts and abandons it on the way out, and the throw skipped
  the abandon: 63 of the 64 in-flight slots were usable afterwards, against 64 on a client that sent
  no such frame. Sixty-four of them and every later request is refused with nothing to say why.

`posted()` carried the same escape into the members whose contract is a boolean, where a throw is
worse still: those callers are taps and teardowns with no catch on them.

Serialization goes back inside the `try`, with the oversized refusal kept in front of the post. The
docstring said the port arm's throw is never `JSON.stringify`'s, which was exactly the assumption
that broke; it now says why the call sits where it does.

The new file is the pin: the rejection's name, the diagnostic, nothing reaching the shell, and the
slot count with a no-cyclic-frame control beside it so the count cannot pass by the cap moving. Both
changed cases red on the serialization moving back out — `expected 'TypeError' to be
'BridgeSendFailedError'` and `expected 63 to be 64`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the page's frame cap to the reader's, at the boundary and by construction (OTA phase C, C4.4 round 1)

Round 1 finding 2. Nothing held the sender's predicate to the reader's. Replacing
`isBridgeFrameWithinCap(json)` with an inline `json.length > BRIDGE_MAX_MESSAGE_BYTES + 1` passed 85
of the 86 mobile-web-shell and source-control test files on this tree, and a frame at exactly cap+1
would then be posted and silently dropped — the hang the refusal exists to end, back for every frame
in that one-unit band.

Two rules, because either alone passes against the defect:

- The boundary. A frame of exactly the cap is posted, arrives at `parseBridgeMessage` and is
  accepted; a frame one byte over is refused with `BridgeRequestOversizedError`, posts nothing, and
  is the same string the reader answers `oversized` to. An off-by-one reds the second.
- The census. The client reaches the cap through the shared predicate and does not name
  `BRIDGE_MAX_MESSAGE_BYTES` at all, and the module that exports the predicate is the module that
  parses inbound frames. A private copy that is correct on the day it is written reds here.

The overhead the boundary frames are built from is itself checked rather than trusted: a frame asked
for at exactly the cap must serialize to exactly the cap, so the constant cannot rot behind an
envelope that grew a field.

Against the mutation both new rules red — `expected null to be 'BridgeRequestOversizedError'` and
the census failing to find the predicate — while the rest of the suite stays green, which is the
finding reproduced.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): count the outbound frame in the unit both shells count it in (OTA phase C, C4.4 round 1)

Round 1 finding 3. Read off both shells rather than assumed, and they agree: iOS gates the inbound
frame on `json.utf8.count` (`MobileWebShellView.swift`, through
`MobileWebShellBridge.acceptsByteCount`) and Android on `json.toByteArray(Charsets.UTF_8).size`
(`MobileWebShellView.kt`, through `acceptsMobileWebShellBridgeByteCount`), both against `640 * 1024`.
UTF-8 bytes on each platform.

The predicate was already right. `isBridgeFrameWithinCap` decides on `utf8ByteLength`, and the
`raw.length` clause in front of it is a cheap refusal in the safe direction, not a second rule: every
code unit encodes to at least one byte, so a string over the cap in units is over it in bytes too.

The diagnostic was not. It reported `json.length` — UTF-16 code units — in a field named `bytes`, so
a frame of CJK text read as a quarter of the cap at the moment it was refused by it. It now reports
`utf8ByteLength(json)`, and the type says which unit that is.

Pinned with a 250,000-character frame of three-byte characters, which is under the cap in code units
and over it in bytes, plus a source case reading the measuring expression out of each shell. Three
mutations, all red: dropping the byte clause from the predicate reds the refusal (`expected null to
be 'BridgeRequestOversizedError'`) and the diagnostic; reporting `json.length` again reds the
diagnostic alone (`expected 250094 to be greater than 655360`), which is the defect this commit
fixes, in the number it would have printed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(config): drop the render check's avatar assertion, which could not fail (OTA phase C, C4.4 round 1)

Round 1 finding 4. The case asserted that no avatar host was requested while both routes painted,
which reads as a proof of the web skip and is not one: no comment card renders on either page,
because the PR chain the file's own closing note names is not scripted. Reproduced here — deleting
the `Platform.OS !== 'web'` guard from `PRCommentCard` leaves the file at 5 passed.

Deleted rather than propped up. Giving the page a presence precondition means five hand-written
fixtures against five Zod schemas inside the shell double, which is exactly what the harness's
docstring says that double must not become. So the only proof of that branch is
`pr-comment-card-web-avatar.test.tsx`, which reds when the check is removed, and the render check now
says so in its header instead of implying otherwise.

What survives is a property of these two closures rather than of that component: not one request
leaves the origin while either route paints, and nothing either paints violates the policy. That one
can fail — planting a `fetch` to a provider host in a module both routes reach reds it twice, on the
console-error case and on the off-origin case, with the `connect-src 'self'` refusal in the output.

The CSP half is unchanged and was never in question: the served header is read from the Kotlin source
and the Swift twin is pinned beside it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 05:32:24 -04:00
Jinwoo Hong 8d42410e01 feat(mobile): render the HTML preview on the page in a sealed srcdoc frame (OTA phase C, C7.10 A) (#21862)
* feat(mobile): offer a cancelled top-frame navigation to the shell's opener

Both shells cancelled every navigation off their own document in silence: iOS
`decidePolicyFor` allowed only `isMainFrame && isDocumentUrl`, Android's
`shouldOverrideUrlLoading` dropped anything whose resolved path was not "/".
Nothing opened. That is the whole of ruling 29's "if they do not": a user tapping
a link inside C7.10's sealed HTML-preview frame reaches the top frame as a
navigation request, and the shell was the only thing that could act on it.

A cancelled main-frame navigation now reaches JS as `onExternalNavigation` and
goes through the same `Linking.openURL` the `externalLink` notify already uses.
The scheme list is not restated natively: the native side caps the string and
says which frame it came from, and `readBridgeExternalLinkUrl` decides what opens
in the half that ships over the air. A subframe navigation is never offered,
because that is the sealed preview loading itself.

swiftc check: OK (`checkCancelledNavigation` added, the whole suite runs).

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): render the HTML preview in a sealed srcdoc frame on the page

C7.6 gave the page the artifact's source, which is the native component's Source
tab and half its job (ruling 8). Ruling 26 makes that debt: the Preview tab comes
back as an `<iframe sandbox srcdoc>` inside the page's own document.

`srcdoc` rather than a `blob:` URL, and no CSP change at all. Measured on Chromium
and WebKit: a `srcdoc` frame has no URL for `frame-src` to match and inherits its
embedder's policy instead, so it is admitted under the shipped `frame-src 'none'`,
while a `blob:` frame is refused by `frame-src` on both and refused a second time
in WebKit by the `frame-ancestors 'none'` it inherits.

Two independent fences seal it, and the render check measures each on its own:
the sandbox grants neither `allow-scripts` nor `allow-same-origin`, and the
inherited `script-src 'self'` refuses the artifact's inline script even when a
control arm grants `allow-scripts`. The inherited `img-src` and `font-src 'none'`
govern its subresources, against a no-header control where the same three are
fetched.

`allow-top-navigation-by-user-activation` is the one token granted (ruling 29), so
a tapped link becomes one top-frame navigation the shell now opens externally,
while a `<meta refresh>`, a form submit, `target="_blank"` and any script-initiated
navigation produce none.

`lucideBarrelPlugin` is exported from the bundle builder so the check builds the
toolbar's icons the way the page does rather than carrying a second shim.

config/scripts suite, this file: 14 passed, 0 errors, exit 0. Control runs: a
literal `sandbox` in the JSX reds 4, an added `allow-scripts` reds the script
fence and the token census.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the preview's sealed frame where the degradation was pinned

The three HTML-preview cases in this file described the state ruling 26 retires:
no toggle, no frame, the source only. They now pin the frame's shape through the
test renderer -- the artifact reaches it as `srcDoc`, the sandbox grants neither
`allow-scripts` nor `allow-same-origin`, both toggle positions exist, and Source
takes the frame away with it -- and the "never renders the html itself" case
becomes "never puts it anywhere but the frame", counted rather than merely absent.
What a browser does with that frame stays in the render check, which is the only
thing that can answer it.

The rich Markdown editor's half is unchanged: it is still the plain field, and
item C is a later PR.

Two mocks added: `Pressable`/`ScrollView` on the react-native double, because the
toggle renders one, and `lucide-react-native`, whose barrel imports a
`LucideProvider` its own context module does not export and so does not load under
vitest at all.

9 passed, exit 0.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): refuse a link-activated top-frame navigation, even to the document

F1, blocking, with F5 and F6 folded in because they are the same decision and
splitting them would mean three rewrites of one function.

F1: `<a href="/" target="_top">` and `href=""` in an artifact resolve against the
embedder's base, so both named the shell's own document URL -- which both shells
ALLOWED (iOS `isDocumentUrl`, Android's path `/`). One tap inside the sealed
preview reloaded the shell's page: bridge target cleared, load state restarted,
page state gone. A navigation a human started is now never allowed, whatever it
names; it is offered instead, and `cancelledShellNavigationTarget` drops
`orca-mobile-web:` in silence exactly as it drops `/h/other`. The page rewriting
its own path carries no gesture and is still allowed.

F5: the OFFER is gated on the same gesture, so a top-page meta refresh or a
redirect is cancelled and never opened externally.

F6: iOS returned early on `shouldPerformDownload` before the offer, so `<a
download>` was dead on iOS and opened on Android. The early return goes; a
download is refused rather than allowed when nothing started it, and a
gesture-started one reaches the opener on both platforms.

The allow half and the offer half are now one function per platform
(`MobileWebShellNavigationPolicy.verdict`, `mobileWebShellNavigationVerdict`), so
they cannot drift. The gesture is the platform's own answer: `.linkActivated` on
iOS, `request.hasGesture()` on Android.

Native tests, both platforms: document URL + gesture refused and offered; document
URL without gesture allowed; foreign + gesture cancelled and offered; foreign
without gesture cancelled and silent; download both ways; subframe never offered.
swiftc OK; control run with the gesture rule removed exits 133. Gradle
MobileWebShellDroppedNavigationTest tests=8 failures=0 errors=0.

Also corrected: the screen comment that claimed the document's own reloads reach
the handler (they never do), and the prop doc, which now states the gesture rule.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): count own-origin top-frame navigations, and drop the goto cap

F2: `page.setDefaultTimeout(4000)` capped `page.goto` at 4 s while every sibling
render check uses the 30 s default, so under load the first WebKit cases redded on
the navigation rather than on anything they assert. The cap goes; the per-action
timeouts that needed to be short are already passed at their call sites.

F1's page-side half: the rig now routes the page's own origin as well as the
foreign one and counts main-frame navigations to each separately, with two cases
pinning that `href="/"` and `href=""` each produce exactly one own-origin
top-frame request. Playwright is not the shell, so what these state is the request
the shell is handed; refusing it is the native tests' job and the docstring names
which ones. The own-origin route is registered after the initial load, because it
aborts main-frame navigations and the first `goto` is one.

The foreign-tap and meta-refresh cases now also assert zero own-origin
navigations, so a fix that merely moved the target would not pass.

16 passed, exit 0, no Errors line.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): wait for the preview frame's own load, never a clock

CI read the child frame before its srcdoc committed: frameUrl came back ''
and the control arm's script as not yet run. The frame list, the frame's URL
and anything read inside it settle at their own moments, and a 900 ms wait
reads whichever of them has happened -- on a loaded runner, none.

Polls for a child frame at about:srcdoc with its load fired, bounded by the
case's own timeout, and an override arm now resolves on the document its
srcdoc assignment commits rather than on the assignment.

Red-first: with a 2.5 s mount delay standing in for a loaded runner, the
paint case failed on both engines before this and all 16 cases pass after.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say whose violations the preview rig reads

The list is the main frame's: securitypolicyviolation does not cross into a
frame, so an empty one says the embedder raised none and says nothing about
the artifact's own style, image or font. A listener inside the frame cannot
be the fix -- the fence under test is that nothing in the artifact runs.

So the comment now claims what the reading supports, and names where the
frame's containment is actually measured: the pixel for its inline style,
the counting server for its img-src and font-src.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): announce which side of the preview toggle is showing

The Preview/Source pair carried a label each and nothing else, so which one
was showing lived only in the active background -- invisible to a screen
reader on both surfaces. Each button is now a tab carrying its selected
state, inside a tablist, and the two files' toolbars stay character-identical
so the page and the phone announce the same thing.

Red-first: the new case renders both siblings and failed on both for the
missing role before this.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): type the WebView mock like the file's other hosts

The anti-slop gate refuses a bare `object` parameter. Takes the same shape as
the react-native mocks beside it, which pass it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): allow only the load the shell itself started

The document URL was allowed whenever the host reported no gesture, so a
navigation the shell never asked for could reload the page out from under the
session. Measured against a real WKWebView off-device: a sandboxed subframe
navigating the top frame to the document URL arrives as `.other` with no
gesture at all, and Chromium's own docs allow hasGesture() to be false for a
request a human started. Census first: nothing in the page navigates the top
frame -- no location assignment, reload, replace, window.open or form -- the
router moves by pushState and replaceState only, so the rule needs no gesture
and no page cooperation.

Both shells now raise a flag around their own load and drop it at commit, and
allow a main-frame navigation only while it is up. Everything else naming the
document is refused and never offered, since offering it would send the user
out of the app. iOS carries the second discriminator the same probe measured:
sourceFrame is the main frame for the shell's own load and the subframe for a
subframe's top navigation, so a subframe can never take the allow path.

Red-first: the Swift checks and the Kotlin tests were written first and failed
to compile against the old signature. 9 Kotlin tests, 54 in the module.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): point the meta-refresh arm at the embedder's own URL

The fixture pointed off-origin, so its own-origin assertion could not move
whatever the frame did. The new arm refreshes to `/`, which resolves against
the embedder's base, and pins zero top-frame requests on a counter the
`href="/"` case proves reads 1 in the same rig.

It also counts what the frame asks for itself, with a presence control that
attributes the fence: with `allow-same-origin` and no policy the same fixture
navigates the frame to the embedder's `/`, and with the policy dropped but the
product's token kept it navigates nothing, so the opaque origin is what
refuses it rather than the CSP.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read what an action produced, not what a clock allowed

The 600 ms after every action is gone. An arm that expects a navigation now
returns the moment the route handler records it, with a deadline only so a
click that missed its target says so instead of spending the case's timeout.
An arm that expects none waits for two painted frames inside the page and one
200 ms drain for the popup queue, which is a browser-process event with no
in-page counterpart; the docstring says why that one is bounded.

Measured and reported rather than claimed: with the new wait replaced by a
no-op every arm still passes, because the reads that follow are each a round
trip. It is insurance against the runner load that produced the frame-commit
race, not a fix for a failure seen here.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): take the settling branch as a ternary

What oxlint's prefer-ternary asks for, and the changed-code gate with it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): find the preview frame by its element, not its URL

CI timed out on all seven preview cases on one engine: the poll waited for a
child frame whose URL reads about:srcdoc, and that browser reports an empty
URL for a srcdoc frame, so every case ran to its own timeout. The same
difference had already shown as `expected '' to be 'about:srcdoc'`.

The frame is now the element: waitForSelector('iframe') then contentFrame(),
with readiness taken from the fixture's own marker inside it. Nothing compares
a frame URL any more -- the paint case reads the element's srcdoc attribute
and the absence of src instead, which is what "parsed inside the frame rather
than fetched into it" actually means. The one arm whose artifact navigates the
frame away says so rather than waiting for a marker that is not coming.

Red-first: with the old poll keyed on a URL the browser never reports, both
engines time out exactly as CI did; the new wait passes 18/18 with the 2.5 s
mount delay still injected.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): make a frame that never becomes ready say what it saw

The runner's Chrome read the preview frame's URL as empty where three
chromium builds here read about:srcdoc: bundled headless, the headless shell,
and --headless=old, all 147. So the difference is not reproducible locally and
the next CI run has to carry its own diagnosis.

The marker wait is bounded well inside the case timeout, and on expiry it
reports the frame's URL, the srcdoc attribute's length and the page's CSP
violation list -- which separates a frame the policy refused from one that was
merely slow, the two readings that look identical from a timeout.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): run the containment arms the comment only claimed

The comment said the fixture navigates nothing with the policy dropped and
the product's token kept, but no arm ran it: the control dropped both fences
at once. Both single-fence arms exist now, either of which would hold.

Measured rather than assumed, and one of them is not what the comment said.
The token alone: the navigation never starts, no request, no violation. The
policy alone, with allow-same-origin granted: the navigation does start and
frame-src refuses it, which the embedder reports as its own violation. The
engines differ only in what is left in the frame -- chromium an error page,
WebKit the artifact -- so neither is asserted; what is asserted is that the
request never reaches the server.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): refuse a download that names the shell's own document

The document branch skipped downloads, so `<a href="/" download>` fell
through to the offer path carrying the shell's own URL. Harmless in practice,
because the opener's scheme list drops it, but it contradicted the policy's
own comment and the prop doc, and it left the one URL that must never be
offered reaching the boundary.

The branch now covers a download too: refused, from either frame, gesture or
not, and never offered. A gesture-started download of anything else still
reaches the opener.

Red-first on both platforms: the Swift checks exited 133 and the Kotlin row
failed against the old policy. 10 navigation tests, 55 in the module.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): drop the own-load flag wherever a document ends

The flag lived beside the load call and had to remember every ending
separately, so iOS missed two: a prop update that fails before it loads, and a
renderer that died. Both left it raised, and a navigation to the document URL
during that window would have been allowed.

It now lives in the load state machine, which every ending already runs
through -- a commit, a failure, a dead renderer, a prop update, a reset -- on
both platforms, so there is nothing left to remember. The view raises it and
reads it, and drops it nowhere.

The Android residual is stated in the policy rather than papered over: between
loadUrl raising the flag and onPageStarted dropping it, a navigation to the
document URL from inside the preview frame would be allowed, because that
callback says nothing about which frame asked and no host discriminator
exists. It needs a generation switch and a tap in that window; iOS closes the
same gap with sourceFrame.

Red-first: the new Swift row failed to compile and the Kotlin row with it.
12 load-state tests, 56 in the module.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): spend the own-load flag on the allow, not on the commit

The flag stayed raised from the load until didCommit, so a second main-frame
action naming the document inside that window was allowed too and replaced the
document. WebKit can decide a second action before the first one starts, so
the commit is too late to be what spends it.

The allow itself spends it now, before the decision goes back, and every
ending still drops it for a load that is allowed and never commits.

Red-first: the new check composes the machine with the policy -- the seam the
flag and the rule meet at -- and failed to compile against the old machine.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): stop raising an own-load flag Android never consults

WebViewClient's javadoc, verbatim: "This callback is not called for all page
navigations. In particular, this is not called for navigations which the app
initiated with loadUrl(): this callback would not serve a purpose in this
case, because the app already knows about the navigation."

So the flag guarded nothing on this platform and, while raised, was the one
thing that could have let a competing request to the document URL through.
The view passes isShellLoad = false always now, the machine drops the field it
had no raiser for, and the policy comment carries the quote. Nothing reaching
that callback is the shell's own load, so nothing naming the document is
allowed there at all -- which also closes the generation-switch window the
residual named, so that paragraph goes.

No red to show: this is a removal, and the behaviour it leaves is the refusal
the existing rows already pin. What a device proof must check is stated in the
policy instead: a WebView that did route its own load here would have it
refused and the load state would sit at loading. 55 tests in the module.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): settle every arm, not only the ones that tap

An arm with no action read its counters as soon as the frame's marker
appeared, so a zero-delay meta refresh could dispatch after the reading. The
arms that pin zero were the ones relying on it.

Every arm settles now, and what it settles on is what it expects: the sealed
refresh arms take the bounded no-navigation path, and the loose arm waits for
a recorded navigation that is neither main-frame nor foreign -- its own
frame's -- rather than the main-frame wait it would never satisfy.

Red-first: with the settling removed and the refresh moved to 2 s, the loose
arm reads 0 on both engines; with it back, 1 on both, the delay still in.
A 0.4 s refresh passes either way, which is why the finding was invisible.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): wait for what the artifact's script wrote, not for the element

The two-fences control asserts the inline script ran, and the marker element
it waited for exists from parse time, so the arm could read window.__ran
before the script had touched it. Under a loaded runner that reads 0, which is
CI's "expected +0 to be 1" on chromium.

Readiness is now per-arm: 'script' waits for the script's own write, 'load'
for the arm whose artifact navigates the frame away, 'artifact' for the rest.

Red-first: with the inline script's write delayed 1.5 s, the old arm fails on
both engines with that exact message and the new one passes, delay still in.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): bound the rig's waits by the case timeout and nothing else

Two inner deadlines, 20 s and 15 s, were racing the outer one they sit
inside, so a slow runner could fail a case on a number this file picked
rather than on the one the case declares.

Both now run to vitest's own `ctx.signal`, which aborts when the case times
out. On abort the rig prints its reading -- the frame's URL, the srcdoc
length, the violation list, or the navigations it did record -- and lets the
case fail as the timeout it is. Nothing is rethrown from that path: a
rejection raised after vitest has given up on a case has nobody left to catch
it, and an unhandled one fails a run whose every test passed.

Red-first: with the marker selector pointed at an element that never appears
and the case timeout cut to 8 s, the diagnostic prints and the case fails as
`Test timed out in 8000ms` rather than hanging in silence.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): ask a stuck preview frame everything it can still answer

The old diagnostic said only that a frame never parsed, and its violation
list was the top document's -- securitypolicyviolation does not cross frames,
so it said nothing about what the frame itself refused.

It now prints the browser version, the arm it came from, the iframe element's
srcdoc length and sandbox, contentDocument.readyState and contentWindow.href
(which answer for a same-origin arm and report `refused` for an opaque one,
so the arm's own origin is in the log), and every Playwright frame with its
url, name, readyState, body length, marker presence, window.__ran and its own
violations. Per frame, because the page's init script installs the collector
in every frame -- measured on both engines -- and CDP evaluates inside an
opaque frame whose scripts are blocked.

Two corrections that the local probes forced. The reading is sampled while
waiting and printed from the last sample: read at the abort it lost its race
with vitest's teardown and printed nothing at all. And two arms had never been
given the case's signal, so their waits could not be bounded or diagnosed.

The diagnosis moves to its own module because the test file is at its line
limit, and because the bound and the reading it prints are one thing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): build a widened control frame instead of relaxing a live one

A live frame cannot be relaxed. Sandbox flags are fixed on a browsing context
when it is created, and Chrome 152 keeps the original ones through a srcdoc
reassignment while still parsing the new document -- so the control arms that
widened the product's own frame stayed sealed on the runner, and CI read a
script that never ran and a refresh that never navigated. Chromium 147 here
honours the relaxation, which is why it passed locally for a year of runs.

The override now clones the element, sets the sandbox on the clone, gives it
the artifact and replaces the product's frame with it, so the widened flags
are there from creation -- the way the product does it, since React sets the
attribute before insertion and never after. The product's own arms are
untouched: a null override still returns immediately.

And the control can no longer pass for the wrong reason on any engine. The
header-keeping arm now reads the violation raised inside the frame: a
script-src refusal can only happen if the sandbox let the script start, so it
separates "the policy held" from "the frame was never widened", which the old
arm could not. The loose arm pins an empty list beside it, the sealed arm pins
an empty one too, and those three readings are the whole fence story. The
violations come from each frame's own collector, because the embedder never
sees them.

Two diagnostic repairs the local probes forced: the browser version is read
once at open, since asking at the abort printed "browser unknown" in the CI
log this exists for, and the reading is sampled immediately as well as every
five seconds, since a wait that only prints "no reading was taken" says
nothing.

Red-first: with the widening disabled, both engines fail exactly as CI did --
180 s timeouts on the script arm -- and the diagnostic names the arm, the
version and the sandbox it actually had.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): put the toggle's selected state where a browser reads it

CodeRabbit is right, and the browser says so: react-native-web's createDOMProps
never reads accessibilityState, so on the page the tab pair emitted role="tab"
and no aria-selected at all. The test renderer could not see it, because it
reports the props the component was handed rather than the DOM they become.

Both siblings carry aria-selected beside accessibilityState now -- the phone's
screen reader takes the latter, the browser the former -- and the toolbars stay
character-identical.

Red-first, in a real browser on both engines: the rig now reads every
[role="tab"] element's aria-selected before and after the tap, and it read null
for both positions before this line existed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 03:32:34 -04:00
Jinwoo Hong e9b180685b feat(mobile): render Mermaid diagrams on the page from one deferred engine artifact (OTA phase C, C7.10 B) (#21871)
* test(mobile): measure mermaid rendered in the page

Red-first for C7.10 item B. The check mounts the real web sibling in
chromium and webkit under the shipped shell CSP and asks four things of
it: that a diagram renders with zero policy violations and zero eval /
new Function calls, that the SVG is the native buildHtml's own output
once the diagram id and xmlns:xlink are normalised away, that a hostile
diagram lands inert, and that a source change, an unmount and a remount
leave exactly one SVG and no listener of the first mount.

The equality oracle is buildHtml itself, bundled for Node behind a
Proxy stub for its native imports and served as its own document in the
same browser, so neither side of the comparison is retyped.

All eight cases fail on this commit: the sibling is still the labelled
source box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): fence the session download rather than its module list

Ruling 28. mobileWebAppRouteClosure reads metafile.inputs, which holds
dynamically imported modules under splitting: true exactly as it does
under splitting: false, so it cannot say "on demand" about anything: an
on-demand mermaid moves the session route's module list 4320 -> 6362
while its download does not move at all.

So the fence moves to entryStaticClosure. The new helper walks the
emitted chunks from the output the route's own module landed in and
follows import-statement edges only, and hands back both halves, because
mermaid's absence from the download is only a measurement while its 66
files are present in the deferred half.

The module list's new total is recorded in the docstring with its reason
and asserted beside the engine's own file count, which moves only when
the pinned mermaid version does.

Red on this commit: no mermaid in the closure yet.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): render mermaid in the page

The web sibling stops being a source box. mermaid is a browser library,
so the page imports it inside the render effect and draws the diagram in
this document: no WebView, no 3.7 MB engine string, and nothing of the
engine downloaded by a session with no diagram on it.

What replaces the sandbox is mermaid's own securityLevel: 'strict',
which runs its serialized SVG through DOMPurify. The native path's
</script> escaping has no analogue here and needs none, because the
source is a JS string argument rather than text spliced into an inline
script. Measured in both engines: a script in a label, a </script>, an
onerror and a javascript: click all land inert.

The configuration is now one object both hosts read, so the theme cannot
drift between the page and the phone; buildHtml serializes it instead of
holding a second copy. It gains suppressErrorRendering, because mermaid
otherwise draws its own error diagram into a temporary element and leaves
that element behind when it rethrows -- an orphan SVG on the page, and on
native a diagram the component is about to replace with the source box
anyway.

The dispose clears the host on unmount and on a source change; the id is
a useId, because mermaid writes it into the stylesheet inside the SVG and
it has to be a CSS identifier.

Also re-records the closure total the previous commit pinned: with the
real component the session route's module list is 6376, not the design
probe's 6362, and the reason is in that file's docstring.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): budget the deferred engine's chunks apart from the routes

Putting mermaid on the page took the app bundle from 69 emitted scripts
to 172, and the asset budget failed: 215 assets against a ceiling of
115. The cause is not a page split running away, which is what that
ceiling is for -- it is that mermaid lazily imports each of its own
diagram types, so one import() lands 103 scripts no route count
predicts.

So the ceiling gains a second term, named and measured (172 scripts with
mermaid against 69 with it aliased to a stub, at 11.17.2), rather than
the route term being raised to cover it. A page split running away still
fails on the route term, and the failure still says which of the two
grew.

The consequence is worth reading twice: the derived ceiling has to stay
inside the 256 assets the shell will load, and with 42 images it now
crosses that at 24 routes instead of 50. The bundle is at 215 today with
14 routes, so there is room for about ten more routes before a green
build produces a manifest no phone will open.

Measured by the config/scripts suite failing on this head, not predicted.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): pre-bundle the page's mermaid into one artifact

import('mermaid') from inside the app bundle emitted 103 scripts, not
one: mermaid lazily imports each of its own diagram types and esbuild
splits along those boundaries. Every one of those scripts sits inside
the OTA generation the phone has already downloaded, so the split moved
no bytes over the wire and spent 103 of the 256 manifest assets the
shell will load -- which is the scarce resource here, and the reason the
previous commit had to invent a second ceiling term.

So a sibling generator bundles the package into one ESM module beside
the WebView engine it already builds, emitted by the same postinstall
run, gitignored and lint-ignored with the others. The page imports that
artifact on demand instead, through a loader whose return type names the
two calls the component makes -- checked against the artifact's own
inferred export rather than cast to it.

Measured, at 14 routes:

  emitted scripts   172 -> 69   (68 with no deferred engine at all)
  manifest assets   215 -> 112  (111 with none)
  session modules  6376 -> 4323 (+3 over main: config, loader, artifact)
  chunks fetched for one graph TD   27 -> 1
  bytes fetched      837,530 -> 3,482,965

The static-closure fence is unchanged in meaning and now reads on the
artifact: absent from every chunk the route reaches by an import
statement, present in the deferred half. The rendered SVG is byte-for-
byte what it was, so the equality against the native document still
holds on both engines.

Also adds the diagram to the webview-consumers list, which is what that
list means: its native component imports the package and its sibling is
what the builder resolves instead.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* revert(mobile): drop the deferred-engine ceiling term, keep the control

With the engine pre-bundled into one artifact the bundle emits 69 scripts
at 14 routes against the route term's 72, so the second term this series
added has nothing left to do and the route count is the only term again.
mobileWebAppBundleMaxChunks and the asset ceiling derived from it are
back to what main has; the shell's 256 assets are crossed at 50 routes
again rather than at 24.

What stays is why. A ceiling raised to admit 172 scripts would have
admitted any split at all, so the budget test gains the control that
holds the line: the single-artifact count passes the ceiling and the
lazily-chunked count fails it, both measured at 14 routes, with mermaid
named as what produced the second.

Red before the term came out: the control failed asserting 172 > 175.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): keep build output out of the raw-request-port census

The census walks mobile/src for AST reaches into the unvalidated request
port, and the pre-bundled mermaid artifact is the first generated file
under src that is executable code rather than a string literal. Two of
its own vendored dependencies contain the token `sendRequest`, so the
walk read minified third-party code as a new call site and asked for an
inventory line nobody can ever migrate.

So `*.generated.ts` joins node_modules and test files in that file's
stated list of what it does not scan, with the reason. The scripts that
emit those artifacts are ordinary source and are still scanned, which is
where a real reach would be.

Two halves to the new control, because a filter that skipped everything
would satisfy either alone: nothing generated is left in the scan, and
the matcher still finds the port when handed one line of code.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): escape the shared config into the native inline script

buildHtml spliced JSON.stringify(MERMAID_DIAGRAM_CONFIG) straight into
the inline <script>, twenty lines below the function that exists because
JSON.stringify leaves `<`, `>`, `&` and the U+2028/9 separators raw. Inert
at today's five hex colours, and not inert for a themeCSS or a font stack,
which is free text going into the same script element.

So the escaping splits from the stringify and both callers use it: the
source keeps its own wrapper, the config gets one. Those characters only
ever appear inside JSON string literals, so escaping them is valid for an
object serialization exactly as it is for a string.

Red first: a config carrying `</script><script>` put four raw closers in
the document where a benign build has two.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): pin the page mermaid type against the package's own

The loader returned the artifact's default as PageMermaid, which checked
that two names exist and nothing about their shapes: the artifact is
minified vendor output and both members infer as `any` there -- a probe
assigning engine.render to a number compiles -- and `any` satisfies every
signature there is.

So the shapes are asserted against the package's `Mermaid`, which is
precise. A PageMermaid member whose signature the engine does not really
have now fails at this line rather than at a call the page makes.

In the product module, not a test: mobile/tsconfig.json excludes test
files, so a type-only assertion in one is never compiled. Underscored
because it is a compile-time statement with no runtime reader, which is
the form the linter asks for.

Control, verified both ways: changing render to (id: number) => Promise<{
svg: number }> reds tsc naming both parameter and return, and the real
signatures compile.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): re-measure the chunk series and say what it does not show

The four-point series was stale and read as a slope it is not. Measured
again on this head, by copying the route tree and dropping routes from
the end of the sorted key list -- both siblings of each, because deleting
a .web.tsx alone leaves the native file for the builder to resolve and
measures an entirely different closure, which is how the first attempt
at this produced 77 scripts for 14 routes:

  8 routes  -> 32 scripts
  10 routes -> 43
  12 routes -> 61
  14 routes -> 69   (the real tree)

Between four and nine more per route depending on which route, so 4r + 16
is a bound and not a fit, and the justification now says that instead of
claiming three per route. It also says the part that matters more: at 14
routes the tree measures 69 against 72, and the last two routes cost the
8 the ceiling grants for two. The fence is at break-even, and the new
assertion states that slope from the function rather than from a comment.

Also records what the generation weighs, since every chunk ships in it
whether or not a phone fetches one: 8,016,714 bytes across 112 assets
against the 9 MiB ceiling, 84.9%, 1,420,470 left. It was 4,539,090 before
item B, and the engine is the difference.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the native fallback under suppressErrorRendering

The shared config reaches the phone too, and it gained a key the native
path did not have. So the native document is now loaded for a diagram
that throws, in both engines, with window.ReactNativeWebView standing in
for the host: mermaid's run still rethrows, the document's own catch
still posts `error`, and that is the message the component turns into the
source box.

Measured both ways, so the case says which half the key owns. Whether
the fallback fires does not depend on it -- `error` is posted with the
key and without it. What depends on it is that nothing is drawn behind
the fallback: removing the key leaves mermaid's own error diagram in the
document and reds this case at 1 SVG against 0, on chromium and webkit
alike.

The control is the same document for a diagram that parses: a height,
not `error`, and one SVG.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one walk for every source census, without build output

Nine censuses under mobile/src each held a copy of the same recursive
walk, and each decided for itself what a source file is: seven had no
opinion about generated files, one excluded them in its own regex, and
one had the exclusion I added last round. So all nine read 7.9 MB of
emitted vendor code -- 3.7 MB of mermaid for the WebView, 3.5 MB of it
for the page -- and the two largest censuses TypeScript-parsed all of it,
looking for call sites nobody wrote and nobody can move.

That is what took rpc-params-contract-type-only-boundary over its 5 s
timeout in CI once the fifth artifact arrived. Measured here, median of
3, import plus tests:

  main, 4 artifacts, no exclusion   1004 ms   (slowest case  831 ms)
  with the 5th, no exclusion        1513 ms   (slowest case 1358 ms)
  with the 5th, this commit          947 ms   (slowest case  788 ms)

So it lands below where main has it, not merely below where I left it.
Across the nine, four more halve: rpc-operation-cast-fence 769 -> 441,
rpc-subscription-boundary 946 -> 468, unchecked-rpc-reader-boundary
1042 -> 538, lifecycle-owner 747 -> 433, reanimated-web-mapper-deps
1028 -> 516. The two that already excluded generated files do not move.

What each census counts as interesting -- extensions, whether test files
are in -- stays its own, because they genuinely disagree. What counts as
a source file at all is now said once.

The control is the file that started it: a *.generated.ts whose text
holds exactly the import a census is hunting, planted beside an ordinary
file carrying the same text. The generated one is not returned and the
ordinary one is, so the absence is a measurement. A second control reads
mobile/.gitignore and holds the predicate to every artifact the tree
generates, and a third fences the walk itself to one spelling, so a tenth
census cannot paste the cost back in.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): correct why the type pin sits in the product module

The comment said a type-only pin in a test file "is never compiled".
That is false: mobile/tsconfig.json excludes *.test.ts, but
tsconfig.test.json is a second program that does check them, run by
check:tests-typecheck and held by the tests-typecheck ratchet.

The conclusion is unchanged and the reason is now the true one. The app's
own typecheck is the unconditional gate and would not cover a pin written
in a test; the test program is real but carries a grandfathered baseline
and a few files held outside it on purpose. And the assertion is about
this module's own type either way, so it belongs beside it.

Comment only; tsc, the ratchet and both lints re-run on the file.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 00:36:51 -04:00
Jinwoo Hong bd5177801b feat(mobile): put the page's pickers, paste and editor fallbacks on the media verbs (OTA phase C, C7.6) (#21795)
* feat(mobile): put media picking behind a platform seam (OTA phase C, C7.6)

The session screen picks images three ways — the photo library, Files, and the
pasteboard — and all three are native modules a page cannot import: the codegen
lookup `expo-image-picker` and `expo-document-picker` run at import throws in a
browser, and the route manifest imports every route, so one of them in a page
closure is the whole bundle down rather than one picker.

`src/platform/media-picker.ts` is the phone's, delegating to the same three
calls the screen already made. `.web.ts` is the page's: `native.media.pick`,
then `read` in order to `eof`, then `release` for every handle it was handed,
including the ones its caller never took — the shell holds eight staged files
at a time and an abandoned pick otherwise waits out the five-minute TTL. A
refusal rejects with the shell's code on it and is never folded into the empty
answer that means the user cancelled.

The bytes are concatenated decoded and encoded once, because the wire promises
`eof` and nothing about the length: a shell answering a range shorter than the
one asked for ends a chunk on a partial base64 group, and a reader joining the
strings would fold that padding into the middle of the file.

The census walks the session route module's own closure — the route is not
registered until C7.7 — and names any module that reaches a picker or
`Clipboard.getImageAsync` directly. Today that is the two modules C7.6's next
commit moves, listed by name so the list goes empty rather than the rule going
quiet.

Inert: nothing calls the seam yet.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): put the session's paste and attach on the media seam (OTA phase C, C7.6)

The terminal paste read the pasteboard through `expo-clipboard` directly and
the two attach paths called `pickMobileImage`/`pickMobileImages`, so the page's
closure carried `expo-image-picker` and `expo-document-picker` — native modules
whose import throws in a browser.

All three now go through the seam. Text is `native.clipboard.read` on the page,
which the clipboard seam gains a reader for: `expo-clipboard` resolves to
`navigator.clipboard` there, which needs a secure context the iOS shell's
custom scheme is not. An image is `pick { source: 'clipboard' }` rather than an
inline value, because a clipboard image is 24 MiB of base64 against an 8 MiB
reply ceiling.

The census over the session closure is empty now and asserts the seam is in it,
so a rule that found nothing is one that had something to find: with the three
call sites restored it names all three.

`mobile-image-source-picker.ts` stays the phone's implementation, reached only
through the seam's native sibling, and resolves out of the web closure entirely.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): resize a clipboard image on the page with a canvas (OTA phase C, C7.6)

The paste hook carried the raster shrink inline over `expo-image-manipulator`
and two `expo-file-system` writes. Both are native: the manipulator has no
browser build, and the temp file exists only to work around an iOS loader that
cannot decode a large base64 data URI, which a browser does not need.

Split into `mobile-clipboard-image-resize.ts`, unchanged, and a `.web.ts` that
decodes one `<img>` from the data URL the shell's `img-src 'self' data:`
already admits, draws it into a canvas at the target size and reads the PNG
back out of `toDataURL`. It reports the canvas's own size rather than the size
asked for, because a browser clamps a canvas past its area limit and the
downscale loop above would otherwise retry a raster that never shrank; and it
awaits `decode()` rather than `onload`, which never fires for a source the
browser cannot read and would leave the paste waiting on a promise nothing
settles.

Measured in Chromium under the shipped header, on a noise PNG because that is
what PNG compresses least: 1400x1000 encodes to 5,476,032 base64 characters and
converges in one pass to 368x263 and 397,220, which is 75.8% of the upload
path's 512 KiB chunk. Zero policy violations and zero page errors. Red under
three mutations: the source returned unchanged, a reported size the canvas did
not draw, and `onload` in place of `decode()`.

`computeMobileClipboardImageDownscale` moves to a leaf for the reason the
upload-chunk constant has one: the check wants the arithmetic and not the
upload path's RPC operations behind it.

The page closure now carries none of `expo-image-picker`,
`expo-document-picker`, `expo-image-manipulator` or `expo-file-system`, pinned
beside the seam census.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): give the two WebView editors their plain web fallbacks (OTA phase C, C7.6)

`MobileRichMarkdownEditor` and `MobileHtmlPreview` are the session closure's
other two `react-native-webview` consumers. On the web that package renders the
line "React Native WebView does not support this platform" where the surface
was, so nothing it was mounted for works and the closure pays for a module that
cannot do its job.

Ruling 8: each gets the plain state it already degrades to, and no second
renderer. The editor renders the Markdown source in one field on the text-input
seam, so the screen around it keeps the text, every edit through `onChange`,
and Save, Discard, Copy and Refresh; the degradation is the formatting toolbar,
whose fifteen commands are the rich document's. The preview renders its own
Source tab; the degradation is the rendered artifact, and the toggle goes with
it, because a control that can only be in one position is a control that lies.

Neither is smaller than a DOM renderer, which is why neither is one here. The
editor's toolbar would need a `contenteditable` implementation with its own
escaping, and the preview has no nested frame to sandbox agent-produced HTML in
at all — the shell's policy carries `frame-src 'none'` and `child-src 'none'`.

`dismissKeyboard` blurs the field rather than calling `Keyboard.dismiss`, which
is a stub on React Native Web; `onKeyboardInsetChange` is never called, because
it exists to correct for a WebView's covered area and on the page
`keyboard-occlusion.web.ts` is the only measurement there is.

The closure census names the one consumer left, `TerminalWebView.tsx`, which is
C7.5's: with both siblings removed it names all three.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): read a destructured clipboard alias in the media census (OTA phase C, C7.6)

The census recognised `Clipboard.getImageAsync` as a property access and
nothing else, so `const { getImageAsync } = Clipboard` reached the same
function without ever writing one and the closure was approved. On the page
that call is `navigator.clipboard`, which needs a secure context the iOS
shell's custom scheme is not, so the approval was for a path that dies at the
browser clipboard API.

Aliases are now resolved to a fixpoint — `const pasteboard = Clipboard` makes
`pasteboard` the module too, and the chain has no length limit — and a
destructuring off any of them is reported at its declaration, which is the line
to delete. The destructured name is read the way the import clause's is, off
`propertyName` when the element renames it, so `{ getImageAsync: readImage }`
is the same offence spelled differently. A binding element's `name` can be a
nested pattern and a `propertyName` can be computed, so the text is taken only
off a node that has one.

Red-first with each shape planted in the scratch tree before the rule moved:
the plain destructuring, the renamed one and the re-destructured chain were all
missed. Dropping the fixpoint afterwards loses the chain; reading the local
name instead of the property loses the rename. `{ getStringAsync } = Clipboard`
stays unreported, because text off the pasteboard is the clipboard seam's and
not this rule's.

The session closure is still empty under the widened rule.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): release every item a clipboard pick answered (OTA phase C, C7.6)

`readClipboardImage` destructured the first staged item and released only that
one, while `pickImage` already guards the same shape through `readPicked`.
`multiple: false` is what the page asks for and not what a shell promises, so a
caller taking the first of several would hold the rest against the
eight-handle cap until the five-minute TTL. Today's shell stages at most one on
the clipboard arm, so this is the seam's own docstring made true rather than a
leak in the field.

Red-first with two staged clipboard items: releasing only the one read leaves
`media-2` held, and the second is now returned without ever being read, which
is what the single-image pick does.

The refusal case is one path over both codes a pick can answer with: the
registry's `native_media_handle_cap`, raised before a picker runs, and ruling
6c's `native_media_too_large`, raised once a picked item has been weighed. A
code outside the seam's vocabulary floors to `native_verb_failed` rather than
crossing verbatim, which is what makes naming the exact code load-bearing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read the upload chunk from its module in the resize check (OTA phase C, C7.6)

The canvas resize check restated `512 * 1024` as the budget it holds a run to.
A check carrying its own copy of a product constant is one that goes on passing
after the upload path's chunk has moved, which is the reason the harness reads
the CSP, the protocol version and the window caps out of their own sources.

`readClipboardImageUploadChunkBase64Chars` joins them, evaluating the product
the way the window caps reader does. Proved live by moving the constant: at
64 MiB the run reds on the fixture no longer being over the budget, and it is
back to 512 KiB here.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): read element access in the media census and correct two claims (OTA phase C, C7.6)

Round 2, four lows.

The census read `Clipboard.getImageAsync` and not `Clipboard['getImageAsync']`,
which is the same call, the spelling a bundler produces, and the one a reader
reaches for to get around a rule about dots. Element access with a string
literal is now read the same way; a computed key is not, because its value is
not in the source and guessing would report a line nobody can act on. The
closure test cannot back this up — `expo-clipboard` legitimately sits in the
session closure — so the scratch fixture is the whole of the evidence, and it
reds with the arm removed.

The fixture also could not tell the alias fixpoint from one source-order pass:
every planted chain happened to be declared in the order a single walk learns
it. `reverse-order-alias.ts` is declared back to front, and is valid at run
time because the destructure sits inside a function the module body finishes
before anything calls. Bounding the loop to one pass now reds it.

The canvas resize justified reading its size back off the element by a browser
clamping past its area limit. That is not what browsers do: the width attribute
reflects whatever it was assigned, so the returned size is always the target.
The real reason is narrower and is now what the comment and the override entry
say — the dimensions and the bytes come from one element, so a caller's
bookkeeping cannot describe a raster that was not encoded.

The override entry also carried a stray apostrophe in `img-src 'self' data:`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): say what the clipboard contract shipped as, and read a backticked key (OTA phase C, C7.6)

Round 3, three lows.

The merge took main's `clipboard.ts` byte for byte, so its reader docstring
still described the state C7.2 shipped: one verb on the web, and a page whose
lack of an image verb degraded into the old path. The design that shipped is
the other one — this seam owns the pasteboard on both platforms and the page's
`readImage` runs `native.media.pick { source: 'clipboard' }` with the chunked
read behind it. The prose now says that, and says that null still means an
empty pasteboard while every other outcome rejects.

The same merge left `clipboard` twice in the paste hook's dependency list, one
from each side. Deduped.

The census read a quoted element-access key and not a backticked one, so
``Clipboard[`getImageAsync`]`` escaped a rule that catches both other
spellings. A template with no substitution is a string literal with a different
quote, and reading only one of the two leaves the other as the way around.

The computed-key plant could not see the literal-kind check at all: its
variable was named `key`, so reading the identifier's text found nothing
either way. It is now named after the method and holds a different one, which
makes dropping the kind check a false positive on a call that reads text.

Red-first: the backticked access planted before the rule moved is missed;
ignoring template keys afterwards misses it again; accepting any key node
reports the computed plant.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold canPickMedia to all three verbs and seed the census from import() (OTA phase C, C7.6)

Two bot findings.

`canPickMedia` answered true on `native.media.pick` and `native.media.read`
alone, but every image read releases what it picked. On a route without
`native.media.release` the release rejects, the cleanup swallows it by design,
and the staged file stays live to the five-minute TTL: eight pastes and the
next pick is refused at the handle cap, with nothing on screen to say why. A
route missing one verb has no working image path, so `contents()` now says so
up front rather than after four of them. Red-first: a route granted pick and
read but not release answered `image: true`.

The census seeded its aliases from static import and export declarations only,
so `const Clipboard = await import('expo-clipboard')` produced no offender —
while the bundler resolves a literal dynamic import into the closure exactly as
a static one. A dynamic import is now read wherever it appears: `await` and
parentheses unwrapped, the assigned identifier seeded as an alias, a
destructuring off one reported at its declaration, and a picker module reported
at the call, since reaching one at all is the offence. A specifier that is not
a literal is left alone, for the reason a computed key is.

Red-first with all three forms planted and the seeding removed: the namespace
alias, the destructuring and the picker import are each missed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): seed the media census from a backticked import() too (OTA phase C, C7.6 bots)

CodeRabbit: `import(`expo-image-picker`)` is as static to the bundler as the
quoted form, but the census read only a string literal specifier, so a
backticked one joined the closure unseen. A no-substitution template literal
now seeds it the same way; the planted fixture is reported at its line and
was unreported before the arm.

pullfrog: the clipboard seam's docstring counted the web read as two verbs
where its web sibling counts one for text and three for an image. It now
counts the same way in both files.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-20 11:24:53 -04:00
Jinwoo Hong e736a9f29b feat(mobile): take the session screen's inputs, links, clipboard and routers through the platform seams (OTA phase C, C7.2) (#21790)
* feat(mobile): put the session screen's nine text inputs on the web font seam (OTA phase C, C7.2)

The landed text-input census, run over `app/h/[hostId]/session/[worktreeId].tsx`,
reports nine sizes that do not come from `TEXT_INPUT_FONT_SIZE`. Six declare the
app's body size and move in place, which is the same number natively. Three do
not — a 22px key-capture field and the chat's two 15px fields — so each gets a
`.web.ts` sibling of the address bar's shape, with a shared base so the two
halves can differ in nothing but the size.

The capture field is the one the move shrinks rather than raises: 22 already
clears the focus-zoom floor, and the census reads the seam as a binding rather
than as a number, so there is no expression that keeps 22 and still says where
the size came from.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): drop the theme import the composer's style split left behind

`oxlint` over the whole tree, which CI runs, reads it as an error.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): open the session screen's three external URLs through the platform seam (OTA phase C, C7.2)

The landed external-link census, run over the session route's closure, reports
three modules reaching react-native's `Linking`: a terminal link tap whose open
mode is the phone's browser, and the two WebView-backed readers, each of which
sends a tapped link to the system browser rather than navigating the artifact
away.

Inside the shell `Linking.openURL` calls `window.open`, which both shells refuse
and which resolves either way, so all three reported success into a tap that did
nothing. The seam also stops swallowing the failure: each site caught and
discarded, and `openExternalLink` names it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): take the session screen's seven clipboard sites through the platform seam (OTA phase C, C7.2)

`expo-clipboard` resolves to `navigator.clipboard` on the web, which needs a
secure context — iOS serves the page from a custom scheme and Android from
`https`, so that path works on one platform and silently not on the other. The
landed census now reports the module out of the route's closure entirely.

The seam grows its reader half, on the landed `native.clipboard.read` verb: text,
a PNG, and a presence probe. Two degradations are recorded rather than implied.
No shell serves an image, so the page answers null and the terminal's paste takes
the branch an empty clipboard already took; and the shell serves no presence verb,
so `contents` answers what this side knows rather than reading to find out, which
would raise iOS's paste-consent prompt on every foreground.

The copy-path sheet gains the failure toast its two neighbours already had: it
showed "Path copied" before the write, and the seam rejects rather than returning
false.

The route parity pin moves with it: five clipboard hooks join the expanded route
and one runtime string joins the sheet.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): take the session domain's three routers through the handoff seam (OTA phase C, C7.2)

Inside the page a screen is one document standing in for one screen, and
`useRouteHandoff` is the only thing that knows which targets the page keeps and
which it hands back to the app. The three holders here are the workspace-missing
bounce, the file-tap preview push, and the pane-tap param consume.

The domain's census is narrower than the two landed ones because it has to be:
eight of its hooks take `useFocusEffect` and two take `useLocalSearchParams`,
neither of which can navigate, so the rule is a closed list of names rather than
a ban on any value import — which also catches expo-router's module-singleton
`router`, a spelling a `useRouter` rule would have read as clean.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): list the two style siblings the session screen's inputs added (OTA phase C, C7.2)

The overrides census fails on an unlisted `.web.*`. One raises the chat's two
15px fields past the focus-zoom floor; the other lowers a 22px capture field onto
the seam, and its entry says why a reduction is the right answer there.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): let the text-input census read a literal already clear of the floor (OTA phase C, C7.2, ruling 12)

The floor is the rule and the seam is the mechanism. A binding rule alone made
the custom-key capture field an offender at 22, where nothing can zoom, and the
only way to satisfy it was to lower a one-character field to 16 — the tail
wagging the dog.

The seam's web half now exports the floor it already computed `Math.max` against,
and the census reads that number out of that file rather than carrying a second
copy of 16. The rule becomes "the seam's binding, or a literal at or above the
floor", with no per-site exemption: a literal under the floor is still reported,
which is the case the seam exists for. A tree whose seam declares no floor is
refused rather than judged against a number the census invented.

So the capture field goes back to 22 on both platforms and its split, its
override entry and its parity test go with it. The chat's two fields stay split,
because 15 is under the floor however it is spelled.

Red-first: with the rule removed, a planted literal 16 and a literal 22 are both
reported and the refusal case does not throw; a literal 15 is reported either
way. All three route closures that run this census — session, source-control,
review — report 0 offenders and 0 unresolved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): answer clipboard reads on the read grant and catch refused writes

`contents()` reported no text on a route that granted `native.clipboard.read`
without the write, because it read `verbs.granted`, which is write AND read. The
verbs hook now exposes the two grants separately and the web seam answers on the
read one; `granted` keeps its meaning for the callers that need both.

The Markdown copy action was the one write of eight in the session domain with
nowhere for a rejection to go: the seam rejects when the pasteboard refused the
text, the callback had no failure branch, and its caller drops the promise, so a
refused write raised an unhandled rejection and still left "Copied" on screen. It
now takes the error haptic and the "Couldn't copy" toast the other copy paths
show. A census over `src/session` fails if any `writeText` call site lacks a
failure branch, so the ninth site cannot arrive without one.

The route parity pin moves with it: one callback body, one runtime string. Its
refresh note claimed six clipboard hook sites for a delta of five; the walk from
`SessionScreen` reaches five, and the terminal's paste is not among them.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): probe both clipboard kinds at once and judge sizes against the floor

Moving the two clipboard probes into an object literal serialised them: the
migrated `contents()` awaited `hasStringAsync` before `hasImageAsync` was called,
where both callers had used `Promise.all`. That path runs on mount, on every
AppState foreground and on every select-mode toggle. Restored, with an ordering
probe that deadlocks unless both probes start before either answers.

The floor case could not fail for the reason it named: its fixture declared 16,
so a census carrying its own copy of 16 passed it. It now plants a seam declaring
20 and a literal 18, the size that is clean under one floor and an offence under
the other.

Two stale wordings from the reverted split: one closure case still said "both
split style modules" over a one-element list.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): buzz the quick-command row when a copy is refused

The last of the seven migrated writes without the error haptic. The row already
said "Couldn't copy" on its own control, in red, for the 1500 ms the toast the
other six show would have lasted, so it never claimed a refused write had landed;
what it had no way to say was anything the thumb still on the button could feel.

Its first test, on the harness its list already uses: the seam rejects when the
pasteboard refuses, and the two cases are the difference between the row that
shows a green check over nothing copied and the row that does not. The list's own
test gains the haptics mock the row's new import needs.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make the clipboard census require the await its rule depends on

`hasFailureBranch` accepted any enclosing `try` with a `catch`, so the one shape
the census exists to stop passed it: `void clipboard.writeText(...)` inside a
try/catch is an unhandled rejection with a handler three lines above it that can
never run, because the block returns before the promise settles. It now requires
the call to be awaited inside the try's own block, or to carry a `.catch` along
its own chain. The boundary walk stopped only at function and method
declarations, so a `catch` outside an arrow answered for the call left running
inside it; every function-like node ends the search now.

Five cases over snippets read through the same reader, because a `void` write
would have to be committed to be tested against the real tree. Control on a real
site: making the Markdown write un-awaited inside its own try reports it.

Two provenance fixes. The runtime-string delta across C7.2 is two literals, not
one: "Couldn't copy path" took the count from 532 to 533 and "Couldn't copy" took
it to 534. And main's C7.4 made `BRIDGE_CLIPBOARD_MIMES` `['text']`, so an image
mime is a value the schema does not admit rather than a refusal the verb spells
out, with `native.media.pick { source: 'clipboard' }` waiting on C7.6; the web
seam and its test said otherwise. Behaviour unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): let an unmounted quick-command row emit nothing on a refusal

The haptic I added ran before the mounted guard, so a copy pressed on a row that
then scrolled out of the list, or a sheet closed over it, still buzzed when the
rejection arrived. A buzz with no row to explain it is feedback for nothing, and
the guard was already there for the feedback state one line below.

Red-first: press, unmount, then reject. The success path already guarded first.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make the clipboard census require a catch with something in it

Any `.catch` property access counted as a failure branch, so two shapes that
handle nothing passed: `clipboard.writeText(text).catch` reads the handler's name
and registers nothing, and `.catch()` swallows the rejection while the caller goes
on to say the write landed. The rule now requires `.catch` to be the callee of a
call carrying at least one argument.

Red-first with both shapes in the snippet reader, the accepting cases unchanged.
Control on the real tree: emptying the notes sheet's handler reports
`MobileSessionSheets.tsx:174`, and restoring it greens.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-20 09:49:59 -04:00
Jinwoo HongandClaude ee61e3bd41 fix(mobile): measure the keyboard from visualViewport inside the page (OTA phase C, C4.2) (#21735)
* feat(mobile): measure the keyboard from visualViewport inside the page (OTA phase C, C4.2)

react-native-web's `Keyboard` is a stub: `addListener` returns a
subscription that never fires and `isVisible()` is always false. A screen
inside the shell's page that waits for `keyboardDidShow` waits for the
life of the document, and the software keyboard covers whatever sits at
the bottom of it. Two C4 screens are text entry at the bottom.

`platform/keyboard-occlusion` is the pair. The native file carries the
source-control hook's logic unchanged, events and clamp and the comment
that travels with it. The web sibling reads `visualViewport`: the layout
viewport keeps its size and the visual one shrinks, so the occluded strip
is `innerHeight - (height + offsetTop)`. `offsetTop` is in it because a
scrolled or pinched visual viewport sits partway down the layout viewport
and the strip below it is not keyboard; dropping the term reds two cases.
It listens on `resize` and `scroll` — the browser scrolling a focused
input into view moves the offset without resizing anything — and reads
once at mount, because a composer opened over an already-raised keyboard
receives no event at all; dropping that read reds a third case.

`useKeyboardAvoidingPadding` is a second name rather than a `Platform.OS`
branch at the call site. Natively it is 0 and subscribes to nothing, so a
composer that asks for it renders exactly as often as it does today;
`KeyboardAvoidingView` has already moved it and padding would move it
twice. On the web it is the whole of the avoidance, that view being
driven by the events this file exists because the page never receives.

No `visualViewport` answers 0 rather than guessing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): lift the commit bar and the note composer inside the page (OTA phase C, C4.2)

The two consumers move onto the seam. The hub's hook becomes one line and
keeps its name, which is what the hub's state calls the number. The note
composer takes the padding as a style on the `KeyboardAvoidingView` it
already had: natively that is 0, so the prop is `undefined` and the phone
renders exactly what it rendered before; inside the page it is the strip
the keyboard covers, which is the only thing that moves the composer
there.

The census is over both future route closures rather than over the two
call sites: `platform/keyboard-occlusion` is the one module in either
closure allowed to name the stub. Red first at the base commit — run in a
throwaway worktree at `9309350864` rather than by setting the fix aside —
it named `use-mobile-source-control-keyboard-lift.ts` as a subscriber
outside the seam and found the seam's web file in neither closure.

`mounted-bottom-drawer.tsx` is exempt by name, and the census asserts the
exemption is really in both closures so it cannot outlive its subject. It
reads more than a height — `Keyboard.metrics()` for a sheet opened over a
raised keyboard, and each event's `duration` to animate with it — which
the seam does not model, and it sits in C1's, C2's, C3's and C5's closures
too, so moving it is a change to every page rather than to this domain.
Its listeners are inert on the web the same way, which is why the composer
inside it takes its own padding rather than inheriting one.

No render-check case: measured, none of the five registered routes reaches
the seam, the commit bar or the composer, and a headless browser cannot
shrink the visual viewport independently of the layout one anyway. C4.4
carries it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): type the keyboard harness instead of asserting its fields (OTA phase C, C4.2)

The changed-code gate flagged the two `as` casts in the hoisted harness.
A return type on the `vi.hoisted` callback says the same thing and is
checked rather than asserted, which is the shape the host-list route test
already uses.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): read a pinch zoom as no keyboard, and test the clamp (OTA phase C, C4.2 round 1)

Round-1 folds plus CodeRabbit's exemption point.

**A pinch zoom read as a keyboard.** A 2x zoom shrinks the visual viewport
by exactly as much as a half-screen keyboard, so the commit bar and the
composer moved on a page nobody was typing into. A `scale` other than 1
answers 0. Geometry alone cannot tell the two apart and a stored "no
keyboard" baseline would be a heuristic, so a keyboard raised while zoomed
is the accepted rare case rather than a guess. `scale` is read defensively
because older WebViews do not implement it, and taking its absence for
zoomed would answer 0 for every keyboard on them; mutating the guard to
key on absence reds both cases.

**The clamp had no test.** A bare subtraction left all nine cases green.
The case is a visual viewport taller than the layout one, which mobile
Safari reports mid-scroll and which would have pushed the commit bar down
the screen instead of up.

**One guard, where the test reaches it.** `occlusion`'s `viewport ===
undefined` arm was unreachable: the effect returns before calling it, and
the absence case exercised that one. Deleted, and the remaining case says
which guard it proves.

**The census exempts two files, not a directory.** `startsWith('src/platform/')`
would wave through a later `src/platform/*.web.ts` that subscribed to the
stub directly, which is the defect this census exists for. Named exactly,
with a planted subscriber beside the seam as the fixture; restoring the
directory filter reds it.

**And the moved comment claimed an inset it never subtracted.** Deleted.
Correcting a comment that was false where it came from is not a rewrite of
the logic the move carried: no statement moved with it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the page at scale 1 so the zoom guard is not the keyboard path (OTA phase C, C4.2 round 2)

Round 2's finding changes what the zoom guard costs. iOS auto-zooms on
focus of any input under 16px; both consumers' inputs are 14px
(`typography.bodySize`), and the page's viewport meta set no
`maximum-scale`. So `scale !== 1` was not the rare pinch the guard was
written for, it was every focus — and the seam would have answered 0 on
the one flow it exists for.

The guard stays and the premise is fixed instead: `maximum-scale=1` in
both places the page's meta is written, the built document in
`build-mobile-web-app-bundle.mjs` and the bootstrap `index.html`. iOS
honours it for the focus auto-zoom and has ignored `user-scalable=no`
since 10, so a deliberate pinch still works; the input sizes are
untouched. C4.6 step i is what settles it on a device.

Three test changes and one correction.

The census took a `rootDir`, as `findWebSiblings` does: it planted
`src/platform/other.web.ts` in the real tree while the overrides census
walks `mobile/src` in a parallel worker and would read it as an unlisted
override. It plants under `mkdtemp` now, and writes the two seam files
there too, so the empty result for them is the name exemption working
rather than those files happening not to subscribe.

A case for the ruling itself: scale 2 with a viewport shrunk past what
the zoom explains answers 0. Dropping the guard reds it and the pinch
case together.

`useKeyboardAvoidingPadding` is rendered through the test renderer now
instead of called outside one, with a counter on `Keyboard.addListener`.
Making the native hook return `useKeyboardOcclusion()` reds it at two
calls; the old shape could not see that, because a hook read outside a
component never runs its effects.

Item 4 did not hold as written. `window.visualViewport ?? undefined` is
not a no-op: the DOM declares the property `VisualViewport | null` and an
older WebView omits it entirely, so the coalesce was normalising both
shapes into one `=== undefined` check. Removing it and testing only for
`null` throws on the absent-viewport case (reproduced: `Cannot read
properties of undefined (reading 'scale')`). The coalesce is gone and the
guard names both shapes instead.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): raise the two page inputs to 16px on web instead of pinning the page scale (OTA phase C, C4.2 round 2)

`maximum-scale=1` is reverted from both metas. It fixed the right problem
in the wrong place: Android WebView honours it and iOS ignores it for
pinch, so the cost of stopping an iOS focus auto-zoom was deliberate
zoom on Android, taken from the users who need it most.

The font size is where it belongs. `src/platform/text-input-font-size.ts`
is the app's body size and `.web.ts` is that raised to 16, the size below
which iOS zooms on focus and does not zoom back. The commit bar and the
review note composer take their `fontSize` from it. A phone renders what
it rendered before: the native constant is `typography.bodySize`, so both
style objects are unchanged there.

`Math.max` rather than the literal, so a theme that raises the body size
past 16 keeps its own value.

The zoom guard stays and its rationale is rewritten to say what now keeps
the ordinary path off it: the inputs clear the floor, so a scale other
than 1 means a user pinched rather than an input took focus.

The pin is a unit case because the render check has no route to open yet.
Three assertions and what reds each: the web constant below 16 reds the
first, and a style going back to `typography.bodySize` reds the third,
which reads the two stylesheets as source because a node test resolves
the native sibling and would otherwise pass while shipping 14px to the
web. The overrides census covers the swap itself.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): put every text input in the two closures on the size seam (OTA phase C, C4.2 round 2)

The 16px floor reached two inputs and the rationale claimed a page. Eight
more text inputs in the same two closures still declared 14px, so a focus
on any of them zoomed the document and the occlusion seam — which reads a
scale other than 1 as no keyboard — stopped lifting for the rest of that
session. "A scale other than 1 means a pinch" was false while they were
there.

All eight go through `TEXT_INPUT_FONT_SIZE`, named by the census before
the change:

  src/components/MobileSearchField.tsx:175
  src/components/SmartWorkspaceAdvancedFields.tsx:84
  src/components/SmartWorkspaceSourceField.tsx:137
  src/components/new-worktree-form-styles.ts:125
  src/components/pr-sidebar/MobileLinkPrForm.tsx:120
  src/components/pr-sidebar/mobile-pr-sidebar-styles.ts:299
  src/components/pr-sidebar/pr-comment-composer-styles.ts:20
  src/components/smart-workspace-source-drawer-styles.ts:60

Every one declared `typography.bodySize`, so there was no input carrying
a size of its own to preserve and the phone is byte-identical again. Each
of those style keys was checked for consumers first: all of them are read
by a `TextInput` and nothing else, so raising the web value moves no
other element.

The census is the rule rather than the list. Over both closures it
resolves each `TextInput`'s style to the module that really declares the
size — following a spread, because both seam-served inputs are reached
through `{ ...base, ...list }` and a walk that stopped at the first
module would have called their offence absent — and names anything not on
the seam as `path:line`. A style with no `fontSize` inherits and is not
an offender. Presence precondition: the seam's web file is in the
closure, so an empty list cannot mean a page with no inputs.

Run against the previous head it prints exactly those eight for both
routes; three fixtures under mkdtemp cover the cross-module line, the
spread, and the two non-offender shapes.

The web test's rationale named `maximum-scale=1`, which is gone; it names
the input floor now.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): make the input census prove its own enumeration (OTA phase C, C4.2 round 2 addendum)

The offender list only says every text input is on the seam if every text
input was read, and the walk could not tell "this key sets no size" from
"I could not follow this style" — both answered nothing, so a resolution
failure would have read as a clean input and the rule would have gone
quietly vacuous.

`resolveStyleKey` answers three ways now: not found, found with no size,
found with one. `unresolvedTextInputStyles` reports the first as
`path:line (key)`, and the census asserts it is empty for both closures
beside asserting the offender list is.

Measured rather than assumed, which is what the addendum asks for. The
two closures hold 12 `TextInput` elements and 13 style references; none
uses an inline style object and none is without a style prop. All 13
resolve, 12 to `TEXT_INPUT_FONT_SIZE` and one — `styles.disabled`,
combined with `styles.input` on the same input — to a style that really
sets no size. The reviewer picker is in that list at
`mobile-pr-sidebar-styles.ts:300`; it was already on the seam from the
previous commit, which enumerated from the closure rather than from the
review.

A fourth fixture plants both shapes side by side: a style with no size,
which is not an offender, and a style reached through a package import,
which is named.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): close three holes in the text-input census (OTA phase C, C4.2 fold 3)

All three of CodeRabbit's findings are on the completeness property the
addendum bought, and all three reproduced before the change: each shape
below answered 0 offenders and 0 unresolved, which is to say it vanished.

Inline style literals. The walk recorded only `object.key` references, so
`style={{ fontSize: 14 }}` was neither an offender nor a hole. Style
props are flattened structurally now — arrays, spreads, `?:`, `&&` and
parentheses down to the expressions that can really land — rather than
walked as a subtree, which had the second bug of descending into an
inline literal's own properties. `&&` is followed because
`[styles.input, disabled && styles.disabled]` is the shape this tree
actually uses; `null`, `undefined` and `false` branches contribute no
style and are dropped rather than called unfollowable. An inline literal
resolves in place, and any other shape — a call, a bare identifier —
lands in the unresolved list.

Source-order precedence. `{ input: safe, ...legacy }` is `legacy.input`
at runtime, and answering direct keys before spreads read `safe` and
called the override clean. Properties are walked in reverse source order
now, direct keys and spreads in one pass, first answer wins.

The seam by binding. `size.text !== SEAM_EXPORT` accepted anything
spelled `TEXT_INPUT_FONT_SIZE`, so a local `const TEXT_INPUT_FONT_SIZE =
14` two lines up passed, and so did an import of that name from any other
module — the regression the seam exists to stop, wearing its name. The
identifier is resolved in the declaring module and accepted only as an
import from `src/platform/text-input-font-size`.

That last one changes what a fixture must say: the existing seam case
spelled the name without importing it, so it plants the seam module and
imports from it now. Six new fixtures, all six red on the previous walk.

Re-measured at this head, both closures: 12 `TextInput` elements, 13
style references, 12 on the seam, 1 sizeless (`styles.disabled`, combined
with `styles.input` on one element), 0 offenders, 0 unresolved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-20 01:29:30 -04:00
Jinwoo HongandClaude f3bda1bf3e refactor(mobile): seam moves and the shared shell route guard for the source-control domain (OTA phase C, C4.1) (#21732)
* refactor(mobile): open PR sidebar URLs through the external-link seam (OTA phase C, C4.1)

The three openers in the PR sidebar called `Linking.openURL` directly:
a check's "open on the web", a comment's permalink, and a link inside
comment Markdown. Inside the shell's WebView react-native-web routes that
to `window.open(url, '_blank', 'noopener')`, which both shells refuse and
which resolves anyway, so the tap reports success and opens nothing. Both
C4 routes reach the sidebar, so both would have shipped that.

The census is the point rather than the three edits. It derives the two
future route closures through `mobileWebAppRouteClosure` and holds every
module in them to the seam, so a module entering either closure later is
ruled without anyone adding it here. Red first it named all three by
`path:line`: CommentMarkdown.tsx:2, PRChecksSection.tsx:2,
PRCommentCard.tsx:2, on both routes.

The walk it runs was the third copy of one function, so it moves into the
seam's own module beside the predicate that module exists to share, and
the files and tasks censuses now call it too. It reports `path:line` where
the copies reported paths; `reachesReactNativeLinking` keeps its name and
its meaning and is now derived from the line list, so there is one rule.
An empty offender list is empty in either shape, which is why repointing
the two landed censuses moves nothing they assert.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): copy through the platform clipboard seam in review and conflicts (OTA phase C, C4.1)

The two copy actions both C4 routes reach called `expo-clipboard`
directly: the conflict section's refresh commands and the review sheet's
notes. On the web that module is `navigator.clipboard`, which needs a
secure context — the iOS shell serves the page from a custom scheme and
Android from https, so the path works on one platform and silently not on
the other. `useClipboardWriter` is the seam C2.4 landed for exactly that.

Red first, the census named both routes: `ExpoClipboard.web.js` in each
closure, and `src/platform/clipboard.web.ts` in neither.

Both call sites also stopped ignoring whether the pasteboard took the
text. The conflict section already returned on a throw, so the seam's
rejection reaches an arm it had. `copyNotes` had none and its only caller
is `void controller.copyNotes()`, so a rejection would have been unhandled
with "Review notes copied" left on screen; it now catches and reports
through the screen's own error line. That is the one behaviour change here
and the reason `clipboard` joins its dependency array.

Its suite mocked `setStringAsync` as resolving `undefined`, which the seam
reads as a pasteboard that refused, so every copy would have gone down the
new refusal arm unseen. The mock now resolves `true` and two cases pin
both arms; mutating the catch away kills the refusal one.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): take the source-control router from the handoff seam (OTA phase C, C4.1)

The hub takes its router once, in the openers hook, and passes it down to
the runners and the panel — so one `useRouter()` is this domain's whole
reach into routing, and it was expo-router's own. Inside the shell's page
that posts no `navigate`, so the hub's push to review would stay in the
document whatever its grants, and its push to a native route would paint
Unmatched over the page. Inert today: no C4 route is registered yet.

`use-mobile-source-control-runners.ts` is the second case and the reason
the rule reads value imports rather than identifiers: it named expo-router
only to write `ReturnType<typeof useRouter>`, a value import in a type
position that keeps the module in the graph. `RouteHandoff` is the seam's
own name for that type.

The census is C3.1's, and its walk moves to `src/navigation` rather than
being copied a second time; each domain keeps only its own evidence, the
list of modules meant to hold a router. Red first it named both modules on
the expo-router rule and reported no handoff caller at all.

The C2.9 hop census is unchanged and cannot move: its targets come from
the call sites, and the derivation over this tree returns the same ten
targets and the same 26 unresolved sites before and after this commit,
byte for byte. Its `HANDED_OFF` pin is over registered routes, of which
this adds none.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): move the shell route guard out of the files domain (OTA phase C, C4.1)

`src/files/mobile-file-shell-route.ts` was never about files: it parses a
route against `BridgeInitRouteSchema` and builds the key a shell screen
remounts on. It moves to `src/mobile-web-shell/shell-screen-route.ts` as
`shellScreenRoute` / `shellScreenRouteKey`, with its test. The move is
pure — with the rename applied and comments stripped, the old file and the
new one diff to nothing.

Three routes had grown their own copy of the call and two had none. The
copies go: `agent-history` and `tasks` now ask the shared predicate, which
is the same schema and the same fallback they already had. `index.tsx` had
no guard at all, so a `.` or `..` host id was handed over and came back as
"Update Orca to open this workspace" painted over the native list behind
the switch; it now stays native. That is the one behaviour change here,
pinned red first and killed by mutation.

`web.tsx` keeps handing that route over on purpose and is exempt by name:
its fallback is a redirect to the route the user came from, so the host's
own verdict is the better answer there, which
`mobile-web-shell-route.test.tsx` already pins. No `key=` expression moved;
the three switches still key differently (host id, pathname, pathname plus
params) and making them agree is a behaviour change for another PR.

The census walks the route tree rather than a list, so a switch added later
is held to both rules without being added here.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): mount the copy cases without a client instead of casting one (OTA phase C, C4.1)

The changed-code gate flagged four type assertions on the two cases added
with the clipboard seam: they stubbed an `RpcClient` the way the file's
older cases do, and the gate reads changed lines. Copying reaches no
client at all, so they mount without one, which is both cast-free and a
truer statement of what the path needs.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): say what the censuses report and sort the red list by line (OTA phase C, C4.1 round 1)

Round-1 folds, four wordings and one ordering.

`externalLinkOffenders` said "every call site" and reports the line the
name enters the module: a named import once, however many times the module
calls `openURL`, because the import is what the rule is about and what has
to go. Only a namespace import reports its uses, there being no single
line to name. The docstring now says that.

Its red list sorted the rendered strings, which puts `:10` before `:2`.
It now sorts by path and then by line as a number. Pinned against a
written fixture rather than the tree, because the case needs a module with
sites either side of line ten and no module in a closure has to keep
having one — the first fixture used lines 11 and 12, where both orders
agree, and the mutation walked straight through it.

`shell-screen-route.test.ts` still named the files screens in its describe
after the guard stopped being theirs; it names what a switch does now.

`router-seam-census.test-support.ts` excluded `.test-support.ts` from the
walk, which the files census it was extracted from never did. Dropped, so
both censuses walk the same set. Inert today: neither `src/files` nor
`src/source-control` holds such a file, so it only decides the next one.

The `web.tsx` exemption claimed a redirect "that looks like nothing
happened". What was measured: adopting the guard there sends a `..` deep
link through `Redirect href="/h/.."` to the host route, which this PR
keeps native, so the developer lands on the host list with nothing said
about why the page did not open. The route is `__DEV__`-only and the
host's own failure screen is the better verdict.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): see every react-native alias, normalise the host id, surface a refused copy (OTA phase C, C4.1 CodeRabbit)

Three bots findings on #21732, all real.

**Every alias, not the first.** `reactNativeLinkingSites` found namespace
imports with `exec` and inspected only the first binding, so a module
importing the namespace twice and calling `Linking.openURL` on the second
reported no site at all. It reads every alias now and counts a line once
however many meet on it. Red first with exactly that fixture.

**The host id can be an array.** `app/h/[hostId]/index.tsx` read it bare,
and Expo Router answers a repeated key with one: `String(['a','b'])` is
`a,b`, `encodeURIComponent` makes that the single segment `a%2Cb`, and the
segment rule accepts it — so the shell opened a page for a host nobody
has. Through `firstParam`, as the other four switches do. Red first it
handed over `/h/host-1%2Chost-2`, and the empty-array case found a second
one: `[]` is truthy, so a bare read built `/h/` and handed that over too;
`firstParam` answers `''` and the route stays native.

That import pulls the source-control screen state, and with it the lucide
barrel whose `LucideProvider` re-export is the gap the web build patches,
so the suite mocks the barrel as the other suites do. It moves no page
closure: the closure resolves `index.web.tsx`, which this does not touch,
and the index route still measures 3426 modules, 289 local, 22 families.

**A refused copy said nothing.** `PRConflictingFilesSection` caught the
rejection and returned: no tick, no message, a tap indistinguishable from
one that copied. The label now carries the third state, reusing the tasks
page's own wording for it, and the component has its first test. Mutating
the failure arm away reds it.

Its prop narrows to `Pick<PRInfo, 'mergeable' | 'conflictSummary'>`, which
is what it reads and what let the test drop a cast the gate flagged; every
caller holds a full `PRInfo` and satisfies it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read the censuses' subjects as code, not as text (OTA phase C, C4.1 round 2)

Round-2 additions. A separate commit because `d4b14d54e5` was already
made and this lane does not amend.

**The seam walk parses now.** Matching `X.Linking` in the text named it
inside a comment that talks about it and inside a string that quotes it,
and the named-import regex did the same for a commented-out import.
Checked against the previous implementation, all three fixtures were red
there: the comment case reported lines 2 and 3, the string case reported
the string's line beside the real call, and `// import { Linking } from
'react-native'` reported line 1. The walk builds a `SourceFile` and reads
import declarations and property accesses, so comments and strings are
gone by construction and the quote styles stop being a special case. Cost
measured on the three closure censuses: 3.3 s, unchanged.

**The route census reads the call, not the import.** A switch that keeps
the import while the call goes — deleted, or moved behind a branch that
never runs — looked exactly like one that asks. It now needs both, proved
by mutation: dropping `shellScreenRoute(` from `tasks.tsx` while leaving
its import names `tasks.tsx`. A fixture carries the same rule in
isolation, since every switch in the tree calls what it imports and the
case would otherwise be unfalsifiable against it.

**And recognises a switch by its import** of `MobileWebShellScreen` rather
than by `<MobileWebShellScreen` in the text, so an alias or a line break
the formatter chose cannot hide one and a comment cannot invent one.

The `app/h/[hostId]` root stays written out: deriving it from the manifest
is not a one-liner from here, the manifest being an `.mjs` this test reads
as text. What ties the two together instead is a new case asserting every
registered pathname starts with that prefix, so a page route outside it
fails rather than going unwalked.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read the imported name, not the local one (OTA phase C, C4.1 CodeRabbit)

`import { Linking as NativeLinking } from 'react-native'` went through the
census untouched: the walk compared the specifier's local binding, which
is `NativeLinking`, while the imported name lives in `propertyName` when
a specifier renames it and only in `name` when it does not. Reproduced
before the fix — the aliased import with a call beside it reported no
site at all.

Reading `propertyName ?? name` closes it in both directions. A module
that renames `Linking` is named at its import line like any other, and a
module that imports `View as Linking` is no longer named for a local
binding that reaches nothing. The second was a false positive the old
comparison had by construction.

One more of the same class, found while checking and verified rather than
assumed: `import RN from 'react-native'` typechecks in this project (tsc
accepts it), and a default binding is the whole namespace exactly as
`* as RN` is, so `RN.Linking.openURL` through it was invisible too. The
default binding joins the alias set, which already reports uses rather
than the import.

Four fixtures. Three red on the previous walk: the renamed import, the
local-only `Linking`, and the default import. The fourth — an alias
imported that never reaches `Linking` — passed before and is here to hold
the other half of the rule, that importing react-native is not itself the
offence.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): parse each module as its own kind, and read re-exports (OTA phase C, C4.1 CodeRabbit)

Two ways a module reached `Linking` past the census, both reproduced
before the change.

Every file was parsed as TSX. In a `.ts` module `const id = <T>(value:
T) => value` is a generic arrow; as TSX it is an unclosed JSX element,
and the parser folds the rest of the file into the error node. A
`RN.Linking.openURL` after one reported nothing, and so did the same call
with its import above the arrow. The file name goes into the parse now
and TypeScript reads the kind off the extension; `externalLinkOffenders`
passes the real path, which it had all along.

`ExportDeclaration` was never inspected, so `export { Linking } from
'react-native'` put the name back in reach of anything importing that
module while the census saw an import list it was not on. All four shapes
are read — named, renamed, `export *` and `export * as` — and reported at
the export statement, which is the line to delete exactly as an import
is. A re-export of another name, or of `Linking` from somewhere that is
not react-native, stays unnamed.

Seven fixtures. Five red on the previous walk: the `.ts` generic arrow
and the four re-export shapes. The two that pass before and after hold
the other half, that re-exporting is not itself the offence.

The named-import and re-export clauses read `propertyName ?? name`
through one helper rather than two spellings of it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-19 21:06:05 -04:00
Jinwoo HongandClaude 65cde9bb80 fix(mobile): name the source-control Back controls and split the dock's Close from Back (OTA phase C, C4.3) (#21739)
* fix(mobile): name the custom-key drawer's Back for the accessibility tree (OTA phase C, C4.3)

`CustomKeyModal`'s Back is a bare `Pressable` with a label and no role, so
a screen reader has nothing to announce it as. It is reachable today from
the session sheets and the terminal shortcut settings, so this is a gap
now, not only inside the page.

It surfaces here because `screenTree` takes a screen's directory: C4.4
registering the review route puts the whole of `src/components` under the
Back rule. Fixing it in the PR that registers would make a route entry
carry unrelated accessibility work, so the census gains a case that holds
the arriving trees to the same two rules before the rows land. Red first
it printed exactly what the rule would:

  src/components/CustomKeyModal.tsx:191 role=none label=Back

Once `PAGE_SERVED_SCREENS` has the two rows, `CONTROLS` covers these trees
and the new case becomes a second reading of the same thing; it says so.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): split the source-control header's Close from its Back (OTA phase C, C4.3)

One `Pressable` served both modes — `onPress={onBack}` with a conditional
label, `X` docked and `ChevronLeft` full-screen. The Back census reads a
control by what its handler does, so it sees that one as a Back and then
requires a label starting with "Back", in a mode where the control
dismisses the dock beside the terminal. The cheap way to go green is to
call a close "Back", which satisfies the rule by making the wording wrong.

So the modes become two controls. Embedded presses `onClose` and is named
"Close source control"; otherwise it presses `onBack` and is named "Back
to session". Both carry the button role. The panel stops choosing by mode
and passes both handlers; the dock keeps exactly the behaviour it had, its
dismiss still `onRequestClose` falling back to a pop.

Probed before it was written: run through the census's own predicate, the
post-split shape yields one back control rather than two — `onClose`
matches neither the handler pattern nor a declaration this file holds,
being a destructured prop — so the Close is invisible to the rule and the
presence precondition still holds on the Back.

The component test is what the census cannot do. Nothing else pins this:
no golden names this component and no parity census covers
`src/source-control`. Red first against the single control, both cases
failed on the missing role. And renaming the Close to "Back to close
source control" — the gaming this split exists to prevent — reds the
component test while leaving the census green, which is the whole argument
for having both.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): assert a Back per arriving tree, and name the refresh control (OTA phase C, C4.3 round 1)

The arriving-trees presence case counted controls over the union of both
trees, so one tree answered for the other. Per tree now, in the shape the
block above it already uses per screen module.

Red first, with the mutation round 1 named: rename the route branch's
handler to `onDismiss` and its label to `Return to session`, and the only
Back in `src/source-control` disappears. The per-tree case names that
tree. The same mutation against the union count passes all six cases,
which is what the change is for.

The refresh control had a label and no role, so react-native-web renders
a `div` carrying `aria-label` and a screen reader announces no control.
Its two siblings in this header already carry one.

Two claims in my round-1 report were wrong and I am the reason the body
carried them. The Close pressing `onBack` reds the census's label rule as
well as the component test, not the component test alone; and no fixture
of the post-split shape exists — the shape is read from the real file.
Both were stated from reasoning rather than from a run.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): assert an arriving Back per screen, not per tree (OTA phase C, C4.3 round 2)

Round 1 moved the arriving-trees presence assertion from the union to
each tree, which was not far enough. `src/components` holds two Backs, so
the tree answers for both: renaming `MobileDiffReviewHeader`'s handler to
`onDismiss` and its label to `Return to session` leaves every case in the
file green, with `CustomKeyModal.tsx:191` standing in for the screen that
just lost its Back. Reproduced before the change — six passed with the
review header's Back gone.

Per screen module now, which is what the table above already does and for
the same reason its docstring gives: a directory with more than one
control cannot say which one a rule was written about. The two modules
are named and the trees derive from them, so the pair C4.4 adds to
`PAGE_SERVED_SCREENS` is the same pair spelled once here.

Both mutations red on the new case and name the file: the review header's
rename names `src/components/MobileDiffReviewHeader.tsx`, and round 1's
source-control rename still names its own, so this does not trade one
cover for another.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-19 20:51:47 -04:00
Jinwoo Hong cef4416115 fix(mobile): name the narrow host header's controls and gate the drawer's hardware back on web (OTA phase C, C2.10) (#21729)
* fix(mobile): name the narrow host header's controls

The header renders two toolbars and the phone sees the narrow one,
whose controls carried neither a role nor a name. A screen reader could
not find them, and C2.9's render check could only assert their absence
at 390 px. The wide toolbar already names every control from the same
state, so the fix is to say the same thing rather than invent wording:
filter, sort, group, accounts, tasks and the search toggle take their
wide sibling's role and label expression verbatim.

The census names a seventh site the plan did not: the Reconnect button
in the status bar above both toolbars, which is shared rather than
narrow and has no wide sibling. It announces through its Text child
today, so it takes the role and the string it already renders.

No layout, style, handler or order changed; the diff is accessibility
props only.

The census parses the file with the TypeScript API and rules that every
Pressable carrying an onPress has a button role and a name, and that a
control both toolbars render is named the same way in both. It keys the
pairing on the handler, because that is what makes two elements the
same control, and asserts each shared handler is found exactly twice,
so a control deleted from one toolbar cannot leave the naming rule
comparing a group of one with itself. Red first, naming all seven
sites by path and line.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): stop the right drawer arming hardware back on web

React Native Web logs "BackHandler is not supported on web and should
not be used." and returns an inert subscription, so inside the shell's
page every open of this drawer put that line on the console and armed
nothing. The gate is the one mounted-bottom-drawer and the file preview
already carry, with the same comment stating the degradation: there is
no hardware back in a WebView, and the shell owns the one the phone has.

The drawer had no render test. This one mocks react-native, the safe
area, gesture handler and Reanimated the way the bottom drawer's
hand-back test does, and reads the call rather than the console: on iOS
and on Android the handler is registered once for 'hardwareBackPress'
and released when the drawer hides, and on web it is never reached. The
three native cases are the control that keeps the web case honest; they
passed before the fix, which is what makes the single red meaningful.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): type the right drawer test's element helper

The tests-typecheck ratchet was red on the previous commit: the drawer's
props declare `children` as required, so passing it as createElement's
third argument left no overload matching. It is a prop here, and the
helper answers a ReactElement rather than a return type borrowed from
createElement.

Test files sit outside `tsc --noEmit`, so only the ratchet sees this;
it is the gate that exists because a type-level pin in an unchecked
test proves nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): write the right drawer test in JSX

Lint was red on the previous commit and I ran it in the same command as
the commit, so it landed: passing `children` as a prop to satisfy the
type checker is exactly what react(no-children-prop) refuses. The
canonical form settles both, so the test is JSX in a .tsx file and the
drawer takes its body as a child again. The StyleSheet mock's generic
needs the trailing comma a .tsx file requires.

Re-proved in this form: with the web gate removed the web case fails
and the three native cases still pass.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): derive the header's naming groups, and read a spread as unknown

Round 1, four folds.

The drawer's comment claimed the console line was observed inside the
shell's page. It was not: the drawer's one caller is the review screen,
whose route C4 serves, so no page closure reaches it today. The gate is
pre-emptive and now says so in its own words rather than borrowing the
bottom drawer's sentence.

The naming rule iterated a hand-written list, so it only ever compared
the six controls both toolbars render. Giving the two
`actions.openFloatingWorkspace` sites different labels left the census
green. The groups are derived from the discovered controls now, keyed
by the handler text, so any handler this header presses from more than
one place is compared and the failure prints both names. The declared
list stays as the precondition it always was: each of the six is found
exactly twice, which is what keeps the derived rule from holding
vacuously over a file with no repeated handler.

The scan read `Pressable` only and dropped any control whose `onPress`
read as empty text, which is what a spread reads as. It reads
`TouchableOpacity` too now, and a spread answers unknown rather than
absent: a control whose handler or whose accessibility props arrive
through one is kept, fails both rules, and prints `spread` rather than
`none`, so it can never be mistaken for a control the scan judged.

Red first on all four: the reviewer's disagreeing-label mutation, a
spread over the a11y props, a spread over the handler, an unlabelled
TouchableOpacity, and a shared control deleted from one toolbar.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): read a braced-empty label as unnamed, from one shared reader

Round 2, two folds.

Only a bare `""` or an omitted attribute read as unnamed, so
`accessibilityLabel={undefined}`, `{''}` and an empty template all left
a control with nothing to announce and the census green. Reproduced on
both Tasks sites with each of the three shapes before the fix. The
reader unwraps a braced expression now: a string or a no-substitution
template answers its own text, and the identifier `undefined` answers
empty, so all three read as unnamed.

That reader was a near-verbatim copy in both censuses, which is how one
of them could have gained this rule and the other kept the hole. It
lives in one module under mobile-web-shell now, named for what it reads
and typechecked by mobile tsc rather than by the ratchet alone. Both
censuses import it and neither changed an assertion; their diffs are
the deleted copies and the import.

Red first, five mutations: the three empty shapes on both Tasks sites
here, and `{undefined}` and `{''}` on the tasks Back, which the page
census now catches too and did not before.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-19 18:33:20 -04:00
Jinwoo Hong ac024d4f05 feat(mobile): serve the files explorer and preview from the page (OTA phase C, C3.1) (#21710)
* refactor(mobile): take the files screens' router from the handoff seam

Inside the shell's page a screen is one document standing in for one screen, so
a target the page does not render has to be handed back to the app that does.
`useRouteHandoff` is where that decision lives, and its web sibling is the only
thing that makes it; both files screens held expo-router's own `useRouter`, so
on the web the explorer's Back and the preview's Back would post nothing and a
target outside the page would paint Unmatched over the page it is on.

Natively this is the same object — `route-handoff.ts` is `useRouter()` — so no
behaviour moves here, and `back()` stays expo-router's until the navigate-back
verb lands and the seam starts wrapping it.

A census rather than a behaviour test: neither screen's own tests can see the
difference, because a push that is never handed off still works for a target
inside the page. It walks this directory, refuses a value import of
expo-router, and names the two screens that must hold a router so a walk that
found nothing fails instead of passing empty.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): let the shell stand in for the two files routes

Both route files take the index.tsx shape — flag, MobileWebShellScreen, native
screen as fallback — and both gain the `.web.tsx` sibling that shape forces.

Inert until the manifest lists these routes: the shell answers `native-route`
for a route the bundle does not name, which is what `fallback` renders, and the
flag is `__DEV__`-only besides. Listing them waits on C2.3 and C2.5.

The sibling is not a precaution. The manifest defers every route behind
`import()`, so a native-only route module is invisible until the page opens
that route; the render check now opens both and, without the siblings, painted
`expo-modules-core.requireNativeViewManager is not available on web` instead of
the screen. That is also why the two cases render the route rather than
asserting a file exists.

The file path never becomes a path segment: only `hostId` and `worktreeId` are
spelled into the pathname, encoded, and everything else — `relativePath`,
`absolutePath`, `cwd`, `pathText` — is a param, which is how a `/`, a space or a
`..` stays out of the segment vocabulary the bridge holds a route to. The
preview render case proves the round trip on `docs/my notes/readme.md`.

`mobileFilePreviewShellParams` drops a param the normalizer left `undefined`
rather than sending it empty, because the page reads these back through
useLocalSearchParams where `line: ''` and no `line` are different screens. Its
test drives the normalizer rather than a hand-written literal: the literal omits
the key entirely, so it held with the filter removed.

The preview case also records what React Native Web says out loud — BackHandler
is inert on web, so Android back inside the page skips the unsaved-draft
prompt. Named in the assertion rather than filtered out, so closing it is a
change to that line.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): ask about an unsaved draft in the screen, not through Alert

React Native Web's `Alert` is `static alert() {}`. Inside the shell's page that
made Back with an unsaved terminal-artifact draft a button that did nothing at
all: no prompt, because the dialog is a no-op, and no navigation either, because
the code took the branch that shows one. Silently, with nothing on the console.

The prompt is now a row under the header. Not `ConfirmModal`, which every other
confirm here uses: that is a `BottomDrawer`, and C1.9 has Reanimated's animated
styles never reaching the DOM node on WKWebView, so on iOS in the page the
drawer parks off-screen and Back would be dead a second way. This paints the
same on every platform with no animation behind it.

Hardware back is registered natively only. React Native Web's
`BackHandler.addEventListener` logs "BackHandler is not supported on web and
should not be used." and hands back an inert subscription, so the guard never
armed there regardless; the render check asserted that console error on main and
now asserts none. The degradation is real and stated rather than hidden: Android
back inside the page pops the native stack without asking, and the page's own
Back control is where the question lives.

The decision moved to a hook so it is testable without a screen: the prompt also
drops itself when the draft it was about is saved or reverted, which is a state
`Alert` had no way to be in.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep expo-haptics' DOM shim out of the page

expo-haptics has a web build, and with no `navigator.vibrate` — iOS Safari,
which is the WebView the page runs in — it fakes a haptic by appending a hidden
`<label><input type="checkbox" switch>` to `document.head`, clicking it, and
removing it, once per call. C1.9 traced a long press that never fired on the
worktree list to exactly that stray click, and the file explorer calls
`triggerSelection` on every row tap, so C3 is the first domain to fire it per
tap rather than per long press.

`haptics.web.ts` answers the same five names with nothing. A phone holding the
page is a phone whose native app is right there with the real haptics, and a
missing tap feedback is worth less than a tap that does not register.

The test reads the shipped bytes rather than the import, because that is the
claim: with the override removed the bundle carries `ariaHidden` and
`pointer: coarse`; with it, neither, nor the `setAttribute("switch"` that does
the clicking. Not `navigator.vibrate` — react-native-web's own Vibration export
calls that and touches no DOM until something invokes it, which cost this test
one wrong red before it was narrowed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the files routes native when the page could not be given one

A file path is a param, so `/`, spaces and `..` all cross safely — but
`BRIDGE_MAX_ROUTE_PARAM_CHARS` is 1024 and a Windows long path is not bounded by
anything the user cannot exceed.

The symptom is not the blank document the design predicted, and the correction
matters: `bridge-host.ts` already parses the route against the page's own schema
and drops it to `null` when it fails, so `init` arrives naming no screen and the
page paints "Update Orca to open this workspace" — a wrong message about a fine
app, over a native screen that works. Deciding before the switch instead leaves
the route native, which is where every route starts.

The schema is the predicate rather than a copy of its bounds, so the rule cannot
drift from the half that matters, which is the half the page reads. The same
call also refuses a `worktreeId` the segment rule will not route: `..` survives
`encodeURIComponent`, which is the C1.8 class.

The tests assert the schema really refuses each input before asserting the guard
does, so neither case can pass by being impossible.

This belongs in the shell beside the schema; it is in the files domain while the
contract files are the C2 lane's.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin what keeps a file path out of the route vocabulary

Seven shapes, one case each rather than a representative: a plain path, a space,
a dot segment, an already-encoded slash, a fragment, non-ASCII, and an absolute
path. Each is checked in the two directions a path travels — the href the shell
writes into the page's history, and the href the page would hand back — for both
the pattern accepting it and the path coming back out of the query unchanged.

The counterfactual is in the file: the same paths spelled as a segment are
refused. Without that, the cases above would hold for a rule that was never
doing any work. Mutating `stringifyRouteHref` to join its query by hand instead
of through `URLSearchParams` fails three of them.

Also fixes two new test files the tests-typecheck ratchet caught: the partial
`react-native` mock needs a typed `addEventListener`, `act` will not take a
callback that returns a value, and `findAllByType('Pressable')` does not
typecheck against `ElementType` — the neighbouring files that do it are
grandfathered, so the tag comparison goes through a helper instead.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): derive the discard prompt instead of clearing it in an effect

Both changed-code gate findings, which the lane had not run until the last
commit. React Doctor is right: the effect that cleared the prompt when the draft
went away adjusted state after a prop changed, so a save landing while the
prompt was up painted one frame still offering to discard nothing. The prompt is
now `asking && hasUnsavedDraft`, which cannot be stale by construction, and the
test that covers it passes unchanged.

The hoisted mock's `as` on a string literal is gone too: the literal narrows on
its own and the tests reassign it, so the holder is annotated instead.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): add the files routes to the hybrid shell flag census

The census pins every file that reads `useMobileWebShellEnabled`, because a
reader nobody listed is how a dark feature stops being dark. C3's two routes are
deliberate entries: each has a native screen behind it as `fallback`, and each
is inert until the manifest lists the route.

Found by the full mobile suite rather than by the files subset this lane had
been running per commit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): serve the files explorer and preview from the page

The last C3 commit: both routes join MOBILE_WEB_PAGE_ROUTES, and the shell
starts rendering the page for them on a phone with the dev flag on.

Grants are not the same for the two, and the difference is the point. Both take
`navigate` (Back pops the native stack, and the explorer's rows open the preview
beside it) and `storage` (the shared components the host layout renders above
them). Only the preview takes `externalLink`: a Markdown preview renders links
and `MobileMarkdown` opens them through the platform seam.

The explorer does not, and measuring is what says so rather than reading. Every
page route reaches `external-link.web.ts` — `/h/[hostId]` and agent-history
included, both granted nothing for it — because the protocol wall in the shared
host layout imports it. So closure membership is not the oracle for a grant; the
question is whether the route's own screens call it, and only the preview's do.
`MobileMarkdown` is in the preview closure and absent from the explorer's, which
the census now asserts in both directions.

Neither route writes a clipboard, so neither takes `native.clipboard.write`;
the census pins that as the absence of both `ExpoClipboard.web.js` and the
clipboard seam, with the tasks closure as the control that the probe can see one
when there is one.

The seam predicate moved into a module both censuses import rather than being
restated per series: two spellings of one rule drift, and this one is a regex.

Red-first: both manifest assertions failed on the new entries before they were
updated, and routing `MobileMarkdown` around the seam fails the preview's census
while leaving the explorer's passing, which is the asymmetry the grants encode.

Closure sizes as the page ships them, extensionless so the `.web.tsx` is what is
measured: explorer 3439 modules / 302 local / 10 under src/files, preview 3667 /
331 / 20.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): read the files route's ids as one value and key the shell on them

Two round-1 findings, both reproduced before the fix.

A repeated query key reaches `useLocalSearchParams` as an array, and the
explorer read `hostId` and `worktreeId` bare. `String(['a','b'])` is `a,b`, so
the template built `/h/host-a%2Chost-b/files/wt-1%2Cwt-2` — a single segment the
bridge's rule accepts, and the shell would open a page for a host nobody has.
Read through `firstParam` now, as the tasks and agent-history switches do. The
preview already went through `singleParam` and is unchanged.

Neither switch keyed `MobileWebShellScreen`, where `index.tsx`, `tasks.tsx` and
agent-history all do. A host captures the grants its session opened with, so a
screen reused across a route change keeps authorising frames under the grants of
the route the page has left; only a remount drops that bridge. Both are keyed on
the route pathname now, with agent-history's reason.

The new route test is the agent-history one's shape. It caught both: the array
case landed on no route at all, because `name` was an array too and the schema
refuses a non-string param value, and the two lifecycle cases saw a prop update
where a remount was owed. It also needs agent-history's `lucide-react-native`
mock, since `firstParam` lives in the source-control barrel.

`name` is now omitted when empty rather than sent as `name=`, matching the two
switches beside it: an absent label lets the panel derive its own.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): confirm a discarded draft with the app's own modal

Round-1 findings 3, 4, 5 and the minor one.

**ConfirmModal, not the bespoke row.** The row existed because C1.9 had
Reanimated's animated styles never reaching the DOM node on WKWebView, which
left every BottomDrawer parked off-screen. C1.10 (`b7c06900e2`, an ancestor of
this branch) fixed that with a dependency array on the mapper hooks, and the
drawer render check now holds it on WebKit as well as Chromium. With the reason
gone the row does not stand on its other merits: `Alert.alert` was modal on
native before the page existed, and the row quietly changed that for phones
too, so the app's own confirm is both the idiom and the closer behaviour.
`MobileFilePreviewDiscardPrompt`, its test and its thirty style keys are gone;
the hook's state machine and its tests are unchanged.

**The encoding test claimed more than it pinned.** Hand-joining the query reds
only three of the seven shapes; `docs/readme.md`, `../etc/passwd`,
`docs/日本語.md` and `/logs/run.txt` are encoding-neutral in the query, whose
pattern half is `[^#\s]*` and admits a slash, a dot segment and non-ASCII
verbatim. Rather than narrow the claim in a comment, the split is now pinned by
behaviour: each neutral shape must survive the query unencoded, each
load-bearing one must not. Moving `docs/readme.md` between the lists fails it.

**The manifest comment named one shared-layout opener and there are two.** The
New Workspace source field, which the sidebar renders on a wide layout, opens a
URL through the seam as well. Both are the shared layout's and every `/h` route
reaches both, `/h/[hostId]` included with no `externalLink`, so the tablet tap
is dead on all of them — recorded here as pre-existing rather than fixed, since
the grants do not move.

**Minor:** the dot-segment case in the guard test now asserts the schema refuses
the route before asserting the guard returns null, as the length case does.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): stop every page drawer logging a BackHandler error when it opens

Round-2 findings.

**The registration belongs to the drawer, and that is where the guard went.**
`mounted-bottom-drawer.tsx` armed `hardwareBackPress` whenever a drawer was
visible and interactive, with no platform check, so the hook's claim to have
dropped that console line held only while its prompt was closed — and every page
drawer since C1 has logged it on open. Platform-gated at the drawer now; the
hook's comment says so rather than claiming the credit.

**Nothing had ever opened a modal in a browser.** The render check next door
mounts both files routes and reads what they paint but taps nothing, so
`ConfirmModal` inside the page — a BottomDrawer, so Reanimated, a portal and a
gesture handler — was unproved. A new render file loads an editable terminal
artifact through the harness's scripted reply, edits it, taps the page's Back,
and asserts the prompt's title is up and no BackHandler line is on the console.
Red first on exactly that line; the prompt itself painted, which is also the
first proof on a browser that C1.10's fix carries a real drawer in the page. A
second case answers Stay and checks the draft survives. Its own file rather than
the render check's, which is at 482 of the 600-line cap; registered in pr.yml.

**The encoding rule was stated wrong.** Two rules decide it and neither is about
paths: the pattern's query half refuses whitespace and `#`, and
`URLSearchParams` is form-urlencoded, so it reinterprets `&`, `+` and a valid
`%XX`. `a+b.ts` reads back `a b.ts` and `a&b.ts` reads back `a`, so both are
load-bearing; `a=b.ts` and `a%b.ts` are not, because only the first `=` splits
the pair and a lone `%` begins no escape. A newline joins the load-bearing list
as the refused shape rather than the altered one.

**The web sibling read its params bare** where the native one uses `firstParam`.
Not reachable — the page only arrives through `init.route`, whose params are
already `Record<string, string>` — but the two files are meant to be one screen.

The preview keys on the pathname alone, and the comment now says why that is
enough: every caller in this tree pushes.

Closures after this: explorer 3441 / 304 / 10, preview 3666 / 330 / 19. The
explorer grew two modules because its web sibling now reaches `firstParam`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): give the explorer the grants the preview needs, and key on the route

Bot findings, one of them a real gap.

**Pullfrog is right, and my grant oracle was half a rule.** Grants resolve once,
from the route the shell opened: `grantsForRoute` reads `session.routePathname`
and `init.grants.native` carries the answer for that session. The explorer's
rows push to the preview, and because the preview is a page route that push
stays inside the same document — no second `init`. So a preview opened that way
runs under the explorer's grants, and a Markdown link in it was refused by
`notifyExternalLink` with nothing on screen to say why. "Does the route's own
screen call it" was right for a route's own screens and wrong for the routes it
reaches in-page, so the explorer now declares `externalLink` as a transitive
grant, with the comment saying that rather than claiming it opens links. The
census pins the pair as a superset; removing the grant reds it.

**The seam regexes matched one quote style.** A double-quoted `react-native`
specifier walked past both censuses unseen. Both styles now, with the predicate
tested directly for the first time.

**The discard request outlived its draft.** `asking` stayed set after a save or
a revert, so the next edit re-showed the prompt with no Back request behind it.
The request is now dropped when the draft it was about goes, adjusted during
render rather than in an effect — the shape React Doctor named in the round-1
fold. Red first: save with the prompt up, edit again, prompt is back.

**CodeRabbit's keying comment is a correctness point, not the question I
answered.** The page learns its route exactly once, out of `init`, so a
same-path param change — another file in the same worktree — left the shell
mounted and the page still showing the file it was opened on. My comment claimed
"the screen reloads the preview from the param either way", which is true only
with the shell absent. Both switches key on the whole route now, params
included; two tests cover the same-path case and both red on a pathname-only
key.

`build-mobile-web-app-bundle.test.mjs` hit 601 of its 600-line cap on the way,
so the two manifest assertions now share one expected list instead of repeating
it. Closures unchanged: explorer 3441 / 304 / 10, preview 3666 / 330 / 19.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make the seam test import the module it is testing

Round 3.

**The blocker is mine and the reviewer's diagnosis is exact.** The seam
predicate test imported an absolute path into this lane's worktree. On CI that
module does not exist and it takes the whole `config/scripts` suite down; here
it resolved to the same file by accident, so the test was green against a tree
rather than against the checkout — which is why reverting the double-quote fix
left it passing and the predicate untested. Relative now, and proved: reverting
the fix in place reds both double-quoted cases, which is the first time this
test has failed for the right reason. Every file this PR touches is grepped for
`/Users/` and `orca-lanes`; none carries a path.

**Three comments outlived the grant change.** The two lists became equal when
the explorer took `externalLink`, so "longer than the explorer's" and "declared
with different grants" were both false. Corrected to what is actually true: the
lists are equal and the reasons are not — the preview has its own consumer in
`MobileMarkdown`, the explorer has none and declares the grant because its rows
push to the preview in-page.

**The duplicated serializer is pinned rather than imported.** `shellRouteHref`
lives in `page-bootstrap.ts` beside the page's RPC client and its document
channel, so a native route file importing it would pull both into the app. The
copy stays, and a test asserts the two agree on three routes; dropping the
empty-search branch reds it.

**Recorded, not fixed:** the sidebar `HostScreen` pushes to `/h/<id>/tasks`
through the handoff, which is local, so on a tablet the tasks page runs without
`native.clipboard.write` from any page route and its copy actions refuse
silently. Pre-existing since C2.1 for the worktree list and agent history. Named
in the explorer's manifest comment as the known remaining hop, with the fix
being a handoff rule in its own PR.

The equality pin needed `it.each<BridgeInitRoute>`: the inferred table is a
union whose members carry `?: undefined`, which the ratchet caught.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-19 15:47:25 -04:00
Jinwoo Hong b6e8b1a7b2 feat(mobile): serve the tasks screen from the page, with its seams (OTA phase C, C2.1 + C2.5) (#21694)
* fix(mobile): encode the host id in the tasks workspace-creation href (OTA phase C, C2.1)

`use-mobile-tasks-workspace-create-actions.tsx` built
`/h/${hostId}/session/...` with the host id interpolated raw — the C1.2 class.
A host id carrying `/`, `#`, `?` or whitespace reaches the wire as an href
`BRIDGE_ROUTE_HREF_PATTERN` refuses, the handoff falls through to the local
router, and expo-router's Unmatched paints over the page.

Deleted rather than patched: `hostNewWorktreeSessionRoute` already builds
this exact href with both segments encoded, and already has the test that
pins it. The screen now calls it.

The census that caught it stays: no module under `src/tasks` may interpolate
into `/h/${...}` without encoding, which is the rule rather than this one
line. Three refactor-parity hashes move with the statement change and are
recorded in that file the way every earlier movement is.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): route the tasks tree's external links through the seam (OTA phase C, C2.1)

Ten of the twelve call sites in the tasks page closure: the nine under
`src/tasks`, swapped by one export in the dependency barrel, and
`MobileMarkdown.tsx`, which imports react-native directly and is edited in
place.

Inside the shell's WebView react-native-web's `openURL` calls
`window.open(url, '_blank')`, which both shells refuse — iOS returns nil from
`createWebViewWith`, Android false from `onCreateWindow` — and resolves
regardless. Every one of these sites would have reported success into a tap
that opened nothing.

The barrel's `Linking` is typed `{ openURL: (url: string) => void }`, so a
`.catch` on it is a compile error rather than a handler for a rejection that
cannot arrive; the seam names its own failures. `MobileMarkdown`'s own
`.catch(() => {})` goes with the swap for the same reason.

No parity hash moved: the barrel and `MobileMarkdown` are outside the
refactor-parity family's source set.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): route the shared screens' external links through the seam, with a census (OTA phase C, C2.1)

The last two of the twelve call sites in the tasks page closure:
`ProtocolBlockScreen.tsx` and the `openExternalUrl` prop wiring at
`host-screen-overlays.tsx`.

Both are shared with native routes and with the already-live `/h/[hostId]`
page, so this changes that page too: its external links go from the measured
`window.open` no-op — which both shells refuse and which resolves anyway — to
a URL handed to the shell. Nothing changes on a phone, where the seam is
`Linking.openURL` unchanged.

The `openExternalUrl` prop chain is retyped `(url: string) => void` with it,
and `SmartWorkspaceSourceField`'s `.catch(() => {})` goes: the seam names its
own failures and never rejects, so that was a handler for a rejection that
cannot arrive.

The census is the rule rather than today's twelve sites: no module in the
tasks page closure may reach react-native's `Linking`, by name or through a
namespace import. It reads the closure from a new builder export —
`metafile.inputs` for `_layout` plus the route, which is one definition of
what a page contains — and checks which module the name comes from, not which
text a call site writes, since the tasks tree still calls `Linking.openURL`
and that `Linking` is now the barrel's seam-backed export. Confirmed to
discriminate: restoring one react-native import turns it red.

A second case pins that the seam is in the closure, so an empty offender list
cannot also mean a page that reaches no link code at all.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): write the tasks clipboard through the shell's verb (OTA phase C, C2.1)

The two `Clipboard.setStringAsync` sites in the tasks page closure move onto
a seam, `src/platform/clipboard.ts` with a `.web.ts` sibling, registered in
the overrides.

A hook rather than a function because the web form needs the page's bridge
client, which is React context. Native is `expo-clipboard` unchanged. Web
calls `native.clipboard.write` through `useNativeVerbs`, because
`expo-clipboard` on the web is `navigator.clipboard` and needs a secure
context: the iOS shell serves the page from a custom scheme and Android from
`https`, so that path would work on one platform and silently not on the
other, with nothing at the call site able to tell.

Both seams reject rather than return false, and both call sites already wrap
the write in a `catch` that puts the message on screen — so a write that did
not land says so instead of showing "Copied". A route that has not declared
`native.clipboard.write` is refused before a frame is sent and lands in that
same `catch`; the route declares it in the entry commit.

Two parity hashes move, the hook list and the statement hash, each by one
entry, and are recorded in that file. `semantics` holds, as do render and
style: no RPC call, method literal or JSX host signature changed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): hand the tasks Back button to the shell (OTA phase C, C2.1)

The tasks header's `router.back()` reached expo-router through the dependency
barrel, and inside the page that moves nothing: the document holds the single
history entry the entry wrote with `replaceState`. The stack with somewhere
to go is the native one the shell pushed the page onto.

One line in the barrel, as with `Linking`: `useRouteHandoff` is router-shaped,
so every call site is unchanged. On a phone it is expo-router. Inside the page
it keeps a route the page renders and posts `navigate-back` for a Back the
document cannot serve — the C2.2 seam, which until now had no consumer.

No parity hash moved: the barrel is outside the refactor-parity source set,
and no call site changed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): render mermaid as its own source box on the web (OTA phase C, C2.5)

`MermaidDiagram` is in the tasks page closure, reached through
`MobileMarkdown`, and it renders the diagram inside a sandboxed `WebView`.
`react-native-webview` is a native component with no browser counterpart:
importing it runs a codegen lookup that throws, and the route manifest imports
every route, so one such import takes the whole page down rather than one
diagram.

The web sibling renders the labelled source box the native component already
falls back to on a parse or render error, with that component's own styles, so
the degradation looks like a state the product already has rather than a
second design.

Not a browser renderer, and the reason is not reach: mermaid is a browser
library and the engine bundle is vendored. It is that the native path's safety
comes from the WebView it runs in — `buildHtml` escapes `</script>` and the
U+2028/U+2029 separators because diagram source is untrusted agent and PR
content — and a DOM path has no such sandbox, so it needs its own escaping and
its own proof. That is a change of its own, not a smaller version of this one.

Registered in the overrides, whose gate fails on an unlisted `.web.*` file.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): turn the tasks route on for the page (OTA phase C, C2.1)

The entry: `/h/[hostId]/tasks` joins `MOBILE_WEB_PAGE_ROUTES`, the route file
becomes the shell's flag switch in `index.tsx`'s shape, and a `.web.tsx`
sibling renders the screen directly, registered in the overrides.

The screen moves to `src/tasks/MobileTasksScreen.tsx` first, verbatim — body
byte-identical, imports rewritten to `./`. It has to: under the builder's
`resolveExtensions` a web sibling importing `./tasks` resolves back to
itself, which is why every other shell route's screen already lives in `src`.

The parity family follows the file rather than the path. `TASKS_ROUTE` leaves
`MOBILE_TASKS_SOURCE_FILES` — `SOURCE_PATTERN` already matches
`MobileTasks*.tsx`, so listing it too would double-count — and the execution
reader points at the new file. Measured rather than predicted: all six
refactor-parity cases pass unchanged. No hash moved, including the family
text and declaration list, because the new name sorts where the route path
sat.

The route declares `navigate`, `storage`, `externalLink` and
`native.clipboard.write`, which the grammar fold made expressible and
per-route scoping makes meaningful: it is granted those and not the rest of
what this shell implements.

The browser check covers what only a browser answers — every module in the
closure evaluating under React Native Web, `taskSource` surviving the
handshake into the page's own URL, and the route's chunk arriving on a
client-side navigation. It states plainly what it does not cover: the three
seams are reached from controls that need provider data the double does not
serve, so a case posting those frames directly would prove the transport and
read as a tap it never performed. Both new checks join the `mobile_web_app`
job.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(config): resolve a route closure the way the bundle ships it (OTA phase C, C2.1)

`mobileWebAppRouteClosure` took the route's explicit `.tsx` path as an entry
point, so esbuild used that file directly and `resolveExtensions` never ran.
For a route with a `.web.tsx` sibling that measured the native switch, which
no browser loads: the tasks closure came back carrying
`MobileWebShellScreen`, and with it a `Linking` import the census then
reported as an offender.

Extensionless now, so the closure is the one the page actually contains:
3775 modules, 428 local, with `external-link.web.ts` and `clipboard.web.ts`
in it and the shell screen out.

The route-manifest pins move with the tasks route joining
`MOBILE_WEB_PAGE_ROUTES`, in both the declaration check and the built
manifest.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): cover the clipboard seam, close two page escapes, share the mermaid props (OTA phase C, C2.1)

Four from round 1.

The clipboard seam shipped untested. Both halves have one now: the native
form rejects when `setStringAsync` answers false and resolves when it does
not, and the web form is driven through the real port pair — resolving on a
reply, rejecting when the shell says the pasteboard refused, and rejecting on
an ungranted route without putting a frame on the wire.

The tasks barrel still re-exported `expo-clipboard` with no consumer, which
kept `ExpoClipboard.web.js` — the `navigator.clipboard` path this series
exists to avoid — inside the page closure. Deleted, and asserted as the
module's absence from that closure rather than as a count of importers: a new
import puts the file back whoever writes it.

`ProtocolBlockScreen` reached expo-router's singleton for its way out to the
host list. A singleton is the one shape the handoff cannot intercept — it is
not a hook, so the page's bridge client is never consulted — and `/` is a
route the page does not carry, so inside the shell that replace rendered the
root route in the WebView instead of leaving it. Pre-existing and live via
`/h/[hostId]`; routed through the handoff now. Two suites' `expo-router`
mocks gain the hook the handoff reads.

`MermaidDiagram.web.tsx` redeclared its props; it imports the native
component's type, so drift fails tsc.

No parity hash moved: none of these files is in the refactor-parity source
set.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(config): use endsWith for the clipboard module check

The changed-code gate refuses a dollar-anchored regex where `String#endsWith`
says the same thing. No behaviour change.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): close the href census gap, read route params through firstParam (OTA phase C, C2.1)

Five from round 2, two of them real.

The raw-interpolation census inspected only the leading `${...}`, so
`` `/h/${encodeURIComponent(hostId)}/session/${worktreeId}` `` passed it — and
a worktree id carrying `/`, `#`, `?` or whitespace breaks the href exactly as
a host id does. It now refuses any hand-built `/h/...` template with any
interpolation left raw, whichever segment it is. Proved against exactly that
shape in a throwaway before the change, which the old rule admitted.

The tasks switch read `hostId` and `taskSource` as plain strings. expo-router
hands back an array for a repeated query key, so a duplicate `?hostId=` built
`/h/host-a%2Chost-b/tasks`; both go through `firstParam` now, as the
agent-history switch does. `index.tsx` is untouched, per the Phase D list.

Three in the render check's prose: the header claimed the browser proves the
three seams fire from a tap, which the file's own closing note denies; a
module count repeated a number the closure test already pins; and a `replies`
parameter was threaded through without ever being supplied.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-19 13:40:26 -04:00
Jinwoo Hong b7c06900e2 fix(mobile): give reanimated mapper hooks the inputs esbuild never writes (OTA phase C, C1.10) (#21592)
* refactor(mobile-web): extract the page render harness

The shell double, the CSP/bridge constant readers and the bundle server were
private to mobile-web-app-render.test.mjs, so a second check against the same
page had no way to reach them. Moved as-is into a module both can import; the
double also gained a `replies` map so a check can answer one method and leave
the refusal in place for everything else.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): give reanimated mapper hooks a dependency array

The bottom drawer never slid onto the screen in the web shell page: `progress`
animated to 1 and `withTiming` reported finished, but the sheet kept the
translateY of the animation's first frame and sat one viewport below the fold,
with its invisible backdrop swallowing the next touch.

Cause, bisected in the browser: `useAnimatedStyle` reads its mapper inputs from
`updater.__closure` (hook/useAnimatedStyle.js), which only Reanimated's Babel
plugin writes. The page is bundled by esbuild, which runs no Babel, so
`__closure` is undefined; with no dependency array either, `inputs` is empty and
`startMapper` registers a mapper that listens to no shared value. It runs once
and never again. Reanimated does throw for exactly this, but behind `__DEV__`,
which the bundle builds out, so the page reports nothing. The rAF loop stopping
after one write is the observable end of it.

Not a WebKit fault. Headless Chromium parks the sheet the same way
(translateY(843) vs WebKit's translateY(841)), so the earlier
JavaScriptCore-vs-V8 reading does not hold, and the pin added here runs on both
engines rather than on Chromium alone. WebKit is downloaded in the
mobile_web_app job for it.

Every mapper-backed call site takes the same array, not just the drawer's:
RightDrawer and DragReorderList are the same defect on the same bundler.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): census the reanimated hooks that need a dependency array

The drawer pin covers MountedBottomDrawer only, and the failure mode is silent:
a new `useAnimatedStyle`, `useAnimatedProps` or `useDerivedValue` without an
array animates once on the phone's native build and freezes in the web page,
with no error on either. Parsed rather than grepped so a call spanning lines,
or one whose second argument is not an array, is still seen.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile-web): state motion-on as the drawer pin's precondition

Under `prefers-reduced-motion: reduce` Reanimated finishes `withTiming` in one
frame, so a mapper that only ever runs once still writes the final translateY
and the pin goes green on the broken build. Measured: the unfixed bundle under
reduced motion lands at translateY(0) with the sheet on screen in both engines,
which is also what the Android emulator does with animator scale off — the same
single write, not a healthy animation. The context now says no-preference and
the page is asked to confirm it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): census useAnimatedReaction, whose deps are its third argument

Same fallback as the other three (hook/useAnimatedReaction.js:26-34), so the
same silent freeze applies. Its shape is not the same: the array is argument
three, behind `prepare` and `react`, and both callbacks run inside the one
mapper it starts, so both count as updaters. Indexing it like the others would
have read the `react` callback as the array. No call site today; this is the
gate for the first one.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): require the dependency array to list every value the updater reads

An array proves a call was written, not that the mapper listens to everything
it reads. On web `inputs` becomes exactly that array
(hook/useAnimatedStyle.js:338-341), so a value read but not listed is a value
the mapper never hears about: the updater stops re-running when only that one
changes. Same freeze as no array at all, in one prop rather than all of them.

Reads only. The first fixture caught this check counting `opacity.value = v` as
a read, which it is not -- a written value is an output, and demanding it in
the array would be noise at every `useAnimatedReaction`. Assignment targets and
increments are excluded; a value both read and written is still required.

Verified against the tree by dropping `translateY` from the bottom drawer's
array, which the census names at mounted-bottom-drawer.tsx:286.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): resolve the hook through the file's imports, not by spelling

Matching the callee's text both missed and invented. `useAnimatedStyle as useAS`
and `Reanimated.useAnimatedStyle` are the same hook wearing another name and
went unchecked; a local helper that happens to be called `useDerivedValue` is
not this hook and would have been flagged. Each local name is now resolved
through the file's imports from `react-native-reanimated`, named, aliased or
namespace member.

A second argument that is not a literal array now counts as present rather than
missing: the hook only needs an array to exist, and this file cannot see what a
hoisted `const deps = [...]` holds, so completeness covers literal arrays only.

Resolution can fail closed, which would read exactly like a clean tree, so the
census now asserts it saw the calls before asserting none are missing. Checked
against the tree by dropping `translateX` from RightDrawer's array, which it
names at RightDrawer.tsx:156.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile-web): select the drawer sheet by name, not by its corner radius

The pin walked up from the handle to the first ancestor with a 16px top radius,
so it found the sheet through a styling token. Change that radius and the pin
reports `sheet: false` -- a red naming the selector rather than the animation it
exists to watch, on a change that broke nothing.

The sheet now says what it is. `testID` on the RN side renders as `data-testid`
on web (react-native-web createDOMProps/index.js:832), which is the one line of
product change this needs.

Re-verified after retargeting: still red on both engines with the dependency
arrays removed (translateY 843.271 chromium, 841.447 webkit), green with them.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile-web): say why the motion option must precede navigation

Reviewer follow-up on the reduced-motion guard. The context option and the
`goto` order are both load-bearing, and nothing in the file said so: Reanimated
reads `matchMedia('(prefers-reduced-motion: reduce)')` once into a module-level
const at import (ReducedMotion.js:8-10), so a `page.emulateMedia()` after
navigation would leave the assertion passing over a value already latched true.
Comment only.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile-web): name the drawer pin's precondition instead of asserting past it

CI's Linux WebKit failed this pin at `matrix(1, 0, 0, 1, 0, 844)` -- exactly the
viewport, the mount-time value, not a first-frame 843.x. Nothing animated there,
so the pin was reporting a parked sheet without being able to say whether the
mapper was subscribed. Two different faults, one message.

`requestAnimationFrame` separates them and sheet writes do not. `withTiming`
schedules a frame per step (valueSetter.js) whether or not a mapper listens, so
frames across the window mean the shared value moved; the assertion now names
that. Counting sheet writes as the precondition inverts the diagnosis: measured
on the broken build, "written more than once" fires first and calls the defect
this pin exists to catch an engine that does not animate.

Sheet writes stay, as a second statement of the subject and as context in the
transform failure, which now reads "1 style write(s) on the sheet across 30
frame(s)" on the broken build.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile-web): wait for the drawer to arrive, not for a clock

The pin paused a fixed 1s after the sheet opened and then read the transform,
which makes it a race on a loaded runner: a healthy engine that is merely slow
reads as parked, and the red names the transform rather than the wait. It now
waits for the settled transform, times out at 15s, and asserts on whatever it
found either way, so a genuinely parked sheet gives the same red with the
timing assumption removed. On the broken build that red now reads "1 style
write(s) on the sheet across 3635 frame(s)", which says the fault in one line.

Aimed at CI's Linux WebKit red rather than proven against it: eight container
runs on the Playwright Linux image never reproduced that failure. See the
report for what the container did and did not show.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile-web): stop asserting on the sheet's style-write count

The count cannot carry an assertion in either direction. Measured under
`--cpus=0.35` in Playwright's Linux image, a healthy page starved of frames
reaches translateY(0) in a single write, because `withTiming` covers the whole
180ms in one step when one step is all the frames it gets. "Written more than
once" would have redded that page, which is a CI runner under load -- the exact
situation this pin keeps meeting.

So the transform is the only subject, `requestAnimationFrame` during the window
is the only precondition, and the write count is context in the failure text.

Also worth recording against the CI log: exactly `matrix(1, 0, 0, 1, 0, 844)`
is reproducible here on the broken build, as the single mapper run landing at
progress 0. It is the mapper's signature as much as a dead engine's, so it does
not on its own say which failed -- the frame and write counts now printed
beside it are what separate them.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): count a value read under a unary operator as a read

`isWriteTarget` took any prefix-unary parent for a write, so `!hidden.value`,
`-offset.value`, `+x.value` and `~x.value` were dropped from the reads the
dependency array has to list. A style that gates on `!hidden.value` would have
passed the census while its mapper never listened to `hidden` -- the exact
freeze this file exists to catch, hidden by the check meant to catch it.

Only `++` and `--` mutate, so the prefix branch is narrowed to those two.
Postfix needs no narrowing: `++` and `--` are the whole set there.

Red-first with a negation fixture and a unary-minus fixture; the increment
fixture holds the other side, that a value only incremented is still not
required. Found by a review bot on #21592.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-19 03:47:44 -04:00
Jinwoo Hong 57fdf68ab3 feat(mobile): let the page hand a link to the device through the shell (OTA phase C, C2.3) (#21597)
* feat(mobile): answer externalLink on the shell side of the bridge (OTA phase C, C2.3)

A page has no way to open a URL outside itself: `Linking.openURL` is the
app's, and inside the shell the page is a document that cannot reach it. Adds
`notify { name: 'externalLink', url }` and a new grant name of its own in
`MOBILE_WEB_SHELL_GRANTS`, rather than a verb of `navigate` — `navigate`
opens a screen this app carries, this hands a URL to whatever the device
opens it with, and a shell implementing one and not the other is a real shell
the route policy has to be able to describe.

`https:`, `http:` and `mailto:` only, and broad inside that: any host, any
path, because a grant that named GitHub would grow a row per provider. The
rule is parsed rather than prefix-matched, since a scheme is what a URL
parser says it is and `startsWith('https:')` reads one out of
`javascript:alert("https://x")`. It is enforced at the frame as well as at
the page's call site, so a page that skipped its own check still cannot reach
the device handler. Bounded by the route href cap, per the ruling.

`BRIDGE_PROTOCOL_VERSION` is not bumped. Inert until a consumer exists: no
call site and no barrel is touched here.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): plumb externalLink through the bridge hook probe (OTA phase C, C2.3)

The hook's own suite builds its caller options inline, so the new required
option made it stop typechecking. `tsc -p tsconfig.json` excludes test files;
only the tests-typecheck ratchet saw it.

Adds the case that goes with it: a URL the page hands over reaches the
caller that can leave the app, and nothing reaches the navigate path.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): let the page post an externalLink over the bridge (OTA phase C, C2.3)

`notifyExternalLink(url)` on the page client, gated on the `externalLink`
grant and on the same scheme rule the frame enforces.

Checked twice on purpose. Nothing crosses back for a notify, so the boolean
is the only answer a tap gets: a page that posted a URL the shell's reader
then dropped would report "opened" into a frame nobody acted on, which is
precisely the dead tap the grant exists to rule out.

False before `init`, false after `close`, and never a throw — the callers are
tap handlers with no catch around them.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): add the external-link seam the tasks call sites will use (OTA phase C, C2.3)

One module with a `.web.ts` sibling, which is the shape every platform gap in
this bundle already takes. Native is `Linking.openURL` with the rejection
swallowed, because every caller is a tap handler and `openURL` rejects for a
URL no installed app claims. Web posts the `externalLink` notify after the
same scheme check the frame enforces, names its refusals and throws nothing.

The opener is published by the entry rather than read from context, for the
reason `publishPageStorage` is: the callers are plain functions in render
trees the provider does not wrap. A document that published none refuses
every URL, which is the right answer for a page with no shell.

Registered in `web-overrides.json`: inside the shell's WebView, react-native
-web's `Linking.openURL` opens the URL in that WebView and replaces the page
rather than handing it to the system browser.

No call site and no barrel is touched. The consumer PR swaps them onto this.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): state the real reason the page cannot open its own links (OTA phase C, C2.3)

The override reason claimed react-native-web's `Linking.openURL` opens the
URL in the shell's WebView and replaces the page. It does not. Read from
react-native-web 0.21.2: `openURL` calls
`window.open(url, '_blank', 'noopener')` and resolves whether or not anything
opened; only a `tel:` URL assigns `window.location`, and none of the three
allowed schemes is one.

The true failure is the worse one and the better argument for the verb. Both
shells refuse `window.open` outright, measured in their own sources: iOS sets
`javaScriptCanOpenWindowsAutomatically = false` and returns nil from
`WKUIDelegate`'s `createWebViewWith`; Android sets the same flag false, calls
`setSupportMultipleWindows(false)` and returns false from `onCreateWindow`.
So nothing opens, `openURL` resolves anyway, and the native path reports
success into a tap that did nothing — precisely the dead tap the grant exists
to rule out.

The seam's own comment now says the same, so the two files cannot drift.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): report a URL the phone could not open (OTA phase C, C2.3)

The shell swallowed `Linking.openURL`'s rejection. Nothing crosses back to
the page for a notify, so an open that failed — a `mailto:` on a phone with
no mail account — was silent on both sides. That is the one dead tap this
verb does not rule out, and it was the only one with no record at all.

Warned with the URL and the error, in the shape the two neighbouring reports
in this screen use, and still not rethrown: this runs on the native frame
handler. The contract now asks for the report rather than only for the
absence of a throw.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): forward the URL the parser read, not the string the page sent (OTA phase C, C2.3)

The scheme check reads the protocol through the WHATWG parser, which strips
tab, LF and CR from anywhere in a URL and trims leading C0 and space before
the scheme is visible. So `ht\ntps://example.com`, `https://example.com/a\r\n`,
`  https://example.com/a  ` and `https:example.com` all passed the check, and
both sides then forwarded the original string. Not a scheme escape — the
parser had already decided the scheme — but the device handler was given a
URL the check never looked at, which is a dead tap through an allowed URL.

`readBridgeExternalLinkUrl` answers the parsed href, and the page posts it
and the host forwards it. Normalizing rather than comparing, because
`https://example.com` differs from its own href by a path slash: refusing
what differs from its normalization would refuse an ordinary URL.

The cap now applies to the normalized form as well as the raw string, since
percent-encoding expands and a string inside the cap can leave it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): exercise the seam's default opener instead of a published one (OTA phase C, C2.3)

The case named for a document that published no opener published one first,
and `post` is module state every earlier case had already set, so the default
at the top of the module was never the thing under test. The only assertion
was `not.toThrow()`, which passes against any implementation.

`vi.resetModules()` and a fresh import, and the refusal reason is asserted.
Confirmed to discriminate: flipping the default to `() => true` turns this
case red and leaves the other three green.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): build the openURL rejection per call, not at mock setup

The case added for a failed open used
`mockReturnValue(Promise.reject(failure))`, which builds the rejected promise
at setup time. Nothing attaches a handler until the notify frame arrives
several awaits later, so the suite reported an unhandled rejection and exited
1 with every test passing — a red run that reads as green in the counts
alone.

A fresh rejection per call closes the window, and `openUrl` is reset between
cases so the mock cannot leak into one that does not expect it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): guard the openURL failure that escapes a tap handler (OTA phase C, C2.3)

`Linking.openURL` validates before it returns anything: `_validateURL` is an
`invariant` that throws for an empty string (react-native 0.83.10,
`Libraries/Linking/Linking.js:117-123`). So the seam's `.catch` was attached
to a promise that, in that case, never existed, and the throw went straight
through a tap handler — contradicting the module's own claim to be safe in
one.

Both failure modes are now caught and reported, and neither is rethrown. The
test double validates the way the real module does, because a mock that only
rejects cannot reproduce the failure that escapes.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): collapse the duplicate openURL wrapper onto the seam (OTA phase C, C2.3)

`mobile-pr-url.ts` was the seam's native body already, byte for byte, written
before it. It now re-exports the seam under its own name, so the empty-URL
guard and the failure report reach its four callers too.

Nothing changes natively: the seam's native form is what those callers were
running. On the web they would now post the notify instead, which is the
behaviour they should have had; it is unreachable today, since none of them
is in a page route's closure.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read externalLink end to end over the port pair (OTA phase C, C2.3)

The pair harness grew `externalLinks` and nothing read it. This is the
`navigate` twin's shape, over the six normalization inputs: what the page put
on the wire and what the shell forwarded are the same strings, read back off
the frames rather than recomputed, and nothing reaches the shell's client.

Two halves on purpose. The page normalizes before it posts, so over the
client the host only ever receives an already-normalized URL and forwarding
it raw would pass — the frame injected straight into the host at the end is
what holds the host to the rule on its own. Confirmed to discriminate:
reverting the host to forward `message.url` turns that assertion red.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(mobile): match the neighbours on three small inconsistencies (OTA phase C, C2.3)

Three of a kind, none behavioural:

- the notify guard's docstring had a 109-character line; reflowed
- the shell's could-not-open warning passed three arguments where the three
  other reports in that screen pass two; it now passes `{ url, error }`
- the native seam's suite built its rejection with `mockReturnValue`, which
  is the shape `52191bab02` removed elsewhere; it now builds one per call,
  through the same double that validates the way the real module does

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-19 03:33:41 -04:00