Commit Graph
4 Commits
Author SHA1 Message Date
Jinwoo Hong a3c6d4266a fix(mobile): admit https: images on the web shell's CSP (OTA phase C, ruling 27) (#21964)
* fix(mobile): admit https: images on the web shell's CSP (OTA phase C, ruling 27)

Native markdown and the native rich editor load images the author referenced
by URL, so the page has to as well or a remote image is a blank where native
paints a picture. `img-src` widens to `img-src 'self' data: https:` on both
platforms; `script-src`, `connect-src`, `object-src`, `frame-src` and
`child-src` do not move.

`http:` stays out, and the pins say so directly rather than by absence: the
Kotlin test's blanket `!contains("http")` could not survive `https:`, so both
native pins now check `http:` (not a substring of `https:`) and check that
`https:` appears in `img-src` and nowhere else, the same shape the `data:`
pin already had.

No behaviour change on released phones: the shell ships in no released tag
(mobile-v0.0.9 predates it), so this reaches devices with the Phase E native
build and not before.

Neither native module has a CI job, so both ran locally: swiftc over the
module plus MobileWebShellChecks, and
`:orca-mobile-web-shell:testDebugUnitTest`. Both were confirmed red against
the old directive first.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): correct what the sealed preview frame is stricter about

The doc comment said the page was deliberately stricter than the native
preview because it loads no remote image and runs no script. Since `img-src`
gained `https:` only the script half is true: the frame loads a remote image
exactly as the native WebView does.

Says instead what an artifact's image URL now is -- a channel that fires on
view and carries whatever its author encoded, with nothing dynamic behind it
because no script runs -- and names `referrerPolicy` as what keeps the
document's own origin out of the request.

Comment only; no behaviour and no test moves.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): measure both halves of the preview frame's image fence

"fetches nothing of the artifact that leaves the origin" stopped being what
the sealed arm proves once `img-src` gained `https:`. The fixture's foreign
origin is `http://127.0.0.1`, so its two images are refused on the scheme
alone and only the font is refused by `font-src 'none'`. Renamed to say
exactly that.

The half that was missing is an https arm. Playwright route interception
answers an `https://…invalid` origin in the page, so the arm needs no TLS
server and no new dependency, and a request only reaches the handler if the
policy let it out. Under the shipped header, on Chromium and WebKit, the
`<img>` and the CSS background are both requested -- `img-src` governs a
background too -- and the font still is not.

`artifact()` takes the subresource origin; the links stay on the cleartext
one so no existing navigation case changes.

Red-first: with `img-src 'self' data:` put back into the parsed Kotlin
policy, the new arm fails on both engines with `expected [] to deeply equal
[ '/css-bg.png', '/img.png' ]`. The directive was restored byte-identical
before this commit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(scripts): split the preview frame's settling out of the render check

The https arm pushed mobile-web-app-html-preview-render.test.mjs to 620
counted lines, over the 600 cap config/scripts carries. Split at a module
boundary rather than bumped: the four wait-and-settle functions are rig
mechanics with no assertion in them, and they now sit beside the diagnosis
module they already reported through.

`waitForLoadedFrame` and `settleAfterMount` are the two the render check
calls; `waitForRecordedNavigation` and `settleWithoutNavigation` stay
internal to the new module.

Move only. Same 20 tests pass on both engines.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): send Referrer-Policy: no-referrer on the shell document

`img-src https:` gave the page somewhere to send a request, and the document
origin is `orca-mobile-web://<sessionId>/`, so a request that carries a
referrer carries the session id to whatever host an artifact or a markdown
document named.

`referrerPolicy="no-referrer"` on the preview iframe does not cover it.
Measured in the render rig against a permissive control policy: WebKit puts
the embedder's URL on a srcdoc frame's image request despite the attribute,
and Chromium sends none. So the guarantee belongs on the document, where one
header covers every request the page makes, and it rides the document alone
with the policy -- the referrer of a request is decided by the document that
made it, so on a subresource response it would govern nothing.

WKWebView under the custom scheme is unverified: the rig is Playwright
WebKit over http, not WKWebView over `orca-mobile-web://`. The header is the
hedge, and it costs nothing if that host never leaked.

Pinned three ways, each confirmed red first:
- Swift, exit 133 with the header removed.
- Kotlin, MobileWebShellResponseHeadersTest "sends the policy on the
  document" FAILED at :17 with it removed.
- The rig, through a new `readShellDocumentHeaders` that parses the Kotlin
  source the way `readShellCsp` does and throws rather than returning an
  empty map. With the value flipped to `unsafe-url` the WebKit arm fails
  `expected [ …(2) ] to deeply equal [ null, null ]`; with the line deleted
  the parse throws "could not parse the shell document headers".

The rig's arm carries its own presence precondition: a third server serves
the shipped policy with `unsafe-url`, so the WebKit reading is the header
doing the work, and Chromium's null either way is pinned as the browser's
behaviour rather than sold as evidence the header arrived.

MobileHtmlPreview.web.tsx said the iframe attribute kept the origin out of
the request. Corrected to name the header, since the measurement above is
what disproved it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): quote the current directive where the old text was written down

Three comments still read `img-src 'self' data:`, so a grep for the old
directive found live prose that no longer matches the header. Each stays
about `data:`, which is what those paths rest on; only the quoted policy
changes.

The two remaining hits in the repo are src/main/browser/doc-preview-protocol,
which is the desktop preview's own policy and not this one.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): name the surfaces img-src https: actually unblocks today

The comment justified `https:` with markdown and the rich editor, and
neither renders a remote image on the page. Verified in the tree:
MobileMarkdown paints `![](...)` as a tappable link at both of its image
branches and never mounts an Image, and it has no `.web` sibling, so that is
what native does too; MobileRichMarkdownEditor.web.tsx is a 92-line
multiline TextInput, still C7.6's plain source field.

What the directive unblocks today is four surfaces, none of them overridden
on the page:
- MobileAgentIcon's favicon, a hardcoded `google.com/s2/favicons` URL, used
  by thirteen callers including the session header and the worktree rows;
- MobileRepoIcon's project icon, a host-named favicon, avatar or upload, on
  the worktree list and the host workspace list;
- PRCommentCard's author avatar, from the review reply schema;
- the sealed HTML preview frame, which inherits the policy.

Markdown and the editor are named as the anticipated surfaces ruling 26
points at, so a later reader does not take the loosening as already covering
them. Both native pins carried the same wrong claim and are corrected.

That comment is the only record of why the policy loosened, so it says what
is true now and what is coming, separately.

Comment only: the parsed header is unchanged, checked through the harness
reader the render suite uses.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): point the new source-control route pin at the current directive

Merge resolution, not a conflict git could see. #21957 landed the
source-control and review page routes on main while this branch was open,
and its render check pins the directive text twice: `cspHeader` by substring,
which survives the widening, and the Swift source by the quoted literal
`"img-src 'self' data:"`, which does not. Two PRs green alone, red on the
merge.

Both pins now read the current directive.

One comment goes with it. "Not one request left the origin, so there is
nothing for the policy to have refused" now needs saying why: `https:` is
admitted, so an empty host list is these two closures fetching nothing
rather than the policy refusing something. The avatar that would fetch needs
provider data this page never gets, which the file's own closing note
already explains.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): wait for the admitted images before reading their hits

CI's Chrome 152 recorded the CSS background and not the `<img>` by the time
the bounded settle returned, so both https arms failed on a count: "expected
[ '/css-bg.png' ] to deeply equal [ '/css-bg.png', '/img.png' ]" and
"expected 1 to be 2". The reads were absence-shaped -- two frames and 200 ms
-- and the claim they carry is a presence.

So the arms wait for their own evidence, the way the `'refusal'` arm already
does. `frameReady: 'images'` polls until both admitted paths are recorded,
bounded by nothing but the case's own `ctx.signal`. It sits after the marker
wait, because an image is requested by a document that has parsed, and the
arm hands its reader in rather than the settling module reaching for state
that belongs to an arm.

One reader now serves the wait and the reading. An arm that waits on one
list and asserts on another has proved nothing about the list it asserts on.

The `/probe.woff2` absence is untouched and is now an absence standing
behind two presences rather than beside them.

What the wait prints when it does not arrive, captured by making the paths
unsatisfiable against a 12 s case:

  [html-preview-render] the arm recorded ["/img.png","/css-bg.png"] of
  ["/css-bg.png","/img.png","/never-arrives.png"]; #remote
  {"complete":true,"naturalWidth":1,
  "currentSrc":"https://artifact-images.invalid/img.png?n=n1",
  "loading":null}: arm csp=shipped sandbox=product frameReady=images
  nonce=n1 | browser 147.0.7727.15 | ... | frames [...]

`complete` with a zero `naturalWidth` is a request that finished and
produced no image; `complete` false is one still in flight. So a Chrome that
never issues the request says which of those it was, instead of a bare count.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): say why an admitted image never arrived, and hand back the context

CI's Chrome 152 read the `<img>` as complete with a zero naturalWidth and a
resolved currentSrc while the route handler never saw the request, and the
CSS background from the same origin did reach it. The diagnosis could say
the image failed but not why, because nothing was watching the request.

Now four sources are, for the `.invalid` origin only, in a module of their
own so the rig file stays under its cap: `request` says whether the page
asked at all, `requestfailed` carries the browser's `errorText`, and CDP's
`Network.loadingFailed` adds `blockedReason` and `corsErrorStatus`, which is
the only place a refusal names itself once the request never reaches a route
handler. `Network.requestWillBeSent` records the resource type, the initiator
and the frame, which separates an image the parser found from one nothing
asked for. They fill arrays while an arm passes and are only read on abort.

Proved by forcing the abort rather than assuming: with the awaited paths made
unsatisfiable, the reading names the font's refusal in both vocabularies at
once, `failed [{"url":".../probe.woff2","errorText":"csp"}]` and `cdp
loadingFailed [{"errorText":"","blockedReason":"csp",...,"type":"Font"}]`,
beside `cdp sent` showing every request's type, initiator and frameId.

Teardown: `open()` now takes an explicit context and closes both the page and
the context in a `finally`. The close used to sit on the happy path, so an
arm whose wait aborted and whose result reads then raced vitest's teardown
left its page and its implicit context open on a browser every later case in
that engine still runs on.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): correct three rationales the widening left wrong

(a) A review comment's avatar is not a surface the widening unblocks.
PRCommentCard renders it only under `Platform.OS !== 'web'` and a component
test pins the skip, so on the page it never renders. Dropped from both native
rationales and moved to the anticipated list beside markdown and the editor,
with the reason each is anticipated rather than current.

(b) The Kotlin rationale quoted the iOS origin. Android serves from
`https://<sha256(sessionId) first 32 hex>.orca-mobile-web.invalid/`, so a
referrer there carries a stable per-session handle and not the id itself,
while iOS serves `orca-mobile-web://<sessionId>/` and carries it verbatim.
Both are something an image host can key on across requests, which is what
the header is for; each file now names its own origin.

(c) "Only the script half of that is stricter than native" overstated it.
`font-src 'none'` and `connect-src 'self'` are stricter too. Images are the
one of the four that stopped being stricter, and the comment now says which
three remain and why.

A fourth, found while checking (a): the skip's own comment justified itself
with `img-src` being `'self' data:`, so a provider avatar would be "one
refused request per card". That is no longer true -- the avatar would load
now -- so the skip is a page capability gap rather than a policy consequence.
Recorded as such at the guard. Whether to lift the guard is a ruling-26
question and not this PR's.

Comments only. The parsed policy and document headers are unchanged, checked
through the harness readers the render suite uses.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): probe why Chrome never asks for the artifact image

CI's read was decisive: on Chrome 152 only the CSS background was requested,
while the `<img>` reported complete with a zero naturalWidth and a resolved
currentSrc. A request that went out and failed cannot produce both readings,
so the next probe asks the frame rather than the network.

On abort it now reads, inside the artifact frame: readyState, the init
script's own moment, document.images.length, every
`performance.getEntriesByType('resource')` name, the navigation entry types,
and for #remote its src, isConnected, complete, naturalWidth, currentSrc and
the outcome of decode(). A resource entry for a URL the rig never saw would
mean the request left the frame and died before reaching it.

Then it issues a `new Image()` at a URL that has never existed and reports two
seconds later whether the rig saw it. That splits the two live explanations: if
the fresh request is seen and the artifact's was not, the frame can fetch and
the parser-inserted element is the cause; if neither is seen, requests from
this frame are not reaching the rig at all. Subframe document commits are
counted from mount, because a second parse is a new window and leaves nothing
behind to count, and a second parse could be meeting a failure the first
cached.

`cdp sent` was empty on CI even for a request Playwright did record, so the
page's own session is blind to the frame. Chromium isolates sandboxed iframes
into their own process, srcdoc included, so flattened Target.setAutoAttach now
puts each child target on the same connection with Network.enable on the
child, and the attached list reports whether the frame is a separate target
at all.

The navigation arm gets the same reading, since CI showed it fails on its own
rather than behind the aborted image arms.

Verified by forcing the abort rather than assumed. Locally the reading prints
one subframe parse, decode resolved, every resource the document fetched, and
`fresh ... issued true seen true`, with the attached list empty, which is
consistent with this Chrome not isolating the frame and its page session
seeing the requests.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): time the artifact image against the frame's attachment

CI's second read showed the frame did issue the request -- it has a
resource-timing entry and decode rejected with EncodingError -- while the rig
saw only the CSS background, and a fresh image created later from the same
frame was both issued and seen. The remaining question is whether the entry
starts before anything was listening to that frame.

So the entry is now reported in full for the element under test:
responseStatus, transferSize, encodedBodySize, nextHopProtocol, startTime and
duration. A zero status with a zero transferSize is a fetch that reached the
network stack and came back with nothing, which is what an unintercepted
request looks like once `.invalid` fails to resolve.

Both sides of the comparison get a wall clock: `Target.attachedToTarget` and
Playwright's own `frameattached` now carry the moment they fired, and every
recorded request carries the moment it was seen. An entry that starts before
the attachment is the race stated rather than inferred.

Abort path only; the passing run is unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): serve the artifact's https assets from a real TLS listener

Interception could not measure what the directive admits. Chrome 152 isolates
the sandboxed srcdoc frame into its own target and the parser-inserted `<img>`
is the document's first fetch, issued before interception attaches there: the
request escaped to the real network, `artifact-images.invalid` did not
resolve, and the rig recorded nothing while the frame's own resource timing
showed the fetch and a later fresh image was both issued and seen.

So the assets come from a listener that is already accepting before the page
exists. It cannot be raced: the request arrives or it does not, and either
answer is the measurement. Hits and referrers are recorded server-side, the
way this rig's cleartext origin already does it, and read per arm by nonce.
`img-src 'self' data: https:` matches on scheme, so `https://127.0.0.1:<port>`
exercises the same directive as any other https host.

Lifecycle: started in beforeAll before any browser, closed in afterAll beside
the other servers. Its certificate is generated per run by openssl into the
suite's own scratch directory under `mobile/.tmp`, which the root gitignore
already covers and into which the server writes a second `.gitignore` as well;
the key never leaves that directory and nothing trusts it, since the context
is created with `ignoreHTTPSErrors`. No arm shares state: one hit list keyed
by each arm's nonce, and the permissive-Referrer-Policy control stays what it
was, a second bundle server serving the page, because the control is the
document's header and not the image host's.

The navigation record moves off interception too. It is now `page.on('request')`,
one subscription over every frame, armed after the rig's own `goto` exactly
where the route used to be registered; the route stays only for what only a
route can do, refuse the navigation. That answers the top-nav arm's `recorded
[]`: its record depended on the same per-target interception.

And the arms stop swallowing their clicks. `click(...).catch(() => {})` made a
tap that never landed and a tap that produced no navigation the same empty
counter; `open()` now records the error and the two top-nav arms assert it is
null before reading any count.

One correction to the reading added in the previous commit. The resource-timing
fields came back zero for a request that had plainly succeeded: they are opaque
cross-origin. The listener now sends `Timing-Allow-Origin`, after which
transferSize, encodedBodySize and nextHopProtocol carry real values.
`responseStatus` still reads zero on a successful request, so the comment names
the three that discriminate rather than the four that are printed.

24/24 on both local engines.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): compare the artifact fetch and the attachment on one clock

The early-or-late comparison spanned two clocks and could not answer the
question it was written for. Every `at` in the request log is Node's
`performance.now()`, counting from process start; the resource entry's
`startTime` is the frame's own, counting from that document's navigation. A
frame entry reads as earlier than a Node attachment by roughly the process
uptime, so the comparison would have reported the race as confirmed on every
run, including runs where there was no race. A green CI would not have caught
it.

So the comparison is stated where both numbers actually live: `asked` against
`attached` in the request log, on the Node clock alone. `startTime` and
`duration` stay, labelled as the frame's own account and explicitly not
comparable to an attachment time. The module docstring says the same, so the
next reading added here starts from the rule rather than rediscovering it.

The commit message of b5e82065f3 carries the same overstatement and is left
as it stands; this is the correction.

Also the stale route-handler references, now that the asset listener records
the secure origin and the navigation record is a page subscription. Three were
in the review; two more were not, and both were stale for the same reason:
`waitForRecordedNavigation`'s docstring still credited the route with
recording a main-frame navigation, which stopped being true when the record
moved off interception, and the request log described a refusal as one the
request never reached a route handler with. The route now only refuses; it
counts nothing. The one remaining mention is the deliberate contrast in the
rig that says the record is the page's event and not the route's.

Comments only. 24/24 on both local engines.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 08:34:38 -04:00
Jinwoo Hong 963da57cc3 test(config): preview rig readiness polls the main world, never the utility world (Chrome 152 hang) (#21963)
* test(mobile): wait for the preview frame in its main world, and probe the world that hung

Three cases spent their whole 180s on CI's Chrome inside `waitForSelector('#marker')` while
the diagnosis reported, from the same frame, `readyState: complete` and `marker: true`.
Those two readings ask in different worlds. `frame.evaluate` needs only the frame's main
execution context; a selector wait needs Playwright's injected script in Chromium's utility
world, an isolated world created per document by a command whose failure the driver swallows
and whose creation event it drops for a frame the driver considers stale. With `timeout: 0`
a world that never arrives is a wait that never ends.

So readiness is main-world polling now: the frame is resolved again from `page.frames()` on
every attempt and the predicate runs through `frame.evaluate`, still bounded by the case's
own `ctx.signal` and still ending in the diagnosis. The evaluate is abandoned after a second
so a frame that never answers cannot outlive its own replacement.

The diagnosis gains the reading that would have settled this in one run: a bounded
`utilityWorld` probe per frame, printed beside the main-world reading, so the split is
measured rather than inferred again. The competing explanation is ruled out in code --
Playwright closes a detached frame's scope with an error that every wait races, so a stale
Frame rejects rather than hangs.

Not proven red-first. Chrome 152 is the only engine that has shown this and it is not
available here; chromium 147 and WebKit 26.4 both build the utility world and both report
`utilityWorld "resolved"` for the sealed `srcdoc` frame. What is proven locally: 18 of 18 on
both engines, and a deliberately marker-less artifact still ends in the diagnosis, with
exactly one line per case naming the wait that hung.

That last part needed a fix of its own: an abort listener left behind by a wait that had
already resolved printed its stale reading at a later wait's timeout, so every timeout spoke
with more voices than it had hung waits. The listener is dropped on the way out.

In-frame `frame.click` needs the utility world too and is left alone: a main-world click is
not a user gesture, and the gesture is what those cases assert on.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): take the preview's refusal from the browser's report, not from a listener in the frame

The utility-world theory is refuted: CI's Chrome answered `utilityWorld "resolved"` on both frames
and the case failed anyway, with the widened frame reporting the artifact parsed, the CSS
background's `img-src` refusal recorded, no `script-src` refusal, and no script run. Two different
things produce exactly that reading. The policy refused the script and the frame's own listener was
not there to see it, or the sandbox refused it first, which raises no violation at all -- and a
listener inside the frame cannot tell them apart, because in the second case there is nothing for it
to hear.

So the evidence moves to where neither depends on timing: the sealed server now appends `report-uri`
to the policy it serves, carrying the arm's nonce, and the rig records what the browser reports. The
override arm's precondition is a `script-src` report from this arm's frame, waited for under
`ctx.signal` and ending in the diagnosis. Measured on both engines: a widened frame is reported for
`script-src` and a sealed one never is, while both are reported for the image the policy refuses. So
the sealed arm now waits for its own `img-src` report, which turns "no script-src refusal here" from
an unguarded absence into one measured beside a presence.

`report-uri` is additive -- it names where a report goes and changes nothing about what is enforced
-- and the first case now pins that by splitting the served header and asserting the rest is the
shipped Kotlin text exactly.

The in-frame collector stays, for the diagnosis only, and it now carries the readings that would
have answered the ordering question in one run: the init script records when it ran in each frame,
the artifact's script records the same on the document element, and the diagnosis prints both. What
the artifact wrote moved off `window` entirely for the same reason -- a page init script owns the
window of every frame it reaches. Locally the init script precedes the artifact's by one
millisecond, in every arm on both engines; the ordering on Chrome 152 is now a reading rather than a
hypothesis.

A measurement worth keeping beside the code: in a frame with no `allow-scripts` the init script runs
and its array exists, and no violation event is ever delivered to it, while the browser reports the
same refusals to the server. That is why the old `violations` assertions could not have caught this.

Red-first, all three locally: with report recording off, with the report endpoint not appended, and
with `script-src` reports alone dropped, the preconditions time out into the diagnosis and the
served-policy assertion reds too. 18 of 18 on both engines, three runs.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): cover the navigation wait's sampling branch, and bind the load wait to the case

Two findings from the bots on the rig, both real.

The navigation wait's five-second sampling branch called `describePreviewFrame` after the import
that supplied it had gone. It fires only when an arm is slow, the name is evaluated before `.catch`
can attach, and `no-undef` is off, so nothing in the file or the lint run had ever executed that
line. Fixed by moving the settle waits into the readiness module, where the call sits beside the
import it needs rather than a file away from it -- the split is what let the reference dangle.

The proof is a case that drives the branch: a navigation the arm will never see, a sampling interval
passed in, and the case's own abort ending it, asserting on the reading it printed rather than on
its own absence of an error. Red-first, with only that branch's callee renamed: 2 failed, 18 passed,
`ReferenceError`. So the case covers the branch and nothing else in the file did.

The load-only arm's `frame.waitForLoadState('load')` was the one wait left that did not observe
`ctx.signal`; after an abort it kept waiting on its own timeout. It is a main-world poll on
`document.readyState` now, re-resolving the frame each attempt like every other wait here, and it
ends in the diagnosis.

20 of 20 on both engines, twice.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 06:16:40 -04:00
Jinwoo Hong a08a4915b3 test(mobile): wait for the frame's script-src refusal before reading it (#21948)
securitypolicyviolation is dispatched as a queued task, so its order against
the frame's load is not guaranteed. On the runner's Chrome the list held only
the blocked background image when the arm read it, and the assertion failed on
an entry that had not landed yet; this PR's own run passed by timing.

The header-keeping arm now waits for its own evidence, bounded by the case's
signal like every other wait here, and the assertion is unchanged. A frame
that was never widened raises no such entry at all, so the wait is the arm's
precondition rather than a convenience, and it ends in the rig's diagnosis
rather than a bare hang.

Red-first, isolated to the new wait: with nothing in the artifact for the
policy to refuse, the case fails at its timeout and prints "the frame never
reported a script-src refusal" with the arm, the browser and the frame's own
readings.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 04:21:43 -04:00
Jinwoo Hong 8d42410e01 feat(mobile): render the HTML preview on the page in a sealed srcdoc frame (OTA phase C, C7.10 A) (#21862)
* feat(mobile): offer a cancelled top-frame navigation to the shell's opener

Both shells cancelled every navigation off their own document in silence: iOS
`decidePolicyFor` allowed only `isMainFrame && isDocumentUrl`, Android's
`shouldOverrideUrlLoading` dropped anything whose resolved path was not "/".
Nothing opened. That is the whole of ruling 29's "if they do not": a user tapping
a link inside C7.10's sealed HTML-preview frame reaches the top frame as a
navigation request, and the shell was the only thing that could act on it.

A cancelled main-frame navigation now reaches JS as `onExternalNavigation` and
goes through the same `Linking.openURL` the `externalLink` notify already uses.
The scheme list is not restated natively: the native side caps the string and
says which frame it came from, and `readBridgeExternalLinkUrl` decides what opens
in the half that ships over the air. A subframe navigation is never offered,
because that is the sealed preview loading itself.

swiftc check: OK (`checkCancelledNavigation` added, the whole suite runs).

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): render the HTML preview in a sealed srcdoc frame on the page

C7.6 gave the page the artifact's source, which is the native component's Source
tab and half its job (ruling 8). Ruling 26 makes that debt: the Preview tab comes
back as an `<iframe sandbox srcdoc>` inside the page's own document.

`srcdoc` rather than a `blob:` URL, and no CSP change at all. Measured on Chromium
and WebKit: a `srcdoc` frame has no URL for `frame-src` to match and inherits its
embedder's policy instead, so it is admitted under the shipped `frame-src 'none'`,
while a `blob:` frame is refused by `frame-src` on both and refused a second time
in WebKit by the `frame-ancestors 'none'` it inherits.

Two independent fences seal it, and the render check measures each on its own:
the sandbox grants neither `allow-scripts` nor `allow-same-origin`, and the
inherited `script-src 'self'` refuses the artifact's inline script even when a
control arm grants `allow-scripts`. The inherited `img-src` and `font-src 'none'`
govern its subresources, against a no-header control where the same three are
fetched.

`allow-top-navigation-by-user-activation` is the one token granted (ruling 29), so
a tapped link becomes one top-frame navigation the shell now opens externally,
while a `<meta refresh>`, a form submit, `target="_blank"` and any script-initiated
navigation produce none.

`lucideBarrelPlugin` is exported from the bundle builder so the check builds the
toolbar's icons the way the page does rather than carrying a second shim.

config/scripts suite, this file: 14 passed, 0 errors, exit 0. Control runs: a
literal `sandbox` in the JSX reds 4, an added `allow-scripts` reds the script
fence and the token census.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the preview's sealed frame where the degradation was pinned

The three HTML-preview cases in this file described the state ruling 26 retires:
no toggle, no frame, the source only. They now pin the frame's shape through the
test renderer -- the artifact reaches it as `srcDoc`, the sandbox grants neither
`allow-scripts` nor `allow-same-origin`, both toggle positions exist, and Source
takes the frame away with it -- and the "never renders the html itself" case
becomes "never puts it anywhere but the frame", counted rather than merely absent.
What a browser does with that frame stays in the render check, which is the only
thing that can answer it.

The rich Markdown editor's half is unchanged: it is still the plain field, and
item C is a later PR.

Two mocks added: `Pressable`/`ScrollView` on the react-native double, because the
toggle renders one, and `lucide-react-native`, whose barrel imports a
`LucideProvider` its own context module does not export and so does not load under
vitest at all.

9 passed, exit 0.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): refuse a link-activated top-frame navigation, even to the document

F1, blocking, with F5 and F6 folded in because they are the same decision and
splitting them would mean three rewrites of one function.

F1: `<a href="/" target="_top">` and `href=""` in an artifact resolve against the
embedder's base, so both named the shell's own document URL -- which both shells
ALLOWED (iOS `isDocumentUrl`, Android's path `/`). One tap inside the sealed
preview reloaded the shell's page: bridge target cleared, load state restarted,
page state gone. A navigation a human started is now never allowed, whatever it
names; it is offered instead, and `cancelledShellNavigationTarget` drops
`orca-mobile-web:` in silence exactly as it drops `/h/other`. The page rewriting
its own path carries no gesture and is still allowed.

F5: the OFFER is gated on the same gesture, so a top-page meta refresh or a
redirect is cancelled and never opened externally.

F6: iOS returned early on `shouldPerformDownload` before the offer, so `<a
download>` was dead on iOS and opened on Android. The early return goes; a
download is refused rather than allowed when nothing started it, and a
gesture-started one reaches the opener on both platforms.

The allow half and the offer half are now one function per platform
(`MobileWebShellNavigationPolicy.verdict`, `mobileWebShellNavigationVerdict`), so
they cannot drift. The gesture is the platform's own answer: `.linkActivated` on
iOS, `request.hasGesture()` on Android.

Native tests, both platforms: document URL + gesture refused and offered; document
URL without gesture allowed; foreign + gesture cancelled and offered; foreign
without gesture cancelled and silent; download both ways; subframe never offered.
swiftc OK; control run with the gesture rule removed exits 133. Gradle
MobileWebShellDroppedNavigationTest tests=8 failures=0 errors=0.

Also corrected: the screen comment that claimed the document's own reloads reach
the handler (they never do), and the prop doc, which now states the gesture rule.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): count own-origin top-frame navigations, and drop the goto cap

F2: `page.setDefaultTimeout(4000)` capped `page.goto` at 4 s while every sibling
render check uses the 30 s default, so under load the first WebKit cases redded on
the navigation rather than on anything they assert. The cap goes; the per-action
timeouts that needed to be short are already passed at their call sites.

F1's page-side half: the rig now routes the page's own origin as well as the
foreign one and counts main-frame navigations to each separately, with two cases
pinning that `href="/"` and `href=""` each produce exactly one own-origin
top-frame request. Playwright is not the shell, so what these state is the request
the shell is handed; refusing it is the native tests' job and the docstring names
which ones. The own-origin route is registered after the initial load, because it
aborts main-frame navigations and the first `goto` is one.

The foreign-tap and meta-refresh cases now also assert zero own-origin
navigations, so a fix that merely moved the target would not pass.

16 passed, exit 0, no Errors line.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): wait for the preview frame's own load, never a clock

CI read the child frame before its srcdoc committed: frameUrl came back ''
and the control arm's script as not yet run. The frame list, the frame's URL
and anything read inside it settle at their own moments, and a 900 ms wait
reads whichever of them has happened -- on a loaded runner, none.

Polls for a child frame at about:srcdoc with its load fired, bounded by the
case's own timeout, and an override arm now resolves on the document its
srcdoc assignment commits rather than on the assignment.

Red-first: with a 2.5 s mount delay standing in for a loaded runner, the
paint case failed on both engines before this and all 16 cases pass after.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say whose violations the preview rig reads

The list is the main frame's: securitypolicyviolation does not cross into a
frame, so an empty one says the embedder raised none and says nothing about
the artifact's own style, image or font. A listener inside the frame cannot
be the fix -- the fence under test is that nothing in the artifact runs.

So the comment now claims what the reading supports, and names where the
frame's containment is actually measured: the pixel for its inline style,
the counting server for its img-src and font-src.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): announce which side of the preview toggle is showing

The Preview/Source pair carried a label each and nothing else, so which one
was showing lived only in the active background -- invisible to a screen
reader on both surfaces. Each button is now a tab carrying its selected
state, inside a tablist, and the two files' toolbars stay character-identical
so the page and the phone announce the same thing.

Red-first: the new case renders both siblings and failed on both for the
missing role before this.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): type the WebView mock like the file's other hosts

The anti-slop gate refuses a bare `object` parameter. Takes the same shape as
the react-native mocks beside it, which pass it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): allow only the load the shell itself started

The document URL was allowed whenever the host reported no gesture, so a
navigation the shell never asked for could reload the page out from under the
session. Measured against a real WKWebView off-device: a sandboxed subframe
navigating the top frame to the document URL arrives as `.other` with no
gesture at all, and Chromium's own docs allow hasGesture() to be false for a
request a human started. Census first: nothing in the page navigates the top
frame -- no location assignment, reload, replace, window.open or form -- the
router moves by pushState and replaceState only, so the rule needs no gesture
and no page cooperation.

Both shells now raise a flag around their own load and drop it at commit, and
allow a main-frame navigation only while it is up. Everything else naming the
document is refused and never offered, since offering it would send the user
out of the app. iOS carries the second discriminator the same probe measured:
sourceFrame is the main frame for the shell's own load and the subframe for a
subframe's top navigation, so a subframe can never take the allow path.

Red-first: the Swift checks and the Kotlin tests were written first and failed
to compile against the old signature. 9 Kotlin tests, 54 in the module.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): point the meta-refresh arm at the embedder's own URL

The fixture pointed off-origin, so its own-origin assertion could not move
whatever the frame did. The new arm refreshes to `/`, which resolves against
the embedder's base, and pins zero top-frame requests on a counter the
`href="/"` case proves reads 1 in the same rig.

It also counts what the frame asks for itself, with a presence control that
attributes the fence: with `allow-same-origin` and no policy the same fixture
navigates the frame to the embedder's `/`, and with the policy dropped but the
product's token kept it navigates nothing, so the opaque origin is what
refuses it rather than the CSP.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read what an action produced, not what a clock allowed

The 600 ms after every action is gone. An arm that expects a navigation now
returns the moment the route handler records it, with a deadline only so a
click that missed its target says so instead of spending the case's timeout.
An arm that expects none waits for two painted frames inside the page and one
200 ms drain for the popup queue, which is a browser-process event with no
in-page counterpart; the docstring says why that one is bounded.

Measured and reported rather than claimed: with the new wait replaced by a
no-op every arm still passes, because the reads that follow are each a round
trip. It is insurance against the runner load that produced the frame-commit
race, not a fix for a failure seen here.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): take the settling branch as a ternary

What oxlint's prefer-ternary asks for, and the changed-code gate with it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): find the preview frame by its element, not its URL

CI timed out on all seven preview cases on one engine: the poll waited for a
child frame whose URL reads about:srcdoc, and that browser reports an empty
URL for a srcdoc frame, so every case ran to its own timeout. The same
difference had already shown as `expected '' to be 'about:srcdoc'`.

The frame is now the element: waitForSelector('iframe') then contentFrame(),
with readiness taken from the fixture's own marker inside it. Nothing compares
a frame URL any more -- the paint case reads the element's srcdoc attribute
and the absence of src instead, which is what "parsed inside the frame rather
than fetched into it" actually means. The one arm whose artifact navigates the
frame away says so rather than waiting for a marker that is not coming.

Red-first: with the old poll keyed on a URL the browser never reports, both
engines time out exactly as CI did; the new wait passes 18/18 with the 2.5 s
mount delay still injected.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): make a frame that never becomes ready say what it saw

The runner's Chrome read the preview frame's URL as empty where three
chromium builds here read about:srcdoc: bundled headless, the headless shell,
and --headless=old, all 147. So the difference is not reproducible locally and
the next CI run has to carry its own diagnosis.

The marker wait is bounded well inside the case timeout, and on expiry it
reports the frame's URL, the srcdoc attribute's length and the page's CSP
violation list -- which separates a frame the policy refused from one that was
merely slow, the two readings that look identical from a timeout.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): run the containment arms the comment only claimed

The comment said the fixture navigates nothing with the policy dropped and
the product's token kept, but no arm ran it: the control dropped both fences
at once. Both single-fence arms exist now, either of which would hold.

Measured rather than assumed, and one of them is not what the comment said.
The token alone: the navigation never starts, no request, no violation. The
policy alone, with allow-same-origin granted: the navigation does start and
frame-src refuses it, which the embedder reports as its own violation. The
engines differ only in what is left in the frame -- chromium an error page,
WebKit the artifact -- so neither is asserted; what is asserted is that the
request never reaches the server.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): refuse a download that names the shell's own document

The document branch skipped downloads, so `<a href="/" download>` fell
through to the offer path carrying the shell's own URL. Harmless in practice,
because the opener's scheme list drops it, but it contradicted the policy's
own comment and the prop doc, and it left the one URL that must never be
offered reaching the boundary.

The branch now covers a download too: refused, from either frame, gesture or
not, and never offered. A gesture-started download of anything else still
reaches the opener.

Red-first on both platforms: the Swift checks exited 133 and the Kotlin row
failed against the old policy. 10 navigation tests, 55 in the module.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): drop the own-load flag wherever a document ends

The flag lived beside the load call and had to remember every ending
separately, so iOS missed two: a prop update that fails before it loads, and a
renderer that died. Both left it raised, and a navigation to the document URL
during that window would have been allowed.

It now lives in the load state machine, which every ending already runs
through -- a commit, a failure, a dead renderer, a prop update, a reset -- on
both platforms, so there is nothing left to remember. The view raises it and
reads it, and drops it nowhere.

The Android residual is stated in the policy rather than papered over: between
loadUrl raising the flag and onPageStarted dropping it, a navigation to the
document URL from inside the preview frame would be allowed, because that
callback says nothing about which frame asked and no host discriminator
exists. It needs a generation switch and a tap in that window; iOS closes the
same gap with sourceFrame.

Red-first: the new Swift row failed to compile and the Kotlin row with it.
12 load-state tests, 56 in the module.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): spend the own-load flag on the allow, not on the commit

The flag stayed raised from the load until didCommit, so a second main-frame
action naming the document inside that window was allowed too and replaced the
document. WebKit can decide a second action before the first one starts, so
the commit is too late to be what spends it.

The allow itself spends it now, before the decision goes back, and every
ending still drops it for a load that is allowed and never commits.

Red-first: the new check composes the machine with the policy -- the seam the
flag and the rule meet at -- and failed to compile against the old machine.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): stop raising an own-load flag Android never consults

WebViewClient's javadoc, verbatim: "This callback is not called for all page
navigations. In particular, this is not called for navigations which the app
initiated with loadUrl(): this callback would not serve a purpose in this
case, because the app already knows about the navigation."

So the flag guarded nothing on this platform and, while raised, was the one
thing that could have let a competing request to the document URL through.
The view passes isShellLoad = false always now, the machine drops the field it
had no raiser for, and the policy comment carries the quote. Nothing reaching
that callback is the shell's own load, so nothing naming the document is
allowed there at all -- which also closes the generation-switch window the
residual named, so that paragraph goes.

No red to show: this is a removal, and the behaviour it leaves is the refusal
the existing rows already pin. What a device proof must check is stated in the
policy instead: a WebView that did route its own load here would have it
refused and the load state would sit at loading. 55 tests in the module.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): settle every arm, not only the ones that tap

An arm with no action read its counters as soon as the frame's marker
appeared, so a zero-delay meta refresh could dispatch after the reading. The
arms that pin zero were the ones relying on it.

Every arm settles now, and what it settles on is what it expects: the sealed
refresh arms take the bounded no-navigation path, and the loose arm waits for
a recorded navigation that is neither main-frame nor foreign -- its own
frame's -- rather than the main-frame wait it would never satisfy.

Red-first: with the settling removed and the refresh moved to 2 s, the loose
arm reads 0 on both engines; with it back, 1 on both, the delay still in.
A 0.4 s refresh passes either way, which is why the finding was invisible.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): wait for what the artifact's script wrote, not for the element

The two-fences control asserts the inline script ran, and the marker element
it waited for exists from parse time, so the arm could read window.__ran
before the script had touched it. Under a loaded runner that reads 0, which is
CI's "expected +0 to be 1" on chromium.

Readiness is now per-arm: 'script' waits for the script's own write, 'load'
for the arm whose artifact navigates the frame away, 'artifact' for the rest.

Red-first: with the inline script's write delayed 1.5 s, the old arm fails on
both engines with that exact message and the new one passes, delay still in.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): bound the rig's waits by the case timeout and nothing else

Two inner deadlines, 20 s and 15 s, were racing the outer one they sit
inside, so a slow runner could fail a case on a number this file picked
rather than on the one the case declares.

Both now run to vitest's own `ctx.signal`, which aborts when the case times
out. On abort the rig prints its reading -- the frame's URL, the srcdoc
length, the violation list, or the navigations it did record -- and lets the
case fail as the timeout it is. Nothing is rethrown from that path: a
rejection raised after vitest has given up on a case has nobody left to catch
it, and an unhandled one fails a run whose every test passed.

Red-first: with the marker selector pointed at an element that never appears
and the case timeout cut to 8 s, the diagnostic prints and the case fails as
`Test timed out in 8000ms` rather than hanging in silence.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): ask a stuck preview frame everything it can still answer

The old diagnostic said only that a frame never parsed, and its violation
list was the top document's -- securitypolicyviolation does not cross frames,
so it said nothing about what the frame itself refused.

It now prints the browser version, the arm it came from, the iframe element's
srcdoc length and sandbox, contentDocument.readyState and contentWindow.href
(which answer for a same-origin arm and report `refused` for an opaque one,
so the arm's own origin is in the log), and every Playwright frame with its
url, name, readyState, body length, marker presence, window.__ran and its own
violations. Per frame, because the page's init script installs the collector
in every frame -- measured on both engines -- and CDP evaluates inside an
opaque frame whose scripts are blocked.

Two corrections that the local probes forced. The reading is sampled while
waiting and printed from the last sample: read at the abort it lost its race
with vitest's teardown and printed nothing at all. And two arms had never been
given the case's signal, so their waits could not be bounded or diagnosed.

The diagnosis moves to its own module because the test file is at its line
limit, and because the bound and the reading it prints are one thing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): build a widened control frame instead of relaxing a live one

A live frame cannot be relaxed. Sandbox flags are fixed on a browsing context
when it is created, and Chrome 152 keeps the original ones through a srcdoc
reassignment while still parsing the new document -- so the control arms that
widened the product's own frame stayed sealed on the runner, and CI read a
script that never ran and a refresh that never navigated. Chromium 147 here
honours the relaxation, which is why it passed locally for a year of runs.

The override now clones the element, sets the sandbox on the clone, gives it
the artifact and replaces the product's frame with it, so the widened flags
are there from creation -- the way the product does it, since React sets the
attribute before insertion and never after. The product's own arms are
untouched: a null override still returns immediately.

And the control can no longer pass for the wrong reason on any engine. The
header-keeping arm now reads the violation raised inside the frame: a
script-src refusal can only happen if the sandbox let the script start, so it
separates "the policy held" from "the frame was never widened", which the old
arm could not. The loose arm pins an empty list beside it, the sealed arm pins
an empty one too, and those three readings are the whole fence story. The
violations come from each frame's own collector, because the embedder never
sees them.

Two diagnostic repairs the local probes forced: the browser version is read
once at open, since asking at the abort printed "browser unknown" in the CI
log this exists for, and the reading is sampled immediately as well as every
five seconds, since a wait that only prints "no reading was taken" says
nothing.

Red-first: with the widening disabled, both engines fail exactly as CI did --
180 s timeouts on the script arm -- and the diagnostic names the arm, the
version and the sandbox it actually had.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): put the toggle's selected state where a browser reads it

CodeRabbit is right, and the browser says so: react-native-web's createDOMProps
never reads accessibilityState, so on the page the tab pair emitted role="tab"
and no aria-selected at all. The test renderer could not see it, because it
reports the props the component was handed rather than the DOM they become.

Both siblings carry aria-selected beside accessibilityState now -- the phone's
screen reader takes the latter, the browser the former -- and the toolbars stay
character-identical.

Red-first, in a real browser on both engines: the rig now reads every
[role="tab"] element's aria-selected before and after the tap, and it read null
for both positions before this line existed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 03:32:34 -04:00