Move public-web seeds into CSV fixtures and share scheduling, command construction, classification, evidence, and artifact handling across top-sites and wild-web suites.
Bind Chromium DCL and response evidence to the current navigation loader, surface navigation errors promptly, rotate target-first order across repeated runs, and normalize HTTP and timeout outcomes so cross-engine results remain comparable.
Apply one absolute deadline across response headers, streaming raw bodies, page creation, and later readiness waits. Report the active timeout phase without extending the caller's budget and reject raw documents from Page APIs before an unusable body can stall.
Give CLI fetch failures a stable two-line Error/Reason presentation while retaining full anyhow chains for non-fetch failures, including raw-document page-wait errors.
Keep physical stylesheet fetches and import graph work shared while installing CSSOM state and terminal events for every link owner. Settle completed clients at admission, retain completed import graphs for late owners, and validate document ownership before asynchronous installation.
This replaces the abandoned turn-exit replay path with an explicit shared-fetch/per-owner lifecycle model and adds coverage for document.write(), data URLs, imports, failures, node mutation, and document replacement.
`<meta http-equiv="Content-Type" content="...">` was parsed as a MIME
parameter list: split on `;`, skip the media type, compare each parameter
name for equality with `charset`, then strip quotes with `trim_matches`.
The HTML Standard instead specifies a keyword search over the whole
attribute value, and that difference is observable in both directions.
Declarations that were dropped, each falling back to windows-1252 and
rendering as mojibake:
charset=utf-8 (no media type at all)
text/html; charset=utf-8 profile=x (label ends at whitespace)
text/html; xcharset=utf-8 (keyword matched as a substring)
Declarations that were wrongly accepted, because `trim_matches` strips
any number of either quote from both ends:
text/html; charset='utf-8 (never closed)
text/html; charset="utf-8' (closed by the other quote)
The first group is the one that shows up in the wild — a `content`
attribute carrying no media type is common on legacy pages, and Moli
silently ignored it.
Implement the algorithm as specified: search for the literal `charset`
from a moving position, allow whitespace either side of the equals sign,
resume the search past a keyword that is not an assignment, and take a
quoted label only when the same quote closes it, otherwise ending an
unquoted label at the first ASCII whitespace or `;`.
Two details worth noting for review. The search resumes at the end of a
non-assignment keyword, which strictly advances and cannot loop. And all
byte indices land on ASCII bytes, which never occur inside a multi-byte
UTF-8 sequence, so slicing the latin1-mapped prescan input stays on
character boundaries.
An empty label now ends the extraction rather than continuing to a later
`charset=` in the same value, matching step 6, and that case is pinned by
a test.
Verified on aarch64-darwin per AGENTS.md: `cargo fmt --all`,
`cargo clippy --workspace --all-targets --all-features -- -D warnings`,
and `cargo nextest run --no-fail-fast` (16392 passed; the 7 failures are
pre-existing on 6f5aa320 and reproduce identically with this change
reverted).
Closes#156
Split screenshot and screencast capture contracts, publish one shared visual resource generation, and collect CSS image references during the real layout traversal. Add asynchronous publication and multi-session screencast regressions.
A document whose only encoding declaration is `<meta charset="utf-16">`
was decoded as UTF-16, folding every byte pair into one CJK code point.
The damage was not confined to text: the tokenizer saw no markup either,
so `<title>` and `<p>` disappeared and the whole document collapsed into
a single text node inside `<body>`.
The prescan reaches a `meta` element only by reading ASCII-compatible
bytes, so a document it can see that declares UTF-16 has necessarily
mislabeled itself. The HTML Standard therefore rewrites the charset
before returning it:
If charset is UTF-16BE/LE, then set charset to UTF-8.
If charset is x-user-defined, then set charset to windows-1252.
Apply both rewrites in `moli-charset-parser`, which covers the
`meta charset` and the `http-equiv=content-type` paths at once. Keeping
them there rather than in `moli-encoding` leaves the two neighbouring
paths that legitimately select UTF-16 untouched: a UTF-16 BOM, and a
UTF-16 charset on the transport layer, which the encoding sniffing
algorithm takes with confidence certain and does not rewrite. Both are
now covered by regression tests, alongside the rewrites themselves and a
check that unrelated labels and labels outside the Encoding Standard are
still resolved exactly as before.
Verified on aarch64-darwin: `cargo test -p moli-charset-parser
-p moli-encoding` passes 55 tests, and `cargo fmt`/`cargo clippy
--all-targets` are clean for both crates.
Closes#152
Use a synchronous prepare/commit/apply boundary for connected stylesheet lifecycle authority, so DocumentRuntime no longer acquires it through raw ContextHost pointers. Keep modulepreload network fetches identity-only and publish terminal link events without a Document load-delay lease.