Accept text/json for JSON modules and use one is_json_mime classifier across
modules, documents, JSONPath, and resource inspection. Remove the module and
document aliases and recognize the seven registered application font types.
Use Fetch MIME extraction over the complete Content-Type header list for
script/style checks, workers, WebAssembly, ORB, and JSONPath. Derive classic
script charsets from the same MIME record, including charset inheritance and
resets when the MIME essence changes, while preserving BOM and fallback order.
Cover valid and invalid MIME types, header ordering, combined fields, quoted
commas, nosniff, and charset selection with regression tests.
Validation:
- cargo fmt --all
- cargo clippy --workspace --all-targets --all-features -- -D warnings
- cargo nextest run --no-fail-fast: 19,005 passed, 16 skipped
- Focused WPT: Moli passes 9 cases / 261 subtests, including the three JSON
regressions and all 31 script Content-Type extraction checks.
Give document decoding explicit HTML, XML, and text declaration policies.
XML ignores HTML meta elements, retains BOM and transport charset precedence,
and selects an initial XML declaration or UTF-16 signature before its UTF-8
default. Keep partial declarations and signatures buffered across chunks.
Cover conflicting meta declarations with and without an XML declaration,
including main documents, iframes, and external raw input. Adapt seven
Chromium XML/XHTML navigation scenarios and three decoder boundary tests
from Chromium a03603fe9af6230a12f1b2fb2c18a7d003a0d937; source paths are
recorded next to the tests. These include KOI8-R script inheritance, CP1251
XML, BOM-less UTF-16, supplementary Unicode characters, and BOM boundaries.
Validation:
- cargo fmt --all
- cargo clippy --workspace --all-targets --all-features -- -D warnings
- XML and Chromium regression tests passed.
- NEXTEST_TEST_THREADS=8 cargo nextest run --no-fail-fast:
18,980 passed, 16 skipped.
Keep valid inherited text encodings ahead of heuristic detection while
preserving BOM, transport charset, and JSON default precedence. Share the
response decoder between streaming main documents, asynchronous child
loads, data URLs, and blocking child snapshots.
Classify navigation XML MIME types independently of DOMParser's allowlist,
including +xml subtypes in parser routing and text exclusions. Add
regressions for inherited UTF-8, literal encoding declarations, XML
namespaces, and DOMParser rejection of navigation-only XML types.
Validation:
- cargo fmt --all
- cargo clippy --workspace --all-targets --all-features -- -D warnings
- NEXTEST_TEST_THREADS=8 cargo nextest run --no-fail-fast
18,952 passed, 16 skipped. Two existing timing tests failed under the
default concurrency, then passed individually and in the full run above.
Carry raw response header values through fetch, caches, redirects,
renderer delivery, workers, and protocol records. Convert explicitly
at WebIDL and protocol text boundaries while preserving opaque bytes.
Keep legacy CacheStorage and service worker metadata readable, and
advance the HTTP cache format for raw header values.
Validation:
- cargo fmt --all
- cargo clippy --workspace --all-targets --all-features -- -D warnings
- cargo nextest run --no-fail-fast
* fix(parser): preserve text document MIME and encoding semantics
* fix(document): share literal text parsing across child document paths
* fix(document): construct projection parser with the public API
A document whose only encoding declaration is `<meta charset="utf-16">`
was decoded as UTF-16, folding every byte pair into one CJK code point.
The damage was not confined to text: the tokenizer saw no markup either,
so `<title>` and `<p>` disappeared and the whole document collapsed into
a single text node inside `<body>`.
The prescan reaches a `meta` element only by reading ASCII-compatible
bytes, so a document it can see that declares UTF-16 has necessarily
mislabeled itself. The HTML Standard therefore rewrites the charset
before returning it:
If charset is UTF-16BE/LE, then set charset to UTF-8.
If charset is x-user-defined, then set charset to windows-1252.
Apply both rewrites in `moli-charset-parser`, which covers the
`meta charset` and the `http-equiv=content-type` paths at once. Keeping
them there rather than in `moli-encoding` leaves the two neighbouring
paths that legitimately select UTF-16 untouched: a UTF-16 BOM, and a
UTF-16 charset on the transport layer, which the encoding sniffing
algorithm takes with confidence certain and does not rewrite. Both are
now covered by regression tests, alongside the rewrites themselves and a
check that unrelated labels and labels outside the Encoding Standard are
still resolved exactly as before.
Verified on aarch64-darwin: `cargo test -p moli-charset-parser
-p moli-encoding` passes 55 tests, and `cargo fmt`/`cargo clippy
--all-targets` are clean for both crates.
Closes#152