Commit Graph
22 Commits
Author SHA1 Message Date
ldm0 ce2edbb9e1 fix(mime): unify JSON/font types and response header extraction
Accept text/json for JSON modules and use one is_json_mime classifier across
modules, documents, JSONPath, and resource inspection. Remove the module and
document aliases and recognize the seven registered application font types.

Use Fetch MIME extraction over the complete Content-Type header list for
script/style checks, workers, WebAssembly, ORB, and JSONPath. Derive classic
script charsets from the same MIME record, including charset inheritance and
resets when the MIME essence changes, while preserving BOM and fallback order.

Cover valid and invalid MIME types, header ordering, combined fields, quoted
commas, nosniff, and charset selection with regression tests.

Validation:
- cargo fmt --all
- cargo clippy --workspace --all-targets --all-features -- -D warnings
- cargo nextest run --no-fail-fast: 19,005 passed, 16 skipped
- Focused WPT: Moli passes 9 cases / 261 subtests, including the three JSON
  regressions and all 31 script Content-Type extraction checks.
2026-09-29 23:17:46 +08:00
ldm0 14cc9792ef fix(encoding): exclude HTML meta from XML encoding selection
Give document decoding explicit HTML, XML, and text declaration policies.
XML ignores HTML meta elements, retains BOM and transport charset precedence,
and selects an initial XML declaration or UTF-16 signature before its UTF-8
default. Keep partial declarations and signatures buffered across chunks.

Cover conflicting meta declarations with and without an XML declaration,
including main documents, iframes, and external raw input. Adapt seven
Chromium XML/XHTML navigation scenarios and three decoder boundary tests
from Chromium a03603fe9af6230a12f1b2fb2c18a7d003a0d937; source paths are
recorded next to the tests. These include KOI8-R script inheritance, CP1251
XML, BOM-less UTF-16, supplementary Unicode characters, and BOM boundaries.

Validation:
- cargo fmt --all
- cargo clippy --workspace --all-targets --all-features -- -D warnings
- XML and Chromium regression tests passed.
- NEXTEST_TEST_THREADS=8 cargo nextest run --no-fail-fast:
  18,980 passed, 16 skipped.
2026-09-29 18:01:12 +08:00
ldm0 a3af75845d fix(document): preserve inherited encoding and XML MIME routing
Keep valid inherited text encodings ahead of heuristic detection while
preserving BOM, transport charset, and JSON default precedence. Share the
response decoder between streaming main documents, asynchronous child
loads, data URLs, and blocking child snapshots.

Classify navigation XML MIME types independently of DOMParser's allowlist,
including +xml subtypes in parser routing and text exclusions. Add
regressions for inherited UTF-8, literal encoding declarations, XML
namespaces, and DOMParser rejection of navigation-only XML types.

Validation:
- cargo fmt --all
- cargo clippy --workspace --all-targets --all-features -- -D warnings
- NEXTEST_TEST_THREADS=8 cargo nextest run --no-fail-fast
  18,952 passed, 16 skipped. Two existing timing tests failed under the
  default concurrency, then passed individually and in the full run above.
2026-09-29 18:01:12 +08:00
lanyue-llk 9d236a20ef fix(parser): preserve text document MIME and encoding semantics 2026-09-29 18:01:12 +08:00
ldm0 6d08975a69 refactor(fetch): preserve response header bytes end to end
Carry raw response header values through fetch, caches, redirects,
renderer delivery, workers, and protocol records. Convert explicitly
at WebIDL and protocol text boundaries while preserving opaque bytes.

Keep legacy CacheStorage and service worker metadata readable, and
advance the HTTP cache format for raw header values.

Validation:
- cargo fmt --all
- cargo clippy --workspace --all-targets --all-features -- -D warnings
- cargo nextest run --no-fail-fast
2026-09-23 02:02:54 +08:00
lanyue-llk bcf420f622 revert: undo split PR 640 merges (#734-#740) 2026-09-22 20:02:07 +08:00
mini llk cba79021aa fix(document): parse text responses with consistent decoding (#740)
* fix(parser): preserve text document MIME and encoding semantics

* fix(document): share literal text parsing across child document paths

* fix(document): construct projection parser with the public API
2026-09-22 19:50:42 +08:00
ldm0 b88999b91e fix(modules): decode fetched source as UTF-8
Extract an independently mergeable topic from wpt-misc-fix at
a70a96f9d1.

Apply UTF-8 decoding directly to the baseline module fetch path without including import-map integrity changes.

Source commits:
- bffc90964d

Validation:
- cargo fmt --all
- cargo clippy --workspace --all-targets --all-features -- -D warnings
- cargo nextest run --no-fail-fast
2026-09-10 14:39:04 +08:00
ldm0 4c00e6dab6 fix(encoding): honor compatibility XML declarations 2026-09-08 16:17:11 +08:00
ldm0 5e173f3a5a refactor(encoding): extract legacy charset detector 2026-08-27 17:53:51 +08:00
ldm0 23acc9d739 feat(encoding): detect unlabeled legacy HTML like Blink 2026-08-27 17:53:51 +08:00
ldm0 48c307f7b3 fix(encoding): match Blink meta charset prescan 2026-08-27 17:53:51 +08:00
ldm0 9b5961086c fix(encoding): separate response charset lookup 2026-08-26 20:28:00 +08:00
ldm0 515549d1e3 fix(encoding): match Blink exact charset lookup 2026-08-26 20:28:00 +08:00
ldm0 d7ce6655ff refactor(content-type): centralize parsing semantics 2026-08-26 20:28:00 +08:00
Athul Nambiar eeb2ae6179 fix(headers): respect quoted strings when splitting header parameters 2026-08-26 03:47:19 +08:00
ldm0 6fa68cadbe fix(encoding): preserve split UTF-16 XML signatures 2026-08-24 02:02:40 +08:00
BibekPathak 0e597fc376 Moved encoding_for_document_xml_declaration check after self.transport_encoding, so the HTML prescan cannot override an explicit HTTP charset and Added transport_charset_wins_over_utf16_xml_signature: sends a UTF-16LE body with <?xml but with a Content-Type: text/html; charset=windows-1252 header 2026-08-23 18:32:00 +08:00
BibekPathak f917e513a0 encoding_for_document_xml_declaration now matches 6-byte <?x patterns per the HTML Standard and Added 6 negative tests 2026-08-23 18:32:00 +08:00
BibekPathak eaafee04fe BOM-less UTF-16 XML detection 2026-08-23 18:32:00 +08:00
Athul Nambiar 787fde9723 fix(encoding): rewrite meta-declared UTF-16 to UTF-8
A document whose only encoding declaration is `<meta charset="utf-16">`
was decoded as UTF-16, folding every byte pair into one CJK code point.
The damage was not confined to text: the tokenizer saw no markup either,
so `<title>` and `<p>` disappeared and the whole document collapsed into
a single text node inside `<body>`.

The prescan reaches a `meta` element only by reading ASCII-compatible
bytes, so a document it can see that declares UTF-16 has necessarily
mislabeled itself. The HTML Standard therefore rewrites the charset
before returning it:

    If charset is UTF-16BE/LE, then set charset to UTF-8.
    If charset is x-user-defined, then set charset to windows-1252.

Apply both rewrites in `moli-charset-parser`, which covers the
`meta charset` and the `http-equiv=content-type` paths at once. Keeping
them there rather than in `moli-encoding` leaves the two neighbouring
paths that legitimately select UTF-16 untouched: a UTF-16 BOM, and a
UTF-16 charset on the transport layer, which the encoding sniffing
algorithm takes with confidence certain and does not rewrite. Both are
now covered by regression tests, alongside the rewrites themselves and a
check that unrelated labels and labels outside the Encoding Standard are
still resolved exactly as before.

Verified on aarch64-darwin: `cargo test -p moli-charset-parser
-p moli-encoding` passes 55 tests, and `cargo fmt`/`cargo clippy
--all-targets` are clean for both crates.

Closes #152
2026-08-22 15:37:25 +08:00
ldm0 27b5135cb6 Public preview 2026-08-11 00:10:12 +08:00