`<meta http-equiv="Content-Type" content="...">` was parsed as a MIME
parameter list: split on `;`, skip the media type, compare each parameter
name for equality with `charset`, then strip quotes with `trim_matches`.
The HTML Standard instead specifies a keyword search over the whole
attribute value, and that difference is observable in both directions.
Declarations that were dropped, each falling back to windows-1252 and
rendering as mojibake:
charset=utf-8 (no media type at all)
text/html; charset=utf-8 profile=x (label ends at whitespace)
text/html; xcharset=utf-8 (keyword matched as a substring)
Declarations that were wrongly accepted, because `trim_matches` strips
any number of either quote from both ends:
text/html; charset='utf-8 (never closed)
text/html; charset="utf-8' (closed by the other quote)
The first group is the one that shows up in the wild — a `content`
attribute carrying no media type is common on legacy pages, and Moli
silently ignored it.
Implement the algorithm as specified: search for the literal `charset`
from a moving position, allow whitespace either side of the equals sign,
resume the search past a keyword that is not an assignment, and take a
quoted label only when the same quote closes it, otherwise ending an
unquoted label at the first ASCII whitespace or `;`.
Two details worth noting for review. The search resumes at the end of a
non-assignment keyword, which strictly advances and cannot loop. And all
byte indices land on ASCII bytes, which never occur inside a multi-byte
UTF-8 sequence, so slicing the latin1-mapped prescan input stays on
character boundaries.
An empty label now ends the extraction rather than continuing to a later
`charset=` in the same value, matching step 6, and that case is pinned by
a test.
Verified on aarch64-darwin per AGENTS.md: `cargo fmt --all`,
`cargo clippy --workspace --all-targets --all-features -- -D warnings`,
and `cargo nextest run --no-fail-fast` (16392 passed; the 7 failures are
pre-existing on 6f5aa320 and reproduce identically with this change
reverted).
Closes#156
A document whose only encoding declaration is `<meta charset="utf-16">`
was decoded as UTF-16, folding every byte pair into one CJK code point.
The damage was not confined to text: the tokenizer saw no markup either,
so `<title>` and `<p>` disappeared and the whole document collapsed into
a single text node inside `<body>`.
The prescan reaches a `meta` element only by reading ASCII-compatible
bytes, so a document it can see that declares UTF-16 has necessarily
mislabeled itself. The HTML Standard therefore rewrites the charset
before returning it:
If charset is UTF-16BE/LE, then set charset to UTF-8.
If charset is x-user-defined, then set charset to windows-1252.
Apply both rewrites in `moli-charset-parser`, which covers the
`meta charset` and the `http-equiv=content-type` paths at once. Keeping
them there rather than in `moli-encoding` leaves the two neighbouring
paths that legitimately select UTF-16 untouched: a UTF-16 BOM, and a
UTF-16 charset on the transport layer, which the encoding sniffing
algorithm takes with confidence certain and does not rewrite. Both are
now covered by regression tests, alongside the rewrites themselves and a
check that unrelated labels and labels outside the Encoding Standard are
still resolved exactly as before.
Verified on aarch64-darwin: `cargo test -p moli-charset-parser
-p moli-encoding` passes 55 tests, and `cargo fmt`/`cargo clippy
--all-targets` are clean for both crates.
Closes#152