276 Commits
Author SHA1 Message Date
ldm0 482a9b891f fix(charset): block fallback after invalid meta charset 2026-08-23 15:50:37 +08:00
ldm0 ed9e65bd31 fix(charset): reject non-ASCII label whitespace 2026-08-23 15:31:43 +08:00
ldm0 08d8461c26 fix(fetch): preserve typed readiness timeout phases 2026-08-23 04:45:50 +08:00
ldm0 7cf5dfbb34 fix(renderer): isolate stale import completions 2026-08-23 04:45:50 +08:00
ldm0 09b70a9b40 fix(benchmark): handle CDP download navigations 2026-08-23 04:45:50 +08:00
ldm0 c461cb1799 refactor(benchmark): unify public web comparisons
Move public-web seeds into CSV fixtures and share scheduling, command construction, classification, evidence, and artifact handling across top-sites and wild-web suites.

Bind Chromium DCL and response evidence to the current navigation loader, surface navigation errors promptly, rotate target-first order across repeated runs, and normalize HTTP and timeout outcomes so cross-engine results remain comparable.
2026-08-23 04:45:50 +08:00
ldm0 54f453cdd0 fix(fetch): preserve readiness deadlines and diagnostics
Apply one absolute deadline across response headers, streaming raw bodies, page creation, and later readiness waits. Report the active timeout phase without extending the caller's budget and reject raw documents from Page APIs before an unusable body can stall.

Give CLI fetch failures a stable two-line Error/Reason presentation while retaining full anyhow chains for non-fetch failures, including raw-document page-wait errors.
2026-08-23 04:45:50 +08:00
ldm0 8e6dd03abc fix(renderer): settle shared stylesheet clients per owner
Keep physical stylesheet fetches and import graph work shared while installing CSSOM state and terminal events for every link owner. Settle completed clients at admission, retain completed import graphs for late owners, and validate document ownership before asynchronous installation.

This replaces the abandoned turn-exit replay path with an explicit shared-fetch/per-owner lifecycle model and adds coverage for document.write(), data URLs, imports, failures, node mutation, and document replacement.
2026-08-23 04:45:50 +08:00
ldm0 ddab89586a fix(layout): keep scrollbar gutters out of overflow 2026-08-23 03:49:48 +08:00
ldm0 897e8e88a2 feat(layout): model physical scrollbar insets 2026-08-23 03:49:48 +08:00
ldm0 ac477b86f7 fix(layout): align viewport and scrollbar behavior 2026-08-23 03:49:48 +08:00
ldm0 48bdcdd104 refactor(renderer): isolate followed navigation flow 2026-08-23 03:49:07 +08:00
ldm0 412faf9153 refactor(renderer): model committed navigation state 2026-08-23 03:49:07 +08:00
ldm0 cb287fb651 fix(cli): apply redirect wait to terminal 3xx pages 2026-08-23 03:02:54 +08:00
ldm0 c86a373bcc fix(cli): preserve HTTP error bodies after readiness 2026-08-23 03:02:03 +08:00
ldm0 e2fdb90600 feat(renderer): allow lifecycle navigation fallback 2026-08-23 03:02:03 +08:00
ldm0 06134f461d fix(renderer): retire failed committed navigations 2026-08-23 02:19:51 +08:00
ldm0 050245f035 fix(cli): keep fetch dump output quiet by default 2026-08-23 02:07:35 +08:00
Athul Nambiar 86839b630c fix(encoding): extract meta content-type charset per the standard
`<meta http-equiv="Content-Type" content="...">` was parsed as a MIME
parameter list: split on `;`, skip the media type, compare each parameter
name for equality with `charset`, then strip quotes with `trim_matches`.
The HTML Standard instead specifies a keyword search over the whole
attribute value, and that difference is observable in both directions.

Declarations that were dropped, each falling back to windows-1252 and
rendering as mojibake:

    charset=utf-8                        (no media type at all)
    text/html; charset=utf-8 profile=x   (label ends at whitespace)
    text/html; xcharset=utf-8            (keyword matched as a substring)

Declarations that were wrongly accepted, because `trim_matches` strips
any number of either quote from both ends:

    text/html; charset='utf-8            (never closed)
    text/html; charset="utf-8'           (closed by the other quote)

The first group is the one that shows up in the wild — a `content`
attribute carrying no media type is common on legacy pages, and Moli
silently ignored it.

Implement the algorithm as specified: search for the literal `charset`
from a moving position, allow whitespace either side of the equals sign,
resume the search past a keyword that is not an assignment, and take a
quoted label only when the same quote closes it, otherwise ending an
unquoted label at the first ASCII whitespace or `;`.

Two details worth noting for review. The search resumes at the end of a
non-assignment keyword, which strictly advances and cannot loop. And all
byte indices land on ASCII bytes, which never occur inside a multi-byte
UTF-8 sequence, so slicing the latin1-mapped prescan input stays on
character boundaries.

An empty label now ends the extraction rather than continuing to a later
`charset=` in the same value, matching step 6, and that case is pinned by
a test.

Verified on aarch64-darwin per AGENTS.md: `cargo fmt --all`,
`cargo clippy --workspace --all-targets --all-features -- -D warnings`,
and `cargo nextest run --no-fail-fast` (16392 passed; the 7 failures are
pre-existing on 6f5aa320 and reproduce identically with this change
reverted).

Closes #156
2026-08-23 00:44:47 +08:00
ldm0 6f5aa320b2 Bump version to 1.0.3 v1.0.3 2026-08-22 20:00:01 +08:00
ldm0 f594c635ac test(cdp): pin classic scrollbar viewport 2026-08-22 19:44:54 +08:00
ldm0 d823999dc0 fix(renderer): fence screencast resource races
Split screenshot and screencast capture contracts, publish one shared visual resource generation, and collect CSS image references during the real layout traversal. Add asynchronous publication and multi-session screencast regressions.
2026-08-22 19:44:54 +08:00
ldm0 7608292c78 fix(renderer): stabilize style observations and screencast polling 2026-08-22 19:44:54 +08:00
ldm0 03a80259eb Reuse composite layout snapshots for iframe input 2026-08-22 19:44:54 +08:00
ldm0 e3d1accf53 Add nested iframe CDP input smoke 2026-08-22 19:44:54 +08:00
ldm0 a7223d4908 Split geometry scrolling from CSSOM metrics 2026-08-22 19:44:54 +08:00
ldm0 2e0140e53a Split geometry queries from hit testing and mock layout 2026-08-22 19:44:54 +08:00
ldm0 3f89e0109c Simplify iframe input and scroll routing 2026-08-22 19:44:54 +08:00
ldm0 4f39a0f0d9 Expand iframe and scrollbar input coverage 2026-08-22 19:44:54 +08:00
ldm0 f481bfa192 Fix transformed iframe input routing 2026-08-22 19:44:54 +08:00
ldm0 3a8e52ba99 Implement classic scrollbars 2026-08-22 19:44:54 +08:00
ldm0 43271abe7e docs(fetch): correct default timeout 2026-08-22 18:33:16 +08:00
Athul Nambiar 787fde9723 fix(encoding): rewrite meta-declared UTF-16 to UTF-8
A document whose only encoding declaration is `<meta charset="utf-16">`
was decoded as UTF-16, folding every byte pair into one CJK code point.
The damage was not confined to text: the tokenizer saw no markup either,
so `<title>` and `<p>` disappeared and the whole document collapsed into
a single text node inside `<body>`.

The prescan reaches a `meta` element only by reading ASCII-compatible
bytes, so a document it can see that declares UTF-16 has necessarily
mislabeled itself. The HTML Standard therefore rewrites the charset
before returning it:

    If charset is UTF-16BE/LE, then set charset to UTF-8.
    If charset is x-user-defined, then set charset to windows-1252.

Apply both rewrites in `moli-charset-parser`, which covers the
`meta charset` and the `http-equiv=content-type` paths at once. Keeping
them there rather than in `moli-encoding` leaves the two neighbouring
paths that legitimately select UTF-16 untouched: a UTF-16 BOM, and a
UTF-16 charset on the transport layer, which the encoding sniffing
algorithm takes with confidence certain and does not rewrite. Both are
now covered by regression tests, alongside the rewrites themselves and a
check that unrelated labels and labels outside the Encoding Standard are
still resolved exactly as before.

Verified on aarch64-darwin: `cargo test -p moli-charset-parser
-p moli-encoding` passes 55 tests, and `cargo fmt`/`cargo clippy
--all-targets` are clean for both crates.

Closes #152
2026-08-22 15:37:25 +08:00
ldm0 0c3a4068bf docs(playground): explain the ChatGPT login bridge 2026-08-22 15:22:06 +08:00
ldm0 37ddcd6952 feat(playground): persist ChatGPT login through Moli 2026-08-22 15:22:06 +08:00
ldm0 194698bd6a fix(playground): harden ChatGPT password form detection 2026-08-22 15:22:06 +08:00
ldm0 8441f9a7d8 feat(cli): split response waits into literal and regex modes 2026-08-22 03:20:24 +08:00
ldm0 ca4f01f6a2 refactor(fetch): drop unused wait criteria equality 2026-08-22 03:20:24 +08:00
ldm0 e7a8a709ef refactor(cli): encapsulate response body regex 2026-08-22 03:20:24 +08:00
ldm0 cba58e3aba refactor(fetch): name response body regex explicitly 2026-08-22 03:20:24 +08:00
ldm0 420cbdcf5d test(cli): cover response body regex end to end 2026-08-22 03:20:24 +08:00
ldm0 3212b02cac fix(renderer): bind WebStorage to receiver window 2026-08-22 03:15:32 +08:00
ldm0 97f72b88d6 test(protocol): select binding event by execution context 2026-08-22 02:08:32 +08:00
ldm0 b70555cb4a fix(renderer): keep child modulepreload from delaying load 2026-08-22 02:08:32 +08:00
ldm0 d0fdcf8fff fix(renderer): classify modulepreload admission by relation 2026-08-22 02:08:32 +08:00
ldm0 8d82acf953 fix(renderer): keep modulepreload from delaying load
Use a synchronous prepare/commit/apply boundary for connected stylesheet lifecycle authority, so DocumentRuntime no longer acquires it through raw ContextHost pointers. Keep modulepreload network fetches identity-only and publish terminal link events without a Document load-delay lease.
2026-08-22 02:08:32 +08:00
ldm0 494ed8356a feat(cli): add dcl wait alias 2026-08-21 22:14:32 +08:00
ldm0 6aa0f4818f fix(cli): trim ambiguous short flags 2026-08-21 22:14:32 +08:00
ldm0 baadb4f734 fix(cli): stop inferring subcommands from flags 2026-08-21 22:14:32 +08:00
ldm0 3b555523ab test(ci): run complete protocol smoke suites by default 2026-08-21 22:14:01 +08:00