moli-benchmark
Python benchmark harness for the Moli benchmark standard in docs/moli-benchmark-standard.md.
This project writes development benchmark artifacts under moli-benchmark/results/<timestamp>/ by default. Use --report-date YYYY-MM-DD for formal artifacts under benchmarks/results/YYYY-MM-DD/. Most suites run their own Python harnesses; Suite C WPT archival intentionally shells out to cargo nextest run -p moli-core --test wpt_compat --release --no-fail-fast so WPT reports use the same release profile as normal repository verification:
environment.jsonversions.jsonsummary.jsonsummary.mdpublish-readiness.jsonreport-data.jsonindex.html- optional
report-diff.{json,csv}when--baseline-reportis provided startup/runs.{csv,json},startup/gate-rows.json, andstartup/summary.jsonsynthetic/runs.{csv,json}andsynthetic/summary.jsonsynthetic-matrix/matrix.{csv,json},synthetic-matrix/gate-rows.json,synthetic-matrix/run-summaries.json, andsynthetic-matrix/summary.jsonwpt/moli-wpt-compat-report.json,wpt/by-tag.csv, optionalwpt/diff.{csv,json}, andwpt/summary.mdcdp-smoke/group-listing.json,cdp-smoke/preflight.json,cdp-smoke/client-rows.json,cdp-smoke/moli-cdp-smoke.json, andcdp-smoke/summary.mdsynthetic-compare/runs.{csv,json},synthetic-compare/summary.json, and top-levelindex.htmlcdp-session/runs.{csv,json}for long-lived CDP sessions.agent-episode/report-data.json, offlineagent-episode/index.html, raw episode/step rows, resource timelines, phase markers, and bounded failure artifacts for deterministic RL-shaped CDP workflows.crawler/raw-runs.csvfor local multi-page crawler runs.amiibo-crawler/raw-runs.csvfor the Python raw-CDP Amiibo crawler.wild-web/raw-runs.csv,wild-web/failures/, and optionalwild-web/replay/captures for real-site seed classification, extraction assertions, failure snapshots, and explicit replay fixtures.top-sites/raw-runs.csv,top-sites/runs.json,top-sites/summary.json, andtop-sites/failures/for public-web fetch/DCL benchmarks. The default source remainsdocs/chinese-community-top100-websites.md(quick= top 20,full= top 100).--target moli-cdp,--target lightpanda-cdp, and--target obscura-cdprun the same URL set through each engine's CDP server and dump DOM after the CDP DOMContentLoaded signal.--target moli-fulland--target moli-full-cdpreuse the Moli binary with--layout --resource; ordinary Moli targets keep the product-default Mock layout and disabled optional resources.--source webfetch-mix --profile webfetchcombines 100 mixed Chinese/global top sites with observed longtail WebFetch URL paths fromdocs/webfetch-longtail-seed-list.md.--source render-qualityuses concrete article/document URLs fromdocs/render-quality-seed-list.mdfor rendered-DOM quality regression checks.--source legacy-encodinguses non-UTF-8 pages fromdocs/legacy-encoding-websites-seed-list.md.render-compare/raw-runs.csv,render-compare/runs.json,render-compare/fetch-runs.json,render-compare/baseline-runs.json,render-compare/baseline-sites.{csv,json},render-compare/summary.json, andrender-compare/failures/for baseline-relative public-web render quality checks. The suite defaults to Chrome as the baseline, first runs the full baseline URL set, filters tobaseline-usablepages, then runs target browsers only for that evaluated set. It compares target visible text with character n-gram containment, key phrase hit rate, visible text ratio, and a combined render quality score. Rows can distinguishrender-match,render-partial,render-mismatch,state-only-content, and fetch-level failures. Baseline-unusable or baseline-thin pages are recorded inbaseline-sitesand skipped before target scoring.
versions.json records detected engine binaries for moli, lightpanda, chrome, and obscura. Suite result rows use engine/driver target variants: CLI fetch targets use moli, moli-full, lightpanda, and obscura; public-web DCL targets also support moli-cdp, moli-full-cdp, lightpanda-cdp, obscura-cdp, and chrome with driver=cdp-dcl; synthetic CDP suites use moli-cdp, moli-full-cdp, lightpanda-cdp, chrome-cdp, and obscura-cdp. Every row and target summary carries engine, driver, label, and binary_key, so report-data.json is independent of the HTML renderer. Defaults come from MOLI_BIN / LIGHTPANDA_BIN / CHROME_BIN / OBSCURA_BIN, repository target/release/moli, and then PATH. Obscura local fixture runs use an external-looking global IPv6 fixture URL when one is available, so Obscura does not see a 127.0.0.1 URL.
The first landing slice covers:
- Suite E startup/deploy subset: binary size, stripped binary size, tar.gz package size, SHA256, daemonless minimal rootfs image tar/tar.gz size,
moli servereadiness, optional CDP first page, optional warm-process CDP page creation, optional idle footprint,moli fetch about:blank, local JS fetch startup,/usr/bin/time -vraw artifacts when GNU time is available, cgroup/procfs raw artifacts, and structured formal gate rows. - Suite B synthetic subset:
static-html,js-xhr-fetch,dynamic-script,dom-heavy,storage-cookie,forms-eventsthrough a local Python fixture server. - Suite B synthetic matrix: repeats the synthetic suite across a concurrency matrix and records median drift for stability checks. The
formalprofile expands default runs/repeats to the P0 floor fromdocs/moli-benchmark-standard.mdand emits structuredgate-rows.json. - Suite D CDP smoke archival: runs or archives
moli-cdp-smokeJSON output.cdp-smoke --profile formalrequires raw CDP, Playwright, and Puppeteer client coverage; Puppeteer coverage requires localnodeandpuppeteer-core. - Suite C WPT archival: runs local WPT compat per-fixture reports through release nextest, or collects existing reports with overall/per-tag pass-rate summaries and optional baseline diffing.
- Synthetic horizontal compare: runs the same fixture set across fetch-style variants
moli,moli-full,lightpanda,chrome, andobscura, then renders a top-level professional staticindex.htmlreport with KPI cards, P0 scorecard, headline charts, drilldown tables, and artifact links. - CDP session compare: starts each engine as a CDP endpoint through
moli-cdp,moli-full-cdp,lightpanda-cdp,chrome-cdp, andobscura-cdp, navigates multiple cases through a reused page session, and records compact console / JavaScript exception / network failure traces undercdp-session/traces/when failures or error events occur. - Crawler and wild-web suites: separate local many-page crawling from real external seed classification.
- Amiibo crawler suite: uses a Python raw-CDP crawler against the Lightpanda demo Amiibo site so the real 933-page workload can be run from the same benchmark harness.
- Synthetic fixture cases are split by domain under
moli_benchmark/synthetic_case_groups/; current groups arebasic,modules,dom, andio. synthetic-compareexits on the selected gate target only. The default gate target ismoli; competitor failures stay visible in JSON/CSV/HTML without blocking report generation. Obscura is included as a local adapter; the benchmark harness gives Obscura an external-looking global IPv6 fixture URL instead of a127.0.0.1URL.
Resource sampling runs at 100 ms and records process-tree PSS when /proc/<pid>/smaps_rollup is available, RSS fallback, and aggregate ps CPU percentage. Startup cases also request /usr/bin/time -v artifacts under startup/time/; if GNU time is unavailable, the runner archives an explicit unavailability marker instead of silently omitting the evidence. Startup cases also archive /proc/<pid>/cgroup plus common cgroup v2/v1 files under startup/cgroup/ when the environment exposes them. Image size artifacts are written under startup/image-size/ by packaging the benchmark binary plus ldd-discovered dynamic dependencies into a minimal rootfs tar and tar.gz without requiring a container daemon. Startup rows explicitly record process_cache_mode and kernel_cache_mode; --drop-os-cache attempts /proc/sys/vm/drop_caches as root and otherwise records an unavailable marker under startup/cache/.
The HTML report is a Chart.js static dashboard generated from report-data.json. The JSON file is the renderer-independent contract; index.html embeds the same payload so the report can still be opened directly from disk without a local web server. When synthetic-compare and cdp-session are present together, report-data.json also includes a derived horizontal_comparisons[0] entry named web-scraping-variants that joins shared cases across fetch-style and CDP variants for one side-by-side chart/table view.
The standalone browser-spider-local/bench.mjs runner samples at 500 ms by
default. Its sampler runs in a Node worker thread, derives interval CPU from
/proc tick deltas, and records complete browser process-tree RSS/PSS alongside
case and site markers. Every run automatically writes
resource-samples.{json,csv}, report-data.json, and a resource/correctness
dashboard at index.html. The dashboard combines resource timelines, case
quality/resource comparisons, process topology, sampler health, site latency,
outcome distribution, and per-site diagnostics. Use
--no-resource-sampling for an explicit
sampling-free run, or --sample-interval-ms N to select a custom interval
(minimum 100 ms).
The detailed contract is in
docs/browser-spider-resource-sampling-report-design-2026-07-26.md.
Pull requests from branches in this repository also run an exact base/HEAD
Spider Bench comparison in GitHub Actions. The deterministic local fixture is
reported as a correctness and stale-page-leakage diagnostic, while the public
48-site run provides noisier performance and compatibility evidence. Neither
result fails the PR: the check fails only when the observer cannot produce its
artifact. A trusted follow-up workflow creates a new PR comment for each
completed current-head run and links the full spider-bench-results artifact.
The execution and permission model is documented in
docs/browser-spider-ci-pr-benchmark-2026-08-03.md.
Expected rows are derived from the selected workload. Public sites default to
five rows each (240 for the full 48-site set); deterministic fixture routes
declare whether they intentionally produce five rows or no rows (40 total).
Consequently fixture failures, --site-limit, and full48 no longer share a
misleading hard-coded denominator. The PR comment opens the public result first
and includes site coverage, bounded outcome counts, and per-category rows.
For a compact leadership-facing local comparison report, use:
uv run moli-benchmark run --profile horizontal --timeout 10
This preset runs synthetic-compare and cdp-session against the default target matrix with 10 runs per case unless --runs is explicitly provided. It is designed to produce the ten-way moli / moli-cdp / moli-full / moli-full-cdp / lightpanda / lightpanda-cdp / chrome / chrome-cdp / obscura / obscura-cdp comparison view quickly. It remains an investigation report until the independent P0 formal gates in publish-readiness.json also pass.
Running
Build moli first:
cargo build --release
Run the default initial suite set:
cd moli-benchmark
python3 -m moli_benchmark run
Write a formal report directory. Generated date directories under benchmarks/results/ are ignored by git and should be uploaded by CI or release tooling:
python3 -m moli_benchmark run --report-date 2026-05-07
Run focused suites:
python3 -m moli_benchmark startup --runs 5 --timeout 30
uv run moli-benchmark startup --runs 5 --include-cdp-first-page --include-cdp-warm-pages --cdp-warm-pages 10 --idle-seconds 1 --idle-seconds 5 --idle-seconds 30 --timeout 30
uv run moli-benchmark startup --profile formal --timeout 30
sudo uv run moli-benchmark startup --runs 5 --drop-os-cache --timeout 30
python3 -m moli_benchmark synthetic --case static-html --case js-xhr-fetch --runs 5 --concurrency 5 --timeout 30
python3 -m moli_benchmark synthetic-matrix --case static-html --runs 5 --matrix-concurrency 1 --matrix-concurrency 5 --matrix-repeats 2 --timeout 30
python3 -m moli_benchmark synthetic-matrix --profile formal --timeout 30
python3 -m moli_benchmark synthetic-compare --runs 5 --concurrency 3 --timeout 30
python3 -m moli_benchmark synthetic-compare --gate-target moli --runs 5 --concurrency 3 --timeout 30
uv run moli-benchmark cdp-session --case static-html --runs 5 --timeout 30
uv run moli-benchmark agent-episode --target moli-cdp --target chrome-cdp --workers 1 --parallelism 1 --runs 1 --step-dwell-ms 14000 --sample-interval-ms 500 --timeout 30 --output-dir ../target/agent-episode-bench
uv run moli-benchmark run --suite synthetic-compare --suite cdp-session --case static-html --runs 5 --timeout 30
uv run moli-benchmark run --profile horizontal --timeout 10
python3 -m moli_benchmark crawler --pages 50 --timeout 30
uv run moli-benchmark amiibo-crawler --target moli --pool 1 --limit 5 --timeout 60
uv run moli-benchmark amiibo-crawler --target moli --pool 2 --limit 5 --amiibo-mode process --timeout 60
uv run moli-benchmark amiibo-crawler --target moli --amiibo-profile formal --timeout 60
python3 -m moli_benchmark wild-web --seed baidu-home --timeout 30
python3 -m moli_benchmark wild-web --seed zhihu-home --capture-replay --timeout 30
python3 -m moli_benchmark top-sites --profile quick --timeout 15
python3 -m moli_benchmark top-sites --profile full --timeout 15 --parallelism 6
uv run moli-benchmark top-sites --target moli --target moli-cdp --target moli-full --target moli-full-cdp --profile full --timeout 30 --parallelism 2
python3 -m moli_benchmark top-sites --source webfetch-mix --profile webfetch --target moli --target moli-cdp --target moli-full --target moli-full-cdp --target obscura --target obscura-cdp --timeout 30 --parallelism 6
uv run moli-benchmark top-sites --source legacy-encoding --limit 6 --target moli --target lightpanda --target chrome --timeout 30 --parallelism 3 --chrome-parallelism 2
./scripts/run-webfetch-mix-benchmark.sh
uv run moli-benchmark render-compare --source webfetch-mix --profile webfetch --limit 50 --target moli --target lightpanda --baseline-target chrome --timeout 30 --parallelism 6
uv run moli-benchmark render-compare --source webfetch-mix --profile webfetch --limit 50 --target moli --target lightpanda --baseline-target chrome --match-threshold 0.65 --key-hit-threshold 0.70 --timeout 30 --parallelism 6
uv run moli-benchmark render-compare --source render-quality --profile quick --limit 12 --target moli --target lightpanda --baseline-target chrome --timeout 30 --parallelism 3
python3 -m moli_benchmark cdp-smoke --group protocol --timeout 30
uv run moli-benchmark cdp-smoke --profile formal --timeout 120
python3 -m moli_benchmark wpt --case abortcontroller-basic --timeout 60
python3 -m moli_benchmark wpt --no-run --baseline ../previous/wpt/moli-wpt-compat-report.json
agent-episode is one fixed local benchmark, not a family of smoke/stress/live
profiles. Both engines consume the same checked-in manifest, fixture,
Runtime.evaluate(awaitPromise=true) expressions, and assertions. The
canonical 14,000 ms dwell is workload idle between steps; readiness still
comes from CDP responses and lifecycle events. Only correct episodes contribute
to latency summaries. report-data.json is the suite-level authority and the
HTML report is a self-contained renderer of that payload.
Compare a new top-level report with a previous report directory or summary.json:
python3 -m moli_benchmark run --report-date 2026-05-07 --baseline-report ../benchmarks/results/2026-05-01
Use an exact binary:
MOLI_BIN=../target/release/moli python3 -m moli_benchmark run --runs 5
LIGHTPANDA_BIN=/usr/local/bin/lightpanda CHROME_BIN=/usr/bin/chromium OBSCURA_BIN=$HOME/.cargo/bin/obscura python3 -m moli_benchmark collect-env
For publishable P0 synthetic data, use at least:
python3 -m moli_benchmark synthetic-matrix --profile formal --timeout 30
--profile formal uses all synthetic cases, the 1/5/10/25/100 concurrency matrix, runs=100, and matrix-repeats=5 unless those values are explicitly overridden. The report marks formal profile requirement failures separately from workload failures.
Remaining P0 Gaps
The current harness is a runnable benchmark skeleton with smoke coverage. It is not yet a publishable P0 benchmark report. Each run writes publish-readiness.json; a report remains investigation until the formal synthetic matrix, formal Amiibo crawler, WPT P0 smoke, CDP, startup/size, wild-web, top-level artifacts, and four-way target matrix checks all pass.
The default execution order is documented in docs/moli-benchmark-standard.md: formalize synthetic first, fill startup/size second, then add the real Amiibo crawler.
Targets Availablemeans the binary was detected and recorded inversions.json; it does not mean every suite measured that target. Each suite summary must be read separately. For example,startup --profile formalis currently Moli-only, whilesynthetic-comparemeasures fetch-style variants andcdp-sessionmeasures*-cdpvariants.- Synthetic has a formal profile for the required
1/5/10/25/100concurrency matrix, repeated stability checks, andRUNS=100floor; P0 still needs an actual completed formal run artifact and follow-up profiling for slow cases. - Startup now has a
formalprofile withruns=10, CDP first page, 10 warm-process page creations, idle footprint at1s/5s/30s, and structuredgate-rows.json. The 2026-05-07 local formal verification passed with 0 failures; P0 still needs release/report artifact retention rather than committing large raw outputs to git. Container image measurement is intentionally out of scope. - CDP smoke now has a formal profile that requires raw CDP, Playwright, and Puppeteer coverage from
moli-cdp-smoke, and CDP session records compact Runtime/Log/Network traces for failed or error-bearing runs. CDP still needs wider Playwright/Puppeteer workflow depth beyond the current smoke gate. - Crawler has a Python raw-CDP Amiibo crawler for the Lightpanda demo workload. Amiibo rows are bounded by
--timeout, CDP page-session setup failures are archived as failure artifacts instead of blocking the queue, and fetched pages assert URL/title/body text/link-count plus Amiibo name/series/image fields. The suite supportssessionmode for one browser process with multiple page sessions andprocessmode for one browser process per worker. The defaultsmokeprofile runspool=1,limit=5, andsession;formalexpands to the full pool matrix, both modes, and all 933 pages. P0 still needs an actual completed formal artifact. - Wild-web has first-pass title/body keyword extraction assertions, failure taxonomy, failure snapshots, and explicit opt-in replay capture via
--capture-replay; it still needs deeper per-site business-field extraction and curated replay usage. - Obscura is now a runnable target adapter for binary discovery, CLI fetch, and CDP serve startup. Local benchmark fixtures use an external-looking global IPv6 URL for Obscura when the host has one, so the target binary stays unmodified.