Files
camoufox/ci
Jake WriterandClaude Opus 5.5 0c6cc0a397 Official TypeScript/JavaScript launcher at parity with pythonlib, published to npm (#785)
* feat(ts): import the TypeScript launcher port from feat/captchakrakenAndJSSupport

CAPTCHA support is left out; this branch is the JS/TS driver only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(ts): drop the CAPTCHA wiring left behind by the import

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ts): port fpgen to TypeScript, on the same pinned model

fpgen is not on npm. The port reads scripts/data/fpgen-model.json and checks
its sha256 with TLS on, never fpgen's own first-release download. Everything
that does not depend on the random draw is identical to Python (network,
value lookups, trace probabilities, conditions, errors); the draws are held to
Python's distributions by chi-square tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ts): identity layer at parity with pythonlib

A bit-exact port of CPython's random.Random, numpy's PCG64 choice and orjson's
serialisation, so identity_salt/identity_seed and every seeded draw (fonts,
voices, media devices, WebGL, noise seeds) come out identical to Python for the
same identity. coherence.py, presets and screen/window fixes are ported, and
golden fixtures recorded from pythonlib hold all of it to exact equality.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ts): launcher at parity with pythonlib's launch_options

launch_options() now produces pythonlib's output byte for byte (CAMOU_CONFIG,
CAMOU_PREFS_N, prefs, env, fontconfig, warnings) over 89 recorded scenarios.
Ports core pinning, geolocation, locales, fontprobe, the async API, and the
pkgman/multiversion integrity checks. An opt-in e2e suite launches a real
build through both launchers and compares what a page sees.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: test the TypeScript package, and publish it to npm like pypi

- ci/run_typescript.py writes the `typescript` gate (typecheck, lint, vitest
  with the pythonlib golden tests) and, with --browser, `typescript_browser`
  (the e2e suite against the browser under test). Both are required by the gate.
- publish-npm.yml mirrors publish-pypi.yml: workflow_dispatch, checks, build,
  scripts/check-pack.mjs (version == pythonlib, every data file shipped, the
  tarball installs and imports), then publish via npm trusted publishing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ts): e2e that holds on any browser, NewContext, and the package on every PR

- Parity (TS == Python on the same binary) stays strict everywhere; whether the
  browser honours the config is asserted only on a binary whose properties.json
  knows every key the launcher sets, and otherwise skips naming the missing keys.
  A driver-only pull request is tested against the published release, which
  lags the launcher (beta.30 predates #779), so this is what makes the suite
  meaningful there instead of red on skew it cannot fix.
- New: NewContext in a real browser -- a per-context identity that differs from
  the launch identity and from a sibling context, and equals Python's.
- python_probe.py keeps stdout for its JSON (pythonlib prints "Skipping unknown
  patch" there), and a non-JSON reply now fails fast instead of hanging 240 s.
- The typescript gate builds the package and runs scripts/check-pack.mjs, so a
  packaging mistake fails the pull request that makes it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): commit the launch fixtures, and fetch the browser from its real directory

- The root .gitignore ignores every path named `launch` (local build output),
  which silently dropped typescript/tests/fixtures/launch/ -- the
  launch_options() goldens -- from the branch. Re-included in
  typescript/.gitignore.
- fetch-browser read camoufox-bin from `camoufox path`, the cache ROOT, but
  multiversion installs each build under browsers/<channel>/<version>/, so the
  job has failed on every driver-only pull request since #772. It now resolves
  the active build as the launcher does (pkgman.camoufox_path).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci(ts): give the typescript job pythonlib, so the cross-language checks run

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ts): fpgen model install is safe across processes

Processes installing into an empty cache at once each downloaded the model,
and one's install deleted the values.dat another had just decompressed, which
then failed its next lookup with ENOENT. Seen with vitest's parallel files on a
cold cache; a worker pool on a fresh machine would hit it too.

- ensureModel() installs under a cross-process lock (an atomic mkdir, stale
  after 10 min) and re-checks what is installed once it holds it.
- values.dat is only removed when the model is actually being replaced.
- The model keeps values.dat open, instead of reopening it on every lookup.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: make local runs and CI see the same test suite

Three ways the suite passed here and not on the runner, each fixed at its cause:

- The root .gitignore's bare `launch` rule (for the Go launcher binary) ignored
  every path segment named launch, so tests/fixtures/launch/ never reached git.
  Anchored to /launch; and ci/run_typescript.py now fails when any file under
  typescript/{src,tests,scripts} is git-ignored, which would have caught it on
  the machine that wrote the fixtures.
- A missing prerequisite (fpgen model, pythonlib venv, fontTools, Xvfb, a font
  directory) skipped its tests, and a skip reads as green. tests/prereq.ts now
  fails them under CI unless the job names the gap in
  CAMOUFOX_TEST_ALLOW_MISSING. The typescript job installs all of them. The
  font-name check read one developer's local browser bundle; it now reads
  /usr/share/fonts (or CAMOUFOX_TEST_FONT_DIR), and CI installs a .ttc set.
- The fpgen install race surfaced only on a cold cache, by accident. It now has
  deterministic tests: a same-model reinstall keeps values.dat (verified to fail
  on the old code), the lock admits one holder and releases on error, and a
  stale lock is reclaimed.

Also: the browser gate runs only the e2e file, and the e2e probe and the
virtual-display test time-box each await, so a hang names its step instead of
reporting a bare 240 s timeout.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ts): hold headless="virtual" to Python's, not to headless

On the runner (no media hardware), published beta.31 never settles
enumerateDevices() in a headful window while headless answers -- the named
timeout in the probe caught it. That is a browser property, so like the other
page-vs-config checks it moves to a test that runs on a binary current with
the launcher; the virtual-display test now requires the same page as Python's
headless="virtual" on the same binary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(build-tester): accept 18 and 22 cores, as real hardware reports

plausibleHWC's list of common core counts lacked 18 and 22 -- Intel Meteor
Lake laptops (Core Ultra 5 125H, Core Ultra 7 155H), and 22 is in 8 recorded
presets. build-tester draws random presets, so a run that picked one of the two
Linux presets reporting 22 failed: about one run in eleven, on any pull
request. A CI self-test now fails if the list rejects any core count pythonlib
can present (the presets and PLAUSIBLE_CORE_COUNTS).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): test on the published release only when it matches this tree

A pull request that did not touch the browser was always tested against the
published release. The patch guards and suites come from the checkout, so once
a browser change was merged but not yet released (#779, on top of beta.31),
every driver-only pull request ran #779's guards against a browser without
#779 -- eight guards failed on #785, which changes no browser source.

resolve now also compares the tree's browser sources with the tag the release
was cut from (v<version>-<release> from upstream.sh), and builds when they
differ or the tag does not exist. Building restores the base branch's cached
browser when its compiled half matches -- main's #779 build, here -- so the
extra cost is a cache restore, not a compile. Self-tests run the workflow's own
scope step in a scratch repo for the four cases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ts): say when CI's browser cannot present the configured locale

With #785 finally tested on a browser current with the tree (main's cached
#779 build), every TS-vs-Python parity check passed and the page-vs-config
check failed: a de-DE/fr-FR identity presented en-US. Python presents the same
on that binary. CI tests the build job's unpackaged dist/bin, whose
res/multilocale.txt lists en-US only -- scripts/package.py injects the
langpacks, and CI never packages. So no CI suite had ever run a non-English
locale on a browser that has one.

The e2e locale assertions now run when the binary under test packages the
configured locale (read from res/multilocale.txt, loose or in omni.ja), and
otherwise go through prerequisite("packaged-locales"), which fails in CI
unless the job names the gap. The typescript (browser) job names it, with the
reason; the rest of the page-vs-config check stays strict. On a packaged #779
build all of it, locale included, passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(guards): judge the query-cost probes on a median, not one sample

stock-parity-probes timed each getter once. On a shared runner one GC pause or
CPU-steal spike decided the verdict: navigator.hardwareConcurrency took 77 ms
against a 50 ms allowance on the same restored build that passed the run
before. Each pair is now timed five times, interleaved, and compared by median.
The regressions these catch (a sync IPC per read, ~240 ms over the loop) cost
extra on every read, so they move the median; verified by giving the getter a
constant ~4 us of extra work per read -- 86 ms median, FAIL -- while the
healthy build passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ts): give the headful e2e page focus before probing it

enumerateDevices() intermittently never settled in the headless="virtual"
test on CI (passed one run, timed out the next, same build). Firefox defers
device enumeration until the document has focus -- LEAKS row 57 recorded the
same for a background tab -- and headless mode fakes focus while a headful
window on a bare Xvfb, with no window manager, only sometimes receives it. A
user's window has focus, so both launchers' virtual-display probes now bring
the page to the front first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: remove build tooling nothing uses

- The developer UI (scripts/developer.py, `make edits`). It depended on
  easygui, which no requirements file declares, and every action it offered is
  a Makefile target: patch, unpatch, workspace, revert, diff. Its two helpers in
  scripts/_mixin.py (is_bootstrap_patch, patch) had no other callers.
- legacy/, the Go launcher deprecated in 2024-11. Nothing built or shipped it.
  Its Makefile targets and scripts/run-pw.py go with it, and so does Go from
  every dependency list and workflow.
- jsonvv/ and settings/camoucfg.jvv. Nothing read the .jvv schema: config is
  validated against settings/properties.json, and the two had already drifted.
  The jsonvv package stays on PyPI.
- Scripts with no caller: bootstrap.py, moztree, setup-wasi-linux.sh,
  package-helper.sh, install-local-build.sh, mozfetch.sh (copied into lw/ but
  never packaged), examples/.
- The pre-ESM Juggler copies JugglerFrameParent.jsm and JugglerFrameChild.jsm,
  and hidden-scrollbars.css. Juggler loads the .sys.mjs actors and deliberately
  no stylesheet, but jar.mn still packaged all three.
- patches/librewolf/*.opt, which list_patches() never picks up; the roverfox
  second pass in patch.py, whose directory no longer exists; the unread
  --no-settings-pane option.
- The CAMOUFOX_PASSWD secret passed to `make fetch` and closedsrc_rev in
  upstream.sh, which nothing reads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(python): remove dead helpers and a stale dependency

None of these had a caller:
- pkgman: is_supported_path, extract_zip, cleanup and set_version, left over
  from the single-directory install. cleanup() would have deleted every
  installed browser version.
- multiversion.get_cached_repo_names, CONSTRAINTS.as_range,
  fingerprints._load_os_voices, utils._clean_locals, and unused imports.

Also:
- The "Apify Fingerprints" row in `camoufox version`, which has read "?"
  since fpgen replaced BrowserForge.
- lxml is no longer a dependency; nothing imports it.
- The geoip extra now names maxminddb, the module geolocation.py actually
  imports, rather than getting it transitively through geoip2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: show the cursor paths humanize=True actually produces

The README's cursor video showed the Bezier generator Camoufox replaced
with Cursory's recorded trajectories. scripts/cursor-demo.py drives a real
build with humanize=True and records every mousemove event the page
receives. It writes assets/humanize-cursor.svg, an animated replay at the
recorded speed, so what the figure shows is what a site sees.

The script cannot change the binary, so ci/browser_inputs.py lists it as
non-native and editing it does not invalidate the cached browser.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(python): stop naming BrowserForge in user-facing text

fpgen replaced BrowserForge, but two LeakWarnings, the NonFirefoxFingerprint
message and the fingerprint_preset docstring still named it. One warning also
linked to a README anchor that no longer exists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: one AGENTS.md for every agent, a roadmap, and docs that match the code

- AGENTS.md holds the engineering rules for any coding agent, plus the
  repo map, build, patch and test commands that CLAUDE.md used to carry.
  CLAUDE.md now only imports it, so there is one set of rules.
  ci/tribal-rules.yml is the record of settled decisions it points to.
- ROADMAP.md lists planned work, each item linked to its issue.
- README:
  - fpgen and the coherence check replace BrowserForge;
  - the patch workflow uses the make targets instead of the removed
    developer UI;
  - letter-spacing noise is described as off by default, as it is.
- docs/:
  - beta-testing-ff146.md removed;
  - patch-upgrading-guide rewritten around the make targets;
  - per-context-patches without the canvas patch that no longer exists,
    and with measured preset counts;
  - playwright-maintenance without the JSM wrapper that does not exist;
  - smaller fixes in MEDIA-DEVICES, input-dispatch and FONTS.
- ci/README: every job, and the real shard, skiplist and entry-point lists.
- pythonlib, tester and patch-dependency READMEs corrected against the code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(pythonlib): handle headless='virtual' in launch_server

launch_server() is documented to take the same arguments as Camoufox(),
but passed headless='virtual' straight to launch_options(), so the server
launched with no Xvfb display. Start a VirtualDisplay the way Camoufox()
does, launch headful on it, and kill it when the server process exits or
the launch fails.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(python): remove fontprobe, which nothing called

fontprobe listed the fonts installed on the host, for a `camoufox fonts`
command that was never added. It has nothing to do with the font bundle
Camoufox serves to pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(license): the Python launcher is MIT; the browser stays MPL-2.0

The Python package has always been published to PyPI as MIT (#727), but
pythonlib/ shipped no licence file, and the repo's LICENSE is the browser's
MPL-2.0. MPL is copyleft per file. It covers the modified Firefox sources, not
a separate launcher that drives the browser over Playwright. So
pythonlib/LICENSE now carries the MIT text its metadata already declares, and
a Licensing section in the README says which part is which.

Closes #727.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): fingerprint_preset=False no longer turns presets on

launch_options checked `fingerprint_preset is not None`, so passing False
drew a random bundled preset, the opposite of what was asked. It now uses a
truthiness check, and a test proves that None and False never draw a preset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: bring the release workflow in line with the tests build job

build.yml had drifted from tests.yml. It ran actions at v1/v2 on a
retired Node runtime, prepared the source tree with bare make calls
that fail the whole release on one dropped connection, and built with
a different Python than every pull request is tested with.

- Pin every action by commit SHA, at the major versions tests.yml uses
  (checkout v4, setup-python v5, upload/download-artifact v4, the same
  remove-unwanted-software SHA), and action-gh-release v2. The release
  job holds contents: write, so it should not follow a movable tag.
- Prepare the tree with `python3 -m ci.run_prepare`, as the tests build
  job does. BUILD_TARGET is set from the matrix so `make dir` writes the
  right mozconfig and Rust targets; multibuild.py then finds _READY and
  builds without re-patching. mach's toolchain bootstrap ignores the
  mozconfig, so running it after `dir` bootstraps the same toolchains.
- Build with Python 3.12, the version the tests build job compiles with.
- Default the workflow to no permissions; the build job gets
  contents: read and the release job keeps contents: write.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): NewContext looks up a proxy's exit IP through the right URL, or fails

NewContext derives the context's WebRTC IP and timezone from the proxy's exit
IP. That lookup had two defects, and both left the context showing the
host's values while its traffic went through the proxy:

- It built its own proxy URL with urlparse, which reads a scheme-less server
  such as "1.2.3.4:8080" (a form Playwright accepts) as scheme "1.2.3.4" with
  no host. urllib could not use a SOCKS proxy at all.
- Any failure was swallowed, and the context opened without the values.

The URL is now built with Proxy.as_string(), which the geoip launch path
already uses (scheme-less means http). The lookup goes through requests,
which handles SOCKS, and a failed lookup raises InvalidIP, naming the two
options that skip it. The tests cover scheme-less, http and socks5 servers
with credentials, both failure modes, and the case where no lookup is needed,
for NewContext and AsyncNewContext.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: stop generating a canvas seed, and drop config keys nothing reads

The browser has not noised the canvas since #528, and no patch reads
canvas:seed (#721). The launcher still drew one on every launch and sent it
through CAMOU_CONFIG, and NewContext called a setCanvasSeed that does not
exist. They no longer do.

For users this changes nothing on any browser since #528: the value was
ignored. A config that still passes canvas:seed gets the usual "Skipping
unknown patch" notice instead of silence. On a browser from before #528, the
launcher no longer turns canvas noise on, which is the behaviour #528 chose.

The same audit found more keys declared in settings/properties.json that no
patch or Juggler file reads, so setting them did nothing:
- canvas:aaOffset, canvas:aaCapOffset
- memorysaver, pdfViewerEnabled, webrtc:localipv4/6
- navigator.onLine, navigator.cookieEnabled, navigator.languages
- navigator.appCodeName, appName, product, productSub. Firefox reports these
  constants itself, so fpgen.yml no longer maps them.
- webGl:parameters:blockIfNotDefined and its WebGL2 twin

test_config_schema now checks this direction too: every declared key must be
read by the browser, unless it is listed with a reason. Three are listed:
locale:script and navigator.doNotTrack, which the launcher applies itself,
and navigator.buildID (#780).

The build-tester grading followed the same wrong premise. It tracked canvas
collisions as an unfixed per-context leak. A canvas that is rendered rather
than noised follows the fonts and GPU, as it does on real machines, so canvas
collisions are now counted with the other device-level values. The tribal rule
that recorded it as an open question is now a settled one,
canvas-is-not-noised, with an automated check.

Closes #721.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(ts): remove dead helpers

None of these had a caller:
- pkgman: isSupportedPath, extractZip, cleanup and setVersion, left over
  from the single-directory install. cleanup() would have deleted every
  installed browser version.
- multiversion getCachedRepoNames and getCachedVersions, CONSTRAINTS.asRange,
  removeMmdb (Python keeps its twins for the GUI) and pycompat pySorted.
- The "Apify Fingerprints" row in `camoufox version`, which read "?".

utils.ts now calls noiseSeedsFromIdentity instead of repeating its two
formulas inline, so the tested function is the one that runs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(ts): remove fontprobe, which nothing called

fontprobe.ts listed the fonts installed on the host, for a `camoufox fonts`
command neither launcher has. It has nothing to do with the font bundle
Camoufox serves to pages. Its parity test goes with it, and so do the CI
prerequisites only that test needed: fonttools and the extra font packages.
(The Python twin is removed in #787.)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(ts): license the launcher MIT, with third-party notices

The TypeScript launcher is a port of pythonlib, which has always been
published to PyPI as MIT (#727); the MPL-2.0 of the browser covers the
modified Firefox sources, not a launcher that drives it over Playwright.

THIRD_PARTY_NOTICES.md ships in the npm package with the notices for the
code the port translates: fpgen (Apache-2.0), CPython's random (the MT19937
BSD notice and the PSF licence), and NumPy's SeedSequence and PCG64 (BSD-3
and MIT).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: document the TypeScript package outside typescript/

The README, CONTRIBUTING, ci/README and the issue templates did not mention
the npm package or its two CI gates. ci/README also still said driver-only
pull requests never build. Since the scope step started comparing browser
sources against the release tag, they build whenever the release is behind.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): repair devicePixelRatio the same way on every launch

The DPR repair snaps an off-grid ratio to the nearest real scaling step and
keeps the first of two equally near steps. The steps were frozenset
literals, and a frozenset literal iterates in one order when the module is
compiled from source and another when it is loaded back from a .pyc. So a
midpoint such as 1.125 became 1.25 on the first launch after an install and
1 on every launch after it: the same pinned identity presented two different
devicePixelRatio values.

The steps are now ascending tuples, so a tie always goes to the lower step.
The test runs the repair in two fresh interpreters that share a bytecode
cache, compiling in the first and loading in the second. It failed before
this change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ts): mirror #787's pythonlib fixes

The TypeScript side of the behaviour #787 changes in pythonlib, so the port
stays at parity:
- fingerprint_preset=false no longer draws a preset.
- NewContext builds the proxy URL with ProxyHelper.asString() (scheme-less
  means http), looks up the exit IP through impit, and throws InvalidIP when
  the lookup fails instead of opening the context with the host's values.
- No canvas seed is generated or sent (#721). noiseSeedsFromIdentity becomes
  audioSeedFromIdentity, and fpgen's constant navigator fields are no longer
  mapped.
- The devicePixelRatio steps are ascending, so a tie goes to the lower step.
- The two LeakWarning texts that named BrowserForge.
- The README's note that Python's launch_server() ignored headless='virtual'
  is gone, because it no longer does.

The golden fixtures are regenerated from #787's pythonlib. The generator now
masks the fontconfig file name the way the test already did. The name hashes
content that embeds the checkout path, so every regeneration from a different
checkout used to rewrite 76 fixtures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: remove the glyph-spacing seed from the browser and the launcher

anti-font-fingerprinting.patch added a seeded amount to every glyph advance,
so that text widths differed per context. No real machine produces those
widths: the same font on the same OS measures the same everywhere. So the
noise was itself a fingerprint, measured in #779 at +1 px per ~100 glyphs
plus fractional deltas on every measureText. #779 defaulted the seed to 0
and kept it as an opt-in, but an opt-in whose only effect is to become
detectable is not worth carrying. Removed:

- The browser side:
  - FontSpacingSeedManager and window.setFontSpacingSeed;
  - the HarfBuzz hook;
  - the plumbing that existed only to carry the context id down to the
    shaper: the userContextId on gfxTextRun, gfxShapedWord and the word-cache
    key, and the extra MakeTextRun argument in nsTextFrame, nsFontMetrics,
    MathML and canvas.
  The font group keeps its userContextId, which font-list-spoofing.patch
  uses to apply the per-context font list. Text is now shaped exactly as
  stock Firefox shapes it.
- The fonts:spacing_seed key. The launcher had been sending 0 on every
  launch, plus a setFontSpacingSeed(0) call in every context's init script.
- tests/patches/config-overrides.py, which tested only the spacing override.
  A pythonlib test now covers config_overrides with another key.

timezone-spoofing, webrtc-ip-spoofing and window-setter-seal change only in
context lines and the setter seal list. Every patch applies cleanly to a
fresh tree, and the result builds. The settled decision is recorded as
no-glyph-spacing-noise, with an automated check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: stock animations and speech by default; drop config keys that freeze live values

Three behaviours a page could detect, changed in one breaking release:

- **Animations run on stock timing.** no-css-animations.patch finished every
  finite animation at once by default, and any page could read it:
  `el.animate(frames, 1000).effect.getComputedTiming().duration` was 0, and a
  500ms transition reported 0. Measured on v152.0.4-beta.31. The speedup is
  now an opt-in, `instantAnimations: True`, which raises a LeakWarning.
  disableInstantAnimations is gone.
- **speak() on a spoofed voice works like a real voice.** It fired `error`
  after 3ms unless voices:fakeCompletion was set, and then start and end in
  the same tick. It now starts and ends after the text's duration at ~150
  words per minute. Both voices:fakeCompletion keys are gone, and so is a
  debug line printed to stderr on every call.
- **Keys removed:**
  - battery:* and window.scrollMinX/Y: Firefox keeps getBattery() and
    scrollMin* chrome-only, so no page could read them.
  - window.scrollMaxX/Y, screen.pageXOffset/pageYOffset,
    window.history.length and document.body.client*: each pinned a live value
    to a constant, so scrolling, navigating or re-laying out never changed it.
    fpgen.yml mapped pageYOffset, so about 15% of identities froze
    window.scrollY at a non-zero value.
  - The body keys' role as an undocumented alias for window.innerWidth/Height
    in browser-init and in the launcher.
  - MaskConfig::GetInt32Rect, which only the body keys used.

New guards, both of which fail on v152.0.4-beta.31:
tests/patches/animation-timing.py and tests/patches/spoofed-voice-speaks.py.
The decisions are recorded as animations-run-on-stock-timing and
spoofed-voices-speak. Every patch applies cleanly to a fresh tree, and the
result builds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python)!: remove dead public API and make `list all --path` work

Breaking changes:

- Remove the exceptions UnknownProperty, InvalidDebugPort and
  MissingDebugPort. Nothing in the package raises them, so code catching
  them was catching nothing.
- Remove the legacy `allow_webgl` keyword of launch_options(). Use
  `block_webgl=True`. The keyword now reaches Playwright as an unknown
  launch option and fails there instead of being silently consumed.

`camoufox list all --path` accepted the flag and ignored it. It now prints
the install path beside each installed build, as `camoufox list --path`
already does for the installed tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(python)!: drop data the package never draws from

voices.json shipped in the wheel, but no code in the package reads it: the
voice draw uses voice-manifests.json and voice-uris.json. Its only readers
are the TypeScript port's golden-fixture generator and data-sync script
(typescript/scripts/golden/identity_golden.py,
typescript/scripts/sync-identity-data.py), which live on another branch and
will need a new source; the last copy is at
676fb3f:pythonlib/camoufox/voices.json. docs/per-context-patches.md
described it as runtime data and now describes the files that are.

webgl_data.db held two rows with zero weight on every OS ("Intel(R) HD
Graphics 400, or similar" from "Intel Inc." and "Radeon R9 200 Series, or
similar" from "ATI Technologies Inc."), left behind when their impossible
macOS weights were zeroed. No draw can reach them. They are deleted with
secure_delete so their blobs do not linger in free pages; the file is not
vacuumed, so the other pages are unchanged. A new test requires every row
to be drawable on at least one OS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): warn whenever an identity falls back to a substitute value

Several draws swallowed their failure and used something else, so an
identity could ship with values the rest of it was not drawn to match and
nobody would hear about it:

- from_preset(): a failed font or voice draw used the preset's recorded
  list, or nothing, on any exception.
- generate_context_fingerprint(): a failed font, voice or WebGL draw was
  `except Exception: pass`, leaving the browser's launch-time values.
- _load_font_groups() / _load_font_bases(): an unreadable file became {},
  i.e. no font additions or no OS-version base.
- launch_options(): a failed font draw used every font in fonts.json, a
  failed voice draw used no voices, and a preset GPU missing from
  webgl_data.db was silently swapped for a drawn one (36 of the 397
  bundled presets).

Each site now catches only the errors its data can raise (OSError and
ValueError for an unreadable or corrupt file, KeyError for a manifest with
no entry for the OS, sqlite3.Error for the WebGL database) and emits a
FallbackWarning. The text names what failed and what the identity uses
instead, then gives a block to paste into an issue (camoufox, browser, OS
and Python versions, the error, and the identity's user agent or GPU),
asking the user to report it on GitHub. It shares LeakWarning's
caller-frame attribution and its template lives in warnings.yml.

The broad excepts had also been hiding a broken fixture:
test_launch_environment's font and voice stubs did not accept `seed`, so
every draw there raised and was swallowed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): give NewContext identities the browser's Firefox version

NewContext() and AsyncNewContext() passed ff_version=None through to
generate_context_fingerprint(), so a context's user agent kept the version
fpgen drew (e.g. Firefox/146) while the browser underneath was 152. They
now default ff_version to the major version of Playwright's
Browser.version, which Juggler reports from MOZ_APP_VERSION_DISPLAY, so the
UA always names the browser the page is actually talking to. An explicit
ff_version still wins.

The docstrings said each context gets "its own real fingerprint preset";
the default has been an fpgen draw, with a preset only when one is passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): send an IPv6 WebRTC address to setWebRTCIPv6

The per-context init script passed every webrtc_ip, IPv6 included, to
window.setWebRTCIPv4(), and never called setWebRTCIPv6(). An IPv6 address
(given directly, or resolved as a proxy's exit IP) was stored as the
context's IPv4 value and the IPv6 slot stayed empty. The script now picks
the setter by address family, and an address that is neither raises
InvalidIP instead of being passed through.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): stop pinning the page's scroll offset from fpgen

fpgen.yml mapped the drawn window.pageYOffset (e.g. 528) to
screen.pageYOffset, and the browser returns that value from scrollY on
every read, so a page saw one scroll position forever whatever the user
did. Real scroll offsets are live page state, not part of a device's
fingerprint, so neither offset is mapped any more.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(python): stop checking a config key that no longer exists

warn_manual_config() looked for navigator.languages, which was removed
from settings/properties.json; validate_config() rejects it before the
check could matter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): close the WebGL database connection on every path

sample_webgl raised its not-found and wrong-OS errors before reaching
conn.close(), leaking a sqlite connection each time a preset named a GPU the
database does not hold.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(data): drop the 23 presets whose GPU has no WebGL data

A preset records only its GPU's name. The WebGL parameters, extensions and
shader precision behind it have to come from somewhere, and for these 23
nothing Camoufox has describes the GPU: fpgen has never seen Firefox report it
on that OS. So each launch paired the name with another device's parameters,
a mismatch any WebGL fingerprinter can see. They were:

- Windows on ARM (Adreno 650);
- Direct3D 10-level GPUs (vs_4_0/vs_4_1);
- "Generic Renderer";
- 945GM and GTX 480 on macOS;
- nouveau/Mesa buckets on Linux;
- one Linux preset pairing NVIDIA's proprietary vendor string with the
  nouveau renderer name.

scripts/clean-fingerprint-data.py now applies the rule, via a shared
fingerprints.firefox_gpus(), and test_shipped_data asserts it. 374 presets
remain, and every OS keeps its presets. ROADMAP.md lists capturing WebGL data
for these GPUs, which would bring them back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ts): stop generating the glyph-spacing seed

Mirrors 676fb3f: the browser no longer has glyph-spacing noise, so the
launcher sends no fonts:spacing_seed and the per-context init script no
longer calls setFontSpacingSeed. config_overrides is now tested with
audio:seed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ts): warn on instantAnimations; stop treating body keys as window size

Mirrors the launcher half of fb21b2e: instantAnimations raises the
instant_animations LeakWarning (warnings.yml copied from pythonlib), the
document.body.client* keys no longer count as window dimensions, and fpgen's
pageYOffset is no longer mapped, so no identity freezes window.scrollY.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ts)!: remove dead public API and make `list all --path` work

Mirrors 4616aa5: drop the UnknownProperty, InvalidDebugPort and
MissingDebugPort exceptions (nothing raises them) and the legacy allow_webgl
option (use block_webgl; allow_webgl now passes through to Playwright like
any unknown option). `camoufox list all --path` prints each installed
build's path, as `list --path` already did.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(ts)!: drop voices.json, which nothing draws from

Mirrors 9a42b6a. The voice draw reads voice-manifests.json and
voice-uris.json; voices.json was only read by the golden generator and the
data-sync script. The voice-URI golden now hashes the URI of every entry in
voice-manifests.json, the list the draw actually picks from.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ts): warn whenever an identity falls back to a substitute value

Mirrors 165e68f for the font and voice draws. Each fallback site in
fromPreset(), generateContextFingerprint(), loadFontGroups(),
loadFontBases() and launchOptions() now catches only the errors its data can
raise and emits a FallbackWarning naming what failed, what the identity uses
instead, and a block to paste into an issue (camoufox, browser, OS and Node
versions, the error, the identity). The message is warnings.yml's
`fallback` template, shared with pythonlib.

Python's except clauses name builtin classes JavaScript lacks, so pycompat
gains OSError, ValueError and KeyError twins and isPyError(): a Node system
error counts as an OSError and JSON.parse's SyntaxError as a ValueError, as
json.JSONDecodeError is. The voice draw now throws ValueError for a malformed
entry and KeyError when the manifest has no macOS entry, as Python does.

The WebGL fallback sites are left for the change that replaces the TS WebGL
source.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ts): give NewContext identities the browser's Firefox version

Mirrors 46e5c8d: without an explicit ff_version, NewContext() kept the
Firefox version fpgen drew, so a context's UA could name 146 on a 152
browser. It now defaults to the major version of Browser.version(). The
option docs now say the default identity is an fpgen draw.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ts): send an IPv6 WebRTC address to setWebRTCIPv6

Mirrors 0bbd152: the per-context init script passed every WebRTC IP to
setWebRTCIPv4(), IPv6 included. It now picks the setter by address family
and raises InvalidIP for an address that is neither. The init-script golden
gains an IPv6 case.

Also ports 7b43112's regression test: a drawn pageXOffset/pageYOffset is
not carried into the config (the mapping went in fb21b2e's mirror).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(ts): stop checking a config key that no longer exists

Mirrors 480789a: warnManualConfig() looked for navigator.languages, which
settings/properties.json no longer has; validateConfig() rejects it first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(ts): sync the pruned presets and the properties.json fixture

Copies fingerprint-presets*.json from pythonlib (40edebb dropped the 23
presets whose GPU has no WebGL data) and refreshes the launch fixture's copy
of settings/properties.json, which lost the keys removed in 676fb3f and
fb21b2e.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(python): draw every identity's WebGL from fpgen

WebGL vendor, renderer, context attributes, extensions, parameters and
shader precisions, for WebGL1 and WebGL2, now come from fpgen's recorded
Firefox devices instead of webgl_data.db, which is deleted with the
camoufox/webgl/ package.

camoufox/webgl.py:
- webgl_for_gpu() traces `webgl` given Firefox, the OS and the GPU, then
  `webgl2` given the chosen `webgl` too, and draws each with one seeded
  random.Random. The GPU and the webgl value are pinned by their fpgen
  lookup index: a dict condition is flattened into leaves that overwrite
  each other, so only the renderer applied and Linux "Mesa" and "AMD"
  Radeon HD 3200 devices came back mixed.
- sample_webgl_for_screen() draws the GPU of a generated identity from
  fpgen's per-OS weights, filtering out software rasterisers, GPUs the OS
  cannot report, discrete GPUs behind a netbook screen and the
  resistFingerprinting "Mozilla" mask before the weighted choice, so there
  is no rejection loop. An empty pool raises.
- The draft/host-dependent extension filter moves over unchanged.

A preset's GPU and a caller's webgl_config pair are looked up as given;
a pair fpgen has never seen from Firefox on that OS raises instead of
falling back to another GPU. generate_context_fingerprint no longer
falls back to the host GPU when the draw fails.

For 10 of the 15 (GPU, OS) pairs the two sources share, one of fpgen's
records converts to exactly the database row on every value the browser
reads. The other five rows (Linux R9 200 and Radeon HD 3200, macOS
Intel HD, and two software rasterisers) are devices fpgen does not carry;
those GPUs now present fpgen's recorded devices instead. The Linux
GTX 980 row is kept as a test fixture.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: say where WebGL comes from now that the database is gone

The per-context guide, the fpgen.yml header and coherence's comments still
named webgl_data.db and sample_webgl(). They now point at camoufox/webgl.py
and fpgen. The guide also claimed presets carry WebGL parameters; they
record only the vendor and renderer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ts): regenerate the golden fixtures for the mirrored pythonlib changes

Regenerates identity_golden.py and launch_golden.py output from this tree's
pythonlib. launch_golden.py drops fonts:spacing_seed and
document.body.clientWidth from its inputs, adds an instantAnimations
scenario, and masks a FallbackWarning's report block in both launchers, since
it names the host and the runtime.

Two launch scenarios still differ: preset_windows_unknown_gpu and
config_webgl_unknown_pair expect the FallbackWarning pythonlib now raises when
a preset's GPU is missing from the WebGL data. That site belongs to the
change replacing the TS WebGL source; the rest of each scenario matches.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ts): draw every identity's WebGL from fpgen

Mirrors 0d859b3. src/webgl.ts is the twin of camoufox/webgl.py:
webglForGpu() traces fpgen's webgl node given Firefox, the OS and the GPU,
then webgl2 given the chosen webgl too; sampleWebglForScreen() first draws
the GPU from fpgen's per-OS weights, filtered (no software rasteriser, no
resistFingerprinting mask, a GPU the OS can report, no discrete GPU behind a
netbook screen) before one weighted choice. Every draw is PyRandom.choices on
one seeded instance, in Python's order. The GPU and webgl value are pinned by
their fpgen lookup index, found from the value's stored JSON, which
TraceResult now carries: re-serialising a parsed value would spell 2**64
differently from orjson.

A preset's GPU and a caller's webgl_config are looked up as given and raise
when fpgen has never seen them, instead of falling back to another GPU;
generateContextFingerprint no longer swallows a failed draw. Removed with
the old source: webgl/sample.ts, the numpy default_rng port
(webgl/nprandom.ts), data-files/webgl_data.json and its export in
sync-identity-data.py. The NumPy notice now covers the pairwise sum in
locales.ts, the one NumPy port left.

launchOptions now throws the pycompat ValueError where Python raises
ValueError. Tests port test_webgl.py and the shipped-data check that every
preset GPU has WebGL data.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ts): pin the fpgen WebGL draws against Python bit for bit

identity_golden.py now records camoufox.webgl's output as sha256 of the
exact orjson bytes (key order, int versus float and all): 1272 screen draws
over every OS and seven screens (60 seeds each, plus four large seeds),
webgl_for_gpu for every GPU fpgen records and every bundled preset GPU, the
unknown-GPU and unknown-OS errors, and to_config's extension filter. The
numpy golden keeps only np.sum, now checked against locales.ts. The renderer
list the coherence golden walks comes from fpgen's traces.

launch_golden.py takes its WebGL pairs from firefox_gpus(), and the
unknown-GPU preset input is a GPU nobody records. The launch goldens are
regenerated; the preset_windows_unknown_gpu and config_webgl_unknown_pair
scenarios now expect Python's ValueError. The golden test no longer maps
ValueError to Error.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(patches): a host's missing speech daemon no longer errors spoofed speech

On a Linux host where speech-dispatcher cannot start, Firefox broadcasts
synth-voices-error, and SpeechSynthesis answers it by firing `error` on every
queued utterance. So a spoofed Windows voice errored about 11ms into speak()
on any host without the daemon: the CI runners, and most servers. It passed
only where the daemon runs.

While Camoufox manages the voice list, the registry no longer forwards a host
backend's error. The spoofed voices do not depend on the host's engine, and a
Windows or macOS identity never raises one. The guard now makes the daemon
unreachable itself, so it tests this case on every machine; on the previous
build it fails every time.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: stop skipping the two click tests that stock animation timing fixed

test_wait_for_stable_position and test_timeout_waiting_for_stable_position
were skipped with humanized travel time as the reason. The real cause was
instant animations. Every finite animation finished at once, so the button
Playwright waits on to stop moving never moved, and the click landed where
upstream does not expect. With animations on stock timing both pass, and the
skiplist audit flagged them as no longer failing. The entries go, and the
counts in ci/README.md drop from 14 to 12.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(build-tester): accept 18 and 22 cores, as real hardware reports

plausibleHWC's list of common core counts lacked 18 and 22 -- Intel Meteor
Lake laptops (Core Ultra 5 125H, Core Ultra 7 155H), and 22 is in 8 recorded
presets. build-tester draws random presets, so a run that picked one of the two
Linux presets reporting 22 failed: about one run in eleven, on any pull
request. A CI self-test now fails if the list rejects any core count pythonlib
can present (the presets and PLAUSIBLE_CORE_COUNTS).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(guards): judge the query-cost probes on a median, not one sample

stock-parity-probes timed each getter once. On a shared runner one GC pause or
CPU-steal spike decided the verdict: navigator.hardwareConcurrency took 77 ms
against a 50 ms allowance on the same restored build that passed the run
before. Each pair is now timed five times, interleaved, and compared by median.
The regressions these catch (a sync IPC per read, ~240 ms over the loop) cost
extra on every read, so they move the median; verified by giving the getter a
constant ~4 us of extra work per read -- 86 ms median, FAIL -- while the
healthy build passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(native): compare the whole fingerprint when two launches must differ

test_two_browsers_get_different_fingerprints compared seven coarse values:
UA, platform, screen size, core count, timezone and language. CI pins the
timezone and language, and real machines share the rest: two draws of a
common Mac (Firefox 152, MacIntel, 2560x1440, 8 cores) matched, and the test
failed on a correct browser.

It now reads the whole fingerprint a site computes, from a script in the page:
- navigator values, screen and window geometry, device pixel ratio, timezone;
- WebGL vendor, renderer, limits and extensions;
- installed fonts, measured by width against the generic fallbacks;
- voices, media-device counts, and an OfflineAudioContext hash.
The page is served from an https URL Playwright fulfils locally, because
mediaDevices exists only in a secure context. The page computes the result
itself because the isolated world may not read audio sample data. The test
then requires the fingerprints to differ, and the audio hash to differ on its
own, since its noise is seeded per identity.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Update README to remove warning, camoufox is now actively maintained

Camoufox will now be actively maintained and improved for the foreseeable future

* ci(ts): hold the golden tests to live pythonlib, not a snapshot of it

The typescript job ran the golden tests against the fixtures committed in
typescript/tests/fixtures/. A pythonlib change that typescript/ did not
mirror left those fixtures untouched, so the tests kept passing -- the
opposite of what the job's comment promised.

`ci.run_typescript --regenerate-golden` now rewrites the fixtures from the
checkout's pythonlib before vitest runs, and CI passes it. The job runs
Python 3.14 because pySum() reproduces sum() as 3.14 computes it; on 3.12
one crafted mixed int/float case differs in its last bit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci(ts): do not redraw fpgen's stats fixture on every run

stats.json is thousands of random fpgen draws that the TS tests compare
statistically. It changes with the pinned model, not with pythonlib, and
redrawing it took eight of the typescript job's eleven minutes on a
runner. The deterministic fpgen fixtures are still regenerated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(native): measure crash growth from a warmed-up parent

test_the_parent_stays_flat_across_content_crashes took its baseline right
after launch. The parent's first context costs it 150-200 MB with no
crash at all, so the warm-up counted as crash growth: 330-370 MB of the
400 MB allowance locally, and 469 MB on a CI runner, failing a PR that
changes nothing in the browser. The baseline now follows one clean
context; each crash still has to stay within the same allowance.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(ts): read pythonlib's data files instead of copying them

typescript/src/data-files/ held byte-identical copies of eleven pythonlib
data files (23k lines), kept in step by a sync script and a test that
failed when a copy drifted. The TS launcher now reads them from
pythonlib/camoufox/ when it runs from the repo, and `pnpm build` copies
them into dist/data-files/ for the npm tarball, so what users install is
unchanged. DATA_FILES in src/paths.ts is the one list; check-pack.mjs
checks each is in the tarball.

The essential-font lists were the one large table both ports hard-coded
(~170 lines of Python, ~770 of TS). They move to
pythonlib/camoufox/essential-fonts.json, which fingerprints.py and
fingerprints.ts both read; gen-fonts-json.py --print-bases writes it and
verify-fonts.py checks it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ts): record the golden fixtures from pythonlib on every run

The goldens were pythonlib's output committed as 14k lines of fixtures,
most of it the same input fingerprint pasted into ~80 launch scenarios,
and CI already rewrote them from pythonlib before each run. They are now
recorded by a vitest globalSetup (tests/golden-setup.ts) from the repo's
.venv or $CAMOUFOX_PYTHON, in about 6 seconds, and git-ignored. A
pythonlib change that typescript/ does not mirror fails `pnpm test`
locally as well as in CI, and ci.run_typescript no longer needs
--regenerate-golden.

Committed inputs stay: launch/inputs.json, the bundle stubs, the addon,
e2e/probe.js, and fpgen/stats.json (random draws tested statistically,
which change with the pinned model, not with pythonlib, and take minutes
to redraw). The one Python-version-sensitive case, sum() over mixed ints
and floats, skips with a named prerequisite when the goldens come from
Python < 3.14.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci(npm): give the publish job the pythonlib its tests record goldens from

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ts): NewContext hands Playwright the identity's user agent, DPR and timezone

generateContextFingerprint already returns Playwright's JS option names;
NewContext re-cased them, and camelCase() lowercases first, so userAgent,
deviceScaleFactor and timezoneId became keys Playwright silently drops.
navigator.userAgent was still spoofed at the C++ level, so the page probe
matched Python's, but the HTTP User-Agent, the DPR and the timezone did
not. The options now pass through as generated.

NewContext also awaits ensureModel(): a browser from connect() or a
custom executable never went through launchOptions(), which fetches the
fpgen model, so an fpgen draw threw ModelNotInstalled where Python works.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ts): a failed browser download rejects instead of killing the process

The download's file stream had no 'error' listener, so a failed write
(disk full) was an uncaught 'error' event: Node exited before
installVersioned()'s catch could remove the partial install and its temp
directory, and the caller had nothing to catch. finished() now listens
from the moment the stream is created, and webdl() surfaces an errored
stream instead of writing into it.

webdl() also waits for 'drain': it ignored write()'s return value, so on
a slow disk the whole archive queued in memory.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ts): concurrent launches no longer read or inherit each other's CPU pin

playwright-core spawns the browser from this process, so pin_cpu_cores
narrows this process's own mask while a launch's browser starts. Python
pins a separate driver and never sees its own mask change. Here, a
second launch started in that window:

- read the pinned mask as the host's cores (pinnedCoreCount, and the
  identity's hardwareConcurrency via availableParallelism()), and
- if it did not pin, spawned its browser without the lock, inheriting the
  first launch's pin while reporting more cores.

The host's core count is now read once, before this process first pins
itself (cpu_affinity.hostCoreCount), and unpinned launches, launchServer
included, take the pin lock once any launch in the process has pinned.
With pin_cpu_cores off (the default) nothing waits.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ts): an Xvfb that fails to start throws CannotExecuteXvfb

A spawn that fails after spawn() returns (EACCES, ENOENT) is an 'error'
event on the child process, which had no listener: Node treated it as
uncaught and exited instead of get() throwing. The child now always has
a listener, readDisplayNumber() rejects with CannotExecuteXvfb on it, and
the display pipe keeps an error listener after the read settles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(scripts): gen-fonts-json refuses to write an incomplete essential-fonts.json

An OS with no bases in the manifest was skipped, and the file written
without its key; fingerprints.py and fingerprints.ts read every OS's list
at import, so `import camoufox` then failed with a KeyError. The script
now exits and keeps the existing file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(release): pythonlib and the npm package to 0.5.7

@camoufox/camoufox 0.5.6 went to npm before the review fixes above, so
they ship as 0.5.7; the two launchers are versioned in lockstep, and
main already carries pythonlib changes from #787 that 0.5.6 on PyPI
does not have.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ts): NewContext's HTTP User-Agent must match navigator.userAgent

The e2e parity check read only navigator.userAgent, which the browser
spoofs itself, so a context that dropped Playwright's userAgent option
still matched Python. The probe server now records the request's
User-Agent header. Against the NewContext before the fix, a context
whose navigator said Windows sent the launch identity's Linux UA.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ts): draw the NewContext option-name test's preset from the v150 bundle

getRandomPreset() without a Firefox version draws from the older bundle,
where some Windows presets carry no devicePixelRatio, so the test failed
on some draws in CI. Every v150 preset has one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ts): wait for the headful e2e page to have focus, not one bringToFront()

On a bare Xvfb, with no window manager, a single bringToFront() before
goto() sometimes left the window unfocused, and Firefox holds
enumerateDevices() until the document has focus, so the virtual-display
probe timed out on some runs. Both launchers' probes now navigate first,
then bring the page to the front until document.hasFocus() is true, and
fail with that reason if it never is.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 03:28:12 +00:00
..

The test pipeline

Everything that runs a suite against Camoufox. Driven identically from a pull request, a push to main, and — through workflow_call — any caller that needs to test a specific browser version, so there is one definition of "the tests pass", not two.

resolve ── static ─────────────── lint, tribal rules, skiplist, self-tests  (seconds)
             ├─ typescript ────── type check, lint, vitest, golden parity
             └─ pythonlib ─────── the package's own tests                   (a minute)
                  └─ build or fetch ─┬─ patch guards ─────── one per spoofing patch, + skiplist audit
                                     ├─ build-tester ─────── 8 fingerprint profiles
                                     ├─ typescript-browser ─ the npm launcher end to end
                                     └─ once guards and build-tester pass:
                                          ├─ playwright × 6 shards   (conformance + our own)
                                          ├─ native ───────── leaks, contexts, crash recovery
                                          ├─ sundial ──────── stealth grade
                                          └─ growth ───────── memory growth  (scheduled only)
                                  │
                               summary ──► one comment on the PR

Which browser, which suite

ci/versions.py answers both, and every entry point uses it:

  • browser — from upstream.sh, or whatever a caller passes in. A caller moving to a new Firefox passes the version it is moving to, which is what lets one pipeline test both a pull request and an upgrade.
  • suite — the newest released playwright-python tag whose pinned Firefox is not ahead of that browser, and which is below the Playwright ceiling pythonlib/pyproject.toml pins.

Newest-not-ahead, rather than an exact match, because Playwright trails Firefox and skips generations: it pinned Firefox 151 and then 153, never 152, so a browser built on 152 has no exact suite and never will. Requiring a match would leave most of a release cycle with no suite at all, and taking a newer one would test against an automation contract that assumes engine work the build does not have. So the suite's Firefox pin being a release or two behind the browser is the normal case, not a misconfiguration — the summary line says which rule picked the tag.

The ceiling matters for the same reason it exists in pythonlib: camoufox.server imports playwright._impl._driver, a private API, and every Playwright minor is free to change Juggler. Testing above the ceiling would exercise a client the shipped package will not install.

python3 -m ci.versions --json                          # what would run
python3 -m ci.versions --browser-version 153.0.4 --json

The version under test has to be the version that gets built. Only suite selection follows --browser-version; the build reads upstream.sh and the fetch path downloads whatever pythonlib considers current. Asking for a version the branch does not pin would therefore compile the old browser and judge it against the new suite — green, meaningless and silent. --check-upstream refuses that, and the workflow passes it.

So an upgrade to a new Firefox is a branch that edits upstream.sh, which is what an upgrade is anyway. Resolution then reads it by default, the build produces it, and the suite is chosen for it — the three cannot disagree. The browser_version input exists for a caller that wants to state the version explicitly; it must match.

The Playwright suite

One suite, fetched fresh per run: upstream playwright-python at the resolved tag. It runs unmodified — ci/pw_camoufox_plugin.py adapts the environment around it rather than editing it, hooking BrowserType at the _impl layer so upstream can refactor its fixtures freely — and ci/suite.py overlays tests/camoufox/ into it.

tests/camoufox/ is small on purpose: behaviour upstream has no test for (that page.route() must not change what a request looks like on the wire), or asserts the opposite of on purpose (that a worker should not inherit the context locale, which stock Firefox gets wrong and Camoufox does not). It is not a fork of anything, so it cannot go stale; it runs against upstream's own conftest and server at whatever tag was resolved.

tests/ used to hold a fork of a ~v1.55-era upstream suite. It was deleted after measuring it against the same binary:

  • 73 of the 74 tests it skipped as "Not supported by Camoufox" pass in upstream's copy. Those skips predated main-world execution and were never revisited, so the fork was asserting the browser was worse than it is.
  • Of its passing tests, eight had no upstream counterpart. Six of those were Camoufox-specific and now live in tests/camoufox/ as three modules; the other two were tests upstream had since renamed.

A re-fetched suite cannot drift, and a deliberate difference from upstream now has to be written down in ci/skiplist.yml with a reason, where it is visible, instead of being encoded as a silent edit to a vendored file.

What "the suite" means

ci/run_playwright.py names its targets explicitly rather than pointing at tests/, so the one thing left out stays visible:

tests/async/ + tests/sync/ one pytest process
tests/common/, tests/test_reference_count_async.py their own process
tests/test_installation.py excluded, with a reason

The isolated pair each call sync_playwright()/async_playwright() inside the test body, which cannot start while the session fixtures already hold a loop (Cannot run the event loop while another loop is running). Run with the others all six fail; run alone all six pass. Skiplisting them for that would have recorded a browser failure that does not exist.

This used to be tests/async/ alone — 722 tests, 31% of the suite, excluded with nothing written down. Not a decision: the vendored fork carried async/ and no sync suite, and this runner was pointed at the same shape without checking what upstream shipped. unclaimed() now fails the run if upstream adds a test path that is neither in TARGETS nor in EXCLUDED with a reason.

Which world, and the skip list

The suite runs isolated first — the configuration Camoufox actually ships — and falls back to the main world only for what fails, counting every test that needed the fallback.

Upstream asserts upstream semantics: tests read globals their own page scripts defined and pass handles into evaluate(), and a test doing that fails under isolation by design. Running the whole suite main-world-only (the previous behaviour) made those pass, which is true but uninformative — it measured a mode nobody ships and produced no number for what isolation costs.

Each group is therefore run up to three times, and normally twice:

Pass CI_WORLD What it establishes
1 isolated the browser as users run it
2 main the failures again with isolation off
3 main only what failed in both, retried once — normally empty

There is deliberately no second isolated pass. One sat between 1 and 2 on the theory that a flake must not be mistaken for a world difference; measured on the first real CI run it cost 7m50s per shard and recovered nothing:

isolated (full)   335s + 331s   35 and 11 failures
isolated retry    205s + 265s   0 recovered      <- deleted
main world         19s +  12s   46 recovered

Two reasons it was never going to earn that. These failures are deterministic — a test reading a global its page script defined does not intermittently see it — and failing that way is slow, because the read returns undefined and the test sits on a Playwright timeout rather than throwing. And upstream's suite already ships pytest-rerunfailures: pass 1 reported 105 rerun, which is each of those 35 failures having been retried three times before the run even reported them. A flake does not survive that.

A test that passes in pass 2 is recorded as a main-world fallback: it counts as a pass for the run, and its identity goes into metrics.main_world_fallbacks, with metrics.main_world_fallback_count on the summary table (summed across shards). A test failing in both worlds and on retry is a plain failure. The fallback count is the isolated-world conformance gap — watch it between runs; a jump means the isolation boundary moved.

Both rerun passes are guarded on the failure set being non-empty, which is load-bearing rather than tidy: pytest declines to filter when nothing it collected previously failed, so an unguarded --last-failed runs the whole group again — in the other world, silently replacing the result it was meant to refine.

Why some isolated failures hang

Not every isolated failure fails. Some wait forever, in both tests/async/test_route_web_socket.py and tests/sync/test_route_web_socket.py — and the two are not equally recoverable, which is the subject of the second half of this section.

The shape recurs, so it is worth stating generally: a Playwright feature implemented by installing something on the page's global lands in the isolated world instead, so anything the page itself originates never reaches the automation. Two instances, both measured directly against a build:

page's own script opens a WebSocket    isolated -> handler never fires   main -> intercepted
page script calls window.exposedFn()   isolated -> HANG                  main -> resolves
evaluate() calls window.exposedFn()    isolated -> resolves              main -> resolves

route_web_socket works by replacing window.WebSocket from an init script; isolated, that replacement lands in the sandbox, so a socket the page opens is never seen. expose_function installs its binding on the sandbox global, so page script calling window.fn() finds nothing — though called from evaluate() it works, which is why that one does not hang here.

They hang rather than fail because the waits involved — a Twisted future from the test server, an asyncio future a binding was meant to resolve — have no Playwright timeout behind them. Everything else isolation breaks fails at Playwright's 30s.

Worth being clear that the route_web_socket half is not a test artifact: a real site's WebSocket is not intercepted either, and the user gets no error saying so. It is not fixable at this layer — the feature works by replacing a page global, which is precisely what an isolated world exists to stop a page from seeing. Tracked in #775; the fix is native interception below the DOM object, which is also the only version of it that stays undetectable.

So pass 1 is cost-bounded, and only pass 1:

Bound Value Why
per-test timeout 90s the slowest test in the whole main-world baseline was 30.4s; only two exceeded 30s and none exceeded 45s. A Playwright action times out at 30s
upstream reruns off (CI="") tests/conftest.py sets reruns = 3 whenever $CI is set — the only thing it reads $CI for. The baseline recorded 2 reruns across all 2295 tests; the isolated pass recorded 138 in one shard, all re-running deterministic world differences

Together that turns a hang from up to 4 × 180s into one 90s wait. A flake missed by not rerunning is not lost — it fails pass 1, passes pass 2, and is counted as a fallback. Note that --reruns 0 as an argument would not work: upstream's conftest overwrites config.option.reruns in pytest_configure, so clearing the environment variable is the only lever that holds.

The ones a timeout cannot bound

That bound is not enough for all of them, and it is worth knowing exactly where it stops working. Measured on run 34799668707 with the 90s bound already in place:

Group Isolated pass Outcome
tests/async/ completed in 296s the bound works
tests/sync/ test_should_work_with_ws_close printed pytest-timeout's +++ Timeout +++ banner at exactly 90s the process then sat for 1h50m, until the job's timeout-minutes killed it

So the signal fires and the test dies; the process does not. pytest-timeout's signal method raises at the next bytecode boundary, and Playwright's sync API is parked in a greenlet switch that never reaches one cleanly — the raise lands inside the dispatcher and wedges it. --timeout-method=thread fires reliably but kills the interpreter, taking the other ~1500 tests in the group with it. There is no per-test timeout value that bounds this.

So those modules are declared, not discovered — ISOLATION_HANGS in ci/run_playwright.py. The isolated pass cannot learn that they hang without hanging, so it is told: they are --ignored out of pass 1 and run directly in the main world (pass 1b), where they pass and are counted as fallbacks exactly as if isolation had failed them honestly. The same tests still run, in the world that can run them.

Why not ci/skiplist.yml. That list means "fails in the most permissive world", and ci/run_skiplist_audit.py enforces it by running every entry with CI_WORLD=main and failing the build on any that pass. A route_web_socket test passes there — the main world is precisely where the feature works — so an entry would be rejected by the audit, and would be untrue as written. The two lists are not interchangeable, and test_isolation_hangs_are_not_in_the_skiplist keeps them apart.

Pass 1b sits above the if not failing: continue guard, deliberately: a group whose isolated pass found nothing would otherwise skip it, and coverage would disappear on exactly the runs that look healthiest.

tests/patches/isolated-evaluate.py still owns the direct coverage of isolated evaluation, and must keep passing regardless. That file is what to check if isolation itself regresses.

ci/run_skiplist_audit.py deliberately runs in the main world: a skiplist entry has to claim a test cannot pass in either world, or the suite would have counted it as a fallback rather than a failure.

Fourteen tests are deselected outright by ci/skiplist.yml, which requires a stated reason per entry — ci/summarize.py fails the run on an unreasoned one.

A reason is not evidence, so the reasons are checked. The first version of this file inherited all nine tests/async/*.disabled modules from the vendored suite and gave each a plausible justification without running any of them: of the 202 tests it skipped, 193 passed, and seven of the nine modules failed nothing at all. A written reason made them look verified, which is worse than leaving them bare.

ci/run_skiplist_audit.py now runs every entry with the skiplist disabled and fails the build if a skipped test passes. It is cheap precisely because a correct skiplist is short — twelve tests, a few seconds — and it is what keeps the list from drifting back into a place failing tests go to disappear.

python3 -m ci.run_skiplist_audit --binary /path/to/camoufox-bin

What remains, 12 tests: two test_keyboard.py tests that assert a shifted character arrives without Shift, which Camoufox presses as a real keyboard would; six client-certificate tests (async and sync) that need the browser to present a certificate during the TLS handshake — the two that go through the Node driver's own request context instead pass, and are not skipped; two upstream expectations that encode a stock-Firefox quirk, replaced by tests/camoufox/; and two popup tests that rely on Playwright shipping Firefox's popup blocker off, which Camoufox keeps on.

That client-certificate split is the audit earning its place. The entry was first written as a whole module, because on a local machine all five fail — Node/OpenSSL there rejects the fixture server outright. In CI two of them pass, and the audit failed the build one run after the entry was written. CI is the authority for what fails; a local run is a hypothesis.

Camoufox's own suite

native-tests/ covers what the Playwright suite cannot ask about:

  • Leaks. Launch browsers, kill them, prove nothing survived — file descriptors, sockets, child processes, X11 lock files. The real assertion is that cost does not scale with launch count, because that is the shape a leak actually has: a scraper that runs fine for six hours and then dies of EMFILE. Scope is honest: this measures resources held by our process and its children, not Gecko's internal heap.
  • Contexts versus browsers. Two contexts in one browser must get different fingerprints; two pages in one context must get the same one. Get this wrong and per-context injection silently degrades to process-global — which passes every single-context test there is. It has happened here before (commit d17c887, "fix screen size leak in contexts").
  • Crashes. Kill the browser, the X server, a content process or the driver mid-run, then check that teardown does not hang, nothing leaks, and a fresh launch still works (test_crash_recovery.py).
  • Memory growth. Drive one mechanism (iframes, canvas readback, WebGL contexts, workers, script compilation, font measurement) N and 4N times and compare the growth: a bounded cost stays flat, a per-iteration leak scales (test_memory_growth.py). It takes over half an hour, so it runs on the schedule and on demand (the growth job, --subset growth), not in the gate.
  • Settled decisions. ci/tribal-rules.yml lists choices this project already made, each with the issue or PR that made it, and native-tests/test_tribal_rules.py asserts them. A comment explaining a decision only works on someone who reads it.

Sundial

On since 2026-09-12. sundial 0.5.0 is deployed and serving score mode, and sundial's master branch now deploys itself on push, so merged does mean deployed. It was off for as long as the live build predated score mode: an older sundial ignores ?score=1 and posts the entire report — every vector's id, name, brief, source and value — to whatever collector asked. Receiving that on a public runner and discarding it afterwards is not the guarantee this section describes; not receiving it is. enabled: false in ci/sundial.yml is still the kill switch, and the resolve job checks it before the credential comes into scope, so flipping it back stops the request rather than just the reporting.

The stealth check reports a letter grade and a count. Nothing else leaves ci/run_sundial.py::redact() — not a vector name, description, measured value, source, and not a per-category breakdown either: a table reading "Graphics 3/17" is the most useful single fact an adversary could take from a public CI log.

There are no per-check rows in the results file at all — not even opaque ones. An HMAC does not name a vector, but a map of them still publishes how many distinct checks fail and lets a reader follow one across releases, which is per-vector data wearing a hash. The instruction was a score, so it is a score: regression detection is per-score, via min_pass_rate and a maximum allowed drop. Scope and thresholds live in ci/sundial.yml; only categories Camoufox actually claims are gated.

Everything that leaves redact() is checked against a whitelist at runtime, not a blacklist — a blacklist only stops the leaks somebody already thought of. Adding a field without adding it to _PUBLISHABLE fails the run:

{ "grade": "A", "checks_total": 412, "checks_passed": 403, "pass_rate": 0.978,
  "out_of_scope_failed": 6, "cross_os_total": 24, "cross_os_passed": 5,
  "os": "linux", "sundial_version": "0.3.1", "schema_version": 1 }

Cross-OS detectors are counted, never scored. They read the host machine rather than the disguise: a browser claiming macOS while running on Linux fails them however good its spoofing is, and Camoufox does not claim byte-identical cross-OS emulation. Folding them into one average would mark it down for a promise nobody made, and would hide a real regression behind noise it cannot control. They are reported separately so a drop there reads as "the host shows through more than it did", which is a different conversation.

The run asks sundial for ?auto=1&score=1, so it receives counts and the vectors never cross the wire at all.

Finding out which checks failed

Worth being precise about, because the answer is "you can't, from CI", and that is deliberate rather than an oversight. Score mode's payload is buckets keyed "<Category>|<class>" holding two integers each. It carries no check names and no ids, so a failing check's identity is not something the CI process discards — it is something sundial never sends. Nothing in the artifact, the log, or the sealed report can recover it.

Two steps down from there, both local only:

# which CATEGORY the failures are in -- works with the credential CI already has
python3 -m ci.run_sundial --explain --binary /path/to/camoufox-bin

# which CHECKS -- needs a role sundial serves full reports to
python3 -m ci.run_sundial --explain --allow-full-report --binary /path/to/camoufox-bin

--explain prints to the terminal and never writes to a result file, and is refused outright under GITHUB_ACTIONS: a category-level table is not a vector, but "Graphics 3/17" is still the most useful single fact an adversary could take from a public log, which is exactly why redact() does not publish one.

Order of operations

?score=1 needs a sundial that has it. An older deployment ignores the unknown parameter and posts the whole report; the numbers still come out right and redact() still discards everything identifying, but nothing is classified, so every cross-OS tally reads 0 — which looks like "no host-OS failures" rather than "nobody sorted them". score_mode: false in the result says which it is, and the gate says so in its notes rather than leaving you to notice.

So the dependency runs one way, and setting the GitHub secrets is the last step, not the first:

  1. deploy sundial's score mode — done; it ships in 0.5.0, and master now deploys on push
  2. gh secret set SUNDIAL_AUTOMATION_KEY -R <repo> for every repository whose CI runs this. Without it the job skips, which is the normal case for a fork pull request
  3. flip enabled: true in ci/sundial.yml — done

Two things keep a vector out of a public log, and it is worth separating them, because only one is enforced by the server:

guarantee
server-side The run logs in as guest, and sundial's middleware refuses guest the private-vector bundle outright — those definitions are never served to the session.
client-side This gate only ever requests /?auto=1&score=1, and redact(require_score_mode=True) fails the run if a full report arrives anyway, rather than folding it down and carrying on.

The stricter option is sundial's score-only ci role, which is refused anything but /?auto=1&score=1 server-side and so cannot be handed a report even if the credential leaks. That role is not in sundial's master branch and is therefore not deployed; when it lands, mint the credential (make pages-ci, then redeploy) and set SUNDIAL_USERNAME=ci. Nothing in this repository changes — the client-side half already behaves as though the server were enforcing it.

There are deliberately no per-vector rows, not even opaque ones. An HMAC names nothing, but a map of them publishes how many distinct checks fail and lets a reader follow the same id from release to release.

The cost is real: regression detection drops from per-vector ("the check that passed last release fails now") to per-score ("we got worse"), covered by min_pass_rate in ci/sundial.yml — and, once an auto-update pipeline exists to compare releases, by a maximum allowed drop in its policy file. To get the per-vector view back for your own debugging, set SUNDIAL_REPORT_AGE_RECIPIENT to an age public key — the full report is then kept encrypted to you and nobody else can open it.

Needs SUNDIAL_AUTOMATION_KEY (the password). SUNDIAL_USERNAME is optional and names the account, which is not a secret — it defaults to guest. Absent the password — a pull request from a fork — the job is skipped and the summary says so.

What actually stops a vector reaching the log

Asking for ?score=1 is a promise the caller makes, and a promise is not a mechanism. Today two things back it:

  • guest cannot load the private vectors, and that is checked. sundial's middleware answers isPrivateVectorAsset paths with an empty stub for that role specifically. admin and private do get them — and /automated?key= resolves to private when handed the private key, which is indistinguishable from the guest one by looking at it. So "we set the right key" stays an assumption until something checks: the gate reads sundial's own /__auth/me and refuses to open the browser at all unless the session is a role the vectors are withheld from. Not knowing the role counts as not safe.
  • A non-score payload fails the run. redact(require_score_mode=True) refuses to process a full report rather than folding it down, so a deployment that ignored score=1 is a red build, not a quiet leak.

What is still missing is a server that refuses the request. sundial's score-only ci role does exactly that — a bare /, auto=1 without score=1, ?mode=raw, ?download=true, ?key=, the /automated and /locale export routes, and any parameter not on its allow-list each get a 403, and on the pages it does serve the report is never written to a global, so there is nothing for page.evaluate, devtools or a screenshot to read. That role is not merged into sundial's master and so is not deployed. When it is, set SUNDIAL_USERNAME=ci; nothing here changes, because this side already behaves as though the server were enforcing it.

The distinction is worth keeping straight: under guest, dropping score=1 by accident would put the whole report in the collector and leave redact() as the only thing between it and a public artifact. Under ci, the same mistake is a 403 at the first request. --allow-full-report exists for a deliberate local run under an account that is allowed one.

Blocking a merge

Branch protection on main requires exactly one check: All tests passed, the gate job. Pointing at one job instead of a dozen means the required-check list does not need editing every time a suite is added, renamed, or resharded.

The gate allows exactly three skips, each for a stated reason:

Skipped Because
build the browser was fetched, not compiled
fetch-browser the browser was compiled, not fetched
sundial disabled in ci/sundial.yml, or a fork pull request with no stealth credentials

Anything else that is not success fails it — including skipped. A suite that did not run has not passed, and quietly skipping one is the cheapest route to a green tick.

build is also the one suite dropped from --require when the browser was fetched rather than compiled: that job writes no result, and requiring a name nothing produces fails a run where everything passed.

One suite may additionally record a skip result without failing the run, named explicitly in --allow-skip: sundial, and only when sundial itself is unreachable. It is a separate service on a separate host, so an outage there means this browser was never measured — neither a pass nor a failure is true, and blocking every merge in the repository on someone else's downtime is the wrong answer. The summary shows it as skipped with the reason. A rejected credential, a role that would be served the private vectors, a full report where a score was requested, or a score under the floor all still fail: those are answers, and an answer gets judged.

The settings live in ci/branch-protection.json so they are reviewable rather than lore. To apply them (needs admin):

gh api -X PUT repos/<owner>/<repo>/branches/main/protection \
  --input ci/branch-protection.json

Two choices worth knowing about:

  • enforce_admins: false — you can still merge when CI itself is broken. Protection should stop mistakes, not lock you out of your own repository.
  • strict: false — a pull request does not have to be rebased onto the latest main before merging. With true, every push to main would invalidate every open pull request and force another build, and a build here is over an hour cold.

Reviews are deliberately not required: a solo maintainer cannot approve their own pull request, so requiring one would block every merge.

Cost control

Each tier gates the next, so a two-second lint failure never reaches the build:

0  static    lint, self-tests, settled decisions        seconds
1  unit      pythonlib, typescript                      ~1 min
2  browser   build  (patches/additions/settings/assets/upstream.sh/Makefile/scripts changed)
             fetch  (anything else, when the published release has this tree's browser sources)
3a smoke     patch guards, skiplist audit, build-tester,
             typescript-browser                         ~15 min
3b full      Playwright x6, leaks, stealth              ~40 min
4  gate      the required check

Driver-only pull requests test the published release, when it matches. There is nothing new to compile, so fetch-browser downloads the published release and the browser suites run against the build users are actually on — a minute instead of seventy. That is only right while the release was built from this tree's browser sources: once a browser change has merged but not been released, the guards in the checkout would judge an older browser. So the scope step compares the browser sources against the release tag, and when they differ it builds instead, which restores the base branch's cached browser.

Changing Juggler's JavaScript does not rebuild the browser. Measured on a real build: ccache reported a 98.63% hit rate, so almost none of those 24 minutes was compiling C++ — it was Rust, linking libxul, and packaging, none of which a .js file affects. And in the unpackaged dist/bin that CI archives there is no omni.ja at all: Juggler is loose files under chrome/juggler/, so delivering new JavaScript is a file copy.

So the build cache is keyed on a hash of the compiled inputs only (ci/browser_inputs.py). If that hash matches, the compiled half is identical by construction — no diff required, and no dependence on what the pull request base happened to contain — and this branch's resources are laid over the restored browser. A Juggler JavaScript change costs about a minute instead of twenty-four.

Two things make that dangerous, and both are closed and mutation-tested:

Trap What closes it
additions/juggler/ is not all JavaScript — it holds the screencast encoder and the debugging pipe (5 .cpp, 5 .h, 2 .idl, 3 components.conf, 4 moz.build) Only files jar.mn actually lists are resources. Everything else — including any extension nobody has considered yet — is native and forces a build. jar.mn itself is native, so removing an entry cannot leave a stale file behind
The source→destination mapping is per-file, not a prefix It is read from jar.mn. TargetRegistry.js → content/TargetRegistry.js (a level added), content/FrameTree.js → content/content/FrameTree.js (preserved), content/JugglerFrameChild.sys.mjs → content/JugglerFrameChild.sys.mjs (dropped). Two files in one source directory landing at different depths is exactly what a prefix rule gets wrong — and it would run stale Juggler while every suite went green

A browser that is already built is not built again. browser_changed is computed against the pull request's base, so it stays true for every push to a branch that touched patches/ even once. That is right — the published release does not contain that branch's browser changes, so it cannot be tested against — but taken alone it meant recompiling a byte-identical browser on every push, 24 minutes at a time, to fix a typo in ci/.

The build job therefore asks a narrower question first: not "does this branch change the browser" but "has the browser changed since the last one we built". The answer is a cache keyed on a hash of every input that can alter the binary — the same path list browser_changed uses, plus this workflow, which pins the toolchain. On a hit, a 634 MB camoufox-dist.tar.zst is restored and every build step is skipped; the run still records a build result saying the browser was restored rather than compiled, because a required suite that reports nothing fails the gate, and "nothing was compiled" should be a fact in the evidence rather than a hole in it.

Deliberately no restore-keys on that cache. Everywhere else a partial match is fine — a partly warm ccache is still warm — but here it would hand the test jobs a browser built from different sources, and every suite would report on it looking perfectly healthy. A self-test asserts the key covers every path browser_changed considers browser-affecting, so the two cannot drift apart.

The ccache is kept warm from main. Pushes to main populate it and a twice-weekly schedule keeps it from being evicted (GitHub drops a cache after seven days unused). Pull requests restore it through restore-keys, so a build in a pull request starts warm even though its own key is new.

A prebuilt image in ghcr.io with the object cache baked in would be warmer still and would not need the eviction guard. It also needs registry credentials and a rebuild pipeline of its own; this is the version that works with no setup.

Preparing the tree retries the network, and nothing else. mach bootstrap pulls toolchains from Taskcluster, and a connection reset there used to fail the whole pull request. ci.run_prepare runs setup-minimal → dir → mozbootstrap, retrying the two that download things and only when the failure text reads as transient. A failed patch hunk or a compile error still fails on the first attempt — retrying a broken tree only spends a runner to reach the same answer, and a retry loop that swallows a real breakage turns a red build into a slow red build. The release workflow (build.yml) prepares its tree the same way, so a tagged build gets the same hardening.

One consequence of cancel-in-progress: pushing to a branch cancels its running build. That is right while iterating, but a 70-minute build will not survive a push made 20 minutes in.

Running a piece by hand

python3 -m ci.run_prepare                                # make setup-minimal, dir, mozbootstrap
python3 -m ci.run_build
python3 -m ci.run_pythonlib                              # no browser needed
python3 -m ci.run_typescript                             # no browser needed
python3 -m ci.run_typescript     --browser path/to/camoufox-bin
python3 -m ci.run_patch_guards   --binary path/to/camoufox-bin
python3 -m ci.run_build_tester   --binary path/to/camoufox-bin
python3 -m ci.run_skiplist_audit --binary path/to/camoufox-bin
python3 -m ci.run_playwright     --binary path/to/camoufox-bin
python3 -m ci.run_playwright     --binary path/to/camoufox-bin --shard 3/6
python3 -m ci.run_native         --subset rules          # no browser needed
python3 -m ci.run_native         --subset browser --binary path/to/camoufox-bin
python3 -m ci.run_native         --subset growth  --binary path/to/camoufox-bin
python3 -m ci.run_sundial        --binary path/to/camoufox-bin
python3 -m ci.summarize          --results-dir .ci-work/results

Each suite runner writes one result file to .ci-work/results/ (run_prepare writes none). ci/summarize.py folds the shards, decides, and renders the table. A required suite that produced no result file is a failure, never a skip — otherwise deleting a job would be the cheapest way to a green tick.

Self-tests

ci/tests/ asserts the pipeline reports honestly: redaction leaks nothing, skips carry reasons, shards partition exactly once, version resolution never picks a suite newer than the browser. These run in the static job on every pull request.