* feat(ts): import the TypeScript launcher port from feat/captchakrakenAndJSSupport CAPTCHA support is left out; this branch is the JS/TS driver only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(ts): drop the CAPTCHA wiring left behind by the import Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(ts): port fpgen to TypeScript, on the same pinned model fpgen is not on npm. The port reads scripts/data/fpgen-model.json and checks its sha256 with TLS on, never fpgen's own first-release download. Everything that does not depend on the random draw is identical to Python (network, value lookups, trace probabilities, conditions, errors); the draws are held to Python's distributions by chi-square tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(ts): identity layer at parity with pythonlib A bit-exact port of CPython's random.Random, numpy's PCG64 choice and orjson's serialisation, so identity_salt/identity_seed and every seeded draw (fonts, voices, media devices, WebGL, noise seeds) come out identical to Python for the same identity. coherence.py, presets and screen/window fixes are ported, and golden fixtures recorded from pythonlib hold all of it to exact equality. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(ts): launcher at parity with pythonlib's launch_options launch_options() now produces pythonlib's output byte for byte (CAMOU_CONFIG, CAMOU_PREFS_N, prefs, env, fontconfig, warnings) over 89 recorded scenarios. Ports core pinning, geolocation, locales, fontprobe, the async API, and the pkgman/multiversion integrity checks. An opt-in e2e suite launches a real build through both launchers and compares what a page sees. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: test the TypeScript package, and publish it to npm like pypi - ci/run_typescript.py writes the `typescript` gate (typecheck, lint, vitest with the pythonlib golden tests) and, with --browser, `typescript_browser` (the e2e suite against the browser under test). Both are required by the gate. - publish-npm.yml mirrors publish-pypi.yml: workflow_dispatch, checks, build, scripts/check-pack.mjs (version == pythonlib, every data file shipped, the tarball installs and imports), then publish via npm trusted publishing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ts): e2e that holds on any browser, NewContext, and the package on every PR - Parity (TS == Python on the same binary) stays strict everywhere; whether the browser honours the config is asserted only on a binary whose properties.json knows every key the launcher sets, and otherwise skips naming the missing keys. A driver-only pull request is tested against the published release, which lags the launcher (beta.30 predates #779), so this is what makes the suite meaningful there instead of red on skew it cannot fix. - New: NewContext in a real browser -- a per-context identity that differs from the launch identity and from a sibling context, and equals Python's. - python_probe.py keeps stdout for its JSON (pythonlib prints "Skipping unknown patch" there), and a non-JSON reply now fails fast instead of hanging 240 s. - The typescript gate builds the package and runs scripts/check-pack.mjs, so a packaging mistake fails the pull request that makes it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ci): commit the launch fixtures, and fetch the browser from its real directory - The root .gitignore ignores every path named `launch` (local build output), which silently dropped typescript/tests/fixtures/launch/ -- the launch_options() goldens -- from the branch. Re-included in typescript/.gitignore. - fetch-browser read camoufox-bin from `camoufox path`, the cache ROOT, but multiversion installs each build under browsers/<channel>/<version>/, so the job has failed on every driver-only pull request since #772. It now resolves the active build as the launcher does (pkgman.camoufox_path). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci(ts): give the typescript job pythonlib, so the cross-language checks run Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ts): fpgen model install is safe across processes Processes installing into an empty cache at once each downloaded the model, and one's install deleted the values.dat another had just decompressed, which then failed its next lookup with ENOENT. Seen with vitest's parallel files on a cold cache; a worker pool on a fresh machine would hit it too. - ensureModel() installs under a cross-process lock (an atomic mkdir, stale after 10 min) and re-checks what is installed once it holds it. - values.dat is only removed when the model is actually being replaced. - The model keeps values.dat open, instead of reopening it on every lookup. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: make local runs and CI see the same test suite Three ways the suite passed here and not on the runner, each fixed at its cause: - The root .gitignore's bare `launch` rule (for the Go launcher binary) ignored every path segment named launch, so tests/fixtures/launch/ never reached git. Anchored to /launch; and ci/run_typescript.py now fails when any file under typescript/{src,tests,scripts} is git-ignored, which would have caught it on the machine that wrote the fixtures. - A missing prerequisite (fpgen model, pythonlib venv, fontTools, Xvfb, a font directory) skipped its tests, and a skip reads as green. tests/prereq.ts now fails them under CI unless the job names the gap in CAMOUFOX_TEST_ALLOW_MISSING. The typescript job installs all of them. The font-name check read one developer's local browser bundle; it now reads /usr/share/fonts (or CAMOUFOX_TEST_FONT_DIR), and CI installs a .ttc set. - The fpgen install race surfaced only on a cold cache, by accident. It now has deterministic tests: a same-model reinstall keeps values.dat (verified to fail on the old code), the lock admits one holder and releases on error, and a stale lock is reclaimed. Also: the browser gate runs only the e2e file, and the e2e probe and the virtual-display test time-box each await, so a hang names its step instead of reporting a bare 240 s timeout. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ts): hold headless="virtual" to Python's, not to headless On the runner (no media hardware), published beta.31 never settles enumerateDevices() in a headful window while headless answers -- the named timeout in the probe caught it. That is a browser property, so like the other page-vs-config checks it moves to a test that runs on a binary current with the launcher; the virtual-display test now requires the same page as Python's headless="virtual" on the same binary. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(build-tester): accept 18 and 22 cores, as real hardware reports plausibleHWC's list of common core counts lacked 18 and 22 -- Intel Meteor Lake laptops (Core Ultra 5 125H, Core Ultra 7 155H), and 22 is in 8 recorded presets. build-tester draws random presets, so a run that picked one of the two Linux presets reporting 22 failed: about one run in eleven, on any pull request. A CI self-test now fails if the list rejects any core count pythonlib can present (the presets and PLAUSIBLE_CORE_COUNTS). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ci): test on the published release only when it matches this tree A pull request that did not touch the browser was always tested against the published release. The patch guards and suites come from the checkout, so once a browser change was merged but not yet released (#779, on top of beta.31), every driver-only pull request ran #779's guards against a browser without #779 -- eight guards failed on #785, which changes no browser source. resolve now also compares the tree's browser sources with the tag the release was cut from (v<version>-<release> from upstream.sh), and builds when they differ or the tag does not exist. Building restores the base branch's cached browser when its compiled half matches -- main's #779 build, here -- so the extra cost is a cache restore, not a compile. Self-tests run the workflow's own scope step in a scratch repo for the four cases. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ts): say when CI's browser cannot present the configured locale With #785 finally tested on a browser current with the tree (main's cached #779 build), every TS-vs-Python parity check passed and the page-vs-config check failed: a de-DE/fr-FR identity presented en-US. Python presents the same on that binary. CI tests the build job's unpackaged dist/bin, whose res/multilocale.txt lists en-US only -- scripts/package.py injects the langpacks, and CI never packages. So no CI suite had ever run a non-English locale on a browser that has one. The e2e locale assertions now run when the binary under test packages the configured locale (read from res/multilocale.txt, loose or in omni.ja), and otherwise go through prerequisite("packaged-locales"), which fails in CI unless the job names the gap. The typescript (browser) job names it, with the reason; the rest of the page-vs-config check stays strict. On a packaged #779 build all of it, locale included, passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(guards): judge the query-cost probes on a median, not one sample stock-parity-probes timed each getter once. On a shared runner one GC pause or CPU-steal spike decided the verdict: navigator.hardwareConcurrency took 77 ms against a 50 ms allowance on the same restored build that passed the run before. Each pair is now timed five times, interleaved, and compared by median. The regressions these catch (a sync IPC per read, ~240 ms over the loop) cost extra on every read, so they move the median; verified by giving the getter a constant ~4 us of extra work per read -- 86 ms median, FAIL -- while the healthy build passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ts): give the headful e2e page focus before probing it enumerateDevices() intermittently never settled in the headless="virtual" test on CI (passed one run, timed out the next, same build). Firefox defers device enumeration until the document has focus -- LEAKS row 57 recorded the same for a background tab -- and headless mode fakes focus while a headful window on a bare Xvfb, with no window manager, only sometimes receives it. A user's window has focus, so both launchers' virtual-display probes now bring the page to the front first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore: remove build tooling nothing uses - The developer UI (scripts/developer.py, `make edits`). It depended on easygui, which no requirements file declares, and every action it offered is a Makefile target: patch, unpatch, workspace, revert, diff. Its two helpers in scripts/_mixin.py (is_bootstrap_patch, patch) had no other callers. - legacy/, the Go launcher deprecated in 2024-11. Nothing built or shipped it. Its Makefile targets and scripts/run-pw.py go with it, and so does Go from every dependency list and workflow. - jsonvv/ and settings/camoucfg.jvv. Nothing read the .jvv schema: config is validated against settings/properties.json, and the two had already drifted. The jsonvv package stays on PyPI. - Scripts with no caller: bootstrap.py, moztree, setup-wasi-linux.sh, package-helper.sh, install-local-build.sh, mozfetch.sh (copied into lw/ but never packaged), examples/. - The pre-ESM Juggler copies JugglerFrameParent.jsm and JugglerFrameChild.jsm, and hidden-scrollbars.css. Juggler loads the .sys.mjs actors and deliberately no stylesheet, but jar.mn still packaged all three. - patches/librewolf/*.opt, which list_patches() never picks up; the roverfox second pass in patch.py, whose directory no longer exists; the unread --no-settings-pane option. - The CAMOUFOX_PASSWD secret passed to `make fetch` and closedsrc_rev in upstream.sh, which nothing reads. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(python): remove dead helpers and a stale dependency None of these had a caller: - pkgman: is_supported_path, extract_zip, cleanup and set_version, left over from the single-directory install. cleanup() would have deleted every installed browser version. - multiversion.get_cached_repo_names, CONSTRAINTS.as_range, fingerprints._load_os_voices, utils._clean_locals, and unused imports. Also: - The "Apify Fingerprints" row in `camoufox version`, which has read "?" since fpgen replaced BrowserForge. - lxml is no longer a dependency; nothing imports it. - The geoip extra now names maxminddb, the module geolocation.py actually imports, rather than getting it transitively through geoip2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: show the cursor paths humanize=True actually produces The README's cursor video showed the Bezier generator Camoufox replaced with Cursory's recorded trajectories. scripts/cursor-demo.py drives a real build with humanize=True and records every mousemove event the page receives. It writes assets/humanize-cursor.svg, an animated replay at the recorded speed, so what the figure shows is what a site sees. The script cannot change the binary, so ci/browser_inputs.py lists it as non-native and editing it does not invalidate the cached browser. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(python): stop naming BrowserForge in user-facing text fpgen replaced BrowserForge, but two LeakWarnings, the NonFirefoxFingerprint message and the fingerprint_preset docstring still named it. One warning also linked to a README anchor that no longer exists. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: one AGENTS.md for every agent, a roadmap, and docs that match the code - AGENTS.md holds the engineering rules for any coding agent, plus the repo map, build, patch and test commands that CLAUDE.md used to carry. CLAUDE.md now only imports it, so there is one set of rules. ci/tribal-rules.yml is the record of settled decisions it points to. - ROADMAP.md lists planned work, each item linked to its issue. - README: - fpgen and the coherence check replace BrowserForge; - the patch workflow uses the make targets instead of the removed developer UI; - letter-spacing noise is described as off by default, as it is. - docs/: - beta-testing-ff146.md removed; - patch-upgrading-guide rewritten around the make targets; - per-context-patches without the canvas patch that no longer exists, and with measured preset counts; - playwright-maintenance without the JSM wrapper that does not exist; - smaller fixes in MEDIA-DEVICES, input-dispatch and FONTS. - ci/README: every job, and the real shard, skiplist and entry-point lists. - pythonlib, tester and patch-dependency READMEs corrected against the code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(pythonlib): handle headless='virtual' in launch_server launch_server() is documented to take the same arguments as Camoufox(), but passed headless='virtual' straight to launch_options(), so the server launched with no Xvfb display. Start a VirtualDisplay the way Camoufox() does, launch headful on it, and kill it when the server process exits or the launch fails. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(python): remove fontprobe, which nothing called fontprobe listed the fonts installed on the host, for a `camoufox fonts` command that was never added. It has nothing to do with the font bundle Camoufox serves to pages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(license): the Python launcher is MIT; the browser stays MPL-2.0 The Python package has always been published to PyPI as MIT (#727), but pythonlib/ shipped no licence file, and the repo's LICENSE is the browser's MPL-2.0. MPL is copyleft per file. It covers the modified Firefox sources, not a separate launcher that drives the browser over Playwright. So pythonlib/LICENSE now carries the MIT text its metadata already declares, and a Licensing section in the README says which part is which. Closes #727. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): fingerprint_preset=False no longer turns presets on launch_options checked `fingerprint_preset is not None`, so passing False drew a random bundled preset, the opposite of what was asked. It now uses a truthiness check, and a test proves that None and False never draw a preset. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: bring the release workflow in line with the tests build job build.yml had drifted from tests.yml. It ran actions at v1/v2 on a retired Node runtime, prepared the source tree with bare make calls that fail the whole release on one dropped connection, and built with a different Python than every pull request is tested with. - Pin every action by commit SHA, at the major versions tests.yml uses (checkout v4, setup-python v5, upload/download-artifact v4, the same remove-unwanted-software SHA), and action-gh-release v2. The release job holds contents: write, so it should not follow a movable tag. - Prepare the tree with `python3 -m ci.run_prepare`, as the tests build job does. BUILD_TARGET is set from the matrix so `make dir` writes the right mozconfig and Rust targets; multibuild.py then finds _READY and builds without re-patching. mach's toolchain bootstrap ignores the mozconfig, so running it after `dir` bootstraps the same toolchains. - Build with Python 3.12, the version the tests build job compiles with. - Default the workflow to no permissions; the build job gets contents: read and the release job keeps contents: write. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): NewContext looks up a proxy's exit IP through the right URL, or fails NewContext derives the context's WebRTC IP and timezone from the proxy's exit IP. That lookup had two defects, and both left the context showing the host's values while its traffic went through the proxy: - It built its own proxy URL with urlparse, which reads a scheme-less server such as "1.2.3.4:8080" (a form Playwright accepts) as scheme "1.2.3.4" with no host. urllib could not use a SOCKS proxy at all. - Any failure was swallowed, and the context opened without the values. The URL is now built with Proxy.as_string(), which the geoip launch path already uses (scheme-less means http). The lookup goes through requests, which handles SOCKS, and a failed lookup raises InvalidIP, naming the two options that skip it. The tests cover scheme-less, http and socks5 servers with credentials, both failure modes, and the case where no lookup is needed, for NewContext and AsyncNewContext. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: stop generating a canvas seed, and drop config keys nothing reads The browser has not noised the canvas since #528, and no patch reads canvas:seed (#721). The launcher still drew one on every launch and sent it through CAMOU_CONFIG, and NewContext called a setCanvasSeed that does not exist. They no longer do. For users this changes nothing on any browser since #528: the value was ignored. A config that still passes canvas:seed gets the usual "Skipping unknown patch" notice instead of silence. On a browser from before #528, the launcher no longer turns canvas noise on, which is the behaviour #528 chose. The same audit found more keys declared in settings/properties.json that no patch or Juggler file reads, so setting them did nothing: - canvas:aaOffset, canvas:aaCapOffset - memorysaver, pdfViewerEnabled, webrtc:localipv4/6 - navigator.onLine, navigator.cookieEnabled, navigator.languages - navigator.appCodeName, appName, product, productSub. Firefox reports these constants itself, so fpgen.yml no longer maps them. - webGl:parameters:blockIfNotDefined and its WebGL2 twin test_config_schema now checks this direction too: every declared key must be read by the browser, unless it is listed with a reason. Three are listed: locale:script and navigator.doNotTrack, which the launcher applies itself, and navigator.buildID (#780). The build-tester grading followed the same wrong premise. It tracked canvas collisions as an unfixed per-context leak. A canvas that is rendered rather than noised follows the fonts and GPU, as it does on real machines, so canvas collisions are now counted with the other device-level values. The tribal rule that recorded it as an open question is now a settled one, canvas-is-not-noised, with an automated check. Closes #721. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(ts): remove dead helpers None of these had a caller: - pkgman: isSupportedPath, extractZip, cleanup and setVersion, left over from the single-directory install. cleanup() would have deleted every installed browser version. - multiversion getCachedRepoNames and getCachedVersions, CONSTRAINTS.asRange, removeMmdb (Python keeps its twins for the GUI) and pycompat pySorted. - The "Apify Fingerprints" row in `camoufox version`, which read "?". utils.ts now calls noiseSeedsFromIdentity instead of repeating its two formulas inline, so the tested function is the one that runs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(ts): remove fontprobe, which nothing called fontprobe.ts listed the fonts installed on the host, for a `camoufox fonts` command neither launcher has. It has nothing to do with the font bundle Camoufox serves to pages. Its parity test goes with it, and so do the CI prerequisites only that test needed: fonttools and the extra font packages. (The Python twin is removed in #787.) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(ts): license the launcher MIT, with third-party notices The TypeScript launcher is a port of pythonlib, which has always been published to PyPI as MIT (#727); the MPL-2.0 of the browser covers the modified Firefox sources, not a launcher that drives it over Playwright. THIRD_PARTY_NOTICES.md ships in the npm package with the notices for the code the port translates: fpgen (Apache-2.0), CPython's random (the MT19937 BSD notice and the PSF licence), and NumPy's SeedSequence and PCG64 (BSD-3 and MIT). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: document the TypeScript package outside typescript/ The README, CONTRIBUTING, ci/README and the issue templates did not mention the npm package or its two CI gates. ci/README also still said driver-only pull requests never build. Since the scope step started comparing browser sources against the release tag, they build whenever the release is behind. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): repair devicePixelRatio the same way on every launch The DPR repair snaps an off-grid ratio to the nearest real scaling step and keeps the first of two equally near steps. The steps were frozenset literals, and a frozenset literal iterates in one order when the module is compiled from source and another when it is loaded back from a .pyc. So a midpoint such as 1.125 became 1.25 on the first launch after an install and 1 on every launch after it: the same pinned identity presented two different devicePixelRatio values. The steps are now ascending tuples, so a tie always goes to the lower step. The test runs the repair in two fresh interpreters that share a bytecode cache, compiling in the first and loading in the second. It failed before this change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ts): mirror #787's pythonlib fixes The TypeScript side of the behaviour #787 changes in pythonlib, so the port stays at parity: - fingerprint_preset=false no longer draws a preset. - NewContext builds the proxy URL with ProxyHelper.asString() (scheme-less means http), looks up the exit IP through impit, and throws InvalidIP when the lookup fails instead of opening the context with the host's values. - No canvas seed is generated or sent (#721). noiseSeedsFromIdentity becomes audioSeedFromIdentity, and fpgen's constant navigator fields are no longer mapped. - The devicePixelRatio steps are ascending, so a tie goes to the lower step. - The two LeakWarning texts that named BrowserForge. - The README's note that Python's launch_server() ignored headless='virtual' is gone, because it no longer does. The golden fixtures are regenerated from #787's pythonlib. The generator now masks the fontconfig file name the way the test already did. The name hashes content that embeds the checkout path, so every regeneration from a different checkout used to rewrite 76 fixtures. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: remove the glyph-spacing seed from the browser and the launcher anti-font-fingerprinting.patch added a seeded amount to every glyph advance, so that text widths differed per context. No real machine produces those widths: the same font on the same OS measures the same everywhere. So the noise was itself a fingerprint, measured in #779 at +1 px per ~100 glyphs plus fractional deltas on every measureText. #779 defaulted the seed to 0 and kept it as an opt-in, but an opt-in whose only effect is to become detectable is not worth carrying. Removed: - The browser side: - FontSpacingSeedManager and window.setFontSpacingSeed; - the HarfBuzz hook; - the plumbing that existed only to carry the context id down to the shaper: the userContextId on gfxTextRun, gfxShapedWord and the word-cache key, and the extra MakeTextRun argument in nsTextFrame, nsFontMetrics, MathML and canvas. The font group keeps its userContextId, which font-list-spoofing.patch uses to apply the per-context font list. Text is now shaped exactly as stock Firefox shapes it. - The fonts:spacing_seed key. The launcher had been sending 0 on every launch, plus a setFontSpacingSeed(0) call in every context's init script. - tests/patches/config-overrides.py, which tested only the spacing override. A pythonlib test now covers config_overrides with another key. timezone-spoofing, webrtc-ip-spoofing and window-setter-seal change only in context lines and the setter seal list. Every patch applies cleanly to a fresh tree, and the result builds. The settled decision is recorded as no-glyph-spacing-noise, with an automated check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: stock animations and speech by default; drop config keys that freeze live values Three behaviours a page could detect, changed in one breaking release: - **Animations run on stock timing.** no-css-animations.patch finished every finite animation at once by default, and any page could read it: `el.animate(frames, 1000).effect.getComputedTiming().duration` was 0, and a 500ms transition reported 0. Measured on v152.0.4-beta.31. The speedup is now an opt-in, `instantAnimations: True`, which raises a LeakWarning. disableInstantAnimations is gone. - **speak() on a spoofed voice works like a real voice.** It fired `error` after 3ms unless voices:fakeCompletion was set, and then start and end in the same tick. It now starts and ends after the text's duration at ~150 words per minute. Both voices:fakeCompletion keys are gone, and so is a debug line printed to stderr on every call. - **Keys removed:** - battery:* and window.scrollMinX/Y: Firefox keeps getBattery() and scrollMin* chrome-only, so no page could read them. - window.scrollMaxX/Y, screen.pageXOffset/pageYOffset, window.history.length and document.body.client*: each pinned a live value to a constant, so scrolling, navigating or re-laying out never changed it. fpgen.yml mapped pageYOffset, so about 15% of identities froze window.scrollY at a non-zero value. - The body keys' role as an undocumented alias for window.innerWidth/Height in browser-init and in the launcher. - MaskConfig::GetInt32Rect, which only the body keys used. New guards, both of which fail on v152.0.4-beta.31: tests/patches/animation-timing.py and tests/patches/spoofed-voice-speaks.py. The decisions are recorded as animations-run-on-stock-timing and spoofed-voices-speak. Every patch applies cleanly to a fresh tree, and the result builds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python)!: remove dead public API and make `list all --path` work Breaking changes: - Remove the exceptions UnknownProperty, InvalidDebugPort and MissingDebugPort. Nothing in the package raises them, so code catching them was catching nothing. - Remove the legacy `allow_webgl` keyword of launch_options(). Use `block_webgl=True`. The keyword now reaches Playwright as an unknown launch option and fails there instead of being silently consumed. `camoufox list all --path` accepted the flag and ignored it. It now prints the install path beside each installed build, as `camoufox list --path` already does for the installed tree. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(python)!: drop data the package never draws from voices.json shipped in the wheel, but no code in the package reads it: the voice draw uses voice-manifests.json and voice-uris.json. Its only readers are the TypeScript port's golden-fixture generator and data-sync script (typescript/scripts/golden/identity_golden.py, typescript/scripts/sync-identity-data.py), which live on another branch and will need a new source; the last copy is at 676fb3f:pythonlib/camoufox/voices.json. docs/per-context-patches.md described it as runtime data and now describes the files that are. webgl_data.db held two rows with zero weight on every OS ("Intel(R) HD Graphics 400, or similar" from "Intel Inc." and "Radeon R9 200 Series, or similar" from "ATI Technologies Inc."), left behind when their impossible macOS weights were zeroed. No draw can reach them. They are deleted with secure_delete so their blobs do not linger in free pages; the file is not vacuumed, so the other pages are unchanged. A new test requires every row to be drawable on at least one OS. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): warn whenever an identity falls back to a substitute value Several draws swallowed their failure and used something else, so an identity could ship with values the rest of it was not drawn to match and nobody would hear about it: - from_preset(): a failed font or voice draw used the preset's recorded list, or nothing, on any exception. - generate_context_fingerprint(): a failed font, voice or WebGL draw was `except Exception: pass`, leaving the browser's launch-time values. - _load_font_groups() / _load_font_bases(): an unreadable file became {}, i.e. no font additions or no OS-version base. - launch_options(): a failed font draw used every font in fonts.json, a failed voice draw used no voices, and a preset GPU missing from webgl_data.db was silently swapped for a drawn one (36 of the 397 bundled presets). Each site now catches only the errors its data can raise (OSError and ValueError for an unreadable or corrupt file, KeyError for a manifest with no entry for the OS, sqlite3.Error for the WebGL database) and emits a FallbackWarning. The text names what failed and what the identity uses instead, then gives a block to paste into an issue (camoufox, browser, OS and Python versions, the error, and the identity's user agent or GPU), asking the user to report it on GitHub. It shares LeakWarning's caller-frame attribution and its template lives in warnings.yml. The broad excepts had also been hiding a broken fixture: test_launch_environment's font and voice stubs did not accept `seed`, so every draw there raised and was swallowed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): give NewContext identities the browser's Firefox version NewContext() and AsyncNewContext() passed ff_version=None through to generate_context_fingerprint(), so a context's user agent kept the version fpgen drew (e.g. Firefox/146) while the browser underneath was 152. They now default ff_version to the major version of Playwright's Browser.version, which Juggler reports from MOZ_APP_VERSION_DISPLAY, so the UA always names the browser the page is actually talking to. An explicit ff_version still wins. The docstrings said each context gets "its own real fingerprint preset"; the default has been an fpgen draw, with a preset only when one is passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): send an IPv6 WebRTC address to setWebRTCIPv6 The per-context init script passed every webrtc_ip, IPv6 included, to window.setWebRTCIPv4(), and never called setWebRTCIPv6(). An IPv6 address (given directly, or resolved as a proxy's exit IP) was stored as the context's IPv4 value and the IPv6 slot stayed empty. The script now picks the setter by address family, and an address that is neither raises InvalidIP instead of being passed through. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): stop pinning the page's scroll offset from fpgen fpgen.yml mapped the drawn window.pageYOffset (e.g. 528) to screen.pageYOffset, and the browser returns that value from scrollY on every read, so a page saw one scroll position forever whatever the user did. Real scroll offsets are live page state, not part of a device's fingerprint, so neither offset is mapped any more. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(python): stop checking a config key that no longer exists warn_manual_config() looked for navigator.languages, which was removed from settings/properties.json; validate_config() rejects it before the check could matter. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): close the WebGL database connection on every path sample_webgl raised its not-found and wrong-OS errors before reaching conn.close(), leaking a sqlite connection each time a preset named a GPU the database does not hold. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(data): drop the 23 presets whose GPU has no WebGL data A preset records only its GPU's name. The WebGL parameters, extensions and shader precision behind it have to come from somewhere, and for these 23 nothing Camoufox has describes the GPU: fpgen has never seen Firefox report it on that OS. So each launch paired the name with another device's parameters, a mismatch any WebGL fingerprinter can see. They were: - Windows on ARM (Adreno 650); - Direct3D 10-level GPUs (vs_4_0/vs_4_1); - "Generic Renderer"; - 945GM and GTX 480 on macOS; - nouveau/Mesa buckets on Linux; - one Linux preset pairing NVIDIA's proprietary vendor string with the nouveau renderer name. scripts/clean-fingerprint-data.py now applies the rule, via a shared fingerprints.firefox_gpus(), and test_shipped_data asserts it. 374 presets remain, and every OS keeps its presets. ROADMAP.md lists capturing WebGL data for these GPUs, which would bring them back. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ts): stop generating the glyph-spacing seed Mirrors676fb3f: the browser no longer has glyph-spacing noise, so the launcher sends no fonts:spacing_seed and the per-context init script no longer calls setFontSpacingSeed. config_overrides is now tested with audio:seed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ts): warn on instantAnimations; stop treating body keys as window size Mirrors the launcher half offb21b2e: instantAnimations raises the instant_animations LeakWarning (warnings.yml copied from pythonlib), the document.body.client* keys no longer count as window dimensions, and fpgen's pageYOffset is no longer mapped, so no identity freezes window.scrollY. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ts)!: remove dead public API and make `list all --path` work Mirrors4616aa5: drop the UnknownProperty, InvalidDebugPort and MissingDebugPort exceptions (nothing raises them) and the legacy allow_webgl option (use block_webgl; allow_webgl now passes through to Playwright like any unknown option). `camoufox list all --path` prints each installed build's path, as `list --path` already did. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(ts)!: drop voices.json, which nothing draws from Mirrors9a42b6a. The voice draw reads voice-manifests.json and voice-uris.json; voices.json was only read by the golden generator and the data-sync script. The voice-URI golden now hashes the URI of every entry in voice-manifests.json, the list the draw actually picks from. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ts): warn whenever an identity falls back to a substitute value Mirrors165e68ffor the font and voice draws. Each fallback site in fromPreset(), generateContextFingerprint(), loadFontGroups(), loadFontBases() and launchOptions() now catches only the errors its data can raise and emits a FallbackWarning naming what failed, what the identity uses instead, and a block to paste into an issue (camoufox, browser, OS and Node versions, the error, the identity). The message is warnings.yml's `fallback` template, shared with pythonlib. Python's except clauses name builtin classes JavaScript lacks, so pycompat gains OSError, ValueError and KeyError twins and isPyError(): a Node system error counts as an OSError and JSON.parse's SyntaxError as a ValueError, as json.JSONDecodeError is. The voice draw now throws ValueError for a malformed entry and KeyError when the manifest has no macOS entry, as Python does. The WebGL fallback sites are left for the change that replaces the TS WebGL source. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ts): give NewContext identities the browser's Firefox version Mirrors46e5c8d: without an explicit ff_version, NewContext() kept the Firefox version fpgen drew, so a context's UA could name 146 on a 152 browser. It now defaults to the major version of Browser.version(). The option docs now say the default identity is an fpgen draw. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ts): send an IPv6 WebRTC address to setWebRTCIPv6 Mirrors0bbd152: the per-context init script passed every WebRTC IP to setWebRTCIPv4(), IPv6 included. It now picks the setter by address family and raises InvalidIP for an address that is neither. The init-script golden gains an IPv6 case. Also ports 7b43112's regression test: a drawn pageXOffset/pageYOffset is not carried into the config (the mapping went in fb21b2e's mirror). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(ts): stop checking a config key that no longer exists Mirrors480789a: warnManualConfig() looked for navigator.languages, which settings/properties.json no longer has; validateConfig() rejects it first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(ts): sync the pruned presets and the properties.json fixture Copies fingerprint-presets*.json from pythonlib (40edebbdropped the 23 presets whose GPU has no WebGL data) and refreshes the launch fixture's copy of settings/properties.json, which lost the keys removed in676fb3fandfb21b2e. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(python): draw every identity's WebGL from fpgen WebGL vendor, renderer, context attributes, extensions, parameters and shader precisions, for WebGL1 and WebGL2, now come from fpgen's recorded Firefox devices instead of webgl_data.db, which is deleted with the camoufox/webgl/ package. camoufox/webgl.py: - webgl_for_gpu() traces `webgl` given Firefox, the OS and the GPU, then `webgl2` given the chosen `webgl` too, and draws each with one seeded random.Random. The GPU and the webgl value are pinned by their fpgen lookup index: a dict condition is flattened into leaves that overwrite each other, so only the renderer applied and Linux "Mesa" and "AMD" Radeon HD 3200 devices came back mixed. - sample_webgl_for_screen() draws the GPU of a generated identity from fpgen's per-OS weights, filtering out software rasterisers, GPUs the OS cannot report, discrete GPUs behind a netbook screen and the resistFingerprinting "Mozilla" mask before the weighted choice, so there is no rejection loop. An empty pool raises. - The draft/host-dependent extension filter moves over unchanged. A preset's GPU and a caller's webgl_config pair are looked up as given; a pair fpgen has never seen from Firefox on that OS raises instead of falling back to another GPU. generate_context_fingerprint no longer falls back to the host GPU when the draw fails. For 10 of the 15 (GPU, OS) pairs the two sources share, one of fpgen's records converts to exactly the database row on every value the browser reads. The other five rows (Linux R9 200 and Radeon HD 3200, macOS Intel HD, and two software rasterisers) are devices fpgen does not carry; those GPUs now present fpgen's recorded devices instead. The Linux GTX 980 row is kept as a test fixture. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: say where WebGL comes from now that the database is gone The per-context guide, the fpgen.yml header and coherence's comments still named webgl_data.db and sample_webgl(). They now point at camoufox/webgl.py and fpgen. The guide also claimed presets carry WebGL parameters; they record only the vendor and renderer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ts): regenerate the golden fixtures for the mirrored pythonlib changes Regenerates identity_golden.py and launch_golden.py output from this tree's pythonlib. launch_golden.py drops fonts:spacing_seed and document.body.clientWidth from its inputs, adds an instantAnimations scenario, and masks a FallbackWarning's report block in both launchers, since it names the host and the runtime. Two launch scenarios still differ: preset_windows_unknown_gpu and config_webgl_unknown_pair expect the FallbackWarning pythonlib now raises when a preset's GPU is missing from the WebGL data. That site belongs to the change replacing the TS WebGL source; the rest of each scenario matches. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(ts): draw every identity's WebGL from fpgen Mirrors0d859b3. src/webgl.ts is the twin of camoufox/webgl.py: webglForGpu() traces fpgen's webgl node given Firefox, the OS and the GPU, then webgl2 given the chosen webgl too; sampleWebglForScreen() first draws the GPU from fpgen's per-OS weights, filtered (no software rasteriser, no resistFingerprinting mask, a GPU the OS can report, no discrete GPU behind a netbook screen) before one weighted choice. Every draw is PyRandom.choices on one seeded instance, in Python's order. The GPU and webgl value are pinned by their fpgen lookup index, found from the value's stored JSON, which TraceResult now carries: re-serialising a parsed value would spell 2**64 differently from orjson. A preset's GPU and a caller's webgl_config are looked up as given and raise when fpgen has never seen them, instead of falling back to another GPU; generateContextFingerprint no longer swallows a failed draw. Removed with the old source: webgl/sample.ts, the numpy default_rng port (webgl/nprandom.ts), data-files/webgl_data.json and its export in sync-identity-data.py. The NumPy notice now covers the pairwise sum in locales.ts, the one NumPy port left. launchOptions now throws the pycompat ValueError where Python raises ValueError. Tests port test_webgl.py and the shipped-data check that every preset GPU has WebGL data. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ts): pin the fpgen WebGL draws against Python bit for bit identity_golden.py now records camoufox.webgl's output as sha256 of the exact orjson bytes (key order, int versus float and all): 1272 screen draws over every OS and seven screens (60 seeds each, plus four large seeds), webgl_for_gpu for every GPU fpgen records and every bundled preset GPU, the unknown-GPU and unknown-OS errors, and to_config's extension filter. The numpy golden keeps only np.sum, now checked against locales.ts. The renderer list the coherence golden walks comes from fpgen's traces. launch_golden.py takes its WebGL pairs from firefox_gpus(), and the unknown-GPU preset input is a GPU nobody records. The launch goldens are regenerated; the preset_windows_unknown_gpu and config_webgl_unknown_pair scenarios now expect Python's ValueError. The golden test no longer maps ValueError to Error. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(patches): a host's missing speech daemon no longer errors spoofed speech On a Linux host where speech-dispatcher cannot start, Firefox broadcasts synth-voices-error, and SpeechSynthesis answers it by firing `error` on every queued utterance. So a spoofed Windows voice errored about 11ms into speak() on any host without the daemon: the CI runners, and most servers. It passed only where the daemon runs. While Camoufox manages the voice list, the registry no longer forwards a host backend's error. The spoofed voices do not depend on the host's engine, and a Windows or macOS identity never raises one. The guard now makes the daemon unreachable itself, so it tests this case on every machine; on the previous build it fails every time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: stop skipping the two click tests that stock animation timing fixed test_wait_for_stable_position and test_timeout_waiting_for_stable_position were skipped with humanized travel time as the reason. The real cause was instant animations. Every finite animation finished at once, so the button Playwright waits on to stop moving never moved, and the click landed where upstream does not expect. With animations on stock timing both pass, and the skiplist audit flagged them as no longer failing. The entries go, and the counts in ci/README.md drop from 14 to 12. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(build-tester): accept 18 and 22 cores, as real hardware reports plausibleHWC's list of common core counts lacked 18 and 22 -- Intel Meteor Lake laptops (Core Ultra 5 125H, Core Ultra 7 155H), and 22 is in 8 recorded presets. build-tester draws random presets, so a run that picked one of the two Linux presets reporting 22 failed: about one run in eleven, on any pull request. A CI self-test now fails if the list rejects any core count pythonlib can present (the presets and PLAUSIBLE_CORE_COUNTS). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(guards): judge the query-cost probes on a median, not one sample stock-parity-probes timed each getter once. On a shared runner one GC pause or CPU-steal spike decided the verdict: navigator.hardwareConcurrency took 77 ms against a 50 ms allowance on the same restored build that passed the run before. Each pair is now timed five times, interleaved, and compared by median. The regressions these catch (a sync IPC per read, ~240 ms over the loop) cost extra on every read, so they move the median; verified by giving the getter a constant ~4 us of extra work per read -- 86 ms median, FAIL -- while the healthy build passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(native): compare the whole fingerprint when two launches must differ test_two_browsers_get_different_fingerprints compared seven coarse values: UA, platform, screen size, core count, timezone and language. CI pins the timezone and language, and real machines share the rest: two draws of a common Mac (Firefox 152, MacIntel, 2560x1440, 8 cores) matched, and the test failed on a correct browser. It now reads the whole fingerprint a site computes, from a script in the page: - navigator values, screen and window geometry, device pixel ratio, timezone; - WebGL vendor, renderer, limits and extensions; - installed fonts, measured by width against the generic fallbacks; - voices, media-device counts, and an OfflineAudioContext hash. The page is served from an https URL Playwright fulfils locally, because mediaDevices exists only in a secure context. The page computes the result itself because the isolated world may not read audio sample data. The test then requires the fingerprints to differ, and the audio hash to differ on its own, since its noise is seeded per identity. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Update README to remove warning, camoufox is now actively maintained Camoufox will now be actively maintained and improved for the foreseeable future * ci(ts): hold the golden tests to live pythonlib, not a snapshot of it The typescript job ran the golden tests against the fixtures committed in typescript/tests/fixtures/. A pythonlib change that typescript/ did not mirror left those fixtures untouched, so the tests kept passing -- the opposite of what the job's comment promised. `ci.run_typescript --regenerate-golden` now rewrites the fixtures from the checkout's pythonlib before vitest runs, and CI passes it. The job runs Python 3.14 because pySum() reproduces sum() as 3.14 computes it; on 3.12 one crafted mixed int/float case differs in its last bit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci(ts): do not redraw fpgen's stats fixture on every run stats.json is thousands of random fpgen draws that the TS tests compare statistically. It changes with the pinned model, not with pythonlib, and redrawing it took eight of the typescript job's eleven minutes on a runner. The deterministic fpgen fixtures are still regenerated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(native): measure crash growth from a warmed-up parent test_the_parent_stays_flat_across_content_crashes took its baseline right after launch. The parent's first context costs it 150-200 MB with no crash at all, so the warm-up counted as crash growth: 330-370 MB of the 400 MB allowance locally, and 469 MB on a CI runner, failing a PR that changes nothing in the browser. The baseline now follows one clean context; each crash still has to stay within the same allowance. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(ts): read pythonlib's data files instead of copying them typescript/src/data-files/ held byte-identical copies of eleven pythonlib data files (23k lines), kept in step by a sync script and a test that failed when a copy drifted. The TS launcher now reads them from pythonlib/camoufox/ when it runs from the repo, and `pnpm build` copies them into dist/data-files/ for the npm tarball, so what users install is unchanged. DATA_FILES in src/paths.ts is the one list; check-pack.mjs checks each is in the tarball. The essential-font lists were the one large table both ports hard-coded (~170 lines of Python, ~770 of TS). They move to pythonlib/camoufox/essential-fonts.json, which fingerprints.py and fingerprints.ts both read; gen-fonts-json.py --print-bases writes it and verify-fonts.py checks it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ts): record the golden fixtures from pythonlib on every run The goldens were pythonlib's output committed as 14k lines of fixtures, most of it the same input fingerprint pasted into ~80 launch scenarios, and CI already rewrote them from pythonlib before each run. They are now recorded by a vitest globalSetup (tests/golden-setup.ts) from the repo's .venv or $CAMOUFOX_PYTHON, in about 6 seconds, and git-ignored. A pythonlib change that typescript/ does not mirror fails `pnpm test` locally as well as in CI, and ci.run_typescript no longer needs --regenerate-golden. Committed inputs stay: launch/inputs.json, the bundle stubs, the addon, e2e/probe.js, and fpgen/stats.json (random draws tested statistically, which change with the pinned model, not with pythonlib, and take minutes to redraw). The one Python-version-sensitive case, sum() over mixed ints and floats, skips with a named prerequisite when the goldens come from Python < 3.14. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci(npm): give the publish job the pythonlib its tests record goldens from Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ts): NewContext hands Playwright the identity's user agent, DPR and timezone generateContextFingerprint already returns Playwright's JS option names; NewContext re-cased them, and camelCase() lowercases first, so userAgent, deviceScaleFactor and timezoneId became keys Playwright silently drops. navigator.userAgent was still spoofed at the C++ level, so the page probe matched Python's, but the HTTP User-Agent, the DPR and the timezone did not. The options now pass through as generated. NewContext also awaits ensureModel(): a browser from connect() or a custom executable never went through launchOptions(), which fetches the fpgen model, so an fpgen draw threw ModelNotInstalled where Python works. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ts): a failed browser download rejects instead of killing the process The download's file stream had no 'error' listener, so a failed write (disk full) was an uncaught 'error' event: Node exited before installVersioned()'s catch could remove the partial install and its temp directory, and the caller had nothing to catch. finished() now listens from the moment the stream is created, and webdl() surfaces an errored stream instead of writing into it. webdl() also waits for 'drain': it ignored write()'s return value, so on a slow disk the whole archive queued in memory. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ts): concurrent launches no longer read or inherit each other's CPU pin playwright-core spawns the browser from this process, so pin_cpu_cores narrows this process's own mask while a launch's browser starts. Python pins a separate driver and never sees its own mask change. Here, a second launch started in that window: - read the pinned mask as the host's cores (pinnedCoreCount, and the identity's hardwareConcurrency via availableParallelism()), and - if it did not pin, spawned its browser without the lock, inheriting the first launch's pin while reporting more cores. The host's core count is now read once, before this process first pins itself (cpu_affinity.hostCoreCount), and unpinned launches, launchServer included, take the pin lock once any launch in the process has pinned. With pin_cpu_cores off (the default) nothing waits. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ts): an Xvfb that fails to start throws CannotExecuteXvfb A spawn that fails after spawn() returns (EACCES, ENOENT) is an 'error' event on the child process, which had no listener: Node treated it as uncaught and exited instead of get() throwing. The child now always has a listener, readDisplayNumber() rejects with CannotExecuteXvfb on it, and the display pipe keeps an error listener after the read settles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(scripts): gen-fonts-json refuses to write an incomplete essential-fonts.json An OS with no bases in the manifest was skipped, and the file written without its key; fingerprints.py and fingerprints.ts read every OS's list at import, so `import camoufox` then failed with a KeyError. The script now exits and keeps the existing file. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(release): pythonlib and the npm package to 0.5.7 @camoufox/camoufox 0.5.6 went to npm before the review fixes above, so they ship as 0.5.7; the two launchers are versioned in lockstep, and main already carries pythonlib changes from #787 that 0.5.6 on PyPI does not have. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ts): NewContext's HTTP User-Agent must match navigator.userAgent The e2e parity check read only navigator.userAgent, which the browser spoofs itself, so a context that dropped Playwright's userAgent option still matched Python. The probe server now records the request's User-Agent header. Against the NewContext before the fix, a context whose navigator said Windows sent the launch identity's Linux UA. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ts): draw the NewContext option-name test's preset from the v150 bundle getRandomPreset() without a Firefox version draws from the older bundle, where some Windows presets carry no devicePixelRatio, so the test failed on some draws in CI. Every v150 preset has one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ts): wait for the headful e2e page to have focus, not one bringToFront() On a bare Xvfb, with no window manager, a single bringToFront() before goto() sometimes left the window unfocused, and Firefox holds enumerateDevices() until the document has focus, so the virtual-display probe timed out on some runs. Both launchers' probes now navigate first, then bring the page to the front until document.hasFocus() is true, and fail with that reason if it never is. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The test pipeline
Everything that runs a suite against Camoufox. Driven identically from a pull
request, a push to main, and — through workflow_call — any caller that needs
to test a specific browser version, so there is one definition of "the tests
pass", not two.
resolve ── static ─────────────── lint, tribal rules, skiplist, self-tests (seconds)
├─ typescript ────── type check, lint, vitest, golden parity
└─ pythonlib ─────── the package's own tests (a minute)
└─ build or fetch ─┬─ patch guards ─────── one per spoofing patch, + skiplist audit
├─ build-tester ─────── 8 fingerprint profiles
├─ typescript-browser ─ the npm launcher end to end
└─ once guards and build-tester pass:
├─ playwright × 6 shards (conformance + our own)
├─ native ───────── leaks, contexts, crash recovery
├─ sundial ──────── stealth grade
└─ growth ───────── memory growth (scheduled only)
│
summary ──► one comment on the PR
Which browser, which suite
ci/versions.py answers both, and every entry point uses it:
- browser — from
upstream.sh, or whatever a caller passes in. A caller moving to a new Firefox passes the version it is moving to, which is what lets one pipeline test both a pull request and an upgrade. - suite — the newest released playwright-python tag whose pinned Firefox is
not ahead of that browser, and which is below the Playwright ceiling
pythonlib/pyproject.tomlpins.
Newest-not-ahead, rather than an exact match, because Playwright trails Firefox and skips generations: it pinned Firefox 151 and then 153, never 152, so a browser built on 152 has no exact suite and never will. Requiring a match would leave most of a release cycle with no suite at all, and taking a newer one would test against an automation contract that assumes engine work the build does not have. So the suite's Firefox pin being a release or two behind the browser is the normal case, not a misconfiguration — the summary line says which rule picked the tag.
The ceiling matters for the same reason it exists in pythonlib: camoufox.server
imports playwright._impl._driver, a private API, and every Playwright minor is
free to change Juggler. Testing above the ceiling would exercise a client the
shipped package will not install.
python3 -m ci.versions --json # what would run
python3 -m ci.versions --browser-version 153.0.4 --json
The version under test has to be the version that gets built. Only suite
selection follows --browser-version; the build reads upstream.sh and the
fetch path downloads whatever pythonlib considers current. Asking for a version
the branch does not pin would therefore compile the old browser and judge it
against the new suite — green, meaningless and silent. --check-upstream
refuses that, and the workflow passes it.
So an upgrade to a new Firefox is a branch that edits upstream.sh, which is
what an upgrade is anyway. Resolution then reads it by default, the build
produces it, and the suite is chosen for it — the three cannot disagree. The
browser_version input exists for a caller that wants to state the version
explicitly; it must match.
The Playwright suite
One suite, fetched fresh per run: upstream playwright-python at the resolved tag.
It runs unmodified — ci/pw_camoufox_plugin.py adapts the environment around it
rather than editing it, hooking BrowserType at the _impl layer so upstream
can refactor its fixtures freely — and ci/suite.py overlays tests/camoufox/
into it.
tests/camoufox/ is small on purpose: behaviour upstream has no test for (that
page.route() must not change what a request looks like on the wire), or asserts
the opposite of on purpose (that a worker should not inherit the context
locale, which stock Firefox gets wrong and Camoufox does not). It is not a fork
of anything, so it cannot go stale; it runs against upstream's own conftest and
server at whatever tag was resolved.
tests/ used to hold a fork of a ~v1.55-era upstream suite. It was deleted after
measuring it against the same binary:
- 73 of the 74 tests it skipped as "Not supported by Camoufox" pass in upstream's copy. Those skips predated main-world execution and were never revisited, so the fork was asserting the browser was worse than it is.
- Of its passing tests, eight had no upstream counterpart. Six of those were
Camoufox-specific and now live in
tests/camoufox/as three modules; the other two were tests upstream had since renamed.
A re-fetched suite cannot drift, and a deliberate difference from upstream now
has to be written down in ci/skiplist.yml with a reason, where it is visible,
instead of being encoded as a silent edit to a vendored file.
What "the suite" means
ci/run_playwright.py names its targets explicitly rather than pointing at
tests/, so the one thing left out stays visible:
tests/async/ + tests/sync/ |
one pytest process |
tests/common/, tests/test_reference_count_async.py |
their own process |
tests/test_installation.py |
excluded, with a reason |
The isolated pair each call sync_playwright()/async_playwright() inside the
test body, which cannot start while the session fixtures already hold a loop
(Cannot run the event loop while another loop is running). Run with the others
all six fail; run alone all six pass. Skiplisting them for that would have
recorded a browser failure that does not exist.
This used to be tests/async/ alone — 722 tests, 31% of the suite, excluded with
nothing written down. Not a decision: the vendored fork carried async/ and no
sync suite, and this runner was pointed at the same shape without checking what
upstream shipped. unclaimed() now fails the run if upstream adds a test path
that is neither in TARGETS nor in EXCLUDED with a reason.
Which world, and the skip list
The suite runs isolated first — the configuration Camoufox actually ships — and falls back to the main world only for what fails, counting every test that needed the fallback.
Upstream asserts upstream semantics: tests read globals their own page scripts
defined and pass handles into evaluate(), and a test doing that fails under
isolation by design. Running the whole suite main-world-only (the previous
behaviour) made those pass, which is true but uninformative — it measured a mode
nobody ships and produced no number for what isolation costs.
Each group is therefore run up to three times, and normally twice:
| Pass | CI_WORLD |
What it establishes |
|---|---|---|
| 1 | isolated |
the browser as users run it |
| 2 | main |
the failures again with isolation off |
| 3 | main |
only what failed in both, retried once — normally empty |
There is deliberately no second isolated pass. One sat between 1 and 2 on the theory that a flake must not be mistaken for a world difference; measured on the first real CI run it cost 7m50s per shard and recovered nothing:
isolated (full) 335s + 331s 35 and 11 failures
isolated retry 205s + 265s 0 recovered <- deleted
main world 19s + 12s 46 recovered
Two reasons it was never going to earn that. These failures are deterministic —
a test reading a global its page script defined does not intermittently see it —
and failing that way is slow, because the read returns undefined and the test
sits on a Playwright timeout rather than throwing. And upstream's suite already
ships pytest-rerunfailures: pass 1 reported 105 rerun, which is each of
those 35 failures having been retried three times before the run even reported
them. A flake does not survive that.
A test that passes in pass 2 is recorded as a main-world fallback: it counts
as a pass for the run, and its identity goes into
metrics.main_world_fallbacks, with metrics.main_world_fallback_count on the
summary table (summed across shards). A test failing in both worlds and on
retry is a plain failure. The fallback count is the isolated-world conformance
gap — watch it between runs; a jump means the isolation boundary moved.
Both rerun passes are guarded on the failure set being non-empty, which is
load-bearing rather than tidy: pytest declines to filter when nothing it
collected previously failed, so an unguarded --last-failed runs the whole
group again — in the other world, silently replacing the result it was meant to
refine.
Why some isolated failures hang
Not every isolated failure fails. Some wait forever, in both
tests/async/test_route_web_socket.py and tests/sync/test_route_web_socket.py
— and the two are not equally recoverable, which is the subject of the second
half of this section.
The shape recurs, so it is worth stating generally: a Playwright feature implemented by installing something on the page's global lands in the isolated world instead, so anything the page itself originates never reaches the automation. Two instances, both measured directly against a build:
page's own script opens a WebSocket isolated -> handler never fires main -> intercepted
page script calls window.exposedFn() isolated -> HANG main -> resolves
evaluate() calls window.exposedFn() isolated -> resolves main -> resolves
route_web_socket works by replacing window.WebSocket from an init script;
isolated, that replacement lands in the sandbox, so a socket the page opens is
never seen. expose_function installs its binding on the sandbox global, so page
script calling window.fn() finds nothing — though called from evaluate() it
works, which is why that one does not hang here.
They hang rather than fail because the waits involved — a Twisted future from the test server, an asyncio future a binding was meant to resolve — have no Playwright timeout behind them. Everything else isolation breaks fails at Playwright's 30s.
Worth being clear that the route_web_socket half is not a test artifact: a
real site's WebSocket is not intercepted either, and the user gets no error
saying so. It is not fixable at this layer — the feature works by replacing a
page global, which is precisely what an isolated world exists to stop a page
from seeing. Tracked in #775; the fix is native interception below
the DOM object, which is also the only version of it that stays undetectable.
So pass 1 is cost-bounded, and only pass 1:
| Bound | Value | Why |
|---|---|---|
| per-test timeout | 90s | the slowest test in the whole main-world baseline was 30.4s; only two exceeded 30s and none exceeded 45s. A Playwright action times out at 30s |
| upstream reruns | off (CI="") |
tests/conftest.py sets reruns = 3 whenever $CI is set — the only thing it reads $CI for. The baseline recorded 2 reruns across all 2295 tests; the isolated pass recorded 138 in one shard, all re-running deterministic world differences |
Together that turns a hang from up to 4 × 180s into one 90s wait. A flake missed
by not rerunning is not lost — it fails pass 1, passes pass 2, and is counted as
a fallback. Note that --reruns 0 as an argument would not work: upstream's
conftest overwrites config.option.reruns in pytest_configure, so clearing
the environment variable is the only lever that holds.
The ones a timeout cannot bound
That bound is not enough for all of them, and it is worth knowing exactly where it stops working. Measured on run 34799668707 with the 90s bound already in place:
| Group | Isolated pass | Outcome |
|---|---|---|
tests/async/ |
completed in 296s | the bound works |
tests/sync/ |
test_should_work_with_ws_close printed pytest-timeout's +++ Timeout +++ banner at exactly 90s |
the process then sat for 1h50m, until the job's timeout-minutes killed it |
So the signal fires and the test dies; the process does not. pytest-timeout's
signal method raises at the next bytecode boundary, and Playwright's sync API is
parked in a greenlet switch that never reaches one cleanly — the raise lands
inside the dispatcher and wedges it. --timeout-method=thread fires reliably but
kills the interpreter, taking the other ~1500 tests in the group with it. There
is no per-test timeout value that bounds this.
So those modules are declared, not discovered — ISOLATION_HANGS in
ci/run_playwright.py. The isolated pass cannot learn that they hang without
hanging, so it is told: they are --ignored out of pass 1 and run directly in
the main world (pass 1b), where they pass and are counted as fallbacks exactly
as if isolation had failed them honestly. The same tests still run, in the world
that can run them.
Why not ci/skiplist.yml. That list means "fails in the most permissive
world", and ci/run_skiplist_audit.py enforces it by running every entry with
CI_WORLD=main and failing the build on any that pass. A route_web_socket
test passes there — the main world is precisely where the feature works — so an
entry would be rejected by the audit, and would be untrue as written. The two
lists are not interchangeable, and test_isolation_hangs_are_not_in_the_skiplist
keeps them apart.
Pass 1b sits above the if not failing: continue guard, deliberately: a
group whose isolated pass found nothing would otherwise skip it, and coverage
would disappear on exactly the runs that look healthiest.
tests/patches/isolated-evaluate.py still owns the direct coverage of isolated
evaluation, and must keep passing regardless. That file is what to check if
isolation itself regresses.
ci/run_skiplist_audit.py deliberately runs in the main world: a skiplist
entry has to claim a test cannot pass in either world, or the suite would have
counted it as a fallback rather than a failure.
Fourteen tests are deselected outright by ci/skiplist.yml, which
requires a stated reason per entry — ci/summarize.py fails the run on an
unreasoned one.
A reason is not evidence, so the reasons are checked. The first version of
this file inherited all nine tests/async/*.disabled modules from the vendored
suite and gave each a plausible justification without running any of them: of
the 202 tests it skipped, 193 passed, and seven of the nine modules failed
nothing at all. A written reason made them look verified, which is worse than
leaving them bare.
ci/run_skiplist_audit.py now runs every entry with the skiplist disabled and
fails the build if a skipped test passes. It is cheap precisely because a
correct skiplist is short — twelve tests, a few seconds — and it is what keeps the
list from drifting back into a place failing tests go to disappear.
python3 -m ci.run_skiplist_audit --binary /path/to/camoufox-bin
What remains, 12 tests: two test_keyboard.py
tests that assert a shifted character arrives without Shift, which Camoufox
presses as a real keyboard would; six client-certificate tests (async and sync)
that need the browser to present a certificate during the TLS handshake —
the two that go through the Node driver's own request context instead pass, and
are not skipped; two upstream expectations that encode a stock-Firefox quirk,
replaced by tests/camoufox/; and two popup tests that rely on Playwright
shipping Firefox's popup blocker off, which Camoufox keeps on.
That client-certificate split is the audit earning its place. The entry was first written as a whole module, because on a local machine all five fail — Node/OpenSSL there rejects the fixture server outright. In CI two of them pass, and the audit failed the build one run after the entry was written. CI is the authority for what fails; a local run is a hypothesis.
Camoufox's own suite
native-tests/ covers what the Playwright suite cannot ask about:
- Leaks. Launch browsers, kill them, prove nothing survived — file descriptors, sockets, child processes, X11 lock files. The real assertion is that cost does not scale with launch count, because that is the shape a leak actually has: a scraper that runs fine for six hours and then dies of EMFILE. Scope is honest: this measures resources held by our process and its children, not Gecko's internal heap.
- Contexts versus browsers. Two contexts in one browser must get different
fingerprints; two pages in one context must get the same one. Get this wrong
and per-context injection silently degrades to process-global — which passes
every single-context test there is. It has happened here before (commit
d17c887, "fix screen size leak in contexts"). - Crashes. Kill the browser, the X server, a content process or the driver
mid-run, then check that teardown does not hang, nothing leaks, and a fresh
launch still works (
test_crash_recovery.py). - Memory growth. Drive one mechanism (iframes, canvas readback, WebGL
contexts, workers, script compilation, font measurement) N and 4N times and
compare the growth: a bounded cost stays flat, a per-iteration leak scales
(
test_memory_growth.py). It takes over half an hour, so it runs on the schedule and on demand (thegrowthjob,--subset growth), not in the gate. - Settled decisions.
ci/tribal-rules.ymllists choices this project already made, each with the issue or PR that made it, andnative-tests/test_tribal_rules.pyasserts them. A comment explaining a decision only works on someone who reads it.
Sundial
On since 2026-09-12. sundial 0.5.0 is deployed and serving score mode, and sundial's master branch now deploys itself on push, so merged does mean deployed. It was off for as long as the live build predated score mode: an older sundial ignores
?score=1and posts the entire report — every vector's id, name, brief, source and value — to whatever collector asked. Receiving that on a public runner and discarding it afterwards is not the guarantee this section describes; not receiving it is.enabled: falseinci/sundial.ymlis still the kill switch, and the resolve job checks it before the credential comes into scope, so flipping it back stops the request rather than just the reporting.
The stealth check reports a letter grade and a count. Nothing else leaves
ci/run_sundial.py::redact() — not a vector name, description, measured value,
source, and not a per-category breakdown either: a table reading "Graphics 3/17"
is the most useful single fact an adversary could take from a public CI log.
There are no per-check rows in the results file at all — not even opaque
ones. An HMAC does not name a vector, but a map of them still publishes how many
distinct checks fail and lets a reader follow one across releases, which is
per-vector data wearing a hash. The instruction was a score, so it is a score:
regression detection is per-score, via min_pass_rate and a maximum allowed
drop. Scope and thresholds live in ci/sundial.yml; only
categories Camoufox actually claims are gated.
Everything that leaves redact() is checked against a whitelist at runtime,
not a blacklist — a blacklist only stops the leaks somebody already thought of.
Adding a field without adding it to _PUBLISHABLE fails the run:
{ "grade": "A", "checks_total": 412, "checks_passed": 403, "pass_rate": 0.978,
"out_of_scope_failed": 6, "cross_os_total": 24, "cross_os_passed": 5,
"os": "linux", "sundial_version": "0.3.1", "schema_version": 1 }
Cross-OS detectors are counted, never scored. They read the host machine rather than the disguise: a browser claiming macOS while running on Linux fails them however good its spoofing is, and Camoufox does not claim byte-identical cross-OS emulation. Folding them into one average would mark it down for a promise nobody made, and would hide a real regression behind noise it cannot control. They are reported separately so a drop there reads as "the host shows through more than it did", which is a different conversation.
The run asks sundial for ?auto=1&score=1, so it receives counts and the
vectors never cross the wire at all.
Finding out which checks failed
Worth being precise about, because the answer is "you can't, from CI", and that
is deliberate rather than an oversight. Score mode's payload is buckets keyed
"<Category>|<class>" holding two integers each. It carries no check names and
no ids, so a failing check's identity is not something the CI process discards
— it is something sundial never sends. Nothing in the artifact, the log, or the
sealed report can recover it.
Two steps down from there, both local only:
# which CATEGORY the failures are in -- works with the credential CI already has
python3 -m ci.run_sundial --explain --binary /path/to/camoufox-bin
# which CHECKS -- needs a role sundial serves full reports to
python3 -m ci.run_sundial --explain --allow-full-report --binary /path/to/camoufox-bin
--explain prints to the terminal and never writes to a result file, and is
refused outright under GITHUB_ACTIONS: a category-level table is not a
vector, but "Graphics 3/17" is still the most useful single fact an adversary
could take from a public log, which is exactly why redact() does not publish
one.
Order of operations
?score=1 needs a sundial that has it. An older deployment ignores the unknown
parameter and posts the whole report; the numbers still come out right and
redact() still discards everything identifying, but nothing is classified,
so every cross-OS tally reads 0 — which looks like "no host-OS failures"
rather than "nobody sorted them". score_mode: false in the result says which
it is, and the gate says so in its notes rather than leaving you to notice.
So the dependency runs one way, and setting the GitHub secrets is the last step, not the first:
- deploy sundial's score mode — done; it ships in 0.5.0, and master now deploys on push
gh secret set SUNDIAL_AUTOMATION_KEY -R <repo>for every repository whose CI runs this. Without it the job skips, which is the normal case for a fork pull request- flip
enabled: trueinci/sundial.yml— done
Two things keep a vector out of a public log, and it is worth separating them, because only one is enforced by the server:
| guarantee | |
|---|---|
| server-side | The run logs in as guest, and sundial's middleware refuses guest the private-vector bundle outright — those definitions are never served to the session. |
| client-side | This gate only ever requests /?auto=1&score=1, and redact(require_score_mode=True) fails the run if a full report arrives anyway, rather than folding it down and carrying on. |
The stricter option is sundial's score-only ci role, which is refused anything
but /?auto=1&score=1 server-side and so cannot be handed a report even if the
credential leaks. That role is not in sundial's master branch and is therefore
not deployed; when it lands, mint the credential (make pages-ci, then
redeploy) and set SUNDIAL_USERNAME=ci. Nothing in this repository changes —
the client-side half already behaves as though the server were enforcing it.
There are deliberately no per-vector rows, not even opaque ones. An HMAC names nothing, but a map of them publishes how many distinct checks fail and lets a reader follow the same id from release to release.
The cost is real: regression detection drops from per-vector ("the check that
passed last release fails now") to per-score ("we got worse"), covered by
min_pass_rate in ci/sundial.yml — and, once an auto-update pipeline exists
to compare releases, by a maximum allowed drop in its policy file. To get the per-vector view back for your own debugging,
set SUNDIAL_REPORT_AGE_RECIPIENT to an age public key — the full report is
then kept encrypted to you and nobody else can open it.
Needs SUNDIAL_AUTOMATION_KEY (the password). SUNDIAL_USERNAME is
optional and names the account, which is not a secret — it defaults to guest.
Absent the password — a pull request from a fork — the job is skipped and the
summary says so.
What actually stops a vector reaching the log
Asking for ?score=1 is a promise the caller makes, and a promise is not a
mechanism. Today two things back it:
guestcannot load the private vectors, and that is checked. sundial's middleware answersisPrivateVectorAssetpaths with an empty stub for that role specifically.adminandprivatedo get them — and/automated?key=resolves toprivatewhen handed the private key, which is indistinguishable from the guest one by looking at it. So "we set the right key" stays an assumption until something checks: the gate reads sundial's own/__auth/meand refuses to open the browser at all unless the session is a role the vectors are withheld from. Not knowing the role counts as not safe.- A non-score payload fails the run.
redact(require_score_mode=True)refuses to process a full report rather than folding it down, so a deployment that ignoredscore=1is a red build, not a quiet leak.
What is still missing is a server that refuses the request. sundial's
score-only ci role does exactly that — a bare /, auto=1 without score=1,
?mode=raw, ?download=true, ?key=, the /automated and /locale export
routes, and any parameter not on its allow-list each get a 403, and on the pages
it does serve the report is never written to a global, so there is nothing for
page.evaluate, devtools or a screenshot to read. That role is not merged into
sundial's master and so is not deployed. When it is, set SUNDIAL_USERNAME=ci;
nothing here changes, because this side already behaves as though the server
were enforcing it.
The distinction is worth keeping straight: under guest, dropping score=1 by
accident would put the whole report in the collector and leave redact() as the
only thing between it and a public artifact. Under ci, the same mistake is a
403 at the first request. --allow-full-report exists for a deliberate local
run under an account that is allowed one.
Blocking a merge
Branch protection on main requires exactly one check: All tests passed,
the gate job. Pointing at one job instead of a dozen means the required-check
list does not need editing every time a suite is added, renamed, or resharded.
The gate allows exactly three skips, each for a stated reason:
| Skipped | Because |
|---|---|
build |
the browser was fetched, not compiled |
fetch-browser |
the browser was compiled, not fetched |
sundial |
disabled in ci/sundial.yml, or a fork pull request with no stealth credentials |
Anything else that is not success fails it — including skipped. A suite
that did not run has not passed, and quietly skipping one is the cheapest route
to a green tick.
build is also the one suite dropped from --require when the browser was
fetched rather than compiled: that job writes no result, and requiring a name
nothing produces fails a run where everything passed.
One suite may additionally record a skip result without failing the run,
named explicitly in --allow-skip: sundial, and only when sundial itself is
unreachable. It is a separate service on a separate host, so an outage there
means this browser was never measured — neither a pass nor a failure is true,
and blocking every merge in the repository on someone else's downtime is the
wrong answer. The summary shows it as skipped with the reason. A rejected
credential, a role that would be served the private vectors, a full report where
a score was requested, or a score under the floor all still fail: those are
answers, and an answer gets judged.
The settings live in ci/branch-protection.json so
they are reviewable rather than lore. To apply them (needs admin):
gh api -X PUT repos/<owner>/<repo>/branches/main/protection \
--input ci/branch-protection.json
Two choices worth knowing about:
enforce_admins: false— you can still merge when CI itself is broken. Protection should stop mistakes, not lock you out of your own repository.strict: false— a pull request does not have to be rebased onto the latestmainbefore merging. Withtrue, every push tomainwould invalidate every open pull request and force another build, and a build here is over an hour cold.
Reviews are deliberately not required: a solo maintainer cannot approve their own pull request, so requiring one would block every merge.
Cost control
Each tier gates the next, so a two-second lint failure never reaches the build:
0 static lint, self-tests, settled decisions seconds
1 unit pythonlib, typescript ~1 min
2 browser build (patches/additions/settings/assets/upstream.sh/Makefile/scripts changed)
fetch (anything else, when the published release has this tree's browser sources)
3a smoke patch guards, skiplist audit, build-tester,
typescript-browser ~15 min
3b full Playwright x6, leaks, stealth ~40 min
4 gate the required check
Driver-only pull requests test the published release, when it matches.
There is nothing new to compile, so fetch-browser downloads the published
release and the browser suites run against the build users are actually on — a
minute instead of seventy. That is only right while the release was built from
this tree's browser sources: once a browser change has merged but not been
released, the guards in the checkout would judge an older browser. So the scope
step compares the browser sources against the release tag, and when they
differ it builds instead, which restores the base branch's cached browser.
Changing Juggler's JavaScript does not rebuild the browser. Measured on a
real build: ccache reported a 98.63% hit rate, so almost none of those 24
minutes was compiling C++ — it was Rust, linking libxul, and packaging, none of
which a .js file affects. And in the unpackaged dist/bin that CI archives
there is no omni.ja at all: Juggler is loose files under chrome/juggler/, so
delivering new JavaScript is a file copy.
So the build cache is keyed on a hash of the compiled inputs only
(ci/browser_inputs.py). If that hash matches, the compiled half is identical
by construction — no diff required, and no dependence on what the pull request
base happened to contain — and this branch's resources are laid over the
restored browser. A Juggler JavaScript change costs about a minute instead of
twenty-four.
Two things make that dangerous, and both are closed and mutation-tested:
| Trap | What closes it |
|---|---|
additions/juggler/ is not all JavaScript — it holds the screencast encoder and the debugging pipe (5 .cpp, 5 .h, 2 .idl, 3 components.conf, 4 moz.build) |
Only files jar.mn actually lists are resources. Everything else — including any extension nobody has considered yet — is native and forces a build. jar.mn itself is native, so removing an entry cannot leave a stale file behind |
| The source→destination mapping is per-file, not a prefix | It is read from jar.mn. TargetRegistry.js → content/TargetRegistry.js (a level added), content/FrameTree.js → content/content/FrameTree.js (preserved), content/JugglerFrameChild.sys.mjs → content/JugglerFrameChild.sys.mjs (dropped). Two files in one source directory landing at different depths is exactly what a prefix rule gets wrong — and it would run stale Juggler while every suite went green |
A browser that is already built is not built again. browser_changed is
computed against the pull request's base, so it stays true for every push to a
branch that touched patches/ even once. That is right — the published release
does not contain that branch's browser changes, so it cannot be tested against —
but taken alone it meant recompiling a byte-identical browser on every push, 24
minutes at a time, to fix a typo in ci/.
The build job therefore asks a narrower question first: not "does this branch
change the browser" but "has the browser changed since the last one we built".
The answer is a cache keyed on a hash of every input that can alter the binary —
the same path list browser_changed uses, plus this workflow, which pins the
toolchain. On a hit, a 634 MB camoufox-dist.tar.zst is restored and every
build step is skipped; the run still records a build result saying the browser
was restored rather than compiled, because a required suite that reports nothing
fails the gate, and "nothing was compiled" should be a fact in the evidence
rather than a hole in it.
Deliberately no restore-keys on that cache. Everywhere else a partial
match is fine — a partly warm ccache is still warm — but here it would hand the
test jobs a browser built from different sources, and every suite would report
on it looking perfectly healthy. A self-test asserts the key covers every path
browser_changed considers browser-affecting, so the two cannot drift apart.
The ccache is kept warm from main. Pushes to main populate it and a
twice-weekly schedule keeps it from being evicted (GitHub drops a cache after
seven days unused). Pull requests restore it through restore-keys, so a build
in a pull request starts warm even though its own key is new.
A prebuilt image in ghcr.io with the object cache baked in would be warmer
still and would not need the eviction guard. It also needs registry credentials
and a rebuild pipeline of its own; this is the version that works with no setup.
Preparing the tree retries the network, and nothing else. mach bootstrap
pulls toolchains from Taskcluster, and a connection reset there used to fail the
whole pull request. ci.run_prepare runs setup-minimal → dir →
mozbootstrap, retrying the two that download things and only when the failure
text reads as transient. A failed patch hunk or a compile error still fails on
the first attempt — retrying a broken tree only spends a runner to reach the
same answer, and a retry loop that swallows a real breakage turns a red build
into a slow red build. The release workflow (build.yml) prepares its tree the same
way, so a tagged build gets the same hardening.
One consequence of
cancel-in-progress: pushing to a branch cancels its running build. That is right while iterating, but a 70-minute build will not survive a push made 20 minutes in.
Running a piece by hand
python3 -m ci.run_prepare # make setup-minimal, dir, mozbootstrap
python3 -m ci.run_build
python3 -m ci.run_pythonlib # no browser needed
python3 -m ci.run_typescript # no browser needed
python3 -m ci.run_typescript --browser path/to/camoufox-bin
python3 -m ci.run_patch_guards --binary path/to/camoufox-bin
python3 -m ci.run_build_tester --binary path/to/camoufox-bin
python3 -m ci.run_skiplist_audit --binary path/to/camoufox-bin
python3 -m ci.run_playwright --binary path/to/camoufox-bin
python3 -m ci.run_playwright --binary path/to/camoufox-bin --shard 3/6
python3 -m ci.run_native --subset rules # no browser needed
python3 -m ci.run_native --subset browser --binary path/to/camoufox-bin
python3 -m ci.run_native --subset growth --binary path/to/camoufox-bin
python3 -m ci.run_sundial --binary path/to/camoufox-bin
python3 -m ci.summarize --results-dir .ci-work/results
Each suite runner writes one result file to .ci-work/results/ (run_prepare
writes none). ci/summarize.py folds the shards, decides, and renders the table. A required suite that produced no result
file is a failure, never a skip — otherwise deleting a job would be the
cheapest way to a green tick.
Self-tests
ci/tests/ asserts the pipeline reports honestly: redaction leaks nothing,
skips carry reasons, shards partition exactly once, version resolution never
picks a suite newer than the browser. These run in the static job on every
pull request.