Files
camoufox/docs/FONTS.md
T
Jake WriterandClaude Opus 5.5 180e6553cc Docs that match the code, one AGENTS.md, dead code removed, and the audit's bug fixes (#787)
* chore: remove build tooling nothing uses

- The developer UI (scripts/developer.py, `make edits`). It depended on
  easygui, which no requirements file declares, and every action it offered is
  a Makefile target: patch, unpatch, workspace, revert, diff. Its two helpers in
  scripts/_mixin.py (is_bootstrap_patch, patch) had no other callers.
- legacy/, the Go launcher deprecated in 2024-11. Nothing built or shipped it.
  Its Makefile targets and scripts/run-pw.py go with it, and so does Go from
  every dependency list and workflow.
- jsonvv/ and settings/camoucfg.jvv. Nothing read the .jvv schema: config is
  validated against settings/properties.json, and the two had already drifted.
  The jsonvv package stays on PyPI.
- Scripts with no caller: bootstrap.py, moztree, setup-wasi-linux.sh,
  package-helper.sh, install-local-build.sh, mozfetch.sh (copied into lw/ but
  never packaged), examples/.
- The pre-ESM Juggler copies JugglerFrameParent.jsm and JugglerFrameChild.jsm,
  and hidden-scrollbars.css. Juggler loads the .sys.mjs actors and deliberately
  no stylesheet, but jar.mn still packaged all three.
- patches/librewolf/*.opt, which list_patches() never picks up; the roverfox
  second pass in patch.py, whose directory no longer exists; the unread
  --no-settings-pane option.
- The CAMOUFOX_PASSWD secret passed to `make fetch` and closedsrc_rev in
  upstream.sh, which nothing reads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(python): remove dead helpers and a stale dependency

None of these had a caller:
- pkgman: is_supported_path, extract_zip, cleanup and set_version, left over
  from the single-directory install. cleanup() would have deleted every
  installed browser version.
- multiversion.get_cached_repo_names, CONSTRAINTS.as_range,
  fingerprints._load_os_voices, utils._clean_locals, and unused imports.

Also:
- The "Apify Fingerprints" row in `camoufox version`, which has read "?"
  since fpgen replaced BrowserForge.
- lxml is no longer a dependency; nothing imports it.
- The geoip extra now names maxminddb, the module geolocation.py actually
  imports, rather than getting it transitively through geoip2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: show the cursor paths humanize=True actually produces

The README's cursor video showed the Bezier generator Camoufox replaced
with Cursory's recorded trajectories. scripts/cursor-demo.py drives a real
build with humanize=True and records every mousemove event the page
receives. It writes assets/humanize-cursor.svg, an animated replay at the
recorded speed, so what the figure shows is what a site sees.

The script cannot change the binary, so ci/browser_inputs.py lists it as
non-native and editing it does not invalidate the cached browser.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(python): stop naming BrowserForge in user-facing text

fpgen replaced BrowserForge, but two LeakWarnings, the NonFirefoxFingerprint
message and the fingerprint_preset docstring still named it. One warning also
linked to a README anchor that no longer exists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: one AGENTS.md for every agent, a roadmap, and docs that match the code

- AGENTS.md holds the engineering rules for any coding agent, plus the
  repo map, build, patch and test commands that CLAUDE.md used to carry.
  CLAUDE.md now only imports it, so there is one set of rules.
  ci/tribal-rules.yml is the record of settled decisions it points to.
- ROADMAP.md lists planned work, each item linked to its issue.
- README:
  - fpgen and the coherence check replace BrowserForge;
  - the patch workflow uses the make targets instead of the removed
    developer UI;
  - letter-spacing noise is described as off by default, as it is.
- docs/:
  - beta-testing-ff146.md removed;
  - patch-upgrading-guide rewritten around the make targets;
  - per-context-patches without the canvas patch that no longer exists,
    and with measured preset counts;
  - playwright-maintenance without the JSM wrapper that does not exist;
  - smaller fixes in MEDIA-DEVICES, input-dispatch and FONTS.
- ci/README: every job, and the real shard, skiplist and entry-point lists.
- pythonlib, tester and patch-dependency READMEs corrected against the code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(pythonlib): handle headless='virtual' in launch_server

launch_server() is documented to take the same arguments as Camoufox(),
but passed headless='virtual' straight to launch_options(), so the server
launched with no Xvfb display. Start a VirtualDisplay the way Camoufox()
does, launch headful on it, and kill it when the server process exits or
the launch fails.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(python): remove fontprobe, which nothing called

fontprobe listed the fonts installed on the host, for a `camoufox fonts`
command that was never added. It has nothing to do with the font bundle
Camoufox serves to pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(license): the Python launcher is MIT; the browser stays MPL-2.0

The Python package has always been published to PyPI as MIT (#727), but
pythonlib/ shipped no licence file, and the repo's LICENSE is the browser's
MPL-2.0. MPL is copyleft per file. It covers the modified Firefox sources, not
a separate launcher that drives the browser over Playwright. So
pythonlib/LICENSE now carries the MIT text its metadata already declares, and
a Licensing section in the README says which part is which.

Closes #727.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): fingerprint_preset=False no longer turns presets on

launch_options checked `fingerprint_preset is not None`, so passing False
drew a random bundled preset, the opposite of what was asked. It now uses a
truthiness check, and a test proves that None and False never draw a preset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: bring the release workflow in line with the tests build job

build.yml had drifted from tests.yml. It ran actions at v1/v2 on a
retired Node runtime, prepared the source tree with bare make calls
that fail the whole release on one dropped connection, and built with
a different Python than every pull request is tested with.

- Pin every action by commit SHA, at the major versions tests.yml uses
  (checkout v4, setup-python v5, upload/download-artifact v4, the same
  remove-unwanted-software SHA), and action-gh-release v2. The release
  job holds contents: write, so it should not follow a movable tag.
- Prepare the tree with `python3 -m ci.run_prepare`, as the tests build
  job does. BUILD_TARGET is set from the matrix so `make dir` writes the
  right mozconfig and Rust targets; multibuild.py then finds _READY and
  builds without re-patching. mach's toolchain bootstrap ignores the
  mozconfig, so running it after `dir` bootstraps the same toolchains.
- Build with Python 3.12, the version the tests build job compiles with.
- Default the workflow to no permissions; the build job gets
  contents: read and the release job keeps contents: write.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): NewContext looks up a proxy's exit IP through the right URL, or fails

NewContext derives the context's WebRTC IP and timezone from the proxy's exit
IP. That lookup had two defects, and both left the context showing the
host's values while its traffic went through the proxy:

- It built its own proxy URL with urlparse, which reads a scheme-less server
  such as "1.2.3.4:8080" (a form Playwright accepts) as scheme "1.2.3.4" with
  no host. urllib could not use a SOCKS proxy at all.
- Any failure was swallowed, and the context opened without the values.

The URL is now built with Proxy.as_string(), which the geoip launch path
already uses (scheme-less means http). The lookup goes through requests,
which handles SOCKS, and a failed lookup raises InvalidIP, naming the two
options that skip it. The tests cover scheme-less, http and socks5 servers
with credentials, both failure modes, and the case where no lookup is needed,
for NewContext and AsyncNewContext.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: stop generating a canvas seed, and drop config keys nothing reads

The browser has not noised the canvas since #528, and no patch reads
canvas:seed (#721). The launcher still drew one on every launch and sent it
through CAMOU_CONFIG, and NewContext called a setCanvasSeed that does not
exist. They no longer do.

For users this changes nothing on any browser since #528: the value was
ignored. A config that still passes canvas:seed gets the usual "Skipping
unknown patch" notice instead of silence. On a browser from before #528, the
launcher no longer turns canvas noise on, which is the behaviour #528 chose.

The same audit found more keys declared in settings/properties.json that no
patch or Juggler file reads, so setting them did nothing:
- canvas:aaOffset, canvas:aaCapOffset
- memorysaver, pdfViewerEnabled, webrtc:localipv4/6
- navigator.onLine, navigator.cookieEnabled, navigator.languages
- navigator.appCodeName, appName, product, productSub. Firefox reports these
  constants itself, so fpgen.yml no longer maps them.
- webGl:parameters:blockIfNotDefined and its WebGL2 twin

test_config_schema now checks this direction too: every declared key must be
read by the browser, unless it is listed with a reason. Three are listed:
locale:script and navigator.doNotTrack, which the launcher applies itself,
and navigator.buildID (#780).

The build-tester grading followed the same wrong premise. It tracked canvas
collisions as an unfixed per-context leak. A canvas that is rendered rather
than noised follows the fonts and GPU, as it does on real machines, so canvas
collisions are now counted with the other device-level values. The tribal rule
that recorded it as an open question is now a settled one,
canvas-is-not-noised, with an automated check.

Closes #721.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): repair devicePixelRatio the same way on every launch

The DPR repair snaps an off-grid ratio to the nearest real scaling step and
keeps the first of two equally near steps. The steps were frozenset
literals, and a frozenset literal iterates in one order when the module is
compiled from source and another when it is loaded back from a .pyc. So a
midpoint such as 1.125 became 1.25 on the first launch after an install and
1 on every launch after it: the same pinned identity presented two different
devicePixelRatio values.

The steps are now ascending tuples, so a tie always goes to the lower step.
The test runs the repair in two fresh interpreters that share a bytecode
cache, compiling in the first and loading in the second. It failed before
this change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: remove the glyph-spacing seed from the browser and the launcher

anti-font-fingerprinting.patch added a seeded amount to every glyph advance,
so that text widths differed per context. No real machine produces those
widths: the same font on the same OS measures the same everywhere. So the
noise was itself a fingerprint, measured in #779 at +1 px per ~100 glyphs
plus fractional deltas on every measureText. #779 defaulted the seed to 0
and kept it as an opt-in, but an opt-in whose only effect is to become
detectable is not worth carrying. Removed:

- The browser side:
  - FontSpacingSeedManager and window.setFontSpacingSeed;
  - the HarfBuzz hook;
  - the plumbing that existed only to carry the context id down to the
    shaper: the userContextId on gfxTextRun, gfxShapedWord and the word-cache
    key, and the extra MakeTextRun argument in nsTextFrame, nsFontMetrics,
    MathML and canvas.
  The font group keeps its userContextId, which font-list-spoofing.patch
  uses to apply the per-context font list. Text is now shaped exactly as
  stock Firefox shapes it.
- The fonts:spacing_seed key. The launcher had been sending 0 on every
  launch, plus a setFontSpacingSeed(0) call in every context's init script.
- tests/patches/config-overrides.py, which tested only the spacing override.
  A pythonlib test now covers config_overrides with another key.

timezone-spoofing, webrtc-ip-spoofing and window-setter-seal change only in
context lines and the setter seal list. Every patch applies cleanly to a
fresh tree, and the result builds. The settled decision is recorded as
no-glyph-spacing-noise, with an automated check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: stock animations and speech by default; drop config keys that freeze live values

Three behaviours a page could detect, changed in one breaking release:

- **Animations run on stock timing.** no-css-animations.patch finished every
  finite animation at once by default, and any page could read it:
  `el.animate(frames, 1000).effect.getComputedTiming().duration` was 0, and a
  500ms transition reported 0. Measured on v152.0.4-beta.31. The speedup is
  now an opt-in, `instantAnimations: True`, which raises a LeakWarning.
  disableInstantAnimations is gone.
- **speak() on a spoofed voice works like a real voice.** It fired `error`
  after 3ms unless voices:fakeCompletion was set, and then start and end in
  the same tick. It now starts and ends after the text's duration at ~150
  words per minute. Both voices:fakeCompletion keys are gone, and so is a
  debug line printed to stderr on every call.
- **Keys removed:**
  - battery:* and window.scrollMinX/Y: Firefox keeps getBattery() and
    scrollMin* chrome-only, so no page could read them.
  - window.scrollMaxX/Y, screen.pageXOffset/pageYOffset,
    window.history.length and document.body.client*: each pinned a live value
    to a constant, so scrolling, navigating or re-laying out never changed it.
    fpgen.yml mapped pageYOffset, so about 15% of identities froze
    window.scrollY at a non-zero value.
  - The body keys' role as an undocumented alias for window.innerWidth/Height
    in browser-init and in the launcher.
  - MaskConfig::GetInt32Rect, which only the body keys used.

New guards, both of which fail on v152.0.4-beta.31:
tests/patches/animation-timing.py and tests/patches/spoofed-voice-speaks.py.
The decisions are recorded as animations-run-on-stock-timing and
spoofed-voices-speak. Every patch applies cleanly to a fresh tree, and the
result builds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python)!: remove dead public API and make `list all --path` work

Breaking changes:

- Remove the exceptions UnknownProperty, InvalidDebugPort and
  MissingDebugPort. Nothing in the package raises them, so code catching
  them was catching nothing.
- Remove the legacy `allow_webgl` keyword of launch_options(). Use
  `block_webgl=True`. The keyword now reaches Playwright as an unknown
  launch option and fails there instead of being silently consumed.

`camoufox list all --path` accepted the flag and ignored it. It now prints
the install path beside each installed build, as `camoufox list --path`
already does for the installed tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(python)!: drop data the package never draws from

voices.json shipped in the wheel, but no code in the package reads it: the
voice draw uses voice-manifests.json and voice-uris.json. Its only readers
are the TypeScript port's golden-fixture generator and data-sync script
(typescript/scripts/golden/identity_golden.py,
typescript/scripts/sync-identity-data.py), which live on another branch and
will need a new source; the last copy is at
676fb3f:pythonlib/camoufox/voices.json. docs/per-context-patches.md
described it as runtime data and now describes the files that are.

webgl_data.db held two rows with zero weight on every OS ("Intel(R) HD
Graphics 400, or similar" from "Intel Inc." and "Radeon R9 200 Series, or
similar" from "ATI Technologies Inc."), left behind when their impossible
macOS weights were zeroed. No draw can reach them. They are deleted with
secure_delete so their blobs do not linger in free pages; the file is not
vacuumed, so the other pages are unchanged. A new test requires every row
to be drawable on at least one OS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): warn whenever an identity falls back to a substitute value

Several draws swallowed their failure and used something else, so an
identity could ship with values the rest of it was not drawn to match and
nobody would hear about it:

- from_preset(): a failed font or voice draw used the preset's recorded
  list, or nothing, on any exception.
- generate_context_fingerprint(): a failed font, voice or WebGL draw was
  `except Exception: pass`, leaving the browser's launch-time values.
- _load_font_groups() / _load_font_bases(): an unreadable file became {},
  i.e. no font additions or no OS-version base.
- launch_options(): a failed font draw used every font in fonts.json, a
  failed voice draw used no voices, and a preset GPU missing from
  webgl_data.db was silently swapped for a drawn one (36 of the 397
  bundled presets).

Each site now catches only the errors its data can raise (OSError and
ValueError for an unreadable or corrupt file, KeyError for a manifest with
no entry for the OS, sqlite3.Error for the WebGL database) and emits a
FallbackWarning. The text names what failed and what the identity uses
instead, then gives a block to paste into an issue (camoufox, browser, OS
and Python versions, the error, and the identity's user agent or GPU),
asking the user to report it on GitHub. It shares LeakWarning's
caller-frame attribution and its template lives in warnings.yml.

The broad excepts had also been hiding a broken fixture:
test_launch_environment's font and voice stubs did not accept `seed`, so
every draw there raised and was swallowed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): give NewContext identities the browser's Firefox version

NewContext() and AsyncNewContext() passed ff_version=None through to
generate_context_fingerprint(), so a context's user agent kept the version
fpgen drew (e.g. Firefox/146) while the browser underneath was 152. They
now default ff_version to the major version of Playwright's
Browser.version, which Juggler reports from MOZ_APP_VERSION_DISPLAY, so the
UA always names the browser the page is actually talking to. An explicit
ff_version still wins.

The docstrings said each context gets "its own real fingerprint preset";
the default has been an fpgen draw, with a preset only when one is passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): send an IPv6 WebRTC address to setWebRTCIPv6

The per-context init script passed every webrtc_ip, IPv6 included, to
window.setWebRTCIPv4(), and never called setWebRTCIPv6(). An IPv6 address
(given directly, or resolved as a proxy's exit IP) was stored as the
context's IPv4 value and the IPv6 slot stayed empty. The script now picks
the setter by address family, and an address that is neither raises
InvalidIP instead of being passed through.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): stop pinning the page's scroll offset from fpgen

fpgen.yml mapped the drawn window.pageYOffset (e.g. 528) to
screen.pageYOffset, and the browser returns that value from scrollY on
every read, so a page saw one scroll position forever whatever the user
did. Real scroll offsets are live page state, not part of a device's
fingerprint, so neither offset is mapped any more.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(python): stop checking a config key that no longer exists

warn_manual_config() looked for navigator.languages, which was removed
from settings/properties.json; validate_config() rejects it before the
check could matter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(python): close the WebGL database connection on every path

sample_webgl raised its not-found and wrong-OS errors before reaching
conn.close(), leaking a sqlite connection each time a preset named a GPU the
database does not hold.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(data): drop the 23 presets whose GPU has no WebGL data

A preset records only its GPU's name. The WebGL parameters, extensions and
shader precision behind it have to come from somewhere, and for these 23
nothing Camoufox has describes the GPU: fpgen has never seen Firefox report it
on that OS. So each launch paired the name with another device's parameters,
a mismatch any WebGL fingerprinter can see. They were:

- Windows on ARM (Adreno 650);
- Direct3D 10-level GPUs (vs_4_0/vs_4_1);
- "Generic Renderer";
- 945GM and GTX 480 on macOS;
- nouveau/Mesa buckets on Linux;
- one Linux preset pairing NVIDIA's proprietary vendor string with the
  nouveau renderer name.

scripts/clean-fingerprint-data.py now applies the rule, via a shared
fingerprints.firefox_gpus(), and test_shipped_data asserts it. 374 presets
remain, and every OS keeps its presets. ROADMAP.md lists capturing WebGL data
for these GPUs, which would bring them back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(python): draw every identity's WebGL from fpgen

WebGL vendor, renderer, context attributes, extensions, parameters and
shader precisions, for WebGL1 and WebGL2, now come from fpgen's recorded
Firefox devices instead of webgl_data.db, which is deleted with the
camoufox/webgl/ package.

camoufox/webgl.py:
- webgl_for_gpu() traces `webgl` given Firefox, the OS and the GPU, then
  `webgl2` given the chosen `webgl` too, and draws each with one seeded
  random.Random. The GPU and the webgl value are pinned by their fpgen
  lookup index: a dict condition is flattened into leaves that overwrite
  each other, so only the renderer applied and Linux "Mesa" and "AMD"
  Radeon HD 3200 devices came back mixed.
- sample_webgl_for_screen() draws the GPU of a generated identity from
  fpgen's per-OS weights, filtering out software rasterisers, GPUs the OS
  cannot report, discrete GPUs behind a netbook screen and the
  resistFingerprinting "Mozilla" mask before the weighted choice, so there
  is no rejection loop. An empty pool raises.
- The draft/host-dependent extension filter moves over unchanged.

A preset's GPU and a caller's webgl_config pair are looked up as given;
a pair fpgen has never seen from Firefox on that OS raises instead of
falling back to another GPU. generate_context_fingerprint no longer
falls back to the host GPU when the draw fails.

For 10 of the 15 (GPU, OS) pairs the two sources share, one of fpgen's
records converts to exactly the database row on every value the browser
reads. The other five rows (Linux R9 200 and Radeon HD 3200, macOS
Intel HD, and two software rasterisers) are devices fpgen does not carry;
those GPUs now present fpgen's recorded devices instead. The Linux
GTX 980 row is kept as a test fixture.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: say where WebGL comes from now that the database is gone

The per-context guide, the fpgen.yml header and coherence's comments still
named webgl_data.db and sample_webgl(). They now point at camoufox/webgl.py
and fpgen. The guide also claimed presets carry WebGL parameters; they
record only the vendor and renderer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(patches): a host's missing speech daemon no longer errors spoofed speech

On a Linux host where speech-dispatcher cannot start, Firefox broadcasts
synth-voices-error, and SpeechSynthesis answers it by firing `error` on every
queued utterance. So a spoofed Windows voice errored about 11ms into speak()
on any host without the daemon: the CI runners, and most servers. It passed
only where the daemon runs.

While Camoufox manages the voice list, the registry no longer forwards a host
backend's error. The spoofed voices do not depend on the host's engine, and a
Windows or macOS identity never raises one. The guard now makes the daemon
unreachable itself, so it tests this case on every machine; on the previous
build it fails every time.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: stop skipping the two click tests that stock animation timing fixed

test_wait_for_stable_position and test_timeout_waiting_for_stable_position
were skipped with humanized travel time as the reason. The real cause was
instant animations. Every finite animation finished at once, so the button
Playwright waits on to stop moving never moved, and the click landed where
upstream does not expect. With animations on stock timing both pass, and the
skiplist audit flagged them as no longer failing. The entries go, and the
counts in ci/README.md drop from 14 to 12.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(build-tester): accept 18 and 22 cores, as real hardware reports

plausibleHWC's list of common core counts lacked 18 and 22 -- Intel Meteor
Lake laptops (Core Ultra 5 125H, Core Ultra 7 155H), and 22 is in 8 recorded
presets. build-tester draws random presets, so a run that picked one of the two
Linux presets reporting 22 failed: about one run in eleven, on any pull
request. A CI self-test now fails if the list rejects any core count pythonlib
can present (the presets and PLAUSIBLE_CORE_COUNTS).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(guards): judge the query-cost probes on a median, not one sample

stock-parity-probes timed each getter once. On a shared runner one GC pause or
CPU-steal spike decided the verdict: navigator.hardwareConcurrency took 77 ms
against a 50 ms allowance on the same restored build that passed the run
before. Each pair is now timed five times, interleaved, and compared by median.
The regressions these catch (a sync IPC per read, ~240 ms over the loop) cost
extra on every read, so they move the median; verified by giving the getter a
constant ~4 us of extra work per read -- 86 ms median, FAIL -- while the
healthy build passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(native): compare the whole fingerprint when two launches must differ

test_two_browsers_get_different_fingerprints compared seven coarse values:
UA, platform, screen size, core count, timezone and language. CI pins the
timezone and language, and real machines share the rest: two draws of a
common Mac (Firefox 152, MacIntel, 2560x1440, 8 cores) matched, and the test
failed on a correct browser.

It now reads the whole fingerprint a site computes, from a script in the page:
- navigator values, screen and window geometry, device pixel ratio, timezone;
- WebGL vendor, renderer, limits and extensions;
- installed fonts, measured by width against the generic fallbacks;
- voices, media-device counts, and an OfflineAudioContext hash.
The page is served from an https URL Playwright fulfils locally, because
mediaDevices exists only in a secure context. The page computes the result
itself because the isolated world may not read audio sample data. The test
then requires the fingerprints to differ, and the audio hash to differ on its
own, since its noise is seeded per identity.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Update README to remove warning, camoufox is now actively maintained

Camoufox will now be actively maintained and improved for the foreseeable future

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 20:13:35 +00:00

14 KiB
Raw Blame History

Bundled fonts: what ships, what is reported, how to verify

Camoufox ships its own fonts so that font enumeration and text rendering are a property of the claimed OS rather than of the host. There are two halves: a font bundle (bundle/fonts/) and one fontconfig per OS (bundle/fontconfig/<os>/fonts.conf). The data the launcher draws from (pythonlib/camoufox/fonts.json, font-bases.json, font-groups.json and the _ESSENTIAL_FONTS_* / _BASE_VARIANT_FONTS_* / _*_MARKER_FONTS constants in fingerprints.py) is generated against that bundle and must be regenerated with it.

Getting the bundle

The bundle is a release asset, not repo content — 1513 faces, 2.16 GB extracted, 843 MB as .tar.xz. It cannot live in git: GitHub caps a tracked file at 100 MiB (Apple Color Emoji.ttc alone is 183 MiB) and an .xz archive cannot be delta-compressed, so every font update would append another ~843 MB blob to history forever. Release assets allow 2 GB per file, which fits the whole bundle with no splitting and no LFS. It is the same trust model the build already uses for the Firefox source itself.

make fetch-fonts     # download + verify the archive (bundle/fonts-bundle-*.tar.xz)
make fonts-extract   # ...and unpack it to bundle/fonts/   (gitignored)
make fonts-check     # verify the archive against the pin; non-zero if absent
make fonts-clean     # drop the extracted tree

scripts/data/font-bundle.json pins the tag, asset name, size and sha256, so a given commit expects one exact bundle and a truncated or substituted download fails loudly instead of silently producing a browser that reports fonts it cannot draw. make fonts-extract stamps the unpacked tree with that sha256 (bundle/fonts/.bundle-sha256), so re-running it is free — which is what lets package-* and stage-fonts depend on it unconditionally.

The asset is hosted by this repository, never by a fork: a build input that lives in someone's personal account breaks the moment that account renames the repo or deletes the release.

Publishing a new bundle: rebuild the archive, python3 scripts/fetch-fonts.py --write-spec <archive> --tag font-bundle-vN, then upload it under that tag. Font-bundle tags are excluded from build.yml, so they do not trigger a browser build or appear among the browser downloads.

Layout: each face stored once

A face used by several OSes is stored once, in a directory named for the set of OSes that use it. bundle/fonts/groups.json records which groups each OS reads.

group faces read by
LMW 328 all three
LM 92 Linux + macOS
LW 71 Linux + Windows
MW 48 macOS + Windows
L 301 Linux only
M 311 macOS only
W 362 Windows only

An OS reads the four groups its letter appears in: Linux L+LM+LMW+LW (792 faces), macOS LM+LMW+M+MW (779), Windows LMW+LW+MW+W (809). Storing a full copy per OS instead made 41% of the bundle byte-identical duplicates (2955 files / 3.96 GB became 1513 / 2.16 GB) while leaving every OS's set unchanged.

Every package ships all seven groups, so the fingerprint is built from the fonts actually available. What differs is how they are gated:

  • Linux packages keep the group subdirectories and groups.json. utils._generate_fontconfig reads it at launch and hands fontconfig one <dir> per group the claimed OS reads. The parent is never named, because fontconfig scans <dir> recursively — naming it would put every other OS's faces on the search path.
  • macOS and Windows packages flatten the groups into one directory: CoreText and DirectWrite activate a single directory and cannot read subfolders. There is no directory gate on those hosts; the allowlist in patches/font-hijacker.patch is what restricts a lookup. Basenames therefore must not collide within a package's group set, which verify-fonts.py checks.

This replaced the 455 <glob> duplicate-face rejects the Windows fontconfig used to carry. They were redundant under grouping and also unsound: 114 bundled filenames contain [ ] (e.g. ReemKufi[wght].ttf), which fontconfig parses as a character class, so those rejects silently matched nothing.

The two invariants

1. Everything reported must be renderable. Every family the launcher may report for an OS must be one the packaged browser can draw under that OS's bundled fontconfig. font-hijacker.patch prunes lookups by family name after fontconfig substitution, so a reported name fontconfig cannot resolve measures as the fallback — a reverse leak. The old fonts.json[mac] carried 70 such names (Hiragino etc.); the current lists carry none.

2. Nothing outside an OS's groups may be reachable. The converse, and the one the group layout exists to enforce: a Windows identity must not be able to render a macOS-only face, or an emoji/CJK glyph can fall back to Segoe UI Emoji or PingFang on a machine claiming Linux. No reported name reveals this, and no draw test can see it — it is checked directly, by asking fontconfig for every file it can reach under each OS's conf and requiring all of them to sit inside that OS's own groups (810 / 780 / 793 faces, including browser-shipped TwemojiMozilla.ttf).

The reverse direction of invariant 1 is deliberately loose: an OS renders more families than it ever reports (529 published vs 335 reported on Windows, 872 vs 567 on macOS, 583 vs 358 on Linux; gen-fonts-json.py --dump-union lists the gap). Those are names a real OS never presents as a family — weight-variant subfamilies DirectWrite folds into their parent ("Barlow Black"), the bare Sitka / Segoe UI Variable umbrella names, and CJK families a stock install does not report. They stay renderable but unreported.

How fonts.json is derived (scripts/gen-fonts-json.py)

reportable[os] = fc-scan(the groups <os> reads)          families fontconfig publishes
               + SCAN_FAMILIES                           Sitka * / Segoe UI Variable * (scan-time matches)
               + ALIASES                                 names fonts.conf rewrites to a bundled target
               + SHIPPED_BY_BROWSER                      Twemoji Mozilla (from the Firefox build)
               ∩ names(scripts/data/font-manifests.json) bases + additions + marker fonts for that OS
               + ALIASES, REPORTABLE_EXTRA               forced in

Namespace matters here and is easy to get wrong: Windows and macOS enumerate the legacy family (nameID 1), while Linux enumerates through fontconfig. system_profiler and WPF report typographic families (nameID 16), a different namespace — a list gathered that way will disagree with what a browser sees.

--print-bases prints the OS base lists intersected with the result, which is what the _ESSENTIAL_FONTS_* constants are pasted from. Those constants are the intersection of an OS's bases, never a superset: a name present in only one version's base must not be forced onto every identity, or that version's base becomes a no-op.

Regenerate together whenever the bundle changes:

python3 scripts/gen-font-groups.py          # font-groups.json + font-bases.json
python3 scripts/gen-fonts-json.py           # fonts.json
python3 scripts/gen-fonts-json.py --print-bases   # paste into fingerprints.py
python3 scripts/verify-fonts.py             # must pass before committing

Deliberate choices:

  • Windows aliases are unconditional and always reported. The 13 GDI-substitution / Light-Semilight rules (Courier -> Courier New, Helvetica -> Arial, Calibri Light -> Calibri, ...) are unconditional in bundle/fontconfig/windows/fonts.conf because there is no per-launch fontconfig hook, and the names are in _ESSENTIAL_FONTS_WINDOWS. That matches reality: every Windows install resolves all of them.
  • macOS TTC weight names are reported. The confs rewrite American Typewriter Semibold, Futura Bold, STIX Two Math Regular, ... to their base family. A real Mac registers those as families but fc-scan cannot see them, so the ones whose target is bundled are forced in.
  • Twemoji Mozilla is reportable on Linux (a CreepJS Linux marker, and Firefox exposes its bundled font to content). Not reported on macOS.
  • One name renders for every Linux identity but is not always reported (Noto Sans Canadian Aboriginal Regular) because a real Ubuntu does not report it. verify-fonts.py warns rather than fails.

The per-launch draw (fingerprints._generate_random_font_subset)

Modelled on how real machines actually differ: one OS-version base, always whole, plus independent per-unit draws. Not a uniform percentage of a pool.

  1. One base, drawn by weight, always complete. A real machine has all of its OS defaults, so a base is never sampled.

    OS bases (weight / families)
    Windows win11 1.00 / 125
    macOS sonoma 0.20 / 499 · tahoe26 0.65 / 364 · macos27 0.15 / 361
    Linux ubuntu2404 0.45 / 257 · ubuntu2604 0.55 / 266

    Windows 10 was dropped entirely. macOS 27 differs from 26.6.2 by only three families and is identified by the absence of Noto Sans Brahmi / Noto Sans CanAborig.

  2. The essential floor (_ESSENTIAL_FONTS_*: 125 / 369 / 274) — the intersection of that OS's bases, so it holds whichever base was drawn.

  3. Per-unit Bernoulli draws, per OS. Each addition is one independent decision at its own measured probability. kind: bundle is atomic — it installs whole or not at all. kind: alacarte is piecemeal: on a hit, a size is drawn from its sizes distribution and that many families are taken. requiresLocale gates a unit on the identity's locale (e.g. cjk-tc on zh-TW).

    unit kind Windows macOS Linux
    office bundle 0.602 — —
    apple-optional bundle — 0.818 —
    hololens bundle 0.592 — —
    distro-i18n bundle — — 0.372
    libreoffice bundle 0.066 — 0.112
    pan-euro bundle 0.028 — —
    cascadia bundle 0.100 0.060 0.080
    meslo bundle 0.040 0.040 0.060
    cjk-tc bundle (zh-TW) 0.900 — —
    dev à la carte 0.038 0.070 0.174
    web à la carte 0.088 0.270 0.195

    apple-optional is one atomic unit rather than base because Apple ships those 57 faces as optional Font Book downloads, and the fpgen corpus puts all of them on exactly 81.8% of Macs.

  4. Marker fonts appended if missing (_ensure_marker_fonts).

Resulting list sizes over 20 draws: 125–281 of 335 (Windows), 369–526 of 567 (macOS), 278–319 of 358 (Linux).

For a native identity — the host's own OS — on macOS and Windows the draw claims the OS base only, because the real system fonts are the right ones there and font-hijacker.patch does not activate the bundle. On a Windows host the Win11 marker families are subtracted when the host cannot render them, rather than added when it can; claiming a marker the host lacks is the leak.

pythonlib/tests/test_font_distribution.py covers this: base completeness and weights, per-unit probability, bundle atomicity, à-la-carte sizing, locale gating, determinism, renderable-only, and that the draw actually varies (distinct lists, no single list dominating). Tolerances are binomial, at 5σ, so they fail on a real shift rather than on noise.

fontconfig parity

Linux: generics resolve as a stock Ubuntu does (sans-serif → Noto Sans, serif → Noto Serif, monospace → DejaVu Sans Mono, cursive → Z003), hintstyle is hintslight, and the stock metric aliases (30-metric-aliases), 49-sansserif, urw-base35 and the non-Latin rule files are merged in stock conf.d order. Windows: sans-serif → Arial, serif → Times New Roman, monospace → Consolas, cursive → Comic Sans MS, plus MS Shell Dlg 2 -> Tahoma, the Sitka / Segoe UI Variable scan-time instance families and the alias block above. macOS: Helvetica / Times / Menlo / Apple Chancery / Zapfino. system-ui is left to Gecko, which resolves it itself. The <dir prefix="cwd">fonts</dir> line that utils.py rewrites is kept in all three — verify-fonts.py fails if it is missing.

Verifying

python3 scripts/verify-fonts.py          # full: builds a fontconfig cache over the bundle, minutes on first run
python3 scripts/verify-fonts.py --quick  # constants, draws and layout only
cd pythonlib && python3 -m pytest -q -k "font or humanize or launch_environment"

The full run checks, per OS: fonts.json[os] ⊆ what fc-list publishes (aliases resolved via fc-match), no file reachable outside that OS's groups, generics resolve to a reportable family, essential / variant / marker ⊆ fonts.json, 20 draws stay inside fonts.json, each face stored exactly once, no subfolders, and no basename collisions within a package's group set.

Known residue

  • The Linux base over-claims against the recorded corpus. A draw reports a median ~287 families where real Linux machines in the corpus report ~100. The cause is that the bundle ships fontconfig's 30-metric-aliases, so Arial/Times/Courier always resolve and must therefore always be reported. Closing it needs either distro base variants or a change to which aliases ship.
  • The macOS base weights (0.20 / 0.65 / 0.15) are estimates. There is no version-share data source behind them, unlike the addition probabilities, which are measured. Sonoma's base is inherited and not re-verifiable on current hardware.
  • Nothing has been checked in a running browser. Both invariants above are verified at the fontconfig layer, which is the gate on Linux packages. On macOS and Windows hosts the gate is the allowlist patch instead, and that has not been exercised against the flattened group layout.