mirror of
https://github.com/daijro/camoufox.git
synced 2026-10-03 16:00:19 +00:00
* chore: remove build tooling nothing uses - The developer UI (scripts/developer.py, `make edits`). It depended on easygui, which no requirements file declares, and every action it offered is a Makefile target: patch, unpatch, workspace, revert, diff. Its two helpers in scripts/_mixin.py (is_bootstrap_patch, patch) had no other callers. - legacy/, the Go launcher deprecated in 2024-11. Nothing built or shipped it. Its Makefile targets and scripts/run-pw.py go with it, and so does Go from every dependency list and workflow. - jsonvv/ and settings/camoucfg.jvv. Nothing read the .jvv schema: config is validated against settings/properties.json, and the two had already drifted. The jsonvv package stays on PyPI. - Scripts with no caller: bootstrap.py, moztree, setup-wasi-linux.sh, package-helper.sh, install-local-build.sh, mozfetch.sh (copied into lw/ but never packaged), examples/. - The pre-ESM Juggler copies JugglerFrameParent.jsm and JugglerFrameChild.jsm, and hidden-scrollbars.css. Juggler loads the .sys.mjs actors and deliberately no stylesheet, but jar.mn still packaged all three. - patches/librewolf/*.opt, which list_patches() never picks up; the roverfox second pass in patch.py, whose directory no longer exists; the unread --no-settings-pane option. - The CAMOUFOX_PASSWD secret passed to `make fetch` and closedsrc_rev in upstream.sh, which nothing reads. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(python): remove dead helpers and a stale dependency None of these had a caller: - pkgman: is_supported_path, extract_zip, cleanup and set_version, left over from the single-directory install. cleanup() would have deleted every installed browser version. - multiversion.get_cached_repo_names, CONSTRAINTS.as_range, fingerprints._load_os_voices, utils._clean_locals, and unused imports. Also: - The "Apify Fingerprints" row in `camoufox version`, which has read "?" since fpgen replaced BrowserForge. - lxml is no longer a dependency; nothing imports it. - The geoip extra now names maxminddb, the module geolocation.py actually imports, rather than getting it transitively through geoip2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: show the cursor paths humanize=True actually produces The README's cursor video showed the Bezier generator Camoufox replaced with Cursory's recorded trajectories. scripts/cursor-demo.py drives a real build with humanize=True and records every mousemove event the page receives. It writes assets/humanize-cursor.svg, an animated replay at the recorded speed, so what the figure shows is what a site sees. The script cannot change the binary, so ci/browser_inputs.py lists it as non-native and editing it does not invalidate the cached browser. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(python): stop naming BrowserForge in user-facing text fpgen replaced BrowserForge, but two LeakWarnings, the NonFirefoxFingerprint message and the fingerprint_preset docstring still named it. One warning also linked to a README anchor that no longer exists. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: one AGENTS.md for every agent, a roadmap, and docs that match the code - AGENTS.md holds the engineering rules for any coding agent, plus the repo map, build, patch and test commands that CLAUDE.md used to carry. CLAUDE.md now only imports it, so there is one set of rules. ci/tribal-rules.yml is the record of settled decisions it points to. - ROADMAP.md lists planned work, each item linked to its issue. - README: - fpgen and the coherence check replace BrowserForge; - the patch workflow uses the make targets instead of the removed developer UI; - letter-spacing noise is described as off by default, as it is. - docs/: - beta-testing-ff146.md removed; - patch-upgrading-guide rewritten around the make targets; - per-context-patches without the canvas patch that no longer exists, and with measured preset counts; - playwright-maintenance without the JSM wrapper that does not exist; - smaller fixes in MEDIA-DEVICES, input-dispatch and FONTS. - ci/README: every job, and the real shard, skiplist and entry-point lists. - pythonlib, tester and patch-dependency READMEs corrected against the code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(pythonlib): handle headless='virtual' in launch_server launch_server() is documented to take the same arguments as Camoufox(), but passed headless='virtual' straight to launch_options(), so the server launched with no Xvfb display. Start a VirtualDisplay the way Camoufox() does, launch headful on it, and kill it when the server process exits or the launch fails. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(python): remove fontprobe, which nothing called fontprobe listed the fonts installed on the host, for a `camoufox fonts` command that was never added. It has nothing to do with the font bundle Camoufox serves to pages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(license): the Python launcher is MIT; the browser stays MPL-2.0 The Python package has always been published to PyPI as MIT (#727), but pythonlib/ shipped no licence file, and the repo's LICENSE is the browser's MPL-2.0. MPL is copyleft per file. It covers the modified Firefox sources, not a separate launcher that drives the browser over Playwright. So pythonlib/LICENSE now carries the MIT text its metadata already declares, and a Licensing section in the README says which part is which. Closes #727. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): fingerprint_preset=False no longer turns presets on launch_options checked `fingerprint_preset is not None`, so passing False drew a random bundled preset, the opposite of what was asked. It now uses a truthiness check, and a test proves that None and False never draw a preset. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: bring the release workflow in line with the tests build job build.yml had drifted from tests.yml. It ran actions at v1/v2 on a retired Node runtime, prepared the source tree with bare make calls that fail the whole release on one dropped connection, and built with a different Python than every pull request is tested with. - Pin every action by commit SHA, at the major versions tests.yml uses (checkout v4, setup-python v5, upload/download-artifact v4, the same remove-unwanted-software SHA), and action-gh-release v2. The release job holds contents: write, so it should not follow a movable tag. - Prepare the tree with `python3 -m ci.run_prepare`, as the tests build job does. BUILD_TARGET is set from the matrix so `make dir` writes the right mozconfig and Rust targets; multibuild.py then finds _READY and builds without re-patching. mach's toolchain bootstrap ignores the mozconfig, so running it after `dir` bootstraps the same toolchains. - Build with Python 3.12, the version the tests build job compiles with. - Default the workflow to no permissions; the build job gets contents: read and the release job keeps contents: write. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): NewContext looks up a proxy's exit IP through the right URL, or fails NewContext derives the context's WebRTC IP and timezone from the proxy's exit IP. That lookup had two defects, and both left the context showing the host's values while its traffic went through the proxy: - It built its own proxy URL with urlparse, which reads a scheme-less server such as "1.2.3.4:8080" (a form Playwright accepts) as scheme "1.2.3.4" with no host. urllib could not use a SOCKS proxy at all. - Any failure was swallowed, and the context opened without the values. The URL is now built with Proxy.as_string(), which the geoip launch path already uses (scheme-less means http). The lookup goes through requests, which handles SOCKS, and a failed lookup raises InvalidIP, naming the two options that skip it. The tests cover scheme-less, http and socks5 servers with credentials, both failure modes, and the case where no lookup is needed, for NewContext and AsyncNewContext. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: stop generating a canvas seed, and drop config keys nothing reads The browser has not noised the canvas since #528, and no patch reads canvas:seed (#721). The launcher still drew one on every launch and sent it through CAMOU_CONFIG, and NewContext called a setCanvasSeed that does not exist. They no longer do. For users this changes nothing on any browser since #528: the value was ignored. A config that still passes canvas:seed gets the usual "Skipping unknown patch" notice instead of silence. On a browser from before #528, the launcher no longer turns canvas noise on, which is the behaviour #528 chose. The same audit found more keys declared in settings/properties.json that no patch or Juggler file reads, so setting them did nothing: - canvas:aaOffset, canvas:aaCapOffset - memorysaver, pdfViewerEnabled, webrtc:localipv4/6 - navigator.onLine, navigator.cookieEnabled, navigator.languages - navigator.appCodeName, appName, product, productSub. Firefox reports these constants itself, so fpgen.yml no longer maps them. - webGl:parameters:blockIfNotDefined and its WebGL2 twin test_config_schema now checks this direction too: every declared key must be read by the browser, unless it is listed with a reason. Three are listed: locale:script and navigator.doNotTrack, which the launcher applies itself, and navigator.buildID (#780). The build-tester grading followed the same wrong premise. It tracked canvas collisions as an unfixed per-context leak. A canvas that is rendered rather than noised follows the fonts and GPU, as it does on real machines, so canvas collisions are now counted with the other device-level values. The tribal rule that recorded it as an open question is now a settled one, canvas-is-not-noised, with an automated check. Closes #721. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): repair devicePixelRatio the same way on every launch The DPR repair snaps an off-grid ratio to the nearest real scaling step and keeps the first of two equally near steps. The steps were frozenset literals, and a frozenset literal iterates in one order when the module is compiled from source and another when it is loaded back from a .pyc. So a midpoint such as 1.125 became 1.25 on the first launch after an install and 1 on every launch after it: the same pinned identity presented two different devicePixelRatio values. The steps are now ascending tuples, so a tie always goes to the lower step. The test runs the repair in two fresh interpreters that share a bytecode cache, compiling in the first and loading in the second. It failed before this change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: remove the glyph-spacing seed from the browser and the launcher anti-font-fingerprinting.patch added a seeded amount to every glyph advance, so that text widths differed per context. No real machine produces those widths: the same font on the same OS measures the same everywhere. So the noise was itself a fingerprint, measured in #779 at +1 px per ~100 glyphs plus fractional deltas on every measureText. #779 defaulted the seed to 0 and kept it as an opt-in, but an opt-in whose only effect is to become detectable is not worth carrying. Removed: - The browser side: - FontSpacingSeedManager and window.setFontSpacingSeed; - the HarfBuzz hook; - the plumbing that existed only to carry the context id down to the shaper: the userContextId on gfxTextRun, gfxShapedWord and the word-cache key, and the extra MakeTextRun argument in nsTextFrame, nsFontMetrics, MathML and canvas. The font group keeps its userContextId, which font-list-spoofing.patch uses to apply the per-context font list. Text is now shaped exactly as stock Firefox shapes it. - The fonts:spacing_seed key. The launcher had been sending 0 on every launch, plus a setFontSpacingSeed(0) call in every context's init script. - tests/patches/config-overrides.py, which tested only the spacing override. A pythonlib test now covers config_overrides with another key. timezone-spoofing, webrtc-ip-spoofing and window-setter-seal change only in context lines and the setter seal list. Every patch applies cleanly to a fresh tree, and the result builds. The settled decision is recorded as no-glyph-spacing-noise, with an automated check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: stock animations and speech by default; drop config keys that freeze live values Three behaviours a page could detect, changed in one breaking release: - **Animations run on stock timing.** no-css-animations.patch finished every finite animation at once by default, and any page could read it: `el.animate(frames, 1000).effect.getComputedTiming().duration` was 0, and a 500ms transition reported 0. Measured on v152.0.4-beta.31. The speedup is now an opt-in, `instantAnimations: True`, which raises a LeakWarning. disableInstantAnimations is gone. - **speak() on a spoofed voice works like a real voice.** It fired `error` after 3ms unless voices:fakeCompletion was set, and then start and end in the same tick. It now starts and ends after the text's duration at ~150 words per minute. Both voices:fakeCompletion keys are gone, and so is a debug line printed to stderr on every call. - **Keys removed:** - battery:* and window.scrollMinX/Y: Firefox keeps getBattery() and scrollMin* chrome-only, so no page could read them. - window.scrollMaxX/Y, screen.pageXOffset/pageYOffset, window.history.length and document.body.client*: each pinned a live value to a constant, so scrolling, navigating or re-laying out never changed it. fpgen.yml mapped pageYOffset, so about 15% of identities froze window.scrollY at a non-zero value. - The body keys' role as an undocumented alias for window.innerWidth/Height in browser-init and in the launcher. - MaskConfig::GetInt32Rect, which only the body keys used. New guards, both of which fail on v152.0.4-beta.31: tests/patches/animation-timing.py and tests/patches/spoofed-voice-speaks.py. The decisions are recorded as animations-run-on-stock-timing and spoofed-voices-speak. Every patch applies cleanly to a fresh tree, and the result builds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python)!: remove dead public API and make `list all --path` work Breaking changes: - Remove the exceptions UnknownProperty, InvalidDebugPort and MissingDebugPort. Nothing in the package raises them, so code catching them was catching nothing. - Remove the legacy `allow_webgl` keyword of launch_options(). Use `block_webgl=True`. The keyword now reaches Playwright as an unknown launch option and fails there instead of being silently consumed. `camoufox list all --path` accepted the flag and ignored it. It now prints the install path beside each installed build, as `camoufox list --path` already does for the installed tree. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(python)!: drop data the package never draws from voices.json shipped in the wheel, but no code in the package reads it: the voice draw uses voice-manifests.json and voice-uris.json. Its only readers are the TypeScript port's golden-fixture generator and data-sync script (typescript/scripts/golden/identity_golden.py, typescript/scripts/sync-identity-data.py), which live on another branch and will need a new source; the last copy is at 676fb3f:pythonlib/camoufox/voices.json. docs/per-context-patches.md described it as runtime data and now describes the files that are. webgl_data.db held two rows with zero weight on every OS ("Intel(R) HD Graphics 400, or similar" from "Intel Inc." and "Radeon R9 200 Series, or similar" from "ATI Technologies Inc."), left behind when their impossible macOS weights were zeroed. No draw can reach them. They are deleted with secure_delete so their blobs do not linger in free pages; the file is not vacuumed, so the other pages are unchanged. A new test requires every row to be drawable on at least one OS. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): warn whenever an identity falls back to a substitute value Several draws swallowed their failure and used something else, so an identity could ship with values the rest of it was not drawn to match and nobody would hear about it: - from_preset(): a failed font or voice draw used the preset's recorded list, or nothing, on any exception. - generate_context_fingerprint(): a failed font, voice or WebGL draw was `except Exception: pass`, leaving the browser's launch-time values. - _load_font_groups() / _load_font_bases(): an unreadable file became {}, i.e. no font additions or no OS-version base. - launch_options(): a failed font draw used every font in fonts.json, a failed voice draw used no voices, and a preset GPU missing from webgl_data.db was silently swapped for a drawn one (36 of the 397 bundled presets). Each site now catches only the errors its data can raise (OSError and ValueError for an unreadable or corrupt file, KeyError for a manifest with no entry for the OS, sqlite3.Error for the WebGL database) and emits a FallbackWarning. The text names what failed and what the identity uses instead, then gives a block to paste into an issue (camoufox, browser, OS and Python versions, the error, and the identity's user agent or GPU), asking the user to report it on GitHub. It shares LeakWarning's caller-frame attribution and its template lives in warnings.yml. The broad excepts had also been hiding a broken fixture: test_launch_environment's font and voice stubs did not accept `seed`, so every draw there raised and was swallowed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): give NewContext identities the browser's Firefox version NewContext() and AsyncNewContext() passed ff_version=None through to generate_context_fingerprint(), so a context's user agent kept the version fpgen drew (e.g. Firefox/146) while the browser underneath was 152. They now default ff_version to the major version of Playwright's Browser.version, which Juggler reports from MOZ_APP_VERSION_DISPLAY, so the UA always names the browser the page is actually talking to. An explicit ff_version still wins. The docstrings said each context gets "its own real fingerprint preset"; the default has been an fpgen draw, with a preset only when one is passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): send an IPv6 WebRTC address to setWebRTCIPv6 The per-context init script passed every webrtc_ip, IPv6 included, to window.setWebRTCIPv4(), and never called setWebRTCIPv6(). An IPv6 address (given directly, or resolved as a proxy's exit IP) was stored as the context's IPv4 value and the IPv6 slot stayed empty. The script now picks the setter by address family, and an address that is neither raises InvalidIP instead of being passed through. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): stop pinning the page's scroll offset from fpgen fpgen.yml mapped the drawn window.pageYOffset (e.g. 528) to screen.pageYOffset, and the browser returns that value from scrollY on every read, so a page saw one scroll position forever whatever the user did. Real scroll offsets are live page state, not part of a device's fingerprint, so neither offset is mapped any more. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(python): stop checking a config key that no longer exists warn_manual_config() looked for navigator.languages, which was removed from settings/properties.json; validate_config() rejects it before the check could matter. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(python): close the WebGL database connection on every path sample_webgl raised its not-found and wrong-OS errors before reaching conn.close(), leaking a sqlite connection each time a preset named a GPU the database does not hold. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(data): drop the 23 presets whose GPU has no WebGL data A preset records only its GPU's name. The WebGL parameters, extensions and shader precision behind it have to come from somewhere, and for these 23 nothing Camoufox has describes the GPU: fpgen has never seen Firefox report it on that OS. So each launch paired the name with another device's parameters, a mismatch any WebGL fingerprinter can see. They were: - Windows on ARM (Adreno 650); - Direct3D 10-level GPUs (vs_4_0/vs_4_1); - "Generic Renderer"; - 945GM and GTX 480 on macOS; - nouveau/Mesa buckets on Linux; - one Linux preset pairing NVIDIA's proprietary vendor string with the nouveau renderer name. scripts/clean-fingerprint-data.py now applies the rule, via a shared fingerprints.firefox_gpus(), and test_shipped_data asserts it. 374 presets remain, and every OS keeps its presets. ROADMAP.md lists capturing WebGL data for these GPUs, which would bring them back. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(python): draw every identity's WebGL from fpgen WebGL vendor, renderer, context attributes, extensions, parameters and shader precisions, for WebGL1 and WebGL2, now come from fpgen's recorded Firefox devices instead of webgl_data.db, which is deleted with the camoufox/webgl/ package. camoufox/webgl.py: - webgl_for_gpu() traces `webgl` given Firefox, the OS and the GPU, then `webgl2` given the chosen `webgl` too, and draws each with one seeded random.Random. The GPU and the webgl value are pinned by their fpgen lookup index: a dict condition is flattened into leaves that overwrite each other, so only the renderer applied and Linux "Mesa" and "AMD" Radeon HD 3200 devices came back mixed. - sample_webgl_for_screen() draws the GPU of a generated identity from fpgen's per-OS weights, filtering out software rasterisers, GPUs the OS cannot report, discrete GPUs behind a netbook screen and the resistFingerprinting "Mozilla" mask before the weighted choice, so there is no rejection loop. An empty pool raises. - The draft/host-dependent extension filter moves over unchanged. A preset's GPU and a caller's webgl_config pair are looked up as given; a pair fpgen has never seen from Firefox on that OS raises instead of falling back to another GPU. generate_context_fingerprint no longer falls back to the host GPU when the draw fails. For 10 of the 15 (GPU, OS) pairs the two sources share, one of fpgen's records converts to exactly the database row on every value the browser reads. The other five rows (Linux R9 200 and Radeon HD 3200, macOS Intel HD, and two software rasterisers) are devices fpgen does not carry; those GPUs now present fpgen's recorded devices instead. The Linux GTX 980 row is kept as a test fixture. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: say where WebGL comes from now that the database is gone The per-context guide, the fpgen.yml header and coherence's comments still named webgl_data.db and sample_webgl(). They now point at camoufox/webgl.py and fpgen. The guide also claimed presets carry WebGL parameters; they record only the vendor and renderer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(patches): a host's missing speech daemon no longer errors spoofed speech On a Linux host where speech-dispatcher cannot start, Firefox broadcasts synth-voices-error, and SpeechSynthesis answers it by firing `error` on every queued utterance. So a spoofed Windows voice errored about 11ms into speak() on any host without the daemon: the CI runners, and most servers. It passed only where the daemon runs. While Camoufox manages the voice list, the registry no longer forwards a host backend's error. The spoofed voices do not depend on the host's engine, and a Windows or macOS identity never raises one. The guard now makes the daemon unreachable itself, so it tests this case on every machine; on the previous build it fails every time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: stop skipping the two click tests that stock animation timing fixed test_wait_for_stable_position and test_timeout_waiting_for_stable_position were skipped with humanized travel time as the reason. The real cause was instant animations. Every finite animation finished at once, so the button Playwright waits on to stop moving never moved, and the click landed where upstream does not expect. With animations on stock timing both pass, and the skiplist audit flagged them as no longer failing. The entries go, and the counts in ci/README.md drop from 14 to 12. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(build-tester): accept 18 and 22 cores, as real hardware reports plausibleHWC's list of common core counts lacked 18 and 22 -- Intel Meteor Lake laptops (Core Ultra 5 125H, Core Ultra 7 155H), and 22 is in 8 recorded presets. build-tester draws random presets, so a run that picked one of the two Linux presets reporting 22 failed: about one run in eleven, on any pull request. A CI self-test now fails if the list rejects any core count pythonlib can present (the presets and PLAUSIBLE_CORE_COUNTS). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(guards): judge the query-cost probes on a median, not one sample stock-parity-probes timed each getter once. On a shared runner one GC pause or CPU-steal spike decided the verdict: navigator.hardwareConcurrency took 77 ms against a 50 ms allowance on the same restored build that passed the run before. Each pair is now timed five times, interleaved, and compared by median. The regressions these catch (a sync IPC per read, ~240 ms over the loop) cost extra on every read, so they move the median; verified by giving the getter a constant ~4 us of extra work per read -- 86 ms median, FAIL -- while the healthy build passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(native): compare the whole fingerprint when two launches must differ test_two_browsers_get_different_fingerprints compared seven coarse values: UA, platform, screen size, core count, timezone and language. CI pins the timezone and language, and real machines share the rest: two draws of a common Mac (Firefox 152, MacIntel, 2560x1440, 8 cores) matched, and the test failed on a correct browser. It now reads the whole fingerprint a site computes, from a script in the page: - navigator values, screen and window geometry, device pixel ratio, timezone; - WebGL vendor, renderer, limits and extensions; - installed fonts, measured by width against the generic fallbacks; - voices, media-device counts, and an OfflineAudioContext hash. The page is served from an https URL Playwright fulfils locally, because mediaDevices exists only in a secure context. The page computes the result itself because the isolated world may not read audio sample data. The test then requires the fingerprints to differ, and the audio hash to differ on its own, since its noise is seeded per identity. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Update README to remove warning, camoufox is now actively maintained Camoufox will now be actively maintained and improved for the foreseeable future --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
406 lines
18 KiB
Python
406 lines
18 KiB
Python
"""What a context is, versus what a browser launch is.
|
|
|
|
Camoufox offers two ways to get an identity, and conflating them is easy:
|
|
|
|
a **browser** launch carries one fingerprint for its whole process, resolved
|
|
before launch and passed through CAMOU_CONFIG;
|
|
|
|
a **context** carries the browser's fingerprint *unless* one is injected into
|
|
it -- `AsyncNewContext()`, or `generate_context_fingerprint()` plus
|
|
`new_context(**context_options)` and `add_init_script(init_script)`, which is
|
|
what build-tester and tests/patches/helpers.py do.
|
|
|
|
Both halves are contracts worth pinning, and they pull in opposite directions:
|
|
|
|
* a plain `new_context()` MUST inherit. Two tabs of the same machine that
|
|
disagreed would be a leak, not a feature.
|
|
* an injected context MUST get its own, and MUST NOT leak into its siblings.
|
|
That is what the per-context patches exist for, and it is the half that
|
|
degrades silently: a value that is really process-global looks correct in
|
|
any single-context test and only shows up when a second context opens. It
|
|
has happened here -- commit d17c887, "fix screen size leak in contexts".
|
|
|
|
The first draft of this file asserted that two *plain* contexts get different
|
|
fingerprints. They do not, by design, and the test failed against a real binary
|
|
the first time it ran. Pinning the contract Camoufox actually offers, rather
|
|
than the one that sounded right, is the whole point.
|
|
|
|
These assert isolation and consistency, never specific values: presets are drawn
|
|
at random, so pinning a number would make the suite a liability.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import asyncio
|
|
import json
|
|
from pathlib import Path
|
|
|
|
import pytest
|
|
|
|
pytestmark = pytest.mark.asyncio
|
|
|
|
# Read-only probes. `page.evaluate` in Camoufox reads values; it does not run
|
|
# script in the page's world, which is the whole point of the fork.
|
|
PROBES = {
|
|
"userAgent": "navigator.userAgent",
|
|
"platform": "navigator.platform",
|
|
"screenWidth": "screen.width",
|
|
"screenHeight": "screen.height",
|
|
"hardwareConcurrency": "navigator.hardwareConcurrency",
|
|
"timezone": "Intl.DateTimeFormat().resolvedOptions().timeZone",
|
|
"language": "navigator.language",
|
|
}
|
|
|
|
|
|
async def probe(page) -> dict:
|
|
out = {}
|
|
for name, expression in PROBES.items():
|
|
try:
|
|
out[name] = await page.evaluate(expression)
|
|
except Exception as exc: # noqa: BLE001
|
|
out[name] = f"<error: {exc}>"
|
|
return out
|
|
|
|
|
|
async def open_page(browser, context=None):
|
|
"""A context (inheriting, unless one is given) with one page on about:blank."""
|
|
context = context or await browser.new_context()
|
|
page = await context.new_page()
|
|
await page.goto("about:blank")
|
|
return context, page
|
|
|
|
|
|
# About half of get_random_preset()'s draws produce a WebGL vendor/renderer
|
|
# combination that is not in the sample data, and AsyncNewContext raises. That is
|
|
# expected -- tests/patches/helpers.py retries the same way -- so retry rather
|
|
# than letting a draw decide whether the suite passes.
|
|
MAX_PRESET_ATTEMPTS = 15
|
|
|
|
|
|
async def injected_context(browser, **kwargs):
|
|
"""A context with its own fingerprint, retrying past unusable draws."""
|
|
from camoufox.async_api import AsyncNewContext
|
|
|
|
last = None
|
|
for _ in range(MAX_PRESET_ATTEMPTS):
|
|
try:
|
|
return await AsyncNewContext(browser, **kwargs)
|
|
except ValueError as exc:
|
|
if "WebGL" not in str(exc):
|
|
raise
|
|
last = exc
|
|
raise AssertionError(f"no usable preset in {MAX_PRESET_ATTEMPTS} draws: {last}")
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
async def test_plain_contexts_inherit_the_browser_fingerprint(binary):
|
|
"""A context is not automatically a new machine.
|
|
|
|
Without an injected fingerprint a context belongs to the browser's identity,
|
|
and two of them must agree. Two tabs of one machine that disagreed would be
|
|
the leak, not the feature.
|
|
"""
|
|
from camoufox.async_api import AsyncCamoufox
|
|
|
|
async with AsyncCamoufox(executable_path=str(binary), headless=True,
|
|
i_know_what_im_doing=True) as browser:
|
|
ctx_a, page_a = await open_page(browser)
|
|
ctx_b, page_b = await open_page(browser)
|
|
a, b = await probe(page_a), await probe(page_b)
|
|
await ctx_a.close()
|
|
await ctx_b.close()
|
|
|
|
differing = {k: (a.get(k), b.get(k)) for k in PROBES if a.get(k) != b.get(k)}
|
|
assert not differing, (
|
|
"two plain contexts in one browser reported different fingerprints: "
|
|
f"{differing}. A context inherits the browser's identity unless one is "
|
|
"injected into it."
|
|
)
|
|
|
|
|
|
async def test_injected_contexts_each_get_their_own_fingerprint(binary):
|
|
"""The core promise of per-context spoofing.
|
|
|
|
If these come back identical, per-context injection has degraded to
|
|
process-global -- which passes every single-context test there is.
|
|
"""
|
|
from camoufox.async_api import AsyncCamoufox
|
|
|
|
async with AsyncCamoufox(executable_path=str(binary), headless=True,
|
|
i_know_what_im_doing=True) as browser:
|
|
ctx_a = await injected_context(browser, os="macos")
|
|
ctx_b = await injected_context(browser, os="linux")
|
|
_, page_a = await open_page(browser, ctx_a)
|
|
_, page_b = await open_page(browser, ctx_b)
|
|
a, b = await probe(page_a), await probe(page_b)
|
|
await ctx_a.close()
|
|
await ctx_b.close()
|
|
|
|
differing = [k for k in PROBES if a.get(k) != b.get(k)]
|
|
assert differing, (
|
|
"two injected contexts reported an identical fingerprint on every probe.\n"
|
|
f" {a}\n"
|
|
"Per-context injection has degraded to process-global (cf. commit d17c887)."
|
|
)
|
|
|
|
|
|
async def test_an_injected_context_does_not_leak_into_a_plain_one(binary):
|
|
"""The d17c887 shape: injecting into one context must not move another.
|
|
|
|
A per-context value implemented as a process-global would change the
|
|
browser's identity for everyone the moment the first context set it.
|
|
"""
|
|
from camoufox.async_api import AsyncCamoufox
|
|
|
|
async with AsyncCamoufox(executable_path=str(binary), headless=True,
|
|
i_know_what_im_doing=True) as browser:
|
|
plain_ctx, plain_page = await open_page(browser)
|
|
before = await probe(plain_page)
|
|
|
|
injected = await injected_context(browser, os="macos")
|
|
_, injected_page = await open_page(browser, injected)
|
|
await probe(injected_page)
|
|
|
|
after = await probe(plain_page)
|
|
await injected.close()
|
|
await plain_ctx.close()
|
|
|
|
moved = {k: (before.get(k), after.get(k)) for k in PROBES if before.get(k) != after.get(k)}
|
|
assert not moved, (
|
|
f"opening an injected context changed an existing context's fingerprint: {moved}. "
|
|
"That is a per-context value implemented as a process-global."
|
|
)
|
|
|
|
|
|
async def test_a_context_is_internally_coherent(binary):
|
|
"""A fingerprint has to agree with itself.
|
|
|
|
Cross-signal inconsistency -- a macOS platform with a Linux user agent -- is
|
|
more detectable than any single wrong value, because it cannot happen on a
|
|
real machine.
|
|
"""
|
|
from camoufox.async_api import AsyncCamoufox
|
|
|
|
async with AsyncCamoufox(executable_path=str(binary), headless=True, os="macos",
|
|
i_know_what_im_doing=True) as browser:
|
|
context, page = await open_page(browser)
|
|
values = await probe(page)
|
|
await context.close()
|
|
|
|
ua = str(values.get("userAgent", ""))
|
|
platform = str(values.get("platform", ""))
|
|
assert "Firefox" in ua, f"user agent does not claim Firefox: {ua!r}"
|
|
if platform.startswith("Mac"):
|
|
assert "Macintosh" in ua, f"platform {platform!r} disagrees with user agent {ua!r}"
|
|
assert int(values.get("screenWidth") or 0) > 1, values
|
|
assert int(values.get("screenHeight") or 0) > 1, (
|
|
f"screen height is {values.get('screenHeight')!r}. A 1x1 virtual display root "
|
|
"must never clamp the generated screen -- see the `not virtual_display` guards "
|
|
"in pythonlib/camoufox/utils.py."
|
|
)
|
|
|
|
|
|
async def test_closing_one_context_does_not_disturb_another(binary):
|
|
from camoufox.async_api import AsyncCamoufox
|
|
|
|
async with AsyncCamoufox(executable_path=str(binary), headless=True,
|
|
i_know_what_im_doing=True) as browser:
|
|
ctx_a, page_a = await open_page(browser)
|
|
ctx_b, page_b = await open_page(browser)
|
|
before = await probe(page_b)
|
|
await ctx_a.close()
|
|
after = await probe(page_b)
|
|
await ctx_b.close()
|
|
|
|
assert before == after, (
|
|
"closing one context changed another context's fingerprint:\n"
|
|
f" before {before}\n after {after}"
|
|
)
|
|
|
|
|
|
# The whole fingerprint a site computes, run by the page itself: Playwright's
|
|
# evaluate runs in Camoufox's isolated world, which may not read audio sample
|
|
# data, so the page writes its result into the DOM for the test to read.
|
|
FONT_CANDIDATES = [
|
|
"Arial", "Arial Black", "Calibri", "Cambria", "Candara", "Consolas", "Constantia",
|
|
"Corbel", "Courier New", "Ebrima", "Franklin Gothic Medium", "Gabriola", "Gadugi",
|
|
"Georgia", "Impact", "Ink Free", "Javanese Text", "Leelawadee UI", "Lucida Console",
|
|
"Lucida Sans Unicode", "Malgun Gothic", "Marlett", "Microsoft Himalaya",
|
|
"Microsoft JhengHei", "Microsoft New Tai Lue", "Microsoft PhagsPa", "Microsoft Sans Serif",
|
|
"Microsoft Tai Le", "Microsoft YaHei", "Microsoft Yi Baiti", "MingLiU-ExtB",
|
|
"Mongolian Baiti", "MS Gothic", "MV Boli", "Myanmar Text", "Nirmala UI",
|
|
"Palatino Linotype", "Segoe Print", "Segoe Script", "Segoe UI", "Segoe UI Emoji",
|
|
"SimSun", "Sitka Small", "Sylfaen", "Symbol", "Tahoma", "Times New Roman",
|
|
"Trebuchet MS", "Verdana", "Webdings", "Wingdings", "Yu Gothic", "Aptos", "Bahnschrift",
|
|
"HoloLens MDL2 Assets", "Cascadia Code", "Cascadia Mono", "Segoe Fluent Icons",
|
|
"American Typewriter", "Andale Mono", "Apple Chancery", "Apple Color Emoji",
|
|
"Apple SD Gothic Neo", "Avenir", "Avenir Next", "Baskerville", "Big Caslon",
|
|
"Chalkboard", "Chalkboard SE", "Charter", "Cochin", "Copperplate", "Didot",
|
|
"Futura", "Geneva", "Gill Sans", "Helvetica", "Helvetica Neue", "Herculanum",
|
|
"Hoefler Text", "Lucida Grande", "Marker Felt", "Menlo", "Monaco", "Noteworthy",
|
|
"Optima", "Papyrus", "PingFang SC", "Rockwell", "SF Pro", "Skia", "Snell Roundhand",
|
|
"Zapfino", "Arial Hebrew", "Hiragino Sans", "Kohinoor Devanagari", "Thonburi",
|
|
"DejaVu Sans", "DejaVu Serif", "DejaVu Sans Mono", "Liberation Sans",
|
|
"Liberation Serif", "Liberation Mono", "Noto Sans", "Noto Serif", "Noto Color Emoji",
|
|
"Noto Sans CJK SC", "Ubuntu", "Ubuntu Mono", "Cantarell", "FreeSans", "FreeSerif",
|
|
"Nimbus Sans", "Nimbus Roman", "C059", "P052", "URW Bookman", "Droid Sans Fallback",
|
|
]
|
|
|
|
FINGERPRINT_PAGE = """<!doctype html><html><body><script>
|
|
(async () => {
|
|
try {
|
|
const within = (p, ms) => Promise.race([p, new Promise(r => setTimeout(() => r(null), ms))]);
|
|
const fp = {};
|
|
const n = navigator;
|
|
Object.assign(fp, {
|
|
userAgent: n.userAgent, platform: n.platform, oscpu: n.oscpu,
|
|
hardwareConcurrency: n.hardwareConcurrency, languages: n.languages.join(","),
|
|
maxTouchPoints: n.maxTouchPoints,
|
|
screen: [screen.width, screen.height, screen.availWidth, screen.availHeight,
|
|
screen.colorDepth].join("x"),
|
|
outer: [outerWidth, outerHeight].join("x"),
|
|
devicePixelRatio: devicePixelRatio,
|
|
timezone: Intl.DateTimeFormat().resolvedOptions().timeZone,
|
|
});
|
|
const gl = document.createElement("canvas").getContext("webgl");
|
|
if (gl) {
|
|
const dbg = gl.getExtension("WEBGL_debug_renderer_info");
|
|
fp.webglVendor = dbg && gl.getParameter(dbg.UNMASKED_VENDOR_WEBGL);
|
|
fp.webglRenderer = dbg && gl.getParameter(dbg.UNMASKED_RENDERER_WEBGL);
|
|
fp.webglLimits = [gl.MAX_TEXTURE_SIZE, gl.MAX_VERTEX_UNIFORM_VECTORS,
|
|
gl.MAX_FRAGMENT_UNIFORM_VECTORS, gl.MAX_VARYING_VECTORS,
|
|
gl.MAX_RENDERBUFFER_SIZE].map(p => gl.getParameter(p)).join(",");
|
|
fp.webglExtensions = (gl.getSupportedExtensions() || []).join(",");
|
|
}
|
|
// Width against each generic fallback, as fingerprinting scripts detect
|
|
// installed fonts. document.fonts.check() only knows web fonts and answers
|
|
// true for any system font name, so it cannot tell.
|
|
const measure = document.createElement("canvas").getContext("2d");
|
|
const text = "mmmmmmmmmmlliWW@#&0123456789";
|
|
const width = family => { measure.font = `72px ${family}`; return measure.measureText(text).width; };
|
|
const bases = ["monospace", "sans-serif", "serif"];
|
|
const baseWidths = bases.map(width);
|
|
fp.fonts = FONTS.filter(f => bases.some((base, i) => width(`"${f}", ${base}`) !== baseWidths[i])).join(",");
|
|
let voices = speechSynthesis.getVoices();
|
|
if (!voices.length) {
|
|
await within(new Promise(r => speechSynthesis.addEventListener("voiceschanged", r, {once: true})), 3000);
|
|
voices = speechSynthesis.getVoices();
|
|
}
|
|
fp.voices = voices.map(v => v.name).join(",");
|
|
const devices = await within(n.mediaDevices.enumerateDevices(), 3000);
|
|
fp.mediaDevices = devices ? ["audioinput", "videoinput", "audiooutput"]
|
|
.map(k => devices.filter(d => d.kind === k).length).join(",") : "timeout";
|
|
const ctx = new OfflineAudioContext(1, 5000, 44100);
|
|
const osc = ctx.createOscillator();
|
|
osc.type = "triangle";
|
|
osc.frequency.value = 10000;
|
|
const comp = ctx.createDynamicsCompressor();
|
|
osc.connect(comp);
|
|
comp.connect(ctx.destination);
|
|
osc.start(0);
|
|
const data = (await ctx.startRendering()).getChannelData(0);
|
|
let sum = 0;
|
|
for (let i = 4500; i < 5000; i++) sum += Math.abs(data[i]);
|
|
fp.audio = sum;
|
|
document.documentElement.dataset.fingerprint = JSON.stringify(fp);
|
|
} catch (e) {
|
|
document.documentElement.dataset.fingerprint = JSON.stringify({error: String(e)});
|
|
}
|
|
})();
|
|
</script></body></html>"""
|
|
|
|
|
|
async def full_fingerprint(page) -> dict:
|
|
"""Everything a fingerprinting script reads, computed in the page.
|
|
|
|
Served from an https URL Playwright fulfils locally: mediaDevices and other
|
|
APIs exist only in a secure context, which about:blank content is not.
|
|
"""
|
|
html = FINGERPRINT_PAGE.replace("FONTS", json.dumps(FONT_CANDIDATES))
|
|
await page.route("https://fingerprint.camoufox.test/",
|
|
lambda route: route.fulfill(content_type="text/html", body=html))
|
|
await page.goto("https://fingerprint.camoufox.test/")
|
|
for _ in range(60):
|
|
result = await page.evaluate("document.documentElement.dataset.fingerprint || null")
|
|
if result:
|
|
fingerprint = json.loads(result)
|
|
assert "error" not in fingerprint, fingerprint["error"]
|
|
return fingerprint
|
|
await asyncio.sleep(0.5)
|
|
raise AssertionError("the fingerprint page produced no result in 30s")
|
|
|
|
|
|
async def test_two_browsers_get_different_fingerprints(binary):
|
|
"""Two launches must not present the same device.
|
|
|
|
Measured on the whole fingerprint a site computes. The coarse values alone
|
|
can legitimately match -- two real Macs share a UA, a 2560x1440 screen and
|
|
8 cores, and CI pins timezone and language -- so a check on those alone
|
|
failed whenever two draws landed on a common machine. Fonts, voices, the
|
|
GPU and the per-identity audio seed together cannot.
|
|
"""
|
|
from camoufox.async_api import AsyncCamoufox
|
|
|
|
async def one() -> dict:
|
|
async with AsyncCamoufox(executable_path=str(binary), headless=True,
|
|
i_know_what_im_doing=True) as browser:
|
|
context, page = await open_page(browser)
|
|
fingerprint = await full_fingerprint(page)
|
|
await context.close()
|
|
return fingerprint
|
|
|
|
a, b = await asyncio.gather(one(), one())
|
|
assert "audio" in a and "audio" in b, (a, b)
|
|
differing = sorted(k for k in a if a.get(k) != b.get(k))
|
|
assert differing, f"two separate browser launches produced an identical fingerprint: {a}"
|
|
# The audio noise is seeded per identity, so it must differ on its own:
|
|
# equal hashes would mean the seed stopped reaching the browser.
|
|
assert a["audio"] != b["audio"], f"both launches rendered audio hash {a['audio']}"
|
|
|
|
|
|
async def test_a_context_survives_its_sibling_browser(binary):
|
|
"""Two browsers are two processes; one closing must not affect the other."""
|
|
from camoufox.async_api import AsyncCamoufox
|
|
|
|
async with AsyncCamoufox(executable_path=str(binary), headless=True,
|
|
i_know_what_im_doing=True) as keeper:
|
|
ctx_keep, page_keep = await open_page(keeper)
|
|
before = await probe(page_keep)
|
|
|
|
async with AsyncCamoufox(executable_path=str(binary), headless=True,
|
|
i_know_what_im_doing=True) as transient:
|
|
ctx_t, page_t = await open_page(transient)
|
|
await probe(page_t)
|
|
await ctx_t.close()
|
|
|
|
after = await probe(page_keep)
|
|
await ctx_keep.close()
|
|
|
|
assert before == after, "closing a second browser perturbed the first one's fingerprint"
|
|
|
|
|
|
async def test_pages_in_one_context_share_its_fingerprint(binary):
|
|
"""A context is the isolation boundary; a page is not.
|
|
|
|
Two pages in one context must agree, or the boundary has been drawn in the
|
|
wrong place and a site could tell two of its own tabs apart.
|
|
"""
|
|
from camoufox.async_api import AsyncCamoufox
|
|
|
|
async with AsyncCamoufox(executable_path=str(binary), headless=True,
|
|
i_know_what_im_doing=True) as browser:
|
|
context = await browser.new_context()
|
|
page_one = await context.new_page()
|
|
await page_one.goto("about:blank")
|
|
page_two = await context.new_page()
|
|
await page_two.goto("about:blank")
|
|
one, two = await probe(page_one), await probe(page_two)
|
|
await context.close()
|
|
|
|
assert one == two, (
|
|
"two pages in the SAME context reported different fingerprints:\n"
|
|
f" page 1 {one}\n page 2 {two}"
|
|
)
|