* feat(humanize): replay recorded human mouse movements (Cursory)
humanize=True used to walk a Bezier curve through two random knots and emit a
point every 10 ms. Both halves are tells: an analytic curve sampled at a fixed
rate has velocity and jerk profiles that separate cleanly from a hand's, and
every movement accelerated through the same easing function.
Juggler now picks one of Cursory's 2357 recorded human movements whose
direction, distance and wander suit the move, morphs it onto the requested
endpoints and replays it with the recording's own timing. The generator is
cursory-js (a bit-exact TypeScript port of Vinyzu/cursory) vendored under
additions/juggler/input/cursory/; it is LGPLv3-or-later, not MPL-2.0, and ships
its LICENSE and NOTICE inside juggler.jar.
MouseTrajectories.hpp and ChromeUtils.camouGetMouseTrajectory are removed.
sendTrajectoryAcked takes per-step pauses, drops points on the pixel the last
dispatch left the cursor on (a zero-displacement move is never acked), and the
humanize guards are updated for the new path shape.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(patches): shared helpers for binary resolution, a private Xvfb and Marionette
resolve_binary() honours the runner's CAMOUFOX_EXECUTABLE_PATH before falling
back to a Linux objdir (search-service-init and touchscreen-digitizer ignored it
and ran the newest objdir, which after a macOS cross build is an arm64 Mach-O),
hidden_display() gives a guard its own Xvfb so nothing ever opens on the user's
display, and a minimal chrome-context Marionette client lets guards inspect
browser UI state.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(juggler): synthesized input carries what real mouse and keyboard input carries
- pointerType was "" for every Playwright mouse event: juggler dispatched with
MOZ_SOURCE_UNKNOWN. It now passes MOZ_SOURCE_MOUSE (#776).
- keyboard.type() never pressed Shift: a shifted character now arrives
bracketed by ShiftLeft keydown/keyup (location 1) with shiftKey set.
- After the pointer was parked off content, pointerover/enter re-entered with
buttons=1 and pressure 0.5; the tracked position is now forgotten on park.
- A Windows identity gets contextmenu after mouseup with buttons=0, as Windows
does; GTK/macOS keep it on press.
- Wheel events are sent as line deltas (DOMMouseScroll.detail 3 per notch
instead of the pixel count).
- The browser rect is measured after the APZ flush await, so a chrome height
change during the wait cannot put a y==0 dispatch one row above content.
- ci/run_sundial.py moves and clicks the mouse so input vectors have data.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(juggler): evaluate() no longer grants user activation
Upstream Playwright runs every evaluate() as handling user input and notifies a
user-gesture activation. Init scripts go through that path at load, so every
page started with navigator.userActivation.hasBeenActive === true, autoplay
allowed and popups permitted before any input. Activation now only comes from
juggler's trusted input events, as in a stock browser.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(juggler): stop hiding scrollbars in headless
The headless agent sheet set scrollbar-width: none !important, which a page
reads back from getComputedStyle and from overflow:scroll gutters. Scrollbar
appearance is left to the platform look-and-feel (the launcher sets
ui.useOverlayScrollbars per claimed OS).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(ui): no visible automation cues in the browser window
- Every Playwright context was a public container, so the URL bar showed a
"JUGGLER <id>" label and a container colour. Contexts are now non-public
identities (tabbrowser renders public identities only); startup cleanup
still removes persisted leftovers.
- showcursor defaulted to true, drawing a red dot that followed the mouse.
It is now opt-in.
tests/patches/visible-automation-cues.py checks both on a private Xvfb.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(fonts): web fonts, local(), per-character fallback and native bundles
- FontFace / @font-face were answered from the font allowlist by the FontFace's
own family name, so every url() web font failed with NS_ERROR_FAILURE and
never rendered, local() of an allowed font failed, and a miss rejected with an
XPCOM code instead of NetworkError (#759). Stock FontFace/FontFaceImpl are
restored; local() is filtered by the RESOLVED family in gfxUserFontSet.
- GlobalFontFallback forced the cmap scan, which skips families whose charmap is
not loaded yet, so any character outside Gecko's script-based common-fallback
table rendered as the primary family's .notdef (U+1E9E on macOS). The platform
fallback chooses again, and its choice is held to the mask.
- For a native macOS/Windows identity the bundled font sets are not activated:
a bundled face of a family the system also has (Papyrus, Helvetica) won the
lookup with different metrics. On Windows the enumerator still keeps Twemoji
Mozilla, the emoji font stock Firefox ships (flag emoji drew nothing).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(fonts): CSS2 system fonts and system-ui follow the claimed OS
- The host's own OS gets no system-ui override (macOS resolved system-ui to
Helvetica instead of -apple-system).
- CSS2 system font keywords use per-keyword faces and sizes; a Linux identity
reports the Ubuntu desktop font; Windows form controls (-moz-button/field/list)
answer "MS Shell Dlg 2" as Windows does, not Segoe UI.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(fonts): per-OS font model and fontconfig parity
fonts.json is now generated from the bundle by scripts/gen-fonts-json.py
(fc-scan + aliases + scan-time families, intersected with the per-OS manifest in
scripts/data/font-manifests.json) so every reportable family is renderable;
scripts/verify-fonts.py checks that invariant, the generics and the reject
globs. font-groups.json lets the draw keep co-shipped families together.
Linux fontconfig: stock metric aliases (Arial -> Liberation Sans, ...),
49-sansserif, urw-base35 and the non-Latin rule files in stock conf.d order,
generics resolving like a stock Ubuntu (Noto Sans / Noto Serif / DejaVu Sans
Mono / Z003), hintslight so advances are not pinned to whole pixels, and weak
<prefer> lists instead of strongly-bound generic pins so lang can promote a
script face. Windows fontconfig: GDI substitution aliases, MS Shell Dlg 2,
cursive/fantasy generics, duplicate-face rejects and Sitka / Segoe UI Variable
optical-size families.
NOTE: generated against a ~3.9 GB target font bundle that is not part of this
change (one file is over GitHub's 100 MB limit); see docs/FONTS.md.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(locale): localize browser strings with the spoofed locale; stop rewriting explicit locales
- With locale="fr-FR", Intl/number/date went French but input.validationMessage
and XML parse errors stayed English, a mix no real Firefox produces. Official
language packs are now baked in as packaged locales (scripts/fetch-langpacks.py,
scripts/inject-locales.py, called by package.py, fetched on demand) and the
launcher selects the UI locale through intl.locale.requested. A langpack
add-on cannot do this: the parent pre-creates those string bundles first.
- locale-spoofing.patch overrode Language/Script/Region on every intl::Locale,
so new Intl.DisplayNames(['en'],{type:'region'}).of('DE') returned the spoofed
region's name and Intl.Locale('ja-Jpan-JP').minimize() returned the spoofed
tag. Only the OS/default locale is spoofed now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(media): enumerate, capture and label the identity's media devices coherently
The fake media engine now enumerates the identity's microphones, cameras and
speakers (labels and group ids from new mediaDevices:*Labels/*Groups config
keys), MediaManager uses it whenever mediaDevices:enabled, and stock exposure
rules apply: before a grant one device per input kind, no outputs, no labels;
after a grant OS-style labels, distinct deviceIds, shared groupIds. So
enumerateDevices(), getUserMedia() tracks and getSettings() ids agree, and a
claimed camera captures instead of throwing NotFoundError. Fixes the
content-process crash on an identity with a camera and no microphone
(InsertElementAt on an empty array). docs/MEDIA-DEVICES.md; guard
tests/patches/media-devices-coherence.py.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(navigator): globalPrivacyControl agrees between window and workers (#760)
The main-thread Navigator getter ignored the config key that
WorkerNavigator::GlobalPrivacyControl honours, so a page read false in the
window and true in a worker. Both read the key the same way now; the launcher
also mirrors it into privacy.globalprivacycontrol.enabled so the Sec-GPC header
agrees.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(timezone): apply the launch-level timezone from the first read (#773)
The timezone config key was only applied lazily from a navigator getter, so
Intl and Date reported the host zone until a page happened to touch navigator.
It is now applied eagerly in every process (nsJSContext::EnsureStatics) and per
realm when a new inner window is created, entering that window's realm rather
than whichever one triggered the navigation. window.setTimezone() still takes
precedence per context.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(screen): the CSS color media feature follows the spoofed colorDepth
screen.colorDepth was spoofed at the WebIDL level only, so on a 10-bit panel a
24-bit identity reported 24 with (color: 10), a pair Gecko cannot produce.
Gecko_MediaFeatures_GetColorDepth now resolves the depth in the same order as
nsScreen::PixelDepth.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(webgl): pass live state through instead of answering it from the table
getParameter answered everything from the sampled table, so state a page had
just set read back wrong (lineWidth(5) read 1, VIEWPORT/SCISSOR_BOX stayed
300x150 on a 64x64 canvas), extension parameters were null (anisotropy, draw
buffers), COMPRESSED_TEXTURE_FORMATS was null instead of [], and
getContextAttributes() ignored the attributes requested ({antialias:false}
still reported 4 samples). Identity and limits still come from the table; live
state and context attributes are real.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(webrtc): ICE gathering completes behind a proxy (#774)
With Playwright's per-context proxy and
media.peerconnection.ice.proxy_only_if_behind_proxy, ICE failed before
gathering started and iceGatheringState stayed "new" forever, where stock
Firefox completes with host candidates. When that happens around the
fabricated candidates the new -> gathering -> complete state walk is replayed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(patches): refresh window-setter-seal.patch offsets
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(windows): embed Firefox's application manifest in camoufox.exe
config/rules.mk embeds <program>.manifest and browser/app only ships
firefox.exe.manifest, so --with-app-name=camoufox produced an exe with no
manifest. Without the Windows 10 supportedOS GUID the process and its children
run as a pre-Windows-10 application and Gecko's Windows-10-gated paths switch off
(MediaCapabilities.decodingInfo powerEfficient false for H.264/VP9 where stock
is true). The new patch adds a byte-for-byte copy as camoufox.exe.manifest; the
old rename hunk in windows-theming-bug-modified.patch is dropped. Guard:
tests/patches/windows-exe-manifest.py.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(settings): stock values for page-observable prefs; launcher prefs at startup
Page-observable defaults that no stock Firefox has, restored:
- the forced built-in dark theme (it also removed the 1 px nav-bar separator,
and was applied ~1 s after startup, resizing the viewport) and
ui.systemUsesDarkTheme (prefers-color-scheme disagreed with the desktop);
- focus rings off, autoplay allowed, popup blocker off;
- gfx.color_management.mode=0 (Playwright's test pref: ICC-tagged images were
drawn unconverted, readable from a canvas pixel);
- ui.use_standins_for_native_colors (non-native system colours);
- GMP updates off (Widevine/OpenH264 never available);
- storage.estimate() quota derived from the raw disk instead of the stock cap.
The HardwareAcceleration:false enterprise policy is removed: it locked software
WebRender with no hardware video decoding on every OS (guard
tests/patches/hardware-acceleration-policy.py). The minimal-theme chrome.css is
emptied: its ~55 px chrome made outerHeight - innerHeight impossible.
Playwright's non-persistent launch writes no user.js, so launcher prefs only
arrived through juggler after startup and anything Gecko reads while starting
raced (on Windows the UI locale lost 3 of 4 launches). camoufox.cfg now applies
the launcher's CAMOU_PREFS_1..N env chunks as default prefs at startup (guard
tests/patches/startup-prefs.py).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(branding): chrome://branding assets match stock Firefox
chrome://branding/content/ is content-accessible. The wordmark SVGs had
different intrinsic sizes (336x48 / 172x48 vs 300x67) and document.ico,
document_pdf.svg and the private-browsing about logos were missing, all
measurable from a page with an <img>.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(pythonlib): identity draws that match real machines and stay stable
Launcher-side fixes found by comparing camoufox against stock Firefox 152.0.4
on Linux, Windows 11 and macOS hosts:
- DNT / GPC: BrowserForge draws doNotTrack "1" on most Firefox samples, but a
stock Firefox 152 reports "unspecified" and globalPrivacyControl false; the
stock defaults are used unless the caller sets them, and both are applied as
prefs so the API, the worker and the DNT / Sec-GPC headers agree (#760).
- Timezone and geolocation: the timezone is passed to the browser, and a
configured position sets permissions.default.geo so permissions.query agrees
with the auto-grant (#769, #773).
- hardwareConcurrency: the reported count is the fingerprint's and the browser
is pinned to that many cores (cpu_affinity.py, Linux/Windows), so worker
timing agrees with it; otherwise the host count snapped into the core counts
real machines ship with (never 2, Firefox's resistFingerprinting value).
- Fonts: the OS base is always present in full, OS-version variants are drawn
all-or-nothing, co-shipped groups stay together, Cascadia is never claimed
off Windows, a native macOS/Windows identity claims only the real OS base,
and gfx.font_rendering.fallback.async is off on Linux so per-character
fallback does not depend on cmap-load timing.
- Speech voices: a per-OS installed-voice model (voice-manifests.json) with
the voiceURI formats each backend really produces (voice-uris.json); no
default voice where stock has none.
- WebGL: extensions a release Firefox never exposes are filtered, but
OVR_multiview2 stays for Windows D3D11 renderers, which expose it.
- Media devices: a seeded draw of common per-OS devices with OS-style labels.
- Windows scrollbars follow the drawn Windows version (overlay on 11).
- Glyph-advance perturbation (fonts:spacing_seed) defaults to off: it moved
every measureText width off the value the same font gives on a real machine.
- Launcher prefs are also exported as CAMOU_PREFS_1..N so camoufox.cfg applies
them at startup, and the browser UI locale follows the spoofed locale.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(pythonlib): per-identity salt for seeded draws; core pinning under concurrency
Found in review of the previous commits:
- identity_seed() hashed only the UA, platform, screen size and core count.
Those take a handful of values per OS, so over 500 launches the seed took
12-30 distinct values and every install drew its fonts, voices, GPU, media
devices and canvas/audio noise seeds from that same short list. The seed now
mixes in identity_salt(): derived from what the caller pinned the identity
with (a Fingerprint, a preset dict, a config naming the UA) so relaunching
that identity reproduces every draw, and random otherwise. A pinned preset
now reproduces its noise seeds too; seeds the caller sets are kept.
- Concurrent AsyncNewBrowser launches on one driver interleaved pin/restore:
one browser inherited the other's mask and the driver could stay pinned.
pin -> launch -> restore is serialized per driver.
- Every pinned browser landed on cores 0..N-1; pins now take N adjacent cores
from a random start.
- A pinnable host with 1-3 cores reported 1, 2 or 3 (2 is the
resistFingerprinting value); the table floor of 4 applies as on other hosts.
- launch_options() callers that launch the browser themselves (launch_server,
direct use) kept the drawn core count although nothing pins the browser;
only Camoufox/AsyncCamoufox pass pin_cpu_cores=True now, everyone else
reports the host's snapped count.
- PLAUSIBLE_CORE_COUNTS gains 18, 22, 28 and 32, all recorded in the -v150
corpus.
- The Windows voice list was drawn before the locale was resolved, so an
fr-FR identity got en-US voices; it is drawn after locale/geoip now.
- macOS "Alex" gets its com.apple.speech.synthesis.voice identifier.
- CAMOU_PREFS env chunks are ASCII-only JSON (Windows getenv goes through the
ANSI code page).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(fonts): canvas accepts CSS2 system-font keywords; local() works on macOS
- GetSpoofedSystemFontForRFP's per-OS branches returned before marking the
result a system font. ComputeSystemFont copies that flag into
FontFamilyList::is_system_font, and without it the canvas font setter could
not serialize the value: ctx.font = 'caption' (or icon, menu, message-box,
small-caption, status-bar) was silently ignored and read back
'10px sans-serif' where stock reads back the keyword.
- CoreTextFontList::LookupLocalFont builds a CTFontEntry with no family name,
and local() sources are held to the spoofed font list by the resolved
family, so on macOS every local() face (Helvetica, Menlo, Arial...) failed
with NetworkError, installed and allowed or not. The entry now carries the
family CoreText resolved. A blocked lookup's entry is released instead of
leaked.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(webgl): only device limits come from the spoofed table
getParameter still answered ~100 state pnames from the table, so state the page
had just changed read back wrong: UNPACK_FLIP_Y_WEBGL / PREMULTIPLY_ALPHA /
COLORSPACE_CONVERSION after pixelStorei, FRAGMENT_SHADER_DERIVATIVE_HINT after
hint(), DRAW_BUFFERi after drawBuffers(), RED/ALPHA/DEPTH/STENCIL_BITS and
IMPLEMENTATION_COLOR_READ_* for the bound framebuffer, and COMPRESSED_TEXTURE_
FORMATS after enabling an extension. UNMASKED_VENDOR/RENDERER_WEBGL came back
without the extension enabled, where stock returns null with INVALID_ENUM.
The table now answers only the MAX_*/ALIASED_*/SUBPIXEL_BITS limits, WebGL 2
limits on WebGL 2 contexts only, and extension limits (anisotropy, draw
buffers, OVR multiview) only once that extension is enabled; everything else is
the real context's answer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(webrtc): fabricate candidates only where a real gather would have them
With webrtc:ipv4/ipv6 set (every geoip launch), new RTCPeerConnection() with no
iceServers produced a srflx carrying the spoofed IP, and after end-of-candidates
a second host set with a different mDNS name. getStats() exposed that srflx as
id 'camou-srflx' and rewrote every candidate address, including .local host
names and the remote peer's candidates.
A srflx is now fabricated only when the page configured an ICE server, host
candidates only when none reached the page (sharing the real UDP host's port
otherwise), the synthetic stats id has the shape real candidate ids have (8 hex digits,
fixed per connection), and only this side's non-mDNS addresses are rewritten.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(media): honour mediaDevices:enabled=false; page fake:true behaves as stock
- MaskConfig::GetBool returns std::optional<bool>, and the checks tested its
presence: "mediaDevices:enabled": false still enabled the fake devices.
- media.navigator.permission.fake=true was page-readable: a page's own
getUserMedia({video: true, fake: true}) prompted and never resolved, where
stock resolves at once with its generic fake device. The pref is off again;
the identity's devices count as real hardware in the capturing checks
instead (prompt, sharing indicator, post-grant labels), and a page's
fake:true request gets stock's generic devices rather than the identity's.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(storage): per-context values are session state, and misses are cached
- Values lived on the user pref branch, which a persistent profile writes to
prefs.js: relaunching with a different timezone (or navigator values) kept
reporting the previous session's in the page, iframes and workers. They now
live on the default branch, which is never saved, and reads ignore user
values an older build left behind.
- A read of an unset key did a synchronous IPC to the parent every time, and in
a launch without per-context values every read is unset:
navigator.hardwareConcurrency, screen.* and (color) media queries measured
~20x slower than stock. A miss is now cached per key until a pref change or a
local put clears it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(timezone): no per-realm override for the process-wide zone; cache DateTimeInfo
With a launch-level timezone every new document and worker got a per-realm
override of the zone the process already reported. Setting one releases all JIT
code in the runtime (hot code after adding an iframe ran ~4x slower), and the
realm rebuilt its DateTimeInfo on every call (getHours() ~40x slower than
stock). The override is applied only when the zone differs from the one
JS::SetTimeZoneOverride applied process-wide, and a realm keeps its DateTimeInfo
until its override changes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(juggler): wheel scrolls in native notches; Shift leads the key it modifies
- A wheel notch now reaches the page as its own 3-line event carrying one
native tick (new WHEEL_EVENT_NATIVE_NOTCHES option in
patches/wheel-native-ticks.patch), so wheelDelta is -120 per notch as with a
physical wheel; it was -396, and a multi-notch scroll arrived as one event.
Several notches are spaced a few tens of ms apart.
- Auto-Shift pressed Shift 0 ms before the character's keydown and released it
0 ms after its keyup; it now leads and trails by a drawn human-scale delay,
and a failing keydown no longer leaves Shift latched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(settings): stock cookie partitioning, preconnect, popup notification; about dialog CSS
- network.cookie.cookieBehavior 4 (Playwright's) -> Firefox's default 5. With 4 a
cross-site iframe (captcha and anti-bot widgets are exactly that) sees
document.hasStorageAccess() true and its first-party cookies and
localStorage, where stock partitions them. Playwright set 4 so storageState
need not carry thirdPartyCookie^ permissions.
- network.http.speculative-parallel-limit 0 turned <link rel=preconnect> into a
no-op, visible in Resource Timing.
- privacy.popups.showBrowserMessage false: stock shows a notification bar for a
blocked popup, which shrinks the viewport and fires resize.
- chrome://branding/content/aboutDialog.css is page-loadable and was empty; it is
the official branding's now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(patches): stock-parity probes for the leaks found in review
One launch (plus two persistent relaunches) checks page-observable invariants
stock Firefox 152 holds: canvas CSS2 system-font keywords, WebGL state readback
and UNMASKED_RENDERER without the extension, no srflx without ICE servers and
no 'camou' stats id, getUserMedia({fake: true}), cross-site storage
partitioning, wheel notches, per-read cost of (color)/hardwareConcurrency and
of local Date getters under a launch timezone, and a persistent profile's
timezone after relaunch (page and worker). Run against the build before these
fixes it fails on every one of them.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(humanize): judge cadence by distinct values and spread, not a share of the gap count
The page clock is clamped to 1 ms and Cursory's recorded steps mostly sit
between 12 and 20 ms, so the number of distinct gap values cannot grow with the
number of gaps. Requiring len(gaps) // 4 made the guard fail on visibly uneven
runs whenever event delivery was steady (3 of 4 runs once the per-read sync IPC
jitter was gone). A fixed 10 ms cadence yields about three values within a few
ms of each other, which the new rule (>= 6 values, >= 20 ms spread) still fails.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(juggler): restore the popup, wheel and history contracts
The Playwright suite went red on this branch, and five of its six shards ran out
their 40-minute budget before reporting, so ~470 sync tests were never run at
all. Three causes, all ours:
- window.open() from page.evaluate() returned null. The popup blocker being ON
(Firefox's default) and evaluate() no longer granting user activation are each
defensible alone; together they block every gesture-less popup. ~35 tests, each
burning 30s x 4 attempts x 2 worlds, which is what exhausted the shards.
The blocker goes back to Playwright's and geckodriver's value. Reading it costs
a detector a popup window the user sees, so it is not a check an anti-bot
script in the page runs -- unlike navigator.userActivation.hasBeenActive, which
is one property read, and which is why the activation half stays.
- mouse.wheel(0, 100) delivered deltaY 114 (or 132, depending on the host's font
metrics) in deltaMode 1. Quantising into native wheel notches is what a
physical wheel does, but it changes the number the caller asked for, so it now
rides behind humanize= with the rest of the humanized input. Default is the
exact requested delta in deltaMode 0.
- page.go_back() did nothing after history.pushState(). canGoBack is the BACK
BUTTON's answer: under browser.navigation.requireUserInteraction it reports
false when every entry behind this one was pushed without the user touching the
page, which is now every entry, because evaluate() grants no activation.
goBack() itself does not skip those entries and neither does history.back(),
so ask canGoBackIgnoringUserInteraction, as Marionette does.
Two keyboard tests are skiplisted rather than fixed: auto-Shift means typing "!"
emits the Shift a US keyboard requires, and upstream asserts the character's
three events with shiftKey false throughout. The character's own key/code/keyCode
are unchanged; what upstream asserts is the absence of a Shift no real typist
could omit.
Full suite against the fixed build: 2223 passed, 6 failed -- the two keyboard
tests above, and four client-certificate tests that fail only on this machine
(Node/OpenSSL rejects the fixture server) and pass in CI.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(pythonlib): stop overriding the corpus on core counts; pinning is opt-in
Two findings from auditing the sweep's fixes against one bar: a difference is a
leak only if a page's JavaScript can actually read it.
hardwareConcurrency 2 was excluded from PLAUSIBLE_CORE_COUNTS because "2 is what
Firefox reports under resistFingerprinting". That has not been true for years:
RuntimeService::ClampedHardwareConcurrency hardcodes 4, and 8 on macOS, both of
which are already in the table. The exclusion protected against nothing and cost
every genuinely dual-core machine -- 20% of the macOS presets in the recorded
corpus, 4.2% of Linux draws. It also made the small-host tail worse: a 3-core
host reported 4, which cannot be pinned, so a page measured 3 while being told 4.
At 2 the pin succeeds.
pin_cpu_cores now defaults to False. What it buys is defence against a page
timing N parallel workers; what it costs is a browser-wide CPU cap, a per-driver
launch lock, and nothing at all on macOS. Unpinned, the host's own snapped count
is reported, so reported and measurable still agree -- the identity just loses
one drawn value. Callers who want the draw kept can still ask for it.
The WebGL sampler keeps rejecting software rasterisers, and its docstring now
says so: it described the opposite of what the code does. llvmpipe as the
presented GPU is a live check on a string every fingerprint script reads, which
is worth ~1.5% of corpus fidelity.
251 pythonlib tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(pythonlib): keep 2 out of the core table -- an Apple M1 is never dual-core
Reverts the PLAUSIBLE_CORE_COUNTS half of fd501e5. The conclusion there was
wrong, even though the fact that prompted it was right.
Right: "2 is what Firefox reports under resistFingerprinting" is false, and has
been for years. RuntimeService::ClampedHardwareConcurrency hardcodes 4, and 8 on
macOS, both already in the table.
Wrong: concluding from that, and from the corpus recording 2 on 20% of macOS
presets, that 2 should be drawable. 85% of macOS identities draw
"Apple M1, or similar" as the WebGL renderer, and no Apple Silicon part has ever
had fewer than 8 cores. A page reading navigator.hardwareConcurrency and
UNMASKED_RENDERER_WEBGL together -- two property reads, both already in every
fingerprint payload -- would see a machine that does not exist.
The corpus frequency is not counter-evidence. Those rows report 2 more often
than 4 on macOS (11 vs 3), which no real hardware population does: the corpus is
scraped from live traffic, so it carries privacy-hardened browsers, 2-vCPU VMs
and other people's bots. The corpus settles what real machines report where the
field is hardware; hardwareConcurrency is a number a browser can be made to say.
The comment now records the true reason, so the next reader does not undo this
by discovering the RFP claim is false -- which is exactly how it came undone.
pin_cpu_cores stays opt-in, and the WebGL software-rasteriser rejection is
unchanged. 251 pythonlib tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(pythonlib): a preset naming an unknown GPU must not fail the launch
The pythonlib job failed on a Windows preset whose GPU is
"ANGLE (Unknown, Adreno (TM) 650 Direct3D11 vs_5_0 ps_5_0)" -- a phone GPU, in
the Windows pool. sample_webgl raised, and launch_options() with it.
The crash is not new here: upstream/main has the same branch, which looks the
preset's vendor/renderer up in webgl_data.db to get the parameters that belong
to it. What is new is a test that draws a RANDOM preset, so it surfaced as a
1-in-11 flake instead of a deterministic failure. 39 of the 435 bundled presets
name a GPU that is not among the 33 in the database, so ~9% of preset launches
have always raised -- and a caller passing their own preset dict had no way to
know which pairs are supported.
The named GPU cannot be kept: parameters, extension list and shader precisions
all have to come from one real recorded device, and there is none for an unknown
renderer. So the fallback draws a GPU that fits the screen and REPLACES the
pair. Replacing it needs the two keys popped first, because merge_into does not
overwrite what the preset already set -- without that the page reads
"Adreno (TM) 650" with a desktop GPU's parameters behind it, which is a louder
mismatch than the one being fixed. Measured on the bundled presets: 30 of 312
now swap, e.g. a macOS preset claiming a 2008 Radeon HD 3200 becomes Apple M1.
Guarded by a parametrised test over EVERY bundled preset in all three pools,
asserting the pair that survives is one the database knows. It fails without the
fix. 254 pythonlib tests pass.
The rows themselves are corpus contamination and worth cleaning separately: the
presets are scraped from live traffic, so the Windows pool carries an Android
GPU and the macOS pool carries GPUs no Mac has shipped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(patches): check both wheel modes, not just the notched one
The stock-parity guard asserted that mouse.wheel(0, 300) arrives as three
notched events. That is now the humanize=True behaviour, not the default, so the
guard failed on the build it was meant to certify -- it encoded one side of a
decision that has two sides.
It now checks both: with humanize off, one event carrying the delta the caller
asked for (deltaMode 0, deltaY 300); with humanize on, three events whose
wheelDeltaY is a multiple of 120, as a physical wheel produces. A second short
launch covers the humanized half, in the style of the timezone relaunch probe.
Verified against the local build: PASS on every probe, with the humanized scroll
arriving as 3 events of wheelDeltaY -120.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(patches): the query-cost probe measured the JIT, not the IPC
query-cost failed in CI with "hardwareConcurrency 19 ms vs userAgent 0 ms". The
0 ms is the tell: the loop reads a getter and discards the result, so the JIT
elided the BASELINE loop entirely on that runner. With the baseline at zero the
check `costHwc > 5 * costUA + 15` collapses to a flat 15 ms allowance for 20000
reads -- 0.75 us each -- while a healthy read of a value that lives in the config
costs ~1 us. It was timing whether the JIT dropped the loop.
Every read is now accumulated into a sink that is returned, so the loop cannot be
optimised away. (userAgent still measures ~0 because the string is cached, hence
the second change.)
The allowance on the two config-read checks goes to 40 ms. The state they guard
against is a sync IPC per read, measured at ~12 us each when it regressed, i.e.
~240 ms over this loop; a healthy read is ~20 ms. 40 sits an order of magnitude
under the defect and clear of a slow runner.
The timezone-cost check keeps its 15 ms allowance and is commented to say why:
its healthy numbers are ~2 ms vs ~1 ms and its regression is ~20 ms over the same
loop, so widening it to 40 would step over the very thing it exists to catch.
Verified against the local build: PASS, costHwc 7-13 ms, costLocalDate 1-2 ms.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(pythonlib): generate fingerprints with fpgen instead of BrowserForge
BrowserForge (and the Apify fingerprint-suite data behind it) is replaced by
fpgen -- scrapfly/fingerprint-generator, Apache-2.0, same author as Camoufox,
trained on Scrapfly's live traffic.
The reason is coverage. Measured over 120-200 Firefox draws per OS:
BrowserForge / Apify fpgen
distinct GPUs 2-3 per OS 6-9 sampled, 999 in the
value space (webgl_data.db
has 33)
fonts 4-22 names 539-802 on macOS, 87
distinct sets on Windows
media devices always empty for Firefox counts (labels are not
collectable -- see below)
voices absent real lists with voiceURIs
system fonts/colours absent per keyword, per OS
WebGL params absent params, extensions,
shaderPrecisionFormats,
contextAttributes
audio absent 1238 distinct hashes
This commit is the swap alone: fpgen supplies exactly what BrowserForge did --
navigator, screen, window geometry, Accept-Encoding -- through a new fpgen.yml
mapping. The richer fields are NOT wired up yet; Camoufox still draws WebGL from
webgl_data.db and fonts/voices from its own catalogues. That is the follow-up,
and the one with the real diversity win in it.
Notes for callers:
- `camoufox.fingerprints.Screen` replaces `browserforge.fingerprints.Screen`,
same four bounds. It becomes an fpgen predicate rather than a filter, and an
unsatisfiable bound (a 1x1 Xvfb) falls back to an unbounded draw, as
BrowserForge did by silently dropping the constraint.
- `fingerprint=` now takes an fpgen dict, not a browserforge Fingerprint.
- `from_browserforge()` is now `from_fpgen()`.
- fpgen fetches ~7 MB of model data from GitHub on first use, so a fully offline
install cannot generate a fingerprint. Presets and caller-supplied configs are
unaffected.
- doNotTrack is deliberately unmapped: Camoufox owns DNT, and a drawn value was
being stripped anyway.
- fpgen's own draws still need the coherence fixes -- conditioned on an Apple M1
renderer it still returns hardwareConcurrency 2 in 12% of draws. Changing the
source does not remove the need for fix_hardware_concurrency and friends.
254 pythonlib tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(pythonlib): every identity passes a whole-identity coherence check
Camoufox assembles an identity from pools sampled independently -- navigator and
screen from the generator, GPU from webgl_data.db, fonts and voices from their
own catalogues. Nothing compared them, so a machine that never existed could be
built out of parts that are each fine on their own. Cleaning the pools cannot
fix that: the incoherence is created at composition.
Measured before this, per 200 generated identities:
- 14% of macOS identities drew "Radeon R9 200 Series, or similar", a desktop
PC card, or "Intel(R) HD Graphics 400", a Braswell Atom IGP. Both are in
webgl_data.db's macOS column at 3.7% and 7.4%; neither shipped in a Mac.
- 7% paired Apple Silicon with colorDepth 24. Deep colour is the macOS
default, measured 30 on the Mac mini.
- 2% of Windows identities reported maxTouchPoints 256.
- 18 of the 312 bundled presets carried a GPU or a screen no desktop has,
including a 736x414 iPhone viewport and a portrait 1440x2560.
coherence.py states each invariant as a rule with the measurement behind it,
repairs what has a determined correct value, and reports what does not. It runs
on every identity whatever built it -- generated, preset, or caller-supplied.
Where a value can be replaced rather than repaired it is dropped BEFORE the pool
that would defer to it (a preset's own GPU pair wins over sampling, so an
impossible pair is dropped and sampling draws a coherent one), and
webgl_data.db is filtered by the same predicate before sampling, so the check
and the draw cannot disagree about what a Mac may claim.
After: 0 violations across 600 generated identities and all 312 presets.
Guarded by tests/test_coherence.py, which checks each rule against the value
that motivated it, asserts the three machines captured on 2026-09-17 pass as
themselves, and walks every bundled preset.
test_preset_screens_are_never_lifted asserted that a preset IS a real device and
its screen must never be rewritten. That premise does not survive the data, so
the test now distinguishes a small desktop panel, which is still kept, from a
phone viewport, which is repaired.
273 pythonlib tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(pythonlib): filter incoherent identities out of the shipped data too
The coherence layer filters when an identity is drawn. This filters the data it
is drawn from, so an impossible row never reaches a build: belt and braces, and
the two use the same rules.
scripts/clean-fingerprint-data.py checks each preset the way a launch converts
it, and each webgl_data.db pair against the OS weights it is offered under.
Presets are DROPPED rather than repaired -- repairing would write an invented
value ("what core count does a 2-core Apple M1 really have?") into a file whose
whole purpose is being real. GPU rows are kept with the impossible OS weight
zeroed, because "Radeon R9 200 Series" is a genuine Linux and Windows card that
simply never shipped in a Mac.
Removed, with --write:
- 38 of 435 presets: 27 with a GPU their OS cannot report, 7 pairing Apple
Silicon with a core count Apple never shipped (2, 18), 4 with a colour depth
their GPU contradicts, 3 with a phone viewport (736x414, 960x540, portrait
1440x2560), 1 claiming 40 touch points, 1 whose renderer is "Mozilla".
- 2 impossible macOS weights in webgl_data.db (Intel HD Graphics 400 at 7.4%,
Radeon R9 200 Series at 3.7%).
macOS loses the most: 97 presets -> 63. Windows 255 -> 251, Linux 83 -> 83.
tests/test_shipped_data.py asserts the files stay clean, so a refresh that
reintroduces a bad row fails CI instead of shipping, and that no OS pool is
emptied by the filter -- one GPU per OS would be its own tell.
Two rules were loosened after checking them against real hardware rather than
against the pools: maxTouchPoints is now a range (0..10) instead of a list of
values seen in scraped data, because fpgen's Windows pool never offers the 5
that win-i9 actually reports; and APPLE_SILICON_CORES gained 11, which is the
M3 Pro's 6P+5E. Over-filtering costs realism as surely as under-filtering.
A new device-pixel-ratio rule snaps scraped artefacts (1.818, 1.09) to the
nearest real scaling step, and rejects fractional DPR on macOS, which has none.
The three captured machines (dpr 1 / 2.5 / 2) pass as themselves.
The writer reproduces each file's own formatting byte-for-byte on a no-op run,
so the diff of a real run is the dropped rows and nothing else. Verified: every
surviving preset is unchanged and the metadata is untouched.
278 pythonlib tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: point CLAUDE.md at the coherence layer and the data cleaner
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: a pythonlib data tool must not cost an hour of compiling
The build cache is keyed on a hash of the inputs that can change compiled
output, and scripts/ is hashed wholesale because it holds the build machinery
(patch.py, copy-additions.sh, package.py). It also holds tools that only rewrite
the PYTHON package's data files, and those pay the same price: adding
scripts/clean-fingerprint-data.py, which edits pythonlib JSON, invalidated a
665 MB cached browser and bought a full rebuild of a browser whose sources had
not moved. Measured on this branch -- the only browser-input file changed across
the last five pushes was that script.
NON_NATIVE_SCRIPTS names the exceptions. Nothing is excluded for looking
unrelated: excluding a script a build DOES run is the dangerous direction, since
the cache would then serve a browser built from different sources while every
suite downstream passed against it. So ci/tests/test_ci.py checks each entry
against the files a build enters through (Makefile, multibuild.py, patch.py,
package.py, copy-additions.sh, _mixin.py) and fails if one is reachable, and a
second test fails if an entry no longer exists -- a stale exclusion stops
excluding anything, quietly.
Verified against the real cache: a clean checkout of 77edaeb hashes to
fa62d4be1bdff96eee6e9018552255c4, which is the browser cached for this pull
request, and a clean checkout of HEAD with this change hashes to the same value.
So the next run restores that browser instead of compiling one.
This does not change `browser_changed`, which asks a different question (build
versus fetch the published release, against the pull request's base) and is
correctly true for every push to a branch that has touched the browser once.
169 ci tests, 278 pythonlib tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(pythonlib): keep the window's inner dimensions real, and its chrome honest
Three patch guards failed after the fpgen switch -- humanize-edge-deadlock timed
out at the full 600s, humanize-mouse-trajectory saw only its first endpoint, and
mouse-boundary-sweep lost 15 of 25 ring points. All three are geometry, and both
causes were mine. The same guards pass on the pre-fpgen pythonlib with the same
binary, which is how the two were separated from a browser fault.
**Inner dimensions must stay real.** Playwright's setViewportSize resizes the
window and then waits for the page to report the size it asked for. A spoofed
innerWidth/innerHeight never changes, so that wait never returns -- the trap
no_viewport already exists for (#666), reached here through an explicit
set_viewport_size() call. BrowserForge hid it by accident: its Firefox samples
carry innerWidth/innerHeight as 0, and _cast_to_properties skips falsy values,
so they were never spoofed. fpgen reports the real numbers, and mapping them
turned an unmapped field into a hang. They are no longer mapped, which also
means the content area a page measures is the one it actually has.
**The claimed chrome height cannot be smaller than the real one.** The window is
sized from window.outerHeight, so the content area that can receive input is
outerHeight minus the browser's own 86px of chrome. An identity claiming
`outerHeight - innerHeight` below that claims a viewport taller than the window
can hold, and the difference is dead: mouse events dispatched into those rows
reach nothing. Measured -- a drawn pair of outer 801 / inner 717 (chrome 84)
left the bottom 2px unreachable, which is exactly what the boundary sweep saw,
always on the bottom two rows. browserforge's pairs were never below 86; fpgen's
sometimes are. New coherence rule, repaired by growing the window where the
screen has room and shrinking the viewport where it does not.
humanize-mouse-trajectory also had a latent bug of its own: it moved to a flat
(1100, 650), which silently tested nothing whenever the drawn window was smaller
-- the move landed outside the content area and the run failed reporting only
the start point. One drawn window was 924x1364, narrower than that x. It now
derives the destination from the viewport it actually got.
22/22 patch guards pass locally (5m37s); 279 pythonlib tests pass, including a
regression test that no generated config carries window.innerWidth/innerHeight.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(fonts): draw a real OS-version base, at measured real-world rates
The font draw modelled a machine as "always-present core + a flat 30-78%
sample of everything else". Neither half held up:
* the sample was UNIFORM, so every addition was equally likely. Office
(on ~60% of real Windows machines) and the Pan-European supplemental
pack (~2.8%) were drawn at the same rate, and `kind`, `prob`,
`requiresLocale` and `sizes` from scripts/data/font-manifests.json were
thrown away when font-groups.json was written by hand.
* `_ESSENTIAL_FONTS_MACOS` was a SUPERSET of one base (553 entries), which
silently forced 131 Sonoma-only families onto every macOS identity. It
has to be the INTERSECTION of that OS's bases, or adding a second base
does nothing.
A machine now draws one OS-version base, whole and never subsetted, then
each addition unit independently at its own probability.
The bases are measured from real sources rather than inherited:
Ubuntu 24.04 / 26.04 official desktop ISO manifest -> the shipped .debs
-> fc-scan. Firefox enumerates via fontconfig on
Linux, NOT nameID 1: the same file reports
"Noto Sans MeeteiMayek" in nameID 1 and
"Noto Sans Meetei Mayek" in fontconfig. The old base
named 26 families Ubuntu does not ship and missed 41
it does.
Windows 11 26200.9457 a real box. Office is installed there, so OS-native
files are separated by WinSxS hardlink (an OS font
has one, an Office font does not): 143 of 340 files.
Scanning English-only drops the 17 localized CJK
names, so all language ids are kept.
macOS 26.6.2 / 27.0 a real Mac, verified clean. They differ by three
families: 27 drops Noto Sans Brahmi and
Noto Sans CanAborig and adds Noto Sans Sunuwar.
macOS Sonoma NOT re-verifiable (that machine has since been
upgraded). Six family names were stored mojibaked
(Shift-JIS read as UTF-16BE) and are repaired here.
Windows 10 is dropped (end of support Oct 2025), so win11 carries weight 1.0
and the Win11 families are part of the base; the NATIVE path now subtracts
them on a Windows 10 host instead of adding them on a Windows 11 one.
Apple ships 189 families as optional Font Book downloads rather than enabled
by default. The fpgen corpus independently puts all 57 reportable ones at
exactly 81.8%, so they are an `apple-optional` unit, not base. Dot-prefixed
macOS families are stripped: CoreText excludes them from enumeration.
Adobe's fonts are removed from the bundle (185 files, 240 MB): Adobe permits
no redistribution, and the corpus puts every Adobe family on under 2% of real
machines, so they bought no realism. The `adobe-cc` unit goes with them,
because reporting a family the bundle cannot render is a reverse leak.
font-groups.json and the new font-bases.json are now GENERATED by
scripts/gen-font-groups.py instead of maintained by hand, which is what keeps
the draw's probabilities equal to the manifest's.
tests/test_font_distribution.py asserts the properties verify-fonts.py cannot:
complete-base containment, base weights, per-unit probability, bundle
atomicity, a-la-carte piecemeal sizing, locale gating, and variation. Checked
by mutation: reintroducing the uniform sample fails 3 tests, collapsing
variation fails 14, truncating a base fails 1.
The font bundle itself is deliberately not in this commit: it is 3.7 GB and
one file is 184 MB, over GitHub's push limit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor(fonts): store each face once, fetch the bundle as a release asset
The bundle was three per-OS directories, so a face used by more than one OS was
stored more than once: 2955 files, 3.96 GB, of which 1.63 GB (41%) was
byte-identical copies. That was not a decision, it was the absence of one --
the DIRECTORY was the only selection mechanism, so fontconfig could be scoped to
an OS only by giving that OS its own full copy of everything it needed.
Each face is now stored once, in a directory named for the set of OSes that use
it (L, M, W, LM, LW, MW, LMW). An OS reads the four groups its letter appears
in; bundle/fonts/groups.json records the mapping and
utils._generate_fontconfig hands fontconfig exactly those directories.
1513 faces, 2.16 GB
lin -> L+LM+LMW+LW (792) mac -> LM+LMW+M+MW (779) win -> LMW+LW+MW+W (809)
package before after
linux 3.96 GB 2.16 GB (-46%)
macos 2.21 GB 1.31 GB (-41%)
windows 2.69 GB 1.79 GB (-34%)
pythonlib/camoufox/fonts.json is BYTE-IDENTICAL across the change: the same
families are renderable and reportable on all three OSes. The saving is pure
redundancy.
The group directory is also a better gate than what it replaces. The 455
<rejectfont> globs in fontconfig/windows/fonts.conf are gone: a face Windows
must not see is simply not in a group Windows reads (verified: 0 of the 243
formerly rejected basenames appear in any Windows group). Those globs were
unsound anyway -- 114 bundled filenames contain [ ] (e.g. ReemKufi[wght].ttf),
which fontconfig parses as a character class, so they silently matched nothing.
macOS (CoreText) and Windows (DirectWrite) cannot read subfolders, so package.py
flattens the groups for those targets and the allowlist gates them, as before.
The bundle itself leaves git. It is 2.16 GB extracted and 843 MB as .tar.xz;
GitHub rejects files over 100 MiB, and an xz archive cannot be delta-compressed,
so committing it -- split into ~9 parts -- would append the whole archive to
history on every font change, paid by everyone who clones the repo. It is now a
release asset in its own tag namespace (font-bundle-*, excluded from build.yml
so it does not trigger a browser build), fetched on demand:
make fetch-fonts download + verify make fonts-extract unpack
make fonts-check verify only make fonts-clean drop the unpack
scripts/data/font-bundle.json pins the asset name, size and sha256, so a commit
still names exactly one bundle and a truncated download fails loudly instead of
producing a browser that reports fonts it cannot render. This is the trust model
the build already uses for the Firefox source (`make fetch` pulls a ~500 MB
tarball from archive.mozilla.org), so a clone was never buildable offline.
Only the compressed archive is kept; bundle/fonts/ is extracted on demand and is
gitignored. Round-trip verified byte-identical. package-* now depend on
fonts-extract, and gen-fonts-json.py / verify-fonts.py fail with the fix
("run make fonts-extract") rather than a stack trace.
642 font blobs (974 MB) leave the index. bundle/fonts/000_README.txt moves to
bundle/FONTS-README.txt and cleanfonts.sh to scripts/, since bundle/fonts/ is
now ignored.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci(fonts): fetch the font bundle before packaging
The bundle stopped being repo content in the previous commit, so a tagged build
would package a browser with NO fonts at all while pythonlib/camoufox/fonts.json
still reports 335-567 families per OS -- every one of them a family the browser
cannot render, which is the reverse leak this whole series exists to remove.
`make fonts-extract` downloads and unpacks it (verified against the sha256 in
scripts/data/font-bundle.json), and verify-fonts.py then asserts the bundle and
the manifest actually agree before anything is packaged. A build that would have
shipped a mismatched font set now fails in CI instead.
Net disk in CI goes DOWN: 3.96 GB of fonts from the checkout becomes a 0.84 GB
archive plus 2.16 GB extracted.
The tests workflow is untouched: it reads the generated JSON, not the bundle.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(fonts): drop the stale README the v1 archive still carries
bundle/fonts/000_README.txt and cleanfonts.sh were inside the bundle when the v1
archive was built, and they moved to bundle/FONTS-README.txt and scripts/ in the
same change that took the bundle out of git. Extracting therefore re-created a
stale copy of a file git now owns -- invisible to git (bundle/fonts/ is ignored)
but contradicting the tracked README, and it would have been folded back in if a
later bundle were rebuilt from an extracted tree.
Prune both after extraction. A future archive will not contain them, at which
point this is a no-op.
Extraction verified reproducible: two consecutive `make fonts-extract` runs
produce a byte-identical 1514-file tree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(fonts): stage the groups, not per-OS dirs that no longer exist
scripts/stage-fonts.sh copied bundle/fonts/{linux,macos,windows} into an
unpackaged build's dist/bin. Neither half of that is true any more: the bundle
stores each face once under a group directory (L, M, W, LM, LW, MW, LMW) and is
a release asset rather than repo content, so on a fresh checkout the source
paths do not exist and bundle/fonts/ itself does not either. Under `set -e` the
cp fails, which fails `make stage-fonts`, which fails the "Package the binary
for the test jobs" step in tests.yml -- so the build job, and with it the
required gate, would have gone red on every pull request. Locally it broke the
patch guards and build-tester the same way.
Stage the group directories verbatim instead, plus groups.json, which is what
utils._generate_fontconfig reads to decide which of them a claimed OS may see.
Staging the groups rather than a flattened copy is what makes the per-OS gate
work against an unpackaged build exactly as it does in a package. groups.json is
copied last, so an interrupted copy leaves no marker and the next run stages
again instead of trusting a partial tree; the group dirs are copied by name so
fetch-fonts.py's .bundle-sha256 bookkeeping file stays out of a browser's font
directory.
`stage-fonts` now depends on `fonts-extract`, since a fresh checkout has nothing
to stage from. To keep that cheap enough to run before every launch,
fetch-fonts.py stamps the unpacked tree with the archive's sha256 and `--extract`
returns immediately when it matches: 30s to 0.03s, and it no longer needs the
843 MB archive to still be on disk. A bundle bump changes the sha256, so a stale
tree is replaced rather than trusted.
The artifact check in tests.yml asserted fonts/linux; it now asserts
fonts/groups.json and fonts/LMW.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(fonts): prove no identity can reach another OS's faces
verify-fonts.py handed every OS the bundle ROOT as its fontconfig <dir>.
fontconfig scans <dir> recursively, so all three OSes were being tested against
all 1513 faces -- which is why each reported an identical "fc-list publishes 1307
families". The per-OS group gate, the thing that replaced the 455 Windows reject
globs, was therefore never checked by anything.
Hand each OS the four group directories it actually reads, as
utils._generate_fontconfig does, and then assert the converse of the existing
invariant: ask fontconfig for every file it can reach under that conf and require
all of them to sit inside that OS's own groups. The existing checks only prove an
OS can render what it reports; nothing proved it cannot reach what it must not,
and no reported name would reveal it -- a Windows-only or macOS-only face on the
search path is a glyph-fallback candidate, so an emoji or CJK glyph could resolve
to Segoe UI Emoji or PingFang on a machine claiming Linux.
The gate holds: 810 / 780 / 793 faces reachable for win / mac / lin, all inside
their own groups, and the per-OS family counts are now honestly different (529 /
872 / 583 published against 335 / 567 / 358 reported). Mutation-tested -- putting
the root back makes all three fail with 1513 reachable faces.
Also in build.yml: install dependencies before fetching the bundle, so the 843 MB
download uses aria2c's parallel connections instead of silently falling back to
single-connection curl, and add fontconfig explicitly since the verify step
resolves every reportable family through fc-list/fc-match. The over-100-MiB
report is no longer a warning: it is the settled reason the bundle is a release
asset, not an outstanding problem.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(fonts): describe the bundle that actually ships
docs/FONTS.md still described a world several changes back: one bundle per OS at
bundle/fonts/{linux,windows,macos}, a "target bundle" not yet in git, a Windows 10
base, Sonoma as the only macOS base, a uniform 30-78% draw, the 455 reject globs,
Adobe CC among the additions, and a Linux package that duplicates fonts. Every one
of those is now wrong, which makes the document worse than no document.
Rewritten against the shipped data, with the figures read out of the files rather
than recalled: the release-asset workflow and the sha256 pin, the group layout and
which four groups each OS reads, both invariants (reported ⊆ renderable, and
nothing outside an OS's groups reachable) and which one each package type relies
on, the per-OS base weights and the per-unit bundle / à-la-carte probabilities as
tables, and the resulting draw sizes. The namespace trap that produced two wrong
font lists during this work -- Windows/macOS enumerate nameID 1, Linux enumerates
through fontconfig, and system_profiler/WPF report nameID 16 -- is written down so
it is not rediscovered a third time.
Known residue is now stated rather than implied: the Linux base over-claims
against the corpus because the bundle ships 30-metric-aliases, the macOS base
weights are estimates where the addition probabilities are measured, nothing has
been checked in a running browser, and font-bundle.json points at a release on
this fork, which has to be re-uploaded and re-pinned before upstream can build.
Also: per-context-patches.md described the font staging as OS subdirectories and
named createRuntimeFontconfig(), which does not exist anywhere in the tree (it is
utils._generate_fontconfig); gen-fonts-json.py's header still said it scanned
bundle/fonts/<os> and listed Adobe CC; licences.py pointed at the README's old
path inside the bundle.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(fingerprints): pin fpgen's model, and close the taskbar gap it exposed
fpgen supplies every synthetic fingerprint, and it does not ship its model --
it downloads one on first import, and again whenever the files are over five
weeks old. Four things are wrong with that fetch, all in fpgen/pkgman.py:
* TLS verification is disabled on BOTH the API call and the download
(verify=False), so anyone on the path can serve the model;
* the archive is never checksummed, and goes straight into extractall()
with no path-traversal guard;
* the GitHub API is called unauthenticated, on a rate limit shared by every
job on the runner's IP;
* it takes the FIRST release the API lists. model-4/2025 and model-2/2026
carry an identical created_at (2025-03-22, inherited from the tag's
commit), so the sort ties and resolves to the lower id -- model-4/2025.
The consequence is that every Camoufox generates from an April-2025 corpus and
cannot be talked into anything newer: newest Firefox 137, newest GPU an RTX 40,
no RDNA4. The 2026 model has been sitting unreachable for seven months.
scripts/pin-fpgen-model.py installs the model named by scripts/data/fpgen-model.json
before anything imports fpgen, with verification on and the sha256 checked, and
refuses any member that escapes the data directory. Finding fpgen's data dir
must not import fpgen -- importing is what triggers the download -- so it reads
the module origin via find_spec without executing it. Wired into all seven CI
jobs that install pythonlib.
Pinning to model-2/2026 then failed tests/test_launch_geometry.py about 30% of
the time, on `availHeight < height`. That turned out to be our bug, not the
model's: fix_screen_no_taskbar only fired when avail equalled screen on BOTH
axes, so `availWidth < width, availHeight == height` -- a Windows taskbar
docked left or right -- passed straight through and the identity claimed no
vertical chrome at all. Rare shape, but the 2026 corpus produces it in ~30% of
draws once conditioned on a small display, against under 1% unconditioned.
Trigger on the height alone: a side dock is far rarer than a Mac menu bar, a
bottom taskbar or a Linux panel, so the vertical delta is worth more than the
few genuine side-docked machines it overwrites.
316 passed, 2 skipped with the 2026 model pinned.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(fonts): host the font bundle upstream, not on a contributor's fork
scripts/data/font-bundle.json pointed at a release in JWriter20/camoufox, so
merging this would have left daijro/camoufox fetching a required build input
from a personal fork -- a build that breaks if that fork is renamed, made
private, or has the release deleted, by someone with no obligation to keep it.
The identical asset is now published at daijro/camoufox under the same
font-bundle-v1 tag, so only `repo` and `url` move here; `size` and `sha256` are
byte-for-byte what they were, which is the point -- the pin proves the bytes did
not change when the host did.
The upstream release is deliberately NOT marked latest: v152.0.4-beta.30 holds
that badge, and a build input must not displace the browser download people
actually come for. `font-bundle-*` tags are already excluded from build.yml, so
publishing it triggered no browser build.
Verified against the new host from scratch: local archive moved aside, `make
fetch-fonts` pulled 843 MB from daijro/camoufox, sha256 matched the pin, a
forced re-extract produced 1515 files, verify-fonts.py passed every check
(1513 faces stored once; 810/780/793 reachable, each inside its own groups),
and pythonlib is 316 passed, 2 skipped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(fonts): the bundle is hosted upstream now, not on a fork
Removes the known-residue entry that said font-bundle.json points at a fork --
fb8a297 moved the asset to daijro/camoufox -- and states the rule that replaced
it, so the next person publishing a bundle does not put it back on a personal
account.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(ci): pin fpgen's model for build-tester too
859000d wired scripts/pin-fpgen-model.py into every job matching
`pip install ... -e pythonlib`. build-tester does not match: it installs
pythonlib -- and so fpgen -- through build-tester/requirements.txt, which
carries `-e ../pythonlib`. So it kept fetching the model itself.
Confirmed rather than assumed: in run 36050915401 the pin logged
"OK: pinned fpgen model model-2/2026" in Patch guards and the other wired jobs,
while build-tester logged "Fetching model files from GitHub..." and no pin
output at all. That job was still reaching the network with TLS verification
off, no checksum, and no way to land on anything newer than April 2025 --
which is the job that generates the fingerprints the anti-detect suite grades,
so it is the one that least wants a different corpus than everything else.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(parity): expose web content the same files stock Firefox does
Twelve chrome/resource packages are flagged `contentaccessible=yes`, and
ALLOW_CHROME in caps/nsScriptSecurityManager.cpp checks the TARGET package's
flag, never the source. So any page, from any origin, can load their files as
subresources -- and a file Camoufox adds, drops or renames there is a build tell
readable in three lines with no permission:
const s = document.createElement('script');
s.onload = () => BUILD_IS_NOT_STOCK;
s.src = 'chrome://browser/content/browser-development-helpers.js';
scripts/gen-contentaccessible-manifest.py records the file set of every such
package from a stock release into scripts/data/contentaccessible-manifest.json,
and tests/patches/contentaccessible-parity.py diffs a build against it and then
actually loads a sample from a page, so the static diff cannot pass while the
real surface differs. The guard refuses to compare across Firefox versions;
re-record on an uplift.
Two existing divergences it found:
* aboutDialog.js imported Services, declaring an extra page-visible global --
and resource://gre/modules/Services.jsm no longer exists in Firefox 152, so
the line threw. Services is already a chrome global; AppConstants moves to
importESModule.
* hide-default-browser.patch DELETED the "set as default" blocks from
preferences/main.js. That file is contentaccessible, so a page can load it
and read the order of the globals it declares, and deleting moved that order
away from stock. It now hides the UI instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(parity): keep Firefox's popup blocker, give only the driver a key
`dom.disable_open_during_load` is ON in stock Firefox, so a page that calls
window.open() without user activation gets null. Playwright and geckodriver both
ship it OFF, and Camoufox inherited that. It is a one-bit tell any page can read
with no permission and nothing to wait for:
const w = window.open('', '_blank');
if (w) { w.close(); /* not a stock browser */ }
The pref goes back to Firefox's default. To keep
`page.evaluate(() => window.open(...))` working, popup-blocker-parity.patch adds
nsIDocShell.driverPopupsAllowed, and Runtime.js lifts the blocker for that
docShell only while a juggler Runtime.evaluate / callFunction is on the stack.
Unlike upstream's setHandlingUserInput() it grants NO user-gesture activation,
so navigator.userActivation and the autoplay policy are untouched -- trading one
tell for another would be no gain.
The window is SYNCHRONOUS on purpose. Holding it across the promise an async
evaluate returns would leave the blocker open for that evaluate's whole life,
and a page polling window.open() on a timer would eventually land inside it --
silent on stock, one popup here. So a popup opened after an `await` inside the
evaluated function is blocked, exactly as on stock without activation. The depth
is counted rather than a boolean: one docShell carries both the isolated and the
main world, and interleaved evaluates must not clear each other's flag.
tests/patches/popup-blocker-parity.py checks both directions, because either
regressing is silent -- losing the driver's key breaks ~35 upstream tests, and
losing the blocker puts the tell back.
One upstream test is skipped, with the reason recorded: it navigates to
/popup/window-open.html, which calls window.open() from the PAGE's own script at
load with no gesture. Stock returns null there, so no page event arrives.
Upstream can rely on it only because Playwright's Firefox ships the blocker off.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(package): ship glxtest and vaapitest, so graphics decisions match stock
Firefox runs two helper binaries at startup -- glxtest (GL/EGL) and vaapitest
(video decode) -- and feeds what they report into nsIGfxInfo. scripts/package.py
listed both under UNNEEDED_PATHS and dropped them to save ~50 KB. Without them
nsIGfxInfo has no data and the driver blocklist refuses EVERY WebGL context:
WebglAllowWindowsNativeGl:false restricts context creation on this system.
Exhausted GL driver options. (FEATURE_FAILURE_WEBGL_EXHAUSTED_DRIVERS)
Measured 2026-09-18 on one machine: stock Firefox 152.0.4 returned a full WebGL
2.0 context from the real GPU; camoufox returned null from
canvas.getContext('webgl'). The launcher hid it with webgl.force-enabled, so the
breakage only surfaced on a launch that did not go through pythonlib -- and
forcing a context is not the same decision stock makes, which is the point of
shipping the probes instead.
tests/patches/gfx-probes-packaged.py checks three things, because each can pass
while the others are broken: the packaging list must not name them, the build
must have them beside the binary, and a bare launch -- no pythonlib, so no
webgl.force-enabled -- must still get a context.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(pythonlib): derive the storage quota from the disk, default media features to stock
Two page-readable values that were constants where stock's are not.
navigator.storage.estimate().quota is not a fixed number: Gecko derives it from
the disk. GetTemporaryStorageLimit() (dom/quota/ActorsParent.cpp) takes
nsIFile::GetDiskCapacity() of the storage directory and halves it. camoufox.cfg
pinned dom.quotaManager.temporaryStorage.fixedLimit to 52428800, so every
Camoufox reported the same 50 MB where a real browser reports half its disk --
identical across every install, and wrong on all of them. The pref goes; the
launcher computes the limit from the host's own disk instead.
The media features are the same problem from the other end. Playwright sets
color_scheme / reduced_motion / forced_colors / contrast on every context, so a
desktop in dark mode reads `(prefers-color-scheme: light)` where stock reads
dark. STOCK_MEDIA_DEFAULTS supplies the no-preference values and
attach_stock_media_defaults wraps context creation to apply them; an explicit
color_scheme= from the caller still wins. The persistent context is created by
the launch itself, so its features come from the launch options rather than from
new_context() -- handled in both async_api and sync_api.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(prefs): stop suppressing speculative loading, and audit the rest
`<link rel=prefetch>` is page-readable: the prefetch is a real request, so it
lands in the document's own PerformanceResourceTiming entries. Measured
2026-09-18 against stock 152.0.4 on the same host -- stock reported the entry
(initiatorType "other", transferSize 200300), camoufox reported none, from three
lines of page JS with no gesture and no permission. network.prefetch-next and
the two dns.disablePrefetch overrides go back to Firefox's defaults;
dns-prefetch is not directly readable but belongs to the same surface.
BFCache stays disabled, now with the reason written down where someone will find
it rather than discovered again. It IS page-visible -- after a scripted
same-origin navigation and history.back(), stock fires pageshow with
persisted=true and does not re-run the document's scripts, while camoufox fires
persisted=false and rebuilds it (LEAKS row 106). Removing it is not a one-line
change: measured 2026-09-18, the document is then cached (SHISTORY logs
`OnPageHide persisted=1`) but the restore emits no navigation, no lifecycle event
and no execution context, so page.go_back() times out waiting for "load" and the
next evaluate throws against the stale context. Closing it needs FrameTree and
PageAgent to treat a persisted pageshow as a navigation.
scripts/pref-diff.py is the tool that makes this answerable in general.
camoufox.cfg sets ~375 prefs and nothing said which of them actually deviate
from stock. It reads what a stock Firefox exposes over Marionette and classifies
every override as REDUNDANT (dead weight), DEVIATION (the surface that matters),
NEW, or LOCKED. `--camoufox BIN` adds a live diff against a built browser, which
also catches prefs set by policies.json, by a patch, or at launch -- the ones
reading camoufox.cfg alone will miss.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(sundial): outrun the compositor, not just the clock
The pointer probe only made paced moves, and paced moves look the same on every
browser. What ls-pointer-move-rate actually judges is whether events are queued
and coalesced the way real input is, which only shows when the pointer outruns
the compositor -- so follow the paced pass with a 60-move burst that has no
pause between moves.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(window): report the real outer size, and undo a startup auto-maximize
window.outerWidth/outerHeight were returned from MaskConfig unconditionally
while innerWidth/innerHeight report the real content area. Whenever the real
window ended up a different size than the drawn one, the pair contradicted
itself and pages could see innerWidth > outerWidth (and > screen.width):
- GNOME mutter auto-maximizes a new window covering >= 80% of the work area,
and clamps one larger than it. Measured on a 1920x1080 display with a
1853x1048 work area: a drawn 1680x1010 window reported outer 1680 with
inner 1853; every bad reading fit the 80% rule (1707x912 = 1.557M px vs a
1.553M threshold).
- An explicit Playwright viewport makes Juggler resize the real window while
outer stayed fixed: headless viewport 1900x1000 on a drawn 1536x824 window
reported inner > outer; 1280x720 reported 256 px of phantom side chrome.
The outer getters now report the real window, which browser-init.js already
resizes to the drawn size, and browser-init.js restores a maximize the WM
applies in the first 5 s and re-asserts the drawn size. Prototyped in an
omni.ja-patched beta.30: bad readings went from 3/24 launches to 0/24.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(guards): let the contentaccessible guard read an unpackaged build
The guard required browser/omni.ja and omni.ja beside the binary. CI tests the
UNPACKAGED dist/bin -- tests.yml says so where it packages the artifact ("there
is no omni.ja to rebuild there"; Juggler ships as loose files under
chrome/juggler/) -- so the guard could never run on the builds it exists to gate.
It failed the gate in run 36066726109 with "browser/omni.ja not found next to the
binary", after its four live page-load checks had already passed.
`mach build` leaves those resources as loose files at exactly the paths they
would occupy inside the jar, under dist/bin/browser/ for browser/omni.ja and
dist/bin/ for omni.ja. package_listing() now reads the jar when it is there and
walks the tree when it is not, and only fails when neither exists -- so a
packaged build and an objdir are both comparable, rather than one of them being
untestable.
Verified against the local unpackaged tree: 5148 and 8310 entries read, all 12
contentaccessible packages compared, 0 differences from the recorded stock
manifest.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: a repo-wide test pipeline, and the suites Camoufox was missing
Nothing checked a pull request before this. `build.yml` runs on tags and takes
about forty minutes, and `lint.yml` ran a single static script, so a change
could reach main having had no browser suite run against it at all.
This adds one pipeline, driven identically from a pull request, a push to main,
and -- through `workflow_call` -- any caller that needs to test a specific
browser version, so there is exactly one definition of "the tests pass".
resolve ──┬─ static ────────── tribal rules, skiplist, self-tests
├─ pythonlib ─────── the package's own tests
└─ build ──┬─ playwright upstream × 6 shards (conformance)
├─ playwright vendored (regression)
├─ native ───────────── leaks, contexts
├─ patch guards ─────── one per spoofing patch
├─ build-tester ─────── 8 fingerprint profiles
└─ sundial ──────────── stealth grade (off, see below)
│
summary ──► one comment on the PR
Two Playwright suites, because they answer different questions. `tests/` is a
frozen ~v1.55-era fork carrying roughly 1800 lines of Camoufox adaptations, so
every test in it has a known prior outcome: that is the regression check. The
upstream suite is fetched fresh at the tag `ci/versions.py` resolves and runs
unmodified, which is the conformance check -- `ci/pw_camoufox_plugin.py` adapts
the environment around it rather than editing it, hooking BrowserType at the
_impl layer so upstream can refactor its fixtures freely.
`native-tests/` covers what neither can ask about: that resource cost does not
scale with launch count (the shape an FD or socket leak actually has), that two
contexts in one browser get different fingerprints while two pages in one
context get the same one (get this wrong and per-context injection silently
degrades to process-global, which passes every single-context test there is),
and that decisions already made stay made -- `ci/tribal-rules.yml` lists them
with the issue or PR that settled each.
Cost is tiered so a two-second lint failure never reaches a build, and a
driver-only pull request never builds at all: it fetches the published release
and tests against the build users are actually running, a minute instead of
seventy. Merges gate on one required check, `All tests passed`, so the
branch-protection list does not need editing every time a suite is added or
resharded; `ci/branch-protection.json` holds the settings so they are reviewable
rather than lore.
**The stealth check ships disabled** (`ci/sundial.yml: enabled: false`). It
drives a private detection suite, and the deployment it talks to predates that
suite's score mode; an older one ignores `?score=1` and posts the entire report
-- every vector's id, name, brief, source and value -- to whatever collector
asked. Receiving that on a public runner and discarding it afterwards is not the
same guarantee as never being sent it, so while the flag is false the job is not
scheduled, no credential enters a runner, and `run_sundial.py` refuses a hand-run
too. When it is enabled, `redact()` publishes a grade and counts against a
runtime whitelist and refuses anything that is not already aggregated.
Also included: the fixes these suites exposed on a clean runner -- build-tester
hashing canvas pixels rather than a prefix of the data URL, the virtdisplay
cleanup when Xvfb has already died, a juggler sandbox released on frame destroy
rather than only on navigation, and the pythonlib geometry and version-floor
corrections. `lint.yml` is removed because the static job absorbed its one check.
Verified locally: ci/tests 68 passed, tribal rules 24 passed, pythonlib 209
passed, input-dispatch clean, `ci.versions` resolves 152.0.4/beta.31 against
playwright v1.61.0, and `ci.summarize` folds a run to "all suites passed".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* ci: make result files survive the trip from job to summary
The first full run failed, and the summary could not say why: five suites came
back "required, but produced no result", including two whose jobs had passed.
Three separate plumbing bugs, none of them in a test.
**Hidden files.** `actions/upload-artifact@v4` excludes dotfiles unless told
otherwise, and every result we write lives under `.ci-work`. The jobs whose
`path:` was a list containing a glob uploaded nothing at all -- the Playwright
suites and the leak suite each wrote their evidence and then had it silently
dropped:
evidence -> .../.ci-work/results/playwright_vendored.json (fail, 1203 tests)
##[warning]No files were found with the provided path: .ci-work/results/
.ci-work/junit-*.xml. No artifacts will be uploaded.
**Common root.** Where a list did upload, the second entry moved
upload-artifact's common root from `.ci-work/results/` up to `.ci-work/`, so the
JSON arrived at `results/build_tester.json` instead of the artifact root. The
summary merges every `results-*` into one directory and `load_all()` globs a
single level, so the file was there and invisible. build_tester passed and was
reported missing.
Every `results-*` artifact now uploads exactly `.ci-work/results/`, with
diagnostics (junit XML, the build-tester graded tree) split into their own
`diagnostics-*` artifacts that the summary's `results-*` pattern ignores.
`include-hidden-files: true` everywhere that touches `.ci-work`.
**A required name nothing writes.** `static` was in the required list, but it is
a job, not a suite -- no runner writes a result by that name, so summarize
reported it missing on every run including a wholly green one. The suites that
job runs are the pipeline self-tests, which write no result, and native_rules,
which is required by name. The job is already covered: the gate fails on any job
that is not success.
Three guards, each verified by reintroducing the bug it catches:
- results-* artifacts upload exactly one path, so nothing nests
- anything touching .ci-work sets include-hidden-files
- every required name is one some runner can actually write
This changes no test. The real failures the first run found -- 6 in the vendored
suite, plus upstream shards 1 and 5 and the leak suite -- were masked by the
above and should now be reported rather than swallowed.
ci/tests 71 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* test: two failures that were the tests' fault, not the browser's
**The leak check waited on the wrong set of processes.**
`test_a_single_launch_leaves_nothing` failed with Gecko's GPU probe still alive:
1 process(es) this test started are still alive: glxtest(2887, now ppid=1)
`settle()` polled `children(recursive=True)`, but `survivors()` judges the
sampled PID set -- deliberately, so that a process reparented to init cannot
hide a leak. Those two sets differ exactly when a process outlives its parent:
it stops being our child, `settle()` sees nothing left and returns at once, and
anything still winding down is reported as leaked. `glxtest` does this on every
launch; it is spawned by Gecko, its parent exits first, and it needs a moment.
So settle on the set the assertion actually uses. This is a grace period, not an
exemption -- a process that is still there when the timeout expires fails the
test exactly as before, and no name is special-cased.
**Playwright renamed a protocol method the tracing tests spelled out.**
`Page.waitForEventInfo` is `Page.__waitInfo__` in newer versions, so two tracing
assertions failed on a name, not on behaviour. The suite is pinned to a range
(`playwright<1.63`), not a version, so hard-coding either spelling is wrong.
Normalised in `get_trace_actions()`, next to the comment about the last time
Playwright moved this data -- the tests care which actions ran and in what
order, not what Playwright calls them this month.
Neither of these was Camoufox misbehaving.
Still failing, and genuinely about the browser or by design -- triaged next:
navigation popup load state, locator handler visibility, clock pause off by 1ms,
websocket close reason, and the three upstream ones (request headers, worker
locale, screencast viewport) which all look like deliberate spoofing divergence
and probably belong in the skiplist with a stated reason.
ci/tests 71 passed, tribal rules 24 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix: the failures the new pipeline found, and the flake that hid them
Eleven gates were red on PR #9. Each one is now either a fixed defect or an
entry that says why the test cannot apply here -- nothing is silenced.
One real browser bug, found by the conformance suite:
The compositor-backed screencast added in fa8a935 reported the *scaled frame*
size as the viewport. Playwright's Firefox delegate maps deviceWidth/Height
straight onto the client-visible viewportWidth/viewportHeight, and every other
backend fills them from the page's viewport -- the native path right below it
sends pageWidth/pageHeight, clamped to the viewport and never scaled. So a
client asking for a 500x400 frame of a 1000x400 page was told the viewport was
500x200, and asking for a frame *larger* than the page reported a viewport
larger than the page. Confirmed against the built binary, fixed, and confirmed
again: 11 screencast and video tests pass, where the requested size no longer
moves the reported viewport at all.
Four stale expectations in the vendored suite. All four pass in upstream's
v1.61 suite against this same binary, which is what identified them as the
suite's problem rather than the browser's:
websocket Firefox no longer collapses a refused handshake to
CLOSE_ABNORMAL; it reports the HTTP status, like the other
engines. The handler also set a settled future twice, which
surfaced as a suite ERROR rather than a failure.
navigation Firefox now reports a window.open('') popup as "complete". The
old assertion also mis-parsed -- the conditional bound to the
whole assert, so the non-Firefox arm compared nothing.
locator expect().to_be_visible() kept re-arming the locator handler the
click was still waiting to see finish, so the check meant to
observe the interstitial kept it alive. One-shot is_visible().
page_clock resume() before reading the clock added the real second spent in
wait_for_timeout back, landing at 1001 -- one millisecond out,
every time.
Two upstream tests that encode a stock-Firefox quirk this fork does not
reproduce, now in the skiplist with the reason:
A worker inheriting the context locale -- upstream expects en-US from a
ru-RU context, citing playwright#38919, because stock Firefox applies the
locale to the page and not to its workers. Camoufox sets it below that layer,
so matching upstream would mean reintroducing a main-thread/worker
disagreement that anything looking in both places gets for free.
"Firefox" in the User-Agent -- the bare binary advertises its own build
token; the Python package replaces it when it injects a fingerprint. The
vendored suite already asserts what this layer can promise.
And three pieces of the harness that were reporting badly:
A profile that asked for llvmpipe is no longer graded as headless. camoufox's
own preset pool ships "llvmpipe, or similar", so when that preset is drawn,
reporting it is the WebGL spoof working -- and grading it a failure made this
gate fail at random depending on which presets the run happened to draw. The
check still fails on a software renderer the profile did not ask for, which
is the case it exists for.
junit ids put a test's class in the path (test_page_clock/TestWhileRunning.py
::test_should_pause), naming a directory that does not exist -- so the id
could not be fed back to pytest and no skiplist entry could match it.
A `test:` skiplist entry was compared for exact equality against a node id
that always ends in [firefox], so every such entry was a silent no-op: the
test went on running and failing while the list read as handled.
Finally, `mach bootstrap` pulls toolchains from Taskcluster, and a connection
reset there failed the whole pull request (run 34673115086). ci.run_prepare
retries the two steps that download things, and only when the failure reads as
transient -- a failed patch hunk still fails on the first try, because retrying
a broken tree only spends a runner to reach the same answer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* docs(ci): say what the stealth gate actually guarantees now
The README described a `ci` role that is not merged into sundial's master and
so is not deployed, and a kill switch that had since been flipped. Both claims
now match what is running.
The substantive change is separating the two halves of the guarantee, because
only one of them is enforced by the server: `guest` is refused the
private-vector bundle by sundial's middleware, and this repository refuses to
process anything that is not a score payload. What is still missing is a server
that refuses the *request* -- which is what the `ci` role adds, and why the
upgrade path is worth keeping written down rather than implied.
Also documents ci.run_prepare, since "the build retries" is the kind of thing
that needs its limits stated: the network, and nothing else.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): authenticate to sundial with either shape of credential
`SUNDIAL_AUTOMATION_KEY` has been two different things over this repository's
life -- an automation key for `/automated?key=`, which is what sundial's
`make pages-automation-keys` mints and what resolves to the `guest` role, and a
password for the login form. They are indistinguishable by inspection, and the
gate only knew how to use it as a password, so a repository holding a key would
have failed to log in with no hint as to why.
Try the key route first, since it needs no username and so nothing has to be
kept in step with whatever `GUEST_USER` was set to, and fall back to the form.
When the secret really is a password the only cost is one extra request that
401s. If both routes fail the error names both attempts and says what the secret
is supposed to be.
Also sends browser headers on every request rather than `User-Agent:
camoufox-harness`. Cloudflare sits in front of this host and refuses a document
request from a non-browser agent before it reaches sundial at all, which
produces a 403 that looks like a permissions problem and is not one.
test_a_disabled_gate_makes_no_request now makes every route out fatal --
authenticate() and both login functions, not just the one main() used to call --
so the kill switch cannot be bypassed through a path the test does not watch.
Mutation-checked: making the gate ignore `enabled:` fails it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): never let run_prepare invent a success
`--attempts 0` would have fallen straight past the loop bound and returned 0
without running the step at all, which is the one answer it must not invent.
Clamps to one attempt and makes falling out of the loop an assertion rather
than a bare success.
Also corrects a direction in the screencast comment: the native path it
contrasts with sits above _startSnapshotScreencast, not below it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): a sundial category nobody has ruled on now fails the gate
ci/sundial.yml partitions sundial's taxonomy: its nine SECTIONS labels
(Identity, Security, JS Engine, Graphics, Display, Locale, Audio, CPU, Network)
are exactly the six gated plus the three ungated, with nothing left over. That
is now asserted, because it is the property the pass rate depends on and
nothing was checking it.
If sundial grows a tenth section, every check in it used to fold into
"out of scope" -- measured, never gated, and indistinguishable in the summary
from a category someone had deliberately decided not to gate. That answers
"does Camoufox claim this?" by default, in the only direction that never fails
a build: a stealth blind spot that reads as a clean run.
The count is now carried separately and fails the gate, with a note saying to
put the new section in one list or the other. Only ever a count -- which
category, like which vector, does not leave redact().
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): stream the prepare log, and make its timeout real
Two problems with capturing the output of `make dir` / `make mozbootstrap`
instead of streaming it. The visible one: the step printed nothing for minutes
while aria2c pulled a 500MB tarball, which looks exactly like a hung job. The
one that mattered: draining the pipe on the calling thread blocks in readline
until EOF and only *then* reaches proc.wait(timeout=...), so a step that wedged
without printing anything -- a stalled download, precisely the failure this
module exists for -- would never have been timed out at all.
The reader now runs on its own thread, so output appears as it arrives and the
timeout covers a silent hang. Both are asserted against real processes:
stdout and stderr both survive into the text the transient-classifier reads,
and `sleep 60` under a 2s timeout dies in 2s with exit 124 rather than
reporting success.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* docs(ci): two comments that said less than they meant to
The skiplist header claimed to quote failure text it only described, and the
sundial header comment had lost the word that made it a sentence.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* docs(ci): the stealth job comment described the role we do not have yet
It claimed the credential is for a role sundial refuses to serve a report to.
That is the `ci` role, which is not deployed. Says what is actually true of
`guest` instead: the request asks for counts, the role cannot load the private
vectors, and the gate refuses a payload that is not a score.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* test(ci): name what this section actually asserts
The heading said "the score-only `ci` role" and the docstring said the gate
runs as a role sundial refuses to serve a report to. Neither is true yet --
that role is not deployed. What the tests actually pin is that score mode is a
requirement rather than a preference: a full report is refused whoever asked
for it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): the sundial credential must not ride out on an error message
Whatever reaches result.note() is written to the results artifact and posted as
a pull request comment. An exception raised inside urllib can carry the URL that
produced it -- and for the token route that URL *is* the credential,
percent-encoded in the query string. A 401 from a stale key would have published
the key.
Every string built from an exception now goes through scrub() first, which
removes both the raw secret and its percent-encoded form. The message still says
which routes were tried and what the secret is supposed to be, so a real
misconfiguration is still diagnosable from the log alone.
Mutation-checked: with the replacement removed, both new tests fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* docs(ci): say which role the private-vector gate actually blocks
"serves that bundle to nobody else" read as though sundial withheld the private
vectors from everyone. It withholds them from `guest` specifically -- admin and
private still get them -- which is the whole reason CI authenticates as the
least privileged account rather than whichever one was to hand. Verified against
the deployed middleware: guests get a 200 with a no-op body, not a 404.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* docs(ci): the Cloudflare 403 is narrower than the comment claimed
It said any non-browser User-Agent is refused. In fact `POST /__auth/login`
worked for months with `User-Agent: camoufox-harness` -- the block tracks
document-shaped requests, which is what `/automated?key=` is. Browser headers
everywhere are still right, but as "cheaper than remembering which hop is
which", not as a fix for an outage that was never happening.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): verify the sundial role instead of assuming it
The CI log for the first real run says `sundial auth OK (automation key)` --
the stored credential is an automation key, not a form password. Which matters
more than it sounds: sundial's `/automated?key=` resolves to the **private**
role when handed the private key and to `guest` when handed the guest one, and
the two are indistinguishable by looking at them. `private` is served the
private vectors.
So the guarantee this gate has been documenting -- "CI runs as a role sundial
withholds the vectors from" -- was resting on whoever set the secret having
picked the right key. redact() does not help here: it governs what this
repository *publishes*, not what the browser is *given*, and the thing being
prevented is a public runner holding the vectors at all.
The gate now reads sundial's own /__auth/me and refuses to open the browser
unless the session is a role the vectors are withheld from. Not knowing the role
counts as not safe: an absent or unreadable endpoint fails closed, because the
alternative is loading them on the assumption that the credential was right.
Probed live: a bogus cookie yields no role and the check refuses, as it should.
The positive case is what the next run confirms -- and if that credential turns
out to be the private key, this goes red, which is the correct and useful
outcome.
Also corrects the Cloudflare comment, which I had just rewritten on a false
premise. The form-login route was never exercised in CI (the gate shipped
disabled), so "POST /__auth/login worked for months with camoufox-harness" was
unfounded. What is measured: /automated?key= answers 401 with a browser
User-Agent and 403 with urllib's default -- and since that is the route the real
credential uses, the browser headers were load-bearing, not cosmetic.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): record the sundial role in the evidence, not just in the log
`result.metrics["sundial_role"]` was set before the scan and then thrown away:
redact()'s output replaces result.metrics wholesale a few lines later. The check
itself was unaffected -- it ran, and it fails closed -- but the saved artifact
did not say which role had been confirmed, which leaves "the vectors were never
served to this session" unverifiable after the fact. That is most of the reason
to record it.
Caught by reading the artifact the live run actually published, not by the 105
tests, none of which exercised gate() end to end. There is now one that does,
and it fails when the re-assignment is removed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* test: delete the vendored Playwright fork, keep only what is ours
`tests/` held a fork of a ~v1.55-era playwright-python suite, run in CI as
`playwright_vendored` alongside upstream's own suite. Measuring the two against
the same binary says the fork was strictly the weaker of them:
* 73 of the 74 tests it skipped as "Not supported by Camoufox" PASS in
upstream's copy. Those skips predate main-world execution and were never
revisited, so the fork was asserting the browser is worse than it is.
* Of its passing tests, eight had no upstream counterpart. Six were genuinely
Camoufox-specific; the other two were tests upstream had since renamed.
* Everything else was upstream code, one generation stale, run twice.
So the fork goes, and the six tests that were doing real work become
`tests/camoufox/` -- three modules that `ci/suite.py` overlays into the fetched
upstream checkout, where they run against its conftest and its server at
whatever tag was resolved. Nothing frozen, so nothing to go stale.
Two of them take over from skiplist entries that had been pointing at the fork:
* the worker locale. Upstream expects a worker NOT to inherit the context
locale (playwright#38919); Camoufox sets it below that layer, so its workers
agree with the main thread. A page/worker disagreement is a free signal, so
the replacement asserts they match rather than hardcoding one string.
* the User-Agent. The bare binary advertises `Camoufox/<version>`, not
`Firefox/<version>`, and only the Python package rewrites it. The
replacement asserts a well-formed Gecko token and, more usefully, that the
wire and the DOM agree on it.
A skiplist entry that hands its job to another test now says so in a
`replaced-by:` field, and `ci/summarize.py` fails the run if that file does not
exist -- otherwise a rename quietly turns "covered elsewhere" into "not covered".
Also here, because deleting the fork exposed them:
* `tests/local-requirements.txt` was the only thing holding the suite below
pythonlib's `playwright = "<1.63"` ceiling, and it is gone. `ci/versions.py`
now applies that ceiling directly, so a future Firefox bump cannot silently
resolve to a client the shipped package refuses to install.
* the pytest header named `<plugin dir>/skiplist.yml` whatever it had actually
read. Since ci/suite.py copies the plugin into the checkout, that was a path
with no file at the end of it -- misleading precisely when someone is chasing
down a skip. It now reports the file it loaded.
* the summary line printed "Camoufox 152.0.4 ... against Playwright v1.61.0,
which targets Firefox 151.0", which reads like a misconfiguration. Playwright
trails Firefox and skips generations -- it pinned 151 then 153, never 152 --
so it now says which rule picked the tag and that the browser is Firefox 152.
`make tests` runs the one suite. Verified: the three overlaid modules pass
against 152.0.4-beta.31 (8/8), the overlay refuses to shadow an upstream module
and is idempotent across a reused checkout, and both new guards were
mutation-checked. ci/tests: 112 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* feat(ci): a local way to ask where the stealth failures are
"Which five sundial checks failed?" has no answer in CI, and the reason is worth
writing down rather than rediscovering: score mode's payload is buckets keyed
"<Category>|<class>" holding two integers each. It carries no check names and no
ids, so a failing check's identity is not something the gate discards -- it is
something sundial never sends. No artifact, log or sealed report can recover it.
What the payload does know is which category the failures are in, and the gate
was throwing that away too. `--explain` prints a per-category breakdown, and:
* it prints, never records -- `result` is untouched, so the artifact still
carries only `_PUBLISHABLE`;
* it is refused outright under GITHUB_ACTIONS, before anything is sent. A
category table is not a vector, but "Graphics 3/17" is the most useful
single fact an adversary could take from a public log, which is precisely
why redact() does not publish one;
* it says in its own output that names need --allow-full-report and a role
sundial serves full reports to, so a reader does not mistake the category
view for the whole answer;
* an unclassified category is flagged there too, on the same rule the gate
uses -- a new sundial section must not default to ignored.
Also corrects ci/README.md, which claimed "identities in the results file are
HMACs". They are not, and have not been: redact() ships no per-check rows at
all, deliberately, because a map of HMACs still says how many distinct checks
fail and lets a reader follow one across releases. The README was describing a
weaker guarantee than the code actually makes.
The CI refusal is mutation-checked. ci/tests: 116 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): 193 of the 202 tests the skiplist skipped actually pass
The skiplist was wrong, and wrong in the way that matters: it claimed the
browser could not pass tests it passes, and the stated reasons made that look
checked.
Where it came from: the first version of ci/skiplist.yml took all nine
`tests/async/*.disabled` modules from the vendored suite -- a 1:1 match, nothing
independently derived -- and gave each a plausible justification without running
any of them. That is the exact thing the file's own header calls "a regression
wearing a disguise", committed in the commit that wrote the header.
Ran every entry with the skiplist disabled, against 152.0.4-beta.31:
test_element_handle.py 59 skipped 0 fail
test_popup.py 26 skipped 0 fail
test_dispatch_event.py 12 skipped 0 fail
test_check.py 9 skipped 0 fail
test_launcher.py 9 skipped 0 fail
test_focus.py 6 skipped 0 fail
test_fill.py 3 skipped 0 fail
test_click.py 71 skipped 2 fail
client_certificates 5 skipped 5 fail
(two individual tests) 2 skipped 2 fail
Seven of nine modules failed nothing at all. 193 of 202 passed. The suite was
reporting 1339 passing while silently excluding 193 more that also pass -- and
nothing would have caught those 193 regressing.
Some reasons were not just over-broad but wrong. test_popup.py was skipped for
"Camoufox resolves the User-Agent from its fingerprint config, so an arbitrary
override does not and must not stick"; the two UA tests in it pass, because the
bare binary under plain Playwright does honour `user_agent=` -- the fingerprint
resolution is in the pythonlib wrapper, which this suite does not use. The
neighbouring test_request_headers_should_work entry gets that distinction right,
four entries earlier. test_dispatch_event.py was skipped for "Camoufox only
emits trusted events"; that is true and those tests never assert isTrusted.
So: 13 entries down to 5. test_click.py narrows from the module to the two tests
that fail -- Playwright's stable-position wait polls the bounding box then
dispatches instantly, and the humanized path spends real time travelling, so an
animating button is clicked mid-flight (offset 100, expected 300). The
client-certificate module stays whole: all five fail at the TLS layer, which is
support that genuinely is not compiled in. The `[chromium]` and `[webkit]`
patterns are deleted -- the runner pins `--browser firefox`, so they matched 0
tests and only made the list look more considered than it was.
And the guard, because a reason is an assertion about the browser and nothing
was checking it: ci/run_skiplist_audit.py runs every entry with the skiplist
disabled and FAILS THE BUILD if a skipped test passes. It is cheap exactly
because a correct skiplist is short -- 9 tests, 8 seconds -- so it runs on every
PR in tier 3a and is a required gate. Mutation-checked both directions: exit 1
naming the newly-passing tests when a stale entry is re-added, exit 0 now.
ci/tests: 118 passed, including that every shipped entry is actually auditable
(a `pattern` entry cannot be, and now fails that test rather than riding along
unverified).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* feat(ci): run the sync suite too, and say what is left out
The Playwright gate ran `tests/async/` and nothing else. Upstream v1.61 collects
2306 tests; that was 1584 of them. The other 722 -- `tests/sync/` (715),
`tests/common/` (4), `test_reference_count_async.py` (2), `test_installation.py`
(1) -- never ran, and unlike a skiplist entry there was nothing anywhere saying
so. No reason was recorded because there was no decision: the vendored fork in
tests/ carried `async/` and `async_imp/` and no sync suite, and this runner was
pointed at the same shape without checking what upstream had. Same inherited
shape as the stale skiplist, one level up.
TARGETS now names `tests/async/` and `tests/sync/`. The sync API is a greenlet
wrapper over the same Juggler traffic, so much of it duplicates async at the
protocol level; it is here because pythonlib ships a sync API users actually
drive, and the wrapper has its own timeout and reentrancy behaviour the async
tests cannot reach.
ISOLATED_TARGETS runs `tests/common/` and `test_reference_count_async.py` in a
second pytest process, on the first shard only. They cannot share a process with
the others: each calls sync_playwright()/async_playwright() inside the test body,
which cannot start while the session fixtures hold a loop. Together all six fail
with "Cannot run the event loop while another loop is running"; alone all six
pass. They are worth the extra invocation rather than dropping, because
ProtocolCallback objects accumulate when the browser never replies to a protocol
message -- and this fork patches Juggler heavily, so that leak can be ours.
EXCLUDED holds the one real exclusion with its reason: test_installation.py
pip-installs playwright to check packaging, which exercises Playwright's release
process and not this browser. A self-test requires every exclusion to carry a
reason.
And unclaimed() closes the level above the skiplist: if upstream adds a test path
that is in neither set, the run errors instead of quietly getting narrower. It
ignores assets/ and golden-*/ fixtures, and a self-test proves it catches a new
tests/integration/.
Verified against the CI-built binary: 2225 passed, 0 failed, 66 skipped, 2291
collected. ci/tests: 122 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): client certs are only unsupported where the BROWSER does the handshake
The skiplist audit failed its first CI run, on an entry written one commit
earlier. It was right to.
`tests/async/test_browsercontext_client_certificates.py` was skipped as a whole
module, reasoned as "client-certificate support is not built into the Camoufox
binary". On the GitHub runner two of its five tests pass:
test_should_throw_with_untrusted_client_certs PASSES
test_should_work_with_global_api_request_context PASSES
test_should_work_with_new_context fails
test_should_work_with_new_context_passing_as_content fails
test_should_work_with_new_persistent_context fails
The split is not noise. Playwright offers client certificates two ways:
playwright.request.new_context(client_certificates=...) the Node driver does
the TLS handshake itself -- works, and those are the two that pass.
browser.new_context(client_certificates=...) the BROWSER does the
handshake -- and this build has nothing to do it with.
So the reason was over-broad rather than wrong, and a module entry was exactly
the shape that hid the distinction. Now six per-test entries, async and sync,
naming the browser-side handshake specifically.
Worth recording how the bad entry got written, because the mechanism matters
more than the entry: locally all five fail, because this machine's Node/OpenSSL
rejects the fixture server outright ("wrong version number"). That looked like
uniform absence of support and it was not. A local run is a hypothesis; CI is
the authority for what fails. The audit is what turned that from an opinion into
a build failure, one run after the mistake.
The sync entries assume the same split from identical test names; the audit will
confirm or correct that on the next run rather than my asserting it from here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): async and sync cannot share a pytest process
Adding tests/sync/ to the same pytest run as tests/async/ produced 50 tests that
passed only on retry. 47 were fixture-setup errors in two fetch modules:
RuntimeError: Runner.run() cannot be called from a running event loop
Upstream's sync suite is a greenlet wrapper; its async suite runs under
pytest-asyncio. In one process, whichever runs second breaks the other's loop.
Measured, against the CI-built binary:
tests/async alone 1526 passed, 1 timing flake
the two fetch modules alone 102 passed
async fetch + one sync module 14 failed
sync module first, then async fetch 46 errors
So sharding was not hiding anything -- async alone is clean -- and the fault is
the mixing, not the size of the run.
What makes this worth a structural fix rather than a retry budget: the damage
lands in async FIXTURE SETUP, so it presents as "the fetch tests are flaky" --
a browser-shaped symptom for a harness-shaped cause -- and the retry pass then
makes it vanish. Left alone, CI goes green with 50 silent retries, and the
natural response to any that stuck would be a skiplist entry recording a browser
failure that does not exist. That is the failure this branch has spent its last
several commits removing.
GROUPS now names three sets, each run in its own process: tests/async/,
tests/sync/, and (tests/common/ + test_reference_count_async.py). The first two
shard; the third is six tests and does not, because splitting it hands some
shard an empty selection and pytest exits 5 for that. The pairing in the third
group is measured, not assumed -- those two run together cleanly (6 passed).
Retries stay inside their own group, for the reason the groups exist.
Verified: 2225 passed, 4 failed, 66 skipped, 2295 collected, and 1 retry-passer
instead of 50. All four failures are the two client-cert tests that go through
playwright.request.new_context(), which fail only on this machine -- its
Node/OpenSSL rejects the fixture server ("wrong version number"). They pass on
the runner, which is why they are not skiplisted; the audit is what will hold
that claim honest. Expect 2229/0 in CI.
ci/tests: 124 passed, three of them pinning this shape -- async and sync in
different groups, no target in two groups, and the small group unsharded.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): refuse to test a Firefox the branch does not pin
Suite selection follows the browser version correctly: ci/versions.py resolves
Firefox 155 to Playwright v1.62.0 and 152 to v1.61.0, and the workflow threads
`browser_version` from workflow_call/dispatch through resolve into the Playwright
job, which fetches that tag. That half works.
The other half does not. Nothing else honours the input:
ci.run_build / make read upstream.sh
fetch-browser `python -m camoufox fetch`, whatever is current
ccache key uses the resolved version (cosmetic)
So `browser_version: 155.0` on a branch pinning 152.0.4 compiles 152 and judges
it against the suite chosen for 155. It passes, it means nothing, and no output
anywhere says the browser and the suite are describing different releases. For a
gate whose entire job is to make an upgrade provable, that is the worst
available outcome.
The legitimate flow never had this in it: an upgrade to a new Firefox is a
branch that edits upstream.sh -- that is what an upgrade is -- and resolution
then reads it by default, so the build, the suite and the input cannot disagree.
The mismatch only arises from a dispatch that asks for a version the branch does
not pin.
`--check-upstream` refuses that combination with a message saying what would
have happened and what to do instead (bump upstream.sh). The workflow passes it,
and a self-test asserts the workflow passes it, because a guard nothing invokes
is decoration. Mutation-checked.
Not fixed here, and deliberately: making the build honour an arbitrary version
would mean synthesising an upstream.sh -- version plus release tag plus
closedsrc_rev -- for a release that may not exist. Refusing the contradiction is
the honest amount of machinery for it.
ci/tests: 126 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): the fetched browser must be the generation the suite was chosen for
A driver-only pull request does not build. It downloads the current release,
which is the right browser to judge a driver change against -- it is what users
run. But the Playwright suite is chosen from upstream.sh, and those two agree
only until an upgrade window opens.
The sequence that breaks: the version agent bumps upstream.sh to Firefox 155 and
that merges. No build of 155 is published yet. The next driver-only PR resolves
its suite for 155, fetches the published 152, and tests the old browser against
the new suite. It passes, and says nothing.
Today they happen to agree exactly -- upstream.sh pins 152.0.4-beta.31 and the
latest release is v152.0.4-beta.31 -- which is timing, not a guarantee, and
precisely the kind of coincidence that hides this until the upgrade it is
supposed to protect.
So the fetch step now reports which build it installed and refuses a mismatch.
Beta drift inside a generation is fine and expected: beta.30 against beta.31
does not change which Playwright tag is right, and demanding an exact match
would fail every run between a bump and a release. A generation apart is not
fine, and that is what is checked.
Fails closed on input it cannot read, rather than passing by accident.
Mutation-checked. A self-test asserts the workflow actually invokes it.
ci/tests: 128 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): a driver-only pull request cannot satisfy a required build
The summary was told to require the `build` suite on every run:
required="pythonlib native_rules"
if [ "$browser_changed" != "skip" ]; then
required="$required build patch_guards ..."
fi
`browser_changed` is only ever `true` or `false` -- nothing emits `skip` -- so
the branch is always taken and `build` is always required. But the build job is
deliberately skipped whenever the browser was fetched rather than compiled, and
only that job writes a `build` result. summarize then reports
`build` is required but produced no result file. A suite that did not run
has not passed.
which is the correct rule applied to a suite that was never supposed to run,
and the gate goes red.
That blocked every pull request touching only pythonlib/, ci/, tests/ or the
documentation -- most of them, and precisely the cheap path ci/README.md
advertises as "driver-only pull requests never build". It was never seen
because this branch edits the Makefile and additions/, so its own runs always
took the build path.
The browser suites stay required either way: fetched or built, the browser is
there and they run against it. `build` is the only one that follows.
`test_build_is_required_only_when_the_browser_was_built` extracts the workflow's
own `required=` assembly and runs it under bash for both values, rather than
pattern-matching the shell -- the bug was a comparison that read as deliberate,
and only running it says what it does. Mutation-checked: restoring `!= "skip"`
fails it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): sundial being down is a skip, not a verdict on the browser
The stealth check drives a service on another host. Any failure reaching it --
DNS, a refused connection, the edge answering 502 for a minute -- came back as
ERROR, failed the job, and failed the merge gate. So an outage over there
blocked every pull request in this repository, including ones that had nothing
to do with stealth, and the only fix available to a contributor was to wait.
That is not a statement about the browser. When sundial cannot be reached the
browser was not measured at all, so neither a pass nor a failure is true. It is
now recorded as SKIP with the reason attached, the job exits 0, and
`ci/summarize.py --allow-skip sundial` tolerates it: shown on the summary table
with its own icon and its reason, and not a merge block.
The line is drawn at whether sundial answered:
down no HTTP reply at all (URLError, timeout, connection reset), or a
5xx, or a 429 -- the origin is broken or is refusing everyone.
answered everything else. 401 is a bad credential, 403 is the edge refusing
a non-browser request, and both are this repository's problem to
fix. Skipping past those would turn a misconfigured stealth gate
into a permanently green one.
And an answer stays an answer further in: a role sundial would serve the
private vectors to, a full report where a score was requested, a pass rate under
the floor -- all still fail, as before.
Both auth routes are tried, and the key route 401s whenever the stored secret is
a password rather than an automation key. One real reply is enough to know
sundial is up, so `authenticate()` reports an outage only when neither route got
an answer.
The skip note goes through `scrub()` like every other published string: the
token route puts the credential in the URL, and a URLError carries the URL that
raised it.
`--allow-skip` is per suite and nothing else is on the list. Every other suite
runs on the runner; none of them has this excuse.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* feat(ci): run the suite isolated first, and count what needs the main world
The conformance suite forced `disableWorldIsolation` for every run. That made
upstream's tests pass -- they assert upstream semantics, reading globals their
own page scripts defined -- but it measured a mode nobody ships, and it produced
no number at all for what isolation costs. Isolation is the reason this fork
exists; running 2229 tests against it turned off says less than it looks like.
Each group is now run up to three times:
1. isolated -- the configuration users get. A test reading a page-defined
global fails here by design.
2. the same failures again, still isolated. A pass is a flake, and a flake
must not be counted as a world difference.
3. what is still failing, with isolation off.
A test that passes in 3 is a **main-world fallback**: it counts as a pass -- the
browser does honour the contract -- and its identity is recorded in
`metrics.main_world_fallbacks`, with the count on the summary table. That count
is the isolated-world conformance gap. It is invisible in the pass/fail totals
by construction, which is exactly why it has to be printed: a jump in it means
the isolation boundary moved, and nothing else in this pipeline would say so.
A test failing in *both* worlds is a plain failure, as before.
`CI_WORLD` selects the world and defaults to isolated, so a plain
`pytest -p pw_camoufox_plugin` by hand measures the browser as it ships.
`_apply_world()` *clears* a stale `disableWorldIsolation` as well as setting it:
the passes are separate processes inheriting one job environment, and a
leftover flag would make the isolated pass quietly measure the main world --
which would silently zero the very number this is for.
Two things this depends on, fixed here:
Each group gets its own pytest cache. All three run in one checkout, and
pytest only drops a `lastfailed` entry when that test is collected again and
passes -- so with the shared cache the async group's rerun was selecting from
a set the sync group had also written into. Depending on which groups had
failed that meant re-running the whole group or selecting nothing at all.
Per-group, `--last-failed` means what it says, and passes 2 and 3 can use it
to name exactly the right tests.
`ci/run_skiplist_audit.py` pins the MAIN world. The suite now counts a test
needing the main world as a fallback rather than a failure, so a skiplist
entry has to claim the test cannot pass in either world -- auditing under
isolation would let an entry justify itself with a failure the suite would
never have counted, which is the same class of untrue-but-plausible reason
the audit exists to catch.
Shard merging sums the fallback counts and unions the identity lists; taking
the first shard's, as the generic metric merge did, would report a sixth of the
number. Mutation-checked, along with the stale-flag clear.
The first CI run on this is what establishes the real fallback count. The
plugin's own note put it at roughly 37; that was measured a while ago and is not
a promise.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* docs: describe the pipeline that exists, not the one that is coming
CONTRIBUTING.md and the pull request template still described the process this
branch replaced: run both suites by hand, screenshot the output, paste it in.
CLAUDE.md was already pointing at CONTRIBUTING.md for "the full pipeline" that
CONTRIBUTING.md did not mention. Since the premise of this work is that the rule
in CONTRIBUTING.md was never enforced, leaving it unchanged left the rule
describing the wrong thing.
Both now say what CI does and what the one required check is. The template no
longer asks for a screenshot: nothing checked that the browser in one was built
from the branch under review, which is the whole reason the pipeline exists, and
CI leaves its own report as a comment. The local commands stay, because running
build-tester by hand while working on a spoofing patch is still the fastest way
to find out whether it did what you meant.
Several comments described this repository as containing an auto-update harness
that is not in it: `ci/results.py` credited `verify.py` and "the repair agent"
with computing the verdict that `ci/summarize.py` computes, and `ci/sundial.yml`
and `ci/build-tester.yml` sited themselves relative to a `harness/policy.yml`
nobody can open. Each now names what actually decides here and marks the harness
as the out-of-repository caller it is. `run_sundial.py waive` says outright that
the file it prints a stanza for is not in this repository, which is worth
knowing before going to look for it.
Also: PEP 8 blank lines around `resolve_verstr()`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* perf(ci): drop the same-world retry -- 7m50s a shard, nothing recovered
The first real run of the isolated-first suite took 21m13s a shard against
4m34s before. Most of that was a pass that could not have worked. Shard 4, both
sharded groups:
isolated (full) 335s + 331s 35 and 11 failures
isolated retry 205s + 265s 0 recovered
main world 19s + 12s 46 recovered
The retry was there so that a flake could not be mistaken for a world
difference. It cannot do that job:
* These failures are deterministic. A test reading a global its own page
script defined does not intermittently stop seeing it.
* Failing that way is slow. The read returns undefined rather than throwing,
so the test sits on a Playwright timeout -- which is why re-running 46
known failures cost nearly eight minutes while proving them in the other
world cost thirty-one seconds.
* Upstream already reruns. The `105 rerun` on that first line is every one of
those 35 failures having been retried three times by the suite's own
pytest-rerunfailures before the run reported them. A flake does not survive
that, so there was nothing left for a fourth and fifth attempt to find.
So: isolated, then the main world, then -- only for what failed in BOTH -- one
retry. That last set is normally empty, so the retry is free on a healthy run
and still answers the one open question on an unhealthy one: a test no world
satisfies is either broken or flaky. It runs in the main world, where a pass
means "not reproducible" rather than "needed isolation off", which is already
known by then.
A flake surviving upstream's three reruns and then passing in the main world
would now be counted as a fallback rather than as a flake. That is the trade,
and it is worth it: the count is reported, not gated, and no verdict moves.
Also adds the guard that was missing between the phases. `--last-failed` with
nothing previously failed does not select nothing -- pytest declines to filter,
and runs the whole group. Unguarded, a group that passed cleanly under
isolation would have been re-run end to end in the main world, silently
replacing the result it was meant to refine. Both rerun passes are now behind a
non-empty check, and a self-test asserts it of every `--last-failed` in the
loop. Mutation-checked, as is the single-isolated-pass rule -- "retry it in the
same world first, just to be safe" reads as obviously correct and costs eight
minutes a shard.
Expected shard time is now ~11 min: the isolated pass, which is the
measurement, plus half a minute to resolve it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* perf(ci): bound the isolated pass, which is where the timeouts come from
Some isolated failures do not fail. They hang for the full 180s per-test
timeout and are then rerun three more times, so one test can cost twelve
minutes. The main-world baseline had zero timeouts across all six shards, so
this arrived with isolation.
Reproduced against a build, and the cause is not the harness:
page script calls window.exposedFn() isolated -> HANG main -> resolves
evaluate() calls window.exposedFn() isolated -> resolves main -> resolves
`expose_function` installs its binding on the isolated world's global. Page
script calling `window.fn()` looks at the page's own window, does not find it,
and the call never reaches Python -- so a test awaiting the future that call was
meant to resolve waits forever, because that await has no Playwright timeout
behind it. evaluate() works because it runs in the same world as the binding.
That is isolation doing exactly what it is for. A page that can reach an
automation binding can detect it, which is the reason this fork exists. These
tests assert a behaviour Camoufox deliberately does not have, and they cannot be
fixed -- only recognised, which pass 2 does in about half a second each.
What can be fixed is the price of recognising them. Pass 1 is a classifier: its
only question is whether a test passes as Camoufox ships, and a test that hangs
has already answered it. So pass 1, and only pass 1, is bounded:
per-test timeout 90s. The slowest test in the entire main-world baseline was
30.4s of 2295; two exceeded 30s and none exceeded 45s. A Playwright action
times out at 30s. Three times the slowest honest thing that happens, and half
the previous bound.
upstream's reruns off, via CI="". tests/conftest.py sets `reruns = 3`
whenever $CI is set, and that is the only thing it reads $CI for. It is
insurance that almost never pays out -- 2 reruns across all 2295 baseline
tests -- and under isolation it turned every deterministic world difference
into four attempts: 138 reruns in a single shard's isolated pass, recovering
nothing. Note that `--reruns 0` as an argument would not work; conftest
overwrites config.option.reruns in pytest_configure, so the environment is the
only lever that holds.
A hang now costs one 90s wait instead of up to 720s. A flake missed by not
rerunning is not lost: it fails pass 1, passes pass 2, and is counted as a
fallback -- noise in a reported metric, not a change in any verdict.
Passes 2 and 3 keep upstream's conditions untouched. They are the ones deciding
what an answer means, and they run against a handful of tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* perf(ci): stop rebuilding a browser that has not changed
Every push to this branch recompiled Firefox, 24 minutes at a time, including
the three in a row that changed nothing but ci/ and documentation.
`browser_changed` is computed against the pull request's BASE, not against the
push, so a branch that touched patches/ or additions/ even once keeps rebuilding
forever. That part is correct and should stay: the published release does not
contain this branch's browser changes, so testing against it would test a
different browser than the one under review. Using the release is not the
alternative.
The alternative is not building the same thing twice. The build job now asks a
narrower question first -- not "does this branch change the browser" but "has
the browser changed since the last one we built" -- and answers it with a cache
keyed on a hash of every input that can alter the binary: the same path list
browser_changed greps for, plus this workflow, which pins the toolchain the
build runs on. On a hit the 634 MB dist is restored and every build step is
skipped, which is the difference between 24 minutes and about one.
Two things that would otherwise have made this quietly wrong:
A hit still has to report a `build` result. It is a required suite whenever
the browser was built rather than fetched, and a required suite that produced
no result is -- correctly -- a failure. Same trap as requiring `build` on a
driver-only pull request, reached from the other side. The hit path writes one
recording that the browser was restored and under which key, so "this run
compiled nothing" is a fact in the evidence rather than an absence in it.
No restore-keys. Everywhere else in this workflow a prefix match is right; a
partly warm ccache is still warm. Here it would hand the test jobs a browser
built from different sources while every suite reported on it looking
perfectly healthy.
The self-tests hold the two lists together: if a path is ever added to the
browser_changed grep without being added to the cache key, a change there would
neither force a build nor invalidate the cache, and the run would silently test
a browser that predates it. Mutation-checked, along with the guards on each
expensive step and the absence of restore-keys.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* perf(ci): a Juggler JavaScript change does not need libxul relinked
Juggler is mostly JavaScript, and JavaScript does not need a compiler. Measured
on a real build, ccache reported a 98.63% hit rate (4167/4225), so almost none
of those 24 minutes was compiling C++ -- it was Rust, linking libxul, and
packaging, none of which a .js file affects.
And there is no omni.ja to rebuild either. CI archives the UNPACKAGED dist/bin;
omni.ja only exists in dist/camoufox/, which `mach package` produces and nothing
here uses. In dist/bin, Juggler is loose files under chrome/juggler/ -- symlinks
into the source tree, dereferenced into the artifact by `tar -ch`. Delivering
new JavaScript is a copy.
So the build cache is now keyed on a hash of the COMPILED inputs only. If that
hash matches, the compiled half is identical by construction and this branch's
resources are laid over the restored browser. The hash is the classification:
there is no "did only JavaScript change?" diff, because a diff answers the wrong
question -- it compares against the pull request's base, while what matters is
whether the cached browser has the same native sources. Same hash, same binary,
whatever the diff says.
Two traps, both closed and both mutation-tested, because either one silently
serves a browser that is not the one under review while every suite reports
green:
additions/juggler/ is not all JavaScript. It also holds the screencast encoder
and the remote-debugging pipe -- 5 .cpp, 5 .h, 2 .idl, 3 components.conf, 4
moz.build -- compiled into libxul. Only the files jar.mn lists are treated as
resources; everything else, including any extension nobody has considered yet,
is native and forces a build. jar.mn is itself native, so a resource removed
from it cannot leave a stale copy behind.
The mapping is per-file, not a prefix. jar.mn maps TargetRegistry.js to
content/TargetRegistry.js (a level added), content/FrameTree.js to
content/content/FrameTree.js (preserved), and content/JugglerFrameChild.sys.mjs
to content/JugglerFrameChild.sys.mjs (dropped). Two files in one source
directory land at different depths. A prefix rule writes one of them to the
wrong path and leaves the old copy in place.
Verified against a real build rather than reasoned about: all 22 jar.mn entries
resolve to files that exist in dist/bin, and overlaying this branch onto a dist
built before it turns FrameTree.js from 0 occurrences of nukeSandbox to 1, with
all 22 resources byte-identical to the branch afterwards.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* docs(ci): name the mechanism that actually hangs, not the one I found first
The comment blamed expose_function. That mechanism is real -- verified against a
build -- but no upstream test has that shape, so it is not what hangs in CI.
Running the full async suite isolated against a build names them: four tests,
always the same four, all in tests/async/test_route_web_socket.py. 225 of 1509
fail under isolation; four of those hang.
The general shape, which is worth stating because it will recur: a Playwright
feature implemented by installing something on the page's global lands in the
isolated world instead, so anything the PAGE originates never reaches the
automation. route_web_socket replaces window.WebSocket from an init script;
isolated, that replacement is in the sandbox and a socket the page's own script
opens is never intercepted. expose_function puts its binding on the sandbox
global, so page script calling window.fn() finds nothing -- called from
evaluate() it works, which is why it does not hang.
They hang rather than fail because the waits involved have no Playwright timeout
behind them: a Twisted future from the test server, an asyncio future a binding
was meant to resolve. Everything else isolation breaks fails at Playwright's 30s.
One thing to flag beyond the test suite: the route_web_socket half is not a test
artifact. Measured with a page whose own script opens a socket -- what a real
site does -- the handler fires in the main world and never fires isolated. A
user calling page.route_web_socket() against a real site gets no interception
and no error. It is not fixable here, because the feature works by replacing a
page global and that is exactly what an isolated world exists to prevent a page
from seeing, but it deserves an issue of its own rather than a comment in a CI
runner.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): the isolated pass cannot bound a hang, so stop asking it to
Four shards ran for two hours each and were killed by timeout-minutes, three
runs in a row. The cause is not the duration of anything.
Measured on run 34799668707, with ISOLATED_TIMEOUT already lowered to 90s:
tests/async/ isolated pass completed in 296s. The bound works.
tests/sync/ test_should_work_with_ws_close printed pytest-timeout's
"+++ Timeout +++" banner at exactly 90s -- and the process
then sat there for the remaining 1h50m.
So the signal fires and the test dies; the process does not. pytest-timeout's
signal method raises at the next bytecode boundary, and Playwright's sync API is
parked in a greenlet switch that never reaches one cleanly, so the raise lands
inside the dispatcher and wedges it. `--timeout-method=thread` fires reliably but
kills the interpreter and takes the other ~1500 tests in the group with it. No
per-test value bounds this, which is why lowering 180 -> 90 changed nothing.
Declared rather than discovered, then. The isolated pass cannot learn that these
hang without hanging, so ISOLATION_HANGS tells it: they are --ignore'd out of
pass 1 and run directly in the main world, where they pass and are counted as
fallbacks exactly as if isolation had failed them honestly. Coverage is not lost
-- the same tests run, in the world that can run them. Verified in the same run:
tests/async/test_route_web_socket.py::test_should_work_without_server
isolated -> Timeout main -> PASSED
Deliberately not ci/skiplist.yml. That list means "fails in the most permissive
world", and run_skiplist_audit.py enforces it by running every entry with
CI_WORLD=main and failing the build on any that pass. These pass there, so an
entry would be rejected by the audit and would be untrue as written. The tests
keep the two lists apart.
The second half is the backstop, because the next unboundable hang will not be
this one. Each pytest invocation was bounded at args.timeout, default 10800s --
three hours, against a job capped at 120 minutes. It could never fire: GitHub
hard-killed the runner first, taking the junit and diagnostics uploads with it.
Now --group-timeout, 1200s, roughly four times the slowest healthy invocation
measured, and the job drops 120 -> 40. A wedge costs twenty minutes and still
reports what it collected, instead of two hours and nothing.
The browser bug underneath is real and is not a test artifact: route_web_socket
replaces window.WebSocket from an init script, which under isolation lands in the
sandbox, so a socket the page's own script opens is never intercepted and the
caller gets no error. Filed as #775, with the native-interception fix that keeps
it undetectable. When that lands these stop hanging and the declaration goes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): a cache-hit build cannot repack the archive over its own input
The prebuilt-browser path failed every time it actually hit:
zstd: camoufox-dist.tar.zst already exists; stdin is an input -
not proceeding.
tar: -: Cannot write: Broken pipe
`zstd -o` refuses an existing destination and exits 1, and the overlay step
repacks to the same filename it just unpacked from. It went unnoticed because
until now every run changed the browser sources and rebuilt instead -- the cache
restored, the overlay ran ("overlaid 22 resource(s)"), and the step died one
line later. The first pull request that did not touch the browser found it.
Repack to a temporary name and mv it into place. That sidesteps the refusal, and
means a repack that dies partway cannot leave a truncated archive where the
restored one was.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
* fix(ci): a six-test module cannot be sharded six ways
The declared-hang pass ran with the shard filter still applied, so a shard that
owned none of the module got:
collected 6 items / 6 deselected / 0 selected
============================ 6 deselected in 0.05s ===========================
pytest exits 5 on an empty selection and writes no junit, which is exactly what
"the module did not run" looks like -- so the guard that exists to catch lost
coverage failed three of six shards instead. The other three owned some of the
six and passed, which is why it looked intermittent.
Run it unsharded on the first shard, the way tests/common/ already is and for
the same reason. Running the module whole also keeps its fallback accounting in
one place rather than spread across shards that each saw a fraction of it.
CI_SHARD is cleared rather than dropped: ci/_util.run() layers env over
os.environ, so an omitted key would still inherit one. parse_shard() reads empty
as "no shard", the same way _NO_UPSTREAM_RERUNS clears $CI.
The rest of the run confirms the mechanism. On shard 3, which completed:
tests/sync/test_route_web_socket.py::test_should_work_with_ws_close PASSED
in 3.07s -- the test that wedged a runner for 1h50m two runs ago -- and the
shard finished 386 passed, 0 failed, 0 errored. No shard hung. The whole run
took 26 minutes against 2h35m, and the build was 53s against 24m now that the
prebuilt cache can repack.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K1UY3f8gm2jA1J23C3ew9s
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
get_screen_cons() was gated on DISPLAY being set, which only ever happens on
Linux, so headful runs on Windows and macOS generated fingerprints with no
monitor bound at all.
Fixes#425
headless='virtual' reaches launch_options as headless=False with
virtual_display set (async_api rewrites it), so the headful gate fired and
clamped the fingerprint to Xvfb's 1x1 stub. fix_screen_no_taskbar then drove
availHeight to -39 and validate_config rejected the launch outright.