Files
warmbly/qa
..

Proof harness

Recorded proof for pull requests. A flow is a short scripted Playwright walk through the dashboard. Running it produces a 1920x1080 H.264 walkthrough and named full-resolution stills, and pnpm share puts them on the PR:

  • first publish: a Proof section appended to the PR description, stills inline, walkthrough videos last
  • every publish after that: a comment with before/after pairs for only the stills that changed since the previous publish, the commits in between, and the videos of the flows that changed. Nothing changed means nothing is posted

It is not a test gate and CI does not run it (CI only typechecks it and checks the flows load). It exists so a reviewer can see a change without checking it out.

When to record

Only for a UI change worth watching. Most PRs need no proof.

  • record: a new page, dialog, drawer or multi-step flow; a redesigned or re-laid-out screen; a changed interaction (new controls, new states, a different path through a task); a visible bug fix where the before was visibly broken; or when the user asks for it
  • skip: anything without a visible change (backend, API, migrations, workers, tests, CI, docs, refactors, dependency bumps) and small visual tweaks (copy, a label, an icon, spacing or colour nudges). Describe those in the PR text. When unsure, skip
  • follow-ups: re-record only when the commit changes what the proof shows

Why scripted

The model never looks at the page while the browser runs. You write the flow from the code you just changed (you know the routes, labels and fields), run it, and read one line: pass, or the failing step. A 15-step flow costs a few thousand tokens to write and nothing to re-run. Driving a browser step by step from screenshots or snapshots costs that much on every step of every run, and leaves dead air in the video while the model thinks.

Be frugal with the machine

Several agents share this machine, and every stack and browser holds real memory. The harness enforces the expensive parts; the rest is on you.

  • one recording at a time, machine-wide. pnpm proof (and pnpm aria) take a lock in the system temp directory; a second one waits for the first instead of starting its own Chromium. Never fan recording out to subagents: they would only queue behind each other while each holds a context
  • one stack per worktree, in the smallest mode the flow needs. lite is enough for anything that only reads or edits records
  • stop it when you are done: pnpm stack down after pnpm share. A stack nobody records against stops itself after 60 minutes (QA_STACK_IDLE_MIN), but do not lean on that
  • the browser and the encoder never run at once: frames are spooled to disk during the run and encoded after the browser exits, one flow at a time, with bounded encoder threads

Measured on this repo's dashboard:

memory
lite stack (backend, dashboard, Redis, NATS, Mailpit) ~300 MB idle, ~500 MB once used
full stack (+ consumer, worker, realtime) ~560 MB idle
sandbox stack (+ tracking, Dovecot, simulator) ~760 MB idle
Chromium during a recording (proportional set size) ~800 MB
ffmpeg encoding, after Chromium has exited ~580 MB

One-time setup

make infra              # repo root: the shared postgres (each stack gets its own database in it)
cd qa && pnpm bootstrap # deps, Playwright's Chromium, the small stack images, then `pnpm check`

pnpm check verifies node 22.18+, ffmpeg with libx264, Chromium (and that it launches), gh 2.99+ and signed in, docker, the shared postgres, the images each mode needs, and this worktree's stack. Each failure prints its fix.

The loop

pnpm stack up                 # this worktree's own stack (lite unless told otherwise)
pnpm proof contacts           # run the flows whose file or title matches
git push && gh pr create ...  # the PR has to exist first
pnpm share                    # publish the last run to this branch's PR
pnpm stack down               # free the memory

After a follow-up commit: push, record the flows your change touched again, and pnpm share. It posts a comment, never a second description section. Record follow-ups with pnpm proof:fresh <filter>, which re-seeds the stack (about 15 seconds) before recording, so the comparison shows what the commit changed and not what background jobs did to the data in the meantime.

pnpm share --dry-run prints the body and the gh command without uploading anything, and works before the PR exists. --no-video posts stills only, --force comments even when no still changed, --replace-body rewrites the description's Proof section instead of commenting (for a re-record nobody has reviewed yet), --allow-failed publishes a run with failing flows.

share refuses a run with failing flows and a run recorded at a different commit than HEAD, and warns when the tree was dirty or HEAD is not pushed.

The stack

pnpm stack up [lite|full|sandbox] runs this worktree's Warmbly natively against the shared postgres, with everything else private to it: its own database (warmbly_qa_<worktree>), Redis, NATS and Mailpit, on ports picked once per worktree and remembered in .artifacts/stack/ports. Nothing crosses between stacks: no cached profile, no worker event, no login code. The Go services run from binaries built once per up (go build, cached), not from one go run toolchain per service. Linux and WSL.

Mode Runs Seed and account
lite backend, dashboard rich seed (SEED_RICH + SEED_FULL), dev@warmbly.com
full lite + consumer, worker, realtime same seed; sends, syncs and live updates work
sandbox full + tracking, Dovecot, simulator Sunrise Labs (sandbox@warmbly.test): live mailboxes, campaign mail, opens, clicks, replies

lite and full share a seed and switch freely. The sandbox is a different organization, so moving to or from it needs pnpm stack reset-data <mode>; up refuses rather than mix them. In the sandbox the simulator keeps changing data while you record, so its stills differ from run to run by design.

Other commands: status (what runs, ports, memory), logs [name], restart (rebuilds the binaries; the dashboard hot-reloads on its own), reset-data [mode] (drop the database, fresh Redis, reseed), ls (every QA stack on the machine), down.

Realtime and tracking run from the local ghcr.io/warmbly/warmbly/*:prod images. up warns when one was built before the last commit to its source, with the command that rebuilds it.

Sign-in happens once per stack through the API and the login code in the stack's Mailpit; the first-run wizard is completed through its own endpoint, and the session is saved under .artifacts/auth. A flow that lands on the sign-in page, onboarding or the workspace picker fails and says why.

Recording against a stack the harness did not start is possible but explicit: QA_WEB_URL and QA_API_URL (and QA_MAILPIT_URL, QA_EMAIL, QA_PASSWORD). Without either, pnpm proof refuses rather than record whatever happens to answer on the default ports, which is usually another session's stack.

Writing a flow

Flows live in flows/<area>.flow.ts, one test per walkthrough. Extend the area's file when your change lands in it; add one when it does not exist.

import { expect, test } from "../lib/proof.ts";

test("search contacts and open a contact's details", async ({ page, proof }) => {
  await page.goto("/app/contacts");
  await expect(page.getByRole("table")).toBeVisible();
  await proof.chapter("Contacts", "Search the list, then open one contact");

  await page.getByRole("textbox", { name: /Search by name/ }).pressSequentially("beth", { delay: 60 });
  await expect(page.getByRole("row", { name: /Beth Chen/ })).toBeVisible();
  await proof.shot("search-results", { caption: "Search narrowed to one contact" });
});
  • proof.chapter(title, description?) shows a title card in the video
  • proof.shot(name, { caption?, target?, ignore? }) saves a still. The name is kebab-case and is the key follow-up comments diff by, so keep it stable across commits. target shoots one element. ignore takes locators whose content changes on its own (relative times, live counters): the still is published untouched, and those regions are skipped when it is compared with the previous publish
  • proof.dwell(ms?) holds a result on screen long enough to read
  • test.use({ seed: "sandbox" }) for a flow written against the Sunrise Labs data. A flow is skipped, with the command that fixes it, when the stack holds the other seed
  • test.use({ signedIn: false }) for a flow that starts on the sign-in pages

Conventions:

  • expect(...) the content before every chapter and shot, so neither lands on a loading state
  • locate by role and accessible name (getByRole, getByLabel, getByText); no CSS classes
  • pressSequentially(text, { delay: 60 }) where typing should be visible, fill where it should not
  • no waitForTimeout except through dwell
  • keep a walkthrough under about 30 seconds; split longer stories into tests
  • the title reads as a sentence about what the reviewer sees; it heads the PR section

When you need a selector you cannot read off the code, print the page's accessibility tree with the saved session instead of opening a browser:

pnpm aria /app/contacts            # whole page
pnpm aria /app/contacts main       # one region, fewer tokens

Do not read the video or screenshot your way around the page. Read a still only when the change is a visual judgment (spacing, colour, layout) that the assertions cannot make.

When a flow fails

The list reporter prints the failing step and its locator. Failures keep a Playwright trace (DOM snapshots, network, console) and the reporter prints its path; open it with pnpm exec playwright show-trace <trace.zip> when the error line is not enough. A failed flow's frames are dropped instead of encoded (QA_VIDEO_ON_FAIL=1 keeps them).

Recording settings

1920x1080 at 30 fps, H.264 CRF 18 (slow preset), yuv420p so every browser's player shows it. A clip over 10 MB is re-encoded smaller. Frames come from Playwright's screencast at full size and are encoded here, because Playwright's built-in video option is VP8 at 1 Mbit/s and smears text. The video starts at the first painted frame, not on the blank page before the app renders. The animated cursor, click ripples and action labels are Playwright's screencast.showActions. Trace screenshots stay off: they would start a smaller screencast first, and the recorder refuses frames below 1080p.

Variable Default
QA_FPS 30 frame rate
QA_CRF 18 quality, lower is better and bigger
QA_PRESET slow x264 preset
QA_ENCODE_THREADS 4 encoder threads, which also bound its memory
QA_SLOWMO 250 ms between actions, so a viewer can follow
QA_MAX_VIDEO_MB 10 re-encode above this
QA_VIDEO_ON_FAIL 1 encodes failed flows too
QA_STACK_IDLE_MIN 60 stop an unused stack after this long; 0 never
QA_LOCK_WAIT_MIN 30 how long a recording waits for another to finish
QA_WEB_URL, QA_API_URL, QA_MAILPIT_URL this worktree's stack record a stack the harness did not start
QA_EMAIL, QA_PASSWORD the mode's seeded account account to sign in as
QA_ALLOW_HOST extra hostnames allowed besides localhost

Safety

The repository is public and so is everything attached to its pull requests. The harness refuses to record anything but localhost (or a host named in QA_ALLOW_HOST), so recordings only ever show seed data. Never point it at production or a customer workspace, and never attach media to a PR any other way: no image hosts, no commits of screenshots, no other repositories. Everything the harness writes stays under the ignored .artifacts/.