Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Proof harness
Recorded proof for pull requests. A flow is a short scripted Playwright walk
through the dashboard. Running it produces a 1920x1080 H.264 walkthrough and
named full-resolution stills, and pnpm share puts them on the PR:
- first publish: a Proof section appended to the PR description, stills inline, walkthrough videos last
- every publish after that: a comment with before/after pairs for only the stills that changed since the previous publish, the commits in between, and the videos of the flows that changed. Nothing changed means nothing is posted
It is not a test gate and CI does not run it (CI only typechecks it and checks the flows load). It exists so a reviewer can see a change without checking it out.
When to record
Only for a UI change worth watching. Most PRs need no proof.
- record: a new page, dialog, drawer or multi-step flow; a redesigned or re-laid-out screen; a changed interaction (new controls, new states, a different path through a task); a visible bug fix where the before was visibly broken; or when the user asks for it
- skip: anything without a visible change (backend, API, migrations, workers, tests, CI, docs, refactors, dependency bumps) and small visual tweaks (copy, a label, an icon, spacing or colour nudges). Describe those in the PR text. When unsure, skip
- follow-ups: re-record only when the commit changes what the proof shows
Why scripted
The model never looks at the page while the browser runs. You write the flow from the code you just changed (you know the routes, labels and fields), run it, and read one line: pass, or the failing step. A 15-step flow costs a few thousand tokens to write and nothing to re-run. Driving a browser step by step from screenshots or snapshots costs that much on every step of every run, and leaves dead air in the video while the model thinks.
Be frugal with the machine
Several agents share this machine, and every stack and browser holds real memory. The harness enforces the expensive parts; the rest is on you.
- one recording at a time, machine-wide.
pnpm proof(andpnpm aria) take a lock in the system temp directory; a second one waits for the first instead of starting its own Chromium. Never fan recording out to subagents: they would only queue behind each other while each holds a context - one stack per worktree, in the smallest mode the flow needs.
liteis enough for anything that only reads or edits records - stop it when you are done:
pnpm stack downafterpnpm share. A stack nobody records against stops itself after 60 minutes (QA_STACK_IDLE_MIN), but do not lean on that - the browser and the encoder never run at once: frames are spooled to disk during the run and encoded after the browser exits, one flow at a time, with bounded encoder threads
Measured on this repo's dashboard:
| memory | |
|---|---|
lite stack (backend, dashboard, Redis, NATS, Mailpit) |
~300 MB idle, ~500 MB once used |
full stack (+ consumer, worker, realtime) |
~560 MB idle |
sandbox stack (+ tracking, Dovecot, simulator) |
~760 MB idle |
| Chromium during a recording (proportional set size) | ~800 MB |
| ffmpeg encoding, after Chromium has exited | ~580 MB |
One-time setup
make infra # repo root: the shared postgres (each stack gets its own database in it)
cd qa && pnpm bootstrap # deps, Playwright's Chromium, the small stack images, then `pnpm check`
pnpm check verifies node 22.18+, ffmpeg with libx264, Chromium (and that it
launches), gh 2.99+ and signed in, docker, the shared postgres, the images each
mode needs, and this worktree's stack. Each failure prints its fix.
The loop
pnpm stack up # this worktree's own stack (lite unless told otherwise)
pnpm proof contacts # run the flows whose file or title matches
git push && gh pr create ... # the PR has to exist first
pnpm share # publish the last run to this branch's PR
pnpm stack down # free the memory
After a follow-up commit: push, record the flows your change touched again,
and pnpm share. It posts a comment, never a second description section.
Record follow-ups with pnpm proof:fresh <filter>, which re-seeds the stack
(about 15 seconds) before recording, so the comparison shows what the commit
changed and not what background jobs did to the data in the meantime.
pnpm share --dry-run prints the body and the gh command without uploading
anything, and works before the PR exists. --no-video posts stills only,
--force comments even when no still changed, --replace-body rewrites the
description's Proof section instead of commenting (for a re-record nobody has
reviewed yet), --allow-failed publishes a run with failing flows.
share refuses a run with failing flows and a run recorded at a different
commit than HEAD, and warns when the tree was dirty or HEAD is not pushed.
The stack
pnpm stack up [lite|full|sandbox] runs this worktree's Warmbly natively
against the shared postgres, with everything else private to it: its own
database (warmbly_qa_<worktree>), Redis, NATS and Mailpit, on ports picked
once per worktree and remembered in .artifacts/stack/ports. Nothing crosses
between stacks: no cached profile, no worker event, no login code. The Go
services run from binaries built once per up (go build, cached), not from
one go run toolchain per service. Linux and WSL.
| Mode | Runs | Seed and account |
|---|---|---|
lite |
backend, dashboard | rich seed (SEED_RICH + SEED_FULL), dev@warmbly.com |
full |
lite + consumer, worker, realtime | same seed; sends, syncs and live updates work |
sandbox |
full + tracking, Dovecot, simulator | Sunrise Labs (sandbox@warmbly.test): live mailboxes, campaign mail, opens, clicks, replies |
lite and full share a seed and switch freely. The sandbox is a different
organization, so moving to or from it needs pnpm stack reset-data <mode>;
up refuses rather than mix them. In the sandbox the simulator keeps changing
data while you record, so its stills differ from run to run by design.
Other commands: status (what runs, ports, memory), logs [name], restart
(rebuilds the binaries; the dashboard hot-reloads on its own), reset-data [mode] (drop the database, fresh Redis, reseed), ls (every QA stack on the
machine), down.
Realtime and tracking run from the local ghcr.io/warmbly/warmbly/*:prod
images. up warns when one was built before the last commit to its source,
with the command that rebuilds it.
Sign-in happens once per stack through the API and the login code in the
stack's Mailpit; the first-run wizard is completed through its own endpoint,
and the session is saved under .artifacts/auth. A flow that lands on the
sign-in page, onboarding or the workspace picker fails and says why.
Recording against a stack the harness did not start is possible but explicit:
QA_WEB_URL and QA_API_URL (and QA_MAILPIT_URL, QA_EMAIL,
QA_PASSWORD). Without either, pnpm proof refuses rather than record
whatever happens to answer on the default ports, which is usually another
session's stack.
Writing a flow
Flows live in flows/<area>.flow.ts, one test per walkthrough. Extend the
area's file when your change lands in it; add one when it does not exist.
import { expect, test } from "../lib/proof.ts";
test("search contacts and open a contact's details", async ({ page, proof }) => {
await page.goto("/app/contacts");
await expect(page.getByRole("table")).toBeVisible();
await proof.chapter("Contacts", "Search the list, then open one contact");
await page.getByRole("textbox", { name: /Search by name/ }).pressSequentially("beth", { delay: 60 });
await expect(page.getByRole("row", { name: /Beth Chen/ })).toBeVisible();
await proof.shot("search-results", { caption: "Search narrowed to one contact" });
});
proof.chapter(title, description?)shows a title card in the videoproof.shot(name, { caption?, target?, ignore? })saves a still. The name is kebab-case and is the key follow-up comments diff by, so keep it stable across commits.targetshoots one element.ignoretakes locators whose content changes on its own (relative times, live counters): the still is published untouched, and those regions are skipped when it is compared with the previous publishproof.dwell(ms?)holds a result on screen long enough to readtest.use({ seed: "sandbox" })for a flow written against the Sunrise Labs data. A flow is skipped, with the command that fixes it, when the stack holds the other seedtest.use({ signedIn: false })for a flow that starts on the sign-in pages
Conventions:
expect(...)the content before everychapterandshot, so neither lands on a loading state- locate by role and accessible name (
getByRole,getByLabel,getByText); no CSS classes pressSequentially(text, { delay: 60 })where typing should be visible,fillwhere it should not- no
waitForTimeoutexcept throughdwell - keep a walkthrough under about 30 seconds; split longer stories into tests
- the title reads as a sentence about what the reviewer sees; it heads the PR section
When you need a selector you cannot read off the code, print the page's accessibility tree with the saved session instead of opening a browser:
pnpm aria /app/contacts # whole page
pnpm aria /app/contacts main # one region, fewer tokens
Do not read the video or screenshot your way around the page. Read a still only when the change is a visual judgment (spacing, colour, layout) that the assertions cannot make.
When a flow fails
The list reporter prints the failing step and its locator. Failures keep a
Playwright trace (DOM snapshots, network, console) and the reporter prints its
path; open it with pnpm exec playwright show-trace <trace.zip> when the error
line is not enough. A failed flow's frames are dropped instead of encoded
(QA_VIDEO_ON_FAIL=1 keeps them).
Recording settings
1920x1080 at 30 fps, H.264 CRF 18 (slow preset), yuv420p so every browser's
player shows it. A clip over 10 MB is re-encoded smaller. Frames come from
Playwright's screencast at full size and are encoded here, because Playwright's
built-in video option is VP8 at 1 Mbit/s and smears text. The video starts at
the first painted frame, not on the blank page before the app renders. The
animated cursor, click ripples and action labels are Playwright's
screencast.showActions. Trace screenshots stay off: they would start a
smaller screencast first, and the recorder refuses frames below 1080p.
| Variable | Default | |
|---|---|---|
QA_FPS |
30 |
frame rate |
QA_CRF |
18 |
quality, lower is better and bigger |
QA_PRESET |
slow |
x264 preset |
QA_ENCODE_THREADS |
4 |
encoder threads, which also bound its memory |
QA_SLOWMO |
250 |
ms between actions, so a viewer can follow |
QA_MAX_VIDEO_MB |
10 |
re-encode above this |
QA_VIDEO_ON_FAIL |
1 encodes failed flows too |
|
QA_STACK_IDLE_MIN |
60 |
stop an unused stack after this long; 0 never |
QA_LOCK_WAIT_MIN |
30 |
how long a recording waits for another to finish |
QA_WEB_URL, QA_API_URL, QA_MAILPIT_URL |
this worktree's stack | record a stack the harness did not start |
QA_EMAIL, QA_PASSWORD |
the mode's seeded account | account to sign in as |
QA_ALLOW_HOST |
extra hostnames allowed besides localhost |
Safety
The repository is public and so is everything attached to its pull requests.
The harness refuses to record anything but localhost (or a host named in
QA_ALLOW_HOST), so recordings only ever show seed data. Never point it at
production or a customer workspace, and never attach media to a PR any other
way: no image hosts, no commits of screenshots, no other repositories.
Everything the harness writes stays under the ignored .artifacts/.