Files
warmbly/qa/README.md
T

244 lines
12 KiB
Markdown

# Proof harness
Recorded proof for pull requests. A flow is a short scripted Playwright walk
through the dashboard. Running it produces a 1920x1080 H.264 walkthrough and
named full-resolution stills, and `pnpm share` puts them on the PR:
- **first publish**: a Proof section appended to the PR description, stills
inline, walkthrough videos last
- **every publish after that**: a comment with before/after pairs for only the
stills that changed since the previous publish, the commits in between, and
the videos of the flows that changed. Nothing changed means nothing is posted
It is not a test gate and CI does not run it (CI only typechecks it and checks
the flows load). It exists so a reviewer can see a change without checking it
out.
## When to record
Only for a UI change worth watching. Most PRs need no proof.
- **record**: a new page, dialog, drawer or multi-step flow; a redesigned or
re-laid-out screen; a changed interaction (new controls, new states, a
different path through a task); a visible bug fix where the before was
visibly broken; or when the user asks for it
- **skip**: anything without a visible change (backend, API, migrations,
workers, tests, CI, docs, refactors, dependency bumps) and small visual
tweaks (copy, a label, an icon, spacing or colour nudges). Describe those in
the PR text. When unsure, skip
- **follow-ups**: re-record only when the commit changes what the proof shows
## Why scripted
The model never looks at the page while the browser runs. You write the flow
from the code you just changed (you know the routes, labels and fields), run
it, and read one line: pass, or the failing step. A 15-step flow costs a few
thousand tokens to write and nothing to re-run. Driving a browser step by step
from screenshots or snapshots costs that much on every step of every run, and
leaves dead air in the video while the model thinks.
## Be frugal with the machine
Several agents share this machine, and every stack and browser holds real
memory. The harness enforces the expensive parts; the rest is on you.
- **one recording at a time, machine-wide.** `pnpm proof` (and `pnpm aria`)
take a lock in the system temp directory; a second one waits for the first
instead of starting its own Chromium. Never fan recording out to subagents:
they would only queue behind each other while each holds a context
- **one stack per worktree**, in the smallest mode the flow needs. `lite` is
enough for anything that only reads or edits records
- **stop it when you are done**: `pnpm stack down` after `pnpm share`. A stack
nobody records against stops itself after 60 minutes (`QA_STACK_IDLE_MIN`),
but do not lean on that
- the browser and the encoder never run at once: frames are spooled to disk
during the run and encoded after the browser exits, one flow at a time, with
bounded encoder threads
Measured on this repo's dashboard:
| | memory |
| --- | --- |
| `lite` stack (backend, dashboard, Redis, NATS, Mailpit) | ~300 MB idle, ~500 MB once used |
| `full` stack (+ consumer, worker, realtime) | ~560 MB idle |
| `sandbox` stack (+ tracking, Dovecot, simulator) | ~760 MB idle |
| Chromium during a recording (proportional set size) | ~800 MB |
| ffmpeg encoding, after Chromium has exited | ~580 MB |
## One-time setup
```bash
make infra # repo root: the shared postgres (each stack gets its own database in it)
cd qa && pnpm bootstrap # deps, Playwright's Chromium, the small stack images, then `pnpm check`
```
`pnpm check` verifies node 22.18+, ffmpeg with libx264, Chromium (and that it
launches), gh 2.99+ and signed in, docker, the shared postgres, the images each
mode needs, and this worktree's stack. Each failure prints its fix.
## The loop
```bash
pnpm stack up # this worktree's own stack (lite unless told otherwise)
pnpm proof contacts # run the flows whose file or title matches
git push && gh pr create ... # the PR has to exist first
pnpm share # publish the last run to this branch's PR
pnpm stack down # free the memory
```
After a follow-up commit: push, record the flows your change touched again,
and `pnpm share`. It posts a comment, never a second description section.
Record follow-ups with `pnpm proof:fresh <filter>`, which re-seeds the stack
(about 15 seconds) before recording, so the comparison shows what the commit
changed and not what background jobs did to the data in the meantime.
`pnpm share --dry-run` prints the body and the gh command without uploading
anything, and works before the PR exists. `--no-video` posts stills only,
`--force` comments even when no still changed, `--replace-body` rewrites the
description's Proof section instead of commenting (for a re-record nobody has
reviewed yet), `--allow-failed` publishes a run with failing flows.
`share` refuses a run with failing flows and a run recorded at a different
commit than HEAD, and warns when the tree was dirty or HEAD is not pushed.
## The stack
`pnpm stack up [lite|full|sandbox]` runs this worktree's Warmbly natively
against the shared postgres, with everything else private to it: its own
database (`warmbly_qa_<worktree>`), Redis, NATS and Mailpit, on ports picked
once per worktree and remembered in `.artifacts/stack/ports`. Nothing crosses
between stacks: no cached profile, no worker event, no login code. The Go
services run from binaries built once per `up` (`go build`, cached), not from
one `go run` toolchain per service. Linux and WSL.
| Mode | Runs | Seed and account |
| --- | --- | --- |
| `lite` | backend, dashboard | rich seed (`SEED_RICH` + `SEED_FULL`), `dev@warmbly.com` |
| `full` | lite + consumer, worker, realtime | same seed; sends, syncs and live updates work |
| `sandbox` | full + tracking, Dovecot, simulator | Sunrise Labs (`sandbox@warmbly.test`): live mailboxes, campaign mail, opens, clicks, replies |
`lite` and `full` share a seed and switch freely. The sandbox is a different
organization, so moving to or from it needs `pnpm stack reset-data <mode>`;
`up` refuses rather than mix them. In the sandbox the simulator keeps changing
data while you record, so its stills differ from run to run by design.
Other commands: `status` (what runs, ports, memory), `logs [name]`, `restart`
(rebuilds the binaries; the dashboard hot-reloads on its own), `reset-data
[mode]` (drop the database, fresh Redis, reseed), `ls` (every QA stack on the
machine), `down`.
Realtime and tracking run from the local `ghcr.io/warmbly/warmbly/*:prod`
images. `up` warns when one was built before the last commit to its source,
with the command that rebuilds it.
Sign-in happens once per stack through the API and the login code in the
stack's Mailpit; the first-run wizard is completed through its own endpoint,
and the session is saved under `.artifacts/auth`. A flow that lands on the
sign-in page, onboarding or the workspace picker fails and says why.
Recording against a stack the harness did not start is possible but explicit:
`QA_WEB_URL` and `QA_API_URL` (and `QA_MAILPIT_URL`, `QA_EMAIL`,
`QA_PASSWORD`). Without either, `pnpm proof` refuses rather than record
whatever happens to answer on the default ports, which is usually another
session's stack.
## Writing a flow
Flows live in `flows/<area>.flow.ts`, one `test` per walkthrough. Extend the
area's file when your change lands in it; add one when it does not exist.
```ts
import { expect, test } from "../lib/proof.ts";
test("search contacts and open a contact's details", async ({ page, proof }) => {
await page.goto("/app/contacts");
await expect(page.getByRole("table")).toBeVisible();
await proof.chapter("Contacts", "Search the list, then open one contact");
await page.getByRole("textbox", { name: /Search by name/ }).pressSequentially("beth", { delay: 60 });
await expect(page.getByRole("row", { name: /Beth Chen/ })).toBeVisible();
await proof.shot("search-results", { caption: "Search narrowed to one contact" });
});
```
- `proof.chapter(title, description?)` shows a title card in the video
- `proof.shot(name, { caption?, target?, ignore? })` saves a still. The name is
kebab-case and is the key follow-up comments diff by, so keep it stable
across commits. `target` shoots one element. `ignore` takes locators whose
content changes on its own (relative times, live counters): the still is
published untouched, and those regions are skipped when it is compared with
the previous publish
- `proof.dwell(ms?)` holds a result on screen long enough to read
- `test.use({ seed: "sandbox" })` for a flow written against the Sunrise Labs
data. A flow is skipped, with the command that fixes it, when the stack holds
the other seed
- `test.use({ signedIn: false })` for a flow that starts on the sign-in pages
Conventions:
- `expect(...)` the content before every `chapter` and `shot`, so neither lands
on a loading state
- locate by role and accessible name (`getByRole`, `getByLabel`, `getByText`);
no CSS classes
- `pressSequentially(text, { delay: 60 })` where typing should be visible,
`fill` where it should not
- no `waitForTimeout` except through `dwell`
- keep a walkthrough under about 30 seconds; split longer stories into tests
- the title reads as a sentence about what the reviewer sees; it heads the PR
section
When you need a selector you cannot read off the code, print the page's
accessibility tree with the saved session instead of opening a browser:
```bash
pnpm aria /app/contacts # whole page
pnpm aria /app/contacts main # one region, fewer tokens
```
Do not read the video or screenshot your way around the page. Read a still
only when the change is a visual judgment (spacing, colour, layout) that the
assertions cannot make.
## When a flow fails
The list reporter prints the failing step and its locator. Failures keep a
Playwright trace (DOM snapshots, network, console) and the reporter prints its
path; open it with `pnpm exec playwright show-trace <trace.zip>` when the error
line is not enough. A failed flow's frames are dropped instead of encoded
(`QA_VIDEO_ON_FAIL=1` keeps them).
## Recording settings
1920x1080 at 30 fps, H.264 CRF 18 (`slow` preset), yuv420p so every browser's
player shows it. A clip over 10 MB is re-encoded smaller. Frames come from
Playwright's screencast at full size and are encoded here, because Playwright's
built-in `video` option is VP8 at 1 Mbit/s and smears text. The video starts at
the first painted frame, not on the blank page before the app renders. The
animated cursor, click ripples and action labels are Playwright's
`screencast.showActions`. Trace screenshots stay off: they would start a
smaller screencast first, and the recorder refuses frames below 1080p.
| Variable | Default | |
| --- | --- | --- |
| `QA_FPS` | `30` | frame rate |
| `QA_CRF` | `18` | quality, lower is better and bigger |
| `QA_PRESET` | `slow` | x264 preset |
| `QA_ENCODE_THREADS` | `4` | encoder threads, which also bound its memory |
| `QA_SLOWMO` | `250` | ms between actions, so a viewer can follow |
| `QA_MAX_VIDEO_MB` | `10` | re-encode above this |
| `QA_VIDEO_ON_FAIL` | | `1` encodes failed flows too |
| `QA_STACK_IDLE_MIN` | `60` | stop an unused stack after this long; `0` never |
| `QA_LOCK_WAIT_MIN` | `30` | how long a recording waits for another to finish |
| `QA_WEB_URL`, `QA_API_URL`, `QA_MAILPIT_URL` | this worktree's stack | record a stack the harness did not start |
| `QA_EMAIL`, `QA_PASSWORD` | the mode's seeded account | account to sign in as |
| `QA_ALLOW_HOST` | | extra hostnames allowed besides localhost |
## Safety
The repository is public and so is everything attached to its pull requests.
The harness refuses to record anything but localhost (or a host named in
`QA_ALLOW_HOST`), so recordings only ever show seed data. Never point it at
production or a customer workspace, and never attach media to a PR any other
way: no image hosts, no commits of screenshots, no other repositories.
Everything the harness writes stays under the ignored `.artifacts/`.