Commit Graph
2 Commits
Author SHA1 Message Date
Brennan Benson e594cb06af test(mobile): record RPC goldens without a pinned commit, and check recorded requests against the desktop's params rules (#23732)
* test(mobile): add rpc:diff to decode what a recording change moved

The RPC recording goldens are content-addressed JSON, so their raw git diff is
pool hashes. `pnpm --dir mobile rpc:diff [<base>]` decodes both sides and prints,
per golden, the checkpoint, field and JSON path that moved with both values,
grouped across checkpoints, plus added and removed goldens. `--summary <file>`
appends a Markdown report capped for GitHub's step-summary limit.

It reads any pooled format, so it can prove the next commit's format change
moves no recorded value. Checkpoints are matched by occurrence because an id can
repeat within one golden.

This commit adds files under the recorder directory, which moves the header
digest every golden pins; the next commit removes that header.

* test(mobile): record RPC goldens without a pinned commit or input digests

Every golden carried a pinned `baseline` commit plus digests of the recorder,
its mount adapter and its scenario, and the record script refused to run unless
the product tree matched the pin. So every behaviour change repinned to its own
branch commit and rewrote all ~790 files, the squash made that commit
unreachable, and main's pin job stayed red until a hand-made repin pull request
landed (22 of them in 12 days). The digests could only fail when an input moved
and the recording did not, which is exactly the change that carries no
information; every run already re-derives each golden from the current tree and
compares it.

Format 6 keeps the format version, operation, family, named deltas, the value
pool and the recording. Removed: the pin and fence, the three digest modules and
their test, the pin guard and its CI job, and the dead scenario `version` field
(the manifest reader now refuses `baseline` and `version` with a message).

- `pnpm --dir mobile rpc:record [<golden-id>...] [--prune]` records all or some
  goldens; orphans are listed, and deleted only with `--prune`. Every derived
  test title now starts with its golden id so an id selects it.
- `compareGolden` reports every difference in one failure (identity fields by
  name, the checkpoint list, each checkpoint/field/path grouped), keeps the
  final byte compare, and ends with the command to re-record that golden.
- `unhandled-recording.test.ts` now drives a detached rejection through
  `runRecording` into a checkpoint and the cleanup checkpoint; no golden carries
  one, and disconnecting the capture passed every suite before.
- Seam rules that existed only to keep a digest honest are gone; the
  mutant-reachability, register-completeness and one-exposure rules stay.
- CI: `Mobile tests on main` runs the whole mobile suite on every merge that
  touches mobile/, src/shared/, the root lockfile or the host RPC paths, since
  `verify` never runs on main. A new `Mobile RPC Recording Replay` workflow
  replays the recordings on pull requests that touch src/shared/ or the root
  lockfile without touching mobile/. `verify` writes the `rpc:diff` report to
  the job summary.

Proof: `rpc:diff` against the parent reports no recorded behaviour moved; each
golden only loses its ten header lines.

* test(mobile): check every recorded request against the host's params contract

The goldens script the host's replies, so a scenario could record a success
for a request the real host would refuse, and a desktop change that tightens a
params schema moved no golden at all.

`recorded-request-params.test.ts` parses every distinct request the corpus puts
on the wire with the host dispatcher's own `parseRpcRequestParams` and the
schema `rpc-params-catalog.generated.ts` binds to that method. It fails on a
method the host lacks, params it refuses, params sent to a method that takes
none (the dispatcher never reads them), and keys the schema silently strips
unless an inventory entry gives the reason; a stale entry fails too. Each rule
is also shown firing on a made-up request, since the corpus has no instance of
three of them. It imports the desktop dispatcher, so it sits beside the other
Node-side tests outside the RN test program, and the params-contract boundary
now exempts test files, which are never bundled.

It found twelve requests the host would refuse, all from invented fixture
values, not product code, fixed at their source:
- git.branchDiff sent `base-oid`/`head-oid`/`merge-base` where the host needs
  full object ids (diff-review and source-control adapters, and the branch
  compare replies in the manifest that feed them);
- an iOS push registration without `apnsEnvironment`, which a real iOS token
  always carries (`push-token.ts`); the adapter now defaults to `production`;
- `settings.update` given Linear's `assigned` filter as a GitHub preset, which
  the product type forbids; the scenario now picks `my-issues`;
- GitLab `projectRef` as a string where the host and the product type take
  `{ host, path }` (7 methods, 5 adapters and the manifest).

46 goldens move, and a decoded comparison of every one of them shows no change
other than those substitutions; `rpc:diff` lists them.

* ci(mobile): detect a mobile change without a SIGPIPE-prone grep pipe

Under the runner's pipefail, grep -q exiting on its first match SIGPIPEs git
diff on a long file list, so a large pull request touching mobile/ read as
uncovered and replayed the recordings a second time.

* test(mobile): drop comments that still describe the golden header and digests

Eleven adapters justified an import rule by the header a golden no longer
carries, and that rule's test is gone. The census failure now names the
rpc:record and --prune commands.

* test(mobile): refuse a golden that keeps a key no recording writes

Decoding dropped unknown top-level keys, so an old header left behind by a
hand-resolved merge conflict passed every compare unseen.

* ci(mobile): summarize RPC recording changes after a failed test step too

* test(mobile): stream rpc:record output instead of capturing it

A captured run stayed silent for its whole duration and clipped its tail,
where the failure summary sits, past 8 MB.

* test(ci): let the Ruby-gate contract skip the always-run RPC summary step

fef088d8f4 gave the summary step an `if: ${{ !cancelled() }}`, and this test lists every gated step in `verify` and expects each to be gated on the Ruby scope.

* test(mobile): replay only a golden file that is exactly what rpc:record writes

Replay compared two re-encodings of decoded values, so anything decoding drops (a leftover header
key, a hand edit) sat in the committed file uncompared; a key allow-list covered one case of that.
Replay now passes only if the file text equals the formatted golden for the run, sharing one
formatter with writeGolden, and keeps the field-level report as the failure message. The allow-list
goes; the value-based compareGolden stays for the bridged run, which has no file.

* test(mobile): end a corrupt or hand-edited golden's failure with the re-record command

A hand edit to a pooled value failed in decode with only "Golden value <hash> does not hash to its
pool key": no golden id and no command to fix it. readGolden now prefixes parse and decode failures
with the golden id and ends them with the rpc:record command. The rpc:diff header also said it
always exits 0; it exits non-zero when git or a golden cannot be read, and now says so.

* ci(mobile): run Mobile Checks on every src/shared and root lockfile change

Replaces the replay-only workflow: mobile imports hundreds of shared modules, so a shared edit can
move a golden or break mobile's typecheck, and the full job catches both before merge. A root
lockfile-only change skips the Ruby release checks, which read no root Node dependency.

* test(mobile): list or prune orphaned goldens even when the recording run fails

Orphans come from the manifest, not the run, so a failed or timed-out rpc:record still reports
them; the exit code stays non-zero. README: say what a failed replay reports (first differing
path per field, capped groups) and what rpc:diff compares with and without a base.

* test(mobile): pin the RPC recording goldens to LF so a CRLF checkout still replays

Replay now requires the committed golden text to equal exactly what rpc:record writes, which is
LF. A Windows checkout with core.autocrlf=true converted every golden to CRLF and failed all 790
with "holds the same recording but is not the file rpc:record writes for it".
2026-09-29 23:26:55 -07:00
Jinwoo Hong 40b2230508 test(mobile): typecheck the test files on a ratchet, and pin the reply enums where tsc looks (#21298)
* fix(mobile): move the last six reply-enum pins where tsc looks

mobile/tsconfig.json excludes *.test.ts, so a `Record<HostUnion, true>`
coverage record in a schema test is never typechecked: the two that existed
(SshConnectionStatus, GitHubProjectOwnerType) checked nothing, and the four
closed enums beside them had only a doc citation of the host type.

Each arm list moves into its schema module as hostUnionArms<Union>(), which
#21269 introduced for the same reason, and each test iterates the exported
list instead of holding its own copy:

- SSH_CONNECTION_STATUS to SshConnectionStatus
- PROJECT_OWNER_TYPE to GitHubProjectOwnerType
- DETAIL_FILE_STATUS to GitHubPRFile['status']
- PUSH_TEST_REFUSAL_REASONS and PUSH_REGISTER_REFUSAL_REASONS to the refusal
  arms of MobilePushTestResult and MobilePushRegisterResult
- SETUP_RUN_POLICIES to SetupRunPolicy

openEnum's parameter widens from a non-empty tuple to `readonly string[]` so
a hostUnionArms list can feed it. z.enum already accepts the same, so the
tuple constraint only excluded callers zod itself takes; behaviour unchanged.

Twelve mutations prove the pins: dropping one arm and adding a bogus one
each fail mobile tsc in all six places. Zero goldens move, the schemas'
behaviour being unchanged, and the 21 recording suites pass at the existing
baseline.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): fix the type errors in eighteen test files

Found by typechecking the tests for the first time (see the config that
follows). All mechanical, none weakens a product type:

- 67 `act(() => vi.advanceTimersByTime(...))` callbacks return VitestUtils
  where act wants void, so each becomes a block. The async ones await only a
  genuinely promise-returning call, so no extra microtask tick is introduced.
- Four fixtures were stale against a product type that gained a required
  member: MobileViewState.alwaysShowDefaultBranch, PrSidebarData.checksError,
  the branch-compare summary's errorMessage, and SessionOptionDescriptor's
  transport, which #20884 added precisely so a producer could not inherit the
  wrong lane's rendering by omission.
- `getLastConnectedAt` on the shared relay fake was typed `() => null`, which
  refused the timestamp two escalation suites assign to it.
- Two holders used before assignment take `!`, one `advance!.kind === ...`
  becomes `advance?.kind`, one widened status arm takes `as const`, and the
  Expo notification fixture keeps `data` required because the dismissal cases
  assign through it.

631 test files pass, 6222 tests, unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): typecheck the test files, on a ratchet

mobile/tsconfig.json excludes *.test.ts so Metro never compiles tests into the
release bundle, and vitest transpiles without checking types. Nothing had ever
typechecked a mobile test, which is why a `Record<HostUnion, true>` pin written
in one proved nothing and why 144 of the 630 test files had drifted.

tsconfig.test.json is that program with the tests put back, behind
`typecheck:tests`. Four files stay out: they import the desktop main process or
src/shared/child-process, which are written against @types/node, and this
program's libs are React Native's, where setTimeout answers a number rather
than a NodeJS.Timeout. Pulling that graph in reports ~280 errors about the
desktop rather than about mobile; vitest runs those four under Node, which is
where they belong.

The CI gate is a ratchet rather than the raw typecheck, modelled on
check-ts-nocheck-ratchet.mjs: 126 files still fail, so the gate freezes that
set and fails when a file that checks today stops checking, or when a baseline
entry starts checking and was not pruned. The list may only shrink.

Why not zero: 180 of the remaining 510 errors are one seam — tests locate
mocked react-native components by string name, which `ElementType` does not
admit — and closing it means either 180 casts or a global JSX declaration for
the mocked names. That is a design decision, not a mechanical fix, so it is
left for a follow-up rather than made here. The rest are smaller clusters of
the same kind: vi.fn mocks assigned into typed slots, call-arg tuple indexing,
and createElement props fixtures.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile-recorder): correct the corpus counts and the salvage claim

The oracle section still quoted the corpus as 368 scenarios and 727 goldens;
it is 393 and 778, and the three replay suites report 781 tests. Each number
now names the command that measures it.

"No golden carries one" was the load-bearing error: 44 goldens carry a
recorded `reply-salvage` today, starting with the push-test unknown-reason
scenario #21176 added for exactly that purpose. The paragraph claimed the
observation pins an absence when on those families it pins a recorded drop.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the tests-typecheck ratchet's parser

The gate reads tsc's output, and tsc indents the "Overload 1 of 2, ..." detail
under an error. Counting those as filenames would write unparseable entries
into the baseline and leave the gate unprunable, so the parser is pinned on
that shape as well as on the added/stale diff.

Written against the gate itself: it flagged this file before the directive it
carried was removed, which is the end-to-end proof the spawn half works.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): await the timer advances the act() rewrite dropped

Rewriting `await act(async () => vi.advanceTimersByTimeAsync(n))` into a
braced body left the returned promise floating at 27 sites, so the advance
was no longer ordered before the assertions that follow it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): unshadow MobileHostCard's .tsx suite

A wildcard `include` keeps only the higher-priority extension, so
MobileHostCard.test.tsx sat outside every tsc program while
MobileHostCard.test.ts existed beside it. Its one error is the same
react-test-renderer seam its sibling is baselined for.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): census every test file into the typecheck program

The ratchet diffs only files that error, so a test excluded from
tsconfig.test.json or shadowed by a sibling extension left the gate
silently. Every *.test.ts(x) on disk must now be in the program or
named in TESTS_OUTSIDE_PROGRAM with its reason.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(shared): make the enum helpers refuse the ways they can prove nothing

openEnum takes a `const` T so a bare literal keeps its arms rather than
widening to string. hostUnionArms blocks inference of U with NoInfer and
defaults it to never, so a call that omits the host union — where the
record would only pin itself — no longer compiles.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): describe the census and correct the baseline count

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): give the push fixture cast its SAFETY rationale

Widening the pre-existing cast made the changed-code gate attribute it as
a new finding.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): build the push fixtures as typed notifications

Replaces the `as unknown as` cast with Expo's own types, filling
FirebaseRemoteMessage and its notification once in two builders, and
passes the data payload in rather than mutating through an optional
member. Typing the fixture showed one assertion comparing the scheduled
content against the whole arriving content, which only held while the
cast let the fixture omit the two members the presenter drops; it now
names the four members the presenter forwards.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): keep the grouped-question advance read non-optional

`advance?.kind` let an absent advance take the null-draft branch instead
of failing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): run the tests-typecheck ratchet on Windows

Spawns tsc's JS entry on this Node instead of the node_modules/.bin
shim, which is a POSIX shell script that Windows resolves to tsc.CMD and
then appends .exe to. Parsed paths are normalised to POSIX so a Windows
run does not read every baseline entry as both stale and added.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): close the ratchet's @ts-nocheck hole and read tsc once

tsc exits 0 on a @ts-nocheck file, so a baselined test could be "fixed"
with one line, pruned, and never checked again; the census now names any
program test file whose leading comment carries the directive.

`--noEmit --listFiles` answers both questions in one pass, so the gate
spawns tsc once rather than twice. Corrects the two stale counts, and
states hostUnionArms' real reason for living in the schema module now
that tests are typechecked.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-17 18:58:25 -04:00