selection_evidence is continue-on-error on both the job and its comparison step,
so it can never fail a PR -- it downloads the shard reports, compares selection
against the full results and uploads a review artifact. But a caller's
`needs: test` waits for every job in the called workflow, so living inside
unit-tests.yml it held verify for ~36s after the last shard finished.
It moves to its own reusable workflow called as a sibling, so it still runs on
every PR and still uploads its artifact, but verify no longer waits for it. It is
deliberately absent from verify's needs, and a contract test pins both that and
its advisory status so it cannot drift back onto the critical path.
Measured on a recent run: the shards finished, then selection_evidence ran 36s,
then verify 3s. Only the last of those gates anything.
A caller's `needs` gate the whole called workflow, so while the plan job lived in
unit-tests.yml it could not start until static analysis and typecheck had both
finished and passed -- and the shard matrix then waited on it. The two hops were
serial when they did not need to be: planning reads the checkout, a git diff
against HEAD^1, the import graph and the checked-in timing baseline in
config/scripts/ci-shard-timings.json, and consumes nothing that static analysis,
typecheck or the native-cache primer produce.
Planning moves to its own reusable workflow so pr.yml can run it against
code_paths alone, overlapping it with the gate. Measured across 99 runs, the
shard matrix is created a median 93s earlier (p25 47s, p90 241s, never later).
Planning stays a required predecessor of the shards, so an empty assignment
cannot expand the matrix.
The gate itself is deliberately left in place. It fires on 22% of runs, and the
shard queue wait knees hard above ~9 concurrent ARM jobs -- 4s median below that
against 218s at 15-19 -- so admitting 8 doomed shards per failed run would cost
more in queue pressure than it returns in latency.
Cost is one 37s ubuntu-latest job, which does not touch the ARM pool the shards
contend for.
A planning failure still fails the PR: the shards are skipped, and verify's
check_job requires success whenever the classifier says tests should run, so it
reports `test: expected success, got skipped`.
* ci: benchmark per-job Node compile caching on full unit shards
* ci: measure unit shards with three and four workers
* ci: benchmark localization extraction CLI patch
* perf(build): reuse identical relay bundles across platforms
* ci: compare Vitest 4 and 5 on complete ARM shards
* perf(ci): upgrade localization extraction to skip irrelevant syntax walks
* perf(ci): use all four ARM cores and remove benchmark workflows
* ci: preserve failures while capturing unit source revision
* fix(ci): preserve commented and escaped localization calls
* ci: remove corrected localization benchmark harness
* ci: stage heavy checks and measure affected-test selection
* fix(ci): exercise the real Git boundary in unit selection planning
* Harden review cancellation and CI demand reporting
* test: reuse isolated mobile bundle fixtures for verifier checks
* ci: run PR unit shards on ARM and overlap independent web builds
* docs: record controlled CI overlap and runner measurements
* test: observe WebRTC packets with the host clock
* ci: isolate Windows installer CIM probe from native test load
* docs: record native probe scheduling validation
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
* ci: balance existing unit and E2E shards using recorded timings
* ci: fix timing refresh units and deferred-menu test traversal
* ci: preserve isolated E2E window launch policy
* ci: keep diagnostic artifact outages from failing tests