A concurrency slot is charged per job, not per core, and the account's cap is the scarce resource: standard runner minutes are free and unlimited on a public repository. Two paths spent slots that bought nothing. The unit matrix ran eight fixed shards averaging 6.5 minutes each, 3384 job-slots a day and 68% of all slot demand, while the arm pool queued 10.5 minutes at p95 — the queue was the oversharding. Five shards run the same work in ~10.5 minutes each for three fewer slots per run. Bun profile persistence escalated to all six platforms on `config/`, `resources/` and `.github/` wholesale, which took 36.5% of the last 1100 commits through the full matrix where a platform-flavoured predicate takes 19%. A pull request now qualifies one platform unless the change is platform- flavoured, and the push to main re-qualifies all six, so an unescalated miss surfaces minutes after merge rather than at the next cron. Missing changed-file evidence and an unavailable dependency graph still fail closed to all six.
9.5 KiB
CI demand rollout
This implements the September 28 runner-demand analysis. The baseline inventory covered September 27 04:00–September 28 04:00 UTC: 4,028 workflow runs, with 463 stratified job samples. Estimated occupancy was 1,081 runner-hours, dominated by PR unit shards (414 hours) and Bun qualification (262 hours). These are sampled sums of job durations across different runner pools, not billing totals or a guaranteed forecast of savings.
What runs now
| Work | Ordinary draft update | Ready PR / final checks | Main reference |
|---|---|---|---|
| Static analysis and types | Immediately | Immediately | Existing workflows |
| Unit suite | Full, with shadow selection evidence | Full | Existing daily Node 24/26 x86 suite |
| Packages | After successful static analysis and types | Same | Existing release workflows |
| Bun persistence | Linux x64 for ordinary runtime changes; all six platforms for sensitive inputs | All six platforms when Bun inputs changed | Full nightly qualification at 11:30 UTC |
| Bun glibc/musl qualification | Sensitive inputs only, after persistence succeeds | Both architectures after persistence succeeds | Both architectures |
| E2E | Existing targeted routing | Existing targeted routing | One complete run at 17:00 UTC |
Bun-sensitive inputs include root/toolchain files, configuration, native code, resources, platform-specific paths, persistence, SQLite, orcad, providers, daemon, SSH, relay, and child-process code. Dependency discovery failure, a missing event or an incomplete diff retains full qualification. An unrelated change still skips Bun through the existing dependency classifier. No native artifact is shared across platforms or ABIs. The routine draft reduction is an explicit coverage-placement change; it does not assert identical per-update coverage. Every non-draft synchronize event and the ready-for-review event restores full qualification.
Expensive PR jobs wait for static/type success. This reduces fan-out for failed or rapidly superseded commits without sleeping on a runner. Successful isolated PRs pay the extra stage latency. Existing per-PR cancellation remains in place. Package assertions, native boundaries, SSH/folder coverage, cache warming and slow-test assertions are retained.
Unit selection rollout
ci-unit-plan.mjs discovers the same include/exclude set as Vitest and follows
static imports, re-exports, literal dynamic imports, CommonJS requires and the
renderer aliases. Consumers of indirect filesystem/process inputs remain in the
candidate set, as do script/tool tests. Global configuration changes, deletions,
renames involving removed paths, unknown inputs and graph failures run the full
suite. A shard verifies the plan's source SHA and complete discovery list before
using it. Missing/stale artifacts fall back to full coverage, even if that means
running the suite on fewer shards.
The initial policy is shadow, with every shard retained. A full run spends
FULL_SHARD_COUNT (five) concurrency slots: the account is charged a slot per job
rather than per core, so eight 6.5-minute shards made this matrix 68% of daily slot
demand while the arm pool queued 10.5 minutes at p95. The first
local inventory matched Vitest exactly (9,950 files at validation); representative
source changes retained roughly 88% of files because of indirect input readers.
That is evidence for conservative coverage, not evidence of the analysis's
hypothetical 50% unit-work reduction. Improvements to indirect dependency
modeling should be demonstrated against full results before expanding selection.
Every shard uploads unit-selection.json, unit-timings.json (including module
outcomes), and its assignment. The evidence job combines these into
unit-selection-review-attempt-N/selection-review.json, reporting:
- Whether every discovered file appeared once across a complete reference run.
- Failures outside the candidate set, including failures in otherwise red runs.
- Measured worker time that selection would omit; worker times overlap and are not runner occupancy or a prediction of wall-clock savings.
Missing/duplicate shards, stale plans, interrupted runs and unhandled errors do not count as complete references. Diagnostic upload/report failures do not make tests pass and do not independently fail successful tests.
After representative complete shadow runs show no missed failures, set repository
variable ORCA_UNIT_SELECTION_MODE=selected to enable selection only for draft
PRs. Keep full ready-PR checks and the daily compatibility suite. Inspect at
least a week's evidence across renderer, main, shared, SSH and fixture changes
before promotion, including red runs rather than only successful examples.
Unknown variable values retain shadow mode. Unset the variable or set it to
shadow to roll back immediately. Selected runs use one to eight timing-balanced
shards based on retained work.
Ready-for-review result reuse includes unit full in its source/workflow
identity. A green selected draft cannot satisfy the final full check, even if
the repository variable changes between runs. An already successful full
identical-source check can still be reused.
To inspect downloaded shard artifacts locally:
node config/scripts/ci-unit-selection-review.mjs ARTIFACT_DIRECTORY
Review automation
Pullfrog recognizes the existing Review #N [id] and
Review new commits on #N [id] dispatch names. Explicit dispatchers may provide
pull_request_number and head_sha. Explicit PR identities share concurrency at the workflow boundary. Legacy review
names use a bounded lookup of the latest 100 dispatches and cancel only lower
run IDs for the same PR; a delayed older scope cannot cancel a newer review.
Unrecognized tasks are never grouped. The scope job alone has Actions write
permission for ordered cancellation. Closed PRs and explicitly stale heads
are skipped. A second head check prevents starting an agent after its queued
head has changed. Unrecognized agent tasks remain independent; lookup failures
also retain an independent task rather than cancelling unrelated work.
This does not introduce a fixed debounce interval or remove final reviews.
Dispatchers should supply head_sha for reliable stale-at-dispatch detection;
legacy names identify a PR but do not prove which head the prompt describes.
E2E signal
The daily reference still executes all shards and keeps original verdicts. Each shard uploads Playwright JSON and publishes expected, skipped, unexpected, flaky and startup-error counts with the failing test names/messages. Targeted PR and manual coverage remain available. This change does not fix the historically red tests or pretend they pass.
config/e2e-failure-tracking.json can separate an evidenced repeated failure
from new failures in the summary. Each entry must have exact file, full
title, project, a nonempty stable message substring, an @owner, a linked
repository issue, and an ISO expires review date. Expired/malformed entries
are ignored and reported; changed error signatures appear as untracked. Entries
never skip a test or change its exit status. The initial list is empty because
the analysis established red workflows but did not establish owners and
reproductions for individual failures. Do not blanket-baseline an entire red run.
Capacity measurements and acceptance
CI runner demand runs daily at 04:23 UTC and can be dispatched manually. It
reads the previous 24 complete hours in hourly pages, samples up to six runs per
workflow/outcome stratum, and fetches job pages with bounded concurrency. An
hour exceeding the API's 1,000-result search cap fails visibly. The report and
raw evidence are retained for 30 days. No extra runner pool is provisioned.
The report measures the full job durations of runs created in the window, not occupancy clipped to the window: earlier runs that overlap it are excluded, and completed sampled jobs may finish after it. This matches the baseline cohort method. Workflow IDs keep ref-qualified paths in one sampling stratum.
The report shows weighted runner-hours and cancelled-run hours per workflow,
runner-minutes per completed PR run, and weighted queue/provisioning p95 per
runner label. It counts latest attempts only, excludes incomplete jobs, and
retains zero-job observations. It does not measure other repositories competing
for organization capacity. Compare equivalent traffic windows, not raw totals
alone. The collector needs only contents: read and actions: read.
After a week, compare runner-minutes per PR run, cancellation occupancy and queue p95 in each affected pool. Count newly added planning/reference overhead. A 25–35% overall reduction remains an experiment target, not an achieved result; selection, coalescing and matrix reductions overlap and cannot simply be added.