mirror of
https://github.com/stablyai/orca.git
synced 2026-10-01 08:01:56 +00:00
A concurrency slot is charged per job, not per core, and the account's cap is the scarce resource: standard runner minutes are free and unlimited on a public repository. Two paths spent slots that bought nothing. The unit matrix ran eight fixed shards averaging 6.5 minutes each, 3384 job-slots a day and 68% of all slot demand, while the arm pool queued 10.5 minutes at p95 — the queue was the oversharding. Five shards run the same work in ~10.5 minutes each for three fewer slots per run. Bun profile persistence escalated to all six platforms on `config/`, `resources/` and `.github/` wholesale, which took 36.5% of the last 1100 commits through the full matrix where a platform-flavoured predicate takes 19%. A pull request now qualifies one platform unless the change is platform- flavoured, and the push to main re-qualifies all six, so an unescalated miss surfaces minutes after merge rather than at the next cron. Missing changed-file evidence and an unavailable dependency graph still fail closed to all six.
149 lines
9.5 KiB
Markdown
149 lines
9.5 KiB
Markdown
# CI demand rollout
|
||
|
||
This implements the September 28 runner-demand analysis. The baseline inventory
|
||
covered September 27 04:00–September 28 04:00 UTC: 4,028 workflow runs, with
|
||
463 stratified job samples. Estimated occupancy was 1,081 runner-hours, dominated
|
||
by PR unit shards (414 hours) and Bun qualification (262 hours). These are
|
||
sampled sums of job durations across different runner pools, not billing totals
|
||
or a guaranteed forecast of savings.
|
||
|
||
## What runs now
|
||
|
||
| Work | Ordinary draft update | Ready PR / final checks | Main reference |
|
||
| ---------------------------- | ------------------------------------------------------------------------------ | --------------------------------------------- | --------------------------------------- |
|
||
| Static analysis and types | Immediately | Immediately | Existing workflows |
|
||
| Unit suite | Full, with shadow selection evidence | Full | Existing daily Node 24/26 x86 suite |
|
||
| Packages | After successful static analysis and types | Same | Existing release workflows |
|
||
| Bun persistence | Linux x64 for ordinary runtime changes; all six platforms for sensitive inputs | All six platforms when Bun inputs changed | Full nightly qualification at 11:30 UTC |
|
||
| Bun glibc/musl qualification | Sensitive inputs only, after persistence succeeds | Both architectures after persistence succeeds | Both architectures |
|
||
| E2E | Existing targeted routing | Existing targeted routing | One complete run at 17:00 UTC |
|
||
|
||
Bun-sensitive inputs include root/toolchain files, configuration, native code,
|
||
resources, platform-specific paths, persistence, SQLite, orcad, providers,
|
||
daemon, SSH, relay, and child-process code. Dependency discovery failure, a
|
||
missing event or an incomplete diff retains full qualification. An unrelated
|
||
change still skips Bun through the existing dependency classifier. No native
|
||
artifact is shared across platforms or ABIs. The routine draft reduction is an
|
||
explicit coverage-placement change; it does not assert identical per-update
|
||
coverage. Every non-draft synchronize event and the ready-for-review event
|
||
restores full qualification.
|
||
|
||
Expensive PR jobs wait for static/type success. This reduces fan-out for failed
|
||
or rapidly superseded commits without sleeping on a runner. Successful isolated
|
||
PRs pay the extra stage latency. Existing per-PR cancellation remains in place.
|
||
Package assertions, native boundaries, SSH/folder coverage, cache warming and
|
||
slow-test assertions are retained.
|
||
|
||
## Unit selection rollout
|
||
|
||
`ci-unit-plan.mjs` discovers the same include/exclude set as Vitest and follows
|
||
static imports, re-exports, literal dynamic imports, CommonJS requires and the
|
||
renderer aliases. Consumers of indirect filesystem/process inputs remain in the
|
||
candidate set, as do script/tool tests. Global configuration changes, deletions,
|
||
renames involving removed paths, unknown inputs and graph failures run the full
|
||
suite. A shard verifies the plan's source SHA and complete discovery list before
|
||
using it. Missing/stale artifacts fall back to full coverage, even if that means
|
||
running the suite on fewer shards.
|
||
|
||
The initial policy is **shadow**, with every shard retained. A full run spends
|
||
`FULL_SHARD_COUNT` (five) concurrency slots: the account is charged a slot per job
|
||
rather than per core, so eight 6.5-minute shards made this matrix 68% of daily slot
|
||
demand while the arm pool queued 10.5 minutes at p95. The first
|
||
local inventory matched Vitest exactly (9,950 files at validation); representative
|
||
source changes retained roughly 88% of files because of indirect input readers.
|
||
That is evidence for conservative coverage, not evidence of the analysis's
|
||
hypothetical 50% unit-work reduction. Improvements to indirect dependency
|
||
modeling should be demonstrated against full results before expanding selection.
|
||
|
||
Every shard uploads `unit-selection.json`, `unit-timings.json` (including module
|
||
outcomes), and its assignment. The evidence job combines these into
|
||
`unit-selection-review-attempt-N/selection-review.json`, reporting:
|
||
|
||
- Whether every discovered file appeared once across a complete reference run.
|
||
- Failures outside the candidate set, including failures in otherwise red runs.
|
||
- Measured worker time that selection would omit; worker times overlap and are
|
||
not runner occupancy or a prediction of wall-clock savings.
|
||
|
||
Missing/duplicate shards, stale plans, interrupted runs and unhandled errors do
|
||
not count as complete references. Diagnostic upload/report failures do not make
|
||
tests pass and do not independently fail successful tests.
|
||
|
||
After representative complete shadow runs show no missed failures, set repository
|
||
variable `ORCA_UNIT_SELECTION_MODE=selected` to enable selection **only for draft
|
||
PRs**. Keep full ready-PR checks and the daily compatibility suite. Inspect at
|
||
least a week's evidence across renderer, main, shared, SSH and fixture changes
|
||
before promotion, including red runs rather than only successful examples.
|
||
Unknown variable values retain shadow mode. Unset the variable or set it to
|
||
`shadow` to roll back immediately. Selected runs use one to eight timing-balanced
|
||
shards based on retained work.
|
||
|
||
Ready-for-review result reuse includes `unit full` in its source/workflow
|
||
identity. A green selected draft cannot satisfy the final full check, even if
|
||
the repository variable changes between runs. An already successful _full_
|
||
identical-source check can still be reused.
|
||
|
||
To inspect downloaded shard artifacts locally:
|
||
|
||
```sh
|
||
node config/scripts/ci-unit-selection-review.mjs ARTIFACT_DIRECTORY
|
||
```
|
||
|
||
## Review automation
|
||
|
||
Pullfrog recognizes the existing `Review #N [id]` and
|
||
`Review new commits on #N [id]` dispatch names. Explicit dispatchers may provide
|
||
`pull_request_number` and `head_sha`. Explicit PR identities share concurrency at the workflow boundary. Legacy review
|
||
names use a bounded lookup of the latest 100 dispatches and cancel only lower
|
||
run IDs for the same PR; a delayed older scope cannot cancel a newer review.
|
||
Unrecognized tasks are never grouped. The scope job alone has Actions write
|
||
permission for ordered cancellation. Closed PRs and explicitly stale heads
|
||
are skipped. A second head check prevents starting an agent after its queued
|
||
head has changed. Unrecognized agent tasks remain independent; lookup failures
|
||
also retain an independent task rather than cancelling unrelated work.
|
||
|
||
This does not introduce a fixed debounce interval or remove final reviews.
|
||
Dispatchers should supply `head_sha` for reliable stale-at-dispatch detection;
|
||
legacy names identify a PR but do not prove which head the prompt describes.
|
||
|
||
## E2E signal
|
||
|
||
The daily reference still executes all shards and keeps original verdicts. Each
|
||
shard uploads Playwright JSON and publishes expected, skipped, unexpected, flaky
|
||
and startup-error counts with the failing test names/messages. Targeted PR and
|
||
manual coverage remain available. This change does not fix the historically
|
||
red tests or pretend they pass.
|
||
|
||
`config/e2e-failure-tracking.json` can separate an evidenced repeated failure
|
||
from new failures in the summary. Each entry must have exact `file`, full
|
||
`title`, `project`, a nonempty stable `message` substring, an `@owner`, a linked
|
||
repository `issue`, and an ISO `expires` review date. Expired/malformed entries
|
||
are ignored and reported; changed error signatures appear as untracked. Entries
|
||
never skip a test or change its exit status. The initial list is empty because
|
||
the analysis established red workflows but did not establish owners and
|
||
reproductions for individual failures. Do not blanket-baseline an entire red run.
|
||
|
||
## Capacity measurements and acceptance
|
||
|
||
`CI runner demand` runs daily at 04:23 UTC and can be dispatched manually. It
|
||
reads the previous 24 complete hours in hourly pages, samples up to six runs per
|
||
workflow/outcome stratum, and fetches job pages with bounded concurrency. An
|
||
hour exceeding the API's 1,000-result search cap fails visibly. The report and
|
||
raw evidence are retained for 30 days. No extra runner pool is provisioned.
|
||
|
||
The report measures the full job durations of runs **created** in the window,
|
||
not occupancy clipped to the window: earlier runs that overlap it are excluded,
|
||
and completed sampled jobs may finish after it. This matches the baseline
|
||
cohort method. Workflow IDs keep ref-qualified paths in one sampling stratum.
|
||
|
||
The report shows weighted runner-hours and cancelled-run hours per workflow,
|
||
runner-minutes per completed PR _run_, and weighted queue/provisioning p95 per
|
||
runner label. It counts latest attempts only, excludes incomplete jobs, and
|
||
retains zero-job observations. It does not measure other repositories competing
|
||
for organization capacity. Compare equivalent traffic windows, not raw totals
|
||
alone. The collector needs only `contents: read` and `actions: read`.
|
||
|
||
After a week, compare runner-minutes per PR run, cancellation occupancy and
|
||
queue p95 in each affected pool. Count newly added planning/reference overhead.
|
||
A 25–35% overall reduction remains an experiment target, not an achieved result;
|
||
selection, coalescing and matrix reductions overlap and cannot simply be added.
|