Files
orca/.github/workflows/unit-plan.yml
T
Neil ec9f35e2ee perf(ci): plan the unit shards before the static-analysis gate instead of behind it (#23743)
A caller's `needs` gate the whole called workflow, so while the plan job lived in
unit-tests.yml it could not start until static analysis and typecheck had both
finished and passed -- and the shard matrix then waited on it. The two hops were
serial when they did not need to be: planning reads the checkout, a git diff
against HEAD^1, the import graph and the checked-in timing baseline in
config/scripts/ci-shard-timings.json, and consumes nothing that static analysis,
typecheck or the native-cache primer produce.

Planning moves to its own reusable workflow so pr.yml can run it against
code_paths alone, overlapping it with the gate. Measured across 99 runs, the
shard matrix is created a median 93s earlier (p25 47s, p90 241s, never later).
Planning stays a required predecessor of the shards, so an empty assignment
cannot expand the matrix.

The gate itself is deliberately left in place. It fires on 22% of runs, and the
shard queue wait knees hard above ~9 concurrent ARM jobs -- 4s median below that
against 218s at 15-19 -- so admitting 8 doomed shards per failed run would cost
more in queue pressure than it returns in latency.

Cost is one 37s ubuntu-latest job, which does not touch the ARM pool the shards
contend for.

A planning failure still fails the PR: the shards are skipped, and verify's
check_job requires success whenever the classifier says tests should run, so it
reports `test: expected success, got skipped`.
2026-09-28 19:22:10 -07:00

49 lines
1.6 KiB
YAML

name: Unit plan
# Why its own workflow: the caller's `needs` gate the whole called workflow, so while this lived
# inside unit-tests.yml it waited on static analysis and typecheck before it could even start —
# and the shards then waited on it. Planning needs neither (its inputs are the checkout, a git
# diff against HEAD^1, the import graph, and the checked-in timing baseline), so hoisting it out
# lets it overlap the gate instead of queueing behind it. Measured: a median 93s off the shard
# matrix's start.
on:
workflow_call:
inputs:
selection_mode:
description: Shadow validates selection; selected applies it only to draft PRs.
required: false
default: shadow
type: string
outputs:
shards:
description: JSON array of shard assignments for the unit matrix.
value: ${{ jobs.plan.outputs.shards }}
permissions:
contents: read
jobs:
plan:
runs-on: ubuntu-latest
timeout-minutes: 5
outputs:
shards: ${{ steps.plan.outputs.shards }}
steps:
- uses: actions/checkout@v6
with:
fetch-depth: 2
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
- name: Plan unit selection
id: plan
env:
ORCA_UNIT_SELECTION_MODE: ${{ inputs.selection_mode }}
run: node config/scripts/ci-unit-plan.mjs
- uses: actions/upload-artifact@v7
continue-on-error: true
with:
name: unit-selection-attempt-${{ github.run_attempt }}
path: ci-shards/unit-selection.json
retention-days: 14