Files
orca/.github/workflows/pr.yml
T
Neil a5796ec8eb refactor(runtime): split OrcaRuntimeService and compatibility tests (#17605)
* refactor(runtime): split OrcaRuntimeService into focused modules

* test(runtime): cover admission tiers and strict worktree reconciliation

* fix(runtime): preserve owner and structured session visibility

* fix(runtime): port post-extraction compatibility fixes

* fix(runtime): preserve skill-share cancellation barrier

* test(runtime): update identity inventory after extraction

* fix(runtime): preserve hook transport environment cleanup

* fix(runtime): consolidate idle probe imports

* test(runtime): retire split file process allowlist entry

* fix(runtime): route child process types through shared boundary

* test(runtime): preserve worktree host metadata precedence

* fix(runtime): update extracted test seams

* fix(runtime): gate the split's ts-nocheck set and restore the stop-confirmed contract

Audit follow-ups for the OrcaRuntimeService split:

- Freeze the 171 @ts-nocheck files behind a ratchet so no new file can disable
  type checking. The split's linear mixin chain cannot express forward
  references yet, so the existing suppressions are grandfathered; the baseline
  may only shrink.
- Drop the stray @ts-nocheck at the end of orca-runtime-get-status.ts. It sat
  after the first statement, where TypeScript ignores it, so the module was
  already checked.
- Restore `retireRejectedPty(ptyId, stopConfirmed: boolean)` as a required
  argument. The split widened it to optional and patched the resulting error
  with `stopConfirmed === true`; an omitted argument would have silently taken
  the unverified-stop path instead of failing to compile.
- Guard that every orca-runtime-tests fragment is imported by the compatibility
  entrypoint. The fragments are .spec.ts, which no Vitest include glob matches,
  so one left out of the list would silently stop running.

* fix(runtime): restore four behaviors the OrcaRuntimeService split dropped

Audit findings against the refactor's true base (ad5ba2572e):

- retirePtyAgentLaunchAuthority collected pane keys after deleting the
  restored-authority receipt instead of before it. collectPaneKeysForPty reads
  that receipt, so a receipt-only pane lost its key and never had its agent-hook
  compatibility authority retired. on-pty-exit.ts already carried a comment
  naming this exact invariant.
- The PTY-exit path kept orchestrationMailboxNotifications.retirePty but lost
  the loop that schedules a debounced mail-pointer repoint for the dead pty's
  terminal handle and any run bound to its panes. Restores the schedule call
  count to 7, matching base.
- subscribeToPtyExit lost isPtyKnownExited's leaf fallback and its
  post-registration lifecycle-generation recheck. leavesByPtyId is rebuilt from
  the renderer graph independently of ptysById, so a leaf can outlive its pty
  record; without the fallback a caller waiting on an already-dead pty never
  gets released.
- The chain root declared `[key: string]: unknown`, which base had nowhere. It
  leaked through the exported runtime type into every consumer, so any misspelled
  member access typechecked as unknown instead of erroring, and it accounted for
  957 of the suppressed errors. Removing it costs zero type errors.

* fix(runtime): restore escalation prose and unscoped automation publication

Two more behaviors the split dropped, each with a regression test that fails
against the pre-fix code:

- The worker-exit escalation stopped deriving its title through
  buildOrchestrationTaskDisplayMetadata and inlined `task.spec` instead. That
  ignored an explicit task_title, dropped the single-line normalization and the
  80-character bound, and turned the no-spec case into a quoted, duplicated id.
  A multi-paragraph spec landed verbatim in the coordinator's banner. The
  existing 11 tests all use short single-line specs, where the derived title and
  the raw spec are identical, so none of them could see it.
  Also reverts an added `if (!handle) return` guard: the dispatch lookup is
  deliberately keyed on the pane as well, because a reminted handle no longer
  matches the row while the pane identity outlives the remint.
- updateAutomation stopped going through automationChangePublications and
  published `source` unconditionally while gating the fallback on a non-null
  destination. A destination the store can no longer name then published only
  the stale source, so subscribers scoped elsewhere kept rendering a row that
  had left them — the exact case the helper documents. The helper had been left
  with zero callers; all three sites use it again.

* fix(skills): stop swallowing lookup errors and hard-erroring on non-ssh hosts

Follow-ups from auditing the skill install path against the refactor's base:

- resolveWorktree wrapped showManagedWorktree in `.catch(() => null)`, so a
  transient git or IO failure surfaced to the user as
  skill-install-workspace-not-found with the real cause discarded. Errors
  propagate again; a genuine id mismatch still returns null.
- resolveSkillSshTarget threw skill-install-workspace-host-unavailable when the
  execution host was neither local nor ssh, on both the repo and folder
  branches. Base gated these on connectionId, so a runtime-owned repo simply
  was not an SSH install and fell through to the local path. Both return null
  again, and the error code the split invented is now unreferenced.
- listManagedSkillInstalls awaited the receipt walk and the worktree resolve in
  sequence. They are independent and either can hit disk, WSL, or an SSH scan,
  so Promise.all is restored.

Deliberately unchanged: resolving the worktree through listResolvedWorktrees
rather than showManagedWorktree, which disambiguates a worktree id colliding
across hosts and is covered by its own test, and the SSH-folder
skill-install-ssh-dispatch-required throw, which matches the repo branch.

* fix(runtime): merge duplicate worktree-logic imports

The #17448 port added a third import from ../ipc/worktree-logic, which the
code-quality oxlint config rejects under --deny-warnings. Plain oxlint does not
flag it, so it only surfaced in CI's static analysis job.

* ci: run the ts-nocheck ratchet in PR checks

pr-workflow-lint-parity requires every leaf command in `pnpm lint` to have a
matching step in pr.yml. The ratchet was wired into lint but not the workflow,
so PR CI would not have enforced it.

* Merge remote-tracking branch 'origin/main' and retry the paired-host launch evaluate

main advanced 9 commits; none touch the orca-runtime.ts this branch splits, so
nothing needed porting.

CI failed twice on `Execution context was destroyed` thrown from
headless-paired-runtime-host's first `evaluate` after launch — a different spec
each run, which is the signature of the flake #17780 describes rather than a
regression. That commit added retryTransientMainEvaluate and adopted it in five
helpers but not this call site, even though its docblock names exactly this
case: the first evaluate after electron.launch() resolves, before the app is
ready. Wrapped it the same way.
2026-08-31 19:34:55 -07:00

981 lines
44 KiB
YAML

name: PR Checks
on:
pull_request:
types:
- opened
- synchronize
- reopened
- ready_for_review
concurrency:
group: pr-checks-${{ github.event.pull_request.number }}
cancel-in-progress: true
permissions:
contents: read
jobs:
# Why: a README/docs-only PR used to start the full matrix (test shards,
# two package jobs, typecheck, git compat, xterm, shell contracts). Path
# filters on `on.pull_request` would drop the `verify` check entirely; this
# detector keeps verify as the required aggregate and skips the expensive jobs.
# Per-job outputs also skip git-compat/xterm/packaging/shell when those
# inputs are unchanged; empty diffs fail closed and run everything.
code_paths:
name: detect code-relevant changes
runs-on: ubuntu-latest
outputs:
should_run: ${{ steps.filter.outputs.should_run }}
native_cache_changed: ${{ steps.filter.outputs.native_cache_changed }}
static_analysis: ${{ steps.filter.outputs.static_analysis }}
typecheck: ${{ steps.filter.outputs.typecheck }}
git_compatibility: ${{ steps.filter.outputs.git_compatibility }}
codex_index_heal_contract: ${{ steps.filter.outputs.codex_index_heal_contract }}
xterm_patch_sync: ${{ steps.filter.outputs.xterm_patch_sync }}
shell_contracts: ${{ steps.filter.outputs.shell_contracts }}
test: ${{ steps.filter.outputs.test }}
orcad_browser: ${{ steps.filter.outputs.orcad_browser }}
cross-version-wire: ${{ steps.filter.outputs.cross-version-wire }}
managed_hook_node18: ${{ steps.filter.outputs.managed_hook_node18 }}
package: ${{ steps.filter.outputs.package }}
package_windows: ${{ steps.filter.outputs.package_windows }}
steps:
- name: Checkout
uses: actions/checkout@v6
with:
# Why blob:none: full history is needed for the merge-base diff, but historical
# file contents are not. Blobs are ~89% of this repo's pack, and Git fetches the
# few this job actually reads on demand.
fetch-depth: 0
filter: blob:none
persist-credentials: false
- name: Classify changed paths
id: filter
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
run: |
set -euo pipefail
# Why --no-renames: name-only rename detection can report only the destination.
# A code file moved under docs/ must still expose its code-side deletion.
CHANGED="$(git diff --name-only --no-renames --diff-filter=ACDMR --merge-base "$BASE_SHA" "$HEAD_SHA")"
echo "Changed paths:"
printf '%s\n' "$CHANGED"
printf '%s\n' "$CHANGED" | node config/scripts/pr-code-change-scope.mjs | tee -a "$GITHUB_OUTPUT"
static_analysis:
name: static analysis
needs: [code_paths]
if: needs.code_paths.outputs.static_analysis == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
# Why blob:none: full history is needed for the merge-base diff, but historical
# file contents are not. Blobs are ~89% of this repo's pack, and Git fetches the
# few this job actually reads on demand.
fetch-depth: 0
filter: blob:none
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: node
- name: Lint
run: pnpm exec oxlint --format github
- name: Enforce focused code-quality plugins
run: pnpm run audit:code-quality:native
- name: Enforce type-aware code-quality baseline
run: pnpm run audit:code-quality:type-aware
- name: Enforce changed-code quality
run: pnpm run check:code-quality:changed -- "${{ github.event.pull_request.base.sha }}"
- name: Enforce React Doctor on changed lines
run: pnpm run check:react-doctor:changed -- "${{ github.event.pull_request.base.sha }}"
- name: Check Zustand selector fan-out budget
run: pnpm run check:zustand-selector-fanout
- name: Check reliability gate manifest
run: pnpm run check:reliability-gates
- name: Check VM runtime rollback compatibility
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
run: |
if git diff --quiet --merge-base "$BASE_SHA" "$HEAD_SHA" -- \
src/shared/ephemeral-vm-runtime-store.ts \
src/shared/ephemeral-vm-runtime-feature-store.ts \
src/shared/ephemeral-vm-runtime-rollback-projection.ts \
src/shared/ephemeral-vm-runtimes.ts \
src/shared/ephemeral-vm-recipes.ts \
src/shared/orca-yaml-hook-types.ts \
src/main/ephemeral-vm-runtime-service.ts \
src/main/ephemeral-vm-runtime-provisioning-persistence.ts \
src/main/ephemeral-vm-failed-start-cleanup.ts; then
echo "VM runtime persistence is unchanged."
exit 0
fi
node config/scripts/run-ephemeral-vm-runtime-store-rollback-repro.mjs \
config/scripts/ephemeral-vm-runtime-store-cross-version.test.ts
- name: Enforce max-lines ratchet
run: pnpm run check:max-lines-ratchet
- name: Enforce ts-nocheck ratchet
run: pnpm run check:ts-nocheck-ratchet
- name: Enforce runtime Electron-import ratchet
run: pnpm run check:runtime-electron-ratchet
# Why both: the ratchet proves nothing reachable from the runtime imports electron,
# which is a property of the import graph. This proves the Node artifact it enables
# actually boots, pairs, creates a worktree and round-trips a real PTY.
- name: Boot orcad and round-trip a terminal
run: pnpm run smoke:orcad-terminal
- name: Verify bundled skill guides
run: pnpm run verify:bundled-skill-guides
- name: Verify skill freshness manifest
run: pnpm run verify:skill-bundle-manifest
- name: Verify localization catalog
run: pnpm run verify:localization-catalog
# Why: extraction writes sorted evidence to an isolated temporary path,
# so feature PRs need one normalized AST pass rather than a three-OS matrix.
- name: Verify localization extraction
run: pnpm run verify:localization-extraction
- name: Verify localization coverage
run: pnpm run verify:localization-coverage
# Why: project-owned type declarations must live in .ts so tsc
# actually checks them. TypeScript's skipLibCheck: true (inherited
# from @electron-toolkit/tsconfig) silently widens unresolved names
# in .d.ts to `any`, which is how #1186 shipped a broken IPC signature
# past typecheck. See .github/CONTRIBUTING.md#type-declarations-prefer-ts-over-dts.
- name: Guard against project-owned .d.ts in preload/shared
run: |
matches=$(find src/preload src/shared -name '*.d.ts' 2>/dev/null || true)
if [ -n "$matches" ]; then
echo "::error::Project-owned .d.ts files are not allowed under src/preload or src/shared."
echo "Move type declarations into a .ts file so skipLibCheck does not hide errors."
echo "See .github/CONTRIBUTING.md#type-declarations-prefer-ts-over-dts."
echo "Found:"
echo "$matches"
exit 1
fi
- name: Check feature wall asset budget
run: pnpm check:feature-wall-assets
- name: Verify macOS entitlements
run: pnpm verify:macos-entitlements
root_directory_guard:
name: root directory guard
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
# Why blob:none: full history is needed for the merge-base diff, but historical
# file contents are not. Blobs are ~89% of this repo's pack, and Git fetches the
# few this job actually reads on demand.
fetch-depth: 0
filter: blob:none
persist-credentials: false
- name: Reject new root-level files and folders
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
run: node .github/scripts/check-root-directory-entries.mjs "$BASE_SHA" "$HEAD_SHA"
typecheck:
needs: [code_paths]
if: needs.code_paths.outputs.typecheck == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
# Why: every project is `composite`, so tsc already writes a .tsbuildinfo that lets
# the next run skip unchanged files. Share one cache entry across commits while the
# PR base stays stable; actions/cache keeps the first successful graph and the
# compiler still invalidates stale files from its content hashes.
- name: Cache TypeScript incremental state
uses: actions/cache@v5
with:
path: config/*.tsbuildinfo
key: tsbuildinfo-${{ runner.os }}-${{ hashFiles('pnpm-lock.yaml', 'config/tsconfig*.json') }}-${{ github.event.pull_request.base.sha }}
restore-keys: |
tsbuildinfo-${{ runner.os }}-${{ hashFiles('pnpm-lock.yaml', 'config/tsconfig*.json') }}-
- run: pnpm run typecheck
git_compatibility:
name: Git compatibility
needs: [code_paths]
if: needs.code_paths.outputs.git_compatibility == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
# Why: the 2.25.5 lane is a source build of a pinned tarball, so it produced the
# same binary on every PR for minutes of runner time. The key carries the version
# because that is the only input; the sha256 assertion below still guards the
# tarball on the miss path that actually builds.
- name: Cache baseline Git build
uses: actions/cache@v5
with:
path: ~/.cache/orca-git-compat/git-2.25.5
key: git-compat-baseline-${{ runner.os }}-${{ runner.arch }}-2.25.5
- name: Verify Git binary compatibility matrix
run: |
pids=()
(
archive="$RUNNER_TEMP/git-2.25.5.tar.gz"
source="$HOME/.cache/orca-git-compat/git-2.25.5"
if [ ! -x "$source/git" ]; then
curl -fsSL https://www.kernel.org/pub/software/scm/git/git-2.25.5.tar.gz -o "$archive"
echo "41662c52fc16fec4963bfc41075e71f8ead6b5e386797eb6f9a1111ff95a8ddf $archive" \
| sha256sum --check
mkdir -p "$source"
tar -xzf "$archive" -C "$source" --strip-components=1
make -C "$source" -j"$(nproc)" \
NO_GETTEXT=YesPlease NO_TCLTK=YesPlease NO_PYTHON=YesPlease git
# Why: the linked binaries are what the next run needs; the objects that
# produced them are most of the tree and would bloat the cache entry.
find "$source" -name '*.o' -delete
fi
ORCA_GIT_COMPAT_BINARY="$source/git" ORCA_GIT_COMPAT_VERSION="2.25.5" \
pnpm exec vitest run --config config/vitest.config.ts \
src/shared/git-binary-compatibility.test.ts
) &
pids+=("$!")
for spec in \
"alpine/git:edge-2.38.1|2.38.1" \
"alpine/git:v2.49.1|2.49.1"; do
(
image="${spec%%|*}"
version="${spec#*|}"
ORCA_GIT_COMPAT_IMAGE="$image" ORCA_GIT_COMPAT_VERSION="$version" \
pnpm exec vitest run --config config/vitest.config.ts \
src/shared/git-binary-compatibility.test.ts
) &
pids+=("$!")
done
status=0
for pid in "${pids[@]}"; do
wait "$pid" || status=1
done
exit "$status"
# Why this job: Orca's session index-heal depends on a Codex behavior — a
# `thread/read` of an unindexed rollout performs a read-repair that inserts the
# `threads` row. Every unit test drives a stub app-server and asserts only that the
# call did not error, so if Codex dropped the repair they would all stay green while
# the subsystem went inert. This runs the pinned real binary and fails when the
# repair stops happening. Pinned because the binary is the thing expected to drift.
codex_index_heal_contract:
name: Codex index-heal contract
needs: [code_paths]
if: needs.code_paths.outputs.codex_index_heal_contract == 'true'
runs-on: ubuntu-latest
env:
CODEX_CLI_VERSION: '0.150.1'
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
- name: Install pinned Codex CLI
run: |
set -euo pipefail
npm install --no-audit --no-fund --prefix "$RUNNER_TEMP/codex-cli" \
"@openai/codex@$CODEX_CLI_VERSION"
- name: Verify Codex index-heal contract
env:
# Why REQUIRED: without a binary the suite skips, and a job that skips
# reports success. This turns a failed or missing install into a red test
# instead of a green no-op.
ORCA_CODEX_CONTRACT_REQUIRED: '1'
ORCA_CODEX_CONTRACT_VERSION: ${{ env.CODEX_CLI_VERSION }}
run: |
set -euo pipefail
ORCA_CODEX_CONTRACT_BINARY="$RUNNER_TEMP/codex-cli/node_modules/.bin/codex" \
pnpm exec vitest run --config config/vitest.config.ts \
src/main/codex/codex-index-heal-binary-contract.test.ts
xterm_patch_sync:
name: xterm patch sync
needs: [code_paths]
if: needs.code_paths.outputs.xterm_patch_sync == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
# Why: the check rebuilds every package in the manifest from a pinned upstream
# commit — @xterm/xterm and the two addons, each built twice (once unmodified to
# prove the toolchain still reproduces the published bundles, once patched). Caching
# the npm metadata and the shallow clone keeps the repeated cost to the builds
# themselves; the key is the manifest, so a commit, package or toolchain bump
# invalidates it.
- name: Restore upstream xterm build inputs
uses: actions/cache@v5
with:
path: |
~/.npm
${{ runner.temp }}/xterm-patch-build/upstream/.git
key: xterm-upstream-${{ hashFiles('config/patches/xterm-upstream.json') }}
- name: Verify xterm patches match the pinned upstream build
env:
WORK_DIR: ${{ runner.temp }}/xterm-patch-build
run: node config/scripts/regenerate-xterm-patches.mjs --check --work-dir="$WORK_DIR"
shell_contracts:
name: shell contracts
needs: [code_paths]
if: needs.code_paths.outputs.shell_contracts == 'true'
runs-on: ubuntu-latest
# Why: this job's cost is almost entirely package download, and a stalled mirror has
# no wall-clock bound of its own. A successful run finishes in ~4.5 minutes, so this
# is generous; it exists so a wedge fails the job instead of holding the whole run
# open for the 6h GitHub default — which also blocks `gh run rerun --failed`.
timeout-minutes: 15
env:
# Why: the suites below gate their live fish tests on the binary, which is
# right on a developer machine and wrong here — this job is a required check
# and its fish lane is the only end-to-end guard for #9993, so a skip would
# report green with nothing exercised. Turns those skips into failures.
ORCA_REQUIRE_FISH: '1'
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
# Why fish: shell-ready.test.ts gates its live fish test on the binary being
# present, so without this the fish barrier is only covered by config-shape
# assertions and never actually exercised.
# Why release-4: DECSET 2031 arming lives in the fish 4.0 Rust tty_handoff, and
# fish-color-scheme-child-stdin.node-pty.test.ts (#9993) needs it. Noble ships
# 3.7, so the PPA is what makes that lane real.
- name: Install zsh and fish
run: |
# Why the update/PPA/fish steps are tolerant: a repo the runner image already
# ships can lack a Release file for this suite, and a failed add-apt-repository
# still leaves its list entry behind — either makes `apt-get update` exit
# non-zero and would red this required check over something unrelated to the
# PR. Every fish outcome is judged by the version gate below instead, so only
# the zsh install (which has no such gate) stays fatal here.
# Why retry only here: adding the PPA is the network-flaky step, and the
# version gate below is fatal, so a transient Launchpad blip would
# otherwise red a required check on PRs unrelated to shells.
# Why -n: add-apt-repository refreshes every configured repo on its own. With
# an update on each side of it this step refreshed them three times over, and
# the Azure archive mirror alone costs ~15-30s a pass. The PPA index is the
# only thing the repo list gains here, and the single update below fetches it.
# Why bound acquisition: measured on a *passing* run, this step spent 40s
# fetching 11.4 MB of index and then 2m17s fetching 8.9 MB of packages at
# 65 kB/s — it is dominated by download throughput, not by work. apt applies
# no wall-clock bound to a stalled mirror, so a slow Launchpad or archive
# host wedges the step for tens of minutes. This job is a required check, so
# a wedge holds the entire run open and blocks `gh run rerun --failed`.
# Bounded timeouts plus retries turn an unbounded hang into a fast, legible
# failure. Set in apt.conf.d rather than on each command line so the two
# invocations below stay exactly as pr-workflow-parallelism.test.mjs parses
# them. Retries are 1, not 3: a first attempt at these bounds already multiplied
# 30s x 3 retries across every index file into a ~15 minute stall on a dead
# mirror, which is worse than failing once and moving on.
sudo tee /etc/apt/apt.conf.d/99-orca-shell-contracts >/dev/null <<'APTCONF'
Acquire::http::Timeout "15";
Acquire::https::Timeout "15";
Acquire::Retries "1";
APTCONF
for attempt in 1 2 3; do
sudo add-apt-repository -y -n ppa:fish-shell/release-4 && break
echo "add-apt-repository attempt ${attempt} failed; retrying" >&2
sudo add-apt-repository -y -n -r ppa:fish-shell/release-4 || true
sleep 5
done
# Why a wall-clock bound on each command: apt's Acquire timeouts are per-connection,
# so a dead mirror costs timeout x retries x every index file. Measured: the archive
# mirror stalled with zero bytes and the step burned 14m26s before the job bound
# killed it. `timeout` is the only thing that bounds the command as a whole.
# The update is already tolerant by design (see above), so bounding it just caps
# what a dead mirror can cost before the install runs against whatever index exists.
timeout 120 sudo apt-get update || true
# Why both shells on one line: pr-workflow-parallelism.test.mjs parses only the
# first install command in this step to prove the lane really installs them.
timeout 300 sudo apt-get install -y zsh fish
# Separate from the install so the failure names the contract, not an apt error.
# ORCA_REQUIRE_FISH re-checks this at test time; this step just fails in seconds
# instead of after a full dependency install.
- name: Require fish 4+
run: |
version="$(fish --version 2>/dev/null || true)"
major="${version##*version }"
major="${major%%.*}"
case "$major" in '' | *[!0-9]*) major=0 ;; esac
echo "${version:-<fish not installed>}"
if [ "$major" -lt 4 ]; then
echo "::error::shell contracts needs fish 4+ (DECSET 2031 arming, #9993) but got '${version:-none}'. Fix the ppa:fish-shell/release-4 install rather than letting the fish lane skip." >&2
exit 1
fi
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: node
- name: Test real shell contracts
run: |
pnpm exec vitest run --config config/vitest.config.ts --maxWorkers=1 \
src/main/daemon/repro-13767-shell-ready-marker-lost-to-exec.test.ts \
src/main/daemon/shell-ready.test.ts \
src/main/daemon/node-pty-fd-leak.test.ts \
src/main/providers/local-pty-shell-ready-zsh-launch-environment.test.ts \
src/main/providers/__tests__/shell-ready-framework-example.test.ts \
src/main/pty/codex-shell-launch-preflight.test.ts \
src/main/pty/omp-shell-wrapper-alias-safety.test.ts \
src/main/pty/omp-shell-wrapper.node-pty.test.ts \
src/main/shell-startup-feature-channel.test.ts \
src/main/terminal-history-fish-session.node-pty.test.ts \
src/main/zsh-scoped-histfile.live-shell.test.ts \
src/main/zsh-startup-hook-user-config-equivalence.live-shell.test.ts \
src/main/zsh-wrapper-version-mismatch.live-shell.test.ts \
src/renderer/src/components/terminal-pane/fish-color-scheme-child-stdin.node-pty.test.ts \
src/shared/fish-query-reply-child-stdin.node-pty.test.ts \
src/shared/pty-reply-echo-shapes.node-pty.test.ts \
src/shared/startup-shell-portability.live-shell.test.ts \
src/shared/posix-command-path-lookup.test.ts
# Cache-key input changes would otherwise make every shard compile the same
# native addon concurrently. Prime the supported Node ABI before the matrix fans out.
test_native_cache:
name: prepare test native cache node 24
needs: [code_paths]
if: needs.code_paths.outputs.native_cache_changed == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: node
node-version: '24'
test:
needs: [code_paths, test_native_cache]
if: >-
always() &&
needs.code_paths.outputs.test == 'true' &&
(needs.test_native_cache.result == 'success' || needs.test_native_cache.result == 'skipped')
uses: ./.github/workflows/unit-tests.yml
with:
node_versions: '["24"]'
# Why a separate job: the test needs a real Chrome, and the sharded `test` matrix
# would pay for it on every shard to run one file in whichever shard it landed in.
orcad_browser:
name: orcad browser provider
needs: [code_paths]
if: needs.code_paths.outputs.orcad_browser == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
# Why no native-runtime: the provider drives the prebuilt agent-browser binary
# shipped in node_modules and never touches node-pty.
- uses: ./.github/actions/install-node-dependencies
# Why the runner's Google Chrome and not its chromium: Ubuntu 24.04 only ships an
# AppArmor userns profile for the Chrome .deb, so chromium dies with "No usable
# sandbox" and the provider passes no --no-sandbox. Why fail instead of skip: an
# unset ORCA_BROWSER_EXECUTABLE is exactly how this test went uncovered for so long.
- name: Resolve Chrome for the browser provider
run: |
set -euo pipefail
chrome="$(command -v google-chrome || command -v google-chrome-stable || true)"
if [ -z "$chrome" ]; then
echo "::error::No Google Chrome on the runner; the browser provider test would silently skip."
exit 1
fi
"$chrome" --version
echo "ORCA_BROWSER_EXECUTABLE=$chrome" >> "$GITHUB_ENV"
- name: Test external Chromium browser provider
run: |
pnpm exec vitest run --config config/vitest.config.ts \
src/main/orcad/external-chromium-browser-process.integration.test.ts
cross-version-wire:
name: cross-version wire compatibility
needs: [code_paths]
if: needs.code_paths.outputs.cross-version-wire == 'true'
runs-on: ubuntu-latest
steps:
# Why fetch-depth 0: the harness extracts the newest release tag to skew
# current code against it. The default shallow clone has no tags, which is
# why this cannot ride along in the sharded `test` job.
- name: Checkout
uses: actions/checkout@v6
with:
# Why blob:none: full history is needed for the merge-base diff, but historical
# file contents are not. Blobs are ~89% of this repo's pack, and Git fetches the
# few this job actually reads on demand.
fetch-depth: 0
filter: blob:none
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: node
# A path filter that matches nothing exits 1 ("No test files found"), so this
# lane cannot report success while running zero tests.
- name: Old/new client and server compatibility journeys
run: >-
pnpm exec vitest run --config config/vitest.config.ts
tests/e2e/cross-version-wire/release-checkout.unit.test.ts
tests/e2e/cross-version-wire/cross-version-browser-placement.unit.test.ts
tests/e2e/cross-version-wire/cross-version-terminal-wire.unit.test.ts
tests/e2e/cross-version-wire/reported-lossy-initial-snapshot.unit.test.ts
tests/e2e/cross-version-wire/cross-version-agent-session-wire.unit.test.ts
managed_hook_node18:
name: managed hooks on Node 18
needs: [code_paths]
if: needs.code_paths.outputs.managed_hook_node18 == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
- name: Build relay companions
run: pnpm run build:relay
- name: Setup Node 18 runtime
uses: actions/setup-node@v6
with:
node-version: '18'
- name: Smoke managed-hook companions
run: node config/scripts/smoke-managed-hook-runtime-node18.mjs
package:
name: package
needs: [code_paths]
if: needs.code_paths.outputs.package == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Cache electron-builder downloads
uses: actions/cache@v5
with:
path: ~/.cache/electron-builder
key: electron-builder-linux-${{ hashFiles('pnpm-lock.yaml') }}
restore-keys: |
electron-builder-linux-
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: electron
# Why --no-file-parallelism: every file here launches a full Electron stack twice, and each
# probe carries its own in-process deadline. Four at once on a 4-vCPU runner starve each other
# past those deadlines; serial, every probe owns the runner.
- name: Test Linux Electron lifecycle boundary
run: >-
xvfb-run --auto-servernum pnpm exec vitest run --config config/vitest.config.ts
--no-file-parallelism
src/main/browser/browser-client-page-renderer-lifecycle.electron.test.ts
src/main/browser/browser-route-tcp-egress.electron.test.ts
src/main/browser/browser-route-webrtc-egress.electron.test.ts
src/main/browser/browser-route-h3-egress.electron.test.ts
src/main/browser/browser-route-dns-prefetch.electron.test.ts
- name: Build package inputs
run: |
status=0
pnpm run build:cli || status=1
scripts=(build:relay build:electron-vite:parallel)
pids=()
for script in "${scripts[@]}"; do
pnpm run "$script" &
pids+=("$!")
done
for pid in "${pids[@]}"; do
wait "$pid" || status=1
done
exit "$status"
- name: Project web client from renderer build
run: pnpm run build:web-from-renderer
- name: Build native components
run: pnpm run build:native
- name: Package unpacked app
env:
ORCA_REUSE_PREPARED_NATIVE_RUNTIME: '1'
run: pnpm exec electron-builder --config config/electron-builder.config.cjs --linux AppImage --x64 --publish never
- name: Verify headless serve signal shutdown
run: node config/scripts/run-headless-serve-shutdown-docker.mjs --appimage dist/orca-linux.AppImage
- name: Smoke packaged CLI
run: node config/scripts/smoke-packaged-cli.mjs --app-dir=dist/linux-unpacked
- name: Smoke packaged hang watchdog worker
run: xvfb-run --auto-servernum node config/scripts/smoke-packaged-hang-watchdog-worker.mjs --app-dir=dist/linux-unpacked
package_windows:
name: package (windows)
needs: [code_paths]
if: needs.code_paths.outputs.package_windows == 'true'
runs-on: windows-2022
timeout-minutes: 30
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Cache electron-builder downloads
uses: actions/cache@v5
with:
path: |
~\AppData\Local\electron\Cache
~\AppData\Local\electron-builder\Cache
key: electron-builder-windows-${{ hashFiles('pnpm-lock.yaml') }}
restore-keys: |
electron-builder-windows-
# Why persist-native-cache false: this job later rebuilds the same path for
# Electron. A post-job save would store the Electron ABI under the Node key.
- uses: ./.github/actions/install-node-dependencies
id: deps
with:
native-runtime: node
persist-native-cache: 'false'
- name: Save compiled Node native modules
if: steps.deps.outputs.native-cache-hit != 'true'
uses: actions/cache/save@v5
with:
path: |
node_modules/.pnpm/node-pty@*/node_modules/node-pty/build
node_modules/.pnpm/windows-native-registry@*/node_modules/windows-native-registry/build
node_modules/.pnpm/@vscode+windows-process-tree@*/node_modules/@vscode/windows-process-tree/build
key: native-modules-${{ runner.os }}-${{ steps.deps.outputs.native-cache-scope }}-${{ runner.arch }}-node-node${{ steps.deps.outputs.node-version }}-${{ hashFiles('pnpm-lock.yaml', '.github/actions/install-node-dependencies/action.yml', 'config/scripts/ensure-native-runtime.mjs', 'config/scripts/rebuild-native-deps.mjs', 'config/patches/node-pty@1.1.0.patch', 'config/patches/@vscode__windows-process-tree@0.8.0.patch') }}
- name: Test Windows-specific boundaries
run: >-
pnpm exec vitest run --config config/vitest.config.ts
config/scripts/rebuild-native-deps.test.mjs
src/main/browser/browser-client-page-renderer-lifecycle.electron.test.ts
src/main/browser/browser-route-tcp-egress.electron.test.ts
src/main/browser/browser-route-webrtc-egress.electron.test.ts
src/main/browser/browser-route-h3-egress.electron.test.ts
src/main/browser/browser-route-dns-prefetch.electron.test.ts
src/main/providers/windows-conpty-wide-char-duplication.node-pty.test.ts
src/main/providers/pty-repaint-wide-char-buffer.node-pty.test.ts
src/shared/child-process/windows-command-line.win32.test.ts
src/main/agent-hooks/windows-hook-payload-delivery.test.ts
src/main/windows/windows-pty-job.win32.test.ts
src/main/windows/windows-host-job.win32.test.ts
src/main/wsl/wsl-runner.test.ts
src/main/wsl/wsl-guest-environment.test.ts
src/main/wsl/wsl-invocation-boundary.test.ts
src/main/wsl/wsl-executable-path.win32.test.ts
src/main/wsl/wsl-w1-w3-contract.test.ts
src/shared/source-scan/source-tree-scan.test.ts
src/main/cli/wsl-cli-powershell-boundary.test.ts
src/main/cursor/hook-service.test.ts
src/main/orca-profiles/profile-index-store.test.ts
src/main/runtime/repo-worktree-admin-fingerprint.test.ts
src/main/runtime/worktree-scan-admin-fingerprint-gate.test.ts
src/shared/secure-file-fsync-flags.test.ts
src/main/ipc/pty-codex-account-attribution.test.ts
src/main/ipc/pty-spawn-env-codex-resume-provenance.test.ts
# Why the :parallel variant: identical to build:release except the three
# electron-vite targets overlap instead of running back to back. The Linux package
# job already packages and smoke-tests an AppImage built that way.
- name: Cache Windows CLI launcher
uses: actions/cache@v5
with:
path: native/windows-cli-launcher/.build
key: windows-cli-launcher-${{ runner.os }}-${{ runner.arch }}-${{ hashFiles('native/windows-cli-launcher/**', 'config/scripts/build-windows-cli-launcher.mjs') }}
- name: Build package inputs
env:
ORCA_REUSE_WINDOWS_CLI_LAUNCHER: '1'
run: pnpm run build:release:parallel
- name: Restore compiled Electron native modules
uses: actions/cache@v5
with:
path: |
node_modules/.pnpm/node-pty@*/node_modules/node-pty/build
node_modules/.pnpm/windows-native-registry@*/node_modules/windows-native-registry/build
node_modules/.pnpm/@vscode+windows-process-tree@*/node_modules/@vscode/windows-process-tree/build
key: native-modules-${{ runner.os }}-${{ steps.deps.outputs.native-cache-scope }}-${{ runner.arch }}-electron-node${{ steps.deps.outputs.node-version }}-${{ hashFiles('pnpm-lock.yaml', '.github/actions/install-node-dependencies/action.yml', 'config/scripts/ensure-native-runtime.mjs', 'config/scripts/rebuild-native-deps.mjs', 'config/patches/node-pty@1.1.0.patch', 'config/patches/@vscode__windows-process-tree@0.8.0.patch') }}
- name: Prepare Electron native runtime
run: node config/scripts/ensure-native-runtime.mjs --runtime=electron
- name: Package unpacked app
env:
ORCA_REUSE_PREPARED_NATIVE_RUNTIME: '1'
run: pnpm exec electron-builder --config config/electron-builder.config.cjs --dir
- name: Smoke packaged Windows PTY native capability
run: pnpm run smoke:windows-pty-native-capability -- --exe=dist/win-unpacked/Orca.exe
- name: Smoke packaged CLI
run: node config/scripts/smoke-packaged-cli.mjs --app-dir=dist/win-unpacked
# Why: PR E2E is advisory and only validates changed specs; scheduled and
# release runs retain full-suite coverage.
e2e-paths:
name: detect changed e2e specs
needs: [code_paths]
runs-on: ubuntu-latest
if: github.event.pull_request.draft != true && needs.code_paths.outputs.should_run == 'true'
# Why: detector only needs to read the checkout; do not inherit repo defaults.
permissions:
contents: read
outputs:
should_run: ${{ steps.filter.outputs.should_run }}
test_files: ${{ steps.filter.outputs.test_files }}
ssh_source_changed: ${{ steps.filter.outputs.ssh_source_changed }}
native_ime_source_changed: ${{ steps.filter.outputs.native_ime_source_changed }}
steps:
- name: Checkout
uses: actions/checkout@v6
with:
# Why blob:none: full history is needed for the merge-base diff, but historical
# file contents are not. Blobs are ~89% of this repo's pack, and Git fetches the
# few this job actually reads on demand.
fetch-depth: 0
filter: blob:none
persist-credentials: false
- name: Filter changed E2E specs
id: filter
run: |
set -euo pipefail
BASE="${{ github.event.pull_request.base.sha }}"
HEAD="${{ github.event.pull_request.head.sha }}"
CHANGED="$(git diff --name-only --diff-filter=AMCR --merge-base "$BASE" "$HEAD")"
# Source routes are executable contracts so a test can prove exact
# authorities, exclusions, and sentinels without evaluating workflow shell.
TEST_FILES_JSON="$(printf '%s\n' "$CHANGED" | node config/scripts/pr-e2e-source-routing.mjs)"
echo "test_files=$TEST_FILES_JSON" >> "$GITHUB_OUTPUT"
# Why a separate signal: the Docker-SSH lane must trigger on SSH source, not on a
# spec name surviving in a route's list. Same routes, so the two cannot drift.
SSH_SOURCE_CHANGED="$(printf '%s\n' "$CHANGED" | node config/scripts/pr-e2e-source-routing.mjs --ssh-source)"
echo "ssh_source_changed=$SSH_SOURCE_CHANGED" >> "$GITHUB_OUTPUT"
echo "SSH source changed: $SSH_SOURCE_CHANGED"
# Why its own signal: the real-IME lane is a whole ibus session, not a spec, so it must
# trigger on IME source rather than on a spec name in some route's list.
NATIVE_IME_SOURCE_CHANGED="$(printf '%s\n' "$CHANGED" | node config/scripts/pr-e2e-source-routing.mjs --native-ime-source)"
echo "native_ime_source_changed=$NATIVE_IME_SOURCE_CHANGED" >> "$GITHUB_OUTPUT"
echo "Native IME source changed: $NATIVE_IME_SOURCE_CHANGED"
if [ "$TEST_FILES_JSON" != '[]' ]; then
echo "should_run=true" >> "$GITHUB_OUTPUT"
echo "Changed E2E specs: $TEST_FILES_JSON"
else
echo "should_run=false" >> "$GITHUB_OUTPUT"
echo "No changed E2E specs"
fi
e2e:
name: e2e
needs: e2e-paths
if: needs.e2e-paths.outputs.should_run == 'true'
# Why: reusable e2e.yml only checkouts, builds, and uploads artifacts.
permissions:
contents: read
uses: ./.github/workflows/e2e.yml
with:
test_files: ${{ needs.e2e-paths.outputs.test_files }}
ssh_source_changed: ${{ needs.e2e-paths.outputs.ssh_source_changed }}
# Why this is not in verify's needs: it is the first PR-gate run of a harness whose reliability
# is only known from nightly main runs (20/20 green, 2026-08-09..2026-08-29, p50 3m25s). It
# reports a red X on the PR without blocking, exactly like `e2e` above. Deliberately no
# continue-on-error: that renders the check green and hides the signal it exists to give. To
# make it blocking, add it to verify.needs, add TERMINAL_IME_NATIVE to the env below, and
# require `success || skipped` outside the strict loop — see the note on `e2e`.
terminal_ime_native:
name: real IME
needs: e2e-paths
if: needs.e2e-paths.outputs.native_ime_source_changed == 'true'
# Why: the reusable workflow only checks out, builds, and uploads artifacts.
permissions:
contents: read
uses: ./.github/workflows/terminal-ime-e2e.yml
verify:
if: always()
needs:
- code_paths
- static_analysis
- root_directory_guard
- typecheck
- git_compatibility
- codex_index_heal_contract
- xterm_patch_sync
- shell_contracts
- test
- orcad_browser
- cross-version-wire
- managed_hook_node18
- package
- package_windows
runs-on: ubuntu-latest
steps:
# Why: e2e is deliberately absent from needs. The suite is currently red on
# main (every scheduled run), so gating merges on it would block any PR that
# touches tests/e2e/** — including the ones fixing the suite. Until it is
# green the job runs and reports for E2E-path PRs without blocking. To flip
# it on: add `e2e` to needs, add E2E to the env below, and require
# `"$E2E" = success || skipped` after the loop — skipped is the normal
# result for a path-filtered job and must keep passing, so it has to be
# checked outside the loop or it would excuse the jobs above.
- name: Require successful checks
env:
CODE_PATHS: ${{ needs.code_paths.result }}
SHOULD_RUN: ${{ needs.code_paths.outputs.should_run }}
STATIC_ANALYSIS: ${{ needs.static_analysis.result }}
STATIC_ANALYSIS_SHOULD_RUN: ${{ needs.code_paths.outputs.static_analysis }}
ROOT_DIRECTORY_GUARD: ${{ needs.root_directory_guard.result }}
TYPECHECK: ${{ needs.typecheck.result }}
TYPECHECK_SHOULD_RUN: ${{ needs.code_paths.outputs.typecheck }}
GIT_COMPATIBILITY: ${{ needs.git_compatibility.result }}
GIT_COMPATIBILITY_SHOULD_RUN: ${{ needs.code_paths.outputs.git_compatibility }}
CODEX_INDEX_HEAL_CONTRACT: ${{ needs.codex_index_heal_contract.result }}
CODEX_INDEX_HEAL_CONTRACT_SHOULD_RUN: ${{ needs.code_paths.outputs.codex_index_heal_contract }}
XTERM_PATCH_SYNC: ${{ needs.xterm_patch_sync.result }}
XTERM_PATCH_SYNC_SHOULD_RUN: ${{ needs.code_paths.outputs.xterm_patch_sync }}
SHELL_CONTRACTS: ${{ needs.shell_contracts.result }}
SHELL_CONTRACTS_SHOULD_RUN: ${{ needs.code_paths.outputs.shell_contracts }}
TEST: ${{ needs.test.result }}
TEST_SHOULD_RUN: ${{ needs.code_paths.outputs.test }}
ORCAD_BROWSER: ${{ needs.orcad_browser.result }}
ORCAD_BROWSER_SHOULD_RUN: ${{ needs.code_paths.outputs.orcad_browser }}
CROSS_VERSION_WIRE: ${{ needs.cross-version-wire.result }}
CROSS_VERSION_WIRE_SHOULD_RUN: ${{ needs.code_paths.outputs.cross-version-wire }}
MANAGED_HOOK_NODE18: ${{ needs.managed_hook_node18.result }}
MANAGED_HOOK_NODE18_SHOULD_RUN: ${{ needs.code_paths.outputs.managed_hook_node18 }}
PACKAGE: ${{ needs.package.result }}
PACKAGE_SHOULD_RUN: ${{ needs.code_paths.outputs.package }}
PACKAGE_WINDOWS: ${{ needs.package_windows.result }}
PACKAGE_WINDOWS_SHOULD_RUN: ${{ needs.code_paths.outputs.package_windows }}
run: |
if [ "$CODE_PATHS" != "success" ]; then
exit 1
fi
if [ "$ROOT_DIRECTORY_GUARD" != "success" ]; then
exit 1
fi
if [ "$SHOULD_RUN" != "true" ]; then
echo "Docs-only change; expensive PR checks skipped."
fi
failed=0
check_job() {
local name="$1" result="$2" should="$3"
if [ "$should" = "true" ]; then
if [ "$result" != "success" ]; then
echo "$name: expected success, got $result"
failed=1
fi
else
if [ "$result" != "skipped" ]; then
echo "$name: expected skipped, got $result"
failed=1
fi
fi
}
# Require success when the PR has code-relevant changes
check_job static_analysis "$STATIC_ANALYSIS" "$STATIC_ANALYSIS_SHOULD_RUN"
check_job typecheck "$TYPECHECK" "$TYPECHECK_SHOULD_RUN"
check_job git_compatibility "$GIT_COMPATIBILITY" "$GIT_COMPATIBILITY_SHOULD_RUN"
check_job codex_index_heal_contract "$CODEX_INDEX_HEAL_CONTRACT" "$CODEX_INDEX_HEAL_CONTRACT_SHOULD_RUN"
check_job xterm_patch_sync "$XTERM_PATCH_SYNC" "$XTERM_PATCH_SYNC_SHOULD_RUN"
check_job shell_contracts "$SHELL_CONTRACTS" "$SHELL_CONTRACTS_SHOULD_RUN"
check_job test "$TEST" "$TEST_SHOULD_RUN"
check_job orcad_browser "$ORCAD_BROWSER" "$ORCAD_BROWSER_SHOULD_RUN"
check_job cross-version-wire "$CROSS_VERSION_WIRE" "$CROSS_VERSION_WIRE_SHOULD_RUN"
check_job managed_hook_node18 "$MANAGED_HOOK_NODE18" "$MANAGED_HOOK_NODE18_SHOULD_RUN"
check_job package "$PACKAGE" "$PACKAGE_SHOULD_RUN"
check_job package_windows "$PACKAGE_WINDOWS" "$PACKAGE_WINDOWS_SHOULD_RUN"
exit "$failed"