Files
orca/.github/workflows/pr.yml
T
Brennan Benson 020cebeff6 Add standalone Agent Client Protocol client layer (#24990)
* Add standalone ACP protocol client and session runtime

* Protect ACP transport teardown from late stream errors

* Retire incoming ACP request ids before publishing responses

* Narrow ACP configuration requests and transport message types

* Remove redundant ACP request handler return unions

* Keep ACP waits caller-owned and preserve protocol extensions

* Preserve open ACP decisions through prompt completion

* Generate open ACP enums and check the generated schema offline

A newer or vendor enum value (tool kind, tool status, option kind, stop
reason) no longer fails the whole message: generated enums accept the known
literals plus any other string, typed so callers can still narrow on the
known ones. The generated header now records the pinned input digests, the
generator digest and a body hash, so `verify:acp-protocol` catches a stale or
hand-edited file without network access; it runs in lint and the PR workflow.

* Land the ACP runtime contract the agent adapters use

- Deliver notifications other than session/update through
  onExtensionNotification, in arrival order with session updates.
- Accept _meta on prompt, setMode, setModel, setConfigOption and cancel.
- cancel() always sends session/cancel once the session runs, since the
  agent can be in a turn it began itself; only a successful send is shared,
  so a failed write is retried.
- Cancel aborts each open agent request's signal and lets its handler send
  its own answer; -32800 only when the handler rejects.
- Permission requests validate only the session, tool call id and options;
  unreadable fields are dropped with a diagnostic, and any answer Orca
  cannot send is `cancelled` instead of a JSON-RPC error. Agent-started
  turns may ask; whether to show it is the caller's decision.
- AcpAgentError marks the agent's own errors; AcpInvalidResponseError keeps
  the raw answer and validation issues for answers Orca could not read.
- Lines over the size limit are classified by prefix (shared with the Codex
  reader): the owed request fails, an oversized agent request is answered
  with an error, and an unattributable response closes the connection.

* Answer every agent request after an ACP cancel

A cancel that lands before a permission handler starts now still runs the
permission path, so the agent gets the `cancelled` outcome rather than a
request-cancelled error. A handler that ignores the abort no longer leaves
the agent waiting: once the abort has run through, any request still
unanswered gets request-cancelled. Handlers that answer on abort keep their
own reply.

Also renames a lint-rejected helper parameter, replaces a Reflect.apply in a
test, and stops the permission diagnostic from firing with an empty list.

* Let each ACP request handler own its answer after a cancel

Removes the next-event-loop-turn fallback that answered request-cancelled
for any handler still silent after a cancel. It raced answers that were
still being saved (an approval mid-journal-write reached the agent as an
error) and made the outcome depend on event-loop timing. The handler that
owns an agent request now always sends its answer, or throws for
request-cancelled; a request it never answers ends when the connection
closes. A permission whose handler had not started still answers
`cancelled`.

* Register the ACP schema verify step in the PR preflight phase test

* feat(acp): a steer's cancel asks once and never ends the agent

The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle,
then close the connection, which ends the agent. A steer used it too, so a slow agent lost its
process just because the person added a message. requestSteerCancel() now sends session/cancel
once per prompt, cancels the agent's open requests and answers later permissions cancelled, and
never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel()
stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel
paths move into acp-prompt-cancel.ts over one cancel channel.

* fix(acp): a repeated steer shares the cancel in flight; say what the caller owns

Per review: a second steer before the first write lands returns that write instead of resolving
early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that
fails instead must not take the steer until the caller rebuilds the session; the Stop's says a
prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives
the runtime a handler that would allow: the open permission's signal aborts and the late one never
reaches it.

* test(ratchet): require src/main/acp now that this PR lands it
2026-10-05 22:25:40 -07:00

1441 lines
68 KiB
YAML

name: PR Checks
run-name: "PR ${{ github.event.pull_request.number }} | source ${{ github.sha }} | workflow ${{ github.workflow_sha }} | unit ${{ github.event.pull_request.draft && vars.ORCA_UNIT_SELECTION_MODE == 'selected' && 'selected' || 'full' }}"
on:
pull_request:
types:
- opened
- synchronize
- reopened
- ready_for_review
concurrency:
group: pr-checks-${{ github.event.pull_request.number }}
cancel-in-progress: true
permissions:
contents: read
jobs:
# Why: a README/docs-only PR used to start the full matrix (test shards,
# two package jobs, typecheck, git compat, xterm, shell contracts). Path
# filters on `on.pull_request` would drop the `verify` check entirely; this
# detector keeps verify as the required aggregate and skips the expensive jobs.
# Per-job outputs also skip git-compat/xterm/packaging/shell when those
# inputs are unchanged; empty diffs fail closed and run everything.
code_paths:
name: detect changes and check repository guards
# Reuse one lightweight checkout for detection and the always-required guards.
runs-on: ubuntu-slim
timeout-minutes: 5
permissions:
contents: read
actions: read
outputs:
should_run: ${{ steps.filter.outputs.should_run }}
reused_run_id: ${{ steps.readiness.outputs.run_id }}
# A proven success masks required work only; advisory routing still uses the full diff.
mobile_dependencies: ${{ steps.filter.outputs.mobile_dependencies }}
mobile_web_app: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.mobile_web_app }}
static_analysis: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.static_analysis }}
typecheck: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.typecheck }}
git_compatibility: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.git_compatibility }}
codex_index_heal_contract: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.codex_index_heal_contract }}
xterm_patch_sync: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.xterm_patch_sync }}
shell_contracts: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.shell_contracts }}
test: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.test }}
orcad_browser: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.orcad_browser }}
cross-version-wire: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.cross-version-wire }}
managed_hook_node18: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.managed_hook_node18 }}
package: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.package }}
package_windows: ${{ steps.readiness.outputs.reused != 'true' && steps.filter.outputs.package_windows }}
e2e_should_run: ${{ steps.e2e_filter.outputs.should_run }}
test_files: ${{ steps.e2e_filter.outputs.test_files }}
e2e_run_changed: ${{ steps.e2e_filter.outputs.e2e_run_changed }}
e2e_needs_build: ${{ steps.e2e_filter.outputs.e2e_needs_build }}
ssh_source_changed: ${{ steps.e2e_filter.outputs.ssh_source_changed }}
native_ime_source_changed: ${{ steps.e2e_filter.outputs.native_ime_source_changed }}
wsl_source_changed: ${{ steps.e2e_filter.outputs.wsl_source_changed }}
steps:
- name: Checkout
uses: actions/checkout@v6
with:
# Why depth 50 and not 0: every diff below resolves to HEAD^1, so only the merge commit
# and a little slack are needed. Fetching all 8127 refs' commit graph cost ~20s here and
# is charged to the start of all 22 jobs, since they all need this one.
# Why blob:none stays: the sparse tree below is ~7 files, so there are no blobs to
# materialize and no promisor refetch. Measured 9.5s -> 1.6s against 20.7s today.
fetch-depth: 50
filter: blob:none
sparse-checkout: |
/config/scripts/git-pull-request-diff-base.mjs
/package.json
/README.md
/docs/readme/
/.github/scripts/check-root-directory-entries.mjs
/config/scripts/check-readme-local-links.mjs
/config/scripts/pr-code-change-scope.mjs
/config/scripts/pr-e2e-source-routing.mjs
/config/scripts/ci-e2e-job-selection.mjs
/config/scripts/pr-ready-check-reuse.mjs
sparse-checkout-cone-mode: false
persist-credentials: false
- name: Reject new root-level files and folders
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
run: |
DIFF_BASE="$(node config/scripts/git-pull-request-diff-base.mjs "$BASE_SHA")"
node .github/scripts/check-root-directory-entries.mjs "$DIFF_BASE" HEAD
# The full Git index retains link targets outside the sparse working tree.
- name: Check README local links
run: node config/scripts/check-readme-local-links.mjs
# Readiness changes eligibility for advisory tests, not the already-tested source.
- name: Find identical successful required checks
id: readiness
if: github.event.action == 'ready_for_review'
env:
GH_TOKEN: ${{ github.token }}
PR_CHECK_WORKFLOW_SHA: ${{ github.workflow_sha }}
run: node config/scripts/pr-ready-check-reuse.mjs
- name: Classify changed paths
id: filter
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
run: |
set -euo pipefail
# Why HEAD^1 and not --merge-base: HEAD is the pull request merge commit, so its first
# parent is the base side already. Computing a merge base instead would require the
# payload base SHA to be in the graph, which is what forced a full-history checkout.
DIFF_BASE="$(node config/scripts/git-pull-request-diff-base.mjs "$BASE_SHA")"
# Why --no-renames: name-only rename detection can report only the destination.
# A code file moved under docs/ must still expose its code-side deletion.
CHANGED="$(git diff --name-only --no-renames --diff-filter=ACDMR "$DIFF_BASE" HEAD)"
echo "Changed paths:"
printf '%s\n' "$CHANGED"
printf '%s\n' "$CHANGED" | node config/scripts/pr-code-change-scope.mjs | tee -a "$GITHUB_OUTPUT"
# Reuse the path-detector checkout instead of queuing another runner.
- name: Filter changed E2E specs
id: e2e_filter
if: github.event.pull_request.draft != true && steps.filter.outputs.should_run == 'true'
run: |
set -euo pipefail
BASE="${{ github.event.pull_request.base.sha }}"
DIFF_BASE="$(node config/scripts/git-pull-request-diff-base.mjs "$BASE")"
CHANGED="$(git diff --name-only --diff-filter=AMCR "$DIFF_BASE" HEAD)"
# Source routes are executable contracts so a test can prove exact
# authorities, exclusions, and sentinels without evaluating workflow shell.
TEST_FILES_JSON="$(printf '%s\n' "$CHANGED" | node config/scripts/pr-e2e-source-routing.mjs)"
echo "test_files=$TEST_FILES_JSON" >> "$GITHUB_OUTPUT"
# Why a separate signal: the Docker-SSH lane must trigger on SSH source, not on a
# spec name surviving in a route's list. Same routes, so the two cannot drift.
SSH_SOURCE_CHANGED="$(printf '%s\n' "$CHANGED" | node config/scripts/pr-e2e-source-routing.mjs --ssh-source)"
echo "ssh_source_changed=$SSH_SOURCE_CHANGED" >> "$GITHUB_OUTPUT"
echo "SSH source changed: $SSH_SOURCE_CHANGED"
printf '%s\n' "$TEST_FILES_JSON" | E2E_SSH_SOURCE_CHANGED="$SSH_SOURCE_CHANGED" node config/scripts/ci-e2e-job-selection.mjs --job-outputs >> "$GITHUB_OUTPUT"
# Why its own signal: the real-IME lane is a whole ibus session, not a spec, so it must
# trigger on IME source rather than on a spec name in some route's list.
NATIVE_IME_SOURCE_CHANGED="$(printf '%s\n' "$CHANGED" | node config/scripts/pr-e2e-source-routing.mjs --native-ime-source)"
echo "native_ime_source_changed=$NATIVE_IME_SOURCE_CHANGED" >> "$GITHUB_OUTPUT"
WSL_CHANGED="$(git diff --name-only --no-renames --diff-filter=ACDMR "$DIFF_BASE" HEAD)"
WSL_SOURCE_CHANGED="$(printf '%s\n' "$WSL_CHANGED" | node config/scripts/pr-e2e-source-routing.mjs --wsl-source)"
echo "wsl_source_changed=$WSL_SOURCE_CHANGED" >> "$GITHUB_OUTPUT"
echo "Native IME source changed: $NATIVE_IME_SOURCE_CHANGED"
SHOULD_RUN="$(printf '%s\n' "$CHANGED" | node config/scripts/pr-e2e-source-routing.mjs --reusable-workflow)"
if [ "$SHOULD_RUN" = true ]; then
echo "should_run=true" >> "$GITHUB_OUTPUT"
echo "Changed E2E specs: $TEST_FILES_JSON"
else
echo "should_run=false" >> "$GITHUB_OUTPUT"
echo "No specs requiring the reusable E2E workflow"
fi
preflight:
name: static analysis and typecheck
needs: [code_paths]
if: needs.code_paths.outputs.static_analysis == 'true' || needs.code_paths.outputs.typecheck == 'true'
# Why ARM: measured 128s against 172s on ubuntu-latest, with every compute step faster --
# type-aware 24s->15s, anti-slop 28->19s, localization extraction 67->46s, the orcad smoke
# 39->14s. Both lint engines ship linux-arm64 and the Bun target follows process.arch, so
# the whole toolchain resolves. Free for public repositories.
runs-on: ubuntu-24.04-arm
outputs:
shards: ${{ steps.unit-plan.outputs.shards }}
steps:
- name: Checkout
uses: actions/checkout@v6
with:
# Why depth 50: the gates below diff against HEAD^1, so no merge base is computed and
# the payload base SHA need not be in the graph.
# Why no blob:none here, unlike code_paths: this job checks out all 30,226 files, and
# the filter then forces a second promisor fetch of nearly every blob. Measured 23s
# blobless against ~11s without it.
fetch-depth: 50
persist-credentials: false
# Why two guarded installs: the mixed root+mobile store entry is 537 MB against
# 321 MB for root alone, and restoring it costs 8.6s against 4.6s. Most PRs skip the
# mobile install below, so they were paying 216 MB for packages they never link. The
# root-only key is also the one the hourly warmer reseeds. Only one of these runs.
- uses: ./.github/actions/install-node-dependencies
if: needs.code_paths.outputs.mobile_dependencies != 'true'
with:
native-runtime: node
node-version: '24'
- uses: ./.github/actions/install-node-dependencies
if: needs.code_paths.outputs.mobile_dependencies == 'true'
with:
native-runtime: node
node-version: '24'
cache-dependency-path: |
pnpm-lock.yaml
mobile/pnpm-lock.yaml
# Keep each check in its own log while sharing this runner.
- name: Lint
if: '!cancelled()'
id: root-lint
background: true
env:
PREFLIGHT_PHASE_SELECTED: ${{ needs.code_paths.outputs.static_analysis == 'true' }}
PREFLIGHT_PRIOR_SUCCESS: ${{ job.status == 'success' }}
run: |
if [ "$PREFLIGHT_PHASE_SELECTED" != true ] || [ "$PREFLIGHT_PRIOR_SUCCESS" != true ]; then exit 0; fi
pnpm exec oxlint --format github
- name: Reject low-evidence patterns
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run audit:anti-slop
- wait: root-lint
- name: Enforce focused code-quality plugins
if: '!cancelled()'
id: native-code-quality
background: true
env:
PREFLIGHT_PHASE_SELECTED: ${{ needs.code_paths.outputs.static_analysis == 'true' }}
PREFLIGHT_PRIOR_SUCCESS: ${{ job.status == 'success' }}
run: |
if [ "$PREFLIGHT_PHASE_SELECTED" != true ] || [ "$PREFLIGHT_PRIOR_SUCCESS" != true ]; then exit 0; fi
pnpm run audit:code-quality:native
- name: Enforce type-aware code-quality baseline
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run audit:code-quality:type-aware
# Mobile installation changes import resolution for the native cycle check.
- wait: native-code-quality
# Why here: the changed-code gate lints mobile files too, and its type-aware pass
# resolves types from mobile/node_modules. Without the install every mobile type
# degrades to an `error` type — reported as phantom findings against the changed lines.
- uses: ./.github/actions/install-mobile-dependencies
if: needs.code_paths.outputs.static_analysis == 'true' && needs.code_paths.outputs.mobile_dependencies == 'true'
- name: Enforce changed-code quality
if: '!cancelled()'
id: changed-code-quality
background: true
env:
PREFLIGHT_PHASE_SELECTED: ${{ needs.code_paths.outputs.static_analysis == 'true' }}
PREFLIGHT_PRIOR_SUCCESS: ${{ job.status == 'success' }}
run: |
if [ "$PREFLIGHT_PHASE_SELECTED" != true ] || [ "$PREFLIGHT_PRIOR_SUCCESS" != true ]; then exit 0; fi
pnpm run check:code-quality:changed -- "${{ github.event.pull_request.base.sha }}"
- name: Enforce React Doctor on changed lines
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run check:react-doctor:changed -- "${{ github.event.pull_request.base.sha }}"
- wait: changed-code-quality
- name: Check Zustand selector fan-out budget
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run check:zustand-selector-fanout
- name: Check reliability gate manifest
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run check:reliability-gates
- name: Enforce dead design-system classes
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run check:dead-classes
- name: Check VM runtime rollback compatibility
if: needs.code_paths.outputs.static_analysis == 'true'
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
run: |
DIFF_BASE="$(node config/scripts/git-pull-request-diff-base.mjs "$BASE_SHA")"
if git diff --quiet "$DIFF_BASE" HEAD -- \
src/shared/ephemeral-vm-runtime-store.ts \
src/shared/ephemeral-vm-runtime-feature-store.ts \
src/shared/ephemeral-vm-runtime-rollback-projection.ts \
src/shared/ephemeral-vm-runtimes.ts \
src/shared/ephemeral-vm-recipes.ts \
src/shared/orca-yaml-hook-types.ts \
src/main/ephemeral-vm-runtime-service.ts \
src/main/ephemeral-vm-runtime-provisioning-persistence.ts \
src/main/ephemeral-vm-failed-start-cleanup.ts; then
echo "VM runtime persistence is unchanged."
exit 0
fi
node config/scripts/run-ephemeral-vm-runtime-store-rollback-repro.mjs \
config/scripts/ephemeral-vm-runtime-store-cross-version.test.ts
- name: Enforce max-lines ratchet
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run check:max-lines-ratchet
- name: Enforce ts-nocheck ratchet
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run check:ts-nocheck-ratchet
- name: Enforce runtime Electron-import ratchet
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run check:runtime-electron-ratchet
- name: Check Node runtime pin
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run check:node-runtime-pin
# Why: extraction writes sorted evidence to an isolated temporary path,
# so feature PRs need one normalized AST pass rather than a three-OS matrix.
- name: Verify localization extraction
if: '!cancelled()'
id: localization-extraction
background: true
env:
PREFLIGHT_PHASE_SELECTED: ${{ needs.code_paths.outputs.static_analysis == 'true' }}
PREFLIGHT_PRIOR_SUCCESS: ${{ job.status == 'success' }}
BASE_SHA: ${{ github.event.pull_request.base.sha }}
run: |
if [ "$PREFLIGHT_PHASE_SELECTED" != true ] || [ "$PREFLIGHT_PRIOR_SUCCESS" != true ]; then exit 0; fi
# Detection failures run the full check; renames retain the removed input path.
DIFF_BASE="$(node config/scripts/git-pull-request-diff-base.mjs "$BASE_SHA")"
if git diff --name-only --no-renames -z "$DIFF_BASE" HEAD > "$RUNNER_TEMP/localization-changes" &&
scope="$(node config/scripts/localization-extraction-change-scope.mjs "$RUNNER_TEMP/localization-changes")" && [ "$scope" = false ]; then
echo "Localization extraction inputs are unchanged."
else
pnpm run verify:localization-extraction
fi
# Why both: the ratchet proves nothing reachable from the runtime imports electron,
# which is a property of the import graph. This proves the Node artifact it enables
# actually boots, pairs, creates a worktree and round-trips a real PTY.
- name: Boot orcad and round-trip a terminal
if: needs.code_paths.outputs.static_analysis == 'true'
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
ORCA_BACKGROUND_LAUNCH: '1'
run: |
# Detection failures run the smoke; renames retain the removed input path.
DIFF_BASE="$(node config/scripts/git-pull-request-diff-base.mjs "$BASE_SHA")"
if git diff --name-only --no-renames -z "$DIFF_BASE" HEAD > "$RUNNER_TEMP/orcad-smoke-changes" &&
scope="$(node config/scripts/orcad-terminal-smoke-change-scope.mjs "$RUNNER_TEMP/orcad-smoke-changes")" && [ "$scope" = false ]; then
echo "Orcad terminal smoke inputs are unchanged."
else
pnpm run smoke:orcad-terminal
fi
- name: Verify the generated RPC params catalog
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run verify:rpc-params-catalog
- name: Verify the generated ACP protocol schema
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run verify:acp-protocol
- name: Verify bundled skill guides
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run verify:bundled-skill-guides
- name: Verify skill freshness manifest
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run verify:skill-bundle-manifest
- name: Verify localization catalogs
if: '!cancelled()'
id: localization-catalogs
background: true
env:
PREFLIGHT_PHASE_SELECTED: ${{ needs.code_paths.outputs.static_analysis == 'true' }}
PREFLIGHT_PRIOR_SUCCESS: ${{ job.status == 'success' }}
run: |
if [ "$PREFLIGHT_PHASE_SELECTED" != true ] || [ "$PREFLIGHT_PRIOR_SUCCESS" != true ]; then exit 0; fi
pnpm run verify:localization-catalogs
- name: Verify localization coverage
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm run verify:localization-coverage
- wait: [localization-catalogs, localization-extraction]
# Why: project-owned type declarations must live in .ts so tsc
# actually checks them. TypeScript's skipLibCheck: true (inherited
# from @electron-toolkit/tsconfig) silently widens unresolved names
# in .d.ts to `any`, which is how #1186 shipped a broken IPC signature
# past typecheck. See .github/CONTRIBUTING.md#type-declarations-prefer-ts-over-dts.
- name: Guard against project-owned .d.ts in preload/shared
if: needs.code_paths.outputs.static_analysis == 'true'
run: |
matches=$(find src/preload src/shared -name '*.d.ts' 2>/dev/null || true)
if [ -n "$matches" ]; then
echo "::error::Project-owned .d.ts files are not allowed under src/preload or src/shared."
echo "Move type declarations into a .ts file so skipLibCheck does not hide errors."
echo "See .github/CONTRIBUTING.md#type-declarations-prefer-ts-over-dts."
echo "Found:"
echo "$matches"
exit 1
fi
- name: Check feature wall asset budget
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm check:feature-wall-assets
- name: Verify macOS entitlements
if: needs.code_paths.outputs.static_analysis == 'true'
run: pnpm verify:macos-entitlements
# Why: every project is `composite`, so tsc already writes a .tsbuildinfo that lets
# the next run skip unchanged files. Share one cache entry across commits while the
# PR base stays stable; actions/cache keeps the first successful graph and the
# compiler still invalidates stale files from its content hashes.
- name: Cache TypeScript incremental state
if: needs.code_paths.outputs.typecheck == 'true'
uses: actions/cache@v5
with:
path: config/*.tsbuildinfo
key: tsbuildinfo-${{ runner.os }}-${{ hashFiles('pnpm-lock.yaml', 'config/tsconfig*.json') }}-${{ github.event.pull_request.base.sha }}
restore-keys: |
tsbuildinfo-${{ runner.os }}-${{ hashFiles('pnpm-lock.yaml', 'config/tsconfig*.json') }}-
# Planning shares setup and stays off the compiler's critical path.
- name: Plan unit selection
if: '!cancelled()'
id: unit-plan
background: true
env:
PREFLIGHT_PHASE_SELECTED: ${{ needs.code_paths.outputs.typecheck == 'true' }}
PREFLIGHT_PRIOR_SUCCESS: ${{ job.status == 'success' }}
ORCA_UNIT_SELECTION_MODE: ${{ vars.ORCA_UNIT_SELECTION_MODE || 'shadow' }}
run: |
if [ "$PREFLIGHT_PHASE_SELECTED" != true ] || [ "$PREFLIGHT_PRIOR_SUCCESS" != true ]; then exit 0; fi
node config/scripts/ci-unit-plan.mjs
- run: pnpm run typecheck
if: needs.code_paths.outputs.typecheck == 'true'
- wait: unit-plan
- uses: actions/upload-artifact@v7
if: needs.code_paths.outputs.typecheck == 'true'
with:
name: unit-selection-attempt-${{ github.run_attempt }}
path: ci-shards/unit-selection.json
retention-days: 14
git_compatibility:
name: Git compatibility
needs: [code_paths, preflight]
if: >-
!cancelled() &&
needs.code_paths.result == 'success' &&
needs.code_paths.outputs.git_compatibility == 'true' &&
needs.preflight.result == 'success'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
- uses: ./.github/actions/prepare-git-compatibility
- name: Verify Git binary compatibility matrix
env:
ORCA_BACKGROUND_LAUNCH: '1'
run: |
specs=(
"alpine/git:edge-2.38.1|2.38.1"
"alpine/git:v2.49.1|2.49.1"
)
# Why pull up front: a lane's first `docker run` otherwise pulls its image
# while the sibling lane is mid-test, and that stall is charged to the test.
for spec in "${specs[@]}"; do
docker pull --quiet "${spec%%|*}"
done
pids=()
(
ORCA_GIT_COMPAT_BINARY="$HOME/.cache/orca-git-compat/git-2.25.5/git" \
GIT_EXEC_PATH="$HOME/.cache/orca-git-compat/git-2.25.5" \
ORCA_GIT_COMPAT_VERSION="2.25.5" \
pnpm exec vitest run --config config/vitest.config.ts \
src/shared/git-binary-compatibility.test.ts \
src/main/git/worktree-safety-real-git.test.ts \
src/main/git/worktree-rebase-update-refs-real-git.test.ts \
src/relay/git-review-draft-binary-compatibility.test.ts
) &
pids+=("$!")
for spec in "${specs[@]}"; do
(
image="${spec%%|*}"
version="${spec#*|}"
ORCA_GIT_COMPAT_IMAGE="$image" ORCA_GIT_COMPAT_VERSION="$version" \
pnpm exec vitest run --config config/vitest.config.ts \
src/shared/git-binary-compatibility.test.ts \
src/main/git/worktree-safety-real-git.test.ts \
src/main/git/worktree-rebase-update-refs-real-git.test.ts \
src/relay/git-review-draft-binary-compatibility.test.ts
) &
pids+=("$!")
done
status=0
for pid in "${pids[@]}"; do
wait "$pid" || status=1
done
exit "$status"
# Why this job: Orca's session index-heal depends on a Codex behavior — a
# `thread/read` of an unindexed rollout performs a read-repair that inserts the
# `threads` row. Every unit test drives a stub app-server and asserts only that the
# call did not error, so if Codex dropped the repair they would all stay green while
# the subsystem went inert. This runs the pinned real binary and fails when the
# repair stops happening. Pinned because the binary is the thing expected to drift.
codex_index_heal_contract:
name: Codex index-heal contract
needs: [code_paths, preflight]
if: >-
!cancelled() &&
needs.code_paths.result == 'success' &&
needs.code_paths.outputs.codex_index_heal_contract == 'true' &&
needs.preflight.result == 'success'
# Why ARM: @openai/codex ships @openai/codex-linux-arm64.
runs-on: ubuntu-24.04-arm
env:
CODEX_CLI_VERSION: '0.150.1'
# Why a second pin: --no-daemon only exists from 0.156, and Orca's codex wrapper relies on it.
CODEX_NO_DAEMON_CLI_VERSION: '0.158.0'
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
- name: Install pinned Codex CLI
run: |
set -euo pipefail
npm install --no-audit --no-fund --prefix "$RUNNER_TEMP/codex-cli" \
"@openai/codex@$CODEX_CLI_VERSION"
- name: Verify Codex index-heal contract
env:
# Why REQUIRED: without a binary the suite skips, and a job that skips
# reports success. This turns a failed or missing install into a red test
# instead of a green no-op.
ORCA_CODEX_CONTRACT_REQUIRED: '1'
ORCA_CODEX_CONTRACT_VERSION: ${{ env.CODEX_CLI_VERSION }}
run: |
set -euo pipefail
ORCA_CODEX_CONTRACT_BINARY="$RUNNER_TEMP/codex-cli/node_modules/.bin/codex" \
pnpm exec vitest run --config config/vitest.config.ts \
src/main/codex/codex-index-heal-binary-contract.test.ts
- name: Install pinned no-daemon Codex CLI
run: |
set -euo pipefail
npm install --no-audit --no-fund --prefix "$RUNNER_TEMP/codex-cli-no-daemon" \
"@openai/codex@$CODEX_NO_DAEMON_CLI_VERSION"
- name: Verify Codex --no-daemon contract
env:
ORCA_CODEX_NO_DAEMON_CONTRACT_REQUIRED: '1'
ORCA_CODEX_NO_DAEMON_CONTRACT_VERSION: ${{ env.CODEX_NO_DAEMON_CLI_VERSION }}
run: |
set -euo pipefail
ORCA_CODEX_NO_DAEMON_CONTRACT_BINARY="$RUNNER_TEMP/codex-cli-no-daemon/node_modules/.bin/codex" \
pnpm exec vitest run --config config/vitest.config.ts \
src/main/pty/codex-no-daemon-binary-contract.test.ts
# Why the no-daemon install: it is the newest pinned Codex, and its project
# lookup decides whether a worktree launch stops at the trust prompt.
- name: Verify Codex project-trust contract
env:
ORCA_CODEX_TRUST_CONTRACT_REQUIRED: '1'
ORCA_CODEX_TRUST_CONTRACT_VERSION: ${{ env.CODEX_NO_DAEMON_CLI_VERSION }}
run: |
set -euo pipefail
ORCA_CODEX_TRUST_CONTRACT_BINARY="$RUNNER_TEMP/codex-cli-no-daemon/node_modules/.bin/codex" \
pnpm exec vitest run --config config/vitest.config.ts \
src/main/agent-trust-presets.test.ts
xterm_patch_sync:
name: xterm patch sync
needs: [code_paths, preflight]
if: >-
!cancelled() &&
needs.code_paths.result == 'success' &&
needs.code_paths.outputs.xterm_patch_sync == 'true' &&
needs.preflight.result == 'success'
# Why ARM: the patch check rebuilds 4 packages x 2 builds and byte-compares against the
# checked-in bundles. Those were generated on Linux x64 and reproduce byte-for-byte on
# darwin-arm64, so the output is neither host-arch nor host-OS dependent.
runs-on: ubuntu-24.04-arm
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
# The generator uses Node built-ins and installs its own upstream toolchain.
- uses: actions/setup-node@v6
with:
node-version-file: package.json
package-manager-cache: false
# Rebuild both pristine and patched bundles; reuse only the pinned toolchain.
- uses: ./.github/actions/prepare-xterm-dependencies
- name: Verify xterm patches match the pinned upstream build
env:
WORK_DIR: ${{ runner.temp }}/xterm-patch-build
run: node config/scripts/regenerate-xterm-patches.mjs --check --work-dir="$WORK_DIR"
shell_contracts:
name: shell contracts
needs: [code_paths, preflight]
if: >-
!cancelled() &&
needs.code_paths.result == 'success' &&
needs.code_paths.outputs.shell_contracts == 'true' &&
needs.preflight.result == 'success'
# Why ARM: fish 4.x is published for noble/arm64 and zsh is in the arm64 archive.
runs-on: ubuntu-24.04-arm
# Why: this job's cost is almost entirely package download, and a stalled mirror has
# no wall-clock bound of its own. A successful run finishes in ~4.5 minutes, so this
# is generous; it exists so a wedge fails the job instead of holding the whole run
# open for the 6h GitHub default — which also blocks `gh run rerun --failed`.
timeout-minutes: 15
env:
# Why: the suites below gate their live fish tests on the binary, which is
# right on a developer machine and wrong here — this job is a required check
# and its fish lane is the only end-to-end guard for #9993, so a skip would
# report green with nothing exercised. Turns those skips into failures.
ORCA_REQUIRE_FISH: '1'
# Why: the runner image ships pwsh, and the codex wrapper's PowerShell lane must not skip.
ORCA_REQUIRE_PWSH: '1'
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
# Why fish: shell-ready.test.ts gates its live fish test on the binary being
# present, so without this the fish barrier is only covered by config-shape
# assertions and never actually exercised.
# Why release-4: DECSET 2031 arming lives in the fish 4.0 Rust tty_handoff, and
# fish-color-scheme-child-stdin.node-pty.test.ts (#9993) needs it. Noble ships
# 3.7, so the PPA is what makes that lane real.
- name: Install zsh and fish
id: shells
background: true
run: |
# Why the update/PPA/fish steps are tolerant: a repo the runner image already
# ships can lack a Release file for this suite, and a failed add-apt-repository
# still leaves its list entry behind — either makes `apt-get update` exit
# non-zero and would red this required check over something unrelated to the
# PR. Every fish outcome is judged by the version gate below instead, so only
# the zsh install (which has no such gate) stays fatal here.
# Why retry only here: adding the PPA is the network-flaky step, and the
# version gate below is fatal, so a transient Launchpad blip would
# otherwise red a required check on PRs unrelated to shells.
# Why -n: add-apt-repository refreshes every configured repo on its own. With
# an update on each side of it this step refreshed them three times over, and
# the Azure archive mirror alone costs ~15-30s a pass. The PPA index is the
# only thing the repo list gains here, and the single update below fetches it.
# Why bound acquisition: measured on a *passing* run, this step spent 40s
# fetching 11.4 MB of index and then 2m17s fetching 8.9 MB of packages at
# 65 kB/s — it is dominated by download throughput, not by work. apt applies
# no wall-clock bound to a stalled mirror, so a slow Launchpad or archive
# host wedges the step for tens of minutes. This job is a required check, so
# a wedge holds the entire run open and blocks `gh run rerun --failed`.
# Bounded timeouts plus retries turn an unbounded hang into a fast, legible
# failure. Set in apt.conf.d rather than on each command line so the two
# invocations below stay exactly as pr-workflow-parallelism.test.mjs parses
# them. Retries are 1, not 3: a first attempt at these bounds already multiplied
# 30s x 3 retries across every index file into a ~15 minute stall on a dead
# mirror, which is worse than failing once and moving on.
sudo tee /etc/apt/apt.conf.d/99-orca-shell-contracts >/dev/null <<'APTCONF'
Acquire::http::Timeout "15";
Acquire::https::Timeout "15";
Acquire::Retries "1";
APTCONF
for attempt in 1 2 3; do
sudo add-apt-repository -y -n ppa:fish-shell/release-4 && break
echo "add-apt-repository attempt ${attempt} failed; retrying" >&2
sudo add-apt-repository -y -n -r ppa:fish-shell/release-4 || true
sleep 5
done
# Why a wall-clock bound on each command: apt's Acquire timeouts are per-connection,
# so a dead mirror costs timeout x retries x every index file. Measured: the archive
# mirror stalled with zero bytes and the step burned 14m26s before the job bound
# killed it. `timeout` is the only thing that bounds the command as a whole.
# The update is already tolerant by design (see above), so bounding it just caps
# what a dead mirror can cost before the install runs against whatever index exists.
timeout 120 sudo apt-get update || true
# Why both shells on one line: pr-workflow-parallelism.test.mjs parses only the
# first install command in this step to prove the lane really installs them.
timeout 300 sudo apt-get install -y zsh fish
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: node
- wait: shells
# Separate from the install so the failure names the contract, not an apt error.
# ORCA_REQUIRE_FISH re-checks this at test time; this step just fails in seconds
# instead of after a full dependency install.
- name: Require fish 4+
run: |
version="$(fish --version 2>/dev/null || true)"
major="${version##*version }"
major="${major%%.*}"
case "$major" in '' | *[!0-9]*) major=0 ;; esac
echo "${version:-<fish not installed>}"
if [ "$major" -lt 4 ]; then
echo "::error::shell contracts needs fish 4+ (DECSET 2031 arming, #9993) but got '${version:-none}'. Fix the ppa:fish-shell/release-4 install rather than letting the fish lane skip." >&2
exit 1
fi
- name: Test real shell contracts
run: |
pnpm exec vitest run --config config/vitest.config.ts --maxWorkers=1 \
src/main/daemon/repro-13767-shell-ready-marker-lost-to-exec.test.ts \
src/main/daemon/shell-ready.test.ts \
src/main/daemon/node-pty-fd-leak.test.ts \
src/main/providers/local-pty-shell-ready-zsh-launch-environment.test.ts \
src/main/providers/__tests__/shell-ready-framework-example.test.ts \
src/main/pty/codex-shell-launch-preflight.test.ts \
src/main/pty/codex-shell-no-daemon.test.ts \
src/main/pty/omp-shell-wrapper-alias-safety.test.ts \
src/main/pty/omp-shell-wrapper.node-pty.test.ts \
src/main/fish-xdg-data-dirs-handoff.test.ts \
src/main/shell-startup-feature-channel.test.ts \
src/main/zsh-deferred-startup-line-init.live-shell.test.ts \
src/main/terminal-history-fish-session.node-pty.test.ts \
src/main/zsh-scoped-histfile.live-shell.test.ts \
src/main/zsh-startup-hook-user-config-equivalence.live-shell.test.ts \
src/main/zsh-wrapper-version-mismatch.live-shell.test.ts \
src/main/runtime/structured-session-cli-login-shell.live-shell.test.ts \
src/renderer/src/components/terminal-pane/fish-color-scheme-child-stdin.node-pty.test.ts \
src/shared/fish-query-reply-child-stdin.node-pty.test.ts \
src/shared/pty-reply-echo-shapes.node-pty.test.ts \
src/shared/startup-shell-portability.live-shell.test.ts \
src/shared/posix-command-path-lookup.test.ts
# Preflight saves the Node cache and publishes the plan before tests fan out.
test:
needs: [code_paths, preflight]
# Cancellation and failed prerequisites stop the expensive matrix.
if: >-
!cancelled() &&
needs.code_paths.outputs.test == 'true' &&
needs.preflight.result == 'success'
uses: ./.github/workflows/unit-tests.yml
with:
node_versions: '["24"]'
runner: ubuntu-24.04-arm
shards: ${{ needs.preflight.outputs.shards }}
# Why a sibling and not part of the test workflow: it is advisory, so it must not delay the
# gate. Inside unit-tests.yml a caller's `needs: test` waited for it, holding verify ~36s
# after the last shard. Deliberately absent from verify's needs for the same reason.
unit_selection_evidence:
needs: [test]
if: ${{ !cancelled() && (needs.test.result == 'success' || needs.test.result == 'failure') }}
uses: ./.github/workflows/unit-selection-evidence.yml
# Why a separate job: the test needs a real Chrome, and the sharded `test` matrix
# would pay for it on every shard to run one file in whichever shard it landed in.
orcad_browser:
name: orcad browser provider
needs: [code_paths, preflight]
if: >-
!cancelled() &&
needs.code_paths.result == 'success' &&
needs.code_paths.outputs.orcad_browser == 'true' &&
needs.preflight.result == 'success'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
# Why no native-runtime: the provider drives the prebuilt agent-browser binary
# shipped in node_modules and never touches node-pty.
- uses: ./.github/actions/install-node-dependencies
# Why the runner's Google Chrome and not its chromium: Ubuntu 24.04 only ships an
# AppArmor userns profile for the Chrome .deb, so chromium dies with "No usable
# sandbox" and the provider passes no --no-sandbox. Why fail instead of skip: an
# unset ORCA_BROWSER_EXECUTABLE is exactly how this test went uncovered for so long.
- name: Resolve Chrome for the browser provider
run: |
set -euo pipefail
chrome="$(command -v google-chrome || command -v google-chrome-stable || true)"
if [ -z "$chrome" ]; then
echo "::error::No Google Chrome on the runner; the browser provider test would silently skip."
exit 1
fi
"$chrome" --version
echo "ORCA_BROWSER_EXECUTABLE=$chrome" >> "$GITHUB_ENV"
- name: Test external Chromium browser provider
run: |
pnpm exec vitest run --config config/vitest.config.ts \
src/main/orcad/external-chromium-browser-process.integration.test.ts
# Why its own job: it needs mobile/node_modules and a real browser, and the sharded `test`
# matrix would pay for both on every shard to run two files. It builds the same bundle the
# package job ships, and adds the render and census checks packaging does not run.
mobile_web_app:
name: mobile web app bundle
needs: [code_paths]
if: needs.code_paths.outputs.mobile_web_app == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
# Why no native-runtime: the builder is esbuild and the render check is a browser. Nothing
# in this job loads node-pty.
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: none
cache-dependency-path: |
pnpm-lock.yaml
mobile/pnpm-lock.yaml
# The drawer check runs on WebKit as well as Chrome, because the shell's iOS WebView is
# WebKit and the Chrome above cannot stand in for it. Downloaded rather than resolved from
# the runner: Ubuntu ships no WebKit build to point at.
- name: Install WebKit for the drawer check
id: webkit
background: true
run: pnpm exec playwright install --with-deps webkit
# The entry lives in mobile/ so one React resolves; without this every RN import is nothing.
- uses: ./.github/actions/install-mobile-dependencies
- name: Prepare mobile route snapshot
id: mobile-routes
background: true
run: node config/scripts/run-mobile-web-app-checks.mjs --prepare-route-snapshot "$RUNNER_TEMP/mobile-routes.json"
# Why the runner's Google Chrome and not a downloaded chromium: same reason as the orcad
# browser job -- Ubuntu 24.04 only ships an AppArmor userns profile for the Chrome .deb.
# Why fail instead of skip: a silently skipped render check is the failure this job exists
# to prevent.
- name: Resolve Chrome for the render check
run: |
set -euo pipefail
chrome="$(command -v google-chrome || command -v google-chrome-stable || true)"
if [ -z "$chrome" ]; then
echo "::error::No Google Chrome on the runner; the render check would silently skip."
exit 1
fi
"$chrome" --version
echo "ORCA_MOBILE_WEB_RENDER_BROWSER=$chrome" >> "$GITHUB_ENV"
- name: Build and verify the app bundle
run: pnpm run build:mobile-web
# Browser checks must observe both a finished install and the built bundle.
- wait: [webkit, mobile-routes]
# The bundling tests skip themselves where mobile dependencies are absent, which is how they
# stay green in the sharded `test` job. This is the job that installs them, so here a missing
# install has to fail rather than skip everything the job exists to run.
# Why a prefix and not a file list: the list this replaces had gone stale twice without
# anyone noticing, because a census whose closure block skips without the env flag below is
# green in the sharded `test` job whether or not it ever runs here. The prefix is the same
# one `pr-code-change-scope.mjs` fires this job on, so naming a test into the family is all
# it takes to have it run. Quoted because these are vitest filename filters, matched as
# substrings against the discovered files, and the shell must not touch them.
# Route censuses share fresh dependency lists for this invocation; scratch builds stay independent.
- name: Builder, override census and render checks
env:
ORCA_MOBILE_WEB_APP_DEPS_REQUIRED: '1'
ORCA_MOBILE_WEB_PREPARED_ROUTE_SNAPSHOT: ${{ runner.temp }}/mobile-routes.json
run: |
node config/scripts/run-mobile-web-app-checks.mjs
cross-version-wire:
name: cross-version wire compatibility
needs: [code_paths, preflight]
if: >-
!cancelled() &&
needs.code_paths.result == 'success' &&
needs.code_paths.outputs.cross-version-wire == 'true' &&
needs.preflight.result == 'success'
# Why ARM: source-only: tagged checkout plus in-process vitest, no docker or browser.
runs-on: ubuntu-24.04-arm
steps:
# Why fetch-depth 0: the harness extracts the newest release tag to skew
# current code against it. The default shallow clone has no tags, which is
# why this cannot ride along in the sharded `test` job.
- name: Checkout
uses: actions/checkout@v6
with:
# Why blob:none: full history is needed for the merge-base diff, but historical
# file contents are not. Blobs are ~89% of this repo's pack, and Git fetches the
# few this job actually reads on demand.
fetch-depth: 0
filter: blob:none
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: node
# R1: an upgrade from the newest release must adopt its live terminal daemon.
- name: Daemon protocol crossing against newest release
run: pnpm run check:daemon-protocol-crossing
# R3: a runtime launcher swap must not also bump the daemon protocol (D7.1).
# The label is read from the event payload, so push again after applying it.
- name: Runtime launcher change does not bump daemon protocol
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
ORCA_ALLOW_RUNTIME_LAUNCHER_PROTOCOL_BUMP: ${{ contains(github.event.pull_request.labels.*.name, 'allow-runtime-launcher-protocol-bump') }}
run: |
DIFF_BASE="$(node config/scripts/git-pull-request-diff-base.mjs "$BASE_SHA")"
pnpm run check:runtime-launcher-protocol-ratchet -- --base "$DIFF_BASE"
# Runs the whole directory so a new file there cannot be left out; an empty match
# exits 1 ("No test files found"), so this lane cannot pass while running zero tests.
- name: Old/new client and server compatibility journeys
run: pnpm exec vitest run --config config/vitest.config.ts tests/e2e/cross-version-wire/
managed_hook_node18:
name: managed hooks on Node 18
needs: [code_paths, preflight]
if: >-
!cancelled() &&
needs.code_paths.result == 'success' &&
needs.code_paths.outputs.managed_hook_node18 == 'true' &&
needs.preflight.result == 'success'
# Why ARM: Node 18 publishes linux-arm64; the per-platform runtime files are read as data.
runs-on: ubuntu-24.04-arm
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
- name: Build relay companions
run: pnpm run build:relay
- name: Setup Node 18 runtime
uses: actions/setup-node@v6
with:
node-version: '18'
- name: Smoke managed-hook companions
run: node config/scripts/smoke-managed-hook-runtime-node18.mjs
package:
name: package
needs: [code_paths, preflight]
if: needs.code_paths.outputs.package == 'true'
runs-on: ubuntu-latest
# Let the serial Docker gates reach their own deadlines and report cleanup failures.
timeout-minutes: 90
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Cache electron-builder downloads
uses: actions/cache/restore@v5
with:
path: ~/.cache/electron-builder
key: electron-builder-linux-${{ hashFiles('pnpm-lock.yaml') }}
restore-keys: |
electron-builder-linux-
- uses: ./.github/actions/prepare-linux-package-fixture
id: daemon-fixture-cache
background: true
with:
fixture: daemon-shutdown-descendants
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: electron
cache-dependency-path: |
pnpm-lock.yaml
mobile/pnpm-lock.yaml
# Why here: electron-builder's beforePack requires out/mobile-web, and the bundle
# build resolves React Native and Expo from mobile/node_modules.
- uses: ./.github/actions/install-mobile-dependencies
# The checkout must pass without depending on a historical baseline that leaks.
- wait: daemon-fixture-cache
- name: Verify Linux daemon shutdown descendant cleanup
env:
ORCA_BACKGROUND_LAUNCH: '1'
ORCA_DAEMON_SHUTDOWN_FIXTURE_CACHE_IMAGE: ${{ steps.daemon-fixture-cache.outputs.image }}
run: node config/scripts/run-daemon-shutdown-descendants-docker.mjs
# Why --no-file-parallelism: every file here launches a full Electron stack twice, and each
# probe carries its own in-process deadline. Four at once on a 4-vCPU runner starve each other
# past those deadlines; serial, every probe owns the runner.
- name: Test Linux Electron lifecycle boundary
run: >-
xvfb-run --auto-servernum pnpm exec vitest run --config config/vitest.config.ts
--no-file-parallelism
src/main/browser/browser-client-page-renderer-lifecycle.electron.test.ts
src/main/browser/browser-route-tcp-egress.electron.test.ts
src/main/browser/browser-route-webrtc-egress.electron.test.ts
src/main/browser/browser-route-h3-egress.electron.test.ts
src/main/browser/browser-route-dns-prefetch.electron.test.ts
# Restore independent fixtures during builds; keep native lifecycle probes serial.
- uses: ./.github/actions/prepare-linux-package-fixture
id: shutdown-fixture-cache
background: true
with:
fixture: headless-serve-shutdown
- uses: ./.github/actions/prepare-linux-package-fixture
id: cli-fixture-cache
background: true
with:
fixture: cli-launch-contract
- name: Install Linux package tooling
id: linux-package-tools
background: true
run: sudo apt-get update && sudo apt-get install -y cpio rpm
- name: Build package inputs
run: |
status=0
pnpm run build:cli || status=1
scripts=(build:relay build:electron-vite:parallel)
pids=()
for script in "${scripts[@]}"; do
pnpm run "$script" &
pids+=("$!")
done
for pid in "${pids[@]}"; do
wait "$pid" || status=1
done
exit "$status"
- name: Project web client from renderer build
id: web-client
background: true
run: pnpm run build:web-from-renderer
# Why here and not inside "Build package inputs": this job assembles packaging inputs step by
# step instead of calling build:release, and electron-builder's beforePack guard hard-fails
# without out/mobile-web. This is the real page, ~8 MB, not a bootstrap document.
- name: Build mobile web bundle
run: pnpm run build:mobile-web
- name: Build native components
run: pnpm run build:native
- wait: [linux-package-tools, web-client]
- name: Package unpacked app
env:
ORCA_BACKGROUND_LAUNCH: '1'
ORCA_REUSE_PREPARED_NATIVE_RUNTIME: '1'
# Prepare once with every hook, then isolate the format-specific mutations.
run: |
pnpm exec electron-builder --config config/electron-builder.config.cjs --linux dir --x64 --publish never
node config/scripts/package-linux-formats.mjs
- name: Verify root-package marker payloads
run: |
set -euo pipefail
version="$(node -p "require('./package.json').version")"
deb="dist/orca-ide_${version}_amd64.deb"
rpm="dist/orca-ide-${version}.x86_64.rpm"
test -s "$deb"
test -s "$rpm"
deb_marker="$(dpkg-deb --fsys-tarfile "$deb" | tar -xOf - ./opt/Orca/resources/package-type)"
rpm_marker="$(rpm2cpio "$rpm" | cpio --quiet --extract --to-stdout ./opt/Orca/resources/package-type)"
[[ "$deb_marker" == deb ]] || { echo "Expected deb marker, got: $deb_marker"; exit 1; }
[[ "$rpm_marker" == rpm ]] || { echo "Expected rpm marker, got: $rpm_marker"; exit 1; }
- wait: shutdown-fixture-cache
- name: Verify headless serve signal shutdown
env:
ORCA_SHUTDOWN_FIXTURE_CACHE_IMAGE: ${{ steps.shutdown-fixture-cache.outputs.image }}
run: >-
node config/scripts/run-headless-serve-shutdown-docker.mjs
--appimage dist/orca-linux.AppImage --all-entrypoints
# A default container reproduces the hostile AppImage launch environment.
- wait: cli-fixture-cache
- name: Verify Linux CLI launch contract
env:
ORCA_CLI_FIXTURE_CACHE_IMAGE: ${{ steps.cli-fixture-cache.outputs.image }}
run: node config/scripts/run-linux-cli-launch-contract-docker.mjs --appimage dist/orca-linux.AppImage
- name: Smoke packaged CLI
run: node config/scripts/smoke-packaged-cli.mjs --app-dir=dist/linux-unpacked
- name: Smoke packaged hang watchdog worker
run: xvfb-run --auto-servernum node config/scripts/smoke-packaged-hang-watchdog-worker.mjs --app-dir=dist/linux-unpacked
package_windows:
name: package (windows)
needs: [code_paths, preflight]
if: needs.code_paths.outputs.package_windows == 'true'
runs-on: windows-2022
timeout-minutes: 30
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Cache electron-builder downloads
uses: actions/cache/restore@v5
with:
path: |
~\AppData\Local\electron\Cache
~\AppData\Local\electron-builder\Cache
# Release builds seed this exact path set; PR-local copies cannot serve other PRs.
key: electron-builder-win-${{ hashFiles('pnpm-lock.yaml') }}
restore-keys: |
electron-builder-win-
# Why persist-native-cache false: this job later rebuilds the same path for
# Electron. A post-job save would store the Electron ABI under the Node key.
- uses: ./.github/actions/install-node-dependencies
id: deps
with:
native-runtime: node
persist-native-cache: 'false'
cache-dependency-path: |
pnpm-lock.yaml
mobile/pnpm-lock.yaml
# Why here: electron-builder's beforePack requires out/mobile-web, and the bundle
# build resolves React Native and Expo from mobile/node_modules.
- uses: ./.github/actions/install-mobile-dependencies
- name: Save compiled Node native modules
if: steps.deps.outputs.native-cache-hit != 'true'
uses: actions/cache/save@v5
with:
path: ${{ steps.deps.outputs.native-cache-path }}
key: ${{ steps.deps.outputs.native-cache-key }}
# vitest runs here directly rather than through `pnpm test`, so the addon
# assertions only hold once install-node-dependencies has rebuilt natives.
- name: Test Windows installer process probe
# Keep cold CIM startup out of the concurrent Electron/native process workload.
run: >-
pnpm exec vitest run --config config/vitest.config.ts
config/scripts/nsis-process-check.test.mjs
- name: Test Windows-specific boundaries
run: >-
pnpm exec vitest run --config config/vitest.config.ts
config/scripts/rebuild-native-deps.test.mjs
config/scripts/rebuild-native-deps-windows-process-tree.test.mjs
config/scripts/rebuild-native-deps-node-pty.test.mjs
config/scripts/ensure-native-runtime-job-ownership.test.mjs
config/scripts/verify-packaged-node-pty-job-ownership.test.mjs
config/scripts/windows-pe-machine.test.mjs
config/scripts/script-module-dependencies.test.mjs
src/main/windows-registry-addon.test.ts
config/scripts/windows-process-tree-gyp-path.test.mjs
config/scripts/windows-process-tree-gyp-rebuild.test.mjs
config/scripts/package-electron-runtime-contract.test.mjs
config/scripts/electron-builder-runtime-resources.test.mjs
src/main/browser/browser-client-page-renderer-lifecycle.electron.test.ts
src/main/browser/browser-route-tcp-egress.electron.test.ts
src/main/browser/browser-route-webrtc-egress.electron.test.ts
src/main/browser/browser-route-h3-egress.electron.test.ts
src/main/browser/browser-route-dns-prefetch.electron.test.ts
src/main/providers/windows-conpty-wide-char-duplication.node-pty.test.ts
src/main/providers/pty-repaint-wide-char-buffer.node-pty.test.ts
src/shared/child-process/windows-command-line.win32.test.ts
src/shared/child-process/windows-cmd-shim-resolution.test.ts
src/shared/child-process/windows-cmd-shim-resolution.win32.test.ts
src/main/agent-hooks/windows-hook-payload-delivery.test.ts
src/main/jcode/hook-gate-script.test.ts
src/main/agent-hooks/windows-direct-cmd-hook-command.test.ts
src/main/codex/windows-hook-command.test.ts
src/main/codex/windows-hook-upgrade.test.ts
src/main/codex/hook-service-managed-install.test.ts
src/main/windows/windows-pty-job.win32.test.ts
src/main/windows/windows-msys-job.win32.test.ts
src/main/providers/agent-foreground-process-git-bash.win32.test.ts
src/main/windows/windows-host-job.win32.test.ts
src/main/windows/windows-process-tree-command-line-patch.test.ts
src/main/windows/windows-process-table-native-addon.win32.test.ts
src/main/persistence/profile-state/profile-state-access-windows-native.win32.test.ts
src/main/windows-live-tree-kill.win32.test.ts
src/main/wsl/wsl-runner.test.ts
src/main/wsl/wsl-guest-environment.test.ts
src/main/wsl/wsl-executable-path.win32.test.ts
src/main/wsl/wsl-w1-w3-contract.test.ts
src/shared/source-scan/source-tree-scan.test.ts
src/main/cli/wsl-cli-powershell-boundary.test.ts
src/main/computer/desktop-script-runtime-host.win32.test.ts
src/main/cursor/hook-service.test.ts
src/main/orca-profiles/profile-index-store.test.ts
src/main/startup/windows-install-dir-acl-repair.win32.test.ts
src/main/runtime/repo-worktree-admin-fingerprint.test.ts
src/main/runtime/worktree-scan-admin-fingerprint-gate.test.ts
src/shared/secure-file-fsync-flags.test.ts
src/shared/secure-path-windows-acl.win32.test.ts
src/main/runtime/unreadable-secret-store-preservation.win32.test.ts
src/main/ipc/pty-codex-account-attribution.test.ts
src/main/ipc/pty-spawn-env-codex-resume-provenance.test.ts
src/main/ipc/preflight-provider-command-selection.test.ts
src/main/ipc/preflight-runnable-local-cli.test.ts
src/relay/windows-port-scan.win32.test.ts
src/main/ssh/ssh-relay-upload-stage-windows-identity.test.ts
src/main/ssh/remote-node-runtime-store-windows.test.ts
# Why the :parallel variant: identical to build:release except the three
# electron-vite targets overlap instead of running back to back. The Linux package
# job already packages and smoke-tests an AppImage built that way.
- name: Cache Windows CLI launcher
uses: actions/cache@v5
with:
# Why the cargo directories ride along: a hit on .build skips the build
# entirely, but a miss otherwise recompiles the resource crate from a
# cold registry.
path: |
native/windows-cli-launcher/.build
native/windows-cli-launcher/target
~/.cargo/registry
key: windows-cli-launcher-${{ runner.os }}-${{ runner.arch }}-${{ hashFiles('native/windows-cli-launcher/**', 'config/scripts/build-windows-cli-launcher.mjs', 'resources/build/icon.ico', 'package.json') }}
# Why an explicit step rather than trusting the runner image: the packaged
# CLI launcher is a native Rust binary, and a host without cargo fails deep
# inside electron-builder's native hook instead of here.
- name: Ensure the Rust toolchain for the Windows CLI launcher
shell: pwsh
run: |
if (-not (Get-Command cargo -ErrorAction SilentlyContinue)) {
rustup toolchain install stable --profile minimal
rustup default stable
}
cargo --version
- name: Build package inputs
run: pnpm run build:release:parallel
- name: Prepare Electron native runtime
uses: ./.github/actions/prepare-native-runtime
with:
native-runtime: electron
node-version: ${{ steps.deps.outputs.node-version }}
- name: Package unpacked app
env:
ORCA_REUSE_PREPARED_NATIVE_RUNTIME: '1'
run: pnpm exec electron-builder --config config/electron-builder.config.cjs --dir
- name: Smoke packaged Windows PTY native capability
run: pnpm run smoke:windows-pty-native-capability -- --exe=dist/win-unpacked/Orca.exe
- name: Smoke packaged CLI
run: node config/scripts/smoke-packaged-cli.mjs --app-dir=dist/win-unpacked
e2e:
name: e2e
needs: [code_paths, preflight]
if: >-
!cancelled() &&
needs.code_paths.result == 'success' &&
needs.code_paths.outputs.e2e_should_run == 'true' &&
(
needs.preflight.result == 'success' ||
(
needs.preflight.result == 'skipped' &&
needs.code_paths.outputs.static_analysis == 'false' &&
needs.code_paths.outputs.typecheck == 'false'
)
)
# Why: reusable e2e.yml only checkouts, builds, and uploads artifacts.
permissions:
contents: read
uses: ./.github/workflows/e2e.yml
with:
# The synthetic pull-request merge ref can disappear while this reusable
# workflow is queued. The head SHA is immutable and works for every PR.
ref: ${{ github.event.pull_request.head.sha }}
test_files: ${{ needs.code_paths.outputs.test_files }}
ssh_source_changed: ${{ needs.code_paths.outputs.ssh_source_changed }}
run_changed_e2e: ${{ needs.code_paths.outputs.e2e_run_changed != 'false' }}
needs_build: ${{ needs.code_paths.outputs.e2e_needs_build != 'false' }}
# Why this is not in verify's needs: it is the first PR-gate run of a harness whose reliability
# is only known from nightly main runs (20/20 green, 2026-08-09..2026-08-29, p50 3m25s). It
# reports a red X on the PR without blocking, exactly like `e2e` above. Deliberately no
# continue-on-error: that renders the check green and hides the signal it exists to give. To
# make it blocking, add it to verify.needs, add TERMINAL_IME_NATIVE to the env below, and
# require `success || skipped` outside the strict loop — see the note on `e2e`.
terminal_ime_native:
name: real IME
needs: code_paths
if: needs.code_paths.outputs.native_ime_source_changed == 'true'
# Why: the reusable workflow only checks out, builds, and uploads artifacts.
permissions:
contents: read
uses: ./.github/workflows/terminal-ime-e2e.yml
windows_wsl:
name: real WSL terminal
needs: code_paths
if: needs.code_paths.outputs.wsl_source_changed == 'true'
permissions:
contents: read
uses: ./.github/workflows/windows-wsl-e2e.yml
with:
ref: ${{ github.event.pull_request.head.sha }}
verify:
if: ${{ !cancelled() }}
needs:
- code_paths
- preflight
- git_compatibility
- codex_index_heal_contract
- xterm_patch_sync
- shell_contracts
- test
- orcad_browser
- mobile_web_app
- cross-version-wire
- managed_hook_node18
- package
- package_windows
# Evaluating job results only needs the lightweight container runner.
runs-on: ubuntu-slim
timeout-minutes: 5
steps:
# Why: e2e is deliberately absent from needs. The suite is currently red on
# main (every scheduled run), so gating merges on it would block any PR that
# touches tests/e2e/** — including the ones fixing the suite. Until it is
# green the job runs and reports for E2E-path PRs without blocking. To flip
# it on: add `e2e` to needs, add E2E to the env below, and require
# `"$E2E" = success || skipped` after the loop — skipped is the normal
# result for a path-filtered job and must keep passing, so it has to be
# checked outside the loop or it would excuse the jobs above.
- name: Require successful checks
env:
CODE_PATHS: ${{ needs.code_paths.result }}
SHOULD_RUN: ${{ needs.code_paths.outputs.should_run }}
PREFLIGHT: ${{ needs.preflight.result }}
PREFLIGHT_SHOULD_RUN: ${{ needs.code_paths.outputs.static_analysis == 'true' || needs.code_paths.outputs.typecheck == 'true' }}
GIT_COMPATIBILITY: ${{ needs.git_compatibility.result }}
GIT_COMPATIBILITY_SHOULD_RUN: ${{ needs.code_paths.outputs.git_compatibility }}
CODEX_INDEX_HEAL_CONTRACT: ${{ needs.codex_index_heal_contract.result }}
CODEX_INDEX_HEAL_CONTRACT_SHOULD_RUN: ${{ needs.code_paths.outputs.codex_index_heal_contract }}
XTERM_PATCH_SYNC: ${{ needs.xterm_patch_sync.result }}
XTERM_PATCH_SYNC_SHOULD_RUN: ${{ needs.code_paths.outputs.xterm_patch_sync }}
SHELL_CONTRACTS: ${{ needs.shell_contracts.result }}
SHELL_CONTRACTS_SHOULD_RUN: ${{ needs.code_paths.outputs.shell_contracts }}
TEST: ${{ needs.test.result }}
TEST_SHOULD_RUN: ${{ needs.code_paths.outputs.test }}
ORCAD_BROWSER: ${{ needs.orcad_browser.result }}
ORCAD_BROWSER_SHOULD_RUN: ${{ needs.code_paths.outputs.orcad_browser }}
MOBILE_WEB_APP: ${{ needs.mobile_web_app.result }}
MOBILE_WEB_APP_SHOULD_RUN: ${{ needs.code_paths.outputs.mobile_web_app }}
CROSS_VERSION_WIRE: ${{ needs.cross-version-wire.result }}
CROSS_VERSION_WIRE_SHOULD_RUN: ${{ needs.code_paths.outputs.cross-version-wire }}
MANAGED_HOOK_NODE18: ${{ needs.managed_hook_node18.result }}
MANAGED_HOOK_NODE18_SHOULD_RUN: ${{ needs.code_paths.outputs.managed_hook_node18 }}
PACKAGE: ${{ needs.package.result }}
PACKAGE_SHOULD_RUN: ${{ needs.code_paths.outputs.package }}
PACKAGE_WINDOWS: ${{ needs.package_windows.result }}
PACKAGE_WINDOWS_SHOULD_RUN: ${{ needs.code_paths.outputs.package_windows }}
run: |
if [ "$CODE_PATHS" != "success" ]; then
exit 1
fi
if [ "$SHOULD_RUN" != "true" ]; then
echo "Docs-only change; expensive PR checks skipped."
fi
failed=0
check_job() {
local name="$1" result="$2" should="$3"
if [ "$should" = "true" ]; then
if [ "$result" != "success" ]; then
echo "$name: expected success, got $result"
failed=1
fi
else
if [ "$result" != "skipped" ]; then
echo "$name: expected skipped, got $result"
failed=1
fi
fi
}
# Require success when the PR has code-relevant changes
check_job preflight "$PREFLIGHT" "$PREFLIGHT_SHOULD_RUN"
check_job git_compatibility "$GIT_COMPATIBILITY" "$GIT_COMPATIBILITY_SHOULD_RUN"
check_job codex_index_heal_contract "$CODEX_INDEX_HEAL_CONTRACT" "$CODEX_INDEX_HEAL_CONTRACT_SHOULD_RUN"
check_job xterm_patch_sync "$XTERM_PATCH_SYNC" "$XTERM_PATCH_SYNC_SHOULD_RUN"
check_job shell_contracts "$SHELL_CONTRACTS" "$SHELL_CONTRACTS_SHOULD_RUN"
check_job test "$TEST" "$TEST_SHOULD_RUN"
check_job orcad_browser "$ORCAD_BROWSER" "$ORCAD_BROWSER_SHOULD_RUN"
check_job mobile_web_app "$MOBILE_WEB_APP" "$MOBILE_WEB_APP_SHOULD_RUN"
check_job cross-version-wire "$CROSS_VERSION_WIRE" "$CROSS_VERSION_WIRE_SHOULD_RUN"
check_job managed_hook_node18 "$MANAGED_HOOK_NODE18" "$MANAGED_HOOK_NODE18_SHOULD_RUN"
check_job package "$PACKAGE" "$PACKAGE_SHOULD_RUN"
check_job package_windows "$PACKAGE_WINDOWS" "$PACKAGE_WINDOWS_SHOULD_RUN"
exit "$failed"