Files
orca/.github/workflows/pr.yml
Jinwoo Hong 211821dc17 feat(mobile): render agent session history from the desktop's bundle (OTA phase C, C5.1) (#21596)
* fix(mobile): refuse a page target the shell will not take instead of opening it here

`useRouteHandoff`'s web sibling answered two things — handed off, or push it
locally — and fell through to the local router for three different reasons. Only
one of them is a page route. An href the protocol's own pattern drops and a shell
that answered no are the page reaching past what this shell can serve, and the
bundle carries every route under `app/h`, so the fallback does not paint
Unmatched: it mounts `session/[worktreeId]` on React Native Web inside the shell.

The outcome is now tri-state. A target outside `pageRoutes` is never pushed
locally; the page stays where it is and names the reason once per client, which
is the bound the other page-side reporters take.

Proved in the render check against the real bundle: with the double granting no
`navigate`, "Back to hosts" left the host route for `/` and painted Unmatched
before this, and now stays put, posts nothing and reports no page fault.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): encode the host id the worktree row's navigation actions build

Both targets this sheet offers interpolated `hostId` raw — the C1.2 class the
C1.8 stack fixed at the route files and `web.tsx`, at the last two sites that
still had it. `useLocalSearchParams` answers the decoded value, so a deep-linked
id carrying `?`, `#` or whitespace stops being one segment.

It matters more from C5.1 on. Inside the page these targets go through
`useRouteHandoff`, which matches the pathname against the shell's `pageRoutes`
before deciding anything, and the id is the segment the pattern is reading.

The worktree id was already encoded at both sites; this makes the host id match,
and the new test pins all four targets rather than only the one that moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): render agent session history from the desktop's bundle (OTA phase C, C5.1)

The second page route. `agent-history/[worktreeId].tsx` already shipped in the
bundle with its own chunk, so listing it adds nothing to the download and moves
no route count: the switch is `index.tsx`'s, and the shell still decides, because
a bundle naming a grant this app lacks renders the native panel instead.

Its `.web.tsx` sibling is required for `index.web.tsx`'s reason — the native file
reaches OrcaMobileWebShellView, whose requireNativeViewManager runs at import and
takes the whole bundle down in a browser, since the manifest imports every route.

First route with two dynamic segments, so both are encoded. Grants are `navigate`
and `storage`: a resumed session opens the native session screen, the worktree
list now reaches this screen without leaving the page, and `app/h/_layout.tsx`
reads the app's own sidebar width above every page route.

The panel's router becomes `useRouteHandoff`, which is the seam that tells those
two apart: agent history is a page route and is pushed here, the session screen is
not and goes to the shell.

The three writes a resume makes needed no page-side handling and have none. What
they needed was a test that the descriptor's handling survives the extra hop, so
each is run through the bridge and against the same fake directly and the two
verdicts compared: a refused create raises the host's message, and a lost reply or
a shell disposed mid-flight stays delivery-unknown rather than becoming a failure
a user would retry blindly. No golden covers those three.

The flag census grows its first entry since C1.3, which is what it is for.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the agent-history Back button to the shell handoff

C5.1 wired this button by swapping the panel's `useRouter` for `useRouteHandoff`;
nothing else was needed, because `RouteHandoff` is the router's own shape and the
seam's web sibling decides `back`. So this commit is the test that would have
caught the wiring being absent, not a fix.

Red before the merge, green after, on the same four cases: on `e897e8123a`, where
`back` was still expo-router's own spread member, 3 failed and 1 passed — the one
that passed is the local-pop case, which is the branch C2.2 did not change. After
the merge brought in C2.2's `back`, all 4 pass. The pre-merge run named the notify
by its literal `'navigate-back'` because the contract constant did not exist yet;
it is the same string `BRIDGE_NAVIGATE_BACK_NOTIFY` holds, so the two runs asked
the same question.

Both module substitutions are the builder's own rather than conveniences: the web
bundle resolves `route-handoff` and `client-context` to their `.web` siblings, so
mocking each to its sibling gives this screen the module graph it has inside the
page. The frames are read off the port pair's lane rather than off a spy, and one
case asserts a frame crossed at all before either absence is read as an answer.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): compare the resume's second write, not only its first

Round 1, finding 1. All three `resumeAiVaultSessionInTerminal` cases settled the
create (`ai-vault-resume-launch.ts:158`), so `terminal.send` had never crossed the
bridge and the half of the resume that types the command into the pane was
uncompared. Three cases now drive both writes: a refused send raises the host's
message, an accepted send reporting `accepted: false` in-band says "Terminal input
is locked", and a send the host takes resolves — the last one being the presence
precondition, since a run that failed at the create would give the same shape of
verdict as one that failed at the send.

Reading `requests[1]` straight after settling the create finds nothing on the
bridged leg: the second write is made only once the first settles, so it is two
more lane round trips away. `nthRequest` waits instead, and says how many it saw
when it gives up, so this cannot pass by proving the opposite of what it says.

The locked reply is `{ send: { accepted: false } }`, not `{ accepted: false }`:
the reader is `reply.send?.accepted !== false` (`review-terminal-reply-schema.ts:65`),
and the flat shape resolves rather than throwing. Written the flat way first, both
legs agreed on "(resolved)", which is the comparison doing its job.

Also finding 1's second half: the file docstring claimed every case runs twice and
differences the verdicts, which was false for the dispose case — a fake RPC client
has no door to shut, so there is no native run to compare against. The docstring
now says so and the case carries the same note.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin what encoding cannot save about a dot-segment id

Round 1, finding 2. The route's docstring listed `/ ? #` and whitespace and the
encoding test covered five ids of that kind, which together implied encoding makes
any id safe. It does not: `encodeURIComponent('..')` is `'..'`, so the pathname
reaches `BRIDGE_ROUTE_PATHNAME_PATTERN` intact, fails the lookahead that stops a
climb out of `/h/` (`bridge-caps.ts:68`, read through `bridge-envelope.ts:117`),
and the shell answers with `reportShellFailure` — a failure screen where the route
would otherwise have rendered the native panel it already has.

Pinned, not fixed, and the docstring now says which. `app/h/[hostId]/index.tsx`
builds its pathname identically and has the same hole, so this series fixing one
of two call sites would leave the shape behind and stop describing it. The new
case asserts both halves — the segment survives encoding unchanged, and the
pattern refuses the pathname — so a later change that starts encoding dots fails
here and has to say which screen it wants instead.

Characterisation, so it was green on the first run rather than red: the claim is
about behaviour that already ships, and the value is that the refusal is on record.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): mount the agent-history route in a real browser

Round 1, finding 3. Nothing rendered this route's real module graph anywhere. The
unit tests mock react-native, safe-area, svg, lucide and the icon assets away —
they have to, since react-native is Flow source vitest cannot parse — so a
component in this closure with no web build would have reached a device before it
reached a test. The render check is the only place the graph meets React Native
Web, and this route was not in it.

Two cases. The first mounts the route from the shell double and reads the screen:
"Agent Session History" and the worktree label the params half carried, no fault,
no console error, no CSP refusal, and the URL the page wrote for itself. That also
proves `init.route.params` end to end on a route that has a dynamic segment too,
which §1 of the design claimed and nothing checked.

The second pins the chunk. C5 is the first series whose success path pulls a
second chunk after the first paint, which on iOS goes through WKURLSchemeHandler
under `script-src 'self'`. The chunk is named from the builder's own route map
rather than guessed from the bytes, and asserted absent from what the first route
loaded, so this says the route came over the wire now.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say the real lifetime of the route-handoff refusal set

Round 1, finding 4. The comment claimed one line per reason "for the life of one
client", borrowing `createPageDiagnosticReporter`'s bound. The set is built inside
the `useMemo` keyed on `[client, router]`, so it is per hook instance: in practice
the memo is not recomputed, because `useRouter()` is expo-router's module
singleton and the page holds one client, but every screen calling the hook gets
its own set and a reason can be reported once per screen rather than once per
document.

Says that now, and why it is not tightened: a per-module set would outlive the
page's client, which is the lifetime the rest of these reporters are scoped to,
and there is no document-wide reporter to join without reaching into a contract
file the C2 lane owns.

Records the other half of the finding too, which came back confirmed rather than
changed: `console.warn` is right here. It is the vocabulary `page-bootstrap.ts:35`
already writes in, and a `fault` notify would be wrong twice — the shell drops the
generation on a page fault, and a navigation the page declined is not a failure.

Comment only; no behaviour change, 25 navigation tests unchanged and green.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): wrap every router member that takes a target, not three of five

Round 2, finding 1. `...router` hands through everything this file does not name,
and two of the members it did not name take an href: `navigate` and `prefetch`.
`navigate` to a route outside `pageRoutes` went straight to expo-router and pushed
it into this document — the hole the tri-state exists to close, reopened under a
name nobody had looked at. No call site uses it today, which is why it shipped.

`navigate` is now wrapped exactly as `push` is: which of push-or-collapse it does
is a decision about this document's stack, and a target outside this document has
no such stack.

`prefetch` is decided the other way, explicitly. It is the one target-taker that
must never reach the shell: a prefetch is a background load, `navigate` is the
only thing the shell can be told, so handing one over would open a screen nobody
asked for. A route this document serves is prefetched here, which is what the
per-route chunk split makes worth doing; every other one is dropped without a
line, because a warm-up that did not happen is not a failure to report.

The docstring's "four members that can leave this document" is now five wrapped
members and a rule for which is which.

A list would rot, so the pin is derived: `HrefTakingRouterMember` reads the
parameter tuple of every member of `RouteHandoff` and `WRAPPED_HREF_MEMBERS` is
asserted equal to it in both directions. It reads the tuple rather than testing
assignability because `() => void` is assignable to `(href: RouterHref) => void`,
which would make `back`, `dismissAll` and `reload` target-takers and prove
nothing. Checked both ways: dropping `prefetch` from the list fails the compile
with "Type 'HrefTakingRouterMember' does not satisfy the constraint", and the
union resolves to exactly the five, with `back` and `setParams` outside it.

The pin is in the product module because `mobile/tsconfig.json` excludes tests.
The runtime test asserts each wrapped member is not the router's own function and
that `setParams` still is, so a hook that wrapped everything fails too.

Red first: 4 of the new cases fail against the previous file, 33 pass now.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): say per hook instance in the title too, not per client

Round 2, finding 2. The source comment was corrected in round 1 and this test's
title was not, so the two disagreed about the bound the refusal set actually has:
the set lives in the `useMemo`, so it is per hook instance, and a title claiming
per client is the stronger promise the code does not make.

Title only; the case and its assertions are unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say what the agent-history render case does not cover

Round 2, finding 3. The docstring claimed the case is where the panel's closure
meets React Native Web, which overstates it. The shell double answers no RPC, so
the session scan fails and the panel paints its "Unable to Load" state: the
session list, its rows, the resume button and the scope tabs never render, and a
render-time gap inside any of them would pass this check.

Now says both halves — import-time evaluation of every module in the closure and
the panel's own chrome are covered, the list subtree is not — and names what
covering the rest would take: a double that answers `aiVault.listSessions`, which
is a different instrument and would put domain behaviour in this file.

Text only. This case moves to its own file on the extracted harness after the
merge with #21592; the corrected text travels with it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): import the handoff module once in its own test

My round-2 fold added `WRAPPED_HREF_MEMBERS` as a second import of
`./route-handoff.web`, which `import(no-duplicates)` fails in the focused-plugins
pass of the changed-code gate. Joined to the existing import below the mocks,
which is where an import of the module under test has to sit in this file.

Found by running the changed-code gate rather than by review: mobile tsc, whole
tree oxlint and the suite were all green with it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): give agent-history its own render file on the extracted harness

The render check gained browser cases from three domain series at once, each
under the `.mjs` cap of 600 counted lines alone and no two of them together: at
39b15e395d the file was 781 raw lines and clean, main's was 786 and clean, and
their merge was 857 raw and 623 counted, which is the CI red on #21596. C1.10
extracted the harness so a domain gets a file instead. This is C5's, and the
render check is back to 639 raw lines and clean.

Five cases. The two that moved — the route mounts and paints, and its chunk is
fetched on navigation — plus three new ones.

Back, twice. With `navigate` granted the page's Back control posts exactly one
`navigate-back` notify and the page does not move; with the grant withheld the
same tap reaches the same handler and posts nothing. The pair is the point: the
document holds the single history entry the entry wrote with `replaceState`, so a
Back this page served itself would also have gone nowhere and looked identical.
This is the first proof of that handoff in a browser rather than against a mocked
router.

And a row. The harness's new `replies` lets the double answer named methods, so
the panel now renders a real session instead of its "Unable to Load" state, which
is the render-time gap the round-2 docstring conceded. Assertions are on the row's
own text and message count, plus the absence of both silent states — the scan
failing, and a session out of scope.

Replies lifted from the corpus, and one of them needed two scenarios. The session
and worktree lists are `aivault-history-screen-listed`'s. Its `status.get` is a
capability list alone, and the first run painted "Update Orca on your computer":
`HostProtocolGate` above every host route reads the same method for fields that
scenario never scripts. The status reply merges those from
`transport-host-status-gates-ready`, and the comment says why two.

`wt-history` is load-bearing, not incidental. The panel opens on the `workspace`
scope and filters by paths from the worktree list, so on any other worktree these
same replies paint "No agent sessions" — green, and proving nothing.

Registered in the `mobile_web_app` job beside the drawer check, which is the job
that makes a missing mobile install fail rather than skip.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): encode the host id at every href the host page builds

Pullfrog on #21596. My earlier commit fixed the two hrefs in the row's navigation
sheet and stopped there; five more sites in the same page interpolate the decoded
id raw — Accounts and Tasks in both header layouts
(`host-screen-header.tsx:188,204,307,318`) and the session target
`openWorktreeSession` builds (`use-host-worktree-actions.ts:192`).

Same C1.2 class. The persisted host store admits any non-empty id and both
`useLocalSearchParams` and the store answer it decoded, so one carrying `/`, `?`,
`#` or whitespace stops being the single segment `matchesRoutePattern` reads.
Inside the page that decides where a tap goes, because the handoff matches the
pathname against the shell's `pageRoutes` before choosing this document or the
native stack.

A census rather than five more assertions: the failure is a habit, not a bug —
each of these was written by copying the one beside it, and the seventh will be
too. It counts `/h/${...}` interpolations across the host page's four source
files and requires `encodeURIComponent` at each, with a presence check so it
cannot pass on an empty list.

Two sites are exempt and stay raw: `use-host-worktree-actions.ts:171` and
`app/h/_layout.tsx:100` compare against a pathname the router answers rather than
building a link, so encoding them would change what a comparison matches instead
of what a tap opens. The exemption is subtracted by count rather than matched
away, so a file that lost its comparison and gained a raw target does not come
out even.

ONE BEHAVIOURAL EDGE, named rather than fixed. `navigateFromHostList` short
-circuits when `pathname` equals the target minus its query. That comparison now
has an encoded target on one side and whatever `usePathname()` answers on the
other, so for a host id that needs encoding the short-circuit stops firing and a
tap on the screen you are already on re-navigates instead of doing nothing. It is
a redundant navigation, not a wrong one, and the guard at :171 is unaffected
because it compares against the same raw form it always did. Left alone because
fixing it means deciding what `usePathname()` returns for an encoded segment,
which is a question worth its own change.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): guard the optional host id the session target encodes

The commit before this one did not typecheck: `useHostWorktreeActions` takes
`hostId` as `string | undefined`, and `encodeURIComponent` does not. I committed
on a green test run without waiting for `tsc`, which is my error and the reason
this is a second commit rather than an amend — the lane forbids rewriting a
commit that exists.

`?? ''` rather than a cast or a non-null assertion. An absent id then builds
`/h//session/...`, an empty segment the shell's own route rule refuses, instead
of the string "undefined", which that rule would accept as a host genuinely named
undefined. Every other member of this hook already guards the same field.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the agent-history route native when the bridge would refuse its id

CodeRabbit on #21596. The persisted host store admits any non-empty id, so `.` or
`..` reaches this route, survives `encodeURIComponent` unchanged, and fails the
bridge's own segment rule. The route handed it over anyway: `bridge-host.ts`
parses the route against `BridgeInitRouteSchema`, drops it to null when it fails,
and the page answers an `init` naming no screen with "Update Orca to open this
workspace". A failure screen, in place of the native panel sitting right behind
this switch.

The route asks the schema first now and stays native when the answer is no, which
is where every route starts. Mirrors C3.1's call for the files routes
(`69e618e19a`), including its reason for using the schema rather than a copy of
its bounds: two spellings of one rule drift, and the half that matters is the
half the page reads.

The pin moves with it. It characterised the refusal before — asserting the
pathname was built and that the pattern rejected it — and now asserts the native
render, for a dot host id and for a dot worktree id, which is the other segment
and was never covered.

`app/h/[hostId]/index.tsx` has the same hole and is not fixed here, as asked: it
builds its pathname the same way and hands it over unchecked. When C3.1 is also
on main the two guards and `mobile-file-shell-route.ts` belong in one module
beside the schema, rather than a third spelling of a one-line call.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): forward navigation options through the wrapped router members

CodeRabbit (major) on #21596. expo-router's `push`, `replace`, `navigate` and
`dismissTo` are `(href, options?)`, and the wrappers took the href alone. A local
push asking for `{ withAnchor: false }` reached the router without it, so inside
the page the router did something other than what the caller wrote — silently,
because dropping an optional argument is not a type error.

Each wrapper forwards both on its local branch now. Nothing in this tree passes
options today, which is why it went unnoticed and exactly why it needed pinning:
the first caller to pass one would have had it dropped without a word.

Options do not cross to the shell, and the docstring says so rather than leaving
it to be discovered. The `navigate` notify carries an href and nothing else, so a
target handed over is opened by the native stack on that stack's own terms. That
is the right shape — the options describe a push inside a document the shell's
target is not in — but it is a loss, and a loss worth naming.

Four existing assertions moved from `toHaveBeenCalledWith(href)` to
`(href, undefined)`. That is what the router now receives when a caller passes
none, and expo-router reads an undefined second argument as absent; the comment
above them says so, so the next reader does not take it for a bug.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): count the wrapped members the way the returned object does

Pullfrog on #21596. The header still described the set as it stood before
`860577cc30`: "the four members that can leave this document", "the three that
carry a target", "the other three". There are six wrapped members now and five
carry a target, so every count in the paragraph was one or two short and a reader
checking the object against the prose would have found neither explained.

Now says six wrapped, five target-takers named and pinned by
`WRAPPED_HREF_MEMBERS`, four decided by the shell's route list, `prefetch` the
fifth and decided differently for a reason the member's own comment gives, and
`back` the sixth carrying no target at all.

Comment only; 36 navigation tests unchanged and green.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-19 05:15:46 -04:00

1177 lines
54 KiB
YAML

name: PR Checks
on:
pull_request:
types:
- opened
- synchronize
- reopened
- ready_for_review
concurrency:
group: pr-checks-${{ github.event.pull_request.number }}
cancel-in-progress: true
permissions:
contents: read
jobs:
# Why: a README/docs-only PR used to start the full matrix (test shards,
# two package jobs, typecheck, git compat, xterm, shell contracts). Path
# filters on `on.pull_request` would drop the `verify` check entirely; this
# detector keeps verify as the required aggregate and skips the expensive jobs.
# Per-job outputs also skip git-compat/xterm/packaging/shell when those
# inputs are unchanged; empty diffs fail closed and run everything.
code_paths:
name: detect code-relevant changes
runs-on: ubuntu-latest
outputs:
should_run: ${{ steps.filter.outputs.should_run }}
native_cache_changed: ${{ steps.filter.outputs.native_cache_changed }}
mobile_dependencies: ${{ steps.filter.outputs.mobile_dependencies }}
mobile_web_app: ${{ steps.filter.outputs.mobile_web_app }}
static_analysis: ${{ steps.filter.outputs.static_analysis }}
typecheck: ${{ steps.filter.outputs.typecheck }}
git_compatibility: ${{ steps.filter.outputs.git_compatibility }}
codex_index_heal_contract: ${{ steps.filter.outputs.codex_index_heal_contract }}
xterm_patch_sync: ${{ steps.filter.outputs.xterm_patch_sync }}
shell_contracts: ${{ steps.filter.outputs.shell_contracts }}
test: ${{ steps.filter.outputs.test }}
orcad_browser: ${{ steps.filter.outputs.orcad_browser }}
cross-version-wire: ${{ steps.filter.outputs.cross-version-wire }}
managed_hook_node18: ${{ steps.filter.outputs.managed_hook_node18 }}
package: ${{ steps.filter.outputs.package }}
package_windows: ${{ steps.filter.outputs.package_windows }}
e2e_should_run: ${{ steps.e2e_filter.outputs.should_run }}
test_files: ${{ steps.e2e_filter.outputs.test_files }}
ssh_source_changed: ${{ steps.e2e_filter.outputs.ssh_source_changed }}
native_ime_source_changed: ${{ steps.e2e_filter.outputs.native_ime_source_changed }}
wsl_source_changed: ${{ steps.e2e_filter.outputs.wsl_source_changed }}
steps:
- name: Checkout
uses: actions/checkout@v6
with:
# Why blob:none: full history is needed for the merge-base diff, but historical
# file contents are not. Blobs are ~89% of this repo's pack, and Git fetches the
# few this job actually reads on demand.
fetch-depth: 0
filter: blob:none
persist-credentials: false
- name: Classify changed paths
id: filter
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
run: |
set -euo pipefail
# Why --no-renames: name-only rename detection can report only the destination.
# A code file moved under docs/ must still expose its code-side deletion.
CHANGED="$(git diff --name-only --no-renames --diff-filter=ACDMR --merge-base "$BASE_SHA" "$HEAD_SHA")"
echo "Changed paths:"
printf '%s\n' "$CHANGED"
printf '%s\n' "$CHANGED" | node config/scripts/pr-code-change-scope.mjs | tee -a "$GITHUB_OUTPUT"
# Reuse the path-detector checkout instead of queuing another runner.
- name: Filter changed E2E specs
id: e2e_filter
if: github.event.pull_request.draft != true && steps.filter.outputs.should_run == 'true'
run: |
set -euo pipefail
BASE="${{ github.event.pull_request.base.sha }}"
HEAD="${{ github.event.pull_request.head.sha }}"
CHANGED="$(git diff --name-only --diff-filter=AMCR --merge-base "$BASE" "$HEAD")"
# Source routes are executable contracts so a test can prove exact
# authorities, exclusions, and sentinels without evaluating workflow shell.
TEST_FILES_JSON="$(printf '%s\n' "$CHANGED" | node config/scripts/pr-e2e-source-routing.mjs)"
echo "test_files=$TEST_FILES_JSON" >> "$GITHUB_OUTPUT"
# Why a separate signal: the Docker-SSH lane must trigger on SSH source, not on a
# spec name surviving in a route's list. Same routes, so the two cannot drift.
SSH_SOURCE_CHANGED="$(printf '%s\n' "$CHANGED" | node config/scripts/pr-e2e-source-routing.mjs --ssh-source)"
echo "ssh_source_changed=$SSH_SOURCE_CHANGED" >> "$GITHUB_OUTPUT"
echo "SSH source changed: $SSH_SOURCE_CHANGED"
# Why its own signal: the real-IME lane is a whole ibus session, not a spec, so it must
# trigger on IME source rather than on a spec name in some route's list.
NATIVE_IME_SOURCE_CHANGED="$(printf '%s\n' "$CHANGED" | node config/scripts/pr-e2e-source-routing.mjs --native-ime-source)"
echo "native_ime_source_changed=$NATIVE_IME_SOURCE_CHANGED" >> "$GITHUB_OUTPUT"
WSL_CHANGED="$(git diff --name-only --no-renames --diff-filter=ACDMR --merge-base "$BASE" "$HEAD")"
WSL_SOURCE_CHANGED="$(printf '%s\n' "$WSL_CHANGED" | node config/scripts/pr-e2e-source-routing.mjs --wsl-source)"
echo "wsl_source_changed=$WSL_SOURCE_CHANGED" >> "$GITHUB_OUTPUT"
echo "Native IME source changed: $NATIVE_IME_SOURCE_CHANGED"
SHOULD_RUN="$(printf '%s\n' "$CHANGED" | node config/scripts/pr-e2e-source-routing.mjs --reusable-workflow)"
if [ "$SHOULD_RUN" = true ]; then
echo "should_run=true" >> "$GITHUB_OUTPUT"
echo "Changed E2E specs: $TEST_FILES_JSON"
else
echo "should_run=false" >> "$GITHUB_OUTPUT"
echo "No specs requiring the reusable E2E workflow"
fi
static_analysis:
name: static analysis
needs: [code_paths]
if: needs.code_paths.outputs.static_analysis == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
# Why blob:none: full history is needed for the merge-base diff, but historical
# file contents are not. Blobs are ~89% of this repo's pack, and Git fetches the
# few this job actually reads on demand.
fetch-depth: 0
filter: blob:none
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: node
cache-dependency-path: |
pnpm-lock.yaml
mobile/pnpm-lock.yaml
- name: Lint
run: pnpm exec oxlint --format github
- name: Reject low-evidence patterns
run: pnpm run audit:anti-slop
- name: Enforce focused code-quality plugins
run: pnpm run audit:code-quality:native
- name: Enforce type-aware code-quality baseline
run: pnpm run audit:code-quality:type-aware
# Why here: the changed-code gate lints mobile files too, and its type-aware pass
# resolves types from mobile/node_modules. Without the install every mobile type
# degrades to an `error` type — reported as phantom findings against the changed lines.
- uses: ./.github/actions/install-mobile-dependencies
if: needs.code_paths.outputs.mobile_dependencies == 'true'
- name: Enforce changed-code quality
run: pnpm run check:code-quality:changed -- "${{ github.event.pull_request.base.sha }}"
- name: Enforce React Doctor on changed lines
run: pnpm run check:react-doctor:changed -- "${{ github.event.pull_request.base.sha }}"
- name: Check Zustand selector fan-out budget
run: pnpm run check:zustand-selector-fanout
- name: Check reliability gate manifest
run: pnpm run check:reliability-gates
- name: Enforce dead design-system classes
run: pnpm run check:dead-classes
- name: Check VM runtime rollback compatibility
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
run: |
if git diff --quiet --merge-base "$BASE_SHA" "$HEAD_SHA" -- \
src/shared/ephemeral-vm-runtime-store.ts \
src/shared/ephemeral-vm-runtime-feature-store.ts \
src/shared/ephemeral-vm-runtime-rollback-projection.ts \
src/shared/ephemeral-vm-runtimes.ts \
src/shared/ephemeral-vm-recipes.ts \
src/shared/orca-yaml-hook-types.ts \
src/main/ephemeral-vm-runtime-service.ts \
src/main/ephemeral-vm-runtime-provisioning-persistence.ts \
src/main/ephemeral-vm-failed-start-cleanup.ts; then
echo "VM runtime persistence is unchanged."
exit 0
fi
node config/scripts/run-ephemeral-vm-runtime-store-rollback-repro.mjs \
config/scripts/ephemeral-vm-runtime-store-cross-version.test.ts
- name: Enforce max-lines ratchet
run: pnpm run check:max-lines-ratchet
- name: Enforce ts-nocheck ratchet
run: pnpm run check:ts-nocheck-ratchet
- name: Enforce runtime Electron-import ratchet
run: pnpm run check:runtime-electron-ratchet
# Why both: the ratchet proves nothing reachable from the runtime imports electron,
# which is a property of the import graph. This proves the Node artifact it enables
# actually boots, pairs, creates a worktree and round-trips a real PTY.
- name: Boot orcad and round-trip a terminal
run: pnpm run smoke:orcad-terminal
- name: Verify the generated RPC params catalog
run: pnpm run verify:rpc-params-catalog
- name: Verify bundled skill guides
run: pnpm run verify:bundled-skill-guides
- name: Verify skill freshness manifest
run: pnpm run verify:skill-bundle-manifest
- name: Verify localization catalog
run: pnpm run verify:localization-catalog
# Why: the renderer ships only the English entries i18next cannot rebuild
# from each call site's inline default, so the generated subset has to
# track en.json and those defaults.
- name: Verify runtime-required localization catalog
run: pnpm run verify:localization-runtime-catalog
# Why: extraction writes sorted evidence to an isolated temporary path,
# so feature PRs need one normalized AST pass rather than a three-OS matrix.
- name: Verify localization extraction
run: pnpm run verify:localization-extraction
- name: Verify localization coverage
run: pnpm run verify:localization-coverage
# Why: project-owned type declarations must live in .ts so tsc
# actually checks them. TypeScript's skipLibCheck: true (inherited
# from @electron-toolkit/tsconfig) silently widens unresolved names
# in .d.ts to `any`, which is how #1186 shipped a broken IPC signature
# past typecheck. See .github/CONTRIBUTING.md#type-declarations-prefer-ts-over-dts.
- name: Guard against project-owned .d.ts in preload/shared
run: |
matches=$(find src/preload src/shared -name '*.d.ts' 2>/dev/null || true)
if [ -n "$matches" ]; then
echo "::error::Project-owned .d.ts files are not allowed under src/preload or src/shared."
echo "Move type declarations into a .ts file so skipLibCheck does not hide errors."
echo "See .github/CONTRIBUTING.md#type-declarations-prefer-ts-over-dts."
echo "Found:"
echo "$matches"
exit 1
fi
- name: Check feature wall asset budget
run: pnpm check:feature-wall-assets
- name: Verify macOS entitlements
run: pnpm verify:macos-entitlements
root_directory_guard:
name: root directory guard
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
# Why blob:none: full history is needed for the merge-base diff, but historical
# file contents are not. Blobs are ~89% of this repo's pack, and Git fetches the
# few this job actually reads on demand.
fetch-depth: 0
filter: blob:none
persist-credentials: false
- name: Reject new root-level files and folders
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
run: node .github/scripts/check-root-directory-entries.mjs "$BASE_SHA" "$HEAD_SHA"
# Why here: the READMEs embed media owned by docs/site and resources/onboarding,
# and the classifier skips static_analysis for docs-only diffs. This job runs
# on every PR and needs no install.
- name: Check README local links
run: node config/scripts/check-readme-local-links.mjs
typecheck:
needs: [code_paths]
if: needs.code_paths.outputs.typecheck == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
# Why: every project is `composite`, so tsc already writes a .tsbuildinfo that lets
# the next run skip unchanged files. Share one cache entry across commits while the
# PR base stays stable; actions/cache keeps the first successful graph and the
# compiler still invalidates stale files from its content hashes.
- name: Cache TypeScript incremental state
uses: actions/cache@v5
with:
path: config/*.tsbuildinfo
key: tsbuildinfo-${{ runner.os }}-${{ hashFiles('pnpm-lock.yaml', 'config/tsconfig*.json') }}-${{ github.event.pull_request.base.sha }}
restore-keys: |
tsbuildinfo-${{ runner.os }}-${{ hashFiles('pnpm-lock.yaml', 'config/tsconfig*.json') }}-
- run: pnpm run typecheck
git_compatibility:
name: Git compatibility
needs: [code_paths]
if: needs.code_paths.outputs.git_compatibility == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
# Why: the 2.25.5 lane is a source build of a pinned tarball, so it produced the
# same binary on every PR for minutes of runner time. The key carries the version
# because that is the only input; the sha256 assertion below still guards the
# tarball on the miss path that actually builds. Only this PR's own later pushes
# can restore it — GitHub scopes a cache written from a pull_request run to that
# ref — so a first push always takes the build path below.
- name: Cache baseline Git build
uses: actions/cache@v5
with:
path: ~/.cache/orca-git-compat/git-2.25.5
key: git-compat-baseline-${{ runner.os }}-${{ runner.arch }}-2.25.5
# Why its own step: this is `make -j$(nproc)` on every core, and the lanes below
# spend their wall clock waiting on container starts, not on Git. Sharing a runner
# with the build stretched one ~1.5s boundary case past Vitest's 30s timeout, so
# the build has to finish before anything timed starts.
- name: Build the baseline Git binary
run: |
archive="$RUNNER_TEMP/git-2.25.5.tar.gz"
source="$HOME/.cache/orca-git-compat/git-2.25.5"
if [ -x "$source/git" ]; then
exit 0
fi
curl -fsSL https://www.kernel.org/pub/software/scm/git/git-2.25.5.tar.gz -o "$archive"
echo "41662c52fc16fec4963bfc41075e71f8ead6b5e386797eb6f9a1111ff95a8ddf $archive" \
| sha256sum --check
mkdir -p "$source"
tar -xzf "$archive" -C "$source" --strip-components=1
make -C "$source" -j"$(nproc)" \
NO_GETTEXT=YesPlease NO_TCLTK=YesPlease NO_PYTHON=YesPlease git
# Why: the linked binaries are what the next run needs; the objects that
# produced them are most of the tree and would bloat the cache entry.
find "$source" -name '*.o' -delete
- name: Verify Git binary compatibility matrix
run: |
specs=(
"alpine/git:edge-2.38.1|2.38.1"
"alpine/git:v2.49.1|2.49.1"
)
# Why pull up front: a lane's first `docker run` otherwise pulls its image
# while the sibling lane is mid-test, and that stall is charged to the test.
for spec in "${specs[@]}"; do
docker pull --quiet "${spec%%|*}"
done
pids=()
(
ORCA_GIT_COMPAT_BINARY="$HOME/.cache/orca-git-compat/git-2.25.5/git" \
ORCA_GIT_COMPAT_VERSION="2.25.5" \
pnpm exec vitest run --config config/vitest.config.ts \
src/shared/git-binary-compatibility.test.ts
) &
pids+=("$!")
for spec in "${specs[@]}"; do
(
image="${spec%%|*}"
version="${spec#*|}"
ORCA_GIT_COMPAT_IMAGE="$image" ORCA_GIT_COMPAT_VERSION="$version" \
pnpm exec vitest run --config config/vitest.config.ts \
src/shared/git-binary-compatibility.test.ts
) &
pids+=("$!")
done
status=0
for pid in "${pids[@]}"; do
wait "$pid" || status=1
done
exit "$status"
# Why this job: Orca's session index-heal depends on a Codex behavior — a
# `thread/read` of an unindexed rollout performs a read-repair that inserts the
# `threads` row. Every unit test drives a stub app-server and asserts only that the
# call did not error, so if Codex dropped the repair they would all stay green while
# the subsystem went inert. This runs the pinned real binary and fails when the
# repair stops happening. Pinned because the binary is the thing expected to drift.
codex_index_heal_contract:
name: Codex index-heal contract
needs: [code_paths]
if: needs.code_paths.outputs.codex_index_heal_contract == 'true'
runs-on: ubuntu-latest
env:
CODEX_CLI_VERSION: '0.150.1'
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
- name: Install pinned Codex CLI
run: |
set -euo pipefail
npm install --no-audit --no-fund --prefix "$RUNNER_TEMP/codex-cli" \
"@openai/codex@$CODEX_CLI_VERSION"
- name: Verify Codex index-heal contract
env:
# Why REQUIRED: without a binary the suite skips, and a job that skips
# reports success. This turns a failed or missing install into a red test
# instead of a green no-op.
ORCA_CODEX_CONTRACT_REQUIRED: '1'
ORCA_CODEX_CONTRACT_VERSION: ${{ env.CODEX_CLI_VERSION }}
run: |
set -euo pipefail
ORCA_CODEX_CONTRACT_BINARY="$RUNNER_TEMP/codex-cli/node_modules/.bin/codex" \
pnpm exec vitest run --config config/vitest.config.ts \
src/main/codex/codex-index-heal-binary-contract.test.ts
xterm_patch_sync:
name: xterm patch sync
needs: [code_paths]
if: needs.code_paths.outputs.xterm_patch_sync == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
# Why: the check rebuilds every package in the manifest from a pinned upstream
# commit — @xterm/xterm and its three addons, each built twice (once unmodified to
# prove the toolchain still reproduces the published bundles, once patched). Caching
# the npm metadata and the shallow clone keeps the repeated cost to the builds
# themselves; the key is the manifest, so a commit, package or toolchain bump
# invalidates it.
- name: Restore upstream xterm build inputs
uses: actions/cache@v5
with:
path: |
~/.npm
${{ runner.temp }}/xterm-patch-build/upstream/.git
key: xterm-upstream-${{ hashFiles('config/patches/xterm-upstream.json') }}
- name: Verify xterm patches match the pinned upstream build
env:
WORK_DIR: ${{ runner.temp }}/xterm-patch-build
run: node config/scripts/regenerate-xterm-patches.mjs --check --work-dir="$WORK_DIR"
shell_contracts:
name: shell contracts
needs: [code_paths]
if: needs.code_paths.outputs.shell_contracts == 'true'
runs-on: ubuntu-latest
# Why: this job's cost is almost entirely package download, and a stalled mirror has
# no wall-clock bound of its own. A successful run finishes in ~4.5 minutes, so this
# is generous; it exists so a wedge fails the job instead of holding the whole run
# open for the 6h GitHub default — which also blocks `gh run rerun --failed`.
timeout-minutes: 15
env:
# Why: the suites below gate their live fish tests on the binary, which is
# right on a developer machine and wrong here — this job is a required check
# and its fish lane is the only end-to-end guard for #9993, so a skip would
# report green with nothing exercised. Turns those skips into failures.
ORCA_REQUIRE_FISH: '1'
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
# Why fish: shell-ready.test.ts gates its live fish test on the binary being
# present, so without this the fish barrier is only covered by config-shape
# assertions and never actually exercised.
# Why release-4: DECSET 2031 arming lives in the fish 4.0 Rust tty_handoff, and
# fish-color-scheme-child-stdin.node-pty.test.ts (#9993) needs it. Noble ships
# 3.7, so the PPA is what makes that lane real.
- name: Install zsh and fish
run: |
# Why the update/PPA/fish steps are tolerant: a repo the runner image already
# ships can lack a Release file for this suite, and a failed add-apt-repository
# still leaves its list entry behind — either makes `apt-get update` exit
# non-zero and would red this required check over something unrelated to the
# PR. Every fish outcome is judged by the version gate below instead, so only
# the zsh install (which has no such gate) stays fatal here.
# Why retry only here: adding the PPA is the network-flaky step, and the
# version gate below is fatal, so a transient Launchpad blip would
# otherwise red a required check on PRs unrelated to shells.
# Why -n: add-apt-repository refreshes every configured repo on its own. With
# an update on each side of it this step refreshed them three times over, and
# the Azure archive mirror alone costs ~15-30s a pass. The PPA index is the
# only thing the repo list gains here, and the single update below fetches it.
# Why bound acquisition: measured on a *passing* run, this step spent 40s
# fetching 11.4 MB of index and then 2m17s fetching 8.9 MB of packages at
# 65 kB/s — it is dominated by download throughput, not by work. apt applies
# no wall-clock bound to a stalled mirror, so a slow Launchpad or archive
# host wedges the step for tens of minutes. This job is a required check, so
# a wedge holds the entire run open and blocks `gh run rerun --failed`.
# Bounded timeouts plus retries turn an unbounded hang into a fast, legible
# failure. Set in apt.conf.d rather than on each command line so the two
# invocations below stay exactly as pr-workflow-parallelism.test.mjs parses
# them. Retries are 1, not 3: a first attempt at these bounds already multiplied
# 30s x 3 retries across every index file into a ~15 minute stall on a dead
# mirror, which is worse than failing once and moving on.
sudo tee /etc/apt/apt.conf.d/99-orca-shell-contracts >/dev/null <<'APTCONF'
Acquire::http::Timeout "15";
Acquire::https::Timeout "15";
Acquire::Retries "1";
APTCONF
for attempt in 1 2 3; do
sudo add-apt-repository -y -n ppa:fish-shell/release-4 && break
echo "add-apt-repository attempt ${attempt} failed; retrying" >&2
sudo add-apt-repository -y -n -r ppa:fish-shell/release-4 || true
sleep 5
done
# Why a wall-clock bound on each command: apt's Acquire timeouts are per-connection,
# so a dead mirror costs timeout x retries x every index file. Measured: the archive
# mirror stalled with zero bytes and the step burned 14m26s before the job bound
# killed it. `timeout` is the only thing that bounds the command as a whole.
# The update is already tolerant by design (see above), so bounding it just caps
# what a dead mirror can cost before the install runs against whatever index exists.
timeout 120 sudo apt-get update || true
# Why both shells on one line: pr-workflow-parallelism.test.mjs parses only the
# first install command in this step to prove the lane really installs them.
timeout 300 sudo apt-get install -y zsh fish
# Separate from the install so the failure names the contract, not an apt error.
# ORCA_REQUIRE_FISH re-checks this at test time; this step just fails in seconds
# instead of after a full dependency install.
- name: Require fish 4+
run: |
version="$(fish --version 2>/dev/null || true)"
major="${version##*version }"
major="${major%%.*}"
case "$major" in '' | *[!0-9]*) major=0 ;; esac
echo "${version:-<fish not installed>}"
if [ "$major" -lt 4 ]; then
echo "::error::shell contracts needs fish 4+ (DECSET 2031 arming, #9993) but got '${version:-none}'. Fix the ppa:fish-shell/release-4 install rather than letting the fish lane skip." >&2
exit 1
fi
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: node
- name: Test real shell contracts
run: |
pnpm exec vitest run --config config/vitest.config.ts --maxWorkers=1 \
src/main/daemon/repro-13767-shell-ready-marker-lost-to-exec.test.ts \
src/main/daemon/shell-ready.test.ts \
src/main/daemon/node-pty-fd-leak.test.ts \
src/main/providers/local-pty-shell-ready-zsh-launch-environment.test.ts \
src/main/providers/__tests__/shell-ready-framework-example.test.ts \
src/main/pty/codex-shell-launch-preflight.test.ts \
src/main/pty/omp-shell-wrapper-alias-safety.test.ts \
src/main/pty/omp-shell-wrapper.node-pty.test.ts \
src/main/shell-startup-feature-channel.test.ts \
src/main/terminal-history-fish-session.node-pty.test.ts \
src/main/zsh-scoped-histfile.live-shell.test.ts \
src/main/zsh-startup-hook-user-config-equivalence.live-shell.test.ts \
src/main/zsh-wrapper-version-mismatch.live-shell.test.ts \
src/renderer/src/components/terminal-pane/fish-color-scheme-child-stdin.node-pty.test.ts \
src/shared/fish-query-reply-child-stdin.node-pty.test.ts \
src/shared/pty-reply-echo-shapes.node-pty.test.ts \
src/shared/startup-shell-portability.live-shell.test.ts \
src/shared/posix-command-path-lookup.test.ts
# Cache-key input changes would otherwise make every shard compile the same
# native addon concurrently. Prime the supported Node ABI before the matrix fans out.
test_native_cache:
name: prepare test native cache node 24
needs: [code_paths]
if: needs.code_paths.outputs.native_cache_changed == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: node
node-version: '24'
test:
needs: [code_paths, test_native_cache]
if: >-
always() &&
needs.code_paths.outputs.test == 'true' &&
(needs.test_native_cache.result == 'success' || needs.test_native_cache.result == 'skipped')
uses: ./.github/workflows/unit-tests.yml
with:
node_versions: '["24"]'
# Why a separate job: the test needs a real Chrome, and the sharded `test` matrix
# would pay for it on every shard to run one file in whichever shard it landed in.
orcad_browser:
name: orcad browser provider
needs: [code_paths]
if: needs.code_paths.outputs.orcad_browser == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
# Why no native-runtime: the provider drives the prebuilt agent-browser binary
# shipped in node_modules and never touches node-pty.
- uses: ./.github/actions/install-node-dependencies
# Why the runner's Google Chrome and not its chromium: Ubuntu 24.04 only ships an
# AppArmor userns profile for the Chrome .deb, so chromium dies with "No usable
# sandbox" and the provider passes no --no-sandbox. Why fail instead of skip: an
# unset ORCA_BROWSER_EXECUTABLE is exactly how this test went uncovered for so long.
- name: Resolve Chrome for the browser provider
run: |
set -euo pipefail
chrome="$(command -v google-chrome || command -v google-chrome-stable || true)"
if [ -z "$chrome" ]; then
echo "::error::No Google Chrome on the runner; the browser provider test would silently skip."
exit 1
fi
"$chrome" --version
echo "ORCA_BROWSER_EXECUTABLE=$chrome" >> "$GITHUB_ENV"
- name: Test external Chromium browser provider
run: |
pnpm exec vitest run --config config/vitest.config.ts \
src/main/orcad/external-chromium-browser-process.integration.test.ts
# Why its own job: it needs mobile/node_modules and a real browser, and the sharded `test`
# matrix would pay for both on every shard to run two files. Dark through Phase C: this proves
# `build:mobile-web:app` on every PR that touches the page, and ships nothing -- packaging still
# builds the Phase A bootstrap via build:mobile-web.
mobile_web_app:
name: mobile web app bundle
needs: [code_paths]
if: needs.code_paths.outputs.mobile_web_app == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
# Why no native-runtime: the builder is esbuild and the render check is a browser. Nothing
# in this job loads node-pty.
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: node
cache-dependency-path: |
pnpm-lock.yaml
mobile/pnpm-lock.yaml
# The entry lives in mobile/ so one React resolves; without this every RN import is nothing.
- uses: ./.github/actions/install-mobile-dependencies
# Why the runner's Google Chrome and not a downloaded chromium: same reason as the orcad
# browser job -- Ubuntu 24.04 only ships an AppArmor userns profile for the Chrome .deb.
# Why fail instead of skip: a silently skipped render check is the failure this job exists
# to prevent.
- name: Resolve Chrome for the render check
run: |
set -euo pipefail
chrome="$(command -v google-chrome || command -v google-chrome-stable || true)"
if [ -z "$chrome" ]; then
echo "::error::No Google Chrome on the runner; the render check would silently skip."
exit 1
fi
"$chrome" --version
echo "ORCA_MOBILE_WEB_RENDER_BROWSER=$chrome" >> "$GITHUB_ENV"
# The drawer check runs on WebKit as well as Chrome, because the shell's iOS WebView is
# WebKit and the Chrome above cannot stand in for it. Downloaded rather than resolved from
# the runner: Ubuntu ships no WebKit build to point at.
- name: Install WebKit for the drawer check
run: pnpm exec playwright install --with-deps webkit
- name: Build and verify the app bundle
run: pnpm run build:mobile-web:app
# The bundling tests skip themselves where mobile dependencies are absent, which is how they
# stay green in the sharded `test` job. This is the job that installs them, so here a missing
# install has to fail rather than skip everything the job exists to run.
- name: Builder, override census and render checks
env:
ORCA_MOBILE_WEB_APP_DEPS_REQUIRED: '1'
run: |
pnpm exec vitest run --config config/vitest.config.ts \
config/scripts/build-mobile-web-app-bundle.test.mjs \
config/scripts/mobile-web-app-web-overrides.test.mjs \
config/scripts/mobile-web-app-render.test.mjs \
config/scripts/mobile-web-app-drawer-render.test.mjs \
config/scripts/mobile-web-app-agent-history-render.test.mjs
cross-version-wire:
name: cross-version wire compatibility
needs: [code_paths]
if: needs.code_paths.outputs.cross-version-wire == 'true'
runs-on: ubuntu-latest
steps:
# Why fetch-depth 0: the harness extracts the newest release tag to skew
# current code against it. The default shallow clone has no tags, which is
# why this cannot ride along in the sharded `test` job.
- name: Checkout
uses: actions/checkout@v6
with:
# Why blob:none: full history is needed for the merge-base diff, but historical
# file contents are not. Blobs are ~89% of this repo's pack, and Git fetches the
# few this job actually reads on demand.
fetch-depth: 0
filter: blob:none
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: node
# A path filter that matches nothing exits 1 ("No test files found"), so this
# lane cannot report success while running zero tests.
- name: Old/new client and server compatibility journeys
run: >-
pnpm exec vitest run --config config/vitest.config.ts
tests/e2e/cross-version-wire/release-checkout.unit.test.ts
tests/e2e/cross-version-wire/cross-version-browser-placement.unit.test.ts
tests/e2e/cross-version-wire/cross-version-terminal-wire.unit.test.ts
tests/e2e/cross-version-wire/reported-lossy-initial-snapshot.unit.test.ts
tests/e2e/cross-version-wire/cross-version-agent-session-wire.unit.test.ts
tests/e2e/cross-version-wire/cross-version-worktree-identity-downgrade.unit.test.ts
tests/e2e/cross-version-wire/cross-version-session-tabs-retirement-proof.unit.test.ts
managed_hook_node18:
name: managed hooks on Node 18
needs: [code_paths]
if: needs.code_paths.outputs.managed_hook_node18 == 'true'
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- uses: ./.github/actions/install-node-dependencies
- name: Build relay companions
run: pnpm run build:relay
- name: Setup Node 18 runtime
uses: actions/setup-node@v6
with:
node-version: '18'
- name: Smoke managed-hook companions
run: node config/scripts/smoke-managed-hook-runtime-node18.mjs
package:
name: package
needs: [code_paths]
if: needs.code_paths.outputs.package == 'true'
runs-on: ubuntu-latest
# Let the serial Docker gates reach their own deadlines and report cleanup failures.
timeout-minutes: 90
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Cache electron-builder downloads
uses: actions/cache@v5
with:
path: ~/.cache/electron-builder
key: electron-builder-linux-${{ hashFiles('pnpm-lock.yaml') }}
restore-keys: |
electron-builder-linux-
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: electron
cache-dependency-path: |
pnpm-lock.yaml
mobile/pnpm-lock.yaml
# Why here: electron-builder's beforePack requires out/mobile-web, and the bundle
# build resolves React Native and Expo from mobile/node_modules.
- uses: ./.github/actions/install-mobile-dependencies
# Why --no-file-parallelism: every file here launches a full Electron stack twice, and each
# probe carries its own in-process deadline. Four at once on a 4-vCPU runner starve each other
# past those deadlines; serial, every probe owns the runner.
- name: Test Linux Electron lifecycle boundary
run: >-
xvfb-run --auto-servernum pnpm exec vitest run --config config/vitest.config.ts
--no-file-parallelism
src/main/browser/browser-client-page-renderer-lifecycle.electron.test.ts
src/main/browser/browser-route-tcp-egress.electron.test.ts
src/main/browser/browser-route-webrtc-egress.electron.test.ts
src/main/browser/browser-route-h3-egress.electron.test.ts
src/main/browser/browser-route-dns-prefetch.electron.test.ts
- name: Build package inputs
run: |
status=0
pnpm run build:cli || status=1
scripts=(build:relay build:electron-vite:parallel)
pids=()
for script in "${scripts[@]}"; do
pnpm run "$script" &
pids+=("$!")
done
for pid in "${pids[@]}"; do
wait "$pid" || status=1
done
exit "$status"
- name: Project web client from renderer build
run: pnpm run build:web-from-renderer
# Why here and not inside "Build package inputs": this job assembles packaging inputs step by
# step instead of calling build:release, and electron-builder's beforePack guard hard-fails
# without out/mobile-web.
- name: Build mobile web bundle
run: pnpm run build:mobile-web
- name: Build native components
run: pnpm run build:native
- name: Install Linux package tooling
run: sudo apt-get update && sudo apt-get install -y cpio rpm
- name: Package unpacked app
env:
ORCA_REUSE_PREPARED_NATIVE_RUNTIME: '1'
# PR artifacts are only inspected locally; gzip avoids release-size xz compression.
run: >-
pnpm exec electron-builder --config config/electron-builder.config.cjs
--linux AppImage deb rpm --x64 --publish never
--config.deb.compression=gz --config.rpm.compression=gzip
- name: Verify root-package marker payloads
run: |
set -euo pipefail
version="$(node -p "require('./package.json').version")"
deb="dist/orca-ide_${version}_amd64.deb"
rpm="dist/orca-ide-${version}.x86_64.rpm"
test -s "$deb"
test -s "$rpm"
deb_marker="$(dpkg-deb --fsys-tarfile "$deb" | tar -xOf - ./opt/Orca/resources/package-type)"
rpm_marker="$(rpm2cpio "$rpm" | cpio --quiet --extract --to-stdout ./opt/Orca/resources/package-type)"
[[ "$deb_marker" == deb ]] || { echo "Expected deb marker, got: $deb_marker"; exit 1; }
[[ "$rpm_marker" == rpm ]] || { echo "Expected rpm marker, got: $rpm_marker"; exit 1; }
- name: Verify headless serve signal shutdown
run: >-
node config/scripts/run-headless-serve-shutdown-docker.mjs
--appimage dist/orca-linux.AppImage --all-entrypoints
# A default container reproduces the hostile AppImage launch environment.
- name: Verify Linux CLI launch contract
run: node config/scripts/run-linux-cli-launch-contract-docker.mjs --appimage dist/orca-linux.AppImage
- name: Smoke packaged CLI
run: node config/scripts/smoke-packaged-cli.mjs --app-dir=dist/linux-unpacked
- name: Smoke packaged hang watchdog worker
run: xvfb-run --auto-servernum node config/scripts/smoke-packaged-hang-watchdog-worker.mjs --app-dir=dist/linux-unpacked
package_windows:
name: package (windows)
needs: [code_paths]
if: needs.code_paths.outputs.package_windows == 'true'
runs-on: windows-2022
timeout-minutes: 30
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
- name: Cache electron-builder downloads
uses: actions/cache@v5
with:
path: |
~\AppData\Local\electron\Cache
~\AppData\Local\electron-builder\Cache
key: electron-builder-windows-${{ hashFiles('pnpm-lock.yaml') }}
restore-keys: |
electron-builder-windows-
# Why persist-native-cache false: this job later rebuilds the same path for
# Electron. A post-job save would store the Electron ABI under the Node key.
- uses: ./.github/actions/install-node-dependencies
id: deps
with:
native-runtime: node
persist-native-cache: 'false'
cache-dependency-path: |
pnpm-lock.yaml
mobile/pnpm-lock.yaml
# Why here: electron-builder's beforePack requires out/mobile-web, and the bundle
# build resolves React Native and Expo from mobile/node_modules.
- uses: ./.github/actions/install-mobile-dependencies
- name: Save compiled Node native modules
if: steps.deps.outputs.native-cache-hit != 'true'
uses: actions/cache/save@v5
with:
path: |
node_modules/.pnpm/node-pty@*/node_modules/node-pty/build
native/windows-registry/build
node_modules/.pnpm/@vscode+windows-process-tre*/node_modules/@vscode/windows-process-tree/build
key: native-modules-${{ runner.os }}-${{ steps.deps.outputs.native-cache-scope }}-${{ runner.arch }}-node-node${{ steps.deps.outputs.node-version }}-${{ hashFiles('pnpm-lock.yaml', '.github/actions/install-node-dependencies/action.yml', 'config/scripts/ensure-native-runtime.mjs', 'config/scripts/rebuild-native-deps.mjs', 'config/patches/node-pty@1.1.0.patch', 'config/patches/@vscode__windows-process-tree@0.8.0.patch', 'native/windows-registry/src/addon.cc', 'native/windows-registry/binding.gyp', 'native/windows-registry/package.json') }}
# vitest runs here directly rather than through `pnpm test`, so the addon
# assertions only hold once install-node-dependencies has rebuilt natives.
- name: Test Windows-specific boundaries
run: >-
pnpm exec vitest run --config config/vitest.config.ts
config/scripts/rebuild-native-deps.test.mjs
config/scripts/rebuild-native-deps-windows-process-tree.test.mjs
config/scripts/rebuild-native-deps-node-pty.test.mjs
config/scripts/ensure-native-runtime-job-ownership.test.mjs
config/scripts/verify-packaged-node-pty-job-ownership.test.mjs
config/scripts/windows-pe-machine.test.mjs
config/scripts/script-module-dependencies.test.mjs
src/main/windows-registry-addon.test.ts
config/scripts/windows-process-tree-gyp-path.test.mjs
config/scripts/windows-process-tree-gyp-rebuild.test.mjs
config/scripts/package-electron-runtime-contract.test.mjs
config/scripts/electron-builder-runtime-resources.test.mjs
src/main/browser/browser-client-page-renderer-lifecycle.electron.test.ts
src/main/browser/browser-route-tcp-egress.electron.test.ts
src/main/browser/browser-route-webrtc-egress.electron.test.ts
src/main/browser/browser-route-h3-egress.electron.test.ts
src/main/browser/browser-route-dns-prefetch.electron.test.ts
src/main/providers/windows-conpty-wide-char-duplication.node-pty.test.ts
src/main/providers/pty-repaint-wide-char-buffer.node-pty.test.ts
src/shared/child-process/windows-command-line.win32.test.ts
src/shared/child-process/windows-cmd-shim-resolution.test.ts
src/shared/child-process/windows-cmd-shim-resolution.win32.test.ts
src/main/agent-hooks/windows-hook-payload-delivery.test.ts
src/main/agent-hooks/windows-direct-cmd-hook-command.test.ts
src/main/codex/windows-hook-command.test.ts
src/main/codex/windows-hook-upgrade.test.ts
src/main/windows/windows-pty-job.win32.test.ts
src/main/windows/windows-msys-job.win32.test.ts
src/main/windows/windows-host-job.win32.test.ts
src/main/windows/windows-process-tree-command-line-patch.test.ts
src/main/windows/windows-process-table-native-addon.win32.test.ts
src/main/windows-live-tree-kill.win32.test.ts
src/main/wsl/wsl-runner.test.ts
src/main/wsl/wsl-guest-environment.test.ts
src/main/wsl/wsl-invocation-boundary.test.ts
src/main/wsl/wsl-executable-path.win32.test.ts
src/main/wsl/wsl-w1-w3-contract.test.ts
src/shared/source-scan/source-tree-scan.test.ts
src/main/cli/wsl-cli-powershell-boundary.test.ts
src/main/computer/desktop-script-runtime-host.win32.test.ts
src/main/cursor/hook-service.test.ts
src/main/orca-profiles/profile-index-store.test.ts
src/main/startup/windows-install-dir-acl-repair.win32.test.ts
src/main/runtime/repo-worktree-admin-fingerprint.test.ts
src/main/runtime/worktree-scan-admin-fingerprint-gate.test.ts
src/shared/secure-file-fsync-flags.test.ts
src/shared/secure-path-windows-acl.win32.test.ts
src/main/runtime/unreadable-secret-store-preservation.win32.test.ts
src/main/ipc/pty-codex-account-attribution.test.ts
src/main/ipc/pty-spawn-env-codex-resume-provenance.test.ts
src/relay/windows-port-scan.win32.test.ts
# Why the :parallel variant: identical to build:release except the three
# electron-vite targets overlap instead of running back to back. The Linux package
# job already packages and smoke-tests an AppImage built that way.
- name: Cache Windows CLI launcher
uses: actions/cache@v5
with:
path: native/windows-cli-launcher/.build
key: windows-cli-launcher-${{ runner.os }}-${{ runner.arch }}-${{ hashFiles('native/windows-cli-launcher/**', 'config/scripts/build-windows-cli-launcher.mjs') }}
- name: Build package inputs
env:
ORCA_REUSE_WINDOWS_CLI_LAUNCHER: '1'
run: pnpm run build:release:parallel
- name: Restore compiled Electron native modules
uses: actions/cache@v5
with:
path: |
node_modules/.pnpm/node-pty@*/node_modules/node-pty/build
native/windows-registry/build
node_modules/.pnpm/@vscode+windows-process-tre*/node_modules/@vscode/windows-process-tree/build
key: native-modules-${{ runner.os }}-${{ steps.deps.outputs.native-cache-scope }}-${{ runner.arch }}-electron-node${{ steps.deps.outputs.node-version }}-${{ hashFiles('pnpm-lock.yaml', '.github/actions/install-node-dependencies/action.yml', 'config/scripts/ensure-native-runtime.mjs', 'config/scripts/rebuild-native-deps.mjs', 'config/patches/node-pty@1.1.0.patch', 'config/patches/@vscode__windows-process-tree@0.8.0.patch', 'native/windows-registry/src/addon.cc', 'native/windows-registry/binding.gyp', 'native/windows-registry/package.json') }}
- name: Prepare Electron native runtime
run: node config/scripts/ensure-native-runtime.mjs --runtime=electron
- name: Package unpacked app
env:
ORCA_REUSE_PREPARED_NATIVE_RUNTIME: '1'
run: pnpm exec electron-builder --config config/electron-builder.config.cjs --dir
- name: Smoke packaged Windows PTY native capability
run: pnpm run smoke:windows-pty-native-capability -- --exe=dist/win-unpacked/Orca.exe
- name: Smoke packaged CLI
run: node config/scripts/smoke-packaged-cli.mjs --app-dir=dist/win-unpacked
e2e:
name: e2e
needs: code_paths
if: needs.code_paths.outputs.e2e_should_run == 'true'
# Why: reusable e2e.yml only checkouts, builds, and uploads artifacts.
permissions:
contents: read
uses: ./.github/workflows/e2e.yml
with:
# The synthetic pull-request merge ref can disappear while this reusable
# workflow is queued. The head SHA is immutable and works for every PR.
ref: ${{ github.event.pull_request.head.sha }}
test_files: ${{ needs.code_paths.outputs.test_files }}
ssh_source_changed: ${{ needs.code_paths.outputs.ssh_source_changed }}
# Why this is not in verify's needs: it is the first PR-gate run of a harness whose reliability
# is only known from nightly main runs (20/20 green, 2026-08-09..2026-08-29, p50 3m25s). It
# reports a red X on the PR without blocking, exactly like `e2e` above. Deliberately no
# continue-on-error: that renders the check green and hides the signal it exists to give. To
# make it blocking, add it to verify.needs, add TERMINAL_IME_NATIVE to the env below, and
# require `success || skipped` outside the strict loop — see the note on `e2e`.
terminal_ime_native:
name: real IME
needs: code_paths
if: needs.code_paths.outputs.native_ime_source_changed == 'true'
# Why: the reusable workflow only checks out, builds, and uploads artifacts.
permissions:
contents: read
uses: ./.github/workflows/terminal-ime-e2e.yml
windows_wsl:
name: real WSL terminal
needs: code_paths
if: needs.code_paths.outputs.wsl_source_changed == 'true'
permissions:
contents: read
uses: ./.github/workflows/windows-wsl-e2e.yml
with:
ref: ${{ github.event.pull_request.head.sha }}
verify:
if: always()
needs:
- code_paths
- static_analysis
- root_directory_guard
- typecheck
- git_compatibility
- codex_index_heal_contract
- xterm_patch_sync
- shell_contracts
- test
- orcad_browser
- mobile_web_app
- cross-version-wire
- managed_hook_node18
- package
- package_windows
runs-on: ubuntu-latest
steps:
# Why: e2e is deliberately absent from needs. The suite is currently red on
# main (every scheduled run), so gating merges on it would block any PR that
# touches tests/e2e/** — including the ones fixing the suite. Until it is
# green the job runs and reports for E2E-path PRs without blocking. To flip
# it on: add `e2e` to needs, add E2E to the env below, and require
# `"$E2E" = success || skipped` after the loop — skipped is the normal
# result for a path-filtered job and must keep passing, so it has to be
# checked outside the loop or it would excuse the jobs above.
- name: Require successful checks
env:
CODE_PATHS: ${{ needs.code_paths.result }}
SHOULD_RUN: ${{ needs.code_paths.outputs.should_run }}
STATIC_ANALYSIS: ${{ needs.static_analysis.result }}
STATIC_ANALYSIS_SHOULD_RUN: ${{ needs.code_paths.outputs.static_analysis }}
ROOT_DIRECTORY_GUARD: ${{ needs.root_directory_guard.result }}
TYPECHECK: ${{ needs.typecheck.result }}
TYPECHECK_SHOULD_RUN: ${{ needs.code_paths.outputs.typecheck }}
GIT_COMPATIBILITY: ${{ needs.git_compatibility.result }}
GIT_COMPATIBILITY_SHOULD_RUN: ${{ needs.code_paths.outputs.git_compatibility }}
CODEX_INDEX_HEAL_CONTRACT: ${{ needs.codex_index_heal_contract.result }}
CODEX_INDEX_HEAL_CONTRACT_SHOULD_RUN: ${{ needs.code_paths.outputs.codex_index_heal_contract }}
XTERM_PATCH_SYNC: ${{ needs.xterm_patch_sync.result }}
XTERM_PATCH_SYNC_SHOULD_RUN: ${{ needs.code_paths.outputs.xterm_patch_sync }}
SHELL_CONTRACTS: ${{ needs.shell_contracts.result }}
SHELL_CONTRACTS_SHOULD_RUN: ${{ needs.code_paths.outputs.shell_contracts }}
TEST: ${{ needs.test.result }}
TEST_SHOULD_RUN: ${{ needs.code_paths.outputs.test }}
ORCAD_BROWSER: ${{ needs.orcad_browser.result }}
ORCAD_BROWSER_SHOULD_RUN: ${{ needs.code_paths.outputs.orcad_browser }}
MOBILE_WEB_APP: ${{ needs.mobile_web_app.result }}
MOBILE_WEB_APP_SHOULD_RUN: ${{ needs.code_paths.outputs.mobile_web_app }}
CROSS_VERSION_WIRE: ${{ needs.cross-version-wire.result }}
CROSS_VERSION_WIRE_SHOULD_RUN: ${{ needs.code_paths.outputs.cross-version-wire }}
MANAGED_HOOK_NODE18: ${{ needs.managed_hook_node18.result }}
MANAGED_HOOK_NODE18_SHOULD_RUN: ${{ needs.code_paths.outputs.managed_hook_node18 }}
PACKAGE: ${{ needs.package.result }}
PACKAGE_SHOULD_RUN: ${{ needs.code_paths.outputs.package }}
PACKAGE_WINDOWS: ${{ needs.package_windows.result }}
PACKAGE_WINDOWS_SHOULD_RUN: ${{ needs.code_paths.outputs.package_windows }}
run: |
if [ "$CODE_PATHS" != "success" ]; then
exit 1
fi
if [ "$ROOT_DIRECTORY_GUARD" != "success" ]; then
exit 1
fi
if [ "$SHOULD_RUN" != "true" ]; then
echo "Docs-only change; expensive PR checks skipped."
fi
failed=0
check_job() {
local name="$1" result="$2" should="$3"
if [ "$should" = "true" ]; then
if [ "$result" != "success" ]; then
echo "$name: expected success, got $result"
failed=1
fi
else
if [ "$result" != "skipped" ]; then
echo "$name: expected skipped, got $result"
failed=1
fi
fi
}
# Require success when the PR has code-relevant changes
check_job static_analysis "$STATIC_ANALYSIS" "$STATIC_ANALYSIS_SHOULD_RUN"
check_job typecheck "$TYPECHECK" "$TYPECHECK_SHOULD_RUN"
check_job git_compatibility "$GIT_COMPATIBILITY" "$GIT_COMPATIBILITY_SHOULD_RUN"
check_job codex_index_heal_contract "$CODEX_INDEX_HEAL_CONTRACT" "$CODEX_INDEX_HEAL_CONTRACT_SHOULD_RUN"
check_job xterm_patch_sync "$XTERM_PATCH_SYNC" "$XTERM_PATCH_SYNC_SHOULD_RUN"
check_job shell_contracts "$SHELL_CONTRACTS" "$SHELL_CONTRACTS_SHOULD_RUN"
check_job test "$TEST" "$TEST_SHOULD_RUN"
check_job orcad_browser "$ORCAD_BROWSER" "$ORCAD_BROWSER_SHOULD_RUN"
check_job mobile_web_app "$MOBILE_WEB_APP" "$MOBILE_WEB_APP_SHOULD_RUN"
check_job cross-version-wire "$CROSS_VERSION_WIRE" "$CROSS_VERSION_WIRE_SHOULD_RUN"
check_job managed_hook_node18 "$MANAGED_HOOK_NODE18" "$MANAGED_HOOK_NODE18_SHOULD_RUN"
check_job package "$PACKAGE" "$PACKAGE_SHOULD_RUN"
check_job package_windows "$PACKAGE_WINDOWS" "$PACKAGE_WINDOWS_SHOULD_RUN"
exit "$failed"