Commit Graph
141 Commits
Author SHA1 Message Date
Jinwoo Hong e6aa90ff36 test(mobile): certify the browser pane's golden families and render it in a page (OTA phase C, C6.5) (#21777)
* test(mobile): pin the browser pane's golden families

The half pin for C6: 4 families, 15 goldens, every verdict the one C2's
rule predicts. Measured per family with vitest `-t` over the full
787-golden corpus, with C1's 103 reproduced golden-for-golden as the
control: 6 byte-identical, 9 result-absent-settlement.

No composed `c6-page-closure.ts`: a composed table is pinned against a
route and the browser is a pane, so C7's route is what composes this
with C1's.

The derivation census does not wait for that route. `mobileWebAppRoute-
Closure` becomes one case of `mobileWebAppModuleClosure`, which takes
any entries, so the pane's own closure can be read from the module. Two
cases: the pane alone reaches exactly the pinned four, and the pane
beside `app/h/_layout` adds exactly those four and no other, with the
layout reproducing C1's 22 as the control for the difference.

Closure at this base: 48 local modules alone, 34 beyond the layout, 30
under `src/browser` and four through the web siblings. The design said
23, all under `src/browser`; it was measured before C6.2 and C6.3 added
those siblings, so the pin carries the re-measured number.

`browser.screencast` has no golden at all, so this certifies the input
path and says nothing about the frame path.

Red first: with `browser.wheel` dropped from the table, both census
cases fail naming the missing family; restored, the file's 10 cases pass
and the parity suite reports "15 goldens in 4 families, 6 byte-identical".

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the frame budget against the shell's real frame

Ruling 2's pin. `binaryEventEnvelopeBytes()` sizes the mobile view's
device scale from a skeleton it builds itself, and until now its only
check was another skeleton of the same shape in the same file: two
copies of one assumption agreeing with each other.

This measures the real thing. A frame with CDP's nine metadata fields
and a real `Page.screencastFrame` timestamp, encoded by C6.1's
`encodeBridgeScreencastFrame` and serialized by the real
`BridgeHostSubscriptions`, posted through the host harness: 303 bytes
besides the image, against a bound of 516.

Held above is not enough on its own — 213 bytes of slack is room for the
shell to grow the envelope by a field the page never hears about — so
the bound is reconstructed exactly instead. Every byte of that slack is
a number this frame prints narrower than a double can; adding those back
gives 516 on the nose.

The budget cases run a generated noise image at the budgeted scale, not
a committed fixture: the worst case is the image JPEG compresses least,
and a photograph sits a tenth of the way to it. 901,161 px at 0.545
bytes per pixel is 491,132 bytes, which the shell posts at 654,857 of
the 655,360-byte cap. One envelope more and the shell drops it, which is
ruling 1 read from the budget's side.

Red first, two ways. Drop the metadata widening from the bound and three
cases fail, the sharpest being the real shell answering the frame the
page thought it could send with zero posts. Add a field to the shell's
own envelope and the reconstruction fails at 516 against 548, where the
existing suite stays green on all 14 — which is the drift this file
exists for.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): record the measured frame bytes, correcting e7cd24ef10

The previous commit message says the shell posts 654,857 bytes for a
frame at the budgeted area. That number was not measured; I wrote it
from the budget arithmetic instead of reading it off the harness. The
measured value is 655,147, which is 213 under the cap rather than 503.

Nothing in the assertions changes — they compare against the cap and
the bound, never against a literal — but the figure now lives in the
file where it was measured rather than only in a message that has it
wrong.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): render the browser pane in a page and paint a real frame

The only place C6's whole frame path runs. Every other check reads one
half: the shell suites drive the host with no page, the page suites
drive the hooks with no shell, and the parity pin certifies the input
path from a recording.

Ruling 4: no route is added. `bundleMobileWebApp` already takes an
`appDir`, so this builds a one-route tree of its own and nothing under
`mobile/app` moves. The shell double grows a screencast lane to serve
it: it accepts a subscribe, posts `event.binary`, and prices each frame
the way `BridgeHostSubscriptions` does, so an over-cap frame is dropped
where the page can watch the stream survive it.

Six cases: the pane subscribes with `wantsBinary` and paints the frame
it is handed; a second frame flips the double buffer; an over-cap frame
is dropped and the next one paints on the same subscription; the grant
withheld produces the update-the-app copy and no subscribe at all; a tap
issues one `browser.mouseClick` at the centre of the source viewport;
and no request leaves the bundle's own origin.

Measured. The frame the pane asks this viewport for is 390x698, which as
noise is 201,924 base64 characters. The phone's mobile-mode frame is
780x1424 and encodes to 811,168, which is 124% of the cap and the reason
the area budget exists; the over-cap case uses 2400x2160 at 3,761,580,
574% of it. The tap maps to (194, 356) against a 390x712 source, one
device pixel off centre because the rendered width is 382.33 CSS pixels
for 390 source pixels.

One finding, recorded rather than fixed because it is not the pane's.
The page files a CSP `script-src` violation on every load, on any route:
Zod 4 feature-detects its compiled path with `new Function('')`, the
shell's `script-src 'self'` blocks it, Zod catches the throw and takes
the interpreted path. The page is correct and the report is filed
anyway. One case names it so a second `eval` is visible, and every other
case asserts no violation beyond it.

Red first, twice, both by reverting behaviour C6.2 landed. Stub out the
decode probe in `whenBrowserFrameDisplayable` and the flip case fails on
two identical frame digests. Point `updateBrowserImageSource` at the
host element instead of the surface child and the paint and flip cases
both fail. The first case's comment is corrected by the first of those:
it claimed a visible layer proved the decode-then-flip, and the frame
still paints with the probe gone, so the flip case is what proves it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): sweep the worst-case JPEG cost instead of taking one point

The reviewer is right, and it is worse than the report says. At 0.545
bytes per pixel, 90 of 143 viewports posted a frame over the cap, and
59 of the 111 the budget claims to fit were dropped outright by the
real shell — 390x712 at scale 1.8 among them.

The old number came from one 2400x2160 frame. A single large frame is
the cheapest per pixel in the whole range, so a worst case measured
there is not a worst case anywhere else.

Swept 143 viewports, widths 320 to 1400 and heights 480 to 1600, each
encoded by Chromium at the scale the real budget picks for it. Across
the 111 the budget fits, the cost ranges 0.54470 to 0.55351 bytes per
pixel. The constant is now 0.56: that maximum plus 0.00649, about 1.2%,
for the encoder version it was not swept on. The docstring carries the
sweep, the range, the margin and the date.

0.56 is a fixed point, not a guess. Raising the constant shrinks the
budget, which lowers the scale, which moves the cost; 0.555, 0.56 and
0.565 all leave the same 31 viewports over the cap, and every one of
those sits at the scale floor of 1, where the module already declines
to go blurrier and C6 ruling 1's drop rule is the protection. The new
test asserts both halves: nothing the budget fits goes over, and the
largest viewport it cannot fit is dropped by the real shell.

The sweep lives in `config/scripts` because it needs Chromium: the
frames are CDP screencast frames, so Chromium's encoder is the oracle
and a Node JPEG library would calibrate against the wrong bytes. It
drives the real budget, the real scale function and the real
`BridgeHostSubscriptions`, and runs in about 4 seconds.

Ruling 2's block in `browser-screencast-budget-at-the-shell.test.ts`
now says plainly what it measures. It feeds `noise(area * theConstant)`,
a byte count the constant itself produced, so it can falsify the
expansion and the drop rule but never the constant. It read as if it
validated the worst case, and it did not.

Red first: put 0.545 back and the sweep fails with 59 viewports, each
naming its scale and reporting `null` — the real shell dropping the
frame rather than posting it over the cap.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): turn Zod's JIT probe off for the page, before any module

The page filed a CSP `script-src` violation on every load: Zod decides
whether it may compile by constructing `new Function('')` and reading
the throw as "no JIT here", the shell's `script-src 'self'` is exactly
that throw, and the browser reports it before Zod catches it. Zod's own
source gates the probe on `jitless` for this case.

`z.config({ jitless: true })` at the entry does not work, and the
reviewer's suggestion of putting it there was measured losing the race.
`$ZodObject` reads `allowsEval` when a schema is constructed, not when
one is parsed, so the first module-scope `z.object(...)` in the bundle
fires the probe — and esbuild evaluates the chunk holding zod and its
callers before the chunk holding any module of ours that imports zod. A
Function-constructor trap in the page put the call under `new ZodObject`
ahead of the entry's first statement.

`globalConfig` is `globalThis.__zod_globalConfig`, which zod adopts with
`??=` rather than replacing, so the banner can set the flag before any
module runs. That is where it now lives, beside the `process` shim and
under the same `MOBILE_WEB_APP_SHIMS` contract, which asserts it is
applied. Nothing is lost: the compiled path was never reachable in a
page under this policy.

The render check's `newCsp()` filter is gone. It dropped violations by
`blockedURI === 'eval'`, which would have hidden a real one, and every
case now asserts zero. The first case walks load and first paint, which
is where the second of the two reports fired. The dead `violations`
array is deleted.

Red first: blank the banner constant and four of the six cases fail,
each naming a `blockedUri: 'eval'` the filter used to swallow.

Finding, not fixed here and reported instead: the page bundles two
copies of zod, mobile's 4.4.3 and the repo root's 4.5.4, because
`src/shared/zod-salvage.ts` resolves upward. That is 808 KB of duplicate
source. Aliasing `zod` to one copy in the builder fixes it and was
measured working, but it changes which zod shared code runs in the
shipped page, which is a call to make on its own rather than inside a
CSP fix.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): give the shell double the whole canCarry rule and the acks

The double reproduced one arm of `BridgeHostSubscriptions.canCarry`, the
message cap, and silently carried anything the other two would have
refused: a window already holding its maximum frames, and a window whose
bytes the frame would push past the limit. It also ignored the page's
`ack` frames, so its window never reopened — which was invisible only
because no case streamed far enough to close it.

Both arms are in now, and the `ack` arm consumes the page's acks exactly
as the host does. The three caps are read out of `bridge-caps.ts` and
`bridge-host-subscriptions.ts` rather than retyped, the same way the
harness already reads the protocol version and the CSP, so a double
carrying a stale number is not possible. The render check's own
`640 * 1024` is gone with them.

One case for it: thirty frames of about 200 KB, roughly 6 MB through a
4 MiB window, nothing over the message cap, so a drop can only come from
the window. Every frame posts, nothing is dropped, and the page's ack
seqs are read back to show the window stayed open because the page acked
rather than because the double was generous.

The file docstring said the double answers no RPC. It serves a
screencast stream now, so it says that instead, and says what it still
is not: it decides no domain behaviour.

The dead `violations` array is gone, folded with the CSP commit.

Red first: make the `ack` arm inert, as it was before this commit, and
the case fails with `Set{'posted','dropped'}` against `Set{'posted'}`.
A first attempt at that mutation left the byte subtraction in place and
stayed green, which is the mutation being wrong rather than the case.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci(mobile): run the whole mobile-web-app family, not a list that goes stale

The `mobile_web_app` job hand-listed ten files. The render check this
chain added was not among them, so it would have skipped in CI — and it
was not the first: three landed censuses were already unlisted, and
their closure blocks only run with `ORCA_MOBILE_WEB_APP_DEPS_REQUIRED=1`,
so they are green in the sharded `test` job whether or not they ever
ran here. Nobody could see it.

The list is now two vitest filename filters, `config/scripts/mobile-web-
app-` and the one builder test outside that prefix. Quoted, because
vitest matches a positional as a substring against the discovered files
rather than expanding a glob: `mobile-web-app-*.test.mjs` finds nothing,
and it fails by reporting no test files rather than by running fewer.
Both forms were tried before this one was written.

It runs 18 files and 205 cases, against 10 files before. With mobile
dependencies absent, 111 of those 205 skip, which is the measure of what
only this job runs. Per file, cases CI has never run:

  browser-pane-render            7 of 7   (this chain)
  source-control-external-links  9 of 9
  source-control-keyboard        6 of 6
  source-control-text-inputs     6 of 19  (C4.2)
  frame-budget-sweep             4 of 4   (this chain)
  route-manifest                 2 of 17

Two more files the filter adds run fully in the sharded job already and
change nothing here: `browser-pane-text-inputs` (C6.4 — it censuses a
hand-written closure and never bundles, so unlike the report it was not
skipping) and `external-link-seam`.

The sweep is renamed into the family for the same reason. As
`mobile-browser-frame-budget-sweep.test.ts` it matched neither the job's
filter nor `pr-code-change-scope.mjs`'s `config/scripts/mobile-web-app-`
prefix, so a change to it alone would not have run the job that runs it.
It is also gated on the dependency check now: it needs no
react-native-web, but it launches Chromium, and that flag is what tells
the job with a browser from the one without. Unguarded it would have
failed the sharded `test` job outright.

That filter is a prefix match. `config/scripts/mobile-web-app-` and
`mobile/src/` both fire this job, and `.github/workflows/pr.yml` is in
GLOBAL_FORCE_PREFIXES, so this commit runs everything.

Also: `postedFrame` in the ruling-2 pin and in the sweep both reached
the binary lane through `?.`, so a subscribe that opened no stream read
as zero posts — indistinguishable from a dropped frame, which is the
verdict both files are about. They throw now. The render check's
restated `640 * 1024` went with the window caps in c73b405f81; the cap
is read from `bridge-caps.ts`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci(mobile): record what the mobile-web-app job costs to run

The filter that replaced the hand list runs 18 files where the list ran
10, so the step's cost is now a function of what anyone names into the
family rather than of what a reviewer remembered to add. Measured on
this machine: 25-30s wall for the whole step, of which the frame-budget
sweep is 2.5s.

The sweep is the one part whose cost is a choice. It encodes 111 noise
JPEGs in Chromium, one per viewport the budget fits, so adding rows to
that set is a decision about this job's runtime and the comment says so
where someone would make it.

Found, not fixed, and reported for its own PR rather than folded here:
the page bundles two copies of zod, mobile's 4.4.3 and the repo root's
4.5.4, reached through `src/shared/zod-salvage.ts`, which resolves
upward while `mobile/src/` resolves to mobile's. That is 808 KB of
duplicate source and two module instances in the shipped page. Aliasing
`zod` in `mobileWebAppBuildOptions` fixes it and was measured working
during this chain; it is reverted and stays reverted, because it changes
which zod shared code runs in the page and that is not a call to make
inside a CI commit. The CSP fix in 71254ab3a9 does not depend on it:
`globalConfig` lives on `globalThis`, so the banner covers both copies.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): load the frame-budget sweep's mobile modules after the dependency guard

vite transforms every file under mobile/ against mobile/tsconfig.json, which extends
expo/tsconfig.base.json; the sharded test job installs no mobile dependencies, so the
sweep's static imports failed the file at load before describe.skip ran. Type-only imports
stay static; the values load in beforeAll behind mobileWebAppDependenciesPresent().

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): certify the frame budget at the quality the pane ships

Round 2: the sweep and the render check encoded fixtures at a retyped 0.72; both now read
BROWSER_FRAME_QUALITY (the sweep from the module, the render check through the harness reader),
so a quality change fails the certification instead of leaving it green. Every render case now
asserts zero CSP violations; the sweep pins the 32 viewports left at scale 1; three references to
a renamed file and a file that never existed are corrected; a shim count comment is made
count-agnostic.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): assert the frame-budget sweep against the constant, not the measured maximum

The margin above the measured 0.55351 is what an encoder drift is allowed to spend; pinning the
measurement made a drift inside the margin fail a budget that still held. The number stays in the
docstring as the sweep's record.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): move the sweep's measured-maximum note beside the assertion it explains

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-20 07:43:14 -04:00
Jinwoo Hong ac024d4f05 feat(mobile): serve the files explorer and preview from the page (OTA phase C, C3.1) (#21710)
* refactor(mobile): take the files screens' router from the handoff seam

Inside the shell's page a screen is one document standing in for one screen, so
a target the page does not render has to be handed back to the app that does.
`useRouteHandoff` is where that decision lives, and its web sibling is the only
thing that makes it; both files screens held expo-router's own `useRouter`, so
on the web the explorer's Back and the preview's Back would post nothing and a
target outside the page would paint Unmatched over the page it is on.

Natively this is the same object — `route-handoff.ts` is `useRouter()` — so no
behaviour moves here, and `back()` stays expo-router's until the navigate-back
verb lands and the seam starts wrapping it.

A census rather than a behaviour test: neither screen's own tests can see the
difference, because a push that is never handed off still works for a target
inside the page. It walks this directory, refuses a value import of
expo-router, and names the two screens that must hold a router so a walk that
found nothing fails instead of passing empty.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): let the shell stand in for the two files routes

Both route files take the index.tsx shape — flag, MobileWebShellScreen, native
screen as fallback — and both gain the `.web.tsx` sibling that shape forces.

Inert until the manifest lists these routes: the shell answers `native-route`
for a route the bundle does not name, which is what `fallback` renders, and the
flag is `__DEV__`-only besides. Listing them waits on C2.3 and C2.5.

The sibling is not a precaution. The manifest defers every route behind
`import()`, so a native-only route module is invisible until the page opens
that route; the render check now opens both and, without the siblings, painted
`expo-modules-core.requireNativeViewManager is not available on web` instead of
the screen. That is also why the two cases render the route rather than
asserting a file exists.

The file path never becomes a path segment: only `hostId` and `worktreeId` are
spelled into the pathname, encoded, and everything else — `relativePath`,
`absolutePath`, `cwd`, `pathText` — is a param, which is how a `/`, a space or a
`..` stays out of the segment vocabulary the bridge holds a route to. The
preview render case proves the round trip on `docs/my notes/readme.md`.

`mobileFilePreviewShellParams` drops a param the normalizer left `undefined`
rather than sending it empty, because the page reads these back through
useLocalSearchParams where `line: ''` and no `line` are different screens. Its
test drives the normalizer rather than a hand-written literal: the literal omits
the key entirely, so it held with the filter removed.

The preview case also records what React Native Web says out loud — BackHandler
is inert on web, so Android back inside the page skips the unsaved-draft
prompt. Named in the assertion rather than filtered out, so closing it is a
change to that line.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): ask about an unsaved draft in the screen, not through Alert

React Native Web's `Alert` is `static alert() {}`. Inside the shell's page that
made Back with an unsaved terminal-artifact draft a button that did nothing at
all: no prompt, because the dialog is a no-op, and no navigation either, because
the code took the branch that shows one. Silently, with nothing on the console.

The prompt is now a row under the header. Not `ConfirmModal`, which every other
confirm here uses: that is a `BottomDrawer`, and C1.9 has Reanimated's animated
styles never reaching the DOM node on WKWebView, so on iOS in the page the
drawer parks off-screen and Back would be dead a second way. This paints the
same on every platform with no animation behind it.

Hardware back is registered natively only. React Native Web's
`BackHandler.addEventListener` logs "BackHandler is not supported on web and
should not be used." and hands back an inert subscription, so the guard never
armed there regardless; the render check asserted that console error on main and
now asserts none. The degradation is real and stated rather than hidden: Android
back inside the page pops the native stack without asking, and the page's own
Back control is where the question lives.

The decision moved to a hook so it is testable without a screen: the prompt also
drops itself when the draft it was about is saved or reverted, which is a state
`Alert` had no way to be in.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep expo-haptics' DOM shim out of the page

expo-haptics has a web build, and with no `navigator.vibrate` — iOS Safari,
which is the WebView the page runs in — it fakes a haptic by appending a hidden
`<label><input type="checkbox" switch>` to `document.head`, clicking it, and
removing it, once per call. C1.9 traced a long press that never fired on the
worktree list to exactly that stray click, and the file explorer calls
`triggerSelection` on every row tap, so C3 is the first domain to fire it per
tap rather than per long press.

`haptics.web.ts` answers the same five names with nothing. A phone holding the
page is a phone whose native app is right there with the real haptics, and a
missing tap feedback is worth less than a tap that does not register.

The test reads the shipped bytes rather than the import, because that is the
claim: with the override removed the bundle carries `ariaHidden` and
`pointer: coarse`; with it, neither, nor the `setAttribute("switch"` that does
the clicking. Not `navigator.vibrate` — react-native-web's own Vibration export
calls that and touches no DOM until something invokes it, which cost this test
one wrong red before it was narrowed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the files routes native when the page could not be given one

A file path is a param, so `/`, spaces and `..` all cross safely — but
`BRIDGE_MAX_ROUTE_PARAM_CHARS` is 1024 and a Windows long path is not bounded by
anything the user cannot exceed.

The symptom is not the blank document the design predicted, and the correction
matters: `bridge-host.ts` already parses the route against the page's own schema
and drops it to `null` when it fails, so `init` arrives naming no screen and the
page paints "Update Orca to open this workspace" — a wrong message about a fine
app, over a native screen that works. Deciding before the switch instead leaves
the route native, which is where every route starts.

The schema is the predicate rather than a copy of its bounds, so the rule cannot
drift from the half that matters, which is the half the page reads. The same
call also refuses a `worktreeId` the segment rule will not route: `..` survives
`encodeURIComponent`, which is the C1.8 class.

The tests assert the schema really refuses each input before asserting the guard
does, so neither case can pass by being impossible.

This belongs in the shell beside the schema; it is in the files domain while the
contract files are the C2 lane's.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin what keeps a file path out of the route vocabulary

Seven shapes, one case each rather than a representative: a plain path, a space,
a dot segment, an already-encoded slash, a fragment, non-ASCII, and an absolute
path. Each is checked in the two directions a path travels — the href the shell
writes into the page's history, and the href the page would hand back — for both
the pattern accepting it and the path coming back out of the query unchanged.

The counterfactual is in the file: the same paths spelled as a segment are
refused. Without that, the cases above would hold for a rule that was never
doing any work. Mutating `stringifyRouteHref` to join its query by hand instead
of through `URLSearchParams` fails three of them.

Also fixes two new test files the tests-typecheck ratchet caught: the partial
`react-native` mock needs a typed `addEventListener`, `act` will not take a
callback that returns a value, and `findAllByType('Pressable')` does not
typecheck against `ElementType` — the neighbouring files that do it are
grandfathered, so the tag comparison goes through a helper instead.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): derive the discard prompt instead of clearing it in an effect

Both changed-code gate findings, which the lane had not run until the last
commit. React Doctor is right: the effect that cleared the prompt when the draft
went away adjusted state after a prop changed, so a save landing while the
prompt was up painted one frame still offering to discard nothing. The prompt is
now `asking && hasUnsavedDraft`, which cannot be stale by construction, and the
test that covers it passes unchanged.

The hoisted mock's `as` on a string literal is gone too: the literal narrows on
its own and the tests reassign it, so the holder is annotated instead.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): add the files routes to the hybrid shell flag census

The census pins every file that reads `useMobileWebShellEnabled`, because a
reader nobody listed is how a dark feature stops being dark. C3's two routes are
deliberate entries: each has a native screen behind it as `fallback`, and each
is inert until the manifest lists the route.

Found by the full mobile suite rather than by the files subset this lane had
been running per commit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): serve the files explorer and preview from the page

The last C3 commit: both routes join MOBILE_WEB_PAGE_ROUTES, and the shell
starts rendering the page for them on a phone with the dev flag on.

Grants are not the same for the two, and the difference is the point. Both take
`navigate` (Back pops the native stack, and the explorer's rows open the preview
beside it) and `storage` (the shared components the host layout renders above
them). Only the preview takes `externalLink`: a Markdown preview renders links
and `MobileMarkdown` opens them through the platform seam.

The explorer does not, and measuring is what says so rather than reading. Every
page route reaches `external-link.web.ts` — `/h/[hostId]` and agent-history
included, both granted nothing for it — because the protocol wall in the shared
host layout imports it. So closure membership is not the oracle for a grant; the
question is whether the route's own screens call it, and only the preview's do.
`MobileMarkdown` is in the preview closure and absent from the explorer's, which
the census now asserts in both directions.

Neither route writes a clipboard, so neither takes `native.clipboard.write`;
the census pins that as the absence of both `ExpoClipboard.web.js` and the
clipboard seam, with the tasks closure as the control that the probe can see one
when there is one.

The seam predicate moved into a module both censuses import rather than being
restated per series: two spellings of one rule drift, and this one is a regex.

Red-first: both manifest assertions failed on the new entries before they were
updated, and routing `MobileMarkdown` around the seam fails the preview's census
while leaving the explorer's passing, which is the asymmetry the grants encode.

Closure sizes as the page ships them, extensionless so the `.web.tsx` is what is
measured: explorer 3439 modules / 302 local / 10 under src/files, preview 3667 /
331 / 20.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): read the files route's ids as one value and key the shell on them

Two round-1 findings, both reproduced before the fix.

A repeated query key reaches `useLocalSearchParams` as an array, and the
explorer read `hostId` and `worktreeId` bare. `String(['a','b'])` is `a,b`, so
the template built `/h/host-a%2Chost-b/files/wt-1%2Cwt-2` — a single segment the
bridge's rule accepts, and the shell would open a page for a host nobody has.
Read through `firstParam` now, as the tasks and agent-history switches do. The
preview already went through `singleParam` and is unchanged.

Neither switch keyed `MobileWebShellScreen`, where `index.tsx`, `tasks.tsx` and
agent-history all do. A host captures the grants its session opened with, so a
screen reused across a route change keeps authorising frames under the grants of
the route the page has left; only a remount drops that bridge. Both are keyed on
the route pathname now, with agent-history's reason.

The new route test is the agent-history one's shape. It caught both: the array
case landed on no route at all, because `name` was an array too and the schema
refuses a non-string param value, and the two lifecycle cases saw a prop update
where a remount was owed. It also needs agent-history's `lucide-react-native`
mock, since `firstParam` lives in the source-control barrel.

`name` is now omitted when empty rather than sent as `name=`, matching the two
switches beside it: an absent label lets the panel derive its own.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): confirm a discarded draft with the app's own modal

Round-1 findings 3, 4, 5 and the minor one.

**ConfirmModal, not the bespoke row.** The row existed because C1.9 had
Reanimated's animated styles never reaching the DOM node on WKWebView, which
left every BottomDrawer parked off-screen. C1.10 (`b7c06900e2`, an ancestor of
this branch) fixed that with a dependency array on the mapper hooks, and the
drawer render check now holds it on WebKit as well as Chromium. With the reason
gone the row does not stand on its other merits: `Alert.alert` was modal on
native before the page existed, and the row quietly changed that for phones
too, so the app's own confirm is both the idiom and the closer behaviour.
`MobileFilePreviewDiscardPrompt`, its test and its thirty style keys are gone;
the hook's state machine and its tests are unchanged.

**The encoding test claimed more than it pinned.** Hand-joining the query reds
only three of the seven shapes; `docs/readme.md`, `../etc/passwd`,
`docs/日本語.md` and `/logs/run.txt` are encoding-neutral in the query, whose
pattern half is `[^#\s]*` and admits a slash, a dot segment and non-ASCII
verbatim. Rather than narrow the claim in a comment, the split is now pinned by
behaviour: each neutral shape must survive the query unencoded, each
load-bearing one must not. Moving `docs/readme.md` between the lists fails it.

**The manifest comment named one shared-layout opener and there are two.** The
New Workspace source field, which the sidebar renders on a wide layout, opens a
URL through the seam as well. Both are the shared layout's and every `/h` route
reaches both, `/h/[hostId]` included with no `externalLink`, so the tablet tap
is dead on all of them — recorded here as pre-existing rather than fixed, since
the grants do not move.

**Minor:** the dot-segment case in the guard test now asserts the schema refuses
the route before asserting the guard returns null, as the length case does.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): stop every page drawer logging a BackHandler error when it opens

Round-2 findings.

**The registration belongs to the drawer, and that is where the guard went.**
`mounted-bottom-drawer.tsx` armed `hardwareBackPress` whenever a drawer was
visible and interactive, with no platform check, so the hook's claim to have
dropped that console line held only while its prompt was closed — and every page
drawer since C1 has logged it on open. Platform-gated at the drawer now; the
hook's comment says so rather than claiming the credit.

**Nothing had ever opened a modal in a browser.** The render check next door
mounts both files routes and reads what they paint but taps nothing, so
`ConfirmModal` inside the page — a BottomDrawer, so Reanimated, a portal and a
gesture handler — was unproved. A new render file loads an editable terminal
artifact through the harness's scripted reply, edits it, taps the page's Back,
and asserts the prompt's title is up and no BackHandler line is on the console.
Red first on exactly that line; the prompt itself painted, which is also the
first proof on a browser that C1.10's fix carries a real drawer in the page. A
second case answers Stay and checks the draft survives. Its own file rather than
the render check's, which is at 482 of the 600-line cap; registered in pr.yml.

**The encoding rule was stated wrong.** Two rules decide it and neither is about
paths: the pattern's query half refuses whitespace and `#`, and
`URLSearchParams` is form-urlencoded, so it reinterprets `&`, `+` and a valid
`%XX`. `a+b.ts` reads back `a b.ts` and `a&b.ts` reads back `a`, so both are
load-bearing; `a=b.ts` and `a%b.ts` are not, because only the first `=` splits
the pair and a lone `%` begins no escape. A newline joins the load-bearing list
as the refused shape rather than the altered one.

**The web sibling read its params bare** where the native one uses `firstParam`.
Not reachable — the page only arrives through `init.route`, whose params are
already `Record<string, string>` — but the two files are meant to be one screen.

The preview keys on the pathname alone, and the comment now says why that is
enough: every caller in this tree pushes.

Closures after this: explorer 3441 / 304 / 10, preview 3666 / 330 / 19. The
explorer grew two modules because its web sibling now reaches `firstParam`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): give the explorer the grants the preview needs, and key on the route

Bot findings, one of them a real gap.

**Pullfrog is right, and my grant oracle was half a rule.** Grants resolve once,
from the route the shell opened: `grantsForRoute` reads `session.routePathname`
and `init.grants.native` carries the answer for that session. The explorer's
rows push to the preview, and because the preview is a page route that push
stays inside the same document — no second `init`. So a preview opened that way
runs under the explorer's grants, and a Markdown link in it was refused by
`notifyExternalLink` with nothing on screen to say why. "Does the route's own
screen call it" was right for a route's own screens and wrong for the routes it
reaches in-page, so the explorer now declares `externalLink` as a transitive
grant, with the comment saying that rather than claiming it opens links. The
census pins the pair as a superset; removing the grant reds it.

**The seam regexes matched one quote style.** A double-quoted `react-native`
specifier walked past both censuses unseen. Both styles now, with the predicate
tested directly for the first time.

**The discard request outlived its draft.** `asking` stayed set after a save or
a revert, so the next edit re-showed the prompt with no Back request behind it.
The request is now dropped when the draft it was about goes, adjusted during
render rather than in an effect — the shape React Doctor named in the round-1
fold. Red first: save with the prompt up, edit again, prompt is back.

**CodeRabbit's keying comment is a correctness point, not the question I
answered.** The page learns its route exactly once, out of `init`, so a
same-path param change — another file in the same worktree — left the shell
mounted and the page still showing the file it was opened on. My comment claimed
"the screen reloads the preview from the param either way", which is true only
with the shell absent. Both switches key on the whole route now, params
included; two tests cover the same-path case and both red on a pathname-only
key.

`build-mobile-web-app-bundle.test.mjs` hit 601 of its 600-line cap on the way,
so the two manifest assertions now share one expected list instead of repeating
it. Closures unchanged: explorer 3441 / 304 / 10, preview 3666 / 330 / 19.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make the seam test import the module it is testing

Round 3.

**The blocker is mine and the reviewer's diagnosis is exact.** The seam
predicate test imported an absolute path into this lane's worktree. On CI that
module does not exist and it takes the whole `config/scripts` suite down; here
it resolved to the same file by accident, so the test was green against a tree
rather than against the checkout — which is why reverting the double-quote fix
left it passing and the predicate untested. Relative now, and proved: reverting
the fix in place reds both double-quoted cases, which is the first time this
test has failed for the right reason. Every file this PR touches is grepped for
`/Users/` and `orca-lanes`; none carries a path.

**Three comments outlived the grant change.** The two lists became equal when
the explorer took `externalLink`, so "longer than the explorer's" and "declared
with different grants" were both false. Corrected to what is actually true: the
lists are equal and the reasons are not — the preview has its own consumer in
`MobileMarkdown`, the explorer has none and declares the grant because its rows
push to the preview in-page.

**The duplicated serializer is pinned rather than imported.** `shellRouteHref`
lives in `page-bootstrap.ts` beside the page's RPC client and its document
channel, so a native route file importing it would pull both into the app. The
copy stays, and a test asserts the two agree on three routes; dropping the
empty-search branch reds it.

**Recorded, not fixed:** the sidebar `HostScreen` pushes to `/h/<id>/tasks`
through the handoff, which is local, so on a tablet the tasks page runs without
`native.clipboard.write` from any page route and its copy actions refuse
silently. Pre-existing since C2.1 for the worktree list and agent history. Named
in the explorer's manifest comment as the known remaining hop, with the fix
being a handoff rule in its own PR.

The equality pin needed `it.each<BridgeInitRoute>`: the inferred table is a
union whose members carry `?: undefined`, which the ratchet caught.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-19 15:47:25 -04:00
Jinwoo Hong 76063e7ab1 test(mobile): certify the tasks page closure, 70 families and 266 goldens (OTA phase C, C2.6) (#21712)
* test(mobile): read a page closure's run totals through one reader

The C5 gate counted the run's classes inline. C2 needs the same count over its
own closure, and two spellings of "what the run tallied" can disagree while both
stay green, so the loop moves next to `pageClosureTotals` where the table-side
count already lives.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): certify the tasks page closure, 70 families and 266 goldens

C2 moves the tasks screen to the web, so the goldens recorded at a call site
inside `app/h/[hostId]/tasks.web.tsx` and `app/h/_layout.tsx` are the ones whose
divergence would be this domain's. Each is pinned by id: the suite's own counts
run over 787, where one of the other 521 can pay for a closure golden that
stopped replaying.

C1's 22 families are inherited verbatim rather than re-derived — C2's rule
disagrees with them on 10 of the 103 — and the rule decides only the 48 this
domain adds. The pin is split at the domain's seam, one work item opened versus
choosing which to open, because the table is 409 lines of data and `max-lines`
is not a thing to disable.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): correct the C2 pin's inheritance count and census scope

Two comments overstated what was measured. The rule disagrees with 13 of C1's
103 inherited pins, not 10 — the 10 was copied from C5's file, which carries the
same error over the same 22 families — and the breakdown is now named so the
number can be re-derived rather than trusted.

The census reads the committed table and does not re-derive the closure, so a
golden arriving in a pinned family is caught while a new family entering the
closure is not. That was true and unsaid, which is the worse of the two.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test: derive the tasks page closure's family set instead of trusting the pins

Round 2 folds, three.

The pin tables walk the families they already hold, so a scenario recorded at a
call site the route already imports lands in a family nobody pinned and every
assertion stays green. `mobileWebAppRouteClosure` runs in a quarter second and
`config/scripts` already imports it, so the derivation is now a test: the family
set the closure reaches must equal the union of the three committed tables.

C2's inheritance check read the object its own table spreads, which cannot
disagree with itself; it now reads C1's file as text. What that does and does
not hold is written down, because a verdict edited inside `c1-page-closure.ts`
is green there either way — C2 inherits whatever C1 commits. The gate's C1 block
gains the run-totals assertion C5 and C2 already had, which is the check that
edit does fail.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci: run the page-closure family check in the job that installs mobile deps

Its closure half asks `mobileWebAppDependenciesPresent()` first, so outside the
`mobile_web_app` job it skips itself and the precondition it exists to be never
runs. That job sets `ORCA_MOBILE_WEB_APP_DEPS_REQUIRED`, which turns the same
question into a failure.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-19 15:17:29 -04:00
Jinwoo Hong b6e8b1a7b2 feat(mobile): serve the tasks screen from the page, with its seams (OTA phase C, C2.1 + C2.5) (#21694)
* fix(mobile): encode the host id in the tasks workspace-creation href (OTA phase C, C2.1)

`use-mobile-tasks-workspace-create-actions.tsx` built
`/h/${hostId}/session/...` with the host id interpolated raw — the C1.2 class.
A host id carrying `/`, `#`, `?` or whitespace reaches the wire as an href
`BRIDGE_ROUTE_HREF_PATTERN` refuses, the handoff falls through to the local
router, and expo-router's Unmatched paints over the page.

Deleted rather than patched: `hostNewWorktreeSessionRoute` already builds
this exact href with both segments encoded, and already has the test that
pins it. The screen now calls it.

The census that caught it stays: no module under `src/tasks` may interpolate
into `/h/${...}` without encoding, which is the rule rather than this one
line. Three refactor-parity hashes move with the statement change and are
recorded in that file the way every earlier movement is.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): route the tasks tree's external links through the seam (OTA phase C, C2.1)

Ten of the twelve call sites in the tasks page closure: the nine under
`src/tasks`, swapped by one export in the dependency barrel, and
`MobileMarkdown.tsx`, which imports react-native directly and is edited in
place.

Inside the shell's WebView react-native-web's `openURL` calls
`window.open(url, '_blank')`, which both shells refuse — iOS returns nil from
`createWebViewWith`, Android false from `onCreateWindow` — and resolves
regardless. Every one of these sites would have reported success into a tap
that opened nothing.

The barrel's `Linking` is typed `{ openURL: (url: string) => void }`, so a
`.catch` on it is a compile error rather than a handler for a rejection that
cannot arrive; the seam names its own failures. `MobileMarkdown`'s own
`.catch(() => {})` goes with the swap for the same reason.

No parity hash moved: the barrel and `MobileMarkdown` are outside the
refactor-parity family's source set.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): route the shared screens' external links through the seam, with a census (OTA phase C, C2.1)

The last two of the twelve call sites in the tasks page closure:
`ProtocolBlockScreen.tsx` and the `openExternalUrl` prop wiring at
`host-screen-overlays.tsx`.

Both are shared with native routes and with the already-live `/h/[hostId]`
page, so this changes that page too: its external links go from the measured
`window.open` no-op — which both shells refuse and which resolves anyway — to
a URL handed to the shell. Nothing changes on a phone, where the seam is
`Linking.openURL` unchanged.

The `openExternalUrl` prop chain is retyped `(url: string) => void` with it,
and `SmartWorkspaceSourceField`'s `.catch(() => {})` goes: the seam names its
own failures and never rejects, so that was a handler for a rejection that
cannot arrive.

The census is the rule rather than today's twelve sites: no module in the
tasks page closure may reach react-native's `Linking`, by name or through a
namespace import. It reads the closure from a new builder export —
`metafile.inputs` for `_layout` plus the route, which is one definition of
what a page contains — and checks which module the name comes from, not which
text a call site writes, since the tasks tree still calls `Linking.openURL`
and that `Linking` is now the barrel's seam-backed export. Confirmed to
discriminate: restoring one react-native import turns it red.

A second case pins that the seam is in the closure, so an empty offender list
cannot also mean a page that reaches no link code at all.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): write the tasks clipboard through the shell's verb (OTA phase C, C2.1)

The two `Clipboard.setStringAsync` sites in the tasks page closure move onto
a seam, `src/platform/clipboard.ts` with a `.web.ts` sibling, registered in
the overrides.

A hook rather than a function because the web form needs the page's bridge
client, which is React context. Native is `expo-clipboard` unchanged. Web
calls `native.clipboard.write` through `useNativeVerbs`, because
`expo-clipboard` on the web is `navigator.clipboard` and needs a secure
context: the iOS shell serves the page from a custom scheme and Android from
`https`, so that path would work on one platform and silently not on the
other, with nothing at the call site able to tell.

Both seams reject rather than return false, and both call sites already wrap
the write in a `catch` that puts the message on screen — so a write that did
not land says so instead of showing "Copied". A route that has not declared
`native.clipboard.write` is refused before a frame is sent and lands in that
same `catch`; the route declares it in the entry commit.

Two parity hashes move, the hook list and the statement hash, each by one
entry, and are recorded in that file. `semantics` holds, as do render and
style: no RPC call, method literal or JSX host signature changed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): hand the tasks Back button to the shell (OTA phase C, C2.1)

The tasks header's `router.back()` reached expo-router through the dependency
barrel, and inside the page that moves nothing: the document holds the single
history entry the entry wrote with `replaceState`. The stack with somewhere
to go is the native one the shell pushed the page onto.

One line in the barrel, as with `Linking`: `useRouteHandoff` is router-shaped,
so every call site is unchanged. On a phone it is expo-router. Inside the page
it keeps a route the page renders and posts `navigate-back` for a Back the
document cannot serve — the C2.2 seam, which until now had no consumer.

No parity hash moved: the barrel is outside the refactor-parity source set,
and no call site changed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): render mermaid as its own source box on the web (OTA phase C, C2.5)

`MermaidDiagram` is in the tasks page closure, reached through
`MobileMarkdown`, and it renders the diagram inside a sandboxed `WebView`.
`react-native-webview` is a native component with no browser counterpart:
importing it runs a codegen lookup that throws, and the route manifest imports
every route, so one such import takes the whole page down rather than one
diagram.

The web sibling renders the labelled source box the native component already
falls back to on a parse or render error, with that component's own styles, so
the degradation looks like a state the product already has rather than a
second design.

Not a browser renderer, and the reason is not reach: mermaid is a browser
library and the engine bundle is vendored. It is that the native path's safety
comes from the WebView it runs in — `buildHtml` escapes `</script>` and the
U+2028/U+2029 separators because diagram source is untrusted agent and PR
content — and a DOM path has no such sandbox, so it needs its own escaping and
its own proof. That is a change of its own, not a smaller version of this one.

Registered in the overrides, whose gate fails on an unlisted `.web.*` file.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): turn the tasks route on for the page (OTA phase C, C2.1)

The entry: `/h/[hostId]/tasks` joins `MOBILE_WEB_PAGE_ROUTES`, the route file
becomes the shell's flag switch in `index.tsx`'s shape, and a `.web.tsx`
sibling renders the screen directly, registered in the overrides.

The screen moves to `src/tasks/MobileTasksScreen.tsx` first, verbatim — body
byte-identical, imports rewritten to `./`. It has to: under the builder's
`resolveExtensions` a web sibling importing `./tasks` resolves back to
itself, which is why every other shell route's screen already lives in `src`.

The parity family follows the file rather than the path. `TASKS_ROUTE` leaves
`MOBILE_TASKS_SOURCE_FILES` — `SOURCE_PATTERN` already matches
`MobileTasks*.tsx`, so listing it too would double-count — and the execution
reader points at the new file. Measured rather than predicted: all six
refactor-parity cases pass unchanged. No hash moved, including the family
text and declaration list, because the new name sorts where the route path
sat.

The route declares `navigate`, `storage`, `externalLink` and
`native.clipboard.write`, which the grammar fold made expressible and
per-route scoping makes meaningful: it is granted those and not the rest of
what this shell implements.

The browser check covers what only a browser answers — every module in the
closure evaluating under React Native Web, `taskSource` surviving the
handshake into the page's own URL, and the route's chunk arriving on a
client-side navigation. It states plainly what it does not cover: the three
seams are reached from controls that need provider data the double does not
serve, so a case posting those frames directly would prove the transport and
read as a tap it never performed. Both new checks join the `mobile_web_app`
job.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(config): resolve a route closure the way the bundle ships it (OTA phase C, C2.1)

`mobileWebAppRouteClosure` took the route's explicit `.tsx` path as an entry
point, so esbuild used that file directly and `resolveExtensions` never ran.
For a route with a `.web.tsx` sibling that measured the native switch, which
no browser loads: the tasks closure came back carrying
`MobileWebShellScreen`, and with it a `Linking` import the census then
reported as an offender.

Extensionless now, so the closure is the one the page actually contains:
3775 modules, 428 local, with `external-link.web.ts` and `clipboard.web.ts`
in it and the shell screen out.

The route-manifest pins move with the tasks route joining
`MOBILE_WEB_PAGE_ROUTES`, in both the declaration check and the built
manifest.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): cover the clipboard seam, close two page escapes, share the mermaid props (OTA phase C, C2.1)

Four from round 1.

The clipboard seam shipped untested. Both halves have one now: the native
form rejects when `setStringAsync` answers false and resolves when it does
not, and the web form is driven through the real port pair — resolving on a
reply, rejecting when the shell says the pasteboard refused, and rejecting on
an ungranted route without putting a frame on the wire.

The tasks barrel still re-exported `expo-clipboard` with no consumer, which
kept `ExpoClipboard.web.js` — the `navigator.clipboard` path this series
exists to avoid — inside the page closure. Deleted, and asserted as the
module's absence from that closure rather than as a count of importers: a new
import puts the file back whoever writes it.

`ProtocolBlockScreen` reached expo-router's singleton for its way out to the
host list. A singleton is the one shape the handoff cannot intercept — it is
not a hook, so the page's bridge client is never consulted — and `/` is a
route the page does not carry, so inside the shell that replace rendered the
root route in the WebView instead of leaving it. Pre-existing and live via
`/h/[hostId]`; routed through the handoff now. Two suites' `expo-router`
mocks gain the hook the handoff reads.

`MermaidDiagram.web.tsx` redeclared its props; it imports the native
component's type, so drift fails tsc.

No parity hash moved: none of these files is in the refactor-parity source
set.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(config): use endsWith for the clipboard module check

The changed-code gate refuses a dollar-anchored regex where `String#endsWith`
says the same thing. No behaviour change.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): close the href census gap, read route params through firstParam (OTA phase C, C2.1)

Five from round 2, two of them real.

The raw-interpolation census inspected only the leading `${...}`, so
`` `/h/${encodeURIComponent(hostId)}/session/${worktreeId}` `` passed it — and
a worktree id carrying `/`, `#`, `?` or whitespace breaks the href exactly as
a host id does. It now refuses any hand-built `/h/...` template with any
interpolation left raw, whichever segment it is. Proved against exactly that
shape in a throwaway before the change, which the old rule admitted.

The tasks switch read `hostId` and `taskSource` as plain strings. expo-router
hands back an array for a repeated query key, so a duplicate `?hostId=` built
`/h/host-a%2Chost-b/tasks`; both go through `firstParam` now, as the
agent-history switch does. `index.tsx` is untouched, per the Phase D list.

Three in the render check's prose: the header claimed the browser proves the
three seams fire from a tap, which the file's own closing note denies; a
module count repeated a number the closure test already pins; and a `replies`
parameter was threaded through without ever being supplied.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-19 13:40:26 -04:00
Jinwoo Hong 211821dc17 feat(mobile): render agent session history from the desktop's bundle (OTA phase C, C5.1) (#21596)
* fix(mobile): refuse a page target the shell will not take instead of opening it here

`useRouteHandoff`'s web sibling answered two things — handed off, or push it
locally — and fell through to the local router for three different reasons. Only
one of them is a page route. An href the protocol's own pattern drops and a shell
that answered no are the page reaching past what this shell can serve, and the
bundle carries every route under `app/h`, so the fallback does not paint
Unmatched: it mounts `session/[worktreeId]` on React Native Web inside the shell.

The outcome is now tri-state. A target outside `pageRoutes` is never pushed
locally; the page stays where it is and names the reason once per client, which
is the bound the other page-side reporters take.

Proved in the render check against the real bundle: with the double granting no
`navigate`, "Back to hosts" left the host route for `/` and painted Unmatched
before this, and now stays put, posts nothing and reports no page fault.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): encode the host id the worktree row's navigation actions build

Both targets this sheet offers interpolated `hostId` raw — the C1.2 class the
C1.8 stack fixed at the route files and `web.tsx`, at the last two sites that
still had it. `useLocalSearchParams` answers the decoded value, so a deep-linked
id carrying `?`, `#` or whitespace stops being one segment.

It matters more from C5.1 on. Inside the page these targets go through
`useRouteHandoff`, which matches the pathname against the shell's `pageRoutes`
before deciding anything, and the id is the segment the pattern is reading.

The worktree id was already encoded at both sites; this makes the host id match,
and the new test pins all four targets rather than only the one that moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): render agent session history from the desktop's bundle (OTA phase C, C5.1)

The second page route. `agent-history/[worktreeId].tsx` already shipped in the
bundle with its own chunk, so listing it adds nothing to the download and moves
no route count: the switch is `index.tsx`'s, and the shell still decides, because
a bundle naming a grant this app lacks renders the native panel instead.

Its `.web.tsx` sibling is required for `index.web.tsx`'s reason — the native file
reaches OrcaMobileWebShellView, whose requireNativeViewManager runs at import and
takes the whole bundle down in a browser, since the manifest imports every route.

First route with two dynamic segments, so both are encoded. Grants are `navigate`
and `storage`: a resumed session opens the native session screen, the worktree
list now reaches this screen without leaving the page, and `app/h/_layout.tsx`
reads the app's own sidebar width above every page route.

The panel's router becomes `useRouteHandoff`, which is the seam that tells those
two apart: agent history is a page route and is pushed here, the session screen is
not and goes to the shell.

The three writes a resume makes needed no page-side handling and have none. What
they needed was a test that the descriptor's handling survives the extra hop, so
each is run through the bridge and against the same fake directly and the two
verdicts compared: a refused create raises the host's message, and a lost reply or
a shell disposed mid-flight stays delivery-unknown rather than becoming a failure
a user would retry blindly. No golden covers those three.

The flag census grows its first entry since C1.3, which is what it is for.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the agent-history Back button to the shell handoff

C5.1 wired this button by swapping the panel's `useRouter` for `useRouteHandoff`;
nothing else was needed, because `RouteHandoff` is the router's own shape and the
seam's web sibling decides `back`. So this commit is the test that would have
caught the wiring being absent, not a fix.

Red before the merge, green after, on the same four cases: on `e897e8123a`, where
`back` was still expo-router's own spread member, 3 failed and 1 passed — the one
that passed is the local-pop case, which is the branch C2.2 did not change. After
the merge brought in C2.2's `back`, all 4 pass. The pre-merge run named the notify
by its literal `'navigate-back'` because the contract constant did not exist yet;
it is the same string `BRIDGE_NAVIGATE_BACK_NOTIFY` holds, so the two runs asked
the same question.

Both module substitutions are the builder's own rather than conveniences: the web
bundle resolves `route-handoff` and `client-context` to their `.web` siblings, so
mocking each to its sibling gives this screen the module graph it has inside the
page. The frames are read off the port pair's lane rather than off a spy, and one
case asserts a frame crossed at all before either absence is read as an answer.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): compare the resume's second write, not only its first

Round 1, finding 1. All three `resumeAiVaultSessionInTerminal` cases settled the
create (`ai-vault-resume-launch.ts:158`), so `terminal.send` had never crossed the
bridge and the half of the resume that types the command into the pane was
uncompared. Three cases now drive both writes: a refused send raises the host's
message, an accepted send reporting `accepted: false` in-band says "Terminal input
is locked", and a send the host takes resolves — the last one being the presence
precondition, since a run that failed at the create would give the same shape of
verdict as one that failed at the send.

Reading `requests[1]` straight after settling the create finds nothing on the
bridged leg: the second write is made only once the first settles, so it is two
more lane round trips away. `nthRequest` waits instead, and says how many it saw
when it gives up, so this cannot pass by proving the opposite of what it says.

The locked reply is `{ send: { accepted: false } }`, not `{ accepted: false }`:
the reader is `reply.send?.accepted !== false` (`review-terminal-reply-schema.ts:65`),
and the flat shape resolves rather than throwing. Written the flat way first, both
legs agreed on "(resolved)", which is the comparison doing its job.

Also finding 1's second half: the file docstring claimed every case runs twice and
differences the verdicts, which was false for the dispose case — a fake RPC client
has no door to shut, so there is no native run to compare against. The docstring
now says so and the case carries the same note.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin what encoding cannot save about a dot-segment id

Round 1, finding 2. The route's docstring listed `/ ? #` and whitespace and the
encoding test covered five ids of that kind, which together implied encoding makes
any id safe. It does not: `encodeURIComponent('..')` is `'..'`, so the pathname
reaches `BRIDGE_ROUTE_PATHNAME_PATTERN` intact, fails the lookahead that stops a
climb out of `/h/` (`bridge-caps.ts:68`, read through `bridge-envelope.ts:117`),
and the shell answers with `reportShellFailure` — a failure screen where the route
would otherwise have rendered the native panel it already has.

Pinned, not fixed, and the docstring now says which. `app/h/[hostId]/index.tsx`
builds its pathname identically and has the same hole, so this series fixing one
of two call sites would leave the shape behind and stop describing it. The new
case asserts both halves — the segment survives encoding unchanged, and the
pattern refuses the pathname — so a later change that starts encoding dots fails
here and has to say which screen it wants instead.

Characterisation, so it was green on the first run rather than red: the claim is
about behaviour that already ships, and the value is that the refusal is on record.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): mount the agent-history route in a real browser

Round 1, finding 3. Nothing rendered this route's real module graph anywhere. The
unit tests mock react-native, safe-area, svg, lucide and the icon assets away —
they have to, since react-native is Flow source vitest cannot parse — so a
component in this closure with no web build would have reached a device before it
reached a test. The render check is the only place the graph meets React Native
Web, and this route was not in it.

Two cases. The first mounts the route from the shell double and reads the screen:
"Agent Session History" and the worktree label the params half carried, no fault,
no console error, no CSP refusal, and the URL the page wrote for itself. That also
proves `init.route.params` end to end on a route that has a dynamic segment too,
which §1 of the design claimed and nothing checked.

The second pins the chunk. C5 is the first series whose success path pulls a
second chunk after the first paint, which on iOS goes through WKURLSchemeHandler
under `script-src 'self'`. The chunk is named from the builder's own route map
rather than guessed from the bytes, and asserted absent from what the first route
loaded, so this says the route came over the wire now.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say the real lifetime of the route-handoff refusal set

Round 1, finding 4. The comment claimed one line per reason "for the life of one
client", borrowing `createPageDiagnosticReporter`'s bound. The set is built inside
the `useMemo` keyed on `[client, router]`, so it is per hook instance: in practice
the memo is not recomputed, because `useRouter()` is expo-router's module
singleton and the page holds one client, but every screen calling the hook gets
its own set and a reason can be reported once per screen rather than once per
document.

Says that now, and why it is not tightened: a per-module set would outlive the
page's client, which is the lifetime the rest of these reporters are scoped to,
and there is no document-wide reporter to join without reaching into a contract
file the C2 lane owns.

Records the other half of the finding too, which came back confirmed rather than
changed: `console.warn` is right here. It is the vocabulary `page-bootstrap.ts:35`
already writes in, and a `fault` notify would be wrong twice — the shell drops the
generation on a page fault, and a navigation the page declined is not a failure.

Comment only; no behaviour change, 25 navigation tests unchanged and green.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): wrap every router member that takes a target, not three of five

Round 2, finding 1. `...router` hands through everything this file does not name,
and two of the members it did not name take an href: `navigate` and `prefetch`.
`navigate` to a route outside `pageRoutes` went straight to expo-router and pushed
it into this document — the hole the tri-state exists to close, reopened under a
name nobody had looked at. No call site uses it today, which is why it shipped.

`navigate` is now wrapped exactly as `push` is: which of push-or-collapse it does
is a decision about this document's stack, and a target outside this document has
no such stack.

`prefetch` is decided the other way, explicitly. It is the one target-taker that
must never reach the shell: a prefetch is a background load, `navigate` is the
only thing the shell can be told, so handing one over would open a screen nobody
asked for. A route this document serves is prefetched here, which is what the
per-route chunk split makes worth doing; every other one is dropped without a
line, because a warm-up that did not happen is not a failure to report.

The docstring's "four members that can leave this document" is now five wrapped
members and a rule for which is which.

A list would rot, so the pin is derived: `HrefTakingRouterMember` reads the
parameter tuple of every member of `RouteHandoff` and `WRAPPED_HREF_MEMBERS` is
asserted equal to it in both directions. It reads the tuple rather than testing
assignability because `() => void` is assignable to `(href: RouterHref) => void`,
which would make `back`, `dismissAll` and `reload` target-takers and prove
nothing. Checked both ways: dropping `prefetch` from the list fails the compile
with "Type 'HrefTakingRouterMember' does not satisfy the constraint", and the
union resolves to exactly the five, with `back` and `setParams` outside it.

The pin is in the product module because `mobile/tsconfig.json` excludes tests.
The runtime test asserts each wrapped member is not the router's own function and
that `setParams` still is, so a hook that wrapped everything fails too.

Red first: 4 of the new cases fail against the previous file, 33 pass now.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): say per hook instance in the title too, not per client

Round 2, finding 2. The source comment was corrected in round 1 and this test's
title was not, so the two disagreed about the bound the refusal set actually has:
the set lives in the `useMemo`, so it is per hook instance, and a title claiming
per client is the stronger promise the code does not make.

Title only; the case and its assertions are unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say what the agent-history render case does not cover

Round 2, finding 3. The docstring claimed the case is where the panel's closure
meets React Native Web, which overstates it. The shell double answers no RPC, so
the session scan fails and the panel paints its "Unable to Load" state: the
session list, its rows, the resume button and the scope tabs never render, and a
render-time gap inside any of them would pass this check.

Now says both halves — import-time evaluation of every module in the closure and
the panel's own chrome are covered, the list subtree is not — and names what
covering the rest would take: a double that answers `aiVault.listSessions`, which
is a different instrument and would put domain behaviour in this file.

Text only. This case moves to its own file on the extracted harness after the
merge with #21592; the corrected text travels with it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): import the handoff module once in its own test

My round-2 fold added `WRAPPED_HREF_MEMBERS` as a second import of
`./route-handoff.web`, which `import(no-duplicates)` fails in the focused-plugins
pass of the changed-code gate. Joined to the existing import below the mocks,
which is where an import of the module under test has to sit in this file.

Found by running the changed-code gate rather than by review: mobile tsc, whole
tree oxlint and the suite were all green with it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): give agent-history its own render file on the extracted harness

The render check gained browser cases from three domain series at once, each
under the `.mjs` cap of 600 counted lines alone and no two of them together: at
39b15e395d the file was 781 raw lines and clean, main's was 786 and clean, and
their merge was 857 raw and 623 counted, which is the CI red on #21596. C1.10
extracted the harness so a domain gets a file instead. This is C5's, and the
render check is back to 639 raw lines and clean.

Five cases. The two that moved — the route mounts and paints, and its chunk is
fetched on navigation — plus three new ones.

Back, twice. With `navigate` granted the page's Back control posts exactly one
`navigate-back` notify and the page does not move; with the grant withheld the
same tap reaches the same handler and posts nothing. The pair is the point: the
document holds the single history entry the entry wrote with `replaceState`, so a
Back this page served itself would also have gone nowhere and looked identical.
This is the first proof of that handoff in a browser rather than against a mocked
router.

And a row. The harness's new `replies` lets the double answer named methods, so
the panel now renders a real session instead of its "Unable to Load" state, which
is the render-time gap the round-2 docstring conceded. Assertions are on the row's
own text and message count, plus the absence of both silent states — the scan
failing, and a session out of scope.

Replies lifted from the corpus, and one of them needed two scenarios. The session
and worktree lists are `aivault-history-screen-listed`'s. Its `status.get` is a
capability list alone, and the first run painted "Update Orca on your computer":
`HostProtocolGate` above every host route reads the same method for fields that
scenario never scripts. The status reply merges those from
`transport-host-status-gates-ready`, and the comment says why two.

`wt-history` is load-bearing, not incidental. The panel opens on the `workspace`
scope and filters by paths from the worktree list, so on any other worktree these
same replies paint "No agent sessions" — green, and proving nothing.

Registered in the `mobile_web_app` job beside the drawer check, which is the job
that makes a missing mobile install fail rather than skip.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): encode the host id at every href the host page builds

Pullfrog on #21596. My earlier commit fixed the two hrefs in the row's navigation
sheet and stopped there; five more sites in the same page interpolate the decoded
id raw — Accounts and Tasks in both header layouts
(`host-screen-header.tsx:188,204,307,318`) and the session target
`openWorktreeSession` builds (`use-host-worktree-actions.ts:192`).

Same C1.2 class. The persisted host store admits any non-empty id and both
`useLocalSearchParams` and the store answer it decoded, so one carrying `/`, `?`,
`#` or whitespace stops being the single segment `matchesRoutePattern` reads.
Inside the page that decides where a tap goes, because the handoff matches the
pathname against the shell's `pageRoutes` before choosing this document or the
native stack.

A census rather than five more assertions: the failure is a habit, not a bug —
each of these was written by copying the one beside it, and the seventh will be
too. It counts `/h/${...}` interpolations across the host page's four source
files and requires `encodeURIComponent` at each, with a presence check so it
cannot pass on an empty list.

Two sites are exempt and stay raw: `use-host-worktree-actions.ts:171` and
`app/h/_layout.tsx:100` compare against a pathname the router answers rather than
building a link, so encoding them would change what a comparison matches instead
of what a tap opens. The exemption is subtracted by count rather than matched
away, so a file that lost its comparison and gained a raw target does not come
out even.

ONE BEHAVIOURAL EDGE, named rather than fixed. `navigateFromHostList` short
-circuits when `pathname` equals the target minus its query. That comparison now
has an encoded target on one side and whatever `usePathname()` answers on the
other, so for a host id that needs encoding the short-circuit stops firing and a
tap on the screen you are already on re-navigates instead of doing nothing. It is
a redundant navigation, not a wrong one, and the guard at :171 is unaffected
because it compares against the same raw form it always did. Left alone because
fixing it means deciding what `usePathname()` returns for an encoded segment,
which is a question worth its own change.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): guard the optional host id the session target encodes

The commit before this one did not typecheck: `useHostWorktreeActions` takes
`hostId` as `string | undefined`, and `encodeURIComponent` does not. I committed
on a green test run without waiting for `tsc`, which is my error and the reason
this is a second commit rather than an amend — the lane forbids rewriting a
commit that exists.

`?? ''` rather than a cast or a non-null assertion. An absent id then builds
`/h//session/...`, an empty segment the shell's own route rule refuses, instead
of the string "undefined", which that rule would accept as a host genuinely named
undefined. Every other member of this hook already guards the same field.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the agent-history route native when the bridge would refuse its id

CodeRabbit on #21596. The persisted host store admits any non-empty id, so `.` or
`..` reaches this route, survives `encodeURIComponent` unchanged, and fails the
bridge's own segment rule. The route handed it over anyway: `bridge-host.ts`
parses the route against `BridgeInitRouteSchema`, drops it to null when it fails,
and the page answers an `init` naming no screen with "Update Orca to open this
workspace". A failure screen, in place of the native panel sitting right behind
this switch.

The route asks the schema first now and stays native when the answer is no, which
is where every route starts. Mirrors C3.1's call for the files routes
(`69e618e19a`), including its reason for using the schema rather than a copy of
its bounds: two spellings of one rule drift, and the half that matters is the
half the page reads.

The pin moves with it. It characterised the refusal before — asserting the
pathname was built and that the pattern rejected it — and now asserts the native
render, for a dot host id and for a dot worktree id, which is the other segment
and was never covered.

`app/h/[hostId]/index.tsx` has the same hole and is not fixed here, as asked: it
builds its pathname the same way and hands it over unchecked. When C3.1 is also
on main the two guards and `mobile-file-shell-route.ts` belong in one module
beside the schema, rather than a third spelling of a one-line call.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): forward navigation options through the wrapped router members

CodeRabbit (major) on #21596. expo-router's `push`, `replace`, `navigate` and
`dismissTo` are `(href, options?)`, and the wrappers took the href alone. A local
push asking for `{ withAnchor: false }` reached the router without it, so inside
the page the router did something other than what the caller wrote — silently,
because dropping an optional argument is not a type error.

Each wrapper forwards both on its local branch now. Nothing in this tree passes
options today, which is why it went unnoticed and exactly why it needed pinning:
the first caller to pass one would have had it dropped without a word.

Options do not cross to the shell, and the docstring says so rather than leaving
it to be discovered. The `navigate` notify carries an href and nothing else, so a
target handed over is opened by the native stack on that stack's own terms. That
is the right shape — the options describe a push inside a document the shell's
target is not in — but it is a loss, and a loss worth naming.

Four existing assertions moved from `toHaveBeenCalledWith(href)` to
`(href, undefined)`. That is what the router now receives when a caller passes
none, and expo-router reads an undefined second argument as absent; the comment
above them says so, so the next reader does not take it for a bug.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): count the wrapped members the way the returned object does

Pullfrog on #21596. The header still described the set as it stood before
`860577cc30`: "the four members that can leave this document", "the three that
carry a target", "the other three". There are six wrapped members now and five
carry a target, so every count in the paragraph was one or two short and a reader
checking the object against the prose would have found neither explained.

Now says six wrapped, five target-takers named and pinned by
`WRAPPED_HREF_MEMBERS`, four decided by the shell's route list, `prefetch` the
fifth and decided differently for a reason the member's own comment gives, and
`back` the sixth carrying no target at all.

Comment only; 36 navigation tests unchanged and green.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-19 05:15:46 -04:00
Jinwoo Hong b7c06900e2 fix(mobile): give reanimated mapper hooks the inputs esbuild never writes (OTA phase C, C1.10) (#21592)
* refactor(mobile-web): extract the page render harness

The shell double, the CSP/bridge constant readers and the bundle server were
private to mobile-web-app-render.test.mjs, so a second check against the same
page had no way to reach them. Moved as-is into a module both can import; the
double also gained a `replies` map so a check can answer one method and leave
the refusal in place for everything else.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): give reanimated mapper hooks a dependency array

The bottom drawer never slid onto the screen in the web shell page: `progress`
animated to 1 and `withTiming` reported finished, but the sheet kept the
translateY of the animation's first frame and sat one viewport below the fold,
with its invisible backdrop swallowing the next touch.

Cause, bisected in the browser: `useAnimatedStyle` reads its mapper inputs from
`updater.__closure` (hook/useAnimatedStyle.js), which only Reanimated's Babel
plugin writes. The page is bundled by esbuild, which runs no Babel, so
`__closure` is undefined; with no dependency array either, `inputs` is empty and
`startMapper` registers a mapper that listens to no shared value. It runs once
and never again. Reanimated does throw for exactly this, but behind `__DEV__`,
which the bundle builds out, so the page reports nothing. The rAF loop stopping
after one write is the observable end of it.

Not a WebKit fault. Headless Chromium parks the sheet the same way
(translateY(843) vs WebKit's translateY(841)), so the earlier
JavaScriptCore-vs-V8 reading does not hold, and the pin added here runs on both
engines rather than on Chromium alone. WebKit is downloaded in the
mobile_web_app job for it.

Every mapper-backed call site takes the same array, not just the drawer's:
RightDrawer and DragReorderList are the same defect on the same bundler.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): census the reanimated hooks that need a dependency array

The drawer pin covers MountedBottomDrawer only, and the failure mode is silent:
a new `useAnimatedStyle`, `useAnimatedProps` or `useDerivedValue` without an
array animates once on the phone's native build and freezes in the web page,
with no error on either. Parsed rather than grepped so a call spanning lines,
or one whose second argument is not an array, is still seen.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile-web): state motion-on as the drawer pin's precondition

Under `prefers-reduced-motion: reduce` Reanimated finishes `withTiming` in one
frame, so a mapper that only ever runs once still writes the final translateY
and the pin goes green on the broken build. Measured: the unfixed bundle under
reduced motion lands at translateY(0) with the sheet on screen in both engines,
which is also what the Android emulator does with animator scale off — the same
single write, not a healthy animation. The context now says no-preference and
the page is asked to confirm it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): census useAnimatedReaction, whose deps are its third argument

Same fallback as the other three (hook/useAnimatedReaction.js:26-34), so the
same silent freeze applies. Its shape is not the same: the array is argument
three, behind `prepare` and `react`, and both callbacks run inside the one
mapper it starts, so both count as updaters. Indexing it like the others would
have read the `react` callback as the array. No call site today; this is the
gate for the first one.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): require the dependency array to list every value the updater reads

An array proves a call was written, not that the mapper listens to everything
it reads. On web `inputs` becomes exactly that array
(hook/useAnimatedStyle.js:338-341), so a value read but not listed is a value
the mapper never hears about: the updater stops re-running when only that one
changes. Same freeze as no array at all, in one prop rather than all of them.

Reads only. The first fixture caught this check counting `opacity.value = v` as
a read, which it is not -- a written value is an output, and demanding it in
the array would be noise at every `useAnimatedReaction`. Assignment targets and
increments are excluded; a value both read and written is still required.

Verified against the tree by dropping `translateY` from the bottom drawer's
array, which the census names at mounted-bottom-drawer.tsx:286.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): resolve the hook through the file's imports, not by spelling

Matching the callee's text both missed and invented. `useAnimatedStyle as useAS`
and `Reanimated.useAnimatedStyle` are the same hook wearing another name and
went unchecked; a local helper that happens to be called `useDerivedValue` is
not this hook and would have been flagged. Each local name is now resolved
through the file's imports from `react-native-reanimated`, named, aliased or
namespace member.

A second argument that is not a literal array now counts as present rather than
missing: the hook only needs an array to exist, and this file cannot see what a
hoisted `const deps = [...]` holds, so completeness covers literal arrays only.

Resolution can fail closed, which would read exactly like a clean tree, so the
census now asserts it saw the calls before asserting none are missing. Checked
against the tree by dropping `translateX` from RightDrawer's array, which it
names at RightDrawer.tsx:156.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile-web): select the drawer sheet by name, not by its corner radius

The pin walked up from the handle to the first ancestor with a 16px top radius,
so it found the sheet through a styling token. Change that radius and the pin
reports `sheet: false` -- a red naming the selector rather than the animation it
exists to watch, on a change that broke nothing.

The sheet now says what it is. `testID` on the RN side renders as `data-testid`
on web (react-native-web createDOMProps/index.js:832), which is the one line of
product change this needs.

Re-verified after retargeting: still red on both engines with the dependency
arrays removed (translateY 843.271 chromium, 841.447 webkit), green with them.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile-web): say why the motion option must precede navigation

Reviewer follow-up on the reduced-motion guard. The context option and the
`goto` order are both load-bearing, and nothing in the file said so: Reanimated
reads `matchMedia('(prefers-reduced-motion: reduce)')` once into a module-level
const at import (ReducedMotion.js:8-10), so a `page.emulateMedia()` after
navigation would leave the assertion passing over a value already latched true.
Comment only.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile-web): name the drawer pin's precondition instead of asserting past it

CI's Linux WebKit failed this pin at `matrix(1, 0, 0, 1, 0, 844)` -- exactly the
viewport, the mount-time value, not a first-frame 843.x. Nothing animated there,
so the pin was reporting a parked sheet without being able to say whether the
mapper was subscribed. Two different faults, one message.

`requestAnimationFrame` separates them and sheet writes do not. `withTiming`
schedules a frame per step (valueSetter.js) whether or not a mapper listens, so
frames across the window mean the shared value moved; the assertion now names
that. Counting sheet writes as the precondition inverts the diagnosis: measured
on the broken build, "written more than once" fires first and calls the defect
this pin exists to catch an engine that does not animate.

Sheet writes stay, as a second statement of the subject and as context in the
transform failure, which now reads "1 style write(s) on the sheet across 30
frame(s)" on the broken build.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile-web): wait for the drawer to arrive, not for a clock

The pin paused a fixed 1s after the sheet opened and then read the transform,
which makes it a race on a loaded runner: a healthy engine that is merely slow
reads as parked, and the red names the transform rather than the wait. It now
waits for the settled transform, times out at 15s, and asserts on whatever it
found either way, so a genuinely parked sheet gives the same red with the
timing assumption removed. On the broken build that red now reads "1 style
write(s) on the sheet across 3635 frame(s)", which says the fault in one line.

Aimed at CI's Linux WebKit red rather than proven against it: eight container
runs on the Playwright Linux image never reproduced that failure. See the
report for what the container did and did not show.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile-web): stop asserting on the sheet's style-write count

The count cannot carry an assertion in either direction. Measured under
`--cpus=0.35` in Playwright's Linux image, a healthy page starved of frames
reaches translateY(0) in a single write, because `withTiming` covers the whole
180ms in one step when one step is all the frames it gets. "Written more than
once" would have redded that page, which is a CI runner under load -- the exact
situation this pin keeps meeting.

So the transform is the only subject, `requestAnimationFrame` during the window
is the only precondition, and the write count is context in the failure text.

Also worth recording against the CI log: exactly `matrix(1, 0, 0, 1, 0, 844)`
is reproducible here on the broken build, as the single mapper run landing at
progress 0. It is the mapper's signature as much as a dead engine's, so it does
not on its own say which failed -- the frame and write counts now printed
beside it are what separate them.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): count a value read under a unary operator as a read

`isWriteTarget` took any prefix-unary parent for a write, so `!hidden.value`,
`-offset.value`, `+x.value` and `~x.value` were dropped from the reads the
dependency array has to list. A style that gates on `!hidden.value` would have
passed the census while its mapper never listened to `hidden` -- the exact
freeze this file exists to catch, hidden by the check meant to catch it.

Only `++` and `--` mutate, so the prefix branch is narrowed to those two.
Postfix needs no narrowing: `++` and `--` are the whole set there.

Red-first with a negation fixture and a unary-minus fixture; the increment
fixture holds the other side, that a value only incremented is still not
required. Found by a review bot on #21592.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-19 03:47:44 -04:00
Jinwoo Hong 381a3da46f feat(build): Route A, the phone's host routes bundled for the web, dark (OTA phase C, C0.7) (#21449)
* refactor(mobile-web): share the bundle manifest assembly with a second builder

Manifest assembly and the on-disk write move to writeMobileWebBundleTree, and
the helpers the Phase C app builder needs become exports. No behaviour change
to the shipped bootstrap bundle.

The CRLF guard grows two exemptions it needs once it is pointed at mobile/src:
the image and font extensions .gitattributes already pins -text, and the
gitignored webview engine modules the postinstall writes.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): web entry for the host route tree, and its two transport siblings

The entry mounts app/h on react-native-web through expo-router's own ExpoRoot.
It lives inside mobile/ so one React resolves, and supplies RpcClientProvider
itself: the route tree starts below the native root layout that owns it.

route-manifest.ts is a real typed module whose body the builder replaces --
esbuild has no require.context. A virtual specifier would need an ambient
declaration and would leave the entry unchecked.

Two .web.* siblings, both listed with a reason in web-overrides.json: the
transport substitution point (a placeholder client until C0.4 lands
BridgeRpcClient) and the device token store, whose native path imports
expo-secure-store, which is {} on web.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(build): build:mobile-web:app, the phone's host routes bundled for the web

Same builder shape as the Phase A bootstrap into a separate out/mobile-web-app,
with the same manifest and the same two-scratch-build determinism check. Dark:
build:mobile-web, packaging and the A2 census are untouched, and C1 is what
flips build:release.

Six shims, each a named Metro or RN Web gap. Images are emitted as same-origin
hashed assets rather than data: URLs, because the shell's CSP sets img-src
'self'; the render check under that exact header is what found it. The script is
referenced root-absolute for the same reason a <base> tag cannot be used: the
document is served at every route depth and base-uri is 'none'.

The budget sits below the contract's per-asset ceiling so growth trips a build
rather than a refused asset on a phone. esbuild splitting does not lower it:
one entry with only static imports emits one chunk (measured).

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): let React Native Web paint under the shell CSP

RN Web 0.21.2 injects its stylesheet at runtime with no nonce support, so
style-src 'self' blocks every rule and the page renders unstyled. Measured, not
predicted: the render check serves the document under this exact header and
reported the violation.

'unsafe-inline' is granted to style-src and nothing else. script-src 'self'
holds, which is the directive that decides whether page code can arrive any way
other than as a fetched same-origin script. The test now pins that scoping
rather than rejecting the token everywhere.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci: prove the Route A app bundle on every PR

A dedicated job, for the same reason the browser provider has one: it needs
mobile/node_modules and a real browser, and the sharded test matrix would pay
for both on every shard. It builds the bundle, verifies it, and runs the
builder, override-census and render suites. It ships nothing.

The mobile_web_app signal is lifted out of should_run the way static_analysis
is. A mobile-only diff is desktop-irrelevant and skips every gated job, and
that is exactly the diff that changes the page this job builds.

Also the C0.6 review follow-up: mobile/package.json and mobile/pnpm-lock.yaml
join the installer cache keys in the two workflows that build an installer off
a hashFiles key, since beforePack requires out/mobile-web and a mobile-only
change must miss those caches rather than reuse a stale build.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(build): pin the shipped builder against the app builder's own module name

The assertion named a specifier that no longer exists, so it held vacuously.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): assert the RN Web style-src grant in the Swift checks

The Swift twin of the Kotlin CSP test still required style-src 'self' and
no unsafe-inline anywhere, so it trapped on the approved grant.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(build): make the Route A render check name what each route paints

The check asserted only "some html, no errors", which expo-router's Unmatched
screen satisfies: pointing HOST_ROUTE at /zzz/not-a-real-prefix stayed green.
Each route now asserts content only its own component produces, and the
unmatched case asserts the screen positively so the negatives discriminate.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(build): read the shell CSP past the comments that quote directives

Both constants document themselves with // comments containing quoted
directive text, which the quoted-string scan picked up as directives. One
parser now drops comment lines, and iOS and Android go through it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(build): honour a .web.* route sibling in the app bundle

Routes were imported by absolute path with the extension, so esbuild's
resolveExtensions never applied and a .web.tsx under app/ was dead code the
census still accepted. The manifest now carries a key and a module: the key
stays the native filename so the URL does not move, and the module is the web
sibling when one exists.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(build): tie each named shim to the esbuild option that implements it

The shim list was asserted against a literal copy of itself, which passes
however the build is configured. Each entry now carries an appliesTo that
reads its own option, checked against the real options object.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore(build): line up the CRLF exemptions, the budget comment, and the job scope

The builder loads .gif as a file but neither .gitattributes nor the CRLF scan
exempted it, so the blanket eol=lf pin would have rewritten one. A test now
keeps the two lists in step. The Phase C byte budget's comment sat on the
asset count, and a root package.json edit could change build:mobile-web:app
without running the job that proves it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(build): satisfy the index-check lint rule in the CSP parser

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci: key the installer caches on the mobile page trees too

beforePack builds the mobile web bundle into the installer. Today those bytes
are Phase A's, which src/** already covers, but once C1 flips the entry to
mobile/app a page-only change would hit a cache holding a stale installer.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(build): skip the bundling tests where mobile dependencies are absent

The sharded `test` job collects config/scripts/**/*.test.mjs and installs no
mobile dependencies, so the two new suites failed there on "Could not resolve
react-native-web". They now skip themselves with a message naming the job that
runs them, and that job sets ORCA_MOBILE_WEB_APP_DEPS_REQUIRED so a missing
install fails it instead of skipping everything it exists to prove.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(build): scan mobile/packages in the .web.* census

The census claimed the app entry never resolves into packages/, but the
dictation hook imports @orca/expo-two-way-audio and the built script carries
ExpoTwoWayAudioModule.web.ts. That file is now listed with its reason, and
planting a .web.* in each scanned tree proves the scan is not passing because
a tree happens to be empty.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(build): assert the route exclusions against a tree that has them

mobile/app holds no test, spec or +api file, so the exclusion rule was
asserted against a tree it could not fire on. A scratch tree plants one of
each; dropping the rule now fails this test.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(build): 404 unknown file paths in the render check's page server

The server answered every path with the document, so pointing publicPath at
/wrong-prefix still rendered three green routes: the script is fetched from
the one prefix that is served. A path naming a file now has to come out of the
bundle, which is what the shell's manifest map does.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(build): cover the app bundle verifier's own checks

The verifier had no test. One doctors the buildId, which the packaged assert
catches; the other rewrites the tree so every digest still agrees and only the
two fresh builds can tell, which is what a stale out/ looks like. Deleting
either check now fails a test.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore(build): tidy the app bundle comments and the job's path prefixes

Drops an export nothing read, merges two comments that had drifted apart from
the constant they describe, and corrects the claim that the job runs on every
PR when it is path-gated. package.json leaves the prefix list because
GLOBAL_FORCE_FILES already forces every job on it; mobile/packages/ joins it,
since the page resolves a .web.ts out of there.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(build): merge the duplicate node:fs/promises import in the census

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): redirect the hybrid shell route on the web page

app/h/[hostId]/web.tsx reaches OrcaMobileWebShellView, whose module calls
requireNativeViewManager at import. In a browser that throws before React
mounts, and the route manifest imports every route statically, so one native
route left the whole page blank at every URL.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(build): fail the render check with the error that stopped the mount

The check waited on "#root has children" with Playwright's animation-frame
polling, so a route module that threw at import read as a bare 30s timeout
naming nothing. It now waits on a mount attribute the entry sets after the
router commits, polls on a timer, and races the wait against the first
uncaught error so the failure carries it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(build): answer the favicon the render browser asks for

CI resolves the runner's Google Chrome, which requests /favicon.ico; the
bundled headless shell does not. The bundle carries no icon, so the server
answers 204 rather than turning a browser habit into a console error the
render assertions read as a page fault.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(build): settle the render check's uncaught-error race without rejecting

The entry throws during goto, before anything awaits the race, so a rejected
promise surfaced as an unhandled rejection beside the real failure. The same
signal now resolves with the error.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore(mobile): list the page transport in the raw request port inventory

The placeholder client implements the port, so the boundary test counts it as
an unlisted file. It belongs under OWNERS until C0.4's BridgeRpcClient
replaces it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-18 09:50:37 -04:00
Jinwoo Hong ca2ae89011 ci: install mobile dependencies in every desktop packaging job, pin mobile page source to LF (OTA phase C, C0.6) (#21425)
* chore(mobile): pin mobile page source to LF so a Windows checkout keeps buildId

The Phase C web bundle hashes every text byte under mobile/src and mobile/app into
its asset digests and from there into buildId. There is no global text=auto, so a
CRLF checkout on Windows would give the Windows release a different buildId for
identical source, the same failure the src/mobile-web pin above exists for.

All 1924 tracked files in those directories are already LF in the index, so the
pin renormalises nothing. mobile/web-entry does not exist yet; the pin is
forward-looking for the Phase C entry point.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci: install mobile dependencies in every desktop packaging job

Ten workflows reach build:release/build:desktop and none of them installs
mobile/node_modules. Root has no react-native, react-native-web or expo, so once
the mobile web bundle builds from mobile/ its packaging jobs would fail at
electron-builder's beforePack with an unresolvable import.

Extract the frozen mobile install that pr.yml's static analysis job already ran
inline into .github/actions/install-mobile-dependencies, and invoke it from every
job the packaging census enumerates, after the root install and before the build.
Same --frozen-lockfile, same lockfile-drift guard, and still no --ignore-scripts:
mobile's postinstall generates the gitignored webview engine modules that tracked
source imports. pr.yml now uses the action too, so there is one definition.

Where a packaging job's setup-node caches the pnpm store, mobile/pnpm-lock.yaml
joins cache-dependency-path so a mobile lockfile change invalidates it. Two jobs
(daemon-relocation-spike, win-update-survival-e2e) do not cache at all and are
left alone.

The census test grows a per-job assertion that the action is present, so a new
packaging job has to add the install deliberately rather than discover it at
beforePack. release-cut's composite-action restore is no longer Windows-only:
every platform consumes this action now, so any of them can be the leg whose cut
ref predates it.

No job builds anything different; this only makes mobile/node_modules present.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(ci): assert the mobile install contract on the shared action

pr.yml's static analysis job no longer carries the install inline, so the scope
test's findIndex by step name resolved to -1. Match the step by the action it
uses, and read working-directory and --frozen-lockfile off the action itself so
the job cannot keep the step while the action stops installing anything.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore(mobile): exempt binary asset types from the mobile LF pin

/mobile/{src,app,web-entry}/** text eol=lf would mark a future PNG or font as
text and rewrite its bytes on a Windows checkout. Exempt the asset types an RN
page carries, the same way src/mobile-web exempts its PNG.

-text after text eol=lf wins: probed a CRLF-bearing .png under the pin, it stays
i/crlf attr/-text while a sibling .ts still normalises to i/lf. No tracked file
changes classification; the 1924 files under mobile/src and mobile/app stay i/lf.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci: gate the mobile install with the build it feeds in the cached lanes

win-crash-survival, win-update-survival and daemon-relocation-spike all skip
electron-builder on an installer/unpacked cache hit, so an unconditional mobile
install spent time on node_modules nothing then consumed. Move each `uses:`
below its cache step and carry the same cache-hit condition as the build.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-18 05:44:11 -04:00
Neil 5c2d3322c1 fix(runtime): name a terminal whose pane a graph republish dropped (#19860)
* fix(runtime): name a terminal whose pane the graph dropped

`buildPtyTerminalSummary` decided `orphaned` from the PTY record's agreement
with itself — `!pty.tabId || !pane || pane.tabId !== pty.tabId`. A record whose
`paneKey` still parses to its own `tabId` passes that forever, including long
after the session graph dropped the pane, so a terminal that had lost its
surface reported `orphaned: false, connected: true, writable: true` and a
`tabId` no tab has: field-for-field identical to a healthy one (#18191).

Consult the leaf topology instead, gated on a graph statement having had the
standing to contradict the record. `graphSequence` counts authoritative graph
statements; every statement re-records the surface of every pane it publishes,
so a pane the current graph holds carries the current stamp and is answered
without touching the leaf map. That covers the two absences that are not
evidence, without a second flag: a surface recorded since the last statement
(spawn records the pane before the graph carrying it arrives, #7587), and a
lost graph clearing every leaf at once without advancing the sequence. A pane
already observed dropped keeps its stale stamp and stays named, because losing
the ability to re-check is not a reason to un-see it.

`orphaned: true` is shipped vocabulary that both consumers already read, so no
capability gate is needed: adoption keys on it (`hasStrongOrphanIdentity`) and
now reaches this population, and the duplicate-surface index
(`indexLiveTerminalSurfaceOwners`) stops recording a destroyed pane as a PTY's
live owner.

* fix(runtime): publish a terminal retirement proof on the exit's own evidence

A paired client may drop a mirrored terminal on exactly two kinds of host
evidence: a `retiredTerminalSurfaces` proof naming the handle, or two
authoritative `terminal.list` inventories that omit it. The second needs two
host publications, and a quiet workspace publishes one, so the proof is the
only evidence that rides the frame carrying the retraction.

That proof was minted only as a byproduct of persistence *accepting a change*,
which made one value carry two meanings: "a change was accepted" and "the PTY
exited". The host renderer's close transaction de-persists the surface and
republishes without it, so when it got there first the exit found nothing left
to accept and the attestation died with it. Measured on a real paired client:
the host retracted in under 500ms, published no proof, then froze its
snapshotVersion for 60s while the client kept a dead pane in its tab bar.

Persistence still gates *removal* — publishing absence before the membership
fence is durable would let a crash resurrect the surface. It no longer gates
the proof: the observed exit is itself the attestation.

The exit-first ordering already had a passing test; the renderer-first ordering
had none, and that is the one users hit. Both orderings are now pinned, with
exit-first as the control that makes the renderer-first failures mean something.

Wire: `retiredTerminalSurfaces` is an existing optional field on an existing
path, already negotiated as `session-tabs.retirement-proof-delta.v1`. This is
Rule 1 — an old client that ignores it degrades to the two-inventory route it
already uses today, so no capability gate is needed. The sentence "the host
starts sending a frame it did not send before" reads like Rule 3; it is not,
because the frame shape, the field, and the reader contract are all unchanged.

* test(runtime): pin the removal frame retiring a still-live publisher

KNOWN RED (`it.fails`), no product change. Found while verifying the close
retraction fix: once the emptying actually reaches paired clients — a state the
previous behaviour never allowed, because nothing propagated — re-adoption of a
later create is flaky. Measured 1 failure in 6 runs of the two-client journey.

`decideWebSessionTabsSnapshot` treats the host's synthetic `removed:<t>`
retraction as a publisher handover: it retires the still-live renderer epoch and
installs the retraction as current, while the removal also clears the live
freshness record. The next frame from that same running publisher then matches
no lineage and reads as a retired generation, so it is outranked and the
publisher is locked out of the worktree until its generation changes.
`local-structured-session-tabs-sync/snapshot-apply.ts` documents this exact
scenario and has a revive escape; the mirror path has none.

The suffix case explains the 1-in-6: `hasRetiredValue` is an exact string match,
so a republication carrying `:headless-merge:` walks past the fence and only a
bare same-epoch republication is locked out.

Not fixed here on purpose. Dropping the retirement makes the red case pass but
breaks `web-session-tabs-sync.test.ts > keeps a removed worktree fenced against
delayed predecessor epochs`, which asserts a same-epoch higher-version frame
after a removal must be rejected. At this layer those are the same frame — this
function holds no `receivedFrame`, so it cannot separate a delayed predecessor
from the live publisher speaking again. The fix belongs in
`shouldApplyRecoveredWebSessionTabsSnapshot`, which does hold that ordering and
currently defers to the same epoch fence. That is a contract change across two
functions and an existing invariant, not a one-liner.

* fix(runtime): a removal retraction is not a publisher handover

The host drops a worktree's entry when its last tab closes and announces it
with a synthetic `removed:<t>` epoch. Both receipt sites treated that as a
publication: `decideWebSessionTabsSnapshot` and
`recordReceivedWebSessionTabsSnapshot` each noted the retraction epoch as
current, which pushed the still-live renderer epoch onto `retired`. The removal
also drops the live freshness record, so the next frame from that same running
publisher matched no lineage, read as a retired generation, and was outranked.
The live publisher was locked out of its own worktree until its generation
changed. That is fail-closed, and it is why re-adoption after an emptying was
flaky once the emptying actually reached paired clients.

A retraction and the live publisher's next frame are the same epoch at a higher
version, so epoch identity cannot separate them and never could. Delivery order
can. `recordReceivedWebSessionTabsRemoval` now records the retraction as the
worktree's newest received evidence instead of deleting the ledger, so
`shouldApplyRecoveredWebSessionTabsSnapshot` — the gate every production apply
path passes before `decideWebSessionTabsSnapshot` — fences a frame that
reserved its received frame before the retraction while admitting one that
arrives after it. The boundary carries the retraction's own epoch, which never
matches a host publication, so a later live frame may still restart its version
counter. `local-structured-session-tabs-sync/snapshot-apply.ts` documents the
same conclusion for the local path: a retired epoch is not proof of a dead
generation.

`keeps a removed worktree fenced against delayed predecessor epochs` pinned the
delayed predecessor at the raw decision layer, which is the same call as the
live publisher's republication. It now pins the identical scenario — same
epoch, higher version, still rejected — through the receive-and-apply path that
actually holds the ordering, plus the composed gate as production spells it.

The committed `it.fails` repro is not sufficient on its own: it records no
received frame, so dropping only the `decideWebSessionTabsSnapshot` retirement
turns it green while the publisher stays locked out on every real path. A
receive-and-apply case is added alongside it to close that gap.

* test(runtime): pin the retraction boundary against a stale inventory omission

Mutation testing left a survivor: writing the boundary unconditionally, instead
of only when it advances the ledger, passed the whole runtime suite. It is not
inert. A visibility-resume inventory reserves its received frame before it
lists, so an omission it reports can be older than a stream frame that landed
meanwhile; without the guard that stale omission rewinds the ledger, forgetting
the stream frame's version, and a delayed list reserved in between is then
readmitted instead of outranked. This pins that ordering.

The one remaining survivor is the boundary's `snapshotVersion`, and it is inert:
the ledger's version is read at exactly two sites, both reachable only when the
incoming frame's epoch equals the stored one, and a retraction epoch never
equals a live publication.

* test(runtime): cover the fences the retraction change narrowed

Two gaps found by mutating the fences themselves rather than the fix.

Deleting the epoch fence in `shouldApplyRecoveredWebSessionTabsSnapshot` passed
the entire runtime suite. It is not unreachable: a superseded generation whose
sibling stream delivers its frame after the handover outranks the successor on
delivery order, and only the retired-epoch check rejects it. Retractions used to
exercise that fence too; now that they no longer retire anything, a genuine
handover is the only thing left that reaches it, and nothing covered that. The
fence is narrower than it was, not dead.

The second case pins rate-independence. The defect surfaced 1 run in 6 because
`hasRetiredValue` is an exact string match while `sameSessionTabsPublicationLineage`
treats `:headless-merge:` as the same publisher, so a merged republication walked
past a fence a bare one hit. The removal path is now asserted over both epoch
shapes through the full path, so a fix that only re-rated the defect instead of
removing it would fail here.

* fix(runtime): give "same publisher" one answer across the epoch fences

Separable from the retraction fix beneath it, and it changes handover-path
behaviour: a superseded generation that republishes under a merged epoch is now
rejected where it was previously accepted. Take it independently or not at all.

`publisher-identity-fences.ts` held two answers to "is this the same publisher".
`noteRetiredValue` treated a `:headless-merge:` epoch as a SUCCESSOR of its base
and retired the base when the merged form became current, while
`sameSessionTabsPublicationLineage` treated the two as ONE publisher. Those are
contradictory, and the retired-value check's exact-string match was the shim
that kept them from ever meeting: a merged frame was a different string, so it
never looked retired no matter what had been retired.

The cost was that the same predecessor was accepted or rejected depending on
which shape it arrived in. A generation a successor had replaced was fenced when
it republished bare and admitted when it republished merged — the fail-open half
of the same disagreement whose fail-closed half was the removal defect, and the
reason that defect reproduced 1 run in 6 rather than every time.

This cannot be fixed in the fence alone. Making the fence lineage-aware while a
merged epoch still retires its base has the generation retire itself: the
rebuild arrives, retires its own base, and the fence then rejects it as a
retired generation. So both sides move together — a lineage sibling advances the
current epoch instead of superseding it, and inherits its generation's
retirement instead of escaping it.

Scoped to the publication-epoch functions. Runtime-id retirement keeps exact
matching, and `local-structured-session-tabs-sync` keeps its own
`hasRetiredValue` call, where a lineage sibling is already excused explicitly
and a retired epoch is deliberately not treated as proof of a dead generation.

* test(e2e): journeys for a reopened client and two clients on one host

Two gaps this suite had no coverage for, both driven end to end against a real
paired desktop client rather than at a seam.

A relaunched client holding a live remote terminal: every paired restart spec
here restarts around a browser pane, none around the terminal the user is
actually mid-work in. The host-side fixture's on-disk sink is the oracle — one
READY for the whole run proves the host never re-spawned the session, and a
recorded line for input sent after the relaunch proves the restored pane is
wired to that same process rather than painted with its scrollback.

Two clients on one host across an emptied workspace: the tombstone is
client-local on the runtime path, so a client that never held a row still seeds
into a workspace another client deliberately emptied. That asymmetry is by
design; a client falling out of step with the host and staying there is not.
Phase 0 is the control — without it a later divergence cannot be attributed to
the emptying rather than to mirroring never having worked.

The input probe goes through `pane.terminal.input`, not `window.api.pty.write`:
a mirrored pane's handle is a `remote:` id that no local PTY answers to, so a
direct write is swallowed and the assertion passes on nothing. The pre-restart
control exists to catch exactly that, and did.

* test(e2e): keep the two-client journey spec type-clean

* test(e2e): pin the close retraction a paired host does not publish

* docs(e2e): say why the red close-retraction spec sits on this PR

The spec was written on a branch carrying neither of this PR's publish-side
fixes, and its own diagnosis -- the fault is the host's publish-after-close,
not any client's mirror -- names exactly what they change. Landing it here
makes CI the measurement rather than leaving a red spec parked on a branch
with no fix in it.

Records the one thing a reader needs to not do: skip-tagging it. And why the
obvious split is not a block move -- phase 2 depends on phase 1b's emptying
and both share the two-client pairing fixture, so splitting means duplicating
the fixture.

* test(e2e): the close-retraction spec is green on this branch, measured

It was written to pin a defect and was red where it was written. On this
branch, with `publish a terminal retirement proof on the exit's own evidence`
and `a removal retraction is not a publisher handover` both present, it passes
-- twice, independently: phase1a A=9ms/B=158ms then A=2ms/B=1ms, against a
prior baseline of "none reached either client within 90 seconds".

So the KNOWN RED header had become the thing it warned about: a test carrying
prose asserting the very behaviour the commits beside it remove. Rewritten to
record the measurement and the numbers to regress against, and to keep the one
instruction that still applies -- if it reddens again, do not skip-tag it; the
failure shape is a 90s timeout on both clients at once while creates still
propagate.

No assertion changed. Comment only.

* test(wire): pair the session-tabs retirement proof across two builds

The stack makes a host start sending a retirement proof on its own frame
when no surface removal carries one. The change argues Rule 1; Rule 3's
fourth bullet covers a frame the host starts sending on an existing path,
so the claim is measured against v1.4.199 rather than accepted.

Neither existing cross-version suite reaches session-tabs: the terminal
one covers the binary stream, the agent-session one covers agentSession.*.

Result: the old client acts on the proof-only frame, because the whole
client half of this surface is unchanged. The old-host cells are pinned
to a release that cannot publish the frame at all, which is what makes
the new-host cells mean something.

* fix(lint): clear the casting gate on the surface-lost inventory

main tightened typescript/consistent-type-assertions to assertionStyle:
never, which the rebase brings onto these added lines. The retraction
read narrows on the property instead of casting; the fixture and
cross-build-import casts carry per-site SAFETY rationales.

* fix(lint): bind the protected-stamp cast to a name

The leading-semicolon parenthesised call put the suppression on a line
oxfmt then reflowed away from the assertion it covers. Naming the
narrowed handle keeps the directive next to the cast.

* fix(runtime): route every non-null surface write through the stamped writer

`ptyHoldsRecordedSurface` trusts a record only while its stamp is current;
after that the leaf map answers. Four writers still named a pane with a bare
`tabId = / paneKey =` — orphan adoption (both branches), split, create on an
adopted stable pane, and TUI-owner recovery — so a record that had already
been contradicted stayed contradicted after the claim, and `terminal list`
reported the just-claimed PTY `orphaned: true` until the renderer's next
graph statement re-recorded it. Before this branch those sites read as
attached at once, so this was a regression window of one round-trip, and
`indexLiveTerminalSurfaceOwners` reads `orphaned` as "unowned".

`recordPtySurface` is now the one writer; the adoption module reaches it
through a port because it has no `graphSequence` of its own. The nulling
writers are untouched: a null surface is never held, stamped or not.

* test(runtime): keep one copy of each publisher-fence case

The removed-frame suite asserted four properties that another case in the
same suite or the lineage suite already pinned:

- the decide-only readmit and the bare full-path readmit are the bare arm of
  the parameterized full-path readmit, verbatim;
- the merged-suffix decide-only readmit is the merged arm of the same loop;
- "still fences a predecessor a successor replaced" is the lineage suite's
  bare arm with different version numbers;
- the recovery-gate handover case is the lineage suite's recovery-gate case
  with a bare late frame instead of a merged one, so that test now runs both
  shapes and this copy goes.

Mutation-checked: reverting each of the five renderer changes on this branch
(retire-on-removal in decide, noting a retraction current, the exact-match
retired fence, merge-supersedes-base, dropping the ledger on removal) still
fails at least one of the remaining ten cases.

Also corrects the suite header: a retraction carries a synthetic `removed:`
epoch, so it is the in-flight predecessor frame, not the retraction, that
shares the live publisher's epoch and needs delivery order to be separated.

* test(e2e): fail the two-client journey when phase 1a cannot run

Phase 1a sat inside `if (beforePartialClose.length > 1)`. A host workspace
that starts with one terminal skipped the control silently while 1b and 2
still ran, and the spec passed green without ever exercising the
close-with-others-open retraction it was written to measure. The skip is now
a recorded failure naming the host count.

* fix(runtime): order every session-tabs apply path against the retraction

A closed terminal came back on the other client because "this worktree was
retracted" was neither durable nor universal:

- `refreshWebRuntimeSessionTabsSnapshot` reached `decide` with no place in
  receipt order at all, so a list the host answered before the close applied
  after the retraction had already cleared the worktree. It is a production
  path for close, create, activation, split and PTY reconnect.
- the boundary lived in a single receipt slot the next stream frame overwrote,
  and in a fence that only existed when a recovery happened to be pending when
  the retraction landed, so a pre-close list could out-rank the republication
  on `snapshotVersion` alone.

Replace both with one raise-only removal watermark per (environment, worktree)
and give the list path a receipt position, reserved by the request and carried
in its answer so a dedupe joiner inherits it rather than minting a newer one.
The pending-recovery fence and its bookkeeping are dead once the boundary is
monotonic. The exact-match retirement check in the receipt ledger becomes the
one lineage-aware predicate, so a `:headless-merge:` rebuild can no longer be
noted as current and retire the live publisher out of its own worktree.

On the main side, `recordPtyWorktree` stamped `surfaceRecordedAtGraphSequence`
at write time, so any `paneKey` write claimed the standing of a fresh graph
statement. The inventory restore in `terminal list` therefore un-dropped the
very pane the read was meant to report, on every listing. A surface claim now
carries no graph standing unless its writer names one: the graph statement,
live leaf output and spawn do, while the inventory restore, the floating
liveness restore and the mobile projection replay do not. Defaulting this way
means a writer that says nothing fails safe and self-corrects, which the type
alone could not guarantee across the projection contract's own `recordPty`.

Spawn claims now span the one graph statement the renderer may already have in
flight, and retirement proofs compare by identity instead of by position, so a
re-delivered exit no longer fans out a `snapshotVersion` bump carrying nothing.

* fix(runtime): stop an unpublished-worktree placeholder retiring the live publisher

A worktree the host has published nothing for still answers a forced list, with
a synthesized `none`/v0 frame that means "ask me later"
(host-session-snapshot-authority.ts). Every post-close list and every
activation of an emptied worktree gets one. Noting it as a publication retired
the renderer generation that is still live, and because that epoch is
per-process, the terminal the user created next never reached this client — the
same lockout the retraction path was already careful to avoid, through a door
it did not cover. `local-structured-session-tabs-sync` already skips the
placeholder for this exact reason; the web mirror now does too, on both the
receipt ledger and the frame decision.

Bound the receipt ledgers by frame age rather than entry count. One bootstrap
inventory records a receipt per worktree under a single reserved frame, so
evicting by insertion order dropped that batch's own earlier entries, and an
absent receipt is what the recovery gate reads as "no evidence for this
worktree". Only a receipt no in-flight frame can still be ranked against is
droppable.

Take the receipt gate off the `web-session-tabs-sync` barrel in the refresh
path. Ordering is that path's gate, not an optional collaborator a caller's
module mock may leave out, and being reachable only through the barrel is how
the path came to have no ordering at all.

* fix(runtime): let the TUI-owner recovery name its pane without claiming the graph holds it

`recoverStructuredTuiOwner` rebinds a recovered PTY from the persisted owner
binding — the same replayed-evidence class as the inventory restore — but
stamped it with the current graph sequence, so a pane the renderer had already
dropped read as attached for one more statement. The guard below it needs the
tabId and paneKey, not the standing.

Also say plainly in `decideWebSessionTabsSnapshot` what the affirms check does
and does not cover: an unpublished-worktree placeholder is withheld from epoch
noting only. It still applies, because rejecting it outright would drop the
terminal reconciliation that legitimately rides on it.

* fix(runtime): keep the retraction boundary out of the receipt bound

Bounding the removal watermark alongside the receipt ledger reintroduced the
defect the watermark exists to prevent: past 512 retracted worktrees, evicting
a boundary readmits every pre-close frame it was fencing, and a delayed list
resurrects the closed tab. A boundary is not a cache. One number per worktree
ever retracted on an environment is the cheaper price, and environment teardown
drains it; only the receipt ledger stays bounded, by frame age.

Split the orphan-adoption port by provenance so the last writer that disagreed
with the surface-standing rule stops disagreeing. `adoptRuntimeTerminalOrphans`
replays the persisted binding when the claim already matches it and writes a
new one otherwise, and both went through a single `recordSurface` that stamped
the current graph sequence — so re-adopting an already-adopted orphan lifted a
dropped pane's stale stamp and reported it attached, in a quiet workspace
possibly forever. The replay now names the pane without standing and the fresh
claim takes spawn standing, like every other writer.

Replace a receipt-count assertion that was vacuous for a map keyed by
environment and worktree with the mirror state and freshness it was standing in
for.

* fix(runtime): keep a closed-tab worktree under the epoch already publishing it

`closeHeadlessMobileTerminalTab` minted `headless:<now>` on every close. Its
sibling headless writers carry the stored `publicationEpoch` forward and mint
only when there is no snapshot to inherit from — because a write to a worktree
is not a claim to publish it. The close was the one writer that claimed.

A paired client retires the epoch a new publisher displaces, and the web
mirror's retirement is final: there is no revive lane, and the per-worktree
tracking teardown deliberately keeps the epoch history. So an ordinary close
published a stranger for a worktree the renderer generation still owned, retired
that generation on every client, and the renderer's next publication — carrying
the epoch the close had just retired — was rejected forever. The user emptied a
workspace, created a terminal, and it never arrived on either machine while
`session.tabs.list` showed the host holding it.

This is the same thesis the retraction path already states, through the door
next to it: a retraction is not a handover, and neither is a close.

Measured on `paired-two-client-emptied-workspace-reseed.spec.ts`, six runs each:
phase 2 failed 3/6 before (`A=null B=null`, both clients blind for the full 30s
budget) and 0/6 after, with both clients adopting in single-digit milliseconds.

* fix(lint): give the fixtures real types instead of casting past them

The casting gate failed on eight assertions this branch added. All eight were
suppressible, but the suppressions were not the problem: the casts were hiding
fixtures that did not match the contracts they stood in for.

`sessionStillHoldingBothPanes` built tabs as `{id, title, type}` — `type` is not
a `TerminalTab` field and eight required ones were missing — and layouts holding
only `ptyIdsByLeafId`. `as never` made both compile. They are now real
`TerminalTab` / `TerminalLayoutSnapshot` values, so the fixture is checked
against the type `listTerminals` actually reads.

`terminalTab` in the epoch suite built a *client* tab (`status`, `terminal`) for
a field typed with *snapshot* tabs, which forced `as never` at the call and a
cast on the snapshot itself. Production reads only `type`, `parentTabId`,
`leafId`, `ptyId` and `parentLayout` from that tab, so the two client-only
fields were inert; dropping them lets the declared
`RuntimeMobileSessionTerminalTab` type the fixture end to end, and the closed tab
is now held by name rather than recovered from `snapshot.tabs[0]`.

The remaining three casts are unchanged in kind and now carry correctly placed
SAFETY rationales: reaching a protected member is the only way to drive these
paths. `graphSequence` folds into the reach-through that was already there
rather than opening a second one, and the map read narrows instead of asserting.

Mutation-tested, all three suites, regression re-introduced for each:
- epoch mint on close restored -> 1 failed | 1 passed
- orphan check reverted to self-consistency -> 4 failed | 4 passed
- placeholder retirement guard removed -> 1 failed | 7 passed

src/main/runtime 8169 passed | 31 skipped; src/renderer/src/runtime 1581 passed.
`check:code-quality:changed` goes 8 findings -> 0. `pnpm tc` clean.

* fix(runtime): stop the headless placeholder graph from dropping every restored pane

A headless server publishes one empty graph at launch so status clients see a
ready server. It names no renderer pane and is never replaced, but it was
counted as an authoritative graph statement all the same: `graphSequence` went
0 -> 1 while the leaf map stayed empty for the life of the process.

Every surface claim written without standing - a persisted replay, an inventory
restore, the TUI-owner recovery - is stamped 0. Against `graphSequence` 1 the
`>=` guard fails, the empty leaf map answers "no pane holds this", and the
terminal reports `orphaned: true` under a `pty:` tabId. Nothing can re-stamp it,
because the only graph that host will ever publish has already been published.
On a headless or SSH host that is permanent, and it is the same lie #18191 is
about, pointed the other way.

The placeholder no longer spends a graph statement. A renderer graph still does,
so a pane a real graph drops is still reported dropped - including on a desktop
window promoted from headless, which the third case pins as a negative control.

Mutation: restoring the unconditional bump fails the first two cases
("expected 1 to be +0", "expected true to be false"); the promoted-window
control passes either way, as a control should.

Also registers tests/e2e/cross-version-wire/cross-version-session-tabs-retirement-proof.unit.test.ts
in the cross-version-wire job. The file matches CROSS_VERSION_WIRE_PREFIXES, so
adding it had switched the job's gate on, but the job runs an explicit file list
that omitted it - the test executed nowhere in CI. It passes 8/8.
2026-09-18 01:56:07 -07:00
Neil 9d1826ae65 fix(session): repoint the rows a worktree re-key strands (latent; producer is flag-disabled) (#20057)
* fix(session): keep a renamed worktree's rows from matching on the id it lost

Three persisted session fields survived a worktree re-key still naming the old
identity. Two of them are suppression records, so a stale id does not read as
residue -- it silently re-admits state the user removed:

- closedTerminalTabTombstonesByTabId: the remote merge only suppresses a host
  tab when the tombstone's worktree equals the tab's, and no snapshot ever
  covers the old id, so the tombstone never retires either.
- clientHostedBrowserCloseIntentsByEnvironment: the replay targets the intent's
  worktree, and an unresolvable selector answers selector_not_found -- which the
  replay reads as definitively gone and uses to DROP the intent.
- clientHostedBrowserPagesByWorktree: keyed by worktree and re-checked against
  the row's own workspaceId, so both halves have to move or the pages are never
  rehydrated.

Fixed on both sides of the rename: the main-process persisted migration and the
renderer's live store, which would otherwise write the stale values straight
back. The coverage test drives off WORKSPACE_SESSION_WORKTREE_REFERENCE_KIND,
the census these three fell out of, with the shipping owner collector as its
oracle.

* docs(session): record why a re-key clobbering an existing target stays unfixed

Not a missing guard -- an unresolvable one. Keeping the target is correct when
it holds a real closed-last-terminal tombstone; keeping the source is correct
when the target row is a stub; nothing records which is newer. The recency map
is the only one that can settle it, because Math.max needs no such ordering.

* test(persistence): measure the downgrade direction for worktree identity

The stack widens migrateWorktreeIdentity to repoint worktreeId inside
session rows. That changes what lands on disk with no wire change, which
is Rule 3's shape applied to persistence, so it is measured against
v1.4.199 rather than reasoned about.

Result: new-build state does not break the old build. The old build
renames over it without throwing and loses no row; the two row kinds it
cannot repoint stay stale, which is exactly what its own renames already
produce.

The numbers are measured. A first draft asserted the old build repointed
no inner rows at all; it repoints two of four, and the probe is what
caught that.

* test(ci): run the worktree-identity downgrade lane instead of describing it

The cross-version job names its files explicitly, so a new one is inert until
it is listed; the sharded unit job excludes the whole directory and the E2E
router only takes `*.spec.ts`. Also pairs the forward-compat case against the
current build — the stack's own field-list walk is the guarantee that matters,
and only the frozen build was exercised.

* refactor(session): drop the type assertions the rename migration leaned on

`consistent-type-assertions` landed on main after this branch last built, and
three of the `as never` fixtures were hiding real contract drift: a browser
workspace row missing six required fields, a tab group naming three fields the
type does not have while omitting the two it requires, and a sleeping-agent row
whose `providerSession` had neither `key` nor `id` and whose `state` was not in
`AgentStatusState`.

Indexing the session by a computed field name is what forced the casts in the
migration, so the four row maps are now spelled out; the census test is what
keeps a fifth from joining silently. The renderer test builds its state from
the real slice instead of casting a four-field partial.

* refactor(test): name the module namespace the skew harness reads

`object` is too broad for the anti-slop gate, and the import helper already
declares what it hands back.

* docs(test): say which maps the harness actually supplies

The two under test live in slices this harness does not mount, so calling it
"the real slice's state" overclaimed.
2026-09-18 01:11:03 -07:00
Jinwoo Hong ad4f26cdd4 feat(build): build, verify and package the mobile web bundle with every desktop release (OTA phase A, 2/5) (#21326)
* feat(mobile-web): add the Phase A bootstrap web source

A peer of src/ so the root workspace owns it and mobile's separate lockfile
stays out of packaging. Four assets across four content types, enough to
exercise multi-asset manifest handling rather than assume it.

The page reads buildId from manifest.json at runtime: buildId hashes the asset
list that index.html belongs to, so injecting it into a hashed asset would make
that asset's hash depend on itself.

Registered as a fourth typecheck project; without it the entry would be the
only TypeScript in a release path that tsc never sees.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(build): build and verify the mobile web bundle from the root workspace

Root esbuild over mobile-web/ into out/mobile-web/, content-addressed as
assets/<sha256>.<ext> with index.html the only stable name. buildId is the
sha256 of the canonical serialization of the sorted asset list, so it is a pure
function of content and usable as a cache key with no further reasoning.

The verifier builds twice into scratch dirs and compares: a timestamp, an
absolute path, or an unstable ordering fails the build when someone introduces
it, not the first time a phone gets a spurious cache miss. It also enforces the
Phase A budget of 16 assets and 256 KiB, separate from the permanent contract
ceiling.

build:release does not call build:desktop, so build:mobile-web is wired into
build:desktop, build:release, and build:release:parallel.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(packaging): fail the release when the mobile web bundle is missing or stale

electron-builder only warns about a missing input, so without a beforePack
guard a release ships an app that advertises the bundle capability and then
errors on every request. The hash check, not the existence check, is what
catches a half-written or stale out/.

The source tree is excluded from app.asar; out/mobile-web ships inside it under
the existing out rules, exactly as out/web does.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile-web): narrow the manifest with `in` instead of a cast

The changed-code casting gate rejects assertions, and `in` narrows the same
untrusted JSON without one.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile-web): move the bundle source under src/ so the root guard passes

.github/scripts/check-root-directory-entries.mjs blocks any new top-level entry
by name, so mobile-web/ could not live at the root.

The source is excluded from app.asar by the existing '!src{,/**/*}' rule; the
explicit '!src/mobile-web{,/**/*}' entry stays as a marker. out/mobile-web is
unaffected and still ships under the out rules like out/web. No tsconfig
includes src/**, so node, web, cli, and relay do not pick the tree up; it is
registered as a knip entry so audit:dead-code does not call it unused.

buildId is unchanged at 9d78435e: the builder hashes content, not paths.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(build): resolve the entry-script guard through pathToFileURL

`file://${process.argv[1]}` never equals import.meta.url on Windows, where that
url is file:///C:/... So the builder exited 0 having written nothing and the
Windows packaging job failed later, at the guard, with no clue why. Every other
script in config/scripts already uses pathToFileURL; this one now does too, via
an exported predicate a posix runner can exercise with a win32 path.

The verify script had no entry guard at all, so importing its budget constants
ran the whole verification — including its process.exit — inside the test
worker. It is now a function behind the same guard.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(ci): build the mobile web bundle in the PR package job

That job assembles packaging inputs step by step instead of calling
build:release, so the new beforePack guard hard-failed it.

The census test added here is the oracle: it walks every workflow job that
invokes electron-builder without --prepackaged (which short-circuits doPack
before beforePack) and requires a bundle-producing script in the same job. It
goes red on exactly pr.yml's package job when this step is removed. Ten jobs
covered; the other nine already ran build:release, build:release:parallel, or
build:desktop.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile-web): pin source line endings, because CRLF changes the buildId

Every text byte under src/mobile-web is hashed into an asset digest and from
there into buildId, so a CRLF checkout produces a different bundle id for the
same commit: 91af2897 instead of 9d78435e. That would make a Windows-built
desktop disagree with a mac-built one about which bundle a phone has cached.

.gitattributes pins eol=lf for the text sources and -text for the PNG, matching
the four trees already pinned for byte-hashing. The verify script asserts no
source file carries a CR, so the build fails if the pin ever stops applying
rather than silently shipping a second bundle identity.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(build): read the test's own path from import.meta.filename

oxlint unicorn/prefer-import-meta-properties.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(test): census packaging jobs over raw workflow text, not re-serialized YAML

yaml.stringify folds long lines, and in dev-channel-win-build.yml's build-win the
fold landed between `electron-builder` and `--config`, so a real packaging job was
invisible to the census: 11 jobs exist, the test saw 10. Slice each job's raw source
by its parsed boundaries instead, and pin the inventory so a new packaging workflow
has to be added here on purpose.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(build): assert the script chain the packaging census trusts

The census only checks that a packaging job invokes one of ten build scripts; that
those scripts still reach build:mobile-web was asserted nowhere, so a dropped link
would leave every job looking covered while packaging failed at beforePack. Resolve
each script for real, and pin pr.yml's hand-rolled step, since that job never calls
build:release.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(build): realpath the entry path before the direct-invocation compare

Node resolves symlinks in import.meta.url but not in argv[1], so `node /tmp/...`
against a /private/tmp realpath compared two different strings: the builder and the
verifier exited 0 having written and checked nothing. Same silent-success shape as
the Windows file:// bug, so the fix sits next to it, with both seams injectable.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(mobile-web): format bootstrap.css with oxfmt

It was the only tracked CSS failing oxfmt --check. The buildId is unchanged at
9d78435e8bb73c3341f833c20aaefbd7bfdfc414b68dadf87c1689d86728fe33, because esbuild's
CSS minifier normalises the whitespace this touches before the asset is hashed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(packaging): reject bundle files the manifest does not list

The guard only walked the manifest, so a dropped assets/stale.js passed: assets are
content-addressed, nothing ever overwrites a stale copy, and it would ship inside
asar unreachable and unverified. Require every file under out/mobile-web to be the
manifest or a listed asset.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(packaging): give beforePack an explicit mobile web bundle root

The bundle guard read the repo's out/mobile-web unconditionally, so the two
arch-aware packaging tests that call the real beforePack went red in the unit-test
job, which never runs build:mobile-web. beforePack now takes the bundle root as a
second parameter defaulting to out/mobile-web, which is what electron-builder gets,
and those tests build a real bundle into a temp dir instead. The guard is neither
skipped nor made tolerant of a missing bundle.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(packaging): census sees script-wrapped packers; dev verify reuses the guard

The workflow census only matched a literal `electron-builder --config` line, so
daemon-relocation-spike's `pnpm run build:unpack` (which packs and runs beforePack) was
invisible to it. Jobs now count when any `pnpm run <script>` they invoke chains to
electron-builder without --prepackaged; the spike joins the pinned list (12 jobs).

verify-mobile-web-bundle.mjs re-implemented a weaker subset of the packaging guard
(no safe-path check, no buildId recompute). It now calls assertMobileWebBundleBuilt, so a
manifest edited after the build fails at `pnpm build:mobile-web` exactly as at beforePack.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-17 22:41:13 -04:00
Neil ea01cd0ccd fix(windows): reject a node-pty addon that predates the MSYS breakaway denial (#20047)
* docs(windows): record the measured MSYS job-breakaway mechanism

The per-PTY job already denies JOB_OBJECT_LIMIT_BREAKAWAY_OK for Cygwin/MSYS
shells (#19068), but nothing records why, and a conpty.node built before that
commit fails windows-msys-job.win32.test.ts in a way that reads as a source
defect. Measured on a real Windows 11 host: both the plain and the exec-
replacement Git Bash shapes leak, the escape is the MSYS runtime's own
spawn/exec (fork keeps membership), and a single-variable A/B on
usesCygwinRuntime flips the result 0/2 -> 4/4.

Also names the gap the failure hid behind: node-pty-job-ownership.cjs asserts
symbol presence, which cannot distinguish patch revisions.

* fix(windows): reject a node-pty addon that predates the MSYS breakaway denial

The native-runtime gate asserted only that terminateJob, listJobProcessIds and
assignCurrentProcessToJob were exported. All three predate the Cygwin/MSYS
breakaway denial, so an addon built before it passes every gate,
isPtyJobOwnershipAvailable() returns true, and windows-pty-job.win32.test.ts
passes 6/6 -- while every Git Bash child is created outside its pane's job and
survives terminatePtyJob.

Read the resolved .node and require the wide msys-2.0.dll literal that
usesCygwinRuntime holds, the way stagedRelayAddonIsUnpatched() already tells a
patched windows-process-tree addon from a published one. An addon the caller
cannot name is refused rather than skipped: a gate that cannot see its subject
is not a gate.

Verified against real binaries on a Windows 11 host: the shared checkout's
pre-#19068 build errors, a build from current patched source passes, a missing
path errors.

Also closes the cross-host packaging skip. The export half has to load the
addon so it cannot run when the packaging host is not the target, which is how
a Windows release built elsewhere could ship this. The marker is a file read
and needs neither; an unrecognised layout warns rather than fails a release
that was packaging fine.

* fix(windows): check the MSYS breakaway denial on the rebuild path too

The Electron probe carried the marker check, but it lives inside
probeElectronNativeModules, which returns early whenever the Electron package
binary is unusable. Covered by another path is not this path checks -- and the
defect this whole change closes was a gate that looked like it checked.

Reading the binary needs neither a loadable Electron nor an executable target
arch, so assert it after the rebuild, beside the windows-process-tree
assertion that exists for the same reason: this is the addon copied into the
packaged app. Absent warns (a cross-platform rebuild need not leave a win32
addon on this disk); present and unmarked is fatal.

The fixtures now write a real addon file, because the gate reads the binary it
was told about rather than trusting the exports. Verified against the two real
binaries measured on the Windows host: the pre-#19068 build fails this path,
the build from current patched source passes.

* fix(windows): check the marker on every ConPTY path the packaged app can load

The packaged marker check read one hard-coded path, `build/Release/conpty.node`,
and warned when it was absent. `loadNativeModule` tries `build/Release`, then
`build/Debug`, then `prebuilds/win32-<arch>`, swallowing each failure, and
`prunePackagedNodePty` drops the published prebuild only when a same-arch
`build/Release` exists to replace it. So the two packages the check was added for
were the two it could not see:

- cross-host: no host but Windows can build conpty.node, so there is no
  `build/Release` and the prebuild is what ships. The check warned and returned.
- cross-arch: `build/Release` is the packaging host's own arch, patched and
  marked, so the check printed OK -- while the target app cannot load it and
  falls through to the unmarked prebuild underneath.

Measured, not assumed: both published Windows prebuilds in the node-pty tarball
contain neither `msys-2.0.dll` nor `cygwin1.dll` in any encoding. They are the
binary that leaks every MSYS pane child out of its job.

It now sweeps every candidate present for the *target* arch and refuses a package
with no candidate at all, which is a package with no ConPTY backend rather than a
layout to shrug at. It runs for every Windows slice instead of only the branch
the export check skips, so deleting the export check cannot silently take it too.
A stale source build keeps the rebuild advice; the prebuild gets the advice that
actually works, which is to package the slice on a Windows host of that arch.

Also: the marker constant was re-typed in four places and was tied to the C++
literal that produces it by nothing at all, so editing the patch would have left
a gate that fails every correctly rebuilt addon and tells the developer to do the
one thing that cannot help. The fixtures now take the constant from the gate, and
a test asserts the patch still adds `L"msys-2.0.dll"` to conpty.cc.

And the rebuild path treated a missing addon as a warning even on the host that
will run the install, where node-pty would fall through to that same prebuild.
The verdict is now a value, so it is tested without a platform gate.

* fix(windows): resolve the packaged ConPTY the way its loader does

Sweeping every candidate and demanding the marker on all of them was wrong in
the one case it was meant to make safe. `beforeBuild` runs
`rebuild-native-deps.mjs --platform=win32 --arch=<target>`, so a cross-arch slice
normally does get a patched `build/Release` for the target; `prunePackagedNodePty`
keeps the prebuild anyway because its guard is `electronArch === process.arch`
rather than the arch of the binary. That package is correct and its leftover
prebuild is never reached, and the sweep failed it -- telling whoever ran it to
package on a Windows arm64 host, which is both the wrong remedy and one no runner
here can offer.

Presence cannot separate that package from the one whose cross-arch rebuild
quietly emitted the host's architecture, because the only difference is the arch
of `build/Release`. So the gate now resolves the addon the way `loadNativeModule`
does -- first candidate whose PE `IMAGE_FILE_HEADER.Machine` matches the target,
walking root-then-lib for each layout in node-pty's own order -- and checks the
marker on the one that will actually run. A package with no candidate, or none of
the target's architecture, is refused: it has no ConPTY backend either way, and
the second is exactly what a silently host-arch cross-build looks like.

The PE machine reader already existed, privately, in the relay addon builder that
needed the same "a cross-build cannot silently emit host arch" guarantee. It is
now shared rather than copied.

Two seams were unreachable from anything but Windows, so nothing tested them:

- the afterPack hook's win32 block was an inline if/else that only a source-text
  assertion could inspect, and that assertion could not tell the difference
  between the check running and the check being wrapped in `try {} catch {}`. It
  is now `verifyPackagedWindowsNodePty`, and "the marker check runs even where
  the export check cannot" is four spied assertions instead of a string match.
- the rebuild path's verdict read `process` directly, so the branch that fires
  only on the host being rebuilt for was dead on every other host. It now takes
  the host as arguments, and the fs checks, the warning and the failure are all
  exercised from macOS.

Fixtures write a real PE header rather than `MZ fake addon`, since the gate now
reads one. The machine table is pinned to the documented IMAGE_FILE_MACHINE
values, because every fixture builds its header from that table and a table wrong
in both entries would otherwise agree with itself.

* fix(windows): say why the packaged ConPTY fell back, not just that it did

The previous commit resolved the addon by architecture but still had one message
for every way the resolution could land on the published prebuild. Those ways
want opposite remedies, and the one it printed was the remedy the commit before
it had just called wrong:

- no source build in the package at all — the slice has to be built somewhere
  that can build node-pty for the target arch.
- a source build that is there but is the packaging host's architecture, because
  the cross-arch rebuild did not honour `--arch` — re-running that rebuild is the
  fix, and "package on a Windows arm64 host" is neither necessary nor possible.

The second is the common one, since node-pty publishes a prebuild for both
Windows arches and prune keeps the target's on every cross-arch package. So the
old text fired mostly on the case it described least. It now reports which source
builds were skipped and the machine field each carried, and names the rebuild
command.

"Nothing the target can load" had the same problem in reverse: a zero-length or
truncated `conpty.node` got a cross-architecture diagnosis. Every candidate is
now named with what was actually read, including "not a PE image".

The rebuild path asserts the architecture too. A rebuild that ignored `--arch`
was otherwise only visible at packaging, two steps from the command that fixes
it. Arches with no known machine value are left unjudged rather than guessed at.

Two things the extraction broke or nearly broke, both found by mutation:

- the shared PE reader answers `null` where the relay builder's private copy
  returned a number, which would have turned its "node-gyp ignored --arch" error
  into a `TypeError`. Both callers now go through `describePeMachine`.
- the rebuild fixtures stage a script's co-located modules by walking its
  imports, and the walker only understood `from '...'` — so the gate's new
  `require('./windows-pe-machine.cjs')` was left behind and every subprocess test
  failed with a resolution error, which is the exact failure its own comment
  warns about. It now follows `require` and bare side-effect `import` as well,
  and has tests; the fixture stages the gate by walking it rather than by naming
  one file.

Fixtures write real PE headers through one shared builder instead of three
hand-rolled ones.

* fix(windows): run the node-pty addon gates on the Windows job that can

`rebuild-native-deps-node-pty.test.mjs` carries four `skipIf(platform !== 'win32')`
tests. The full suite runs on ubuntu, and the Windows PR job runs an explicit
file list that never named this file -- so those tests were skipped on Linux and
never reached anywhere else. Three of them predate this branch. The Windows job
is added the four node-pty addon suites plus the module-walker one; the comment
above that list already says why it is the right place, which is that the addon
assertions only hold once natives have been rebuilt. Running the path-joining
suites there also covers the separator this gate's candidate list is built from.

The rest is round-three review:

- the rebuild-time arch assertion told a reader "node-gyp did not honour --arch"
  about a file that was not a PE image at all, which is a truncated or
  quarantined artifact and a different command to run. The two now read
  differently, and neither claims the other's cause. Same fix the packaged gate
  had one commit ago, in the place that had not had it yet.
- the missing-addon error said node-pty "would load" a prebuild without checking
  it is there. It says "fall through to" now, which is true either way.
- `isLoadableByArch` had no caller left once the packaged gate started needing
  the raw machine field for its message. Removed rather than kept warm.
- each candidate's header is read once instead of up to three times.
- the module walker's comment claimed every shape that reaches a co-located
  module; it does not follow `projectRequire`/`requireLocal`, and it must not --
  those specifiers resolve against the project root, so following one stages the
  wrong path and the copy fails. Proven by trying: widening the pattern to
  require-shaped names broke nine tests on
  `projectRequire('./config/scripts/...')`. The comment now says what it follows
  and why it stops there.
- a new test resolved a file URL with `.pathname`, which keeps the drive-letter
  slash on Windows -- the very job this commit adds it to.

* docs(windows): put the superseded export-only gate in the past tense

It describes what used to pass a broken addon, so present tense reads as a
description of the gate the same document then explains replacing it.

* fix(windows): repair what running the node-pty suites on Windows exposed

Putting these files on the Windows job turned four assertions red on the first
run. Three of them were in tests that carried `skipIf(platform !== 'win32')` and
had therefore never executed anywhere, on any branch.

- `writeFakeElectronRebuild` emitted the `windows-process-tree` addon a real
  rebuild leaves but never node-pty's, so every Windows test of the rebuild path
  ran against a tree no real rebuild can produce: node-pty "rebuilt" with nothing
  in `build/Release`. The new same-host check reads that state correctly and said
  so. The fake rebuild now writes `build/Release/conpty.node` when it was asked
  to rebuild node-pty for win32, with the marker and the target machine.
- `mkTempProject` never staged `windows-process-tree-creation-time.cjs`. The
  rebuild script reaches it through `projectRequire`, which resolves against the
  project root, so the module walker cannot follow it and must not try. Staged by
  name, with a comment saying which of the two it is. Without it the
  windows-process-tree probe failed to load its own checker and the module joined
  `modulesToRebuild`, which is the second and third red assertion.
- the two `nodePtyAddonPath` cases compared against a literal POSIX string.
  `resolve` returns a drive letter and backslashes on Windows, so they could only
  ever pass off it. Built from segments now, which still pins the `..` traversal
  that is the point of the test.

Verified on macOS: ensure-native-runtime-job-ownership,
verify-packaged-node-pty-job-ownership, windows-pe-machine,
script-module-dependencies, rebuild-native-deps-node-pty, rebuild-native-deps,
rebuild-native-deps-windows-process-tree, ensure-native-runtime -- 109 passed, 6
skipped. The 6 are the Windows-gated rebuild tests, which is the job this change
is aimed at; Windows CI is the arbiter.

* fix(windows): give the packaged fallback a third verdict, for a file that is no image

The packaged gate had two remedies for landing on the published prebuild and
picked between them on `!prebuilt`, which puts a truncated, empty or quarantined
`build/Release/conpty.node` in the cross-arch bucket: "the source build beside it
is the wrong architecture ... re-run with --arch". It is not the wrong
architecture, it is not an architecture, and `--arch` is not the command. The
rebuild-path gate was split for exactly this a commit ago; this is the same split
in the place that had not had it.

Also from review of the settled state:

- the stale-source-build branch ended in a call that happened to throw, so a
  reader could not see it was terminal and the file was read twice to get there.
  The verdict is now an Error the caller throws, built once from the read it
  already did, and shared with `assertCygwinBreakawayDenied` rather than copied.
- four injection seams had no consumer in production or in tests
  (`deniesBreakaway`, `peMachine`, and `exists`/`peMachine` on the rebuild
  verdict). An unused seam is a way for the tested path and the real one to drift
  apart; the tests drive both with real files. Removed.
- the loader table existed in a docblock and in the reference doc, already
  disagreeing about row four. The docblock cites the doc now.
- `peImage` stamped machine `0x0000` for an arch it had no value for, because
  `writeUInt16LE(undefined)` coerces to zero. A fixture that quietly invents the
  field the gates read is the same species of silent lie the gates exist to
  catch; it throws, and a test holds it to that.
- a test named for refusing an unreadable candidate asserted only that something
  threw. Renamed to what it proves.

* fix(windows): make the rebuild fixtures represent a tree that can exist

Second round of what running these suites on Windows exposed. The module the
walker could not stage is now staged, so the probe reached its own checker and
the real reasons surfaced:

- `writeFakeWindowsProcessTree` exported `{}`. The creation-time gate reads
  `supportedProcessDataFlags` off the addon and calls its absence "the tarball
  prebuilt, not a build of the patched source" — correctly. The fixture predates
  that gate and, being Windows-only, never met it. The healthy fake now reports
  the flag, taken from the gate's own constant. Two tests were failing on this,
  the second only because the module then joined `modulesToRebuild`.
- `rebuilds a loadable ConPTY native that lacks Orca job ownership` asked for a
  node-pty rebuild in a tree where node-pty had none of the payload its package
  ships. It gets `writeFakeNodePtyConptyPayload` like its two siblings.

I also tried making the fake rebuild emit `build/Release/conpty.node` the way a
real one does, and backed it out: `restoreNodePtyWindowsConptyRuntime` keys off
that file and then reads `third_party/conpty`, so emitting it in a tree without
the package payload turns one honest gap into an ENOENT two steps away. The
payload fixture is where "node-pty has its addon" belongs.

macOS: ensure-native-runtime-job-ownership, verify-packaged-node-pty-job-ownership,
windows-pe-machine, script-module-dependencies, rebuild-native-deps-node-pty,
rebuild-native-deps, rebuild-native-deps-windows-process-tree,
ensure-native-runtime — 112 passed, 6 skipped. The 6 are the Windows-gated
rebuild tests; Windows CI is the arbiter and is why they are on that job now.

* fix(windows): register the node-pty addon suites in the scope list too

Putting the five suites in the Windows lane's vitest argv gets them run once the
job starts; `WINDOWS_PACKAGE_TESTS` in `pr-code-change-scope.mjs` is what decides
whether the job starts at all. Only the argv was updated, so a PR touching just
`rebuild-native-deps-node-pty.test.mjs` would not have started the Windows job,
and its four Windows-only cases — including the same-host-absent one added here —
would have run on no machine for that PR. Exactly the shape of gap this branch is
about. Both lists now name all five, and `windows-pe-machine`,
`windows-pe-image-fixture` and `script-module-dependencies` join
`NATIVE_RUNTIME_PREFIXES` so a change to the modules themselves starts it too.

`win32-test-lane-registration.test.mjs` exists to catch precisely this and did
not, because its matcher only recognises suite-level gates (`describe.runIf` /
`describe.skipIf`) and a `.win32.` filename. These tests gate per `it`. Widening
it is not this branch's change to make: about thirty files across the repo carry
per-`it` Windows gates and are unregistered, so the ratchet would move far beyond
node-pty. Flagged rather than done.

Message repairs from the same review:

- the non-PE arm of the rebuild-time arch error read "... is not a PE image, so
  nothing can load it, so node-pty would fall back ...". The shared consequence
  clause already opens with ", so".
- the no-source-build packaging error ended "Package this Windows slice on such a
  host", which is wrong advice for the case where the host IS such a host and the
  rebuild simply left nothing — reachable when the artifact is removed before
  prune runs. It now names both readings and points at the beforeBuild output.
- the relay-addon builder blamed `--arch` for a build output that is not a PE at
  all, the same guess the node-pty gate was taught to stop making.
- the patch-drift assertion was a bare `toBe(true)`, so a real drift read as
  "expected false to be true". It now names the two things that can have drifted
  and what happens until they agree.
2026-09-16 22:23:30 -07:00
Neil d62328aa4d fix(codex): remove redundant Windows hook launcher for Unicode profiles (#20952)
* fix(codex): reuse the Windows hook shell for Unicode profile paths

* test(codex): register Unicode hook tests in Windows CI

* test(codex): pin trust hash replacement during Windows upgrade

* test(codex): retry transient Windows teardown locks
2026-09-16 00:30:52 -07:00
Neil 11180fa532 chore(lint): add anti-slop oxlint plugin (pinned, all rules off) (#20726)
* chore(lint): add anti-slop oxlint plugin (all rules off)

Vendors dmmulroy/anti-slop (MIT) plus no-call-only-assertions and
no-pass-through-type-alias from maharshi365/deslop (MIT). Every rule starts
"off"; each follow-up PR fixes one rule's violations and flips it to "error".

* fix(lint): actually exclude the vendored plugin from the anti-slop audit

oxlint does not honour ignorePatterns supplied via --config, so the
config/oxlint-plugins/anti-slop/** entry never matched and the vendored rule
source was being linted as first-party code (505 violations). Move the exclusion
to the --ignore-pattern CLI flag in audit:anti-slop, which does work, and drop
the entry that gave a false sense of coverage.

Keeping vendored source unlinted matters because anti-slop is updated by
three-way merge against the upstream snapshot; reformatting it locally would
conflict on every update.

* chore(lint): pin anti-slop instead of vendoring it; drop deslop

Replaces the ~5k vendored lines with a git-pinned devDependency:
  oxlint-plugin-anti-slop: github:dmmulroy/anti-slop#c44ef22

anti-slop ships raw .ts with no build step, and Node refuses to type-strip
anything under node_modules (ERR_UNSUPPORTED_NODE_MODULES_TYPE_STRIPPING), so
oxlint cannot load it from there -- which is why upstream says to vendor it. A
postinstall step copies the pinned package's source to .anti-slop-plugin/
(gitignored), which Node will type-strip because it sits outside node_modules.
Upgrading is now a SHA bump rather than a re-vendor and three-way merge.

Verified byte-identical rule output to the vendored copy across all 16 rules
that fire.

Drops maharshi365/deslop and its two rules (no-call-only-assertions,
no-pass-through-type-alias). It is not on npm either, so it would need a second
git pin and copy step, and it is a 5-star single-maintainer repo that is itself
a re-namespaced copy of anti-slop. One upstream is enough.

* ci(lint): run audit:anti-slop in PR CI

config/scripts/pr-workflow-lint-parity.test.mjs requires every step in
`pnpm lint` to have a matching step in .github/workflows/pr.yml; adding
audit:anti-slop to lint without the workflow step failed that ratchet.

Also makes audit:anti-slop sync the plugin itself before linting. The generated
.anti-slop-plugin/ directory is gitignored and otherwise only created by
postinstall, so a cached install that skips postinstall would leave oxlint
unable to load the plugin.
2026-09-14 21:42:37 -07:00
Neil b61a2347b9 feat(design-system): gate renderer UI with @shadcn/lint (#20731)
* feat(design-system): gate renderer UI with @shadcn/lint

Wires shadcn-ui/lint's Oxlint plugin into the two places this repo already
ratchets: the changed-lines PR gate for rules the renderer can't satisfy
today, and `pnpm lint` for the one that is already at zero.

- config/oxlint-design-system.json: no-restyle (layout allowed),
  no-raw-colors, require-static-classes -- scoped to src/renderer/**/*.tsx,
  run over added lines only. Measured at 10 findings across the last 60
  commits (771 changed files), so it holds the line without a migration.
- config/oxlint-dead-classes.json: no-unknown-classes repo-wide, with the
  renderer's plain-CSS hook namespaces allow-listed. Now at zero.
- no-inline-styles and no-arbitrary-values stay off; STYLEGUIDE says why.

Fixes the three live bugs the linter found:

- `--editor-surface` never reached `@theme inline`, so `bg-editor-surface`
  generated no CSS -- 12 editor/artifact/notebook panes fell through to the
  page background instead of #1e1e1e in dark mode.
- `scrollbar-none` is not a Tailwind utility and was declared nowhere, so
  the remote file browser breadcrumbs showed the scrollbar they meant to
  hide. Declared as a real `@utility`.
- Notebook markdown cells used `markdown-preview-body`, which no stylesheet
  defines; the styled class is `markdown-body`. They rendered unstyled.

* ci: run the dead-class gate in PR CI

`pnpm lint` gained check:dead-classes, and pr-workflow-lint-parity requires
every `pnpm lint` step to have a matching step in pr.yml.

* fix(notebook): keep markdown theme selectors working
2026-09-14 17:52:21 -07:00
Neil 20794ee785 ci: keep the baseline build off the compatibility matrix lanes (#20733)
The compatibility gate started the pinned 2.25.5 source build inside the same
step that runs the three measured lanes, so `make -j$(nproc)` competed with two
container lanes whose wall clock is container starts, not Git. A boundary case
that costs ~1.5s stretched past Vitest's 30s timeout and failed the job.

Build the binary in its own step before the matrix, and pull both images before
any lane starts so a lazy pull cannot stall whichever test its sibling is timing.
2026-09-14 17:32:33 -07:00
Neil e86cba888b build: reduce native dependency installs to the host platform (#20420)
* Reduce native dependency installs to the host platform

* Remove install policy documentation

* Guard cross-arch packaging and scope release installs to the runner

electron-builder only logs a warning for a missing extraResources source,
so a host-only install silently shipped a foreign-arch slice without its
natives — `pnpm build:mac` on Apple Silicon produced an x64 DMG with no
sherpa-onnx-darwin-x64 and no @parcel/watcher-darwin-x64. The previous
beforePack hook covered only win32.

- Add assertPackagedNativeVariantsInstalled, an arch-aware check over the
  target's sherpa-onnx, @parcel/watcher, and (on Windows) node-gyp addons.
  beforePack now runs it for every platform, with remedies split: another
  architecture comes from install:release, the os:win32 addons need a
  Windows host.
- Drop --os from the release installs. Every packaging job already runs on
  a runner whose OS matches its target, so only the macOS lanes need extra
  breadth, and only on CPU for their x64+arm64 config. Windows and Linux
  packaging return to a plain host-only install.
- Add --frozen-lockfile to install:release so a bare run cannot rewrite
  the lockfile.
- Restore the install policy reference doc and the CONTRIBUTING note, plus
  the rationale comments dropped from the runtime contract test.
- Gate the packaging-closure assertions on whether the Windows addons are
  installed rather than on the host OS, so a cross-arch install exercises
  them off Windows too.
- Make the workflow contract test read `run:` steps as well as retry-action
  commands, and enforce host-only scoping on the non-macOS packaging lanes.
- Remove the unreferenced install measurement script; its numbers live in
  the policy doc.

* Track the install policy doc and index it from AGENTS.md

docs/** is ignored behind a per-file allow-list, so the new reference doc
was only committed via git add -f and future edits would be skipped. Add
it to the allow-list and give it an AGENTS.md entry like every other
tracked reference doc, so the host-only install rule is discoverable
before someone packages a second architecture.

* Route Windows-lane removals through the retrying helper

Adding these four specs to the PR Windows lane pulled them into the
windows-lane-tree-removal-boundary ratchet, which failed on 20 raw
recursive removals. On Windows a bare rmSync races a handle the OS has
not released, throwing EPERM after the assertions already passed and
reporting a green test as a lane failure.

* Adapt the packaging guard to the vendored Windows registry addon

main vendored windows-native-registry as the workspace package
@orca/windows-registry (#20438). A workspace link resolves on every
host, so including it in the installed-Windows-addons checks proved
nothing. @vscode/windows-process-tree is the only os: win32 npm addon
left, so it alone decides whether the win32 resource plan resolves.
2026-09-12 21:25:03 -07:00
Neil bd0f8826ea fix(ci): match the truncated windows-process-tree virtual store dir (#20447)
On Windows, pnpm shortens the virtual store directory to
@vscode+windows-process-tre_<hash>, cutting into the package name before
the @, so the @vscode+windows-process-tree@* glob matched nothing and the
addon recompiled on every Windows job. node-pty escapes this because its
truncation lands after node-pty@, which the glob still matches.

Widening the prefix to @vscode+windows-process-tre* matches both the full
name kept on macOS/Linux and the truncated Windows one.
2026-09-12 21:11:39 -07:00
Neil 411843f633 fix(ci): cache the vendored addon where node-gyp actually writes it (#20445)
The workspace link means pnpm never creates a .pnpm/@orca+windows-registry@*
entry, so all four native-cache blocks globbed a path that cannot exist and
the addon was recompiled on every Windows job.

Also hardens the addon itself: RegEnumValueW reports a byte count and the
registry does not enforce whole WCHARs for string types, so an odd count let
Napi's auto-length scan run past the value; and a value named __proto__ would
reassign the result object's prototype instead of becoming an entry.
2026-09-12 20:45:42 -07:00
Neil 5127d1eb3b refactor(windows): vendor the registry addon as @orca/windows-registry (#20438)
* refactor(windows): vendor the registry addon as @orca/windows-registry

windows-native-registry@3.2.2 was last published in 2023 by a single
maintainer. Orca called two of its exports, both read-only, so the whole
dependency is replaced by a local N-API addon under native/.

The vendored addon is read-only by construction: setValue, createKey and
deleteKey are gone, so RegDeleteTreeW no longer ships in the app. Two
upstream defects are also fixed rather than carried over — the name/data
scratch buffers were file-scope statics that concurrent reads would
scribble over, and createKey/deleteKey called .c_str() on a temporary.

Build wiring keeps the existing shape: still an optionalDependency gated
to win32, still excluded from pnpm's allowBuilds so only Orca's own
Windows rebuild runs node-gyp for it, still copied into the packaged
resources. The CI native caches now key on the vendored sources so an
addon.cc edit cannot restore a stale .node.

* test(windows): check the vendored registry addon against reg.exe

The addon is vendored source, so no upstream release proves it still
decodes values the way Orca's PATH readers expect. reg.exe is the only
independent oracle on the box.

* ci(windows): register the registry addon test on the Windows runner

A Windows-gated file self-skips on ubuntu, so without both registrations
it reports success while running on no machine at all.

* fix(build): link the registry addon as a workspace package, not file:

As a `file:` dependency pnpm re-resolved and re-linked the package on
every install, including `--frozen-lockfile` (measured: "added 1" on a
repeat no-op install). That virtual-store churn ran concurrently with
node-gyp reading the same tree and cost @vscode/windows-process-tree its
binding.gyp mid-rebuild, failing package (windows) whenever the native
cache hit and only that module needed building. The linux packaging job
hit the same race from the other side, as a pnpm staging move failure.

A workspace link resolves once and leaves the store alone; repeat
installs are now 55ms no-ops. native/windows-registry is listed
explicitly so `packages:` still does not auto-discover mobile/.

* fix(build): stop tracking node-gyp output for the vendored addon

The build/ tree is generated per host and ABI; the committed copy was
macOS-specific gyp scaffolding from a local build and would have shipped
stale Makefiles to every checkout.

* chore: ignore the vendored addon's node-gyp bin output too

node-gyp also emits bin/<platform>-<abi>/ beside build/; both are per-host
generated output that must never be committed.
2026-09-12 20:10:49 -07:00
Neil 7b0701aefa chore: remove 20.7 MiB of duplicate and unused media (#20416)
* chore: remove duplicate and unused documentation media

* chore: guard README local links and refresh tile-01 vendor metadata

- Add config/scripts/check-readme-local-links.mjs: every local src/srcset/href
  in README.md and docs/readme/*.md must resolve to a tracked file. Runs in the
  ungated root_directory_guard job so docs-only diffs (which skip static_analysis)
  still catch a deleted docs-site or feature-wall asset the README embeds.
- Refresh tile-01.recorded-at.json to what vendor-feature-wall-assets.mjs now
  emits for the tab-split source path.
- Drop the pr-19217 evidence prose that cited the removed screenshots.

* fix: accept single-quoted attributes in README local link check

The parser only matched double-quoted src/srcset/href, so <img src='missing.gif'>
was skipped and the guard passed a README that GitHub renders with a broken image.
Regression test fails without the parser change.
2026-09-12 19:40:18 -07:00
Neil 4aa9329e99 ci: narrow pnpm cache keys and shallow development checkouts (#20370)
* ci: cache dependency downloads and shallow development checkouts

* ci: defer mobile caches after measuring restore overhead

* ci: retain existing release signing cache behavior
2026-09-12 01:31:04 -07:00
Neil de8e5421ff ci: reuse immutable package setup across shutdown checks (#20368) 2026-09-12 01:23:05 -07:00
Jinwoo Hong c84007c541 feat(rpc): generate a shared params catalog from the host registry, gated on parse parity (#19961) 2026-09-10 21:18:39 -07:00
Neil 314506003a fix: retain MSYS shell descendants in their terminal job (#19068)
* fix: retain MSYS shell descendants in their terminal job

* test: complete MSYS regression CI registration and teardown contract

* fix(windows): deny job breakaway for the whole Cygwin/MSYS shell family

The per-PTY job probed only msys-2.0.dll, and only for bash.exe/sh.exe.
Cygwin ships the same spawn.cc breakaway logic under cygwin1.dll, and an
MSYS2 zsh escapes exactly like its bash does, so both kept the orphan bug.

Probe the runtime DLL on the shell's own search path instead of matching
shell names: that is the property that decides whether the runtime will
ask for CREATE_BREAKAWAY_FROM_JOB, and it drops the name special-casing.

* chore(patch): restore the conpty.cc index line

The earlier hand-edit dropped it while every sibling section kept one.
Recomputed against the real blobs: applying this patch to 7b286d3d
yields exactly 4b06d185, so git apply -3 has its fallback back.
2026-09-07 00:35:55 -07:00
Brennan BensonandMerge Sim 1478101342 fix(windows): unblock structured native chat by exposing process creation time (#18986)
* fix(windows): guard process creation times

* fix(windows): ask the relay's bare addon for creation times too

The relay addon build now emits creationTimeMs, but the runtime binding
for the bare addon still declared only CommandLine, so a Windows relay
host requested flag 2 and every row came back without a creation time.
That leaves captureWindowsDescendantSnapshot returning null and
verifyWindowsProcessIdentity false forever on those hosts -- the relay
half of the patch was unreachable.

Naming CreationTime in the adapter is safe because the bare addon is a
content-hashed relay artifact: it ships in the same immutable relay
directory as the bundle reading it, so it can never be older than the
code asking for the bit.

Also bound the win32 guard test on our own row, which the addon can
never fail to answer, so an unconverted FILETIME or a 1601-epoch stamp
fails instead of satisfying a bare count.

* fix(windows): make the compiled addon prove its own CreationTime support

CI caught the real defect: the win32 guard test read
isWindowsProcessStartTimeAvailable() as true and then found 0 rows
carrying creationTimeMs. Unlike node-pty, this package publishes a
prebuilt .node at the same build/Release path node-gyp writes to, so
pnpm patches the source tree and leaves that binary alone. A host then
holds a patched lib/index.js -- ProcessDataFlag.CreationTime and all --
over a binary that ignores flag 4, and neither a load check nor a path
check can see the difference.

So the binary now says so itself: addon.cc exports
supportedProcessDataFlags, lib/index.js re-exports it, and

  - windows-process-tree-creation-time.cjs asserts it during install,
    which is what forces a from-source rebuild. It is shared by the Node
    probe in ensure-native-runtime.mjs and the Electron probe in
    rebuild-native-deps.mjs, exactly as node-pty-job-ownership.cjs is --
    the Electron half matters because that probe decides onlyModules, so
    without it the packaged app would ship the stale prebuilt.
  - isWindowsProcessStartTimeAvailable() gates on the reported bit, not
    the enum. Believing the enum is worse than reporting false: the
    descendant snapshot returns null forever and the exit proof latches
    unverifiable while structured chat believes it has a reaper.

rebuildNodeRuntimeModules could not actually have rebuilt this package:
the patched binding.gyp includes deps/node-addon-api, which the tarball
does not ship, and node-gyp must run from the physical dir.

Also closes the relay repair path's divergence: repairCreationTimeSources
wrote the C++ but not the buildNode splat or the tree-node typing, and
assertPatchApplied checked neither, so a repaired tree passed as patched
with buildProcessTree silently dropping the field.

The guard test is unchanged.

* fix(windows): keep the process-tree patch LF-only

windows-process-tree-patch-contract.test.mjs requires the patch file to
carry no CR bytes. Regenerating through pnpm patch-commit emitted 199 of
them, because the creation-time change is the first to touch files the
package ships as CRLF (src/process.h, src/process_worker.cc,
src/addon.cc, lib/index.js, lib/index.ts, the typings) -- and #17886's
own hunks over binding.gyp and src/process_commandline.cc carry the rest.

Stripping them is safe and changes nothing the lockfile records: pnpm
hashes patches CRLF-normalized, so the digest stays
e66202cc623996d02040c93449eb9ae353fddadf426cb53202a59ee710ee6fe7 and now
equals the file's plain sha256 too. It also still applies -- verified
against a deleted store entry, not a warm one -- and the precedent was
already there: the previous patch was LF-only and had been patching
those same CRLF files all along.

ensure-native-runtime.test.mjs stages the siblings the script loads at
module scope into its temp project. The import walk added by #17886 sees
`from './x.mjs'` only, so the createRequire'd .cjs siblings still have to
be named, and this PR adds a second one.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 13:59:59 -07:00
Neil f952f1ac96 test: run real WSL terminal launch and paste in PR CI (#19072)
* test: continuously exercise real WSL terminal launch and paste

* test: establish live WSL reader before changing default shell

* ci: pin WSL kernel installer and participation selectors

* ci: route deleted WSL paths and record immutable run evidence

* test: require exactly three WSL repetitions in lane contract
2026-09-06 05:24:54 -07:00
OrcaWinandOrca Worker c252d855ac fix(windows): resolve npm/pnpm .cmd shims past cmd.exe (#17869)
* fix(windows): resolve npm/pnpm .cmd shims past cmd.exe

A `.cmd` target forces every spawn through `cmd.exe /c` with each argument
caret-escaped, and Microsoft Defender for Endpoint scores a long `cmd.exe /c`
line carrying caret-escaped natural language as obfuscation. `codex.cmd` is
named in the spawn cluster of the MDE incident this addresses.

npm's `cmd-shim` and pnpm's `@zkochan/cmd-shim` generate files whose whole body
is "find node, run this script". Read one, and the spawn can go straight to
`node.exe <script> <args>` — no cmd.exe, no caret escaping. Anything the parser
does not recognise exactly, or whose target cannot be confirmed on disk, keeps
the existing cmd.exe path.

Incidentally fixes a real bug: cmd ends its command at a raw CR/LF whatever the
quote state, so a multi-line agent prompt through a `.cmd` shim had to be
rejected. Resolved shims have no such limit.

* fix(windows): refuse drive-relative shim paths and run the win32 tests in CI

Two blocking findings from review.

A drive-relative path defeated the absolute-path guard:
`win32.isAbsolute('D:evil.js')` is false, but `win32.resolve` reads the drive
letter and lands on `D:\evil.js`, outside the shim directory. cmd would have
built `C:\shim\D:evil.js` and failed; we would have executed the wrong file.
Adding `:` to the unsafe-character set closes it, and the alternate-data-stream
spelling `a.js:zone` with it. It costs no coverage: 84 of the 91 real shims on
this box still resolve, the same seven fall back.

Neither `windows-cmd-shim-resolution.test.ts` nor its `.win32` sibling was in
the Windows package job's file list, so the whole filesystem/resolution half and
the real-spawn equivalence suite ran nowhere. Both are now in
`WINDOWS_PACKAGE_TESTS` and in the pr.yml step.

Also from review: clear `windowsVerbatimArguments` explicitly on the resolved
branch rather than inheriting it, since there is no caller-built command line
there; document the kill switch and the PTY/hook-wrapper scope limits in
docs/reference; and cover drive-relative, BOM, line-ending, casing and `%*`
tampering in the platform-independent half of the tests.

* docs(windows): justify the shim-path colon guard from the filesystem rule

The guard was argued empirically ("none of the 91 shims on this box has one"),
which invites a future reader to relax it for a shim we have not seen. Windows
reserves `:` within a path segment, so a relative path cannot carry one at all:
the only spellings that can are drive-qualified, an alternate data stream, or a
`\?\` device path, and the last is already refused as absolute. That makes a
false refusal impossible rather than unobserved.

* refactor(child-process): move resolveSpawn into its own module

The merge with main pushed run-process.ts one line past the 300-line cap:
both sides grew it. The spawn-argv decision is already a pure, separately
tested unit, so it moves out rather than the cap moving up. run-process.ts
re-exports it, so no caller changes.

* perf(child-process): cache the shim interpreter lookup

The parse cache spared the shim read but not the PATH walk, so a second
resolution of the same .cmd did 0 reads and one statSync per PATH entry --
30 on a 30-entry PATH, synchronous on resolveSpawn, where one dead network
mount blocks the calling thread on every spawn.

Keyed by shim directory AND PATH, since the shim's own rule is
%~dp0\node.exe first then PATH, and a PATH edit between spawns must miss.
Corrects the stat comment, which accounted only for the shim itself.

* fix(child-process): revalidate a cached shim interpreter before using it

The node cache was held for process life and never rechecked, so a cached
node.exe that was later uninstalled -- or dropped from PATH by a version
manager -- was still handed to resolveSpawn, failing the spawn with ENOENT.
An uncached process in the same state returns null and falls back to
cmd.exe successfully, so the cache was strictly worse than no cache.

One statSync on a non-null hit, not one per PATH entry, so the walk this
cache exists to skip is still skipped. The stale-null direction stays
uncorrected on purpose: it only keeps the working cmd.exe fallback. Both
directions are now stated in the comment, along with the known miss for
callers that vary PATH per spawn.

* fix(child-process): honour PATHEXT when resolving the shim interpreter

The doc claimed a node.com/.bat/.cmd on PATH returned null and fell back to
cmd.exe. The scan actually skipped those entries and kept looking for a
node.exe, so PATH=C:\A;C:\B with C:\A\node.com and C:\B\node.exe resolved to
B's node.exe while the shim runs A's node.com -- a different binary, chosen
silently, on the one axis this module must not get wrong.

The scan now follows cmd's rule: first PATH directory holding any PATHEXT
spelling wins, PATHEXT order decides within it, and only an .exe winner is
returned. Anything else gives up and keeps the cmd.exe path, which restores
the strict-subset-of-cmd property everywhere except the documented cwd case.

PATHEXT is read from the child's env and joined into the cache key, since it
now changes the answer. Costs one stat per PATHEXT entry per node-less
directory, paid once per process behind the cache.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
2026-09-05 21:33:16 -07:00
0cbb01ef4b fix(security): apply the Windows path-hardening ACL that never ran (#17884)
* fix(security): apply the Windows path-hardening ACL that never ran

`buildWindowsRestrictAclArgs` invoked the hardening script as
`powershell.exe -Command <script> <path> <sid> <isDir>`. `-Command` does
not populate `$args`; it appends the trailing tokens to the command text.
The script therefore read `$args[1]` as `$null`, threw `NullArrayIndex` at
`$allowedSids[$sidText] = $true` under `$ErrorActionPreference = 'Stop'`,
and exited 1. Both callers swallowed that: the async callback was empty and
`applySecurePathRestriction` returned `true` regardless, while the sync
`catch` returned `false` and nobody logged. Every Windows secure path has
been left on its inherited ACL since the ACL was introduced (#5006), and
nothing said so.

Replace PowerShell with `icacls.exe`, which takes plain argv. That removes
the quoting surface entirely rather than escaping it: interpolating a path
into the command text would have turned a dead no-op into arbitrary
PowerShell on a filesystem path, since `-Command` executes what it appends.
It also drops the execution-policy dependency and the `powershell.exe`
spawn an EDR flags, and runs ~25x faster than the PowerShell cold start.

Hardening is now three passes: `/reset` to purge explicit ACEs that
`/inheritance:r` leaves behind, `/inheritance:r` plus a `/grant:r` per
allowed SID, then a read-back that checks the DACL is protected and grants
only the intended rights. The predecessor's verification block was equally
dead, and an apply that is never read back is only half a control.

Failures stay non-fatal — non-NTFS volumes, network paths and restricted
tokens fail legitimately and must not break startup — but they are no
longer invisible: every failure is logged, and a failed async apply now
evicts its cache entry so the next call retries instead of trusting a
success that never happened.

Routing through `runProcess`/`runProcessSync` also retires this file's
`node:child_process` allowlist entry.

* fix(security): verify the hardened ACL by identity, not by shape

Review found the bug class this PR fixes surviving inside the fix. The
verify pass checked rule count, absence of the inherited marker, and exact
rights — never *who* the rules named. Granting Everyone full control
satisfies all three, so hardening reported success on a DACL that handed
the credential to every local account, and most of the real-filesystem
tests still passed.

Verification now reads the descriptor back with `icacls /save`, which emits
SDDL with raw SIDs, and compares the principal set exactly. That is also
locale-independent by construction: the previous parse read localized
account names out of icacls' OEM-codepage stdout, where a non-ASCII path
survived by accident rather than by the documented mechanism. SDDL parsing
moves to `windows-security-descriptor.ts`.

Two further self-inflicted problems, both measured:

The post-rename re-harden led with `/reset`, which re-widened a DACL that
was already correct — the staged file's protected DACL survives the rename,
so the pass had nothing to do but open a window. Polling an external
process during a write into a relocated root caught it: the e2ee keypair
dropped to `BUILTIN\Users:(RX)` plus `Authenticated Users:(M)` — read *and*
write — before tightening again. Hardening now verifies first and returns
early when the DACL already reads back correct, which closes the window and
cuts the steady state from three spawns to one. Re-measured: 158 samples,
one DACL state, zero broad.

Evicting the cache on every failed async apply reintroduced #4901. The env
store re-hardens on the read path at ~2/s, so on a host where hardening
cannot work (FAT32, network path, restricted token) that was two icacls
spawns and two warnings a second, forever. Async retries now take a retry
floor and a hard per-path attempt cap. The write path keeps retrying
unthrottled — it is user-driven, and a failed credential ACL must still be
retried on the next write.

Also: failures route through a reporter hook that the main process points
at the diagnostic tracer, because `console.warn` reaches nothing in a
packaged GUI-subsystem build; `writeSecureFile` returns whether hardening
took, and the async branch reports `pending` rather than claiming `applied`;
a transient `whoami` failure no longer disables hardening for the process
lifetime, and the SID is shape-validated; the `/c` guard now covers the
synchronous runner too.

* fix(security): re-probe hardening instead of latching a transient failure

The per-process attempt cap added for the read-path storm was a permanent
latch: one AV scan, momentary lock or %TEMP% blip and every later credential
write in that session went unhardened, silently, on a host where hardening
would now succeed. Same defect class as #17858's computer-use host, and
worse here because what stops happening is security hardening on credential
files and nothing said so.

The retry budget now bounds the *rate*, not the lifetime: at most three
attempts per path per minute, re-probing in every later window, forever. The
transition is announced in both directions — `throttled` once per window on
entry, `recovered` when a rate-limited path hardens again — so a host stuck
in the degraded state is diagnosable rather than merely quiet. The reporter
type covers both, and the main process ends the `recovered` span
successfully rather than failing it.

Extracted to secure-path-hardening-retry-budget.ts, which keeps
secure-file.ts under its line cap without a max-lines disable.

Also confirms the second flagged risk rather than assuming it: a real
unwritable %TEMP% is now covered by a test proving verification fails
closed, reports at the `verify` stage, and still leaves the ACL applied —
so that path loses proof, not protection, and with the lifetime cap gone it
can no longer combine into a permanent-off state.

* fix(security): verify a directory's whole inheritance flag set

The flag check tested only that `OI` was present — never that `CI` was, nor
that nothing else was. That was harmless while `/reset` + `/grant` ran on
every pass and repaired whatever was there. The verify-first short-circuit
made it load-bearing: what verification accepts is now left alone, so a
latent under-check went live because a different fix started depending on
it.

Two directory DACLs passed while being wrong — both protected, three
non-inherited full-control rules, correct SIDs, differing from correct only
in their flags:

  (OI)(F)        - no CI, so subdirectories are left unprotected
  (OI)(CI)(IO)   - inherit-only, so the directory object itself grants
                   nobody anything; the next writeFileSync into it fails
                   with EPERM, on a directory just cached as hardened

Verification now compares the whole flag set, which also rejects IO and NP,
and names the offending flags in the failure. Both shapes are planted in
real-filesystem regression tests, including an assertion that a write into
the repaired directory succeeds and its child inherits. Confirmed both tests
fail against the old check and pass against this one.

* fix(security): back the hardening retry off exponentially

The fixed one-minute window bounded the retry rate but left a standing floor
of three attempts per path per minute on a host where hardening can never
succeed — FAT32/exFAT, a network path, a redirected profile. That budget is
per path and there are several secure files, so the floor multiplied into
tens of thousands of icacls spawns a day for work guaranteed to fail.

The delay now doubles after each consecutive failure, from a one-minute
floor to a thirty-minute ceiling, and the attempt cap is gone entirely: once
the backoff elapses the path is re-probed however long it has been failing.
A permanently incapable host settles at ~2 attempts/hour.

Slowing the backstop costs almost nothing, because it is not the recovery
mechanism: the synchronous write path is deliberately unthrottled, so a host
that recovers hardens on its very next credential write regardless of what
the read-path budget says.

The `throttled`/`recovered` reports are unchanged and matter more here,
since the quiet periods between probes are now much longer.

The curve is pinned in a new unit test against the exported delay function
rather than a copy of its constants, covering the doubling, the ceiling
holding at 5000 consecutive failures, a 30-day failing path still
re-probing, one announcement per degraded episode, and per-path isolation.
The integration tests keep only what they uniquely prove: that the read path
is wired to the budget, and that a day of failures still re-probes.
Confirmed four of these fail against a reinstated lifetime cap.

* ci(windows): run the real-icacls DACL suite in CI

The win32 suite only self-skips off Windows, so it passed vacuously in
every lane. Register it the way the cmd-shim suite is registered.

* fix(security): describe the cache's real cost, which is icacls now

Both cache comments still justified themselves with PowerShell -- "~1-1.5s" and
"a PowerShell spawn every read" -- in the same file whose PR removed PowerShell
from this path. The caches are still right, but for different numbers, and the
old ones are the kind an engineer would reasonably delete a cache over.

The real shape: hardening verifies first and returns early, so an already-correct
DACL costs one synchronous icacls spawn and a rewrite costs four (verify, reset,
grant, verify). Still worth caching on the read path, which polls at ~2/s.

* test(security): make the DACL suite safe to schedule

Registering this spec in the Windows lane put it under two rules it had
never been measured against.

Teardown now goes through `removeTreeSync`, which the lane's boundary test
requires, and repairs the DACLs the suite plants on purpose first: those
retries only cover transient locks, so a regressed `(OI)(CI)(IO)` repair
leaves the root un-removable and `afterAll` throws EPERM.

And the no-permission case decides by elevation before it writes anything.
`windows-2022` runs elevated, where hardening succeeds: the old branch
asserted nothing about denial and instead replaced the `hosts` DACL, then
`icacls /reset` -- which is not a restore, it drops the explicit
`SYSTEM:(F)` that file ships with. Ephemeral in CI; permanent for a
developer running the lane from an elevated shell. Now it asserts or it
skips. The probe reads the token integrity SID rather than `icacls /save`,
which succeeds unelevated (`BUILTIN\Users:(RX)` carries READ_CONTROL) and
would have skipped the case on every machine.

* fix(security): measure the hardening latches on a clock that cannot go backwards

`mayAttemptHardening` compared wall-clock times, so any backwards step --
an NTP correction, a VM snapshot restore, a user changing the clock --
made the elapsed time negative and held every failing path below its delay
until the clock caught up. Measured at the 30-minute ceiling with the clock
stepped back a year, the path was refused at +0d, +1d, +30d, +180d and
+364d, and re-probed only at +366d. That is the permanent latch the
exponential backoff was added to remove, and it contradicts the module's
own "bounds the rate without ever bounding the lifetime".

The SID lookup's own one-minute window had the identical shape and is
worse: a failed lookup makes `planFor` return null, which disables the
synchronous *write* path too, so the write-path exemption that recovers
the read-path budget cannot recover it. Both now measure elapsed monotonic
time, following the repo's existing `monotonicNowMs` spelling.

Two things the write path was not doing, both found in the same pass:

- A successful synchronous apply now records the outcome. It is exempt
  from the budget, but it was also invisible to it, so a host that had
  demonstrably recovered kept the read path backing off for up to 30
  minutes and no `recovered` transition ever came from that lane. Only
  success is recorded; recording failure would put the exempt lane back
  under the budget.
- `writeSecureFile`'s JSDoc now says its boolean covers the file only. The
  directory harden is fire-and-forget and answers `pending` on Windows
  regardless, so a `true` says nothing about the directory's ACL.

* fix(security): stop the hardening test doubles from faking a no-op

Three CI failures on this branch, one failure shape: hardening silently
does nothing and the check that should have caught it agrees.

The auth critical-path test hand-rolled a `node:child_process` factory with
`execFileSync`/`execFile`. The rewritten ACL path goes through
`runProcessSync`, i.e. `spawnSync`, which the factory never returned — so
every spawn threw into the SID lookup's bare catch, `planFor` returned null,
and hardening no-opped. It mocks `child-process/run-process` now, the
boundary production code actually calls and the one sibling ACL tests
already mock: an export missing there fails loudly by name instead of
returning undefined. Its fake icacls writes a real UTF-16LE SDDL file, so
the pinned spawn count per write is a property of the ACL path rather than
of the double. The test forces `platform='win32'`, so this failed on every
platform, Linux CI included.

`windowsSystem32Binary` is a production bug, not a test bug: it builds a
Windows path with the host `join`, which off-platform yields the mixed
`C:\Windows/System32/whoami.exe`. On Windows the two joins agree, which is
why it survived; on Linux the SID lookup's whoami match missed and 27 of
secure-file's 32 tests exercised a lane that never ran. These are always
Windows paths, so `path.win32.join` is what it should have been.

The import-boundary pin still read 160 after this branch migrated
secure-path-windows-acl.ts off `node:child_process`; the ratchet correctly
refuses a pin left above reality.

* fix(security): resolve the machine-relative SDDL alias, and stop a denied read destroying the file

Path hardening verified the DACL it wrote by comparing the SIDs `icacls /save`
reports. SDDL substitutes two-letter aliases for well-known SIDs, and the
resolution table could only hold constants -- but `LA` and `LG` name an account
by RID inside the *machine's own* SID, so on a box whose user is the built-in
Administrator (a CI runner, an Administrator-only install) the current user read
back as `LA`, matched nothing, and hardening reported failure for every path.
Resolve those two against the machine authority derived from the user SID;
without one they stay unresolved and the comparison still fails closed.

Three secret stores treated any read failure as "malformed -- regenerate" and
overwrote. A hardened file granting a SID this process does not hold reads as
EPERM while its directory stays writable, so the overwrite succeeds: renaming
over an unreadable file needs FILE_DELETE_CHILD on the parent, not DELETE on the
file. That destroyed the E2EE secret key, every paired device's bearer token,
and the plugin vault. Distinguish EPERM/EACCES from a parse failure and refuse.

Also close the async lane's unhandled rejection: `void p.then(onSettled)` turned
a throw from `onSettled` into a dead main process, and the retry budget it calls
threw whenever nothing had configured it -- a contract held only by import
order. The budget now defaults its own bounds.

* test(windows): say which ACEs icacls listed when a planted DACL fails

`toHaveLength` reports only a count and vitest elides the array, so three
preconditions failing on the CI runner said "expected 3, got 6" and nothing
about what the sixth entry was. Name the entries in the failure.

* fix(security): stop three more stores overwriting what they were denied

Same swallow-default-overwrite shape as the readers already fixed, found by
sweeping every store that reads under a hardened root.

- plugin-storage-store.ts returned `{}` on any read failure and set()/delete()
  wrote it back, losing the plugin KV store. It is the secrets store's shape
  line for line, so the two now behave identically.
- relay-revoke-outbox.ts returned [] and save() wrote it, dropping revocations
  that never reached the relay -- a revoked device stays live.
- profile-cloud-session-store.ts mapped an EPERM read onto `decrypt-failed`,
  which fails the `status === 'found'` guard in clearCloudSessionIfUnchanged and
  falls through to an rmSync of the account session. A denied read now reports
  `unreadable`, which licenses nothing; the refresh path bails on it and the
  auth status surfaces it rather than reporting a bare reconnect.

All reuse isPermissionDeniedError. The predicate stays an EPERM/EACCES allow
list rather than "ENOENT defaults, everything else throws": these stores are
meant to self-heal a truncated or malformed file, and inverting it would turn a
corrupt keypair into an app that cannot start. The distinction that matters is
"could not read it" versus "read it and it was garbage".

* test(windows): plant fixture DACLs that cannot inherit what they did not plant

%TEMP% grants [SYSTEM, Administrators, <user>] (OI)(CI)(F) by default, and those
propagate into every fixture. Three preconditions read back 4 and 6 ACEs where 3
were planted, and the extras looked like Orca's own hardening because the shape
is identical -- on a runner whose user is the built-in Administrator, the
inherited trio IS the trio production grants.

Combining /inheritance:r with /grant:r leaves the argument order to icacls, and
that combined form drops the inherited ACEs on Windows 11 but keeps them as
explicit ones on the Windows Server runner. Removing inheritance in its own
invocation makes the grant the whole DACL on either host, and the fixture root
is de-inherited once up front so nothing propagates in.

Rooting the fixtures outside %TEMP% would not have fixed this: any directory
inherits from wherever it lives. The fix is to stop inheriting, not to move.

No assertion is relaxed -- the counts stay exact.

* test(windows): pick a foreign SID that stays foreign on an elevated runner

`S-1-5-32-544` is only foreign to a token that is not an administrator. The CI
runner is elevated AND logged in as the built-in Administrator, so granting
Administrators granted the reader full control: the file stayed readable, and
all six preservation assertions went vacuous rather than proving anything.

BUILTIN\Guests is resolvable everywhere and no interactive token is a member,
so the read is denied on an unelevated developer box and on the runner alike.
An unresolvable SID would have been the stronger choice but icacls rejects one
with ERROR_NONE_MAPPED (1332).

The premise guard is what caught this -- it asserted the file was actually
unreadable instead of trusting the grant, and named elevation as the suspect.

* fix(security): refuse on any read that never reached the contents, not just a denied one

isPermissionDeniedError becomes isUnreadableError, because "permission denied"
was never the concept -- "could not read it", as opposed to "read it and it was
garbage", is. EBUSY, EMFILE, ENFILE and EIO say exactly as little about a file's
contents as EACCES does, and they fell into the branch that regenerates and
overwrites. On Windows EBUSY is the likelier of the two: antivirus holding a
credential open at the moment of a startup read produces it, which makes it a
commoner path to the same permanent loss than the ACL case that motivated the
original fix.

Still an allow list, deliberately: ENOENT keeps licensing a create, and a parse
failure keeps self-healing. The stores are built to recover from a truncated
write, and turning that into a refusal would trade a recoverable state for an
unrecoverable one on the startup path.

Also fixes the regression suite's own premise on an elevated runner:
makeUnreadable combined /inheritance:r with /grant:r, and that form keeps
%TEMP%'s inherited [SYSTEM, Administrators, user] as explicit ACEs on Windows
Server -- so the file stayed readable and all six assertions were vacuous. Same
split-the-invocation fix as the ACL suite's planter.

* test(windows): skip the preservation suite where a read cannot be denied

An elevated token logged in as the built-in Administrator reads straight through
a DACL that grants it nothing -- confirmed on the CI runner against both
BUILTIN\Administrators and BUILTIN\Guests, and with the grant split into its own
icacls invocation so the DACL really was the planted one. On such a host the
premise these tests rest on does not hold, and every assertion would pass while
proving nothing.

So probe once at module scope and skip rather than assert vacuously -- the same
trade the ACL suite already makes for its unelevated-only case. The gate stays
in the compound `<win32 check> && <flag>` form the win32 lane ratchet detects, so
the file stays registered in both lane lists.

Coverage is not lost where it counts: isUnreadableError has unit tests that run
on every platform and every host, and the stores' refusal is exercised in full on
any machine where a denial is reproducible -- which is every developer box.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-05 21:13:06 -07:00
OrcaWinandOrca Worker 975bbdedcc fix(windows): scan ports natively instead of encoded PowerShell (#17861)
* fix(windows): scan ports natively instead of encoded PowerShell

Microsoft Defender for Endpoint scored the relay's Windows port scan as
suspicious PowerShell plus network discovery (T1049). The command line was
`-ExecutionPolicy Bypass -EncodedCommand <base64>` around a
Get-NetTCPConnection/Get-Process join -- base64 next to a policy override is
the highest-weighted token pair on a PowerShell command line, and netstat only
ever ran as its fallback.

Invert the chain. `netstat.exe -ano` is now the primary reader and the owning
process name comes from the shared native process table, which exists to keep
PID lookups off PowerShell. The payload survives only as a last resort, and
without the override: execution policy gates script files, never `-Command`,
so nothing needed it (verified: `-ExecutionPolicy Restricted -Command` runs).

Drop `-p tcp` while inverting: on Windows that protocol name means IPv4 only,
so as a primary reader it would have hidden every `[::]` listener the payload
used to report. Names arrive as `sshd.exe` from the table and are published as
`sshd`, keeping the sshd filter and old clients' rendering intact.

Routes both spawns through runProcess, removing the file from the
child_process and windowsHide ratchets.

* fix(windows): read netstat state by shape and refuse a truncated table

Review of the port-scan inversion found two ways the new primary path could
be silently wrong, both of which would have kept the flagged PowerShell
payload running on exactly the hosts this change targets.

`LISTENING` is not in netstat.exe. It lives in System32\<locale>\netstat.exe.mui
and MUI selection follows the UI language, so the pinned-locale env in
relay-command-env.ts cannot reach it -- a German host prints `ABHOEREN` and the
word test parsed zero rows. The zero-listeners guard then read that as a
blocked reader and ran `Get-NetTCPConnection` every 12-30s forever, or returned
nothing at all where PowerShell is also restricted. Keep the word as the fast
path and, when it finds nothing over output that did contain TCP rows, re-read
by shape: only a listening socket has no peer. Measured on this host across all
four states present (LISTENING 47, ESTABLISHED 49, CLOSE_WAIT 29, TIME_WAIT
213): zero non-listening rows with a zero peer, zero listening rows without
one, and the same 47 rows parse after substituting the German state words.
Shape stays the fallback because `BOUND` also prints a zero peer.

Truncation was invisible: createOutputSink discards overflow, ProcessResult
carries no flag, so a capped read still exits 0 and its head still parses.
netstat orders IPv4 TCP, then IPv6 TCP, then UDP, so a host with tens of
thousands of TIME_WAIT rows would have lost every `[::]` listener -- the exact
loss dropping `-p tcp` exists to prevent, and one the zero-listeners guard
cannot see. Refuse the read instead. A `truncated` flag on the shared sink
would be cleaner and is left as a follow-up rather than widened into this PR.

Also: decline to wait on the shared process table once the request is aborted
(it takes no signal and must not be cancelled for other callers); note the
name lookup as best-effort, since a TTL-cached snapshot can hand a recycled
PID its previous owner name; log once on either fall-through, because both are
permanent and invisible when wrong; and drop a stderr assertion that any
PowerShell autoload banner would redden.

Correcting the cost claim in the previous commit: the aggregate win holds with
the native addon (netstat 21ms vs the retired payload 860ms at 532 processes),
not without it. The addon is optional, the snapshot TTL is 500ms and the scan
cadence is 12-30s, so a relay with no active agent pane never warms its own
cache and pays ~1.4s cold on the CIM path -- slower than what it replaced.

* fix(windows): log the port-scan fall-through on the relay diagnostic stream

Checked where this code actually runs before trusting the log. `console.warn`
did reach a file, but relayLogLine is the right call and the reasoning is worth
recording.

`scanWindowsListeningPorts` runs only in the detached relay daemon: relay.ts
returns early for --connect and --orca-cli, so PortScanHandler is reached only
through runRelayDaemon, and both launchers start it detached with a log file
(POSIX `> relay.log 2>&1`, Windows `1>relay.log 2>relay.err.log` via
Win32_Process.Create). installRelayLogRotation then wraps both streams into
relay.log, which is the file the documented diagnostics tail reads. Verified by
installing the real rotation over a temp path and reading the file back.

So the line surfaced -- but untimestamped, in a log whose format exists so
reconnect flaps can be correlated with the events around them (#7773).
relayLogLine is that format and the relay idiom in 41 other places, and
"since when has this host been stuck on PowerShell" is most of what this line
is for. The test spies on process.stderr to pin the stream and the ISO stamp
rather than just asserting something was called, since a fall-through logged
somewhere unread is the failure being guarded against.

Also fixes a comment that ended its own block early: `relay-*/relay.log` in a
doc comment contains `*/`.

* fix(windows): keep the dominant zero-peer state when reading a localized netstat

Shape alone promoted any zero-peer TCP row, not just listeners. `BOUND` and
`CLOSED` print a zero peer too, and on a localized host their state words are
exactly as unreadable as the listening one -- so a German host with listeners
plus one BOUND socket published a phantom listener. Reachable on an English
host too: with zero listeners a lone BOUND row is promoted AND, because the
result is then non-empty, it suppresses the blocked-reader fall-through.

Group the zero-peer rows by state word and keep only the largest group. A
transient BOUND or CLOSED socket cannot outnumber the listeners (51 against 0
on this host), so this removes the class rather than special-casing the words,
which would just be the localization bug again. An exact tie keeps every tied
group rather than guessing -- no worse than reading shape alone.

Verified against real netstat output: injecting a BOUND row into the localized
capture leaves the result identical to the English answer (47 rows, no phantom
65001). The new test has teeth -- reverting the grouping fails it and nothing
else.

Corrects two claims that were slightly wrong: the docblock said shape was the
fallback because BOUND prints a zero peer, which described the hazard without
saying it was unhandled; and a test comment said an English host "never sees a
bound socket", true only when it has at least one readable LISTENING row.

Also gates the fall-through log per reason instead of per module, so a host
that parses nothing today and truncates tomorrow reports both faults. Same
one-shot cost, and the vocabulary is two fixed strings so the set cannot grow.
That guard matters more than it looks: --log-file rotates stdout only, so the
file stderr can land in is unrotated.

* docs(windows): note the direction the zero-peer majority rule can fail in

The docblock described the tie case and stopped there, which reads as a
complete account of the limits when it is not: a majority rule inverts if the
majority is wrong, and enough transient zero-peer sockets would publish the
phantoms and drop the real listeners. Someone would reasonably have concluded
the rule was safe in both directions.

Trigger numbers and the repro stay in the PR discussion; the code only needs
the reader to know the rule has a direction, and the hatch (defer to the
PowerShell reader, which reads the state word instead of inferring it) since
that is the part a future editor would otherwise re-derive.

* ci(windows): run the real-netstat port scan suite in CI

The win32 suite only self-skips off Windows, so it passed vacuously in
every lane. Register it the way the cmd-shim suite is registered.

* test(windows): lower both child-process ratchets to the ground this PR took

Migrating the port scan off `node:child_process` onto `runProcess` drops
`src/relay/windows-port-scan.ts` from both allowlists, so both offender
counts fall by one. Each ratchet pins the count from below as well as
above, so a pin left above reality fails and re-opens room for the next
direct import to land for free.

* docs(windows): qualify the no-PowerShell claim on the netstat scan

The scan starts no PowerShell of its own, but no released relay carries the
optional `windows-process-tree.node` addon (only dev-channel-win-build.yml
builds it), so the shared process-table read falls back to a CIM scan that
forks one `powershell.exe`. The EDR win is the removal of the
`-EncodedCommand` / `-ExecutionPolicy Bypass` shape, not the elimination of
PowerShell. Comment-only.

* docs(windows): record the identity-reader follow-up and the perf table's addon

attachWindowsProcessNames reads only `name`, so it should move to
`readWindowsProcessIdentityTable` once #17866 lands -- on that PR's detailed
reader it would open per-process handles for a field it discards. The reader
does not exist on this branch, so the call stays as-is with the follow-up
recorded rather than pulling #17866 in.

The process-table perf table's two Toolhelp32 rows assume the optional
`windows-process-tree.node` addon. The desktop bundles it; no released relay
does, so on an SSH host the CIM row is the operative number. Comment-only.

* docs(windows): state the CIM scan as the relay's normal path, not a fallback

No released relay carries the optional `windows-process-tree.node` addon --
release-cut.yml has zero references to it and only dev-channel-win-build.yml
builds it -- so the PowerShell CIM scan is what every SSH host runs. The
call-site docstring read as a conditional fallback standalone. Comment-only.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
2026-09-05 21:12:59 -07:00
bfc6a262a7 fix(windows): read command lines from the kernel, not each process's PEB (#17886)
* fix(windows): read command lines from the kernel, not each process's PEB

MDE incident D scored Orca for suspicious memory activity: the vendored
`@vscode/windows-process-tree` recovered every process's command line by
opening it with `PROCESS_QUERY_INFORMATION | PROCESS_VM_READ` and chaining
three `ReadProcessMemory` calls through the PEB and
`RTL_USER_PROCESS_PARAMETERS`. On a 750ms/2s cadence over the whole table that
is the credential-dumping primitive, whatever the intent.

Windows 8.1 added `NtQueryInformationProcess`'s `ProcessCommandLineInformation`
class (60), which returns the same string as a kernel-built `UNICODE_STRING`
under `PROCESS_QUERY_LIMITED_INFORMATION` alone. Electron's floor is Windows
10, so every supported OS has it. The PEB reader stays behind a process-wide
latch that only `STATUS_INVALID_INFO_CLASS`/`NOT_SUPPORTED`/`NOT_IMPLEMENTED`
can set; a pid that merely denied a handle does not re-arm it, because
`PROCESS_QUERY_INFORMATION` implicitly grants the limited right and so cannot
be obtained where the weaker open already failed.

The same hunk drops `PROCESS_VM_READ` from `GetProcessMemoryUsage` and
`GetCpuUsage`, which acquired it and never read an address space.

Measured on Windows 11 (514 processes), counted in-process by swapping the
addon's import table entries for counting stubs, per CommandLine scan:
`ReadProcessMemory` 1128 -> 0, desired access 0x0410 -> 0x1000, p50 12.7ms ->
9.3ms. Command lines were byte-identical on every process both readers
recovered (376/376, 379/379 across runs), including a 24,068-character argv
with quotes, non-ASCII and trailing whitespace, and a WOW64 target. Three
processes that refused the old rights granted the new one; none went the other
way.

* chore(deps): refresh the windows-process-tree patch hash in the lockfile

* fix(windows): drop the PEB fallback and detect the unpatched prebuilt

Review of #17886 found three ways the reader could still perform, or silently
resume, the primitive it exists to remove.

The class-missing latch was a permanent, process-wide, one-way downgrade back
to the PEB read, and any single target returning STATUS_INVALID_INFO_CLASS /
NOT_SUPPORTED / NOT_IMPLEMENTED could trip it. On an EDR-hooked ntdll -- the
entire premise of this change -- a hook that does not recognise class 60 would
have restored PROCESS_VM_READ plus three ReadProcessMemory per pid per scan for
the life of the process, unobservably, on precisely the machines this was
written for. The fallback is deleted rather than guarded: GetProcessCommandLine
now returns false and leaves the command line empty, which callers already
handle, so the addon imports no ReadProcessMemory at all.

That absence is what makes the property checkable on the artifact. The
published 0.8.0 tarball ships a loadable prebuilt built from unpatched source;
it is node-addon-api, so a bare require() accepts it, allowBuilds is false and
CI installs with --ignore-scripts, and a rebuild that soft-exits on a Windows
file lock leaves it in place. Source-text guards could never see it.
windowsProcessTreeAddonReadsProcessMemory() checks the compiled binary instead,
and is wired into the install check, the rebuild, and the relay build.

The repair itself never worked: `git apply` run inside a work tree prefixes
patch paths with the cwd-relative prefix, skips what does not match, and exits
0, so the branch always fell through to its own post-check throw. The package
dir is always under the project root, while the fixture that covered it was in
%TEMP%, outside any repo. Blinding git with GIT_DIR fixes it, and the test now
runs inside a real work tree.

Also from review: bounds-check the returned UNICODE_STRING against the
allocation (not the size the second query clobbers) and cap the probe so a
bogus length cannot bad_alloc a whole scan; test NT_SUCCESS explicitly; value-
initialize ProcessInfo, which left `memory` as stack garbage -- measured, 82
processes reported the same bogus working set; and correct a comment in
windows-process-table.ts that still described the command line as a PEB read.

Re-measured on Windows 11 (543 processes): ReadProcessMemory 1128 -> 0, with
the symbol absent from the import table so the IAT hook finds no slot to
count; desired access 0x0410 -> 0x1000 on all 543 opens; p50 13.5 -> 12.3ms;
405/405 command lines byte-identical including a 24,087-character quoted
non-ASCII argv and a WOW64 target; 3 processes recovered only by the new path,
0 only by the old.

* chore(deps): refresh the windows-process-tree patch hash in the lockfile

* test(scripts): stage a script's local imports into the native-runtime fixture

ensure-native-runtime.mjs gained an import of windows-process-tree-gyp-rebuild.mjs,
but the fixture copied only the script itself, so every case in the suite died
with ERR_MODULE_NOT_FOUND before reaching its own assertions. copyScriptWithLocalModules
already walks a script's co-located imports for exactly this reason -- its own doc
comment names this failure -- so use it rather than listing files by hand.

The two Windows cases still fail here, on a missing node-pty ConPTY runtime that
also fails on main; this only stops a resolution error from standing in front of
whatever they were meant to catch.

* fix(windows): route a locked stale addon to the Windows file-lock message

`pnpm install` with Orca running aborted with a raw EPERM stack. The stale-binary
guard -- which deletes an addon that still imports ReadProcessMemory so a skipped
rebuild cannot use it -- ran outside the try whose catch classifies Windows file
locks, and whose message is literally "Close running Orca/Electron/dev processes
for this worktree": exactly this situation.

Measured rather than assumed: rmSync against a loaded (memory-mapped) addon throws
EPERM, and `force: true` does not help, since it only swallows ENOENT. Cold copies
of the same file delete fine. So the delete threw a page before the handler that
knows what it means.

Moving the guard inside the try is the whole fix; the classifier already matches
the EPERM text. The new case runs the real script against a temp project whose
stale addon is held open by a live child process, and fails against the old
placement with the raw `syscall: 'rm'` stack the report described.

* feat(windows): warn once when command-line recovery is refused host-wide

Removing the PEB fallback removed a total-defeat vector, but it left a cliff: if
NtQueryInformationProcess(ProcessCommandLineInformation) is refused -- a hooked
ntdll that does not know class 60 -- every command line comes back empty and
agent identity matching silently degrades to image names. The addon still loads
and still enumerates, so every health check the app has stays green. A cliff
nobody can see is the failure mode this area keeps producing.

The querying process is the unambiguous probe. A process can always open itself
with PROCESS_QUERY_LIMITED_INFORMATION, so its own command line coming back empty
means the query is refused for every process -- not that some target denied a
handle, which is normal for roughly a quarter of the table. Keying on our own row
rather than a fraction means no threshold to tune and no false positive on a
hardened box where most processes deny.

One warning per session, gated on the CommandLine flag actually being requested so
a future identity-only reader cannot trip it. The suite's own SELF fixture gains a
command line for the same reason: a self row without one is the alarm, not a
detail.

* fix(windows): check the relay's staged addon at load, and answer tri-state

Two gaps in the ReadProcessMemory check, both about what it does not see.

It only ever looked at node_modules/@vscode/windows-process-tree. A relay host
has no node_modules of ours: it loads ./windows-process-tree.node staged beside
the bundle. The relay build asserts the symbol on the artifact it produces, but a
bundle and the addon beside it redeploy independently, so a host that has not
taken a new bundle keeps whatever binary is already there -- and the published
prebuilt is node-addon-api, so it binds cleanly and then walks every process's
address space. loadWindowsProcessTree now checks that file too and refuses it,
falling back to the CIM scan: slower, but not the thing an EDR quarantines a host
for. The predicate is duplicated rather than imported, because the config-script
copy is install-time tooling that drags in node-gyp and child_process, and this
module is bundled into the app and the relay.

And it returned false for a binary that is not there. All three callers happened
to be safe, but the name read as a safety predicate, so a future caller would take
a missing binary as verified. inspectWindowsProcessTreeAddon() now answers
clean/unpatched/missing over an explicit binary path -- which is also what lets
the relay's staged addon be checked at all -- and each caller states which state
it acts on.

Both are covered by cases that fail against the old code: without the load-time
check the unpatched staged addon is bound and the CIM fallback never runs, and
with 'missing' folded back into 'clean' the absence case fails outright.

* test(windows): load the addon in beforeAll, not at collection time

loadAddon() ran while the file was being collected, so on a Windows checkout with
no built addon the require threw before any case existed and took the seven
patch-text cases down with it -- cases that read only the patch file and need no
binary at all. Verified both ways against a deliberately unresolvable addon path:
at collection time vitest reports "no tests" for the file; from beforeAll the
seven text cases pass and only the three addon cases go.

* fix(deps): normalize the windows-process-tree patch to LF and let pnpm own its hash

`pnpm install --frozen-lockfile` failed on this branch on every platform with
ERR_PNPM_LOCKFILE_CONFIG_MISMATCH, which breaks CI and the release build.

Two coupled defects. The patch file was committed with CRLF -- 174 CR bytes,
against zero on main -- and `.gitattributes` pins `/config/patches/*.patch -text`
precisely so checkout cannot convert it, so those bytes reached every runner. And
pnpm hashes a patch **LF-normalized**, so the raw sha256 of a CRLF file is a value
pnpm never computes:

  raw sha256      322965470c05f63d8527f7d8e892ee26ee444136b66b57fd64c362a9f2ff05d1
  LF-normalized   f8ea245391c94da5770045aeea01fa6de466c2199c6ef46b5b769b398aa9823e

The lockfile carried the raw one, at all three sites. It is the only one of the
seven patches where the two digests differ, which is why the other six passed.

Normalized the patch to LF and took pnpm's own value from
`pnpm install --no-frozen-lockfile`; nothing here is hand-computed. With the file
LF-only the two interpretations coincide, so the lockfile, the contract test's
no-CR assertion and its hash assertion all agree at one number -- and
`config/scripts/windows-process-tree-patch-contract.test.mjs`, which was red on
this branch for the same reason, is green again. The lockfile diff is exactly the
three hash lines.

The regression check is the installer, not a digest. Two separate reviews
"verified" the shipped hash by recomputing sha256(patchBytes) and matching the
lockfile; both were wrong, because both repeated the same wrong assumption about
which bytes pnpm hashes. A check that reproduces the original mistake is not
independent. So the new case runs `pnpm install --frozen-lockfile --lockfile-only
--ignore-scripts` against a copy of the manifest, lockfile and patches, and
asserts exit 0 -- verified by deletion: restoring the shipped hash fails it with
the exact ERR_PNPM_LOCKFILE_CONFIG_MISMATCH from the branch's package (windows)
job.

Also corrected the `.gitattributes` comment claiming pnpm hashes patches
byte-for-byte. The `-text` setting is right -- `git apply` needs the exact bytes --
but that sentence is the claim that produced the wrong hash twice.

* ci(windows): run the process-tree patch suites in CI

Both suites only self-skip off Windows, so the binary-level check that the
addon carries no ReadProcessMemory passed vacuously in every lane.

* fix(windows): force core.autocrlf=input for the patch repair

My LF normalization of the windows-process-tree patch broke the `git apply`
repair path introduced in this PR. The two are coupled and I checked only one.

Those 174 CR bytes were not editor noise. They sat on exactly the pre-image
lines and nowhere else -- 107/107 in src/process.cc, 67/67 in
src/process_commandline.cc, 0 on every added or context line -- because
@vscode/windows-process-tree@0.8.0 ships those two sources as CRLF. Normalizing
the patch made its pre-image stop matching the file it is applied against.

Measured, reconstructing the true CRLF pre-image from the pre-normalization
blob and applying the current LF patch:

  core.autocrlf   plain   -c core.autocrlf=input
  true            exit 0  exit 0
  input           exit 0  exit 0
  false           exit 1  exit 0

`false` is Git's own built-in default and what "checkout as-is" selects in the
Git for Windows installer -- on this box the `true` that hides it comes from the
installer's system gitconfig, not from anything in the repo. There the repair
throws, ensureWindowsProcessTreeCommandLinePatch reports "still reads the PEB,
and repairing it ... failed", isWindowsNativeLockError does not match that text,
and `pnpm install` dies with no path forward.

Forcing the mode rather than `--ignore-whitespace`: both fix every cell and both
leave the applied file fully LF, but `input` relaxes line endings only, so a hunk
whose real content drifted is still rejected. The repair rewrites a
security-relevant source file; it should stay strict about everything except the
thing that is legitimately ambiguous.

Not reverting the patch to CRLF: windows-process-tree-patch-contract.test.mjs
(pre-existing on main) forbids CR bytes in it, and pnpm computes the same hash
either way. LF plus the forced mode is the end state.

The suite could not have caught this. The fixture built its pre-image from the
patch itself and joined with '\n', so fixture and patch agreed by construction on
any encoding -- once again a test that passes without its fix. It now emits the
CRLF the real package ships, and the case runs under both autocrlf modes pinned
through a temp HOME gitconfig, because the repair blinds git to the repo and so
reads global config. Verified by deletion in both directions: with the flag
removed the autocrlf=false case fails with the exact "still reads the PEB" dead
end while autocrlf=true still passes, and with the fixture back on LF all eight
cases pass with no fix present at all.

Also corrected the .gitattributes comment I added last commit. It said `git
apply` needs the bytes the patch was written against, which is now false -- the
pinned bytes are LF and the bytes it was written against are CRLF. That is the
same class of confident-and-wrong claim that produced the bad hash twice.

* fix(windows): assert the rebuilt addon, and install the patch for real in tests

Three follow-ups from review.

**The packaged binary had no check.** The relay build asserts its own artifact
and ensure-native-runtime asserts what it loads, but nothing looked at the addon
copied into the packaged app -- so a rebuild that silently produced the upstream
reader shipped. `rebuild-native-deps.mjs` now asserts `clean` on it after
`rebuild()`. This is also the caller D4's tri-state was missing: every existing
site branches on `=== 'unpatched'`, so `missing` still behaved exactly like
`clean` everywhere, which was the thing making it a state rather than a boolean.
Here both non-clean states fail, and they fail differently: after a rebuild that
reported success, an absent binary is a broken build, not an absence to shrug at.

The fake `rebuild()` had to start producing a binary for that to mean anything,
so it now emits stand-in bytes and takes `addon: 'clean' | 'unpatched' | 'none'`.
Verified by deletion: with the assertion removed both new cases pass.

**The frozen-install case could not see a patch at all.** `--lockfile-only`
resolves and never applies one, so its coverage stops at hash consistency. Added
a case that installs `@vscode/windows-process-tree@0.8.0` for real with the patch
and asserts the materialized `src/process_commandline.cc` carries the marker and
no longer carries `ReadProcessMemory` -- about 1.5s for the pair.

Correcting the brief on that one: it does **not** catch the `git apply` breakage
from the previous commit. Measured -- with `-c core.autocrlf=input` removed it
passes cleanly, because `pnpm install` uses pnpm's own patch applier and never
runs our repair script. What it does catch is a patch pnpm can no longer apply:
corrupting one pre-image line fails both cases. The repair path stays covered by
the CRLF fixture in rebuild-native-deps-node-pty.test.mjs.

Worth recording, since it decides whether the LF normalization was safe at all:
pnpm applies the LF patch to the CRLF tarball sources without complaint, and
materializes them as LF with the marker present and `ReadProcessMemory` absent.
The primary install path was never affected -- only the `git apply` fallback was.

**Dead timeout.** The frozen-install case passed `timeoutMs: 300_000` to the
spawn while vitest capped the case itself at 30s, so on a cold runner vitest
would have killed it first. Both cases now declare the budget they use.

* test(windows): route the frozen-install check through the pnpm invocation owner

The new patched-dependencies check hand-rolled a PATH walk naming 'pnpm.cmd',
which the windows batch shim spawn boundary ratchet rejects: pnpm-cli-invocation
already owns that decision for every other script, and its allowlist only
shrinks.

Reuse resolvePnpmCliInvocation for the command and prefixArgs, and the shared
resolveCliCommand for the presence check, so no shim name is spelled here. Its
`shell` flag is dropped because runProcessSync refuses it and already drives a
shim through the interpreter itself.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-05 21:12:47 -07:00
OrcaWinandOrca Worker 687a22e1ee fix(computer-use): run the Windows runtime as one persistent helper (#17858)
* fix(computer-use): run the Windows runtime as one persistent helper

Microsoft Defender for Endpoint raised multi-stage Execution + Collection
incidents against Orca on Windows ("Screenshots were taken unexpectedly on
this device... Screen capture code was found in a script launched by
powershell.exe", factor "Executes suspicious MSIL code"). The desktop script
provider spawned a fresh powershell.exe per operation, so a single computer-use
session produced a burst of short-lived PIDs and re-emitted runtime.ps1's
inline Add-Type P/Invoke assembly on every click.

runtime.ps1 gains a -Serve mode that loads its assemblies once and then reads
NDJSON requests from stdin, and a new DesktopScriptRuntimeHost owns one
long-lived child: lazy spawn, strict serialization, a 30s per-request timeout,
restart on crash, a 120s idle shutdown, and dispose() on provider teardown. The
one-shot -OperationPath path stays as the fallback, and Linux keeps its python3
bridge unchanged.

Both Windows spawn sites now use -ExecutionPolicy RemoteSigned instead of
Bypass, falling back once to Bypass (and logging) when a Restricted host
refuses the unsigned script.

* fix(computer-use): recover the runtime host instead of latching it off

Review follow-up on the persistent Windows computer-use helper.

A helper that died before producing a line set an unavailable flag nothing ever
cleared, and the client then dropped the host for the life of the session. One
transient bad spawn — a Defender scan, a locked CSC temp directory — silently
restored the per-click powershell.exe burst and per-operation MSIL emission this
work exists to remove, with computer use still working so nothing looked wrong.
Start failures are now retried, then cool down for 60s, then re-probed; the
client keeps the host so it can come back. Repeated post-answer crashes cool
down too, and a single reply no longer clears the failure count.

The one-shot bridge decided its execution-policy retry from a message that fell
back to stdout, so a window title containing "SecurityError" could replay a
non-idempotent operation — a double click, keystroke or paste — and stick the
session on Bypass. The retry now requires empty stdout and a matching stderr.

Serve-mode replies carry an echoed request id. Without one a single stray stdout
line would make every later response answer the previous request, acting on
stale element indexes with no error raised; a mismatch now kills the child.
Non-JSON noise is ignored rather than counted as the helper having answered.

Also: warnings reach the main process over the sidecar's IPC channel rather than
its piped, unread stdio; the child is watched on close rather than exit; dispose
latches so a queued request cannot respawn during teardown; and the host is
split into a serve channel and an availability policy to stay under max-lines.

* fix(computer-use): prove a helper never started before replaying its request

The retry that replaced the permanent-latch bug could deliver unrequested
input. send() re-sent the same request whenever the helper died without
replying, but "no reply came back" is not "the operation did not run":
runtime.ps1 synthesizes the click and only then builds the snapshot, which
allocates a full-window bitmap and walks the UIA tree — a native GDI+/UIA fault
there is uncatchable, and leaves the click already delivered. A deterministic
fault meant three clicks from the host plus a fourth from the one-shot bridge,
surfaced as a single failed operation.

-Serve now writes one {"ready":true} line after its Add-Type work and before
its first read, so "never started" is a fact rather than an inference. A request
is replayed only when the helper died before announcing. A runtime.ps1 that
predates the announcement — reachable through the provider path override — is
covered by an observation-tool allowlist until a ready line proves otherwise.

Host-detected aborts (timeout, desynchronised reply, oversized line) suppress
the exit handler, so they were bypassing failure accounting entirely and a
helper failing that way was respawned once per operation forever. They now
count and are logged.

Also stop charging twice for one outage: entering the cooldown resets the
failure count, so the first death after recovery no longer re-enters a full
cooldown and an interleaved workload cannot be stranded on the one-shot bridge.

* fix(computer-use): ignore a stdin write callback from a torn-down helper

stop() destroys stdin, so a write still queued at teardown calls back with
ERR_STREAM_DESTROYED. The callback carried no channel or request identity and
write() had no closed guard, so it ran abortChannel a second time: stopChannel
no-opped but recordFailure and the warning did not, charging two failures for
one operation and reaching the 3-strike cooldown at half the intended rate.
That feeds the same accounting that keeps a persistently broken helper from
respawning once per operation.

The same root also allowed a late callback landing after a replacement channel
existed to stop that channel and reject a different request with the previous
one's error. Node fires the destroyed-stream callback on the next tick, well
before a new request arrives, so the double-count is the reachable effect;
binding the callback closes both.

write() now drops payloads and error reports once closed, and the host ignores
any report whose channel or request id is no longer current.

* test(computer-use): pin each stale-write guard independently

The channel's closed guard and the host's request-identity check are redundant
by design, and the existing tests only failed when both were absent. Someone
deleting one, believing the other was the covered one, would have got a green
suite and a live regression — the same shape as a test that passes without the
fix it was written for.

Each is now pinned on its own. The channel's half is tested against the channel
directly: after stop() it takes no writes and reports no error from one already
queued, which the host cannot observe because it drops the channel at the same
moment. The host's half is pinned by the case the channel cannot see — a live
channel whose request was already answered, where backpressure delivers a write
callback for a request that is no longer pending.

Removing either guard alone now fails a test. Both carry a comment saying they
are deliberately redundant and separately pinned, so the next reader does not
have to rediscover this from the diff.

* ci(windows): run the computer-use runtime host suite in CI

The win32 suite only self-skips off Windows, so it passed vacuously in
every lane. Register it the way the cmd-shim suite is registered.

* fix(computer-use): time the runtime host cooldown on a monotonic clock

The start-failure cooldown was a wall-clock deadline, so a backwards step —
an NTP correction, a VM snapshot restore, a user changing the clock — left
`remainingCooldown()` returning the cooldown plus the whole step. A one-hour
step measured 3,660,000ms, and ten real minutes later still 3,060,000ms.

Nothing shortens it from there. Only `recordSuccess()` clears the cooldown on
a non-dispose path, and no request can reach a helper to succeed while it
holds, so every `send()` throws `runtime_host_unavailable` first. The host is
built with no `now` override and its lifecycle is a module-level singleton
that shuts down at process exit, so the latch held for the sidecar's life —
computer use kept working via the one-shot bridge while the per-click
powershell.exe burst this host exists to remove came back silently.

Store the instant the cooldown began and compare elapsed monotonic time,
following the two fixes in #17884. The field is `number | null` rather than
sentinel 0 because `performance.now()` legitimately returns 0.

Both new tests leave `now` unset, because the bug was in the default the host
picks and a test that injects a clock cannot see it.

* fix(computer-use): give a queued request its own deadline

The 30s request timeout was armed only in `sendOnce`, once a request reached
a helper. A request behind N timing-out ones therefore waited roughly N times
that with no deadline of its own: bounded, but the caller sees an `await` that
looks hung for minutes and gets no error to act on.

Move the serialization tail into its own class and arm a deadline at enqueue
time. Only the wait is bounded — a request that reaches a helper still gets
its full execution budget, so nothing that used to succeed now fails. An
expired request is dropped rather than sent late: the caller has already been
told it failed, and a click delivered after that is worse than no click.

The tail keeps its never-rejecting shape and chains on the turn rather than on
the raced promise, so a caller giving up early cannot release the next request
while its predecessor is still in flight.

* fix(computer-use): stop reading a locked file as an execution policy block

`UnauthorizedAccess` is the FullyQualifiedErrorId PowerShell reports for a
policy block, and it is also a strict prefix of `UnauthorizedAccessException`,
which .NET raises for any ordinary locked or ACL-denied file. The predicate
matched the token unanchored, so an AV scan holding runtime.ps1 or a locked
CSC temp directory was read as a policy block.

Two consequences, both bad. `escalateExecutionPolicy()` has no path back, so
one false match spent the rest of the session on `-ExecutionPolicy Bypass` —
the exact command line token this stack exists to stop emitting. And on the
one-shot path `isPolicyBlockedStart` re-runs the operation: one-shot mode
writes stdout only after the operation returns, so a crash partway through an
action is indistinguishable from a helper that never started, and the click
lands twice.

Measured on Windows against all three records, which the test carries verbatim
as fixtures:

  policy/Restricted      FullyQualifiedErrorId: UnauthorizedAccess
  policy/RemoteSigned    FullyQualifiedErrorId: UnauthorizedAccess
  genuine access denied  FullyQualifiedErrorId: UnauthorizedAccessException

`\b` is the whole discriminator: between `s` and `E` both sides are word
characters, so no boundary exists there and the exception cannot match.

Dropped two alternatives that measurement showed were wrong. `PSSecurityException`
never appears — the record surfaces through a native-command wrapper and reports
`ParentContainsErrorRecordException`. The prose is wrong three times over: it
differs by policy, it is localized, and PowerShell hard-wraps it mid-sentence.

Anchoring on the `FullyQualifiedErrorId:`/`CategoryInfo:` labels would be more
precise again, but those labels are localized where the values are not, so it
would lose a real block on a non-English host and strand it with no fallback.
Matching the values with word boundaries keeps both directions; a fixture with
translated labels pins it.

The escalation stays sticky. With the predicate correct, it only fires on a
machine that really does block, where re-probing the preferred policy would buy
a guaranteed failed spawn per operation.

* fix(computer-use): route a malformed request back to the request that caused it

`ConvertFrom-Json` throws before `$requestId` is read, so the serve loop
answered an unparseable request with an untagged error. On the client that is
not an error at all: `deliver()` sees no matching id, calls `abortChannel`,
kills the helper and charges a failure — and the helper's own message is
discarded. A parse failure was reported as a stream desync with no trace of
the real cause, and three of them walked into the 60s cooldown behind three
misleading "did not match" messages.

Recover the id from the raw line when the parse fails. No wire change: the
response shape is untouched and `BridgeResponse.requestId` already documents
this echo. It is the same shape the helper already returns for `not_a_tool`,
where the id survives because it is read before the operation runs. Both
mixed pairings degrade safely — a new script with an old client resolves the
error normally, and an old script with a new client still aborts, but now
reports what the helper said.

When the line is mangled past recovering an id, the desync abort is the honest
outcome, so keep it and carry the helper's text into it rather than replacing
it. A line the helper could not tag is usually the only account of the cause.

Proven against the real `runtime.ps1 -Serve`: the host can only write
well-formed JSON, so the parse-failure branch is unreachable through it and
the test drives the channel directly.

* fix(computer-use): keep the Bypass escalation only when Bypass actually works

AppLocker and WDAC constrained language mode raise PSSecurityException under
the same SecurityError category a real execution-policy block uses, so the
predicate matches them - correctly, on the evidence available. But those block
the script at parse time, which `-ExecutionPolicy Bypass` cannot lift. The
escalation was sticky unconditionally, so on a WDAC host we misdiagnosed,
retried, failed again, and then latched: every later command line carried the
most heavily weighted MDE token there is, on exactly the hardened, monitored
enterprise machine that is watching for it.

Treat the escalation as the diagnosis it is. A fallback that cannot start a
helper either disproves it - the policy was not what stopped the first attempt
- so revert to RemoteSigned instead of latching. When Bypass does start a
helper the diagnosis is confirmed and it stays sticky exactly as before, so a
genuinely Restricted machine still never pays a re-probe per operation.

The revert lands inside the outage rather than only at its end, so a
misdiagnosis costs one Bypass command line instead of one per attempt, and an
escalation that never proved itself does not outlive the cooldown that ends
the outage. Deliberately not a permanent "fallback is useless" flag: a Bypass
attempt that failed for a transient reason would then disable the fallback for
the session, which is the same latch in the other direction.

Only `runtime_host_unavailable` proves no helper started, so only that reverts;
a helper that started and then died proves Bypass works. That also makes the
policy branch reachable on a final attempt for the first time, so it now
rejects as unavailable rather than a generic error - that code is what routes
the operation to the one-shot bridge, which carries its own policy fallback,
and without it an all-blocked host would fail operations outright instead of
degrading. The pre-existing "reports itself unavailable when Bypass is also
refused" test pins that.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
2026-09-05 21:12:33 -07:00
Neil 6031c19e9f ci: reduce dependency, checkout, and test deadline overhead (#18968)
* ci: reduce dependency, checkout, and test deadline overhead

* ci: avoid generic E2E jobs for native-only IME changes
2026-09-05 18:20:12 -07:00
OrcaWinandOrca Worker b6ca8dad99 fix(hooks): register the Claude hook script directly on Windows (#18875) (#18905)
* fix(hooks): register the Claude hook script directly on Windows (#18875)

The Windows Claude Code lifecycle hook was registered as
`powershell.exe -NoProfile -EncodedCommand <...>` whose entire decoded payload
was a `Test-Path` and a call to `~/.orca/agent-hooks/claude-hook.cmd`. Every
hook event paid a full PowerShell start-up to reach a script that exits at its
first `ORCA_PANE_KEY` guard, so sessions outside Orca paid it to do nothing.

Register the script path itself instead, with `|| echo {}` for the
neutral-JSON-when-missing contract (#14818). Measured on Windows 11, invoked as
Claude Code invokes it (`printf payload | bash -c -l "<command>"`):

  idle (n=12)          baseline 177ms | before 471ms | after 213ms
  10-way conc (n=40)            --    | before 656ms | after 296ms
  p95 under load                --    | before 696ms | after 337ms

It also drops an interpreter from the chain the hook's timeout kill must tear
down. Killing the hook does not kill its PowerShell grandchild, which still
holds the stdout handle the agent reads to EOF -- measured, EOF arrived 352ms
AFTER the kill, when the orphan exited by itself. msys2 creates children
suspended and resumes them after, so a kill landing in that window strands one
that never exits and EOF never comes; that is the reported frozen session.

The encoded launcher stays as the fallback for profile paths the shells cannot
carry bare (space, `%`, `^`, `&`, non-ASCII) and for hosts where Git Bash is not
resolvable, because PowerShell 5.1 rejects `||`. Every other agent's hook is
untouched, as is the remote/SSH path.

Not adopted from the report: `cmd.exe /d /c <path>` (MSYS rewrites the `/c`
under Git Bash -- measured, the invocation fails), and raising the 10s timeout
(the orphan survives the kill regardless; the fast path puts the hook 30x under
the budget so the kill effectively stops firing).

* fix(build): list the new hook launcher modules in the CLI tsconfig project

config/tsconfig.cli.json enumerates its files explicitly, so the two new
imports reached by src/main/claude/hook-settings.ts failed tc:cli with TS6307.
src/main/git-bash.ts pulls in only node:fs, node:path and a shared constant,
so it adds nothing heavy to the CLI project.

* fix(hooks): address review of the direct Windows Claude hook launcher

- Make the Windows hook suites host-independent. A box with a cmd.exe AutoRun
  (HKCU\...\Command Processor\AutoRun) failed them at HEAD too: the tests
  redirect USERPROFILE, the AutoRun target vanishes, and MSYS spawns a .cmd
  without /d so AutoRun runs and lands on the hook's stderr. Seed an empty
  target, including under the deliberately-absent profile.
- Note in managed-hook-stdin-lifecycle why the "missing managed script" case no
  longer exercises the fallback for the direct shape (it carries an absolute
  path, so a redirected profile changes nothing); that path is covered live in
  windows-direct-cmd-hook-command.test.ts.
- Keep the direct shape off UNC profiles: WINDOWS_CMD_SAFE_PATH admits them, but
  //server/share/... is not a command cmd.exe reliably starts.
- Correct the comments: `|| echo {}` also fires when cmd.exe itself exits
  non-zero (failing AutoRun), printing {} twice. The encoded launcher exited 1
  on that same box, so neither shape is clean there.
- Test the contract that replaced runtime %USERPROFILE% resolution (STA-3348): a
  stale absolute path reports not_installed and is rewritten on install.
- Record the standing unmeasured assumption in windows-edr-posture.md: `||` does
  not parse in Windows PowerShell 5.1, so a compat consumer that hosts hook
  strings there would fail closed. Measure before widening to another agent.
- Trim the launcher comments per AGENTS.md; the numbers live in the doc.

* test(win32): register the new Windows-gated hook test in the CI lane

win32-test-lane-registration guards against exactly this: a Windows-gated file
that self-skips on ubuntu and reports success, so it runs on no machine. The new
windows-direct-cmd-hook-command.test.ts needs both entries — WINDOWS_PACKAGE_TESTS
decides whether package_windows runs for a diff, and the workflow argv decides
whether the file runs once that job started.

* test(win32): remove the hook temp tree through the retrying helper

windows-lane-tree-removal-boundary scans exactly the specs in the Windows CI
lane, so registering windows-direct-cmd-hook-command.test.ts subjected it to the
rule: cmd.exe and bash have just exited in that tree, and a raw recursive rm
throws EPERM on Windows while their handles drain, turning a green spec into a
lane failure. Use removeTreeSync, which carries the repo's maxRetries policy.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
2026-09-05 17:50:33 -07:00
Neil 7bb54cc2f7 ci: reduce runner overhead and disposable package compression (#18948)
* ci: reduce PR runner overhead and package compression time

* ci: validate mobile when its dependency action changes
2026-09-05 16:56:57 -07:00
Neil 0a821e5bc8 fix(crash-reporting): make the own-Chromium gate a real choke point, and stop a refusal leaking the root (#18459)
* fix(crash-reporting): make the own-Chromium gate a real choke point

Round-3 review found the guard was not the choke point its own comments
claimed: six pid-addressed `taskkill /pid <pid> /t /f` families in main were
ungated and uninstrumented, so the stale-pid shape stayed producible and a
`selfInitiatedTreeKillCount: 0` could read as exculpatory when it was not.

- Gate the remaining main-process families: the git command-runner abort, the
  notebook-cell and automation-precheck timeouts.
- Turn the `src/shared` seam into the gate itself (`process-tree-kill-gate`), so
  the runProcess choke point, the codex app-server deadline kill and the
  ephemeral-VM recipe kill ask the same decision. Those three are compiled into
  the CLI/relay too and cannot import main; main installs the guard at preflight.
- Ratchet (`main-process-tree-kill-gate.test.ts`): a new pid-addressed taskkill
  in main that skips the gate fails, and the allowlist entries must still exist.
- Give pid-addressed kills eviction priority in the 32-entry ring: 32 routine
  `win-pty-job` teardowns from a window-close burst no longer evict the one
  entry that discriminates a self-kill from an external one.
- Correct the coverage doc, which described the uninstrumented Windows sites as
  POSIX `process.kill(-pid)` group kills and omitted the git and codex paths.

* fix(crash-reporting): keep a refused tree-kill from leaking the root it owns

A refusal must block the pid-addressed tree walk, not the termination. Five of
the six gated sites returned on refusal with no fallback, so a refused
`taskkill /pid /t /f` left git.exe, a timed-out notebook cell, an automation
precheck or an ephemeral-VM recipe running while the caller reported it stopped.
The root kill is addressed by the child handle, which cannot reach the recycled
pid the refusal is about, so it stays correct and required on that path.

Also fixes the ring eviction the scope preference introduced: with the ring
saturated by pid-addressed kills, the only non-pid-addressed entry is the one
just pushed, so the splice evicted itself and the detail came back `{}` --
byte-identical to the external-kill arm, in the window-close case the guard
exists for. Eviction now excludes the newest entry and falls back to FIFO.

Tests: refusal now asserts the root kill at all six sites, and the ring covers
the saturated-pid ordering as well as round 3's group-burst ordering.

* fix(crash-reporting): stop a refused tree-kill leaking the commit-message agent, and count call sites

Two round-5 blocking findings, both open on main and on both branches.

`killSourceControlAgentProcess` had no root-kill fallback on its win32 arm: the
taskkill was the only termination, so once the own-Chromium gate could refuse it
the promise resolved having killed nothing. Both callers do
`terminationComplete ??= killSourceControlAgentProcess(child)` and then release
the managed-home lock on that promise, so a refusal left the local Codex/Claude
commit-message agent running while the caller reported it stopped -- the
lock-contention failure the taskkill was added for. Same fix as the six sibling
sites: the handle-addressed root kill cannot reach the recycled pid the refusal
is about, so it stays correct and required on that path.

The ratchet was file-granular, not call-site granular: one gate mention anywhere
in a file exempted every taskkill in it, which left the six files that now ask
the gate ratchet-blind -- the inverse of what it is for. It now counts `/pid`
call sites against gate admissions per file, so a second ungated kill inside an
existing family fails. Keying on the `/pid` argument rather than a quoted
`taskkill` also catches a kill whose program name comes from a constant. The
three comments that claimed more than the old scan enforced now state the rule
and its two remaining blind spots.

Also: the recording in `admitSelfInitiatedTreeKill` is now wrapped the way the
`admitProcessTreeKill` seam already wraps it, with the refusal decision taken
before anything that can throw so a diagnostics failure cannot flip it; and
`orca-chromium-process-pids` documents the false-positive direction (a stale
`getAppMetrics()` entry plus pid reuse refuses a live unrelated child), which is
the mechanism the root-kill fallback exists to bound.

Tests: refusal now asserts the root kill at all seven sites; the ratchet asserts
call-site counting and the constant-program form.

* test(crash-reporting): run the own-Chromium gate against real Windows trees

Nothing on this branch had ever executed on Windows. The unit tests pin the
gate's decision against a mocked taskkill, which cannot show that the decision
does anything to a real process: that `/T /F` reaps a detached grandchild, that
a refusal leaves that tree standing, or that the handle-addressed root kill the
refusal path falls back to reaps the root while orphaning descendants.

Adds a win32-gated live test covering all four, registered in both the
`package_windows` CI lane and `WINDOWS_PACKAGE_TESTS` as
`win32-test-lane-registration` requires.

Also completes the coverage doc's "never instrumented" list, which omitted the
macOS keyboard-input-source probe's POSIX group kill in `ipc/app.ts`.

* fix(crash-reporting): pin the commit-message root kill on the Windows arm

The first Windows run of this branch found nine failures the macOS suite
cannot see: `commit-message-text-generation-test-harness` asserts
`expect(child.kill).not.toHaveBeenCalled()` on `process.platform === 'win32'`,
which is the contract the previous commit deliberately replaced — and it
branches on the real platform, so it is dead code everywhere CI runs today.

The harness now asserts the handle-addressed root kill on every platform. On
win32 it lands after the tree walk, so the expectation waits rather than reading
one tick early, and its ten call sites await it. Red against the pre-fix arm at
all seven sites; the production code is unchanged.

* test(crash-reporting): remove the Windows lane marker tree through the retrying helper

The new win32 spec teardown used a raw rmSync, which the windows-lane-tree-removal
boundary ratchet rejects — and which is exactly the EPERM the ratchet exists to
prevent, since this spec's marker directory is written by processes it has just
force-killed.

* fix(crash-reporting): only refuse pid-addressed tree walks, disclose the handle-less codex site

The own-Chromium gate refused the POSIX process-group arm of
signalProcessTree as well, which was new macOS/Linux behaviour: a stale
getAppMetrics() entry plus pid reuse would orphan a group that main reaps
today. A POSIX group only holds what Orca put in it, so the refusal is now
scoped to win-taskkill-tree and the POSIX arm is recorded and admitted like
the other group kills in main. That also drops the synchronous
getAppMetrics() read from every POSIX termination.

codex-turn-added-roots kills roots found by a table walk, so a refusal has
no handle to fall back to. Pin that the refusal is visible - crumb written,
turn reported as not cancelled - rather than fixing what cannot be fixed.

* test(crash-reporting): detach the Windows survival fixture and observe real spawns
2026-09-04 16:42:42 -07:00
f36c03e84a fix(windows): make the install-dir ACL repair rescue the launch it runs in (#18361)
* fix(windows): repair the poisoned install-dir ACL before the window, not after

The install-dir LPAC ACL poison (electron/electron#51761) still costs every
affected machine at least one crash: the probe that detects it is
setImmediate-deferred and answers 0.9-3.0s in, while createMainWindow runs
synchronously in the same frame and its renderer dies at init 48-1373ms later.

- Persist the poison verdict the moment the probe reports it, and await the
  repair (bounded at 20s) before any window is created on a launch that already
  carries the marker.
- Do not engage the GPU safe-graphics fallback while the install-dir ACL verdict
  is poisoned or still outstanding. Safe graphics does not rescue a poisoned
  tree, and --in-process-gpu removes the GPU child, erasing the sibling-death
  evidence that identifies the shape (4 field reports landed in 'misc' this way).
- Clear the safe-graphics marker once the repair lands, so a repaired machine
  stops launching software-rendered for the rest of that build.
- Give the repair marker a bounded retry budget: it was written on failure and
  matched regardless of outcome, so one transient failure pinned a machine to
  'marker-hit' for the life of that version.

* test(windows): pin the install-dir ACL repair against the real icacls binary

* fix(windows): stop the install-DACL verdict from outliving the evidence

Adversarial review round 1. Five blocking findings, all addressed.

1. gpu-lifecycle guard had only a source grep (green with the polarity
   inverted). The stated justification -- that gpu-lifecycle's import graph
   cannot be driven in-process -- was wrong: mocking `electron` plus
   `@electron-toolkit/utils` imports it fine. Replaced with
   gpu-lifecycle-install-dir-acl-guard.test.ts, which drives the real
   handleGpuChildCrash against a stub tracker. All four cases go red when the
   guard is flipped to `if (!isInstallDirAclSuspect())`.

2. A clean probe verdict retired the on-disk marker but not the in-memory
   `poison` verdict, so a machine the probe just proved healthy kept
   suppressing the GPU safe-graphics fallback and kept the dialog accusing the
   install folder -- permanently, since a `status:'failed'` probe deliberately
   keeps the marker. A positive clean reading now latches `installDirReadClean`,
   drops the verdict, and outranks a repair result that lands after it (a
   'failed' from a repair with nothing left to fix must not re-accuse).
   'repaired' is kept: it is not a contradiction and it is what tells the user
   to reload.

3. `noteWindowsInstallDirAclProbePending()` ran on every `openMainWindow` while
   the probe is once-per-process, so every tray/second-instance reopen armed a
   15s window in which `recordGpuCrash` was never called at all -- on healthy
   machines. `probeWindowsInstallDirAcl` now reports whether THIS call
   dispatched, and only a dispatch arms the grace window.

4. The pre-window ordering guarantee was defeatable and untested.
   `focusExistingMainWindow` opens a window whenever there is none and the app
   is ready -- true for the whole 20s gate, which is exactly when a user
   double-clicks the shortcut again. Added a `canOpenWindow` seam (same
   'pending' semantics as the existing `!app.isReady()` case) wired to
   `isBlockingInstallDirAclRepairInFlight()`, plus
   windows-install-dir-acl-startup-wiring.test.ts pinning the await ahead of
   both window-creation paths and both new call sites.

5. windows-install-dir-acl-repair.win32.test.ts was absent from the pr.yml
   win32 allowlist, so it ran nowhere. Added.

Also from the non-blocking list:
- The repair no longer clears a `userConfirmed: true` safe-graphics marker;
  "keep safe graphics" is a user choice, not Orca's automatic latch.
- `repairWindowsInstallDirPackageAcl` now reports its dispatch too, so a second
  entry into the gate resolves immediately instead of eating the full 20s
  budget waiting on an `onDone` that is never coming.
- The gate is wrapped in try/catch/finally, matching the contract the probe
  documents as mandatory for anything upstream of window creation.

Rebutted, not applied:
- "Gate should be conditioned on app.isPackaged." A dev launch only carries the
  poison marker if a dev launch actually probed that tree and found the
  signature, in which case the dev renderer is dying the same way and the
  repair is exactly what is needed. The adjacent `isPackaged` check guards a
  packaged-only early-window optimisation, not a correctness boundary.
- "Fold the poison marker into the repair marker's `outcome`." They answer
  different questions with different lifetimes. The repair marker is a retry
  budget (`attempts >= 3` disables the repair for that version) and is never
  cleared; the poison marker is cleared by a successful repair and by a clean
  probe. A `'pending'` outcome written before the attempt would bump `attempts`,
  so three launches killed mid-repair would permanently disable a repair that
  never once ran icacls to completion.

* fix(windows): keep counting GPU crashes while the install-DACL verdict is pending

Adversarial review round 2. Both blocking findings addressed.

1. handleGpuChildCrash early-returned on isInstallDirAclSuspect() BEFORE
   recordGpuCrash, so the crash left no trace in the 30s rolling window. The
   suspect window is armed on every win32 non-serve launch, and the field
   bundles put it at 0.8-1.7s after main_window_created on hosts whose DACL is
   clean (matchesPoisonSignature=false) -- squarely inside the 2.1-6.2s
   bad-driver bursts this repo already pinned in
   gpu-crash-fallback-field-sessions.test.ts. A healthy machine with a failing
   driver could lose an entire coalesced burst and never engage safe graphics.

   The crash is now always recorded; only the engagement consults the verdict,
   and it waits for the verdict rather than acting on the suspicion
   (waitForInstallDirAclVerdict, resolved by the probe's onDone or by the
   existing 15s grace, whichever lands first).

   Deviation from the review's suggested shape, deliberately: awaiting the
   verdict before persisting anything reintroduces the exact race
   gpu-fallback-engagement.ts documents -- Chromium aborts the whole browser
   process on the 6th GPU crash, ~1.3s after the 3rd, which is less than the
   probe takes to answer. So the unconfirmed marker is written up front and
   withdrawn if the verdict comes back poisoned. A machine killed mid-wait
   still comes back software-rendered, and its marker is unconfirmed, which is
   the state the repair's own clear already retires.

   gpu-lifecycle-install-dir-acl-guard.test.ts now drives the real
   GpuCrashFallbackTracker and the real engagement path (the restart prompt
   firing is the signal) instead of a stub tracker, and covers the case the
   previous suite could not express: a burst that lands entirely inside the
   pending window still engages once the probe reports clean. Four reverts go
   red -- restoring the pre-record guard (2 tests), dropping the wait, dropping
   the post-wait re-check, and dropping the pre-wait marker write (2 tests).

2. The round-1 evidence block quoted commits, a test name and pass counts that
   no longer exist, and its real-icacls Windows run predated the commit that
   rewrote the gate. Re-run at this commit; counts and the live-Windows result
   are restated in the handoff rather than carried forward.

Also from the non-blocking list:
- 'marker-hit' conflated "already repaired" with "retry budget spent", because
  hasMarkerFor matches outcome === 'repaired' too. The result now carries
  alreadyRepaired, and the recovery maps that to stage 'repaired' -- so a launch
  killed between a successful repair and its marker clear no longer tells the
  user the folder needs an administrator, no longer latches
  isInstallDirAclSuspect() for the session, and does retire the poison marker.

Not applied, with reasoning:
- "clearGpuFallbackMarker narrowed to userConfirmed === false leaves the target
  population software-rendered after a repair." The summary was overstated and
  is corrected, but the narrowing stands: a userConfirmed marker now requires a
  clean DACL verdict, because the restart prompt that writes it is exactly what
  the gate above withholds while the install is a suspect. The population this
  family targets can no longer reach confirmMarker while poisoned.
- "writeInstallDirAclPoisonMarker re-stamps on a budget-exhausted machine
  forever." True, but on that machine the tree really is still poisoned and the
  gate resolves immediately ('skipped', no icacls spawn, no 20s wait), so the
  marker is telling the truth. Retiring it would be wrong; only a clean probe
  reading should.

* fix(windows): register the real-icacls spec and stop its teardown racing icacls

Two ratchets were red:
- windows-lane-tree-removal-boundary: the win32 spec's afterAll used raw
  rmSync on a tree two icacls.exe children had just rewritten DACLs on, which
  is the EPERM race removeTreeSync exists for.
- win32-test-lane-registration: the spec was in the pr.yml argv but not in
  WINDOWS_PACKAGE_TESTS, so a future diff touching only test files would not
  select package_windows and the spec would self-skip on ubuntu and report
  success.

* fix(windows): re-arm the GPU fallback latch when the install-DACL verdict withholds it

recordGpuCrash reports the threshold crossing exactly once and latches `engaged`.
handleGpuChildCrash consumes that report before consulting the DACL verdict, and
installDirAclClearsGpuFallback then discards it — so nothing could ever engage
safe graphics again in that process. A machine whose tree the repair fixes and
whose driver is genuinely broken stayed hardware-accelerated through an unbounded
crash loop, with no prompt and no marker.

disengage() releases only the one-shot latch; the crash window is untouched, so a
real driver burst is still never erased. Test is RED without the re-arm.

* fix(windows): keep the safe-graphics marker while an install-DACL repair is in flight

The gate dispatches a repair without arming the probe clock, so
waitForInstallDirAclVerdict() returns immediately and the withdrawal deleted the
marker inside Chromium's FATAL window (crash 6 lands ~1.3s after crash 3, well
inside the 20s gate). The process then died mid-repair, spent no attempt, and
relaunched hardware accelerated into the same gate — spawning the same GPU
children, FATALing again, forever.

Hold the marker while poison.stage is 'pending' so that launch comes back
software rendered and the next gate runs to completion. Still not engaged this
launch, so --in-process-gpu does not erase the sibling-death evidence. A
terminal verdict has no next step to rescue, so it still withdraws. Both new
tests are RED without the retention.

* fix(windows): stop a repaired marker outranking a fresh poison verdict

The probe reads the install DACL and finds it poisoned; `startRepair` dispatches;
`markerHitFor` sees a repair marker recording `outcome: 'repaired'` for the same
installDir+appVersion and reports `alreadyRepaired`, which the recovery module maps
to stage 'repaired'. So the launch that just proved the tree poisoned runs no icacls,
deletes the poison marker that arms the next launch's pre-window gate, clears the
suspect flag so `--in-process-gpu` can engage on a tree safe graphics cannot rescue,
and tells the user "Orca repaired the permissions."

Reachable whenever the tree is re-poisoned after one successful repair of the same
version, and whenever a repair reports success without clearing the tree — the silent
icacls no-op this module exists to document.

A DACL reading taken this launch now outranks the marker: `probeConfirmedPoisoned`
stops `outcome: 'repaired'` short-circuiting the repair. The attempt budget still
bounds it, so an unrepairable tree does not re-spawn icacls forever. The pre-window
gate does not set the flag — it acts on a marker from an earlier launch, not on
evidence of its own, so a recorded repair still outranks it there.

Also drives the GPU-fallback re-arm test through a repair that actually completes
'repaired', rather than a later clean probe, which is the route the review exercised.

* fix(windows): make the pre-window ACL gate act on the poison evidence it fired on

The gate fired on a poison marker — an earlier launch's DACL reading that nothing has
retired — but withheld `probeConfirmedPoisoned` from the repair, so a repair marker
recording an older success still short-circuited it. On the three-launch shape the gate
exists for (repair succeeds; tree is re-poisoned; the next launch's probe records the
poison but dies before writing its repair marker) the gate ran no icacls, deleted the
poison marker that arms every later gate, un-suspected the tree so --in-process-gpu could
engage, and told the user "Orca repaired the permissions." `applyInstallDirAclProbeVerdict`
then swallowed that launch's own reading behind `if (poison) return`.

Both callers of `startRepair` hold outstanding poison evidence, so the flag is now
unconditional (renamed `poisonEvidenceOutstanding`) and `marker-hit` means only that the
attempt budget is spent. The probe guard is narrowed to an in-flight gate repair: a reading
taken after the gate finished re-arms the poison marker and downgrades a claimed repair.

Also: withholding safe graphics now ends with the repair budget. A machine whose attempts
are spent while the signature persists was denied safe graphics on every launch for the
life of that appVersion — and had its marker deleted each time — including the healthy
installs the probe's flag-blind ACE match over-matches, where the driver really is broken.

Non-blocking, same lane: re-read `isQuitting` after the up-to-15s verdict wait, and skip
the recovered-launch prompt when the ACL gate retired the marker read before whenReady.

* fix(windows): stop a timed-out gate repair outranking a later poison reading

The gate's 20s budget expires while icacls runs on under its own 120s cap, so
the probe can read the tree poisoned while that repair is still in flight. Its
success claim then deleted the poison marker, un-suspected the tree and told the
user their permissions were fixed. The reading is now latched and outranks it.

* fix(windows): stop a gate repair claim pre-empting this launch's probe reading

Round-7 adversarial findings, both driven against the real modules:

- isInstallDirAclSuspect returned false the moment the pre-window gate set
  stage 'repaired', short-circuiting ahead of the probe-pending grace check.
  The GPU children die 48-1373ms after window creation while the probe
  answers 0.9-3.0s in, so an icacls that silently no-opped (exit 0, tree
  untouched) opened exactly that interval to --in-process-gpu on a
  still-poisoned tree - and a 'keep safe graphics' answer then pinned a
  userConfirmed marker no later repair may clear, with the poison marker
  already deleted so no later launch gates. The claim now stays provisional
  until this launch's probe corroborates it or the grace window lapses.

- A probe reading that disproves a 'repaired' claim re-armed the poison
  marker but never restored the unconfirmed safe-graphics marker the claim
  had cleared, so the next launch relaunched hardware-accelerated into the
  re-armed gate. The clear is now captured and handed back on disproof.

* test(windows): pin the nested and update-inherited grants against real icacls

The live spec asserted the grant landed on the root-level module file only.
It now also pins that the flagless /T pass reaches a nested file carrying
its own protected DACL (the shape app.asar.unpacked and node_modules have),
and that a file written after the repair inherits the (OI)(CI) root grant -
the stated reason that grant form exists.

* fix(windows): keep the recovered-launch prompt silent while the tree is the suspect

Round-8 fresh-eyes finding, driven against the real modules: the prompt
re-read the marker the pre-window gate may have retired, but never consulted
isInstallDirAclSuspect() - so after a FAILED gate (tree still a live suspect,
window blank behind the 10s reveal fallback, Keep as both defaultId and
cancelId) a 'keep it' answer pinned a userConfirmed marker no later repair
may clear, on the exact victim class the repair cannot help. The guard now
covers both gate outcomes; staying silent leaves the marker unconfirmed,
which a successful repair still retires.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
Co-authored-by: OrcaWin <alpha-eng@stably.ai>
2026-09-03 21:39:34 -07:00
Neil 8463dcb7b9 fix(terminal): make wrapped-line search rewind iterative and bound its scans (#18402)
Patches @xterm/addon-search so one very long un-newlined line no longer overflows the stack, freezes the renderer, or goes unsearched. Submitted upstream as xtermjs/xterm.js#6149 (issue #6148); drop the patch once a release ships it. See the PR for measurements and the differential fuzz.
2026-09-03 17:59:16 -07:00
Brennan BensonandMerge Sim 98e77ef1a7 feat(mobile): structured native Codex chat (#18074)
* feat(mobile): finalize structured native Codex chat

* fix(mobile): close structured chat lifecycle gaps

* wip(mobile): fence stale structured inventory and bound operation-id retention

Fence local structured-session inventory and subscription responses with a
sync generation so a toggle-off clear, reconnect restore, or retry cannot
apply a mirror from a superseded instance. Bound mobile ambiguous
operation-ID retention at 128 with unmount cleanup.

Staged on the reconcile branch only: the sync module is now 312 lines and
needs a real split before this can reach the PR head.

* fix(ci): split the structured session-tabs sync and give static analysis mobile types

The local structured session-tabs sync module outgrew the 300-line cap once it
took on generation fencing, so split it along its real seams instead of raising
the cap: the generation/cursor fence, snapshot projection, snapshot apply,
inventory refresh, and the subscription loop. The original path stays as a
barrel so no importer moves.

Repoint the host-session-mirror settle census at the apply module, which owns
two receipts now — the snapshot it mirrors in, and the toggle-off teardown that
retracts what it published. The teardown receipt is named rather than anonymous
so the pin says which direction it settles.

The changed-code quality gate lints mobile files and resolves their types from
mobile/node_modules, but mobile is a separate pnpm project that the root install
never populates, so every mobile type degraded to an `error` type and the gate
reported phantom findings. Install mobile dependencies in static analysis when
the diff touches mobile, gated on a new classifier output.

* fix(mobile): let a slow capability handshake still reach connected

The mobile capability update is an advisory whose result is discarded, yet an
unanswered one was fatal while an explicit rejection was tolerated. A 5s timeout
on the direct client force-closed the socket, and on the relay path it failed
`confirmResume` before `connected` was ever published, so a consistently slow
link redialled forever. Both paths now share one helper that settles every
ambiguous outcome (timeout, mid-flight drop) like a rejection and rejects only
when the frame never reached the wire — the one case nothing else recovers from,
since the socket's own desync force-close is gated on already being connected.
The generation guard still keeps a replaced session from connecting.

Retained structured-session operation ids were capped at 128 with oldest-first
eviction, but every retained id belongs to a send whose outcome is unknown, so
eviction turned a user's retry into a second message on the host. Bound the map
by expiry against the id's own embedded timestamp instead, mirroring the host's
operation ledger, so no id is released while the host would still honour it.

Also give the mobile CI install the root install's lockfile drift guard (mobile's
lockfile carries patchedDependencies a silent rewrite would drop), gate
mobile_dependencies on should_run, and key the pnpm store cache on both lockfiles.

* refactor(mobile): extract the relay pending-request registry

The merge composed two independently-sized changes — this branch's capability
handshake settle and main's dial-stage tracking — pushing the relay session file
to 304 lines against a 300 cap. Neither side broke it alone.

Move the in-flight request registry (id generation, tracking, settlement, and
reject-all with its delivery-ambiguity marking) into RelayPendingRequests,
matching the existing collaborator pattern alongside RelayDialStageTracker and
RpcSessionLivenessWatchdog. No behavior change.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-03 15:19:26 -07:00
Neil 968dbd905f perf(renderer): take the English catalog and the xterm WebGL addon off the boot graph (#18326)
* perf(renderer): take the English catalog, xterm WebGL addon and emoji data off the boot graph

The renderer's boot graph — the entry chunk plus its 331 modulepreload links,
all fetched and evaluated before first paint — carried three payloads nothing
needs at that moment.

`en.json` (644 KB) was an eager i18next resource, but every renderer string
goes through `translate(key, fallback)` and `en` resolves that inline default,
so most of the catalog was dead weight. The renderer now bundles a generated
`en-runtime-required.json` holding only the 2,583 of 13,828 entries a default
cannot reproduce: plural-suffixed keys, keys whose catalog value differs from a
call site's default, and keys no call site references with a literal default.
`en.json` stays the translator source and the input to the four lazy catalogs.

`@xterm/addon-webgl` (243.6 KB) and `emojibase-data` (170 KB) are now primed
right after the React root renders instead of statically imported. The load
stays eager and `attachWebgl` stays synchronous — it reads the resolved
constructor — so no terminal ever falls back to the DOM renderer for a frame.

`isPluginPanelTabKey`/`isQualifiedPluginKey` move to schema-free sibling
modules, re-exported from `plugin-manifest.ts`. This evicts the plugin manifest
schema graph from the boot chunk but measures ~0 KB, because six other shared
modules still put zod on the boot path.

Boot graph: 332 chunks / 5107.2 KB -> 336 chunks / 4161.5 KB (-945.7 KB, -18.5%).

A new ratchet parses the built index.html and fails if `en.json`,
`@xterm/addon-webgl` or `emojibase-data` is preloaded again; it runs at the end
of every `build:electron-vite`.

* chore(i18n): pin the generated English subset to LF and mark it generated

* fix(i18n): make the runtime-catalog gate merge-robust and prime emoji data in tests

CI builds the merge of a PR with main, so a byte-for-byte comparison against a
committed generated file fails the moment any unrelated PR adds a translate()
call — which is what happened here. The check now asserts the property that
actually matters instead of byte equality: every runtime-required entry is
shipped, and nothing shipped disagrees with en.json. Entries that stopped being
required are dead weight, never a wrong string, so they are reported and
tolerated. Failures now name the offending keys rather than saying "stale".

The generator itself was already deterministic (plain code-unit sort, no
locale collation, order-independent set construction); a test now pins that a
reversed call-site walk produces byte-identical output.

Test fixes for the catalog prune and the deferred emoji load:
- browser-search / NativeChatSupportedAgents asserted key presence on the
  renderer's runtime resource. The durable contract is en.json — the renderer
  deliberately no longer bundles entries a call site default reproduces — so
  they assert against the translator catalog.
- Four emoji tests typed a shortcode in the same tick as mount, before the
  catalog the hook primes on mount resolves. Not reachable by a human; the
  tests now await the prime.

* revert(renderer): keep the emoji shortcode catalog statically imported

Deferring emojibase-data introduced a window that did not exist before: until
the dynamic import settled, getPrimedEmojiShortcodeEntries returned [], so
exactShortcodeIndex built an empty map and replaceCompletedWorkspaceEmojiShortcode
returned null — leaving a typed `:wink:` in the field literally, and persisting
it as the workspace display name.

Pre-change the shared catalog was statically imported, so the first call at any
tick returned full data. The window is reachable by anything that dispatches
input in the same task as the field's mount effect — Playwright/CDP in the e2e
suite and agent automation both do, and the WorktreeMetaDialog test failure was
exactly that, producing 'Feature 😉' instead of 'Feature 😉'.

Nothing that resolves a shortcode can be async without that race, and a wrong
persisted name is not an acceptable trade for 166.7 KB, so the deferral is
reverted rather than papered over in the tests. The boot-graph ratchet drops
its emojibase-data probe and records why.

Boot graph: 5108.9 KB -> 4329.9 KB (-779.0 KB, -15.2%), down from -945.7 KB.

* fix(terminal): make the deferred WebGL addon load recoverable and refit on late attach

Two defects the deferral introduced, neither possible with a static import.

A failed load latched the DOM renderer for the whole session. `.then(onOk,
onError)` settles fulfilled, so the memoized promise was cached forever with a
null constructor: attachWebgl's re-prime got the cached promise back, and
resetTerminalWebglSuggestion — the documented "GPU setting changed, retry" path
— could not clear it either. The rejection path now clears the memo, latches the
queued panes the way a failed construction does so they retry at a recovery
boundary rather than every frame, and caps attempts so a genuinely missing chunk
is not re-fetched forever. The recovery boundary re-arms it.

The queued-attach drain skipped the refit. Every other late-attach path pairs
attach with a refit because the grid was measured under DOM cell metrics and
WebGL floors the device cell width. Post-deferral, openTerminal's attachWebgl
queued and returned, the initial fit rAF then measured DOM metrics and sized the
PTY from them, and the addon attached with no refit — a persistently narrow PTY
and an unpainted right gutter, not a one-frame flicker. Both paths now go
through one attachWebglAndRefit pairing so they cannot diverge again.

Regression tests cover both, and each was verified to fail without its fix.

The addon-load state machine moves to terminal-webgl-addon-loader.ts and the
viewport presentation helpers to pane-viewport-present.ts, keeping
pane-webgl-renderer.ts under the 300-line budget without a suppression.
2026-09-03 00:26:30 -07:00
Jinjing 6c66487fca ci: checkout PR head for reusable E2E (#18230) 2026-09-02 11:20:15 -07:00
Neil f37d2fec97 fix(linux): land the reviewed Linux packaging stack on main (#18100)
* fix(linux): give the CLI one entrypoint by extracting the AppImage once

* refactor(linux): trim AppImage CLI registration seams

* test(cli): assert registration lock serialization

* fix(linux): fence AppImage terminal shim mounts

* fix(linux): accept extracted AppImage runtimes with APPDIR only

* docs(linux): make headless AppImage extraction runnable

* refactor(linux): import bundled launcher directly

* fix(linux): reclaim superseded AppImage payloads and packaged symlinks

Pruning removed 3215 of 3216 files from a superseded generation and always
stranded resources/app.asar, leaking ~105 MB per version update. Electron's
asar shim reports a *.asar file as a directory, so the recursive remove tried
to rmdir a real file and failed with ENOTEMPTY; the .catch(() => {}) hid it.
Reproduced end to end on Ubuntu 24.04: 519M -> 623M across one update, and
519M again once the payload is actually reclaimed.

removeExtractedAppImagePayload holds process.noAsar for the removal, counted
so overlapping removals cannot hand the shim back early, and the prune site
now warns with the path instead of swallowing the rejection. All three
removal sites use it -- staging cleanup and displaced roots leaked the same
way.

Also reclaim symlinks left by a packaged deb/rpm install, which the
extracted-cache-only rule turned into a hard conflict on a deb -> AppImage
migration, and name the remedy in the conflict error.

* fix(linux): bound the CLI registration lock wait

`retries: 1000` caps the attempt count, not elapsed time, so at up to 1s per
attempt an IPC-driven registration could hang ~16 minutes against a wedged
holder with no feedback.

A legitimate holder is bounded by the extraction timeout, so wait that plus
slack and then fail with a message naming the lock file, rather than hanging.
`maxRetryTime` is forwarded verbatim to the `retry` package by proper-lockfile.

* fix(linux): stop re-extracting the AppImage on inode metadata churn

The extracted-payload cache key hashed ctime alongside dev/ino/size/mtime.
ctime moves on any inode metadata write -- `chmod +x`, which every AppImage
user is told to run, plus `chown`, an ACL or SELinux relabel, and a backup
restore -- none of which alter a byte of the payload.

Measured on Ubuntu 24.04: `chmod +x` leaves dev, ino, size and mtime
identical and moves ctime alone, so the key changed and the next launch paid
a full ~519 MB re-extraction and a multi-second stall to rebuild a payload it
already had, then pruned the old generation.

Key on content identity instead. An in-place content change moves mtime and
almost always size; a replacement moves the inode. The existing
replace-in-place test still passes.

* fix(linux): stop CLI commands from falling through to Chromium startup

* refactor(cli): remove redundant command membership check

* test(cli): cover command-named project selectors

* fix(cli): redirect the open-url command before startup

* test(linux): cover AUR serve wrapper flags

* fix(linux): tighten CLI launch detection

* fix(linux): respect CLI flag value boundaries

* fix(linux): strip injected Chromium switches from CLI args

* fix(linux): report a missing display instead of dying in uv_close

* refactor(linux): read display locks without a preflight race

* fix(linux): preserve unverified external displays

* chore: format reliability gate manifest

* test(packaging): split runtime resource checks

* fix(linux): fail serve when no display is available

* fix(linux): do not treat a lockless X socket as a dead display

An X server writes its lock beside its socket and both survive a crash
(verified against Xvfb under SIGKILL), so a socket with no lock was never
left by a crashed server. It is an endpoint published from elsewhere: a
container bind-mounting only /tmp/.X11-unix, WSLg, or a foreign PID
namespace. Declaring those dead made the desktop gate exit(1) on displays
that work, with no workaround, and the serve gate refuse to start.

Liveness now splits by ownership. A foreign DISPLAY trusts a lockless
socket; Orca's own :99 does not, because removeStaleDisplayArtifacts
unlinks the lock before the socket and so manufactures that state itself --
adopting it would resurrect the orphan-socket bug and stop the cleanup from
self-healing. The stale-lock rejection is unchanged.

Also correct four doc statements this behaviour falsified.

* fix(linux): fail closed when a stale socket blocks the Xvfb rebind

Readiness only checked that /tmp/.X11-unix/X99 exists. A stale socket we
could not unlink still exists after our own Xvfb refused to bind, so Orca set
DISPLAY to a dead server and Chromium died in Ozone init.

Measured on Ubuntu 24.04 against the pre-fix build: with a leftover :99
socket and no lock, serve exits 139 (SIGSEGV), the socket inode is unchanged
before and after, and no lock is recreated -- it neither cleaned up nor
respawned. To a user that is a crash, not a misconfiguration.

This is reachable in the documented topology, where orca-xvfb.service has no
User= and runs as root while serve runs as User=orca: /tmp is sticky, so the
orca uid cannot unlink a root-owned socket, rmSync fails, and Xvfb exits with
the display already active.

Readiness now requires the display to actually be live -- our socket plus a
lock naming a running process -- so the same state reports an unusable
display and exits 1 with the existing diagnosis.

* fix(linux): recognise abstract X sockets and inherited Wayland fds

Two display setups this gate could not prove were refused outright, and on the
desktop path that is app.exit(1) with no workaround.

An X server may bind only the abstract namespace (`@/tmp/.X11-unix/X0`), which
leaves no filesystem socket to stat. Abstract addresses are kernel-owned and
vanish the moment the owner exits, so an entry in /proc/net/unix is proof of a
live server -- no lock file needed and no stale entry possible. Verified on
Ubuntu 24.04, where 139 such addresses were present.

WAYLAND_SOCKET is an already-connected fd handed over by the compositor, so
there is no path to stat and WAYLAND_DISPLAY may be unset entirely. Its
presence is the display.

Both are consulted only after the filesystem-socket check fails, so no
existing verdict changes.

* fix(linux): never treat Orca's own display number as a foreign endpoint

Recognising a lockless X socket as live is correct for an endpoint published
from elsewhere -- a container bind mount, WSLg -- because an X server writes
its lock beside its socket and both survive a crash. It is wrong for
VIRTUAL_DISPLAY_NUMBER, because Orca's own teardown unlinks the lock before
the socket and so manufactures that exact state.

The managed branch was already strict, but a caller that sets DISPLAY=:99
explicitly takes the foreign path and skipped it, accepting a dead display
left by Orca's own interrupted cleanup. Route the managed number through the
strict probe on both paths.

Found by an adversarial audit of the asymmetry introduced earlier in this
branch; the documented systemd topology is unaffected because its Xvfb writes
a real lock.

* test(linux): add a packaged-artifact contract for the CLI launch paths

* test(linux): avoid buffered serve readiness detection

* test(linux): signal AppImage serve owner directly

* test(linux): tolerate readiness timeout boundary

* test(linux): add startup margin to shutdown oracle

* ci(linux): give package contracts timeout headroom

* fix(ci): route all Linux packaging contract changes

* test(linux): poll shutdown readiness without tail leaks

* test(linux): bound shutdown cleanup grace

* test(linux): assert on CLI output, not the harness's own control lines

run-cli-case.sh echoes `RESULT status=N case=<name>`, and the two cases named
*-skills asserted `expectOutput: 'skills'`. That substring was satisfied by
the case name in the harness's own line, so 2 of 8 cases asserted nothing
about the command -- gutting `skills` entirely would still have gone green.

Control lines are now excluded before matching, and both cases assert the
rendered help header, which only real help output produces. Verified on an
Ubuntu 24.04 host: 8/8 still pass against a stack-tip AppImage.

Also register the gate in reliability-gates.jsonc, which #15085 added a CI
Docker gate without. Red/green is recorded from a stock release AppImage
failing 4 of 8, three of them at status 133 (SIGTRAP).

* fix(linux): require static AppImage runtimes (#17319)

* test(linux): reject a wrong-architecture native binary at packaging time

Cross-building the arm64 slice on an x64 host silently packed an x86-64
`pty.node` -- the rebuild logged "Forcing native rebuild for linux-arm64" and
shipped the host's binary anyway. Every gate here inspects symbol versions,
which are perfectly valid on the wrong architecture, so nothing noticed.

Observed on a Raspberry Pi 5: the packaged app loaded, then failed with
"Failed to load native module: pty.node", and the launch contract reported
3 of 8 cases crashed rather than naming the cause. Swapping in the aarch64
`pty.node` took the same build to 8/8.

Compare ELF `e_machine` against the slice being packaged and fail with the
offending path. Checked before the glibc pass, because a wrong-architecture
binary's symbol versions are valid but meaningless and would send the reader
down the wrong path.

Release CI builds arm64 on a native runner, so this guards local and future
cross-builds rather than a shipped artifact.

* test(linux): judge per-arch vendored binaries against their own path

The first CI run of the architecture gate failed the x64 package job on
`@parcel/watcher-linux-arm64-glibc/watcher.node`. That binary is arm64 on
purpose: the package ships every architecture and its loader picks the match,
so its presence in an x64 build is correct.

Judge a binary against the architecture its own path names, falling back to
the slice when the path names none. That keeps the case this gate exists for
-- `bin/linux-arm64-*/node-pty.node` holding an x86-64 binary, which is what
shipped to a Raspberry Pi 5 -- while letting multi-arch dependencies through.

Dry-run over the real dependency tree flags nothing for either target arch.

* fix(linux): move deb/rpm update installation outside Orca (#17318)

* fix(linux): complete deb/rpm package metadata

* fix(linux): preserve CLI link during package upgrades

* docs(linux): document local RPM build prerequisites

* fix(linux): move deb/rpm update installation outside Orca

* fix(updater): preserve Linux recovery across stale events

* fix(updater): fence stale downloaded events by active target

* fix(updater): preserve active Linux package recovery

* test(linux): keep workflow order assertion in scope

* test(updater): assert stale recovery stays silent

* fix(updater): preserve Linux package recovery after checks

* refactor(updater): keep Linux marker message with status

* fix(linux): describe the right manual update path for deb/rpm hosts

A remote host installed from .deb or .rpm now reports
manual-service-update-required, and the guidance told the operator to
"update through the service manager that starts this server" -- which is
correct for unsupported-headless-serve but wrong for a package install,
where nothing about the remedy involves the service manager.

Say both, keyed on how the host was installed.

* docs(linux): document orcad update restart safety

* docs(linux): scope restart census omissions

* docs(linux): use absolute service CLI launcher

* fix(serve): validate in-process serve options before startup (#17683)

* fix(linux): stop offering updates a distro-managed install cannot apply (#17918)

Closes #17702.

The resources/package-type marker is authoritative but never checked against
the host, so any repackager that unpacks Orca's .deb -- AUR, Nix, a container
rebuild -- inherits `deb` verbatim. Install feasibility was then computed
after a ~165 MB download, so those users got check -> download -> a card
promising an install command -> a dead end.

Validate the marker against the host: a deb/rpm marker with no matching
package manager in the trusted directories means a package manager owns this
install. This reuses the exact lists and resolver that
buildLinuxPackageInstallCommand already loops over, so a false positive is
impossible by construction -- any host flagged here would have failed with
no-package-manager after the download anyway. The gate only moves that
verdict earlier. Verified across Debian 12, Ubuntu 24.04, Arch, Fedora 40 and
openSUSE Leap: no false positive on a real deb host, correct on every
repackaging host.

The release is still reported, because the user does want to know 1.4.194
exists and to update through their distro; only the download path is closed.
`externallyManaged` is an additive optional field on the existing `available`
status, so older paired clients decode it unchanged. downloadUpdate() refuses
authoritatively, since main owns this verdict rather than the card, and
unwinds any pinned-build state first -- a Linux pinned jump resolves to
'release', and stranding isPinnedBuildActive would silently kill every
background check for the rest of the process.

Note the fix the issue suggests cannot work: electron-updater builds a
PacmanUpdater whose doDownloadUpdate looks for a .pacman asset Orca does not
publish, then dereferences undefined.

* style(cli): restore prettier wrapping on install error copy

* test(linux): re-pin the child-process ratchets and the batch-shim allowlist after the merge
2026-09-02 03:08:01 -07:00
Neil a5796ec8eb refactor(runtime): split OrcaRuntimeService and compatibility tests (#17605)
* refactor(runtime): split OrcaRuntimeService into focused modules

* test(runtime): cover admission tiers and strict worktree reconciliation

* fix(runtime): preserve owner and structured session visibility

* fix(runtime): port post-extraction compatibility fixes

* fix(runtime): preserve skill-share cancellation barrier

* test(runtime): update identity inventory after extraction

* fix(runtime): preserve hook transport environment cleanup

* fix(runtime): consolidate idle probe imports

* test(runtime): retire split file process allowlist entry

* fix(runtime): route child process types through shared boundary

* test(runtime): preserve worktree host metadata precedence

* fix(runtime): update extracted test seams

* fix(runtime): gate the split's ts-nocheck set and restore the stop-confirmed contract

Audit follow-ups for the OrcaRuntimeService split:

- Freeze the 171 @ts-nocheck files behind a ratchet so no new file can disable
  type checking. The split's linear mixin chain cannot express forward
  references yet, so the existing suppressions are grandfathered; the baseline
  may only shrink.
- Drop the stray @ts-nocheck at the end of orca-runtime-get-status.ts. It sat
  after the first statement, where TypeScript ignores it, so the module was
  already checked.
- Restore `retireRejectedPty(ptyId, stopConfirmed: boolean)` as a required
  argument. The split widened it to optional and patched the resulting error
  with `stopConfirmed === true`; an omitted argument would have silently taken
  the unverified-stop path instead of failing to compile.
- Guard that every orca-runtime-tests fragment is imported by the compatibility
  entrypoint. The fragments are .spec.ts, which no Vitest include glob matches,
  so one left out of the list would silently stop running.

* fix(runtime): restore four behaviors the OrcaRuntimeService split dropped

Audit findings against the refactor's true base (ad5ba2572e):

- retirePtyAgentLaunchAuthority collected pane keys after deleting the
  restored-authority receipt instead of before it. collectPaneKeysForPty reads
  that receipt, so a receipt-only pane lost its key and never had its agent-hook
  compatibility authority retired. on-pty-exit.ts already carried a comment
  naming this exact invariant.
- The PTY-exit path kept orchestrationMailboxNotifications.retirePty but lost
  the loop that schedules a debounced mail-pointer repoint for the dead pty's
  terminal handle and any run bound to its panes. Restores the schedule call
  count to 7, matching base.
- subscribeToPtyExit lost isPtyKnownExited's leaf fallback and its
  post-registration lifecycle-generation recheck. leavesByPtyId is rebuilt from
  the renderer graph independently of ptysById, so a leaf can outlive its pty
  record; without the fallback a caller waiting on an already-dead pty never
  gets released.
- The chain root declared `[key: string]: unknown`, which base had nowhere. It
  leaked through the exported runtime type into every consumer, so any misspelled
  member access typechecked as unknown instead of erroring, and it accounted for
  957 of the suppressed errors. Removing it costs zero type errors.

* fix(runtime): restore escalation prose and unscoped automation publication

Two more behaviors the split dropped, each with a regression test that fails
against the pre-fix code:

- The worker-exit escalation stopped deriving its title through
  buildOrchestrationTaskDisplayMetadata and inlined `task.spec` instead. That
  ignored an explicit task_title, dropped the single-line normalization and the
  80-character bound, and turned the no-spec case into a quoted, duplicated id.
  A multi-paragraph spec landed verbatim in the coordinator's banner. The
  existing 11 tests all use short single-line specs, where the derived title and
  the raw spec are identical, so none of them could see it.
  Also reverts an added `if (!handle) return` guard: the dispatch lookup is
  deliberately keyed on the pane as well, because a reminted handle no longer
  matches the row while the pane identity outlives the remint.
- updateAutomation stopped going through automationChangePublications and
  published `source` unconditionally while gating the fallback on a non-null
  destination. A destination the store can no longer name then published only
  the stale source, so subscribers scoped elsewhere kept rendering a row that
  had left them — the exact case the helper documents. The helper had been left
  with zero callers; all three sites use it again.

* fix(skills): stop swallowing lookup errors and hard-erroring on non-ssh hosts

Follow-ups from auditing the skill install path against the refactor's base:

- resolveWorktree wrapped showManagedWorktree in `.catch(() => null)`, so a
  transient git or IO failure surfaced to the user as
  skill-install-workspace-not-found with the real cause discarded. Errors
  propagate again; a genuine id mismatch still returns null.
- resolveSkillSshTarget threw skill-install-workspace-host-unavailable when the
  execution host was neither local nor ssh, on both the repo and folder
  branches. Base gated these on connectionId, so a runtime-owned repo simply
  was not an SSH install and fell through to the local path. Both return null
  again, and the error code the split invented is now unreferenced.
- listManagedSkillInstalls awaited the receipt walk and the worktree resolve in
  sequence. They are independent and either can hit disk, WSL, or an SSH scan,
  so Promise.all is restored.

Deliberately unchanged: resolving the worktree through listResolvedWorktrees
rather than showManagedWorktree, which disambiguates a worktree id colliding
across hosts and is covered by its own test, and the SSH-folder
skill-install-ssh-dispatch-required throw, which matches the repo branch.

* fix(runtime): merge duplicate worktree-logic imports

The #17448 port added a third import from ../ipc/worktree-logic, which the
code-quality oxlint config rejects under --deny-warnings. Plain oxlint does not
flag it, so it only surfaced in CI's static analysis job.

* ci: run the ts-nocheck ratchet in PR checks

pr-workflow-lint-parity requires every leaf command in `pnpm lint` to have a
matching step in pr.yml. The ratchet was wired into lint but not the workflow,
so PR CI would not have enforced it.

* Merge remote-tracking branch 'origin/main' and retry the paired-host launch evaluate

main advanced 9 commits; none touch the orca-runtime.ts this branch splits, so
nothing needed porting.

CI failed twice on `Execution context was destroyed` thrown from
headless-paired-runtime-host's first `evaluate` after launch — a different spec
each run, which is the signature of the flake #17780 describes rather than a
regression. That commit added retryTransientMainEvaluate and adopted it in five
helpers but not this call site, even though its docblock names exactly this
case: the first evaluate after electron.launch() resolves, before the app is
ready. Wrapped it the same way.
2026-08-31 19:34:55 -07:00
OrcaWin 1ec13cbda2 Speed up CI dependency and computer E2E setup (#17513) 2026-08-30 18:19:08 -07:00
Neil e84042572c Upgrade xterm to 6.1.0-beta.303 and generate addon patches
* Upgrade xterm to 6.1.0-beta.303 and generate the addon patches

Takes the current xterm beta line: xterm 287 -> 303, addon-webgl 286 -> 299,
addon-serialize 287 -> 300, headless 302, the remaining addons -> 300, and the
same set on mobile. All four packages stamp upstream commit d3e32b3.

The reasons are upstream #6042/#6043/#6055 (a shared glyph atlas no longer
garbles sibling panes on a page merge, clear, or sampler-budget overflow) and
Note that core 303 is not image-addon-only over 302: it carries the buffer perf
work, including the new BufferLineStringCache.

addon-webgl and addon-serialize move into the patch generator
--------------------------------------------------------------
Both were hand-edited minified bundles, which is what the Known Gaps section of
docs/reference/xterm-patch-regeneration.md described. Both reproduce byte for
byte from the pinned commit, so they are now manifest entries generated from a
source patch like @xterm/xterm already was. Their sourcemaps now move with their
bundles; before this they shipped maps whose offsets did not match the code
beside them.

The webgl patch shrinks from a 1.06 MB hand-edited bundle to a 6.6 KB source
patch, because upstream took the invalidation half Orca had backported. What is
left is only what upstream still lacks: the fragment-shader else branch for a
v_texpage past the sampler budget, the clearTexture guard that no-ops once a
merged page holds index 0, spending the merge retry budget before beginFrame
latches the version it saw, and Orca's font-weight probe.

The serialize source patch is byte-for-byte the same fixes as before; upstream
changed nothing in that addon between 287 and 300.

Generator fixes, each of which failed silently
----------------------------------------------
- `--relative` was appended after the `--` separator in CHECKOUT_DIFF_FLAGS, so
  git read it as a pathspec and kept repo-root-relative paths, dropping every
  source hunk from an addon's patch.
- `git apply` run from a package subdirectory still resolves patch paths from
  the repo root, skips every hunk and exits 0. It now runs from the root with
  `--directory=<packageDir>`, and a source patch that leaves the checkout
  unchanged is a hard failure rather than an empty patch.
- An addon's own `tsgo -p .` has empty files/include and only project
  references, so it emits nothing and the addon webpack then fails on a missing
  ./out/. The root build now runs first.
- versionStampFile is optional; publish.js stamps an addon's package.json, which
  overlayBuildOutput never patches.
- On a version bump the lockfile has no entry under the new key yet, so --write
  reports the gap instead of aborting mid-run. --check still fails on it.

Adding the two addons pushed the generator and the Electron packaging contract
test over max-lines, so the patch-text helpers move to xterm-patch-text.mjs
(pure text: no checkout, no build) and the vendored-xterm assertions move out of
the packaging contract into xterm-webgl-runtime-contract.test.mjs.

Tests
-----
Four tests asserted upstream bugs that are now fixed, not Orca behaviour:

- xterm-user-scrolling-contract pinned headless and core by version string.
  Upstream bumps each package only when its own output changes, so headless 302
  and core 303 are the same source. It now asserts they share a commit.
- Five CSI 3 J assertions expected a reader stranded at the top after an erase.
  Upstream #6081 clears isUserScrolling there, so the erase releases them to the
  bottom instead. Orca's pin still lands them correctly, because its parser
  handler observes the erase before xterm's own handler runs.
- The IME transaction test hard-coded the xterm version; it now reads the
  installed package, since the point is that bundle, map and version agree.
- The Electron runtime contract asserted Orca's old clearModelGeneration. Shared
  atlas invalidation is upstream's now, so it asserts pageLayoutVersion on the
  resolved dependency, plus the Orca-only hunks on the patch.

Verified: 66,008 unit tests, mobile's 3,863, the four WebGL atlas e2e specs, and
`regenerate-xterm-patches.mjs --check` in sync on all three packages.

Left alone deliberately: resetAllTerminalWebglAtlases still fans out globally
even though clearTexture now self-heals siblings, and upstream #6068
(WebglAddon.dispose leaks the GL context) is still open.

* Drop the two unused WebGL atlas fan-out exports

resetAllTerminalWebglAtlases and presentAllTerminalPanesWithoutAtlasClear have
no callers, and had none at cadfc55102 either — the last call site went in
#6949, which routed reveal recovery through
resetAndRefreshAllTerminalWebglAtlases instead. Only a comment in
pane-manager.ts still named the first one; it now points at the live entry
point. scheduleRevealPresent leaves the registry's structural type with them,
though the manager method stays: terminal-visibility-resume.ts calls it
directly.

This is dead-code removal, not a consequence of the xterm bump. The live
recovery path is unchanged.

resetAndRefreshAllTerminalWebglAtlases stays, and so does the reveal-time
escalation in pane-reveal-repaint.ts. Upstream 299 does make a pane-local
clearTexture bump pageLayoutVersion so siblings rebuild on their next frame,
which is the bug the escalation was written for, but I could not demonstrate
that removing it is safe: with the escalation removed,
floating-workspace-shared-glyph-atlas.spec.ts still passed headful, and it also
passed with upstream's mechanism deliberately disabled (pageLayoutVersion
pinned to 0 in the installed bundle, verified present in the built renderer).
A guard that passes with the fix disabled cannot license removing the
workaround, so the escalation stays until that spec can reproduce the garbling.

Verified: pane-manager and terminal-pane suites (4,713 tests), typecheck, the
headful shared-atlas spec, and the three headless WebGL specs.

* Give the shared glyph atlas spec a trigger that can fail

floating-workspace-shared-glyph-atlas.spec.ts guards the corruption where one
terminal wiping the module-global atlas leaves sibling terminals drawing from
stale texture coordinates. Both of its tests drive that through a floating
panel reveal, and Orca's reveal paths escalate to a registry-wide atlas reset
that repaints every pane — so the recovery under test heals the damage before
the assertion runs, and the tests pass whether or not xterm propagates the
invalidation at all.

The new test clears the shared atlas straight through the floating manager with
the panel closed, so nothing else repaints the workspace terminal, then repaints
it with terminal.refresh(). That is the load-bearing detail: _updateModel skips
cells whose content is unchanged, so the refresh reuses vertices baked against
the pages that were just wiped, which is exactly the state the fix has to
recover from.

Verified as a discriminator rather than assumed. Pinning ITextureAtlas's
pageLayoutVersion getter to 0 in the installed bundle, which disables the
per-renderer invalidation upstream added in addon-webgl 0.20.0-beta.299, and
confirming that reached the built renderer:

  fix intact:   siblingClearIntact=true   1 passed
  fix disabled: siblingClearIntact=false  1 failed

The failure renders the workspace terminal completely blank — stale coordinates
into a wiped atlas sample nothing. The two reveal tests pass unchanged in both
configurations, which is the gap this closes.

* Compare shared-atlas screenshots with tolerance instead of byte equality

Byte equality fails on sub-pixel antialiasing noise that leaves every glyph
legible, so the headful spec flaked under xterm 303. Reuse the existing
compareTerminalScreenshots helper: real stale-model corruption blanks the
terminal at ~3% of pixels, twice the helper's 1.5% threshold, so the looser
oracle keeps its teeth. Log the ratio so failures are diagnosable.

* fix(xterm): cancel empty deferred IME compositions

* test(xterm): strengthen runtime patch contracts
2026-08-30 15:14:49 -07:00
Brennan BensonandMerge Sim 585b4086d3 test(codex): pin Codex read-repair with a real-binary contract check (#17300)
* test(codex): pin Codex read-repair with a real-binary contract check

Orca's session index-heal depends on a Codex behavior: a `thread/read` of an
unindexed rollout performs a read-repair that inserts the `threads` row. All 55
existing heal tests drive a stub app-server and assert "healed" as "the call did
not error", so if Codex ever dropped the repair they would all stay green while
the subsystem went silently inert.

Adds a real-binary contract check built to the same shape as the Git binary
compatibility contract (src/shared/git-binary-compatibility.test.ts): env-gated
test file, version asserted against the binary, dedicated path-filtered PR job.

Pins only the four arms ablation established Orca relies on:
  - a read of an unindexed rollout inserts the state row
  - a session with no read inserts nothing (the negative control that makes the
    insert causal rather than incidental)
  - re-reading an indexed thread inserts nothing
  - an archived thread stays archived rather than being resurrected

Written against codex-cli 0.150.1. The job sets ORCA_CODEX_CONTRACT_REQUIRED=1
so a missing or failed CLI install fails red instead of silently skipping.

Existing heal tests are unchanged.

* test(codex): register the contract job in the verify aggregate contract

`pr-workflow-parallelism.test.mjs` pins `verify.needs` exactly, so adding the
job to pr.yml without updating that list failed the shard. Adds the entry, and
adds a workflow contract test mirroring `git-binary-compatibility-workflow.test.mjs`:

  - the pinned CODEX_CLI_VERSION is the single source for both the npm install
    and the runtime version assertion, so the two cannot drift apart
  - the install prefix and the binary path the test is pointed at are the same tree
  - ORCA_CODEX_CONTRACT_REQUIRED=1 is set, so a failed install fails red rather
    than turning the job into a green no-op

Removing the REQUIRED env from pr.yml reddens the new test, confirming it is live.

* test(codex): make binary version guard exact and bounded

* ci(codex): cover index-heal transport dependencies

* test(ci): pin Codex contract dependency coverage

* test(codex): align contract watchdog with child deadlines

* test(codex): cover three-session contract watchdog

* fix(codex): add sqlite sync-database to index-heal scope

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-30 14:39:46 -07:00
Neil 7b467bd0a6 ci: gate PRs on a real input method, and prove the lane engaged one (#17365)
* ci: gate PRs on a real input method, and prove the lane engaged one

No job on the PR gate has ever run a real input method. pr.yml and e2e.yml are
ubuntu-latest with CDP `Input.imeSetComposition`, which is a synthetic
composition; the only job that drives ibus-hangul through xdotool is
terminal-ime-e2e.yml, and it is schedule + dispatch only. A PR could turn the
real-IME path red and merge green.

Route IME source to that lane from pr.yml through the existing
pr-e2e-source-routing mechanism, so it runs on IME-touching PRs and nothing
else. The lane stays out of verify.needs — advisory, like `e2e` — because its
reliability is known only from nightly main runs. Deliberately no
continue-on-error: that reports green and hides the signal.

The harness fails open in ways that all look like success: Playwright reports a
skipped test as a pass, so an unset ORCA_E2E_NATIVE_IBUS_HANGUL, a renamed test,
or a session with no engine all exit 0 having exercised nothing. The specs now
append an engagement receipt only after observing real composition events, and
the runner requires one per expected test before the lane may report success.

Also drop the native spec from changed-e2e: it was already routed there by its
own filename, where it self-skips for want of an ibus session and reported that
skip as coverage.

* ci: let the real-IME step report even when the synthetic step failed
2026-08-30 01:58:32 -07:00
Jinwoo Hong 252dbd60ea fix(terminal): restore lossy initial remote snapshots (#17113)
* fix(terminal): restore lossy initial remote snapshots

* test(terminal): strengthen lossy snapshot causal oracle
2026-08-30 03:11:55 -04:00
Brennan Benson 6bef6d2727 fix(ci): stop the Linux Electron probe step from starving its own probes (#17071)
* fix(ci): stop the Linux Electron probe step from starving its own probes

The package job's "Test Linux Electron lifecycle boundary" step ran five
Electron probe files under Vitest's default file parallelism, so four full
Electron stacks competed for a 4-vCPU runner. Each probe carries its own
in-process deadline (20s for WebRTC, 25s for H3), and every observed failure
was one of those deadlines expiring: exit code 2, "no result", with every
sibling file in the same run slower than its own green maximum.

Run the step with --no-file-parallelism so each probe owns the runner, and
stop each probe nesting a private `xvfb-run --auto-servernum` X server inside
the step's own xvfb-run: reuse an inherited DISPLAY, and only own one when
there is none (shards, local dev), which leaves those lanes unchanged.

Also set each Docker-SSH E2E step's Playwright output aside before the next
step starts, because Playwright empties test-results/ on every run and only
the last lane's traces survived to the artifact.

* fix(ci): route the persisted-worker probe through the same display resolver
2026-08-29 01:38:17 -07:00
Brennan Benson fd9125ea8c feat(native-chat): Codex structured native chat restructure (#16729)
* feat(native-chat): port structured Codex sessions from restructure-recovery

Rebuilds the desktop structured native-chat implementation from
brennanb2025/native-chat-restructure-recovery (tip 4e31c08db3) on top of
current main as a single commit, scoped to the local Codex path.

Ported:
- Structured agent-session core: durable record store + single-writer lease,
  canonical journal, agent-session wire host/attach/eviction/subscribers,
  `agentSession.*` RPC surface (registered via ALL_RPC_METHODS; host-side
  mobile allowlist included for wire compat), pty write gate, transcript
  additions, and the Codex app-server adapter/launch resolution.
- Renderer: NativeChatStructuredSession view/composer stack, structured
  launch path with the single-flight guard, local structured session tabs
  sync, activation gate + structured inventory (read-only
  `agentSession.handoffStatus` probe), agent-session tabs in the tab strip,
  AI-vault structured session activation, and the settings pane with the
  parent Experimental Chat UI toggle plus the nested "Use updated structured
  native chat" toggle. New sessions require both flags, agent codex, no
  prompt, and a local non-WSL, non-Windows-host execution host
  (structured-native-chat-availability).
- Fixes 72c013cea6 (verified Codex launch recovery), 8ddbaf5e3d (defer
  native terminal view switching affordances), and 4e31c08db3 (release the
  launch gate after a visibility retry) with their regression tests,
  including the third-launch-after-retry guard case.
- Cross-version agent-session wire test + CI lane, packaging entries
  (proper-lockfile, agent-tooling asar excludes), and the wire-compat doc
  section.

Deliberately not ported: mobile/ changes, the Claude structured runtime
(only the claude-transcript-branch-proof and claude-structured-owner-identity
leaf modules remain, backing the kept TUI-recovery arms), the terminal↔chat
adoption/handoff flow (`agentSession.adoptTerminal`/`requestHandoff`, the
handoff request engine, TUI adoption machinery, orca-runtime adoption
methods), renderer switching affordances and their dead leftovers, the
hook/subagent-status refactor cluster, and unrelated branch changes. The
crash-during-acquisition recovery path (restart handoff adjudication,
restore/reverse re-acquire, lease schema handoff keys) is kept because every
plain direct launch depends on it; a trimmed handoff coordinator exposes
only status/restore/close.

Branch edits that targeted files main has since split (ipc/pty.ts,
worktrees.ts, rpc/methods/terminal.ts, useIpcEvents, pty-connection,
store/slices/terminals.ts, runtime-types, web preload) were re-applied to
the split modules, preserving main's newer logic (Windows CIM fallback,
browser tab close rework, cold-restore resume flow, dispatcher threading).

Known seam: the mobile clipboard image-provenance CONSUMER gate ships
(agentSession.send refuses unproven mobile image refs with
agent_session_image_untrusted) but the producer hunk in
rpc/methods/clipboard.ts stays with the unported mobile cluster, so mobile
image sends into structured chat fail closed until that side ports.

* fix(native-chat): trust only authenticated local image uploads

* fix(build): preserve Windows process-tree patch application

* test(windows): include process creation time in addon fixture

* fix(build): run windows-process-tree node-gyp from the physical package dir

gyp expands the node-addon-api dependency by probing node, whose cwd
resolves to the package's physical directory in the store, so the emitted
target is a store-relative ../../../../node-addon-api@... hop. gyp then
resolves that hop against the rebuild cwd; from the node_modules
symlink/junction it escapes the store and configure fails with
"node_addon_api.gyp not found" (run 32999886072).

Rebuild from realpath(package dir) so both bases agree, matching how the
package manager itself runs native install scripts. The regression test
replays gyp's expansion+resolution against the planned cwd and fails
without the fix.

* fix(native-chat): keep chat tabs visible through terminal closes and empty-worktree launches

Two proven blockers in the native Codex tab contract:

closeTerminalTab pre-empted the canonical unified close. With one terminal
left it deactivated the worktree on a terminal/editor/browser-only check,
blanking a workspace that still held a renderable agent-session tab; with
two or more it pre-picked a successor from terminal entities only,
re-stamping the group active before closeUnifiedTab's MRU/neighbor repair
could land on the chat tab. Successor choice now defers to the unified
contract whenever the terminal has a unified row, and deactivation is
gated on the unified renderable count (matching leaveWorktreeIfEmpty),
with the legacy pre-pick kept only for terminals without a unified row.

A structured session created on an empty worktree was published into the
host's headless group while preserveLocalLayout froze the local layout,
leaving the tab in store but permanently off screen. A preserveLocalLayout
owner now always takes client-owned placement — repairing a rendered
leaf whose group record is missing, or materializing a rendered group on a
truly empty worktree — and applies the client-derived layout repair while
still rejecting host-authored layout.

Regression tests drive the real store through closeTerminalTab (git
worktree and folder workspace) and the real snapshot applier for the
empty-worktree adoption states; all fail without the fixes.

* fix(native-chat): close stale turns and retry rejected sends

* fix(native-chat): retire hosted rows on structured tab activation

* fix(native-chat): preserve rpc defaults across main merge

* chore: format remote wire compatibility guide

* test(native-chat): cover retry after unconfirmed send

* fix(native-chat): reload outbox on session switch

* docs(settings): disclose structured chat platform limits

* fix(native-chat): await Codex launch-home preparation

* fix(codex): align child-process allowlist with async trust bridge

* test(identity): update inventory for tab surface refactor

* fix(windows): preserve process-tree CRLF patch sources

* fix(native-chat): anchor an unmatched chat echo where it was sent (#16117)

* fix(native-chat): anchor an unmatched chat echo where it was sent

The reported symptom was old user messages replaying below every new turn, so the
conversation read as scrambled. The cause was not that the echo failed to match a
transcript row. Claude consumes a mid-turn send through a `queued_command`
attachment and writes no `type:"user"` record for it, so some echoes can never
match, and no amount of matching will change that. The cause was WHERE an
unmatched echo rendered: buildMobileNativeChatTransientData appended every pending
item after the entire transcript, so it re-read below each turn that landed
afterwards.

Render each echo directly after the transcript row it was sent against, using the
baseline the send already captures. An unmatched echo is then at worst a duplicate
in the right position rather than a scrambled one, and it stays visible. Echoes
sharing an anchor keep send order; a send with no baseline, or one whose anchor
folding dropped, still falls back to the tail.

Deliberately NOT fixed by deleting the echo. Inferring from send ordering that an
echo can never match, then removing it, loses the user's own text for a message
the agent did receive, and it cannot fire in the common case anyway - measured
drain groups are 1,017 of size 1 against 55 larger. It also escalates an existing
gap: the count pass has no baseline-tail guard, unlike the glue pass, while
`messages` is a 40-row window that head-trims, resets on reconnect and grows at
the front on loadEarlier, so a false landing there would license deleting a
DIFFERENT outstanding message.

That count-pass gap is real and left for a separate change; anchoring makes its
worst case a duplicate in place rather than a scrambled conversation.

* fix(native-chat): preserve folded echo anchors

* fix(native-chat): preserve forward-folded echo anchors

* fix(native-chat): keep leading folded echoes in place

* fix(workspace-cleanup): show git status for every row (#16690)

* fix(native-chat): refuse structured chat on every Windows execution path

canUseStructuredNativeChat only refused win32 when a project runtime
resolved, so folder-workspace keys (and other keys with no project
runtime) failed open into structured chat on Windows. Fail closed on
win32 unconditionally after the host check, matching the settings copy:
local macOS/Linux only; Windows/WSL/SSH stay on terminal chat.

* fix(native-chat): restore runtime refusals behind the win32 gate

506d375de3 replaced the project-runtime checks with a bare platform test,
so a WSL or repair-required runtime resolution would no longer refuse
structured chat off-win32. Keep the unconditional win32 refusal and
re-run the runtime resolution after it, so the gate does not depend on
the resolver's own platform guard. Tests inject WSL and repair-required
resolutions on darwin/linux and fail against the regressed gate.

* fix structured session journal durability

* fix structured tab active pointer after restart

* fix(native-chat): await optional lease renewal callbacks

* refactor(skills): extract install error messages

* fix(agent-session): harden recovery ownership

* fix(native-chat): retain panes across tab activation

* fix(native-chat): address round-one review findings

* test(native-chat): align integration coverage after main merge

* fix(native-chat): harden round-two reliability

* fix(native-chat): harden round-three reliability

* fix(native-chat): close round-four recovery gaps

* fix(native-chat): separate bounded journal key forms

* fix(native-chat): reset outbox error in render on session switch

The switch effect adjusted error state after the sessionId prop changed,
tripping react-doctor's no-adjust-state-on-prop-change on the changed-code
gate and flashing the old session's banner for a frame. Reset it with the
render-time previous-value guard instead.

* fix(native-chat): invalidate stale outbox settlements

* test(native-chat): restore settled-error session-switch regression

a6e2379bd1 replaced this test with the in-flight settlement race test,
leaving the render-time error reset unpinned: deleting the reset block
still passed the whole native-chat suite. Keep both scenarios pinned;
they are distinct (settled error clears on switch vs stale settlement
invalidated in the commit-to-passive window).

* test(wire): make release checkouts race safe

* test(wire): pin cross-process checkout single-flight and importer specifier contract

* test(wire): harden release checkout lifecycle

* fix(build): drop CR-byte residue from windows-process-tree patch

The two trailing CR bytes on the patch's deletion lines are a proven
no-op: pnpm hashes patches CRLF-normalized (both forms hash to the
lockfile's 946ffb2b) and materializes this package without applying the
patch in either form, so the load-bearing build edits come solely from
applyWindowsProcessTreeBuildFixes() (#16947), which handles both source
EOL forms. Restore byte-identity with main and repin the contract test
to the post-#16947 reality: LF-only patch bytes plus lockfile hash sync.

* fix(native-chat): skip empty startup recovery
2026-08-28 16:45:58 -07:00