Commit Graph
187 Commits
Author SHA1 Message Date
Baekspace 78121c2dd5 fix: support CJK-adjacent Markdown emphasis 2026-09-18 00:20:39 -07:00
Neil 2569a71ce8 fix(deps): update vulnerable dependencies without new overrides 2026-09-16 20:36:24 -07:00
Neil 11180fa532 chore(lint): add anti-slop oxlint plugin (pinned, all rules off) (#20726)
* chore(lint): add anti-slop oxlint plugin (all rules off)

Vendors dmmulroy/anti-slop (MIT) plus no-call-only-assertions and
no-pass-through-type-alias from maharshi365/deslop (MIT). Every rule starts
"off"; each follow-up PR fixes one rule's violations and flips it to "error".

* fix(lint): actually exclude the vendored plugin from the anti-slop audit

oxlint does not honour ignorePatterns supplied via --config, so the
config/oxlint-plugins/anti-slop/** entry never matched and the vendored rule
source was being linted as first-party code (505 violations). Move the exclusion
to the --ignore-pattern CLI flag in audit:anti-slop, which does work, and drop
the entry that gave a false sense of coverage.

Keeping vendored source unlinted matters because anti-slop is updated by
three-way merge against the upstream snapshot; reformatting it locally would
conflict on every update.

* chore(lint): pin anti-slop instead of vendoring it; drop deslop

Replaces the ~5k vendored lines with a git-pinned devDependency:
  oxlint-plugin-anti-slop: github:dmmulroy/anti-slop#c44ef22

anti-slop ships raw .ts with no build step, and Node refuses to type-strip
anything under node_modules (ERR_UNSUPPORTED_NODE_MODULES_TYPE_STRIPPING), so
oxlint cannot load it from there -- which is why upstream says to vendor it. A
postinstall step copies the pinned package's source to .anti-slop-plugin/
(gitignored), which Node will type-strip because it sits outside node_modules.
Upgrading is now a SHA bump rather than a re-vendor and three-way merge.

Verified byte-identical rule output to the vendored copy across all 16 rules
that fire.

Drops maharshi365/deslop and its two rules (no-call-only-assertions,
no-pass-through-type-alias). It is not on npm either, so it would need a second
git pin and copy step, and it is a 5-star single-maintainer repo that is itself
a re-namespaced copy of anti-slop. One upstream is enough.

* ci(lint): run audit:anti-slop in PR CI

config/scripts/pr-workflow-lint-parity.test.mjs requires every step in
`pnpm lint` to have a matching step in .github/workflows/pr.yml; adding
audit:anti-slop to lint without the workflow step failed that ratchet.

Also makes audit:anti-slop sync the plugin itself before linting. The generated
.anti-slop-plugin/ directory is gitignored and otherwise only created by
postinstall, so a cached install that skips postinstall would leave oxlint
unable to load the plugin.
2026-09-14 21:42:37 -07:00
Neil b61a2347b9 feat(design-system): gate renderer UI with @shadcn/lint (#20731)
* feat(design-system): gate renderer UI with @shadcn/lint

Wires shadcn-ui/lint's Oxlint plugin into the two places this repo already
ratchets: the changed-lines PR gate for rules the renderer can't satisfy
today, and `pnpm lint` for the one that is already at zero.

- config/oxlint-design-system.json: no-restyle (layout allowed),
  no-raw-colors, require-static-classes -- scoped to src/renderer/**/*.tsx,
  run over added lines only. Measured at 10 findings across the last 60
  commits (771 changed files), so it holds the line without a migration.
- config/oxlint-dead-classes.json: no-unknown-classes repo-wide, with the
  renderer's plain-CSS hook namespaces allow-listed. Now at zero.
- no-inline-styles and no-arbitrary-values stay off; STYLEGUIDE says why.

Fixes the three live bugs the linter found:

- `--editor-surface` never reached `@theme inline`, so `bg-editor-surface`
  generated no CSS -- 12 editor/artifact/notebook panes fell through to the
  page background instead of #1e1e1e in dark mode.
- `scrollbar-none` is not a Tailwind utility and was declared nowhere, so
  the remote file browser breadcrumbs showed the scrollbar they meant to
  hide. Declared as a real `@utility`.
- Notebook markdown cells used `markdown-preview-body`, which no stylesheet
  defines; the styled class is `markdown-body`. They rendered unstyled.

* ci: run the dead-class gate in PR CI

`pnpm lint` gained check:dead-classes, and pr-workflow-lint-parity requires
every `pnpm lint` step to have a matching step in pr.yml.

* fix(notebook): keep markdown theme selectors working
2026-09-14 17:52:21 -07:00
6cb5643241 fix(deps): migrate Tiptap security updates with Markdown compatibility guards (#19376)
* chore(deps): evaluate coordinated Tiptap security migration

* fix(editor): adapt link ranking and initialization for Tiptap 3.31

* fix(editor): preserve literal Markdown through Tiptap serialization

* test(editor): cover literal saves in local folder and paired workspaces

* test(editor): reselect folder after closing its final tab

* perf(editor): avoid repeated inline source-marker lookahead scans

* refactor(editor): inline redundant HTML match wrapper

* test(chat): await Tiptap React skill-pill rendering

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-14 16:36:06 -07:00
nireak b87a6c0f23 fix(pty): pace the EAGAIN write retry so a stalled reader can't saturate the daemon thread (#15319)
node-pty's CustomWriteStream retries an EAGAIN write with setImmediate, which
re-attempts within microseconds. A pty whose child has stopped draining stdin
keeps that branch EAGAIN-ing, so the retry becomes a busy-loop on the thread
that owns every pty on the runtime. Measured against this commit's parent on
macOS arm64: 121,316 EAGAIN/s at 101.6% CPU, versus 805/s at 4.1% with the
retry paced to 1ms.

The delay is 1ms rather than longer because the cost lands on readers that
drain in bursts -- what an agent does between event-loop ticks. Delivering 2MB
to a reader that drains 20ms out of every 100ms: 689ms unpaced, 907ms at 1ms,
1414ms at 5ms. 1ms keeps essentially all of the CPU saving without the
delivery regression.

clearImmediate -> clearTimeout in dispose() is required, not cosmetic: once the
handle is a Timeout, clearImmediate does not cancel it and a pending retry can
fire after dispose. The disposal guards that make that harmless (_fd = -1, queue
drop) are already on main; this mirrors them into src/unixTerminal.ts so the
TypeScript twin no longer drifts from the compiled lib.

Scope: this fixes the CPU saturation. It does not stop other terminals from
being serviced -- a second live pty kept answering echo round-trips throughout
the storm in every configuration tested (1 and 8 stalled writers, macOS and
Linux, 8 CPUs and 1), with throughput down ~20-50% rather than hung. The
"every terminal froze" symptom in #11178 has another cause and that issue
stays open.

Upstream chose setImmediate deliberately (microsoft/node-pty#831, #833) to fix
large-paste latency, and rejected polling POLLOUT because it reports writable
rather than flushed. That reasoning targets a per-write delay in an interactive
terminal; this delays only the EAGAIN branch in a long-lived daemon. Pastes to
a draining reader are unaffected (0-3 EAGAINs per MB in every arm).

Verified: patch applies to a pristine node-pty@1.1.0 tarball, the patched
src/unixTerminal.ts compiles byte-identical to the patched lib/unixTerminal.js,
patch_hash matches the file, and on Windows the changed code never executes
(WindowsTerminal, 0 EAGAINs on a 300KB conpty write).
2026-09-14 02:32:08 -07:00
OrcaWinandm4air 09187fcad8 fix(ai-vault): stream oversized remote session transcripts (#20455)
* fix(ai-vault): stream oversized remote session transcripts

* fix(build): bundle streamed JSON parser in desktop main

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-13 15:43:21 -07:00
Neil 62ae09d947 chore(deps): pin serve-sim to an exact version (#20446)
serve-sim ships platform binaries that are bundled into the app, and it
publishes frequently in the 0.1.x range, so a routine install could change
them. The resolved version is unchanged at 0.1.40; only the range is
narrowed, so upgrades become a deliberate edit.
2026-09-12 20:41:02 -07:00
Neil 5127d1eb3b refactor(windows): vendor the registry addon as @orca/windows-registry (#20438)
* refactor(windows): vendor the registry addon as @orca/windows-registry

windows-native-registry@3.2.2 was last published in 2023 by a single
maintainer. Orca called two of its exports, both read-only, so the whole
dependency is replaced by a local N-API addon under native/.

The vendored addon is read-only by construction: setValue, createKey and
deleteKey are gone, so RegDeleteTreeW no longer ships in the app. Two
upstream defects are also fixed rather than carried over — the name/data
scratch buffers were file-scope statics that concurrent reads would
scribble over, and createKey/deleteKey called .c_str() on a temporary.

Build wiring keeps the existing shape: still an optionalDependency gated
to win32, still excluded from pnpm's allowBuilds so only Orca's own
Windows rebuild runs node-gyp for it, still copied into the packaged
resources. The CI native caches now key on the vendored sources so an
addon.cc edit cannot restore a stale .node.

* test(windows): check the vendored registry addon against reg.exe

The addon is vendored source, so no upstream release proves it still
decodes values the way Orca's PATH readers expect. reg.exe is the only
independent oracle on the box.

* ci(windows): register the registry addon test on the Windows runner

A Windows-gated file self-skips on ubuntu, so without both registrations
it reports success while running on no machine at all.

* fix(build): link the registry addon as a workspace package, not file:

As a `file:` dependency pnpm re-resolved and re-linked the package on
every install, including `--frozen-lockfile` (measured: "added 1" on a
repeat no-op install). That virtual-store churn ran concurrently with
node-gyp reading the same tree and cost @vscode/windows-process-tree its
binding.gyp mid-rebuild, failing package (windows) whenever the native
cache hit and only that module needed building. The linux packaging job
hit the same race from the other side, as a pnpm staging move failure.

A workspace link resolves once and leaves the store alone; repeat
installs are now 55ms no-ops. native/windows-registry is listed
explicitly so `packages:` still does not auto-discover mobile/.

* fix(build): stop tracking node-gyp output for the vendored addon

The build/ tree is generated per host and ABI; the committed copy was
macOS-specific gyp scaffolding from a local build and would have shipped
stale Makefiles to every checkout.

* chore: ignore the vendored addon's node-gyp bin output too

node-gyp also emits bin/<platform>-<abi>/ beside build/; both are per-host
generated output that must never be committed.
2026-09-12 20:10:49 -07:00
Neil 56fcb544e0 fix(browser): move cookie scoping off psl's stale suffix list (#20421)
* fix(browser): move cookie scoping off psl's stale suffix list

psl@1.15.0 is its latest release and ships a Dec-2024 snapshot of the
public suffix list. Measured against the current upstream list, it fails
to recognise 600 of 10,030 suffixes; tldts misses 2.

That gap is a cookie-isolation bug. psl does not know `api.br` is a
suffix, so it falls back to the `br` rule and maps foo.api.br, bar.api.br
and example.api.br all onto the single family `api.br`. Unrelated
registrants then share a removal scope, and a replace-mode import for one
clears the others' cookies. The same holds for seg.ar, co.az, gov.cz and
~597 more.

tldts is called with allowPrivateDomains, without which the PSL's PRIVATE
section is ignored and every *.github.io / *.s3.amazonaws.com / *.vercel.app
tenant collapses into one family — 21 of 49 probed hosts changed family
under the default. The new test pins that boundary.

One deliberate behaviour change: hosts under `.local` (not in the PSL)
were their own family under psl, which returned an all-null parse for
them; they now resolve to the two-label boundary (app.orca.local ->
orca.local), matching what Chromium treats as the registrable domain.

* fix(build): bundle tldts into the main process like psl was

psl sat in BUNDLED_MAIN_DEPENDENCIES, so it was inlined into the main
bundle rather than externalized and copied into resources/node_modules.
Swapping the dependency without moving that entry left a bare tldts
import that afterPack's runtime-closure check rejects.

* fix(build): point the output contract at tldts and drop the psl shim

The contract test still asserted psl was in BUNDLED_MAIN_DEPENDENCIES, so
it failed once the entry became tldts. src/types/psl.ts declared a module
that no longer resolves; tldts ships its own types.

* test(browser): pin the suffix boundaries the tldts swap moved

Three semantic changes shipped untested:

- `.local` is unlisted, and the libraries disagreed on what that means. psl
  returned an all-null parse so every `*.orca.local` host was its own family;
  tldts stops at `orca.local`. The consequence is wider than the family name —
  importDomainAncestors now yields the shared parent, so a replace-mode import
  of one host clears non-host-only cookies every sibling shares.
- psl's snapshot had `compute.amazonaws.com` as a literal PRIVATE suffix; the
  current list only carries the wildcard, so the bare host is ICANN now.
- The renderer's `psl.isValid` gate had no direct test at all — nothing imported
  the module from a test.

Also drops comments that explained a boundary in terms of psl's internals. One
was wrong under tldts: bracketed IPv6 does not reach an error branch, it parses
with the brackets stripped and falls through the unlisted path.
2026-09-12 16:00:54 -07:00
Neil 53eb639983 refactor(preload): drop the unused raw electron IPC bridge (#20419)
* refactor(preload): drop the unused raw electron IPC bridge

`@electron-toolkit/preload` was used only to expose `window.electron`,
which hands the renderer unrestricted `ipcRenderer` send/invoke/on for any
channel — bypassing the typed per-domain bridges in `src/preload/api/`.

Nothing consumed it. The only references were the assignment itself, the
web client's empty fallback, and a test asserting that fallback has no
keys — i.e. the web build already ran with it empty.

* chore(build): drop the dangling @electron-toolkit/preload vite exclude

The package is gone from package.json and source; leaving it in the
preload externalizeDeps exclude list points at a package that no longer
resolves.
2026-09-12 16:00:19 -07:00
Neil 9c918b93a6 fix(deps): take Electron 43.7.0 for the glibc environ use-after-free (#20089) 2026-09-11 15:43:43 -07:00
Neil bba68b1bdd fix(pi): finish the dialog-wait signal on every surface (#19533)
* fix(pi): carry modal waits to mobile and stop losing the dialog close

Follow-ups to #18836, from its readiness review.

- Paint pi's `!` needs-input state marker while a dialog is open, so the
  80ms spinner frame stops repainting a working title over a mid-turn
  wait. Mobile and the CLI read the title, so they saw `working` where
  the desktop already showed `waiting`.
- Keep the assistant reply that lands while a dialog is open. The modal
  guard cleared tool fields and the `message_end` capture with them, so
  a turn ending under a dialog left the preview on the previous message.
- Report `ui_prompt_end` even when `ctx.isIdle()` throws on a runner the
  modal itself invalidated; the lost post stranded the pane on `waiting`.
- Declare the `esbuild` the runtime smoke tool imports.

* fix(pi): hold the needs-input marker until the dialog actually closes

From review of the previous commit.

- Settling under an open dialog no longer retires the marker. stopAnimation
  painted the plain title unconditionally, so agent_settled, a resolved
  agent_end, or an idle auto_compaction_end erased it mid-dialog — and
  because that also cleared the timer, the close then painted the plain
  title again and the wait was lost for good.
- Track the dialog as a boolean, not a depth counter. Pi does its own
  nesting accounting and emits one pair per stack, which is what the status
  extension already assumes; two files disagreeing on that would have let an
  inner close release the outer wait.
- Reset the flag on agent_start in both extensions. A turn cannot begin under
  a dialog holding input focus, so it is the one boundary that can recover a
  close that never arrived instead of pinning the pane forever.
- Leave OMP to its approval events: it reports waits through those already,
  and painting the marker there too would put title and hook in disagreement.

* fix(pi): do not ring the completion bell for a dialog that lost its close

From review of the previous commit.

- Report working, not done, when ui_prompt_end's isIdle() throws. done is
  not cosmetic: it reaches dispatchCompletion and fires the pane's finished
  notification, so a turn that is still running would announce itself. The
  real done still arrives from agent_end/agent_settled.
- Keep the idle-maintenance frame cap accruing while a dialog holds the
  title, so a dialog left open cannot suspend the guard that stops a
  compaction spinner whose end event never came.
- Guard the dialog handlers against a ctx without ui. The source is
  generated and untypechecked, and pi does not document the ctx it passes
  these two events; a TypeError there would surface on every dialog.

* fix(pi): let a turn still complete after a dialog loses its runner

From review of the previous commit.

- Re-arm the completion report when ui_prompt_end's isIdle() throws. The
  fallback posts working, but the finished turn had already reported its
  end, so nothing further would ever fire and an idle pane sat spinning.
- Count dialog depth in both extensions instead of trusting pi to emit one
  pair per stack. The guarantee is undocumented, and if it ever does emit a
  pair per dialog, an inner close would release the wait the outer dialog
  still holds. A counter costs nothing and drops the dependency.

* fix(pi): decide a dialog close from turn state, not from a guess

From review of the previous commit.

- Fall back to agentEndReported when ctx.isIdle is unavailable or throws.
  The previous guess of working stranded the common case — a dialog opened
  at idle — because no later event was coming to correct it, and the
  agentEndReported re-arm it relied on could not fire either. A turn that
  already reported its end is not still running, and that is knowledge this
  process holds without needing ctx at all.
- Only suppress spinner frames once the marker is actually painted. Pi may
  pass a ctx with no ui, and freezing the title on its last working frame
  is the opposite of what the marker is for.
- Gate the titlebar dialog handlers on the OMP runtime too, not just the
  installed kind: a bare-shell OMP launch runs inside a pi-kind pane, and
  the status extension already defers there. Extracted that check so both
  extensions share it rather than carrying two copies.

* fix(pi): treat a pane that never ran a turn as idle, not busy

From review of the previous commit.

- Track turn-in-flight separately from agentEndReported. That flag also
  dedupes the completion post, so it starts false on a pane that has not
  run a turn — which read as still-running and left a dialog opened before
  the first prompt spinning forever.
- Retry the marker paint on each dialog open instead of only the outermost,
  so an outer ctx without ui cannot decide the whole nested stack goes
  unmarked.
- Fall back to the opening ctx when the close carries no ui. Nothing else
  clears the needs-input marker, so the pane would have kept asking for
  attention until the next turn.

* fix(pi): keep a dying dialog ctx from stranding the needs-input marker

The close path paints through the ctx captured at open time, which is the
one a session-switching modal is most likely to have invalidated. Guard
both paint sites so a throw cannot reject the handler and leave the title
on the needs-input marker, and make local turn state the floor for the
status extension's idleness verdict instead of a fallback.

* fix(pi): hold the dialog wait against pi's own title writes and lost closes

Reviewed against real Pi 0.85.1 source rather than inference:

- ctx.ui is a getter that calls assertActive() and throws once a session-
  replacing dialog invalidates the runner, so optional chaining never
  screened it out and the probe sat outside the try. A throw landed after
  the depth decrement but before markerPainted cleared, stranding the
  needs-input marker until the next turn.
- Pi writes the same terminal title from its own writers with no event we
  observe, so the marker is now re-asserted rather than merely not
  overwritten, on a slow timer that outlives the spinner and its cap.
- resetExtensionUI drops an open dialog without resolving its promise, so
  a replaced or reloaded session never emits the matching ui_prompt_end.
  Both extensions now release the wait on session_start and shutdown.

* fix(pi): build the title inside the guard, not as an argument to it

paintTitle caught the setTitle throw but not the two calls one argument to
its left: pi.getSessionName() asserts runner liveness the same way ctx.ui
does, and process.cwd() throws ENOENT once the worktree is unlinked under a
live pane. Four of the six call sites are timer callbacks, where an escape
is an uncaught exception and pi exits(1) through its own handler — so the
cwd route was reachable today. paintTitle now takes a builder and runs it
inside the existing try.

* fix(pi): let only the pane-owning process assert the needs-input marker

The spinner is harmlessly per-process, but the marker is status the pane
reports, and child agents inherit ORCA_PANE_KEY. Gate the two dialog
handlers on a PID claim, mirroring ORCA_PI_STATUS_OWNED in the status hook.
2026-09-08 03:06:54 -07:00
OrcaWinandm4air de0a91b99f fix(deps): update Electron to reviewed 43.6 runtime (#19369)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-07 19:12:12 -07:00
BingZandNeil 5bd0247aaa fix(xterm): remove scrollback decorations by identity (#13178)
* fix(xterm): fire Marker dispose before clearing line (#10879)

Scrollback trim under search highlights was O(k²) because dispose set
marker.line to -1 before onDispose, collapsing SortedList keys. Fire
listeners first so delete still sees the real line, then clear the line.

Fixes #10879

* fix(xterm): remove scrollback decorations by identity

* perf(xterm): avoid index arrays for unique decorations

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-07 18:18:04 -07:00
OrcaWinandm4air 102402e41e fix(deps): update DOMPurify sanitizer hardening (#19377)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-07 17:43:28 -07:00
OrcaWinandm4air b8f6c7cabe fix(deps): update react-i18next for TypeScript 7 and parser fixes (#19378)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-07 17:43:18 -07:00
OrcaWinandm4air cfd59f2ca0 fix(deps): update desktop parser security dependencies (#19361)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-07 17:42:29 -07:00
Neil 314506003a fix: retain MSYS shell descendants in their terminal job (#19068)
* fix: retain MSYS shell descendants in their terminal job

* test: complete MSYS regression CI registration and teardown contract

* fix(windows): deny job breakaway for the whole Cygwin/MSYS shell family

The per-PTY job probed only msys-2.0.dll, and only for bash.exe/sh.exe.
Cygwin ships the same spawn.cc breakaway logic under cygwin1.dll, and an
MSYS2 zsh escapes exactly like its bash does, so both kept the orphan bug.

Probe the runtime DLL on the shell's own search path instead of matching
shell names: that is the property that decides whether the runtime will
ask for CREATE_BREAKAWAY_FROM_JOB, and it drops the name special-casing.

* chore(patch): restore the conpty.cc index line

The earlier hand-edit dropped it while every sibling section kept one.
Recomputed against the real blobs: applying this patch to 7b286d3d
yields exactly 4b06d185, so git apply -3 has its fallback back.
2026-09-07 00:35:55 -07:00
Brennan BensonandMerge Sim 1478101342 fix(windows): unblock structured native chat by exposing process creation time (#18986)
* fix(windows): guard process creation times

* fix(windows): ask the relay's bare addon for creation times too

The relay addon build now emits creationTimeMs, but the runtime binding
for the bare addon still declared only CommandLine, so a Windows relay
host requested flag 2 and every row came back without a creation time.
That leaves captureWindowsDescendantSnapshot returning null and
verifyWindowsProcessIdentity false forever on those hosts -- the relay
half of the patch was unreachable.

Naming CreationTime in the adapter is safe because the bare addon is a
content-hashed relay artifact: it ships in the same immutable relay
directory as the bundle reading it, so it can never be older than the
code asking for the bit.

Also bound the win32 guard test on our own row, which the addon can
never fail to answer, so an unconverted FILETIME or a 1601-epoch stamp
fails instead of satisfying a bare count.

* fix(windows): make the compiled addon prove its own CreationTime support

CI caught the real defect: the win32 guard test read
isWindowsProcessStartTimeAvailable() as true and then found 0 rows
carrying creationTimeMs. Unlike node-pty, this package publishes a
prebuilt .node at the same build/Release path node-gyp writes to, so
pnpm patches the source tree and leaves that binary alone. A host then
holds a patched lib/index.js -- ProcessDataFlag.CreationTime and all --
over a binary that ignores flag 4, and neither a load check nor a path
check can see the difference.

So the binary now says so itself: addon.cc exports
supportedProcessDataFlags, lib/index.js re-exports it, and

  - windows-process-tree-creation-time.cjs asserts it during install,
    which is what forces a from-source rebuild. It is shared by the Node
    probe in ensure-native-runtime.mjs and the Electron probe in
    rebuild-native-deps.mjs, exactly as node-pty-job-ownership.cjs is --
    the Electron half matters because that probe decides onlyModules, so
    without it the packaged app would ship the stale prebuilt.
  - isWindowsProcessStartTimeAvailable() gates on the reported bit, not
    the enum. Believing the enum is worse than reporting false: the
    descendant snapshot returns null forever and the exit proof latches
    unverifiable while structured chat believes it has a reaper.

rebuildNodeRuntimeModules could not actually have rebuilt this package:
the patched binding.gyp includes deps/node-addon-api, which the tarball
does not ship, and node-gyp must run from the physical dir.

Also closes the relay repair path's divergence: repairCreationTimeSources
wrote the C++ but not the buildNode splat or the tree-node typing, and
assertPatchApplied checked neither, so a repaired tree passed as patched
with buildProcessTree silently dropping the field.

The guard test is unchanged.

* fix(windows): keep the process-tree patch LF-only

windows-process-tree-patch-contract.test.mjs requires the patch file to
carry no CR bytes. Regenerating through pnpm patch-commit emitted 199 of
them, because the creation-time change is the first to touch files the
package ships as CRLF (src/process.h, src/process_worker.cc,
src/addon.cc, lib/index.js, lib/index.ts, the typings) -- and #17886's
own hunks over binding.gyp and src/process_commandline.cc carry the rest.

Stripping them is safe and changes nothing the lockfile records: pnpm
hashes patches CRLF-normalized, so the digest stays
e66202cc623996d02040c93449eb9ae353fddadf426cb53202a59ee710ee6fe7 and now
equals the file's plain sha256 too. It also still applies -- verified
against a deleted store entry, not a warm one -- and the precedent was
already there: the previous patch was LF-only and had been patching
those same CRLF files all along.

ensure-native-runtime.test.mjs stages the siblings the script loads at
module scope into its temp project. The import walk added by #17886 sees
`from './x.mjs'` only, so the createRequire'd .cjs siblings still have to
be named, and this PR adds a second one.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 13:59:59 -07:00
bfc6a262a7 fix(windows): read command lines from the kernel, not each process's PEB (#17886)
* fix(windows): read command lines from the kernel, not each process's PEB

MDE incident D scored Orca for suspicious memory activity: the vendored
`@vscode/windows-process-tree` recovered every process's command line by
opening it with `PROCESS_QUERY_INFORMATION | PROCESS_VM_READ` and chaining
three `ReadProcessMemory` calls through the PEB and
`RTL_USER_PROCESS_PARAMETERS`. On a 750ms/2s cadence over the whole table that
is the credential-dumping primitive, whatever the intent.

Windows 8.1 added `NtQueryInformationProcess`'s `ProcessCommandLineInformation`
class (60), which returns the same string as a kernel-built `UNICODE_STRING`
under `PROCESS_QUERY_LIMITED_INFORMATION` alone. Electron's floor is Windows
10, so every supported OS has it. The PEB reader stays behind a process-wide
latch that only `STATUS_INVALID_INFO_CLASS`/`NOT_SUPPORTED`/`NOT_IMPLEMENTED`
can set; a pid that merely denied a handle does not re-arm it, because
`PROCESS_QUERY_INFORMATION` implicitly grants the limited right and so cannot
be obtained where the weaker open already failed.

The same hunk drops `PROCESS_VM_READ` from `GetProcessMemoryUsage` and
`GetCpuUsage`, which acquired it and never read an address space.

Measured on Windows 11 (514 processes), counted in-process by swapping the
addon's import table entries for counting stubs, per CommandLine scan:
`ReadProcessMemory` 1128 -> 0, desired access 0x0410 -> 0x1000, p50 12.7ms ->
9.3ms. Command lines were byte-identical on every process both readers
recovered (376/376, 379/379 across runs), including a 24,068-character argv
with quotes, non-ASCII and trailing whitespace, and a WOW64 target. Three
processes that refused the old rights granted the new one; none went the other
way.

* chore(deps): refresh the windows-process-tree patch hash in the lockfile

* fix(windows): drop the PEB fallback and detect the unpatched prebuilt

Review of #17886 found three ways the reader could still perform, or silently
resume, the primitive it exists to remove.

The class-missing latch was a permanent, process-wide, one-way downgrade back
to the PEB read, and any single target returning STATUS_INVALID_INFO_CLASS /
NOT_SUPPORTED / NOT_IMPLEMENTED could trip it. On an EDR-hooked ntdll -- the
entire premise of this change -- a hook that does not recognise class 60 would
have restored PROCESS_VM_READ plus three ReadProcessMemory per pid per scan for
the life of the process, unobservably, on precisely the machines this was
written for. The fallback is deleted rather than guarded: GetProcessCommandLine
now returns false and leaves the command line empty, which callers already
handle, so the addon imports no ReadProcessMemory at all.

That absence is what makes the property checkable on the artifact. The
published 0.8.0 tarball ships a loadable prebuilt built from unpatched source;
it is node-addon-api, so a bare require() accepts it, allowBuilds is false and
CI installs with --ignore-scripts, and a rebuild that soft-exits on a Windows
file lock leaves it in place. Source-text guards could never see it.
windowsProcessTreeAddonReadsProcessMemory() checks the compiled binary instead,
and is wired into the install check, the rebuild, and the relay build.

The repair itself never worked: `git apply` run inside a work tree prefixes
patch paths with the cwd-relative prefix, skips what does not match, and exits
0, so the branch always fell through to its own post-check throw. The package
dir is always under the project root, while the fixture that covered it was in
%TEMP%, outside any repo. Blinding git with GIT_DIR fixes it, and the test now
runs inside a real work tree.

Also from review: bounds-check the returned UNICODE_STRING against the
allocation (not the size the second query clobbers) and cap the probe so a
bogus length cannot bad_alloc a whole scan; test NT_SUCCESS explicitly; value-
initialize ProcessInfo, which left `memory` as stack garbage -- measured, 82
processes reported the same bogus working set; and correct a comment in
windows-process-table.ts that still described the command line as a PEB read.

Re-measured on Windows 11 (543 processes): ReadProcessMemory 1128 -> 0, with
the symbol absent from the import table so the IAT hook finds no slot to
count; desired access 0x0410 -> 0x1000 on all 543 opens; p50 13.5 -> 12.3ms;
405/405 command lines byte-identical including a 24,087-character quoted
non-ASCII argv and a WOW64 target; 3 processes recovered only by the new path,
0 only by the old.

* chore(deps): refresh the windows-process-tree patch hash in the lockfile

* test(scripts): stage a script's local imports into the native-runtime fixture

ensure-native-runtime.mjs gained an import of windows-process-tree-gyp-rebuild.mjs,
but the fixture copied only the script itself, so every case in the suite died
with ERR_MODULE_NOT_FOUND before reaching its own assertions. copyScriptWithLocalModules
already walks a script's co-located imports for exactly this reason -- its own doc
comment names this failure -- so use it rather than listing files by hand.

The two Windows cases still fail here, on a missing node-pty ConPTY runtime that
also fails on main; this only stops a resolution error from standing in front of
whatever they were meant to catch.

* fix(windows): route a locked stale addon to the Windows file-lock message

`pnpm install` with Orca running aborted with a raw EPERM stack. The stale-binary
guard -- which deletes an addon that still imports ReadProcessMemory so a skipped
rebuild cannot use it -- ran outside the try whose catch classifies Windows file
locks, and whose message is literally "Close running Orca/Electron/dev processes
for this worktree": exactly this situation.

Measured rather than assumed: rmSync against a loaded (memory-mapped) addon throws
EPERM, and `force: true` does not help, since it only swallows ENOENT. Cold copies
of the same file delete fine. So the delete threw a page before the handler that
knows what it means.

Moving the guard inside the try is the whole fix; the classifier already matches
the EPERM text. The new case runs the real script against a temp project whose
stale addon is held open by a live child process, and fails against the old
placement with the raw `syscall: 'rm'` stack the report described.

* feat(windows): warn once when command-line recovery is refused host-wide

Removing the PEB fallback removed a total-defeat vector, but it left a cliff: if
NtQueryInformationProcess(ProcessCommandLineInformation) is refused -- a hooked
ntdll that does not know class 60 -- every command line comes back empty and
agent identity matching silently degrades to image names. The addon still loads
and still enumerates, so every health check the app has stays green. A cliff
nobody can see is the failure mode this area keeps producing.

The querying process is the unambiguous probe. A process can always open itself
with PROCESS_QUERY_LIMITED_INFORMATION, so its own command line coming back empty
means the query is refused for every process -- not that some target denied a
handle, which is normal for roughly a quarter of the table. Keying on our own row
rather than a fraction means no threshold to tune and no false positive on a
hardened box where most processes deny.

One warning per session, gated on the CommandLine flag actually being requested so
a future identity-only reader cannot trip it. The suite's own SELF fixture gains a
command line for the same reason: a self row without one is the alarm, not a
detail.

* fix(windows): check the relay's staged addon at load, and answer tri-state

Two gaps in the ReadProcessMemory check, both about what it does not see.

It only ever looked at node_modules/@vscode/windows-process-tree. A relay host
has no node_modules of ours: it loads ./windows-process-tree.node staged beside
the bundle. The relay build asserts the symbol on the artifact it produces, but a
bundle and the addon beside it redeploy independently, so a host that has not
taken a new bundle keeps whatever binary is already there -- and the published
prebuilt is node-addon-api, so it binds cleanly and then walks every process's
address space. loadWindowsProcessTree now checks that file too and refuses it,
falling back to the CIM scan: slower, but not the thing an EDR quarantines a host
for. The predicate is duplicated rather than imported, because the config-script
copy is install-time tooling that drags in node-gyp and child_process, and this
module is bundled into the app and the relay.

And it returned false for a binary that is not there. All three callers happened
to be safe, but the name read as a safety predicate, so a future caller would take
a missing binary as verified. inspectWindowsProcessTreeAddon() now answers
clean/unpatched/missing over an explicit binary path -- which is also what lets
the relay's staged addon be checked at all -- and each caller states which state
it acts on.

Both are covered by cases that fail against the old code: without the load-time
check the unpatched staged addon is bound and the CIM fallback never runs, and
with 'missing' folded back into 'clean' the absence case fails outright.

* test(windows): load the addon in beforeAll, not at collection time

loadAddon() ran while the file was being collected, so on a Windows checkout with
no built addon the require threw before any case existed and took the seven
patch-text cases down with it -- cases that read only the patch file and need no
binary at all. Verified both ways against a deliberately unresolvable addon path:
at collection time vitest reports "no tests" for the file; from beforeAll the
seven text cases pass and only the three addon cases go.

* fix(deps): normalize the windows-process-tree patch to LF and let pnpm own its hash

`pnpm install --frozen-lockfile` failed on this branch on every platform with
ERR_PNPM_LOCKFILE_CONFIG_MISMATCH, which breaks CI and the release build.

Two coupled defects. The patch file was committed with CRLF -- 174 CR bytes,
against zero on main -- and `.gitattributes` pins `/config/patches/*.patch -text`
precisely so checkout cannot convert it, so those bytes reached every runner. And
pnpm hashes a patch **LF-normalized**, so the raw sha256 of a CRLF file is a value
pnpm never computes:

  raw sha256      322965470c05f63d8527f7d8e892ee26ee444136b66b57fd64c362a9f2ff05d1
  LF-normalized   f8ea245391c94da5770045aeea01fa6de466c2199c6ef46b5b769b398aa9823e

The lockfile carried the raw one, at all three sites. It is the only one of the
seven patches where the two digests differ, which is why the other six passed.

Normalized the patch to LF and took pnpm's own value from
`pnpm install --no-frozen-lockfile`; nothing here is hand-computed. With the file
LF-only the two interpretations coincide, so the lockfile, the contract test's
no-CR assertion and its hash assertion all agree at one number -- and
`config/scripts/windows-process-tree-patch-contract.test.mjs`, which was red on
this branch for the same reason, is green again. The lockfile diff is exactly the
three hash lines.

The regression check is the installer, not a digest. Two separate reviews
"verified" the shipped hash by recomputing sha256(patchBytes) and matching the
lockfile; both were wrong, because both repeated the same wrong assumption about
which bytes pnpm hashes. A check that reproduces the original mistake is not
independent. So the new case runs `pnpm install --frozen-lockfile --lockfile-only
--ignore-scripts` against a copy of the manifest, lockfile and patches, and
asserts exit 0 -- verified by deletion: restoring the shipped hash fails it with
the exact ERR_PNPM_LOCKFILE_CONFIG_MISMATCH from the branch's package (windows)
job.

Also corrected the `.gitattributes` comment claiming pnpm hashes patches
byte-for-byte. The `-text` setting is right -- `git apply` needs the exact bytes --
but that sentence is the claim that produced the wrong hash twice.

* ci(windows): run the process-tree patch suites in CI

Both suites only self-skip off Windows, so the binary-level check that the
addon carries no ReadProcessMemory passed vacuously in every lane.

* fix(windows): force core.autocrlf=input for the patch repair

My LF normalization of the windows-process-tree patch broke the `git apply`
repair path introduced in this PR. The two are coupled and I checked only one.

Those 174 CR bytes were not editor noise. They sat on exactly the pre-image
lines and nowhere else -- 107/107 in src/process.cc, 67/67 in
src/process_commandline.cc, 0 on every added or context line -- because
@vscode/windows-process-tree@0.8.0 ships those two sources as CRLF. Normalizing
the patch made its pre-image stop matching the file it is applied against.

Measured, reconstructing the true CRLF pre-image from the pre-normalization
blob and applying the current LF patch:

  core.autocrlf   plain   -c core.autocrlf=input
  true            exit 0  exit 0
  input           exit 0  exit 0
  false           exit 1  exit 0

`false` is Git's own built-in default and what "checkout as-is" selects in the
Git for Windows installer -- on this box the `true` that hides it comes from the
installer's system gitconfig, not from anything in the repo. There the repair
throws, ensureWindowsProcessTreeCommandLinePatch reports "still reads the PEB,
and repairing it ... failed", isWindowsNativeLockError does not match that text,
and `pnpm install` dies with no path forward.

Forcing the mode rather than `--ignore-whitespace`: both fix every cell and both
leave the applied file fully LF, but `input` relaxes line endings only, so a hunk
whose real content drifted is still rejected. The repair rewrites a
security-relevant source file; it should stay strict about everything except the
thing that is legitimately ambiguous.

Not reverting the patch to CRLF: windows-process-tree-patch-contract.test.mjs
(pre-existing on main) forbids CR bytes in it, and pnpm computes the same hash
either way. LF plus the forced mode is the end state.

The suite could not have caught this. The fixture built its pre-image from the
patch itself and joined with '\n', so fixture and patch agreed by construction on
any encoding -- once again a test that passes without its fix. It now emits the
CRLF the real package ships, and the case runs under both autocrlf modes pinned
through a temp HOME gitconfig, because the repair blinds git to the repo and so
reads global config. Verified by deletion in both directions: with the flag
removed the autocrlf=false case fails with the exact "still reads the PEB" dead
end while autocrlf=true still passes, and with the fixture back on LF all eight
cases pass with no fix present at all.

Also corrected the .gitattributes comment I added last commit. It said `git
apply` needs the bytes the patch was written against, which is now false -- the
pinned bytes are LF and the bytes it was written against are CRLF. That is the
same class of confident-and-wrong claim that produced the bad hash twice.

* fix(windows): assert the rebuilt addon, and install the patch for real in tests

Three follow-ups from review.

**The packaged binary had no check.** The relay build asserts its own artifact
and ensure-native-runtime asserts what it loads, but nothing looked at the addon
copied into the packaged app -- so a rebuild that silently produced the upstream
reader shipped. `rebuild-native-deps.mjs` now asserts `clean` on it after
`rebuild()`. This is also the caller D4's tri-state was missing: every existing
site branches on `=== 'unpatched'`, so `missing` still behaved exactly like
`clean` everywhere, which was the thing making it a state rather than a boolean.
Here both non-clean states fail, and they fail differently: after a rebuild that
reported success, an absent binary is a broken build, not an absence to shrug at.

The fake `rebuild()` had to start producing a binary for that to mean anything,
so it now emits stand-in bytes and takes `addon: 'clean' | 'unpatched' | 'none'`.
Verified by deletion: with the assertion removed both new cases pass.

**The frozen-install case could not see a patch at all.** `--lockfile-only`
resolves and never applies one, so its coverage stops at hash consistency. Added
a case that installs `@vscode/windows-process-tree@0.8.0` for real with the patch
and asserts the materialized `src/process_commandline.cc` carries the marker and
no longer carries `ReadProcessMemory` -- about 1.5s for the pair.

Correcting the brief on that one: it does **not** catch the `git apply` breakage
from the previous commit. Measured -- with `-c core.autocrlf=input` removed it
passes cleanly, because `pnpm install` uses pnpm's own patch applier and never
runs our repair script. What it does catch is a patch pnpm can no longer apply:
corrupting one pre-image line fails both cases. The repair path stays covered by
the CRLF fixture in rebuild-native-deps-node-pty.test.mjs.

Worth recording, since it decides whether the LF normalization was safe at all:
pnpm applies the LF patch to the CRLF tarball sources without complaint, and
materializes them as LF with the marker present and `ReadProcessMemory` absent.
The primary install path was never affected -- only the `git apply` fallback was.

**Dead timeout.** The frozen-install case passed `timeoutMs: 300_000` to the
spawn while vitest capped the case itself at 30s, so on a cold runner vitest
would have killed it first. Both cases now declare the budget they use.

* test(windows): route the frozen-install check through the pnpm invocation owner

The new patched-dependencies check hand-rolled a PATH walk naming 'pnpm.cmd',
which the windows batch shim spawn boundary ratchet rejects: pnpm-cli-invocation
already owns that decision for every other script, and its allowlist only
shrinks.

Reuse resolvePnpmCliInvocation for the command and prefixArgs, and the shared
resolveCliCommand for the presence check, so no shim name is spelled here. Its
`shell` flag is dropped because runProcessSync refuses it and already drives a
shim through the interpreter itself.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-05 21:12:47 -07:00
a65332a8bd feat(claude): move structured native chat onto the Claude Agent SDK and enable it on macOS and Linux (#18560)
* Join structured attach teardown through journal bind

* fix: restore structured chat parity

* feat: add Claude structured session adapter

* fix: harden Claude structured adapter

* fix: close Claude adapter edge cases

* fix: start Claude init deadline after launch

* feat: wire Claude structured sessions

* fix: harden Claude structured runtime

* fix: fence Claude structured compatibility

* fix: preserve Claude free-text prompt answers

* fix: decode addressed Claude prompt text

* feat: enable Claude structured chat on mobile

* fix(mobile): keep structured chat provider-aware

* fix(mobile): negotiate Claude structured tabs

* fix: keep scoped RPC tests native-free

* fix: secure mobile structured image delivery

* fix: close structured session data-loss gaps

* fix: prove real Claude structured startup

* fix: consume pre-spawn proof before retry

* feat(native-chat): add desktop structured sessions

* fix(native-chat): satisfy structured session cleanup gates

* fix(native-chat): keep structured renders pure

* fix(native-chat): open composer pickers upward

* fix(native-chat): use existing view for structured sessions

* fix: harden structured desktop status projection

* fix: close structured desktop lifecycle gaps

* fix: fence structured AI Vault resumes

* fix: fence structured AI Vault resumes

* fix: preserve structured tabs during activation

* feat: toggle structured sessions between chat and TUI

* fix: harden structured session handoffs

* fix: bind structured TUI before rollout proof

* fix: complete structured chat round trips

* fix: align structured TUI return readiness

* fix(native-chat): make reverse handoff transactional

* Add Claude structured TUI handoff seams

* fix(native-chat): clear sticky handoff recovery

* fix(native-chat): complete mobile reverse after TUI exit

* fix(native-chat): keep TUI transcripts readable

* fix(native-chat): recover TUI transcript gaps

* fix(native-chat): recover claimed TUI owners

* fix(native-chat): retain cold TUI proof authority

* fix(native-chat): preserve Claude handoff authority

* fix(native-chat): recover TUI transcripts read-only

* fix(native-chat): harden Claude handoff recovery

* fix(native-chat): serialize structured handoff recovery

* fix(native-chat): close handoff admission races

* fix(native-chat): validate pinned launch environment

* fix(native-chat): revalidate restored and retried owners

* fix(native-chat): gate restart recovery publications

* fix(i18n): catalog Claude session controls

* fix(native-chat): wait for structured TUI process proof

* fix(native-chat): queue stale idle TUI handoffs

* fix(native-chat): route structured Codex options directly

* fix(native-chat): persist structured session options

* fix(native-chat): hydrate resumed structured options

* fix(native-chat): preserve options across structured handoffs

* fix(native-chat): replay pending option mutations

* fix(native-chat): rotate settled handoff operations

* fix(native-chat): rotate refused send operations

* test(native-chat): derive refusal retry state from host

* test(native-chat): give the host-oracle matrix test an explicit timeout

* fix(native-chat): keep Claude option controls idle

* fix mobile structured first-send hydration race

* fix(native-chat): preserve handoff launch authority

* fix(native-chat): harden shared handoff recovery

* fix(native-chat): serialize structured handoff recovery

* fix(native-chat): close handoff admission races

* fix(native-chat): validate pinned launch environment

* fix(native-chat): revalidate restored and retried owners

* fix(native-chat): gate restart recovery publications

* fix(i18n): catalog structured session recovery control

* fix(native-chat): wait for structured TUI process proof

* fix(native-chat): queue stale idle TUI handoffs

* fix(native-chat): keep structured recovery provider-neutral

* fix(native-chat): drop local terminal topology from structured sync

* fix structured outbox and tab restore races

* fix(native-chat): preserve Claude question groups

* fix structured provider visibility and request handling

* fix structured session TUI handoff recovery

* fix reverse structured session handoff

* fix(native-chat): recover Claude outbox and resume state

* chore(mobile): preserve the working-tree lockfile state before the main merge

Carries the pre-existing uncommitted mobile/pnpm-lock.yaml modification into history so the
main merge cannot overwrite it. Verified benign pnpm drift (babel 7.29.7->7.29.8 transitives
plus deprecation metadata); drops no patchedDependencies (the mobile lockfile declares none).

* test(native-chat): drop orphaned Claude handoff-auth test left by the main merge

'pins Claude handoff auth through the terminal provider boundary' is absent from main and its
production counterpart preserveClaudeAuthEnv no longer exists outside this test - orphaned residue
of the terminal/native handoff work this PR excludes by scope.

Removed rather than repaired: the failure was a renamed field (providerHome -> providerRoot), and
renaming it would have carried out-of-scope handoff code into the merge. Body preserved as evidence
and logged in CLAUDE-STRUCTURED-DISPOSITION-TABLE.md.

* Fix mobile structured turn state

* fix Claude structured session blockers

* fix claude structured lane blockers

* fix Claude acquisition exit proof

* fix(claude): route stream-json launch through process wrapper

* fix(claude): gate structured chat support

* Fix Claude structured launch gating

* fix(claude): split session acquisition and prune mobile scope

* test(claude): align structured session fixtures

* fix(agent-session): preserve handoff launch arguments

* fix(claude): open journals through the factory after origin/main split

The journal opener moved to journal-store-factory on main; retarget the
Claude structured tests that still imported the old path.

* fix(claude): resolve Claude structured launch args, auth, and win32 proof

The origin/main merge re-expressed the lane's Claude wiring onto main's split
orca-runtime facade and dropped three wires past green typecheck and lint.

- resolveLaunchArgs discarded its provider parameter, so structured Claude
  sessions were launched with Codex app-server flags; Claude exits on
  --dangerously-bypass-approvals-and-sandbox, and a Codex arg-parse throw
  could block Claude session creation outright.
- resolveClaudeLaunchEnv was no longer supplied, so the launch resolver fell
  back to the whole process env as configuredEnv and
  buildClaudeChildProcessEnv re-applied every auth var it had just stripped.
  The resolver now merges the Claude overlay onto a strip-applied copy of the
  inherited env, which also keeps PATH intact for withCliRuntimeOnPath.
- The windowsProcessStartTimeAvailable producer was gone while the contract
  field and both consumers survived, so the renderer gate fail-closed and
  structured native chat was unreachable on every win32 host.

Separately, structured Claude pinned CLAUDE_CONFIG_DIR unconditionally. An
explicit pin makes the CLI abandon the macOS Keychain even when it names the
CLI's own default, so a default claude.ai account could not authenticate where
the legacy Claude terminal could. Pin only a home the CLI would not resolve on
its own, matching ClaudeRuntimePathResolver, and compare against the env the
child would otherwise inherit so a diverging overlay cannot outrank the
record's account home.

Also await the now-async revealNativeSession in its regression test, and set
the native status before revealing so a rejecting reveal cannot leave a
session released but never marked native.

Claude-Session: https://claude.ai/code/session_013UqKCRB6k5e8UaYhXUHeWY

* fix(claude): scrub case-insensitive Windows auth env

* fix(native-chat): settle handoff outcome-write failures instead of leaking them

A store write failure while recording a handoff outcome escaped the flow
runner's catch handler, so the client never received the failure and the
flow surfaced as an unhandled rejection (seen as an intermittent
agent_session_store_corrupt error in the proven-dead-retry suite, whose
teardown raced the flow's trailing outcome write). Record the failed
outcome best-effort, and drain the coordinator before that test's
teardown removes the store root.

Claude-Session: https://claude.ai/code/session_011aXkcHyeiRJuezupQdjZaM

* fix(native-chat): make the structured close-failure toast provider-neutral

The structuredSessionCloseFailed toast fires for any structured session,
but its copy said 'Codex chat', so a Claude structured session that fails
to close showed the wrong provider name. The launch-failure toast is only
reachable behind the agent === 'codex' gate, so its copy stays as is.

Claude-Session: https://claude.ai/code/session_013ugSpCx4AWkySaJb69BQax

* fix(native-chat): wire structured handoff proof recovery

* fix(native-chat): wire structured handoff proof recovery

* fix(native-chat): correct the structured chat opt-in copy

The one `experimentalStructuredNativeChat` toggle gates both providers —
`useStructuredAgentSessionCreate` runs `canUseStructuredNativeChat` for
`'claude'` as well as `'codex'` — but its description named only Codex.

Its scope line also said Windows keeps using terminal chat, while the gate
refuses win32 only until the host proves it can read a process start time.
`structured-native-chat-availability.test.ts` already pins that Windows is
allowed once the proof is cached, so the two contradicted each other.

Claude-Session: https://claude.ai/code/session_01RJFsidQWmKYFmeoUuVu4Tp

* test(claude): pin @anthropic-ai/claude-agent-sdk 0.3.251 contracts against a scripted CLI

PR 1 of the SDK migration: dependency + test-only harness, no product wiring.

- Pin @anthropic-ai/claude-agent-sdk to exactly 0.3.251 — not the newest
  release — because 0.3.251 (published 2026-08-28) clears the repo's 3-day
  minimumReleaseAge supply-chain gate with no exclusion, while the newest
  release was minutes old and would have required excluding a brand-new
  publish from the exact control built to catch brand-new malicious
  publishes. Every contract this design depends on was verified identical
  on 0.3.251: the full option surface, no pid on SpawnedProcess (custom
  spawner stays mandatory), env defaulting to process.env when omitted, and
  --replay-user-messages appearing only via extraArgs.
- Exclude all eight bundled CLI platform binaries via
  ignoredOptionalDependencies. The setting lives in pnpm-workspace.yaml
  because pnpm 12 no longer reads the package.json "pnpm" field (it warns
  and ignores it; verified by install ablation). Excluding the binaries is
  what makes Orca's pathToClaudeCodeExecutable override mandatory rather
  than merely preferred. Note: pnpm 12.0.0 honors the ignore list when
  reconciling an existing lockfile but not on fresh resolution of a new
  dependency, so the lockfile's SDK entry was pinned surgically; both
  'pnpm install' and 'pnpm install --frozen-lockfile' verify clean and
  stable against the committed lockfile.
- Contract-pin suite drives the real SDK against a scripted fake CLI and pins:
  unknown type/field/content-block pass-through (and keep_alive interception),
  spawner env fidelity plus the omitted-env process.env inheritance sharp edge,
  extraArgs producing --replay-user-messages, argument parity for every
  CLAUDE_STRUCTURED_BASE_ARGS entry plus --session-id/--resume/
  --resume-session-at, canUseTool wire request_id stability and abort on
  control_cancel_request, one spawn per query, pathToClaudeCodeExecutable
  honored by the default spawner, the exact SDK version, and the eight platform
  binaries staying uninstalled.

Claude-Session: https://claude.ai/code/session_01FGCRfYUnb4hbvfTAHGtJKQ

* feat(claude): drive the structured transport through the agent SDK

Replaces the hand-rolled `claude -p --input-format stream-json` transport with
@anthropic-ai/claude-agent-sdk 0.3.251, keeping the existing connection
interface for this commit so the acquisition path changes minimally. The
control-plane rewrite is a separate change.

Orca still supplies the process. `spawnClaudeCodeProcess` routes through
`spawnProcess`, retains the child and its pid — the triple the durable lease
adjudicates on — drains stderr so exit errors keep their tail, and hands `.cmd`
shims to Orca's Windows argument encoder rather than the SDK's plain spawn.
`close()` keeps Orca's own bounded tree-kill and exit deadline, so it still
resolves true only after an observed exit.

Launch resolution emits an SDK options object instead of argv; durable
`launchArgs` translate to a typed option where one exists and to `extraArgs`
otherwise, refusing a token neither can carry rather than dropping it. The
child env is always passed explicitly — omitting it would let the SDK inherit
`process.env` and reintroduce the ambient `ANTHROPIC_*` leak. The stdout line
parser is deleted; the SDK owns framing, and unknown frames still reach the
translator verbatim.

Claude-Session: https://claude.ai/code/session_01JMhFjh9HEnkcJ5YTfCdgD3

* fix(claude): settle the frame the SDK pulled but never wrote

The SDK's input pump is `for await (frame of prompt) { await transport.write(frame) }`.
When that write rejects — the child dies between Orca's liveness guard and the
write — the for-await ends abruptly and calls the generator's `return()`, so the
code after `yield` never runs. The frame was already shift()ed out of `queued`,
so the later `fail()` from the exit path could not reach it and `send()` never
settled: `dispatchClaudeTurn` awaits that send before it can return `unknown`,
wedging the caller and the durable outbox. The pre-SDK transport rejected on the
stdin write callback instead.

Retain the in-flight entry and settle it from the generator's cleanup, and let
fail() reach it too for the pump that never resumes at all.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): keep the agent SDK behind the structured-Claude boundary

The ordinary OrcaRuntimeService graph statically reaches the Claude adapter and
so the transport module, whose first line imported @anthropic-ai/claude-agent-sdk.
The SDK is evaluated whenever the regular runtime loads, before any structured
Claude session is chosen: it sets process.env.NoDefaultCurrentDirectoryInExePath,
changing Windows executable resolution for later subprocesses, and a missing or
incompatible install would break normal runtime startup — for a user who never
leaves the terminal/TUI path.

Defer the SDK to the connection, memoized so it loads once per process, and add
the import-graph ratchet: a walk from the Electron main entry that fails on any
static import of the package, plus a clean-fork check that loading the runtime
leaves the Windows search variable untouched and a child-process pin that the
side effect is still real.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): answer list_models so the picker stops serving the seed

sendControlRequest had no list_models case, so every request hit the default
reject; readClaudeStructuredSessionOptions swallows that with .catch(() => null)
and falls back to the static catalog. Every structured session therefore served a
hardcoded model list with no per-model effort levels, no resolvedModel and no
default detection, and nothing surfaced the failure. The pre-SDK transport got the
live catalog from the CLI.

Route it through the SDK's supportedModels(), wrapped in the { models } envelope
the existing parser reads.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): reap the child's descendants before killing it

The forced step of the exit ladder went through the Codex helper, which spawns
`pkill -KILL -P <pid>` and SIGKILLs the parent in the same tick: the parent
usually dies first, the descendants reparent to pid 1, and `-P` matches nothing.
An MCP or launcher descendant of a stubborn Claude child was left running. The
test named for that requirement declined to assert it and killed the survivor by
hand instead, so it could not fail for the thing it was named after.

Route the Claude reap through Orca's existing sweep, which snapshots descendants
while their parent link still exists and signals them before the root goes, and
on Windows uses the identity-gated `taskkill /T /F`. The test now asserts the
descendant is dead; the manual kill stays only as a failure-safe. close() still
returns true only on an observed exit.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(native-chat): merge the duplicated handoff type import

CI's static-analysis lint (`oxlint --config
config/oxlint-code-quality-native-plugins.json src config tests mobile
--deny-warnings`) exits 1 on the two separate `import type` statements from the
same module.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): answer a permission callback whose signal already aborted

settleFrom registered the abort listener and then delivered the request. A
callback that arrives already aborted never fires that event, so the promise
stayed pending behind a durable prompt with no cancel path. Check the signal
first, emit the cancel, and resolve the SDK's null sentinel without registering.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* test(claude): wait for the child to record the frame, not just for its report

The scripted CLI writes its report at startup, so `until(readReport)` returned a
report with no user messages whenever the child had not yet read the line. The
assertion then failed under parallel load. Poll for the frame instead of for the
file.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): coalesce partial deltas onto one assistant item and stop painting result frames

Under --include-partial-messages every stream_event frame carries its own
uuid, and the final assistant frame for a block carries yet another; only
message.id ties them. The translator keyed each delta by its frame uuid, so a
reply painted as one bubble per delta chunk followed by a complete duplicate
under the final frame's uuid. The block's first stream frame now mints the
claude:(sessionId, uuid) identity, deltas coalesce onto it through the shared
60ms seam, and the final frame reconciles onto that same item.

Known SDK bookkeeping no longer reaches the provider-fallback row: result
subtypes are catalogued and settled by the turn lifecycle, an empty thinking
block (redacted thinking) is a modeled kind, a string-content user replay is a
text block, and an empty user frame paints nothing. An unmodeled result
subtype or content kind still lands on the bounded fallback row.

Claude-Session: https://claude.ai/code/session_01GaP5HpYQbvy2hYehVhwfEW

* fix(claude): prove descendant exit at the close boundary instead of on an unref'd timer

close() reported proven=true as soon as the direct child exited while the
descendant sweep's SIGKILL sat on an unref'd 2 s timer, so a SIGTERM-resistant
MCP server outlived the lease release. The reaper now composes the same shared
primitives the Codex structured provider uses: snapshot, verified bounded
descendant termination on POSIX, taskkill /T /F on Windows. The proof is false
whenever descendants outlive the deadline, a retried close re-verifies the
retained snapshot rather than trusting the dead root, and the raw pipe child no
longer goes through the PTY job sweep it never owned a job for.

Measured on macOS: a killed child of a SIGSTOPped parent stays a matching zombie
row in ps, so the root is killed while verification runs rather than stopped
first as the Codex non-group path does.

Claude-Session: https://claude.ai/code/session_0161QFm3KVRNJKfdzWVGVNWk

* feat(claude): replace the hand-rolled control plane with the SDK's native surface

PR 3 of the Claude structured SDK migration removes the wire-frame scaffolding
PR 2 kept, so Orca drives the SDK's typed control surface directly.

Inbound permissions move from a rebuilt control_request dispatch to the SDK's
canUseTool / onUserDialog callbacks. The prompt registry now carries the
callback's own resolver: a decodable can_use_tool becomes a durable prompt whose
answer settles the callback; a malformed one is denied without registering; the
SDK's abort signal (fired on control_cancel_request, which the SDK matches and
dedups itself) forgets the prompt and settles it null, and a late answer after
abort finds no prompt and is refused. Closing settles every in-flight callback so
no promise dangles. The claude-agent-sdk-control-bridge that rebuilt the wire
frame is deleted.

Outbound control maps to Query methods: interrupt() for cancel, setModel /
setPermissionMode / applyFlagSettings for options, supportedModels for the model
list, initializationResult() for init proof, each under Orca's own request
deadline and error classification. Cancel is interrupt-receipt aware: a CLI
advertising interrupt_cancel_queued_v1 gets cancel_queued in one round trip,
otherwise the receipt's still_queued uuids are swept with cancel_async_message so
a cancelled turn cannot spawn a later unexpected turn; older CLIs resolve no
receipt. Init keeps the 10s deadline and the unauthenticated-startup guidance.

Every behavior is failing-first and ablation-proven; the toggle-off import
boundary and the accepted loss of unknown-control visibility rows are unchanged.

Claude-Session: https://claude.ai/code/session_01Pqjduxt5G4rr9aYvtp7rNm

* fix(claude): arm the descendant snapshot before stdin closes and make the tree verdict unproven by default

A healthy Claude root leaves within the graceful window, and the close ladder
only snapshotted descendants when the root was still alive after that window.
So the common close never looked at the tree: `treeExited` stayed null,
`!== false` passed it, and close() reported a proven exit with an MCP child
still running. A root that died before the walk made the snapshot vacuous too.

The proof is now unproven by default. The reaper holds one verdict in Orca's
vocabulary (exited / live / unverifiable), assigned in exactly one place from
the bounded verification, and close() returns true only on `exited`. The
snapshot is armed before stdin closes, while the root can still be walked, and
is verified after the root exits; a root that left before any snapshot could
be armed stays unverifiable rather than vouching for descendants it never
showed us. The shared verifier gains the three-way verdict behind its boolean
face, and the connection reports the root and tree verdicts separately along
with the child's exit status.

One verification per close attempt: the retried close re-verifies, so the
intra-attempt re-reap is gone from the teardown budget.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): verify the Windows tree after taskkill instead of trusting that it ran

`terminateWindowsProcessTree` resolves from taskkill's callback whatever the
error says, so a timeout, an access denial, a recycled root and a surviving
descendant all looked identical to the reaper — which then returned a proven
exit unconditionally. close() reported true and the lease was released with an
MCP descendant potentially still live.

The Windows branch now snapshots the root's descendants while it is alive and,
after taskkill, polls a fresh process table to a bounded deadline: a row still
matching by pid AND creation time is `live`, an unreadable table is
`unverifiable`, and only a table with no match is `exited`. Creation time is
the PID-reuse guard the POSIX path gets from ps lstart, so a descendant that
denied a creation-time query is omitted rather than signalled on a bare pid.
A root already observed exited is never taskkilled: `/T /F` on a recycled pid
would take an unrelated tree down with it.

The captured tree is tagged by platform so neither verifier can be handed the
other's rows.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): release a reservation on a first-hand root exit instead of latching it into manual recovery

Making close() strict about the descendant tree exposed a second defect at the
same boundary. A create-time acquisition has no ownerProcess until publication,
so an unproven cleanup mapped to handoffStage `manual-recovery`, and
adjudication then refuses every later attach with agent_session_ownership_unknown.
A user who was merely signed out, or whose --resume the CLI rejected, wedged the
session id permanently.

Each question now answers from its own evidence. close() is unchanged and stays
strict about the tree. Separately, the lease is keyed on the root's pid and
start time, so when Orca's own child handle observed that root exit and no
descendant snapshot was ever admissible, the reservation is released and the
CLI's exit code and stderr reach the user. A descendant observed still alive,
or a root Orca never saw leave, stays unproven and keeps the reservation.

The settlement records only what was observed: the released lease says the
provider process exited and its descendants were not verifiable, rather than
reusing the wording that claims cleanup proved no child remains.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): surface an API error a result frame reports instead of settling the turn on it

The SDK models an API failure as a SUCCESS-subtype result whose `result` string
is the user-facing error text, with no assistant frame behind it. The translator
suppressed every catalogued result subtype as turn bookkeeping, so that turn
tombstoned its lifecycle and showed the user a completed, empty reply with no
sign anything had failed.

Suppression is now by meaning. A result reporting a failure routes to the
bounded provider-error surface, leading with the provider's own sentence and
keeping the raw frame behind the row's disclosure; ordinary successful results
stay off the timeline as before. A turn the user aborted also stays suppressed:
its interrupt frame already says so, and its execution diagnostic would only be
noise on every stop.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): drop the stream state of turns that never received their final frame

Every streamed delta recorded its block's identity, latest text and checkpoint
length. Only the final assistant frame removed them, so an interrupted turn left
its whole accumulated reply reachable until the session was disposed, and a long
session with repeated interruptions grew those maps without bound. The partial
text was already journaled by the flush that precedes settlement, so the live
copy was pure retention.

That state now lives in its own module, named for what it does — grow a streamed
block's journal row between its deltas and its final frame — and turn settlement
drops every block still awaiting a final. The translator reports how many remain,
which is the invariant: a settled turn leaves none.

Also makes a timed-out process-table read retryable while the root is still
alive. A loaded host can miss the table's one-second deadline, and latching that
as "no descendants" both lost the descendant sweep and, on a busy machine, made
the close ladder report unproven for a tree it never actually looked at. Only
the root's death still makes a missing snapshot final.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* perf(claude): capture the Windows descendant tree from one process-table read

The capture walked the descendant tree and then read the table again for the
creation times the walk's projection drops. Each read is bounded in seconds and
both run inside the close ladder's budget, so the second one cost the worst-case
teardown three seconds for data the first read already held.

The walk is now exported from the module that owns it and runs over rows the
caller has already read, which is also what lets the snapshot keep the
PID-reuse guard the projection cannot carry.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(pty): spend the descendant verification window instead of surrendering on one slow table read

The verification abandoned the whole check the first time a process-table read
missed its own one-second deadline, with seconds of its window still unspent.
On a loaded host that reported a tree unverifiable without ever having looked at
it, which the Claude close ladder then turned into an unproven close and a
retried teardown. It also made the descendant-exit tests flake under a parallel
suite run, for the same reason and with the same honest-but-premature verdict.

A read that missed its deadline is now simply not an answer: the loop waits and
reads again until its own deadline, and only a window that ends without a
readable table reports unverifiable. This can only turn a premature verdict into
one backed by evidence; it never manufactures a proof.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): never let a later failed look collapse an observed live descendant into unverifiable

The reaper's single assignment site latched only 'exited', so a second reap
whose table reads all missed their deadline overwrote an earlier completed
verification's 'live' with 'unverifiable'. The acquisition release gate
discriminates on exactly that pair, so a root exit after such a decay released
the lease over a descendant that had been observed alive. The latch is now
monotone in trust order: exited is final, and live is only ever raised to exited.

Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP

* fix(claude): never prove a Windows tree gone while a descendant denied identification

The Windows snapshot dropped rows that denied the creation-time query, and an
emptied snapshot was judged exited without any table read: a descendant Orca was
refused information about was treated as one that had left. The snapshot now
counts the unidentified rows it saw, and verification caps its verdict at
unverifiable while any exist. Nothing is ever signalled on a bare pid, as before.

Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP

* fix(claude): classify cleanup after a first-hand exit as a root exit instead of a proven tree

When the CLI died between a successful acquire and the host's commit or proof
of the lease, handleExit had already removed the session, so releaseAcquisition
found nothing and reported true. The attach flow then settled exit-proven with
deathEvidence claiming cleanup proved no provider child remains, though the
tree was never verified. The adapter now keeps the exit that removed a
published session until the session is acquired again; acquisition cleanup runs
that connection's close ladder and classifies its verdict exactly as a
start-time failure would be, so the record reads root-exit-observed. The wire
helper keeps that typed classification and its provider diagnostic instead of
wrapping it as unproven, and the router gives up its owner even when the
release throws.

Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP

* fix(claude): integrate SDK teardown and picker lifecycle fixes

* fix(claude): preserve resume leaf and settle processless spawns

* fix(claude): reacquire from persisted resume leaf

* fix(native-chat): restore Claude grouped question handling

* fix(claude): persist only resumable transcript leaves

* fix(claude): recover structured session exits safely

* fix(claude): close remaining structured session P1s

* fix(claude): harden transcript branch proof

* Remove superseded root fix reports

* fix(windows): restore indexed descendant row walk

* fix(router): forward force-close lifecycle

* fix(claude): fence stale turn cancellations

* fix(claude): fence cancellation after unknown dispatch

* fix(claude): fence replay and option recovery races

* fix(claude): block replay fallback after waiter eviction

* fix(claude): fence evicted slash results

* fix(claude): fence ambiguous results and restore options safely

* fix(claude): scrub SDK child env and localize pending launch

* fix(claude): pin transcript roots and exit recovery proofs

* fix(claude): retain unproven SDK exits

* fix(claude): settle retained exit before reacquire

* fix(claude): resume from settled retained cursor

* chore: remove tracked review artifact

* fix: harden Claude SDK transport session cleanup

* fix: close Claude sessions safely

* fix(claude): close races with fresh child snapshots

* fix(claude): fail closed on recycled child identities

* fix(claude): gate root cleanup on process identity

* fix(claude): fence same-second root identity reuse

* fix(claude): restore the root SIGKILL fallback the identity gate took away

The direct root kill goes through the handle Node owns, not through a pid:
libuv drops that handle in the same turn it reaps, so the signal either
reaches the process Orca spawned or reaches nothing at all. Gating it on a
process-table probe therefore bought no safety and cost the tree its only
fallback whenever the probe declined -- a first capture landing in the fork's
own second, a recycled descendant pid voiding the snapshot, or a process table
that could not be read on either platform.

Identity verification stays where a bare pid is genuinely addressed: Windows
`taskkill /T /F`, and the descendant sweep's own revalidation before it signals.

Also stops a declined root probe from collapsing an observed `live` or `exited`
descendant verdict into `unverifiable`, and stops a successful taskkill from
reporting `unverifiable` because a later probe found the root correctly dead.

* docs(claude): rewrap the root-kill ordering comment

* Match the Claude structured launch to the terminal path's managed-account auth rules

The SDK path stripped ambient Anthropic auth unconditionally, let an explicit
agentDefaultEnv override beat a pinned managed account, and had no account-switch
guard. Reuse the terminal preflight's own predicate and messages so both transports
strip, refuse, and report identically, and cover the CLI transcript location that
mobile native chat depends on.

* Reach the Claude structured chat lane from the desktop UI

The main process has had a complete, correctly gated Claude Agent SDK lane for
a while, but no renderer ever asked for it: the launch route accepted only
`codex`, and the create path was typed `agent: 'codex'` end to end.

Widen both to the structured provider union that already exists
(`AgentSessionHandleProvider`), and generalize the codex-named create path
instead of adding a Claude twin beside it. The pending-launch registry is now
keyed by agent as well as workspace — a shared key handed a second caller the
first agent's intent, so a Claude and a Codex launch in one worktree collided.

Windows, per agent. Codex's client-side win32 refusal is deliberate and settled
elsewhere, so it stays exactly as it was. Claude's answer is no longer guessed
from the client's platform: a structured session fences its provider child on
that child's process start time, and only the executing host knows whether it
can read one. `agentSession.createSupport` already answers precisely that, per
agent, and had no renderer caller — so the Claude create path asks it before
creating and turns a "no", or a probe it cannot get answered, into the
definitive refusal the launch fallback already handles. Fail closed either way.

That refusal mapping also closes a real gap: the host reports an unsupported
location by throwing `structured_agent_session_unsupported`, which reaches the
client as a transport rejection rather than a refusal envelope, so
`StructuredAgentSessionCreateRefusalError` never fired. The launch would retry
the create, strand itself in `visibilityUnknown`, run no legacy fallback, and
show an error toast.

Close a fail-open hole while Claude and win32 become reachable: `create` with a
client-supplied location, and `ensure`, both skip the worktree-resolving support
check. They now ask the executing host the same question directly, so a host
that cannot fence a provider child no longer creates one on a client's say-so.

Also deletes `structured-agent-session-provider-routing.ts`, a duplicate of
`structured-agent-session-provider-support.ts` with no importers.

WSL, SSH and paired hosts, floating workspaces, draft prompt delivery, explicit
TUI customization and initial session options all keep refusing; folder
workspaces keep working.

* P1-1: make the structured Claude auth policy required and testable

The optional dep plus a {stripAuthEnv:false} fallback meant a dropped wiring
under-stripped silently. Required at all three hops, asserted at install time for
the @ts-nocheck caller, and the settings-to-policy mapping is now a named tested
function.

* P2-3: mobile's default Claude transcript root must follow CLAUDE_CONFIG_DIR

session-file-resolver's default ignored the variable the pinned account home
follows, so a CLAUDE_CONFIG_DIR launch wrote one tree and mobile read another. The
Task-4 test now resolves with no root override (mobile's own call) and checks the
answer against the root the CLI itself reports, instead of mirroring the code under
test's own expression.

* P2-1/P2-2/P3: close the teardown window, join the live-auth gate, align the refusal

P2-1: a switch beginning inside the acquire teardown left a dead chat and no
replacement. Past that point the launch waits the swap out and refuses only if it
never settles; the entry guard still refuses outright, because nothing is torn down
there yet.
P2-2: structured children now hold the same OAuth-refresh gate a Claude PTY does,
so a managed refresh cannot rotate the token out from under a live turn.
P3: the refusal now matches the strip it guards (case-folded on win32, presence not
truthiness), and the dead structured-to-TUI builder states its auth policy instead
of silently signing a system-auth user out.

* Make the live-auth gate tests independent of sibling connection teardown order

* Do not offer structured Claude under a WSL-only managed account

Structured Claude launches against the ambient Claude config, which the account
service keeps in sync with the selected HOST account. A WSL-bound managed
account lives inside the distro and is never synced there, so on Windows a
structured session would authenticate as whatever the ambient identity happens
to be while the UI names the WSL account — the user is told one identity and
given another.

That was unreachable only because nothing offered structured Claude on win32.
Enabling it makes it reachable, so gate it here rather than patching the auth
layer: refuse the structured path when the active managed Claude account is
WSL-bound, and let the terminal-backed path — which resolves the account per
runtime — handle that account shape.

The answer rides the agentSession.createSupport seam the renderer already
consumes, so no new capability and no renderer knowledge of account internals.
A create the host declines becomes the definitive refusal the launch fallback
already turns into a legacy native chat tab, with no error toast.

Unknown answers refuse. An install with no managed accounts claims no identity
and is fine, but an active selection that cannot be resolved — or account state
that cannot be read at all — is not evidence that the ambient identity is right.

Claude only. Codex resolves its account through a different path and its
createSupport answer is untouched, as is every Codex routing decision.

* Read the structured Claude account gate through the auth policy's accessor

The gate resolved the active account from the account-service snapshot's
runtime map; the auth policy resolves it with
getSelectedClaudeAccountIdForTarget(settings, { runtime: 'host' }). Those are
two sources and two resolution rules, and they disagree on a legacy settings
blob that carries the selection only in the flat activeClaudeManagedAccountId:
the accessor falls through to it, a direct read of the runtime map does not. The
gate would then refuse a launch the policy would have run under host-1 — and in
the mirror case a session could be admitted under a policy computed from a
different account than the gate approved.

Read the same settings through the same accessor so agreement is structural
rather than coincidental, and drop the controller accessor that existed only to
reach the snapshot.

No behaviour change for any state both already agreed on; Codex is untouched.

* Round-3 review fixes: N-1 empty-value regression, N-2 gate leak window, N-4 lost history

N-1: my presence-based conflict predicate refused a terminal launch that works
today. 'ANTHROPIC_API_KEY=' is how a user blanks a variable and the settings
pipeline preserves that empty value; an empty override cannot beat the pinned
account and the strip removes the name anyway. Back to truthiness for the value,
keeping the win32 case folding.
N-2: enter the live-auth gate only after the exit/close handlers that release it,
so no throw in between can leave an entry nothing reconciles.
N-4: the Claude transcript resolver searches config-dir-then-default and de-dupes,
matching the Codex sibling in the same file, so adopting CLAUDE_CONFIG_DIR no
longer hides history written before it.

* Run the managed-account gate on every Claude acquisition, not just create

createSupport gates the create path, but a session's account state can change
while it lives. A reacquire after an unexpected child exit re-resolves the
launch and re-derives auth, with nothing re-checking the gate — so a session
created while supported could come back up in the refused shape. With the strip
predicate keyed on there being an active non-WSL account, the WSL-only user's
normalized steady state (accounts exist, none active) does not strip, and that
reacquire reaches the child with ambient auth while the UI names the account.

Gate at resolveLaunch, the one choke point every acquisition passes through,
refusing with the pre-spawn error the caller already handles. Same predicate as
create-time, now sharing one settings reader so the two cannot drift.

Claude only; Codex resolves its account on a different path and is untouched.

The runtime class that wires this does not typecheck its own `this` calls — a
missing hookup compiles clean — so the wiring is pinned behaviourally rather
than trusted to the compiler.

* Move the structured Claude gate out of the @ts-nocheck runtime files

Both call sites of the managed-account gate sat in files whose first line is
`// @ts-nocheck`, so neither was typechecked: three arguments to a one-argument
function plus an undeclared identifier compiled clean. New auth-identity
decision logic had no compiler behind it.

Move the verdict into a checked module that takes the two facts the runtime
owns — the adapter's answer and a settings getter — and decides. The runtime
class now only forwards. Move the gate reader's construction into the checked
installer too, so the nocheck file passes a plain settings closure and never
names a gate symbol.

Every reference to the gate predicate and its reader now lives in a checked
file, so the ablation that used to pass silently is a compile error at both the
create-support and reacquire sites.

Removing the file-level @ts-nocheck is a separate, larger job and is not
attempted here.

* Derive the gate test's auth policy from the settings under test

A hardcoded stripAuthEnv asserts a gate/policy pairing production cannot
produce, and false additionally lets launch.env inherit the runner's real
process.env. Derive via claudeStructuredAuthPolicyForSettings instead: the
gate settings type is the same Pick the policy takes, and both resolve the
account through getSelectedClaudeAccountIdForTarget.

* Pin the absent-vs-empty distinction in the managed-account gate

An empty claudeManagedAccounts array is a real answer: the user has no managed
accounts, nothing claims an identity, and the ambient path is legitimate. A
readable settings object with no such field is settings we failed to parse —
the same unknown as unreadable — so it refuses.

The two are one character apart in the code and the difference is invisible
without the reasoning, so record it at the branch and pin both sides. The test
fails under the obvious "consistency fix" of treating a missing field as empty.

* fix(claude): keep command queue bookkeeping out of the transcript

Claude Code 2.1.258 emits a `command_lifecycle` frame for every uuid-stamped
command it starts, completes or cancels. The frame carries a command uuid and a
state and no content, and the CLI keeps it out of its own transcript -- but it
is absent from the SDK's SDKMessage union and so from Orca's frame catalogue,
where an uncatalogued kind defaults to a substantive row. Every structured turn
therefore painted raw JSON rows into the user-visible transcript.

Catalogue it and disposition it as status chrome. The unknown-kind default stays
`timeline-substantive`: a kind we have never seen is likelier to carry content
than to be chrome, and a visible row we can catalogue later beats content we
silently dropped. A lifecycle state that reads as a failure still surfaces,
because the payload error check in `classifyProviderFrame` outranks the
catalogue.

* fix(claude): let a re-walked descendant become eligible for the forced sweep

A descendant first observed by a capture inside its own birth second could never
be SIGKILLed: `ps lstart` is second-resolution, so that capture cannot rule out
a pid recycled later in the same second, and the merge pinned each retained row
to the boundary of the walk that first saw it. SIGTERM-resistant children forked
in that window were signalled and then never escalated -- they survived close,
quit and restart, reparented to init, and had to be killed by hand.

Advancing that boundary on any later capture would be unsound: a later capture
matching pid, pgid and start-second is exactly what an impostor would also show.
But a capture is not a match -- it is a fresh ppid walk from a root Node pins
through its own handle, so a row it re-derives is proved ours at that instant
without appealing to its start time. Chain the fence from there instead, and
take that walk at the close boundary while the root certainly still lives: the
root may leave inside the grace window, and the post-timeout refresh never runs.

A row absent from the later walk still keeps its earlier boundary, and a row no
walk has ever re-derived in a later second is still never escalated.

* Treat an absent managed-account list as empty, not as unreadable

An empty claudeManagedAccounts array and a missing one are the same answer:
this user has no managed Claude accounts, so nothing claims an identity and
ambient auth is the truth. Refusing on absence strands any profile that simply
never wrote the key, and it disagrees with the auth policy, whose own predicate
takes `(accounts ?? [])` for exactly this reason.

Only settings that cannot be READ stay unknown, and those still refuse — as do
a WSL-bound active account and a selection naming an account the list does not
explain.

The earlier reasoning treated a missing field as settings we failed to parse.
That conflated "not present" with "not readable"; only the second is unknown.

* Support structured Claude when accounts are registered but none is selected

Registered-but-deselected Claude accounts were refused, which is behaviourally
identical to having no accounts at all: the auth policy does not strip, ambient
auth is the truth, and the UI names no host identity. A user who deselected
their accounts silently got legacy chat with nothing explaining why.

Nothing selected for the host runtime is two states the settings cannot tell
apart after the fact, because pruneInvalidClaudeRuntimeSelection empties the
host slot and persists null in the second one:

  honest deselection      -> ambient auth, UI names nothing   -> SUPPORTED
  the WSL-only steady state -> ambient auth, UI names the WSL account -> REFUSED

The presence of any WSL-bound account in the list decides. Simplifying this to
"none active -> supported" re-opens the auth-identity misrepresentation, so the
tests fail loudly on exactly that: five of them, across the unit rule and the
createSupport path.

* Stop treating an unanswerable create-support probe as a refusal

A worktree is not resolvable for a beat after createWorktree resolves, so a
probe fired immediately after creation fails the RPC with selector_not_found
instead of answering. The catch collapsed that into `supported = false`, so the
composer refused and quietly built a terminal session — the gate never said no,
it was never asked successfully. Elapsed time was the only input that decided
whether a Claude launch went structured.

"Could not answer" and "answered no" are different states and only the second
is a verdict. Retry while the host cannot yet resolve the selector, with a
bounded backoff that covers the measured window with margin, and keep refusing
on the first ask for everything else. Fail-closed is unchanged: a probe that
still cannot be answered when the budget is spent refuses.

The retry is narrowed with the shared error-code matcher, which classifies a
token that transports re-wrap into a longer message without matching prose that
merely mentions it.

Codex never probes, so this race has never been able to refuse a Codex launch —
the race itself is identical for it. Recorded at the early return, because
whoever gives Codex a probe inherits the bug.

* fix(claude): fence the forced sweep on re-derivation, not on lstart's second

A descendant forked in the same wall-clock second as every walk that sees it was
signalled with SIGTERM and then never escalated, so a SIGTERM-resistant child
survived tab close, app quit and a full relaunch. Two children of one parent
96ms apart across a second boundary took opposite paths. The leak predates this
branch: it reproduces with the change reverted.

`ps lstart` has one-second resolution, so a walk landing inside a row's birth
second can never rule out a pid recycled later in that same second. But a walk
is not a match: a ppid walk only reaches what the root actually parents, and the
root is pinned by Node's own handle, so a row the walk re-derived is ours
whatever second it was born in -- a stranger would have to have been forked into
our tree, and then it is not a stranger. Fence the escalation on that.

Rows a merge retained from an earlier walk are not re-derived and still answer
to the start-time fence, which remains correct for them.

Scoped to callers that revalidate identity before signalling, which is the
Claude close path. Codex teardown reaches this same verifier and is unchanged;
the argument holds there too, but widening it is its own deliberate change.

Also reverts two changes from the previous attempt at this leak. Advancing the
capture boundary on a later walk is inert once the sweep fences on re-derivation
-- both key on the same set of rows, so the new term short-circuits for exactly
the rows whose boundary it advanced. The extra ladder refresh was a duplicate
full process-table read: close() already awaits tree.refresh() immediately
before proveClaudeChildExit, on the only path that reaches it.

Known property: the kill lands roughly a grace window after the walk that proved
membership, so a pid recycled inside that gap could in principle be signalled.
It is bounded -- matchingSnapshotRows already requires the live row to carry the
same start-second and pgid, so an impostor must be born in the remainder of that
one second, land on that exact pid, and sit in the same process group, and it
has already received the unfenced SIGTERM from the same loop.

* Run the Claude structured integration suite as a runtime client

The suite exercises agentSession.* for Claude, not the mobile surface: nothing
in it asserts anything mobile-specific and its sibling integration suites use
'runtime'. Mobile now additionally requires the experimental structured-chat
setting, which structured-agent-session.test.ts pins in both states, so the
stale 'mobile' fixture was claiming coverage it never had.

* fix(claude): report effort from get_settings, which is the only frame that has it

The composer's Effort pill rendered blank in every structured session. This is
not a missing source: the publication reads `effortLevel` off the `system/init`
frame, and that frame has never carried an effort of any kind, while the correct
value is already fetched at acquisition and thrown away on the auth diagnostic.
Verified two ways -- a live get_settings probe against Claude Code 2.1.258, and
the shipped binary's own init frame construction, which lists `model` and no
effort. So `reportedOptions.effort` was always empty, the options reader dropped
the key, and the pill had no value. Model survived only because
`currentModelId()` has a fallback chain.

The get_settings call acquisition already makes reports the session's current
effort as `effective.effortLevel`; pass that into the publication instead.
Selecting an effort already worked, so this is the arrival value only.

The legacy PTY path is unaffected and must not be "fixed" to match: it reads its
effort by parsing the startup banner (`CLAUDE_MODEL_EFFORT` in
src/renderer/src/components/native-chat/claude-terminal-session-options.ts),
which is why it shows a value where the structured path does not.

Also removes the fixture that hid this: the fake init frame invented
`effortLevel: 'high'`, a field the CLI does not send, which is why every gate
stayed green over a value that is always empty in production. The fixture's
get_settings now returns the real {applied, effective, sources} shape instead of
a bare `{env: {}}`, so the two adapter tests that asserted an effort keep
asserting it through the path production actually uses.

The reader returns null rather than defaulting: an effort nothing measured would
repeat the fixture's mistake, and a blank pill is the honest degradation if the
provider ever renames the key.

* fix(claude): only record an effort the child confirms it adopted

apply_flag_settings answers `success` for an effort it then ignores. Measured
against Claude Code 2.1.258: applying `bogus-effort-xyz` returns
subtype "success" with no error while `applied.effort` stays at its previous
value, and a valid `low` moves it. The option write treated the absence of a
throw as adoption and recorded the requested value unconditionally, so Orca
would show and persist an effort the child was not using, with nothing anywhere
reporting a problem.

Read the effort back after applying it, through the same reader the arrival
value uses, and reject when the child reports a different one. A readback that
could not be taken is not evidence of a refusal -- the apply itself succeeded --
so it still records; only a readback that disagrees rejects.

Not reachable from today's picker, which offers catalog values only, but the
CLI's effort catalog is server-delivered and has changed before, so a retired id
would otherwise become a pill confidently displaying a setting that never took.

* test(claude): assert the effort contract against the real binary

The blank pill survived every gate because the only tests that touched it were
fixture-backed, and the fixture invented the field. A test that pins the shape
we read cannot catch the provider renaming the key, which is the failure mode
that produced this defect.

Asserts both halves against a live authenticated CLI: that no frame it publishes
carries an effort at all, and that the session's current effort arrives through
get_settings. Which frame proves the session varies by host -- this machine
proves it with a SessionStart hook rather than a system/init frame -- so the
negative half asserts over every published frame rather than picking one.

Skips with the rest of the file when no authenticated CLI is present.

* fix(claude): stop the synthesised content-part kinds leaking into the transcript

Sending an image put a bare `claude · message:user:content:image` row between
the user's bubble and the answer. Two causes, and only the second is a family.

An image part counted as modelled only when `source.type === 'url'`, but
claudeDispatchMessageContent sends a local attachment as a base64 source and the
CLI replays that shape back, so every attached image was classified unmodelled.
Accept the base64 and file sources Orca itself sends.

The family is the real defect. `message:<role>:content:<type>` kinds are
synthesised at runtime from whatever `part.type` arrives, so unlike the
top-level frame catalogue they can never be enumerated ahead of time -- the
`?? 'timeline-substantive'` default then prints the synthesised name at a user
who cannot act on it. That default is right for top-level frames, where
"substantive" means show the frame; here it meant show our own vocabulary, which
drops the content AND leaks the opcode.

So an unrenderable part now renders a sentence saying exactly that, with the
kind and payload still on the row's disclosure. A part that carries its own
readable sentence keeps it -- the placeholder is a fallback, not an override.

An unknown future part type is therefore visible, never silently dropped and
never printed as a kind: the same principle as the effort readback, which
records only what the provider confirms.

* Declare agentSession.requestHandoff on the cross-version wire surface

The manifest is a ratchet for cross-version reachability, so the method is
declared with real HandoffParams rather than counted. requestHandoff is
capability-gated through requireStructuredHost and has no client caller, so
declaring it is the whole of the change.

Also model two host capabilities the harness omitted: the stub host's
supportsCreate, and the fake adapter's, without which adapterSupportsCreate
falls through to a supportsLocation the fake also lacks. Every ensure was
refused for the harness's silence rather than for its location.

* Gate structured Claude session tabs on the client capability that names them

The Claude structured lane deleted the projection's `agent !== 'codex'`
filter and added CLAUDE_STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY in the
same commit, but never wired the constant to anything. Paired clients then
received agent-session tabs for Claude, which no shipped client renders --
mobile's resolveMobileNativeChat returns null for every agent but codex, so
the row listed and selected into a pane with neither chat nor terminal.

Restore the filter behind the declared capability instead of the bare agent
name. No client advertises it yet, so this matches main's behaviour today
and becomes a negotiation a future client can opt into.

* Confirm the structured Claude model against the model the CLI reports

set_model answers success for any string, including a model it cannot
resolve — the failure only surfaces when the turn runs — and get_settings
reports the settings-file model, not the session's. The init frame that
opens each turn is the only channel carrying the adopted model, so keep
the session's reported model current from it instead of reading it once
at acquisition.

Also stop rejecting an effort the readback cannot represent: max is
session-scoped and excluded from the persisted effortLevel, so a readback
reporting the level underneath it is an absence of evidence, not a refusal.

* Clear the session-option hedge when the provider confirms the value

The pill claimed every option was unconfirmed for the life of the session:
the renderer recorded each write as dispatched and nothing ever moved it,
so a model the CLI had already reported back still read as unconfirmed.

Carry the provider's own confirmation to the surface. Main reports which
option ids the provider named rather than merely accepted, and the client
re-reads options as a turn changes, because the frame that opens a turn is
where the adopted model arrives. A value the provider has not reported
stays hedged, including an effort whose readback could not be taken.

The confirmed list is optional on the wire: a host that predates it sends
nothing and the client keeps hedging, which is the behaviour it had.

* Keep the model report current across an acquisition fence bump

* Show the picked session-option value and let the provider report correct it

The pill showed a "not confirmed" second tooltip line for any value we had sent
but not yet seen reported back. Nothing acts on it, and for the PTY lane it was
permanent — that transport has no report channel. The pill now shows the picked
value immediately and the provider's per-turn report corrects it when the two
disagree; a newer local write still outranks a report that precedes it.

`dispatched` stays as a provenance member rather than collapsing into `applied`:
it is produced independently by the PTY lane, and it is where the `confirmed`
wire field lands, which would otherwise be unobservable.

Effort keeps its readback and its rejection path. That matters more now, not
less: with the hedge gone the rejection is the only user-visible failure signal
on this surface, so a spurious one would be the loudest bug here. Skipping the
readback for an effort the settings response structurally cannot echo is what
prevents it — the response carries the persisted level, so reading it back for a
session-scoped value would report the level underneath and fail a valid write.

* Hedge a session-option value only when the terminal transport sent it

Both lanes emit `dispatched`, so it could never say which one produced a value.
The descriptor now carries the transport that built it, set once in the shared
snapshot builder from a parameter that is required rather than defaulted — the
builder is the only place a descriptor is constructed, so a new producer has to
name its lane or fail to compile.

The structured lane confirms every value from the provider's own per-turn report,
which makes the hedge transient noise there. The terminal lane can only learn an
outcome by parsing the screen back, and only for Claude: every other agent's
`dispatched` value stays unconfirmed for the life of the session, so the line is
the only signal that we sent something we never saw land.

* Refuse an effort the session's model advertises no control for

* Refuse tab mutations on a Claude row the client never negotiated

The branch added a case asserting a client advertising only
agent-session.structured.v1 may mutate a claude row. That is the same
ungated behaviour the projection gate removes, encoded a second time —
mutation authorization reads the projection, so hiding the row refuses
the write. Assert that contract instead, and add the positive case for a
client that does negotiate Claude rows.

* Resolve the Claude session's current model in one place so the effort guard and the pill agree

* Record an effort the child did not adopt instead of refusing the write

apply_flag_settings answers success for an effort it then ignores, so the
readback exists to detect that. Refusing on it made the detection a veto,
and a veto is only correct if the readback can never be wrong about which
model is current -- which it was, twice. The pre-flight guard already
refuses a level the model advertises no control for, so the veto guarded a
door that is now locked upstream.

Keep the detection, drop the refusal: a disagreement records the child's
own answer and omits the option from confirmed, so main stops vouching for
a value the provider rejected without blocking the user's write.

* Stop a slow whole-machine ps from being read as an absent process

`ps -axo ...command=` pays a per-pid argv read: measured 1.15s for 1,948
processes (0.03s without `command=`), and CPU contention stretched the same
capture to 6.0s. Two budgets sized for a cheap look then misreport a readable
machine.

The reader's 3s ceiling killed 6 of 20 consecutive captures at load 27, so
every consumer answered "unverifiable" about a table it could read. Raise it
to 15s, and stamp the capture instant at ps START so `capturedAgeMs` is the
upper bound its contract promises -- a 6s capture used to report itself as
freshly taken, understating staleness against a 5s kill gate. The TTL keys on
completion so a slow capture still coalesces instead of forking ps per caller.

`readStructuredTuiProcessIdentity` then spent its whole 5s wait inside one
capture and concluded "no exact child" after a single look taken before the
child existed (observed landing at ~3.5s). Absence needs a look that did not
race the spawn, so require two captures before the deadline can end the loop.

Both surfaced by the real-binary Claude TUI resume test, which failed ~1 in 5
under load; 14/14 now, 8 of those runs containing a capture the old 3s budget
would have killed.

* Let the desktop renderer negotiate Claude structured tabs

The paired-client gate hides agent-session rows an agent the client cannot
render. The desktop renderer's own IPC dispatches as clientKind 'runtime'
advertising only agent-session.structured.v1, so the gate hid Claude rows
from the surface this feature ships on. It renders them; it should say so.

* Stop a slow process table from silently blinding every freshness gate

Stamping `capturedAgeMs` at ps START made the number honest, and honest broke
both consumers that read it. `ps -axo ...command=` measured 2.5-9.0s on an idle
2,002-process laptop and 4.0-18.6s at load 46, so the age it now reports lands
past every budget: `planRelayPtySweep` refuses the stop as "too old", and the
renderer's `admitRemoteForegroundEvidence` refuses the record outright. That
second one is the expensive half and was outside the diff -- a refusal bumps
`consecutiveInspectionErrors`, the poll scheduler backs off to its 10s floor,
and agent-completion detection stops for the pane. The subsystem went blind on
exactly the loaded hosts the honest stamp was meant to serve.

The evidence-publishing read now gives up at 1,200ms instead of waiting out
`PS_TIMEOUT_MS`. It is one budget for one question: these consumers ask whether
an observation describes NOW, and past this it does not -- a late answer is
refused by the age gate anyway, having first blocked a polled path for the whole
capture, so a prompt `unverifiable` is both the truthful verdict and the cheap
one. Both relay call sites already produce it from a rejection, and an admitted
`unverifiable` costs a poll where a refusal costs the cadence. Identity proof
keeps the full 15s through `getFreshProcessTableSnapshot`, because it asks
whether a process EXISTS and must never read slow as absent. The budget bounds
the wait, never the capture: the reader coalesces, so an abandoned wait leaves
its capture running to fill the cache rather than forking a second whole-machine
`ps` on the host that can least afford one.

1,200ms is bracketed rather than picked. The floor is the capture's own cost --
`command=` measured 1.15s for 1,948 processes on an idle host, and a budget
under that answers `unverifiable` about a machine nobody is straining. The
ceiling is the consumer's: 2,000ms, less the 500ms a TTL-shared capture may
already have aged, leaves 1,500ms, and transit takes the rest.

That ceiling only fits once the capture stops being charged twice. `ps` runs
inside the RPC round trip, so its duration is already in `receiveDelay`, and
`capturedAgeMs` is that same duration on the host's clock; summing them halved
the budget this gate grants a host from ~2.0s of `ps` to ~1.0s, which is why a
1.2s capture arriving at 1.3s read as 2.5s old and was refused. Admission now
takes the larger of the two. The sweep's gate keeps its sum, which is correct
there: `evidenceAgeSinceListingMs` is stamped after the listing ARRIVES, so it
measures planning time and overlaps nothing.

A stated limit rather than an assumed one: 15s is not proven sufficient for
identity proof. The same capture reached 18.6s at load 46, so that path can
still time out and answer "no exact child" about a host it simply could not read
in time. Narrowing it needs a cheaper question than a whole-machine argv read,
not a larger number.

The one test guarding this field could not fail. `beginPtyHandlerTest` installs
fake timers, so `Date.now()` is frozen, the real reader reports exactly +0, and
`0 <= 500` held identically for a hardcoded zero, for completion-stamping and
for start-stamping -- while the real reader on that host returns thousands of
ms. It now drives a measured age in and asserts the handler publishes it rather
than restamping; that the reader MEASURES it correctly stays pinned separately,
against a controllable clock. Both consumers get boundary coverage either side,
and each new gate was ablated red before it went green.

* Keep the compatibility fields off the capture the budget just abandoned

inspectProcess falls back to processHasChildren and listProcesses to
getForegroundProcessName, and both read the same TTL-shared capture with
no budget of their own. On a slow host they joined the in-flight capture
the budgeted evidence read had just given up on, so the call still blocked
for the full 6-18s and the budget bought nothing -- once for inspectProcess
and once per managed PTY for listProcesses.

Use the degraded answers those helpers already give for an unreadable
table, reached promptly. pty.hasChildProcesses keeps its unbudgeted fresh
probe: it is a one-shot destructive gate that can afford to wait.

---------

Co-authored-by: Merge Sim <merge-sim@local>
Co-authored-by: Merge Sim <sim@local>
2026-09-04 15:55:20 -07:00
Neil f7e3af254a fix(pty): close the pseudoconsole and dispose the conout worker on Windows self-exit (F24) (#18635)
* fix(pty): close the pseudoconsole when a Windows shell exits by itself

`ClosePseudoConsole` is the only thing that reaps a ConPTY's console host.
node-pty calls it from one place, `PtyKill`, which starts by looking the baton
up by id -- and the exit watcher in `SetupExitCallback` erased that baton the
moment the shell died. So on the self-exit path (typing `exit`, how panes
usually close) the lookup missed, `PtyKill` did nothing at all, and the
pseudoconsole was never closed.

The baton now survives until BOTH the shell has exited and `kill()` has run;
whichever arrives second frees it. `PtyKill` copies `hpc` out under the lock and
closes it afterwards, guards `TerminateProcess` on a shell handle the watcher
may already have closed, and duplicates that handle rather than reordering, so
upstream's close-then-terminate sequence is unchanged.

Measured on Windows 11, 20 self-exit cycles driven exactly as Orca drives them
(`onExit -> destroy()`), handles bucketed by NT object type:

  relay spawn (no useConptyDll)   225 -> 285 (+1 Process +2 File/term)
                            after 219 -> 219  FLAT

  desktop spawn (useConptyDll)    239 -> 439 (+1 Process +2 Thread +5 File/term)
                            after 235 -> 395 (+2 Thread +4 File/term)

The desktop residue is a separate defect in the `useConptyDll` branch of
`WindowsPtyAgent.kill()`, which disposes the conout worker only from an
`_outSocket.on('data')` handler -- and no data arrives after the shell has gone.
Fixing that line as well takes the desktop to 222 -> 222 FLAT, but it lives in
the `kill()` hunk owned by F23, so it is left to that change.

Refs F24.

* fix(pty): dispose the conout worker when a Windows shell exits by itself

Second, independent defect on the same self-exit path, and the larger half of
the desktop's leak. The `useConptyDll` branch of `WindowsPtyAgent.kill()`
disposed the conout worker only from an `_outSocket.on('data')` handler -- and
once the shell has gone no more data ever arrives, so the worker was never
disposed. The non-DLL branch three lines above already disposed unconditionally,
which is why only the desktop (the only spawner that sets `useConptyDll`) hit it.

Measured on Windows 11, 20 cycles, handles bucketed by NT object type, totals:

  self-exit,     relay spawn     225 -> 285   now  219 -> 219  FLAT
  self-exit,     desktop spawn   239 -> 439   now  222 -> 222  FLAT
  explicit kill, relay spawn     225 -> 285   now  219 -> 219  FLAT
  explicit kill, desktop spawn   235 -> 395   now  219 -> 219  FLAT

Neither fix alone is enough on the desktop: the pseudoconsole close is worth
+1 Process +1 File per terminal, this dispose +2 Thread +4 File.

The relay asset (config/relay-assets/node-pty-1.1.0-windows-pty-teardown-patch.cjs)
deliberately gets no counterpart: the relay takes the non-DLL branch, where the
dispose is already unconditional. Its reconstruction table needs the new hunk
though, or un-applying the desktop hunks no longer yields published node-pty.

Taken over from F23 at win-relay-qa's request after they verified that the
desktop never executes the non-DLL branch F23 was scoped around.

Refs F24.

* fix(pty): harden PtyKill against a failed handle duplication and a missing DLL

Both from review of #18635.

DuplicateHandle's result was dropped. On the live explicit-kill path a failed
duplication left hShellDup null, which the guard below could not tell apart from
the self-exit case, so TerminateProcess was skipped and the shell kept running
after its pane closed -- a worse outcome than the handle leak this patch exists
to fix. The failure now terminates through handle->hShell under the lock, where
it is valid and where TerminateProcess does not block. The only cost is that the
rare path kills before the console closes instead of after.

LoadConptyDll is now resolved BEFORE any baton state is touched, matching what
PtyConnect already does for the same reason. It throws when conpty.dll is
missing, and a throw after consoleClosed was set would strand the pseudoconsole
permanently: the retry finds the work claimed and does nothing.

Also corrects three comments the earlier commits made stale:
  - the ptyJobMutex note still said PtyKill reads the table unlocked
  - PtyListJobProcessIds said the baton is gone once the shell exits; it now
    outlives the shell, and the nulled hJob is what makes the answer null
  - windows-pty-job.ts said node-pty drops its handle record on exit

Re-measured on Windows 11 with the rebuilt binary, 20 cycles, all four paths
still flat: self-exit relay 219->219, self-exit desktop 222->222, explicit-kill
relay 219->219, explicit-kill desktop 219->219. Both explicit-kill runs report
22/22 shells exited, so the kill still lands.

Refs F24.
2026-09-04 05:32:19 -07:00
Neil 8463dcb7b9 fix(terminal): make wrapped-line search rewind iterative and bound its scans (#18402)
Patches @xterm/addon-search so one very long un-newlined line no longer overflows the stack, freezes the renderer, or goes unsearched. Submitted upstream as xtermjs/xterm.js#6149 (issue #6148); drop the patch once a release ships it. See the PR for measurements and the differential fuzz.
2026-09-03 17:59:16 -07:00
Neil 39330c5aca fix(relay): retire PTYs the host proves are gone, and stop two per-poll scan storms (#17832)
* fix(relay): stop three CPU growth terms in a long-running remote session

pty.resize gated only on `managed.disposed`, which is bookkeeping rather than
liveness. A shell that exits without node-pty's `onExit` leaves an undisposed
entry holding a closed master fd, and UnixTerminal.resize has no fd guard, so
the ioctl threw `ioctl(2) failed, EBADF` into the dispatcher's generic
parse-error catch. Nothing retired the entry, so it stayed advertised and kept
activePtyCount above zero -- which is what stops a relay with an unlimited
grace from reaching its idle-no-ptys exit (#12423). Probe liveness with the
same helper attach/listProcesses use, retire a provably dead pid, and contain
an ioctl failure over a live-or-unverifiable process.

processHasChildren forked `pgrep -P` per pane per inspection poll, uncached.
procps-ng opens six procfs files per process to resolve one ppid, so each call
cost O(host process count). Answer from the TTL-cached `ps` table the same RPC
already captured for the foreground lookup (#13537).

The remote AI Vault scanner had no parse cache at all, so every forced rescan
re-read and re-parsed the whole transcript corpus, including files untouched
for a month. Give it the mtime+size keyed memo the local scanner has (#13753).

* fix(pty): invalidate the descriptor when node-pty gives up the handle (#17930)

Carried forward from PR #17930, which merged into this branch. Rebased onto
current main; main's newer node-pty-fd-leak test is kept as-is.

* fix(ai-vault): refresh codex titles on the remote parse-cache reuse path

The remote cache keys on the transcript's (mtime, size, host), but codex
titles live in $CODEX_HOME/session_index.jsonl and are written after the
rollout — so a cache hit froze the fallback title forever. Mirrors the
local scanner's existing reuse-path refresh via a shared core.

* fix(relay): publish the exit a reap performs, and rescan for close decisions

Two review findings on the CPU work.

reapExitedPty told only the relay-internal exit listener, so a retirement left
the client's pane mounted against a session the relay had already forgotten --
the next attach answered `PTY "<id>" not found` with nothing before it to
explain why. Pre-existing on three probe paths; resize made it user-triggered.
Publish the same pending-exit the natural onExit path publishes, carrying -1
("gone, status unrecoverable"), and skip it when onExit already reported the
real code.

processHasChildren now answers from a 500ms TTL-cached table. That is right for
pty.inspectProcess, which every tracked pane polls, but pty.hasChildProcesses
gates the window-close confirmation and workspace cleanup's idle evidence --
one destructive decision per answer, where a child started inside the window
would be killed unasked. Give that RPC a fresh scan; pgrep used to.

* fix(relay): publish a reap's exit only on proven-exited evidence

The publication is a verdict the client acts on by retiring the pane, so it
must not be reachable from the disposed-record sweep, which retires off our own
bookkeeping rather than the host's process table. Only ESRCH earns it.

* fix(i18n): restore the activity-options key the rebase dropped

* fix(i18n): union en.json with main so the rebase cannot drop keys
2026-09-02 15:14:44 -07:00
Neil 8197268956 fix(pty,remote): close the pty master fd leak, and two remote-terminal defects (#17914)
* fix(pty,remote): close the pty master fd leak and two remote-terminal defects

so on Linux every later child of the process -- both later pty children and
plain child_process spawns -- inherits it and keeps the /dev/pts device alive.
Measured on Linux with stock node-pty 1.1.0: master fd flags 0404002
(cloexec=false), and 17 -> /dev/pts/ptmx present in both a later pty child's
/proc/self/fd and a later child_process child's. Extend the existing node-pty
patch with pty_cloexec() on both PtyFork spawn paths; after the patch the flags
read 02404002 (cloexec=true) and neither child sees the master. This covers the
app and terminal daemon only -- the SSH relay installs node-pty from npm on the
remote host, so it stays exposed (see the report).

rejecting inspection as a renderer-global unhandledrejection, which an
unreachable runtime produced on every cadence tick.

path cleared the close intent for it exactly like a dropped connection, so a
host that keeps republishing the dead surface re-materialized the pane the user
just closed. Keep that intent and drop its TTL. Also route the banner's
"Remote terminal was closed." line through translate() so it stops mixing
English into a localized banner.

* test(pty,remote): make the fd-leak evidence positive and size the close intent to its RPC

The Linux 'does not hand an earlier pty master to a later pty child' case only asserted that ptmx was absent from the captured listing, so any run that produced no listing passed without inspecting a single fd. Block the child on stdin, emit a sentinel, and assert both the sentinel and a real /dev/pts fd row before the negative assertion. Verified in node:24-bookworm: passes with the patch, and with pty_cloexec() reverted it fails on four inherited /dev/pts/ptmx rows.

The close intent's TTL was a 10s literal while the close RPC that can still answer tab_not_found had its own 15s literal. A host that answered slowly while republishing the surface had its intent evicted by the republish path's own pending-check, so makeWebSessionCloseIntentDurable found nothing to flip and #9194 reproduced. Derive the TTL from the shared session.tabs RPC timeout so the two cannot cross, with an invariant test and a regression test for the slow answer.
2026-09-02 02:23:45 -07:00
NeilandBrennan Benson fbe94ceff6 fix: close readiness gaps found by merged-change audit (#17159)
* fix(ssh): fence stale kills and retired pane replay

* fix(ssh): support cancellable interactive authentication

* fix(ssh): await remote catalog before snapshot adoption

* fix(pty): contain Windows ConPTY input failures

* fix(power): avoid redundant macOS display blocking

* perf(editor): narrow markdown override subscriptions

* fix(quick-open): close directory handles after reads

* refactor(linux): remove unused proc socket scanner

* fix(usage): apply flat Sonnet 4.6 pricing

* ci: prime Node next native test cache

* docs(skills): resolve snapshot cleanup data path

* fix(ssh): recover install locks after host reboot

* test(ssh): recognize boot-aware install locks

* test(ssh): prove previous-boot lock recovery live

* test(wire): pin pre-metadata release coverage

* fix(terminal): preserve remote tab ownership through recovery races

* test(runtime): fence replaced terminal handles in agent guard

* fix(ssh): preserve remote snapshot authority across polls

* fix(pty): contain late ConPTY output EPIPE

* test(pty): register Windows exit watcher before kill

* fix: close SSH and tab readiness race gaps

* fix(tabs): retain headless order and placeholder titles

* fix(build): avoid parallel electron-vite config race

* test(windows): avoid MSYS temp path rewriting

* test(windows): avoid killing exited PTY

* fix(pty): avoid late ConPTY input teardown race

* fix(terminal): sync reconnect error ownership after commit

* fix(runtime): use canonical worktree identity comparison

* test(ssh): assert complete cold-hydration baseline

* test(windows): invoke quoted retention fixture via PowerShell

* test(windows): read ConPTY grid through mode con

* fix(terminal): publish PTY replacements atomically

* fix(terminal): infer stale identity on reattach

* fix(terminal): fence stale pane PTY callbacks

* fix(terminal): fence stale pane binds after rebind

* fix(terminal): reject stale pane transport callbacks

* fix(terminal): fence mirrored reattach spawn callbacks

* fix(terminal): replace stale pane PTYs on remount

* fix(ci): size the Windows launcher-compile test budget from measurement

`native-smoke (windows-latest)` fails ~4.5% of runs on
`preserves a multiline argument through the compiled remote launcher`
with "Test timed out in 15000ms" — on unrelated PRs, for reasons that
have nothing to do with them. Across 176 sampled attempts it is the only
red that job produced, and it hit seven different PRs in two days:
#16900, #16904, #16915, #16955 (twice), #16979, #17014, #17085.

The test is six process creations: powershell.exe forks csc.exe, then
the freshly compiled orca.exe forks node.exe, twice. Hosted Windows
runners periodically slow process creation down, and this test amplifies
that far harder than anything else in the job. Comparing the 80 attempts
where it ran under 3s against the 12 where it ran over 12s, its own
median goes 2198ms -> 15917ms (7.2x) while the same file's
powershell-only test moves 556 -> 686ms (1.2x), the cmd.exe and Git Bash
process tests in the neighbouring file move 1.4x, and the other 35 files
put together move 1.5x.

Measured across those 176 attempts: 1881ms to 35438ms, p50 4264ms,
correlation +0.881 with the job's total Vitest duration. 8 of 176 (4.5%)
exceeded the 15s cap; 2 of 176 (1.1%) also exceeded the shared 30s
testTimeout, so deleting the override and inheriting the config is not
enough on its own. 60s clears all 176 with 1.7x headroom on the worst.

This is slow, not hung. Every body here is synchronous spawnSync, so
Vitest cannot interrupt one — the timer fires only after the body
returns and the reported duration is real elapsed time. That is why a
failure reads `× ... 22464ms` under `Test timed out in 15000ms`. The
work finished; the stopwatch was short. Seven reruns at one identical
head measured 2053 / 4680 / 5551 / 8732 / 13506 / 14868 / 21937ms — the
last of those would have been red on code that had not changed.

The 15s came from #8897, which raised this test off Vitest's built-in 5s
default because the job then ran bare `pnpm vitest run`. #8909 landed
3h27m later and pointed the job at config/vitest.config.ts, which is the
real fix for that. The constant stayed behind and has been the binding
budget ever since.

* fix(terminal): fence stale remount reattach ownership

* fix(terminal): reconcile mounted pane identity after replacement

* fix(terminal): fence stale reattach fallback ownership

* fix(terminal): fence deferred SSH reattach ownership

* fix(terminal): fence stale split pane ownership callbacks

* fix(terminal): keep stale spawns from consuming startup

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-31 08:17:40 -07:00
Neil c09810b641 perf(rpc): restore compiled Zod request schemas without override (#17374)
* perf(rpc): compile Zod request schemas lazily

* chore(deps): pin zod 4.5.4 and except it from the release-age gate

4.5.4 is the first release fixing isRecursiveSchema (upstream 84e416f, #6500),
which compile() calls on every schema — on 4.5.0 it fired .default() factories
at compile time. Verified: compile-time factory calls 0 on 4.5.4, 1 on 4.5.0.
2026-08-31 00:47:22 -07:00
Neil e84042572c Upgrade xterm to 6.1.0-beta.303 and generate addon patches
* Upgrade xterm to 6.1.0-beta.303 and generate the addon patches

Takes the current xterm beta line: xterm 287 -> 303, addon-webgl 286 -> 299,
addon-serialize 287 -> 300, headless 302, the remaining addons -> 300, and the
same set on mobile. All four packages stamp upstream commit d3e32b3.

The reasons are upstream #6042/#6043/#6055 (a shared glyph atlas no longer
garbles sibling panes on a page merge, clear, or sampler-budget overflow) and
Note that core 303 is not image-addon-only over 302: it carries the buffer perf
work, including the new BufferLineStringCache.

addon-webgl and addon-serialize move into the patch generator
--------------------------------------------------------------
Both were hand-edited minified bundles, which is what the Known Gaps section of
docs/reference/xterm-patch-regeneration.md described. Both reproduce byte for
byte from the pinned commit, so they are now manifest entries generated from a
source patch like @xterm/xterm already was. Their sourcemaps now move with their
bundles; before this they shipped maps whose offsets did not match the code
beside them.

The webgl patch shrinks from a 1.06 MB hand-edited bundle to a 6.6 KB source
patch, because upstream took the invalidation half Orca had backported. What is
left is only what upstream still lacks: the fragment-shader else branch for a
v_texpage past the sampler budget, the clearTexture guard that no-ops once a
merged page holds index 0, spending the merge retry budget before beginFrame
latches the version it saw, and Orca's font-weight probe.

The serialize source patch is byte-for-byte the same fixes as before; upstream
changed nothing in that addon between 287 and 300.

Generator fixes, each of which failed silently
----------------------------------------------
- `--relative` was appended after the `--` separator in CHECKOUT_DIFF_FLAGS, so
  git read it as a pathspec and kept repo-root-relative paths, dropping every
  source hunk from an addon's patch.
- `git apply` run from a package subdirectory still resolves patch paths from
  the repo root, skips every hunk and exits 0. It now runs from the root with
  `--directory=<packageDir>`, and a source patch that leaves the checkout
  unchanged is a hard failure rather than an empty patch.
- An addon's own `tsgo -p .` has empty files/include and only project
  references, so it emits nothing and the addon webpack then fails on a missing
  ./out/. The root build now runs first.
- versionStampFile is optional; publish.js stamps an addon's package.json, which
  overlayBuildOutput never patches.
- On a version bump the lockfile has no entry under the new key yet, so --write
  reports the gap instead of aborting mid-run. --check still fails on it.

Adding the two addons pushed the generator and the Electron packaging contract
test over max-lines, so the patch-text helpers move to xterm-patch-text.mjs
(pure text: no checkout, no build) and the vendored-xterm assertions move out of
the packaging contract into xterm-webgl-runtime-contract.test.mjs.

Tests
-----
Four tests asserted upstream bugs that are now fixed, not Orca behaviour:

- xterm-user-scrolling-contract pinned headless and core by version string.
  Upstream bumps each package only when its own output changes, so headless 302
  and core 303 are the same source. It now asserts they share a commit.
- Five CSI 3 J assertions expected a reader stranded at the top after an erase.
  Upstream #6081 clears isUserScrolling there, so the erase releases them to the
  bottom instead. Orca's pin still lands them correctly, because its parser
  handler observes the erase before xterm's own handler runs.
- The IME transaction test hard-coded the xterm version; it now reads the
  installed package, since the point is that bundle, map and version agree.
- The Electron runtime contract asserted Orca's old clearModelGeneration. Shared
  atlas invalidation is upstream's now, so it asserts pageLayoutVersion on the
  resolved dependency, plus the Orca-only hunks on the patch.

Verified: 66,008 unit tests, mobile's 3,863, the four WebGL atlas e2e specs, and
`regenerate-xterm-patches.mjs --check` in sync on all three packages.

Left alone deliberately: resetAllTerminalWebglAtlases still fans out globally
even though clearTexture now self-heals siblings, and upstream #6068
(WebglAddon.dispose leaks the GL context) is still open.

* Drop the two unused WebGL atlas fan-out exports

resetAllTerminalWebglAtlases and presentAllTerminalPanesWithoutAtlasClear have
no callers, and had none at cadfc55102 either — the last call site went in
#6949, which routed reveal recovery through
resetAndRefreshAllTerminalWebglAtlases instead. Only a comment in
pane-manager.ts still named the first one; it now points at the live entry
point. scheduleRevealPresent leaves the registry's structural type with them,
though the manager method stays: terminal-visibility-resume.ts calls it
directly.

This is dead-code removal, not a consequence of the xterm bump. The live
recovery path is unchanged.

resetAndRefreshAllTerminalWebglAtlases stays, and so does the reveal-time
escalation in pane-reveal-repaint.ts. Upstream 299 does make a pane-local
clearTexture bump pageLayoutVersion so siblings rebuild on their next frame,
which is the bug the escalation was written for, but I could not demonstrate
that removing it is safe: with the escalation removed,
floating-workspace-shared-glyph-atlas.spec.ts still passed headful, and it also
passed with upstream's mechanism deliberately disabled (pageLayoutVersion
pinned to 0 in the installed bundle, verified present in the built renderer).
A guard that passes with the fix disabled cannot license removing the
workaround, so the escalation stays until that spec can reproduce the garbling.

Verified: pane-manager and terminal-pane suites (4,713 tests), typecheck, the
headful shared-atlas spec, and the three headless WebGL specs.

* Give the shared glyph atlas spec a trigger that can fail

floating-workspace-shared-glyph-atlas.spec.ts guards the corruption where one
terminal wiping the module-global atlas leaves sibling terminals drawing from
stale texture coordinates. Both of its tests drive that through a floating
panel reveal, and Orca's reveal paths escalate to a registry-wide atlas reset
that repaints every pane — so the recovery under test heals the damage before
the assertion runs, and the tests pass whether or not xterm propagates the
invalidation at all.

The new test clears the shared atlas straight through the floating manager with
the panel closed, so nothing else repaints the workspace terminal, then repaints
it with terminal.refresh(). That is the load-bearing detail: _updateModel skips
cells whose content is unchanged, so the refresh reuses vertices baked against
the pages that were just wiped, which is exactly the state the fix has to
recover from.

Verified as a discriminator rather than assumed. Pinning ITextureAtlas's
pageLayoutVersion getter to 0 in the installed bundle, which disables the
per-renderer invalidation upstream added in addon-webgl 0.20.0-beta.299, and
confirming that reached the built renderer:

  fix intact:   siblingClearIntact=true   1 passed
  fix disabled: siblingClearIntact=false  1 failed

The failure renders the workspace terminal completely blank — stale coordinates
into a wiped atlas sample nothing. The two reveal tests pass unchanged in both
configurations, which is the gap this closes.

* Compare shared-atlas screenshots with tolerance instead of byte equality

Byte equality fails on sub-pixel antialiasing noise that leaves every glyph
legible, so the headful spec flaked under xterm 303. Reuse the existing
compareTerminalScreenshots helper: real stale-model corruption blanks the
terminal at ~3% of pixels, twice the helper's 1.5% threshold, so the looser
oracle keeps its teeth. Log the ratio so failures are diagnosable.

* fix(xterm): cancel empty deferred IME compositions

* test(xterm): strengthen runtime patch contracts
2026-08-30 15:14:49 -07:00
Neil 7f822a73e3 fix(terminal): render the IME caret and give the candidate anchor one owner (#17170)
* fix(terminal): render IME caret without placeholder overlap

* fix(terminal): preserve dim mid-line composition tails

* fix(terminal): keep IME caret visible at row edge

* fix(terminal): harden IME overlay lifecycle and layout

* test(terminal): type final-cell layout mock

* fix(terminal): keep final-cell IME anchor on-screen

* fix(terminal): bind IME masking to composer ownership

* fix(terminal): bound IME placeholder session ownership

* fix(terminal): track latest IME placeholder session

* test(terminal): share IME session event fixture

* fix(terminal): keep both writers of the IME candidate anchor in agreement

`textarea.style.left` has two writers: xterm's patched CompositionHelper and
Orca's terminal-ime-candidate-anchor.ts. The anchor module listens on
terminal.element, so within a composition event it writes after xterm's textarea
listener and reverted the final-column clamp the patch had just applied.

Moving the clamp into the anchor module and dropping the patch hunk does not fix
it, and the rendered e2e caught that: CoreBrowserTerminal.ts:444 drives
updateCompositionElements from onRender as well, so xterm re-asserts the textarea
position on every repaint, with no composition event for that module to hear. The
anchor survived only when no render happened to follow — measured as a flake at the
final column, 1561.28px against a 1557px screen edge, the fully unclamped value.

So both writers now compute the same clamp. The patch keeps it, because it is the
writer on the render path and already holds cursorLeft, maxWidth and the preedit
bounds. The anchor module applies the same one, so its composition-event write no
longer reverts the correction in the window before the next render. Both halves are
individually necessary and both are mutation-tested.

Also restores _getRowRemainderText's expression from main: translateToString(true,
x, line.length) and translateToString(false, x, getTrimmedLength()) are the same
call, since upstream does endCol = min(endCol, getTrimmedLength()) under trimRight.

Adds the two missing tests — one installing both anchor writers in a single rig, one
driving a render under an open composition — plus disposal cleanup and clamp-bound
coverage, and moves the Codex/Claude placeholder mask to a follow-up PR.
2026-08-30 02:23:04 -07:00
Neil 4bc2085271 Revert "perf(rpc): compile Zod request schemas lazily" (#17368) 2026-08-30 01:21:55 -07:00
Neil 7b86833120 perf(rpc): compile Zod request schemas lazily (#17353)
* perf(rpc): compile Zod request schemas lazily

* test: align window reveal assertion
2026-08-30 01:03:54 -07:00
Neil 4bb9dd5b89 chore(deps): bump electron 43.4.1 and other meaningful runtime deps (#17330)
Take the high-value desktop and mobile upgrades that fix crashes, jank,
or security holes. Leave Electron 44, Lucide 1, Reanimated 4.6, Expo
56/57, and xterm betas for later.

Desktop: electron 43.4.1, @tanstack/react-virtual 3.14.10, mermaid
11.17.2, ws 8.21.3, react 19.2.8, pdfjs-dist 6.3.289, vitest 4.1.11,
happy-dom 20.11.8.

Mobile: Expo SDK 55 patch train, react-native 0.83.10 (IME patch
ported), reanimated 4.3.4, webview 13.16.2 (thread-safe decision
manager; restore WebView generic default so TS 6 does not collapse
props to never).

Electron 43.4 dropped marginType from PrintToPDFMargins; CDP print
mapping now supplies the four sides only.
2026-08-29 20:44:43 -07:00
Neil 2dfaa676d8 chore: update oxlint and oxfmt (#17150) 2026-08-29 14:13:35 -07:00
Neil b17f60d744 build: upgrade to pnpm 12 (#17156) 2026-08-29 14:13:26 -07:00
Neil 0bf5361c92 perf(markdown): update code highlighting incrementally (#17147) 2026-08-29 13:52:16 -07:00
Neil eb00123a81 perf(markdown): skip unmatched list tokenizer scans (#17134) 2026-08-29 13:43:27 -07:00
Brennan Benson fd9125ea8c feat(native-chat): Codex structured native chat restructure (#16729)
* feat(native-chat): port structured Codex sessions from restructure-recovery

Rebuilds the desktop structured native-chat implementation from
brennanb2025/native-chat-restructure-recovery (tip 4e31c08db3) on top of
current main as a single commit, scoped to the local Codex path.

Ported:
- Structured agent-session core: durable record store + single-writer lease,
  canonical journal, agent-session wire host/attach/eviction/subscribers,
  `agentSession.*` RPC surface (registered via ALL_RPC_METHODS; host-side
  mobile allowlist included for wire compat), pty write gate, transcript
  additions, and the Codex app-server adapter/launch resolution.
- Renderer: NativeChatStructuredSession view/composer stack, structured
  launch path with the single-flight guard, local structured session tabs
  sync, activation gate + structured inventory (read-only
  `agentSession.handoffStatus` probe), agent-session tabs in the tab strip,
  AI-vault structured session activation, and the settings pane with the
  parent Experimental Chat UI toggle plus the nested "Use updated structured
  native chat" toggle. New sessions require both flags, agent codex, no
  prompt, and a local non-WSL, non-Windows-host execution host
  (structured-native-chat-availability).
- Fixes 72c013cea6 (verified Codex launch recovery), 8ddbaf5e3d (defer
  native terminal view switching affordances), and 4e31c08db3 (release the
  launch gate after a visibility retry) with their regression tests,
  including the third-launch-after-retry guard case.
- Cross-version agent-session wire test + CI lane, packaging entries
  (proper-lockfile, agent-tooling asar excludes), and the wire-compat doc
  section.

Deliberately not ported: mobile/ changes, the Claude structured runtime
(only the claude-transcript-branch-proof and claude-structured-owner-identity
leaf modules remain, backing the kept TUI-recovery arms), the terminal↔chat
adoption/handoff flow (`agentSession.adoptTerminal`/`requestHandoff`, the
handoff request engine, TUI adoption machinery, orca-runtime adoption
methods), renderer switching affordances and their dead leftovers, the
hook/subagent-status refactor cluster, and unrelated branch changes. The
crash-during-acquisition recovery path (restart handoff adjudication,
restore/reverse re-acquire, lease schema handoff keys) is kept because every
plain direct launch depends on it; a trimmed handoff coordinator exposes
only status/restore/close.

Branch edits that targeted files main has since split (ipc/pty.ts,
worktrees.ts, rpc/methods/terminal.ts, useIpcEvents, pty-connection,
store/slices/terminals.ts, runtime-types, web preload) were re-applied to
the split modules, preserving main's newer logic (Windows CIM fallback,
browser tab close rework, cold-restore resume flow, dispatcher threading).

Known seam: the mobile clipboard image-provenance CONSUMER gate ships
(agentSession.send refuses unproven mobile image refs with
agent_session_image_untrusted) but the producer hunk in
rpc/methods/clipboard.ts stays with the unported mobile cluster, so mobile
image sends into structured chat fail closed until that side ports.

* fix(native-chat): trust only authenticated local image uploads

* fix(build): preserve Windows process-tree patch application

* test(windows): include process creation time in addon fixture

* fix(build): run windows-process-tree node-gyp from the physical package dir

gyp expands the node-addon-api dependency by probing node, whose cwd
resolves to the package's physical directory in the store, so the emitted
target is a store-relative ../../../../node-addon-api@... hop. gyp then
resolves that hop against the rebuild cwd; from the node_modules
symlink/junction it escapes the store and configure fails with
"node_addon_api.gyp not found" (run 32999886072).

Rebuild from realpath(package dir) so both bases agree, matching how the
package manager itself runs native install scripts. The regression test
replays gyp's expansion+resolution against the planned cwd and fails
without the fix.

* fix(native-chat): keep chat tabs visible through terminal closes and empty-worktree launches

Two proven blockers in the native Codex tab contract:

closeTerminalTab pre-empted the canonical unified close. With one terminal
left it deactivated the worktree on a terminal/editor/browser-only check,
blanking a workspace that still held a renderable agent-session tab; with
two or more it pre-picked a successor from terminal entities only,
re-stamping the group active before closeUnifiedTab's MRU/neighbor repair
could land on the chat tab. Successor choice now defers to the unified
contract whenever the terminal has a unified row, and deactivation is
gated on the unified renderable count (matching leaveWorktreeIfEmpty),
with the legacy pre-pick kept only for terminals without a unified row.

A structured session created on an empty worktree was published into the
host's headless group while preserveLocalLayout froze the local layout,
leaving the tab in store but permanently off screen. A preserveLocalLayout
owner now always takes client-owned placement — repairing a rendered
leaf whose group record is missing, or materializing a rendered group on a
truly empty worktree — and applies the client-derived layout repair while
still rejecting host-authored layout.

Regression tests drive the real store through closeTerminalTab (git
worktree and folder workspace) and the real snapshot applier for the
empty-worktree adoption states; all fail without the fixes.

* fix(native-chat): close stale turns and retry rejected sends

* fix(native-chat): retire hosted rows on structured tab activation

* fix(native-chat): preserve rpc defaults across main merge

* chore: format remote wire compatibility guide

* test(native-chat): cover retry after unconfirmed send

* fix(native-chat): reload outbox on session switch

* docs(settings): disclose structured chat platform limits

* fix(native-chat): await Codex launch-home preparation

* fix(codex): align child-process allowlist with async trust bridge

* test(identity): update inventory for tab surface refactor

* fix(windows): preserve process-tree CRLF patch sources

* fix(native-chat): anchor an unmatched chat echo where it was sent (#16117)

* fix(native-chat): anchor an unmatched chat echo where it was sent

The reported symptom was old user messages replaying below every new turn, so the
conversation read as scrambled. The cause was not that the echo failed to match a
transcript row. Claude consumes a mid-turn send through a `queued_command`
attachment and writes no `type:"user"` record for it, so some echoes can never
match, and no amount of matching will change that. The cause was WHERE an
unmatched echo rendered: buildMobileNativeChatTransientData appended every pending
item after the entire transcript, so it re-read below each turn that landed
afterwards.

Render each echo directly after the transcript row it was sent against, using the
baseline the send already captures. An unmatched echo is then at worst a duplicate
in the right position rather than a scrambled one, and it stays visible. Echoes
sharing an anchor keep send order; a send with no baseline, or one whose anchor
folding dropped, still falls back to the tail.

Deliberately NOT fixed by deleting the echo. Inferring from send ordering that an
echo can never match, then removing it, loses the user's own text for a message
the agent did receive, and it cannot fire in the common case anyway - measured
drain groups are 1,017 of size 1 against 55 larger. It also escalates an existing
gap: the count pass has no baseline-tail guard, unlike the glue pass, while
`messages` is a 40-row window that head-trims, resets on reconnect and grows at
the front on loadEarlier, so a false landing there would license deleting a
DIFFERENT outstanding message.

That count-pass gap is real and left for a separate change; anchoring makes its
worst case a duplicate in place rather than a scrambled conversation.

* fix(native-chat): preserve folded echo anchors

* fix(native-chat): preserve forward-folded echo anchors

* fix(native-chat): keep leading folded echoes in place

* fix(workspace-cleanup): show git status for every row (#16690)

* fix(native-chat): refuse structured chat on every Windows execution path

canUseStructuredNativeChat only refused win32 when a project runtime
resolved, so folder-workspace keys (and other keys with no project
runtime) failed open into structured chat on Windows. Fail closed on
win32 unconditionally after the host check, matching the settings copy:
local macOS/Linux only; Windows/WSL/SSH stay on terminal chat.

* fix(native-chat): restore runtime refusals behind the win32 gate

506d375de3 replaced the project-runtime checks with a bare platform test,
so a WSL or repair-required runtime resolution would no longer refuse
structured chat off-win32. Keep the unconditional win32 refusal and
re-run the runtime resolution after it, so the gate does not depend on
the resolver's own platform guard. Tests inject WSL and repair-required
resolutions on darwin/linux and fail against the regressed gate.

* fix structured session journal durability

* fix structured tab active pointer after restart

* fix(native-chat): await optional lease renewal callbacks

* refactor(skills): extract install error messages

* fix(agent-session): harden recovery ownership

* fix(native-chat): retain panes across tab activation

* fix(native-chat): address round-one review findings

* test(native-chat): align integration coverage after main merge

* fix(native-chat): harden round-two reliability

* fix(native-chat): harden round-three reliability

* fix(native-chat): close round-four recovery gaps

* fix(native-chat): separate bounded journal key forms

* fix(native-chat): reset outbox error in render on session switch

The switch effect adjusted error state after the sessionId prop changed,
tripping react-doctor's no-adjust-state-on-prop-change on the changed-code
gate and flashing the old session's banner for a frame. Reset it with the
render-time previous-value guard instead.

* fix(native-chat): invalidate stale outbox settlements

* test(native-chat): restore settled-error session-switch regression

a6e2379bd1 replaced this test with the in-flight settlement race test,
leaving the render-time error reset unpinned: deleting the reset block
still passed the whole native-chat suite. Keep both scenarios pinned;
they are distinct (settled error clears on switch vs stale settlement
invalidated in the commit-to-passive window).

* test(wire): make release checkouts race safe

* test(wire): pin cross-process checkout single-flight and importer specifier contract

* test(wire): harden release checkout lifecycle

* fix(build): drop CR-byte residue from windows-process-tree patch

The two trailing CR bytes on the patch's deletion lines are a proven
no-op: pnpm hashes patches CRLF-normalized (both forms hash to the
lockfile's 946ffb2b) and materializes this package without applying the
patch in either form, so the load-bearing build edits come solely from
applyWindowsProcessTreeBuildFixes() (#16947), which handles both source
EOL forms. Restore byte-identity with main and repin the contract test
to the post-#16947 reality: LF-only patch bytes plus lockfile hash sync.

* fix(native-chat): skip empty startup recovery
2026-08-28 16:45:58 -07:00
Brennan Benson 4bb337741c feat(terminal): weight-layer forensics for the bold-collapse bug (STA-4042) (#16868)
* feat(terminal): weight-layer forensics for the bold-collapse bug (STA-4042)

Field instrumentation to name the writer behind regular-text-renders-bold:
- metric-weight-change crumbs at the writePaneMetricOptions funnel
  (prev/next/reason; weights never change in normal operation)
- terminal-weight-parity-mismatch audit on every visibility resume
- sentinel weightProbe capture fields: live options vs atlas captured
  config vs renderer-buffer bold census
- Cmd/Ctrl+Shift+click unconditional capture (no divergence gate, no
  recovery) for states the missing-ink detector cannot see
- patched addon-webgl ctx.font readback probe: detects failed font
  assignments that rasterize glyphs at a stale weight

* fix(terminal): treat canvas weight-700-serializes-as-bold as a match in the atlas font probe

Found by live validation: Chromium's ctx.font getter normalizes numeric 700
to the keyword 'bold', which made every legitimate bold rasterization count
as a failed assignment (124 false positives in one session).

* chore: update patch hash for the font-probe normalization fix

* fix(terminal): bound bold glitch diagnostics

* fix(terminal): cover serialized WebGL probe state

* feat(settings): hidden staff toggle to arm terminal render diagnostics

Replaces the reserved hidden-experimental placeholder slot with a real
switch (Shift-click the Experimental sidebar entry to reveal). It arms
and disarms the render-desync capture sentinel live — no localStorage
incantation, no reload — for the bold-glitch investigation. The passive
probes stay always-on; only the capture gestures are gated.

* fix(settings): make render diagnostics disarm exact

* chore(settings): rename hidden group to 'Hidden experimental settings', drop its description

* feat(settings): unlock hidden experimental group via Option-click on the Experimental page title

Replaces the Shift-click-sidebar unlock with the Updates-header idiom:
Option-click the Experimental page title toggles the hidden group.
Removes the now-unused click-modifier plumbing from the settings sidebar.
2026-08-27 23:36:44 -07:00
Jinwoo Hong 8dd7d6060c fix(release): stabilize native builds across CI platforms (#16947) 2026-08-27 21:42:34 -07:00
Neil 0096e47850 fix(windows): keep windows-process-tree gyp paths absolute under pnpm (#16688)
* fix(windows): keep windows-process-tree gyp paths absolute under pnpm

Hourly Windows builds have failed since #16598 at
`build-windows-process-tree-relay-addon`: `require('node-addon-api').targets`
is cwd-relative, so node-gyp evaluates it from the pnpm store realpath and
then loads it from the `node_modules` symlink. That resolves
`node_addon_api.gyp` outside the repo.

Use `require.resolve` for an absolute path, matching the node-pty patch.

* i18n: keep ja skill-filter labels on the catalog's Agent brand

#16682 merged with a failing localization catalog: ja used エージェント
in three new skill-filter strings, and repair-locale-catalog rewrites
those to Agent. Match the rest of ja.json so static analysis can pass.
2026-08-26 16:15:03 -07:00
Neil 2d500278b4 build(windows): refuse unpatched node-pty prebuilds
Merged after clean CI, Windows packaging verification, and readiness review.
2026-08-24 15:48:47 -07:00
Neil 2b1254d681 fix(windows): own PTY process trees with job objects (#15755)
* fix(windows): own PTY process trees with job objects

Teardown used to answer 'is this tree mine, and how do I kill it?' by
scraping the process table, walking parent pids back to Orca, and running
taskkill /T /F only if the walk said yes. Every step is a guess, and the
code said so itself: windows-pty-root-identity.ts:35 already named the
fix -- 'an inherited handle / Job Object'.

The guesses fail in the ways users report. A pid walk cannot survive pid
reuse, so teardown refused whenever it could not prove ownership, and a
refused kill is an orphaned agent tree holding the worktree directory
open (#9045, #10475, #10087). A descendant that reparented is invisible
to the walk. The scrape itself could be blocked by policy, which read as
'no evidence'.

node-pty now creates a job object per ConPTY and assigns the shell under
CREATE_SUSPENDED, before it can spawn anything -- assigning afterwards
leaves a window in which a fast child escapes. Termination is one
TerminateJobObject; liveness is QueryInformationJobObject.

Verified on Windows 11 against a shell whose grandchild was spawned
detached: job membership came back [shell, grandchild] and one call
killed both. Neither a parent-pid walk nor GetConsoleProcessList sees
that grandchild -- it leaves the console and reparents, which is exactly
the claude.exe/node.exe/cmd.exe orphan in #9045.

KILL_ON_JOB_CLOSE means a daemon that dies without unwinding no longer
strands shells (#9195, #10415). The job is the daemon's, not the app's,
so an app-main crash still leaves sessions alive -- the guarantee
win-crash-survival-e2e asserts.

Both entry points report unavailable rather than a false success when a
pty has no job: an outer job without BREAKAWAY_OK can refuse the
assignment, and a pty from an older build has none. Reading 'we could
not tell' as 'already dead' is the original bug, so the old probe stays
as the fallback.

* test(windows): pin job ownership against a real detached grandchild

The unit tests pin the contract; this pins what the contract is for. A
grandchild spawned detached leaves the pane's console and reparents, so
GetConsoleProcessList and a parent-pid walk both miss it -- that is the
process that outlived its pane and held the worktree directory open.

Includes a guard that this build actually has job support, so a node-pty
rebuilt from unpatched sources fails loudly instead of letting every
assertion pass vacuously.

* fix(windows): correct the job liveness contract to what Windows actually does

I claimed an emptied tree would report [] and that this was the evidence
a stale registry entry lacks (#15549). Running it on Windows 11 showed
otherwise: node-pty drops its handle record and closes the job when the
shell exits, so a dead tree reports null.

Null therefore means unverifiable in the sense of
docs/reference/ssh-execution-boundary.md -- no job support, not a ConPTY,
or no longer tracked -- and is never evidence that processes died. A
caller reading it as proof of death would have been right by accident
after a normal exit and wrong on a host that refused the assignment.

What the API does add is descendant liveness for a tree that is still
tracked, including children that detached from the console.

* fix(windows): stop a clean shell exit from reaping backgrounded processes

Measured on Windows 11: with JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE on the
per-PTY job, releasing the handle when the shell exits also killed
whatever the user had backgrounded. Typing 'exit' in a pane reaped a
detached server that survived before this patch.

That is a behaviour change nobody asked for. The approved change was
that killing the terminal daemon reaps its shells -- not that a clean
exit reaps your background job. The job's purpose is to make an EXPLICIT
teardown exact, which TerminateJobObject still does.

Reaping a dead daemon's shells now needs the daemon-level job the design
called for: the daemon assigns itself, children inherit membership, and
its closure on daemon death reaps them without touching clean-exit
semantics. Not in this PR; noted in the reference doc.

* test(windows): pin that a clean exit leaves backgrounded work alone

The counterpart to the tree-kill test. Without it, re-adding
JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE would look like a tightening rather
than the regression it is.

* fix(windows): stop a winpty pty id from matching a ConPTY job

winpty.cc and conpty.cc each mint their 'pty' id from an independent
counter, and windowsPtyAgent stores both in the same _pty field. So a
winpty-backed terminal's id can collide with a live ConPTY baton -- and
closing that pane would have terminated an unrelated pane's entire
process tree.

Both job entry points now take the shell pid and the native side refuses
unless GetProcessId(hShell) matches, which makes the id unforgeable.

Two more from the same read-through:
- ResumeThread's failure was ignored. A shell left suspended is a pane
  that never prints and never exits, which is far harder to diagnose
  than a failed spawn; it now cleans up and throws.
- handle->hJob was assigned before LoadConptyDll, which can throw. A
  baton carrying a job but never reaching SetupExitCallback has nothing
  left to close it, so the assignment moved down beside hShell.

* docs(windows): record the unsynchronised node-pty baton table

Pre-existing upstream -- the exit thread erases while the main thread
reads -- but terminatePtyJob adds an instance of it, so it belongs in
writing rather than in someone's head.

* fix(windows): close four gaps found in review

BREAKAWAY. The per-PTY job set no limits, so a child asking for
CREATE_BREAKAWAY_FROM_JOB was refused with ERROR_ACCESS_DENIED.
Installers, msiexec and some updater and service-control paths spawn
that way deliberately -- they worked before this patch and would have
failed only inside an Orca terminal, which is the worst shape a bug
report can take. JOB_OBJECT_LIMIT_BREAKAWAY_OK restores it; a child
still has to ask, so ordinary descendants stay owned.

EMPTY IS NOT UNAVAILABLE. The native reader returns an empty list --
not an error -- when CreateToolhelp32Snapshot fails, which is what an
EDR hook or a restricted token produces. Callers read that as 'nothing
is running' and teardown concludes a live PTY root is already gone. The
snapshot must contain the querying process; nothing else is
unfalsifiable, and one predicate catches empty, truncated and
permission-filtered tables alike.

NO DEADLINE. Replacing execFile dropped its 3s timeout. The vendored
reader latches a module-global while a request is in flight and clears
it only after draining its callbacks, with no try/catch -- so one wedge
leaves every later call queued behind a promise that never settles, and
the process table is dead for the life of the app. The bound is back.

GUESSED IMAGE PATH. executablePath was derived from the first
space-delimited token, which reads 'C:\Program' out of an unquoted
'C:\Program Files\nodejs\node.exe ...'. Wrong evidence is worse than
none, and the only consumer already had the full path in , so
the field is gone rather than repaired.

Also: remove_pty_baton no longer sits inside assert(), which NDEBUG
would compile away along with the call, and the job accessors hold a
lock across lookup and use -- handle values are recycled, so an
unguarded read could pass the shell-pid check against an unrelated
process and terminate the wrong job.

* fix(windows): apply the job lock once per accessor

The patch script matched a string its own replacement still contained, so
PtyTerminateJob got two lock_guards named guard and PtyListJobProcessIds
got none. MSVC caught it: error C2374 redefinition.

* test(windows): pin that a child can still break away from the job

Verified on Windows 11: 'start /b' writes its marker and no access-denied
appears. Without JOB_OBJECT_LIMIT_BREAKAWAY_OK this fails, and it fails
only inside an Orca terminal -- so the failure would look like Orca
corrupting unrelated software rather than like a job-object change.

* fix(windows): stop the ownership guard from reading a closing handle

The guard called GetProcessId(hShell) to prove identity, but the exit
watcher closes hShell on another thread -- so the guard could read a
closed handle, and under strict handle checks that is fatal rather than
merely wrong. Worse, it widened the gap between validating hJob and
using it from two instructions to a kernel round-trip, and handle values
recycle: the likeliest occupant of a freshly recycled value in this
process is another pane's job.

The pid never needed a handle. It is captured at spawn and compared as a
DWORD, so the guard touches no handle at all, and hShell is now closed
inside the same lock as hJob.

Also from review:
- reject CR/LF in a cmd argument. cmd ends the command at a raw line
  break whatever the quote state, so there is no escape for it; encoding
  one anyway truncates the argument and can leave the remainder to run
  as a command. Agent prompts are this encoder's motivating input.
- ask the process table only for the fields a caller needs. Memory and
  CommandLine each cost an OpenProcess per process, inline, for every
  process on the box -- and the 1024 bound is patched out. Ancestry
  reads now skip both.
- corpus gains the degenerate quote-only and two-quote arguments.
- PtyListJobProcessIds' docblock still taught the empty-list contract
  that was corrected on the TS side, and now records that the ConPTY
  console host is never a job member.
- drop a write to NumberOfAssignedProcesses, which is output-only.
- pty_baton::hShell is initialised; ownsShell was only safe because &&
  short-circuited ahead of it.

The backgrounded-child test is rescoped: 'start /b' uses
CREATE_NEW_CONSOLE, not CREATE_BREAKAWAY_FROM_JOB, so it proves job
membership does not block backgrounding -- not that BREAKAWAY_OK works.
That flag rests on the Win32 contract, and I have said so rather than
letting the test imply coverage it does not have.

* fix(windows): bound retries after the process table wedges

The 3s deadline stops a caller hanging, but the timed-out call leaves its
callback in the vendored module's queue -- and that queue drains only
when the latched request completes, which in this wedge never happens.
Retrying at the caller's poll rate would add a closure per tick forever.
A 30s cooldown bounds it to one probe, and a late callback clears the
cooldown because it proves the reader recovered.

Also pins the deadlock invariant in the patch: the exit thread's lock
must close before tsfn.BlockingCall, because that waits on the JS thread
and the JS thread can be waiting on the same mutex inside
PtyTerminateJob. Correct today by scoping; a comment so a later refactor
does not widen it.

* revert(windows): drop the field-selection API, which cannot pay off

I added it for a real perf finding -- Memory and CommandLine each cost an
OpenProcess per process -- and then never wired a caller, so the claim
that ancestry reads skip them was wrong.

Wiring it would have been worse than leaving it dead. The only ancestry
consumer is the teardown identity probe, which needs a snapshot that
started AFTER it asked, for pid-recycle detection. Bypassing the shared
reader to get narrow fields would let that request join a scan already in
flight -- trading a correctness guarantee for milliseconds.

Field selection only pays off if callers can ask for less, and they
cannot: one shared snapshot serves every caller so a 32-wide teardown
collapses into a single scan, which means it has to carry every field.
The reasoning now lives next to the flags instead of in a dead export.

* fix(process): three P1s from review — a crash vector and two wedge bugs

STDIN EPIPE COULD TAKE DOWN THE MAIN PROCESS. A child that exits without
reading makes the queued write fail with EPIPE, and an unhandled error on
a stream is an uncaught exception. The child's own error listener does
not cover its stdin stream, so runProcess({ input }) against a
short-lived child was a crash, not a failed call.

THE COOLDOWN LEAKED A BATCH PER CYCLE INSTEAD OF BOUNDING IT. At expiry
every concurrent caller passed the check before any of them re-armed it,
so each enqueued a callback into the still-latched native queue and each
cycle leaked another batch. The cooldown is now re-armed BEFORE probing,
so exactly one caller gets through.

A SYNCHRONOUS THROW LEFT ITS DEADLINE RUNNING. The timer was declared
inside the try, so catch could not clear it; it fired later and wedged a
reader that had already recovered. Hoisted and cleared, and wedge state
now carries a generation so a request that lost its deadline cannot
mutate it on behalf of the one that replaced it.

Found by review once the prompts were short enough for the reviewer to
finish -- the previous two rounds died on prompt length.

* fix(process): stop a stream error from crashing the main process

Same class as the stdin EPIPE finding, two instances further on: stdout
and stderr had data listeners and no error listeners, and an unhandled
error on a stream is an uncaught exception.

Scoped to runProcess, which owns the child outright. spawnProcess hands
the streams to its caller, and a blanket handler there defeats callers
that track and remove their own listeners -- the SSH ProxyCommand
transport does exactly that, and its cleanup test caught the attempt.
Documented on spawnProcess so the boundary is explicit rather than
inferred.

* fix(windows): validate the ConPTY DLL before creating the process

LoadConptyDll throws when conpty.dll is missing -- a real state, and one
this branch hit during development. It ran after CreateProcessW and
ResumeThread but before the baton and the exit watcher were installed,
so a throw leaked the job, process and thread handles and left an
untracked shell tree running. Once per attempt, so a broken install
accumulates orphan shells on every retry.

Resolving the DLL first costs nothing and leaves exactly two throws
after creation: the CreateProcessW failure, where nothing exists yet,
and the resume failure, which already cleans up after itself.

This also closes the same leak for hProcess and hThread, which predates
the job work.

* feat(windows): add the daemon-level job the design called for

The plan specified two nested jobs and I built one. That gap is why
dropping KILL_ON_JOB_CLOSE from the per-PTY job cost the approved
guarantee that a dead daemon reaps its shells -- I had one job trying to
answer two questions, and the two answers conflict.

They are separate jobs. The per-PTY job answers 'kill exactly this
pane's tree, now', and cannot be kill-on-close because its handle is
released when the shell exits, which would reap whatever the user
backgrounded. The daemon assigns itself to a second job that IS
kill-on-close; its handle is released only when the daemon dies.
Children inherit membership, so every pty is covered and the per-PTY
jobs nest inside it.

Daemon, never app: an app-main crash must still leave sessions alive,
which win-crash-survival-e2e asserts. Both jobs carry BREAKAWAY_OK, or a
child asking to break away is refused at whichever level lacks it.

Restores #9195 and #10415, which I withdrew from this PR earlier.

* docs(windows): record what the host job does not cover

An app-hosted PTY gets a per-PTY job but no crash reaping, because the
alternative is a kill-on-close job on the app -- which is precisely what
the crash-survival guarantee forbids.

* ci(windows): run the win32 suites in the PR windows job

Both were skip-on-non-win32 and had only ever run on one machine I drive
by hand -- which went unreachable at exactly the moment I needed to
verify the percent-escaping fix. Verification that depends on one box is
not verification.

The job already builds node-pty from patched source and already runs a
useConptyDll test, so the ConPTY runtime files are in place by this
step. This also makes the encoder a gate: the corpus is the only thing
standing between an agent prompt and a mangled argv, and it now runs
against real cmd.exe on every PR.

* fix(deps): refresh the lockfile for the current patch hashes

pnpm records a hash per patched dependency, and I regenerated both
patches repeatedly across the review rounds without refreshing the
lockfile. Every local run used --frozen-lockfile's looser sibling, so
nothing caught it until CI did:

  ERR_PNPM_LOCKFILE_CONFIG_MISMATCH  Cannot proceed with the frozen
  installation. The current "patchedDependencies" configuration doesn't
  match the value found in the lockfile

Verified with pnpm install --frozen-lockfile locally this time.

* ci(windows): build node-pty from source before the win32 suites

CI proved the encoder fix on real cmd.exe -- 26/26 -- and in the same run
proved the job suite had been testing an unpatched binary. node-pty
prefers its upstream prebuild, which does not contain this patch, so
every job-object export was absent and isPtyJobOwnershipAvailable() was
false.

That guard is why the failure was loud rather than a vacuous pass, and
it is the reason the assertion exists.

Packaging was never affected: rebuild-native-deps.mjs already builds
node-pty from source for Electron and restores the ConPTY runtime files.
The gap was the node-runtime test environment only.

Not changing requiresPatchedNodePtySourceBuild's win32 exemption here.
Its premise -- that the patch is Unix-only -- is now false, but lifting
it also needs pnpm rebuild to force a source build, and I cannot
validate that on macOS and Linux from here. Recorded as a follow-up
instead of changed blind.

* test(windows): gate the host-job guarantee in CI

The daemon-level job had one hand-run proof and no automated coverage --
the same shape of gap that let an unpatched node-pty go unnoticed until
CI caught it.

It needs a real second process, because the assertion is about what
happens when that process is force-killed: a host in a kill-on-close job
must strand neither its pty nor a grandchild spawned detached, which is
the process a parent-pid walk cannot see.

Runs in the Windows PR job alongside the per-pty and encoder suites, so
both halves of the two-job design are now gated rather than asserted.

* fix(windows): serialise host-job creation

Two callers racing PtyAssignCurrentProcessToJob would each create a job,
put the process in both, and leak the first handle -- and the handle is
what keeps a kill-on-close job alive, so a leaked one is never released.
'Only JS calls it' is not a guarantee: a worker thread with its own
N-API env shares these statics.

Also records the ordering requirement it depends on.
AssignProcessToJobObject adds only the named process; children inherit
membership, but a pty that already exists does not join retroactively
and would not be reaped. The daemon assigns at startup, before the
ConPTY warmup and before any session, which is correct today and now
stated rather than implied.

* fix(daemon): keep the host job off the startup path

Assigning the host job at daemon startup resolves the node-pty native
module, which loads the ConPTY addon -- and paying that before the
endpoint is published delayed readiness enough that daemon-boot-smoke
failed on windows-latest, deterministically.

windows-conpty-warmup already carries the comment for this exact
hazard ('setImmediate keeps the ready/handshake path ahead of the
warm-up') and I put an eager load in front of it anyway.

Moved to the pty spawn path, which already pays ConPTY cost, and
memoised. Children inherit job membership, so assigning immediately
before the first spawn still covers every pty -- and nothing can spawn
one before the endpoint exists.
2026-08-21 22:31:36 -07:00
Neil 057fbfcffc perf(windows): read the process table natively instead of forking PowerShell (#15749)
* perf(windows): read the process table natively instead of forking PowerShell

Seven independent readers each forked powershell.exe to run
Get-CimInstance Win32_Process, with a wmic fallback that Windows 11 24H2
has removed. On a domain-joined host with PowerShell Transcription
enabled by policy, one of them running every ~2s recorded ~289GB across
1.4 million files (#15209). The same scan cost ~700ms and ran per pane
(#15036), and a Group Policy or AV block turned it into 'unavailable',
which callers read as 'no evidence' -- which is how a PTY tree survives
its own teardown (#9045, #10475).

A Toolhelp32 snapshot answers the same question with no child process.
Measured on Windows 11 with 1050 processes, p50/p95:

  pid+ppid+name          15.9 / 17.5 ms
  +memory +command line  30.6 / 33.7 ms
  Get-CimInstance         706 / 723  ms

Two upstream defects needed patching, both found by running it on real
hardware. The binding requires Spectre-mitigated libraries our agents do
not carry (node-pty is patched the same way). And enumeration stopped
after 1024 processes: on a host with 1051 the module returned exactly
1024, and the querying process was itself among the 27 missing -- a
truncated snapshot silently hides the descendants teardown is looking
for, which is the failure this whole change exists to remove.

Migrated: the foreground/descendant reader (the #15209 scraper and the
teardown identity gate) and the port scanner's PID attribution. NOT
migrated: the memory collector and three identity probes, which need
Win32_Process.CreationDate and have no native equivalent. Start time is
a proxy for identity anyway; an inherited job handle is the real answer,
so those belong with the job-object work rather than here.

Packaging follows the windows-native-registry contract exactly:
optional, absent from onlyBuiltDependencies so macOS/Linux never run
node-gyp, win32-only in the packaged runtime. Asserted by the existing
contract test, which also stops pinning a whole source literal that only
tested its own formatting.

* chore(process): ratchet the child_process allowlist down

windows-foreground-process-rows.ts no longer spawns anything, so its
allowlist line is stale. The guard fails on a stale entry as well as a
new one, precisely so a migrated file cannot keep a slot open and hide
the next regression in the same path.

* fix(ports): import the process-table reader the scanner uses

Missing import: the migration replaced the PowerShell call but the new
symbol was never imported, so tsc failed. Vitest transpiles without
typechecking, which is why the port-scanner suite stayed green.

* fix(deps): sync this branch's lockfile with its patch set

Same class as the fix on the tip branch: pnpm records a hash per patched
dependency, and this branch introduces the windows-process-tree patch
without its lockfile entry matching. Every job here failed at install
with ERR_PNPM_LOCKFILE_CONFIG_MISMATCH.

Verified with --frozen-lockfile, which is what CI runs and what my local
runs were not.

* test(relay): drive the relay's Windows fixtures from the native snapshot

Two relay cases fed a PowerShell CIM payload through a mocked execFile.
That reader is gone, so both failed -- deterministically, on every PR
run for this branch and the one above it.

I did not catch it because my own verification sweep was
'src/main src/shared config/scripts' and never included src/relay. The
relay is a first-class consumer of the process table; leaving it out of
the sweep is how a deterministic failure survived six review rounds.
2026-08-21 21:54:57 -07:00
Neil fdd4091ebd fix(hooks): isolate lint-staged backups per worktree (#15388) 2026-08-19 22:37:02 -07:00
Neil d143922561 fix(terminal): deliver an IME commit the deferred textarea diff missed (#15198)
Picking a single Chinese character from the candidate window with a
number key loses it. The character flashes and disappears. Picking the
same candidate with the mouse works, and picking multi-character words
with number keys works.

Two paths can deliver an IME commit, and this falls between them. A
keydown the input method consumed routes into a setTimeout(0) diff of
the helper textarea, and that diff is what normally delivers the commit;
xterm's _keyDownSeen guard exists to defer to it. When the commit
arrives after that timer has already run, neither path delivers. Mouse
selection works because no key is down, and a real composition session
works because it takes a different path entirely. That narrows it to an
input method whose commit round-trips asynchronously and which shows no
in-application preedit.

Track that a consumed keydown still owes its commit, and deliver only
when the diff did not. The upstream guard and its single read site are
untouched, which is what keeps the duplicate-commit behaviour it was
added for sealed.

Not doing the obvious repairs deliberately: clearing the flag, skipping
it for keyCode 229, or setting it after the composition short-circuit
each unblock the input path without retiring the diff, and all three
were measured emitting the character twice.

The patch and the lockfile hash here are generated. Review
config/patches/xterm-src/@xterm__xterm@6.1.0-beta.287.src.patch, which
is the hand-written source of the change; the shipped patch and both
minified bundles are the regenerator's output from the pinned upstream
build, so nothing in this change was hand-transcribed into a bundle.

Refs xtermjs/xterm.js#6036
Closes #12099
2026-08-17 22:13:32 -07:00
Neil 49752477a6 build(xterm): restore the patch regeneration harness and gate it in CI (#15223)
* build(xterm): restore the patch regeneration harness and gate it in CI

docs/reference/ime-architecture.md says "Never hand-edit the bundles in
the patch" and links to docs/reference/xterm-patch-regeneration.md. That
doc does not exist, and neither does the harness it describes.

Both landed in 29117bf776 and were deleted by 17cfc968cf, a revert of
the composition-ownership change, which swept up a build tool and a CI
gate as collateral. The rule survived; its enforcement did not. Every
xterm patch since has had to hand-edit minified bundles to comply with
the surrounding architecture, because everything resolves to
lib/xterm.mjs at runtime and under vitest, so a src-only edit is inert.

The shipped bundles were therefore not the output of any build, and this
restores them to build output. Comparing identifier multisets against a
pristine build of the pinned commit finds hand-written names a minifier
never emits ($rl, $hp, $tid), const in an otherwise let-only esbuild
bundle, !! where the source reads Boolean(), an escaped LRM where esbuild
emits the literal, and a return block esbuild collapses to void(...).
Every remaining token difference is a minifier local reallocating.

The old source patch could not be reused. It described the reverted
composition-ownership architecture, so restoring it would have re-applied
an abandoned design on top of dropping three accumulated fixes. It is
re-derived from the shipped patch instead, and the derivation is a fixed
point.

Two deliberate departures from the deleted version. Sourcemaps are
included rather than deleted, because a live test reads lib/*.map and
asserts the mapped version matches the runtime version. The source-patch
superset carve-out is gone, so a source hunk the shipped patch cannot
name now fails loudly instead of being carved out silently.

The doc's claim that the webgl and serialize addons reproduce byte for
byte was half wrong. Their ESM output does reproduce at the pinned
commit, but both also publish CJS that the root package script never
builds, so folding either in needs a build step this harness lacks.
Recorded as a blocker rather than a confident sentence.

xterm_patch_sync runs the regenerator in --check mode, so a patch that
does not match a rebuild of the pinned upstream now fails PR CI.

The -diff -text attribute is required, not cosmetic: pnpm hashes the
patch byte-for-byte, so a CRLF checkout breaks the install outright.

Not verified: the CI job has not run on a real runner, the addon CJS
bundles are unreproduced, and the generator is untested on Windows and
Linux.

* build(xterm): make the regenerator runnable on Windows and drop dead paths

Readiness review on the restore found one blocking gap and two cheap
cleanups. None of them change the emitted patch, which is byte-identical
before and after.

The generator could not run on Windows at all. Three sites called npm
through execFileSync with shell:false, but npm ships as npm.cmd there,
execFile applies no PATHEXT, and since CVE-2024-27980 it refuses a .cmd
target without a shell. That matters because this harness arms a
blocking gate whose documented remedy is --write, so a Windows
contributor who tripped the gate had no remedy except hand-editing a 7MB
minified bundle, which is the practice the gate exists to abolish. Four
sibling scripts in config/scripts already handle this; the fix follows
them and lands in run(), so the manifest-driven build step is covered
too. git and tar are real executables in System32 and keep resolving
without a shell, which avoids quoting exposure on paths with spaces.

deleteGeneratedSourcemaps was unreachable, since the policy is include.
Deleting it left "delete" as a legal policy value that nothing honoured,
so a manifest asking for it would have silently shipped sourcemaps that
do not match the bundle. The enum is narrowed and an unrecognised policy
now throws rather than falling through.

generatedHunks moved into the test file rather than being dropped; its
partition assertion, that generated and source hunks reconstruct the
whole patch, is worth keeping.

The -text attribute now covers all five patch files. pnpm hashes each of
them byte-for-byte, so the CRLF hazard the xterm patch was protected
from applies equally to node-pty and the three addons. All five were
already LF in the object DB, so this pins existing behaviour. -diff
stays scoped to the xterm patch, since the others are readable.

The doc's claim that the addons reproduce byte for byte is now dated and
marked a one-off measurement rather than an invariant, because nothing
re-runs it.

Effective lines fall from 591 to 568 against the 600 budget. Still the
largest file in config/scripts, and adding a second package to the
manifest would need a split first.
2026-08-17 21:31:20 -07:00
Neilandrayim 453237cc57 fix(terminal): render the row tail the IME preedit overlay covers (#15014)
* fix(terminal): render the covered row tail inside the IME preedit overlay

Closes #12545.

Composing mid-line hid the character at the cursor for the whole composition.
The preedit overlay is an opaque box anchored to the cursor cell, and nothing
reaches the pty while composing, so those cells still held their characters —
the box simply covered them.

`CompositionHelper` now draws the rest of the row after the preedit inside the
view, so the composition reads as inserted text pushing the tail right. Four
details come with it:

- The view is start-anchored while it carries a tail, so the preedit stays put
  and the pushed tail clips at the right edge; alone, `rtl` still keeps a long
  preedit's end in view.
- It is themed from `options.theme` instead of the stock `#000`/`#FFF`, with any
  alpha dropped — the view masks the cells it draws over, so a see-through
  background would re-expose the very characters the tail stands in for.
- The helper textarea syncs to the preedit's own bounds, so IME candidate
  dialogs anchor to the composing text rather than past the rendered tail.
- A TUI can repaint the row under an open composition, so
  `updateCompositionElements` — which already runs on every render — re-reads
  the remainder and re-renders on change. A string compare adds no layout read.

The tail is read with an explicit end column: the cacheable form of
`translateToString` arms the line string cache's self-renewing idle-clear timer,
and the composition path must own no timers.

Geometry is not the cause. Two mature reference terminal implementations compose
marked text into the grid rather than into a floating box, and both still blank
the cells under it — one of them literally substitutes the marked characters
into the row's character array before rasterizing. Moving off the overlay would
not have fixed this report; rendering the covered tail is what does.

The e2e arm asserts the invariant an opaque overlay owes the grid: it must
render every committed cell its bounding rect covers. That is measured from the
real rect against the real cell grid, so it fails on the unfixed build with
`covers "하" / renders "가"`.

Known limitation: the rendered tail is plain-styled while composing (theme
foreground on theme background, no per-cell colors); colors return on commit.
This is inherent to the overlay, and drawing the preedit into the cell renderer
instead would be a far larger change.

Co-authored-by: rayim <rayim@fxy.global>

* test(e2e): assert the occlusion invariant, not the runner's cell width

CI covered four columns where this machine covers two — 34.4px over an 8.43px
grid against 12.3px over an 8px grid — so pinning the covered text verbatim
pinned the font metrics rather than the behaviour. Assert instead that every
committed cell the overlay covers appears in what it draws, which is the actual
invariant and holds at any cell width.

Still fails against main: covers "하" / renders "가".

* fix(terminal): keep the rendered tail's spacing on the grid

The composition view is white-space: nowrap, which collapses runs of spaces
exactly like normal — it only suppresses wrapping. So a committed tail carrying
padding drew its trailing glyph cells left of where the grid has them: measured
in Chromium with xterm's own rule, twenty spaces plus a border rendered two
cells wide instead of twenty-one.

The visible case is Orca's most common IME context — composing inside an agent
TUI input box, where the row is a prompt, padding, then a real border glyph the
trim cannot drop. A stray border appeared a cell after the preedit while the
real one stayed put.

xterm sets white-space: pre on its grid rows for this reason; the view was only
nowrap-safe while it held preedit text alone.

The existing fixtures are all space-free, and the e2e invariant is that the
overlay renders everything it covers — collapsing makes it cover less, so both
stayed green. Pinned with a padded-row fixture.

---------

Co-authored-by: rayim <rayim@fxy.global>
2026-08-17 00:18:11 -07:00
Neil 17ef6ccce6 fix(terminal): clear the preedit overlay when an IME cancels a composition (#14758)
Backspacing over the last radical of a Cangjie composition empties the IME's
marked text without reaching compositionend, and the vendored xterm
CompositionHelper only dropped the overlay's `active` class there. The box
stayed painted with whatever glyph it last held (#11951).

Clear on the state rather than on the key, as native terminals do: an empty
`compositionupdate` now hides the overlay instead of only ever showing it, and
a key the IME swallows re-derives the preedit from the textarea once it settles
so a composition emptied with no composition event at all is cancelled too.
2026-08-16 15:53:16 -07:00
Neil 991a3fe963 chore(lint): update oxlint to 1.77 and enable no-op cleanup rules (#13901)
Enable eleven oxlint rules that simplify code without changing behavior, and fix
every existing violation. Each candidate was gated on measured cost rather than
assumption, so rules that regressed runtime performance or type checking were
dropped instead of suppressed.

typescript/no-redundant-type-constituents is the largest addition: 113 sites, no
autofix. Dead constituents are deleted. Where the redundant literal existed to
document intent (`string | 'all'`), it is preserved as `(string & {})`, which
keeps the autocomplete hint the original code was reaching for instead of
flattening it away. The rule also caught a broken import —
remote-shared-control-retirement-probe.ts pulled RuntimeStatus from
src/shared/types, which does not export it, so the type silently degraded to
`any`; no tsconfig covers that file, so tsc never saw it.

oxlint stays at 1.77.0 rather than 1.78.0 because .npmrc sets
minimum-release-age=4320 and 1.78.0 is younger than that window.

Rules evaluated and rejected, with what disqualified each:
- prefer-string-raw: String.raw is a runtime call, not a literal (184x slower)
- prefer-string-replace-all: 26% slower
- text-encoding-identifier-case: ~5% slower, reproducible
- prefer-spread: [...str] is 110% slower than split('') and differs on surrogates
- no-implicit-coercion: `!!x` narrows types and `Boolean(x)` does not (22 tsc errors)
- prefer-arrow-callback: arrows are not constructible, breaking `new` on mocks
- object-shorthand: rewrites source text asserted by a tracked reliability gate
- switch-case-braces: pushes ten files past max-lines, which cannot be suppressed
- no-useless-switch-case: drops `case undefined:` that switch-exhaustiveness-check needs
- arrow-body-style: 115 violations have no fix, and it breaks max-lines
- newline-after-import: false-positives on the leading-semicolon ASI idiom

electron-vite-output-contract asserted on the literal
Object.prototype.hasOwnProperty.call text; retarget it to Object.hasOwn, which
rejects inherited keys identically.
2026-08-11 18:19:43 -07:00