Commit Graph
682 Commits
Author SHA1 Message Date
Neil 4fab8e2f15 Fix Smart create retaining a task checkout hash with Create more (#18727) 2026-09-04 15:56:38 -07:00
a65332a8bd feat(claude): move structured native chat onto the Claude Agent SDK and enable it on macOS and Linux (#18560)
* Join structured attach teardown through journal bind

* fix: restore structured chat parity

* feat: add Claude structured session adapter

* fix: harden Claude structured adapter

* fix: close Claude adapter edge cases

* fix: start Claude init deadline after launch

* feat: wire Claude structured sessions

* fix: harden Claude structured runtime

* fix: fence Claude structured compatibility

* fix: preserve Claude free-text prompt answers

* fix: decode addressed Claude prompt text

* feat: enable Claude structured chat on mobile

* fix(mobile): keep structured chat provider-aware

* fix(mobile): negotiate Claude structured tabs

* fix: keep scoped RPC tests native-free

* fix: secure mobile structured image delivery

* fix: close structured session data-loss gaps

* fix: prove real Claude structured startup

* fix: consume pre-spawn proof before retry

* feat(native-chat): add desktop structured sessions

* fix(native-chat): satisfy structured session cleanup gates

* fix(native-chat): keep structured renders pure

* fix(native-chat): open composer pickers upward

* fix(native-chat): use existing view for structured sessions

* fix: harden structured desktop status projection

* fix: close structured desktop lifecycle gaps

* fix: fence structured AI Vault resumes

* fix: fence structured AI Vault resumes

* fix: preserve structured tabs during activation

* feat: toggle structured sessions between chat and TUI

* fix: harden structured session handoffs

* fix: bind structured TUI before rollout proof

* fix: complete structured chat round trips

* fix: align structured TUI return readiness

* fix(native-chat): make reverse handoff transactional

* Add Claude structured TUI handoff seams

* fix(native-chat): clear sticky handoff recovery

* fix(native-chat): complete mobile reverse after TUI exit

* fix(native-chat): keep TUI transcripts readable

* fix(native-chat): recover TUI transcript gaps

* fix(native-chat): recover claimed TUI owners

* fix(native-chat): retain cold TUI proof authority

* fix(native-chat): preserve Claude handoff authority

* fix(native-chat): recover TUI transcripts read-only

* fix(native-chat): harden Claude handoff recovery

* fix(native-chat): serialize structured handoff recovery

* fix(native-chat): close handoff admission races

* fix(native-chat): validate pinned launch environment

* fix(native-chat): revalidate restored and retried owners

* fix(native-chat): gate restart recovery publications

* fix(i18n): catalog Claude session controls

* fix(native-chat): wait for structured TUI process proof

* fix(native-chat): queue stale idle TUI handoffs

* fix(native-chat): route structured Codex options directly

* fix(native-chat): persist structured session options

* fix(native-chat): hydrate resumed structured options

* fix(native-chat): preserve options across structured handoffs

* fix(native-chat): replay pending option mutations

* fix(native-chat): rotate settled handoff operations

* fix(native-chat): rotate refused send operations

* test(native-chat): derive refusal retry state from host

* test(native-chat): give the host-oracle matrix test an explicit timeout

* fix(native-chat): keep Claude option controls idle

* fix mobile structured first-send hydration race

* fix(native-chat): preserve handoff launch authority

* fix(native-chat): harden shared handoff recovery

* fix(native-chat): serialize structured handoff recovery

* fix(native-chat): close handoff admission races

* fix(native-chat): validate pinned launch environment

* fix(native-chat): revalidate restored and retried owners

* fix(native-chat): gate restart recovery publications

* fix(i18n): catalog structured session recovery control

* fix(native-chat): wait for structured TUI process proof

* fix(native-chat): queue stale idle TUI handoffs

* fix(native-chat): keep structured recovery provider-neutral

* fix(native-chat): drop local terminal topology from structured sync

* fix structured outbox and tab restore races

* fix(native-chat): preserve Claude question groups

* fix structured provider visibility and request handling

* fix structured session TUI handoff recovery

* fix reverse structured session handoff

* fix(native-chat): recover Claude outbox and resume state

* chore(mobile): preserve the working-tree lockfile state before the main merge

Carries the pre-existing uncommitted mobile/pnpm-lock.yaml modification into history so the
main merge cannot overwrite it. Verified benign pnpm drift (babel 7.29.7->7.29.8 transitives
plus deprecation metadata); drops no patchedDependencies (the mobile lockfile declares none).

* test(native-chat): drop orphaned Claude handoff-auth test left by the main merge

'pins Claude handoff auth through the terminal provider boundary' is absent from main and its
production counterpart preserveClaudeAuthEnv no longer exists outside this test - orphaned residue
of the terminal/native handoff work this PR excludes by scope.

Removed rather than repaired: the failure was a renamed field (providerHome -> providerRoot), and
renaming it would have carried out-of-scope handoff code into the merge. Body preserved as evidence
and logged in CLAUDE-STRUCTURED-DISPOSITION-TABLE.md.

* Fix mobile structured turn state

* fix Claude structured session blockers

* fix claude structured lane blockers

* fix Claude acquisition exit proof

* fix(claude): route stream-json launch through process wrapper

* fix(claude): gate structured chat support

* Fix Claude structured launch gating

* fix(claude): split session acquisition and prune mobile scope

* test(claude): align structured session fixtures

* fix(agent-session): preserve handoff launch arguments

* fix(claude): open journals through the factory after origin/main split

The journal opener moved to journal-store-factory on main; retarget the
Claude structured tests that still imported the old path.

* fix(claude): resolve Claude structured launch args, auth, and win32 proof

The origin/main merge re-expressed the lane's Claude wiring onto main's split
orca-runtime facade and dropped three wires past green typecheck and lint.

- resolveLaunchArgs discarded its provider parameter, so structured Claude
  sessions were launched with Codex app-server flags; Claude exits on
  --dangerously-bypass-approvals-and-sandbox, and a Codex arg-parse throw
  could block Claude session creation outright.
- resolveClaudeLaunchEnv was no longer supplied, so the launch resolver fell
  back to the whole process env as configuredEnv and
  buildClaudeChildProcessEnv re-applied every auth var it had just stripped.
  The resolver now merges the Claude overlay onto a strip-applied copy of the
  inherited env, which also keeps PATH intact for withCliRuntimeOnPath.
- The windowsProcessStartTimeAvailable producer was gone while the contract
  field and both consumers survived, so the renderer gate fail-closed and
  structured native chat was unreachable on every win32 host.

Separately, structured Claude pinned CLAUDE_CONFIG_DIR unconditionally. An
explicit pin makes the CLI abandon the macOS Keychain even when it names the
CLI's own default, so a default claude.ai account could not authenticate where
the legacy Claude terminal could. Pin only a home the CLI would not resolve on
its own, matching ClaudeRuntimePathResolver, and compare against the env the
child would otherwise inherit so a diverging overlay cannot outrank the
record's account home.

Also await the now-async revealNativeSession in its regression test, and set
the native status before revealing so a rejecting reveal cannot leave a
session released but never marked native.

Claude-Session: https://claude.ai/code/session_013UqKCRB6k5e8UaYhXUHeWY

* fix(claude): scrub case-insensitive Windows auth env

* fix(native-chat): settle handoff outcome-write failures instead of leaking them

A store write failure while recording a handoff outcome escaped the flow
runner's catch handler, so the client never received the failure and the
flow surfaced as an unhandled rejection (seen as an intermittent
agent_session_store_corrupt error in the proven-dead-retry suite, whose
teardown raced the flow's trailing outcome write). Record the failed
outcome best-effort, and drain the coordinator before that test's
teardown removes the store root.

Claude-Session: https://claude.ai/code/session_011aXkcHyeiRJuezupQdjZaM

* fix(native-chat): make the structured close-failure toast provider-neutral

The structuredSessionCloseFailed toast fires for any structured session,
but its copy said 'Codex chat', so a Claude structured session that fails
to close showed the wrong provider name. The launch-failure toast is only
reachable behind the agent === 'codex' gate, so its copy stays as is.

Claude-Session: https://claude.ai/code/session_013ugSpCx4AWkySaJb69BQax

* fix(native-chat): wire structured handoff proof recovery

* fix(native-chat): wire structured handoff proof recovery

* fix(native-chat): correct the structured chat opt-in copy

The one `experimentalStructuredNativeChat` toggle gates both providers —
`useStructuredAgentSessionCreate` runs `canUseStructuredNativeChat` for
`'claude'` as well as `'codex'` — but its description named only Codex.

Its scope line also said Windows keeps using terminal chat, while the gate
refuses win32 only until the host proves it can read a process start time.
`structured-native-chat-availability.test.ts` already pins that Windows is
allowed once the proof is cached, so the two contradicted each other.

Claude-Session: https://claude.ai/code/session_01RJFsidQWmKYFmeoUuVu4Tp

* test(claude): pin @anthropic-ai/claude-agent-sdk 0.3.251 contracts against a scripted CLI

PR 1 of the SDK migration: dependency + test-only harness, no product wiring.

- Pin @anthropic-ai/claude-agent-sdk to exactly 0.3.251 — not the newest
  release — because 0.3.251 (published 2026-08-28) clears the repo's 3-day
  minimumReleaseAge supply-chain gate with no exclusion, while the newest
  release was minutes old and would have required excluding a brand-new
  publish from the exact control built to catch brand-new malicious
  publishes. Every contract this design depends on was verified identical
  on 0.3.251: the full option surface, no pid on SpawnedProcess (custom
  spawner stays mandatory), env defaulting to process.env when omitted, and
  --replay-user-messages appearing only via extraArgs.
- Exclude all eight bundled CLI platform binaries via
  ignoredOptionalDependencies. The setting lives in pnpm-workspace.yaml
  because pnpm 12 no longer reads the package.json "pnpm" field (it warns
  and ignores it; verified by install ablation). Excluding the binaries is
  what makes Orca's pathToClaudeCodeExecutable override mandatory rather
  than merely preferred. Note: pnpm 12.0.0 honors the ignore list when
  reconciling an existing lockfile but not on fresh resolution of a new
  dependency, so the lockfile's SDK entry was pinned surgically; both
  'pnpm install' and 'pnpm install --frozen-lockfile' verify clean and
  stable against the committed lockfile.
- Contract-pin suite drives the real SDK against a scripted fake CLI and pins:
  unknown type/field/content-block pass-through (and keep_alive interception),
  spawner env fidelity plus the omitted-env process.env inheritance sharp edge,
  extraArgs producing --replay-user-messages, argument parity for every
  CLAUDE_STRUCTURED_BASE_ARGS entry plus --session-id/--resume/
  --resume-session-at, canUseTool wire request_id stability and abort on
  control_cancel_request, one spawn per query, pathToClaudeCodeExecutable
  honored by the default spawner, the exact SDK version, and the eight platform
  binaries staying uninstalled.

Claude-Session: https://claude.ai/code/session_01FGCRfYUnb4hbvfTAHGtJKQ

* feat(claude): drive the structured transport through the agent SDK

Replaces the hand-rolled `claude -p --input-format stream-json` transport with
@anthropic-ai/claude-agent-sdk 0.3.251, keeping the existing connection
interface for this commit so the acquisition path changes minimally. The
control-plane rewrite is a separate change.

Orca still supplies the process. `spawnClaudeCodeProcess` routes through
`spawnProcess`, retains the child and its pid — the triple the durable lease
adjudicates on — drains stderr so exit errors keep their tail, and hands `.cmd`
shims to Orca's Windows argument encoder rather than the SDK's plain spawn.
`close()` keeps Orca's own bounded tree-kill and exit deadline, so it still
resolves true only after an observed exit.

Launch resolution emits an SDK options object instead of argv; durable
`launchArgs` translate to a typed option where one exists and to `extraArgs`
otherwise, refusing a token neither can carry rather than dropping it. The
child env is always passed explicitly — omitting it would let the SDK inherit
`process.env` and reintroduce the ambient `ANTHROPIC_*` leak. The stdout line
parser is deleted; the SDK owns framing, and unknown frames still reach the
translator verbatim.

Claude-Session: https://claude.ai/code/session_01JMhFjh9HEnkcJ5YTfCdgD3

* fix(claude): settle the frame the SDK pulled but never wrote

The SDK's input pump is `for await (frame of prompt) { await transport.write(frame) }`.
When that write rejects — the child dies between Orca's liveness guard and the
write — the for-await ends abruptly and calls the generator's `return()`, so the
code after `yield` never runs. The frame was already shift()ed out of `queued`,
so the later `fail()` from the exit path could not reach it and `send()` never
settled: `dispatchClaudeTurn` awaits that send before it can return `unknown`,
wedging the caller and the durable outbox. The pre-SDK transport rejected on the
stdin write callback instead.

Retain the in-flight entry and settle it from the generator's cleanup, and let
fail() reach it too for the pump that never resumes at all.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): keep the agent SDK behind the structured-Claude boundary

The ordinary OrcaRuntimeService graph statically reaches the Claude adapter and
so the transport module, whose first line imported @anthropic-ai/claude-agent-sdk.
The SDK is evaluated whenever the regular runtime loads, before any structured
Claude session is chosen: it sets process.env.NoDefaultCurrentDirectoryInExePath,
changing Windows executable resolution for later subprocesses, and a missing or
incompatible install would break normal runtime startup — for a user who never
leaves the terminal/TUI path.

Defer the SDK to the connection, memoized so it loads once per process, and add
the import-graph ratchet: a walk from the Electron main entry that fails on any
static import of the package, plus a clean-fork check that loading the runtime
leaves the Windows search variable untouched and a child-process pin that the
side effect is still real.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): answer list_models so the picker stops serving the seed

sendControlRequest had no list_models case, so every request hit the default
reject; readClaudeStructuredSessionOptions swallows that with .catch(() => null)
and falls back to the static catalog. Every structured session therefore served a
hardcoded model list with no per-model effort levels, no resolvedModel and no
default detection, and nothing surfaced the failure. The pre-SDK transport got the
live catalog from the CLI.

Route it through the SDK's supportedModels(), wrapped in the { models } envelope
the existing parser reads.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): reap the child's descendants before killing it

The forced step of the exit ladder went through the Codex helper, which spawns
`pkill -KILL -P <pid>` and SIGKILLs the parent in the same tick: the parent
usually dies first, the descendants reparent to pid 1, and `-P` matches nothing.
An MCP or launcher descendant of a stubborn Claude child was left running. The
test named for that requirement declined to assert it and killed the survivor by
hand instead, so it could not fail for the thing it was named after.

Route the Claude reap through Orca's existing sweep, which snapshots descendants
while their parent link still exists and signals them before the root goes, and
on Windows uses the identity-gated `taskkill /T /F`. The test now asserts the
descendant is dead; the manual kill stays only as a failure-safe. close() still
returns true only on an observed exit.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(native-chat): merge the duplicated handoff type import

CI's static-analysis lint (`oxlint --config
config/oxlint-code-quality-native-plugins.json src config tests mobile
--deny-warnings`) exits 1 on the two separate `import type` statements from the
same module.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): answer a permission callback whose signal already aborted

settleFrom registered the abort listener and then delivered the request. A
callback that arrives already aborted never fires that event, so the promise
stayed pending behind a durable prompt with no cancel path. Check the signal
first, emit the cancel, and resolve the SDK's null sentinel without registering.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* test(claude): wait for the child to record the frame, not just for its report

The scripted CLI writes its report at startup, so `until(readReport)` returned a
report with no user messages whenever the child had not yet read the line. The
assertion then failed under parallel load. Poll for the frame instead of for the
file.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): coalesce partial deltas onto one assistant item and stop painting result frames

Under --include-partial-messages every stream_event frame carries its own
uuid, and the final assistant frame for a block carries yet another; only
message.id ties them. The translator keyed each delta by its frame uuid, so a
reply painted as one bubble per delta chunk followed by a complete duplicate
under the final frame's uuid. The block's first stream frame now mints the
claude:(sessionId, uuid) identity, deltas coalesce onto it through the shared
60ms seam, and the final frame reconciles onto that same item.

Known SDK bookkeeping no longer reaches the provider-fallback row: result
subtypes are catalogued and settled by the turn lifecycle, an empty thinking
block (redacted thinking) is a modeled kind, a string-content user replay is a
text block, and an empty user frame paints nothing. An unmodeled result
subtype or content kind still lands on the bounded fallback row.

Claude-Session: https://claude.ai/code/session_01GaP5HpYQbvy2hYehVhwfEW

* fix(claude): prove descendant exit at the close boundary instead of on an unref'd timer

close() reported proven=true as soon as the direct child exited while the
descendant sweep's SIGKILL sat on an unref'd 2 s timer, so a SIGTERM-resistant
MCP server outlived the lease release. The reaper now composes the same shared
primitives the Codex structured provider uses: snapshot, verified bounded
descendant termination on POSIX, taskkill /T /F on Windows. The proof is false
whenever descendants outlive the deadline, a retried close re-verifies the
retained snapshot rather than trusting the dead root, and the raw pipe child no
longer goes through the PTY job sweep it never owned a job for.

Measured on macOS: a killed child of a SIGSTOPped parent stays a matching zombie
row in ps, so the root is killed while verification runs rather than stopped
first as the Codex non-group path does.

Claude-Session: https://claude.ai/code/session_0161QFm3KVRNJKfdzWVGVNWk

* feat(claude): replace the hand-rolled control plane with the SDK's native surface

PR 3 of the Claude structured SDK migration removes the wire-frame scaffolding
PR 2 kept, so Orca drives the SDK's typed control surface directly.

Inbound permissions move from a rebuilt control_request dispatch to the SDK's
canUseTool / onUserDialog callbacks. The prompt registry now carries the
callback's own resolver: a decodable can_use_tool becomes a durable prompt whose
answer settles the callback; a malformed one is denied without registering; the
SDK's abort signal (fired on control_cancel_request, which the SDK matches and
dedups itself) forgets the prompt and settles it null, and a late answer after
abort finds no prompt and is refused. Closing settles every in-flight callback so
no promise dangles. The claude-agent-sdk-control-bridge that rebuilt the wire
frame is deleted.

Outbound control maps to Query methods: interrupt() for cancel, setModel /
setPermissionMode / applyFlagSettings for options, supportedModels for the model
list, initializationResult() for init proof, each under Orca's own request
deadline and error classification. Cancel is interrupt-receipt aware: a CLI
advertising interrupt_cancel_queued_v1 gets cancel_queued in one round trip,
otherwise the receipt's still_queued uuids are swept with cancel_async_message so
a cancelled turn cannot spawn a later unexpected turn; older CLIs resolve no
receipt. Init keeps the 10s deadline and the unauthenticated-startup guidance.

Every behavior is failing-first and ablation-proven; the toggle-off import
boundary and the accepted loss of unknown-control visibility rows are unchanged.

Claude-Session: https://claude.ai/code/session_01Pqjduxt5G4rr9aYvtp7rNm

* fix(claude): arm the descendant snapshot before stdin closes and make the tree verdict unproven by default

A healthy Claude root leaves within the graceful window, and the close ladder
only snapshotted descendants when the root was still alive after that window.
So the common close never looked at the tree: `treeExited` stayed null,
`!== false` passed it, and close() reported a proven exit with an MCP child
still running. A root that died before the walk made the snapshot vacuous too.

The proof is now unproven by default. The reaper holds one verdict in Orca's
vocabulary (exited / live / unverifiable), assigned in exactly one place from
the bounded verification, and close() returns true only on `exited`. The
snapshot is armed before stdin closes, while the root can still be walked, and
is verified after the root exits; a root that left before any snapshot could
be armed stays unverifiable rather than vouching for descendants it never
showed us. The shared verifier gains the three-way verdict behind its boolean
face, and the connection reports the root and tree verdicts separately along
with the child's exit status.

One verification per close attempt: the retried close re-verifies, so the
intra-attempt re-reap is gone from the teardown budget.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): verify the Windows tree after taskkill instead of trusting that it ran

`terminateWindowsProcessTree` resolves from taskkill's callback whatever the
error says, so a timeout, an access denial, a recycled root and a surviving
descendant all looked identical to the reaper — which then returned a proven
exit unconditionally. close() reported true and the lease was released with an
MCP descendant potentially still live.

The Windows branch now snapshots the root's descendants while it is alive and,
after taskkill, polls a fresh process table to a bounded deadline: a row still
matching by pid AND creation time is `live`, an unreadable table is
`unverifiable`, and only a table with no match is `exited`. Creation time is
the PID-reuse guard the POSIX path gets from ps lstart, so a descendant that
denied a creation-time query is omitted rather than signalled on a bare pid.
A root already observed exited is never taskkilled: `/T /F` on a recycled pid
would take an unrelated tree down with it.

The captured tree is tagged by platform so neither verifier can be handed the
other's rows.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): release a reservation on a first-hand root exit instead of latching it into manual recovery

Making close() strict about the descendant tree exposed a second defect at the
same boundary. A create-time acquisition has no ownerProcess until publication,
so an unproven cleanup mapped to handoffStage `manual-recovery`, and
adjudication then refuses every later attach with agent_session_ownership_unknown.
A user who was merely signed out, or whose --resume the CLI rejected, wedged the
session id permanently.

Each question now answers from its own evidence. close() is unchanged and stays
strict about the tree. Separately, the lease is keyed on the root's pid and
start time, so when Orca's own child handle observed that root exit and no
descendant snapshot was ever admissible, the reservation is released and the
CLI's exit code and stderr reach the user. A descendant observed still alive,
or a root Orca never saw leave, stays unproven and keeps the reservation.

The settlement records only what was observed: the released lease says the
provider process exited and its descendants were not verifiable, rather than
reusing the wording that claims cleanup proved no child remains.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): surface an API error a result frame reports instead of settling the turn on it

The SDK models an API failure as a SUCCESS-subtype result whose `result` string
is the user-facing error text, with no assistant frame behind it. The translator
suppressed every catalogued result subtype as turn bookkeeping, so that turn
tombstoned its lifecycle and showed the user a completed, empty reply with no
sign anything had failed.

Suppression is now by meaning. A result reporting a failure routes to the
bounded provider-error surface, leading with the provider's own sentence and
keeping the raw frame behind the row's disclosure; ordinary successful results
stay off the timeline as before. A turn the user aborted also stays suppressed:
its interrupt frame already says so, and its execution diagnostic would only be
noise on every stop.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): drop the stream state of turns that never received their final frame

Every streamed delta recorded its block's identity, latest text and checkpoint
length. Only the final assistant frame removed them, so an interrupted turn left
its whole accumulated reply reachable until the session was disposed, and a long
session with repeated interruptions grew those maps without bound. The partial
text was already journaled by the flush that precedes settlement, so the live
copy was pure retention.

That state now lives in its own module, named for what it does — grow a streamed
block's journal row between its deltas and its final frame — and turn settlement
drops every block still awaiting a final. The translator reports how many remain,
which is the invariant: a settled turn leaves none.

Also makes a timed-out process-table read retryable while the root is still
alive. A loaded host can miss the table's one-second deadline, and latching that
as "no descendants" both lost the descendant sweep and, on a busy machine, made
the close ladder report unproven for a tree it never actually looked at. Only
the root's death still makes a missing snapshot final.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* perf(claude): capture the Windows descendant tree from one process-table read

The capture walked the descendant tree and then read the table again for the
creation times the walk's projection drops. Each read is bounded in seconds and
both run inside the close ladder's budget, so the second one cost the worst-case
teardown three seconds for data the first read already held.

The walk is now exported from the module that owns it and runs over rows the
caller has already read, which is also what lets the snapshot keep the
PID-reuse guard the projection cannot carry.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(pty): spend the descendant verification window instead of surrendering on one slow table read

The verification abandoned the whole check the first time a process-table read
missed its own one-second deadline, with seconds of its window still unspent.
On a loaded host that reported a tree unverifiable without ever having looked at
it, which the Claude close ladder then turned into an unproven close and a
retried teardown. It also made the descendant-exit tests flake under a parallel
suite run, for the same reason and with the same honest-but-premature verdict.

A read that missed its deadline is now simply not an answer: the loop waits and
reads again until its own deadline, and only a window that ends without a
readable table reports unverifiable. This can only turn a premature verdict into
one backed by evidence; it never manufactures a proof.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): never let a later failed look collapse an observed live descendant into unverifiable

The reaper's single assignment site latched only 'exited', so a second reap
whose table reads all missed their deadline overwrote an earlier completed
verification's 'live' with 'unverifiable'. The acquisition release gate
discriminates on exactly that pair, so a root exit after such a decay released
the lease over a descendant that had been observed alive. The latch is now
monotone in trust order: exited is final, and live is only ever raised to exited.

Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP

* fix(claude): never prove a Windows tree gone while a descendant denied identification

The Windows snapshot dropped rows that denied the creation-time query, and an
emptied snapshot was judged exited without any table read: a descendant Orca was
refused information about was treated as one that had left. The snapshot now
counts the unidentified rows it saw, and verification caps its verdict at
unverifiable while any exist. Nothing is ever signalled on a bare pid, as before.

Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP

* fix(claude): classify cleanup after a first-hand exit as a root exit instead of a proven tree

When the CLI died between a successful acquire and the host's commit or proof
of the lease, handleExit had already removed the session, so releaseAcquisition
found nothing and reported true. The attach flow then settled exit-proven with
deathEvidence claiming cleanup proved no provider child remains, though the
tree was never verified. The adapter now keeps the exit that removed a
published session until the session is acquired again; acquisition cleanup runs
that connection's close ladder and classifies its verdict exactly as a
start-time failure would be, so the record reads root-exit-observed. The wire
helper keeps that typed classification and its provider diagnostic instead of
wrapping it as unproven, and the router gives up its owner even when the
release throws.

Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP

* fix(claude): integrate SDK teardown and picker lifecycle fixes

* fix(claude): preserve resume leaf and settle processless spawns

* fix(claude): reacquire from persisted resume leaf

* fix(native-chat): restore Claude grouped question handling

* fix(claude): persist only resumable transcript leaves

* fix(claude): recover structured session exits safely

* fix(claude): close remaining structured session P1s

* fix(claude): harden transcript branch proof

* Remove superseded root fix reports

* fix(windows): restore indexed descendant row walk

* fix(router): forward force-close lifecycle

* fix(claude): fence stale turn cancellations

* fix(claude): fence cancellation after unknown dispatch

* fix(claude): fence replay and option recovery races

* fix(claude): block replay fallback after waiter eviction

* fix(claude): fence evicted slash results

* fix(claude): fence ambiguous results and restore options safely

* fix(claude): scrub SDK child env and localize pending launch

* fix(claude): pin transcript roots and exit recovery proofs

* fix(claude): retain unproven SDK exits

* fix(claude): settle retained exit before reacquire

* fix(claude): resume from settled retained cursor

* chore: remove tracked review artifact

* fix: harden Claude SDK transport session cleanup

* fix: close Claude sessions safely

* fix(claude): close races with fresh child snapshots

* fix(claude): fail closed on recycled child identities

* fix(claude): gate root cleanup on process identity

* fix(claude): fence same-second root identity reuse

* fix(claude): restore the root SIGKILL fallback the identity gate took away

The direct root kill goes through the handle Node owns, not through a pid:
libuv drops that handle in the same turn it reaps, so the signal either
reaches the process Orca spawned or reaches nothing at all. Gating it on a
process-table probe therefore bought no safety and cost the tree its only
fallback whenever the probe declined -- a first capture landing in the fork's
own second, a recycled descendant pid voiding the snapshot, or a process table
that could not be read on either platform.

Identity verification stays where a bare pid is genuinely addressed: Windows
`taskkill /T /F`, and the descendant sweep's own revalidation before it signals.

Also stops a declined root probe from collapsing an observed `live` or `exited`
descendant verdict into `unverifiable`, and stops a successful taskkill from
reporting `unverifiable` because a later probe found the root correctly dead.

* docs(claude): rewrap the root-kill ordering comment

* Match the Claude structured launch to the terminal path's managed-account auth rules

The SDK path stripped ambient Anthropic auth unconditionally, let an explicit
agentDefaultEnv override beat a pinned managed account, and had no account-switch
guard. Reuse the terminal preflight's own predicate and messages so both transports
strip, refuse, and report identically, and cover the CLI transcript location that
mobile native chat depends on.

* Reach the Claude structured chat lane from the desktop UI

The main process has had a complete, correctly gated Claude Agent SDK lane for
a while, but no renderer ever asked for it: the launch route accepted only
`codex`, and the create path was typed `agent: 'codex'` end to end.

Widen both to the structured provider union that already exists
(`AgentSessionHandleProvider`), and generalize the codex-named create path
instead of adding a Claude twin beside it. The pending-launch registry is now
keyed by agent as well as workspace — a shared key handed a second caller the
first agent's intent, so a Claude and a Codex launch in one worktree collided.

Windows, per agent. Codex's client-side win32 refusal is deliberate and settled
elsewhere, so it stays exactly as it was. Claude's answer is no longer guessed
from the client's platform: a structured session fences its provider child on
that child's process start time, and only the executing host knows whether it
can read one. `agentSession.createSupport` already answers precisely that, per
agent, and had no renderer caller — so the Claude create path asks it before
creating and turns a "no", or a probe it cannot get answered, into the
definitive refusal the launch fallback already handles. Fail closed either way.

That refusal mapping also closes a real gap: the host reports an unsupported
location by throwing `structured_agent_session_unsupported`, which reaches the
client as a transport rejection rather than a refusal envelope, so
`StructuredAgentSessionCreateRefusalError` never fired. The launch would retry
the create, strand itself in `visibilityUnknown`, run no legacy fallback, and
show an error toast.

Close a fail-open hole while Claude and win32 become reachable: `create` with a
client-supplied location, and `ensure`, both skip the worktree-resolving support
check. They now ask the executing host the same question directly, so a host
that cannot fence a provider child no longer creates one on a client's say-so.

Also deletes `structured-agent-session-provider-routing.ts`, a duplicate of
`structured-agent-session-provider-support.ts` with no importers.

WSL, SSH and paired hosts, floating workspaces, draft prompt delivery, explicit
TUI customization and initial session options all keep refusing; folder
workspaces keep working.

* P1-1: make the structured Claude auth policy required and testable

The optional dep plus a {stripAuthEnv:false} fallback meant a dropped wiring
under-stripped silently. Required at all three hops, asserted at install time for
the @ts-nocheck caller, and the settings-to-policy mapping is now a named tested
function.

* P2-3: mobile's default Claude transcript root must follow CLAUDE_CONFIG_DIR

session-file-resolver's default ignored the variable the pinned account home
follows, so a CLAUDE_CONFIG_DIR launch wrote one tree and mobile read another. The
Task-4 test now resolves with no root override (mobile's own call) and checks the
answer against the root the CLI itself reports, instead of mirroring the code under
test's own expression.

* P2-1/P2-2/P3: close the teardown window, join the live-auth gate, align the refusal

P2-1: a switch beginning inside the acquire teardown left a dead chat and no
replacement. Past that point the launch waits the swap out and refuses only if it
never settles; the entry guard still refuses outright, because nothing is torn down
there yet.
P2-2: structured children now hold the same OAuth-refresh gate a Claude PTY does,
so a managed refresh cannot rotate the token out from under a live turn.
P3: the refusal now matches the strip it guards (case-folded on win32, presence not
truthiness), and the dead structured-to-TUI builder states its auth policy instead
of silently signing a system-auth user out.

* Make the live-auth gate tests independent of sibling connection teardown order

* Do not offer structured Claude under a WSL-only managed account

Structured Claude launches against the ambient Claude config, which the account
service keeps in sync with the selected HOST account. A WSL-bound managed
account lives inside the distro and is never synced there, so on Windows a
structured session would authenticate as whatever the ambient identity happens
to be while the UI names the WSL account — the user is told one identity and
given another.

That was unreachable only because nothing offered structured Claude on win32.
Enabling it makes it reachable, so gate it here rather than patching the auth
layer: refuse the structured path when the active managed Claude account is
WSL-bound, and let the terminal-backed path — which resolves the account per
runtime — handle that account shape.

The answer rides the agentSession.createSupport seam the renderer already
consumes, so no new capability and no renderer knowledge of account internals.
A create the host declines becomes the definitive refusal the launch fallback
already turns into a legacy native chat tab, with no error toast.

Unknown answers refuse. An install with no managed accounts claims no identity
and is fine, but an active selection that cannot be resolved — or account state
that cannot be read at all — is not evidence that the ambient identity is right.

Claude only. Codex resolves its account through a different path and its
createSupport answer is untouched, as is every Codex routing decision.

* Read the structured Claude account gate through the auth policy's accessor

The gate resolved the active account from the account-service snapshot's
runtime map; the auth policy resolves it with
getSelectedClaudeAccountIdForTarget(settings, { runtime: 'host' }). Those are
two sources and two resolution rules, and they disagree on a legacy settings
blob that carries the selection only in the flat activeClaudeManagedAccountId:
the accessor falls through to it, a direct read of the runtime map does not. The
gate would then refuse a launch the policy would have run under host-1 — and in
the mirror case a session could be admitted under a policy computed from a
different account than the gate approved.

Read the same settings through the same accessor so agreement is structural
rather than coincidental, and drop the controller accessor that existed only to
reach the snapshot.

No behaviour change for any state both already agreed on; Codex is untouched.

* Round-3 review fixes: N-1 empty-value regression, N-2 gate leak window, N-4 lost history

N-1: my presence-based conflict predicate refused a terminal launch that works
today. 'ANTHROPIC_API_KEY=' is how a user blanks a variable and the settings
pipeline preserves that empty value; an empty override cannot beat the pinned
account and the strip removes the name anyway. Back to truthiness for the value,
keeping the win32 case folding.
N-2: enter the live-auth gate only after the exit/close handlers that release it,
so no throw in between can leave an entry nothing reconciles.
N-4: the Claude transcript resolver searches config-dir-then-default and de-dupes,
matching the Codex sibling in the same file, so adopting CLAUDE_CONFIG_DIR no
longer hides history written before it.

* Run the managed-account gate on every Claude acquisition, not just create

createSupport gates the create path, but a session's account state can change
while it lives. A reacquire after an unexpected child exit re-resolves the
launch and re-derives auth, with nothing re-checking the gate — so a session
created while supported could come back up in the refused shape. With the strip
predicate keyed on there being an active non-WSL account, the WSL-only user's
normalized steady state (accounts exist, none active) does not strip, and that
reacquire reaches the child with ambient auth while the UI names the account.

Gate at resolveLaunch, the one choke point every acquisition passes through,
refusing with the pre-spawn error the caller already handles. Same predicate as
create-time, now sharing one settings reader so the two cannot drift.

Claude only; Codex resolves its account on a different path and is untouched.

The runtime class that wires this does not typecheck its own `this` calls — a
missing hookup compiles clean — so the wiring is pinned behaviourally rather
than trusted to the compiler.

* Move the structured Claude gate out of the @ts-nocheck runtime files

Both call sites of the managed-account gate sat in files whose first line is
`// @ts-nocheck`, so neither was typechecked: three arguments to a one-argument
function plus an undeclared identifier compiled clean. New auth-identity
decision logic had no compiler behind it.

Move the verdict into a checked module that takes the two facts the runtime
owns — the adapter's answer and a settings getter — and decides. The runtime
class now only forwards. Move the gate reader's construction into the checked
installer too, so the nocheck file passes a plain settings closure and never
names a gate symbol.

Every reference to the gate predicate and its reader now lives in a checked
file, so the ablation that used to pass silently is a compile error at both the
create-support and reacquire sites.

Removing the file-level @ts-nocheck is a separate, larger job and is not
attempted here.

* Derive the gate test's auth policy from the settings under test

A hardcoded stripAuthEnv asserts a gate/policy pairing production cannot
produce, and false additionally lets launch.env inherit the runner's real
process.env. Derive via claudeStructuredAuthPolicyForSettings instead: the
gate settings type is the same Pick the policy takes, and both resolve the
account through getSelectedClaudeAccountIdForTarget.

* Pin the absent-vs-empty distinction in the managed-account gate

An empty claudeManagedAccounts array is a real answer: the user has no managed
accounts, nothing claims an identity, and the ambient path is legitimate. A
readable settings object with no such field is settings we failed to parse —
the same unknown as unreadable — so it refuses.

The two are one character apart in the code and the difference is invisible
without the reasoning, so record it at the branch and pin both sides. The test
fails under the obvious "consistency fix" of treating a missing field as empty.

* fix(claude): keep command queue bookkeeping out of the transcript

Claude Code 2.1.258 emits a `command_lifecycle` frame for every uuid-stamped
command it starts, completes or cancels. The frame carries a command uuid and a
state and no content, and the CLI keeps it out of its own transcript -- but it
is absent from the SDK's SDKMessage union and so from Orca's frame catalogue,
where an uncatalogued kind defaults to a substantive row. Every structured turn
therefore painted raw JSON rows into the user-visible transcript.

Catalogue it and disposition it as status chrome. The unknown-kind default stays
`timeline-substantive`: a kind we have never seen is likelier to carry content
than to be chrome, and a visible row we can catalogue later beats content we
silently dropped. A lifecycle state that reads as a failure still surfaces,
because the payload error check in `classifyProviderFrame` outranks the
catalogue.

* fix(claude): let a re-walked descendant become eligible for the forced sweep

A descendant first observed by a capture inside its own birth second could never
be SIGKILLed: `ps lstart` is second-resolution, so that capture cannot rule out
a pid recycled later in the same second, and the merge pinned each retained row
to the boundary of the walk that first saw it. SIGTERM-resistant children forked
in that window were signalled and then never escalated -- they survived close,
quit and restart, reparented to init, and had to be killed by hand.

Advancing that boundary on any later capture would be unsound: a later capture
matching pid, pgid and start-second is exactly what an impostor would also show.
But a capture is not a match -- it is a fresh ppid walk from a root Node pins
through its own handle, so a row it re-derives is proved ours at that instant
without appealing to its start time. Chain the fence from there instead, and
take that walk at the close boundary while the root certainly still lives: the
root may leave inside the grace window, and the post-timeout refresh never runs.

A row absent from the later walk still keeps its earlier boundary, and a row no
walk has ever re-derived in a later second is still never escalated.

* Treat an absent managed-account list as empty, not as unreadable

An empty claudeManagedAccounts array and a missing one are the same answer:
this user has no managed Claude accounts, so nothing claims an identity and
ambient auth is the truth. Refusing on absence strands any profile that simply
never wrote the key, and it disagrees with the auth policy, whose own predicate
takes `(accounts ?? [])` for exactly this reason.

Only settings that cannot be READ stay unknown, and those still refuse — as do
a WSL-bound active account and a selection naming an account the list does not
explain.

The earlier reasoning treated a missing field as settings we failed to parse.
That conflated "not present" with "not readable"; only the second is unknown.

* Support structured Claude when accounts are registered but none is selected

Registered-but-deselected Claude accounts were refused, which is behaviourally
identical to having no accounts at all: the auth policy does not strip, ambient
auth is the truth, and the UI names no host identity. A user who deselected
their accounts silently got legacy chat with nothing explaining why.

Nothing selected for the host runtime is two states the settings cannot tell
apart after the fact, because pruneInvalidClaudeRuntimeSelection empties the
host slot and persists null in the second one:

  honest deselection      -> ambient auth, UI names nothing   -> SUPPORTED
  the WSL-only steady state -> ambient auth, UI names the WSL account -> REFUSED

The presence of any WSL-bound account in the list decides. Simplifying this to
"none active -> supported" re-opens the auth-identity misrepresentation, so the
tests fail loudly on exactly that: five of them, across the unit rule and the
createSupport path.

* Stop treating an unanswerable create-support probe as a refusal

A worktree is not resolvable for a beat after createWorktree resolves, so a
probe fired immediately after creation fails the RPC with selector_not_found
instead of answering. The catch collapsed that into `supported = false`, so the
composer refused and quietly built a terminal session — the gate never said no,
it was never asked successfully. Elapsed time was the only input that decided
whether a Claude launch went structured.

"Could not answer" and "answered no" are different states and only the second
is a verdict. Retry while the host cannot yet resolve the selector, with a
bounded backoff that covers the measured window with margin, and keep refusing
on the first ask for everything else. Fail-closed is unchanged: a probe that
still cannot be answered when the budget is spent refuses.

The retry is narrowed with the shared error-code matcher, which classifies a
token that transports re-wrap into a longer message without matching prose that
merely mentions it.

Codex never probes, so this race has never been able to refuse a Codex launch —
the race itself is identical for it. Recorded at the early return, because
whoever gives Codex a probe inherits the bug.

* fix(claude): fence the forced sweep on re-derivation, not on lstart's second

A descendant forked in the same wall-clock second as every walk that sees it was
signalled with SIGTERM and then never escalated, so a SIGTERM-resistant child
survived tab close, app quit and a full relaunch. Two children of one parent
96ms apart across a second boundary took opposite paths. The leak predates this
branch: it reproduces with the change reverted.

`ps lstart` has one-second resolution, so a walk landing inside a row's birth
second can never rule out a pid recycled later in that same second. But a walk
is not a match: a ppid walk only reaches what the root actually parents, and the
root is pinned by Node's own handle, so a row the walk re-derived is ours
whatever second it was born in -- a stranger would have to have been forked into
our tree, and then it is not a stranger. Fence the escalation on that.

Rows a merge retained from an earlier walk are not re-derived and still answer
to the start-time fence, which remains correct for them.

Scoped to callers that revalidate identity before signalling, which is the
Claude close path. Codex teardown reaches this same verifier and is unchanged;
the argument holds there too, but widening it is its own deliberate change.

Also reverts two changes from the previous attempt at this leak. Advancing the
capture boundary on a later walk is inert once the sweep fences on re-derivation
-- both key on the same set of rows, so the new term short-circuits for exactly
the rows whose boundary it advanced. The extra ladder refresh was a duplicate
full process-table read: close() already awaits tree.refresh() immediately
before proveClaudeChildExit, on the only path that reaches it.

Known property: the kill lands roughly a grace window after the walk that proved
membership, so a pid recycled inside that gap could in principle be signalled.
It is bounded -- matchingSnapshotRows already requires the live row to carry the
same start-second and pgid, so an impostor must be born in the remainder of that
one second, land on that exact pid, and sit in the same process group, and it
has already received the unfenced SIGTERM from the same loop.

* Run the Claude structured integration suite as a runtime client

The suite exercises agentSession.* for Claude, not the mobile surface: nothing
in it asserts anything mobile-specific and its sibling integration suites use
'runtime'. Mobile now additionally requires the experimental structured-chat
setting, which structured-agent-session.test.ts pins in both states, so the
stale 'mobile' fixture was claiming coverage it never had.

* fix(claude): report effort from get_settings, which is the only frame that has it

The composer's Effort pill rendered blank in every structured session. This is
not a missing source: the publication reads `effortLevel` off the `system/init`
frame, and that frame has never carried an effort of any kind, while the correct
value is already fetched at acquisition and thrown away on the auth diagnostic.
Verified two ways -- a live get_settings probe against Claude Code 2.1.258, and
the shipped binary's own init frame construction, which lists `model` and no
effort. So `reportedOptions.effort` was always empty, the options reader dropped
the key, and the pill had no value. Model survived only because
`currentModelId()` has a fallback chain.

The get_settings call acquisition already makes reports the session's current
effort as `effective.effortLevel`; pass that into the publication instead.
Selecting an effort already worked, so this is the arrival value only.

The legacy PTY path is unaffected and must not be "fixed" to match: it reads its
effort by parsing the startup banner (`CLAUDE_MODEL_EFFORT` in
src/renderer/src/components/native-chat/claude-terminal-session-options.ts),
which is why it shows a value where the structured path does not.

Also removes the fixture that hid this: the fake init frame invented
`effortLevel: 'high'`, a field the CLI does not send, which is why every gate
stayed green over a value that is always empty in production. The fixture's
get_settings now returns the real {applied, effective, sources} shape instead of
a bare `{env: {}}`, so the two adapter tests that asserted an effort keep
asserting it through the path production actually uses.

The reader returns null rather than defaulting: an effort nothing measured would
repeat the fixture's mistake, and a blank pill is the honest degradation if the
provider ever renames the key.

* fix(claude): only record an effort the child confirms it adopted

apply_flag_settings answers `success` for an effort it then ignores. Measured
against Claude Code 2.1.258: applying `bogus-effort-xyz` returns
subtype "success" with no error while `applied.effort` stays at its previous
value, and a valid `low` moves it. The option write treated the absence of a
throw as adoption and recorded the requested value unconditionally, so Orca
would show and persist an effort the child was not using, with nothing anywhere
reporting a problem.

Read the effort back after applying it, through the same reader the arrival
value uses, and reject when the child reports a different one. A readback that
could not be taken is not evidence of a refusal -- the apply itself succeeded --
so it still records; only a readback that disagrees rejects.

Not reachable from today's picker, which offers catalog values only, but the
CLI's effort catalog is server-delivered and has changed before, so a retired id
would otherwise become a pill confidently displaying a setting that never took.

* test(claude): assert the effort contract against the real binary

The blank pill survived every gate because the only tests that touched it were
fixture-backed, and the fixture invented the field. A test that pins the shape
we read cannot catch the provider renaming the key, which is the failure mode
that produced this defect.

Asserts both halves against a live authenticated CLI: that no frame it publishes
carries an effort at all, and that the session's current effort arrives through
get_settings. Which frame proves the session varies by host -- this machine
proves it with a SessionStart hook rather than a system/init frame -- so the
negative half asserts over every published frame rather than picking one.

Skips with the rest of the file when no authenticated CLI is present.

* fix(claude): stop the synthesised content-part kinds leaking into the transcript

Sending an image put a bare `claude · message:user:content:image` row between
the user's bubble and the answer. Two causes, and only the second is a family.

An image part counted as modelled only when `source.type === 'url'`, but
claudeDispatchMessageContent sends a local attachment as a base64 source and the
CLI replays that shape back, so every attached image was classified unmodelled.
Accept the base64 and file sources Orca itself sends.

The family is the real defect. `message:<role>:content:<type>` kinds are
synthesised at runtime from whatever `part.type` arrives, so unlike the
top-level frame catalogue they can never be enumerated ahead of time -- the
`?? 'timeline-substantive'` default then prints the synthesised name at a user
who cannot act on it. That default is right for top-level frames, where
"substantive" means show the frame; here it meant show our own vocabulary, which
drops the content AND leaks the opcode.

So an unrenderable part now renders a sentence saying exactly that, with the
kind and payload still on the row's disclosure. A part that carries its own
readable sentence keeps it -- the placeholder is a fallback, not an override.

An unknown future part type is therefore visible, never silently dropped and
never printed as a kind: the same principle as the effort readback, which
records only what the provider confirms.

* Declare agentSession.requestHandoff on the cross-version wire surface

The manifest is a ratchet for cross-version reachability, so the method is
declared with real HandoffParams rather than counted. requestHandoff is
capability-gated through requireStructuredHost and has no client caller, so
declaring it is the whole of the change.

Also model two host capabilities the harness omitted: the stub host's
supportsCreate, and the fake adapter's, without which adapterSupportsCreate
falls through to a supportsLocation the fake also lacks. Every ensure was
refused for the harness's silence rather than for its location.

* Gate structured Claude session tabs on the client capability that names them

The Claude structured lane deleted the projection's `agent !== 'codex'`
filter and added CLAUDE_STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY in the
same commit, but never wired the constant to anything. Paired clients then
received agent-session tabs for Claude, which no shipped client renders --
mobile's resolveMobileNativeChat returns null for every agent but codex, so
the row listed and selected into a pane with neither chat nor terminal.

Restore the filter behind the declared capability instead of the bare agent
name. No client advertises it yet, so this matches main's behaviour today
and becomes a negotiation a future client can opt into.

* Confirm the structured Claude model against the model the CLI reports

set_model answers success for any string, including a model it cannot
resolve — the failure only surfaces when the turn runs — and get_settings
reports the settings-file model, not the session's. The init frame that
opens each turn is the only channel carrying the adopted model, so keep
the session's reported model current from it instead of reading it once
at acquisition.

Also stop rejecting an effort the readback cannot represent: max is
session-scoped and excluded from the persisted effortLevel, so a readback
reporting the level underneath it is an absence of evidence, not a refusal.

* Clear the session-option hedge when the provider confirms the value

The pill claimed every option was unconfirmed for the life of the session:
the renderer recorded each write as dispatched and nothing ever moved it,
so a model the CLI had already reported back still read as unconfirmed.

Carry the provider's own confirmation to the surface. Main reports which
option ids the provider named rather than merely accepted, and the client
re-reads options as a turn changes, because the frame that opens a turn is
where the adopted model arrives. A value the provider has not reported
stays hedged, including an effort whose readback could not be taken.

The confirmed list is optional on the wire: a host that predates it sends
nothing and the client keeps hedging, which is the behaviour it had.

* Keep the model report current across an acquisition fence bump

* Show the picked session-option value and let the provider report correct it

The pill showed a "not confirmed" second tooltip line for any value we had sent
but not yet seen reported back. Nothing acts on it, and for the PTY lane it was
permanent — that transport has no report channel. The pill now shows the picked
value immediately and the provider's per-turn report corrects it when the two
disagree; a newer local write still outranks a report that precedes it.

`dispatched` stays as a provenance member rather than collapsing into `applied`:
it is produced independently by the PTY lane, and it is where the `confirmed`
wire field lands, which would otherwise be unobservable.

Effort keeps its readback and its rejection path. That matters more now, not
less: with the hedge gone the rejection is the only user-visible failure signal
on this surface, so a spurious one would be the loudest bug here. Skipping the
readback for an effort the settings response structurally cannot echo is what
prevents it — the response carries the persisted level, so reading it back for a
session-scoped value would report the level underneath and fail a valid write.

* Hedge a session-option value only when the terminal transport sent it

Both lanes emit `dispatched`, so it could never say which one produced a value.
The descriptor now carries the transport that built it, set once in the shared
snapshot builder from a parameter that is required rather than defaulted — the
builder is the only place a descriptor is constructed, so a new producer has to
name its lane or fail to compile.

The structured lane confirms every value from the provider's own per-turn report,
which makes the hedge transient noise there. The terminal lane can only learn an
outcome by parsing the screen back, and only for Claude: every other agent's
`dispatched` value stays unconfirmed for the life of the session, so the line is
the only signal that we sent something we never saw land.

* Refuse an effort the session's model advertises no control for

* Refuse tab mutations on a Claude row the client never negotiated

The branch added a case asserting a client advertising only
agent-session.structured.v1 may mutate a claude row. That is the same
ungated behaviour the projection gate removes, encoded a second time —
mutation authorization reads the projection, so hiding the row refuses
the write. Assert that contract instead, and add the positive case for a
client that does negotiate Claude rows.

* Resolve the Claude session's current model in one place so the effort guard and the pill agree

* Record an effort the child did not adopt instead of refusing the write

apply_flag_settings answers success for an effort it then ignores, so the
readback exists to detect that. Refusing on it made the detection a veto,
and a veto is only correct if the readback can never be wrong about which
model is current -- which it was, twice. The pre-flight guard already
refuses a level the model advertises no control for, so the veto guarded a
door that is now locked upstream.

Keep the detection, drop the refusal: a disagreement records the child's
own answer and omits the option from confirmed, so main stops vouching for
a value the provider rejected without blocking the user's write.

* Stop a slow whole-machine ps from being read as an absent process

`ps -axo ...command=` pays a per-pid argv read: measured 1.15s for 1,948
processes (0.03s without `command=`), and CPU contention stretched the same
capture to 6.0s. Two budgets sized for a cheap look then misreport a readable
machine.

The reader's 3s ceiling killed 6 of 20 consecutive captures at load 27, so
every consumer answered "unverifiable" about a table it could read. Raise it
to 15s, and stamp the capture instant at ps START so `capturedAgeMs` is the
upper bound its contract promises -- a 6s capture used to report itself as
freshly taken, understating staleness against a 5s kill gate. The TTL keys on
completion so a slow capture still coalesces instead of forking ps per caller.

`readStructuredTuiProcessIdentity` then spent its whole 5s wait inside one
capture and concluded "no exact child" after a single look taken before the
child existed (observed landing at ~3.5s). Absence needs a look that did not
race the spawn, so require two captures before the deadline can end the loop.

Both surfaced by the real-binary Claude TUI resume test, which failed ~1 in 5
under load; 14/14 now, 8 of those runs containing a capture the old 3s budget
would have killed.

* Let the desktop renderer negotiate Claude structured tabs

The paired-client gate hides agent-session rows an agent the client cannot
render. The desktop renderer's own IPC dispatches as clientKind 'runtime'
advertising only agent-session.structured.v1, so the gate hid Claude rows
from the surface this feature ships on. It renders them; it should say so.

* Stop a slow process table from silently blinding every freshness gate

Stamping `capturedAgeMs` at ps START made the number honest, and honest broke
both consumers that read it. `ps -axo ...command=` measured 2.5-9.0s on an idle
2,002-process laptop and 4.0-18.6s at load 46, so the age it now reports lands
past every budget: `planRelayPtySweep` refuses the stop as "too old", and the
renderer's `admitRemoteForegroundEvidence` refuses the record outright. That
second one is the expensive half and was outside the diff -- a refusal bumps
`consecutiveInspectionErrors`, the poll scheduler backs off to its 10s floor,
and agent-completion detection stops for the pane. The subsystem went blind on
exactly the loaded hosts the honest stamp was meant to serve.

The evidence-publishing read now gives up at 1,200ms instead of waiting out
`PS_TIMEOUT_MS`. It is one budget for one question: these consumers ask whether
an observation describes NOW, and past this it does not -- a late answer is
refused by the age gate anyway, having first blocked a polled path for the whole
capture, so a prompt `unverifiable` is both the truthful verdict and the cheap
one. Both relay call sites already produce it from a rejection, and an admitted
`unverifiable` costs a poll where a refusal costs the cadence. Identity proof
keeps the full 15s through `getFreshProcessTableSnapshot`, because it asks
whether a process EXISTS and must never read slow as absent. The budget bounds
the wait, never the capture: the reader coalesces, so an abandoned wait leaves
its capture running to fill the cache rather than forking a second whole-machine
`ps` on the host that can least afford one.

1,200ms is bracketed rather than picked. The floor is the capture's own cost --
`command=` measured 1.15s for 1,948 processes on an idle host, and a budget
under that answers `unverifiable` about a machine nobody is straining. The
ceiling is the consumer's: 2,000ms, less the 500ms a TTL-shared capture may
already have aged, leaves 1,500ms, and transit takes the rest.

That ceiling only fits once the capture stops being charged twice. `ps` runs
inside the RPC round trip, so its duration is already in `receiveDelay`, and
`capturedAgeMs` is that same duration on the host's clock; summing them halved
the budget this gate grants a host from ~2.0s of `ps` to ~1.0s, which is why a
1.2s capture arriving at 1.3s read as 2.5s old and was refused. Admission now
takes the larger of the two. The sweep's gate keeps its sum, which is correct
there: `evidenceAgeSinceListingMs` is stamped after the listing ARRIVES, so it
measures planning time and overlaps nothing.

A stated limit rather than an assumed one: 15s is not proven sufficient for
identity proof. The same capture reached 18.6s at load 46, so that path can
still time out and answer "no exact child" about a host it simply could not read
in time. Narrowing it needs a cheaper question than a whole-machine argv read,
not a larger number.

The one test guarding this field could not fail. `beginPtyHandlerTest` installs
fake timers, so `Date.now()` is frozen, the real reader reports exactly +0, and
`0 <= 500` held identically for a hardcoded zero, for completion-stamping and
for start-stamping -- while the real reader on that host returns thousands of
ms. It now drives a measured age in and asserts the handler publishes it rather
than restamping; that the reader MEASURES it correctly stays pinned separately,
against a controllable clock. Both consumers get boundary coverage either side,
and each new gate was ablated red before it went green.

* Keep the compatibility fields off the capture the budget just abandoned

inspectProcess falls back to processHasChildren and listProcesses to
getForegroundProcessName, and both read the same TTL-shared capture with
no budget of their own. On a slow host they joined the in-flight capture
the budgeted evidence read had just given up on, so the call still blocked
for the full 6-18s and the budget bought nothing -- once for inspectProcess
and once per managed PTY for listProcesses.

Use the degraded answers those helpers already give for an unreadable
table, reached promptly. pty.hasChildProcesses keeps its unbudgeted fresh
probe: it is a one-shot destructive gate that can afford to wait.

---------

Co-authored-by: Merge Sim <merge-sim@local>
Co-authored-by: Merge Sim <sim@local>
2026-09-04 15:55:20 -07:00
Neil e42c60e8a3 fix(ssh): resolve a pane's binding from the target partition, not the stale local copy (#18546)
One SSH pane accumulated one extra reattachable lease per relay restart (2, 3, 4,
5, 6 across five), and every one of them costs a `pty.attach` round trip on every
later connect, forever. Nothing prunes `sshRemotePtyLeases`, so the fan-out only
grows.

`supersedeSiblingLeasesForPane` is fenced on the PTY the pane is durably bound to,
and `durablyBoundPtyIdForPane` read `state.workspaceSession` (local) before
`workspaceSessionsByHostId['ssh:<target>']`. But `persistPtyBinding(binding, hostId)`
updates ONLY the host partition:

  AFTER-PERSIST  local= ssh:t@@pty2:old:1   host= ssh:t@@pty2:new:1

So for the length of a reconnect the local copy still names the predecessor, the
fence resolved to it, supersession took an already-`expired` lease as its winner,
and returned having marked nothing. Both partitions agree again once the renderer
republishes its layout, which is why the settled store looks consistent and hid
this.

Read both partitions as an ordered list, target's own first, and test the fence by
membership rather than by equality with whichever was read first. Pick the winner
preferring a lease this client still has a route to, since the stale partition
names an expired one. Never retire a lease that is both bound and live, so a
partition disagreement can't strand a running remote process.

Superseded predecessors stay `expired` and are never `terminated`: losing a lease
is not evidence the shell died (docs/reference/ssh-execution-boundary.md). A pane
with no binding is skipped rather than pruned, so a genuine orphan stays askable.

Also re-runs supersession from the binding side after each spawn commit's binding
write, so the lease/binding order at a call site no longer decides, and reconciles
every pane for a target immediately before `reattachKnownPtys` reads the set it
feeds to `pty.attach` — that repairs stores which already accumulated these rows.

The guard suite could not catch this: every assertion bound the pane BEFORE
upserting the lease, an order no caller uses. Rewritten to the spawn commits' real
order (lease, then binding, then the binding-side trigger); it fails 8 assertions
without this change. Added a suite that drives the real `persistPtyIpcSpawnCommit`
rather than the store primitives, including the exact stale-partition state written
by production's own binding writer.

Verified on the Docker SSH lane: five `relay.js` SIGKILLs with recovery between
each, reattachable leases flat at one per pane.

Note: this bounds the reattach SET, not the store. `sshRemotePtyLeases` still has
no cap or TTL and rows still accumulate; pruning is left alone deliberately, since
an `expired` row without `supersededBy` is a genuine orphan and must not be dropped
on age.
2026-09-03 16:47:32 -07:00
Brennan BensonandMerge Sim 98e77ef1a7 feat(mobile): structured native Codex chat (#18074)
* feat(mobile): finalize structured native Codex chat

* fix(mobile): close structured chat lifecycle gaps

* wip(mobile): fence stale structured inventory and bound operation-id retention

Fence local structured-session inventory and subscription responses with a
sync generation so a toggle-off clear, reconnect restore, or retry cannot
apply a mirror from a superseded instance. Bound mobile ambiguous
operation-ID retention at 128 with unmount cleanup.

Staged on the reconcile branch only: the sync module is now 312 lines and
needs a real split before this can reach the PR head.

* fix(ci): split the structured session-tabs sync and give static analysis mobile types

The local structured session-tabs sync module outgrew the 300-line cap once it
took on generation fencing, so split it along its real seams instead of raising
the cap: the generation/cursor fence, snapshot projection, snapshot apply,
inventory refresh, and the subscription loop. The original path stays as a
barrel so no importer moves.

Repoint the host-session-mirror settle census at the apply module, which owns
two receipts now — the snapshot it mirrors in, and the toggle-off teardown that
retracts what it published. The teardown receipt is named rather than anonymous
so the pin says which direction it settles.

The changed-code quality gate lints mobile files and resolves their types from
mobile/node_modules, but mobile is a separate pnpm project that the root install
never populates, so every mobile type degraded to an `error` type and the gate
reported phantom findings. Install mobile dependencies in static analysis when
the diff touches mobile, gated on a new classifier output.

* fix(mobile): let a slow capability handshake still reach connected

The mobile capability update is an advisory whose result is discarded, yet an
unanswered one was fatal while an explicit rejection was tolerated. A 5s timeout
on the direct client force-closed the socket, and on the relay path it failed
`confirmResume` before `connected` was ever published, so a consistently slow
link redialled forever. Both paths now share one helper that settles every
ambiguous outcome (timeout, mid-flight drop) like a rejection and rejects only
when the frame never reached the wire — the one case nothing else recovers from,
since the socket's own desync force-close is gated on already being connected.
The generation guard still keeps a replaced session from connecting.

Retained structured-session operation ids were capped at 128 with oldest-first
eviction, but every retained id belongs to a send whose outcome is unknown, so
eviction turned a user's retry into a second message on the host. Bound the map
by expiry against the id's own embedded timestamp instead, mirroring the host's
operation ledger, so no id is released while the host would still honour it.

Also give the mobile CI install the root install's lockfile drift guard (mobile's
lockfile carries patchedDependencies a silent rewrite would drop), gate
mobile_dependencies on should_run, and key the pnpm store cache on both lockfiles.

* refactor(mobile): extract the relay pending-request registry

The merge composed two independently-sized changes — this branch's capability
handshake settle and main's dial-stage tracking — pushing the relay session file
to 304 lines against a 300 cap. Neither side broke it alone.

Move the in-flight request registry (id generation, tracking, settlement, and
reject-all with its delivery-ambiguity marking) into RelayPendingRequests,
matching the existing collaborator pattern alongside RelayDialStageTracker and
RpcSessionLivenessWatchdog. No behavior change.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-03 15:19:26 -07:00
Brennan BensonandMerge Sim d5803bdbc4 feat(ssh): host-stamped remote foreground identity (#18078)
* docs: add SSH agent identity implementation plan

* feat(ssh): host-stamped remote foreground identity

* fix(runtime): preserve unfenced inspect call shape

* perf(ssh): traverse foreground descendants linearly

* fix(ssh): bound retired PTY evidence records

* test(ssh): cover retired incarnation retention

* fix(ssh): make remote process inspection total

* Split SSH identity build hot spots

* Fix process table snapshot module split

* test(ssh): update process inspection expectations

* docs: drop the SSH identity plan from the PR

The design doc does not belong in the product repo; it stays out of the
shipped tree while the implementation carries its own comments.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-02 23:32:41 -07:00
Neil 6cd477a2f1 test(e2e): un-rot the SSH freeze repro and probe two failure modes nothing covered (#17940)
Test-only. No production code.

## The freeze repro was rotted in three ways, not one

#16764 tracks four stale call sites. There were three separate problems:

1. **Stale call sites** — `execInTerminal` gained a `ptyId` and
   `splitActiveTerminalPane` gained a direction. (`startDockerSshRelayTarget`'s
   missing `testInfo` was the third; #18257 has since landed it on main.)
2. **It connected before session restore settled**, so the seeded tab never bound
   to a remote PTY and the terminal sat on "Connecting…" forever.
3. **It could never have passed, even once.** It waited for a one-shot `READY:`
   line through a 4000-char terminal window while its own 2 KB-every-8 ms flood
   buries that line within ~16 ms. Readiness is now keyed on the repeating `BG:`
   flood marker, which is strictly stronger — it proves the pane is streaming
   rather than merely started.

It now runs end to end and prints a measurement instead of dying on a call site:

```
[freeze-repro R2] hiddenFloodMaxLagMs 2.1  bulkOpenMaxLagMs 41.5
                  interactionProbeMs 53.6  softFreeze false  hardFreeze false
```

**It is still not CI-gateable, and the exclusion comment now says so.** The same
spec on the same commit measured `bulkOpen 2575.6ms / interaction 3464.2ms` on a
GitHub ubuntu runner against a 2500 ms soft budget — a ~60x spread on the number
the budget reads, with the relay still streaming. That is the budget failing, not
the product. The earlier draft of this comment claimed "repaired and passing",
which was true only of the host it was measured on; gating this needs a
host-relative oracle, not a bigger constant.

## New: a half-open link is judged, not wedged

The fixture image has no `iptables` and the container has no `NET_ADMIN`, so
`docker pause` is used instead — a harder case, because the container's TCP stack
keeps ACKing: no FIN, no RST, and the socket looks perfectly healthy. Only an
application-level probe can detect it.

```
[half-open] {"verdict":"reconnecting","verdictMs":25135,"budgetMs":90000}
```

Nothing in the suite covered the failure mode behind the "SSH hangs until I
restart Orca" reports.

## New: resource accumulation measured on the remote host

6 terminals, then 5 reconnect cycles, counted on the container itself:

```
open:       pts 1->6 (exactly 1/terminal), relay fds 25->30 (exactly 1/terminal)
reconnect:  pts flat at 6, relay procs flat at 1, node procs flat at 3
```

`leakedMasterFdCount` is now **asserted**, not merely recorded. It counts PTY
master fds held by non-relay processes: without `FD_CLOEXEC` a master is inherited
by every later child, so terminal k adds k of them — the triangular signature
measured as 15 across 5 terminals before the fix. #17914 patched the app and
daemon and #17920 shipped the same patch to the relay host, and both are now on
main, so the correct value is 0 and the probe holds it there:

```
baseline    leakedMasterFdCount 0
6 terminals leakedMasterFdCount 0    (holders: only relay.js, n=6)
reconnects  leakedMasterFdCount 0 across all 5 cycles
```

Any growth here means the relay's node-pty rebuild did not take on that host,
which is exactly what a remote-host probe exists to catch — and it is the half of
#17914's claim that no unit test can reach.

## Routing

Both new probes are claimed by `run-ssh-docker-e2e.mjs` (a Docker-gated spec no
runner names self-skips everywhere and still reports green) **and** by the
`ssh-terminal-source` route in `pr-e2e-source-routing.mjs`, so they run when the
relay and SSH code they guard changes rather than only on a scheduled lane.
2026-09-02 15:58:42 -07:00
Neil 074a2366bf test(e2e): pass testInfo to startDockerSshRelayTarget in the freeze repro (#18257)
The spec called startDockerSshRelayTarget() with no argument while the
helper signature is (testInfo: TestInfo) and dereferences
testInfo.workerIndex, so it threw before any Orca code ran and took the
Docker SSH lane red on every PR.

Fixes #16764
2026-09-02 15:14:58 -07:00
Jinjing 4ff96df2b5 Auto e2e tests autofix scheduled ci 1h run 32 20260902T0700 (#18227)
* Fix flaky e2e tests with improved locators and synchronization

Add explicit waits, use more robust element selectors, and simplify test
setup to reduce race conditions. Replace file-based fixtures with
programmatic browser creation, use parent-scoped locators for menu
interactions, and poll for stable state before assertions.

* Add E2E failure triage report for run 33564563164

- Reconciles 14 failed tests against job logs and trace artifacts
- Categorizes failures: 8 product bugs, 2 flaky tests, 4 test updates
- Documents test-maintenance fixes and diagnostic findings
- Files 8 Linear issues with owners and fresh recurrence evidence
- Provides next actions for product owners and repository maintenance

* rm artifact notes

* Refactor browser creation E2E test to use UI interactions

- Click through menu instead of manipulating internal store state
- Use Playwright's locator and toBeVisible() assertion patterns

* Record E2E browser creation pageId before barrier check

Move createdPageId assignment before the barrier arm/fire checks. This
ensures the pageId is recorded unconditionally when tracking is enabled,
allowing tests to distinguish between creations rejected before the host
attempt vs those that failed after creation.

* Remove browser page reclamation assertion from restart test

Simplifies test by removing page ID tracking and poll checking
if pages persist after paired runtime restart.
2026-09-02 14:05:29 -07:00
Jinjing c2fce80289 Fix agent dashboard setting configure (#18245)
* Make agents activity always-on; toggle via bell icon

- Remove optional showAgentsSidebar setting
- Replace sidebar view-toggle with bell-button for activity access
- Agents activity now always accessible in sidebar
- Preserve migration flag for introduction to existing users
- Remove visibility inference utilities

* Simplify sidebar when agents view active: hide workspace options, add to

- Hide workspace options menu and add project button when agents view is
  active, reducing UI clutter in that mode
- Add tooltip to the activity bell button for better discoverability
- Localize sidebar search field text
- Move search and filter toggles to local state in SidebarAgentsList,
  removing unused callbacks from thread list components
- Manage search input focus properly when opening
2026-09-02 13:43:06 -07:00
Jinjing e3de6b2ce8 Add automation runs dashboard with pagination and filtering (#18226)
* Add automation runs dashboard with pagination and filtering

Adds a new Runs view in the Automations page that lets users browse all runs across automations with status/host filtering, search, and pagination support. Includes virtualized table rendering for efficient handling of large run histories and summary cards showing 24h/7d success/failure counts.

* Fix missing dependencies in useCallback hooks and imports

Missing dependencies in useCallback can cause stale closure bugs. This
adds missing state setters to dependency arrays and consolidates type
imports for consistency.

* Use keyset pagination for stable automation runs pages

Pagination now uses createdAt:id boundaries instead of offsets, so new
runs arriving between pages don't shift the window. Maintains backwards
compatibility with legacy offset cursors.

Move pagination to shared module, fix outcome counting for future-dated
runs, and improve hook state tracking on authority re-pairing or target
changes.

* Extract automation run details to top-level page view

Moves run display from detail pane to dedicated page, establishing
three-level navigation (Automations → Runs → Run Details) and simplifying
the detail pane component.

* Fix pagination stability when automation runs share createdAt

- Define a stable total order with createdAt and id tiebreaker to prevent runs tied on createdAt from being dropped when the boundary run is pruned between page requests
- Retain cursor on failed pagination so pages remain retryable
- Update ownerNotice type to AutomationActionNotice

* Extract automations list panel and worktree map logic

Split AutomationsPageSurface into smaller, focused modules for better maintainability and reusability. Move list panel UI rendering to AutomationsPageListPanel component and worktree map selection logic to a standalone utility function.

* Add i18n strings for automation runs dashboard

Adds localized strings for the automation runs dashboard view, including search, filtering by host and status, run counts for 24h/7d windows, and empty state messaging across all supported languages.

* fix missing translation

* fix missing translation
2026-09-02 13:42:14 -07:00
Jinjing ae1dab40d6 Sta 6308 add copy session id option to terminal tab context menu (#18070)
* Move Copy Session ID from tab to terminal pane context menu

- Relocates session ID copy to the exact pane that owns it, not the tab's active pane
- Adds support for durable sleeping agent sessions as fallback for cleared live status
- Generalizes copy-rejection guards to handle any identity type, not just pane IDs
- Updates e2e test to verify pane-specific session ID copying

* Gate session ID liveness by shell foreground state

Once OSC 133;D proves a pane is back at the shell, don't return the
session ID even if a durable record survived the exit. This prevents
treating exited sessions as still active when the user is typing at
the prompt.

* Update hook order parity test for session-ID projection hook

The pane session-ID projection adds a render hook to TerminalPane.
Update the expected hook count from 229 to 230 and the corresponding
SHA256 hash.
2026-09-02 13:05:24 -07:00
Jinjing 61e010079f New agent dashboard (#18222)
* more obvious toggle

* more obvious toggle

* feat(activity): redesign thread rows and add child agent filtering

- Emphasize task title and last activity in row layout over metadata
- Add child agent toggle; hide orchestration workers by default
- Support collapsible groups and ungrouped view mode
- Improve orchestration worker message handling to surface replies
- Add sidebar search and filter controls for agent activity

* periodic checkin

* feat(activity): add "Clear completed" action and performance improvement

- Add "Clear completed" action for activity threads with undo window; clears completed and interrupted rows from view, persists across restart
- Virtualize activity thread list to render only viewport-bounded rows
- Cache activity thread search text to prevent recomputation on every keystroke
- Cache dashboard bucket counts per-worktree for selective invalidation on unrelated changes
- Use useDeferredValue for activity search filtering to keep input responsive
- Make compact mode the default display for activity threads
- Add activity-cleared-at persisted state tracking (per-pane cutoff timestamps)

* improve style

* minor change

* feat(activity): add persisted host and project filters to agents view

Agents scope filters are deliberately separate from workspace-nav filters so a monitoring surface never inherits workspace context silently. Filters survive restarts and always display an active-filter chips row with hidden count, making filtering visible and reversible.

* Graduate Agents view from experimental, refine activity handling

- Agents Dashboard moves from experimental to standard feature with showAgentsSidebar setting controlling visibility
- Add identity-checked cache eviction (dropPersisted IPC) to prevent newer runs from being evicted when UI clears older status, fixing clear-completed safety
- Extract ActivityThreadHoverCardSummary and ActivityThreadListToolbar components for better organization and reusability
- Implement mark-thread-read as separate action from select with clickable bell icon
- Add hasActivityThreadWorkspace helper for checking workspace availability across hosts (SSH/runtime targets)
- Preserve scope filter array identity during hydration for memo optimization
- Track manually-unread turns in auto-ack to prevent re-acknowledgement
- Clean up activity cleared-at cutoffs on pane retirement
- Remove activity-thread-hover-card max-lines lint override (code refactored below threshold)

* Refactor agent cache identity to use timing fields only

- Simplify AgentStatusCacheIdentity: keep only paneKey, receivedAt, stateStartedAt
- This fixes silent no-ops where renderer-enriched fields diverged from main's cache
- Add worktree-jump-navigation for navigating activity to workspaces
- Add manual mark-unread protection separate from auto-ack
- Optimize activity owner resolution with per-build memoization
- Optimize detected worktree lookup with indexed search

* Remove sticky header, add scroll position persistence

Replace the floating sticky header overlay with scroll position memory via
a ref. This preserves the user's scroll location when switching between
threads or remounting the agents list, improving UX without requiring
React state.

* Implement sticky group headers in activity thread list

Keep group headers visible at the top while scrolling when threads are grouped. Headers stick to the viewport while their section is in view, then unstick as the next header approaches.

* add blue flash

* update settings appearnce

* Extracted activity acknowledgement/clearance actions from the oversized UI slice.
  - Removed dead sidebar search/menu props and the unused search ref.
  - Removed the unnecessary sidebar visibility bitmask.
  - Replaced hardcoded sidebar toggle colors with design-system tokens.
  - Removed duplicate “mark all read / clear completed” controls in the sidebar.
  - Preserved manual-unread state correctly across pane retire, transfer, and drop.
  - Made clear-completed cutoffs monotonic so clock skew cannot resurrect old activity.
  - Fixed blank workspace names in hover cards with the existing fallback helper.
  - Added missing localization entries and stabilized hydrated filter array identity.
  - Updated misleading Agents setting copy to describe both sidebar surfaces.

* add onboarding guide for the new agents panel

* Add activity clearance tracking and synced agent view settings

Agent view filters and presentation settings now sync across paired clients.
Preserves per-pane activity clearance cutoffs in persistent state. Improves
activity thread row accessibility with proper ARIA roles, and preserves
terminal host ownership after pane teardown via retained terminal handle.

* rm html

* Graduate Agents from experimental and improve activity visibility

- Migrate `showAgentsSidebar` setting from legacy experimental flags; default new profiles to the agents sidebar
- Replace scoped-thread filtering with visible-thread filtering so bulk actions (mark all read, clear completed) only affect rendered rows
- Rewrite child agent classification as a set of visible pane keys to fix orphan promotion and parent-cycle handling
- Improve activity cleared-at cutoff lifecycle: preserve on row dismissal (pane may still be live) but clear on pane removal
- Add pagehide flush for pending clear-completed evictions so quit/reload cannot replay cleared activity
- Polish agents sidebar: unread count badge, expand button, onboarding intro for migrated/new users
- Extract shared time-ago formatting to a library module
- Fix scroll restoration to defer until content can contain the saved offset
- Improve stable message hold for compact agent rows using state instead of refs
- Add worktree filter-visibility check to distinguish collapsed-but-unfiltered from filtered-hidden

* Graduate Agents from experimental and improve activity visibility

- Remove the deprecated full-page Agents view; fix settings navigation fallback
- Refactor bulk action bindings and separate mark-all-read from visible threads
- Preserve sidebar collapse state across remounts; fix child-agent badge filtering
- Add safety window for scroll-restore and improve worktree host-qualified filtering

* Graduate Agents from experimental and add manual unread tracking

- Move Agents sidebar from experimental settings to standard feature with intro flow
- Add persistent manual unread turn tracking for activity feed
- Consolidate workspace activation through activateAndRevealWorkspace dispatcher
- Improve sidebar view toggle with radio semantics and arrow-key navigation

* Graduate Agents sidebar and separate dashboard experiment

The Agents tab now has its own `showAgentsSidebar` setting (defaults on) independent from the dashboard popout experiment. Activity unread counting is simplified to count all events uniformly without mode-specific filtering. Dashboard visibility is now controlled solely by `experimentalAgentDashboardPopout`, with its own UI in the Experimental settings pane. Migration path updated: only `experimentalActivity=true` graduates to the sidebar; the dashboard experiment remains separate.

* Add agent-session tab support to activity tracking

Build activity event contexts from structured agent-session tabs and
worktree-attributed status entries. When activating a thread, try
agent-session tab activation before falling back to terminal pane.

* • The workspace sidebar tab is now a static Spaces
  label—no grouping-based “Projects” label or hidden
  width-reservation span.

* Show unread count badge and prioritize attention-needing agent threads

Activity group order now surfaces threads needing attention (blocked,
waiting, interrupted) before working/done so they're never buried. The
Agents tab shows an unread count badge while viewing Spaces, since the
open Agents list already highlights unread rows.

Also improves UX text ("Hide Agents" vs "Maybe later"), accessibility
with proper ARIA labels, and handles edge cases: preserves read state
for retained panes on SSH reconnect and handles deleted worktrees
gracefully in navigation.

* Batch agent-status evictions and optimize activity pane rebuilds

- Add dropPersistedStatusEntries batch API; consolidate evictions into one persist
- Implement fallback timeout in clear-completed for unseen toast callbacks
- Project only activity-relevant tabs; memoize terminal tab derivations
- Stabilize activity virtualizer key to prevent unnecessary item measurements

* Remove unread count badge from Agents sidebar tab

Simplify useActivityUnreadCount by removing the enabled parameter and
conditional logic, as the badge is no longer displayed in the UI.

* Deduplicate activity unread counts across source overlaps

Live pane status is the primary source; retained and migration entries
serve as fallback caches that may briefly overlap it during lifecycle
transitions. Count each pane only once by tracking seen keys, prioritizing
the live status as the canonical source.

Also fix monitoring state display: it's a distinct agent state, not a
tool-running row state, so exclude it from tool preview checks.

* Update activity pane tests to remove unread badge assertions

- Remove ActivityPaneVisibility type and readActivityPaneVisibility() helper
- Update agentsSidebarButton selector to match badge-less state
- Simplify assertions to check pane focus instead of visibility isolation
- Remove test for unread badge acknowledgement flow

* Fix activity pane workspace resolution and localization handling

- Thread defaultHostId through activity operations for correct host resolution
- Add language-aware caching for standalone terminal names with cache invalidation
- Fix scroll restoration bounds calculation for tall viewports
- Add focus management to sidebar radio group keyboard navigation
- Refresh localized sidebar content on language changes
- Preserve activity state across heartbeats to prevent history loss
- Improve host-id strictness in worktree jump navigation

* Preserve activity view when settings fetch fails

A failed window.api.settings.get() leaves settings null, which was
incorrectly treated as opt-out. Add the missing null check so the
activity-view gate only applies when settings are available.

Includes tests for this scenario and related edge cases in keyboard
navigation, worktree jumping, and session state handling.
2026-09-02 11:00:24 -07:00
Jinjing 6062edf296 test: simplify remote pane link routing to server-hosted placement (#18219)
Remote-pane links are now explicitly server-hosted regardless of generic
client-hosted preference. Remove client-hosted placement verification,
placement-switching test acts, and related type definitions. Focus the
test on verifying the core invariant: links stay server-hosted on their
owning runtime.
2026-09-02 10:39:46 -07:00
Neil f737f3499f fix(relay): stream an oversized fs.listFiles reply instead of refusing it (#17954)
Opening Orca's own checkout over SSH cannot list its files in one response frame.
22,617 tracked paths average 58 characters, so the 20,001-row page the client asks
for serializes to 1,223,415 bytes — past `DISPATCHER_CONTROL_QUEUE_MAX_BYTES`, so
`sendResponse` demotes it to the `legacy-response` lane, where an unrelated
producer backlog can refuse it as an opaque `ResponseOverCapacity`. Break-even is
around 49 characters of average path; any `packages/<name>/src/...` monorepo is
over the line.

Picking a ceiling to refuse at does not fix that, it just moves where it shows up
and refuses listings that would have been delivered. `__streamResponse` already
exists for exactly this on the git methods, and it is its own negotiation in both
directions: an old client never sends it and gets the plain array on the
legacy-response lane as before, and an old relay ignores it and answers plainly,
which the client detects by the sentinel marker being absent. So fs.listFiles opts
into it — no new method, no new opcode, nothing to advertise — and the size of a
listing stops being a correctness question.

The response-stream registry becomes one per relay, shared by FsHandler and
GitHandler. A second registry is not an option and the header of
git-response-stream.ts says why: a client keys reassembly on `streamId` alone, so
two would hand out the same id and cross-feed chunks, and only the handler that
registers `git.responseAck` can credit the window a pump parks on.

Also declares `maxResults` on the runtime-RPC `files.listAll` and forwards it.
The mechanism "the client names its cap, so a full page reads as truncation" was
wired only on the Electron IPC hop; web and mobile were saved incidentally by
`remoteFileContentBudget` defaulting the cap inside `listRuntimeFiles`. A new
optional field is additive in both directions (wire rule 1).

The new Docker-gated spec is claimed by run-ssh-docker-e2e.mjs. The sharded e2e
lanes set no ORCA_E2E_SSH_DOCKER, so a Docker-gated spec that no runner names
self-skips everywhere and still reports green — pr-e2e-gate-contract enforces that.

Closes #12547
2026-09-02 05:36:54 -07:00
Jinjing 32e4c6be4a Auto-focus editor when opening new markdown file (#18071)
* feat(markdown-preview): autofocus editor when opening new markdown file

Users should be able to start typing immediately after creating a markdown file without an extra click.

* add e2e tests

* add e2e tests
2026-09-02 00:22:59 -07:00
Neil 0dbe9d0504 test(ssh): dockerized relay fault injection with verdict assertions (#18017)
* test(ssh): add a dockerized SSH fault-injection lane with four fault shapes

The existing SSH reconnect specs all reconnect by calling ssh.disconnect() then
ssh.connect() - a clean cycle the client knows is coming. Nothing covered the
faults the reconnect machinery exists for.

Four shapes, each documented with why it is not the others: killing sshd's
per-connection forks (transport dies, relay survives), `docker pause` (silence
with TCP still established), SIGKILLing every relay.js (the only fault where
`exited` is the correct verdict), and a 48MB flood with nobody attached.

The relay-kill case is the one that makes the rest meaningful: every other case
asserts the session survived, which only means something if a genuinely dead
session is distinguishable. It is the only case where replacing the pane is
correct, so it pins the boundary in
docs/reference/ssh-execution-boundary.md rather than just testing reconnection.

The `docker pause` case pins the other side of that boundary: after 30s of
silence from a healthy host the pane keeps its PTY and its scrollback, because
loss of contact is never evidence of death.

No network-blackhole fault: reconnecting the fixture does not restore its
published port mapping, so that fault is not reversible on this container and
would strand the worker it ran on.

* test(ssh): fixme the flood case pending #18018

It fails in CI on its first real run: the pane keeps its PTY and repaints,
but a command run after the flood produces no output within the poll budget.
Same shape as #18018 and not caused by this spec. The three verdict
assertions around it stay enforced.
2026-09-01 23:35:47 -07:00
Jinjing 0352c239c2 Add Copy Session ID menu item to terminal tabs (#18039)
* Add Copy Session ID menu item to terminal tabs

Adds a menu item to copy the active pane's agent session ID when available.
The item only appears when the session is still live and has reported an ID.

* Add Copy Session ID i18n strings and e2e test

- Add localized strings for Session ID context menu item
- Add e2e test coverage for copying session ID from terminal tabs
- Fix dev build permissions when copying private Electron app bundles

* Drop the Electron dev-bundle fix from this branch

It landed on main as 519af49a58, which restores write permission inside
copyPrivateTree itself rather than at the dev runner's call site, so every
caller of the private-copy contract is covered and not just this one. That
commit also fixes the test that should have caught the crash: the wrapper ran
with stdio: 'ignore', so a hard failure presented as a bare timeout.

This branch predated that commit and carried a narrower duplicate, mixed into
an i18n/e2e commit where it did not belong.

* refactor: use dedicated i18n keys for copy session ID toasts

Replace auto-generated translation keys with specific, dedicated keys
for copy session ID success and error messages. This improves
maintainability and makes the strings easier to translate across all
supported languages.
2026-09-01 19:20:48 -07:00
Neil f8a3f2c7c0 test(e2e): do not treat a destroyed renderer as a relaunched runtime (#17785)
* test(e2e): do not treat a destroyed renderer as a relaunched runtime

waitForRelaunchedRuntime polled refreshAuthorityRuntimeId with
expect.not.stringMatching(previousId). Playwright treats null as a
non-match, so an Execution-context-destroyed miss ended the wait as if
the client had already reconnected. Poll until a non-null id that
differs from the pre-restart process.

* test(e2e): wrap cookie-survival restart evaluates as pending misses

The cookie spec still opened a post-restart page with a raw evaluate
poll. A recycled renderer then timed out as "never materialized". Use
the shared fixture helpers so destroyed-context is a miss, not a fail.

* test(e2e): leave cookie-survival on its own wait for this PR

The relaunch-wait fix made the cookie spec's post-restart echo render
time out in CI. Keep that spec out of this change so the destroyed-
context wait can land on the helpers restart-survival actually uses.
2026-08-31 19:47:27 -07:00
Neil a5796ec8eb refactor(runtime): split OrcaRuntimeService and compatibility tests (#17605)
* refactor(runtime): split OrcaRuntimeService into focused modules

* test(runtime): cover admission tiers and strict worktree reconciliation

* fix(runtime): preserve owner and structured session visibility

* fix(runtime): port post-extraction compatibility fixes

* fix(runtime): preserve skill-share cancellation barrier

* test(runtime): update identity inventory after extraction

* fix(runtime): preserve hook transport environment cleanup

* fix(runtime): consolidate idle probe imports

* test(runtime): retire split file process allowlist entry

* fix(runtime): route child process types through shared boundary

* test(runtime): preserve worktree host metadata precedence

* fix(runtime): update extracted test seams

* fix(runtime): gate the split's ts-nocheck set and restore the stop-confirmed contract

Audit follow-ups for the OrcaRuntimeService split:

- Freeze the 171 @ts-nocheck files behind a ratchet so no new file can disable
  type checking. The split's linear mixin chain cannot express forward
  references yet, so the existing suppressions are grandfathered; the baseline
  may only shrink.
- Drop the stray @ts-nocheck at the end of orca-runtime-get-status.ts. It sat
  after the first statement, where TypeScript ignores it, so the module was
  already checked.
- Restore `retireRejectedPty(ptyId, stopConfirmed: boolean)` as a required
  argument. The split widened it to optional and patched the resulting error
  with `stopConfirmed === true`; an omitted argument would have silently taken
  the unverified-stop path instead of failing to compile.
- Guard that every orca-runtime-tests fragment is imported by the compatibility
  entrypoint. The fragments are .spec.ts, which no Vitest include glob matches,
  so one left out of the list would silently stop running.

* fix(runtime): restore four behaviors the OrcaRuntimeService split dropped

Audit findings against the refactor's true base (ad5ba2572e):

- retirePtyAgentLaunchAuthority collected pane keys after deleting the
  restored-authority receipt instead of before it. collectPaneKeysForPty reads
  that receipt, so a receipt-only pane lost its key and never had its agent-hook
  compatibility authority retired. on-pty-exit.ts already carried a comment
  naming this exact invariant.
- The PTY-exit path kept orchestrationMailboxNotifications.retirePty but lost
  the loop that schedules a debounced mail-pointer repoint for the dead pty's
  terminal handle and any run bound to its panes. Restores the schedule call
  count to 7, matching base.
- subscribeToPtyExit lost isPtyKnownExited's leaf fallback and its
  post-registration lifecycle-generation recheck. leavesByPtyId is rebuilt from
  the renderer graph independently of ptysById, so a leaf can outlive its pty
  record; without the fallback a caller waiting on an already-dead pty never
  gets released.
- The chain root declared `[key: string]: unknown`, which base had nowhere. It
  leaked through the exported runtime type into every consumer, so any misspelled
  member access typechecked as unknown instead of erroring, and it accounted for
  957 of the suppressed errors. Removing it costs zero type errors.

* fix(runtime): restore escalation prose and unscoped automation publication

Two more behaviors the split dropped, each with a regression test that fails
against the pre-fix code:

- The worker-exit escalation stopped deriving its title through
  buildOrchestrationTaskDisplayMetadata and inlined `task.spec` instead. That
  ignored an explicit task_title, dropped the single-line normalization and the
  80-character bound, and turned the no-spec case into a quoted, duplicated id.
  A multi-paragraph spec landed verbatim in the coordinator's banner. The
  existing 11 tests all use short single-line specs, where the derived title and
  the raw spec are identical, so none of them could see it.
  Also reverts an added `if (!handle) return` guard: the dispatch lookup is
  deliberately keyed on the pane as well, because a reminted handle no longer
  matches the row while the pane identity outlives the remint.
- updateAutomation stopped going through automationChangePublications and
  published `source` unconditionally while gating the fallback on a non-null
  destination. A destination the store can no longer name then published only
  the stale source, so subscribers scoped elsewhere kept rendering a row that
  had left them — the exact case the helper documents. The helper had been left
  with zero callers; all three sites use it again.

* fix(skills): stop swallowing lookup errors and hard-erroring on non-ssh hosts

Follow-ups from auditing the skill install path against the refactor's base:

- resolveWorktree wrapped showManagedWorktree in `.catch(() => null)`, so a
  transient git or IO failure surfaced to the user as
  skill-install-workspace-not-found with the real cause discarded. Errors
  propagate again; a genuine id mismatch still returns null.
- resolveSkillSshTarget threw skill-install-workspace-host-unavailable when the
  execution host was neither local nor ssh, on both the repo and folder
  branches. Base gated these on connectionId, so a runtime-owned repo simply
  was not an SSH install and fell through to the local path. Both return null
  again, and the error code the split invented is now unreferenced.
- listManagedSkillInstalls awaited the receipt walk and the worktree resolve in
  sequence. They are independent and either can hit disk, WSL, or an SSH scan,
  so Promise.all is restored.

Deliberately unchanged: resolving the worktree through listResolvedWorktrees
rather than showManagedWorktree, which disambiguates a worktree id colliding
across hosts and is covered by its own test, and the SSH-folder
skill-install-ssh-dispatch-required throw, which matches the repo branch.

* fix(runtime): merge duplicate worktree-logic imports

The #17448 port added a third import from ../ipc/worktree-logic, which the
code-quality oxlint config rejects under --deny-warnings. Plain oxlint does not
flag it, so it only surfaced in CI's static analysis job.

* ci: run the ts-nocheck ratchet in PR checks

pr-workflow-lint-parity requires every leaf command in `pnpm lint` to have a
matching step in pr.yml. The ratchet was wired into lint but not the workflow,
so PR CI would not have enforced it.

* Merge remote-tracking branch 'origin/main' and retry the paired-host launch evaluate

main advanced 9 commits; none touch the orca-runtime.ts this branch splits, so
nothing needed porting.

CI failed twice on `Execution context was destroyed` thrown from
headless-paired-runtime-host's first `evaluate` after launch — a different spec
each run, which is the signature of the flake #17780 describes rather than a
regression. That commit added retryTransientMainEvaluate and adopted it in five
helpers but not this call site, even though its docblock names exactly this
case: the first evaluate after electron.launch() resolves, before the app is
ready. Wrapped it the same way.
2026-08-31 19:34:55 -07:00
Jinwoo Hong 1efd4e1a97 test(e2e): seed source control diff before opening panel (#17784) 2026-08-31 22:22:17 -04:00
Neil f116d2ca2a test(ci): retry Windows teardown EPERM and restart evaluate misses (#17780)
Restart-survival polls treated a recycled renderer as a hard failure.
Wrap those evaluates so "Execution context was destroyed" is a pending
miss. Windows package-lane teardowns after a force-kill used rmSync
with force:true only, which does not absorb EPERM; put them on the
shared maxRetries:8 policy.
2026-08-31 18:53:01 -07:00
Neil eff317939a fix(terminal): mount one surface per workspace id in the workbench (STA-4846) (#17432)
* fix(terminal): mount one surface per workspace id in the workbench (STA-4846)

* test(terminal): pin the workbench projection against under-selecting

Losing a surface unmounts live terminals, which is worse than the
duplicate mount STA-4846 fixes, so cover every catalog shape that reaches
the workbench: local-only rows that name no host, an unqualified row
colliding with a host-qualified one, two SSH hosts on one id, folder rows
across three hosts, folder ids alongside git worktree ids, and a
whole-catalog assertion that the emitted id set equals the distinct input
id set. Also pin the `useAllWorktrees` -> `useWorktreeMap` swap: both read
the same WeakMap-cached snapshot, so the zustand compare is unchanged.

Harden the folder tie-break to require the row to name its own host.
`getCatalogOwnerHostId` defaults an unstamped row to `local`, which would
let a row that never named a host win the `local` tie and mount another
host's path; it now keeps first-wins instead of guessing.

* fix(terminal): surface the unresolvable folder-surface collision

When two hosts publish the same folder-workspace id and the active workspace's
host cannot be resolved, the projection drops one row's folderPath first-wins.
That path is the PTY cwd for any tab without a startupCwd, so the drop was
silent. Warn on it, and pin the two tie-break branches the unit tests missed:
a colliding row that is not the active workspace, and the same collision with
the rows in swapped order (a host reconnect re-appends its rows, flipping which
row is first mid-session).

* test(e2e): ride out Playwright's spurious main-process evaluate rejection

`e2e / changed e2e specs` failed on `pr11346-selected-runtime-add.spec.ts`
with "Execution context was destroyed, most likely because of a navigation"
from the paired client's first `app.evaluate` — the isolated-HOME assert that
runs one millisecond after `electron.launch()` resolves, which is before the
app is `ready`. Nothing navigates there: Playwright raises that message for
any main-process CDP failure that is neither a JS error nor a closed session,
and `ElectronApplication.evaluate` is unreliable on Electron 27+
(microsoft/playwright#33737). Reproduced locally, and a plain re-run of the
same commit went green.

Extract the retry `installTerminalPtyWriteSpy` already carried for this exact
message into `retryTransientMainEvaluate`, and use it for the launch-time home
read in all three launchers. The read is idempotent and a real boundary escape
still throws on the first successful read.

Also forward the paired client's process logs before the assert instead of
after: this failure reached CI with none of the client's own output, because
forwarding had not started yet.

* test(e2e): wait on the owning group before asserting a Cmd-J browser tab is active

`changed e2e specs` then failed at the remote browser-page step: the store poll
had already seen `activeBrowserTabId` land on the mirrored workspace, but
`[data-tab-id=...][data-active="true"]` never appeared. `data-active` on a
`BrowserTab` is the strip's active tab, which comes from the owning group's
`activeTabId` — not from `activeBrowserTabId` — so the DOM assert was racing an
activation the poll never waited for. The simulator rows in the same spec
already poll the group; the two browser-page rows did not.

Poll the same triple for them, so a genuinely stuck group fails with the ids it
ended on instead of a bare "element(s) not found".
2026-08-31 18:45:52 -07:00
Neil c558d7e083 Activate terminal splits before inherited CWD resolution (#17601)
* perf(terminal): activate splits before cwd resolution

* test(terminal): prove split focus before cwd publish

* fix(terminal): release stale split cwd fence

* test(terminal): add visible split activation latency benchmark

* docs(reliability): clarify split benchmark provenance

* fix: preserve deferred split handoffs across remounts

* fix: fence late deferred split closes

* docs(reliability): record exact split benchmark runs

* test(reliability): fail benchmark on artifact write errors

* test(reliability): attribute split activation phases

* docs(reliability): record schema-v2 split benchmark

* refactor(terminal): collapse duplicated split-handoff and write-queue paths

- Drop the discardDeferredSplitPaneHandoff alias for its identical clear twin.
- Fold the deferred-cwd resolve/reject settle handlers into one applier.
- Extract settlePaneCwdDeferredSpawn for the repeated read-clear-write pattern.
- Share one head-index FIFO primitive between the ordinary and reply queues.

* fix(terminal): stop retaining a promise reaction per acknowledged write

Racing every accepted write against one queue-lifetime cancel promise kept a
reaction record alive until that promise settled: 200k acknowledged writes
retained 88.6MB, now 0.1MB. Give each in-flight write its own cancel, and
split the shared FIFO primitive into its own module.

Also sanitize the split-latency benchmark report at its single serialization
point so shared artifacts no longer carry the machine-local repo path or
unbounded cleanup error text.

* fix(terminal): settle deferred split input when the spawn is abandoned

An abandoned deferred spawn returns before transport.connect(), so nothing
drained the pre-connect buffer: sendInputAccepted's promise never settled and
a paste into that pane hung forever. Clear the buffer on the abandon fence.

Also re-derive the pre-connect retention cap from the clipboard-paste ceiling
rather than the 16MB single-write ceiling; it is held twice per pane across up
to 64 deferred splits, so 5.59M code units guarded the wrong thing.

* fix(terminal): release the deferred cwd fence on a rejected reattach

A daemon createOrAttach can turn an apparent fresh spawn into a reattach; when
that reattach is refused the spawn ends with deferredSplitSpawn/pendingCwd
still set, permanently arming the pre-bind detach refusal. The release no-ops
when a PTY did bind, so it only fires where the fence would otherwise leak.

The stale-generation return above is deliberately left alone: a newer connect
already owns the pane there, and the fence is not generation-scoped.
2026-08-31 16:45:36 -07:00
Jinwoo Hong bbbb59e18c test: cover quick commands, catalog links, and long discard dialogs (#17489) 2026-08-31 14:56:15 -04:00
Jinjing d5d3c4898a perf(diff): defer large diffs until user loads them (#17521)
* perf(diff): defer large diffs until user loads them

Rendering very large diffs would freeze the UI. Diffs exceeding
MAX_AUTOMATIC_DIFF_CHANGED_LINES now show a prompt allowing users
to load them on demand instead of automatically rendering.

* perf(diff): defer large diffs until user loads them

Diffs with >10,000 changed lines are now deferred and only rendered when
the user explicitly clicks "Load diff" in a prompt. This improves initial
render performance for large file changes while maintaining full access
when needed.

* perf(diff): defer large diffs until user loads them

Prevents UI freeze when opening files with very large diffs by
deferring render until the user explicitly loads them.

* fix(diff-view): defer loading large untracked files and refactor fallbac

Split on-demand load decision logic to distinguish tracked vs untracked files — large untracked files now properly defer loading while untracked images remain automatic. Extract fallback height computation into a dedicated function to centralize the logic for render-limited and in-flight-loading states, reducing code duplication and clarifying when to use bounded fallback heights.

* fix(diff-view): defer loading large SVG files

SVG renders as source text in the diff view rather than a preview, so should defer like other text files. Also fix Windows e2e test cleanup by using post-Electron shutdown.
2026-08-31 09:20:12 -07:00
NeilandBrennan Benson fbe94ceff6 fix: close readiness gaps found by merged-change audit (#17159)
* fix(ssh): fence stale kills and retired pane replay

* fix(ssh): support cancellable interactive authentication

* fix(ssh): await remote catalog before snapshot adoption

* fix(pty): contain Windows ConPTY input failures

* fix(power): avoid redundant macOS display blocking

* perf(editor): narrow markdown override subscriptions

* fix(quick-open): close directory handles after reads

* refactor(linux): remove unused proc socket scanner

* fix(usage): apply flat Sonnet 4.6 pricing

* ci: prime Node next native test cache

* docs(skills): resolve snapshot cleanup data path

* fix(ssh): recover install locks after host reboot

* test(ssh): recognize boot-aware install locks

* test(ssh): prove previous-boot lock recovery live

* test(wire): pin pre-metadata release coverage

* fix(terminal): preserve remote tab ownership through recovery races

* test(runtime): fence replaced terminal handles in agent guard

* fix(ssh): preserve remote snapshot authority across polls

* fix(pty): contain late ConPTY output EPIPE

* test(pty): register Windows exit watcher before kill

* fix: close SSH and tab readiness race gaps

* fix(tabs): retain headless order and placeholder titles

* fix(build): avoid parallel electron-vite config race

* test(windows): avoid MSYS temp path rewriting

* test(windows): avoid killing exited PTY

* fix(pty): avoid late ConPTY input teardown race

* fix(terminal): sync reconnect error ownership after commit

* fix(runtime): use canonical worktree identity comparison

* test(ssh): assert complete cold-hydration baseline

* test(windows): invoke quoted retention fixture via PowerShell

* test(windows): read ConPTY grid through mode con

* fix(terminal): publish PTY replacements atomically

* fix(terminal): infer stale identity on reattach

* fix(terminal): fence stale pane PTY callbacks

* fix(terminal): fence stale pane binds after rebind

* fix(terminal): reject stale pane transport callbacks

* fix(terminal): fence mirrored reattach spawn callbacks

* fix(terminal): replace stale pane PTYs on remount

* fix(ci): size the Windows launcher-compile test budget from measurement

`native-smoke (windows-latest)` fails ~4.5% of runs on
`preserves a multiline argument through the compiled remote launcher`
with "Test timed out in 15000ms" — on unrelated PRs, for reasons that
have nothing to do with them. Across 176 sampled attempts it is the only
red that job produced, and it hit seven different PRs in two days:
#16900, #16904, #16915, #16955 (twice), #16979, #17014, #17085.

The test is six process creations: powershell.exe forks csc.exe, then
the freshly compiled orca.exe forks node.exe, twice. Hosted Windows
runners periodically slow process creation down, and this test amplifies
that far harder than anything else in the job. Comparing the 80 attempts
where it ran under 3s against the 12 where it ran over 12s, its own
median goes 2198ms -> 15917ms (7.2x) while the same file's
powershell-only test moves 556 -> 686ms (1.2x), the cmd.exe and Git Bash
process tests in the neighbouring file move 1.4x, and the other 35 files
put together move 1.5x.

Measured across those 176 attempts: 1881ms to 35438ms, p50 4264ms,
correlation +0.881 with the job's total Vitest duration. 8 of 176 (4.5%)
exceeded the 15s cap; 2 of 176 (1.1%) also exceeded the shared 30s
testTimeout, so deleting the override and inheriting the config is not
enough on its own. 60s clears all 176 with 1.7x headroom on the worst.

This is slow, not hung. Every body here is synchronous spawnSync, so
Vitest cannot interrupt one — the timer fires only after the body
returns and the reported duration is real elapsed time. That is why a
failure reads `× ... 22464ms` under `Test timed out in 15000ms`. The
work finished; the stopwatch was short. Seven reruns at one identical
head measured 2053 / 4680 / 5551 / 8732 / 13506 / 14868 / 21937ms — the
last of those would have been red on code that had not changed.

The 15s came from #8897, which raised this test off Vitest's built-in 5s
default because the job then ran bare `pnpm vitest run`. #8909 landed
3h27m later and pointed the job at config/vitest.config.ts, which is the
real fix for that. The constant stayed behind and has been the binding
budget ever since.

* fix(terminal): fence stale remount reattach ownership

* fix(terminal): reconcile mounted pane identity after replacement

* fix(terminal): fence stale reattach fallback ownership

* fix(terminal): fence deferred SSH reattach ownership

* fix(terminal): fence stale split pane ownership callbacks

* fix(terminal): keep stale spawns from consuming startup

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-31 08:17:40 -07:00
Jinwoo Hong 46fa1a98d0 fix(browser): show reload loading feedback (#17635) 2026-08-31 02:46:17 -04:00
Jinjing 3060cf73b9 fix(tasks): restore scroll position when reopening GitHub item details (#17524)
Track scroll-restore generation to invalidate stale callbacks that were
resetting the scroll position to 0 when reopening a detail page. Prevent
restoration while a detail page is open.

Update automation test to use runtime.call RPC instead of removed preload
CRUD method. Hide browser import hint in E2E profile to prevent overlay from
intercepting test setup clicks.
2026-08-30 20:28:35 -07:00
Jinjing 0d8785b916 Prevent stale search commits from timer race conditions (#17495)
Validate scheduled values match current state before executing idle
timeout callbacks. Use useLayoutEffect to synchronously update refs,
preventing outdated searches when rapid keystrokes overwrite timers.
2026-08-30 19:07:01 -07:00
Brennan Benson f7d8d7f77a test(e2e): make the cold-hydration spec verify its own captured snapshot (#17031)
`adds no tab when the host workspace snapshot stalls across a relaunch` replays
the bytes it reads off the relay, but only ever waited for the snapshot FILE to
exist -- never for it to carry the tabs the test had just seeded. A capture that
missed the baseline produced a failure that reads as a product regression and is
not one: an empty `session.tabsByWorktreePath` places nothing, so it reports
nothing unplaced, so `remote-workspace-snapshot-apply.ts` marks the target
hydrated and `hydrateTabsSession` replaces the worktree's tabs with none. That is
exactly the observed `baseline=3 duringStall=3 afterHydration=0`, and it is
correct behaviour for a host snapshot that genuinely holds no tabs.

Assert the precondition where it belongs -- on the capture, before the relaunch
that consumes it -- so an empty or unparseable fixture names itself instead of
surfacing later as a tab count the product appears to have lost.

No assertion is weakened: `afterHydration` still has to equal the baseline, and
no retry, sleep, or timeout was added.
2026-08-30 15:56:21 -07:00
Neil e84042572c Upgrade xterm to 6.1.0-beta.303 and generate addon patches
* Upgrade xterm to 6.1.0-beta.303 and generate the addon patches

Takes the current xterm beta line: xterm 287 -> 303, addon-webgl 286 -> 299,
addon-serialize 287 -> 300, headless 302, the remaining addons -> 300, and the
same set on mobile. All four packages stamp upstream commit d3e32b3.

The reasons are upstream #6042/#6043/#6055 (a shared glyph atlas no longer
garbles sibling panes on a page merge, clear, or sampler-budget overflow) and
Note that core 303 is not image-addon-only over 302: it carries the buffer perf
work, including the new BufferLineStringCache.

addon-webgl and addon-serialize move into the patch generator
--------------------------------------------------------------
Both were hand-edited minified bundles, which is what the Known Gaps section of
docs/reference/xterm-patch-regeneration.md described. Both reproduce byte for
byte from the pinned commit, so they are now manifest entries generated from a
source patch like @xterm/xterm already was. Their sourcemaps now move with their
bundles; before this they shipped maps whose offsets did not match the code
beside them.

The webgl patch shrinks from a 1.06 MB hand-edited bundle to a 6.6 KB source
patch, because upstream took the invalidation half Orca had backported. What is
left is only what upstream still lacks: the fragment-shader else branch for a
v_texpage past the sampler budget, the clearTexture guard that no-ops once a
merged page holds index 0, spending the merge retry budget before beginFrame
latches the version it saw, and Orca's font-weight probe.

The serialize source patch is byte-for-byte the same fixes as before; upstream
changed nothing in that addon between 287 and 300.

Generator fixes, each of which failed silently
----------------------------------------------
- `--relative` was appended after the `--` separator in CHECKOUT_DIFF_FLAGS, so
  git read it as a pathspec and kept repo-root-relative paths, dropping every
  source hunk from an addon's patch.
- `git apply` run from a package subdirectory still resolves patch paths from
  the repo root, skips every hunk and exits 0. It now runs from the root with
  `--directory=<packageDir>`, and a source patch that leaves the checkout
  unchanged is a hard failure rather than an empty patch.
- An addon's own `tsgo -p .` has empty files/include and only project
  references, so it emits nothing and the addon webpack then fails on a missing
  ./out/. The root build now runs first.
- versionStampFile is optional; publish.js stamps an addon's package.json, which
  overlayBuildOutput never patches.
- On a version bump the lockfile has no entry under the new key yet, so --write
  reports the gap instead of aborting mid-run. --check still fails on it.

Adding the two addons pushed the generator and the Electron packaging contract
test over max-lines, so the patch-text helpers move to xterm-patch-text.mjs
(pure text: no checkout, no build) and the vendored-xterm assertions move out of
the packaging contract into xterm-webgl-runtime-contract.test.mjs.

Tests
-----
Four tests asserted upstream bugs that are now fixed, not Orca behaviour:

- xterm-user-scrolling-contract pinned headless and core by version string.
  Upstream bumps each package only when its own output changes, so headless 302
  and core 303 are the same source. It now asserts they share a commit.
- Five CSI 3 J assertions expected a reader stranded at the top after an erase.
  Upstream #6081 clears isUserScrolling there, so the erase releases them to the
  bottom instead. Orca's pin still lands them correctly, because its parser
  handler observes the erase before xterm's own handler runs.
- The IME transaction test hard-coded the xterm version; it now reads the
  installed package, since the point is that bundle, map and version agree.
- The Electron runtime contract asserted Orca's old clearModelGeneration. Shared
  atlas invalidation is upstream's now, so it asserts pageLayoutVersion on the
  resolved dependency, plus the Orca-only hunks on the patch.

Verified: 66,008 unit tests, mobile's 3,863, the four WebGL atlas e2e specs, and
`regenerate-xterm-patches.mjs --check` in sync on all three packages.

Left alone deliberately: resetAllTerminalWebglAtlases still fans out globally
even though clearTexture now self-heals siblings, and upstream #6068
(WebglAddon.dispose leaks the GL context) is still open.

* Drop the two unused WebGL atlas fan-out exports

resetAllTerminalWebglAtlases and presentAllTerminalPanesWithoutAtlasClear have
no callers, and had none at cadfc55102 either — the last call site went in
#6949, which routed reveal recovery through
resetAndRefreshAllTerminalWebglAtlases instead. Only a comment in
pane-manager.ts still named the first one; it now points at the live entry
point. scheduleRevealPresent leaves the registry's structural type with them,
though the manager method stays: terminal-visibility-resume.ts calls it
directly.

This is dead-code removal, not a consequence of the xterm bump. The live
recovery path is unchanged.

resetAndRefreshAllTerminalWebglAtlases stays, and so does the reveal-time
escalation in pane-reveal-repaint.ts. Upstream 299 does make a pane-local
clearTexture bump pageLayoutVersion so siblings rebuild on their next frame,
which is the bug the escalation was written for, but I could not demonstrate
that removing it is safe: with the escalation removed,
floating-workspace-shared-glyph-atlas.spec.ts still passed headful, and it also
passed with upstream's mechanism deliberately disabled (pageLayoutVersion
pinned to 0 in the installed bundle, verified present in the built renderer).
A guard that passes with the fix disabled cannot license removing the
workaround, so the escalation stays until that spec can reproduce the garbling.

Verified: pane-manager and terminal-pane suites (4,713 tests), typecheck, the
headful shared-atlas spec, and the three headless WebGL specs.

* Give the shared glyph atlas spec a trigger that can fail

floating-workspace-shared-glyph-atlas.spec.ts guards the corruption where one
terminal wiping the module-global atlas leaves sibling terminals drawing from
stale texture coordinates. Both of its tests drive that through a floating
panel reveal, and Orca's reveal paths escalate to a registry-wide atlas reset
that repaints every pane — so the recovery under test heals the damage before
the assertion runs, and the tests pass whether or not xterm propagates the
invalidation at all.

The new test clears the shared atlas straight through the floating manager with
the panel closed, so nothing else repaints the workspace terminal, then repaints
it with terminal.refresh(). That is the load-bearing detail: _updateModel skips
cells whose content is unchanged, so the refresh reuses vertices baked against
the pages that were just wiped, which is exactly the state the fix has to
recover from.

Verified as a discriminator rather than assumed. Pinning ITextureAtlas's
pageLayoutVersion getter to 0 in the installed bundle, which disables the
per-renderer invalidation upstream added in addon-webgl 0.20.0-beta.299, and
confirming that reached the built renderer:

  fix intact:   siblingClearIntact=true   1 passed
  fix disabled: siblingClearIntact=false  1 failed

The failure renders the workspace terminal completely blank — stale coordinates
into a wiped atlas sample nothing. The two reveal tests pass unchanged in both
configurations, which is the gap this closes.

* Compare shared-atlas screenshots with tolerance instead of byte equality

Byte equality fails on sub-pixel antialiasing noise that leaves every glyph
legible, so the headful spec flaked under xterm 303. Reuse the existing
compareTerminalScreenshots helper: real stale-model corruption blanks the
terminal at ~3% of pixels, twice the helper's 1.5% threshold, so the looser
oracle keeps its teeth. Log the ratio so failures are diagnosable.

* fix(xterm): cancel empty deferred IME compositions

* test(xterm): strengthen runtime patch contracts
2026-08-30 15:14:49 -07:00
Neil fd52e942bd fix(tasks): keep the remembered GitHub scroll offset instead of clobbering it (STA-5949) (#17433) 2026-08-30 11:49:12 -07:00
Neil 6677ae4e5e test: correct 8 stale specs surfaced by the test-detected-bugs sweep (#17434) 2026-08-30 11:45:21 -07:00
Neil 70df6f0224 fix(terminal): mask the agent composer's dim placeholder during a preedit (#17377)
Split out of #17170, which now carries only the xterm composition-overlay work.

Codex and Claude draw an all-dim, full-row ghost placeholder. The opaque preedit
overlay reproduces the committed row tail it covers, so without this the ghost is
repeated to the right of the composing syllable instead of staying masked. The
binding keys off the `.xterm-composition-remainder` class that #17170 adds and
hides it through CSS while a composition owns a structurally verified placeholder
row — bold prompt glyph plus a dimmed model footer below a blank gap for Codex, a
frame line above the prompt for Claude. Arbitrary dim output, shell lookalikes,
and any row carrying typed text keep their tail visible.

readTerminalCursorLineContext moves from src/main/daemon to src/shared because the
renderer now needs the same reader the daemon uses; the move is import-only.

Depends on #17170.
2026-08-30 03:12:22 -07:00
Neil 7f822a73e3 fix(terminal): render the IME caret and give the candidate anchor one owner (#17170)
* fix(terminal): render IME caret without placeholder overlap

* fix(terminal): preserve dim mid-line composition tails

* fix(terminal): keep IME caret visible at row edge

* fix(terminal): harden IME overlay lifecycle and layout

* test(terminal): type final-cell layout mock

* fix(terminal): keep final-cell IME anchor on-screen

* fix(terminal): bind IME masking to composer ownership

* fix(terminal): bound IME placeholder session ownership

* fix(terminal): track latest IME placeholder session

* test(terminal): share IME session event fixture

* fix(terminal): keep both writers of the IME candidate anchor in agreement

`textarea.style.left` has two writers: xterm's patched CompositionHelper and
Orca's terminal-ime-candidate-anchor.ts. The anchor module listens on
terminal.element, so within a composition event it writes after xterm's textarea
listener and reverted the final-column clamp the patch had just applied.

Moving the clamp into the anchor module and dropping the patch hunk does not fix
it, and the rendered e2e caught that: CoreBrowserTerminal.ts:444 drives
updateCompositionElements from onRender as well, so xterm re-asserts the textarea
position on every repaint, with no composition event for that module to hear. The
anchor survived only when no render happened to follow — measured as a flake at the
final column, 1561.28px against a 1557px screen edge, the fully unclamped value.

So both writers now compute the same clamp. The patch keeps it, because it is the
writer on the render path and already holds cursorLeft, maxWidth and the preedit
bounds. The anchor module applies the same one, so its composition-event write no
longer reverts the correction in the window before the next render. Both halves are
individually necessary and both are mutation-tested.

Also restores _getRowRemainderText's expression from main: translateToString(true,
x, line.length) and translateToString(false, x, getTrimmedLength()) are the same
call, since upstream does endCol = min(endCol, getTrimmedLength()) under trimRight.

Adds the two missing tests — one installing both anchor writers in a single rig, one
driving a render under an open composition — plus disposal cleanup and clamp-bound
coverage, and moves the Codex/Claude placeholder mask to a follow-up PR.
2026-08-30 02:23:04 -07:00
Neil 7b467bd0a6 ci: gate PRs on a real input method, and prove the lane engaged one (#17365)
* ci: gate PRs on a real input method, and prove the lane engaged one

No job on the PR gate has ever run a real input method. pr.yml and e2e.yml are
ubuntu-latest with CDP `Input.imeSetComposition`, which is a synthetic
composition; the only job that drives ibus-hangul through xdotool is
terminal-ime-e2e.yml, and it is schedule + dispatch only. A PR could turn the
real-IME path red and merge green.

Route IME source to that lane from pr.yml through the existing
pr-e2e-source-routing mechanism, so it runs on IME-touching PRs and nothing
else. The lane stays out of verify.needs — advisory, like `e2e` — because its
reliability is known only from nightly main runs. Deliberately no
continue-on-error: that reports green and hides the signal.

The harness fails open in ways that all look like success: Playwright reports a
skipped test as a pass, so an unset ORCA_E2E_NATIVE_IBUS_HANGUL, a renamed test,
or a session with no engine all exit 0 having exercised nothing. The specs now
append an engagement receipt only after observing real composition events, and
the runner requires one per expected test before the lane may report success.

Also drop the native spec from changed-e2e: it was already routed there by its
own filename, where it self-skips for want of an ibus session and reported that
skip as coverage.

* ci: let the real-IME step report even when the synthetic step failed
2026-08-30 01:58:32 -07:00
Neil 2259e06ff6 fix(tests): match showInactive() in paired-client-window-reveal spec (#17362)
PR #17347 switched the reveal helper from window.show() to
window.showInactive() and updated the thrown message, but left the
unit test's regex/title matching the old show() wording — failing
deterministically in CI (which builds against current main) while
passing on any stale checkout that predates #17347.
2026-08-30 00:57:14 -07:00
Neil 09429768c5 test(cross-version-wire): compare published fields per frame (#17301) 2026-08-30 00:26:17 -07:00
Jinwoo Hong 252dbd60ea fix(terminal): restore lossy initial remote snapshots (#17113)
* fix(terminal): restore lossy initial remote snapshots

* test(terminal): strengthen lossy snapshot causal oracle
2026-08-30 03:11:55 -04:00
Jinwoo Hong ae0f3675a1 fix(remote): focus host-delegated split panes (#16886)
* fix(remote): focus host-delegated split panes

Return the authoritative leaf identity from terminal.split, record viewer-local focus intent behind the captured pairing revision, and replay the mirrored layout before focusing the exact pane. Preserve old-host fallback and prevent delayed split responses from stealing focus after the viewer moves away.

Add deterministic runtime, renderer, concurrency, compatibility, and headed paired-Electron coverage for Cmd+D, header splits, and immediate PTY input routing.

Fixes #16510

* fix(remote): preserve split focus across tab groups

Resolve the initiating source tab and leaf from the remote PTY, while keeping the viewer's current focus as a separate anti-steal baseline. This lets context-menu/header splits from non-focused group tabs focus their result without allowing delayed responses to override a later navigation.

* test(remote): drive split focus with key events

* test(remote): use the platform split shortcut

* fix(remote): fence concurrent split focus intent

* fix(remote): harden split focus ordering

* fix(remote): preserve split focus after runtime refactor

* fix(remote): fence stale split focus gestures

* test(remote): keep split focus regression within line budget
2026-08-30 03:07:20 -04:00
Neil 5ea9daba97 fix(window): keep automated Electron launches out of the foreground (#17347) 2026-08-29 23:55:00 -07:00
Jinwoo Hong f572ba34bc feat(browser): address-bar convergence — previews and browser tabs convert in place (STA-5681) (#16998) 2026-08-29 22:38:59 -07:00
Jinjing 73ff003147 test(e2e): cover session upgrade and Windows terminal recovery (#17289)
* coverage report

* rm test coverage

* test(e2e): cover session upgrade and Windows terminal recovery

* fix stub
2026-08-29 22:37:05 -07:00
Neil c6641152f1 Split relay dispatcher layers (#17174)
* Split speech session lifecycle

* Split terminal output scheduler pipeline

* Split mobile browser pane modules

* Prune resolved max-lines suppressions

* Split pane tree equalization logic

* Extract mobile troubleshoot screen styles

* Split external automation manager

* Split main window service attachments

* Split hosted review creation checks

* Split automation dispatch event handling

* Split settings navigation metadata

* Split daemon initialization lifecycle

* Split GitLab item dialog

* Split relay dispatcher layers

* Fix F3-speech for #17123

* Fix F1-cycle for #17131

* Fix F4-navtest for #17157

* Fix F2-allowlist for #17161
2026-08-29 20:15:10 -07:00
Jinjing 2214d29f15 fix(browser): close guest-owned split tab (#17281)
* fix(browser): close guest-owned split tab

* fix: check sourceId before toggling floating panel on close

The empty-panel toggle is the ambient fallback only. Guest-initiated
closes (with sourceId) target the main workspace and should not toggle
the panel.

* test(browser-split-shortcuts): remove terminal-mirrors close test and un

Removes test case that verified Cmd+W closes guest-owned browser splits when
active-tab mirrors point to a terminal, along with the helper function and
unused fixture properties that only that test required.
2026-08-29 15:43:36 -07:00
Neil 2dfaa676d8 chore: update oxlint and oxfmt (#17150) 2026-08-29 14:13:35 -07:00
Neil b17f60d744 build: upgrade to pnpm 12 (#17156) 2026-08-29 14:13:26 -07:00
Brennan BensonandBrennan Benson 11d8673112 test(cross-version-wire): derive skew expectations from the baseline under test (#17178)
* test(cross-version-wire): derive skew expectations from the baseline under test

The cross-version wire job pairs current code against whichever release tag is
newest, so a hand-written "the old side does not have X" assertion expires by
itself: v1.4.192 was the first tag containing the SnapshotStart `terminalOwner`
field, and cutting it turned the new-client/old-server pairing red on unrelated
pull requests with no code change anywhere.

Read what each build publishes from that build. Each host is now paired against
a client of its own version to produce a reference, and the skewed pairings are
compared against that reference, so the expectation is whatever the release
actually shipped. The same class of assertion in the agent-session suite —
"the old build advertises no structured capability and registers no structured
method" — becomes "each build's advertisement agrees with what it registers",
and the "client too old to know this capability" is derived by removing the
capability from the baseline's own list.

The guard is unchanged in strength: a field the old host still publishes may not
be dropped, skew may not change what a host puts on the wire, and a new pairing
asserts the oracle still stalls when a peer cannot decode an opcode the other
side sends.

* test(cross-version-wire): exercise release structured methods

* test(cross-version-wire): load the registered method manifest

* test(cross-version-wire): assert execution, not registration, on both host gates

The release-shaped checkout gate accepted any reply that was not
method_not_found, so a registered-but-throwing handler passed it. The
capability gate asserted a shared host spy had been called at all, so the
second method mapped to that spy could stop reaching the host unnoticed.

* test(cross-version): make the release-shaped skew cover the whole agent-session manifest

The release-shaped checkout is the only place the "registered means usable"
claim is executable today — the baseline release registers none of these
methods — and it was exercising one of sixteen. A handler registered and
returning an execution error passed the suite.

- Declare each method's result in the manifest, so "answered" is the contract
  rather than "did not say method_not_found".
- Give each build a seam to install a host into its own module slot; a release
  checkout has its own copy, so the working tree's host was never this
  dispatcher's, and every host-backed method answered
  structured_agent_session_unsupported — the capability gate's own words.
- Run one execution contract over both skews instead of two divergent loops.
- Pair the AI Vault never-called spy with a positive control; renaming the
  runtime method it watches left it green.

---------

Co-authored-by: Brennan Benson <brennanbenson@Brennans-MacBook-Pro.local>
2026-08-29 13:42:50 -07:00
Neil 92ab618a11 Repair scheduled computer-use CI (#17122)
* Repair scheduled computer-use CI

* Make Calculator E2E Windows-version neutral

* Handle classic Calculator accessibility panes

* Update Calculator E2E source contract
2026-08-29 01:50:38 -07:00
Brennan Benson fd9125ea8c feat(native-chat): Codex structured native chat restructure (#16729)
* feat(native-chat): port structured Codex sessions from restructure-recovery

Rebuilds the desktop structured native-chat implementation from
brennanb2025/native-chat-restructure-recovery (tip 4e31c08db3) on top of
current main as a single commit, scoped to the local Codex path.

Ported:
- Structured agent-session core: durable record store + single-writer lease,
  canonical journal, agent-session wire host/attach/eviction/subscribers,
  `agentSession.*` RPC surface (registered via ALL_RPC_METHODS; host-side
  mobile allowlist included for wire compat), pty write gate, transcript
  additions, and the Codex app-server adapter/launch resolution.
- Renderer: NativeChatStructuredSession view/composer stack, structured
  launch path with the single-flight guard, local structured session tabs
  sync, activation gate + structured inventory (read-only
  `agentSession.handoffStatus` probe), agent-session tabs in the tab strip,
  AI-vault structured session activation, and the settings pane with the
  parent Experimental Chat UI toggle plus the nested "Use updated structured
  native chat" toggle. New sessions require both flags, agent codex, no
  prompt, and a local non-WSL, non-Windows-host execution host
  (structured-native-chat-availability).
- Fixes 72c013cea6 (verified Codex launch recovery), 8ddbaf5e3d (defer
  native terminal view switching affordances), and 4e31c08db3 (release the
  launch gate after a visibility retry) with their regression tests,
  including the third-launch-after-retry guard case.
- Cross-version agent-session wire test + CI lane, packaging entries
  (proper-lockfile, agent-tooling asar excludes), and the wire-compat doc
  section.

Deliberately not ported: mobile/ changes, the Claude structured runtime
(only the claude-transcript-branch-proof and claude-structured-owner-identity
leaf modules remain, backing the kept TUI-recovery arms), the terminal↔chat
adoption/handoff flow (`agentSession.adoptTerminal`/`requestHandoff`, the
handoff request engine, TUI adoption machinery, orca-runtime adoption
methods), renderer switching affordances and their dead leftovers, the
hook/subagent-status refactor cluster, and unrelated branch changes. The
crash-during-acquisition recovery path (restart handoff adjudication,
restore/reverse re-acquire, lease schema handoff keys) is kept because every
plain direct launch depends on it; a trimmed handoff coordinator exposes
only status/restore/close.

Branch edits that targeted files main has since split (ipc/pty.ts,
worktrees.ts, rpc/methods/terminal.ts, useIpcEvents, pty-connection,
store/slices/terminals.ts, runtime-types, web preload) were re-applied to
the split modules, preserving main's newer logic (Windows CIM fallback,
browser tab close rework, cold-restore resume flow, dispatcher threading).

Known seam: the mobile clipboard image-provenance CONSUMER gate ships
(agentSession.send refuses unproven mobile image refs with
agent_session_image_untrusted) but the producer hunk in
rpc/methods/clipboard.ts stays with the unported mobile cluster, so mobile
image sends into structured chat fail closed until that side ports.

* fix(native-chat): trust only authenticated local image uploads

* fix(build): preserve Windows process-tree patch application

* test(windows): include process creation time in addon fixture

* fix(build): run windows-process-tree node-gyp from the physical package dir

gyp expands the node-addon-api dependency by probing node, whose cwd
resolves to the package's physical directory in the store, so the emitted
target is a store-relative ../../../../node-addon-api@... hop. gyp then
resolves that hop against the rebuild cwd; from the node_modules
symlink/junction it escapes the store and configure fails with
"node_addon_api.gyp not found" (run 32999886072).

Rebuild from realpath(package dir) so both bases agree, matching how the
package manager itself runs native install scripts. The regression test
replays gyp's expansion+resolution against the planned cwd and fails
without the fix.

* fix(native-chat): keep chat tabs visible through terminal closes and empty-worktree launches

Two proven blockers in the native Codex tab contract:

closeTerminalTab pre-empted the canonical unified close. With one terminal
left it deactivated the worktree on a terminal/editor/browser-only check,
blanking a workspace that still held a renderable agent-session tab; with
two or more it pre-picked a successor from terminal entities only,
re-stamping the group active before closeUnifiedTab's MRU/neighbor repair
could land on the chat tab. Successor choice now defers to the unified
contract whenever the terminal has a unified row, and deactivation is
gated on the unified renderable count (matching leaveWorktreeIfEmpty),
with the legacy pre-pick kept only for terminals without a unified row.

A structured session created on an empty worktree was published into the
host's headless group while preserveLocalLayout froze the local layout,
leaving the tab in store but permanently off screen. A preserveLocalLayout
owner now always takes client-owned placement — repairing a rendered
leaf whose group record is missing, or materializing a rendered group on a
truly empty worktree — and applies the client-derived layout repair while
still rejecting host-authored layout.

Regression tests drive the real store through closeTerminalTab (git
worktree and folder workspace) and the real snapshot applier for the
empty-worktree adoption states; all fail without the fixes.

* fix(native-chat): close stale turns and retry rejected sends

* fix(native-chat): retire hosted rows on structured tab activation

* fix(native-chat): preserve rpc defaults across main merge

* chore: format remote wire compatibility guide

* test(native-chat): cover retry after unconfirmed send

* fix(native-chat): reload outbox on session switch

* docs(settings): disclose structured chat platform limits

* fix(native-chat): await Codex launch-home preparation

* fix(codex): align child-process allowlist with async trust bridge

* test(identity): update inventory for tab surface refactor

* fix(windows): preserve process-tree CRLF patch sources

* fix(native-chat): anchor an unmatched chat echo where it was sent (#16117)

* fix(native-chat): anchor an unmatched chat echo where it was sent

The reported symptom was old user messages replaying below every new turn, so the
conversation read as scrambled. The cause was not that the echo failed to match a
transcript row. Claude consumes a mid-turn send through a `queued_command`
attachment and writes no `type:"user"` record for it, so some echoes can never
match, and no amount of matching will change that. The cause was WHERE an
unmatched echo rendered: buildMobileNativeChatTransientData appended every pending
item after the entire transcript, so it re-read below each turn that landed
afterwards.

Render each echo directly after the transcript row it was sent against, using the
baseline the send already captures. An unmatched echo is then at worst a duplicate
in the right position rather than a scrambled one, and it stays visible. Echoes
sharing an anchor keep send order; a send with no baseline, or one whose anchor
folding dropped, still falls back to the tail.

Deliberately NOT fixed by deleting the echo. Inferring from send ordering that an
echo can never match, then removing it, loses the user's own text for a message
the agent did receive, and it cannot fire in the common case anyway - measured
drain groups are 1,017 of size 1 against 55 larger. It also escalates an existing
gap: the count pass has no baseline-tail guard, unlike the glue pass, while
`messages` is a 40-row window that head-trims, resets on reconnect and grows at
the front on loadEarlier, so a false landing there would license deleting a
DIFFERENT outstanding message.

That count-pass gap is real and left for a separate change; anchoring makes its
worst case a duplicate in place rather than a scrambled conversation.

* fix(native-chat): preserve folded echo anchors

* fix(native-chat): preserve forward-folded echo anchors

* fix(native-chat): keep leading folded echoes in place

* fix(workspace-cleanup): show git status for every row (#16690)

* fix(native-chat): refuse structured chat on every Windows execution path

canUseStructuredNativeChat only refused win32 when a project runtime
resolved, so folder-workspace keys (and other keys with no project
runtime) failed open into structured chat on Windows. Fail closed on
win32 unconditionally after the host check, matching the settings copy:
local macOS/Linux only; Windows/WSL/SSH stay on terminal chat.

* fix(native-chat): restore runtime refusals behind the win32 gate

506d375de3 replaced the project-runtime checks with a bare platform test,
so a WSL or repair-required runtime resolution would no longer refuse
structured chat off-win32. Keep the unconditional win32 refusal and
re-run the runtime resolution after it, so the gate does not depend on
the resolver's own platform guard. Tests inject WSL and repair-required
resolutions on darwin/linux and fail against the regressed gate.

* fix structured session journal durability

* fix structured tab active pointer after restart

* fix(native-chat): await optional lease renewal callbacks

* refactor(skills): extract install error messages

* fix(agent-session): harden recovery ownership

* fix(native-chat): retain panes across tab activation

* fix(native-chat): address round-one review findings

* test(native-chat): align integration coverage after main merge

* fix(native-chat): harden round-two reliability

* fix(native-chat): harden round-three reliability

* fix(native-chat): close round-four recovery gaps

* fix(native-chat): separate bounded journal key forms

* fix(native-chat): reset outbox error in render on session switch

The switch effect adjusted error state after the sessionId prop changed,
tripping react-doctor's no-adjust-state-on-prop-change on the changed-code
gate and flashing the old session's banner for a frame. Reset it with the
render-time previous-value guard instead.

* fix(native-chat): invalidate stale outbox settlements

* test(native-chat): restore settled-error session-switch regression

a6e2379bd1 replaced this test with the in-flight settlement race test,
leaving the render-time error reset unpinned: deleting the reset block
still passed the whole native-chat suite. Keep both scenarios pinned;
they are distinct (settled error clears on switch vs stale settlement
invalidated in the commit-to-passive window).

* test(wire): make release checkouts race safe

* test(wire): pin cross-process checkout single-flight and importer specifier contract

* test(wire): harden release checkout lifecycle

* fix(build): drop CR-byte residue from windows-process-tree patch

The two trailing CR bytes on the patch's deletion lines are a proven
no-op: pnpm hashes patches CRLF-normalized (both forms hash to the
lockfile's 946ffb2b) and materializes this package without applying the
patch in either form, so the load-bearing build edits come solely from
applyWindowsProcessTreeBuildFixes() (#16947), which handles both source
EOL forms. Restore byte-identity with main and repin the contract test
to the post-#16947 reality: LF-only patch bytes plus lockfile hash sync.

* fix(native-chat): skip empty startup recovery
2026-08-28 16:45:58 -07:00