Commit Graph
2191 Commits
Author SHA1 Message Date
Neil 057fbfcffc perf(windows): read the process table natively instead of forking PowerShell (#15749)
* perf(windows): read the process table natively instead of forking PowerShell

Seven independent readers each forked powershell.exe to run
Get-CimInstance Win32_Process, with a wmic fallback that Windows 11 24H2
has removed. On a domain-joined host with PowerShell Transcription
enabled by policy, one of them running every ~2s recorded ~289GB across
1.4 million files (#15209). The same scan cost ~700ms and ran per pane
(#15036), and a Group Policy or AV block turned it into 'unavailable',
which callers read as 'no evidence' -- which is how a PTY tree survives
its own teardown (#9045, #10475).

A Toolhelp32 snapshot answers the same question with no child process.
Measured on Windows 11 with 1050 processes, p50/p95:

  pid+ppid+name          15.9 / 17.5 ms
  +memory +command line  30.6 / 33.7 ms
  Get-CimInstance         706 / 723  ms

Two upstream defects needed patching, both found by running it on real
hardware. The binding requires Spectre-mitigated libraries our agents do
not carry (node-pty is patched the same way). And enumeration stopped
after 1024 processes: on a host with 1051 the module returned exactly
1024, and the querying process was itself among the 27 missing -- a
truncated snapshot silently hides the descendants teardown is looking
for, which is the failure this whole change exists to remove.

Migrated: the foreground/descendant reader (the #15209 scraper and the
teardown identity gate) and the port scanner's PID attribution. NOT
migrated: the memory collector and three identity probes, which need
Win32_Process.CreationDate and have no native equivalent. Start time is
a proxy for identity anyway; an inherited job handle is the real answer,
so those belong with the job-object work rather than here.

Packaging follows the windows-native-registry contract exactly:
optional, absent from onlyBuiltDependencies so macOS/Linux never run
node-gyp, win32-only in the packaged runtime. Asserted by the existing
contract test, which also stops pinning a whole source literal that only
tested its own formatting.

* chore(process): ratchet the child_process allowlist down

windows-foreground-process-rows.ts no longer spawns anything, so its
allowlist line is stale. The guard fails on a stale entry as well as a
new one, precisely so a migrated file cannot keep a slot open and hide
the next regression in the same path.

* fix(ports): import the process-table reader the scanner uses

Missing import: the migration replaced the PowerShell call but the new
symbol was never imported, so tsc failed. Vitest transpiles without
typechecking, which is why the port-scanner suite stayed green.

* fix(deps): sync this branch's lockfile with its patch set

Same class as the fix on the tip branch: pnpm records a hash per patched
dependency, and this branch introduces the windows-process-tree patch
without its lockfile entry matching. Every job here failed at install
with ERR_PNPM_LOCKFILE_CONFIG_MISMATCH.

Verified with --frozen-lockfile, which is what CI runs and what my local
runs were not.

* test(relay): drive the relay's Windows fixtures from the native snapshot

Two relay cases fed a PowerShell CIM payload through a mocked execFile.
That reader is gone, so both failed -- deterministically, on every PR
run for this branch and the one above it.

I did not catch it because my own verification sweep was
'src/main src/shared config/scripts' and never included src/relay. The
relay is a first-class consumer of the process table; leaving it out of
the sweep is how a deterministic failure survived six review rounds.
2026-08-21 21:54:57 -07:00
Denis Darii d5667376b0 feat(dashboard): add a keyboard shortcut to toggle the Agent Dashboard (#15353)
Adds a configurable, unbound-by-default `dashboard.toggle` action that toggles the Agent Dashboard (in-window drawer or pop-out, per the existing mode setting).

- Wired through window-shortcut-policy, main-window dispatch, browser-guest dispatch, preload, and the renderer IPC handler.
- Opening the in-window drawer reveals the sidebar first; closing leaves it alone.
- Gated on the `experimentalAgentDashboardPopout` experiment, and the Settings shortcut row is hidden while that experiment is off.
2026-08-21 21:12:12 -07:00
Neil b7e79b7ca6 fix(windows): one chokepoint for every child process (#15746)
* feat(process): add the Windows-correct child-process chokepoint

Six decisions have to be made every time Orca starts a child process --
console visibility, argument quoting, .cmd interpretation, binary
resolution, timeout policy, and how the tree is later terminated. POSIX
forgives all six. Windows punishes each differently, and made per-call
site across 172 files they were right in some and wrong in others.

runProcess/spawnProcess make them once:
- windowsHide unconditionally, shell:false unconditionally (shell:true
  concatenates argv unescaped and silently disables windowsHide)
- .cmd/.bat routed through cmd.exe /d /v:off /s /c with a verbatim line,
  because Node refuses to spawn them otherwise (EINVAL)

The encoding was derived by measurement on Windows 11, not from the
docs. An embedded quote is written "" rather than \" so cmd's naive
quote count stays even -- with \" the parity flips and every later &
| < > on the line stops being data. Measured before the fix, argv
["a b", 'c"d', "e%F%g", "h&i", "j^k"] arrived as
["a b", 'c"d', "e^%F^%g", "h"]: the & truncated the argument and
ran its remainder as a command. Each % is broken out of the quoted run
as "^%" because %VAR% expands even inside quotes.

The import-boundary test is a ratchet seeded at today's 172 files; it
only shrinks.

* fix(process): route the console-flashing spawn sites through the chokepoint

The ssh -G config probe fires on every connect and reconnect, and ssh.exe
is console-subsystem, so a GUI-subsystem parent gets a fresh visible
conhost that takes foreground -- keystrokes typed into an Orca terminal
at that moment go into the black box (#10488, #14543). Same for the
ProxyJump tunnel, the ProxyCommand cmd.exe wrapper, the font enumeration
and the DPAPI cookie decrypt.

Also stops spawning powershell by bare name: PATH under Electron is not
the user's, so where policy has pruned the System32 entry the spawn fails
and the font picker silently reports five hardcoded families rather than
an error (#11771).

Deletes system-fonts' 40-line bespoke execFileText -- timeout, output cap
and kill are the chokepoint's job now. Adds runProcessSync so the sync
callers have a compliant path; without one the ratchet could never
reach zero.

The three suites that mocked child_process directly now mock runProcess,
which is the point: how a process gets started is no longer each
module's business. Ratchet 173 -> 170.

* fix(process): do not report a deliberately killed child as timed out

runProcessSync inferred a timeout from signal === 'SIGTERM'. Measured:
a real timeout sets error.code ETIMEDOUT and kills with SIGTERM, but so
does anything else that terminates the child -- and those cases set no
error at all. Reading the signal alone reports a process someone stopped
on purpose as having timed out, which callers retry.

* refactor(process): hold the ratchet as data and migrate the pwsh probes

The allowlist and the adversarial argument corpus are read only by tests,
so they were production modules in name only; they move to __fixtures__.

pwsh.ts carried isTimeoutError() purely to reconcile two spellings of the
same event -- execFileSync reports a timeout as ETIMEDOUT, the execFile
callback as a SIGTERM kill with no code. runProcess reports one timedOut
flag, so the helper and the reasoning behind it both go.

Its sync probe also spawned without windowsHide, which flashes a console
and steals foreground on every cold cache read.

* refactor(process): migrate five more spawn sites onto the chokepoint

Each one deletes a hand-rolled promise/timeout/kill wrapper and stops
re-deciding console visibility for itself. Ratchet 170 -> 164.

Two things this surfaced, both kept:

runProcess now accepts string chunks as well as buffers. A stream someone
called setEncoding on emits strings, and concatenating those as buffers
throws inside a data handler -- where the rejection has nowhere to go and
the caller simply hangs rather than failing.

ProcessSpec keeps its AbortSignal. I had removed it as unused; the macOS
PAM preflight passes one through from its own caller.

ipc/app.ts is deliberately NOT migrated. Its probe spawns a three-stage
 pipeline detached so a timeout can reap the group with one
negative-pid SIGKILL; runProcess kills only the root, which would orphan
the plutil stages. Migrating it needs the chokepoint to own POSIX
process-group termination first -- the same guarantee job objects give on
Windows. Reverted and left on the ratchet.

* test(process): do not assert a POSIX signal on Windows

Windows has no signals, so the same deliberate kill reports an exit code
there and a signal on POSIX. What has to hold on both is that neither
shape reads as a timeout. Caught by running the suite on Windows.

(cherry picked from commit 0a6e9902a22a369a0e85e113ea8d87b726f82e1f)

* fix(process): settle a timed-out run even when the child ignores the kill

close only fires once the child is actually gone, so a child that traps
SIGTERM never emits it and the promise outlives its own deadline
forever. That is the same wedge shape just fixed for the process table,
and it is worse here: pwsh.ts and the snapshot reader both cache an
in-flight probe, so one unkillable child hands every later caller the
same dead promise.

After the deadline it now escalates to SIGKILL and settles regardless,
reporting timedOut with whatever output arrived.

(cherry picked from commit 78ac169197c4e6faee1b9310a7186029cc11acbc)

* fix(process): escalate an aborted child too, not just a timed-out one

The grace escalation I added covered the timeout path and left abort on
the old one, so an aborted caller with an unkillable child still waited
forever -- the same defect, one path over. The macOS PAM preflight is a
real caller that passes an AbortSignal.

Both paths now share one stop-and-settle, and the result reports
timedOut honestly: false when the caller aborted.

(cherry picked from commit 7e9523a9e31172bb8183661b56f04c3ab6a03d0d)

* fix(windows): stop percent escaping from forging an escaped quote

escapePercentForCmd ran as a post-pass over the quoted string, so it
inserted a quote wherever a percent was -- including straight after a
backslash. CommandLineToArgvW reads backslash-quote as an escaped quote,
so C:\Users\%USERNAME%\x arrived corrupted. That is about as common as
Windows paths get, and my 20-case corpus had no backslash-before-percent
entry to catch it.

Percent handling is now part of the quoting loop, where the backslash
run is known and can be doubled before the inserted quote. Two corpus
cases cover the shape.

The program path gets the same treatment. It was quoted but not
percent-escaped, so a launcher under C:\Users\%USERNAME%\ had its own
path expanded on the cmd hop.

quoteWindowsArgument no longer takes a boolean. Passing it to
values.map() handed map's index in as the flag -- which is how the first
version of this fix was written, and the corpus test caught it.

Separately: an AbortSignal that was already aborted never fires the
event, so runProcess ran the child to its full timeout for a caller who
had already given up.

(cherry picked from commit f7e2e56b1ee1f27ab6d1035dde4501b38f95b374)
2026-08-21 21:05:24 -07:00
Neil 080c95940f fix(ssh): let fork-PR worktrees add their contributor remote via the relay (#15827)
* fix(ssh): let fork-PR worktrees add their contributor remote via the relay

Creating a workspace from a fork PR on an SSH host failed with "Destructive
git remote operations are not allowed via exec". The relay's git.exec
allowlist blocked every `remote` write subcommand, but SSH fork-PR creation
has to run `git remote add <fork> <url>` on the host before it can fetch and
track the contributor's branch, so the whole create aborted.

Allow exactly the two shapes that flow needs -- `remote add <name> <url>` and
`remote remove <name>` -- validated with the same remote-name and URL rules
the relay already applies to every pushTarget-carrying RPC. Everything else
(set-url, rename, prune, extra operands, flags before the action) stays
blocked, and the URL must be a github.com clone/ssh URL, so no new reach is
granted beyond what push/fetch already accept.

`remote remove` was blocked too, which silently leaked fork remotes on SSH
hosts: worktree removal swallows the cleanup error. It works again now.

A host still running an older relay gets an actionable "reconnect to deploy
the latest relay" message instead of the raw policy error.

* test(git-exec): pin remote read/write mutation classification

Misclassifying `git remote` / `remote get-url` as mutating would flush the
relay and SSH provider git read caches on every remote probe, so pin both
directions.
2026-08-21 14:44:08 -07:00
Brennan Benson 3fca1d1648 fix(linear): unbound list-issues by default, surface truncation, bind cursor workspace (#15824)
Fixes STA-5076.

list-issues capped at 50 by default and hard-clamped at 250, with hasMore buried
under result.meta and no stderr warning for --json, so a page that stopped early
read as a complete answer. Omitting --limit now walks Linear's pages until they
run out (meta.limit is null), and --limit <n> is the only cap, paging past
Linear's 250-per-request maximum to reach it. result.truncated sits next to
result.issues and is set only when a cap actually held results back; human output
prints "truncated: showing N".

The read still has to fit the CLI's 60s RPC budget, so a 20s wall-clock deadline
and a 200-page ceiling stop the walk early and report truncated with a
continuation cursor rather than failing the command.

Also:
- issued --cursor values bind the resolved workspace, so call -> nextCursor ->
  call works without --workspace; raw Linear cursors still need one and now carry
  nextSteps
- issued cursors whose payload smuggles back `all` or an empty workspace are
  rejected at decode, since either would widen the read past the bound workspace
- JSON issue rows carry priorityLabel (none/urgent/high/medium/low), matching
  orca linear priority set
- truncated and priorityLabel are optional on the wire, so a host that predates
  either is not read as "complete"; readers fall back to meta.hasMore
- the truncation line prints the rows actually rendered, so a remote result with
  no meta.returned cannot print "showing undefined"
2026-08-21 14:28:55 -07:00
OrcaWinandBrennan Benson 2a68b78bb3 fix(worktree): let a configured worktree base outrank a built-in visibility source (#15232) (#15430)
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-21 12:20:04 -07:00
OrcaWin 2c24025e1d fix(worktrees): honor absolute Linux worktree base paths for WSL repos (STA-4772) (#15384) 2026-08-20 19:10:28 -07:00
JinjingandJinwoo-H 6e25a90085 fix(terminal): keep a quick command queued until its own spawn takes it (#15630)
* Increase shell readiness timeout to match daemon barrier

Slow interactive rc files can take longer than 1.5s to initialize. Raise
the startup command readiness timeout from 1.5s to 15s to match the daemon
barrier and prevent queued commands from executing mid-startup.

* fix(terminal): keep a quick command queued until its own spawn takes it (STA-4876)

Triggering a quick command opened a terminal tab titled with the command's
label, the shell started and drew its prompt, and the command never ran.

TerminalPane snapshots `pendingStartupByTabId[tabId]` in a useState lazy
initializer, and a mount effect deleted the entry immediately. The pane's key
is `${tab.id}-${tab.generation ?? 0}`, so anything that bumps generation before
the command reaches a shell — the stall-recovery remount fired from
`requestTerminalPaneRecovery`, or the allDead activation regeneration — mounted
a second pane that re-read an emptied slot and spawned with no command at all.
The loss was permanent, which is why every scope failed alike: repo, global and
agent-prompt all funnel through the same queue-then-snapshot sequence.

Spend the entry at `onPtySpawn` instead, which is the one point that proves this
pane's own fresh spawn exists. A pane retired mid-connect never reaches it, so
the command stays queued for the next mount; reattach skips `onPtySpawn`, so it
cannot spend a command it never delivers.

Three details are load-bearing:

- Ownership is reference identity (`paneOwnsQueuedStartup`). Setup and issue
  splits borrow the same `deps.startup` field for their own one-shot payload, and
  that payload can be structurally identical to the queued command, so a
  truthiness test would let a split pane spend a command it never runs.
- The consume runs after `bindActivePanePty`. While the tab still has no ptyId,
  the queued entry is the only thing holding its worktree out of the
  retention-budget force-park, so dropping it first unmounts the pane mid-spawn.
- The callback is one-shot. `onPtySpawn` fires on every fresh spawn a pane makes,
  including hibernation wake and the respawn ladder, and a command queued after
  the first launch belongs to that later launch.

Known residual, documented at the call site: the consume tracks "a pty exists",
not "the command ran". Windows embeds short commands in the shell argv, so they
execute before the spawn resolves and a pane retired in that window re-delivers
on remount; on POSIX the write waits for shell-ready, so a pty that dies in that
window loses a command already spent. Closing either needs a delivery signal
from main rather than this callback. Both windows are narrow, and both are
strictly better than losing the command unconditionally.

* fix(terminal): guard the queued-startup wiring the review found untested

Follow-ups from the final review pass on this branch.

- Collapse the ownership + one-shot decision into `createQueuedStartupConsumer`
  so the call site is a single call rather than inline logic no test could
  reach. Two mutants survived the whole suite before this: relaxing ownership to
  a truthiness check, and dropping the one-shot guard. Both now fail.
- Rewrite the throwing-consume test. It asserted `updateTabPtyId` had been
  called, which runs *before* the callback, so it passed with the try/catch
  deleted. It now asserts the throw does not escape into the connect promise,
  which is the invariant the try/catch actually provides.
- Correct the `onQueuedStartupSpawned` docblock. It claimed the callback is "the
  first moment the command is guaranteed to reach a shell"; the diff's own caveat
  says otherwise, since Windows runs an argv-embedded command before this fires
  and a POSIX shell can die before the shell-ready write. It marks a live shell,
  not delivery.

No behavior change: the consumer is the same predicate and the same one-shot,
moved behind one exported seam.

* fix(terminal): roll shell wrapper isolation into a fresh daemon

* Revert "Increase shell readiness timeout to match daemon barrier"

This reverts commit 6ab273ef36.

* Condense queued startup spawn comment

Simplify the multi-paragraph explanation into a concise summary that captures the key points: spawn must wait until after the pane is bound to preserve the worktree from force-parking, and the behavior differs between POSIX and Windows for delivery timing.

* fix(terminal): prevent consuming replaced queued startup commands

When a queued startup is replaced before the pane's first spawn,
the one-shot guard alone still allows consuming the replacement.
Add isStillQueued callback to verify the slot still holds the
originally captured command (STA-4876).

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
2026-08-20 12:15:08 -07:00
Jinjing d8e9fa1bb9 Revert "fix(terminal): apply pane padding on all four edges (#15544)" (#15623)
This reverts commit 4b2ed5ddd4.
2026-08-20 09:57:55 -07:00
NeilandBrennan c92f394cde fix(pty): delete the reply-withholding scheduler (#15578)
* fix(pty): answer a terminal colour query in its own turn

Root-cause follow-up to #15559, which stopped a CPR overtaking a deferred
colour reply but left the deferral itself in place.

Orca answers terminal queries by writing to the PTY master, which a line
discipline in ECHO copies straight back out as junk on a cooked prompt
(#12112). The guard was to withhold the write until an `stty` subprocess
proved ECHO clear — and forking is what forced the decision to be async.
Any deferral, however short, lets a reply written later in the same turn
overtake this one, so the async probe was the bug's root cause.

Read the bit synchronously instead. Linux and the BSDs redirect a
master's mode ioctls to the slave, so a `tcgetattr` on the master fd
node-pty already owns answers for the slave with no fork: measured 0.26us
against 2403us for the subprocess. With a verdict available inline, a
querying program that already cleared ECHO — every raw-mode prober,
including the colour probe behind the `gh auth login` report — is
answered in its own turn and can never be reordered.

The deferral stays for the genuinely cooked case, and the ordering
guarantee stays underneath it: hosts whose node-pty predates this patch
get no sync probe and fall back to the deferred path, which mixed
client/host versions make a live production path.

Reply routing is all-or-nothing: a payload needing neither containment
nor ordering stays on the host's own path, so a CPR answered during shell
startup cannot pass the daemon's post-ready flush gate and splice into
the buffered startup command.

Native side is fail-safe: a kernel that did not redirect would answer
from the master's own termios, whose ECHO defaults set, so the degraded
verdict is "echoing" — never a false "quiet". The JS half ships in the
pnpm patch while the binding needs a source build, so
ORCA_REQUIRE_NODE_PTY_ECHO_STATE=1 makes CI fail rather than silently
skip when it is handed an upstream prebuild.

Co-authored-by: Brennan <brennanb2025@users.noreply.github.com>

* fix(pty): keep the flush ordered under synchronous re-entry

Three defects found in external review of the reply-ordering work.

node-pty delivers onData inside the master write, so a query can be
answered while the queue is mid-flush. `flushPendingWrites` spliced the
array off before writing, so that reply saw an empty queue, took the
same-turn path, and landed ahead of entries the loop had not written yet
— reproduced as 01, 99, 02, 03. It now shifts one entry at a time so a
re-entrant reply queues behind the rest, bounded by the length at entry
so a re-entrant push cannot spin the loop.

An overflow flush can re-enter as far as teardown. `answer` did not
re-check `closed` afterwards, so it queued behind a closed delivery,
returned true, and the reply was never written and never reported.

The payload router's ownership comment overstated its guarantee. The
`any` semantics are deliberate — returning false after a constituent was
already written would have the caller re-write the whole payload and
duplicate it into the child's stdin — so the residual mixed-failure drop
is now documented rather than implied away.

* fix(pty): delete the reply-withholding scheduler

Orca answered a terminal query by withholding the write until a probe
proved the slave's ECHO bit was clear. That was the wrong mechanism, and
it is now gone: replies are written in the caller's turn and their echo
is contained on the output side, where it always was.

Withholding never removed an echo. The wait was bounded and always ended
in a write, so the output-side projections were doing the work the whole
time — including the readline rewrite, which happens with the tty already
raw and which therefore no reading of the ECHO bit can predict. What
withholding did add was an asynchronous write path, and that is what let
one reply overtake another and land in the next program's stdin (#15559),
what produced a re-entrancy inversion inside its own flush, and what four
rounds of regressions have lived in.

The last thing it covered was the verbatim echo of a `stty -echoctl` tty.
That shape is now projected directly. It starts with ESC, so it is
matched only when complete and never held as a partial: holding it would
take a bare trailing ESC from the query parser and an expired hold would
release it raw, so a query torn at its own ESC would never be answered.
Complete-match-only is what makes the shape safe to project at all.

Measured on a real pty: a cooked-mode master write is both echoed AND
delivered — ECHO copies the bytes without consuming them from the slave's
input queue, so a program arming raw mode with TCSANOW/TCSADRAIN (libuv's
setRawMode, hence every Node agent) still reads them. Only a TCSAFLUSH
switcher discards it, which it does on every terminal, none of which
gates a reply on termios state.

Deletes the pending-write queue, the async stty probe, the poll budget
and probe rate limit, the deadline-driven flush, and the answer/
answerInOrder split. Replies now leave in call order by construction.
No packaging, native or CI surface is touched.

* test(pty): restore stty-probe coverage and pin the duplicate-query retry

Archaeology on how withholding got here, and what its tests were really
protecting.

Deleting the ECHO probe took four tests with it that were not about the
probe at all: they cover createSttyProbe, which the shell-readiness
line-editor probe still uses — in-flight sharing, the per-platform stty
flag, and transient-versus-permanent failure latching. Restored against
the line-editor probe, which is now their only caller.

Also pins the property that answers the one case an immediate write
cannot serve. A program that queries while cooked and then arms raw mode
with TCSAFLUSH discards the reply with the rest of its input queue.
Nothing can prevent that from the terminal side, and no terminal tries.
What matters is that such a program re-queries after its own timeout: the
ingress declines to answer an already-answered slot but forwards the
duplicate downstream, so the renderer's emulator answers the retry, by
which point the program is raw. The retry path is the recovery, not
withholding.

* ci(pty): keep the fish real-PTY test in the shell-contracts lane only

Reverting pr.yml to main dropped the exclusion for the fish query-reply
test, which this branch keeps, so it would have run in the sharded lane
as well. Restores it to the shell-contracts include list and the shard
exclude list, and drops the parallelism expectations for the deleted
cooked-querier suite and the echo-state env guard.

---------

Co-authored-by: Brennan <brennanb2025@users.noreply.github.com>
2026-08-20 02:15:42 -07:00
Brennan Benson 4b2ed5ddd4 fix(terminal): apply pane padding on all four edges (#15544)
* fix(terminal): apply pane padding on all four edges

Move the configured inset onto xterm so the terminal fills its pane while the fit calculation accounts for both sides of each axis. Add a geometry golden that forces cell remainders and verifies dynamic padding without relying on renderer pixels.

* fix(terminal): normalize imported padding for fitting

* fix(terminal): align stored and fitted padding
2026-08-20 01:07:56 -07:00
OrcaWin 471bc9d8ce Ship the WSL transcript helper with the Windows relay (STA-4831) (#15529) 2026-08-20 00:16:04 -07:00
Neil fdd4091ebd fix(hooks): isolate lint-staged backups per worktree (#15388) 2026-08-19 22:37:02 -07:00
Neil d7a23c84a9 fix(pty): keep a CPR reply from overtaking a deferred colour reply (#15559)
A background-colour probe writes `OSC 11 ;? ST` then `CSI 6n` and reads
exactly one response, using the CPR as its sentinel: a non-OSC first
response means "unsupported" and it stops draining. #13309 routed live
cooked-echo-risk replies through the ECHO-probe deferral while CPR kept
the immediate path, so the CPR overtook the colour reply, the prober gave
up, and the stray `ESC ]` was left in the tty for the next program —
`gh auth login` died on it with an escape-sequence error.

Queue a reply that needs no echo containment behind ones that do, FIFO,
and only while something is actually deferred, so latency-critical
replies stay immediate on every other path. Windows is unaffected: only
posix-pty defers, so the queue is always empty there.

Also: an in-flight echo probe is already the write continuation, so
re-arming the timer for a queued reply would fork a second stty and throw
away the first verdict; and teardown now hands queued uncontained writes
to the pty best-effort instead of dropping bytes the caller was told were
sent.

Known scope limit, pinned by tests and tracked for follow-up: the
guarantee is FIFO among recognised query replies, not over every byte —
a reply coalesced with a keystroke, and ordinary typed input, still
bypass the queue. Both were unordered before this change too.
2026-08-19 20:46:12 -07:00
Jinjing d541982b9c Improve cmd j search usefulness (#15551)
* Show last active time for workspace tabs in cmd+j palette

Adds session-age formatting and activity tracking to help users find
recently-used tabs. Replaces host badge display with last-active timestamps
that reflect either agent activity or worktree PTY activity, whichever is
more recent.

* Improve cmd+j search ranking with direct fields and recency

Prioritize results matching direct fields (titles, content) over
container fields (worktree, branch, repo). Use tab focus time to break
ranking ties. Makes search more useful for quick navigation.

* Extract path flavor logic to cross-platform-path utility

- Remove local pathFlavor function in favor of shared cross-platform utilities
- Simplify buildExcludePathPrefixes to use relativePathInsideRoot and resolveRuntimePath
- Ensures consistent path handling for both local and remote roots

* Improve cmd+j search ranking with recency-based tiebreaking

Track lastFocusedAt on tab creation/focus and use it to break ties between equally-ranked search results. This surfaces recently-used items first, improving search utility. Also fixes hasDirectHit to check field matches directly rather than evidence metadata.
2026-08-19 19:46:45 -07:00
Brennan Benson 30bf2647fc fix(mobile): replay a delivery-ambiguous worktree.create instead of failing it (#15472)
* fix(mobile): replay a delivery-ambiguous worktree.create instead of failing it

A socket close or response timeout rejects an in-flight worktree.create as
delivery-unknown: the frame reached the wire, so the host may already have
built the worktree. The client only replayed connection-migration cutovers,
so every other ambiguity surfaced as a create failure for a create that may
well have succeeded. Replay on the same clientMutationId — which the host
already dedupes — after waiting for the transport to come back.

* fix(mobile): bound the ambiguous worktree.create replay by the host's dedupe window

The replay was bounded only by a retry count, but what makes a replay reconcile
instead of building a second worktree is wall clock: the host drops a settled
create's dedupe record 60s after it resolves, and past that the replay is just a
fresh create that the host's suffix loop happily duplicates — for a folder
workspace, into a second workspace with the very same name and no collision
check at all.

Two paths ran past that window:

- The request-timeout path. A silently dropped response frame leaves the socket
  alive, so nothing rejects until WORKTREE_CREATE_TIMEOUT_MS — ten minutes, with
  no bound at all on when the host actually resolved. This was previously the
  path that replayed *soonest*, short-circuiting the reconnect wait because the
  transport still looked healthy. Invert it: every path that reports a real drop
  has already left 'connected' by the time the rejection surfaces, so still being
  'connected' identifies the timeout and is now refused.
- The reported-drop path. Worst-case detection is a full liveness idle period
  plus the missed-probe budget before the client even learns the socket is dead,
  and the old 20s wait on top of that overran the record. Derive the wait from
  the watchdog constants and the TTL instead of hardcoding it, and anchor a
  single deadline at the first ambiguity so a second wait gets the remainder
  rather than restarting.

The TTL now has one definition shared by both processes, so the client asserts
its budget against the host's real window instead of a copied literal.

* fix(mobile): end the reconnect wait on a revoked pairing, and pin the wait's behavior

waitForRpcClientReconnected resolves only on 'connected' or the timeout, but an
'auth-failed' client never reaches 'connected' — so a create interrupted by a
revoked pairing sat out the full wait before surfacing the error it already had.
Treat auth-failed as a terminal answer on both the fast path and the listener.

The helper also shipped with no tests of its own: its already-connected fast path,
its timeout path, and the synchronous-notification-during-subscribe teardown were
only ever exercised indirectly through the retry suite, and neither RpcClient
implementation notifies synchronously, so that branch had no coverage at all. Add
a direct suite covering all of them, asserting listener and timer teardown rather
than just the resolved value.

Also give the fake-timer tests an explicit timeout. advanceTimersByTimeAsync
yields through real macrotasks between ticks while vitest's own budget runs on
real time, so on a loaded runner the default 5s is reachable — observed once as a
spurious timeout in this suite.

* fix(mobile): bound the ambiguous replay in wall clock, not timer time

The replay window was derived from the liveness watchdog's own budget
(idle + missed probes x probe timeout). That is a bound on how long the
watchdog takes to *fire*, not on how much wall clock passed. iOS and
Android suspend JS timers while the app is backgrounded, so across a
background cycle the socket dies silently and the pending create rejects
delivery-unknown minutes later with the timer-derived ceiling still
reading ~44s. The replay then lands well past the host's 60s dedupe
record and the suffix loop builds a SECOND worktree - for a folder
workspace, one with the very same name and no collision check at all.

Anchor the deadline on the watchdog's lastInboundAt instead: a wall-clock
stamp of a frame that really arrived, so it stays honest across a
suspension. Fall back to the send time when the transport can't vouch for
one (relay sessions run with idleProbeMs: null), which errs toward
refusing the replay.

Also restore the delivery-unknown discrimination test that the
still-connected guard had made vacuous, pin the still-connected guard
itself against a live inbound stamp, and pin the deadline against being
re-read from a fresher replacement session.
2026-08-19 18:10:39 -07:00
Neil 36d78e88af fix(agent-hooks): stop Antigravity's Windows hook from spawning PowerShell on every event (#15520)
Antigravity was the last agent posting hook status through Windows PowerShell 5.1. Every
hook event — roughly one every 2-6s during an active session — paid a ~300ms interpreter
cold start, which is what made the console the agent allocates for each hook last long
enough to be seen as continuous flashing.

Move the Windows POST to the shared curl.exe builder every other agent already uses, via
the `extraFormLines` escape hatch for the `hook_event_name` field Antigravity uniquely
needs. Measured on Windows 11: 326ms -> 134ms per event.

Because the curl line percent-expands its arguments, the script also needs
`setlocal DisableDelayedExpansion` (#9358/#9941) so a `!` in a pane key or worktree path
is not eaten as a delayed reference.

curl omits a `--data-urlencode name@-` field entirely when stdin is empty, so accept an
absent or blank Antigravity payload as `{}` at the ingest boundary — the POSIX script
substitutes `{}` before posting and PowerShell did the same, and without this a
payload-less event lost the status transition its `hook_event_name` still carried. Scoped
to that source; every other agent keeps rejecting a body it cannot parse.

Adds a cross-agent guard asserting the invariant the original drift violated: a managed
Windows .cmd hook posts through fully-qualified curl.exe and spawns no interpreter.
Generated under a mocked win32 platform so the POSIX CI legs guard it too.

Validated on a real Windows 11 host, not an emulated platform check.

Fixes #15117
2026-08-19 17:36:33 -07:00
Neil cb95582cea feat(release): build unsigned Windows artifacts for the dev channels (#15465) 2026-08-19 17:34:27 -07:00
Jinwoo Hong 0e5b348414 fix(repos): keep the desktop-owned manual project order authoritative across paired clients and SSH catalog publishes (STA-4850) (#15538) 2026-08-19 17:26:16 -07:00
Neil a61b39a9a6 fix(runtime): stamp a runtime's own project setups as local, and report remote status about the remote (STA-4792) (#15376)
* fix(runtime): stamp a runtime's own project setups as local, and report remote status about the remote (STA-4792)

Two independent frame-of-reference bugs, both from code describing one machine
while labelled as another.

#15366 — projectHostSetup.* persisted the caller's host id verbatim. Those
`runtime:<environment-id>` ids are minted by the calling client's own pairing
store, so they name a machine only relative to that client. A client sending
one is addressing this runtime, and runtimes do not proxy these calls onward,
so the host it names is us. Storing the client's spelling made one machine look
like a different host to every other client, hid its rows from them, and
defeated the (projectId, hostId) duplicate check — two laptops paired to one
server each created their own setup for the same checkout. Re-spell it as
`local` at the RPC boundary. Rows written earlier keep their old stamp; readers
already project `local` back to `runtime:<their-id>`, so the client-visible
model is unchanged and no ids are rewritten.

STA-4792 defect 4 — `status --environment <name>` hardcoded app.running:false
to mean "no desktop on THIS machine" while every other field in the same object
described the target, including a desktopWindowStatus echoed straight from it.
The result contradicted itself and read as "that run was headless" when the
remote GUI was up. `app` now describes the target, keyed off the one window
status that requires a live renderer, and the result names its own subject so
the frame can't be misread again. The remote pid is not knowable, so it stays
null.

STA-4792 defect 2 gets a regression test rather than a fix: routing already
made the client remote, which is what stops a Windows destination being joined
to the local cwd. The test pins the exact reported invocation.

* fix(status): share the remote app projection with the SSH host passthrough, and name the version gap on project host setup

Two review follow-ups.

The SSH host passthrough answered `app.running: true` unconditionally for the
Orca host a caller reached over SSH, claiming a desktop app even for a headless
`serve`. That is the same defect as the paired-server path, one transport over,
so the projection moved to shared and both now answer the question the same way.

`--host runtime:<id>` routes project commands to a paired server, which means a
client can reach a server that predates project host setup without meaning to.
That answered a raw `method_not_found`, which reads as an Orca bug rather than a
version gap; the CLI now names it the way the desktop already does.

Reverted a third change: making the persistence duplicate check treat `local`
and `runtime:*` as one machine. That assumption holds at the RPC boundary, where
a `runtime:` host means the runtime being addressed, but not in the store, which
also records independent provisioning metadata for machines that are not itself.
An existing test covers exactly that, and it was right. The duplicate
convergence therefore stays bounded to rows written after the normalization.
2026-08-19 17:12:17 -07:00
Brennan Benson 6415511a82 feat(remote): add a positional file read to the filesystem provider (#15517)
* feat(remote): add a positional file read to the filesystem provider

Following a growing remote file means re-reading it from the top on every
poll: the relay exposes only whole-file reads, so tailing an append-only log
over SSH costs O(size) per tick.

Adds fs.readFileRange plus a rangedReadVersion capability, and an optional
readFileRange on IFilesystemProvider -- matching how lstat/
supportsQuickOpenSearch already declare degradable capabilities.

Three deliberate choices:

- The relay loops until the requested length is satisfied or the file truly
  ends, and REJECTS an over-cap request rather than clamping it. A clamped
  read is indistinguishable from EOF, so a caller advancing a cursor by
  bytesRead would silently skip data.
- Bytes cross the wire base64-encoded. A range boundary can split a UTF-8
  sequence at either edge, and a utf-8 round trip would substitute U+FFFD and
  shift every subsequent offset.
- The provider throws a typed FileRangeReadUnsupportedError against an older
  relay instead of quietly falling back to a whole-file read. A tailing caller
  issues several reads per snapshot, so a per-call fallback is quadratic;
  callers probe supportsFileRangeRead once and snapshot instead.

The response is validated before use -- a byte count disagreeing with the
payload would shift every downstream offset while looking like success.

Terminal-artifact reads/writes move to their own module, mirroring the relay's
existing fs-handler-terminal-artifact split; the provider was at the max-lines
ceiling and this was the cohesive piece to extract.

* fix(remote): size the ranged read to what the relay writer can deliver

The 4 MiB cap was justified against MAX_MESSAGE_SIZE (16 MiB), but that is
the frame DECODER bound. Responses are gated by the writer's admission
budget: a frame over DISPATCHER_CONTROL_QUEUE_MAX_BYTES (1 MiB) is demoted
to the legacy-response lane, which is refused once the producer queue passes
2 MiB. A 4 MiB window is ~5.46 MiB of base64, so it was never admissible --
it came back as an opaque ResponseOverCapacity (-33008), which is neither of
the PR's typed errors, and above ~1.4 MiB the outcome depended on unrelated
queued traffic. Cap at STREAM_CHUNK_SIZE (256 KiB), the house per-frame
budget for file bytes, which stays in the control lane unconditionally.

Also:
- Hoist the cap and offset validation into src/shared/file-range-read.ts so
  the client rejects an out-of-contract request locally instead of paying a
  round trip for an error that does not survive the wire as a type.
- Validate filePath in the relay handler; a missing one threw a TypeError
  out of expandTilde despite the comment claiming hand-validated params.
- Collapse the two fs.getCapabilities probes onto one cached fetch per
  multiplexer. They read one document, so probing per feature spent an extra
  round trip per connection and duplicated the eviction logic.
- Reuse readFullStreamChunk instead of a second copy of the short-read fill.
- allocUnsafe the window; only subarray(0, bytesRead) escapes, so a tailing
  poll no longer memsets the whole window per call.
- Plain methods for readFileRange/supportsFileRangeRead rather than
  constructor-assigned arrows; both are unconditional, unlike downloadFolder.

Tests: cover the dispatch path and fs.getCapabilities (neither was
exercised), param validation at both boundaries, EOF at and past the end,
and a full-cap read over a real RelayDispatcher. The transport guard fails
at 4 MiB with the real -33008.

* test(remote): pin the ranged-read cap to real control-queue headroom

The cap comment claimed a full-cap window stays in the control lane
"unconditionally" and the guard test only asserted one frame fits the
lane, so a raise to 384-768 KiB stayed green while two concurrent
full-cap responses would already overflow the shared control queue --
which for a response closes the client. Pin the two-deep headroom and
state the real bound, including that widening the cap is a wire change
against a host still advertising rangedReadVersion 1.

Also cover the two behaviours the suite claimed but did not exercise: a
regular file answers a full-cap read in one syscall, so the fill loop
was untested (both mutations of readFullStreamChunk stayed green), and
the merged capability document made the abort-does-not-evict guard
load-bearing without any test reaching it.

* fix(remote): harden ranged-read validation and retry
2026-08-19 16:45:42 -07:00
Neil 3ffab9a6b3 feat(terminal): read the rendered screen with terminal read --screen (STA-4792) (#15380)
* feat(terminal): read the rendered screen with `terminal read --screen` (STA-4792)

`terminal read` returns accumulated pty output with escape sequences stripped.
That is the right answer for "what happened over time" and the wrong one for
"what is on screen": any program that repaints a line comes back as stacked
fragments, so one `clear` typed key by key reads as `cclclecleaclear`, and a
prompt that draws a space by moving the cursor loses it. Nothing in the output
said which question had been answered, so it was used as rendering evidence and
produced false conclusions.

The runtime already knew how to render — it replays the byte stream through a
headless emulator — but only as a fallback for blank reads, alternate screen,
and never-attached ptys. A normal attached terminal never reached it. `--screen`
asks for it directly.

Every read now reports its source, which also surfaces the pre-existing
snapshot fallback that until now swapped rendered lines into an ordinary read
with no indication. `screen-unavailable` distinguishes "asked for a screen,
none could be rendered, here is the stream" from a stream the caller asked for,
and an absent source means the host predates the field. Because an older host
strips the unknown param and answers with its ordinary read, `--screen` against
one fails with that explanation rather than passing the stream off as a screen.

`--screen` and `--cursor` are mutually exclusive: a screen is the current frame
and has nothing behind it to page.

* refactor(terminal): stamp the screen source where rendered lines enter the read

Inferring it from tail array identity worked but made a load-bearing contract
out of reference equality; any later path spreading the read would silently
mislabel. Rendered lines only enter through one builder, so it stamps there and
anything still unlabelled is the stream.
2026-08-18 22:51:37 -07:00
Brennan Benson 6386b540ad fix(quick-open): stop node:path from white-screening the renderer (#15405)
* fix(quick-open): stop node:path from white-screening the renderer

#15158 pulled `quick-open-filter` into the renderer graph, but the module
still used `import { posix, win32 } from 'node:path'`. The bundler stubs
`node:path` in the renderer with an object that throws on any member read,
and named bindings resolve at module evaluation — so the renderer threw
before React mounted. `pnpm dev` white-screened on main.

Switch to a namespace import and have `pathFlavor` return the flavor NAME
instead of the module, so `node:path` stays untouched until the out-of-root
`relative` fallback actually needs it. A namespace import alone is not
enough: `pathFlavor` is called unconditionally by `buildExcludePathPrefixes`,
so it still threw on any Quick Open with nested-worktree excludes.

No behavior change for the main process or the relay.

* fix(quick-open): keep exclude containment browser-safe
2026-08-18 18:47:53 -07:00
Brennan Benson c228030516 fix(tabs): keep an open diff focused while agents stream (STA-4697) (#15390)
* refactor(tabs): put the visible-tab-type projection in one place

Three copies of toVisibleTabType had drifted: the runtime one omits 'simulator'.
Move the canonical projection next to the two unions it maps between and replace
the two copies that are already identical to it. The runtime copy is left alone
on purpose - unifying it would change behavior, so it goes with the follow-up.

* fix(tabs): keep an open diff focused while agents stream (STA-4697)

resolveWebSessionVisibleTabId answered 'which tab is the user looking at' by
inverting a many-to-one projection: it compared tab.contentType against the
coarse activeTabType. Diff tabs open with activeTabType 'editor' but carry
contentType 'diff', so the match never succeeded and the guard returned null -
which is the reconciler's signal to fall through and activate a terminal. Every
agent status echo republished the snapshot, so the diff lost focus ~300ms after
opening, once per click. Same for conflict-review and check-details.

Resolve the visible tab from group state instead, which is what is actually on
screen and is the rule deriveActiveSurfaceForWorktree already uses. The coarse
address survives only when there are no group records, now projected rather than
compared exactly.

Also follow the entity within the group when reconcile rematerializes the visible
tab under a new id, and teach the browser-create focus guard to observe the group
records the resolver now reads.
2026-08-18 18:20:56 -07:00
Neil 1b5cc00870 fix(shell): content-address shell wrapper trees so builds stop clobbering each other (#15285)
Every writer sharing a userData dir -- main's local PTY path, the daemon fork,
and the daemons of other builds that outlive the app that spawned them -- wrote
one fixed `shell-ready/` tree. Last writer won, and the guard only re-checked
that the files were present, never that they were this build's. A daemon whose
spawn env no longer agreed with the wrapper on disk kept launching shells it
could not read: the ready marker never fired and every startup command waited
out the full 15s timeout, silently, with restarting the app powerless to fix it
because the daemon outlives the app.

Measured 15.17s vs 1.10s once the tree matched. The five live daemons on the
machine this was found on spanned app versions 1.4.181-1.4.185, all from the
same installed app across auto-updates, so this reaches ordinary installs.

Name each tree after a hash of its contents, at
`<userData>/shell-wrappers/<hash>/shell-ready/`. Different bytes are a different
directory, so "present" means "written by this build" again. The hash sits above
the `shell-ready` leaf because ZDOTDIR self-reference guards match that exact
suffix. Nothing collects old trees: ~48KB each, a couple of MB a year against a
userData dir in the tens of GB, not worth an `rm -rf` on the spawn path.

Also publishes the resolved root to WSL over WSLENV `/p`, since the in-guest
script cannot derive a hash, and reports readiness failures to the daemon's
NDJSON log rather than a console the detached daemon discards.

Generated wrapper content is byte-identical; snapshots unchanged.
2026-08-18 15:59:17 -07:00
Jinwoo Hong 79be5b7fde feat(orchestration): report a worker blocked on a human prompt (STA-4513, STA-3714) (#15261)
* feat(orchestration): report a worker blocked on a human prompt (STA-4513, STA-3714)

A lane parked on an approval, trust, or permission prompt looked exactly like a
lane that was thinking or inside a long tool call. On origin/main, driving a real
cursor-agent through Orca:

  surface                        running `sleep 60`   awaiting approval
  worktree ps agents[].state     working              working
  terminal show / list           no such field        no such field
  terminal wait --for tui-idle   satisfied: true      satisfied: true
  worker-show                    no agent state       no agent state

The runtime already fuses hook state, OSC title, and matched prompt text into a
`permission` verdict inside getTerminalAgentStatus — it was reachable only from the
renderer, and it was blind to cursor-agent approvals. Two gaps, one boundary.

Exposure: getTerminalInteractiveWait publishes that same fusion, minus the async
foreground probe, as `agentWait` on `terminal show` and on `worker-show`'s
observation. It carries the evidence that proved the wait (hook, prompt-text, or
title) so a coordinator can weigh it. Null means no proof; a missing field means
the host predates it — absence is never read as "not waiting".

Detection: cursor-agent's hook set has no approval event and beforeShellExecution
fires identically for auto-allowed commands, so its rendered menu is the only
authority. Matched on the key-bound choices rather than the prose, requiring two,
and self-clearing when the follow-up input line returns. Its live spinner title is
exempted from the staleness rule that clears startup modals, because cursor keeps
spinning while it waits.

Falls out of routing it through the shared verdict: `dispatch --inject` into a
cursor pane on an approval now refuses with agent_prompt_blocked instead of typing
the preamble into the dialog.

Fixtures are captured verbatim from cursor-agent 2026.08.11-e8db854 driven through
Orca; the same case matrix was replayed live against a built runtime.

terminal list stays untouched: its rows would each need a full tail scan, and
STA-4694 owns the one-call-per-run aggregate.

* fix(orchestration): only call a Cursor approval live while it owns the screen

Independent review found the approval detector trusted one dismissal string, so
any later output that did not contain cursor's follow-up line left the menu
reading as a live wait. Reproduced: a tail of the real menu followed by two lines
of ordinary output returned agent-approval-prompt, which fails tui-idle and
refuses prompt injection on a healthy lane.

Replaced with the structural property the string was standing in for: a live
dialog owns the bottom of the screen, so the last choice may sit at most one line
above the end of the retained tail. That tolerates a status footer or a partial
line mid-redraw without admitting scrollback, and it drops the vendor prose.

Being bottom-of-screen is also the dating this reason needed, so it no longer
requires waitBlockedAt. A tail restored from terminal history carries none, and a
lane parked on a prompt emits no bytes — so before this, an Orca restart made
exactly the lane both issues are about go quiet for good. The startup modals keep
the timestamp rule: their text lingers in scrollback with nothing to say whether
it was answered.

Also from review:
- worker-show and federationShow reuse the verdict showTerminal already computed
  rather than rescanning the tail, so the two can no longer disagree.
- The worker-show test now drives a real runtime, real PTY tail, and the real
  detector; it previously mocked getTerminalInteractiveWait, so it would have
  passed with detection permanently returning null.
- The guard claim is now asserted against the guard: a blocked pane rejects both
  assertTerminalAgentSendable and sendTerminalAgentPrompt, and a working pane
  still passes.
- Added a non-local (connectionId) pane case, since the verdict is derived from
  retained tail and title state on every host.

* fix(agent-status): stop a hook wait from outliving its agent

A third reviewer caught that the hook branch proved agent ownership from the pane
title alone, while the shared verdict it claimed to reuse also probes the
foreground process. A shell that takes a pane back usually sets something like
`user@host: ~/repo`, which no title rule recognizes, and a hook row stays fresh
for AGENT_STATUS_STALE_AFTER_MS — so a dead agent could be reported as waiting on
a human for half an hour.

Hook evidence now goes through getTerminalAgentStatus, which is the only thing
that can answer whether an agent still owns this PTY. The two prompt branches skip
it: a matched prompt is on the pane's screen now, so it proves itself. That makes
the probe cost fall exactly where correctness needs it, and getTerminalInteractiveWait
async, which only showTerminal had to absorb.

Also trims the comments the same reviewer flagged as longer than the repo's rule.

* test(agent-status): pin that a dead pane stops reporting a human wait

A fourth reviewer noted the approval menu sits at the bottom of a dead pane's tail
forever, and that no test covered process exit with no trailing output. The
snapshot already refuses an exited pane, and worker-show gates agentWait on proven
identity — this pins both so neither can drift into reporting a worker that needs
intervention as one that needs an answer.

* fix(orchestration): never report an unchecked worker as not waiting

Automated review caught that the three worker paths which return before the wait
is ever evaluated — unattached, missing, and identity_changed — then had their
undefined coerced to null by the emitters. A worker whose process was replaced was
reported as `agentWait: null`, which reads as "Orca looked and nobody is waiting"
when Orca never looked. That is the false negative this field exists to remove.

The field is now emitted only when it was evaluated, so a present null is a claim
about the pane and an absent one means nobody looked — because the host predates
the field, or the worker's identity could not be verified. The CLI and the
worker-show note say that rather than blaming an old host.

Covered on the context-only path, where the regression test fails against the
previous behavior; the supervised and federated emitters take the identical
one-line change.

Also trims the two test-file headers to one statement of purpose.

* fix(agent-status): tighten the Cursor menu match and stop guessing on unknowns

Fourth review round, three findings, each reproduced before acting.

Matching each choice marker with an independent lastIndexOf let text outside the
menu carry the anchor. An agent narrating "next time I'll suggest Run Everything"
after the menu was answered pulled the match down to the bottom of the screen and
revived it. The match is now confined to the last lines of the tail, and a choice
is a line that ends in the key that picks it — prose writes the same words but not
the same shape.

The one line of slack under the dialog went with it. It was a guess; every capture
of a live dialog ends on its last choice, and one line is exactly enough room for
that narration. A redraw caught mid-flight now reads as no wait until the next
poll, which is the safe way to be wrong.

The hook branch awaited a foreground probe that reaches a PTY controller which may
be a remote host, so a wedged probe stalled every caller of showTerminal — a path
that never probed before. It is bounded now, and a timeout leaves the wait
unevaluated rather than claiming there is none.

Which is the same distinction the previous commit only fixed one level up:
getTerminalInteractiveWait itself turned an unreadable pane into `null`, so
showTerminal published "looked, nobody waiting" for a pane it could not read. It
returns undefined there, showTerminal omits the key, and worker-show's text output
prints unknown rather than rendering it the same as none.

* fix(agent-status): bound the wedged probe's cost and stop matching prose keys

Fifth review round. No correctness defects in the shipped behaviour this time; two
robustness holes and the documentation of the contract.

The bounded probe abandoned the wait but not the request, so a coordinator watching
a wedged remote host added one live probe on every poll. It is single-flighted per
PTY now, the way the leaf-absence probe already is.

The trailing-key rule that separates a menu row from the agent narrating a choice
was written as a character class, and any lowercase run up to twelve characters
satisfied it — "…suggest Run Everything (as before)" passed. Spelled out as key
names instead, which also lets the glyph forms of those keys through.

The contract wording said an absent agentWait meant an old host or an unverifiable
identity. It also covers an unreadable pane and a probe that did not answer, and a
reader diagnosing an old peer from that would be wrong. Corrected on the type, the
worker-show note, and in docs/reference/remote-wire-compatibility.md, which had no
entry for a field whose absent and null states mean different things.

Also strengthens the worker-show agreement test, which compared the terminal and
observation payloads without asserting either held the expected wait, so it passed
when both were absent.
2026-08-18 14:19:20 -07:00
Brennan Benson 0b80a773a4 fix(codex): stop overwriting and deleting Codex files that were merely unreadable (STA-4737) (#15287)
* fix(codex): stop overwriting and deleting Codex files that were merely unreadable (STA-4737)

Three modules shared by the host and WSL Codex lanes decided a file was absent
from a read that had only failed, and then wrote over it or removed it.

- `codex-config-mirror`: `existsSync` on the RUNTIME config.toml returned false
  for a locked file exactly as for an absent one, so the mirror took the
  "seed a fresh runtime config" branch and replaced the user's config wholesale.
- `config-settings-promotion`: an unreadable ~/.codex/config.toml counted as
  having no promoted settings, and the write path then rebuilt the user's
  canonical Codex config from Orca's runtime copy.
- `codex-home-paths`: both delete branches in `linkSystemCodexResource` remove
  Orca's mirrored copy because the system resource "is not there". `existsSync`
  and `systemResourceIsRegularFile`'s `catch { return false }` both reported
  that for a source nobody could read, so one denied read on ~/.codex/AGENTS.md
  removed the managed copy on the next launch.

`src/shared/definitive-filesystem-absence.ts` now owns the one errno allowlist —
ENOENT and ENOTDIR, with every other code including unrecognised ones treated as
indeterminate — and `host-codex-managed-home-ownership.ts` drops its private
copy rather than letting the two drift. `codex-path-observation.ts` builds the
three-valued observation on top of it.

The resource sync's two `existsSync`/`statSync` probes collapse into one
resolved stat, which answers reachability and regular-file-ness together and
closes the window between them.

`config-settings-promotion.ts` crossed its max-lines budget, so the write-target
resolution moves to its own module rather than taking a lint exemption.

Deliberately not here: the hook-service trust writes that run after a refused
mirror, and the promotion write target's own classification, which is
unreachable because it always resolves to the same file the read above already
refused. Both are noted in comments rather than half-built.

* fix(codex): preserve resource copies on indeterminate reads
2026-08-18 14:10:32 -07:00
Jinwoo Hong a77a2f93f7 fix(remote): search Quick Open paths on the host (#15158) 2026-08-18 13:32:48 -07:00
Jinwoo Hong bef76953d7 fix(orchestration): record why a terminal's process is gone (STA-4603, STA-4536) (#15244) 2026-08-18 13:17:42 -07:00
Brennan Benson 2a760e310b fix(computer): report unasserted accessibility actions (#15028)
* fix(computer): report unasserted accessibility actions

* fix(computer): fail closed on missing action metadata

* Fix merged tab search test fixture
2026-08-18 11:29:26 -07:00
Brennan Benson 2f0f9a8a39 Revert "fix(agent-hooks): bind agent status to the pane its session was spawned into (STA-2069) (#14615)" (#15295)
Reverts #14615. Its premise does not reproduce, it does not reach the failure that does, and the correction it installs can misattribute status on a path that worked before.

1. PREMISE FALSE. #14615 asserts Claude Code >= 2.1.206 hosts TUI sessions under a shared daemon. On 2.1.233 `claude daemon status` reports "not running" with 69 live interactive sessions, and every client is a direct child of its own pane's shell. Measured across the fleet: 68 distinct pane keys, zero collisions. Foreground attribution was never broken.

2. DOES NOT FIX THE REAL BUG. The failure in #9236 is real but scoped to BACKGROUNDED sessions, whose workers inherit the dispatching pane's whole ORCA_* set. #14615 mints a binding only for launches Orca constructs, so a typed `claude --bg` produces none. Fixed properly in #15304.

3. INTRODUCES A MISATTRIBUTION. Bindings are removed only on PTY death, and a user who exits Claude keeps the pane's PTY. Resuming that session in another pane does not rebind (`--resume` is a session selector, so the pin declines), and resolveBoundPaneOverride then rewrites paneKey and tabId onto the ORIGINAL pane despite a correct posted key. Demonstrated with a failing test against main; causation isolated to resolveBoundPaneOverride.

Kept #14706's observations.rebind() in the conflicting hunk — it postdates #14615 and is not part of this revert.
2026-08-18 03:20:48 -07:00
OrcaWinandOrcaWin b7f2e17712 fix(git): fence buffered WSL login-shell reads (#15060)
* fix(git): fence buffered WSL login-shell reads

Two git paths force the login shell unconditionally and buffer its whole
stdout, so the distro's rc banner lands in front of the payload:

- gitExecFileAsyncBuffer backs `git show :<path>` blob reads and hands
  the bytes straight to the diff/blob viewer, so the banner is prepended
  to displayed file content.
- buildNetworkSshPolicyEnv probes `core.sshCommand` and treats any
  non-empty answer as a user-configured wrapper. A banner reads as
  configured, so the code skips the `ssh -o BatchMode=yes` fallback and
  silently disarms the guard that keeps non-interactive SSH from
  hanging on a prompt.

Fencing is opt-in per call site rather than applied to the login-shell
branch as a whole: streaming consumers (`git grep`, `ls-files -z`) parse
records as they arrive, so an opening marker would be glued onto their
first record. Only these two, both buffered by construction, opt in.

Blob content can be binary, so the payload is sliced out of the raw
bytes; decoding to find the fence would corrupt it. The markers are
exposed on the captured command because the shared module is bundled for
the renderer and cannot reference Buffer.

Note this path is not a rare fallback: `preferWslDirectGit` is only set
by gitStatusReadOptionsForWorktree, so every other WSL-routed git call
takes the login shell on every invocation.

* test(git): use findLast for the ssh-policy call lookup

Satisfies the code-quality rule that flags filter-then-index.

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-18 02:40:06 -07:00
OrcaWinandOrcaWin 3a9f40ed70 fix(wsl): read machine output from a fenced login shell (#15290)
* fix(wsl): read machine output from a fenced login shell

Orca runs WSL reads through the distro's *interactive* login shell so
PATH matches the user's own terminal (nvm, mise and asdf only install
into rc files interactive shells read). An interactive shell also runs
the distro's rc/motd, and stock Ubuntu 24.04 writes its "run a command
as administrator" hint to stdout -- no user customization required.

Every caller parsing that stream was reading the banner as data:

  statPath  -> "To run a command as administrator...\n\ndirectory"
  readPath  -> banner prepended to the contents of every file read
  preflight -> banner prepended to `gh --version` / auth output

`.trim()` cannot recover any of these, so a WSL worktree's file
explorer sees no valid entry types and file reads return junk.

Three call sites had independently grown their own marker to survive
this (`__ORCA_AGENT_PATH__`, `ORCA_WSL_GIT_READ_ENV_V1`, and a
`>/dev/null` fd dance), which is the tell that it belongs in one place.

Fence the payload once, in the shared builder, and hand callers a
reader that returns just their bytes. The fence carries a per-call
nonce so `cat`-ing a file that happens to quote a marker is not
truncated. Exit status is preserved, so the ENOENT mapping still works.

wsl-git-read-environment drops its bespoke marker and parsing.

* test(wsl): fence the login-shell path-lookup boundary test

It asserted a raw interactive login-shell read matched an absolute path,
so the distro rc banner made it fail on any stock Ubuntu. It is part of
the shell-contracts CI gate, where it skips on Linux and hid the break.

* docs(wsl): record the guest command-execution contract

Both failure modes are silent - the command runs, exits 0, and returns
the wrong bytes - so the rules need to live somewhere a reader will
find them before writing the next wsl.exe call site.

* fix(codex): fence the WSL Codex identity probe

buildWslCodexBinaryStamp reads the login shell's stdout positionally --
path before the first newline, version after -- through an interactive
login shell. On a stock Ubuntu the rc banner lands ahead of the payload,
so the first newline falls inside the banner and the stamp becomes
path="To run a command as administrator..." with the rest as version.
Both halves are non-empty, so nothing throws: the stamp is silently
wrong, and an unstable stamp reads as "the Codex binary changed" and
reissues the trust grant.

The identity script ends in `exec`, so it never writes a closing fence;
the reader returns everything after the opening one, which is exactly
this case.

buildWslCodexIdentityArgs becomes buildWslCodexIdentityProbe and returns
the reader with the argv so the two cannot drift apart. The other three
WSL Codex commands are deliberately left unfenced: availability is
exit-code only, and app-server/login hand stdout to a long-running
program.

* fix(wsl): harden the capture fence after review

- readStdout now takes the LAST opening fence, matching the lastIndexOf
  the wsl-git-read-environment marker used deliberately: a login shell
  can echo the command text before running it, repeating the fence.
- local-worktree-filesystem throws instead of falling back to raw stdout
  when the fence is missing. The fallback silently reinstated the bug
  being fixed -- statPath would return the banner as a file type and
  readPath would return banner+contents, with no signal. Preflight keeps
  its fallback; its matchers scan the whole blob and tolerate a prefix.
- The exit-status test asserted only that the script CONTAINS `exit $?`,
  which is true for any input and never executed those lines. It now
  runs a real distro and asserts status 2 reaches the caller, which is
  what statPath's ENOENT mapping depends on.
- Corrected the doc: a sed backreference has no `$`, so `--` never
  rewrote it. Replaced with the positional and shell-local cases that
  were measured to differ.

* fix(wsl): stop running a login shell for filesystem reads

statPath/readPath/rm run coreutils at standard paths and shell builtins.
They need nothing from the user's PATH, so there was never a reason to
start a login shell -- and starting one is what put the distro's rc/motd
on the stdout these callers parse.

Fencing that output treated the symptom. Using a plain `sh -c` removes
the cause: no profile, no rc, no banner, by construction. The fence and
its missing-fence error go away with it.

The fence stays where it is actually needed: the three places that must
run the user's shell to resolve their PATH (the preflight CLI probe, the
WSL git environment probe, and the Codex identity probe).

Net -12 lines.

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-18 02:30:02 -07:00
OrcaWinandOrcaWin 6e8da1df8d fix(wsl): pass guest argv verbatim through --exec (#15039)
* fix(wsl): pass guest argv verbatim through --exec

`wsl.exe <...> -- <argv>` expands `$name` in every argument against the
guest environment before the guest ever runs. It does this even when no
shell is involved, so `-- /usr/bin/printf %s '$HOME'` prints /home/you.

Every WSL invocation went through that preprocessor, so scripts arrived
already rewritten: `awk '{print $2}'` lost its field reference, and a
POSIX script asking for the literal `$HOME` got the expanded path.

`escapeWslShCommandForWindows` tried to compensate by escaping `$`, but
it skipped any `$` preceded by a backslash, so a script containing `\$`
was still corrupted -- and half the call sites never applied it at all.

Route every invocation through `--exec`, which passes argv through
untouched, and delete the escaper. The direct-git path already used
`--exec`, so this is not a new compatibility dependency.

A guard test fails if the `--` form reappears anywhere in the tree.

Net -50 lines of production code.

* test(wsl): drop remaining escaped-dollar assertions

* fix(wsl): cover the --exec migration's blind spots

An audit of every wsl.exe invocation found sites the first pass missed,
including two it actively broke:

- config/scripts/wsl-git-shell-benchmark.mjs imported
  escapeWslShCommandForWindows, which no longer exists, so the script
  threw on startup. Its wslShellArgs helper also still used `--`; the
  file already had an --exec helper, so route both call sites there.
- classifySubprocessCommand unwrapped `wsl.exe <...> -- <binary>` by
  breaking on `--` alone. With every Orca spawn now on --exec it never
  found the guest binary and bucketed all WSL subprocesses as plain
  "wsl", losing the git/gh/glab breakdown. Break on either separator,
  since foreign wsl.exe processes still use `--`.

CliSkillRuntimeSetup builds its setup command as a template literal
rather than an argv array, so no array-shaped search could see it. Its
decoder accepts both separators so commands persisted before this
change still decode.

The guard now scans config/ and tests/ as well as src/, and checks the
command-string spelling alongside the argv one — the two shapes that
have each shipped a regression. It skips comment lines so prose about
the old form stays allowed, and asserts it scanned a plausible file
count so a bad root cannot make it vacuous.

* fix(wsl): restore the guard's multi-line sensitivity

The guard matched line by line, so `'--',\s*'bash'` could not span a
newline -- and every argv array in this repo is formatted one element
per line, which is exactly the shape it exists to catch. Measured
against the pre-migration tree it caught 17 files before and 9 fewer
after. It now strips comment lines and matches the rejoined text, with
a case that pins the multi-line shape so this cannot silently return.

The program list is wider than shells now, which surfaced a false
positive: tmux takes a `--` separator followed by a program too
(`split-window ... -- cat`). Matching is scoped to files that mention
WSL rather than narrowing the list back.

Also:
- Replaced the `sed` regression case, which was vacuous. A backreference
  contains no `$`, so it returned `bac` under both separators and would
  have passed without the fix. The block claimed every case proved the
  bug. Swapped in a positional argument and a shell local, both measured
  to differ -- the positional is the shape `wslUncDirectoryExists` uses,
  where `--` blanked `$1` so every existing directory probed as missing.
- windows-shell-args.test.ts derived its expected argv from
  buildWslExecArgs, the helper under test, so six assertions would still
  pass if it regressed to `--`. Spelled the expectation out.
- Dropped two comments citing the removed `--` behavior as rationale.

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-18 02:12:21 -07:00
Brennan Benson 7ba336099c fix(sidebar): name each split-pane agent row from its own pane title (STA-2811) (#14707)
Agent rows are per pane, but their conversation name came from `tab.title`,
which carries only the FOCUSED pane's title. In a split tab every row showed
one pane's name, and all of them changed when the user clicked a sibling.

Rows on a multi-pane tab now resolve their own leaf's runtime pane title via
the existing `resolveRuntimePaneTitleForLeaf`, and fall back to no live title
rather than a sibling's. Single-pane tabs pass `undefined` and are unchanged.

Extracts the subagent grouping out of build-dashboard-snapshot.ts, which was
exactly at the 300-line cap.
2026-08-18 00:47:14 -07:00
Brennan Benson 26bdfc0fe4 feat(agent-status): stamp observation provenance at every status ingress (STA-4293) (#14706)
Add an optional `observation` facet to agent status rows recording the origin
(hook | osc | title | process | launch | orchestration), the authority that
sequenced it, a per-pane incarnation, a monotonic revision, and the authority's
own clock. Stamp it at every ingress; no consumer reads it.

Boundary is stamped from the hook listener's existing per-provider
`isNewTurnEvent`, not a second list of event-name literals. Identity-only
(`providerSessionOnly`) rows are tagged `kind: 'identity-only'` so future
consumers do not each rediscover that they are not turn transitions.

The staleness-decay contract is documented at the type: staleness must be
computed against the same authority clock that stamped `observedAt`, or
replicas must decay on local receipt time. Not fixed here.

Behavior-neutral: optional field on existing JSON, never persisted, never
published to paired clients, and never inherited across writes.
2026-08-18 00:46:41 -07:00
Neil 9b1f0373eb fix(relay): scope shell history for Windows -> WSL panes (STA-4682) (#15236)
`injectRelayHistoryEnv` matched only bash*/zsh*, so a relay pane launched
through `wsl.exe` got no HISTFILE at all and every WSL worktree shared one
global history.

The history file stays on the relay host under the existing flat root, so
`deleteRelayHistory` remains the deletion counterpart unchanged; only the
exported path is translated to drvfs for the guest, and WSLENV carries it
across the boundary.

Guest fish stays out of scope on purpose: its history file lives inside the
distro, where the relay has no deletion path.
2026-08-17 22:39:02 -07:00
Neil 0e96b82e44 fix(mobile): keep phone tab selection across host snapshots
* fix(mobile): keep phone tab selection across host snapshots

Preserve device-owned tab focus across ordinary host republications while explicit follow navigation remains authoritative. Retire closed selections across clients so stale snapshots cannot resurrect tabs.

* fix(mobile): acknowledge session tab closes

* fix(mobile): avoid tombstones for uncommitted closes

* fix(web): implement session close IPC stubs

* refactor: simplify mobile tab close flow

* fix: bound session tab close confirmation
2026-08-17 18:26:52 -07:00
Brennan Benson 6ee265e579 feat(agent-status): surface the model each Codex subagent is running (#8251) (#14627)
Codex child rows have carried a model field end-to-end since #9637, but the
transcript reader never populated it, so every transcript-discovered child
rendered with an empty model chip. Read the child's own turn_context.model
from the rollout records already fetched for completion detection, so the
sidebar can distinguish an orchestrator model from a subagent model.

No added file I/O and no added rows: the model is parsed from records the
reconcile pass already read, and both row components already render
entry.model.
2026-08-17 17:26:52 -07:00
Brennan BensonandBrian Dai 15efc87e35 fix(agent-hooks): bind agent status to the pane its session was spawned into (STA-2069) (#14615)
* fix(agent-hooks): bind agent status to the pane the session was spawned into (STA-2069)

Claude Code >= 2.1.206 hosts TUI sessions as workers under a shared daemon,
and the daemon forwards only its own allowlisted env — so hook posts carry
whichever pane first started the daemon, not the pane the user is in.

Pin a minted --session-id at spawn where Orca still knows the pane, record
sessionId -> pane, and correct the posted key at both hook ingest seams.

Co-authored-by: Brian Dai <43929761+BrianDai22@users.noreply.github.com>

* fix(agent-hooks): pin the session id in root-option position, not appended

Appending `--session-id <uuid>` broke every `claude <subcommand>` launch:
`--session-id` is a ROOT option, so `claude mcp list --session-id <uuid>`
exits with "error: unknown option '--session-id'". Splice it immediately
after the executable token instead, which is valid for both a bare session
and a subcommand, and is already before claude's own `--` terminator.

Also write the binding-key separator as an escape rather than a raw
NUL byte, which made the file a binary blob in git.

Close three hunks that no test could fail on: the pty.ts spawn call site
that records the binding, the relay seam's worktreeId override, and the
already-correct-pane early-return that suppresses a worktree restamp.

---------

Co-authored-by: Brian Dai <43929761+BrianDai22@users.noreply.github.com>
2026-08-17 16:38:59 -07:00
Brennan Benson 9f4ea42493 fix(agent-status): say what an OpenCode permission request is waiting on (STA-3160) (#14614)
* fix(agent-status): say what an OpenCode permission request is waiting on (STA-3160)

A permission.asked arrives as hook_event_name PermissionRequest, but
extractOpenCodeToolFields had no branch for it, so the pane reported a bare
{state:'waiting'} with no tool or command. The user could see that OpenCode was
blocked but not on what.

Read the fields @opencode-ai/sdk fixes for EventPermissionAsked: 'permission'
names the request, and metadata/patterns carry the command or paths it covers.
The normalizer is shared with mimo-code, so both are covered.

* fix(agent-status): show the OpenCode permission on the row, and retire it after (STA-3160)

Live validation against opencode 1.18.18 showed the original change populated
toolName/toolInput on a `waiting` entry that no surface rendered, while leaving
the answered permission cached for the rest of the pane's session.

Retire the tool fields on every OpenCode-family event except PermissionRequest.
isNewTurnEvent is false for this family, so resolveToolState otherwise inherits
one answered permission onto every later frame and the row reads a resolved
command as the live tool. Reproduced end to end: after approving `rm -rf build/`,
an unrelated later turn still reported it.

Read `filepath` from permission metadata. The SDK types metadata as
Record<string, unknown>, so its keys come from each tool; a live opencode 1.18.18
sends `filepath` (one word) for `edit`, which the previous key list missed. The
fallback to `patterns` covered it by accident, and the test that claimed to cover
it used `metadata: {}` — a shape OpenCode never emits. Tests now use captured
payloads for bash, edit and webfetch.

Show tool fields on `waiting` as well as `working`. All three consumers gated on
`working`, so a permission request rendered nothing at all; before/after of the
sidebar was pixel-identical. The rule now lives in one place (showsAgentToolPreview)
because a gate duplicated across three surfaces is a gate that drifts.
2026-08-17 16:37:07 -07:00
Brennan BensonandQA 64de8dd637 fix(workspaces): delete on the confirmed host, and make both hosts' rows selectable (STA-4343) (#15013)
* fix(workspaces): host-qualified workspace deletion (STA-4343, STA-4448)

Squashed integration of PR #14606 + the codex review-loop output, replayed
onto current main. Granular history preserved on brennanb2025/sta-4343-review-full.

Fixes the regression from #13413: a workspace id is repoId::path with no host
component, so the same repo at the same path on two hosts published one id for
two workspaces, and deletion routed by that id landed on whichever host routing
preferred - usually the ACTIVE one, not the row the user confirmed.

- removeWorktree takes a REQUIRED host-qualified WorktreeRemovalTarget; omitting
  the host is a type error. All destructive callers migrated.
- Projections dedup on (host, id), so two hosts render as two selectable rows
  while the createWorktree/fetchWorktrees race duplicate still collapses.
- Ephemeral VM cleanup is host-scoped. It matched on bare workspaceId, so the
  host-scoped delete path destroyed the SURVIVING host's VM and its unpushed
  filesystem - a leak fix that had become data destruction.
- Selection, keyboard routing, lineage grouping and Space rows carry host
  identity end to end; fixing the executor dedupe alone would have turned
  one-row intent into deleting both hosts.

Files split to stay under max-lines rather than raising any cap.

* refactor: split files that crossed max-lines

The review-loop commits used --no-verify, so the pre-commit hook never
enforced the caps. Extracted cohesive units rather than raising any limit:
renderer teardown, delete-with-toast, pinned-group rows, host-scope helpers,
workspace-kind predicates, filter actions, kanban drag selection, the
renderer removal result type, and the native-chat persistence tests.

* refactor(workspaces): extract cleanup deletion-phase selector

Clears the last max-lines violation and the import-type side effect the
changed-code gate flagged.

* refactor(sidebar): track the delete-dialog extraction modules

* fix(workspaces): preserve host identity across remaining surfaces

* fix(sidebar): re-carry host through the rewritten palette result model

#15170 replaced PaletteSearchResult while this PR was open. Re-applied the
host qualification on top of the new model instead of taking either side:
results carry worktreeHostId again, and the board filter keys its matched
set on host identity rather than the bare id.

Known gap, documented in the board test rather than deleted: searchWorktrees
resolves evidence through a `documents` map keyed by BARE worktree id, so two
same-id host rows collapse before this code sees them. Closing that belongs
with the palette work.

* test(cmd-j): pin the palette collision gap instead of asserting the old model

The palette collision test asserted two host-qualified rows, which #15170's
rewrite made unreachable: item ids are bare again and worktreeMap is id-keyed.

Rewritten to assert what holds — activation always names a host — and to pin
the defect it exposes: two same-id rows render on ONE command value, so React
sees duplicate keys and a click on the first row activates the second row's
host. That reproduces on main, so it is pre-existing, not from this PR. Pinned
rather than deleted so fixing it must update this test.

---------

Co-authored-by: QA <qa@local>
2026-08-17 15:57:07 -07:00
Neil 1412ae2d91 Revert "fix(terminal): inset the grid inside the xterm surface (#14583)" (#15181)
This reverts commit 226cf88ba6.
2026-08-17 15:38:42 -07:00
Brennan Benson 619ee2cc90 fix(agent-hooks): detect IDS-truncated hook POSTs instead of failing open silently (STA-2870) (#14625) 2026-08-17 15:15:14 -07:00
Brennan Benson 32ee3b0536 reland(browser): route every cookie-import write through CDP identities, and never clear what it will not write back (#15030)
* reland(browser): restore CDP-identity cookie-import writes (#14729)

Reverts the revert 3c8410d927 (#14942) to restore the reviewed bf6dc6fcba
tree. The defect that caused that revert is NOT fixed by this commit — it is
fixed in the commits that follow, so the delta a reviewer must scrutinise
stays small instead of hiding inside a 2000-line re-add.

Conflict resolutions, all keep-both:
- browser-cookie-import.ts: two hunks around the zero-import early return.
  #14683's undecryptable warning and decryptedCookies.length === 0 condition
  are kept alongside the reland's old-client gate and partitionSkippedCookies.
  The old-client gate stays above the early return, and so above any mutation.
- en.json: both string sets (#14683's undecryptable copy and partitionSkipped).
- crashpad-capture.test.ts: took HEAD wholesale; unrelated to this ticket.

getStoragePath, and its positive-write assertion moves from cookies.set to the
CDP identity store — adapted, not weakened, and now strictly stronger because
it also asserts cookies.set carries no imported user data. Every decryption
assertion is untouched.

* fix(browser): add registrableFamily, one definition of a cookie family

STA-4300, change 1 of the reland fixes. No behaviour change yet — this only
adds the helper the later commits derive every skip scope from.

Deriving "family" inline in several places is what let the removal scope and
the write set disagree in STA-4090 and STA-4170, so there is exactly one.

The IP check runs on normalizeCookieDomain's OUTPUT, never the raw string.
psl treats an IPv4 literal as a dotted DNS name — psl.parse('127.0.0.1').domain
is '0.1' — and Chromium accepts many spellings of one address. new URL() inside
normalizeCookieDomain canonicalises 127.1 / 2130706433 / 0x7f.1 / 010.0.0.1 and
a trailing dot to a dotted quad first, so isIP() then catches all of them.
Mutation-proved: moving the check ahead of normalisation reddens the five
alternate-spelling cases and nothing else.

Bracketed IPv6 needs its own branch because isIP('[::1]') is 0; without it the
value falls through to psl, which throws, which returns the host — the right
answer for the wrong reason.

Returns null for a bare public suffix: naming `com` as a family would preserve
an entire TLD from removal and silently turn an import into a no-op.

* fix(browser): plan every cookie's fate before any jar mutation

STA-4300, change 2. Adds planImportWrites — pure, no I/O — so the write set and
every removal scope can derive from one value instead of being computed twice.
Not yet wired into the import paths; that is the next commits.

TWO passes, and the second is not optional. Family-atomic skip is a property of
the whole input: with a readable mixed.example row BEFORE an unreadable
sub.mixed.example row, a per-row guard emits the readable one before anything
knows the family will be skipped. Pass 1 classifies and collects the skipped
families; pass 2 re-filters the provisional writes.

Mutation-proved with a faithful one-pass guard (consult skippedFamilies as you
go): the readable-before-unreadable case goes red and the unreadable-before-
readable case still passes. That asymmetry is why both orders are tested and
why only the first is the named detector.

Family closure is at the registrable boundary because that is the boundary the
removal scope actually expands to: importedDomainScopes turns an imported
mixed.example into descendant roots, so a readable apex cookie drags a skipped
subdomain's live session into the removal scope (STA-4300 §2b).

hasUnrepresentableSkip surfaces a skip whose family cannot be named. A family we
cannot name is one we cannot exclude from the removal plan, and clearing a
family we cannot protect is the P0 this ticket exists to stop — so the caller
refuses before mutating rather than proceeding and hoping.

* fix(browser): derive path A's removal scope from the write plan (STA-4300 §2b)

Change 3. importValidatedCookies now plans before it opens the jar, and the
replacement scope and the write set are the SAME array.

bf6dc6fcba filtered replacementDomains per exact cookie
(partition.status !== 'unreadable'). That is not sufficient:
replaceCookiesForImportedDomains expands each imported domain into its
descendant roots, so a readable cookie on mixed.example pulls sub.mixed.example
into the removal scope — and a skipped cookie living there has its session
removed with nothing written back. Same erasure as the P0, different path.

plan.writes is family-closed, and because path A passes the very same array to
replaceCookiesForImportedDomains and to writeImportedCookies, the two sets
cannot drift apart. That is stronger than keeping them in sync: there is only
one set.

Also here, both before any mutation:
- an unrepresentable skip (registrableFamily → null) refuses the import, because
  a family we cannot name is one we cannot exclude from the removal scope;
- the old-client gate now keys off plan.skips rather than re-deriving the
  condition.

Counters: partitionSkippedCookies is a BREAKDOWN of skippedCookies — the
unreadable rows plus their family-suppressed siblings — added in exactly once,
so totalCookies === importedCookies + skippedCookies keeps holding. Counting it
separately is how a summary silently stops adding up.

Full src/main/browser suite: 69 files, 729 tests, green.

* fix(browser): split path B into scan and emit, and never stage a skip (STA-4300)

Change 4 — this is the commit that fixes the shipped P0.

bf6dc6fcba did:
    decryptedCookies.push({ ..., partition })   // write set built HERE
    if (partition.status === 'unreadable') { ... continue }   // guard AFTER

so an unreadable row discovered late could not retract a sibling already
emitted, while removeTransplantableCookies cleared the ENTIRE jar. The mutation
set was the whole jar and the write set a strict subset of it, which is how a
mixed readable/unreadable source emptied a populated jar and repopulated only
part of it.

Now the row loop SCANS only. Nothing is emitted inside it — not decryptedCookies,
not domainSet, not a staging row, not the imported count. Between scan and emit,
planImportWrites closes the skip over whole registrable families, and emit walks
the plan. Each scanned candidate carries its raw source row because
buildChromiumCookieInsertParams needs it; a record holding only the derived
fields compiles and then silently cannot stage.

Also between scan and emit, both before any mutation:
- an unrepresentable skip refuses the import;
- any skipped family calls disableStaging. A staged image is a whole-database
  replacement on the next start, so it cannot express "preserve this family";
  main's existing memoryFailed arm then reports restart-fallback-unavailable.

Test contract change, deliberately stronger: "never stages a partitioned row
whose ancestor bit is unreadable" asserted that an image WAS registered and
merely omitted the row. That image would still have erased the preserved family
on replay. It now asserts no image is registered at all — plus a new paired case
proving a no-skip import still stages, so "disable on skip" cannot be
implemented as "disable always" unnoticed.

Honest note on detection: reverting the emit filter currently reddens only the
staging assertion, because every end-to-end fixture in this module still starts
with an EMPTY jar — where clear-then-write-all and clear-then-write-some are
indistinguishable. That is the blind spot that let this ship. The populated-jar
suite in the next commit is what actually detects the erasure.

Full src/main/browser: 69 files, 730 tests, green.

* fix(browser): never remove a family the import declined to write (STA-4300 I2)

Change 5, and the one that makes preservation observable.

removeTransplantableCookies takes preserveFamilies. removableCookieEntries
filters those families out, which keeps them out of the removal plan AND — since
the CDP snapshot is taken from that same list — out of the restore set, so a
preserved coordinate is never submitted to any mutation at all.

When anything is preserved, the bulk clearData path is not used. clearData
removes everything outside excludeOrigins, and its own comment concedes a
rejection may already have emptied part of the jar; a preserved family is
deliberately absent from the snapshot, so a partial delete followed by a
rejection would destroy it with no identity able to restore it. Handing a
dynamically derived preserve list to that primitive cannot be made safe, so a
skip-bearing clear runs the frozen per-coordinate plan — which already excludes
those families — as its primary path instead. Nothing is preserved on an
ordinary import, so that path keeps today's single clearData call unchanged.

New populated-jar suite. Every end-to-end fixture in this module starts with
cookies.get returning [], and against an empty jar "clear then write all" and
"clear then write some" are indistinguishable — which is why four independent
gates missed the P0. These fixtures start populated.

Mutation-proved, and this is the first detector in the change set that catches
the actual erasure rather than a proxy:
- drop the preserved-family filter          → 6 of 8 red
- use bulk clearData while preserving       → 6 of 8 red, including
  "leaves a preserved family untouched", which is the erasure itself

IPv4-literal and single-label host families are covered explicitly: psl reads
127.0.0.1 as the dotted DNS name '0.1', so a wrong family here would fail to
match the preserve set and erase the live loopback session.

Full src/main/browser: 70 files, 738 tests, green.

* test(browser): pin STA-4300 end to end in real Electron, both write paths

Ports the pre-implementation reproduction into the PR. Verified RED against
main's product code and GREEN with the fix, on both paths — I re-ran it both
ways rather than assuming.

Two assertions are INVERTED rather than deleted, and both are now stronger. The
repro proved the defect by showing cookies.set was called WITHOUT a partitionKey;
after the reland the import write path does not touch cookies.set for user data
at all, so seeing the imported cookie there is itself the regression. The
fixture's internal guard is inverted the same way.

Non-vacuity, asserted in-run rather than in a report: ok:true with a nonzero
importedCookies, the importer reaching its terminal step, the source really
carrying its partition fields, and a CONTROL cookie written via CDP *with* a
partitionKey and read back with it — so "partitionKey absent" cannot be confused
with "the oracle cannot see partitions". electronVersion proves Electron ran.

Fixture correction worth recording, because the original reproduction was
invalid on this path and I did not catch it on first read: readJsonCookiePartition
reads a NESTED partitionKey object, but the fixture wrote topLevelSite and
hasCrossSiteAncestor at the entry's top level, where the importer never looks.
The file/paste case therefore failed on main for a fixture-shape reason rather
than for the defect — and its own non-vacuity check validated what the fixture
wrote instead of what the parser reads, which is exactly how a check that looks
rigorous proves nothing. Only the native half was ever a real reproduction.

Partition keys use a schemeful SITE, not an origin: Chromium canonicalises
https://app.example.com to https://example.com, measured via CDP.

* fix(browser): fail closed on lossy cookie identity recovery

* fix(browser): preserve unreadable families before decrypt

* fix(browser): reject malformed cookie partition sites

* fix(browser): validate recovered cookie partition sites

* fix(browser): reject empty JSON partition identities

* fix(browser): fail closed and serialize cookie imports

* fix(browser): reject malformed CDP opaque flags

* fix(browser): roll back outcome-unknown cookie writes

* revert(browser): move import serialization out to STA-4601

Reverts only the concurrency-serialization half of f7c27b71ab so #15030 stays
scoped to the STA-4300 reland. That commit mixed two concerns; this removes one
of them and keeps the other.

REMOVED (moves to STA-4601, a separate pre-existing P1):
- acquireCookieMutationLock / withCookieMutationLock and the mutationLocks map;
  clear.ts goes back to withCookieClearLock
- the import-wide lock in path A (mutationOwner on the target, acquire/release
  around replace + writes + rollback)
- the widened lock around path B's clear + memory writes
- browser-cookie-import-concurrency.test.ts

KEPT (genuinely STA-4300 scope, not concurrency):
- partitionKeyOpaque handling in readJsonCookiePartition and the CDP snapshot.
  An opaque partition key cannot be represented faithfully, so it reads as
  unreadable and its family is preserved — exactly the rule this ticket adds.
  194e57a628's refinement of that shape stays too.

The concurrent-import interleaving is real and reachable (nothing serialises
imports per partition, and the clear lock is released before the memory writes),
but it predates this change and is not caused or worsened by it. It gets its own
PR and its own review rather than riding a P0 reland whose history is that every
additional fix introduced a new defect.

src/main/browser: 71 files, 772 tests, green. Electron oracle still passes on
both write paths and still goes red against main's product code.
2026-08-17 12:44:39 -07:00
Brennan Benson 60805f5c45 fix(agent-status): preserve restored child provenance (#15082)
* fix(agent-status): preserve restored child provenance

* fix(agent-status): preserve restored completion context

* fix(agent-status): retain child boundary across OSC
2026-08-17 11:10:38 -07:00
Brennan Benson 7ae6aedc02 fix(codex): stop a transient filesystem error from logging out the active account (#15046)
* fix(codex): stop a transient filesystem error from logging out the active account

A single unreadable read of a managed Codex home's ownership marker cleared the
user's active account selection, permanently. On Windows any exclusive lock —
Defender real-time scanning, a backup agent, a sync client — makes every read of
that marker fail with EBUSY, and the background rate-limit poll runs every 15
minutes plus once at every app start.

Root cause: the ownership gate answered two very different questions through one
channel. "This home is not ours" (a successful observation that failed a trust
check) and "we could not read it" both surfaced as a throw, which the caller
flattened to null, which three call sites took as proof the home was
untrustworthy and wrote activeCodexManagedAccountId: null.

Refusing to USE an unverified home is correct. Erasing the user's account
selection because a file was briefly locked is not.

The gate now returns a tri-state verdict. `untrusted` comes only from a proven
trust failure or a definitive ENOENT/ENOTDIR where absence is itself the
verdict; every other filesystem exception is `indeterminate`. Only `untrusted`
may touch persisted state.

Because `null` already meant "fall through to the system default" on both the
launch and poll paths, not-clearing on its own would have run a DIFFERENT
account behind a UI still showing the selected one. So the refusal needed real
channels rather than a sentinel:

- the poll returns an explicit skip; returning null would not have skipped at
  all, since the fetcher maps null to ~/.codex and would have spawned a
  token-refreshing app-server inside the user's real credential home
- pane launch throws a typed temporary-unavailability error that both PTY
  implementations convert into a clean refusal with a retry message, including
  the re-resolution after the async auth-readiness wait
- automatic session resume resolves the selected home eagerly, so an unreadable
  account can no longer be silently replaced by another one in the ranking
- config-sync status reports a distinct managed-home-unavailable stall instead
  of "synced", with a bounded renderer retry so it clears on its own

Also fixes the ticket's second symptom. The status bar's Sign in button called a
re-auth that captured the selection before login and restored it after, so
re-authenticating a deselected account restored `null` — a successful login that
left the account inactive, with no success toast to distinguish it from failure.
It now activates the account it just signed in, but only when the pre-login
selection was empty, so it cannot silently switch accounts for multi-account
users, and it runs the same restart prompt an explicit switch does.

No retry or grace window inside the synchronous gate: it runs on the Electron
main process in a loop over accounts, so a sleep there would freeze the UI.
Recovery is simply the next readable evaluation.

The WSL lane has the same class of defect, including one path that deletes a
credential mirror. It is pre-existing, unreachable from these host code paths,
and deliberately left for its own change; the host clearing sites cannot reach a
WSL account because getSelfContainedManagedHostAccount excludes them.

Fixes STA-4422

* test(codex): cover pending reset home ownership
2026-08-17 02:19:57 -07:00
Brennan Benson 1aa51f6914 fix(computer): preserve accessibility value types (#15031)
* fix(computer): preserve AX value types

* fix(computer): preserve exact integer values

* fix(computer): preserve exact integer exponent forms
2026-08-17 01:20:31 -07:00
Brennan Benson 6387e5b8d3 fix(folder-workspaces): keep the broken-folder marker when a host sends a new reason (#15027)
FolderWorkspacePathStatus is cast, not decoded, off the runtime RPC wire --
runtime-rpc-envelope declares result: z.unknown(), so unwrapRuntimeRpcResult hands
back whatever the host sent. Both title and description switch on status.reason with
no runtime guard, so a newer host publishing a fifth reason matched nothing and
returned undefined. FolderPathStatusIndicator's `!title` check then dropped the whole
indicator, and a broken folder workspace rendered as healthy -- worse than the blank
toast #15002 fixed, because there the warning was empty and here it is gone.

Guard before each switch, the shape #15002 landed. A default: arm is not available:
the type-aware config sets allowDefaultCaseForExhaustiveSwitch:false and rejects one
with switch-exhaustiveness-check.

Extract that guard into isHandledWireDiscriminant instead of hand-writing a third and
fourth copy, and move #15002's two bespoke guards onto it. It takes unknown and checks
typeof before Object.hasOwn -- hasOwn coerces its key, so a host that widened the field
to an array sends ['missing'], which a hasOwn-only guard admits before the switch drops
it straight back out. That was the P1 found in review on #15002; one implementation
makes it structural instead of tribal.

An unrecognized reason gets its own copy rather than reusing 'unavailable'. The
unavailable remedy -- "Check the runtime or SSH connection and try again" -- is a false
lead here: the host did check and reported the folder unusable, so retrying and
inspecting a healthy connection wastes the user's time. Update Orca is the real remedy.

Adding a fifth reason still fails typecheck in two places: TS2741 on the Record and
TS2366 plus switch-exhaustiveness-check on both switches.
2026-08-17 01:01:28 -07:00
Jinwoo Hong 7c798907c5 fix(skills): harden cross-host bundle installs (#15000) 2026-08-17 00:55:04 -07:00