Commit Graph
1613 Commits
Author SHA1 Message Date
Jinwoo Hong bef76953d7 fix(orchestration): record why a terminal's process is gone (STA-4603, STA-4536) (#15244) 2026-08-18 13:17:42 -07:00
Brennan Benson 2a760e310b fix(computer): report unasserted accessibility actions (#15028)
* fix(computer): report unasserted accessibility actions

* fix(computer): fail closed on missing action metadata

* Fix merged tab search test fixture
2026-08-18 11:29:26 -07:00
Brennan Benson 2f0f9a8a39 Revert "fix(agent-hooks): bind agent status to the pane its session was spawned into (STA-2069) (#14615)" (#15295)
Reverts #14615. Its premise does not reproduce, it does not reach the failure that does, and the correction it installs can misattribute status on a path that worked before.

1. PREMISE FALSE. #14615 asserts Claude Code >= 2.1.206 hosts TUI sessions under a shared daemon. On 2.1.233 `claude daemon status` reports "not running" with 69 live interactive sessions, and every client is a direct child of its own pane's shell. Measured across the fleet: 68 distinct pane keys, zero collisions. Foreground attribution was never broken.

2. DOES NOT FIX THE REAL BUG. The failure in #9236 is real but scoped to BACKGROUNDED sessions, whose workers inherit the dispatching pane's whole ORCA_* set. #14615 mints a binding only for launches Orca constructs, so a typed `claude --bg` produces none. Fixed properly in #15304.

3. INTRODUCES A MISATTRIBUTION. Bindings are removed only on PTY death, and a user who exits Claude keeps the pane's PTY. Resuming that session in another pane does not rebind (`--resume` is a session selector, so the pin declines), and resolveBoundPaneOverride then rewrites paneKey and tabId onto the ORIGINAL pane despite a correct posted key. Demonstrated with a failing test against main; causation isolated to resolveBoundPaneOverride.

Kept #14706's observations.rebind() in the conflicting hunk — it postdates #14615 and is not part of this revert.
2026-08-18 03:20:48 -07:00
OrcaWinandOrcaWin b7f2e17712 fix(git): fence buffered WSL login-shell reads (#15060)
* fix(git): fence buffered WSL login-shell reads

Two git paths force the login shell unconditionally and buffer its whole
stdout, so the distro's rc banner lands in front of the payload:

- gitExecFileAsyncBuffer backs `git show :<path>` blob reads and hands
  the bytes straight to the diff/blob viewer, so the banner is prepended
  to displayed file content.
- buildNetworkSshPolicyEnv probes `core.sshCommand` and treats any
  non-empty answer as a user-configured wrapper. A banner reads as
  configured, so the code skips the `ssh -o BatchMode=yes` fallback and
  silently disarms the guard that keeps non-interactive SSH from
  hanging on a prompt.

Fencing is opt-in per call site rather than applied to the login-shell
branch as a whole: streaming consumers (`git grep`, `ls-files -z`) parse
records as they arrive, so an opening marker would be glued onto their
first record. Only these two, both buffered by construction, opt in.

Blob content can be binary, so the payload is sliced out of the raw
bytes; decoding to find the fence would corrupt it. The markers are
exposed on the captured command because the shared module is bundled for
the renderer and cannot reference Buffer.

Note this path is not a rare fallback: `preferWslDirectGit` is only set
by gitStatusReadOptionsForWorktree, so every other WSL-routed git call
takes the login shell on every invocation.

* test(git): use findLast for the ssh-policy call lookup

Satisfies the code-quality rule that flags filter-then-index.

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-18 02:40:06 -07:00
OrcaWinandOrcaWin 3a9f40ed70 fix(wsl): read machine output from a fenced login shell (#15290)
* fix(wsl): read machine output from a fenced login shell

Orca runs WSL reads through the distro's *interactive* login shell so
PATH matches the user's own terminal (nvm, mise and asdf only install
into rc files interactive shells read). An interactive shell also runs
the distro's rc/motd, and stock Ubuntu 24.04 writes its "run a command
as administrator" hint to stdout -- no user customization required.

Every caller parsing that stream was reading the banner as data:

  statPath  -> "To run a command as administrator...\n\ndirectory"
  readPath  -> banner prepended to the contents of every file read
  preflight -> banner prepended to `gh --version` / auth output

`.trim()` cannot recover any of these, so a WSL worktree's file
explorer sees no valid entry types and file reads return junk.

Three call sites had independently grown their own marker to survive
this (`__ORCA_AGENT_PATH__`, `ORCA_WSL_GIT_READ_ENV_V1`, and a
`>/dev/null` fd dance), which is the tell that it belongs in one place.

Fence the payload once, in the shared builder, and hand callers a
reader that returns just their bytes. The fence carries a per-call
nonce so `cat`-ing a file that happens to quote a marker is not
truncated. Exit status is preserved, so the ENOENT mapping still works.

wsl-git-read-environment drops its bespoke marker and parsing.

* test(wsl): fence the login-shell path-lookup boundary test

It asserted a raw interactive login-shell read matched an absolute path,
so the distro rc banner made it fail on any stock Ubuntu. It is part of
the shell-contracts CI gate, where it skips on Linux and hid the break.

* docs(wsl): record the guest command-execution contract

Both failure modes are silent - the command runs, exits 0, and returns
the wrong bytes - so the rules need to live somewhere a reader will
find them before writing the next wsl.exe call site.

* fix(codex): fence the WSL Codex identity probe

buildWslCodexBinaryStamp reads the login shell's stdout positionally --
path before the first newline, version after -- through an interactive
login shell. On a stock Ubuntu the rc banner lands ahead of the payload,
so the first newline falls inside the banner and the stamp becomes
path="To run a command as administrator..." with the rest as version.
Both halves are non-empty, so nothing throws: the stamp is silently
wrong, and an unstable stamp reads as "the Codex binary changed" and
reissues the trust grant.

The identity script ends in `exec`, so it never writes a closing fence;
the reader returns everything after the opening one, which is exactly
this case.

buildWslCodexIdentityArgs becomes buildWslCodexIdentityProbe and returns
the reader with the argv so the two cannot drift apart. The other three
WSL Codex commands are deliberately left unfenced: availability is
exit-code only, and app-server/login hand stdout to a long-running
program.

* fix(wsl): harden the capture fence after review

- readStdout now takes the LAST opening fence, matching the lastIndexOf
  the wsl-git-read-environment marker used deliberately: a login shell
  can echo the command text before running it, repeating the fence.
- local-worktree-filesystem throws instead of falling back to raw stdout
  when the fence is missing. The fallback silently reinstated the bug
  being fixed -- statPath would return the banner as a file type and
  readPath would return banner+contents, with no signal. Preflight keeps
  its fallback; its matchers scan the whole blob and tolerate a prefix.
- The exit-status test asserted only that the script CONTAINS `exit $?`,
  which is true for any input and never executed those lines. It now
  runs a real distro and asserts status 2 reaches the caller, which is
  what statPath's ENOENT mapping depends on.
- Corrected the doc: a sed backreference has no `$`, so `--` never
  rewrote it. Replaced with the positional and shell-local cases that
  were measured to differ.

* fix(wsl): stop running a login shell for filesystem reads

statPath/readPath/rm run coreutils at standard paths and shell builtins.
They need nothing from the user's PATH, so there was never a reason to
start a login shell -- and starting one is what put the distro's rc/motd
on the stdout these callers parse.

Fencing that output treated the symptom. Using a plain `sh -c` removes
the cause: no profile, no rc, no banner, by construction. The fence and
its missing-fence error go away with it.

The fence stays where it is actually needed: the three places that must
run the user's shell to resolve their PATH (the preflight CLI probe, the
WSL git environment probe, and the Codex identity probe).

Net -12 lines.

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-18 02:30:02 -07:00
OrcaWinandOrcaWin 6e8da1df8d fix(wsl): pass guest argv verbatim through --exec (#15039)
* fix(wsl): pass guest argv verbatim through --exec

`wsl.exe <...> -- <argv>` expands `$name` in every argument against the
guest environment before the guest ever runs. It does this even when no
shell is involved, so `-- /usr/bin/printf %s '$HOME'` prints /home/you.

Every WSL invocation went through that preprocessor, so scripts arrived
already rewritten: `awk '{print $2}'` lost its field reference, and a
POSIX script asking for the literal `$HOME` got the expanded path.

`escapeWslShCommandForWindows` tried to compensate by escaping `$`, but
it skipped any `$` preceded by a backslash, so a script containing `\$`
was still corrupted -- and half the call sites never applied it at all.

Route every invocation through `--exec`, which passes argv through
untouched, and delete the escaper. The direct-git path already used
`--exec`, so this is not a new compatibility dependency.

A guard test fails if the `--` form reappears anywhere in the tree.

Net -50 lines of production code.

* test(wsl): drop remaining escaped-dollar assertions

* fix(wsl): cover the --exec migration's blind spots

An audit of every wsl.exe invocation found sites the first pass missed,
including two it actively broke:

- config/scripts/wsl-git-shell-benchmark.mjs imported
  escapeWslShCommandForWindows, which no longer exists, so the script
  threw on startup. Its wslShellArgs helper also still used `--`; the
  file already had an --exec helper, so route both call sites there.
- classifySubprocessCommand unwrapped `wsl.exe <...> -- <binary>` by
  breaking on `--` alone. With every Orca spawn now on --exec it never
  found the guest binary and bucketed all WSL subprocesses as plain
  "wsl", losing the git/gh/glab breakdown. Break on either separator,
  since foreign wsl.exe processes still use `--`.

CliSkillRuntimeSetup builds its setup command as a template literal
rather than an argv array, so no array-shaped search could see it. Its
decoder accepts both separators so commands persisted before this
change still decode.

The guard now scans config/ and tests/ as well as src/, and checks the
command-string spelling alongside the argv one — the two shapes that
have each shipped a regression. It skips comment lines so prose about
the old form stays allowed, and asserts it scanned a plausible file
count so a bad root cannot make it vacuous.

* fix(wsl): restore the guard's multi-line sensitivity

The guard matched line by line, so `'--',\s*'bash'` could not span a
newline -- and every argv array in this repo is formatted one element
per line, which is exactly the shape it exists to catch. Measured
against the pre-migration tree it caught 17 files before and 9 fewer
after. It now strips comment lines and matches the rejoined text, with
a case that pins the multi-line shape so this cannot silently return.

The program list is wider than shells now, which surfaced a false
positive: tmux takes a `--` separator followed by a program too
(`split-window ... -- cat`). Matching is scoped to files that mention
WSL rather than narrowing the list back.

Also:
- Replaced the `sed` regression case, which was vacuous. A backreference
  contains no `$`, so it returned `bac` under both separators and would
  have passed without the fix. The block claimed every case proved the
  bug. Swapped in a positional argument and a shell local, both measured
  to differ -- the positional is the shape `wslUncDirectoryExists` uses,
  where `--` blanked `$1` so every existing directory probed as missing.
- windows-shell-args.test.ts derived its expected argv from
  buildWslExecArgs, the helper under test, so six assertions would still
  pass if it regressed to `--`. Spelled the expectation out.
- Dropped two comments citing the removed `--` behavior as rationale.

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-18 02:12:21 -07:00
Brennan Benson 7ba336099c fix(sidebar): name each split-pane agent row from its own pane title (STA-2811) (#14707)
Agent rows are per pane, but their conversation name came from `tab.title`,
which carries only the FOCUSED pane's title. In a split tab every row showed
one pane's name, and all of them changed when the user clicked a sibling.

Rows on a multi-pane tab now resolve their own leaf's runtime pane title via
the existing `resolveRuntimePaneTitleForLeaf`, and fall back to no live title
rather than a sibling's. Single-pane tabs pass `undefined` and are unchanged.

Extracts the subagent grouping out of build-dashboard-snapshot.ts, which was
exactly at the 300-line cap.
2026-08-18 00:47:14 -07:00
Brennan Benson 26bdfc0fe4 feat(agent-status): stamp observation provenance at every status ingress (STA-4293) (#14706)
Add an optional `observation` facet to agent status rows recording the origin
(hook | osc | title | process | launch | orchestration), the authority that
sequenced it, a per-pane incarnation, a monotonic revision, and the authority's
own clock. Stamp it at every ingress; no consumer reads it.

Boundary is stamped from the hook listener's existing per-provider
`isNewTurnEvent`, not a second list of event-name literals. Identity-only
(`providerSessionOnly`) rows are tagged `kind: 'identity-only'` so future
consumers do not each rediscover that they are not turn transitions.

The staleness-decay contract is documented at the type: staleness must be
computed against the same authority clock that stamped `observedAt`, or
replicas must decay on local receipt time. Not fixed here.

Behavior-neutral: optional field on existing JSON, never persisted, never
published to paired clients, and never inherited across writes.
2026-08-18 00:46:41 -07:00
Neil 9b1f0373eb fix(relay): scope shell history for Windows -> WSL panes (STA-4682) (#15236)
`injectRelayHistoryEnv` matched only bash*/zsh*, so a relay pane launched
through `wsl.exe` got no HISTFILE at all and every WSL worktree shared one
global history.

The history file stays on the relay host under the existing flat root, so
`deleteRelayHistory` remains the deletion counterpart unchanged; only the
exported path is translated to drvfs for the guest, and WSLENV carries it
across the boundary.

Guest fish stays out of scope on purpose: its history file lives inside the
distro, where the relay has no deletion path.
2026-08-17 22:39:02 -07:00
Neil 0e96b82e44 fix(mobile): keep phone tab selection across host snapshots
* fix(mobile): keep phone tab selection across host snapshots

Preserve device-owned tab focus across ordinary host republications while explicit follow navigation remains authoritative. Retire closed selections across clients so stale snapshots cannot resurrect tabs.

* fix(mobile): acknowledge session tab closes

* fix(mobile): avoid tombstones for uncommitted closes

* fix(web): implement session close IPC stubs

* refactor: simplify mobile tab close flow

* fix: bound session tab close confirmation
2026-08-17 18:26:52 -07:00
Brennan Benson 6ee265e579 feat(agent-status): surface the model each Codex subagent is running (#8251) (#14627)
Codex child rows have carried a model field end-to-end since #9637, but the
transcript reader never populated it, so every transcript-discovered child
rendered with an empty model chip. Read the child's own turn_context.model
from the rollout records already fetched for completion detection, so the
sidebar can distinguish an orchestrator model from a subagent model.

No added file I/O and no added rows: the model is parsed from records the
reconcile pass already read, and both row components already render
entry.model.
2026-08-17 17:26:52 -07:00
Brennan BensonandBrian Dai 15efc87e35 fix(agent-hooks): bind agent status to the pane its session was spawned into (STA-2069) (#14615)
* fix(agent-hooks): bind agent status to the pane the session was spawned into (STA-2069)

Claude Code >= 2.1.206 hosts TUI sessions as workers under a shared daemon,
and the daemon forwards only its own allowlisted env — so hook posts carry
whichever pane first started the daemon, not the pane the user is in.

Pin a minted --session-id at spawn where Orca still knows the pane, record
sessionId -> pane, and correct the posted key at both hook ingest seams.

Co-authored-by: Brian Dai <43929761+BrianDai22@users.noreply.github.com>

* fix(agent-hooks): pin the session id in root-option position, not appended

Appending `--session-id <uuid>` broke every `claude <subcommand>` launch:
`--session-id` is a ROOT option, so `claude mcp list --session-id <uuid>`
exits with "error: unknown option '--session-id'". Splice it immediately
after the executable token instead, which is valid for both a bare session
and a subcommand, and is already before claude's own `--` terminator.

Also write the binding-key separator as an escape rather than a raw
NUL byte, which made the file a binary blob in git.

Close three hunks that no test could fail on: the pty.ts spawn call site
that records the binding, the relay seam's worktreeId override, and the
already-correct-pane early-return that suppresses a worktree restamp.

---------

Co-authored-by: Brian Dai <43929761+BrianDai22@users.noreply.github.com>
2026-08-17 16:38:59 -07:00
Brennan Benson 9f4ea42493 fix(agent-status): say what an OpenCode permission request is waiting on (STA-3160) (#14614)
* fix(agent-status): say what an OpenCode permission request is waiting on (STA-3160)

A permission.asked arrives as hook_event_name PermissionRequest, but
extractOpenCodeToolFields had no branch for it, so the pane reported a bare
{state:'waiting'} with no tool or command. The user could see that OpenCode was
blocked but not on what.

Read the fields @opencode-ai/sdk fixes for EventPermissionAsked: 'permission'
names the request, and metadata/patterns carry the command or paths it covers.
The normalizer is shared with mimo-code, so both are covered.

* fix(agent-status): show the OpenCode permission on the row, and retire it after (STA-3160)

Live validation against opencode 1.18.18 showed the original change populated
toolName/toolInput on a `waiting` entry that no surface rendered, while leaving
the answered permission cached for the rest of the pane's session.

Retire the tool fields on every OpenCode-family event except PermissionRequest.
isNewTurnEvent is false for this family, so resolveToolState otherwise inherits
one answered permission onto every later frame and the row reads a resolved
command as the live tool. Reproduced end to end: after approving `rm -rf build/`,
an unrelated later turn still reported it.

Read `filepath` from permission metadata. The SDK types metadata as
Record<string, unknown>, so its keys come from each tool; a live opencode 1.18.18
sends `filepath` (one word) for `edit`, which the previous key list missed. The
fallback to `patterns` covered it by accident, and the test that claimed to cover
it used `metadata: {}` — a shape OpenCode never emits. Tests now use captured
payloads for bash, edit and webfetch.

Show tool fields on `waiting` as well as `working`. All three consumers gated on
`working`, so a permission request rendered nothing at all; before/after of the
sidebar was pixel-identical. The rule now lives in one place (showsAgentToolPreview)
because a gate duplicated across three surfaces is a gate that drifts.
2026-08-17 16:37:07 -07:00
Brennan BensonandQA 64de8dd637 fix(workspaces): delete on the confirmed host, and make both hosts' rows selectable (STA-4343) (#15013)
* fix(workspaces): host-qualified workspace deletion (STA-4343, STA-4448)

Squashed integration of PR #14606 + the codex review-loop output, replayed
onto current main. Granular history preserved on brennanb2025/sta-4343-review-full.

Fixes the regression from #13413: a workspace id is repoId::path with no host
component, so the same repo at the same path on two hosts published one id for
two workspaces, and deletion routed by that id landed on whichever host routing
preferred - usually the ACTIVE one, not the row the user confirmed.

- removeWorktree takes a REQUIRED host-qualified WorktreeRemovalTarget; omitting
  the host is a type error. All destructive callers migrated.
- Projections dedup on (host, id), so two hosts render as two selectable rows
  while the createWorktree/fetchWorktrees race duplicate still collapses.
- Ephemeral VM cleanup is host-scoped. It matched on bare workspaceId, so the
  host-scoped delete path destroyed the SURVIVING host's VM and its unpushed
  filesystem - a leak fix that had become data destruction.
- Selection, keyboard routing, lineage grouping and Space rows carry host
  identity end to end; fixing the executor dedupe alone would have turned
  one-row intent into deleting both hosts.

Files split to stay under max-lines rather than raising any cap.

* refactor: split files that crossed max-lines

The review-loop commits used --no-verify, so the pre-commit hook never
enforced the caps. Extracted cohesive units rather than raising any limit:
renderer teardown, delete-with-toast, pinned-group rows, host-scope helpers,
workspace-kind predicates, filter actions, kanban drag selection, the
renderer removal result type, and the native-chat persistence tests.

* refactor(workspaces): extract cleanup deletion-phase selector

Clears the last max-lines violation and the import-type side effect the
changed-code gate flagged.

* refactor(sidebar): track the delete-dialog extraction modules

* fix(workspaces): preserve host identity across remaining surfaces

* fix(sidebar): re-carry host through the rewritten palette result model

#15170 replaced PaletteSearchResult while this PR was open. Re-applied the
host qualification on top of the new model instead of taking either side:
results carry worktreeHostId again, and the board filter keys its matched
set on host identity rather than the bare id.

Known gap, documented in the board test rather than deleted: searchWorktrees
resolves evidence through a `documents` map keyed by BARE worktree id, so two
same-id host rows collapse before this code sees them. Closing that belongs
with the palette work.

* test(cmd-j): pin the palette collision gap instead of asserting the old model

The palette collision test asserted two host-qualified rows, which #15170's
rewrite made unreachable: item ids are bare again and worktreeMap is id-keyed.

Rewritten to assert what holds — activation always names a host — and to pin
the defect it exposes: two same-id rows render on ONE command value, so React
sees duplicate keys and a click on the first row activates the second row's
host. That reproduces on main, so it is pre-existing, not from this PR. Pinned
rather than deleted so fixing it must update this test.

---------

Co-authored-by: QA <qa@local>
2026-08-17 15:57:07 -07:00
Neil 1412ae2d91 Revert "fix(terminal): inset the grid inside the xterm surface (#14583)" (#15181)
This reverts commit 226cf88ba6.
2026-08-17 15:38:42 -07:00
Brennan Benson 619ee2cc90 fix(agent-hooks): detect IDS-truncated hook POSTs instead of failing open silently (STA-2870) (#14625) 2026-08-17 15:15:14 -07:00
Brennan Benson 32ee3b0536 reland(browser): route every cookie-import write through CDP identities, and never clear what it will not write back (#15030)
* reland(browser): restore CDP-identity cookie-import writes (#14729)

Reverts the revert 3c8410d927 (#14942) to restore the reviewed bf6dc6fcba
tree. The defect that caused that revert is NOT fixed by this commit — it is
fixed in the commits that follow, so the delta a reviewer must scrutinise
stays small instead of hiding inside a 2000-line re-add.

Conflict resolutions, all keep-both:
- browser-cookie-import.ts: two hunks around the zero-import early return.
  #14683's undecryptable warning and decryptedCookies.length === 0 condition
  are kept alongside the reland's old-client gate and partitionSkippedCookies.
  The old-client gate stays above the early return, and so above any mutation.
- en.json: both string sets (#14683's undecryptable copy and partitionSkipped).
- crashpad-capture.test.ts: took HEAD wholesale; unrelated to this ticket.

getStoragePath, and its positive-write assertion moves from cookies.set to the
CDP identity store — adapted, not weakened, and now strictly stronger because
it also asserts cookies.set carries no imported user data. Every decryption
assertion is untouched.

* fix(browser): add registrableFamily, one definition of a cookie family

STA-4300, change 1 of the reland fixes. No behaviour change yet — this only
adds the helper the later commits derive every skip scope from.

Deriving "family" inline in several places is what let the removal scope and
the write set disagree in STA-4090 and STA-4170, so there is exactly one.

The IP check runs on normalizeCookieDomain's OUTPUT, never the raw string.
psl treats an IPv4 literal as a dotted DNS name — psl.parse('127.0.0.1').domain
is '0.1' — and Chromium accepts many spellings of one address. new URL() inside
normalizeCookieDomain canonicalises 127.1 / 2130706433 / 0x7f.1 / 010.0.0.1 and
a trailing dot to a dotted quad first, so isIP() then catches all of them.
Mutation-proved: moving the check ahead of normalisation reddens the five
alternate-spelling cases and nothing else.

Bracketed IPv6 needs its own branch because isIP('[::1]') is 0; without it the
value falls through to psl, which throws, which returns the host — the right
answer for the wrong reason.

Returns null for a bare public suffix: naming `com` as a family would preserve
an entire TLD from removal and silently turn an import into a no-op.

* fix(browser): plan every cookie's fate before any jar mutation

STA-4300, change 2. Adds planImportWrites — pure, no I/O — so the write set and
every removal scope can derive from one value instead of being computed twice.
Not yet wired into the import paths; that is the next commits.

TWO passes, and the second is not optional. Family-atomic skip is a property of
the whole input: with a readable mixed.example row BEFORE an unreadable
sub.mixed.example row, a per-row guard emits the readable one before anything
knows the family will be skipped. Pass 1 classifies and collects the skipped
families; pass 2 re-filters the provisional writes.

Mutation-proved with a faithful one-pass guard (consult skippedFamilies as you
go): the readable-before-unreadable case goes red and the unreadable-before-
readable case still passes. That asymmetry is why both orders are tested and
why only the first is the named detector.

Family closure is at the registrable boundary because that is the boundary the
removal scope actually expands to: importedDomainScopes turns an imported
mixed.example into descendant roots, so a readable apex cookie drags a skipped
subdomain's live session into the removal scope (STA-4300 §2b).

hasUnrepresentableSkip surfaces a skip whose family cannot be named. A family we
cannot name is one we cannot exclude from the removal plan, and clearing a
family we cannot protect is the P0 this ticket exists to stop — so the caller
refuses before mutating rather than proceeding and hoping.

* fix(browser): derive path A's removal scope from the write plan (STA-4300 §2b)

Change 3. importValidatedCookies now plans before it opens the jar, and the
replacement scope and the write set are the SAME array.

bf6dc6fcba filtered replacementDomains per exact cookie
(partition.status !== 'unreadable'). That is not sufficient:
replaceCookiesForImportedDomains expands each imported domain into its
descendant roots, so a readable cookie on mixed.example pulls sub.mixed.example
into the removal scope — and a skipped cookie living there has its session
removed with nothing written back. Same erasure as the P0, different path.

plan.writes is family-closed, and because path A passes the very same array to
replaceCookiesForImportedDomains and to writeImportedCookies, the two sets
cannot drift apart. That is stronger than keeping them in sync: there is only
one set.

Also here, both before any mutation:
- an unrepresentable skip (registrableFamily → null) refuses the import, because
  a family we cannot name is one we cannot exclude from the removal scope;
- the old-client gate now keys off plan.skips rather than re-deriving the
  condition.

Counters: partitionSkippedCookies is a BREAKDOWN of skippedCookies — the
unreadable rows plus their family-suppressed siblings — added in exactly once,
so totalCookies === importedCookies + skippedCookies keeps holding. Counting it
separately is how a summary silently stops adding up.

Full src/main/browser suite: 69 files, 729 tests, green.

* fix(browser): split path B into scan and emit, and never stage a skip (STA-4300)

Change 4 — this is the commit that fixes the shipped P0.

bf6dc6fcba did:
    decryptedCookies.push({ ..., partition })   // write set built HERE
    if (partition.status === 'unreadable') { ... continue }   // guard AFTER

so an unreadable row discovered late could not retract a sibling already
emitted, while removeTransplantableCookies cleared the ENTIRE jar. The mutation
set was the whole jar and the write set a strict subset of it, which is how a
mixed readable/unreadable source emptied a populated jar and repopulated only
part of it.

Now the row loop SCANS only. Nothing is emitted inside it — not decryptedCookies,
not domainSet, not a staging row, not the imported count. Between scan and emit,
planImportWrites closes the skip over whole registrable families, and emit walks
the plan. Each scanned candidate carries its raw source row because
buildChromiumCookieInsertParams needs it; a record holding only the derived
fields compiles and then silently cannot stage.

Also between scan and emit, both before any mutation:
- an unrepresentable skip refuses the import;
- any skipped family calls disableStaging. A staged image is a whole-database
  replacement on the next start, so it cannot express "preserve this family";
  main's existing memoryFailed arm then reports restart-fallback-unavailable.

Test contract change, deliberately stronger: "never stages a partitioned row
whose ancestor bit is unreadable" asserted that an image WAS registered and
merely omitted the row. That image would still have erased the preserved family
on replay. It now asserts no image is registered at all — plus a new paired case
proving a no-skip import still stages, so "disable on skip" cannot be
implemented as "disable always" unnoticed.

Honest note on detection: reverting the emit filter currently reddens only the
staging assertion, because every end-to-end fixture in this module still starts
with an EMPTY jar — where clear-then-write-all and clear-then-write-some are
indistinguishable. That is the blind spot that let this ship. The populated-jar
suite in the next commit is what actually detects the erasure.

Full src/main/browser: 69 files, 730 tests, green.

* fix(browser): never remove a family the import declined to write (STA-4300 I2)

Change 5, and the one that makes preservation observable.

removeTransplantableCookies takes preserveFamilies. removableCookieEntries
filters those families out, which keeps them out of the removal plan AND — since
the CDP snapshot is taken from that same list — out of the restore set, so a
preserved coordinate is never submitted to any mutation at all.

When anything is preserved, the bulk clearData path is not used. clearData
removes everything outside excludeOrigins, and its own comment concedes a
rejection may already have emptied part of the jar; a preserved family is
deliberately absent from the snapshot, so a partial delete followed by a
rejection would destroy it with no identity able to restore it. Handing a
dynamically derived preserve list to that primitive cannot be made safe, so a
skip-bearing clear runs the frozen per-coordinate plan — which already excludes
those families — as its primary path instead. Nothing is preserved on an
ordinary import, so that path keeps today's single clearData call unchanged.

New populated-jar suite. Every end-to-end fixture in this module starts with
cookies.get returning [], and against an empty jar "clear then write all" and
"clear then write some" are indistinguishable — which is why four independent
gates missed the P0. These fixtures start populated.

Mutation-proved, and this is the first detector in the change set that catches
the actual erasure rather than a proxy:
- drop the preserved-family filter          → 6 of 8 red
- use bulk clearData while preserving       → 6 of 8 red, including
  "leaves a preserved family untouched", which is the erasure itself

IPv4-literal and single-label host families are covered explicitly: psl reads
127.0.0.1 as the dotted DNS name '0.1', so a wrong family here would fail to
match the preserve set and erase the live loopback session.

Full src/main/browser: 70 files, 738 tests, green.

* test(browser): pin STA-4300 end to end in real Electron, both write paths

Ports the pre-implementation reproduction into the PR. Verified RED against
main's product code and GREEN with the fix, on both paths — I re-ran it both
ways rather than assuming.

Two assertions are INVERTED rather than deleted, and both are now stronger. The
repro proved the defect by showing cookies.set was called WITHOUT a partitionKey;
after the reland the import write path does not touch cookies.set for user data
at all, so seeing the imported cookie there is itself the regression. The
fixture's internal guard is inverted the same way.

Non-vacuity, asserted in-run rather than in a report: ok:true with a nonzero
importedCookies, the importer reaching its terminal step, the source really
carrying its partition fields, and a CONTROL cookie written via CDP *with* a
partitionKey and read back with it — so "partitionKey absent" cannot be confused
with "the oracle cannot see partitions". electronVersion proves Electron ran.

Fixture correction worth recording, because the original reproduction was
invalid on this path and I did not catch it on first read: readJsonCookiePartition
reads a NESTED partitionKey object, but the fixture wrote topLevelSite and
hasCrossSiteAncestor at the entry's top level, where the importer never looks.
The file/paste case therefore failed on main for a fixture-shape reason rather
than for the defect — and its own non-vacuity check validated what the fixture
wrote instead of what the parser reads, which is exactly how a check that looks
rigorous proves nothing. Only the native half was ever a real reproduction.

Partition keys use a schemeful SITE, not an origin: Chromium canonicalises
https://app.example.com to https://example.com, measured via CDP.

* fix(browser): fail closed on lossy cookie identity recovery

* fix(browser): preserve unreadable families before decrypt

* fix(browser): reject malformed cookie partition sites

* fix(browser): validate recovered cookie partition sites

* fix(browser): reject empty JSON partition identities

* fix(browser): fail closed and serialize cookie imports

* fix(browser): reject malformed CDP opaque flags

* fix(browser): roll back outcome-unknown cookie writes

* revert(browser): move import serialization out to STA-4601

Reverts only the concurrency-serialization half of f7c27b71ab so #15030 stays
scoped to the STA-4300 reland. That commit mixed two concerns; this removes one
of them and keeps the other.

REMOVED (moves to STA-4601, a separate pre-existing P1):
- acquireCookieMutationLock / withCookieMutationLock and the mutationLocks map;
  clear.ts goes back to withCookieClearLock
- the import-wide lock in path A (mutationOwner on the target, acquire/release
  around replace + writes + rollback)
- the widened lock around path B's clear + memory writes
- browser-cookie-import-concurrency.test.ts

KEPT (genuinely STA-4300 scope, not concurrency):
- partitionKeyOpaque handling in readJsonCookiePartition and the CDP snapshot.
  An opaque partition key cannot be represented faithfully, so it reads as
  unreadable and its family is preserved — exactly the rule this ticket adds.
  194e57a628's refinement of that shape stays too.

The concurrent-import interleaving is real and reachable (nothing serialises
imports per partition, and the clear lock is released before the memory writes),
but it predates this change and is not caused or worsened by it. It gets its own
PR and its own review rather than riding a P0 reland whose history is that every
additional fix introduced a new defect.

src/main/browser: 71 files, 772 tests, green. Electron oracle still passes on
both write paths and still goes red against main's product code.
2026-08-17 12:44:39 -07:00
Brennan Benson 60805f5c45 fix(agent-status): preserve restored child provenance (#15082)
* fix(agent-status): preserve restored child provenance

* fix(agent-status): preserve restored completion context

* fix(agent-status): retain child boundary across OSC
2026-08-17 11:10:38 -07:00
Brennan Benson 7ae6aedc02 fix(codex): stop a transient filesystem error from logging out the active account (#15046)
* fix(codex): stop a transient filesystem error from logging out the active account

A single unreadable read of a managed Codex home's ownership marker cleared the
user's active account selection, permanently. On Windows any exclusive lock —
Defender real-time scanning, a backup agent, a sync client — makes every read of
that marker fail with EBUSY, and the background rate-limit poll runs every 15
minutes plus once at every app start.

Root cause: the ownership gate answered two very different questions through one
channel. "This home is not ours" (a successful observation that failed a trust
check) and "we could not read it" both surfaced as a throw, which the caller
flattened to null, which three call sites took as proof the home was
untrustworthy and wrote activeCodexManagedAccountId: null.

Refusing to USE an unverified home is correct. Erasing the user's account
selection because a file was briefly locked is not.

The gate now returns a tri-state verdict. `untrusted` comes only from a proven
trust failure or a definitive ENOENT/ENOTDIR where absence is itself the
verdict; every other filesystem exception is `indeterminate`. Only `untrusted`
may touch persisted state.

Because `null` already meant "fall through to the system default" on both the
launch and poll paths, not-clearing on its own would have run a DIFFERENT
account behind a UI still showing the selected one. So the refusal needed real
channels rather than a sentinel:

- the poll returns an explicit skip; returning null would not have skipped at
  all, since the fetcher maps null to ~/.codex and would have spawned a
  token-refreshing app-server inside the user's real credential home
- pane launch throws a typed temporary-unavailability error that both PTY
  implementations convert into a clean refusal with a retry message, including
  the re-resolution after the async auth-readiness wait
- automatic session resume resolves the selected home eagerly, so an unreadable
  account can no longer be silently replaced by another one in the ranking
- config-sync status reports a distinct managed-home-unavailable stall instead
  of "synced", with a bounded renderer retry so it clears on its own

Also fixes the ticket's second symptom. The status bar's Sign in button called a
re-auth that captured the selection before login and restored it after, so
re-authenticating a deselected account restored `null` — a successful login that
left the account inactive, with no success toast to distinguish it from failure.
It now activates the account it just signed in, but only when the pre-login
selection was empty, so it cannot silently switch accounts for multi-account
users, and it runs the same restart prompt an explicit switch does.

No retry or grace window inside the synchronous gate: it runs on the Electron
main process in a loop over accounts, so a sleep there would freeze the UI.
Recovery is simply the next readable evaluation.

The WSL lane has the same class of defect, including one path that deletes a
credential mirror. It is pre-existing, unreachable from these host code paths,
and deliberately left for its own change; the host clearing sites cannot reach a
WSL account because getSelfContainedManagedHostAccount excludes them.

Fixes STA-4422

* test(codex): cover pending reset home ownership
2026-08-17 02:19:57 -07:00
Brennan Benson 1aa51f6914 fix(computer): preserve accessibility value types (#15031)
* fix(computer): preserve AX value types

* fix(computer): preserve exact integer values

* fix(computer): preserve exact integer exponent forms
2026-08-17 01:20:31 -07:00
Brennan Benson 6387e5b8d3 fix(folder-workspaces): keep the broken-folder marker when a host sends a new reason (#15027)
FolderWorkspacePathStatus is cast, not decoded, off the runtime RPC wire --
runtime-rpc-envelope declares result: z.unknown(), so unwrapRuntimeRpcResult hands
back whatever the host sent. Both title and description switch on status.reason with
no runtime guard, so a newer host publishing a fifth reason matched nothing and
returned undefined. FolderPathStatusIndicator's `!title` check then dropped the whole
indicator, and a broken folder workspace rendered as healthy -- worse than the blank
toast #15002 fixed, because there the warning was empty and here it is gone.

Guard before each switch, the shape #15002 landed. A default: arm is not available:
the type-aware config sets allowDefaultCaseForExhaustiveSwitch:false and rejects one
with switch-exhaustiveness-check.

Extract that guard into isHandledWireDiscriminant instead of hand-writing a third and
fourth copy, and move #15002's two bespoke guards onto it. It takes unknown and checks
typeof before Object.hasOwn -- hasOwn coerces its key, so a host that widened the field
to an array sends ['missing'], which a hasOwn-only guard admits before the switch drops
it straight back out. That was the P1 found in review on #15002; one implementation
makes it structural instead of tribal.

An unrecognized reason gets its own copy rather than reusing 'unavailable'. The
unavailable remedy -- "Check the runtime or SSH connection and try again" -- is a false
lead here: the host did check and reported the folder unusable, so retrying and
inspecting a healthy connection wastes the user's time. Update Orca is the real remedy.

Adding a fifth reason still fails typecheck in two places: TS2741 on the Record and
TS2366 plus switch-exhaustiveness-check on both switches.
2026-08-17 01:01:28 -07:00
Jinwoo Hong 7c798907c5 fix(skills): harden cross-host bundle installs (#15000) 2026-08-17 00:55:04 -07:00
Brennan Benson 7afce2ea41 fix(ssh): stop reporting a confirmed kill when the SSH provider is gone (#14977)
* fix(ssh): stop reporting a confirmed kill when the SSH provider is gone

A detached relay PTY is designed to outlive the provider that addressed it
(it ignores SIGHUP and ships with an unlimited grace), so "the SSH provider
is no longer registered" is lost contact, never evidence the remote process
stopped. Both stop primitives in the PTY controller returned `true` from
that branch, and every caller downstream reported the fabricated success:
the CLI printed "PTY killed.", worker-stop settled the dispatch as stopped,
and — because the stop "succeeded" — the unstopped-PTY gate never ran, so
worktree removal walked straight past a live remote agent.

`kill`/`stopAndWait` now still tombstone the local lease but report an
unconfirmed stop and record why, using the three-verdict vocabulary the
worktree teardown gate already spoke (`live` / `unverifiable` / `exited`),
promoted out of that module into `src/shared/pty-liveness-verdict.ts`.
The close receipt, the CLI wording, worker-stop and the removal gate all
read that verdict instead of inferring an exit from silence.

The same rule fixes the mirror-image defect: the aggregate inventory only
enumerates registered providers, so a dropped relay clears `connected` for
every remote PTY at once. The sweep now separates the provider answering
"absent" (an exit) from no provider being able to answer (lost contact), so
worker-stop stops claiming `exited` from a disconnect.

The `connected` wire field is unchanged in meaning and shape.

* fix(orchestration): apply the same honesty to the federation stop path

The federation host runs its own copy of the worker observation and stop
logic, with the same two defects: `inspectRemoteAttachment` read a dropped
relay's `connected: false` as `exited`, and `federationStop` settled the
dispatch as stopped from a close it never confirmed — relaying a fabricated
success all the way home to the coordinator.

Two guards also had to move so the honest verdict does not become a new
refusal. `federationRead` gated on `status !== 'running'`, which would have
rejected a connected terminal the moment a stop lost contact with it; it now
gates on `status === 'exited'`, which is equivalent for every pre-existing
status given the two guards beside it. Local `workerStop` likewise still
attempts the close when the verdict is `unverifiable` — losing contact is a
reason to report the outcome honestly, never a reason to stop trying.

The show observations now carry the reason alongside the status, so a bare
`unverifiable` is actionable. Both are new optional fields.

* fix(ssh): preserve unconfirmed stop verdicts across consumers

* fix(ssh): use canonical live verdict wording

* fix(ssh): refuse wrong-host teardown verification

* test(orchestration): confirm worker release teardown

* fix(orchestration): negotiate honest worker stop receipts

* fix(agent-teams): fence uncertain teammate respawns

* fix(ssh): avoid duplicate missing-provider teardown

* fix(orchestration): preserve archives across release retries

* fix(ssh): preserve verdicts across synthetic kill exits

* fix(ssh): preserve liveness evidence across teardown

* fix(agent-teams): replace panes only after confirmed stop

* fix(ssh): distinguish host exits from relay loss

* fix(ssh): narrow concurrent inventory verdicts

* fix(orchestration): serve archives after uncertain release

* fix(orchestration): expose unverifiable read liveness

* test(ssh): align liveness assertions with verdicts

* fix(ssh): preserve host scope across inventory failures
2026-08-17 00:11:19 -07:00
Neil 3bb87ff93b reland(shell): one portable Unix startup dialect, with both revert causes fixed (#15018)
* reland: portable startup-shell dialect, with the two revert causes fixed

Relands #14863 (reverted by #14975) with fixes for both regressions the
revert cited.

1. History GC deleted folder-workspace shell history. The live set was built
   from `getAllWorktreeMeta()` alone, but a folder workspace's PTY carries
   `folder:<id>` as its worktree id, so every live folder workspace looked
   orphaned. `getKnownWorktreeIdsForHistoryGc` now unions in
   `getFolderWorkspaces()`. Both consumers — the history-directory prune and
   the fish-history sweep — read that one set, so the fix covers bash, zsh and
   fish history alike. The directory prune had this gap since #1524; #14863
   only widened its blast radius to fish files.

2. A copied Codex resume command aborted under `set -u`. Its leading clear
   statement has to test `$fish_pid`, and that unbound expansion takes the
   whole line — including the agent launch — down with it. Copied text runs in
   a shell Orca never spawned, so nothing can seed that variable first. The
   removal now rides on the agent itself as `env -u`, which needs no shell
   syntax and no expansion. Verified byte-identical under `set -u` in sh,
   bash, zsh, dash, ksh and fish.

   `env` cannot run the `cd` builtin, and a child `cd` would not move the
   agent, so the prefix is placed on the agent rather than on the whole
   `cd … && agent` chain. cmd and PowerShell have no nounset hazard and keep
   their clear ahead of the `cd`, which preserves `cd … && agent` — a failed
   `cd` still cannot launch the agent in the wrong directory.

* fix(history-gc): stop three more paths from deleting live shell history

Found by adversarial review of the reland. All three are the same class as
the bug that caused the revert: a live set that is missing a category of
real workspace, so the GC reads it as orphaned.

1. Profiles. The history root is `userData/terminal-history`, which has no
   profile segment, but the Store the GC consults is per-profile. So after a
   profile switch the live set condemned every other profile's history — and
   fish history, which lands in the user's own fish data dir, is shared by
   every profile on the machine. The live set now unions in the inactive
   profiles' worktrees and folder workspaces, read from their data files. A
   profile whose ids cannot be read reports the empty set rather than one
   that condemns real history.

2. No empty-set guard on the tree scan. `sweepOrphanedFishHistoryFiles`
   refuses an empty live set because it cannot be told apart from a store
   that failed to hydrate; the directory scan, which deletes more, had no
   such guard. A store that fell back to default state would have taken
   every worktree's bash and zsh history with it, across all roots including
   WSL. Four existing tests passed `new Set()` and relied on "empty means
   everything is orphaned" — exactly the behavior being removed — so they
   now pass a real live set.

3. Relay fish history. The relay isolates its history tree under its own
   root but wrote fish history into the shared fish data dir under the
   desktop naming, keyed by the CLIENT's worktree ids. On a machine running
   both Orca and a relay host, the desktop sweep deleted remote sessions'
   history once it went stale. Relay files are now `orca_relay_<hash>`,
   which the sweep's pattern deliberately does not match; the relay still
   deletes them by exact name when the worktree goes away.

* fix(resume): enforce the env-removal invariants instead of documenting them

Both found by adversarial review; both were unreachable from today's callers
and silent if reached, which is exactly how they would survive to a caller
that does reach them.

- A pinned CODEX_HOME and the removal named the same variable, and `env -u`
  strips what the assignment just set — so the agent would have resumed
  against the real home and not found the session. The removal list now
  excludes any name the prefix pins, keeping the assignment authoritative as
  the old `clear…; CODEX_HOME=x agent` ordering did. Same fix in the git-bash
  twin. The PowerShell branch already clears before it assigns, so it was
  never affected.

- Placement was keyed on the platform while the grammar it selects is keyed
  on the shell, so `platform: 'linux'` with `shell: 'powershell'` emitted
  POSIX `env -u` into a PowerShell line. PowerShell now routes to the
  PowerShell builder whatever the host, and the POSIX/cmd split below asks
  the shell rather than the platform.
2026-08-16 23:51:53 -07:00
Jinwoo Hong 77ef6bb9ee fix(terminal): verify agent prompt submission (#14962) 2026-08-16 23:34:18 -07:00
Brennan Benson 8ca4ed945e feat(terminal): report execution host and listing scope in terminal list (#14973)
* feat(terminal): report execution host and listing scope in terminal list

`orca terminal list` returned rows with no host identity and no statement
of what the listing covered, so a scoped listing that saw nothing read as
"nothing exists anywhere" — an agent reported a live remote worker dead.

Each row now carries an optional `executionHostId` derived from the PTY id
(SSH and paired-runtime ids embed their owner), and the result carries an
optional `hostScope` naming the hosts covered and the known hosts skipped.
Both are surfaced in `--json` and in the human-readable CLI output, where
an absent field renders as `unknown` rather than `local`.

Both row builders route through one resolver, so the rule lives in one place.

* fix(terminal): preserve unverifiable host scope

* fix(terminal): fail closed on unverifiable hosts

* test(terminal): name unverifiable scope explicitly

* perf(terminal): keep graph hydration host scans narrow

* fix(terminal): reject blank foreign host owners

* fix(terminal): validate inferred inventory hosts

* fix(terminal): preserve paired folder host scope

* fix(terminal): keep inventory host inference typed

* fix(terminal): disclose paired folder hosts
2026-08-16 22:13:03 -07:00
Neil 71bbab72e1 fix(commit-message): keep Windows paths intact in agent command overrides (#14984)
* fix(commit-message): keep Windows paths intact in agent command overrides

`tokenizeCustomCommandTemplate` applies POSIX backslash-escape rules on every
platform. On Windows `\` is the path separator, so a native absolute path in an
agent command override is silently destroyed:

  C:\Windows\System32\WindowsPowerShell\v1.0\powershell.exe
  -> C:WindowsSystem32WindowsPowerShellv1.0powershell.exe

which is then reported as not found on PATH. The agent *startup* path already
routes Windows shells to the Windows tokenizer, but the commit-message AI path
still calls the generic tokenizer directly, so overrides, extra CLI args and
custom commands there are all affected.

The tokenizer gains an explicit `'escape' | 'literal'` mode rather than reading
`process.platform`, because the same template can be parsed on one host and
executed on another. `'escape'` stays the default, so POSIX behaviour — where
`foo\ bar` is deliberately one token — is unchanged.

`'literal'` is selected only where the command provably runs on native Windows:
a LOCAL target, on win32, with no WSL distro. A WSL target runs a Linux binary
inside the distro, and a remote target runs on a host whose platform this
process cannot see; both keep POSIX escaping.

Fixes #11375

* test: pin the platform decision for literal-backslash parsing

commandBackslashMode is the only place that reads the platform, so it is where
this can be wrong in the direction that matters — applying Windows rules to a
command that will actually run under a POSIX shell. WSL and remote targets are
pinned explicitly; both were previously untested.
2026-08-16 21:12:14 -07:00
Brennan Benson 226cf88ba6 fix(terminal): inset the grid inside the xterm surface (#14583)
* fix(terminal): inset the grid inside the xterm surface (#13252)

Padding X/Y was applied as start-edge container margin, so the cell grid
stayed flush on the trailing edges and a fractional background opacity
stacked a darker gutter around the viewport. Put the setting on .xterm
so FitAddon insets both axes and the themed background fills the pad.

* fix(terminal): normalize padding before fit

* fix(terminal): align stored and fitted padding

* test(terminal): lock padding before fit

* fix terminal padding opacity compositing

* fix live terminal padding backgrounds

* fix(terminal): preserve source-over alpha blending

* fix(terminal): restore WebGL alpha blending

* test(terminal): complete hidden retention pane fixture
2026-08-16 20:49:45 -07:00
Brennan Benson 5e9e38fa75 fix(agent-status): announce Claude turn complete while background work runs (#14580)
* fix(agent-status): announce Claude turn complete while background work runs

Lead Stop/StopFailure already ends the turn, but resolveClaudePaneState keeps
the pane working for subagents, background shells, and session crons. That
erases the working→done edge that mints the completion banner: subagent turns
notify late with a stale body, and shells/crons never notify at all.

Stamp turnCompletedAt on the gated lead Stop, announce immediately from that
row, and pair the later all-clear done to the same end time so it cannot
double-fire or collapse consecutive turns onto the pinned stateStartedAt.

Fixes #13245

* fix(agent-status): suppress stamped turn replays

* fix(agent-status): notify paired clients at turn end

* Fix late-paired completion notification arming

* test(agent-status): make the notification-id test fail on the pre-fix ordering

The stored row inherited the helper's default codex agentType while the event
named claude, so agentSnapshotMatchesExplicitTitle dropped it and
freshStoredAgentStatus was undefined — the assertion held under either side of
the `??`. Name the stored row's agent so the pinned working row survives and
the snapshot-first precedence is what the test actually pins.

* Suppress stamped completion tail replays

* Prevent cross-coordinator title replays

* Preserve stamped tails across fallback signals

* Keep remount replay state while sibling lives

* Scope OSC turn stamp preservation

* Forward paired host completion stamps

* Deduplicate paired completion tails

* Bind paired completion tails to their turn

* Preserve paired completion tail ownership

* Keep paired tail replay state across remounts

* Seed paired recovery without replaying completions

* Seed startup replay and release stale fallback dedupe

* fix(notifications): preserve stamped OSC repaints

* fix(notifications): retain paired client turn boundary

* fix notifications module import safety

* chore: keep main integration focused

* fix: preserve completion re-enable boundary
2026-08-16 20:47:02 -07:00
Brennan Bensonandmanuaudio 886dec1d2a fix(browser): report cookies an import could not decrypt (#14683)
* fix(browser): report cookies an import could not decrypt

Supersedes #13193, which reported only the Windows v20 case.

Nothing distinguished "decryption failed" from "no cookies present". A row that
would not decrypt was folded into the generic `skipped` counter, and a profile
whose rows all failed returned ok:true with importedCookies:0 and no warning —
a green "Imported 0 cookies from Google Chrome." The two situations produce
opposite result shapes and the worse one reported success.

Attribute the cause at the point of failure, while the version prefix is still
in hand, and surface it as one `cookies-undecryptable` warning carrying the
reason. Covers all three known causes rather than one prefix:

- app-bound-encryption: Chrome/Edge 140+ on Windows write `v20`, which only the
  writing browser can unwrap. The version gate is a FORMAT check (`/^v\d\d$/`),
  so v20 passed it and failed inside AES like corruption.
- linux-keyring-unavailable: getLinuxEncryptionKey derived the v11 key from an
  empty password when both secret-tool lookups failed, so it never returned null
  and the "Could not access encryption key" guard was unreachable on Linux.
- unknown: any other cause still warns instead of reporting success.

Deliberately not a hard failure on Linux: Chrome falls back to the "peanuts" v10
key precisely when no keyring exists, so those profiles still import. Pinned by
a regression test.

Refs #13192, #14181

* fix(browser): attribute decrypt failures exactly and gate CBC by version

Review-loop findings on the initial commit, all fixed here.

- CORRECTNESS: v11 rows were attempted with the v10 key when the keyring was
  unavailable. AES-128-CBC is unauthenticated, so a wrong key that yields valid
  PKCS#7 padding was accepted — roughly 1 in 256 per row. Garbage values were
  written into the jar as real cookies, and because those rows counted as
  successes the warning this PR adds could never fire. Key eligibility is now
  explicit per version rather than implicit in key ordering.

- CORRECTNESS: the CBC path returned an empty Buffer for a prefix-only value
  BEFORE checking eligibility. An empty Buffer is truthy, so an ineligible row
  counted as imported and reached the live-jar clear. Eligibility now precedes
  that branch and empty CBC ciphertext is rejected as malformed.

- ACCURACY: a named cause reported the TOTAL failure count, so one v20 row plus
  one corrupt row claimed both failed to app-bound encryption. Counts are now
  exact per cause, with the remainder reported separately and a tie falling back
  to 'unknown'. Exact-count approach carried over from #13193.

- The app-bound copy no longer dead-ends. It names the existing in-app file
  import without describing how to produce the file — Chrome has no native
  decrypted-cookie export, so concrete guidance would send users to an
  extension that can read their whole session jar.

- Direct prefix edge tests carried forward from #13193.

Repo-wide search found no second multi-key unauthenticated-CBC first-success
site, so this pattern was one occurrence rather than a class.

Co-authored-by: manuaudio <manuaudio@users.noreply.github.com>

* fix(pr): preserve split worktree slice

Remove the unrelated rollback of the worktree-slice split and its forbidden max-lines baseline addition from this cookie-import PR.

* fix(browser): match the unknown decrypt reason explicitly

CI's type-aware code-quality gate flagged the reason switch as non-exhaustive:
the 'unknown' member was handled by `default:` rather than matched.

Matching it explicitly keeps the behaviour identical today and makes the gate
enforce the thing that matters — adding a new reason to the union now fails the
switch instead of falling silently into a generic message that would not
describe it.

This gate is separate from `oxlint` and is not covered by running oxlint on the
changed files, which is why it only surfaced in CI.

---------

Co-authored-by: manuaudio <manuaudio@users.noreply.github.com>
2026-08-16 18:43:11 -07:00
Brennan Benson 5e189d6081 feat(browser): add a WebAuthn account picker (#14687)
* fix(browser): prompt for WebAuthn account selection

* fix(browser): scope WebAuthn cancellation to session

* fix(renderer): keep WebAuthn render phase pure
2026-08-16 18:32:20 -07:00
Brennan Benson a324ee20d4 Reset terminal SGR state around restored output (#14700)
* fix(terminal): reset SGR around restored output

* fix(terminal): preserve live replay styling

* fix(terminal): ground dead reattach fallback
2026-08-16 17:08:10 -07:00
Neil 85565a9302 reland(workspace): set project location from the create-worktree host picker (#14965)
* feat(workspace): reland set project location from the create-worktree host picker

Relands #14868 (reverted by #14912) with a fix for the regression that caused the
revert: setting a project location could change the path before Orca used it.

The retarget-after-setup path read the raw store record to find a just-created
setup, because the memoized picker options had not refreshed yet:

    useAppStore.getState().projectHostSetups.find(
      (candidate) => candidate.id === setupId && candidate.setupState === 'ready'
    )

That hand-rolls a second selection path that skips every rule the option builder
applies — repo eligibility, ephemeral-VM and runtime-owned SSH host exclusion,
and the one-setup-per-host dedupe whose own comment notes that
resolveWorkspaceCreationTarget takes the first project+host match and ignores the
rest. So the composer could be retargeted at a setup other than the canonical one
for that host, pointing creation at a different location than the one chosen.

Resolves through buildProjectHostSetupOptions against fresh store state instead,
so the fallback and the steady-state picker agree by construction.

STA-4547

* fix(workspace): sanitize the clone prefill and drop an abandoned set-location

Review follow-ups on this PR.

The "Clone from URL" prefill seeded the field with the verbatim `git remote` URL,
which can embed a PAT (`https://x-access-token:ghp_...@github.com/...`). The
clone then runs on the *target* host, writing that token into its .git/config —
a credential the user never typed into this flow, now readable by anyone on a
shared host. Strip it with the same sanitizer `getProvisionedRootRecipeRepoUrl`
already applies to the ephemeral-VM recipe URL. Extracted to
resolveProjectCloneUrlPrefill so the rule is directly testable.

The dialog also stays dismissable while a submit is in flight, and an SSH clone
is unbounded. A clone the user backed out of minutes earlier still called
onReady, silently moving the run target and resetting start-from under a form
they had since pointed at another host. Drop the result if the dialog went away.

* fix(workspace): re-arm the abandoned guard on mount

StrictMode runs mount/cleanup/mount, so latching `abandoned` on the first
cleanup left it true for the rest of the session and permanently suppressed
onReady — the app wraps its root in StrictMode. Reset it on mount.
2026-08-16 16:50:58 -07:00
Neil 3f58d5cf9a fix(daemon): bound cwd validation per UNC route (#14967)
Async cwd validation dedupes by exact path but had no concurrency bound, and a
dead UNC share answers `stat` in ~21s while holding one of libuv's 4 default fs
threads. Four distinct paths on one unreachable server therefore starved every
other async fs read in the daemon — including the cold-restore history replay
running alongside them — which moves the head-of-line stall #14848 removed from
the event loop into the thread pool.

Adds a per-route lane of 2, reusing PrioritySemaphore and matching the per-distro
lane in rate-limits/auth-filesystem-operation.ts. Keyed by the host that has to
answer (WSL distro, or the `\\server` prefix) so many dead subdirectories of one
share fold into one lane. Local-disk paths bypass the lane entirely: a global cap
would queue a healthy local spawn behind a dead share.

Moves PrioritySemaphore to src/shared. It has no imports, and reaching into
src/main/daemon from src/main/providers inverted the dependency direction that
already runs daemon -> providers.

Note this bounds pool occupancy, which cancellation cannot: an aborted `stat`
still holds its libuv thread until the OS returns.

STA-4543
2026-08-16 16:50:26 -07:00
Jinjing f070033156 Revert "refactor(shell): one portable Unix startup dialect instead of shell d…" (#14975)
This reverts commit b6ea3f17a9.
2026-08-16 16:48:55 -07:00
Brennan Benson 8e0dad8e21 fix(crash-reporting): decode POSIX wait statuses in crash-report display and record OS session ends (#14659)
* fix(crash-reporting): decode POSIX wait statuses for display and record Windows session-end reasons

Chromium on POSIX hands render/child-process-gone the raw waitpid() status, so
crash reports read "Exit code: 61696" where exit status 241 is meant (field:
61696=exit 241, 9=SIGKILL, 133=SIGTRAP+core, darwin crashed 5=SIGTRAP). Decode
at the display layer only: the stored exitCode stays raw, Windows codes and
launch-failed launch-error codes render unchanged, and the process-gone span
gains a crash.exit_code_decoded attribute.

Also durably record a system_session_end breadcrumb (with WindowSessionEndEvent
reasons) when Windows session-end fires, so bundles can tell OS shutdown from a
user task-kill in killed/exit-1 sweeps.

* test(crash-reporting): pin the exit(0) no-suffix rendering

Adversarial mutation review: removing the exit-0 suppression in
formatCrashReportExitCode survived the suite — nothing pinned that a clean
exit(0) renders without an '(exit status 0)' suffix.

* test(crash-reporting): decode-attribute tests use synchronous child kills

Renderer killed events gain a 250ms sibling-kill settle once the
correlation branch lands, which (a) defers the span past the test's
platform stub so the decode gate reads the real host platform, and (b)
adds a deferred span that breaks the exact sink assertion. The decode
gate is source-agnostic, and a non-recoverable child kill persists
synchronously on every branch of the stack, so coverage is unchanged and
the platform stub is deterministic on any CI host.

* fix(crash-reporting): keep session-end reasons type-safe
2026-08-16 15:45:38 -07:00
Neil b6ea3f17a9 refactor(shell): one portable Unix startup dialect instead of shell detection (#14863)
Orca had to guess which shell would parse a queued command line, then emit
syntax for it. Guessing is unreliable for a remote or WSL host, and every
dialect-dependent function is a place to get it wrong.

Replace the guess. Everything emitted for a Unix shell is now built to be
correct in sh, bash, zsh, dash, ksh and fish alike, so no detection is needed:

- quoteStartupArg emits backslashes as "\\" and apostrophes as "'" between
  single-quoted runs. Both families read that identically, unlike the sh '\''
  idiom, which fish silently halves and which makes a trailing backslash a
  hard syntax error.
- clearEnvCommand emits a self-contained fish/sh branch. It deliberately does
  NOT call a helper defined by Orca's shell wrappers: Orca wraps only zsh, bash
  and fish, so an `sh`/`dash`/`ksh` login shell launches unwrapped — and the
  same text is copied to the clipboard and pasted into shells Orca never
  spawned. In both, a helper would be `command not found`, which is the exact
  failure this exists to avoid. Two guarded statements rather than `A && B ||
  C`, because fish's `set -e` returns non-zero for an already-unset variable
  and would fall through to the sh branch; a trailing `true` pins the status,
  since this is the last statement of a launch line and the prompt renders it.
- One tokenizer for Unix. The input is a settings string the shell never
  parses, so parsing it per-shell only made the same setting mean different
  things in different workspaces.

AgentStartupShell loses its 'fish' and 'unix' members, and the three
login-shell resolvers, the fish tokenizer and the agentEnv.SHELL probe go with
them.

Per-worktree shell history now actually works:

- zsh on macOS was a no-op. /etc/zshrc assigns HISTFILE unconditionally before
  any wrapper Orca controls, so the injected value was already gone — and with
  ZDOTDIR still pointing at Orca's wrapper dir, history landed inside it. The
  intended path rides ORCA_HISTFILE and is restored after user config.
  Fixes #11044.
- fish keeps history in its own data dir keyed by session name, since it
  ignores HISTFILE and has no custom-directory knob. Files are deleted rather
  than truncated, a symlinked ~/.local/share no longer disables cleanup, and a
  GC sweep reclaims orphans whose meta.json is gone. The sweep refuses an empty
  live-worktree set (indistinguishable from a store that failed to hydrate) and
  skips files younger than GC_MIN_AGE_MS, mirroring the tree GC's guard against
  the live-set snapshot race.

Verified against real shells rather than asserted as strings:
startup-shell-portability.live-shell.test.ts runs 194 assertions across
sh/bash/zsh/dash/ksh/fish, and zsh-scoped-histfile.live-shell.test.ts drives a
real login zsh through /etc/zshrc. Both are vacuity-checked. The same quoting
corpus was replayed byte-exact on Linux, where /bin/sh is dash.
2026-08-16 15:28:50 -07:00
Jinwoo Hong fa9b20cb41 feat(skills): reland private bundle sharing safely (#14934) 2026-08-16 13:45:54 -07:00
Neil 9f3a912c1e fix(terminal): type Option-composed ASCII instead of reporting it as a chord (#14743)
* fix(terminal): preserve Option-composed ASCII input

* fix(terminal): preserve Option keyboard protocol semantics

* fix(terminal): complete Option keyboard event encoding

* fix(terminal): harden Option input encoding

* fix(terminal): close keyboard protocol fallback gaps

* test(terminal): prove Option-composed ASCII reaches the pty end to end

The Option-compose fix had unit coverage only. This drives a live Electron
pane whose kitty flags are armed by the application's own CSI > 1 u and
asserts the bytes at the pty boundary: composed `@` and Shift-layer `\`
arrive as text, configured Option-as-Alt still reports the layout-resolved
chord, and a non-ASCII glyph still reaches the app as its alt hotkey.
Restoring the pre-fix policy fails exactly the two composed-text scenarios.

Also records the ASCII rule's rationale where the rule lives, not only in a
test comment.

* refactor(terminal): drop the unread Option layers from the layout snapshot

The native helper computed an Option and Option+Shift character for every
key, shipped both over IPC, validated them in the parser and cached them in
the renderer — but no production caller ever asked for them. Only the base
and Shift layers are read, and Shift is the one the web layout map cannot
supply, which is why the helper exists at all.

Removing them halves the helper's UCKeyTranslate work per key and drops the
option parameter that six signatures were threading through for nobody.
2026-08-16 12:49:02 -07:00
Jinjing 3c8410d927 Revert "fix(browser): route every cookie-import write through CDP identities …" (#14942)
This reverts commit bf6dc6fcba.
2026-08-16 12:33:45 -07:00
Jinjing ef1224c4f7 Revert "Preserve OpenCode session across command completion, control SessionS…" (#14943)
This reverts commit 1da1bdc01c.
2026-08-16 12:33:28 -07:00
Brennan Benson bf6dc6fcba fix(browser): route every cookie-import write through CDP identities (STA-4300) (#14729)
* test(browser): repro STA-4300 CHIPS partition downgrade on native cookie import success path

* fix(browser): route every cookie-import write through CDP identities (STA-4300)

* test(browser): cover partition fidelity on both import write paths (STA-4300)

* fix(browser): never stage unreadable cookie partitions (STA-4300)

* fix(browser): gate partition skips on client support (STA-4300)

* test(browser): anchor the client partition-skip capability assertion (STA-4300)

* fix(browser): skip Firefox CHIPS cookies whose Chromium identity cannot be rebuilt (STA-4300)

* fix(browser): tolerate a Firefox schema without originAttributes (STA-4300)

* docs(browser): correct a comment that still named the removed cookies.set write

* fix(browser): block lossy Firefox cookie import on old clients

* fix(browser): read Firefox CHIPS from schema flag
2026-08-16 11:46:45 -07:00
Jinjing b60df2e3d6 Revert "feat(workspace): set project location from the create-worktree host p…" (#14912)
This reverts commit e4e54a17d0.
2026-08-16 10:41:39 -07:00
Jinjing 763b1febeb Revert "feat(skills): add private bundle sharing (#14401)" (#14913)
This reverts commit 757fae28d7.
2026-08-16 10:39:57 -07:00
Jinjing 1da1bdc01c Preserve OpenCode session across command completion, control SessionStart emission (#14866)
* Preserve OpenCode session across command completion

- Add session start events and launch token tracking to establish session boundaries
- Defer retiring launch authority until OpenCode process actually exits, not just when a command finishes
- Fence previous tokens after restarts to prevent status updates from stale sessions
- Maps SessionStart as a session boundary for proper turn/state management

* Emit SessionStart only from OpenCode, not mimo-code

Restrict SessionStart lifecycle events to OpenCode exclusively. Mimo-code no longer emits SessionStart, as it should rely on OpenCode for session boundary signals. This prevents duplicate lifecycle events that could interfere with pane authority tracking and session state management. Also tighten foreground process result validation to reject stale results after title observation changes, fixing a race where a delayed foreground read from a previous cycle would incorrectly retire authority.
2026-08-16 10:01:27 -07:00
Neil e4e54a17d0 feat(workspace): set project location from the create-worktree host picker (#14868)
* feat(workspace): set project location from the create-worktree host picker

Hosts that still say "Project location not set" now get an inline Set location action. It opens a nested dialog over Create worktree so the in-progress form stays put.

* fix(workspace): replace unset-location status copy with a button

Drop the redundant "Project location not set" caption and show a Set project location action with a hover tooltip instead.

* refactor(workspace): tighten the set-project-location dialog

- reuse CreateProjectParentBrowser instead of a second host-filesystem browse view
- drop ProjectLocationBrowseTarget; parseExecutionHostId already models it
- single setLocation path in RunTargetCombobox (row, button, Enter)
- memoize the default clone URL instead of scanning on every store update
- fix missing required props in the new composer-card test

* fix(workspace): close the correctness gaps in set-project-location

- Escape in the host browser now backs out to the form instead of dismissing
  the dialog and discarding the half-filled path/clone URL. Radix dismisses
  from a document-capture listener, so only preventDefault can stop it.
- Drop the stopPropagation guards: window-capture (the composer's Escape
  handler) already ran by then, so they never protected it — the nestedDialogOpen
  gate does. They did silently kill RemoteFileBrowser's own key handling.
- Hide Set project location for host-local repo:<id> projects (folder projects,
  git repos with no remote). Linking on another host matches by project identity,
  which those have none of, so the call could only ever toast an error.
- Drop a standalone placeholder setup once a repo projection covers the same
  project+host, restoring the (projectId, hostId) uniqueness invariant. Setting a
  location on a host with a pending setup was leaving a ghost that sorts first and
  reads back as 'not set up'.
- Existing-folder submit label matched a catalog string reading 'Importing...'

* test(composer): follow the renamed local in the host-retarget source assertion
2026-08-16 03:30:47 -07:00
Jinwoo HongandE2E Test 757fae28d7 feat(skills): add private bundle sharing (#14401)
Co-authored-by: E2E Test <e2e@test.local>
2026-08-16 02:36:18 -07:00
Neil fd1dba9db9 fix(daemon): validate spawn cwd asynchronously so one dead share cannot freeze every terminal (#14848)
* fix(daemon): validate spawn cwd asynchronously so one dead share cannot freeze every terminal

createOrAttach validated the working directory synchronously on the daemon's
only thread. Measured on Windows 11 + Ubuntu-24.04:

  existsSync on an unreachable UNC share   21,022 ms
  wsl.exe probe, cold distro                1,266 ms
  wsl.exe probe, warm distro                   59 ms
  existsSync/statSync on healthy \\wsl.localhost  4 ms / 1 ms

A single unreachable share therefore blocks the whole RPC loop past the
client's 30s request ceiling, so every other terminal stalls behind it and
reports `DaemonProtocolError: Request createOrAttach timed out after 30000ms`.
The main process already validates asynchronously and passes prevalidatedCwd
(ipc/pty.ts); the daemon never got the same treatment.

Add validateWorkingDirectoryAsync (one stat, not exists-then-stat, so an
unreachable share is not paid for twice) and await it from the daemon spawn
preflights. spawnSubprocess now returns SubprocessHandle | Promise<...>, which
existing sync stubs still satisfy.

Deliberately not bounding the stat with a timeout: the 30s ceiling comes from
blocking the shared loop, not from the duration. A timeout cannot tell "slow
share" from "gone share", so it would fail spawns that succeed today at 3-8s
on a cold VPN mount, and trade an accurate "working directory does not exist"
for a guess.

The new await opened a race: it sits between the "already exists?" check and
the sessions.set that publishes the session, so two concurrent creates for one
session id both spawned. Gate creation per session id; distinct ids still spawn
in parallel.

STA-4470

* fix(daemon): fence async spawn lifecycle
2026-08-16 01:16:11 -07:00
Neil 6cf6a7faff feat(crash-reporting): capture Crashpad minidumps and name the failing CHECK (#14823)
* feat(crash-reporting): capture Crashpad minidumps and name the failing CHECK

40% of renderer deaths report exit 0x80000003 (STATUS_BREAKPOINT) — a Chromium
CHECK/DCHECK — and we captured only the exit code, so the cause was structurally
unknowable. Nothing in the tree wired crashReporter at all.

Start Crashpad pre-whenReady and lift the text signature out of the dump:
Chromium stores the fatal log line in the LOG_FATAL annotation, so the check
name, file and line are recoverable with no symbols and no minidump_stackwalk.

Upload stays off. The existing transport is a user-initiated 4 MiB text bundle;
raw dumps are multi-MB binary carrying process memory. Dumps stay on disk and
only the signature rides the existing crash-report flow.

- minidump-stream-reader: bounds-checked view; a truncated dump degrades
- minidump-crashpad-annotations: allowlisted keys (switch-N carries command lines)
- minidump-crash-signature: LOG_FATAL -> file/line, exception, faulting module
- crashpad-capture: polls for the dump, which races process-gone delivery

STA-4469

* refactor(crash-reporting): claim dumps once and match them to the dead process

A newer dump from a different process could be paired to the wrong report, and
two reports in one crash burst could both claim the same file. Match the dump's
own ptype against the process Electron said died, and claim each dump once.

Also let the fatal line use the stack-length budget: a CHECK message truncated
at 240 chars can lose the condition, which is the diagnosis.

Fix the child-type test to match the classifier: GPU exits are recoverable
churn and never become reports, so they must not burn a dump poll either.

STA-4469

* fix(crash-reporting): recover real Electron CHECK logs

Electron 43 Windows dumps carry the Chromium CHECK line in captured memory but omit the claimed LOG_FATAL annotation. Recover only bounded Chromium-formatted fatal/CHECK lines, and prune raw dumps after crash events so suppressed child crash storms stay within the 128 MiB budget.

* fix(crash-reporting): satisfy not-found lint
2026-08-15 23:45:32 -07:00
Neil 83117f2860 refactor(integrations): split issue-tracker clients under the max-lines budget (#14704)
The GitLab, GitHub, Jira and Linear integration modules, their two IPC
registrars, and the shared GitHub project types each carried a file-level
`eslint-disable max-lines` and ran 351-614 counted lines against a 300-line
budget. AGENTS.md calls for splitting rather than suppressing, and
config/max-lines-baseline.txt is a shrink-only ratchet, so this removes all
eight suppressions and prunes their entries (341 -> 333).

Pure move, no behavior change. Each client is cut along the seam it already
had: per-operation modules for the issue APIs (create / update / comment /
field options), and for Jira the request queue, site credential store,
authenticated request, and site identity. The two IPC registrars keep their own
handlers and delegate the rest to per-domain sub-registrars, so they remain
real entry points rather than re-export shims.

The IPC surface is proved intact rather than assumed: comparing (method,
channel) multisets between HEAD and the split gives 52 registrations across 52
distinct channels on both sides.

Provider-neutrality is preserved -- GitLab and GitHub keep separate, parallel
module layouts rather than being merged behind a shared abstraction.

Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(the one remaining failure is a pre-existing load flake in an untouched file,
green when re-run serially), no new runtime import cycles among 744 modules,
and no lint suppression added anywhere.
2026-08-15 18:17:20 -07:00