Commit Graph
733 Commits
Author SHA1 Message Date
Brennan Benson 6c4625d7ff fix(terminal): main records every tab close, and seeding reads the records (#22955)
* fix(terminal): every explicit terminal close commits through one main transaction

A renderer save cannot shrink terminal membership once main owns a repo's
topology, so desktop tab and pane closes, CLI split-pane closes and mobile
split-pane closes only became durable when the killed process's exit retired
the surface. A close whose kill failed or threw, or whose exit was never
certified, came back after a reload.

Every close now reaches closeTerminalSurface: the renderer sends an explicit
intent for user and cleanup closes, the CLI and mobile split-pane closes commit
the pane after their stop, and the headless and relayed mobile closes reuse the
same commit. A failed flush keeps the in-memory removal and no longer cancels
the kill. Exit retirement is unchanged.

* fix(terminal): tell the desktop renderer to drop a split pane main closed

A CLI or mobile close of one pane in a split commits the pane in main, but the
desktop kept showing it until reload when no exit arrived to remove it. The
close now sends a leaf-addressed notice: a mounted pane closes by leaf id, and
a parked tab collapses its stored layout. Addressing by leaf makes the notice
and the renderer's exit handling no-ops after each other, which replaces the
numeric pane-id notice that could close the whole tab when the exit won.

* fix(terminal): a pane close never widens into a whole-tab close

A leaf-addressed close fell through to the whole-tab close whenever main's layout no longer
held that leaf as one of several. Main's exit handling retires an exited split pane from the
saved layout, so closing that pane afterwards (the exited-pane overlay's Close, or a CLI close
whose stop delivers the exit first) removed the whole tab, live sibling included, and the
next renderer save could not restore it. A pane close is now a no-op unless its leaf is in a
multi-pane layout.

Also updates two mobile split-close assertions to expect the leaf-addressed notice, and adds a
test that a relayed mobile close of a renderer-listed tab still reaches the renderer's pin guard.

* test(terminal): cover the PTY-handle branch of a CLI split-pane close

The existing CLI split test resolves its handle through the renderer graph, so the branch
that closes a runtime-owned pane by its PTY handle had no test failing without its commit.

* fix(terminal): main records every terminal tab close, and seeding reads the records

* test(terminal): pin the renderer close mirror against an early SSH pull, and main's record writes against later saves

* fix(terminal): a CLI pane close with an unconfirmed stop closes only that pane

`orca terminal close <handle>` on one pane of a split used to close the
whole tab, live sibling included, whenever that pane's stop could not be
confirmed (for example an unreachable SSH host). An unconfirmed stop is
unverifiable, not a reason to drop siblings: the close now commits only
that leaf, tells the renderer to drop that leaf, and leaves the owed kill
to the controller's existing SSH pending-kill path.

On a host where no renderer lists the tab, main now also removes the
closed pane from the paired-client snapshot (with its retirement proof),
since no exit may arrive to do it.

* docs(terminal): correct the pull merge's record-safety comment

The mirror is written by closeTab before main commits, so the claim that only a
main-committed close writes a record was inaccurate; also reflow a split comment.

* refactor(terminal): one resolver decides whether a pane close becomes a tab close

Every explicit close now states its target as `{kind:'tab'}` or `{kind:'pane', leafId}`; no
optional leaf id silently means the whole tab. Main resolves a close it started in exactly one
place, reading the copy of the tab's panes its layout owner holds (the renderer-published layout
for tabs the desktop renderer lists, main's session layout otherwise). Only `last-pane` escalates,
through the existing tab path so the renderer's pin guard still runs; an unknown pane never widens.

- The CLI and phone paths drop their per-site sibling counts for the resolver.
- The notifier splits into a tab-only close and a leaf-addressed pane close.
- The headless tab closer takes a parent tab id, so a pane row cannot reach it.
- A phone close of one pane on a host with no desktop window now stops and closes only that pane.
- A phone close of one pane with no live process record closes that pane, not its tab.

* fix(cli): an unverifiable stop says the close happened

`orca terminal close` still exits 1 when the process stop cannot be verified, but its message now
says the terminal was closed and names the host's reason, instead of "close failed". It promises
that the kill retries on reconnect only when the SSH relay itself never answered the stop, the one
case a recorded kill order backs.

* fix(terminal): a phone pane close commits even when its kill fails

A paired client's close of one pane threw `terminal_close_failed` before committing anything when
the controller reported the kill failed, so the pane stayed. The kill is now best-effort, as it is
for a whole-tab close: the pane's removal always commits and the failure stays on the PTY's
liveness verdict.

* fix(terminal): a pane close widens only when a copy shows it is the last pane

The close resolver read an owner copy that records no panes as "the tab has
one pane", so a CLI close of one pane of a split, addressed while the
renderer listed the tab before publishing its panes, closed the whole tab.

Every copy now counts only if it records at least one pane, read in the
owner's order with the published rows as the last fallback, and a pane
close widens only when a copy lists that pane as the tab's only one. An
unsplit tab whose saved layout predates its pane still closes: its
published row names the pane.

* fix(cli): promise a kill retry only when the host recorded the kill

The close receipt inferred "the kill retries when the host reconnects" from
the stop reason's text, which a new transport message or a reworded error
would silently break.

An explicit close now records the replayable kill order when its stop goes
unconfirmed, before sending the follow-up kill (whose own failure is
recorded only once its RPC settles), and reports that on the receipt as an
optional `pendingKillRecorded`. The CLI promises the retry only from that
field, so an older host, which never sends it, gets no promise.

* test(pty): justify the controller cast the recorded-stop tests extend

* fix(terminal): parse the close target with typed narrowing

The low-evidence lint gate rejects Reflect.get and broad object parameters,
which failed static analysis. Narrow with 'in' checks instead and cover the
boundary parser's accept and reject cases.

* test(terminal): drive the close-record tests through the durable store

* fix(terminal): a desktop tab close is not refused by a split that bound while it waited

The renderer has already removed and killed a tab it closes, so its close intent now skips the
owner fence phone and CLI closes use. Before, a split pane whose binding was admitted between the
close request and its durable write made main refuse the close, and the tab came back on the next
launch whenever its processes did not exit.

* chore(terminal): note that closedByLayoutOwner goes away once main owns the terminal layout

* chore: take the base branch's lockfile, which a merge had reverted

* test(terminal): reload the close-intent fixture through the SQLite profile store

Main now requires a SQLite profile-state authority for a writable Store, so the save-and-reload
close tests build and reopen their store through the shared SQLite test harness.

* docs(terminal): state the close-record rule in the merge's active-workspace comment

* fix(terminal): the first close of a tab keeps its record, and every lookup honours the TTL

* test(terminal): give the relay reattach close record a recent close time

* fix(terminal): a close that removes a listed tab records itself; only an echo is skipped

The echo of a close main already made never finds the tab listed, and the close transaction already skips it when a live record exists. The record helper kept the first record as well, which only ever applied to a close that did remove a listed tab, and there it kept a stale reason and TTL.
2026-09-28 12:40:26 -07:00
8b410b4893 feat: add first-class Qoder CLI support (#23581)
feat: add first-class Qoder CLI support

Integrate Qoder launch, identity, canonical hook status, trust and resume.
Verify with captured Qoder 1.1.64 transcripts and hidden Electron sidebar checks.

Builds on and cross-reviews #7502, #8611, #9655, #12910, #13311 and #15291.

Co-authored-by: dalveytech-vincent <vincent@dalveytech.com>
Co-authored-by: Eridanus117 <45489268+Eridanus117@users.noreply.github.com>
Co-authored-by: xingqingzzp-gif <xingqingzzp-gif@users.noreply.github.com>
Co-authored-by: jyang2004 <jyang2004@users.noreply.github.com>
Co-authored-by: yunqian <yunqian@alibaba-inc.com>
Co-authored-by: huzhening.hzn <huzhening.hzn@alibaba-inc.com>
2026-09-28 02:59:50 -07:00
Brennan Benson c53ed030b2 fix(terminal): one intentional-stop register, and per-run spawn and input facts (#22989)
* fix(terminal): one intentional-stop register and per-run spawn and input facts

Main now keeps one register of PTY stops it made on purpose, with an owner
count and a kind: reversible (sleep, hibernation) or replaced (a restart
handing the pane to a new process). It replaces four separate markers, and
every exit path reads it, so a hibernated or restarted pane keeps its tab and
binding through the exit instead of being retired and grafted back.

Main also records two facts per process run, keyed by incarnation: whether the
run was a fresh spawn (from the spawn-commit origin, which now counts a cold
restore as a reattach) and when a client first sent it input a person
produced. Renderer keystroke, paste, drop and quick-command writes carry a
userInput flag; terminal.send, stream input and dispatched agent prompts are
recorded by the runtime, excluding terminal query replies.

* test(terminal): a stopped process's synthetic and late provider exits keep its pane

One process can reach the runtime's exit handler twice: main's synthetic exit
when the kill reply overtakes the stream, then the provider's own exit, which
certifies the death. Both must read the same stop, a later process on the
same id must not, and the stop is forgotten when the duplicate-exit window
closes.

* test(terminal): a runtime worktree sleep keeps its tabs and wake bindings

* fix(terminal): keep every overlapping stop kind on one exit

A sleep and a restart that stop the same process now each keep their
label: the exit carries both renderer flags, and a runtime kill during
the sleep still records its SSH stop as reversible. A later stop of an
already-stopped process joins its entry, so a failed repeat cannot erase
the stop that landed.

The cold-restore rule moves into the run facts, so the binding span keeps
labelling a cold restore as a spawn.

* test(terminal): IME and Hangul commits reach the PTY as user input

Drives a real xterm through the pane's input handling: a macOS
input-source substitution, an iPadOS Hangul syllable, and a composition
the route delivers after a pane switch each reach onData flagged as user
input, which is what tags the write main records.

* test(terminal): state why the IME provenance test's canvas stub is safe

* fix(terminal): a new process ends an unpinned stop, and focus reports are not input

- A landed stop that no exit pinned to an incarnation now ends when a new
  process commits on the same id, so that process's exit is not read as
  the stop. Both spawn-commit funnels report through one runtime method.
- A spawn commit with no incarnation starts its run facts clean, since it
  cannot be told from a new process.
- Stream input that is only focus reports no longer counts as user input;
  the desktop renderer already excludes them through xterm's own signal.
- The IME provenance test uses typed fakes: the pane input and composition
  route installers now take only the fields they read.

* chore: restore pnpm-lock.yaml to the base revision

* fix(terminal): typing in the dashboard preview is a run's user input

The dashboard popout's terminal preview writes to the PTY through its own
runtime path, which recorded no input, so a pane typed into only from the
preview read as untyped. It now records before the write, after the
mobile-driver check. One classifier, shared with terminal.send, decides
which provenance-free bytes nobody typed: a whole terminal reply or only
focus reports, which also covers the focus-in a desktop renderer sends
when it reattaches a remote pane.

* test(terminal): the IME provenance test removes the navigator stubs it adds

happy-dom serves navigator.userAgent from the prototype, so the test found no
own descriptor to restore and left the iPad user agent in place. The pane-switch
case then ran as a Mac pane only because it followed the Hangul case. Deleting
the own-property stub restores the default platform for every case.

* fix(terminal): every PTY write names its input kind, and one record point reads it

A run's first input was recorded by opposite defaults: the renderer's
pty:write recorded nothing unless a writer opted in, terminal.send recorded
everything that was not a reply, and main's own controller writes recorded
nothing. Each unclassified producer silently took its transport's default, so
the worktree-create draft counted from the renderer and not from main, and
mailbox pointers never counted.

Every host write entry point now takes a required kind: driving, launch or
query-reply. That covers the runtime controller's write and settled write,
the terminal writer, sendTerminal, sendTerminalAgentPrompt, pty:write and
pty:writeAccepted, the renderer transport and the runtime input helpers. The
fact is recorded once, just before the provider write, in the controller
funnel and the pty:write funnel: only driving bytes count, and a payload that
is only a reply or focus reports never does. The per-producer records are
gone.

Launch writes (create-time drafts and follow-ups, the agent launch prompt,
restored and cold-restore startup commands, the SSH background launch) do not
count. Mailbox pointers, dispatch, plugins and agent-team sends do. The
runtime mixins are unchecked by the compiler, and the pane session is an any
bag, so two source scans fail on any write there that leaves out its kind.

* fix(terminal): type the pane session's transport so the compiler checks every write's kind

The pane-connection session is an `any` bag, so a transport write there that
left out its input kind compiled, and only a text scan over the renderer
caught it. Declaring `transport: PtyTransport` on the session puts those writes
under the compiler, so the scan is gone. One hidden-delivery guard now narrows
a null PTY id itself instead of relying on an untyped predicate.

The main-process scan over `@ts-nocheck` runtime files missed five files whose
directive sits on the second line, below a lint directive, and missed a write
through a local alias of the PTY controller. It now finds both.

* fix(terminal): record a runtime spawn's run facts only after its binding save succeeds

The runtime spawn funnel reported the commit before its host-session binding save, so a spawn discarded for a failed save still recorded run facts and cleared a landed stop. Report it after the save, and separately on the adopted return, which makes no save. Also align tests and fixtures with main: the removed stop-owner map, required write input kinds, SQLite-backed stores and async binding saves.

* test(terminal): pin where the renderer spawn funnel records its commit

The merge moved the renderer funnel's spawn-commit report into the publish step, after the binding save, but no test covered it: deleting the call or moving it back before the save left every suite green. Cover both: a committed spawn records its run facts and supersedes an unpinned landed stop, and a spawn discarded for a failed binding save records nothing and leaves the stop.

* fix(terminal): record a runtime spawn's commit only once its registration succeeds

The runtime funnel reported each commit before registering the process, so a spawn that exited during start still recorded run facts and could end a landed stop, while the renderer funnel reports only after registration. Report after registration on both the ordinary and the adopted branch, so only a committed, registered run has facts and an unknown run keeps reading as not fresh.

* test(terminal): pin that a runtime adoption supersedes a landed stop no exit pinned

* fix(terminal): read an unconfirmed explicit stop's reversibility from the intentional-stop registry
2026-09-27 23:13:24 -07:00
Brennan Benson 993183afd7 fix(terminal): every explicit terminal close commits through one main transaction (#22929)
* fix(terminal): every explicit terminal close commits through one main transaction

A renderer save cannot shrink terminal membership once main owns a repo's
topology, so desktop tab and pane closes, CLI split-pane closes and mobile
split-pane closes only became durable when the killed process's exit retired
the surface. A close whose kill failed or threw, or whose exit was never
certified, came back after a reload.

Every close now reaches closeTerminalSurface: the renderer sends an explicit
intent for user and cleanup closes, the CLI and mobile split-pane closes commit
the pane after their stop, and the headless and relayed mobile closes reuse the
same commit. A failed flush keeps the in-memory removal and no longer cancels
the kill. Exit retirement is unchanged.

* fix(terminal): tell the desktop renderer to drop a split pane main closed

A CLI or mobile close of one pane in a split commits the pane in main, but the
desktop kept showing it until reload when no exit arrived to remove it. The
close now sends a leaf-addressed notice: a mounted pane closes by leaf id, and
a parked tab collapses its stored layout. Addressing by leaf makes the notice
and the renderer's exit handling no-ops after each other, which replaces the
numeric pane-id notice that could close the whole tab when the exit won.

* fix(terminal): a pane close never widens into a whole-tab close

A leaf-addressed close fell through to the whole-tab close whenever main's layout no longer
held that leaf as one of several. Main's exit handling retires an exited split pane from the
saved layout, so closing that pane afterwards (the exited-pane overlay's Close, or a CLI close
whose stop delivers the exit first) removed the whole tab, live sibling included, and the
next renderer save could not restore it. A pane close is now a no-op unless its leaf is in a
multi-pane layout.

Also updates two mobile split-close assertions to expect the leaf-addressed notice, and adds a
test that a relayed mobile close of a renderer-listed tab still reaches the renderer's pin guard.

* test(terminal): cover the PTY-handle branch of a CLI split-pane close

The existing CLI split test resolves its handle through the renderer graph, so the branch
that closes a runtime-owned pane by its PTY handle had no test failing without its commit.

* fix(terminal): a CLI pane close with an unconfirmed stop closes only that pane

`orca terminal close <handle>` on one pane of a split used to close the
whole tab, live sibling included, whenever that pane's stop could not be
confirmed (for example an unreachable SSH host). An unconfirmed stop is
unverifiable, not a reason to drop siblings: the close now commits only
that leaf, tells the renderer to drop that leaf, and leaves the owed kill
to the controller's existing SSH pending-kill path.

On a host where no renderer lists the tab, main now also removes the
closed pane from the paired-client snapshot (with its retirement proof),
since no exit may arrive to do it.

* refactor(terminal): one resolver decides whether a pane close becomes a tab close

Every explicit close now states its target as `{kind:'tab'}` or `{kind:'pane', leafId}`; no
optional leaf id silently means the whole tab. Main resolves a close it started in exactly one
place, reading the copy of the tab's panes its layout owner holds (the renderer-published layout
for tabs the desktop renderer lists, main's session layout otherwise). Only `last-pane` escalates,
through the existing tab path so the renderer's pin guard still runs; an unknown pane never widens.

- The CLI and phone paths drop their per-site sibling counts for the resolver.
- The notifier splits into a tab-only close and a leaf-addressed pane close.
- The headless tab closer takes a parent tab id, so a pane row cannot reach it.
- A phone close of one pane on a host with no desktop window now stops and closes only that pane.
- A phone close of one pane with no live process record closes that pane, not its tab.

* fix(cli): an unverifiable stop says the close happened

`orca terminal close` still exits 1 when the process stop cannot be verified, but its message now
says the terminal was closed and names the host's reason, instead of "close failed". It promises
that the kill retries on reconnect only when the SSH relay itself never answered the stop, the one
case a recorded kill order backs.

* fix(terminal): a phone pane close commits even when its kill fails

A paired client's close of one pane threw `terminal_close_failed` before committing anything when
the controller reported the kill failed, so the pane stayed. The kill is now best-effort, as it is
for a whole-tab close: the pane's removal always commits and the failure stays on the PTY's
liveness verdict.

* fix(terminal): a pane close widens only when a copy shows it is the last pane

The close resolver read an owner copy that records no panes as "the tab has
one pane", so a CLI close of one pane of a split, addressed while the
renderer listed the tab before publishing its panes, closed the whole tab.

Every copy now counts only if it records at least one pane, read in the
owner's order with the published rows as the last fallback, and a pane
close widens only when a copy lists that pane as the tab's only one. An
unsplit tab whose saved layout predates its pane still closes: its
published row names the pane.

* fix(cli): promise a kill retry only when the host recorded the kill

The close receipt inferred "the kill retries when the host reconnects" from
the stop reason's text, which a new transport message or a reworded error
would silently break.

An explicit close now records the replayable kill order when its stop goes
unconfirmed, before sending the follow-up kill (whose own failure is
recorded only once its RPC settles), and reports that on the receipt as an
optional `pendingKillRecorded`. The CLI promises the retry only from that
field, so an older host, which never sends it, gets no promise.

* test(pty): justify the controller cast the recorded-stop tests extend

* fix(terminal): parse the close target with typed narrowing

The low-evidence lint gate rejects Reflect.get and broad object parameters,
which failed static analysis. Narrow with 'in' checks instead and cover the
boundary parser's accept and reject cases.

* fix(terminal): a desktop tab close is not refused by a split that bound while it waited

The renderer has already removed and killed a tab it closes, so its close intent now skips the
owner fence phone and CLI closes use. Before, a split pane whose binding was admitted between the
close request and its durable write made main refuse the close, and the tab came back on the next
launch whenever its processes did not exit.

* chore(terminal): note that closedByLayoutOwner goes away once main owns the terminal layout

* test(terminal): reload the close-intent fixture through the SQLite profile store

Main now requires a SQLite profile-state authority for a writable Store, so the save-and-reload
close tests build and reopen their store through the shared SQLite test harness.
2026-09-27 15:29:22 -07:00
93d8b1f042 fix(ssh): complete keyboard-interactive MFA prompt handling (#15588)
Honor SSH keyboard-interactive prompt echo and empty responses, reuse login
passwords without replaying rejected values, and stop cancelled or stale
credential requests from continuing authentication or restoring the cache.

Original implementation: Junho Kim (#8750).
Port and follow-up work: Allen (#15588).

Verified with 2,773 SSH/credential tests, full typecheck, changed-code quality,
a production Electron build, real-socket MFA fixtures, and rendered UI checks.

Fixes #8622

Co-authored-by: Junho Kim <arkimjh@illinois.edu>
Co-authored-by: microdaery <microdaery@gapp.nthu.edu.tw>
2026-09-27 01:12:29 -07:00
Jinjing f5f537ef14 Revert "Support mouse Back/Forward buttons in shortcuts (#23287)" (#23350)
This reverts commit a86fae0889.
2026-09-26 22:58:01 -07:00
Neil a86fae0889 Support mouse Back/Forward buttons in shortcuts (#23287)
* feat: support mouse Back and Forward shortcut bindings

* fix: ignore duplicate mouse shortcut presses until release
2026-09-26 17:53:58 -07:00
OrcaWin 82412dab8b Persist profile state in SQLite with background writes (#22612)
Migrate profile state to SQLite and move writes and backups into a background worker. Acknowledge terminal, SSH and automation changes only after durable saves. Preserve JSON import, recovery, rollback and compatibility exports.

Validate migration, worker failures, maintenance, cross-profile moves and terminal lifetime races with unit, integration and end-to-end coverage.
2026-09-25 22:47:33 -07:00
Brennan Benson c5f33bd139 fix(ipynb): run no workspace interpreter until the notebook is trusted (#22962) 2026-09-25 20:08:24 -07:00
8846987c99 feat(rate-limits): add Cursor usage tracking (#22633)
* feat(rate-limits): add Cursor usage tracking

## ELI5

If you use Cursor, Orca now shows how much of your monthly Cursor plan you
have used, next to the Claude, Codex and Grok meters, and in Settings →
Accounts. It reads the sign-in Cursor already saved on this computer and never
changes it.

## What changed

Cursor becomes a rate-limit provider like Grok: a status-bar meter (default-on,
with its own toggle), a row in the usage roster, and a Settings → Accounts
section naming the signed-in account.

The credential is read from whichever of three stores has it, first match wins,
all read-only:

- the macOS login keychain item `cursor-access-token` / `cursor-user`, which is
  where `cursor-agent` 2026.06+ keeps the session;
- `~/.cursor/auth.json` and its platform variants, used by older CLIs;
- the Cursor IDE's `state.vscdb` (`cursorAuth/accessToken`), for people who
  never run the CLI.

The keychain entry is the one current CLIs use, and reading only `auth.json`
finds nothing on an up-to-date macOS install. A locked keychain cannot mask a
readable `auth.json`, and a locked `state.vscdb` cannot mask either.
`~/.cursor/cli-config.json` supplies the account's email and display name; it
never holds a token.

Usage comes from the dashboard route the Cursor web dashboard itself reads,
because Cursor documents no individual-user usage API — every documented API is
team- or Enterprise-scoped. Per Cursor's pricing docs an individual plan has two
pools, Cursor Models and Other Models, both resetting with the billing cycle,
plus optional on-demand spend; each becomes a named bucket. The headline
percentage prefers `used / limit` over the sibling percentage fields, which are
pre-rounded for the dashboard's own copy. Because the route is undocumented the
mapping is defensive: an unrecognised payload resolves to `unavailable` and
hides the bar rather than publishing a zero that reads as "no usage".

Orca never runs `cursor-agent login` and never writes, refreshes or rotates a
Cursor credential. An expired token short-circuits to an actionable
"run cursor-agent login" instead of spending a request that can only 401 — not a
rare case, since `cursor-agent status` still reports `isAuthenticated: true`
against a token that expired months ago.

## Why this shape

Six open PRs implement this feature and none reads the keychain, so each finds
nothing for a large share of users; this takes the auth layer further and keeps
what those PRs verified live. The bar is not gated on `cursor-agent` being on
PATH, unlike other CLI providers, because an IDE-only session is real usage with
no CLI to detect.

`readKeychainPassword` moved out of the Claude keychain reader into
`src/main/macos-keychain/generic-password.ts` so both providers share one
`security(1)` wrapper. It is a byte-for-byte relocation, so Claude's credential
path is unchanged; the two child_process allowlists move the entry with it and
neither ratchet count changes.

Co-authored-by: Preschian Febryantara <preschian@users.noreply.github.com>
Co-authored-by: Qwesdy <qwezdi@proton.me>
Co-authored-by: ivo922 <github.concur614@passmail.net>
Co-authored-by: Mihail Vratchanski <mivrkiki@gmail.com>
Co-authored-by: Tauri-EPO <enrico.pin@gmail.com>
Co-authored-by: Raajik <44516546+Raajik@users.noreply.github.com>

* test(rate-limits): name the JWT helper's segment type in the Cursor tests

The anti-slop gate rejects a bare `object` parameter; the fixtures build a
claims record, so say that.

* fix(rate-limits): render Cursor's pools and keep its plan total visible

Review of the first commit found the meter effectively blank for a healthy
account, which the screenshots missed because the only Cursor session on hand
had expired and never reached the success path.

- The verbose status-bar segment filtered buckets through an allowlist written
  for Gemini's experimental models, so both Cursor pools were dropped and the
  fallback needed a `session` window Cursor never reports. A signed-in account
  rendered an icon and no number. The allowlist now admits Cursor's pools, and
  the fallback accepts a monthly window.
- `getWindowSections` dropped `monthly` whenever buckets existed. Cursor puts
  the plan total there and its sub-pools in buckets, so a plan at 92% showed as
  50% in the roster, the tooltip, and the tightest-usage pick.
- A plan reporting `enabled: false` still published its 0% pools, painting a
  healthy meter for a pool the account does not own and skipping the
  request-quota fallback.
- `redirect: 'error'` turned the dashboard's bounce to /login into a generic
  network failure, hiding the actionable sign-in message.
- A busy `state.vscdb` (the IDE holds it open) surfaced as a provider error,
  which would pin an alert bar on Cursor IDE users who never set Cursor up in
  Orca. It falls through to "no credential" instead.
- Refreshing the Accounts section read the keychain twice for one update.

* fix(rate-limits): pin the platform in the Cursor keychain tests

Review caught three cases that assumed macOS: the keychain source is behind an
explicit `process.platform` check, so on the Linux CI runner the mocked read was
never reached and the tests read the CLI file instead. They now set the platform
they mean, and two new cases assert the off-macOS fall-through.

Also track the credentials reference doc (docs/** is ignored by default, so a
new reference needs its own allowlist entry) and give the visibility fixtures
their own provider id instead of Grok's.

* fix(rate-limits): prefer a live Cursor session and report a failed refresh

Review round two, from CodeRabbit and Pullfrog.

- Credential precedence returned the first token that parsed, so an expired
  keychain token in front of a fresh Cursor IDE session reported "sign-in
  expired" on every poll while a usable session sat one source below. A live
  session now wins; the expired one is returned only when nothing live exists,
  so the actionable message still appears in that case.
- The usage schema took `.optional()` where the route sends `null` for an absent
  sub-object, so one null pool failed the parse for the whole body and threw
  away valid pools and the billing cycle with it.
- Cursor usage could survive an account switch: a failed refresh for account B
  kept account A's figures beside B's name in Accounts. The snapshot now carries
  a hashed account fingerprint, and a known-and-changed identity clears the
  previous reading. A refresh that names no account still keeps its own.
- The Accounts section rendered nothing at all when a signed-in account's fetch
  failed, and could repaint an older account when two status reads overlapped.
  It now states the failure — beside the numbers when a stale snapshot remains —
  and ignores superseded reads.
- A web client claimed "not signed in" for a host it cannot read, contradicting
  the meter beside it; it now says the detail is host-only.
- Signed-out copy named `cursor-agent login` as the only way in, though an IDE
  sign-in works just as well.
- The census comment ended at 4219 after the pacer squash without naming the two
  modules #22616 added; recorded them, re-measured on a clean origin/main.
- Narrowed the docs claim: Cursor documents all-plan APIs, but no individual
  usage endpoint.

* fix(i18n): localize the web client's Cursor host-only notice

It reaches the Accounts pane like any other string, so the coverage gate is
right to want it in the catalog rather than allowlisted.

* fix(rate-limits): name the Cursor account on failed refreshes, and ship the reworded copy

Review round three. Both findings say an earlier fix did not actually take.

- The account-switch guard reads `authProvenance` off the fresh result, but the
  fetcher stamped it only on success and network failures. The `stale-token`,
  429, 5xx and parse results omitted it, and so did the expired-session branch —
  so a switch whose first refresh failed, which is precisely the case the guard
  exists for, still rendered the previous account's figures under the new name.
  Every failure holding a readable session now names its account; a missing or
  unreadable credential still names none. The service test also fed a result
  shape the fetcher never produces, so it proved nothing; it now uses the real
  stale-token shape, and the fetcher test asserts provenance across 401/429/5xx
  and expiry.
- The reworded signed-out copy never rendered: a present catalog value beats the
  `translate()` fallback, and `sync:localization-catalog` only adds missing keys
  rather than updating changed defaults. Updated both strings in en.json, which
  also prunes them from the runtime-required catalog now that they match.

* docs: keep the Cursor credentials reference out of the tree

Its content lives in the PR description instead; docs/** stays ignored rather
than gaining an allowlist entry for this branch.

* test(mobile): drop the census note main no longer pins

main removed `SESSION_ROUTE_MODULES` and re-pinned this lane on a different
count, so the paragraph this branch added documents a number series that is
gone. The branch touches nothing in this file now.

---------

Co-authored-by: Preschian Febryantara <preschian@users.noreply.github.com>
Co-authored-by: Qwesdy <qwezdi@proton.me>
Co-authored-by: ivo922 <github.concur614@passmail.net>
Co-authored-by: Mihail Vratchanski <mivrkiki@gmail.com>
Co-authored-by: Tauri-EPO <enrico.pin@gmail.com>
Co-authored-by: Raajik <44516546+Raajik@users.noreply.github.com>
2026-09-25 18:54:11 -07:00
Brennan BensonandClaude d443320af2 refactor(native-chat): remove the unused terminal handoff (#22783)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-25 10:17:36 -07:00
745cde69cd fix(feedback): send text-only report when screenshots exceed the upload limit (#22508)
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>
2026-09-25 02:49:52 -07:00
Neil 69839c253e feat(zcode): explain a ZCode build that has no terminal UI (#22730)
* feat(zcode): explain a ZCode build that has no terminal UI

ZCode ships one agent runtime behind two front ends. The desktop app bundles it
without `@zcode/tui`, because it draws its own window in Electron. Put that
bundle on PATH as `zcode` and it answers `--version`, runs `-p` headlessly, and
passes `zcode doctor` — so Orca detects it, launches it, and installs hooks
against it, all successfully. Only the interactive session fails, leaving a bare
Node stack trace in the pane that reads as a broken Orca integration.

Watch a freshly launched ZCode pane's first output and replace that with an
explanation: Orca's hooks are fine, this `zcode` just cannot open a session,
install one that ships the TUI.

The rule keys on Node's own module-resolution error rather than on the healthy
build's "TUI requires an interactive terminal." message, because ZCode localizes
the latter (`TUI 需要交互式终端。` in zh-CN) and matching it would miss every
non-English user. Node's error is not translated and names the package.

Scoped so it costs a healthy pane nothing: it runs only for a pane Orca launched
as `zcode`, and only over the first 8 KiB, because a module-resolution failure
happens before the runtime renders anything.

Evidence: `src/main/runtime/__fixtures__/zcode-missing-tui.txt`, a recorded PTY
capture of the desktop bundle refusing to start, per
docs/reference/agent-pty-transcript-capture.md.

Reported-by: JWu527

* refactor(zcode): ask the CLI if it can open a session instead of watching for the failure

The stream watcher this replaces never fired. Before/after screenshots were
identical and instrumentation showed the hook never ran, so the sidecar was
both misplaced and racing a failure that lands ~440ms after spawn.

Replace it with a direct question, answered once per run and cached.

Reading zai-org/ZCode shows why running it is the only way to ask, and why the
answer is unambiguous. `--version` and `doctor` are byte-identical in shape
between a build that has the terminal UI and one that does not, because the TUI
is only ever touched on the `tui` command path. There, `runTuiCommand` calls
`loadTuiRuntime()` before anything else, and `runTui` checks for a TTY only
after that module is already loaded. So with stdin at EOF:

  - no TUI  -> fails in the loader  -> Node's module-resolution error
  - has TUI -> loads, then declines -> "TUI requires an interactive terminal."

The module error is therefore present exactly when the terminal UI is absent.
All three shipping shapes land correctly: an npm/node-bundle install resolves
`@zcode/tui` as a real package (esbuild marks it external, so it is never
inlined), a SEA build always carries it as embedded assets, and the desktop
app's bundled runtime carries neither.

Verified against both real builds on this machine rather than a mock: the
desktop bundle answers `missing-tui`, a CLI built from source answers
`interactive`, and a command that does not exist answers `unknown` — the probe
fails open so an unrelated spawn failure never accuses a working CLI.

* feat(zcode): warn at launch when the installed zcode cannot open a session

Wires the capability probe to the one place a ZCode launch is first known:
terminal tab creation, which runs before the pane connects, so the explanation
reaches the screen alongside the failure rather than after it.

- main exposes the cached probe over `preflight:zcodeInteractiveCapability`,
  beside the other "what can the installed CLIs do" answers
- the web preload stub answers `unknown`, because a paired client has no
  business deciding anything about the host's CLI install
- the renderer notice is advisory: a probe that cannot run never blocks a launch

Verified in the running app against the real desktop bundle: creating a ZCode
workspace now shows "This ZCode build has no terminal UI" next to the stack
trace, where before the trace stood alone.
2026-09-25 02:28:40 -07:00
Brennan BensonandClaude 5c45337a6f fix(terminal): make Codex restart replace the pane's process instead of reattaching it (#22737)
* fix(terminal): make Codex restart replace the pane's process instead of reattaching it

A spawn for a pane that still has a live process is treated as a reattach.
Both Codex restart paths raced that: the open-pane restart killed the old
PTY without waiting and then spawned, and the unmounted-tab restart spawned
before killing. Either way the "fresh" spawn could re-adopt the old Codex
(dialog comes back) or hit the half-killed session and leave a plain shell.

pty:spawn now accepts replacesPtyId. Main stops that PTY and waits for the
exit before the spawn resolves the pane owner, so the existing dead-owner
path launches fresh and records the current Codex home. Both restart paths
send it and no longer kill the old PTY themselves.

Fixes #18174

* fix(terminal): send the replaced PTY once and keep the hidden-pane replacement

Two follow-ups to the restart handoff:

- The replaced PTY id rode the transport options, which outlive the first
  spawn. A later fresh spawn from the same pane (for example a hibernation
  wake) re-sent it, so main stopped that id again; SSH relay ids restart at
  pty-1, so it could name another pane's PTY. It is now a connection input
  consumed by the first fresh spawn only.
- The hidden-tab restart still treated a changed binding after the spawn as a
  reason to stand down and reap the replacement. Main has already stopped the
  old PTY by then, so its exit can clear the tab binding (a background-launch
  exit observer does) or a pane can mount, and standing down left the pane
  with no process. It now keeps the replacement unless the tab or leaf was
  actually taken over.

* fix(terminal): hold the pane while a restart stops its old process

A restart spawn stopped the pane's old process before reserving the pane,
so a hidden tab revealed during the stop could reattach the dying process
(and the restart could then join that reattach and return the old id).
The replacing spawn now reserves the pane before the stop, so any spawn
for the pane in that window joins the replacement; a failed stop still
settles the reservation.

The hidden-tab restart also tombstones the old PTY's buffered exit before
spawning, so a reveal mid-restart reconnects by pane identity instead of
replaying that exit into the pane.

* fix(terminal): label a restart's replaced-PTY exit and keep its stop owed until sent

Main now marks the PTY it stops for a replacing spawn and stamps
`replacedByRestart` on that PTY's exit (provider-observed, synthetic, and
SSH-unregistered stops all reach the renderer through the same finalize
step). A failed stop clears the mark with nothing sent; an undelivered mark
expires. The renderer classifies the labeled exit before any consumer runs:
parked-tab watchers end their subscription instead of collapsing the leaf or
closing the tab, the pre-attach buffer discards instead of queuing a death
for a later mount, and a mounted pane treats it as an intentional restart.

With that, the hidden-tab restart no longer tears down its parked watchers
and buffered output before spawning; it releases them only after the swap,
so a stop that fails leaves the still-running Codex fully observed.

The visible restart's pane session now holds the replaced PTY until a spawn
request actually carries it (the IPC transport claims it as it sends). If the
pane is closed or parked first, disposal stops that PTY with an ordinary
kill instead of orphaning it.

* test(terminal): type restart test fakes and merge a duplicate import

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-25 00:10:48 -07:00
Jinwoo Hong 5610b11703 feat(ipynb): create a .venv when pip is locked out, and show ipykernel setup progress (#22710)
* feat(ipynb): set up ipykernel in a new .venv when pip is locked out, and show install progress in the dialog

* fix(ipynb): drop the retired installFailed string from translated catalogs

* refactor(ipynb): drive the setup dialog from one setup state; fix review findings

- Kernel status now only describes the kernel; a single `setup` object (base, offer, phase, error) drives the dialog, replacing the extra statuses and the externallyManaged/setupError fields.
- The picker's 'Create virtual environment…' opens the same dialog; success switches through selectEnvironment, so a running kernel is only replaced once the venv exists.
- Async setup results are dropped when the dialog they belong to is gone (tab closed/reopened).
- main verifies ipykernel imports after pip, reuses an existing .venv instead of re-running venv over it, and explains failures that printed nothing.
- Windows copy command guards the install with if ($?); notebooks at a filesystem root get a correct .venv parent; 'Try again' shows for both retry paths.

* fix(ipynb): close the setup prompt when another Python is picked
2026-09-24 16:22:34 -04:00
Jinwoo Hong 8ba7f829ac feat(ipynb): run notebook cells in a persistent Jupyter kernel (#22581)
* feat(ipynb): run notebook cells in a persistent Jupyter kernel

Replaces the fresh-process runner (which silently re-ran every earlier cell)
with the user's own ipykernel, driven by a small bundled Python bridge over
line-framed JSON. One kernel per open notebook: started on first Run, shut
down when its tab closes or Orca exits (stdin EOF), and the kernel's own
parent poller reaps it if the bridge dies.

The header gains a kernel pill (workspace .venv/.conda recommended, PATH
interpreters, Browse), Interrupt/Restart/Run all/Clear all, and a one-time
missing-ipykernel dialog with Install. Outputs stream live per cell and are
written into the document when the run finishes.

* test(ipynb): cover the stalled-interrupt restart offer

* refactor(ipynb): disable the kernel pill while settling; merge its classes with cn

* refactor(ipynb): tie kernels to their renderer document and simplify the run flow

- Main keys kernels per renderer, so a reload, renderer crash or closed
  window shuts them down, and two windows never share one notebook kernel.
  One close-driven cleanup replaces the separate exit and start-failure
  deletions.
- The first run uses the nearest workspace env, else the first Python on
  PATH; the picker no longer opens itself, so its open state stays in the
  toolbar. Closing the picker brings the missing-ipykernel dialog back
  instead of dropping the queue, which also keeps a Browse pick's run.
- Discovery marks the kernel starting, so a second run during it queues
  instead of starting a second kernel, and a tab closed mid-discovery no
  longer leaks one.
- Running an nbformat 4.4 notebook gives its cells ids (upgrading to 4.5),
  so moving a cell mid-run cannot misroute its output.
- The death notice drops stderr from before the kernel was ready (the
  unencrypted-TCP warning).
- SSH and non-Python runs toast instead of writing a notice into the cell.
- Windows conda envs are named after their folder.

* fix(ipynb): install ipykernel into envs without pip

uv-created venvs ship without pip, so Install failed with 'No module named
pip' there. When pip is missing, bootstrap it with the stdlib's ensurepip
and retry. The install moves beside the other interpreter probes, and the
copyable install command comes from one helper.

* fix(ipynb): address PR review comments on stream errors, old jupyter_client and the Windows venv hint

- Swallow stdout/stderr stream errors on the bridge child, as spawnProcess
  requires, so a broken pipe cannot crash main.
- The bridge exits (reporting the death) even when cleanup_resources is
  missing (jupyter_client < 6.1.5) or raises.
- The install-failure hint suggests `py -m venv .venv` on Windows.

* feat(ipynb): add Cancel to the missing-ipykernel dialog

It does what Esc does: drops the cells waiting on the kernel.

* fix(ipynb): recover from a rejected kernel start; quote the install command per shell

- Discovery moves into start, so one catch turns a rejected
  listPythonEnvironments or startKernel into the usual failed start: the
  session returns to off with the error in the cell, instead of sticking
  at starting.
- The copyable ipykernel command quotes the interpreter only when its path
  has whitespace, prefixing PowerShell's call operator on Windows. Install
  itself still spawns without a shell.

* fix(ipynb): always shell-quote the copyable ipykernel install command

Quote the interpreter path for every path, not only ones with whitespace,
so paths with shell metacharacters like & copy as a working command.
Single quotes are literal in POSIX shells and PowerShell; embedded quotes
are escaped per shell, and PowerShell keeps its & call operator.
2026-09-23 23:38:33 -04:00
Jinwoo Hong f2ac9f29b2 fix(browser): let pixel capture hold its own page drawn, without the desktop window (#22534)
* fix(browser): let pixel capture hold its own page drawn, without the desktop window

Screenshots were the last browser commands that still borrowed the desktop
window: they took the per-page automation-visibility lease, which waits for
two desktop-window animation frames (capped at 2 s) and never arrives when the
window is minimized or throttled. Only pixel capture actually needs a page
drawn — input, scripts, layout, the accessibility tree and PDF all work on a
hidden page.

Capture now takes a main-owned paint hold: a synchronous, per-page ref-count
that tells the renderer one way (no reply awaited) to keep the page drawn and
keeps the desktop renderer unthrottled while held. Both Orca's full-page
capture and the agent-browser helper's screenshots take it in
cdp-screenshot.ts and retry on a bounded schedule until the page answers with
a frame; a CDP error fails fast.

Deleted: the queue's needsPaint lease, the executeJavaScript acquire path and
its two racing 2 s timeouts and late-token cleanup, the renderer's rAF wait and
window bridge, the capture commands' own leases, the fixed 300/500 ms settle
waits, and the global one-screenshot-at-a-time lock.

Rebased onto main after #22528 landed; content identical to the reviewed
branch head 7b390ed6a8.

* fix(browser): probe for a frame instead of repeating the full capture

Retrying a capture resent the caller's full request, so on an already
drawn tall page (a full-page capture takes ~0.5 s) the 250 ms retry
started a second full beyond-viewport capture while the first was still
running. Measured on Electron 43: any later request makes a held page
produce a frame, and that frame answers every pending capture with a
full, correct image. So the capture is sent once and 1x1 probes follow
until it answers; their results are ignored.

Also report a detached debugger as detached rather than destroyed, and
give the layout-metrics timeout its own "did not respond" message, since
that request doesn't need a drawn page.
2026-09-23 21:17:46 -04:00
Neil eb18eaf2b6 feat(usage): add Muse Code local usage provider (#22379)
* feat(usage): add Muse Code local usage provider

Scan Muse session logs (including subagent logs, which hold usage the parent
log does not) for model_completed token events and surface them as a fourth
local usage provider: shared scan worker, persisted per-file cache reused by
mtime/size, cross-log dedupe, Stats tab, and Usage Overview integration.
Muse logs carry no price, so the provider reports tokens only.

* fix(usage): name Muse in Stats & Usage copy; skip partial-cost warning when nothing is priced

* fix(usage): surface unreadable Muse sessions root; name Muse in remaining Stats & Usage copy

* fix(usage): count distinct same-content Muse records within one log
2026-09-22 22:44:10 -07:00
Neil 83dd047fd9 fix(explorer): find files by name in large local workspaces (#22369)
* fix(explorer): search local workspaces by file name across every file

The Explorer name filter only searched remote workspaces directly; local
workspaces still filtered the first 20,001 listed files, so files beyond
that cap never matched in large repos. Local name queries now rank the
whole workspace on the host, and fall back to an uncapped git listing
when ripgrep is not installed.

* fix(explorer): filter capped local listings on the host with the Explorer word rule

Replaces the Quick Open fuzzy top-32 routing, which dropped multi-word
matches and capped visible results. The Explorer keeps its instant
renderer-side filter; only when the local listing hits its cap does it
re-list on the host with the same word rule applied before the cap.

* fix(explorer): keep capped matches when host name filtering fails

- Fall back to the capped listing (and stop re-listing) if the host scan fails
- Keep primary matches when the ignored-file pass fails during a filtered scan
- Key host scans on normalized filter words; reset capped state per filter session
- Bound nameFilter size at the IPC boundary; drop the double readdir walk

* fix(explorer): match name filters without locale-dependent lowercasing

* fix(explorer): avoid render-time ref writes in the host name filter fallback
2026-09-22 19:39:52 -07:00
Brennan Benson 1b85be67d8 feat(native-chat): notify on every settled structured turn (#22105)
* feat(native-chat): notify on every settled structured turn

A structured chat that finished while you were elsewhere lit the sidebar
but never raised an OS notification, and a notification that did arrive
for one could not open the chat it came from.

Unread and delivery now come out of the single resolveAgentAttention
decision the terminal lane already uses: the structured dispatcher calls
applyAgentAttention instead of applyAgentAttentionUnread, so the same
policy that decides what to light also decides what to deliver, through
the same sound and blocked-permission tail.

Every settled turn notifies, as the CLI lane does. Success says
"finished"; failure and cancellation say "stopped" through the shipped
agentInterrupted flag rather than a second vocabulary. A turn whose
outcome the host never stated stays unknown and lights nothing.

The host now dedupes mobile fan-out by event identity (scope, session,
turn) beside the existing per-workspace burst cooldown, so a completion
two windows both saw reaches the phone once while each window still
decides its own banner. Clicking a structured notification reveals the
chat tab: its pane key's leaf is synthetic, so focusTerminal would hunt
a split-layout leaf that does not exist.

* fix(notifications): spend each mobile gate only when it actually notifies

Two review findings on the structured-chat notification lane, both real.

The mobile event gate consumed its reservation before the per-workspace
burst cooldown ran. Two chats in one workspace share that cooldown key,
so the second chat's completion could burn its event key and then lose
the cooldown to the first chat — never announced, yet permanently marked
as announced, so a later window dispatching it could no longer reach the
phone. The gate now peeks first and records the event at dispatch, which
also keeps a known duplicate from burning the cooldown slot.

A notification id is minted from the status row's stateStartedAt, and the
row re-projects that field as the turn settles: the working episode's
start moves into stateHistory and the settled start takes its place. A
banner raised in the window before that re-projection therefore carried
an id acknowledgement never rebuilt, leaving it on screen for good.
Acknowledgement now collects ids for the row's left episodes too — the
same episodes the unread check beside it already scanned, so the two
halves finally read the same turns. Lane-neutral: the terminal lane
mints its ids the same way and had the same gap.

* fix(notifications): drop the mobile event gate and reveal chats in folder workspaces

The per-event mobile dedupe defended against one completion being
dispatched by several Orca windows. Only one renderer mounts the
structured attention bridge, the completion feed is live-only with no
replay, and any in-process duplicate lands inside the existing 5s
per-workspace burst cooldown, which already collapses mobile and
desktop alike. The gate never acted on a real sequence, so the wire
field, the shared ledger and its tests go; mobile delivery is back to
main's behavior.

A folder workspace id ("folder:<id>") has no "repoId::" prefix, so the
click binding was skipped and clicking a chat notification there did
nothing. The chat route selects its workspace itself through
ui:focusEditorTab, so it now binds without a repoId; the terminal
route is unchanged.

* fix(notifications): retire the banner ids actually dispatched, not ids rebuilt from a moved row

A banner's id is minted from the status row's stateStartedAt at dispatch, and that field moves
afterwards: a completion can outrun the settled re-projection, and a settled structured row is
re-stamped with no history entry by any later journal row (a cancel appends a status note after
the turn settles). Rebuilding ids from the row's episodes at acknowledgement missed the second
case and fanned out up to 21 mobile dismissals per pane for ids never raised.

The shared delivery tail now records each dispatched id per subject; acknowledgement retires
those plus the current-row rebuild it always had. The acknowledgement collector is back to
main's single-field form.

* refactor(notifications): retire announced notifications by subject in main

Main now records, per pane, the ids it actually announced (a desktop banner
shown or a phone alert sent) and an acknowledgement passes the acknowledged
pane keys so main retires all of them. This replaces the renderer-side record
of dispatched ids: main is where the announcement happens, so it records only
real announcements, including phone alerts whose desktop banner focus
suppressed. The id rebuilt from the current row stays as the fallback after
a restart empties the in-memory record.
2026-09-22 18:35:18 -07:00
Jinwoo Hong dfff3915c4 fix(browser): scope back/forward/reload/zoom/grab shortcuts to the originating split (#22340)
* fix(browser): scope back/forward/reload/zoom/grab shortcuts to the originating split

With two browser panes visible in a split, Back, Forward, Reload, Hard
Reload, page zoom and Focus Address Bar fired in every visible pane. Main
forwarded these guest chords without the page id, and each split's active
pane subscribed. The renderer-side listeners for the same chords were also
window-wide per pane, so a key pressed in the toolbar (or in a terminal in
another split) reached every active browser pane.

Guest-forwarded chords now carry the originating browserPageId; preload
admits only well-formed payloads and each pane ignores ids that aren't its
own. Toolbar-path listeners use the same focused-split scope Find already
uses. The streamed remote pane's history chord moves onto that scoped hook.

Cmd/Ctrl+C grab (STA-3319) gets the same scope and no longer arms while a
text selection exists outside the browser pane, so copying from the native
chat transcript works again.

* refactor(browser): simplify split shortcut scoping per review

Drop the preload payload admission (main and preload ship together), fold
the three inline scope checks into browserChromeShortcutOwnsEvent, and
replace the outside-overlay selection check with a plain live-selection
rule so Cmd+C copies from surfaces that do not move split focus.

* refactor(browser): share one zoom command type and tidy shortcut comments

BrowserPageZoomEventDetail and BrowserPageZoomCommand were the same shape;
keep one in shared/browser-page-zoom.ts and route guest and local zoom
through a single handler.

* refactor(browser): narrow the zoom event with instanceof instead of a cast

* test(e2e): pin split-scoped browser shortcuts

Two browser splits (and a terminal beside a browser) now prove that Back,
Forward, Reload, Hard Reload, page zoom, Focus Address Bar, and the element
grab chord act only on the split that sent them, from both the guest page and
the browser toolbar. A native chat selection proves Cmd/Ctrl+C copies instead
of arming grab. Split fixtures move to a shared helper so both specs reuse them.
2026-09-22 20:14:19 -04:00
Jinwoo Hong 7650abe224 fix(macos): tell the user when Orca's terminal service can't read their folder, and walk them through the fix (#21923)
* fix(macos): tell the user when Orca's terminal service can't read their folder

On macOS, a terminal daemon that survived an app update can be refused access to
a workspace under Documents, Desktop, or Downloads while the Orca app itself can
still read it. Terminals opened there die with "Operation not permitted" and
nothing on screen explains why. The daemon has reported `cwdReadableByDaemon` on
every create since #18043 and main has emitted `daemon_pty_cwd_denied` on proven
divergence since then; the field data says 1,438 users hit it in 21 days. What
was missing was the notice.

The verdict itself moves off `access()`. A grant-less probe on an affected
machine showed a TCC mode where `access(R_OK|X_OK)` passes on `~/Documents` and
`opendir` still fails, so the check now does what a shell listing its cwd does:
`opendirSync`, one `readSync`, `closeSync`. Only EPERM/EACCES reads as denial —
a missing path, a non-directory, or an unexpected error still reads as readable,
so a non-permission failure can never masquerade as one. The same probe is what
the app side compares with, through one oracle shared by the telemetry emitter
and the notice, so the spawn path reads the directory once.

Proven divergence now also records evidence in main: one entry, keyed by the
daemon's pid, start time and launch nonce, carrying an opaque digest of that
identity and the folder class. No path leaves main. The existing focus-time
`macTccAttribution` poll carries it to the renderer, which raises a second toast
latched per daemon scope: dismissed stays dismissed, and a restart mints a new
identity so the poll returns null and the toast clears with no post-restart
probe. If the replacement daemon is denied too, about 31% of cases, the next
spawn re-records under the new scope and the notice returns, now with the
re-allow sentence doing the work.

No new IPC channel, no daemon protocol field, no polling change, and nothing new
on the spawn path beyond one `opendir`. `daemon_folder_access_notice` counts
shown, dismissed and open_manage_sessions against `daemon_pty_cwd_denied` as the
denominator; `shown` is emitted from main the first time a scope leaves the IPC
handler, so the renderer carries no telemetry plumbing for it.

* fix(macos): clear folder-access evidence only when the same folder class reads back

A readable spawn in ~/code said nothing about a Documents denial but was
hiding the notice; retire the evidence only when the daemon reads a folder
of the class it was denied on.

* fix(macos): say what a terminal-service restart actually does

The Manage Sessions restart confirmation still described the product as it was
before agents resumed themselves: it promised panes showing "Process exited"
that the user reopens by hand, and mentioned legacy-protocol sessions nobody
outside the daemon code can act on. Open terminals and agents come back on
their own now, so the old copy made a routine remedy sound like data loss.

It also called the thing a "daemon". The same restart is about to be offered
from a user-facing fix dialog, so both surfaces now say "terminal service", and
the confirm button is just "Restart".

The new body adds the one fact the old one never stated: terminals on remote
hosts are not affected. Translations of the two changed strings are dropped so
the five non-English locales fall back to English rather than keep showing copy
that is now wrong.

* feat(macos): give the denied-folder notice a fix the user can follow

The folder-access toast told the user their terminal service could not read
Documents and then handed them a paragraph: restart from Manage Sessions, and
if that does not work, re-allow Orca in System Settings. Both halves were
guesses. Roughly a third of restarts do not fix it, and the user had no way to
know which case they were in before spending every open terminal on finding
out.

Main can now answer that. `daemon-folder-access-probe.ts` forks a short-lived
child of the app binary the same way the daemon itself is forked, runs one
opendir/readdir/closedir against the denied path, and prints a single JSON
line. macOS attributes a TCC grant to the process that forked the child, so a
child of the app running now answers exactly the question the running daemon
cannot: would a replacement daemon get in? The child goes through the shared
child-process wrapper, never a shell, with a 3s deadline, a 1KB output cap and
an environment scrubbed to PATH/HOME/TMPDIR. Every failure — timeout, bad
output, spawn error — reads as `unknown`, never as a verdict.

That answer rides out as `restartWillHelp` on the evidence the existing
focus-time poll already carries, and the toast becomes a title and two buttons:
Fix… and Not now. Fix opens a dialog with the two real steps. When the grant is
already in place, step one is shown as done and Restart is live. When it is
not, step one is open and Restart is disabled until it completes — which it
does by itself, because the poll re-probes while the answer is still no, and
returning from System Settings is the moment that lands. An unanswered probe
never accuses the user of a missing grant; it leaves both steps open.

Restart calls the management API directly rather than stacking the Manage
Sessions confirmation on top, since the dialog already states the consequence.
Success replaces the steps with a done line and takes the toast down; failure
says so inline and leaves the button usable.

System Settings opens through the existing developer-permissions pane opener,
which takes an id rather than a URL, with Files and Folders added to it. The
event's action enum now also counts fix_opened, settings_opened,
restart_clicked and — emitted from main when a replacement daemon's first spawn
lands in the folder class the previous one was denied on — whether the restart
actually worked.

* fix(macos): let the folder-access notice return after a poll that read no daemon

A daemon identity reads as null during any reconnect blip, and the poll reports that as
"no mismatch". The notice dismissed itself and then never showed again for that daemon,
because the once-per-daemon latch still held its scope. Only "Not now" should latch.

* fix(macos): say what the folder-access notice costs the user

One line read like a stray warning. The toast now says who is blocked and what fails,
and still leaves the steps to the fix dialog.

* fix(macos): give the folder-access toast one action and the X, like every other toast

"Fix" is the only button; the X dismisses. Sonner fires onDismiss for programmatic
dismissals too, so the post-restart takedown now goes through the store and the hook,
and only a user's X is counted as dismissed.

* fix(macos): keep the fix dialog's steps a checklist and put the one action in the footer

Buttons inside each step made the list look like a form, and a footer Close duplicated
the X. The footer now carries the active step's action, with a ghost Cancel; a probe
that could not answer says so under step 1 instead of showing a check.

* fix(macos): let the checklist show the fix landed instead of saying so

A hedged sentence addressed to the user read like chat. On success both steps check
off and the footer offers Done; the unanswered-probe helper is a status, not advice.

* chore(i18n): drop the fix dialog's unused close key

* Revert "chore(i18n): drop the fix dialog's unused close key"

This reverts commit 365915df48.

* chore(i18n): drop the fix dialog's unused close key

* fix(macos): tell step 1 what to do when the folder toggle is already on

Users who need step 1 usually find Orca already allowed in System Settings; the grant
is recorded but not honoured for the daemon. Re-toggling re-records it.

* fix(macos): drop the unverified toggle instruction from step 1

Nothing has been confirmed to fix a grant that is already on, so the step says only
what the probe knows.

* refactor(macos): share the tccutil reset and bundle-id read behind one module

Clearing a macOS TCC row is about to have a second caller: the daemon
folder-access fix (STA-7948) needs the exact `tccutil reset` the computer-use
helper already issues. Extract both it and the PlistBuddy bundle-id read into
src/main/macos-tcc-reset.ts so the two remedies cannot drift apart.

The extracted calls go through runProcessSync rather than a fresh
node:child_process import: the spawn chokepoint's ratchet holds the direct
importer count at a pin, and a new module with its own spawnSync would raise it.
Behaviour is unchanged except that both calls now carry a 10s bound, and the
computer-use test asserts the same argv against the chokepoint's options.

* feat(macos): offer a permission reset when restarting the terminal service cannot help

About a third of the users who see the folder-access notice are still denied by
a freshly forked daemon even though Orca itself is allowed under Files and
Folders, so the restart the dialog offers cannot fix anything for them. That
state previously had one action: open System Settings, where the toggle they
would look for is already on.

The denied state now offers "Reset permission". Main clears Orca's TCC row for
that folder class with tccutil, then reads the folder from the app itself so
macOS raises its prompt against Orca rather than the daemon, then forces a
fresh-daemon re-probe that bypasses the poll's reuse interval. The dialog
re-renders from that verdict: allowed turns step one green and offers Restart,
still denied says so, and a refused reset points back at System Settings.

Nobody has confirmed this remedy on an affected machine, which is why main emits
the re-probe's verdict as reset_outcome_allowed/still_denied/unknown. Those
three, plus reset_clicked, are the evidence that decides whether the feature
stays.

* fix(macos): say what the permission reset does, and keep System Settings as the fallback

Step 1 was labelled like a Settings task while the button did something else, with two routes
in the footer for one step. The denied state now names the step for what the reset does,
explains it under the step, and shows System Settings only after a reset fails or leaves
things blocked.

* fix(macos): count a folder-access restart only against evidence that survived

The stored denial is the prior denial, so a second copy of it outlived the
one event that retires it: a daemon that read its own folder back cleared the
entry but left the copy, and the next daemon's first denial was then reported
as a restart that had never happened.

Track the outcome on the entry itself, drop the spawn-path probe (ten denied
terminals forked ten probe children the focus-time poll re-runs anyway), and
stop emitting `shown` from a getter the reset path calls for data. The
renderer's toast latch is what decides a scope is shown, so it emits it.

Both accessors now read one identity-matched entry.

* refactor(macos): name the folder-access verdict instead of encoding it as a tri-state

`restartWillHelp: boolean | null` re-encoded a verdict the probe already
returns as a named union, so every reader had to remember that `false` meant
"Orca itself must be re-allowed" and `null` meant "no answer".

`freshDaemonAccess: 'allowed' | 'denied' | 'unknown'` says it, end to end
through main, the IPC payload, the preload mirror and the dialog. The reset's
outcome event becomes a lookup. No user-visible string changes.

* refactor(macos): give the folder-access notice one latch instead of three

Two refs in the hook and a field in the store tracked the same fact, and the
dialog reached the hook through a store field plus an effect just to take its
own toast down before sonner echoed the dismissal back.

The store now holds the visible scope and the scopes the user closed, and
exposes the three things that happen to a notice: it is shown, someone else
retires it, or the user dismisses it. The dialog calls retire directly and the
effect is gone. `settingsIsFallback` loses an argument that was always true at
its only call site, so it becomes the local it always was.

* refactor(macos): stop blocking main on the tccutil reset

Two spawnSync calls with a ten-second timeout sat inside an async IPC handler,
so clearing a TCC row held main's event loop for as long as either binary took.

Both now run through runProcess. The computer-use caller that shared them was
already async, so it awaits them.

* test(macos): run the folder-access probe script against real paths

Every other test mocks the spawn away, so the minified child script — the one
piece that duplicates enumerateDirectoryOnce's errno mapping — had no oracle.
It now runs against a temp directory, an absent path, a file, and a directory
whose mode withholds it, which is skipped for root and on Windows.

* refactor(macos): read the folder-access entry through one identity match

All four callers that ask "is this evidence still this daemon's?" now go
through the same private accessor, so the rule the canonical path depends on
lives in one place.

* fix(macos): keep folder evidence through a failed health read, and make a forced re-probe always probe

A rejected attribution-health read nulled the folder evidence on the same poll, which the
renderer read as "cleared". A forced refresh after a reset returned early on an older
settled verdict. The dialog also closes when a reset finds the evidence gone, and stops
showing the unverified helper once the restart is done.

* fix(macos): name the folder in the access-notice scope

One daemon denied two protected folders kept one scope, so the toast, the
fix dialog, and the tccutil reset could each be about a different folder.

* refactor(macos): derive the folder-access dialog from the latest verdict

The store held an `open` flag and a mismatch frozen at the moment the toast
was raised, so the dialog could open on a stale verdict and its remedy state
could survive a close. It now keeps the latest verdict and the scope the user
opened, and the dialog is shown only while the two agree.

* fix(macos): offer the permission reset only where there is a row to reset

A workspace symlinked out of Documents or on an external volume can be denied
too, and the dialog offered a reset that main refuses. One shared list of the
TCC-backed folder classes now decides both.

* fix(macos): give the permission prompt's read a deadline

An unanswered macOS sheet blocks the app's folder read for as long as the user
ignores it, and the fix dialog is modal and busy until that read returns. The
wait now ends after a minute and reports an unknown outcome rather than
probing under the sheet.

* fix(macos): count the folder-access notice once per scope

A reconnect blip reports no daemon, which takes the toast down and lets the
same scope raise it again. Both raises counted as separate notices, inflating
the denominator behind the affected-user rate. The two latches are now one
map from scope to phase, and the count follows first insertion.

* fix(macos): drop the restart warning once the restart is done

Step two ticked green while its helper still warned that open terminals and
agents would restart, which had already happened.

* fix(macos): keep the folder-access toast up when the fix dialog opens

Sonner deletes a toast after its action button runs unless the handler
prevents the event, and it does so without calling onDismiss. Clicking Fix
therefore took the notice off screen while the scope stayed latched as
visible, so cancelling the dialog left no way back to it.

* refactor(preload): reuse the shared daemon cwd class instead of copying it

The five folder classes were hand-mirrored in preload behind a comment saying
preload cannot depend on main-only modules. The enum lives in src/shared,
which preload already imports from elsewhere, so the copy could drift.

* refactor(macos): close the fix dialog when its evidence disappears

A null verdict left the opened scope set, so the same scope coming back
remounted a checklist nobody had opened. Clearing it on a null verdict also
makes the dialog's scope key redundant, so it goes.

* fix(macos): let each fix-dialog button report its own work

The footer swaps the reset for a restart as soon as a poll says the grant
landed, which can happen while the reset is still running. Both buttons read
their label off the dialog being busy at all, so the restart button appeared
spinning as "Restarting…" for a restart nobody had started.

* fix(macos): clear the reset failure once the permission is granted

"Couldn't reset the permission" stayed on screen after the user granted it in
System Settings and the probe read allowed, contradicting the ticked step
above it. Its sibling line was already gated on the same verdict.

* fix(macos): end the folder-access remedy with the evidence it is about

Two ways out were missing. A reset that cleared the evidence closed the dialog
but left the toast on screen, because only the poll retired it; the store now
retires the notice whenever a verdict comes back null, so both callers get it
and the hook's own branch goes. And the opened scope survived a verdict for a
different scope, so the original one returning later reopened the dialog with
nobody having asked for it.

* fix(macos): keep the folder prompt off main's spawn path

The app-side readability check moved from accessSync to opendir when the
notice was added. TCC lets accessSync through but gates opendir, so on a
machine that has never granted Orca the folder, spawning a terminal there
raised the macOS sheet and froze main until the user answered it. The read is
async now and the spawn no longer waits for it. The blocking variant keeps a
name that says so, and the reset module's own copy of the read is gone.

* fix(macos): only say a folder is still blocked when something re-read it

Two paths reached "Still blocked after the reset." with no verdict behind it:
an unanswered prompt, where the reset returns the verdict stored before it
ran, and a re-probe that could not answer. The reset now returns the same
access it reports to telemetry, and the line waits for a real denial.

* refactor(macos): let the folder-access refresh decide when to skip itself

The poll handler re-implemented the refresh's own two guards, a null entry
and a settled allowed verdict, so each had to be kept in step by hand.

* fix(macos): stop the daemon blocking on its own folder read

The daemon reads the requested cwd before forking a shell to report whether
it can list it. That read is the one macOS gates, so on a folder the daemon
is refused it could hold the daemon's event loop behind a prompt. It is
awaited now, which leaves the blocking enumerator with no callers.
2026-09-22 01:09:26 -04:00
eb92222e7f feat: support Antigravity as supervised worker (#21705)
* feat: add supervised Antigravity worker support

* fix: address Antigravity worker review findings

* fix: stabilize Antigravity readiness detection

* fix: allow Antigravity resume footer after readiness

* fix(antigravity): make agy reach worker_done as a supervised worker

Three defects each blocked `orchestration worker-start --agent antigravity
--worktree new-child` at the agent_readiness stage.

1. Readiness never fired. The composer check required the trimmed line to be
   exactly one character, but agy 1.2.7 launches in accept-edits mode and paints
   it into the caret row (`> Accept-edits mode: ...`). Widened narrowly to a bare
   `>` or `> <name> mode:`; matching any `> <text>` would make every menu dialog
   read as ready, since they all prefix their highlighted row the same way.

2. No trust artifact for agy. Added markAntigravityWorkspaceTrusted, writing
   ~/.gemini/antigravity-cli/settings.json under `trustedWorkspaces` — verified
   empirically against agy 1.2.7, and distinct from the Gemini CLI's
   trustedFolders.json, which agy does not consult. Trust is exact-path and not
   inherited by subdirectories, so each child worktree needs its own entry.

3. The orchestration path skipped the preset. Orca has two trust dispatch
   chains: the renderer's preflightAgentTrust and the main-process
   markLocalWorktreeTrusted. worker-start only takes the second, which matched
   cursor/copilot/codex and fell through for antigravity, so the trust write
   never happened while renderer-side tests passed.

Verified live end to end: the dispatch settles `succeeded` with worker_done
carrying the right task and dispatch ids, and the worktree is appended to agy's
settings with sibling keys untouched.

Known gap: remote-agent-trust-presets.ts has no antigravity branch. The SSH
artifact path is unverified, so agy over SSH still stalls at agent_readiness.
Recorded in a comment there rather than guessed at.

* fix(antigravity): wire trust preset through preload safely

* fix: preserve Antigravity readiness across transcript tails

---------

Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: LielinaH <lielinah@gmail.com>
2026-09-21 20:22:08 -07:00
Jinjing 7da9788c83 fix(updater): send the gh token and cache the release picker's build list (#21902)
* fix(updater): send the gh token and cache the release picker's build list

The dev build picker listed releases through api.github.com with no
Authorization header, so it spent GitHub's 60/hour per-IP bucket that every
unauthenticated caller on the same network shares, and it refetched on every
settings mount and channel click. When that bucket ran dry the picker showed
"No builds found" with a rate-limit line even though GitHub was healthy and
the user's own token had its full quota.

Attach the local `gh auth token` when there is one so the request draws from
the user's 5000/hour bucket, fall back to unauthenticated on a rejected token
or a spent token bucket, cache the list per channel for five minutes in the
main process (the refresh button forces a reload), classify 403 by the
rate-limit headers, and say when the limit resets.

Fixes #21898

* fix(updater): don't trip breaker for secondary rate limits

GitHub sends x-ratelimit-remaining: 0 on both primary and secondary
limits. Secondary limits carry Retry-After and shouldn't block all core
gh commands — only the primary limit should trip the shared breaker.

* Scope gh rate limits to execution environment

* Add build list cache hint to release channel settings

Inform users that build lists are cached for 5 minutes and they can
refresh to check for new builds immediately. This makes the cache
behavior visible and explains why a manual refresh is necessary to
bypass the cache.
2026-09-21 11:52:37 -07:00
Neil e8a7be4ce2 fix(omp): recover retired pane status with validated restart authority
Merged after fresh run 35448889017 passed all required checks, including static analysis, typecheck, package jobs, all test shards, changed E2E, Docker SSH E2E, and verify.
2026-09-19 08:05:33 -07:00
Neilandstevelliu ea02d90704 fix(omp): answer startup Kitty queries before renderer handoff (#20620)
* fix(omp): answer startup Kitty queries before renderer handoff

Forward actual renderer capability through local and remote spawn. Preserve source ranges and following keyboard mode pushes, and retain independent ConPTY color authority.

Refs #17081. Secondary review: #17082.

Co-authored-by: stevelliu <stevelliu@tencent.com>

* test(omp): cover fragmented keyboard modes and ConPTY handoff

* fix: preserve keyboard startup intent without terminal colors

* fix: negotiate keyboard support for host-authoritative agent launches

* fix: keep terminal creation within line budget

* fix(omp): negotiate keyboard support for paired web launches

* test: remove obsolete message type import after main integration

* fix: validate paired launch results and retry incomplete SSH test snapshots

* fix(omp): negotiate keyboard support for background paired launches

---------

Co-authored-by: stevelliu <stevelliu@tencent.com>
2026-09-19 01:29:16 -07:00
c34b944136 feat(github): bind projects to a specific gh account (#13664)
* feat(github): bind projects to a specific gh account

Adds per-project `Repo.ghAccount` so repo-scoped gh calls (create-worktree
issue/PR search, work items, hosted-review reads and mutations) run as the bound
account via ephemeral child-env token injection instead of the globally active
gh login. Multi-account resolution is capability-gated (gh >= 2.40) and fails
closed when the bound account or host is unavailable; Project View stays
ambient by design.

Repository settings gains a section for selecting or clearing a keyring-backed
account (shadcn `Select`), with mixed-version "not enforced" handling for older
remote runtimes. Attached `-Rhost/owner/repo` forms are covered by the host-drift
guard and its tests; es/ja/ko/zh catalogs carry the section's strings.

`getLocalProjectGhExecOptions` centralizes the binding lookup so every gh
execution path picks it up, including the Electron `hostedReview:*` handlers
that previously stayed on the ambient login. `gh auth token` (a keyring read)
is exempt from the rate-limit breaker gate so a tripped bucket cannot turn a
bound-token resolve into a false "unavailable".

The `ghAccount` update field and the two binding RPC methods live in the shared
RPC params contract; the generated catalog is regenerated.

Fixes #13612

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012B3QEP5iP4WGGEpPLtkHqA

* fix(settings): make GitHub account refresh secondary

* fix(github): satisfy strict casting quality checks

* test(rpc): use runtime fixture for repo binding

* fix(github): preserve project account for PR worktree lookups

* test(rpc): avoid incomplete runtime settings fixture

* fix(i18n): add GitHub account refresh label

* fix(i18n): refresh runtime required catalog

* fix(windows): preserve mobile patch bytes

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Neil <neil@stably.ai>
2026-09-18 20:39:52 -07:00
1ff4fe677c fix(main,preload): tear down renderer relay and preload listeners (#20909)
* Clean up renderer relay listeners on teardown

* fix(main): guard empty markdown relay results

* test: document relay window test double safety

* fix(relay): retain web contents through window destruction

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-18 00:18:19 -07:00
Neil 4e3170a76e fix(accounts): free the account queue when a sign-in is abandoned, and show the Codex sign-in link (#21372)
* fix(accounts): free the account queue when a sign-in is abandoned

Closing Settings mid sign-in left the `codex login` / `claude auth login`
child running, and every account mutation shares one FIFO queue, so the
next Add Account sat behind it for the login's whole deadline and then
inherited the abandoned call's timeout toast.

Cancel the pending login before enqueueing the next add or reauth (never
inside the queue the abandoned login owns), give Codex the cancel handle
and Cancel button Claude already had, and stop reporting a cancellation
as a failure.

Also surface the sign-in link Codex prints, with copy and open, so the
flow can be finished in a private window or another browser profile.

* test(accounts): drop the bare casts CI's changed-code gate rejects

The service doubles still need a cast; one documented helper per file
carries the SAFETY rationale instead of nine bare `as never`s.

* fix(codex): a cancel must not discard a sign-in that already succeeded

The Windows post-auth watcher gives a lingering codex login five seconds
to exit after it writes auth.json. A cancel arriving in that window
rejected the login, and the caller's rollback then deleted the managed
home that had just authenticated.

Refuse the cancel once new credential bytes exist: there is nothing left
to cancel, and the close handler already treats that state as success.

Found by review of #21372.

* fix(codex): keep a refused cancel cancellable, and require the sign-in notice

Review of the auth-aware cancel guard found two holes it opened:

- The outer handle latched `cancelled` before asking the session, so a
  refusal killed cancellation for the rest of the deadline. On a host
  with no post-auth watcher that reinstated the very stall this PR
  removes. Latch only when the cancel is accepted.
- WSL never reads a pre-spawn baseline, so the guard read the auth.json
  that was already there and refused from the first click, making a WSL
  reauthentication uncancellable. Require a baseline before refusing.

Also from review: publish the sign-in link from a stdout-only buffer, so
an interleaved stderr chunk cannot truncate it; require codex's own
"navigate to this URL" notice rather than offering the first link in the
output; hide the notice in a remote account scope, where it would name a
login running on this desktop; and share the cancellation message
instead of matching a duplicated literal.

The Claude case joins the login-process suite that already owns the two
neighbouring cancel cases, and the auth-snapshot helpers move out of the
session file, which the additions pushed over the line cap.

* refactor(codex): cut the sign-in-link plumbing to its smallest form

Review found the change correct but larger than it needs to be:

- The pending-link store was a class with one permanent subscriber, a
  never-called unsubscribe and a try/catch that could not fire. It is a
  field and a listener set on the service, beside the cancel handle it
  already owned — and the service now clears both in one place.
- The optional login-session dependencies were always supplied.
- The parser's https check could not fail; the pattern already fixed the
  scheme. The renderer's unmount guard inside a synchronous IPC listener
  could not fire either.
- The broadcast channel and the cancellation message are single sources
  of truth in src/shared now, rather than exported next to a hardcoded
  copy of themselves.
- The duplicated seven-line rationale in both services says the same
  thing in three, including why only add and reauthenticate supersede.
- The codex suite reuses its own factory, and unmocks once.

Also reverts four reformat hunks the formatter pulled in around edits.

* fix(accounts): free the queue for a switch, not only for another add

Switching or removing an account shares the mutation queue an abandoned
sign-in was holding, so the commonest thing a user does after giving up
— pick a different account — still spun for the whole deadline while Add
recovered instantly. Both now supersede, as does the Claude side.

Every caller is a person: the two IPC handlers and the mobile RPC
methods. No poll, sync or CLI path reaches them, and a sign-in that
already wrote credentials refuses the cancel, so a switch cannot discard
one that succeeded.

Also from review: the Cancel button regains the gap its Claude twin has
(layout is allowed by the design-system rule; only the colour override
was not), and the URL subscription says what it is — registration for
the process's lifetime, with no teardown to hand back.
2026-09-17 23:08:59 -07:00
0e3b71f605 fix(session): give an SSH workspace one owning partition so its tabs stop round-tripping as deletions (#19572)
* fix(session): give an SSH workspace one owning partition so its tabs stop round-tripping as deletions

`workspaceSessionPartitionHostId` answered differently depending on who asked: the
renderer mapped an SSH worktree's session to the `local` blob, the main-process
runtime read-modify-wrote `ssh:<targetId>`. One workspace's session lived in two
stores and no reader reunited them, so whatever landed on the unread side did not
read as unknown — it round-tripped as absence. The remote-workspace upload is a
`replace-session` patch, which turned that absence into deletion on the host, and
the next pull applied the deletion locally and re-poisoned the snapshot.

Collapse the two answers into one: every non-'local' host owns its partition.
Boot hydration and the export fallback now read the SSH partition, and rows a
shipping build left in `local` are folded back in once, gap-filling only — an
empty tab row is a gap, never proof that anything was closed.

Folder workspaces deliberately keep their existing 'local' routing: boot
discovers SSH partitions from the repo catalog, so an SSH target that owns only a
folder workspace has no partition any reader enumerates. They are still adopted
back out of an SSH partition when a repo does name the host.

Fixes #12721
Supersedes #12722

Co-authored-by: Robert Nisipeanu <github@nisipeanu.com>
Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>

* test(session): pin the old-client empty-publish skew direction

* fix(session): adopt every workspace the host partition names, not only tabbed ones

Review caught that gating adoption on `host.tabsByWorktree[key].length > 0` traded
the #12721 deletion for a narrower one. The write path routes EVERY worktree-scoped
field to the owning partition, so an SSH workspace with open editor files or browser
tabs and no terminals had all of it dropped on every restart — and unlike terminal
state it cannot be recovered from the host snapshot, which carries terminal fields
only, so an unsaved `dirtyDraftContent` was destroyed outright.

The defect was not a missing field. It was a hand-maintained field list deciding what
the read recovers while the write used the ownership table, so the two could disagree.
Adoption now walks `WORKSPACE_SESSION_FIELD_OWNERSHIP` with an exhaustive switch, and
a new ownership kind is a compile-time decision rather than a silent omission.

Session keys are normalized through the shared `normalizeWorkspaceSessionKeyToWorkspaceId`
so host-qualified visit recency (`ssh:target|worktreeId`) reaches its workspace, and the
regression is pinned by feeding the shipping split's own output back through the real
boot read rather than a hand-built fixture.

Co-authored-by: Robert Nisipeanu <github@nisipeanu.com>
Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>

* fix(session): stop adoption overwriting rows it was never told about

Three losses, one cause: the reader walks its own description of the
partition layout while the writer walks another, so the two agree on
which ownership kinds exist and not on what a kind means.

- An empty host row replaced a populated base row, destroying an unsaved
  dirtyDraftContent the header comment says must never be destroyed. The
  host holding nothing is not evidence the base is wrong.
- A contested bare id was adopted as if local and ssh:<target> were one
  workspace written twice, which is exactly the id where that premise is
  false. The read already reached that verdict and adoption could not ask
  for it, so it is passed in; contested keys are gap-filled, never
  replaced. mergeWorkspaceSessionsWithHostShadow now reports the real
  contested set, which primaryHostBySessionKey never was.
- Tab-, pane- and file-keyed rows are adopted through the split's own
  indexes, so unified-only tabs come back and the pane key is parsed once.
- A bare lastVisitedAtByWorktreeId key only fills a gap; the split has a
  dedicated branch for that field and the reader had none.

* test(session): pin the tombstone/gap boundary the two readings meet at

An explicit empty tabsByWorktree row means the user closed the last
terminal; adoption reads an empty base row as a gap to fill. Same value,
opposite readings, so the boundary is asserted rather than argued: the
tombstone lands in the owning partition, restores as a present empty row
rather than a deleted key, is declined by the real seeding predicate, is
published as an empty list, and the legacy-transition resurrection
happens once and cannot recur.

* docs(reliability): record the adoption guards and the tombstone boundary in the gate

* test(e2e): read the SSH restart assertions from the partition that owns them

ssh-cold-activation-restore asserted persistence through session.get()
with no host, which is the local partition an SSH worktree's rows no
longer live in. The invariant it means to check is that the state is
persisted where the boot read will find it, so it now unions local and
ssh:<targetId> and stays correct on both layouts.

Confirmed the product invariant separately rather than by the edit: the
behavioural half of both tests - the full app restart, the active
worktree, the eager terminal remount and the PTY-owner reclaim against a
real Docker OpenSSH host - runs after this check and passes. 2 passed in
48.9s.

* test(e2e): read ssh-restart-tab-accumulation from the owning partition too

Same layout-coupled read as ssh-cold-activation-restore: the pre-quit
flush asserted through session.get() with no host. Verified against a
real Docker OpenSSH target - both repeated quit/relaunch cycles keep
exactly the restored SSH tabs, no accumulation and no loss. 2 passed in
52.9s.

* fix(lint): clear the casting gate on the partition adoption

main tightened typescript/consistent-type-assertions to assertionStyle:
never, which the rebase brings onto these added lines. Most of the
round-trip fixtures did not need a cast at all -- three were hiding
wrong-shaped literals (a browser workspace keyed 'name', a unified tab
keyed 'type', a layout keyed 'direction'), now written as the types they
stand for. The adoption reads narrow through an isRecord predicate
instead of casting, which also stops a null entry throwing out of
Object.keys. What is left is dynamic-field writes and unknown-typed IPC
returns, each with its own SAFETY rationale.

* fix(session): give an SSH folder workspace one owning partition boot can find

The partition owner rule already names `ssh:<targetId>` for a repo-backed worktree, but
`getFolderWorkspacePartitionHostId` still answered 'local' for a folder workspace while
main's `RuntimeWorkspaceSessionController.getPreferredHostId` answered `ssh:<targetId>`
for the same key. That is #12723 unfixed for folder workspaces, and once the renderer
started writing `ssh:*` at all it got worse: a save's field-level patch carries only the
rows routed to that partition, so a `tabsByWorktree` write without the folder row erased
the row main had put there.

The reason the renderer could not route there was real - boot discovered SSH partitions
from the repo catalog, which cannot name a target whose only workspace is a folder. So
persistence now answers that directly over `session:list-host-ids`, and boot reads the
partitions that exist rather than the ones a catalog implies. Removing a folder workspace
prunes its rows from the owning partition too, or the census would adopt them back on the
next launch as a workspace the user already deleted.

Adoption now decides from the repo catalog instead of from co-presence. Two partitions
holding one bare `repoId::path` is not evidence of a collision - that is the exact shape
the repair exists for - so the verdict comes from `resolveWorktreeExecutionHost`: a repo id
registered on more than one host is contested and may only be gap-filled, and one the
catalog positively resolves to a different host is residue this partition does not own and
is not adopted at all. Without the second rule a stale partition sorting first won the read
and was then written into the live one. Nothing is deleted either way; the rows stay where
they are.

Finally, a workspace adopted out of a partition now routes back to that partition. Routing
used to re-derive an owner from the catalog, so a boot whose repos had not hydrated moved
the rows it had just reunited back into 'local' and re-stranded them. Contested ids are
withheld from that override, because routing the whole bare id to one host is the loss the
gap-fill prevents.

The publish path resolves each workspace's owner once for the whole publish, shared with the
projection, so the per-target catalog attribution does not repeat it per connected host.

* fix(session): drop a deleted workspace from every partition, not just the local blob

Adversarial review of the previous commit found three ways the partition census - which now
reads whatever persistence holds rather than what the repo catalog implies - keeps rows alive
that nothing should keep alive.

`deleteProjectGroup` pruned only the local blob, so every folder workspace under a deleted
group left its rows in `ssh:<targetId>`; the next boot adopted them back, named that partition
their owner and wrote them there again, forever. `removeFolderWorkspace` had the same hole for
a workspace whose partition its host expression could not name: main never persists a folder
workspace's `executionHostId`, and `RuntimeWorkspaceSessionController` can infer a connection
from the group's repos that the workspace row itself does not carry. Deriving the partition at
delete time is the wrong question - a deleted workspace owns nothing anywhere - so both paths
now remove it from every partition.

The third is on the read side. A contested id is deliberately withheld from the read-source
override so the write cannot carry one host's rows into another's partition, but the routing
that then re-derives an owner answers 'local' for an id the catalog cannot name. Adopting such
a row moved it out of the partition that owns it and into the blob: the two-store split this
change exists to remove. A contested id the assembled session holds no row for is therefore not
adopted at all. Gap-filling stays available for a contested id the session already names, since
that row's own partition is what the write follows. Declining to adopt leaves a row invisible
for one boot; it never deletes one.

Also: the folder-key guard in both catalog attributions was dead, because
`getRepoIdFromWorktreeId` hands back the whole key rather than nothing when there is no `::`.
The verdict was right and the resolution wasted; it now skips by shape. And the two type
assertions the changed-code casting gate rejected are gone rather than suppressed.

* fix(session): park the rows a partition read declines instead of letting the next write erase them

A partition write replaces each field with exactly what the unified session routed there. So a row
the read left out of that session is erased from its own partition the moment any sibling workspace
writes the same one - and with SSH partitions now the owning store, that row is then in no partition
at all. Three separate decisions produce such rows: residue the catalog attributes to another host,
a contested id withheld so the write cannot carry one host's rows into another's partition, and a
workspace the base already holds the live copy of. Declining to show a row was quietly deleting it.

The machinery for this already exists. `attachHostSessionShadow` writes a contested runtime
co-claimant's parked rows straight back into its own slice before the write, so the primary's write
cannot erase them; the ssh partitions simply were not among the slices the contention split
arbitrates. The read now parks everything it is not returning to an ssh partition into that same
shadow, and the existing re-attach puts it back. Leak, never kill - docs/reference/ssh-execution-
boundary.md - and a row no partition holds is unrecoverable.

Second, the contested branch of the tab adoption read `Object.hasOwn` as "the base has tabs here".
An empty list satisfies it, so whenever a legacy id happened to be contested, #12721's empty local
row won over the host's real one - the exact reading the module's own header, and the gate invariant
it is pinned by, say is wrong. An empty row is the gap this repair fills, so it is now treated as
one.

* test(session): pin the empty-base-row gap for a contested id

Mutation testing found the assertion missing: reverting the gate to `Object.hasOwn` left all 39
assertions passing, which makes the fix that reads an empty base tab row as a gap unguarded. The
#12721 shape does not stop being a gap because the id happens to be contested.

---------

Co-authored-by: Robert Nisipeanu <github@nisipeanu.com>
Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
2026-09-16 22:24:33 -07:00
Brennan Benson aee98ccaa0 fix(browser): make the browser identity one process-wide choice (#13822) (#20767)
* feat(browser): process-wide browser identity, chosen before ready

Electron resolves worker identity from a single process-global default, so two
coherent identities cannot coexist in one process. This makes clean/native one
app-wide decision read before `ready`, instead of a per-profile one that leaves
documents on one identity and every worker request on the other.

Both identities are load-bearing, measured across four origins at five reps:
the cleaned identity clears an embedded Turnstile widget and WhatsApp's browser
check where native is refused; native clears a full-page Cloudflare interstitial
that the cleaned identity never clears.

Base commit only: removing the per-profile field, its settings surface, and the
migration notice follow.

* test(browser): cover cross-context UA wire identity

* refactor(browser): make user agent identity app-wide

* test(browser): repair process identity wire fixture

* Fix browser identity startup migration failures

* WIP: rescue in-flight reduced-design work from a dead worker

Worker ctx_cb5b1262d7fe stopped ~2h ago mid-implementation (last heartbeat
2026-09-14T22:48:06Z) leaving this uncommitted. Committed unverified to make it
recoverable; not reviewed, not necessarily green.

* fix(browser): repair the rescued identity work so it typechecks

Finishes the interrupted edits in 7db9c54b54:

- browser-user-agent-migration-notice.ts was truncated mid-write; close the
  then() callback so the file parses.
- Register browser.identity.get/set in the generated RPC params catalog so the
  params type-parity gate is satisfied.
- Retire the persistence assertions for the superseded design: a
  migratedNativeProfileIds event map, a notice-acknowledgement clear, and a
  global persistence-failure accessor. Legacy userAgentMode bytes are retained
  now, so these assert retention plus a failed notice write still hydrating.
- The in-memory fs fixture threw a codeless ENOENT, which reads as "unreadable"
  rather than "missing" and made every identity write refuse. Carry the code.
- Use the segmented control's per-option disabled rather than adding a
  control-level prop it does not have.

* refactor(browser): make the identity store the only writer

The rescued work already serialized identity writes, but the writer lived beside the pre-ready reader, so nothing stopped a second caller from writing the record directly -- which is the shape of the bug this change set removes.

browser-identity-mode-record.ts is now read-only: record shape, path, parsing and the pre-ready synchronous read. browser-identity-mode-store.ts owns every mutation behind one queue, holds the snapshot and listeners, and derives restartRequired from appliedMode vs configuredMode rather than storing it. Consumers move to the store.

The two identity RPC methods also move out of browser-core.ts into browser-identity-rpc.ts: they read and write this host's own process identity rather than driving a page, and browser-core.ts was over its line cap. The generated params catalog is byte-identical.

* feat(browser): make resetting unhealthy identity data explicit and lossless

A corrupt or newer-version record left the identity unchangeable with no way out. An explicit reset now copies the old bytes verbatim to a fresh unique path before publishing a replacement, and refuses the whole operation if that backup cannot be written -- so the reset can never be the thing that loses the data. Nothing resets automatically.

Future-version data says update Orca rather than reporting corruption. Reset is opt-in via browser.identity.set and orca browser identity set --reset.

ProfileCreate and BrowserIdentitySet move to browser-identity-params.ts: both carry the per-profile to app-wide identity move, and browser-params.ts was over its line cap.

Also registers browser as a top-level CLI name so the Windows launch redirect covers it -- without it orca browser identity get boots the GUI and exits silently there -- and adds the canonical browser identity show alias the CLI vocabulary policy requires.

* feat(browser): advertise the identity capability only where it exists

browser.identity.v1 was static, so every host claimed it including one that never initialized the identity store, where both methods can only throw. It now follows the browser.headless.v1 precedent and is pushed at status time when the store is actually initialized.

Also covers the retired profileCreate userAgentMode field at the dispatcher rather than only at the schema, so an older client provably gets the changed-semantics rejection over the wire instead of a success with the field quietly dropped.

* refactor(browser): delete the identity write queue and guard backup uniqueness

The queue could not be falsified by any test: writeRecord is synchronous end to end, so two calls cannot interleave and removing serialization entirely left every store test green. Carrying machinery whose guard is unconstructible is what the design review told us to cut, so it is gone. If durable writes ever become async, serialization comes back with the change that makes it testable.

The test that claimed to prove serialization now states what it actually pins -- the later of two selections is the one that survives -- and the module doc no longer claims a queue that is not there.

Adds the guard that was missing on reset: two resets across separate launches must produce two distinct backups, each holding its own original bytes. Verified discriminating -- a fixed backup filename fails it.

* test(browser): guard the identity capability and harden two weak assertions

Pins the mixed-version guarantee that had no test: browser.identity.v1 is advertised when the identity store is initialized and absent when it is not. Verified discriminating -- advertising it unconditionally fails the test.

The profileCreate rejection test asserted ok:false against a runtime with no browserProfileCreate, so that assertion passed even when the retired field was accepted. It now stubs a working runtime method, making ok:false load-bearing, and asserts the runtime is never reached.

Removes the persistence fixture's dead failIdentityWrite branch on writeFileAtomically: nothing on that path calls it, so it implied a second write mechanism that does not exist. Failure is injected through node:fs, which is what the identity write actually uses.

* test(browser): classify the identity channels on the preview seam

The channel split is asserted total, so adding browser:identity:get/set left it
short by two. They manage the host's own process-wide user-agent choice rather
than acting on a guest the reader is looking at, so they sit with the session
and profile channels, not the preview tools.

* test(browser): audit the identity rig's global-fetch call sites

The wire probe server and CDP collector arrived with the cross-context coverage
and were never added to the audit list. The collector's two real call sites are
safe: the poll cancels its unread body and the version probe consumes it through
response.json(). Every hit in the probe server is inside an injected page or
worker script source string, not a call this process makes.

* fix(browser): strip an app name that contains a space

app.setName decides the app token in the user agent, and dev sets "Orca Dev".
The cleaner matched a single whitespace-delimited token, which cannot span that
space, so the replace failed outright and every dev build presented
"Orca Dev/1.4.203" on the wire — the exact token class that gets transplanted
sessions revoked.

Anchoring on the engine comment and consuming lazily up to Chrome/ removes any
number of app tokens. A user agent without that comment is returned unchanged
rather than mangled, because over-stripping is worse than under-stripping.

The function had no unit test at all; it was only exercised through the
real-Electron wire tests, which run with a single-token fixture name. That is
why this survived.

* fix(browser): anchor the cleaner on the gap before Chrome/

My first attempt anchored on the engine comment, which broke a startup fixture
whose platform comment is "(Test)" with no "(KHTML, like Gecko)" at all — the app
token survived and the ordering test went red.

Anchoring on the nearest ")" before Chrome/ and consuming only non-")" tokens
keeps the match inside that gap, so it handles a multi-word app name, a synthetic
platform comment, and an already-clean identity alike. A user agent with no such
gap is still returned unchanged.

The fixture shape is now a test case, since it is what caught the first attempt.

* test(browser): repair the cleaner's case table

A missing comma between two it.each elements was reformatted into an index
expression, collapsing the table so every case ran with undefined input.

* test(browser): make a CI-only capture failure diagnosable

This probe passes locally and fails on CI with an empty receipt set, an empty
CDP diagnostic list, and a fixture that still exits 0 — so the assertion message
carried nothing usable. Thread the fixture's own result and stderr into the
capture assertion so the next run says what the fixture actually did.

* fix(browser): let an explicit choice retire the migration notice for good

The retired per-profile userAgentMode bytes are retained on disk by design, so
every launch rediscovers them and re-arms the notice — including the launch
right after the user answers it, and every launch after that. Documented as
one-time, it was permanent.

The record already carries explicitSelection, which is exactly the fact that
should end the notice. Gate the mark at the single writer rather than deleting
the legacy key, so the retained bytes stay untouched and disk never claims a
notice is pending beside a choice the user already made.

The new test pushed the persistence suite past max-lines, so the in-memory fs
and module mocks move to a named fixture module and the retired-identity tests
move beside them in their own file.

* fix(browser): stop reporting an unhydratable profile as a retired choice

A profile that fails validation for a reason unrelated to identity — a non-UUID
id, a mismatched partition — armed both the notice and its degraded flag. Since
hydrateFromPersisted skips such entries silently and nothing ever repairs them,
the user got "an old browser identity choice could not be inspected" forever,
about a profile that never carried one.

Key the notice on the presence of userAgentMode instead, and use validation only
to decide whether the choice that was found is inspectable. Refusing to hydrate
an entry and finding a retired choice are now separate facts.

The old case table asserted the defect for null, 42 and 'broken', so it is
replaced by two tables stating the new contract rather than adapted to pass.

* fix(browser): stop rewriting worker requests for viewport emulation

A worker request carries no webContentsId, so it always took the session-wide
branch and picked up the mobile UA if any tab in the session had a mobile
preset. That made a single context disagree with itself: a desktop tab's shared
worker reported a desktop navigator.userAgent — the per-target CDP override
cannot reach a worker — while its fetches left as CriOS. It also leaked across
tabs, and closing the emulated tab silently reverted it.

On main the divergence was between contexts, each internally coherent. Making
one context internally inconsistent is worse by this PR's own standard, so
accept that viewport emulation reaches documents only. Workers keep the session
identity on the wire, which is the identity they report in JavaScript.

That left hasSessionMobileViewportIntent with no reader, so the map it fed and
its three accessors go too, rather than leaving a dead latch behind the guard.

The electron fixture models this rule in its own header hook, so its hook and
both mobile arms are rewritten around the invariant that each context's wire
identity equals the identity its own JavaScript reports — not adapted to keep
the old path list passing.

* test(browser): point the identity tests at keys and writers that exist

browserUserAgentMode appears in zero production files and zero commits on main;
`git log -S` finds nothing. The retired key is profile.userAgentMode inside
browser-session-meta.json. Two tests were built on the invented one.

The global-settings test is deleted rather than repointed: no browser identity
key has ever lived in global settings, and stripRetiredGlobalSettings strips
only three unrelated keys, so the test asserted that an arbitrary unknown key
survives an object spread — a fact about the normalizer, not about identity.

The ready-phase test asserted on writeFileAtomically while the identity store
writes through writeFileDurableSync, so it could not go red for the write it
existed to forbid. It now watches the real writer, matched on the record path so
an unrelated durable write cannot fail it for the wrong reason, and the invented
settings key is gone from the Store mock.

Proven by ablation: injecting a byte-identical rewrite of the record into ready
composition leaves every snapshot and record assertion green and is caught only
by the new assertion, while writeFileAtomically is never called.

* fix(browser): let an unavailable process identity reject instead of throwing

installBrowserSessionPartitionPolicies returned Promise<void> without being
async, and configures the user agent policy before any suspension point.
getBrowserProcessUserAgentIdentity throws when the process identity was never
initialized, so that throw escaped synchronously past every caller's handler:
`void install(...).catch(...)` in the registry, and a bare `void install(...)`
in the route policies, which has no handler at all.

Bookkeeping must never gate a user action. Session startup would have died on a
failure its callers were already written to absorb and report.

* docs(browser): scope the meta-store claim about dropped legacy keys

The comment said persistMeta drops legacy keys on the next write because the
loader no longer carries them. That holds for the top-level userAgent keys it
describes, but not for the retired per-profile userAgentMode: it sits inside
each BrowserSessionProfile in `profiles`, which is carried through untouched, so
those bytes survive every write.

Retaining them is deliberate — it is what makes rollback and data-loss machinery
unnecessary, and the startup notice keys on their presence — so the comment read
as broader cover than it provided, in the one place someone would look before
deciding it was safe to strip them.

* test(browser): pin the unmapped-webContents path beside an emulated tab

A popup carries a webContentsId that maps to no registered tab, so it resolves
through the same branch as a worker request that carries none at all. The branch
already handled both, but only the absent-id case was covered.

* test(browser): make the ordering fixture exhibit a multi-word app name

This file sets the dev app name to "Orca Development" and then used a
single-token user agent fixture, so it set up the multi-word scenario and used a
fixture that could not exhibit it — which is how the multi-word app-name leak
got through. The fixture now carries a two-word app token, matching what
app.setName produces in dev, and the assertion names both words: a single \S+
match would leave "Orca" on the wire and still pass a one-token check.

* test(settings): cover the local branch of the browser identity setting

The only existing test covered the remote-host branch. The local branch — load,
select, refused write, and reset-required — had none, and that is the path the
retired-identity notice sends users down to make the choice that retires it.

Covers the selected-mode render, the commit that reports restartRequired, a
refused write surfacing its message without showing the mode as changed, and the
reset-required state offering no control.

* test(browser): run the real registry path in the ready identity pin

The test stubbed browser-session-startup and browser-session-registry, which are
the one ready-phase path that can write the identity record, so the record
content assertion could not fail for the write it existed to forbid.

Both are now real. Only the pieces hanging off the identity path are stubbed —
partition policies, route sessions, cookie staging, webauthn — so the meta load,
the retired-choice inspection, the identity store and the durable write all run
for real against temp directories. The canonical path mock moves to
persistence/loading-store/user-data-path, which is where the registry reads it;
mocking persistence alone left the registry pointed elsewhere. The active
profile directory is now a real temp dir, so the seeded browser-session-meta.json
is actually found — against the old /test-profile literal the meta load found
nothing and the whole exercise would have been vacuous.

A third case proves the path is live: with no explicit choice, the same retired
profile arms the notice through ready and lands migrationNoticePending on disk.
The two authority cases assert the opposite, that an explicit choice leaves the
record untouched.

initializeBrowserSessionsForApp latches on module state, so each case resets
modules and imports ready dynamically.

Ablated: disabling the explicitSelection gate turns both authority cases red on
the record content assertion while the arming case stays green.

* fix(browser): reject an unrecognized identity mode at the IPC door

normalizeBrowserUserAgentMode turned any unrecognized value into 'clean', so the
IPC door reported success for a mode it had quietly replaced, while the RPC door
validates against z.enum(['clean', 'native']) and rejects. One concept answered
an unknown value two different ways, and a future mode name was silently
downgraded rather than refused.

The handler now rejects, which is what the RPC door does and what the renderer
already handles — its catch puts the message in the error slot. Returning a
result instead would have meant inventing a fourth error code for a case no
legitimate caller can reach.

normalizeBrowserUserAgentMode had no other consumer, so it goes with the change:
leaving a coercion helper called "normalize" in shared/ invites the behaviour
straight back in.

* fix(settings): name the reset command where identity data is unusable

When configuredMode is null the setting says identity data must be reset
explicitly and then offers no control, because the reset overwrites data that
may belong to a newer Orca. The only escape is the CLI, which the message never
named — so it told the user to do something and gave them no way to do it.

Copy only: one line naming the command, no control and no destructive action in
the UI. The command goes in a new key beside the existing sentence rather than
expanding its default, which keeps the already-translated string valid.

No en.json entry: this component has no catalog entries for any of its keys, so
English resolves from the call-site defaults and adding one only for the new key
would be inconsistent with its siblings.

* fix(i18n): add the browser identity keys to the localization catalog

* fix(i18n): regenerate the runtime-required English catalog

* fix(browser): attach nested CDP targets paused before enabling Network

An OOPIF or dedicated worker was reached only through Target.targetCreated plus
an explicit attachToTarget, which never pauses the target. The frame could issue
its subresource fetch before Network.enable took effect, so the capture came back
empty and the cross-context assertion failed under CI load.

Re-arm auto-attach on each attached session, filtered to nested target types, so
an OOPIF or worker arrives waiting for the debugger and its enables are ordered
ahead of the resume. Drop the explicit attach, which is now both redundant and
the racy path.

* fix(settings): localize the browser identity search keywords

* fix(browser): await route policy setup

* fix(browser): satisfy strict static analysis

* test(browser): update live identity fixture API

* test(browser): preserve native UA in live probe

* fix(browser): close the open review findings on the identity revert

- drop a stray JSDoc left over from the removed per-profile setting
- leave user agents without a Chromium engine comment byte-identical
  instead of anchoring the app-token strip on the OS comment and
  destroying a real engine token
- localize the browser identity unavailable error
- correct the worker comment: only shared and service worker requests
  carry no webContentsId, so emulation still reaches dedicated workers
- retire the session user agent policy when a profile is deleted

* test(browser): model a real Electron fallback in the startup UA fixture

The ordering fixture carried no "(KHTML, like Gecko)" engine comment, a
shape app.userAgentFallback cannot actually produce. That unfaithfulness
was what made the old over-stripping look correct, and it broke once the
cleaner started leaving non-Chromium identities alone.

Add the engine comment, keeping the two-word "Orca Development" app token
so the multi-word leak this test exists to catch is still caught. Both
assertions are unchanged.
2026-09-16 10:31:01 -07:00
Jinwoo Hong a01027697c feat(session-search): enable indexing on paired servers from a client (#20886)
* feat(session-search): add ranked history panel search and consent

* test: wait for initial session indexing before refreshing results

* feat(session-history): add local search settings and index controls

* Use shared local host identifier for session index status

* feat(session-search): merge all-computers search across hosts

The `all` scope on `aiVault:searchSessions` now fans out from the desktop
to every host the session list enumerates and merges the pages into one.
Legs run in parallel: the local index through the search service, SSH and
runtime hosts through the existing remote search client.

Two fixed orders, because relevance scores from independent indexes are
not comparable. `newest` asks every leg for recency and k-way merges on
`updatedAt`, nulls last, ties broken on execution host id. `relevance`
rotates hosts in host-id order by their own rank.

The merged cursor is an opaque base64url payload holding each host's
cursor, how many of its current page were already emitted, and the
generation that offset counts into, plus the page size and sort the
cursor belongs to. A host whose index moved is fenced to `stale` and
stops contributing; the rest keep paging. Per-host outcomes ride back on
one new optional `hosts` field on the results response.

`aiVault:searchStatus` with `all` stays refused, and neither the runtime
RPC nor the CLI gains the scope, so a fan-out is never two hops.

* fix(preload): let the search bridge address the all-computers scope

* feat(settings): live index status, enable confirm, advanced delete

* feat(session-search): search every computer from the history panel

The panel's "All computers" scope produced no request: the hook parsed the
scope into a single host id and stopped when that was null, so the panel
answered "Choose one computer to search its sessions." The desktop already
merges every enumerated host behind `aiVault:searchSessions`, so pass the
scope straight through and stamp each hit with the host it came back on.

Hosts the merge could not search are named under the results header with a
short reason, since a silent partial answer reads as "no such session".

(cherry picked from commit c6b9179316)

* feat(session-search): enable indexing on paired servers from a client

Adds `aiVault.setSearchEnabled` so a desktop can turn a paired Orca server's
transcript index on or off and have the server apply it without a restart.

The runtime method refuses any caller without a `pairedDeviceId` with a
`forbidden`-class error, writes the whole resolved policy through the runtime
store so retention rides along untouched, then reaches the index through a
host-supplied hook: `applySessionSearchSettingsChange` on the desktop, the
in-process instance's new `apply` on orcad. The relay is unchanged.

Wire compatibility is Rule 1 shaped: a new optional method. A server that
predates it answers method-not-found, which the desktop IPC handler maps to an
error whose message is exactly `host-too-old`. Old clients never call it. The
method is deliberately absent from the mobile allowlist, and `aiVaultSearch`
stays out of the paired settings projection.

(cherry picked from commit 640c715fbd)

* fix(session-search): report a paired server without session search as host-too-old on status reads

(cherry picked from commit 1463e8a4bc)

* fix(settings): let Button and Collapsible own their spacing and type
2026-09-16 12:50:59 -04:00
Jinwoo Hong 8153ec2306 feat(session-search): search every computer from the history panel (#20885)
* feat(session-search): add ranked history panel search and consent

* test: wait for initial session indexing before refreshing results

* feat(session-history): add local search settings and index controls

* Use shared local host identifier for session index status

* feat(session-search): merge all-computers search across hosts

The `all` scope on `aiVault:searchSessions` now fans out from the desktop
to every host the session list enumerates and merges the pages into one.
Legs run in parallel: the local index through the search service, SSH and
runtime hosts through the existing remote search client.

Two fixed orders, because relevance scores from independent indexes are
not comparable. `newest` asks every leg for recency and k-way merges on
`updatedAt`, nulls last, ties broken on execution host id. `relevance`
rotates hosts in host-id order by their own rank.

The merged cursor is an opaque base64url payload holding each host's
cursor, how many of its current page were already emitted, and the
generation that offset counts into, plus the page size and sort the
cursor belongs to. A host whose index moved is fenced to `stale` and
stops contributing; the rest keep paging. Per-host outcomes ride back on
one new optional `hosts` field on the results response.

`aiVault:searchStatus` with `all` stays refused, and neither the runtime
RPC nor the CLI gains the scope, so a fan-out is never two hops.

* fix(preload): let the search bridge address the all-computers scope

* feat(settings): live index status, enable confirm, advanced delete

* feat(session-search): search every computer from the history panel

The panel's "All computers" scope produced no request: the hook parsed the
scope into a single host id and stopped when that was null, so the panel
answered "Choose one computer to search its sessions." The desktop already
merges every enumerated host behind `aiVault:searchSessions`, so pass the
scope straight through and stamp each hit with the host it came back on.

Hosts the merge could not search are named under the results header with a
short reason, since a silent partial answer reads as "no such session".

(cherry picked from commit c6b9179316)

* fix(settings): let Button and Collapsible own their spacing and type
2026-09-16 12:40:42 -04:00
Jinwoo Hong 46ed53b88a feat(session-search): merge all-computers search across hosts (#20670)
* feat(session-history): add local search settings and index controls

* Use shared local host identifier for session index status

* feat(session-search): merge all-computers search across hosts

The `all` scope on `aiVault:searchSessions` now fans out from the desktop
to every host the session list enumerates and merges the pages into one.
Legs run in parallel: the local index through the search service, SSH and
runtime hosts through the existing remote search client.

Two fixed orders, because relevance scores from independent indexes are
not comparable. `newest` asks every leg for recency and k-way merges on
`updatedAt`, nulls last, ties broken on execution host id. `relevance`
rotates hosts in host-id order by their own rank.

The merged cursor is an opaque base64url payload holding each host's
cursor, how many of its current page were already emitted, and the
generation that offset counts into, plus the page size and sort the
cursor belongs to. A host whose index moved is fenced to `stale` and
stops contributing; the rest keep paging. Per-host outcomes ride back on
one new optional `hosts` field on the results response.

`aiVault:searchStatus` with `all` stays refused, and neither the runtime
RPC nor the CLI gains the scope, so a fan-out is never two hops.

* fix(preload): let the search bridge address the all-computers scope

* feat(settings): live index status, enable confirm, advanced delete

* fix(settings): let Button and Collapsible own their spacing and type
2026-09-16 12:29:11 -04:00
Jinwoo Hong dec0e2cd56 feat(session-history): add local search settings and index controls (#20582)
* feat(session-history): add local search settings and index controls

* Use shared local host identifier for session index status

* feat(settings): live index status, enable confirm, advanced delete

* fix(settings): let Button and Collapsible own their spacing and type
2026-09-16 11:59:22 -04:00
Jinjing 47bb473ec6 Remove agent map from dashboard popout (#20929)
The agent map view was not functional and its components have been removed entirely. The dashboard popout now only supports the kanban board view, with all map-related code, utilities, types, and translations cleaned up accordingly.
2026-09-15 21:55:06 -07:00
Neil f7b2736d6d fix(worktree): block removal when the archive hook fails (#20153)
* fix(worktree): block removal when the archive hook fails

A repo's orca.yaml archive hook is the user's last chance to save work off a
checkout Orca is about to delete. A failed hook was logged as advisory and
stepped over, so the removal went ahead with nothing archived — and the caller
could still be told it succeeded.

The hook is now a blocking precondition, evaluated while the checkout, its Git
registration, its agents and Orca's ownership evidence are all still intact: it
sits ahead of the registration re-read, the lock/dirty preflights, stopPtys()
and removeWorktree in every orchestrator that runs it.

Failure is typed (worktree_archive_hook_failed) and carries the worktree path,
outcome, exit code where one was observed, and the hook's output. unverifiable
stays distinct from exited, so loss of contact is never read as a pass. The
waiver rides its own field at every layer and is never implied by --force, which
already carries the PTY-stop waiver; when used, the waived failure comes back on
result.archiveHookOverride rather than being swallowed.

worktree.archive-failure-blocking.v1 is advertised so an integration can tell
"accepts --run-hooks" from "safely propagates a failing hook" without risking the
data loss to find out. The runtime's SSH path cannot run a hook at all, so rather
than delete with the archive step silently skipped it refuses — waivable like
every other refusal here. #18563 retires that gate by making the path run the
hook for real.

Stacked on #20559, which makes a timed-out hook report honestly; without it a
hook that traps SIGTERM and exits 0 would defeat this gate.

Fixes #19334

* fix(worktree): close the skip-confirm dead end and the client/hook timeout gap

Four review findings on the gate.

A retry from the failure toast could fail for a DIFFERENT reason than the one
the user had just answered, and that second failure got a bare toast with no
buttons. With skipDeleteWorktreeConfirm set, the delete helpers pass no force, so
waiving a failed archive hook on a dirty checkout landed on the dirty preflight
and stopped there. Retry failures now re-enter the same failure toast, so every
retry stays as actionable as the first attempt. Third instance of this class.

The renderer gave worktree.rm a 60s budget while an archive hook may run for
120s. A hook that took 90s and succeeded timed the client out and reported
failure while the host went on to delete — telling the user their delete failed
and their checkout was gone. The budget is now derived from the hook's, and only
when a hook can run.

The SSH fail-open is logged rather than silent, and the capability's doc comment
scopes what it claims: a hook that RUNS and fails cannot delete the checkout; it
is not a promise the hook was found.

The SSH owner-resolution test now reads a real remote orca.yaml through a stubbed
provider and asserts the returned script is the remote one. It previously stopped
at the lookup key, which is the coverage that let this path break twice. It fails
against the row-only resolution.

* fix(worktree): name a signalled hook exit, and state why prunable cleanup skips the gate

Two things the rebase onto #20617 and #20576 surfaced, both found by rerunning
the real-repo harness rather than by reading the diff.

- #20617 added a registration-cleanup branch that returns before the archive
  gate. That ordering is correct — both of its arms describe a row with no
  checkout behind it, so there is nothing to archive and running the hook would
  fail on the missing cwd — but the gate's ordering invariant is documented, so
  the exception should be too.
- A signalled hook reported `Command failed with exit code null.`, which reads
  as a reporting glitch rather than the `unverifiable` verdict it is about to
  produce. It now says the command was terminated without reporting an exit
  code. Introduced by #20576; the withheld `exitCode` itself was always right.

Fixes #19334
2026-09-15 01:19:32 -07:00
Jinjing b8554f1c59 fix(composer): clarify failed attachment drops (#20704)
* refactor(renderer): give the IPC error reader a clamped and an unclamped shape

* fix(composer): name the attachments a drop could not add, in one toast

* fix(composer, source-control): use one stable failure toast slot

- Replace per-worktree toast IDs with single slot that replaces on each failure
- Remove destructive retry actions; discard must confirm in dialog
- Consolidate filesystem import types to shared location
- Add compactIpcErrorMessage for string error handling

* refactor: centralize filesystem import types and clarify failure naming

Move import result types from main/ipc to shared layer so they're available
across preload and renderer. Rename uniformFailure → commonFailure and
skippedOrFailed → failureCount for clarity. Simplify preload/API type
definitions by reusing shared types directly instead of duplicating inlined
union shapes.

* Reuse single toast slot for composer drop failures

Multiple drop failures now replace the previous toast instead of
stacking, preventing notification clutter. Uses a dedicated toast ID
separate from Source Control's stage/discard notifications.
2026-09-14 15:22:05 -07:00
mmarabelandNeil 68f0b2e835 feat(runtime): stream file uploads instead of buffering whole files (#16106)
* feat(runtime): stream file uploads instead of buffering whole files

Staging read each dropped file whole with readFile(), base64-encoded it
(a 4/3 expansion), and passed the string through IPC to the renderer,
which re-chunked it. Peak memory was ~2.3x the file size before a byte
moved, so a 25 MB per-file cap existed to protect the heap.

Staging now records identity only. The byte pump moves into main, where
the file handle and the runtime socket both live: 384 KiB slices (512 KiB
once base64-encoded, matching the chunk size the renderer used) appended
through the existing files.writeBase64Chunk RPC. Peak memory is one slice
regardless of file size, so the ceilings become user-safety limits on an
unattended transfer — 2 GB per file, 8 GB per drop — and over-limit errors
name both the size and the limit.

Because staging and streaming are separate calls, the staged entry carries
size, inode, device and mtime, and the streamer re-checks all four against
the pre-open lstat and against the handle it actually reads. A source
replaced or rewritten at the same size between the two calls is refused
rather than uploaded under the original name. The post-read check compares
mtime as well as size, so an in-place rewrite mid-transfer aborts before
commitUpload renames anything into place.

O_NOFOLLOW, realpath containment and stat identity are preserved, and the
pairing revision plus the runtime id ride every chunk, so a re-pair or a
replacement runtime aborts instead of appending the rest of the file to a
different host.

No wire change: files.writeBase64Chunk and its params are untouched, so
old and new hosts behave identically. The SSH import path is separate and
unchanged. The web client has no local filesystem to stream from and says
so instead of failing obscurely.

* fix(runtime): close the empty-upload and per-drop budget holes

Two gaps the first pass left open.

A zero-byte source returned before the post-transfer identity check, so a
file that gained content during the empty write's round trip committed as
an empty file at the user's chosen name. The empty chunk now falls through
to the same final check the slice loop uses.

Each staged source also started its own byte counter, so the 8 GB ceiling
capped one source rather than the drop: five 2 GB files staged cleanly at
10 GB total. The IPC handler now carries one budget across sourcePaths and
adds only what each source actually staged. The per-file ceiling is still
re-enforced where the bytes move; the drop total holds at staging because
identity enforcement means each file streams exactly the bytes measured.

* docs(runtime): name the invariants the upload helpers carry

* fix(runtime): name the source in errors and stop uploads with their window

Three problems an independent review turned up.

A dropped file's relative path is '', so the over-limit error read "'' is
3 GB, over the 2 GB per-file remote import limit" — the message this change
exists to fix, naming nothing. Errors now fall back to the file's own name;
the staged entry keeps '' so the destination path is unaffected. The
streamer had the same shape, falling back to the hidden .orca-upload-<nonce>
temp destination, a path the user never chose.

The byte loop used to live in the renderer and died with it. Moving it into
main meant closing or reloading the window left the rest of a multi-GB
transfer running, with the renderer's temp cleanup never reaching its
finally. An AbortSignal now rides the caller's lifetime and every chunk, is
re-checked per slice, and main sweeps the abandoned temp path itself when
the renderer is no longer there to do it.

Upload failures also reached the import result wrapped in Electron's
"Error invoking remote method '...'" prefix, because the throw crossed IPC
instead of happening in-renderer; extractIpcErrorMessage unwraps it.

An existing staging test asserted the empty-name message, so it encoded the
bug rather than catching it; it now asserts the file name.

* test(runtime): cover the containment check and the per-chunk host guards

The "escapes the dropped root" test only reached the lstat symlink guard,
so assertEntryInsideRoot had no coverage at all. The shape that actually
needs it is a regular file under a symlinked intermediate directory: lstat
sees a plain file, and realpath containment is the only thing that refuses
it. Disabling the guard now fails this test and nothing else.

Nothing asserted that the SSH target, connection generation and execution
host reach the writeBase64Chunk params either — the renderer tests stop at
the IPC boundary, so the streamer's half of that contract was untested.

* fix(runtime): survive a straggling append when sweeping an aborted upload

Aborting rejects the in-flight chunk locally, but the host may still apply
that append, and appends open with flag 'a' — which recreates the file the
sweep just deleted. The delete and the straggler also race: they are
separate calls on a queue that is not ordered between them.

Slices are strictly sequential, so at most one append can be outstanding.
A second pass after it has had time to land is therefore sufficient, not
merely a heuristic. The sweep moves out of filesystem-mutations.ts into its
own module so the behaviour is testable directly.

Found by an independent review pass, which also pointed out that the
"escapes the dropped root" test only reached the lstat symlink guard.

* fix(runtime): abort uploads only when the document commits, and honour manual disconnect per chunk

did-start-navigation fires before will-navigate blocks an external link or a
stray file drop, and the renderer survives those (verified against Electron 43
with a hidden window). Aborting there killed a healthy upload with a misleading
'window went away' error. did-navigate fires only once a new document has
replaced the caller.

The renderer's per-chunk calls used to go through the IPC handler that refuses
a manually disconnected environment; the loop in main made no such check, so a
disconnect mid-upload kept pushing the rest of the file. The handler now
resolves the selector to an environment id and the streamer checks it per slice.

Adds slice-boundary coverage against the real chunk schema and host write
flags, staging-to-stream on a real filesystem, and handler-level lifetime tests.

---------

Co-authored-by: Neil <neil@stably.ai>
2026-09-14 14:08:31 -07:00
dngur6344andNeil 9cf0a6c37f perf(remote): avoid repeated capability probes during file imports (#14555)
* perf: avoid repeated remote import capability probes

* test: cover cold remote import compatibility probe

* fix(remote): fence imports across runtime reconnects

* fix(remote): bind import proof to connection

* fix(remote): fence import routing by runtime identity

* test(remote): remove unsafe import fixture assertions

- type remote RPC mocks at declaration so call arguments stay checked
- narrow upload params before reusing generated temp paths

---------

Co-authored-by: Neil <neil@stably.ai>
2026-09-13 20:51:15 -07:00
Brennan BensonandMerge Sim 5e70014da8 feat(native-chat): support file drag and drop (#20494)
* feat(native-chat): support workspace file drops

* fix(native-chat): report OS file drops that attach nothing

#15782 is a silent failure on the Finder route, and that route still
swallowed every way it could fail:

- the preload handler returned with no feedback when the OS handed us
  file items `webUtils.getPathForFile` could read no path from (promised
  or virtual files). It now sends the existing `rejected` payload with a
  new `unresolved-paths` reason, which the global drop toast names.
- the composer's external-attach path dropped the batch with no notice
  when every path failed authorization, when an upload came back empty,
  and (new in this branch) when the owner changed mid-flight. Each exit
  now sets a notice; only a disabled composer stays quiet, because it has
  no notice surface.

Also stops `resolveNativeChatAttachmentOwnerForWorktree` throwing out of a
drop/IME handler when an SSH connection's generation is gone mid-attach —
that is an unknown owner, which the resolver already models as
`not-ready`.

* refactor(native-chat): one owner-identity check for composer attachments

The branch had two near-identical "is this still the same owner" helpers,
one per attach route, and they disagreed: the workspace-drop copy ignored
the SSH connection generation, so a reconnect between the drop and the IME
flush read as the same owner and the path landed on a new connection.

Collapses both onto one predicate in the pure ownership module (the
store/toast-free seam both routes already depend on), which compares the
full SSH expectation and never treats `not-ready` as a match.

* perf(file-explorer): resolve drag ownership at dragstart, not per render

The virtualized row list resolved the selection's source execution host on
every render — the virtualizer re-renders on every scroll frame, so a large
multi-selection paid a full projection scan plus a route allocation per
selected path per frame, and per visible row on top of that. Only
`onDragStart` ever read the result.

Rows now receive a resolver they call with the paths they are about to
drag. The three copies of the "stamp only if both halves resolve" guard
(explorer row, both combined-diff row shapes) collapse into one helper next
to the writer.

* fix(native-chat): refuse a guarded composer drop visibly

The drop handlers claimed the drag (preventDefault + stopPropagation) before
checking `disabled`, so a guarded composer told the browser it accepted the
drop, left the copy cursor up, and then did nothing — the same silent swallow
this branch exists to remove.

Dragover now answers `none` when the composer is guarded, so the cursor refuses
and no drop event follows. It still claims the event either way: the composer
sits inside the terminal surface, which accepts the same drag and would paste
the paths into the shell instead.

Drops `stopImmediatePropagation`. The capture-phase `stopPropagation` already
keeps the event off the editor below, so the stronger form only risked
suppressing unrelated listeners on the React root.

The fake DataTransfer in the test now starts at a dropEffect we never write, so
asserting `none` or `copy` proves the handler set it.

* fix(native-chat): decide attachment ownership per path, not per batch

A queued batch can mix sources — a workspace drop the target host owns and a
client-local paste it cannot read — because IME composition holds both until it
settles. Collapsing the batch to one verdict refused the whole thing on a remote
target, including the drop the user was entitled to make.

The verdict now follows the path it belongs to: owned paths attach, client-local
ones are refused, and the refusal is reported rather than dropped. A stale owner
still refuses everything, since that means the target moved under all of them.
Also guards the empty-batch case, which previously read as "every path owned".

* refactor(combined-diff): resolve drag ownership from the live workspace

The combined diff captured an execution host into the open-file record at tab
open and drilled it through three components to reach the row. That host was
never persisted, so after a restart every drag from a restored diff was refused
until the tab was reopened, and the capture failure was swallowed into an
undefined source with no trace.

Rows now resolve the owner the same way the source-control rows already do, from
the workspace the diff belongs to at the moment of the drag. That deletes the
prop drilling, the store capture and its bare catch, and leaves one way to
answer "who owns these paths" for every live listing.

The file explorer keeps its per-node owner: its tree is a cache that can still be
showing a previous host's listing, which is exactly what that field records.

* revert(file-explorer): drop the workspace-id tree reset

Resetting and reloading the tree when the workspace id changes at an unchanged
path is not needed for the drag source to be correct. The tree already records
the workspace whose root listing it committed, so a cache left over from a
previous workspace stamps that workspace and the composer refuses the drop —
the intended answer, reached without touching the reset rule.

That rule clears selection, the name filter and undo history, which is more
file-explorer behaviour change than this feature asked for.

* test(native-chat): stop the external-attach mock hiding new notices

The hook's test replaced the whole attachment-owner module with a hand-written
stub, so the two notices added alongside the owner-change guards resolved to
undefined. Calling them threw inside the async attach loop — an unhandled
rejection, which leaves every test in the file reported as passing while the run
as a whole fails. CI caught it; a local run reporting only pass/fail counts does
not.

The mock now spreads the real module, so a notice added later cannot go missing
from it, and both owner-change tests assert the string a user would read instead
of only asserting that nothing attached.

* test(native-chat): guard the last-path owner change on a one-file drop

The owner flipping while the final path is authorizing has no next loop
iteration to catch it, so the post-loop check is all that stands between a
single-file drop and a path attached to a host that no longer owns it — and a
one-file drop is the ordinary shape. No test covered that exit.

Removing the post-loop check now turns this red; before it, only the
multi-path exit was guarded.

* fix(native-chat): keep a mixed attachment batch in attach order

applyResolvedPaths partitioned a queued batch into a target-owned half and
a client-local half and concatenated them. An IME-delayed batch that mixed
a workspace drop with a paste made earlier in the same composition was
therefore inserted owned-first, so the dropped reference jumped ahead of
the pasted one in the draft.

Filter against the two verdicts in place instead. Membership is unchanged,
the order the user attached in survives, and the two intermediate arrays go
away.

* fix(file-explorer): name the owner of a dragged path whose row is hidden

A multi-selection outlives the rows that showed it. Nothing prunes
selectedPaths when a directory collapses, when the name filter narrows, or
when dotfiles are hidden, and the drag still carries every selected path.
Drag-source resolution read those owners from the row projection, which is
built from visible rows only, so one hidden path collapsed the whole drag to
an unstamped one and the composer refused it as coming from another
workspace.

The owner was never unknowable — the dir cache the projection is built from
still records which host listed that path. Fall back to it when the path has
no visible row. A path in neither (a name-filter synthetic node for a
directory that was never listed) still fails closed.

* fix(native-chat): ask which workspace the composer serves now

The IME-flush ownership check compared the workspace id captured when the
drop happened against the same captured value, so for a structured pane the
comparison could only ever hold. The live protection came from the host and
owner checks beside it; this one asked nothing.

Read the id through a ref so the check means what it reads as. A pane whose
structured target moves between the drop and the composition settling now
refuses the queued path instead of attaching it.

* fix(native-chat): ask which workspace an external attach lands on

The post-await ownership gate resolved the owner through the render closure, so
it re-asked the workspace the attach started in and compared the answer with
itself. A tab moved to another workspace mid-authorization passed the gate, and
the paths landed in a composer that no longer served that workspace.

Read the pane through a ref and compare the workspace identity as well as the
owner: two workspaces can both report a local owner, so the owner alone cannot
tell them apart.

* test(native-chat): read the real notice on a workspace drop

The drop tests hand-built their attachment-upload mock and hand-copied the
not-ready wording into it, so the assertion tracked the copy rather than the
string a user reads: rewording the real notice left all 15 tests green.

Spread the real module and override only the owner resolver, matching the two
sibling test files in this directory. Rewording the notice now fails the test.

* docs(native-chat): restore the hook's doc comment to the hook

The workspace comparison landed between the doc block and the function it
describes, leaving the comment attached to a type alias.

* test(native-chat): cover the upload window for a moved pane

The workspace-currency gate guards two windows and only the authorize loop was
covered. The upload window is the longer one: the paths go to the worktree the
attach captured, so a pane that moved workspaces meanwhile must not receive
remote paths living under the workspace it left.

* test(native-chat): pin the two untested attachment refusals

Refusing an already-blocked target at the drop rather than queueing it had no
test: queued paths that can never attach still spend the pending budget, and the
next legitimate drop is then turned away for being one too many.

Also pins the immediate already-false ownership verdict. Today's only caller
settles ownership synchronously so it cannot arrive false, but the hook exports
this entry point and the fallback is not a refusal — a false verdict is not
"owned", so a remote target blames client-local attachments for an ownership
failure. Verified: removing the branch reports the wrong notice.

* docs(native-chat): say which rule the ownership refusal follows

The per-path comment sat directly above the batch-wide ownership refusal while
describing the blocked-target logic below it, so the refusal read as a
contradiction of the line under it rather than as the file's stated rule.

Name the rule at the refusal: a failed ownership verdict refuses the whole
completion, the same way the pending-limit rejection does.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-13 16:41:51 -07:00
241fb9ed9d perf(terminal): batch file-link checks on their owning host (#20463)
* perf(terminal): batch file-link existence checks on their owning host

* test(relay): allow additive filesystem capabilities

* fix(web): keep terminal file links working under batched existence checks

createShellApi omitted pathsExist, so withFallback answered the new batch
call with a truthy proxy resolving to undefined and the whole hover batch
rejected — dropping every link on lines with an out-of-worktree path.

* test(web): assert the shim without type assertions

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <neil@stably.ai>
2026-09-13 16:27:06 -07:00
Jinwoo Hong 1f7655f3e3 feat(ai-vault-search): public session search contract and transports (#20277)
* feat(ai-vault-search): define public contract and service seam

* feat(ai-vault-search): add IPC runtime relay and web transports

* fix(ai-vault-search): register search IPC at the core handler site

ai-vault.ts was two lines over the 300-line max-lines limit; the search
handlers belong with the other register*Handlers calls anyway.

* fix(ai-vault-search): withhold degraded-root paths from relay status

Status carried local filesystem paths over the relay while hits redact
theirs. redactStatusForTransport applies the same policy at the same
boundary: relay callers keep each root's reason and the array length as
the count, so the type only makes root optional.

* fix(ai-vault-search): close diagnostic path leak and remove test casts

* feat(ai-vault-search): carry an execution host id and per-host outcomes on hits

* feat(ai-vault-search): route desktop search by execution host scope, including runtimes

* feat(preload): accept an execution host scope on session search

* feat(web): answer only for the paired runtime on session search

* docs(ai-vault-search): describe execution-host routing and the all-hosts merge

* test(ai-vault-search): cover every host scope, the all-hosts merge and wire compat

* fix(ai-vault-search): resume every host mid-page so a merged page never drops a hit

* fix(ai-vault-search): decode the merged cursor with a schema instead of casts

CI's type-aware audit refuses type assertions; a zod record validates the
per-host entries and yields the typed map without one.

* refactor(ai-vault-search): defer cross-host merged search
2026-09-13 17:53:50 -04:00
Neil 53eb639983 refactor(preload): drop the unused raw electron IPC bridge (#20419)
* refactor(preload): drop the unused raw electron IPC bridge

`@electron-toolkit/preload` was used only to expose `window.electron`,
which hands the renderer unrestricted `ipcRenderer` send/invoke/on for any
channel — bypassing the typed per-domain bridges in `src/preload/api/`.

Nothing consumed it. The only references were the assignment itself, the
web client's empty fallback, and a test asserting that fallback has no
keys — i.e. the web build already ran with it empty.

* chore(build): drop the dangling @electron-toolkit/preload vite exclude

The package is gone from package.json and source; leaving it in the
preload externalizeDeps exclude list points at a package that no longer
resolves.
2026-09-12 16:00:19 -07:00
Jinwoo HongandOmar Shahine 3b82d8de64 fix(runtime): let connections own host status recovery (#20003)
* fix(runtime): let connections own host status recovery

Verify runtime status after authenticated connection recovery and publish
ordered snapshots to desktop and browser viewers. Consolidate failed-status
retries in the connection owner and remove renderer retry/diagnostics merging.

Adapt sidebar host-state derivation and regression coverage from Omar
Shahine's original fix in https://github.com/stablyai/orca/pull/19163.

Co-authored-by: Omar Shahine <10343873+omarshahine@users.noreply.github.com>

* fix(runtime): show blocked hosts honestly and remove obsolete status options

* fix(runtime): preserve timeout guidance and update IPC test fixtures

* fix(runtime): preserve status evidence and address review gaps

* test(sidebar): assert workspace host icons dimming and recovery tooltips

* fix(palette): require available hosts before adding implicit badges

* fix: retain disconnected host snapshots for new renderers

---------

Co-authored-by: Omar Shahine <10343873+omarshahine@users.noreply.github.com>
2026-09-11 03:19:01 -04:00
Jinwoo Hong 74cc9b5039 feat(desktop): native mobile push integration (2/3) (#19935)
* feat(desktop): integrate native mobile push delivery and lifecycle

* fix(desktop): preserve notification replay policy and review invariants

* fix(desktop): correct notification locale namespace and auto-ack tests
2026-09-10 20:16:56 -04:00
f242c99af4 perf: scope activation inventory to the owning host and workspace (#19447)
* perf: scope activation inventory to the owning host and workspace

* fix(activation): keep an unscoped census fallback when the owning host is unnameable

Scoping the activation inventory made resolveActivationPtyListScope throw for
paired-runtime workspaces and made a detached relay reject the scoped list, and
both collapse to a 'blocked' gate. 'blocked' skips the sleeping-agent resume and
the caller's reseed, so an SSH target on the bounded offline floor lost its
initial pane and peer workspaces stopped resuming.

Fall back to the unscoped inventory that shipped in exactly those two cases; the
scoped fast path still covers local, folder and attached-SSH workspaces. Also OR
the host-reported worktreeId with the id-prefix match instead of preferring it,
because a relay seeds worktreeId from the host's own ORCA_WORKTREE_ID and a
session dropped from the census is one the gate forks a second writer onto.

* test(activation): update forkbomb fakes to the scoped session.tabs.list shape

The gate now asks the host for one workspace's snapshot instead of the whole session.tabs.listAll inventory and refuses an answer that does not name its scope, so the old snapshots-array fakes made it block instead of resume.

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-08 19:42:25 -07:00
Jinwoo Hong 12f53da542 Remove settled-worker automatic resume and hibernation fences (#19544)
* Remove settled-worker automatic resume and hibernation fences

* test: retirement rollback case follows the no-fence policy

Case 4 seeded and asserted automaticResumeBlockedBy, which this branch
deletes. A rolled-back settled worker is now an ordinary done record that
wake clears as passive evidence, same as any finished agent pane.

* chore(i18n): regenerate the runtime-required catalog for the contrast floor strings

* test(orchestration): give the stopping-worker guard fixtures a Run
2026-09-08 05:14:59 -04:00
Brennan BensonandMerge Sim 6a47d2831f fix(native-chat): scope composer file drops to the pane that received them (#19328)
* fix(native-chat): scope composer file drops to the pane that received them

A native OS file drop resolving to `target: 'composer'` carried no pane
identity, so the window-wide payload was attached by every mounted composer.
Because inactive chat tabs stay mounted (hidden), one drop populated every
chat pane's attachment cache, and those chips replayed whenever the user
returned to a tab they never dropped into. The workspace-creation composer
and chat composers also leaked into each other, since neither could tell
which surface actually received the drop.

Composer drops now carry a `scopeKey` the way a terminal drop carries its
tab and pane leaf id: the composer publishes its pane key as
`data-composer-scope-key`, the preload harvests it during the composedPath
walk, and each composer attaches only its own. The workspace composer's
last-wins ownership stack now claims unscoped payloads only.

* test(native-chat): supersede the bug-asserting drop repro with the scoping test

The repro that landed on main asserts the pre-fix behavior (a drop reaching
every mounted composer), so it fails once drops are scoped to the pane that
received them. Its scoping cases now live in
native-chat-composer-drop-scope.test.tsx, which keeps its editor-target
control case verbatim and adds coverage for unscoped composers and a scope
key published inside the drop-target marker.

* test(native-chat): cover workspace composer drop isolation

* fix(native-chat): authorize external attachment paths before preview

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-07 13:34:23 -07:00
Jinwoo Hong 5857357fcf feat(relay): log the region probe and name the assigned cell (#19307)
* feat(relay): log the region probe and name the assigned cell

A desktop silently pinned itself to a far relay region for a day and every
phone connect paid the round trip. Nothing in the desktop logs said which
regions were probed, what they measured, why one was rejected, or which cell
the host landed on, so the only way to diagnose it was a bench harness.

The resolver now emits one line per outcome. A refresh carries every region's
probe origins, the discarded warm-up, the kept samples, the minimum, the
spread, and a verdict, then the chosen region or no-hint with the reason it
withheld one. Cache hits, diagnostic overrides, and a director that cannot
list its regions each get their own line so a quiet run is never ambiguous.
Self-heal logs the cached region, the best measured region, the assigned
cell's round trip, and whether it kept or deleted the cache. Only a refresh
reports a catalog failure; a self-heal never chose a region, so a line saying
it withheld a hint would be a lie.

Relay status now carries the assigned cell so the pairing panel can name it.
The field is optional because an offline host holds no assignment and the web
client answers from a stub that never has one.

Splitting catalog fetching out of the preference module keeps both files
inside the line budget without a lint disable.

* fix(relay): drop the assigned cell from statuses not served on it

The origin pool publishes offline while it still holds the assignment it is
about to rotate, so the panel kept naming a cell nothing was served from. The
same class of bug hid a second instance: the coordinator republishes
registered right after the broker announces its cell, and that republish
carried no cell, blanking the value moments after it was set. The cell would
never have reached the panel in the real flow.

Deriving the cell from the status at each publisher removes both. The rule
lives beside the status type because it defines when the optional field is
populated, and the coordinator reads the owned broker's endpoint rather than
trusting a call site to remember to pass it.

* i18n: add the relay cell label to the English catalog

* test(relay): audit the relocated region catalog fetch call site

* fix(relay): report a self-heal whose catalog request failed instead of staying silent
2026-09-07 13:45:48 -04:00