* perf: coalesce git upstream status reads
* fix(git): repair upstream lease key imports and guard its field list
The read owner imported the shared/types barrel deleted by #14447, and its
push-target key hand-enumerated fields, so a new GitPushTarget field would
silently share a lease between two different targets. The destructure now
fails to compile if a field is added. Lease tests moved into their own file
after #14728 split ssh-git-provider.test.ts.
* test(git): enforce native upstream coalescing in CI
The 10-caller benchmark only runs under ORCA_GIT_UPSTREAM_COALESCING_BENCH_JSON,
so nothing in CI failed when the native/WSL lease was bypassed. Route status.ts
through invalidateGitUpstreamStatusReads so the export has a production caller.
* ci: color added/deleted LoC counts in PR summary
* ci: use GitHub color-swatch dots for added/deleted LoC counts
* ci: color LoC counts with LaTeX textsf
* ci: bold LoC counts; render zero in white
* ci: render LoC counts large bold sans-serif
* ci: use bold math font for LoC counts
* ci: color only the + and - signs on LoC counts
* refactor: split editor.ts under 400 lines
Move editor slice types, file-id/tab helpers, and action factories under
src/renderer/src/store/slices/editor/. The source file is now a public
barrel. No intentional behavior change.
* refactor: split editor-chrome-slice into state modules
- Move EditorDraftState, ExplorerDirState, and RightSidebarState type definitions into their respective action files
- Simplify state creator return types from Pick<EditorSlice, ...> to specific state types
- Improve modularity by colocating types with implementations
* refactor(persistence): extract modules to half persistence.ts
* refactor(persistence): tighten the extracted operations seam
Review follow-ups on the module extraction, all behavior-neutral.
The extracted operations read and mutate the Store's state object in place, but
every seam typed it as a bare PersistedState, so nothing at the boundary said a
caller must pass the live reference — a future caller handing over a clone would
have its writes silently dropped. Name that contract: StoreOwnedPersistedState
carries it to every operations interface and every mutating free function.
normalizePersistedPaneIdentityState and backfillFolderScopeConnectionIds stay on
PersistedState; they build a fresh state rather than mutating the Store's.
The six *PersistenceOperations wrappers were constructed per delegate call. They
are stateless today, so this was inert, but any future instance state would be
lost between calls. Memoize them, and mark state and gitUsernameCache readonly
so the compiler enforces the single-assignment invariant memoizing them relies
on.
Also: restore flushSshPtyConsumerRecovery, whose inlining left its rationale
duplicated at both call sites; document that migrateWorktreeIdentity's boolean
gates the caller's save, since the extracted function kept no docs of its own;
and merge a duplicate shared/types import that was failing lint under
--deny-warnings.
* delete plan doc
* refactor(persistence): add error recovery and improve field cleanup
- Rollback failed migrations to prevent corrupted state that blocks retry
- Gracefully skip malformed entries in normalization instead of aborting
- Strip retired fields to prevent orphaned state and sync issues
* refactor(persistence): drop the redundant persistence- filename prefix
The extracted modules already live in src/main/persistence/, so name
them after the domain they own. Point leftover shared/types imports
at the real type modules while touching those files.
* refactor(persistence): optimize lookups and fix unsanitized updates
- Use Maps instead of repeated array searches for O(1) lookups
- Apply sanitized updates instead of raw input in ui-state-update
- Compare fields directly rather than JSON strings to avoid false dirty states from persisted key ordering differences
* refactor(persistence): group modules into lifecycle folders
Move the 42 flat persistence modules into six folders named for what the
module does, and lift the Store class out of the barrel so persistence.ts
becomes an 8-line public surface.
Bodies are unchanged: every moved file diffs clean against HEAD once import
blocks are excluded. Only import specifiers were rewritten, by resolving each
one to an absolute path and mapping it through the move map.
Store keeps its existing max-lines suppression; its baseline entry is repathed
rather than re-added. Its 119-method public API sets a ~525-line floor, so it
cannot meet the 400-line cap without breaking the API for 153 importers.
* Sanitize worktree visibility sources and preferences on hydration
Ensure invalid or corrupted data from disk (untracked whitespace,
relative paths, bogus preference values) is cleaned during load
rather than corrupting the in-memory store.
`source-control-dropdown-items.ts` was 524 counted lines behind an
`eslint-disable max-lines`. It splits along the seams the resolver already had:
- `source-control-dropdown-item-types` — the row union, consumed by CommitArea,
the composer and the action dispatcher without pulling in the state machine.
- `source-control-dropdown-labels` — count/label/title wording.
- `source-control-dropdown-action-context` — the branch, upstream and review
facts every row reads, derived once so rows cannot disagree about them.
- `source-control-dropdown-commit-items` / `-remote-items` / `-review-items` —
the three row groups, each keeping its own disabled-reason ladder intact.
`resolveDropdownItems` is now just the entry order plus the conflict-abort and
hosted-review-busy passes.
Verified output-identical to the pre-split resolver: a differential harness ran
both implementations over 16,380 generated input combinations (every upstream
shape × PR state × conflict operation × blocked reason × staged count ×
provider) and compared entries deeply. The harness was scaffolding and is not
committed.
Splits the nine oversized modules in the daemon/provider/runtime domain into
focused per-concern files and drops their max-lines baseline entries.
- daemon: `Session` decomposes into an output plane (emulator, pending-output
buffer, client fan-out), a producer-pause controller, a shell-ready barrier and
a termination controller; `DaemonClient` into socket connect, hello handshake,
ndjson readers, pending-request settlement, listener registry and notify
settlement; `daemon-health` into pid-file parsing, process identity,
stale-kill, TCC attribution and bundle staleness; `shell-ready` into the marker
constant and the bash/zsh rcfile generators.
- providers: local-pty shell-ready wrapper generation, wrapper root, startup
command and bash rcfile split out of local-pty-shell-ready.
- runtime: `Coordinator` sheds DAG convergence, decision gates, escalation
triage, the runtime contract, the stale-base flag and task dispatch; the
files/git/github rpc modules split into per-domain method groups.
Behavior-preserving: the extracted units keep their original construction order,
guards and timer lifetimes, and every RPC method name is still registered.
Test `vi.mock` surfaces were re-partitioned to follow the moved symbols.
The two tab-group hooks, the pane manager, worktree activation, and the terminal
pane context menu each carried a file-level `eslint-disable max-lines` and ran
461-745 counted lines against a 300-line budget. AGENTS.md calls for splitting
rather than suppressing, and config/max-lines-baseline.txt is a shrink-only
ratchet, so this removes all five suppressions and prunes their entries
(341 -> 335).
Pure move, no behavior change. useTabDragSplit is cut into gesture lifecycle,
hover preview and drop commit; useTabGroupWorkspaceModel into item projections
plus the tab-close, close-scope, activation and creation command sets; the pane
manager into host, tree mutations, pane creation, drag wiring, reparent frame
tracking, layout sweeps and rendering diagnostics.
react-hooks exhaustive-deps stays at zero warnings, matching HEAD. Dependency
additions are only stable identifiers -- refs and callbacks that became
parameters -- and no `.current` dereference was added to any dependency array.
Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(the three remaining failures are pre-existing load flakes in untouched files,
each green when re-run serially), no new runtime import cycles among 1020
modules, no barrel files, and no lint suppression added anywhere.
The four editor modules, the diff-comment decorator and the file-type icon table
each carried a file-level `eslint-disable max-lines` and ran 319-806 counted
lines against 300/400-line budgets. AGENTS.md calls for splitting rather than
suppressing, and config/max-lines-baseline.txt is a shrink-only ratchet, so this
removes all six suppressions and prunes their entries (341 -> 334).
Pure move, no behavior change. MonacoEditor is cut along its own seams -- mount,
input bindings, markdown annotations, decorations, content sync, view-state
persistence and reveal scheduling -- with the markdown overlay becoming its own
component. useEditorPanelContentState splits into file and diff content loaders
plus the active-tab load and reload triggers.
When the mount module came in at 370 counted lines, over the 300 ceiling, it was
split again into its parameter types and its input bindings rather than carrying
a suppression.
Hook usage is identical to HEAD across all three React split families: the same
counts of every hook type between each original and its extracted modules, so no
hook was added, dropped, or converted to a plain function. react-hooks
exhaustive-deps stays at zero warnings, matching HEAD.
Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(the four remaining failures are pre-existing load flakes in untouched files,
each green when re-run serially), no new runtime import cycles among 884
modules, and no lint suppression added anywhere.
The six right-sidebar modules and the remote file browser each carried a
file-level `eslint-disable max-lines` and ran 347-797 counted lines against
300/400-line budgets. AGENTS.md calls for splitting rather than suppressing, and
config/max-lines-baseline.txt is a shrink-only ratchet, so this removes all seven
suppressions and prunes their entries (341 -> 333).
Pure move, no behavior change.
Two renderer-specific hazards were found and fixed rather than shipped.
First, effect and ref LIFETIME. FileExplorer's `if (!worktreePath) return` sits
above the files pane, so moving the worktree-reset effect into that pane made its
guard ref `lastResetWorktreePathRef` die on any render where worktreePath was
transiently null (workspace-list refresh, store rehydrate, remote worktree
reload). On remount the guard read null, so the reset fired even when returning
to the SAME worktree -- wiping dirCache, collapsing every expanded directory,
clearing the name filter and undo history, and forcing a full re-read over SSH.
The tree-load effects now live in a hook called from FileExplorer above the early
return, and the pane is purely presentational with zero hooks. That also restores
the original parent-effect ordering, which had shifted because React flushes
child effects before parent effects.
Second, extracting a hook silently degrades dependency analysis: `setX` setters
that the linter knew were stable when created locally become opaque parameters,
producing 8 new react-hooks/exhaustive-deps warnings where src/renderer had zero.
Those are fixed by listing the genuinely stable identifiers (useState setters and
ref OBJECTS). No `.current` dereference was added to any dependency array, since
that would change callback identity as the ref mutates.
Verified: oxlint clean with exhaustive-deps back to zero, ratchet passes,
typecheck clean, full unit suite green on the first pass, no new runtime import
cycles, no lint suppression added, and hook usage identical to HEAD across both
split families.
The three usage scanners and their stores, plus the renderer usage-overview
model, each carried a file-level `eslint-disable max-lines` and had grown to
338-769 counted lines against a 300-line budget. AGENTS.md calls for splitting
rather than suppressing, and config/max-lines-baseline.txt is a shrink-only
ratchet, so this removes all seven suppressions and prunes their entries
(341 -> 334).
Each file is cut along the seams it already had -- and that several of the
suppression comments named out loud: filesystem discovery / record parsing /
attribution / aggregation for the scanners, and pricing policy / scope filters /
rollups / session rows / automation attribution for the stores.
Pure move, no behavior change. Code is relocated verbatim; the only edits are
import plumbing and, where a private class method became a free function, the
mechanical `this.state` -> `state` parameter threading. Every converted call
site passes `this.state` at call time and the automation path takes a live
`getState: () => this.state` getter, so no state is snapshotted. No barrel
exports: each new module owns real logic and importers point at the owner.
Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(remaining failures are pre-existing load flakes in untouched files, each green
when re-run serially), no import cycles among the 64 affected modules, and a
statement-level diff of every split confirms the moves are verbatim.
The four agent hook services, the main hooks module, and the two relay modules
each carried a file-level `eslint-disable max-lines` and ran 365-628 counted
lines against a 300-line budget. AGENTS.md calls for splitting rather than
suppressing, and config/max-lines-baseline.txt is a shrink-only ratchet, so this
removes all seven suppressions and prunes their entries (341 -> 334).
Pure move, no behavior change. Each hook service splits into its managed script
source, its config/bundle serialization, and its remote-install path, keeping the
per-agent integrations independent: copilot, amp, antigravity and hermes each
retain their own getManagedScript rather than sharing one, because each emits a
different script body for a different agent. Merging them by name would have
been a behavior change, not a refactor.
For antigravity the suppression's stated rationale -- that local install, Windows
wrapper generation, status cleanup, and SSH remote install must share one event
list and managed-command matcher so stale-hook cleanup cannot drift by platform
-- is now enforced structurally instead: both install paths call
buildInstalledConfig + createAntigravityManagedCommandMatcher over the single
ANTIGRAVITY_EVENTS catalog, with the graph a strict DAG.
Also registers the six new antigravity/ and copilot/ modules in
config/tsconfig.cli.json. That project uses a curated `include` list rather than
a glob, so an unlisted module fails `tsc -p config/tsconfig.tc.cli.json` with
TS6307 even though the entire unit suite passes.
Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(remaining failures are pre-existing load flakes in untouched files, green when
re-run serially), no new runtime import cycles, and no lint suppression added.
The GitLab, GitHub, Jira and Linear integration modules, their two IPC
registrars, and the shared GitHub project types each carried a file-level
`eslint-disable max-lines` and ran 351-614 counted lines against a 300-line
budget. AGENTS.md calls for splitting rather than suppressing, and
config/max-lines-baseline.txt is a shrink-only ratchet, so this removes all
eight suppressions and prunes their entries (341 -> 333).
Pure move, no behavior change. Each client is cut along the seam it already
had: per-operation modules for the issue APIs (create / update / comment /
field options), and for Jira the request queue, site credential store,
authenticated request, and site identity. The two IPC registrars keep their own
handlers and delegate the rest to per-domain sub-registrars, so they remain
real entry points rather than re-export shims.
The IPC surface is proved intact rather than assumed: comparing (method,
channel) multisets between HEAD and the split gives 52 registrations across 52
distinct channels on both sides.
Provider-neutrality is preserved -- GitLab and GitHub keep separate, parallel
module layouts rather than being merged behind a shared abstraction.
Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(the one remaining failure is a pre-existing load flake in an untouched file,
green when re-run serially), no new runtime import cycles among 744 modules,
and no lint suppression added anywhere.
The six oversized src/main/ipc modules each carried a file-level
`eslint-disable max-lines` and ran 427-671 counted lines against a 300-line
budget. AGENTS.md calls for splitting rather than suppressing, and
config/max-lines-baseline.txt is a shrink-only ratchet, so this removes all six
suppressions and prunes their entries (341 -> 335).
Pure move, no behavior change. Each file is cut along the seams it already had:
pet splits into format allowlist / storage paths / symlink-safe copy / bundle
manifest + import; filesystem-auth into path-containment primitives, the
config-derived allow-list, and the git-registered root cache; notifications into
sound selection, native lifecycle, permission probe, and burst cooldown;
crash-reporting into renderer error reports, breadcrumbs, and sender.
The IPC surface is proved intact rather than assumed: comparing (method,
channel) multisets between HEAD and the split gives 49 registrations across 49
distinct channels on both sides. filesystem-auth's security boundary keeps its
acyclic layering -- containment primitives, then allow-list, then root cache,
then path-resolution orchestration -- with no layer gaining a back-edge.
Also keeps clipboard-ipc-handlers.test.ts under the 800-line test budget. The
split had briefly added a redundant vi.mock for isENOENT (byte-identical to the
real implementation) that pushed it to 801; the mock is dropped in favor of the
real function, with realpath added to the existing node:fs/promises mock.
Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(the one remaining failure is a pre-existing load flake in an untouched file,
green when re-run serially), no new runtime import cycles among 617 modules,
and no lint suppression added anywhere.
The five oversized src/main/browser modules and src/main/ipc/browser.ts each
carried a file-level `eslint-disable max-lines` and ran 377-654 counted lines
against a 300-line budget. AGENTS.md calls for splitting rather than
suppressing, and config/max-lines-baseline.txt is a shrink-only ratchet, so
this removes all six suppressions and prunes their entries (341 -> 335).
Pure move, no behavior change. cdp-ws-proxy is decomposed into collaborating
objects rather than free functions because its state is genuinely
per-connection: every collaborator is a private readonly instance field built
in the constructor with live closures over `this`, so per-connection state
stays per-connection. Likewise the screencast pacer's isClosed/isStopping and
snapshot capture's getSeq are live thunks, not values captured at wiring time,
so guards inside already-armed timers still observe a later stop().
browser-guest-ui.ts is renamed to browser-guest-shortcut-forwarding.ts: after
the split it exports exactly one function, setupGuestShortcutForwarding, so the
old name no longer described its contents.
Also restores a single `webContents.debugger` read in the screencast path. The
extraction had left three reads where the original had one; the accessor is
stable today, so this is not a behavior fix but it removes a latent divergence.
Verified: oxlint clean, ratchet passes, typecheck clean, full unit suite green
(remaining failures are pre-existing load flakes in untouched files, each green
when re-run serially), no new runtime import cycles, and the IPC channel set
diffed identical before/after with all 23 handlers still trust-gated.
* docs(source-control): plan half-size extraction
* refactor(source-control): extract modules to half SourceControl.tsx
* move git decoration token comment to correct component
* delete plan doc
* test: add useSourceControlBranchCompare and git-history hook tests
Comprehensive unit tests covering the scheduling, stale response filtering,
and visibility logic of the extracted branch-compare and git-history hooks.
* refactor(source-control): internationalize UI strings
Add translate() support for all hardcoded strings throughout source control UI,
extract reusable SourceControlTreeDirectoryHeader component, improve error
handling in bulk operations with logging and user-facing toasts, and add proper
return type annotations to hooks.
* fix(source-control): satisfy react-doctor rules in extracted modules
Reset worktree-scoped hook state during render instead of in an effect,
and give dropdown separators stable ids so the changed-code quality gate
stops flagging the extracted SourceControl modules.
* docs: add JSDoc comments to source-control hooks and components
Clarify the purpose, behavior, and constraints of test-harness functions,
directory-row components, and the git-history hook to help maintainers
understand the extracted and refactored source-control module.
* refactor: organize source-control into lifecycle dest folders
* fix(source-control): clear remaining react-doctor findings
Reset worktree- and history-scoped state during render, keep Cmd/Ctrl
selection updates free of setter side effects, and key graph paths by
swimlane/parent id. Also add the PR LoC helper scripts the quality
workflow fetches from the branch head.
* fix(source-control): refetch git history when owner host changes
Track activeRuntimeEnvironmentId as a stable key in useSourceControlGitHistory so that when the owner host changes but the worktree and path remain the same, the git history panel correctly refetches from the new host instead of keeping stale commits from the previous one. Add ownerHostKey to the useEffect dependency array to trigger refetch on host changes. Include JSDoc documentation for related components and expand test coverage to verify the host-change scenario.
* fix(source-control): refetch git history when owner host changes
Track activeRuntimeEnvironmentId as a stable key in useSourceControlGitHistory so that when the owner host changes but the worktree and path remain the same, the git history panel correctly refetches from the new host instead of keeping stale commits from the previous one. Add ownerHostKey to the useEffect dependency array to trigger refetch on host changes. Include JSDoc documentation for related components and expand test coverage to verify the host-change scenario.
* refactor(source-control): consolidate bulk mutation error handling
Extracts repeated error reporting into a dedicated helper function and applies
it consistently across all bulk stage/unstage handlers, including two that were
previously missing error handling.
* refactor(sidebar): group worktree-list files by domain
Follow-up to #14465 / #14467. Keep the landed extract and reorganize the
flat worktree-list dump into drag/, headers/, reveal/, rows/, scroll/,
and viewport/. Fold tiny modules into their owners, move leftover
sidebar-root files into the module, and retarget imports and source-path
tests. Layout-only; no behavior change.
* fix(sidebar): merge duplicate virtual-rows imports
Inlining virtual-row-dom-attributes left a second import from the same
module, which fails audit:code-quality:native --deny-warnings.
* refactor(sidebar): condense indentation comments
Shorten explanations to focus on the essential why, removing redundant
detail and improving readability without changing functionality.
* refactor: organize worktree-list into lifecycle dest folders
* fix react doctor
* fix: update reliability-gates path after worktree-list reorg
host-filtering.test.ts moved from viewport/ to listing/; keep the
runtime-routing.active-server-preference gate pointing at the real file.
* Extract workspace status colors to design tokens
Define theme-aware color tokens for workspace PR-state indicators (done, in-review, in-progress) to ensure consistent identity across theme switches. Update references to use the new tokens and refactor EmptyState button to use the Button component.
* fix(sidebar): stop mutating refs during worktree-list render
React Doctor fails static analysis when refs are written in render.
Commit reused array identity and the Smart live-signal latch after
paint, and return the attention map from the sort memo instead of
stashing it on a render-time ref.
`requestHiddenOutputRestoreIfNeeded` tracked its in-flight task with a `.finally`
handler that re-armed the restore. The task body is `while (!disposed)`, so once
the pane is disposed it exits immediately, the handler runs at once and re-arms
again — an unbounded self-feeding promise chain that consumed ~4GB in ~12s and
starved the microtask queue.
The pane is gone at that point and there is nothing to restore, so the handler
now returns early when disposed.
This was only invisible because the 25k-line pty-connection suite happened to run
a later test that tore the loop down; splitting that file into focused suites left
the arming tests at the end of a file and turned it into a reproducible OOM.
The regression test drives the restore into its armed state, disposes the pane and
counts how many times the chain re-reads `isVisibleRef`. Before this change it
cycled 196 times and OOM-killed the worker; after, it settles immediately.
* feat(shortcuts): warn when macOS Mission Control captures digit chords
Mission Control's Switch to Desktop shortcuts (Ctrl+digit by default,
present whenever the user has multiple Spaces) are consumed by
WindowServer before the app receives the event, so Orca's digit-range
shortcuts silently do nothing and the app can never observe the press.
Detect the conflict instead: probe com.apple.symbolichotkeys through
the live-prefs pipeline on Shortcuts pane mount and surface a standing
conflict warning on the affected rows, counted into the Conflicts stat.
Non-Darwin and the web client return no chords, and any probe failure
yields an empty result so a missing signal can never show a false
warning.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(shortcuts): harden Mission Control conflict warnings
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
The two cwd-repair tests really did `process.chdir()` into a temp directory and
delete it, so for that window every other test file sharing the worker process saw
a missing cwd. That was already a hazard; splitting pty-subprocess.test.ts into ten
files made the window overlap far more work, and the tests started failing
intermittently in full-suite runs while passing in isolation.
Stub the boundary instead: report a non-existent daemon cwd via `process.cwd()` and
spy on `process.chdir()` to capture the repair target. The assertion gets stricter
rather than weaker — it now names the directory the repair chose instead of
observing where the process happened to land — and no process-global state moves.
Both cwd-repair tests had failed at least once in prior full-suite runs; the suite
is green across repeated runs with this change.
Replace the per-editor-tab Array.find over the worktree's unified tabs with a
lazily-built id/fileId map, and swap two O(n^2) includes-in-a-loop scans for
Set membership. The map is only materialized when a snapshot actually carries
a mirrored editor tab, so terminal-only snapshots pay nothing.
* perf: coalesce SSH git status reads
* fix(ssh): key status leases by branch-line-total fork point
The request payload carries branchLineTotalMergeBase but the lease key did
not, so a strict refresh asking for the total could join an in-flight poll
that omitted it (blanking the branch-header chip) or reuse a pre-commit
fork point. Mirrors the guard already on the local path in git/status.ts.
#14738 landed a template literal interpolating an untyped fetch url, which
fails audit:code-quality:type-aware. That audit runs in the static analysis
job, so main is currently red and every open PR inherits the failure.
* fix(workspace-cleanup): refuse removal when the owning host is not certain (STA-4343)
* fix(workspace-cleanup): distinguish host collisions
* fix(workspace-cleanup): recheck host at removal boundary
* fix(e2e): unblock golden file-link hover and Windows worktree activate
Mac/Windows tmp paths wrap across xterm rows, so locateLink never
found the full absolute path. Print ./package.json instead.
createGoldenWorktree used os.tmpdir() (Windows 8.3 RUNNER~1) while
Git listed the long path, so activateGoldenWorktree never matched.
Realpath after worktree add and compare on the Node side.
* fix(e2e): handle realpath failures in golden worktree creation
Ensure half-built worktrees and branches are rolled back when
realpathSync fails, preventing leaks into later test runs. Extract
error handling into rollbackGoldenWorktree() for consistent cleanup.
* refactor(tests): split oversized test files off the max-lines suppression list
Every `*.test.ts`/`*.spec.ts` that carried an `eslint/oxlint-disable max-lines`
directive is now split into focused, behavior-scoped suites that fit the 800-line
test budget, with shared setup extracted into co-located `*-test-harness.ts` /
`*-test-fixtures.ts` modules (300-line budget). 83 files became ~930; the largest
output is 797 effective lines. `orca-runtime.test.ts` is intentionally untouched.
Test bodies were moved by scripted line-range slicing rather than retyped, so
assertions are byte-identical. The only permitted body edits were mechanical
rebinding where a shared value moved into a harness (e.g. `tmpHome` ->
`homes.tmpHome`).
Registries that enumerate test files were updated in lockstep:
- config/max-lines-baseline.txt: pruned 341 -> 258 entries (all 83 removed).
- config/reliability-gates.jsonc: 33 gates repointed at the split files, with
assertionRefs split per file where a gate's coverage now spans several.
- .github/workflows/pr.yml: the real-zsh lane now lists the 4 split files that
actually exercise zsh, so they keep running in the dedicated shell lane.
Also renamed agent-hooks `server-test-fixtures.ts` to `server.test-fixtures.ts`
so the global-fetch call-site audit keeps skipping it, and added `.js` extensions
to the CLI suites' dynamic harness imports (node16 resolution) to unbreak
`build:cli`.
Verification: full suite 52,449 passing vs 52,448 at baseline with zero
assertions lost; `pnpm lint`, `pnpm typecheck`, and `pnpm build:cli` all exit 0;
the terminal-pane e2e spec runs 31/31 headless.
* refactor(tests): split hook-idle arbitration suite that oxfmt pushed over budget
The pre-commit oxfmt pass reflowed pty-connection-hook-idle-arbitration.test.ts
to 811 effective lines, 11 over the test budget. Split the hook-completion side
effect and replacement-agent veto cases into their own suite; both files now sit
well under the cap and the 15 tests are unchanged.
* test: port upstream test changes into the split files after rebase
Rebasing onto main surfaced 27 tests that main had added to files this branch
deleted, plus edits to tests that had already moved. Taking the deletion side of
those modify/delete conflicts would have dropped that coverage silently, so each
upstream change is ported into the split file that now owns the behavior — for
example main's six orchestration mailbox tests land across orchestration-runs,
-send, and -check.
Also repoints `orchestration.notification-mailbox-consistency`, a gate main added
after this branch's gate remap, at those same three split files, and re-prunes
the max-lines baseline against main's (257 entries).
Verified: all 27 upstream test titles present; full suite 52,761 passing with the
only diff vs baseline being 12 tests main itself removed and 3 that moved from
skipped to passing; lint and typecheck exit 0.
* fix(test): flush pending continuations before tearing down terminal test globals
CI shard 5/16 failed on both Node 24 and 26 with `ReferenceError: window is not
defined` from pty-connection.ts, surfacing through
pty-connection-daemon-snapshot-replay.test.ts.
The reattach/settle chains `await` a real promise and then touch `window.api`.
Under fake timers those continuations cannot run, so they only become schedulable
once restoreTerminalTestGlobals() switches back to real timers — which previously
happened immediately before `delete globalThis.window`, so a late continuation
threw and failed the whole file. Flush async ticks in that window instead.
This is latent in the source rather than new: the pre-split 25k-line file kept
running other tests after these, which gave the chains time to settle before
teardown. Splitting the file moved teardown directly behind them.
* fix(test): keep an inert window after terminal test teardown instead of deleting it
The async-tick flush was not enough: the reattach/settle chain can resolve after
teardown regardless of how long we drain, so CI shard 5/16 still failed with
`ReferenceError: window is not defined` from pty-connection.ts.
A real renderer never loses `window`, so deleting it was the artificial part.
Swap in an inert proxy whose properties resolve to callables and whose calls
resolve to undefined, making a late `window.api.pty.*` call a harmless no-op.
The next test replaces it wholesale via installTerminalTestGlobals(), and no test
asserts that `window` is absent.
* feat(computer-use): support macOS middle click and gate the AX click path
`--mouse-button middle` already validated end-to-end through the CLI, the
zod schema, and the provider validator, and both the Windows and Linux
providers honored it. Only the macOS provider rejected it outright with
"middle-click is not yet supported", so the flag was a dead end on the one
platform that has no fallback.
Two changes:
- Add `.middle` to the macOS button mapping. macOS has no dedicated middle
event family, so it rides `otherMouseDown`/`otherMouseUp` with the button
number carried by `mouseButton: .center`; that constructor argument is
honored for exactly the `otherMouse*` types, so no extra field write is
needed.
- Validate the requested button before the accessibility fast path, and skip
that path for buttons it cannot express. Previously the raw string was read
unvalidated, and `performClickAction` only special-cased `right`, so
`click --mouse-button middle --element-index N` (no modifiers, count 1) fell
through to `AXPress` — a left click — and reported success with
`path: "accessibility"`. Any unrecognized button string did the same. This
matches guards the Windows and Linux providers already had.
The button enum moves into `OrcaComputerUseMacOSCore` so it is unit-testable;
`main.swift` keeps only the CoreGraphics mapping.
Also documents `--mouse-button` in the computer-use skill guide, which never
mentioned the flag, so agents on Windows and Linux had no way to discover it.
* test(computer-use): cover macOS middle click in the real-desktop e2e suite
* test(computer-use): prove macOS middle-click delivery
* fix(terminal): hold a cursor chord until the composing syllable commits
The composed glyph reaches the pty from the composition session-end handler,
which runs after the chord's keydown. Only Enter was held for that, so every
other chord went straight out on the transport and overtook the text it was
typed after: with 가나 on the line, typing 가나다 and pressing Cmd+Left left
다가나, the composing 다 landed at the cursor's destination.
Defer any sendInput chord while a composition is live or its session has not
yet flushed. Korean 2-Set shows the shape most clearly — the platform replays
the chord unmarked after keyup, so isComposing is already false while the
session is still pending.
No fallback timer on this path. A newline arriving late still arrives, which is
what that timer is for; a chord arriving mid-preedit is the corruption the wait
exists to prevent, and a conversion can hold its candidate window open for
seconds. Dropping the chord costs one keypress, firing early costs a line.
Pane commands are unaffected: they are not sendInput actions.
Fixes#12871
* test(e2e): pin the composing-chord order at the pty
The unit coverage asserts the handler's ordering against a synthetic transport.
This asserts it where it is actually observable: the committed glyph and the
chord reach the pty by two different routes, and only their merged order is
visible to the shell.
Verified to discriminate — against keyboard-handlers.ts from main the same spec
reads 01 eb8ba4 0a, the chord ahead of the syllable, which is the reported
corruption byte-for-byte.
* refactor(terminal): add the composing-chord deferral without touching the Enter path
Nesting the new branch inside the Enter condition re-indented the whole Enter
block, which is the kind of diff that can silently change it. Keeping them as
sibling conditions leaves the Enter path out of the diff entirely.
* test(e2e): pin the renderer to macOS for the Cmd+Left chord
Cmd+Left resolves to \x01 only under the macOS branch of the shortcut policy, so
on a Linux shard the chord produced no byte and the spec passed by measuring
nothing — it failed in CI for that reason, not for the behaviour under test.
Pinning the platform is the established pattern for these specs, and
expectImePlatformPolicy fails loudly if the override does not take.
* fix(mobile-native-chat): retire pending bubbles glued into one transcript row
Mobile's native chat retires an optimistic pending bubble only when a
transcript user turn matches its normalized text at the expected ordinal.
When two rapid sends collapse into a single glued user row neither key
matches, so both bubbles pin below every newer reply for the rest of the
session — mobile has a parallel implementation with no glue handling at all.
Trim the send body once at the send seam so the bytes the host writes
verbatim and the reconciliation key describe the same message on every send
path, then add a bounded glue matcher: a greedy cursor walk that may only
consider transcript turns strictly AFTER each send's captured tail, so an
older turn that happens to read like the concatenation can never retire a
newer queued send.
Refs #14262
* fix(mobile-native-chat): harden glued pending retirement
* fix(mobile-native-chat): preserve pending image previews
* fix(mobile-native-chat): bound glue to loaded transcripts
Removes the appVersion!==packagedAppVersion check that returned 'severed' after every update, and re-expresses stale-daemon retirement under its own stale_bundle reason.
Field-validated on a machine in the failing state: real ShipIt update cycle, both daemons still version-mismatched, no toast, 200 terminals preserved.
* fix(worktree): never reissue a generated workspace name
Generated workspace names were deduped only against currently-live
worktrees, so deleting a workspace returned its name to the pool. A later
workspace could draw the same name, land on the same directory path, and
inherit the previous occupant's agent conversation history — coding-agent
CLIs key their prompt history and transcripts by cwd.
Names are now retired permanently per repo. The registry is written in
main with the name Git actually used (the create loop can advance past a
requested name on collision), and seeded once per run from workspace
directories and surviving agent transcript buckets so already-spent names
are excluded from the start. Suggestions degrade to -2, -3 variants
instead of recycling, and those variants retire too.
User-typed names are untouched: retirement filters suggestions only.
* fix(mobile): honor retired workspace names, on one shared implementation
Mobile hand-duplicated the desktop name-suggestion algorithm and deduped
only against live workspaces, so a phone could still be offered a name
whose deleted workspace left agent conversation state behind at that path.
Both platforms now call one shared selector in src/shared, so the two can
no longer drift. The host publishes retired names as an optional field on
the existing worktree.list response, and mobile fetches them per selected
repo while the create sheet is open — mirroring the desktop hook.
Mobile never calls worktree.list for its catalog (it uses worktree.ps,
which carries rows only), so this is a targeted request rather than a
change to the catalog or its cache. Hosts predating the field omit it and
mobile falls back to live-only dedupe, which is the pre-change behavior.
* fix(worktree): close retirement consistency gaps
* test(worktree): cover retirement runtime contracts
* fix(worktree): retire generated collision names
* fix(worktree): enforce retired names at creation
* refactor(ai-vault): extract the Claude project-dir encoder
The bucket-name encoder and its scope-boundary check were private to the
session scanner, so a second consumer had to reimplement them — and got the
per-character encoding wrong. Move both to a shared module with direct tests.
* fix(worktree): make the retirement seed scan actually match buckets
The bucket encoder collapsed runs of non-alphanumerics while the real one
emits a dash per character, so every dot-path bucket missed and the Windows
default workspace root (C:\...) matched nothing at all. Reuse the shared
encoder and its boundary check, which also stops a repo absorbing a sibling
whose path merely shares its prefix.
Also:
- Derive the workspace leaf by stripping the known encoded parent instead of
guessing from trailing dash segments, which retired the parent directory's
name whenever a workspace was named numerically.
- Reuse isAutoGeneratedCreatureBranchName so the -10 and -100 tiers retire.
- Drop the .codex/sessions root: Codex keeps the cwd inside the transcript
rather than in a directory name, so the scan could only ever see a year
folder. Reading transcript contents is not a trade this feature justifies,
so the gap is documented instead.
- Honor CLAUDE_CONFIG_DIR, which relocates the bucket root.
- Delete the unused retirableLeafName export.
Tests write buckets with the real per-character encoding against a fake home,
covering POSIX, dot-directory, Windows drive and WSL UNC roots; all three
platform cases fail against the previous encoder.
* fix(worktree): retire only generated names, keyed by cwd namespace
Two problems in the host-side registry.
Retirement fired for every create, including names the user typed. The
creature pool contains ordinary words — orca, runner, sole, molly, oscar — so
typing a retired 'nautilus' silently produced directory and branch
'nautilus-2' and burned the name for good. Creates now carry an explicit
nameWasGenerated flag; both the skip and the retire are gated on it, and it
defaults to false so CLI and automation callers are unaffected.
The registry was keyed by repo id, but both readers already discarded the id
and unioned by the cwd collision key, because the collision this prevents is
on the path. Keying by that namespace directly fixes several things at once:
entries no longer orphan when a repo is removed, remove/re-add no longer loses
every retirement for an unchanged path, the missing removeProject prune is
moot, and the backfill promise no longer merges into only the first repo id it
saw. The feature is unreleased, so no migration is needed.
Also:
- Memoize the collision key. It runs computeWorktreePath, which for a WSL repo
is a blocking execFileSync('wsl.exe') whose failure path is uncached, and
the previous code recomputed it once per repo on every create and every
listRetiredNames call.
- Drop retiredNamesByRepo from the worktree list result. It had no readers and
leaked onto 'orca worktree list --json', and its awaited backfill sat on CLI
selector resolution. The dedicated listRetiredNames RPC keeps its consumers.
- Make the three RuntimeStore methods required. RuntimeStore is file-private
with two constructors, so the 'older embedders' the optionality protected do
not exist, and the optional chain silently returned no retirements.
- Revert the unrelated forceDeleteBranch rewrite, and make room under the
file's line budget by extracting the create-args mapping instead.
* fix(worktree): send name provenance and stop gating Create on the fetch
Desktop and mobile now mark a create as generated-name only when the user
typed nothing and the composer fell back to the suggestion, so the host knows
which names it may retire.
Remove the retired-names loading gate from every create path. The host already
skips retired candidates before doing any git work, so the client gate bought
nothing while it could disable Create for the length of a full mobile
reconnect ladder (the wait had no timeout) and blank the desktop button
between queued creates. The suggestion still waits; the button never does.
Also make the web client call worktree.listRetiredNames instead of hardcoding
an empty list — the method is registered and mobile-allowlisted, so the
comment claiming no wire call existed was wrong — and filter the mobile
response to strings so a malformed row cannot throw during normalization.
* fix(worktree): key retirement by repo id and prune it with the repo
Reverts the collision-key storage key. It was a function of workspaceDir,
nestWorkspaces, worktreeBasePath and repo.path, so toggling any one of those
orphaned every retirement for every affected repo at once — trading a rare
churn (remove/re-add) for a common one. The read path already unions by cwd
namespace at query time, so cross-repo sharing never depended on the storage
key.
Instead, address the growth and orphaning directly:
- Drop the registry in removeProject, and in removeProjectForHost once the last
host's copy of the repo id is gone, alongside the sparse-preset deletes that
already follow this convention.
- Bound each repo's registry. The cap sits far above the 552-name pool because
evicting inside it would reissue a name whose agent state is still on disk;
only -2/-3 tier accumulation can ever reach it.
- Carry retirements through profile transfer, re-keyed to the destination repo
id and dropped from the source, mirroring sparsePresetsByRepo.
Separately, fix the backfill merge: the scan promise is cached per cwd
namespace, but it closed over the first repo id that triggered it, so a second
repo in the same namespace received nothing. The scan stays shared; the merge
moves out of the cached promise and runs for whichever repo asked.
Local repos re-seed on re-add through that backfill. SSH repos do not — the
scan cannot see the execution host — which is now stated in the module.
* docs(worktree): spell out why the retirement bound sits above the pool
Names the trap directly: the neighbouring 50/200 bounds cap histories, so
lowering this one to match them would silently start reissuing names whose
agent state is still on disk. Also states that oldest-first eviction is a
deliberate least-bad choice rather than a neutral one.
* fix(worktree): send name provenance from the web runtime client
This client hand-enumerates worktree.create params, so the new optional field
was silently dropped and typecheck could not see it. On web and paired-desktop
the host therefore never received it: generated names were never retired, and
the host-side skip that backstops a stale suggestion was disabled too. The same
client does fetch retired names for suggestions, so it was filtering against a
registry nothing ever wrote to.
The test asserts both directions, and fails without the fix.
* fix(worktree): retire names that took more than one collision suffix
isAutoGeneratedCreatureBranchName strips exactly one trailing -N, which is
right for auto-rename eligibility but wrong here. Once the pool is spent the
suggester emits nautilus-2, and a collision on that yields nautilus-2-3 —
which a single strip leaves as nautilus-2, not a pool name, so retirement
no-opped at exactly the tier where every base name is already gone. Strip
repeated suffixes locally rather than moving the auto-rename predicate.
* perf(worktree): keep the retirement backfill off the blocking WSL probe
The backfill runs on composer repo-select, not just at create time, and it
derived the probe path synchronously — which for a WSL repo with a mirrored
workspace dir reaches getWslHome and its blocking execFileSync('wsl.exe').
A stopped distro froze the main process for up to 5s on composer open.
Adds an async twin of computeWorktreePath and uses it for the probe. Resolving
the home there also warms the shared cache, so later sync callers are free.
Also stops memoizing the collision key when the WSL home is still unresolved:
only the success path is cached upstream, so caching the fallback namespace
would strand the repo there for the rest of the session.
* fix(worktree): hold retired names across a refresh instead of blanking
refreshKey changes on every workspace-list mutation, so create-multiple
refetches after each create and the hook returned an empty list until the
refetch landed — precisely the window in which resetForNextCreate clears the
name field and a fresh suggestion is drawn. Keep the previous answer while
revalidating and reset only when the repo changes; a failed refresh keeps what
was already loaded rather than un-retiring everything.
Also makes the returned array referentially stable, so the suggestion memo
downstream stops rerunning on every refetch.
* refactor(worktree): put the retired-name cache rules on one implementation
The desktop and mobile hooks that fetch retired names had already drifted
four ways. The transports genuinely differ (IPC vs RPC), but the caching
rules must not, and mobile's copy reset to [] on any error -- which
un-retires every name for the rest of the sheet session, the one outcome
retirement exists to prevent.
Moves the rules into src/shared/worktree/retired-name-cache: response
normalization, the never-leak-across-repos rule, and the hold-previous-on-
failure rule. Pure, no React, because src/shared is on the main process's
import graph. Each platform keeps its own transport and effect.
Mobile moves up to desktop's behavior: it now holds the previous answer
through a failed refresh, and refetches when the workspace list changes
instead of never refetching after mount.
Also drops the unused `loading` return. Neither platform consumed it; its
only consumer was the Create-button gate reviewed out earlier, and removing
it makes that regression unexpressible.
* fix(worktree): import shared types from their real modules
Main dropped the src/shared/types barrel, so the retirement module's import
resolved locally but not against the PR's merge base.
* refactor(worktree): bound the retirement registry by tier compaction, not eviction
Retirement is a correctness guarantee — a spent name's directory may still hold
agent conversation state keyed by that cwd — so the 2000-entry cap was the wrong
shape: reaching it handed a name back. At the owner's measured rate (~6.6 pool
names retired per day in one repo) the cap was ~9 months out.
Names come from a fixed 552-entry pool and the suggester only reaches tier N+1
once every tier-N name is taken, so a completed tier is exactly a set that no
longer needs listing. A row is now a watermark plus the names above it: reads
answer at-or-below the watermark with no lookup, and compaction drops the 552
entries the watermark now covers. Bounded at one pool per repo forever, with no
eviction and nothing un-retired.
Tiers can complete out of order (a create-time collision can spend `nautilus-2`
while tier 1 is open), so compaction loops and higher-tier names simply wait.
The RPC result carries the watermark beside the names as a new field; a client
predating it reads the names only and under-retires the compacted tiers, which
degrades to the pre-retirement behavior rather than breaking.
* fix(worktree): preserve generated name retirement across failures
* fix(native-chat): stop glued rapid sends from pinning queued bubbles
Trim the draft at the send boundary so the PTY body and the optimistic
echo's match key agree, and bound the glue matcher to rows after the
oldest open echo's send boundary.
Fixes#14262
* fix(native-chat): preserve exact prompt payloads
* fix(skills): cancel and bound abandoned skill discovery scans
A root on a stalled network mount never settles its readdir, so after 30s
the coalescer starts a replacement walk. The abandoned walk kept running
with no cancellation, and nothing counted it, so live filesystem work
accumulated for the life of the process.
Superseding a scan for age now aborts it, and `findSkillFiles` plus the
candidate tasks bail on the signal — that cannot unblock a syscall already
in the kernel, but it stops an abandoned walk issuing more of them. A
budget caps how many abandoned scans may be live at once; past it a
replacement is shed rather than started, and the stalled entry is left in
place so the root recovers on its own once the mount answers.
The budget counts abandoned scans rather than live ones on purpose: one
discovery legitimately walks a dozen-plus roots at once, so a cap on live
scans would shed healthy roots and empty the picker.
A `refresh` still does not abort what it supersedes — on a healthy root
that scan is about to finish and its callers want an answer, and the
existing publish fence already stops it writing a pre-mutation result.
Shed roots report the new `unavailable` skipped reason so they stay
distinct from roots that genuinely are not there.
* fix(skills): propagate a walk abort thrown through a symlinked directory
The broken-link `catch` around the symlink branch also wrapped the nested
`visit`, so an abort thrown from inside a symlinked subtree was swallowed.
When that link was the last entry, nothing afterwards re-checked the
signal and the walk returned a truncated listing as success — the exact
outcome the abort path exists to prevent.
Only the `stat` is guarded now. The test pins the narrow window by
aborting inside the symlink's stat, after the entry loop's check and
before the nested visit's.
* fix(skills): degrade an aborted root instead of failing the whole discovery
Aborting a scan abandoned for age gave the coalescer a way to reject that
it did not have before: every error inside the walk and the candidate
tasks was caught locally, so the task effectively could not fail. Callers
already waiting on that scan now see it reject, and `scanRootShared`
re-threw anything that was not a shed — so one slow root failed the
entire discovery and emptied the picker for every healthy root beside it,
which is exactly what the shed path was written to avoid.
Both endings now mean the same thing to a caller that can degrade a
single root, behind one predicate: shed before the walk began, or aborted
after it was abandoned.
`discoverSkillsOnTarget` runs a second coalescer over the whole target,
where there is no partial answer to degrade to, so it converts both into
a retryable error rather than leaking the internal class to IPC/RPC. It
must not answer with an empty result — zero skills reads as "nothing
installed" and re-offers installs for skills that are present.
Also drops the scan key from the shed message, which carried an absolute
workspace path to a paired client and the renderer's error string, and
corrects the budget comment: it is global and per-abandonment, so one
wedged root can spend it alone, one replacement per 30s window.
* fix(skills): keep the original error as the cause of a stalled-target error
The target layer replaces the internal error with a user-facing one, which
discarded what actually went wrong. Attaching it as `cause` keeps that in
logs while the message stays free of the host path.
Also records why name-matching is narrow rather than broad: the walk and
the candidate tasks catch every filesystem error locally, so the only
AbortError that can escape a scan is the one its own signal raised.
* test(skills): pin that a real abort produces the name the predicate matches
`isSkillRootUnavailableError` decides on `error.name`, and the discovery
and target tests fabricate that shape rather than driving a real abort. So
nothing pinned the linkage: if `throwIfAborted()` stopped producing an
`AbortError` that is an `instanceof Error`, every test would still pass
while a stalled root began failing whole discoveries again.
The walk test already drives a genuine abort, so it asserts the name, the
Error subclassing DOMException relies on, and the predicate itself.
* fix(repos): forget remote-identity deadlines for removed repo locations
`probeRetryAfterByLocation` is keyed by connection and path and was
written for both resolved and unresolved probes, but never pruned — a
removed repo or a retired SSH host kept its deadline for the life of the
process. `isIdentityRefreshDue` also seeds an entry for every resolved
repo on first sight, so the map grew with repository and host churn even
with no probe activity.
The candidate sweep already enumerates every live repo, so it now
reconciles the deadline map against those location keys. `getRepos()`
reads a hydrated in-memory array, so a repo is never transiently absent
mid-sweep and cannot lose its startup delay or its backoff. Locations
with a probe still in flight are kept, since that probe re-adds its own
key when it settles.
* refactor(repos): drop the no-op in-flight exemption from the deadline prune
The exemption's own comment named the reason it was unnecessary: a probe
re-adds its deadline when it settles, so skipping its key produced the
same map state as deleting it, and only live repos are ever candidates so
nothing read the entry in between.
Removing it also stops a probe that never settles from pinning its
deadline forever — the location's `git remote -v` can already hang with
no timeout, and the exemption turned that one stranded entry into two.
* fix(workspaces): gate the Jira palette match on the issue's tenant
Pasting a Jira issue URL matched any worktree whose linked item carried
the same `jiraIdentifier`, with no comparison of the site it came from.
Jira issue keys are per-project, not per-tenant, so every tenant with a
PROJ project has a PROJ-123 — pasting one tenant's URL could jump to a
worktree tracking a different tenant's issue entirely.
The stored linked URL is the only tenant evidence available here, so
where it exists it now decides, comparing origin and site path the way
the fallback already did. This also covers path-scoped Jira Server
installs sharing one host. The bare identifier still matches when no URL
was stored, since it is then the only evidence there is.
This mirrors the check `isWorkspaceLinkedItemSourceContextMatch` already
makes on the same fields.
* test(workspaces): cover the reachable Jira identifier fallback
The fallback fixture used a blank url, a shape `normalizeWorkspaceLinkedItem`
rejects outright, so it pinned a state that cannot reach the palette. The
reachable way to have no tenant evidence is a url that is present but is not
a Jira browse link, which is now what the test uses.
Also covers a pasted url carrying Jira's `atlOrigin` tracking query, since
the matcher compares pathname only.
* fix: preserve agent badges when searching for tabs
A searched-for tab is exactly when its agent status matters—the map
covers all open tabs, not just recent ones.
* test: scope recent-tabs search test assertions to matched tab
- Make setCommandQuery checking explicit to prevent skipped assertions
- Scope badge verification to the specific queried tab instead of global query
- Add check that only the searched tab appears in results
A terminal tab published as `status: 'pending-handle'` renders the session
screen's spinner. Leaving it requires a snapshot that carries the materialized
handle, but a certified-live tabs stream parks `poll()` unless
`hasRecoveryNeed()` says otherwise — and that predicate never considered a
pending terminal. A host that mints the handle without republishing therefore
stranded the pane on its spinner forever: measured live, zero further
`session.tabs.list` calls over 90s while `terminal.list` kept firing every 2s.
Mirrors the existing native-chat recovery-need pattern. Client-only; no wire
change.