* perf: select highest usage totals without full sorting
* test(usage): pin first-inserted tie-break contract for highestUsageKey
Document why the strict > and the NaN sort fallback are load-bearing, and
cover the tie/re-set ordering the replaced stable sort guaranteed.
---------
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
* perf: lazily index case-insensitive Windows environment keys
* test(windows): pin env expansion fallback against the per-miss lookup oracle
Adds zero-enumeration, first-case-variant-wins, prototype-chain and 4,000-case
randomized differential coverage, and groups the new cases under their describe.
---------
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
`extractPartialEscapeTail` broke its own fold invariant
(extract(a + b) === extract(extract(a) + b)) in the oscEsc/stringEsc states,
so a PTY read that split there produced a different pending tail than the
same bytes delivered whole — the tail snapshots append after a restore.
Two causes, both in the "ESC did not terminate the string" branch:
- CAN/SUB were routed through `stateAfterEscByte`, which maps them back to
`esc` instead of aborting to ground. `extractPartialEscapeTail('\x1bPx\x1b\x18X0abc')`
returned '\x1b\x18X0abc'; the chunk-split fold returned ''.
- A second ESC opened its new sequence at `i - 1` rather than at itself.
`extractPartialEscapeTail('\x1b] \x1b\x1b^')` returned '\x1b\x1b^' whole but
'\x1b^' folded. The fold was right — xterm starts the sequence at the second ESC.
The existing fuzz only asserted the fold as `advance(extract(pending), chunk)`,
which is a tautology because every PENDINGS entry is already a tail. Replaced
with a sweep that re-splits the combined stream at every code-unit boundary,
and extended the alphabet (NUL, 0x20 intermediate, CJK) and SEQUENCES with
CAN/SUB and doubled-ESC-inside-string cases. A 1.25M-split fold fuzz over a
VT alphabet goes from 421 failures to 0.
* perf(native-chat): preserve historical tool rows while streaming
* perf(native-chat): short-circuit identical rows and lock producer immutability
Most folded rows come back as the input object, so compare identity before
scanning fields and blocks. Add a regression test for the invariant the reuse
cache depends on: ordering and folding never rewrite producer-owned messages
or blocks, which reused rows alias.
* perf(native-chat): bound journal reads during paged catch-up
* perf(native-chat): reduce the journal once per catch-up run, not per page
Bounding the SQL read per page left the JS side still O(total items) per
page: every page re-reduced the whole timeline and rebuilt the live-item
map, and the byte-shrink loop rebuilt it again on each halving.
A catch-up run is a synchronous loop with no await between pages, so the
reduced timeline is loop-invariant. `createAgentSessionCatchUpReader`
holds one snapshot for the run and re-reduces only if the journal cursor
actually moved, and the projection's live-item / alias / submission-byte
indexes memoize on the snapshot arrays the reducer rebuilds on change.
Per catch-up over a 8,000-message backlog: 40 timeline reductions to 1,
reduce+project time 24.3ms to 3.5ms, end-to-end 134.6ms to 111.3ms.
* fix(remote): keep terminal tabs syncing after orphan recovery
* fix(remote): validate recovery snapshots (#19065)
Address CodeRabbit feedback discussion_r3943677920 by validating the complete session-tabs payload before orphan recovery can publish it. Reject malformed rows and metadata as a whole while preserving optional and unknown additive fields.
Add validation and recovery/mirror regressions proving invalid follow-up reads retain the previous inventory and retry successfully.
Validation: 805 tests passed across 49 files; all four changed files pass Oxlint 1.80.0.
Note: pre-existing web typecheck errors in psl/emojibase-data resolution and export-let-function-initializer-ban.test.ts are unchanged from upstream.
* fix(remote): test reachable pending recovery states (#19065)
Address Pullfrog feedback discussion_r3943701696 by removing the retirement guard the host projection cannot reach and validating both affected fixtures through real host finalization.
Cover exact retirement, ready rebinding, pending/no-proof retention, and newer pending rows surviving prior authoritative removal. Preserve the host wire format and retain-on-unverifiable policy.
Validation: 807 tests passed across 50 files; all five changed files pass Oxlint 1.80.0.
Note: pre-existing web typecheck errors in psl/emojibase-data resolution and export-let-function-initializer-ban.test.ts remain unchanged.
* fix(remote): keep recovery reads tolerant of newer hosts (#19065)
Narrow the post-adoption snapshot validator to the fields recovery and the
mirror's coordinate logic actually consume. The previous schema closed every
enum and discriminant on the session-tab channel, so a host that published an
unknown agent name, status state, or tab kind failed the whole parse and
recovery retained forever - the same permanently-invisible-terminal symptom
this PR fixes. Unknown labels now pass through; structural defects in consumed
fields (coordinates, handles, groups, layouts, active selection) still fail
closed, and the three adoption regressions pinning that keep passing.
Replace the cyclic-layout test, which assumed a zod v3 stack overflow that
zod v4 cycle-detects away, with a throwing-accessor case that exercises the
same fail-closed branch.
Thread expectedRuntimeId through refreshWebRuntimeSessionTabsSnapshot so the
fifth recovery call site fences its post-adoption read like the other four.
Takes over stablyai/orca#19065 from its original author.
Co-authored-by: Shahar Mor <shaharmor1@gmail.com>
---------
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
* Update PR checks fix prompt to verify failure causality before fixing
Revise the prompt to classify failures as caused by this branch, not caused,
or uncertain before making changes. Only proceed autonomously for confirmed
failures; ask the user for guidance on uncertain or unrelated issues to avoid
fixing failures that weren't caused by the branch.
* Update PR checks fix prompt to verify failure causality before fixing
- Emphasize investigation phase by reframing prompt: "Investigate" rather than "Fix"
- Extend untrusted-data warning to all investigation sources (repository files, commit messages, diffs, CI output)
- Add test verifying injection safety: malicious input confined to JSON payloads, never as prompt instructions
* Refactor buildFixChecksPrompt test to focus on field mapping
The wrapper's only responsibility is renaming mobile PR fields onto the
shared prompt builder. Remove assertions about prompt wording, which are
already covered by the builder's own test suite. Simplify the test to
verify the field mapping contract and nothing else.
- Update automations API to use runtime.call pattern with automation.create
- Refactor browser creation flow to use state helpers instead of file explorer
- Simplify Playwright selectors and context menu interactions
- Remove fixture file creation from test setup
Monitoring turns emit no working event (4b2e3dded0), so the titlebar
badge must not count them either, or it lights with no unread row to
clear. The build cache's cached-events ternary could never take its
cached path because the early return above already covers it.
* Fix native PTY I/O failures disabling session termination
I/O failures on write or resize were incorrectly treated as exit evidence,
which disabled further termination attempts and producer flow control.
Separate ioFailed state from dead state; I/O errors suppress operations
but keep kill, forceKill, and signal available. Publish physical exit
before notifying listeners to prevent reentrant cleanup attempts from
accessing the retired native process.
* Fix PTY I/O cleanup tests and config syntax error
- Fixed missing closing brace and comma in reliability-gates.jsonc
- Added clear() operation to PTY mock fixture and test coverage
- Enhanced assertions to verify operation suppression during I/O failures
- Improved kill operation error handling with better promise-chain assertions
- Updated test result summaries in reliability gate documentation
* refactor(activity): rank status groups by attention level
Establishes consistent group ordering by introducing an attention-based ranking system, ensuring status groups maintain a fixed order regardless of thread recency. Consolidates thread status classification logic into `activityThreadStatusId` and simplifies group key naming.
* refactor(activity): emit working state for live agent turns
Activity events now emit working state for current turns,
enabling attention ranking above historical states.
* fix(activity): preserve working turns and count as unread
- Remove working-state events from cap logic so live turns stay visible
- Count fresh working/monitoring as unread in Activity badge
- Extract state-checking to activity-event-state module
- Use agentStatusEpoch for freshness-based invalidation
* fix(activity): subscribe only to epoch for unread count, not status map
The unread receipt is keyed on turn boundaries (stateStartedAt), not
heartbeats (updatedAt). Only the epoch matters; read the status map
directly via getState() to avoid wasteful re-renders on same-turn
heartbeats.
* fix(activity): prevent monitoring turns from emitting working events
Monitoring turns should surface only via the 'monitoring' snapshot in
the live state, not as separate working events that would contradict
the snapshot signal.
* Stabilize scrollbar gutter to prevent message list layout shift
- Add `scrollbar-gutter:stable` to prevent reflow when scrollbar appears
- Adjust scroll container padding to properly accommodate the scrollbar
- Add 5px horizontal inset to content for alignment with composer field
* Simplify message list padding and update scrollbar-gutter
emitPinnedGroup was the one section emitter that appended worktree rows
without hostContextLabelByWorktreeIdentity, and the mixed-host map it
would have received was computed over naturalWorktrees, which under the
default pinned policy has the pinned worktrees filtered out. Under that
policy a pinned worktree renders only in the Pinned section, so a pinned
remote workspace had no host badge anywhere.
Compute the mixed-host map over the full worktree set and thread it into
the Pinned emitter. Single-host sidebars still draw no badge.
Fixes#18472
A paired runtime host's first status publication counted as a connection
change, advancing the connection generation. Worktree scans already in
flight against that same connection were then discarded, so the sidebar
showed a strict subset of the host's worktrees until an unrelated refresh.
Two independent defects, both fixed:
- `connectionChanged` conflated "no entry yet" with "recorded unreachable".
Only the latter is a reconnect. The provider-session bump keeps the
broader predicate, since a first publication is a real session start for
integration-readiness caches.
- A stale-generation result was thrown away with no retry, so even a
genuine mid-flight reconnect silently dropped completed work. The scan is
now re-read once against the new generation.
* test: cover keyboard input in five simultaneously flooding SSH panes
* test: capture pane focus and buffers on flood input failure
* test: capture pane focus and buffers on flood input failure
* test: capture pane focus and buffers on flood input failure
* test: record replay input loss and application fix dependency
* test: record merged replay-input fix in the five-pane flood gate
* fix(pi): show input modals as waiting instead of working
* test(pi): verify real input dialogs through Electron CDP
---------
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
* perf(renderer): avoid per-second spinner animation events
* fix(bench): ensure the Electron runtime before bench:spinners
The script launches Electron via Playwright but skipped ensure:electron-runtime,
which every other Electron-launching bench script runs first.
* docs(renderer): scope spinner pixel-tolerance claim to paused-animation checks
---------
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: pullfrog[bot] <226033991+pullfrog[bot]@users.noreply.github.com>
* fix: preserve terminal retirement proof across renderer publications
* refactor: share the live-surface filter between retirement proof preservation and projection
The publication projection already dropped proofs whose surface is live;
reuse that as one helper instead of a second inline scan.
* fix: emit stored retirement proofs from host-authored snapshot writes
Three callers built a snapshot, stored it, then emitted the pre-store object. Storing grafts on the preserved proofs, so those frames carried the stored snapshotVersion without the proofs; subscribers dedupe on version and never saw them.
* fix: send terminal retirement proofs once per stream and fence them by occupant
Proofs are pinned per worktree for the host's lifetime, so every snapshot publication — including a 50ms title tick — re-shipped up to 64 proofs (~17 KB on realistic ids) to every paired client.
Negotiate session-tabs.retirement-proof-delta.v1: the host projects each session-tabs stream to send a proof only the first time that stream carries it, and a capable renderer keeps the union in a ledger keyed by (environment, worktree) with the same 64-entry bound and the same live-surface drop rule as the host, reset on removed frames and on a new connection generation. Legacy clients keep receiving the full list; CLI and mobile do not advertise the capability.
Also inherit worktreeInstanceId onto identity-less host writes so a host write between two renderer occupants can no longer launder one occupant's proofs into the next.
* fix: keep an empty proof delta distinguishable from a proof-less host
A negotiated stream now sends retiredTerminalSurfaces: [] when nothing is new instead of omitting the field. Absence is the host's "I hold no proofs" signal — which is also what a recreated worktree's fresh host entry publishes — so the client ledger forgets on absence and a successor occupant never inherits its predecessor's proofs, even when the removed frame was missed.
* test: pin ledger visibility against a legacy full-list host
An old host sends the full proof list whenever it holds any and omits the field when it holds none. Prove the new client ledger shows exactly what a legacy client would see across that sequence, so forgetting on absence is verified not to regress the mixed-version case.
* Unify sidebar create actions into a single dropdown menu
- Combine "New workspace" and "Add project" under a unified "Create" button
- Remove layout logic that split these actions based on sidebar width
- Normalize "Add Project" to "Add project" (lowercase) throughout the UI
* Use null instead of 'Unassigned' for unassigned shortcut labels
Add formatOptionalPrimaryShortcutLabel that returns null when a
shortcut is unassigned, enabling simpler conditional rendering in
dropdown menus. Remove associated translation strings.
* fix: preserve user input while terminal scrollback replays
* test: model multiple xterm user-input subscribers
* fix: keep mouse reports suppressed during replay and bind forwarders once
Real keystrokes now survive the replay guard, but xterm flags pointer
reports as user input too, and replayed bytes can leave mouse tracking
armed until the guarded mode reset lands. Keep those suppressed so a
click on restoring scrollback cannot print SGR fragments on the prompt.
Hoist the two provenance-bound forwarders out of the per-keystroke path.
* fix: keep wheel cursor keys off a replayed alt-screen frame
xterm turns a wheel notch into cursor up/down when the active buffer has
no scrollback, and flags it as user input. During a dead-TUI restore that
frame is replayed on the alt buffer and only leaves it when the guarded
?1049l lands, so forwarding those arrows would recall shell history at
the fresh prompt. Suppress them on the alt buffer only; the same bytes on
the normal buffer can only be a keyboard arrow and still survive replay.
Group the pointer-derived predicates in terminal-pointer-input-sequences.
* feat(native-chat): read a tool batch as a group
A run of several tool calls collapsed to one joined string: names and
arguments run together, separated by a middle dot that also occurs inside
`browser.open` and `tools/read`, with the overflow cut mid-token. Opened,
the member rows sat flush with the header and with the message content
around them, so the batch had no visible end.
Two presentation changes, no new derivation:
- Each member gets its own bounded pill in the collapsed header, carrying
its own category glyph, so the boundary between calls is a shape rather
than a character. Pills wrap instead of truncating, and members past the
summary cap are counted in `+N more` rather than dropped silently.
- Opened members are indented under the header, which is what marks where
the run ends.
`toolRunSummaryMembers` keeps the run's leading calls apart instead of
pre-joining them; `summarizeToolRun` now derives its string from it, so
mobile's header is byte-identical and the two cannot disagree about which
calls speak for a run.
Two existing behaviours are pinned by test rather than changed, both being
naming decisions rather than layout ones: the header still prints the raw
`mcp__linear__list_issues` while the row beneath prints the split name, and
a call carrying only a `url` still falls through to a JSON preview clipped
at 28 characters.
* fix(native-chat): bundle hidden tool count copy
* fix(native-chat): drop the filled pill for a glyph-led member list
Rendered in the app, the filled chips were wrong twice over. `bg-accent` is
reserved for hover/active row backgrounds, and the only full-strength use of
it in native chat is on payload and diff surfaces — so each member read as a
shrunken content block, and a run became the loudest thing in the transcript.
Worse, `flex-wrap` degenerated: at a 297px pane each member is 274-288px, so
every one took its own line, the header grew 24px to 72px, and the `5x` count
centred against the block landed beside the second member as though it counted
that call alone.
The glyph already marks where a member starts, so the fill was carrying no
information the icon wasn't. Members are now inline, glyph-led, and separated
by spacing; the list stays one line and truncates as a whole, as it did before
this branch. `+N more` moves outside the truncating span so the count of what
is not shown survives a pane too narrow to print the list.
Members carry `data-tool-run-member` rather than being found by their fill.
* fix(native-chat): let the run summary size to its content
`flex-1` on the truncating member list made it claim the header's slack, so
`+N more` was pushed to the far right edge with a gap between it and the last
member it counts. Without it the span still shrinks and truncates — `min-w-0`
plus the default shrink is what drives the ellipsis, which is how the header
worked before this branch — and the count now sits directly after the list at
every width.
* fix(native-chat): separate run-header members with real whitespace
An `ml-3` margin marks the boundary on screen but is invisible to a copied
selection and to the button's accessible name, so the header read
`ls -latools/read`. Adds a space text node between members and trims the
margin to pay for its width. `+N more` also picks up the hover transition
every other header segment already had.
---------
Co-authored-by: Merge Sim <sim@local>
* fix(xterm): fire Marker dispose before clearing line (#10879)
Scrollback trim under search highlights was O(k²) because dispose set
marker.line to -1 before onDispose, collapsing SortedList keys. Fire
listeners first so delete still sees the real line, then clear the line.
Fixes#10879
* fix(xterm): remove scrollback decorations by identity
* perf(xterm): avoid index arrays for unique decorations
---------
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
`pr-code-change-scope.mjs` read stdin with `readFileSync(0, 'utf8')`, a single
read of fd 0. Once the writer outgrows the 64 KB pipe buffer that read returns
early or throws EAGAIN, the script exits 0 having emitted no `name=value` pairs,
and `tee -a "$GITHUB_OUTPUT"` records nothing -- so every lane the classifier
gates is silently skipped rather than failing loudly.
A PR opened long ago carries a stale `pull_request.base.sha`, so the gate's
merge-base diff spans the whole base branch. PR #13178 diffed 13,294 files
(773 KB) against a base 1,592 commits behind main and lost typecheck, test,
static analysis, xterm patch sync, package and e2e to this.
Stream stdin instead, matching how the sibling `pr-e2e-source-routing.mjs`
already reads the same list in the same workflow.