Commit Graph
10509 Commits
Author SHA1 Message Date
Merge Sim 9634184fdf Separate launch cancellation persistence from delivery settlement 2026-09-08 00:25:02 -07:00
Merge Sim e9c73bd572 fix: settle failed outbox mutations and retire launch callers 2026-09-07 23:58:12 -07:00
Merge Sim daf83559a1 Type retained settlement handles in regression coverage 2026-09-07 23:15:51 -07:00
Merge Sim 37ca13790c Preserve uncertain dispatch budgets and observe shared outbox settlement 2026-09-07 23:13:55 -07:00
Merge Sim 94c6bd06ae Keep current outbox revision stamping linear in queue size 2026-09-07 22:40:57 -07:00
Merge Sim ae5492df00 Read live dispatch state before mounting another outbox subscriber 2026-09-07 22:31:37 -07:00
Merge Sim 97a17d6118 Cover superseded refusal against a live successor attempt 2026-09-07 22:30:33 -07:00
Merge Sim 373a7703bf Serialize structured outbox transitions and fence operation attempts 2026-09-07 22:29:36 -07:00
Merge Sim 951997b810 Restore quota spies even when a recovery assertion fails 2026-09-07 22:07:13 -07:00
Merge Sim 14b01d944b Own structured send recovery budgets in the durable outbox 2026-09-07 22:02:05 -07:00
Jinjing ba5f708290 Rank activity status groups by attention level (#19329)
* refactor(activity): rank status groups by attention level

Establishes consistent group ordering by introducing an attention-based ranking system, ensuring status groups maintain a fixed order regardless of thread recency. Consolidates thread status classification logic into `activityThreadStatusId` and simplifies group key naming.

* refactor(activity): emit working state for live agent turns

Activity events now emit working state for current turns,
enabling attention ranking above historical states.

* fix(activity): preserve working turns and count as unread

- Remove working-state events from cap logic so live turns stay visible
- Count fresh working/monitoring as unread in Activity badge
- Extract state-checking to activity-event-state module
- Use agentStatusEpoch for freshness-based invalidation

* fix(activity): subscribe only to epoch for unread count, not status map

The unread receipt is keyed on turn boundaries (stateStartedAt), not
heartbeats (updatedAt). Only the epoch matters; read the status map
directly via getState() to avoid wasteful re-renders on same-turn
heartbeats.

* fix(activity): prevent monitoring turns from emitting working events

Monitoring turns should surface only via the 'monitoring' snapshot in
the live state, not as separate working events that would contradict
the snapshot signal.
2026-09-07 21:50:14 -07:00
Jinjing 1a8640adb6 Stabilize scrollbar gutter to prevent message list layout shift (#19332)
* Stabilize scrollbar gutter to prevent message list layout shift

- Add `scrollbar-gutter:stable` to prevent reflow when scrollbar appears
- Adjust scroll container padding to properly accommodate the scrollbar
- Add 5px horizontal inset to content for alignment with composer field

* Simplify message list padding and update scrollbar-gutter
2026-09-07 20:24:59 -07:00
Jinjing 217125338e Log error details on session kill failure (#19381)
When a session kill operation fails, capture the error name and message
in the log to aid debugging and performance issue investigation.
2026-09-07 20:20:00 -07:00
Neil 7adb5b3dc4 fix(sidebar): label pinned rows with their host on a multi-host sidebar (#19351)
emitPinnedGroup was the one section emitter that appended worktree rows
without hostContextLabelByWorktreeIdentity, and the mixed-host map it
would have received was computed over naturalWorktrees, which under the
default pinned policy has the pinned worktrees filtered out. Under that
policy a pinned worktree renders only in the Pinned section, so a pinned
remote workspace had no host badge anywhere.

Compute the mixed-host map over the full worktree set and thread it into
the Pinned emitter. Single-host sidebars still draw no badge.

Fixes #18472
2026-09-07 19:54:30 -07:00
Neil 1e693edee4 fix(clipboard): route runtime-owned SSH image paste through the runtime (#17679) (#19352) 2026-09-07 19:54:16 -07:00
Neil fc78a7d9ca fix(runtime): stop a first status publication retiring in-flight worktree scans (#19357)
A paired runtime host's first status publication counted as a connection
change, advancing the connection generation. Worktree scans already in
flight against that same connection were then discarded, so the sidebar
showed a strict subset of the host's worktrees until an unrelated refresh.

Two independent defects, both fixed:

- `connectionChanged` conflated "no entry yet" with "recorded unreachable".
  Only the latter is a reconnect. The provider-session bump keeps the
  broader predicate, since a first publication is a real session start for
  integration-readiness caches.
- A stale-generation result was thrown away with no retry, so even a
  genuine mid-flight reconnect silently dropped completed work. The scan is
  now re-read once against the new generation.
2026-09-07 19:54:01 -07:00
Neil e182930670 test: cover input in five simultaneously flooding SSH panes (#19071)
* test: cover keyboard input in five simultaneously flooding SSH panes

* test: capture pane focus and buffers on flood input failure

* test: capture pane focus and buffers on flood input failure

* test: capture pane focus and buffers on flood input failure

* test: record replay input loss and application fix dependency

* test: record merged replay-input fix in the five-pane flood gate
2026-09-07 19:51:44 -07:00
Jinjing 66420537b7 fix e2e create menu races (#19448) 2026-09-07 19:45:54 -07:00
gatsby74andNeil a278d84a4e fix(pi): show input modals as waiting instead of working (#18836)
* fix(pi): show input modals as waiting instead of working

* test(pi): verify real input dialogs through Electron CDP

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-07 19:35:19 -07:00
aeddfa463d perf(renderer): avoid per-second spinner animation events (#19407)
* perf(renderer): avoid per-second spinner animation events

* fix(bench): ensure the Electron runtime before bench:spinners

The script launches Electron via Playwright but skipped ensure:electron-runtime,
which every other Electron-launching bench script runs first.

* docs(renderer): scope spinner pixel-tolerance claim to paused-animation checks

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: pullfrog[bot] <226033991+pullfrog[bot]@users.noreply.github.com>
2026-09-07 19:29:53 -07:00
Neil da836faeef fix: preserve terminal retirement proof across renderer publications (#19002)
* fix: preserve terminal retirement proof across renderer publications

* refactor: share the live-surface filter between retirement proof preservation and projection

The publication projection already dropped proofs whose surface is live;
reuse that as one helper instead of a second inline scan.

* fix: emit stored retirement proofs from host-authored snapshot writes

Three callers built a snapshot, stored it, then emitted the pre-store object. Storing grafts on the preserved proofs, so those frames carried the stored snapshotVersion without the proofs; subscribers dedupe on version and never saw them.

* fix: send terminal retirement proofs once per stream and fence them by occupant

Proofs are pinned per worktree for the host's lifetime, so every snapshot publication — including a 50ms title tick — re-shipped up to 64 proofs (~17 KB on realistic ids) to every paired client.

Negotiate session-tabs.retirement-proof-delta.v1: the host projects each session-tabs stream to send a proof only the first time that stream carries it, and a capable renderer keeps the union in a ledger keyed by (environment, worktree) with the same 64-entry bound and the same live-surface drop rule as the host, reset on removed frames and on a new connection generation. Legacy clients keep receiving the full list; CLI and mobile do not advertise the capability.

Also inherit worktreeInstanceId onto identity-less host writes so a host write between two renderer occupants can no longer launder one occupant's proofs into the next.

* fix: keep an empty proof delta distinguishable from a proof-less host

A negotiated stream now sends retiredTerminalSurfaces: [] when nothing is new instead of omitting the field. Absence is the host's "I hold no proofs" signal — which is also what a recreated worktree's fresh host entry publishes — so the client ledger forgets on absence and a successor occupant never inherits its predecessor's proofs, even when the removed frame was missed.

* test: pin ledger visibility against a legacy full-list host

An old host sends the full proof list whenever it holds any and omits the field when it holds none. Prove the new client ledger shows exactly what a legacy client would see across that sequence, so forgetting on absence is verified not to regress the mixed-version case.
2026-09-07 19:28:17 -07:00
Jinjing c056c6f9ac Unify sidebar create actions into single dropdown menu (#19375)
* Unify sidebar create actions into a single dropdown menu

- Combine "New workspace" and "Add project" under a unified "Create" button
- Remove layout logic that split these actions based on sidebar width
- Normalize "Add Project" to "Add project" (lowercase) throughout the UI

* Use null instead of 'Unassigned' for unassigned shortcut labels

Add formatOptionalPrimaryShortcutLabel that returns null when a
shortcut is unassigned, enabling simpler conditional rendering in
dropdown menus. Remove associated translation strings.
2026-09-07 19:24:34 -07:00
Neil 98b0c329ff fix: preserve user input during terminal scrollback replay (#19075)
* fix: preserve user input while terminal scrollback replays

* test: model multiple xterm user-input subscribers

* fix: keep mouse reports suppressed during replay and bind forwarders once

Real keystrokes now survive the replay guard, but xterm flags pointer
reports as user input too, and replayed bytes can leave mouse tracking
armed until the guarded mode reset lands. Keep those suppressed so a
click on restoring scrollback cannot print SGR fragments on the prompt.
Hoist the two provenance-bound forwarders out of the per-keystroke path.

* fix: keep wheel cursor keys off a replayed alt-screen frame

xterm turns a wheel notch into cursor up/down when the active buffer has
no scrollback, and flags it as user input. During a dead-TUI restore that
frame is replayed on the alt buffer and only leaves it when the guarded
?1049l lands, so forwarding those arrows would recall shell history at
the fresh prompt. Suppress them on the alt buffer only; the same bytes on
the normal buffer can only be a keyboard arrow and still survive replay.
Group the pointer-derived predicates in terminal-pointer-input-sequences.
2026-09-07 19:15:08 -07:00
OrcaWinandm4air de0a91b99f fix(deps): update Electron to reviewed 43.6 runtime (#19369)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-07 19:12:12 -07:00
Brennan BensonandMerge Sim c1e15c4008 feat(native-chat): read a tool batch as a group (#19372)
* feat(native-chat): read a tool batch as a group

A run of several tool calls collapsed to one joined string: names and
arguments run together, separated by a middle dot that also occurs inside
`browser.open` and `tools/read`, with the overflow cut mid-token. Opened,
the member rows sat flush with the header and with the message content
around them, so the batch had no visible end.

Two presentation changes, no new derivation:

- Each member gets its own bounded pill in the collapsed header, carrying
  its own category glyph, so the boundary between calls is a shape rather
  than a character. Pills wrap instead of truncating, and members past the
  summary cap are counted in `+N more` rather than dropped silently.
- Opened members are indented under the header, which is what marks where
  the run ends.

`toolRunSummaryMembers` keeps the run's leading calls apart instead of
pre-joining them; `summarizeToolRun` now derives its string from it, so
mobile's header is byte-identical and the two cannot disagree about which
calls speak for a run.

Two existing behaviours are pinned by test rather than changed, both being
naming decisions rather than layout ones: the header still prints the raw
`mcp__linear__list_issues` while the row beneath prints the split name, and
a call carrying only a `url` still falls through to a JSON preview clipped
at 28 characters.

* fix(native-chat): bundle hidden tool count copy

* fix(native-chat): drop the filled pill for a glyph-led member list

Rendered in the app, the filled chips were wrong twice over. `bg-accent` is
reserved for hover/active row backgrounds, and the only full-strength use of
it in native chat is on payload and diff surfaces — so each member read as a
shrunken content block, and a run became the loudest thing in the transcript.
Worse, `flex-wrap` degenerated: at a 297px pane each member is 274-288px, so
every one took its own line, the header grew 24px to 72px, and the `5x` count
centred against the block landed beside the second member as though it counted
that call alone.

The glyph already marks where a member starts, so the fill was carrying no
information the icon wasn't. Members are now inline, glyph-led, and separated
by spacing; the list stays one line and truncates as a whole, as it did before
this branch. `+N more` moves outside the truncating span so the count of what
is not shown survives a pane too narrow to print the list.

Members carry `data-tool-run-member` rather than being found by their fill.

* fix(native-chat): let the run summary size to its content

`flex-1` on the truncating member list made it claim the header's slack, so
`+N more` was pushed to the far right edge with a gap between it and the last
member it counts. Without it the span still shrinks and truncates — `min-w-0`
plus the default shrink is what drives the ellipsis, which is how the header
worked before this branch — and the count now sits directly after the list at
every width.

* fix(native-chat): separate run-header members with real whitespace

An `ml-3` margin marks the boundary on screen but is invisible to a copied
selection and to the button's accessible name, so the header read
`ls -latools/read`. Adds a space text node between members and trims the
margin to pay for its width. `+N more` also picks up the hover transition
every other header segment already had.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-07 18:24:30 -07:00
BingZandNeil 5bd0247aaa fix(xterm): remove scrollback decorations by identity (#13178)
* fix(xterm): fire Marker dispose before clearing line (#10879)

Scrollback trim under search highlights was O(k²) because dispose set
marker.line to -1 before onDispose, collapsing SortedList keys. Fire
listeners first so delete still sees the real line, then clear the line.

Fixes #10879

* fix(xterm): remove scrollback decorations by identity

* perf(xterm): avoid index arrays for unique decorations

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-07 18:18:04 -07:00
Neil cbc7bb418c fix(ci): read the changed-path list past the first pipe buffer (#19409)
`pr-code-change-scope.mjs` read stdin with `readFileSync(0, 'utf8')`, a single
read of fd 0. Once the writer outgrows the 64 KB pipe buffer that read returns
early or throws EAGAIN, the script exits 0 having emitted no `name=value` pairs,
and `tee -a "$GITHUB_OUTPUT"` records nothing -- so every lane the classifier
gates is silently skipped rather than failing loudly.

A PR opened long ago carries a stale `pull_request.base.sha`, so the gate's
merge-base diff spans the whole base branch. PR #13178 diffed 13,294 files
(773 KB) against a base 1,592 commits behind main and lost typecheck, test,
static analysis, xterm patch sync, package and e2e to this.

Stream stdin instead, matching how the sibling `pr-e2e-source-routing.mjs`
already reads the same list in the same workflow.
2026-09-07 17:57:05 -07:00
OrcaWinandm4air 102402e41e fix(deps): update DOMPurify sanitizer hardening (#19377)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-07 17:43:28 -07:00
OrcaWinandm4air b8f6c7cabe fix(deps): update react-i18next for TypeScript 7 and parser fixes (#19378)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-07 17:43:18 -07:00
OrcaWinandm4air a4e5106e14 chore(docs): patch brace expansion resource exhaustion fixes (#19382)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-07 17:42:55 -07:00
OrcaWinandm4air 1dae024ab2 chore(mobile): patch xmldom security fixes in plist tooling (#19380)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-07 17:42:48 -07:00
OrcaWinandm4air 91fbc1529a fix(deps): harden cloud HTTP and WebSocket dependencies (#19362)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-07 17:42:37 -07:00
OrcaWinandm4air cfd59f2ca0 fix(deps): update desktop parser security dependencies (#19361)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-07 17:42:29 -07:00
Brennan BensonandMerge Sim 3b8128df04 fix(orchestration): index per-PTY mailbox reservation cleanup (#19390)
Co-authored-by: Merge Sim <sim@local>
2026-09-07 17:42:23 -07:00
Neil 588240043e fix(sidebar): reveal collapsed workspaces without clearing filters (#19398) 2026-09-07 17:14:54 -07:00
Brennan BensonandMerge Sim 9f044031fc fix(native-chat): render compaction notices, plan documents, and images (#19228)
* fix(native-chat): render compaction notices, plan documents, and images

* fix(native-chat): avoid repeating notice text in details

* fix(native-chat): journal canonical and legacy compaction events

* test: add digest to native chat notice payload fixture

* chore(native-chat): drop the planning doc from the PR

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-07 16:21:47 -07:00
Neil 2265fce591 chore(mobile): remove stale max-lines exceptions (#19366) 2026-09-07 14:54:18 -07:00
Brennan BensonandMerge Sim d0506bf5de feat(native-chat): add execution details and tool row identity (#19226)
* feat(native-chat): annotate tool rows with execution and source details

* fix(native-chat): require explicit MCP identity for tool annotations

* test: add required state to MCP projection fixture

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-07 13:51:47 -07:00
Brennan BensonandMerge Sim 0b3af9ddbc Show native chat message timestamps on hover and keyboard focus (#19218)
Co-authored-by: Merge Sim <sim@local>
2026-09-07 13:46:21 -07:00
Jinwoo Hong bcd4076dd1 fix(relay): never cache a region hint from a one-region catalog (#19349)
* fix(relay): never cache a region hint from a one-region catalog

The director lists only regions with a serving cell, so a roll wave shortens the
catalog to one entry. The resolver required only every *listed* region to be
measured, so that lone region won against nothing and was cached for 24 h: a US
desktop refreshing while US cells rolled published asia-east2 for a day, the
incident #19233 was written to end. Now fewer listed regions than the fleet serves
withholds the hint (1 h no-hint TTL), the same outcome as an unmeasurable peer.

* test(relay): give the unstable-probe case a two-region catalog so it has one cause
2026-09-07 16:44:33 -04:00
Jinwoo Hong d936d8da82 revert(mobile): pull the relay connect-speed mobile pass pending a smaller, verified re-land (#19348)
* Revert "feat(mobile): time relay dial stages so diagnostics say where a slow connect went (#19245)"

This reverts commit 83b1558ecc.

* Revert "perf(mobile): race the direct and relay dials from t=0 on every reconnect (#19308)"

This reverts commit ceafdcad2f.

* Revert "feat(mobile): draw the last known tab strip while a session reconnects (mobile pass) (#19281)"

This reverts commit 643571def6.

* Revert "perf(mobile): open a session with parallel startup RPCs and a pre-warmed terminal engine (#19260)"

This reverts commit c37413271e.

* Revert "perf(mobile): cut the relay reconnect critical path and admit dead sockets faster (mobile pass) (#19280)"

This reverts commit e628090ad4.

* chore: keep the react-doctor suppression for the startup timers

The pattern it covers (a variable number of timers cleared through one cleanup)
predates #19260 and is unchanged by the revert; dropping the entry only re-exposed
a pre-existing finding to the changed-code gate.
2026-09-07 16:44:26 -04:00
Brennan BensonandMerge Sim 6a47d2831f fix(native-chat): scope composer file drops to the pane that received them (#19328)
* fix(native-chat): scope composer file drops to the pane that received them

A native OS file drop resolving to `target: 'composer'` carried no pane
identity, so the window-wide payload was attached by every mounted composer.
Because inactive chat tabs stay mounted (hidden), one drop populated every
chat pane's attachment cache, and those chips replayed whenever the user
returned to a tab they never dropped into. The workspace-creation composer
and chat composers also leaked into each other, since neither could tell
which surface actually received the drop.

Composer drops now carry a `scopeKey` the way a terminal drop carries its
tab and pane leaf id: the composer publishes its pane key as
`data-composer-scope-key`, the preload harvests it during the composedPath
walk, and each composer attaches only its own. The workspace composer's
last-wins ownership stack now claims unscoped payloads only.

* test(native-chat): supersede the bug-asserting drop repro with the scoping test

The repro that landed on main asserts the pre-fix behavior (a drop reaching
every mounted composer), so it fails once drops are scoped to the pane that
received them. Its scoping cases now live in
native-chat-composer-drop-scope.test.tsx, which keeps its editor-target
control case verbatim and adds coverage for unscoped composers and a scope
key published inside the drop-target marker.

* test(native-chat): cover workspace composer drop isolation

* fix(native-chat): authorize external attachment paths before preview

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-07 13:34:23 -07:00
Brennan BensonandMerge Sim 5cefb440bf fix(native-chat): stop an unanswered host from reading as one that refuses structured chat (#19321)
* fix(native-chat): stop an unanswered host from reading as one that refuses structured chat

`readLocalRuntimeCapabilities()` returned `[]` both before the first status probe
landed and after one failed, so "not asked yet" and "host says no" were the same
value. Every structured-chat launch route consumed it, and an unprobed host was
routed to legacy chat exactly as a refusing one is.

Keep the two apart: the cache holds `null` until a probe succeeds, a failed probe
leaves it `null` rather than emptying it, and the launch route names the case with
its own blocker instead of borrowing `runtime-capability`.

No routing outcome changes — both cases still decline structured chat. The point is
that the reason is now truthful, which is what the routing work needs to build on:
once a launch can target a runtime peer, capabilities come from that host, and an
unanswered remote must not be indistinguishable from one that refuses.

`hostCapabilities` on the launch route stays local-only at every call site; a
per-target resolver replaces it when the route learns to reach a peer.

* test: cover unknown runtime capability lifecycle and launch fallback

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-07 13:25:40 -07:00
Jinjing 1d1b73c408 Defer inactive browser pages across worktree switches (#19326)
* Defer inactive browser tabs while retaining their viewport slots

Restore worktrees and tabs on demand instead of mounting the full tree.
Only render active pages and those required by automation, mobile drivers,
or remote viewers. Inactive panes stay deferred with persistent viewport
slots so their webview guests survive chrome unmounts, reducing memory
overhead when opening workspaces with many tabs.

* Defer browser pages until active and recover if evicted

Pages defer rendering until active, then retain state when inactive.
Add recovery logic to restore guests evicted by workspace memory
pressure when pages are reactivated.

* Stop retaining browser content when worktree is inactive

- Browser panes and pages now unmount when their worktree transitions to inactive, except for pages claimed by automation/mobile/viewer consumers
- Prevents unwanted restoration of all hidden browser tabs when switching between worktrees
- Tests verify proper cleanup at scale and correct page lifecycle across worktree switches

* Preserve document-preview guests when switching browser tab profiles

Document previews use a fixed partition and should not be recreated when
the profile changes. Only URL-based pages need their webviews destroyed
and rebuilt with the new profile. Includes test coverage.

* Create browser pages cold to defer guest initialization

Pages created in the background now start with loading: false, since they
don't own a guest until first shown. Only live guests can report loading
status, so background tabs sit idle until activation triggers navigation.

* Prevent document preview from swallowing pointer events during drag

Move webview registration to attachDocPreviewWebview before append,
ensuring it's enrolled in drag passthrough before becoming hittable.
When a document preview tab remounts mid-drag, the previous hook-based
enrollment landed too late. Also refactor mountEligible into
isBrowserPagePanePaintable for clarity.
2026-09-07 12:20:27 -07:00
Brennan BensonandMerge Sim ce4a3a4186 feat(chat): add structured session rewind backend (#19235)
* feat(chat): add structured session rewind backend

* fix(chat): make interrupted session rewinds recover safely

* fix(native-chat): negotiate rewind runtime capability

* fix(native-chat): consolidate remaining adapter imports

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-07 12:20:24 -07:00
Jinjing 9fed61e5c2 Persist agents sidebar search visibility as pairing-local preference (#19313)
* Persist agents sidebar search field visibility as pairing-local preferen

- Add `agentsShowSearch` to workspace UI state with default on
- Include in pairing-local fields so preference syncs across clients
- Convert search from menu action to checkbox menu item for explicit toggle
- Update activity thread options menu to reflect checkbox state
- Add localization strings across all supported languages
- Update RPC schemas and preference persistence layer
- Includes readiness validation reports confirming feature is clean

* rm review

* fix documentation
2026-09-07 11:27:40 -07:00
Jinwoo Hong 83b1558ecc feat(mobile): time relay dial stages so diagnostics say where a slow connect went (#19245)
* feat(mobile): time relay dial stages so diagnostics say where a slow connect went

A 10s connect was unattributable from a shared report. Relay dial stages carried
no timestamps, so nothing could tell "the cell never answered relay-hello" from
"the E2EE handshake was slow", and the per-state dweltMs the client already
computed went only to console.log — invisible without a debug build.

RelayDialStageTracker now stamps each stage entry from a monotonic clock
(performance.now where present, wall clock otherwise) and returns the duration of
the stage it just left. The session logs one entry per stage, and settles the
in-flight stage on connect, failure, or close, so a dial that dies mid-way still
names the stage it never finished. dweltMs joins the same buffer as a structured
field instead of console.

Durations ride the existing per-host log buffer and its cap, so memory is
unchanged and no new storage appears. The report derives two lines from them: the
latest dial's stage breakdown (a reconnect loop must not average away the attempt
being reported) and total dwell per connection state. Both are numbers and
closed-enum names, and the entries still pass through the existing redaction.

* fix(mobile): never let a diagnostics sink break a dial, and pin timing names to their enums

Review follow-ups on the dial-stage timing work.

The stage timing emitted on the confirm's success path ran inside the try that
calls fail(), so an onLog sink that threw would have turned a good connect into a
failed session. The same hazard existed on the direct path, where the dwell emit
sits in publish() ahead of the listener loop and the connect waiters. Both sink
calls are now isolated: a broken sink loses a log line and nothing else.

The persisted-log validator accepted any string as a timing name, and the report
echoes that name unredacted. Names are now checked against the closed enum for
their kind, backed by Record<Union, true> tables so adding a stage or a state
breaks the build rather than silently widening what a corrupted store can inject.

Entry volume: every reconnect cycle walks four connection states, so logging each
one would roughly double what a slow-connect report holds against the unchanged
200-entry per-host cap. Transitions under 100ms are therefore not buffered. They
cannot be where a slow connect spent its time, and console still shows all of
them. States that flap slowly, which is the case support cares about, still land
in the log.

RpcClientConnectionState takes an optional clock so dwell thresholds are testable
without sleeping.

* fix(mobile): reject a negative stored stage duration when hydrating the log

A persisted timing only had to be finite to survive hydration, so a corrupted
`ms: -1` reached the diagnostics report, where the dial summary sums the stage
durations and a negative would subtract from the total. Producers clamp at 0
(`elapsedMs`), so anything below it is corruption. 0 itself still hydrates: a
stage the dial passes through instantly is real.

* refactor(mobile): move the relay liveness profile out of the session so the dial log fits

* fix(mobile): never let the liveness-timeout log line keep a dead relay connected

* test(mobile): prove the throwing timeout sink was actually reached
2026-09-07 13:52:15 -04:00
Jinwoo Hong ceafdcad2f perf(mobile): race the direct and relay dials from t=0 on every reconnect (#19308)
* perf(mobile): race the direct and relay dials from t=0 on every reconnect

A foreground reconnect gave the direct dial a fixed 2.5s head start, and while
that dial sat in 'connecting'/'handshaking' the supervisor refused to open a
relay socket at all. A phone that is off the LAN paid the full head start on
every reconnect and got nothing for it, and a phone whose relay dropped could
only return to the LAN through three hysteresis probes.

Both dials now start together and the first authenticated socket is adopted
through the existing migrateTo cutover. Nothing about the migration machinery
changes: only who is allowed to start a dial.

- The relay dial now yields to a live session and to nothing else. An unfinished
  direct dial is progress on the other runner, not a reason to stand still.
- The direct return probe grows a second adoption policy. Against a live relay
  hysteresis still has to prove direct stable; during a reconnect there is no
  session to protect, so an authenticated direct socket wins outright. probeNow
  pre-empts a pending 15s tick so that dial starts with the relay dial, not
  after it, and the dial itself no longer waits for the operation mutex — a
  relay dial holding it is exactly the case the race exists for.
- A loser closes and books nothing. The relay dial withdraws inside migrateTo
  and returns 'aborted', so no backoff is booked against it; a direct socket
  that loses leaves the promotion streak untouched. Only a reconnect that both
  paths lose books a failure, once, on the relay cadence.

Kept: the 30s background grace and the foreground gate, because a backgrounded
phone must not open a billed relay splice; the shared failure cooldown, because
a genuine relay failure still has to be paced; the hysteresis dwell after a
migration, because it is what stops a marginal LAN flapping a healthy session.

The accepted cost is one relay socket per reconnect for a phone that is on its
LAN. It closes as soon as the direct path authenticates, before the resume
confirm, because migrateTo only checks the abort predicate after E2EE auth.

Tests that encoded the removed rules:
- 'fails over when the direct retry loop publishes reconnecting' asserted no
  relay dial while direct was handshaking. The failover now precedes the direct
  client giving up, so it asserts the dial instead of its absence.
- 'does not spend a queued relay retry while direct authentication is
  progressing' encoded the block outright; it now asserts the retry runs on the
  failure cadence while a handshake drags on.
- The four grace-race cases move to mobile-endpoint-reconnect-race.test.ts as
  t=0, direct-wins, background/resume and both-lose cases.
- Five relay-bookkeeping cases now state their premise with unreachableDirect.
  They describe a phone with no LAN, which used to be implicit and is now
  load-bearing: with a reachable LAN the direct socket wins those reconnects.

* fix(mobile): withdraw a lost relay dial pre-handshake and damp blip races

Review follow-up to 4e31130471. Racing both paths from t=0 was correct but
charged the LAN case twice: once per reconnect in cell work, and again whenever
the LAN flapped.

Withdraw before the handshake. migrateTo only consults its abort predicate after
E2EE authentication, so a dial that had already lost still made the cell reserve
a splice and the desktop finish a key exchange. The establisher now watches the
logical client across the dial and closes the cell socket the moment direct
authenticates. In the common window, after relay-auth is on the wire and before
the hello lands, nothing of the key exchange has started, so the withdrawal
costs the desktop nothing. The dial still reports itself aborted and still books
nothing. The watch is dropped once migrateTo returns, because past the cutover
this session is the active path and a later direct promotion must not read as a
reason to close the client's own socket.

Damp races that a blip started. relayDialAllowed yields only to a live session
and a lost race books nothing, so a flapping LAN drove one cell socket per blip
with only the relay's per-host rate limiter as a backstop, and reaching that
limiter would have converted a benign race into a booked relay failure. After a
race is lost to direct, the next unforced race is suppressed for 2s, doubling
per consecutive loss to a 30s cap. This is not backoff and is kept separate from
it: a forced replacement is never damped, a relay dial that wins clears the
streak, and a foreground resume clears it too, so the path the user is watching
never waits. The window arms its own lapse timer, so a LAN that dies inside the
window still reaches relay without a new trigger.

A superseded cutover no longer escapes probe() as an unhandled rejection. Only
the probe timer calls it, and it discards the promise, so the routine end of a
lost race would have surfaced as one.

Credential rotation moves to MobileRelayCredentialRefresh. The supervisor
crossed the 300-line cap; rotation is a self-contained responsibility that only
runs over a live direct connection, so it splits cleanly instead of taking a
max-lines bump.

* fix(mobile): end a damper window as soon as the direct path is really gone

Round-2 review follow-up to 9a21da4326. The damper armed its window when direct
won the race, and nothing shortened it. A LAN that died inside that window left
the phone waiting out the whole thing, up to 30s at the cap, with only a log
line to show for it. My previous commit body claimed the path the user watches
never waits; that was true only of a foreground resume, and it is corrected
here.

Losing the direct path now collapses the wait to a 250ms floor, so the next
recovery races almost at once. The floor is not zero because the reason the
damper exists is a LAN that drops and comes straight back, and a disconnect is
how such a blip begins. So the rest of the window is kept aside rather than
spent: if direct returns inside the floor it was a blip and the window resumes,
and if the floor lapses with direct still gone it was an outage and the held
window is void. Without the second half, one blip would have bought a flapping
LAN a free pass on every race that followed, which is the case the damper was
added for.

The streak itself is untouched by the clamp. A LAN that flaps all afternoon
still escalates toward the cap; only the current wait is cut short.

record() now takes the same forceReplacement guard as suppresses(), so a forced
replacement that stands down cannot grow the streak or be read as a loss to
direct. A lease rotation or a reconsidered network change is not a LAN that
flapped.

Also documents that a genuine relay failure deliberately does not reset the
streak, and that the damper and the failure backoff serialize rather than stack:
a damped attempt never reaches the dial that would book a cooldown.

* fix(test): give the direct-probe fixture the race-era hooks

The phase-1 probe test predates canDial and adoptsOutright, so its hooks
literal threw at the first dial. These cases model a live relay session.

* docs(mobile): say why a finished credential refresh races relay instead of waiting on direct
2026-09-07 13:45:51 -04:00
Jinwoo Hong 5857357fcf feat(relay): log the region probe and name the assigned cell (#19307)
* feat(relay): log the region probe and name the assigned cell

A desktop silently pinned itself to a far relay region for a day and every
phone connect paid the round trip. Nothing in the desktop logs said which
regions were probed, what they measured, why one was rejected, or which cell
the host landed on, so the only way to diagnose it was a bench harness.

The resolver now emits one line per outcome. A refresh carries every region's
probe origins, the discarded warm-up, the kept samples, the minimum, the
spread, and a verdict, then the chosen region or no-hint with the reason it
withheld one. Cache hits, diagnostic overrides, and a director that cannot
list its regions each get their own line so a quiet run is never ambiguous.
Self-heal logs the cached region, the best measured region, the assigned
cell's round trip, and whether it kept or deleted the cache. Only a refresh
reports a catalog failure; a self-heal never chose a region, so a line saying
it withheld a hint would be a lie.

Relay status now carries the assigned cell so the pairing panel can name it.
The field is optional because an offline host holds no assignment and the web
client answers from a stub that never has one.

Splitting catalog fetching out of the preference module keeps both files
inside the line budget without a lint disable.

* fix(relay): drop the assigned cell from statuses not served on it

The origin pool publishes offline while it still holds the assignment it is
about to rotate, so the panel kept naming a cell nothing was served from. The
same class of bug hid a second instance: the coordinator republishes
registered right after the broker announces its cell, and that republish
carried no cell, blanking the value moments after it was set. The cell would
never have reached the panel in the real flow.

Deriving the cell from the status at each publisher removes both. The rule
lives beside the status type because it defines when the optional field is
populated, and the coordinator reads the owned broker's endpoint rather than
trusting a call site to remember to pass it.

* i18n: add the relay cell label to the English catalog

* test(relay): audit the relocated region catalog fetch call site

* fix(relay): report a self-heal whose catalog request failed instead of staying silent
2026-09-07 13:45:48 -04:00
Jinwoo Hong 643571def6 feat(mobile): draw the last known tab strip while a session reconnects (mobile pass) (#19281)
* feat(mobile): draw the last known tab strip while a session reconnects

Reopening a workspace the phone has already visited threw away everything
it knew. The route clears its tabs on mount, so until the reconnect lands
and the first snapshot is applied the session screen has an empty header
and a bare spinner, even though the strip it is about to be handed is the
one it drew a minute ago.

Persist the four fields the strip actually draws -- id, type, title, agent
-- per host and workspace, and add a reconnecting-with-cache shape to the
route state so those rows render immediately, disabled, under the ids the
live snapshot will reuse. Live tabs always outrank the cache, so a
mid-session drop keeps its mounted terminals; an exhausted retry loop or a
rejected pairing outranks it the other way, because a strip the user cannot
reach is worse than the existing offline affordance. With nothing cached
the screen behaves exactly as before.

The body stays a placeholder. Replaying stored scrollback into the terminal
WebView would double-render the same rows once the live stream replays them,
so the strip is the cached content and the body waits for the stream.

* fix(mobile): keep shell titles and unpaired hosts out of the cached tab strip

Review of the reconnect strip cache found two ways it leaked.

A terminal's title is whatever the shell last set, which is routinely the
command line: a psql URL with an inline password, a curl with a bearer
token. Both fit well inside the 64-character cap and both were written to
plaintext AsyncStorage verbatim. Browser tabs carried their page title the
same way. Terminals and browsers now collapse to a fixed label, with a
resolved agent naming itself because that lookup is a closed enum. The rule
lives in the storage module rather than its caller, so it holds for entries
an older build already wrote, and a tab type this build cannot draw is
dropped instead of having its title trusted.

The cache also survived forgetting a host. Nothing expired an entry, and
the module-global memory map meant a later save from any surviving host
serialized the forgotten host's rows straight back to disk. Both cleanup
paths now evict by host, dropping the in-memory rows and rewriting storage,
with a pending debounced write cancelled so it cannot restore them.

Also: the storage key digests the workspace id, which ended in a filesystem
path, and cached rows carry the same de-emphasis as the disabled tab-bar
buttons beside them, so an inert row does not pass for a live one.

* fix(mobile): make a forgotten host's cached tab strip actually leave disk

Review finding on this PR, fixed here so it rides along with the rest.

writeFile swallowed its own rejection, so deleteCachedSessionTabStripForHost
resolved successfully while the unpaired host's plaintext tab titles stayed on
disk, and removeHostAndCloseClient discarded the promise with void so nothing
could have observed the failure anyway.

The write now throws. The debounced save keeps a best-effort catch, since a
dropped cache refresh costs one repaint and the next save rewrites the whole
map, so only the deletion path needs the failure. Host removal awaits the
deletion and logs a failure but never rethrows: the metadata removal has
committed and the client is closed by that point, so reporting a finished
removal as failed would be wrong. The unpaired-host credential sweep already
awaited the deletion and now sees the rejection, consistent with its sibling
credential deletions.

Two ways the rows could come back are closed as well. The cache refuses saves
for a host it has been told to forget, so a snapshot racing the deletion cannot
re-insert it, and the deletion awaits any debounced write already on the wire,
since that write built its blob from the map as it was and would otherwise race
the purge for the last word on disk. The refusal lasts for the process, so
re-pairing the same host caches again from the next app launch, which is the
cheap direction for a deletion the user asked for.

* fix(mobile): order the tab-strip cache writes so a purge is the last word

Two debounced writes could sit on the AsyncStorage bridge at once, and the
second replaced the in-flight handle. A host purge then awaited only the newer
write, so the older blob -- snapshotted while the forgotten host was still in
the map -- could commit after it and restore the host's titles to disk. Writes
now queue behind one chain and the purge queues last.

The unpaired-credential sweep also aborted on a cache-purge failure, stranding
the write revision and onDeleted after every credential was already deleted. It
now warns and finishes, as removeHostAndCloseClient already did.
2026-09-07 13:40:35 -04:00