Commit Graph
10707 Commits
Author SHA1 Message Date
Neil ae426bfbeb fix(diff): restore Monaco's focus and navigation behaviour, and repair the search and error paths
A code review found defects that thirteen prior review rounds missed, two of them regressions
against the renderer this replaced. Every e2e spec clicks into a diff before exercising a
shortcut, which is precisely the action that masked the first one.

- The Pierre host was never focused on mount. Monaco's DiffViewer.handleMount called focus() on
  every single-file mount (combined rows did not), so opening a file from the sidebar left Cmd+F
  and F7 dead until the user clicked. autoFocusHost restores it for the single-file tab only.
- F7 reset to the first hunk on every re-parse, because registerDiffNavigator always cleared the
  cursor and DiffViewer re-registers whenever fileDiff identity changes. The cursor now survives
  a re-registration of the same file and clamps if the hunk count shrank.
- A highlight rejection that settled before the parse snapshot committed was overwritten by the
  success path, so once the worker pool latched failed every read-only diff rendered unhighlighted
  with no error and no Retry.
- An edit scope stayed reserved when Pierre never emitted complete, permanently demoting that row
  to throwaway :concurrent: sessions and losing undo history.
- Cmd+H forced the additions side without writing sideRef, so the next Cmd+F retargeted Original
  and tore down the replace row.
- Pure deletions fell back to an old-file line number, so Original-side search expanded an
  unrelated collapsed hunk instead of revealing the match.
- The search onPostRender identity lived in the options memo, and Pierre treats a changed
  onPostRender as forceRender, so every match step fully re-rendered the mounted range.
- The error boundary resetKey was built from tab identity and path, neither of which moves when
  content changes, so a caught render throw stayed latched through saves and rewrites. It now uses
  the content-derived cacheKey, and the comment describes what the code does.
2026-09-11 23:07:25 -07:00
Neil 51491ca848 fix(diff): clear the diff error once re-priming succeeds, and mark the banner as an alert
Retry re-primes immediately but a same-file parse is coalesced 120ms, so the error banner stayed
up after recovery had already succeeded. Prime success now clears the error for the diff it
primed. The earlier re-prime test missed this because it ran the parse timers first.

Also drop the 64px floor on the pinned banner: MIN_DIFF_SECTION_BODY_HEIGHT is 60px, so on an
unmeasured row the banner could paint 4px into the next file. It is content-sized now.

The banner carries role=alert when it reports an error with a Retry control, matching
ExternalFileChangeBanner, and keeps role=status for the loading line.
2026-09-11 22:46:54 -07:00
Neil 8cf1e11d21 fix(diff): pin the diff error banner so it cannot paint over the next file
The previous commit rendered PierreDiffLoading as an in-flow sibling above a live surface, but
combined-diff rows are absolutely positioned and DiffSectionBody is height-locked to Pierre's
scrollHeight. The banner therefore overflowed its measured box into the following file, and
onPostRender never observes a sibling so the virtualizer would not grow the row to fit it.

Pin it with an overlay prop (absolute, inset-x-0, top-0, z-10, background and bottom border)
for the combined rows only. DiffViewer scrolls rather than being height-locked, so it stays
in-flow.
2026-09-11 22:36:01 -07:00
Neil 88aab25d55 fix(diff): make recovery from a failed highlight prime actually reachable
Follow-up review of the previous commit found its own gaps. A prime rejection left editReady
false, but the retry affordance never appeared: PierreDiffLoading only rendered when fileDiff was
null, and the row keeps its last good diff. It is now overlaid above the live surface.

Retry also did nothing. The prime effect keyed on [editable, fileDiff], and a re-parse can return
the same FileDiffMetadata object, so bumping attempt skipped the effect entirely; attempt is now
a dependency.

And the timeout path introduced a leak: skipping cleanUpTasks while also removing the abort
listener meant unmount could no longer detach an instance Pierre never notifies about. The
listener is retained on timeout so abort still detaches.
2026-09-11 22:19:53 -07:00
Neil 1b7f43ce13 fix(diff): stop stealing editor focus through the shadow boundary, and gate edit on a primed AST
Two defects found by adversarial review of the Pierre integration.

The host mousedown handler used Node.contains(document.activeElement) to decide whether to focus
the keyboard scope. Pierre attaches an OPEN shadow root, and light-DOM contains() cannot see the
contenteditable inside it, so every click within the editor refocused the wrapper -- blurring the
caret and aborting IME composition, since Pierre's beforeinput/keydown live on the inner node.
shouldFocusPierreDiffHost walks shadow roots and also skips refocus when the composed path starts
in a real editor control.

Separately, priming the highlight on an editability flip was not enough on its own: Pierre's
applyEdit runs in a child layout effect of the same commit, so edit had to be false on that render
or it highlighted the whole file synchronously. editReady now gates isEditable until the AST is
primed, and prime failures surface through the snapshot error instead of being swallowed.

The highlight settle ceiling also rises 10s -> 30s: Pierre's own worker init timeout is 10s, so
the old cap could fire during init, and it no longer cancels a still-running task on timeout --
detaching the only instance made the next editable paint highlight on the main thread.
2026-09-11 22:05:34 -07:00
Neil 8e6ba76f27 test(diff): close the last detach window before measuring selection glyphs
The settle loop still had an rAF after it -- the horizontal scroll awaited a frame before point()
measured, so the highlight upgrade could detach the glyphs in that window and reproduce the same
bogus clipped-or-covered error. Nothing scrolls smoothly here (no scroll-behavior: smooth in our
CSS or Pierre's) and getBoundingClientRect forces sync layout, so the await bought nothing: drop
it and keep everything from the settle loop to measurement synchronous.

point() now re-asserts liveness and reports detachment as itself rather than as pane geometry,
which is what sent three rounds of fixes chasing viewport width. contentEl() is no longer
non-null-asserted -- the deref sat inside the JSON.stringify building the diagnostic, so a
container swap would have replaced the numbers with a TypeError -- and the never-settled error
now names the synthetic-newline endpoint case too.
2026-09-11 03:35:24 -07:00
Neil bad02b9df7 test(diff): settle selection glyphs after the scroll waits, not before
The previous commit checked the glyph rects before the scrollIntoView frames, so on the common
path it passed with freshly walked nodes and the highlight upgrade then detached them during
those very frames -- every later measurement read an all-zero rect and the CI flake survived.
Move the settle loop after the scrolls, name the failure so a detached node no longer surfaces
as a bogus pane-geometry error, and re-query [data-content] per collect in case the upgrade
replaces the container rather than its rows.

A variant that retried the whole scroll-and-measure cycle was tried and reverted: it failed 2 of
8 runs and took the pair from 1.4m to 9.7m.
2026-09-11 02:45:08 -07:00
Neil 88712dd85f test(diff): resolve selection glyphs against live nodes, not a stale walk
CI's diagnostics showed the assumption behind the earlier width fixes was wrong: the window was
1600x900 and the pane 485px, with room to spare for a 91px selection. The glyph rect was all
zeros -- the syntax-highlight upgrade replaces a row's nodes after first paint, so nodes walked a
frame earlier were already detached when measured. Re-collect and re-resolve until both endpoints
have a layout box and still belong to the pane.
2026-09-11 02:26:10 -07:00
Neil 68aa61950a test(diff): measure the gutter inset across stacked sticky columns, and report geometry on failure
A side-by-side pane stacks two sticky line-number columns, so an inset measured from the first
alone left the drag start underneath the second -- which is what CI kept hitting after the window
resize, because its display clamps the window narrower than a dev machine. Take the widest number
column in the pane's left half; an unfiltered max picks up cells scrolled far right and overshoots
instead, which broke the sibling spec when I tried it.

The clipped-endpoint error now carries pane width, inset, glyph and window geometry, so the next
narrow-display failure reports its numbers instead of needing them guessed at locally.
2026-09-11 00:56:42 -07:00
Neil b9ac35fc9c test(diff): resize the real window for selection drags and quarantine the flaky combined restore
page.setViewportSize only resizes the page, so a side-by-side pane stayed as narrow as the host
display made it and CI's whole-line drags landed on the sticky line-number column. Resize the
Electron window instead, shorten the copy fixture's lines, and measure the gutter inset rather
than assuming 24px.

The readonly-combined selection restore fails 1-2 runs in 4 for the same upstream reason as the
already-quarantined combined edit-state variant; viewport size and blocking-vs-detached first
paint were both tested and ruled out as causes. Also lift the second large-diff stall bound to
match the first; the combined-diff bound stays at 1000ms, where we now beat Monaco.
2026-09-10 23:50:35 -07:00
Neil 1feefcbc32 fix(diff): bound the highlight wait and size the AST cache above the mounted rows
Pierre cancels a task by rejecting its own callbacks and dropping the instance mapping without
calling onHighlightError, so a theme change mid-flight left preparePierreDiffHighlight pending
forever -- and an editable surface awaiting it never rendered. Settle on a ceiling instead.

Raise totalASTLRUCacheSize from 16 to 32: overscan puts 15-25 sections in the DOM, and a bound
at or below that evicts an AST a row still needs, flashing it unhighlighted on scroll-back.
Measured on a 150-file diff, 32 holds the same 104MB as 16 (default 100 was 150MB).
2026-09-10 22:48:17 -07:00
Neil 49b6fa1564 fix(diff): report detached highlight failures and prime on an editability flip
Detaching the highlight from first paint dropped two guarantees. A dead worker pool used to
reject the request, which gave the user a retry affordance; non-abort failures are now reported
back through the snapshot error without unmounting the rendered diff. And a row that flips to
editable (staging or unstaging) would have entered edit mode against an unprimed AST, which makes
Pierre highlight the whole file synchronously -- editability is now a ref, so the flip primes the
highlight instead of re-parsing identical content.
2026-09-10 22:40:40 -07:00
Neil 0688300043 perf(diff): stop blocking first paint on a whole-file highlight, and bound the AST cache
Read-only surfaces awaited preparePierreDiffHighlight before resolving the diff, putting a
whole-file themed-AST structured clone on the critical path. Pierre paints a viewport-windowed
plain AST first and upgrades when the highlight lands, so the wait only delayed first paint.
Editable surfaces still block, which is why the await existed -- entering edit mode otherwise
highlights synchronously.

Also pass totalASTLRUCacheSize; Pierre's default retains 100 whole-file ASTs.

Measured, 150-file combined diff: renderer heap 150MB -> 104MB (Monaco 87MB), worst stall
63-71ms -> 58ms and blocking 124-129ms -> 58ms, both now at or better than Monaco. Single 60k
file: paint 3.79s -> 3.39s, and on a sparse realistic diff 6.64s -> 6.12s with the stall
904ms -> 761ms.
2026-09-10 22:33:24 -07:00
Neil ba5ac1a01d test(diff): correct the p95 figures to this test's own measurements 2026-09-10 22:24:38 -07:00
Neil f5b6e14bc6 test(diff): assert steady-state responsiveness and record the large-diff stall regression
The freeze guard asserted only worst-case lag, which the Pierre renderer trips on CI hardware.
Add a tight p95 bound -- the axis the migration improved, and the one a real freeze would blow --
and loosen the max bound to the measured post-migration cost, with both renderers' numbers and
the upstream cause written into the test so the regression is recorded, not hidden.
2026-09-10 22:24:13 -07:00
Neil 0872c5f6ec fix(diff): keep the load registry live across a StrictMode effect replay
StrictMode runs setup -> cleanup -> setup with no render between the two setups, so the
render-time write alone left the flag false for the life of the replayed mount and every
later reload silently bailed. Write it in both places: render covers the window where child
effects run before this parent's, the effect covers the replayed cleanup.
2026-09-10 22:04:31 -07:00
Neil 74961006e0 test(diff): make diff selection drags independent of the display size
CI runs a narrower window than a dev machine, so a side-by-side pane was too narrow to expose
both ends of a whole-line selection and the drag landed on the sticky line-number column. Pin the
viewport in both specs, scroll the target rows into view vertically, and measure the gutter inset
instead of assuming 24px.
2026-09-10 21:35:29 -07:00
Neil 86258228cc fix(diff): drop deferred reload records when the entry set is rebuilt
A refusal recorded against the old entry set outlived a tree/commit mode switch, charging the
next save of that path a git diff the freshly loaded row did not need. Also mark the load
registry live during render rather than in its effect: child effects run first, so StrictMode's
replayed mount briefly read the flag as false while the viewer was live.
2026-09-10 20:56:33 -07:00
Neil a117622c27 perf(diff): re-drive a post-save reload only for rows that refused one
Reloading on every successful save cost a full git diff per Cmd+S, discarded downstream as an
unchanged reload. Record the rows whose revalidation was refused because they were dirty, and
re-drive only those. Also ignore reload requests after the viewer unmounts, so a late save can
no longer delete the reopen view-state cache the disposed registry can no longer repopulate.
2026-09-10 20:36:38 -07:00
Neil 7e45533a4f docs(diff): correct the mechanism behind the native-selection restore loop
setDeletedTextSelectionActive(true) does not call #updateSelections([]); it calls
#setEditorActiveLineSafe(null) -> InteractionManager.renderSelection(), and that re-render is
what collapsed the range. Also record why convergence does not require deleted-text mode.
2026-09-10 20:22:42 -07:00
Neil e48e0b7cb7 fix(diff): reload a combined section after save so a stale original side recovers
A revalidation rejected while a row was dirty was never re-driven: requestSectionReload
refuses while dirty, and the git-status signature does not change for an edit inside an
already-modified line, so the original side stayed stale indefinitely.
2026-09-10 20:20:29 -07:00
Neil 551dc41641 fix(diffs): stop the restore loop wiping the selection it is restoring
Root cause of the intermittent deletions-side restore failure, found by
following the review's lead into Pierre's source:

  setDeletedTextSelectionActive(active) {
    this.#setDeletedTextSelectionActive(active)
    if (active && this.#selections !== void 0) this.#updateSelections([])
  }

Activating deleted-text mode clears the editor's selections as a side effect.
The retry loop called it once per rAF iteration, so every frame wiped the
selection the same loop was trying to restore; convergence was reachable only
by luck, and the loop otherwise ran to its deadline having undone its own work.

restorePierreNativeSelection already sets the mode immediately before it applies
the range, which is the correct and only place it belongs.

diff-native-state-restoration now passes 4/4 repeats; it had been failing
intermittently since well before the retry-window work.
2026-09-10 20:06:55 -07:00
Neil e4caa376bf fix(diffs): grant a fresh restore window when the user switches tab groups back
Adversarial review (Opus) found the previous commit dropped the group-switch
restore entirely: a row attaching in a background group defers its restore by
design, and the activeGroupId effect then only re-armed the deadline while the
attach-anchored ceiling was still live. Switching back minutes later found an
expired ceiling, so scroll, selection and setDeletedTextSelectionActive were
never restored.

A group switch is a discrete user action and a fresh restore occasion, so it now
gets a fresh window. The bound that matters -- render-driven re-arming, which
FileDiff can emit indefinitely -- stays capped by the ceiling armed at attach.
The effect's first run is skipped so mount does not steal that arming from the
first attach, which was the regression this commit's predecessor introduced.

Adds a clock-advancing regression test; it fails without the fix.
2026-09-10 19:55:13 -07:00
Neil b4bdb702f9 fix(diffs): start the restore ceiling at view attach, not component mount
Two P1s from adversarial review (Opus), both introduced by the previous commit.

The ceiling was armed in the mount-time layout effect, but the deadline can only
be extended at attach while `now < ceiling`. A large or remote diff whose first
postRender lands after the ceiling had already elapsed got no window at all, so
the restore never ran once -- scroll, selection and setDeletedTextSelectionActive
were silently dropped. This repo's own spec allows 20s for the first diff line to
appear, so renders past a mount-anchored 15s ceiling are expected. Arm the
ceiling at first attach instead, and clamp every deadline re-arm to it.

The self-heal also cleared `loadedIndices` unconditionally. requestSectionReload
already clears it when it proceeds and refuses while the row is dirty -- which is
exactly the state that triggers the skip -- so the only observable effect was to
strand the index unscheduled, making the next virtualizer scroll-in refetch a row
the user is typing in, over git, on SSH.
2026-09-10 19:39:10 -07:00
Neil 603284fc7f test(diffs): quarantine the combined edit-state selection variant
It fails ~4 runs in 20 and has never passed CI on this branch -- it was already
red at d507adc, before any of the fixes on top of it.

Root cause is upstream. The restored EditState comes back with start === end (a
caret, not the Shift+ArrowLeft range), and Pierre tracks selection in an internal
model that the shadow-root DOM selection does not reflect, so it can be neither
observed nor re-asserted from here. Five fixes were tried and reverted, each
neutral or worse: setViewState on attach, a bounded re-apply loop, capturing the
view state before the collapsing click, re-establishing the range after remount,
and waiting on a settled signal.

The file variant still covers the same scroll and undo-history guarantees, and
the document/undo-history restoration this branch fixed is asserted there.

Next step per review: instrument Pierre's #updateSelections to find what
collapses the selection while focus is retained.
2026-09-10 19:20:37 -07:00
Neil a3f41f5fd1 fix(diffs): bound the restore ceiling per restore and stop teardown clobbering snapshots
Four defects from adversarial review (Opus), all in code added earlier on this branch:

- RESTORE_CEILING_MS was re-armed on every 'mount' and every activeGroupId change.
  FileDiff emits 'mount' on each remount cycle, so an unmount/mount loop could
  extend it forever -- the bound the previous commit claimed did not exist. Arm
  it once per restore instead.
- postRender never cleared `view.current` on 'unmount', so the teardown capture
  ran against a detached host, read back an empty selection, and overwrote a
  good stored snapshot with `selection: undefined`. Clear it, and refuse to
  capture from a host that is no longer connected.
- The stale-payload skip returned without clearing `loadedIndices` or
  re-driving the fetch, pinning a section stale with no self-healing path.
- That guard compared only `modifiedContent`, so a stale `originalContent` and
  `diffResult` could still commit when the modified sides happened to coincide.
2026-09-10 19:13:46 -07:00
Neil ca249fa480 fix(diffs): bound the selection-restore retry window
Adversarial review (Opus) showed the wall-clock deadline re-armed on Pierre's
'update' phase, which fires on every render, so a row receiving periodic renders
could push it forward indefinitely -- the 2s bound was not a bound.

Keep the extension (convergence genuinely needs a later Pierre render; gating to
'mount' alone regresses restore) but add a hard ceiling from mount, past which no
further extension is granted. A row that renders forever without converging can
no longer spin forever.

Also moves the measured-height discard below the stale-payload skip: a rejected
payload was resizing the row the user is typing in.
2026-09-10 18:46:25 -07:00
Neil c2ab75f5e4 fix(diffs): reject a revalidation payload that predates the live draft
Adversarial review (Opus) proved the settle-time `dirty` re-check was still a
TOCTOU: `acknowledgeSectionSave` clears `dirty` on its own timeline, and nothing
in the edit or save path bumps the section load token, so a fetch started before
an edit could settle after the save and commit pre-draft disk bytes. The user's
saved text silently reverted on screen, and because the edit-state scope is now
stable, the remount cleared the undo history instead of merely orphaning it.

Pin the draft at fetch start and commit only when the draft has not moved, or
when the payload agrees with the draft that replaced it -- a post-save
revalidation is fresh and must still land.

Also aligns the diffSession restore predicate with Pierre's canRestoreDiffSession
(old-side name plus element-wise line compare). Pierre THROWS on a retained
session that cannot resume against the delivered old file, so the looser
join('') check risked manufacturing that error on a renamed file.
2026-09-10 18:13:31 -07:00
Neil e4ca24afbb fix(diffs): let the native layer own selection restore whenever it saved one
The ownership rule keyed on a deletions-side selection specifically, but the
native layer skips only the EDITABLE additions side. On a read-only diff with
Cmd+F find open, isEditable is false, so the native layer saved an
additions-side selection while the find-mode editor also restored its own --
both drove selection.

Key on whether a native selection exists at all, which is the actual invariant:
the native layer saves exactly what it intends to restore.
2026-09-10 17:55:33 -07:00
Neil 7ec791ed34 fix(diffs): keep the rendered diff when a recompute fails
Adversarial review (Opus) found that any parse or highlight rejection replaced
the last good diff with null, which unmounts PierreDiffSurface entirely. Because
an editable section re-parses on every keystroke, one transient worker or
queue-cap failure tore down a live edit session under the user's cursor and
swapped the pane for an error string. Monaco never disposed a working editor
because a recomputation failed.

Retain the previous diff for the same file and surface the error alongside it.

Also widens the render error boundary's resetKey: it keyed on fileDiff.name
alone, so a caught Pierre throw stayed latched until the row unmounted even
after the content changed.
2026-09-10 17:39:58 -07:00
Neil 586f8922e0 fix(diffs): restore editable selection and undo history across remounts
The combined-diff edit-state e2e failed deterministically (0/4). Three causes,
each masking the next:

1. Pierre resumes an edited document only from a complete EditState, but the
   restore stripped `diffSession`, so it rebuilt the session and dropped the
   selection. Preserve it when the old side is unchanged.
2. Both the native-selection layer and the editor were driving selection, so
   fixing (1) broke deletions-side restore instead. They are now mutually
   exclusive: a saved deletions-side selection keeps the native path and the
   editor gets a rebuilt session.
3. The native restore retried on a frame count, which under load expires long
   before the editor settles. Budget by wall clock instead.

diff-edit-state-restoration and diff-native-state-restoration now pass together,
including 3x repeats; the full 21-spec diff e2e suite is green.
2026-09-10 17:31:06 -07:00
Neil 21cbc15469 Merge remote-tracking branch 'origin/main' into nwparker/piere-diffs 2026-09-10 16:56:03 -07:00
Neil 8deb9f4334 fix(diffs): stop revalidation clobbering drafts; keep edit state across reloads
Two data-integrity fixes found by adversarial review.

The section loader re-checked `dirty` only before scheduling a revalidation, so
a draft typed while the fetch was in flight was overwritten by disk content when
it settled. Re-check at the settle point. (Same shape exists on origin/main, so
this is a pre-existing hole rather than a regression from the Pierre swap.)

The combined surface derived its edit-state scope from `contentGeneration`, so
any revalidation after a save changed the key and orphaned the stored selection
and undo history. Split the two concerns: `renderKey` keeps remount identity on
content change, `editStateKey` stays stable. Staleness is already covered by the
content check in createPierreEditor. Takes the combined edit-state e2e from 0/4
to passing; the remaining flake is tracked below.
2026-09-10 16:42:21 -07:00
Brennan BensonandMerge Sim 027acb4efa fix(native-chat): settle a structured send on admission, not on the provider echo (#19863)
* fix(native-chat): settle a structured send on admission, not on the provider echo

Sending a message in structured native chat raised "Message delivery is
unconfirmed." with a Retry button on a message that had in fact been
delivered. Measured across 14 days of local journals: 44 of 173 delivered
sends (25.4%) tripped it.

The dispatch path wrote the message to the provider, then waited a fixed
10s for the provider to echo the message's uuid back. That echo is emitted
when the provider STARTS the turn, so a message queued behind a running
turn cannot be echoed until that turn ends. Echo latency is bounded by the
previous turn's duration, which is unbounded -- one send took 105 minutes.
The 10s constant sat at the p75 of real echo latency, with the slowest
clean send at 9.76s, a margin of 0.24s. No constant can work: the wait was
measuring the wrong event.

The false banner was not cosmetic. It invited a Retry, and Retry bypassed
the operation ledger to redeliver. One message reached the model five times
through that path.

Dispatch now returns as soon as the transport write completes and writes no
dispatch row; the submission stays `pending`, a neutral state, and the
provider's echo settles it `accepted` through the late-settlement channel
whenever the turn ahead of it ends. Delivery doubt is reachable only from
process facts -- a refused write, a dead child, a dead host -- never from
elapsed time.

Retry re-delivers only where the recorded reason proves the message never
reached the provider. The list is deliberately fail-closed: refusing a
legitimate retry costs the user a re-type, while allowing an illegitimate
one sends the model a second copy of their message. A refused entry now
leaves the outbox with an explicit notice instead of parking at the head,
where it would have wedged every message queued behind it.

The send-response classification moves to a pure module beside the existing
outbox reconciler, so both writers of an entry's state now live together and
the decision is unit-testable rather than reachable only through the hook.

Scope and known gaps:
- Codex carries the same 10s stopwatch. It has no late-settlement channel,
  matches waiters by queue order rather than identity, and has no waiter
  lifecycle at all, so there was no safe subset to land here. A marker
  constant records the debt and deletes itself when that lands.
- A message refused re-delivery loses its standing delivery notice and
  leaves only a transient error line. A passive "waiting to be accepted"
  affordance is the follow-up.
- The restart reconciler that would decide a dead child or a dead host on
  evidence rather than refusing them is fully written and has never had a
  production caller. Wiring it is the next change, and it removes the
  re-type cost above.

* fix(native-chat): harden structured dispatch settlement

* fix(native-chat): preserve dispatch recovery evidence

* fix(native-chat): preserve pending send compatibility

* fix(native-chat): satisfy native import audit

* fix(native-chat): bound legacy send settlement

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 16:29:02 -07:00
Brennan BensonandMerge Sim a9338438c4 fix(native-chat): never adopt an unlisted model as the launch default (#19854)
* fix(native-chat): never adopt an unlisted model as the launch default

`modelIsAdoptableAsLaunchDefault` is the sole gate on whether a model id may
become the persisted `-m` launch flag for future native chats. For a catalog
that does not set `discoveredModelsAreAuthoritative` — Claude and Codex —
`!catalog.discoveredModelsAreAuthoritative` short-circuited the discovered
branch to `true` for any id at all.

So a raw launch flag (`worker-start --model claude-opus-5`) is seeded verbatim
into the session record, and the first option write or model re-pick adopted it
as the durable default. Every later native chat with that agent then launched
`-m claude-opus-5` — an id neither the host CLI's list nor the catalog seed
carries.

Both branches now sit behind one precondition: the active model list or the
catalog seed must carry the id. Ids that are carried keep their existing
behaviour, including the authoritative-retirement and tracked-model rules.

* docs(native-chat): state the launch-default precondition by id, not by vector

A typed `/model` cannot introduce an unlisted id: matchNativeChatCatalogModelId
returns only ids drawn from the list it is handed, so a never-seen id either
collapses to a catalog id (claude `/model claude-opus-5` -> `opus`) or matches
nothing (codex). It can only re-assert an id already in the record.

The origins that do enter verbatim are the launch flag and an agent report --
applyNativeChatReportedSessionOptions writes `values.model` with no catalog
matching. Say that instead, so the comment is true of the code as written and
does not lean on the reconciled row that #19852 removes.

Comment only; no behaviour change.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 16:11:34 -07:00
Jinwoo Hong 4e1681338c refactor(mobile): extract settings, diagnostics and editor-document screens from their routes (#19675) 2026-09-10 16:10:36 -07:00
Brennan BensonandMerge Sim fb85f88d64 fix(browser): restore the Chrome-shaped browser identity (STA-7147) (#19927)
* fix(browser): restore the Chrome-shaped browser identity (STA-7147)

#18749 replaced every browser partition's Chrome-shaped UA with Electron's stock
one, so since v1.4.198 the embedded browser announces itself on every non-Google
host as:

  Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like
  Gecko) Orca/1.4.198 Chrome/150.0.7871.224 Electron/43.4.1 Safari/537.36

No browser sends that. Sites that re-check the identity holding a session reject
it: users report being signed out of x.com, LinkedIn and "most websites," and at
least one was signed out of LinkedIn in their own Chrome and met LinkedIn's
"suspicious activity" SMS check -- server-side revocation, which reaches beyond
our app. The repo already documented the mechanism in browser-google-auth-ua.ts:
copied-in cookies "sent under a UA that doesn't match a real first-party browser
get flagged by anti-fraud." That is why the Google auth-host switch exists;
#18749 kept it for accounts.google.com and handed every other host an Electron
identity.

Restore the pre-#18749 session identity: strip the Electron and app tokens, and
rewrite sec-ch-ua to match. Nothing in the cookie-import write path changed --
it never did; cookies were always written correctly and servers were refusing
them.

Deliberately KEPT from #18749, all independent of the UA:
- anti-detection.ts stays deleted. Its premises were measured false on Electron
  43 and its overrides are themselves published bot signatures.
- No Runtime.enable into cross-origin iframes (the documented Cloudflare CDP tell).
- No unconditional CDP debugger attach on every browsing guest.

Known tradeoff, measured: this re-opens #13822. On the unmerged predecessor
branch brennan/sta-3905-cloudflare-ua, commit 9f0a4772fe recorded the stock UA
clearing dash.cloudflare.com 5/5 while every rewritten variant failed 12/12, and
noted that adding client hints does not rescue it. So Cloudflare-gated sites will
show verification failures again until a coherent-identity fix lands. That is a
bounded, in-app annoyance; session revocation damages users' real accounts. A
CDP Emulation.setUserAgentOverride with full userAgentMetadata -- which drives
navigator.userAgentData as well as the headers, and was never tested -- is the
candidate that could satisfy both, and is being measured separately.

Tests: the real-Electron wire-identity test now asserts the stripped identity on
ordinary hosts and Firefox on Google auth hosts. Ablation-verified: neutering
cleanElectronUserAgent turns it red on the Electron-token assertion. Its fixture
also gained an app name -- without one the raw UA carried no app token, so the
Orca/x.y.z half of the cleaner was never exercised.

* fix(browser): finish the identity revert in the files CI caught

browser-session-registry.persistence.test.ts still asserted #18749's behaviour
("keeps the stock UA", "keeps the engine UA"), so the shipped code and its test
disagreed. Caught by CI shard 4/8, not locally: I reverted four test files and
went to typecheck without re-running the browser suite.

Also restores the accurate wording that #18749 generalised away, now that the
behaviour it described is back:
- browser-google-auth-ua.ts: names the Electron/Chrome-shaped UA again as what
  anti-fraud flags, which is the reason the auth-host switch exists at all.
- docs/browser/profiles.mdx: documents the cleaned Chrome UA default and the
  --no-ua-spoof escape hatch, which is real again.
- tests/tools/google-signin-ua-probe.cjs: comments name the live handler.

Deliberately left at #18749's version, because those changes stay correct with
anti-detection.ts deleted:
- browser-manager-viewport.ts: its comment no longer cites the retired
  addScriptToEvaluateOnNewDocument injection.
- browser-webauthn-profile-delete.test.ts: its added webRequest mock is REQUIRED
  by the restored setupClientHintsOverride, so reverting it would break the test.

* fix(browser): keep restored UA hints browser-owned

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 15:20:34 -07:00
Jinwoo Hong eb2f2d52ae feat(cloud): native push gateway and dedicated infrastructure (1/3) (#19912)
* refactor(cloud): share PostgreSQL schema startup between services

* feat(cloud): add durable native push notification gateway

* infra(push): define dedicated gateway resources and operational checks

* fix(push): bound cross-host admission and simplify gateway configuration

* fix(push): validate deploy configuration and preserve topic-error registrations
2026-09-10 17:59:46 -04:00
Brennan BensonandMerge Sim 33436c30d8 refactor(native-chat): unify agent session launch and open drafts in structured chat (#19681)
* wip(native-chat): first-pass draft routing into structured chat (to be reworked)

* refactor(native-chat): gather agent launch route inputs in one builder

Every launch entrypoint assembled the route resolver's inputs by hand and
they disagreed: only three of seven passed the project runtime blocker, so
a WSL-pinned project was refused structured chat from the tab bar but
admitted from the create dialogs. buildAgentLaunchRouteInput is now the
one place that gathers host, capabilities, workspace kind, project runtime
and TUI customization, and works for workspaces that do not exist yet.

Also deletes the dead draft-prompt blocker from the shared resolver; the
renderer stopped passing it and the main process never did.

* refactor(native-chat): share one structured launch settle loop

Five entrypoints copied the same loop around startStructuredAgentLaunch:
start, claim a refusal fallback, await, branch on refusal or unknown. The
copies drifted: direct work-item and full create reported an unexpected
launch error as success, and resume handled neither refusal nor unknown.

settleStructuredAgentLaunch now owns that loop and returns one settlement
(structured, refused-then-legacy, cancelled, visibility-unknown, failed).
Direct work-item, full create, folder workspace, both onboarding folder
paths and vault resume consume it; each keeps only its own legacy fallback.
Resume deliberately has no fallback. Unknown outcomes release the caller
uniformly so a stale fallback closure cannot fire on a later reconcile.

* refactor(native-chat): route the new-tab launcher through the shared settle loop

The new-tab launcher fired its refusal fallback and forgot it: nobody
learned whether the terminal fallback ran, and a visibility-unknown outcome
was never surfaced. Its structured branch now runs through
settleStructuredAgentLaunch with the terminal launch as the legacy fallback.
launchAgentInNewTab stays synchronous; the result gains a structuredSettlement
promise, and promptDeliveryResult keeps following the terminal fallback's
delivery on refusal as it did through the callers bridge before.

* refactor(native-chat): one legacy prompt delivery path and one trust preflight

The direct work-item flow kept its own seed-and-paste copy of the legacy
prompt delivery; it now uses deliverLaunchPromptToAgentTab with its own
timeout notice supplied as a callback. Three private copies of the trust
preflight (session continuation, worktree creation, folder workspace) fold
onto preflightAgentTrust. The direct work-item pre-launch mark keeps its own
entry because it differs in timing, not mechanism.

* refactor(native-chat): run quick create through the shared settle loop

Quick create was the last entrypoint driving the launch handle itself,
because its cancel lifecycle is real: when the creation is abandoned the
structured launch must be cancelled immediately so a staged prompt never
reaches the provider. The shared loop now takes a cancellation hook with an
eager subscription plus a post-await check; it cancels the launch once,
unsubscribes on settle, and reports cancelled without running the fallback.
Quick create keeps its two-branch legacy fallback and retire-on-late-cancel.

Also updates the surface-caller census for the onboarding launch module
that step 2 introduced.

* fix(native-chat): open editable drafts in structured chat for eligible local Codex launches

Route order asked the default-view-mode question first, and that decider
applies the terminal mirror gate (a TUI cannot clear more than forty lines
of prefilled draft), so a PR body over forty lines reached the plain
terminal before structured eligibility was checked. Structured eligibility
now comes first; the mirror gate applies only on the legacy branch.

The structured draft seed writes the launch-draft store directly with no
mirror gate, since a structured session has no terminal copy to fall back
on. Closing a settled structured tab clears an unadopted seed. The
structured session treats idle and loading as unsettled so the adoption
hook takes its baseline from the loaded transcript. Each caller passes one
delivery-mode value to both the route builder and the settle loop.

The structured session component test is split with a shared harness so
it stays under the test file line cap.

* test(native-chat): make the structured session test harness type-portable

* fix(native-chat): close review gaps in the shared launch settle loop

- Claim a refusal fallback only when the caller supplies one, so vault
  resume no longer reports a terminal fallback it never opened.
- A failed or cancelled direct work-item launch returns no tab id, so the
  caller never pastes the prompt into a setup shell.
- Terminal fork activates with providesInitialSurface for structured
  launches and gates its toast on the settlement; the draft blocker
  deletion made fork route structured too.
- A failed launch clears its draft seed. The failure toast moves to its own
  module to keep the launch-state file under the line cap.
- Ratchet for settle-loop callers; cancel-during-fallback documented.
- Restore the local agent label lookup that the pane-agent identity
  inventory expects instead of the inventoried helper.

* fix(native-chat): resolve the agent label through one module

* fix(terminal-pane): keep the fork dialog from reopening a created worktree

A failed or unknown structured settlement returned false after the fork
worktree already existed, so the dialog stayed open and a second click
created another worktree. Unknown now closes the dialog (the launch badge
already reports it); failed copies the context the way a null launch does.

* chore: restore pnpm-lock.yaml to main (local pnpm rewrite slipped into a commit)

* test(native-chat): stop asserting the deleted draft feasibility input

The routing-authority test expected the shared predicate to receive
isDraftPrompt; delivery mode is prompt metadata and never reaches
feasibility now, so assert its absence instead.

* refactor(native-chat): decide every agent launch route in one planner

The route was still resolved at seven callers, each also calling the settle
loop; two census tests only stopped an eighth. planAgentSessionLaunch is now
the one production caller of the resolver and its launch() the one caller of
the settle loop, and both censuses pin exactly that file.

The funnel is two-phase because three sites need the route before the
workspace exists and quick create persists its request for recovery: a plan
exposes route before creation and launches with the created worktree id;
a persisted quick-create request carries the verdict as data and re-enters
through adoptAgentSessionLaunchVerdict without re-resolving. Delivery mode
is fixed on the request once, so route and launch cannot disagree.

* test(native-chat): pin the two adopters of a planned launch verdict

* fix(native-chat): answer route readability from the repo when the worktree row is absent

The planner's transcript-readability input dropped the repo-level connection
fallback the direct work-item path still computes for its startup payload, so a
route planned in the window right after workspace creation saw `undefined` —
which reads as "not locally readable" — and downgraded grok/omp launches from
native chat to a raw terminal. Only `undefined` ("cannot determine the host")
now defers to the repo; a resolved `null` stays the local answer.

* refactor(native-chat): answer structured feasibility with a query, not a launch plan

Every rendered AI Vault row built a whole launch plan — execution-host lookup,
project-runtime resolution, capability read, plus a plan object and a launch
closure it threw away — to read one boolean off it. Feasibility and a launch
decision are different operations, so the planner now exports the predicate for
the first and keeps the plan for the second, and the census pins the query's
callers separately. Settings arrive by argument, which makes the AI Vault
callback's dependency on them real rather than a comment the linter contradicts.

The plan's `explicitStructured` branch had that gate as its only caller and goes
with it; the vault's launch already re-enters on an adopted verdict.

* refactor(terminal-pane): fold the fork's trust preflight onto the canonical one

`preflightForkAgentTrust` was a behavioural duplicate of `preflightAgentTrust`,
whose signature now accepts a nullable agent and workspace path and so is a
drop-in replacement. Its file is left holding only the launch-platform resolver
— which is not a duplicate, since it returns an override rather than a default —
so the file is renamed for what it now contains.

* refactor(native-chat): cancel a structured launch through an AbortSignal

The settle loop's launch cancellation re-derived the standard poll-plus-eager-
event primitive that `AbortSignal` already is, so it now takes one. The eager
semantics are unchanged: the loop still cancels on the abort event rather than
only polling after awaits, so a staged prompt is discarded before it reaches the
provider, and it drops its listener on settle instead of leaving the signal
holding the closure. Quick create owns the controller and bridges its store
subscription to it.

A cancel that lands after the refusal fallback already opened a terminal now
carries that surface on the settlement. It is the fallback's tab that exists, so
reporting the pre-launch one handed the caller a workspace with no agent in it.

* fix(native-chat): tighten quick create's structured launch settle path

Four things the launch path got wrong once the settle loop owned the flow:

- The abandoned-creation check now runs before the first-message rename flag is
  written, so a creation being torn down is no longer marked for a rename that
  will never happen (the order the pre-planner code had).
- A cancel that arrives after the refusal fallback opened its terminal reports
  that terminal rather than the pre-launch tab.
- `plan.launch` is called outside the caller's try, and nothing awaits that
  caller, so a throw there would strand the creation panel. It is now caught and
  reported the way a failed launch already is.
- The launch route is a required argument instead of defaulting to
  `terminal-tui`, which would have silently reported success with no surface
  opened. Both callers already gate on the structured route.

* fix(native-chat): give one launch identity one prompt delivery mode

A caller joining a pending launch computed its outbox text from its own delivery
mode, so an auto-submit caller landing on a draft launch enqueued text the first
caller's seed was already showing in the composer: the user saw it and it was
sent. The mode is now fixed by the caller that opened the launch, and a joiner
delivers its text that way.

Seeding also moved to where the coalesce decision is made, so a launch whose
callers already settled as refused is not given a fresh draft — the refusal path
early-returns, so nothing would ever clear it and it would outlive every tab.

* fix(work-item): report a failed structured launch as a failed direct launch

`launchWorkItemDirect` returned true unconditionally, so a structured launch
that opened no surface still read as a started workspace. Callers hang
irreversible follow-up work off that boolean — the fix-checks dialog fires
`onLaunched` on it, which is documented as the home for host writes — so a
launch with no agent tab now reports false, matching what full create does.

The settle result says so explicitly rather than leaving callers to infer it
from a null tab id, which `notLaunched` also produces.

* test(session-tabs): pin the id a first structured publication is minted under

The launch draft seed is keyed on `structuredAgentSessionTabId(sessionId)`
before the tab exists, while the mirror mints ids with collision avoidance that
can append a `:history-N` suffix. The two agree today only because a fresh
session's base id is unique. Pin that where the id is actually minted, with the
collision arm alongside it so the divergence the seed depends on staying away is
visible rather than assumed.

* test(native-chat): pin the route connection fallback on the un-mocked resolver

The suite that covers the builder stages `getConnectionIdFromState`, so it can
characterize the fallback but cannot catch a defect that lives in owner
resolution itself. This one runs the real resolution over real store rows: two
repos publishing the same worktree id on different hosts, which is the
documented case where the owner cannot be named and `undefined` is returned.
Red with both fix files at the previous head, green with them.

Reverts the two caller pins added to the route census — the feasibility
predicate is exported from the planner, which the census already permits, so it
passes unedited and needs no permit clause.

* fix(native-chat): keep the structured launch's own agent eligibility check

Quick create's structured launch narrowed its guard to a bare `agent` presence
check, so a creation carrying an agent that cannot hold a structured session
reported itself cancelled once dismissed, where it previously reported that it
had done nothing. Unreachable through both callers today, but it is the last
local eligibility check in a module that otherwise trusts its callers for the
route, so it is restored rather than left to the required-route typing — which
says nothing about the agent.

Also corrects two comments that called the quick-create request "persisted".
It lives in renderer session memory and dies with the renderer; calling it
persisted made the plan/adopt split read as restart recovery, when what it
actually buys is a route decided before the worktree exists.

* fix(native-chat): keep the structured feasibility query typecheck-clean

The query threaded its narrow settings through the store, but the route
store's settings must satisfy the full GlobalSettings that two of its
resolvers require, so the narrow copy never fit. Ride the named settings
on the built input instead: the caller still names them, so a React memo
still depends on them, and no store-shaped object is needed.

Also give the launch state its delivery mode unconditionally; the key is
required, and a conditional spread makes it optional under
exactOptionalPropertyTypes.

* docs(native-chat): name the feasibility query's one remaining settings asymmetry

The builder reads launch customization off the store while the routing gate
reads the named settings, so one answer has two settings sources. It cannot
diverge with the single caller passing the object the store already holds, but a
PR about removing split sources should not leave that unstated.

* fix(native-chat): keep a coalesced joiner's draft unsent

joinLaunchDelivery stripped the joiner's delivery mode when the launch it
joined had established none, and an absent mode reads as submit. A joiner
that asked for a draft therefore had its text sent — the send-without-
consent this PR exists to prevent. Fall back to the joiner's own mode only
when nothing was established, so the first caller still wins otherwise.

* chore: re-trigger CI

GitHub created no workflow run for e935ea5e42 — the pull_request
synchronize event was dropped. No content change.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 14:54:09 -07:00
Brennan BensonandMerge Sim 2626e2eca4 Make the structured turn lifecycle row durable so completed durations survive (#19695)
* Make the structured turn lifecycle row durable so completed durations survive

A structured-chat turn used to end by tombstoning its running lifecycle item,
which threw away the only durable record of when the turn ended. Completed
"Worked for" labels therefore depended on the renderer having observed the
turn finish, and vanished on reopen.

The lifecycle item is now revised in place, never tombstoned:
- running, with startedAt, at the provider's turn start
- completed or interrupted, with completedAt, at the provider's terminal frame,
  a user stop, or a child exit the host observed
- unverifiable, with no end, when a cold acquire finds a running row from a
  generation whose exit nobody observed

Both timestamps are the execution host's clock at receipt, captured before the
deferred sink, so the completed value is identical on every client and needs
no client clock. Codex history restore uses the provider's own second-granular
endpoints for turns that predate this change. Desktop and mobile read settled
durations off the journal through one shared selector, and anchor the live
counter on the host start with the client's local receipt so a skewed client
clock never leaks into the label. Locally observed durations remain the
fallback for hosts that still tombstone.

Timestamps live inside the existing turnLifecycle field, which old clients
strip, and every working-state consumer keys on state === 'running', so no
capability negotiation is needed.

* native-chat: avoid stale working status on settled turns

* test: align settled turn status expectations

* Name settled lifecycle rows by their terminal state

An interrupted or unverifiable turn must not read as completed for any
consumer that renders status text raw. One shared helper builds the text for
both providers from the lifecycle state.

* test: deduplicate turn lifecycle suites

Each behavior keeps one test; duplicated harnesses and restated cases go.

* Key lifecycle rows to their user item and record the provider's measured duration

A lifecycle row now names the user item that opened the turn by its provider
key, so clients attribute timing explicitly and fall back to journal order
only for rows from older hosts. A provider-initiated turn with no prompt can
no longer claim the previous prompt's duration.

When the provider measures the turn itself (Codex turn.durationMs, Claude
result.duration_ms) the terminal row records it and clients prefer it over the
host interval, so a turn shows the same number live and after a history
restore. Host receipt times remain the live-counter anchor and the fallback.

* Record a turn as a first-class journal item

The turn record is now its own item kind rather than a status row carrying a
lifecycle field: no text to misuse, and the fold matches the durable turn
record other systems keep. Rows that carry it are stamped journal schema v3;
every other row stays v2, so an older host keeps reading them and latches
read-only at the first v3 row instead of truncating the epoch.

Clients that predate the item would paint an unknown kind as a text bubble,
so the host publishes the legacy status form to any client that does not
advertise agent-session.turn-item.v1, through the same per-client seam
background tasks use. The downgrade is transitional and goes once no
supported release lacks the capability. The shared projection now renders
unknown item kinds as nothing, so later kinds need no gate. One shared reader
handles both forms for old journals and old hosts.

* Preserve observed turn end across settlement retries

* Retain turn attribution for loaded chat history

* Preserve Codex exit receipt across close retries

* Register completed turn duration reliability gate

* Keep earlier turns through a Codex rewind and count a mid-turn attach from the real start

Findings from an independent adversarial review of the typed turn record:

- A Codex rewind adopted the provider's item list as the new epoch, and the
  provider never returns the host's own turn rows, so every duration before
  the rewind point vanished. The host's turn rows are now spliced back beside
  the item each followed, and recovery no longer expects the provider to
  prove rows it never owned.
- The epoch row was stamped with the current schema version, so an older host
  latched read-only at row 1 of every new session, defeating the mixed
  version design. It carries no body and stays at v2; a stored-row test now
  reads SQLite directly, because the reader upcasts every row on read.
- A send Codex folds into a running turn shares the opening prompt's provider
  key, and the alias map credited the duration to the later prompt. The
  earliest submission naming a key now wins.
- The live counter anchored on first sight, so a client attaching mid-turn
  counted from zero. Published frames now carry the host's clock, the reducer
  keeps the last sample with its local receipt time, and both clients anchor
  on how long the host says the turn has run.

* Correct turn duration gate assertion reference

* Respect authoritative unknown native chat duration

* Preserve unverifiable timing across older host upgrade

* Record final completed turn duration reliability evidence

* Fix the CI failures the merge left behind

- A merged import list named the same module twice, which the native code
  quality plugin fails on.
- A running turn is now reported by the host with no duration, so the settled
  map carries an explicit null for it; the hook test still expected the entry
  to be absent.
- main gave the older-page action a cursor with a head-trim guard, so the
  retention test's epoch-only action no longer typechecks; it now passes an
  unbounded sequence, which is what the old shape meant.
- The roster comparator moved into the extracted module, leaving its import
  unused in the reducer.

* Split two files back under the line cap after the merge

Merging main put both one effective line over 300, and the cap forbids a
disable or a shave. The wire module's refusal vocabulary moves to its own file
and is re-exported, so its consumers are untouched; the host's four thin
mutation delegates move next to the functions they call.

* Advertise the turn-item capability on every client transport

Local IPC and mobile advertised it; the remote and web transports did not, so a
desktop paired to a remote host, the CLI, and web silently ran on the legacy
carrier forever and the canonical row was never exercised there. The renderer
that paints it is the same build on every transport.

* Update the web auth-frame expectation for the new capability

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 14:32:50 -07:00
Brennan BensonandMerge Sim 721a269289 test(native-chat): split structured question fixtures (#19924)
Co-authored-by: Merge Sim <sim@local>
2026-09-10 13:35:15 -07:00
Merge Sim 4f5a8275e8 Revert "test(native-chat): split structured question fixtures"
This reverts commit 68e207ca2f.
2026-09-10 13:26:08 -07:00
Merge Sim 68e207ca2f test(native-chat): split structured question fixtures 2026-09-10 13:23:20 -07:00
Brennan BensonandMerge Sim 6e9de5fa58 fix(orchestration): revalidate an attempted Enter instead of resending it (#19911)
When a PTY retires mid-delivery, every staged message was marked undelivered,
which made all of them redeliverable. That is right for a pointer whose Enter
never fired, but an Enter that was already written may have landed: redelivering
it types the same mail into the pane a second time.

The Enter timer is cleared at the top of retirement, so a RESERVED or
WRITE_ATTEMPTED pointer provably never submitted and is released. An
ENTER_ATTEMPTED pointer is ambiguous and now stays at its phase for the resume
path to revalidate, matching the policy mailbox-pointer-submit.ts already
documents for an unverifiable settlement.

Co-authored-by: Merge Sim <sim@local>
2026-09-10 13:19:20 -07:00
Jinwoo Hong a6e6de93c4 fix(relay): keep failed rehome polls out of the durable failure budget (#19915)
* fix(relay): keep failed rehome polls out of the durable failure budget

The regional rehome worker polls claimRegionalRehome about once a second.
Any error thrown before an attempt was claimed - in practice a director pool
timeout on the pre-claim control read, 52-74 a day against a pool of 3 - was
charged to relay_region_rehome_worker_state.consecutive_failures, which
durably disables the control at three. That counter only ever resets on a
drain receipt, so while the control is disabled it never resets: production
sits at 1068 and still climbing. Enabling the control leaves the stale
counter in place, so the next pool timeout latches it straight back off.
That is what ended the 2026-08-28 enable after ten minutes.

- A poll that never claimed an attempt drained nothing, so it no longer feeds
  the dispatch-failure budget and logs .._poll_failed instead of
  .._dispatch_failed. recordRegionalRehomeWorkerFailure had no other caller
  and is removed.
- Enabling the control clears consecutive_failures and paused_until, so a
  budget spent under a previous enable cannot kill a fresh one. The dispatch
  interval in next_dispatch_at is deliberately left alone.
- The budget's auto-disable now emits
  orca_relay_regional_rehome_failure_budget_disabled, matching the existing
  .._safety_disabled precedent. It wrote no event before, which is why this
  went unnoticed for two weeks.

No change to region selection, the candidate query, or host eligibility.

* fix(relay): serialize rehome failure accounting with control updates
2026-09-10 16:04:25 -04:00
Brennan BensonandMerge Sim 5a96158849 feat(native-chat): focus the message box when a chat appears (#19868)
* feat(native-chat): focus the message box when a chat appears

Opening a native chat left focus nowhere, so you had to click the
composer before typing. Nothing in the chat surface focused it on open;
the only existing focus calls were reactive (typing on the bridge pane
background, picker acceptance, attachments, dictation), and the
structured pane had none of those.

useNativeChatComposerRevealFocus focuses the composer on the reveal
edge, covering a new chat tab, a worktree-create landing in chat, the
chat-view toggle, and switching back to an existing chat tab. Mount is
the wrong signal: retained panes hide with display:none + inert and
never unmount on a tab switch. It reuses the existing composer handle
and shouldPreserveEditableFocus rather than adding a parallel path, and
retries across a bounded run of frames because Tiptap publishes its
adapter after mount and Radix restores a closing dialog's trigger in a
setTimeout(0).

Two supporting changes:

- isFocusedGroup, from activeGroupIdByWorktree. On worktree activation
  both columns of a split flip visible in the same commit, so without it
  two revealed chats fight over the caret. The bridge route already had
  this bit as controller.isActive; only the structured overlay needed it.

- focusRuntimeTerminalSurface bails on a chat-covered pane. Its DOM-path
  twin already declines chat view via data-terminal-chat-view, but the
  runtime path focused the covered xterm unconditionally and pulled the
  caret out of the composer. Returns true, not false: false sends the
  caller to the DOM fallback, which for a structured tab id focuses an
  unrelated tab's xterm.

* fix(native-chat): preserve reveal focus ownership

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 12:38:06 -07:00
Brennan BensonandMerge Sim 4408fe897a feat(sidebar): show native-chat subagents as sidebar child rows, like CLI agents already do (#19807)
* feat(sidebar): indent native-chat subagents under their session row

Stacked on #19311, which adds the background-task channel this reads. The
bridge maps agent-kind background tasks into AgentStatusEntry.subagents, and
the renderer status feed confirms per connection so a reconnect cannot leave a
child asserting live from a stream that ended.

* fix(sidebar): avoid completed age for unverifiable subagents

* fix(sidebar): preserve unverifiable child verdicts

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 12:29:30 -07:00
f2af92b2fa feat(native-chat): show live background work and name each row by kind (#19705)
* fix(codex): reserve the label's share of a qualified command row

A child's label is raw provider text and was spliced into the command
row unbounded, then the pair clipped to the description cap. A label at
or past that cap clipped the command away entirely, leaving a row of
kind 'command' that named an agent and showed no command - the failure
qualification exists to remove, inverted. The same clip could also cut a
surrogate pair, which boundSubagentField already guards against on the
agent row two lines away.

Give the label a reserved share and clip it the way the agent row does.

* feat(native-chat): show live background work and name each row by kind

The strip suppressed itself in three places: the Claude tracker blanked
its roster for the whole of any turn, the Codex tracker returned nothing
while a primary turn was open, and the renderer view gated on
`turnId === null`. Between them, work in flight was never shown — and a
task backgrounded in an earlier turn vanished from the strip as soon as
the next prompt was sent. Claude additionally dropped every foreground
subagent, so a fan-out reported nothing at all.

Report work while it is live, in all three layers. Foreground Claude
work is turn-scoped, so `result` retires it — that is the provider's own
outcome for a task it marked foreground, not a roster sweep. Nothing
settles a Codex child on turn end: those keep reporting well past their
parent, so turn frames only prompt a republish.

Name each ROW by kind — Subagent, Shell command, Workflow, Monitor —
instead of a generic "Background <kind>", each drawing the glyph the
shared tool-icon table already uses for that category. A row that
carries a provider description still shows it unchanged. The collapsed
header summary is deliberately untouched; it is owned elsewhere.

The conversation-command gate is unchanged in effect: an open turn
already refuses first, and Claude foreground work never reaches the
backgrounded set the gate reads.

* fix(native-chat): withhold the row stop Claude foreground work cannot honour

The strip now publishes foreground rows, but `stoppableTaskIds` still filters
on `backgrounded`, so `stopClaudeBackgroundTasks` resolved an empty target list
and returned `{ cancelled: false }` that no renderer reads: the user clicked
"Stop Subagent" and nothing ever happened.

Carry stoppability per row instead of widening the stop to a target the SDK has
no way to reach. `AgentSessionBackgroundTask.stoppable` is absent-means-yes, so
hosts that predate it keep their working control, Claude emits `false` only on
foreground rows, and the strip hides that row's button the same way it already
hides the stop-all a provider cannot honour.

* fix(claude): scope aggregate-roster authority to the work it enumerates

`background_tasks_changed` lists BACKGROUNDED tasks, so a foreground subagent
can never appear in it. Treating it as the whole world meant any such frame
cleared every live foreground row mid-flight and then dropped every later
foreground `task_started` for the rest of the session, killing the in-turn
fan-out the strip exists to show in any session that ever backgrounds anything.

Decide `backgrounded` before the staleness guard and apply the guard only to a
backgrounded start, and retain live foreground entries across a roster replace.
Retained rows count against MAX_TRACKED_TASKS, so the map stays bounded, and a
stale backgrounded start the roster no longer lists is still dropped.

* test(native-chat): pin the strip's monitor amber to the constant that defines it

`MONITOR_GLYPH_COLOR`'s comment claimed a test held it and AgentStateDot's amber
together, but no test imported it — the assertions hardcoded 'text-yellow-500',
so the two could drift with every test still green. Read the colour from the
module, which is what the comment always said was happening. Drop the unused
`BackgroundTaskGlyph` export too: nothing outside the module names it.

* fix(native-chat): keep the task list open across a gap in live work

The strip is now mounted on live work, so a sequential fan-out unmounts it
between one subagent finishing and the next starting: local `useState` meant
the expanded list collapsed itself on every such gap, on top of the strip
flickering above the composer.

Hand the disclosure to the session, keyed by session id so it does not leak
across a session switch. The strip is now controlled and holds no state of its
own, which is what makes it survive its own mount churn.

* fix(codex): route every command-row cut through one surrogate-safe clip

`boundLabel` avoided splitting a pair, then `qualifiedDescription` re-cut the
COMPOSED string with a raw slice: label (<=96) plus separator plus description
(<=512) is up to 611 chars, so that second cut landed at an arbitrary index
inside the description and could publish a lone high surrogate — lossy through
any non-JSON UTF-8 hop. `parse` had the identical hazard on an unqualified
primary-thread command.

One `boundText` helper now owns all three cuts, so no path in the file can emit
a lone surrogate from well-formed input.

* fix(claude): keep terminal evidence for ids an aggregate roster never lists

Narrowing the admission guard to backgrounded starts left a finished FOREGROUND
id with no defence: `replaceAggregateRoster` wiped `terminalTaskIds` wholesale,
so after any `background_tasks_changed` a replayed `task_started` revived a task
whose completion had already been seen — and only a later `result` could settle
it again.

Scope the wipe the same way the guard was scoped: delete only the ids the
incoming roster actually enumerates. A roster still overrules terminal evidence
for the work it lists, which is what that behaviour was added for.

* fix(claude): keep retained rows in place and evict the stalest, not the newest

Re-adding retained foreground entries after the roster made a live row the user
is reading jump below the backgrounded rows on every `background_tasks_changed`,
and the cap `break` kept the STALEST retained rows while dropping the newest.

Merge in the tracked map's own order so a surviving row holds its position, and
count the overflow up front so eviction takes the oldest retained rows. Roster
entries are never starved and the map stays bounded either way.

* fix(claude): retire leftover foreground rows when the next turn starts

A foreground `task_started` arriving with no turn open has no `result` coming
to retire it, so it sat in the strip indefinitely — with no per-row stop, since
foreground rows are not stoppable — and refused conversation commands behind an
instruction nobody could follow.

Settle on turn start as well as on `result`. This is cleanup only: visibility
never consults `startsTurn`, so a missed one degrades to today's behaviour and
can never switch the feature off. It shortens the row's life to the next turn;
the case where no further turn is ever sent is filed separately.

* fix(agent-session): withhold unstoppable rows from readers that predate them

Rule 3 of remote-wire-compatibility: changing what the host publishes reaches
old clients with no wire change. The Claude host published no foreground rows
before this feature; it does now, and a client that cannot read `stoppable`
draws a per-row Stop on every one of them — Claude always sets
`supportsTaskStop` — which filters to the backgrounded ids, stops nothing, and
returns a result no renderer inspects. That is the dead button `stoppable` was
added to remove, reappearing across a version skew.

Negotiate it. A client can advertise the existing background-task-stop
capability and still predate `stoppable`, so this needs its own constant.
Readers that do not advertise it get unstoppable rows dropped, and a state whose
every row is dropped becomes no strip — exactly their pre-feature view.

RUNTIME_PROTOCOL_VERSION is not bumped: this adds an optional field and a new
negotiated capability, and changes no existing field's meaning, which is the
explicit do-not-bump case in protocol-version.ts.

* test(agent-session): name the projected rows so the fixture typechecks

An indexed lookup into the fixture's task list is possibly-undefined under
`pnpm tc`; the rows are more readable named anyway.

* test(web): advertise the row-stop capability in the e2ee auth expectation

The web e2ee handshake started sending
AGENT_SESSION_BACKGROUND_TASK_ROW_STOP_CAPABILITY, and this test asserts the
advertised list by deep equality, so it went red on CI while every targeted
test run stayed green. Add the capability in the position the router sends it.

* test(claude): pin why the roster empties mid-turn in a sequential fan-out

The strip unmounting between two sequential subagents is truthful, not a swept
row: A leaves on the provider's own terminal frame, B does not exist yet, and
backgrounded work spanning the same gap holds the roster open — so an empty
roster is never work the strip is hiding.

Also pins the previous-turn rule against the one the subagent roster already
applies on the same frame: a still-working FOREGROUND child becomes
`unverifiable` there and a backgrounded one is left alone, so the strip drops
the first and keeps the second rather than asserting `live` for either.

---------

Co-authored-by: Merge Sim <merge-sim@users.noreply.github.com>
Co-authored-by: Merge Sim <sim@local>
2026-09-10 12:16:30 -07:00
Jinwoo Hong 4b1b7178ad fix(orchestration): scope @ group addresses to the sender's Run (#19783)
* fix(orchestration): scope @ group addresses to the sender's Run

`@all`, `@idle`, and the agent-name groups (`@claude`, `@codex`, ...) resolved
against every terminal on the host. A coordinator meaning "my three reviewers"
reached 126 agents across every open project, twice in one day, and every
unrelated agent burned a turn discarding mail that was never for it.

Every group except `@worktree:<id>` now means the live Dispatches of the
sender's own Run, each addressed as `dispatch:<id>` so delivery is durable
even when the worker terminal is not attached yet. A sender bound to no Run is
refused with `invalid_argument` naming `run:<id>` / `dispatch:<id>`; there is
no host-wide fallback and the host's terminals are never enumerated for it.
`@idle` and the agent-name groups filter within that set by the same terminal
status and host-resolved identity as before. `ask --to @group` returns the
same code and points at the owning Run mailbox.

Federated Dispatches read relayed control mail rather than a local mailbox,
so a Run-scoped fan-out skips them with a `recipient_unreachable` warning
naming the direct `dispatch:<id>` address.

Group addresses are resolved host-side, so no RPC or stream shape changes; an
older CLI sending `@all` to a new host gets the Run-scoped meaning.

Claude-Session: run-scoped-group-addresses

* fix(orchestration): revalidate legacy takeover before the recipient verdict

A legacy coordinator taken over while `listTerminals` was in flight reported
`runtime_error` instead of `legacy_read_only`: Run scoping made "no live
Dispatch in this Run" the first thing the group send could fail on, and that
threw before the takeover check ran. Takeover is a precondition, not a
commit-time detail — the sender must be told it is read-only whatever else is
wrong with its recipient set.

Revalidation moves to immediately after the only `await` in the path.
Everything below it is synchronous, so the commit-time window it used to guard
is unchanged; only the error paths now see it.

The legacy partition test gave `term_current_worker` no Dispatch, so under Run
scoping it is correctly not a recipient. It now holds a real current-contract
Dispatch in the same adopted Run, which is what the test is named for: one
`legacy_direct` and one `current_delivery` recipient in one fan-out.

Claude-Session: run-scoped-group-addresses

* fix(orchestration): address the Run a nested coordinator created, not its parent

A nested coordinator is both a worker of its parent Run and the coordinator of
the Run it created. `resolveMessageRun` answers with the parent, correctly,
because that is where its own `worker_done` belongs — but audience is a
different question. Scoping `@all` to that Run sent a nested coordinator's
"shared context" to the siblings it was started beside instead of the workers
it started, and reported success, so it never learned its sub-workers heard
nothing. Before Run scoping the host-wide fan-out reached the sub-workers by
accident; this turned an over-broad delivery into a wrong-audience one, the
exact failure class the change exists to remove.

Group audience now resolves off the Run the sender coordinates, falling back
to its Dispatch's Run. A leaf worker coordinates nothing and is unaffected.
This is a separate question from `routing.run`, not a second answer to the
same one, so `resolveMessageRun` keeps its meaning for point-to-point mail.

Also: when every live Dispatch in a Run is federated, the fan-out skipped them
all and threw a bare `Error` that discarded the warnings naming those remote
workers and how to address each one. The sender was told "no recipients" while
three remote workers existed. That throw now carries a code and the skip
explanations.

Claude-Session: run-scoped-group-addresses

* docs(orchestration): say that no group address reaches a coordinator

A coordinator is not a Dispatch, so Run-scoped groups never include one. That
follows from the rule, but nothing said it, and the old host-wide meaning did
include the coordinator — a worker sending `@all` to raise a blocker would be
heard by its siblings and by nobody who can act. The guide, the CLI note, and
the docs page now say to use `run:<id>` for that, and that a worker which
created its own Run addresses that Run's workers.

Also restores the `@cursor` case dropped when the group tests moved: a Claude
pane titled "Fix the text cursor blink" must not receive Cursor's mail. That
hazard was recorded from real titles and `@droid` alone did not cover it.

Claude-Session: run-scoped-group-addresses

* fix(orchestration): preserve group audience and mailbox identity

* fix(orchestration): validate group scope before dispatch routing

* fix(orchestration): preserve pane identity and exclude coordinator dispatches
2026-09-10 15:06:03 -04:00
github-actions[bot] 26f9fd8ea1 Update README downloads badge 2026-09-10 18:29:07 +00:00