Commit Graph
10731 Commits
Author SHA1 Message Date
Neil 83369d0ffc test(relay): pin RelayControlClient teardown before, and twice after, connect
Characterisation only — no defect found, which is the result worth recording,
because the same audit found the service-level fence was not hard.

Covers closeNow() at 'idle' before any socket exists, connect() refused after
teardown, a control request issued before connect, and closeNow() twice from
both 'opening' and 'active'. Asserts the specific fields rather than absence of
errors: socket handle dropped so the second pass cannot re-terminate, onClose
not re-fired, pendingRequestCount back to 0, and vi.getTimerCount() 0 so
neither the connect deadline nor the silence watchdog outlives the close.

Mutation-verified: dropping the idle guard in connect(), `this.socket = null`,
requests.rejectAll, or silenceWatchdog.stop() each fails one of these.
2026-09-10 18:35:37 -07:00
Neil ed35db892b fix(relay): stop a mid-read profile switch reporting as a sign-out
Same taxonomy error 8972d744 fixed for the session file, in the other null
exit. readRelayAuthContext re-reads the active profile after the refresh,
because refresh and org selection can rewrite cloud linkage in flight. It then
collapsed two unrelated outcomes into null:

  if (!cloud || refreshed.profile.id !== active.profile.id) return null

A profile switch landing inside that window is not "the cloud session is gone",
which is the contract the coordinator's own comment says null carries. Measured
before: offlineReason "signed-out" — the terminal wire reason every paired
phone latches as "sign in on your desktop to reconnect", arming no retry — for
a race the switch's own authMutated() re-reads moments later. After:
auth_unavailable, which is retryable.

Split the condition: throw for the id mismatch, keep null for a profile that
genuinely has no cloud linkage.

Still unfixed, and not reachable from here: ensureActiveOrcaProfile treats an
unreadable profile index the same as a corrupt one — readProfileIndexFile
catches every error including EACCES/EMFILE, index and .bak fail together under
descriptor exhaustion, and it fabricates a default local profile and writes it
over the unreadable file. readRelayAuthContext then sees no cloud and reports a
sign-out honestly, because by that point the index really has been replaced.
That one belongs in profile-index-store.ts and changes app-wide sign-in
semantics, the same caveat 8972d744 raised for decrypt-failed.
2026-09-10 18:35:23 -07:00
Neil f1eb8037c8 fix(relay): drain the revoke outbox on one broker, not one open per item
flushAll runs unawaited from inside the coordinator's openBroker, so the broker
it is draining is not registered yet. Refreshing demand after every removal
therefore reconciled against a null ownership: each reconcile opened a fresh
director assignment and abandoned the broker still working through the loop.

Measured with three queued revocations and a real RelayRevokeOutbox on disk:
3 items produced 3 RelaySessionBroker.connect calls and 2 abandoned brokers,
against a relay that rate-limits assignment. After: 1 connect, 3 revokes, one
`connecting` transition, the draining broker never closed, outbox empty.

Reconcile once after the loop instead, and only when something was actually
removed. flushItem keeps its own refresh because onDeviceRevokeQueued calls it
directly on an already-registered broker, where retiring the last demand is
exactly what should reach standby.

Tests use the durable file as the assertion rather than a mock, and cover the
fence landing before the open resolves, the fence landing between two items,
the multi-item drain, the last-demand revoke reaching standby, and the same
reqId queued twice (dedup is the control layer's job — duplicate_relay_request_id
— so the service must only settle both without stranding the item).

Mutation note: moving onDrained one await later, inside the loop after
flushItem returns, survives. The extra microtask lets the coordinator register
the broker first, so no extra open occurs and the variant is genuinely inert
here; the assertion does catch the real pre-fix placement, synchronous inside
the revoke after remove.
2026-09-10 18:35:01 -07:00
Neil 4733b21cdc fix(relay): make the quit fence hard, and route quit to the terminal stop()
DesktopRelayService.stop() has zero production callers. Quit, relaunch and
pre-sign-out all go through fenceAndCloseNow() (main-process-quit.ts,
main-window-core-services.ts). That matters because the two are not
equivalent: stop() sets `stopped` and clears both timers, while the fence
cleared only `livenessTimer`, left `demandExpiryTimer` armed, and left
`stopped` false. Its comment — "the next auth mutation re-arms via
refreshDemand" — is right for sign-out and exactly wrong for quit. The path
production takes was the one that permitted a resurrection, and the suite was
exercising the other one, which is why this survived.

Three doors re-armed the fence, each measured against a fresh service:

  - a pending invite-expiry wake: getTimerCount() is 1 immediately after the
    fence, and the timer fires refreshDemand and reopens the broker
  - a mint settling after the fence, through withTransientDemand's finally:
    0 timers at the fence, 1 once the mint settles, and connect runs again
  - a power-resume ensureLive(), gated only on `stopped`: connect called twice,
    a new relay control session opened after quit

Two intents, kept apart rather than unified. fenceAndCloseNow gains a `fenced`
latch cleared only by start()/authMutated(), so sign-out and relaunch keep the
re-arm the old comment protects while the dangerous window between the
pre-sign-out fence and the profile wipe stays shut. Quit calls stop() instead,
which is terminal; the wire behaviour is unchanged, since coordinator.stop()
fences reasonlessly exactly as before. The outbox flush gets the same latch,
because it runs unawaited from inside the broker open and a fence can land
between two revokes.

Deliberate on a vetoed quit: Relay stays down for the session. before-quit
already nulls the mobile pairing provider unconditionally, so pairing is dead
there regardless — a re-armed broker would be a zombie holding a control
session nothing can use.

Not fixed here: `fenced` is not set by stop(), so the two latches stay
independent and a future caller can still pick the wrong one.
2026-09-10 18:34:26 -07:00
Neil cd175054c5 refactor(relay): extract the revoke-outbox flush into its own module
Behaviour-preserving move, no logic change: flushRevokeOutbox and flushRevoke
leave DesktopRelayService as RelayRevokeOutboxFlusher.flushAll/flushItem. The
two fixes that follow each add lines to this file, and it has no headroom under
the 300-line cap; a max-lines disable is not an option.

Why a new file rather than reusing something: nothing existing drains the
outbox. RelayRevokeOutbox is the durable store — making it drain itself over a
control would give the persistence layer a dependency on a broker.
RelayControlRequests dedups in-flight reqIds on a single socket, not queued
items across reconnects. RelayDemandLedger reads the outbox for demand and
never mutates it. The drain only ever existed inline here, so this is an
extraction, not a parallel implementation; the only new code is the options
type and the class shell.
2026-09-10 18:33:41 -07:00
Neil a3512fcd69 fix(runtime): stop the epoch-history cap evicting a fence that can still be beaten
The LRU cap added in 20fb0b3f6e bounded memory by discarding a safety property.
Reported as a review concern, reproduced before believing it, and confirmed:

`sessionTabsPublicationEpochHistoryByWorktree` is deliberately RETAINED after a
worktree's live record is dropped — it is the tombstone fence that stops a
sibling stream's late frame from a retired publisher being applied. A pure LRU
evicts exactly those tombstones, by construction: a still-publishing worktree
renotes its epoch on every accepted frame, so the eviction victim is always a
fence.

Absence is fail-OPEN on both read paths. `hasRetiredValue` answers false for a
missing entry, so `isRetiredSessionTabsPublicationEpoch` cannot tell "never seen"
from "fenced and forgotten" — and `recordReceivedWebSessionTabsSnapshot` then
RE-NOTES the evicted epoch as current, resurrecting the publisher the fence
retired.

MEASURED, not argued. Fence a worktree, note 512+ others, deliver the retired
publisher's late frame: with the tombstone retained it is rejected; after
eviction it is ACCEPTED.

The first version of that repro passed and was wrong. It gave each worktree its
own runtime id, so the retired RUNTIME-ID fence rejected the frame and the epoch
fence was never consulted — a test green for an unrelated reason. A real sibling
stream on one environment shares the runtime id; with one shared id throughout,
the acceptance reproduces. Recorded here because the fixture detail is the whole
difference between a passing test and a real one.

THE FIX: the cap yields to the fence rather than the other way round. Each entry
carries `notedAt`, and eviction stops at the first entry younger than
SESSION_TABS_PUBLICATION_FENCE_RETENTION_MS. The map may exceed 512 while every
entry is still young; memory is then bounded by worktree churn WITHIN the
retention window instead of by count, which is the bound that can be held
without discarding a live fence.

512 IS NOT THE JUSTIFIED NUMBER, and it was right to challenge it — it was a
memory target with nothing behind it. The justified number is the retention
window: 4x the tab RPC budget. A frame in flight longer than that budget has
already been abandoned by the transport, so a fence older than 4x it has outlived
anything it could fence against. The count cap is now only the memory target, and
it is the one that gives way.

Both growth tests are updated to advance the clock, which also makes them
faithful: a long-running session is long in TIME, and a count-only fixture
measures the map at an instant where nothing is evictable and no bound can hold
without dropping a live fence.

Also writes the structural-assertion lesson into the teardown test's header
rather than leaving it in a commit message: the original perf assertion was a
wall-clock ratio that flaked at 22.3x against a 20x bound, and was replaced with
a `.keys()` spy, because the enumeration IS the cost. A duration assertion fails
for reasons unrelated to the property it protects and gets retried away, taking
the real regression with it.

Mutation: removing the retention guard (back to pure LRU) kills exactly the
"keeps a young fence past the cap" assertion, and leaves every bound assertion
passing — the bound tests and the fence test are cleanly separated.

Verification: 190 files / 1595 tests pass, log grepped = 0. pnpm tc and oxlint
clean.
2026-09-10 18:13:33 -07:00
Neil 20fb0b3f6e perf(runtime): index session-tabs tracking keys by environment, and bound the epoch history
Two defects that only appear at scale or over time, in one commit because the
fix for the second calls into the structure the first introduces — see the
coupling note at the end.

1. THE PER-ENVIRONMENT TEARDOWN SWEEP WAS QUADRATIC IN ENVIRONMENT COUNT.
`clearWebSessionTabsTrackingForEnvironment` prefix-scanned every key of eight
module maps, each holding E*W entries, plus one global walk — so a catalog
update clearing E environments cost 8*E^2*W + E*W. It runs on every
pairing-revision change.

Measured with a counting probe on the real function (exact key visits, no
clock, so a loaded machine cannot contaminate it):

  E=10 W=50    500/map    17,000 visits -> 500    (34x)
  E=20 W=50   1000/map    58,740 visits -> 1,000  (59x)
  E=40 W=50   2000/map   202,220 visits -> 2,000  (101x)
  E=50 W=20   1000/map   141,600 visits -> 1,000  (142x)

Doubling E at fixed W: 17,000 -> 58,740 -> 202,220, i.e. 3.45x and 3.44x per
doubling — quadratic, not linear-but-large. The indexed column over the same
sweep is exactly 2x per doubling. Doubling W at fixed E is 1.84x and 1.78x,
linear as expected. These are freshly-populated fixtures with nothing stale in
them, so the quadratic term is a property of the scan itself and not of a leak.
Wall clock for cross-reference only: 1.92ms -> 0.64ms at 20x50, 37.1ms -> 4.8ms
at 100x50.

The fix is the index shape this file already uses for
`hostSessionTabMappingKeysByEnvironmentAndWorktree` — that precedent is why this
is worth indexing rather than tolerating.

The regression assertion is STRUCTURAL, not a timing threshold. A wall-clock
ratio bound was tried first and flaked at 22.3x against a <20x limit under suite
load, which is exactly the failure mode a timing assertion has on this machine.
It now spies on `.keys()` for the nine per-worktree maps and asserts teardown
never enumerates any of them — enumerating them IS the sweep, so it pins the
property directly and is load-independent. Mutation-checked: restoring the
prefix sweep kills it.

2. `sessionTabsPublicationEpochHistoryByWorktree` WAS UNBOUNDED IN KEY COUNT.
Only its inner retired array was capped (at 8); the map itself grew 1:1 with
worktree lifecycles, since the entry is deliberately retained as a tombstone
fence after the live record is dropped. 10,000 create/remove cycles left 10,000
entries. Now LRU-capped at 512, re-inserting on every accepted frame so a
still-publishing worktree is never the eviction victim — the tombstones, which
are what actually accumulate, are evicted first.

WHY ONE COMMIT: the eviction added by (2) calls
`releaseSessionTabsEnvironmentKeyedWorktree`, the index added by (1), because an
evicted tombstone must also release its index entry. Split, the first commit
would leak the index on eviction or the second would not build. The coupling is
real rather than incidental, so they land together.

Verification: 191 files / 1611 tests pass, log grepped for ELIFECYCLE and
failures = 0. pnpm tc and oxlint clean.
2026-09-10 18:04:03 -07:00
Neil cbf18d46d4 fix(runtime): key the wake-respawn latch per environment, like its twin
The wake-respawn latch is the initial-terminal bootstrap latch's twin — both stop
one focus from issuing two creates for a workspace, both are consulted from the
same subscription closure, and they sit one line apart in the same teardown. The
bootstrap latch got two fixes this stack; this one got neither, and still carried
both defects:

1. NOT KEYED BY ENVIRONMENT. It was a bare `Set<worktreeId>`. A worktree id is
   `repoId::path` with no host component, so the same id can be live on two
   paired runtimes at once (STA-4343) — which is exactly why the bootstrap latch
   was re-keyed. Keyed by worktree alone, one runtime's respawn suppressed the
   other runtime's, and one runtime's `end` released the other's claim.

2. A CLEAR-EVERYTHING INSIDE A PER-ENVIRONMENT TEARDOWN.
   `clearWebSessionTabsTrackingForEnvironment` called
   `clearAllWebRuntimeWakeTerminalRespawn()`, so tearing one environment down
   freed every other environment's in-flight claim and a fresh closure for the
   sibling could issue a second respawn while the first create was still running.
   That is STA-6173, and the bootstrap latch's own docstring names it: "clearing
   every environment's latch would release a sibling environment's pending create
   and let a new subscription for it seed a duplicate, which is this very bug
   through another door." The correctly-scoped bootstrap clear is the very next
   line.

Found by reading the teardown function as a list rather than as prose — a
clear-everything is a one-line call that looks identical to a scoped one at a
glance, which is how it sat next to the fix for its own defect.

Both latches now have the same shape: `Map<environmentId, Set<worktreeId>>`, a
per-worktree release, and a per-environment clear that touches only its own keys.

The existing wake-respawn test is threaded with an environment id rather than
rewritten: its contract was correct, only the signature moved. The new file
pins the two defects themselves.

Also documents `clearWebSessionTabsTrackingForEnvironment` as what it has become
— the per-environment teardown registry. It now carries seven module clears, and
a new per-environment map belongs in that list rather than behind a trigger of
its own. The doc says each clear must be scoped to THIS environment and names
the wake-respawn latch as the one that was not, because the next clear-everything
will look just as harmless.

Mutations: reverting the per-environment clear to a clear-all kills exactly the
sibling-claim assertion; making the skip check ignore the environment (the old
worktree-only keying) kills exactly the two-environments assertion and the
sibling-claim one. No survivors.
2026-09-10 17:52:10 -07:00
Neil 9de4e9dd21 fix(runtime): drain a removed environment's handle-gap verdicts on teardown
`expiredGenerationByPane` is pruned only by rules that run when a verdict is
RECORDED — the stale-generation sweep here, and the tab-death sweep added
separately (8f16641130, env-scoped in c0e44238ea). An environment that is
REMOVED records nothing ever again, so neither rule can reach its rows and they
survive for the life of the session. Two orphan classes on one map; neither
prune subsumes the other, because both are driven by a recording.

Severity is a leak, not a correctness bug, and the commit pins WHY so nobody
re-derives it: removing an environment advances its connection generation, so a
stranded verdict can never match again even if the id returns. That test exists
to stop the generation advance being "optimised" away later, since it is the
only thing making the stranded row inert.

Hung off `clearWebSessionTabsTrackingForEnvironment` because that is the only
caller that fires for an environment that is going away.

Clears VERDICTS ONLY. Parked waiters deliberately survive, matching
`clearHostSessionMirrorHydration`: a re-pair or effect restart replaces the
connection's evidence, it does not cancel the recovery this client still owes
the pane. A waiter left behind is bounded by its own deadline and replays its
sweep exactly as it would have. Clearing them here would silently drop a parked
resume that nothing else replays.

A measurement worth recording, because it argued me out of a change I was about
to make: on the unfixed map the per-expiry rescan is super-linear — 500/1000/
2000/4000 sequential expiries cost 7.2/15.3/51.8/173.1 ms, doubling ratios
converging on ~3.35 against 4.0 for quadratic. That looked like a case for
reshaping the map to `Map<env, {generation, Set<tabId>}>`. It is not: the
quadratic is a property of the LEAK, not of the scan. Once the tab-death prune
holds the map at roughly one entry per environment the scan is over ~1 entry,
and a counting probe on the fixed map (summing `map.size` across N expiries,
which IS the iteration count and needs no clock) gives exactly N-1 — linear, and
2000x fewer iterations than quadratic at N=4000. The flat prefix loop used here
is the established pattern in this subsystem and needs no restructure.

Two methodology traps this cost, recorded for the next person measuring in this
repo: `vi.useFakeTimers()` fakes `process.hrtime` and `performance.now` as well,
so a timing harness reports the advanced deadline rather than work done — fake
only the timer surface under test. And expiring N panes in one burst measures
the fake-timer harness clearing N timers, not product code; 1000 panes "cost"
~1s that way and almost none of it was ours.

Mutations: a clear that drops nothing kills exactly the two assertions that
claim it drains, and correctly leaves the waiter-survival and generation-advance
tests passing. An UNSCOPED clear kills the same two, via their sibling-
environment half.
2026-09-10 17:44:39 -07:00
Neil c3d46edfea test(relay): pin the mid-mint policy flip under interleaving, not just in sequence
The mid-mint LAN-flip fix (a97c9e2, "name the LAN flip that lands mid-mint, not
relay_control_not_active") was pinned with the flip sequenced BEFORE the mint.
That leaves the case it was actually written for untested: the flip landing at
an await point inside the mint.

Both interleavings now run — the flip landing while the mint is parked on an
in-flight broker open, and landing mid `create_pairing_relay` after the broker is
already live. Both name `relay_disabled_for_device`, so the fix holds under
interleaving and not merely in sequence.

No production change. Committed separately because it verifies an existing fix
rather than carrying one, and routes to whoever owns that fix.
2026-09-10 17:28:34 -07:00
Neil 193119c18d fix(relay): scope nextPendingExpiry the way hasDemand is scoped
`hasDemand` requires four things of a standing binding — mobile scope, matching
owner identity, matching relay host, and allowed by the live pairing policy.
`nextPendingExpiry` applied NONE of them: it took the minimum `inviteExpiresAt`
across every device in the registry. So it armed the `demandExpiryTimer` wake in
`refreshDemand` off state that provably cannot produce demand.

Three cases measured returning a wake time where `hasDemand` was already false:
a binding for a different relayHostId, a runtime-scope (non-mobile) device, and
a phone the live LAN policy excludes.

Low severity on its own — the fired timer just reconciles to no demand — but it
is spurious wakeup churn from the class that owns the correct predicate three
lines above, which is the kind of drift that stops being harmless later. The
shared clause is now `demandCandidateBinding`; the owner check stays in
`hasDemand` because `nextPendingExpiry` takes no identity. The `hasDemand` side
is a pure conjunction reorder, confirmed behaviour-preserving by mutation rather
than by eye.

Re-arming after a policy exclusion is safe: `pairingPolicyChanged()` calls
`refreshDemand()`, so a flip back to automatic restores the timer. That round
trip is asserted rather than assumed.

Also records two things next to the code that would otherwise be lost:

- Transient refs carry NO OWNER IDENTITY, so `hasDemand`'s transient loop answers
  true for any signed-in identity while the two branches below it filter on
  `ownerIdentityKey`. Nothing defends the current behaviour OR that regression:
  scoping the loop by owner fails exactly one test across the whole relay suite,
  the characterisation test added for it. Deliberately not fixed here, and the
  comment says why the in-file version is unsafe —
  `withTransientDemand('provision')` calls `setMobileRelayBinding` INSIDE the
  operation, so during a re-pair the device still holds the old owner's binding
  and an inferred filter would drop demand mid-provision, tearing the broker down
  under the operation holding the ref. The real fix threads identity through
  `acquireTransient`, whose call site is desktop-relay-service.ts.

- The `if (!current) return` guard in the release closure is unreachable today
  and mutating it away breaks nothing. Kept and labelled INERT rather than
  deleted: it goes load-bearing the moment the ledger grows a dispose()/clear(),
  and nothing in the suite would catch that.

The ref-counting core itself came back clean under every hazard exercised:
policy flip between acquire and release (the release closure is policy-blind by
construction, and the demand filter is a live pull per call rather than an
acquire-time snapshot, so neither a permanent pin nor a dropped-but-needed
demand); double release; opposite-order release of two refs on one key; and
cross-device release.
2026-09-10 17:28:21 -07:00
Neil 1673716c6d fix(relay): bound the live-broker wait over a chain of superseding reconciles
The budget bounded the armed-retry chain only. A waiter that kept following
fresh reconciles — which the previous two commits made it correctly do — rode
that chain with no deadline check at all.

MEASURED, on a fake clock: budget 1,000 ms, reconciles landing every 200 ms,
opens taking 400 ms and rejecting. Pre-fix the waiter NEVER SETTLED within a
300,000 ms horizon. Post-fix it settles at exactly 1,000 ms, the budget the
caller asked for.

Reachable in production because `withTransientDemand` reconciles on BOTH acquire
and release, so a busy host supersedes a waiter's reconcile faster than opens
settle, and the caller rides the chain indefinitely while holding its demand ref.

The deadline check is gated on `joinedReconcile` so the contract that matters is
preserved: the one open the waiter ARRIVED on is still never cut short, because
cutting a slow-but-succeeding open short would fail a pairing that was about to
work. The existing "slow first open past the budget still wins" test passes
unchanged, which is the assertion that proves the gate.

This is the third defect in this loop and all three share one root cause:
`await pending` had no escape once the reconcile the waiter captured stopped
being the one that mattered.
2026-09-10 17:28:00 -07:00
Neil 6cfc1f2a4a fix(relay): release a live-broker waiter parked on an open that stop() abandons
`stop()` fences and closes, but did not wake waiters. A waiter parked on an open
that never settles therefore had nothing left to release it — the wait is
deliberately unbounded across the open it joined, and the coordinator it was
waiting on is gone. Measured: the promise stayed pending for the life of the
process.

This is promise-shaped, not timer-shaped, which is why it survived the leak
checks: `vi.getTimerCount()` reads zero at teardown and always did. The new test
asserts on the settled value, and keeps the timer-count assertion so the two are
not confused again.

`fenceAndCloseNow` now fires the same authority-change signal a fresh reconcile
does, which is all a parked waiter needs to re-read the world and return the
settled cause. One line, because the wake mechanism it reuses landed in the
previous commit.

The sibling case — a waiter parked on an ARMED RETRY when the coordinator stops
— was already clean; pinned in the same file so the distinction between the two
teardown positions stays visible.
2026-09-10 17:27:31 -07:00
Neil 8904488d3b fix(relay): wake a live-broker waiter when a newer reconcile takes the authority
A waiter joined `latestReconcile` and awaited it unbounded, on the reasoning
that a reconcile always settles. True, but incomplete: once that reconcile is
superseded its result is DISCARDED, so the waiter sits out an open nobody will
use — measurably still parked after a newer reconcile had already registered a
live broker.

The wait now races the joined reconcile against an authority-change signal that
`beginReconcile` fires, so a waiter follows the reconcile that actually matters
instead of the one it happened to arrive on. The open the waiter arrived on is
still never cut short.

WORTH KNOWING FOR ANYONE TOUCHING THIS LOOP: the generation check that guards
this — `if (pending !== source.reconcile()) continue` — was completely
unprotected. Deleting it survived ALL 226 existing relay tests. It is the thing
that keeps a waiter from answering out of a reconcile that no longer speaks for
the coordinator, and nothing in the suite noticed its removal. The new test
"follows a superseding reconcile that is still opening rather than answering
from the old one" pins it, and is the only test that fails when it is deleted.

The loop moved to relay-live-broker-wait.ts because the fix pushed
relay-auth-coordinator.ts past the 300-line cap and AGENTS.md forbids disabling
max-lines. The source is a record of getters rather than a snapshot because the
wait re-reads all of it after each await.

Also pins three interleavings that came back CLEAN, so the scope of "clean" is
on the record: two waiters on one armed schedule (the short-budget waiter
neither cancels the retry nor strands the long one); a retry firing on the exact
tick the budget expires (the armed retry takes the tie deterministically and the
waiter returns that attempt's cause, not a null one); and a terminal cause
landing while a retryable wait is armed.
2026-09-10 17:27:06 -07:00
Neil 6f514e4022 fix(runtime): isolate one worktree's replay from the mirror-hydration drain
The same fan-out hazard as the handle-gap drain, one module up. Settling an
environment drains every worktree parked on it in a single loop, called from the
frame apply, with `waiter.run()` unguarded. One replay that throws strands every
waiter queued behind it and surfaces in the caller applying the frame.

Found by looking for the sibling of a defect rather than by a separate
interleaving: both modules park a `run` callback and drain N of them from one
event, so both have the same blast radius. Kept as its own commit because the
two modules route independently.

Mutation: dropping the guard kills exactly the one new assertion.
2026-09-10 17:26:01 -07:00
Neil 15f3401415 fix(runtime): isolate one pane's replay from the handle-gap drain
One store write releases every due pane, and the drain runs synchronously inside
a zustand subscriber. `waiter.run()` was unguarded, so a single pane's replay
reached two things it has no business touching:

  - the throw escapes out of `useAppStore.setState`, meaning the mirror apply
    that published the PTY handle throws at its own call site;
  - every pane queued behind the thrower is stranded — waiter still parked,
    deadline still armed — and then decides on a connection whose evidence
    landed long ago.

The deadline path fans out the same way, so a throwing replay also escaped the
timer callback.

Reachable: `resumeSleepingAgentSessionsForWorktree` reaches `state.createTab`
with no guard of its own. The panes in a drain are strangers to each other and
to the frame that released them; none of them should be able to see another's
failure.

The new tests live in their own file because
host-mirror-handle-gap-resume.test.ts drives the waiter through the real resume
sweep and so cannot choose what a replay DOES. Note for anyone extending that
file: per its header, "did the waiter release" is not an observable here — a
spurious release is re-parked immediately and reads identically one tick later.
These tests assert on timer count and on the deadline instead.

Also records two findings next to the code, so they are not rediscovered:
`expiredGenerationByPane` is never pruned for a removed environment (bounded and
inert, since removal advances the generation, but it does not drain — and a
DIFFERENT leak in that same map is being fixed concurrently, so reconcile rather
than patch around it); and sustained reconnect churn holding a pane parked
indefinitely is CORRECT, not the latch-that-never-releases defect, because under
churn liveness genuinely is unverifiable and ssh-execution-boundary.md forbids
resolving that to `exited`. It has the shape of the defect and will eventually
be "fixed" by someone who does not know that.

Mutation: dropping the guard kills exactly the three new assertions and leaves
all twelve existing waiter tests passing.
2026-09-10 17:25:48 -07:00
Neil bf497512d8 fix(runtime): stop a re-pair freeing an in-flight initial-terminal claim
The three initial-terminal suppression defects already fixed were all exits of
`dispatchWebRuntimeInitialTerminalBootstrap`. Those exits ARE exhaustive, and
this change is not a fourth one: `WebRuntimeTerminalCreateOutcome` is a closed
union of `created | failed`, so the helper's five paths — claim denied, thrown
failure, returned failure, success with a mirrored row, success without one —
cover every way it can return. Read alongside the three prior fixes this looks
redundant; it is not, because the gap is a release from OUTSIDE the helper,
which the helper cannot see and cannot guard.

`clearWebSessionTabsTrackingForEnvironment` dropped the environment's whole
latch map regardless of phase, and it runs on a pairing-revision change — which
is simultaneously a dependency of the active session-tabs effect. So a re-pair
freed an in-flight claim and re-armed a closure whose `requestedInitialTerminal`
is false in the same tick, and the next empty frame owned a second create with
the first still unresolved. That is STA-6173 restored through a door one level
up from the latch.

Teardown now drops a parked claim — nothing will answer an `awaiting-mirror` key
once its subscription is gone — but only MARKS an in-flight one, as
`creating-after-teardown`. It still blocks, and the create's own dispatch
releases it on every exit it has. The inverse hazard is real, so a torn-down
claim that resolves without a row is released rather than parked: no frame is
coming to release it.

Holding is bounded: every RPC on the create path carries `timeoutMs: 15_000`,
the placement settle a 10s deadline, and the whole body sits in a try/catch, so
the create always settles.

Two existing tests are rewritten rather than deleted because they asserted the
BUGGY contract — that teardown must free an in-flight claim, to stop "a create
RPC that never settles" from suppressing the next bootstrap. That premise is
unreachable for the bounds above, and what the free actually did was hand the
claim to the closure the same teardown re-armed. Their replacements pin the
bounded invariant instead, plus the parked-claim half that teardown must still
drop.

Mutations: freeing the in-flight claim at teardown kills exactly three
assertions; parking a torn-down claim kills exactly one. No survivors.
2026-09-10 17:25:23 -07:00
Neil 09be96a2b6 test(runtime): pin sleeping-agent resume on a failed SSH target
The terminal-state floor in workspace-terminal-host-authority.ts has three
consumers: initial-terminal seeding, the startup terminal watcher, and
sleeping-agent resume. Seeding is covered end to end by
worktree-agent-activation-seam.test.ts. Resume was covered only at the
predicate, so nothing failed if the floor stopped reaching it — and the
floor's own comment says the cost of losing it is a failed target left
terminal-less with unresumable agents for the rest of the app session.

Pins the resume half directly: an SSH git worktree on a target whose sync
terminated in offline/error with an empty hydrated set resumes its sleeping
agent. Two controls keep the floor from widening into "resume whenever we
are unsure" — an in-flight 'pulling' sync and no sync status at all both stay
unverifiable and resume nothing.

Verified by mutation: emptying TERMINATED_WITHOUT_ANSWER_PHASES fails exactly
the two floor assertions and leaves both controls passing.

Routes independently of the two fixes on this branch: the floor predates this
stack (#16750), and this only closes a coverage gap in it.
2026-09-10 16:43:38 -07:00
Neil a97c9e2144 fix(relay): name the LAN flip that lands mid-mint, not relay_control_not_active
withTransientDemand checks the host pairing policy once, at entry. But
RelayDemandLedger.hasDemand filters the in-flight operation's OWN transient
ref through the live policy, so a flip to local-only while createPairingRelay
or provisionRelay is awaiting withdraws the demand that operation is holding.
The coordinator then reaches its no-demand branch and publishes 'standby',
publish() clears offlineReason to null for any non-offline status, and the
broker wait returns with no cause at all — relay_control_not_active, the
generic code this stack exists to remove.

The flip is exactly the cause and relay_disabled_for_device is already its
word; the entry gate just could not see a flip that had not happened yet.
Re-ask the policy on the failure path rather than trusting the thrown code.

Reproduced first: the added test failed with "expected ... to throw error
including 'relay_disabled_for_device' but got 'relay_control_not_active'".
Controls cover the two ways this could over-reach — a failure the policy had
nothing to do with is not rewritten, and a grant that succeeded under a flip
is left alone.

Top follow-up found alongside this and deliberately not fixed here:
flushRevoke swallows every error with a bare catch, no attempt cap and no
item expiry, while the revoke check in hasDemand is deliberately not filtered
through the policy. A revoke that fails permanently server-side therefore
holds demand open forever and defeats the LAN pick entirely — the same
symptom this change is part of closing, by a different route, with no test
coverage.
2026-09-10 16:43:38 -07:00
Neil c5a38e6383 fix(relay): stop an unreadable session file reporting as a sign-out
The coordinator's own comment states the contract: readContext "throws on
transient failures and returns null solely when the cloud session is gone
(absent, or cleared by a 401)". The implementation did not honour it.

readFreshOrcaCloudSession collapsed readOrcaCloudSession's `unreadable`
status into `reconnect-required`, so readRelayAuthContext returned null and
the coordinator published RELAY_HOST_CLOSE_REASON.SIGNED_OUT, closed the
broker with that wire reason, and armed no retry because a terminal cause
arms none. `unreadable` fires on EPERM/EACCES/EBUSY/EMFILE/ENFILE/EIO —
descriptor exhaustion on a busy machine, a Windows AV file lock. The phone
latches setHostSignedOut and shows "Desktop signed out — sign in to Orca on
your desktop to reconnect" for a failure the next read would have cleared.

The contrast is the argument: `unreadable` is the one status the session
store's own doc says "licenses nothing", and clearCloudSessionIfUnchanged
already refuses to delete a session because of it — "a session we were
denied is not a session we may delete". The relay taxonomy was the single
place treating it as evidence the user signed out.

Give it its own arm on FreshCloudSessionResult and throw for it in
readRelayAuthContext, which lands it on the retryable auth_unavailable path
the coordinator already has. Every other caller tests `status !== 'found'`,
so their behaviour is unchanged. Correct the coordinator comment to say what
it actually depends on: null-means-gone is a contract readRelayAuthContext
owes it, not something that branch can verify.

Measured before the fix: offlineReason "signed-out", mint code
relay_signed_out. After: auth_unavailable.

Not fixed here, and worth a separate look: `decrypt-failed` when
safeStorage.isEncryptionAvailable() is false (a Linux keyring still locked at
login) takes the same route to SIGNED_OUT, but reclassifying it changes
sign-in semantics well beyond the relay.
2026-09-10 16:43:37 -07:00
Neil 614ea3f6a9 fix(session): an emptied workspace still survives the reconnect merge
localActiveWorkspaceSurvives decided whether the workspace the user is
standing in still exists by counting its tabs. A workspace they had just
closed the last terminal in reads as zero, so it "did not survive", the
host's null activeWorktreeId was taken literally, and the reconnect
dropped them onto the home screen -- for closing a tab.

An explicit empty tabsByWorktree row is the record that the user closed
the last terminal, not the absence of a workspace (initial-terminal.ts).
hasLocalTabsRow two hundred lines up in this same function already draws
that distinction with Object.hasOwn; this line was simply missed.

The counterweight is pinned too: presence must not turn into "never
follow the host", so a host that does name an active worktree still wins
over the emptied local one.
2026-09-10 16:43:35 -07:00
Neil 6bd12fc3e9 fix(terminal): a live pane owns its transcript in any workspace
The resume dedup was scoped to the record's own workspace on both terms
-- the entry's tab had to be in worktreeTabIds AND entry.worktreeId had
to match -- while the completed-turn widening only relaxed the status
term. A cross-workspace record whose peer pane is `done` and still holds
a live PTY therefore matched nothing, and the sweep launched a second
agent onto a transcript the peer is still writing.

The two ids really do drift. canonicalizeTerminalSessionWorktreeId
re-keys tabsByWorktree, tabGroups, tabGroupLayouts, activeTabIdByWorktree
and activeGroupIdByWorktree onto the canonical worktree id, and does NOT
re-key sleepingAgentSessionsByPaneKey, whose records carry worktreeId
inside them. So adopting an orphaned terminal is a direct producer of a
record naming one workspace while its pane and status row name another.

Split into two arms rather than widening the existing condition. The new
arm carries no workspace scope but demands hard evidence: a provider
session id names one transcript, so a pane whose exact PTY is live right
now already owns it wherever that pane sits, and no workspace boundary
makes a live PTY less live. The scoped arm keeps its scope and its
state !== 'done' term, because a status row with no live PTY is a claim
about the past and must not reach across workspaces.

Pins both directions: the live peer is not forked, and the same peer
without a live PTY still resumes.
2026-09-10 16:43:35 -07:00
Neil c2f14d1cd7 fix(runtime): void a handle-gap verdict the reconnect made stale
The per-pane park bounds itself with one deadline per connection, but the
waiter never recorded WHICH connection it was armed on. A wait armed on
generation 0 that fires after a reconnect stamps its expiry against the
current generation, so hasHostMirrorHandleWaitExpired agrees, the mirror
lookup returns null, and the pane is resumed after 1ms on a connection
that has had no chance to publish the handle. That is #19735's fork with
an extra step, reached through the guard that exists to prevent it.

The module's own doc comment claims the opposite -- "a reconnect bumps
the connection generation and arms a fresh wait" -- and that is true only
for a wait which had ALREADY expired, which is precisely the case the
existing test covered. The test and the comment agreed with each other
and both were wrong about the live case.

The waiter now carries the generation it was armed on and records no
verdict when the generation has moved; the replay re-parks through the
existing machinery and the new connection gets its own full budget. Still
bounded per connection generation, which is what was documented all along.

Also pins the three sibling attacks on the same window: two panes in one
environment where only one handle lands, a handle published by a foreign
environment, and an environment tearing its rows down mid-park (which
leaves no waiter and no scheduled timer).

The test file now leads with how to assert on this module at all, because
the obvious shape cannot fail. "Did the waiter release" is not an
observable here -- a waiter released for the wrong reason is re-parked by
the replayed sweep, so the store reads identically one tick later, and a
mutation releasing every waiter on any tab's handle survived twelve
assertions written that way. What a spurious release costs is the
deadline, so the assertions advance the clock and require the pane to
decide on the ORIGINAL schedule.
2026-09-10 16:43:35 -07:00
Neil 798a37bcba merge: origin/main (fb85f88d64) into ssh-remote-integration-stack 2026-09-10 15:35:50 -07:00
Brennan BensonandMerge Sim fb85f88d64 fix(browser): restore the Chrome-shaped browser identity (STA-7147) (#19927)
* fix(browser): restore the Chrome-shaped browser identity (STA-7147)

#18749 replaced every browser partition's Chrome-shaped UA with Electron's stock
one, so since v1.4.198 the embedded browser announces itself on every non-Google
host as:

  Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like
  Gecko) Orca/1.4.198 Chrome/150.0.7871.224 Electron/43.4.1 Safari/537.36

No browser sends that. Sites that re-check the identity holding a session reject
it: users report being signed out of x.com, LinkedIn and "most websites," and at
least one was signed out of LinkedIn in their own Chrome and met LinkedIn's
"suspicious activity" SMS check -- server-side revocation, which reaches beyond
our app. The repo already documented the mechanism in browser-google-auth-ua.ts:
copied-in cookies "sent under a UA that doesn't match a real first-party browser
get flagged by anti-fraud." That is why the Google auth-host switch exists;
#18749 kept it for accounts.google.com and handed every other host an Electron
identity.

Restore the pre-#18749 session identity: strip the Electron and app tokens, and
rewrite sec-ch-ua to match. Nothing in the cookie-import write path changed --
it never did; cookies were always written correctly and servers were refusing
them.

Deliberately KEPT from #18749, all independent of the UA:
- anti-detection.ts stays deleted. Its premises were measured false on Electron
  43 and its overrides are themselves published bot signatures.
- No Runtime.enable into cross-origin iframes (the documented Cloudflare CDP tell).
- No unconditional CDP debugger attach on every browsing guest.

Known tradeoff, measured: this re-opens #13822. On the unmerged predecessor
branch brennan/sta-3905-cloudflare-ua, commit 9f0a4772fe recorded the stock UA
clearing dash.cloudflare.com 5/5 while every rewritten variant failed 12/12, and
noted that adding client hints does not rescue it. So Cloudflare-gated sites will
show verification failures again until a coherent-identity fix lands. That is a
bounded, in-app annoyance; session revocation damages users' real accounts. A
CDP Emulation.setUserAgentOverride with full userAgentMetadata -- which drives
navigator.userAgentData as well as the headers, and was never tested -- is the
candidate that could satisfy both, and is being measured separately.

Tests: the real-Electron wire-identity test now asserts the stripped identity on
ordinary hosts and Firefox on Google auth hosts. Ablation-verified: neutering
cleanElectronUserAgent turns it red on the Electron-token assertion. Its fixture
also gained an app name -- without one the raw UA carried no app token, so the
Orca/x.y.z half of the cleaner was never exercised.

* fix(browser): finish the identity revert in the files CI caught

browser-session-registry.persistence.test.ts still asserted #18749's behaviour
("keeps the stock UA", "keeps the engine UA"), so the shipped code and its test
disagreed. Caught by CI shard 4/8, not locally: I reverted four test files and
went to typecheck without re-running the browser suite.

Also restores the accurate wording that #18749 generalised away, now that the
behaviour it described is back:
- browser-google-auth-ua.ts: names the Electron/Chrome-shaped UA again as what
  anti-fraud flags, which is the reason the auth-host switch exists at all.
- docs/browser/profiles.mdx: documents the cleaned Chrome UA default and the
  --no-ua-spoof escape hatch, which is real again.
- tests/tools/google-signin-ua-probe.cjs: comments name the live handler.

Deliberately left at #18749's version, because those changes stay correct with
anti-detection.ts deleted:
- browser-manager-viewport.ts: its comment no longer cites the retired
  addScriptToEvaluateOnNewDocument injection.
- browser-webauthn-profile-delete.test.ts: its added webRequest mock is REQUIRED
  by the restored setupClientHintsOverride, so reverting it would break the test.

* fix(browser): keep restored UA hints browser-owned

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 15:20:34 -07:00
Neil c37a342b0e merge: PR #19870 into integration stack (substack 3/3) 2026-09-10 15:04:56 -07:00
Neil 9c137aeb2e merge: PR #19436 into integration stack (substack 2/3) 2026-09-10 15:04:50 -07:00
Neil 5ba7473a06 merge: PR #19659 into integration stack (substack 1/3) 2026-09-10 15:04:44 -07:00
Neil e8e705ef71 merge: PR #19736 into integration stack 2026-09-10 15:04:35 -07:00
Neil eaf4bf8871 merge: PR #19877 (agent-status repro harness doc) into integration stack
Conflict: .gitignore - both sides added a docs/reference allowlist entry at
the same spot; kept both in alphabetical order.
2026-09-10 15:04:30 -07:00
Neil 1c4fa5f520 merge: PR #19878 into integration stack 2026-09-10 15:04:00 -07:00
Neil 54ed30a8d6 merge: PR #19879 into integration stack 2026-09-10 15:03:53 -07:00
Neil b90c4bb01f merge: PR #19860 into integration stack 2026-09-10 15:03:47 -07:00
Neil e95ffa8c07 merge: PR #19882 into integration stack 2026-09-10 15:03:42 -07:00
Neil 2cd6466463 merge: PR #19861 into integration stack 2026-09-10 15:03:36 -07:00
Neil 3b254283db merge: PR #19865 into integration stack 2026-09-10 15:03:30 -07:00
Neil a010da5d71 merge: PR #19873 into integration stack 2026-09-10 15:03:24 -07:00
Neil 8afc97002b merge: PR #18513 into integration stack 2026-09-10 15:03:16 -07:00
Neil 1aa426aa48 merge: PR #19572 (ssh host partition adoption) into integration stack
Conflict: config/reliability-gates.jsonc - both sides appended a new gate
entry at the same array position; kept both.
2026-09-10 15:03:11 -07:00
Jinwoo Hong eb2f2d52ae feat(cloud): native push gateway and dedicated infrastructure (1/3) (#19912)
* refactor(cloud): share PostgreSQL schema startup between services

* feat(cloud): add durable native push notification gateway

* infra(push): define dedicated gateway resources and operational checks

* fix(push): bound cross-host admission and simplify gateway configuration

* fix(push): validate deploy configuration and preserve topic-error registrations
2026-09-10 17:59:46 -04:00
Brennan BensonandMerge Sim 33436c30d8 refactor(native-chat): unify agent session launch and open drafts in structured chat (#19681)
* wip(native-chat): first-pass draft routing into structured chat (to be reworked)

* refactor(native-chat): gather agent launch route inputs in one builder

Every launch entrypoint assembled the route resolver's inputs by hand and
they disagreed: only three of seven passed the project runtime blocker, so
a WSL-pinned project was refused structured chat from the tab bar but
admitted from the create dialogs. buildAgentLaunchRouteInput is now the
one place that gathers host, capabilities, workspace kind, project runtime
and TUI customization, and works for workspaces that do not exist yet.

Also deletes the dead draft-prompt blocker from the shared resolver; the
renderer stopped passing it and the main process never did.

* refactor(native-chat): share one structured launch settle loop

Five entrypoints copied the same loop around startStructuredAgentLaunch:
start, claim a refusal fallback, await, branch on refusal or unknown. The
copies drifted: direct work-item and full create reported an unexpected
launch error as success, and resume handled neither refusal nor unknown.

settleStructuredAgentLaunch now owns that loop and returns one settlement
(structured, refused-then-legacy, cancelled, visibility-unknown, failed).
Direct work-item, full create, folder workspace, both onboarding folder
paths and vault resume consume it; each keeps only its own legacy fallback.
Resume deliberately has no fallback. Unknown outcomes release the caller
uniformly so a stale fallback closure cannot fire on a later reconcile.

* refactor(native-chat): route the new-tab launcher through the shared settle loop

The new-tab launcher fired its refusal fallback and forgot it: nobody
learned whether the terminal fallback ran, and a visibility-unknown outcome
was never surfaced. Its structured branch now runs through
settleStructuredAgentLaunch with the terminal launch as the legacy fallback.
launchAgentInNewTab stays synchronous; the result gains a structuredSettlement
promise, and promptDeliveryResult keeps following the terminal fallback's
delivery on refusal as it did through the callers bridge before.

* refactor(native-chat): one legacy prompt delivery path and one trust preflight

The direct work-item flow kept its own seed-and-paste copy of the legacy
prompt delivery; it now uses deliverLaunchPromptToAgentTab with its own
timeout notice supplied as a callback. Three private copies of the trust
preflight (session continuation, worktree creation, folder workspace) fold
onto preflightAgentTrust. The direct work-item pre-launch mark keeps its own
entry because it differs in timing, not mechanism.

* refactor(native-chat): run quick create through the shared settle loop

Quick create was the last entrypoint driving the launch handle itself,
because its cancel lifecycle is real: when the creation is abandoned the
structured launch must be cancelled immediately so a staged prompt never
reaches the provider. The shared loop now takes a cancellation hook with an
eager subscription plus a post-await check; it cancels the launch once,
unsubscribes on settle, and reports cancelled without running the fallback.
Quick create keeps its two-branch legacy fallback and retire-on-late-cancel.

Also updates the surface-caller census for the onboarding launch module
that step 2 introduced.

* fix(native-chat): open editable drafts in structured chat for eligible local Codex launches

Route order asked the default-view-mode question first, and that decider
applies the terminal mirror gate (a TUI cannot clear more than forty lines
of prefilled draft), so a PR body over forty lines reached the plain
terminal before structured eligibility was checked. Structured eligibility
now comes first; the mirror gate applies only on the legacy branch.

The structured draft seed writes the launch-draft store directly with no
mirror gate, since a structured session has no terminal copy to fall back
on. Closing a settled structured tab clears an unadopted seed. The
structured session treats idle and loading as unsettled so the adoption
hook takes its baseline from the loaded transcript. Each caller passes one
delivery-mode value to both the route builder and the settle loop.

The structured session component test is split with a shared harness so
it stays under the test file line cap.

* test(native-chat): make the structured session test harness type-portable

* fix(native-chat): close review gaps in the shared launch settle loop

- Claim a refusal fallback only when the caller supplies one, so vault
  resume no longer reports a terminal fallback it never opened.
- A failed or cancelled direct work-item launch returns no tab id, so the
  caller never pastes the prompt into a setup shell.
- Terminal fork activates with providesInitialSurface for structured
  launches and gates its toast on the settlement; the draft blocker
  deletion made fork route structured too.
- A failed launch clears its draft seed. The failure toast moves to its own
  module to keep the launch-state file under the line cap.
- Ratchet for settle-loop callers; cancel-during-fallback documented.
- Restore the local agent label lookup that the pane-agent identity
  inventory expects instead of the inventoried helper.

* fix(native-chat): resolve the agent label through one module

* fix(terminal-pane): keep the fork dialog from reopening a created worktree

A failed or unknown structured settlement returned false after the fork
worktree already existed, so the dialog stayed open and a second click
created another worktree. Unknown now closes the dialog (the launch badge
already reports it); failed copies the context the way a null launch does.

* chore: restore pnpm-lock.yaml to main (local pnpm rewrite slipped into a commit)

* test(native-chat): stop asserting the deleted draft feasibility input

The routing-authority test expected the shared predicate to receive
isDraftPrompt; delivery mode is prompt metadata and never reaches
feasibility now, so assert its absence instead.

* refactor(native-chat): decide every agent launch route in one planner

The route was still resolved at seven callers, each also calling the settle
loop; two census tests only stopped an eighth. planAgentSessionLaunch is now
the one production caller of the resolver and its launch() the one caller of
the settle loop, and both censuses pin exactly that file.

The funnel is two-phase because three sites need the route before the
workspace exists and quick create persists its request for recovery: a plan
exposes route before creation and launches with the created worktree id;
a persisted quick-create request carries the verdict as data and re-enters
through adoptAgentSessionLaunchVerdict without re-resolving. Delivery mode
is fixed on the request once, so route and launch cannot disagree.

* test(native-chat): pin the two adopters of a planned launch verdict

* fix(native-chat): answer route readability from the repo when the worktree row is absent

The planner's transcript-readability input dropped the repo-level connection
fallback the direct work-item path still computes for its startup payload, so a
route planned in the window right after workspace creation saw `undefined` —
which reads as "not locally readable" — and downgraded grok/omp launches from
native chat to a raw terminal. Only `undefined` ("cannot determine the host")
now defers to the repo; a resolved `null` stays the local answer.

* refactor(native-chat): answer structured feasibility with a query, not a launch plan

Every rendered AI Vault row built a whole launch plan — execution-host lookup,
project-runtime resolution, capability read, plus a plan object and a launch
closure it threw away — to read one boolean off it. Feasibility and a launch
decision are different operations, so the planner now exports the predicate for
the first and keeps the plan for the second, and the census pins the query's
callers separately. Settings arrive by argument, which makes the AI Vault
callback's dependency on them real rather than a comment the linter contradicts.

The plan's `explicitStructured` branch had that gate as its only caller and goes
with it; the vault's launch already re-enters on an adopted verdict.

* refactor(terminal-pane): fold the fork's trust preflight onto the canonical one

`preflightForkAgentTrust` was a behavioural duplicate of `preflightAgentTrust`,
whose signature now accepts a nullable agent and workspace path and so is a
drop-in replacement. Its file is left holding only the launch-platform resolver
— which is not a duplicate, since it returns an override rather than a default —
so the file is renamed for what it now contains.

* refactor(native-chat): cancel a structured launch through an AbortSignal

The settle loop's launch cancellation re-derived the standard poll-plus-eager-
event primitive that `AbortSignal` already is, so it now takes one. The eager
semantics are unchanged: the loop still cancels on the abort event rather than
only polling after awaits, so a staged prompt is discarded before it reaches the
provider, and it drops its listener on settle instead of leaving the signal
holding the closure. Quick create owns the controller and bridges its store
subscription to it.

A cancel that lands after the refusal fallback already opened a terminal now
carries that surface on the settlement. It is the fallback's tab that exists, so
reporting the pre-launch one handed the caller a workspace with no agent in it.

* fix(native-chat): tighten quick create's structured launch settle path

Four things the launch path got wrong once the settle loop owned the flow:

- The abandoned-creation check now runs before the first-message rename flag is
  written, so a creation being torn down is no longer marked for a rename that
  will never happen (the order the pre-planner code had).
- A cancel that arrives after the refusal fallback opened its terminal reports
  that terminal rather than the pre-launch tab.
- `plan.launch` is called outside the caller's try, and nothing awaits that
  caller, so a throw there would strand the creation panel. It is now caught and
  reported the way a failed launch already is.
- The launch route is a required argument instead of defaulting to
  `terminal-tui`, which would have silently reported success with no surface
  opened. Both callers already gate on the structured route.

* fix(native-chat): give one launch identity one prompt delivery mode

A caller joining a pending launch computed its outbox text from its own delivery
mode, so an auto-submit caller landing on a draft launch enqueued text the first
caller's seed was already showing in the composer: the user saw it and it was
sent. The mode is now fixed by the caller that opened the launch, and a joiner
delivers its text that way.

Seeding also moved to where the coalesce decision is made, so a launch whose
callers already settled as refused is not given a fresh draft — the refusal path
early-returns, so nothing would ever clear it and it would outlive every tab.

* fix(work-item): report a failed structured launch as a failed direct launch

`launchWorkItemDirect` returned true unconditionally, so a structured launch
that opened no surface still read as a started workspace. Callers hang
irreversible follow-up work off that boolean — the fix-checks dialog fires
`onLaunched` on it, which is documented as the home for host writes — so a
launch with no agent tab now reports false, matching what full create does.

The settle result says so explicitly rather than leaving callers to infer it
from a null tab id, which `notLaunched` also produces.

* test(session-tabs): pin the id a first structured publication is minted under

The launch draft seed is keyed on `structuredAgentSessionTabId(sessionId)`
before the tab exists, while the mirror mints ids with collision avoidance that
can append a `:history-N` suffix. The two agree today only because a fresh
session's base id is unique. Pin that where the id is actually minted, with the
collision arm alongside it so the divergence the seed depends on staying away is
visible rather than assumed.

* test(native-chat): pin the route connection fallback on the un-mocked resolver

The suite that covers the builder stages `getConnectionIdFromState`, so it can
characterize the fallback but cannot catch a defect that lives in owner
resolution itself. This one runs the real resolution over real store rows: two
repos publishing the same worktree id on different hosts, which is the
documented case where the owner cannot be named and `undefined` is returned.
Red with both fix files at the previous head, green with them.

Reverts the two caller pins added to the route census — the feasibility
predicate is exported from the planner, which the census already permits, so it
passes unedited and needs no permit clause.

* fix(native-chat): keep the structured launch's own agent eligibility check

Quick create's structured launch narrowed its guard to a bare `agent` presence
check, so a creation carrying an agent that cannot hold a structured session
reported itself cancelled once dismissed, where it previously reported that it
had done nothing. Unreachable through both callers today, but it is the last
local eligibility check in a module that otherwise trusts its callers for the
route, so it is restored rather than left to the required-route typing — which
says nothing about the agent.

Also corrects two comments that called the quick-create request "persisted".
It lives in renderer session memory and dies with the renderer; calling it
persisted made the plan/adopt split read as restart recovery, when what it
actually buys is a route decided before the worktree exists.

* fix(native-chat): keep the structured feasibility query typecheck-clean

The query threaded its narrow settings through the store, but the route
store's settings must satisfy the full GlobalSettings that two of its
resolvers require, so the narrow copy never fit. Ride the named settings
on the built input instead: the caller still names them, so a React memo
still depends on them, and no store-shaped object is needed.

Also give the launch state its delivery mode unconditionally; the key is
required, and a conditional spread makes it optional under
exactOptionalPropertyTypes.

* docs(native-chat): name the feasibility query's one remaining settings asymmetry

The builder reads launch customization off the store while the routing gate
reads the named settings, so one answer has two settings sources. It cannot
diverge with the single caller passing the object the store already holds, but a
PR about removing split sources should not leave that unstated.

* fix(native-chat): keep a coalesced joiner's draft unsent

joinLaunchDelivery stripped the joiner's delivery mode when the launch it
joined had established none, and an absent mode reads as submit. A joiner
that asked for a draft therefore had its text sent — the send-without-
consent this PR exists to prevent. Fall back to the joiner's own mode only
when nothing was established, so the first caller still wins otherwise.

* chore: re-trigger CI

GitHub created no workflow run for e935ea5e42 — the pull_request
synchronize event was dropped. No content change.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 14:54:09 -07:00
Brennan BensonandMerge Sim 2626e2eca4 Make the structured turn lifecycle row durable so completed durations survive (#19695)
* Make the structured turn lifecycle row durable so completed durations survive

A structured-chat turn used to end by tombstoning its running lifecycle item,
which threw away the only durable record of when the turn ended. Completed
"Worked for" labels therefore depended on the renderer having observed the
turn finish, and vanished on reopen.

The lifecycle item is now revised in place, never tombstoned:
- running, with startedAt, at the provider's turn start
- completed or interrupted, with completedAt, at the provider's terminal frame,
  a user stop, or a child exit the host observed
- unverifiable, with no end, when a cold acquire finds a running row from a
  generation whose exit nobody observed

Both timestamps are the execution host's clock at receipt, captured before the
deferred sink, so the completed value is identical on every client and needs
no client clock. Codex history restore uses the provider's own second-granular
endpoints for turns that predate this change. Desktop and mobile read settled
durations off the journal through one shared selector, and anchor the live
counter on the host start with the client's local receipt so a skewed client
clock never leaks into the label. Locally observed durations remain the
fallback for hosts that still tombstone.

Timestamps live inside the existing turnLifecycle field, which old clients
strip, and every working-state consumer keys on state === 'running', so no
capability negotiation is needed.

* native-chat: avoid stale working status on settled turns

* test: align settled turn status expectations

* Name settled lifecycle rows by their terminal state

An interrupted or unverifiable turn must not read as completed for any
consumer that renders status text raw. One shared helper builds the text for
both providers from the lifecycle state.

* test: deduplicate turn lifecycle suites

Each behavior keeps one test; duplicated harnesses and restated cases go.

* Key lifecycle rows to their user item and record the provider's measured duration

A lifecycle row now names the user item that opened the turn by its provider
key, so clients attribute timing explicitly and fall back to journal order
only for rows from older hosts. A provider-initiated turn with no prompt can
no longer claim the previous prompt's duration.

When the provider measures the turn itself (Codex turn.durationMs, Claude
result.duration_ms) the terminal row records it and clients prefer it over the
host interval, so a turn shows the same number live and after a history
restore. Host receipt times remain the live-counter anchor and the fallback.

* Record a turn as a first-class journal item

The turn record is now its own item kind rather than a status row carrying a
lifecycle field: no text to misuse, and the fold matches the durable turn
record other systems keep. Rows that carry it are stamped journal schema v3;
every other row stays v2, so an older host keeps reading them and latches
read-only at the first v3 row instead of truncating the epoch.

Clients that predate the item would paint an unknown kind as a text bubble,
so the host publishes the legacy status form to any client that does not
advertise agent-session.turn-item.v1, through the same per-client seam
background tasks use. The downgrade is transitional and goes once no
supported release lacks the capability. The shared projection now renders
unknown item kinds as nothing, so later kinds need no gate. One shared reader
handles both forms for old journals and old hosts.

* Preserve observed turn end across settlement retries

* Retain turn attribution for loaded chat history

* Preserve Codex exit receipt across close retries

* Register completed turn duration reliability gate

* Keep earlier turns through a Codex rewind and count a mid-turn attach from the real start

Findings from an independent adversarial review of the typed turn record:

- A Codex rewind adopted the provider's item list as the new epoch, and the
  provider never returns the host's own turn rows, so every duration before
  the rewind point vanished. The host's turn rows are now spliced back beside
  the item each followed, and recovery no longer expects the provider to
  prove rows it never owned.
- The epoch row was stamped with the current schema version, so an older host
  latched read-only at row 1 of every new session, defeating the mixed
  version design. It carries no body and stays at v2; a stored-row test now
  reads SQLite directly, because the reader upcasts every row on read.
- A send Codex folds into a running turn shares the opening prompt's provider
  key, and the alias map credited the duration to the later prompt. The
  earliest submission naming a key now wins.
- The live counter anchored on first sight, so a client attaching mid-turn
  counted from zero. Published frames now carry the host's clock, the reducer
  keeps the last sample with its local receipt time, and both clients anchor
  on how long the host says the turn has run.

* Correct turn duration gate assertion reference

* Respect authoritative unknown native chat duration

* Preserve unverifiable timing across older host upgrade

* Record final completed turn duration reliability evidence

* Fix the CI failures the merge left behind

- A merged import list named the same module twice, which the native code
  quality plugin fails on.
- A running turn is now reported by the host with no duration, so the settled
  map carries an explicit null for it; the hook test still expected the entry
  to be absent.
- main gave the older-page action a cursor with a head-trim guard, so the
  retention test's epoch-only action no longer typechecks; it now passes an
  unbounded sequence, which is what the old shape meant.
- The roster comparator moved into the extracted module, leaving its import
  unused in the reducer.

* Split two files back under the line cap after the merge

Merging main put both one effective line over 300, and the cap forbids a
disable or a shave. The wire module's refusal vocabulary moves to its own file
and is re-exported, so its consumers are untouched; the host's four thin
mutation delegates move next to the functions they call.

* Advertise the turn-item capability on every client transport

Local IPC and mobile advertised it; the remote and web transports did not, so a
desktop paired to a remote host, the CLI, and web silently ran on the legacy
carrier forever and the canonical row was never exercised there. The renderer
that paints it is the same build on every transport.

* Update the web auth-frame expectation for the new capability

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 14:32:50 -07:00
Brennan BensonandMerge Sim 721a269289 test(native-chat): split structured question fixtures (#19924)
Co-authored-by: Merge Sim <sim@local>
2026-09-10 13:35:15 -07:00
Merge Sim 4f5a8275e8 Revert "test(native-chat): split structured question fixtures"
This reverts commit 68e207ca2f.
2026-09-10 13:26:08 -07:00
Merge Sim 68e207ca2f test(native-chat): split structured question fixtures 2026-09-10 13:23:20 -07:00
Brennan BensonandMerge Sim 6e9de5fa58 fix(orchestration): revalidate an attempted Enter instead of resending it (#19911)
When a PTY retires mid-delivery, every staged message was marked undelivered,
which made all of them redeliverable. That is right for a pointer whose Enter
never fired, but an Enter that was already written may have landed: redelivering
it types the same mail into the pane a second time.

The Enter timer is cleared at the top of retirement, so a RESERVED or
WRITE_ATTEMPTED pointer provably never submitted and is released. An
ENTER_ATTEMPTED pointer is ambiguous and now stays at its phase for the resume
path to revalidate, matching the policy mailbox-pointer-submit.ts already
documents for an unverifiable settlement.

Co-authored-by: Merge Sim <sim@local>
2026-09-10 13:19:20 -07:00
Jinwoo Hong a6e6de93c4 fix(relay): keep failed rehome polls out of the durable failure budget (#19915)
* fix(relay): keep failed rehome polls out of the durable failure budget

The regional rehome worker polls claimRegionalRehome about once a second.
Any error thrown before an attempt was claimed - in practice a director pool
timeout on the pre-claim control read, 52-74 a day against a pool of 3 - was
charged to relay_region_rehome_worker_state.consecutive_failures, which
durably disables the control at three. That counter only ever resets on a
drain receipt, so while the control is disabled it never resets: production
sits at 1068 and still climbing. Enabling the control leaves the stale
counter in place, so the next pool timeout latches it straight back off.
That is what ended the 2026-08-28 enable after ten minutes.

- A poll that never claimed an attempt drained nothing, so it no longer feeds
  the dispatch-failure budget and logs .._poll_failed instead of
  .._dispatch_failed. recordRegionalRehomeWorkerFailure had no other caller
  and is removed.
- Enabling the control clears consecutive_failures and paused_until, so a
  budget spent under a previous enable cannot kill a fresh one. The dispatch
  interval in next_dispatch_at is deliberately left alone.
- The budget's auto-disable now emits
  orca_relay_regional_rehome_failure_budget_disabled, matching the existing
  .._safety_disabled precedent. It wrote no event before, which is why this
  went unnoticed for two weeks.

No change to region selection, the candidate query, or host eligibility.

* fix(relay): serialize rehome failure accounting with control updates
2026-09-10 16:04:25 -04:00
Brennan BensonandMerge Sim 5a96158849 feat(native-chat): focus the message box when a chat appears (#19868)
* feat(native-chat): focus the message box when a chat appears

Opening a native chat left focus nowhere, so you had to click the
composer before typing. Nothing in the chat surface focused it on open;
the only existing focus calls were reactive (typing on the bridge pane
background, picker acceptance, attachments, dictation), and the
structured pane had none of those.

useNativeChatComposerRevealFocus focuses the composer on the reveal
edge, covering a new chat tab, a worktree-create landing in chat, the
chat-view toggle, and switching back to an existing chat tab. Mount is
the wrong signal: retained panes hide with display:none + inert and
never unmount on a tab switch. It reuses the existing composer handle
and shouldPreserveEditableFocus rather than adding a parallel path, and
retries across a bounded run of frames because Tiptap publishes its
adapter after mount and Radix restores a closing dialog's trigger in a
setTimeout(0).

Two supporting changes:

- isFocusedGroup, from activeGroupIdByWorktree. On worktree activation
  both columns of a split flip visible in the same commit, so without it
  two revealed chats fight over the caret. The bridge route already had
  this bit as controller.isActive; only the structured overlay needed it.

- focusRuntimeTerminalSurface bails on a chat-covered pane. Its DOM-path
  twin already declines chat view via data-terminal-chat-view, but the
  runtime path focused the covered xterm unconditionally and pulled the
  caret out of the composer. Returns true, not false: false sends the
  caller to the DOM fallback, which for a structured tab id focuses an
  unrelated tab's xterm.

* fix(native-chat): preserve reveal focus ownership

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 12:38:06 -07:00
Brennan BensonandMerge Sim 4408fe897a feat(sidebar): show native-chat subagents as sidebar child rows, like CLI agents already do (#19807)
* feat(sidebar): indent native-chat subagents under their session row

Stacked on #19311, which adds the background-task channel this reads. The
bridge maps agent-kind background tasks into AgentStatusEntry.subagents, and
the renderer status feed confirms per connection so a reconnect cannot leave a
child asserting live from a stream that ended.

* fix(sidebar): avoid completed age for unverifiable subagents

* fix(sidebar): preserve unverifiable child verdicts

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 12:29:30 -07:00
f2af92b2fa feat(native-chat): show live background work and name each row by kind (#19705)
* fix(codex): reserve the label's share of a qualified command row

A child's label is raw provider text and was spliced into the command
row unbounded, then the pair clipped to the description cap. A label at
or past that cap clipped the command away entirely, leaving a row of
kind 'command' that named an agent and showed no command - the failure
qualification exists to remove, inverted. The same clip could also cut a
surrogate pair, which boundSubagentField already guards against on the
agent row two lines away.

Give the label a reserved share and clip it the way the agent row does.

* feat(native-chat): show live background work and name each row by kind

The strip suppressed itself in three places: the Claude tracker blanked
its roster for the whole of any turn, the Codex tracker returned nothing
while a primary turn was open, and the renderer view gated on
`turnId === null`. Between them, work in flight was never shown — and a
task backgrounded in an earlier turn vanished from the strip as soon as
the next prompt was sent. Claude additionally dropped every foreground
subagent, so a fan-out reported nothing at all.

Report work while it is live, in all three layers. Foreground Claude
work is turn-scoped, so `result` retires it — that is the provider's own
outcome for a task it marked foreground, not a roster sweep. Nothing
settles a Codex child on turn end: those keep reporting well past their
parent, so turn frames only prompt a republish.

Name each ROW by kind — Subagent, Shell command, Workflow, Monitor —
instead of a generic "Background <kind>", each drawing the glyph the
shared tool-icon table already uses for that category. A row that
carries a provider description still shows it unchanged. The collapsed
header summary is deliberately untouched; it is owned elsewhere.

The conversation-command gate is unchanged in effect: an open turn
already refuses first, and Claude foreground work never reaches the
backgrounded set the gate reads.

* fix(native-chat): withhold the row stop Claude foreground work cannot honour

The strip now publishes foreground rows, but `stoppableTaskIds` still filters
on `backgrounded`, so `stopClaudeBackgroundTasks` resolved an empty target list
and returned `{ cancelled: false }` that no renderer reads: the user clicked
"Stop Subagent" and nothing ever happened.

Carry stoppability per row instead of widening the stop to a target the SDK has
no way to reach. `AgentSessionBackgroundTask.stoppable` is absent-means-yes, so
hosts that predate it keep their working control, Claude emits `false` only on
foreground rows, and the strip hides that row's button the same way it already
hides the stop-all a provider cannot honour.

* fix(claude): scope aggregate-roster authority to the work it enumerates

`background_tasks_changed` lists BACKGROUNDED tasks, so a foreground subagent
can never appear in it. Treating it as the whole world meant any such frame
cleared every live foreground row mid-flight and then dropped every later
foreground `task_started` for the rest of the session, killing the in-turn
fan-out the strip exists to show in any session that ever backgrounds anything.

Decide `backgrounded` before the staleness guard and apply the guard only to a
backgrounded start, and retain live foreground entries across a roster replace.
Retained rows count against MAX_TRACKED_TASKS, so the map stays bounded, and a
stale backgrounded start the roster no longer lists is still dropped.

* test(native-chat): pin the strip's monitor amber to the constant that defines it

`MONITOR_GLYPH_COLOR`'s comment claimed a test held it and AgentStateDot's amber
together, but no test imported it — the assertions hardcoded 'text-yellow-500',
so the two could drift with every test still green. Read the colour from the
module, which is what the comment always said was happening. Drop the unused
`BackgroundTaskGlyph` export too: nothing outside the module names it.

* fix(native-chat): keep the task list open across a gap in live work

The strip is now mounted on live work, so a sequential fan-out unmounts it
between one subagent finishing and the next starting: local `useState` meant
the expanded list collapsed itself on every such gap, on top of the strip
flickering above the composer.

Hand the disclosure to the session, keyed by session id so it does not leak
across a session switch. The strip is now controlled and holds no state of its
own, which is what makes it survive its own mount churn.

* fix(codex): route every command-row cut through one surrogate-safe clip

`boundLabel` avoided splitting a pair, then `qualifiedDescription` re-cut the
COMPOSED string with a raw slice: label (<=96) plus separator plus description
(<=512) is up to 611 chars, so that second cut landed at an arbitrary index
inside the description and could publish a lone high surrogate — lossy through
any non-JSON UTF-8 hop. `parse` had the identical hazard on an unqualified
primary-thread command.

One `boundText` helper now owns all three cuts, so no path in the file can emit
a lone surrogate from well-formed input.

* fix(claude): keep terminal evidence for ids an aggregate roster never lists

Narrowing the admission guard to backgrounded starts left a finished FOREGROUND
id with no defence: `replaceAggregateRoster` wiped `terminalTaskIds` wholesale,
so after any `background_tasks_changed` a replayed `task_started` revived a task
whose completion had already been seen — and only a later `result` could settle
it again.

Scope the wipe the same way the guard was scoped: delete only the ids the
incoming roster actually enumerates. A roster still overrules terminal evidence
for the work it lists, which is what that behaviour was added for.

* fix(claude): keep retained rows in place and evict the stalest, not the newest

Re-adding retained foreground entries after the roster made a live row the user
is reading jump below the backgrounded rows on every `background_tasks_changed`,
and the cap `break` kept the STALEST retained rows while dropping the newest.

Merge in the tracked map's own order so a surviving row holds its position, and
count the overflow up front so eviction takes the oldest retained rows. Roster
entries are never starved and the map stays bounded either way.

* fix(claude): retire leftover foreground rows when the next turn starts

A foreground `task_started` arriving with no turn open has no `result` coming
to retire it, so it sat in the strip indefinitely — with no per-row stop, since
foreground rows are not stoppable — and refused conversation commands behind an
instruction nobody could follow.

Settle on turn start as well as on `result`. This is cleanup only: visibility
never consults `startsTurn`, so a missed one degrades to today's behaviour and
can never switch the feature off. It shortens the row's life to the next turn;
the case where no further turn is ever sent is filed separately.

* fix(agent-session): withhold unstoppable rows from readers that predate them

Rule 3 of remote-wire-compatibility: changing what the host publishes reaches
old clients with no wire change. The Claude host published no foreground rows
before this feature; it does now, and a client that cannot read `stoppable`
draws a per-row Stop on every one of them — Claude always sets
`supportsTaskStop` — which filters to the backgrounded ids, stops nothing, and
returns a result no renderer inspects. That is the dead button `stoppable` was
added to remove, reappearing across a version skew.

Negotiate it. A client can advertise the existing background-task-stop
capability and still predate `stoppable`, so this needs its own constant.
Readers that do not advertise it get unstoppable rows dropped, and a state whose
every row is dropped becomes no strip — exactly their pre-feature view.

RUNTIME_PROTOCOL_VERSION is not bumped: this adds an optional field and a new
negotiated capability, and changes no existing field's meaning, which is the
explicit do-not-bump case in protocol-version.ts.

* test(agent-session): name the projected rows so the fixture typechecks

An indexed lookup into the fixture's task list is possibly-undefined under
`pnpm tc`; the rows are more readable named anyway.

* test(web): advertise the row-stop capability in the e2ee auth expectation

The web e2ee handshake started sending
AGENT_SESSION_BACKGROUND_TASK_ROW_STOP_CAPABILITY, and this test asserts the
advertised list by deep equality, so it went red on CI while every targeted
test run stayed green. Add the capability in the position the router sends it.

* test(claude): pin why the roster empties mid-turn in a sequential fan-out

The strip unmounting between two sequential subagents is truthful, not a swept
row: A leaves on the provider's own terminal frame, B does not exist yet, and
backgrounded work spanning the same gap holds the roster open — so an empty
roster is never work the strip is hiding.

Also pins the previous-turn rule against the one the subagent roster already
applies on the same frame: a still-working FOREGROUND child becomes
`unverifiable` there and a backgrounded one is left alone, so the strip drops
the first and keeps the second rather than asserting `live` for either.

---------

Co-authored-by: Merge Sim <merge-sim@users.noreply.github.com>
Co-authored-by: Merge Sim <sim@local>
2026-09-10 12:16:30 -07:00