Files
orca/mobile/src
Jinwoo Hong ceafdcad2f perf(mobile): race the direct and relay dials from t=0 on every reconnect (#19308)
* perf(mobile): race the direct and relay dials from t=0 on every reconnect

A foreground reconnect gave the direct dial a fixed 2.5s head start, and while
that dial sat in 'connecting'/'handshaking' the supervisor refused to open a
relay socket at all. A phone that is off the LAN paid the full head start on
every reconnect and got nothing for it, and a phone whose relay dropped could
only return to the LAN through three hysteresis probes.

Both dials now start together and the first authenticated socket is adopted
through the existing migrateTo cutover. Nothing about the migration machinery
changes: only who is allowed to start a dial.

- The relay dial now yields to a live session and to nothing else. An unfinished
  direct dial is progress on the other runner, not a reason to stand still.
- The direct return probe grows a second adoption policy. Against a live relay
  hysteresis still has to prove direct stable; during a reconnect there is no
  session to protect, so an authenticated direct socket wins outright. probeNow
  pre-empts a pending 15s tick so that dial starts with the relay dial, not
  after it, and the dial itself no longer waits for the operation mutex — a
  relay dial holding it is exactly the case the race exists for.
- A loser closes and books nothing. The relay dial withdraws inside migrateTo
  and returns 'aborted', so no backoff is booked against it; a direct socket
  that loses leaves the promotion streak untouched. Only a reconnect that both
  paths lose books a failure, once, on the relay cadence.

Kept: the 30s background grace and the foreground gate, because a backgrounded
phone must not open a billed relay splice; the shared failure cooldown, because
a genuine relay failure still has to be paced; the hysteresis dwell after a
migration, because it is what stops a marginal LAN flapping a healthy session.

The accepted cost is one relay socket per reconnect for a phone that is on its
LAN. It closes as soon as the direct path authenticates, before the resume
confirm, because migrateTo only checks the abort predicate after E2EE auth.

Tests that encoded the removed rules:
- 'fails over when the direct retry loop publishes reconnecting' asserted no
  relay dial while direct was handshaking. The failover now precedes the direct
  client giving up, so it asserts the dial instead of its absence.
- 'does not spend a queued relay retry while direct authentication is
  progressing' encoded the block outright; it now asserts the retry runs on the
  failure cadence while a handshake drags on.
- The four grace-race cases move to mobile-endpoint-reconnect-race.test.ts as
  t=0, direct-wins, background/resume and both-lose cases.
- Five relay-bookkeeping cases now state their premise with unreachableDirect.
  They describe a phone with no LAN, which used to be implicit and is now
  load-bearing: with a reachable LAN the direct socket wins those reconnects.

* fix(mobile): withdraw a lost relay dial pre-handshake and damp blip races

Review follow-up to 4e31130471. Racing both paths from t=0 was correct but
charged the LAN case twice: once per reconnect in cell work, and again whenever
the LAN flapped.

Withdraw before the handshake. migrateTo only consults its abort predicate after
E2EE authentication, so a dial that had already lost still made the cell reserve
a splice and the desktop finish a key exchange. The establisher now watches the
logical client across the dial and closes the cell socket the moment direct
authenticates. In the common window, after relay-auth is on the wire and before
the hello lands, nothing of the key exchange has started, so the withdrawal
costs the desktop nothing. The dial still reports itself aborted and still books
nothing. The watch is dropped once migrateTo returns, because past the cutover
this session is the active path and a later direct promotion must not read as a
reason to close the client's own socket.

Damp races that a blip started. relayDialAllowed yields only to a live session
and a lost race books nothing, so a flapping LAN drove one cell socket per blip
with only the relay's per-host rate limiter as a backstop, and reaching that
limiter would have converted a benign race into a booked relay failure. After a
race is lost to direct, the next unforced race is suppressed for 2s, doubling
per consecutive loss to a 30s cap. This is not backoff and is kept separate from
it: a forced replacement is never damped, a relay dial that wins clears the
streak, and a foreground resume clears it too, so the path the user is watching
never waits. The window arms its own lapse timer, so a LAN that dies inside the
window still reaches relay without a new trigger.

A superseded cutover no longer escapes probe() as an unhandled rejection. Only
the probe timer calls it, and it discards the promise, so the routine end of a
lost race would have surfaced as one.

Credential rotation moves to MobileRelayCredentialRefresh. The supervisor
crossed the 300-line cap; rotation is a self-contained responsibility that only
runs over a live direct connection, so it splits cleanly instead of taking a
max-lines bump.

* fix(mobile): end a damper window as soon as the direct path is really gone

Round-2 review follow-up to 9a21da4326. The damper armed its window when direct
won the race, and nothing shortened it. A LAN that died inside that window left
the phone waiting out the whole thing, up to 30s at the cap, with only a log
line to show for it. My previous commit body claimed the path the user watches
never waits; that was true only of a foreground resume, and it is corrected
here.

Losing the direct path now collapses the wait to a 250ms floor, so the next
recovery races almost at once. The floor is not zero because the reason the
damper exists is a LAN that drops and comes straight back, and a disconnect is
how such a blip begins. So the rest of the window is kept aside rather than
spent: if direct returns inside the floor it was a blip and the window resumes,
and if the floor lapses with direct still gone it was an outage and the held
window is void. Without the second half, one blip would have bought a flapping
LAN a free pass on every race that followed, which is the case the damper was
added for.

The streak itself is untouched by the clamp. A LAN that flaps all afternoon
still escalates toward the cap; only the current wait is cut short.

record() now takes the same forceReplacement guard as suppresses(), so a forced
replacement that stands down cannot grow the streak or be read as a loss to
direct. A lease rotation or a reconsidered network change is not a LAN that
flapped.

Also documents that a genuine relay failure deliberately does not reset the
streak, and that the damper and the failure backoff serialize rather than stack:
a damped attempt never reaches the dial that would book a cooldown.

* fix(test): give the direct-probe fixture the race-era hooks

The phase-1 probe test predates canDial and adoptsOutright, so its hooks
literal threw at the first dial. These cases model a live relay session.

* docs(mobile): say why a finished credential refresh races relay instead of waiting on direct
2026-09-07 13:45:51 -04:00
..