mirror of
https://github.com/stablyai/orca.git
synced 2026-09-23 00:02:29 +00:00
* fix(mobile): retry the stored assignment when the director reports no newer move A director answering /v1/connect can only reply relay-moved with the stored assignment; it has no 'assignment unchanged' verb, and sticky assignments make equal-epoch replies the steady state. Treating every non-newer move as fatal made pairing recovery unwinnable for any transient cell dial failure (DRAINING, 1006), which bricked off-LAN pairing on Android 0.0.44. A non-newer move now confirms the stored assignment: the candidate re-dials it with a 250ms floor instead of abandoning the relay path. The move is never adopted or persisted, so the anti-rollback contract (requireStrictlyNewerEpoch for persisted moves) is unchanged. 4429 stays out of director recovery: each cell dial burns an invite attempt server-side and a director hop cannot relieve cell load. * fix(mobile): honor relay director Retry-After when pacing recovery Mobile /v1/resolve collapsed every non-OK status into a generic error and discarded Retry-After, so overloaded windows produced hammering instead of paced retries. RelayDirectorHttpError now carries status and retryAfterMs (clamped to 120s), and the reconnect controller floors its existing transport delay with it — no new timers or retry state. The Retry-After parser is extracted from the desktop relay client into src/shared and reused by both. * fix(mobile): attribute pairing log lines to their candidate path The pairing race interleaves the direct LAN and relay candidates into one PAIRING LOG pane; direct lines (WebSocket closed, Reconnecting 10.x.x.x:6768) carried no path label and repeatedly read as Relay retrying a private IP — misleading users and two investigations. The coordinator now wraps each candidate's sink with an idempotent Direct:/Relay: prefix at the one seam where both paths are known. * fix(relay): gate desktop /v1/assign at the per-host rate limit The director rate-limits /v1/assign per host at 5s, but every desktop retry path could fire immediately: both schedulers draw full jitter from [0, cap] (floor 0), first attempts after a drain are undelayed, the 400-fallbacks issue up to 3 assigns per round trip, and reconcile() cancels the armed Retry-After timer from ~8 refreshDemand callers. Production shows hosts permanently rejected at ~100-200 rejects per success. A shared per-host gate now lives inside requestRelayAssignment — the single assign call site — so every path books a >=5s (+jitter) slot. Retry-After raises the gate persistently, surviving the coordinator's timer cancellation. Concurrent callers serialize through a per-key chain. Callers with staleness fencing pass isCurrent; a superseded caller aborts after the wait instead of spending the host's slot. Internal 400-fallback retries stay one logical attempt and do not re-enter the gate. * refactor(mobile): rename the log-only assignment-echo predicate isCurrentAssignmentMove no longer gates control flow — every non-newer move retries the stored assignment — so the name overstated its role. * fix(relay): honor mid-wait raises and cap the assign gate's inline wait Review findings on the per-host assign gate: the deadline was read once before sleeping, so a sibling's Retry-After landing mid-wait was ignored (the exact storm the gate exists for), and the sleep was uncancellable — a booked five-minute Retry-After could park pairing IPC, which awaits reconcile inline, for its full duration. The wait now runs in 1s slices, re-reading the deadline and the caller's isCurrent fence each slice. Remaining waits beyond 15s fail fast as a RelayHttpError 429 carrying the remainder, so the existing schedulers pace with it while the gate keeps the deadline. Staleness aborts are classified non-retryable. Also from review: the broker's isCurrent wiring and the shared-gate default are now pinned by tests, the 4429 comment states the reservation-order rationale precisely, the mobile Retry-After ceiling is renamed to avoid colliding with the desktop's 5-minute one, and a past-HTTP-date header case is covered. * fix(relay): tag locally paced assigns and warn about frozen test clocks Review polish: the synthesized 429 for a beyond-cap local wait now carries a distinct message (relay_assignment_locally_paced_429) so log censuses can tell it from a real director 429, with the comment stating the invariant that makes the translation honest (local booking alone never exceeds ~5.5s). The gate's sleep option documents that test fakes must advance the clock — the slice loop re-reads it and never terminates against a frozen one. * fix(relay): fence superseded callers at the assign send boundary reserve() checks staleness while waiting, but a caller superseded after booking — or between the 400 field-fallback retries — could still spend one to two requests on an assignment nobody consumes. Re-check isCurrent at the top of sendRelayAssignment so the fallback recursion is fenced too.