The shell collapsed six transport states into four, and every page route
expanded them back. Both mappings are gone; the page reads the state the
transport reported.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Two classes and two Sets tracked the same thing: a page-minted id the broker
has already honoured. Request and subscription ids share one window now. It
stays a module rather than a broker field because the broker sits at its
max-lines ceiling.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The hand-written MobileWebSpeechRuntime interface described exactly what the
native factory returns, so the alias now comes from the factory. The startup
wake takes only the call it makes.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Neither module is a shell/page contract. The mermaid frame belongs with the
diagram component that renders it, and the editor script hash with the editor
document that declares it.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The OSC module existed only to be re-exported. Inlining it pushed the file
past max-lines, so the sequencing rules both sides apply to terminal frames
now live beside the frames rather than inside the schema module, and the four
copies of decodedBase64Length collapse to one.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Both schemas parsed an avatar URL and then transformed it away, so the page
only ever saw null. The page CSP allows img-src 'self' data: blob:, so a
remote avatar could never render there either. The strip() objects still
discard an incoming URL, which the contract test now asserts by absence.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The native views serve every mobile-web document through their own scheme
handler and set Content-Security-Policy as a response header, so the meta
copy the packager wrote was never the enforced policy.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Deleting the grant took the mode check with it, so `workspace.hostRequest`
would forward a subscribe method. The desktop answers that with `streaming:
true` frames, and `RpcStreamRegistry.handleResponse` claims every one of them
before the pending-request map sees it, so the unary promise hangs to its 15s
deadline while the desktop subscription stays open with nothing to cancel it.
On the relay transport the pending map wins instead and the promise settles on
a partial first frame, which is wrong in a different way.
The naming rule already answers this without a grant: a method has a derivable
cancel name exactly when it is a subscribe method. `prepareMobileWebHostRequest`
now returns that derivation, `hostRequest` refuses a method that has one, and
`hostSubscribe` refuses a method that does not. Both raise
`unsupported_capability`, the code the grant-mode mismatch used.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The RPC deletion removed a check the native app really made. On the hosted page
readability was constant true, but the native shell resolved the workspace's
repo and hid grok and omp tabs whose transcripts the runtime cannot read, which
is every repo behind a non-runtime-owned SSH connection. Deleting it changed
released native behavior.
Restore the gate on the native path only: `isMobileNativeChatTranscriptReadable`
and the `nativeChatRequiresLocalTranscript` branch in `resolveMobileNativeChat`,
the readability hook, and the flag through the controller, active resolution and
terminal action sheet. The hosted page implements the same operation as a local
`true`, so it stays ungated and still makes no bridge call.
Readability is sourced from `repo.list` as before. The only native repo store is
a 60-second host-scoped cache written by the host screen and the new-workspace
dialog, which a deep-linked session route cannot rely on.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Three suites still named `workspace.hostCatalog`: a page mock branch, a
saturation slot that the grant-index filter already skipped, and a ceiling
assertion. None of them reached a registered operation any more.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The page reached a desktop method by first reading `mobileWeb.host.catalog` for
a grant, then forwarding through it. Every grant field duplicated something the
desktop already enforced: the allowlist is the mobile-scope socket gate, the
byte budgets are the bridge envelope, `workspaceParam` was `worktree` in every
entry, and `scope` was "did the page send a workspaceId".
The shell now forwards `workspace.hostRequest` and `workspace.hostSubscribe`
straight through. It rewrites the opaque workspace handle when the payload
carries one, checks one envelope in both directions, and derives a stream's
desktop cancel name from its subscribe name: `X.watch` to `X.unwatch` and
`X.subscribe` to `X.unsubscribe`. Unary versus stream is decided by which shell
operation the page called. `files.unwatch` and `nativeChat.unsubscribe` keep
their names for the released native app and gain `mobileWeb.`-prefixed aliases
sharing the same handler.
Shell feature negotiation goes with it: `init.shellFeatures` had no consumer
beyond its own tests, so the protocol version is the bridge's only compat gate.
That constant and the package bridge range move into `bridge-limits.ts`, and
`bridge-protocol-version.ts`, `bridge-release-policy.ts` and
`shell-feature-contract.ts` are deleted.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The hosted page owned two shell destinations. `terminalSettings` was already
dead: the page's own device operations push `/terminal-settings`, which
`host-web-app/` serves. `connectionLog` was the last live one, and the hosted
`/connection-log` route renders the same `ConnectionDiagnosticsScreen` over the
same bridge-backed native diagnostics operations, so the host screen's
diagnostics link now navigates in place instead of tearing down the session
view to reach the shell copy.
With no destinations left, `MobileWebNativeRouteHandoff` and everything that
carried a request id to it go: the route enum members, the handoff record and
replay, the broker-message wrapper that existed only to complete it, and the
page-document and navigation-authority plumbing. `handleMobileWebBrokerMessage`
collapsed to a single broker `handle` call, so the shell calls that directly.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
`mobileWeb.nativeChat.readability` answered `true` for every workspace: the
host treats native-chat eligibility as its own and each transcript read
validates its execution provider independently. The page still paid a bridge
round trip per workspace to learn nothing.
Delete the RPC, its page-catalog grant, both page clients, the operation
contract, and the shell hook that consumed it. With the flag gone the
`nativeChatRequiresLocalTranscript` pre-gate in `resolveMobileNativeChat` is
dead, so grok and omp tabs now resolve like every other supported agent and a
transcript the host cannot read surfaces as a read error rather than a hidden
chat.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
nativeChat.readSession answers an unreachable SSH transcript with `{ error }`
rather than throwing. The removed shell translator turned that into a
retryable host_error; the desktop adapter forwarded it, so the page parsed
it as a non-retryable invalid_message and never retried after reconnect.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The max-lines split of the Docker SSH relay helper changed the perf spec's
import lines, which routed a spec main deliberately keeps out of CI into this
branch's changed-spec lane. Re-exporting the reconnect helper from the
original module restores main's import so the spec no longer differs.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The packaging fixture's document predates the verify rule that every packaged
script is referenced once with `defer`, and the Docker SSH transcript spec
still called the removed opaque-resource `nativeChat.read`; it now reads by
tab id like the page does.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The three placeholder cases pinned the unbounded sanitizer output; now that
the desktop clips reads to the page contract they prove the clip is required
and sufficient.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The host sanitizer caps text blocks at 64 KiB and bounds neither block
count nor identifier length, but the page's read schema is strict at
4200 characters, 64 blocks and a 1024 character id. Each overrun was a
different silent loss: an over-long text block is an unclassifiable
member of an array of unions and vanished, while an over-long id or an
over-count turn failed its message, and messages is not a union array,
so one bad message failed the whole read with a non-retryable
invalid_message.
Clipping now runs on every read and stream event, ahead of the byte
budget that only engaged above 512 KiB. Unknown keys and unknown block
types still pass through, because the shell forwards this payload
without parsing it.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Replaces the removed implementation worklog with a decision record that a
reader on main can act on, and corrects the architecture note now that the
page addresses host tabs by host ids and the catalog is cached per connection.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* fix(relay): bound control RTT samples per ping and per flush window
An authenticated host chose how many round-trip samples a cell recorded: every
pong carrying a recent plausible `t` was forwarded to the process-wide window,
which grew unbounded until the 30s flush copied and sorted it for percentiles.
Time a pong only when it echoes the `t` of the ping still outstanding on that
session, so a flood yields at most one sample per ping the cell actually sent.
A pong that lost the race to the next ping is dropped for timing but still
counts as proof of life for the silence watchdog. Bound the process-wide window
with a 1024-sample reservoir (Algorithm R) so the percentiles stay unbiased,
keep `controlRttSamplesDelta` meaning round trips observed, and publish
`controlRttSamplesDroppedDelta` for the ones the reservoir did not keep.
Replace the leak guard's blanket `"credential":` string rewrite with an exact,
path-scoped rename of the two schema keys that spell a policed word, and make
the guard case-insensitive now that nothing legitimate trips it.
Follow-up to #19232.
* test(relay): prove the RTT reservoir samples the whole window
The cell announces a connection with a single conn-open. When the desktop's
control socket dies mid-accept the phone waited out the 10s attach deadline and
was closed HOST_OFFLINE, even though the desktop was online. host-hello-ack
already restates those connections in pendingConns, but only by connId and
connTicket, which is not enough for the desktop to dial: kind and relayDeviceId
decide the pairing authority a connection carries and the E2EE device binding,
so neither may be guessed.
The cell now states kind and relayDeviceId on each pending entry, but only to a
host that advertised it can read them: a shipped host parses those entries
strictly, so an unannounced key fails the whole ack parse and kills a working
control. The advertisement rides the control upgrade as
x-orca-host-capabilities, not host-hello, because HostHelloSchema is strict on
the cell too and any new hello key is refused by every already-deployed cell.
The capability is keyed by socket, not by session: a rebind can land a successor
whose decoder is older or newer than the one that opened the session, and the
ack must follow the socket that will actually read it.
With no capable host in the fleet the emitted ack is byte-identical to today's.
The desktop half that consumes the new fields is #19238.
`relay_region_rehome_source_cell_ids` listed only the 16 US cells, and that
list is the sole thing that stamps ORCA_RELAY_REHOME_DIRECTOR_SERVICE_ACCOUNT
and ORCA_RELAY_REHOME_AUDIENCE into a cell's startup script. A cell reports
regionalRehomeProtocol 1 only when both are present, so c27-c29 have always
reported 0. That leaves them ineligible as rehome sources and, once the worker
is bidirectional, as targets too, which strands the US desktops homed there.
This is a prerequisite only. Merge and roll it ONLY AFTER the bidirectional
rehome director change is deployed. Two live gates still hard-code the primary
region and would reject an Asia source no matter what the template stamps:
`cloud/apps/relay/src/app.ts` line 610 fails the trust probe with 409 when the
source cell's region is not RELAY_DEFAULT_REGION, and
`cloud/apps/relay/src/assignment-store.ts` line 5476 skips such a cell as
source_ineligible during rehome source selection. The bidirectional lane
removes both.
The topology check asserted every source sits in the primary region. That
mirrored those two gates rather than protecting anything Terraform owns, so it
is now advisory: it requires only a configured, unfenced cell with an explicit
connection limit, and the comment records that region eligibility belongs to
the director's own source and target predicates. Every cell's region is
already constrained by the assert above it.
The same-cap census test cross-checked membership against us-central1. Every
reviewed serving cell now carries the trust, so it asserts protocol 1 for all,
plus one non-source cell to keep the validator's protocol-0 branch covered.
Roll sequencing, because this apply is not self-contained:
- After the apply the Asia templates carry the two rehome lines, and the
`unexpectedRehome` rule at `cloud/dev/scripts/validate-relay-capacity-plan.mjs`
lines 243-247 rejects a protocol-0 plan that contains them. So c27-c29 have
no dispatchable protocol-0 same-cap roll until the director gate is gone or
this is reverted.
- The same-cap job runs the per-host trust probe after isolate, drain, and the
targeted apply. A 409 there leaves the cell serving but isolated and
migration-only, which is what happened to c13 on 2026-09-06.
- The only safe path: deploy the bidirectional rehome director, then dispatch
`Deploy Relay Production Same-Cap` canary-apply for one Asia cell with
target-rehome-protocol 1 and rollback-rehome-protocol 0, then batch-apply the
remaining two. That job runs its own targeted template and MIG apply.
- Never reach these cells with an untargeted root apply. The current plan
carries 60 changes and 50 destroys of unrelated standing drift.
The Relay integration harness and the remaining shell fixtures still
faked bind and page-session round trips, so they proved a protocol that
no longer exists. They now assert the host ids the page actually sends.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* Revert "feat(mobile): draw the last known tab strip while a session reconnects (#19258)"
This reverts commit 0ba7f8dc8d.
* Revert "perf(mobile): cut the relay reconnect critical path and admit dead sockets faster (#19236)"
This reverts commit 23df74d85a.
The shell mirrored the desktop's opaque handles with a browser authority,
a native-chat binding cache and a page-session lifetime subscription, and
the page bundle bound a resource before every read, write and file
action. With host ids in the session snapshot none of that maps anything:
the page's tab id is the host tab id and its session id is the provider
session id.
Chat and terminal actions now travel in a single host call, the shell
resolves a chat session against the live tab list when it needs a native
terminal handle, and the image ledger keeps its own bound rather than
riding on the deleted binding cache.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
`syncOrchestrationFederation()` coalesces onto an already-in-flight relay-tick
sync, which may have pulled from the peer before the caller's mutation existed.
Tests used it as a barrier, so `keeps a timed-out remote question resumable`
could reply against a home DB that had never imported the worker's question:
the reply failed with `Message not found`, no `to_worker` relay was enqueued,
and the resume ask surfaced it 5s later as a spurious timeout.
Add `syncFederationBarrier()`, which chains each active dispatch past the
current round via `syncOrchestrationFederatedDispatchAfterCurrent`, and use it
at every barrier-purpose sync site. The two tests whose subject is the sync
machinery itself keep the raw call. Also assert the reply response, so a failed
reply fails at the reply instead of masquerading as a timeout.
Production is unaffected: `syncOrchestrationFederation` has no production
callers, real read-after-write paths already use the after-current sync, and
relay ticks retry every second.
* feat(relay): alert on far-cell placement and skewed region hints
US desktops were homed on asia-east2 cells for weeks in 2026-08 with every
existing relay alert green. Roughly 226 of 332 hosts on those cells were
non-APAC, and a phone connect took ~10 s there against ~0.6 s in region, but
nothing in Cloud Monitoring could see distance: the connection, queue, heap,
and SQL bars all measure a cell's own health, which was fine.
Three policies close that gap. Two read distance per cell, from the accept
and control-RTT timing added in the parent commit: phone-accept p95 above
2 s, and control ping p50 above 150 ms. The third reads the cause fleet-wide,
as the asia-east2 share of the region hints desktops send the director, so a
mis-picking client probe is visible before it lands anyone on a far cell.
All three are MQL rather than the metric filters the other relay policies
use. Every runtime metric is a DELTA DISTRIBUTION, and a filter condition can
only align one with a percentile; each alert needs the sum of the extracted
values as a volume floor so a sparse window cannot page. None of these
metrics exists in the project yet, so what was checked against production is
the query shape: the same MQL run over existing metrics of the same kind.
The skew denominator needs one log-based metric per hint key, so
`requestedRegionsDelta` now has one per relay region plus the unhinted
bucket. Those ride the existing snapshot metric family, which adds map
entries without touching the live metrics. A ratchet test pins the key list
to relay-contract's RELAY_REGIONS: a region added there without a metric
would shrink the denominator, so the test fails rather than letting the share
quietly inflate.
* fix(relay): compare hinted regions against placed ones, not a fixed share
Review found the skew alert inverted at both ends. A fixed 40% bar on the
asia-east2 share of region hints was silent through the exact broken state it
was written for, and would page forever once the desktop probe is fixed and
the genuine APAC share rises past it. An absolute share cannot separate those
because it has no reference point.
The hint share now has one: the share of assignments the director actually
placed in that region during the same hour. Measured over twelve hours on
2026-09-07, while the probe was still mis-picking, asia-east2 was 33.8% of
33,800 hinted requests and 7.9% of 45,364 assignments. That is a 4.27x
divergence and a 25.9-point gap, so the alert fires above 2x and 15 points,
inside the broken state and outside a healthy one. Both bars must hold: the
ratio alone blows up on tiny placement counts, the gap alone misses a
proportionally large skew at low volume. The reviewer proposed either bar
alone; requiring both keeps each one meaningful and still clears today's
numbers with room.
`unhinted` requests leave the denominator. They were 27% of all requests, so
a client that always sends a hint would move the number from 21.9% to 35.0%
with no behaviour change at all.
The comparison needs per-region placement counters, so `selectedRegionsDelta`
gets log-based metrics alongside the requested ones. Rather than extract four
hyphenated map keys through quoted field paths, which nothing in the project
does and which cannot be checked without applying, the relay now also
publishes flat `requestedRegion<Region>Delta` and `selectedRegion<Region>Delta`
fields next to the untouched maps. They are emitted as zeros in every
interval, so no series can drop out of the alert's inner join in an hour with
no asia placements, which is exactly the hour the skew is worst. Additive
only: metricVersion is unchanged, the maps still carry anything outside the
catalog, and the emitter's leak guard still passes.
Two corrections to what the previous commit claimed. None of these metrics
exist in the project yet, so the code, the doc and this message now say what
was actually checked against production: the query shapes, run over existing
metrics of the same kind. And the control-RTT policy records that EU desktops
on us-central1 sit at 100-130 ms, so a European-heavy cell can approach the
150 ms bar while correctly homed.
The skew alert will stay lit after a client fix until the backlog is rehomed.
Sticky assignment never re-consults the hint, so a desktop already on an asia
cell keeps landing there whatever it now asks for. The policy description and
the doc both say so, so nobody reads a slow clear as a failed fix.
* fix(relay): cross-multiply the skew bars so a zero placement share still fires
`hint_share / placement_share` is undefined in the hour that matters most.
When the director placed nobody in the region, MQL returns no rows for either
0/0 or x/0, so the series disappears before the gap and volume clauses run and
the alert stays silent. That hour is not hypothetical: it is every desktop
asking for a region while the director puts nobody there, which is what a
drained, fenced, or full region looks like, and it is the most extreme skew
the alert can see.
The condition is now cross-multiplied, `hint_share > 2 * placement_share`,
which is well defined at zero. Both forms were run read-only against
production surrogates chosen so the placement denominator is exactly zero:
the ratio form returned no rows, the cross-multiplied form returned the series
with the condition true on every point. A second surrogate pass with a tiny
hint share returned the series with the condition false, so the gap clause
still suppresses the healthy shape rather than the query silently matching
everything.
The flat field names are no longer derived on either side. Terraform title
cased each dash-separated part and the emitter upper cased each part's first
character, so the ratchet had to pin two source expressions by regex, which a
reformat would break and which never compared the actual rendered names. Both
sides now declare a literal map, relay-contract's
RELAY_REGION_METRIC_SEGMENTS and Terraform's relay_region_field_segments, and
the test compares the two declarations against each other and against the
expected names. `satisfies Record<RelayRegion, string>` makes a region added
without a segment a compile error rather than a silent gap in the alert's
denominators.
Both ratchets were checked by mutation: a wrong Terraform segment, a contract
region with no Terraform entry, and a revert to the ratio form each fail the
node test, and the new region fails the contract build.
Unpairing cleared metadata, overlays, and credentials but left the
host-scoped page preference blob behind forever, since nothing else is
keyed by the pairing key. Remove it when no remaining host still holds
that key, best-effort so it cannot fail an authoritative removal.
The host list mutation queue moves to its own module to keep host-store
under its line budget; it is the same chain, only named.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The native route inlined 245 lines of orchestration the hosted screen had
already factored behind DiagnosticsDeviceOperations, so the two paths
gathered, redacted, and submitted the same report differently. The screen
now takes the device operations, a clipboard writer, and a host picker
slot; the native route supplies the existing native operations and its
host chips.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The hosted adapter rejected a configure whose echoed receipt differed from
the request while the native adapter accepted any receipt. The paired
desktop is trusted and ships the page bundle, so the extra check only
produced a failure the native path does not have.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The native and hosted settings routes each hardcoded the same seven rows
and had already diverged: Troubleshooting carried a different icon in each
and the availability flags only existed on one side. One builder now owns
the rows and takes the hosted shell/page-preference availability, so the
routes keep only what is genuinely theirs — credential cleanup on native,
external links on hosted.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The load path set the error but left busy pinned, and every control is
disabled while busy, so a single failed preference read left the screen
inert with no way to retry. Clear busy in a finally as the voice settings
screen does.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The broker parsed the stored blob before dispatching on the action, so a
corrupt or oversized blob failed every request forever — including the
clear that would have repaired it. A reset now runs before the size check
and the parse, and a read serves empty values with a single warning rather
than a hard unavailable. Writes still refuse to overwrite data they cannot
read.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Every getItem crossed the shell bridge as its own one-key multiGet, so the
five preference reads a settings screen issues on mount exhausted the
shell's four-concurrent request grant. Adjacent same-action calls in a tick
now fold into one request bounded by the contract's 64-key limit, and
flushGetRequests dispatches the pending batch instead of doing nothing.
That removes the reason product code hand-serialized its reads, so the
sequential awaits in the terminal preference loaders and the ad-hoc
in-flight accessory dedupe go with it.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* feat(mobile): draw the last known tab strip while a session reconnects
Reopening a workspace the phone has already visited threw away everything
it knew. The route clears its tabs on mount, so until the reconnect lands
and the first snapshot is applied the session screen has an empty header
and a bare spinner, even though the strip it is about to be handed is the
one it drew a minute ago.
Persist the four fields the strip actually draws -- id, type, title, agent
-- per host and workspace, and add a reconnecting-with-cache shape to the
route state so those rows render immediately, disabled, under the ids the
live snapshot will reuse. Live tabs always outrank the cache, so a
mid-session drop keeps its mounted terminals; an exhausted retry loop or a
rejected pairing outranks it the other way, because a strip the user cannot
reach is worse than the existing offline affordance. With nothing cached
the screen behaves exactly as before.
The body stays a placeholder. Replaying stored scrollback into the terminal
WebView would double-render the same rows once the live stream replays them,
so the strip is the cached content and the body waits for the stream.
* fix(mobile): keep shell titles and unpaired hosts out of the cached tab strip
Review of the reconnect strip cache found two ways it leaked.
A terminal's title is whatever the shell last set, which is routinely the
command line: a psql URL with an inline password, a curl with a bearer
token. Both fit well inside the 64-character cap and both were written to
plaintext AsyncStorage verbatim. Browser tabs carried their page title the
same way. Terminals and browsers now collapse to a fixed label, with a
resolved agent naming itself because that lookup is a closed enum. The rule
lives in the storage module rather than its caller, so it holds for entries
an older build already wrote, and a tab type this build cannot draw is
dropped instead of having its title trusted.
The cache also survived forgetting a host. Nothing expired an entry, and
the module-global memory map meant a later save from any surviving host
serialized the forgotten host's rows straight back to disk. Both cleanup
paths now evict by host, dropping the in-memory rows and rewriting storage,
with a pending debounced write cancelled so it cannot restore them.
Also: the storage key digests the workspace id, which ended in a filesystem
path, and cached rows carry the same de-emphasis as the disabled tab-bar
buttons beside them, so an inert row does not pass for a live one.
* perf(mobile): cut the relay reconnect critical path and admit dead sockets faster
Phone medians put E2EE authentication at ~424ms but `connected` at ~630ms,
because the session serialized two RPC round trips behind it: the resume
confirm (`pairing.getEndpoints`) and the capability advisory. Both now ride
the authenticated socket concurrently and off the critical path, so the
session publishes `connected` as soon as E2EE authenticates. Peer identity
is already proven by then — the confirm carries credential/lease bookkeeping
and the cell assignment check, and it still fails the session on a bad answer
or a foreign relayHostId, only later. `persistResumeConfirmation` awaits the
new `whenResumeConfirmed()` instead of assuming the answer is present at
`connected`.
Foreground liveness on a retained relay: `notifyForeground('app-resume')`
now probes past the 10s voluntary minimum on urgent bounds (2s, one miss),
so a socket that died while the process was suspended is admitted in ~2s
instead of ~8s. Focus and network nudges keep the old minimum and bounds.
Relay sessions also gain a 25s idle sweep, gated on foreground so a
backgrounded app spends no probes.
Recovery is no longer blocked by the direct return probe. The probe's 12s
dial is a pure observation on its own socket, so it takes the supervisor's
operation mutex only for the cutover; a relay recovery landing during a
foreground return now starts immediately instead of waiting the budget out.
Requests that do land during the cutover are queued in a new
RelayRecoveryIntentQueue and replayed on release — an owning forced
replacement keeps its intent, everything else replays as a plain recovery.
Tests updated deliberately, for the new ordering:
- 'sends no periodic traffic while an authenticated relay is idle' asserted
the absence of any relay idle probe, which is exactly the gap D3 closes.
Replaced by a sweep test plus a backgrounded no-probe test.
- 'rate-limits foreground sequences without suppressing a retry' asserted
that app-resume was suppressed inside the 10s minimum. An app resume is
now the one nudge that must never be rate-limited.
- the session helpers waited for the confirm answer before `connected`;
they now authenticate, read both concurrent frames, and settle them.
* fix(mobile): book backoff when a relay resume confirm fails after the cutover
Review round 1 on 352bfd2300.
P1: publishing `connected` at E2EE authentication made `migrateTo` resolve
before the resume confirm answered, so a confirm that failed afterwards —
a `relayHostId` mismatch from a rehomed desktop is the live case — was still
reported as an `established` dial. registerFailure was skipped, no cooldown
was booked, recordMigration()/setActiveSession() ran for a dying session, and
the queued-recovery replay redialled immediately: a tight loop with a
connected→disconnected blip per pass. The establisher now awaits
whenResumeConfirmed() after the cutover and, if the session is no longer
connected, reports a failed dial (or an aborted one when direct won or the
supervisor went inactive) exactly as a rejected migrateTo used to. The UI
still connects early; only the supervisor's bookkeeping waits.
The state check, rather than getFailure(), is the oracle: a live session can
carry a latched failure without having failed yet, and "is this session still
alive once the confirm settled" is precisely the question migrateTo used to
answer.
P2: the resume probe profile goes to two 2s misses instead of one. The first
frame after a resume rides a cold radio and a possibly distant cell, so one
slow answer is not proof of a dead link; the verdict still lands at 4s rather
than the previous 8s.
Nits: the direct probe's two early returns no longer close the candidate the
finally also closes (the second shape pre-existed); RelayRecoveryIntentQueue
is cleared in the supervisor's stop().
Mutex-hold note: persistResumeConfirmation, and now the establisher's own
await, are bounded by the confirm's request timeout. That would have been the
session's 30s default, so the confirm is pinned to RELAY_CONFIRM_TIMEOUT_MS
(12s) — the same bound migrateTo's waitForAuthenticated applied before.
Test: a supervisor-level case where every dial authenticates then fails the
confirm must book 250/500/1000ms backoff with no immediate redial, and must
never record a migration. It fails on the pre-fix establisher.
* fix(relay): rehome hosts to their preferred region in either direction
The regional-rehome worker only moved hosts from a us-central1 cell to an
asia-east2 one, so a host whose desktop later records us-central1 stays where
it was put. Rehoming now compares the fresh preference against the region of
the cell the host is on and moves it to a general cell in the preferred
region either way, through the same drain, migrate, safety, and rate-limit
machinery.
- relay_region_rehome_attempts.preferred_region accepts both regions; existing
databases are upgraded in place by an idempotent named-constraint swap that
is safe when several directors start at once.
- A target must carry the drain protocol too: moving a host onto a cell it
can never be drained off again is the trap this change exists to undo. The
fleet whose health gates a rehome is now every general drainable cell,
which is exactly the set of legal sources and targets.
- The trust probe accepts a source cell in any region.
No wire change, and no behaviour change while the durable control is off.
* fix(relay): bound bidirectional rehoming with a per-host cooldown
Moving hosts in both directions removed the property that made the old
one-way worker self-terminating: a desktop whose region probe flips would be
dragged back and forth, one full drain and migrate per flip, because the
preference age never expires while the host keeps reconnecting.
- relay_region_rehome_control gains host_cooldown_ms, an operator input
plumbed like preference_max_age_ms (workflow, ops script, admin route,
durable row) and defaulted to seven days. A host with any attempt row
inside the window, whichever way that move went, is not a candidate; the
claim re-reads it under lock so an attempt landing between scan and claim
cannot start a second move. Skips are named host_cooldown, and the lookup
rides a new index on (user_id, relay_host_id, created_at).
- The candidate scan now also requires the target cell to be enabled, so it
mirrors the claim-time filter exactly and stops spending batch slots on
candidates that are certain to be skipped.
- Region CHECK lists are rendered from the shared region list instead of
being written out four times.
- The operations runbook states that cells without the drain protocol are
neither sources, targets, nor members of the safety gate.
* fix(relay): keep rehome reads and brakes working across the cooldown rollout
The ops script validated hostCooldownMs on every inspected control, so
against any director image predating the field inspect, pause, disable, and
failed-enable recovery all threw client-side. The workflow always runs from
main while the director image is operator-supplied, so that window opened at
merge and reopened on every rollback: the operator lost read-only visibility
and both emergency brakes while the worker could still be enabled.
The field is now validated only when the director reports it, and every apply
body that echoes an inspected control omits the key when that control lacks
it, so a legacy director never sees an unknown key. The write path stays
fail-closed the other way: enable refuses up front, before any mutation, when
the director does not report a cooldown it could honour.
Also replaces two bare 'us-central1' defaults with RELAY_DEFAULT_REGION.
* feat(relay): time successful client accepts and control round trips
A 6s accept on a cross-region cell was invisible: only the abandoned path
was timed. Record per-stage durations across acceptClient and acceptHostData
(assignment/credential/activity/attach), emit one completed log line per
accept, and aggregate p50/p95/max into the runtime metrics event.
Sample control ping round trips from the pong echo so a host sitting on a
distant cell is visible fleet-wide and per host, rate-limited to one log
line an hour per session.
* fix(relay): review round 1 on accept and control-RTT timing
Omit the accept and RTT percentiles from windows with no samples: accepts
are sparse, so a zero point every 30s would pin the p50 at 0 and collapse
the p95. The *Delta counts still publish, and say when the omission is
expected. Control-renewal output is unchanged.
Add a `basis` stage for the splice lease and connection-basis writes that
run between the host data leg and relay-hello, and start `attach` where the
activity stage ended, so the stages now tile the whole accept and their sum
equals totalMs. Clamp every stage at zero against a backwards clock step.
Carry role/cellId/region on both new log lines, flatten the stage p95 field
names so the log-metric extractors stay top-level, and record that only the
RTT median reads as distance: the desktop echoes the pong on its main
thread, so the p95 and max track desktop stalls.
The long-lived shell note was an agent worklog, not a decision record. The
simplification audit cited /tmp run directories that no reader can open, and
opened with a validation hold that duplicated what the validation section says.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Nothing checked that what the desktop sanitizer emits is readable by the page,
even though a page parse failure is permanent. The new contract test pins the
shapes that survive and records three that do not: an over-long text block is
silently dropped, and a turn over the block limit or an over-long message id
fails the whole read. The desktop caps are the released native app's, so the
missing bound belongs to the mobile-web read adapter, not the sanitizer.
The Source Control host subscription had only happy-path coverage from a
roundtrip test. It now has its own cases for watcher failure, end-of-watch,
retirement, overflow batching and an unreadable frame.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The desktop's page-method catalog is a module constant, so its grants can only
change when the desktop process restarts, which tears down the socket and makes
the broker replace the client. Reading it over the wire before every forwarded
request cost a full round-trip per host call and, at mount, several identical
reads at once. Cache grants per client with in-flight dedupe, keyed by method so
a sibling's cancellation cannot fail a peer, and clear it wherever the broker
already clears per-connection authority state.
The host-request concurrency ceiling lived twice: a literal 4 in the accounting
module and a grant limit the shell never read. Take it from the grant alone and
raise it to what a product page needs, counting only one-shot forwards, since
subscriptions have their own ledger cap and catalog reads no longer reach the
wire per request.
Byte ceilings disagreed in both directions. The grant schema let a desktop
advertise a size a shipped shell hard-rejects, so cap it at the shared bridge
envelope, and give the read-heavy desktop methods that whole envelope instead of
an arbitrary 512 KiB. Forwarding also serialized each payload up to four times
per direction; one walk now yields both the verdict and the length.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb