Commit Graph
10881 Commits
Author SHA1 Message Date
Jinwoo-H 0018fd041f refactor(mobile): delete the workspace and account contracts nothing imports
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 15:43:19 -04:00
Jinwoo-H fb69f030af docs(mobile): record what the workspace translator still owns
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 15:25:31 -04:00
Jinwoo-H 298375cbdd refactor(mobile): move the account domain onto the host lane
accounts.list, the three selection methods and accounts.subscribe now reach
the desktop through the generic lane, and the page parses the snapshot with
its own schema. The shell keeps only the reset-credit arms, which mint a
native idempotency key and so cannot be a plain forward.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 15:23:32 -04:00
Jinwoo-H 671aecd4ff test(mobile): retire the shell echo checks the host lane no longer performs
The page synthesizes its own workspace handle for activate/update/remove now,
so there is no shell echo left to mismatch.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 15:08:50 -04:00
Jinwoo-H 418bb65e1c refactor(mobile): move workspace reads and mutations onto the host lane
worktree.activate/set/sleep/rm, repo.list and ui.get/ui.set now reach the
desktop through workspace.hostRequest, and a mobileWeb.workspace.subscribe
wrapper reduces the client-event firehose to bare catalog change types. The
shell keeps only the worktree.ps catalog read that mints workspace handles;
repo handles are gone from that path, so the page addresses repos by host id.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 15:07:28 -04:00
Jinwoo-H b68f2246e9 Merge ota-p1-routes into ota-p1-integrate
Resolve the catalog rename against the readability deletion: drop the old
host-catalog module and remove the readability entry from the allowlist.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 14:37:11 -04:00
Jinwoo-H 2dd40295a3 fix(mobile): refuse a lane mismatch on the generic host request
Deleting the grant took the mode check with it, so `workspace.hostRequest`
would forward a subscribe method. The desktop answers that with `streaming:
true` frames, and `RpcStreamRegistry.handleResponse` claims every one of them
before the pending-request map sees it, so the unary promise hangs to its 15s
deadline while the desktop subscription stays open with nothing to cancel it.
On the relay transport the pending map wins instead and the promise settles on
a partial first frame, which is wrong in a different way.

The naming rule already answers this without a grant: a method has a derivable
cancel name exactly when it is a subscribe method. `prepareMobileWebHostRequest`
now returns that derivation, `hostRequest` refuses a method that has one, and
`hostSubscribe` refuses a method that does not. Both raise
`unsupported_capability`, the code the grant-mode mismatch used.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 14:36:31 -04:00
Jinwoo-H 7a1b9d2a65 fix(mobile): restore the native-app native-chat readability gate
The RPC deletion removed a check the native app really made. On the hosted page
readability was constant true, but the native shell resolved the workspace's
repo and hid grok and omp tabs whose transcripts the runtime cannot read, which
is every repo behind a non-runtime-owned SSH connection. Deleting it changed
released native behavior.

Restore the gate on the native path only: `isMobileNativeChatTranscriptReadable`
and the `nativeChatRequiresLocalTranscript` branch in `resolveMobileNativeChat`,
the readability hook, and the flag through the controller, active resolution and
terminal action sheet. The hosted page implements the same operation as a local
`true`, so it stays ungated and still makes no bridge call.

Readability is sourced from `repo.list` as before. The only native repo store is
a 60-second host-scoped cache written by the host screen and the new-workspace
dialog, which a deep-linked session route cannot rely on.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 14:32:52 -04:00
Jinwoo-H 313845e3d0 test(mobile): drop the last references to the deleted hostCatalog operation
Three suites still named `workspace.hostCatalog`: a page mock branch, a
saturation slot that the grant-index filter already skipped, and a ceiling
assertion. None of them reached a registered operation any more.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 14:28:00 -04:00
Jinwoo-H a04c0ae995 refactor(mobile): delete the host catalog and shell feature negotiation
The page reached a desktop method by first reading `mobileWeb.host.catalog` for
a grant, then forwarding through it. Every grant field duplicated something the
desktop already enforced: the allowlist is the mobile-scope socket gate, the
byte budgets are the bridge envelope, `workspaceParam` was `worktree` in every
entry, and `scope` was "did the page send a workspaceId".

The shell now forwards `workspace.hostRequest` and `workspace.hostSubscribe`
straight through. It rewrites the opaque workspace handle when the payload
carries one, checks one envelope in both directions, and derives a stream's
desktop cancel name from its subscribe name: `X.watch` to `X.unwatch` and
`X.subscribe` to `X.unsubscribe`. Unary versus stream is decided by which shell
operation the page called. `files.unwatch` and `nativeChat.unsubscribe` keep
their names for the released native app and gain `mobileWeb.`-prefixed aliases
sharing the same handler.

Shell feature negotiation goes with it: `init.shellFeatures` had no consumer
beyond its own tests, so the protocol version is the bridge's only compat gate.
That constant and the package bridge range move into `bridge-limits.ts`, and
`bridge-protocol-version.ts`, `bridge-release-policy.ts` and
`shell-feature-contract.ts` are deleted.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 14:26:39 -04:00
Jinwoo-H 906f7491a5 refactor(mobile): retarget page diagnostics in-page and delete the native route handoff
The hosted page owned two shell destinations. `terminalSettings` was already
dead: the page's own device operations push `/terminal-settings`, which
`host-web-app/` serves. `connectionLog` was the last live one, and the hosted
`/connection-log` route renders the same `ConnectionDiagnosticsScreen` over the
same bridge-backed native diagnostics operations, so the host screen's
diagnostics link now navigates in place instead of tearing down the session
view to reach the shell copy.

With no destinations left, `MobileWebNativeRouteHandoff` and everything that
carried a request id to it go: the route enum members, the handoff record and
replay, the broker-message wrapper that existed only to complete it, and the
page-document and navigation-authority plumbing. `handleMobileWebBrokerMessage`
collapsed to a single broker `handle` call, so the shell calls that directly.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 14:21:08 -04:00
Jinwoo-H adb5c85a65 refactor(mobile): delete the constant-true native-chat readability probe
`mobileWeb.nativeChat.readability` answered `true` for every workspace: the
host treats native-chat eligibility as its own and each transcript read
validates its execution provider independently. The page still paid a bridge
round trip per workspace to learn nothing.

Delete the RPC, its page-catalog grant, both page clients, the operation
contract, and the shell hook that consumed it. With the flag gone the
`nativeChatRequiresLocalTranscript` pre-gate in `resolveMobileNativeChat` is
dead, so grok and omp tabs now resolve like every other supported agent and a
transcript the host cannot read surfaces as a read error rather than a hidden
chat.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 14:14:42 -04:00
Jinwoo-H 48e3a12363 fix(desktop): fail the page chat read when the host transcript is unreachable
nativeChat.readSession answers an unreachable SSH transcript with `{ error }`
rather than throwing. The removed shell translator turned that into a
retryable host_error; the desktop adapter forwarded it, so the page parsed
it as a non-retryable invalid_message and never retried after reconnect.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 06:10:02 -04:00
Jinwoo-H 4d172e9818 test(e2e): keep the relay perf spec byte-identical to main
The max-lines split of the Docker SSH relay helper changed the perf spec's
import lines, which routed a spec main deliberately keeps out of CI into this
branch's changed-spec lane. Re-exporting the reconnect helper from the
original module restores main's import so the spec no longer differs.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 05:53:38 -04:00
Jinwoo-H c418ead885 test: follow the host-id native-chat read and the deferred RNW script rule
The packaging fixture's document predates the verify rule that every packaged
script is referenced once with `defer`, and the Docker SSH transcript spec
still called the removed opaque-resource `nativeChat.read`; it now reads by
tab id like the page does.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 05:49:14 -04:00
Jinwoo-H 89d733a9ca Merge remote-tracking branch 'origin/main' into mobile-rearch
# Conflicts:
#	config/reliability-gates.jsonc
#	src/main/git/command-runner/gh-exec-file.ts
2026-09-07 05:13:00 -04:00
Jinwoo-H 446b668817 test(main): assert the native-chat page-contract clip closes the sanitizer gaps
The three placeholder cases pinned the unbounded sanitizer output; now that
the desktop clips reads to the page contract they prove the clip is required
and sufficient.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 05:09:50 -04:00
Jinwoo-H 3e452c2924 Merge branch 'fix-ledger' into fix-integrate 2026-09-07 05:09:00 -04:00
Jinwoo-H 2cabf0b0bc fix(desktop): clip native-chat reads to the page wire contract
The host sanitizer caps text blocks at 64 KiB and bounds neither block
count nor identifier length, but the page's read schema is strict at
4200 characters, 64 blocks and a 1024 character id. Each overrun was a
different silent loss: an over-long text block is an unclassifiable
member of an array of unions and vanished, while an over-long id or an
over-count turn failed its message, and messages is not a union array,
so one bad message failed the whole read with a non-retryable
invalid_message.

Clipping now runs on every read and stream event, ahead of the byte
budget that only engaged above 512 KiB. Unknown keys and unknown block
types still pass through, because the shell forwards this payload
without parsing it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 05:08:34 -04:00
Jinwoo-H 89b365ca16 docs(mobile): record the long-lived shell decisions and update the architecture note
Replaces the removed implementation worklog with a decision record that a
reader on main can act on, and corrects the architecture note now that the
page addresses host tabs by host ids and the catalog is cached per connection.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 05:08:12 -04:00
Jinwoo Hong 9c8f4c398c fix(relay): bound control RTT samples per ping and per flush window (#19268)
* fix(relay): bound control RTT samples per ping and per flush window

An authenticated host chose how many round-trip samples a cell recorded: every
pong carrying a recent plausible `t` was forwarded to the process-wide window,
which grew unbounded until the 30s flush copied and sorted it for percentiles.

Time a pong only when it echoes the `t` of the ping still outstanding on that
session, so a flood yields at most one sample per ping the cell actually sent.
A pong that lost the race to the next ping is dropped for timing but still
counts as proof of life for the silence watchdog. Bound the process-wide window
with a 1024-sample reservoir (Algorithm R) so the percentiles stay unbiased,
keep `controlRttSamplesDelta` meaning round trips observed, and publish
`controlRttSamplesDroppedDelta` for the ones the reservoir did not keep.

Replace the leak guard's blanket `"credential":` string rewrite with an exact,
path-scoped rename of the two schema keys that spell a policed word, and make
the guard case-insensitive now that nothing legitimate trips it.

Follow-up to #19232.

* test(relay): prove the RTT reservoir samples the whole window
2026-09-07 05:06:42 -04:00
Jinwoo Hong 91d7783f2b fix(relay): state pending-conn details to hosts that advertise the capability (cell side) (#19266)
The cell announces a connection with a single conn-open. When the desktop's
control socket dies mid-accept the phone waited out the 10s attach deadline and
was closed HOST_OFFLINE, even though the desktop was online. host-hello-ack
already restates those connections in pendingConns, but only by connId and
connTicket, which is not enough for the desktop to dial: kind and relayDeviceId
decide the pairing authority a connection carries and the E2EE device binding,
so neither may be guessed.

The cell now states kind and relayDeviceId on each pending entry, but only to a
host that advertised it can read them: a shipped host parses those entries
strictly, so an unannounced key fails the whole ack parse and kills a working
control. The advertisement rides the control upgrade as
x-orca-host-capabilities, not host-hello, because HostHelloSchema is strict on
the cell too and any new hello key is refused by every already-deployed cell.

The capability is keyed by socket, not by session: a rebind can land a successor
whose decoder is older or newer than the one that opened the session, and the
ack must follow the socket that will actually read it.

With no capable host in the fleet the emitted ack is byte-identical to today's.
The desktop half that consumes the new fields is #19238.
2026-09-07 05:06:38 -04:00
Jinwoo-H db127197ef Merge branch 'fix-ledger' into fix-integrate
# Conflicts:
#	docs/reference/plans/2026-09-06-long-lived-mobile-shell.md
#	mobile/src/mobile-web/mobile-web-capability-authorities.ts
#	mobile/src/mobile-web/mobile-web-capability-execution-arms.ts
#	mobile/src/mobile-web/mobile-web-host-native-chat-mutation-roundtrip.test.ts
#	mobile/src/mobile-web/mobile-web-host-native-chat-roundtrip.test.ts
#	mobile/src/mobile-web/mobile-web-host-requests.ts
#	mobile/src/transport/mobile-relay-hosted-bridge.integration.test.ts
#	src/main/runtime/rpc/methods/mobile-web-host-catalog.ts
#	src/main/runtime/rpc/methods/mobile-web-session-snapshot.ts
2026-09-07 05:03:50 -04:00
Jinwoo Hong db13cff832 relay: give the asia-east2 cells the regional rehome identity (#19239)
`relay_region_rehome_source_cell_ids` listed only the 16 US cells, and that
list is the sole thing that stamps ORCA_RELAY_REHOME_DIRECTOR_SERVICE_ACCOUNT
and ORCA_RELAY_REHOME_AUDIENCE into a cell's startup script. A cell reports
regionalRehomeProtocol 1 only when both are present, so c27-c29 have always
reported 0. That leaves them ineligible as rehome sources and, once the worker
is bidirectional, as targets too, which strands the US desktops homed there.

This is a prerequisite only. Merge and roll it ONLY AFTER the bidirectional
rehome director change is deployed. Two live gates still hard-code the primary
region and would reject an Asia source no matter what the template stamps:
`cloud/apps/relay/src/app.ts` line 610 fails the trust probe with 409 when the
source cell's region is not RELAY_DEFAULT_REGION, and
`cloud/apps/relay/src/assignment-store.ts` line 5476 skips such a cell as
source_ineligible during rehome source selection. The bidirectional lane
removes both.

The topology check asserted every source sits in the primary region. That
mirrored those two gates rather than protecting anything Terraform owns, so it
is now advisory: it requires only a configured, unfenced cell with an explicit
connection limit, and the comment records that region eligibility belongs to
the director's own source and target predicates. Every cell's region is
already constrained by the assert above it.

The same-cap census test cross-checked membership against us-central1. Every
reviewed serving cell now carries the trust, so it asserts protocol 1 for all,
plus one non-source cell to keep the validator's protocol-0 branch covered.

Roll sequencing, because this apply is not self-contained:

- After the apply the Asia templates carry the two rehome lines, and the
  `unexpectedRehome` rule at `cloud/dev/scripts/validate-relay-capacity-plan.mjs`
  lines 243-247 rejects a protocol-0 plan that contains them. So c27-c29 have
  no dispatchable protocol-0 same-cap roll until the director gate is gone or
  this is reverted.
- The same-cap job runs the per-host trust probe after isolate, drain, and the
  targeted apply. A 409 there leaves the cell serving but isolated and
  migration-only, which is what happened to c13 on 2026-09-06.
- The only safe path: deploy the bidirectional rehome director, then dispatch
  `Deploy Relay Production Same-Cap` canary-apply for one Asia cell with
  target-rehome-protocol 1 and rollback-rehome-protocol 0, then batch-apply the
  remaining two. That job runs its own targeted template and MIG apply.
- Never reach these cells with an untargeted root apply. The current plan
  carries 60 changes and 50 destroys of unrelated standing drift.
2026-09-07 05:03:26 -04:00
Jinwoo-H 239ebe9045 test(mobile): retire the resource-ledger fixtures and record the removal
The Relay integration harness and the remaining shell fixtures still
faked bind and page-session round trips, so they proved a protocol that
no longer exists. They now assert the host ids the page actually sends.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 05:00:17 -04:00
Jinwoo Hong d74f8cb787 revert(mobile): hold the relay reconnect path and cache-first reconnect for a separate mobile pass (#19265)
* Revert "feat(mobile): draw the last known tab strip while a session reconnects (#19258)"

This reverts commit 0ba7f8dc8d.

* Revert "perf(mobile): cut the relay reconnect critical path and admit dead sockets faster (#19236)"

This reverts commit 23df74d85a.
2026-09-07 04:56:13 -04:00
Jinwoo-H d478600f27 refactor(mobile): drop the shell and page resource ledgers
The shell mirrored the desktop's opaque handles with a browser authority,
a native-chat binding cache and a page-session lifetime subscription, and
the page bundle bound a resource before every read, write and file
action. With host ids in the session snapshot none of that maps anything:
the page's tab id is the host tab id and its session id is the provider
session id.

Chat and terminal actions now travel in a single host call, the shell
resolves a chat session against the live tab list when it needs a native
terminal handle, and the image ledger keeps its own bound rather than
riding on the deleted binding cache.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:56:03 -04:00
Jinwoo Hong fede3eb2ff fix(test): give federation tests a real read-after-write sync barrier (#19262)
`syncOrchestrationFederation()` coalesces onto an already-in-flight relay-tick
sync, which may have pulled from the peer before the caller's mutation existed.
Tests used it as a barrier, so `keeps a timed-out remote question resumable`
could reply against a home DB that had never imported the worker's question:
the reply failed with `Message not found`, no `to_worker` relay was enqueued,
and the resume ask surfaced it 5s later as a spurious timeout.

Add `syncFederationBarrier()`, which chains each active dispatch past the
current round via `syncOrchestrationFederatedDispatchAfterCurrent`, and use it
at every barrier-purpose sync site. The two tests whose subject is the sync
machinery itself keep the raw call. Also assert the reply response, so a failed
reply fails at the reply instead of masquerading as a timeout.

Production is unaffected: `syncOrchestrationFederation` has no production
callers, real read-after-write paths already use the after-current sync, and
relay ticks retry every second.
2026-09-07 04:54:28 -04:00
Jinwoo Hong e068947d4c feat(relay): alert on far-cell placement and skewed region hints (#19253)
* feat(relay): alert on far-cell placement and skewed region hints

US desktops were homed on asia-east2 cells for weeks in 2026-08 with every
existing relay alert green. Roughly 226 of 332 hosts on those cells were
non-APAC, and a phone connect took ~10 s there against ~0.6 s in region, but
nothing in Cloud Monitoring could see distance: the connection, queue, heap,
and SQL bars all measure a cell's own health, which was fine.

Three policies close that gap. Two read distance per cell, from the accept
and control-RTT timing added in the parent commit: phone-accept p95 above
2 s, and control ping p50 above 150 ms. The third reads the cause fleet-wide,
as the asia-east2 share of the region hints desktops send the director, so a
mis-picking client probe is visible before it lands anyone on a far cell.

All three are MQL rather than the metric filters the other relay policies
use. Every runtime metric is a DELTA DISTRIBUTION, and a filter condition can
only align one with a percentile; each alert needs the sum of the extracted
values as a volume floor so a sparse window cannot page. None of these
metrics exists in the project yet, so what was checked against production is
the query shape: the same MQL run over existing metrics of the same kind.

The skew denominator needs one log-based metric per hint key, so
`requestedRegionsDelta` now has one per relay region plus the unhinted
bucket. Those ride the existing snapshot metric family, which adds map
entries without touching the live metrics. A ratchet test pins the key list
to relay-contract's RELAY_REGIONS: a region added there without a metric
would shrink the denominator, so the test fails rather than letting the share
quietly inflate.

* fix(relay): compare hinted regions against placed ones, not a fixed share

Review found the skew alert inverted at both ends. A fixed 40% bar on the
asia-east2 share of region hints was silent through the exact broken state it
was written for, and would page forever once the desktop probe is fixed and
the genuine APAC share rises past it. An absolute share cannot separate those
because it has no reference point.

The hint share now has one: the share of assignments the director actually
placed in that region during the same hour. Measured over twelve hours on
2026-09-07, while the probe was still mis-picking, asia-east2 was 33.8% of
33,800 hinted requests and 7.9% of 45,364 assignments. That is a 4.27x
divergence and a 25.9-point gap, so the alert fires above 2x and 15 points,
inside the broken state and outside a healthy one. Both bars must hold: the
ratio alone blows up on tiny placement counts, the gap alone misses a
proportionally large skew at low volume. The reviewer proposed either bar
alone; requiring both keeps each one meaningful and still clears today's
numbers with room.

`unhinted` requests leave the denominator. They were 27% of all requests, so
a client that always sends a hint would move the number from 21.9% to 35.0%
with no behaviour change at all.

The comparison needs per-region placement counters, so `selectedRegionsDelta`
gets log-based metrics alongside the requested ones. Rather than extract four
hyphenated map keys through quoted field paths, which nothing in the project
does and which cannot be checked without applying, the relay now also
publishes flat `requestedRegion<Region>Delta` and `selectedRegion<Region>Delta`
fields next to the untouched maps. They are emitted as zeros in every
interval, so no series can drop out of the alert's inner join in an hour with
no asia placements, which is exactly the hour the skew is worst. Additive
only: metricVersion is unchanged, the maps still carry anything outside the
catalog, and the emitter's leak guard still passes.

Two corrections to what the previous commit claimed. None of these metrics
exist in the project yet, so the code, the doc and this message now say what
was actually checked against production: the query shapes, run over existing
metrics of the same kind. And the control-RTT policy records that EU desktops
on us-central1 sit at 100-130 ms, so a European-heavy cell can approach the
150 ms bar while correctly homed.

The skew alert will stay lit after a client fix until the backlog is rehomed.
Sticky assignment never re-consults the hint, so a desktop already on an asia
cell keeps landing there whatever it now asks for. The policy description and
the doc both say so, so nobody reads a slow clear as a failed fix.

* fix(relay): cross-multiply the skew bars so a zero placement share still fires

`hint_share / placement_share` is undefined in the hour that matters most.
When the director placed nobody in the region, MQL returns no rows for either
0/0 or x/0, so the series disappears before the gap and volume clauses run and
the alert stays silent. That hour is not hypothetical: it is every desktop
asking for a region while the director puts nobody there, which is what a
drained, fenced, or full region looks like, and it is the most extreme skew
the alert can see.

The condition is now cross-multiplied, `hint_share > 2 * placement_share`,
which is well defined at zero. Both forms were run read-only against
production surrogates chosen so the placement denominator is exactly zero:
the ratio form returned no rows, the cross-multiplied form returned the series
with the condition true on every point. A second surrogate pass with a tiny
hint share returned the series with the condition false, so the gap clause
still suppresses the healthy shape rather than the query silently matching
everything.

The flat field names are no longer derived on either side. Terraform title
cased each dash-separated part and the emitter upper cased each part's first
character, so the ratchet had to pin two source expressions by regex, which a
reformat would break and which never compared the actual rendered names. Both
sides now declare a literal map, relay-contract's
RELAY_REGION_METRIC_SEGMENTS and Terraform's relay_region_field_segments, and
the test compares the two declarations against each other and against the
expected names. `satisfies Record<RelayRegion, string>` makes a region added
without a segment a compile error rather than a silent gap in the alert's
denominators.

Both ratchets were checked by mutation: a wrong Terraform segment, a contract
region with no Terraform entry, and a revert to the ratio form each fail the
node test, and the new region fails the contract build.
2026-09-07 04:51:48 -04:00
Jinwoo-H 0b052a1e5e Merge branch 'fix-transport' into fix-integrate 2026-09-07 04:48:37 -04:00
Jinwoo-H ffcbfa842a Merge branch 'fix-storage' into fix-integrate 2026-09-07 04:48:37 -04:00
Jinwoo-H 27742de73b Merge branch 'fix-hygiene' into fix-integrate 2026-09-07 04:48:37 -04:00
Jinwoo-H 27a194c5b7 fix(mobile): drop hosted page preferences when a pairing is removed
Unpairing cleared metadata, overlays, and credentials but left the
host-scoped page preference blob behind forever, since nothing else is
keyed by the pairing key. Remove it when no remaining host still holds
that key, best-effort so it cannot fail an authoritative removal.

The host list mutation queue moves to its own module to keep host-store
under its line budget; it is the same chain, only named.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:46:31 -04:00
Jinwoo-H 7d8a829f3d refactor(mobile): run both connection log routes on one screen
The native route inlined 245 lines of orchestration the hosted screen had
already factored behind DiagnosticsDeviceOperations, so the two paths
gathered, redacted, and submitted the same report differently. The screen
now takes the device operations, a clipboard writer, and a host picker
slot; the native route supplies the existing native operations and its
host chips.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:46:31 -04:00
Jinwoo-H ade5b90ba5 refactor(mobile): drop the page-only voice settings confirmation
The hosted adapter rejected a configure whose echoed receipt differed from
the request while the native adapter accepted any receipt. The paired
desktop is trusted and ships the page bundle, so the extra check only
produced a failure the native path does not have.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:46:22 -04:00
Jinwoo-H 5f8494f7eb refactor(mobile): share one settings row list across both routes
The native and hosted settings routes each hardcoded the same seven rows
and had already diverged: Troubleshooting carried a different icon in each
and the availability flags only existed on one side. One builder now owns
the rows and takes the hosted shell/page-preference availability, so the
routes keep only what is genuinely theirs — credential cleanup on native,
external links on hosted.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:46:22 -04:00
Jinwoo-H dae554eb7e fix(mobile): release the terminal settings controls after a failed load
The load path set the error but left busy pinned, and every control is
disabled while busy, so a single failed preference read left the screen
inert with no way to retry. Clear busy in a finally as the voice settings
screen does.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:46:13 -04:00
Jinwoo-H 4bc5a1697d fix(mobile): recover hosted preferences from an unusable blob
The broker parsed the stored blob before dispatching on the action, so a
corrupt or oversized blob failed every request forever — including the
clear that would have repaired it. A reset now runs before the size check
and the parse, and a read serves empty values with a single warning rather
than a hard unavailable. Writes still refuse to overwrite data they cannot
read.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:46:13 -04:00
Jinwoo-H c3c9b94c55 perf(mobile): batch hosted page preference requests
Every getItem crossed the shell bridge as its own one-key multiGet, so the
five preference reads a settings screen issues on mount exhausted the
shell's four-concurrent request grant. Adjacent same-action calls in a tick
now fold into one request bounded by the contract's 64-key limit, and
flushGetRequests dispatches the pending batch instead of doing nothing.

That removes the reason product code hand-serialized its reads, so the
sequential awaits in the terminal preference loaders and the ad-hoc
in-flight accessory dedupe go with it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:46:05 -04:00
Jinwoo Hong 0ba7f8dc8d feat(mobile): draw the last known tab strip while a session reconnects (#19258)
* feat(mobile): draw the last known tab strip while a session reconnects

Reopening a workspace the phone has already visited threw away everything
it knew. The route clears its tabs on mount, so until the reconnect lands
and the first snapshot is applied the session screen has an empty header
and a bare spinner, even though the strip it is about to be handed is the
one it drew a minute ago.

Persist the four fields the strip actually draws -- id, type, title, agent
-- per host and workspace, and add a reconnecting-with-cache shape to the
route state so those rows render immediately, disabled, under the ids the
live snapshot will reuse. Live tabs always outrank the cache, so a
mid-session drop keeps its mounted terminals; an exhausted retry loop or a
rejected pairing outranks it the other way, because a strip the user cannot
reach is worse than the existing offline affordance. With nothing cached
the screen behaves exactly as before.

The body stays a placeholder. Replaying stored scrollback into the terminal
WebView would double-render the same rows once the live stream replays them,
so the strip is the cached content and the body waits for the stream.

* fix(mobile): keep shell titles and unpaired hosts out of the cached tab strip

Review of the reconnect strip cache found two ways it leaked.

A terminal's title is whatever the shell last set, which is routinely the
command line: a psql URL with an inline password, a curl with a bearer
token. Both fit well inside the 64-character cap and both were written to
plaintext AsyncStorage verbatim. Browser tabs carried their page title the
same way. Terminals and browsers now collapse to a fixed label, with a
resolved agent naming itself because that lookup is a closed enum. The rule
lives in the storage module rather than its caller, so it holds for entries
an older build already wrote, and a tab type this build cannot draw is
dropped instead of having its title trusted.

The cache also survived forgetting a host. Nothing expired an entry, and
the module-global memory map meant a later save from any surviving host
serialized the forgotten host's rows straight back to disk. Both cleanup
paths now evict by host, dropping the in-memory rows and rewriting storage,
with a pending debounced write cancelled so it cannot restore them.

Also: the storage key digests the workspace id, which ended in a filesystem
path, and cached rows carry the same de-emphasis as the disabled tab-bar
buttons beside them, so an inert row does not pass for a live one.
2026-09-07 04:42:59 -04:00
Jinwoo Hong 23df74d85a perf(mobile): cut the relay reconnect critical path and admit dead sockets faster (#19236)
* perf(mobile): cut the relay reconnect critical path and admit dead sockets faster

Phone medians put E2EE authentication at ~424ms but `connected` at ~630ms,
because the session serialized two RPC round trips behind it: the resume
confirm (`pairing.getEndpoints`) and the capability advisory. Both now ride
the authenticated socket concurrently and off the critical path, so the
session publishes `connected` as soon as E2EE authenticates. Peer identity
is already proven by then — the confirm carries credential/lease bookkeeping
and the cell assignment check, and it still fails the session on a bad answer
or a foreign relayHostId, only later. `persistResumeConfirmation` awaits the
new `whenResumeConfirmed()` instead of assuming the answer is present at
`connected`.

Foreground liveness on a retained relay: `notifyForeground('app-resume')`
now probes past the 10s voluntary minimum on urgent bounds (2s, one miss),
so a socket that died while the process was suspended is admitted in ~2s
instead of ~8s. Focus and network nudges keep the old minimum and bounds.
Relay sessions also gain a 25s idle sweep, gated on foreground so a
backgrounded app spends no probes.

Recovery is no longer blocked by the direct return probe. The probe's 12s
dial is a pure observation on its own socket, so it takes the supervisor's
operation mutex only for the cutover; a relay recovery landing during a
foreground return now starts immediately instead of waiting the budget out.
Requests that do land during the cutover are queued in a new
RelayRecoveryIntentQueue and replayed on release — an owning forced
replacement keeps its intent, everything else replays as a plain recovery.

Tests updated deliberately, for the new ordering:
- 'sends no periodic traffic while an authenticated relay is idle' asserted
  the absence of any relay idle probe, which is exactly the gap D3 closes.
  Replaced by a sweep test plus a backgrounded no-probe test.
- 'rate-limits foreground sequences without suppressing a retry' asserted
  that app-resume was suppressed inside the 10s minimum. An app resume is
  now the one nudge that must never be rate-limited.
- the session helpers waited for the confirm answer before `connected`;
  they now authenticate, read both concurrent frames, and settle them.

* fix(mobile): book backoff when a relay resume confirm fails after the cutover

Review round 1 on 352bfd2300.

P1: publishing `connected` at E2EE authentication made `migrateTo` resolve
before the resume confirm answered, so a confirm that failed afterwards —
a `relayHostId` mismatch from a rehomed desktop is the live case — was still
reported as an `established` dial. registerFailure was skipped, no cooldown
was booked, recordMigration()/setActiveSession() ran for a dying session, and
the queued-recovery replay redialled immediately: a tight loop with a
connected→disconnected blip per pass. The establisher now awaits
whenResumeConfirmed() after the cutover and, if the session is no longer
connected, reports a failed dial (or an aborted one when direct won or the
supervisor went inactive) exactly as a rejected migrateTo used to. The UI
still connects early; only the supervisor's bookkeeping waits.

The state check, rather than getFailure(), is the oracle: a live session can
carry a latched failure without having failed yet, and "is this session still
alive once the confirm settled" is precisely the question migrateTo used to
answer.

P2: the resume probe profile goes to two 2s misses instead of one. The first
frame after a resume rides a cold radio and a possibly distant cell, so one
slow answer is not proof of a dead link; the verdict still lands at 4s rather
than the previous 8s.

Nits: the direct probe's two early returns no longer close the candidate the
finally also closes (the second shape pre-existed); RelayRecoveryIntentQueue
is cleared in the supervisor's stop().

Mutex-hold note: persistResumeConfirmation, and now the establisher's own
await, are bounded by the confirm's request timeout. That would have been the
session's 30s default, so the confirm is pinned to RELAY_CONFIRM_TIMEOUT_MS
(12s) — the same bound migrateTo's waitForAuthenticated applied before.

Test: a supervisor-level case where every dial authenticates then fails the
confirm must book 250/500/1000ms backoff with no immediate redial, and must
never record a migration. It fails on the pre-fix establisher.
2026-09-07 04:40:40 -04:00
Jinwoo Hong f5be177e44 fix(relay): rehome hosts to their preferred region in either direction (#19241)
* fix(relay): rehome hosts to their preferred region in either direction

The regional-rehome worker only moved hosts from a us-central1 cell to an
asia-east2 one, so a host whose desktop later records us-central1 stays where
it was put. Rehoming now compares the fresh preference against the region of
the cell the host is on and moves it to a general cell in the preferred
region either way, through the same drain, migrate, safety, and rate-limit
machinery.

- relay_region_rehome_attempts.preferred_region accepts both regions; existing
  databases are upgraded in place by an idempotent named-constraint swap that
  is safe when several directors start at once.
- A target must carry the drain protocol too: moving a host onto a cell it
  can never be drained off again is the trap this change exists to undo. The
  fleet whose health gates a rehome is now every general drainable cell,
  which is exactly the set of legal sources and targets.
- The trust probe accepts a source cell in any region.

No wire change, and no behaviour change while the durable control is off.

* fix(relay): bound bidirectional rehoming with a per-host cooldown

Moving hosts in both directions removed the property that made the old
one-way worker self-terminating: a desktop whose region probe flips would be
dragged back and forth, one full drain and migrate per flip, because the
preference age never expires while the host keeps reconnecting.

- relay_region_rehome_control gains host_cooldown_ms, an operator input
  plumbed like preference_max_age_ms (workflow, ops script, admin route,
  durable row) and defaulted to seven days. A host with any attempt row
  inside the window, whichever way that move went, is not a candidate; the
  claim re-reads it under lock so an attempt landing between scan and claim
  cannot start a second move. Skips are named host_cooldown, and the lookup
  rides a new index on (user_id, relay_host_id, created_at).
- The candidate scan now also requires the target cell to be enabled, so it
  mirrors the claim-time filter exactly and stops spending batch slots on
  candidates that are certain to be skipped.
- Region CHECK lists are rendered from the shared region list instead of
  being written out four times.
- The operations runbook states that cells without the drain protocol are
  neither sources, targets, nor members of the safety gate.

* fix(relay): keep rehome reads and brakes working across the cooldown rollout

The ops script validated hostCooldownMs on every inspected control, so
against any director image predating the field inspect, pause, disable, and
failed-enable recovery all threw client-side. The workflow always runs from
main while the director image is operator-supplied, so that window opened at
merge and reopened on every rollback: the operator lost read-only visibility
and both emergency brakes while the worker could still be enabled.

The field is now validated only when the director reports it, and every apply
body that echoes an inspected control omits the key when that control lacks
it, so a legacy director never sees an unknown key. The write path stays
fail-closed the other way: enable refuses up front, before any mutation, when
the director does not report a cooldown it could honour.

Also replaces two bare 'us-central1' defaults with RELAY_DEFAULT_REGION.
2026-09-07 04:40:37 -04:00
Jinwoo Hong ecfcc0d833 feat(relay): time successful client accepts and control round trips (#19232)
* feat(relay): time successful client accepts and control round trips

A 6s accept on a cross-region cell was invisible: only the abandoned path
was timed. Record per-stage durations across acceptClient and acceptHostData
(assignment/credential/activity/attach), emit one completed log line per
accept, and aggregate p50/p95/max into the runtime metrics event.

Sample control ping round trips from the pong echo so a host sitting on a
distant cell is visible fleet-wide and per host, rate-limited to one log
line an hour per session.

* fix(relay): review round 1 on accept and control-RTT timing

Omit the accept and RTT percentiles from windows with no samples: accepts
are sparse, so a zero point every 30s would pin the p50 at 0 and collapse
the p95. The *Delta counts still publish, and say when the omission is
expected. Control-renewal output is unchanged.

Add a `basis` stage for the splice lease and connection-basis writes that
run between the host data leg and relay-hello, and start `attach` where the
activity stage ended, so the stages now tile the whole accept and their sum
equals totalMs. Clamp every stage at zero against a backwards clock step.

Carry role/cellId/region on both new log lines, flatten the stage p95 field
names so the log-metric extractors stay top-level, and record that only the
RTT median reads as distance: the desktop echoes the pong on its main
thread, so the p95 and max track desktop stalls.
2026-09-07 04:40:34 -04:00
Jinwoo-H 24bc25e685 docs: retire the mobile shell worklog and drop machine-local evidence paths
The long-lived shell note was an agent worklog, not a decision record. The
simplification audit cited /tmp run directories that no reader can open, and
opened with a validation hold that duplicated what the validation section says.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:40:30 -04:00
Jinwoo-H 3d62049d93 test: cover the native-chat page contract and the Source Control host watch
Nothing checked that what the desktop sanitizer emits is readable by the page,
even though a page parse failure is permanent. The new contract test pins the
shapes that survive and records three that do not: an over-long text block is
silently dropped, and a turn over the block limit or an over-long message id
fails the whole read. The desktop caps are the released native app's, so the
missing bound belongs to the mobile-web read adapter, not the sanitizer.

The Source Control host subscription had only happy-path coverage from a
roundtrip test. It now has its own cases for watcher failure, end-of-watch,
retirement, overflow batching and an unreadable frame.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:40:30 -04:00
Jinwoo-H 86706dddbd refactor(mobile): cache the host catalog per connection and unify host ceilings
The desktop's page-method catalog is a module constant, so its grants can only
change when the desktop process restarts, which tears down the socket and makes
the broker replace the client. Reading it over the wire before every forwarded
request cost a full round-trip per host call and, at mount, several identical
reads at once. Cache grants per client with in-flight dedupe, keyed by method so
a sibling's cancellation cannot fail a peer, and clear it wherever the broker
already clears per-connection authority state.

The host-request concurrency ceiling lived twice: a literal 4 in the accounting
module and a grant limit the shell never read. Take it from the grant alone and
raise it to what a product page needs, counting only one-shot forwards, since
subscriptions have their own ledger cap and catalog reads no longer reach the
wire per request.

Byte ceilings disagreed in both directions. The grant schema let a desktop
advertise a size a shipped shell hard-rejects, so cap it at the shared bridge
envelope, and give the read-heavy desktop methods that whole envelope instead of
an arbitrary 512 KiB. Forwarding also serialized each payload up to four times
per direction; one walk now yields both the verdict and the length.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:40:23 -04:00
Jinwoo-H 68cf24a488 refactor(mobile-web): drop the unused features argument from the terminal bindings
The bindings never read the shell feature set, so the only caller was passing
state the callee ignored.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:40:21 -04:00
Jinwoo-H 3383be4ca3 style(main): format the native-chat mobile-web files and merge a split import
The two native-chat files were committed unformatted, and the session snapshot
imported the same module twice, which import/no-duplicates flags.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:40:21 -04:00
Jinwoo-H d7400d250b test(mobile): drop the collapsed host/shell roundtrip matrices
Each `it.each([[true, true]])` ran a single case while advertising a matrix, and
the branches guarded by `host && shell` were dead. Naming the case says what the
test proves, and deleting the constant branches leaves only the lane that runs.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:40:20 -04:00
Jinwoo-H 70298a5cfa test(mobile): give the host-origin Source Control journey a real oracle
Under --adversarial-content the journey had no native baseline, so it derived
the expected changed-file label from the page it was checking and then asserted
the page contained it. That assertion could never fail. The expected label now
comes from the adversarial repository fixture that created the changed file, and
the journey refuses to run when neither an independent path nor a native
baseline is available.

The workspace-row tap that follows the journey moved next to it, so the step
owns the row it returns to and the runner stays under the .mjs line ceiling.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-07 04:40:10 -04:00