The refused-body case had no test after the round-5 revert; reintroducing the defect
passed the suite. Also removes the stream types stranded when the terminal subscribe
adapter was deleted, and a truncated field the directory result type no longer declares.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The file explorer's reconnect ternary called the same function on both sides,
and its comment claimed the null branch mattered. Retry now calls
forceReconnect directly, as main did, and `reconnect` leaves the explorer
operations interface, where a UI action did not belong.
Comments corrected: the ratchet's regex doc no longer claims aliasing cannot
hide a subscribe, and names the two shapes it misses; the deliberate-inline
comment says those four are in the routed screens, not the whole remaining
surface; and the chat stop's layout effect names the gate that requires it.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The project target carried a `slug` field the type does not declare and omitted
owner/repo/host, so repoPayload read them as undefined and the oracle only ever
exercised the prRepo:null path. It now uses the real shape and asserts the row
repository reaches the wire; blanking prRepo fails it.
The connection gate asserts TERMINAL_INPUT_SEND_OPTIONS inside sendInput rather
than anywhere in the counted set; removing it from that call fails the gate.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
- listRepositories dereferences result.repos, so a repo.list answer without it
throws into the caller's catch as it did before.
- projectResult requires an explicit ok:true again; a result that omits the
flag is not a confirmed success.
- readRuntimeSettings returns null for a settings-less answer, so the create
sheet keeps what it already had instead of committing {}. The task-create
path keeps its own {} default, which is what it had.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Each had gained a non-null target requirement that demanded a number and a
recognised item type where main's guard did not, so a row main would have sent
for now returns silently or reports an error.
File-viewed and the comment reply gate on main's conditions and build the
never-null identity target after them. Metadata update, issue type and comment
add gate on the slug and the number, as main did, through the slug target,
which does not read the item type.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Each was confirmed to have zero non-test readers before and after:
- HostSessionTerminalOperations.subscribe, its adapter body and its test. The
live stream still calls subscribeMobileTerminalSafely directly, as on main,
so the method existed only to be tested.
- pasteImages and releaseImages, declared on the native-chat interface and
never implemented or called.
- explicitQuickCommandScope, exported with no callers.
- The quick-command snapshot's totalCount and repoId, which nothing read, and
with them the workspaceId argument that existed only to compute repoId. The
hook takes { operations, enabled } again, as main took { client, enabled }.
- The web-artifact result variant, line and column on the resolve request,
rename's third parameter and the directory variant's truncated. The
legacy-list variant's truncated stays: the panel reads it.
Deleting the terminal subscribe also brings the parity oracle's identity
fields to 14, matching main exactly.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
main handled `!response.ok` after the retry helper returned. Throwing inside
the try made a refusal replayable, because the cutover predicate matches any
Error whose message is exactly the migration string, so a host refusal carrying
that text was retried five times.
The helper returns the response again and the caller checks it. The new test
fails under the old shape with six attempts instead of one.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
sparseCheckout was added to the composer create args and forwarded into the
worktree.create params; its only caller never passes it. warning was added to
WorktreeCreateResult and read off the response; nothing reads it, and the one
screen that shows a create warning is on the unchanged inline path with its own
copy.
worktree-create-retry.ts and source-workspace-create.ts are now byte-identical
to main.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Stop routed through the shared chat write, which refuses to start under a 2s
residual budget. The call it replaced tried on anything above zero and could be
accepted, so Stop now has its own send: main's params, main's
{ timeoutMs, budgetSpansConnect: true }, no enter field, the same worker
takeover report on accept and the same unknown/rejected classification.
The parity oracle also follows the two adapters the route's identity fields
moved into, since the visitor only expands bare-identifier calls and could
never reach operations.terminal.sendInput. Identity fields are 15 rather than
main's 14: the seam turned two hook-side deviceToken fields into three
adapter-side client:{id} builders. Dropping the binding in the adapter now
fails the oracle.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
main checked only response.ok, pruned the tab, wrote a tombstone and let the
host's republished snapshot restore it if the close was refused. The adapter
had mapped a refusal to an outcome the caller skipped every state commit for,
leaving the tab on screen, and had added a `closed !== true` precondition main
never had, so a host answering ok without that field would strand the tab.
close() returns a boolean again and the caller prunes on it. The test that
pinned the refusal projection is replaced by two that pin main's contract: ok
with an empty result still reports acceptance, a transport refusal does not.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
main had no clearInputFirst. The flag prefixed a clear byte onto the body of a
terminal.send, which is the shape the same file documents as observed broken in
the field: the burst reached the agent as literal control characters with the
draft still parked. No production caller set it.
mobile-native-chat-send.ts now has no diff against main. The assertions that
existed only to prove nobody set the flag are gone; the case that pins verbatim
text stays and fails if a prefix comes back.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The adapter sources are concatenated, and the last method in each file closes with a bare
brace, so the previous delimiter search could run into the next file. The slice now ends at
the enclosing method's closing brace and fails outright when none exists.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
- HostSessionNativeChatOperations sendMessage, respond and attachImage. Every
shipped send path calls sendMobileNativeChatMessageWithOutcome or
typeMobileNativeChatCommandWithOutcome directly, as on main.
- HostWorkspaceCreationOperations createBlankWorkspace and
createWorkspaceFromSource. use-new-workspace-create-submit imports
createBlankWorkspace from the source module instead.
With those two gone the native creation factory only spread the RPC
operations and RpcWorkspaceCreationOperations only excluded them, so both
collapse. The arg and result types that existed solely for the five methods go
with them, as does the dead SessionTabsResult re-export.
Each name was grepped against non-test, non-adapter callers before and after.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The per-call slice was still a fixed 700 characters, so deleting host from
updateComment passed on deleteComment's host sitting inside the window. It now
ends at the enclosing method, matching the prRepo loop, and accepts the
shorthand property form that github.project.listAccessible uses.
Verified: dropping `host: target.host` from updateComment fails the oracle.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
- The Project host oracle no longer accepts a bare `payload` identifier, which
every call site satisfies by construction. Inline payloads must show a host
or a slug builder; the four typed-payload reads are named explicitly and
their payload types are pinned to carry the host.
- `prRepo` is pinned per mutation and bounded to that adapter method, so
dropping it from one of six fails instead of needing all six.
`github.addPRReviewComment` is pinned again, with the row repository.
- The smart-source fan-out fake now delegates to the shipped search functions
over a scripted transport, so the repoId stamping under test is the real one
rather than the fake's copy.
- The retired-name read's method and `id:` prefix are asserted again, in the
adapter test, since the hook now takes a callback.
- The session adapters join the route parity family, bringing the runtime
strings and identity fields that moved out of the hooks back into scope.
Each restored assertion was verified to fail under the mutation it guards.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
read.bootstrap() folded status.get and the settings fan-out behind one await,
so the supported commit ran only after all five requests succeeded. A
post-connect timeout on any of the four left tasksSupportState at 'unknown',
which the list surface renders as a bare spinner, and the effect deps do not
change on that failure so it never retried.
The probe is now read.tasksSupported() and the fan-out is read.bootstrap(),
with the caller committing supported and clearing the error between them, as
it did before the seam existed.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
resolveTerminalPath took a tabId the native adapter never sent, and main's
wire had none either, so the caller's sourceTabId computation fed nothing.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The native chat subscribe declared an onError the adapter never invoked, since
the transport has none: a stream error arrives as an error frame through the
listener, which is where the caller already handles it. The hook's second
callback was unreachable, so it goes with the parameter.
The terminal adapter's acknowledge was a no-op with no caller here or on main.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Only back, forward and reload have a consumer; the browser pane still issues
its own RPCs. subscribe, navigate, scroll, click, insertText, keypress and
dialog go, with the event and frame-comparison types that existed only for
them. click's fallback had also lost main's move/down/up sequence and cited a
caller that does not exist here.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Returning the promise from inside the try let a rejected count escape the
per-repo catch, reject the whole batch and reach the caller, which fires this
without a handler. The inline call it replaced awaited and reported zero.
The new test drives the real batching shape and fails if the await is removed.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
- Clear terminal consumes the boolean the adapter returns, so a clear that
never reached the host no longer toasts success.
- A refused `@`-autocomplete search returns null instead of an empty list, so
the failure is not cached and the next keystroke retries.
- The Linear workspace picker reloads the host's context on both outcomes, and
`selectWorkspace` detects a refusal by the id the host echoes back, since it
answers a refusal with its unchanged status and no `ok` field.
- `read.loadLinearContext` splits into `linearStatus` and `linearTeams`, so the
screen commits the workspace list and the resolved selection before it asks
for teams, as it did before the seam existed.
- Resolving a review thread reports the host's message for a transport failure
and the caller's wording only for a refusal, in both the item and the
Projects lane.
- The RPC ratchet matches the identifier rather than `sendRequest(`, counts
client subscriptions, walks every source directory and all of `mobile/app`,
classifies adapters by path, and pins each remaining site by file. Verified
against bind, bracket access, a space before the paren, an aliased client, a
helper in another directory and an adapter-shaped new file.
- Deletes the adapter families nothing constructs (file preview, workspace
catalog and the workspace group), the dead device wiring, and the three
methods that shadowed a live call with different behavior.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Two refusals the old inline calls dropped are now reported, because dropping
them is a bug rather than behavior worth preserving:
- linear.selectWorkspace: a refused switch reaches the picker's error instead
of leaving it showing a workspace the host never selected.
- worktree.listRetiredNames: a refused read fails rather than reading as an
empty registry, so the create sheet keeps its previous answer instead of
offering a name that is already spent. This is the policy
src/shared/worktree/retired-name-cache.ts documents.
Every other adapter still answers exactly as the call it replaced did.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
An adapter body has to be the inline call it replaced, so each place where a
carried adapter answered differently from main is corrected here:
- Linear: per-method timeouts and error strings as main sent them, no result
check on selectWorkspace or updateIssue, "Sub-issue not found" and
"Failed to create sub-issue" restored.
- browser.tabCreate: the host's error message, main's 30s budget, and a
tolerated missing page id.
- terminal.clearBuffer: an ok:false reply keeps main's success toast.
- Quick commands: only the load retries a logical cutover, and an empty error
message falls back to main's per-operation copy.
- Project rows: the pull-request mutations no longer require a repository slug;
a row without one forwards prRepo: null, as main did.
- worktree.listRetiredNames: a refused read reads as an empty registry again.
- repo.hooks: the create sheet keeps its state on a refused read, while task
create still fails on one.
The paced-Escape route ref moves from render into a layout effect so the
changed-code React Doctor gate passes; a timer cannot observe the ref between
commit and layout, so the guard's timing is unchanged.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Pins the screens this PR routed at zero inline `sendRequest`, pins the four
calls deliberately left inline at their exact count, and caps the total
screen-side call sites so the remaining debt can only shrink.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The session hooks take HostSessionOperations, the file explorer takes
HostFileExplorerOperations, and the accounts route takes HostAccountOperations.
Terminal, tab, markdown, file, quick-command, browser and native-chat requests
now live in their native-host-session-* adapters.
Four session RPCs stay inline because no adapter method means the same thing:
terminal close, note creation's two calls, and the buffered terminal submit that
needs the raw response to classify a partial write. Worker-terminal takeover
reporting is re-issued at the call sites the adapter does not report from, so
the send-site ratchet still covers all of them.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
The Tasks screens and the new-workspace composer now take an operations
adapter instead of an RpcClient. Every RPC envelope check, response cast and
payload literal that lived in a hook body moved into the matching
native-host-task-* adapter, and the composer's search and create requests moved
into HostWorkspaceCreationOperations.
Behavior is unchanged apart from four adapter-owned differences noted in the PR
body. `worktree.create` from a task item stays inline: the creation adapter
builds the composer's params and adds name-collision retry, which is a
different request, not the same one behind a seam.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
Six operations interfaces the mobile screens will be handed instead of an
RpcClient: HostSessionOperations, HostTaskOperations, HostWorkspaceOperations,
HostFileOperations, HostAccountOperations and DeviceOperations. Each host domain
groups its per-concern contracts as namespaces so a screen takes one prop while
every sub-adapter keeps its own file, and each has one native factory built over
the socket RpcClient.
No screen is wired yet, so this carries no behavior change.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
* fix(orchestration): fence worker release on mobile keystrokes
A settled worker's terminal stayed ownership_state='owned' unless a takeover was
recorded, and the only recorder was orchestration.workerTerminalUserInput, which
only the desktop/web xterm input signal and the native-chat composer call. Mobile
input arrives as terminal.send / stream input frames instead of a report, so a
phone user typing in a settled worker's pane never fenced anything: worker-list
kept recommending release and worker-release closed the PTY under them.
Give the host one definition of "a human typed into this terminal" and route every
lane through it. The mobile input floor claim is that definition and already exists
on both byte lanes: it is taken only for deliberate phone input, never for the
emulator's own query replies, and never for an agent's `orca terminal send`, which
names itself a desktop client and so is indistinguishable from a keystroke at this
layer. Settling that claim after an accepted write now records the takeover through
the same code the RPC reporter uses, throttled to one write per pane per 30s so a
keystroke does not pay for an immediate transaction. The record lands on the runtime
that owns both the terminal and the orchestration database, so SSH-hosted and remote
workers behave exactly like local ones.
No mobile change: mobile already sends client.type (mobile/src/terminal/terminal-send-request.ts:24).
* fix(orchestration): ask the database, do not remember, whether a pane is fenced
The keystroke throttle armed on the attempt rather than on the outcome, so a
zero-row or thrown record poisoned the pane for 30s. A phone keystroke during
the worker-start readiness wait lands before prepareStartingWorkerAuthority
creates the owned resource; a real keystroke seconds later was then suppressed,
the worker settled, and workerRelease closed the terminal under the phone user.
A SQLITE_BUSY on the first write did the same, with no retry.
The cache was the defect, not its arming condition. Its precondition is the set
of owned resources on the pane, which changes underneath it, and any cache keyed
on ownership identity would have to read the database to learn that identity --
which is the whole question. So the input lane now asks: a read using the same
predicate the write uses answers "is anything still fenceable here?" without
taking BEGIN IMMEDIATE, and only then is the write attempted. Ordinary typing
costs a lookup instead of a write lock, a failed write is retried by the next
keystroke, and a takeover writes once per ownership epoch rather than once per
window, because the flip to user_owned removes the pane from the candidate set.
Sharing the predicate keeps the probe from drifting from the writer.
Adds the two escape cases as permanent regressions, drives the mocked send
through the real RuntimeTerminalWriter, and asserts a mobile takeover lifts the
settled-worker resume fence, which no test covered.
* refactor(orchestration): let the database dedupe the takeover, drop the read probe
The probe was meant to keep keystrokes off BEGIN IMMEDIATE, so it had to earn
that with a number. Measured against a real WAL database it costs more than the
write it avoids: at 25 live workers the probe is 0.19ms and the no-op write is
0.10ms, because the probe runs the same candidate selection with each statement
taking its own read snapshot instead of sharing the transaction's. It is a
compensating mechanism with negative value, so it goes, along with the database
method and the predicate extraction it needed.
owned -> user_owned is one-way and scoped to a resource, so the database is
already the dedupe: every deliberate human write attempts the transition, the
second attempt matches no row, and the fence sweep runs only on changed > 0.
Nothing is remembered between keystrokes, so no state can outlive the ownership
it described -- a keystroke before the worker's authority attaches, a write the
database refuses, and a re-dispatch onto the same pane all resolve against the
rows as they are at that instant. An attempt costs about 0.1ms at typical fleet
size and 0.34ms at 100 live workers, on mobile writes only.
Replaces the write-count test, which asserted the old mechanism, with the
invariant: many keystrokes settle into one takeover and one fence sweep. Adds
the re-dispatch case, where a pane's next worker is fenced on its own merits.
* refactor(terminal): name the provenance rule the takeover fence hangs off
The fence rode the mobile input floor claim, with only a comment tying the two
together. The floor is arbitration -- who may write next -- while the fence needs
provenance -- who produced the bytes. They agree today, so anyone reweighing the
floor would have moved the fence without noticing.
isDeliberateHumanInput states the provenance rule on its own terms, and both byte
lanes decide with it when they open a write: the claim carries the verdict beside
the handle, and settlement records the takeover only when a human produced the
bytes. No behavior change -- afterWrite is wired only where the predicate already
answers true -- and the rule is now pinned by its own cases, so a future
arbitration change has to answer this question again rather than inherit it.
* test(orchestration): prove the unary lane classifies a metadata-less phone
A phone build older than client.type is recognised only by its pane's mobile
driver, which the unary lane passes as the provenance evidence. Nothing proved
it did: replacing that argument with false left all 17 tests green while a
shipped phone silently stopped fencing worker release. The new case drives a
clientless send on a mobile-driven pane and fails under that mutation.
The stream lane now passes false outright. Its isMobile is read off the same
client object it carries, so the metadata-less phone cannot reach it, and
passing the flag suggested a legacy path that does not exist there.
Also states what the per-keystroke cost scales with. A pane owning no resource
misses the pane_key index and falls through to a scan of owned resources, so the
figure is tens of microseconds at realistic worker counts rather than a flat
0.1ms, and it grows with rows that are never released.
* fix(terminal): let provenance alone decide the takeover, on every accepted write
A phone older than client.type sends no client metadata, and both stream
initializers derive isMobile from that metadata alone, so such a subscription
reported false and took the stream lane's uninstrumented branch: provenance was
computed and then never consumed. Bytes from a real person landed through both
frame adapters and the resource stayed owned, so workerRelease closed the PTY
under them. The unary lane already fenced that population off the pane's mobile
driver, which is the host's standing reading of clientless input, so the two byte
lanes disagreed at the destructive boundary.
The predicate was still subordinate to floor plumbing: it could only be consulted
where a floor client id existed. Now the accepted-write callback attaches on both
lanes regardless of whether a floor was reserved, and humanInput alone decides
recording; a write holding no claim commits nothing. Arbitration keeps its own
condition around reserveWrite, where it belongs, and the unary lane's duplicate
outer provenance filter is gone. The stream lane reads clientless provenance from
the pane's driver, the same policy the unary lane uses.
The claim holder is now TerminalInputWrite, carrying the verdict beside an
optional floorClaim, so the structure says what the doc said: a write may fence
without holding the floor.
Regressions drive both real frame adapters, clientless direct delivery, and the
paired-web desktop negative. Metadata-only provenance fails 3 on the stream lane
and 1 on the unary lane; gating the callback on a reservation fails the same 3.
* fix(runtime): resolve retained handles before mobile input provenance
A renderer reload clears transient handles while retaining runtime-owned
PTY identities. Legacy mobile provenance saw no leaf, then sendTerminal
restored the same handle and delivered an unfenced key. Normalize through
getLivePtyForHandle at the shared live-leaf resolver entry so classification
and writes agree, preserving existing leaf generation/incarnation checks.
Caller audit:
- terminal-send-method: driver, query-reply authority, lock and floor checks
now resolve the retained PTY before sending.
- terminal-input-delivery: legacy mobile classification and exact-PTY
binding now see the same target as the writer; equality checks remain.
- terminal-multiplex-subscribe-resolution: retained PTYs resolve directly
without a spurious missing-terminal wait.
- terminal-lifecycle-methods resize and terminal-viewport-methods display
mode, restore-fit and updateViewport retain their original PTY target.
- inspectTerminalProcess: avoids false terminal_gone after reload while
preserving provider inspection and incarnation fences.
- getLivePaneKeyForTerminalHandle and getOrchestrationDispatchAuthority:
unaffected because both already call getLivePtyForHandle first.
No wire/schema changes, host fallback, process-death inference, or Git
workspace assumptions; SSH providers keep ownership of execution evidence.
Validation:
- Unmodified round-3 reviewer probe: reproduced 2/2 failures, then 2/2 pass.
- Unmodified round-2 reviewer probes: 13/13 pass.
- Checked-in takeover suites: 24/24 pass. Removing only the resolver call
fails both new reload cases; source restored afterward.
- RPC orchestration + terminal, aggregate runtime handle registry,
handle incarnation, mobile tab mount, stale geometry, and reload probe:
2027 passed, 1 skipped (89 files).
- tc:node and check:code-quality:changed pass; background launch enabled.
* test(rpc): require unconditional terminal afterWrite callbacks
Update exact sendTerminal expectations for the round-2 accepted-write
contract. Preserve beforeWrite expectations, absence of reserveWrite,
byte payloads and call-count checks; require afterWrite to be a function.
Reproduced the requested two-file run: 5 failed, 31 passed. The full RPC
suite exposed the same stale shape in ACK budget/overflow, desktop resize
(including its later retry), and agent-prompt fallback assertions. Update
those too, for 11 assertions across six test files. No production changes.
Validation: ORCA_BACKGROUND_LAUNCH=1 full src/main/runtime/rpc suite:
264 files passed; 2292 tests passed, 1 skipped. Changed-code quality and
staged oxlint/React Doctor/oxfmt checks passed. Ran lint-staged --no-stash
manually to honor checkout safety rather than its default backup hook.
* fix(mobile): report worker takeover outside terminal byte delivery
New phones announce accepted real user input through the existing worker
report RPC, addressed by terminal handle. Share a per-client/per-handle
30-second gate with one bounded retry; report through the same RPC client
as the input. Cover live commits and dictation via their shared sender,
accessory keys, gestures, buffered submit, paste and accepted native chat.
Query replies, attachment heals, triage and diff-review sends do not report.
Phones predating this build do not fence release.
Remove byte provenance and takeover callbacks from host delivery. Restore
both lanes' pre-PR floor-claim plumbing and the original options assertions.
Keep the host recorder uncached with its conditional resume-fence sweep.
No DB schema or stream change; terminal is an optional report address.
Retain the shared resolver recovery independently of takeover: the new
SSH inspection test fails without it during renderer reload. Other callers
still benefit for subscription, resize, viewport and exact-PTY binding;
unary driver/lock checks see the retained PTY. Pane routing and dispatch
authority already recover through getLivePtyForHandle and are unaffected.
Existing leaf generation checks and first-PTY adoption remain unchanged.
No other input-plumbing hunk is retained relative to the PR base.
Replace byte-takeover tests with handle-addressed local/SSH report and
unknown-handle tests, plus real unary/stream writes asserting zero SQL
prepare/exec calls. Mobile send-site integration covers reports, exclusions,
rejected writes and gate counts. Desktop report tests are unchanged.
Register replacement coverage in the settled-worker release manifest.
Validation (all background): host/RPC/runtime 3541 passed, 2 skipped;
mobile session/terminal 2045 passed; node and mobile typechecks, changed
quality, mobile oxlint, reliability manifest and max-lines ratchet passed.
All five requested mutations fail assertions; resolver revert also fails
independent inspection. Staged checks run manually with --no-stash.
Final src diff against PR base: 5 files, +165/-13 (previously +839/-85).
* fix(runtime): allow the takeover report from mobile-scoped tokens
The mobile RPC allow-list gates every phone request before dispatch and the
reporter swallows a refusal, so without this entry every phone shipped
unfenced. Pin it beside the report tests, and pin the once-per-takeover
fence sweep the replaced byte-lane suite used to assert.
* fix(mobile): a no-op takeover report does not arm the gate; Stop reports too
A key during worker startup reports before the resource is owned; caching
that zero-change reply for 30 s suppressed the report that would have fenced
the worker once it attached. Native-chat Stop is deliberate input and now
reports on an accepted Escape.
* fix(mobile): takeover gate ignores the host answer, like desktop
Reopening the gate on a zero-change reply made every accepted key on an
ordinary terminal an RPC plus a host write transaction (round 6: 100 for
100). The startup window it closed is unreachable: the agent has no prompt
to accept input until after its resource row exists. Plain terminals now
pay one report per 30 s window; the native-chat Stop report stays.
Send-site fixture answers the report RPC with a changed count; the draft
test filters to terminal.send calls.
* docs(runtime): say why resolveLiveLeafForHandle re-links before lookup
* chore(i18n): regenerate the runtime-required catalog for the contrast floor strings
* test(orchestration): give the stopping-worker guard fixtures a Run
* test(orchestration): drop fence-sweep assertions retired by the settled-worker policy
* test(orchestration): pin the mid-boot phone takeover that #19608 makes possible
A handle-addressed report during the worker's tui-idle wait now finds the
custody row written at terminal creation, so it flips the pane to user_owned
and worker-release retains it instead of closing it under the user.
* feat(terminal): make the contrast floor user-configurable (#10754)
The xterm minimumContrastRatio floor was hardcoded (3 on dark backgrounds,
4.5 on light) and applied to every pane with no way out, so TUIs that use
deliberately low contrast were rewritten: Powerline separators drawn in the
neighbouring segment's background became visible seams, and dimmed secondary
text lost its hierarchy.
Adds an optional `terminalMinimumContrastRatio` setting under Settings ->
Terminal -> Rendering. Blank keeps today's automatic, background-luminance
gated floor; 1 disables correction entirely (matching VS Code's documented
`terminal.integrated.minimumContrastRatio` and iTerm2's off-by-default
Minimum Contrast); values are clamped to xterm's 1-21 range.
The floor is resolved in one place, so live panes, the Appearance preview
and the dashboard terminal preview all follow it, and the existing
value-gated write still avoids clearing xterm's contrast cache on no-op
re-applies. The clamp also lives at the persistence boundary that every
writer crosses, so a hand-edited profile or CLI write can never hand xterm
a non-finite option. Mobile mirrors the desktop gate, so the resolved floor
travels with the terminal theme payload as a new optional field; hosts that
omit it leave older and newer clients on the luminance gate.
Fixes#10754.
Co-authored-by: Nyanako <44753291+Nanako0129@users.noreply.github.com>
* fix(terminal): refresh mobile payload fixture and clarify contrast target
* feat(terminal): make contrast controls intent-based with custom tuning
---------
Co-authored-by: Nyanako <44753291+Nanako0129@users.noreply.github.com>
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
* Update PR checks fix prompt to verify failure causality before fixing
Revise the prompt to classify failures as caused by this branch, not caused,
or uncertain before making changes. Only proceed autonomously for confirmed
failures; ask the user for guidance on uncertain or unrelated issues to avoid
fixing failures that weren't caused by the branch.
* Update PR checks fix prompt to verify failure causality before fixing
- Emphasize investigation phase by reframing prompt: "Investigate" rather than "Fix"
- Extend untrusted-data warning to all investigation sources (repository files, commit messages, diffs, CI output)
- Add test verifying injection safety: malicious input confined to JSON payloads, never as prompt instructions
* Refactor buildFixChecksPrompt test to focus on field mapping
The wrapper's only responsibility is renaming mobile PR fields onto the
shared prompt builder. Remove assertions about prompt wording, which are
already covered by the builder's own test suite. Simplify the test to
verify the field mapping contract and nothing else.
* Revert "feat(mobile): time relay dial stages so diagnostics say where a slow connect went (#19245)"
This reverts commit 83b1558ecc.
* Revert "perf(mobile): race the direct and relay dials from t=0 on every reconnect (#19308)"
This reverts commit ceafdcad2f.
* Revert "feat(mobile): draw the last known tab strip while a session reconnects (mobile pass) (#19281)"
This reverts commit 643571def6.
* Revert "perf(mobile): open a session with parallel startup RPCs and a pre-warmed terminal engine (#19260)"
This reverts commit c37413271e.
* Revert "perf(mobile): cut the relay reconnect critical path and admit dead sockets faster (mobile pass) (#19280)"
This reverts commit e628090ad4.
* chore: keep the react-doctor suppression for the startup timers
The pattern it covers (a variable number of timers cleared through one cleanup)
predates #19260 and is unchanged by the revert; dropping the entry only re-exposed
a pre-existing finding to the changed-code gate.
* feat(mobile): time relay dial stages so diagnostics say where a slow connect went
A 10s connect was unattributable from a shared report. Relay dial stages carried
no timestamps, so nothing could tell "the cell never answered relay-hello" from
"the E2EE handshake was slow", and the per-state dweltMs the client already
computed went only to console.log — invisible without a debug build.
RelayDialStageTracker now stamps each stage entry from a monotonic clock
(performance.now where present, wall clock otherwise) and returns the duration of
the stage it just left. The session logs one entry per stage, and settles the
in-flight stage on connect, failure, or close, so a dial that dies mid-way still
names the stage it never finished. dweltMs joins the same buffer as a structured
field instead of console.
Durations ride the existing per-host log buffer and its cap, so memory is
unchanged and no new storage appears. The report derives two lines from them: the
latest dial's stage breakdown (a reconnect loop must not average away the attempt
being reported) and total dwell per connection state. Both are numbers and
closed-enum names, and the entries still pass through the existing redaction.
* fix(mobile): never let a diagnostics sink break a dial, and pin timing names to their enums
Review follow-ups on the dial-stage timing work.
The stage timing emitted on the confirm's success path ran inside the try that
calls fail(), so an onLog sink that threw would have turned a good connect into a
failed session. The same hazard existed on the direct path, where the dwell emit
sits in publish() ahead of the listener loop and the connect waiters. Both sink
calls are now isolated: a broken sink loses a log line and nothing else.
The persisted-log validator accepted any string as a timing name, and the report
echoes that name unredacted. Names are now checked against the closed enum for
their kind, backed by Record<Union, true> tables so adding a stage or a state
breaks the build rather than silently widening what a corrupted store can inject.
Entry volume: every reconnect cycle walks four connection states, so logging each
one would roughly double what a slow-connect report holds against the unchanged
200-entry per-host cap. Transitions under 100ms are therefore not buffered. They
cannot be where a slow connect spent its time, and console still shows all of
them. States that flap slowly, which is the case support cares about, still land
in the log.
RpcClientConnectionState takes an optional clock so dwell thresholds are testable
without sleeping.
* fix(mobile): reject a negative stored stage duration when hydrating the log
A persisted timing only had to be finite to survive hydration, so a corrupted
`ms: -1` reached the diagnostics report, where the dial summary sums the stage
durations and a negative would subtract from the total. Producers clamp at 0
(`elapsedMs`), so anything below it is corruption. 0 itself still hydrates: a
stage the dial passes through instantly is real.
* refactor(mobile): move the relay liveness profile out of the session so the dial log fits
* fix(mobile): never let the liveness-timeout log line keep a dead relay connected
* test(mobile): prove the throwing timeout sink was actually reached
* perf(mobile): race the direct and relay dials from t=0 on every reconnect
A foreground reconnect gave the direct dial a fixed 2.5s head start, and while
that dial sat in 'connecting'/'handshaking' the supervisor refused to open a
relay socket at all. A phone that is off the LAN paid the full head start on
every reconnect and got nothing for it, and a phone whose relay dropped could
only return to the LAN through three hysteresis probes.
Both dials now start together and the first authenticated socket is adopted
through the existing migrateTo cutover. Nothing about the migration machinery
changes: only who is allowed to start a dial.
- The relay dial now yields to a live session and to nothing else. An unfinished
direct dial is progress on the other runner, not a reason to stand still.
- The direct return probe grows a second adoption policy. Against a live relay
hysteresis still has to prove direct stable; during a reconnect there is no
session to protect, so an authenticated direct socket wins outright. probeNow
pre-empts a pending 15s tick so that dial starts with the relay dial, not
after it, and the dial itself no longer waits for the operation mutex — a
relay dial holding it is exactly the case the race exists for.
- A loser closes and books nothing. The relay dial withdraws inside migrateTo
and returns 'aborted', so no backoff is booked against it; a direct socket
that loses leaves the promotion streak untouched. Only a reconnect that both
paths lose books a failure, once, on the relay cadence.
Kept: the 30s background grace and the foreground gate, because a backgrounded
phone must not open a billed relay splice; the shared failure cooldown, because
a genuine relay failure still has to be paced; the hysteresis dwell after a
migration, because it is what stops a marginal LAN flapping a healthy session.
The accepted cost is one relay socket per reconnect for a phone that is on its
LAN. It closes as soon as the direct path authenticates, before the resume
confirm, because migrateTo only checks the abort predicate after E2EE auth.
Tests that encoded the removed rules:
- 'fails over when the direct retry loop publishes reconnecting' asserted no
relay dial while direct was handshaking. The failover now precedes the direct
client giving up, so it asserts the dial instead of its absence.
- 'does not spend a queued relay retry while direct authentication is
progressing' encoded the block outright; it now asserts the retry runs on the
failure cadence while a handshake drags on.
- The four grace-race cases move to mobile-endpoint-reconnect-race.test.ts as
t=0, direct-wins, background/resume and both-lose cases.
- Five relay-bookkeeping cases now state their premise with unreachableDirect.
They describe a phone with no LAN, which used to be implicit and is now
load-bearing: with a reachable LAN the direct socket wins those reconnects.
* fix(mobile): withdraw a lost relay dial pre-handshake and damp blip races
Review follow-up to 4e31130471. Racing both paths from t=0 was correct but
charged the LAN case twice: once per reconnect in cell work, and again whenever
the LAN flapped.
Withdraw before the handshake. migrateTo only consults its abort predicate after
E2EE authentication, so a dial that had already lost still made the cell reserve
a splice and the desktop finish a key exchange. The establisher now watches the
logical client across the dial and closes the cell socket the moment direct
authenticates. In the common window, after relay-auth is on the wire and before
the hello lands, nothing of the key exchange has started, so the withdrawal
costs the desktop nothing. The dial still reports itself aborted and still books
nothing. The watch is dropped once migrateTo returns, because past the cutover
this session is the active path and a later direct promotion must not read as a
reason to close the client's own socket.
Damp races that a blip started. relayDialAllowed yields only to a live session
and a lost race books nothing, so a flapping LAN drove one cell socket per blip
with only the relay's per-host rate limiter as a backstop, and reaching that
limiter would have converted a benign race into a booked relay failure. After a
race is lost to direct, the next unforced race is suppressed for 2s, doubling
per consecutive loss to a 30s cap. This is not backoff and is kept separate from
it: a forced replacement is never damped, a relay dial that wins clears the
streak, and a foreground resume clears it too, so the path the user is watching
never waits. The window arms its own lapse timer, so a LAN that dies inside the
window still reaches relay without a new trigger.
A superseded cutover no longer escapes probe() as an unhandled rejection. Only
the probe timer calls it, and it discards the promise, so the routine end of a
lost race would have surfaced as one.
Credential rotation moves to MobileRelayCredentialRefresh. The supervisor
crossed the 300-line cap; rotation is a self-contained responsibility that only
runs over a live direct connection, so it splits cleanly instead of taking a
max-lines bump.
* fix(mobile): end a damper window as soon as the direct path is really gone
Round-2 review follow-up to 9a21da4326. The damper armed its window when direct
won the race, and nothing shortened it. A LAN that died inside that window left
the phone waiting out the whole thing, up to 30s at the cap, with only a log
line to show for it. My previous commit body claimed the path the user watches
never waits; that was true only of a foreground resume, and it is corrected
here.
Losing the direct path now collapses the wait to a 250ms floor, so the next
recovery races almost at once. The floor is not zero because the reason the
damper exists is a LAN that drops and comes straight back, and a disconnect is
how such a blip begins. So the rest of the window is kept aside rather than
spent: if direct returns inside the floor it was a blip and the window resumes,
and if the floor lapses with direct still gone it was an outage and the held
window is void. Without the second half, one blip would have bought a flapping
LAN a free pass on every race that followed, which is the case the damper was
added for.
The streak itself is untouched by the clamp. A LAN that flaps all afternoon
still escalates toward the cap; only the current wait is cut short.
record() now takes the same forceReplacement guard as suppresses(), so a forced
replacement that stands down cannot grow the streak or be read as a loss to
direct. A lease rotation or a reconsidered network change is not a LAN that
flapped.
Also documents that a genuine relay failure deliberately does not reset the
streak, and that the damper and the failure backoff serialize rather than stack:
a damped attempt never reaches the dial that would book a cooldown.
* fix(test): give the direct-probe fixture the race-era hooks
The phase-1 probe test predates canDial and adoptsOutright, so its hooks
literal threw at the first dial. These cases model a live relay session.
* docs(mobile): say why a finished credential refresh races relay instead of waiting on direct
* feat(mobile): draw the last known tab strip while a session reconnects
Reopening a workspace the phone has already visited threw away everything
it knew. The route clears its tabs on mount, so until the reconnect lands
and the first snapshot is applied the session screen has an empty header
and a bare spinner, even though the strip it is about to be handed is the
one it drew a minute ago.
Persist the four fields the strip actually draws -- id, type, title, agent
-- per host and workspace, and add a reconnecting-with-cache shape to the
route state so those rows render immediately, disabled, under the ids the
live snapshot will reuse. Live tabs always outrank the cache, so a
mid-session drop keeps its mounted terminals; an exhausted retry loop or a
rejected pairing outranks it the other way, because a strip the user cannot
reach is worse than the existing offline affordance. With nothing cached
the screen behaves exactly as before.
The body stays a placeholder. Replaying stored scrollback into the terminal
WebView would double-render the same rows once the live stream replays them,
so the strip is the cached content and the body waits for the stream.
* fix(mobile): keep shell titles and unpaired hosts out of the cached tab strip
Review of the reconnect strip cache found two ways it leaked.
A terminal's title is whatever the shell last set, which is routinely the
command line: a psql URL with an inline password, a curl with a bearer
token. Both fit well inside the 64-character cap and both were written to
plaintext AsyncStorage verbatim. Browser tabs carried their page title the
same way. Terminals and browsers now collapse to a fixed label, with a
resolved agent naming itself because that lookup is a closed enum. The rule
lives in the storage module rather than its caller, so it holds for entries
an older build already wrote, and a tab type this build cannot draw is
dropped instead of having its title trusted.
The cache also survived forgetting a host. Nothing expired an entry, and
the module-global memory map meant a later save from any surviving host
serialized the forgotten host's rows straight back to disk. Both cleanup
paths now evict by host, dropping the in-memory rows and rewriting storage,
with a pending debounced write cancelled so it cannot restore them.
Also: the storage key digests the workspace id, which ended in a filesystem
path, and cached rows carry the same de-emphasis as the disabled tab-bar
buttons beside them, so an inert row does not pass for a live one.
* fix(mobile): make a forgotten host's cached tab strip actually leave disk
Review finding on this PR, fixed here so it rides along with the rest.
writeFile swallowed its own rejection, so deleteCachedSessionTabStripForHost
resolved successfully while the unpaired host's plaintext tab titles stayed on
disk, and removeHostAndCloseClient discarded the promise with void so nothing
could have observed the failure anyway.
The write now throws. The debounced save keeps a best-effort catch, since a
dropped cache refresh costs one repaint and the next save rewrites the whole
map, so only the deletion path needs the failure. Host removal awaits the
deletion and logs a failure but never rethrows: the metadata removal has
committed and the client is closed by that point, so reporting a finished
removal as failed would be wrong. The unpaired-host credential sweep already
awaited the deletion and now sees the rejection, consistent with its sibling
credential deletions.
Two ways the rows could come back are closed as well. The cache refuses saves
for a host it has been told to forget, so a snapshot racing the deletion cannot
re-insert it, and the deletion awaits any debounced write already on the wire,
since that write built its blob from the map as it was and would otherwise race
the purge for the last word on disk. The refusal lasts for the process, so
re-pairing the same host caches again from the next app launch, which is the
cheap direction for a deletion the user asked for.
* fix(mobile): order the tab-strip cache writes so a purge is the last word
Two debounced writes could sit on the AsyncStorage bridge at once, and the
second replaced the in-flight handle. A host purge then awaited only the newer
write, so the older blob -- snapshotted while the forgotten host was still in
the map -- could commit after it and restore the host's titles to disk. Writes
now queue behind one chain and the purge queues last.
The unpaired-credential sweep also aborted on a cache-purge failure, stranding
the write revision and onDeleted after every credential was already deleted. It
now warns and finishes, as removeHostAndCloseClient already did.
Startup RPCs now fan out in parallel and the xterm engine pre-warms inside the
real terminal frame while they are in flight, so the first pane inherits a warm
WebView and an already-measured viewport instead of paying a round trip for it.
The pre-warm opens its engine before measuring: web-ready only reports that the
bundle loaded, and the WebView answers a measure with null until a terminal
exists. It also pre-warms at the user's saved text size, because cell size is
what the frame height gets divided by.
Host writes such as worktree.activate wait for an evaluated status.get reply.
Navigation still fails open when a host cannot answer one, but that fallback no
longer reads as a passing compatibility verdict.
* perf(mobile): cut the relay reconnect critical path and admit dead sockets faster
Phone medians put E2EE authentication at ~424ms but `connected` at ~630ms,
because the session serialized two RPC round trips behind it: the resume
confirm (`pairing.getEndpoints`) and the capability advisory. Both now ride
the authenticated socket concurrently and off the critical path, so the
session publishes `connected` as soon as E2EE authenticates. Peer identity
is already proven by then — the confirm carries credential/lease bookkeeping
and the cell assignment check, and it still fails the session on a bad answer
or a foreign relayHostId, only later. `persistResumeConfirmation` awaits the
new `whenResumeConfirmed()` instead of assuming the answer is present at
`connected`.
Foreground liveness on a retained relay: `notifyForeground('app-resume')`
now probes past the 10s voluntary minimum on urgent bounds (2s, one miss),
so a socket that died while the process was suspended is admitted in ~2s
instead of ~8s. Focus and network nudges keep the old minimum and bounds.
Relay sessions also gain a 25s idle sweep, gated on foreground so a
backgrounded app spends no probes.
Recovery is no longer blocked by the direct return probe. The probe's 12s
dial is a pure observation on its own socket, so it takes the supervisor's
operation mutex only for the cutover; a relay recovery landing during a
foreground return now starts immediately instead of waiting the budget out.
Requests that do land during the cutover are queued in a new
RelayRecoveryIntentQueue and replayed on release — an owning forced
replacement keeps its intent, everything else replays as a plain recovery.
Tests updated deliberately, for the new ordering:
- 'sends no periodic traffic while an authenticated relay is idle' asserted
the absence of any relay idle probe, which is exactly the gap D3 closes.
Replaced by a sweep test plus a backgrounded no-probe test.
- 'rate-limits foreground sequences without suppressing a retry' asserted
that app-resume was suppressed inside the 10s minimum. An app resume is
now the one nudge that must never be rate-limited.
- the session helpers waited for the confirm answer before `connected`;
they now authenticate, read both concurrent frames, and settle them.
* fix(mobile): book backoff when a relay resume confirm fails after the cutover
Review round 1 on 352bfd2300.
P1: publishing `connected` at E2EE authentication made `migrateTo` resolve
before the resume confirm answered, so a confirm that failed afterwards —
a `relayHostId` mismatch from a rehomed desktop is the live case — was still
reported as an `established` dial. registerFailure was skipped, no cooldown
was booked, recordMigration()/setActiveSession() ran for a dying session, and
the queued-recovery replay redialled immediately: a tight loop with a
connected→disconnected blip per pass. The establisher now awaits
whenResumeConfirmed() after the cutover and, if the session is no longer
connected, reports a failed dial (or an aborted one when direct won or the
supervisor went inactive) exactly as a rejected migrateTo used to. The UI
still connects early; only the supervisor's bookkeeping waits.
The state check, rather than getFailure(), is the oracle: a live session can
carry a latched failure without having failed yet, and "is this session still
alive once the confirm settled" is precisely the question migrateTo used to
answer.
P2: the resume probe profile goes to two 2s misses instead of one. The first
frame after a resume rides a cold radio and a possibly distant cell, so one
slow answer is not proof of a dead link; the verdict still lands at 4s rather
than the previous 8s.
Nits: the direct probe's two early returns no longer close the candidate the
finally also closes (the second shape pre-existed); RelayRecoveryIntentQueue
is cleared in the supervisor's stop().
Mutex-hold note: persistResumeConfirmation, and now the establisher's own
await, are bounded by the confirm's request timeout. That would have been the
session's 30s default, so the confirm is pinned to RELAY_CONFIRM_TIMEOUT_MS
(12s) — the same bound migrateTo's waitForAuthenticated applied before.
Test: a supervisor-level case where every dial authenticates then fails the
confirm must book 250/500/1000ms backoff with no immediate redial, and must
never record a migration. It fails on the pre-fix establisher.
* fix(mobile): close three relay probe and liveness gaps from review
Review findings on this PR, fixed here so they ride along with the rest.
Direct return probe: schedule() guarded only the pending timer, so a caller
asking for an immediate probe while a dial was in flight started a second one
that overwrote activeProbe. stop() then reached only the newest socket and left
the earlier dial running out its 12s budget. Releasing the operation mutex for
the dial removed the only thing that had been serializing probes, and the
background bounce hits it directly: background() cancels the timer but leaves an
in-flight dial alone, and the matching foreground return asks for a probe at
once. The in-flight probe now owns the next slot and re-arms on the soonest
delay any caller asked for, so an urgent request is deferred rather than dropped
on the 15s floor.
Liveness watchdog: both retry paths in handleProbeTimeout, the tolerated-miss
one and the unfair-window one, retried without rechecking shouldIdleProbe. An
idle-sweep probe that started in the foreground could therefore keep spending
probes after the app backgrounded and terminate a healthy relay on misses that
were really iOS suspending the socket, which is the exact reading the foreground
gate exists to prevent. Probes now carry their origin, and an idle-sweep probe
that times out while backgrounded clears its state and re-arms the sweep with no
misses carried forward. Caller probes still reach a verdict.
Relay RPC session: whenResumeConfirmed() handed a pre-authentication caller an
already-resolved promise, so the documented contract only held after
authentication. No caller can reach that window today, since publishAuthenticated
assigns the promise before publishing 'connected' and both readers run after
migrateTo resolves, but the type comment promised more than the code delivered.
The deferred now exists from construction and settles on the confirm, on fail(),
and on close(), which are the only ways the session can end. Both endings had to
settle it and already shared nearly all of their teardown, so they are unified
behind one terminate().
* fix(mobile): give a resume probe its own miss budget
A resume probe supersedes an ordinary probe already in flight, but startProbe
carried the ordinary profile's missedProbes across the switch. Relay uses 2
misses for both profiles, so one earlier 4s miss plus a single slow 2s answer
terminated the session -- consuming the tolerated cold-radio answer the urgent
profile exists to provide. Switching profile now resets the count.
* Revert "feat(mobile): draw the last known tab strip while a session reconnects (#19258)"
This reverts commit 0ba7f8dc8d.
* Revert "perf(mobile): cut the relay reconnect critical path and admit dead sockets faster (#19236)"
This reverts commit 23df74d85a.
* feat(mobile): draw the last known tab strip while a session reconnects
Reopening a workspace the phone has already visited threw away everything
it knew. The route clears its tabs on mount, so until the reconnect lands
and the first snapshot is applied the session screen has an empty header
and a bare spinner, even though the strip it is about to be handed is the
one it drew a minute ago.
Persist the four fields the strip actually draws -- id, type, title, agent
-- per host and workspace, and add a reconnecting-with-cache shape to the
route state so those rows render immediately, disabled, under the ids the
live snapshot will reuse. Live tabs always outrank the cache, so a
mid-session drop keeps its mounted terminals; an exhausted retry loop or a
rejected pairing outranks it the other way, because a strip the user cannot
reach is worse than the existing offline affordance. With nothing cached
the screen behaves exactly as before.
The body stays a placeholder. Replaying stored scrollback into the terminal
WebView would double-render the same rows once the live stream replays them,
so the strip is the cached content and the body waits for the stream.
* fix(mobile): keep shell titles and unpaired hosts out of the cached tab strip
Review of the reconnect strip cache found two ways it leaked.
A terminal's title is whatever the shell last set, which is routinely the
command line: a psql URL with an inline password, a curl with a bearer
token. Both fit well inside the 64-character cap and both were written to
plaintext AsyncStorage verbatim. Browser tabs carried their page title the
same way. Terminals and browsers now collapse to a fixed label, with a
resolved agent naming itself because that lookup is a closed enum. The rule
lives in the storage module rather than its caller, so it holds for entries
an older build already wrote, and a tab type this build cannot draw is
dropped instead of having its title trusted.
The cache also survived forgetting a host. Nothing expired an entry, and
the module-global memory map meant a later save from any surviving host
serialized the forgotten host's rows straight back to disk. Both cleanup
paths now evict by host, dropping the in-memory rows and rewriting storage,
with a pending debounced write cancelled so it cannot restore them.
Also: the storage key digests the workspace id, which ended in a filesystem
path, and cached rows carry the same de-emphasis as the disabled tab-bar
buttons beside them, so an inert row does not pass for a live one.
* perf(mobile): cut the relay reconnect critical path and admit dead sockets faster
Phone medians put E2EE authentication at ~424ms but `connected` at ~630ms,
because the session serialized two RPC round trips behind it: the resume
confirm (`pairing.getEndpoints`) and the capability advisory. Both now ride
the authenticated socket concurrently and off the critical path, so the
session publishes `connected` as soon as E2EE authenticates. Peer identity
is already proven by then — the confirm carries credential/lease bookkeeping
and the cell assignment check, and it still fails the session on a bad answer
or a foreign relayHostId, only later. `persistResumeConfirmation` awaits the
new `whenResumeConfirmed()` instead of assuming the answer is present at
`connected`.
Foreground liveness on a retained relay: `notifyForeground('app-resume')`
now probes past the 10s voluntary minimum on urgent bounds (2s, one miss),
so a socket that died while the process was suspended is admitted in ~2s
instead of ~8s. Focus and network nudges keep the old minimum and bounds.
Relay sessions also gain a 25s idle sweep, gated on foreground so a
backgrounded app spends no probes.
Recovery is no longer blocked by the direct return probe. The probe's 12s
dial is a pure observation on its own socket, so it takes the supervisor's
operation mutex only for the cutover; a relay recovery landing during a
foreground return now starts immediately instead of waiting the budget out.
Requests that do land during the cutover are queued in a new
RelayRecoveryIntentQueue and replayed on release — an owning forced
replacement keeps its intent, everything else replays as a plain recovery.
Tests updated deliberately, for the new ordering:
- 'sends no periodic traffic while an authenticated relay is idle' asserted
the absence of any relay idle probe, which is exactly the gap D3 closes.
Replaced by a sweep test plus a backgrounded no-probe test.
- 'rate-limits foreground sequences without suppressing a retry' asserted
that app-resume was suppressed inside the 10s minimum. An app resume is
now the one nudge that must never be rate-limited.
- the session helpers waited for the confirm answer before `connected`;
they now authenticate, read both concurrent frames, and settle them.
* fix(mobile): book backoff when a relay resume confirm fails after the cutover
Review round 1 on 352bfd2300.
P1: publishing `connected` at E2EE authentication made `migrateTo` resolve
before the resume confirm answered, so a confirm that failed afterwards —
a `relayHostId` mismatch from a rehomed desktop is the live case — was still
reported as an `established` dial. registerFailure was skipped, no cooldown
was booked, recordMigration()/setActiveSession() ran for a dying session, and
the queued-recovery replay redialled immediately: a tight loop with a
connected→disconnected blip per pass. The establisher now awaits
whenResumeConfirmed() after the cutover and, if the session is no longer
connected, reports a failed dial (or an aborted one when direct won or the
supervisor went inactive) exactly as a rejected migrateTo used to. The UI
still connects early; only the supervisor's bookkeeping waits.
The state check, rather than getFailure(), is the oracle: a live session can
carry a latched failure without having failed yet, and "is this session still
alive once the confirm settled" is precisely the question migrateTo used to
answer.
P2: the resume probe profile goes to two 2s misses instead of one. The first
frame after a resume rides a cold radio and a possibly distant cell, so one
slow answer is not proof of a dead link; the verdict still lands at 4s rather
than the previous 8s.
Nits: the direct probe's two early returns no longer close the candidate the
finally also closes (the second shape pre-existed); RelayRecoveryIntentQueue
is cleared in the supervisor's stop().
Mutex-hold note: persistResumeConfirmation, and now the establisher's own
await, are bounded by the confirm's request timeout. That would have been the
session's 30s default, so the confirm is pinned to RELAY_CONFIRM_TIMEOUT_MS
(12s) — the same bound migrateTo's waitForAuthenticated applied before.
Test: a supervisor-level case where every dial authenticates then fails the
confirm must book 250/500/1000ms backoff with no immediate redial, and must
never record a migration. It fails on the pre-fix establisher.