Commit Graph
10478 Commits
Author SHA1 Message Date
Merge Sim bd0807e1e5 fix(native-chat): honour notify on the relabel path and emit the stored snapshot
The startup sweep republishes every restored session with notify: false, but
the relabel branch never consulted it and emitted a full tab list per named
session — the desktop renderer is a real subscriber on that feed. It also
emitted the pre-store candidate, although the store can hand back a different
object, so a client mirror could retain a replaced tab under an identical
snapshotVersion; the sibling replace path already got this right.

Also corrects the merge comment that claimed a background republish cannot
re-surface a lost mirror, moves the doc block that landed between an existing
comment and the field it described, and routes the replacement tab's
placeholder through defaultAgentChatLabel.
2026-09-07 11:31:06 -07:00
Merge Sim 12e1f913f8 fix(native-chat): strip unrenderable characters from a conversation name
The collapse used \s, which leaves C0 controls, zero-width runs and bidi
overrides in a name that reaches the tab strip and the sidebar row — a U+202E
renders a label that reads as text the name does not contain. Truncation used a
raw slice, which can strand a lone high surrogate. Reuses sliceAtCodeUnitLimit,
moved to src/shared so both callers can reach it.
2026-09-07 11:28:39 -07:00
Merge Sim 97830d40a1 fix(native-chat): harden the Codex naming thread and narrow its request gate
Open the throwaway naming thread with approvalPolicy 'never' and sandbox
'read-only'. The prompt embeds untrusted user message text and the thread's
frames are diverted from the journal, so any tool the host would auto-approve
ran where the user could never see it; refusing server requests only covered
the tools that ask.

Require an exact naming-thread id on the server-REQUEST path. During the
thread/start window no naming thread has a turn running, so the broad pre-id
rule could only ever match a genuine sub-agent, whose approval request was
auto-refused with -32001 instead of reaching the user.

Also caps the structured answer before parsing it, normalizes the generated
name before it is written to the user's real thread, and releases the
throwaway thread with the protocol's own thread/unsubscribe.
2026-09-07 11:25:12 -07:00
Merge Sim c868b9f05e fix(native-chat): stop a text-free first message burning the naming attempt
Both providers claimed namingAttempted before deriving the prompt text, so a
caption-free screenshot as the first message left the chat on its placeholder
for the session's whole life. Claim the flag after the text is in hand; the
prompt reader is now total by construction, so deriving it before the promise
cannot fail a delivered message. Also reattaches the JSDoc the max-lines
extraction left on the wrapper, and softens a 'never titled on its own' claim
to what was actually observed.
2026-09-07 11:22:34 -07:00
Merge Sim 5f4b07a63a fix(native-chat): stop the transcript tail scan corrupting and mis-clearing names
Carry the chunk boundary partial as bytes, so a multi-byte character straddling
a 64KiB boundary is no longer decoded to U+FFFD on both sides and written
durably as the tab's name. Report whether the scan reached the start of the
file, and only read an emptied custom title as a clear when it did — an
ai-title beyond the bound is still the conversation's name. Type-guard
aiTitle instead of coercing it, which turned an object into [object Object].
2026-09-07 11:19:54 -07:00
Merge Sim 136fd21da7 Merge remote-tracking branch 'origin/main' into brennanb2025/native-chat-conversation-name
# Conflicts:
#	src/main/claude/claude-structured-dispatch.test.ts
#	src/main/claude/claude-structured-options.test.ts
#	src/main/claude/claude-structured-session-publication.ts
#	src/main/claude/claude-structured-session-state.ts
#	src/main/native-chat/agent-session-wire/structured-agent-session-host-teardown.ts
#	src/main/native-chat/agent-session-wire/structured-agent-session-host.ts
#	src/main/runtime/agent-session-record-store.ts
#	src/main/runtime/agent-session-visible-tab-index.ts
#	src/main/runtime/orca-runtime-get-worktree-ps.ts
#	src/main/runtime/orca-runtime-restore-structured-agent-session-tabs-once.ts
#	src/main/runtime/structured-agent-session-runtime.ts
#	src/main/runtime/structured-claude-runtime-adapter.ts
#	src/renderer/src/app-shell/app-command-handlers-tab-rename.test.ts
#	src/renderer/src/app-shell/app-command-handlers.ts
#	src/renderer/src/components/native-chat/StructuredAgentSessionStatusBridge.tsx
#	src/shared/agent-session-record.ts
2026-09-07 10:08:43 -07:00
Brennan BensonandMerge Sim bffdad9f05 fix(native-chat): make structured chat tabs renameable (#19153)
* fix(native-chat): let a structured chat tab be renamed

Renaming a native chat tab accepted the text and silently did nothing:
setTabCustomTitle only scanned terminal tabs and only bridged to unified
tabs whose contentType was 'terminal', so the agent-session tab it was
keyed to never matched. Any label that did land was then re-nulled by the
next host snapshot, which preserved color/createdAt/isPinned but not
customLabel.

Also routes both placeholder sites through one helper so a Claude chat
stops falling back to 'Codex Chat'.

* test(native-chat): cover structured chat tab rename and label fallback

* chore: drop the local @pnpm/exe lockfile artifact

Swept in accidentally; running pnpm here adds @pnpm/exe to the root
lockfile, which fails CI's frozen-lockfile guard.

* fix(native-chat): reach the rename shortcut and tab color too

Review found the first fix covered only the context-menu path. The
tab.rename shortcut gated on activeTabType === 'terminal', so on a
structured chat tab it stayed the silent no-op this branch set out to
fix. setTabColor carried the identical terminal-only lookup one function
below the one that was fixed.

Both lookups now share one resolver instead of two copies.

* fix(native-chat): stop unknown agents reading as Codex, cover the terminal path

Review found the placeholder helper encoded "unknown means Codex": its
signature accepts null/undefined and Tab.agentSessionAgent is the open
AgentType, so the first caller passing a Tab would label gemini or grok
as "Codex Chat". Routed through the shared agent-name table instead.

Also adds the missing regression test that a terminal rename still
resolves through its entityId now that both rename and color share one
resolver, and a guard on a test that passed with the fix reverted.

* fix(native-chat): degrade instead of throwing on a null tab title

A stacked branch can publish title: null when a conversation name is
cleared. The wire type says string, so this consumer trusted it and
threw inside the store patch that applies the snapshot. Fall back to the
placeholder — the producer bug is fixed separately, but a consumer of
wire data should not crash on a contract violation.

* fix(native-chat): rename the focused structured tab, not a background terminal

* fix(native-chat): cycle terminals from the structured tab, not a stale terminal

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-07 09:36:14 -07:00
6ae5418a89 Add localization for activity view and sidebar (#18589)
* i18n: add localization for activity view and sidebar

Wrap activity thread state labels, interrupted status, and sidebar
title in translate() calls. Add localization keys to all five locale
catalogs (en, es, ja, ko, zh) to enable translation support.

* i18n: refactor to static keys for activity and sidebar

Convert dynamic translation key construction to static literal keys, enabling proper i18n catalog registration. This ensures activity state labels and sidebar strings are bundled in the boot catalog with their complete translations.

* i18n: change permission state label to 'Needs attention'

- Rename state label for semantic clarity across all locales
- Remove strings now using static keys (per i18n refactor to static keys)

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-07 09:33:23 -07:00
Brennan BensonandMerge Sim bcb703fb4c test(ssh): isolate the MFA fixture from the developer's real ~/.ssh (#19300)
The multi-stage cases pass `resolved: null`, so `resolvePrivateKeys` falls
through to `findDefaultKeyFile`, which reads `~/.ssh/id_*` via `homedir()`.
On a machine with an encrypted default key ssh2 rejects with "Cannot parse
privateKey" before authentication is exercised, so two cases failed locally
while staying green on hosted CI, which has no key. Point home at the existing
fixture directory so default-key discovery stays in the test's control.

Co-authored-by: Merge Sim <sim@local>
2026-09-07 09:32:14 -07:00
a899f92402 feat(windows): enable structured Codex chat on native Windows (#18519)
* feat(native-chat): enable Windows structured sessions

* fix(codex): prove native Windows process identity

* style(codex): format Windows session seam

* fix Windows structured Codex admission

* fix(windows): reprobe missing process identity capability

* fix(windows): decide folder-workspace WSL routing before the click

Review found pathUsesWslUnc exported but unused, and the folder composer
hardcoding worktreeUsesWslPath:false. Together those meant a folder picked
under a \\wsl.localhost\ parent routed to structured chat, then got refused
by the host and fell back AFTER the click -- which defeats the lane's own
design goal that create cannot fail after the click.

The group's parentPath is in scope at submit and the workspace is created
under it, so the parent decides WSL-ness pre-click. Wires pathUsesWslUnc
there and adds tests for the helper, including the unhydrated-store case
that previously threw.

* fix(windows): collapse the gate derivation to one call, restoring max-lines

CI static analysis failed: launch-agent-in-new-tab.ts crossed the 300-line
oxlint ceiling. Adding a max-lines disable is forbidden, so the two gate
derivations collapse into one readWindowsStructuredGateInputs() call --
a store-backed site now adds one line and one import name instead of two.
Better shape anyway: one derivation entry point rather than two reads a
call site must remember to pair.

* fix(windows): engage the legacy fallback when the host THROWS a refusal

Review found a P1 this merge composes: neither parent could reach it. At the
lane head the only structured entry was launch-agent-in-new-tab (full
store-backed WSL check); on main all win32 was refused. The merge enables
win32 in creation flows that pass no projectRuntime, so a WSL folder
workspace, a WSL-configured repo, or a repair-required runtime now routes
structured -- and the host refuses correctly, but by THROWING rather than
returning {ok:false, refusal}.

Callers engage their legacy-terminal fallback on the refusal CLASS, so an
unmapped throw arrives as a generic RPC rejection: no fallback, empty
workspace, error toast, prompt stranded in the launch outbox. Pre-merge the
same action opened a legacy terminal agent.

Map the host's thrown definitive refusals onto the refusal class at the
launch boundary, so every creation flow -- present and future -- degrades to
the legacy terminal instead of stranding. Narrow predicate: unrelated
failures (ECONNRESET, empty message, non-Error) still propagate untouched.

Ablation-proven: removing the mapping reddens the fallback test.

* fix(windows): teach the mobile RPC double the status probe the lane added

CI's first-ever run on this lane caught a pre-existing lane defect. The lane
changed status.get to resolve through
runtime.getStatusAfterWindowsProcessStartTimeProbe(), but never taught the
mobile-surface runtime double about it, so status.get failed for mobile
clients with "not a function". The lane's own test list did not include this
file and the lane had zero CI, so nothing ever ran it.

The real runtime always implements the method; the double omitted it.

* chore: merge current main and regenerate the localization runtime catalog

CI static analysis failed on a stale en-runtime-required.json: main added
onboarding integration-capability keys, and the generated catalog is checked
against the PR MERGE result, not the branch alone -- so it read clean locally
while failing in CI. Merging current main (90780acb85) and regenerating.

Gates after the merge: pnpm tc 0, oxlint 0, changed-code quality 0/56,
7 gate/lane test files 69 tests green.

* fix: route structured launches by execution host platform

* fix: recover paired structured session mirror on host swap

* Revert "fix: recover paired structured session mirror on host swap"

This reverts commit 81bfca0007.

* Revert "fix: route structured launches by execution host platform"

This reverts commit 47abbd354a.

* fix(windows): refuse structured chat in a paired web client

Reverts the two review-loop commits (restoring a tree byte-identical to the
validated head) and closes the hole they were aiming at, without their cost.

A paired web client's `platform` describes the browser's machine, not the host
that will run the agent, so the Windows gate cannot be evaluated there. Before
this, a browser on macOS driving a Windows runtime read "not win32", skipped the
creation-time proof entirely and allowed structured chat — fail-OPEN, the
dangerous direction, bypassing the guarantee this lane is built on.

`isWebClient` is a required input like the other gate fields, so the compiler
enumerated all seven call sites. Refusal is synchronous and fail-closed: no
async round-trip, no null window, no cache to invalidate — unlike keying on an
asynchronously-fetched host platform, which would have made every desktop
launch wait on a round-trip to fix a paired-web-only hole.

Paired web therefore gets the legacy chat until the host publishes eligibility
itself; that is the proper fix and belongs in its own PR.

Ablation-proven: removing the guard reddens both refusal tests; the
desktop-unaffected test is a preservation check and passes either way.
Gates: tc 0, oxlint 0.

Known open: repos-onboarding-folder-startup.test.ts fails on this branch and
passes on plain main — under investigation, NOT caused by this commit.

* test(onboarding): mock the web-client check the store path now reaches

The web-client refusal added `isWebClientLocation()` to the launch-route
inputs, which this suite's store path reaches while adding the FIRST folder.
The suite stubs `window` as `{ api }` with no `location`, so the function
cleared its `typeof window === 'undefined'` guard and then threw on
`window.location.pathname`.

That threw inside addNonGitFolder's own catch, so folder-1 never activated;
folder-2 then returned early (a project already existed) before reaching the
call at all, leaving exactly one activation with no startup seed.

Test artifact, not a product defect: a real renderer always has
`window.location`, so the seeding path is intact for users. Mocking the module
is the convention 7 other suites already use, and keeps product code free of
defensive branches that only exist to satisfy a stub.

Ablation-proven: removing the mock reproduces the original failure exactly.

* fix(renderer): make the web-client check total over a partial window

isWebClientLocation() guarded `typeof window === 'undefined'` and then assumed
`window.location` existed. A window stubbed without a location cleared the
guard and threw on `.pathname`.

That matters because this branch put the call on the launch-routing path,
where the throw is swallowed by the caller's catch and silently becomes a
FAILED LAUNCH rather than a visible error. CI caught it as 9 failures in
launch-work-item-direct.test.ts.

I previously "fixed" this by mocking the module in the one suite I knew about.
That was whack-a-mole against an unbounded set, and it missed this one. The
defect is the partial-window assumption, so fix it there: the mock is removed
from the onboarding suite and both suites now pass on the hardening alone.

Ablation-proven: reverting to the unguarded form reddens 11 tests across the
new unit suite and launch-work-item-direct.

Gates: tc 0, oxlint 0, changed-code quality 0/58.

* Move Codex's Windows structured-chat eligibility onto the host createSupport probe

The renderer no longer decides Codex win32 eligibility: launchStructuredAgentSession
probes agentSession.createSupport for both providers, the host answers via
supportsCodexStructuredLocation (process start-time proof + WSL refusal), and the
create path re-checks live. Deletes the client-side windows gate module and its
routing inputs (windowsProcessStartTime, worktreeUsesWslPath, isWebClient, platform)
from six call sites. Splits killCodexAppServerProcessTree out of
codex-app-server-session to hold the max-lines ceiling without a disable.

* fix(ci): keep pnpm lockfile stable

* test(windows): align foreground snapshot flags

* Restore main's pane-snapshot flag contract

Main asks for CreationTime on both projections; this branch's hot-path
isolation went away with the async probe it served.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
Co-authored-by: Merge Sim <sim@local>
Co-authored-by: Merge Sim <merge@localhost>
2026-09-07 09:18:38 -07:00
Jinwoo Hong 8cd0abf76a test(relay): prove the capability header reaches acceptControl over a real upgrade (#19274)
The unit tests cover parseRelayHostCapabilities, the sendHelloAck gating, and the
header literal separately, but nothing joined them: a typo in the header name
read off the upgrade request passed the entire suite. This drives a real control
upgrade carrying the header, leaves an invite connection pending, and asserts the
rebound control's ack. Renaming the header the server reads fails it.
2026-09-07 07:39:38 -04:00
Neil 9d29e6878e fix(codex): distinguish personal and enterprise accounts sharing an email (#19279)
* fix(codex): distinguish same-email accounts in the switcher

* fix(codex): scope switcher disambiguation to the visible runtime group

Review follow-ups: wrap labels at word boundaries instead of mid-word,
disambiguate against the accounts a group actually renders, and tolerate a
missing email arriving from persisted settings or a remote summary.
2026-09-07 03:28:33 -07:00
Jinwoo Hong 1bf30670d4 fix(relay-ops): let the rehome trust probe approve the asia-east2 cells (#19275) 2026-09-07 05:48:19 -04:00
Jinwoo Hong 9c8f4c398c fix(relay): bound control RTT samples per ping and per flush window (#19268)
* fix(relay): bound control RTT samples per ping and per flush window

An authenticated host chose how many round-trip samples a cell recorded: every
pong carrying a recent plausible `t` was forwarded to the process-wide window,
which grew unbounded until the 30s flush copied and sorted it for percentiles.

Time a pong only when it echoes the `t` of the ping still outstanding on that
session, so a flood yields at most one sample per ping the cell actually sent.
A pong that lost the race to the next ping is dropped for timing but still
counts as proof of life for the silence watchdog. Bound the process-wide window
with a 1024-sample reservoir (Algorithm R) so the percentiles stay unbiased,
keep `controlRttSamplesDelta` meaning round trips observed, and publish
`controlRttSamplesDroppedDelta` for the ones the reservoir did not keep.

Replace the leak guard's blanket `"credential":` string rewrite with an exact,
path-scoped rename of the two schema keys that spell a policed word, and make
the guard case-insensitive now that nothing legitimate trips it.

Follow-up to #19232.

* test(relay): prove the RTT reservoir samples the whole window
2026-09-07 05:06:42 -04:00
Jinwoo Hong 91d7783f2b fix(relay): state pending-conn details to hosts that advertise the capability (cell side) (#19266)
The cell announces a connection with a single conn-open. When the desktop's
control socket dies mid-accept the phone waited out the 10s attach deadline and
was closed HOST_OFFLINE, even though the desktop was online. host-hello-ack
already restates those connections in pendingConns, but only by connId and
connTicket, which is not enough for the desktop to dial: kind and relayDeviceId
decide the pairing authority a connection carries and the E2EE device binding,
so neither may be guessed.

The cell now states kind and relayDeviceId on each pending entry, but only to a
host that advertised it can read them: a shipped host parses those entries
strictly, so an unannounced key fails the whole ack parse and kills a working
control. The advertisement rides the control upgrade as
x-orca-host-capabilities, not host-hello, because HostHelloSchema is strict on
the cell too and any new hello key is refused by every already-deployed cell.

The capability is keyed by socket, not by session: a rebind can land a successor
whose decoder is older or newer than the one that opened the session, and the
ack must follow the socket that will actually read it.

With no capable host in the fleet the emitted ack is byte-identical to today's.
The desktop half that consumes the new fields is #19238.
2026-09-07 05:06:38 -04:00
Jinwoo Hong db13cff832 relay: give the asia-east2 cells the regional rehome identity (#19239)
`relay_region_rehome_source_cell_ids` listed only the 16 US cells, and that
list is the sole thing that stamps ORCA_RELAY_REHOME_DIRECTOR_SERVICE_ACCOUNT
and ORCA_RELAY_REHOME_AUDIENCE into a cell's startup script. A cell reports
regionalRehomeProtocol 1 only when both are present, so c27-c29 have always
reported 0. That leaves them ineligible as rehome sources and, once the worker
is bidirectional, as targets too, which strands the US desktops homed there.

This is a prerequisite only. Merge and roll it ONLY AFTER the bidirectional
rehome director change is deployed. Two live gates still hard-code the primary
region and would reject an Asia source no matter what the template stamps:
`cloud/apps/relay/src/app.ts` line 610 fails the trust probe with 409 when the
source cell's region is not RELAY_DEFAULT_REGION, and
`cloud/apps/relay/src/assignment-store.ts` line 5476 skips such a cell as
source_ineligible during rehome source selection. The bidirectional lane
removes both.

The topology check asserted every source sits in the primary region. That
mirrored those two gates rather than protecting anything Terraform owns, so it
is now advisory: it requires only a configured, unfenced cell with an explicit
connection limit, and the comment records that region eligibility belongs to
the director's own source and target predicates. Every cell's region is
already constrained by the assert above it.

The same-cap census test cross-checked membership against us-central1. Every
reviewed serving cell now carries the trust, so it asserts protocol 1 for all,
plus one non-source cell to keep the validator's protocol-0 branch covered.

Roll sequencing, because this apply is not self-contained:

- After the apply the Asia templates carry the two rehome lines, and the
  `unexpectedRehome` rule at `cloud/dev/scripts/validate-relay-capacity-plan.mjs`
  lines 243-247 rejects a protocol-0 plan that contains them. So c27-c29 have
  no dispatchable protocol-0 same-cap roll until the director gate is gone or
  this is reverted.
- The same-cap job runs the per-host trust probe after isolate, drain, and the
  targeted apply. A 409 there leaves the cell serving but isolated and
  migration-only, which is what happened to c13 on 2026-09-06.
- The only safe path: deploy the bidirectional rehome director, then dispatch
  `Deploy Relay Production Same-Cap` canary-apply for one Asia cell with
  target-rehome-protocol 1 and rollback-rehome-protocol 0, then batch-apply the
  remaining two. That job runs its own targeted template and MIG apply.
- Never reach these cells with an untargeted root apply. The current plan
  carries 60 changes and 50 destroys of unrelated standing drift.
2026-09-07 05:03:26 -04:00
Jinwoo Hong d74f8cb787 revert(mobile): hold the relay reconnect path and cache-first reconnect for a separate mobile pass (#19265)
* Revert "feat(mobile): draw the last known tab strip while a session reconnects (#19258)"

This reverts commit 0ba7f8dc8d.

* Revert "perf(mobile): cut the relay reconnect critical path and admit dead sockets faster (#19236)"

This reverts commit 23df74d85a.
2026-09-07 04:56:13 -04:00
Jinwoo Hong fede3eb2ff fix(test): give federation tests a real read-after-write sync barrier (#19262)
`syncOrchestrationFederation()` coalesces onto an already-in-flight relay-tick
sync, which may have pulled from the peer before the caller's mutation existed.
Tests used it as a barrier, so `keeps a timed-out remote question resumable`
could reply against a home DB that had never imported the worker's question:
the reply failed with `Message not found`, no `to_worker` relay was enqueued,
and the resume ask surfaced it 5s later as a spurious timeout.

Add `syncFederationBarrier()`, which chains each active dispatch past the
current round via `syncOrchestrationFederatedDispatchAfterCurrent`, and use it
at every barrier-purpose sync site. The two tests whose subject is the sync
machinery itself keep the raw call. Also assert the reply response, so a failed
reply fails at the reply instead of masquerading as a timeout.

Production is unaffected: `syncOrchestrationFederation` has no production
callers, real read-after-write paths already use the after-current sync, and
relay ticks retry every second.
2026-09-07 04:54:28 -04:00
Jinwoo Hong e068947d4c feat(relay): alert on far-cell placement and skewed region hints (#19253)
* feat(relay): alert on far-cell placement and skewed region hints

US desktops were homed on asia-east2 cells for weeks in 2026-08 with every
existing relay alert green. Roughly 226 of 332 hosts on those cells were
non-APAC, and a phone connect took ~10 s there against ~0.6 s in region, but
nothing in Cloud Monitoring could see distance: the connection, queue, heap,
and SQL bars all measure a cell's own health, which was fine.

Three policies close that gap. Two read distance per cell, from the accept
and control-RTT timing added in the parent commit: phone-accept p95 above
2 s, and control ping p50 above 150 ms. The third reads the cause fleet-wide,
as the asia-east2 share of the region hints desktops send the director, so a
mis-picking client probe is visible before it lands anyone on a far cell.

All three are MQL rather than the metric filters the other relay policies
use. Every runtime metric is a DELTA DISTRIBUTION, and a filter condition can
only align one with a percentile; each alert needs the sum of the extracted
values as a volume floor so a sparse window cannot page. None of these
metrics exists in the project yet, so what was checked against production is
the query shape: the same MQL run over existing metrics of the same kind.

The skew denominator needs one log-based metric per hint key, so
`requestedRegionsDelta` now has one per relay region plus the unhinted
bucket. Those ride the existing snapshot metric family, which adds map
entries without touching the live metrics. A ratchet test pins the key list
to relay-contract's RELAY_REGIONS: a region added there without a metric
would shrink the denominator, so the test fails rather than letting the share
quietly inflate.

* fix(relay): compare hinted regions against placed ones, not a fixed share

Review found the skew alert inverted at both ends. A fixed 40% bar on the
asia-east2 share of region hints was silent through the exact broken state it
was written for, and would page forever once the desktop probe is fixed and
the genuine APAC share rises past it. An absolute share cannot separate those
because it has no reference point.

The hint share now has one: the share of assignments the director actually
placed in that region during the same hour. Measured over twelve hours on
2026-09-07, while the probe was still mis-picking, asia-east2 was 33.8% of
33,800 hinted requests and 7.9% of 45,364 assignments. That is a 4.27x
divergence and a 25.9-point gap, so the alert fires above 2x and 15 points,
inside the broken state and outside a healthy one. Both bars must hold: the
ratio alone blows up on tiny placement counts, the gap alone misses a
proportionally large skew at low volume. The reviewer proposed either bar
alone; requiring both keeps each one meaningful and still clears today's
numbers with room.

`unhinted` requests leave the denominator. They were 27% of all requests, so
a client that always sends a hint would move the number from 21.9% to 35.0%
with no behaviour change at all.

The comparison needs per-region placement counters, so `selectedRegionsDelta`
gets log-based metrics alongside the requested ones. Rather than extract four
hyphenated map keys through quoted field paths, which nothing in the project
does and which cannot be checked without applying, the relay now also
publishes flat `requestedRegion<Region>Delta` and `selectedRegion<Region>Delta`
fields next to the untouched maps. They are emitted as zeros in every
interval, so no series can drop out of the alert's inner join in an hour with
no asia placements, which is exactly the hour the skew is worst. Additive
only: metricVersion is unchanged, the maps still carry anything outside the
catalog, and the emitter's leak guard still passes.

Two corrections to what the previous commit claimed. None of these metrics
exist in the project yet, so the code, the doc and this message now say what
was actually checked against production: the query shapes, run over existing
metrics of the same kind. And the control-RTT policy records that EU desktops
on us-central1 sit at 100-130 ms, so a European-heavy cell can approach the
150 ms bar while correctly homed.

The skew alert will stay lit after a client fix until the backlog is rehomed.
Sticky assignment never re-consults the hint, so a desktop already on an asia
cell keeps landing there whatever it now asks for. The policy description and
the doc both say so, so nobody reads a slow clear as a failed fix.

* fix(relay): cross-multiply the skew bars so a zero placement share still fires

`hint_share / placement_share` is undefined in the hour that matters most.
When the director placed nobody in the region, MQL returns no rows for either
0/0 or x/0, so the series disappears before the gap and volume clauses run and
the alert stays silent. That hour is not hypothetical: it is every desktop
asking for a region while the director puts nobody there, which is what a
drained, fenced, or full region looks like, and it is the most extreme skew
the alert can see.

The condition is now cross-multiplied, `hint_share > 2 * placement_share`,
which is well defined at zero. Both forms were run read-only against
production surrogates chosen so the placement denominator is exactly zero:
the ratio form returned no rows, the cross-multiplied form returned the series
with the condition true on every point. A second surrogate pass with a tiny
hint share returned the series with the condition false, so the gap clause
still suppresses the healthy shape rather than the query silently matching
everything.

The flat field names are no longer derived on either side. Terraform title
cased each dash-separated part and the emitter upper cased each part's first
character, so the ratchet had to pin two source expressions by regex, which a
reformat would break and which never compared the actual rendered names. Both
sides now declare a literal map, relay-contract's
RELAY_REGION_METRIC_SEGMENTS and Terraform's relay_region_field_segments, and
the test compares the two declarations against each other and against the
expected names. `satisfies Record<RelayRegion, string>` makes a region added
without a segment a compile error rather than a silent gap in the alert's
denominators.

Both ratchets were checked by mutation: a wrong Terraform segment, a contract
region with no Terraform entry, and a revert to the ratio form each fail the
node test, and the new region fails the contract build.
2026-09-07 04:51:48 -04:00
Jinwoo Hong 0ba7f8dc8d feat(mobile): draw the last known tab strip while a session reconnects (#19258)
* feat(mobile): draw the last known tab strip while a session reconnects

Reopening a workspace the phone has already visited threw away everything
it knew. The route clears its tabs on mount, so until the reconnect lands
and the first snapshot is applied the session screen has an empty header
and a bare spinner, even though the strip it is about to be handed is the
one it drew a minute ago.

Persist the four fields the strip actually draws -- id, type, title, agent
-- per host and workspace, and add a reconnecting-with-cache shape to the
route state so those rows render immediately, disabled, under the ids the
live snapshot will reuse. Live tabs always outrank the cache, so a
mid-session drop keeps its mounted terminals; an exhausted retry loop or a
rejected pairing outranks it the other way, because a strip the user cannot
reach is worse than the existing offline affordance. With nothing cached
the screen behaves exactly as before.

The body stays a placeholder. Replaying stored scrollback into the terminal
WebView would double-render the same rows once the live stream replays them,
so the strip is the cached content and the body waits for the stream.

* fix(mobile): keep shell titles and unpaired hosts out of the cached tab strip

Review of the reconnect strip cache found two ways it leaked.

A terminal's title is whatever the shell last set, which is routinely the
command line: a psql URL with an inline password, a curl with a bearer
token. Both fit well inside the 64-character cap and both were written to
plaintext AsyncStorage verbatim. Browser tabs carried their page title the
same way. Terminals and browsers now collapse to a fixed label, with a
resolved agent naming itself because that lookup is a closed enum. The rule
lives in the storage module rather than its caller, so it holds for entries
an older build already wrote, and a tab type this build cannot draw is
dropped instead of having its title trusted.

The cache also survived forgetting a host. Nothing expired an entry, and
the module-global memory map meant a later save from any surviving host
serialized the forgotten host's rows straight back to disk. Both cleanup
paths now evict by host, dropping the in-memory rows and rewriting storage,
with a pending debounced write cancelled so it cannot restore them.

Also: the storage key digests the workspace id, which ended in a filesystem
path, and cached rows carry the same de-emphasis as the disabled tab-bar
buttons beside them, so an inert row does not pass for a live one.
2026-09-07 04:42:59 -04:00
Jinwoo Hong 23df74d85a perf(mobile): cut the relay reconnect critical path and admit dead sockets faster (#19236)
* perf(mobile): cut the relay reconnect critical path and admit dead sockets faster

Phone medians put E2EE authentication at ~424ms but `connected` at ~630ms,
because the session serialized two RPC round trips behind it: the resume
confirm (`pairing.getEndpoints`) and the capability advisory. Both now ride
the authenticated socket concurrently and off the critical path, so the
session publishes `connected` as soon as E2EE authenticates. Peer identity
is already proven by then — the confirm carries credential/lease bookkeeping
and the cell assignment check, and it still fails the session on a bad answer
or a foreign relayHostId, only later. `persistResumeConfirmation` awaits the
new `whenResumeConfirmed()` instead of assuming the answer is present at
`connected`.

Foreground liveness on a retained relay: `notifyForeground('app-resume')`
now probes past the 10s voluntary minimum on urgent bounds (2s, one miss),
so a socket that died while the process was suspended is admitted in ~2s
instead of ~8s. Focus and network nudges keep the old minimum and bounds.
Relay sessions also gain a 25s idle sweep, gated on foreground so a
backgrounded app spends no probes.

Recovery is no longer blocked by the direct return probe. The probe's 12s
dial is a pure observation on its own socket, so it takes the supervisor's
operation mutex only for the cutover; a relay recovery landing during a
foreground return now starts immediately instead of waiting the budget out.
Requests that do land during the cutover are queued in a new
RelayRecoveryIntentQueue and replayed on release — an owning forced
replacement keeps its intent, everything else replays as a plain recovery.

Tests updated deliberately, for the new ordering:
- 'sends no periodic traffic while an authenticated relay is idle' asserted
  the absence of any relay idle probe, which is exactly the gap D3 closes.
  Replaced by a sweep test plus a backgrounded no-probe test.
- 'rate-limits foreground sequences without suppressing a retry' asserted
  that app-resume was suppressed inside the 10s minimum. An app resume is
  now the one nudge that must never be rate-limited.
- the session helpers waited for the confirm answer before `connected`;
  they now authenticate, read both concurrent frames, and settle them.

* fix(mobile): book backoff when a relay resume confirm fails after the cutover

Review round 1 on 352bfd2300.

P1: publishing `connected` at E2EE authentication made `migrateTo` resolve
before the resume confirm answered, so a confirm that failed afterwards —
a `relayHostId` mismatch from a rehomed desktop is the live case — was still
reported as an `established` dial. registerFailure was skipped, no cooldown
was booked, recordMigration()/setActiveSession() ran for a dying session, and
the queued-recovery replay redialled immediately: a tight loop with a
connected→disconnected blip per pass. The establisher now awaits
whenResumeConfirmed() after the cutover and, if the session is no longer
connected, reports a failed dial (or an aborted one when direct won or the
supervisor went inactive) exactly as a rejected migrateTo used to. The UI
still connects early; only the supervisor's bookkeeping waits.

The state check, rather than getFailure(), is the oracle: a live session can
carry a latched failure without having failed yet, and "is this session still
alive once the confirm settled" is precisely the question migrateTo used to
answer.

P2: the resume probe profile goes to two 2s misses instead of one. The first
frame after a resume rides a cold radio and a possibly distant cell, so one
slow answer is not proof of a dead link; the verdict still lands at 4s rather
than the previous 8s.

Nits: the direct probe's two early returns no longer close the candidate the
finally also closes (the second shape pre-existed); RelayRecoveryIntentQueue
is cleared in the supervisor's stop().

Mutex-hold note: persistResumeConfirmation, and now the establisher's own
await, are bounded by the confirm's request timeout. That would have been the
session's 30s default, so the confirm is pinned to RELAY_CONFIRM_TIMEOUT_MS
(12s) — the same bound migrateTo's waitForAuthenticated applied before.

Test: a supervisor-level case where every dial authenticates then fails the
confirm must book 250/500/1000ms backoff with no immediate redial, and must
never record a migration. It fails on the pre-fix establisher.
2026-09-07 04:40:40 -04:00
Jinwoo Hong f5be177e44 fix(relay): rehome hosts to their preferred region in either direction (#19241)
* fix(relay): rehome hosts to their preferred region in either direction

The regional-rehome worker only moved hosts from a us-central1 cell to an
asia-east2 one, so a host whose desktop later records us-central1 stays where
it was put. Rehoming now compares the fresh preference against the region of
the cell the host is on and moves it to a general cell in the preferred
region either way, through the same drain, migrate, safety, and rate-limit
machinery.

- relay_region_rehome_attempts.preferred_region accepts both regions; existing
  databases are upgraded in place by an idempotent named-constraint swap that
  is safe when several directors start at once.
- A target must carry the drain protocol too: moving a host onto a cell it
  can never be drained off again is the trap this change exists to undo. The
  fleet whose health gates a rehome is now every general drainable cell,
  which is exactly the set of legal sources and targets.
- The trust probe accepts a source cell in any region.

No wire change, and no behaviour change while the durable control is off.

* fix(relay): bound bidirectional rehoming with a per-host cooldown

Moving hosts in both directions removed the property that made the old
one-way worker self-terminating: a desktop whose region probe flips would be
dragged back and forth, one full drain and migrate per flip, because the
preference age never expires while the host keeps reconnecting.

- relay_region_rehome_control gains host_cooldown_ms, an operator input
  plumbed like preference_max_age_ms (workflow, ops script, admin route,
  durable row) and defaulted to seven days. A host with any attempt row
  inside the window, whichever way that move went, is not a candidate; the
  claim re-reads it under lock so an attempt landing between scan and claim
  cannot start a second move. Skips are named host_cooldown, and the lookup
  rides a new index on (user_id, relay_host_id, created_at).
- The candidate scan now also requires the target cell to be enabled, so it
  mirrors the claim-time filter exactly and stops spending batch slots on
  candidates that are certain to be skipped.
- Region CHECK lists are rendered from the shared region list instead of
  being written out four times.
- The operations runbook states that cells without the drain protocol are
  neither sources, targets, nor members of the safety gate.

* fix(relay): keep rehome reads and brakes working across the cooldown rollout

The ops script validated hostCooldownMs on every inspected control, so
against any director image predating the field inspect, pause, disable, and
failed-enable recovery all threw client-side. The workflow always runs from
main while the director image is operator-supplied, so that window opened at
merge and reopened on every rollback: the operator lost read-only visibility
and both emergency brakes while the worker could still be enabled.

The field is now validated only when the director reports it, and every apply
body that echoes an inspected control omits the key when that control lacks
it, so a legacy director never sees an unknown key. The write path stays
fail-closed the other way: enable refuses up front, before any mutation, when
the director does not report a cooldown it could honour.

Also replaces two bare 'us-central1' defaults with RELAY_DEFAULT_REGION.
2026-09-07 04:40:37 -04:00
Jinwoo Hong ecfcc0d833 feat(relay): time successful client accepts and control round trips (#19232)
* feat(relay): time successful client accepts and control round trips

A 6s accept on a cross-region cell was invisible: only the abandoned path
was timed. Record per-stage durations across acceptClient and acceptHostData
(assignment/credential/activity/attach), emit one completed log line per
accept, and aggregate p50/p95/max into the runtime metrics event.

Sample control ping round trips from the pong echo so a host sitting on a
distant cell is visible fleet-wide and per host, rate-limited to one log
line an hour per session.

* fix(relay): review round 1 on accept and control-RTT timing

Omit the accept and RTT percentiles from windows with no samples: accepts
are sparse, so a zero point every 30s would pin the p50 at 0 and collapse
the p95. The *Delta counts still publish, and say when the omission is
expected. Control-renewal output is unchanged.

Add a `basis` stage for the splice lease and connection-basis writes that
run between the host data leg and relay-hello, and start `attach` where the
activity stage ended, so the stages now tile the whole accept and their sum
equals totalMs. Clamp every stage at zero against a backwards clock step.

Carry role/cellId/region on both new log lines, flatten the stage p95 field
names so the log-metric extractors stay top-level, and record that only the
RTT median reads as distance: the desktop echoes the pong on its main
thread, so the p95 and max track desktop stalls.
2026-09-07 04:40:34 -04:00
Jinwoo Hong a3e67365a3 fix(orchestration): recover Codex idle after completion title race (#19243)
* fix(orchestration): recover Codex idle after completion title race

* test(native-chat): enable structured sessions in adoption replay fixture

* test(orchestration): cover deferred pointer recovery after prolonged unknown status

* fix(orchestration): fence completion recovery by process generation
2026-09-07 04:35:56 -04:00
Neil c3a70082c6 Fix MiniMax credential-expiry reporting, region sync, and refresh (#19250)
* Fix MiniMax credential-expiry reporting, region sync, and refresh

Three defects from #14929:

1. The usage endpoint answers an expired cookie or key with HTTP 200 and
   base_resp.status_code 1004, never 401/403 (confirmed against both regional
   hosts). The stale-token branch was therefore unreachable, so expired
   credentials surfaced as 'usage-unavailable' with the raw upstream string,
   and stale policy kept showing old numbers as if the failure were transient.
   Classify 1004 as an expired credential.

2. minimaxEndpoint reached the SettingsUpdate schema and the web store but was
   never projected by RuntimeClientSettingsController.get(), so a paired client
   fell back to 'overseas' regardless of the host's region and rendered the
   wrong console link. Add it to the projection and the store contract.

3. Changing the region persisted without refreshing usage, leaving the previous
   host's snapshot in the status bar until the next poll. Invalidate and refetch
   when the endpoint, group id, or model list changes.

The RPC-level tests mock the controller, so the projection had no real
coverage; the new test fails against the pre-fix projection.

* Localize the MiniMax credential-expiry copy

Classifying 1004 as stale-token made the status bar show the raw English
error verbatim: the new wording matches none of USAGE_AUTH_ERROR_PATTERNS,
whereas the old upstream text ('...log in again') matched and was replaced
with localized copy. That traded a localized-but-misleading message for an
actionable English-only one, which is the wrong trade for the CN users this
work targets.

Tag the error with credentialSource so the renderer can pick the right
localized string per credential kind, and add the three catalog entries.
2026-09-07 01:22:08 -07:00
Neil ffff6eaca2 fix(test): admit the adoption-replay create fixture through the structured gate (#19246)
Semantic conflict between two green PRs. #19176 added this replay test while
`agentSession.*` still admitted a `runtime` client on its negotiated capability
alone; #18700 then made `experimentalStructuredNativeChat` one rule for every
caller. Neither branch saw the other, and main runs no post-merge test gate, so
`agentSession.create` started refusing at the envelope level and the test's
`ok: true` expectation broke.

#18700's rule is the intended behaviour and `create` starts work, so it belongs
behind the gate. The fixture is what is stale: it builds a real
`OrcaRuntimeService` whose client settings are unset. Enable the setting the way
#18700 already did for the sibling pre-commit fixture. The assertions about
durable-identity replay are untouched and now actually run.
2026-09-07 01:01:49 -07:00
Neil 314506003a fix: retain MSYS shell descendants in their terminal job (#19068)
* fix: retain MSYS shell descendants in their terminal job

* test: complete MSYS regression CI registration and teardown contract

* fix(windows): deny job breakaway for the whole Cygwin/MSYS shell family

The per-PTY job probed only msys-2.0.dll, and only for bash.exe/sh.exe.
Cygwin ships the same spawn.cc breakaway logic under cygwin1.dll, and an
MSYS2 zsh escapes exactly like its bash does, so both kept the orphan bug.

Probe the runtime DLL on the shell's own search path instead of matching
shell names: that is the property that decides whether the runtime will
ask for CREATE_BREAKAWAY_FROM_JOB, and it drops the name special-casing.

* chore(patch): restore the conpty.cc index line

The earlier hand-edit dropped it while every sibling section kept one.
Recomputed against the real blobs: applying this patch to 7b286d3d
yields exactly 4b06d185, so git apply -3 has its fallback back.
2026-09-07 00:35:55 -07:00
Neil 374c676f6d fix: repaint hidden output overflow after answered restore deadline (#18904) 2026-09-07 00:25:16 -07:00
Neil 3f4793b6c9 Reorganize MiniMax modules and de-duplicate shared test state (#19197)
* Move MiniMax quota fetch modules into rate-limits/minimax

The five MiniMax fetch/transport modules sat flat among ~110 files covering
eight providers. Nest them so the provider's fetch surface is one directory;
credential stores (main/minimax) and the IPC handler (main/ipc) stay where
their siblings are.

* Build rate-limit and settings test state from shared factories

RateLimitState was hand-copied in 9 places and the full GlobalSettings object
in 2 more, so adding one provider field forced edits in unrelated providers'
files -- which is how MiniMax fields ended up in codex-accounts and the Grok
usage-pane test.

Add createEmptyRateLimitState and createGlobalSettingsFixture and route the
copies through them. Values that deviated from the defaults are passed as
explicit overrides, so the fixtures produce what they produced before.

rate-limit-types.test.ts keeps its literal (it exists to assert the shape) and
service-state.ts keeps its own (InternalRateLimitState is a subset, not the
same type).

* Share the codex-account settings fixture between both harnesses

The two codex-account fixtures still carried the same 30-line override block
verbatim, which is the duplication the shared fixture was meant to remove.
Move it into one createCodexAccountSettings and have both call it.

Also drop the hardcoded POSIX workspaceDir default; callers supply the real
directory and a '/tmp' literal would be a trap on Windows.
2026-09-07 00:12:32 -07:00
Jinwoo Hong a62cfedad8 Resolve push source archive from repository root (#19231) 2026-09-07 03:04:19 -04:00
Jinwoo Hong e4770d712f Restore independent push gateway deployment (#19225)
* Restore isolated push gateway deployment workflow

* Register push deployment in the shared SQL lease census

* Restore push workflow inventory and identity contracts
2026-09-07 02:58:32 -04:00
Neil 2ccf35b135 fix: avoid quadratic trimming during fullscreen terminal redraws (#19214) 2026-09-06 23:41:49 -07:00
Brennan BensonandMerge Sim fa5ef99885 fix(native-chat): settle structured chat turns stranded by a restart (#19122)
* fix: settle structured chat turns after restart

* fix: preserve unconfirmed turn cancellation state

* test: preserve unconfirmed turn lifecycle

* test: narrow unconfirmed cancellation coverage

* fix: keep intentional TUI closes out of recovery

* test: keep branch rename journal mock current

* fix: settle dead TUI handoffs before reacquire

* fix: preserve handoff stage after retry settlement

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 23:34:50 -07:00
Brennan BensonandMerge Sim ba4e79c250 fix(runtime): apply the structured-chat setting to every RPC caller (#18700)
* fix(runtime): apply the structured-chat setting to every RPC caller

supportsStructuredAgentSessions only consulted experimentalStructuredNativeChat
when clientKind === 'mobile', so identical host settings admitted desktop and
in-process callers while refusing a phone. The server branched on client surface.

The setting is now one rule for every caller. The negotiated capability stays a
wire term asked of remote clients only, so a capability-less in-process caller is
still admitted on the setting alone.

Making the projection's structuredNativeChatEnabled argument required surfaced
eight call sites that passed `undefined` for non-mobile clients; they now read the
host setting, so tab projection follows the same single rule.

Announced behaviour change: with the flag off, session.tabs.list/listAll no longer
restore structured tabs for desktop. The desktop renderer already discards them in
that state, and startup record/lease reconciliation is unaffected.

* fix(runtime): keep structured session cleanup available

* test(runtime): enable structured chat in desktop projection fixture

* test(agent-session): settle merged fixtures against the all-clients structured policy

The merge with main left three fixtures written for the old mobile-only rule:
a duplicate getClientSettings key, a create fixture with no host settings at
all, and a projection call whose 'old client' is now the mobile fallback-title
case.

* fix(native-chat): let an admitted caller close a chat after the setting is off

Turning `experimentalStructuredNativeChat` off revoked admission for every
`agentSession.*` method, including `close`. A chat opened while the setting was
on stays mounted, so its owner was left with a live provider child and an X
button that answered `structured_agent_session_unsupported`.

Split the surface by what a method does to work in flight rather than by how it
sounds, and write that rule where the gate lives so the next method lands on the
right side: starting, extending, retaining or reading needs admission; stopping
or retiring work the caller already owns does not. Moves `close` and `cancel`
onto the cleanup gate alongside `unsubscribe` and `release`.

The tightening is unchanged - the cleanup gate still demands the negotiated wire
capability and never creates a host, so an incapable client still cannot see the
surface and no method that starts work is reachable with the setting off.

Extracts the dispatcher harness and the method-to-gate table into fixtures so
the new admission suite can share them without a max-lines disable.

* Drop a duplicate lastActivityAt key carried in from main

The main commit this branch merged (fb322046e8) had two lastActivityAt
properties in the same object literal at both journal stubs, which fails
TS1117 and oxlint. Upstream has since kept only the later value; match it.

Not introduced here, but merged in, so it has to be fixed here.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 23:32:00 -07:00
c300913f90 fix(mobile): stop double-scaling commit timestamps in history rows (#17731)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
2026-09-06 23:27:13 -07:00
Brennan BensonandMerge Sim 1ae7aa8bb4 feat(native-chat): resume an Agent Session History row into a new structured chat (#19176)
* feat(native-chat): resume an Agent Session History row into a new structured chat

A Claude or Codex row in Agent Session History gains "Resume in New Chat": it
opens a new structured native-chat tab that continues that provider
conversation, with the prior turns already in the journal. Until now those rows
could only be resumed into a PTY terminal; the structured branch could reveal a
chat Orca already owned but could not adopt one it had never held.

Almost all of the machinery existed. Both lanes already resume from the record's
provider handle chain, the journal already has a transcript importer, and the
handle chain already models `adopted` as an origin. The gap was that a create
always minted an empty chain, so the adapters started a fresh conversation. This
seeds that chain.

The client names only the conversation. `agentSession.create` is reachable by
paired mobile clients, so the transcript path and the account home are derived
by the executing host and validated against the account homes it recognises —
a client-supplied path would choose which file the host imports and which
credential directory the provider child launches against.

Failure refuses rather than degrades. A transcript that cannot be found refuses
before anything is created; one that fails or decodes empty *after* the provider
has resumed fails the attach, tearing the child down and publishing no tab,
because an empty journal beside a context-carrying agent claims a continuity the
provider never gave.

Codex can resume into any workspace since it is handed the rollout path; Claude
resolves transcripts under a project key derived from the launch cwd, so it is
offered only for the workspace the conversation was recorded in.

* fix(native-chat): widen adopted-home discovery and keep ordinary launches untouched

Three corrections from review of the first commit.

The adoption's account-home candidates now include the extra Codex homes session
discovery already scans. A row this host listed could otherwise refuse to
resume, which reads as the feature being broken rather than as a scope.

Ordinary launches call `createStructuredAgentSessionLaunchIntent` with two
arguments again. Passing the resume source unconditionally appended a trailing
`undefined` that four existing call-site assertions had to absorb; the churn was
the caller's fault, not the tests'.

The transactional adoption guard's comment claimed the self-exemption is what
lets a committed create replay. It is not: replay is settled earlier by the
operation ledger, and an adoption always arrives with a null expected fence, so
a request naming an existing session id is refused a few lines below either way.
The exemption is part of what "another record" means, and the comment now says
that instead.

* fix: preserve history adoption through create and retries

* fix: replay committed history adoption from durable identity

* fix: validate history before claiming adopted sessions

* fix: extract AI vault resume domains

* fix: recognize typed history resume refusals

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 23:24:21 -07:00
bixandm4air bd242a0158 fix(editor): support Shift+wheel scrolling in combined diffs (#11756)
* fix(editor): support Shift+wheel in combined diffs

* add active modified-pane test

* fix(editor): skip shift-wheel capture when a diff pane cannot scroll sideways

Word-wrapped panes never overflow horizontally, so consuming the gesture
left it dead instead of reaching the outer combined-diff list.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-06 23:02:12 -07:00
Merge Sim 53ce684c48 fix(native-chat): classify a naming turn by why it ended, not by a null answer
A null answer carried three different facts and was read as one, wrongly in
both directions.

A model that ignored the output schema and answered in prose parsed to no
title and was marked SETTLED, so that conversation became permanently
unnameable — switching to a schema-capable model later never rescued it. And a
model that completed having said nothing is a real decline, but was marked
UNSETTLED, so it was re-asked, and paid for, on every later acquisition.

The collector now reports how the turn ended — answered, declined, failed, or
timed out — and the flow classifies on that. Declines settle; a host or model
that could not be asked properly does not. This also defuses a sub-agent reply
landing in the pre-id window and becoming the naming answer: that costs one
wasted attempt now rather than forfeiting the conversation for good.

Clearing a name also marks it attempted. A chat the user titled by hand in the
CLI and then deleted had no name and no marker, so their very next message
generated a replacement — undoing the deletion, which is the exact failure the
durable marker exists to prevent and which its own comment already promised.

Comments corrected rather than left overclaiming: the pre-id window is bounded
by the request deadline rather than instantaneous; leak protection is now exact
id matching and so relies on the app-server not retagging naming frames; the
attempted marker's write moved after the await, which widens the window where
an eviction can start a second turn; and an unreadable thread id still leaks a
persisted throwaway thread past the new cleanup path.

Two tests encoded the old classification by using prose as a stand-in for a
decline. The fake can now complete a turn having said nothing, which is the
fact they meant to exercise.
2026-09-06 22:26:40 -07:00
Merge Sim a207c17f6d fix(native-chat): stop the clear path wiping a name that is still set
An emptied `custom-title` returned `cleared` immediately, so a transcript
holding an ai-title plus a later emptied custom title wiped the durable name
and reverted the tab to its placeholder while the CLI still displayed the
ai-title. The file's own stated precedence says custom outranks ai — an EMPTY
custom slot falls back to the ai slot, which is what the code did before this
round. Only an emptied custom slot with nothing to fall back to is a clear.

It also read `customTitle ?? ''`, so an absent, null, or renamed field counted
as a deliberate clear. That is the opposite of the discipline applied on the
Codex side, where a reply shape the build cannot read is refused rather than
acted on. Wrongly clearing destroys a name the user can still see.

The only clear test stubbed the reader and exercised the reporter, so the
mechanism was bound by nothing: reverting the reader left every test green.
The new tests write real transcript records — set-then-emptied, emptied with
no fallback, and three unreadable field shapes — and the ablation now fails
four of them.

The reporter's deps are mapped explicitly beside the reporter rather than
handed the adapter's object by reference, which carried an `onError` key the
adapter type never declared and would have dropped silently the day that
object was narrowed.

Also: the sub-agent approval fixture had no `itemId`, so the registry declined
it and the request was still refused, just differently. It now carries one and
the test asserts a durable prompt reaches the user, which is the behaviour the
narrowed gate restored. And the header comment claimed the scan is not
repeated; it runs on every acquisition, deliberately, because that is how a
CLI-side rename reaches Orca at all.
2026-09-06 22:03:51 -07:00
Merge Sim c022089236 docs(native-chat): record why the naming turn does not pin a model
The turn runs on the session's own model. The app-server catalog carries no
structured signal for a small, fast one — `modelSpecialty` is null on every
entry — so selecting one would mean matching a description string or a vendor
id shape like `-mini`, and neither survives a different provider. Naming a
model the account cannot use fails the turn outright, which is worse than a
turn that works and costs slightly more.

The cost is bounded on the other axes instead: lowest reasoning effort, the
prompt capped, and a schema that caps the answer at 36 characters. It runs
once per conversation and off the send path.

Checked while deciding: every model in the catalog supports `low`, and
`turn/start` accepts an unknown effort silently rather than rejecting it, so
no capability fallback is warranted here.
2026-09-06 21:43:06 -07:00
Merge Sim f2ce4e25d4 fix(native-chat): stop naming from crashing, eating sub-agents, or self-forfeiting
Clearing a name crashed the renderer. `applyStructuredAgentSessionConversationName`
declared its input non-null and assigned straight onto `tab.title`, but the
caller passes null on a clear, and `title` is a required string every client's
snapshot builder calls `.trim()` on. That file is @ts-nocheck, so a green
typecheck proved nothing about it. It falls back to the agent's placeholder
now, which is what its own doc comment already claimed.

The naming gate ate Codex SUB-AGENT threads. It treated any thread that was
not the user's as a naming thread for as long as a naming turn was open, and
that window opens at the first turn's ack. So a sub-agent's rows were dropped
from the transcript, its reply could be taken as the naming answer, and its
approval request was auto-refused with a reason that was untrue. The broad
rule now applies only before `thread/start` returns the throwaway id; after
that it is an exact match.

The durable "we already asked" marker was written BEFORE generation, so any
host failure forfeited naming permanently — a Claude CLI without the title
request, or an app-server that rejected the throwaway thread, marked every
conversation it touched and none of them could be named after an upgrade. It
is written only on a settled answer now: a title, or a model that answered
without one. The Claude control surface reports `unsupported` separately from
`declined` so the two are distinguishable at all.

Also: Claude gained the clear path Codex already had, keyed on a positively
emptied title record rather than on absence, since the tail scan is bounded
and absence is not evidence. The blocks guard covers elements, not just the
array. Naming failures log `error.message` rather than the raw error, which an
app-server could echo the user's prompt into. Two comments corrected and two
tests retitled to stop claiming coverage they do not have.

Base branch merged first, so this is tested against what will land.
2026-09-06 21:34:54 -07:00
Jinwoo Hong d53cbed43f revert: hold mobile push feature for user testing (#19203)
Reverts 3160b54c69. Restore through a separate draft PR after user validation.
2026-09-07 00:30:21 -04:00
Brennan BensonandMerge Sim f1d8545024 feat(chat): support structured /clear and /compact commands (#19164)
* feat(chat): support structured clear and compact commands

* fix(chat): authorize mobile commands and bound clear-chain projection

* fix(chat): localize conversation command send errors

* fix(chat): retain clear pane identity with reopened history

* test: account for combined structured session RPC additions

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 21:17:00 -07:00
Merge Sim 22685044ea Merge remote-tracking branch 'origin/brennanb2025/native-chat-tab-rename' into brennanb2025/native-chat-conversation-name 2026-09-06 21:06:55 -07:00
Merge Sim d52a3bb787 fix(sidebar): stop laundering a chat's name through the pty-title sanitizer
A structured chat has no TerminalTab, so its sidebar row is synthesized from
the status entry and its name arrived as `terminalTitle` — which
getAgentRowConversationName reads as a LIVE title and runs through
conversationNameFromLiveTitle. Those heuristics exist to reject OSC titles
scraped off a pty, where a bare path or a status word is not a name. They are
wrong for a conversation name, which is a deliberate value from a known-good
source.

So a chat named `auth/login` or `src/renderer` showed in the tab strip and
vanished from the sidebar, which fell back to the last-message label:
isCwdLikeTitle nulls any single word containing a slash. `done`, `Claude` and
`Terminal 4` went the same way, and a manual rename tripped the same
predicates as a generated name.

The name now travels as its own authoritative field from the bridge to both
row synthesizers, which surface it where the resolver returns it verbatim.
Nothing about the sanitizer changes: a real scraped title still gets it.

Tested with the values that actually trip each predicate rather than a name
that would pass either way.
2026-09-06 21:06:07 -07:00
Jinwoo Hong fb322046e8 skills: rewrite and trim the seven non-orchestration guides (#19128)
* skills: rewrite the seven non-orchestration guides to one outcome-first standard

Every guide leads with Result / Done / Safe failure, states conditions instead of case lists, keeps one done bar and one autonomy envelope, and loads references at the point of use via `skills get <topic> --full`. orca-cli drops from 424 to 260 always-loaded lines with three references; orca-per-workspace-env from 794 to 397 with five.

Defects fixed in shipped guides: `emulator camera` (no such command), iOS `permissions` (backend refuses it), Android pane described as in development, `relayGracePeriodSeconds: 0` documented as immediate teardown (it is unbounded), doctor `ok: true` hiding `warn`, an SSH exemplar setting both `jumpHost` and `proxyCommand`, a provisioned-root fetch from `origin`, and the Linear unconfirmed-write rule keyed on four verbs when ten emit it.

The resolver ladder, placeholder rule, and older-binary fallback shared by every installable SKILL.md now come from one skill-stubs/_shared/cli-resolution.md fragment composed by the generator, which also bundles per-guide references into --full. New guards: every ORCA invocation and flag resolves against COMMAND_SPECS, descriptions carry no angle-bracket tokens, reference routing is checked both ways, and an always-loaded size ratchet (300 lines) that guides may leave but never join.

* skills: address review on the SSH recipe and the parity guard

- ssh-host create script: route the bootstrap ssh through the chosen jump host or proxy command, refuse both at once, use StrictHostKeyChecking=accept-new instead of a blind ssh-keyscan append, and pass gh_token/project_root/repo_url/repo_ref to the remote bash via printf %q so a quote in a value cannot break out of the command.
- per-workspace-env envelope: the step-10 workspace test the user asked for is no longer forbidden by the same paragraph.
- linear guides: name the full verb, ORCA linear list-issues.
- parity guard: a prefix reference such as ORCA linear --help or ORCA emulator --webcam now has its flags checked against every command under that prefix; only an exact path or an explicit ... was checked before.

* skills: tighten prose in the seven rewritten guides

Shorter outcome spines, one idea per sentence, no restated rationale after a rule. No rule, command, or pinned phrase changes; 47 net lines fewer across the guides and references.

* skills: route orca-cli and per-workspace-env gates through --reference

Both guides told agents to load --full at a gate because the per-reference
selector did not exist when they were written. Now that main serves
`skills get <topic> --reference references/<file>.md`, load only the
named file and keep --full as the fallback for an older CLI, matching the
orchestration kernel.

* skills: drop outcome-spine boilerplate from the CLI-wrapper guides

The Result/Done/Safe-failure preambles and Next Action closers restated
rules the body already carries. Agents stop fine without them, and for
a CLI wrapper the command surface is the guide. Keeps the one substantive
rule computer-use's Done block added (never report unverified as success)
inside Action Rules. orchestration and per-workspace-env keep theirs:
those are multi-step workflows where the done bar is load-bearing.

(cherry picked from commit 44a74baf73)

* skills: trim the guides and stubs to what agents actually need

- Drop the Result/Done/Safe-failure preambles and Next Action closers from
  the six CLI-wrapper guides; the one substantive rule (never report an
  unverified computer-use action as success) moves into Action Rules.
- Drop the 'guide may be stale, trust --help' lines: the guide is served by
  the binary that runs the commands, so it cannot be stale relative to it.
- Drop the status --json / open --json preflight from every guide; the stub
  no-guessing paragraph now says to start Orca only when a command reports
  it is not running.
- Cut the ORCA placeholder paragraph in each guide to one line that points
  back at the stub's resolution.
- Trim the orchestration, orca-cli, and computer-use descriptions to trigger
  phrases plus one line of scope.
- Remove the older-binary fallback section from every stub (and its two
  shared blocks); a binary without skills get gets one sentence.
- Remove the guide size ratchet test.

* skills: apply independent review cleanup

* skills: clarify guide loading and Linear command discovery

* skills: harden environment recipe examples

* test: complete branch rename journal doubles

* skills: clarify custom Codex launch and refresh model example

* test: deduplicate journal fix now present on main
2026-09-07 00:03:48 -04:00
Merge Sim 154b1bb333 fix(native-chat): degrade instead of throwing on a null tab title
A stacked branch can publish title: null when a conversation name is
cleared. The wire type says string, so this consumer trusted it and
threw inside the store patch that applies the snapshot. Fall back to the
placeholder — the producer bug is fixed separately, but a consumer of
wire data should not crash on a contract violation.
2026-09-06 21:02:41 -07:00
Merge Sim cf456eb601 fix(native-chat): stop unknown agents reading as Codex, cover the terminal path
Review found the placeholder helper encoded "unknown means Codex": its
signature accepts null/undefined and Tab.agentSessionAgent is the open
AgentType, so the first caller passing a Tab would label gemini or grok
as "Codex Chat". Routed through the shared agent-name table instead.

Also adds the missing regression test that a terminal rename still
resolves through its entityId now that both rename and color share one
resolver, and a guard on a test that passed with the fix reverted.
2026-09-06 20:46:59 -07:00
Merge Sim 3f4126a8fb fix(native-chat): reach the rename shortcut and tab color too
Review found the first fix covered only the context-menu path. The
tab.rename shortcut gated on activeTabType === 'terminal', so on a
structured chat tab it stayed the silent no-op this branch set out to
fix. setTabColor carried the identical terminal-only lookup one function
below the one that was fixed.

Both lookups now share one resolver instead of two copies.
2026-09-06 20:42:51 -07:00
Merge Sim 7424066eec fix(native-chat): make conversation naming a durable, thread-scoped decision
Twelve review findings, most of which collapse into two causes.

"Have we named this conversation" was live-session state. Both providers
rebuild their session object on every acquisition, so the flag reset and the
next message re-titled the chat — re-imposing a name the user had cleared,
and paying for a model call to do it. The marker now lives on the record
beside the name, so an eviction, a restart and a re-acquisition all see the
same answer. The record gained a clear path too: a name deleted in another
client no longer lingers here and keeps rendering. Both once-per-conversation
tests now construct a SECOND session, which is what re-acquisition does and
what the old same-object tests could not catch.

The naming turn's isolation was per-frame-kind and time-boxed. Server
requests bypassed the gate entirely, so an approval from the naming turn
became a durable prompt in the user's chat for a command they never asked
for, pending forever once the turn was abandoned; it is refused now, which
also settles the turn. Unhandled frames bypassed it too. And the gate closed
when the collector settled, including on its own timeout, so a late frame
carrying the title arrived with the gate down. The throwaway thread id is
retained for the life of the session instead, which makes the leak
structurally impossible rather than a matter of timing.

Also: a revealed or reopened chat showed the placeholder forever, because
only the startup sweep passed a title and an unchanged name never fans out —
every publication now labels itself from the record. The re-read before
`thread/name/set` failed OPEN on a reply shape this build cannot parse, and
would have overwritten a name a person chose; it fails closed. The collector
waited on a `turn/failed` that does not exist, so a refused turn stalled for
its full timeout; it settles on a non-retryable `error`, matched on thread id
— a retryable one explicitly does not interrupt the turn. The throwaway
thread is deleted when the app-server persisted it despite `ephemeral`,
verified against codex 0.153.4, which refuses to delete a truly ephemeral one
and so is not asked to. The Claude title read moved onto the bounded reverse
tail scan the transcript readers already use, extracted so both share it. And
every naming failure now logs, so a host whose app-server refuses to name
threads is distinguishable from a model that declined.

`republishStatus` was a no-op once the status summary stopped carrying a name;
removed with its comment. Prompt-text extraction moved inside the protective
promise and guards its input, because only the RPC send path validates blocks
and a throw there turns a delivered message into a failed one.

The session-id pattern forbids ':', so a prefixed key cannot reach the host's
map; the tab id and the record it is labelled from are now derived from one
stripped value so they cannot disagree.
2026-09-06 20:41:57 -07:00