A joiner is subtracted from UNRENDERABLE_RUN by design, so it breaks that run
in two and each half collapses to its own space. JOINERS_ONLY was anchored on
the joiners alone, so an interposed space made the guard miss and a zero-ink
name such as ZWJ TAB ZWJ was accepted, persisted, and — because every display
site falls back only on nullish — won over the real fallback label.
Widen the guard to blank-only, and replace the hand-picked hostile-input list
with a machine sweep of the whole invisible block so the next variant of this
cannot land.
The third walking resolver uses only files[0] but still paid a full-tree walk.
Traversal order is unchanged, so the file it returns is the same one; its
artifact-dir pruning already kept a subagent transcript from winning. Pin the
call site alongside the other two.
Three paths leave before collector.answer is awaited — a refused thread/start,
an opened thread that is unusable, and a refused turn/start — and each left the
60s timer armed, holding the collector closure and its latest answer. Dispose
from the finally so every exit clears it. Settling semantics are unchanged:
disposal resolves an already-resolved promise, which is a no-op.
Subtract the two joiners from \p{Cf} instead of enumerating the characters to
strip. The enumeration re-admitted 154 code points when only U+200C/U+200D were
the point, so a run of U+2062, tag characters, U+180E, U+FFF9-FFFB or
U+206A-206F normalized to a non-empty label that renders as nothing, and a
tag-encoded payload appended to a real title survived into the user's own Codex
history via thread/name/set.
Also reject a name left as bare joiners, which the subtraction newly admits, and
trim a joiner the length cut strands mid-emoji-sequence.
No fixture had two MATCHING files in one directory, so deleting the file-level
early return left the whole suite green — the directory-level one covered for
it. And nothing asserted that the Claude and Codex resolvers pass
stopAfterFirstMatch at all, so the perf wiring could be dropped silently.
The collector resolved its promise and left the 60s timer pending. It is
unref'd and its callback resolves an already-settled promise, so nothing
misbehaves — but the closure is retained for the rest of the timeout after
the session is gone, which is what the earlier commit claimed to have fixed.
`\p{Cf}` swept up U+200C and U+200D, which are not hostile formatting: they
hold multi-part emoji together and are orthographic in Persian and Hindi.
"Fix 👨👩👧 layout" came out as three separate people, and "میخواهم" lost its
ZWNJ — while the naming prompt asks the model to write in the user's own
language. Name the bidi controls and zero-width marks the comment already
claimed to target instead, so the hardening stays and the joiners survive.
The test that looked like it covered this asserted on "Résumé du fil ☕",
characters in categories the regex never touched.
The widening guard tested for "no newline at all", but after the first
iteration `end` always sits one byte past a newline, so an over-long line
leaves the block with exactly one newline: its last byte. `from` then equals
the block length, nothing is yielded, and `end` lands where it started —
neither loop-exit condition can ever change. The scan spun forever with the
fd held open, hanging both the conversation-name read and the pre-existing
TUI leaf-uuid read on any transcript carrying a ~120KB tool_result line.
Widen whenever the window holds no COMPLETE line, and widen geometrically:
each round re-reads the whole window, so arithmetic growth is quadratic in
the length of the long line.
Ending each read one byte past a newline means no line ever straddles a block,
so the boundary partial needs no carry at all — removing the loop-carried
Buffer.concat the quality gate flags as quadratic. Re-reading the one partial
line costs a bounded overlap and no allocation.
Naming reads the persisted conversation name on every Claude acquisition, which
resolves the transcript path through walkSessionFiles. That collected EVERY
match across ~/.claude/projects before taking files[0]: measured here at 84-103ms
warm over 2665 directories and 5911 transcripts, competing with the attach it
runs alongside. Return at the first match, which is the only one either caller
reads and leaves traversal order untouched — 86ms to 0.75ms when the transcript
is found early, unchanged when it is found late or not at all.
A session closed with a naming turn in flight held its collector until the 60s
deadline and then ran its cleanup against a dead connection. Settle it as a
host failure so the conversation stays askable. readCodexThreadId now accepts
the snake_case spelling its sibling name reader already does: an envelope whose
name is readable but whose id is not would attribute another thread's name to
this chat through the caller's session-thread fallback.
The startup sweep republishes every restored session with notify: false, but
the relabel branch never consulted it and emitted a full tab list per named
session — the desktop renderer is a real subscriber on that feed. It also
emitted the pre-store candidate, although the store can hand back a different
object, so a client mirror could retain a replaced tab under an identical
snapshotVersion; the sibling replace path already got this right.
Also corrects the merge comment that claimed a background republish cannot
re-surface a lost mirror, moves the doc block that landed between an existing
comment and the field it described, and routes the replacement tab's
placeholder through defaultAgentChatLabel.
The collapse used \s, which leaves C0 controls, zero-width runs and bidi
overrides in a name that reaches the tab strip and the sidebar row — a U+202E
renders a label that reads as text the name does not contain. Truncation used a
raw slice, which can strand a lone high surrogate. Reuses sliceAtCodeUnitLimit,
moved to src/shared so both callers can reach it.
Open the throwaway naming thread with approvalPolicy 'never' and sandbox
'read-only'. The prompt embeds untrusted user message text and the thread's
frames are diverted from the journal, so any tool the host would auto-approve
ran where the user could never see it; refusing server requests only covered
the tools that ask.
Require an exact naming-thread id on the server-REQUEST path. During the
thread/start window no naming thread has a turn running, so the broad pre-id
rule could only ever match a genuine sub-agent, whose approval request was
auto-refused with -32001 instead of reaching the user.
Also caps the structured answer before parsing it, normalizes the generated
name before it is written to the user's real thread, and releases the
throwaway thread with the protocol's own thread/unsubscribe.
Both providers claimed namingAttempted before deriving the prompt text, so a
caption-free screenshot as the first message left the chat on its placeholder
for the session's whole life. Claim the flag after the text is in hand; the
prompt reader is now total by construction, so deriving it before the promise
cannot fail a delivered message. Also reattaches the JSDoc the max-lines
extraction left on the wrapper, and softens a 'never titled on its own' claim
to what was actually observed.
Carry the chunk boundary partial as bytes, so a multi-byte character straddling
a 64KiB boundary is no longer decoded to U+FFFD on both sides and written
durably as the tab's name. Report whether the scan reached the start of the
file, and only read an emptied custom title as a clear when it did — an
ai-title beyond the bound is still the conversation's name. Type-guard
aiTitle instead of coercing it, which turned an object into [object Object].
* fix(native-chat): let a structured chat tab be renamed
Renaming a native chat tab accepted the text and silently did nothing:
setTabCustomTitle only scanned terminal tabs and only bridged to unified
tabs whose contentType was 'terminal', so the agent-session tab it was
keyed to never matched. Any label that did land was then re-nulled by the
next host snapshot, which preserved color/createdAt/isPinned but not
customLabel.
Also routes both placeholder sites through one helper so a Claude chat
stops falling back to 'Codex Chat'.
* test(native-chat): cover structured chat tab rename and label fallback
* chore: drop the local @pnpm/exe lockfile artifact
Swept in accidentally; running pnpm here adds @pnpm/exe to the root
lockfile, which fails CI's frozen-lockfile guard.
* fix(native-chat): reach the rename shortcut and tab color too
Review found the first fix covered only the context-menu path. The
tab.rename shortcut gated on activeTabType === 'terminal', so on a
structured chat tab it stayed the silent no-op this branch set out to
fix. setTabColor carried the identical terminal-only lookup one function
below the one that was fixed.
Both lookups now share one resolver instead of two copies.
* fix(native-chat): stop unknown agents reading as Codex, cover the terminal path
Review found the placeholder helper encoded "unknown means Codex": its
signature accepts null/undefined and Tab.agentSessionAgent is the open
AgentType, so the first caller passing a Tab would label gemini or grok
as "Codex Chat". Routed through the shared agent-name table instead.
Also adds the missing regression test that a terminal rename still
resolves through its entityId now that both rename and color share one
resolver, and a guard on a test that passed with the fix reverted.
* fix(native-chat): degrade instead of throwing on a null tab title
A stacked branch can publish title: null when a conversation name is
cleared. The wire type says string, so this consumer trusted it and
threw inside the store patch that applies the snapshot. Fall back to the
placeholder — the producer bug is fixed separately, but a consumer of
wire data should not crash on a contract violation.
* fix(native-chat): rename the focused structured tab, not a background terminal
* fix(native-chat): cycle terminals from the structured tab, not a stale terminal
---------
Co-authored-by: Merge Sim <sim@local>
* i18n: add localization for activity view and sidebar
Wrap activity thread state labels, interrupted status, and sidebar
title in translate() calls. Add localization keys to all five locale
catalogs (en, es, ja, ko, zh) to enable translation support.
* i18n: refactor to static keys for activity and sidebar
Convert dynamic translation key construction to static literal keys, enabling proper i18n catalog registration. This ensures activity state labels and sidebar strings are bundled in the boot catalog with their complete translations.
* i18n: change permission state label to 'Needs attention'
- Rename state label for semantic clarity across all locales
- Remove strings now using static keys (per i18n refactor to static keys)
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
The multi-stage cases pass `resolved: null`, so `resolvePrivateKeys` falls
through to `findDefaultKeyFile`, which reads `~/.ssh/id_*` via `homedir()`.
On a machine with an encrypted default key ssh2 rejects with "Cannot parse
privateKey" before authentication is exercised, so two cases failed locally
while staying green on hosted CI, which has no key. Point home at the existing
fixture directory so default-key discovery stays in the test's control.
Co-authored-by: Merge Sim <sim@local>
* feat(native-chat): enable Windows structured sessions
* fix(codex): prove native Windows process identity
* style(codex): format Windows session seam
* fix Windows structured Codex admission
* fix(windows): reprobe missing process identity capability
* fix(windows): decide folder-workspace WSL routing before the click
Review found pathUsesWslUnc exported but unused, and the folder composer
hardcoding worktreeUsesWslPath:false. Together those meant a folder picked
under a \\wsl.localhost\ parent routed to structured chat, then got refused
by the host and fell back AFTER the click -- which defeats the lane's own
design goal that create cannot fail after the click.
The group's parentPath is in scope at submit and the workspace is created
under it, so the parent decides WSL-ness pre-click. Wires pathUsesWslUnc
there and adds tests for the helper, including the unhydrated-store case
that previously threw.
* fix(windows): collapse the gate derivation to one call, restoring max-lines
CI static analysis failed: launch-agent-in-new-tab.ts crossed the 300-line
oxlint ceiling. Adding a max-lines disable is forbidden, so the two gate
derivations collapse into one readWindowsStructuredGateInputs() call --
a store-backed site now adds one line and one import name instead of two.
Better shape anyway: one derivation entry point rather than two reads a
call site must remember to pair.
* fix(windows): engage the legacy fallback when the host THROWS a refusal
Review found a P1 this merge composes: neither parent could reach it. At the
lane head the only structured entry was launch-agent-in-new-tab (full
store-backed WSL check); on main all win32 was refused. The merge enables
win32 in creation flows that pass no projectRuntime, so a WSL folder
workspace, a WSL-configured repo, or a repair-required runtime now routes
structured -- and the host refuses correctly, but by THROWING rather than
returning {ok:false, refusal}.
Callers engage their legacy-terminal fallback on the refusal CLASS, so an
unmapped throw arrives as a generic RPC rejection: no fallback, empty
workspace, error toast, prompt stranded in the launch outbox. Pre-merge the
same action opened a legacy terminal agent.
Map the host's thrown definitive refusals onto the refusal class at the
launch boundary, so every creation flow -- present and future -- degrades to
the legacy terminal instead of stranding. Narrow predicate: unrelated
failures (ECONNRESET, empty message, non-Error) still propagate untouched.
Ablation-proven: removing the mapping reddens the fallback test.
* fix(windows): teach the mobile RPC double the status probe the lane added
CI's first-ever run on this lane caught a pre-existing lane defect. The lane
changed status.get to resolve through
runtime.getStatusAfterWindowsProcessStartTimeProbe(), but never taught the
mobile-surface runtime double about it, so status.get failed for mobile
clients with "not a function". The lane's own test list did not include this
file and the lane had zero CI, so nothing ever ran it.
The real runtime always implements the method; the double omitted it.
* chore: merge current main and regenerate the localization runtime catalog
CI static analysis failed on a stale en-runtime-required.json: main added
onboarding integration-capability keys, and the generated catalog is checked
against the PR MERGE result, not the branch alone -- so it read clean locally
while failing in CI. Merging current main (90780acb85) and regenerating.
Gates after the merge: pnpm tc 0, oxlint 0, changed-code quality 0/56,
7 gate/lane test files 69 tests green.
* fix: route structured launches by execution host platform
* fix: recover paired structured session mirror on host swap
* Revert "fix: recover paired structured session mirror on host swap"
This reverts commit 81bfca0007.
* Revert "fix: route structured launches by execution host platform"
This reverts commit 47abbd354a.
* fix(windows): refuse structured chat in a paired web client
Reverts the two review-loop commits (restoring a tree byte-identical to the
validated head) and closes the hole they were aiming at, without their cost.
A paired web client's `platform` describes the browser's machine, not the host
that will run the agent, so the Windows gate cannot be evaluated there. Before
this, a browser on macOS driving a Windows runtime read "not win32", skipped the
creation-time proof entirely and allowed structured chat — fail-OPEN, the
dangerous direction, bypassing the guarantee this lane is built on.
`isWebClient` is a required input like the other gate fields, so the compiler
enumerated all seven call sites. Refusal is synchronous and fail-closed: no
async round-trip, no null window, no cache to invalidate — unlike keying on an
asynchronously-fetched host platform, which would have made every desktop
launch wait on a round-trip to fix a paired-web-only hole.
Paired web therefore gets the legacy chat until the host publishes eligibility
itself; that is the proper fix and belongs in its own PR.
Ablation-proven: removing the guard reddens both refusal tests; the
desktop-unaffected test is a preservation check and passes either way.
Gates: tc 0, oxlint 0.
Known open: repos-onboarding-folder-startup.test.ts fails on this branch and
passes on plain main — under investigation, NOT caused by this commit.
* test(onboarding): mock the web-client check the store path now reaches
The web-client refusal added `isWebClientLocation()` to the launch-route
inputs, which this suite's store path reaches while adding the FIRST folder.
The suite stubs `window` as `{ api }` with no `location`, so the function
cleared its `typeof window === 'undefined'` guard and then threw on
`window.location.pathname`.
That threw inside addNonGitFolder's own catch, so folder-1 never activated;
folder-2 then returned early (a project already existed) before reaching the
call at all, leaving exactly one activation with no startup seed.
Test artifact, not a product defect: a real renderer always has
`window.location`, so the seeding path is intact for users. Mocking the module
is the convention 7 other suites already use, and keeps product code free of
defensive branches that only exist to satisfy a stub.
Ablation-proven: removing the mock reproduces the original failure exactly.
* fix(renderer): make the web-client check total over a partial window
isWebClientLocation() guarded `typeof window === 'undefined'` and then assumed
`window.location` existed. A window stubbed without a location cleared the
guard and threw on `.pathname`.
That matters because this branch put the call on the launch-routing path,
where the throw is swallowed by the caller's catch and silently becomes a
FAILED LAUNCH rather than a visible error. CI caught it as 9 failures in
launch-work-item-direct.test.ts.
I previously "fixed" this by mocking the module in the one suite I knew about.
That was whack-a-mole against an unbounded set, and it missed this one. The
defect is the partial-window assumption, so fix it there: the mock is removed
from the onboarding suite and both suites now pass on the hardening alone.
Ablation-proven: reverting to the unguarded form reddens 11 tests across the
new unit suite and launch-work-item-direct.
Gates: tc 0, oxlint 0, changed-code quality 0/58.
* Move Codex's Windows structured-chat eligibility onto the host createSupport probe
The renderer no longer decides Codex win32 eligibility: launchStructuredAgentSession
probes agentSession.createSupport for both providers, the host answers via
supportsCodexStructuredLocation (process start-time proof + WSL refusal), and the
create path re-checks live. Deletes the client-side windows gate module and its
routing inputs (windowsProcessStartTime, worktreeUsesWslPath, isWebClient, platform)
from six call sites. Splits killCodexAppServerProcessTree out of
codex-app-server-session to hold the max-lines ceiling without a disable.
* fix(ci): keep pnpm lockfile stable
* test(windows): align foreground snapshot flags
* Restore main's pane-snapshot flag contract
Main asks for CreationTime on both projections; this branch's hot-path
isolation went away with the async probe it served.
---------
Co-authored-by: Orca Worker <orca-worker@localhost>
Co-authored-by: Merge Sim <sim@local>
Co-authored-by: Merge Sim <merge@localhost>
* fix(codex): distinguish same-email accounts in the switcher
* fix(codex): scope switcher disambiguation to the visible runtime group
Review follow-ups: wrap labels at word boundaries instead of mid-word,
disambiguate against the accounts a group actually renders, and tolerate a
missing email arriving from persisted settings or a remote summary.
`syncOrchestrationFederation()` coalesces onto an already-in-flight relay-tick
sync, which may have pulled from the peer before the caller's mutation existed.
Tests used it as a barrier, so `keeps a timed-out remote question resumable`
could reply against a home DB that had never imported the worker's question:
the reply failed with `Message not found`, no `to_worker` relay was enqueued,
and the resume ask surfaced it 5s later as a spurious timeout.
Add `syncFederationBarrier()`, which chains each active dispatch past the
current round via `syncOrchestrationFederatedDispatchAfterCurrent`, and use it
at every barrier-purpose sync site. The two tests whose subject is the sync
machinery itself keep the raw call. Also assert the reply response, so a failed
reply fails at the reply instead of masquerading as a timeout.
Production is unaffected: `syncOrchestrationFederation` has no production
callers, real read-after-write paths already use the after-current sync, and
relay ticks retry every second.
* Fix MiniMax credential-expiry reporting, region sync, and refresh
Three defects from #14929:
1. The usage endpoint answers an expired cookie or key with HTTP 200 and
base_resp.status_code 1004, never 401/403 (confirmed against both regional
hosts). The stale-token branch was therefore unreachable, so expired
credentials surfaced as 'usage-unavailable' with the raw upstream string,
and stale policy kept showing old numbers as if the failure were transient.
Classify 1004 as an expired credential.
2. minimaxEndpoint reached the SettingsUpdate schema and the web store but was
never projected by RuntimeClientSettingsController.get(), so a paired client
fell back to 'overseas' regardless of the host's region and rendered the
wrong console link. Add it to the projection and the store contract.
3. Changing the region persisted without refreshing usage, leaving the previous
host's snapshot in the status bar until the next poll. Invalidate and refetch
when the endpoint, group id, or model list changes.
The RPC-level tests mock the controller, so the projection had no real
coverage; the new test fails against the pre-fix projection.
* Localize the MiniMax credential-expiry copy
Classifying 1004 as stale-token made the status bar show the raw English
error verbatim: the new wording matches none of USAGE_AUTH_ERROR_PATTERNS,
whereas the old upstream text ('...log in again') matched and was replaced
with localized copy. That traded a localized-but-misleading message for an
actionable English-only one, which is the wrong trade for the CN users this
work targets.
Tag the error with credentialSource so the renderer can pick the right
localized string per credential kind, and add the three catalog entries.
Semantic conflict between two green PRs. #19176 added this replay test while
`agentSession.*` still admitted a `runtime` client on its negotiated capability
alone; #18700 then made `experimentalStructuredNativeChat` one rule for every
caller. Neither branch saw the other, and main runs no post-merge test gate, so
`agentSession.create` started refusing at the envelope level and the test's
`ok: true` expectation broke.
#18700's rule is the intended behaviour and `create` starts work, so it belongs
behind the gate. The fixture is what is stale: it builds a real
`OrcaRuntimeService` whose client settings are unset. Enable the setting the way
#18700 already did for the sibling pre-commit fixture. The assertions about
durable-identity replay are untouched and now actually run.
* fix: retain MSYS shell descendants in their terminal job
* test: complete MSYS regression CI registration and teardown contract
* fix(windows): deny job breakaway for the whole Cygwin/MSYS shell family
The per-PTY job probed only msys-2.0.dll, and only for bash.exe/sh.exe.
Cygwin ships the same spawn.cc breakaway logic under cygwin1.dll, and an
MSYS2 zsh escapes exactly like its bash does, so both kept the orphan bug.
Probe the runtime DLL on the shell's own search path instead of matching
shell names: that is the property that decides whether the runtime will
ask for CREATE_BREAKAWAY_FROM_JOB, and it drops the name special-casing.
* chore(patch): restore the conpty.cc index line
The earlier hand-edit dropped it while every sibling section kept one.
Recomputed against the real blobs: applying this patch to 7b286d3d
yields exactly 4b06d185, so git apply -3 has its fallback back.
* Move MiniMax quota fetch modules into rate-limits/minimax
The five MiniMax fetch/transport modules sat flat among ~110 files covering
eight providers. Nest them so the provider's fetch surface is one directory;
credential stores (main/minimax) and the IPC handler (main/ipc) stay where
their siblings are.
* Build rate-limit and settings test state from shared factories
RateLimitState was hand-copied in 9 places and the full GlobalSettings object
in 2 more, so adding one provider field forced edits in unrelated providers'
files -- which is how MiniMax fields ended up in codex-accounts and the Grok
usage-pane test.
Add createEmptyRateLimitState and createGlobalSettingsFixture and route the
copies through them. Values that deviated from the defaults are passed as
explicit overrides, so the fixtures produce what they produced before.
rate-limit-types.test.ts keeps its literal (it exists to assert the shape) and
service-state.ts keeps its own (InternalRateLimitState is a subset, not the
same type).
* Share the codex-account settings fixture between both harnesses
The two codex-account fixtures still carried the same 30-line override block
verbatim, which is the duplication the shared fixture was meant to remove.
Move it into one createCodexAccountSettings and have both call it.
Also drop the hardcoded POSIX workspaceDir default; callers supply the real
directory and a '/tmp' literal would be a trap on Windows.
* fix(runtime): apply the structured-chat setting to every RPC caller
supportsStructuredAgentSessions only consulted experimentalStructuredNativeChat
when clientKind === 'mobile', so identical host settings admitted desktop and
in-process callers while refusing a phone. The server branched on client surface.
The setting is now one rule for every caller. The negotiated capability stays a
wire term asked of remote clients only, so a capability-less in-process caller is
still admitted on the setting alone.
Making the projection's structuredNativeChatEnabled argument required surfaced
eight call sites that passed `undefined` for non-mobile clients; they now read the
host setting, so tab projection follows the same single rule.
Announced behaviour change: with the flag off, session.tabs.list/listAll no longer
restore structured tabs for desktop. The desktop renderer already discards them in
that state, and startup record/lease reconciliation is unaffected.
* fix(runtime): keep structured session cleanup available
* test(runtime): enable structured chat in desktop projection fixture
* test(agent-session): settle merged fixtures against the all-clients structured policy
The merge with main left three fixtures written for the old mobile-only rule:
a duplicate getClientSettings key, a create fixture with no host settings at
all, and a projection call whose 'old client' is now the mobile fallback-title
case.
* fix(native-chat): let an admitted caller close a chat after the setting is off
Turning `experimentalStructuredNativeChat` off revoked admission for every
`agentSession.*` method, including `close`. A chat opened while the setting was
on stays mounted, so its owner was left with a live provider child and an X
button that answered `structured_agent_session_unsupported`.
Split the surface by what a method does to work in flight rather than by how it
sounds, and write that rule where the gate lives so the next method lands on the
right side: starting, extending, retaining or reading needs admission; stopping
or retiring work the caller already owns does not. Moves `close` and `cancel`
onto the cleanup gate alongside `unsubscribe` and `release`.
The tightening is unchanged - the cleanup gate still demands the negotiated wire
capability and never creates a host, so an incapable client still cannot see the
surface and no method that starts work is reachable with the setting off.
Extracts the dispatcher harness and the method-to-gate table into fixtures so
the new admission suite can share them without a max-lines disable.
* Drop a duplicate lastActivityAt key carried in from main
The main commit this branch merged (fb322046e8) had two lastActivityAt
properties in the same object literal at both journal stubs, which fails
TS1117 and oxlint. Upstream has since kept only the later value; match it.
Not introduced here, but merged in, so it has to be fixed here.
---------
Co-authored-by: Merge Sim <sim@local>
* feat(native-chat): resume an Agent Session History row into a new structured chat
A Claude or Codex row in Agent Session History gains "Resume in New Chat": it
opens a new structured native-chat tab that continues that provider
conversation, with the prior turns already in the journal. Until now those rows
could only be resumed into a PTY terminal; the structured branch could reveal a
chat Orca already owned but could not adopt one it had never held.
Almost all of the machinery existed. Both lanes already resume from the record's
provider handle chain, the journal already has a transcript importer, and the
handle chain already models `adopted` as an origin. The gap was that a create
always minted an empty chain, so the adapters started a fresh conversation. This
seeds that chain.
The client names only the conversation. `agentSession.create` is reachable by
paired mobile clients, so the transcript path and the account home are derived
by the executing host and validated against the account homes it recognises —
a client-supplied path would choose which file the host imports and which
credential directory the provider child launches against.
Failure refuses rather than degrades. A transcript that cannot be found refuses
before anything is created; one that fails or decodes empty *after* the provider
has resumed fails the attach, tearing the child down and publishing no tab,
because an empty journal beside a context-carrying agent claims a continuity the
provider never gave.
Codex can resume into any workspace since it is handed the rollout path; Claude
resolves transcripts under a project key derived from the launch cwd, so it is
offered only for the workspace the conversation was recorded in.
* fix(native-chat): widen adopted-home discovery and keep ordinary launches untouched
Three corrections from review of the first commit.
The adoption's account-home candidates now include the extra Codex homes session
discovery already scans. A row this host listed could otherwise refuse to
resume, which reads as the feature being broken rather than as a scope.
Ordinary launches call `createStructuredAgentSessionLaunchIntent` with two
arguments again. Passing the resume source unconditionally appended a trailing
`undefined` that four existing call-site assertions had to absorb; the churn was
the caller's fault, not the tests'.
The transactional adoption guard's comment claimed the self-exemption is what
lets a committed create replay. It is not: replay is settled earlier by the
operation ledger, and an adoption always arrives with a null expected fence, so
a request naming an existing session id is refused a few lines below either way.
The exemption is part of what "another record" means, and the comment now says
that instead.
* fix: preserve history adoption through create and retries
* fix: replay committed history adoption from durable identity
* fix: validate history before claiming adopted sessions
* fix: extract AI vault resume domains
* fix: recognize typed history resume refusals
---------
Co-authored-by: Merge Sim <sim@local>
* fix(editor): support Shift+wheel in combined diffs
* add active modified-pane test
* fix(editor): skip shift-wheel capture when a diff pane cannot scroll sideways
Word-wrapped panes never overflow horizontally, so consuming the gesture
left it dead instead of reaching the outer combined-diff list.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
A null answer carried three different facts and was read as one, wrongly in
both directions.
A model that ignored the output schema and answered in prose parsed to no
title and was marked SETTLED, so that conversation became permanently
unnameable — switching to a schema-capable model later never rescued it. And a
model that completed having said nothing is a real decline, but was marked
UNSETTLED, so it was re-asked, and paid for, on every later acquisition.
The collector now reports how the turn ended — answered, declined, failed, or
timed out — and the flow classifies on that. Declines settle; a host or model
that could not be asked properly does not. This also defuses a sub-agent reply
landing in the pre-id window and becoming the naming answer: that costs one
wasted attempt now rather than forfeiting the conversation for good.
Clearing a name also marks it attempted. A chat the user titled by hand in the
CLI and then deleted had no name and no marker, so their very next message
generated a replacement — undoing the deletion, which is the exact failure the
durable marker exists to prevent and which its own comment already promised.
Comments corrected rather than left overclaiming: the pre-id window is bounded
by the request deadline rather than instantaneous; leak protection is now exact
id matching and so relies on the app-server not retagging naming frames; the
attempted marker's write moved after the await, which widens the window where
an eviction can start a second turn; and an unreadable thread id still leaks a
persisted throwaway thread past the new cleanup path.
Two tests encoded the old classification by using prose as a stand-in for a
decline. The fake can now complete a turn having said nothing, which is the
fact they meant to exercise.
An emptied `custom-title` returned `cleared` immediately, so a transcript
holding an ai-title plus a later emptied custom title wiped the durable name
and reverted the tab to its placeholder while the CLI still displayed the
ai-title. The file's own stated precedence says custom outranks ai — an EMPTY
custom slot falls back to the ai slot, which is what the code did before this
round. Only an emptied custom slot with nothing to fall back to is a clear.
It also read `customTitle ?? ''`, so an absent, null, or renamed field counted
as a deliberate clear. That is the opposite of the discipline applied on the
Codex side, where a reply shape the build cannot read is refused rather than
acted on. Wrongly clearing destroys a name the user can still see.
The only clear test stubbed the reader and exercised the reporter, so the
mechanism was bound by nothing: reverting the reader left every test green.
The new tests write real transcript records — set-then-emptied, emptied with
no fallback, and three unreadable field shapes — and the ablation now fails
four of them.
The reporter's deps are mapped explicitly beside the reporter rather than
handed the adapter's object by reference, which carried an `onError` key the
adapter type never declared and would have dropped silently the day that
object was narrowed.
Also: the sub-agent approval fixture had no `itemId`, so the registry declined
it and the request was still refused, just differently. It now carries one and
the test asserts a durable prompt reaches the user, which is the behaviour the
narrowed gate restored. And the header comment claimed the scan is not
repeated; it runs on every acquisition, deliberately, because that is how a
CLI-side rename reaches Orca at all.
The turn runs on the session's own model. The app-server catalog carries no
structured signal for a small, fast one — `modelSpecialty` is null on every
entry — so selecting one would mean matching a description string or a vendor
id shape like `-mini`, and neither survives a different provider. Naming a
model the account cannot use fails the turn outright, which is worse than a
turn that works and costs slightly more.
The cost is bounded on the other axes instead: lowest reasoning effort, the
prompt capped, and a schema that caps the answer at 36 characters. It runs
once per conversation and off the send path.
Checked while deciding: every model in the catalog supports `low`, and
`turn/start` accepts an unknown effort silently rather than rejecting it, so
no capability fallback is warranted here.
Clearing a name crashed the renderer. `applyStructuredAgentSessionConversationName`
declared its input non-null and assigned straight onto `tab.title`, but the
caller passes null on a clear, and `title` is a required string every client's
snapshot builder calls `.trim()` on. That file is @ts-nocheck, so a green
typecheck proved nothing about it. It falls back to the agent's placeholder
now, which is what its own doc comment already claimed.
The naming gate ate Codex SUB-AGENT threads. It treated any thread that was
not the user's as a naming thread for as long as a naming turn was open, and
that window opens at the first turn's ack. So a sub-agent's rows were dropped
from the transcript, its reply could be taken as the naming answer, and its
approval request was auto-refused with a reason that was untrue. The broad
rule now applies only before `thread/start` returns the throwaway id; after
that it is an exact match.
The durable "we already asked" marker was written BEFORE generation, so any
host failure forfeited naming permanently — a Claude CLI without the title
request, or an app-server that rejected the throwaway thread, marked every
conversation it touched and none of them could be named after an upgrade. It
is written only on a settled answer now: a title, or a model that answered
without one. The Claude control surface reports `unsupported` separately from
`declined` so the two are distinguishable at all.
Also: Claude gained the clear path Codex already had, keyed on a positively
emptied title record rather than on absence, since the tail scan is bounded
and absence is not evidence. The blocks guard covers elements, not just the
array. Naming failures log `error.message` rather than the raw error, which an
app-server could echo the user's prompt into. Two comments corrected and two
tests retitled to stop claiming coverage they do not have.
Base branch merged first, so this is tested against what will land.
A structured chat has no TerminalTab, so its sidebar row is synthesized from
the status entry and its name arrived as `terminalTitle` — which
getAgentRowConversationName reads as a LIVE title and runs through
conversationNameFromLiveTitle. Those heuristics exist to reject OSC titles
scraped off a pty, where a bare path or a status word is not a name. They are
wrong for a conversation name, which is a deliberate value from a known-good
source.
So a chat named `auth/login` or `src/renderer` showed in the tab strip and
vanished from the sidebar, which fell back to the last-message label:
isCwdLikeTitle nulls any single word containing a slash. `done`, `Claude` and
`Terminal 4` went the same way, and a manual rename tripped the same
predicates as a generated name.
The name now travels as its own authoritative field from the bridge to both
row synthesizers, which surface it where the resolver returns it verbatim.
Nothing about the sanitizer changes: a real scraped title still gets it.
Tested with the values that actually trip each predicate rather than a name
that would pass either way.
* skills: rewrite the seven non-orchestration guides to one outcome-first standard
Every guide leads with Result / Done / Safe failure, states conditions instead of case lists, keeps one done bar and one autonomy envelope, and loads references at the point of use via `skills get <topic> --full`. orca-cli drops from 424 to 260 always-loaded lines with three references; orca-per-workspace-env from 794 to 397 with five.
Defects fixed in shipped guides: `emulator camera` (no such command), iOS `permissions` (backend refuses it), Android pane described as in development, `relayGracePeriodSeconds: 0` documented as immediate teardown (it is unbounded), doctor `ok: true` hiding `warn`, an SSH exemplar setting both `jumpHost` and `proxyCommand`, a provisioned-root fetch from `origin`, and the Linear unconfirmed-write rule keyed on four verbs when ten emit it.
The resolver ladder, placeholder rule, and older-binary fallback shared by every installable SKILL.md now come from one skill-stubs/_shared/cli-resolution.md fragment composed by the generator, which also bundles per-guide references into --full. New guards: every ORCA invocation and flag resolves against COMMAND_SPECS, descriptions carry no angle-bracket tokens, reference routing is checked both ways, and an always-loaded size ratchet (300 lines) that guides may leave but never join.
* skills: address review on the SSH recipe and the parity guard
- ssh-host create script: route the bootstrap ssh through the chosen jump host or proxy command, refuse both at once, use StrictHostKeyChecking=accept-new instead of a blind ssh-keyscan append, and pass gh_token/project_root/repo_url/repo_ref to the remote bash via printf %q so a quote in a value cannot break out of the command.
- per-workspace-env envelope: the step-10 workspace test the user asked for is no longer forbidden by the same paragraph.
- linear guides: name the full verb, ORCA linear list-issues.
- parity guard: a prefix reference such as ORCA linear --help or ORCA emulator --webcam now has its flags checked against every command under that prefix; only an exact path or an explicit ... was checked before.
* skills: tighten prose in the seven rewritten guides
Shorter outcome spines, one idea per sentence, no restated rationale after a rule. No rule, command, or pinned phrase changes; 47 net lines fewer across the guides and references.
* skills: route orca-cli and per-workspace-env gates through --reference
Both guides told agents to load --full at a gate because the per-reference
selector did not exist when they were written. Now that main serves
`skills get <topic> --reference references/<file>.md`, load only the
named file and keep --full as the fallback for an older CLI, matching the
orchestration kernel.
* skills: drop outcome-spine boilerplate from the CLI-wrapper guides
The Result/Done/Safe-failure preambles and Next Action closers restated
rules the body already carries. Agents stop fine without them, and for
a CLI wrapper the command surface is the guide. Keeps the one substantive
rule computer-use's Done block added (never report unverified as success)
inside Action Rules. orchestration and per-workspace-env keep theirs:
those are multi-step workflows where the done bar is load-bearing.
(cherry picked from commit 44a74baf73)
* skills: trim the guides and stubs to what agents actually need
- Drop the Result/Done/Safe-failure preambles and Next Action closers from
the six CLI-wrapper guides; the one substantive rule (never report an
unverified computer-use action as success) moves into Action Rules.
- Drop the 'guide may be stale, trust --help' lines: the guide is served by
the binary that runs the commands, so it cannot be stale relative to it.
- Drop the status --json / open --json preflight from every guide; the stub
no-guessing paragraph now says to start Orca only when a command reports
it is not running.
- Cut the ORCA placeholder paragraph in each guide to one line that points
back at the stub's resolution.
- Trim the orchestration, orca-cli, and computer-use descriptions to trigger
phrases plus one line of scope.
- Remove the older-binary fallback section from every stub (and its two
shared blocks); a binary without skills get gets one sentence.
- Remove the guide size ratchet test.
* skills: apply independent review cleanup
* skills: clarify guide loading and Linear command discovery
* skills: harden environment recipe examples
* test: complete branch rename journal doubles
* skills: clarify custom Codex launch and refresh model example
* test: deduplicate journal fix now present on main
A stacked branch can publish title: null when a conversation name is
cleared. The wire type says string, so this consumer trusted it and
threw inside the store patch that applies the snapshot. Fall back to the
placeholder — the producer bug is fixed separately, but a consumer of
wire data should not crash on a contract violation.
Review found the placeholder helper encoded "unknown means Codex": its
signature accepts null/undefined and Tab.agentSessionAgent is the open
AgentType, so the first caller passing a Tab would label gemini or grok
as "Codex Chat". Routed through the shared agent-name table instead.
Also adds the missing regression test that a terminal rename still
resolves through its entityId now that both rename and color share one
resolver, and a guard on a test that passed with the fix reverted.
Review found the first fix covered only the context-menu path. The
tab.rename shortcut gated on activeTabType === 'terminal', so on a
structured chat tab it stayed the silent no-op this branch set out to
fix. setTabColor carried the identical terminal-only lookup one function
below the one that was fixed.
Both lookups now share one resolver instead of two copies.