mirror of
https://github.com/stablyai/orca.git
synced 2026-10-01 00:02:10 +00:00
a0eee2fc21aeea36169717e98cc58b024f4e8b9d
10725
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a0eee2fc21 |
feat(mobile): forward host-advertised subscriptions
Add bounded generic streams with Desktop cleanup metadata, direct and relay cancellation, and page-owned source-control invalidation. Preserve legacy fallback and protocol 2. Required code gates and page export pass; platform journeys remain tracked. |
||
|
|
42c031d85e |
fix(mobile-web): preserve chat bindings across unavailable host reads
Keep authority on failed or malformed tab snapshots. Exercise headless terminal file links through line-addressed hosted previews, since renderer-backed file tabs require a Desktop renderer. Record the observed native tap and host resolution evidence. |
||
|
|
d39be4951d |
fix(mobile): repair hosted journey accessibility and route checks
Label standalone toolbar actions, match WebView accessibility values, and fence workspace return checks by pathname. Reuse Android labelled navigation and support adversarial journeys without native screenshot baselines. Record passing iOS file parity and the unresolved terminal file-link gate on both platforms. |
||
|
|
31024ff031 |
feat(mobile-web): serve page-safe file searches and text reads from desktop
Reuse host file methods behind Desktop-owned identity redaction and move list/text presentation into the hosted page. Preserve older shells and Desktop versions through legacy fallback, without changing the generic bridge contract. |
||
|
|
a8bbed52da |
feat(mobile-web): move directory and chunk reads to the generic bridge
Advertise page-safe file reads from Desktop and project their raw results in the hosted page. Preserve legacy handlers and fallback for older shells and oversized replies. Track the remaining long-lived shell implementation and platform evidence. |
||
|
|
9910fccc29 |
feat(mobile-web): forward desktop-advertised source-control reads
Add authenticated method catalog queries and a bounded generic unary workspace lane. Preserve opaque workspace bindings across awaits and keep actual host calls counted after page cancellation. Move source-control read presentation to code the hosted page can run, with legacy fallback for older shells, older desktops, and oversized raw responses. Preserve v2 and all legacy handlers. Update dispatch, reauthorization, and response corpus ratchets deliberately. Full requested gates pass; the mobile suite passed after rerunning outside concurrent web export. Broader identifier families, generic subscriptions, and native route migration remain separate work. |
||
|
|
11646f11e0 |
test(mobile-web): ratchet the installed bridge protocol floor
Pin version 2 independently of release constants, keep installed APK package admission, and preserve cached version-2 page admission. Add a native shell-init regression using a literal shipped protocol version. All requested TypeScript, lint, quality, mobile and runtime tests, web package build, and Kotlin unit gates pass. |
||
|
|
0eb6cadae2 |
Merge origin/main into mobile-rearch
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
9fbca91320 |
Merge parity-0906: hosted agent status fields and transcript projection onto the wire
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
15d0f8aedf |
skills: rewrite the seven non-orchestration guides to one outcome-first standard (#18724)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 6 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$544 | $\color{#cf222e}{\Huge{\mathbf{−}}}$49 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$495 |
| Prod | 36 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$1719 | $\color{#cf222e}{\Huge{\mathbf{−}}}$1703 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$16 |
<!-- /orca-pr-loc -->
## ELI5
Orca ships eight skill guides that agents read before running the CLI. Seven of them (everything except `orchestration`, which #16904 rewrites) were command catalogs that had drifted from the binary. This PR rewrites them so an agent reads the outcome, the done bar, and the safe-failure rule first, loads reference material only at the step that needs it, and never sees a command or flag the installed CLI does not define.
## What changed
- **Seven guides rewritten** to one standard: outcome spine first (Result / Done / Safe failure), conditions instead of case lists, one done bar, one autonomy envelope, references loaded at the point of use via `skills get <topic> --full`, every runnable invocation spelled `ORCA`. `orca-cli` is 424→260 always-loaded lines with three references (browser, automations, publishing); `orca-per-workspace-env` is 794→397 with five (provider-vercel, ssh-host, docker-ssh, windows-scripts, failure-modes).
- **Defects fixed in shipped guides:** `emulator camera` (no such command), iOS `permissions` (backend refuses it), Android pane described as "in development" (shipped in June), `relayGracePeriodSeconds: 0` documented as immediate teardown (it is unbounded), doctor `ok: true` hiding `warn`, an SSH exemplar setting both `jumpHost` and `proxyCommand`, a provisioned-root fetch from `origin`, the Linear unconfirmed-write rule keyed on four verbs when ten emit it. Linear and emulator descriptions dropped embedded commands and angle-bracket placeholders (651→329, 732→404 chars).
- **Generator bundles references.** `skill-guides/<name>/references/*.md` is appended to `--full`; `skills get` help says compact by default, full with references.
- **Stubs single-authored.** The resolver ladder, placeholder rule, and older-binary fallback shared by all eight installable `SKILL.md` files come from one `skill-stubs/_shared/cli-resolution.md` fragment composed by the generator. Projections were byte-identical before the content fixes.
- **Guards:** every `ORCA <cmd>` and flag in every guide and reference resolves against `COMMAND_SPECS` (this found the camera defect); descriptions ≤1024 chars with no angle-bracket tokens; reference routing checked both directions; an always-loaded size ratchet (300 lines) that guides may leave but never join. `orchestration` (440 lines on main) is recorded as an exception until #16904 lands its kernel.
## Relationship to #16904
Split out of #16904 so that PR carries only the orchestration guide. On main, `terminal send` has no `--wait-submit` / `--retry-request` and the orchestration kernel still carries the resolver ladder and worktree-selector rule, so this branch pins `accepted: true` for handoff receipts and leaves the orchestration pins where main has them. The merge in either direction is mechanical: #16904 rebased on this becomes a one-file `orchestration.md` change plus dropping the two exceptions.
## Standard
Compound Engineering's portable skill-authoring guidance (outcome spine, conditions not cases, pinned fragile commands with an ordered hatch, references at point of use). NVIDIA SkillEvaluator Tier 1 (`schema,pii,license,quality,unicode,lint`) was run on every guide; its deterministic checks pass, its template nudges (Instructions/Examples sections, 50–150 char descriptions) do not apply to Orca's stub architecture and were not applied.
## Testing
- `pnpm typecheck:tsc:cli` clean; `check:code-quality:changed` and `check:react-doctor:changed` 0 findings
- `pnpm verify:bundled-skill-guides` and skill-bundle manifest verify clean
- vitest over `config/scripts`, `src/cli/skill-guide-cli-parity.test.ts`, `src/cli/skills.test.ts`, `src/cli/specs/skills.test.ts`, `src/cli/help.test.ts`, `src/main/skills`: 240 files / 2,019 pass
- Live smoke on the built CLI of every `skills get <topic>` and `--full`, every emulator, linear, and vm verb named in the guides, and every projection's resolver, GNOME warning, and bounded fallback (done on the #16904 branch before the split; the guide bodies are identical here except the send-receipt vocabulary noted above)
## Deferred product decisions
Merging `orca-emulator` and `orca-emulator-android` into one skill with a platform branch; collapsing `linear-tickets` to a guide alias; a `skills get --reference <name>` selector so a gate table can load one file; a fresh-agent routing eval before trimming the `orca-cli` (1,015 chars) and `orchestration` descriptions, whose quoted triggers each fixed a routing misroute.
|
||
|
|
1cdfef70ce |
fix(mobile-web): project host transcript content onto the wire instead of refusing it
The shell parsed the host's transcript reply against a .strict() page contract. The host already publishes three things that contract has never carried, all on the transcript lane the hosted page uses: - Claude's resolved edit hunks (editPatch), on every Edit tool result since #18765; - a structured providerFrame on a text block; - text up to MOBILE_TEXT_BLOCK_CHAR_CAP (64 KiB), where the wire allows 4200. Each was an unrecognized key or an over-long string, so the parse threw. On read the page got 'Transcript read failed' and rendered nothing; on subscribe sanitizeEvent returned null and the ledger cancelled with invalid_message, which is not retryable, so the live transcript died mid-turn. The native app reads the same host method directly and showed all three fine, so this was a hosted-only blank screen rather than a missing feature. The shell now projects each message onto the contract: text and tool output are clipped with the truncation marker the host itself uses, the tool-call lifecycle state is carried, a block whose shape the page cannot render is dropped without its siblings, and a message whose id would break dedup is dropped rather than clipped. A projected corpus is asserted to always satisfy the schema, and the .parse stays behind it as a fence. editPatch and providerFrame are dropped, not widened: mobile renders edits through diffFromToolCall and has no provider-frame row, so carrying them would be wire weight no component reads. Both are additive shell->page fields if a mobile renderer for them ever lands. Supersedes sanitizeMobileWebNativeChatMessages, which only reached tool inputs. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
08b96ed1b3 |
Seed Cmd-J filter from sidebar scope (#19036)
* feat(palette): seed Cmd+J filter from sidebar show scope When opening Cmd+J, the palette's host and project filters now initialize from the sidebar's current Show scope, so results match the user's sidebar view. The palette can still be cleared or changed per open; sidebar never reads back palette filters. * refactor: pass app state to palette filter builder Let the builder function extract the sidebar scope it needs instead of requiring callers to destructure and pass individual properties. This reduces coupling and simplifies the data flow through the palette initialization lifecycle. * Make palette filter repo-granular to preserve sidebar scope Filter options now list individual repositories instead of grouping multi-repo projects into single rows. This preserves the exact repository scope shown in the sidebar when opening Cmd+J, rather than widening selections to entire projects. Removes per-field selection cap and stale-value reconciliation, simplifying the filter lifecycle. * Clarify filter naming and seed from sidebar scope on palette open - Rename projects→repositories in PaletteFilterModel for semantic accuracy - Rename rawFilter→filterState for clearer intent - Initialize filter from sidebar scope in local state, refresh on open - Remove redundant filter reset from selection lifecycle * Seed Cmd-J filter from sidebar scope and reset on close The palette now opens with the sidebar's host and repository scope applied. Filter changes are temporary: closing discards them, and reopening reseeds from the sidebar's current state. - Repository filtering is now granular (individual repos) - Support shared repository IDs across multiple hosts - Disambiguate duplicate repository names by path * Add comment clarifying Projects terminology Document the naming convention for repository-granular filter choices to help future maintainers understand why "Projects" is used as the user-facing term. * Remove redundant Escape press from worktree palette filter test |
||
|
|
20104bfd54 |
feat(mobile-web): carry working mode, tool-output and model on the hosted agent status
The hosted page renders the same Expo chat components as the native app, so a field the shell drops is a behaviour difference, not a missing view. Three were dropped. `workingMode` never crossed the bridge, so `isMobileNativeChatAgentWorking` could not see a monitoring agent and the page showed a busy indicator, Stop and a streaming bubble the app deliberately withholds. `lastAssistantMessageIsToolOutput` never crossed it either, so the page streamed raw tool output into the chat bubble. `model` was already in the contract but the shell never projected it, so the session-option controller always read a null model. Also accepts the tool-call lifecycle `state` the structured lane publishes. The contract refused the key outright, so the first producer to set it on this lane would have failed the whole message rather than been ignored. All four are Rule 1 optional shell->page fields. The two closed sets collapse to absent under the tolerant page parse, and absent is each field's pre-existing reading: a foreground agent, and a call whose liveness the turn's working flag decides. Splits the agent-status projection and the host value bounds out of mobile-web-session-snapshot.ts, which had 16 lines of max-lines headroom left. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
33462b8926 |
Merge cleanup-0906: one shared terminal link matcher, image-echo ordinal fix, speech ledger on the base, writeIfUnchanged external-writer pin
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
94c76c0657 |
refactor(terminal-links): make one shared file-link matcher serve desktop and mobile
src/shared/terminal-file-link-matcher.ts was a branch-only copy of the renderer matcher, frozen at its 2026-07-27 shape. Main has changed the renderer twice since, including widening LOCAL_PATH_REGEX to Unicode property classes, so mobile terminal taps missed every non-ASCII path that desktop already linked. Moves the renderer's file-link cluster to src/shared and deletes the copy. Nothing in it was renderer-only: the eight modules import each other plus file-link-location and terminal-file-url-target, which already live in shared. Mobile's tap and web link provider now import the same module, so they also pick up the file-uri and bare-filename passes the copy never had. matchTerminalFileLinkAtColumn moves in with them, replacing both the copy's version and the renderer test's local helper. The WebView-injected copy has to stay a copy, since that script is interpolated into HTML and cannot import. Its path regex gets the same Unicode widening, and the drift that let it diverge in the first place is closed by two non-ASCII cases in the shared conformance suite, which all three implementations run. Reverting either widening reddens them. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
cdb52b14f8 |
Merge csp-0906: allow https images in the rich markdown and preview frames
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
d2d825b670 |
fix(mobile): allow https images in the rich markdown and preview frames
main shipped both surfaces with no CSP at all: the rich markdown editor
document had no Content-Security-Policy meta, and MobileHtmlPreview fed raw
agent HTML into a react-native-webview with originWhitelist={['*']} and no
sanitizer. Remote https images therefore rendered on main. The branch added
'img-src data:' to both frames, which silently broke them.
Restores parity by adding https: to img-src in both frame policies, and by
letting the preview sanitizer keep an img src that parses as https:. The
editor's own markdown renderer already admitted https urls, so only its CSP
needed widening. Plaintext http: stays blocked in both: the preview sanitizer
strips it and neither policy lists it. script-src, connect-src and frame-src
are unchanged, as are the img attributes the sanitizer copies.
The mobile-web shell policies are deliberately untouched. Their second policy
is the mermaid frame policy, keyed on mermaid-frame.html, so it never governs
these frames; and the shell blocks every http(s) load below CSP anyway via
blockNetworkLoads on Android and a WKContentRuleList on iOS, so widening it
would loosen isolation for no rendering gain.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
|
||
|
|
1fbd496371 |
Merge wire-0906: terminal reply authority as a host fact, tolerant init resumeRoute, gzip ceiling client note
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
e5450ac8b8 |
docs(mobile-web): record who the package gzip ceiling's client actually is
Investigated whether the gzip residual left by
|
||
|
|
0995ced242 |
fix(mobile): count image echoes by the normalized key, not a raw trim
mobile-native-chat-pending-echo.ts was reported dead. It is not: five of its exports are live through use-mobile-native-chat-pending-deliveries. What was dead is one export, appendMobileNativeChatPending, which the delivery hook had re-inlined instead of calling — and the two copies disagreed. The hook picks the image-echo branch when the send's normalized text is empty, but counted prior image echoes with pending.text.trim() === ''. A marker-only caption like '[Image #1]' normalizes to empty yet does not trim to empty, so it took ordinal 1 and was then not counted, handing the next caption-less photo ordinal 1 as well. Both echoes then reconciled against the same transcript row. The dead export already shared the discriminator correctly, so the hook now calls it. Adds the drafts-level regression test; both it and the existing pending-echo case yield [1, 1] against the inlined version. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
266d6eded1 |
fix(mobile-web): degrade an unknown init resumeRoute instead of failing the page
`init.resumeRoute` is MobileWebResumeRouteSchema.optional(), a discriminated union whose `kind` is a closed literal set. The shell->page tolerant parse rescued unknown members inside arrays and unknown values for an optional closed set, but isClosedSet only accepts enums and literals, so an optional discriminated union of objects got no relaxation: a page built before a kind existed failed the whole init. That is the worst frame to drop -- init is the page's only grant delivery, so one unrecognized route cost it every capability. The transform now treats an optional/nullable discriminated union like an optional closed set, because its discriminant is one. Scoped deliberately to an UNRECOGNIZED discriminant rather than a blanket .catch on the wrapper: a member the page can name but whose fields break their bounds is a sender bug, not version skew, and still fails loudly. The existing 'rejects unbounded resume routes' assertion (a 241-character workspaceName on a known 'session' route) therefore keeps failing the parse, and the PII strip on hostPath is unchanged. Page->shell stays strict, which is what fences the shell's route memory: useMobileWebResumeRouteMemory only stores what a strict routeState parse produced, so the shell can never remember a kind its own build cannot replay. The persisted cold-resume record (mobile-web-cold-resume-route) stores hostIdentity and hostWorkspaceIdentity with no kind at all, so it cannot carry one either. The only way the shell holds a route the current page rejects is a mid-session page downgrade, and that now degrades to the page's default route on every boot rather than bricking it. Tests: the transform collapses an unrecognized discriminant and still rejects a malformed known member and a non-object; a real init carrying kind 'someFutureKind' parses through both page entry points with resumeRoute absent and both grants intact; the page channel opens workspaceList and keeps its client; a routeState naming an unknown kind is refused. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
3f069dfc9c |
refactor(mobile-web): put the speech ledger on the shared subscription base
The speech ledger was the one shell ledger still hand-rolling its record map, cancellation and posting because it is push-driven. It maps onto MobileWebSubscriptionLedger with no base-class change: admit plus a records.set stands in for open, and enqueue carries the broadcast. The authority now takes a MobileWebSubscriptionLedgerConfig, which retires the per-subscription post/closed plumbing, and the broker hands all three consumers one shared subscriptionPosts() sender. Speech stays out of replaceClient's closeAll on purpose: its feed is the shell's device runtime, not the host RPC client, so it survives the swap and a terminal closure frame would wedge the dictation hook in error with no resubscribe. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
63cb1ea524 |
test(rpc): pin the external-writer conflict on files.writeIfUnchanged
The reported gap does not exist: the host re-reads the file inside the serialised section and compares sha256 against expectedRevision before writing, so an external editor between the client read and the write is already rejected. Nothing pinned that, so a refactor could drop the hash compare and leave runKeyedSerializedOperation as the only fence, which only orders Orca's own RPCs. Adds a regression test where a queued save carries a stale revision while an external writer rewrites the file. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
16e06f2c3d |
refactor(mobile-web): make the terminal stream's reply authority a host fact
`inputFloor` and `queryReplyAuthority` on the shell->page terminal stream were both
literals the shell fabricated ('held' / true) on every subscribe. Traced each to what
the desktop host actually publishes.
inputFloor: deleted. The desktop's mobile input floor is claimed lazily at write time
(RuntimeTerminalDriverController.beginMobileInputFloor, from terminal-input-delivery)
and is never published on subscribe or on any stream event. Opcode 17 WriteUnavailable
is a per-write refusal with no regain frame, and the mobile-web shell does not even
advertise writeUnavailable:1 in its subscribe capabilities, so it never receives one.
'read-only' was therefore unreachable and canSendInput reduces to the hostReady flag
the scheduler already tracks. The scheduler's own invisibility reset (setVisible(false))
is the real revocation path and still clears hostReady, so behaviour is unchanged.
queryReplyAuthority: kept, renamed queryReplyNegotiated, and sourced from the host. The
host already echoes capabilities.queryReply:1 on the multiplex 'subscribed' frame (the
Rule 2 handshake for opcode 18), and the shell already reads it into
record.supportsQueryReply. That echo is a negotiation, not the election verdict --
isMobileTerminalQueryReplyAuthority is re-evaluated per frame on the host and never
sent -- so the old name asserted something no host computes. Made optional in the
shell->page schema: absent means a shell that predates the field and cannot prove
negotiation, so the page must not attempt a reply. No new host->client field was needed.
Also fixed the downgrade this exposed: when the host had not echoed the capability, the
shell sent the reply bytes under opcode 0 (Input). Those hosts strip inputKind and take
reply bytes as floor-taking shell input -- the exact hazard
TERMINAL_QUERY_REPLY_INPUT_RUNTIME_CAPABILITY documents and the native path already
drops for. The shell now drops instead, and the page stops sending in the first place.
Lease-only streams publish queryReplyNegotiated:false; they negotiate no output
multiplex and every input request on them already fails not_found.
Metadata events carried the same two fields and are never emitted by the shell today;
both are gone from that event, leaving it the displayMode carrier it is.
Compatibility: host->client is untouched (no new field, no changed frame). shell->page
is branch-local; the new field is optional and the page's tolerant parse reads absence
as not negotiated.
Tests: page reads an omitted queryReplyNegotiated as false; the shell reports false and
drops the reply when a host omits the capability echo, and true when it sends it.
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
|
||
|
|
311ea8bf9c |
refactor(mobile-web): move one-shot terminal requests to their own request client
The bridge client crossed the max-lines budget when the shell-feature query joined it. The two terminal one-shot requests were the only operations still inlined there; every other capability already has a request client. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
cfc2856ac3 |
Merge origin/main into mobile-rearch: structured native Claude chat
Merges #18741, which opens the structured agent-session lane to Claude on mobile. Three conflicts, all in mobile/src/session: - mobile-native-chat-eligibility.ts: import block only. Kept the branch's MobileWebNativeChatAgentStatus/AgentWorkingMode pair and added main's isAgentSessionHandleProvider. Main's AgentStatusEntry import is dropped because the branch no longer references it. Main's generalized agent-session resolution auto-merged unchanged. - use-mobile-session-terminal-create-actions.ts: kept the branch's hosted page-adapter prelude, then main's generalized bare-launch gate whole (isAgentSessionHandleProvider + createMobileStructuredAgentSession). - mobile-session-route-parity.test.ts: took the branch's pins and recomputed from the test's own printed values. Ablated first: use-mobile-session-terminal-create-actions.ts is the only changed file in MOBILE_SESSION_ROUTE_SOURCE_FILES, matching main's own ablation. Runtime strings 476 -> 475 (the dropped 'codex' literal) and the nested function body re-froze; JSX, style, identity, navigation and capability digests were all unaffected. The hybrid page path carries no agent-session surface at all, for Codex or Claude, so nothing in src/shared/mobile-web or mobile/src/mobile-web needed a wire change. Left as an open parity item. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
bd995fc9e8 |
refactor(mobile): build the shell init envelope outside the hybrid route
hybrid.tsx crossed the mobile max-lines budget when the shell-features list joined the init message. The envelope is the shell's grant and feature declaration, so it gets its own module. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
d07c47593d |
feat(mobile): structured native Claude chat (#18741)
* feat(mobile): structured native Claude chat Mobile already spoke the structured agent-session protocol for Codex, and the host already had a Claude capability gate — mobile just never advertised it, so `projectAgentSessionTabsOut` stripped every Claude tab before it left the desktop. The structured lane in mobile/ turned out to be agent-agnostic already (shared reducer, message projection, option catalog, prompt tokens), so this opens the gate rather than building a second lane: - advertise `agent-session.structured.claude.v1` - resolve any structured provider in `resolveMobileNativeChat` via the shared `isAgentSessionHandleProvider`, instead of a `'codex'` literal - widen the `agent-session` route type off `'codex'` - route bare Claude launches through `agentSession.createSupport` like Codex, which still degrades to a terminal when the host refuses (remote, WSL, win32, managed-account mismatch, or structured chat switched off) Deduplicate the create envelope. Renderer and mobile each assembled the `agentSession.create` params by hand; the fingerprint has to be computed over the same fields the host recomputes, so both now build it in one shared `structuredAgentSessionCreateParams`. Mobile's Codex-only launcher becomes `createMobileStructuredAgentSession(client, worktreeId, agent)` and reuses the shared display-name map; two copies of a random-UUID fallback collapse into one. Answer grouped Claude questions. A Claude AskUserQuestion carrying more than one question — or one multi-select question — is emitted with the real content in `body.questions` and the flat `options` left EMPTY, so mobile rendered a card with nothing to tap and the turn stalled with no way out. Codex never emits this shape. The phone has room for one question at a time, so the group is answered as steps and submitted once, reusing the shared `encodeAgentSessionQuestionAnswers` / `isValidAgentSessionQuestionAnswers` rather than a second encoding. Prompt responses move into `useMobileStructuredPromptResponses` because grouped questions carry a multi-step draft the rest of the session does not touch, and the session hook was at the 300-line cap. Pin the mobile capability list against the host's parser bounds: it fails closed to NO capabilities when the array exceeds 64 entries, which would look exactly like an old client. Re-pin mobile-session-route-parity: the create-actions edit drops one runtime string literal and changes one nested function body. Ablated to confirm that file is the sole cause. * fix(mobile): derive the grouped-question draft instead of clearing it in an effect The React Doctor gate flagged the session-change reset as a state adjustment after a prop change, which renders the stale draft for a frame. Store the session the answers were collected in alongside them and check it on read, so a session switch drops the draft during render with no effect at all. * test(mobile): pin that grouped steps key apart when the questions read identically Claude can ask the same text twice in one group (once per file, say). The view keys the question card by its projected content, so identical wording must still key apart or step 1's checkboxes would be submitted as step 2's answer. * fix(mobile): harden grouped Claude question answers * fix(mobile): retry transient structured support probes * fix(mobile): preserve grouped prompt response compatibility * fix(mobile): preserve tokenless duplicate choice identity * fix(mobile): point the launch tests at the generalized create API The rebase onto #18697 brought its definitive-refusal tests in cleanly, but they call the pre-rename createMobileStructuredCodexSession, and mobile tsc excludes test files so nothing caught it. Retarget them and give the agent-copy test a code that is actually in the definitive allowlist - agent_session_refused now correctly stays unknown, so it never reached the failure copy it asserted. * test(mobile): re-pin route parity after the rebase onto main Main moved its own runtime-string pin to 547; this branch drops the 'codex' literal from the create-actions gate. Ablated against main's pins to confirm that file is the sole cause before re-deriving. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
1735c2de9e |
Merge origin/main into mobile-rearch
Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
ab9b9c4831 |
Merge fix-review-0906: resolve the three open review threads
gzip chunk ceiling from deflateBound + stored-block fallback; followedByText behind a shell feature; baseline-mode guard pinned. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
2283f8ba4e |
docs(orchestration): never pick a worker model the user did not name (#19109)
The sonnet examples were added for a test cohort. Orchestration must not choose a model on the user's behalf: pass --model only when the user named one, otherwise inherit the configured agent default. |
||
|
|
0de9d7bacf |
test(mobile): pin the development-only guard on the native baseline flag
The production case asserted a value a release build returns anyway, so it stayed green with the developmentBuild guard removed. Assert it against a release build that opted into hybrid, where the flag is the only thing that could flip the result. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
978d1f1fef |
fix(mobile-web): gate the paste-images following-text hint behind a shell feature
Page->shell payloads are strict by design, so an optional field is not free: a page served by a newer desktop sending followedByText to an older APK gets invalid_request and the image paste fails outright. The shell now advertises its own understood behaviors in init as opaque strings, the page reads them off the bridge client, and it sends the hint only when the shell says it understands it. A shell that does not writes no trailing separator, which is what every shell did before the field existed. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
298571ad9f |
fix(codex): uncap app-server stdio records (#18590)
Co-authored-by: Merge Sim <sim@local> |
||
|
|
ebaa01e42c |
Recover branch compare on visibility change (#19021)
* Recover branch compare on visibility change Add recovery mode that reuses cached branch comparison data when the window regains focus instead of clearing results and forcing a refresh. This preserves the diff display during operations like rebasing that may cause the window to go to the background. * Retry failed branch comparison results Cached branch comparison results with error status are now excluded from the cache-hit check, ensuring they are retried rather than silently reused. This fixes missing diffs during rebasing. * Decouple branch compare recovery from refresh kinds Recovery is now a dedicated callback invoked independently on visibility changes, rather than a refresh kind. This allows pending recoveries to queue during in-flight requests, improving handling when the window regains focus during rebasing or other operations. |
||
|
|
6654e83588 |
fix(mobile-web): bound the package gzip chunk ceiling by zlib's real worst case
A full-range read of an incompressible asset (PNG, woff2, wasm) gzips larger than its source, so the flat MAX_RANGE_BYTES + 64 ceiling failed the host's own response schema and aborted the whole package download. Derive the ceiling from zlib's deflateBound instead, and have the host fall back to stored blocks when level 6 does not shrink the range, so it never emits a stream larger than it has to. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
b0a8ff3684 |
Merge origin/main (orchestration v3 + relay endpoint credential) into merge-main-0906
origin/main advanced two commits mid-merge. One conflict: ssh-relay-session-terminal-error.test.ts, where both sides appended a test at the same position. Kept both. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
b8b8c62f53 |
Merge origin/main into merge-main-0906
Resolves eight conflicts from main's SSH e2e lane restructure and the mobile relay stream-cancellation fix. - e2e.yml / run-ssh-docker-e2e.mjs / pr-e2e-gate-contract.test.mjs: keep main's lane structure and re-express only the hosted-mobile-webview SSH exclusion. - reliability-gates.jsonc: main's file plus the branch's mobile-hybrid gate. - mobile-relay-rpc-streams.ts: main's cancellation machinery replaces the branch's equivalent, generalized to every server-assigned-id method. - docker-ssh-relay-connection.ts: main's delegation to connectSshTestTarget, with the branch's connect timeout moved into that shared helper. - mobile-session-route-parity.test.ts: digest re-frozen for main's #12772. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
0c33f58e8a |
fix(ssh-relay): daemon owns the endpoint credential; a losing start never rotates it (#19052)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 19 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$962 | $\color{#cf222e}{\Huge{\mathbf{−}}}$136 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$826 |
| Prod | 18 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$295 | $\color{#cf222e}{\Huge{\mathbf{−}}}$116 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$179 |
<!-- /orca-pr-loc -->
## Symptom
Live 2026-09-05 (Orca 1.4.198 client, Ubuntu host): both relay processes `kill -STOP`ped for 20 s, then `-CONT`. The client redeployed while the host was frozen. Its fresh daemon lost the socket bind (`Socket path already in use`) but had **already rewritten** `relay-<id>.sock.credential`. The surviving daemon kept its in-memory credential, so every later `--connect` got `Endpoint credential mismatch; closing socket`, then `Grace started … timeoutMs=0 … ptys=1, clients=0` every ~20 s, forever. Only a manual `kill -TERM` cleared it. Receipts: `review-archive/orchestration-v3-pr16904/smoke-receipts-t012b/E16,E17,E18,E24`.
Three independent defects kept the wedge alive; each is fixed at its own seam.
## Fix
**1. The relay daemon owns credential publication (race-free under two concurrent starters).**
`relay-daemon.ts` binds the socket first, then publishes via the new `src/relay/relay-endpoint-credential-publication.ts`: adopt a valid pre-existing file (older clients still pre-write), else mint 32 random bytes and write temp+rename at 0600. A start that loses the bind exits inside `listen()` and never reaches the file. Why this option and not restore-on-loss or a client-side write: the only process that can *prove* ownership is the one whose `listen()` succeeded, and that proof is atomic with the bind. The client-side pre-write (`ssh-relay-endpoint-credential.ts`) and the launch-command `chmod 600`/`icacls` are removed on POSIX and Windows. The racing test also exposed that macOS reports a mid-bind collision as `EEXIST` rather than `EADDRINUSE`; `relay-socket-ownership.ts` now treats both as "held or stale".
**2. The client distinguishes "no daemon" from "daemon present but not answering", and never rewrites.**
A credential refusal is now typed on the wire: the daemon replies `orca-relay-handshake-credential-mismatch` (same frame type, no new opcode) and the bridge exits **43**; `waitForSentinel` maps it to `RelayCredentialMismatchError`, which the takeover treats as handshake-refusal evidence exactly like exit 42. A relay that holds the endpoint but **never refused** (the stalled-host shape: kernel backlog accepts the probe, handshake gets no answer) is now `RelayEndpointUnresponsiveError`, routed to the relay-lost backoff instead of the terminal Reset Relay path. Silence is not a decision (`docs/reference/ssh-execution-boundary.md`).
**2b. Deploy honours the verdict.** The 40 s live run exposed that the `--connect` catch block in `deployAndLaunchRelay` predates the incumbent probe and swallowed both verdicts as "probe failed, launch fresh", so a fresh daemon was still launched over the live one (it lost the bind by luck, which is exactly the collision in the incident). Held and Unresponsive now propagate; the session backs off on Unresponsive and surfaces Reset Relay on Held. Red-first in `ssh-relay-deploy-incumbent-verdict.test.ts`.
**3. The daemon cannot be wedged by a rotated file, because nothing can rotate it.**
The credential lives in the content-hashed relay dir, and after (1) the only writer is the daemon that owns the socket, so the "file changed under a live daemon" state the incident depended on is no longer reachable in-product. The credential is therefore fixed for the daemon's lifetime, as a plain secret should be. A hand-edited file is refused with the typed reply until restored (tested). Startup adoption of a pre-written file applies an owner-only + same-uid rule (review finding): anything else is replaced by a fresh mint. An earlier revision of this PR also re-read the file on mismatch and adopted it; that was removed as unreachable machinery that turned the credential into a per-handshake file-ownership check.
**3b. Fail closed between bind and publication.** A client that arrives after `listen()` resolves but before the credential is set is refused, not admitted as `unproved`. Nothing can be delivered in that window today; the guard makes the boundary structural instead of an event-loop ordering fact. Red-first in `relay-reconnect-listener-credential-gate.test.ts`.
**Wire compat.** New optional handshake reply only; an old `--connect` hits `Unknown handshake type` and exits 1 pre-sentinel, which it already treated as a generic failure. New daemon adopts an old client's pre-written file; new client still passes `--credential-file` so an old daemon reads it as before. Absence of exit 43 is never used as evidence.
**Also.** `terminal create` on a reconnecting SSH host now says what to do instead of a bare `No PTY provider for connection "<id>"` (prefix preserved; the renderer matches it).
## Tests (red first)
- `src/relay/subprocess.test.ts`: two `--detached` starts race one socket + credential file → exactly one reaches the sentinel, loser exits 1 with `Socket path already in use`, file valid + 0600, a `--connect` reading it reaches `relay.status` and reports the winner's pid. Red before (both starters died: daemon required a pre-existing file), green 6/6 after.
- `src/relay/relay-endpoint-credential-publication.test.ts`: mints after bind; adopts a pre-written 0600 file; replaces a pre-written 0644 file with a fresh mint; refuses a stale credential with exit 43 while still serving the real one, and keeps refusing a rewritten file until it is restored.
- `src/relay/relay-reconnect-listener-credential-gate.test.ts`: a client in the bind-to-publish window is refused and never attached; after publication the right credential is accepted and a wrong one refused; a daemon launched without a credential file is not gated. Red without the guard.
- `ssh-relay-deploy-incumbent-verdict.test.ts`: live-but-silent incumbent → `RelayEndpointUnresponsiveError`, refused → `RelayEndpointHeldError`, and in neither case is `--detached` launched; a failed `test -S` probe still launches fresh. Red 2/3 without the deploy change.
- `ssh-relay-deploy-helpers.test.ts` (exit 43), `ssh-relay-endpoint-takeover.test.ts` (refused → Held even with no `lsof`; silent → Unresponsive, nothing unlinked or signalled), `ssh-relay-session-terminal-error.test.ts` (Unresponsive → `onRelayLost`, not terminal). Deploy/namespace/native-deps tests updated to assert the client writes **no** credential.
## Live proof
New `tests/e2e/ssh-docker-relay-stall-credential.spec.ts` (claimed in `run-ssh-docker-e2e.mjs` and PR source routing), two cases: `kill -STOP` every relay pid in the container, send input during the freeze, hold **20 s** (the incident's duration, which races the mux liveness timeout) or **40 s** (past it for sure), `kill -CONT`; assert status back to `connected`, same pty, same daemon pid, same credential inode and content, relay.log did not shrink (a relaunch truncates it) and has zero `Endpoint credential mismatch` / `Socket path already in use` lines, in-stall input delivered at most once.
Run output (local, fixture image `orca-e2e-ssh-relay:3a864c665ba2cefd`, `ORCA_E2E_SSH_DOCKER=1 SKIP_BUILD=1 ORCA_E2E_FORWARD_APP_LOGS=1 … --project electron-headless --workers=1`, head `c2c20fd994`; re-run identically on the final head after the credential-lifetime change, 2 passed (1.7m), same annotations, and the bind-to-publish refusal never fired):
```
✓ keeps the same daemon and credential across a 20s relay freeze (38.3s)
relay-processes-stopped: 2 relay-processes-continued: 2
bridge-pids-before-after: 480 -> 480
socket-clients-accepted-before-after: 1 -> 1
in-stall-input-delivered: 1
✓ backs off and reattaches, never relaunching, across a 40s relay freeze (57.5s)
relay-processes-stopped: 2 relay-processes-continued: 4
bridge-pids-before-after: 480 -> 1202
socket-clients-accepted-before-after: 1 -> 3
in-stall-input-delivered: 1
2 passed (1.6m)
```
Client log in the 40 s case shows the new path end to end: `Relay channel lost … reconnect attempt 1/6` → `Socket probe result: "ALIVE"` → `Socket reconnect failed … Relay failed to start within 10s` → `Relay endpoint incumbent: … verdict=live evidence=accepted-connection holders=unenumerable` → `Failed to re-establish relay … A relay still owns … but did not answer the handshake … Orca will retry` → `reconnect attempt 2/6` → `Reconnected to existing relay via socket`. The 20 s case never left the frozen bridge (same bridge pid, one accept), so it exercises the "silence is not death" side of the same race. The 20 s case passed 6/6 across the session; the 40 s case was red on the prior head (`Socket path already in use` + `Startup failed: listen EADDRINUSE` in relay.log from the swallowed verdict) and is green after 2b. Before the fix the same injection produced a fresh daemon that rewrote the credential and a survivor refusing every client.
The `relay-processes-continued` count exceeds `stopped` in the 40 s case because the timed-out client's `--connect` bridge and the loser-side processes are parked behind the frozen listener when `CONT` runs; they exit on their own once it resumes.
## Gates
`pnpm test src/relay src/main/ssh` 332 files / 3884 tests pass · `pnpm typecheck:tsc:node` clean · `check:code-quality:changed` 0 findings · `check:react-doctor:changed` 0 findings · `pr-e2e-gate-contract.test.mjs` 42 pass · no lint disables or max-lines bumps added.
## Noted, not fixed here
- `terminal list` `orphaned:false` / `terminal close` `ptyKilled:true` for a pane whose relay is gone (`orca-runtime-stop-explicitly-closed-tab-ptys.ts`): different seam, `@ts-nocheck` characterization-covered file.
- On a host with no `lsof`, a stalled relay still cannot be enumerated as the holder; it is now retried rather than declared held, but a relay frozen past the backoff budget still ends in the existing "reconnect manually" banner.
|
||
|
|
06a607a1d7 |
feat(orchestration): make multi-agent workflows durable (#16904)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 225 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$21666 | $\color{#cf222e}{\Huge{\mathbf{−}}}$2820 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$18846 |
| Prod | 348 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$17107 | $\color{#cf222e}{\Huge{\mathbf{−}}}$4706 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$12401 |
<!-- /orca-pr-loc -->
## ELI5
Orca now treats orchestration like a durable control plane instead of inferring success from terminal keystrokes. Agents can tell whether a prompt was accepted or a turn started, replay an ambiguous request without sending twice, and recover coordinator mail after a crash. Completed workers can be inspected, released, or retained, and their panes no longer auto-resume as if the work were still running.
## What changed
- **Run receipts** from `run-create/use/current/show/list` are the row without routing plumbing (`home_database`, `coordinator_pane_key`) and without the duplicate `binding` object.
- **`terminal send` receipts are honest and idempotent.** `input_accepted` and `turn_started` are the only stages; `--wait-submit` observes without resending; `--retry-request <uuid>` replays the exact request against the same process incarnation. A transport timeout keeps the retry ID; only a different runtime answering strips it. Value-less or non-UUID `--retry-request` is rejected on the CLI and the SSH shim.
- **Mailbox delivery is committed before wakeup.** Pointer writes are staged in the DB before any PTY byte, replayed once after restart, and never emit a naked Enter. The watermark that parks concurrent deliveries is released with the DB reservation. Restart rescans pointer-pending and `dispatch:` mailboxes.
- **Lifecycle is a guarded transition graph** (`lifecycle-transition.ts`) with a table-driven test over every caller edge. Task reopen/overturn stays in the public contract. A PTY exit during `worker-stop` is the stop succeeding, not a failure.
- **Worker lifecycle CLI:** `worker-start` (`--spec` creates Task + attempt in one call), `worker-show`, `worker-read` (provider transcript first, bounded terminal fallback with a typed reason, local/WSL/SSH), `worker-stop`, `worker-abandon`, `worker-release`, `worker-retain`, `worker-list` (rowid-fenced pagination, fleet liveness, `attention`, literal `nextAction`).
- **Release is an explicit ownership table** (`decideWorkerTerminalRelease`): only an `owned` resource can be settled, the archive is mandatory where reachable, and an owner whose process is proven exited can always get out of `retained` via `archive_status: unavailable`. User-taken-over, external, and transferred panes stay retained.
- **Settled-worker resume fence** (folds in #17651): a settled dispatch whose pane is still open is fenced at settlement, on stop/abandon/exit, and at startup; lifted on release, retain, takeover, and pane reuse.
- **Liveness is `live` / `unverifiable` / `exited` only**, from execution-host evidence. Fleet projection reads the evidence clock, not the relay delivery clock. A host-certified exit outranks the worker's settled state. `unverifiable` never authorizes stop, abandon, retry, or release, in code or in the guide.
- **Federation:** structured reads negotiate by `method_not_found` so every shipped host keeps transcript-first output; exited remote workers are closed before being reported closed; epoch fencing holds across peer restart, downgrade, and pairing rotation; no per-second forced capability probe.
- **Schema v35:** repairs databases stamped v34 by the pre-fix branch (mailbox_handle default, index predicates), drops the write-only `lifecycle_transition_receipts` ledger and five never-read v31 identity columns.
- **Schema v36:** `dispatch:<id>` mailboxes get a real consumer generation on `dispatch_contexts` and `remote_dispatch_attachments`, bumped and fenced in the same transaction on every re-attach (manual inject, worker-start, federated attach). A stale worker whose Dispatch moved to another process now gets `consumer_fenced` instead of silently acking the new worker's Delivery. Run mailboxes already worked this way.
- **Schema v37:** `dispatch_contexts` records its creator (`creator_handle`, `creator_pane_key`), so a coordinator's context-only self-dispatch is bookkeeping rather than a nesting parent; before this, one self-dispatch made every later `worker-start` from that coordinator fail the depth cap. Pre-v37 rows keep counting (fails closed).
- **Dispatch-mailbox ownership is checked, not inferred.** A `check` from a process whose pane no longer holds the Dispatch, or whose last Attempt was abandoned/failed and moved to another terminal, gets `consumer_fenced` instead of an empty inbox that reads as "no mail yet". `--peek`/`--all` stay readable. A paneless caller still gets `stable_pane_required` with the rebind recovery.
- **Liveness certification is stricter:** a `process_exited` stage whose termination reason is `unknown` (a stop that was issued but never observed) projects `unverifiable`, not `exited`. Federated `worker-show` carries the execution host's verdict and host kind instead of a local guess. A live, ready worker with nothing pending has `nextAction: none` rather than pointing at the `worker-show` that produced it.
- **Wire:** `workerShow` keeps `dispatch.task_id` next to `taskId` for shipped CLIs. `ask --json` uses the standard `{ok, result}` envelope like every sibling verb.
- **Migration start-version detection** treats the two v32 recovery columns as versioned. Before this, every shipped database stamped below 32 resolved to the v6 floor and replayed the whole chain (the v23 backfill synthesized 68 phantom retained workers on a real v30 profile). Verified on a copy of a real 62 MB v30 profile: starts at 30, no row delta, integrity ok, 11 ms.
- **Skill guide** rewritten as a ≤200-line kernel plus seven references, to the outcome-first standard (Result / Done / Safe failure first, conditions not case lists, one done bar, references loaded at the point of use). The canonical loop uses `worker-start --spec`, names `worker-list` for completion accounting, documents `--retry-request` / `request-show` / `--wait-submit`, and requires positive evidence before any stall action. The other seven guides get the same treatment in #18724, split out so this PR stays orchestration-only.
- **`rpc/methods/orchestration-*`** (126 flat files) regrouped into `orchestration/{worker,federation,messaging,runs,gates}/`.
## Why
User reports showed the same boundary failures: false `agent_prompt_stalled` causing duplicate sends (#15180), coordinators unable to trust screen scrapes, cold-parked terminals receiving a pointer without the submit, settled workers accumulating as live tabs and auto-resuming after restart, and no way to tell a stalled worker from a working one.
## Linked issues
Fixes #15180. Fixes #17935 (orchestration skill description is 866 characters; a guard now caps every bundled skill at 1,024). Supersedes #17651 (fence folded in). Advances #16660, #16522, #14907, #13047.
## Review record
This PR was reviewed adversarially after revival: eight independent lenses (lifecycle, mailbox, send, worker, federation, transcript, complexity, live ergonomics), each required to prove findings with a failing test. That produced 16 proven blockers, all fixed with red-then-green regression tests, followed by two re-review rounds and a third fix wave that caught 3 regressions introduced by the fixes and 7 fixes that missed their target; all closed. A final pass (five lenses incl. a live built-runtime smoke, then a re-review of the fix wave) found and fixed seven more, chiefly the stale-worker mailbox steal, the self-dispatch depth wedge, and the unproven-exit certification. Three independent Codex (gpt-6-astra) passes followed: the first found nothing new, the second found and fixed 3 defects (task-status reachability, WSL-local host classification, peer-capability epoch), the third found and fixed 6 (production PTY controller never installed settled writes, ambiguous in-flight pointer failures allowed duplicate replay, SSH/relay deadlines cut off a valid `--wait-submit`, stop-vs-exit race during inspection, and two release-recovery paths for vanished or exited terminals). The full record (findings, proof tests, triage, declines with reasons) is archived outside the repo.
**Rework after the live smoke.** A first live cross-host run on the shipped adhoc build (this Mac, a paired Windows host on the same build, a paired Mac on 1.4.195, and an SSH host) found a P1: a running local worker read `unverifiable`/`missing_status` because the fleet snapshot rows lacked the terminal handle the matcher keyed on. A 59-row failure table over every bug fixed during review showed the same two classes recurring: a fact dropped in transit through optional fields, and two authorities for one fact. Two blind designs (Opus, Codex) converged on the same mechanisms, and the scoped tranches landed here with red-then-green seam tests from the real producer to the real consumer, faults injected only at the transport or hook-ingest boundary:
- **Settlement (data-loss class):** one three-valued `WriteSettlement` (`accepted | refused{reason} | unverifiable{reason, bytesHandedToTransport}`) from the SSH multiplexer through daemon client, providers, controller, to pointer staging. No boolean, no rejection-as-third-state. The two silent degrades that fabricated a handoff are deleted; a provider that cannot settle refuses before any effect. Pointer text and Enter share the contract; a partial flush is `unverifiable`, never `refused`.
- **Evidence identity (false-liveness class):** fleet agent-status evidence is a tagged union (`binding: worker | pane | unresolved{reason}`, `clock: observed | delivery`) minted once at ingest, so a hook row captured on one process incarnation can never bind to a later dispatch on the same pane. The matcher's `!worker.paneKey ||` defaults are gone. One host-scope parser replaces two.
- **Small pre-merge items:** `capability_unsupported` from an old peer is no longer relabelled `host_unavailable`; a producer census test asserts every agent-status consumer path projects a pane-only hook row as `live`.
Two ergonomics defects the second live run surfaced on a real database are fixed here too: a pre-v3 dispatch already marked `completed` projected as `outcome_unknown` / `requiresAction: true` forever (three copies of the outcome ladder disagreed on legacy rows; now one resolver, legacy `completed` reads `succeeded` with nothing to act on, legacy `failed` stays actionable on the failure), and an unscoped `worker-list` enumerated the entire database (now defaults to the Run bound to the calling terminal, `--run` overrides, and the receipt's additive `scope` field says which).
A third live round on the shipped adhoc build of `b082443e1f` (same four hosts) plus an unscripted run in the user's own prompt style (a plain Claude Code shell, `/orchestration`, three workers, zero errors, bound-Run default confirmed) found two more branch defects, fixed with red-then-green tests: a worker freshly started on a paired server projected `unverifiable`/`host_indeterminate` with `requiresAction` for ~3 minutes, including after its own `worker_done`, because the host's federation observation returned `missing_liveness_verdict` for any PTY the liveness register had not yet swept (the host now reads a connected pane it owns locally as `live`; disconnected or SSH-scoped panes stay `unverifiable`); and six pre-v3 completed rows still carried an `input` category because settling through the task-status path or `failDispatch` never closed the Dispatch's pending question threads (both paths close them now, and schema v38 closes threads already pending on settled rows). The guide's `worker-start` examples now show `--model sonnet`, since an omitted model inherits the launcher's default.
A Codex adversarial pass on the tranche diff found one real design hole (identity minted at read time instead of ingest, now closed) and two daemon settlement paths that threw instead of settling (fixed). Two `@ts-nocheck` runtime mixins on these paths were extracted into checked modules; the repo-wide `@ts-nocheck` count is unchanged at 171.
Deletions during review: ~1,900 lines (write-only ledger, unread columns, dead v1 archive path, test harnesses shipped in prod, duplicated liveness and state-machine copies, self-capability checks that were compile-time true).
## Testing
- `pnpm typecheck:tsc:node|cli|web` clean
- `pnpm run check:code-quality:changed` 0 findings; `check:react-doctor:changed` 0
- `pnpm verify:bundled-skill-guides`, `verify:skill-bundle-manifest`
- full `pnpm test` on the integrated head: 72,332 pass / 292 skipped; the only failures were three non-PR files (two zsh live-shell suites hit a node-pty spawn-helper ENOENT while a concurrent native rebuild ran, 44/44 in isolation; `release-checkout.unit.test.ts` is a known 30 s load timeout that passes in isolation on `origin/main` too).
- CI on
|
||
|
|
7ac194a634 |
Add persistent turn-scoped chat activity indicator (#19044)
* feat(chat): show turn-scoped activity tail * fix(chat): keep turn activity broad --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
c36c23df5f |
fix(chat): stop terminal focus recovery from stealing Cmd+C in the Chat UI (#18751)
* fix(chat): preserve message copy focus * fix(chat): scope covered-xterm focus guard to the chat leaf Chat view mode is a tab flag, but only the chat leaf's xterm is covered. In a split chat tab with a terminal leaf active, the tab-level guard skipped the terminal's resume focus and the tab-wide deferred focus then landed on the covered chat xterm. Decide per pane: resume and window-wake read the active pane's container, and the surface focus query skips leaves hosting the chat root. * fix(chat): close covered terminal focus fallbacks --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
3be526c5e6 |
test: cover SSH reattach replay and enable deterministic Codex CI (#19106)
* test: cover SSH replay replies and run deterministic Codex restore scenarios * test: register replay probe unit command in reliability gate |
||
|
|
6aa0aaee6b |
test: isolate source-control generation repositories per scenario (#19105)
* test: isolate source control generation repositories per scenario * test: explain scenario repository fixture scope |
||
|
|
4cccadcb95 | test: keep Activity pane selection in the retained sidebar (#19107) | ||
|
|
b44aaf20c6 |
test: reuse authoritative SSH connection readiness in localhost fixture (#19102)
* test: reuse authoritative SSH connection readiness in localhost fixture * test: retain localhost SSH setup diagnostics |
||
|
|
da48ad2b47 |
Bump mobile Android versionCode to 16 to match the 0.0.48 release (#19101)
Co-authored-by: Merge Sim <sim@local> |
||
|
|
4d9e963ffd |
test: enable localhost SSH terminal and hook journey in CI (#19097)
* test: run localhost SSH terminal and hooks in CI * test: isolate localhost SSH session fixtures across repetitions * test: route remote agent hook source changes to localhost journey * test: record localhost SSH reliability evidence and remaining gaps * test: route the real SSH session hook authority |
||
|
|
b459b8f16d |
test: repair nested SSH fixture after HUB restart (#19098)
* test: restore paired nested SSH fixture after HUB restart * test: cover failed re-pair selection and background window safety * test: use required braces in re-pair regression fixture * test: use current paired runtime identity after re-pairing |
||
|
|
5dab495655 |
docs(relay): record Roll 2 cell roll (4916ed67 fleet-wide) and tick checklist (#19096)
All 19 general cells on 4916ed67, selector gen 148 -> 186, 0 serving-process exits across the roll. Three waves used the no-restart mode=rollback resume (c13 transient trust-probe 409; c26/c21 post-apply runtime-status 503 shed). Checklist: 2.3, 4.1, 4.3 relay side deployed; status header 2026-09-06. |