mirror of
https://github.com/stablyai/orca.git
synced 2026-09-29 16:02:50 +00:00
f169cd1cdeb208dab83b385fbaf06ae46af882bd
10323
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f169cd1cde |
fix(pty): require a settled writer on every PTY provider
The production controller never installed `writeWithSettlement`, so every mailbox
pointer write took the fire-and-forget fallback and provider settlement never
controlled the pointer transition (R3-C2). Making the method required is what
turns that into a compile error, so `IPtyProvider.writeWithSettlement` is now
non-optional and `LocalPtyProvider` implements it synchronously: in-process
node-pty is its own sole owner, so its answer is the settlement.
Required-ness alone is satisfiable by a lie. `degraded-daemon-pty-provider` used
to answer it with `provider.write(...) !== false`, which reproduces the bug
through the fix, so its conditional is deleted outright and it delegates.
`RuntimePtyController.writeWithSettlement` stays optional: 809 `setPtyController`
call sites across 164 files construct partial controllers, and requiring it there
would be a mechanical edit of all of them for no added fence. Instead the
orchestration pointer path demands the narrow settled-writer shape and returns
`refused{provider_cannot_settle}` before any effect when it is absent.
Two gates, both red against the fabricated handoff:
- `registerPtyHandlers` with a real controller and a provider lacking settlement
refuses before any byte reaches `provider.write`.
- A census pins the five production `IPtyProvider` classes, proves each instance
exposes the writer, and reads every declaring body to reject a settlement
synthesized from `.write(`.
|
||
|
|
825288f62b |
refactor(orchestration): typecheck the fleet agent-status snapshot
The accessor that mints fleet evidence from hook rows lived in a `@ts-nocheck` runtime mixin, so nothing on it was checked — including the hop that published pane identity into a matcher comparing terminal identity. The body moves to a checked module and the mixin keeps a one-line delegate. Census updated for the new producer file and for the minter token. |
||
|
|
97a5ecc1a5 |
fix(orchestration): bind agent-status evidence on process incarnation
A hook row carries no incarnation, so a row replayed after a runtime restart was indistinguishable from a current one by pane key and terminal handle alone, and the reminted-pane arm accepted a dispatch-labelled row on the mere PRESENCE of a durable `process_incarnation`. Both false binds report a dead process as `live`. The pane's incarnation at mint time now has to equal the incarnation the durable resource named. A worker with no materialized resource keeps binding on pane and handle: absence there is not a contradicting authority, and fencing it out would report a running unsupervised worker as missing. |
||
|
|
f9aaade3b8 |
refactor(orchestration): mint fleet agent-status evidence once
Terminal identity and the observation clock were optional fields on the row the fleet matcher reads, so an unenriched producer published a pane key the matcher silently failed to identify and a missing observation clock silently degraded to the delivery clock. Both are now arms of a union minted at one boundary. - `FleetAgentStatusEvidence` carries `binding` (worker | pane | unresolved), `clock` (observed | delivery) and the activity facts, and exposes no optional identity, no optional clock and no raw payload. - `resolveAgentStatusBinding` is the single identity lookup; both the IPC payload and the fleet evidence derive from it. `agentStatus:set` and `agentStatus:getSnapshot` publish the same fields as before. - `createFleetStatusIndex` / `statusForFleetWorker` consume the union. The `!worker.paneKey ||` and `!worker.agentTerminalHandle ||` defaults are replaced by `FleetWorkerIdentity`, so a worker that names no handle matches nothing. - `keepFreshest` orders on the named `deliveredAt`, never on the staleness clock. |
||
|
|
a6290085ad |
fix(pty): three-valued write settlement from transport to mailbox pointer
An SSH write whose settlement was lost mid-flight was indistinguishable from a
pre-write refusal, so the durable phase-2 mailbox reservation was cleared and a
reconnect wrote the pointer bytes a second time (R3-C1). The fact was present at
the multiplexer (`writeAttempted`) and flattened to one bit at three consecutive
hops: `writeToSshPtyWithSettlement` resolved `false` for both cases, the pointer
accessor swallowed rejections, and `mailbox-pointer-stage` cleared on `false`.
`WriteSettlement` in `src/shared/pty-write-settlement.ts` is now the single type
on that path: `accepted | refused{reason} | unverifiable{reason,
bytesHandedToTransport}`. Ambiguity is a value, never a rejection and never a
boolean. The transport writer is the only boundary that knows whether an entry
was still queued or already handed over, so it mints `refused` versus
`unverifiable`; `MultiplexerTransportWriteResult` keeps the two-valued shape the
socket can actually prove.
Deleted: the rejection channel in `writeToSshPtyWithSettlement` and
`writeNotifyWithSettlement`, `MultiplexerWriteSettlement.writeAttempted`, the
`.then(finish, preserveAmbiguous)` convention, the `?.`/`??` fallback chain in
`writeOrchestrationPointerPty`, and both fabricated handoffs
(`degraded-daemon-pty-provider` and `ipc/pty/runtime/operations`) that answered
`provider.write(...) !== false` for a settled write. A provider that cannot
settle now returns `refused{provider_cannot_settle}` before any effect.
Test doubles adapt through one shared `settledWriteStub`; no contract was
loosened to keep a boolean fake compiling.
|
||
|
|
ff53bbad55 | test(orchestration): census every agent-status producer path | ||
|
|
023ae45872 | fix(orchestration): keep capability_unsupported distinct from host_unavailable | ||
|
|
71d2bc6d4b |
docs(orchestration): scope worker-list recipes and rank liveness evidence
Every worker-list recipe now passes --run <run_id>. Unscoped, a real profile returned 2103 rows and none of the live workers, so the stall recipe the kernel prescribes could not find them. Reword the disagreement rule: the fleet verdict still decides, except against an unverifiable reason that names a client-side gap (missing_status, host_unavailable), where a worker-show verdict from the execution host is the better evidence. Note that a peer lacking the fleet-snapshot capability also reports host_unavailable on the row while the host warning names capability_unsupported. unverifiable still authorizes nothing. |
||
|
|
2bd094c103 |
fix(orchestration): let a failed start release the terminal it created
A start that died before prepareStartingWorkerAuthority recorded no ownership row, so worker-release answered no_owned_resource, worker-stop reported already settled, and the PTY it created had no owner that could ever close it. Record the ownership the successful path would have recorded, gated on provable pane and process identity and on no other live resource accounting for that process. worker-release then proceeds through the existing table, worker-list reports the row reclaimable, and the failure receipt names the command. Extract the parseTaskDeps copy the local and federated start paths duplicated. |
||
|
|
969b2ef40c |
fix(orchestration): restore terminal identity on fleet status rows
The hook server publishes status rows carrying only a pane key, but the fleet matcher compares terminal handle and dispatch id. Every local worker therefore failed identity and projected unverifiable/missing_status while running. Apply the same enrichment the renderer IPC boundary already applies, so the shared matcher sees the identity it was written against. |
||
|
|
3979c5a159 | fix(ssh): preserve terminal prompt observation deadlines | ||
|
|
f06b1a3f3b | fix(orchestration): recover released and exited terminal leases | ||
|
|
26ff7ea00a | fix(orchestration): credit worker exit during stop inspection | ||
|
|
53a4209334 | fix(orchestration): preserve uncertain pointer writes through production transports | ||
|
|
630bc6642c |
feat(cli): add skills get --reference and --references selectors
An action gate names one reference, but --full was the only way to reach it and returned the kernel plus every reference. The selector serves one document, and the orchestration guide now teaches it with --full as the older-CLI fallback. |
||
|
|
4748b24a03 |
feat(skills): embed each bundled guide reference under its own name
The generated guide table carried only the concatenated fullMarkdown, so a per-reference CLI selector had no addressable unit to serve. |
||
|
|
c02eb193cc |
Merge remote-tracking branch 'origin/main' into orchestration-v3
# Conflicts: # src/main/runtime/orchestration/__snapshots__/preamble.test.ts.snap # src/main/runtime/orchestration/preamble.test.ts # src/main/runtime/orchestration/preamble.ts |
||
|
|
436ef827dd |
fix(browser): present Electron's own user agent so Cloudflare Turnstile clears (#18749)
Orca rewrote every browser session's UA to look like plain Chrome by stripping the Electron and app tokens. That rewrite is what Cloudflare rejects: a Chrome UA that ships no client hints reads as a spoof and Turnstile returns 600010, while the same binary on the same IP clears every challenge with its stock UA. PR #885 added the rewrite to fix 600010 and was treating a symptom it created; issue #11518 later found the same rewrite is what broke Google sign-in. - Keep the stock Electron UA on every partition. The webRequest handler now only owns the host-scoped Google auth Firefox switch, which stays unchanged. - Delete the anti-detection script. Measured on Electron 43: plugins are already a real PluginArray, window.chrome exists, and navigator.webdriver is false even with the debugger attached, so three of its four premises were wrong, and the overrides it installed (instance-level webdriver, non-native Permissions.query, stubbed chrome.csi/loadTimes) are themselves published bot signatures. - Stop attaching a CDP debugger to every browsing guest. Only the auth-UA detach listener remains, because a detach clears Chromium's standing UA override. - Stop sending Runtime.enable into cross-origin iframes when the agent bridge auto-attaches. The challenge widget is one, nothing reads iframe Runtime events, and the Runtime domain's serialization side effect is the documented Cloudflare CDP tell. - Add a real-Electron test proving the wire identity: stock UA to ordinary hosts, Firefox with no client hints to accounts.google.com. Verified in the dev build: dash.cloudflare.com/login no longer shows "There was a problem with verification" and scrapingcourse.com's managed challenge clears, both failing deterministically before. Fixes #13822 |
||
|
|
040c3e5b32 |
fix(browser): match loading surfaces to the Orca theme (#18738)
* fix(browser): theme unavailable guest surfaces without recoloring pages * test(browser): keep generated loading evidence out of the PR diff * test(browser): freeze recovery clock during artificial attach gate * test(browser): await painted content after network recovery |
||
|
|
974acc901c | fix(relay-ops): retry freshness-only preflight failures on the first same-cap wave too (#18778) | ||
|
|
cb7f7dd11a |
fix(native-chat): tell old mobile builds why a structured chat is missing (#18756)
* fix(native-chat): tell old mobile builds why a structured chat is missing A structured native chat started on desktop was simply absent on a paired phone running any shipped App Store build. The host strips every `agent-session` tab from a client that does not advertise `agent-session.structured.v1`, and no released mobile build advertises it — so the chat had no representation at all and no way to explain itself. Keep the row and retitle it instead of deleting it. The shipped client does not filter unknown tab types and renders whatever title the host sends, so an old build now shows the chat's slot with a title naming the fix. Nothing is removed, so the tab order, groups and layout it belonged to are left intact. The prompt is keyed on the capability for that specific agent, not on the combined policy boolean: a capable phone whose desktop simply has the experiment off would otherwise be told to take an update that cannot help it. Claude rows are prompted too — mobile cannot render them yet and a later build can, so the message is true for that client as well. Restore is no longer gated on the caller's capability. It stayed gated on the host setting, which is what decides whether there is anything to reach at all, but gating on capability left an old client with nothing to project after a desktop restart: neither the chat nor the prompt. Tab titles are capped at 128px on one line in every shipped build, so the string is sized for ~15 characters rather than a sentence. Prompted rows are visible rows, so the host now permits all five session-tab mutations on them, close included. That is intended: a mobile close runs the same teardown as the desktop's own Close button. * fix(native-chat): keep fallback tabs safe and truthful --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
5a1acfec17 |
fix(native-chat): make document paths and links clickable in chat (#18712)
* fix(chat): make assistant file paths clickable * fix(chat): tighten native file link handling * fix(chat): link prose-joined relative paths separately * fix(chat): preserve complete Unicode file links * fix(chat): preserve links before sentence punctuation * test(chat): align structured session link props * fix(native-chat): harden generated file links * fix(native-chat): reject reference-number false positives * test(native-chat): align structured session parity * perf(native-chat): bound file-link detection --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
e2b70a5eba | fix(relay-ops): retry transient admin-endpoint failures in same-cap verify and rehome jobs (#18769) | ||
|
|
df7b1028dd |
fix: use fallback shell readiness and assert fish startup timing (#18755)
* fix: classify effective startup shell and pin fish delivery timing * test: clarify fallback fixture and remove unused readiness wrapper |
||
|
|
30d7542bc5 |
fix(terminal): stop a hidden pane's unmeasured 80x24 from overwriting a live PTY's size on reattach (#18706)
* fix(terminal): stop a hidden pane's unmeasured 80x24 from overwriting a live PTY's size on reattach A pane that mounts while display:none (app relaunch or update with the floating terminal panel closed, a non-active floating tab, any background tab) cannot fit its container, so it reattaches with xterm's default 80x24. Main wrote those placeholder dims into `ptySizes` unconditionally, both before and after the attach. A daemon attach never resizes the live session, so the real PTY stayed at its wide grid while main's hidden headless model was created (or reflowed after renderer hydration) at 80 columns. Every byte the agent emitted while hidden was parsed 80 wide; reveal restored that image into the pane: rows clamped at column 80 with CHA fill, the status bar interleaved into response text, and for alt-screen TUIs the whole screen stuck in an 80x24 corner until a real resize forced a repaint. Scrollback damage was permanent. Fix, main-side only: - Pre-attach: seed `ptySizes` only for a genuinely fresh session id, or a measured request with nothing cached. A hidden reattach writes nothing. - Commit: on reattach, record the provider's proven grid (`attachedGrid`, set only by the local provider whose attach really resizes), then the reply's `snapshotCols/Rows` (the daemon emulator's grid), then the size main already held, and only then the request. - Reflow an already-created model to that grid after the seed block, so bytes that arrived before the reply no longer leave an 80x24 model. Both the ipc and runtime spawn paths take the same authority module. Renderer and wire formats are unchanged; `PtySpawnResult` is main-internal. Reproduced deterministically: close the floating panel with Claude Code streaming at 211x57, kill only the Electron main process so the daemon survives, relaunch. Main's cache read 80x24 against an applied 211x57 and the reveal snapshot was 80 columns wide; replaying the recorded bytes through an 80-column emulator reproduced the field screenshot. Relaunch with the panel open, and a fresh spawn, keep the wide grid. * fix(terminal): commit the adopted-claim reattach grid and reject non-integer provider grids Review follow-ups. The runtime spawn path's adopted-claim branch returned before the size commit, so an adoption attaching to a live session kept whatever the caller requested; it now commits and reflows like every other reattach. The grid validator requires integers so a malformed provider grid falls through to the cached size instead of reaching xterm. * fix(terminal): reflow main's headless model onto the committed grid for every spawn, not only reattaches A hidden attach whose daemon restarted comes back as a fresh session, and the pre-attach seed is now withheld for unmeasured attaches, so a live byte that created the model at 80x24 before the reply would have kept it there forever. * refactor(terminal): let the provider's reattach flag pick the adopted-claim grid source * fix(terminal): derive the adopted-claim reattach flag once for the size commit and the reservation The SSH relay's adopted reply carries no isReattach, so the size commit would have taken the request while the reservation was told it was an attach. Normalize once so both agree. |
||
|
|
38bde20121 | Update README downloads badge | ||
|
|
74ad08ec66 | fix(relay-ops): accept monitor evidence from an ancestor commit with identical monitor code (#18754) | ||
|
|
0f5f5e6979 | fix(relay-ops): retry a failed MIG inventory read once before calling a cell's power state unknown (#18740) | ||
|
|
746a6b4870 |
fix(orchestration): fence the dispatch CLI preamble so it stops rendering as headings (#18718)
* fix(orchestration): fence dispatch CLI preamble * fix(orchestration): keep optional preamble sections out of Markdown headings The sub-dispatch and base-drift sections end with a bare rule directly under a paragraph, which Markdown parses as a setext H2, so the Chat UI rendered the section's last sentence as a heading. The unfenced sub-dispatch commands also lost their angle-bracket placeholders to the raw-HTML pass. Fence those commands like the main CLI block and put a blank line before each closing rule. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
8096cb2803 |
fix(native-chat): render Claude structured chat through the same UI as Codex (#18743)
* fix(native-chat): render Claude structured chat through the same UI as Codex Structured native chat is one shared, agent-agnostic component tree, but three Codex-hardcoded terms on that path made a Claude session render differently. - `showTurnStatus` was `agent === 'codex'`, which gated the whole structured presentation: the live tool-progress row, the completion check and activity grouping (`structuredActivityUi`), and the Thinking / Working for N / Worked for N status rows. Claude fell back to the legacy compact chrome. - `runtimeContext` was likewise Codex-only, so a Claude transcript rendered images as filename chips instead of previews. - The composer's structured slash menu always served the Codex catalog, ignoring `agent`. That disagreed with the dispatcher, which does branch per agent: Codex-only tokens offered to Claude missed the command guard and were sent to the model as literal prompt text, where Codex shows an error. The first two gates landed Codex-first (#17597, #18266) before the Claude structured lane existed; they were rollout scoping, not capability limits. Neither `useNativeChatTurnStatus` nor `useNativeChatImageRuntimeContext` has any agent-specific logic. The menu now reads `structuredSlashCommands(agent)`, the function the dispatcher already used, so both read one list. No new rendering logic: two gates removed and one existing shared function reused. There are now zero `agent === 'codex' | 'claude'` branches in any native-chat component. * docs(native-chat): update structured turn status contract --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
86cd327749 |
Answer the structured-session support probe without installing the host (#18695)
* Answer the structured-session support probe without installing the host `getStructuredAgentSessionCreateSupport` called `ensureStructuredAgentSessionHost()` before answering, so a read-only "can you create a Codex session here?" question performed the create route's lifecycle work: the first install opens the durable agent-session record store, attaches the PTY write-gate record lookup and starts the orphan-child reaper. Ask the pure predicate instead. `supportsCreate` on the installed host resolves to `adapterSupportsCreate`, which for the Codex adapter is exactly `agent === 'codex' && supportsCodexStructuredLocation(location)` — no adapter instance is needed to answer it. Nothing is lost: the create/attach route still installs via `ensureStructuredHostInstalled`, and startup restoration still installs and reconciles when a store is already persisted. * test(runtime): cover structured support probe parity --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
2e80972450 | chore: update in-app Android APK link (#18745) | ||
|
|
0a821e5bc8 |
fix(crash-reporting): make the own-Chromium gate a real choke point, and stop a refusal leaking the root (#18459)
* fix(crash-reporting): make the own-Chromium gate a real choke point
Round-3 review found the guard was not the choke point its own comments
claimed: six pid-addressed `taskkill /pid <pid> /t /f` families in main were
ungated and uninstrumented, so the stale-pid shape stayed producible and a
`selfInitiatedTreeKillCount: 0` could read as exculpatory when it was not.
- Gate the remaining main-process families: the git command-runner abort, the
notebook-cell and automation-precheck timeouts.
- Turn the `src/shared` seam into the gate itself (`process-tree-kill-gate`), so
the runProcess choke point, the codex app-server deadline kill and the
ephemeral-VM recipe kill ask the same decision. Those three are compiled into
the CLI/relay too and cannot import main; main installs the guard at preflight.
- Ratchet (`main-process-tree-kill-gate.test.ts`): a new pid-addressed taskkill
in main that skips the gate fails, and the allowlist entries must still exist.
- Give pid-addressed kills eviction priority in the 32-entry ring: 32 routine
`win-pty-job` teardowns from a window-close burst no longer evict the one
entry that discriminates a self-kill from an external one.
- Correct the coverage doc, which described the uninstrumented Windows sites as
POSIX `process.kill(-pid)` group kills and omitted the git and codex paths.
* fix(crash-reporting): keep a refused tree-kill from leaking the root it owns
A refusal must block the pid-addressed tree walk, not the termination. Five of
the six gated sites returned on refusal with no fallback, so a refused
`taskkill /pid /t /f` left git.exe, a timed-out notebook cell, an automation
precheck or an ephemeral-VM recipe running while the caller reported it stopped.
The root kill is addressed by the child handle, which cannot reach the recycled
pid the refusal is about, so it stays correct and required on that path.
Also fixes the ring eviction the scope preference introduced: with the ring
saturated by pid-addressed kills, the only non-pid-addressed entry is the one
just pushed, so the splice evicted itself and the detail came back `{}` --
byte-identical to the external-kill arm, in the window-close case the guard
exists for. Eviction now excludes the newest entry and falls back to FIFO.
Tests: refusal now asserts the root kill at all six sites, and the ring covers
the saturated-pid ordering as well as round 3's group-burst ordering.
* fix(crash-reporting): stop a refused tree-kill leaking the commit-message agent, and count call sites
Two round-5 blocking findings, both open on main and on both branches.
`killSourceControlAgentProcess` had no root-kill fallback on its win32 arm: the
taskkill was the only termination, so once the own-Chromium gate could refuse it
the promise resolved having killed nothing. Both callers do
`terminationComplete ??= killSourceControlAgentProcess(child)` and then release
the managed-home lock on that promise, so a refusal left the local Codex/Claude
commit-message agent running while the caller reported it stopped -- the
lock-contention failure the taskkill was added for. Same fix as the six sibling
sites: the handle-addressed root kill cannot reach the recycled pid the refusal
is about, so it stays correct and required on that path.
The ratchet was file-granular, not call-site granular: one gate mention anywhere
in a file exempted every taskkill in it, which left the six files that now ask
the gate ratchet-blind -- the inverse of what it is for. It now counts `/pid`
call sites against gate admissions per file, so a second ungated kill inside an
existing family fails. Keying on the `/pid` argument rather than a quoted
`taskkill` also catches a kill whose program name comes from a constant. The
three comments that claimed more than the old scan enforced now state the rule
and its two remaining blind spots.
Also: the recording in `admitSelfInitiatedTreeKill` is now wrapped the way the
`admitProcessTreeKill` seam already wraps it, with the refusal decision taken
before anything that can throw so a diagnostics failure cannot flip it; and
`orca-chromium-process-pids` documents the false-positive direction (a stale
`getAppMetrics()` entry plus pid reuse refuses a live unrelated child), which is
the mechanism the root-kill fallback exists to bound.
Tests: refusal now asserts the root kill at all seven sites; the ratchet asserts
call-site counting and the constant-program form.
* test(crash-reporting): run the own-Chromium gate against real Windows trees
Nothing on this branch had ever executed on Windows. The unit tests pin the
gate's decision against a mocked taskkill, which cannot show that the decision
does anything to a real process: that `/T /F` reaps a detached grandchild, that
a refusal leaves that tree standing, or that the handle-addressed root kill the
refusal path falls back to reaps the root while orphaning descendants.
Adds a win32-gated live test covering all four, registered in both the
`package_windows` CI lane and `WINDOWS_PACKAGE_TESTS` as
`win32-test-lane-registration` requires.
Also completes the coverage doc's "never instrumented" list, which omitted the
macOS keyboard-input-source probe's POSIX group kill in `ipc/app.ts`.
* fix(crash-reporting): pin the commit-message root kill on the Windows arm
The first Windows run of this branch found nine failures the macOS suite
cannot see: `commit-message-text-generation-test-harness` asserts
`expect(child.kill).not.toHaveBeenCalled()` on `process.platform === 'win32'`,
which is the contract the previous commit deliberately replaced — and it
branches on the real platform, so it is dead code everywhere CI runs today.
The harness now asserts the handle-addressed root kill on every platform. On
win32 it lands after the tree walk, so the expectation waits rather than reading
one tick early, and its ten call sites await it. Red against the pre-fix arm at
all seven sites; the production code is unchanged.
* test(crash-reporting): remove the Windows lane marker tree through the retrying helper
The new win32 spec teardown used a raw rmSync, which the windows-lane-tree-removal
boundary ratchet rejects — and which is exactly the EPERM the ratchet exists to
prevent, since this spec's marker directory is written by processes it has just
force-killed.
* fix(crash-reporting): only refuse pid-addressed tree walks, disclose the handle-less codex site
The own-Chromium gate refused the POSIX process-group arm of
signalProcessTree as well, which was new macOS/Linux behaviour: a stale
getAppMetrics() entry plus pid reuse would orphan a group that main reaps
today. A POSIX group only holds what Orca put in it, so the refusal is now
scoped to win-taskkill-tree and the POSIX arm is recorded and admitted like
the other group kills in main. That also drops the synchronous
getAppMetrics() read from every POSIX termination.
codex-turn-added-roots kills roots found by a table walk, so a refusal has
no handle to fall back to. Pin that the refusal is visible - crumb written,
turn reported as not cancelled - rather than fixing what cannot be fixed.
* test(crash-reporting): detach the Windows survival fixture and observe real spawns
|
||
|
|
0b9d8586b3 |
fix: show agent commands once without startup polling (#18729)
* fix: wait for terminal line editor before agent startup input * fix: submit agent startup at the prompt without polling * test: update generated Bash prompt readiness contract |
||
|
|
367859bb06 | test(orchestration): audit migration stamps and documented commands | ||
|
|
586e6eaf8f | fix(orchestration): advance capability epochs after peer restart | ||
|
|
3def928a62 | fix(orchestration): accept local WSL fleet liveness evidence | ||
|
|
2cc2adc7a5 | fix(orchestration): preserve public task status corrections | ||
|
|
8064d1f991 | fix(ui): unmount project selector before dialog handoff (#18730) | ||
|
|
9a885b80c8 |
feat(native-chat): expose split and move-to-pane actions in Chat UI mode (#18714)
* Expose split actions in native chat * Fix native chat split regressions * Test chat split palette availability --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
6aba81f90a |
infra(relay): drop the unapplied region label from runtime log metrics (#18734)
The live metrics have role and cell_id only. A label change on a log metric is delete+create, so applying the region label would replace all 21 metrics, reset their history, and blank the relay alert policies during the swap. Matching Terraform to live state makes the targeted plan create-only (8 renewal metrics never applied, plus the incident dashboard from #18717). |
||
|
|
18922c803e | Merge remote-tracking branch 'origin/main' into orchestration-v3 | ||
|
|
4fab8e2f15 | Fix Smart create retaining a task checkout hash with Create more (#18727) | ||
|
|
a65332a8bd |
feat(claude): move structured native chat onto the Claude Agent SDK and enable it on macOS and Linux (#18560)
* Join structured attach teardown through journal bind * fix: restore structured chat parity * feat: add Claude structured session adapter * fix: harden Claude structured adapter * fix: close Claude adapter edge cases * fix: start Claude init deadline after launch * feat: wire Claude structured sessions * fix: harden Claude structured runtime * fix: fence Claude structured compatibility * fix: preserve Claude free-text prompt answers * fix: decode addressed Claude prompt text * feat: enable Claude structured chat on mobile * fix(mobile): keep structured chat provider-aware * fix(mobile): negotiate Claude structured tabs * fix: keep scoped RPC tests native-free * fix: secure mobile structured image delivery * fix: close structured session data-loss gaps * fix: prove real Claude structured startup * fix: consume pre-spawn proof before retry * feat(native-chat): add desktop structured sessions * fix(native-chat): satisfy structured session cleanup gates * fix(native-chat): keep structured renders pure * fix(native-chat): open composer pickers upward * fix(native-chat): use existing view for structured sessions * fix: harden structured desktop status projection * fix: close structured desktop lifecycle gaps * fix: fence structured AI Vault resumes * fix: fence structured AI Vault resumes * fix: preserve structured tabs during activation * feat: toggle structured sessions between chat and TUI * fix: harden structured session handoffs * fix: bind structured TUI before rollout proof * fix: complete structured chat round trips * fix: align structured TUI return readiness * fix(native-chat): make reverse handoff transactional * Add Claude structured TUI handoff seams * fix(native-chat): clear sticky handoff recovery * fix(native-chat): complete mobile reverse after TUI exit * fix(native-chat): keep TUI transcripts readable * fix(native-chat): recover TUI transcript gaps * fix(native-chat): recover claimed TUI owners * fix(native-chat): retain cold TUI proof authority * fix(native-chat): preserve Claude handoff authority * fix(native-chat): recover TUI transcripts read-only * fix(native-chat): harden Claude handoff recovery * fix(native-chat): serialize structured handoff recovery * fix(native-chat): close handoff admission races * fix(native-chat): validate pinned launch environment * fix(native-chat): revalidate restored and retried owners * fix(native-chat): gate restart recovery publications * fix(i18n): catalog Claude session controls * fix(native-chat): wait for structured TUI process proof * fix(native-chat): queue stale idle TUI handoffs * fix(native-chat): route structured Codex options directly * fix(native-chat): persist structured session options * fix(native-chat): hydrate resumed structured options * fix(native-chat): preserve options across structured handoffs * fix(native-chat): replay pending option mutations * fix(native-chat): rotate settled handoff operations * fix(native-chat): rotate refused send operations * test(native-chat): derive refusal retry state from host * test(native-chat): give the host-oracle matrix test an explicit timeout * fix(native-chat): keep Claude option controls idle * fix mobile structured first-send hydration race * fix(native-chat): preserve handoff launch authority * fix(native-chat): harden shared handoff recovery * fix(native-chat): serialize structured handoff recovery * fix(native-chat): close handoff admission races * fix(native-chat): validate pinned launch environment * fix(native-chat): revalidate restored and retried owners * fix(native-chat): gate restart recovery publications * fix(i18n): catalog structured session recovery control * fix(native-chat): wait for structured TUI process proof * fix(native-chat): queue stale idle TUI handoffs * fix(native-chat): keep structured recovery provider-neutral * fix(native-chat): drop local terminal topology from structured sync * fix structured outbox and tab restore races * fix(native-chat): preserve Claude question groups * fix structured provider visibility and request handling * fix structured session TUI handoff recovery * fix reverse structured session handoff * fix(native-chat): recover Claude outbox and resume state * chore(mobile): preserve the working-tree lockfile state before the main merge Carries the pre-existing uncommitted mobile/pnpm-lock.yaml modification into history so the main merge cannot overwrite it. Verified benign pnpm drift (babel 7.29.7->7.29.8 transitives plus deprecation metadata); drops no patchedDependencies (the mobile lockfile declares none). * test(native-chat): drop orphaned Claude handoff-auth test left by the main merge 'pins Claude handoff auth through the terminal provider boundary' is absent from main and its production counterpart preserveClaudeAuthEnv no longer exists outside this test - orphaned residue of the terminal/native handoff work this PR excludes by scope. Removed rather than repaired: the failure was a renamed field (providerHome -> providerRoot), and renaming it would have carried out-of-scope handoff code into the merge. Body preserved as evidence and logged in CLAUDE-STRUCTURED-DISPOSITION-TABLE.md. * Fix mobile structured turn state * fix Claude structured session blockers * fix claude structured lane blockers * fix Claude acquisition exit proof * fix(claude): route stream-json launch through process wrapper * fix(claude): gate structured chat support * Fix Claude structured launch gating * fix(claude): split session acquisition and prune mobile scope * test(claude): align structured session fixtures * fix(agent-session): preserve handoff launch arguments * fix(claude): open journals through the factory after origin/main split The journal opener moved to journal-store-factory on main; retarget the Claude structured tests that still imported the old path. * fix(claude): resolve Claude structured launch args, auth, and win32 proof The origin/main merge re-expressed the lane's Claude wiring onto main's split orca-runtime facade and dropped three wires past green typecheck and lint. - resolveLaunchArgs discarded its provider parameter, so structured Claude sessions were launched with Codex app-server flags; Claude exits on --dangerously-bypass-approvals-and-sandbox, and a Codex arg-parse throw could block Claude session creation outright. - resolveClaudeLaunchEnv was no longer supplied, so the launch resolver fell back to the whole process env as configuredEnv and buildClaudeChildProcessEnv re-applied every auth var it had just stripped. The resolver now merges the Claude overlay onto a strip-applied copy of the inherited env, which also keeps PATH intact for withCliRuntimeOnPath. - The windowsProcessStartTimeAvailable producer was gone while the contract field and both consumers survived, so the renderer gate fail-closed and structured native chat was unreachable on every win32 host. Separately, structured Claude pinned CLAUDE_CONFIG_DIR unconditionally. An explicit pin makes the CLI abandon the macOS Keychain even when it names the CLI's own default, so a default claude.ai account could not authenticate where the legacy Claude terminal could. Pin only a home the CLI would not resolve on its own, matching ClaudeRuntimePathResolver, and compare against the env the child would otherwise inherit so a diverging overlay cannot outrank the record's account home. Also await the now-async revealNativeSession in its regression test, and set the native status before revealing so a rejecting reveal cannot leave a session released but never marked native. Claude-Session: https://claude.ai/code/session_013UqKCRB6k5e8UaYhXUHeWY * fix(claude): scrub case-insensitive Windows auth env * fix(native-chat): settle handoff outcome-write failures instead of leaking them A store write failure while recording a handoff outcome escaped the flow runner's catch handler, so the client never received the failure and the flow surfaced as an unhandled rejection (seen as an intermittent agent_session_store_corrupt error in the proven-dead-retry suite, whose teardown raced the flow's trailing outcome write). Record the failed outcome best-effort, and drain the coordinator before that test's teardown removes the store root. Claude-Session: https://claude.ai/code/session_011aXkcHyeiRJuezupQdjZaM * fix(native-chat): make the structured close-failure toast provider-neutral The structuredSessionCloseFailed toast fires for any structured session, but its copy said 'Codex chat', so a Claude structured session that fails to close showed the wrong provider name. The launch-failure toast is only reachable behind the agent === 'codex' gate, so its copy stays as is. Claude-Session: https://claude.ai/code/session_013ugSpCx4AWkySaJb69BQax * fix(native-chat): wire structured handoff proof recovery * fix(native-chat): wire structured handoff proof recovery * fix(native-chat): correct the structured chat opt-in copy The one `experimentalStructuredNativeChat` toggle gates both providers — `useStructuredAgentSessionCreate` runs `canUseStructuredNativeChat` for `'claude'` as well as `'codex'` — but its description named only Codex. Its scope line also said Windows keeps using terminal chat, while the gate refuses win32 only until the host proves it can read a process start time. `structured-native-chat-availability.test.ts` already pins that Windows is allowed once the proof is cached, so the two contradicted each other. Claude-Session: https://claude.ai/code/session_01RJFsidQWmKYFmeoUuVu4Tp * test(claude): pin @anthropic-ai/claude-agent-sdk 0.3.251 contracts against a scripted CLI PR 1 of the SDK migration: dependency + test-only harness, no product wiring. - Pin @anthropic-ai/claude-agent-sdk to exactly 0.3.251 — not the newest release — because 0.3.251 (published 2026-08-28) clears the repo's 3-day minimumReleaseAge supply-chain gate with no exclusion, while the newest release was minutes old and would have required excluding a brand-new publish from the exact control built to catch brand-new malicious publishes. Every contract this design depends on was verified identical on 0.3.251: the full option surface, no pid on SpawnedProcess (custom spawner stays mandatory), env defaulting to process.env when omitted, and --replay-user-messages appearing only via extraArgs. - Exclude all eight bundled CLI platform binaries via ignoredOptionalDependencies. The setting lives in pnpm-workspace.yaml because pnpm 12 no longer reads the package.json "pnpm" field (it warns and ignores it; verified by install ablation). Excluding the binaries is what makes Orca's pathToClaudeCodeExecutable override mandatory rather than merely preferred. Note: pnpm 12.0.0 honors the ignore list when reconciling an existing lockfile but not on fresh resolution of a new dependency, so the lockfile's SDK entry was pinned surgically; both 'pnpm install' and 'pnpm install --frozen-lockfile' verify clean and stable against the committed lockfile. - Contract-pin suite drives the real SDK against a scripted fake CLI and pins: unknown type/field/content-block pass-through (and keep_alive interception), spawner env fidelity plus the omitted-env process.env inheritance sharp edge, extraArgs producing --replay-user-messages, argument parity for every CLAUDE_STRUCTURED_BASE_ARGS entry plus --session-id/--resume/ --resume-session-at, canUseTool wire request_id stability and abort on control_cancel_request, one spawn per query, pathToClaudeCodeExecutable honored by the default spawner, the exact SDK version, and the eight platform binaries staying uninstalled. Claude-Session: https://claude.ai/code/session_01FGCRfYUnb4hbvfTAHGtJKQ * feat(claude): drive the structured transport through the agent SDK Replaces the hand-rolled `claude -p --input-format stream-json` transport with @anthropic-ai/claude-agent-sdk 0.3.251, keeping the existing connection interface for this commit so the acquisition path changes minimally. The control-plane rewrite is a separate change. Orca still supplies the process. `spawnClaudeCodeProcess` routes through `spawnProcess`, retains the child and its pid — the triple the durable lease adjudicates on — drains stderr so exit errors keep their tail, and hands `.cmd` shims to Orca's Windows argument encoder rather than the SDK's plain spawn. `close()` keeps Orca's own bounded tree-kill and exit deadline, so it still resolves true only after an observed exit. Launch resolution emits an SDK options object instead of argv; durable `launchArgs` translate to a typed option where one exists and to `extraArgs` otherwise, refusing a token neither can carry rather than dropping it. The child env is always passed explicitly — omitting it would let the SDK inherit `process.env` and reintroduce the ambient `ANTHROPIC_*` leak. The stdout line parser is deleted; the SDK owns framing, and unknown frames still reach the translator verbatim. Claude-Session: https://claude.ai/code/session_01JMhFjh9HEnkcJ5YTfCdgD3 * fix(claude): settle the frame the SDK pulled but never wrote The SDK's input pump is `for await (frame of prompt) { await transport.write(frame) }`. When that write rejects — the child dies between Orca's liveness guard and the write — the for-await ends abruptly and calls the generator's `return()`, so the code after `yield` never runs. The frame was already shift()ed out of `queued`, so the later `fail()` from the exit path could not reach it and `send()` never settled: `dispatchClaudeTurn` awaits that send before it can return `unknown`, wedging the caller and the durable outbox. The pre-SDK transport rejected on the stdin write callback instead. Retain the in-flight entry and settle it from the generator's cleanup, and let fail() reach it too for the pump that never resumes at all. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * fix(claude): keep the agent SDK behind the structured-Claude boundary The ordinary OrcaRuntimeService graph statically reaches the Claude adapter and so the transport module, whose first line imported @anthropic-ai/claude-agent-sdk. The SDK is evaluated whenever the regular runtime loads, before any structured Claude session is chosen: it sets process.env.NoDefaultCurrentDirectoryInExePath, changing Windows executable resolution for later subprocesses, and a missing or incompatible install would break normal runtime startup — for a user who never leaves the terminal/TUI path. Defer the SDK to the connection, memoized so it loads once per process, and add the import-graph ratchet: a walk from the Electron main entry that fails on any static import of the package, plus a clean-fork check that loading the runtime leaves the Windows search variable untouched and a child-process pin that the side effect is still real. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * fix(claude): answer list_models so the picker stops serving the seed sendControlRequest had no list_models case, so every request hit the default reject; readClaudeStructuredSessionOptions swallows that with .catch(() => null) and falls back to the static catalog. Every structured session therefore served a hardcoded model list with no per-model effort levels, no resolvedModel and no default detection, and nothing surfaced the failure. The pre-SDK transport got the live catalog from the CLI. Route it through the SDK's supportedModels(), wrapped in the { models } envelope the existing parser reads. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * fix(claude): reap the child's descendants before killing it The forced step of the exit ladder went through the Codex helper, which spawns `pkill -KILL -P <pid>` and SIGKILLs the parent in the same tick: the parent usually dies first, the descendants reparent to pid 1, and `-P` matches nothing. An MCP or launcher descendant of a stubborn Claude child was left running. The test named for that requirement declined to assert it and killed the survivor by hand instead, so it could not fail for the thing it was named after. Route the Claude reap through Orca's existing sweep, which snapshots descendants while their parent link still exists and signals them before the root goes, and on Windows uses the identity-gated `taskkill /T /F`. The test now asserts the descendant is dead; the manual kill stays only as a failure-safe. close() still returns true only on an observed exit. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * fix(native-chat): merge the duplicated handoff type import CI's static-analysis lint (`oxlint --config config/oxlint-code-quality-native-plugins.json src config tests mobile --deny-warnings`) exits 1 on the two separate `import type` statements from the same module. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * fix(claude): answer a permission callback whose signal already aborted settleFrom registered the abort listener and then delivered the request. A callback that arrives already aborted never fires that event, so the promise stayed pending behind a durable prompt with no cancel path. Check the signal first, emit the cancel, and resolve the SDK's null sentinel without registering. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * test(claude): wait for the child to record the frame, not just for its report The scripted CLI writes its report at startup, so `until(readReport)` returned a report with no user messages whenever the child had not yet read the line. The assertion then failed under parallel load. Poll for the frame instead of for the file. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * fix(claude): coalesce partial deltas onto one assistant item and stop painting result frames Under --include-partial-messages every stream_event frame carries its own uuid, and the final assistant frame for a block carries yet another; only message.id ties them. The translator keyed each delta by its frame uuid, so a reply painted as one bubble per delta chunk followed by a complete duplicate under the final frame's uuid. The block's first stream frame now mints the claude:(sessionId, uuid) identity, deltas coalesce onto it through the shared 60ms seam, and the final frame reconciles onto that same item. Known SDK bookkeeping no longer reaches the provider-fallback row: result subtypes are catalogued and settled by the turn lifecycle, an empty thinking block (redacted thinking) is a modeled kind, a string-content user replay is a text block, and an empty user frame paints nothing. An unmodeled result subtype or content kind still lands on the bounded fallback row. Claude-Session: https://claude.ai/code/session_01GaP5HpYQbvy2hYehVhwfEW * fix(claude): prove descendant exit at the close boundary instead of on an unref'd timer close() reported proven=true as soon as the direct child exited while the descendant sweep's SIGKILL sat on an unref'd 2 s timer, so a SIGTERM-resistant MCP server outlived the lease release. The reaper now composes the same shared primitives the Codex structured provider uses: snapshot, verified bounded descendant termination on POSIX, taskkill /T /F on Windows. The proof is false whenever descendants outlive the deadline, a retried close re-verifies the retained snapshot rather than trusting the dead root, and the raw pipe child no longer goes through the PTY job sweep it never owned a job for. Measured on macOS: a killed child of a SIGSTOPped parent stays a matching zombie row in ps, so the root is killed while verification runs rather than stopped first as the Codex non-group path does. Claude-Session: https://claude.ai/code/session_0161QFm3KVRNJKfdzWVGVNWk * feat(claude): replace the hand-rolled control plane with the SDK's native surface PR 3 of the Claude structured SDK migration removes the wire-frame scaffolding PR 2 kept, so Orca drives the SDK's typed control surface directly. Inbound permissions move from a rebuilt control_request dispatch to the SDK's canUseTool / onUserDialog callbacks. The prompt registry now carries the callback's own resolver: a decodable can_use_tool becomes a durable prompt whose answer settles the callback; a malformed one is denied without registering; the SDK's abort signal (fired on control_cancel_request, which the SDK matches and dedups itself) forgets the prompt and settles it null, and a late answer after abort finds no prompt and is refused. Closing settles every in-flight callback so no promise dangles. The claude-agent-sdk-control-bridge that rebuilt the wire frame is deleted. Outbound control maps to Query methods: interrupt() for cancel, setModel / setPermissionMode / applyFlagSettings for options, supportedModels for the model list, initializationResult() for init proof, each under Orca's own request deadline and error classification. Cancel is interrupt-receipt aware: a CLI advertising interrupt_cancel_queued_v1 gets cancel_queued in one round trip, otherwise the receipt's still_queued uuids are swept with cancel_async_message so a cancelled turn cannot spawn a later unexpected turn; older CLIs resolve no receipt. Init keeps the 10s deadline and the unauthenticated-startup guidance. Every behavior is failing-first and ablation-proven; the toggle-off import boundary and the accepted loss of unknown-control visibility rows are unchanged. Claude-Session: https://claude.ai/code/session_01Pqjduxt5G4rr9aYvtp7rNm * fix(claude): arm the descendant snapshot before stdin closes and make the tree verdict unproven by default A healthy Claude root leaves within the graceful window, and the close ladder only snapshotted descendants when the root was still alive after that window. So the common close never looked at the tree: `treeExited` stayed null, `!== false` passed it, and close() reported a proven exit with an MCP child still running. A root that died before the walk made the snapshot vacuous too. The proof is now unproven by default. The reaper holds one verdict in Orca's vocabulary (exited / live / unverifiable), assigned in exactly one place from the bounded verification, and close() returns true only on `exited`. The snapshot is armed before stdin closes, while the root can still be walked, and is verified after the root exits; a root that left before any snapshot could be armed stays unverifiable rather than vouching for descendants it never showed us. The shared verifier gains the three-way verdict behind its boolean face, and the connection reports the root and tree verdicts separately along with the child's exit status. One verification per close attempt: the retried close re-verifies, so the intra-attempt re-reap is gone from the teardown budget. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * fix(claude): verify the Windows tree after taskkill instead of trusting that it ran `terminateWindowsProcessTree` resolves from taskkill's callback whatever the error says, so a timeout, an access denial, a recycled root and a surviving descendant all looked identical to the reaper — which then returned a proven exit unconditionally. close() reported true and the lease was released with an MCP descendant potentially still live. The Windows branch now snapshots the root's descendants while it is alive and, after taskkill, polls a fresh process table to a bounded deadline: a row still matching by pid AND creation time is `live`, an unreadable table is `unverifiable`, and only a table with no match is `exited`. Creation time is the PID-reuse guard the POSIX path gets from ps lstart, so a descendant that denied a creation-time query is omitted rather than signalled on a bare pid. A root already observed exited is never taskkilled: `/T /F` on a recycled pid would take an unrelated tree down with it. The captured tree is tagged by platform so neither verifier can be handed the other's rows. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * fix(claude): release a reservation on a first-hand root exit instead of latching it into manual recovery Making close() strict about the descendant tree exposed a second defect at the same boundary. A create-time acquisition has no ownerProcess until publication, so an unproven cleanup mapped to handoffStage `manual-recovery`, and adjudication then refuses every later attach with agent_session_ownership_unknown. A user who was merely signed out, or whose --resume the CLI rejected, wedged the session id permanently. Each question now answers from its own evidence. close() is unchanged and stays strict about the tree. Separately, the lease is keyed on the root's pid and start time, so when Orca's own child handle observed that root exit and no descendant snapshot was ever admissible, the reservation is released and the CLI's exit code and stderr reach the user. A descendant observed still alive, or a root Orca never saw leave, stays unproven and keeps the reservation. The settlement records only what was observed: the released lease says the provider process exited and its descendants were not verifiable, rather than reusing the wording that claims cleanup proved no child remains. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * fix(claude): surface an API error a result frame reports instead of settling the turn on it The SDK models an API failure as a SUCCESS-subtype result whose `result` string is the user-facing error text, with no assistant frame behind it. The translator suppressed every catalogued result subtype as turn bookkeeping, so that turn tombstoned its lifecycle and showed the user a completed, empty reply with no sign anything had failed. Suppression is now by meaning. A result reporting a failure routes to the bounded provider-error surface, leading with the provider's own sentence and keeping the raw frame behind the row's disclosure; ordinary successful results stay off the timeline as before. A turn the user aborted also stays suppressed: its interrupt frame already says so, and its execution diagnostic would only be noise on every stop. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * fix(claude): drop the stream state of turns that never received their final frame Every streamed delta recorded its block's identity, latest text and checkpoint length. Only the final assistant frame removed them, so an interrupted turn left its whole accumulated reply reachable until the session was disposed, and a long session with repeated interruptions grew those maps without bound. The partial text was already journaled by the flush that precedes settlement, so the live copy was pure retention. That state now lives in its own module, named for what it does — grow a streamed block's journal row between its deltas and its final frame — and turn settlement drops every block still awaiting a final. The translator reports how many remain, which is the invariant: a settled turn leaves none. Also makes a timed-out process-table read retryable while the root is still alive. A loaded host can miss the table's one-second deadline, and latching that as "no descendants" both lost the descendant sweep and, on a busy machine, made the close ladder report unproven for a tree it never actually looked at. Only the root's death still makes a missing snapshot final. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * perf(claude): capture the Windows descendant tree from one process-table read The capture walked the descendant tree and then read the table again for the creation times the walk's projection drops. Each read is bounded in seconds and both run inside the close ladder's budget, so the second one cost the worst-case teardown three seconds for data the first read already held. The walk is now exported from the module that owns it and runs over rows the caller has already read, which is also what lets the snapshot keep the PID-reuse guard the projection cannot carry. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * fix(pty): spend the descendant verification window instead of surrendering on one slow table read The verification abandoned the whole check the first time a process-table read missed its own one-second deadline, with seconds of its window still unspent. On a loaded host that reported a tree unverifiable without ever having looked at it, which the Claude close ladder then turned into an unproven close and a retried teardown. It also made the descendant-exit tests flake under a parallel suite run, for the same reason and with the same honest-but-premature verdict. A read that missed its deadline is now simply not an answer: the loop waits and reads again until its own deadline, and only a window that ends without a readable table reports unverifiable. This can only turn a premature verdict into one backed by evidence; it never manufactures a proof. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * fix(claude): never let a later failed look collapse an observed live descendant into unverifiable The reaper's single assignment site latched only 'exited', so a second reap whose table reads all missed their deadline overwrote an earlier completed verification's 'live' with 'unverifiable'. The acquisition release gate discriminates on exactly that pair, so a root exit after such a decay released the lease over a descendant that had been observed alive. The latch is now monotone in trust order: exited is final, and live is only ever raised to exited. Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP * fix(claude): never prove a Windows tree gone while a descendant denied identification The Windows snapshot dropped rows that denied the creation-time query, and an emptied snapshot was judged exited without any table read: a descendant Orca was refused information about was treated as one that had left. The snapshot now counts the unidentified rows it saw, and verification caps its verdict at unverifiable while any exist. Nothing is ever signalled on a bare pid, as before. Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP * fix(claude): classify cleanup after a first-hand exit as a root exit instead of a proven tree When the CLI died between a successful acquire and the host's commit or proof of the lease, handleExit had already removed the session, so releaseAcquisition found nothing and reported true. The attach flow then settled exit-proven with deathEvidence claiming cleanup proved no provider child remains, though the tree was never verified. The adapter now keeps the exit that removed a published session until the session is acquired again; acquisition cleanup runs that connection's close ladder and classifies its verdict exactly as a start-time failure would be, so the record reads root-exit-observed. The wire helper keeps that typed classification and its provider diagnostic instead of wrapping it as unproven, and the router gives up its owner even when the release throws. Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP * fix(claude): integrate SDK teardown and picker lifecycle fixes * fix(claude): preserve resume leaf and settle processless spawns * fix(claude): reacquire from persisted resume leaf * fix(native-chat): restore Claude grouped question handling * fix(claude): persist only resumable transcript leaves * fix(claude): recover structured session exits safely * fix(claude): close remaining structured session P1s * fix(claude): harden transcript branch proof * Remove superseded root fix reports * fix(windows): restore indexed descendant row walk * fix(router): forward force-close lifecycle * fix(claude): fence stale turn cancellations * fix(claude): fence cancellation after unknown dispatch * fix(claude): fence replay and option recovery races * fix(claude): block replay fallback after waiter eviction * fix(claude): fence evicted slash results * fix(claude): fence ambiguous results and restore options safely * fix(claude): scrub SDK child env and localize pending launch * fix(claude): pin transcript roots and exit recovery proofs * fix(claude): retain unproven SDK exits * fix(claude): settle retained exit before reacquire * fix(claude): resume from settled retained cursor * chore: remove tracked review artifact * fix: harden Claude SDK transport session cleanup * fix: close Claude sessions safely * fix(claude): close races with fresh child snapshots * fix(claude): fail closed on recycled child identities * fix(claude): gate root cleanup on process identity * fix(claude): fence same-second root identity reuse * fix(claude): restore the root SIGKILL fallback the identity gate took away The direct root kill goes through the handle Node owns, not through a pid: libuv drops that handle in the same turn it reaps, so the signal either reaches the process Orca spawned or reaches nothing at all. Gating it on a process-table probe therefore bought no safety and cost the tree its only fallback whenever the probe declined -- a first capture landing in the fork's own second, a recycled descendant pid voiding the snapshot, or a process table that could not be read on either platform. Identity verification stays where a bare pid is genuinely addressed: Windows `taskkill /T /F`, and the descendant sweep's own revalidation before it signals. Also stops a declined root probe from collapsing an observed `live` or `exited` descendant verdict into `unverifiable`, and stops a successful taskkill from reporting `unverifiable` because a later probe found the root correctly dead. * docs(claude): rewrap the root-kill ordering comment * Match the Claude structured launch to the terminal path's managed-account auth rules The SDK path stripped ambient Anthropic auth unconditionally, let an explicit agentDefaultEnv override beat a pinned managed account, and had no account-switch guard. Reuse the terminal preflight's own predicate and messages so both transports strip, refuse, and report identically, and cover the CLI transcript location that mobile native chat depends on. * Reach the Claude structured chat lane from the desktop UI The main process has had a complete, correctly gated Claude Agent SDK lane for a while, but no renderer ever asked for it: the launch route accepted only `codex`, and the create path was typed `agent: 'codex'` end to end. Widen both to the structured provider union that already exists (`AgentSessionHandleProvider`), and generalize the codex-named create path instead of adding a Claude twin beside it. The pending-launch registry is now keyed by agent as well as workspace — a shared key handed a second caller the first agent's intent, so a Claude and a Codex launch in one worktree collided. Windows, per agent. Codex's client-side win32 refusal is deliberate and settled elsewhere, so it stays exactly as it was. Claude's answer is no longer guessed from the client's platform: a structured session fences its provider child on that child's process start time, and only the executing host knows whether it can read one. `agentSession.createSupport` already answers precisely that, per agent, and had no renderer caller — so the Claude create path asks it before creating and turns a "no", or a probe it cannot get answered, into the definitive refusal the launch fallback already handles. Fail closed either way. That refusal mapping also closes a real gap: the host reports an unsupported location by throwing `structured_agent_session_unsupported`, which reaches the client as a transport rejection rather than a refusal envelope, so `StructuredAgentSessionCreateRefusalError` never fired. The launch would retry the create, strand itself in `visibilityUnknown`, run no legacy fallback, and show an error toast. Close a fail-open hole while Claude and win32 become reachable: `create` with a client-supplied location, and `ensure`, both skip the worktree-resolving support check. They now ask the executing host the same question directly, so a host that cannot fence a provider child no longer creates one on a client's say-so. Also deletes `structured-agent-session-provider-routing.ts`, a duplicate of `structured-agent-session-provider-support.ts` with no importers. WSL, SSH and paired hosts, floating workspaces, draft prompt delivery, explicit TUI customization and initial session options all keep refusing; folder workspaces keep working. * P1-1: make the structured Claude auth policy required and testable The optional dep plus a {stripAuthEnv:false} fallback meant a dropped wiring under-stripped silently. Required at all three hops, asserted at install time for the @ts-nocheck caller, and the settings-to-policy mapping is now a named tested function. * P2-3: mobile's default Claude transcript root must follow CLAUDE_CONFIG_DIR session-file-resolver's default ignored the variable the pinned account home follows, so a CLAUDE_CONFIG_DIR launch wrote one tree and mobile read another. The Task-4 test now resolves with no root override (mobile's own call) and checks the answer against the root the CLI itself reports, instead of mirroring the code under test's own expression. * P2-1/P2-2/P3: close the teardown window, join the live-auth gate, align the refusal P2-1: a switch beginning inside the acquire teardown left a dead chat and no replacement. Past that point the launch waits the swap out and refuses only if it never settles; the entry guard still refuses outright, because nothing is torn down there yet. P2-2: structured children now hold the same OAuth-refresh gate a Claude PTY does, so a managed refresh cannot rotate the token out from under a live turn. P3: the refusal now matches the strip it guards (case-folded on win32, presence not truthiness), and the dead structured-to-TUI builder states its auth policy instead of silently signing a system-auth user out. * Make the live-auth gate tests independent of sibling connection teardown order * Do not offer structured Claude under a WSL-only managed account Structured Claude launches against the ambient Claude config, which the account service keeps in sync with the selected HOST account. A WSL-bound managed account lives inside the distro and is never synced there, so on Windows a structured session would authenticate as whatever the ambient identity happens to be while the UI names the WSL account — the user is told one identity and given another. That was unreachable only because nothing offered structured Claude on win32. Enabling it makes it reachable, so gate it here rather than patching the auth layer: refuse the structured path when the active managed Claude account is WSL-bound, and let the terminal-backed path — which resolves the account per runtime — handle that account shape. The answer rides the agentSession.createSupport seam the renderer already consumes, so no new capability and no renderer knowledge of account internals. A create the host declines becomes the definitive refusal the launch fallback already turns into a legacy native chat tab, with no error toast. Unknown answers refuse. An install with no managed accounts claims no identity and is fine, but an active selection that cannot be resolved — or account state that cannot be read at all — is not evidence that the ambient identity is right. Claude only. Codex resolves its account through a different path and its createSupport answer is untouched, as is every Codex routing decision. * Read the structured Claude account gate through the auth policy's accessor The gate resolved the active account from the account-service snapshot's runtime map; the auth policy resolves it with getSelectedClaudeAccountIdForTarget(settings, { runtime: 'host' }). Those are two sources and two resolution rules, and they disagree on a legacy settings blob that carries the selection only in the flat activeClaudeManagedAccountId: the accessor falls through to it, a direct read of the runtime map does not. The gate would then refuse a launch the policy would have run under host-1 — and in the mirror case a session could be admitted under a policy computed from a different account than the gate approved. Read the same settings through the same accessor so agreement is structural rather than coincidental, and drop the controller accessor that existed only to reach the snapshot. No behaviour change for any state both already agreed on; Codex is untouched. * Round-3 review fixes: N-1 empty-value regression, N-2 gate leak window, N-4 lost history N-1: my presence-based conflict predicate refused a terminal launch that works today. 'ANTHROPIC_API_KEY=' is how a user blanks a variable and the settings pipeline preserves that empty value; an empty override cannot beat the pinned account and the strip removes the name anyway. Back to truthiness for the value, keeping the win32 case folding. N-2: enter the live-auth gate only after the exit/close handlers that release it, so no throw in between can leave an entry nothing reconciles. N-4: the Claude transcript resolver searches config-dir-then-default and de-dupes, matching the Codex sibling in the same file, so adopting CLAUDE_CONFIG_DIR no longer hides history written before it. * Run the managed-account gate on every Claude acquisition, not just create createSupport gates the create path, but a session's account state can change while it lives. A reacquire after an unexpected child exit re-resolves the launch and re-derives auth, with nothing re-checking the gate — so a session created while supported could come back up in the refused shape. With the strip predicate keyed on there being an active non-WSL account, the WSL-only user's normalized steady state (accounts exist, none active) does not strip, and that reacquire reaches the child with ambient auth while the UI names the account. Gate at resolveLaunch, the one choke point every acquisition passes through, refusing with the pre-spawn error the caller already handles. Same predicate as create-time, now sharing one settings reader so the two cannot drift. Claude only; Codex resolves its account on a different path and is untouched. The runtime class that wires this does not typecheck its own `this` calls — a missing hookup compiles clean — so the wiring is pinned behaviourally rather than trusted to the compiler. * Move the structured Claude gate out of the @ts-nocheck runtime files Both call sites of the managed-account gate sat in files whose first line is `// @ts-nocheck`, so neither was typechecked: three arguments to a one-argument function plus an undeclared identifier compiled clean. New auth-identity decision logic had no compiler behind it. Move the verdict into a checked module that takes the two facts the runtime owns — the adapter's answer and a settings getter — and decides. The runtime class now only forwards. Move the gate reader's construction into the checked installer too, so the nocheck file passes a plain settings closure and never names a gate symbol. Every reference to the gate predicate and its reader now lives in a checked file, so the ablation that used to pass silently is a compile error at both the create-support and reacquire sites. Removing the file-level @ts-nocheck is a separate, larger job and is not attempted here. * Derive the gate test's auth policy from the settings under test A hardcoded stripAuthEnv asserts a gate/policy pairing production cannot produce, and false additionally lets launch.env inherit the runner's real process.env. Derive via claudeStructuredAuthPolicyForSettings instead: the gate settings type is the same Pick the policy takes, and both resolve the account through getSelectedClaudeAccountIdForTarget. * Pin the absent-vs-empty distinction in the managed-account gate An empty claudeManagedAccounts array is a real answer: the user has no managed accounts, nothing claims an identity, and the ambient path is legitimate. A readable settings object with no such field is settings we failed to parse — the same unknown as unreadable — so it refuses. The two are one character apart in the code and the difference is invisible without the reasoning, so record it at the branch and pin both sides. The test fails under the obvious "consistency fix" of treating a missing field as empty. * fix(claude): keep command queue bookkeeping out of the transcript Claude Code 2.1.258 emits a `command_lifecycle` frame for every uuid-stamped command it starts, completes or cancels. The frame carries a command uuid and a state and no content, and the CLI keeps it out of its own transcript -- but it is absent from the SDK's SDKMessage union and so from Orca's frame catalogue, where an uncatalogued kind defaults to a substantive row. Every structured turn therefore painted raw JSON rows into the user-visible transcript. Catalogue it and disposition it as status chrome. The unknown-kind default stays `timeline-substantive`: a kind we have never seen is likelier to carry content than to be chrome, and a visible row we can catalogue later beats content we silently dropped. A lifecycle state that reads as a failure still surfaces, because the payload error check in `classifyProviderFrame` outranks the catalogue. * fix(claude): let a re-walked descendant become eligible for the forced sweep A descendant first observed by a capture inside its own birth second could never be SIGKILLed: `ps lstart` is second-resolution, so that capture cannot rule out a pid recycled later in the same second, and the merge pinned each retained row to the boundary of the walk that first saw it. SIGTERM-resistant children forked in that window were signalled and then never escalated -- they survived close, quit and restart, reparented to init, and had to be killed by hand. Advancing that boundary on any later capture would be unsound: a later capture matching pid, pgid and start-second is exactly what an impostor would also show. But a capture is not a match -- it is a fresh ppid walk from a root Node pins through its own handle, so a row it re-derives is proved ours at that instant without appealing to its start time. Chain the fence from there instead, and take that walk at the close boundary while the root certainly still lives: the root may leave inside the grace window, and the post-timeout refresh never runs. A row absent from the later walk still keeps its earlier boundary, and a row no walk has ever re-derived in a later second is still never escalated. * Treat an absent managed-account list as empty, not as unreadable An empty claudeManagedAccounts array and a missing one are the same answer: this user has no managed Claude accounts, so nothing claims an identity and ambient auth is the truth. Refusing on absence strands any profile that simply never wrote the key, and it disagrees with the auth policy, whose own predicate takes `(accounts ?? [])` for exactly this reason. Only settings that cannot be READ stay unknown, and those still refuse — as do a WSL-bound active account and a selection naming an account the list does not explain. The earlier reasoning treated a missing field as settings we failed to parse. That conflated "not present" with "not readable"; only the second is unknown. * Support structured Claude when accounts are registered but none is selected Registered-but-deselected Claude accounts were refused, which is behaviourally identical to having no accounts at all: the auth policy does not strip, ambient auth is the truth, and the UI names no host identity. A user who deselected their accounts silently got legacy chat with nothing explaining why. Nothing selected for the host runtime is two states the settings cannot tell apart after the fact, because pruneInvalidClaudeRuntimeSelection empties the host slot and persists null in the second one: honest deselection -> ambient auth, UI names nothing -> SUPPORTED the WSL-only steady state -> ambient auth, UI names the WSL account -> REFUSED The presence of any WSL-bound account in the list decides. Simplifying this to "none active -> supported" re-opens the auth-identity misrepresentation, so the tests fail loudly on exactly that: five of them, across the unit rule and the createSupport path. * Stop treating an unanswerable create-support probe as a refusal A worktree is not resolvable for a beat after createWorktree resolves, so a probe fired immediately after creation fails the RPC with selector_not_found instead of answering. The catch collapsed that into `supported = false`, so the composer refused and quietly built a terminal session — the gate never said no, it was never asked successfully. Elapsed time was the only input that decided whether a Claude launch went structured. "Could not answer" and "answered no" are different states and only the second is a verdict. Retry while the host cannot yet resolve the selector, with a bounded backoff that covers the measured window with margin, and keep refusing on the first ask for everything else. Fail-closed is unchanged: a probe that still cannot be answered when the budget is spent refuses. The retry is narrowed with the shared error-code matcher, which classifies a token that transports re-wrap into a longer message without matching prose that merely mentions it. Codex never probes, so this race has never been able to refuse a Codex launch — the race itself is identical for it. Recorded at the early return, because whoever gives Codex a probe inherits the bug. * fix(claude): fence the forced sweep on re-derivation, not on lstart's second A descendant forked in the same wall-clock second as every walk that sees it was signalled with SIGTERM and then never escalated, so a SIGTERM-resistant child survived tab close, app quit and a full relaunch. Two children of one parent 96ms apart across a second boundary took opposite paths. The leak predates this branch: it reproduces with the change reverted. `ps lstart` has one-second resolution, so a walk landing inside a row's birth second can never rule out a pid recycled later in that same second. But a walk is not a match: a ppid walk only reaches what the root actually parents, and the root is pinned by Node's own handle, so a row the walk re-derived is ours whatever second it was born in -- a stranger would have to have been forked into our tree, and then it is not a stranger. Fence the escalation on that. Rows a merge retained from an earlier walk are not re-derived and still answer to the start-time fence, which remains correct for them. Scoped to callers that revalidate identity before signalling, which is the Claude close path. Codex teardown reaches this same verifier and is unchanged; the argument holds there too, but widening it is its own deliberate change. Also reverts two changes from the previous attempt at this leak. Advancing the capture boundary on a later walk is inert once the sweep fences on re-derivation -- both key on the same set of rows, so the new term short-circuits for exactly the rows whose boundary it advanced. The extra ladder refresh was a duplicate full process-table read: close() already awaits tree.refresh() immediately before proveClaudeChildExit, on the only path that reaches it. Known property: the kill lands roughly a grace window after the walk that proved membership, so a pid recycled inside that gap could in principle be signalled. It is bounded -- matchingSnapshotRows already requires the live row to carry the same start-second and pgid, so an impostor must be born in the remainder of that one second, land on that exact pid, and sit in the same process group, and it has already received the unfenced SIGTERM from the same loop. * Run the Claude structured integration suite as a runtime client The suite exercises agentSession.* for Claude, not the mobile surface: nothing in it asserts anything mobile-specific and its sibling integration suites use 'runtime'. Mobile now additionally requires the experimental structured-chat setting, which structured-agent-session.test.ts pins in both states, so the stale 'mobile' fixture was claiming coverage it never had. * fix(claude): report effort from get_settings, which is the only frame that has it The composer's Effort pill rendered blank in every structured session. This is not a missing source: the publication reads `effortLevel` off the `system/init` frame, and that frame has never carried an effort of any kind, while the correct value is already fetched at acquisition and thrown away on the auth diagnostic. Verified two ways -- a live get_settings probe against Claude Code 2.1.258, and the shipped binary's own init frame construction, which lists `model` and no effort. So `reportedOptions.effort` was always empty, the options reader dropped the key, and the pill had no value. Model survived only because `currentModelId()` has a fallback chain. The get_settings call acquisition already makes reports the session's current effort as `effective.effortLevel`; pass that into the publication instead. Selecting an effort already worked, so this is the arrival value only. The legacy PTY path is unaffected and must not be "fixed" to match: it reads its effort by parsing the startup banner (`CLAUDE_MODEL_EFFORT` in src/renderer/src/components/native-chat/claude-terminal-session-options.ts), which is why it shows a value where the structured path does not. Also removes the fixture that hid this: the fake init frame invented `effortLevel: 'high'`, a field the CLI does not send, which is why every gate stayed green over a value that is always empty in production. The fixture's get_settings now returns the real {applied, effective, sources} shape instead of a bare `{env: {}}`, so the two adapter tests that asserted an effort keep asserting it through the path production actually uses. The reader returns null rather than defaulting: an effort nothing measured would repeat the fixture's mistake, and a blank pill is the honest degradation if the provider ever renames the key. * fix(claude): only record an effort the child confirms it adopted apply_flag_settings answers `success` for an effort it then ignores. Measured against Claude Code 2.1.258: applying `bogus-effort-xyz` returns subtype "success" with no error while `applied.effort` stays at its previous value, and a valid `low` moves it. The option write treated the absence of a throw as adoption and recorded the requested value unconditionally, so Orca would show and persist an effort the child was not using, with nothing anywhere reporting a problem. Read the effort back after applying it, through the same reader the arrival value uses, and reject when the child reports a different one. A readback that could not be taken is not evidence of a refusal -- the apply itself succeeded -- so it still records; only a readback that disagrees rejects. Not reachable from today's picker, which offers catalog values only, but the CLI's effort catalog is server-delivered and has changed before, so a retired id would otherwise become a pill confidently displaying a setting that never took. * test(claude): assert the effort contract against the real binary The blank pill survived every gate because the only tests that touched it were fixture-backed, and the fixture invented the field. A test that pins the shape we read cannot catch the provider renaming the key, which is the failure mode that produced this defect. Asserts both halves against a live authenticated CLI: that no frame it publishes carries an effort at all, and that the session's current effort arrives through get_settings. Which frame proves the session varies by host -- this machine proves it with a SessionStart hook rather than a system/init frame -- so the negative half asserts over every published frame rather than picking one. Skips with the rest of the file when no authenticated CLI is present. * fix(claude): stop the synthesised content-part kinds leaking into the transcript Sending an image put a bare `claude · message:user:content:image` row between the user's bubble and the answer. Two causes, and only the second is a family. An image part counted as modelled only when `source.type === 'url'`, but claudeDispatchMessageContent sends a local attachment as a base64 source and the CLI replays that shape back, so every attached image was classified unmodelled. Accept the base64 and file sources Orca itself sends. The family is the real defect. `message:<role>:content:<type>` kinds are synthesised at runtime from whatever `part.type` arrives, so unlike the top-level frame catalogue they can never be enumerated ahead of time -- the `?? 'timeline-substantive'` default then prints the synthesised name at a user who cannot act on it. That default is right for top-level frames, where "substantive" means show the frame; here it meant show our own vocabulary, which drops the content AND leaks the opcode. So an unrenderable part now renders a sentence saying exactly that, with the kind and payload still on the row's disclosure. A part that carries its own readable sentence keeps it -- the placeholder is a fallback, not an override. An unknown future part type is therefore visible, never silently dropped and never printed as a kind: the same principle as the effort readback, which records only what the provider confirms. * Declare agentSession.requestHandoff on the cross-version wire surface The manifest is a ratchet for cross-version reachability, so the method is declared with real HandoffParams rather than counted. requestHandoff is capability-gated through requireStructuredHost and has no client caller, so declaring it is the whole of the change. Also model two host capabilities the harness omitted: the stub host's supportsCreate, and the fake adapter's, without which adapterSupportsCreate falls through to a supportsLocation the fake also lacks. Every ensure was refused for the harness's silence rather than for its location. * Gate structured Claude session tabs on the client capability that names them The Claude structured lane deleted the projection's `agent !== 'codex'` filter and added CLAUDE_STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY in the same commit, but never wired the constant to anything. Paired clients then received agent-session tabs for Claude, which no shipped client renders -- mobile's resolveMobileNativeChat returns null for every agent but codex, so the row listed and selected into a pane with neither chat nor terminal. Restore the filter behind the declared capability instead of the bare agent name. No client advertises it yet, so this matches main's behaviour today and becomes a negotiation a future client can opt into. * Confirm the structured Claude model against the model the CLI reports set_model answers success for any string, including a model it cannot resolve — the failure only surfaces when the turn runs — and get_settings reports the settings-file model, not the session's. The init frame that opens each turn is the only channel carrying the adopted model, so keep the session's reported model current from it instead of reading it once at acquisition. Also stop rejecting an effort the readback cannot represent: max is session-scoped and excluded from the persisted effortLevel, so a readback reporting the level underneath it is an absence of evidence, not a refusal. * Clear the session-option hedge when the provider confirms the value The pill claimed every option was unconfirmed for the life of the session: the renderer recorded each write as dispatched and nothing ever moved it, so a model the CLI had already reported back still read as unconfirmed. Carry the provider's own confirmation to the surface. Main reports which option ids the provider named rather than merely accepted, and the client re-reads options as a turn changes, because the frame that opens a turn is where the adopted model arrives. A value the provider has not reported stays hedged, including an effort whose readback could not be taken. The confirmed list is optional on the wire: a host that predates it sends nothing and the client keeps hedging, which is the behaviour it had. * Keep the model report current across an acquisition fence bump * Show the picked session-option value and let the provider report correct it The pill showed a "not confirmed" second tooltip line for any value we had sent but not yet seen reported back. Nothing acts on it, and for the PTY lane it was permanent — that transport has no report channel. The pill now shows the picked value immediately and the provider's per-turn report corrects it when the two disagree; a newer local write still outranks a report that precedes it. `dispatched` stays as a provenance member rather than collapsing into `applied`: it is produced independently by the PTY lane, and it is where the `confirmed` wire field lands, which would otherwise be unobservable. Effort keeps its readback and its rejection path. That matters more now, not less: with the hedge gone the rejection is the only user-visible failure signal on this surface, so a spurious one would be the loudest bug here. Skipping the readback for an effort the settings response structurally cannot echo is what prevents it — the response carries the persisted level, so reading it back for a session-scoped value would report the level underneath and fail a valid write. * Hedge a session-option value only when the terminal transport sent it Both lanes emit `dispatched`, so it could never say which one produced a value. The descriptor now carries the transport that built it, set once in the shared snapshot builder from a parameter that is required rather than defaulted — the builder is the only place a descriptor is constructed, so a new producer has to name its lane or fail to compile. The structured lane confirms every value from the provider's own per-turn report, which makes the hedge transient noise there. The terminal lane can only learn an outcome by parsing the screen back, and only for Claude: every other agent's `dispatched` value stays unconfirmed for the life of the session, so the line is the only signal that we sent something we never saw land. * Refuse an effort the session's model advertises no control for * Refuse tab mutations on a Claude row the client never negotiated The branch added a case asserting a client advertising only agent-session.structured.v1 may mutate a claude row. That is the same ungated behaviour the projection gate removes, encoded a second time — mutation authorization reads the projection, so hiding the row refuses the write. Assert that contract instead, and add the positive case for a client that does negotiate Claude rows. * Resolve the Claude session's current model in one place so the effort guard and the pill agree * Record an effort the child did not adopt instead of refusing the write apply_flag_settings answers success for an effort it then ignores, so the readback exists to detect that. Refusing on it made the detection a veto, and a veto is only correct if the readback can never be wrong about which model is current -- which it was, twice. The pre-flight guard already refuses a level the model advertises no control for, so the veto guarded a door that is now locked upstream. Keep the detection, drop the refusal: a disagreement records the child's own answer and omits the option from confirmed, so main stops vouching for a value the provider rejected without blocking the user's write. * Stop a slow whole-machine ps from being read as an absent process `ps -axo ...command=` pays a per-pid argv read: measured 1.15s for 1,948 processes (0.03s without `command=`), and CPU contention stretched the same capture to 6.0s. Two budgets sized for a cheap look then misreport a readable machine. The reader's 3s ceiling killed 6 of 20 consecutive captures at load 27, so every consumer answered "unverifiable" about a table it could read. Raise it to 15s, and stamp the capture instant at ps START so `capturedAgeMs` is the upper bound its contract promises -- a 6s capture used to report itself as freshly taken, understating staleness against a 5s kill gate. The TTL keys on completion so a slow capture still coalesces instead of forking ps per caller. `readStructuredTuiProcessIdentity` then spent its whole 5s wait inside one capture and concluded "no exact child" after a single look taken before the child existed (observed landing at ~3.5s). Absence needs a look that did not race the spawn, so require two captures before the deadline can end the loop. Both surfaced by the real-binary Claude TUI resume test, which failed ~1 in 5 under load; 14/14 now, 8 of those runs containing a capture the old 3s budget would have killed. * Let the desktop renderer negotiate Claude structured tabs The paired-client gate hides agent-session rows an agent the client cannot render. The desktop renderer's own IPC dispatches as clientKind 'runtime' advertising only agent-session.structured.v1, so the gate hid Claude rows from the surface this feature ships on. It renders them; it should say so. * Stop a slow process table from silently blinding every freshness gate Stamping `capturedAgeMs` at ps START made the number honest, and honest broke both consumers that read it. `ps -axo ...command=` measured 2.5-9.0s on an idle 2,002-process laptop and 4.0-18.6s at load 46, so the age it now reports lands past every budget: `planRelayPtySweep` refuses the stop as "too old", and the renderer's `admitRemoteForegroundEvidence` refuses the record outright. That second one is the expensive half and was outside the diff -- a refusal bumps `consecutiveInspectionErrors`, the poll scheduler backs off to its 10s floor, and agent-completion detection stops for the pane. The subsystem went blind on exactly the loaded hosts the honest stamp was meant to serve. The evidence-publishing read now gives up at 1,200ms instead of waiting out `PS_TIMEOUT_MS`. It is one budget for one question: these consumers ask whether an observation describes NOW, and past this it does not -- a late answer is refused by the age gate anyway, having first blocked a polled path for the whole capture, so a prompt `unverifiable` is both the truthful verdict and the cheap one. Both relay call sites already produce it from a rejection, and an admitted `unverifiable` costs a poll where a refusal costs the cadence. Identity proof keeps the full 15s through `getFreshProcessTableSnapshot`, because it asks whether a process EXISTS and must never read slow as absent. The budget bounds the wait, never the capture: the reader coalesces, so an abandoned wait leaves its capture running to fill the cache rather than forking a second whole-machine `ps` on the host that can least afford one. 1,200ms is bracketed rather than picked. The floor is the capture's own cost -- `command=` measured 1.15s for 1,948 processes on an idle host, and a budget under that answers `unverifiable` about a machine nobody is straining. The ceiling is the consumer's: 2,000ms, less the 500ms a TTL-shared capture may already have aged, leaves 1,500ms, and transit takes the rest. That ceiling only fits once the capture stops being charged twice. `ps` runs inside the RPC round trip, so its duration is already in `receiveDelay`, and `capturedAgeMs` is that same duration on the host's clock; summing them halved the budget this gate grants a host from ~2.0s of `ps` to ~1.0s, which is why a 1.2s capture arriving at 1.3s read as 2.5s old and was refused. Admission now takes the larger of the two. The sweep's gate keeps its sum, which is correct there: `evidenceAgeSinceListingMs` is stamped after the listing ARRIVES, so it measures planning time and overlaps nothing. A stated limit rather than an assumed one: 15s is not proven sufficient for identity proof. The same capture reached 18.6s at load 46, so that path can still time out and answer "no exact child" about a host it simply could not read in time. Narrowing it needs a cheaper question than a whole-machine argv read, not a larger number. The one test guarding this field could not fail. `beginPtyHandlerTest` installs fake timers, so `Date.now()` is frozen, the real reader reports exactly +0, and `0 <= 500` held identically for a hardcoded zero, for completion-stamping and for start-stamping -- while the real reader on that host returns thousands of ms. It now drives a measured age in and asserts the handler publishes it rather than restamping; that the reader MEASURES it correctly stays pinned separately, against a controllable clock. Both consumers get boundary coverage either side, and each new gate was ablated red before it went green. * Keep the compatibility fields off the capture the budget just abandoned inspectProcess falls back to processHasChildren and listProcesses to getForegroundProcessName, and both read the same TTL-shared capture with no budget of their own. On a slow host they joined the in-flight capture the budgeted evidence read had just given up on, so the call still blocked for the full 6-18s and the budget bought nothing -- once for inspectProcess and once per managed PTY for listProcesses. Use the degraded answers those helpers already give for an unreadable table, reached promptly. pty.hasChildProcesses keeps its unbudgeted fresh probe: it is a one-shot destructive gate that can afford to wait. --------- Co-authored-by: Merge Sim <merge-sim@local> Co-authored-by: Merge Sim <sim@local> |
||
|
|
264c9ed8d2 |
fix(browser-pane): stop a dead client-hosted guest from killing the workbench (#18334)
* fix(browser): stop a dead client-hosted guest taking down the workbench ClientHostedBrowserPagePane called raw <webview> methods from two React effects, so a guest that is gone throws out of a commit phase and unwinds the terminal.workbench error boundary instead of showing the pane's own unavailable notice. Two runtime conditions, two guards: - Guest destroyed in main while the tag is still in the DOM: the tag keeps its guestInstanceId, so every read throws 'Invalid guestInstanceId'. The metadata read is now total and the attach effect degrades to browser_client_page_guest_unavailable. - Retained tag removed from the DOM while the pane stays mounted: contentWindow is null, so focus() throws a TypeError. The activation-focus hook now goes through the BrowserPageGuestFocus wrapper the pane already builds, which has carried that guard since STA-3448. Follow-up, not in this change: the registry's liveness check compares readBrowserClientPageAttachedGuestId(webview) to page.webContentsId, which still matches after main destroys the guest, so a stale 'attached' page can linger. * fix(browser): keep the dead-guest degrade honest — no spinner, no silent swallow Review remediation for the guest guards. - The attach bail now writes `loading: false` before setting browser_client_page_guest_unavailable, matching what retryGuestRecoveryRef already does on the pane's own route into that state. A page that died mid-load carries `loading: true` in the store, so without it the unavailable notice rendered beside a spinner nothing would ever stop. Uses the existing updatePageStateFromGuest effect event, not a new setter. - The dead-guest condition is no longer silent: the metadata read logs the caught error under the subsystem's `[browser-client-page]` warn convention, and the attach bail records a `browser_client_page_guest_unavailable` crash breadcrumb via the existing recordRendererCrashBreadcrumb. Renderer diagnostics only capture window error/rejection, so the breadcrumb is what puts this on a channel crash reports actually carry — the registry liveness defect this change deliberately does not fix stays measurable, and a read failure that is not `Invalid guestInstanceId` is no longer indistinguishable from a dead guest. - onFailLoad no longer pays five sync IPCs per discarded event: resolveBrowserWebviewLoadFailure accepts a lazy fallbackUrl and resolves it after the subframe/ERR_ABORTED filter. Covered by a new case in browser-webview-load-failure.test.ts. Correction to the previous commit's narrative: only the metadata half reaches browser_client_page_guest_unavailable. The focus half reaches no state at all — the guarded BrowserPageGuestFocus wrapper returns false and the pane stays mounted over a webview the registry already removed from the DOM, with no notice and no reopen-on-server escape. It stops the crash; it does not diagnose the page. Not changed, with reasons: - The two sibling attach bails (renderer-unavailable, attach threw) omit the same `loading` write. That predates this branch and neither is reached by the dead-guest path; fixing them is a separate change. - recordHistoryFromGuest still passes a raw `webview.getTitle()`. Substituting `metadata.title` is not behaviour-preserving: metadata.title falls back to the URL, so an untitled page would be filed in history under its URL instead of "New Tab". The call runs only after a five-read succeeded and sits in a DOM event listener, which cannot unwind a React commit. * fix(browser): route every client-hosted guest death to the unavailable notice A dead guest could still leave the pane mute (retained tag fenced on render-process-gone/destroyed with no signal to the pane), spinning forever (did-start-loading read bailing after loading:true), or frozen at stale chrome (navigation reads bailing silently). All of those now go through one watcher that detaches, releases the webview ref and enters the pane's existing browser_client_page_guest_unavailable recovery state, so the user always sees the notice with its reopen-on-server escape. The total catch in the guest reader now records a browser_client_page_guest_read_failed breadcrumb with the error name/message, so a swallowed failure that is not guest death stays distinguishable in diagnostics; the guest_unavailable breadcrumb carries the loss reason. * fix(browser): finish dead guest cleanup and guard history reads |
||
|
|
c16913ab65 |
fix(native-chat): clarify active progress (#18705)
Co-authored-by: Merge Sim <sim@local> |
||
|
|
d3501f7ad6 |
fix(worktrees): stop a failed worktree scan from being recorded as an authoritative empty listing (#18456)
* fix(worktrees): stop a failed worktree scan from pruning as an authoritative empty listing A `git worktree list` that could not run at all — a WSL distro that stopped resolving, a hung mount, a git binary that errored — was softened to `[]` by the lenient listing path, so the detected scan published it as `authoritative: true, worktrees: []`. That disables the #1158 retention guard, drops every persisted tab for the repo, and the pruned session is written back to disk on the next launch. The loss is permanent, not a transient glitch. Route the detected scan through a strict listing that still reports the two genuinely empty states (repo path gone, not a Git repo) as `[]`. Everything else rejects, so the existing catch answers `authoritative: false` and the destructive halves (`rememberLocalWorktreeRoots`, `pruneLineageForMissingRepoWorktrees`) never see a failed scan. Also make the failure readable: wsl.exe reports its own launch failures as exit 0xFFFFFFFF with an EMPTY stderr and the `Wsl/Service/WSL_E_*` line on stdout as UTF-16LE, which is why the field bundle carried a git error with no text. Set WSL_UTF8 for WSL-routed git (matching the wsl runner, #9010) and attach that stdout diagnostic to the error so `git.exec` spans name the cause. * fix(worktrees): surface a failed worktree scan's cause on the repo header with a retry A failed scan now travels with its reason (optional unavailableReason on DetectedWorktreeListResult), the repo header shows it with click-to-retry, and the WSL deleted-guest-directory shape measured on a real Windows host is pinned as retained-not-pruned. |
||
|
|
b6f453df06 |
perf(relay): per-cell inventory locks, delta counters, and a pool statement timeout (#18722)
The sticky refresh path and reservation reconciliation both took the fleet-wide `relay_cells ... FOR UPDATE` scan to mutate one or two rows, so one busy cell queued unrelated reconnects and migration completions behind it. Both now lock only the rows they touch, in the same ascending cell_id order, and the sticky grant moves its counter by a delta instead of writing back a snapshot value. Placement keeps the ordered inventory lock: choosing the least-loaded cell is a genuinely fleet-wide decision, and dynamically locking only the selected target is what allowed cross-cell cycles before. The pool's statement_timeout becomes env-configurable and a 57014 now reaches the bounded transaction retry instead of surfacing as a terminal failure. Schema DDL moves to its own `max: 1`, statement_timeout-free pool that is ended before the serving pool opens, so a slow CREATE INDEX cannot inherit a request deadline it will never fit inside. |
||
|
|
901c6771ff |
fix(agents): detect agent CLIs installed outside a version manager (#18336)
* fix(agents): detect agent CLIs installed outside a version manager The install-dir fallback that answers "is this agent installed?" when the login-shell PATH probe fails listed only version-manager bin dirs, so codex/opencode/cursor-agent installed by Homebrew, npm's default global prefix, snap, nix, or the CLI's own installer read as not installed. Where this is decisive, corrected from the previous message: the `orca` CLI, whose detectSkillsCliAgentKeys (src/cli/handlers/skills.ts) calls detectCommandsInInstallDirs directly and whose entry point seeds no PATH, plus unpackaged/dev runs. NOT the packaged desktop app: patchPackagedProcessPath (configure-process.ts:114) appends /opt/homebrew/bin, /usr/local/bin, ~/.opencode/bin and the Linux/nix prefixes onto process.env.PATH at main-process-preflight.ts:147, before any detection, and mergePathSegments preserves them, so the PATH scan reaches those dirs first and this fallback never fires for them. That means this does NOT explain the packaged macOS v1.4.194 report of codex/cursor-agent/opencode all undetected -- that report stays open and uninvestigated. Second correction: the fallback now carries the prefixes Homebrew actually uses on Linux (/home/linuxbrew/.linuxbrew/bin), plus /snap/bin and the two nix profile dirs, matching what patchPackagedProcessPath already seeds. The WSL guest prelude gains the same entries. Leaving them out closed the native/WSL asymmetry on darwin only, on the platforms where the fallback is decisive. Appended last so a version-manager install still wins, and kept out of getVersionManagerBinPaths, whose result is PREPENDED to PATH (#18234). Lives in its own module so node-cli-command-resolution.ts stays under max-lines. The test stages every path through `join` and asserts via detectCommandsInInstallDirs as well as resolveCliCommands, so it holds on a Windows dev machine and pins the "absolute path means installed" contract. * fix(agents): align system install-dir order across the three PATH lists Round-2 review remediation. The blocking finding was about the handoff artifact, not the code: the summary handed to review described a 2-file/+38 change with 3 new macOS dirs, while HEAD is 4 files/+227 with 6 lookup dirs plus 5 new WSL-guest prelude entries, and the quoted failing test names never existed. Restated against HEAD in the handoff; no rebuttal, the reviewer was right. Justification for the entries the summary never described: the fallback exists to close the native/WSL asymmetry for a CLI no version manager installed, and patchPackagedProcessPath already seeds Linuxbrew, /snap/bin and both nix profile dirs (configure-process.ts:141-158). Shipping only the darwin subset would have left a Linux or WSL user with a snap/nix install still reading as not installed while the packaged macOS user did not -- the asymmetry the change is for. Code changes, all from the non-blocking list: - The three lists disagreed on order while claiming to be kept in step, so a CLI in both /usr/local/bin and /snap/bin could resolve to a different binary than the seeded PATH scan or the WSL guest probe found. All three now use the seed's relative order, pinned by a new duplicate-install test and an offset assertion on the prelude. The prelude's system block also moved after the nvm glob so a version manager still wins in the guest, as it does natively. - The parity docstring asserted "the same set patchPackagedProcessPath appends, minus the sbin dirs and the generic ~/bin", which was false: it also omits ~/.vite-plus/bin (seeded by configure-process.ts:156, but no probed agent command maps to it) and /opt/homebrew off darwin. All three gaps are now named as deliberate. - win32 returns [] and stays that way, but the branch now says why: %USERPROFILE%\.opencode\bin has never had install-dir coverage in either list, and the seed's system block is POSIX-only too. Pre-existing, unchanged. - The detectCommandsInInstallDirs case read the ambient process.env.PATH, so on a box with /usr/local/bin on PATH only the opencode assertion exercised the fallback. It now stubs the GUI-launch PATH, so both do. Unchanged and restated: this does NOT explain the packaged macOS v1.4.194 report of codex/cursor-agent/opencode all undetected, and must not close G6 report #3. patchPackagedProcessPath returns early unless app.isPackaged and seeds all six dirs before any detection, and mergePathSegments never deletes them, so the packaged PATH scan reaches them first and this fallback never fires there. Decisive only for the `orca` CLI's detectSkillsCliAgentKeys and unpackaged/dev runs. Verification: the suite is red without the production hunks (5 failed/3 passed) and green with them (8 passed); 11 related suites pass (197 tests); tc:node, tc:cli, oxlint, oxfmt and the max-lines ratchet are clean. * fix(agents): correct the ordering-parity claims and widen the seed-leak guard Round-3 review found two docstring claims that are false as written and one guard that only asserted a third of its list. - system-cli-install-dirs.ts claimed a CLI in two of these dirs resolves the same here as in the packaged PATH seed. True inside the system block, false across it: `claude` in both ~/.local/bin and /opt/homebrew/bin resolves to ~/.local/bin via the fallback (getBaseVersionManagerDirectories leads) and to /opt/homebrew/bin via the seed, which appends ~/.local/bin last. Scope the claim to the block and name the gap instead of asserting it away. No behavior change: closing it would hoist a system dir over a version-manager one (#18234). - posix-version-manager-bin-dirs.ts justified moving "/usr/local/bin" after the nvm glob with "a version manager still wins in the guest, as it does natively". The glob expands lexicographically; native orders nvm dirs default-alias-first (#10932), so that is not parity. Record the move as the one behavior change in the file and bound it: entries are appended behind a resolved login PATH, and both consumers only test presence. - The #18234 seed-leak guard asserted only /opt/homebrew/bin and /usr/local/bin were absent from getVersionManagerBinPaths, leaving the four other new dirs unpinned. It now spells out all seven across darwin and linux -- spelled out rather than derived from getSystemCliInstallDirectories, which would pass vacuously against exactly the refactor it guards. Tests 8 passed (8); 5 failed / 3 passed with the production hunks reverted to origin/main. tsc node + cli clean, oxlint clean. * fix(agents): find Pi's own installer dir in the CLI install-dir fallback The fallback added `~/.opencode/bin` but skipped `~/.vite-plus/bin` on the claim that no probed agent command maps to it. False: `pi` is a probed detect command on every runtime (`tui-agent-config.ts`, no `detectUnsupportedRuntimes`) and `~/.vite-plus/bin` is the Pi installer's default — the two dirs #829 named and `patchPackagedProcessPath` seeds together. Added to both the native fallback and the WSL guest prelude, so a Pi installed by its own script is found by the `orca` CLI and in WSL, not just on a seeded packaged PATH. Also, all narrower: - `/snap/bin` + Linuxbrew now gate on `linux` like the seed does, instead of every non-darwin posix. - Docstring: `/opt/homebrew` off darwin is the one remaining seed gap and says why; the "lookup-only" paragraph names the `withCliRuntimeOnPath` exception. - New seed-order test derives the expected order from `getSystemCliInstallDirectories`, so reordering either list fails. The PR body's claim that SSH hosts benefit is dropped: they answer `preflight.detectAgents` from `src/relay/preflight-handler.ts` via `isCommandOnPathForRelay`, a separate bundle that never reaches this module. * fix(agents): build the CLI install-dir fallback order once and pin it on both resolvers resolveCliCommand (every spawn site) and resolveCliCommands (detection) each spelled the nvm -> version-manager -> system-dir order by hand, which is how the native and WSL lists drifted apart before. One getCliInstallDirectories now feeds both, and the test pins system dirs LAST on both resolvers, on darwin and linux, plus a derived check that the WSL guest prelude keeps every native version-manager dir ahead of the native system block. |
||
|
|
0032074fd7 |
skills: keep only the orchestration guide rewrite in this PR
The seven non-orchestration guide rewrites, the shared stub fragment, and their guards move to a separate PR. The orca-cli cross-guide pin follows the worktree-selector rule into the orchestration placement reference. |