mirror of
https://github.com/stablyai/orca.git
synced 2026-09-28 16:02:45 +00:00
ae426bfbeb024fc4d0f3c7ac0dafc84dfcdacc4e
1047
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
21cbc15469 | Merge remote-tracking branch 'origin/main' into nwparker/piere-diffs | ||
|
|
2626e2eca4 |
Make the structured turn lifecycle row durable so completed durations survive (#19695)
* Make the structured turn lifecycle row durable so completed durations survive A structured-chat turn used to end by tombstoning its running lifecycle item, which threw away the only durable record of when the turn ended. Completed "Worked for" labels therefore depended on the renderer having observed the turn finish, and vanished on reopen. The lifecycle item is now revised in place, never tombstoned: - running, with startedAt, at the provider's turn start - completed or interrupted, with completedAt, at the provider's terminal frame, a user stop, or a child exit the host observed - unverifiable, with no end, when a cold acquire finds a running row from a generation whose exit nobody observed Both timestamps are the execution host's clock at receipt, captured before the deferred sink, so the completed value is identical on every client and needs no client clock. Codex history restore uses the provider's own second-granular endpoints for turns that predate this change. Desktop and mobile read settled durations off the journal through one shared selector, and anchor the live counter on the host start with the client's local receipt so a skewed client clock never leaks into the label. Locally observed durations remain the fallback for hosts that still tombstone. Timestamps live inside the existing turnLifecycle field, which old clients strip, and every working-state consumer keys on state === 'running', so no capability negotiation is needed. * native-chat: avoid stale working status on settled turns * test: align settled turn status expectations * Name settled lifecycle rows by their terminal state An interrupted or unverifiable turn must not read as completed for any consumer that renders status text raw. One shared helper builds the text for both providers from the lifecycle state. * test: deduplicate turn lifecycle suites Each behavior keeps one test; duplicated harnesses and restated cases go. * Key lifecycle rows to their user item and record the provider's measured duration A lifecycle row now names the user item that opened the turn by its provider key, so clients attribute timing explicitly and fall back to journal order only for rows from older hosts. A provider-initiated turn with no prompt can no longer claim the previous prompt's duration. When the provider measures the turn itself (Codex turn.durationMs, Claude result.duration_ms) the terminal row records it and clients prefer it over the host interval, so a turn shows the same number live and after a history restore. Host receipt times remain the live-counter anchor and the fallback. * Record a turn as a first-class journal item The turn record is now its own item kind rather than a status row carrying a lifecycle field: no text to misuse, and the fold matches the durable turn record other systems keep. Rows that carry it are stamped journal schema v3; every other row stays v2, so an older host keeps reading them and latches read-only at the first v3 row instead of truncating the epoch. Clients that predate the item would paint an unknown kind as a text bubble, so the host publishes the legacy status form to any client that does not advertise agent-session.turn-item.v1, through the same per-client seam background tasks use. The downgrade is transitional and goes once no supported release lacks the capability. The shared projection now renders unknown item kinds as nothing, so later kinds need no gate. One shared reader handles both forms for old journals and old hosts. * Preserve observed turn end across settlement retries * Retain turn attribution for loaded chat history * Preserve Codex exit receipt across close retries * Register completed turn duration reliability gate * Keep earlier turns through a Codex rewind and count a mid-turn attach from the real start Findings from an independent adversarial review of the typed turn record: - A Codex rewind adopted the provider's item list as the new epoch, and the provider never returns the host's own turn rows, so every duration before the rewind point vanished. The host's turn rows are now spliced back beside the item each followed, and recovery no longer expects the provider to prove rows it never owned. - The epoch row was stamped with the current schema version, so an older host latched read-only at row 1 of every new session, defeating the mixed version design. It carries no body and stays at v2; a stored-row test now reads SQLite directly, because the reader upcasts every row on read. - A send Codex folds into a running turn shares the opening prompt's provider key, and the alias map credited the duration to the later prompt. The earliest submission naming a key now wins. - The live counter anchored on first sight, so a client attaching mid-turn counted from zero. Published frames now carry the host's clock, the reducer keeps the last sample with its local receipt time, and both clients anchor on how long the host says the turn has run. * Correct turn duration gate assertion reference * Respect authoritative unknown native chat duration * Preserve unverifiable timing across older host upgrade * Record final completed turn duration reliability evidence * Fix the CI failures the merge left behind - A merged import list named the same module twice, which the native code quality plugin fails on. - A running turn is now reported by the host with no duration, so the settled map carries an explicit null for it; the hook test still expected the entry to be absent. - main gave the older-page action a cursor with a head-trim guard, so the retention test's epoch-only action no longer typechecks; it now passes an unbounded sequence, which is what the old shape meant. - The roster comparator moved into the extracted module, leaving its import unused in the reducer. * Split two files back under the line cap after the merge Merging main put both one effective line over 300, and the cap forbids a disable or a shave. The wire module's refusal vocabulary moves to its own file and is re-exported, so its consumers are untouched; the host's four thin mutation delegates move next to the functions they call. * Advertise the turn-item capability on every client transport Local IPC and mobile advertised it; the remote and web transports did not, so a desktop paired to a remote host, the CLI, and web silently ran on the legacy carrier forever and the canonical row was never exercised there. The renderer that paints it is the same build on every transport. * Update the web auth-frame expectation for the new capability --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
4b1b7178ad |
fix(orchestration): scope @ group addresses to the sender's Run (#19783)
* fix(orchestration): scope @ group addresses to the sender's Run `@all`, `@idle`, and the agent-name groups (`@claude`, `@codex`, ...) resolved against every terminal on the host. A coordinator meaning "my three reviewers" reached 126 agents across every open project, twice in one day, and every unrelated agent burned a turn discarding mail that was never for it. Every group except `@worktree:<id>` now means the live Dispatches of the sender's own Run, each addressed as `dispatch:<id>` so delivery is durable even when the worker terminal is not attached yet. A sender bound to no Run is refused with `invalid_argument` naming `run:<id>` / `dispatch:<id>`; there is no host-wide fallback and the host's terminals are never enumerated for it. `@idle` and the agent-name groups filter within that set by the same terminal status and host-resolved identity as before. `ask --to @group` returns the same code and points at the owning Run mailbox. Federated Dispatches read relayed control mail rather than a local mailbox, so a Run-scoped fan-out skips them with a `recipient_unreachable` warning naming the direct `dispatch:<id>` address. Group addresses are resolved host-side, so no RPC or stream shape changes; an older CLI sending `@all` to a new host gets the Run-scoped meaning. Claude-Session: run-scoped-group-addresses * fix(orchestration): revalidate legacy takeover before the recipient verdict A legacy coordinator taken over while `listTerminals` was in flight reported `runtime_error` instead of `legacy_read_only`: Run scoping made "no live Dispatch in this Run" the first thing the group send could fail on, and that threw before the takeover check ran. Takeover is a precondition, not a commit-time detail — the sender must be told it is read-only whatever else is wrong with its recipient set. Revalidation moves to immediately after the only `await` in the path. Everything below it is synchronous, so the commit-time window it used to guard is unchanged; only the error paths now see it. The legacy partition test gave `term_current_worker` no Dispatch, so under Run scoping it is correctly not a recipient. It now holds a real current-contract Dispatch in the same adopted Run, which is what the test is named for: one `legacy_direct` and one `current_delivery` recipient in one fan-out. Claude-Session: run-scoped-group-addresses * fix(orchestration): address the Run a nested coordinator created, not its parent A nested coordinator is both a worker of its parent Run and the coordinator of the Run it created. `resolveMessageRun` answers with the parent, correctly, because that is where its own `worker_done` belongs — but audience is a different question. Scoping `@all` to that Run sent a nested coordinator's "shared context" to the siblings it was started beside instead of the workers it started, and reported success, so it never learned its sub-workers heard nothing. Before Run scoping the host-wide fan-out reached the sub-workers by accident; this turned an over-broad delivery into a wrong-audience one, the exact failure class the change exists to remove. Group audience now resolves off the Run the sender coordinates, falling back to its Dispatch's Run. A leaf worker coordinates nothing and is unaffected. This is a separate question from `routing.run`, not a second answer to the same one, so `resolveMessageRun` keeps its meaning for point-to-point mail. Also: when every live Dispatch in a Run is federated, the fan-out skipped them all and threw a bare `Error` that discarded the warnings naming those remote workers and how to address each one. The sender was told "no recipients" while three remote workers existed. That throw now carries a code and the skip explanations. Claude-Session: run-scoped-group-addresses * docs(orchestration): say that no group address reaches a coordinator A coordinator is not a Dispatch, so Run-scoped groups never include one. That follows from the rule, but nothing said it, and the old host-wide meaning did include the coordinator — a worker sending `@all` to raise a blocker would be heard by its siblings and by nobody who can act. The guide, the CLI note, and the docs page now say to use `run:<id>` for that, and that a worker which created its own Run addresses that Run's workers. Also restores the `@cursor` case dropped when the group tests moved: a Claude pane titled "Fix the text cursor blink" must not receive Cursor's mail. That hazard was recorded from real titles and `@droid` alone did not cover it. Claude-Session: run-scoped-group-addresses * fix(orchestration): preserve group audience and mailbox identity * fix(orchestration): validate group scope before dispatch routing * fix(orchestration): preserve pane identity and exclude coordinator dispatches |
||
|
|
373a670aba |
perf(terminal): skip ordinary output in preview control scans (#19400)
* perf(terminal): skip ordinary output in preview control scans * test(terminal): cover scan resumption after parsed controls |
||
|
|
1e2ed20b35 |
perf(terminal): release oversized backing strings behind pending controls (#19396)
* perf(terminal): release oversized backing strings behind pending controls * test(terminal): record reproducible pending-storage gate evidence * perf(terminal): own retained control fragments with a fast copy primitive The pending-control ownership landed with a charCodeAt block copier (10 us at 4 Ki, 170 us at 64 Ki), so it needed a "copy only when discarded output dominates the tail" gate to stay affordable. That gate was the whole cost problem: on adversarial streams it fires every chunk and pays the slow copy (+45% on 16 Ki ANSI chunks, +81..111% on 194 Ki status chunks), and it also skipped ownership on fragments too small to be sliced strings anyway. ownRetainedString replaces it with a Buffer utf16le round trip (0.57 us at 4 Ki, 21.9 us at 64 Ki) and returns anything below V8's SlicedString kMinLength unchanged. Buffer is absent in the renderer and on mobile, so the copier is resolved once behind a lone-surrogate round-trip self-check and falls back to the block copier. With a ~1 us copy the gate is unnecessary: ownership is now unconditional at all three retention sites and the adversarial cases land within noise of the un-owned parsers. The three forced-GC threshold fixtures are replaced by one forced-GC test for the primitive plus deterministic spy assertions that each site routes its retained value through ownRetainedString. All fidelity and differential coverage is kept. * fix(terminal): escape the NUL in the round-trip probe A raw NUL byte in the source made git treat the file as binary, so its diffs and blame were unreadable. Escapes are equivalent at runtime. |
||
|
|
042cc5266c |
perf(terminal): skip plain text between partial escape sequences (#19393)
* perf(terminal): skip plain text between partial escape sequences * perf(terminal): take the ESC at hand before searching for one The unconditional ground-state indexOf regressed dense back-to-back SGR/CSI streams, where the code unit at the cursor is already the ESC and the search pays call plus SIMD setup to find it in place. Check the current unit first and fall back to the native search otherwise. 0.9 MiB dense SGR/CSI medians: 2.04 ms before this PR, 2.56 ms with the unconditional search, 1.92 ms with the hybrid. Sparse colored logs and plain text keep the full search win (0.36 / 0.016 ms vs 1.54 / 1.41 ms baseline). Differential over the full VT alphabet with lone and split surrogates matched 388,416 cases against both prior implementations with zero mismatches; the ground-scan work budget now records 16 inspected code units and 2 native searches. |
||
|
|
23f38ddf7f |
perf(workspaces): skip unchanged heartbeat subscriber work (#19392)
* perf(workspaces): reuse unchanged heartbeat status projections * perf(workspaces): skip unchanged title-sync session collections * test(workspaces): cover constant-size key swaps in status projection reuse Also point the title-input gate at a tracked source file; the previous motivatingLink referenced an untracked .agents skill path. |
||
|
|
db66ab4ac8 |
perf(runtime): skip hibernation inventories without completed agents (#19391)
* perf(runtime): skip hibernation inventories without completed agents * perf(runtime): skip hibernation status scan when no runtime owners * fix(runtime): require host evidence for workspaces resolved mid-inventory `runtimeLivenessRequiredWorktreeIds` was sampled before the runtime inventory await, while the plan is built from the state after it. A workspace that gained tabs or resolved its runtime owner during that window was therefore absent from the required set, so the planner did not demand fresh host evidence for it and fell back to client PTYs — client bookkeeping answering for the execution host. Union the post-await targets into the required set inside `snapshotFromState`. Union rather than replace: the set only ever grows, so the planner can only skip more workspaces, never authorize a hibernation it would previously have refused. An absent inventory stays a skip; nothing reads it as an exited PTY. Extract the coordinator test fixtures so the regression lives in its own file without pushing the coordinator suite past the 800-line test cap. * fix(types): annotate hibernation fixture mock exports for declaration emit TS2883: the inferred `Mock<Procedure>` types of the fixture's exported `vi.fn()` bindings reference `Procedure` from a transitive `@vitest/spy` path that cannot be named. * chore: keep local-file-sink-memory test formatting as on main The merge commit's pre-commit hook reformatted a file this branch does not own. |
||
|
|
7197593e31 |
fix(lint): preserve deliberate collator benchmark baselines (#19686)
Co-authored-by: Merge Sim <sim@local> |
||
|
|
3a801d213d |
fix(orchestration): worker-start settles readiness on observed turn start, not write acceptance (#19423)
* fix(orchestration): worker-start settles readiness on observed turn start, not write acceptance A dispatched PTY worker whose agent wedged at startup (six codex workers on 2026-09-07) was reported 'ok: true, state: ready, stage: input_accepted': the preamble write was acknowledged with observationTimeoutMs: 0 and nothing ever verified a turn began. The corpse and the healthy worker produced identical receipts. worker-start now runs the existing second-stage prompt observer (observeTerminalAgentPrompt) after acceptance, inside the 30s window the client RPC grace already budgets for (orchestration-worker-start-prompt-budget): - turn observed (or provider ack for structured sessions) -> ready - permission prompt -> ready; positive liveness, surfaced in the receipt - provider without a turn-start signal -> ready; observation: unsupported - observation supported and nothing started -> worker state start_unknown, response state outcome_unknown with nextCommands. Honest 'unverifiable', never a death claim: the capability and terminal are kept, and worker-report settlement already reconnects a start_unknown worker that recovers and reports. Also fixes the effect-verb lie that misdirected the first diagnosis of this incident: agent-first worktree creation labeled its own brand-new agent terminal 'reused_agent_terminal' (a role test picking a lifecycle verb) on both the local and federation paths. It now says 'created'; readers keep accepting the retired verb for rows persisted before the rename. * fix: preserve worker authority through start observation --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
72befaf360 |
feat(native-chat): show Claude subagent activity on the shared carrier (#18806)
* feat(native-chat): show Claude subagent activity on the shared carrier Claude's `message:system:task_*` frames are classified `status-chrome` and reach the transcript as nothing at all, so a turn that spawns subagents renders as an idle turn. The journal translator now reads them into the shared subagent-group carrier — no new UI, and the frames stay `status-chrome` so nothing prints a raw opcode row. `local_agent`, `local_workflow` and `local_bash` tasks share that channel and all carry a `tool_use_id`, so `task_type` is the discriminator and a backgrounded `sleep 20` stays out of the roster; `subagent_type` covers releases that predate `task_type`. `skip_transcript` tasks never render, `is_backgrounded` children survive the turn-end sweep, and a resumed task re-announced under a fresh tool id is aliased onto its `task_id` rather than duplicated. A child still reported as working when the turn — or the session — ends becomes `unverifiable`: contact was lost, which is not evidence it exited. * fix(native-chat): stop the Claude subagent roster dropping its own rows The roster published under the same coalescing key it appends the row with, and the sink queue replaces any queued operation sharing a key regardless of kind: once a write was in flight, each new append evicted the pending publish and the next publish evicted that append, so the body never reached the journal and `lastSerialized` had already moved past it. Publish now takes the sink's own slot, as the Codex streams do. A tombstoned row could never come back: the non-batch item-row builder derived its revision from `items` alone, so a re-add was built at revision 1 against a tombstone at 2 and the reducer discarded it forever. It now takes the same `max(items, tombstones)` the batch builder already used — reachable here because an announcement that reveals a `local_bash` task empties and tombstones the group row that a genuine subagent later in the turn reuses. `settleTurn` swept whatever group the key named at the time it ran, so children rostered before any turn key existed were never swept, and a turn whose result never arrives was left working forever. The ending turn's key is now an argument, a superseding turn start settles the turn it replaces, and every turn end also sweeps the outside-turn group. Teardown without an `ended` event, and eviction past the group bound, both lose contact instead of stranding a row at `working`. Label ordinals are a high-water mark now: releasing one on a re-label handed the next child an ordinal that was already on screen. * fix(native-chat): bound subagent-group blocks on every wire that carries one Adding a fifth arm to `NativeChatBlock` made every consumer that assumed four wrong. Two of them ended in `return block`, so they compiled while handing a roster straight through: the mobile RPC sanitizer shipped it unclipped past both mobile char caps, and the legacy transcript import stored an untrusted roster unbounded. Both now clip each label and cap the entry count the way they bound their other blocks. The remaining three sites did not compile at all. The worker transcript payload and the live-session benchmark get real arms rather than casts — a cast would have turned the transcript one into a third silent passthrough inside the wire byte budget — and the CLI worker output renders a roster with its shared summary instead of `[image omitted]`. The mobile sanitizer moves to a sibling module beside the image-block one: the file sat exactly on the max-lines bound, and the block bounds are a self-contained concern with their own caps. Also caps the roster's `subagent_type` label fallback, which reached the journal uncapped, and covers the new block type in the schema audit. * test(native-chat): cover the capped subagent_type label The roster stores the frame's label verbatim, so the cap on the `subagent_type` fallback is the only thing bounding it. * fix(native-chat): type the roster fixture so the suite typechecks The mobile-cap test built its entries with an inferred `state: string`, which is not a `NativeChatSubagentState` — the only typecheck failure on the branch. * fix(native-chat): stop child traffic rostering an id Claude never announced `observeChildActivity` minted a provisional row for any `parent_tool_use_id` outside the excluded set. An id that was never announced is never excluded, so a nested Task, a workflow child, or a grandchild parented to a tool id inside the sidechain each produced a permanently unlabelled `subagent` row that could only ever end `unverifiable`. The bounded exclusion set cannot cover an id no frame ever declared, and in a long session it can forget a genuine exclusion. Track instead whether this CLI announces tasks at all — set by ANY `task_started`, including one the subagent filter rejects. Once it has, an undeclared child is provably not a new subagent, so no row is created. The provisional path now serves only releases that announce no task frames, which is what its comment already said it was for. The label-ordinal test moves to an announcement-driven removal, the scenario that path now actually reaches; it still fails if `remove` releases the ordinal. * fix(native-chat): outrank the tombstone when building one too `buildJournalTombstoneRow` still built its revision from `items` alone, leaving it asymmetric with the item builder. It is correct today only because `upsertItem` clears the tombstone whenever a re-add wins — an invariant that lives in the reducer and was not pinned. Apply the same `Math.max`, and pin the invariant so the reducer cannot drop it silently. * fix(native-chat): stop an unrelated turn end settling an outside-turn child `settleTurn` swept the `outside-turn` group on every turn end, so a child Claude announced while no turn was live — a frame trailing the previous turn's result, or one that arrives before the first turn starts — was marked `unverifiable` by the next, unrelated turn ending. That state is terminal and latches, so the `task_updated: completed` that followed was discarded: loss of contact was recorded as the child's outcome on evidence that was never about that child. A turn end now sweeps exactly the group its key names. `outside-turn` belongs to no turn, so only an end with no key of its own reaches it, and what no turn end reaches `settleSession` does — reliably, since teardown without an `ended` event also routes through it. The cost is a child outside every turn showing `working` a little longer; the alternative prints a wrong outcome that nothing can revise. Also pins that a subagent announced after a task the filter rejected still rosters: the announcement path was never what the child-traffic gate closes. * fix(native-chat): bound a subagent entry's id, not just its label Every site that bounds a `subagent-group` block clipped the label and handed the id through whole. From the Claude producer the id is bounded upstream, but the legacy transcript import reads an untrusted file, so an oversized id survived into the journal and then out to every wire that replays it — 64 entries of it, since only the entry count was capped. Each site now clips the id with the helper it already uses for its other bounded fields: the journal's inline-text bound on import, the transcript payload's metadata clip, and the mobile char cap (renamed, since it is no longer a label-only cap). * fix(native-chat): surface an adverse subagent outcome in the fallback sentence The roster row's plain-text stand-in counted only `working`, so a fan-out whose children all latched `unverifiable` (or `failed`, or `stopped`) rendered as "Ran 3 subagents" — a completion claim. Mobile and paired web have no roster renderer, so that write-time-frozen sentence is the entire row there, and collapsing `unverifiable` into something that reads like success is exactly what the SSH execution boundary forbids. It now appends the worst adverse count, worst-first across failed/stopped/ unverifiable, and shows it even while siblings still work — matching the Codex lane's shared `subagentGroupFallbackText` verbatim so collapsing the two copies later is a deletion, not a behaviour change. Also bounds the provisional entry id. `observeChildActivity` wrote the `parent_tool_use_id` straight into the entry's durable id with no length cap, while the announced path already rejects an over-long id via `claudeTaskId`. Both now share `isBoundedClaudeTaskId`, and the provisional path rejects rather than truncates, as the announced one does. * fix(native-chat): stop a subagent label ordinal and a clipped roster key colliding - claimLabel probes the labels the group actually rendered instead of a per-base counter, so a generated `Audit 2` cannot duplicate a provider's own `Audit 2`. - Bound `NativeChatSubagentEntry.id` with a head plus a digest of the whole id at every site that bounds it. The id is the roster key: a prefix clip merged two distinct children onto one entry. - Correct a stale journal-reducer test comment: tombstone cleanup is a map-state invariant, no longer load-bearing for revision ordering. * fix(claude): merge duplicate unhandled-provider-frame imports * fix(claude): preserve subagent lifecycle and bounded invocation identity --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
c3a415487b |
perf(terminal): avoid quadratic status-marker scans in output bursts (#19373)
* perf(terminal): reuse forward OSC status terminator searches * test(terminal): pin cached OSC terminator reuse across BEL frames Document that forward match reuse requires a monotonic search offset and cover a distant ST held across many intervening BEL frames. |
||
|
|
8f78c28248 |
fix(orchestration): fence worker release on mobile keystrokes (#19337)
* fix(orchestration): fence worker release on mobile keystrokes A settled worker's terminal stayed ownership_state='owned' unless a takeover was recorded, and the only recorder was orchestration.workerTerminalUserInput, which only the desktop/web xterm input signal and the native-chat composer call. Mobile input arrives as terminal.send / stream input frames instead of a report, so a phone user typing in a settled worker's pane never fenced anything: worker-list kept recommending release and worker-release closed the PTY under them. Give the host one definition of "a human typed into this terminal" and route every lane through it. The mobile input floor claim is that definition and already exists on both byte lanes: it is taken only for deliberate phone input, never for the emulator's own query replies, and never for an agent's `orca terminal send`, which names itself a desktop client and so is indistinguishable from a keystroke at this layer. Settling that claim after an accepted write now records the takeover through the same code the RPC reporter uses, throttled to one write per pane per 30s so a keystroke does not pay for an immediate transaction. The record lands on the runtime that owns both the terminal and the orchestration database, so SSH-hosted and remote workers behave exactly like local ones. No mobile change: mobile already sends client.type (mobile/src/terminal/terminal-send-request.ts:24). * fix(orchestration): ask the database, do not remember, whether a pane is fenced The keystroke throttle armed on the attempt rather than on the outcome, so a zero-row or thrown record poisoned the pane for 30s. A phone keystroke during the worker-start readiness wait lands before prepareStartingWorkerAuthority creates the owned resource; a real keystroke seconds later was then suppressed, the worker settled, and workerRelease closed the terminal under the phone user. A SQLITE_BUSY on the first write did the same, with no retry. The cache was the defect, not its arming condition. Its precondition is the set of owned resources on the pane, which changes underneath it, and any cache keyed on ownership identity would have to read the database to learn that identity -- which is the whole question. So the input lane now asks: a read using the same predicate the write uses answers "is anything still fenceable here?" without taking BEGIN IMMEDIATE, and only then is the write attempted. Ordinary typing costs a lookup instead of a write lock, a failed write is retried by the next keystroke, and a takeover writes once per ownership epoch rather than once per window, because the flip to user_owned removes the pane from the candidate set. Sharing the predicate keeps the probe from drifting from the writer. Adds the two escape cases as permanent regressions, drives the mocked send through the real RuntimeTerminalWriter, and asserts a mobile takeover lifts the settled-worker resume fence, which no test covered. * refactor(orchestration): let the database dedupe the takeover, drop the read probe The probe was meant to keep keystrokes off BEGIN IMMEDIATE, so it had to earn that with a number. Measured against a real WAL database it costs more than the write it avoids: at 25 live workers the probe is 0.19ms and the no-op write is 0.10ms, because the probe runs the same candidate selection with each statement taking its own read snapshot instead of sharing the transaction's. It is a compensating mechanism with negative value, so it goes, along with the database method and the predicate extraction it needed. owned -> user_owned is one-way and scoped to a resource, so the database is already the dedupe: every deliberate human write attempts the transition, the second attempt matches no row, and the fence sweep runs only on changed > 0. Nothing is remembered between keystrokes, so no state can outlive the ownership it described -- a keystroke before the worker's authority attaches, a write the database refuses, and a re-dispatch onto the same pane all resolve against the rows as they are at that instant. An attempt costs about 0.1ms at typical fleet size and 0.34ms at 100 live workers, on mobile writes only. Replaces the write-count test, which asserted the old mechanism, with the invariant: many keystrokes settle into one takeover and one fence sweep. Adds the re-dispatch case, where a pane's next worker is fenced on its own merits. * refactor(terminal): name the provenance rule the takeover fence hangs off The fence rode the mobile input floor claim, with only a comment tying the two together. The floor is arbitration -- who may write next -- while the fence needs provenance -- who produced the bytes. They agree today, so anyone reweighing the floor would have moved the fence without noticing. isDeliberateHumanInput states the provenance rule on its own terms, and both byte lanes decide with it when they open a write: the claim carries the verdict beside the handle, and settlement records the takeover only when a human produced the bytes. No behavior change -- afterWrite is wired only where the predicate already answers true -- and the rule is now pinned by its own cases, so a future arbitration change has to answer this question again rather than inherit it. * test(orchestration): prove the unary lane classifies a metadata-less phone A phone build older than client.type is recognised only by its pane's mobile driver, which the unary lane passes as the provenance evidence. Nothing proved it did: replacing that argument with false left all 17 tests green while a shipped phone silently stopped fencing worker release. The new case drives a clientless send on a mobile-driven pane and fails under that mutation. The stream lane now passes false outright. Its isMobile is read off the same client object it carries, so the metadata-less phone cannot reach it, and passing the flag suggested a legacy path that does not exist there. Also states what the per-keystroke cost scales with. A pane owning no resource misses the pane_key index and falls through to a scan of owned resources, so the figure is tens of microseconds at realistic worker counts rather than a flat 0.1ms, and it grows with rows that are never released. * fix(terminal): let provenance alone decide the takeover, on every accepted write A phone older than client.type sends no client metadata, and both stream initializers derive isMobile from that metadata alone, so such a subscription reported false and took the stream lane's uninstrumented branch: provenance was computed and then never consumed. Bytes from a real person landed through both frame adapters and the resource stayed owned, so workerRelease closed the PTY under them. The unary lane already fenced that population off the pane's mobile driver, which is the host's standing reading of clientless input, so the two byte lanes disagreed at the destructive boundary. The predicate was still subordinate to floor plumbing: it could only be consulted where a floor client id existed. Now the accepted-write callback attaches on both lanes regardless of whether a floor was reserved, and humanInput alone decides recording; a write holding no claim commits nothing. Arbitration keeps its own condition around reserveWrite, where it belongs, and the unary lane's duplicate outer provenance filter is gone. The stream lane reads clientless provenance from the pane's driver, the same policy the unary lane uses. The claim holder is now TerminalInputWrite, carrying the verdict beside an optional floorClaim, so the structure says what the doc said: a write may fence without holding the floor. Regressions drive both real frame adapters, clientless direct delivery, and the paired-web desktop negative. Metadata-only provenance fails 3 on the stream lane and 1 on the unary lane; gating the callback on a reservation fails the same 3. * fix(runtime): resolve retained handles before mobile input provenance A renderer reload clears transient handles while retaining runtime-owned PTY identities. Legacy mobile provenance saw no leaf, then sendTerminal restored the same handle and delivered an unfenced key. Normalize through getLivePtyForHandle at the shared live-leaf resolver entry so classification and writes agree, preserving existing leaf generation/incarnation checks. Caller audit: - terminal-send-method: driver, query-reply authority, lock and floor checks now resolve the retained PTY before sending. - terminal-input-delivery: legacy mobile classification and exact-PTY binding now see the same target as the writer; equality checks remain. - terminal-multiplex-subscribe-resolution: retained PTYs resolve directly without a spurious missing-terminal wait. - terminal-lifecycle-methods resize and terminal-viewport-methods display mode, restore-fit and updateViewport retain their original PTY target. - inspectTerminalProcess: avoids false terminal_gone after reload while preserving provider inspection and incarnation fences. - getLivePaneKeyForTerminalHandle and getOrchestrationDispatchAuthority: unaffected because both already call getLivePtyForHandle first. No wire/schema changes, host fallback, process-death inference, or Git workspace assumptions; SSH providers keep ownership of execution evidence. Validation: - Unmodified round-3 reviewer probe: reproduced 2/2 failures, then 2/2 pass. - Unmodified round-2 reviewer probes: 13/13 pass. - Checked-in takeover suites: 24/24 pass. Removing only the resolver call fails both new reload cases; source restored afterward. - RPC orchestration + terminal, aggregate runtime handle registry, handle incarnation, mobile tab mount, stale geometry, and reload probe: 2027 passed, 1 skipped (89 files). - tc:node and check:code-quality:changed pass; background launch enabled. * test(rpc): require unconditional terminal afterWrite callbacks Update exact sendTerminal expectations for the round-2 accepted-write contract. Preserve beforeWrite expectations, absence of reserveWrite, byte payloads and call-count checks; require afterWrite to be a function. Reproduced the requested two-file run: 5 failed, 31 passed. The full RPC suite exposed the same stale shape in ACK budget/overflow, desktop resize (including its later retry), and agent-prompt fallback assertions. Update those too, for 11 assertions across six test files. No production changes. Validation: ORCA_BACKGROUND_LAUNCH=1 full src/main/runtime/rpc suite: 264 files passed; 2292 tests passed, 1 skipped. Changed-code quality and staged oxlint/React Doctor/oxfmt checks passed. Ran lint-staged --no-stash manually to honor checkout safety rather than its default backup hook. * fix(mobile): report worker takeover outside terminal byte delivery New phones announce accepted real user input through the existing worker report RPC, addressed by terminal handle. Share a per-client/per-handle 30-second gate with one bounded retry; report through the same RPC client as the input. Cover live commits and dictation via their shared sender, accessory keys, gestures, buffered submit, paste and accepted native chat. Query replies, attachment heals, triage and diff-review sends do not report. Phones predating this build do not fence release. Remove byte provenance and takeover callbacks from host delivery. Restore both lanes' pre-PR floor-claim plumbing and the original options assertions. Keep the host recorder uncached with its conditional resume-fence sweep. No DB schema or stream change; terminal is an optional report address. Retain the shared resolver recovery independently of takeover: the new SSH inspection test fails without it during renderer reload. Other callers still benefit for subscription, resize, viewport and exact-PTY binding; unary driver/lock checks see the retained PTY. Pane routing and dispatch authority already recover through getLivePtyForHandle and are unaffected. Existing leaf generation checks and first-PTY adoption remain unchanged. No other input-plumbing hunk is retained relative to the PR base. Replace byte-takeover tests with handle-addressed local/SSH report and unknown-handle tests, plus real unary/stream writes asserting zero SQL prepare/exec calls. Mobile send-site integration covers reports, exclusions, rejected writes and gate counts. Desktop report tests are unchanged. Register replacement coverage in the settled-worker release manifest. Validation (all background): host/RPC/runtime 3541 passed, 2 skipped; mobile session/terminal 2045 passed; node and mobile typechecks, changed quality, mobile oxlint, reliability manifest and max-lines ratchet passed. All five requested mutations fail assertions; resolver revert also fails independent inspection. Staged checks run manually with --no-stash. Final src diff against PR base: 5 files, +165/-13 (previously +839/-85). * fix(runtime): allow the takeover report from mobile-scoped tokens The mobile RPC allow-list gates every phone request before dispatch and the reporter swallows a refusal, so without this entry every phone shipped unfenced. Pin it beside the report tests, and pin the once-per-takeover fence sweep the replaced byte-lane suite used to assert. * fix(mobile): a no-op takeover report does not arm the gate; Stop reports too A key during worker startup reports before the resource is owned; caching that zero-change reply for 30 s suppressed the report that would have fenced the worker once it attached. Native-chat Stop is deliberate input and now reports on an accepted Escape. * fix(mobile): takeover gate ignores the host answer, like desktop Reopening the gate on a zero-change reply made every accepted key on an ordinary terminal an RPC plus a host write transaction (round 6: 100 for 100). The startup window it closed is unreachable: the agent has no prompt to accept input until after its resource row exists. Plain terminals now pay one report per 30 s window; the native-chat Stop report stays. Send-site fixture answers the report RPC with a changed count; the draft test filters to terminal.send calls. * docs(runtime): say why resolveLiveLeafForHandle re-links before lookup * chore(i18n): regenerate the runtime-required catalog for the contrast floor strings * test(orchestration): give the stopping-worker guard fixtures a Run * test(orchestration): drop fence-sweep assertions retired by the settled-worker policy * test(orchestration): pin the mid-boot phone takeover that #19608 makes possible A handle-addressed report during the worker's tui-idle wait now finds the custody row written at terminal creation, so it flips the pane to user_owned and worker-release retains it instead of closing it under the user. |
||
|
|
12f53da542 |
Remove settled-worker automatic resume and hibernation fences (#19544)
* Remove settled-worker automatic resume and hibernation fences * test: retirement rollback case follows the no-fence policy Case 4 seeded and asserted automaticResumeBlockedBy, which this branch deletes. A rolled-back settled worker is now an ordinary done record that wake clears as passive evidence, same as any finished agent pane. * chore(i18n): regenerate the runtime-required catalog for the contrast floor strings * test(orchestration): give the stopping-worker guard fixtures a Run |
||
|
|
98fdbc4ade | fix(orchestration): file federated worker mail under the coordinator Run (#19542) | ||
|
|
ed881889c4 |
fix(terminal): fold-safe CAN/SUB and double-ESC handling in partial-escape tail (#19521)
`extractPartialEscapeTail` broke its own fold invariant
(extract(a + b) === extract(extract(a) + b)) in the oscEsc/stringEsc states,
so a PTY read that split there produced a different pending tail than the
same bytes delivered whole — the tail snapshots append after a restore.
Two causes, both in the "ESC did not terminate the string" branch:
- CAN/SUB were routed through `stateAfterEscByte`, which maps them back to
`esc` instead of aborting to ground. `extractPartialEscapeTail('\x1bPx\x1b\x18X0abc')`
returned '\x1b\x18X0abc'; the chunk-split fold returned ''.
- A second ESC opened its new sequence at `i - 1` rather than at itself.
`extractPartialEscapeTail('\x1b] \x1b\x1b^')` returned '\x1b\x1b^' whole but
'\x1b^' folded. The fold was right — xterm starts the sequence at the second ESC.
The existing fuzz only asserted the fold as `advance(extract(pending), chunk)`,
which is a tautology because every PENDINGS entry is already a tail. Replaced
with a sweep that re-splits the combined stream at every code-unit boundary,
and extended the alphabet (NUL, 0x20 intermediate, CJK) and SEQUENCES with
CAN/SUB and doubled-ESC-inside-string cases. A 1.25M-split fold fuzz over a
VT alphabet goes from 421 failures to 0.
|
||
|
|
28d936e030 |
perf(native-chat): bound journal reads during paged catch-up (#19360)
* perf(native-chat): bound journal reads during paged catch-up * perf(native-chat): reduce the journal once per catch-up run, not per page Bounding the SQL read per page left the JS side still O(total items) per page: every page re-reduced the whole timeline and rebuilt the live-item map, and the byte-shrink loop rebuilt it again on each halving. A catch-up run is a synchronous loop with no await between pages, so the reduced timeline is loop-invariant. `createAgentSessionCatchUpReader` holds one snapshot for the run and re-reduces only if the journal cursor actually moved, and the projection's live-item / alias / submission-byte indexes memoize on the snapshot arrays the reducer rebuilds on change. Per catch-up over a 8,000-message backlog: 40 timeline reductions to 1, reduce+project time 24.3ms to 3.5ms, end-to-end 134.6ms to 111.3ms. |
||
|
|
31b84cc71f |
Fix native PTY I/O failures disabling session termination (#19523)
* Fix native PTY I/O failures disabling session termination I/O failures on write or resize were incorrectly treated as exit evidence, which disabled further termination attempts and producer flow control. Separate ioFailed state from dead state; I/O errors suppress operations but keep kill, forceKill, and signal available. Publish physical exit before notifying listeners to prevent reentrant cleanup attempts from accessing the retired native process. * Fix PTY I/O cleanup tests and config syntax error - Fixed missing closing brace and comma in reliability-gates.jsonc - Added clear() operation to PTY mock fixture and test coverage - Enhanced assertions to verify operation suppression during I/O failures - Improved kill operation error handling with better promise-chain assertions - Updated test result summaries in reliability gate documentation |
||
|
|
e182930670 |
test: cover input in five simultaneously flooding SSH panes (#19071)
* test: cover keyboard input in five simultaneously flooding SSH panes * test: capture pane focus and buffers on flood input failure * test: capture pane focus and buffers on flood input failure * test: capture pane focus and buffers on flood input failure * test: record replay input loss and application fix dependency * test: record merged replay-input fix in the five-pane flood gate |
||
|
|
aeddfa463d |
perf(renderer): avoid per-second spinner animation events (#19407)
* perf(renderer): avoid per-second spinner animation events * fix(bench): ensure the Electron runtime before bench:spinners The script launches Electron via Playwright but skipped ensure:electron-runtime, which every other Electron-launching bench script runs first. * docs(renderer): scope spinner pixel-tolerance claim to paused-animation checks --------- Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local> Co-authored-by: pullfrog[bot] <226033991+pullfrog[bot]@users.noreply.github.com> |
||
|
|
5bd0247aaa |
fix(xterm): remove scrollback decorations by identity (#13178)
* fix(xterm): fire Marker dispose before clearing line (#10879) Scrollback trim under search highlights was O(k²) because dispose set marker.line to -1 before onDispose, collapsing SortedList keys. Fire listeners first so delete still sees the real line, then clear the line. Fixes #10879 * fix(xterm): remove scrollback decorations by identity * perf(xterm): avoid index arrays for unique decorations --------- Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> |
||
|
|
cbc7bb418c |
fix(ci): read the changed-path list past the first pipe buffer (#19409)
`pr-code-change-scope.mjs` read stdin with `readFileSync(0, 'utf8')`, a single read of fd 0. Once the writer outgrows the 64 KB pipe buffer that read returns early or throws EAGAIN, the script exits 0 having emitted no `name=value` pairs, and `tee -a "$GITHUB_OUTPUT"` records nothing -- so every lane the classifier gates is silently skipped rather than failing loudly. A PR opened long ago carries a stale `pull_request.base.sha`, so the gate's merge-base diff spans the whole base branch. PR #13178 diffed 13,294 files (773 KB) against a base 1,592 commits behind main and lost typecheck, test, static analysis, xterm patch sync, package and e2e to this. Stream stdin instead, matching how the sibling `pr-e2e-source-routing.mjs` already reads the same list in the same workflow. |
||
|
|
070fcf381a | fix(diff): retain native selections and restore visible selection colors | ||
|
|
ad6a51eca5 | Merge origin/main and preserve Shift-wheel scrolling in Pierre diffs | ||
|
|
cfe243af69 | fix(diff): restore read-only and original-side search with cancellable matching | ||
|
|
46bdddb53a | fix: preserve diff note gutter eligibility and labels | ||
|
|
b51d1780d1 | fix(diff): keep parsing, highlighting, and large edits responsive | ||
|
|
2265fce591 | chore(mobile): remove stale max-lines exceptions (#19366) | ||
|
|
9fed61e5c2 |
Persist agents sidebar search visibility as pairing-local preference (#19313)
* Persist agents sidebar search field visibility as pairing-local preferen - Add `agentsShowSearch` to workspace UI state with default on - Include in pairing-local fields so preference syncs across clients - Convert search from menu action to checkbox menu item for explicit toggle - Update activity thread options menu to reflect checkbox state - Add localization strings across all supported languages - Update RPC schemas and preference persistence layer - Includes readiness validation reports confirming feature is clean * rm review * fix documentation |
||
|
|
c37413271e |
perf(mobile): open a session with parallel startup RPCs and a pre-warmed terminal engine (#19260)
Startup RPCs now fan out in parallel and the xterm engine pre-warms inside the real terminal frame while they are in flight, so the first pane inherits a warm WebView and an already-measured viewport instead of paying a round trip for it. The pre-warm opens its engine before measuring: web-ready only reports that the bundle loaded, and the WebView answers a measure with null until a terminal exists. It also pre-warms at the user's saved text size, because cell size is what the frame height gets divided by. Host writes such as worktree.activate wait for an evaluated status.get reply. Navigation still fails open when a host cannot answer one, but that fallback no longer reads as a passing compatibility verdict. |
||
|
|
a899f92402 |
feat(windows): enable structured Codex chat on native Windows (#18519)
* feat(native-chat): enable Windows structured sessions
* fix(codex): prove native Windows process identity
* style(codex): format Windows session seam
* fix Windows structured Codex admission
* fix(windows): reprobe missing process identity capability
* fix(windows): decide folder-workspace WSL routing before the click
Review found pathUsesWslUnc exported but unused, and the folder composer
hardcoding worktreeUsesWslPath:false. Together those meant a folder picked
under a \\wsl.localhost\ parent routed to structured chat, then got refused
by the host and fell back AFTER the click -- which defeats the lane's own
design goal that create cannot fail after the click.
The group's parentPath is in scope at submit and the workspace is created
under it, so the parent decides WSL-ness pre-click. Wires pathUsesWslUnc
there and adds tests for the helper, including the unhydrated-store case
that previously threw.
* fix(windows): collapse the gate derivation to one call, restoring max-lines
CI static analysis failed: launch-agent-in-new-tab.ts crossed the 300-line
oxlint ceiling. Adding a max-lines disable is forbidden, so the two gate
derivations collapse into one readWindowsStructuredGateInputs() call --
a store-backed site now adds one line and one import name instead of two.
Better shape anyway: one derivation entry point rather than two reads a
call site must remember to pair.
* fix(windows): engage the legacy fallback when the host THROWS a refusal
Review found a P1 this merge composes: neither parent could reach it. At the
lane head the only structured entry was launch-agent-in-new-tab (full
store-backed WSL check); on main all win32 was refused. The merge enables
win32 in creation flows that pass no projectRuntime, so a WSL folder
workspace, a WSL-configured repo, or a repair-required runtime now routes
structured -- and the host refuses correctly, but by THROWING rather than
returning {ok:false, refusal}.
Callers engage their legacy-terminal fallback on the refusal CLASS, so an
unmapped throw arrives as a generic RPC rejection: no fallback, empty
workspace, error toast, prompt stranded in the launch outbox. Pre-merge the
same action opened a legacy terminal agent.
Map the host's thrown definitive refusals onto the refusal class at the
launch boundary, so every creation flow -- present and future -- degrades to
the legacy terminal instead of stranding. Narrow predicate: unrelated
failures (ECONNRESET, empty message, non-Error) still propagate untouched.
Ablation-proven: removing the mapping reddens the fallback test.
* fix(windows): teach the mobile RPC double the status probe the lane added
CI's first-ever run on this lane caught a pre-existing lane defect. The lane
changed status.get to resolve through
runtime.getStatusAfterWindowsProcessStartTimeProbe(), but never taught the
mobile-surface runtime double about it, so status.get failed for mobile
clients with "not a function". The lane's own test list did not include this
file and the lane had zero CI, so nothing ever ran it.
The real runtime always implements the method; the double omitted it.
* chore: merge current main and regenerate the localization runtime catalog
CI static analysis failed on a stale en-runtime-required.json: main added
onboarding integration-capability keys, and the generated catalog is checked
against the PR MERGE result, not the branch alone -- so it read clean locally
while failing in CI. Merging current main (
|
||
|
|
314506003a |
fix: retain MSYS shell descendants in their terminal job (#19068)
* fix: retain MSYS shell descendants in their terminal job * test: complete MSYS regression CI registration and teardown contract * fix(windows): deny job breakaway for the whole Cygwin/MSYS shell family The per-PTY job probed only msys-2.0.dll, and only for bash.exe/sh.exe. Cygwin ships the same spawn.cc breakaway logic under cygwin1.dll, and an MSYS2 zsh escapes exactly like its bash does, so both kept the orphan bug. Probe the runtime DLL on the shell's own search path instead of matching shell names: that is the property that decides whether the runtime will ask for CREATE_BREAKAWAY_FROM_JOB, and it drops the name special-casing. * chore(patch): restore the conpty.cc index line The earlier hand-edit dropped it while every sibling section kept one. Recomputed against the real blobs: applying this patch to 7b286d3d yields exactly 4b06d185, so git apply -3 has its fallback back. |
||
|
|
2ccf35b135 | fix: avoid quadratic trimming during fullscreen terminal redraws (#19214) | ||
|
|
fb322046e8 |
skills: rewrite and trim the seven non-orchestration guides (#19128)
* skills: rewrite the seven non-orchestration guides to one outcome-first standard
Every guide leads with Result / Done / Safe failure, states conditions instead of case lists, keeps one done bar and one autonomy envelope, and loads references at the point of use via `skills get <topic> --full`. orca-cli drops from 424 to 260 always-loaded lines with three references; orca-per-workspace-env from 794 to 397 with five.
Defects fixed in shipped guides: `emulator camera` (no such command), iOS `permissions` (backend refuses it), Android pane described as in development, `relayGracePeriodSeconds: 0` documented as immediate teardown (it is unbounded), doctor `ok: true` hiding `warn`, an SSH exemplar setting both `jumpHost` and `proxyCommand`, a provisioned-root fetch from `origin`, and the Linear unconfirmed-write rule keyed on four verbs when ten emit it.
The resolver ladder, placeholder rule, and older-binary fallback shared by every installable SKILL.md now come from one skill-stubs/_shared/cli-resolution.md fragment composed by the generator, which also bundles per-guide references into --full. New guards: every ORCA invocation and flag resolves against COMMAND_SPECS, descriptions carry no angle-bracket tokens, reference routing is checked both ways, and an always-loaded size ratchet (300 lines) that guides may leave but never join.
* skills: address review on the SSH recipe and the parity guard
- ssh-host create script: route the bootstrap ssh through the chosen jump host or proxy command, refuse both at once, use StrictHostKeyChecking=accept-new instead of a blind ssh-keyscan append, and pass gh_token/project_root/repo_url/repo_ref to the remote bash via printf %q so a quote in a value cannot break out of the command.
- per-workspace-env envelope: the step-10 workspace test the user asked for is no longer forbidden by the same paragraph.
- linear guides: name the full verb, ORCA linear list-issues.
- parity guard: a prefix reference such as ORCA linear --help or ORCA emulator --webcam now has its flags checked against every command under that prefix; only an exact path or an explicit ... was checked before.
* skills: tighten prose in the seven rewritten guides
Shorter outcome spines, one idea per sentence, no restated rationale after a rule. No rule, command, or pinned phrase changes; 47 net lines fewer across the guides and references.
* skills: route orca-cli and per-workspace-env gates through --reference
Both guides told agents to load --full at a gate because the per-reference
selector did not exist when they were written. Now that main serves
`skills get <topic> --reference references/<file>.md`, load only the
named file and keep --full as the fallback for an older CLI, matching the
orchestration kernel.
* skills: drop outcome-spine boilerplate from the CLI-wrapper guides
The Result/Done/Safe-failure preambles and Next Action closers restated
rules the body already carries. Agents stop fine without them, and for
a CLI wrapper the command surface is the guide. Keeps the one substantive
rule computer-use's Done block added (never report unverified as success)
inside Action Rules. orchestration and per-workspace-env keep theirs:
those are multi-step workflows where the done bar is load-bearing.
(cherry picked from commit
|
||
|
|
1848855515 |
test: enable software WebGL for Linux CI headful specs (#19001)
* test: enable CI WebGL and route GPU-dependent regressions * test: retain headful atlas cases in terminal rendering goldens * test: reuse golden command in project coverage assertions |
||
|
|
1e301ab1df |
test: cover native Wayland Hangul in isolated CI (#19174)
* test: exercise native Wayland Hangul in isolated CI session * test: wait for nested compositor socket before selecting IBus * test: align Wayland IBus discovery with GNOME environment filtering * test: assert Wayland launch and register native Hangul evidence |
||
|
|
2e8fa3fe9b |
test: exercise packaged browser compatibility in scheduled CI (#19157)
* test: exercise packaged browser compatibility in scheduled CI * test: record final packaged workflow participation evidence * test: expose manual packaged revision and simplify executable check * test: reject missing package checksum assertion |
||
|
|
20eea184cc |
feat(native-chat): offer the link-action popover for chat links (#19130)
* feat(native-chat): offer the link-action popover for chat links A plain click on an http(s) link in a native chat transcript opened the system browser outright, ignoring the link-routing preference the same link honors in the terminal. Chat now shows the terminal's destination popover, with the modifier chords routing straight to a destination. The popover, its request type, the destination policy and the routed open move out of terminal-pane so both surfaces share one implementation; the catalog keys keep their original namespace because they carry shipped translations. Chat resolves its link owner from the session workspace (runtime, then SSH, unresolved stays unknown) so a remote transcript only offers Orca Browser when that host's managed browser route is eligible. The existing toggle now governs both surfaces, so it is retitled; with it off a chat link still opens on a plain click instead of going dead. * Fix native chat link popover lifecycle and keyboard anchoring * test(native-chat): use one store mock for link actions * fix: update reliability gate for shared link popover tests --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
4120501979 |
perf(store): detect Zustand rerender churn the current audit cannot see (#19059)
* perf(store): detect Zustand rerender churn the current audit cannot see The app-store-performance audit only understood inline selectors passed to a hook imported literally as `useAppStore`, so three shapes went unlinted: - a selector referenced by name (`useAppStore(selectRows)`), including one hoisted below its call site — resolved now via a Program:exit pass - the sibling store hooks (`usePluginPanelsStore` and friends), matched by the use<Name>Store convention on local imports; React's `useSyncExternalStore` matches that shape and is excluded - a fresh reference nested inside a `useShallow` projection, which is the worst case of the three: the comparator runs on every write and can never match, so the memo silently buys nothing `no-nested-fresh-under-shallow` covers the last one. `src` is clean against all four rules today, so this is a ratchet rather than a cleanup. The write side stays undecidable statically — whether a `set()` reallocated for nothing depends on the payload — so it gets a runtime probe instead. withStoreIdentityChurnProbe counts writes that replace a field's reference while its value stays equal, and can name the calling site. Cost when disarmed is one boolean load per write, matching react-commit-cascade-write-probe. * perf(store): scope the churn probe's scan to the write's own keys recordWrite iterated Object.keys of the full post-write state, so the armed cost scaled with the store's top-level field count (hundreds) rather than the size of the write. `set(partial)` merges, so no field outside the partial can have changed. The wrapper now resolves a functional updater itself and iterates the resolved partial's keys. Same function, same argument, called once — there is a test pinning that, since calling it twice would double any work a slice does inside its own updater. A replace write drops absent fields, so that path still scans every field. Disarmed cost is unchanged: one boolean load. * perf(store): follow a selector one hop into its helper Review feedback: both the lint rule and the manual sweep it was checked against only looked at the inline selector body, so neither could see a fresh allocation made inside a helper the selector calls — and delegating to a module-scope helper is the idiomatic shape here. Two methods sharing a blind spot is not corroboration. The two fresh-reference rules now resolve a single hop into a module-scope helper. The predicate used across that hop is deliberately stricter than the inline one: it requires EVERY returned expression to allocate unconditionally, so the common `cache.get(k) ?? buildFresh(state)` identity-caching shape is not flagged. An unresolvable helper is left alone rather than guessed at. Still zero hits across 20,330 files, so this stays a ratchet. * perf(store): keep the churn probe off the shipped write path Review hardening for the churn probe and the widened lint rules. Probe: it no longer resolves a functional updater itself. Zustand keeps sole ownership of when and with what argument an updater runs, so the middleware cannot double-invoke it or hand it a stale state. Object partials still scope the scan to the write's own keys; updater and replace writes fall back to the full field list, which costs one Object.is per untouched field and nothing more, since the deep compare only runs on replaced references. store/index.ts installs the probe only when import.meta.env.DEV or e2eConfig.exposeStore is set, the same gate as __store exposure. Nothing in the app arms it, so a shipped build was paying a wrapper frame per write for a diagnostic it could never read. The cascade probe stays unconditional because crash telemetry arms it in the field. Site capture now skips any *-probe.ts frame; under the real composition the first non-node_modules frame was the cascade probe's wrapper, so every churn was attributed to react-commit-cascade-write-probe.ts:32 instead of the caller. Plugin: named-selector recording is restricted to module scope. A component-local `const selectRows = ...` used to overwrite the entry for a same-named imported selector and flag an unrelated useAppStore(selectRows). The any-branch and every-branch allocation predicates are one function with a flag, the Object.* static list is a Set, and import recording is a single pass. Tests: updater called once with live state, identical-state writes ignored, disarmed path forwards exact arguments without calling get(), full composition with the cascade probe (no drop, no double, correct site), and the module-scope shadowing case for the plugin. |
||
|
|
b8311d509a |
Revert "skills: rewrite the seven non-orchestration guides to one outcome-first standard (#18724)" (#19126)
This reverts commit
|
||
|
|
1478101342 |
fix(windows): unblock structured native chat by exposing process creation time (#18986)
* fix(windows): guard process creation times
* fix(windows): ask the relay's bare addon for creation times too
The relay addon build now emits creationTimeMs, but the runtime binding
for the bare addon still declared only CommandLine, so a Windows relay
host requested flag 2 and every row came back without a creation time.
That leaves captureWindowsDescendantSnapshot returning null and
verifyWindowsProcessIdentity false forever on those hosts -- the relay
half of the patch was unreachable.
Naming CreationTime in the adapter is safe because the bare addon is a
content-hashed relay artifact: it ships in the same immutable relay
directory as the bundle reading it, so it can never be older than the
code asking for the bit.
Also bound the win32 guard test on our own row, which the addon can
never fail to answer, so an unconverted FILETIME or a 1601-epoch stamp
fails instead of satisfying a bare count.
* fix(windows): make the compiled addon prove its own CreationTime support
CI caught the real defect: the win32 guard test read
isWindowsProcessStartTimeAvailable() as true and then found 0 rows
carrying creationTimeMs. Unlike node-pty, this package publishes a
prebuilt .node at the same build/Release path node-gyp writes to, so
pnpm patches the source tree and leaves that binary alone. A host then
holds a patched lib/index.js -- ProcessDataFlag.CreationTime and all --
over a binary that ignores flag 4, and neither a load check nor a path
check can see the difference.
So the binary now says so itself: addon.cc exports
supportedProcessDataFlags, lib/index.js re-exports it, and
- windows-process-tree-creation-time.cjs asserts it during install,
which is what forces a from-source rebuild. It is shared by the Node
probe in ensure-native-runtime.mjs and the Electron probe in
rebuild-native-deps.mjs, exactly as node-pty-job-ownership.cjs is --
the Electron half matters because that probe decides onlyModules, so
without it the packaged app would ship the stale prebuilt.
- isWindowsProcessStartTimeAvailable() gates on the reported bit, not
the enum. Believing the enum is worse than reporting false: the
descendant snapshot returns null forever and the exit proof latches
unverifiable while structured chat believes it has a reaper.
rebuildNodeRuntimeModules could not actually have rebuilt this package:
the patched binding.gyp includes deps/node-addon-api, which the tarball
does not ship, and node-gyp must run from the physical dir.
Also closes the relay repair path's divergence: repairCreationTimeSources
wrote the C++ but not the buildNode splat or the tree-node typing, and
assertPatchApplied checked neither, so a repaired tree passed as patched
with buildProcessTree silently dropping the field.
The guard test is unchanged.
* fix(windows): keep the process-tree patch LF-only
windows-process-tree-patch-contract.test.mjs requires the patch file to
carry no CR bytes. Regenerating through pnpm patch-commit emitted 199 of
them, because the creation-time change is the first to touch files the
package ships as CRLF (src/process.h, src/process_worker.cc,
src/addon.cc, lib/index.js, lib/index.ts, the typings) -- and #17886's
own hunks over binding.gyp and src/process_commandline.cc carry the rest.
Stripping them is safe and changes nothing the lockfile records: pnpm
hashes patches CRLF-normalized, so the digest stays
e66202cc623996d02040c93449eb9ae353fddadf426cb53202a59ee710ee6fe7 and now
equals the file's plain sha256 too. It also still applies -- verified
against a deleted store entry, not a warm one -- and the precedent was
already there: the previous patch was LF-only and had been patching
those same CRLF files all along.
ensure-native-runtime.test.mjs stages the siblings the script loads at
module scope into its temp project. The import walk added by #17886 sees
`from './x.mjs'` only, so the createRequire'd .cjs siblings still have to
be named, and this PR adds a second one.
---------
Co-authored-by: Merge Sim <sim@local>
|
||
|
|
59fe8266bd |
fix(orchestration): keep worker lineage across app restart (STA-6366) (#19121)
* fix(orchestration): keep worker lineage across app restart (STA-6366) Terminal handles are minted per process, so after a restart the projected parent (coordinator or creator) named a handle no live row carried and every worker rendered as a top-level row. The projection now resolves the parent from the durable pane keys (runs.coordinator_pane_key, tasks.created_by_pane_key) whenever the stored handle is not one this process minted, re-resolves it to the live handle for that pane, and omits stale handles so they cannot mismatch a row. The creator-pane incarnation gate is untouched: it still decides mutation authority, and display lineage no longer depends on it. Dispatch lookup also passes the pane identity so a worker's own dispatch resolves once its handle is reminted. * test(orchestration): compare lineage without the merged attention field |
||
|
|
15d0f8aedf |
skills: rewrite the seven non-orchestration guides to one outcome-first standard (#18724)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 6 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$544 | $\color{#cf222e}{\Huge{\mathbf{−}}}$49 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$495 |
| Prod | 36 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$1719 | $\color{#cf222e}{\Huge{\mathbf{−}}}$1703 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$16 |
<!-- /orca-pr-loc -->
## ELI5
Orca ships eight skill guides that agents read before running the CLI. Seven of them (everything except `orchestration`, which #16904 rewrites) were command catalogs that had drifted from the binary. This PR rewrites them so an agent reads the outcome, the done bar, and the safe-failure rule first, loads reference material only at the step that needs it, and never sees a command or flag the installed CLI does not define.
## What changed
- **Seven guides rewritten** to one standard: outcome spine first (Result / Done / Safe failure), conditions instead of case lists, one done bar, one autonomy envelope, references loaded at the point of use via `skills get <topic> --full`, every runnable invocation spelled `ORCA`. `orca-cli` is 424→260 always-loaded lines with three references (browser, automations, publishing); `orca-per-workspace-env` is 794→397 with five (provider-vercel, ssh-host, docker-ssh, windows-scripts, failure-modes).
- **Defects fixed in shipped guides:** `emulator camera` (no such command), iOS `permissions` (backend refuses it), Android pane described as "in development" (shipped in June), `relayGracePeriodSeconds: 0` documented as immediate teardown (it is unbounded), doctor `ok: true` hiding `warn`, an SSH exemplar setting both `jumpHost` and `proxyCommand`, a provisioned-root fetch from `origin`, the Linear unconfirmed-write rule keyed on four verbs when ten emit it. Linear and emulator descriptions dropped embedded commands and angle-bracket placeholders (651→329, 732→404 chars).
- **Generator bundles references.** `skill-guides/<name>/references/*.md` is appended to `--full`; `skills get` help says compact by default, full with references.
- **Stubs single-authored.** The resolver ladder, placeholder rule, and older-binary fallback shared by all eight installable `SKILL.md` files come from one `skill-stubs/_shared/cli-resolution.md` fragment composed by the generator. Projections were byte-identical before the content fixes.
- **Guards:** every `ORCA <cmd>` and flag in every guide and reference resolves against `COMMAND_SPECS` (this found the camera defect); descriptions ≤1024 chars with no angle-bracket tokens; reference routing checked both directions; an always-loaded size ratchet (300 lines) that guides may leave but never join. `orchestration` (440 lines on main) is recorded as an exception until #16904 lands its kernel.
## Relationship to #16904
Split out of #16904 so that PR carries only the orchestration guide. On main, `terminal send` has no `--wait-submit` / `--retry-request` and the orchestration kernel still carries the resolver ladder and worktree-selector rule, so this branch pins `accepted: true` for handoff receipts and leaves the orchestration pins where main has them. The merge in either direction is mechanical: #16904 rebased on this becomes a one-file `orchestration.md` change plus dropping the two exceptions.
## Standard
Compound Engineering's portable skill-authoring guidance (outcome spine, conditions not cases, pinned fragile commands with an ordered hatch, references at point of use). NVIDIA SkillEvaluator Tier 1 (`schema,pii,license,quality,unicode,lint`) was run on every guide; its deterministic checks pass, its template nudges (Instructions/Examples sections, 50–150 char descriptions) do not apply to Orca's stub architecture and were not applied.
## Testing
- `pnpm typecheck:tsc:cli` clean; `check:code-quality:changed` and `check:react-doctor:changed` 0 findings
- `pnpm verify:bundled-skill-guides` and skill-bundle manifest verify clean
- vitest over `config/scripts`, `src/cli/skill-guide-cli-parity.test.ts`, `src/cli/skills.test.ts`, `src/cli/specs/skills.test.ts`, `src/cli/help.test.ts`, `src/main/skills`: 240 files / 2,019 pass
- Live smoke on the built CLI of every `skills get <topic>` and `--full`, every emulator, linear, and vm verb named in the guides, and every projection's resolver, GNOME warning, and bounded fallback (done on the #16904 branch before the split; the guide bodies are identical here except the send-receipt vocabulary noted above)
## Deferred product decisions
Merging `orca-emulator` and `orca-emulator-android` into one skill with a platform branch; collapsing `linear-tickets` to a guide alias; a `skills get --reference <name>` selector so a gate table can load one file; a fresh-agent routing eval before trimming the `orca-cli` (1,015 chars) and `orchestration` descriptions, whose quoted triggers each fixed a routing misroute.
|
||
|
|
298571ad9f |
fix(codex): uncap app-server stdio records (#18590)
Co-authored-by: Merge Sim <sim@local> |
||
|
|
0c33f58e8a |
fix(ssh-relay): daemon owns the endpoint credential; a losing start never rotates it (#19052)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 19 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$962 | $\color{#cf222e}{\Huge{\mathbf{−}}}$136 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$826 |
| Prod | 18 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$295 | $\color{#cf222e}{\Huge{\mathbf{−}}}$116 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$179 |
<!-- /orca-pr-loc -->
## Symptom
Live 2026-09-05 (Orca 1.4.198 client, Ubuntu host): both relay processes `kill -STOP`ped for 20 s, then `-CONT`. The client redeployed while the host was frozen. Its fresh daemon lost the socket bind (`Socket path already in use`) but had **already rewritten** `relay-<id>.sock.credential`. The surviving daemon kept its in-memory credential, so every later `--connect` got `Endpoint credential mismatch; closing socket`, then `Grace started … timeoutMs=0 … ptys=1, clients=0` every ~20 s, forever. Only a manual `kill -TERM` cleared it. Receipts: `review-archive/orchestration-v3-pr16904/smoke-receipts-t012b/E16,E17,E18,E24`.
Three independent defects kept the wedge alive; each is fixed at its own seam.
## Fix
**1. The relay daemon owns credential publication (race-free under two concurrent starters).**
`relay-daemon.ts` binds the socket first, then publishes via the new `src/relay/relay-endpoint-credential-publication.ts`: adopt a valid pre-existing file (older clients still pre-write), else mint 32 random bytes and write temp+rename at 0600. A start that loses the bind exits inside `listen()` and never reaches the file. Why this option and not restore-on-loss or a client-side write: the only process that can *prove* ownership is the one whose `listen()` succeeded, and that proof is atomic with the bind. The client-side pre-write (`ssh-relay-endpoint-credential.ts`) and the launch-command `chmod 600`/`icacls` are removed on POSIX and Windows. The racing test also exposed that macOS reports a mid-bind collision as `EEXIST` rather than `EADDRINUSE`; `relay-socket-ownership.ts` now treats both as "held or stale".
**2. The client distinguishes "no daemon" from "daemon present but not answering", and never rewrites.**
A credential refusal is now typed on the wire: the daemon replies `orca-relay-handshake-credential-mismatch` (same frame type, no new opcode) and the bridge exits **43**; `waitForSentinel` maps it to `RelayCredentialMismatchError`, which the takeover treats as handshake-refusal evidence exactly like exit 42. A relay that holds the endpoint but **never refused** (the stalled-host shape: kernel backlog accepts the probe, handshake gets no answer) is now `RelayEndpointUnresponsiveError`, routed to the relay-lost backoff instead of the terminal Reset Relay path. Silence is not a decision (`docs/reference/ssh-execution-boundary.md`).
**2b. Deploy honours the verdict.** The 40 s live run exposed that the `--connect` catch block in `deployAndLaunchRelay` predates the incumbent probe and swallowed both verdicts as "probe failed, launch fresh", so a fresh daemon was still launched over the live one (it lost the bind by luck, which is exactly the collision in the incident). Held and Unresponsive now propagate; the session backs off on Unresponsive and surfaces Reset Relay on Held. Red-first in `ssh-relay-deploy-incumbent-verdict.test.ts`.
**3. The daemon cannot be wedged by a rotated file, because nothing can rotate it.**
The credential lives in the content-hashed relay dir, and after (1) the only writer is the daemon that owns the socket, so the "file changed under a live daemon" state the incident depended on is no longer reachable in-product. The credential is therefore fixed for the daemon's lifetime, as a plain secret should be. A hand-edited file is refused with the typed reply until restored (tested). Startup adoption of a pre-written file applies an owner-only + same-uid rule (review finding): anything else is replaced by a fresh mint. An earlier revision of this PR also re-read the file on mismatch and adopted it; that was removed as unreachable machinery that turned the credential into a per-handshake file-ownership check.
**3b. Fail closed between bind and publication.** A client that arrives after `listen()` resolves but before the credential is set is refused, not admitted as `unproved`. Nothing can be delivered in that window today; the guard makes the boundary structural instead of an event-loop ordering fact. Red-first in `relay-reconnect-listener-credential-gate.test.ts`.
**Wire compat.** New optional handshake reply only; an old `--connect` hits `Unknown handshake type` and exits 1 pre-sentinel, which it already treated as a generic failure. New daemon adopts an old client's pre-written file; new client still passes `--credential-file` so an old daemon reads it as before. Absence of exit 43 is never used as evidence.
**Also.** `terminal create` on a reconnecting SSH host now says what to do instead of a bare `No PTY provider for connection "<id>"` (prefix preserved; the renderer matches it).
## Tests (red first)
- `src/relay/subprocess.test.ts`: two `--detached` starts race one socket + credential file → exactly one reaches the sentinel, loser exits 1 with `Socket path already in use`, file valid + 0600, a `--connect` reading it reaches `relay.status` and reports the winner's pid. Red before (both starters died: daemon required a pre-existing file), green 6/6 after.
- `src/relay/relay-endpoint-credential-publication.test.ts`: mints after bind; adopts a pre-written 0600 file; replaces a pre-written 0644 file with a fresh mint; refuses a stale credential with exit 43 while still serving the real one, and keeps refusing a rewritten file until it is restored.
- `src/relay/relay-reconnect-listener-credential-gate.test.ts`: a client in the bind-to-publish window is refused and never attached; after publication the right credential is accepted and a wrong one refused; a daemon launched without a credential file is not gated. Red without the guard.
- `ssh-relay-deploy-incumbent-verdict.test.ts`: live-but-silent incumbent → `RelayEndpointUnresponsiveError`, refused → `RelayEndpointHeldError`, and in neither case is `--detached` launched; a failed `test -S` probe still launches fresh. Red 2/3 without the deploy change.
- `ssh-relay-deploy-helpers.test.ts` (exit 43), `ssh-relay-endpoint-takeover.test.ts` (refused → Held even with no `lsof`; silent → Unresponsive, nothing unlinked or signalled), `ssh-relay-session-terminal-error.test.ts` (Unresponsive → `onRelayLost`, not terminal). Deploy/namespace/native-deps tests updated to assert the client writes **no** credential.
## Live proof
New `tests/e2e/ssh-docker-relay-stall-credential.spec.ts` (claimed in `run-ssh-docker-e2e.mjs` and PR source routing), two cases: `kill -STOP` every relay pid in the container, send input during the freeze, hold **20 s** (the incident's duration, which races the mux liveness timeout) or **40 s** (past it for sure), `kill -CONT`; assert status back to `connected`, same pty, same daemon pid, same credential inode and content, relay.log did not shrink (a relaunch truncates it) and has zero `Endpoint credential mismatch` / `Socket path already in use` lines, in-stall input delivered at most once.
Run output (local, fixture image `orca-e2e-ssh-relay:3a864c665ba2cefd`, `ORCA_E2E_SSH_DOCKER=1 SKIP_BUILD=1 ORCA_E2E_FORWARD_APP_LOGS=1 … --project electron-headless --workers=1`, head `c2c20fd994`; re-run identically on the final head after the credential-lifetime change, 2 passed (1.7m), same annotations, and the bind-to-publish refusal never fired):
```
✓ keeps the same daemon and credential across a 20s relay freeze (38.3s)
relay-processes-stopped: 2 relay-processes-continued: 2
bridge-pids-before-after: 480 -> 480
socket-clients-accepted-before-after: 1 -> 1
in-stall-input-delivered: 1
✓ backs off and reattaches, never relaunching, across a 40s relay freeze (57.5s)
relay-processes-stopped: 2 relay-processes-continued: 4
bridge-pids-before-after: 480 -> 1202
socket-clients-accepted-before-after: 1 -> 3
in-stall-input-delivered: 1
2 passed (1.6m)
```
Client log in the 40 s case shows the new path end to end: `Relay channel lost … reconnect attempt 1/6` → `Socket probe result: "ALIVE"` → `Socket reconnect failed … Relay failed to start within 10s` → `Relay endpoint incumbent: … verdict=live evidence=accepted-connection holders=unenumerable` → `Failed to re-establish relay … A relay still owns … but did not answer the handshake … Orca will retry` → `reconnect attempt 2/6` → `Reconnected to existing relay via socket`. The 20 s case never left the frozen bridge (same bridge pid, one accept), so it exercises the "silence is not death" side of the same race. The 20 s case passed 6/6 across the session; the 40 s case was red on the prior head (`Socket path already in use` + `Startup failed: listen EADDRINUSE` in relay.log from the swallowed verdict) and is green after 2b. Before the fix the same injection produced a fresh daemon that rewrote the credential and a survivor refusing every client.
The `relay-processes-continued` count exceeds `stopped` in the 40 s case because the timed-out client's `--connect` bridge and the loser-side processes are parked behind the frozen listener when `CONT` runs; they exit on their own once it resumes.
## Gates
`pnpm test src/relay src/main/ssh` 332 files / 3884 tests pass · `pnpm typecheck:tsc:node` clean · `check:code-quality:changed` 0 findings · `check:react-doctor:changed` 0 findings · `pr-e2e-gate-contract.test.mjs` 42 pass · no lint disables or max-lines bumps added.
## Noted, not fixed here
- `terminal list` `orphaned:false` / `terminal close` `ptyKilled:true` for a pane whose relay is gone (`orca-runtime-stop-explicitly-closed-tab-ptys.ts`): different seam, `@ts-nocheck` characterization-covered file.
- On a host with no `lsof`, a stalled relay still cannot be enumerated as the holder; it is now retried rather than declared held, but a relay frozen past the backoff budget still ends in the existing "reconnect manually" banner.
|
||
|
|
06a607a1d7 |
feat(orchestration): make multi-agent workflows durable (#16904)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 225 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$21666 | $\color{#cf222e}{\Huge{\mathbf{−}}}$2820 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$18846 |
| Prod | 348 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$17107 | $\color{#cf222e}{\Huge{\mathbf{−}}}$4706 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$12401 |
<!-- /orca-pr-loc -->
## ELI5
Orca now treats orchestration like a durable control plane instead of inferring success from terminal keystrokes. Agents can tell whether a prompt was accepted or a turn started, replay an ambiguous request without sending twice, and recover coordinator mail after a crash. Completed workers can be inspected, released, or retained, and their panes no longer auto-resume as if the work were still running.
## What changed
- **Run receipts** from `run-create/use/current/show/list` are the row without routing plumbing (`home_database`, `coordinator_pane_key`) and without the duplicate `binding` object.
- **`terminal send` receipts are honest and idempotent.** `input_accepted` and `turn_started` are the only stages; `--wait-submit` observes without resending; `--retry-request <uuid>` replays the exact request against the same process incarnation. A transport timeout keeps the retry ID; only a different runtime answering strips it. Value-less or non-UUID `--retry-request` is rejected on the CLI and the SSH shim.
- **Mailbox delivery is committed before wakeup.** Pointer writes are staged in the DB before any PTY byte, replayed once after restart, and never emit a naked Enter. The watermark that parks concurrent deliveries is released with the DB reservation. Restart rescans pointer-pending and `dispatch:` mailboxes.
- **Lifecycle is a guarded transition graph** (`lifecycle-transition.ts`) with a table-driven test over every caller edge. Task reopen/overturn stays in the public contract. A PTY exit during `worker-stop` is the stop succeeding, not a failure.
- **Worker lifecycle CLI:** `worker-start` (`--spec` creates Task + attempt in one call), `worker-show`, `worker-read` (provider transcript first, bounded terminal fallback with a typed reason, local/WSL/SSH), `worker-stop`, `worker-abandon`, `worker-release`, `worker-retain`, `worker-list` (rowid-fenced pagination, fleet liveness, `attention`, literal `nextAction`).
- **Release is an explicit ownership table** (`decideWorkerTerminalRelease`): only an `owned` resource can be settled, the archive is mandatory where reachable, and an owner whose process is proven exited can always get out of `retained` via `archive_status: unavailable`. User-taken-over, external, and transferred panes stay retained.
- **Settled-worker resume fence** (folds in #17651): a settled dispatch whose pane is still open is fenced at settlement, on stop/abandon/exit, and at startup; lifted on release, retain, takeover, and pane reuse.
- **Liveness is `live` / `unverifiable` / `exited` only**, from execution-host evidence. Fleet projection reads the evidence clock, not the relay delivery clock. A host-certified exit outranks the worker's settled state. `unverifiable` never authorizes stop, abandon, retry, or release, in code or in the guide.
- **Federation:** structured reads negotiate by `method_not_found` so every shipped host keeps transcript-first output; exited remote workers are closed before being reported closed; epoch fencing holds across peer restart, downgrade, and pairing rotation; no per-second forced capability probe.
- **Schema v35:** repairs databases stamped v34 by the pre-fix branch (mailbox_handle default, index predicates), drops the write-only `lifecycle_transition_receipts` ledger and five never-read v31 identity columns.
- **Schema v36:** `dispatch:<id>` mailboxes get a real consumer generation on `dispatch_contexts` and `remote_dispatch_attachments`, bumped and fenced in the same transaction on every re-attach (manual inject, worker-start, federated attach). A stale worker whose Dispatch moved to another process now gets `consumer_fenced` instead of silently acking the new worker's Delivery. Run mailboxes already worked this way.
- **Schema v37:** `dispatch_contexts` records its creator (`creator_handle`, `creator_pane_key`), so a coordinator's context-only self-dispatch is bookkeeping rather than a nesting parent; before this, one self-dispatch made every later `worker-start` from that coordinator fail the depth cap. Pre-v37 rows keep counting (fails closed).
- **Dispatch-mailbox ownership is checked, not inferred.** A `check` from a process whose pane no longer holds the Dispatch, or whose last Attempt was abandoned/failed and moved to another terminal, gets `consumer_fenced` instead of an empty inbox that reads as "no mail yet". `--peek`/`--all` stay readable. A paneless caller still gets `stable_pane_required` with the rebind recovery.
- **Liveness certification is stricter:** a `process_exited` stage whose termination reason is `unknown` (a stop that was issued but never observed) projects `unverifiable`, not `exited`. Federated `worker-show` carries the execution host's verdict and host kind instead of a local guess. A live, ready worker with nothing pending has `nextAction: none` rather than pointing at the `worker-show` that produced it.
- **Wire:** `workerShow` keeps `dispatch.task_id` next to `taskId` for shipped CLIs. `ask --json` uses the standard `{ok, result}` envelope like every sibling verb.
- **Migration start-version detection** treats the two v32 recovery columns as versioned. Before this, every shipped database stamped below 32 resolved to the v6 floor and replayed the whole chain (the v23 backfill synthesized 68 phantom retained workers on a real v30 profile). Verified on a copy of a real 62 MB v30 profile: starts at 30, no row delta, integrity ok, 11 ms.
- **Skill guide** rewritten as a ≤200-line kernel plus seven references, to the outcome-first standard (Result / Done / Safe failure first, conditions not case lists, one done bar, references loaded at the point of use). The canonical loop uses `worker-start --spec`, names `worker-list` for completion accounting, documents `--retry-request` / `request-show` / `--wait-submit`, and requires positive evidence before any stall action. The other seven guides get the same treatment in #18724, split out so this PR stays orchestration-only.
- **`rpc/methods/orchestration-*`** (126 flat files) regrouped into `orchestration/{worker,federation,messaging,runs,gates}/`.
## Why
User reports showed the same boundary failures: false `agent_prompt_stalled` causing duplicate sends (#15180), coordinators unable to trust screen scrapes, cold-parked terminals receiving a pointer without the submit, settled workers accumulating as live tabs and auto-resuming after restart, and no way to tell a stalled worker from a working one.
## Linked issues
Fixes #15180. Fixes #17935 (orchestration skill description is 866 characters; a guard now caps every bundled skill at 1,024). Supersedes #17651 (fence folded in). Advances #16660, #16522, #14907, #13047.
## Review record
This PR was reviewed adversarially after revival: eight independent lenses (lifecycle, mailbox, send, worker, federation, transcript, complexity, live ergonomics), each required to prove findings with a failing test. That produced 16 proven blockers, all fixed with red-then-green regression tests, followed by two re-review rounds and a third fix wave that caught 3 regressions introduced by the fixes and 7 fixes that missed their target; all closed. A final pass (five lenses incl. a live built-runtime smoke, then a re-review of the fix wave) found and fixed seven more, chiefly the stale-worker mailbox steal, the self-dispatch depth wedge, and the unproven-exit certification. Three independent Codex (gpt-6-astra) passes followed: the first found nothing new, the second found and fixed 3 defects (task-status reachability, WSL-local host classification, peer-capability epoch), the third found and fixed 6 (production PTY controller never installed settled writes, ambiguous in-flight pointer failures allowed duplicate replay, SSH/relay deadlines cut off a valid `--wait-submit`, stop-vs-exit race during inspection, and two release-recovery paths for vanished or exited terminals). The full record (findings, proof tests, triage, declines with reasons) is archived outside the repo.
**Rework after the live smoke.** A first live cross-host run on the shipped adhoc build (this Mac, a paired Windows host on the same build, a paired Mac on 1.4.195, and an SSH host) found a P1: a running local worker read `unverifiable`/`missing_status` because the fleet snapshot rows lacked the terminal handle the matcher keyed on. A 59-row failure table over every bug fixed during review showed the same two classes recurring: a fact dropped in transit through optional fields, and two authorities for one fact. Two blind designs (Opus, Codex) converged on the same mechanisms, and the scoped tranches landed here with red-then-green seam tests from the real producer to the real consumer, faults injected only at the transport or hook-ingest boundary:
- **Settlement (data-loss class):** one three-valued `WriteSettlement` (`accepted | refused{reason} | unverifiable{reason, bytesHandedToTransport}`) from the SSH multiplexer through daemon client, providers, controller, to pointer staging. No boolean, no rejection-as-third-state. The two silent degrades that fabricated a handoff are deleted; a provider that cannot settle refuses before any effect. Pointer text and Enter share the contract; a partial flush is `unverifiable`, never `refused`.
- **Evidence identity (false-liveness class):** fleet agent-status evidence is a tagged union (`binding: worker | pane | unresolved{reason}`, `clock: observed | delivery`) minted once at ingest, so a hook row captured on one process incarnation can never bind to a later dispatch on the same pane. The matcher's `!worker.paneKey ||` defaults are gone. One host-scope parser replaces two.
- **Small pre-merge items:** `capability_unsupported` from an old peer is no longer relabelled `host_unavailable`; a producer census test asserts every agent-status consumer path projects a pane-only hook row as `live`.
Two ergonomics defects the second live run surfaced on a real database are fixed here too: a pre-v3 dispatch already marked `completed` projected as `outcome_unknown` / `requiresAction: true` forever (three copies of the outcome ladder disagreed on legacy rows; now one resolver, legacy `completed` reads `succeeded` with nothing to act on, legacy `failed` stays actionable on the failure), and an unscoped `worker-list` enumerated the entire database (now defaults to the Run bound to the calling terminal, `--run` overrides, and the receipt's additive `scope` field says which).
A third live round on the shipped adhoc build of `b082443e1f` (same four hosts) plus an unscripted run in the user's own prompt style (a plain Claude Code shell, `/orchestration`, three workers, zero errors, bound-Run default confirmed) found two more branch defects, fixed with red-then-green tests: a worker freshly started on a paired server projected `unverifiable`/`host_indeterminate` with `requiresAction` for ~3 minutes, including after its own `worker_done`, because the host's federation observation returned `missing_liveness_verdict` for any PTY the liveness register had not yet swept (the host now reads a connected pane it owns locally as `live`; disconnected or SSH-scoped panes stay `unverifiable`); and six pre-v3 completed rows still carried an `input` category because settling through the task-status path or `failDispatch` never closed the Dispatch's pending question threads (both paths close them now, and schema v38 closes threads already pending on settled rows). The guide's `worker-start` examples now show `--model sonnet`, since an omitted model inherits the launcher's default.
A Codex adversarial pass on the tranche diff found one real design hole (identity minted at read time instead of ingest, now closed) and two daemon settlement paths that threw instead of settling (fixed). Two `@ts-nocheck` runtime mixins on these paths were extracted into checked modules; the repo-wide `@ts-nocheck` count is unchanged at 171.
Deletions during review: ~1,900 lines (write-only ledger, unread columns, dead v1 archive path, test harnesses shipped in prod, duplicated liveness and state-machine copies, self-capability checks that were compile-time true).
## Testing
- `pnpm typecheck:tsc:node|cli|web` clean
- `pnpm run check:code-quality:changed` 0 findings; `check:react-doctor:changed` 0
- `pnpm verify:bundled-skill-guides`, `verify:skill-bundle-manifest`
- full `pnpm test` on the integrated head: 72,332 pass / 292 skipped; the only failures were three non-PR files (two zsh live-shell suites hit a node-pty spawn-helper ENOENT while a concurrent native rebuild ran, 44/44 in isolation; `release-checkout.unit.test.ts` is a known 30 s load timeout that passes in isolation on `origin/main` too).
- CI on
|
||
|
|
3be526c5e6 |
test: cover SSH reattach replay and enable deterministic Codex CI (#19106)
* test: cover SSH replay replies and run deterministic Codex restore scenarios * test: register replay probe unit command in reliability gate |
||
|
|
4d9e963ffd |
test: enable localhost SSH terminal and hook journey in CI (#19097)
* test: run localhost SSH terminal and hooks in CI * test: isolate localhost SSH session fixtures across repetitions * test: route remote agent hook source changes to localhost journey * test: record localhost SSH reliability evidence and remaining gaps * test: route the real SSH session hook authority |
||
|
|
f5960cec00 |
test: enable Docker SSH browser network route coverage in CI (#19095)
* test: enable Docker SSH browser network route journeys in CI * test: register Docker browser job in token permissions contract * test: declare SSH client dependency and narrow browser fixture routing |
||
|
|
3631f886a7 |
test: enable direct and client-hosted SSH browser coverage (#19090)
* test: enable direct and client-hosted SSH browser coverage * test: record twelve passing SSH browser journey repetitions * test: distinguish SSH journey evidence from unit runtime budget |