mirror of
https://github.com/stablyai/orca.git
synced 2026-09-22 16:02:32 +00:00
44a74baf73b2bcbfa0a9727fa4d0505bc17f2386
13
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
08b96ed1b3 |
Seed Cmd-J filter from sidebar scope (#19036)
* feat(palette): seed Cmd+J filter from sidebar show scope When opening Cmd+J, the palette's host and project filters now initialize from the sidebar's current Show scope, so results match the user's sidebar view. The palette can still be cleared or changed per open; sidebar never reads back palette filters. * refactor: pass app state to palette filter builder Let the builder function extract the sidebar scope it needs instead of requiring callers to destructure and pass individual properties. This reduces coupling and simplifies the data flow through the palette initialization lifecycle. * Make palette filter repo-granular to preserve sidebar scope Filter options now list individual repositories instead of grouping multi-repo projects into single rows. This preserves the exact repository scope shown in the sidebar when opening Cmd+J, rather than widening selections to entire projects. Removes per-field selection cap and stale-value reconciliation, simplifying the filter lifecycle. * Clarify filter naming and seed from sidebar scope on palette open - Rename projects→repositories in PaletteFilterModel for semantic accuracy - Rename rawFilter→filterState for clearer intent - Initialize filter from sidebar scope in local state, refresh on open - Remove redundant filter reset from selection lifecycle * Seed Cmd-J filter from sidebar scope and reset on close The palette now opens with the sidebar's host and repository scope applied. Filter changes are temporary: closing discards them, and reopening reseeds from the sidebar's current state. - Repository filtering is now granular (individual repos) - Support shared repository IDs across multiple hosts - Disambiguate duplicate repository names by path * Add comment clarifying Projects terminology Document the naming convention for repository-granular filter choices to help future maintainers understand why "Projects" is used as the user-facing term. * Remove redundant Escape press from worktree palette filter test |
||
|
|
06a607a1d7 |
feat(orchestration): make multi-agent workflows durable (#16904)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 225 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$21666 | $\color{#cf222e}{\Huge{\mathbf{−}}}$2820 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$18846 |
| Prod | 348 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$17107 | $\color{#cf222e}{\Huge{\mathbf{−}}}$4706 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$12401 |
<!-- /orca-pr-loc -->
## ELI5
Orca now treats orchestration like a durable control plane instead of inferring success from terminal keystrokes. Agents can tell whether a prompt was accepted or a turn started, replay an ambiguous request without sending twice, and recover coordinator mail after a crash. Completed workers can be inspected, released, or retained, and their panes no longer auto-resume as if the work were still running.
## What changed
- **Run receipts** from `run-create/use/current/show/list` are the row without routing plumbing (`home_database`, `coordinator_pane_key`) and without the duplicate `binding` object.
- **`terminal send` receipts are honest and idempotent.** `input_accepted` and `turn_started` are the only stages; `--wait-submit` observes without resending; `--retry-request <uuid>` replays the exact request against the same process incarnation. A transport timeout keeps the retry ID; only a different runtime answering strips it. Value-less or non-UUID `--retry-request` is rejected on the CLI and the SSH shim.
- **Mailbox delivery is committed before wakeup.** Pointer writes are staged in the DB before any PTY byte, replayed once after restart, and never emit a naked Enter. The watermark that parks concurrent deliveries is released with the DB reservation. Restart rescans pointer-pending and `dispatch:` mailboxes.
- **Lifecycle is a guarded transition graph** (`lifecycle-transition.ts`) with a table-driven test over every caller edge. Task reopen/overturn stays in the public contract. A PTY exit during `worker-stop` is the stop succeeding, not a failure.
- **Worker lifecycle CLI:** `worker-start` (`--spec` creates Task + attempt in one call), `worker-show`, `worker-read` (provider transcript first, bounded terminal fallback with a typed reason, local/WSL/SSH), `worker-stop`, `worker-abandon`, `worker-release`, `worker-retain`, `worker-list` (rowid-fenced pagination, fleet liveness, `attention`, literal `nextAction`).
- **Release is an explicit ownership table** (`decideWorkerTerminalRelease`): only an `owned` resource can be settled, the archive is mandatory where reachable, and an owner whose process is proven exited can always get out of `retained` via `archive_status: unavailable`. User-taken-over, external, and transferred panes stay retained.
- **Settled-worker resume fence** (folds in #17651): a settled dispatch whose pane is still open is fenced at settlement, on stop/abandon/exit, and at startup; lifted on release, retain, takeover, and pane reuse.
- **Liveness is `live` / `unverifiable` / `exited` only**, from execution-host evidence. Fleet projection reads the evidence clock, not the relay delivery clock. A host-certified exit outranks the worker's settled state. `unverifiable` never authorizes stop, abandon, retry, or release, in code or in the guide.
- **Federation:** structured reads negotiate by `method_not_found` so every shipped host keeps transcript-first output; exited remote workers are closed before being reported closed; epoch fencing holds across peer restart, downgrade, and pairing rotation; no per-second forced capability probe.
- **Schema v35:** repairs databases stamped v34 by the pre-fix branch (mailbox_handle default, index predicates), drops the write-only `lifecycle_transition_receipts` ledger and five never-read v31 identity columns.
- **Schema v36:** `dispatch:<id>` mailboxes get a real consumer generation on `dispatch_contexts` and `remote_dispatch_attachments`, bumped and fenced in the same transaction on every re-attach (manual inject, worker-start, federated attach). A stale worker whose Dispatch moved to another process now gets `consumer_fenced` instead of silently acking the new worker's Delivery. Run mailboxes already worked this way.
- **Schema v37:** `dispatch_contexts` records its creator (`creator_handle`, `creator_pane_key`), so a coordinator's context-only self-dispatch is bookkeeping rather than a nesting parent; before this, one self-dispatch made every later `worker-start` from that coordinator fail the depth cap. Pre-v37 rows keep counting (fails closed).
- **Dispatch-mailbox ownership is checked, not inferred.** A `check` from a process whose pane no longer holds the Dispatch, or whose last Attempt was abandoned/failed and moved to another terminal, gets `consumer_fenced` instead of an empty inbox that reads as "no mail yet". `--peek`/`--all` stay readable. A paneless caller still gets `stable_pane_required` with the rebind recovery.
- **Liveness certification is stricter:** a `process_exited` stage whose termination reason is `unknown` (a stop that was issued but never observed) projects `unverifiable`, not `exited`. Federated `worker-show` carries the execution host's verdict and host kind instead of a local guess. A live, ready worker with nothing pending has `nextAction: none` rather than pointing at the `worker-show` that produced it.
- **Wire:** `workerShow` keeps `dispatch.task_id` next to `taskId` for shipped CLIs. `ask --json` uses the standard `{ok, result}` envelope like every sibling verb.
- **Migration start-version detection** treats the two v32 recovery columns as versioned. Before this, every shipped database stamped below 32 resolved to the v6 floor and replayed the whole chain (the v23 backfill synthesized 68 phantom retained workers on a real v30 profile). Verified on a copy of a real 62 MB v30 profile: starts at 30, no row delta, integrity ok, 11 ms.
- **Skill guide** rewritten as a ≤200-line kernel plus seven references, to the outcome-first standard (Result / Done / Safe failure first, conditions not case lists, one done bar, references loaded at the point of use). The canonical loop uses `worker-start --spec`, names `worker-list` for completion accounting, documents `--retry-request` / `request-show` / `--wait-submit`, and requires positive evidence before any stall action. The other seven guides get the same treatment in #18724, split out so this PR stays orchestration-only.
- **`rpc/methods/orchestration-*`** (126 flat files) regrouped into `orchestration/{worker,federation,messaging,runs,gates}/`.
## Why
User reports showed the same boundary failures: false `agent_prompt_stalled` causing duplicate sends (#15180), coordinators unable to trust screen scrapes, cold-parked terminals receiving a pointer without the submit, settled workers accumulating as live tabs and auto-resuming after restart, and no way to tell a stalled worker from a working one.
## Linked issues
Fixes #15180. Fixes #17935 (orchestration skill description is 866 characters; a guard now caps every bundled skill at 1,024). Supersedes #17651 (fence folded in). Advances #16660, #16522, #14907, #13047.
## Review record
This PR was reviewed adversarially after revival: eight independent lenses (lifecycle, mailbox, send, worker, federation, transcript, complexity, live ergonomics), each required to prove findings with a failing test. That produced 16 proven blockers, all fixed with red-then-green regression tests, followed by two re-review rounds and a third fix wave that caught 3 regressions introduced by the fixes and 7 fixes that missed their target; all closed. A final pass (five lenses incl. a live built-runtime smoke, then a re-review of the fix wave) found and fixed seven more, chiefly the stale-worker mailbox steal, the self-dispatch depth wedge, and the unproven-exit certification. Three independent Codex (gpt-6-astra) passes followed: the first found nothing new, the second found and fixed 3 defects (task-status reachability, WSL-local host classification, peer-capability epoch), the third found and fixed 6 (production PTY controller never installed settled writes, ambiguous in-flight pointer failures allowed duplicate replay, SSH/relay deadlines cut off a valid `--wait-submit`, stop-vs-exit race during inspection, and two release-recovery paths for vanished or exited terminals). The full record (findings, proof tests, triage, declines with reasons) is archived outside the repo.
**Rework after the live smoke.** A first live cross-host run on the shipped adhoc build (this Mac, a paired Windows host on the same build, a paired Mac on 1.4.195, and an SSH host) found a P1: a running local worker read `unverifiable`/`missing_status` because the fleet snapshot rows lacked the terminal handle the matcher keyed on. A 59-row failure table over every bug fixed during review showed the same two classes recurring: a fact dropped in transit through optional fields, and two authorities for one fact. Two blind designs (Opus, Codex) converged on the same mechanisms, and the scoped tranches landed here with red-then-green seam tests from the real producer to the real consumer, faults injected only at the transport or hook-ingest boundary:
- **Settlement (data-loss class):** one three-valued `WriteSettlement` (`accepted | refused{reason} | unverifiable{reason, bytesHandedToTransport}`) from the SSH multiplexer through daemon client, providers, controller, to pointer staging. No boolean, no rejection-as-third-state. The two silent degrades that fabricated a handoff are deleted; a provider that cannot settle refuses before any effect. Pointer text and Enter share the contract; a partial flush is `unverifiable`, never `refused`.
- **Evidence identity (false-liveness class):** fleet agent-status evidence is a tagged union (`binding: worker | pane | unresolved{reason}`, `clock: observed | delivery`) minted once at ingest, so a hook row captured on one process incarnation can never bind to a later dispatch on the same pane. The matcher's `!worker.paneKey ||` defaults are gone. One host-scope parser replaces two.
- **Small pre-merge items:** `capability_unsupported` from an old peer is no longer relabelled `host_unavailable`; a producer census test asserts every agent-status consumer path projects a pane-only hook row as `live`.
Two ergonomics defects the second live run surfaced on a real database are fixed here too: a pre-v3 dispatch already marked `completed` projected as `outcome_unknown` / `requiresAction: true` forever (three copies of the outcome ladder disagreed on legacy rows; now one resolver, legacy `completed` reads `succeeded` with nothing to act on, legacy `failed` stays actionable on the failure), and an unscoped `worker-list` enumerated the entire database (now defaults to the Run bound to the calling terminal, `--run` overrides, and the receipt's additive `scope` field says which).
A third live round on the shipped adhoc build of `b082443e1f` (same four hosts) plus an unscripted run in the user's own prompt style (a plain Claude Code shell, `/orchestration`, three workers, zero errors, bound-Run default confirmed) found two more branch defects, fixed with red-then-green tests: a worker freshly started on a paired server projected `unverifiable`/`host_indeterminate` with `requiresAction` for ~3 minutes, including after its own `worker_done`, because the host's federation observation returned `missing_liveness_verdict` for any PTY the liveness register had not yet swept (the host now reads a connected pane it owns locally as `live`; disconnected or SSH-scoped panes stay `unverifiable`); and six pre-v3 completed rows still carried an `input` category because settling through the task-status path or `failDispatch` never closed the Dispatch's pending question threads (both paths close them now, and schema v38 closes threads already pending on settled rows). The guide's `worker-start` examples now show `--model sonnet`, since an omitted model inherits the launcher's default.
A Codex adversarial pass on the tranche diff found one real design hole (identity minted at read time instead of ingest, now closed) and two daemon settlement paths that threw instead of settling (fixed). Two `@ts-nocheck` runtime mixins on these paths were extracted into checked modules; the repo-wide `@ts-nocheck` count is unchanged at 171.
Deletions during review: ~1,900 lines (write-only ledger, unread columns, dead v1 archive path, test harnesses shipped in prod, duplicated liveness and state-machine copies, self-capability checks that were compile-time true).
## Testing
- `pnpm typecheck:tsc:node|cli|web` clean
- `pnpm run check:code-quality:changed` 0 findings; `check:react-doctor:changed` 0
- `pnpm verify:bundled-skill-guides`, `verify:skill-bundle-manifest`
- full `pnpm test` on the integrated head: 72,332 pass / 292 skipped; the only failures were three non-PR files (two zsh live-shell suites hit a node-pty spawn-helper ENOENT while a concurrent native rebuild ran, 44/44 in isolation; `release-checkout.unit.test.ts` is a known 30 s load timeout that passes in isolation on `origin/main` too).
- CI on
|
||
|
|
51eed5a1bc |
feat(cli): report SSH host platforms (#18896)
* feat(cli): report SSH host platforms * feat(cli): include SSH connection status * fix(cli): preserve unknown SSH connection state |
||
|
|
436ef827dd |
fix(browser): present Electron's own user agent so Cloudflare Turnstile clears (#18749)
Orca rewrote every browser session's UA to look like plain Chrome by stripping the Electron and app tokens. That rewrite is what Cloudflare rejects: a Chrome UA that ships no client hints reads as a spoof and Turnstile returns 600010, while the same binary on the same IP clears every challenge with its stock UA. PR #885 added the rewrite to fix 600010 and was treating a symptom it created; issue #11518 later found the same rewrite is what broke Google sign-in. - Keep the stock Electron UA on every partition. The webRequest handler now only owns the host-scoped Google auth Firefox switch, which stays unchanged. - Delete the anti-detection script. Measured on Electron 43: plugins are already a real PluginArray, window.chrome exists, and navigator.webdriver is false even with the debugger attached, so three of its four premises were wrong, and the overrides it installed (instance-level webdriver, non-native Permissions.query, stubbed chrome.csi/loadTimes) are themselves published bot signatures. - Stop attaching a CDP debugger to every browsing guest. Only the auth-UA detach listener remains, because a detach clears Chromium's standing UA override. - Stop sending Runtime.enable into cross-origin iframes when the agent bridge auto-attaches. The challenge widget is one, nothing reads iframe Runtime events, and the Runtime domain's serialization side effect is the documented Cloudflare CDP tell. - Add a real-Electron test proving the wire identity: stock UA to ordinary hosts, Firefox with no client hints to accounts.google.com. Verified in the dev build: dash.cloudflare.com/login no longer shows "There was a problem with verification" and scrapingcourse.com's managed challenge clears, both failing deterministically before. Fixes #13822 |
||
|
|
f4b207bc38 | docs(ssh): document keep-alive-until-reset as the default grace (#18383) | ||
|
|
573537ecd4 |
feat(cli): make terminal close the canonical workspace teardown (#18073)
* fix(runtime): recover stale session owners and await retirement * fix(runtime): preserve session hydration and smoke compatibility * test(runtime): cover empty and unindexed session owners * feat(cli): make terminal close the canonical workspace teardown * fix(preload): align ssh termination result type * test(runtime): assert folder hydration owner * fix(runtime): fence legacy terminal stop by worktree host * fix(preload): reconcile ssh result import with main * fix(runtime): keep same-id sibling hosts out of workspace close The stale-owner fallback in the session controller re-routed any worktree whose catalog partition had no tabs to whichever other partition held tabs. Only `runtime:` environment ids rotate across relay restarts; `repoId::path` legitimately repeats across hosts, so an SSH workspace close could retire the local copy's tabs and resume records, or flip owners mid-close and strand the SSH PTY. Restrict the fallback to runtime hosts, and pin the session partition once per workspace close so record clearing targets the partition that owned the tabs. * test(runtime): give the cross-host close fixture a real resume record * fix(preload): take main's ssh-bridge import order so the merge stays duplicate-free |
||
|
|
e3de6b2ce8 |
Add automation runs dashboard with pagination and filtering (#18226)
* Add automation runs dashboard with pagination and filtering Adds a new Runs view in the Automations page that lets users browse all runs across automations with status/host filtering, search, and pagination support. Includes virtualized table rendering for efficient handling of large run histories and summary cards showing 24h/7d success/failure counts. * Fix missing dependencies in useCallback hooks and imports Missing dependencies in useCallback can cause stale closure bugs. This adds missing state setters to dependency arrays and consolidates type imports for consistency. * Use keyset pagination for stable automation runs pages Pagination now uses createdAt:id boundaries instead of offsets, so new runs arriving between pages don't shift the window. Maintains backwards compatibility with legacy offset cursors. Move pagination to shared module, fix outcome counting for future-dated runs, and improve hook state tracking on authority re-pairing or target changes. * Extract automation run details to top-level page view Moves run display from detail pane to dedicated page, establishing three-level navigation (Automations → Runs → Run Details) and simplifying the detail pane component. * Fix pagination stability when automation runs share createdAt - Define a stable total order with createdAt and id tiebreaker to prevent runs tied on createdAt from being dropped when the boundary run is pruned between page requests - Retain cursor on failed pagination so pages remain retryable - Update ownerNotice type to AutomationActionNotice * Extract automations list panel and worktree map logic Split AutomationsPageSurface into smaller, focused modules for better maintainability and reusability. Move list panel UI rendering to AutomationsPageListPanel component and worktree map selection logic to a standalone utility function. * Add i18n strings for automation runs dashboard Adds localized strings for the automation runs dashboard view, including search, filtering by host and status, run counts for 24h/7d windows, and empty state messaging across all supported languages. * fix missing translation * fix missing translation |
||
|
|
e5a1e79e8e |
docs(linux): say which package to install and how updates arrive (#18123)
* docs(linux): say which package to install and how updates arrive Closes #5188. Closes #10987. The install guide's entire Linux section was "AppImage and `.deb` builds are available. See the Releases page for details." It named two of the three published packages, gave no basis for choosing between them, and said nothing about updating -- which is the one thing that actually differs between them. Separately, nothing human-facing said the Linux CLI is `orca-ide`; only skills/orca-cli/SKILL.md carried it, which agents read and humans do not. Install page now picks the package by update behaviour: the AppImage self-updates, deb/rpm report the new version and hand over the install command, and a repackaged build is not offered a download it cannot apply. Records that Orca never escalates privileges for the package install, and points at #18086 for the signed repo as planned, not shipped. Adds .rpm to the download list. Release CI builds it (release-cut.yml: `--linux AppImage deb rpm`) and verify-release-required-assets.mjs requires the artifact, so omitting it was just wrong. The CLI command name is now stated where humans hit it -- the CLI reference and overview -- with the GNOME Orca collision as the reason, plus the two places bare `orca` does work: inside Orca-managed terminals (PTY PATH shim) and on a packaged `orca serve` host (the ~/.local/bin dispatcher). The headless guide gains the same note, which is what makes its `orca skills install` lines correct rather than a typo. * docs(linux): fix install ordering, CLI verification, and serve bootstrap Readiness review found ten defects. Two would have had a reader run the wrong program, and one would have had them install a .deb over a live app. Install ordering was reversed. The page said "run it, then quit and reopen Orca"; the ref this is gated to land with says the opposite in four places (linux-package-downloaded-status.ts LINUX_PACKAGE_MANUAL_INSTALL_MESSAGE, "Quit Orca before running the system package install command", plus the recovery card's title, summary and explainer). That wording came from main's older run-then-quit card, which the stack deliberately reversed when it retitled the card to "Manual Install Required". Now: quit first. CLI verification put the Linux caveat *below* `command -v orca`. That check succeeds on any GNOME desktop and resolves to the screen reader, so the reader got a confident hit from the page's own verification step and then invoked the wrong program. Caveat moved above, and the block now spells `orca-ide` literally instead of asking the reader to substitute. The serve bootstrap was circular: the bare-`orca` dispatcher is written *during* serve startup (main-process-runtime-launch.ts), so it can never be the command that starts serve. First launch is `orca-ide serve`. Fixed here and in the two pages this links to. Accuracy: the install command now matches what the code emits -- absolute paths resolved from the trusted directories and a POSIX-single-quoted package path, as pinned by linux-package-install-command.test.ts -- and names the manager fallbacks (dpkg; zypper/dnf/yum/rpm) rather than presenting apt as the only form. The pending path honours XDG_CACHE_HOME. rpm arch tokens are x86_64 and aarch64, not deb's amd64/arm64. arm64 AppImage is linked. Dropped the container example: isExternallyManagedLinuxInstall() needs a root marker AND no trusted package manager, and a Debian-based container has apt, so it is not flagged. |
||
|
|
42d9ac1767 |
chore(deps): resolve 81 of 83 Dependabot alerts in docs/site and mobile (#18061)
docs/site: bump next 16.2.1 -> 16.3.4 (with eslint-config-next) and vercel 50.37.0 -> 59.11.1, then refresh transitives. The 16.3.x jump is required: 16.2.x hard-pins the vulnerable postcss@8.4.31 and sharp@^0.34.5, while 16.3.x pins postcss@8.5.23 and sharp@^0.35.4. Five packages are exact-pinned by vercel's own subpackages, so they get scoped overrides. Scoped rather than blanket because a bare undici override would drag the 6.x/7.x consumers in the tree down to 5.x. mobile: bump browserslist 4.28.2 -> 4.28.8. Two alerts stay open, both in mobile: - decode-uri-component@0.2.2 (#285). An override to 0.5.0 breaks the tree: 0.5.0 is ESM-only with a default export, but query-string@7.1.3 is CJS and does `require('decode-uri-component')`, so parse() throws "decodeComponent is not a function" and takes URL parsing in expo-router and @react-navigation/core with it. Both pin query-string@^7.1.3; the fix has to come from upstream moving to query-string 8+. - image-size@1.2.1 (#179, #180) via metro. No patched version exists on any release line, so there is nothing to override to. Verified: docs/site build, tests, lint, tsc and frozen install; mobile typecheck, 3985 tests and frozen install. |
||
|
|
bf62abc3c9 | fix(docs): align OSS docs presentation and SEO (#17809) | ||
|
|
45c4823109 |
Format documentation with consistent line wrapping and table alignment (#17765)
Standardize MDX files across docs with: - Remove trailing semicolons from import statements - Wrap long lines and multi-line component props for readability - Align Markdown table column separators - Normalize text and JSX formatting for consistency |
||
|
|
649c188cc4 |
docs: clarify release cutover ordering
Clarifies that marketing rewrites are prepared but remain disabled until a stable release-backed docs deployment is verified. |
||
|
|
6aba202d5e |
feat(docs): publish OSS docs with stable releases
Publish the standalone docs site under docs/site and deploy it on stable desktop releases. |