From 16cd1e0c75773b9ae5dc76df0d4a4fe194a43cd4 Mon Sep 17 00:00:00 2001 From: Neil <4138956+nwparker@users.noreply.github.com> Date: Wed, 7 Oct 2026 15:43:39 -0700 Subject: [PATCH] docs: remove stale internal reference documentation (#26328) * docs: remove stale internal reference documentation * test: avoid pooled relay hook sockets across fake clock advances * test: remove checks for deleted headless server documentation --- .github/CONTRIBUTING.md | 1 - .github/pull_request_template.md | 3 +- .gitignore | 39 +- AGENTS.md | 30 +- README.md | 1 - config/reliability-gates.jsonc | 3 +- docs/readme/README.fr.md | 1 - docs/readme/README.ko.md | 1 - docs/readme/README.pt.md | 1 - docs/reference/admin-agent-skill-sharing.md | 97 - .../reference/agent-pty-transcript-capture.md | 131 - .../agent-session-search-contract.md | 172 -- .../agent-session-search-query-tuning.md | 229 -- docs/reference/agent-skill-provider-paths.md | 38 - .../agent-skill-sharing-threat-model.md | 101 - .../agent-skill-sharing-upstream-boundary.md | 43 - docs/reference/agent-status-store.md | 577 ---- docs/reference/antigravity-native-accounts.md | 90 - .../antigravity-readiness-evidence.md | 332 --- .../antivirus-prerelease-clearance.md | 117 - docs/reference/ci-demand-rollout.md | 220 -- docs/reference/ci-runner-efficiency.md | 2629 ----------------- ...line-and-prime-agent-readiness-evidence.md | 85 - docs/reference/codebuddy-harness.md | 44 - docs/reference/deepseek-build-observation.md | 28 - docs/reference/dsh-harness-integration.md | 44 - docs/reference/git-compatibility.md | 94 - docs/reference/headless-linux-server.md | 1019 ------- docs/reference/ime-regression-checklist.md | 146 - docs/reference/linux-glibc-compatibility.md | 182 -- docs/reference/macos-press-and-hold.md | 58 - ...malformed-worktree-registration-removal.md | 43 - docs/reference/managed-data-accounts.md | 23 - .../reference/monaco-language-associations.md | 34 - docs/reference/omp-fresh-launch.md | 51 - docs/reference/omp-history-titles.md | 34 - .../omp-resume-transcript-locator.md | 28 - .../omp-runtime-session-provenance.md | 23 - docs/reference/omp-session-roots.md | 53 - docs/reference/omp-startup-keyboard-query.md | 15 - docs/reference/omp-status-input-redaction.md | 39 - docs/reference/orcad-operations.md | 399 --- .../orchestration-configured-agent-aliases.md | 22 - docs/reference/pnpm-install-policy.md | 92 - docs/reference/qoder-integration.md | 90 - docs/reference/relay-regional-placement.md | 39 - docs/reference/remote-wire-compatibility.md | 397 --- .../renderer-agent-status-performance.md | 340 --- docs/reference/runtime-file-base64-padding.md | 92 - docs/reference/sharing-agent-skills.md | 88 - .../spinner-rendering-performance.md | 203 -- docs/reference/ssh-execution-boundary.md | 129 - docs/reference/ssh-host-key-verification.md | 424 --- .../ssh-reconnect-source-recovery.md | 159 - .../task-provider-identity-validation.md | 81 - .../terminal-artifact-grant-integrity.md | 85 - .../terminal-perf-latency-investigation.md | 106 - .../reference/terminal-perf-report-budgets.md | 59 - docs/reference/terminal-startup-timing.md | 30 - docs/reference/windows-cmd-shim-resolution.md | 77 - .../windows-daemon-host-relocation.md | 132 - docs/reference/windows-edr-posture.md | 562 ---- docs/reference/windows-msys-job-breakaway.md | 165 -- docs/reference/windows-process-enumeration.md | 672 ----- docs/reference/windows-setup-shell.md | 86 - docs/reference/windows-signing-runner-time.md | 137 - .../windows-terminal-shell-selection.md | 54 - docs/reference/worktree-scan-fingerprint.md | 274 -- docs/reference/wsl-command-execution.md | 173 -- docs/reference/wsl-managed-cli.md | 37 - docs/reference/wsl-probe-failure-semantics.md | 66 - docs/reference/wsl-runner-verification.md | 27 - docs/reference/xterm-patch-regeneration.md | 320 -- ...erver-claude-task-wakeup-lifecycle.test.ts | 2 + .../antigravity-readiness-transcripts.test.ts | 26 +- ...single-instance-lock-headless-exit.test.ts | 50 - tests/AGENTS.md | 2 +- 77 files changed, 25 insertions(+), 12571 deletions(-) delete mode 100644 docs/reference/admin-agent-skill-sharing.md delete mode 100644 docs/reference/agent-pty-transcript-capture.md delete mode 100644 docs/reference/agent-session-search-contract.md delete mode 100644 docs/reference/agent-session-search-query-tuning.md delete mode 100644 docs/reference/agent-skill-provider-paths.md delete mode 100644 docs/reference/agent-skill-sharing-threat-model.md delete mode 100644 docs/reference/agent-skill-sharing-upstream-boundary.md delete mode 100644 docs/reference/agent-status-store.md delete mode 100644 docs/reference/antigravity-native-accounts.md delete mode 100644 docs/reference/antigravity-readiness-evidence.md delete mode 100644 docs/reference/antivirus-prerelease-clearance.md delete mode 100644 docs/reference/ci-demand-rollout.md delete mode 100644 docs/reference/ci-runner-efficiency.md delete mode 100644 docs/reference/cline-and-prime-agent-readiness-evidence.md delete mode 100644 docs/reference/codebuddy-harness.md delete mode 100644 docs/reference/deepseek-build-observation.md delete mode 100644 docs/reference/dsh-harness-integration.md delete mode 100644 docs/reference/git-compatibility.md delete mode 100644 docs/reference/headless-linux-server.md delete mode 100644 docs/reference/ime-regression-checklist.md delete mode 100644 docs/reference/linux-glibc-compatibility.md delete mode 100644 docs/reference/macos-press-and-hold.md delete mode 100644 docs/reference/malformed-worktree-registration-removal.md delete mode 100644 docs/reference/managed-data-accounts.md delete mode 100644 docs/reference/monaco-language-associations.md delete mode 100644 docs/reference/omp-fresh-launch.md delete mode 100644 docs/reference/omp-history-titles.md delete mode 100644 docs/reference/omp-resume-transcript-locator.md delete mode 100644 docs/reference/omp-runtime-session-provenance.md delete mode 100644 docs/reference/omp-session-roots.md delete mode 100644 docs/reference/omp-startup-keyboard-query.md delete mode 100644 docs/reference/omp-status-input-redaction.md delete mode 100644 docs/reference/orcad-operations.md delete mode 100644 docs/reference/orchestration-configured-agent-aliases.md delete mode 100644 docs/reference/pnpm-install-policy.md delete mode 100644 docs/reference/qoder-integration.md delete mode 100644 docs/reference/relay-regional-placement.md delete mode 100644 docs/reference/remote-wire-compatibility.md delete mode 100644 docs/reference/renderer-agent-status-performance.md delete mode 100644 docs/reference/runtime-file-base64-padding.md delete mode 100644 docs/reference/sharing-agent-skills.md delete mode 100644 docs/reference/spinner-rendering-performance.md delete mode 100644 docs/reference/ssh-execution-boundary.md delete mode 100644 docs/reference/ssh-host-key-verification.md delete mode 100644 docs/reference/ssh-reconnect-source-recovery.md delete mode 100644 docs/reference/task-provider-identity-validation.md delete mode 100644 docs/reference/terminal-artifact-grant-integrity.md delete mode 100644 docs/reference/terminal-perf-latency-investigation.md delete mode 100644 docs/reference/terminal-perf-report-budgets.md delete mode 100644 docs/reference/terminal-startup-timing.md delete mode 100644 docs/reference/windows-cmd-shim-resolution.md delete mode 100644 docs/reference/windows-daemon-host-relocation.md delete mode 100644 docs/reference/windows-edr-posture.md delete mode 100644 docs/reference/windows-msys-job-breakaway.md delete mode 100644 docs/reference/windows-process-enumeration.md delete mode 100644 docs/reference/windows-setup-shell.md delete mode 100644 docs/reference/windows-signing-runner-time.md delete mode 100644 docs/reference/windows-terminal-shell-selection.md delete mode 100644 docs/reference/worktree-scan-fingerprint.md delete mode 100644 docs/reference/wsl-command-execution.md delete mode 100644 docs/reference/wsl-managed-cli.md delete mode 100644 docs/reference/wsl-probe-failure-semantics.md delete mode 100644 docs/reference/wsl-runner-verification.md delete mode 100644 docs/reference/xterm-patch-regeneration.md diff --git a/.github/CONTRIBUTING.md b/.github/CONTRIBUTING.md index 43339372bcb..98c8756f585 100644 --- a/.github/CONTRIBUTING.md +++ b/.github/CONTRIBUTING.md @@ -29,7 +29,6 @@ pnpm dev Ordinary installs include native optional dependencies for the current OS and CPU only. Before a cross-architecture build (including `pnpm build:mac`, which produces both x64 and arm64 artifacts by default), run `pnpm install:release` to add the other CPU's variants. -See [the install policy](../docs/reference/pnpm-install-policy.md). ## Branch Naming diff --git a/.github/pull_request_template.md b/.github/pull_request_template.md index 7cfc408ebf2..b1dac35c0d4 100644 --- a/.github/pull_request_template.md +++ b/.github/pull_request_template.md @@ -11,6 +11,7 @@ ## Linked Issue + _If you do not have one and are an outside contributors, your PR **wiil** be ignored. Refs is not sufficient. Link an actual issue_ @@ -39,7 +40,7 @@ Fixes # ## Agent skill upstream boundary -- [ ] Not applicable, or this change follows `docs/reference/agent-skill-sharing-upstream-boundary.md` and copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation. +- [ ] Not applicable, or this change copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation. ## Notes diff --git a/.gitignore b/.gitignore index ffa4bc4f600..c15e2b4437e 100644 --- a/.gitignore +++ b/.gitignore @@ -90,10 +90,9 @@ design-docs/ # Machine-local agent hook endpoint files may contain auth tokens. /agent-hooks/ -# Local-only design/planning docs (not checked in), including most of docs/reference/. +# Local-only design/planning docs (not checked in). # Durable docs that should be tracked must live in one of the allow-listed -# locations below (assets, readme, STYLEGUIDE, mobile terminal shortcut bar, -# and the tracked reference docs linked from AGENTS.md / README.md). +# locations below (assets, readme, STYLEGUIDE, and mobile terminal shortcut bar). docs/** !docs/ # The deployable docs app is source, not local engineering notes. @@ -117,40 +116,6 @@ docs/** !docs/audits/crashpad-read-limit/source-hashes.json !docs/agent-skill-sharing-implementation-checklist.md !docs/mobile-terminal-shortcut-bar.md -!docs/reference/ -!docs/reference/agent-pty-transcript-capture.md -!docs/reference/agent-session-search-query-tuning.md -!docs/reference/agent-session-search-contract.md -!docs/reference/agent-status-store.md -!docs/reference/antigravity-readiness-evidence.md -!docs/reference/antivirus-prerelease-clearance.md -!docs/reference/cline-and-prime-agent-readiness-evidence.md -!docs/reference/git-compatibility.md -!docs/reference/headless-linux-server.md -!docs/reference/ime-regression-checklist.md -!docs/reference/jcode-hook-events.md -!docs/reference/linux-glibc-compatibility.md -!docs/reference/macos-press-and-hold.md -!docs/reference/orcad-operations.md -!docs/reference/pnpm-install-policy.md -!docs/reference/relay-grace-time-reconfiguration.md -!docs/reference/windows-cmd-shim-resolution.md -!docs/reference/windows-daemon-host-relocation.md -!docs/reference/windows-edr-posture.md -!docs/reference/windows-msys-job-breakaway.md -!docs/reference/windows-process-enumeration.md -!docs/reference/wsl-runner-verification.md -!docs/reference/remote-wire-compatibility.md -!docs/reference/renderer-agent-status-performance.md -!docs/reference/ssh-execution-boundary.md -!docs/reference/ssh-host-key-verification.md -!docs/reference/ssh-reconnect-source-recovery.md -!docs/reference/windows-setup-shell.md -!docs/reference/windows-terminal-shell-selection.md -!docs/reference/worktree-scan-fingerprint.md -!docs/reference/wsl-command-execution.md -!docs/reference/wsl-probe-failure-semantics.md -!docs/reference/xterm-patch-regeneration.md # Stably CLI (only docs/ are tracked) .stably/* diff --git a/AGENTS.md b/AGENTS.md index 451345c8a0c..e4161628503 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -73,24 +73,24 @@ Orca targets macOS, Linux, and Windows. Keep all platform-dependent behavior beh - **Keyboard shortcuts**: Never hardcode `e.metaKey`. Use a platform check (`navigator.userAgent.includes('Mac')`) to pick `metaKey` on Mac and `ctrlKey` on Linux/Windows. Electron menu accelerators should use `CmdOrCtrl`. - **Shortcut labels in UI**: Display `⌘` / `⇧` on Mac and `Ctrl+` / `Shift+` on other platforms. - **File paths**: Use `path.join` or Electron/Node path utilities — never assume `/` or `\`. -- **Windows terminal shells**: `--shell` picks the shell a terminal _is_; `--command` is typed into whatever shell the host spawned, so a shell choice routed through `command` silently becomes a child process. See [`docs/reference/windows-terminal-shell-selection.md`](./docs/reference/windows-terminal-shell-selection.md). -- **Windows setup scripts**: the setup/issue-command runner is a `.cmd` batch file unless the script starts with a `#!` line — never derive that from the user's terminal-shell preference, and never launch a `.cmd` runner with a bare `cmd.exe /c` from a Git Bash pane (MSYS rewrites the `/c`). See [`docs/reference/windows-setup-shell.md`](./docs/reference/windows-setup-shell.md). -- **Windows child processes**: start them through `runProcess`/`spawnProcess` in `src/shared/child-process/` — never `child_process` directly. It pins `windowsHide`, refuses `shell: true`, and encodes `.cmd`/`.bat` arguments so neither `CommandLineToArgvW` nor `cmd.exe` mangles them. Recognised npm/pnpm `.cmd` shims are resolved to their real target so the spawn skips `cmd.exe` entirely; see [`docs/reference/windows-cmd-shim-resolution.md`](./docs/reference/windows-cmd-shim-resolution.md) before adding a shim shape or debugging one. +- **Windows terminal shells**: `--shell` picks the shell a terminal _is_; `--command` is typed into whatever shell the host spawned, so a shell choice routed through `command` silently becomes a child process. +- **Windows setup scripts**: the setup/issue-command runner is a `.cmd` batch file unless the script starts with a `#!` line — never derive that from the user's terminal-shell preference, and never launch a `.cmd` runner with a bare `cmd.exe /c` from a Git Bash pane (MSYS rewrites the `/c`). +- **Windows child processes**: start them through `runProcess`/`spawnProcess` in `src/shared/child-process/` — never `child_process` directly. It pins `windowsHide`, refuses `shell: true`, and encodes `.cmd`/`.bat` arguments so neither `CommandLineToArgvW` nor `cmd.exe` mangles them. Recognised npm/pnpm `.cmd` shims are resolved to their real target so the spawn skips `cmd.exe` entirely. - **Ripgrep**: Orca bundles `rg` for every platform, WSL, and SSH remotes. Spawn it through `spawnBundledRipgrep` (main) or `resolveRelayRipgrepCommand` (relay), never a bare `'rg'` — Windows resolves a bare name in the spawn cwd before PATH. Don't add git/readdir fallbacks locally; the relay's chain exists only for hosts an upload never reached. -- **Windows process enumeration**: read the table through `src/main/windows/windows-process-table.ts`, never by forking `powershell.exe`. See [`docs/reference/windows-process-enumeration.md`](./docs/reference/windows-process-enumeration.md). -- **Windows MSYS/Git Bash panes**: their children break away from the per-PTY job unless it is created without `JOB_OBJECT_LIMIT_BREAKAWAY_OK`, and a `conpty.node` built before that fix passes every existing gate. Before changing the per-PTY job or debugging `windows-msys-job.win32.test.ts`, read [`docs/reference/windows-msys-job-breakaway.md`](./docs/reference/windows-msys-job-breakaway.md). -- **Windows daemon-host relocation**: the terminal daemon runs from a copy of the app runtime under `%LOCALAPPDATA%`, which is what survives an auto-update. Before touching that copy, its exe name, or the NSIS uninstall macro, read [`docs/reference/windows-daemon-host-relocation.md`](./docs/reference/windows-daemon-host-relocation.md). -- **Windows EDR signal**: don't add `-ExecutionPolicy Bypass`, `-EncodedCommand`, `cmd.exe /c` with escaped free text, per-operation interpreter spawning, or runtime `Add-Type` compilation without reading [`docs/reference/windows-edr-posture.md`](./docs/reference/windows-edr-posture.md) first — behavioural EDR scores each of those, and being signed does not clear them. For file verdicts on the bytes we ship — antivirus false positives, and the vendor programs that clear a release before users meet the detection — see [`docs/reference/antivirus-prerelease-clearance.md`](./docs/reference/antivirus-prerelease-clearance.md). -- **WSL commands**: build argv with `buildWslExecArgs` (always `--exec` — under `--`, `wsl.exe` expands `$name` in every argument and silently rewrites the script), and fence anything whose stdout you parse with `buildWslCapturedLoginShellCommand`, because the interactive login shell prints the distro banner to stdout. See [`docs/reference/wsl-command-execution.md`](./docs/reference/wsl-command-execution.md). -- **Linux native modules**: keep the glibc floor at Ubuntu 20.04 / glibc 2.31. A module compiled from source on a newer runner can reference symbol versions absent on the floor and crash the app on startup. See [`docs/reference/linux-glibc-compatibility.md`](./docs/reference/linux-glibc-compatibility.md); packaging fails if a bundled native binary needs newer glibc. +- **Windows process enumeration**: read the table through `src/main/windows/windows-process-table.ts`, never by forking `powershell.exe`. +- **Windows MSYS/Git Bash panes**: their children break away from the per-PTY job unless it is created without `JOB_OBJECT_LIMIT_BREAKAWAY_OK`, and a `conpty.node` built before that fix passes every existing gate. +- **Windows daemon-host relocation**: the terminal daemon runs from a copy of the app runtime under `%LOCALAPPDATA%`, which is what survives an auto-update. +- **Windows EDR signal**: don't add `-ExecutionPolicy Bypass`, `-EncodedCommand`, `cmd.exe /c` with escaped free text, per-operation interpreter spawning, or runtime `Add-Type` compilation without reviewing the EDR impact first — behavioural EDR scores each of those, and being signed does not clear them. +- **WSL commands**: build argv with `buildWslExecArgs` (always `--exec` — under `--`, `wsl.exe` expands `$name` in every argument and silently rewrites the script), and fence anything whose stdout you parse with `buildWslCapturedLoginShellCommand`, because the interactive login shell prints the distro banner to stdout. +- **Linux native modules**: keep the glibc floor at Ubuntu 20.04 / glibc 2.31. A module compiled from source on a newer runner can reference symbol versions absent on the floor and crash the app on startup. Packaging fails if a bundled native binary needs newer glibc. ## Native Dependency Installs -Ordinary `pnpm install` covers the host OS and CPU only. Before packaging for another architecture — including `pnpm build:mac`, which builds x64 and arm64 by default — run `pnpm install:release`. electron-builder only warns on a missing `extraResources` source, so the `beforePack` guard is what turns a thin install into a build failure instead of a silently broken artifact; see [`docs/reference/pnpm-install-policy.md`](./docs/reference/pnpm-install-policy.md). +Ordinary `pnpm install` covers the host OS and CPU only. Before packaging for another architecture — including `pnpm build:mac`, which builds x64 and arm64 by default — run `pnpm install:release`. electron-builder only warns on a missing `extraResources` source, so the `beforePack` guard is what turns a thin install into a build failure instead of a silently broken artifact. ## SSH Use Case -All changes must consider the SSH use case. Don't assume local-only execution. Before changing anything that reports on, stops, or lists remote work, follow [`docs/reference/ssh-execution-boundary.md`](./docs/reference/ssh-execution-boundary.md): the execution host owns everything that touches execution, and loss of contact is never evidence of process death — the verdict vocabulary is `live` / `unverifiable` / `exited`, with no synonyms. +All changes must consider the SSH use case. Don't assume local-only execution. Before changing anything that reports on, stops, or lists remote work, the execution host owns everything that touches execution, and loss of contact is never evidence of process death — the verdict vocabulary is `live` / `unverifiable` / `exited`, with no synonyms. ## Folder Workspace Use Case @@ -98,19 +98,19 @@ All changes must consider folder workspaces as well as git worktrees. Don't assu ## Agent Status -The execution host owns agent status in one store, the hook server's, and every reader (sidebar, `worktree ps`, mobile, dashboard) subscribes to it. Before adding a producer, a cache, or a reader-side precedence rule, read [`docs/reference/agent-status-store.md`](./docs/reference/agent-status-store.md): new producers write into that store, and readers keep only presentation policy. +The execution host owns agent status in one store, the hook server's, and every reader (sidebar, `worktree ps`, mobile, dashboard) subscribes to it. New producers write into that store, and readers keep only presentation policy. ## Agent Terminal Screens -A rule that reads what an agent CLI paints on a terminal — readiness, blocked prompts, idle — must be written against a captured transcript, not a remembered screen. Record one with [`docs/reference/agent-pty-transcript-capture.md`](./docs/reference/agent-pty-transcript-capture.md), which keeps escapes and wrapping intact and scrubs account identifiers before they reach git. Antigravity readiness has no transcript yet and five failed attempts without one; before touching it, read [`docs/reference/antigravity-readiness-evidence.md`](./docs/reference/antigravity-readiness-evidence.md). +A rule that reads what an agent CLI paints on a terminal — readiness, blocked prompts, idle — must be written against a captured transcript, not a remembered screen. Use `config/scripts/capture-agent-pty-transcript.mjs` to keep escapes and wrapping intact and scrub account identifiers before they reach git. ## Remote Wire Compatibility -Clients and remote Orca servers update independently, so mixed versions are the normal state. Before changing anything a paired client and host exchange — RPC params, stream frames, or the content either side publishes over them — follow [`docs/reference/remote-wire-compatibility.md`](./docs/reference/remote-wire-compatibility.md). A new optional field is safe; a new stream opcode must be capability-negotiated because decoders drop unknown opcodes silently; and changing what the host publishes reaches old clients even with no wire change. +Clients and remote Orca servers update independently, so mixed versions are the normal state. Before changing anything a paired client and host exchange — RPC params, stream frames, or the content either side publishes over them — preserve compatibility with older peers. A new optional field is safe; a new stream opcode must be capability-negotiated because decoders drop unknown opcodes silently; and changing what the host publishes reaches old clients even with no wire change. ## Git Binary Compatibility -Orca runs the user's Git binary on native, WSL, and SSH hosts, which may all have different versions. Treat Git 2.25 as the core-workflow baseline and follow [`docs/reference/git-compatibility.md`](./docs/reference/git-compatibility.md). +Orca runs the user's Git binary on native, WSL, and SSH hosts, which may all have different versions. Treat Git 2.25 as the core-workflow baseline. When adding or changing a Git command: diff --git a/README.md b/README.md index 4c215157da9..906fbc2d657 100644 --- a/README.md +++ b/README.md @@ -218,7 +218,6 @@ Works with **any CLI agent** — if it runs in a terminal, it runs in Orca. - **[Download from onOrca.dev](https://onorca.dev/download)** - Or grab a build directly: [macOS Apple Silicon](https://github.com/stablyai/orca/releases/latest/download/orca-macos-arm64.dmg) · [macOS Intel](https://github.com/stablyai/orca/releases/latest/download/orca-macos-x64.dmg) · [Windows (.exe)](https://github.com/stablyai/orca/releases/latest/download/orca-windows-setup.exe) · [Linux AppImage](https://github.com/stablyai/orca/releases/latest/download/orca-linux.AppImage) · [All builds](https://github.com/stablyai/orca/releases/latest) -- Running `orca serve` on a headless Linux server? See the [headless Linux server guide](docs/reference/headless-linux-server.md). _Or via a package manager:_ diff --git a/config/reliability-gates.jsonc b/config/reliability-gates.jsonc index 273995d4399..3b415f60d58 100644 --- a/config/reliability-gates.jsonc +++ b/config/reliability-gates.jsonc @@ -5486,8 +5486,7 @@ "coveredProviders": ["remote-runtime"], "coverageNotes": "Deterministic main-IPC contract tests cover disconnect-driven close delivery, exactly-once close, per-subscription teardown isolation against a failing socket close and a throwing liveness probe, containment of a throwing renderer send on the unguarded host-close path, and continued suppression of stale payloads from a retired transport. A headed paired-server journey (real Orca host plus a separate paired Orca desktop client) covers hidden-but-mounted reveal, cold-parked reveal, and cold-parked reveal across a disconnect/reconnect. Live Linux and Windows paired-server evidence and real sleep/wake transport loss remain uncollected.", "motivatingLinks": [ - "tests/e2e/paired-remote-terminal-parked-reveal-interactivity.spec.ts", - "docs/reference/headless-linux-server.md" + "tests/e2e/paired-remote-terminal-parked-reveal-interactivity.spec.ts" ], "invariant": "Every renderer-held runtime subscription receives exactly one terminal close event when its transport is retired, including when the retirement advanced the transport generation first, and a single failing teardown never abandons that environment's remaining subscriptions nor escapes into the transport that reported the close. Payload frames from a retired transport stay suppressed. A revealed remote terminal therefore reattaches over a live multiplex connection: its buffer restores, typed input reaches the host PTY, the echo paints without a tab flip, and the PTY converges on the revealed pane grid.", "oracle": "The main IPC contract test subscribes terminal.multiplex through the real handler, disconnects the environment, and asserts the renderer received exactly one {type: close} subscription event. Two isolation tests subscribe a second stream to the same environment and make the first one fail -- in its socket close, and in the liveness probe inside notifyClosed -- then assert the disconnect does not throw, both transports closed, and every close the renderer could still receive was delivered. A third drives a host-initiated close through the transport callback, which is the one notifyClosed call site with no surrounding guard, with a renderer send that throws, and asserts it cannot escape into the WebSocket close handler. A fourth test asserts that after retirement a late response frame is not forwarded and a late transport close does not re-send. The paired-server journey runs three reveal scenarios against one real host and one real paired desktop client, and for each records buffer restore, host-side receipt of the typed marker through an out-of-band host sink file, live paint without a tab flip, and PTY-versus-pane grid convergence.", diff --git a/docs/readme/README.fr.md b/docs/readme/README.fr.md index 71689cc0ceb..4958c969144 100644 --- a/docs/readme/README.fr.md +++ b/docs/readme/README.fr.md @@ -222,7 +222,6 @@ Fonctionne avec **n'importe quel agent CLI** — s'il tourne dans un terminal, i - **[Télécharger depuis onOrca.dev](https://onorca.dev/download)** - Ou récupérez un build directement : [macOS Apple Silicon](https://github.com/stablyai/orca/releases/latest/download/orca-macos-arm64.dmg) · [macOS Intel](https://github.com/stablyai/orca/releases/latest/download/orca-macos-x64.dmg) · [Windows (.exe)](https://github.com/stablyai/orca/releases/download/v1.4.147-rc.3/orca-windows-setup.exe) · [Linux AppImage](https://github.com/stablyai/orca/releases/latest/download/orca-linux.AppImage) · [Tous les builds](https://github.com/stablyai/orca/releases/latest) - **Sous Windows :** utilisez la [dernière RC (`v1.4.147-rc.3`)](https://github.com/stablyai/orca/releases#release-v1.4.147-rc.3) — elle inclut des correctifs Windows absents de la stable. -- Vous lancez `orca serve` sur un serveur Linux headless ? Consultez le [guide serveur Linux headless](../reference/headless-linux-server.md). _Ou via un gestionnaire de paquets :_ diff --git a/docs/readme/README.ko.md b/docs/readme/README.ko.md index e984ceffa30..1207d226da0 100644 --- a/docs/readme/README.ko.md +++ b/docs/readme/README.ko.md @@ -217,7 +217,6 @@ diff의 어느 줄에든 코멘트를 남기고 에이전트에게 바로 보내 - **[onOrca.dev에서 다운로드](https://onorca.dev/download)** - 또는 빌드를 직접 받기: [macOS Apple Silicon](https://github.com/stablyai/orca/releases/latest/download/orca-macos-arm64.dmg) · [macOS Intel](https://github.com/stablyai/orca/releases/latest/download/orca-macos-x64.dmg) · [Windows (.exe)](https://github.com/stablyai/orca/releases/latest/download/orca-windows-setup.exe) · [Linux AppImage](https://github.com/stablyai/orca/releases/latest/download/orca-linux.AppImage) · [전체 빌드](https://github.com/stablyai/orca/releases/latest) -- headless Linux 서버에서 `orca serve`를 실행하시나요? [Headless Linux 서버 가이드](../reference/headless-linux-server.md)를 확인하세요. _또는 패키지 매니저로 설치:_ diff --git a/docs/readme/README.pt.md b/docs/readme/README.pt.md index 3e3d0e6b1a5..7a9c7dc3d19 100644 --- a/docs/readme/README.pt.md +++ b/docs/readme/README.pt.md @@ -217,7 +217,6 @@ Funciona com **qualquer agente CLI** — se roda em um terminal, roda no Orca. - **[Baixe em onOrca.dev](https://onorca.dev/download)** - Ou baixe um build diretamente: [macOS Apple Silicon](https://github.com/stablyai/orca/releases/latest/download/orca-macos-arm64.dmg) · [macOS Intel](https://github.com/stablyai/orca/releases/latest/download/orca-macos-x64.dmg) · [Windows (.exe)](https://github.com/stablyai/orca/releases/latest/download/orca-windows-setup.exe) · [Linux AppImage](https://github.com/stablyai/orca/releases/latest/download/orca-linux.AppImage) · [Todos os builds](https://github.com/stablyai/orca/releases/latest) -- Rodando `orca serve` em um servidor Linux headless? Veja o [guia de servidor Linux headless](../reference/headless-linux-server.md). _Ou por um gerenciador de pacotes:_ diff --git a/docs/reference/admin-agent-skill-sharing.md b/docs/reference/admin-agent-skill-sharing.md deleted file mode 100644 index 1ea9a65d8ef..00000000000 --- a/docs/reference/admin-agent-skill-sharing.md +++ /dev/null @@ -1,97 +0,0 @@ -# Administer agent skill sharing - -This guide describes the first-release access, lifecycle, retention, and recovery contract for -Orca skill sharing. The operator runbook remains the source of truth for incident commands and -environment-specific procedures. - -## Access model - -- Shared bundles are unlisted bearer resources. Orca provides no public browse, search, recipient - inventory, or package index. -- Anyone with an active, unexpired link can inspect the package and request a short-lived download - grant without signing in. -- Publishing, package/version management, owned-link inventory, revocation, and deletion require - an authenticated package owner with current organization access. -- Missing, expired, revoked, deleted, and unauthorized resources return the same non-disclosing - response. -- Desktop and remote runtimes receive no GCP identity or long-lived storage credential. - -The durable share ID is a credential. Do not put it in tickets, logs, analytics, or support -bundles. Use revocation if a link may have reached an unintended recipient. - -## Revocation and deletion - -Revoking a share immediately blocks new resolution and download grants. A generation-bound grant -issued before revocation can work until its five-minute expiry. Already installed skills remain on -recipient machines. - -Package deletion follows this order: - -1. Mark the package deleted and revoke its active shares. -2. Dereference retained versions transactionally. -3. Delete only an object generation that no retained version references. -4. Reconcile bounded pending deletions after partial database or GCS failures. - -A version cannot be deleted while an active pinned share references it. Deletion uses the exact -recorded GCS generation and never overwrites an immutable published key. - -## User and organization departure - -Packages belong to an owner tenant and record the publishing user. In an organization tenant, -another current member can manage the package after its publisher leaves; Orca does not rewrite -the recorded creator. Removing a user does not automatically revoke the organization's links, -delete packages, or remove installed copies. - -Before deleting an organization tenant: - -1. Disable new grants for the tenant. -2. Have an authorized operator inventory and revoke active shares. -3. Decide whether packages transfer to another authorized owner, remain retained, or are deleted. -4. Resolve legal hold, erasure, and audit-retention requirements. -5. Apply the coordinated metadata and object lifecycle; do not bypass reference checks. - -The product does not yet encode a universal ownership-transfer or legal-retention policy. Privacy, -security, and the organization owner must approve the applicable policy before external rollout. -Until that decision is recorded, preserve metadata and soft-deleted generations rather than -guessing. - -## Retention contract - -| Data | Default behavior | -| -------------------------------------- | --------------------------------------------------------- | -| Upload policy and pending upload row | Expires after 15 minutes | -| Abandoned `uploads/` quarantine object | Deleted by GCS after one day | -| Published immutable package object | No age-based deletion; retained while referenced | -| Issued download grant | Expires after five minutes | -| Deleted package object | Recoverable through GCS soft delete for seven days | -| PostgreSQL metadata | Covered by backups and seven-day point-in-time recovery | -| Installed recipient copy | Independent local data; Cloud deletion does not remove it | -| Audit event | Follows the approved audit-retention policy | - -Organization retention, legal hold, and erasure requirements take precedence over product rollback -retention. Product deletion is not a legal-hold mechanism. - -## Audit and privacy - -Audit records may include package/version IDs, actor IDs, event category, outcome, and timestamp. -They must not include skill contents, filenames, manifests, organization membership lists, local -paths, durable share URLs, upload policies, download grants, or credentials. Anonymous abuse -controls must not persist raw requester IP addresses. - -Normal Cloud logs are limited to route, method, status, duration, and bounded failure categories. -Use seeded privacy canaries when validating staging logs and diagnostic exports. - -## Recovery and incident controls - -Upload, download, and remote-install operations have independent kill switches. Disable the -narrowest affected operation; existing local discovery and installs continue to work. - -Coordinate PostgreSQL point-in-time recovery with GCS generation recovery. Restore metadata into -an isolated database, identify exact referenced generations, restore only matching soft-deleted -objects, verify archive and package identities, then transactionally repoint metadata. Keep grants -disabled until bearer preview and a generation-bound download pass. - -See the Orca Cloud `docs/skill-sharing-runbook.md` for deployment controls, reconciliation, -saturation, signing failures, database outages, and the guarded restore workflow. Security -invariants and unresolved release gates are recorded in -[Agent skill sharing threat model](./agent-skill-sharing-threat-model.md). diff --git a/docs/reference/agent-pty-transcript-capture.md b/docs/reference/agent-pty-transcript-capture.md deleted file mode 100644 index 4b0f41c46d3..00000000000 --- a/docs/reference/agent-pty-transcript-capture.md +++ /dev/null @@ -1,131 +0,0 @@ -# Capturing an agent PTY transcript - -Orca's readiness and blocked-prompt rules are text rules over what an agent CLI paints on a -terminal. They are only as good as the screens they were written against. This is how to record -one, byte for byte, so a rule can be pinned to evidence instead of to a remembered screen. - -Related: [`antigravity-readiness-evidence.md`](./antigravity-readiness-evidence.md) names the -specific Antigravity transcripts that are still missing and what each one decides. - -## The recorder - -``` -node config/scripts/capture-agent-pty-transcript.mjs --name [options] -- [args...] -``` - -It allocates a real PTY, spawns the agent inside it, mirrors the session to your terminal so you -can drive it by hand, and appends every byte it receives to -`src/main/runtime/__fixtures__/.txt`. It does not strip escapes, fold `\r`, rewrap -lines, or normalise anything — the file is what the terminal received. - -- **Ending a capture:** press Ctrl+]. The recorder consumes that key and - never forwards it, which is the only way to end a capture _while a dialog still owns the - screen_. Quitting the agent instead would first dismiss the dialog you came to record. -- `--cols N --rows M` pin the PTY size (default: your terminal's). Wrapping is part of the - evidence, so record the size — the sidecar does it for you. -- `--duration S` stops unattended after S seconds, for a screen that needs no interaction. -- `--send ":"` types into the PTY at a fixed offset, repeatable, with `\r` `\n` `\t` `\e` - escapes. A dialog capture has to be driven, and an unattended run (CI, or an agent) has no TTY to - type into; the keystrokes ride the same PTY a human's would. For example, the committed - `antigravity-dialog-model-picker.txt` was recorded with - `--duration 24 --send "14000:/model" --send "16000:\r"`, which leaves the picker owning the - screen when the capture stops. -- `--note ""` records the account type, plan, model and CLI version in the sidecar. -- `--out ` writes outside the fixture directory (use it for a first dry run). - -Each capture also writes `.meta.json` with the timestamp, platform, command, -PTY size, note and exit code. Commit it with the transcript; the version and account type behind -a screen are not recoverable from the bytes. - -**Prerequisite:** `node-pty` must be built for plain Node: - -``` -node config/scripts/ensure-native-runtime.mjs --runtime=node -``` - -Orca itself does not need to be running, and the recorder never touches Orca state. - -### Platform notes - -- **macOS / Linux:** nothing special. `TERM=xterm-256color` is set for the child. -- **Windows:** run it from Windows Terminal / PowerShell, not a Git Bash (MSYS) pane — MSYS - rewrites arguments that start with `/`, which mangles the `cmd.exe /c` hand-off. A `.cmd` or - `.bat` agent shim cannot be spawned by node-pty directly, so the recorder routes those through - `cmd.exe` for you. -- **WSL:** capture _inside_ the distro (run the recorder from the distro's checkout). Recording - `wsl.exe` from the Windows side adds the login-shell banner to the transcript. -- **SSH:** record on the execution host. A transcript recorded locally is not evidence about what - a remote agent prints. - -## Privacy: scrub before committing - -A live agent screen routinely contains things that must not enter git history: - -| Scrub | Why | -| ---------------------------------------------------------------------- | ---------------------------------------------------- | -| Account email / sign-in identifier | The account row on a ready screen prints it verbatim | -| Org, tenant or team name | Identifies a customer | -| Machine hostname and OS username | Appear in prompts, paths and the OSC title | -| Absolute home paths (`/Users/`, `C:\Users\`) | Contain the username | -| JWTs, `AIza…` keys, `1//…` refresh tokens, `Bearer …`, `sk-…`, `ghp_…` | Live credentials; a sign-in screen can echo one | -| Private repo, branch and ticket names | Leak roadmap detail | -| Anything you pasted into the agent during the capture | You typed it; it is in the transcript | - -The recorder scans the file as soon as the capture ends and prints every hit with a line and -column. To scrub: - -``` -node config/scripts/capture-agent-pty-transcript.mjs --scan src/main/runtime/__fixtures__/.txt --redact -``` - -Redaction replaces each finding with a **same-length** placeholder (`u…u@example.com`, `XXXX…`). -Length matters: a transcript's value is its exact wrapping and column alignment, and a shorter -replacement reflows the screen and destroys the evidence. - -### Verify it is gone - -1. `node config/scripts/capture-agent-pty-transcript.mjs --scan src/main/runtime/__fixtures__/.txt` - must print `clean` and exit `0`. It recognises its own placeholders, so a scrubbed file passes. -2. Grep for the specifics the scanner cannot know: - `rg -n -i -- "$(whoami)|||" src/main/runtime/__fixtures__/.txt` -3. Read it once with escapes visible: `LC_ALL=C cat -v src/main/runtime/__fixtures__/.txt`. - The scanner matches shapes; only a human catches a project name. -4. Check the sidecar too — `--note` text is free-form and is committed. - -`config/scripts/pty-transcript-secret-scan.test.mjs` re-scans every committed -`__fixtures__/*.txt`, so a transcript that skips step 1 fails the suite. - -## Consuming a transcript in a test - -Feed the raw bytes through the runtime rather than into a matcher directly: escape handling, -tail retention and title tracking all live in `onPtyData`, and a rule tested on pre-normalised -text is tested on something no pane ever sees. - -`src/main/runtime/agent-transcript-pane-test-harness.ts` builds the pane; -`src/main/runtime/terminal-interactive-wait-visibility.test.ts` (cursor-agent) and -`src/main/runtime/antigravity-readiness-transcripts.test.ts` (Antigravity) are examples. An agent -whose readiness is read off the live screen gets its suite from -`src/main/runtime/screen-ruled-agent-transcript-suite.ts`. - -## Worked example: the Antigravity captures - -The six committed `antigravity-*.txt` fixtures were recorded this way on macOS against -`agy` 1.1.25. Two points generalise: - -- **Reach a state without mutating the operator's config.** The ready-screen captures ran in a - directory the CLI already trusted, so no trust answer was written. Where a dialog could only be - reached by signing the operator out or deleting their settings, it was left uncaptured and - recorded as such rather than forced. -- **An environment variable is a legitimate capture knob** where a setting is not. - `AGY_CLI_HIDE_ACCOUNT_INFO=1` produced a second ready screen with no account row, which is - evidence no amount of reasoning about the first screen could have supplied. It changes nothing - on disk. - -## Known gap in the existing captures - -The three `cursor-agent-*.txt` fixtures contain **no escape bytes and no carriage returns**. -Whatever produced them went through a renderer and a clipboard, so they preserve wording and -box-drawing glyphs but not the caret, the cursor moves, the repaints, or whether the CLI uses the -alternate screen buffer. They are good enough for the wording-based rules built on them and are -not evidence for anything else. New captures made with this recorder keep those bytes; the -Antigravity scaffold asserts their presence so a pasted screen cannot pass as a capture. diff --git a/docs/reference/agent-session-search-contract.md b/docs/reference/agent-session-search-contract.md deleted file mode 100644 index daaef0fb618..00000000000 --- a/docs/reference/agent-session-search-contract.md +++ /dev/null @@ -1,172 +0,0 @@ -# Agent session search contract - -`AiVaultSearchRequest`, `AiVaultSearchResponse`, `AiVaultSearchHit`, and -`AiVaultSearchStatus` are defined in `src/shared/ai-vault-search-types.ts` and -validated by `src/shared/ai-vault-search-contract.ts`. - -## Search and pagination - -- Tool output beyond 3,072 characters per row is not indexed and not searchable; user and assistant text is indexed in full. -- A page cursor outstanding during a retention purge is refused once as `stale-cursor`; the client re-issues page 1. -- A phrase match across a chunk boundary of a long message is not supported. - -`aiVault.searchSessions(request)` accepts `query`, optional `scope` -(`conversation` or `all`, default `all`), `freshness` (`indexed` or -`wait-until-current`, default `indexed`), `limit`, opaque `cursor`, `filters`, -and `debug` (default false). Conversation scope searches user and assistant text. -Filters accept `agents`, `scopePaths`, ISO `since`, and `sort` (`relevance` or -`newest`). Paths refer to the execution host and work for folders without Git. -Legacy `tier` and `refresh` fields are accepted and discarded; they do not change -the defaults. Limits use the engine's resolver: default 20, integers clamped to -1–100, fractional numbers use the default. Long queries reach the engine so it -can report truncation rather than fail validation. - -Results contain `kind: 'results'`, `hits`, `page: { cursor, hasMore }`, -`generation`, `truncated: { candidates, snippets, query, freshness }`, and -`durationMs`. `snippets` is a count; the other truncation fields are booleans. -`durationMs` measures the engine search, excluding any reconciliation wait. -`debug: true` adds `debug: { route, repairedTerms?, plannerReport }`; the report -contains `route`, optional `repairedTerms`, and `scope`. Diagnostics never appear -at the top level. Status is never attached to search results. - -A cursor belongs to one query, one host's index generation, and an opaque persisted -index incarnation. Query, scope, filters, and sorting must remain the same; page -size may change. Writes that advance the generation can invalidate it, including -retention purges. Clearing or rebuilding the database invalidates it even when the -new generation counter matches. A refused cursor yields -`{ kind: 'stale-cursor', generation, expectedGeneration? }` and the -client discards it and issues page 1 without a cursor. Reusing that refused cursor -continues to fail; there is no server-side cursor acknowledgement state. -Malformed cursors and cursors for a different query yield -`{ kind: 'malformed-cursor' }`. Generation checks also reject a first page if the -index changes during retrieval. Generation is a fence, not a retained snapshot: -a client cannot ask the host to recreate a previous generation. - -Pages are per host only. Ordering is local to that host's query. Clients must -discard cursors when changing hosts. All-computers search, merged ordering, -per-host aggregate outcomes, and merged cursors are deferred to a separate PR. -That follow-up must define generation fencing, page-size changes, unavailable -hosts, and bounded parallel retrieval before exposing a combined result list. - -## Execution host routing - -Search and status address one execution host: `local`, `ssh:`, or -`runtime:`. An omitted host means this desktop's local index. -Invalid IDs and `all` are refused; neither can widen a request to other hosts. - -- `local` searches this machine's index over desktop IPC. -- `ssh:` asks that relay session and nothing else. -- `runtime:` asks that paired runtime over its RPC. A paired - runtime answers for itself and never forwards through another desktop. - -Each hit may carry `executionHostId`. The desktop stamps remote answers with -the host it addressed rather than trusting an ID returned by that host. -Local answers and older hosts may omit attribution. - -## Evidence and exposure - -Each hit carries agent, session ID, title, cwd, branch, updated time, message -count, score, source, and evidence. Evidence contains snippet, role, and timestamp; -it is null for operator-only matches that have no text evidence. Snippet matches -use `[[` and `]]` markers. Source presence is `present`, `unverifiable`, or -`missing`; the current engine emits the first two. Loss of contact does not prove -a source missing. - -`redactForTransport(hit, transport)` is the exposure policy: - -| Transport | filePath / codexHome | resumeCommand | Status `degradedRoots[].root` | -| ----------------------------------------- | ------------------------------------ | --------------------------------- | ----------------------------- | -| Desktop IPC on the same machine | Included when known, under source | Included only for present sources | Included | -| Runtime RPC on the same machine | Included when known, under source | Included only for present sources | Included | -| Relay or paired runtime/web/mobile client | Withheld; source keeps presence only | Withheld | Withheld | - -`cwd`, titles, snippets, and other hit metadata remain visible to paired clients. -Snippets cross the authenticated transport as indexed; this contract does not -apply an observability redactor to transcript content. A missing Codex home is -omitted. Resume commands reuse the command stored by the transcript reader, -constructed by the sidebar's `buildAiVaultResumeCommand`; this layer does not -construct commands or execute them. The runtime uses its authenticated -`clientKind` context to distinguish paired clients from same-machine RPC, never -a request-supplied locality flag. The receiving remote client also applies the -same exposure function. - -## Status, freshness, and availability - -`aiVault.searchStatus()` returns `enabled`, `phase` (`idle`, `indexing`, `current`, -`degraded`, or `closed`), `filesIndexed`, `filesDue`, `filesFailed`, `degradedRoots` -(`root` and `reason`), `lastReconcileAt`, `lastSweepCompletedAt`, and `generation`. -Times are milliseconds since epoch or null. These are the indexer's observations; -an indexed row is not a new filesystem verification. A degraded root's `root` and -raw `reason` can both contain host filesystem paths. Desktop IPC and same-machine runtime RPC receive the full diagnostic; -relay and paired clients receive only the fixed reason "Source root could not -be verified." for each degraded root. The array length retains the count. - -`wait-until-current` calls `service.reconcile()` before searching. The adapter -uses `indexer.reconcile({ full: false })`. After five seconds the endpoint searches -anyway and sets `truncated.freshness: true` on results. It does not cancel the -host's reconciliation. Completion before the deadline leaves the flag false; -a reconciliation error before the deadline propagates. The indexer's bounded -recent pass is not a promise that the entire historical corpus was swept. - -Search unavailability is a value: -`{ kind: 'unavailable', reason: 'disabled' | 'not-ready' | 'no-service' }`. -No registered service returns `no-service`. Status without a service has -`enabled: false`, `phase: 'idle'`, zero counts and generation, empty degraded roots, -and null timestamps. This is a sentinel for an absent service, not a claim of an -empty, current index. A registered service may report disabled or not-ready. - -## Boundaries and compatibility - -- Desktop: `aiVault:searchSessions` and `aiVault:searchStatus`, via preload. -- Runtime and relay: `aiVault.searchSessions` and `aiVault.searchStatus`. -- CLI: `orca search` calls both over the runtime RPC, against the host that - `--environment` / `--pairing-code` selects and no other. It reuses - `createSessionSearchClient`, so an old host's refusal reaches the caller as - `unavailable/no-service` rather than an error, and needs no new capability. - In an Orca SSH terminal, the forwarded CLI defaults to the controlling Orca - runtime's index. `--path` filters that index; it does not select the SSH host. - `--environment` / `--pairing-code` can explicitly select a paired runtime. -- Desktop preload optionally accepts an execution host scope as a separate - routing argument. It addresses exactly that host; missing connections never - fall back to the local index. The web preload addresses its own paired runtime, - answers `all` and any other host with - `unavailable/no-service` rather than an error. - -Requests and responses are parsed where received from another process. Existing -relay JSON-RPC request/response framing needs no new stream opcode. Following the -existing relay method probe pattern, an old host's explicit `-32601` refusal (or -runtime `method_not_found`) maps to `unavailable/no-service` on the client; status -uses the absent-service sentinel above. Transport failures, authentication errors, -and invalid payloads remain errors. Unknown request fields are stripped for wire -compatibility. - -The process-local `setSessionSearchService(service | null)` registry connects -these endpoints to the production service installed by PR 3b. Desktop indexing -runs in the scanner child; orcad and the SSH relay register their own in-process -services. Registration alone does not grant consent. - -## Desktop index controls (PR8) - -Settings → Agent Session History controls this desktop's persisted -`aiVaultSearch { enabled, historyDays }` policy. It stays local even when another -execution host is selected. Paired clients cannot grant consent or clear an index -through this surface; SSH relay registration remains disabled without a separate -host consent mechanism. - -`aiVault.clearSearchIndex()` is a no-argument desktop-only preload operation over -`aiVault:clearSearchIndex`. It addresses the local scanner child, not the selected -remote host. The child's existing interactive request lane executes `searchClear` -through `SessionSearchInstance.clear()`: close the indexer and database handles, -remove the SQLite database and sidecars, then reconstruct only if consent remains -enabled. Errors propagate to the settings pane. This operation never deletes -original transcripts. There is no new runtime or relay method. Opaque cursors also -carry a persistent database identity: clearing creates a new identity, so a -pre-clear or legacy cursor returns `stale-cursor` even if the rebuilt numeric -generation happens to match. Reopening the same database preserves its identity. - -Disabling closes the indexer and keeps the index copy; clearing deletes the copy. -Changing retention reuses the existing close-and-construct policy application. -The settings pane reads status only while enabled and visible, polls only an -observed indexing phase, and stops on completion or error. Opening the pane, -changing policy, or pressing Refresh obtains a new observation. The indexer's own -schedule does not depend on the pane. diff --git a/docs/reference/agent-session-search-query-tuning.md b/docs/reference/agent-session-search-query-tuning.md deleted file mode 100644 index e7895e0382d..00000000000 --- a/docs/reference/agent-session-search-query-tuning.md +++ /dev/null @@ -1,229 +0,0 @@ -# Agent session search: query tuning - -What a search costs, and what the knobs in `src/main/ai-vault-search/session-search-engine.ts` -buy. Every number here comes from `config/scripts/session-search-query-benchmark.ts` -over the synthetic corpus in `session-search-synthetic-corpus.ts`, except the -`conversation_fts` shoot-out, which writes its own corpus because the answer -turns on how much of a transcript is tool output. Nothing in this file was -measured against a real transcript, and neither benchmark must ever be pointed -at one. - -## Running it - -The benchmark is a top-level-await module that imports the main-process tree by -extensionless path, so it needs a bundler-backed runner rather than bare `node`: - -```sh -cat > src/main/ai-vault-search/zz-bench.test.ts <<'EOF' -import { it } from 'vitest' -it('runs', { timeout: 1_800_000 }, async () => { - await import('../../../config/scripts/session-search-query-benchmark') -}) -EOF -BENCH_OUT=/tmp/ss-query-bench.json pnpm test src/main/ai-vault-search/zz-bench.test.ts -rm src/main/ai-vault-search/zz-bench.test.ts -``` - -The `conversation_fts` shoot-out below runs the same way, importing -`config/scripts/session-search-conversation-fts-benchmark` instead, with -`CORPUS_MB` and `TOOL_SHARE` to size and shape its corpus. `config/scripts` is -not inside any typecheck project, so while that throwaway test exists `tsc` -reports TS6307 for each script it pulls in; delete it and the run is clean -again. - -`BENCH_OUT` exists because vitest intercepts `console.log`; the report is written -to that path as well as printed. - -## Scope: what the second FTS table buys a reader - -Corpus: 40 synthetic Claude transcripts, 10.5 MB, 9,600 messages, indexed through -the real store. Eight queries, one per rung of the route ladder plus the two -shapes that skip it; 5 warm-up runs and 25 samples each. Apple silicon, warm page -cache, machine otherwise idle. Milliseconds, and p95 over 25 samples moves -several milliseconds run to run if anything else is competing for the disk. - -| Scope | p50 | p95 | -| -------------- | ---- | ---- | -| `all` | 7.22 | 8.94 | -| `conversation` | 5.33 | 7.86 | - -Per query, `all` then `conversation` (p50 / p95): - -| Query | `all` | `conversation` | -| ------------------------------------------------ | ------------ | -------------- | -| `"terminal reattach"` (phrase) | 5.24 / 8.42 | 2.97 / 3.24 | -| `resolveTerminalPath` (identifier) | 7.55 / 8.94 | 6.47 / 6.72 | -| `src/main/…/session-transcript-reader.ts` (path) | 8.69 / 10.12 | 7.78 / 8.04 | -| `why is the daemon snapshot stale` (prose) | 7.84 / 8.57 | 5.90 / 7.01 | -| `reattahc worktre` (typo repair) | 7.30 / 7.39 | 5.53 / 5.89 | -| `index` (common term) | 5.45 / 5.66 | 3.81 / 4.02 | -| `repo:app-3` (operator only) | 0.12 / 0.16 | 0.10 / 0.10 | -| `worktree` scoped to one cwd | 1.47 / 1.63 | 1.25 / 1.49 | - -Reading it: - -- `conversation` is about 1.4x faster at p50 and 1.1x at p95, and it is a column - filter over the same table rather than a table of its own. Narrowing to the - two prose columns is what buys the gap: fewer postings to score. It is also - the scope where a match is something a person wrote rather than something a - tool printed. -- A `scopePaths` query is the cheapest real search on the page. It is the one - narrowing SQL can express exactly, so it seeks `sessions_cwd_key` and hands - ranking a small candidate set. -- The operator-only figure is a floor, not a typical cost. `repo:` and `path:` - are applied in JS over retrieved rows (see `session-search-row-filter` for why - they cannot be pushed into SQL), so their cost tracks how many sessions the - walk has to read before it fills a candidate set. This corpus has 40 sessions, - which is one page of that walk; an index where few sessions match the operator - will read up to the ceiling in `session-search-retrieval` instead. - -## What the conversation scope costs at real corpus size - -`conversation` was a second FTS table holding a copy of the two prose columns. -It is a column filter now — `{user_text assistant_text}: (…)` with bm25 weights -that zero the other two — and PR 2 deleted the table on the strength of the -shoot-out this section used to hold: the filter came in at 1.16-1.36x the p95 of -the dedicated table, under the 2x bar, while the table cost a tenth of the index -to maintain. What follows is what the shipped schema actually does, measured -again on the same corpus after the table went and tool rows were capped. - -Corpus: Claude transcripts from `config/scripts/session-search-tool-heavy-corpus.ts`, -105 MB, indexed through the real store, at two points in the 80-97% band a real -transcript tree sits in. Half the tokens in tool output are words the -conversation also uses, so a conversation term really does have postings the -filter must discard. Twenty queries per rung, both scopes interleaved query by -query, warm cache; `config/scripts/session-search-scope-benchmark.ts`, run twice. - -| Tool share | Rung | `all` p50 / p95 | `conversation` p50 / p95 | -| ---------- | ------ | --------------- | ------------------------ | -| 86% | phrase | 16.69 / 17.48 | 13.08 / 13.52 | -| 86% | or | 31.91 / 35.74 | 22.25 / 23.87 | -| 86% | and | 70.04 / 74.00 | 53.47 / 59.39 | -| 93% | phrase | 9.14 / 13.36 | 7.23 / 8.51 | -| 93% | or | 16.46 / 18.70 | 12.34 / 14.88 | -| 93% | and | 39.65 / 43.44 | 31.05 / 32.92 | - -Three things to read out of it. - -**The filter is a win, not a cost.** Every rung is faster narrow than wide, by -1.2x to 1.4x at p50. The shoot-out compared the filter against a table built for -exactly this query; against the wide table it replaces, it does what the second -table did, which is read fewer postings. - -**The `and` rung is where the corpus size shows.** Those queries are eight terms, -chosen so no ordered run that long occurs and the phrase rung has to miss; a -real two-term AND sits nearer the phrase row. It is also the noisiest: the -second run's p95 reached 140 ms on one bucket, which is what twenty samples of a -70 ms query buys. Read the p50 column. - -**The index is far smaller than the shoot-out's was.** 57 MB at 93% tool output -and 103 MB at 86%, against roughly 150 MB for `messages_fts` alone before PR 2 -capped an indexed tool row at 3,072 characters. Most of a tool-heavy transcript -is now not in the index at all, which moves every number above and is the larger -effect of the two. - -What is **not** measured here is relevance, and the column filter does carry one -ranking difference the deleted table did not. FTS5's bm25 normalises by the -whole row's length and has no per-column length, so two rows with identical -prose score differently when one also holds tool output. The rowid set is -unchanged, which is what the deletion was decided on; the order within it can -move. `session-search-engine.test.ts` pins the direction. - -## `sessionCandidateLimit` - -The reviewer's F13: this is a tunable default, not a constant. It bounds how many -sessions the SQL hands ranking, so it bounds both retrieval cost and how deep a -caller can page before the answer simply stops. - -The limit only costs anything once more sessions match than the limit allows, so -this is measured over a second corpus: 2,500 one-turn transcripts, 10.9 MB, every -one of them matching the query. Limits are interleaved sample by sample, because -run back to back the first configuration pays for every page the OS cache had not -seen and the ordering alone moves p95 further than the limit does. - -| Limit | p50 | p95 | Pages of 20 a caller can reach | -| ----- | ----- | ----- | ------------------------------ | -| 200 | 6.85 | 7.21 | 10 | -| 600 | 7.93 | 8.36 | 30 | -| 1200 | 9.55 | 10.53 | 60 | -| 2400 | 12.32 | 13.45 | 120 | - -600 is the default: it costs about 16% over 200 at p50 and buys three times the -reachable depth, and the curve only turns steep past 1200. A host with a much -larger index can raise it; the result's `truncated.candidates` says when the limit -was the thing that cut the answer, so a caller never has to guess. - -What is **not** measured here is relevance. These numbers say what a limit costs, -not what it retrieves. The MRR figures quoted in the BM25 weights -(`session-search-retrieval.ts`) and in the identifier shadow column -(`session-search-identifier-split.ts`) come from the original retrieval shoot-out -on real transcripts and are not reproducible from this repository. Any change to -the limit justified on relevance grounds needs an eval set, not this benchmark. - -## What typo repair costs - -The repair is the one rung whose cost tracks the size of the vocabulary rather -than the size of a result. It only runs for a term the scope has no posting for, -so an ordinary query never pays it; a query of nonsense pays it once per term. - -Measured over a synthetic vocabulary of 1.6 M distinct terms, every term in two -rows so none is filtered out: - -| Query | p50 | -| -------------------------------------- | ------ | -| one known term (no repair) | 11 ms | -| one unknown term | 10 ms | -| 39 unknown 12-character terms (480 ch) | 387 ms | -| 12 unknown 40-character terms | 99 ms | - -Two things follow. The cost is linear in unknown terms and in vocabulary size, -and `search` is synchronous, so a 512-character query of nonsense holds the -thread for a third of a second on an index that large. And the scoped-count fix -made this cheaper rather than dearer — it was 737 ms before — because ordering -the vocabulary scan by term drops the sort that ordering by `doc` required, and -the counts it added are at most eight bounded probes per prefix. A cap on -unknown terms per query is recorded as a follow-up in the split plan. - -## Page warmup, dropped - -PR 2 deferred `warm()` — a sliced read of `messages` that pulls its pages into -the OS cache before the first query — to whoever knew which pages a read -touches. It is not re-added here, for two reasons. The measurement that -justified it (first query 1.3 s to 0.45 s) was on a 4 GB index, and neither -corpus in this file is within an order of magnitude of that, so PR 4 cannot -show a win: removing the call moved the 10.5 MB corpus's p50 by less than the -run-to-run spread. And it is a cancellable background pass, which needs an owner -with a lifecycle; a query library that holds no timers has nothing to hang the -`stopped()` on, and a fire-and-forget async read from a synchronous `search` is -a rejection nothing can supervise. It belongs with the indexer in PR 3b, which -already owns starting and stopping work. - -## Not settled here - -Which process may open, unlink and rebuild the index is PR 3b's decision. A -second handle that finds an older schema version replaces the file while a live -store keeps answering from the unlinked inode, and this PR is what first makes -that reachable, because it is the first thing that reads. What PR 4 does is -refuse to make it worse. The engine restores its derived vocabulary and generation -triggers before a search. A missing `messages_fts` fails clearly; the connection -owner must rebuild the source index. There is no degraded-search capability state -or query logging. Logging can be added by a caller when an evaluation consumer exists. - -Each search checks the generation before retrieval and after its final content -read. A concurrent commit rejects the page with `stale-generation`, including a -first page without a cursor. The caller can retry from page one. No long-lived -read transaction is needed, and a mixed page is never returned as a valid snapshot. - -Repository/path operators are applied before a phrase or AND route is accepted. -Candidate truncation is reported by the rung that answered, not by every rung -tried. Each rung of the ladder matches a superset of the one before it, so a -rung that reached its cap with no eligible sessions is always followed by one -that reaches it too: a full candidate set stays explicit either way. - -The phrase and AND rungs run for prose as well as for literal-looking input, -over the query's tokens as typed rather than the stop-word-stripped OR body. A -sentence pasted out of a transcript is ordinary words in order; over OR its -common words fill the candidate limit with recent sessions and the old session -holding the sentence never reaches ranking. The cost is two FTS queries that -usually miss, which on the corpus above sits inside this harness's run-to-run -noise. A one-token query still takes the rung only when it looked literal. diff --git a/docs/reference/agent-skill-provider-paths.md b/docs/reference/agent-skill-provider-paths.md deleted file mode 100644 index 293229ad70e..00000000000 --- a/docs/reference/agent-skill-provider-paths.md +++ /dev/null @@ -1,38 +0,0 @@ -# Agent skill provider paths - -Last verified: 2026-08-11. - -V1 supports only providers whose paths are independently established by official documentation. -The registry is deliberately small; it is not copied or synchronized from a community path table. - -| Provider | Detection | Global canonical support | Workspace support | Orca placement | -| --- | --- | --- | --- | --- | -| Codex | `codex` CLI found through Orca's host-owned PATH detection | Reads `$HOME/.agents/skills` directly | Reads `.agents/skills` from the current directory through the repository root | Canonical copy only | -| Claude Code | `claude` CLI found through Orca's host-owned PATH detection | Reads `$HOME/.claude/skills` | Reads `.claude/skills` from the launch directory through the repository root, plus nested directories as files are accessed | Relative directory symlink on POSIX, directory junction on Windows, or verified independent-copy fallback | - -Codex locations and symlink behavior are documented in the official OpenAI documentation: -[Build skills](https://learn.chatgpt.com/docs/build-skills#where-codex-loads-local-skills). - -Claude Code locations, precedence, parent traversal, live detection, and symlink behavior are -documented in the official Anthropic documentation: -[Extend Claude with skills](https://code.claude.com/docs/en/skills#where-skills-live). - -Codex therefore needs no provider-specific placement. Claude Code does not document -`.agents/skills` as a discovery root, so Orca reconciles its documented `.claude/skills` path back -to the canonical copy. Orca never replaces a path it does not own. If alias creation is unavailable, -the verified copy fallback is tracked in the install receipt so update and removal can detect drift. - -## Registry change process - -Every registry change requires normal code review and all of the following evidence: - -1. Link current official provider documentation for global and workspace paths. -2. Record whether the provider reads `.agents/skills` directly and its documented link behavior. -3. Verify global and folder-workspace discovery on macOS, Linux, native Windows, and WSL where the - provider supports those platforms. -4. Exercise local, paired-runtime, and SSH host-owned path resolution. -5. Test alias denial, broken owned aliases, independent-copy drift, update, rollback, and removal. -6. Update mixed-version capability evidence if the placement contract changes. - -Do not add automated upstream path-table synchronization. A provider release that changes discovery -semantics must enter through this review process. diff --git a/docs/reference/agent-skill-sharing-threat-model.md b/docs/reference/agent-skill-sharing-threat-model.md deleted file mode 100644 index efccfd7119b..00000000000 --- a/docs/reference/agent-skill-sharing-threat-model.md +++ /dev/null @@ -1,101 +0,0 @@ -# Agent skill sharing threat model - -Status: implementation baseline for security and privacy review. This document does not constitute -security approval. - -## Scope and trust model - -This model covers private skill packaging, Orca Cloud publication and authorization, durable share -resolution, local and remote installation, provider placement, update, rollback, removal, and -operator recovery. It applies to macOS, Linux, native Windows, WSL, paired Orca runtimes, and SSH -targets. - -Private means unlisted: an unpredictable active share ID is a bearer credential, while publication -and management remain access-controlled to authenticated Orca users and organizations. V1 is not -end-to-end encrypted from Orca Cloud operators. A skill is code from its author: `SKILL.md` can -change agent behavior and packaged scripts may be executed later by a user or agent, although the -installer itself never executes package content. - -## Protected assets - -- Skill contents, filenames, manifests, package metadata, release notes, and author identity. -- Organization membership, selected-user ACLs, durable share identifiers, and package existence. -- Signed upload policies, download grants, authentication tokens, database credentials, and GCP - service identities. -- Existing local skills, provider configuration, user modifications, provenance, and transaction - recovery state. -- Cloud package metadata, immutable object generations, audit records, and deletion state. -- Availability and cost of Cloud Run, Cloud SQL, GCS, connected runtimes, and local filesystems. - -## Actors and boundaries - -Actors include an authorized publisher, an authorized recipient, an authenticated but unauthorized -Orca user, a malicious skill author, a compromised renderer, an untrusted remote RPC caller, an -attacker controlling a network endpoint, and an operator with GCP or database access. - -Trust boundaries are: - -1. Source skill directory to the owner-private package staging directory. -2. Desktop renderer to the main process and its authenticated Cloud client. -3. Orca client to a paired runtime or SSH host across independently versioned protocols. -4. Orca Cloud API authorization to short-lived GCS access. -5. GCS quarantine to validated immutable publication and PostgreSQL metadata. -6. Extracted staging to canonical destination and provider placements. -7. Local diagnostic records to a user-reviewed support-bundle upload. -8. Terraform and deployment identities to staging and production resources. - -## Security invariants - -- Package identity is deterministic and binds normalized paths, exact bytes, executable state, and - immutable package/version IDs. -- Publication and installation accept only the `manifest.json` plus `skill/` envelope and never - execute archive content. -- Archive validation completes before destination mutation. -- Destination paths are resolved by the runtime that owns the host or workspace. -- Unowned or modified local content is never silently replaced or deleted. -- Final GCS objects are immutable, generation-fenced, and reachable only through fresh ACL checks - followed by short-lived grants. -- No remote caller can turn a desktop-local path into remote filesystem authority. -- Interrupted transactions converge to a verified old or requested version. -- Logs and support bundles exclude package contents and private authorization or filesystem data. -- Cloud sharing, downloading, and remote installation have independent kill switches. - -## Threat register - -| ID | Threat | Required controls and current evidence | Residual release gate | -| ----- | ------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- | -| TM-01 | Source changes after preview or a link swaps bytes during packaging. | Observe the specific source directory, copy without following links, re-observe staged bytes, compare every identity, and bind preview to the final digest. Source-drift and link/special-file tests cover rejection and cleanup. | Repeat race and permission tests on every supported filesystem. | -| TM-02 | Archive traversal, drive paths, links, devices, duplicate paths, Unicode/case collisions, or decompression/resource exhaustion escape staging. | Streaming parser rejects unsafe entry classes and normalized collisions; compressed, extracted, entry, file, depth, and per-file limits are enforced during parsing and extraction. Boundary, malformed, checksum, and fuzz tests run before destination mutation. | Keep package safety suites required in release CI. | -| TM-03 | A forged manifest lies about names, bytes, executable state, or package identity. | Parse the manifest before trusting entries, require `skill/SKILL.md`, hash extracted bytes independently, recompute package identity, and compare package/version IDs and digest. | Cross-platform identical-byte digest evidence remains required. | -| TM-04 | A client chooses another home, workspace, WSL distro, SSH path, or escapes a destination root. | The executing runtime resolves home and workspace identities, realpath-checks directories, uses platform-native joins, and rejects client-supplied remote paths. `local-file` ingress is trusted in-process only. | Real SSH and mixed-version topology tests remain required. | -| TM-05 | Installation overwrites or deletion removes unowned or locally modified content. | Planner distinguishes missing, unchanged, clean, modified, unowned, external-link, broken-link, and collision states. Provenance lives outside installed content. Replacement requires an explicit conflict decision; removal revalidates ownership and digest. | Complete the remaining real junction, external-link, copy-drift, and permission matrix. | -| TM-06 | A crash between renames or receipt publication loses both old and new versions. | Same-filesystem staging, durable journals, backups, flushed receipt replacement, bounded recovery, and ownership tokens protect every commit boundary. Failure injection covers every journal transition. | Real process termination and disk/antivirus contention tests remain required. | -| TM-07 | A grant leaks through redirects, userinfo, an insecure scheme, DNS/host confusion, or an oversized stream. | Cloud requests reject redirects. Package downloads require configured origins and HTTPS, reject URL credentials and cross-origin redirects, cap redirect count, recheck expiry, stream exact expected bytes, and verify archive/package digests. | Validate approved production origins and exercise malicious network cases in staging. | -| TM-08 | Reuse of an upload ID, wrong tenant, wrong key, stale generation, or quarantine object publishes attacker-selected bytes. | Random tenant-bound upload rows, signed POST conditions, expiry, exact metadata/key/type/size validation, generation-fenced reads, streamed validation, and idempotent finalization fail closed. | Staging GCS integration must cover stale generations and lifecycle deletion. | -| TM-09 | IDOR, stale organization membership, or a guessed identifier exposes package management data or bearer-protected content. | Management requests authenticate and re-evaluate current ownership. Recipient preview and grants require the exact unpredictable, active bearer share ID and ignore legacy ACL rows. Not-found responses hide unauthorized existence. | Human authorization review and organization-departure policy approval remain required. | -| TM-10 | Content-addressed deduplication discloses another tenant's package through response shape or timing. | Tenant-scoped APIs never expose GCS keys, generations, or whether an object already existed; existing-object reuse verifies both archive and logical package identity. Cross-tenant tests prove identical archives use separate tenant-hashed objects and response shapes disclose no reuse. | Keep tenant-isolation and response-disclosure coverage required in release CI. | -| TM-11 | A compromised renderer, old client, or arbitrary RPC caller sends credentials, local paths, unknown opcodes, or unsupported operations to a host. | Main owns auth tokens and grants; remote requests use strict schemas and capabilities; package transfer is separate from install; no stream opcode was added; mixed versions fail with update-required results. | Real client-newer/server-newer and SSH parity gates remain required. | -| TM-12 | Transfer replay, overlap, disconnect, or abandoned staging consumes disk or commits different bytes. | Session count, idle time, total bytes, and chunk size are bounded. Offsets are monotonic, identical retry is idempotent, changed replay fails, commit hashes the exact staged file, and cancellation/disconnect cleanup is bounded. | Exercise offline-GCS and real disconnects at every transfer boundary. | -| TM-13 | Provider aliases or junctions escape canonical storage or point at external content. | POSIX aliases are relative from real parents, Windows junctions use absolute canonical targets, targets are revalidated, and unowned/broken/external links are conflicts. Copy fallback is independently hashed and receipt-owned. Real Windows and WSL coverage includes existing, broken, external, denied, and drifted placement behavior. | Preserve the physical placement matrix in release coverage. | -| TM-14 | Skill instructions or scripts are mistaken for trusted Orca code or executed during install. | Share/install previews identify author, organization, scripts, executable files, digest, and version. Installation never runs scripts. Trust copy says the package is code from its author. | Security and design must approve the final trust wording and accessibility behavior. | -| TM-15 | Telemetry, logs, or support bundles leak instructions, filenames, paths, ACLs, grants, policies, or credentials. | Desktop install diagnostics map values to bounded categories before tracing. Deployed staging logs contain route templates and bounded request metadata only, and the logging exclusion removes bearer URLs. Support-bundle tests inject private canaries and prove they are absent from collected output. | Preserve deployed-log and support-bundle privacy checks for future changes. | -| TM-16 | Permissive staging permissions expose package bytes to another local user. | Package archives, downloads, extraction, relayed uploads, locks, journals, and receipts use owner-private modes on POSIX; Windows uses owner-profile paths and inherited ACLs. Existing POSIX download roots are tightened before use. Real Windows, WSL, Ubuntu-floor, and SSH validation passed. | Preserve owner-private staging checks across supported hosts. | -| TM-17 | Deletion, revocation, retention, or user departure leaves unauthorized grants or unrecoverable metadata/blob divergence. | Revocation blocks new grants immediately; existing grants expire within five minutes. Database references govern final deletion, quarantine lifecycle cleans abandonment, and GCS soft delete supplies recovery. Local installs remain independent. | Approve departure/legal-retention policy and exercise coordinated database/GCS recovery. | -| TM-18 | Broad IAM, public bucket access, service-account keys, or a compromised deployment identity bypasses application authorization. | Uniform bucket access, public-access prevention, bucket-scoped object access, service-account-scoped signing, skill-secret-only access, Cloud SQL client role, no desktop IAM, and no long-lived keys are Terraform-defined. | Review the staging and production plans and verify live IAM before rollout. | -| TM-19 | Unbounded validation or request concurrency causes memory, CPU, database, storage, or egress denial of service. | Fixed streaming buffers, package limits, per-instance finalization semaphore, rate/quota limits, bounded transfer sessions, Cloud Run instance limits, lifecycle cleanup, and independent kill switches constrain work. | Complete load testing, dashboards, alerts, and budget thresholds. | -| TM-20 | Operator recovery, diagnostics, or legal workflows bypass tenant isolation or leak content. | Runbooks require generation-specific recovery, coordinated PostgreSQL/GCS restoration, audited lifecycle actions, and no package contents in normal logs. | Security/privacy approval and restricted break-glass procedure remain required. | - -## Required review evidence - -Security and privacy approval must not rely on this document alone. Reviewers need: - -- Package/admission schemas and stable failure categories. -- Archive parser, extraction containment, transaction, provenance, and removal tests. -- Cloud authorization, object-generation, tenant-isolation, deletion, and recovery tests. -- Terraform plans plus live staging IAM, bucket, Cloud Run, Secret Manager, and Cloud SQL evidence. -- Real macOS, Linux-floor, Windows, WSL, paired-runtime, mixed-version, and SSH results. -- Captured staging logs, metrics, traces, and support bundles with seeded private canaries absent. -- Load-test results and independently tested upload, download, and remote-install kill switches. - -Approval owners record findings and accepted residual risks outside this implementation document. -The external rollout gate remains closed until those findings are resolved or explicitly accepted. diff --git a/docs/reference/agent-skill-sharing-upstream-boundary.md b/docs/reference/agent-skill-sharing-upstream-boundary.md deleted file mode 100644 index bf2323d2895..00000000000 --- a/docs/reference/agent-skill-sharing-upstream-boundary.md +++ /dev/null @@ -1,43 +0,0 @@ -# Agent skill sharing upstream boundary - -Status: proposed for formal engineering and legal review. - -Date: 2026-08-11. - -## Context - -Orca needs private, durable, cross-machine skill sharing with bounded archive ingestion, -host-owned destination resolution, crash-safe transactions, provenance, and mixed-version remote -support. `vercel-labs/skills` exposes a CLI and does not provide the Cloud authorization or local -transaction contract Orca requires. - -The behavioral assessment used upstream commit -`c6f69c631292444cc541ac6d91e2226b0ff247da`. - -## Decision - -Orca implements its package, Cloud, installation, provider-placement, and recovery behavior -independently. The upstream project is a behavioral reference only. - -Do not copy or mechanically translate upstream source, tests, fixtures, registry entries, provider -path tables, comments, or documentation. Derive provider paths from official provider -documentation and verify them with real installations. Orca does not depend on the upstream CLI, -npm package, or an unsupported programmatic API. - -If a future change proposes incorporating upstream material, stop and review the exact material, -license, attribution, notices, and maintenance implications before implementation or merge. - -## Consequences - -- Orca owns stability, security, compatibility, and maintenance of this narrower installer. -- There is no automatic upstream synchronization job. -- Similar behavior is acceptable when independently derived from requirements and official - provider contracts; textual or structural copying is not. -- Provider registry changes require normal code review plus official-documentation and real-host - evidence. -- The pull request template requires reviewers to confirm this boundary for relevant changes. - -## Review record - -Product chose the reference-only approach during planning. Formal engineering and legal reviewers -remain to be named before external rollout. diff --git a/docs/reference/agent-status-store.md b/docs/reference/agent-status-store.md deleted file mode 100644 index 912b14d51fe..00000000000 --- a/docs/reference/agent-status-store.md +++ /dev/null @@ -1,577 +0,0 @@ -# Agent status store - -## Status - -The current boundary is PR 2A: structured sessions use the hook server's fully -scoped canonical store; unbound PTY/relay evidence remains in an isolated legacy -adapter. Do not remove the renderer bridge or its publication filters in this -slice: they still carry native-chat child rows. - -The sections below record the original 2026-09-09 rollout. Its PR 1a and PR 1b -have landed; its proposed PR 2/3 sequence is superseded by that boundary: - -1. main-only: every producer writes into one store and `worktree ps` reads it, - split into 1a (structured sessions join the store) and 1b (the runtime's - duplicate retained store is deleted); -2. renderer: the sidebar becomes a subscriber and stops re-deriving rows; -3. shared: one worktree-status rollup and one freshness rule for every reader. - -## The problem this solves - -Orca shows "what is this agent doing" in four places: the desktop sidebar, the -`orca worktree ps` command, the mobile app, and the agent dashboard. Before -#19217 those readers did not even share their inputs. After #19217 they share -the structured-session mapping and nothing else. - -An audit on 2026-09-09 found six producers and three consumers, and three -separate copies of the same row inside the main process alone: - -| Main-process copy | Keyed by | Owned by | Persisted | Evicted | -| --------------------------------- | --------- | --------------------------------------------------------------------------------- | ------------------ | ---------------------------------------------- | -| hook server `lastStatusByPaneKey` | paneKey | `src/main/agent-hooks/server.ts` | `last-status.json` | tab close, pty exit, hydrate, worktree removal | -| runtime `RuntimeAgentRowStore` | paneKey | `runtime-agent-row-store.ts` (deleted in PR 1b) | no | pty exit only | -| structured feed `published` | sessionId | `src/main/native-chat/agent-session-wire/structured-agent-session-status-feed.ts` | no | never (a broadcast cache) | - -The second copy is a duplicate write: the OSC status parsed in main is -forwarded to the hook server _and_ retained in the runtime store from the same -call (`orca-runtime-create-terminal-side-effect-command-code-detector.ts`). -The third copy is keyed differently and never reaches the hook server at all, -which is why `worktree ps` grew its own adapter for it in #19217. - -Each reader then applies its own precedence and freshness rules, so the same -pane can legitimately read differently on the desktop, on the phone, and in -the CLI. - -## The rule - -**The execution host owns agent status, in one store, and every reader -subscribes to it.** This follows the boundary in -[`ssh-execution-boundary.md`](./ssh-execution-boundary.md): the host that runs -the process is the only party that can observe it, and the client is never -authoritative for execution state. - -Three consequences: - -- One store per execution host. A remote host keeps its own store and the - client mirrors it down, as the web-session mirror already does. Mirroring is - not merging: a client never writes its observations back to a host. -- Precedence is decided once, at write time, with provenance recorded on the - row. Readers never re-adjudicate hook versus terminal versus structured. -- Readers keep only presentation policy and user facts: the 30-minute display - decay, acknowledgements, dismissals, unread. Those stay reader-side but - become one shared implementation (PR 3). - -## The store already exists - -The hook server's state is that store today for every PTY-based agent. The -audit established: - -- hook HTTP posts, the WSL and SSH relay receivers, and main's own OSC parse - all converge on the same `applyNormalizedStatus` path, stamped with the - authority id `main-agent-hooks`; -- it alone holds pane authority: launch tokens and their hashed commitments, - retired-pane fences, pane-key aliases, per-connection ordering watermarks, - and the evidence-age map that must outlive a transport clear; -- it alone persists, with a seven-day hydrate window and the - `restoredUnconfirmed` stamp that keeps a hydrated row from ever reading as - live truth; -- it already fans out to both renderer windows over `agentStatus:set` and - `agentStatus:clear`, and serves `agentStatus:getSnapshot`. - -Nothing else in main carries those guarantees, and building a second store -with them would be the wrong direction. So the design is not "add a store". It -is: **route the two producers that bypass the hook server through it, then -delete the copies.** - -## PR 1a: structured sessions publish into the store - -No renderer behavior changes. The sidebar keeps receiving the same IPC events -it receives today, plus structured-session rows it currently derives itself. - -### Structured sessions publish into the hook server - -The structured feed keeps its job of projecting a session's journal into a -summary and streaming it to subscribers. On every publish it additionally -ingests the summary into the hook server as a status row: - -| Row field | From | -| --------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `paneKey` | `structuredAgentSessionPaneKey(tabId, sessionId)`, the key the renderer already uses; its leaf is UUID-shaped so pane-key validation accepts it | -| `tabId` | `structuredAgentSessionTabId(sessionId)` | -| `worktreeId` | `summary.workspaceId` (a folder workspace id is a valid value) | -| `state` | `structuredAgentSessionAgentStatus(summary).state`: the lead's own status folded with the live child records the store holds for the session (not the summary's task list), so a settled lead whose subagent still runs reads `working` | -| `workingMode` | `'monitoring'` from the same fold when watch loops are the only live child work; omitted otherwise, which clears it on the row | -| `mainAgent` | the main agent's own state before the fold, its last-turn verdict (`summary.turnOutcome`, present only while idle) and its own clock; see "The main agent fact" below | -| `structuredHost` | `'owned'` while `summary.hostExecutionOwned` is set, otherwise `'held'`; `worktree ps` derives its row's `structuredHostOwned` from it | -| prompt, tool, last message, model, provider session | the summary's fields | - -Sessions with no request (`status === null`) produce no row. A request is a -turn record, an assistant message, a user message the provider journaled itself -(history, an older host), an accepted or unanswered send, or a send the agent or -its start refused; a send that was withdrawn, or left undelivered by a -restart or a close, fails nobody and makes nothing listable. -`summary.turnOutcome` is the latest request's verdict: its turn's outcome, or -`failure` for a send the agent or its start refused (a send that joined a running -turn is answered by that turn). A turn the provider gave no outcome reads as its -host-observed end through `agentTurnVerdict`: `interruption` for an `interrupted` -lifecycle, `unconfirmed` for an `unverifiable` one. It is derived on each read, -never journaled. The row also publishes `interrupted` from -`mainAgent.outcome`, exactly as the hook lanes do. When the host revokes live ownership the row is re-set -without the flag; when the host closes or evicts the session the row is -dropped. Both already exist as feed events (`revokeLive` and the roster -filter in `liveSessionSummaries`); PR 1 turns them into store writes. - -Dropping the session from the host's map and dropping its row are one -operation, `forgetStructuredAgentSession`. The store keeps a row until told, -and a host-owned row bypasses the staleness check, so a deletion path that -forgot the row would strand a permanently working-looking agent. - -Two rules the ingest must keep: - -- **Never persist a structured row.** The journal is the durable truth for a - structured session and the host republishes on restore. A structured row in - `last-status.json` would hydrate as `restoredUnconfirmed` and then fight the - live republish. The serializer skips rows carrying `structuredHost`, and - hydrate drops any such row found on disk. Applying one therefore also skips - the persist schedule: the walk and stringify could only reproduce the file - that is already on disk, once per debounce window for every streaming chat. -- **Never let it fight a hook row.** A structured session has no PTY, so no - hook or OSC event carries its pane key. The ingest still goes through the - disposition gate so a retired pane key is refused like any other. - -Applying one does still run both status fan-outs, and that is intended rather -than incidental. `notifyStatusChangeListeners` is what feeds -`agentAwakeService`'s power-save blocker, and `subscribeEnrichedStatus` is what -feeds `AgentSessionTransitionRecorder`'s stats, so joining the store enrolls -native chats in both. A working native chat is real work and should hold the -machine awake exactly like a PTY agent does. - -The drop side routes through `dropStatusEntry`, not `clearPaneState`: a -pane-status-clear reaches the renderer, and until PR 2 the renderer's own feed -bridge is that pane key's writer. It also passes `preserveResumeIdentity: -false` — the `providerSessionOnly` remnant a dismissed pane keeps exists so the -agent can be resumed in that pane, and a structured session has no pane and -keeps its resume identity in the record store. Like every other -`dropStatusEntry` caller, it emits no pane clear, so a session dropped -mid-`working` leaves `AgentSessionTransitionRecorder` holding an open stats -session until its LRU evicts it; that gap is shared with the user-dismissal -path and is not specific to structured rows. - -The ingest lives in the feed, not in `structured-agent-session-host.ts`, which -sits at the file-length cap. - -### `worktree ps` becomes a reader - -The structured adapter added in #19217 is deleted, and structured rows reach -`worktree ps` through the same snapshot as every other row. The -retained-versus-hook reconciliation in `collectRuntimeWorktreePtyAgentSources` -stayed until PR 1b removed the store that fed it. What this step settles is -the admission gate that decides which rows a worktree listing may show: - -- a hook or OSC row needs its tab mirrored or a connected pty, as today, and - SSH rows stay exempt because their tabs may exist only remotely; -- a row carrying `structuredHost` is admitted while the host holds the session, and - the host's drop on close is what removes it. No tab-mirror requirement: a - structured session's tab lives in the renderer's own tab state, and a - headless host has no renderer to mirror it from. That argument only holds if - the headless host is itself wired to the store, which is a separate - obligation per entry point: the Electron hosts (desktop and `orca serve`) - share `main-process-runtime-service.ts`, and `orcad` constructs its own - runtime in `src/main/orcad/orcad-entry.ts`. A host missing that wiring lists - no agents at all, not just no structured ones, because `worktree ps` reads - the same snapshot for every row. - -The freshness bypass for host-owned structured rows already exists in -`isFreshNonDoneAgentStatus`; with the flag now on the row it becomes the only -path, and the hand-rolled check in `runtime-worktree-agent-rows.ts` goes. - -### Wire compatibility - -`AgentStatusIpcPayload` gains one optional field, `structuredHost`, and the -`worktree ps` row gains `structuredHostOwned`. Under rule 1 of -[`remote-wire-compatibility.md`](./remote-wire-compatibility.md) both are safe: -an old client ignores them. `worktree ps` rows keep their shape and vocabulary, -so the mobile app sees no change. - -Until PR 2 the main process does not forward structured rows to the renderer -over `agentStatus:set` or `agentStatus:getSnapshot`. The renderer's feed -bridge still writes those rows itself, and forwarding them too would give one -pane key two writers. Removing that filter is the first step of PR 2. - -### The main agent fact - -Claude, Codex and Grok hook rows and structured-session rows publish the combined -`state` and, beside it, the main agent's own state as `payload.mainAgent`. Other agents' -rows and terminal-title-only rows carry none, and readers fall back to `state`: - -```ts -mainAgent?: { state: AgentStatusState; outcome?: AgentTurnOutcome; stateStartedAt: number } -``` - -`state` still answers "what should the user see" and folds live child work in, -so a settled main agent whose subagent still runs reads `working`. `mainAgent` answers -"what is the main agent itself doing", which the fold used to destroy at publish -time; every guard that reconstructed a fragment of it (`fromChildWork`, the -persisted `claudeLeadBoundaryChildOnly` flag) now reads `mainAgent` instead of a -stored copy. A Claude row whose `mainAgent` is `done` while a child agent still -works (including a child's permission wait) refuses OSC, which carries no child -identity; the children's own lifecycle hooks settle it. `outcome` is the recorded verdict on -the main agent's most recent finished turn, present only while `mainAgent.state` is -`done`. It is reported by the provider, or is a `cancellation` Orca inferred -from the user's own interrupt keystroke, or a `superseded` the host recorded when -a newer request replaced a structured Claude turn before it ended (it names no -sender, and sets no legacy flag), or, on a structured row whose turn the -provider gave no verdict, is what the host observed of its end: `interruption` -(a proven death nobody asked for) or `unconfirmed` (an end it cannot prove, -never success). The journal's turn outcome, by contrast, stores only recorded verdicts: the provider's, a `cancellation`, or the host's `superseded`; `interruption` and `unconfirmed` are derived from the turn's lifecycle state and never stored. A plain end of turn carries none, because absent -means unknown and a provider that omits its interrupt flag must not turn a -cancel into a success. -In the Claude hook lane the cancellation comes primarily from Orca's own -inferred interrupt (`markClaudeLeadTurnInterrupted`), because current Claude -sends no hook at all on a cancel and no `is_interrupt` on Stop; that flag on a -turn boundary remains a secondary source for builds that send it, and -`StopFailure` maps to `failure`. - -Readers decode the verdict through one accessor, `agentMainAgentVerdict`, which -reads the main agent's own state, not the combined row's: `mainAgent.outcome` -while `mainAgent.state` is `done`, then the legacy `interrupted` flag as a -cancellation, which alone needs the combined `done`. So a main agent that -failed while its subagents still run has a verdict on a `working` row. Every -copy of a row (state-history entries, sleep records, `worktree ps` rows) takes -the verdict through `agentVerdictFields`, which carries `interrupted` and the -whole `mainAgent` (state, outcome and its own clock) together, so a copy agrees -with the row and can date a failure by `mainAgent.stateStartedAt`. - -Display reads the verdict through `agentVerdictDisplayMark`. A fault marks the -agent failed whatever the combined state, because it is news the user must see -even while subagents run: a `failure`, and an `interruption`, a turn cut short -by anything other than the user or a newer request. A user's stop (`cancellation`) -marks it interrupted, drawn in the muted tone with the row text "Interrupted by user"; -a turn a newer request replaced (`superseded`) marks it interrupted in the same muted -tone with the row text "Interrupted"; and `unconfirmed` marks it unconfirmed, all only -on a `done` row, so a stopped -or finished main agent with live child work still reads working. The folded -turn header follows the same mark: "Failed after N", "Interrupted after N", or -"Worked for N". -Each subagent keeps its own row and state. Container rollups (worktree card, -terminal tab, Cmd+J) rank a pending question first, then a failure, then live -work, then an unconfirmed end, then a user's stop, then done. On the worktree -card, a failure retained after its agent's pane went away has no expiry, so it -ranks below live work and above an unconfirmed end. Lifecycle waiters keep -reading the combined `state`. - -Policy splits the verdict two ways. Clean-finish policy (hibernation, pane -ownership, the star-nag value moment) treats a failure, an interruption and an -unconfirmed end like a cancellation (`agentTurnEndedUncleanly`). Attention -(completion time, Smart Sort, sticky retention, Cmd+J Recent) demotes only a -turn ended on purpose, the user's stop or a newer request that replaced it -(`agentTurnEndedOnPurpose`); a failure, an interruption or -an unconfirmed end ranks like a completion. - -Admission is one function, `normalizeAgentStatusPayload`, on the relay wire, -IPC and disk. A malformed `mainAgent` drops the field and keeps the row. Old hosts -send none and readers fall back to `state`. Hook rows persist it inside the -payload; hydration maps an older row's `claudeLeadBoundaryChildOnly: true` -onto `mainAgent: { state: 'done' }` when the row has no `mainAgent`, and never writes the -flag again. Hydration seeds the Claude main agent record straight from a saved -`mainAgent` that is `done`, so the children's drain can still settle the row after -a restart. `claudeRunningNonAgentTask` is persisted alongside because it is the one -child-work fact `mainAgent` cannot express: a shell running beside the main agent, -whose liveness hydration does not restore. Hydration seeds only a row that says -`false`; a row silent about it stays unseeded. The row builder pairs the two facts in -one place: a listener event restates the shell fact, and any other write (an OSC -repaint, an inferred answer) keeps it only while `mainAgent` is unchanged. A child's -sticky permission prompt still records the main agent's own progress and background -evidence in the held row, and pushes the held row to subscribers when `mainAgent` changes. - -Every lane, Codex included, combines through the fold. A child waiting on a -human is a fold input (`childWorkLiveness: 'waiting'`, derived from the child's -own `waiting` state; a child's `blocked` means it failed and stays live work) -and makes the row wait whatever the main agent is doing, unless the main agent -is itself asking. The Codex hook lane feeds it from its child transcripts, and -the structured lanes from child records, which read `waiting` for a Codex child -thread's approval or input flag and for a Claude subagent's open permission -request. Known -divergences, pinned by name in the parity table -(`src/shared/main-agent-status-parity.test.ts`) where they are reachable, so a -reader does not mistake them for drift: - -- The Claude hook lane holds a child's permission wait in one slot on the - displaced main agent record (`waitingAgentId`, `stateBeforeWait`), not on - the child. It publishes the displaced state as `mainAgent`, but the next - main agent event overwrites the slot, so the row stops reading `waiting` - while the child is still asking, and a second asking child replaces the - first. -- In the structured lane a child's pending prompt also makes the session - `attention`, which reads as the main agent's own `blocked`: one needs-input - state whoever asked. A Claude subagent reads `waiting` only while the - journal holds its card pending: from after the card's row is written until - just before anyone closes it, so every publish that shows the child waiting - also shows the session's `attention`, and the row never reads `waiting` for - a Claude subagent's request. -- The Codex hook lane drops its roster on a root `Stop` when it tracks no - child transcripts, so a still-running or still-asking child stops holding - the row. - -How the main agent's turn ended is not a fold input. A cancel is a verdict on -the main agent, carried as `mainAgent.outcome: 'cancellation'` (and, for -readers that predate `mainAgent`, as the row's `interrupted` flag on a `done` -row); it never retires a shell, scheduled check or subagent the turn left -running. That work leaves the row only when its own inventory omits it or the -session ends, so a cancelled turn with a still-running shell reads -`monitoring` in every lane, and the parity table in -`src/shared/main-agent-status-parity.test.ts` drives that story through all of -them. The same rule governs the cancel Orca infers from Ctrl+C: for any row -that publishes `mainAgent`, the inference is admitted only when -`mainAgent.state` is `working`, so Orca does not treat a Ctrl+C at the idle -prompt of a row held open by child work as a turn cancel (Codex also keeps the -child-evidence guard, and a row without `mainAgent` keeps only that guard). -The keypress itself is not inert, though: measured live, Claude 2.1.280 stops -its background subagents on a single idle-prompt Ctrl+C (shells survive) and -Codex 0.156.1 quits outright, so refusing the inference can leave the row -showing a subagent its CLI already stopped. The synthesized row is the fold -of the cancelled main agent with the child work the pane's owner can see: the -local listener's roster for a local pane, the row's own subagents and shell fact -for a relayed one, whose provider records live on the relay. - -The store holds that verdict against restatements that predate it -(`server-cancel-verdict-latch.ts`), because a relay never learns of a cancel -the desktop infers and some TUIs emit late same-turn hooks. The hold is read -off the row (`mainAgent.outcome: 'cancellation'`), never stored beside it, and -dies on a new turn (a main agent prompt submission, a changed or explicit -prompt, a session start) or the provider's own settled `mainAgent`. Child and -replayed events under the hold keep the cancelled main agent and are re-folded -with their own child evidence. - -## PR 1b: the runtime's retained row store is deleted - -Landed. `RuntimeAgentRowStore` is gone, and with it the retained-versus-hook -reconciliation in `collectRuntimeWorktreePtyAgentSources`. The hook server's -store is now the only main-process copy of a PTY agent's row. - -### The five call sites - -| Call site | Before | After | -| ------------------------------------------------------------------------------ | --------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- | -| `orca-runtime-create-terminal-side-effect-command-code-detector.ts` `retain()` | second write of the OSC payload already sent to the hook server | deleted; the event now carries the pane's `terminalHandle` and the hook ingest keeps the only copy | -| `...command-code-detector.ts` `clearPty()` | drops rows on pty exit | deleted; pane teardown already clears the hook row | -| `orca-runtime-get-worktree-ps.ts` `values()` | fed `retainedSnapshots` | deleted; the reader keeps only `hookSnapshots` | -| `orca-runtime-serialize-agent-prompt-submission.ts` `getFreshExplicit()` | retained row first, hook rows second | `selectFreshExplicitAgentStatus`, hook rows only | -| `orca-runtime-prune-mobile-session-tab-group-layout.ts` `getFreshForMobile()` | pane key, then pty id | `selectFreshAgentRowForMobileTab`: pane key, then `terminalHandle` | - -Both readers moved into `runtime-hook-agent-row-selection.ts`, which also owns -`RuntimeAgentRowSnapshot` now that nothing retains one. - -### `terminalHandle` is the row's join back to its terminal - -The retained store's only real extra was the pty id, and two readers used it. -The plan said to stamp the event's `ptyId` into `terminalHandle`; that was -wrong. A terminal handle (`term_`) and a pty id are different -identifiers, and `getFreshExplicit` was already comparing hook rows against a -real handle. What landed instead: - -- `AgentHookEventPayload` and the runtime's terminal-status event gained an - optional `terminalHandle`. The detector resolves it once per chunk through - `getAgentStatusTerminalHandleForPaneKey` — the same lookup the renderer-facing - IPC boundary already runs for every row, so the two surfaces cannot disagree - about which terminal a pane is. -- `applyNormalizedStatus` carries the handle forward when an incoming event - resolves none. Only main's OSC parse can resolve one, so an HTTP hook post for - the same pane would otherwise erase it. -- It is never persisted. A handle belongs to the runtime that issued it, and a - hydrated one could only rejoin a row to somebody else's terminal. -- `toAgentStatusIpcPayload` publishes it, which also makes `getFreshExplicit`'s - long-dead handle comparison live: the runtime reads raw snapshot rows, and - before this nothing ever stamped the field on them. - -`worktree ps` uses it too. `ConnectedPtyEvidence` traded its flat `ptyIds` set -for `ptyIdByTerminalHandle`, so a row still resolves the connected PTY behind -it — which is both the working-terminal rollup's match key and the last rescue -for a row whose pane binding was nulled by a controller incarnation change. - -### The change detector had to move with the store - -`retain()` was not only a store: its boolean return was the signal that -republished `session.tabs` for a status-only transition, which no title change -covers (#7970). `hook-status-session-tabs-invalidation.ts` already mirrors that -projection change set, including restore provenance and terminal-handle joins, -so the replacement was to route the signal off the store rather than build a -second comparator. -`installHookStatusSessionTabsRepublish` now owns all three arms — enriched -status, pane clear, and the status-drop tap a dismissal emits — and both hosts -install it. - -### Both hosts, not just the desktop one - -`orcad` constructed its runtime with no `onTerminalAgentStatus`, so main's OSC -parse never reached the store there and the retained copy was the only carrier. -Deleting it without wiring orcad would have made a headless host list no PTY -agents at all. `orcad-entry.ts` now binds the producer and installs the -republish signal, alongside the snapshot and structured sink it already had. - -### The intended behavior change - -A row the user dismisses on the desktop leaves `worktree ps` and the phone at -once, instead of lingering until the pty exits. One store means one dismissal. - -Legacy numeric pane keys remain a bounded compatibility case. Persisted layouts -register aliases to their stable leaf owners; an in-process OSC observation may -also retain a numeric key only when the runtime supplies the matching tab, PTY, -and terminal handle. HTTP and relay ingress still require a stable key or a -registered alias, and numeric rows are never persisted. - -## PR 2: the renderer subscribes - -With structured rows arriving over `agentStatus:set`, the renderer's -`StructuredAgentSessionStatusBridge` no longer needs to write status; its -unmount cleanup becomes a tab-close signal to the host. The IPC applicator is -the single writer for observed status. The 2026-09-09 audit sorted the other -writers: - -| Writer | Disposition | -| ----------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- | -| Command Code output seeds, parked-pane seeds, pty-exit removal | delete; main already emits the same facts | -| structured bridge status writes | delete; main now publishes the row | -| structured bridge failed-start row (the host refused the create) | keep; a refused create leaves the host no session, so the bridge writes it from the launch record | -| launch placeholder seeds (a user launched an agent with a prompt) | keep for now; main holds the launch config and can seed later | -| dismissal, acknowledgement, unmount | keep; user facts and component lifecycle | -| remote-runtime OSC parse (bytes never transit local main) | keep, fenced behind the host's published row once the host is new enough; rule 3 of the wire doc applies | -| web-session mirror receipt clock | keep; the decay rule needs both clocks from one machine | - -The Command Code done-settle window is renderer policy with no main -equivalent. PR 2 either moves it into main's detector or leaves it, and says -which. - -## PR 3: one rollup, one clock - -The worktree card status is derived three times: `lib/worktree-status.ts` in -the renderer, `runtime-worktree-status-projection.ts` in main, and -`agent-row-display.ts` in mobile, which hand-copies the 30-minute constant. -PR 3 moves the rollup and the decay into `src/shared` and makes all three -call it. - -## Readiness reads the store - -`terminal wait --for tui-idle` is a reader too. Before STA-9100 hook state -reached it only through the ` ready` titles the window writes, so a -headless `orca serve` never saw it (#16095). Now an agent whose rule file says -`profile.hooks: "authoritative"` (OpenCode, OpenCode 2, Pi, OMP) or -`"turn-end"` (Codex) has its fresh row read straight from the store, through the same -`selectFreshExplicitAgentStatusRow` join prompt-receipt verification uses -(`src/main/runtime/tui-idle-hook-lane.ts`): - -- the main agent's turn, not the combined row, decides: `mainAgent.state` when - published, so a subagent's Stop does not end the lead turn. `done` settles - the wait, `working` holds it, and a permission wait never settles. The tail's - blocked text goes through the existing permission arbiter with the turn as - its explicit status, so a denied prompt's dialog left in the tail no longer - blocks a turn the hook says ended; -- the row joins on any pane key or terminal handle the PTY owns; a pane neither - reaches, a stale or restored row, a session-start `done`, a row from before - the PTY respawned, and a `done` received before the pane's latest input all - leave the decision to the screen and text rules, which is also how startup - readiness works before an agent's first hook. The input is the PTY run's - `lastInputAt` (`terminal-run-facts.ts`), which both write funnels record, so - a key the user typed counts like a prompt Orca sent: the next turn's first - hook may still be in flight, and an agent restarted in the same shell has - not posted one. A shell command marker is no process boundary: Pi paints - OSC 133 zones itself; -- every other agent stays `identity-only`: Claude sends no event when an - approval is denied or Esc stops a tool, so its row can sit at `waiting` or - `working` forever, and the rules keep deciding; -- Codex is `turn-end`: only a `done` decides (it settles the wait), and a - `working` or permission row leaves the decision to the rules. Before its - `Interrupt` hook an Esc mid-turn can leave the row `working`, and an older - TUI can hand its hooks to a newer shared app server, so no version check - tells which Codex posts it. A `done` is a real turn end on every version, so - trusting only that one keeps the headless gain (no quiet window after the - turn) without letting a missing cancel hang the wait. - -The titles stay for display; remote clients read them. - -## What does not change - -- The hook scripts, the OSC 9999 wire format, and the relay protocol. -- The status vocabulary. `working / blocked / done` for rows, - `working / attention / idle` for structured summaries, mapped once. -- The `live / unverifiable / exited` verdicts for remote work. Loss of contact - clears nothing; the SSH exemptions in the admission gate stay. -- Hydration honesty: a restored non-done row is `restoredUnconfirmed` and is - never fresh. - -## PR 1b reliability contract - -- **Invariant (`agent-session.status-host-ownership`):** each execution host has - one agent-status store; OSC, hooks, and structured sessions write it, while - desktop, `worktree ps`, and mobile only project it. Dismissal, certified PTY - exit, and provider-generation replacement remove the same row everywhere; - transport loss alone removes nothing. -- **Failure source:** the deleted runtime row store duplicated OSC observations, - keyed them by a different terminal identity, and outlived a dismissal from the - hook store. Relay replay could also make old evidence look fresh when readers - used its new delivery timestamp. -- **Oracle:** one OSC observation appears through the hook snapshot in - `worktree ps` and mobile, and one store dismissal removes it from both without - stopping the PTY. Focused tests also require leaf/incarnation-handle rejoin, - legacy numeric-pane compatibility, certified-exit and provider-generation - cleanup, evidence-age freshness, and exactly-once startup/stop teardown. -- **Gate:** `terminal-performance.osc-status-scan-budget` covers the unchanged - bounded OSC parser and the runtime projection. There is not yet a dedicated - blocking multi-surface status-store gate; the focused suites below are the - accepted gap until they accumulate reliability-gate soak evidence. -- **Provider/platform coverage:** local and daemon-backed PTYs are covered by - runtime tests, and SSH relay loss/replay semantics by relay integration tests. - The projection is shared by git worktrees and folder workspaces. WSL uses the - same store and admission code but has no live run here; Linux and Windows - runtime execution, native mobile clients, and mixed-version paired clients - remain validation gaps. -- **Performance budget:** publication stays event-driven with no new polling or - subprocesses. One mobile projection clones the status snapshot once, builds - pane/handle indexes once, and has a deterministic call-count test; lifecycle - cleanup is bounded by the existing status and handle inventories, and orcad - tests prove listeners clean up once on failed startup and repeated stop. -- **Diagnostics:** existing hook-listener errors name the pane and PTY, while - status-store tests pin delivery versus evidence clocks. No new telemetry or - raw terminal data is emitted. -- **Residual gaps:** rendered Electron/mobile behavior, live SSH reconnect, and - Linux/Windows/WSL execution require the platform QA pass. The current - cross-version gate does not cover `session.tabs` content. - -## Verification - -- Unit: ingest a structured summary and read it back through - `getStatusSnapshot`, `worktree ps`, and the mobile projection; assert the - serializer never writes a row carrying `structuredHost`; assert a hydrated - file that somehow contains one is dropped. -- Unit: the `worktree ps` suites written against the retained store are rewired - to a real `AgentHookServer` (`agent-status-store-wiring.test-fixture.ts`) - rather than deleted, so each still asserts the listing behavior it named. The - dismissal change is pinned end to end in - `orca-runtime-tests/worktree-ps-agent-row-dismissal.spec.ts`, which fails with - the retained store restored. -- Live: the parity check from #19217 (working, done, close, reload) repeated - against the merged store, with both surfaces read from the one row. - -## Retired OMP pane recovery - -A desktop renderer retirement carries an optional UUID through the existing -`agentStatus:retirePaneAuthority` IPC message. The hook server retains it with -its bounded retirement fence. A validated live OMP new turn consumes that UUID -and echoes `authorityRestartId` only in the live notification. Cached rows, -persistence and startup replay never carry the acknowledgement. Older peers -omit or ignore it and retain explicit attach restoration. - -The renderer keeps the UUID in its existing non-persisted retirement tombstone; -every re-retirement mints a new one. A matching acknowledgement may clear that -tombstone only with a successful status write for the existing pane and matching -workspace/connection. Closed tombstones remain `true`, including after the tab -LRU evicts its entry. Closing a retired physical alias revokes its whole group. -This is control-plane retirement correlation, not a second agent-status store. - -Fallback restores the hook server's recorded status aliases through the existing -attach-restoration path. The accepted renderer write restores the matching status -alias routes too, preserving group membership for the next retirement. It does -not restore orchestration or launch credentials. -It is scoped to the requesting desktop renderer. A different window's retirement -UUID cannot be cleared by the acknowledgement, and web mirrors keep their existing -host-snapshot/attach behavior. diff --git a/docs/reference/antigravity-native-accounts.md b/docs/reference/antigravity-native-accounts.md deleted file mode 100644 index 29fe943afde..00000000000 --- a/docs/reference/antigravity-native-accounts.md +++ /dev/null @@ -1,90 +0,0 @@ -# Native Antigravity Accounts - -Accounts reads the credential authority on the runtime that owns execution. A client chooses -an owning Orca runtime and a host/distro target before sending an operation; it never replaces -the client's Mac Keychain item for another host. The RPC capability is -`accounts.antigravity-native.v1`. Older paired hosts are refused before account mutations. -The RPC returns account summaries only, never credential JSON, access tokens or refresh tokens. -Displayed quota is tied to the subject and authentication method observed during its refresh; -an external identity change hides the previous account's quota without an automatic fetch. - -## Supported authority - -Normal macOS agy uses service `gemini`, account `antigravity`. Its go-keyring values use the -base64 or legacy hex wrapper. Orca passes writes through `security -i` stdin, validates bounded -output and reads the entire native value back. The command buffer limit is checked before -writing. A missing native item falls back to the CLI-specific -`~/.gemini/antigravity-cli/antigravity-oauth-token` file. The distinct legacy jetski fallback -is not imported. - -The compiled CLI bypasses keyring storage when SSH/WSL environment detectors or WSL kernel -identity apply. A runtime running under that evidenced bypass reads/writes its own CLI file; -it does not contact the client keychain. The file must be private and regular. A macOS -`cache/antigravity-keyring-unavailable` marker makes authority uncertain: Orca refuses instead -of assuming that the keychain or file wins. - -Native Windows Credential Manager, native Linux Secret Service, and operations directed from -Windows Orca to a selected WSL distro are explicitly unsupported pending verified adapters. -Windows file bypass is also refused until private ACL protection is verified. -Windows' `gemini:antigravity` raw blob and 2560-byte limit are different from the Mac wrapper; -Linux uses the login collection with `service=gemini`, `username=antigravity`. No dependency, -PowerShell compilation, credential-home flag, or cross-host fallback is invented here. -A separate SSH relay has no Accounts RPC; use a paired owning runtime that implements it. - -## Identity and snapshots - -A Google ID token supplies the normalized Google issuer and stable subject. The authentication -method also scopes identity. The label uses a verified email when available; email is never the -identity key. Account record IDs are random and survive token, expiry, refresh-token and email -rotation. Profiles without a stable subject can be displayed but cannot be saved for switching. - -Snapshots preserve the exact native JSON, including fields that Orca does not interpret. The -host's vault under `userData/antigravity-accounts/vault` requires meaningful OS encryption and -private permissions. Weak or unavailable encryption is refused. Unreadable/corrupt ciphertext -is preserved; it is never treated as an empty vault. This does not migrate the experimental -candidate's incompatible array vault or token-hash IDs. - -One host service serializes Add, Select, Remove, launch checks and refresh reconciliation. -It re-reads the vault after asynchronous native reads and captures external CLI refreshes into -the same stable account. Selection reconciles the outgoing snapshot, checks the expected native -bytes before writing, and checks native readback before publishing the selected ID. It avoids -writing an old snapshot over an already-active account. The current or selected account cannot -be removed; deletion checks the latest native value again before committing. - -A selected account is checked before new Orca PTY launches, including desktop daemon and -headless runtime paths. An externally changed native identity blocks the launch and asks the -user to select again. Existing sessions can retain their original credentials in memory. -Shell commands typed manually into a running terminal are outside the Orca launch guard. - -## Sign-in and concurrency limits - -Sign-in uses the supported ordinary agy browser/code flow. Users run agy on the owning host; -to add a different account they use its `/logout` command, complete the next sign-in, then save -the actual resulting account in Orca. This implementation does not advertise an Orca-managed -login or invent an agy `login`/`--login` flag. Browser completion and a second real Google -account remain user-driven; tests do not sign out or change the developer's real native item. - -Native keyring does not expose compare-and-swap. Orca's queue serializes its own calls, and -bounded before/after checks detect observed conflicts; another independently running agy or -Orca process can still write between the final check and the write or launch. A failed -verification may mean the native item changed but selection was not persisted. Refresh and -explicit selection resolve that state; automatic rollback could destroy a newer CLI refresh -and is deliberately avoided. The file backend has the same external-writer limit. - -## Evidence and contributor credit - -The foundation adapts the reviewed codec/macOS adapter from #21784 and account-service concepts -from #21797 (nwparker), with fresh identity, persistence, serialization and conflict handling. -The signed-in Accounts card and quota-error visibility acknowledge #19588 by @artile; quota -transport is reused from current main rather than its obsolete extraction code. Targeted -multi-account UI/target concepts acknowledge #23761 by @Tai-DT, replacing its placeholder login -and unused settings selection. The Accounts legacy-Gemini clarification acknowledges #21682 -and the original relevant migration contribution by @siddqamar, as requested in #17345. -No stale development stack was cherry-picked. - -Live proof uses a disposable Mac service/account item, a fully isolated hidden Electron home, -and synthetic accounts. A private task-only copy was also selected through the real service; -installed agy 1.2.14 consumed that verified file credential under its SSH bypass and returned -`command.name=usage`, `num_turns=0`, no conversation. The real native item remained unchanged. -This proves the Mac adapter mechanics and actual CLI file authority, not a second-account -native-keychain switch, native Windows/Linux switching, or WSL/SSH relay deployment. diff --git a/docs/reference/antigravity-readiness-evidence.md b/docs/reference/antigravity-readiness-evidence.md deleted file mode 100644 index 18317d7e5ce..00000000000 --- a/docs/reference/antigravity-readiness-evidence.md +++ /dev/null @@ -1,332 +0,0 @@ -# Antigravity readiness: what the transcripts show - -Antigravity readiness lives in `src/main/runtime/agent-state-rules/antigravity.json`: a screen rule -over the trusted grid, and a text anchor that runs the named scan -`findAntigravityComposerIndex` (`agent-state-rules/antigravity-text-composer.ts`) over the -line-folded tail when no trusted grid exists. That text scan decides whether a pane is ready for a -prompt from its tail alone. It has been written five times, each version tuned -against a five-line screen typed from memory into a `.spec.ts` fixture. Three of the first four -were found worse than the bug they replaced, and the fifth was reverted. - -Real transcripts now exist. They were recorded from a live `agy` on macOS with -[`agent-pty-transcript-capture.md`](./agent-pty-transcript-capture.md) and are committed under -`src/main/runtime/__fixtures__/`. `src/main/runtime/antigravity-readiness-transcripts.test.ts` -replays them through the runtime. - -**Headline (STA-8741, agy 1.2.14): readiness is now read off the live screen, not the text tail.** -The screen's bottom four rows at an idle composer are rule, caret, rule, `? for shortcuts`. The -caret alone is not enough: agy keeps it painted mid-turn and behind the `/model` picker. The hint -row is what changes. [Attempt seven](#attempt-seven-the-live-screen-sta-8741) has the recordings; -the sections after it are the 1.2.0 history that led there. - -## Attempt seven: the live screen (STA-8741) - -Recorded 2026-09-30 on macOS with `agy` 1.2.14 (binary and banner agree), a signed-in Google -account, `AGY_CLI_HIDE_ACCOUNT_INFO=1`, in an already-trusted workspace. All are 120x40 except -where the name says otherwise. `src/main/runtime/antigravity-screen-readiness-transcripts.test.ts` -replays them. - -| Fixture (`antigravity-1-2-14-*.txt`) | Screen at the end | Screen rule | -| ------------------------------------ | --------------------------------------------------- | ----------- | -| `ready` | settled startup, bare `>` | ready | -| `ready-accept-edits` | `> Accept-edits mode: …` (`--mode accept-edits`) | ready | -| `ready-plan` | `> Plan mode: …` (`--mode plan`) | ready | -| `ready-80x24` | settled startup on an 80x24 PTY | ready | -| `picker-dismissed` | `/model` opened, then Esc | ready | -| `turn-ended` | a short turn has ended | ready | -| `model-picker` | `Switch Model` open; the bare `>` is still above it | not ready | -| `command-palette` | `> /` with the palette, hint `esc to cancel` | not ready | -| `busy-thinking` | `Generating...` spinner, bare `>`, `esc to cancel` | not ready | -| `busy-streaming` | answer streaming, bare `>`, `esc to cancel` | not ready | -| `trust-dialog` | untrusted folder, alternate-screen trust menu | not ready | -| `draft` | unsent text in the composer; the hint row is blank | not ready | - -What they show: - -- **The text tail misses three ready screens.** On `ready-accept-edits`, `ready-plan` and - `turn-ended` the line-folded tail never satisfies `findAntigravityComposerIndex`; the screen - rule does. (`turn-ended` is the 1.2.14 form of the old `busy-turn-ended` known defect.) -- **A caret rule is wrong on the screen.** The grid keeps the bare `>` through a turn and behind - the picker. The text tail happened to lose it mid-turn (section 8); the screen does not. -- **The rule can read ready for a moment mid-turn.** Replayed in 64-byte chunks, the submit repaint - clears the composer a moment before `? for shortcuts` becomes `esc to cancel`. So a pane with an - output clock is held to the same 3s quiescence as Codex (tier 1b in `tui-idle-evidence.ts`); an - idle agy goes silent within a second and a turn keeps repainting its spinner. A restored pane with no - clock settles from the screen alone. -- **A name-only title settled the picker.** With the pane titled `agy`, the old ranking reached - its weak title lane and settled with `/model` open. When a trustworthy screen exists it now - decides, and that lane stays shut. -- **A mismatched grid garbles the chrome.** At 80x24, 100x30 or 60x20 the 120x40 recordings lose - the four-row shape, and resizing the model does not make the TUI repaint. So the live screen - counts only when its grid matches the PTY's reported size and is not a reflow the TUI never - repainted for (a re-attach that learned the real size late; a later PTY resize off that size - repaints it). Otherwise every pre-existing lane (text rules, title, quiet process) decides, as - before this change. -- **A readable screen outranks the text.** When the grid is trustworthy and refuses, the text rules - do not overrule it, in the ready-prompt tier or the quiet one; they decide only when there is no - trustworthy grid. The recorded picker-dismissed text over a model-picker screen proves it. - -The visible-read probe's Antigravity branch is retired. It read the provider screen with a looser -rule (any caret after the banner) whenever the pane was Antigravity. The probe now runs the shared -rule for every screen-ruled agent, only for a pane with no output clock. A clocked pane settles -through the poll. The probe's screen read is the draft-blanking read projection, so it restores the -blanked composer row before the rule reads it (Cline's `❯ Ask anything...` otherwise reads as a -bare `❯`). - -Still not captured: a tool-permission prompt (the operator's `toolPermission` is -`always-proceed`, and changing it means editing their settings), and the sign-in, theme, privacy and -update dialogs, for the reasons in the table below. Windows and Linux are unrecorded. - -The 1.2.0 fixtures below were recorded before the recorder stopped writing ahead of shutdown. Their -last bytes are agy's exit teardown (`ESC[J ESC[?2004l`), so their final screens have no hint row. - -## Versions - -| Thing | Value | -| ------------------------- | ----------------------------- | -| `agy --version` | `1.1.25` | -| Banner printed by the TUI | `Antigravity CLI 1.2.0` | -| Captured | 2026-09-10, macOS, 120x40 PTY | - -The binary and its own banner disagree. Any rule keyed to a version string must read the banner, -not `--version`, and must tolerate the two disagreeing. - -## What the captures are - -| Fixture | What it is | -| -------------------------------------------- | --------------------------------------------------------- | -| `antigravity-ready-api-key-gemini-model.txt` | Ready screen, API-key identity, Gemini 3.7 Flash (Low) | -| `antigravity-ready-account-info-hidden.txt` | The same ready screen with `AGY_CLI_HIDE_ACCOUNT_INFO=1` | -| `antigravity-dialog-trust-workspace.txt` | Workspace trust dialog, live and unanswered | -| `antigravity-dialog-model-picker.txt` | `/model` picker, live and unanswered | -| `antigravity-dialog-command-palette.txt` | Slash-command palette, live and unanswered | -| `antigravity-dialog-dismissed.txt` | `/model` picker dismissed with esc, then settled | -| `antigravity-busy-mid-turn.txt` | A real turn, recording stopped while the spinner was live | -| `antigravity-busy-turn-ended.txt` | The same turn after it ended and the composer returned | - -## What could not be captured, and why - -Nothing below was faked. Each is a case the recorder could not reach without changing the -operator's account state or configuration, which is out of bounds. - -| Missing | Why | -| ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| `antigravity-ready-business-non-gemini.txt` | This machine has no OAuth session — the CLI prints _"You are currently not signed in"_ and authenticates from `GEMINI_API_KEY`. Reaching a Business ready screen means signing someone in. | -| A non-Gemini model on any ready screen | `agy models` offers 11 models, all Gemini, and `settings.json` pins `modelProvider: gemini`. A non-Gemini row is not reachable from this account. | -| `antigravity-dialog-sign-in.txt` | Unsetting `GEMINI_API_KEY` does not reach the sign-in dialog; the CLI refuses to start because `modelProvider` is pinned. Reaching it means editing the operator's `settings.json`. | -| `antigravity-dialog-theme-picker.txt` | There is no `/theme` command in 1.2.0 (`Unknown command: /theme`). The picker appears only in first-run onboarding, which means deleting the operator's config. | -| `antigravity-dialog-privacy-notice.txt` | First-run onboarding, as above. | -| `antigravity-dialog-update-banner.txt` | Cannot be forced; no update was pending during the session. | - -Each remains as a named, skipping case in the suite so it is visible rather than forgotten. - -## What the transcripts show - -### 1. The ready screen's model row is not at the start of a line - -The ready screen prints a block-glyph logo down the left, and the identity, model and path rows are -painted **on the same physical lines as the logo**. What Orca derives is: - -``` -▀▀▀▀▀▀ Gemini API key -▀▀▀▀▀▀▀▀ Gemini 3.7 Flash (Low) -▄▀▀ ▀▀▄ ~ -``` - -The earlier detector required `normalized.startsWith('gemini', trimmedStart)` on a trimmed line. The -trimmed line starts with `▀`, so that rule never matched. The shipped detector uses the bare -composer caret instead. Measured three ways on the real screen: - -| Input | `isKnownReadyPromptPreview` | -| ------------------------------------------------------ | --------------------------- | -| Real ready screen | `true` | -| The same screen with the logo glyphs stripped | `true` | -| Real ready screen followed by the live `/model` picker | `false` | - -So the logo — decoration, and suppressible with `AGY_CLI_HIDE_LOGO` — no longer decides readiness, -and the live dialog cannot reuse the stale composer caret as a ready signal. - -### 2. The dialog used to satisfy the model rule - -`/model` prints its options one per line: - -``` -Gemini 3.8 Flash -> Gemini 3.7 Flash (current) -Gemini 3.1 Pro -``` - -Those lines _do_ begin with `Gemini`, and a bare `>` composer line sits earlier in the same tail -from before the picker opened. Both halves of the old rule were satisfied **while a dialog owned the -screen**, and the pane read ready. The shipped detector now recognizes the active `Switch Model` -surface and rejects that stale caret until it sees `Exited /model command`. - -### 3. `>` is the dialog selection marker, not only the composer caret - -Every dialog uses `>` to mark the highlighted row: `> Yes, I trust this folder`, -`> Gemini 3.7 Flash (current)`, `> /add-dir`. The idle composer is a line whose whole trimmed -content is `>`. That distinction is the only thing separating them, which means the relaxation -proposed in PRs #15840 and #15852 — accept any line _beginning_ with `>` — would make the trust -dialog and the model picker read as ready. On 1.2.0 the idle composer is a bare `>`; those PRs' -1.1.17 mode-banner claim could not be reproduced here and may be mode-specific. - -### 4. There is no email account row, and the row can be switched off entirely - -For an API-key user the identity row reads literally `Gemini API key`. There is no `@`, no -domain, nothing an account-row rule can key on. Separately, `AGY_CLI_HIDE_ACCOUNT_INFO=1` — a -supported environment variable in the binary — removes the row from a fully ready screen, which -`antigravity-ready-account-info-hidden.txt` captures. - -### 5. Dialogs are drawn two different ways, and the banner is never reprinted - -The trust dialog and the sign-in splash take the **alternate screen** (`ESC[?1049h` … `ESC[?1049l`). -The model picker and command palette are drawn **in place on the main screen** with erase-to-EOL. -After dismissal the CLI prints `⎿ Exited /model command` and redraws the composer — it does **not** -reprint the banner. The header stays where it was at startup. - -### 6. Rows are positioned with cursor addressing, not newlines - -The status row is written with absolute and relative moves (`ESC[13;99H`, `ESC[83X ESC[83C`), so -`? for shortcuts` and `Gemini 3.7 Flash · low` end up on one derived line. Any rule that assumes -one screen row equals one `\n`-delimited line is reading a different document than the user sees. - -## 8. Busy frames park the caret exactly like idle frames — the spinner is what differs - -The frame that ends a turn-in-progress and the frame that ends an idle screen park the cursor with -the **same bytes**. Only the hint row differs, and the park erases it: - -``` -idle: ? for shortcuts ESC[83X ESC[83C Gemini 3.7 Flash · low CR ESC[2A ESC[2C ESC[?25h -busy: esc to cancel ESC[85X ESC[85C Gemini 3.7 Flash · low CR ESC[2A ESC[2C ESC[?25h -``` - -So a rule that keys on "the caret is the last thing in the tail" cannot tell busy from idle **on the -frame alone**. What saves it is what comes next. Each spinner tick is its own repaint with its own -park, two rows higher than the frame's: - -``` -ESC[?25l CR ESC[2A ⣯ Generating ESC[11D ESC[?25h -ESC[?25l CR ESC[2A ⣟ Generating. ESC[12D ESC[?25h -``` - -That second `CR ESC[2A` splices the composer row away, so the retained tail during a live turn ends -on the spinner row, not on the caret. Measured on `antigravity-busy-mid-turn.txt`: - -| Capture | last retained line | bare `>` line present | -| -------------------------------------------- | ------------------ | --------------------- | -| `antigravity-ready-api-key-gemini-model.txt` | `>` | **yes** | -| `antigravity-busy-mid-turn.txt` | `⣟ Generating...` | **no** | - -**Consequence for a caret-based rule:** it already answers "not ready" for a real mid-turn capture, -because there is no bare caret in the tail to match. A constructed input that keeps the park bytes -and only edits the status text is not faithful to a live turn — a live turn has a spinner row -repainting _below_ the composer. - -**The residual window, and the clause it implies.** Between a frame park and the next spinner tick -the tail does end on the bare caret and is indistinguishable from idle. The gap is one tick -interval. Any readiness path gated on sustained quiescence is safe, because ticks keep arriving and -the pane is never quiet; a path that only inspects retained text is not. For those paths the -evidence supports one clause, and only one: - -> **A braille glyph (U+2800–U+28FF) on the last visible line of the retained tail means working.** - -That predicate already exists in this file for cursor-agent (`CURSOR_BUSY_SPINNER_RE`) and should be -reused rather than reinvented. It must be scoped to the **last visible line**, not the whole tail: -a first-run transcript prints `⠾ Signing in...` during startup, which would otherwise pin a ready -screen as busy forever. - -Nothing else in the capture distinguishes the two states. The hint row (`esc to cancel` versus -`? for shortcuts`) is erased by the park in both cases, the park offsets are identical, and -`ESC[?25l`/`ESC[?25h` fencing appears around every repaint, idle or busy. - -## Confirmed / refuted, by attempt - -Evidence column names the fixture; all quoted text is from the committed transcripts. - -### Attempt 1 — the rule at HEAD - -| # | Claim | Verdict | Evidence | -| ---- | -------------------------------------------------------- | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | -| 1.1 | A ready screen prints the banner `Antigravity CLI` | **Confirmed** | `Antigravity CLI 1.2.0` in both ready fixtures | -| 1.1b | …and its last occurrence in the tail is the live one | **Refuted** | The trust dialog's own body says _"Antigravity CLI requires permission to read, edit, and execute files here"_, so `lastIndexOf` lands inside the dialog | -| 1.2 | The model row begins with the vendor word `Gemini` | **Refuted** | `▀▀▀▀▀▀▀▀ Gemini 3.7 Flash (Low)` — the logo precedes it; never at line start | -| 1.3 | The caret line's whole trimmed content is `>` | **Confirmed** on 1.2.0 idle | bare `>` in both ready fixtures | -| 1.3b | …and only the composer prints `>` | **Refuted** | `> Yes, I trust this folder`, `> Gemini 3.7 Flash (current)`, `> /add-dir` | -| 1.4 | A ready screen prints the workspace path on its own line | **Refuted** | the path shares its line with logo glyphs (`▄▀▀ ▀▀▄ ~`) | - -### Attempt 2 (loop 1) — blacklist the model line - -| # | Claim | Verdict | Evidence | -| --- | ------------------------------------------ | ----------- | ---------------------------------------------------------------------------------------------------------------- | -| 2.1 | Dialog model-row wording is enumerable | **Refuted** | the palette lists 50+ commands with free-form descriptions; the picker prints whatever models the account offers | -| 2.2 | A dialog never reproduces a real model row | **Refuted** | the `/model` picker prints four real model rows, one per line, at line start | - -### Attempt 3 (loop 2) — structural ordering on `headerIndex` - -| # | Claim | Verdict | Evidence | -| --- | -------------------------------------------------- | ---------------------------------- | ---------------------------------------------------------------------------------------------------- | -| 3.1 | A live dialog is printed below the ready chrome | **Confirmed** for in-place dialogs | picker and palette append below the composer | -| 3.2 | The banner is reprinted when a dialog is dismissed | **Refuted** | `antigravity-dialog-dismissed.txt` shows `⎿ Exited /model command` and a redrawn composer, no banner | -| 3.3 | Antigravity does not use the alternate screen | **Refuted** | `ESC[?1049h` opens the trust dialog and the sign-in splash | -| 3.4 | No full repaint per keystroke | **Partly refuted** | typing `/mod` repaints the palette region on each keystroke with `ESC[K` | - -Because of 3.2, `headerIndex` cannot be the anchor: it never advances. Ordering can only be -expressed against the model/caret positions, which is what 1.2 and 1.3b just invalidated. - -### Attempt 4 (loop 3) — require a positive account row - -| # | Claim | Verdict | Evidence | -| --- | ---------------------------------------------------- | ---------------------- | --------------------------------------------------------------------------------------------------------------------------- | -| 4.1 | Every ready screen prints an account row | **Refuted, twice** | API-key identity prints `Gemini API key` (no `@`); `AGY_CLI_HIDE_ACCOUNT_INFO=1` removes the row entirely | -| 4.2 | A startup dialog never contains an `@`-and-`.` token | **Not reachable here** | none of the captured dialogs contains one, but the palette shows free-form skill descriptions, which are user-authored text | -| 4.3 | The account row is distinguishable from prose | **Refuted** | the row is not a distinct line; it shares one with the logo | - -### Attempt 5 (PR #19749, reverted) — ordering + account row - -| # | Claim | Verdict | Evidence | -| --- | -------------------------------------------------------- | ----------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| 5.1 | Ordering plus an account row separates ready from dialog | **Refuted** | the account row is optional (4.1) and the ordering anchor never moves (3.2) | -| 5.2 | Executing both builds was sufficient verification | **Refuted** | the executed input was the hand-written fixture, so the check reproduced the fixture's assumptions. The real screen disagrees with that fixture on the model row, the path row and the account row | -| 5.3 | The wedge is a model-name problem | **Refuted** | it is a line-start problem. Even `Gemini 3.7 Flash (Low)` — a Gemini model — fails, because a logo glyph precedes it | - -### Cross-cutting - -| # | Question | Answer | -| --- | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- | -| X1 | Does `agy` set an OSC title distinguishing busy from idle? | **No.** Not one OSC title sequence appears in any capture. Title-based readiness is unavailable for this agent | -| X2 | Does it repaint with bare `\r`? | **Yes**, constantly, plus `ESC[K` and absolute cursor moves | -| X3 | Does the caret survive in the tail? | **Yes** — a bare `>` line is present in every ready capture | -| X4 | Banner-to-caret distance | ~8 derived lines on a 120x40 PTY; the banner falls outside the 6-line preview window, so only the full retained tail can see it | -| X5 | Pane title on the trust screen versus ready | Identical: none | - -## Attempt six is shipped - -Yes — but not as a variation on any of the five. Every one of them refined a predicate over -`\n`-delimited lines, and that is the layer where the evidence says the information is not. - -What the captures support and the shipped detector now does: - -- **The one stable, dialog-free ready marker is a line whose entire trimmed content is `>`.** It is - present in every ready capture and absent from every dialog capture, because a dialog's `>` always - carries its selected row's label. This is a much narrower rule than any attempt used, and it is - the only one that survived contact with the transcripts. -- **Drop the model-row requirement.** The model rows match dialogs and not the ready screen, so - keeping that requirement inverted the detector. -- **Do not require an account row.** It is optional by environment variable and carries no email for - API-key users. -- **Veto an active model picker.** `Switch Model` followed by a labeled selection row means the - bare caret belongs to the composer behind the picker; readiness resumes after `Exited /model -command`. -- **Do not anchor on `headerIndex`.** The banner is printed once and never reprinted. -- **The blocked-signal path already works** for the trust dialog: `antigravity-dialog-trust-workspace.txt` - is correctly refused today, by wording, not by structure. - -What is still unknown and should be captured: the sign-in, theme, privacy and update dialogs, and -any ready screen where the composer is not idle (accept-edits and plan mode, which PRs #15840 and -#15852 describe from a screenshot). A bare-`>` rule is only as good as the claim that those modes -still end on a bare `>`; that claim is untested. - -The honest summary is that this is a screen-shaped problem being solved with line-shaped tools. A -rule over the derived tail can be made much better than what ships today, but the durable fix is to -ask the terminal emulator what the bottom row of the screen actually is, rather than inferring it -from a byte stream that was written with cursor addressing. Attempt seven does that. diff --git a/docs/reference/antivirus-prerelease-clearance.md b/docs/reference/antivirus-prerelease-clearance.md deleted file mode 100644 index 9efcfb3aaa0..00000000000 --- a/docs/reference/antivirus-prerelease-clearance.md +++ /dev/null @@ -1,117 +0,0 @@ -# Antivirus clearance for future releases - -Orca collects a steady stream of antivirus and EDR false positives — see the -tracking issue for the current grouping. This document covers the part of that -problem worth engineering effort: **stopping the next release from being -flagged.** - -Clearing a *historic* release is explicitly not a goal. A user sitting on a -flagged build should update to a cleared one, not wait for a vendor to whitelist -a version we no longer ship. Retroactive submissions cost the same effort per -vendor and expire the moment we cut a new version. - -For the behavioural side of the problem — the process-tree shapes EDR scores, -which no whitelist fixes — read -[`windows-edr-posture.md`](./windows-edr-posture.md) instead. This document is -about file verdicts on the bytes we ship. - -## The two mechanisms, and only one of them scales - -**Sample submission** clears one build. You send the flagged file to a vendor's -analyst portal, they confirm it is clean, and the verdict is dropped from their -next definition update. This is reactive and per-release: cutting a new version -produces new bytes, new hashes, and a fresh chance of the same heuristic firing. -Doing this every release, across every vendor, is not sustainable. - -**Signer and product whitelisting** clears every future build. The vendor -records the publisher identity or enrolls the product in a dynamic allowlist, and -subsequent releases inherit that trust without another submission. Enrollment is -one-time work per vendor, and it is the only lever that scales with our release -cadence. - -Prefer enrollment. Use submission only to clear a live incident while enrollment -is pending, and only for a release users are actually expected to install. - -## Prerequisites that make enrollment possible - -None of these programs will accept an unsigned or anonymous binary, so these -come first: - -1. **Every shipped PE is Authenticode-signed, and CI fails the release if not.** - Done as of v1.4.217 — the release workflow requires a valid SignPath - Foundation signature on the inner binaries and no longer fails open. -2. **Every shipped PE carries real provenance** — company, product, version, - description, and an explicit `asInvoker` manifest. An anonymous binary scores - worse than an identified one, and several portals reject submissions that - carry no version metadata. -3. **One stable signer identity.** Vendor allowlists key on the certificate - subject. Rotating signers resets accrued reputation, so a certificate change - is a re-enrollment event, not a transparent swap. - -## Vendor programs - -Enrollment state is deliberately left as a task here rather than asserted — fill -each in as it is confirmed, and record the account that owns it so a lapsed -enrollment is traceable. - -| Vendor | Mechanism | Scope | State | -| ------------------------- | ---------------------------------------------------------------------------- | ------------------------ | ----- | -| **VirusTotal** | Monitor — paid; builds rescanned daily, developer and vendor both notified | ~70 engines at once | TODO | -| **Microsoft** | Defender Security Intelligence submission, as a software developer | Defender, Defender FP EP | TODO | -| **Microsoft** | Trusted Signing, or an EV certificate, for SmartScreen and Smart App Control | Reputation gates | TODO | -| **Kaspersky** | Whitelist Program — vendors submit builds for the Dynamic Allowlist | Endpoint, all platforms | TODO | -| **Trend Micro** | Certified Safe Software Service — pre-release software whitelisting | Endpoint, Virus Buster | TODO | -| **Bitdefender** | False-positive submission for software vendors | Endpoint, ATD | TODO | -| **Avast / AVG / Norton** | Gen Digital false-positive and whitelisting channels | Consumer suites | TODO | -| **ESET** | False-positive sample submission | Endpoint | TODO | -| **Tencent iOA** | No public developer channel found; needs a support relationship | iOA, macOS and Windows | TODO | - -VirusTotal Monitor is the highest-leverage single entry, because it is the only -channel built for exactly this workflow: uploads sit in a private store, get -rescanned daily against every engine's current signatures, and when one flags a -file **both we and that vendor are notified automatically**. Pre-publish upload -is a supported use, which is precisely the future-release posture we want. It is -a paid service, monetised on developers and free to the antivirus vendors. - -Be honest about its limit: VirusTotal states plainly that Monitor is not a free -pass to get a file whitelisted. Vendors sometimes keep a detection. What it -reliably buys is *early notice and a real contact path* instead of discovering a -verdict from a user's issue report weeks later. - -Do not treat a plain VirusTotal *scan* as equivalent. A scan tells us a verdict -exists; Monitor is what routes it to someone who can drop it. - -If the subscription is not worth it, the free fallback is the community-maintained -false-positive contact directory (`yaronelh/False-Positive-Center` on GitHub), -which collects the submission addresses and forms each vendor actually reads. -That replaces the hardest part of a submission — finding the right contact — but -keeps the per-release effort that Monitor removes. - -## Where this lands in the release flow - -The check belongs at RC time, not after a stable cut — a verdict discovered after -publication is a verdict users already hit. - -`config/scripts/scan-release-artifacts-antivirus.mjs` reports the current -detection state of built artifacts by hash. Run it against an RC's artifacts, and -treat any engine verdict as a release-blocking question rather than an automatic -stop: these are third-party ML classifiers, so a hard gate on their output would -fail the release for reasons outside our control. Read the report, decide, and -submit if the flagged build is one we intend to ship. - -The script looks up hashes by default and never transmits artifact bytes. Passing -`--upload` sends the file to VirusTotal, which distributes samples to partner -vendors — that is the intended outcome for clearance work, but it is a -publication, so it stays opt-in and out of any automated path. - -## What not to do - -- **Do not ask users to add exclusions** as the resolution. It suppresses the - symptom on one machine, and in several reports here the exclusion did not even - hold because the detection was behavioural rather than path-based. -- **Do not dispute a verdict without a sample.** Every report worth acting on in - this project came with a hash that we verified bit-identical to the published - release asset. That verification is what makes a submission credible. -- **Do not chase a vendor whose detection we cannot reproduce or name.** Route - those back to the reporter for the detection string and the exact flagged path - first. diff --git a/docs/reference/ci-demand-rollout.md b/docs/reference/ci-demand-rollout.md deleted file mode 100644 index f0eaa861e84..00000000000 --- a/docs/reference/ci-demand-rollout.md +++ /dev/null @@ -1,220 +0,0 @@ -# CI demand rollout - -This implements the September 28 runner-demand analysis. The baseline inventory -covered September 27 04:00–September 28 04:00 UTC: 4,028 workflow runs, with -463 stratified job samples. Estimated occupancy was 1,081 runner-hours, dominated -by PR unit shards (414 hours) and Bun qualification (262 hours, since replaced by the -pinned-Node headless lanes). These are -sampled sums of job durations across different runner pools, not billing totals -or a guaranteed forecast of savings. - -## What runs now - -| Work | Ordinary draft update | Ready PR / final checks | Main reference | -| -------------------------------------- | ------------------------------------------ | -------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | -| Static analysis and types | Immediately | Immediately | Existing workflows | -| Unit suite | Full, with shadow selection evidence | Full | Existing daily Node 24/26 x86 suite | -| Packages | After successful static analysis and types | Same | Existing release workflows | -| Headless Node persistence | Deferred until ready | Linux x64 for ordinary runtime changes; explicit platform families or all six for sensitive inputs | All six platforms for relevant main pushes; full nightly qualification at 11:30 UTC | -| Headless Node glibc/musl qualification | Deferred until ready | Linux-specific or full qualification, after persistence succeeds | Both architectures for relevant main pushes and nightly qualification | -| E2E | Existing targeted routing | Existing targeted routing | One complete run at 17:00 UTC | - -Headless draft updates carry no verdict; readiness starts the checks. Relevant -ready PRs retain Linux x64 smoke coverage. Explicit Windows or macOS paths add both -architectures in that family, while Linux paths add Linux ARM and the glibc/musl -lanes. Root/toolchain inputs, native build inputs, shared execution/storage paths, -SSH, providers and relay changes retain every platform. Missing or incomplete -change evidence and failed import analysis also retain full qualification. -Unrelated changes skip through the dependency classifier. Relevant main pushes -qualify all six platforms and both Linux compatibility architectures; detection -uses the whole push before qualification can supersede an older relevant run. - -The detector checks known build inputs with the pinned Node toolchain first. -Other paths install dependencies for import analysis; an uncertain result never -skips qualification. This changes setup cost, not the qualification policy. - -Windows server prebuilds are cached separately by architecture and pinned build -inputs, including the runner image. Only successful qualification on main publishes -them. Consumers validate the payload and still run the pinned-Node load/spawn smoke; -a cache miss or invalid payload builds fresh. Nightly, manual and release builds -remain fresh. No native artifact is shared across platforms or ABIs. - -Expensive PR jobs wait for static/type success. This reduces fan-out for failed -or rapidly superseded commits without sleeping on a runner. Successful isolated -PRs pay the extra stage latency. Existing per-PR cancellation remains in place. -Package assertions, native boundaries, SSH/folder coverage, cache warming and -slow-test assertions are retained. - -The daemon running-work test imports the shared probe directly, with the daemon's -process inspector supplied as its callback. The renderer keeps its existing -adapter and forwarding tests. This removes a mocked renderer dependency from the -headless graph without changing the probe algorithm or skipping backend tests. -At validation, the graph fell from 6,018 inputs (1,070 renderer inputs) to 4,879 -inputs (no renderer inputs), including nine added shared-probe cases. Renderer -adapter changes no longer qualify the headless matrix; shared probe and daemon -test changes still do. Future actual renderer imports remain discoverable. - -## Unit selection rollout - -PR planning runs alongside typechecking after their shared dependency setup; an -explicit join publishes its artifact before the unit matrix can start. Static -analysis's Node 24 install also prepares the native cache before matrix fan-out. -The daily compatibility workflow retains its separate planner and cache primer. - -`ci-unit-plan.mjs` discovers the same include/exclude set as Vitest and follows -static imports, re-exports, literal dynamic imports, CommonJS requires and the -renderer aliases. Consumers of indirect filesystem/process inputs remain in the -candidate set, as do script/tool tests. Global configuration changes, deletions, -renames involving removed paths, unknown inputs and graph failures run the full -suite. A shard verifies the plan's source SHA and complete discovery list before -using it. Missing/stale artifacts fall back to full coverage, even if that means -running the suite on fewer shards. - -The initial policy was **shadow**, with every shard retained and five concurrency -slots for a full run. The account is charged a slot per job rather than per core; -in the September baseline, eight 6.5-minute shards made this matrix 68% of daily -slot demand while the ARM pool queued 10.5 minutes at p95. The first local -inventory matched Vitest exactly (9,950 files at validation); representative source -changes retained roughly 88% of files because of indirect input readers. -That is evidence for conservative coverage, not evidence of the analysis's -hypothetical 50% unit-work reduction. Improvements to indirect dependency -modeling should be demonstrated against full results before expanding selection. - -### Controlled full-PR sharding - -Full PR unit runs now request ten four-worker ARM shards through the existing -planner. Daily Node 24/26 reference runs retain five shards per Node version, and -selected draft runs retain their five-shard cap. Coverage, isolation, setup files, -worker flags and the protected actual-Node runtime projects remain unchanged. - -Set the repository Actions variable `ORCA_UNIT_FULL_SHARD_COUNT` to `5` to roll -back full PR runs. Remove the override or set it to `10` to restore ten. Only `5` -and `10` are valid; invalid values fail planning. If change/graph evidence is -unavailable, planning still retains every discovered file at the requested full -count. Confirm the actual matrix and complete selection evidence on a new run -when changing the variable; do not infer the effective count from the setting. - -The ordinary [five-shard run](https://github.com/stablyai/orca/actions/runs/37657068640) -and [ten-shard run](https://github.com/stablyai/orca/actions/runs/37657071738) used -identical definition/source trees on pinned main `7d6d7ca6`. Both first attempts -passed all required jobs and ran 11,678 files exactly once, with identical module -counts, states and actual runtime routes: 112,615 passed cases, one expected -failure and 1,052 stock skips. - -| Observed boundary or resource | Five shards | Ten shards | Change | -| ------------------------------------------------- | ----------: | ---------: | -----------------------: | -| Longest unit test step | 571s | 322s | 43.61% less time | -| Unit prerequisite release to last unit completion | 618s | 362s | 41.42% less time | -| Unit release to required verification | 628s | 371s | 40.92% less time | -| PR creation to required verification | 819s | 545s | 33.46% less time; 1.503× | -| Aggregate unit test-step time | 2,649s | 2,852s | 7.66% more | -| Aggregate held ARM runner time | 2,827s | 3,206s | 13.41% more | -| Aggregate setup before tests | 157s | 317s | 101.91% more | -| Aggregate dependency installation | 57s | 112s | 96.49% more | - -Nominal peak unit worker slots increased from twenty to forty. The comparison was -one observational pair on different hosts and cache states; unit runners started -6–8 seconds after allocation. Four additional diagnostic ARM jobs started after -all fifteen comparison runners had been allocated. This is not a quiet-fleet, -representative queue-tail or historical twofold-speedup result. Stock artifacts -prove module counts/states/routes, not complete individual case identities. - -This is a controlled latency rollout with a resource tradeoff. The representative -week-long capacity evaluation below remains pending; one successful pair does not -satisfy that fleet gate or erase the earlier oversharding concern. Monitor queue -and provisioning delays, PR-to-verification latency and aggregate ARM runner time -in equivalent traffic windows, including cancellations and planning/reference -costs. Use the five-shard rollback if queue delay erases the latency gain. -Temporary benchmark PRs #26271 and #26272 never merge. - -Every shard uploads `unit-selection.json`, `unit-timings.json` (including module -outcomes), and its assignment. The evidence job combines these into -`unit-selection-review-attempt-N/selection-review.json`, reporting: - -- Whether every discovered file appeared once across a complete reference run. -- Failures outside the candidate set, including failures in otherwise red runs. -- Measured worker time that selection would omit; worker times overlap and are - not runner occupancy or a prediction of wall-clock savings. - -Missing/duplicate shards, stale plans, interrupted runs and unhandled errors do -not count as complete references. Diagnostic upload/report failures do not make -tests pass and do not independently fail successful tests. - -After representative complete shadow runs show no missed failures, set repository -variable `ORCA_UNIT_SELECTION_MODE=selected` to enable selection **only for draft -PRs**. Keep full ready-PR checks and the daily compatibility suite. Inspect at -least a week's evidence across renderer, main, shared, SSH and fixture changes -before promotion, including red runs rather than only successful examples. -Unknown variable values retain shadow mode. Unset the variable or set it to -`shadow` to roll back immediately. Selected runs use one to five timing-balanced -shards based on retained work. - -Ready-for-review result reuse includes `unit full` in its source/workflow -identity. A green selected draft cannot satisfy the final full check, even if -the repository variable changes between runs. An already successful _full_ -identical-source check can still be reused. - -To inspect downloaded shard artifacts locally: - -```sh -node config/scripts/ci-unit-selection-review.mjs ARTIFACT_DIRECTORY -``` - -## Review automation - -Pullfrog recognizes the existing `Review #N [id]` and -`Review new commits on #N [id]` dispatch names. Explicit dispatchers may provide -`pull_request_number` and `head_sha`. Explicit PR identities share concurrency at the workflow boundary. Legacy review -names use a bounded lookup of the latest 100 dispatches and cancel only lower -run IDs for the same PR; a delayed older scope cannot cancel a newer review. -Unrecognized tasks are never grouped. The scope job alone has Actions write -permission for ordered cancellation. Closed PRs and explicitly stale heads -are skipped. A second head check prevents starting an agent after its queued -head has changed. Unrecognized agent tasks remain independent; lookup failures -also retain an independent task rather than cancelling unrelated work. - -This does not introduce a fixed debounce interval or remove final reviews. -Dispatchers should supply `head_sha` for reliable stale-at-dispatch detection; -legacy names identify a PR but do not prove which head the prompt describes. - -## E2E signal - -The daily reference still executes all shards and keeps original verdicts. Each -shard uploads Playwright JSON and publishes expected, skipped, unexpected, flaky -and startup-error counts with the failing test names/messages. Targeted PR and -manual coverage remain available. This change does not fix the historically -red tests or pretend they pass. - -`config/e2e-failure-tracking.json` can separate an evidenced repeated failure -from new failures in the summary. Each entry must have exact `file`, full -`title`, `project`, a nonempty stable `message` substring, an `@owner`, a linked -repository `issue`, and an ISO `expires` review date. Expired/malformed entries -are ignored and reported; changed error signatures appear as untracked. Entries -never skip a test or change its exit status. The initial list is empty because -the analysis established red workflows but did not establish owners and -reproductions for individual failures. Do not blanket-baseline an entire red run. - -## Capacity measurements and acceptance - -`CI runner demand` runs daily at 04:23 UTC and can be dispatched manually. It -reads the previous 24 complete hours in hourly pages, samples up to six runs per -workflow/outcome stratum, and fetches job pages with bounded concurrency. An -hour exceeding the API's 1,000-result search cap fails visibly. The report and -raw evidence are retained for 30 days. No extra runner pool is provisioned. - -The report measures the full job durations of runs **created** in the window, -not occupancy clipped to the window: earlier runs that overlap it are excluded, -and completed sampled jobs may finish after it. This matches the baseline -cohort method. Workflow IDs keep ref-qualified paths in one sampling stratum. - -The report shows weighted runner-hours and cancelled-run hours per workflow, -runner-minutes per completed PR _run_, and weighted queue/provisioning p95 per -runner label. It counts latest attempts only, excludes incomplete jobs, and -retains zero-job observations. It does not measure other repositories competing -for organization capacity. Compare equivalent traffic windows, not raw totals -alone. The collector needs only `contents: read` and `actions: read`. - -After a week, compare runner-minutes per PR run, cancellation occupancy and -queue p95 in each affected pool. Count newly added planning/reference overhead. -A 25–35% overall reduction remains an experiment target, not an achieved result; -selection, coalescing and matrix reductions overlap and cannot simply be added. diff --git a/docs/reference/ci-runner-efficiency.md b/docs/reference/ci-runner-efficiency.md deleted file mode 100644 index aa8146649bb..00000000000 --- a/docs/reference/ci-runner-efficiency.md +++ /dev/null @@ -1,2629 +0,0 @@ -# CI efficiency and runner capacity - -The [September 28 demand rollout](ci-demand-rollout.md) documents staged checks, -unit-selection evidence, headless runtime qualification, review cancellation and daily occupancy reports. - -## Headless server follow-up - -[PR #24527](https://github.com/stablyai/orca/pull/24527) adds dependency detection to -main pushes. Unrelated pushes skip qualification; relevant pushes still run all -six persistence targets and five Linux compatibility jobs. Explicit Windows or -Mac PR paths select both architectures plus a Linux smoke, while shared -execution/storage changes, SSH/provider/relay inputs, native inputs, manifests, and incomplete evidence -retain the full matrix. Main pushes use the same validated exact Windows slot -cache as PRs; nightly and release builds still compile freshly. - -Main detection runs cannot cancel each other. Only eligible qualification jobs -share main concurrency groups, so an unrelated push cannot cancel needed tests. -Release templates, explicit refs, and nightly runs remain isolated. - -Draft PRs have no server verdict, so their detector is also skipped. The existing -`ready_for_review` event performs detection and qualification once the PR is ready. -This removes the checkout and dependency installation for a result whose platform -jobs were already ineligible. - -The glibc 2.28 prerequisite step checks all five tools before installing anything. -The pinned ARM image already supplies them, including Git 2.55.0 built under -`/usr/local/bin`; installing the Git RPM does not change the Git on PATH. A -missing-tool fallback still installs the original package list and disables EPEL -for that one command. In -the baseline x64 log, EPEL metadata took 4 minutes 50 seconds to download although -every installed package came from AlmaLinux BaseOS or AppStream. The package list, -compiler image, libc floor, native smoke, and persistence tests stay unchanged. -The [DNF command reference](https://dnf.readthedocs.io/en/stable/command_ref.html) -defines `--disablerepo` as a temporary command-level filter, so later commands -retain the image's repository configuration. - -The [PR qualification run](https://github.com/stablyai/orca/actions/runs/36972824276/job/110734327685) -passed the x64 floor native smoke and persistence suite. Its prerequisite step -finished within GitHub's one-second timing resolution, compared with 5 minutes -37 seconds in the baseline job; the supplied tools needed no package install. -This measures that step, not the complete workflow or its queue time. - -A [completed main run](https://github.com/stablyai/orca/actions/runs/36962172614) -used 42 aggregate runner-minutes across 11 test jobs. The -[latest daily demand report](https://github.com/stablyai/orca/actions/runs/36965354205) -estimates 34.9 headless runner-hours, including 23.4 in cancelled runs. These are -baseline observations; post-merge savings have not yet been measured. - -## October 4 clock, byte and import test fixtures - -These changes retain production behavior, original case names and platform -outcomes. The paired pilots use three alternating one-worker Node 24 invocations -on Ubuntu 24 ARM. Every median below is a complete focused test invocation; -they do not establish whole-shard savings or queue-delay improvements. - -| Workload | Baseline median | Candidate median | Reduction | Hosted evidence | -| ---------------------------------------------------- | --------------- | ---------------- | --------- | ------------------------------------------------------------------------ | -| Codex settlement and Claude stop deadlines, 35 cases | 35.914s | 10.142s | 71.8% | [37186232658](https://github.com/stablyai/orca/actions/runs/37186232658) | -| Native-chat delivery, 15 cases | 22.321s | 7.376s | 67.0% | [37183731823](https://github.com/stablyai/orca/actions/runs/37183731823) | -| Six profile-storage byte suites, 118 cases | 12.629s | 9.978s | 21.0% | [37183141654](https://github.com/stablyai/orca/actions/runs/37183141654) | -| Encrypted account storage, six cases | 18.277s | 0.958s | 94.8% | [37184007241](https://github.com/stablyai/orca/actions/runs/37184007241) | -| SSH remote commands, 27 cases | 6.840s | 6.078s | 11.1% | [37184454532](https://github.com/stablyai/orca/actions/runs/37184454532) | -| OpenCode subscription, 28 cases | 47.347s | 6.284s | 86.7% | [37184858823](https://github.com/stablyai/orca/actions/runs/37184858823) | -| Window-service attachment, 31 cases | 11.910s | 1.619s | 86.4% | [37186340840](https://github.com/stablyai/orca/actions/runs/37186340840) | -| OpenCode 2 TUI ownership, 41 cases | 16.811s | 2.429s | 85.6% | [37187699312](https://github.com/stablyai/orca/actions/runs/37187699312) | -| Title-send authorization, nine cases | 29.545s | 12.995s | 56.0% | [37188423318](https://github.com/stablyai/orca/actions/runs/37188423318) | -| Range selection, 15 cases | 9.737s | 7.130s | 26.8% | [37188996198](https://github.com/stablyai/orca/actions/runs/37188996198) | - -Provider tests wait for the real fake-child write/ready barrier and drain the -host stream before installing a scoped parent clock. The original 2.5/5-second -Codex and 3-second Claude deadlines remain; before/at assertions check their -boundaries. Early rejection and refusal resolve without waiting on an unreachable -barrier; finally and suite teardown restore clocks and spies even after a native -Vitest timeout. Four healthy failure-path controls pass; old-helper hangs and -removed teardown are caught. Missing-child failure retains a real ten-second observation window. -Six faults for early/late stop deadlines, late queued resend and missing child -output fail the intended assertions. Native-chat tests still use the real React -outbox hooks. Their scoped clock retains probe, retry, churn and target-switch -windows, drains async act work, and unmounts before clock restoration. All eight -hook faults fail; two boundary faults pass the old coarse tests and fail the new -before/at assertions. Node/web typecheck, lint and formatting pass. - -Native Buffer.equals replaces deep per-byte assertion traversal. It checks the -complete original bytes and length; fixtures, SQLite operations, encryption and -processes are unchanged. The six profile suites retain 114 passes and four -existing Linux case-sensitivity skips; all 118 pass locally on macOS. Seven -profile last-byte/length faults and two encrypted-vault last-byte/length faults -fail their exact byte assertions. All six encrypted-vault cases still exercise -52 accounts near the 4 MiB encrypted cap, refused growth, restart/readback and -private permissions. - -The old SSH fixture created 15,197 short stage paths but never exceeded the real -1,048,576 UTF-16-character transport tail cap; its two valid entries also left -the 64-result assertion vacuous. The replacement uses about 1,300 real excluded -stage directories under long Unicode path components, a real shell channel and -the production execCommand limiter. Actual find output exceeds that cap, while -the generated filter retains both original valid entries. A separate population -of 65 valid directories proves the first-64 limit and native find ordering. File, -symlink, nested-install and failed-enumeration checks remain. All 27 case outcomes -match across treatments (23 passes and four unavailable PowerShell 5.1 skips on -the hosted image). Eight cap, ordering, filtering and failure faults are caught. -Paths use platform utilities; local PowerShell availability retains its original -skip policy. The deadline and SSH fixtures reuse existing process/stream code. -OpenCode subscription tests batch only uninterrupted fixture writes between the -original observation barriers. Both SQLite schema versions keep every row, rowid, -read cap, poll, clock position and case. No durability pragma or production code -changes. All 28 cases pass in every pair. Separate captures compare ordered -schema and rows, transaction state, pragmas, signals, page requests/results and -subscriber callbacks: 124,754,101 payload bytes match, canonical digest -`ba0eb6f09281746071d73fae88e2e8eb45332f36892b90998b374b1b8c59b3e3`. -Missing rows, collapsed frontier rowids, read-cap overruns, missing commit and -missing rollback each fail the intended case in both schemas. Positive rollback -controls pass. The unmeasured full-suite timing report ranked this fixture at -134.229 seconds; the paired 47.347-second figure above is the relevant focused -baseline, and the two figures must not be mixed into a claimed saving. - -Window attachment tests mock five unrelated registrar modules using their actual -types. The existing window ownership, reload, media permission, native file drop, -hydration barrier and updater scheduling cases remain real. Exact store/runtime/ -window arguments and daemon registration after the PTY handler are now asserted; -nine actual production wiring/order faults fail. The registrar implementations -retain separate handler tests. Concrete existing stubs moved to one fixture to -stay within the line limit. All 31 cases and statuses match across all pairs. The -final source differs from the timed source only by a required type-assertion -safety comment. -OpenCode 2 TUI tests use a scoped async clock only after the first real native -module import. The unchanged generated plugin runs against the existing fake TUI -and fetch. Original poll, permission, retry, preview, slow POST and endpoint -windows remain. Before/at assertions pin the 100 ms poll, 500 ms permission, -120 ms slow POST and 5-second endpoint boundaries. All 41 original case outcomes -match. Separate ordered captures preserve 678 TUI/event rows and 362 complete -POST start/completion rows; wall timestamps and cross-stream interleaving are -excluded from equivalence. Eleven generated-source faults fail their intended -identity, reload, order, deadline or timer-cleanup assertions. Four import/setup/ -disposer rejection and timeout controls confirm restoration of clocks, fetch, -argv and environment. Teardown holds its fake clock through bounded native cleanup and pending-timer checks, then restores real clocks. The [teardown qualification](https://github.com/stablyai/orca/actions/runs/37195736834) passes all 41 cases and 12 failure and restoration control invocations, including interval and retry leaks missed by the earlier teardown. -Title-send authorization tests retain real terminal creation, graph binding and -positive evidence paths. Negative process-evidence probes use scoped clocks -after setup: both 150 ms polling loops retain their 6,500 ms budget, crossed at -6,600 ms, and the send guard retains its 1,050 ms deadline. Before/at assertions -check 6,599/6,600 and 1,049/1,050 ms. All nine original cases remain. Actual -early/late wrapper and guard deadlines, false spinner identity and unknown-agent -authorization faults fail their intended assertions; native clocks and runtime -instance spies are restored after each failure. -Range-selection tests stub only the saved-note send menu, which their empty -comment populations never render. A facade typed from the actual menu props -throws if invoked, and an afterEach assertion verifies no call with unconditional -mock clearing in finally. The real hook, Monaco constants/model, line and range -drag behavior, draft-card lifecycle and open-inline-card chord remain. All 15 -original test bodies are byte-unchanged. Three actual range/hunk/draft faults -fail their original assertions; a real draft-render menu call hits the facade -and sentinel, and an outer cleanup check proves mock state cleared after failure. -The actual saved-note menu retains its independent component tests. - -## October 2 headless detector compiler cache - -The deferred detector already avoids dependency setup for known build inputs. -For changes that need import analysis, the collector marks package imports external; -only esbuild and its platform binary are needed. A small compiler archive can replace -root dependency setup for this analysis, while qualification jobs still install normally. - -The existing Linux x64 warmer packs these two packages after its frozen, -script-free, policy-checked install. Only main publishes. Readers use an exact key -covering Node/platform/architecture, manifests, install policy, patches and the cache -implementation. The producer and reader use the same archive path. File hashes, -identity and a compiler smoke are checked before availability is reported; missing, -invalid or failed restores use the original full installer. Graph analysis retains -its existing conservative full-qualification verdict on errors. - -A [three-pair hosted comparison](https://github.com/stablyai/orca/actions/runs/37071724200) -passed on Ubuntu x64 with Node 24.21.0 and esbuild 0.28.2. Every pair produced the -same 6,018 source inputs. Sample 2 ran the compiler-only treatment first; samples 1 -and 3 ran the existing installer first. Each used a fresh dependency tree, and the -compiler treatment required a real cache hit and validated its bytes and smoke. - -| Sample | Full installer + graph | Compiler restore + graph | Paired saving | -| ------ | ---------------------- | ------------------------ | ------------- | -| 1 | 11.730s | 2.805s | 8.925s | -| 2 | 14.702s | 4.706s | 9.996s | -| 3 | 14.666s | 3.750s | 10.916s | - -The median paired saving is 9.996 seconds. Inter-step overhead, archive transfer, -validation and the real graph are included. Checkout, initial Node setup, dependency -resets, seed work, post-job cache saves, tests and queues are excluded. These are -warm detector measurements, not whole-workflow or billing savings. The trial uses -the same package layout and validation as the production helper; production also -resolves its policy fingerprint. Cold or changed identities still install fully. - -## SSH Windows slot reuse - -The SSH Windows host workflow uses the same server-slot preparation action as -headless qualification. Its four PR jobs can restore the exact slot published by -fully qualified main runs, then validate its inventory and hashes and run the -required-slot and pinned-Node smoke checks. Misses or invalid payloads compile -freshly. Only main headless qualification publishes; manual SSH qualification -still builds freshly. Both sshd versions, both architectures, all three host -cells, the process-table addon build, and the template/relay builds remain. - -The action is part of the cache fingerprint, so this extraction starts a new -namespace that needs a successful main seed. The existing hosted measurements -below suggest about 160 aggregate runner-seconds saved across four warm SSH jobs; -that is a conditional estimate, not a measured improvement of this consumer. -Private sshd installation and host execution still dominate this workflow. - -## Prepared relay addon reuse - -A [completed SSH Windows run](https://github.com/stablyai/orca/actions/runs/36972043877) -rebuilt the process-table relay addon after native dependency preparation. From -its builder's start message to the validated staged artifact, x64 took 85.5 seconds -and ARM64 took 135.6 seconds. These are single-run observations, not medians. - -An opt-in reuse path checks the same binary architecture, patched reader and -launcher exports as staging, then runs the existing native-load and CreationTime -probe. Repaired source or incomplete evidence requires a fresh build. SSH PRs -request reuse only following an exact prepared native-cache hit; manual SSH and -all release builders retain fresh compilation. Subsequent staging checks still -run. Hosted validation and the reuse interval remain to be measured. - -## Windows root download stores: registry installs finish sooner - -Three paired samples on each Windows architecture compared the existing exact -main download-store restore with a fresh registry install. Each treatment used a -fresh dependency tree, store and pnpm metadata, with registry-first ordering in -sample 2. Both restored the same policy-checked verification record before timing. -All six pairs used Node 24.21.0 and pnpm 12.8.1; manifest digests and installed -lockfile digests matched, and both retained frozen, script-free installation. - -| Runner | Cached totals (seconds) | Registry totals (seconds) | Paired median saving | -| ------------- | --------------------------- | --------------------------- | -------------------- | -| Windows x64 | 26.820 / 28.885 / 27.751 | 13.644 / 14.126 / 12.908 | 14.759 seconds | -| Windows ARM64 | 216.540 / 288.492 / 189.342 | 119.856 / 238.562 / 115.611 | 73.731 seconds | - -The x64 samples are the three successful Windows 2022 jobs in -[run 37064549378](https://github.com/stablyai/orca/actions/runs/37064549378). -Its ARM cleanup guard rejected pnpm's setup-owned store path before measurement; -those incomplete ARM jobs are excluded. The corrected -[ARM-only run](https://github.com/stablyai/orca/actions/runs/37065220916) passed all -three samples. Earlier rejected measurements also stopped before installation -because an optional config file was absent; none count toward these timings. - -Intervals include actual store lookup/restore, inter-step overhead and root -installation. Checkout, toolchain setup, tree/store reset, verification-record -restoration and native preparation are excluded. ARM variation is substantial; -these samples do not measure whole-workflow, queue or billing savings. - -Root-only Windows x64/ARM64 PR installs now skip the download-store restore. -The existing x64 mixed-install exception remains. An explicit store opt-out also -lets Windows headless persistence and SSH jobs avoid the archive on main or -manual runs. Frozen installs, verification records, native caches and every -qualification check remain. Other lockfile sets and platforms keep their -existing policy. Default non-PR writers, including the warmer, still seed stores -for direct setup-node consumers and workflows that run package scripts. - -## October 2 Linux root store comparison - -A [six-job hosted comparison](https://github.com/stablyai/orca/actions/runs/37073978443) -measured the actual main root-store archive against direct registry installation, -with three fresh-runner pairs on each Linux architecture. All six jobs passed. -The middle sample on each architecture reversed treatment order. Between treatments, -the driver removed the dependency tree, store and pnpm metadata, then restored the -same policy-checked verification record before timing. Frozen, script-free installs -preserved policy files and produced identical installed lockfile digests in each pair. - -| Architecture/sample | Store restore + install | Direct registry install | Paired saving | -| ------------------- | ----------------------- | ----------------------- | ------------- | -| x64 / 1 | 6.516s | 5.372s | 1.144s | -| x64 / 2 | 6.743s | 3.950s | 2.793s | -| x64 / 3 | 6.600s | 3.985s | 2.615s | -| ARM64 / 1 | 7.632s | 3.346s | 4.286s | -| ARM64 / 2 | 5.504s | 3.575s | 1.929s | -| ARM64 / 3 | 5.450s | 3.364s | 2.086s | - -Median paired savings are 2.615 seconds on x64 and 2.086 seconds on ARM64; means -are 2.184 and 2.767 seconds. Both used Node 24.21.0 and pnpm 12.8.1. Actual store -lookup/transfer/restore, inter-step overhead and installation are timed. Checkout, -initial toolchain/dependency setup, preparing the existing process wrapper, -dependency resets, verification-record restores, native work, tests, post-job cache -saves and queues are excluded. Package services have already been used by initial -setup. These are warm-policy setup measurements, not workflow or billing savings. - -The shared installer consequently skips root-only Linux x64/ARM64 store restores -on PRs. It still installs and checks every package through pnpm. Mixed mobile and -custom lockfile sets, other architectures, verification/native caches, -main store writers and release installation policies keep their existing behavior. -The measured Windows exceptions remain. Mac restores were retained at this stage; -the following comparison supersedes that policy. No periodic job or cache is added. - -## October 2 macOS root store comparison - -A [six-job hosted comparison](https://github.com/stablyai/orca/actions/runs/37078232553) -used the same paired method on macOS 15 Intel and Apple Silicon. All six jobs -passed with actual main store cache hits. The middle sample reversed treatment -order. Each treatment started with a reset dependency tree, store and pnpm -metadata, followed by the same verification-record restore. Policy files and -installed lockfile digests matched within every pair. - -| Architecture/sample | Store restore + install | Direct registry install | Paired saving | -| ------------------- | ----------------------- | ----------------------- | ------------- | -| x64 / 1 | 87.270s | 60.584s | 26.686s | -| x64 / 2 | 103.280s | 52.640s | 50.640s | -| x64 / 3 | 61.564s | 40.768s | 20.796s | -| ARM64 / 1 | 26.024s | 14.375s | 11.649s | -| ARM64 / 2 | 37.000s | 19.726s | 17.274s | -| ARM64 / 3 | 34.616s | 14.888s | 19.728s | - -Median paired savings are 26.686 seconds on x64 and 17.274 seconds on ARM64; -means are 32.707 and 16.217 seconds. Node matched within each pair: 24.19.0 on -Intel and 24.20.0 on Apple Silicon, as resolved by the existing installer. Both -used pnpm 12.8.1. Timing includes actual cache lookup/transfer/restore, inter-step -overhead and installation. Checkout, initial setup, process-wrapper preparation, -resets, verification restores, native work, tests, cache saves and queues are -excluded. Initial setup has already used package services. These measurements -do not establish whole-workflow or billing savings. - -The existing root-only PR exception now also covers macOS x64/ARM64. Frozen, -script-free installs and pnpm policy checks still run. Mixed/custom lockfile sets, -other architectures, verification/native caches, main/manual store writers and -release installation policies retain their existing behavior. No periodic job -or cache is added. - -## October 2 store producers: keep caches without downloading hits - -The optional `cache-pnpm-store-lookup-only` installer input uses -[`actions/cache` lookup-only](https://github.com/actions/cache#inputs) on non-PR -runs. An exact hit refreshes cache access without extracting the archive; a miss -still installs from the registry and publishes the populated store at successful -job completion. The default remains the existing `setup-node` cache behavior. -The four Linux/Windows dependency warmers and Linux/macOS persistence producers -opt in. Windows persistence retains its existing store opt-out, and PR restore -policies are unchanged. This adds no recurring job or extra cache family. - -A [tiny framework control](https://github.com/stablyai/orca/actions/runs/37082688033) -proved that lookup left the payload absent, refreshed the existing cache's access -time, and published a miss that a fresh job restored. A -[nested composite control](https://github.com/stablyai/orca/actions/runs/37084946789) -then saved and restored a fresh payload using the actual environment-path pattern. -The installer exports its resolved store path through `GITHUB_ENV`: twice-nested composite post-job saves -cannot resolve their internal step outputs. The primary key is captured -by the cache action before cleanup. Paths, architecture and lockfile keys match -`setup-node`, so existing default-branch archives remain reusable. - -The [six-platform installer screen](https://github.com/stablyai/orca/actions/runs/37084946789), -[Linux repeats](https://github.com/stablyai/orca/actions/runs/37085164277), and -[corrected Windows repeats](https://github.com/stablyai/orca/actions/runs/37085248976) -compared the complete shared installer, including toolchain setup, cache actions, -policy verification, frozen installation and native probes where requested. -Each treatment reset dependencies, the store, pnpm metadata and the Windows -registry build directory. Treatment order reversed across architectures and -repeats. Every qualified pair required real main store cache hits, matching -policy/installed-lockfile digests and Node/pnpm versions, plus exact native-cache -hits on Linux and Windows. The initial Windows x64 screen stopped before timing -because its benchmark guard rejected the standard `D:\.pnpm-store` path; that -unqualified job is excluded. - -| Platform / sample | Restore + installer | Lookup + installer | Paired saving | -| ------------------------ | ------------------: | -----------------: | ------------: | -| macOS ARM64 | 37.785s | 20.024s | 17.761s | -| macOS x64 | 75.610s | 46.638s | 28.972s | -| Linux ARM64 / 1 | 10.005s | 7.242s | 2.763s | -| Linux ARM64 / 2 | 8.817s | 6.672s | 2.145s | -| Linux ARM64 / 3 | 8.800s | 6.581s | 2.219s | -| Linux x64 / 1 | 12.309s | 9.735s | 2.574s | -| Linux x64 / 2 | 10.131s | 8.690s | 1.441s | -| Linux x64 / 3 | 13.741s | 9.485s | 4.256s | -| Windows ARM64 / repeat 1 | 145.818s | 78.729s | 67.089s | -| Windows ARM64 / repeat 2 | 152.730s | 106.475s | 46.255s | -| Windows ARM64 / screen | 294.092s | 193.957s | 100.135s | -| Windows x64 / 1 | 30.936s | 20.256s | 10.680s | -| Windows x64 / 2 | 31.021s | 20.098s | 10.923s | - -All 13 qualified pairs improved. Median paired savings were 2.574 seconds on -Linux x64, 2.219 on Linux ARM64, 10.802 on Windows x64 and 67.089 on Windows -ARM64. Each macOS architecture had one pair; its 28.972 / 17.761 second savings -are a screen, supported by the earlier three-pair root-store comparisons. - -All pairs used pnpm 12.8.1. Node was 24.21.0 on Linux and Windows, 24.19.0 on -macOS Intel and 24.20.0 on macOS ARM. Source dependency policies were frozen for -this screen; later main dependency changes do not extend these measurements. -Timing excludes checkout, initial service/bootstrap use, wrapper compilation, -resets, result validation, post-job saves and queues. These are installation -measurements, not whole-workflow or billing savings. Cold publication is verified -separately by the small controls; no large synthetic store cache was uploaded. - -## October 1 Windows and dependency cache follow-up - -[PR #24355](https://github.com/stablyai/orca/pull/24355) merged at `197ea3a3`. -The [next hosted trial](https://github.com/stablyai/orca/actions/runs/36917210453) -ran four alternating pairs on each Windows architecture and both Mac -architectures. All six jobs passed. The exact temporary workflow and drivers -remain available at `ab24952ab5af0f7d8e53a544896ac2486edfb244`; completed trial -tooling is removed from ordinary PR CI. - -### Windows server slots - -The existing dependency-native cache and the server's N-API 8 slot serve different -consumers. Cache the small server slot separately, using the exact compiler image, -architecture, dependency/patch/runtime inputs and compilation/validation source. -PR and main-push qualification restore it. Nightly qualification still compiles -freshly; main saves after persistence/lifecycle tests and the existing x64 Node 18 handoff. -Templates and explicit-ref calls continue to compile freshly. - -| Hosted runner | Fresh build median | Restore median | Difference | -| ---------------- | ------------------ | -------------- | ---------- | -| Windows 2022 x64 | 18.624s | 4.677s | 13.947s | -| Windows 11 ARM64 | 72.002s | 6.111s | 65.891s | - -Both used Node 24.21.0. Each comparison includes actual GitHub cache restoration, -inter-step time, payload validation, required-slot checks and pinned-Node load/spawn -smoke. Every fresh build uses the builder's freshly cleared compilation directory. -Dependency installation, initial seed work, shared download warmup, full qualification -tests and queues are outside these timing intervals. The x64 payload was 2,935,529 -bytes and ARM64 3,533,557 bytes. These are conditional warm-hit gains, rather than a -whole-workflow improvement or the roughly 95–117 seconds seen in earlier cold samples. - -Restored bytes must match current module/version/headers/N-API/host metadata, -the complete inventory and hashes, and current vendored ConPTY files. Existing -patch, PE architecture, post-baseline N-API and MSYS breakaway checks are reused. -Failed or partial restoration and invalid payloads clear only `out/orcad-prebuilds` -and fall back to normal compilation. Both fresh seeds and fresh restored consumers -passed the full existing Windows qualification; x64 also passed Node 18 handoff. -The final key additionally includes the ten transitive process-wrapper sources. - -A bounded sample of 50 first-parent main commits through `197ea3a3` had 45 of 49 -adjacent transitions with identical source inputs, including those ten files. -This one-day sample holds the new validator and runner image constant. Actual -image rotation, seed availability and changed PR inputs can reduce reuse. - -Keep the original Windows native dependency preparation: artifact-mode tests still -import checkout `node-pty`, registry and process-reader addons, and the Windows -artifact builder stages the patched process reader. The cached server slot does -not replace those dependencies. - -### pnpm verification records - -Reuse pnpm's existing policy-checked record on Windows x64/ARM64 and Mac Intel, -while retaining Linux's existing behavior. The exact OS/architecture/pnpm/policy -key, explicit opt-out, frozen installs and main-only production writes remain. -Each treatment resets links and registry metadata, retains the identical warm -download store, and checks unchanged manifests, installed versions and installed -lockfile bytes. Every platform passed real missing/corrupt-record fallback and -changed-policy/changed-integrity rejection controls, plus authentic cross-path reuse. - -| Platform | Baseline install | Cached install | Baseline interval | Cached interval | -| ---------------------- | ---------------- | -------------- | ----------------- | --------------- | -| Windows x64 | 13.105s | 9.767s | 17.717s | 15.146s | -| Windows ARM64 | 96.723s | 88.443s | 164.620s | 157.877s | -| Mac Intel | 31.368s | 21.042s | 47.727s | 41.302s | -| Mac ARM, kept disabled | 11.419s | 8.086s | 17.905s | 15.787s | - -Install columns time pnpm alone. Interval columns also include link reset, -path validation, real cache restoration and inter-step overhead; they exclude -seed setup, parity checks and queues. Key-resolution overhead is outside the paired -trial; the real Windows ARM production-shaped step took 0.296s. Windows ARM's first -pair was slower with the record, while the next three improved, so its median is -not a guaranteed per-job saving. Macs used their hosted Node 24.19.0 Intel and -24.20.0 ARM toolchains; Windows used 24.21.0, and all used pnpm 12.0.0. - -An earlier complete Mac ARM trial saved only 0.681s in its interval median. -That small, variable margin does not justify enabling the extra lookup there. -Only the tiny pnpm-owned record is restored; registry metadata and download-store -policy are unchanged. Changing the shared action also produces a one-time cold -native dependency cache key, whose existing hash includes the action bytes. - -### Further local screens rejected - -A fixed 256-file, 2,471-case cohort preserved every result in twelve invocations. -Lazy per-file user data showed a noisy 2.0% wall difference with no setup-time -improvement; combining DOM setup was 1.8% slower. A separate seven-file leaf-import -screen preserved all 50 assertions and public/hook controls, saving only 0.903 -worker-seconds while elapsed time rose 2.55%. All experiments were reverted. -Documentation-result reuse also lacked positive eligible demand in the bounded -sample, and the apparent example changed the actual tested merge/base inputs. -These results do not justify adding a new coverage-selection or result-reuse policy. - -## October 1 PR concurrency follow-up - -### Where the next gains are - -The [September 30 demand report](https://github.com/stablyai/orca/actions/runs/36816009362) -samples 282 of 4,194 runs across workflow/conclusion strata. It estimates full job -duration for runs created in the reporting window, rather than occupancy clipped -to that window. PR CI accounts for about 537 runner-hours, including about 321 -hours of unit shards, 73 hours of E2E, 41 hours of Windows packaging, 35 hours of -Linux packaging, 18 hours of static analysis, and 11 hours of typechecking. These -are weighted estimates, not exact billing totals. Unassigned and incomplete jobs -are excluded. The report predates the merged planning/setup change below. - -The [full reference run](https://github.com/stablyai/orca/actions/runs/36841821670) -ran 10,310 files exactly once. Across its five shards Vitest reports about 5,063 -worker-seconds importing modules, 2,838 running tests, 685 transforming source, -267 setting up tests, and 390 preparing environments. Workers overlap, so these -figures cannot be added to predict job elapsed time. They identify repeated -imports and real-time test waits as larger targets than line-count reporting, -which uses about 1.1 runner-hours in the same demand sample. - -Parallel steps share the job's CPU and memory. Their -[background/wait support](https://github.blog/changelog/2026-06-25-actions-steps-can-now-be-run-in-parallel/) -saves repeated runner setup when independent checks fit together, but does not -increase machine resources or the account's concurrent-job allowance. Cheap -preflight checks still gate expensive unit and package jobs. - -Organization metadata reported the Team plan on October 1. GitHub documents -[60 standard concurrent jobs by default](https://docs.github.com/en/actions/reference/limits#job-concurrency-limits-for-github-hosted-runners) -and allows support requests for increases. The effective configured allowance -was not exposed by the API. A seven-second, repository-only sample at 10:48 UTC -found 32 queued jobs and 70 assigned job records marked in progress, including -one Blacksmith label. Those non-atomic records, runner turnover, other repositories -and provider labels cannot establish the actual allowance or simultaneous usage. -These changes reduce demand; they do not change account settings. If queues -remain, ask GitHub Support to confirm the effective organization limit before -choosing a larger allowance or paid runners. - -The [follow-up PR](https://github.com/stablyai/orca/pull/24355) measures virtual -readiness deadlines in captured-transcript tests, E2E allocations with no general -consumer, renderer projection, native setup, Docker fixtures, and Windows store -restoration. Alternating hosted comparisons check output parity. Completed -benchmark workflows and drivers are removed; their trial commits retain the -exact reproduction code. - -A [three-pair hosted transcript comparison](https://github.com/stablyai/orca/actions/runs/36849458799) -on one four-worker ARM runner measured baseline invocations at 127.001 / 114.307 / -114.265 seconds, versus 28.903 / 28.663 / 28.455 seconds with virtual readiness -deadlines. Median elapsed time for these five files fell about 75%. All 259 -original named tests passed in every baseline and candidate, and candidates -also passed three repaint checks. Summed test-body time fell from a median 325.85 -to 31.80 worker-seconds. This comparison includes Vitest startup/import work but -excludes checkout, dependency setup, and queues; it is not a measured percentage -improvement in the full unit matrix. Real emulator setup and drains remain real, -and the same captured bytes, readiness/refusal deadlines, and assertions run. - -The same hosted trial passed three forced Electron native rebuilds while the -external node-gyp path pointed to a nonexistent file. A separate fresh consumer -then restored the native cache, required a real hit, and passed the existing -Electron binary probe. The Linux Node-runtime workaround remains intact. - -For a network-only E2E selection in that trial, both dedicated network jobs -passed. The previous allocations consumed 87 seconds for the Electron build, -34 for the native primer, and 67 for a general job whose log confirmed that -every selected spec belonged to a dedicated lane. The candidate skipped those -three jobs before runner allocation, avoiding 188 runner-seconds in this case. -This single-case measurement excludes queues and does not predict savings for -mixed selections; the existing dedicated SSH, IME, and ordinary E2E routes remain. - -A [controlled cancellation trial](https://github.com/stablyai/orca/actions/runs/36853246785) -verified both condition outcomes. After intentional cancellation, `always()` -started another 60-second follow-up, while `!cancelled()` skipped it. Cleanup -and artifact uploads succeeded in both treatments. A separate deliberately -failed test still ran its later `!cancelled()` test and cleanup. The 60 seconds -are synthetic condition evidence, not a measurement of a real SSH test's cost. -Completed comparison workflows are removed after recording their evidence; -the exact drivers and workflow remain reproducible at the trial's source commit. - -The [first full PR validation](https://github.com/stablyai/orca/actions/runs/36849458648) -passed all five unit shards and both Linux/Windows package checks. Its reports -contain 10,326 unique files, each once, with zero unhandled errors. Full shard -job durations ranged from 544 to 581 seconds. They ran a different merged source -on different allocations from the earlier reference, so comparing their totals -does not establish an end-to-end speedup. The alternating transcript comparison -above is the controlled timing evidence. - -The [updated full PR validation](https://github.com/stablyai/orca/actions/runs/36855833565) -passed all required gates and both package checks. All 10,335 discovered files -appear once across five passing timing reports, with zero unhandled errors; -122 modules have the existing expected skipped status. Named merge-tree discovery -and saved assignment replay match both successful validation runs. The latest -unit jobs took 520–558 seconds. Refreshing weights with that same measurement set -would reduce the largest projected load from 1,756.706 to 1,708.187 worker-seconds -(2.76%), while retaining the same 8,540.831 total. Applying the earlier successful -run's proposed weights to the latest measurements improves the maximum only -1.54% and the median 0.68%. These small, variable projections do not establish -an elapsed-time gain, so the existing weights remain. - -A [three-pair Windows store comparison](https://github.com/stablyai/orca/actions/runs/36853246494) -used fresh dependency trees, stores and pnpm metadata before each treatment. The -middle pair reversed order. Cached totals include archive restoration and both -unchanged frozen installs; the mobile install ran every existing postinstall -generator. Setup/reset time and runner queues are excluded. - -| Pair | First treatment | Cached total | Registry total | Registry saving | -| ---- | --------------- | ------------ | -------------- | --------------- | -| 1 | Cached | 71.092s | 37.500s | 33.592s | -| 2 | Registry | 73.556s | 39.935s | 33.621s | -| 3 | Cached | 75.660s | 37.001s | 38.659s | - -Treatment medians were 73.556s cached and 37.500s registry; median paired saving -was 33.621s. Cache restore and step overhead alone cost a median 27.955s. Policy -and all six generated-output digests matched across all six treatments. Windows -x64 PR jobs using the mixed root/mobile key now skip its download-store restore; -PRs already skip store saves. Root-only Windows stores, Windows ARM64/x86, other -operating systems, non-PR writers, native/Electron caches and frozen-install policy retain their existing -behavior. The trial covers this mixed install on Windows 2022, not every Windows -dependency key or a whole PR's elapsed time. - -A [three-pair daemon fixture comparison](https://github.com/stablyai/orca/actions/runs/36853246776) -reset Docker build caches and the fixture/base image before every treatment. -Warm totals include the existing action's archive restore/load and the unchanged -daemon descendant oracle; reset time and runner queues are excluded. The middle -pair again reversed order. - -| Pair | First treatment | Cold total | Warm total | Warm saving | -| ---- | --------------- | ---------- | ---------- | ----------- | -| 1 | Cold | 28.984s | 20.673s | 8.311s | -| 2 | Warm | 21.798s | 28.953s | -7.155s | -| 3 | Cold | 28.964s | 17.614s | 11.350s | - -Treatment medians were 28.964s cold and 20.673s warm; median paired saving was -8.311s. Oracle medians fell from 28.886s to 4.907s, while archive restore/load -cost 13.025–24.024s (15.766s median) for a 714,643,456-byte archive. BuildKit -confirmed warm provisioning was cached and cold provisioning was not; every -treatment used the same immutable base image, reaped the descendant and kept -the canary alive. The PR keeps restoration in the background during root/mobile -installation. One serial pair was slower, so the roughly 24s oracle reduction -is not a guaranteed total runner saving. The drivers and exact workflows remain -available at source commit `9231d1be6c76ccc1d2fef741a4e68ae29735a5c8`. - -A [three-pair AppImage compression comparison](https://github.com/stablyai/orca/actions/runs/36855833100) -packaged the same complete Linux x64 app with the pinned 1.0.3 toolset and -mksquashfs 4.6.1. Each timing includes the private app copy, electron-builder -and blockmap generation. Tool download, app compilation, extraction, parity -checks and queues are excluded; the middle pair reversed order. - -| Pair | First treatment | Default zstd 15 | PR zstd 3 | Saving | -| ---- | --------------- | --------------- | --------- | ------- | -| 1 | Default | 22.580s | 11.760s | 10.820s | -| 2 | PR | 22.612s | 12.560s | 10.052s | -| 3 | Default | 22.512s | 11.815s | 10.697s | - -Median packaging time fell from 22.580s to 11.815s (47.7%); median package size -grew from 213,873,601 to 237,570,611 bytes (11.1%). All six extracted manifests -matched every path, byte, mode and symlink target: 4,096 files and 601,194,650 -file bytes. The runtime prefix matched the pinned runtime exactly, and stored -SquashFS options confirmed the actual level-3 override. All static checks passed; -the representative baseline and candidate each passed the unchanged headless -and CLI journeys, including all entrypoints and both shutdown signals. The -directory build passed the existing glibc floor checks on all 19 native binaries. - -Only the PR Linux x64 AppImage child receives the private tool overlay. It resolves -the existing custom-tool override first, checks the pinned tool/version and zstd -configuration, and reuses the original runtime, validator and libraries. Cleanup -waits for every package worker even on failure. Release settings, Linux ARM, -Debian/RPM packaging and all native/package gates retain their existing behavior. -The isolated AppImage result does not establish the full three-format job's gain. - -The same hosted comparison projected one already-built renderer three times per -treatment, again alternating order. Baseline times were 8.079 / 8.219 / 8.129s; -candidate times were 2.568 / 2.477 / 2.403s. Median projection fell from 8.129s to -2.477s (69.5%, 5.653s saved). All six web snapshots matched all 1,137 files and -51,957,014 bytes, and the renderer input remained unchanged after every run. -These timings include the projector process but exclude renderer compilation, -checkout, setup and queues. The drivers and workflow remain available at source -commit `8ba5c9bf9f734d585f5e89945519aef4f607face`. - -Two local cache screens do not justify enabling Node's compile cache. A 96-file -screen with an explicit worker flush produced a small, noisy difference. A larger -256-file screen retained all 2,088 tests: baseline elapsed times were -54.630 / 55.105 / 55.171 seconds, fresh caches 53.719 / 54.577, and a warm cache -53.732. The roughly 1.7% median difference is too small to justify cache transfer -and another test hook without stronger hosted evidence. - -Vitest 4.1.11's experimental filesystem module cache is more promising locally, -but raw reuse is unsafe. A 96-file screen fell from about 8.4 to 5.9 seconds with -a warm cache, while a cold cache cost about 3%. Negative controls then reproduced -false passes after adding a preferred import extension, retargeting a symlink, -changing package exports, or changing transform inputs. Cache-disabled controls -failed correctly. A cache key must cover resolution and transform inputs as well -as file contents before any production trial; source hashes alone do not suffice. - -Affected-test selection remains in shadow mode. Its first merge was September -28, so October 1 cannot satisfy the documented week of evidence. Seven sampled -complete reference reports included one red run, but only two evaluated a smaller -candidate set; each omitted about 177–179 worker-seconds out of 8,073–8,392. Five -full fallbacks are not selection-validation evidence, and two other sampled red -runs had no review artifact. These samples support keeping the conservative -policy, rather than claiming that omitting roughly 12% of files would omit the -same fraction of work. - -### Shared planning and typechecking - -PR planning now shares checkout and dependency setup with typechecking. Planning -runs in the background, with an explicit failure-propagating join before its -artifact is published. Static analysis remains on a separate runner. Its existing -native dependency install is pinned to Node 24 so it also fills the unit matrix's -native cache, replacing the separate conditional PR primer. Daily Node 24/26 -reference planning and priming remain independent. - -The planner also avoids constructing the import graph when global changes or -missing change evidence already require the full suite. This keeps the same -fallback reason, discovery list and execution coverage. Five PR unit shards, -static/type gates, package checks and the shadow-selection policy are retained. - -A [three-sample hosted comparison](https://github.com/stablyai/orca/actions/runs/36833782900) -measured standalone type jobs at 35 / 26 / 42 seconds and planning jobs at -36 / 25 / 24 seconds. Shared type/planning jobs took 38 / 37 / 27 seconds. -The sum of the separate job medians fell from 60 to 37 seconds. Using that estimate with -the 139-second median static job models about 12% less preflight occupancy. -This is not a measured reduction in total CI time or queue delay. All twelve -plans matched, and three deliberately fresh native-cache producers were reused -by three successful consumers with the real native dependency probe. - -After updating to the current main branch, a -[six-pair alternating comparison](https://github.com/stablyai/orca/actions/runs/36840171700) -reset compiler state before every measurement. Direct compiler/planner command -time was 11.59 / 11.31 / 10.96 seconds sequentially versus -6.82 / 6.52 / 6.33 seconds together with restored state. With state deleted, -it was 67.36 / 67.05 / 61.75 versus 55.52 / 60.06 / 55.68 seconds. -All commands passed and plans matched within each comparison. These command -timings exclude setup and queues; compiler variation contributes to the cold -difference. The retained arrangement showed no cold compiler penalty. - -The sequential gate in [October 4 shared PR preflight capacity](#october-4-shared-pr-preflight-capacity) supersedes the earlier rejection below. - -Combining static analysis too was rejected. An -[alternating same-runner comparison](https://github.com/stablyai/orca/actions/runs/36835091650) -saved runner occupancy, but cold compilation slowed from 61–64 to 83–89 seconds -under contention. The combined gate would finish roughly 40 seconds after the -separate static gate. An -[earlier-compiler trial](https://github.com/stablyai/orca/actions/runs/36837523594) -did not remove that penalty. Small localization scheduling changes also lacked -a repeatable gain. The temporary pilot workflows and benchmark tooling are -available in those runs' commits, rather than retained in normal PR CI. - -Line-count reporting stays separate: it uses a small runner and trusted scripts -with PR write permissions. Sharing that job's credentials with PR-source checks -would provide little resource benefit. Reducing unit shards would trade saved -setup for a longer critical path; enabling selected tests requires the existing -week of representative shadow evidence. - -## September 27 follow-up - -### Shared E2E CLI output - -E2E consumers previously compiled the CLI individually even though they downloaded -shared Electron, web, and relay output. The producer now compiles the CLI once, -in parallel with web projection after Electron has finished clearing `out/main`. -Consumers repair executable permissions and install their own dev launcher with -the same preparation script used by local CLI builds. Older refs without that -script retain their original per-consumer compilation. - -An [eight-sample comparison](https://github.com/stablyai/orca/actions/runs/36307081200) -measured producer time increasing from 26.3–27.9s to 38.8–41.0s, while consumer -CLI compilation fell from 20.9–21.2s to 0.06–0.07s of direct preparation. All 5,519 -output files matched byte-for-byte, and every sample passed the CLI help smoke. -A four-sample -[final implementation comparison](https://github.com/stablyai/orca/actions/runs/36307382635) -also passed parity and CLI smoke checks. Consumer compilation took 4.7 / 12.8s -versus 0.08 / 0.06s of preparation; producer time increased by 0.2 / 7.1s in -the paired trials. Across both runs this models roughly 1.1–4.7 aggregate runner -minutes saved across 14 consumers, before artifact transfer overhead. Runner -variation is substantial; this is not a measured workflow wall-time reduction. -Test coverage and deadlines stay intact. - -[PR #23368](https://github.com/stablyai/orca/pull/23368) overlaps shell installation -with dependency setup, starts localization extraction before the orcad smoke, -and prepares mobile route snapshots while WebKit and the bundle are being built. -Its 27 checks passed without retries; seven existing conditional checks skipped. - -Same-runner comparisons in both orders measured: - -| Work | Before | After | Evidence | -| ----------------------------------- | -------------- | -------------- | ---------------------------------------------------------------------------------- | -| Static block | 63.5 / 74.7s | 38.9 / 55.6s | [Full comparisons](https://github.com/stablyai/orca/actions/runs/36302208990) | -| Shell job, downloads warmed equally | 76.8 / 70.3s | 64.4 / 64.0s | [Controlled shell runs](https://github.com/stablyai/orca/actions/runs/36302612626) | -| Mobile preparation | 19.8–21.6s | 18.0–18.5s | [Eight measurements](https://github.com/stablyai/orca/actions/runs/36302877583) | -| Web projection and mobile build | 17.2–17.3s | 11.7–12.1s | [Eight measurements](https://github.com/stablyai/orca/actions/runs/36302974324) | -| Mobile verifier fixture suite | 25.47 / 25.35s | 20.04 / 19.92s | [Four full-suite runs](https://github.com/stablyai/orca/actions/runs/36302692823) | -| E2E build outputs | 28.6–30.4s | 25.8–27.8s | [Eight measurements](https://github.com/stablyai/orca/actions/runs/36304001325) | - -Full mobile-job timings were dominated by first-run apt installation and browser -test variation; the controlled preparation measurement is the scheduling evidence. -All 1,248 web/mobile output files matched byte-for-byte in the build comparison. -The fixture suite kept all 47 tests, isolated mutable copies, and the verifier's -two fresh builds. No deadline, isolation, or worker-count changes were needed. - -E2E builds reuse the existing isolated main/preload/renderer build wrapper, now -forwarding `--mode e2e` to each target. All 2,640 output files matched byte-for-byte -in both execution orders, including the exposed test store and relay artifacts. -This saves a few build seconds; it does not speed up the E2E tests themselves. - -The existing unit assignment was already balanced at about 919 historical -worker-seconds per shard; fresh x86 elapsed times still ranged from 254 to 433s. -Refreshing weights alone would encode runner variation rather than resolve it. -An [identical-source architecture pilot](https://github.com/stablyai/orca/actions/runs/36302250920) -ran shards 1 and 8 on both four-CPU hosted runners. Complete jobs improved from -449 to 404s and 461 to 384s on ARM, including setup; test and skip counts matched. -A [full ARM run](https://github.com/stablyai/orca/actions/runs/36302906752) then passed -all eight shards in 329–373 test seconds (365–412 job seconds). Uploaded reports -matched the same complete x86 assignment: 9,876 files, each exactly once, no -unhandled errors. These are elapsed samples excluding queue time, not a guarantee -that every ARM allocation is faster than every x86 allocation. - -PR unit shards and their cache primer now use ARM; a main-branch warmer seeds -that architecture's existing native and pnpm cache keys. Native, package, and -relay gates continue on x86. The daily workflow retains complete x86 coverage on -both Node 24 and Node 26, including relay integration. Thus PR unit architecture -changes, while x86 unit coverage remains scheduled; this is an explicit coverage -placement tradeoff rather than a claim of identical per-PR host coverage. - -PR validation exposed a WebRTC probe timeout inside a hidden renderer. Isolated -and four-concurrent probes passed on both architectures; the original cause is -unproven. The probe now uses Electron main for the same three-second observation -interval. A [fault-injection comparison](https://github.com/stablyai/orca/actions/runs/36304349257) -passed with renderer timers unavailable on both architectures, while the original -probe failed the negative control. Packet assertions and deadlines are unchanged. -A subsequent Windows run timed out in the installer's real CIM process query -after verifying restricted policy. Its unchanged probe now runs before the -concurrent native suite, removing that source of contention without relaxing -the twenty-second process deadline or dropping either PowerShell architecture. - -Replacing Vitest deep comparisons with Node assertions in the status-store -oracle saved only about one local second in an initial trial. The change was -not retained: that evidence did not justify changing assertion semantics. - -The combined root/mobile pnpm cache is now present on main and was restored in -the September 27 static comparison, so another warmer for that key is unnecessary. - -A [fixture-warmer overlap trial](https://github.com/stablyai/orca/actions/runs/36303613587) -ran faster after initialization but exposed a first-use action-download race: -both background composites downloaded `actions/cache@v5` simultaneously, and -one briefly could not find `restore/action.yml`. The existing cache fallback -rebuilt the image and the job passed, but that recovery erased the speedup. -Keep the dedicated warmer serial. PR package restores remain safe from this -observed first-use race because their earlier top-level cache action is loaded -before the composites start; the workflow contract now preserves that ordering. - -## Four follow-up changes - -- Keep the readiness event, but reuse required checks only after an Actions API - lookup proves that the same PR head, tested merge commit, and workflow commit - already completed successfully. A changed base, missing proof, failed lookup, - or still-running check falls back to the full checks. Advisory tests retain - their normal readiness routing. The mobile and line-count workflows have no - draft-dependent work, so they no longer run again when a draft becomes ready. -- Route the headless-runtime matrix using the actual headless build and selected tests' - transitive imports, with conservative inclusion for dynamic workers, native - inputs, fixtures, and toolchain changes. A graph failure runs the full matrix; - manual dispatch still runs all ten platform jobs. The shared test selectors - retain the same 85 files. Unrelated shard timings and mobile-test tooling can - skip the matrix; shared shortcut definitions remain real runtime dependencies - and still run it. Building the graph does not execute the imported modules. -- Restore pnpm stores on PRs using setup-node's existing key and store path, - without publishing more PR-private copies. Non-PR setup-node caching and - native/TypeScript caches keep their existing behavior. A missing main store - still installs with the frozen lockfile. The mixed root/mobile store may miss - repeatedly because the existing main warmer only seeds the root lockfile. -- Batch only the PowerShell quota-fixture reservations within each test, using - the original generated scripts in fresh local scopes. Commands under test - retain separate processes, real file identities, and existing race assertions. - A traced local run confirms 44 PowerShell starts become 24, with all 21 cases - passing. Alternating after/before/after elapsed times were 50.00/59.28/37.00 - seconds on a shared macOS arm64 host; that variance does not justify a precise - percentage or hosted runner-time claim. Test budgets and worker counts are - unchanged. - -The reproducible pnpm-store comparison is -`ORCA_BACKGROUND_LAUNCH=1 node config/scripts/ci-pnpm-store-benchmark.mjs --samples=3`. -On macOS arm64 with BSD tar, three alternating fresh-store pairs eliminated a -median 332,746,995-byte archive per miss. Median install time was 17.33 seconds -before and 16.94 after; the removed archive step alone took 53.51 seconds. -Those local disk/CPU measurements exclude uploads and are not a prediction of -Linux or Windows hosted savings. Restore cost is common to both policies. - -## September 26 verification - -[PR #23053](https://github.com/stablyai/orca/pull/23053) was merged before its -latest full run finished. That run, -[36221874572](https://github.com/stablyai/orca/actions/runs/36221874572), ultimately -failed the mobile pending-frame precondition, just as the previous run had. -All eight unit shards passed, but the aggregate did not. An unchanged assertion -was not enough evidence to label the failure an unrelated flake. - -Main subsequently received the deterministic frame hold in PR #22635. The real -terminal refit now queues a frame that the recorder holds until disposal has -finished, then releases surviving work against a remounted terminal. This -follow-up adds a negative control: cancellation is disabled only during disposal, -and the same recorder must report a document-owned callback after disposal. -The normal case retains its pending-work and zero-leak assertions. Both cases -passed five fresh headless Chromium runs locally; the complete terminal-render -file passed all 13 tests. Browser dependencies were required, so these were real -render checks rather than skipped bundles. - -Unit model tests now import Monaco's editor API directly, preserving real models -and undo stacks without loading every language contribution. The registry bridge -requires only the editor and URI interfaces it actually uses. The full application -still imports its existing Monaco entry point; no production runtime behavior, -assertions, worker counts, timeouts or isolation settings changed. -A controlled local comparison (macOS arm64, Node 26.6.0, Vitest 4.1.11, -`--maxWorkers=1` only for this comparison) kept six files and all 32 tests. -Three warm original samples took 14.53/12.84/11.62 seconds; three editor-API -samples took 12.95/9.81/7.99 seconds, alternating back to the original imports -between measurements. Median elapsed time fell 12.84 to 9.81 seconds (23.6%); -median import time fell 10.73 to 7.94 seconds (26.0%). The initial cold original -sample, 20.29 seconds, is excluded. Shared-host variance remains; this is a -focused measurement, not a claim of a 23.6% improvement to the full unit suite. - -The default-branch warmer -[36221917346](https://github.com/stablyai/orca/actions/runs/36221917346) successfully -published native modules and TypeScript state. Fresh PRs -[#23101](https://github.com/stablyai/orca/pull/23101) and -[#23104](https://github.com/stablyai/orca/pull/23104) restored both on their first -runs. Scope inventories showed neither PR had a private copy; the TypeScript -key existed only on main. Native restore took 0.45/0.54 seconds; TypeScript restore -took 0.38/1.31 seconds, with compiler steps of 8/39 seconds versus the warmer's -80-second cold compiler step. PR #23104 used the prefix fallback after its base -advanced, confirming reuse across commits as well as PRs. - -The same warmer saved a 22.5 MB Git cache, but the exact key disappeared before -it was reused. Quota eviction is plausible, not proven: the usage API reported -17.48 GiB while a separate live 100-entry sample contained 15.15 GiB of pnpm -stores alone. These rapidly changing inventories are not atomic. The root-only -warmer does not seed the root-plus-mobile download-store key used by static -analysis, so both fresh PRs saved another roughly 350 MiB store. Controlling that -cache duplication is a remaining opportunity; hourly warming alone cannot -promise retention. Git's checksum-verified cold-build fallback remains required. - -The unit scheduling baseline now comes from all eight successful Node 24 shards in -[run 36294142683](https://github.com/stablyai/orca/actions/runs/36294142683). -All 9,847 measurements match current discovery exactly once; the previous baseline -had 165 unmeasured files and one deleted path. Applying the same measurements to -both assignments reduces the largest projected load from 1,043.811 to 919.695 -worker-seconds (11.9%). This is a scheduling projection, not an elapsed-time claim; -runner variation remains visible in the source run. See -[provenance and reproduction](../../config/scripts/ci-shard-timings.md). - -## Recording compilation reuse - -[Benchmark run 36295773765](https://github.com/stablyai/orca/actions/runs/36295773765) -compared the complete mobile suite on two hosted runners in opposite orders. -Compilation reuse reduced elapsed time from 439.002 to 138.308 seconds and from -560.926 to 200.751 seconds (64–68%). Each run preserved all 9,629 original test -verdicts and passed four additional cache regression tests. Compiled code is -bounded to 512 entries; exports, dependencies, and scenario state remain fresh. - -Splitting family recordings across four files took 145.165 and 217.226 seconds, -5–8% slower than compilation reuse alone, so the original suite structure stays. -All 787 goldens were regenerated from the unchanged pinned product tree; only -the recorder digest changed, with identical recording bodies and value pools. - -Desktop validation in [run 36295671576](https://github.com/stablyai/orca/actions/runs/36295671576) -passed all eight shards. The longest test step was 419 seconds versus 429 in the -source run; summed test time was 3,028 versus 3,031 seconds. Runner variation -prevents attributing that small elapsed-time difference solely to the weights. - -## September 25 follow-up - -The current queue is a bigger part of PR latency than setup. Successful full PR -[36212793136](https://github.com/stablyai/orca/actions/runs/36212793136) used -**74.6 aggregate runner-minutes**, including **55.5** for its eight unit shards. -Those jobs ran for 315–448 seconds but waited 229–1,076 seconds to start. The -three-second final `verify` job waited another 264 seconds. These are observed -job creation-to-start and start-to-completion intervals, not billing figures. - -This follow-up keeps the existing tests, isolation, eight unit shards, platform -coverage, and release behavior: - -- Make the reusable unit-test call and final aggregate respect cancellation. - Their old `always()` conditions kept superseded work alive despite workflow - cancellation. In [36215475069](https://github.com/stablyai/orca/actions/runs/36215475069), - a newer push cancelled ordinary jobs while all eight unit shards remained - queued and the replacement workflow remained pending. `!cancelled()` still - evaluates after failed/skipped dependencies, without resisting cancellation. - Two other superseded PR runs reproduced this: at 04:27 UTC on September 26, - [36216165253](https://github.com/stablyai/orca/actions/runs/36216165253) and - [36215543186](https://github.com/stablyai/orca/actions/runs/36215543186) still held - 11 runners, with two more obsolete jobs queued. They had consumed another - 71.3 runner-minutes after replacement pushes. After rechecking current PR heads - and replacement runs, both obsolete runs were force-cancelled; their replacements - left the blocked pending state. -- Combine root/README guards with change detection. A real sparse checkout kept - all 29,487 index entries while materializing only 12 files (192 KB). This - eliminates one runner allocation and checkout per PR. README link checks still - see tracked targets outside the working tree and run on docs-only PRs. -- Use free `ubuntu-slim` containers for small guard/aggregate/API jobs; trial - free `ubuntu-24.04-arm` for typechecking, which needs no native runtime. - Both share the account's standard concurrency limit. Different labels do - **not** grant extra concurrent jobs; hosted timings determine their value. -- Cancel superseded PR attempts in the Git termination, Pi owner, and Pi provider - runtime workflows, retaining independent manual runs. -- Fetch only complete HEAD ancestry for cloud secret scanning. The old checkout - fetched every branch and tag and took 55 seconds in - [36214130174](https://github.com/stablyai/orca/actions/runs/36214130174). - Real Git fixtures prove both merge parents and deleted historical contents - still produce the identical scanned patches. -- Remove the cloud lockfile from the eight unit shards' download-cache key; the - dedicated relay integration job still includes it. Record actual per-file - environment, setup, import, and test durations for the existing shard planner. - Shard 4 spent 535 worker-seconds importing and 357 executing tests; a uniform - per-file import estimate misses that cost. See [timing refresh](../../config/scripts/ci-shard-timings.md). -- Seed Node 24 native modules, the pinned Git compatibility binary, and TypeScript - state on the default branch when dependency/toolchain inputs change, with - scheduled recovery (originally hourly; now every six hours). One ten-minute-bounded hosted job - reuses existing cache keys and skips typechecking an already-cached commit. - New PRs can restore default-branch caches, while caches saved by another PR - are inaccessible. The audit found 80 entries totaling 10.67 GiB, including - 9.31 GiB of pnpm stores, but no main-branch Node 24 native or TypeScript state. - Seven PRs held separate copies of the same pnpm key (2.36 GB combined). - Git preparation now has one shared action with the unchanged cache key, - checksum, and build command. In - [36212101873](https://github.com/stablyai/orca/actions/runs/36212101873), a new PR - spent 40 seconds compiling the same Git 2.25.5 binary; main-branch warming - makes that cache available to new PRs too. - -The hosted trial also removes repeated work in mobile bundle checks. The -builder keeps private snapshots of the default real output for read-only checks -(15 identical builds become two), and haptics checks reuse each route closure -(32 builds become eight). Determinism, custom inputs, malformed routes, stale -outputs, and tampered manifests retain independent builds. A mutation regression -proves Buffer/manifest consumers cannot change another assertion's fixture. -The grant census reuses parsed references for unchanged file contents, still -reads source every time, and has a real-file edit invalidation regression. -The first hosted mobile lane used 400 seconds for its 448 tests; the grant -census alone took 259 seconds. Locally, the same one-worker invocation of the -three changed files fell from 243.07 seconds (101 passing tests) to 49.49 seconds -(103 passing tests). Hosted run -[36217238462](https://github.com/stablyai/orca/actions/runs/36217238462) retained -the same 43 files and passed all 450 tests: the original 448 plus two regressions. -Its test step fell from 400.26 to 205.17 seconds, and the complete job fell from -498 to 298 seconds (40% less runner time). Builder, haptics, and grant-census file -times fell from 73.70/88.94/258.58 seconds to 29.24/32.98/22.58 seconds. -The later default-worker verification retained the same 450-test coverage, but -its unchanged terminal-render test failed twice on the pre-existing timing -assertion that a frame must be pending at disposal (`expected 0 to be greater -than 0`). Its other 449 tests passed; no mobile source or assertion was relaxed. - -A four-worker experiment reduced aggregate unit job time from 3,386 to 3,133 -seconds (7.5%), but the repeat run exceeded the palette matcher's existing -180ms performance budget at 235ms. The override was removed; retain Vitest's -default worker count, isolation, timeouts, retries, and coverage. The first trial -also found an outdated hook-order snapshot after main added three layout-persistence -hooks. That snapshot was refreshed only after comparing the exact old and merged -hook sequences. Final verification is linked from -[PR #23053](https://github.com/stablyai/orca/pull/23053). Failed timing reports -never replace the checked-in baseline. - -No account settings, paid services, or runner entitlements changed. Standard -public-repository runners remain free. GitHub documents plan concurrency limits -of Free 20, Pro 40, Team 60, Enterprise 500, and permits support requests for -increases. The organization's actual entitlement was not exposed by the API. -See [runner specifications](https://docs.github.com/en/actions/reference/runners/github-hosted-runners) -and [concurrency limits](https://docs.github.com/en/actions/reference/limits). - -Hosted observations from [PR run 36215718607](https://github.com/stablyai/orca/actions/runs/36215718607): - -| Check | Earlier sample | Trial | Result | -| ----------------------------------------- | -------------: | -----------: | -------------------- | -| Detection plus repository guards | 54s, two jobs | 26s, one job | Passed | -| Typecheck (whole job) | 109s x64 | 81s ARM | Passed | -| Typecheck command, cold incremental state | 76s x64 | 51s ARM | Passed | -| Cloud secret scan (whole job) | 69s | 55s | Passed | -| Cloud checkout/history fetch | 55s | 40s | Identical scan scope | - -The [warmup trial](https://github.com/stablyai/orca/actions/runs/36215718295) -passed in 54 seconds. Its x64 compiler restored the ARM job's incremental state -and rechecked the same source in eight seconds. Unit shard 3 subsequently -restored its pnpm/native cache keys successfully. Actual sharing across different -PRs requires the producer to land on the default branch; this trial validates -commands and key compatibility, not a completed default-branch rollout. -The follow-up warmup built the Git binary in 40 seconds; the PR Git check restored -that exact key and passed in 62 seconds overall. Its TypeScript refresh took -seven seconds after restoring earlier incremental state. - -These are small observational samples from different revisions, not controlled -benchmarks. The cold typecheck log confirms an incremental-cache miss. Queue -changes must be separated from active duration and concurrent account traffic. -At the time of that trial, checked-in shard weights were unchanged; the September -26 refresh above now uses the complete successful unit reports. - -## September 5 audit - -Audit date: September 5, 2026. No paid capacity or provider configuration changed. - -## Measurements and changes - -Three recent successful PR runs used 54.6–64.9 aggregate runner minutes: -[33998366568](https://github.com/stablyai/orca/actions/runs/33998366568), -[33998220287](https://github.com/stablyai/orca/actions/runs/33998220287), and -[33998181502](https://github.com/stablyai/orca/actions/runs/33998181502). -These are sums of active job durations, excluding skipped jobs; they are not -billing minutes or queue time. This small sample is not a historical average. - -- Consolidate E2E routing into the existing code-path detector. The removed - detector occupied 20–22 seconds and required another runner allocation and - full-history checkout per nondraft code PR. The same routing commands remain, - including SSH and native IME selection; actual E2E results remain advisory. - A routing-script error now fails the required code-path detector. -- Use gzip for PR-only Debian/RPM artifacts. The two sampled Linux packaging - jobs took 8m10s and 8m19s overall; one spent 3m47s in electron-builder. Its - default Debian/RPM compression is xz. PR artifacts are inspected on the same - runner, so their download size offers no benefit. Keep all AppImage, Debian, - RPM, payload, launcher, and shutdown checks. Release compression is unchanged. - Hosted validation in [33999422341](https://github.com/stablyai/orca/actions/runs/33999422341) - reduced the package-build step to 2m13s and the full Linux job to 6m17s, with - all existing checks passing. This is a small observational sample. -- Cancel superseded Mobile Checks and Skill update round-trip PR runs. The - skill matrix has 13 jobs. Preserve non-cancelling main/merge-group skill runs, - with separate concurrency groups per event. -- Reuse the existing script-free root dependency action in Mobile Checks, - including the pnpm cache keyed by both root and mobile lockfiles. The root - install remains necessary because mobile types import root dependencies. - -The repository already has eight unit shards, path-scoped platform checks, -native caches, one shared E2E build, PR cancellation, incremental TypeScript -caching, and changed-spec E2E routing. Increasing shards would increase setup -work and simultaneous runner demand. Do not adjust the count without comparing -critical-path time and aggregate job time on the same commit. - -## Follow-up savings - -- Move the hourly main/release freshness lookup to a five-minute Ubuntu - preflight without a checkout. In unchanged run - [33986205749](https://github.com/stablyai/orca/actions/runs/33986205749), - Blacksmith macOS was occupied for 40 seconds, including a 30-second checkout, - before skipping. The new job-level gate avoids that Mac allocation. Actual - builds gain an Ubuntu scheduling hop; pin the Mac checkout and downstream - Windows identity to the SHA that the preflight checked. -- Avoid global `npm install -g node-gyp` for validated Linux Node-runtime cache - hits. Use the existing native-module load/provenance check before skipping; - misses, broken addons, and Electron jobs still install the rebuild toolchain. - The action file participates in cache keys, so this rollout creates fresh - native caches once. No measured warm-cache seconds are claimed yet. - -## Runner recommendations - -The repository is **public**, verified using the GitHub API. Standard -GitHub-hosted Linux, Windows, and macOS runners have free compute minutes for -public repositories. Queue pressure and third-party provider allowances still -matter; artifact storage and larger runners have separate billing rules. -See [GitHub Actions billing](https://docs.github.com/en/billing/concepts/product-billing/github-actions). - -1. Keep standard GitHub-hosted runners as the default. Ask GitHub Support for a - higher concurrent-job limit before paying for more capacity. The documented - standard limits depend on the account plan (Free: 20 total/5 macOS; Team: - 60/5; Enterprise: 500/50), and increases are subject to approval. The actual - account entitlement was not verified. See [limits](https://docs.github.com/en/actions/reference/limits). -2. Reserve existing Blacksmith allowance for macOS if that is the priority. - Blacksmith documents 3,000 free x64 2-vCPU-equivalent minutes per organization; - a 6-vCPU Mac minute consumes 20 equivalents, or 150 actual Mac minutes if - it uses the entire free pool. Cloud workflows also use Blacksmith Linux. - Moving Linux to hosted GitHub saves shared allowance, but does not necessarily - free Mac hardware capacity. Account-specific contracts and usage were not - inspected. See [Blacksmith runners](https://docs.blacksmith.sh/blacksmith-runners/overview). -3. Treat Ubicloud as an optional small Linux overflow trial. Its documented - $2.50 monthly credit buys 1,250 premium 2-vCPU minutes at $0.002/minute, or - 2,000 standard 2-vCPU minutes at $0.00125/minute. New accounts default to - premium and require a credit card. No enforceable hard spending cap was - verified, so changing runner labels cannot guarantee the no-spend constraint. - One PR's roughly 55–65 runner minutes also makes clear how small this pool - is relative to repository activity (hardware speeds differ). - See [pricing](https://ubicloud.com/docs/about/pricing) and - [setup](https://ubicloud.com/docs/github-actions-integration/quickstart). - -### A bounded Ubicloud candidate - -The Linux leg of `performance-contracts.yml` took 48 seconds in -[33994756657](https://github.com/stablyai/orca/actions/runs/33994756657). -Its daily schedule and 20-minute timeout make it a small candidate: 31 ordinary -scheduled attempts permit at most 620 job-runtime minutes, before runner -startup/cleanup billing. Actual timings on Ubicloud's 2-vCPU hardware still need -measurement; the GitHub timing is only a sizing reference. - -If enabled later, route only the first attempt of the scheduled Linux job to -Ubicloud; keep PRs, manual dispatches, reruns, and macOS/Windows on GitHub. This -avoids spending the allowance on unpredictable PR volume. Check other account -usage and available credit before enabling; a workflow timeout is not an -account-wide billing cap. On September 5, the organization's GitHub App -installation list contained Blacksmith but no Ubicloud installation, so this -follow-up leaves runner selection on GitHub rather than queueing work against -an unprovisioned label. - -## Machines that also run coding agents - -Do not register the credentialed host directly as a public-PR runner. A PR can -execute arbitrary build/test code, and a persistent host lets it access local -credentials or affect subsequent jobs. Docker alone is not adequate isolation -when it exposes the host home, Docker socket, SSH agent, or office network. - -A possible no-new-hardware experiment is a disposable VM per job, preferably on -a dedicated spare machine, with a just-in-time single-job runner, no shared -home/keychain/SSH agent or host mounts, restricted network access, and CPU/RAM -limits that leave room for coding agents. Destroy the VM after every job; -ephemeral runner registration by itself does not clean the machine. Start with -trusted branch/manual workloads and keep public fork PRs on hosted runners. -Provisioning and ongoing patching are real operational costs even when the -machine is already owned. See GitHub's -[self-hosted runner security guidance](https://docs.github.com/en/actions/security-for-github-actions/security-guides/security-hardening-for-github-actions). - -## Release waits - -The latest successful sampled Windows release used 13m59s of a 21m56s job in -signing wait/download steps. The same release held an Ubuntu job for 11m38s -polling the isolated Mac build. These are stronger occupancy opportunities than -small checkout savings, especially when approval takes hours. - -[Windows signing without occupying a runner](windows-signing-runner-time.md) -describes a staged, same-run design, required protected environments, and -rehearsal criteria. No callback integration or protected Windows signing -environments currently exist. An environment-gated design adds a GitHub -approval after each SignPath approval and changes the current automatic inner -signing timeout fallback; those are explicit release-policy decisions, so this -PR leaves production signing behavior unchanged. - -## Second audit and hosted trials - -- Cloud Verify ran 100 times in a sampled 39-hour window (84 PR and 16 push - runs). Move its four Ubuntu 22.04 jobs from Blacksmith to standard hosted - Ubuntu 22.04, preserving Postgres, secret scanning, build, tests, and Terraform - validation. Baseline [34001538145](https://github.com/stablyai/orca/actions/runs/34001538145) - used 64/72/26/19 seconds for security/test/build/Terraform respectively. - This conserves the shared provider allowance; hosted latency must be checked. -- Keep full tag history for the 13-job skill round-trip matrix, but fetch blobs - lazily. Only two historical SKILL.md files are materialized. Baseline - [33999994876](https://github.com/stablyai/orca/actions/runs/33999994876) - spent 42–84 seconds per checkout, about 14 aggregate runner minutes. A hosted - trial must verify historical blob fetches on all three operating systems. -- Use the existing Electron/native dependency cache for native IME CI. Keep - both deterministic boundary and real IBus tests. Add pnpm store caching to - terminal perf and release golden/evidence lanes; retain their raw installs - because manually selected older refs may not contain the shared action. -- Disable ZIP recompression only for already-compressed NSIS installers sent - to SignPath. Installer contents, release compression, and signing stay intact. -- Advance existing placement and startup deadlines with scoped fake timers in - three renderer test files. All 34 tests pass in 62 ms of local test execution, - versus 65.182 seconds in the sampled hosted baseline. Imports and transforms - still dominate invocation time; this is not a claim of equal PR wall savings. - -Eight unit shards already have balanced 260–296-second sample durations. -Reducing shards or removing test isolation lacks evidence of a net gain. Real -subprocess tests intentionally cover lifecycle behavior and retain real clocks. -The 14-way E2E split retains headroom after earlier 12-way timeouts. Lowering -coverage or schedule frequency is outside this efficiency pass. Cache complexity -for a seven-second docs install is unlikely to pay back. Release build reuse -across modes risks differing telemetry identities and native platform artifacts. - -Terminal Perf's baseline [33955846492](https://github.com/stablyai/orca/actions/runs/33955846492) -failed waiting 30 seconds for workspaceSessionReady in its shared-page fixture, -before measuring terminal performance. Compare hosted trials against that known -failure rather than attributing it to dependency cache changes. - -Hosted trials for the second audit: - -- [Cloud Verify 34002295216](https://github.com/stablyai/orca/actions/runs/34002295216) - passed all four jobs on standard hosted Ubuntu: security 57s, test 102s, build - 35s, Terraform 19s. The test lane is 30s slower than the Blacksmith sample; - retain this modest latency tradeoff to conserve shared allowance. -- [Skill matrix 34002295221](https://github.com/stablyai/orca/actions/runs/34002295221) - passed all 13 legs, including historical blob materialization. Checkout took - 18–20s on Linux, 39–45s on macOS, and 49–58s on Windows, versus the earlier - 42–84s range across platforms. These are observational samples. -- [Native IME 34002299594](https://github.com/stablyai/orca/actions/runs/34002299594) - passed both deterministic and real IBus checks. Shared dependency setup took - 29s, versus 35s for the old install/toolchain steps in the sampled baseline. -- Native-IME-only source/spec changes no longer allocate the reusable E2E - build, cache, and consumer jobs just to filter out the native spec. The - separate native workflow still runs; SSH-only and mixed spec lists still - allocate the reusable workflow. Routing contracts exercise these cases. -- [Hourly 34001816449](https://github.com/stablyai/orca/actions/runs/34001816449) - exercised the new five-second preflight and successfully published macOS. - The Windows follow-up failed in its unchanged input-vetting fetch because - remote refs differ only by case on its case-insensitive filesystem. The - requested SHA was correct; this does not validate an unchanged-main skip yet. - -Moving the daily Mac freshness check has lower expected value than hourly: -only one potential idle allocation per day, and active development usually -requires that build. Defer another release-graph change until skip frequency -justifies it. The substantive remaining release occupancy opportunity is the -separately documented asynchronous signing policy decision. - -## Persistent Vitest transform cache: rejected for now - -A local 96-file import-heavy sample with Vitest 4.1.11 took 8.31/8.53 seconds -without its filesystem module cache, 8.63/8.69 seconds cold, and 5.91/5.91 -seconds warm: about 30% faster warm. The cache held 3,770 modules and 83 MiB. -These timings exclude hosted cache transfer and do not establish a PR saving. - -Correctness probes found eight changes that incorrectly kept a test passing -against the old transformed import or compiler output: - -| Change after warming | Raw cache | Startup fingerprint | -| ----------------------------------------------------------------- | ---------- | ------------------- | -| Add preferred `value.js` beside previously resolved `value.ts` | False pass | Correctly fails | -| Retarget a source symlink while its old target still exists | False pass | Correctly fails | -| Change an inlined package's `exports` to another existing file | False pass | Correctly fails | -| Add a preferred extension in a generated source directory | False pass | Correctly fails | -| Change TypeScript's JSX factory in `tsconfig.json` | False pass | Correctly fails | -| Create the preferred file from setup after startup fingerprinting | False pass | False pass | -| Add a preferred file inside an external symlinked directory | False pass | False pass | -| Change an external file read by a transform plugin | False pass | False pass | - -The fingerprint included file names/types, symlink targets, package/config/ -TypeScript metadata contents, and effective alias/define options. Following -external symlink inventories and hashing declared transform inputs repaired the -last two rows, but did not repair files created after fingerprinting. All ten -cache-disabled changed-input controls failed correctly; initial and repeated -warm controls passed. Effective alias and simple define changes also invalidated -correctly without the added fingerprint. - -The [Vitest 4.1.11 documentation](https://github.com/vitest-dev/vitest/blob/v4.1.11/docs/config/experimental.md#known-issues) -documents incomplete plugin-input tracking. Its -[cache implementation](https://github.com/vitest-dev/vitest/blob/v4.1.11/packages/vitest/src/node/cache/fsModuleCache.ts) -hashes the module and selected configuration, but retains previously resolved -import URLs. [Upstream fix #11381](https://github.com/vitest-dev/vitest/pull/11381) -merged September 29 and revalidates those URLs. A disposable Vitest 5.0.3 probe, -which contains that fix, reproduced all eight false-pass categories: the old -target still exists, so checking its resolved URL misses a newly preferred file -or changed package export. Upgrading alone does not make reuse safe. - -To reproduce the simplest negative control outside the worktree: - -1. Create `value.ts` containing `export const value = 1`, and a test importing - `./value` and asserting `expect(value).toBe(1)`. Use an isolated config/cache, - one fork worker, `ORCA_BACKGROUND_LAUNCH=1`, and - `NODE_DISABLE_COMPILE_CACHE=1`. -2. Run the installed CLI with `--experimental.fsModuleCache=true` to warm it. - Add `value.js` containing `export const value = 2`; retain `value.ts` and the - unchanged test. The same cached invocation incorrectly passes. -3. Repeat with `--experimental.fsModuleCache=false`. The assertion correctly - fails. Vitest 5 uses `--fsModuleCache` for the equivalent controls. -4. For the startup-inventory control, use unchanged setup code that creates - `value.js` only when a runtime environment switch is enabled. Remove that - file before each invocation/fingerprint; warm with the switch off, then run - with it on. The cached importer still points at `value.ts`, while the fresh - module graph correctly fails. This is an additional persistent-cache error, - not a claim that normal in-process module reuse supports arbitrary mutation. - -Orca currently resolves only pinned Vitest/Vite built-in transform plugins. -Its setup files install runtime guards/shims and temporary user data, rather -than custom transforms. A narrower policy could cache only proven immutable -source/dependency inputs and leave tests, setup, virtual modules, external -fixtures, and unknown plugins cold. That requires a validated transitive input -boundary and mutation policy; hashing every source tree on each lookup would -also spend the gain. Until that policy and hosted transfer cost are measured, -the local warm result does not justify adding a persistent cache to CI. - -## Test fixture imports: reuse the existing narrow builder - -The pointer-drag test imported only `makeWorktree` from `store-test-helpers`, -which also loads the real store slices. Its existing identical export in -`worktrees-slice-test-fixtures` supplies the same defaults without that graph. -Changing this single import preserves the five tests, fork workers, -isolation, and disabled filesystem/Node compile caches. - -Three local interleaved before/after pairs took 2.066/2.047/2.031 seconds versus -0.351/0.388/0.353 seconds: the isolated median fell 82.7%, from 2.047 to 0.353 -seconds. Transformed modules fell from 1,086 to 15. This is an isolated test -result, not a whole-shard estimate: other tests need the store modules anyway. -A broader 20-file screening sample saved only 0.295 seconds at its median, -which does not justify splitting the fixture module across those consumers. - -Two further import-only reuses passed the same six-run controls. The kanban -lane test mocks its card component, so switching its builder import reduced -the isolated median from 2.202 to 0.495 seconds (77.5%) and transformed modules -from 1,082 to 13, with all six tests unchanged. The autosave fixture needs -the real editor slice, but not every store slice: the same import change across -its three consuming suites reduced the median from 3.027 to 1.301 seconds -(57.0%), with 1,107 to 320 modules and all 17 tests unchanged. These -results also measure isolated file groups; they are not additive shard savings. -The remaining inspected builder-only imports already load the full store as -their subject, or use builders whose defaults differ from existing exports. - -## Vitest threads: retain forks after the scoped pilot - -Three local interleaved comparisons kept four workers, `isolate: true`, both -persistent caches disabled, and the same test assertions/module graph. The -94-file happy-dom renderer cohort passed all 570 tests: forks took -17.480/17.382/17.657 seconds and threads 15.048/14.890/15.013 seconds, a 14.1% -median reduction. A 23-file shared JavaScript cohort passed all 248 tests, -with its median falling from 1.435 to 1.274 seconds (11.2%). No main-process -module or native addon loaded; guards reject native loading, `chdir`, and -process signals. These Mac/Node 24 timings motivated the hosted comparison. - -The [pinned Vitest pool documentation](https://github.com/vitest-dev/vitest/blob/v4.1.11/docs/config/pool.md) -defaults to forks and documents thread limitations around process APIs and -native libraries. [Node's worker documentation](https://nodejs.org/docs/latest-v24.x/api/worker_threads.html#new-workerfilename-options) -also excludes V8 flags from worker `execArgv`. Orca's `--expose-gc` worker flag -fails with `ERR_WORKER_INVALID_EXEC_ARGV` under threads. The experiment starts -both parent processes with that flag, retains it on fork workers, and removes -it only from thread worker arguments; GC availability is checked in every -test environment. Process, native, lifecycle, and GC-retention tests stay out -of this comparison. - -The [hosted pilot](https://github.com/stablyai/orca/actions/runs/36855833033) -passed on Linux ARM64, four CPUs, Node 24.21.0, and Ubuntu image -`20260927.135.1`. Three alternating pairs preserved source hashes and complete -module graphs, with no main-process module or native addon loaded: - -| Audited cohort | Forks, seconds | Threads, seconds | Median saving | -| -------------------------------------- | ------------------------ | ------------------------ | --------------------- | -| 94 renderer files / 570 tests | 48.591 / 48.144 / 47.299 | 42.486 / 42.402 / 42.221 | 5.742 seconds (11.9%) | -| 23 shared JavaScript files / 248 tests | 3.386 / 3.330 / 3.424 | 3.203 / 3.237 / 3.278 | 0.149 seconds (4.4%) | - -Each timed group also included two isolation sentinels: totals were 96 files / -572 tests and 25 files / 250 tests, respectively, with no skips. -Their graph hashes matched in every pair (4,015 and 251 modules). Separate -single-worker positive controls passed both sentinels in each pool. With -isolation disabled, the second sentinel correctly failed on leaked state. -Missing GC failed setup, and native loading, `chdir`, and process-signal probes -failed at the guard in both pools. Those expected failures did not pass silently. - -The fixed 117-file sample accounts for about 1.5% of the baseline's aggregate -module time; its isolated gains are not a whole-suite estimate. Maintaining -that exact file list for this benefit is not justified. A broader route needs -a safe eligibility policy, Node 26/Windows evidence, and a mixed full-shard -comparison: separate Vitest projects can repeat shared transforms and erase -the pool-startup saving. A renderer path alone does not prove that future -imports avoid process or native behavior. Production retains forks, and the -temporary workflow, driver, and cohort list were removed after measurement. - -## Oxlint scan consolidation: rejected - -The [hosted comparison](https://github.com/stablyai/orca/actions/runs/36855833063) -used one Linux ARM64/four-CPU runner, Node 24.21.0, and three alternating -baseline/candidate pairs. The baseline kept root lint and anti-slop in parallel, -then native and type-aware audits in parallel. The candidate merged the first -three scans and ran the unchanged type-aware audit alongside them. - -| Pair/order | Baseline stage | Candidate stage | Change | -| ------------------ | -------------- | --------------- | ------ | -| 1: baseline first | 49.409 s | 53.197 s | +7.7% | -| 2: candidate first | 50.249 s | 60.297 s | +20.0% | -| 3: baseline first | 50.040 s | 56.672 s | +13.3% | -| Median | 50.040 s | 56.672 s | +13.3% | - -These complete stage timings include anti-slop synchronization in both variants -and candidate configuration generation. Candidate preparation took only -0.141–0.255 seconds. The unchanged type-aware scan took 16.386–17.279 seconds -in the baseline wave, versus 33.561–55.736 seconds beside the merged scan; -these timings are consistent with contention on the four-CPU runner. - -The corrected local Mac/16-CPU comparison had reduced the median from 16.676 -to 14.381 seconds (13.8%). Both comparisons limited each Oxlint invocation to -four threads and used identical source configuration hashes. The hosted result -shows why the local gain did not justify adoption on the actual CI runner. - -Coverage controls passed: the merged scan matched the exact 28,621-file union -with no missing or extra files. Thirty-one fault fixture files produced the -exact 16-diagnostic union, including all seven active root JavaScript rules. -Eighteen focused controls preserved nested-mobile exemptions, type-aware -exclusions, and exit behavior. The native audit's warnings still failed its -original `--deny-warnings` gate and became errors in the merged scan; root -warnings remained non-fatal. Every full-repository scan passed cleanly. - -Keep the existing production waves. The temporary workflow and 601-line -benchmark driver were removed after recording this rejected result. - -## Native cache ownership: retain the extraction - -Native restoration, toolchain recovery, and preparation now belong to -`.github/actions/prepare-native-runtime/action.yml`. The installer forwards -its requested key and three build paths; Windows packaging saves the Node -build before calling the same action for Electron. Existing native load, -patched-build, Windows job-ownership, registry, and process-table probes remain -unchanged on restored consumers. Exact keys still separate OS/image or Linux -container libc, architecture, runtime, resolved Node version, and actual pnpm -version, without partial-key restoration. - -The source hash covers the dedicated action, `pnpm-lock.yaml`, -`pnpm-workspace.yaml`, `.npmrc`, `.pnpmfile.cjs`, both native dependency patches, -and these complete build/probe inputs: - -- `config/scripts/ensure-native-runtime.mjs`, `rebuild-native-deps.mjs`, - `node-pty-job-ownership.cjs`, `windows-pe-machine.cjs`, - `windows-process-tree-gyp-rebuild.mjs`, and - `windows-process-tree-creation-time.cjs`; -- `config/scripts/install-electron-package-binary.mjs`, - `electron-platform-path.mjs`, `zip-extractor-command.mjs`, - `shared-electron-dist-cache.mjs`, `space-sharing-copy.mjs`, and - `src/shared/zip-extractor-command.ts`; -- `native/windows-registry/src/addon.cc`, `binding.gyp`, `package.json`, and - `index.js`. - -The patches are `config/patches/node-pty@1.1.0.patch` and -`config/patches/@vscode__windows-process-tree@0.8.0.patch`. Root app version and -script metadata are excluded; installed package versions remain owned by the -full lockfile, and the external node-gyp pin belongs to the native action. -A negative control changing only the installer's -pnpm verification condition preserves the native key and paths. Every declared -native input mutation changes the key, and main warming watches those inputs. - -This policy creates one cold namespace. The bounded 50-head main sample has -49 adjacent transitions and three native-key changes under both the old and -expanded policies: the added node-pty helper export still invalidates #24448. -There is no measured historical net saving. - -The [cold warming run](https://github.com/stablyai/orca/actions/runs/36945655208/attempts/1) -published all four exact Node keys, and its -[warm rerun](https://github.com/stablyai/orca/actions/runs/36945655208/attempts/2) -restored them on fresh runners with the same frozen inputs, Node 24.21.0, and -pnpm 12.0.0. All five jobs passed in both attempts. These times cover the entire -native action: runtime validation, key resolution, restore, any toolchain -recovery, and the unchanged native preparation probes. - -| Native Node lane | Cold action | Warm action | Cold post-job save | -| ---------------- | ----------- | ----------- | ------------------ | -| Linux x64 | 18.288 s | 0.993 s | 0.387 s | -| Linux ARM64 | 11.540 s | 1.062 s | 1.113 s | -| Windows x64 | 104.482 s | 1.381 s | 2.424 s | -| Windows ARM64 | 256.662 s | 3.832 s | 1.243 s | - -Cold Windows jobs rebuilt all three native addons. Warm jobs loaded and probed -the restored builds; Linux also ran the existing check-only probe before -skipping the external node-gyp installation. Warm post-job steps recognized -their primary keys and did not save again. An earlier trial exposed unavailable -nested composite outputs during post-job saving; both cache variants now use -the same literal path inventory as the requested output, and the fixed cold -jobs published their caches without missing-path warnings. - -Both [PR package jobs](https://github.com/stablyai/orca/actions/runs/36945659474) -passed. Windows packaging consumed the same-run Node seed in 1.374 seconds -before its Node tests and Electron transition. Its Electron cache initially -missed while the modules were already healthy, so that stage does not establish -an avoided compilation. Linux's Electron cache was also published, and all 19 -bundled native binaries passed the existing glibc floor check. - -The [first six-platform headless run](https://github.com/stablyai/orca/actions/runs/36945658897) -ran every persistence lane: five passed, while Mac Intel failed waiting for a -cancel-test worker's ready file before its 500 ms timeout. That lane deliberately -uses `native-runtime: none`; its separate slot build and smoke passed. The failure -blocked the five downstream Linux glibc/musl qualifications, so this run does not -establish complete headless qualification. All six persistence lanes, Node 18 -handoffs, and Linux floor/musl gates remain; final qualification is tracked in -the [PR's latest checks](https://github.com/stablyai/orca/pull/24476/checks). - -These are single cold/warm observations, not paired medians or a measured -whole-workflow saving. They demonstrate usable exact-key reuse after publication; -future savings depend on cache availability and unchanged native inputs. The -trial seeds belong to this PR's merge ref. Other PRs require a main-branch seed -after merging this new namespace; the existing main-push and scheduled warming -jobs provide that seed. - -## Separate mobile install verification: retain the current policy - -Three local paired pnpm 12 mobile installs reduced the median from 17.155 to -15.871 seconds, a 1.284-second difference before cache transfer and postinstall -scripts. That narrow margin does not establish a net hosted saving, so the -separate mobile verification record was not adopted. - -## Unit shard weights: retain the current allocation - -The latest five shard wall times were 526/495/503/508/510 seconds. Reweighting -projected roughly a 4% reduction in the slowest shard without reducing total -CPU work; the evidence across runs was weak. That estimate does not justify -changing allocation, so the current weights remain. - -## Serializer oracle allocations: retain the change - -The serializer round-trip oracle now reloads one xterm cell per buffer traversal -and writes flag digits directly, avoiding fresh cell objects and flag arrays for -every comparison. Independent replay terminals, cell descriptors, transcript -fixtures, resize schedules, ConPTY modes and seeds remain unchanged. - -The [hosted comparison](https://github.com/stablyai/orca/actions/runs/36944080887) -used one Ubuntu 24.04 ARM64 runner, image 20260927.135.1, Node 24.21.0 and one -isolated fork. The baseline formatter was frozen from f69052e. Byte-parity capture -ran separately; these three alternating pairs had no payload instrumentation. -Times cover the complete Vitest invocation, including startup and shutdown. - -| Pair/order | Baseline | Candidate | Change | -| ------------------ | -------- | --------- | ------ | -| 1: baseline first | 70.631 s | 61.661 s | -12.7% | -| 2: candidate first | 70.619 s | 62.284 s | -11.8% | -| 3: baseline first | 70.750 s | 62.091 s | -12.2% | -| Median | 70.631 s | 62.091 s | -12.1% | - -Median test-body time fell from 69.519 to 60.953 seconds. All eight full-cohort -invocations preserved the same 89 passes and two existing conditional skips -across three files. Separate baseline/candidate captures produced identical -95,017,559-byte payloads for all 1,435 scenarios and 7,649 checkpoints, with zero -source crashes; both SHA256 hashes matched the local captures. - -Seven focused controls compare against the original allocating oracle, including -all 128 text-flag combinations, styled blanks, wide cells, cell reuse and immutable -snapshots. Five deliberate faults were detected: stale cell contents, a missing -bold flag, changed empty-cell policy, removed scratch reuse and a source-parser -crash. The last control also proved that crash returns enter the capture. - -This measures the three-file oracle cohort. Whole-shard timings include other -test bodies, imports and transforms, so a whole-suite saving needs separate -measurement. - -## Cache warming: let scheduled ticks wait for active work - -The hourly warmer previously cancelled an active warmer, even when both used -the same source. On October 2, the [merge-triggered run](https://github.com/stablyai/orca/actions/runs/36965832780) -at 8ff6296 was interrupted by the [hourly run](https://github.com/stablyai/orca/actions/runs/36966367896) -at the same commit. The Windows ARM dependency installation had run for 356 -seconds before cancellation; its native verification was skipped. The other -four lanes had already succeeded. - -Scheduled events now wait in the existing concurrency group. Push, PR and manual -events still replace active work. This keeps one active workflow and the default -single pending slot, using GitHub's documented -[conditional cancellation](https://docs.github.com/en/actions/how-tos/write-workflows/choose-when-workflows-run/control-workflow-concurrency). -All cache probes, platforms and publication rules remain. - -This avoids the observed discarded installation. It does not remove the next -scheduled run or its repeated successful lanes, and pending replacement still -applies regardless of the cancellation expression. The bounded 20-run sample -contains this collision; it does not establish a recurring or whole-CI saving. - -## Cache warming: six-hour recovery interval - -Scheduled warming now runs at 00:41, 06:41, 12:41 and 18:41 UTC instead of hourly. -Main pushes that change cache inputs still seed immediately, and manual dispatch -remains available. All five jobs, probes, keys and publication rules remain. -This removes 20 scheduled workflows and 100 scheduled job starts per day (83%). - -Four consecutive October 2 scheduled runs used the same source. The -[18:50 UTC run](https://github.com/stablyai/orca/actions/runs/37050194510) used 474 -aggregate runner-seconds across five jobs, including 242 seconds on Windows ARM. -That job restored exact package, verification and native caches; package-store -restore alone took about 70 seconds. Repeating that observed duration twenty -fewer times would avoid about 158 runner-minutes daily, but this one-run estimate -is not a billing forecast or measured post-rollout saving. - -The longer interval can delay background repair after eviction or runner-image -changes. Existing consumers retain cold-cache installation/build fallback, and -normal cache reads update last access. Storage was near the repository limit -when audited, so retention and unchanged hit rates are not guaranteed. Observe -misses before reducing the recovery frequency further. - -## Daemon shutdown fixture: remove build tools after compilation - -The fixture now removes compiler and Python build dependencies, plus npm and -node-gyp caches, in the same Docker layer that installs node-pty. It restores the base image's manual -package marks, keeps procps and util-linux, and retains the packages owning the -shared libraries used by Node and the actually loaded PTY addon. This extends -the [official Node image's package-ownership approach](https://github.com/nodejs/docker-node/blob/main/22/bookworm-slim/Dockerfile) -to native addons. Dependency checks and a real PTY spawn fail the build if cleanup -breaks the runtime. - -The [hosted comparison](https://github.com/stablyai/orca/actions/runs/36969120786) -used Ubuntu x64, Docker 28.0.4 and the same resolved Node 22.23.3 base digest for -both images. All 3,418 common entries under `/usr/local` retained their bytes, -modes and symlink targets; all seven resolved runtime libraries matched. Only -two directory-only Python paths disappeared. The 52 removed Debian packages -were build dependencies; retained package versions stayed identical. - -| Image payload | Baseline | Candidate | Reduction | -| ------------------ | ------------- | ------------- | --------- | -| Docker archive | 714,643,456 B | 329,967,616 B | 53.8% | -| Compressed archive | 205,819,558 B | 91,888,840 B | 55.4% | - -Each timed arm started a separate Docker daemon with an empty image store, -decompressed the archive, loaded it, rebuilt from its inline cache and ran the -unchanged descendant/canary shutdown check. Every provisioning layer was cached. - -| Pair/order | Baseline | Candidate | Saving | -| ------------------ | -------- | --------- | ------- | -| 1: baseline first | 16.320 s | 10.198 s | 6.122 s | -| 2: candidate first | 16.389 s | 10.141 s | 6.248 s | -| 3: baseline first | 16.351 s | 10.223 s | 6.129 s | -| Median | 16.351 s | 10.198 s | 6.153 s | - -Median decompression fell from 1.059 to 0.493 seconds and loading from 10.723 to -5.125 seconds. Cached rebuild and shutdown times stayed close. Both seed images -and all six restored consumers passed the original shutdown/canary assertions; -both deliberate no-op disposal controls failed with the descendant still live. -Byte and retained-directory mode faults also failed the inventory comparator. -All six owned daemons stopped gracefully without a forced kill. - -These private daemons used separate classic overlay2 stores and the untouched -host containerd service. Production storage settings were not captured, the -filesystem cache was not flushed, and network transfer is excluded. Production -restores overlap dependency installation, so this 37.6% fixture-sequence saving -does not establish a six-second PR wall-time improvement. Single cold image -builds took 19.313 and 22.623 seconds. The Dockerfile change creates one new -fixture key; the existing main warmer seeds it after merge. - -## WebRTC egress fixture: avoid GPU initialization for the data channel - -The Linux-only probe disables hardware acceleration before Electron readiness. -It still creates two independent processes/profiles, a real data channel, offer -and local description, and checks the exact proxy and UDP policy. The three-second -host observation, 500 ms drain and 20/30/45-second deadlines remain unchanged. - -Two fresh Ubuntu x64 runners compared three alternating pairs each, with identical -phase instrumentation. The [first trial](https://github.com/stablyai/orca/actions/runs/36967511524) -started with the baseline; the [second trial](https://github.com/stablyai/orca/actions/runs/36969120786) -started with the candidate. Their first baseline peer constructors took 4.059 -and 2.546 seconds and logged the GPU command-buffer error seen in an earlier -package timeout. In the second trial, that baseline delay followed the cold -candidate's 2.3 ms constructor. Every candidate constructor took 2.0–2.6 ms. - -Typical process time stayed near 8.3 seconds: baseline/candidate medians were -8.288/8.273 seconds in the first trial and 8.298/8.380 in the reverse trial. -The evidence supports removing an avoidable startup delay, without a measured -typical throughput gain or an estimate of future timeout frequency. - -All 12 full case invocations preserved actual unprotected UDP and zero protected -UDP. Both trials rejected seven faults: missing policy, a packet at 2.9 seconds, -a packet during the drain, a missing peer factory or local description, a broken -packet counter and a hung renderer. The original assertions and deadlines caught -each fault. The normal package gate runs the uninstrumented fixture. - -## Remote resync fixture: keep coalesced frames in one decoder pass - -The first [full PR run](https://github.com/stablyai/orca/actions/runs/36971375340) -passed both package jobs but failed one remote-workspace ordering assertion: -it observed revisions `[2, 3]` where the fixture expected `[3]`. The decoder can -yield between two frames after its four-millisecond work budget. Under slow -scheduling, the first response's promise can publish revision 2 before the -second frame's revision 3 notification is dispatched. - -The fixture now holds its delivery clock at the actual timestamp from multiplexer -construction through the first coalesced-buffer delivery, following the existing -decoder test pattern. It restores the clock before asynchronous assertions and -again before disposal. All four source/order cases retain their exact cache, -publication, client-identity and follow-up-read assertions. Production decoding, -its fairness budget and remote messages are unchanged. - -Normal focused runs passed all 18 tests. Advancing the clock by four milliseconds -per call reproduced `[2, 3]` in both original response-first cases; the fixed -fixture passed all 18 under the same control. Removing the freeze reproduced both -failures. Bypassing the production read-safety guard still caused `[3, 2]` rollback -in all four ordering cases and eight failing tests overall. All controls preserved -the same 18 test identities. This corrects a reproducible fixture assumption; -one CI failure does not establish a failure-rate reduction. - -## Windows server cache metadata: retain the current key - -The bounded 50-head main sample ending at 8ff6296 contained no root package -metadata changes. Removing app-version metadata from the Windows server cache -key would not improve reuse in that sample, so the key remains unchanged. - -## Windows ARM SSH: prepare the inbox capability during independent builds - -The ARM inbox lane starts guarded Windows capability preparation after the pure -provisioning self-test and waits for it before any private SSH server or host cell -runs. Dependency installation and the unchanged native artifacts can run during -that preparation. Preview and x64 lanes keep their existing serial provisioning; -the registered background step completes without mutation in those lanes. - -The preparation and the foreground provider use the same installer and isolation -guards. The receipt must match the source, run, attempt, runner, image and native -architecture. The foreground provider still reads the installed capability and -verifies every native binary and Microsoft signature. Account ownership, ACLs, -DefaultShell, private service identity, host cells and cleanup remain independent -checks. A background failure propagates through the unconditional native wait. - -Two full four-lane pairs used frozen source refs and the same dependency and -native-install policy. The [first baseline](https://github.com/stablyai/orca/actions/runs/36986929163) -ran before the [first candidate](https://github.com/stablyai/orca/actions/runs/36986970976); -the [second candidate](https://github.com/stablyai/orca/actions/runs/36991232037) -was dispatched before the [second baseline](https://github.com/stablyai/orca/actions/runs/36991234729). -Runner image versions matched within each platform in both pairs. - -| Active job, seconds | First baseline | First candidate | Second baseline | Second candidate | -| ------------------- | -------------: | --------------: | --------------: | ---------------: | -| ARM inbox | 2,403 | 1,644 | 2,353 | 1,667 | -| ARM preview | 1,002 | 935 | 886 | 872 | -| x64 inbox | 636 | 732 | 616 | 620 | -| x64 preview | 562 | 561 | 623 | 566 | - -The ARM inbox observations improved by 759 and 686 seconds. Baseline dependency -installation and artifact builds consumed 501 and 498 seconds before capability -installation could start. Candidate capability installation ran during that -work, but also took about 261 and 232 seconds less than the baseline. Candidate -dependency installation was slower, particularly in the second pair. These -observations support overlap on ARM; they do not establish a guaranteed 11–13 -minute saving, a reduction in queue time, or the cause of installer variability. -The x64 lane showed no repeatable gain, so it keeps serial preparation. - -All 16 actual Windows providers and 48 host-cell verdicts passed across the two -pairs. Receipts verify native machine identity, private service absence, owned -process exit, account removal and key removal. Loaded profile disposition remains -separate from those required cleanup checks. Hosted execution also verified the -native background/wait syntax; older actionlint versions do not recognize it. - -### Overlap the private profile observation budgets - -After service deletion and owned process exit, profile cleanup polls each owned -SID with its own full 30-second monotonic budget. Independent budgets now run -together. Every deletion follows a fresh targeted read; loaded profiles remain -for disposable VM destruction. Service identity, PID ownership, process exit, -account removal and key removal still fail the complete provider on error. - -The maintained diagnostics self-test executes the actual cleanup try/catch with -scoped Windows API and clock controls. Eight positive cases cover full windows, -late unload, reload, query overhead, mixed states and missing SIDs; ten specific -failure cases cover foreign profiles and the required cleanup gates. Disposable -shortened-deadline and stale-snapshot mutations fail those controls. A separate -mocked real-clock observation took 30.179 seconds for three loaded profiles, -compared with about 90 seconds for serial full budgets. This measures polling, -not an actual Windows provider or the entire job. - -The third profile no longer gains incidental extra time while earlier profiles -consume their budgets. A profile unloading at 45 seconds may therefore remain -where serial cleanup removed it. This uses the existing disposable-VM fallback; -it does not remove a loaded profile or relax mandatory account/key cleanup. -Hosted qualification of the combined workflow remains pending. - -## Coordinator mail tests: advance observation windows without removing them - -Six cases advance their original six 1,500 ms and ten 100 ms observation windows -with a scoped clock. Real filesystem, SQLite, journal, RPC and runtime work still -finishes asynchronously. The original journal-read gate and all counter and -operation assertions remain. Cancellation during delayed startup and the -Date-only age case retain real timers. Teardown stops the host and closes the -database before advancing the known 2,000 ms orphan repair, then asserts no fake -timers remain and restores the clock in `finally`. - -Two opposite-order local pairs passed the same 23 cases and unchanged source -hashes. Selected-case totals fell from 13.674 to 3.318 seconds and from 13.276 to -6.323 seconds. Whole-file test totals fell from 24.845 to 10.829 seconds and from -21.500 to 19.410 seconds. Process wall times were 41.488/37.810 seconds and -42.140/78.450 seconds; the reverse candidate spent 56.31 seconds importing under -unrelated local load. Local process-wall savings were inconclusive. - -The later [hosted x64 and ARM comparison](https://github.com/stablyai/orca/actions/runs/37001891871) -passed the same 23 cases in `structured-chat-coordinator-mail.test.ts` in both -orders on each architecture, with frozen case and policy hashes. Median full-file -wall time was 33.506 → 23.149 seconds on x64 and 33.438 → 22.504 seconds on ARM. -Median test-body totals were 19.329 → 9.343 and 19.695 → 9.103 seconds, respectively. -These measurements qualify this file; they do not measure whole-PR time. - -Injected extra deliveries at 1,499 ms and 99 ms still fail the original assertions -in both clock modes. The latter candidate fails the unchanged journal-read gate -with the same extra provider start. A separate control confirms the orphan repair -actually executes against the closed database and leaves no fake timers. The -change retains all 121 original expectation sites and adds one teardown check; -it does not shorten the runtime's observation interval or claim a whole-PR gain. - -## Stub child shutdown clocks: Codex and Claude - -[Merged Codex change #24893](https://github.com/stablyai/orca/pull/24893) scopes timeout -clocks to two synthetic-child cases in `codex-app-server-connection.test.ts`. -The full platform graceful deadline and 1,000 ms forced wait remain; the test -waits for the actual stub SIGKILL before advancing the forced window. Streams, -process-table reads, Date and immediate callbacks remain real. Fault controls -still detect late exit, missing EPIPE, missing exit proof and unwanted notification. - -The [hosted ARM comparison](https://github.com/stablyai/orca/actions/runs/37074124526) -passed the same 32 full-file cases in baseline/candidate and candidate/baseline -order. File wall times were 13.671 / 13.674 seconds originally and 1.658 / 1.649 -seconds with scoped clocks. Installer time is excluded; generated caches remain -across the disclosed order. Real-child coverage and production shutdown code remain. - -[Merged Claude change #24897](https://github.com/stablyai/orca/pull/24897) changes only -two synthetic-child cases in `claude-agent-sdk-exit-proof.test.ts`. Both full -33-case runs passed, including the unchanged five real-child cases. In one local -macOS pair, the two bodies took 2,503 / 1,502 ms originally and 1.37 / 0.39 ms with -scoped clocks. They cross a real immediate callback before advancing the complete -1,500 ms graceful and 1,000 ms forced windows, restore timers in `finally`, and -retain the original false exit verdicts. Fault controls detect either deadline -shortened by one millisecond, an unproved true verdict and a leftover timer. -[Normal PR CI](https://github.com/stablyai/orca/actions/runs/37075819218) passed; -these local body measurements do not establish hosted or whole-PR time savings. - -## Sequential static analysis and typecheck: retain separate jobs - -The earlier recommendation below is superseded by [October 4 shared PR preflight capacity](#october-4-shared-pr-preflight-capacity). - -A four-trial hosted screen kept the slim router unchanged and compared the two -independent ARM jobs with one ARM job running their unchanged checks sequentially. -The [compiler/planner census](https://github.com/stablyai/orca/actions/runs/37069472888) -matched all compiler inputs and the full 10,477-file unit inventory in separate, -shared root-only and shared mixed-install states. The [safety qualification](https://github.com/stablyai/orca/actions/runs/37075043747) -verified native joins after compiler failure and a real late action-post failure; -all four guarded downstream sentinels skipped and the audit passed. - -| Trial | Mode | Active ARM seconds | Router finish to heavy finish, seconds | -| -------------------------------------------------------------- | -------- | -----------------: | -------------------------------------: | -| [1](https://github.com/stablyai/orca/actions/runs/37075574191) | Separate | 157 | 124 | -| [2](https://github.com/stablyai/orca/actions/runs/37075887937) | Combined | 134 | 150 | -| [3](https://github.com/stablyai/orca/actions/runs/37076222202) | Combined | 129 | 134 | -| [4](https://github.com/stablyai/orca/actions/runs/37076786281) | Separate | 159 | 194 | - -Both pairs saved active ARM time: 23 and 30 seconds, or 14.6% and 18.9%, with one -heavy admission instead of two. The active critical path was 15 and 8 seconds -longer. Downstream eligibility changed by +26 and −60 seconds; observed ready-to-start -delay differences of +11 and −68 seconds explain that reversal. Created-to-start -delay is recorded separately and does not establish a quota or queue cause. - -Retain separate jobs for now. This screen shows a capacity saving, with a longer -active critical path and no repeatable latency gain. All trials used frozen -`cc73c8e1a72b0e9ee9c29e57458ce307f5f019c2` source, manual workflow dispatches, -Node 24.21.0 and the same four exact primary cache hits. Main's later -[Linux PR root-store policy change #24896](https://github.com/stablyai/orca/pull/24896) -is outside this screen. The trial ran actual heavy checks and proved unit and both -package eligibility, without launching those downstream matrices or measuring a -whole-PR speedup. - -## Linux headless runtime build overlap - -The historical pinned Bun artifact now builds in a native background step while -current native preparation and Node bundling run in the foreground. An -unconditional join precedes the unchanged artifact and cross-runtime tests. Bun -setup stays Linux-only; other platforms register and join a successful no-op. -The producer publishes step outputs consumed only by those tests. The existing -selector, native floors, template builders and cache policies remain. - -A [hosted alternating comparison](https://github.com/stablyai/orca/actions/runs/37072923774) -ran four serial/overlap arms on each of two Linux VMs: - -| Architecture | Serial preparation, seconds | Overlapped preparation, seconds | -| ------------ | --------------------------: | ------------------------------: | -| x64 | 20.831 / 19.576 | 11.149 / 10.914 | -| ARM64 | 15.155 / 14.396 | 8.566 / 8.553 | - -Every arm passed the same 961 cases across 92 files: 930 passed and 31 skipped. -Both cross-runtime persistence cases passed. The two live daemon-handover cases -kept their existing protocol-version skips. All four x64 arms passed actual Node -18 loading and pinned-runtime handoff. Installed/source inputs and artifact -inventories matched; each normal owned-process ledger was clean before cleanup. -Common native compiler warmup preceded timing and retained its generated Python -caches in the strict installed ledger. These are preparation savings of 5.8–9.7 -seconds, excluding setup, cold installs, runner start delays and whole-PR time. - -Actual [Bun failure](https://github.com/stablyai/orca/actions/runs/37078015921) and -[Node failure](https://github.com/stablyai/orca/actions/runs/37078021568) controls -qualified genuine compiler errors with fresh live opposite builders, native joins, -skipped consumers, restored inputs and verified exits. A [normal cancellation -control](https://github.com/stablyai/orca/actions/runs/37079655167) received SIGINT -while the actual Bun builder was freshly live; both builders and the detached -owned child had simultaneous earlier readiness. All three native joins had terminal dispositions of cancelled, success and -cancelled, and every consumer skipped. The temporary observer retired its owned processes; -the collector independently verified their absence and unchanged inputs. This -proves signal delivery and observer-owned retirement, without establishing -runner-only descendant cleanup at the join. The unchanged historical builder -starts finite build/smoke work, and its children retain GitHub's normal orphan -tracking marker. - -Earlier cancellation trials remain excluded from live-build qualification: one -collector stopped its observer before signal routing, and the corrected trial -received the signal after both builders finished. The qualifying trial requested -normal cancellation earlier in the same preparation sequence to account for -observed delivery delay; no workload, wait or proof predicate was shortened. - -## October 3 Terminal Perf dependency preparation - -The daily/manual Terminal Perf workflow still installed current dependencies through -raw lifecycle scripts and a global node-gyp installation. Its historical `ref` -input also accepts revisions that lack the shared installer, so replacing that -path unconditionally would break older runs. The current-profile path now uses -the existing shared installer with explicit Electron preparation and archive -caching. A guard requires GitHub-hosted Linux x64, Node 24/pnpm 12.8.1, the -native-only root postinstall and the needed local action inputs/files. Other -profiles and historical revisions keep their original frozen install. - -The [hosted comparison](https://github.com/stablyai/orca/actions/runs/37101695800) -ran both preparation paths in each of two Linux x64 jobs, reversing their order. -Legacy/shared preparation took 25.164/16.956 seconds and 27.434/18.032 seconds: -8.208 and 9.402 seconds saved. Both used Node 24.21.0, pnpm 12.8.1 and Electron -43.7.5. Both shared native-module cache lookups missed, so this improvement did -not depend on a warm native build. Electron archive and root pnpm cache lookups -hit. Dependency trees, pnpm data and Electron archives were reset between paths; -compiler headers and external services were not. Bootstrap, resets, validation, -post-job cleanup, queueing and the production guard step are outside those times. -These are preparation measurements, not whole-workflow or billing savings. - -Both paths passed a native-module probe inside the actual Electron executable -with `ELECTRON_RUN_AS_NODE=1`, and built the same Electron-vite e2e application. -The candidate's 18 focused routing/fallback tests, workflow actionlint and changed -code-quality checks passed. Performance tests, budgets and report uploads remain -unchanged. The [existing October 2 run](https://github.com/stablyai/orca/actions/runs/36985792125) -failed the same-workspace 50/100-terminal budgets (46.9/50.2 ms against 25 ms). -This dependency change does not claim to resolve those application regressions. - -The [full candidate integration](https://github.com/stablyai/orca/actions/runs/37104625474) -passed on `df71ad849cd854a232f7063562785563743b641a`: current preparation was -selected, its native cache missed and rebuilt, the app built and all 32 report -annotation rows passed the unchanged budget checker. The downloaded report also -passed the same checker locally. This is integration evidence; it does not -attribute application latency changes to dependency preparation. Subsequent -rebases resolved report documentation and incorporated fixture teardown fixes. -Workflow, installer-action and toolchain content stayed unchanged. Main also -added an import and a Windows-only MSBuild setting to the native-runtime script: -the imported helper has no top-level side effects, and the Linux rebuild branch -is unchanged. Focused tests verify its Linux/macOS no-op behavior. Final-head PR -checks qualify separately. - -## October 4 reusable cells for terminal context scans - -Terminal cursor-context scans now request one reusable cell per invocation when -the adapter offers getNullCell, and pass it through all unchanged text/style -scans. Adapters without that optional method keep the existing allocating path. -The scratch cell is local and no cell reference escapes into returned context. -Browser composer/readiness text, colors, bold flags and wrapping are unchanged. - -Three alternating one-worker ARM pairs in -[37182789677](https://github.com/stablyai/orca/actions/runs/37182789677) -ran all 19 original cases from readiness census suite 2. Baseline complete -invocations were 37.141 / 37.879 / 37.090 seconds; candidate invocations were -33.887 / 33.387 / 32.916 seconds. Median 37.141 to 33.387 seconds saves 10.1%. -This is a focused workload measurement, not a whole-shard or queue-delay claim. - -Separate baseline/candidate captures retained all 192 cases across six census -suites. Every context and visible projection matched: 643,926 of each, with -7,465,308,324 complete length-prefixed payload bytes hashed per test/type/order. -The canonical capture digest was -`f7440c0f1b5bbb57127cd29245530029415c8e9e243c1744359f330b3c7ace19`. -These captures run outside the timing samples. All 41 cursor/composer/browser -consumer checks passed. Seven faults for lost dim filtering, wide continuation, -bold prompt, custom foreground, wrap preservation, adapter fallback and scratch -reuse failed their intended assertions. Node and web typecheck, lint and format -passed. Two added controls prove per-call scratch lifetime and adapter parity. - -## October 3 producer follow-up: automatic selection for the measured profile - -The first producer rollout in [#24927](https://github.com/stablyai/orca/pull/24927) -passed all 46 PR checks, all five manual warmers and all 11 manual Headless -qualifications on `a2c489c0cca5e46d24333a4d40ba910af0de0208`. The same root installer -also serves recurring unit, browser and performance workflows that had not opted -in. The follow-up defaults the existing input to `auto`, reusing lookup mode for -non-PR root-only installs on GitHub-hosted Linux/macOS/Windows x64/ARM64 runners, -with no job container, the manifest's Node 24/pnpm 12.8.1 profile and no conflicting -Node override. Explicit `true` and `false` retain their previous meanings. Mixed -lockfiles, other toolchains, containers and self-hosted runners retain full cache -restoration; PR policies are unchanged. The manifest check runs only when the -context is potentially eligible, before setup-node chooses its cache behavior. - -A second cleanup audit distinguished nesting depth. The -[twice-nested control](https://github.com/stablyai/orca/actions/runs/37087090689) -published the environment-path payload and lost the output-path payload with an -`Input required and not supplied: path` warning. The -[direct control](https://github.com/stablyai/orca/actions/runs/37087211236) published -and restored both payloads. Current Electron archive callers are direct, so they -need no cache-path change. Keeping the producer's exported path also makes its -new lookup mode safe for callers that nest the shared installer. These tiny -controls establish publication behavior, not installer time savings. - -The [actual automatic-mode cold publisher control](https://github.com/stablyai/orca/actions/runs/37097980789) -passed both jobs on `7b8858bdc8f`. A twice-nested wrapper called the installer -without overriding its default input. The writer selected lookup, missed its -unique root-lockfile key, completed the frozen policy-checked install and saved -that key during cleanup. A fresh reader restored the exact key and installed the -same dependency successfully. The fixture retained the manifest toolchain and -applicable workspace policies; its one dependency keeps the publication check -small. Two earlier trials failed fixture assertions (the pnpm multi-document -header placement, then its empty cache-miss output), and are excluded. This proves -automatic selection and cold publication, not a new timing result. Local -verification passed eight suites / 184 tests, the changed-code quality gate and -compiled-composite actionlint. - -## October 4 shared PR preflight capacity - -Static analysis and the unchanged compiler now share one ARM runner and guarded -Node 24 install. Static checks finish and all background work joins before the -compiler starts; unit planning still overlaps compilation. Each phase keeps its -classifier output. Successful no-op background bodies register every required -join when a phase is unselected or an earlier step failed. Unit and package -consumers depend on physical job success, including action cleanup. - -Three counterbalanced pairs in -[37180613601](https://github.com/stablyai/orca/actions/runs/37180613601) -used the same frozen checkout `f199a20c3acd`, Node 24.21.0, pnpm 12.8.1, -policy hashes, native cache hits, warm TypeScript cache and 10,787-file unit plan. -Both arms used the PR root-only download-store policy. Total active job time was -152 / 153 / 151 seconds separately and 138 / 133 / 129 combined. Excluding the -extra measurement-only evidence steps gives 151 / 151 / 149 versus -136 / 132 / 128 seconds: median 151 to 132, saving 19 seconds (12.6%). -Two heavy runner admissions become one. This saves capacity; it does not prove a -whole-PR latency or queue gain. The median active dependency barrier increases -from 116 to 132 seconds because compilation follows static checks. - -The separate physical-failure run -[37180755694](https://github.com/stablyai/orca/actions/runs/37180755694) -proved that an included TypeScript error failed the actual compiler, its planner -still joined, and unit/package admissions skipped. A registered late action post -failure also blocked both consumers after successful foreground checks and -published shards. All 12 unselected/prior-failure no-op backgrounds joined, and -the downstream audit passed. Local workflow contracts passed 239 tests across -12 suites; lint and formatting passed. - -## October 3 retired-cache collection observation - -The same owner-collection assertion failed in unit shard 3 of -[37098089274](https://github.com/stablyai/orca/actions/runs/37098089274/attempts/1) -and [37100365037](https://github.com/stablyai/orca/actions/runs/37100365037/attempts/1), -requiring a full shard retry despite the focused suite passing locally. Its -three-turn collection budget was shorter than the six-turn plus final yield -pattern already used by the GitLab known-host retirement tests. - -The fixture now uses that existing observation budget. All seven tests, their assertions, -expiry clocks and production code are unchanged. The focused suite passes. A -local fault control changed only the production timer callback to hold its owner -strongly: the owner-collection assertion failed, with the other six tests passing. -The source was restored afterward. Extra collection turns therefore preserve the -strong-retention oracle. Hosted qualification is still required; these observations -do not prove a particular VM-retention cause or quantify avoided retries. - -## October 4 terminal oracle execution - -Three measured test-support changes preserve the original seeds, payloads, -chunk boundaries and meaningful assertions. Serializer comparisons reuse cells -and format only the first mismatch instead of allocating descriptors for every -cell. The terminal parity writer submits every original chunk in FIFO order and -awaits the final parser callback. The independent legacy frame oracle memoizes -measured code-point widths. Its discarded algebra-only case never called -production and still passed when production always threw. - -Three alternating one-worker hosted ARM pairs measured complete invocations: - -| Cohort | Baseline median | Candidate median | Saving | -| ---------------------------------------- | --------------- | ---------------- | ------ | -| Serializer replay/fuzz/descriptor checks | 71.675s | 46.581s | 35.0% | -| Emulator/reconciliation/color parity | 24.095s | 5.411s | 77.5% | -| Frame equivalence | 18.472s | 13.736s | 25.6% | - -[37180517143](https://github.com/stablyai/orca/actions/runs/37180517143) -retained 116 timed serializer passes and three existing/paired-control skips. -Separate captures matched all 190,796,645 raw bytes over 1,611 scenarios and -8,617 checkpoints (SHA256 `00ab219cfb31456af2ecd5e766d1b82d47abc751f6f2de0d7f795e36a936d3c7`), -including complete outputs and diagnostic payloads. Twenty candidate controls -passed; formatting/color/blank/clipping fault controls detected regressions. - -[37181073275](https://github.com/stablyai/orca/actions/runs/37181073275) -retained all 16 parity cases and default fuzz counts. Captures matched 2,325 -batches, 28,182 original chunks and 1,698,285 input bytes, with identical -terminal state and serialization per terminal/batch. Independent terminal -completion order differs, so comparison uses canonical per-terminal ordering -(SHA256 `c38ac1dbbefb9f6dc33ecfe7c495d65b707c1664614544622af93cfc1850e421`). -All 73 callback/parser/other-consumer controls passed; first-callback, reversed -chunks, missing empty boundary and early-completion faults failed. - -The frame candidate passed all 19 retained cases directly against the original -uncached legacy oracle, preserving 4,000 short and 800 near-cap seeded trials. -Sequence, surrogate width, byte width and span-transform faults failed real -assertions. A part-array alternative was rejected after adding time locally. -Hosted Node typecheck passed. These are focused workload savings, not measured -whole-shard or queue-delay improvements; application behavior is unchanged. - -## October 3 unit-selection evidence: include failed references - -The caller's `needs.test.result == 'success'` condition prevented the advisory -collector from reading failed unit runs, despite the reviewer's existing support -for failed tests. A six-run screen from the October 3 occupancy sample found only -one review artifact; it was a full fallback, so it did not validate selection. -Missing artifacts cannot establish that selection catches red tests. - -The caller now permits both success and failure while excluding cancellation and -skipped tests. The collector remains advisory and absent from `verify` dependencies. -Incomplete, interrupted or inconsistent shard records still cannot become complete -reference evidence. Existing omitted-failure tests preserve that negative control. - -The five artifacts from failed [run 37098089274, attempt 1](https://github.com/stablyai/orca/actions/runs/37098089274/attempts/1) -were reviewed locally using the unchanged script. It recognized a complete failed -reference covering 10,606 files and 9,270,307 worker-ms. Its candidate was the full -fallback, so `selectionEvaluated` remained false and no selection promotion is -justified by this control. Focused workflow/reviewer checks passed 24 tests, -including actual caller-expression outcomes for success, failure, skipped and -cancelled states. This repair supplies needed evidence for a later optimization; -it claims no runner-time savings and does not enable selected tests. - -The updated caller also passed the hosted red-run control in -[37100365037](https://github.com/stablyai/orca/actions/runs/37100365037). -The collector succeeded after one unit shard failed, while required verification -remained red. Its review recognized all five shards as a complete reference -(10,608 files, 8,965,977 worker-ms). This was again a full fallback with -`selectionEvaluated: false`, not evidence for enabling selected tests. - -## October 4 runtime imports and recovery fixtures - -Three helper-only tests now import the existing terminal modules directly rather -than initializing the runtime service. Ten copied-loop cases never exercised -runtime memoization: they passed with its cache, timestamp update or prune -invalidation disabled. Two actual helper checks remain. The existing runtime -prune suite now exercises real leaf cache reuse, split prompt timestamps, -ordinary output, fresh prompts and detection after retained-history eviction. -Each of those three production faults fails a real runtime assertion. - -Recovery tests now seed three exact fixture variants once, after the seed child -has closed. Each crash still receives an independent byte-for-byte copy of the -entire database/WAL family and remapped paths. Buffer.equals retains exact byte -comparison without recursive matcher overhead. All 46 original crash boundaries -and retries remain. Four additional copy-isolation/WAL checks run, and teardown -requires that all seed bytes remain unchanged after the full suite. - -Three alternating one-worker hosted ARM pairs in -[37182181976](https://github.com/stablyai/orca/actions/runs/37182181976) -measured these complete invocations: - -| Cohort | Baseline seconds | Candidate seconds | Median saving | -| --------------------------------- | ------------------------ | ------------------------ | ------------- | -| Three imports only, same 15 tests | 19.257 / 19.167 / 19.363 | 1.769 / 1.768 / 1.768 | 90.8% | -| Final four-file runtime cohort | 22.312 / 22.122 / 21.969 | 13.494 / 13.793 / 13.601 | 38.5% | -| Recovery crash boundaries | 24.082 / 24.075 / 24.814 | 8.061 / 8.105 / 9.074 | 66.3% | - -The final runtime cohort has seven real cases versus 16 including the copied -loops; its new runtime case is included in candidate timing. Recovery has 50 -passes versus the original 46. Hosted Node typecheck passed. Recovery faults for -last-byte database/WAL corruption, shared database paths, missing WAL copies and -accepted/unaccepted seed collision failed the intended assertions. These are -focused workload savings, not measured whole-shard or queue-delay gains. - -An independent local cache screen left both caches disabled. Across 14 unchanged -files and 92 cases, a warm Vitest transform cache reduced median invocation time -3.090 to 1.948 seconds, excluding archive costs; its cold arm increased time to -3.281 seconds. Node compilation caching showed no gain. Controls reproduced stale -transforms after TypeScript configuration or plugin-option changes, so persisted -reuse requires a complete transform-input stamp and hosted net-cost evidence. -A separate 130,000-pane leaf-collection optimization was restored: its complete -migration-file timing stayed within noise. The regression fixture remains. - -## October 3 removal fixture cleanup ordering - -[37105566358](https://github.com/stablyai/orca/actions/runs/37105566358) -failed unit shard 4 with `ENOTEMPTY` removing the failed-removal fixture's temporary -directory; the other four shards passed. A client's removal reply intentionally -precedes the detached job's final record persistence. This fixture reset tracking -and removed the directory before waiting for that persistence, allowing a writer -to race cleanup. Its teardown now awaits the existing settlement helper before -resetting tracking or deleting the fixture. Production removal behavior and all -assertions are unchanged. - -All 1,348 runtime tests passed (one existing skip). A temporary controlled queue -held the final record write after the client replied: waiting before reset stayed -pending and passed; resetting before waiting lost the tracked job and failed the -same ordering assertion. The gate was released, both controls drained the captured -job, and the instrumentation was removed. Changed-code quality passed. This proves -the teardown ordering mechanism, not a measured avoided-retry saving. Final-head -hosted qualification remains required. - -## October 4 store oracle and retention fixtures - -The randomized in-place-store test validated the copying oracle twice after -accepted mutations and compared snapshots through the same production parser. -Its 5,000-step retention fixture generated enough tombstones to hit the count -limit, but never reached the 4,096-revision age boundary. - -The test retains all four seeds and 1,500 mutations per seed, removes the duplicate -validation, and projects snapshots directly from the copying oracle's validated -maps. Separate fixtures now check the revision before, at and after expiry and -count overflow. Production code is unchanged. - -Three alternating one-worker pairs on `ubuntu-24.04-arm` in -[37180517143](https://github.com/stablyai/orca/actions/runs/37180517143) -measured baseline invocation times 33.551 / 33.304 / 33.529 seconds and candidate -13.848 / 13.816 / 13.875 seconds: median 33.529 to 13.848 seconds, saving 19.681 -seconds (58.7%). Baseline passed seven tests; candidate passed eight. This is a -focused test saving, not a measured whole-shard or queue-delay change. - -Hosted Node typecheck passed. Separate fault controls failed the intended -assertion for early, late and disabled age expiry, disabled count compaction, -and a snapshot that drops child descriptions. The description fault passes with -the original parser-sharing oracle and fails with the independent projection. - -## October 4 Git contention and remaining readiness waits - -The full Git admission benchmark compared a disabled arm with no correctness -assertions to an enabled arm with structural ledger checks. Its default CI test -now saturates the real base and headroom budgets with FIFO-gated child processes, -queues older background and newer interactive work, releases base slots, and -requires interactive priority, matching outputs and complete permit release. -The full original diagnostic remains opt-in through -`ORCA_GIT_ADMISSION_STORM_MEASUREMENT=1`; both opt-in tests passed locally. -The existing Windows real-Git parity tests remain unchanged; this fixture retains -its existing POSIX platform scope. - -Two remaining Antigravity transcript tests used real 5,000ms refusal windows. -They now use the existing scoped `waitForTranscriptIdle` timer harness after the -emulator drains. All 60 tests, original captured transcripts, deadlines and -readiness assertions remain. - -Three alternating one-worker hosted ARM pairs in -[37180614492](https://github.com/stablyai/orca/actions/runs/37180614492) -measured these complete focused invocations: - -| Suite | Baseline seconds | Candidate seconds | Median saving | -| --------------------- | ------------------------ | ------------------------ | --------------- | -| Git admission storm | 26.619 / 26.635 / 26.582 | 1.017 / 1.018 / 1.016 | 25.602s (96.2%) | -| Antigravity readiness | 27.347 / 27.910 / 27.550 | 13.855 / 13.894 / 13.800 | 13.695s (49.7%) | - -Each candidate passed its original meaningful checks. Hosted Node typecheck -passed. Separate scheduler faults for bypassed admission, withheld release and -FIFO-only priority failed the queued-contention or interactive-start assertion. -Two additional local transcript faults failed the original picker-rejection and -repaint-readiness assertions. These are focused suite savings; whole-shard time -and queue delay were not measured by this experiment. - -## October 4 aggregate unit-test comparison - -A [counterbalanced hosted comparison](https://github.com/stablyai/orca/actions/runs/37197643399) -measured 128.605 seconds less summed test-process time (4.36%) and a 37.694-second -reduction in the slowest shard (5.98%). It compares the accepted optimizations -with their original file snapshots on the same source, five fixed shard -assignments, Node 24, Ubuntu ARM and four workers per process. This is one paired -trial, not a population estimate or a measurement of PR queue delay. - -| Test process | Original snapshots | Accepted optimizations | -| ----------------------------- | ------------------ | ---------------------- | -| Shard 1 | 574.221s | 533.116s | -| Shard 2 | 572.553s | 569.593s | -| Shard 3 | 630.629s | 592.935s | -| Shard 4 | 565.412s | 546.759s | -| Shard 5 | 609.410s | 581.218s | -| Sum: runner time during tests | 2952.225s | 2823.621s | -| Maximum: test critical path | 630.629s | 592.935s | - -Both arms cover exactly 10,838 timed modules on source `9574c8adb253` and tree -`4d5487e8b827`. One added eight-case batching qualification passes in a separate -0.770-second invocation outside the table, completing the 10,839-module ordinary -census. The complete timed case and outcome -comparison accounts for eight approved coverage changes: 105,231 original cases -versus 105,242 candidate cases. Removed copied simulations and an algebra-only -case are accompanied by real runtime, retention, recovery and reusable-cell -regressions. No unexpected case or outcome difference is accepted. - -Each arm starts with distinct empty transform and result caches; Node compilation -caching is disabled. Cold-cache execution ordering remains Vitest's default and -can change with source size. Source snapshots, assignments, raw reports and case outcomes -are checked. Only three case-title fields containing random temporary paths or -a UUID use stable identities, bound to the exact two test-source hashes; raw -titles remain in the artifacts. - -The five dependency setups total 84.284 seconds and are shared by both arms. -The actual paired jobs consumed 5,969 seconds and spanned 2,031 seconds from the -first start to the last completion. Those job figures include both treatments, -setup, uploads and staggered starts; they cannot be assigned to either arm or -used as a workflow saving. The table measures test-process wall time, not CPU -time or the complete CI workflow. Focused-suite percentages elsewhere in this -report are separate measurements and must not be summed into these results. - -The [earlier aggregate trial](https://github.com/stablyai/orca/actions/runs/37193799646) -is rejected because a real test failed; its timings do not qualify a gain. Two -local invocations sharing one XDG directory reproduced the Muse refresher failure. -The fixture now isolates and restores that setting in both arms. The historical -serializer case ledger was also independently corrected from the original source -before this fresh trial; its seven original cases and twenty candidate cases are -an explicit coverage change rather than an assumed equal census. - -## October 4 shard-weight holdouts - -Fresh shard weights were generated with the production importer from a complete -successful run of the accepted source. Two subsequent hosted holdouts used those -same weights and assignments without retraining. The -[first pair](https://github.com/stablyai/orca/actions/runs/37199891967) alternated -existing and fresh assignments across the five jobs; the -[second pair](https://github.com/stablyai/orca/actions/runs/37201939057) reversed -each job's treatment order. - -| Test-process measurement | First: existing | First: fresh | Reversed: existing | Reversed: fresh | -| ------------------------ | --------------- | ------------ | ------------------ | --------------- | -| Sum across five shards | 2813.595s | 2776.421s | 2924.689s | 2878.481s | -| Maximum shard wall | 578.735s | 584.709s | 603.995s | 614.408s | - -Fresh weights reduced summed test-process time by 1.32% and 1.58%, but increased -the slowest shard's time by 1.03% and 1.72%. The small capacity saving comes with -a repeated critical-path regression, so the existing weights remain. The 3.01% -improvement projected from training module durations is not a measured speed -gain. - -Every arm covers the same 10,839 modules and 105,250 case outcomes on source -`9574c8adb253`, tree `4d5487e8b827`, Node 24.21.0, Ubuntu ARM and four workers. -Both trials use cold caches and the same training data, weights, plans and case -identity rules. Complete raw reports, assignments, hashes and opposite treatment -orders are checked before combining the results. Two pairs supply no statistical -confidence or account-wide queue measurement. The second run's jobs started 119 -seconds apart; that stagger and the paired jobs' setup and upload costs are -separate from the treatment timings above. - -## October 4 remaining unit-test opportunities - -The audit retained real child-process, PTY, SSH and crash-boundary tests. It -removed copied simulations or algebra-only cases after fault controls showed -that they could pass with production behavior broken. The retained or replacement -tests exercise production behavior directly. The focused timings in this report -use the final qualified checks. They must not be added together to estimate a -whole-workflow saving. - -Several further changes did not justify promotion: - -- A synchronous readiness-clock screen retained all 49 shard cases and their - outcomes, but complete invocation time changed only from 34.210 to 33.518 - seconds in one local pair. That 2% result was too small to ship without a - stronger result; the original implementation remains. -- A larger local transform-cache screen reduced warm test execution. Separately - measured medians for warm tests (17.251 seconds), extraction (3.539 seconds) and - archive creation (3.092 seconds) sum to 23.882 seconds, versus 23.473 seconds - with caching disabled. This component estimate excludes transfer costs; it is not an - end-to-end measurement. Unkeyed plugin options, inherited configuration and - import priority also produced stale reuse. Persisted test transforms remain - disabled. -- Two successful full-run shadow references identified about 1.9% of - recorded worker time as omittable. That is advisory worker time, not measured - runner occupancy. It does not supply the failed-reference evidence or a - complete selected-run comparison needed to enable test selection. -- Six sampled failed PR runs contained no failed unit job that could trigger - unit-matrix fail-fast. Successful unit siblings of failures in other jobs - cannot be counted as savings from that policy. The sample is too small to - establish a population-wide rate, and the policy remains unchanged. - -These screens reject the examined changes; they do not establish that every -future optimization is exhausted. Shorter admitted jobs and one shared preflight -reduce demand on the existing runner allowance. They do not increase that -allowance or prove lower queue delay under different account traffic. - -## October 5 remaining import, diagnostic and checkout work - -The [hosted comparison](https://github.com/stablyai/orca/actions/runs/37247027814) -used three alternating pairs per treatment at source -`b18b174cd463f852051dc22adc8478b82dee4d1f`. Focused tests ran as fresh one-worker -processes on the four-core, 16 GB Linux ARM runner with Node 24.21.0, pnpm 12.8.1 -and Vitest 4.1.11; persisted transforms and Node compile caches were disabled. - -| Work | Original median | Candidate median | Median paired saving | -| ------------------------------------------------- | --------------- | ---------------- | -------------------- | -| Real incumbent-process test file | 16.478s | 5.976s | 10.520s | -| Commentable-line lifecycle test file | 9.595s | 6.880s | 2.714s | -| Agent-status diagnostic plus semantic test cohort | 11.547s | 8.235s | 3.361s | -| Windows x64 SSH checkout | 22.052s | 16.884s | 5.020s | -| Windows ARM SSH checkout | 46.758s | 33.910s | 13.456s | - -The incumbent cases own distinct sockets, shim directories and process groups. -Running them concurrently preserves all nine outcomes and every real five-second -lsof deadline. Both original and candidate still reject an unreaped helper and -false claims of clean enumeration. Fault receipts prove the actual modified probe -was imported and all owned helpers, groups, sockets and directories were cleaned. -The Windows platform gate retains all nine skips. - -The renderer lifecycle tests keep their six original bodies and real decorator, -zone and model behavior. Only the unrelated saved-note delivery menu is replaced -by a typed throwing facade. Its cleanup assertion rejects unexpected use. Actual -faults in memoization, value-equal refreshes and model replacement still fail their -original assertions; a real menu call fails the facade and the cleanup guard. - -The agent-status benchmark reports counters and timings but asserts only a -nonempty status map and positive elapsed time. It remains available through -`ORCA_BACKGROUND_LAUNCH=1 pnpm exec vitest run --config config/vitest.agent-status-benchmark.config.ts`. -Ordinary discovery removes exactly that reporting test. Its 15 meaningful routing, -index-retention and batch cases remain unchanged. Three real production faults -pass the old reporting test and fail those retained contracts. Manual execution -preserves its JSON schema and all nine deterministic counters at 423 worktrees, -634 tabs and 1,000 events. - -Windows SSH hosts retain complete main, shared, relay, type, configuration, -native, resource and test trees, plus root files. Six actual checkouts per -architecture pinned candidate source `ce69b84d675b0a8b4763b03329c9d6549ac723f1`; -all 16,568 retained files match the full checkout's contents, modes and index. -The original 5,209 required inputs and six further inputs added by the source -rebase are present. Local full/sparse builds match 54 generated artifacts; -explicit test discovery is identical. Missing imported source, a named test or a -required descriptor still fails. Full Windows native/provisioning lanes remain -a separate qualification gate. - -Test timings exclude dependency setup, checkout and queues. Checkout timings -include the action and shell/runner observation boundaries but exclude later -builds and provisioning. These scoped savings must not be summed or treated as -measured changes to full-shard occupancy or PR latency. - -## October 5 mobile typecheck overlap - -The [hosted comparison](https://github.com/stablyai/orca/actions/runs/37249591578) -ran three alternating pairs on the four-core, 16 GB Linux ARM runner, with Node -24.21.0 and TypeScript 6.0.3. The production compiler and test-typecheck ratchet -use installed tools after dependency installation. Both resolved compiler -programs have `noEmit: true`, with no incremental, composite or build-info writes. -The change moves the existing unconditional production join after the foreground -ratchet and before tests; it changes no compiler command or failure policy. - -| Complete typecheck stage | Original serial | Overlapped | -| ------------------------ | --------------- | ---------- | -| Pair 1 | 68.284s | 41.582s | -| Pair 2, reversed order | 65.385s | 41.042s | -| Pair 3 | 64.240s | 40.418s | -| Median | 65.385s | 41.042s | - -The median paired saving is 24.343 seconds, or 37.2% of this stage. The stage -bracket includes native background/wait boundaries and observation overhead, -but excludes dependency installation and the later test suite. All six stages -preserve the three workers' compiler/ratchet verdicts and stdout/stderr hashes. -The test compiler's existing diagnostics remain subject to the unchanged -ratchet. Neither compiler attempted native loading or recorded filesystem -mutations. Maximum aggregate owned-process RSS sampled every 200 milliseconds -was 5.137 GiB; this is a sampled value rather than a kernel peak. The ratchet now -runs even when the concurrently running production compiler later fails. - -Two real, independent type faults still prevent tests from starting. The -production fault preserves the unaffected ratchet and child output; the test -fault preserves the production output and fails the ratchet. Their explicit -control receipts pass even though their intentionally failing jobs are allowed -to finish collecting evidence. - -The separate [external cancellation control](https://github.com/stablyai/orca/actions/runs/37251289372) -held the two real installed compiler entrypoints before checking, while retaining -the production-background/ratchet-foreground topology. All three owned workers -were observed alive 14.275 seconds before the cancellation request. Both native -step outcomes became cancelled, tests did not start, and the runner's final -cleanup log names all three exact worker PIDs. The attempted earlier assertion -of PID absence failed: GitHub performs orphan cleanup after the always-tail -observer and artifact upload. Post-cleanup absence was not observed and is not -claimed. This control supplies no compiler-completion or timing measurement. - -## October 5 explicit RPC registry test setup - -The [hosted comparison](https://github.com/stablyai/orca/actions/runs/37251277897) -ran three alternating pairs of the complete 156-file mixed cohort on source -`d1e08ddd666e3099fc51c79f01f0305f5162186d`, using four workers on the four-core -Linux ARM runner, Node 24.21.0, pnpm 12.8.1 and Vitest 4.1.11. Persisted module -transforms and Node compile caches were disabled. All 153 candidate test files -pass explicit method lists to their dispatchers. A shared throwing fixture -prevents those tests from loading the unused default method catalog and rejects -any accidental iteration of it. The three real catalog subjects remain -unchanged and run beside the candidate subjects in every arm. - -| Complete cohort process | Original | Candidate | Paired saving | -| ----------------------- | -------- | --------- | ------------- | -| Pair 1 | 142.184s | 89.238s | 52.946s | -| Pair 2, reversed order | 142.833s | 88.933s | 53.900s | -| Pair 3 | 141.274s | 90.186s | 51.088s | -| Median | 142.184s | 89.238s | 52.946s | - -The median paired saving is 37.2% of this fixed cohort. All six arms preserve -the complete ordered ledger, including duplicate parameterized case names: -1,469 passes and one existing setup-dependent skip, totaling 1,470 outcomes. -The real catalog subjects contribute 63 of those cases. Candidate bodies and -assertions are unchanged. One file at its line limit uses the dispatch method's -parameter type in place of its equivalent type-only import; its emitted code -matches the measured candidate. - -Seventeen control invocations preserve actual terminal-handler fault detection, -reject unexpected default-catalog consumption, and still detect a missing real -catalog registration. A plain consumption counter prevents global mock-history -resets from erasing the guard; actual `clearAllMocks` and `resetAllMocks` controls -demonstrate that distinction and preserve unrelated call history. The isolated -driver restores all source files after success, faults and a real cancellation. -Independent review also confirms all 153 candidate sources, the fixture, three -catalog subjects and declared qualification inputs survive the rebase unchanged. - -These are fresh-process wall times for this cohort, excluding dependency setup, -queues and other test shards. They do not measure full-shard balance, total PR -runner demand, or a change to the dashboard's PR runtime percentiles. - -## October 5 fused terminal cursor row scans - -Readiness checks repeatedly read terminal rows to recognize composer text. -The shared reader now collects undimmed text and the first visible glyph's style -in one pass. The cursor suffix remains a separate scan; dim glyph attributes, -empty cells, wide characters and wrapped spaces keep their existing behavior. -No grid data is retained between calls. - -The [hosted comparison](https://github.com/stablyai/orca/actions/runs/37263358478) -used source `1bec53ceb23b5f296a38ffa0e085778d32fc7849`, Linux ARM, -Node 24.21.0, pnpm 12.8.1 and one worker. All six fresh processes ran the same -49-case readiness census module, with separate empty Vite caches, filesystem -transform caching disabled and Node compile caching disabled. - -| Complete focused process | Original | Candidate | Paired saving | -| ------------------------ | -------- | --------- | ------------- | -| Pair 1 | 66.412s | 60.464s | 5.948s | -| Pair 2, reversed order | 66.880s | 61.061s | 5.819s | -| Pair 3 | 68.056s | 65.821s | 2.234s | - -The median paired saving is 5.819 seconds, or 8.7% of this fixed workload. -All 294 timed case outcomes passed. Separate qualification preserves complete -context/composer outputs on 98 recordings and the logical cursor/projection -outputs of all 192 Runtime census cases. Nine actual scanner faults fail the -intended assertions; the 89-case IME/composer slice also passes. Blank-row tests -bound cell reads to 72 instead of the original 132 for a 12-column, five-row grid, -with and without a reusable cell adapter. - -The raw timer-driven polling trace differed and was excluded from equivalence -evidence. Logical per-frame output captures match; no raw polling-count equality -is claimed. These measurements exclude setup, queues and other modules and do -not establish a change in full-shard balance, PR percentiles or runner demand. - -## October 5 Qoder test import guards - -The direct Qoder Runtime tests now import the existing unused-default-RPC guard -before their Runtime fixture. Their complete test bodies remain unchanged. -The guard rejects an unexpected registry access instead of loading the full -default-method graph. The three real registry catalogs remain unmocked. - -The [hosted comparison](https://github.com/stablyai/orca/actions/runs/37266875139) -used fixed source `1bec53ceb23b5f296a38ffa0e085778d32fc7849`, Linux ARM, -Node 24.21.0, pnpm 12.8.1 and four isolated fork workers. Every fresh process ran -both Qoder files and all three catalogs: 85 cases across five files. Each invocation -used a distinct empty Vite cache, with results, filesystem transform and Node -compile caching disabled. All 510 timed outcomes passed. - -| Complete five-file process | Original | Candidate | Paired saving | -| -------------------------- | -------- | --------- | ------------- | -| Pair 1 | 17.677s | 16.323s | 1.354s | -| Pair 2, reversed order | 16.827s | 16.021s | 0.806s | -| Pair 3 | 17.324s | 16.211s | 1.113s | - -The median paired saving is 1.113 seconds, or 6.4% of this fixed workload. -Separate actual faults in retained launch recipes, Qoder command selection and -method registration fail the same intended assertions before and after the -imports. Generated launch IDs and shifted stack lines differ in the raw failure -messages; they were preserved and are not claimed byte-identical. - -The single hosted trial occupied 134 runner-seconds including all six samples, -shared setup and upload. That is trial cost, not a production saving. These -focused process measurements do not establish full-shard savings, PR runtime -percentiles, queue relief or a change in the organization's runner allowance. - -## October 5 localization gate work - -The localization coverage audit now classifies only AST nodes that can emit string -parts, and reuses the class-property exclusion set. It still visits every child and -handles JSX text separately. The extraction scope skips only source extensions and -test paths explicitly excluded by the real extractor configuration. Catalogs, -assets, unknown extensions, configuration changes and mixed production changes -still select extraction; both sides of a rename remain covered. - -The [hosted comparison](https://github.com/stablyai/orca/actions/runs/37360932616) -used fixed source `761d63a4e52ff7915d4432c004df13182db63010`, Linux ARM, -Node 24.21.0 and pnpm 12.8.1. Dependency setup was shared outside the samples. -Each timed check was a fresh process with Node compile caching disabled. - -| Complete coverage audit | Original | Candidate | Paired saving | -| ----------------------- | -------- | --------- | ------------- | -| Pair 1 | 9.182s | 4.924s | 4.258s | -| Pair 2, reversed order | 9.133s | 4.924s | 4.209s | -| Pair 3 | 8.930s | 4.924s | 4.007s | - -The median paired saving is 4.209 seconds, or -46.1% of this audit. All three pairs improved; the original range was -0.252 seconds. Complete JSON finding inventories match byte-for-byte, -including all 13 findings and their source locations. All 53 candidate tests pass. -Four real audit mutations fail their intended assertions. An ignored test fixture -passes the real extraction gate; renaming it into production source selects the -gate and fails on the deliberately missing translation key. - -For an existing ignored test path, three conditional comparisons avoid a median -27.628 seconds of extraction work. Original arms really run the full gate -and pass; candidate extraction arms are explicitly `not_selected`. The routing -tests pin the actual extractor inputs and exclusions, so configuration drift fails -the tests. This saving applies only when every changed source path is ignored. - -Relay integration also omits desktop native preparation and the Electron archive -cache. Its existing two files use Node's SQLite, HTTP and WebSocket paths; general -unit shards retain their desktop setup. The existing 16 relay cases pass locally -on Node 24.20.0 and 26.7.0. Removed setup work is not a measured timing saving. - -These comparisons exclude queues and other checks. Extraction already overlaps -other preflight work, so conditional avoidance does not translate directly into -preflight wall time. No change to full-PR percentiles or the concurrency allowance -is established. diff --git a/docs/reference/cline-and-prime-agent-readiness-evidence.md b/docs/reference/cline-and-prime-agent-readiness-evidence.md deleted file mode 100644 index 6509c185f7b..00000000000 --- a/docs/reference/cline-and-prime-agent-readiness-evidence.md +++ /dev/null @@ -1,85 +0,0 @@ -# Cline and Prime Agent readiness: what the transcripts show - -Both agents paint their composer with cursor addressing on the alternate screen, so the line-folded -text tail cannot see it (#23268, #22153). Their readiness is read off the live screen by the -`composer_ready` rules in `src/main/runtime/agent-state-rules/cline.json` and `prime-agent.json`, -through the same tiering as -Antigravity ([`antigravity-readiness-evidence.md`](./antigravity-readiness-evidence.md)): a pane -with an output clock is believed only once quiet, and when a trustworthy screen exists it decides, -so the quiet-process lane cannot settle a dialog it cannot see. Without a readable screen (a -re-attached pane whose grid is untrusted) both keep the quiet-process lane they had before; no -rest-signal entry closes it. Recordings follow -[`agent-pty-transcript-capture.md`](./agent-pty-transcript-capture.md) and are replayed by -`cline-screen-readiness-transcripts.test.ts` and `prime-agent-screen-readiness-transcripts.test.ts`. - -## Cline - -Recorded 2026-09-30 on macOS, `cline` 3.0.66, OpenRouter's free router, with an isolated -`--config`/`--data-dir`. `cline-3-0-65-win32-startup.txt` is a Windows capture from PR #23269. - -| Fixture (`cline-3-0-66-*.txt`) | Screen at the end | Rule | -| ------------------------------ | --------------------------------------------------- | ----------- | -| `ready`, `ready-80x24` | startup composer, `❯ What can I do for you?` | ready | -| `ready-plan` | Plan mode, `❯ Plan something...` | ready | -| `turn-ended` | after a turn, `❯ Ask anything...` | ready | -| `promo` | "Introducing Cline Desktop" drawn over the composer | not ready | -| `permission` | `Approve tool call?` with `[y] Approve [n] Deny` | not ready | -| `slash-menu` | `❯ /` with the command list under it | not ready | -| `draft` | unsent text in the composer | not ready | -| `busy-streaming` | a reply streaming, spinner scrolled off the top | reads ready | - -- **The streaming screen is the idle screen.** Once a long reply scrolls its spinner row away, the - grid is the same composer box as at rest. Only quiescence separates them. A clockless restored - pane still settles from the screen, like Antigravity and Prime: `onPtyData` stamps - `lastOutputAt` on every chunk, so a streaming pane has a clock from its first byte after attach, - and only a pane that has printed nothing since attach is judged on the screen alone. A reply that - stalls for 3s with its spinner off screen would read ready; that is not captured and not ruled - out. -- **The placeholder is not fixed.** It changes with mode and history, so the rule accepts the three - captured placeholders and nothing else. A typed draft looks the same to the read projection, which - is why screen-ruled agents read raw rows (`readScreenRuledLines`, which also requires the PTY's - own grid). Every other agent keeps `readLiveTerminalScreenLines` exactly as before: replaying all - 93 other fixture/grid pairs frame by frame gives identical verdicts on this branch and its base. -- **The promo popup appears about 40ms after the composer** and returns on each launch until it - is dismissed once (`cli-notices.json`). The quiet lane covers that race; the popup carries no - blocked wording. -- **The approval prompt is quiet and unworded.** No blocked rule matches it, so before the screen - decided, the quiet-process lane would have settled it. That lane stays open only while the pane - has no readable screen. -- **Windows:** the reported bug (#23268) is Windows, where the text tail reorders rows. Only the - 3.0.65 contributor capture covers it, and it reads ready from the screen. The rule depends on - rendered rows, not byte order, but no Windows turn or dialog is recorded. - -## Prime Agent - -Recorded 2026-09-30 on macOS, `prime-agent` 0.9.8, OpenRouter -`inclusionai/ling-3.0-flash-sante:free`, with an isolated `HOME` (the first-launch question only -reappears in a fresh one). `prime-agent-0-9-5-*.txt` are 120x35 captures from PR #22154. - -| Fixture (`prime-agent-0-9-8-*.txt`) | Screen at the end | Rule | -| ----------------------------------- | ---------------------------------------------- | --------- | -| `ready`, `ready-80x24` | bare `>` over the `← manage` footer | ready | -| `ready-after-question` | trace question answered "Not now" | ready | -| `turn-ended`, `tool-turn` | a turn (one with a Python tool call) has ended | ready | -| `trace-question` | animated "Share agent traces" question | not ready | -| `slash-menu` | `> /` with the command list | not ready | -| `busy-streaming` | `⠦ Writing · 6s` status row above the composer | not ready | -| `draft` | unsent text in the composer | not ready | - -- **The footer and caret stay up mid-turn.** Only the braille status row Prime keeps directly above - them says a turn is running, so the rule vetoes on it. -- **The rule reads ready for moments it should not.** Replayed in 64-byte chunks, Prime erases that - status row before redrawing it, and on first launch it paints the idle composer a few tens of - milliseconds before the trace question covers it. Both are covered by quiescence: the spinner and the question's - animation repaint continuously, and an idle Prime is silent. -- **No permission prompt exists to capture.** With default settings Prime ran the tool call without - asking. -- **Why the Prime captures are large.** Prime does no cell diffing. Each synchronized frame - (`ESC[?2026h`…`ESC[?2026l`) erases and rewrites every row it touches, so a streaming turn costs - about 4 KB per spinner tick or token (337 frames, 8,593 `ESC[2K` in 1.45 MB). The first-launch - welcome animates a full-screen dotted background with a colour code per glyph, about 10 KB a - frame at roughly ten frames a second. `busy-streaming` and `trace-question` are truncated to the - first frame that shows the screen their tests need (see each `.meta.json`). - `ready-after-question` cannot be: the animation precedes the answer, and truncation only drops - the end. -- 0.9.4's `← agents/resume` layout is not supported; the rule needs 0.9.5 or later. diff --git a/docs/reference/codebuddy-harness.md b/docs/reference/codebuddy-harness.md deleted file mode 100644 index e59eaec855e..00000000000 --- a/docs/reference/codebuddy-harness.md +++ /dev/null @@ -1,44 +0,0 @@ -# CodeBuddy terminal harness - -CodeBuddy uses its own identity in the desktop and mobile agent catalogs, process -detection, telemetry, hooks, resume records and AI Vault. `codebuddy` and `cbc` -identify the interactive CLI; print, server, ACP and detached background modes do -not identify an interactive pane. Model and effort selections use `--model` and -`--effort`; the CLI resolves its stable model aliases for the current account. - -Managed hooks use the existing Claude-compatible installer and transport, with -CodeBuddy's own `.codebuddy/settings.json` and hook source. Installation preserves -user hooks and statusline configuration. Local and SSH installers share the same -plan; Windows explicitly selects CodeBuddy's supported PowerShell hook shell. -Status flows through the execution host's canonical hook store. - -## Observed lifecycle - -Verified with the authenticated CodeBuddy 2.159.0 CLI on macOS: - -- `UserPromptSubmit` starts work. `SessionStart` can arrive afterward and must not - settle that work. -- An unanswered `AskUserQuestion` emits a `Notification` with - `notification_type: permission_prompt` and the message - `needs your permission to use AskUserQuestion`. -- In this version, `PreToolUse(AskUserQuestion)` arrives **after** the user answers; - it resumes working, followed by `PostToolUse` and `Stop`. -- `Stop` supplies the final assistant message and the provider session identity - supplies the resume command. - -The sanitized real event sequence is -`src/shared/__fixtures__/codebuddy-question-hooks.jsonl`. The regression test -replays it through the shared hook listener. Qoder keeps its existing lifecycle -semantics while sharing the common event projection. - -AI Vault reads CodeBuddy's `type: message` JSONL records with top-level roles and -`input_text` / `output_text` content blocks. Local, WSL and SSH discovery use the -provider's `.codebuddy/projects` tree and the same incremental parser. - -## Validation scope - -Live macOS checks exercised launch through Orca's agent menu, model switching, -question waiting, answer submission, working and completion indicators, history -discovery and a resumed session recalling its earlier answer. Hidden-renderer CDP -screenshots record the working, question and completed states. Windows, Linux, -WSL and SSH runtime execution have not been exercised live on this machine. diff --git a/docs/reference/deepseek-build-observation.md b/docs/reference/deepseek-build-observation.md deleted file mode 100644 index 50686cf308d..00000000000 --- a/docs/reference/deepseek-build-observation.md +++ /dev/null @@ -1,28 +0,0 @@ -# DeepSeek Build terminal identity - -DeepSeek Build is the third-party [`innocarpe/deepseek-build`](https://github.com/innocarpe/deepseek-build) product, published as `@innocarpe/deepseek-build`. It is distinct from official DeepSeek Harness (`@deepseek-ai/dsh`), Reasonix, DSH Console and generic DeepSeek TUI wrappers. - -Orca recognizes manually started Build terminals through its existing process and title observations. `TerminalAgent` includes `dsb`; the launchable `TuiAgent` registry does not. No Build launcher, hook, readiness profile, resume command or history reader is registered. Existing generic terminal input remains available. - -The source and actual macOS release were checked at **v6.9.0**, source commit `74df67a56988e9a32845c4565cc62b021ea68c7d`. The darwin-arm64 release tarball SHA-256 is `a57f225a537fc5c027ac4592e3f37f7bdc28cc2d6e27a366934511ea565cb874`. - -- `package.json` publishes `dsb.js` and `deepseek-build.js` npm shims; the native child is `deepseek-build-agent`. -- `crates/dsb-cli/src/main.rs` separates the full-screen entry from `run`. Global value options such as `--cwd` can precede `run`; those invocations remain excluded from interactive recognition. -- `crates/dsb-cli/src/agent_launch.rs` emits the product OSC 0 title. The vendored pager's `notifications/title.rs` composes spinner, activity and product segments with ` - ` separators. -- The committed `dsb-6-9-0-folder` PTY fixture records the released binary's welcome screen and actual title in an isolated home and plain folder. It makes no successful-authentication or completed-model-turn claim. Its runtime test feeds raw chunks through `onPtyData` with foreground inspection unavailable. - -Explicit native owner markers retain their existing precedence. A Claude task merely mentioning Build is not a Build identity. Runtime publication reuses the existing optional `agentIdentity` string; no new RPC, stream opcode or status producer is added. Older hosts can omit identity, while older readers retain their existing unknown-agent handling. Execution-host process/title observations work without a Git repository; local source tests do not establish native Windows, Linux or SSH device coverage. - -For rendered proof, isolate both Electron and the actual PTY. On macOS, `login(1)` can replace the shell's inherited home. Test-only `ORCA_DISABLE_MACOS_LOGIN_SHELL=1` avoids that wrapper; do not change production launch policy for a proof. Require a nonce-bound file written by a helper executed in the spawned PTY, containing its actual `HOME`, `USERPROFILE`, `DEEPSEEK_BUILD_HOME`, `GROK_HOME` and trust-RPC flag, and verify it before agent launch. A terminal-text assertion can match command echo and is not isolation evidence. Explicit provider environment at the final execution boundary protects the test even after shell startup. - -The observation-type propagation and title/process recognition adapt Wooseong Kim's (`innocarpe`) [PR #23485](https://github.com/stablyai/orca/pull/23485), with source-backed corrections for the second npm shim and value options before `run`. Keep that predecessor open until a reviewed successor merges. - -Independent review follow-up: upstream 6.9.0 outer `Commands::Agent` forwards native PagerArgs options. Native `-p`/`--single` (alias `--print`), `--prompt-json` and `--prompt-file` are one-shot forms and are excluded from interactive process/foreground identity, including equals/compact short forms, npm wrappers and preceding value options. Positional interactive prompt text, native option values and the native `--` terminator stay distinct. Actual release native `--help` confirms exposed flags; source alias and forwarding are pinned above. - -Title follow-up confines Gemini identity/normalization and status sniffing before inspecting Build activity text. A verified Build title uses its leading own braille frame for working and leading `⚠ Action Required - ` for permission; embedded Gemini glyphs in activity/session/cwd text do not change its identity or status. Source-backed frame tests cover wrapped and alert variants, plus OSC input through the actual runtime/listing path. Native Gemini and other provider corpus contracts stay covered. - -Actual released outer `dsb agent -- --help` prints the native TUI help (`outer-forwarded-help.txt`), confirming Clap consumes the outer separator before forwarding. The observer distinguishes this from the native `--`: `dsb agent -- --print task` is one-shot, whereas `dsb agent -- -- --print` and direct native `-- --print` retain literal interactive prompt text. - -Native grammar follow-up: only the outer wrapper's `run` subcommand is one-shot. Native `deepseek-build-agent run`, forwarded `dsb agent run`, and `--leader-socket run` remain interactive; the last consumes `run` as a path value. Native `-c` is boolean, so Clap accepts `-cp task` and `-cptask` as continue plus single-turn prompt. Attached `-m`/`-r`/`-s`/`-w` values (including after `c`) do not expose a prompt flag. The released binary accepted the five review argument topologies with `--help` under an executed private-child environment assertion (`native-grammar-oracle.json`); this proves parsing/help, not successful model generation. - -Attached prompt values can begin with hyphens: native `-p-` and `-cp--print` consume `-` and `--print` as the single-turn prompt. The observer accepts the entire remainder after `p`, while the attached m/r/s/w value shields stay covered. The released native binary and outer `agent` wrapper accepted both forms with `--help` in a nonce-asserted private child (`attached-p-oracle.json`). diff --git a/docs/reference/dsh-harness-integration.md b/docs/reference/dsh-harness-integration.md deleted file mode 100644 index 9a850e06900..00000000000 --- a/docs/reference/dsh-harness-integration.md +++ /dev/null @@ -1,44 +0,0 @@ -# DeepSeek Harness integration - -Orca detects the community `@deepseek-harness-tui/dsh-tui` launcher (`dsh-tui`, alias -`dst`) and requires the official `@deepseek-ai/dsh` executable too. The launcher -boots the `dsh-tui` profile; Orca passes `.` to select the current workspace and -reach its composer on the first launch. The official Harness does not bundle this community TUI. -DSH Console and DeepSeek Build are separate products and are not interchangeable -with this launch contract. - -The official DSH 0.2 CLI accepts both `dsh --profile headless` and `dsh headless`. -Orca excludes the known `web`, `headless`, `sdk`, `sdk-minimal`, `acp`, and `desktop` -profiles from interactive process recognition, along with plugin management and -configuration dumps. Custom profile names remain eligible because profiles are -user configurable. Only launcher arguments are inspected; app prompts, resume IDs, -and patch filenames cannot change the selected profile's identity. - -Status hooks use the official `@deepseek-ai/dsh-hooks-claude-code` plugin, installed -as an owned block in `$DSH_HOME/cordis.patch.yml`. User entries outside the block are -preserved. Local installation respects `DSH_HOME`; the existing SSH installer uses -the execution host's default `~/.dsh` because SFTP cannot read its environment. -Hooks report session start, prompt submission, tool start/end, and stopping through -Orca's host status store. Approval has no dedicated hook; it is not inferred from -an uncaptured screen. Subagent lifecycle events are ignored for parent-pane status. - -DSH 0.2 still emits an empty `transcript_path` in Claude-compatible hooks. Its -session persistence defaults to compressed JSONL under `$DSH_HOME/sessions`. -Orca can resume a hook-associated session through `dsh-tui --resume `, but -currently does not discover DSH logs in Agent Session History. Resume support alone -does not establish transcript-history support. - -## Reproduce the official launcher check - -Install `@deepseek-ai/dsh@0.2.0-rc.2` into a disposable prefix, then run: - -```sh -ORCA_BACKGROUND_LAUNCH=1 ORCA_REAL_DSH_CLI=/path/to/prefix/node_modules/.bin/dsh \ - pnpm test src/shared/dsh-real-cli.test.ts -``` - -The opt-in test checks published version, composed profile configurations, and -headless help in an isolated home and working folder without a model request. -Interactive readiness is separately pinned to the captured community TUI transcript -in `src/main/runtime/__fixtures__/dsh-tui-ready-no-key.txt`; that older capture is -not proof of current TUI compatibility or paid generation. diff --git a/docs/reference/git-compatibility.md b/docs/reference/git-compatibility.md deleted file mode 100644 index bbfb4a84876..00000000000 --- a/docs/reference/git-compatibility.md +++ /dev/null @@ -1,94 +0,0 @@ -# Git Compatibility Policy - -## Scope - -Orca executes the user's Git binary on three kinds of execution host: native, -WSL, and SSH. Each host can have a different Git version, so compatibility -state must be scoped to the host that actually runs the command. - -Git 2.25 is the core-workflow compatibility baseline for command selection. It -is the oldest line that covers Orca's baseline use of porcelain v2, `branch ---show-current`, `restore`, and sparse checkout. Optional features that need a -newer Git must degrade safely and cache the missing capability. Orca does not -currently block older Git at startup, but new command construction should not -assume features introduced after this baseline. - -## Capability Rules - -When a newer Git feature materially improves correctness or performance: - -1. Keep a baseline-compatible command or parser as the fallback. -2. Detect rejection with a narrow predicate for that option or subcommand. -3. Run the preferred command through `GitCapabilityCache` so a rejection is - remembered for the native host, WSL distro, or SSH provider that produced it. -4. Retry after the cache interval so an in-place Git upgrade self-heals without - restarting Orca. -5. Test the first fallback, later calls that skip the rejected probe, concurrent - probe coalescing, and execution-host isolation where applicable. - -Do not branch only on a parsed `git --version`. Vendor builds can backport -features, and wrappers can report a host version that differs from the binary -used inside WSL or SSH. A behavior probe plus a precise fallback is the final -authority. - -## Current Capabilities - -| Capability | Preferred behavior | Compatibility behavior | -| --------------------------- | ----------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `fetch-no-write-fetch-head` | Fetch a private rebase ref without changing worktree-local `FETCH_HEAD` | Serialize all Orca fetch/pull operations per worktree Git directory before Git 2.29 | -| `worktree-list-z` | NUL-delimited worktree paths with `prunable` marks | Line-block parser for Git before `worktree list -z` (2.36); the `prunable`/`locked` annotations still parse on Git 2.31–2.35, and a path-existence probe restores `prunable` detection for Git before 2.31 | -| `worktree-add-lock-reason` | Create a prepared checkout with its ownership marker already present | Before Git 2.33, add without checkout, then exclusively create the same reason marker before materializing files | -| `rev-parse-path-format` | Absolute repo metadata paths | Resolve legacy relative output against the scanned repo | -| `for-each-ref-exclude` | Exclude remote HEAD before the output limit | Request extra refs, then filter remote HEAD in Orca | -| `merge-tree-write-tree` | Derive real-merge conflicts and no-op tree proofs | Omit the conflict summary and keep conservative branch cleanup behavior before Git 2.38 | -| `merge-tree-merge-base` | Supply the already-resolved merge base | Use the older two-commit `merge-tree --write-tree` form | - -Prepared creation registers and locks without checking out files while its shared -exact-base fetch runs. A cancellable in-process barrier waits for fetch settlement, -including offline failure, then resolves the current commit OID on the owning host -and materializes files once. Existing preparations queue a tip refresh on that same -barrier before a create can claim them. Finalization still resolves the latest base -and runs the post-checkout hook only when attaching the requested branch. This uses -baseline-compatible `rev-parse` and `reset --hard`; the barrier never enters Git -transport options or the remote wire. - -### Placeholders That Fail Open - -`GitCapabilityCache` records commands Git _rejects_. A `git log --format` -placeholder Git does not know is not rejected: Git echoes it verbatim and exits -zero, so there is no error to remember and no probe to cache. Ask for both forms -in one record and pick at parse time. - -| Placeholder | Preferred behavior | Compatibility behavior | -| --------------- | ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- | -| `%(decorate:…)` | Git 2.43 separates commit decorations with `\x1f`, so ref names containing commas survive | The same record also carries `%D` (Git 2.10); an unexpanded `%(decorate` placeholder selects it, at the cost of comma-splitting | - -## Why Not `simple-git` - -`simple-git` is a process wrapper around the installed Git binary. Its custom -options and `raw` API pass arguments through to Git, so it cannot make a newer -flag work on an older binary or choose Orca's semantic fallback automatically. -It provides version reporting and subprocess queueing, but Orca already needs -its own WSL/SSH routing, cancellation, tracing, redaction, process cleanup, and -bounded output handling. Replacing the runner would move—not remove—the -capability problem. - -## CI Contract - -PR checks run the capability contract against real Git 2.25.5, 2.38.1, and -2.49.1 binaries. This spans the pre-2.29 serialized `FETCH_HEAD` fallback, the transitional -`merge-tree --write-tree` behavior before `--merge-base`, and current Git. - -The three lanes run in parallel and each Git call in the container lanes costs a -container start, so their wall clock is runner contention, not Git. Build the -2.25.5 binary and pull the images before the lanes start: anything heavy left -running alongside them is charged to whichever boundary case is in flight and -surfaces as a Vitest timeout rather than as a slow setup step. - -The idle maintenance contract also verifies `multi-pack-index write` and packed -object reads. The command arrived in Git 2.20 and needs no newer-Git fallback; -Orca uses only index metadata writes, respecting `core.multiPackIndex=false`. - -Keep the unit tests alongside that matrix. They cover concurrent probes, -native/WSL/SSH/relay isolation, and error-stream shapes that a single real -binary invocation cannot exercise deterministically. diff --git a/docs/reference/headless-linux-server.md b/docs/reference/headless-linux-server.md deleted file mode 100644 index 8bbfc0af6d8..00000000000 --- a/docs/reference/headless-linux-server.md +++ /dev/null @@ -1,1019 +0,0 @@ -# Headless Linux Server - -Use this guide when you want to run `orca serve` on a Linux machine without a -desktop session, such as an Ubuntu VPS or a remote build box. - -`orca serve` starts the Orca runtime without opening the desktop window. On -Linux, the packaged AppImage still needs the libraries that Electron expects at -startup. Current Orca builds start Xvfb automatically for `orca serve` when no -`DISPLAY` is set, but Xvfb must be installed first. A separate D-Bus session is -not required. When `DISPLAY` is set, Orca uses that display instead of starting -a competing Xvfb process, provided the display is usable: its socket must exist, -and if an X lock file is present it must name a running process. A `DISPLAY` -whose lock names a dead process is refused rather than replaced, and `orca serve` -exits — unset `DISPLAY` to let Orca start its own Xvfb. A socket published with -no lock at all (a container bind-mounting `/tmp/.X11-unix`, or WSLg) is accepted. - -The supported deployment matrix covers Ubuntu 20.04, 22.04, and 24.04 and -current Debian stable — anything with glibc 2.31 or newer (see -[Linux glibc compatibility](./linux-glibc-compatibility.md)). Package names can -differ on other Debian-derived releases. - -## Ubuntu and Debian prerequisites - -Install the CLI tools, Xvfb, and the shared libraries Electron links against. -A minimal server or container image ships none of the Electron libraries, and -`orca serve` then fails before Electron starts: - -```bash -sudo apt-get update -sudo apt-get install -y \ - curl file jq xvfb zlib1g-dev ca-certificates git \ - libgtk-3-0t64 libnss3 libatk1.0-0t64 libatk-bridge2.0-0t64 libgbm1 libasound2t64 \ - libxtst6 libcups2t64 libdrm2 libxkbcommon0 libpango-1.0-0 libcairo2 libatspi2.0-0t64 \ - libxcomposite1 libxdamage1 libxfixes3 libxrandr2 libxrender1 libx11-xcb1 \ - libxcb-dri3-0 libxss1 -``` - -That command is for Ubuntu 24.04 and newer and Debian 13 and newer. Those -releases carried out the 64-bit `time_t` transition, which renamed six of the -packages with a `t64` suffix. On Ubuntu 20.04, Ubuntu 22.04, and Debian 12, -substitute the unsuffixed names: - -- `libgtk-3-0t64` becomes `libgtk-3-0` -- `libatk1.0-0t64` becomes `libatk1.0-0` -- `libatk-bridge2.0-0t64` becomes `libatk-bridge2.0-0` -- `libasound2t64` becomes `libasound2` -- `libcups2t64` becomes `libcups2` -- `libatspi2.0-0t64` becomes `libatspi2.0-0` - -The other names are identical on every supported release. The substitution is not -symmetric, so use the list that matches the release. A `t64` name on Ubuntu 20.04, -Ubuntu 22.04, or Debian 12 fails immediately with `E: Unable to locate package -libgtk-3-0t64`. In the other direction the old names mostly still resolve, because -each renamed package declares `Provides:` its unsuffixed name — except `libasound2` -on Ubuntu 24.04, where `liboss4-salsa-asound2` in `universe` claims that name too. -apt will not choose between two providers and exits with `E: Package 'libasound2' -has no installation candidate`, which aborts the entire install line and leaves none -of the libraries installed. - -On Ubuntu 20.04 and 22.04, install `libfuse2` to execute the AppImage through -FUSE. On Ubuntu 24.04 and Debian 13 the package is `libfuse2t64`, though the plain -`libfuse2` name also resolves there because nothing else provides it. FUSE is -optional: without it, use the AppImage's supported extraction path. CLI -registration does this once automatically, so registered commands do not need -FUSE. - -Download and make the AppImage executable: - -```bash -sudo mkdir -p /opt/orca -sudo curl -L https://github.com/stablyai/orca/releases/latest/download/orca-linux.AppImage \ - -o /opt/orca/orca-linux.AppImage -sudo chmod +x /opt/orca/orca-linux.AppImage -``` - -To extract it without FUSE, run the extraction as root because the installation -directory is root-owned: - -```bash -cd /opt/orca -sudo ./orca-linux.AppImage --appimage-extract -sudo chmod -R a+rX /opt/orca/squashfs-root -/opt/orca/squashfs-root/AppRun serve --port 6768 -``` - -The `chmod` is required whenever the extraction runs as a different user than -the server: `--appimage-extract` creates `squashfs-root` as `drwx------` owned by -the extracting user, so anyone else — including a dedicated service user — cannot -even traverse it, and the run fails before Electron starts. - -Docker commonly has no FUSE device. Use `--appimage-extract` once or -`--appimage-extract-and-run`; neither requires a privileged container. The -extract-and-run wrapper can print extracted paths before Orca starts, so -automation that requires stdout to contain only the ready JSON should extract -once and invoke `squashfs-root/AppRun`. - -If `Xvfb` was installed somewhere other than `/usr/bin`, confirm systemd can -find it later: - -```bash -command -v Xvfb -``` - -## Run In The Foreground - -Start with a foreground run before creating a service: - -```bash -LIBGL_ALWAYS_SOFTWARE=1 /opt/orca/orca-linux.AppImage serve --port 6768 -``` - -For remote clients, pass the address they should use to reach this server. A -Tailscale address is usually the safest option for private servers: - -```bash -LIBGL_ALWAYS_SOFTWARE=1 /opt/orca/orca-linux.AppImage serve \ - --port 6768 \ - --pairing-address 100.64.1.20 -``` - -`--pairing-address` is only the address advertised to clients. It does not -change the listener bind address. Orca binds its WebSocket listener, then -combines the actual bound port with the advertised host when the address omits -a port. Use a reachable LAN/Tailscale hostname or IP, or a complete reverse -proxy URL such as `https://orca.example.com/runtime` (`http(s)` is normalized -to `ws(s)`). Wildcard addresses such as `*`, `0.0.0.0`, and `::` cannot be -advertised. - -The command writes one ready block to stdout after the listener bind and -pairing initialization complete: - -```text -Orca server ready -Bound endpoint: ws://0.0.0.0:6768 -Advertised endpoint: ws://100.64.1.20:6768 -Pairing URL: orca://pair?code=... -``` - -For supervisors, request the versioned single-line JSON contract: - -```bash -/opt/orca/orca-linux.AppImage serve --port 6768 \ - --pairing-address 100.64.1.20 --json -``` - -The actual output is one compact line; this example is pretty-printed for -readability: - -```json -{ - "type": "orca_server_ready", - "schemaVersion": 1, - "runtimeId": "...", - "endpoint": "ws://0.0.0.0:6768", - "boundEndpoint": "ws://0.0.0.0:6768", - "advertisedEndpoint": "ws://100.64.1.20:6768", - "managedWslCliReconciliation": "settled", - "pairing": { - "available": true, - "url": "orca://pair?code=...", - "endpoint": "ws://100.64.1.20:6768", - "deviceId": "...", - "webClientUrl": "...", - "scope": "runtime", - "qr": null - } -} -``` - -`endpoint` remains a compatibility alias for `boundEndpoint`; new automation -should use the explicit bound and advertised fields. - -When the server remains usable but cannot mint an offer, `pairing` remains an -object with `available:false`, a stable `reason`, and operator `guidance`; it is -never silently omitted. `--recipe-json` is stricter and exits with that reason -because its contract requires a pairing URL. Stop a foreground server with -`Ctrl+C`. Stable reasons are `disabled_by_operator`, `websocket_unavailable`, -`device_registry_unavailable`, `e2ee_key_unavailable`, and -`invalid_advertised_endpoint`. - -## Systemd Service - -Create a dedicated service user and install directory. Run the service as this -user instead of root so the AppImage can keep Chromium's sandbox enabled. Keep -the install directory root-owned: the service needs to read and execute the -AppImage, but must not be able to replace it or the rollback artifacts. - -```bash -sudo useradd --system --create-home --shell /usr/sbin/nologin orca -sudo chown root:root /opt/orca /opt/orca/orca-linux.AppImage -sudo chmod 755 /opt/orca /opt/orca/orca-linux.AppImage -# Only if you ran --appimage-extract: extraction leaves squashfs-root root-only. -sudo chmod -R a+rX /opt/orca/squashfs-root -``` - -The last line matters because the two halves of this guide combine badly without -it. `--appimage-extract` writes `squashfs-root` as `drwx------ root root`, so the -`orca` service user cannot read or traverse the extracted tree and the unit fails -at startup. `chmod 755 /opt/orca` alone does not reach into it. - -For most hosts, one `orca serve` service is enough because Orca starts Xvfb on -display `:99` when no display exists: - -```ini -# /etc/systemd/system/orca-serve.service -[Unit] -Description=Orca runtime server -After=network-online.target -Wants=network-online.target -StartLimitIntervalSec=300 -StartLimitBurst=5 - -[Service] -Type=simple -User=orca -WorkingDirectory=/home/orca -Environment=LIBGL_ALWAYS_SOFTWARE=1 -ExecStart=/opt/orca/orca-linux.AppImage serve --port 6768 --pairing-address 100.64.1.20 -StandardOutput=journal -StandardError=journal -KillMode=mixed -Restart=on-failure -RestartPreventExitStatus=3 -RestartSec=5 - -[Install] -WantedBy=multi-user.target -``` - -Replace `100.64.1.20` with the LAN, Tailscale, tunnel, or public hostname that -clients should use. - -`KillMode=mixed` sends the graceful stop signal only to Orca's main process, -then `SIGKILL`s whatever is still in the cgroup the instant that main process -exits — `TimeoutStopSec` only governs how long systemd waits for the main -process itself, never a grace window for the cgroup's remains. This lets Orca -keep its owned Xvfb alive until Electron disconnects cleanly. - -The detached terminal daemon is preserved by a different mechanism: it is -launched through `systemd-run --user --scope`, so it and its PTYs live in their -own transient `orca-daemon-.scope` unit rather than in -`orca-serve.service`'s cgroup. A `systemctl stop` or `restart` of this unit -leaves that scope running, so live terminals and agent processes survive the -restart and the successor adopts them. - -That requires a reachable systemd **user** manager for the service account. -With `User=orca` and no interactive login there is none by default, so enable -lingering once: - -```bash -sudo loginctl enable-linger orca -``` - -Without it — or on a host without systemd as PID 1, or without `systemd-run` -on `PATH` — the daemon falls back to launching directly inside -`orca-serve.service`'s cgroup, and is then killed when the stop completes: -every `systemctl stop` or `restart` ends live terminals and agent processes, -even though their persisted layout and terminal history remain. Check which -case a running host is in with the `cgroupUnit` field of the daemon health -payload: a `orca-daemon-*.scope` value means isolated, `null` means the -unscoped fallback. - -None of this applies inside a Docker container. There the capability probe -fails closed (no `/run/systemd/system`), but that is the least of it: a -`docker restart` tears down the container's PID namespace, so no in-container -setting — lingering, kill mode, or scope — preserves the daemon or its PTYs -across it. Run the container with `--init` so a real PID 1 reaps exited PTY -subprocesses; without it, Orca is PID 1 and those children accumulate as -zombies because nothing reaps them. - -Exit status `3` means another process already owns this userData profile, so -`RestartPreventExitStatus=3` stops the unit instead of retrying a launch that -cannot succeed. Any other permanent startup fault is capped at 5 starts per -5 minutes; systemd's defaults (10s window, 5 starts) can never trip at -`RestartSec=5`, which is how one bad launch could restart thousands of times. -The start limit counts operator-initiated starts too, so once it trips systemd -refuses a plain `systemctl start` until the 5-minute window rolls over. Run -`sudo systemctl reset-failed orca-serve.service` first to clear it — the -[Upgrade](#upgrade-steps) and [Roll back](#roll-back) scripts already do. -On systemd older than 230 those two directives are spelled -`StartLimitInterval=`/`StartLimitBurst=` and belong in `[Service]`; Ubuntu -20.04, Orca's oldest supported base, ships systemd 245. - -Enable the service: - -```bash -sudo systemctl daemon-reload -sudo systemctl enable --now orca-serve.service -sudo journalctl -u orca-serve.service -f -``` - -`journalctl -o cat` removes journal metadata but still mixes the service's -stdout and stderr. Parse each line as JSON and require the readiness type and -schema before treating the service as ready: - -```bash -sudo journalctl -u orca-serve.service -o cat \ - | jq -Rrc 'fromjson? | select(.type == "orca_server_ready" and .schemaVersion == 1)' -``` - -A bounded health check should require that contract within its startup timeout; -otherwise inspect earlier diagnostics for the precise pairing reason, listener -error, or missing library. - -## Managed Xvfb Service - -If you prefer to own the virtual display lifecycle in systemd, run Xvfb as a -separate service and set `DISPLAY=:99` for Orca. - -```ini -# /etc/systemd/system/orca-xvfb.service -[Unit] -Description=Virtual X display for Orca -After=network-online.target -Wants=network-online.target - -[Service] -Type=simple -ExecStart=/usr/bin/Xvfb :99 -screen 0 1280x1024x24 -nolisten tcp -Restart=on-failure -RestartSec=5 - -[Install] -WantedBy=multi-user.target -``` - -If `command -v Xvfb` returned a different path, update `ExecStart` to that -absolute path. - -Then add the display dependency to the Orca service: - -```ini -# /etc/systemd/system/orca-serve.service -[Unit] -Description=Orca runtime server -After=network-online.target orca-xvfb.service -Wants=network-online.target orca-xvfb.service -StartLimitIntervalSec=300 -StartLimitBurst=5 - -[Service] -Type=simple -User=orca -WorkingDirectory=/home/orca -Environment=DISPLAY=:99 -Environment=LIBGL_ALWAYS_SOFTWARE=1 -ExecStart=/opt/orca/orca-linux.AppImage serve --port 6768 --pairing-address 100.64.1.20 -KillMode=mixed -Restart=on-failure -RestartPreventExitStatus=3 -RestartSec=5 - -[Install] -WantedBy=multi-user.target -``` - -`KillMode=mixed` matters as much here as in the single-service unit: without it -the unit silently defaults to `KillMode=control-group`, which `SIGTERM`s the -whole cgroup at once and then stalls the full `TimeoutStopSec` before the -`SIGKILL`. - -Enable both units: - -```bash -sudo systemctl daemon-reload -sudo systemctl enable --now orca-xvfb.service orca-serve.service -``` - -## CLI Install Note - -The registered Linux CLI command is `orca-ide`, not `orca`, to avoid shadowing -the GNOME Orca screen reader. Desktop-managed terminals receive a -terminal-scoped bare-`orca` shim. A packaged headless `orca serve` also makes a -best-effort dispatcher at `$HOME/.local/bin/orca` for the service user's own -shell, so the Claude Teams launcher can resolve its bare command; it does not -replace another user's `orca`. From an ordinary shell outside that service -user's managed environment, substitute `orca-ide` for `orca` in commands below. - -On a headless host, you do not need to open the desktop UI just to run the -server. Invoke the AppImage directly: - -```bash -/opt/orca/orca-linux.AppImage serve --help -``` - -Running an AppImage as root requires Chromium's `--no-sandbox` switch before -the command: - -```bash -/opt/orca/orca-linux.AppImage --no-sandbox serve --port 6768 -``` - -This disables a security boundary. Prefer a dedicated unprivileged service -user, especially when the listener is reachable beyond localhost. - -The Linux CLI is named `orca-ide`, not `orca`, so it never shadows the GNOME -Orca screen reader at `/usr/bin/orca`. The `.deb` and `.rpm` packages put -`orca-ide` on `PATH` themselves at install time; with the AppImage it arrives -as `~/.local/bin/orca-ide` when the CLI is registered. - -A packaged `orca serve` start also writes a bare `orca` into `~/.local/bin` -that execs the same launcher, which is why the skills commands below can be -typed as `orca`. It writes it while starting, so it is never the command that -starts the server — the first launch is `orca-ide serve`, or the AppImage -invoked directly as above. The write is best-effort: it is gated on a packaged -build, it is skipped when no bundled launcher resolves, and it is skipped when -a file Orca does not own already holds that name (ownership is a marker on the -second line of the file). A host that really does run the screen reader keeps -its own `orca`. - -## Pairing troubleshooting - -- A pairing offer is a capability containing a device credential and E2EE - material. Share it only with the intended client and do not put it in proxy - access logs. -- `boundEndpoint` is where the process listens; `advertisedEndpoint` is what a - client dials. A valid-looking offer still cannot connect if DNS, firewall, - Docker port publishing, Tailscale policy, or a reverse proxy does not route - the advertised endpoint to the bound port. -- An omitted advertised port uses the actual bound port, including a fallback - port selected after a collision. An explicit proxy port is preserved. A port - mismatch therefore means the supplied external routing is wrong, not that - Orca changes it. -- Reverse proxies must support WebSocket upgrade and route the advertised path. - Use `wss://` or `https://` when TLS terminates at the proxy; do not advertise - `ws://` through an HTTPS-only endpoint. -- Hostnames, IPv4, bracketed IPv6, and raw IPv6 literals are supported. IPv6 - still requires an IPv6-reachable listener/network path. -- Background push notifications to a paired phone do not fire from a headless - server: agent-completion detection runs in the desktop renderer, which is not started in serve - mode, so nothing reaches the push gateway even though the phone - registers successfully. -- `xvfb-run` and `dbus-run-session -- xvfb-run` remain valid diagnostic launch - shapes, but neither should be needed when `Xvfb` is installed and no display - is configured. Repeated D-Bus messages without a ready block indicate startup - did not reach serve mode; confirm the AppImage version and exact argument - order, especially `--no-sandbox serve`. - -If you later install the desktop CLI from Orca settings, use that CLI for normal -shell workflows. Keep the AppImage path in systemd so service restarts do not -depend on an interactive shell profile. - -## Upgrade - -`orca serve` never updates itself. In headless mode Orca wires up no auto-updater -at all — the built-in updater only runs in the desktop GUI, and no paired mobile -or web client can trigger it remotely. Upgrading is always a deliberate step: -replace the AppImage and restart the service. - -Two facts make the persisted-state transition predictable: - -- **State lives in the service user's home, not next to the binary.** Persisted - data is under `/home/orca/.config/` (Orca uses both an `orca` and an `Orca` - directory there), fully independent of `/opt/orca/orca-linux.AppImage`. - Replacing the binary never touches projects, worktree metadata, terminal - history, orchestration state, or paired-device keys — so mobile and web - clients reconnect after an upgrade without re-pairing. -- **New builds migrate old state on load.** Orca imports older `orca-data.json` - into each profile's `profile-state.db`, so a forward upgrade needs no manual - data step. Normal writes, shutdown and profile switching save SQLite; they - no longer refresh the legacy JSON file. - -These guarantees preserve live processes only when the daemon is in its own -`orca-daemon-*.scope`, as reported by `health.terminalDaemon.cgroupUnit`. The -unscoped fallback remains destructive: a service restart kills every terminal -and agent in the service cgroup; an agent conversation may be resumable, but -its current process and any in-flight command are gone. Treat a stop as -destructive unless `health.terminalDaemon.cgroupUnit` names an -`orca-daemon-*.scope` on that host. - -When `cgroupUnit` is `null` or unverifiable, immediately before stopping the -service, obtain a fresh census as the service's -OS account and home. Use the installer's absolute launcher path so `sudo`'s -`secure_path` cannot hide a per-user registration: -`sudo -Hu orca /home/orca/.local/bin/orca-ide terminal list --json`. -Replace both `orca` and `/home/orca` with the service account and home used by -your unit; for an extracted deployment, use its absolute `resources/bin/orca-ide` -launcher instead. Proceed only when the result is -untruncated, has an explicit `hostScope`, covers every execution host affected -by this service stop, and lists no terminals on those hosts. Every -`omittedHostIds` entry must be explicitly accounted for outside this service's -execution boundary. A separately paired runtime is outside that boundary; local -execution and SSH hosts reached through this runtime are not. An affected or -unknown omission, missing scope, failed request or lost connection is -`unverifiable`, so defer the restart. Do not allow new work between that census -and the stop; Orca does not yet provide an atomic census-and-stop fence. - -Rolling back is the case that needs care — see [Roll back](#roll-back). - -### Record the version you deploy - -The bundled CLI launcher prints the Orca build with `orca-ide --version`. For an -extracted deployment, that launcher is -`squashfs-root/resources/bin/orca-ide`; deb/rpm installs and CLI registration put -it on `PATH`. Do not use `orca-linux.AppImage --version` for this audit because -Electron owns the direct binary's version flags and may report its own runtime -version. For an AppImage service, choose a release tag explicitly and record it -next to the binary. The steps below keep that record in `/opt/orca/VERSION`. - -### Upgrade steps - -Never download straight onto `/opt/orca/orca-linux.AppImage`. The AppImage is -FUSE-mounted, so overwriting it in place while the service runs can crash or -corrupt the live process — and even with the service stopped, a failed or partial -download would clobber the working binary. Instead download to a temporary name -on the same filesystem, verify it, then swap it in with an atomic rename. - -Check capacity before starting: - -```bash -sudo chown root:root /opt/orca -sudo chmod 755 /opt/orca -sudo test ! -L /opt/orca/orca-linux.AppImage -sudo chown root:root /opt/orca/orca-linux.AppImage -sudo chmod 755 /opt/orca/orca-linux.AppImage -# Clear predictable staging names left by an older attempt after locking the directory -sudo rm -f /opt/orca/orca-linux.AppImage.new /opt/orca/VERSION.new \ - /opt/orca/orca-linux.AppImage.recovering /opt/orca/VERSION.recovering -sudo du -sh /home/orca/.config -df -h /opt/orca /home/orca -``` - -`/opt/orca` needs room for the compressed Orca profile archive, the staged -build, and the rollback binary. A rollback extracts the old profile and preserves -the post-upgrade Orca profile directories, so `/home` needs room for both copies. - -Run the following block as one Bash script so its fail-fast and recovery traps -remain active for the whole operation: - -```bash -set -euo pipefail - -# Replace this example with the release tag you intend to deploy -ORCA_VERSION=v1.4.147 - -# Select the release asset on the server where Orca runs -case "$(uname -m)" in - x86_64) - ORCA_ASSET=orca-linux.AppImage - ORCA_FILE_MACHINE=x86-64 - ;; - aarch64 | arm64) - ORCA_ASSET=orca-linux-arm64.AppImage - ORCA_FILE_MACHINE='ARM aarch64' - ;; - *) - echo "Unsupported architecture: $(uname -m)" >&2 - exit 1 - ;; -esac - -ORCA_ROLLBACK_NEW= -ORCA_ROLLBACK= -ORCA_SERVICE_STOPPED=0 -ORCA_BINARY_PROMOTED=0 -recover_failed_upgrade() { - exit_status=$? - trap - EXIT - set +e - if ((exit_status != 0)); then - sudo rm -f /opt/orca/orca-linux.AppImage.new /opt/orca/VERSION.new \ - /opt/orca/orca-linux.AppImage.recovering /opt/orca/VERSION.recovering - fi - if ((exit_status != 0)) && [[ -n "$ORCA_ROLLBACK_NEW" ]] && \ - sudo test -d "$ORCA_ROLLBACK_NEW"; then - sudo rm -rf -- "$ORCA_ROLLBACK_NEW" - fi - if ((exit_status != 0 && ORCA_SERVICE_STOPPED)); then - recovery_ok=1 - if ((ORCA_BINARY_PROMOTED)); then - if ! sudo cp -a "$ORCA_ROLLBACK/orca-linux.AppImage" \ - /opt/orca/orca-linux.AppImage.recovering || \ - ! sudo mv -f /opt/orca/orca-linux.AppImage.recovering \ - /opt/orca/orca-linux.AppImage; then - recovery_ok=0 - fi - if sudo test -f "$ORCA_ROLLBACK/VERSION"; then - if ! sudo cp -a "$ORCA_ROLLBACK/VERSION" /opt/orca/VERSION.recovering || \ - ! sudo mv -f /opt/orca/VERSION.recovering /opt/orca/VERSION; then - recovery_ok=0 - fi - elif ! sudo rm -f /opt/orca/VERSION; then - recovery_ok=0 - fi - fi - sudo rm -f /opt/orca/orca-linux.AppImage.recovering \ - /opt/orca/VERSION.recovering - if ((recovery_ok)); then - # A tripped StartLimitBurst refuses a plain start - sudo systemctl reset-failed orca-serve.service || true - sudo systemctl start orca-serve.service || true - else - echo 'Upgrade recovery failed; service remains stopped' >&2 - fi - fi - exit "$exit_status" -} -trap recover_failed_upgrade EXIT - -# 1. Stage and verify the new build while the server stays online -sudo curl -fL --retry 3 "https://github.com/stablyai/orca/releases/download/${ORCA_VERSION}/${ORCA_ASSET}" \ - -o /opt/orca/orca-linux.AppImage.new -sudo chown root:root /opt/orca/orca-linux.AppImage.new -sudo chmod 755 /opt/orca/orca-linux.AppImage.new - -# Both checks must match; either grep stops this fail-fast block otherwise -ORCA_FILE_INFO=$(LC_ALL=C file /opt/orca/orca-linux.AppImage.new) -grep 'ELF .* executable' <<<"$ORCA_FILE_INFO" -grep -F "$ORCA_FILE_MACHINE" <<<"$ORCA_FILE_INFO" - -# 2. Assemble the prior binary and version in a root-only rollback bundle -ORCA_ROLLBACK_BASE=/opt/orca/orca-rollback-$(date +%F-%H%M%S-%N) -ORCA_ROLLBACK_NEW=${ORCA_ROLLBACK_BASE}.new -ORCA_ROLLBACK=${ORCA_ROLLBACK_BASE}.ready -sudo install -d -m 700 "$ORCA_ROLLBACK_NEW" -sudo cp -a /opt/orca/orca-linux.AppImage "$ORCA_ROLLBACK_NEW/orca-linux.AppImage" -if sudo test -f /opt/orca/VERSION; then - sudo cp -a /opt/orca/VERSION "$ORCA_ROLLBACK_NEW/VERSION" -fi - -# Stage the new version record before the stop window -printf '%s\n' "$ORCA_VERSION" | sudo tee /opt/orca/VERSION.new >/dev/null -sudo chown root:root /opt/orca/VERSION.new -sudo chmod 644 /opt/orca/VERSION.new - -# 3. Stop the server so the profile backup is consistent -ORCA_SERVICE_STOPPED=1 -sudo systemctl stop orca-serve.service - -# Add only Orca-owned profile directories, then publish the complete bundle -ORCA_PROFILE_DIRS=() -for profile_dir in orca Orca; do - if sudo test -L "/home/orca/.config/$profile_dir"; then - echo "Refusing symlinked Orca profile: /home/orca/.config/$profile_dir" >&2 - exit 1 - fi - if sudo test -d "/home/orca/.config/$profile_dir"; then - if [[ "$profile_dir" == Orca ]] && \ - sudo test /home/orca/.config/orca -ef /home/orca/.config/Orca; then - continue - fi - ORCA_PROFILE_DIRS+=("$profile_dir") - fi -done -if ((${#ORCA_PROFILE_DIRS[@]} == 0)); then - echo 'No Orca profile directory found under /home/orca/.config' >&2 - exit 1 -fi -sudo tar czf "$ORCA_ROLLBACK_NEW/profile.tgz" \ - -C /home/orca/.config "${ORCA_PROFILE_DIRS[@]}" -sudo chmod 600 "$ORCA_ROLLBACK_NEW/profile.tgz" -sudo mv "$ORCA_ROLLBACK_NEW" "$ORCA_ROLLBACK" - -# 4. Atomically replace the binary and version record, then start -ORCA_BINARY_PROMOTED=1 -sudo mv -f /opt/orca/orca-linux.AppImage.new /opt/orca/orca-linux.AppImage -sudo mv -f /opt/orca/VERSION.new /opt/orca/VERSION -# Clears a start-limit hit left by the version being replaced -sudo systemctl reset-failed orca-serve.service -sudo systemctl start orca-serve.service -ORCA_SERVICE_STOPPED=0 -trap - EXIT -``` - -The profile archive created in step 3 captures both Orca profile directory names -when present without rewinding unrelated tools under `/home/orca/.config`. The -`.ready` suffix is published only after the prior binary, version record, and -profile archive are complete. If you run the managed Xvfb unit, only -`orca-serve.service` needs restarting — leave `orca-xvfb.service` running. - -### Verify - -```bash -sudo journalctl -u orca-serve.service -f -``` - -A healthy start prints one `Orca server ready` block with the actual bound and -advertised endpoints. Verify those values rather than assuming the configured -port, because a collision can select a fallback port. -Confirm a client reconnects before you discard the backup. The timestamped -rollback bundles are not pruned automatically. After the new version satisfies -your retention policy, select and inspect the newest complete bundle before -removing it: - -```bash -shopt -s nullglob -ORCA_ROLLBACK_SETS=(/opt/orca/orca-rollback-*.ready) -((${#ORCA_ROLLBACK_SETS[@]} > 0)) -ORCA_ROLLBACK=${ORCA_ROLLBACK_SETS[${#ORCA_ROLLBACK_SETS[@]} - 1]} -printf 'Removing rollback bundle: %s\n' "$ORCA_ROLLBACK" -sudo test -d "$ORCA_ROLLBACK" -sudo rm -rf -- "$ORCA_ROLLBACK" -``` - -Each `.ready` directory is a self-contained rollback generation; never combine -files from different bundles. - -### Roll back - -A rollback is **not** binary-only safe. New builds save profile state in -SQLite, while a JSON-only build would read a stale legacy `orca-data.json`. -The rollback below restores the complete pre-upgrade profile directories from -step 3 and swaps the binary back. It deliberately returns to the pre-upgrade -state rather than preserving changes made since the upgrade. Run this block -as one Bash script: - -To preserve the latest state when moving to a JSON-only build, stop the service -and use the newer build's CLI to run -`orca profile state rollback --latest-json --profile-id ` for **every** -profile the older build may open, then install the older binary manually. -Without `--profile-id`, the command prepares only the active profile. This -explicit handoff archives the SQLite authority and publishes the current JSON; -ordinary exports write revisioned `.sqlite-export.N.json` files and do not -prepare a downgrade. Orca's updater refuses JSON-only targets automatically. - -```bash -set -euo pipefail - -# Select and validate one complete generation before taking the service offline -shopt -s nullglob -ORCA_ROLLBACK_SETS=(/opt/orca/orca-rollback-*.ready) -((${#ORCA_ROLLBACK_SETS[@]} > 0)) -ORCA_ROLLBACK=${ORCA_ROLLBACK_SETS[${#ORCA_ROLLBACK_SETS[@]} - 1]} -sudo test -f "$ORCA_ROLLBACK/orca-linux.AppImage" -sudo tar tzf "$ORCA_ROLLBACK/profile.tgz" >/dev/null - -# Extract and validate the old profile while the current server stays online -sudo test ! -L /home -ORCA_HOME_OWNER=$(sudo stat -c %u /home) -ORCA_HOME_MODE=$(sudo stat -c %a /home) -if [[ "$ORCA_HOME_OWNER" != 0 ]] || ((8#$ORCA_HOME_MODE & 0022)) || \ - sudo -u orca test -w /home; then - echo 'Refusing rollback because /home is not root-controlled' >&2 - exit 1 -fi -ORCA_RESTORE=$(sudo mktemp -d /home/.orca-restore.XXXXXX) -ORCA_SERVICE_STOPPED=0 -ORCA_MOVED_CURRENT_DIRS=() -ORCA_INSTALLED_RESTORE_DIRS=() -ORCA_CURRENT_BINARY_MOVED=0 -ORCA_CURRENT_VERSION_MOVED=0 -ORCA_VERSION_REPLACEMENT_STARTED=0 -ORCA_POST_UPGRADE= -ORCA_ROLLBACK_BINARY_STAGED= -ORCA_ROLLBACK_VERSION_STAGED= -ORCA_ROLLBACK_HAS_VERSION=0 -restart_after_rollback_error() { - exit_status=$? - trap - EXIT - set +e - if ((exit_status != 0 && ORCA_SERVICE_STOPPED)); then - recovery_ok=1 - if ((${#ORCA_INSTALLED_RESTORE_DIRS[@]})); then - for profile_dir in "${ORCA_INSTALLED_RESTORE_DIRS[@]}"; do - if sudo test -d "/home/orca/.config/$profile_dir"; then - if ! sudo mv "/home/orca/.config/$profile_dir" \ - "$ORCA_RESTORE/$profile_dir.failed"; then - recovery_ok=0 - fi - fi - done - fi - if ((${#ORCA_MOVED_CURRENT_DIRS[@]})); then - for profile_dir in "${ORCA_MOVED_CURRENT_DIRS[@]}"; do - if sudo test -d "$ORCA_POST_UPGRADE/$profile_dir"; then - if ! sudo mv "$ORCA_POST_UPGRADE/$profile_dir" /home/orca/.config/; then - recovery_ok=0 - fi - elif ! sudo test -d "/home/orca/.config/$profile_dir"; then - recovery_ok=0 - fi - done - fi - if [[ -n "$ORCA_POST_UPGRADE" ]]; then - sudo rmdir "$ORCA_POST_UPGRADE" 2>/dev/null || true - fi - if ((ORCA_CURRENT_BINARY_MOVED)); then - if sudo test -f "$ORCA_CURRENT_BINARY"; then - if ! sudo mv -f "$ORCA_CURRENT_BINARY" /opt/orca/orca-linux.AppImage; then - recovery_ok=0 - fi - elif ! sudo test -f /opt/orca/orca-linux.AppImage; then - recovery_ok=0 - fi - fi - if ((ORCA_CURRENT_VERSION_MOVED)); then - if sudo test -f "$ORCA_CURRENT_VERSION"; then - if ! sudo mv -f "$ORCA_CURRENT_VERSION" /opt/orca/VERSION; then - recovery_ok=0 - fi - elif ! sudo test -f /opt/orca/VERSION; then - recovery_ok=0 - fi - elif ((ORCA_VERSION_REPLACEMENT_STARTED)); then - if ! sudo rm -f /opt/orca/VERSION; then - recovery_ok=0 - fi - fi - if ((recovery_ok)); then - # A tripped StartLimitBurst refuses a plain start - sudo systemctl reset-failed orca-serve.service || true - sudo systemctl start orca-serve.service || true - else - echo 'Rollback recovery failed; service remains stopped' >&2 - fi - fi - if [[ -n "$ORCA_ROLLBACK_BINARY_STAGED" ]]; then - sudo rm -f -- "$ORCA_ROLLBACK_BINARY_STAGED" - fi - if [[ -n "$ORCA_ROLLBACK_VERSION_STAGED" ]]; then - sudo rm -f -- "$ORCA_ROLLBACK_VERSION_STAGED" - fi - sudo rm -rf -- "$ORCA_RESTORE" - exit "$exit_status" -} -trap restart_after_rollback_error EXIT - -if [[ "$(sudo stat -c %d "$ORCA_RESTORE")" != \ - "$(sudo stat -c %d /home/orca/.config)" ]]; then - echo 'Refusing rollback because staging and the Orca profile are on different filesystems' >&2 - exit 1 -fi -sudo tar xzf "$ORCA_ROLLBACK/profile.tgz" -C "$ORCA_RESTORE" -ORCA_RESTORE_DIRS=() -for profile_dir in orca Orca; do - if sudo test -L "$ORCA_RESTORE/$profile_dir"; then - echo "Rollback bundle contains a symlinked profile: $profile_dir" >&2 - exit 1 - fi - if sudo test -d "$ORCA_RESTORE/$profile_dir"; then - if [[ "$profile_dir" == Orca ]] && \ - sudo test "$ORCA_RESTORE/orca" -ef "$ORCA_RESTORE/Orca"; then - continue - fi - ORCA_RESTORE_DIRS+=("$profile_dir") - fi -done -if ((${#ORCA_RESTORE_DIRS[@]} == 0)); then - echo "Rollback bundle has no Orca profile directories: $ORCA_ROLLBACK" >&2 - exit 1 -fi -for profile_dir in "${ORCA_RESTORE_DIRS[@]}"; do - sudo chown -R orca:orca "$ORCA_RESTORE/$profile_dir" -done - -ORCA_ROLLBACK_STAMP=$(date +%F-%H%M%S-%N) -ORCA_ROLLBACK_BINARY_STAGED=/opt/orca/orca-linux.AppImage.rollback-staged-$ORCA_ROLLBACK_STAMP -sudo cp -a "$ORCA_ROLLBACK/orca-linux.AppImage" "$ORCA_ROLLBACK_BINARY_STAGED" -if sudo test -f "$ORCA_ROLLBACK/VERSION"; then - ORCA_ROLLBACK_HAS_VERSION=1 - ORCA_ROLLBACK_VERSION_STAGED=/opt/orca/VERSION.rollback-staged-$ORCA_ROLLBACK_STAMP - sudo cp -a "$ORCA_ROLLBACK/VERSION" "$ORCA_ROLLBACK_VERSION_STAGED" -fi - -ORCA_SERVICE_STOPPED=1 -sudo systemctl stop orca-serve.service - -# Preserve and replace only Orca-owned profile directories -ORCA_CURRENT_DIRS=() -for profile_dir in orca Orca; do - if sudo test -L "/home/orca/.config/$profile_dir"; then - echo "Refusing symlinked Orca profile: /home/orca/.config/$profile_dir" >&2 - exit 1 - fi - if sudo test -d "/home/orca/.config/$profile_dir"; then - if [[ "$profile_dir" == Orca ]] && \ - sudo test /home/orca/.config/orca -ef /home/orca/.config/Orca; then - continue - fi - ORCA_CURRENT_DIRS+=("$profile_dir") - fi -done -ORCA_POST_UPGRADE=/home/orca/.config/orca-rollback-$ORCA_ROLLBACK_STAMP -sudo install -d -o orca -g orca -m 700 "$ORCA_POST_UPGRADE" -if ((${#ORCA_CURRENT_DIRS[@]})); then - for profile_dir in "${ORCA_CURRENT_DIRS[@]}"; do - ORCA_MOVED_CURRENT_DIRS+=("$profile_dir") - sudo mv "/home/orca/.config/$profile_dir" "$ORCA_POST_UPGRADE/" - done -fi -for profile_dir in "${ORCA_RESTORE_DIRS[@]}"; do - ORCA_INSTALLED_RESTORE_DIRS+=("$profile_dir") - sudo mv "$ORCA_RESTORE/$profile_dir" /home/orca/.config/ -done - -ORCA_CURRENT_BINARY=/opt/orca/orca-linux.AppImage.rollback-current-$ORCA_ROLLBACK_STAMP -ORCA_CURRENT_BINARY_MOVED=1 -sudo mv /opt/orca/orca-linux.AppImage "$ORCA_CURRENT_BINARY" -sudo mv -f "$ORCA_ROLLBACK_BINARY_STAGED" /opt/orca/orca-linux.AppImage - -ORCA_CURRENT_VERSION=/opt/orca/VERSION.rollback-current-$ORCA_ROLLBACK_STAMP -if sudo test -f /opt/orca/VERSION; then - ORCA_CURRENT_VERSION_MOVED=1 - sudo mv /opt/orca/VERSION "$ORCA_CURRENT_VERSION" -fi -ORCA_VERSION_REPLACEMENT_STARTED=1 -if ((ORCA_ROLLBACK_HAS_VERSION)); then - sudo mv -f "$ORCA_ROLLBACK_VERSION_STAGED" /opt/orca/VERSION -else - sudo rm -f /opt/orca/VERSION -fi -# The crash-looping build you are rolling back from tripped StartLimitBurst -sudo systemctl reset-failed orca-serve.service -sudo systemctl start orca-serve.service -ORCA_SERVICE_STOPPED=0 -sudo rm -rf -- "$ORCA_RESTORE" -trap - EXIT -``` - -Restoring the complete backup is required for this rollback procedure: swapping -only the binary leaves the newer SQLite authority and stale legacy JSON in -place. Keep the pre-upgrade backup until the new version is proven -on your host. The `orca-rollback-*` directory inside `.config` is also retained -deliberately. The post-upgrade binary and version record are retained in -`/opt/orca` with the same `rollback-current-` suffix. Inspect these -artifacts and remove them according to your retention policy after the rollback -is resolved. - -## Installing Agent Skills Without A Desktop - -Orca's agent skills (CLI usage, orchestration, computer use, etc.) are normally -installed from Orca Settings, which pre-fills an `npx skills add ... --global` -command in a terminal for you to run. A headless host has no Settings UI, so -use `orca skills install` instead: - -```bash -orca skills install # list installable skills -orca skills install --skill orca-cli --skill orchestration # install globally (default) -orca skills install --skill orca-cli --local # install into the current project only -orca skills install --all # install every bundled skill -orca skills install --all --dry-run # print the npx command without running it -``` - -This resolves the same `npx skills add --skill ...` command -Settings would show you (adding `--global` unless `--local` is passed), then -runs it and forwards its output and exit code. It requires `node`/`npx` on the -host; it does not need a running Orca runtime. - -Unlike the command Settings shows, the spawned one adds `npx --yes` and `-y`. -Without them the `skills` CLI opens an interactive agent picker and blocks -forever on any allocated TTY — which includes a normal `ssh` session. Use -`--dry-run` to see the exact command that will run. - -Settings keeps that picker deliberately, because choosing which agents get a -skill is a real decision. A headless run cannot answer it, so instead of dropping -the choice Orca makes it explicitly: it passes an `--agent` list built from the -coding agents it detects on the host, plus the shared `.agents/skills` directory -it reads itself. Left to decide on its own with no agent detected, the `skills` -CLI installs into all ~75 agents it knows and leaves a config directory for each. -Override the targets yourself, or narrow to the shared directory alone: - -```bash -orca skills install --skill orca-cli --agent claude-code,codex -orca skills install --skill orca-cli --agent universal -``` - -If Orca detects no agent at all, `orca skills install` stops and asks for -`--agent` rather than guessing. - -To refresh already-installed skills, `orca skills update` mirrors the same -selection flags (`--skill`, `--all`, `--local`, `--dry-run`) and resolves to -`npx skills update ` with a matching scope flag — `--global`, or -`--project` when you pass `--local`: - -```bash -orca skills update --all # update every bundled skill globally -orca skills update --skill orca-cli --dry-run # print the npx command without running it -``` - -`orca skills update` only refreshes skills that are already installed — it exits -0 without doing anything for a skill that is missing, so install it first. More -generally, a 0 exit means the `skills` CLI ran without erroring, not that it -wrote anything; read its output to confirm what changed. - -`--json` covers the skill listing and `--dry-run`. A real run streams the -`skills` CLI's own non-JSON output and rejects `--json`. - -Both commands install onto the machine that runs them. In an Orca SSH workspace -or the WSL bridge the `orca` shim forwards commands to the Orca host, so they -refuse to run there and print the command to run on the machine you want. - -## Troubleshooting - -- `dlopen(): error loading libfuse.so.2`: install `libfuse2`. -- `Missing X server or $DISPLAY`: install `xvfb`, or start the managed Xvfb - service and set `DISPLAY=:99`. -- `[serve] Xvfb failed to start` or `[serve] Could not start Xvfb`: confirm - `command -v Xvfb` and that it is on the service `PATH`. -- GPU or DRI warnings on a VPS: keep `LIBGL_ALWAYS_SOFTWARE=1` in the service - environment. -- Chromium sandbox errors: confirm the service is running as the non-root - `orca` user and that `/opt/orca` is readable by that user, including - `/opt/orca/squashfs-root` if you extracted the AppImage. -- Clients cannot connect: make sure `--pairing-address` is an address reachable - from the client, and make sure firewalls allow the selected `--port`. -- Journal shows `Another Orca instance is already running for this userData -profile` and the unit exits `3`: another process already owns the profile, so - `RestartPreventExitStatus=3` leaves the unit `failed` on purpose. Find the - owner with `systemctl status orca-serve` and `pgrep -af orca`. Stop it (or - keep it and leave the unit down), then run - `sudo systemctl reset-failed orca-serve && sudo systemctl start orca-serve` — - `reset-failed` clears the failed state and any start-limit counter. If no owner - exists, the lock is stale (Chromium recorded a pid that - has since been reused): remove `SingletonLock` and `SingletonSocket` from the - userData directory and start again. If an earlier crash-loop already leaked - AppImage mounts, list them with `findmnt -rn -t fuse.orca-linux.AppImage` and - release only the ones with no live owner using `fusermount -uz ` (or - `umount -l `), leaving the running instance's mount alone. -- Service crash-loops right after an upgrade: use [Roll back](#roll-back) with - the pre-upgrade `.ready` bundle. Do not rerun the upgrade first; doing so would - make the crashing version the next rollback binary. The loop trips - `StartLimitBurst`, so any manual `systemctl start` outside that script needs - `sudo systemctl reset-failed orca-serve.service` first. -- Diagnosing other missing libraries: extract the AppImage without launching it - with `./orca-linux.AppImage --appimage-extract`, then run - `ldd squashfs-root/orca-ide` to list any shared libraries the host is missing. - The Electron binary is `orca-ide`, not `orca`; `ldd` on a path that does not - exist prints nothing and exits cleanly, which reads as a clean result in - exactly the situation where you are hunting a missing library. diff --git a/docs/reference/ime-regression-checklist.md b/docs/reference/ime-regression-checklist.md deleted file mode 100644 index bad36127dd5..00000000000 --- a/docs/reference/ime-regression-checklist.md +++ /dev/null @@ -1,146 +0,0 @@ -# IME Regression Checklist - -## Follow-up issue ledger - -These reports define durable acceptance contracts, not only the symptoms from -one machine. - -| Issue | Root cause | Ownership invariant | Required evidence | -| ----------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| [#16911](https://github.com/stablyai/orca/issues/16911) Native Chat preedit overwritten by streaming or attachment settlement | React reconciliation or an asynchronously resolved attachment writes an application draft while the browser owns the composing textarea; duplicate settlement can then re-adopt stale DOM. | From `compositionstart` until the first `compositionend` or blur, the browser owns the textarea. Resolved paths queue at the shared semantic sink; settlement adopts the browser DOM before flushing once. | Repeated stale streaming rerenders preserve the element and preedit; idle external drafts still synchronize; concurrent SSH completions preserve completion order and duplicates; both settlement orders run once; disable discards queued work; blur does not steal focus; native composition commits once. | -| [#16949](https://github.com/stablyai/orca/issues/16949) terminal preedit has no visible cursor | The opaque composition overlay covers the renderer cursor; at the final cell an over-wide inline preedit can also place its caret beyond the clipped screen. | The existing xterm `CompositionHelper` owns a visible caret after the preedit and before any row remainder; final-cell composition end-aligns within the screen while mid-line composition stays left-anchored. | Start, update, arbitrary-width final-cell containment, mid-line remainder placement, cleanup, and update-without-start are covered; preview and normal terminals inherit the live cursor theme. | -| [#16950](https://github.com/stablyai/orca/issues/16950) typing diagnostic records no CJK samples | The probe observes echoing keydowns but not reconciled composition commits, then guesses which queued input owns opaque TUI output. | A reconciled composition is observed even when `compositionend.data` is empty; only an isolated input enters exact percentiles, while overlap or a dropped-input gap produces one aggregate ambiguous burst. | Recorded Linux IBus empty-data commit, isolated direct and IME samples, mixed-source ambiguity, timeout/cap gaps, UTF-8 output bytes, and stop/drain cleanup are covered. | -| [#17104](https://github.com/stablyai/orca/issues/17104) Korean preedit repeats the Codex placeholder | Generic xterm row-tail reproduction exposed an application-semantic Codex or Claude composer placeholder that presentation style cannot identify safely. | Xterm always preserves generic covered row text. Orca's existing structural composer classifier masks only a verified placeholder during the exact active composition session; repaint reclassification runs only while composing, and end, blur, or disposal clears ownership, class, and listeners. Arbitrary dim output and shell lookalikes remain visible. | Codex prompt/footer and Claude prompt/frame classification, arbitrary all-dim and shell-lookalike negatives, repaint entry and exit, end/blur/disposal cleanup, and rendered Electron proof at cursor column 2 preserving generic row text are covered. | - -## Preedit cell advances (#19315) - -Single-codepoint CJK graphemes use the active Unicode provider's cell width and -measured font advance. Ordinary inline spans preserve browser bidi and baseline -layout; equal corrections share a run. Keep glyphs unscaled and the underline, -caret, and candidate textarea aligned with the rendered preedit. Appending ASCII, -emoji, or another script must not change an existing CJK prefix's correction. -Combining sequences, emoji, other scripts, and whitespace retain native shaping. -Font loading, typography changes, and renderer metric changes must update an open -composition; row-tail repaints preserve its unchanged nodes. - -Cold font measurements and styled runs share a fixed work budget. Repeated CJK -can remain one corrected run; after the budget is exhausted, the remaining text -keeps its native advance. This deliberately leaves the original spacing mismatch -in the tail of unusually varied long compositions, without switching the prefix -back to native spacing or rebuilding thousands of spans. - -`terminal-ime-xterm-preedit-cell-grid.test.ts` covers text preservation, native -clusters, lifecycle, and bounded work. `terminal-ime-preedit-cell-grid.spec.ts` -checks rendered glyph origins, caret/textarea geometry, underlines, font changes, -and native shaping at DPR 1, 1.25, and 2 with WebGL on/off. -`terminal-ime-preedit-continuity.spec.ts` covers mixed suffixes and budget crossings. -These checks use Chromium composition through CDP; they do not replace native OS -IME evidence. - -## Bounded-state and ownership contracts - -Every transient collection and ownership tracker must have an explicit lifetime and bound: - -- Native Chat uses `NATIVE_FILE_DROP_MAX_PATHS` (`256`). If a resolved completion would cross the cap, the whole batch is rejected atomically and the overflow notice remains visible through settlement; accepted paths keep order and duplicates. The queue is cleared before re-entry and on disable or pane-owner remount. -- The terminal placeholder mask tracks one scalar `activeSessionId` because xterm renders one composition view. A newer start supersedes an older one, a stale end cannot clear the latest owner, and blur or disposal clears it. The composition route keeps its per-ID reference-counted map intentionally for transport ownership; it is not replaced by the scalar. -- Typing diagnostics cap pending and ignored echo candidates at `MAX_PENDING_ECHO_CANDIDATES` (`64`), cap pending user-input signals, drain timed-out candidates, and clear all series on pane detach. Overflow becomes an explicitly ambiguous burst rather than an arbitrary attribution. - -## Native Chat asynchronous attachment settlement - -Attachment resolution is an external semantic write, including local file -selection, pasted-image temp saves, and SSH uploads. While composition is active, -it must not replace the browser-owned textarea value. - -- Start two concurrent SSH uploads, resolve the second first, and return one path - twice. Preserve completion order and both duplicates. -- On the first settlement event, adopt the browser DOM before flushing queued - paths. Exercise `compositionend` then blur and blur then `compositionend` in - one React batch; both orders must adopt and flush exactly once. -- If the composer becomes disabled before an upload resolves or before the queue - flushes, discard that result. -- If `compositionend` is omitted, blur performs the same one-time settlement - without focusing the textarea or stealing focus back. - -## Cross-platform verification - -Synthetic DOM events prove Orca's event and rendering contracts, but they do -not exercise the operating system's input method. Changes must also cover: - -| Environment | Native evidence | -| ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| macOS | A native Korean 2-set composition in Native Chat and a terminal; preedit survives external renders, the caret remains visible, and commit occurs once. | -| Windows | Microsoft Korean IME over an untouched Codex placeholder; the preedit is the only visible text, the placeholder returns after cancel, and ordinary mid-line content remains visible. | -| Linux / SSH | IBus Hangul with an SSH-hosted PTY; an empty-data `compositionend` still produces one diagnostic sample and one committed syllable. | - -For remote evidence, `live` means the owning host reported the current -verification session or process identity. `exited` requires positive -host-owned evidence that the same identity terminated or is absent. Any -transport failure, stale identity, timeout, or inability to ask the owning host -makes the result `unverifiable`; it is never evidence that the composition or -PTY process exited. - -## Code elegance gate - -Each fix must pass all of these checks: - -- Reuse the component that already owns the state or overlay; do not install a - second composition state machine. -- Make browser, renderer, and PTY ownership boundaries explicit. Provisional - text must not leak into committed state or PTY input. -- Route local selection, pasted-image temp saves, and SSH upload results through - one resolved-attachment sink; queue semantic paths, never whole draft snapshots. -- Keep correctness changes separate from unrelated micro-optimizations. -- Use bounded per-composition state and work. Dispose every listener, timer, - observer, and DOM node with its owner. -- Preserve ordinary Latin input, mixed styled terminal content, local and SSH - PTYs, preview terminals, and folder workspaces with paired negative tests. -- Keep platform quirks behind event contracts or runtime platform checks; do - not branch on an IME vendor, language, or terminal agent name. -- Treat the canonical xterm source patch as the only hand-edited source, then - regenerate its bundle patch and lockfile together. -- Prefer deterministic replay or state-transition tests. Native evidence is a - second layer, never a substitute for regression coverage. - -## Enter in application text fields (#25035) - -Use `Input`, `Textarea`, or `CommandInput` for styled fields. Existing unstyled -fields with keyboard actions use `ImeInput` / `ImeTextarea` from -`lib/ime-text-field.tsx`; those preserve the DOM element, styles, refs, and -composition callbacks. They share `useImeEnterGestureOwnership` and keep -IME-owned keys out of both field actions and bubbling form/menu shortcuts. -Overlay primitives also reject IME-marked Escape in document capture, where -field-level propagation guards cannot intercept dismissal. -Do not add a second tracker at a call site already using a guarded field. -Native Chat and the File Explorer inline name field retain their existing -trackers because they also own specialized composition or element lifetimes. - -Required cases: - -- `isComposing`, `keyCode: 229` without `isComposing`, and `Process/229` must - never submit, choose a suggestion, or dismiss the field. -- The unmarked Enter redispatch stays owned on either side of keyup, including - a `Process/229` release. A - subsequent ordinary typing/navigation key ends that carry immediately; - hidden renderers may defer animation frames, and typing a filename suffix - must not cause the next deliberate Enter to disappear. -- Composition callbacks, blur, refs, and keyed remount cleanup still work. - Normal Enter, modifier submits, and Shift+Enter newlines remain available. -- Test the actual shared field when a consumer delegates IME handling to it; - a mock that replaces `CommandInput` with a raw input removes the protection. - -`ime-text-field.test.tsx` covers primitives, raw fields, parent handlers, and -command selection. File Explorer component tests cover all three operations -and input replacement. `file-explorer-ime-enter.spec.ts` drives Chromium -composition in New File, New Folder, and Rename in a folder workspace, then -checks the complete name in the Explorer and on disk, with both continued typing -and a redispatch followed by deliberate Enter. Overlay tests cover IME Escape -and ordinary dismissal; Markdown tests preserve an unmarked save shortcut while -composition state lingers. These are CDP event -contracts, not native OS keyboard evidence. - -The audit also covers settings and title fields, issue/review creation and -pickers, comments and annotations, search fields, Native Chat questions, -notebook execution shortcuts, and Markdown menu handlers. Terminal input keeps -its existing xterm/PTY ownership; mobile native fields use `onSubmitEditing` -instead of desktop DOM keydown actions. Remote workspaces use the same renderer -fields; file-operation routing and mixed-version wire contracts are unchanged. diff --git a/docs/reference/linux-glibc-compatibility.md b/docs/reference/linux-glibc-compatibility.md deleted file mode 100644 index f1dcdda4710..00000000000 --- a/docs/reference/linux-glibc-compatibility.md +++ /dev/null @@ -1,182 +0,0 @@ -# Linux glibc Compatibility - -Orca's Linux builds target **stock Ubuntu 20.04 and newer** — glibc 2.31 and -libstdc++ `GLIBCXX_3.4.28` (also Debian 11, RHEL 9), on both x64 and arm64. -Packaging enforces this floor automatically; keep it in mind when adding or -upgrading native dependencies. (The optional speech feature is the one -exception — see below.) - -## Local package build prerequisites - -`pnpm run build:linux` produces AppImage, deb, and RPM artifacts. The RPM target -requires `rpmbuild` on `PATH`; install `rpm` on Ubuntu/Debian, `rpm-build` on -Fedora/RHEL, or `rpm` through Homebrew on macOS, then verify it with -`rpmbuild --version` before packaging. Cross-host builds have the same -requirement. - -## Why this needs attention - -A native module (`.node`) links against the glibc of the machine that compiled -it. Our release CI compiles node-pty from source on GitHub's `ubuntu-latest` -runner, whose glibc rises over time as the image is bumped. A binary compiled on -a newer glibc can reference symbol versions that do not exist on an older target, -and the dynamic loader then refuses to load it: - -``` -/lib/x86_64-linux-gnu/libc.so.6: version `GLIBC_2.34' not found (required by .../pty.node) -``` - -Because the Orca main process loads node-pty at startup, that failure crashes the -whole app before a window appears — this is exactly what shipped in v1.4.150 and -broke launch on Ubuntu 20.04 ([#9902](https://github.com/stablyai/orca/issues/9902)). - -The specific trap is glibc's 2.32–2.34 "libpthread/libutil merge", which moved -several long-stable functions into libc under brand-new symbol versions: - -| Symbol | New version | node-pty use | -| ----------------- | ------------ | ----------------------- | -| `pthread_sigmask` | `GLIBC_2.32` | reset child signal mask | -| `openpty` | `GLIBC_2.34` | allocate the pty | -| `forkpty` | `GLIBC_2.34` | fork the shell | - -Electron itself (glibc 2.25) and the other bundled native modules -(`sherpa-onnx`, `@parcel/watcher`, both prebuilt on old glibc) stay well under -the floor, so node-pty was the sole blocker. - -## How we keep the floor - -**1. Pin the relocated symbols (the fix).** -[`config/patches/node-pty@1.1.0.patch`](../../config/patches/node-pty@1.1.0.patch) -adds a `.symver` shim in `src/unix/pty.cc` that binds `openpty`, `forkpty`, and -`pthread_sigmask` to their pre-merge version node — `GLIBC_2.2.5` on x64, -`GLIBC_2.17` on arm64 (each architecture's baseline glibc). glibc still ships -those as compatibility aliases, so the reference resolves on both new build hosts -and old targets. - -The catch: gcc defaults to `--as-needed` and, since the pinned symbols now -resolve from libc's compat aliases at build time, it drops `libutil`/`libpthread` -from `DT_NEEDED`. On the target those libraries are where the symbols actually -live, so the patch's `binding.gyp` `ldflags` force -`-Wl,--no-as-needed,-l:libutil.so.1,-l:libpthread.so.0` back into `DT_NEEDED`. -The shim is guarded by `#if defined(__linux__)`; macOS and Windows are untouched. - -**2. Gate packaging (the regression guard).** -[`config/scripts/verify-linux-glibc-floor.cjs`](../../config/scripts/verify-linux-glibc-floor.cjs) -runs in the electron-builder `afterPack` hook for Linux. It reads every bundled -native binary's version needs (`objdump -p` "Version References" — the -authoritative load-time list, which also captures symbol-less markers like -`GLIBC_ABI_DT_RELR`) and fails the build if any strong `GLIBC_`/`GLIBCXX_`/ -`CXXABI_` node is newer than stock Ubuntu 20.04 provides, naming the file and the -offending node. Weak needs are ignored (the loader tolerates them). It also -asserts the flip side of the `.symver` fix: any binary that imports -`openpty`/`forkpty` must keep `libutil.so.1` in `DT_NEEDED` — otherwise the -pinned `openpty@GLIBC_2.2.5` resolves from libc's compat alias at build time (so -the version check passes) yet fails to load on 20.04, where those functions live -only in libutil. A future runner bump, a new native dependency, or a dropped -ldflag therefore fails the release build instead of shipping a Linux app that -crashes on launch. - -> The gate is a static invariant, not an integration test. The load path was -> verified by hand for this fix (real Ubuntu 20.04, x64 + arm64: `require` -> node-pty and spawn a shell). A CI smoke test that loads the packaged -> `pty.node` in a glibc-2.31 container and spawns a shell is the recommended -> follow-up — it would make the load path self-verifying and stay valid even if -> the build ever moves to an old-glibc sysroot. - -The one carve-out is the `sherpa-onnx` speech prebuilt, which already requires -`GLIBCXX_3.4.29` (GCC 11). It loads lazily in the speech worker -(`src/main/speech/stt-worker.ts`), never at app launch, so it is exempt from the -libstdc++ floor — its glibc needs are still checked. Speech-to-text therefore -needs a host with libstdc++ from GCC 11+ (Ubuntu 21.10 / 22.04 LTS or newer); the -app itself still launches on stock 20.04. - -**3. Check before loading, on hosts that ship without a compiler (`orcad`).** -The two gates above protect the packaged desktop app, where the binary is built and -verified by the same pipeline. `orcad` is deployed to hosts Orca never built on, so it -adds a runtime precondition -([`src/main/orcad/node-pty-precondition.ts`](../../src/main/orcad/node-pty-precondition.ts)), -run from `main.ts` before anything requires `node-pty`. It loads the addon in a **child -process**, so a binary the loader refuses — or one that aborts outright — is data rather -than this process's death, and the operator gets a sentence naming the host's libc, its -Node ABI, its prebuild slot and the command to run. A proven-unloadable binary exits 78 -(`EX_CONFIG`) instead of reaching the `require`; a probe that never answered is reported -as unverifiable and boots anyway, because a silent probe is not evidence. Whatever it -finds is published in `status.get`'s `degradations[]` under `terminal_unavailable`. - -**4. Ship the binary, built from patched sources.** -[`config/scripts/build-orcad-prebuilds.mjs`](../../config/scripts/build-orcad-prebuilds.mjs) -(`pnpm run build:orcad-prebuilds`, before `build:orcad`, which copies its target's slot into -the package's `node_modules/node-pty/build/Release`) compiles node-pty for the current -host against the pinned Node's hash-verified headers at N-API 8, and files it under -`out/orcad-prebuilds//`, where a slot is `linux-{x64,arm64}-{glibc,musl}`, -`darwin-{x64,arm64}` or `win32-{x64,arm64}`. Its `manifest.json` records each file's -sha256, the N-API level and, for glibc slots, the highest `GLIBC_` version the binary -needs; the loader checks N-API, libc, arch and that glibc version before it installs a -slot. glibc slots are built in a `manylinux_2_28` (AlmaLinux 8) container and pass the -same gate at a **glibc 2.28 / `GLIBCXX_3.4.25`** floor instead of the desktop's 2.31, because -the pinned Node they ship beside already runs on 2.28 and a 2.31 slot would leave 2.28–2.30 -hosts with a runtime but no terminal (design D6). The container's gcc-toolset supplies C++20 -and links newer libstdc++ symbols statically, so the slot needs only RHEL 8's system -libstdc++. musl slots skip the gate, since they never meet glibc's libraries. libc is part -of the slot name because node-pty's own loader falls back to `prebuilds/-` -and cannot tell glibc from musl — a glibc binary parked there is loaded on Alpine and dies at `dlopen`. -The script refuses to compile a tree where `config/patches/node-pty@1.1.0.patch` is not -applied: without the patch the prebuilt is a #9902 crash shipped as an artifact rather -than a first-connect error. CI runs it once per slot inside the matching container -(`--slot=` forces the label), merges the trees, and `--require-slots` fails a release with -a hole in the matrix; `--require-slots ` checks one slot's files against their -hashes and `--smoke` loads it under the pinned Node and spawns a PTY -(`.github/workflows/node-server-tests.yml` runs both on every slot's runner). - -The opt-in `linux-x64-glibc217` compat slot (rung B, not part of the default matrix) is -built in `manylinux2014_x86_64` (glibc 2.17, devtoolset C++20) with `-static-libstdc++`, -gated at a glibc 2.17 floor, refused if `libstdc++.so`/`libgcc_s.so` remains in -`DT_NEEDED`, and smoked under the unofficial glibc-217 Node pinned in -`NODE_RUNTIME_COMPAT_ASSETS`. Nothing installs it yet: the loader and the SSH deploy -still choose only default slots. - -## Adding or upgrading a native dependency - -- Prefer packages that ship prebuilt binaries compiled against an old toolchain - (manylinux / `glibc 2.17`-class), like `@parcel/watcher`. -- For a module we compile from source, if the gate flags it, either pin the - offending symbols the way node-pty does, or build it in an old-glibc container. -- To check locally on a Linux host, list what a binary requires (skipping the - weak `0x02`-flagged needs the loader tolerates): - - ```bash - objdump -p path/to/module.node | sed -n '/Version References/,/^$/p' - ``` - - No strong `GLIBC_` node may exceed `2.31`, and no `GLIBCXX_`/`CXXABI_` node may - exceed `3.4.28`/`1.3.12` — what stock Ubuntu 20.04 ships. - -## Runtime floor: the `environ` race below glibc 2.41 (Electron ≥ 43.7.0) - -Separate from the build floor above, one glibc runtime bug constrains which -Electron we may ship. Before glibc 2.41, `setenv`/`unsetenv` reallocate the -`environ` array and **free** the old one, so a concurrent `getenv()` on another -thread reads freed memory. Ubuntu 20.04–24.04 (2.31–2.39) are all below that -line, so every Linux target we support is exposed. - -Electron 43.5.0 made that latent race reachable on every launch: it started -setting `GDK_GL=disable` around `gtk_init()` and unsetting it right after, while -in the same change moving FontConfig warm-up onto a thread-pool thread that runs -concurrently and calls `getenv()` constantly -([electron#53070](https://github.com/electron/electron/pull/53070)). The result -is a browser-process use-after-free about a second into startup — no window, no -GPU child involved, and the corruption surfaces wherever the next allocation -lands, which is why reports name unrelated frames (`gtk_widget_realize`, -libxcb-dri3, FontConfig/expat). Orca 1.4.199/1.4.200 shipped that runtime and -died on launch on Ubuntu + NVIDIA/X11 -([#20081](https://github.com/stablyai/orca/issues/20081)). - -Electron 43.7.0 fixes it by overriding `setenv`/`unsetenv`/`putenv`/`clearenv` -so a published `environ` is never freed, deferring to glibc on 2.41+ -([electron#53491](https://github.com/electron/electron/pull/53491), backported -to 42/43/44/45). **Do not downgrade Electron below 43.7.0, or move to another -line, without confirming that backport is in the target release** — -`config/scripts/electron-runtime-floor.test.ts` fails the suite if the pin drops -below the floor. Orca itself writes `process.env` during early startup -(`patchPackagedProcessPath`, `configureOrcaUserDataPathEnv`, -`hydrate-shell-path`), so it is a first-class trigger, not just a bystander. diff --git a/docs/reference/macos-press-and-hold.md b/docs/reference/macos-press-and-hold.md deleted file mode 100644 index 946a1ff7125..00000000000 --- a/docs/reference/macos-press-and-hold.md +++ /dev/null @@ -1,58 +0,0 @@ -# macOS press-and-hold and key repeat - -macOS opens the accent picker when a key is held unless an application opts out in its preferences -domain. That prevents held keys from repeating in terminal applications such as vim. On the first -eligible launch, Orca writes: - -```sh -defaults write com.stablyai.orca ApplePressAndHoldEnabled -bool false -``` - -The write is scoped to Orca's packaged bundle domain. Bare Electron development bundles and -non-macOS platforms are left untouched. A fresh write is conservatively treated as taking effect -on the next launch. - -## Precedence and decision record - -Orca checks for an explicit domain value before writing. Either `true` or `false` is treated as a -user choice and preserved. Only an unset key receives the `false` default. - -The decision is stored once in -`/macos-press-and-hold-default.json`. An `applied` or -`kept-user-preference` decision prevents future launches from touching the domain again. -Probe and write failures remain retryable so a transient failure does not permanently disable the -fix. - -`defaults read ` is used instead of -`systemPreferences.getUserDefault`: the Electron API cannot distinguish an unset key from an -explicit `false`. Only the missing-key exit status is interpreted as unset; spawn failures, -timeouts, and other exit statuses leave the preference alone. - -## Restoring the accent picker - -Set the preference explicitly, then restart Orca: - -```sh -defaults write com.stablyai.orca ApplePressAndHoldEnabled -bool true -``` - -After Orca has recorded its one-time decision, deleting the key also restores the macOS default -without Orca recreating it: - -```sh -defaults delete com.stablyai.orca ApplePressAndHoldEnabled -``` - -Development and prerelease channels may use a channel-suffixed Orca bundle identifier; use that -domain instead when applicable. - -## Reverting - -Deleting the startup code is not enough. AppKit reads the persisted preference, so a code revert -must also arrange to delete the key for users who ran an affected build. - -## Test coverage - -Unit tests cover platform guards, explicit-value preservation, retry behavior, domain ownership, -record persistence, and subprocess exit interpretation on CI. A macOS-only test additionally pins -the real `defaults(1)` behavior, but current PR CI does not execute tests on macOS. diff --git a/docs/reference/malformed-worktree-registration-removal.md b/docs/reference/malformed-worktree-registration-removal.md deleted file mode 100644 index 34a32e58caf..00000000000 --- a/docs/reference/malformed-worktree-registration-removal.md +++ /dev/null @@ -1,43 +0,0 @@ -# Malformed worktree registration removal - -Git can report a linked worktree at `/.git` when its administrative -`gitdir` backlink incorrectly ends in `.git/.git`. That reproduces #17316's -validation error. The reproduction establishes the malformed registration, not -which program created it; current OMP uses ordinary `git worktree add`. - -Orca's desktop and runtime removal entry points use registration-only recovery -when Git positively marks the row prunable, the row has a named local branch and -HEAD, it is neither main nor locked, and the execution filesystem confirms the -selected `.git` path is a regular file. Missing or unknown evidence does not -permit this recovery. A symlink or directory is not a regular-file proof. - -Recovery reuses `git worktree prune` followed by a strict worktree listing that -must confirm the selected registration is gone. It does not delete the selected -file, infer a parent path for deletion, or delete the branch. Archive hooks and -checkout teardown are skipped because the selected row is not a checkout. - -Two consequences are intentional: - -- Git's prune also clears other stale, unlocked registrations in the repository; - it is not a path-scoped command. Live and locked registrations remain Git's - responsibility, and Orca verifies that the requested registration disappeared. -- The surviving checkout's `.git` file points at removed administrative metadata. - Files and its named branch are preserved; recovery removes the broken navigation - entry and does not repair or claim to restore that checkout. - -Native and WSL checks use the existing execution-filesystem accessor. WSL prune -and verification use the same selected distro. Paired runtimes run the recovery -on their owning host. Direct SSH does not enter this local recovery: its current -provider has no registration-only removal operation, and a failed remote removal -never authorizes a local fallback. - -The Git commands already exist in the 2.25-compatible cleanup path. On an older -Git that cannot positively attest this file-shaped registration as prunable, Orca -refuses this recovery. Deferred deletion independently rejects non-directory and -symlink targets, so force cannot move a `.git` file into deletion trash. - -Regression coverage is in `worktree-prunable-git-file.test.ts`, -`worktrees-removal-recovery.test.ts`, and -`worktree-deferred-removal-real-git.test.ts`. The latter reproduces the exact -malformation against the installed Git binary in a disposable repository and -checks surviving file contents, branch HEAD, and removed registration. diff --git a/docs/reference/managed-data-accounts.md b/docs/reference/managed-data-accounts.md deleted file mode 100644 index c0a78d001d5..00000000000 --- a/docs/reference/managed-data-accounts.md +++ /dev/null @@ -1,23 +0,0 @@ -# Managed OpenCode and Devin accounts - -Run enrollment in a terminal on the machine running Orca: - -```sh -orca account add --agent opencode --label Work -orca account add --agent opencode --integration opencode-go --label Work -orca account add --agent devin --label Work -orca account list --agent opencode --json -orca account select --agent opencode --account -orca account select --agent opencode --account system -orca account rm --agent opencode --account -``` - -OpenCode enrollment requires OpenCode 2 and runs its official `auth login --standalone` command. Devin runs `auth login --force-manual-token-flow`; obtain the enrollment token through Devin's supported login flow. These commands neither reuse a guessed token nor sign out the system account. Settings → AI Provider Accounts provides the enrollment command, refresh, selection, and removal for the selected Orca host. - -Each profile belongs to the execution host. OpenCode's SQLite credentials and Devin's credential TOML stay in private Orca user-data directories. Enrollment isolates XDG data/config/cache/state, copies only authenticated credentials, and then deletes the temporary directory. OpenCode capture rejects databases containing conversations and includes SQLite WAL contents. RPC summaries contain labels, IDs, and integration names, never tokens or credential paths. Only the authenticated local runtime socket can import a credential directory; paired clients cannot ask the host to read arbitrary paths. - -Selection affects newly launched explicit OpenCode/Devin commands and agent launches. It redirects XDG data and state; OpenCode inline-auth/database overrides cannot bypass the profile. Shell wrappers restore this selection after user startup files. Existing provider configuration and environment-based integrations remain available. Running terminals retain their current profile. Stop agents before removing a profile: removal also deletes conversations created in that private profile, without changing the system login. - -For SSH, enroll by running the command on a headless Orca runtime on the remote machine. The remote runtime owns its profiles and selection; a desktop client's credential paths never cross SSH. Direct SSH relay launches and Windows-hosted WSL panes do not consume the desktop host's profiles. Run a headless runtime inside that execution environment instead. Folder workspaces use the same host account store as git worktrees. Older Orca hosts reject new operations before login through capability negotiation. - -Validation covers OpenCode 2.0.16 on macOS and Linux arm64, including isolated official enrollment, selected and System background-terminal credential checks, reselection, and profile deletion. Linux checks used the Node headless runtime in an Ubuntu 24.04 container. Devin 3000.10.31 saved-login recognition was checked on macOS. Fresh Devin manual-token enrollment, a physical SSH host, Linux desktop UI, Windows, and Windows-hosted WSL still require verification. diff --git a/docs/reference/monaco-language-associations.md b/docs/reference/monaco-language-associations.md deleted file mode 100644 index ee985d85959..00000000000 --- a/docs/reference/monaco-language-associations.md +++ /dev/null @@ -1,34 +0,0 @@ -# Monaco filename associations - -Orca imports Monaco's full `editor.main.js` entry point, which registers the built-in -languages and loads their grammars on demand. Filename detection must not load the -editor itself: it also runs during session restoration and before the editor mounts. - -`monaco-language-associations.json` contains the registration metadata from the -installed package's entry point, plus curated Ruby associations in the generator. -Those add `.rake`, `.ru`, `.jbuilder`, `.thor`, `Guardfile`, `Capfile`, `Podfile`, -`Brewfile` and `Vagrantfile` to the existing Ruby grammar. Change these in -`config/scripts/generate-monaco-associations.mjs`, not the generated JSON. -Regenerate after changing the curated associations or upgrading Monaco: - -```sh -node config/scripts/generate-monaco-associations.mjs -pnpm exec oxfmt --write src/renderer/src/lib/monaco-language-associations.json -``` - -The generator reads syntax trees without executing contributions or grammar loaders. -Its test compares the checked-in metadata to the installed package and curated associations. The original -list was verified against a clone of `microsoft/monaco-editor`, tag `v0.55.1`, commit -`516f350bdaf7a82f6731bd128a9ec86a6e5fa47d` (`src/basic-languages` and `src/language`). - -Existing Orca filename and extension choices take precedence. This preserves custom -Vue, Svelte, Astro, Nim, Typst, JSONL, notebook and preview handling, as well as the -Markdown mapping for MDX. The fallback matches exact filenames before the longest -extension, case-insensitively, and resolves duplicate associations in upstream -registration order (the last registration wins, so `.pp` selects Ruby over Pascal). - -Monaco 0.55.1 has 90 registrations, 81 with filenames or extensions. Registrations -without either remain available in Monaco but cannot be inferred from a path. This -does not add VS Code extensions or language servers, guess from file contents, or -change mobile's separate lowlight grammar set. The renderer uses only the basename, -so local, Windows, SSH and folder-workspace paths share the same detection. diff --git a/docs/reference/omp-fresh-launch.md b/docs/reference/omp-fresh-launch.md deleted file mode 100644 index ea667938172..00000000000 --- a/docs/reference/omp-fresh-launch.md +++ /dev/null @@ -1,51 +0,0 @@ -# Fresh OMP launches - -Orca's new-session and draft launch plans apply a one-time `--config` overlay -containing `autoResume: false`. OMP's configured session directory, settings, -authentication and extensions remain in their usual locations. Saved launch -configuration omits the overlay so explicit resume keeps its normal semantics. -Custom commands with session selectors, unknown flags, positional arguments or -shell compounds are left unchanged. - -The execution host creates the overlay. Local and WSL terminals use Orca userData -(with WSLENV path translation); SSH relays use their own managed directory. Config -creation is independent of status-hook preferences and does not require plugin -source installation. An unavailable file produces a terminal diagnostic and skips -OMP. Filesystem failures do not prevent unrelated agents or bare shells starting. - -The guard invokes OMP in the current shell, preserving functions, aliases and the -managed status wrapper. A nonzero agent exit never triggers a second launch. -POSIX commands also support fish; environment presence is checked before expansion -so an old host under `set -u` reports the same missing-settings diagnostic. - -## Mixed versions - -The existing command and environment transport carries the launch unchanged; no -new RPC or stream opcode is introduced. A new host exports the path before shell -startup, including bare shells that receive an OMP command later. An old client -continues its existing launch behavior on a new host. A new fresh-launch command -on an old host without the managed environment fails visibly and requests a host -update and terminal restart. It must not silently fall back to OMP auto-resume. - -## Verification - -`src/shared/omp-fresh-launch-shell.test.ts` runs actual available bash, zsh and fish -shells, checking exact argv, a single invocation, nonzero exit, deleted settings -and an absent environment variable. `src/relay/omp-fresh-launch-environment.test.ts` -checks the guarded command through relay environment assembly and the actual OMP -shell wrapper, retaining both extension and config arguments and prefill. -Local host assembly tests recognize guarded POSIX, cmd and PowerShell commands. - -Run the actual OMP storage smoke against a read-only OMP checkout: - -```sh -ORCA_BACKGROUND_LAUNCH=1 bun tests/tools/omp-fresh-session-runtime-smoke.mjs /path/to/oh-my-pi -``` - -For Windows, bundle `tests/tools/omp-fresh-launch-windows-smoke.ts` with -`bun build --target=node --outfile=/tmp/omp-fresh-windows-smoke.mjs`, transfer the -bundle to the host and run it using Node with `ORCA_BACKGROUND_LAUNCH=1`. -The smoke uses temporary files and process-local environment only. Both cmd and -PowerShell must pass existing/missing/unset/directory settings cases, preserving -exit 17 for the single successful launch and returning exit 1 without launching -when settings are unavailable. This passed on Windows host `awin` on 2026-09-14. diff --git a/docs/reference/omp-history-titles.md b/docs/reference/omp-history-titles.md deleted file mode 100644 index f238574ef2a..00000000000 --- a/docs/reference/omp-history-titles.md +++ /dev/null @@ -1,34 +0,0 @@ -# OMP history titles - -The message-graph scanner uses persisted OMP names ahead of the first user prompt: -`session.title`, version-1 `title` slots, `title_change.title`, and legacy -`session_info.name`. Empty or unsupported metadata leaves the previous name or -prompt fallback intact. Non-OMP graph parsing keeps its existing title policy. - -Explicit user names outrank automatic names. Within the same source, timestamps -prevent the current first-line title slot from being replaced by older rename -entries later in the file. Newer appended renames still update the row. Legacy -records without timestamps retain file-order handling. - -The graph fold stores title authority alongside the existing accumulator. Clones -retain it without sharing mutable accumulator or preview state, while preserving -the existing identity and message-consumer contracts. Cached append parsing uses -the normal durable offset; no extra scan, process, poll or watcher is introduced. - -The parser is shared by local and remote content readers and uses transcript data -from the execution host. It performs no client-side path lookup and changes no -wire shape. Folder workspaces require no git metadata. - -Run actual persistence and cache validation with a read-only OMP checkout: - -```sh -ORCA_BACKGROUND_LAUNCH=1 bun tests/tools/omp-history-title-smoke.mjs /path/to/oh-my-pi -``` - -The smoke persists a first prompt, performs a real OMP user rename, and verifies -both cold and incrementally cached scans. It checks one full parse, one append -parse and identical-object reuse on an unchanged scan. All home/config/data roots -are disposable; no model requests are made. - -This is the OMP subset of the history-name behavior proposed in PR #15696 by -Brennan Benson. Pi naming and title changes in the terminal are separate concerns. diff --git a/docs/reference/omp-resume-transcript-locator.md b/docs/reference/omp-resume-transcript-locator.md deleted file mode 100644 index 75e38b798cb..00000000000 --- a/docs/reference/omp-resume-transcript-locator.md +++ /dev/null @@ -1,28 +0,0 @@ -# OMP recorded transcript resume - -A sleeping OMP session can retain its transcript path from a hook without an -explicit `launchConfig.ompResumeFilePath`. Both cold-restore startup and generic -sleeping-session launch already forward the provider metadata to -`getAgentResumeArgv`; that builder must keep the recorded path. - -Resolution order is explicit launch path, recorded transcript path, then UUID. -The existing shell-aware builder quotes the selected argument for the execution -host. An older metadata record without a path retains UUID fallback. OMP provider -claim keys and equality remain UUID-based, so a later hook adding the path does -not create a second automatic-resume identity. - -`tests/tools/omp-resume-transcript-locator-smoke.mjs` creates an actual OMP session -outside its default session store. UUID lookup fails there; the absolute path and -Orca's generated argv resume the original session. Run it with Bun and a read-only -OMP checkout as argv[2], under `ORCA_BACKGROUND_LAUNCH=1`. It uses a disposable home -and makes no model requests. - -This bounded correction follows the resume-locator portion of -[PR #16276](https://github.com/stablyai/orca/pull/16276) by @CodeHourra. It does not -adopt that PR's reattach injection or title changes. The reattach proposal treats -missing snapshot/replay as permission to type a resume command, but -`daemon-pty-spawn-result.ts` explicitly permits `isReattach: true` without a -snapshot. That payload absence is not positive evidence of a newly created shell. -The proposal also adds the path to OMP claim identity, which separates UUID-only -metadata from a later path-enriched record for the same provider session. Those -changes require separate evidence and are outside this patch's review scope. diff --git a/docs/reference/omp-runtime-session-provenance.md b/docs/reference/omp-runtime-session-provenance.md deleted file mode 100644 index 8d6e33577d9..00000000000 --- a/docs/reference/omp-runtime-session-provenance.md +++ /dev/null @@ -1,23 +0,0 @@ -# OMP runtime session provenance - -OMP computes whether a runtime session is a task child, but the released -`ExtensionContext` does not expose that value. The status extension therefore uses -the session manager's parent header and nested task transcript path only when a -root owner is already known. A nested transcript with no known owner remains -eligible because it may have been resumed directly as the pane's main session. - -The remaining child-first case is inherently ambiguous to Orca: task children and -resumed child transcripts have the same public session-manager shape. A complete -child-first fence requires OMP to expose its computed `agentKind` through -`ExtensionRunner.createContext`; until then the conservative fallback avoids -silencing valid resumed sessions. - -Older runtimes retain the manager-identity guard. That guard assumes the main -session reaches Orca's callback before any child. An earlier user extension can -initialize a child during session_start and violate that assumption. Keep the -ownership merge assessment conditional until the runtime API is available and the -combined flow is validated. Neither callback timeouts, UI presence, nor transcript -paths establish runtime ownership. - -The guard remains scoped to one pane and launch token. It does not define how -several independent SDK/ACP roots sharing one process and pane should be attributed. diff --git a/docs/reference/omp-session-roots.md b/docs/reference/omp-session-roots.md deleted file mode 100644 index 481d5f205c2..00000000000 --- a/docs/reference/omp-session-roots.md +++ /dev/null @@ -1,53 +0,0 @@ -# OMP transcript roots - -Native chat lookup, history discovery and the history path allowlist use -`src/main/ai-vault/omp-session-root.ts` on the execution host. The resolver follows -the active runtime environment rather than scanning every profile. - -The behavior matches OMP's `packages/utils/src/dirs.ts`: - -- On Linux and macOS, an existing `$XDG_DATA_HOME/omp` selects that app's `sessions` - directory, even when legacy transcripts coexist. The sessions directory itself - need not exist yet. There is no implicit `~/.local/share` fallback. -- A named profile uses XDG only when its own `omp/profiles/` path exists; - otherwise it uses the profile's config-root `agent/sessions` directory. -- `OMP_PROFILE` takes precedence over `PI_PROFILE`, including an explicitly empty - canonical value. Named profiles ignore custom `PI_CODING_AGENT_DIR` values. - Default mode respects custom agent directories, except an inherited agent path - derived from the lower-priority profile. `PI_CONFIG_DIR` selects the config root - relative to the owning host's home, as upstream specifies. -- Orca retains its legacy `OMP_CODING_AGENT_DIR` sessions-root override and prefix - normalization. Explicit scan roots override environment discovery. Empty or - filesystem-root scan overrides and invalid profile names refuse discovery; - they never fall back to a different profile or the process working directory. - -The desktop scanner child allowlist forwards only the required directory/profile -variables. SSH relay discovery still builds legacy roots from its host-owned home -and does not gain XDG/profile discovery here. Client XDG/profile values are not -applied to WSL home roots. Exact hook -paths and existing WSL attestation/refusal remain authoritative. No wire fields -or opcodes change; older clients receive the existing session record shape. - -This resolves environment-visible configuration. Per-command `--profile` choices -or directory values loaded only inside the agent are not inferred by a runtime -that never received them; hook-reported transcript paths remain the exact route. - -Run the read-only upstream parity smoke with: - -```sh -ORCA_BACKGROUND_LAUNCH=1 bun tests/tools/omp-session-root-upstream-smoke.mjs /path/to/oh-my-pi -``` - -The smoke uses disposable home/data roots and compares Orca's result with OMP's -actual directory resolver. It makes no model requests. Unit tests also cover -Windows XDG exclusion, legacy override normalization and refusal paths. - -For actual persistence-to-reader validation, run: - -```sh -ORCA_BACKGROUND_LAUNCH=1 bun tests/tools/omp-transcript-root-reader-smoke.mjs /path/to/oh-my-pi -``` - -This creates default and named-profile transcripts through OMP's SessionManager, -with legacy directories still present, then resolves and decodes each by session -ID through Orca's native reader. All files use disposable roots; no model runs. diff --git a/docs/reference/omp-startup-keyboard-query.md b/docs/reference/omp-startup-keyboard-query.md deleted file mode 100644 index 7a675e4e0b2..00000000000 --- a/docs/reference/omp-startup-keyboard-query.md +++ /dev/null @@ -1,15 +0,0 @@ -# OMP startup keyboard capability query - -OMP's ProcessTerminal sends `CSI ? u` and then a DA1 sentinel before selecting its keyboard encoding. A fresh desktop terminal already advertises Kitty support, but a direct New Tab launch can query before that renderer owns replies. Previously startup ingress answered only OSC color queries. - -The renderer now supplies optional `terminalKittyKeyboardProtocol: true` from its actual xterm `vtExtensions.kittyKeyboard` setting. The existing local/SSH/paired spawn route places it in startup ingress as optional `kittyKeyboardProtocol`. Missing or false capability leaves behavior unchanged, including panes that deliberately withhold Kitty on native Windows. The ingress version and stream opcodes do not change. Terminal creation accepts additive fields, but host-authoritative `terminal.createAgentSession` and `terminal.ensureAgentSession` use strict schemas. Clients send the keyboard flag on those methods only after the host advertises `agent-session.keyboard.v1`. The negotiated payload stays fixed across launch retries; old hosts receive the original payload and retain the renderer fallback. New hosts accept older clients that omit the flag. Paired background launches use the same negotiated support and default paired-terminal advertisement; legacy terminal creation receives the additive flag. - -Source ingress answers only the exact first `CSI ? u` before its deadline/renderer handoff. It uses the existing mode tracker for preceding flag pushes and the existing reply-delivery echo guard. Its transformed source span consumes the query once, while the following DA1 and Kitty mode-setting bytes retain their sequence ranges and reach the renderer. Keyboard intent does not require theme colors. Color and Kitty authority end independently: answering both colors does not end Kitty handling, and ConPTY's persistent color ownership does not retain Kitty ownership after handoff. - -Run the actual OMP protocol smoke with a read-only reference checkout: - -```sh -ORCA_BACKGROUND_LAUNCH=1 bun tests/tools/omp-startup-keyboard-smoke.mjs /path/to/oh-my-pi > /tmp/omp-startup-keyboard.json -``` - -This uses OMP's real ProcessTerminal with intercepted process-local stdin/stdout, a disposable HOME, and no model call. It verifies negotiation before any renderer attaches, a single reply, preserved mode push, and contiguous raw sequence coverage. It does not constitute live Windows/SSH or rendered shortcut proof. diff --git a/docs/reference/omp-status-input-redaction.md b/docs/reference/omp-status-input-redaction.md deleted file mode 100644 index 80df07dbb6c..00000000000 --- a/docs/reference/omp-status-input-redaction.md +++ /dev/null @@ -1,39 +0,0 @@ -# Pi/OMP status tool-input redaction - -Generated tool_call and tool_execution_start hooks sanitize inputs on the agent's -execution host before passing them to the existing status transport. Inputs that -reference `.ssh`, `.ssh-mcp`, `.mcp-secrets.env`, or -`.omp-backups-archive/omp-bak-keyfile` become `{ redacted: true }`. Matching includes -nested values, property names, Windows separators, case variants and shell token -boundaries. Sibling names such as `.ssh-backup` remain ordinary data. - -The sanitizer copies data descriptors into objects without prototypes. It never -passes the source object's toJSON or getters to the transport. Cycles, accessors, -class instances, symbols, functions and non-JSON primitives redact the whole input. -Repeated ordinary object references are allowed. Depth, visited values, reserved -array slots (including holes) and inspected text have conservative bounds to avoid -moving unbounded work into synchronous JSON serialization. - -This policy targets credential-path references in status tool inputs. It does not -scan transcript files, tool outputs, prompts or arbitrary secret values, and does -not erase previously persisted status. Proxy reflection traps can still run when -JavaScript inspects a proxy; the sanitizer is not an isolation boundary against a -malicious extension in the same process. - -Ordinary question envelopes and preview input shapes remain unchanged. Older -clients already accept object-valued tool_input; no RPC fields or opcodes change. -The same generated code runs on local, WSL and SSH agent hosts; no local filesystem -lookup or substitution is introduced. Folder workspaces require no special path. - -This follows PR #9554's credential-reference policy, with descriptor copying and -bounded serialization correcting its validation-then-original-object approach. - -Run the actual OMP loader/native HTTP smoke with a read-only checkout: - -```sh -ORCA_BACKGROUND_LAUNCH=1 bun tests/tools/omp-status-input-redaction-smoke.mjs /path/to/oh-my-pi -``` - -It loads Orca's generated extension through OMP, invokes synthetic tool events, -and inspects three real loopback HTTP payloads. Home/config/data roots are -disposable; it makes no model requests and does not claim an interactive tool run. diff --git a/docs/reference/orcad-operations.md b/docs/reference/orcad-operations.md deleted file mode 100644 index 961b7fa1e00..00000000000 --- a/docs/reference/orcad-operations.md +++ /dev/null @@ -1,399 +0,0 @@ -# Running orcad - -`orcad` is the Orca runtime served from plain Node. This is the contract between it and -whatever supervises it: what it binds, what it owns on disk, who restarts what, and what its -readiness payload actually proves. - -## Two long-lived processes, not one - -A deployment is **orcad** plus **the terminal daemon**. - -| | orcad | terminal daemon | -| ---------- | -------------------------------- | ------------------------------------- | -| Started by | the supervisor | orcad, detached | -| Owns | RPC, git, worktrees, persistence | every local PTY | -| Lifetime | one supervised run | detached from orcad, not its service | -| Endpoint | `ws://:` | `/daemon/daemon-v.sock` | - -orcad detaches the daemon and calls `disconnectDaemon()`, never `shutdownDaemon()`. The -built-in remote deployment path stops only the recorded orcad PID, so the daemon and its PTYs -survive. The successor adopts the current endpoint and routes supported previous protocol -versions through legacy adapters. This makes a PID-scoped update, rollback or restart -non-destructive to live work. - -Process detachment is not service isolation. A daemon that orcad launches directly, and every -PTY it owns, remain in the same systemd service cgroup. `KillMode=mixed` does **not** preserve -them: it sends the graceful stop signal only to the main process, then sends `SIGKILL` to every -process remaining in the cgroup the moment that main process exits — `TimeoutStopSec` never gets -the chance to apply. `KillMode=control-group` is destructive too. `KillMode=process` leaves -service-owned processes unmanaged and is not a supported preservation mechanism. - -Service-restart survival therefore requires a separately supervised cgroup, and orcad now asks -for one: on Linux it launches the daemon through `systemd-run --user --scope`, which places the -daemon and its PTYs in their own transient `orca-daemon-.scope` unit under the -user slice instead of the caller's service cgroup. A stop or restart of the service unit then -leaves that scope — and the live terminals in it — running, and the successor adopts the -endpoint as it always has. - -A newly launched private daemon scope also follows the daemon's own lifetime. A small -detached shell holds an input pipe from the daemon; after that pipe closes and `/proc` -confirms the daemon PID is gone, it asks the user manager to stop that exact scope. -Systemd sends remaining processes SIGTERM and escalates after five seconds. This includes -children that double-forked or called `setsid` and can no longer be found by parent PID. -Disconnecting or restarting the runtime does not close the pipe: the daemon owns it. - -The cleanup only arms on a fresh scoped launch with a matching launch nonce. Adopted -legacy scopes can contain GUI processes and are never armed retroactively. Unscoped -launches and children deliberately moved into another systemd unit remain outside this -cleanup. `nohup`, `disown`, and `tmux` alone do not move a process out of its cgroup, so -those children now end when their terminal daemon dies. Work intended to outlive that -daemon needs its own service or scope. - -The scope is requested only where it can work. All of these must hold: - -- **Linux with systemd as PID 1** (`/run/systemd/system` exists). -- **A reachable user bus** — a connectable `bus` socket in the per-UID runtime dir - (`/run/user/`, or whatever `XDG_RUNTIME_DIR` points at). For a service account that is - not otherwise logged in, that means `loginctl enable-linger `; a unit whose - `RuntimeDirectory=` hardening moves `XDG_RUNTIME_DIR` off the per-UID path is handled, because - the real per-UID path is probed first. -- **`systemd-run` on `PATH`** and answering `--version`. - -Any of those missing, or a `StartTransientUnit` call that fails anyway, falls back to the -direct launch — and in that unscoped fallback case the paragraph above still describes reality: -the daemon shares the service cgroup and a combined-unit stop ends live terminals. Read the -`cgroupUnit` field in the daemon health payload to tell the two cases apart on a running host; -it is populated from `/proc/self/cgroup`, so it reports the isolation the daemon actually has -rather than what the launcher intended. - -## `orca serve` on this machine - -`orca serve` runs on the local orcad slot by default. The CLI asks the app's -`out/main/orcad/orcad-local-serve-selection-entry.js` (run as plain Node on the app's executable) -which host to use. Any reason orcad cannot serve falls back to Electron serve with one -`[serve] using Electron serve: ` line on stderr. Those reasons are: no slot for this host, -no template in the install, the pinned Node could not be fetched, or a failed native preflight. - -- `ORCA_SERVE_RUNTIME=electron` keeps Electron serve and skips the question. `orcad` (or unset) is - the default, and any other value falls back with a reason. -- Packaged macOS stays on Electron: only Electron serve, supervised by the CLI, can take a remote - app update there, and orcad has no updater. Recipe-JSON serve has no handoff and uses orcad. -- Windows serves on orcad too. Both hosts share `\daemon`, so the daemon pipe name - (hashed from that path) is the same, and the relocated Electron daemon host changes only the - executable, not the pipe. The `orcad-serve-mode-switch-windows` e2e job checks D7 there, in the - daily run and on PRs routed to it; it does not block merges. - -The slot and its pinned Node live under the desktop's `/orcad-artifacts`. - -## Bind policy - -`--bind `, **default `127.0.0.1`**. - -Only literal IPs are accepted; hostnames are refused because DNS would decide which -interface got bound. `localhost` maps to `127.0.0.1`. `0.0.0.0` / `::` are the explicit -opt-ins to network reach, and the startup log says so on every launch. - -The bind is **pinned**, not defaulted. Two things widen the desktop's listener on their own — -`orca serve`'s wide default, and a startup where some device has connected before — and an -unattended host's exposure must be exactly what the operator asked for on every launch. A -mobile pairing offer, which normally rebinds to all interfaces, is refused while the bind is -pinned to loopback and reports `network_exposure_failed` rather than advertising an endpoint -nothing can reach. - -Under the shipping design a client reaches a remote orcad over an SSH local port-forward, so -loopback is the correct default and the pairing credential travels over SSH. - -A host whose sshd refuses forwarding (`AllowTcpForwarding no`) is reached through the stdio -bridge instead: the client keeps the same local port, and each connection to it opens one SSH -exec channel running a small script on the host's pinned Node that dials orcad's loopback port. -Windows hosts run it as the host script's `stdio-bridge` op and frame bytes as base64 lines, -because a PowerShell DefaultShell re-decodes native output. Bridges are capped below OpenSSH's -default `MaxSessions` of 10 per connection; further connections wait for a free one. The choice -is made each time the tunnel starts (`orcad-managed-tunnel-transport.ts`), so nothing is -recorded per host, and only a host where even the bridge cannot run keeps the relay, recorded as -`ssh_tunnel_unavailable`. - -## Data root and the instance lock - -The data root is `$ORCA_USER_DATA`, else `$XDG_DATA_HOME/Orca`, else `~/.orca`. - -Before the profile index or the store is touched, orcad takes `/orcad.lock`. -It refuses to start when: - -| Code | Meaning | -| -------------------------------------- | ------------------------------------------------------------- | -| `orcad_data_root_wrong_owner` | the root is owned by another uid (POSIX) | -| `orcad_data_root_shared` | the root is group/world accessible and could not be tightened | -| `orcad_instance_lock_held` | another live orcad owns this root | -| `orcad_instance_lock_foreign_identity` | the lock belongs to a different identity | -| `orcad_data_root_unusable` | the root cannot be created, stat'd or written | - -A root that is merely too permissive and that we own is tightened to `0700` rather than -refused — orcad stores credentials there unsealed (no OS keyring on this host), so the goal -is a private root, and refusing when we could just fix it helps nobody. We refuse when the -permissions are not ours to fix. Windows has no owner or mode check, because ACLs are not -expressible as a POSIX mode and `statSync().mode` there reports a synthesized one. Instead -orcad restricts the root's ACL to its own user with `icacls` (the same verified restriction -`secure-file.ts` applies to credential files) and refuses with `orcad_data_root_shared` when -that cannot be applied. - -A dead holder's record is reclaimed (PID plus process start time, so a recycled PID does not -read as alive). On Windows the start time is the kernel creation time read through the -process-tree addon the slot stages; without the addon it is null and the PID alone fences, -which errs toward "held". A record belonging to a different identity is never reclaimed. - -**The lock scopes one role — who is the runtime.** It deliberately says nothing about the -daemon, which lives under `/daemon` and fences its own endpoint with its own PID -record. A lock that asked "is any process using this root" would refuse exactly the restarts -a live daemon makes worthwhile. - -## Supervision - -### Process-scoped and cgroup-wide stops - -The built-in remote updater performs a PID-scoped stop and keeps the daemon's install version -pinned while it owns sessions. A combined-unit systemd stop or restart is different: unless the -daemon holds a durable cgroup scope of its own (see -[Two long-lived processes, not one](#two-long-lived-processes-not-one)), it reaps the daemon and -every live terminal after the graceful window. Treat a stop as destructive unless -`health.terminalDaemon.cgroupUnit` names an `orca-daemon-*.scope` on that host. - -Before a cgroup-wide stop, obtain a fresh `orca-ide terminal list --json` result using the same OS -account and home as the daemon. Invoke the installer's absolute launcher path so `sudo`'s -`secure_path` cannot hide a per-user registration (for example, -`sudo -Hu orca /home/orca/.local/bin/orca-ide terminal list --json`). Replace both `orca` and -`/home/orca` with the service account and home used by the unit; an extracted deployment may use -its absolute `resources/bin/orca-ide` launcher instead. A safe empty census is untruncated, has an explicit `hostScope`, covers every -execution host affected by the stop, and lists no terminals on those hosts. Every -`omittedHostIds` entry must be explicitly accounted for outside the target service's execution -boundary. A separately paired runtime is outside that boundary; local execution and SSH hosts -reached through this runtime are not. An affected or unknown omission, missing scope, -truncation, a failed request or lost contact makes the result `unverifiable`: defer the stop. Do -not admit new work after the census. Orca does not yet provide an atomic census-and-stop fence. - -### Who supervises orcad - -An external supervisor (systemd, launchd, a process manager). orcad conforms to it: - -- **Readiness.** One JSON line on stdout (`--json`), `type: "orca_server_ready"`, published - after the listener is bound and the daemon verdict is in. There is no separate readiness - socket; the line is the signal. Set the supervisor's start timeout generously — the daemon - launch has its own retries and can take tens of seconds on a cold host. -- **Shutdown.** `SIGTERM` or `SIGINT` starts one graceful stop. Repeated signals share - that stop because a supervisor may signal both the launcher and its child. A 15s deadline - exits with code 1 if teardown stalls. The bundled runtime also stops gracefully if its - launcher's IPC channel closes. On POSIX, both the launcher and runtime ignore `SIGHUP`, - so terminal hangups do not stop a headless host. Use `SIGTERM` or `SIGINT` to stop it. -- **Stop requests.** A file stops orcad the same way `SIGTERM` does, without a PID that may - since have been reused by another process: - - `.orcad-stop-request` beside `orcad.js` in the running slot. orcad deletes it and stops. - - An instance-bound request in the data root, named - `.orcad-managed-stop-request.`. orcad stops only when it - names this orcad's version, runtime ID, PID, start time and lock nonce, and while the - instance lock still holds that record. The file is kept as evidence. - - `orcad --complete-managed-stop ''` writes that request, waits for the - instance to exit, and prints one JSON line whose `verdict` is `live`, `unverifiable` or - `exited`. `exited` needs proof: no process with that PID, or a PID whose start time shows - it now belongs to another process. On `exited` it writes - `/orcad-stop-receipts/.json`. It exits 0 whenever it printed a - verdict, 64 for a malformed invocation, and 1 for a failure before any verdict, which is - never evidence of exit. - - A request with `retireIdleDaemon: true` asks orcad to retire the terminal daemon too. This - is best effort and never blocks or fails the stop: - - The daemon is retired only when it proves it owns no live session across every - generation. - - A busy daemon (`live`) or one whose state cannot be proven (`unverifiable`) stays up with - its terminals, and orcad reopens new-terminal admission before exiting. - - The completed-stop receipt records `retirement` as `retired`, `live` or `unverifiable`. - If orcad exits without recording an outcome, the receipt says `unverifiable`. - - `orcad --cancel-managed-stop ''` withdraws a request orcad has not acted on. - orcad and the canceller each try to create `.decision.json` exclusively, - so exactly one wins. `canceled` means orcad keeps running and the request file is removed; - `dispatched` means orcad already began stopping, and only the completion can say how it - ended. - - A build advertises all of the above with `health.stopRequests: 1` in its readiness line. - Clients stop such a build through the slot request file and older builds with `SIGTERM`, - after corroborating the PID with readiness either way. A launch clears a slot request - that the previous process never consumed. -- **Decommissioning a managed slot.** An Orca client decommissions through the same activation - journal and fence as deploy and rollback. It refuses while the terminal census reports live - or uncounted terminals, stops the instance with a managed request that also asks to retire - the daemon, and records that no version is active only after `exited` is proven. A stop - that did not finish is cancelled; if orcad already acted on it, or the host cannot answer, - the fence stays for recovery. -- **Instance lock.** `/orcad.lock` names the running orcad. A record that is - unreadable, malformed or over 64 KiB is never reclaimed: orcad exits 78 until an operator - removes it. A shutdown whose teardown failed keeps the lock until the process exits, so a - second orcad cannot start beside a writer that may still be running. -- **Exit codes.** - - | Code | Meaning | Supervisor should | - | ---- | ------------------------------------------------------------ | -------------------- | - | 0 | clean shutdown | restart per policy | - | 1 | startup or shutdown failure | restart with backoff | - | 78 | configuration fault (bind address, data root, instance lock) | **not** restart | - - 78 is `EX_CONFIG`. Put it in systemd's `RestartPreventExitStatus`: restarting on a data - root owned by someone else is a restart-spin, not a recovery. - -- **Logs.** orcad writes human-readable diagnostics to **stderr** and its readiness contract - to **stdout**; the supervisor owns capture and rotation. The daemon, being detached, writes - its own NDJSON lifecycle log to `/logs/daemon.log` (suppressed by - `ORCA_DIAGNOSTICS_DISABLED=1`). Rotation of that file is not implemented — see - [What is not covered](#what-is-not-covered). orcad records every trace span it emits - (git commands, worktree paths, terminal spawns, structured-chat failures and the rest) to - `/logs/orcad.trace.ndjson`, rotated at 10 MB × 10 files, private to its user and - redacted for secret-shaped strings. It stays on the host: a desktop's diagnostics bundle does - not collect it. `ORCA_DIAGNOSTICS_DISABLED=1` turns it off, and a logs folder orcad cannot - open leaves it off with one stderr warning rather than stopping orcad. - -### orcad supervising the daemon - -- **Launch.** On Linux, through `systemd-run --user --scope` so the daemon gets its own - transient cgroup and survives a service-unit restart; everywhere else, and wherever that - scope is unavailable, forked detached. Either way it runs `daemon-entry.js` beside - `orcad.js` with its own PID record, token and socket under `/daemon`. -- **Adoption before spawn.** A daemon already answering the endpoint is adopted, not - replaced, unless it is unhealthy, foreign, or built from a superseded bundle _and_ owns no - live sessions. Replacing a healthy daemon kills its PTYs, so code freshness always defers - to live work. -- **Restart.** The adapter respawns the daemon on death, transparently to callers. -- **Crash-loop containment.** At most **5 launches per 60s rolling window** per orcad run; - past that, launches are refused with `daemon_crash_loop` and terminals fail with that - message instead of the process forking forever. The window slides, so a repaired host - recovers without restarting orcad. An operator-initiated daemon restart clears it — that - is the deliberate "try again". -- **No macOS login-session watch.** That watch retires the daemon when the spawning GUI login - session dies. An orcad daemon must survive its SSH session ending. -- **Shutdown.** orcad never stops the daemon. A daemon that was never adopted retires itself - after its adoption window; an adopted one stays resident (see Decommissioning). - -### Decommissioning - -After a PID-scoped stop, an adopted daemon stays resident so the next orcad can reattach. A -combined-unit systemd stop also leaves a scope-isolated daemon resident, but kills one that -fell back to the service cgroup. To retire a process-scoped deployment, apply the census rule -above, stop orcad, then stop the daemon named by `health.terminalDaemon.pid`. -Only report it `exited` after verification on the execution host; loss of contact is -`unverifiable`. - -### Windows hosts - -What differs on a Windows SSH host, and what deliberately does not: - -- **Stop path.** A signal is TerminateProcess on Windows: no flush, no lock release. The - slot's `.orcad-stop-request` file (and the managed, instance-bound request) is therefore the - only graceful stop. A detached orcad receives no console control events, so the listener - (`fs.watch` plus a one-second poll) is what stops it; the packaged-slot test proves it exits - cleanly within the 15 s shutdown deadline on every server lane, Windows included. -- **Exit proof.** `--complete-managed-stop` proves a reused PID by the addon's creation time. - Without the addon a live PID stays `live` or `unverifiable`, never `exited`. -- **Daemon endpoint.** The terminal daemon listens on a named pipe - (`\\?\pipe\orca-terminal-host-v-`), not a socket under the data root. -- **Leaving sshd's job.** orcad is started outside the SSH session's kill-on-close job, so the - daemon it forks inherits no such job and outlives the connection the same way. -- **Per-PTY jobs.** Each ConPTY child gets its own job (`windows-pty-job.ts`), and Git Bash / - MSYS panes follow [`windows-msys-job-breakaway.md`](./windows-msys-job-breakaway.md) - unchanged. A ConPTY smoke test runs inside a process started exactly that way (breakaway, - no window) on the Windows server lanes. -- **No daemon-host relocation.** The desktop copies its runtime to `%LOCALAPPDATA%` because - the NSIS updater deletes the install directory under a running daemon - ([`windows-daemon-host-relocation.md`](./windows-daemon-host-relocation.md)). orcad slots are - versioned directories that nothing deletes while a process runs from them: Windows refuses - to delete a running image, and GC treats an in-use slot as live. - -## Idle exit (client-managed orcad only) - -An orcad that a desktop client launched over SSH stops itself, like the relay, once its host has -been unused for 15 minutes. The client's launch sets `ORCA_ORCAD_MANAGED_ACTIVATION_ROOT`; an -orcad started by hand, by a supervisor, or as a paired server never carries it and never idles -out. - -"Unused" means every one of these held on every check for the whole period: - -- no client socket open and no RPC request running; -- no terminal in the PTY provider, and the daemon answered with zero live sessions (a daemon - that does not answer keeps orcad up); -- no agent reporting `working`; -- no staged migration into this server; -- no enabled automation and no automation run still in flight (nothing on the host would start - orcad again for the next scheduled run, so a server with an enabled automation never idles out); -- no activation fence on the host (an update, rollback, decommission or recovery in flight). - -The stop is the ordinary graceful shutdown, which disconnects from the daemon and never shuts it -down, so it cannot kill a terminal. It then asks the daemon to retire only if the daemon itself -proves it holds no session. Before stopping, orcad writes `/orcad-idle-stop.json`; -the next start reports it once as `health.previousIdleStop` and removes it, so a later crash is -never read as an idle stop. A managed start with no record reports `previousIdleStop: null`. - -The client starts a stopped server again, whatever stopped it (an idle stop, a kill, a host -reboot): on every connect, on every fresh tunnel (including after the client wakes from sleep), -and before a call through an environment the client restored at launch. A server that does not -answer is checked on the host; only a proven exit starts the activated slot, under the activation -fence, and the status line shows "Starting managed server…". A daemon that survived is adopted -with its terminals; after a reboot both start fresh. A process that is live or cannot be proven -gone is left alone, and a start that fails keeps the host managed with the reason and orcad.log's -tail, never as a verdict about its terminals. `ORCA_E2E_ORCAD_IDLE_TIMEOUT_MS` shortens the idle -period for tests; the client forwards it to the servers it launches. - -## Health - -The readiness payload carries a `health` object: - -``` -buildHash sha256 (16 hex) of the running orcad bundle — build identity that a version - string cannot give, so a rollback that did not replace the file is visible -buildVersion ORCA_VERSION -nodeVersion / nodeAbi process.versions.node / .modules — the ABI native addons must match -platform / arch / pid -terminalDaemon: - state live | degraded | absent - ownsFreshSessions whether NEW terminals are daemon-owned; this supports PID-scoped - restart recovery, not supervisor or service-cgroup isolation - pid the live daemon's pid, from its own PID record - buildVersion the build the LIVE daemon was forked from (may legitimately predate - this orcad after an update — reporting orcad's version for both would - hide exactly that) - entryPath / protocolVersion - selfTest { ok, coverage, verdict, durationMs } -``` - -### What the self-test proves - -`selfTest` runs `checkDaemonHealth` against the daemon's socket. It is green only when the -daemon **opened its socket, completed the protocol handshake, and ran `ptySpawnHealth` — a -real short-lived PTY spawned inside the daemon's own process**. It therefore spans both -processes: orcad drives it, the daemon performs it, the verdict crosses the socket. - -- `coverage: 'pty-spawn'` — the full round trip above. -- `coverage: 'handshake'` — **win32 only**, where `checkPtySpawnHealth` returns without - spawning anything. A green verdict there covers the handshake and nothing more. It is - reported separately rather than folded into `ok` so nobody reads it as a PTY round trip. - -`state` is `live` only when the self-test passed **and** `ownsFreshSessions` is true. A -daemon that answers but has fallen back to local spawning for new terminals is `degraded`, -because those terminals die with orcad. A daemon that answered and then failed its spawn -probe is also `degraded`, not `absent`: it still holds live sessions, and calling those -exited would be the verdict `ssh-execution-boundary.md` forbids guessing. - -## What is not covered - -Named here so nothing reads as implemented that is not: - -- **A continuous health endpoint.** `health` is published once, in the readiness payload. A - supervisor's periodic liveness/readiness probe needs an HTTP or RPC surface over the same - `collectOrcadHealth()`; that surface does not exist yet. -- **Supervision of an unscoped fallback daemon.** When the durable `systemd-run --user --scope` - launch is unavailable (see [above](#two-long-lived-processes-not-one)) orcad and its daemon - share one service cgroup, and a combined-unit stop cannot preserve live terminals. There is - no mechanism that re-isolates such a daemon after the fact. -- **libc slot.** There is no honest health value to publish until native libc detection owns - it. -- **`degradations[]`.** The readiness contract does not publish this collection yet. -- **Credential administration** (list / revoke / rotate devices, expiring pending offers, - structured security logging). -- **Pinned-port fail-closed.** A pinned `--port` still falls back to an OS-assigned port on - conflict. -- **Reconciling `webClientUrl` with reachability** under the loopback default. -- **State-schema rollback rules.** -- **Daemon log rotation.** `/logs/daemon.log` grows unbounded. diff --git a/docs/reference/orchestration-configured-agent-aliases.md b/docs/reference/orchestration-configured-agent-aliases.md deleted file mode 100644 index 2dc95ccd100..00000000000 --- a/docs/reference/orchestration-configured-agent-aliases.md +++ /dev/null @@ -1,22 +0,0 @@ -# Configured command aliases for orchestration workers - -A worker can use the name of a direct executable configured in **Settings → Agents**. -Choose the built-in agent whose command-line interface the executable implements, -then set that agent's command override to the executable name or quoted full path. -For example, configure Codex's command as `codex-fugu`, then select it with -`orca orchestration worker-start --agent codex-fugu` and the normal placement options. - -The execution host resolves its own configuration. Launch receipts use the canonical -agent (`codex` in this example), and model/effort handling reuses that agent's existing -launch rules. A receipt records applied launch preferences; it does not prove provider -entitlement, successful generation, or an arbitrary vendor's model selection behavior. - -Aliases require a single executable token. Commands containing interpreter arguments, -environment assignments, or shell wrappers are not aliases. Multiple built-in agents -configured with the same executable name are ambiguous and require the canonical agent -ID. Disabled launchers remain disabled. An unconfigured name is refused even if it is -on PATH; Orca cannot infer a compatible launch interface from a process name. - -The same rule applies to folder workspaces and git worktrees. For a remote worker, -configure the command on its execution host. Older hosts may refuse aliases they do not -support. Configuring a command does not create new status producers or grant permissions. diff --git a/docs/reference/pnpm-install-policy.md b/docs/reference/pnpm-install-policy.md deleted file mode 100644 index 6e53ecac77a..00000000000 --- a/docs/reference/pnpm-install-policy.md +++ /dev/null @@ -1,92 +0,0 @@ -# Native dependency install policy - -Ordinary `pnpm install` installs optional native dependencies for the current OS -and CPU only. This applies to local development and root-project CI jobs, -including jobs using `.github/actions/install-node-dependencies`. Mobile and -cloud projects with their own workspace configuration are separate. - -The one cross-target build in the repo is macOS: `pnpm build:mac` and the four -macOS packaging workflows produce both x64 and arm64 artifacts from an arm64 -runner. Before packaging for another architecture, widen the CPU set: - -```sh -pnpm install:release -``` - -This runs `pnpm install --frozen-lockfile --cpu=current,x64,arm64`. It never -widens the OS set: every packaging job runs on a runner whose OS matches its -target, so cross-OS installs are never needed. The macOS workflows pass -`--cpu=current,x64,arm64` directly; Windows and Linux packaging jobs use a plain -host-only `pnpm install --frozen-lockfile`. Keeping `current` in the list -preserves the host's build tools alongside the target resources. An install for -another target does not itself cross-compile native addons. - -## Packaging guard - -electron-builder only logs a warning for a missing `extraResources` source and -continues, so without a check a foreign-architecture slice would ship silently -broken. `beforePack` in -[`config/electron-builder.config.cjs`](../../config/electron-builder.config.cjs) -therefore calls `assertPackagedNativeVariantsInstalled` in -[`config/packaged-runtime-node-modules.cjs`](../../config/packaged-runtime-node-modules.cjs), -which fails the build when the target platform/architecture's native variants -are not installed: `sherpa-onnx-*`, `@parcel/watcher-*`, and on Windows the -node-gyp addon `@vscode/windows-process-tree`. The error names every missing -package and gives the remedy that fits: another architecture's variants come -from `pnpm install:release`, the Windows addon does not (see below). - -Windows packaging requires a Windows host. `@vscode/windows-process-tree` is an -`os: win32` npm addon, so it is installed only where that matches; -`@orca/windows-registry` is a workspace package that links on every host, but -its native binary is still compiled only on Windows. Both are compiled only by -the Windows-only rebuild in `config/scripts/rebuild-native-deps.mjs` -(`allowBuilds` in `pnpm-workspace.yaml` keeps pnpm itself from running node-gyp -for them). The guard checks `@vscode/windows-process-tree` alone because the -workspace link is present everywhere and proves nothing. `pnpm install:release` -does not help on macOS or Linux because it does not widen the OS set. - -Tests that inspect installed Windows addons and their packaging closure run on -Windows, where those dependencies are required. The PR Windows lane explicitly -includes them. Loading the packaging config tolerates absent Windows addons; only -`beforePack` rejects Windows packaging until they are installed. Patch-source -assertions, fixture tests, and the isolated real patch-install test continue to -run on other hosts. - -## Existing checkouts - -A narrowing incremental install can leave previously installed variants behind. -Stop development processes using this checkout, remove **this checkout's** -`node_modules`, then run `pnpm install --frozen-lockfile` to obtain a fresh host -install. Do not use `--force`: pnpm 12 documents that it installs optional -packages even when their OS/CPU/libc do not match. The shared pnpm download store -is separate; this change does not clear it. - -## Measurement: macOS arm64, pnpm 12.0.0 - -Measured 2026-09-12 with the same package manifest, lockfile, patch files, and -existing download store in two fresh install directories, both with -`--frozen-lockfile --ignore-scripts --offline`. One used the earlier broad -policy (`--os=current,darwin,linux,win32 --cpu=current,x64,arm64`); the other -used `--os=current --cpu=current`. Numbers are logical file bytes in the virtual -store, **not unique disk usage**: pnpm hardlinks and APFS clones may share -storage. - -| Measure | Broad install | Host install | Reduction | -| ------------------------------------------ | ------------: | ------------: | --------------------: | -| Package directories | 1,296 | 1,206 | 90 | -| Regular files | 53,972 | 52,915 | 1,057 | -| Logical package bytes | 2,485,153,773 | 1,159,396,649 | 1,325,757,124 (53.3%) | -| Logical package GiB | 2.31 | 1.08 | 1.23 | -| Install wall time, single warm-cache trial | 11.98 s | 11.25 s | 0.73 s | - -| Native family | Broad MiB | Host MiB | -| ------------------- | --------: | -------: | -| Canvas | 235.7 | 25.8 | -| Sherpa speech | 235.2 | 71.6 | -| oxlint and tsgolint | 234.0 | 32.7 | -| SWC | 225.8 | 24.4 | -| TypeScript | 158.4 | 26.2 | - -Electron downloads and native rebuild outputs are absent from both trials. The -single warm-cache timing pair does not establish a speedup, and Linux/Windows -results were not measured; do not present the macOS numbers as CI savings. diff --git a/docs/reference/qoder-integration.md b/docs/reference/qoder-integration.md deleted file mode 100644 index 5aa5c55147e..00000000000 --- a/docs/reference/qoder-integration.md +++ /dev/null @@ -1,90 +0,0 @@ -# Qoder CLI integration - -Orca registers `qodercli` as `qoder`: detection, picker/settings, desktop and mobile -identity, prompt launch, permission flags, managed hooks, workspace trust, and -session resume use the existing TUI-agent and hook-status paths. - -## Verified contract - -Verified on macOS with Qoder CLI 1.1.64, including its versioned executable -`qodercli-1.1.64`: - -- `--prompt-interactive` starts an interactive prompt; `--resume ` resumes a - session; `--dangerously-skip-permissions` is the permission bypass flag. -- Hooks use the Claude-shaped nested configuration in `~/.qoder/settings.json`. - Orca registers its own `/hook/qoder` source and preserves user hooks. -- Real SessionStart, UserPromptSubmit, Notification and SessionEnd events are - captured in `src/shared/__fixtures__/qoder-no-account-hooks.jsonl`, including - `source: resume` with the original session ID. -- Trust is `permissions.trustDirectories` in that settings file, using the - canonical workspace path. A workspace `.trusted` marker did not bypass the - trust dialog in this version. Remote trust uses the execution host filesystem. -- Qoder emits its `◇ … | Ready` title while the trust menu owns input. Readiness - therefore requires the live composer text, not merely this title or silence. - Raw PTY fixtures under `src/main/runtime/__fixtures__/qoder-*` cover trust, - unauthenticated prompt handling and the ready composer. -- A hidden Orca dev instance launched Qoder from New Tab in a folder workspace. - Its icon, label, terminal rendering and failed-turn indicator were inspected. - A harmless prompt reached Qoder, which reported its credit usage limit; the - canonical hook store recorded Qoder identity, session metadata and failure. - -## Compatibility and limits - -Windows hooks explicitly select Qoder's documented PowerShell shell, avoiding -an assumption that Qoder uses Git Bash just because it is installed. They reuse -the existing managed `.cmd` payload without adding an encoded PowerShell hop. -The configuration shape and preservation are tested; Windows, Linux and SSH -execution have not been exercised live here. - -Resume requests to remote hosts require `agent-session.qoder-resume.v1`, so an -older host is not sent an agent enum it cannot accept. Remote hook installation -and trust preservation have automated coverage. - -Successful model output, tool execution and live permission dialogs remain -unverified because the available account has exhausted its credits. Hook event -mapping for these paths follows the official documentation. The China executable -`qoderclicn` is not registered; it was not available for verification. - -Every launch Orca starts (New Tab, workspace and draft launches, automations) -pre-trusts the workspace at spawn while the agent-wide "Trust the folder when -Orca starts an agent" setting is on. A hand-typed `qodercli` in a plain terminal -can still show Qoder's trust dialog. - -## Sources - -- [Official hooks documentation](https://docs.qoder.com/cli/hooks) -- [#15291](https://github.com/stablyai/orca/pull/15291): registration and hook proposal -- [#13311](https://github.com/stablyai/orca/pull/13311): icon asset (commit - `b56197025530adb1b82d97a364a8c3d4a02c0d42`) and agent integration -- [#8611](https://github.com/stablyai/orca/pull/8611) and - [#12910](https://github.com/stablyai/orca/pull/12910): contributor implementations -- [#16540](https://github.com/stablyai/orca/issues/16540): Qoder misidentified as Gemini - -Contributor patches informed the integration; they were not applied wholesale. -In particular, the trust marker proposal was replaced using the installed CLI's -observed settings write, and no renderer workaround was added because the live -WebGL terminal rendered correctly. - -## Cross-review of open proposals - -Reviewed the current patches for all six open Qoder PRs on 2026-09-28: - -| PR | Incorporated or confirmed | Deliberate differences | -| --- | --- | --- | -| #7502 (vincent-lxc) | Agent registration, mobile parity, identity before Gemini | Use structural titles and the current public permission flag. | -| #8611 (Eridanus117, building on #7502) | Qoder hook source, notification/permission mapping, session resume | 1.1.64 supports `--prompt-interactive`; preserve startup session identity and ignore compact restarts. | -| #9655 (xingqingzzp-gif) | Cross-checked minimum registration coverage | Placeholder icon and detection-only scope are superseded by the fuller integration; no code copied. | -| #12910 (jyang2004) | Reuse the Claude-compatible installer with an event subset, plus remote installation | Use observed `.qoder` path and `/hook/qoder`; Qwen and the China build remain separate work. | -| #13311 (sorrycc) | Icon, structural title disambiguation, headless detection, display glyph cleanup | No DOM renderer override: 1.1.64 renders correctly in the live WebGL terminal. No unverified executable alias. | -| #15291 (adlternative) | Versioned binary recognition, trust preflight wiring, skills mapping and resume | Trust comes from settings, not `.trusted`; include notifications, permission and failure events with Qoder-specific normalization. | - -The rendered left sidebar was checked in the hidden Electron app using a temporary -folder workspace and real Qoder hooks. The workspace fixture needed a parent path -on its project group to appear in the sidebar. No agent status was injected into -the renderer. A real prompt updated the row text and returned a credit-limit error; -the row and workspace card showed failure with the Qoder icon. A DOM observer -recorded the prompt-row transition before failure. Permission/waiting and successful -completion remain automated event-mapping tests, not live model/tool verification. - -Commit co-authors credit all six proposal authors for the implementation and -registration groundwork, including xingqingzzp-gif’s minimal registration proposal. diff --git a/docs/reference/relay-regional-placement.md b/docs/reference/relay-regional-placement.md deleted file mode 100644 index 2e984ab0789..00000000000 --- a/docs/reference/relay-regional-placement.md +++ /dev/null @@ -1,39 +0,0 @@ -# Relay regional placement - -Orca selects a Relay region in the Electron main process before requesting a new assignment. The -director publishes an allowlisted region catalog containing only HTTPS cell subdomains of that -director. Orca discards one warm-up `/health` request per probe origin — a cold request pays TCP and -TLS setup that can exceed the round trip it measures — then takes three bounded samples and compares -regions by their minimum. A wide spread still rejects a region, but only a genuinely flapping one. -The stable choice is cached for 24 hours, and a cached region changes only when the alternative is -materially faster. - -A region wins only against a measured competitor. If any region in the catalog is rejected or cannot -be measured, Orca sends no hint rather than selecting the sole survivor. Sending no hint is not -neutral placement: the director assigns `preferredRegion ?? RELAY_DEFAULT_REGION`, and the default -is `us-central1`. So an `asia-east2` user whose `us-central1` probe fails or flaps once is placed in -`us-central1` for that refresh. That trade is accepted because the relay database is -`us-central1`-only, and it is bounded: the withheld hint is cached for one hour, not the 24 hours a -chosen region gets, so the next hour re-measures. An origin that fails its warm-up probe is dropped -before the sampling rounds, so an unreachable region costs one probe timeout rather than four. - -After a control socket registers, Orca probes the cell it actually landed on, once per cell URL per -process. The cache is deleted only when it names a region other than the best measured one and the -assigned cell is more than three times slower than that region — a far cell under a cache that still -names the best region means the director declined the hint, and re-measuring would return the same -answer. Self-heal skips an absent, expired, or no-hint cache, and never runs under -`ORCA_RELAY_REGION_OVERRIDE`. - -The assignment request sends only `preferredRegion`. It does not send latency, IP address, country, -pairing data, or credentials. Catalog, probe, and cache failures fall back to an assignment without -a region preference. A rolled-back director that rejects the new field is retried once without -only that field while preserving reconnect behavior. - -The selection measures the desktop network path. Folder workspaces and SSH workspaces share the -same local broker and do not run probes on remote hosts. The phone continues to connect to the -exact cell URL in the desktop pairing payload, so its location is not measured independently and -no mobile protocol update is required. - -For deterministic local diagnostics, set `ORCA_RELAY_REGION_OVERRIDE` to `us-central1` or -`asia-east2` before launching Orca. The override is not an end-user setting and is not written to -the preference cache. diff --git a/docs/reference/remote-wire-compatibility.md b/docs/reference/remote-wire-compatibility.md deleted file mode 100644 index 6e3be37219d..00000000000 --- a/docs/reference/remote-wire-compatibility.md +++ /dev/null @@ -1,397 +0,0 @@ -# Remote wire compatibility - -Orca's remote-server feature pairs a desktop client to a remote Orca runtime, and -users update the two independently. **Mixed versions are the normal state**, not an -edge case. This page is the contract for changing anything a paired client and host -exchange: the runtime RPC envelope, the terminal binary stream, and the content -either side publishes over them. - -`src/shared/protocol-version.ts` says when to bump `RUNTIME_PROTOCOL_VERSION`. This -page covers the changes that do _not_ bump it and are therefore easy to get wrong. - -## Rule 1 — a new optional JSON field on an existing frame is safe - -Every JSON payload is parsed with a decoder that ignores unknown keys (zod `.strip()` -on RPC params, `JSON.parse` on stream frames). An older peer that has never heard of -the field simply does not read it. - -Safe: - -```ts -// host adds a field; older clients ignore it -encodeTerminalStreamJson({ kind, cols, rows, hiddenOutputReason }) -``` - -**The field is safe only for as long as every reader treats it as optional.** The -moment a newer client _requires_ it, that client is broken against every host that -predates the field — which is the same defect as removing a field, just discovered -later. If new behavior depends on the field being present, that is Rule 2: negotiate -it, or make the reader fall back. - -## Rule 2 — a new stream opcode is NOT safe; negotiate it - -`decodeTerminalStreamFrame` returns `null` for an opcode it does not know, and -`runtime-rpc.ts` drops that frame without an error: - -```ts -const frame = decodeTerminalStreamFrame(bytes) -if (!frame) { - return // silently dropped — the sender never learns -} -``` - -So a new opcode sent to an older peer does not fail loudly. It vanishes, and the -feature behind it appears to hang. Input sent under a new opcode is swallowed. - -A new opcode must be announced in the subscribe handshake and sent only after the -peer confirms it. The existing pattern is `SetOutputPaused` (opcode 16): - -- the client advertises support in the `Subscribe` frame's `capabilities`; -- the host echoes `capabilities: { outputPause: 1 }` on the `subscribed` event; -- the client sends opcode 16 only after that echo (`stream.supportsOutputPause`); -- the host only acts on opcode 16 when it negotiated it (`stream.supportsOutputPause`). - -Reuse an existing opcode with a new optional payload field (Rule 1) whenever that -expresses the change; reach for a new opcode only when framing genuinely differs. - -Opcode numbers are permanent. See the `Ack = 13` and `ClaimViewport = 14` comments -in `src/shared/terminal-stream-protocol.ts` for why a shipped number cannot be -reused even if the feature behind it is removed. - -## Rule 3 — changing what the host publishes breaks old clients with no wire change - -The frame shape can be untouched and the skew still real, because clients react to -frame _content_. PR #12641 is the worked example: the host stopped synthesizing a -finished agent status, and clients running older code saw different content in an -identical frame. - -Treat these as wire changes even though nothing in the codec moves: - -- a field the host stops populating (an old client reading it now sees `undefined`); -- a value whose meaning, units, or nullability changes; -- content the host stops synthesizing, trims, or starts deriving from a new source; -- a frame the host stops sending, or starts sending, on an existing path. - -If old clients cannot interpret the new projection correctly, gate it behind a -runtime capability the same way Rule 2 gates an opcode. - -## Rule 4 — an enum arm set is a wire surface; unknown arms must degrade, never reject - -A closed `z.enum` in a client-side reply schema is a version claim: it asserts the host -will never send an arm this build has not heard of. A newer host that adds one arm then -has its whole reply refused, or has the row carrying it silently dropped, even though -every member the client actually reads is present and well-formed. - -Declare the arm set open instead, with `openEnum` in `src/shared/zod-salvage.ts`: - -```ts -// unknown arm degrades to a member the reader already handles; a non-string stays fatal -status: openEnum(GIT_BRANCH_COMPARE_STATUS, 'error') -``` - -Do not reach for `.catch()`. It swallows absence and the wrong type as well, which turns -a member the reader depends on into a silent default. - -A fallback does not have to be an arm. `git.status`'s `area` is the worked example: -`staged`, `unstaged` and `untracked` each grant an affordance, so coercing an unknown area -to one of them offers stage, unstage or commit against a row the client cannot place. It -degrades to absent instead, which withholds all three — every reader is an equality check -against a known arm, so the row lands in no section — while keeping the row itself. That -last part is the point. Dropping the row would also drop it from the unresolved-conflict -gate, and a conflicted worktree that looks clean is granted a hosted-review create it -should not have. Withholding an affordance is a degrade; removing the evidence a gate -reads is not. - -A fallback is only ever allowed to shape a _reading_. If the member is sent back to the -host — a token the client echoes into a later call's params — pass it through as -`z.string()` and let the send site keep it verbatim. `hostedReview`'s `provider` is the -case: the eligibility reply names it and the create call returns it, so an -`openEnum(..., 'unsupported')` there does not soften how the client reads a newer host's -provider, it puts `unsupported` on the wire and makes that host refuse its own. A -reply-schema fallback must never shape a param. Gate on the token instead, where the -client decides what it is willing to do with an arm it does not know. - -## Worktree activation belongs to a viewer - -Worktree creation and activation default to the host desktop for host/CLI requests, -and to the caller for paired desktop/web requests. Mobile creation still activates -the host renderer to provision setup and default tabs. A headless host does not -borrow a paired client's view. Catalog and session updates continue to reach every -subscriber independently of navigation. -Headless host/CLI creates provision their shells, setup, and default tabs on the -execution host in the background, without depending on an observer to open them. -The same fallback applies when an attached host renderer is unavailable or -reloading, including explicit `all` requests that target the host; availability is -checked after creation, rather than before its awaits. - -`activateWorktree` client events carry an optional `navigation` field. Only explicit -`clients` or `all` requests publish these events; updated clients ignore events -without that address, including implicit broadcasts from older hosts. This is an -optional JSON field (Rule 1): older clients ignore it and still understand explicit -activation. The new host's default publication changes deliberately remove implicit -navigation for older clients too, rather than retaining the unwanted behavior. -An older host cannot express explicit follow intent to an updated client, so that -client must open the workspace itself. -Accepted navigation is also fenced through repository/worktree discovery; bridge -cleanup, reconnection, or re-pairing revokes any activation still awaiting a fetch. - -Older CLIs hardcode `navigation: 'all'` for `--activate` and `--run-hooks`. The host -recognizes their `cliProvenanceRequest` and normalizes that automatic target to -`host`. An API caller without the CLI marker can still explicitly request `clients` -or `all`. Neither creation nor activation navigates remote clients by default. - -Coverage: `multi-client-navigation-isolation.integration.test.ts` exercises real -paired WebSocket clients and host/headless creation, the renderer bridge tests -exercise older unaddressed publications, and -`paired-worktree-activation-isolation.spec.ts` checks a paired desktop remaining in -a local workspace while its remote catalog updates. - -## Session search agent negotiation - -`aiVault.searchStatus` optionally advertises `supportedAgents`; current search clients -send their own `supportedAgents` with `aiVault.searchSessions`. These are string lists, -so a future provider name does not make a peer reject the capability reply. The client -narrows explicit agent filters to the host's list before calling its request parser. -The host narrows retrieval to the client's list before publishing a page. - -A peer without this field uses the frozen v1.4.211 search vocabulary. The existing -`supportsQoderHistory` flag proves CodeBuddy, ZCode, and Qoder support; -`supportsJcodeHistory` independently proves Jcode support. An explicit list takes -precedence over both flags. If only the status method is missing, the client still -searches the conservative legacy subset; other status errors propagate. Empty host -intersections keep the requested filters and use the existing no-match scope, so -consent, readiness, and unknown-scope results retain their normal precedence. -Local IPC advertises this build's full list, and every remote leg negotiates separately. - -## Enforcement - -`tests/e2e/cross-version-wire/cross-version-terminal-wire.unit.test.ts` runs the real -host RPC methods and the real renderer multiplexer from two builds against each -other — current working tree against the newest release tag, in both skew -directions — over one scripted terminal journey (subscribe, input, hide/reveal -snapshot, drop, reconnect). - -Run it with: - -```bash -pnpm exec vitest run --config config/vitest.config.ts tests/e2e/cross-version-wire/cross-version-terminal-wire.unit.test.ts -``` - -It fails when a frame is refused by the receiving build's decoder (Rule 2), when the -observed frame sequence changes (Rule 3), or when published snapshot content or -negotiated capabilities differ from the contract. Repeated frame shapes are compared -by corresponding journey occurrence (initial, reveal, reconnect), so a field removed -from one occurrence cannot hide behind a sibling that still publishes it. Adding an -optional field keeps the suite green (Rule 1); making a client depend on that field -turns the new-client/old-host pairing red. - -### Never write down what the old side has - -The baseline is whichever release tag is newest, so it moves on every cut. An -expectation of the form "the old side does not have X" — a `not.toHaveProperty`, a -`not.toContain`, a hard-coded field list — stops being true the first time a release -ships X. The suite then reddens on whatever pull request is in flight, with no code -change anywhere, and the job trains people to ignore it. That is worse than no test, -because a rolling baseline eventually contains every additive field the wire has, and -adding one is the sanctioned way to evolve it. - -Derive the expectation from the baseline that was actually checked out: - -- for a published frame, pair each build against a client of its own version and - compare the skewed pairing against that same-version reference, so the expectation - is whatever that build publishes today. Compare repeated frames by corresponding - occurrence with `comparePublishedFieldOccurrences` in - `tests/e2e/cross-version-wire/published-field-shape.ts`; never union keys across - initial, reveal, and reconnect frames, because a sibling can mask one occurrence's - removed field; -- for a negotiated surface, read the old build's advertised capabilities and - registered method names from its checkout, and assert they agree with each other - rather than asserting the old build lacks them; -- for a "client too old to know X", derive that client's advertised list by removing - X from the baseline's own list, so the gate stays exercised after X ships. - -Name the direction in the assertion. `new client against old server` and `old client -against new server` fail for different reasons, and the host is the only side that -authors a published frame — the terminal `terminalOwner` false positive on 2026-08-29 -was misread as a new client sending an unknown field when the old server was -publishing it. Two things are still safe to state literally: the current build's own -contract, and an invariant that holds for every version. - -Pinning a legacy ref is the fallback when a contract genuinely needs a release from -before a feature shipped, as `cross-version-browser-placement.unit.test.ts` does with -`LEGACY_BROWSER_PLACEMENT_RELEASE_REF`. It does not rot on a cut, but it is -hand-maintained, so prefer deriving. - -`tests/e2e/cross-version-wire/cross-version-agent-session-wire.unit.test.ts` pairs the -same two builds over the structured `agentSession.*` surface. Because a released build -cannot name a capability string its own source never contains, the old side's advertised -list and registered method names are read from the extracted checkout rather than -hand-written. It covers the three skews that surface can fail on: - -- an old client — advertising the baseline's list minus this capability — is told the - whole surface does not exist and reaches no host method; -- a new client against the old dispatcher always gets an answer rather than silence, - and `method_not_found` for every method that release does not register, so the - absence is visible during negotiation instead of by calling; -- a cursor survives a host restart: a reattach at the client's fence is refused as stale - with the live one attached, a write still carrying that fence is delivered (writes are - named by their target, and every released client still sends a fence), and resuming from - the held cursor replays only what it missed. - -Run it with: - -```bash -pnpm exec vitest run --config config/vitest.config.ts tests/e2e/cross-version-wire/cross-version-agent-session-wire.unit.test.ts -``` - -The harness covers the terminal stream and the structured agent-session surface. It does -**not** cover the session-tab sync channel, legacy agent-session publications, file or Git -RPCs, mobile/E2EE framing, or the relay transport. A change on those paths still needs its -own reasoning against the three rules above. - -## Worked example: `agentWait` on terminal and worker reads - -`terminal.show`, `orchestration.workerShow` and `orchestration.federationShow` carry an -optional `agentWait` naming a pane parked on a prompt only a human can answer. It is Rule 1 — -a new optional field — but it has a second state that Rule 1 alone does not describe, and -getting that wrong turns a skew into a false "nothing is blocked". - -- **present object** — this pane is waiting, with the evidence that proved it. -- **present `null`** — the host evaluated this pane and nothing proves a wait. -- **absent** — the host never evaluated it: it predates the field, the worker identity was - unverifiable, the pane was unreadable, or the agent probe did not answer in time. - -A new client against an old host sees the field absent, which is why absence must read as -_unknown_ and never as _not waiting_. Collapsing absent into `null` at any hop — including a -convenience `?? null` in an RPC handler — makes an old or unreachable peer indistinguishable -from a healthy idle worker, which is the exact failure the field exists to remove. - -An old client against a new host ignores the key, as Rule 1 allows. New members added to -`RuntimeTerminalWaitBlockedReason` are also Rule 1: no consumer switches exhaustively on it, -and both the CLI and worker-start interpolate it as an opaque string. - -## Worked example: the `turn` journal item and its transitional downgrade - -The structured chat journal records a turn as a first-class item, -`{ kind: 'turn', turnId, state, userItemId?, startedAt?, completedAt?, durationMs? }`, where it -used to write `{ kind: 'status', text, turnLifecycle }`. Nothing in the codec moves, but it is -Rule 3: a client that predates the item does not know the kind and renders it as a text bubble -with no text. So the item is gated on a client capability, `agent-session.turn-item.v1`. - -The gate lives at the RPC boundary only, in -`src/main/runtime/rpc/methods/structured-agent-session-turn-item-capability.ts`, composed -around `agentSession.history` and `agentSession.subscribe` next to the background-task -projection. A client that does not advertise the capability receives every `turn` item -rewritten to the legacy status form with the full lifecycle under `turnLifecycle`; a client that -advertises it, and any in-process caller, receives the canonical body. The journal, the status -feed, and every host-side reader keep the `turn` item; `readAgentJournalTurn` in -`src/shared/agent-session-turn-record.ts` reads either form, so a new client against an old -host that still writes the status row also works. - -An old client against a new host sees the status row it always did. A new client against an old -host advertises a capability the host ignores and reads the status row through the shared -reader. The downgrade is transitional: once no supported release lacks the capability, delete -the projection module and the capability check, and leave the reader. - -The cross-version suite derives the old client's list by removing this capability from the -baseline's own list, per the rule above, so the downgrade stays exercised after a release ships -it. - -## Known debt: JSON-RPC errors drop Node's string code - -An error raised on an SSH host crosses the relay as JSON-RPC, and -`ssh-channel-multiplexer` rebuilds it with the TRANSPORT's numeric `code`. Node's -string code — `'ENOENT'`, `'EACCES'` — does not survive, so a caller on this side -cannot ask what kind of failure it was. - -`isENOENT` in `src/main/ipc/filesystem-path-containment.ts` pays for that by also -matching Node's canonical message text, which is what makes remote worktree creation -work. The cost is that a host can make an unrelated failure read as "absent" by -putting that sentence in a message. - -The exit is Rule 1: carry the original string code in a new optional field on the -error payload and read that instead. An old host omits it and the message match still -covers them; once hosts that send it are the floor, the message match can be deleted -rather than lived with at its ~10 call sites. Narrowing `isENOENT` back to `.code` -without doing this reinstates the bug — the transport has already overwritten it. - -## Known hazard: clients ignore host-published failure fields on client-placed pages - -`RuntimeMobileSessionBrowserTab` — the browser tab a host publishes on the session-tab sync -channel — permits `placement`, `loadError` and `certificateFailure` together. But for a tab -whose `placement.kind` is `'client'` the engine runs in the client's own app: the failure is -raised by the local guest webview, and the host has no view of it (`RuntimeBrowserClientPage`, -what the registry actually publishes from, carries neither field). Clients from -this version on therefore refuse host ownership of both records for client-placed pages -(`web-session-tabs-sync.ts`, the `placement?.kind !== 'client'` carve-outs) — without that, -each metadata snapshot deletes the locally recorded failure and the page's failure overlay -disappears mid-navigation. - -The hazard is forward-facing and Rule 3 shaped. A host that later starts publishing -`loadError` or `certificateFailure` for a client-placed page reaches these clients as content -they silently drop, so the host would see no error and no effect. Publishing it has to be -capability-gated, with the carve-out narrowed to clients that did not negotiate the -capability. Note the cross-version harness does not exercise the session-tab sync channel, so -nothing fails if this is forgotten — this note is the only record. - -A related carve-out covers `title`, `url`, `loading`, `canGoBack` and `canGoForward` -(`resolveMirroredBrowserPageContent`), and for those the hazard is already live rather than -forward-facing: the host does publish them, from a `RuntimeBrowserClientPage` it can only learn -about second-hand through the client's own `browser.clientHost.pageMetadata` calls. Its copy -therefore starts at the registry defaults (`'Browser'`, the create-time url), and while those -publishes are failing it never leaves them. - -That copy is not simply behind, though, and a client must not treat it as such. When a lease -reattaches, the host refreshes the page from the client host's own inventory -(`runtime-browser-client-page-recovery.ts`), which reads the live guest — so it can be strictly -fresher than a local row whose pane is unmounted and whose metadata publisher was disposed with -it. A client that ignores the host url is relying on its own guest to re-answer on remount, -which `ClientHostedBrowserPagePane`'s mount-time `syncNavigation` is what makes true. - -These five are therefore refused only by the client whose guest actually runs the page: -`placement.browserHostClientId` is compared against this client's own host id -(`readBrowserClientHostId`). Main stamps that id into the guest-hosting window's -`additionalArguments` at creation, and the preload reads it back out of its own argv — the answer -has to be there before the first snapshot is interpreted, which is earlier than any IPC handler a -renderer could wait on. Every other viewer — a second desktop, the web client, which installs no -page renderer at all, the dashboard pop-out, which is deliberately left unstamped — keeps tracking -the host, which is the only reason a mirrored viewer shows anything but its first snapshot -forever. Improving what a _second_ client sees still means fixing the publish, not the carve-out; -the carve-out no longer stands in the way of it. - -The two failure fields above are deliberately left on the looser `placement?.kind !== 'client'` -predicate. It is unobservable today — the host publishes neither field for a client-placed page at -all, so a mirror has nothing to take either way. If the capability-gated publish this section -anticipates ever lands, narrow them the same way rather than by placement kind: a mirror should -take a failure it cannot otherwise see, and only the hosting client should refuse it. - -## Known hazard: on the mobile surface a scope refusal is not a missing method - -The agent-session harness above asserts that a peer probing an unknown method is told -`method_not_found`. That holds for the runtime-scoped surface and **not for the mobile one**. - -`runtime-rpc-websocket-dispatch.ts` checks `MOBILE_RPC_METHOD_ALLOWLIST` and answers -`forbidden` _before_ it calls the dispatcher. A method a desktop predates is on neither the -allowlist nor the registry, and the gate answers first, so a phone never sees -`method_not_found` for it. `method_not_found` would reach a mobile-scoped device only for a -method that is allowlisted but unregistered, and `src/main/runtime/mobile-rpc-allowlist.test.ts` -requires every method the phone calls to be both, so that combination cannot ship. The reverse -— a registered method that no release has allowlisted yet — is the skew that does occur. - -So a phone-side "is this host new enough to serve X?" probe must read **both** codes as -absence. Keying it on `method_not_found` alone compiles and passes every same-version test -while never firing. The Files and Git fallbacks have read both since they shipped -(`isMobileMethodUnavailableError`, `isMobileGitUnavailable`); the Relay pairing probes were the -outlier, and the cost was a write-once direct-relay upgrade journal, holding a pending resume -secret, that an old desktop could never retire -(`mobile/src/transport/pairing-relay-rpc-unavailable.ts`, with the gate pinned desktop-side by -`src/main/runtime/runtime-rpc-mobile-unknown-method-scope-refusal.test.ts`). - -Widening is safe only where `forbidden` cannot also mean a real authorization failure. Prove -that per probe rather than globally: the pairing handlers cannot emit it (an unwired provider -answers `runtime_error`), a bad or revoked token answers `unauthorized`, and the gate is one -of only two places in `src/main` that emits the code at all. A probe whose handler _can_ -refuse by authorization must not be widened. - -The cross-version harness dispatches as a mobile client but calls the dispatcher directly, so -it never runs this gate and nothing reddens if any of the above is forgotten. diff --git a/docs/reference/renderer-agent-status-performance.md b/docs/reference/renderer-agent-status-performance.md deleted file mode 100644 index cffad695d25..00000000000 --- a/docs/reference/renderer-agent-status-performance.md +++ /dev/null @@ -1,340 +0,0 @@ -# Renderer agent-status performance - -## Status - -This design is adopted for the renderer's high-frequency agent-status path. It -keeps the existing event semantics while bounding the amount of synchronous -store fanout performed for one IPC burst. - -It lands in slices. This document and the `bench:idle-cpu` harness land first so -the store changes can be reviewed against a baseline someone else measured. Until -the store slice lands, the `setAgentStatuses` / `transactAgentStatuses` actions -and the agent-status write workload described below are not yet on `main`; the -harness measures scale, listener census, and raw publication fanout only. - -## Context - -Orca can display a large expanded worktree lineage inside one virtualized list -row. Virtualizing the root row does not virtualize its descendants, so a -100-worktree lineage can mount 100 `WorktreeCard` instances at once. - -Agent-status IPC events are bursty. The renderer already groups live events into -a 33 ms window, but the original flush applied every queued event with a -separate Zustand write. Zustand synchronously visits every listener for every -publication. The resulting work therefore grew with both the number of status -events and the number of mounted subscriptions: - -```text -burst work ~= status events x store listeners x selector work -``` - -A production trace captured the renderer repeatedly entering -`flushLiveAgentStatusBurst -> applyAgentStatus -> setAgentStatus -> setState` -through `Set.forEach`. A deterministic 100-worktree fixture reproduces the -structural multiplier; see "Baseline on `main`" below for the currently measured -listener count and publication cost. - -The production app later recovered substantially when all configured remote -hosts were removed. Host removal can stop relay/reconnect traffic, remove -mounted remote worktrees, or both, depending on host type and removal options. -That observation identifies remote presence as the production trigger but does -not by itself distinguish traffic volume from mounted-listener fanout. - -A read-only reconnect audit ruled out systematic double status emission from a -full PTY replay: replay bytes bypass OSC status parsing. Reconnect still causes -a full terminal-buffer repaint for every attached remote pane, which is a -separate source of renderer work and remains a follow-up investigation. - -## Goals - -- Keep a large expanded lineage responsive during dense agent-status traffic. -- Preserve every ordered status transition, including repeated updates for one - pane inside the same burst. -- Publish agent-status state once for a deferred live burst. -- Preserve selector identity and child render isolation. -- Make the regression reproducible without relying on a user's production data. - -## Non-goals - -- Changing the client/server status payload or remote protocol. -- Deduplicating status events by pane. -- Changing agent freshness, retention, history, title, completion, or provider - session behavior. -- Changing remote reconnect, PTY replay, or terminal repaint behavior. -- Redesigning lineage presentation or collapsing worktrees automatically. - -## Design - -### Bound mounted subscription fanout - -Sidebar components select cohesive state bundles with shallow equality instead -of registering one listener per field. Derived arrays and maps retain their -existing shallow identity behavior so unrelated store writes do not rerender a -card. Full agent-list mode keeps its child-level subscription boundary; compact -mode passes the already selected rows to avoid selecting the same inputs twice. - -The deterministic 100-worktree fixture pins the resulting listener budget: - -| Surface | Listener budget | -| ------------------------------ | --------------: | -| Worktree card state and caches | 2 | -| Agent-row inputs | 1 | -| Worktree activity status | 1 | -| Closed context menu | 1 | - -Unmount tests require the listener count to return to its prior baseline. In the -bundled prototype, the fixture without seeded agents fell from 8,518 listeners -to 1,218; with 100 visible agent rows the candidate mounted 1,618. Compare -against the census in "Baseline on `main`", which the harness reports directly. - -### Share working-spinner phase without synchronous mount queries - -Working rows keep the existing compositor-driven CSS animation and shared -visual phase. `animationstart` anchors each animation to document time zero. -Deferring the animation query until that event avoids a synchronous style flush -at each mount and restores the shared phase after `display:none` or a motion -preference change. A negative mount-time delay cannot preserve that phase after -an animation restarts. - -### Bound spinner animation overhead - -Working rings keep compositor-driven CSS animation, but repeat the animation -once per day rather than once per second. The same 12 steps per second now -avoid recurring React animation-iteration dispatch. The existing stationary -wrapper and ring rendering stay unchanged. Offscreen containment was evaluated -and rejected after a pixel regression at low zoom on 1x displays. - -The history, isolated measurements, full-app workspace/agent/subagent benchmark, -and limitations are documented in [Spinner rendering performance](./spinner-rendering-performance.md). - -### Fold a burst in event order - -The store exposes the single-update action and two batch forms: - -- `setAgentStatus(paneKey, payload, ...)` retains the positional single-update - API for the immediate live path. -- `setAgentStatuses(updates)` applies a prebuilt ordered list, while - `transactAgentStatuses(operation)` lets IPC derive each update against the - exact staged state before the single commit. - -Both entry points reuse the same single-update state transition. The batch -reducer passes each resulting state into the next update, so a sequence such as -`working -> waiting -> done` retains the same history and timestamps as three -sequential calls. Updates are never keyed or deduplicated before the fold. - -The live IPC queue is spliced before it is processed. This preserves the -existing reentrancy guarantee: a synchronous subscriber can enqueue another -event without causing the current queue to be drained recursively. The first -event outside an active burst remains immediate; events accumulated within the -33 ms window are applied as one ordered transaction. Startup snapshots and -bounded pending-hydration retries use the same transaction path instead of -publishing once per restored pane. - -Each transaction builds pane-routing ownership once with the same first-match -semantics as the standalone resolver. Split-layout leaf membership is indexed -once per layout root, so a large snapshot performs linear tab and leaf work -instead of rescanning every mounted worktree for every pane. - -### Run effects after the transaction - -Generated-title work that requires committed state is deferred until after the -transaction. Accepted updates also request freshness scheduling; the outer batch -coalesces those requests and schedules the shared freshness timer once after its -single commit. Generated-title requests are folded in event order and published -together, including first-write and forced-replacement semantics. Resolved tab -titles are projected while the transaction folds, then final title changes are -published together. Completion-triggered review refreshes remain deferred -microtasks. - -Bulk title application preserves event order and duplicate-tab behavior while -indexing owners once, cloning each changed owner array once, and replacing each -top-level map once. This keeps the post-commit title phase linear in mounted -tabs plus changed titles. - -This separation is important: invoking store actions from inside a Zustand -updater would re-enter the store, while running an effect before the commit -would let it observe stale state. - -## Semantic invariants - -Sequential and batched application must agree on: - -- live and retained agent maps; -- state history, `updatedAt`, and `stateStartedAt`; -- agent identity, model, prompt, tools, assistant messages, and subagents; -- orchestration and provider-session continuity; -- sleeping-session and launch-config recovery records; -- retired/closed-pane rejection and inherited-status suppression; -- retention cleanup and live-map eviction; -- `agentStatusEpoch` and `sortEpoch`; -- automation completion observation across intermediate transitions; -- generated-title inputs, freshness scheduling, and completion refreshes. - -Equivalence tests use fixed timestamps and include repeated same-pane -transitions. A publication-count test subscribes to the real store and requires -one notification for a non-empty batch and none for an empty batch. - -## Benchmark contract - -The benchmark launches an E2E-mode Electron build with the store exposed only -for instrumentation. It creates 100 worktrees in one expanded lineage, verifies -100 mounted cards, captures the store listener census, and then applies seeded -ordered agent-status traffic through the real store action. - -The benchmark measures the synchronous store action, not the live IPC leading -edge or post-commit notification path. A real-store snapshot test covers the -end-to-end budget for 100 panes with auto-generated titles enabled: one status, -one bulk generated-title, and one bulk resolved-title publication. Disabling -generated titles removes that middle publication, independent of pane count. - -The artifact records only fixed diagnostic fields needed for comparison: - -- requested and completed batches and updates; -- store action calls and observed publications; -- elapsed time, throughput, and scheduling drift; -- final-state verification; -- renderer mean, p95, and maximum CPU; -- renderer timer drift and long tasks; -- mounted-card and listener counts. - -Raw process inventories, temporary paths, pane identifiers, and DOM text are -diagnostic-only and must not be embedded in the shareable report. - -Run baseline and candidate on the same machine and OS with the same Electron -build mode. CPU samples from macOS and Linux are comparable within that -constraint; Windows process CPU collection currently cannot support this -comparison. - -## Harness - -`pnpm run bench:idle-cpu` drives `config/scripts/run-idle-cpu-benchmark.mjs`, -which composes four modules: - -| Module | Responsibility | -| ------------------------------------- | ------------------------------------------------------------------------------------------------- | -| `idle-cpu-renderer-scale-fixture.mjs` | Seeds the lineage, agent rows, and sidebar view state; takes the mounted-card and listener census | -| `idle-cpu-renderer-timing-probe.mjs` | In-page timer drift and long-task probe; runs the no-op publication workload | -| `idle-cpu-process-sampling.mjs` | Classifies the Electron process tree and samples per-role CPU/RSS | -| `idle-cpu-synthetic-spinners.mjs` | Measurement-only visible spinners | - -The sampling window extends past `--sample-ms` until the workload settles, and -fails the run rather than reporting a truncated window if the workload overruns -the guard. That is why a 2,000-publication run reports a measured window longer -than the requested one. - -`--zustand-publications` publishes an empty partial through the real store, so -each publication costs exactly one full subscriber visit and nothing else. It -isolates the `listeners x selector work` half of the burst-cost model from -agent-status payload work, and it is store-API independent — it measures the -same thing before and after the batching slice. - -The agent-status write workload (`--agent-status-batches`, -`--agent-status-write-mode`) is not in this harness yet. It depends on -`setAgentStatuses`, so it lands with the store slice. - -## Baseline on `main` - -Measured on `main` at `077f5a11cd4` (macOS, arm64, 16 CPUs), Electron built with -`electron-vite --mode e2e`, headless, 100 worktrees at lineage depth 99 with 100 -seeded agent rows, 10 s warmup and a 30 s sampling window. - -Fixture scale is confirmed by the census rather than assumed: 100 store -worktrees, 100 mounted cards, 100 mounted agent rows, and **9,279 store -listeners**. That listener count is the multiplier the design targets. - -2,000 no-op store publications at a 1 ms cadence, three repetitions. The -listener census was 9,279 in every run. - -| Measure | Median | Runs | -| ------------------------ | ----------: | ------------------------------ | -| Wall time to complete | 12,325.7 ms | 13,148.6 / 12,325.7 / 11,870.5 | -| p50 scheduling drift | 5,150.3 ms | 5,432.3 / 5,150.3 / 4,945.2 | -| p95 scheduling drift | 9,802.9 ms | 10,570.4 / 9,802.9 / 9,390.0 | -| Renderer mean CPU | 18.25% | 20.25 / 18.25 / 15.88 | -| Renderer p95 CPU | 32.59% | 40.07 / 32.59 / 31.28 | -| Renderer timer drift p95 | 7.0 ms | 7.3 / 5.3 / 7.0 | - -2,000 publications requested over 2 s take about 12 s, so the renderer sustains -roughly 160 publications per second at this scale. Each publication is -individually short - the long-task observer recorded zero entries in all three -runs - so the cost surfaces as scheduling drift and sustained CPU rather than as -discrete long tasks. Compare drift and CPU here, not long-task counts. - -The idle control at the same scale with `--zustand-publications 0` reports 6.63% -renderer mean CPU, 17.11% p95, and 1.6 ms p95 timer drift. Roughly 11.6 points -of mean renderer CPU are therefore attributable to publication fanout rather than -to the mounted fixture itself. 200 spinner animations run in both cases, so the -control also bounds the animation cost out of the comparison. - -Reproduce with: - -```bash -pnpm run bench:idle-cpu -- --worktrees 100 --lineage-depth 99 \ - --agents-per-worktree 1 --warmup-ms 10000 --sample-ms 30000 \ - --zustand-publications 2000 --zustand-publication-interval-ms 1 \ - --output /tmp/idle-cpu-baseline.json -``` - -## Results - -Three repetitions used 100 mounted worktrees, lineage depth 99, 100 seeded -agent rows, and verified final state. Medians from the regenerated evidence set -are: - -| Single 2,000-update burst | Sequential | Batched | -| ------------------------- | ---------: | ---------: | -| Status-state publications | 2,000 | 1 | -| Store action time | 3,692.0 ms | 188.7 ms | -| Update throughput | 541.7/s | 10,598.8/s | -| Renderer mean CPU | 36.2% | 2.9% | -| Renderer p95 CPU | 107.3% | 8.2% | -| p95 long task | 4,653 ms | 216 ms | - -The direct store transaction performs 99.95% fewer status-state publications, -spends 94.9% less time in the store action, and processes updates 19.6 times -faster. Renderer mean CPU falls 92.0%, renderer p95 CPU falls 92.4%, and the p95 -long task falls 95.4%. - -The 60-burst × 32-update case at 33 ms is a sustained saturation stress, not a -real-time production SLO. Publications fall from 1,920 to 60 and median store -action time falls from 2,791.9 ms to 323.3 ms. Median completion time falls from -5,298.2 ms to 2,710.6 ms, p95 scheduling drift falls from 3,073.0 ms to 664.9 -ms, and long-task count falls from 57 to 1. Renderer p95 CPU remains saturated -and noisy in this cadence, so it is not used as the discriminating measure. - -The 20-pane artificial OpenCode regression passes with 12.4 ms median key echo, -25.2 ms worst key echo, 19.4 ms maximum timer drift, and zero dropped renderer -backlogs. - -These figures come from the bundled prototype and are restated here as the -target. They are re-measured with the harness when the store slice lands. - -## Acceptance criteria - -- The 100-worktree fixture stays at or below the pinned listener budgets. -- A deferred transaction performs one status-state publication while preserving - ordered final state, including live-map eviction at the 500-row cap. -- A 100-pane startup snapshot performs one status and one bulk resolved-title - publication with generated titles disabled; enabling generated titles adds at - most one ordered bulk publication while preserving final statuses and titles. -- Sequential-versus-batch equivalence tests pass across same-pane transitions - and side-effect-bearing updates. -- Renderer CPU tails and scheduling drift improve in repeated candidate runs. -- The 20-pane artificial terminal test reports no dropped output backlog and no - material typing-latency regression. -- Web typecheck, focused unit tests, lint, max-lines ratchet, and E2E build pass. - -## Compatibility - -This is renderer-local. It adds no RPC field, stream opcode, persisted data, Git -command, or provider-specific contract. Native, WSL, SSH, relay, folder -workspace, and git-worktree status events enter the same renderer action. Mixed -client/server versions therefore need no capability negotiation. - -## Failure containment - -The first live event remains immediate. Startup replay and bounded pending -retries fold synchronously without waiting for the 33 ms live-burst window, but -publish their accepted updates together. Empty batches are no-ops. If an update -is stale or targets retired authority, the reducer skips only that update and -continues folding later events in order. diff --git a/docs/reference/runtime-file-base64-padding.md b/docs/reference/runtime-file-base64-padding.md deleted file mode 100644 index 3715855eaa0..00000000000 --- a/docs/reference/runtime-file-base64-padding.md +++ /dev/null @@ -1,92 +0,0 @@ -# Runtime file Base64 padding - -Padded runtime file writes must have a length divisible by four. Empty strings and -unpadded Base64 with length modulo four equal to zero, two, or three remain valid. -The change rejects exactly the previously accepted strings containing trailing -padding whose total length modulo four is two or three. It does not enforce -canonical unused pad bits or change the alphabet. - -## Boundary evidence - -| Input | Before | After | -| ------------------------- | ------ | ------ | -| `A=` | Accept | Reject | -| `AA==` | Accept | Accept | -| `AAA=` | Accept | Accept | -| `AAAA` | Accept | Accept | -| `''` | Accept | Accept | -| `A` | Reject | Reject | -| `==` | Accept | Reject | -| `AA=A` (interior padding) | Reject | Reject | -| `AA=` | Accept | Reject | -| `A==` | Accept | Reject | -| `AAAA==` | Accept | Reject | -| `AA`, `AAA` (unpadded) | Accept | Accept | - -`Buffer.from('A=', 'base64')` decodes to zero bytes. Rejecting malformed padding at -the RPC boundary prevents an accepted request from silently writing different bytes. - -## Caller census - -| Caller / surface | Reachability and compatibility verdict | -| ------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Desktop `runtime-file-import-client.ts` → `uploadRuntimeFileWithoutClobber` → `writeRuntimeBase64File` | The only production producer of `files.writeBase64` and `files.writeBase64Chunk`. Staging in `filesystem-runtime-upload-staging.ts` encodes the complete file with `buffer.toString('base64')`; newly rejected values cannot be produced. | -| Desktop single-frame uploads | Sends the staged string unchanged when its length is at most 512 × 1024 characters. Standard Node Base64 always has length divisible by four. Empty files remain accepted. | -| Desktop chunked uploads | Slices the encoded stream at 512 × 1024 = 524,288 characters, divisible by four. Every offset and every complete chunk is quartet-aligned. The final chunk is the difference of two multiples of four, including when it ends in `=` or `==`. No separately assembled final chunk or per-chunk padding is added. | -| Web implementation of `stageExternalPathsForRuntimeUpload` | Returns an empty source list; no file-write payload is produced. | -| Mobile file editor | Uses `files.writeTerminalArtifact` with text content and its separate schema. Does not reach this predicate. | -| Mobile clipboard / image attachments | Uses `clipboard.startImageUpload`, `clipboard.appendImageUploadChunk`, `clipboard.commitImageUpload`, and the `clipboard.saveImageAsTempFile` fallback. Their validator is `isValidBase64` in `clipboard-params.ts`, not `isValidRuntimeFileBase64`. Unchanged, including the mobile normalizer's existing permissive padding behavior. | -| CLI | No producer of either Base64 file-write method. File commands in `src/cli/handlers/file.ts` call `files.open` / `files.openDiff`; other CLI RPC call sites do not construct Base64 file writes. | -| Generated params catalog | References both schemas; `RpcParams` consumers use inferred types. Mobile's entry point is `export type` only, so no new client-side parsing occurs. | -| RPC dispatch | `files-mutation-methods.ts` registers both schemas. The chunk schema extends the whole-file schema; these are the only runtime consumers of the predicate. Direct runtime/provider calls do not parse these schemas. | - -Repository-wide searches covered method names, schema names, the predicate and its -pattern, and all callers of the upload/staging functions. Targeted history search -on `HEAD` under `mobile/src` found no introduction/removal of the affected methods -or predicate. The local release refs `mobile-ios-v0.0.27` and `mobile-v0.0.13` also -contain no callers of either Base64 file-write method; the iOS ref uses the separate -clipboard and terminal-artifact methods above. No shipped mobile producer of a -newly rejected value was found in this source/history audit. - -## Remote and workspace compatibility - -Old desktop clients using the audited producer send valid quartets to a new host. -A new client still sends the same bytes to an old host. No method, field, opcode, -or host-published content changes. This follows the mixed-version requirements in -[remote-wire-compatibility.md](./remote-wire-compatibility.md). - -The RPC validation runs before workspace resolution and provider selection, so the -same rule applies to folder workspaces, git worktrees, local hosts, and SSH hosts. -SSH ownership fences and provider writes are unchanged. Arbitrary external RPC -callers sending malformed padding will now receive a validation error; valid -padded and unpadded payloads remain accepted. - -## Regression evidence - -`src/main/runtime/rpc/methods/files-base64-padding.test.ts` exercises both actual RPC -registrations, asserts rejected input never reaches the writer, and verifies -accepted content is forwarded unchanged. With the original predicate, the test -run produced **16 failed / 16 passed**; all 16 failures were newly rejected padding -shapes accepted by the old implementation. This was run before editing the predicate. - -The existing desktop external-import test now uses a final `AA==` chunk after a -524,288-character first chunk, pinning padded final-chunk forwarding in the real -upload path. No producer changes or clipboard validation changes were necessary. - -## Validation results - -All test/typecheck commands used `ORCA_BACKGROUND_LAUNCH=1`. - -- `pnpm tc`: exit 0; all root typecheck projects passed. -- `pnpm exec vitest run src/main/runtime/rpc`: 277 files passed, one failed; - 2,419 tests passed, two timed out, one skipped. Both timeouts were in the unchanged - `terminal-output-frame-chunks-equivalence.test.ts` (5s surrogate-range test and - 30s 800-payload fuzz test). -- `pnpm --dir mobile typecheck`: exit 0 (`tsc --noEmit`). -- `pnpm run check:code-quality:changed`: exit 0; code quality, type-aware code - quality, and React Doctor each reported zero new findings across three source files. -- Focused run with `--config config/vitest.config.ts --maxWorkers=2`: all three - files / 58 tests passed, covering padding, desktop external imports, and the - terminal-output equivalence file that timed out in the initial run. -- Full RPC rerun with `--maxWorkers=2`: exit 0; all 278 files passed, - 2,421 tests passed and one skipped (198.34s). No timeout overrides were needed. diff --git a/docs/reference/sharing-agent-skills.md b/docs/reference/sharing-agent-skills.md deleted file mode 100644 index 9b9b53706a3..00000000000 --- a/docs/reference/sharing-agent-skills.md +++ /dev/null @@ -1,88 +0,0 @@ -# Share and install agent skills - -Orca can put one skill or a bundle of skills behind one unlisted, revocable link. Shared bundles -do not appear in search, a catalog, or a public index. Anyone who has an active link can inspect -and install its contents without signing in, so treat the link like a credential. - -## Share skills - -Publishing and link management require an Orca account in the desktop app. - -1. Open **Skills** and choose **Share skills**. -2. Select one or more skills. One link can contain a large collection, such as 30 skills. -3. Review the bundle name, included skills and files, scripts, executable files, digest, account, - and optional release notes. -4. Choose **Publish skill**, **Publish bundle**, or **Publish new version**, then copy the link. - -Orca publishes an immutable version. Later changes do not silently alter a link's current bytes; -publish a new version to update the Cloud package. - -Use **Settings → Share Skills** to copy or revoke active links. Revocation blocks new previews and -download grants. A grant issued immediately before revocation can remain usable for up to five -minutes, and revocation does not remove copies that recipients already installed. - -## Install from a link - -Opening an Orca skill link shows a preview before changing any files. You can also open **Skills**, -choose **Install from link**, and paste the URL. - -1. Verify the author and organization. -2. Review the version, release notes, included skills, scripts, executable files, and digest. -3. Select all skills or only the ones you want. -4. Choose the destination machine and either global or workspace scope. -5. Review new, unchanged, updated, and conflicting skills, then choose **Install N skills**. - -Supported destinations include the local machine, paired Orca runtimes, WSL, and SSH hosts. The -destination runtime resolves its own home and workspace paths, so folder workspaces and remote -filesystems do not borrow paths from the client machine. - -Orca keeps one canonical installed copy and places it where supported agents can discover it. -Current provider coverage is documented in -[Agent skill provider paths](./agent-skill-provider-paths.md). - -## Conflicts, updates, and rollback - -**Keep local** is the default when an existing skill differs. Orca replaces modified content only -after you explicitly choose to discard it. - -Open **Skills → Manage installs** to inspect managed skills and their immutable version history. -Installing the latest version performs an update; selecting an older retained version performs a -rollback. Both use the same protected install transaction. If a bundle changes between versions, -Orca updates only the selected skills that still exist in that version. - -An interrupted install is recovered on restart. If Orca reports a conflict or partial result, -review the named skill and retry; completed skills do not need to be installed again. - -## Remove an installed skill - -Use **Skills → Manage installs → Remove**. Orca removes only copies and provider placements that it -owns and can verify. Modified or unowned files are preserved and reported. Discarding modified -content requires a separate explicit confirmation. - -Removing a local install does not revoke its share or delete its Cloud package. Likewise, -revoking or deleting Cloud data does not reach into recipients' machines. - -## Retention and deletion - -- Upload grants expire after 15 minutes. -- Abandoned upload bytes are removed from quarantine after one day. -- Published versions have no automatic age-based deletion. -- Deleting a package revokes its links before unreferenced objects are deleted. -- Deleted GCS objects remain operator-recoverable through a seven-day soft-delete window. -- Installed copies remain until someone removes them on each destination machine. - -Organization legal or retention requirements can override normal rollback and deletion timing. - -## Trust and privacy - -A skill is code from its author. `SKILL.md` can change agent behavior, and included scripts or -executables may run later when a person or agent uses the skill. Orca validates the package and -never executes its contents during installation, but you should install only from people you trust -and review unexpected scripts or executable files. - -Orca records bounded operational identifiers and outcomes. Normal logs, telemetry, and support -bundles exclude skill contents, filenames, manifests, local paths, share URLs, upload policies, -download grants, credentials, and access lists. - -If a link no longer works, ask its owner for an active link. Missing, expired, revoked, and deleted -links intentionally show the same response so Orca does not disclose private package existence. diff --git a/docs/reference/spinner-rendering-performance.md b/docs/reference/spinner-rendering-performance.md deleted file mode 100644 index 4340027474b..00000000000 --- a/docs/reference/spinner-rendering-performance.md +++ /dev/null @@ -1,203 +0,0 @@ -# Spinner rendering performance - -## ELI5 - -Imagine a wheel that tells the front desk every time it completes a lap. The -front desk is also handling your typing. CSS already turns the wheel for us, -but React still receives its once-per-second lap notifications. - -We put a day's worth of laps into one animation. The wheel moves at the same -speed, while sending one lap notification a day. Drawing visible wheels still -costs something. This removes recurring bookkeeping from the input thread; it -does not make rendering or the rest of Orca free. - -## How this builds on earlier changes - -| Change | What it achieved | Remaining cost | -| ------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | -| [#9380](https://github.com/stablyai/orca/pull/9380): shared JavaScript clock | Reduced frame-pipeline CPU in the original one-agent measurement | Wrote each spinner's style 12 times per second on the input thread | -| [#12359](https://github.com/stablyai/orca/pull/12359): compositor CSS rotation | Removed those recurring JavaScript style writes; fixed the reported typing regression | React still receives CSS iteration events | -| [#13987](https://github.com/stablyai/orca/pull/13987): synchronize on animationstart | Avoided a synchronous style query at every mount | Steady-state animation overhead stayed the same | -| This change | Preserves both later fixes and removes almost all iteration boundaries | Compositing, other app work, mount/reveal work, and a daily iteration boundary remain | - -The historical measurements in #12359 reported 41 rings causing about 490 style -writes per second, with typing input-delay p90 of 363 ms versus 19 ms when those -writes stopped. Those are historical production measurements, not numbers from -this benchmark or a direct comparison with today's app. - -## Implementation - -The production change is entirely in CSS. `AgentWorkingSpinner`, its callers, -markup, border, animation-start handler, and reduced-motion behavior stay the -same. No DOM node, pseudo-element, containment boundary, timer, observer, or -JavaScript animation loop is added. - -The transform travels 86,400 turns in 86,400 seconds with 1,036,800 steps: exactly -one revolution and 12 steps per second. `animationstart` sets `startTime = 0` as -before, preserving shared phase after mount and animation restart. The step -count is a timing-function parameter, not a million-entry keyframe list. - -React installs delegated `animationiteration` listeners even when the component -has no iteration handler. A native 2.2-second trace of 200 isolated rings counted -400 iteration events and 800 JavaScript calls before the change, versus zero of -either with the long cycle. That trace installed no animation-event listener. -These are event dispatches, not component rerenders or 400 separate OS wakeups. - -## Full-app benchmark - -The opt-in Playwright benchmark launches a fresh, hidden Orca app for each -scenario. It creates real Git workspaces and seeds working statuses through the -existing renderer fixture, including in-process subagent data. It renders the -normal sidebar, virtualizer, lineage, agent rows, tabs, and terminal. - -| Scenario | Git workspaces | Root agents | Subagents | Mounted / visible rings | Layout | -| ------------- | -------------: | ----------: | --------: | ----------------------: | -------------------------------------------- | -| `one-agent` | 1 | 1 | 0 | 3 / 3 | One working agent | -| `one-family` | 1 | 2 | 4 | 8 / 8 | All family rows expanded | -| `200-flat` | 200 | 400 | 800 | 162 / 15 | Normal virtualization; 23 workspaces mounted | -| `200-lineage` | 200 | 400 | 800 | 1,401 / 15 | Expanded lineage; all 200 workspaces mounted | - -Measurement-only styles switch between the original one-second cycle and the -new long cycle on the same elements. The real React root, callers, status data, -and app stay the same. The reported run alternates A/B and B/A, with four -ten-second CPU samples per variant after warmup. CPU samples use cumulative -Electron process CPU and CDP main-thread task/script/style/layout metrics. No -renderer polling, screenshots, or benchmark iteration listeners run during -those CPU windows. No samples are discarded. - -Typing is measured separately using the existing paced terminal-typing probe: -64 keys at 113 ms cadence, twice per variant, after two seconds of warmup with -status traffic. Status updates arrive in groups of up to eight every 200 ms. -Keys pass through the DOM, real PTY, and xterm. A sidecar timestamps arrival at -the PTY, and a bounded terminal-buffer scan observes each echo. Missing input -or echoes fail the benchmark. Echo measurements include the 10 ms scan interval; -they do not measure native display presentation. Native animation traces also -run separately from CPU and typing samples. - -The statuses are deterministic test data, not hundreds of paid model sessions. -The test exercises UI cost under agent-status traffic, not the compute or network -cost of model inference, SSH traffic, or hundreds of streaming PTYs. - -## Results - -CPU values are medians of four samples. "CPU ms/s" means milliseconds of -processor time used in one wall-clock second: 100 ms/s is about 10% of one CPU -core. Renderer + GPU-process CPU includes their other app work and CPU used by -the graphics process; it is not GPU hardware utilization or whole-machine CPU. -The main thread handles input and is included in renderer CPU, not extra work. -Echo p90 means 90% of sampled keys were observed within that time; ranges show -the two runs, not confidence intervals. No keys or echoes were missing. - -| Scenario | Renderer + GPU CPU ms/s, old → new | Main-thread ms/s, old → new | Echo p90 ms, old → new | -| ------------- | ---------------------------------: | --------------------------: | ---------------------- | -| `one-agent` | 37.0 → 38.2 | 5.4 → 3.5 | 19 → 18–19 | -| `one-family` | 46.2 → 44.6 | 7.8 → 4.4 | 17–19 → 18–19 | -| `200-flat` | 141.6 → 122.8 | 28.2 → 16.3 | 26–28 → 26–28 | -| `200-lineage` | 324.6 → 295.6 | 140.8 → 70.0 | 159–239 → 93–160 | - -The consistent gain is less main-thread work: about 35%, 43%, 42%, and 50% -less in these four scenarios. Native 2.2-second traces counted 6, 16, 324, and -2,802 iteration events before, and zero in each new variant, without adding an -iteration listener. That avoided work also exists in Orca itself, independently -of the isolated fixture and CPU noise. - -Total CPU was roughly unchanged in the one-worktree cases. In this run it fell -13% with normal virtualization and 9% with expanded lineage; seven of eight -paired large-case CPU samples favored the change. These percentages are not -universal: a shorter three-variant ablation measured flat-list CPU at 89.0 ms/s before and -108.0 ms/s with the long cycle, while main-thread time still fell from 26.3 to -17.4 ms/s. The repeatable main-thread reduction is stronger evidence than a -single total-CPU percentage. - -Typing was similar in the small and flat-list cases. Expanded-lineage echo p90 -improved in the final run, but a shorter ablation had similar before/after -latencies. No general typing speedup or statistical non-regression guarantee -is established by these short experiments. - -### All CPU samples - -Values are rounded to one decimal and listed by round, with no outliers removed. -The first new small-case samples were higher than their paired baselines; they -remain included. CPU and typing were sampled separately. - -| Scenario | Version | Renderer + GPU CPU ms/s | Main-thread ms/s | -| ------------- | ------- | -------------------------- | -------------------------- | -| `one-agent` | Old | 37.8, 36.2, 26.2, 39.6 | 6.5, 5.2, 4.8, 5.7 | -| `one-agent` | New | 53.3, 35.8, 37.5, 39.0 | 6.7, 2.8, 3.0, 4.1 | -| `one-family` | Old | 46.3, 46.1, 47.9, 44.6 | 7.8, 7.6, 9.8, 7.7 | -| `one-family` | New | 53.8, 45.4, 42.5, 43.8 | 6.8, 4.5, 3.1, 4.3 | -| `200-flat` | Old | 142.4, 140.8, 147.0, 136.5 | 28.3, 28.1, 32.2, 27.0 | -| `200-flat` | New | 122.9, 97.2, 122.7, 126.7 | 19.2, 10.7, 16.6, 16.0 | -| `200-lineage` | Old | 317.9, 385.2, 315.9, 331.2 | 134.4, 159.6, 133.9, 147.3 | -| `200-lineage` | New | 318.6, 256.9, 296.5, 294.7 | 89.2, 60.8, 71.6, 68.3 | - -## Reproduce - -```sh -ORCA_BACKGROUND_LAUNCH=1 pnpm bench:spinners --sample-ms=5000 -ORCA_BACKGROUND_LAUNCH=1 pnpm bench:spinners --verify-only --scale-factor=1 -ORCA_BACKGROUND_LAUNCH=1 pnpm bench:spinners --verify-only --scale-factor=2 -ORCA_BACKGROUND_LAUNCH=1 ORCA_SPINNER_BENCH=1 ORCA_SPINNER_KEYS=64 \ - pnpm test:e2e spinner-workspace-perf.spec.ts --workers=1 -``` - -The full-app command rebuilds in `e2e` mode. For a fresh build already made with -`pnpm exec electron-vite build --mode e2e`, `SKIP_BUILD=1` reuses it. Do not reuse -an old launch-policy build. `ORCA_SPINNER_SAMPLE_MS`, `ORCA_SPINNER_ROUNDS`, -`ORCA_SPINNER_KEYS`, `ORCA_SPINNER_KEY_CADENCE_MS`, `ORCA_SPINNER_VARIANTS`, and -`ORCA_SPINNER_OUTPUT` control the experiment. `ORCA_SPINNER_CPU=0` repeats only -typing; `--grep one-agent` selects one scenario. Reports, native traces, typing -sidecars, and CDP screenshots are written under `.bench-fixtures/`. Run one -benchmark at a time, without concurrent builds or tests. - -The optional `contained` variant retains the rejected offscreen experiment for -ablation. It adds `content-visibility:auto` to the existing wrapper through -measurement-only styles. It is not enabled in production or the default -benchmark comparison. - -## Visual and behavioral checks - -Both 1x and 2x display-density checks passed 720 ring comparisons each: 6/8 px -rings, light/dark themes, supported zoom extremes, all 12 phases, long elapsed -times, and the daily wrap. The comparison pauses each animation and sets its -`currentTime`, so the long-elapsed and daily-wrap cases exercise the deterministic -style path rather than a running compositor animation. Against that path the -tolerance is one channel level for floating-point antialias rounding. A running -animation at multi-hour ages can differ by a few channels on the ring edge — a -fraction-of-a-pixel antialias difference at large accumulated angles, not a phase -or shape change. Checks also cover shared phase, reduced motion, initial offscreen -reveal, repeated scroll-away/reveal, and `display:none` restoration. - -## Limits and rejected approaches - -Adding `content-visibility:auto` to the existing stationary wrapper saved more -CPU at large mounted counts, but a 1x display check found a one-pixel shift at -the minimum UI zoom. That containment change is excluded. A previous -pseudo-element version also regressed typing latency in the virtualized list. -Neither prototype's CPU or typing numbers describe the final patch. - -An initial typing run used a 100 ms key cadence, which can repeatedly align with -200 ms status bursts. Follow-up runs use 113 ms, more keys, and two seconds of -warmup under status traffic. This reduces timing bias; it does not excuse a -regression. CPU measurements run separately and do not depend on key cadence. - -An early isolated test suggested a 31% process-CPU reduction that a longer audit -did not reproduce. The longer isolated audit measured original 104.04 versus -long-cycle 92.32 CPU ms/s, and main-thread 10.08 versus 0.24 ms/s. A fixture with -every ring far offscreen and containment enabled could also approach idle; that -is not representative of Orca with visible animations. Neither result justifies -claiming "free spinners" or a universal CPU percentage. Virtualized, unmounted -rows already cost nothing, and this patch does not add offscreen culling. - -All local measurements use an Apple M4 (10 cores), macOS, Electron 43.4.1 / -Chromium 150.0.7871.224. Native windows stay hidden and unfocused; -benchmark-only settings disable background throttling to exercise the frame -pipeline. These are not visible-window power measurements. No battery benefit -is established. Linux/Windows need their own runtime measurements. The -renderer-only change does not alter SSH execution, wire data, status semantics, -Git operations, or folder-workspace ownership. - -Animated PNGs, masks, layer promotion, CSS sprites, individual `rotate`, and -containment on the rotating element were also explored. Shared images added -raster work and regressed the single-ring case; sprites reintroduced per-frame -style work. They did not meet the appearance and responsiveness requirements. diff --git a/docs/reference/ssh-execution-boundary.md b/docs/reference/ssh-execution-boundary.md deleted file mode 100644 index 62ecb83838d..00000000000 --- a/docs/reference/ssh-execution-boundary.md +++ /dev/null @@ -1,129 +0,0 @@ -# SSH Execution Boundary - -How Orca splits work between your machine and an SSH host, what survives a disconnect, and how to keep `unverifiable` distinct from `exited`. Nothing under `docs/` stated this before; agents and humans were inferring it from error strings and getting it wrong. - -## The rule - -**The execution host owns everything that touches execution** — tools, credentials, identity, environment, processes, and artifacts. The client owns the UI, transport, and Orca control-plane state, but is not authoritative for execution state. - -Two consequences, both non-negotiable: - -1. **No silent substitution.** An operation on a remote `repoPath` must never fall back to running on the client. A missing SSH provider is not permission to answer locally — a local run can answer for the _wrong repository_. -2. **No asserting what you cannot observe.** Loss of contact is not evidence of `exited`. Report `unverifiable`, never `exited`. - -The vocabulary is fixed: **`live` / `unverifiable` / `exited`**, taken from the incumbent `UnstoppedPtyVerdict`. Do not introduce synonyms, and never collapse `unverifiable` into either neighbour. `exited` requires positive evidence of absence from the host that owns the process; a transport failure can only ever produce `unverifiable`. - -Rule 1 is stated at `src/main/source-control/repo-default-branch.ts:76-78`, `src/main/repo-worktrees.ts:45-48`, `OrcaRuntimeService.probeWorktreeDrift` in `src/main/runtime/orca-runtime.ts`, and `src/renderer/src/lib/connection-context.ts:22-24`. It is enforced throughout `src/main/runtime/orca-runtime-git.ts` by `requireRuntimeGitProvider` in `src/main/runtime/runtime-git-command-target.ts`, and throughout the runtime filesystem commands by `requireRuntimeFileProvider` in `src/main/runtime/runtime-file-command-target.ts`. Both route on the target's resolved `executionHostId` rather than on a repo row's `connectionId`: they throw the provider-unavailable message when an SSH host has no registered provider, throw `ExecutionHostNotDispatchableError` for a `runtime:` host this process does not execute, and return `null` only for `local`. Grep those names for the current call sites rather than trusting a count. - -`src/main/runtime/unstopped-pty-verification.ts:12-16` is the reference implementation of rule 2: it keeps `live` / `unverifiable` / `exited` as three distinct verdicts, and treats "we could not ask" as its own answer. - -## What runs where - -| Concern | Executes on | Notes | -| -------------------------------------------------------------- | ------------------ | -------------------------------------------------------------------- | -| PTYs, agent CLIs | **remote** | children of the detached relay daemon, not of the ssh channel | -| git (status, diff, log, fetch, push, commit, branch, worktree) | **remote** | via `src/relay/git-handler.ts` | -| filesystem, watching, search | **remote** | | -| repo setup hooks (`--setup`) | **remote** | identical policy to local | -| commit-message / PR-field AI generation | **remote** | uses the remote agent CLI and its auth | -| `gh` / GitHub API, `glab` / GitLab | **client** | inconsistent with the rule; PRs carry the client's identity | -| the `orca` CLI inside a remote terminal | **client runtime** | control plane only — your files and processes stay remote; see below | - -## Survival: what a disconnect does _not_ do - -By default, remote work survives your machine going away. The relay is a detached daemon (`nohup … .sock`, from `getDaemonSocketPath` in `src/main/daemon/daemon-spawner.ts`), every earlier protocol version stays attachable (`PROTOCOL_VERSION` in `src/main/daemon/daemon-protocol-version.ts`), and a daemon holding live sessions is preserved across a version change instead of replaced (`shouldPreserveDaemonWithLiveSessions` in `src/main/daemon/daemon-replacement-preflight.ts`). - -## Control plane - -On an SSH host, `orca` is a shim (`~/.orca-relay/bin/orca`) that proxies **back to the client's runtime** over the relay socket. Your repository, processes, and files remain remote — only the control plane is on the client. This is correct for an SSH target, but it has a consequence worth stating plainly: - -> When the client disconnects, every `orca …` command run on the SSH host fails with `No owning Orca client is connected to the relay`. The PTY stays `live`; its control plane does not. - -Orchestration state (Runs, Tasks, Dispatches, mailboxes) is client-resident for the same reason. An agent on an SSH host should not depend on `orca` for anything it must finish while you are away. **Commit and push early** — unpushed work on a remote box is unavailable to the client until it reconnects. - -## Distinguishing `unverifiable` from `exited` - -A verdict needs evidence from the host that owns the process. Apply these tests in order. - -**Was the signal produced by the owning host, or by the client's own bookkeeping?** Absence from a client-side set, a lookup that threw, a socket that closed, a command that timed out — none of these observe the process. They are `unverifiable` by construction, whatever the field is named. - -**Did every remote PTY on that target go quiet at once?** A transport drop takes them all together. Simultaneous silence across a host indicates a lost link, not simultaneous death. - -**Does the termination event match the current identity?** A host-delivered exit for the live PTY incarnation and provider generation, while its siblings still report, establishes `exited`. A stale event, an event for a superseded incarnation, or one quiet terminal with no host evidence does not. - -**Did the answer carry its evidence, or only the same wording?** `pty.attach` refuses with `PTY "" not found` both for a pid the relay probed and found gone and for an id its session map never had — which is every id minted before a relay restart, since ids carry a per-start mint epoch. Only the probed refusal carries `PTY_ATTACH_PROVEN_EXITED_MARKER` (`src/shared/pty-attach-absence-evidence.ts`) and reaches the client as `SshPtyProvenExitedOnRelayError`; the unmarked union arrives as `SshPtyAbsentFromRelayError`, which licenses retiring the client's own route to the PTY and nothing more. A missing marker is never evidence — an older relay omits it too. - -**Is a returned status actually a claim of success?** An operation that reports failure may have succeeded, and one that reports success may not have run — check the durable state it should have changed rather than trusting the return. - -Anything short of positive host evidence is `unverifiable`. Reporting it as `exited` is the error this document exists to prevent: it orphans live work and can cold-start a duplicate over the same worktree. - -## Host contact is a different question from process liveness - -The `live` / `unverifiable` / `exited` triple above answers one question: is this PTY running. It has -no synonyms, and nothing below adds any. - -A second, narrower question — can we currently reach the host at all, and what is its last answer -worth — is answered by `RuntimeHostContact` (`src/shared/runtime-host-contact.ts`), whose arms are -`live` / `unverifiable` / `refused` / `retired`. These are **not** extra process verdicts and must -never be mapped onto one: - -- `refused` is the host answering and turning us away — unauthorized, a protocol mismatch, a status - method it does not implement. That is positive evidence about the _connection_, and it says - nothing whatever about whether the host's PTYs are running. They almost certainly still are. -- `retired` is the pairing being ended by explicit user action. Same point: the client stops having - a route, the remote work is unaffected. - -Both are reasons to stop _trusting a cached answer_, never reasons to report a process `exited`. A -reader that needs a process verdict must still get it from the host that owns the process, by the -tests above. - -Why the extra arms exist at all: the renderer previously expressed every non-answer as one nullable -`status`, so a probe in flight, a probe that failed, a host that refused us and a retired pairing -all reached readers as the same `null` — and readers spent that `null` on decisions of very -different weight, including destructive ones. Folding `refused` and `retired` back into -`unverifiable` to match this document's triple would recreate exactly that collapse. The vocabularies -are deliberately separate because the questions are. - -## Deciding a remote pane is idle - -The orphan-PTY sweep is the one flow that turns an observation into a SIGKILL, so its idleness evidence has to be measured against the same thing the signal reaches. It is not the terminal. - -`forceKillPosixPtyProcessGroups` (`src/main/pty/posix-pty-process-groups.ts`) collects every process group on the pane's tty and `killpg`s each one. The blast radius is therefore _(process groups on the tty) × (members of those groups, wherever they are)_, and the second factor is not bounded by the terminal at all. Two facts make that gap reachable: - -- **Job control can be off.** With `set +m` a background job does not get its own process group — it keeps the shell's. `ps` then shows one process group on the tty, running a build. Nothing in a tty-shaped predicate can see it. -- **A group member can leave the terminal.** `ioctl(TIOCNOTTY)` without `setsid` drops the controlling terminal but keeps the pgid, so the process reports `tpgid == -1`, never appears in `ps -t `, and is still killed by `killpg(shellPgid)`. A double-forked grandchild similarly keeps the pgid while reparenting to pid 1, so no walk by `ppid` from the PTY root can name it either. - -So `shellOwnsEveryTtyProcessGroup` (`src/main/providers/agent-foreground-process-batch.ts`) requires both measurements: every process group on the tty is the shell's own with none stopped, **and** the shell's own process group has no other member anywhere in the host's process table. The name is tty-shaped for wire-compatibility reasons only. - -Two residuals remain, and neither is removable here. The capture is a snapshot, so work started between the `ps` and the signal is invisible — bounded by `RELAY_PTY_SWEEP_MAX_EVIDENCE_AGE_MS` on the reading side, not eliminated. And a process the host's own `ps` cannot enumerate (another PID namespace, `hidepid=2`, a table truncated by a permission boundary) is unobservable while `killpg` still reaches it. - -The general rule this instantiates: **evidence must be measured in the unit the destructive action operates on.** Evidence in a different unit is `unverifiable` no matter how precise it looks. - -## Reading artifacts instead of process state - -Artifacts are stronger evidence than liveness signals, but they answer a narrower question than they appear to. - -A matching commit from `git ls-remote --heads origin ` or a PR head lookup proves **that commit reached the remote** — not that the current run pushed it, and not that the latest work was included. An absent result proves nothing was found, not that nothing was pushed: the ref may have been deleted, the PR closed, or the query may simply have failed. - -A listing is only evidence about the hosts it actually covered. When a result does not name its scope, an empty answer is not evidence that nothing is running elsewhere. A clean **local** worktree says nothing at all about the remote one. - -## One host, one model - -An SSH host and a paired runtime (`orca environment`) imply opposite boundaries: the first is a dumb execution host driven by your client, the second is a peer that owns its own control plane. Registering the same machine both ways splits its worktrees across two identities, makes `terminal list` return different sets depending on `--environment`, and reliably confuses both humans and agents. Pick one per machine. - -For work that must continue while you are offline, use the peer/headless-runtime model on the remote host instead of the direct-SSH model. Its control plane is host-local, and its daemon-backed PTYs can stay `live` across a PID-scoped runtime restart so the runtime can reattach. A service manager that reaps the runtime's cgroup, or an explicit daemon shutdown, makes them `exited`; see [Running orcad](./orcad-operations.md#process-scoped-and-cgroup-wide-stops). Do not register the same machine through both models. A detached agent process outside Orca can also survive a control-plane outage, but it has no stdin, so its instructions cannot be amended mid-run. diff --git a/docs/reference/ssh-host-key-verification.md b/docs/reference/ssh-host-key-verification.md deleted file mode 100644 index 73955bdc7df..00000000000 --- a/docs/reference/ssh-host-key-verification.md +++ /dev/null @@ -1,424 +0,0 @@ -# SSH host key verification (STA-4319) - -Revised after security and migration review. Where a first draft was wrong, the correction is kept -visible rather than quietly edited out — the reasoning matters for anyone changing this later. - -## The defect - -`src/main/ssh/ssh-connection.ts:1184` installs a `hostVerifier` that records a SHA-256 fingerprint -and then `return true`. Every ssh2 connection accepts every host key. There is no `known_hosts` -consult, no trust record, and no change detection anywhere in `src/main/ssh/`. There is exactly one -ssh2 `Client` construction site, so the fix has a single chokepoint. - -Scope is per-connection, not per-feature: one `SshConnection` per target serves exec, SFTP, port -forwarding, the filesystem watcher and relay deploy. - -### Threat model, corrected - -Traffic is still encrypted, so a passive observer gets nothing. The exposure is an **active** -attacker who can redirect the connection — ARP/DNS spoofing, hostile Wi-Fi, a hijacked internal name. - -Three corrections to the first draft: - -- **Jump hosts are NOT the worst case; they are already safe.** `shouldUseSystemSshTransport` - (`ssh-transport-selection.ts:71-91`) returns true for exactly the conditions under which - `resolveEffectiveProxy` (`ssh-proxy-command.ts:17-38`) returns a proxy — the two branch on the same - inputs in the same order — and `attemptConnect` returns unconditionally after the system probe - (`ssh-connection.ts:670-673`). So ProxyJump/ProxyCommand go through OpenSSH and are already - verified. The ssh2 proxy-spawn at `:697` is effectively unreachable. Good news for migration, and - the first draft's motivating example was simply wrong. -- **Agent forwarding was overstated.** `agentForward` is gated on the user's `ForwardAgent yes` - (`ssh-connection-utils.ts:203-205`). `config.agent` is always set, but that is agent _auth_, whose - signatures bind the session id and cannot be replayed onward. The risk applies to users who opted - into `ForwardAgent`, not everyone. -- **Credential theft was understated, and the relay claim was backwards.** `isAgentFallbackError` - treats _any_ auth error as agent fallback (`ssh-connection-utils.ts:59-61`), so a MITM that rejects - publickey walks the user to the password prompt (`ssh-connection.ts:844`) and the private-key - **passphrase** prompt (`:834`), and `cachedPassword` is replayed without prompting on every - reconnect (`:709`). Meanwhile the relay upload matters less than assumed — the attacker already - owns their machine. The real client-side impact is the **return** direction: the attacker becomes - the host our workspace trusts, driving relay protocol frames, landing SFTP content in local - worktrees, and feeding agent-hook payloads in. - -## Decisions - -### D1. Read the user's `known_hosts`; write only to our own store - -Consult the user's real `known_hosts` as a trust source — most developers already have their hosts -there from `ssh` and `git`, which is the entire migration story. Do **not** write to it: that file is -shared with every other SSH tool on the machine, and appending brings line-endings, permissions, -concurrent writers, and a corruption blast radius well beyond us. - -Two consequences to own rather than discover: - -- **Revocation does not propagate into our store.** `ssh-keygen -R host` clears `known_hosts` but not - our record. That is survivable for the ordinary rotation, because a `known_hosts` MATCH is now - decided before our store's mismatch — running the remedy we print and reconnecting works, which it - did not when the store was consulted first. What it does not cure is a host we only ever knew - ourselves, never written to `known_hosts`: there is no `ssh-keygen -R` for that one, so the - rejection names the store file directly. The "forget" action (D5) replaces that with a button; its - helper is deliberately absent until then, since an exported API nothing can reach is unverified in - production. Mismatch messaging must keep naming which source disagreed. -- **`ssh -G` on the HOME-divergent `-F` path suppresses `/etc/ssh/ssh_config`** - (`ssh-g-config-resolution.ts:44-52`), hiding site-wide `StrictHostKeyChecking yes` and - `GlobalKnownHostsFile`. On that path we must fail **strict**, never laxer than `ssh` would. - - There is no `ssh`-only way out of this: `-F /dev/null` does NOT invert the exclusion, it reports - built-in defaults, so a probe built on it looks permissive on every machine. Verified against - OpenSSH 10.2p1. So the file is read directly, answering a deliberately weaker question — _could_ - the site config be restricting host keys — where anything ambiguous (unreadable, an unresolvable - `Include`, the directive present at all) keeps the refusal. Only a site config that demonstrably - says nothing about host keys clears it, which is what stops the rule punishing every devcontainer, - `su` shell and Nix shell. - -### D2. Ask `ssh -G`, do not reimplement config resolution - -`ssh -G` reports `userknownhostsfile`, `globalknownhostsfile`, `stricthostkeychecking`, -`checkhostip`, `hostkeyalgorithms`, `fingerprinthash`, `hashknownhosts`, `updatehostkeys` and -**`hostkeyalias`** — with `Match` and `Include` already applied. `resolveWithSshG` exists and simply -does not read them yet. - -`userknownhostsfile` is a space-separated list on one line, may contain `~`, and may contain -double-quoted paths with spaces. When `ssh -G` is unavailable (no `ssh`, non-zero exit, >5s timeout) -fall back to `~/.ssh/known_hosts` + `known_hosts2` — never to accept. - -`HostKeyAlias` must be honoured: users tunnelling bastions through `localhost:port` depend on it and -would otherwise hit spurious mismatches. It appears nowhere in `src/main/ssh/` today. - -**Lookup key.** Config resolution uses `configHost || label` (`ssh-connection.ts:660`) while ssh2 -dials `effectiveHost` (`ssh-connection-utils.ts:188`). The `known_hosts` lookup must use -`HostKeyAlias` if set, else the **resolved hostname** — keying on the Orca label would miss every -existing entry. - -**Two ordered lookup passes, not one candidate set.** Verified against OpenSSH 10.2p1: a non-default -port looks up `[host]:port` first, and if that finds nothing it retries the **bare** host. Crucially, -on that second pass a wrong key is downgraded to `unknown` rather than reported as changed. So the -passes are `[['[host]:port'], ['host']]`, and the fallback pass can only yield `match` or `unknown`. -Collapsing them into one set would give a spurious first-contact prompt to anyone who has a bare -line and connects on a non-default port; treating the fallback as authoritative would raise a false -change-of-key alarm. - -**The entry condition to that second pass is the part that bites.** ssh runs it only when the -port-qualified lookup matched no plain entry of ANY key type — not "no match". Gating it on -"no match and no same-type mismatch" reaches the bare line when an off-port entry of another type -exists, and returns `match` where ssh prints `IDENTIFICATION HAS CHANGED`: an accept-a-changed-key -path, reproduced live. And the observations from each pass must not leak into the other, or an entry -found only on the fallback refuses a host ssh accepts as first contact. - -**`HostKeyAlias` suppresses the port entirely.** ssh looks the alias up bare and never brackets it, -so an alias gets ONE pass regardless of port. Combined with the rule above, a stale `[alias]:port` -line would otherwise block the bare lookup ssh actually performs — turning the bastion case this -feature cites `HostKeyAlias` for into a hard failure. - -**Hashed entries hash the candidate form, not the bare host** — `[example.com]:2222` is what gets -HMAC'd for a bracketed entry, so each candidate must be hashed separately. - -**Multiple files union.** Any exact hit in any file wins; a disagreeing entry in another file does -not make it a mismatch. Confirmed live in both orderings. - -**A `@cert-authority` line whose key equals the presented plain host key is not a match** — a CA line -only validates certificates. A normal line alongside it still decides. But ssh's verdict for a -CA-covered host presenting a plain key is `HOST_NEW`, not a failure: it connects. See D4. - -### D3. Six outcomes, and type scoping is only safe with algorithm ordering - -`match | mismatch | revoked | ca-only | unknown-type-known-host | unknown`. - -Mismatch is scoped to the same key type: a host with only an RSA entry that presents ed25519 is not -"changed". Without scoping we would false-alarm nearly every RSA-era user on their first upgraded -connect, training them to dismiss the one warning that matters. - -> **Corrected against a live client.** The premise above is wrong about OpenSSH, though the -> conclusion survives. `check_key_in_hostkeys` is not type-scoped at all: ANY non-marker entry for -> the host that is not byte-equal produces `HOST_CHANGED`. Verified on 127.0.0.1:2223 — `known_hosts` -> holding only `ssh-rsa` against an ed25519-only server prints `IDENTIFICATION HAS CHANGED` and -> refuses. So ssh does not avoid the false alarm by scoping; it avoids the _situation_ via -> `order_hostkeyalgs`, and hard-fails when the situation arises anyway. Our split into `mismatch` -> and `unknown-type-known-host` therefore only chooses the wording — both refuse, which is ssh's -> action. What the ordering below buys us is what it buys ssh: the situation mostly never arises. - -**But scoping alone is a downgrade vector, and this is the correction that most changes the design.** -OpenSSH is safe here only because `order_hostkeyalgs()` reorders the client's proposed host-key -algorithms to put the types already in `known_hosts` first, and RFC 4253 gives the _client's_ order -priority — so a server cannot choose a type the client deprioritised. ssh2 negotiates ed25519 first -regardless. An attacker who cannot forge the RSA key on file simply presents ed25519 and receives a -friendly first-contact prompt instead of a hard failure. - -Therefore: **set ssh2's `algorithms.serverHostKey` to lead with the key types already known for that -host.** Type scoping without algorithm ordering is not a safe design. - -And when the presented type is unknown _while other types are known for this host_, that is -`unknown-type-known-host` — never a plain TOFU prompt. It must say we already hold a different key -for this host. - -### D4. Outcomes - -- **match** → connect silently. -- **unknown** → trust-on-first-use (see the phasing below for whether that is silent or prompted). -- **mismatch** → hard fail, no override in the failure surface. -- **revoked** → hard fail, always. -- **ca-only** → ~~hard fail~~ **REVERSED: treated as first contact.** See below. -- **unknown-type-known-host** → treat as suspicious, not first contact. - -`StrictHostKeyChecking` is honoured: `no`/`off` accepts unknown but **never persists** and still -hard-fails changed and revoked; `accept-new` persists silently; `yes` denies unknown. - -> **`ssh -G` does not report the spelling the user wrote.** StrictHostKeyChecking is rendered through -> `fmt_multistate_int`, which prints the first entry of `multistate_strict_hostkey`, and that table -> lists true/false before yes/no. So `yes` arrives as `true`, `no` and `off` both as `false`; only -> `ask` and `accept-new` pass through unchanged. Matching on `yes`/`no`/`off` matches nothing a real -> config can produce. `UpdateHostKeys` has the same shape (`true`, not `yes`). - -**ca-only, reversed after review.** The rejection was stricter than ssh, and the blast radius was -mispriced. An SSH CA user holds ONE line — very often `@cert-authority *` — which matches every -candidate, so EVERY target failed, not just CA-signed ones, including on-demand runtime VMs, and -`StrictHostKeyChecking=no` did not help. `ORCA_SSH_FORCE_SYSTEM_TRANSPORT=1` is read from the -process environment, which an Electron app launched from the Dock or Start Menu does not have, so -the documented escape was unreachable for exactly the people who needed it. And OpenSSH's own -verdict for a CA-covered host presenting a plain key is `HOST_NEW`: it connects. ssh2 cannot -validate certificates at all, so refusing conceded nothing ssh was not already conceding. - -The residual risk is accepted, not resolved: for a CA-protected host we take a plain key we cannot -tie to the CA. Certificate support is Phase 2 work. The `ca-only` outcome is still produced and -carried through the decision so the log shows a CA line was involved. - -**An unreadable known_hosts connects but records nothing.** A file that EXISTS and will not open is -the absence of evidence, and the common trigger is not exotic — a Windows OneDrive Known Folder Move -placeholder while offline fails with a cloud-file error, not ENOENT. Refusing there broke an -ordinary corporate laptop while blaming a config file that was fine, and was asymmetric with our own -store, which degrades to "nothing trusted" and connects. ssh warns and treats the host as unknown; -so do we — but we write no record, so a first contact we could not check never becomes durable -trust. An ABSENT file is not this case: that is the normal state for a fresh profile and genuinely -means nothing is known. - -### D5. Recovery must not live in the failure dialog - -A "forget this host key" button _in_ the mismatch dialog is D4's rejected "trust anyway" with one -extra click. Recovery lives in target settings: a separate, deliberate surface, no auto-retry, and it -shows the stored fingerprint so the user is choosing knowingly. - -Offer it only when **our** store is what disagreed; when `known_hosts` disagrees, forgetting our -record cannot unblock the connect. Messages, written to avoid naming internals: - -> **Ours disagreed** — "The host key for `build-01` changed since you last connected from Orca. If you -> rebuilt or reprovisioned this machine, this is expected." → _Forget the saved key_ / _Cancel_ - -> **`known_hosts` disagreed** — "The host key for `build-01` does not match the entry in -> `~/.ssh/known_hosts`. `ssh` and `git` will refuse this host too. Run `ssh-keygen -R build-01`." → -> no button, because a button would not help. - -### D6. Never prompt on a background reconnect - -A prompt only means something when a human initiated the connect. `userInitiated` does not exist on -the connect path today and must be threaded through `connect → attemptConnect → doSsh2Connect`, -defaulting **false**. - -Two traps: `useAutomationDispatchEvents.ts:203` and `pty-connection.ts:857` reach `ssh:connect` -without a human click — automation must pass `false`, but **terminal-pane focus reconnects must count -as user-initiated** or terminals die silently. And the denial string must avoid "authentication -failed"/"permission denied", or `isAgentFallbackError`/`isAuthError` -(`ssh-connection-utils.ts:46-61`) misclassifies it and the reconnect ladder retries a decision that -will never change. - -### D7. Fail closed — three known fail-open shapes - -1. The existing generation/disposed guard at `:1185` has the fail-open shape today: skip recording, - still `return true`. Post-fix that branch must **deny**. -2. A synchronous throw inside the verifier may not be caught by ssh2 — wrap and `verify(false)`. -3. Any non-`undefined` return accepts immediately (see Traps). - -Plus: no prompt channel registered → deny (the load-bearing default lives in `doSsh2Connect`, not in -IPC, so a caller that forgets to wire it cannot accidentally accept); no window → deny; timeout → -deny; dialog dismissed → deny. - -### D8. Store shape and scope - -Accepted keys are scoped to **host + port + key type**, not target id — aliases point at different -machines, two targets can name one machine, and a re-created target must not lose trust. - -The store is a **dedicated file**, not the main persistence blob (`persistence.ts:7088`): a settings -restore or rollback must not silently reset trust. Accept and mismatch events are logged. - -`hostKeyFingerprint` is now security-relevant _and_ wire-relevant — it is an isolation namespace sent -to the host (`ssh-relay-session.ts:1298`, `managed-hook-owner-identity.ts:187`). It is `undefined` on -the system transport, so **no trust logic may key off it**, and its format must not change (see -Traps). - -## Phasing — ship the defence before the dialog - -Review made the case that the riskiest part of this change is not the security model but the modal. -Startup restore fires eager connects for _all_ previously-active targets in parallel (`App.tsx:1041`) -with a 15s timeout, while a prompt would live 120s — N unknown hosts means N stacked dialogs -outliving the timeout that already deferred them. Runtime-owned ephemeral VMs -(`ephemeral-vm-runtime-ssh.ts:31`) dial a freshly provisioned host with a brand-new key on every -launch. Paired-web connects run on the _host desktop_ (`runtime/rpc/methods/ssh.ts:32`), so the -dialog would open on someone else's screen while the web user watches a spinner. - -**Phase 1 — no new modal.** Consult `known_hosts` + our store. `match` connects. `unknown` persists -silently with `accept-new` semantics and a passive notification naming the host and fingerprint. -`mismatch` (same type) and `revoked` hard-fail. This is the entire MITM defence with zero prompts, -zero startup storms and zero web hang. - -**Phase 2** — the TOFU dialog, `StrictHostKeyChecking` honouring, `ca-only`, `userInitiated` -plumbing, and the D5 settings surface. - -Carve-outs required before Phase 1 ships: - -- **Runtime-owned ephemeral targets are exempt from persistence** — a new key every launch is - expected, not suspicious, and recording one would accumulate a row per launch that eventually - reads as a spurious change. Implemented via `target.owner?.type === 'on-demand-runtime'`. -- **RPC-originated connects: NOT needed in Phase 1, required in Phase 2.** The review asked for - these to fail fast rather than leave a paired-web user watching a spinner for the 120s prompt - timeout. That hang is only reachable if a prompt exists, and Phase 1 has none — the decision - function is pinned by a test asserting it never returns `prompt`. An RPC connect therefore behaves - exactly like a local one: it accepts and records on first contact, or fails immediately with the - host-key reason. Adding a fail-fast path now would introduce a failure mode for a hang that cannot - occur. It becomes load-bearing the moment the dialog lands, and is listed in Phase 2. - - Worth noting for Phase 2: `runtime/rpc/methods/ssh.ts` already swallows the specific error and - rethrows `getPublicSshError(status)`, so a web client sees a generic failure rather than the - host-key reason. Pre-existing, but it means the Phase 2 message will not reach the web user - without a change there too. - -## Traps - -Each of these makes the fix silently do nothing. All confirmed in our tree. - -1. **An `async` verifier defeats it entirely.** ssh2 does - `const ret = hashCb(key, verify); if (ret !== undefined) verify(ret)`. An async function returns a - Promise — not `undefined`, and truthy — so ssh2 accepts before our callback settles. -2. **Do not set ssh2's `hostHash`.** It hands the callback a hex digest and discards the raw blob we - must compare — and it would change `hostKeyFingerprint`'s format, which is a cross-version state - break, not a local refactor. -3. **The existing test mock calls `hostVerifier(key)` with one argument** and ignores the return - (`ssh-connection.test.ts:86-91`). Under an async verifier every connect test there breaks. The - mock must change — flagged deliberately, not rewritten silently. -4. **Validate the blob**: embedded algorithm name must match the line's key-type field; reject empty - decodes, empty salts, and hashed entries whose hash is not 20 bytes. -5. **`ssh-relay-live-connect.test.ts:59`** constructs a connection with no credential callback — - headless with no prompt channel must deny, not hang. - -## Scope - -**In scope, corrected:** IPv6 literals and `[host]:port` bracket parsing. Review was right that this -is a _parser_ requirement, not a scope call — getting it wrong means hosts `ssh` knows come back -`unknown`, which is the prompt-training harm D3 exists to avoid. - -**Out of scope, with consequences stated:** - -- **`CheckHostIP`** — OpenSSH defaults it off; we form candidates from the hostname only. -- **WSL** — `src/main/ssh/` has no WSL awareness; a distro's `known_hosts` is unreachable, so WSL - users get first-contact treatment for hosts they already verified. -- **`UpdateHostKeys`** — we read it and use nothing, so we never learn a rotated key, which makes D5 - the routine path for key rotation rather than an exception. -- **Moving SFTP to the system transport** — correct direction, separate change. - -## Test plan - -**Parser** (against the file format, not our code's shape): plain lines, `host,host2` lists, -`[host]:port` used only when port ≠ 22, IPv6 literals, hashed `|1|salt|hash` with a real computable -vector, `@revoked`, `@cert-authority`, `*`/`?` globs, `!` negation vetoing a whole line, unrecognised -`@marker` skipping the line, malformed lines skipped not fatal, multiple keys per host, CRLF, blank -lines, comments, user file and global file disagreeing. - -**Decision function**: all six outcomes; type scoping; revocation resolved before match regardless of -line order; every `StrictHostKeyChecking` value; `no`/`off` never persists. - -**Algorithm ordering**: `algorithms.serverHostKey` leads with types on file — the test that makes D3 -safe rather than merely scoped. - -**Wiring**: unknown persists (Phase 1) without a prompt; match never notifies; mismatch fails with no -accept path; revoked fails; background reconnect denies; aborted connect settles pending verify -false; no prompt channel denies; runtime-owned targets are exempt; the denial string does not match -`isAuthError`; and — catching the worst regression — **the verifier returns nothing**, so a refactor -to `async` reddens a test rather than reaching a user. - -**Checked against a live client, not just the file format.** Two assumptions the design leans on were -verified by running an OpenSSH 10.2p1 client against a real `sshd` on `127.0.0.1:2222` and recording -its verdict: - -- **The bare-host fallback pass never reports a change.** With `StrictHostKeyChecking=accept-new`, a - bare line holding a _different_ key, dialed on a non-default port, made ssh connect and append a - new `[127.0.0.1]:2222` line — first contact, no `IDENTIFICATION HAS CHANGED`. Reporting `mismatch` - on that pass would refuse hosts ssh connects to happily, and would have looked like the cautious - choice. -- **`unknown-type-known-host` is ssh's own behaviour.** known_hosts holding `ssh-rsa` while the - server offers ed25519 makes ssh print `IDENTIFICATION HAS CHANGED` and refuse. So the rejection is - neither stricter nor laxer than ssh — and treating it as first contact, which a naive type-scoped - lookup does, is the laxer mistake. It also means `ssh-keygen -R` is the right remedy to name there. - -## What Phase 1 shipped, and what review changed - -The design above survived implementation. Every defect found afterwards was in the wiring, and the -pattern is worth recording because it repeats: **each one made us either blind or unusable, never -subtly wrong.** - -Fixed after review: - -1. **Our own store was type-downgradable.** The inline lookup filtered by key type first and could - only answer match/mismatch/unknown, so a record of a _different_ type read as `unknown`. D3's - downgrade, applied to the records we create ourselves. Stored types now also feed the algorithm - ordering — without that the guard is only half present. -2. **We keyed on the Orca label.** `ssh -G` echoes its own argument back as `hostname` when no Host - block matches, so for a manual target `resolved.hostname` _is_ the label — the one name D2 - forbids. We consulted no entries at all. -3. **A refused key still walked the credential ladder.** ssh2 reports a denial as a generic auth - failure, so we went on to prompt for the passphrase and hand it to the host we had just refused. - Rejections are now a typed error recognised before any fallback. -4. **Fail-closed nearly became fail-always.** "No readable `known_hosts`" counted a _missing_ file the - same as an unreadable one, so a profile that had never connected — everyone's first run — would - have been refused, and the suite passed only because dev machines have a `known_hosts`. -5. **Ephemeral runtimes were refused for a policy they cannot satisfy.** The carve-out sat below the - incomplete-sources check, so a HOME-divergent environment turned on-demand runtimes off entirely. - -A second review round, run against a live OpenSSH client and sshd rather than against the source, -found five more — and the pattern held: the two that mattered most were both cases where we refused -a host `ssh` connects to, and the worst single defect was that **`StrictHostKeyChecking` had never -been read correctly at all**, so a config saying `yes` was silently accepted AND persisted. See the -D2/D3/D4 corrections above. The lesson worth keeping: every one of these was invisible to unit tests -that fed the code the value a human writes, rather than the value the tool emits. - -## Action items (STA-4319) - -**Where the message actually lands.** Traced end to end, because a rejection the user cannot read is -a half-shipped feature. Fixed in this branch: the settings card clamped it to one line with no -tooltip, and the terminal reconnect overlay never asked for it at all. Still open: - -- **The "Remote Hosts" status bar shows only `Error`.** `SshTargetStatusRow` does not receive the - error, so the status bar is a dead end for the most likely place a user notices the failure. -- **"Connect again" is the wrong advice for a decision that will never change.** The terminal overlay - now prints the reason underneath, but its call to action still invites an action that cannot - succeed. Telling a permanent rejection from a transient fault in the renderer needs a typed reason - on the wire rather than a string — a remote-wire-compatibility decision, so deliberately deferred. -- **Toasts carry Electron's `Error invoking remote method 'ssh:connect':` prefix.** The repo has - strippers for exactly this; no SSH call site uses one. Also worth noting sonner auto-dismisses in - 4s, which is short for a message ending in a command the user is meant to copy. - -**Before Phase 2:** - -- **`UpdateHostKeys` (out of scope above, now the highest-value gap).** We read it and use nothing, - so a rotated key is a hard failure the user must resolve by hand. Combined with D5 this is the - routine path for key rotation, and it will be the most common way a legitimate user meets a - rejection. Decide whether Phase 2 honours it or D5's recovery surface absorbs it. -- **The web user never sees the reason.** `runtime/rpc/methods/ssh.ts` rethrows - `getPublicSshError(status)` on all three paths, and push events are redacted through - `getPublicSshState`, so a paired-web client always sees exactly `SSH connection unavailable`. - Pre-existing, but it makes the Phase 2 dialog message unreachable there without a change. Note the - redaction is not web-only: any target owned by a paired runtime environment is redacted, so a - _desktop_ user viewing a remote-Orca-server-owned host gets the same generic string. -- **RPC fail-fast** becomes load-bearing the moment the dialog exists (see Phasing). - -**Known gaps that Phase 1 accepts, listed so they are choices and not surprises:** - -- **WSL** — a distro's `known_hosts` is unreachable, so WSL users get first-contact treatment for - hosts they already verified through `ssh` inside the distro. -- **`CheckHostIP`** — candidates are formed from the hostname only. -- **Certificate validation** is still absent — a CA-covered host is now accepted on first contact - rather than refused (D4), so those users connect, but the CA itself verifies nothing for us. -- **`DEFAULT_SERVER_HOST_KEY_ALGORITHMS`** is a hand-copy of an ssh2 internal. A test pins it, so an - ssh2 upgrade that changes it fails CI rather than shipping — but the pin has to be honoured, not - deleted, because ssh2 throws `Unsupported algorithm` and every target stops connecting. - -**Rollout:** the first release carrying this is the first time Orca can refuse an SSH connection at -all. Worth a staged rollout or a kill switch: the failure modes we could not find are, by the shape -of the five above, far more likely to be "a legitimate host is refused" than "a bad key is accepted". diff --git a/docs/reference/ssh-reconnect-source-recovery.md b/docs/reference/ssh-reconnect-source-recovery.md deleted file mode 100644 index e8c12cf5189..00000000000 --- a/docs/reference/ssh-reconnect-source-recovery.md +++ /dev/null @@ -1,159 +0,0 @@ -# SSH reconnect: why the pane retry gets a byte tail, and what would actually change it - -Status: investigation result. The obvious follow-up to PR #14844 was traced and **rejected**, and -tracing it turned up the actual root cause: checkpointed source recovery has never run on an SSH -reconnect. Both are recorded here — the rejected shape so nobody re-proposes it, and the verified -cause with the fix it implies. - -## The shape of the problem - -A reconnect remounts the pane (`tab.generation` is its React key), so the xterm is disposed with its -buffer and something must repaint it. Today that is a **byte tail**: `reattachSshPtySession` sends -`requireReplay: true` and the relay returns `RecentPtyOutputBuffer.read()` — the last 100KB, read -non-destructively, with no notion of what this client already consumed. - -Two costs follow. Main's `@xterm/headless` model never sees those bytes (the tail bypasses -`onPtyData`), so it is stale by exactly the outage — which is what forces -`sshReconnectPaintsFromModel` to restrict the grid repaint to the alternate screen. And a shell loses -outage output past 100KB permanently. - -## The proposal that does not work - -"Make the pane-retry path request source recovery like `reattachKnownPtys` does." Mechanically this -is trivial — `sourceRecovery` is already an optional `pty.attach` param the relay parses, Path C -already calls the same `requestSshPtyAttach` helper and already parses the response field. The -required checkpoint state also survives a transport drop, in the module-level `recoveryByTarget` map -(`ssh-pty-consumer-recovery.ts:17`), reachable from `connectionId` because `connectionId === targetId`. - -It still fails, three ways: - -1. **The relay answers `'existing'` before it looks at the recovery argument.** - `relay-pty-source-publication.ts:99-109` short-circuits on a same-`clientId` attach, and a - reconnected client presents the same id — see the root-cause section below, where this turns out - to be the whole story rather than an obstacle specific to this proposal. -2. **A failed `reattachKnownPtys` deletes the checkpoint on purpose** (`ssh-relay-session.ts:3006-3007`) - and detaches the lease (`:3008`). The pane retry runs _after_ that, so it would present - `checkpointUnavailable`, which the relay converts to `restoreRequired` - (`relay-pty-source-publication.ts:124-130`) and the provider converts to - `SSH_SESSION_EXPIRED_ERROR` (`ssh-pty-provider.ts:103-107`). We would trade a blank-pane-with-tail - for a **killed session**. -3. **Wrong payload shape.** Recovery replays only the post-checkpoint delta - `(acceptedSourceEndSu → receivedEndSu]`. The byte tail is a screen snapshot for a _fresh, empty_ - xterm. Even a successful recovery returns roughly nothing in the common case, and the pane stays - blank. - -These two mechanisms answer different questions. Recovery keeps main's model whole; the tail repaints -a new terminal. Substituting one for the other is a category error. - -## A correction worth recording - -The motivating argument was "`requireReplay` is optional, so older relays ignore it and still show -blank panes." **That is wrong for the SSH relay.** The client deploys and launches its own relay -build into a version-scoped directory (`ssh-relay-deploy.ts:231`, `:594`), and `validateGrant` -rejects any grant whose `serverBuildId` differs from the expected one -(`ssh-pty-consumer-session.ts:58-65`, rationale in-code: _"client and relay ship in one build"_). -Client and SSH relay are version-locked; mixed versions do not occur on this channel. The -independent-update rule in `remote-wire-compatibility.md` still governs remote _runtime_ hosts — just -not this one. - -So there is no old-host population to rescue, and the urgency that argument created was false. - -## ANSWERED: checkpointed recovery never runs on an SSH reconnect - -The question above was "does the reconnecting client present a new `clientId`?" It does not, and the -consequence is that the whole checkpoint mechanism is dead on this path. Every link verified: - -1. **The client keeps its id.** A reconnect calls `Dispatcher.setWrite` - (`src/relay/dispatcher.ts:149-157`), which reuses `this.primaryClient` — including its `id` — and - replaces only the writer. The dispatcher refuses to detach the primary. This is already stated - in-repo at `src/relay/pty-handler.ts:1736-1742`. -2. **So `activate()` short-circuits.** `relay-pty-source-publication.ts:99` tests - `current?.clientId === context.clientId` and returns `'existing'` at `:108`. The `rotateDelivery` - recovery path at `:118-142` is reachable **only** when the ids differ — i.e. never, here. -3. **So the relay returns no `sourceRecovery`.** -4. **So the client abandons.** `finishSourceRecovery` (`ssh-relay-session.ts:2766-2785`) fails its - `!pendingRecovery` guard, calls `abandonPtySourceRecovery`, and returns false — which cancels the - delivery and deletes the checkpoint (`:3006-3008`). -5. **So the pane retry opens fresh and gets the byte tail**, via the `requireReplay` fix. - -The byte tail is therefore not a fallback. It is the only path that has ever run for an SSH -reconnect, and the flow-control/checkpoint machinery is inert on this path. - -That also explains the original blank-pane bug exactly: the relay concluded "this client already -holds the stream" because, by its own identity rule, it does. - -### The fix this implies - -Give a reconnected primary a distinguishable identity — a transport generation on the client record, -bumped in `setWrite` — and have `activate()` compare it alongside `clientId`, so a reconnect takes -`rotateDelivery` instead of `'existing'`. - -Why this is the tractable shape: - -- **No wire change.** `RequestContext`, `setWrite` and the publication are all relay-internal. -- **No compatibility exposure.** Client and relay ship in one build and are version-locked. -- **It does not disturb the invariant that broke three earlier attempts.** Deliveries still outlive - their clients; nothing retires on `onClientDetached`. The delivery is _rotated on re-attach_, - which is what the recovery design already intends and what its tests already cover. - -**UNVERIFIED and to be checked before implementing:** that `rotateDelivery`'s preconditions hold at -that moment (the checkpoint's `deliveryToken`, `clientGeneration`, `ownerGeneration` and -`ptyIncarnation` must match the live identity, `:124-128`); that `outputFlowControl` is granted on -the reconnected session; and what a rotation implies for the _renderer_, which still remounts with an -empty xterm and needs a screen, not a post-checkpoint delta. Recovery keeps main's model whole — it -does not by itself repaint a fresh terminal, so the tail may still be wanted for the pane even once -the model stops going stale. - -## Do not start at `onClientDetached` - -Three attempts failed there, each plausible until run: - -- Retiring the delivery on `dispatcher.onClientDetached` **breaks checkpoint recovery** (10 tests). - A delivery outliving its client is deliberate — it is what lets a client resume from a checkpoint. -- Retiring without `session.cancelDelivery()` orphans the credit ledger's one-upstream-owner slot; - the next open throws `PTY source delivery already has an upstream owner`. Seen live as a toast and - a blank pane. -- Comparing `record.identity.clientGeneration` to the request is impossible: that value is - client-supplied via `pty.openClient`, and `RequestContext` carries no generation of its own. - -## The lead that survives - -`reattachRejectedPty` (`ssh-relay-session.ts:1957-2004`) is an existing **single-PTY** entry point -into the `reattachKnownPtys` machinery, taking `(relayPtyId, mux, providerGeneration, mode)` and -driving recovery with `targetedDeliveryRecovery`. If per-pane recovery is wanted, that is the hook — -and it does not involve the pane-retry path at all. Unverified whether it is reachable at the moment -the renderer retries. - -## Preconditions, unchanged - -The SSH e2e lane must be green and triggering on **source** changes before any of this is attempted. -It was skipping for 15 specs; four regressions reached a user during that window. - -## Resolved: a disposed pane killed its successor's new shell - -A pane rebuilt during its first spawn uses the same reservation key, so main can return the -same PTY to both transports (#19386, #22578). The disposed transport must keep that shell -while its tab and layout leaf remain and its execution host's workspace is not being deleted. -A live transport refusing the id still retires it (#11003). This applies to local, WSL and SSH -IPC terminals; the remote-runtime transport has no corresponding kill. - -The remount trigger in the Scan-22 user report remains unknown. A retained shell can outlive -its tab if the tab closes before a successor binds it; keeping potentially owned work follows -the SSH execution boundary. - -## Open: the pane behind a preserved tab does not always rebind - -The merge now keeps a local tab the host has never been told about, so the tab and its title survive -a reconnect. The reattach behind it does not, reliably — measured at three runs in four against the -Docker-SSH lane. When it misses, the store holds the tab, the tab bar renders it, and the pane never -rebinds: the "frozen tab" shape the original report described, one layer down from the deletion that -used to cause it. - -Deliberately NOT asserted in `ssh-reconnect-tab-destruction.spec.ts`. A one-in-four flake in the lane -that exists to catch this class costs more than it proves — the lane stops being trusted, which is -exactly how the earlier silent-skip failure happened. Tab survival is asserted there and is -deterministic; the liveness gap is recorded here instead. - -Worth checking first, since it is the same shape as everything else in this file: the tab is absent -from the host snapshot, so whatever drives the per-tab reattach after an apply may simply not know to -reattach a tab the snapshot never mentioned. diff --git a/docs/reference/task-provider-identity-validation.md b/docs/reference/task-provider-identity-validation.md deleted file mode 100644 index 4dba0760aab..00000000000 --- a/docs/reference/task-provider-identity-validation.md +++ /dev/null @@ -1,81 +0,0 @@ -# Task provider identity RPC validation - -The automation RPC identity schema follows `src/shared/task-provider-identity.ts`, -re-exported by `src/shared/task-source-context.ts`. Only GitHub requires fields beyond -`provider`: `owner` and `repo` are strings, and `host` is optional. GitLab, Linear, -and Jira fields are optional nullable strings. Requiring a GitLab project or a -Linear workspace/Jira site would contradict the domain type and account-wide scopes. - -The discriminated union validates these existing field types without trimming, -coercing, or stripping identity fields. Unknown fields pass through as they did under -`z.custom`, including fields from newer clients. Absent and explicit-null identities -remain distinct. The schema does not infer providers from owner/repo or require a -git worktree, repository slug, or local execution host for a source context. - -## Producer census - -Paths below are relative to the repository root. Searches covered production -`providerIdentity`, `TaskProviderIdentity`, `sourceContext`, and -`linkedTaskSourceContext` uses across desktop, shared code, mobile, and CLI. - -| Producer or forwarding path | Populated verdict | -| ------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `src/shared/project-host-setup-projection.ts`: `getProjectProviderIdentity` | GitHub owner/repo populated together, or no identity. Supplies project identities consumed by desktop and migration. | -| `src/renderer/src/components/task-page-source-context.tsx`: `getTaskPageRepoSourceContext` | GitHub fields populated through projection, or null. GitLab uses explicit provider and `buildGitLabProviderIdentity`; projectId/project/webUrl populated, namespace can be null. | -| Same file: `buildGitLabProviderIdentity` | GitLab fields come from project path/host; missing path components become null. No required GitLab field is invented. | -| `src/renderer/src/hooks/composer-state/source-context-state.ts` | Derived GitHub context carries a complete project identity or null. Jira folder/project-group context explicitly has null identity. Draft/linked contexts are forwarded. | -| `src/renderer/src/components/use-task-page-source-availability.ts` | Linear workspaceId/name and Jira siteId/URL can be null for account-wide selections; team/project fields are not populated. Valid under the existing optional-field contract. | -| `src/renderer/src/components/task-page-jira-item-source-context.ts` | Bound issue context populates siteId, siteUrl, projectKey. | -| `src/renderer/src/components/new-workspace/use-jira-url-source.ts` | Bound URL issue context populates siteId, siteUrl, projectKey. | -| `src/renderer/src/components/worktree-jump-palette-create-worktree.ts` | Linear teamId/key populated; workspaceId/name may be null. | -| `src/shared/task-source-context.ts`: normalize/build functions | GitHub missing owner/repo becomes null identity; other providers' missing fields become null. Provider mismatch becomes null, never an inferred provider. | -| `src/cli/handlers/automation-handler-flags.ts`, `src/cli/handlers/automations.ts` | Explicit JSON source-context input is normalized before create/update. GitHub fields populated or null identity; other fields nullable. Omitted/null context preserved by flag handling. | -| `src/main/persistence/scheduling-automations/automation-context-migration.ts` | Builds source context from projected complete GitHub identity, or null context. | -| Desktop automation save/scoped-list/host clients and web transport | Forward existing source contexts, not new identity constructors. `automation-orca-save.ts` forwards the current automation context or null. Legacy arbitrary malformed RPC input is deliberately rejected by the new schema. | -| Mobile | No task-provider identity/source-context constructor or sender found. `mobile/src/components/new-workspace-project-targets.ts` uses project identity solely for display. | - -## Compatibility evidence and limits - -Read `docs/reference/remote-wire-compatibility.md` before changing validation. -A source search against release tag `v1.4.199` also finds no mobile -`sourceContext`/`linkedTaskSourceContext` sender; its sole `providerIdentity` use -is the display-only project target above. The released CLI flag reader also calls -`normalizeTaskSourceContext`. This is source-level evidence for the checked release, -not a claim to have executed every historical mobile binary. - -No shipped mobile producer with a newly rejected payload was found. No new required -field was added to the domain contract. GitLab, Linear, and Jira discriminant-only -identities remain valid. Folder-workspace null/absent identities remain valid on -both local and SSH hosts. No execution/status logic or client-side parsing changed. - -## Regression evidence - -`src/main/runtime/rpc/methods/task-provider-identity.test.ts` checks unchanged valid -identities for all four providers, required GitHub fields, every declared field's -type, optional/null non-GitHub fields, unknown-field preservation, explicit GitLab -discrimination with owner/repo present, local/SSH folder contexts, and update patches. - -The focused run passed 81 tests. Temporarily replacing GitHub's field validators -with optional `z.unknown()` validators (discriminant-only acceptance) caused 17 -failures and 64 passes. The mutation was restored before running the gates. - -Counts re-measured at `cf4f77f` after the blank-field commit added seven tests; -the earlier 74/16/58 figures described the commit before it. - -## Gate results - -All commands ran with `ORCA_BACKGROUND_LAUNCH=1`. - -- `pnpm tc`: exit 0; completed the repository typecheck runner. -- `pnpm exec vitest run src/main/runtime/rpc`: exit 1; 277 files passed, - one failed; 2,462 tests passed, one failed, one skipped. The only failure was - the unrelated `structured-agent-session-adoption-replay.test.ts` hitting its - 5,000 ms timeout. All identity tests passed. -- `pnpm --dir mobile typecheck`: exit 0; `tsc --noEmit` passed. -- `pnpm run check:code-quality:changed`: exit 0; zero new code-quality, - type-aware, or React Doctor findings across the two changed code files. -- Isolated retry of `structured-agent-session-adoption-replay.test.ts`: exit 0; - one test passed, with the test body completing in 358 ms. -- Full RPC retry with `pnpm exec vitest run src/main/runtime/rpc --maxWorkers=4`: - exit 0; all 278 files passed, 2,470 tests passed, one skipped (97.24 seconds). - The bounded-concurrency rerun resolved the timeout without changing test code. diff --git a/docs/reference/terminal-artifact-grant-integrity.md b/docs/reference/terminal-artifact-grant-integrity.md deleted file mode 100644 index ea04d35deef..00000000000 --- a/docs/reference/terminal-artifact-grant-integrity.md +++ /dev/null @@ -1,85 +0,0 @@ -# Terminal artifact grant integrity - -Clicking a path in terminal output mints a short-lived grant over one file in a -world-writable temp directory. Between minting the grant and using it, anything on the -host can replace that file. This page is the contract for the checks that catch such a -replacement — and, more importantly, for the windows they do **not** close. - -Read this before weakening a check in -`src/main/runtime/runtime-file-commands-terminal-artifact-access.ts`, before assuming -the sequence is atomic, or before extending the guarantee to remote hosts. - -## The stat identity is weaker than it looks - -A grant pins the file as `dev:ino:nlink:size:mtimeMs`. Four of those five fields -survive an unlink-and-recreate at the same path, measured on Linux (Node 24) by -replaying that exact sequence: - -| Field | After `rm` + recreate in the same directory | -| --------- | ------------------------------------------------------------------------ | -| `dev` | unchanged — same filesystem | -| `ino` | **reused 100% of the time** (3000/3000); ext4 hands back the freed inode | -| `nlink` | `1` before and after | -| `size` | unchanged whenever the replacement is the same length | -| `mtimeMs` | **quantised to 1 ms** — the kernel's coarse clock advances once a tick | - -So for a same-length replacement the whole identity string collapses to a 1 ms -timestamp race. The full string collided in 63.7% of back-to-back iterations, 19.8% -with a 0.5 ms gap, and 0% at ≥1 ms. The same probe on macOS collided 0 times in 2000 — -no inode reuse, nanosecond mtimes — which is why this only ever showed up on Linux CI. - -`ino` contributes no discriminating power against the exact case it is there to catch. -Do not add `ctimeMs` or `birthtimeMs` hoping to fix this: they come from the same -coarse clock and collide in the same window. `bigint: true` stats do not help either — -the precision loss is in the stored kernel timestamp, not in Node's `Number`. - -## The content digest sits alongside the identity, and cannot replace it - -Local grants also pin a sha256 of the artifact's bytes, read from the **same open -handle** as the stat so the two describe one inode with no gap between them. - -The identity string stays exactly as it is because it is a wire contract, not a -host-local detail: `src/relay/fs-handler-terminal-artifact.ts` recomputes that same -`dev:ino:nlink:size:mtimeMs` format field-for-field from its own stat and compares it -against the `expectedStatIdentity` the host sends. Changing the format would have a new -host publishing a string an older relay can never reproduce, failing every remote -artifact read as `terminal_file_grant_stale` — a break that reaches old peers with no -wire-schema change at all. See [remote wire compatibility](./remote-wire-compatibility.md). - -## What is still open - -The digest narrows these checks. It does not make any of them atomic. - -- **Read and preview — closed.** The bytes returned to the caller are the same - in-memory buffer that was digested, with no re-read in between, and the handle pins - the inode for the whole operation. A swap landing after the `open` leaves the handle - on the granted inode, so the granted content is what is served; a swap landing before - it fails the digest. -- **Write — a window survives.** Between the final pre-rename verification and the - `rename()` itself, the target path can still be swapped, and the rename clobbers - whatever is there. POSIX `rename()` has no "only if the target is still inode X" - form; Linux's `renameat2(RENAME_EXCHANGE)` would close it but is not portable and is - not exposed by Node. Narrowing this further means changing the commit strategy, not - adding another check before it. -- **Artifacts over 10 MB fall back to stat-only.** They digest to `null`, so only the - identity guards them. Nothing leaks: every read, preview, and write path rejects on - size before reading. The weakness is unreachable, not fixed — a later cap change - could expose it. -- **Remote and SSH grants are untouched.** They keep the stat-only check, with the full - 1 ms weakness on whatever filesystem the relay runs. Closing it needs a negotiated - capability so an older relay is never sent a digest it cannot verify. - -## What is proven, and what is inferred - -The filesystem numbers above are direct measurements. That identical stats plus changed -content previously returned the swapped bytes, and now do not, is pinned by -`orca-runtime-files-terminal-artifact-swap-detection.test.ts`, which replays the first -stat seen for a path so the collision is deterministic rather than a 1 ms coin flip. - -The join between the two is a chain, not a single observation. This defect surfaced as -an intermittent failure of `orca-runtime-files-terminal-artifact-io.test.ts` on -`rejects stale absolute terminal artifact previews before returning changed content`, -which swaps an 8-byte artifact for 8 different bytes. No one has instrumented a Linux -runner to prove that a specific CI failure was a same-tick mtime collision; the -conclusion rests on every ingredient being measured separately. Treat it accordingly if -a future failure does not fit. diff --git a/docs/reference/terminal-perf-latency-investigation.md b/docs/reference/terminal-perf-latency-investigation.md deleted file mode 100644 index 4e61ab3529c..00000000000 --- a/docs/reference/terminal-perf-latency-investigation.md +++ /dev/null @@ -1,106 +0,0 @@ -# Terminal latency investigation (2026-09-21) - -The September 21 scheduled report has five latency violations: three restores -above 1,000 ms, a worst key around 2,091 ms, and timer drift around 2,034 ms. -These remain failures under the [historically calibrated budgets](terminal-perf-report-budgets.md). -Passing the looser Electron assertions does not establish that the report passed. - -## Historical boundary - -Slow restores predate September: the July 20 scheduled run measured a 1,226.6 ms -Latin restore; August 1 measured 1,492 ms. Earlier June/July logs did not contain -usable summary rows and their artifacts expired. There is no established good/bad -application revision boundary, and no completed git bisect. A newly visible report -failure is not, by itself, evidence of a newly introduced application regression. - -The scheduled workflow uses one Playwright worker. Parallel Electron workers do -not explain these particular failures. Repeated macOS runs did not reproduce the -Linux stalls (15 targeted samples: restore 136–236 ms, hidden worst key ≤23.6 ms). - -## Controlled Linux experiments - -All comparisons use the same application build within their run, one worker, -real PTYs, and the original terminal workload. Diagnostic tracing can perturb -measurements, so it identifies the mechanism rather than setting new budgets. - -| Comparison | Evidence | Finding | -| --- | --- | --- | -| Default versus disabled background throttling | [35657300611](https://github.com/stablyai/orca/actions/runs/35657300611) | 14/20 restores exceed 1 second; four typing measurements approach 1 second. `setBackgroundThrottling(false)` does not remove the stalls. | -| Native browser trace | [35658595456](https://github.com/stablyai/orca/actions/runs/35658595456) | Renderer waits roughly 960–1,010 ms in `LayerTreeHost::WaitForCommitCompletion`. Some typing samples contain consecutive waits. | -| Current flags versus flags preceding `c64777d1bcd` versus SwiftShader | [35659601151](https://github.com/stablyai/orca/actions/runs/35659601151) | All three configurations still trigger undrawn-frame throttling. Reverting the May 26 flags is not a demonstrated fix. | - -In the graphics comparison, current flags had one of six restores above 1 second; -the earlier flags had three and a 1,036 ms worst key. SwiftShader had six measured -restores between 301 and 396 ms, but still contained one-second native waits after -the restore measurement ended, and one 202 ms timer-drift violation. Its faster -restore numbers are insufficient evidence of a fix. - -## Native mechanism - -The trace shows the renderer blocked inside: - -``` -ProxyMain::BeginMainFrame - Commit - ProxyMain::BeginMainFrame::commit - LayerTreeHost::WaitForCommitCompletion -``` - -During the gap, Viz repeatedly emits `SendBeginFrameDecision` with -`reason: ThrottleUndrawnFrames` and `should_send: false`. The graphics comparison -recorded 346 such decisions with current flags, 626 with the earlier flags, and -401 with SwiftShader. Raster work was already ready before the wait ended. - -The [matching Chromium source](https://github.com/chromium/chromium/blob/150.0.7871.250/components/viz/service/frame_sinks/compositor_frame_sink_support.cc) -limits begin frames to once per second when too many submitted frames remain -undrawn. Input can block the renderer waiting for the next compositor commit; -this is not a one-second xterm parse or proof of a second of CPU consumption. -The same throttle is present in Chromium 148.0.7778.218 (Electron 42.3.3) and -146.0.7680.177. This source comparison does not prove identical runtime behavior. - -## Visibility control - -[Run 35660847455](https://github.com/stablyai/orca/actions/runs/35660847455) -compared ten hidden-window samples with ten visible-window samples on the same -isolated Xvfb runner, using SwiftShader in both modes. Actual window visibility -was recorded. The background terminal panes remained hidden in both modes. - -| Measurement | Hidden window | Visible window | -| --- | ---: | ---: | -| Undrawn-frame throttle decisions | 1,251 | 0 | -| Largest worst-key latency | 3,062.8 ms | 30.5 ms | -| Largest timer drift | 3,111.8 ms | 67.0 ms | -| Restore range | 213.8–1,862.2 ms (9 completed) | 223.4–734.3 ms (10 completed) | -| Electron tests passed | 9/10 | 10/10 | - -All ten visible-window samples satisfy the existing latency limits. This isolates -the never-presented Linux test window as the trigger for the reproduced native -stalls. It does not establish a newly introduced application-code regression or -prove that every historical outlier had the same cause. - -## Full scale validation - -[Run 35662787327](https://github.com/stablyai/orca/actions/runs/35662787327) -passed all 21 scenarios and all 32 strict report rows, with zero skipped, -unexpected, or retried tests. All 21 scenarios recorded successful isolated-display -presentation. This run restored the original graphics flags and removed profiling. - -| Metric | Largest measurement | Unchanged report limit | -| --- | ---: | ---: | -| Median typing | 15.5 ms | 25 ms | -| Worst key | 45.7 ms | 300 ms | -| Hidden-output restore | 395.7 ms | 1,000 ms | -| Worktree revisit | 50.3 ms | 300 ms | -| Scroll | 54.7 ms | 150 ms | - -Timer drift and queue/drop checks also passed. Coverage includes 100-pane -same-workspace and cross-workspace redraws, 50 real PTYs under held-ACK pressure, -and the original plain/Latin/title/rich-model hidden-output scenarios. No workload -or performance limit changed. - -The correction presents the benchmark window only when explicitly enabled inside -`xvfb-run` on a GitHub-hosted Linux runner. Ordinary local automation stays -windowless; production launch policy and hidden-terminal delivery remain unchanged. -For comparable Linux latency evidence, use the Terminal Perf workflow: a never- -presented local Linux window can still encounter the same compositor throttle. -Temporary profiling hooks and the comparison workflow were removed before the PR. diff --git a/docs/reference/terminal-perf-report-budgets.md b/docs/reference/terminal-perf-report-budgets.md deleted file mode 100644 index 54d26c3d2a5..00000000000 --- a/docs/reference/terminal-perf-report-budgets.md +++ /dev/null @@ -1,59 +0,0 @@ -# Terminal performance report budgets - -The saved-report gate is a performance regression gate, deliberately stricter than -some Electron test timeouts. Passing the Electron suite does not establish that a -slow sample is acceptable. CLI and HTML reports use the same policy in -`config/scripts/terminal-perf-report-budgets.mjs`. - -## Historical evidence (2026-09-21) - -Sample: 11 full scheduled Ubuntu runs, 32 annotation rows per run, August 1 through -September 21 (352 rows). Runs include failures, rather than selecting only green -runs. These are samples across different revisions and shared runners, not a -controlled A/B experiment or a statistical tail-latency estimate. Original June -logs returned HTTP 410 and could not establish an original baseline. - -Values below are milliseconds except peak queue chars (JavaScript character -counts, not process memory bytes). Maximum median spans every typing scenario; -peak queue spans the active/revisit ACK-pressure scenarios. - -| Date / run | Baseline median | Maximum typing median | Latin restore | Hidden 25-pane worst key | Peak queue chars | -| ----------------------------------------------------------------------- | --------------: | --------------------: | ------------: | -----------------------: | ---------------: | -| [2026-08-01](https://github.com/stablyai/orca/actions/runs/30693362353) | 11.7 | 12.7 | 1492.0 | 61.2 | 1867776 | -| [2026-08-03](https://github.com/stablyai/orca/actions/runs/30802935899) | 7.1 | 9.4 | 414.0 | 12.8 | 294912 | -| [2026-08-15](https://github.com/stablyai/orca/actions/runs/31875140053) | 8.9 | 13.8 | 1297.5 | 15.4 | 2523136 | -| [2026-08-24](https://github.com/stablyai/orca/actions/runs/32708128219) | 12.2 | 12.2 | 463.6 | 16.3 | 3227648 | -| [2026-09-01](https://github.com/stablyai/orca/actions/runs/33488510218) | 6.6 | 6.8 | 302.2 | 11.5 | 360448 | -| [2026-09-07](https://github.com/stablyai/orca/actions/runs/34102348134) | 7.5 | 11.3 | 328.2 | 269.6 | 2818048 | -| [2026-09-11](https://github.com/stablyai/orca/actions/runs/34580494139) | 9.5 | 10.4 | 233.5 | 186.2 | 2441216 | -| [2026-09-15](https://github.com/stablyai/orca/actions/runs/34948597774) | 10.3 | 12.6 | 282.2 | 1178.4 | 2818048 | -| [2026-09-17](https://github.com/stablyai/orca/actions/runs/35201162712) | 10.3 | 12.7 | 1182.2 | 15.7 | 2998272 | -| [2026-09-19](https://github.com/stablyai/orca/actions/runs/35432611564) | 7.6 | 10.3 | 236.7 | 1435.5 | 2588672 | -| [2026-09-21](https://github.com/stablyai/orca/actions/runs/35579708852) | 7.8 | 11.6 | 1640.7 | 2090.6 | 2523136 | - -## Decisions - -- Median typing: **25 ms**, tightened from 75 ms. The largest observed median was - 13.8 ms, leaving about 81% headroom without accepting a sustained 5x slowdown. -- Worst key: retain **300 ms**, including stress scenarios. Typical per-scenario - worst-key samples were tens of milliseconds; isolated 1–3 second samples are - failures to investigate, not a reason to adopt the e2e 3–3.5 second ceiling. -- Revisit: retain **300 ms**. Median of the 11 revisit samples was 160.7 ms; - the 1,030.2 ms outlier remains a failure. -- Restore: retain **1,000 ms**. Median restore per scenario ranged from 123 to - 493.9 ms. Observed outliers up to 1,706.4 ms do not justify a 4 second budget. -- Timer drift: retain **150 ms** and the pre-existing **3,500 ms** allowance only - for injected same/cross-workspace redraw scenarios. No new scenario receives - the broad allowance. This retains the CLI gate's existing policy in HTML too. -- Scroll: retain **150 ms**. Dropped backlogs: retain **zero**. -- Current queue: retain **2 Mi characters** everywhere. Only the transient peak - in active/revisit ACK-pressure scenarios gets **3.5 Mi characters** (3,670,016). - The maximum observed peak was 3,227,648 (3.08 Mi), leaving about 14% headroom. - The old 2 Mi peak budget rejects ordinary deliberately held-ACK bursts; the - proposed 5 Mi e2e ceiling was unnecessarily loose. Other scenarios keep 2 Mi. - -For the linked September 21 report, the two peak-queue failures are corrected; -the three slow restores, worst-key stall, and timer stall still fail. This change -does not claim to fix those stalls or make that run green. Re-evaluate future -budget changes against recorded measurements; do not set limits just above a new -failure or mirror relaxed test timeouts. diff --git a/docs/reference/terminal-startup-timing.md b/docs/reference/terminal-startup-timing.md deleted file mode 100644 index bf79200fdfc..00000000000 --- a/docs/reference/terminal-startup-timing.md +++ /dev/null @@ -1,30 +0,0 @@ -# Terminal startup timing - -For #19333, enable the renderer's opt-in recorder in its DevTools console before opening a new terminal: - -```js -localStorage.setItem('orca:terminal-startup-timing', '1') -``` - -Remove the key to disable it. Existing sessions are unaffected. To capture the existing host spawn phases, start the host with `ORCA_PTY_SPAWN_TIMING=1`. Do not restart a host with active work just to enable diagnostics. - -The renderer emits one `terminal_startup_timing` breadcrumb per transport callback generation through the existing local diagnostic channel. In `main.trace.ndjson`, find the `renderer.breadcrumb` record whose `breadcrumb.name` matches. The host's existing console timing line also becomes a `pty.spawn.timing` trace record. Correlate available PTY IDs; renderer generation distinguishes retries. Compare elapsed durations within each process, not wall clocks across hosts. - -Renderer offsets are monotonic milliseconds from callback-generation creation immediately before a transport operation: - -| Field | Observation | -|---|---| -| connected | Transport connection callback accepted for the current generation | -| liveData | First nonempty live delivery, including control-only output | -| submitted | First live batch sent to the renderer output scheduler | -| writeStarted | Scheduler invokes the batch's pre-write callback | -| parsed | Xterm invokes that batch's completion callback | -| renderEvent | First public xterm render event after the batch starts writing | - -A render event can precede the parse callback. These observations do not establish the first printable glyph, physical screen presentation, React mount time or click-to-paint latency. Replay and synthetic reset writes do not claim the first live batch. A replay or resize can still contribute to a render event after a live write, so the event is temporal evidence rather than attribution to exact content. Hidden or restored panes may never submit a live batch; missing fields remain missing. A queue-cap warning can inherit the pre-write callback while discarding the original batch’s parse callback. In that case writeStarted/renderEvent describe incomplete pre-write activity, not successful delivery of the original batch; outcome cannot be observed without parsed. - -The recorder ends after connection, parse and render observations, or on replacement, disposal, error or a ten-second diagnostic deadline. It retains only phase numbers and identifiers, with one timer and at most one render listener while enabled. It does not retain terminal text, commands, credentials or transcript buffers. Disabled recording adds no listeners or timers. - -Host `phaseDurations` preserve the current phase boundaries: the timer starts after initial ownership lookups and logs before all commit/serializer work finishes. `totalMs` is that measured interval, not full IPC latency. `provider_spawn` includes provider call and surrounding reconciliation; it is not raw process creation time. The enclosing trace record is a diagnostic snapshot, not a span covering that interval. - -Reliability invariant: diagnostics must not change terminal output, delivery credits, provider ownership or spawn outcome. Failure source: Windows OMP first-paint report #19333. Oracle: opt-in recorder tests distinguish queued, parsed and render milestones; existing live-delivery and synchronized-output suites preserve output behavior. No matching startup diagnostic reliability gate exists; full click-to-physical-presentation remains an explicit validation gap. Native, daemon, WSL and SSH execution remain host-owned; this adds no wire fields or remote process queries. Mobile has no recorder change. macOS/Linux/Windows renderer timing uses the same public xterm events; physical-device timing requires a separate capture. diff --git a/docs/reference/windows-cmd-shim-resolution.md b/docs/reference/windows-cmd-shim-resolution.md deleted file mode 100644 index c17380e800d..00000000000 --- a/docs/reference/windows-cmd-shim-resolution.md +++ /dev/null @@ -1,77 +0,0 @@ -# Resolving Windows `.cmd` shims past cmd.exe - -Node refuses to spawn a `.cmd`/`.bat` target without a shell (the -CVE-2024-27980 mitigation), so `resolveSpawn` has to make `cmd.exe` the program -and hand it `/d /v:off /s /c ""`. For an agent CLI that -means a long `cmd.exe /c` line whose caret-escaped payload is natural-language -prompt text — which Microsoft Defender for Endpoint's command-line model scores -as obfuscation. `codex.cmd` appeared in the spawn cluster of an MDE incident -against Orca for exactly this reason. - -`src/shared/child-process/windows-cmd-shim-resolution.ts` sidesteps it. npm's -`cmd-shim` and pnpm's `@zkochan/cmd-shim` generate files whose entire body is -"find a Node interpreter and run this script". Reading one lets `resolveSpawn` -spawn `node.exe