docs: remove stale internal reference documentation (#26328)

* docs: remove stale internal reference documentation

* test: avoid pooled relay hook sockets across fake clock advances

* test: remove checks for deleted headless server documentation
This commit is contained in:
Neil
2026-10-07 15:43:39 -07:00
committed by GitHub
parent b57bb7c351
commit 16cd1e0c75
77 changed files with 25 additions and 12571 deletions
-1
View File
@@ -29,7 +29,6 @@ pnpm dev
Ordinary installs include native optional dependencies for the current OS and CPU only.
Before a cross-architecture build (including `pnpm build:mac`, which produces both x64 and
arm64 artifacts by default), run `pnpm install:release` to add the other CPU's variants.
See [the install policy](../docs/reference/pnpm-install-policy.md).
## Branch Naming
+2 -1
View File
@@ -11,6 +11,7 @@
<!-- What problem does this solve, and why is this approach better than the alternatives you considered? -->
## Linked Issue
_If you do not have one and are an outside contributors, your PR **wiil** be ignored. Refs is not sufficient. Link an actual issue_
<!-- Link the issue this PR addresses, there should ALWAYS be one (for outside contributors) -->
<!-- SPECIAL CASE: If you are a maintainer (member of stablyai org) AVOID opening needless issues. Only attach pre-existing ones -->
@@ -39,7 +40,7 @@ Fixes #
## Agent skill upstream boundary
- [ ] Not applicable, or this change follows `docs/reference/agent-skill-sharing-upstream-boundary.md` and copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.
- [ ] Not applicable, or this change copies or mechanically translates no upstream skill-installer source, tests, fixtures, registry entries, path tables, comments, or documentation.
## Notes
+2 -37
View File
@@ -90,10 +90,9 @@ design-docs/
# Machine-local agent hook endpoint files may contain auth tokens.
/agent-hooks/
# Local-only design/planning docs (not checked in), including most of docs/reference/.
# Local-only design/planning docs (not checked in).
# Durable docs that should be tracked must live in one of the allow-listed
# locations below (assets, readme, STYLEGUIDE, mobile terminal shortcut bar,
# and the tracked reference docs linked from AGENTS.md / README.md).
# locations below (assets, readme, STYLEGUIDE, and mobile terminal shortcut bar).
docs/**
!docs/
# The deployable docs app is source, not local engineering notes.
@@ -117,40 +116,6 @@ docs/**
!docs/audits/crashpad-read-limit/source-hashes.json
!docs/agent-skill-sharing-implementation-checklist.md
!docs/mobile-terminal-shortcut-bar.md
!docs/reference/
!docs/reference/agent-pty-transcript-capture.md
!docs/reference/agent-session-search-query-tuning.md
!docs/reference/agent-session-search-contract.md
!docs/reference/agent-status-store.md
!docs/reference/antigravity-readiness-evidence.md
!docs/reference/antivirus-prerelease-clearance.md
!docs/reference/cline-and-prime-agent-readiness-evidence.md
!docs/reference/git-compatibility.md
!docs/reference/headless-linux-server.md
!docs/reference/ime-regression-checklist.md
!docs/reference/jcode-hook-events.md
!docs/reference/linux-glibc-compatibility.md
!docs/reference/macos-press-and-hold.md
!docs/reference/orcad-operations.md
!docs/reference/pnpm-install-policy.md
!docs/reference/relay-grace-time-reconfiguration.md
!docs/reference/windows-cmd-shim-resolution.md
!docs/reference/windows-daemon-host-relocation.md
!docs/reference/windows-edr-posture.md
!docs/reference/windows-msys-job-breakaway.md
!docs/reference/windows-process-enumeration.md
!docs/reference/wsl-runner-verification.md
!docs/reference/remote-wire-compatibility.md
!docs/reference/renderer-agent-status-performance.md
!docs/reference/ssh-execution-boundary.md
!docs/reference/ssh-host-key-verification.md
!docs/reference/ssh-reconnect-source-recovery.md
!docs/reference/windows-setup-shell.md
!docs/reference/windows-terminal-shell-selection.md
!docs/reference/worktree-scan-fingerprint.md
!docs/reference/wsl-command-execution.md
!docs/reference/wsl-probe-failure-semantics.md
!docs/reference/xterm-patch-regeneration.md
# Stably CLI (only docs/ are tracked)
.stably/*
+15 -15
View File
@@ -73,24 +73,24 @@ Orca targets macOS, Linux, and Windows. Keep all platform-dependent behavior beh
- **Keyboard shortcuts**: Never hardcode `e.metaKey`. Use a platform check (`navigator.userAgent.includes('Mac')`) to pick `metaKey` on Mac and `ctrlKey` on Linux/Windows. Electron menu accelerators should use `CmdOrCtrl`.
- **Shortcut labels in UI**: Display `⌘` / `⇧` on Mac and `Ctrl+` / `Shift+` on other platforms.
- **File paths**: Use `path.join` or Electron/Node path utilities — never assume `/` or `\`.
- **Windows terminal shells**: `--shell` picks the shell a terminal _is_; `--command` is typed into whatever shell the host spawned, so a shell choice routed through `command` silently becomes a child process. See [`docs/reference/windows-terminal-shell-selection.md`](./docs/reference/windows-terminal-shell-selection.md).
- **Windows setup scripts**: the setup/issue-command runner is a `.cmd` batch file unless the script starts with a `#!` line — never derive that from the user's terminal-shell preference, and never launch a `.cmd` runner with a bare `cmd.exe /c` from a Git Bash pane (MSYS rewrites the `/c`). See [`docs/reference/windows-setup-shell.md`](./docs/reference/windows-setup-shell.md).
- **Windows child processes**: start them through `runProcess`/`spawnProcess` in `src/shared/child-process/` — never `child_process` directly. It pins `windowsHide`, refuses `shell: true`, and encodes `.cmd`/`.bat` arguments so neither `CommandLineToArgvW` nor `cmd.exe` mangles them. Recognised npm/pnpm `.cmd` shims are resolved to their real target so the spawn skips `cmd.exe` entirely; see [`docs/reference/windows-cmd-shim-resolution.md`](./docs/reference/windows-cmd-shim-resolution.md) before adding a shim shape or debugging one.
- **Windows terminal shells**: `--shell` picks the shell a terminal _is_; `--command` is typed into whatever shell the host spawned, so a shell choice routed through `command` silently becomes a child process.
- **Windows setup scripts**: the setup/issue-command runner is a `.cmd` batch file unless the script starts with a `#!` line — never derive that from the user's terminal-shell preference, and never launch a `.cmd` runner with a bare `cmd.exe /c` from a Git Bash pane (MSYS rewrites the `/c`).
- **Windows child processes**: start them through `runProcess`/`spawnProcess` in `src/shared/child-process/` — never `child_process` directly. It pins `windowsHide`, refuses `shell: true`, and encodes `.cmd`/`.bat` arguments so neither `CommandLineToArgvW` nor `cmd.exe` mangles them. Recognised npm/pnpm `.cmd` shims are resolved to their real target so the spawn skips `cmd.exe` entirely.
- **Ripgrep**: Orca bundles `rg` for every platform, WSL, and SSH remotes. Spawn it through `spawnBundledRipgrep` (main) or `resolveRelayRipgrepCommand` (relay), never a bare `'rg'` — Windows resolves a bare name in the spawn cwd before PATH. Don't add git/readdir fallbacks locally; the relay's chain exists only for hosts an upload never reached.
- **Windows process enumeration**: read the table through `src/main/windows/windows-process-table.ts`, never by forking `powershell.exe`. See [`docs/reference/windows-process-enumeration.md`](./docs/reference/windows-process-enumeration.md).
- **Windows MSYS/Git Bash panes**: their children break away from the per-PTY job unless it is created without `JOB_OBJECT_LIMIT_BREAKAWAY_OK`, and a `conpty.node` built before that fix passes every existing gate. Before changing the per-PTY job or debugging `windows-msys-job.win32.test.ts`, read [`docs/reference/windows-msys-job-breakaway.md`](./docs/reference/windows-msys-job-breakaway.md).
- **Windows daemon-host relocation**: the terminal daemon runs from a copy of the app runtime under `%LOCALAPPDATA%`, which is what survives an auto-update. Before touching that copy, its exe name, or the NSIS uninstall macro, read [`docs/reference/windows-daemon-host-relocation.md`](./docs/reference/windows-daemon-host-relocation.md).
- **Windows EDR signal**: don't add `-ExecutionPolicy Bypass`, `-EncodedCommand`, `cmd.exe /c` with escaped free text, per-operation interpreter spawning, or runtime `Add-Type` compilation without reading [`docs/reference/windows-edr-posture.md`](./docs/reference/windows-edr-posture.md) first — behavioural EDR scores each of those, and being signed does not clear them. For file verdicts on the bytes we ship — antivirus false positives, and the vendor programs that clear a release before users meet the detection — see [`docs/reference/antivirus-prerelease-clearance.md`](./docs/reference/antivirus-prerelease-clearance.md).
- **WSL commands**: build argv with `buildWslExecArgs` (always `--exec` — under `--`, `wsl.exe` expands `$name` in every argument and silently rewrites the script), and fence anything whose stdout you parse with `buildWslCapturedLoginShellCommand`, because the interactive login shell prints the distro banner to stdout. See [`docs/reference/wsl-command-execution.md`](./docs/reference/wsl-command-execution.md).
- **Linux native modules**: keep the glibc floor at Ubuntu 20.04 / glibc 2.31. A module compiled from source on a newer runner can reference symbol versions absent on the floor and crash the app on startup. See [`docs/reference/linux-glibc-compatibility.md`](./docs/reference/linux-glibc-compatibility.md); packaging fails if a bundled native binary needs newer glibc.
- **Windows process enumeration**: read the table through `src/main/windows/windows-process-table.ts`, never by forking `powershell.exe`.
- **Windows MSYS/Git Bash panes**: their children break away from the per-PTY job unless it is created without `JOB_OBJECT_LIMIT_BREAKAWAY_OK`, and a `conpty.node` built before that fix passes every existing gate.
- **Windows daemon-host relocation**: the terminal daemon runs from a copy of the app runtime under `%LOCALAPPDATA%`, which is what survives an auto-update.
- **Windows EDR signal**: don't add `-ExecutionPolicy Bypass`, `-EncodedCommand`, `cmd.exe /c` with escaped free text, per-operation interpreter spawning, or runtime `Add-Type` compilation without reviewing the EDR impact first — behavioural EDR scores each of those, and being signed does not clear them.
- **WSL commands**: build argv with `buildWslExecArgs` (always `--exec` — under `--`, `wsl.exe` expands `$name` in every argument and silently rewrites the script), and fence anything whose stdout you parse with `buildWslCapturedLoginShellCommand`, because the interactive login shell prints the distro banner to stdout.
- **Linux native modules**: keep the glibc floor at Ubuntu 20.04 / glibc 2.31. A module compiled from source on a newer runner can reference symbol versions absent on the floor and crash the app on startup. Packaging fails if a bundled native binary needs newer glibc.
## Native Dependency Installs
Ordinary `pnpm install` covers the host OS and CPU only. Before packaging for another architecture — including `pnpm build:mac`, which builds x64 and arm64 by default — run `pnpm install:release`. electron-builder only warns on a missing `extraResources` source, so the `beforePack` guard is what turns a thin install into a build failure instead of a silently broken artifact; see [`docs/reference/pnpm-install-policy.md`](./docs/reference/pnpm-install-policy.md).
Ordinary `pnpm install` covers the host OS and CPU only. Before packaging for another architecture — including `pnpm build:mac`, which builds x64 and arm64 by default — run `pnpm install:release`. electron-builder only warns on a missing `extraResources` source, so the `beforePack` guard is what turns a thin install into a build failure instead of a silently broken artifact.
## SSH Use Case
All changes must consider the SSH use case. Don't assume local-only execution. Before changing anything that reports on, stops, or lists remote work, follow [`docs/reference/ssh-execution-boundary.md`](./docs/reference/ssh-execution-boundary.md): the execution host owns everything that touches execution, and loss of contact is never evidence of process death — the verdict vocabulary is `live` / `unverifiable` / `exited`, with no synonyms.
All changes must consider the SSH use case. Don't assume local-only execution. Before changing anything that reports on, stops, or lists remote work, the execution host owns everything that touches execution, and loss of contact is never evidence of process death — the verdict vocabulary is `live` / `unverifiable` / `exited`, with no synonyms.
## Folder Workspace Use Case
@@ -98,19 +98,19 @@ All changes must consider folder workspaces as well as git worktrees. Don't assu
## Agent Status
The execution host owns agent status in one store, the hook server's, and every reader (sidebar, `worktree ps`, mobile, dashboard) subscribes to it. Before adding a producer, a cache, or a reader-side precedence rule, read [`docs/reference/agent-status-store.md`](./docs/reference/agent-status-store.md): new producers write into that store, and readers keep only presentation policy.
The execution host owns agent status in one store, the hook server's, and every reader (sidebar, `worktree ps`, mobile, dashboard) subscribes to it. New producers write into that store, and readers keep only presentation policy.
## Agent Terminal Screens
A rule that reads what an agent CLI paints on a terminal — readiness, blocked prompts, idle — must be written against a captured transcript, not a remembered screen. Record one with [`docs/reference/agent-pty-transcript-capture.md`](./docs/reference/agent-pty-transcript-capture.md), which keeps escapes and wrapping intact and scrubs account identifiers before they reach git. Antigravity readiness has no transcript yet and five failed attempts without one; before touching it, read [`docs/reference/antigravity-readiness-evidence.md`](./docs/reference/antigravity-readiness-evidence.md).
A rule that reads what an agent CLI paints on a terminal — readiness, blocked prompts, idle — must be written against a captured transcript, not a remembered screen. Use `config/scripts/capture-agent-pty-transcript.mjs` to keep escapes and wrapping intact and scrub account identifiers before they reach git.
## Remote Wire Compatibility
Clients and remote Orca servers update independently, so mixed versions are the normal state. Before changing anything a paired client and host exchange — RPC params, stream frames, or the content either side publishes over them — follow [`docs/reference/remote-wire-compatibility.md`](./docs/reference/remote-wire-compatibility.md). A new optional field is safe; a new stream opcode must be capability-negotiated because decoders drop unknown opcodes silently; and changing what the host publishes reaches old clients even with no wire change.
Clients and remote Orca servers update independently, so mixed versions are the normal state. Before changing anything a paired client and host exchange — RPC params, stream frames, or the content either side publishes over them — preserve compatibility with older peers. A new optional field is safe; a new stream opcode must be capability-negotiated because decoders drop unknown opcodes silently; and changing what the host publishes reaches old clients even with no wire change.
## Git Binary Compatibility
Orca runs the user's Git binary on native, WSL, and SSH hosts, which may all have different versions. Treat Git 2.25 as the core-workflow baseline and follow [`docs/reference/git-compatibility.md`](./docs/reference/git-compatibility.md).
Orca runs the user's Git binary on native, WSL, and SSH hosts, which may all have different versions. Treat Git 2.25 as the core-workflow baseline.
When adding or changing a Git command:
-1
View File
@@ -218,7 +218,6 @@ Works with **any CLI agent** — if it runs in a terminal, it runs in Orca.
- **[Download from onOrca.dev](https://onorca.dev/download)**
- Or grab a build directly: [macOS Apple Silicon](https://github.com/stablyai/orca/releases/latest/download/orca-macos-arm64.dmg) · [macOS Intel](https://github.com/stablyai/orca/releases/latest/download/orca-macos-x64.dmg) · [Windows (.exe)](https://github.com/stablyai/orca/releases/latest/download/orca-windows-setup.exe) · [Linux AppImage](https://github.com/stablyai/orca/releases/latest/download/orca-linux.AppImage) · [All builds](https://github.com/stablyai/orca/releases/latest)
- Running `orca serve` on a headless Linux server? See the [headless Linux server guide](docs/reference/headless-linux-server.md).
_Or via a package manager:_
+1 -2
View File
@@ -5486,8 +5486,7 @@
"coveredProviders": ["remote-runtime"],
"coverageNotes": "Deterministic main-IPC contract tests cover disconnect-driven close delivery, exactly-once close, per-subscription teardown isolation against a failing socket close and a throwing liveness probe, containment of a throwing renderer send on the unguarded host-close path, and continued suppression of stale payloads from a retired transport. A headed paired-server journey (real Orca host plus a separate paired Orca desktop client) covers hidden-but-mounted reveal, cold-parked reveal, and cold-parked reveal across a disconnect/reconnect. Live Linux and Windows paired-server evidence and real sleep/wake transport loss remain uncollected.",
"motivatingLinks": [
"tests/e2e/paired-remote-terminal-parked-reveal-interactivity.spec.ts",
"docs/reference/headless-linux-server.md"
"tests/e2e/paired-remote-terminal-parked-reveal-interactivity.spec.ts"
],
"invariant": "Every renderer-held runtime subscription receives exactly one terminal close event when its transport is retired, including when the retirement advanced the transport generation first, and a single failing teardown never abandons that environment's remaining subscriptions nor escapes into the transport that reported the close. Payload frames from a retired transport stay suppressed. A revealed remote terminal therefore reattaches over a live multiplex connection: its buffer restores, typed input reaches the host PTY, the echo paints without a tab flip, and the PTY converges on the revealed pane grid.",
"oracle": "The main IPC contract test subscribes terminal.multiplex through the real handler, disconnects the environment, and asserts the renderer received exactly one {type: close} subscription event. Two isolation tests subscribe a second stream to the same environment and make the first one fail -- in its socket close, and in the liveness probe inside notifyClosed -- then assert the disconnect does not throw, both transports closed, and every close the renderer could still receive was delivered. A third drives a host-initiated close through the transport callback, which is the one notifyClosed call site with no surrounding guard, with a renderer send that throws, and asserts it cannot escape into the WebSocket close handler. A fourth test asserts that after retirement a late response frame is not forwarded and a late transport close does not re-send. The paired-server journey runs three reveal scenarios against one real host and one real paired desktop client, and for each records buffer restore, host-side receipt of the typed marker through an out-of-band host sink file, live paint without a tab flip, and PTY-versus-pane grid convergence.",
-1
View File
@@ -222,7 +222,6 @@ Fonctionne avec **n'importe quel agent CLI** — s'il tourne dans un terminal, i
- **[Télécharger depuis onOrca.dev](https://onorca.dev/download)**
- Ou récupérez un build directement : [macOS Apple Silicon](https://github.com/stablyai/orca/releases/latest/download/orca-macos-arm64.dmg) · [macOS Intel](https://github.com/stablyai/orca/releases/latest/download/orca-macos-x64.dmg) · [Windows (.exe)](https://github.com/stablyai/orca/releases/download/v1.4.147-rc.3/orca-windows-setup.exe) · [Linux AppImage](https://github.com/stablyai/orca/releases/latest/download/orca-linux.AppImage) · [Tous les builds](https://github.com/stablyai/orca/releases/latest)
- **Sous Windows :** utilisez la [dernière RC (`v1.4.147-rc.3`)](https://github.com/stablyai/orca/releases#release-v1.4.147-rc.3) — elle inclut des correctifs Windows absents de la stable.
- Vous lancez `orca serve` sur un serveur Linux headless ? Consultez le [guide serveur Linux headless](../reference/headless-linux-server.md).
_Ou via un gestionnaire de paquets :_
-1
View File
@@ -217,7 +217,6 @@ diff의 어느 줄에든 코멘트를 남기고 에이전트에게 바로 보내
- **[onOrca.dev에서 다운로드](https://onorca.dev/download)**
- 또는 빌드를 직접 받기: [macOS Apple Silicon](https://github.com/stablyai/orca/releases/latest/download/orca-macos-arm64.dmg) · [macOS Intel](https://github.com/stablyai/orca/releases/latest/download/orca-macos-x64.dmg) · [Windows (.exe)](https://github.com/stablyai/orca/releases/latest/download/orca-windows-setup.exe) · [Linux AppImage](https://github.com/stablyai/orca/releases/latest/download/orca-linux.AppImage) · [전체 빌드](https://github.com/stablyai/orca/releases/latest)
- headless Linux 서버에서 `orca serve`를 실행하시나요? [Headless Linux 서버 가이드](../reference/headless-linux-server.md)를 확인하세요.
_또는 패키지 매니저로 설치:_
-1
View File
@@ -217,7 +217,6 @@ Funciona com **qualquer agente CLI** — se roda em um terminal, roda no Orca.
- **[Baixe em onOrca.dev](https://onorca.dev/download)**
- Ou baixe um build diretamente: [macOS Apple Silicon](https://github.com/stablyai/orca/releases/latest/download/orca-macos-arm64.dmg) · [macOS Intel](https://github.com/stablyai/orca/releases/latest/download/orca-macos-x64.dmg) · [Windows (.exe)](https://github.com/stablyai/orca/releases/latest/download/orca-windows-setup.exe) · [Linux AppImage](https://github.com/stablyai/orca/releases/latest/download/orca-linux.AppImage) · [Todos os builds](https://github.com/stablyai/orca/releases/latest)
- Rodando `orca serve` em um servidor Linux headless? Veja o [guia de servidor Linux headless](../reference/headless-linux-server.md).
_Ou por um gerenciador de pacotes:_
@@ -1,97 +0,0 @@
# Administer agent skill sharing
This guide describes the first-release access, lifecycle, retention, and recovery contract for
Orca skill sharing. The operator runbook remains the source of truth for incident commands and
environment-specific procedures.
## Access model
- Shared bundles are unlisted bearer resources. Orca provides no public browse, search, recipient
inventory, or package index.
- Anyone with an active, unexpired link can inspect the package and request a short-lived download
grant without signing in.
- Publishing, package/version management, owned-link inventory, revocation, and deletion require
an authenticated package owner with current organization access.
- Missing, expired, revoked, deleted, and unauthorized resources return the same non-disclosing
response.
- Desktop and remote runtimes receive no GCP identity or long-lived storage credential.
The durable share ID is a credential. Do not put it in tickets, logs, analytics, or support
bundles. Use revocation if a link may have reached an unintended recipient.
## Revocation and deletion
Revoking a share immediately blocks new resolution and download grants. A generation-bound grant
issued before revocation can work until its five-minute expiry. Already installed skills remain on
recipient machines.
Package deletion follows this order:
1. Mark the package deleted and revoke its active shares.
2. Dereference retained versions transactionally.
3. Delete only an object generation that no retained version references.
4. Reconcile bounded pending deletions after partial database or GCS failures.
A version cannot be deleted while an active pinned share references it. Deletion uses the exact
recorded GCS generation and never overwrites an immutable published key.
## User and organization departure
Packages belong to an owner tenant and record the publishing user. In an organization tenant,
another current member can manage the package after its publisher leaves; Orca does not rewrite
the recorded creator. Removing a user does not automatically revoke the organization's links,
delete packages, or remove installed copies.
Before deleting an organization tenant:
1. Disable new grants for the tenant.
2. Have an authorized operator inventory and revoke active shares.
3. Decide whether packages transfer to another authorized owner, remain retained, or are deleted.
4. Resolve legal hold, erasure, and audit-retention requirements.
5. Apply the coordinated metadata and object lifecycle; do not bypass reference checks.
The product does not yet encode a universal ownership-transfer or legal-retention policy. Privacy,
security, and the organization owner must approve the applicable policy before external rollout.
Until that decision is recorded, preserve metadata and soft-deleted generations rather than
guessing.
## Retention contract
| Data | Default behavior |
| -------------------------------------- | --------------------------------------------------------- |
| Upload policy and pending upload row | Expires after 15 minutes |
| Abandoned `uploads/` quarantine object | Deleted by GCS after one day |
| Published immutable package object | No age-based deletion; retained while referenced |
| Issued download grant | Expires after five minutes |
| Deleted package object | Recoverable through GCS soft delete for seven days |
| PostgreSQL metadata | Covered by backups and seven-day point-in-time recovery |
| Installed recipient copy | Independent local data; Cloud deletion does not remove it |
| Audit event | Follows the approved audit-retention policy |
Organization retention, legal hold, and erasure requirements take precedence over product rollback
retention. Product deletion is not a legal-hold mechanism.
## Audit and privacy
Audit records may include package/version IDs, actor IDs, event category, outcome, and timestamp.
They must not include skill contents, filenames, manifests, organization membership lists, local
paths, durable share URLs, upload policies, download grants, or credentials. Anonymous abuse
controls must not persist raw requester IP addresses.
Normal Cloud logs are limited to route, method, status, duration, and bounded failure categories.
Use seeded privacy canaries when validating staging logs and diagnostic exports.
## Recovery and incident controls
Upload, download, and remote-install operations have independent kill switches. Disable the
narrowest affected operation; existing local discovery and installs continue to work.
Coordinate PostgreSQL point-in-time recovery with GCS generation recovery. Restore metadata into
an isolated database, identify exact referenced generations, restore only matching soft-deleted
objects, verify archive and package identities, then transactionally repoint metadata. Keep grants
disabled until bearer preview and a generation-bound download pass.
See the Orca Cloud `docs/skill-sharing-runbook.md` for deployment controls, reconciliation,
saturation, signing failures, database outages, and the guarded restore workflow. Security
invariants and unresolved release gates are recorded in
[Agent skill sharing threat model](./agent-skill-sharing-threat-model.md).
@@ -1,131 +0,0 @@
# Capturing an agent PTY transcript
Orca's readiness and blocked-prompt rules are text rules over what an agent CLI paints on a
terminal. They are only as good as the screens they were written against. This is how to record
one, byte for byte, so a rule can be pinned to evidence instead of to a remembered screen.
Related: [`antigravity-readiness-evidence.md`](./antigravity-readiness-evidence.md) names the
specific Antigravity transcripts that are still missing and what each one decides.
## The recorder
```
node config/scripts/capture-agent-pty-transcript.mjs --name <fixture-name> [options] -- <command> [args...]
```
It allocates a real PTY, spawns the agent inside it, mirrors the session to your terminal so you
can drive it by hand, and appends every byte it receives to
`src/main/runtime/__fixtures__/<fixture-name>.txt`. It does not strip escapes, fold `\r`, rewrap
lines, or normalise anything — the file is what the terminal received.
- **Ending a capture:** press <kbd>Ctrl</kbd>+<kbd>]</kbd>. The recorder consumes that key and
never forwards it, which is the only way to end a capture _while a dialog still owns the
screen_. Quitting the agent instead would first dismiss the dialog you came to record.
- `--cols N --rows M` pin the PTY size (default: your terminal's). Wrapping is part of the
evidence, so record the size — the sidecar does it for you.
- `--duration S` stops unattended after S seconds, for a screen that needs no interaction.
- `--send "<ms>:<text>"` types into the PTY at a fixed offset, repeatable, with `\r` `\n` `\t` `\e`
escapes. A dialog capture has to be driven, and an unattended run (CI, or an agent) has no TTY to
type into; the keystrokes ride the same PTY a human's would. For example, the committed
`antigravity-dialog-model-picker.txt` was recorded with
`--duration 24 --send "14000:/model" --send "16000:\r"`, which leaves the picker owning the
screen when the capture stops.
- `--note "<text>"` records the account type, plan, model and CLI version in the sidecar.
- `--out <path>` writes outside the fixture directory (use it for a first dry run).
Each capture also writes `<fixture-name>.meta.json` with the timestamp, platform, command,
PTY size, note and exit code. Commit it with the transcript; the version and account type behind
a screen are not recoverable from the bytes.
**Prerequisite:** `node-pty` must be built for plain Node:
```
node config/scripts/ensure-native-runtime.mjs --runtime=node
```
Orca itself does not need to be running, and the recorder never touches Orca state.
### Platform notes
- **macOS / Linux:** nothing special. `TERM=xterm-256color` is set for the child.
- **Windows:** run it from Windows Terminal / PowerShell, not a Git Bash (MSYS) pane — MSYS
rewrites arguments that start with `/`, which mangles the `cmd.exe /c` hand-off. A `.cmd` or
`.bat` agent shim cannot be spawned by node-pty directly, so the recorder routes those through
`cmd.exe` for you.
- **WSL:** capture _inside_ the distro (run the recorder from the distro's checkout). Recording
`wsl.exe` from the Windows side adds the login-shell banner to the transcript.
- **SSH:** record on the execution host. A transcript recorded locally is not evidence about what
a remote agent prints.
## Privacy: scrub before committing
A live agent screen routinely contains things that must not enter git history:
| Scrub | Why |
| ---------------------------------------------------------------------- | ---------------------------------------------------- |
| Account email / sign-in identifier | The account row on a ready screen prints it verbatim |
| Org, tenant or team name | Identifies a customer |
| Machine hostname and OS username | Appear in prompts, paths and the OSC title |
| Absolute home paths (`/Users/<you>`, `C:\Users\<you>`) | Contain the username |
| JWTs, `AIza…` keys, `1//…` refresh tokens, `Bearer …`, `sk-…`, `ghp_…` | Live credentials; a sign-in screen can echo one |
| Private repo, branch and ticket names | Leak roadmap detail |
| Anything you pasted into the agent during the capture | You typed it; it is in the transcript |
The recorder scans the file as soon as the capture ends and prints every hit with a line and
column. To scrub:
```
node config/scripts/capture-agent-pty-transcript.mjs --scan src/main/runtime/__fixtures__/<name>.txt --redact
```
Redaction replaces each finding with a **same-length** placeholder (`u…u@example.com`, `XXXX…`).
Length matters: a transcript's value is its exact wrapping and column alignment, and a shorter
replacement reflows the screen and destroys the evidence.
### Verify it is gone
1. `node config/scripts/capture-agent-pty-transcript.mjs --scan src/main/runtime/__fixtures__/<name>.txt`
must print `clean` and exit `0`. It recognises its own placeholders, so a scrubbed file passes.
2. Grep for the specifics the scanner cannot know:
`rg -n -i -- "$(whoami)|<your-email>|<your-org>|<your-hostname>" src/main/runtime/__fixtures__/<name>.txt`
3. Read it once with escapes visible: `LC_ALL=C cat -v src/main/runtime/__fixtures__/<name>.txt`.
The scanner matches shapes; only a human catches a project name.
4. Check the sidecar too — `--note` text is free-form and is committed.
`config/scripts/pty-transcript-secret-scan.test.mjs` re-scans every committed
`__fixtures__/*.txt`, so a transcript that skips step 1 fails the suite.
## Consuming a transcript in a test
Feed the raw bytes through the runtime rather than into a matcher directly: escape handling,
tail retention and title tracking all live in `onPtyData`, and a rule tested on pre-normalised
text is tested on something no pane ever sees.
`src/main/runtime/agent-transcript-pane-test-harness.ts` builds the pane;
`src/main/runtime/terminal-interactive-wait-visibility.test.ts` (cursor-agent) and
`src/main/runtime/antigravity-readiness-transcripts.test.ts` (Antigravity) are examples. An agent
whose readiness is read off the live screen gets its suite from
`src/main/runtime/screen-ruled-agent-transcript-suite.ts`.
## Worked example: the Antigravity captures
The six committed `antigravity-*.txt` fixtures were recorded this way on macOS against
`agy` 1.1.25. Two points generalise:
- **Reach a state without mutating the operator's config.** The ready-screen captures ran in a
directory the CLI already trusted, so no trust answer was written. Where a dialog could only be
reached by signing the operator out or deleting their settings, it was left uncaptured and
recorded as such rather than forced.
- **An environment variable is a legitimate capture knob** where a setting is not.
`AGY_CLI_HIDE_ACCOUNT_INFO=1` produced a second ready screen with no account row, which is
evidence no amount of reasoning about the first screen could have supplied. It changes nothing
on disk.
## Known gap in the existing captures
The three `cursor-agent-*.txt` fixtures contain **no escape bytes and no carriage returns**.
Whatever produced them went through a renderer and a clipboard, so they preserve wording and
box-drawing glyphs but not the caret, the cursor moves, the repaints, or whether the CLI uses the
alternate screen buffer. They are good enough for the wording-based rules built on them and are
not evidence for anything else. New captures made with this recorder keep those bytes; the
Antigravity scaffold asserts their presence so a pasted screen cannot pass as a capture.
@@ -1,172 +0,0 @@
# Agent session search contract
`AiVaultSearchRequest`, `AiVaultSearchResponse`, `AiVaultSearchHit`, and
`AiVaultSearchStatus` are defined in `src/shared/ai-vault-search-types.ts` and
validated by `src/shared/ai-vault-search-contract.ts`.
## Search and pagination
- Tool output beyond 3,072 characters per row is not indexed and not searchable; user and assistant text is indexed in full.
- A page cursor outstanding during a retention purge is refused once as `stale-cursor`; the client re-issues page 1.
- A phrase match across a chunk boundary of a long message is not supported.
`aiVault.searchSessions(request)` accepts `query`, optional `scope`
(`conversation` or `all`, default `all`), `freshness` (`indexed` or
`wait-until-current`, default `indexed`), `limit`, opaque `cursor`, `filters`,
and `debug` (default false). Conversation scope searches user and assistant text.
Filters accept `agents`, `scopePaths`, ISO `since`, and `sort` (`relevance` or
`newest`). Paths refer to the execution host and work for folders without Git.
Legacy `tier` and `refresh` fields are accepted and discarded; they do not change
the defaults. Limits use the engine's resolver: default 20, integers clamped to
1–100, fractional numbers use the default. Long queries reach the engine so it
can report truncation rather than fail validation.
Results contain `kind: 'results'`, `hits`, `page: { cursor, hasMore }`,
`generation`, `truncated: { candidates, snippets, query, freshness }`, and
`durationMs`. `snippets` is a count; the other truncation fields are booleans.
`durationMs` measures the engine search, excluding any reconciliation wait.
`debug: true` adds `debug: { route, repairedTerms?, plannerReport }`; the report
contains `route`, optional `repairedTerms`, and `scope`. Diagnostics never appear
at the top level. Status is never attached to search results.
A cursor belongs to one query, one host's index generation, and an opaque persisted
index incarnation. Query, scope, filters, and sorting must remain the same; page
size may change. Writes that advance the generation can invalidate it, including
retention purges. Clearing or rebuilding the database invalidates it even when the
new generation counter matches. A refused cursor yields
`{ kind: 'stale-cursor', generation, expectedGeneration? }` and the
client discards it and issues page 1 without a cursor. Reusing that refused cursor
continues to fail; there is no server-side cursor acknowledgement state.
Malformed cursors and cursors for a different query yield
`{ kind: 'malformed-cursor' }`. Generation checks also reject a first page if the
index changes during retrieval. Generation is a fence, not a retained snapshot:
a client cannot ask the host to recreate a previous generation.
Pages are per host only. Ordering is local to that host's query. Clients must
discard cursors when changing hosts. All-computers search, merged ordering,
per-host aggregate outcomes, and merged cursors are deferred to a separate PR.
That follow-up must define generation fencing, page-size changes, unavailable
hosts, and bounded parallel retrieval before exposing a combined result list.
## Execution host routing
Search and status address one execution host: `local`, `ssh:<target>`, or
`runtime:<environmentId>`. An omitted host means this desktop's local index.
Invalid IDs and `all` are refused; neither can widen a request to other hosts.
- `local` searches this machine's index over desktop IPC.
- `ssh:<target>` asks that relay session and nothing else.
- `runtime:<environmentId>` asks that paired runtime over its RPC. A paired
runtime answers for itself and never forwards through another desktop.
Each hit may carry `executionHostId`. The desktop stamps remote answers with
the host it addressed rather than trusting an ID returned by that host.
Local answers and older hosts may omit attribution.
## Evidence and exposure
Each hit carries agent, session ID, title, cwd, branch, updated time, message
count, score, source, and evidence. Evidence contains snippet, role, and timestamp;
it is null for operator-only matches that have no text evidence. Snippet matches
use `[[` and `]]` markers. Source presence is `present`, `unverifiable`, or
`missing`; the current engine emits the first two. Loss of contact does not prove
a source missing.
`redactForTransport(hit, transport)` is the exposure policy:
| Transport | filePath / codexHome | resumeCommand | Status `degradedRoots[].root` |
| ----------------------------------------- | ------------------------------------ | --------------------------------- | ----------------------------- |
| Desktop IPC on the same machine | Included when known, under source | Included only for present sources | Included |
| Runtime RPC on the same machine | Included when known, under source | Included only for present sources | Included |
| Relay or paired runtime/web/mobile client | Withheld; source keeps presence only | Withheld | Withheld |
`cwd`, titles, snippets, and other hit metadata remain visible to paired clients.
Snippets cross the authenticated transport as indexed; this contract does not
apply an observability redactor to transcript content. A missing Codex home is
omitted. Resume commands reuse the command stored by the transcript reader,
constructed by the sidebar's `buildAiVaultResumeCommand`; this layer does not
construct commands or execute them. The runtime uses its authenticated
`clientKind` context to distinguish paired clients from same-machine RPC, never
a request-supplied locality flag. The receiving remote client also applies the
same exposure function.
## Status, freshness, and availability
`aiVault.searchStatus()` returns `enabled`, `phase` (`idle`, `indexing`, `current`,
`degraded`, or `closed`), `filesIndexed`, `filesDue`, `filesFailed`, `degradedRoots`
(`root` and `reason`), `lastReconcileAt`, `lastSweepCompletedAt`, and `generation`.
Times are milliseconds since epoch or null. These are the indexer's observations;
an indexed row is not a new filesystem verification. A degraded root's `root` and
raw `reason` can both contain host filesystem paths. Desktop IPC and same-machine runtime RPC receive the full diagnostic;
relay and paired clients receive only the fixed reason "Source root could not
be verified." for each degraded root. The array length retains the count.
`wait-until-current` calls `service.reconcile()` before searching. The adapter
uses `indexer.reconcile({ full: false })`. After five seconds the endpoint searches
anyway and sets `truncated.freshness: true` on results. It does not cancel the
host's reconciliation. Completion before the deadline leaves the flag false;
a reconciliation error before the deadline propagates. The indexer's bounded
recent pass is not a promise that the entire historical corpus was swept.
Search unavailability is a value:
`{ kind: 'unavailable', reason: 'disabled' | 'not-ready' | 'no-service' }`.
No registered service returns `no-service`. Status without a service has
`enabled: false`, `phase: 'idle'`, zero counts and generation, empty degraded roots,
and null timestamps. This is a sentinel for an absent service, not a claim of an
empty, current index. A registered service may report disabled or not-ready.
## Boundaries and compatibility
- Desktop: `aiVault:searchSessions` and `aiVault:searchStatus`, via preload.
- Runtime and relay: `aiVault.searchSessions` and `aiVault.searchStatus`.
- CLI: `orca search` calls both over the runtime RPC, against the host that
`--environment` / `--pairing-code` selects and no other. It reuses
`createSessionSearchClient`, so an old host's refusal reaches the caller as
`unavailable/no-service` rather than an error, and needs no new capability.
In an Orca SSH terminal, the forwarded CLI defaults to the controlling Orca
runtime's index. `--path` filters that index; it does not select the SSH host.
`--environment` / `--pairing-code` can explicitly select a paired runtime.
- Desktop preload optionally accepts an execution host scope as a separate
routing argument. It addresses exactly that host; missing connections never
fall back to the local index. The web preload addresses its own paired runtime,
answers `all` and any other host with
`unavailable/no-service` rather than an error.
Requests and responses are parsed where received from another process. Existing
relay JSON-RPC request/response framing needs no new stream opcode. Following the
existing relay method probe pattern, an old host's explicit `-32601` refusal (or
runtime `method_not_found`) maps to `unavailable/no-service` on the client; status
uses the absent-service sentinel above. Transport failures, authentication errors,
and invalid payloads remain errors. Unknown request fields are stripped for wire
compatibility.
The process-local `setSessionSearchService(service | null)` registry connects
these endpoints to the production service installed by PR 3b. Desktop indexing
runs in the scanner child; orcad and the SSH relay register their own in-process
services. Registration alone does not grant consent.
## Desktop index controls (PR8)
Settings → Agent Session History controls this desktop's persisted
`aiVaultSearch { enabled, historyDays }` policy. It stays local even when another
execution host is selected. Paired clients cannot grant consent or clear an index
through this surface; SSH relay registration remains disabled without a separate
host consent mechanism.
`aiVault.clearSearchIndex()` is a no-argument desktop-only preload operation over
`aiVault:clearSearchIndex`. It addresses the local scanner child, not the selected
remote host. The child's existing interactive request lane executes `searchClear`
through `SessionSearchInstance.clear()`: close the indexer and database handles,
remove the SQLite database and sidecars, then reconstruct only if consent remains
enabled. Errors propagate to the settings pane. This operation never deletes
original transcripts. There is no new runtime or relay method. Opaque cursors also
carry a persistent database identity: clearing creates a new identity, so a
pre-clear or legacy cursor returns `stale-cursor` even if the rebuilt numeric
generation happens to match. Reopening the same database preserves its identity.
Disabling closes the indexer and keeps the index copy; clearing deletes the copy.
Changing retention reuses the existing close-and-construct policy application.
The settings pane reads status only while enabled and visible, polls only an
observed indexing phase, and stops on completion or error. Opening the pane,
changing policy, or pressing Refresh obtains a new observation. The indexer's own
schedule does not depend on the pane.
@@ -1,229 +0,0 @@
# Agent session search: query tuning
What a search costs, and what the knobs in `src/main/ai-vault-search/session-search-engine.ts`
buy. Every number here comes from `config/scripts/session-search-query-benchmark.ts`
over the synthetic corpus in `session-search-synthetic-corpus.ts`, except the
`conversation_fts` shoot-out, which writes its own corpus because the answer
turns on how much of a transcript is tool output. Nothing in this file was
measured against a real transcript, and neither benchmark must ever be pointed
at one.
## Running it
The benchmark is a top-level-await module that imports the main-process tree by
extensionless path, so it needs a bundler-backed runner rather than bare `node`:
```sh
cat > src/main/ai-vault-search/zz-bench.test.ts <<'EOF'
import { it } from 'vitest'
it('runs', { timeout: 1_800_000 }, async () => {
await import('../../../config/scripts/session-search-query-benchmark')
})
EOF
BENCH_OUT=/tmp/ss-query-bench.json pnpm test src/main/ai-vault-search/zz-bench.test.ts
rm src/main/ai-vault-search/zz-bench.test.ts
```
The `conversation_fts` shoot-out below runs the same way, importing
`config/scripts/session-search-conversation-fts-benchmark` instead, with
`CORPUS_MB` and `TOOL_SHARE` to size and shape its corpus. `config/scripts` is
not inside any typecheck project, so while that throwaway test exists `tsc`
reports TS6307 for each script it pulls in; delete it and the run is clean
again.
`BENCH_OUT` exists because vitest intercepts `console.log`; the report is written
to that path as well as printed.
## Scope: what the second FTS table buys a reader
Corpus: 40 synthetic Claude transcripts, 10.5 MB, 9,600 messages, indexed through
the real store. Eight queries, one per rung of the route ladder plus the two
shapes that skip it; 5 warm-up runs and 25 samples each. Apple silicon, warm page
cache, machine otherwise idle. Milliseconds, and p95 over 25 samples moves
several milliseconds run to run if anything else is competing for the disk.
| Scope | p50 | p95 |
| -------------- | ---- | ---- |
| `all` | 7.22 | 8.94 |
| `conversation` | 5.33 | 7.86 |
Per query, `all` then `conversation` (p50 / p95):
| Query | `all` | `conversation` |
| ------------------------------------------------ | ------------ | -------------- |
| `"terminal reattach"` (phrase) | 5.24 / 8.42 | 2.97 / 3.24 |
| `resolveTerminalPath` (identifier) | 7.55 / 8.94 | 6.47 / 6.72 |
| `src/main/…/session-transcript-reader.ts` (path) | 8.69 / 10.12 | 7.78 / 8.04 |
| `why is the daemon snapshot stale` (prose) | 7.84 / 8.57 | 5.90 / 7.01 |
| `reattahc worktre` (typo repair) | 7.30 / 7.39 | 5.53 / 5.89 |
| `index` (common term) | 5.45 / 5.66 | 3.81 / 4.02 |
| `repo:app-3` (operator only) | 0.12 / 0.16 | 0.10 / 0.10 |
| `worktree` scoped to one cwd | 1.47 / 1.63 | 1.25 / 1.49 |
Reading it:
- `conversation` is about 1.4x faster at p50 and 1.1x at p95, and it is a column
filter over the same table rather than a table of its own. Narrowing to the
two prose columns is what buys the gap: fewer postings to score. It is also
the scope where a match is something a person wrote rather than something a
tool printed.
- A `scopePaths` query is the cheapest real search on the page. It is the one
narrowing SQL can express exactly, so it seeks `sessions_cwd_key` and hands
ranking a small candidate set.
- The operator-only figure is a floor, not a typical cost. `repo:` and `path:`
are applied in JS over retrieved rows (see `session-search-row-filter` for why
they cannot be pushed into SQL), so their cost tracks how many sessions the
walk has to read before it fills a candidate set. This corpus has 40 sessions,
which is one page of that walk; an index where few sessions match the operator
will read up to the ceiling in `session-search-retrieval` instead.
## What the conversation scope costs at real corpus size
`conversation` was a second FTS table holding a copy of the two prose columns.
It is a column filter now — `{user_text assistant_text}: (…)` with bm25 weights
that zero the other two — and PR 2 deleted the table on the strength of the
shoot-out this section used to hold: the filter came in at 1.16-1.36x the p95 of
the dedicated table, under the 2x bar, while the table cost a tenth of the index
to maintain. What follows is what the shipped schema actually does, measured
again on the same corpus after the table went and tool rows were capped.
Corpus: Claude transcripts from `config/scripts/session-search-tool-heavy-corpus.ts`,
105 MB, indexed through the real store, at two points in the 80-97% band a real
transcript tree sits in. Half the tokens in tool output are words the
conversation also uses, so a conversation term really does have postings the
filter must discard. Twenty queries per rung, both scopes interleaved query by
query, warm cache; `config/scripts/session-search-scope-benchmark.ts`, run twice.
| Tool share | Rung | `all` p50 / p95 | `conversation` p50 / p95 |
| ---------- | ------ | --------------- | ------------------------ |
| 86% | phrase | 16.69 / 17.48 | 13.08 / 13.52 |
| 86% | or | 31.91 / 35.74 | 22.25 / 23.87 |
| 86% | and | 70.04 / 74.00 | 53.47 / 59.39 |
| 93% | phrase | 9.14 / 13.36 | 7.23 / 8.51 |
| 93% | or | 16.46 / 18.70 | 12.34 / 14.88 |
| 93% | and | 39.65 / 43.44 | 31.05 / 32.92 |
Three things to read out of it.
**The filter is a win, not a cost.** Every rung is faster narrow than wide, by
1.2x to 1.4x at p50. The shoot-out compared the filter against a table built for
exactly this query; against the wide table it replaces, it does what the second
table did, which is read fewer postings.
**The `and` rung is where the corpus size shows.** Those queries are eight terms,
chosen so no ordered run that long occurs and the phrase rung has to miss; a
real two-term AND sits nearer the phrase row. It is also the noisiest: the
second run's p95 reached 140 ms on one bucket, which is what twenty samples of a
70 ms query buys. Read the p50 column.
**The index is far smaller than the shoot-out's was.** 57 MB at 93% tool output
and 103 MB at 86%, against roughly 150 MB for `messages_fts` alone before PR 2
capped an indexed tool row at 3,072 characters. Most of a tool-heavy transcript
is now not in the index at all, which moves every number above and is the larger
effect of the two.
What is **not** measured here is relevance, and the column filter does carry one
ranking difference the deleted table did not. FTS5's bm25 normalises by the
whole row's length and has no per-column length, so two rows with identical
prose score differently when one also holds tool output. The rowid set is
unchanged, which is what the deletion was decided on; the order within it can
move. `session-search-engine.test.ts` pins the direction.
## `sessionCandidateLimit`
The reviewer's F13: this is a tunable default, not a constant. It bounds how many
sessions the SQL hands ranking, so it bounds both retrieval cost and how deep a
caller can page before the answer simply stops.
The limit only costs anything once more sessions match than the limit allows, so
this is measured over a second corpus: 2,500 one-turn transcripts, 10.9 MB, every
one of them matching the query. Limits are interleaved sample by sample, because
run back to back the first configuration pays for every page the OS cache had not
seen and the ordering alone moves p95 further than the limit does.
| Limit | p50 | p95 | Pages of 20 a caller can reach |
| ----- | ----- | ----- | ------------------------------ |
| 200 | 6.85 | 7.21 | 10 |
| 600 | 7.93 | 8.36 | 30 |
| 1200 | 9.55 | 10.53 | 60 |
| 2400 | 12.32 | 13.45 | 120 |
600 is the default: it costs about 16% over 200 at p50 and buys three times the
reachable depth, and the curve only turns steep past 1200. A host with a much
larger index can raise it; the result's `truncated.candidates` says when the limit
was the thing that cut the answer, so a caller never has to guess.
What is **not** measured here is relevance. These numbers say what a limit costs,
not what it retrieves. The MRR figures quoted in the BM25 weights
(`session-search-retrieval.ts`) and in the identifier shadow column
(`session-search-identifier-split.ts`) come from the original retrieval shoot-out
on real transcripts and are not reproducible from this repository. Any change to
the limit justified on relevance grounds needs an eval set, not this benchmark.
## What typo repair costs
The repair is the one rung whose cost tracks the size of the vocabulary rather
than the size of a result. It only runs for a term the scope has no posting for,
so an ordinary query never pays it; a query of nonsense pays it once per term.
Measured over a synthetic vocabulary of 1.6 M distinct terms, every term in two
rows so none is filtered out:
| Query | p50 |
| -------------------------------------- | ------ |
| one known term (no repair) | 11 ms |
| one unknown term | 10 ms |
| 39 unknown 12-character terms (480 ch) | 387 ms |
| 12 unknown 40-character terms | 99 ms |
Two things follow. The cost is linear in unknown terms and in vocabulary size,
and `search` is synchronous, so a 512-character query of nonsense holds the
thread for a third of a second on an index that large. And the scoped-count fix
made this cheaper rather than dearer — it was 737 ms before — because ordering
the vocabulary scan by term drops the sort that ordering by `doc` required, and
the counts it added are at most eight bounded probes per prefix. A cap on
unknown terms per query is recorded as a follow-up in the split plan.
## Page warmup, dropped
PR 2 deferred `warm()` — a sliced read of `messages` that pulls its pages into
the OS cache before the first query — to whoever knew which pages a read
touches. It is not re-added here, for two reasons. The measurement that
justified it (first query 1.3 s to 0.45 s) was on a 4 GB index, and neither
corpus in this file is within an order of magnitude of that, so PR 4 cannot
show a win: removing the call moved the 10.5 MB corpus's p50 by less than the
run-to-run spread. And it is a cancellable background pass, which needs an owner
with a lifecycle; a query library that holds no timers has nothing to hang the
`stopped()` on, and a fire-and-forget async read from a synchronous `search` is
a rejection nothing can supervise. It belongs with the indexer in PR 3b, which
already owns starting and stopping work.
## Not settled here
Which process may open, unlink and rebuild the index is PR 3b's decision. A
second handle that finds an older schema version replaces the file while a live
store keeps answering from the unlinked inode, and this PR is what first makes
that reachable, because it is the first thing that reads. What PR 4 does is
refuse to make it worse. The engine restores its derived vocabulary and generation
triggers before a search. A missing `messages_fts` fails clearly; the connection
owner must rebuild the source index. There is no degraded-search capability state
or query logging. Logging can be added by a caller when an evaluation consumer exists.
Each search checks the generation before retrieval and after its final content
read. A concurrent commit rejects the page with `stale-generation`, including a
first page without a cursor. The caller can retry from page one. No long-lived
read transaction is needed, and a mixed page is never returned as a valid snapshot.
Repository/path operators are applied before a phrase or AND route is accepted.
Candidate truncation is reported by the rung that answered, not by every rung
tried. Each rung of the ladder matches a superset of the one before it, so a
rung that reached its cap with no eligible sessions is always followed by one
that reaches it too: a full candidate set stays explicit either way.
The phrase and AND rungs run for prose as well as for literal-looking input,
over the query's tokens as typed rather than the stop-word-stripped OR body. A
sentence pasted out of a transcript is ordinary words in order; over OR its
common words fill the candidate limit with recent sessions and the old session
holding the sentence never reaches ranking. The cost is two FTS queries that
usually miss, which on the corpus above sits inside this harness's run-to-run
noise. A one-token query still takes the rung only when it looked literal.
@@ -1,38 +0,0 @@
# Agent skill provider paths
Last verified: 2026-08-11.
V1 supports only providers whose paths are independently established by official documentation.
The registry is deliberately small; it is not copied or synchronized from a community path table.
| Provider | Detection | Global canonical support | Workspace support | Orca placement |
| --- | --- | --- | --- | --- |
| Codex | `codex` CLI found through Orca's host-owned PATH detection | Reads `$HOME/.agents/skills` directly | Reads `.agents/skills` from the current directory through the repository root | Canonical copy only |
| Claude Code | `claude` CLI found through Orca's host-owned PATH detection | Reads `$HOME/.claude/skills` | Reads `.claude/skills` from the launch directory through the repository root, plus nested directories as files are accessed | Relative directory symlink on POSIX, directory junction on Windows, or verified independent-copy fallback |
Codex locations and symlink behavior are documented in the official OpenAI documentation:
[Build skills](https://learn.chatgpt.com/docs/build-skills#where-codex-loads-local-skills).
Claude Code locations, precedence, parent traversal, live detection, and symlink behavior are
documented in the official Anthropic documentation:
[Extend Claude with skills](https://code.claude.com/docs/en/skills#where-skills-live).
Codex therefore needs no provider-specific placement. Claude Code does not document
`.agents/skills` as a discovery root, so Orca reconciles its documented `.claude/skills` path back
to the canonical copy. Orca never replaces a path it does not own. If alias creation is unavailable,
the verified copy fallback is tracked in the install receipt so update and removal can detect drift.
## Registry change process
Every registry change requires normal code review and all of the following evidence:
1. Link current official provider documentation for global and workspace paths.
2. Record whether the provider reads `.agents/skills` directly and its documented link behavior.
3. Verify global and folder-workspace discovery on macOS, Linux, native Windows, and WSL where the
provider supports those platforms.
4. Exercise local, paired-runtime, and SSH host-owned path resolution.
5. Test alias denial, broken owned aliases, independent-copy drift, update, rollback, and removal.
6. Update mixed-version capability evidence if the placement contract changes.
Do not add automated upstream path-table synchronization. A provider release that changes discovery
semantics must enter through this review process.
@@ -1,101 +0,0 @@
# Agent skill sharing threat model
Status: implementation baseline for security and privacy review. This document does not constitute
security approval.
## Scope and trust model
This model covers private skill packaging, Orca Cloud publication and authorization, durable share
resolution, local and remote installation, provider placement, update, rollback, removal, and
operator recovery. It applies to macOS, Linux, native Windows, WSL, paired Orca runtimes, and SSH
targets.
Private means unlisted: an unpredictable active share ID is a bearer credential, while publication
and management remain access-controlled to authenticated Orca users and organizations. V1 is not
end-to-end encrypted from Orca Cloud operators. A skill is code from its author: `SKILL.md` can
change agent behavior and packaged scripts may be executed later by a user or agent, although the
installer itself never executes package content.
## Protected assets
- Skill contents, filenames, manifests, package metadata, release notes, and author identity.
- Organization membership, selected-user ACLs, durable share identifiers, and package existence.
- Signed upload policies, download grants, authentication tokens, database credentials, and GCP
service identities.
- Existing local skills, provider configuration, user modifications, provenance, and transaction
recovery state.
- Cloud package metadata, immutable object generations, audit records, and deletion state.
- Availability and cost of Cloud Run, Cloud SQL, GCS, connected runtimes, and local filesystems.
## Actors and boundaries
Actors include an authorized publisher, an authorized recipient, an authenticated but unauthorized
Orca user, a malicious skill author, a compromised renderer, an untrusted remote RPC caller, an
attacker controlling a network endpoint, and an operator with GCP or database access.
Trust boundaries are:
1. Source skill directory to the owner-private package staging directory.
2. Desktop renderer to the main process and its authenticated Cloud client.
3. Orca client to a paired runtime or SSH host across independently versioned protocols.
4. Orca Cloud API authorization to short-lived GCS access.
5. GCS quarantine to validated immutable publication and PostgreSQL metadata.
6. Extracted staging to canonical destination and provider placements.
7. Local diagnostic records to a user-reviewed support-bundle upload.
8. Terraform and deployment identities to staging and production resources.
## Security invariants
- Package identity is deterministic and binds normalized paths, exact bytes, executable state, and
immutable package/version IDs.
- Publication and installation accept only the `manifest.json` plus `skill/` envelope and never
execute archive content.
- Archive validation completes before destination mutation.
- Destination paths are resolved by the runtime that owns the host or workspace.
- Unowned or modified local content is never silently replaced or deleted.
- Final GCS objects are immutable, generation-fenced, and reachable only through fresh ACL checks
followed by short-lived grants.
- No remote caller can turn a desktop-local path into remote filesystem authority.
- Interrupted transactions converge to a verified old or requested version.
- Logs and support bundles exclude package contents and private authorization or filesystem data.
- Cloud sharing, downloading, and remote installation have independent kill switches.
## Threat register
| ID | Threat | Required controls and current evidence | Residual release gate |
| ----- | ------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| TM-01 | Source changes after preview or a link swaps bytes during packaging. | Observe the specific source directory, copy without following links, re-observe staged bytes, compare every identity, and bind preview to the final digest. Source-drift and link/special-file tests cover rejection and cleanup. | Repeat race and permission tests on every supported filesystem. |
| TM-02 | Archive traversal, drive paths, links, devices, duplicate paths, Unicode/case collisions, or decompression/resource exhaustion escape staging. | Streaming parser rejects unsafe entry classes and normalized collisions; compressed, extracted, entry, file, depth, and per-file limits are enforced during parsing and extraction. Boundary, malformed, checksum, and fuzz tests run before destination mutation. | Keep package safety suites required in release CI. |
| TM-03 | A forged manifest lies about names, bytes, executable state, or package identity. | Parse the manifest before trusting entries, require `skill/SKILL.md`, hash extracted bytes independently, recompute package identity, and compare package/version IDs and digest. | Cross-platform identical-byte digest evidence remains required. |
| TM-04 | A client chooses another home, workspace, WSL distro, SSH path, or escapes a destination root. | The executing runtime resolves home and workspace identities, realpath-checks directories, uses platform-native joins, and rejects client-supplied remote paths. `local-file` ingress is trusted in-process only. | Real SSH and mixed-version topology tests remain required. |
| TM-05 | Installation overwrites or deletion removes unowned or locally modified content. | Planner distinguishes missing, unchanged, clean, modified, unowned, external-link, broken-link, and collision states. Provenance lives outside installed content. Replacement requires an explicit conflict decision; removal revalidates ownership and digest. | Complete the remaining real junction, external-link, copy-drift, and permission matrix. |
| TM-06 | A crash between renames or receipt publication loses both old and new versions. | Same-filesystem staging, durable journals, backups, flushed receipt replacement, bounded recovery, and ownership tokens protect every commit boundary. Failure injection covers every journal transition. | Real process termination and disk/antivirus contention tests remain required. |
| TM-07 | A grant leaks through redirects, userinfo, an insecure scheme, DNS/host confusion, or an oversized stream. | Cloud requests reject redirects. Package downloads require configured origins and HTTPS, reject URL credentials and cross-origin redirects, cap redirect count, recheck expiry, stream exact expected bytes, and verify archive/package digests. | Validate approved production origins and exercise malicious network cases in staging. |
| TM-08 | Reuse of an upload ID, wrong tenant, wrong key, stale generation, or quarantine object publishes attacker-selected bytes. | Random tenant-bound upload rows, signed POST conditions, expiry, exact metadata/key/type/size validation, generation-fenced reads, streamed validation, and idempotent finalization fail closed. | Staging GCS integration must cover stale generations and lifecycle deletion. |
| TM-09 | IDOR, stale organization membership, or a guessed identifier exposes package management data or bearer-protected content. | Management requests authenticate and re-evaluate current ownership. Recipient preview and grants require the exact unpredictable, active bearer share ID and ignore legacy ACL rows. Not-found responses hide unauthorized existence. | Human authorization review and organization-departure policy approval remain required. |
| TM-10 | Content-addressed deduplication discloses another tenant's package through response shape or timing. | Tenant-scoped APIs never expose GCS keys, generations, or whether an object already existed; existing-object reuse verifies both archive and logical package identity. Cross-tenant tests prove identical archives use separate tenant-hashed objects and response shapes disclose no reuse. | Keep tenant-isolation and response-disclosure coverage required in release CI. |
| TM-11 | A compromised renderer, old client, or arbitrary RPC caller sends credentials, local paths, unknown opcodes, or unsupported operations to a host. | Main owns auth tokens and grants; remote requests use strict schemas and capabilities; package transfer is separate from install; no stream opcode was added; mixed versions fail with update-required results. | Real client-newer/server-newer and SSH parity gates remain required. |
| TM-12 | Transfer replay, overlap, disconnect, or abandoned staging consumes disk or commits different bytes. | Session count, idle time, total bytes, and chunk size are bounded. Offsets are monotonic, identical retry is idempotent, changed replay fails, commit hashes the exact staged file, and cancellation/disconnect cleanup is bounded. | Exercise offline-GCS and real disconnects at every transfer boundary. |
| TM-13 | Provider aliases or junctions escape canonical storage or point at external content. | POSIX aliases are relative from real parents, Windows junctions use absolute canonical targets, targets are revalidated, and unowned/broken/external links are conflicts. Copy fallback is independently hashed and receipt-owned. Real Windows and WSL coverage includes existing, broken, external, denied, and drifted placement behavior. | Preserve the physical placement matrix in release coverage. |
| TM-14 | Skill instructions or scripts are mistaken for trusted Orca code or executed during install. | Share/install previews identify author, organization, scripts, executable files, digest, and version. Installation never runs scripts. Trust copy says the package is code from its author. | Security and design must approve the final trust wording and accessibility behavior. |
| TM-15 | Telemetry, logs, or support bundles leak instructions, filenames, paths, ACLs, grants, policies, or credentials. | Desktop install diagnostics map values to bounded categories before tracing. Deployed staging logs contain route templates and bounded request metadata only, and the logging exclusion removes bearer URLs. Support-bundle tests inject private canaries and prove they are absent from collected output. | Preserve deployed-log and support-bundle privacy checks for future changes. |
| TM-16 | Permissive staging permissions expose package bytes to another local user. | Package archives, downloads, extraction, relayed uploads, locks, journals, and receipts use owner-private modes on POSIX; Windows uses owner-profile paths and inherited ACLs. Existing POSIX download roots are tightened before use. Real Windows, WSL, Ubuntu-floor, and SSH validation passed. | Preserve owner-private staging checks across supported hosts. |
| TM-17 | Deletion, revocation, retention, or user departure leaves unauthorized grants or unrecoverable metadata/blob divergence. | Revocation blocks new grants immediately; existing grants expire within five minutes. Database references govern final deletion, quarantine lifecycle cleans abandonment, and GCS soft delete supplies recovery. Local installs remain independent. | Approve departure/legal-retention policy and exercise coordinated database/GCS recovery. |
| TM-18 | Broad IAM, public bucket access, service-account keys, or a compromised deployment identity bypasses application authorization. | Uniform bucket access, public-access prevention, bucket-scoped object access, service-account-scoped signing, skill-secret-only access, Cloud SQL client role, no desktop IAM, and no long-lived keys are Terraform-defined. | Review the staging and production plans and verify live IAM before rollout. |
| TM-19 | Unbounded validation or request concurrency causes memory, CPU, database, storage, or egress denial of service. | Fixed streaming buffers, package limits, per-instance finalization semaphore, rate/quota limits, bounded transfer sessions, Cloud Run instance limits, lifecycle cleanup, and independent kill switches constrain work. | Complete load testing, dashboards, alerts, and budget thresholds. |
| TM-20 | Operator recovery, diagnostics, or legal workflows bypass tenant isolation or leak content. | Runbooks require generation-specific recovery, coordinated PostgreSQL/GCS restoration, audited lifecycle actions, and no package contents in normal logs. | Security/privacy approval and restricted break-glass procedure remain required. |
## Required review evidence
Security and privacy approval must not rely on this document alone. Reviewers need:
- Package/admission schemas and stable failure categories.
- Archive parser, extraction containment, transaction, provenance, and removal tests.
- Cloud authorization, object-generation, tenant-isolation, deletion, and recovery tests.
- Terraform plans plus live staging IAM, bucket, Cloud Run, Secret Manager, and Cloud SQL evidence.
- Real macOS, Linux-floor, Windows, WSL, paired-runtime, mixed-version, and SSH results.
- Captured staging logs, metrics, traces, and support bundles with seeded private canaries absent.
- Load-test results and independently tested upload, download, and remote-install kill switches.
Approval owners record findings and accepted residual risks outside this implementation document.
The external rollout gate remains closed until those findings are resolved or explicitly accepted.
@@ -1,43 +0,0 @@
# Agent skill sharing upstream boundary
Status: proposed for formal engineering and legal review.
Date: 2026-08-11.
## Context
Orca needs private, durable, cross-machine skill sharing with bounded archive ingestion,
host-owned destination resolution, crash-safe transactions, provenance, and mixed-version remote
support. `vercel-labs/skills` exposes a CLI and does not provide the Cloud authorization or local
transaction contract Orca requires.
The behavioral assessment used upstream commit
`c6f69c631292444cc541ac6d91e2226b0ff247da`.
## Decision
Orca implements its package, Cloud, installation, provider-placement, and recovery behavior
independently. The upstream project is a behavioral reference only.
Do not copy or mechanically translate upstream source, tests, fixtures, registry entries, provider
path tables, comments, or documentation. Derive provider paths from official provider
documentation and verify them with real installations. Orca does not depend on the upstream CLI,
npm package, or an unsupported programmatic API.
If a future change proposes incorporating upstream material, stop and review the exact material,
license, attribution, notices, and maintenance implications before implementation or merge.
## Consequences
- Orca owns stability, security, compatibility, and maintenance of this narrower installer.
- There is no automatic upstream synchronization job.
- Similar behavior is acceptable when independently derived from requirements and official
provider contracts; textual or structural copying is not.
- Provider registry changes require normal code review plus official-documentation and real-host
evidence.
- The pull request template requires reviewers to confirm this boundary for relevant changes.
## Review record
Product chose the reference-only approach during planning. Formal engineering and legal reviewers
remain to be named before external rollout.
-577
View File
@@ -1,577 +0,0 @@
# Agent status store
## Status
The current boundary is PR 2A: structured sessions use the hook server's fully
scoped canonical store; unbound PTY/relay evidence remains in an isolated legacy
adapter. Do not remove the renderer bridge or its publication filters in this
slice: they still carry native-chat child rows.
The sections below record the original 2026-09-09 rollout. Its PR 1a and PR 1b
have landed; its proposed PR 2/3 sequence is superseded by that boundary:
1. main-only: every producer writes into one store and `worktree ps` reads it,
split into 1a (structured sessions join the store) and 1b (the runtime's
duplicate retained store is deleted);
2. renderer: the sidebar becomes a subscriber and stops re-deriving rows;
3. shared: one worktree-status rollup and one freshness rule for every reader.
## The problem this solves
Orca shows "what is this agent doing" in four places: the desktop sidebar, the
`orca worktree ps` command, the mobile app, and the agent dashboard. Before
#19217 those readers did not even share their inputs. After #19217 they share
the structured-session mapping and nothing else.
An audit on 2026-09-09 found six producers and three consumers, and three
separate copies of the same row inside the main process alone:
| Main-process copy | Keyed by | Owned by | Persisted | Evicted |
| --------------------------------- | --------- | --------------------------------------------------------------------------------- | ------------------ | ---------------------------------------------- |
| hook server `lastStatusByPaneKey` | paneKey | `src/main/agent-hooks/server.ts` | `last-status.json` | tab close, pty exit, hydrate, worktree removal |
| runtime `RuntimeAgentRowStore` | paneKey | `runtime-agent-row-store.ts` (deleted in PR 1b) | no | pty exit only |
| structured feed `published` | sessionId | `src/main/native-chat/agent-session-wire/structured-agent-session-status-feed.ts` | no | never (a broadcast cache) |
The second copy is a duplicate write: the OSC status parsed in main is
forwarded to the hook server _and_ retained in the runtime store from the same
call (`orca-runtime-create-terminal-side-effect-command-code-detector.ts`).
The third copy is keyed differently and never reaches the hook server at all,
which is why `worktree ps` grew its own adapter for it in #19217.
Each reader then applies its own precedence and freshness rules, so the same
pane can legitimately read differently on the desktop, on the phone, and in
the CLI.
## The rule
**The execution host owns agent status, in one store, and every reader
subscribes to it.** This follows the boundary in
[`ssh-execution-boundary.md`](./ssh-execution-boundary.md): the host that runs
the process is the only party that can observe it, and the client is never
authoritative for execution state.
Three consequences:
- One store per execution host. A remote host keeps its own store and the
client mirrors it down, as the web-session mirror already does. Mirroring is
not merging: a client never writes its observations back to a host.
- Precedence is decided once, at write time, with provenance recorded on the
row. Readers never re-adjudicate hook versus terminal versus structured.
- Readers keep only presentation policy and user facts: the 30-minute display
decay, acknowledgements, dismissals, unread. Those stay reader-side but
become one shared implementation (PR 3).
## The store already exists
The hook server's state is that store today for every PTY-based agent. The
audit established:
- hook HTTP posts, the WSL and SSH relay receivers, and main's own OSC parse
all converge on the same `applyNormalizedStatus` path, stamped with the
authority id `main-agent-hooks`;
- it alone holds pane authority: launch tokens and their hashed commitments,
retired-pane fences, pane-key aliases, per-connection ordering watermarks,
and the evidence-age map that must outlive a transport clear;
- it alone persists, with a seven-day hydrate window and the
`restoredUnconfirmed` stamp that keeps a hydrated row from ever reading as
live truth;
- it already fans out to both renderer windows over `agentStatus:set` and
`agentStatus:clear`, and serves `agentStatus:getSnapshot`.
Nothing else in main carries those guarantees, and building a second store
with them would be the wrong direction. So the design is not "add a store". It
is: **route the two producers that bypass the hook server through it, then
delete the copies.**
## PR 1a: structured sessions publish into the store
No renderer behavior changes. The sidebar keeps receiving the same IPC events
it receives today, plus structured-session rows it currently derives itself.
### Structured sessions publish into the hook server
The structured feed keeps its job of projecting a session's journal into a
summary and streaming it to subscribers. On every publish it additionally
ingests the summary into the hook server as a status row:
| Row field | From |
| --------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `paneKey` | `structuredAgentSessionPaneKey(tabId, sessionId)`, the key the renderer already uses; its leaf is UUID-shaped so pane-key validation accepts it |
| `tabId` | `structuredAgentSessionTabId(sessionId)` |
| `worktreeId` | `summary.workspaceId` (a folder workspace id is a valid value) |
| `state` | `structuredAgentSessionAgentStatus(summary).state`: the lead's own status folded with the live child records the store holds for the session (not the summary's task list), so a settled lead whose subagent still runs reads `working` |
| `workingMode` | `'monitoring'` from the same fold when watch loops are the only live child work; omitted otherwise, which clears it on the row |
| `mainAgent` | the main agent's own state before the fold, its last-turn verdict (`summary.turnOutcome`, present only while idle) and its own clock; see "The main agent fact" below |
| `structuredHost` | `'owned'` while `summary.hostExecutionOwned` is set, otherwise `'held'`; `worktree ps` derives its row's `structuredHostOwned` from it |
| prompt, tool, last message, model, provider session | the summary's fields |
Sessions with no request (`status === null`) produce no row. A request is a
turn record, an assistant message, a user message the provider journaled itself
(history, an older host), an accepted or unanswered send, or a send the agent or
its start refused; a send that was withdrawn, or left undelivered by a
restart or a close, fails nobody and makes nothing listable.
`summary.turnOutcome` is the latest request's verdict: its turn's outcome, or
`failure` for a send the agent or its start refused (a send that joined a running
turn is answered by that turn). A turn the provider gave no outcome reads as its
host-observed end through `agentTurnVerdict`: `interruption` for an `interrupted`
lifecycle, `unconfirmed` for an `unverifiable` one. It is derived on each read,
never journaled. The row also publishes `interrupted` from
`mainAgent.outcome`, exactly as the hook lanes do. When the host revokes live ownership the row is re-set
without the flag; when the host closes or evicts the session the row is
dropped. Both already exist as feed events (`revokeLive` and the roster
filter in `liveSessionSummaries`); PR 1 turns them into store writes.
Dropping the session from the host's map and dropping its row are one
operation, `forgetStructuredAgentSession`. The store keeps a row until told,
and a host-owned row bypasses the staleness check, so a deletion path that
forgot the row would strand a permanently working-looking agent.
Two rules the ingest must keep:
- **Never persist a structured row.** The journal is the durable truth for a
structured session and the host republishes on restore. A structured row in
`last-status.json` would hydrate as `restoredUnconfirmed` and then fight the
live republish. The serializer skips rows carrying `structuredHost`, and
hydrate drops any such row found on disk. Applying one therefore also skips
the persist schedule: the walk and stringify could only reproduce the file
that is already on disk, once per debounce window for every streaming chat.
- **Never let it fight a hook row.** A structured session has no PTY, so no
hook or OSC event carries its pane key. The ingest still goes through the
disposition gate so a retired pane key is refused like any other.
Applying one does still run both status fan-outs, and that is intended rather
than incidental. `notifyStatusChangeListeners` is what feeds
`agentAwakeService`'s power-save blocker, and `subscribeEnrichedStatus` is what
feeds `AgentSessionTransitionRecorder`'s stats, so joining the store enrolls
native chats in both. A working native chat is real work and should hold the
machine awake exactly like a PTY agent does.
The drop side routes through `dropStatusEntry`, not `clearPaneState`: a
pane-status-clear reaches the renderer, and until PR 2 the renderer's own feed
bridge is that pane key's writer. It also passes `preserveResumeIdentity:
false` — the `providerSessionOnly` remnant a dismissed pane keeps exists so the
agent can be resumed in that pane, and a structured session has no pane and
keeps its resume identity in the record store. Like every other
`dropStatusEntry` caller, it emits no pane clear, so a session dropped
mid-`working` leaves `AgentSessionTransitionRecorder` holding an open stats
session until its LRU evicts it; that gap is shared with the user-dismissal
path and is not specific to structured rows.
The ingest lives in the feed, not in `structured-agent-session-host.ts`, which
sits at the file-length cap.
### `worktree ps` becomes a reader
The structured adapter added in #19217 is deleted, and structured rows reach
`worktree ps` through the same snapshot as every other row. The
retained-versus-hook reconciliation in `collectRuntimeWorktreePtyAgentSources`
stayed until PR 1b removed the store that fed it. What this step settles is
the admission gate that decides which rows a worktree listing may show:
- a hook or OSC row needs its tab mirrored or a connected pty, as today, and
SSH rows stay exempt because their tabs may exist only remotely;
- a row carrying `structuredHost` is admitted while the host holds the session, and
the host's drop on close is what removes it. No tab-mirror requirement: a
structured session's tab lives in the renderer's own tab state, and a
headless host has no renderer to mirror it from. That argument only holds if
the headless host is itself wired to the store, which is a separate
obligation per entry point: the Electron hosts (desktop and `orca serve`)
share `main-process-runtime-service.ts`, and `orcad` constructs its own
runtime in `src/main/orcad/orcad-entry.ts`. A host missing that wiring lists
no agents at all, not just no structured ones, because `worktree ps` reads
the same snapshot for every row.
The freshness bypass for host-owned structured rows already exists in
`isFreshNonDoneAgentStatus`; with the flag now on the row it becomes the only
path, and the hand-rolled check in `runtime-worktree-agent-rows.ts` goes.
### Wire compatibility
`AgentStatusIpcPayload` gains one optional field, `structuredHost`, and the
`worktree ps` row gains `structuredHostOwned`. Under rule 1 of
[`remote-wire-compatibility.md`](./remote-wire-compatibility.md) both are safe:
an old client ignores them. `worktree ps` rows keep their shape and vocabulary,
so the mobile app sees no change.
Until PR 2 the main process does not forward structured rows to the renderer
over `agentStatus:set` or `agentStatus:getSnapshot`. The renderer's feed
bridge still writes those rows itself, and forwarding them too would give one
pane key two writers. Removing that filter is the first step of PR 2.
### The main agent fact
Claude, Codex and Grok hook rows and structured-session rows publish the combined
`state` and, beside it, the main agent's own state as `payload.mainAgent`. Other agents'
rows and terminal-title-only rows carry none, and readers fall back to `state`:
```ts
mainAgent?: { state: AgentStatusState; outcome?: AgentTurnOutcome; stateStartedAt: number }
```
`state` still answers "what should the user see" and folds live child work in,
so a settled main agent whose subagent still runs reads `working`. `mainAgent` answers
"what is the main agent itself doing", which the fold used to destroy at publish
time; every guard that reconstructed a fragment of it (`fromChildWork`, the
persisted `claudeLeadBoundaryChildOnly` flag) now reads `mainAgent` instead of a
stored copy. A Claude row whose `mainAgent` is `done` while a child agent still
works (including a child's permission wait) refuses OSC, which carries no child
identity; the children's own lifecycle hooks settle it. `outcome` is the recorded verdict on
the main agent's most recent finished turn, present only while `mainAgent.state` is
`done`. It is reported by the provider, or is a `cancellation` Orca inferred
from the user's own interrupt keystroke, or a `superseded` the host recorded when
a newer request replaced a structured Claude turn before it ended (it names no
sender, and sets no legacy flag), or, on a structured row whose turn the
provider gave no verdict, is what the host observed of its end: `interruption`
(a proven death nobody asked for) or `unconfirmed` (an end it cannot prove,
never success). The journal's turn outcome, by contrast, stores only recorded verdicts: the provider's, a `cancellation`, or the host's `superseded`; `interruption` and `unconfirmed` are derived from the turn's lifecycle state and never stored. A plain end of turn carries none, because absent
means unknown and a provider that omits its interrupt flag must not turn a
cancel into a success.
In the Claude hook lane the cancellation comes primarily from Orca's own
inferred interrupt (`markClaudeLeadTurnInterrupted`), because current Claude
sends no hook at all on a cancel and no `is_interrupt` on Stop; that flag on a
turn boundary remains a secondary source for builds that send it, and
`StopFailure` maps to `failure`.
Readers decode the verdict through one accessor, `agentMainAgentVerdict`, which
reads the main agent's own state, not the combined row's: `mainAgent.outcome`
while `mainAgent.state` is `done`, then the legacy `interrupted` flag as a
cancellation, which alone needs the combined `done`. So a main agent that
failed while its subagents still run has a verdict on a `working` row. Every
copy of a row (state-history entries, sleep records, `worktree ps` rows) takes
the verdict through `agentVerdictFields`, which carries `interrupted` and the
whole `mainAgent` (state, outcome and its own clock) together, so a copy agrees
with the row and can date a failure by `mainAgent.stateStartedAt`.
Display reads the verdict through `agentVerdictDisplayMark`. A fault marks the
agent failed whatever the combined state, because it is news the user must see
even while subagents run: a `failure`, and an `interruption`, a turn cut short
by anything other than the user or a newer request. A user's stop (`cancellation`)
marks it interrupted, drawn in the muted tone with the row text "Interrupted by user";
a turn a newer request replaced (`superseded`) marks it interrupted in the same muted
tone with the row text "Interrupted"; and `unconfirmed` marks it unconfirmed, all only
on a `done` row, so a stopped
or finished main agent with live child work still reads working. The folded
turn header follows the same mark: "Failed after N", "Interrupted after N", or
"Worked for N".
Each subagent keeps its own row and state. Container rollups (worktree card,
terminal tab, Cmd+J) rank a pending question first, then a failure, then live
work, then an unconfirmed end, then a user's stop, then done. On the worktree
card, a failure retained after its agent's pane went away has no expiry, so it
ranks below live work and above an unconfirmed end. Lifecycle waiters keep
reading the combined `state`.
Policy splits the verdict two ways. Clean-finish policy (hibernation, pane
ownership, the star-nag value moment) treats a failure, an interruption and an
unconfirmed end like a cancellation (`agentTurnEndedUncleanly`). Attention
(completion time, Smart Sort, sticky retention, Cmd+J Recent) demotes only a
turn ended on purpose, the user's stop or a newer request that replaced it
(`agentTurnEndedOnPurpose`); a failure, an interruption or
an unconfirmed end ranks like a completion.
Admission is one function, `normalizeAgentStatusPayload`, on the relay wire,
IPC and disk. A malformed `mainAgent` drops the field and keeps the row. Old hosts
send none and readers fall back to `state`. Hook rows persist it inside the
payload; hydration maps an older row's `claudeLeadBoundaryChildOnly: true`
onto `mainAgent: { state: 'done' }` when the row has no `mainAgent`, and never writes the
flag again. Hydration seeds the Claude main agent record straight from a saved
`mainAgent` that is `done`, so the children's drain can still settle the row after
a restart. `claudeRunningNonAgentTask` is persisted alongside because it is the one
child-work fact `mainAgent` cannot express: a shell running beside the main agent,
whose liveness hydration does not restore. Hydration seeds only a row that says
`false`; a row silent about it stays unseeded. The row builder pairs the two facts in
one place: a listener event restates the shell fact, and any other write (an OSC
repaint, an inferred answer) keeps it only while `mainAgent` is unchanged. A child's
sticky permission prompt still records the main agent's own progress and background
evidence in the held row, and pushes the held row to subscribers when `mainAgent` changes.
Every lane, Codex included, combines through the fold. A child waiting on a
human is a fold input (`childWorkLiveness: 'waiting'`, derived from the child's
own `waiting` state; a child's `blocked` means it failed and stays live work)
and makes the row wait whatever the main agent is doing, unless the main agent
is itself asking. The Codex hook lane feeds it from its child transcripts, and
the structured lanes from child records, which read `waiting` for a Codex child
thread's approval or input flag and for a Claude subagent's open permission
request. Known
divergences, pinned by name in the parity table
(`src/shared/main-agent-status-parity.test.ts`) where they are reachable, so a
reader does not mistake them for drift:
- The Claude hook lane holds a child's permission wait in one slot on the
displaced main agent record (`waitingAgentId`, `stateBeforeWait`), not on
the child. It publishes the displaced state as `mainAgent`, but the next
main agent event overwrites the slot, so the row stops reading `waiting`
while the child is still asking, and a second asking child replaces the
first.
- In the structured lane a child's pending prompt also makes the session
`attention`, which reads as the main agent's own `blocked`: one needs-input
state whoever asked. A Claude subagent reads `waiting` only while the
journal holds its card pending: from after the card's row is written until
just before anyone closes it, so every publish that shows the child waiting
also shows the session's `attention`, and the row never reads `waiting` for
a Claude subagent's request.
- The Codex hook lane drops its roster on a root `Stop` when it tracks no
child transcripts, so a still-running or still-asking child stops holding
the row.
How the main agent's turn ended is not a fold input. A cancel is a verdict on
the main agent, carried as `mainAgent.outcome: 'cancellation'` (and, for
readers that predate `mainAgent`, as the row's `interrupted` flag on a `done`
row); it never retires a shell, scheduled check or subagent the turn left
running. That work leaves the row only when its own inventory omits it or the
session ends, so a cancelled turn with a still-running shell reads
`monitoring` in every lane, and the parity table in
`src/shared/main-agent-status-parity.test.ts` drives that story through all of
them. The same rule governs the cancel Orca infers from Ctrl+C: for any row
that publishes `mainAgent`, the inference is admitted only when
`mainAgent.state` is `working`, so Orca does not treat a Ctrl+C at the idle
prompt of a row held open by child work as a turn cancel (Codex also keeps the
child-evidence guard, and a row without `mainAgent` keeps only that guard).
The keypress itself is not inert, though: measured live, Claude 2.1.280 stops
its background subagents on a single idle-prompt Ctrl+C (shells survive) and
Codex 0.156.1 quits outright, so refusing the inference can leave the row
showing a subagent its CLI already stopped. The synthesized row is the fold
of the cancelled main agent with the child work the pane's owner can see: the
local listener's roster for a local pane, the row's own subagents and shell fact
for a relayed one, whose provider records live on the relay.
The store holds that verdict against restatements that predate it
(`server-cancel-verdict-latch.ts`), because a relay never learns of a cancel
the desktop infers and some TUIs emit late same-turn hooks. The hold is read
off the row (`mainAgent.outcome: 'cancellation'`), never stored beside it, and
dies on a new turn (a main agent prompt submission, a changed or explicit
prompt, a session start) or the provider's own settled `mainAgent`. Child and
replayed events under the hold keep the cancelled main agent and are re-folded
with their own child evidence.
## PR 1b: the runtime's retained row store is deleted
Landed. `RuntimeAgentRowStore` is gone, and with it the retained-versus-hook
reconciliation in `collectRuntimeWorktreePtyAgentSources`. The hook server's
store is now the only main-process copy of a PTY agent's row.
### The five call sites
| Call site | Before | After |
| ------------------------------------------------------------------------------ | --------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| `orca-runtime-create-terminal-side-effect-command-code-detector.ts` `retain()` | second write of the OSC payload already sent to the hook server | deleted; the event now carries the pane's `terminalHandle` and the hook ingest keeps the only copy |
| `...command-code-detector.ts` `clearPty()` | drops rows on pty exit | deleted; pane teardown already clears the hook row |
| `orca-runtime-get-worktree-ps.ts` `values()` | fed `retainedSnapshots` | deleted; the reader keeps only `hookSnapshots` |
| `orca-runtime-serialize-agent-prompt-submission.ts` `getFreshExplicit()` | retained row first, hook rows second | `selectFreshExplicitAgentStatus`, hook rows only |
| `orca-runtime-prune-mobile-session-tab-group-layout.ts` `getFreshForMobile()` | pane key, then pty id | `selectFreshAgentRowForMobileTab`: pane key, then `terminalHandle` |
Both readers moved into `runtime-hook-agent-row-selection.ts`, which also owns
`RuntimeAgentRowSnapshot` now that nothing retains one.
### `terminalHandle` is the row's join back to its terminal
The retained store's only real extra was the pty id, and two readers used it.
The plan said to stamp the event's `ptyId` into `terminalHandle`; that was
wrong. A terminal handle (`term_<uuid>`) and a pty id are different
identifiers, and `getFreshExplicit` was already comparing hook rows against a
real handle. What landed instead:
- `AgentHookEventPayload` and the runtime's terminal-status event gained an
optional `terminalHandle`. The detector resolves it once per chunk through
`getAgentStatusTerminalHandleForPaneKey` — the same lookup the renderer-facing
IPC boundary already runs for every row, so the two surfaces cannot disagree
about which terminal a pane is.
- `applyNormalizedStatus` carries the handle forward when an incoming event
resolves none. Only main's OSC parse can resolve one, so an HTTP hook post for
the same pane would otherwise erase it.
- It is never persisted. A handle belongs to the runtime that issued it, and a
hydrated one could only rejoin a row to somebody else's terminal.
- `toAgentStatusIpcPayload` publishes it, which also makes `getFreshExplicit`'s
long-dead handle comparison live: the runtime reads raw snapshot rows, and
before this nothing ever stamped the field on them.
`worktree ps` uses it too. `ConnectedPtyEvidence` traded its flat `ptyIds` set
for `ptyIdByTerminalHandle`, so a row still resolves the connected PTY behind
it — which is both the working-terminal rollup's match key and the last rescue
for a row whose pane binding was nulled by a controller incarnation change.
### The change detector had to move with the store
`retain()` was not only a store: its boolean return was the signal that
republished `session.tabs` for a status-only transition, which no title change
covers (#7970). `hook-status-session-tabs-invalidation.ts` already mirrors that
projection change set, including restore provenance and terminal-handle joins,
so the replacement was to route the signal off the store rather than build a
second comparator.
`installHookStatusSessionTabsRepublish` now owns all three arms — enriched
status, pane clear, and the status-drop tap a dismissal emits — and both hosts
install it.
### Both hosts, not just the desktop one
`orcad` constructed its runtime with no `onTerminalAgentStatus`, so main's OSC
parse never reached the store there and the retained copy was the only carrier.
Deleting it without wiring orcad would have made a headless host list no PTY
agents at all. `orcad-entry.ts` now binds the producer and installs the
republish signal, alongside the snapshot and structured sink it already had.
### The intended behavior change
A row the user dismisses on the desktop leaves `worktree ps` and the phone at
once, instead of lingering until the pty exits. One store means one dismissal.
Legacy numeric pane keys remain a bounded compatibility case. Persisted layouts
register aliases to their stable leaf owners; an in-process OSC observation may
also retain a numeric key only when the runtime supplies the matching tab, PTY,
and terminal handle. HTTP and relay ingress still require a stable key or a
registered alias, and numeric rows are never persisted.
## PR 2: the renderer subscribes
With structured rows arriving over `agentStatus:set`, the renderer's
`StructuredAgentSessionStatusBridge` no longer needs to write status; its
unmount cleanup becomes a tab-close signal to the host. The IPC applicator is
the single writer for observed status. The 2026-09-09 audit sorted the other
writers:
| Writer | Disposition |
| ----------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
| Command Code output seeds, parked-pane seeds, pty-exit removal | delete; main already emits the same facts |
| structured bridge status writes | delete; main now publishes the row |
| structured bridge failed-start row (the host refused the create) | keep; a refused create leaves the host no session, so the bridge writes it from the launch record |
| launch placeholder seeds (a user launched an agent with a prompt) | keep for now; main holds the launch config and can seed later |
| dismissal, acknowledgement, unmount | keep; user facts and component lifecycle |
| remote-runtime OSC parse (bytes never transit local main) | keep, fenced behind the host's published row once the host is new enough; rule 3 of the wire doc applies |
| web-session mirror receipt clock | keep; the decay rule needs both clocks from one machine |
The Command Code done-settle window is renderer policy with no main
equivalent. PR 2 either moves it into main's detector or leaves it, and says
which.
## PR 3: one rollup, one clock
The worktree card status is derived three times: `lib/worktree-status.ts` in
the renderer, `runtime-worktree-status-projection.ts` in main, and
`agent-row-display.ts` in mobile, which hand-copies the 30-minute constant.
PR 3 moves the rollup and the decay into `src/shared` and makes all three
call it.
## Readiness reads the store
`terminal wait --for tui-idle` is a reader too. Before STA-9100 hook state
reached it only through the `<Agent> ready` titles the window writes, so a
headless `orca serve` never saw it (#16095). Now an agent whose rule file says
`profile.hooks: "authoritative"` (OpenCode, OpenCode 2, Pi, OMP) or
`"turn-end"` (Codex) has its fresh row read straight from the store, through the same
`selectFreshExplicitAgentStatusRow` join prompt-receipt verification uses
(`src/main/runtime/tui-idle-hook-lane.ts`):
- the main agent's turn, not the combined row, decides: `mainAgent.state` when
published, so a subagent's Stop does not end the lead turn. `done` settles
the wait, `working` holds it, and a permission wait never settles. The tail's
blocked text goes through the existing permission arbiter with the turn as
its explicit status, so a denied prompt's dialog left in the tail no longer
blocks a turn the hook says ended;
- the row joins on any pane key or terminal handle the PTY owns; a pane neither
reaches, a stale or restored row, a session-start `done`, a row from before
the PTY respawned, and a `done` received before the pane's latest input all
leave the decision to the screen and text rules, which is also how startup
readiness works before an agent's first hook. The input is the PTY run's
`lastInputAt` (`terminal-run-facts.ts`), which both write funnels record, so
a key the user typed counts like a prompt Orca sent: the next turn's first
hook may still be in flight, and an agent restarted in the same shell has
not posted one. A shell command marker is no process boundary: Pi paints
OSC 133 zones itself;
- every other agent stays `identity-only`: Claude sends no event when an
approval is denied or Esc stops a tool, so its row can sit at `waiting` or
`working` forever, and the rules keep deciding;
- Codex is `turn-end`: only a `done` decides (it settles the wait), and a
`working` or permission row leaves the decision to the rules. Before its
`Interrupt` hook an Esc mid-turn can leave the row `working`, and an older
TUI can hand its hooks to a newer shared app server, so no version check
tells which Codex posts it. A `done` is a real turn end on every version, so
trusting only that one keeps the headless gain (no quiet window after the
turn) without letting a missing cancel hang the wait.
The titles stay for display; remote clients read them.
## What does not change
- The hook scripts, the OSC 9999 wire format, and the relay protocol.
- The status vocabulary. `working / blocked / done` for rows,
`working / attention / idle` for structured summaries, mapped once.
- The `live / unverifiable / exited` verdicts for remote work. Loss of contact
clears nothing; the SSH exemptions in the admission gate stay.
- Hydration honesty: a restored non-done row is `restoredUnconfirmed` and is
never fresh.
## PR 1b reliability contract
- **Invariant (`agent-session.status-host-ownership`):** each execution host has
one agent-status store; OSC, hooks, and structured sessions write it, while
desktop, `worktree ps`, and mobile only project it. Dismissal, certified PTY
exit, and provider-generation replacement remove the same row everywhere;
transport loss alone removes nothing.
- **Failure source:** the deleted runtime row store duplicated OSC observations,
keyed them by a different terminal identity, and outlived a dismissal from the
hook store. Relay replay could also make old evidence look fresh when readers
used its new delivery timestamp.
- **Oracle:** one OSC observation appears through the hook snapshot in
`worktree ps` and mobile, and one store dismissal removes it from both without
stopping the PTY. Focused tests also require leaf/incarnation-handle rejoin,
legacy numeric-pane compatibility, certified-exit and provider-generation
cleanup, evidence-age freshness, and exactly-once startup/stop teardown.
- **Gate:** `terminal-performance.osc-status-scan-budget` covers the unchanged
bounded OSC parser and the runtime projection. There is not yet a dedicated
blocking multi-surface status-store gate; the focused suites below are the
accepted gap until they accumulate reliability-gate soak evidence.
- **Provider/platform coverage:** local and daemon-backed PTYs are covered by
runtime tests, and SSH relay loss/replay semantics by relay integration tests.
The projection is shared by git worktrees and folder workspaces. WSL uses the
same store and admission code but has no live run here; Linux and Windows
runtime execution, native mobile clients, and mixed-version paired clients
remain validation gaps.
- **Performance budget:** publication stays event-driven with no new polling or
subprocesses. One mobile projection clones the status snapshot once, builds
pane/handle indexes once, and has a deterministic call-count test; lifecycle
cleanup is bounded by the existing status and handle inventories, and orcad
tests prove listeners clean up once on failed startup and repeated stop.
- **Diagnostics:** existing hook-listener errors name the pane and PTY, while
status-store tests pin delivery versus evidence clocks. No new telemetry or
raw terminal data is emitted.
- **Residual gaps:** rendered Electron/mobile behavior, live SSH reconnect, and
Linux/Windows/WSL execution require the platform QA pass. The current
cross-version gate does not cover `session.tabs` content.
## Verification
- Unit: ingest a structured summary and read it back through
`getStatusSnapshot`, `worktree ps`, and the mobile projection; assert the
serializer never writes a row carrying `structuredHost`; assert a hydrated
file that somehow contains one is dropped.
- Unit: the `worktree ps` suites written against the retained store are rewired
to a real `AgentHookServer` (`agent-status-store-wiring.test-fixture.ts`)
rather than deleted, so each still asserts the listing behavior it named. The
dismissal change is pinned end to end in
`orca-runtime-tests/worktree-ps-agent-row-dismissal.spec.ts`, which fails with
the retained store restored.
- Live: the parity check from #19217 (working, done, close, reload) repeated
against the merged store, with both surfaces read from the one row.
## Retired OMP pane recovery
A desktop renderer retirement carries an optional UUID through the existing
`agentStatus:retirePaneAuthority` IPC message. The hook server retains it with
its bounded retirement fence. A validated live OMP new turn consumes that UUID
and echoes `authorityRestartId` only in the live notification. Cached rows,
persistence and startup replay never carry the acknowledgement. Older peers
omit or ignore it and retain explicit attach restoration.
The renderer keeps the UUID in its existing non-persisted retirement tombstone;
every re-retirement mints a new one. A matching acknowledgement may clear that
tombstone only with a successful status write for the existing pane and matching
workspace/connection. Closed tombstones remain `true`, including after the tab
LRU evicts its entry. Closing a retired physical alias revokes its whole group.
This is control-plane retirement correlation, not a second agent-status store.
Fallback restores the hook server's recorded status aliases through the existing
attach-restoration path. The accepted renderer write restores the matching status
alias routes too, preserving group membership for the next retirement. It does
not restore orchestration or launch credentials.
It is scoped to the requesting desktop renderer. A different window's retirement
UUID cannot be cleared by the acknowledgement, and web mirrors keep their existing
host-snapshot/attach behavior.
@@ -1,90 +0,0 @@
# Native Antigravity Accounts
Accounts reads the credential authority on the runtime that owns execution. A client chooses
an owning Orca runtime and a host/distro target before sending an operation; it never replaces
the client's Mac Keychain item for another host. The RPC capability is
`accounts.antigravity-native.v1`. Older paired hosts are refused before account mutations.
The RPC returns account summaries only, never credential JSON, access tokens or refresh tokens.
Displayed quota is tied to the subject and authentication method observed during its refresh;
an external identity change hides the previous account's quota without an automatic fetch.
## Supported authority
Normal macOS agy uses service `gemini`, account `antigravity`. Its go-keyring values use the
base64 or legacy hex wrapper. Orca passes writes through `security -i` stdin, validates bounded
output and reads the entire native value back. The command buffer limit is checked before
writing. A missing native item falls back to the CLI-specific
`~/.gemini/antigravity-cli/antigravity-oauth-token` file. The distinct legacy jetski fallback
is not imported.
The compiled CLI bypasses keyring storage when SSH/WSL environment detectors or WSL kernel
identity apply. A runtime running under that evidenced bypass reads/writes its own CLI file;
it does not contact the client keychain. The file must be private and regular. A macOS
`cache/antigravity-keyring-unavailable` marker makes authority uncertain: Orca refuses instead
of assuming that the keychain or file wins.
Native Windows Credential Manager, native Linux Secret Service, and operations directed from
Windows Orca to a selected WSL distro are explicitly unsupported pending verified adapters.
Windows file bypass is also refused until private ACL protection is verified.
Windows' `gemini:antigravity` raw blob and 2560-byte limit are different from the Mac wrapper;
Linux uses the login collection with `service=gemini`, `username=antigravity`. No dependency,
PowerShell compilation, credential-home flag, or cross-host fallback is invented here.
A separate SSH relay has no Accounts RPC; use a paired owning runtime that implements it.
## Identity and snapshots
A Google ID token supplies the normalized Google issuer and stable subject. The authentication
method also scopes identity. The label uses a verified email when available; email is never the
identity key. Account record IDs are random and survive token, expiry, refresh-token and email
rotation. Profiles without a stable subject can be displayed but cannot be saved for switching.
Snapshots preserve the exact native JSON, including fields that Orca does not interpret. The
host's vault under `userData/antigravity-accounts/vault` requires meaningful OS encryption and
private permissions. Weak or unavailable encryption is refused. Unreadable/corrupt ciphertext
is preserved; it is never treated as an empty vault. This does not migrate the experimental
candidate's incompatible array vault or token-hash IDs.
One host service serializes Add, Select, Remove, launch checks and refresh reconciliation.
It re-reads the vault after asynchronous native reads and captures external CLI refreshes into
the same stable account. Selection reconciles the outgoing snapshot, checks the expected native
bytes before writing, and checks native readback before publishing the selected ID. It avoids
writing an old snapshot over an already-active account. The current or selected account cannot
be removed; deletion checks the latest native value again before committing.
A selected account is checked before new Orca PTY launches, including desktop daemon and
headless runtime paths. An externally changed native identity blocks the launch and asks the
user to select again. Existing sessions can retain their original credentials in memory.
Shell commands typed manually into a running terminal are outside the Orca launch guard.
## Sign-in and concurrency limits
Sign-in uses the supported ordinary agy browser/code flow. Users run agy on the owning host;
to add a different account they use its `/logout` command, complete the next sign-in, then save
the actual resulting account in Orca. This implementation does not advertise an Orca-managed
login or invent an agy `login`/`--login` flag. Browser completion and a second real Google
account remain user-driven; tests do not sign out or change the developer's real native item.
Native keyring does not expose compare-and-swap. Orca's queue serializes its own calls, and
bounded before/after checks detect observed conflicts; another independently running agy or
Orca process can still write between the final check and the write or launch. A failed
verification may mean the native item changed but selection was not persisted. Refresh and
explicit selection resolve that state; automatic rollback could destroy a newer CLI refresh
and is deliberately avoided. The file backend has the same external-writer limit.
## Evidence and contributor credit
The foundation adapts the reviewed codec/macOS adapter from #21784 and account-service concepts
from #21797 (nwparker), with fresh identity, persistence, serialization and conflict handling.
The signed-in Accounts card and quota-error visibility acknowledge #19588 by @artile; quota
transport is reused from current main rather than its obsolete extraction code. Targeted
multi-account UI/target concepts acknowledge #23761 by @Tai-DT, replacing its placeholder login
and unused settings selection. The Accounts legacy-Gemini clarification acknowledges #21682
and the original relevant migration contribution by @siddqamar, as requested in #17345.
No stale development stack was cherry-picked.
Live proof uses a disposable Mac service/account item, a fully isolated hidden Electron home,
and synthetic accounts. A private task-only copy was also selected through the real service;
installed agy 1.2.14 consumed that verified file credential under its SSH bypass and returned
`command.name=usage`, `num_turns=0`, no conversation. The real native item remained unchanged.
This proves the Mac adapter mechanics and actual CLI file authority, not a second-account
native-keychain switch, native Windows/Linux switching, or WSL/SSH relay deployment.
@@ -1,332 +0,0 @@
# Antigravity readiness: what the transcripts show
Antigravity readiness lives in `src/main/runtime/agent-state-rules/antigravity.json`: a screen rule
over the trusted grid, and a text anchor that runs the named scan
`findAntigravityComposerIndex` (`agent-state-rules/antigravity-text-composer.ts`) over the
line-folded tail when no trusted grid exists. That text scan decides whether a pane is ready for a
prompt from its tail alone. It has been written five times, each version tuned
against a five-line screen typed from memory into a `.spec.ts` fixture. Three of the first four
were found worse than the bug they replaced, and the fifth was reverted.
Real transcripts now exist. They were recorded from a live `agy` on macOS with
[`agent-pty-transcript-capture.md`](./agent-pty-transcript-capture.md) and are committed under
`src/main/runtime/__fixtures__/`. `src/main/runtime/antigravity-readiness-transcripts.test.ts`
replays them through the runtime.
**Headline (STA-8741, agy 1.2.14): readiness is now read off the live screen, not the text tail.**
The screen's bottom four rows at an idle composer are rule, caret, rule, `? for shortcuts`. The
caret alone is not enough: agy keeps it painted mid-turn and behind the `/model` picker. The hint
row is what changes. [Attempt seven](#attempt-seven-the-live-screen-sta-8741) has the recordings;
the sections after it are the 1.2.0 history that led there.
## Attempt seven: the live screen (STA-8741)
Recorded 2026-09-30 on macOS with `agy` 1.2.14 (binary and banner agree), a signed-in Google
account, `AGY_CLI_HIDE_ACCOUNT_INFO=1`, in an already-trusted workspace. All are 120x40 except
where the name says otherwise. `src/main/runtime/antigravity-screen-readiness-transcripts.test.ts`
replays them.
| Fixture (`antigravity-1-2-14-*.txt`) | Screen at the end | Screen rule |
| ------------------------------------ | --------------------------------------------------- | ----------- |
| `ready` | settled startup, bare `>` | ready |
| `ready-accept-edits` | `> Accept-edits mode: …` (`--mode accept-edits`) | ready |
| `ready-plan` | `> Plan mode: …` (`--mode plan`) | ready |
| `ready-80x24` | settled startup on an 80x24 PTY | ready |
| `picker-dismissed` | `/model` opened, then Esc | ready |
| `turn-ended` | a short turn has ended | ready |
| `model-picker` | `Switch Model` open; the bare `>` is still above it | not ready |
| `command-palette` | `> /` with the palette, hint `esc to cancel` | not ready |
| `busy-thinking` | `Generating...` spinner, bare `>`, `esc to cancel` | not ready |
| `busy-streaming` | answer streaming, bare `>`, `esc to cancel` | not ready |
| `trust-dialog` | untrusted folder, alternate-screen trust menu | not ready |
| `draft` | unsent text in the composer; the hint row is blank | not ready |
What they show:
- **The text tail misses three ready screens.** On `ready-accept-edits`, `ready-plan` and
`turn-ended` the line-folded tail never satisfies `findAntigravityComposerIndex`; the screen
rule does. (`turn-ended` is the 1.2.14 form of the old `busy-turn-ended` known defect.)
- **A caret rule is wrong on the screen.** The grid keeps the bare `>` through a turn and behind
the picker. The text tail happened to lose it mid-turn (section 8); the screen does not.
- **The rule can read ready for a moment mid-turn.** Replayed in 64-byte chunks, the submit repaint
clears the composer a moment before `? for shortcuts` becomes `esc to cancel`. So a pane with an
output clock is held to the same 3s quiescence as Codex (tier 1b in `tui-idle-evidence.ts`); an
idle agy goes silent within a second and a turn keeps repainting its spinner. A restored pane with no
clock settles from the screen alone.
- **A name-only title settled the picker.** With the pane titled `agy`, the old ranking reached
its weak title lane and settled with `/model` open. When a trustworthy screen exists it now
decides, and that lane stays shut.
- **A mismatched grid garbles the chrome.** At 80x24, 100x30 or 60x20 the 120x40 recordings lose
the four-row shape, and resizing the model does not make the TUI repaint. So the live screen
counts only when its grid matches the PTY's reported size and is not a reflow the TUI never
repainted for (a re-attach that learned the real size late; a later PTY resize off that size
repaints it). Otherwise every pre-existing lane (text rules, title, quiet process) decides, as
before this change.
- **A readable screen outranks the text.** When the grid is trustworthy and refuses, the text rules
do not overrule it, in the ready-prompt tier or the quiet one; they decide only when there is no
trustworthy grid. The recorded picker-dismissed text over a model-picker screen proves it.
The visible-read probe's Antigravity branch is retired. It read the provider screen with a looser
rule (any caret after the banner) whenever the pane was Antigravity. The probe now runs the shared
rule for every screen-ruled agent, only for a pane with no output clock. A clocked pane settles
through the poll. The probe's screen read is the draft-blanking read projection, so it restores the
blanked composer row before the rule reads it (Cline's `❯ Ask anything...` otherwise reads as a
bare `❯`).
Still not captured: a tool-permission prompt (the operator's `toolPermission` is
`always-proceed`, and changing it means editing their settings), and the sign-in, theme, privacy and
update dialogs, for the reasons in the table below. Windows and Linux are unrecorded.
The 1.2.0 fixtures below were recorded before the recorder stopped writing ahead of shutdown. Their
last bytes are agy's exit teardown (`ESC[J ESC[?2004l`), so their final screens have no hint row.
## Versions
| Thing | Value |
| ------------------------- | ----------------------------- |
| `agy --version` | `1.1.25` |
| Banner printed by the TUI | `Antigravity CLI 1.2.0` |
| Captured | 2026-09-10, macOS, 120x40 PTY |
The binary and its own banner disagree. Any rule keyed to a version string must read the banner,
not `--version`, and must tolerate the two disagreeing.
## What the captures are
| Fixture | What it is |
| -------------------------------------------- | --------------------------------------------------------- |
| `antigravity-ready-api-key-gemini-model.txt` | Ready screen, API-key identity, Gemini 3.7 Flash (Low) |
| `antigravity-ready-account-info-hidden.txt` | The same ready screen with `AGY_CLI_HIDE_ACCOUNT_INFO=1` |
| `antigravity-dialog-trust-workspace.txt` | Workspace trust dialog, live and unanswered |
| `antigravity-dialog-model-picker.txt` | `/model` picker, live and unanswered |
| `antigravity-dialog-command-palette.txt` | Slash-command palette, live and unanswered |
| `antigravity-dialog-dismissed.txt` | `/model` picker dismissed with esc, then settled |
| `antigravity-busy-mid-turn.txt` | A real turn, recording stopped while the spinner was live |
| `antigravity-busy-turn-ended.txt` | The same turn after it ended and the composer returned |
## What could not be captured, and why
Nothing below was faked. Each is a case the recorder could not reach without changing the
operator's account state or configuration, which is out of bounds.
| Missing | Why |
| ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `antigravity-ready-business-non-gemini.txt` | This machine has no OAuth session — the CLI prints _"You are currently not signed in"_ and authenticates from `GEMINI_API_KEY`. Reaching a Business ready screen means signing someone in. |
| A non-Gemini model on any ready screen | `agy models` offers 11 models, all Gemini, and `settings.json` pins `modelProvider: gemini`. A non-Gemini row is not reachable from this account. |
| `antigravity-dialog-sign-in.txt` | Unsetting `GEMINI_API_KEY` does not reach the sign-in dialog; the CLI refuses to start because `modelProvider` is pinned. Reaching it means editing the operator's `settings.json`. |
| `antigravity-dialog-theme-picker.txt` | There is no `/theme` command in 1.2.0 (`Unknown command: /theme`). The picker appears only in first-run onboarding, which means deleting the operator's config. |
| `antigravity-dialog-privacy-notice.txt` | First-run onboarding, as above. |
| `antigravity-dialog-update-banner.txt` | Cannot be forced; no update was pending during the session. |
Each remains as a named, skipping case in the suite so it is visible rather than forgotten.
## What the transcripts show
### 1. The ready screen's model row is not at the start of a line
The ready screen prints a block-glyph logo down the left, and the identity, model and path rows are
painted **on the same physical lines as the logo**. What Orca derives is:
```
▀▀▀▀▀▀ Gemini API key
▀▀▀▀▀▀▀▀ Gemini 3.7 Flash (Low)
▄▀▀ ▀▀▄ ~
```
The earlier detector required `normalized.startsWith('gemini', trimmedStart)` on a trimmed line. The
trimmed line starts with `▀`, so that rule never matched. The shipped detector uses the bare
composer caret instead. Measured three ways on the real screen:
| Input | `isKnownReadyPromptPreview` |
| ------------------------------------------------------ | --------------------------- |
| Real ready screen | `true` |
| The same screen with the logo glyphs stripped | `true` |
| Real ready screen followed by the live `/model` picker | `false` |
So the logo — decoration, and suppressible with `AGY_CLI_HIDE_LOGO` — no longer decides readiness,
and the live dialog cannot reuse the stale composer caret as a ready signal.
### 2. The dialog used to satisfy the model rule
`/model` prints its options one per line:
```
Gemini 3.8 Flash
> Gemini 3.7 Flash (current)
Gemini 3.1 Pro
```
Those lines _do_ begin with `Gemini`, and a bare `>` composer line sits earlier in the same tail
from before the picker opened. Both halves of the old rule were satisfied **while a dialog owned the
screen**, and the pane read ready. The shipped detector now recognizes the active `Switch Model`
surface and rejects that stale caret until it sees `Exited /model command`.
### 3. `>` is the dialog selection marker, not only the composer caret
Every dialog uses `>` to mark the highlighted row: `> Yes, I trust this folder`,
`> Gemini 3.7 Flash (current)`, `> /add-dir`. The idle composer is a line whose whole trimmed
content is `>`. That distinction is the only thing separating them, which means the relaxation
proposed in PRs #15840 and #15852 — accept any line _beginning_ with `>` — would make the trust
dialog and the model picker read as ready. On 1.2.0 the idle composer is a bare `>`; those PRs'
1.1.17 mode-banner claim could not be reproduced here and may be mode-specific.
### 4. There is no email account row, and the row can be switched off entirely
For an API-key user the identity row reads literally `Gemini API key`. There is no `@`, no
domain, nothing an account-row rule can key on. Separately, `AGY_CLI_HIDE_ACCOUNT_INFO=1` — a
supported environment variable in the binary — removes the row from a fully ready screen, which
`antigravity-ready-account-info-hidden.txt` captures.
### 5. Dialogs are drawn two different ways, and the banner is never reprinted
The trust dialog and the sign-in splash take the **alternate screen** (`ESC[?1049h` … `ESC[?1049l`).
The model picker and command palette are drawn **in place on the main screen** with erase-to-EOL.
After dismissal the CLI prints `⎿ Exited /model command` and redraws the composer — it does **not**
reprint the banner. The header stays where it was at startup.
### 6. Rows are positioned with cursor addressing, not newlines
The status row is written with absolute and relative moves (`ESC[13;99H`, `ESC[83X ESC[83C`), so
`? for shortcuts` and `Gemini 3.7 Flash · low` end up on one derived line. Any rule that assumes
one screen row equals one `\n`-delimited line is reading a different document than the user sees.
## 8. Busy frames park the caret exactly like idle frames — the spinner is what differs
The frame that ends a turn-in-progress and the frame that ends an idle screen park the cursor with
the **same bytes**. Only the hint row differs, and the park erases it:
```
idle: ? for shortcuts ESC[83X ESC[83C Gemini 3.7 Flash · low CR ESC[2A ESC[2C ESC[?25h
busy: esc to cancel ESC[85X ESC[85C Gemini 3.7 Flash · low CR ESC[2A ESC[2C ESC[?25h
```
So a rule that keys on "the caret is the last thing in the tail" cannot tell busy from idle **on the
frame alone**. What saves it is what comes next. Each spinner tick is its own repaint with its own
park, two rows higher than the frame's:
```
ESC[?25l CR ESC[2A ⣯ Generating ESC[11D ESC[?25h
ESC[?25l CR ESC[2A ⣟ Generating. ESC[12D ESC[?25h
```
That second `CR ESC[2A` splices the composer row away, so the retained tail during a live turn ends
on the spinner row, not on the caret. Measured on `antigravity-busy-mid-turn.txt`:
| Capture | last retained line | bare `>` line present |
| -------------------------------------------- | ------------------ | --------------------- |
| `antigravity-ready-api-key-gemini-model.txt` | `>` | **yes** |
| `antigravity-busy-mid-turn.txt` | `⣟ Generating...` | **no** |
**Consequence for a caret-based rule:** it already answers "not ready" for a real mid-turn capture,
because there is no bare caret in the tail to match. A constructed input that keeps the park bytes
and only edits the status text is not faithful to a live turn — a live turn has a spinner row
repainting _below_ the composer.
**The residual window, and the clause it implies.** Between a frame park and the next spinner tick
the tail does end on the bare caret and is indistinguishable from idle. The gap is one tick
interval. Any readiness path gated on sustained quiescence is safe, because ticks keep arriving and
the pane is never quiet; a path that only inspects retained text is not. For those paths the
evidence supports one clause, and only one:
> **A braille glyph (U+2800–U+28FF) on the last visible line of the retained tail means working.**
That predicate already exists in this file for cursor-agent (`CURSOR_BUSY_SPINNER_RE`) and should be
reused rather than reinvented. It must be scoped to the **last visible line**, not the whole tail:
a first-run transcript prints `⠾ Signing in...` during startup, which would otherwise pin a ready
screen as busy forever.
Nothing else in the capture distinguishes the two states. The hint row (`esc to cancel` versus
`? for shortcuts`) is erased by the park in both cases, the park offsets are identical, and
`ESC[?25l`/`ESC[?25h` fencing appears around every repaint, idle or busy.
## Confirmed / refuted, by attempt
Evidence column names the fixture; all quoted text is from the committed transcripts.
### Attempt 1 — the rule at HEAD
| # | Claim | Verdict | Evidence |
| ---- | -------------------------------------------------------- | --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1.1 | A ready screen prints the banner `Antigravity CLI` | **Confirmed** | `Antigravity CLI 1.2.0` in both ready fixtures |
| 1.1b | …and its last occurrence in the tail is the live one | **Refuted** | The trust dialog's own body says _"Antigravity CLI requires permission to read, edit, and execute files here"_, so `lastIndexOf` lands inside the dialog |
| 1.2 | The model row begins with the vendor word `Gemini` | **Refuted** | `▀▀▀▀▀▀▀▀ Gemini 3.7 Flash (Low)` — the logo precedes it; never at line start |
| 1.3 | The caret line's whole trimmed content is `>` | **Confirmed** on 1.2.0 idle | bare `>` in both ready fixtures |
| 1.3b | …and only the composer prints `>` | **Refuted** | `> Yes, I trust this folder`, `> Gemini 3.7 Flash (current)`, `> /add-dir` |
| 1.4 | A ready screen prints the workspace path on its own line | **Refuted** | the path shares its line with logo glyphs (`▄▀▀ ▀▀▄ ~`) |
### Attempt 2 (loop 1) — blacklist the model line
| # | Claim | Verdict | Evidence |
| --- | ------------------------------------------ | ----------- | ---------------------------------------------------------------------------------------------------------------- |
| 2.1 | Dialog model-row wording is enumerable | **Refuted** | the palette lists 50+ commands with free-form descriptions; the picker prints whatever models the account offers |
| 2.2 | A dialog never reproduces a real model row | **Refuted** | the `/model` picker prints four real model rows, one per line, at line start |
### Attempt 3 (loop 2) — structural ordering on `headerIndex`
| # | Claim | Verdict | Evidence |
| --- | -------------------------------------------------- | ---------------------------------- | ---------------------------------------------------------------------------------------------------- |
| 3.1 | A live dialog is printed below the ready chrome | **Confirmed** for in-place dialogs | picker and palette append below the composer |
| 3.2 | The banner is reprinted when a dialog is dismissed | **Refuted** | `antigravity-dialog-dismissed.txt` shows `⎿ Exited /model command` and a redrawn composer, no banner |
| 3.3 | Antigravity does not use the alternate screen | **Refuted** | `ESC[?1049h` opens the trust dialog and the sign-in splash |
| 3.4 | No full repaint per keystroke | **Partly refuted** | typing `/mod` repaints the palette region on each keystroke with `ESC[K` |
Because of 3.2, `headerIndex` cannot be the anchor: it never advances. Ordering can only be
expressed against the model/caret positions, which is what 1.2 and 1.3b just invalidated.
### Attempt 4 (loop 3) — require a positive account row
| # | Claim | Verdict | Evidence |
| --- | ---------------------------------------------------- | ---------------------- | --------------------------------------------------------------------------------------------------------------------------- |
| 4.1 | Every ready screen prints an account row | **Refuted, twice** | API-key identity prints `Gemini API key` (no `@`); `AGY_CLI_HIDE_ACCOUNT_INFO=1` removes the row entirely |
| 4.2 | A startup dialog never contains an `@`-and-`.` token | **Not reachable here** | none of the captured dialogs contains one, but the palette shows free-form skill descriptions, which are user-authored text |
| 4.3 | The account row is distinguishable from prose | **Refuted** | the row is not a distinct line; it shares one with the logo |
### Attempt 5 (PR #19749, reverted) — ordering + account row
| # | Claim | Verdict | Evidence |
| --- | -------------------------------------------------------- | ----------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 5.1 | Ordering plus an account row separates ready from dialog | **Refuted** | the account row is optional (4.1) and the ordering anchor never moves (3.2) |
| 5.2 | Executing both builds was sufficient verification | **Refuted** | the executed input was the hand-written fixture, so the check reproduced the fixture's assumptions. The real screen disagrees with that fixture on the model row, the path row and the account row |
| 5.3 | The wedge is a model-name problem | **Refuted** | it is a line-start problem. Even `Gemini 3.7 Flash (Low)` — a Gemini model — fails, because a logo glyph precedes it |
### Cross-cutting
| # | Question | Answer |
| --- | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| X1 | Does `agy` set an OSC title distinguishing busy from idle? | **No.** Not one OSC title sequence appears in any capture. Title-based readiness is unavailable for this agent |
| X2 | Does it repaint with bare `\r`? | **Yes**, constantly, plus `ESC[K` and absolute cursor moves |
| X3 | Does the caret survive in the tail? | **Yes** — a bare `>` line is present in every ready capture |
| X4 | Banner-to-caret distance | ~8 derived lines on a 120x40 PTY; the banner falls outside the 6-line preview window, so only the full retained tail can see it |
| X5 | Pane title on the trust screen versus ready | Identical: none |
## Attempt six is shipped
Yes — but not as a variation on any of the five. Every one of them refined a predicate over
`\n`-delimited lines, and that is the layer where the evidence says the information is not.
What the captures support and the shipped detector now does:
- **The one stable, dialog-free ready marker is a line whose entire trimmed content is `>`.** It is
present in every ready capture and absent from every dialog capture, because a dialog's `>` always
carries its selected row's label. This is a much narrower rule than any attempt used, and it is
the only one that survived contact with the transcripts.
- **Drop the model-row requirement.** The model rows match dialogs and not the ready screen, so
keeping that requirement inverted the detector.
- **Do not require an account row.** It is optional by environment variable and carries no email for
API-key users.
- **Veto an active model picker.** `Switch Model` followed by a labeled selection row means the
bare caret belongs to the composer behind the picker; readiness resumes after `Exited /model
command`.
- **Do not anchor on `headerIndex`.** The banner is printed once and never reprinted.
- **The blocked-signal path already works** for the trust dialog: `antigravity-dialog-trust-workspace.txt`
is correctly refused today, by wording, not by structure.
What is still unknown and should be captured: the sign-in, theme, privacy and update dialogs, and
any ready screen where the composer is not idle (accept-edits and plan mode, which PRs #15840 and
#15852 describe from a screenshot). A bare-`>` rule is only as good as the claim that those modes
still end on a bare `>`; that claim is untested.
The honest summary is that this is a screen-shaped problem being solved with line-shaped tools. A
rule over the derived tail can be made much better than what ships today, but the durable fix is to
ask the terminal emulator what the bottom row of the screen actually is, rather than inferring it
from a byte stream that was written with cursor addressing. Attempt seven does that.
@@ -1,117 +0,0 @@
# Antivirus clearance for future releases
Orca collects a steady stream of antivirus and EDR false positives — see the
tracking issue for the current grouping. This document covers the part of that
problem worth engineering effort: **stopping the next release from being
flagged.**
Clearing a *historic* release is explicitly not a goal. A user sitting on a
flagged build should update to a cleared one, not wait for a vendor to whitelist
a version we no longer ship. Retroactive submissions cost the same effort per
vendor and expire the moment we cut a new version.
For the behavioural side of the problem — the process-tree shapes EDR scores,
which no whitelist fixes — read
[`windows-edr-posture.md`](./windows-edr-posture.md) instead. This document is
about file verdicts on the bytes we ship.
## The two mechanisms, and only one of them scales
**Sample submission** clears one build. You send the flagged file to a vendor's
analyst portal, they confirm it is clean, and the verdict is dropped from their
next definition update. This is reactive and per-release: cutting a new version
produces new bytes, new hashes, and a fresh chance of the same heuristic firing.
Doing this every release, across every vendor, is not sustainable.
**Signer and product whitelisting** clears every future build. The vendor
records the publisher identity or enrolls the product in a dynamic allowlist, and
subsequent releases inherit that trust without another submission. Enrollment is
one-time work per vendor, and it is the only lever that scales with our release
cadence.
Prefer enrollment. Use submission only to clear a live incident while enrollment
is pending, and only for a release users are actually expected to install.
## Prerequisites that make enrollment possible
None of these programs will accept an unsigned or anonymous binary, so these
come first:
1. **Every shipped PE is Authenticode-signed, and CI fails the release if not.**
Done as of v1.4.217 — the release workflow requires a valid SignPath
Foundation signature on the inner binaries and no longer fails open.
2. **Every shipped PE carries real provenance** — company, product, version,
description, and an explicit `asInvoker` manifest. An anonymous binary scores
worse than an identified one, and several portals reject submissions that
carry no version metadata.
3. **One stable signer identity.** Vendor allowlists key on the certificate
subject. Rotating signers resets accrued reputation, so a certificate change
is a re-enrollment event, not a transparent swap.
## Vendor programs
Enrollment state is deliberately left as a task here rather than asserted — fill
each in as it is confirmed, and record the account that owns it so a lapsed
enrollment is traceable.
| Vendor | Mechanism | Scope | State |
| ------------------------- | ---------------------------------------------------------------------------- | ------------------------ | ----- |
| **VirusTotal** | Monitor — paid; builds rescanned daily, developer and vendor both notified | ~70 engines at once | TODO |
| **Microsoft** | Defender Security Intelligence submission, as a software developer | Defender, Defender FP EP | TODO |
| **Microsoft** | Trusted Signing, or an EV certificate, for SmartScreen and Smart App Control | Reputation gates | TODO |
| **Kaspersky** | Whitelist Program — vendors submit builds for the Dynamic Allowlist | Endpoint, all platforms | TODO |
| **Trend Micro** | Certified Safe Software Service — pre-release software whitelisting | Endpoint, Virus Buster | TODO |
| **Bitdefender** | False-positive submission for software vendors | Endpoint, ATD | TODO |
| **Avast / AVG / Norton** | Gen Digital false-positive and whitelisting channels | Consumer suites | TODO |
| **ESET** | False-positive sample submission | Endpoint | TODO |
| **Tencent iOA** | No public developer channel found; needs a support relationship | iOA, macOS and Windows | TODO |
VirusTotal Monitor is the highest-leverage single entry, because it is the only
channel built for exactly this workflow: uploads sit in a private store, get
rescanned daily against every engine's current signatures, and when one flags a
file **both we and that vendor are notified automatically**. Pre-publish upload
is a supported use, which is precisely the future-release posture we want. It is
a paid service, monetised on developers and free to the antivirus vendors.
Be honest about its limit: VirusTotal states plainly that Monitor is not a free
pass to get a file whitelisted. Vendors sometimes keep a detection. What it
reliably buys is *early notice and a real contact path* instead of discovering a
verdict from a user's issue report weeks later.
Do not treat a plain VirusTotal *scan* as equivalent. A scan tells us a verdict
exists; Monitor is what routes it to someone who can drop it.
If the subscription is not worth it, the free fallback is the community-maintained
false-positive contact directory (`yaronelh/False-Positive-Center` on GitHub),
which collects the submission addresses and forms each vendor actually reads.
That replaces the hardest part of a submission — finding the right contact — but
keeps the per-release effort that Monitor removes.
## Where this lands in the release flow
The check belongs at RC time, not after a stable cut — a verdict discovered after
publication is a verdict users already hit.
`config/scripts/scan-release-artifacts-antivirus.mjs` reports the current
detection state of built artifacts by hash. Run it against an RC's artifacts, and
treat any engine verdict as a release-blocking question rather than an automatic
stop: these are third-party ML classifiers, so a hard gate on their output would
fail the release for reasons outside our control. Read the report, decide, and
submit if the flagged build is one we intend to ship.
The script looks up hashes by default and never transmits artifact bytes. Passing
`--upload` sends the file to VirusTotal, which distributes samples to partner
vendors — that is the intended outcome for clearance work, but it is a
publication, so it stays opt-in and out of any automated path.
## What not to do
- **Do not ask users to add exclusions** as the resolution. It suppresses the
symptom on one machine, and in several reports here the exclusion did not even
hold because the detection was behavioural rather than path-based.
- **Do not dispute a verdict without a sample.** Every report worth acting on in
this project came with a hash that we verified bit-identical to the published
release asset. That verification is what makes a submission credible.
- **Do not chase a vendor whose detection we cannot reproduce or name.** Route
those back to the reporter for the detection string and the exact flagged path
first.
-220
View File
@@ -1,220 +0,0 @@
# CI demand rollout
This implements the September 28 runner-demand analysis. The baseline inventory
covered September 27 04:00–September 28 04:00 UTC: 4,028 workflow runs, with
463 stratified job samples. Estimated occupancy was 1,081 runner-hours, dominated
by PR unit shards (414 hours) and Bun qualification (262 hours, since replaced by the
pinned-Node headless lanes). These are
sampled sums of job durations across different runner pools, not billing totals
or a guaranteed forecast of savings.
## What runs now
| Work | Ordinary draft update | Ready PR / final checks | Main reference |
| -------------------------------------- | ------------------------------------------ | -------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- |
| Static analysis and types | Immediately | Immediately | Existing workflows |
| Unit suite | Full, with shadow selection evidence | Full | Existing daily Node 24/26 x86 suite |
| Packages | After successful static analysis and types | Same | Existing release workflows |
| Headless Node persistence | Deferred until ready | Linux x64 for ordinary runtime changes; explicit platform families or all six for sensitive inputs | All six platforms for relevant main pushes; full nightly qualification at 11:30 UTC |
| Headless Node glibc/musl qualification | Deferred until ready | Linux-specific or full qualification, after persistence succeeds | Both architectures for relevant main pushes and nightly qualification |
| E2E | Existing targeted routing | Existing targeted routing | One complete run at 17:00 UTC |
Headless draft updates carry no verdict; readiness starts the checks. Relevant
ready PRs retain Linux x64 smoke coverage. Explicit Windows or macOS paths add both
architectures in that family, while Linux paths add Linux ARM and the glibc/musl
lanes. Root/toolchain inputs, native build inputs, shared execution/storage paths,
SSH, providers and relay changes retain every platform. Missing or incomplete
change evidence and failed import analysis also retain full qualification.
Unrelated changes skip through the dependency classifier. Relevant main pushes
qualify all six platforms and both Linux compatibility architectures; detection
uses the whole push before qualification can supersede an older relevant run.
The detector checks known build inputs with the pinned Node toolchain first.
Other paths install dependencies for import analysis; an uncertain result never
skips qualification. This changes setup cost, not the qualification policy.
Windows server prebuilds are cached separately by architecture and pinned build
inputs, including the runner image. Only successful qualification on main publishes
them. Consumers validate the payload and still run the pinned-Node load/spawn smoke;
a cache miss or invalid payload builds fresh. Nightly, manual and release builds
remain fresh. No native artifact is shared across platforms or ABIs.
Expensive PR jobs wait for static/type success. This reduces fan-out for failed
or rapidly superseded commits without sleeping on a runner. Successful isolated
PRs pay the extra stage latency. Existing per-PR cancellation remains in place.
Package assertions, native boundaries, SSH/folder coverage, cache warming and
slow-test assertions are retained.
The daemon running-work test imports the shared probe directly, with the daemon's
process inspector supplied as its callback. The renderer keeps its existing
adapter and forwarding tests. This removes a mocked renderer dependency from the
headless graph without changing the probe algorithm or skipping backend tests.
At validation, the graph fell from 6,018 inputs (1,070 renderer inputs) to 4,879
inputs (no renderer inputs), including nine added shared-probe cases. Renderer
adapter changes no longer qualify the headless matrix; shared probe and daemon
test changes still do. Future actual renderer imports remain discoverable.
## Unit selection rollout
PR planning runs alongside typechecking after their shared dependency setup; an
explicit join publishes its artifact before the unit matrix can start. Static
analysis's Node 24 install also prepares the native cache before matrix fan-out.
The daily compatibility workflow retains its separate planner and cache primer.
`ci-unit-plan.mjs` discovers the same include/exclude set as Vitest and follows
static imports, re-exports, literal dynamic imports, CommonJS requires and the
renderer aliases. Consumers of indirect filesystem/process inputs remain in the
candidate set, as do script/tool tests. Global configuration changes, deletions,
renames involving removed paths, unknown inputs and graph failures run the full
suite. A shard verifies the plan's source SHA and complete discovery list before
using it. Missing/stale artifacts fall back to full coverage, even if that means
running the suite on fewer shards.
The initial policy was **shadow**, with every shard retained and five concurrency
slots for a full run. The account is charged a slot per job rather than per core;
in the September baseline, eight 6.5-minute shards made this matrix 68% of daily
slot demand while the ARM pool queued 10.5 minutes at p95. The first local
inventory matched Vitest exactly (9,950 files at validation); representative source
changes retained roughly 88% of files because of indirect input readers.
That is evidence for conservative coverage, not evidence of the analysis's
hypothetical 50% unit-work reduction. Improvements to indirect dependency
modeling should be demonstrated against full results before expanding selection.
### Controlled full-PR sharding
Full PR unit runs now request ten four-worker ARM shards through the existing
planner. Daily Node 24/26 reference runs retain five shards per Node version, and
selected draft runs retain their five-shard cap. Coverage, isolation, setup files,
worker flags and the protected actual-Node runtime projects remain unchanged.
Set the repository Actions variable `ORCA_UNIT_FULL_SHARD_COUNT` to `5` to roll
back full PR runs. Remove the override or set it to `10` to restore ten. Only `5`
and `10` are valid; invalid values fail planning. If change/graph evidence is
unavailable, planning still retains every discovered file at the requested full
count. Confirm the actual matrix and complete selection evidence on a new run
when changing the variable; do not infer the effective count from the setting.
The ordinary [five-shard run](https://github.com/stablyai/orca/actions/runs/37657068640)
and [ten-shard run](https://github.com/stablyai/orca/actions/runs/37657071738) used
identical definition/source trees on pinned main `7d6d7ca6`. Both first attempts
passed all required jobs and ran 11,678 files exactly once, with identical module
counts, states and actual runtime routes: 112,615 passed cases, one expected
failure and 1,052 stock skips.
| Observed boundary or resource | Five shards | Ten shards | Change |
| ------------------------------------------------- | ----------: | ---------: | -----------------------: |
| Longest unit test step | 571s | 322s | 43.61% less time |
| Unit prerequisite release to last unit completion | 618s | 362s | 41.42% less time |
| Unit release to required verification | 628s | 371s | 40.92% less time |
| PR creation to required verification | 819s | 545s | 33.46% less time; 1.503× |
| Aggregate unit test-step time | 2,649s | 2,852s | 7.66% more |
| Aggregate held ARM runner time | 2,827s | 3,206s | 13.41% more |
| Aggregate setup before tests | 157s | 317s | 101.91% more |
| Aggregate dependency installation | 57s | 112s | 96.49% more |
Nominal peak unit worker slots increased from twenty to forty. The comparison was
one observational pair on different hosts and cache states; unit runners started
6–8 seconds after allocation. Four additional diagnostic ARM jobs started after
all fifteen comparison runners had been allocated. This is not a quiet-fleet,
representative queue-tail or historical twofold-speedup result. Stock artifacts
prove module counts/states/routes, not complete individual case identities.
This is a controlled latency rollout with a resource tradeoff. The representative
week-long capacity evaluation below remains pending; one successful pair does not
satisfy that fleet gate or erase the earlier oversharding concern. Monitor queue
and provisioning delays, PR-to-verification latency and aggregate ARM runner time
in equivalent traffic windows, including cancellations and planning/reference
costs. Use the five-shard rollback if queue delay erases the latency gain.
Temporary benchmark PRs #26271 and #26272 never merge.
Every shard uploads `unit-selection.json`, `unit-timings.json` (including module
outcomes), and its assignment. The evidence job combines these into
`unit-selection-review-attempt-N/selection-review.json`, reporting:
- Whether every discovered file appeared once across a complete reference run.
- Failures outside the candidate set, including failures in otherwise red runs.
- Measured worker time that selection would omit; worker times overlap and are
not runner occupancy or a prediction of wall-clock savings.
Missing/duplicate shards, stale plans, interrupted runs and unhandled errors do
not count as complete references. Diagnostic upload/report failures do not make
tests pass and do not independently fail successful tests.
After representative complete shadow runs show no missed failures, set repository
variable `ORCA_UNIT_SELECTION_MODE=selected` to enable selection **only for draft
PRs**. Keep full ready-PR checks and the daily compatibility suite. Inspect at
least a week's evidence across renderer, main, shared, SSH and fixture changes
before promotion, including red runs rather than only successful examples.
Unknown variable values retain shadow mode. Unset the variable or set it to
`shadow` to roll back immediately. Selected runs use one to five timing-balanced
shards based on retained work.
Ready-for-review result reuse includes `unit full` in its source/workflow
identity. A green selected draft cannot satisfy the final full check, even if
the repository variable changes between runs. An already successful _full_
identical-source check can still be reused.
To inspect downloaded shard artifacts locally:
```sh
node config/scripts/ci-unit-selection-review.mjs ARTIFACT_DIRECTORY
```
## Review automation
Pullfrog recognizes the existing `Review #N [id]` and
`Review new commits on #N [id]` dispatch names. Explicit dispatchers may provide
`pull_request_number` and `head_sha`. Explicit PR identities share concurrency at the workflow boundary. Legacy review
names use a bounded lookup of the latest 100 dispatches and cancel only lower
run IDs for the same PR; a delayed older scope cannot cancel a newer review.
Unrecognized tasks are never grouped. The scope job alone has Actions write
permission for ordered cancellation. Closed PRs and explicitly stale heads
are skipped. A second head check prevents starting an agent after its queued
head has changed. Unrecognized agent tasks remain independent; lookup failures
also retain an independent task rather than cancelling unrelated work.
This does not introduce a fixed debounce interval or remove final reviews.
Dispatchers should supply `head_sha` for reliable stale-at-dispatch detection;
legacy names identify a PR but do not prove which head the prompt describes.
## E2E signal
The daily reference still executes all shards and keeps original verdicts. Each
shard uploads Playwright JSON and publishes expected, skipped, unexpected, flaky
and startup-error counts with the failing test names/messages. Targeted PR and
manual coverage remain available. This change does not fix the historically
red tests or pretend they pass.
`config/e2e-failure-tracking.json` can separate an evidenced repeated failure
from new failures in the summary. Each entry must have exact `file`, full
`title`, `project`, a nonempty stable `message` substring, an `@owner`, a linked
repository `issue`, and an ISO `expires` review date. Expired/malformed entries
are ignored and reported; changed error signatures appear as untracked. Entries
never skip a test or change its exit status. The initial list is empty because
the analysis established red workflows but did not establish owners and
reproductions for individual failures. Do not blanket-baseline an entire red run.
## Capacity measurements and acceptance
`CI runner demand` runs daily at 04:23 UTC and can be dispatched manually. It
reads the previous 24 complete hours in hourly pages, samples up to six runs per
workflow/outcome stratum, and fetches job pages with bounded concurrency. An
hour exceeding the API's 1,000-result search cap fails visibly. The report and
raw evidence are retained for 30 days. No extra runner pool is provisioned.
The report measures the full job durations of runs **created** in the window,
not occupancy clipped to the window: earlier runs that overlap it are excluded,
and completed sampled jobs may finish after it. This matches the baseline
cohort method. Workflow IDs keep ref-qualified paths in one sampling stratum.
The report shows weighted runner-hours and cancelled-run hours per workflow,
runner-minutes per completed PR _run_, and weighted queue/provisioning p95 per
runner label. It counts latest attempts only, excludes incomplete jobs, and
retains zero-job observations. It does not measure other repositories competing
for organization capacity. Compare equivalent traffic windows, not raw totals
alone. The collector needs only `contents: read` and `actions: read`.
After a week, compare runner-minutes per PR run, cancellation occupancy and
queue p95 in each affected pool. Count newly added planning/reference overhead.
A 25–35% overall reduction remains an experiment target, not an achieved result;
selection, coalescing and matrix reductions overlap and cannot simply be added.
File diff suppressed because it is too large Load Diff
@@ -1,85 +0,0 @@
# Cline and Prime Agent readiness: what the transcripts show
Both agents paint their composer with cursor addressing on the alternate screen, so the line-folded
text tail cannot see it (#23268, #22153). Their readiness is read off the live screen by the
`composer_ready` rules in `src/main/runtime/agent-state-rules/cline.json` and `prime-agent.json`,
through the same tiering as
Antigravity ([`antigravity-readiness-evidence.md`](./antigravity-readiness-evidence.md)): a pane
with an output clock is believed only once quiet, and when a trustworthy screen exists it decides,
so the quiet-process lane cannot settle a dialog it cannot see. Without a readable screen (a
re-attached pane whose grid is untrusted) both keep the quiet-process lane they had before; no
rest-signal entry closes it. Recordings follow
[`agent-pty-transcript-capture.md`](./agent-pty-transcript-capture.md) and are replayed by
`cline-screen-readiness-transcripts.test.ts` and `prime-agent-screen-readiness-transcripts.test.ts`.
## Cline
Recorded 2026-09-30 on macOS, `cline` 3.0.66, OpenRouter's free router, with an isolated
`--config`/`--data-dir`. `cline-3-0-65-win32-startup.txt` is a Windows capture from PR #23269.
| Fixture (`cline-3-0-66-*.txt`) | Screen at the end | Rule |
| ------------------------------ | --------------------------------------------------- | ----------- |
| `ready`, `ready-80x24` | startup composer, `❯ What can I do for you?` | ready |
| `ready-plan` | Plan mode, `❯ Plan something...` | ready |
| `turn-ended` | after a turn, `❯ Ask anything...` | ready |
| `promo` | "Introducing Cline Desktop" drawn over the composer | not ready |
| `permission` | `Approve tool call?` with `[y] Approve [n] Deny` | not ready |
| `slash-menu` | `❯ /` with the command list under it | not ready |
| `draft` | unsent text in the composer | not ready |
| `busy-streaming` | a reply streaming, spinner scrolled off the top | reads ready |
- **The streaming screen is the idle screen.** Once a long reply scrolls its spinner row away, the
grid is the same composer box as at rest. Only quiescence separates them. A clockless restored
pane still settles from the screen, like Antigravity and Prime: `onPtyData` stamps
`lastOutputAt` on every chunk, so a streaming pane has a clock from its first byte after attach,
and only a pane that has printed nothing since attach is judged on the screen alone. A reply that
stalls for 3s with its spinner off screen would read ready; that is not captured and not ruled
out.
- **The placeholder is not fixed.** It changes with mode and history, so the rule accepts the three
captured placeholders and nothing else. A typed draft looks the same to the read projection, which
is why screen-ruled agents read raw rows (`readScreenRuledLines`, which also requires the PTY's
own grid). Every other agent keeps `readLiveTerminalScreenLines` exactly as before: replaying all
93 other fixture/grid pairs frame by frame gives identical verdicts on this branch and its base.
- **The promo popup appears about 40ms after the composer** and returns on each launch until it
is dismissed once (`cli-notices.json`). The quiet lane covers that race; the popup carries no
blocked wording.
- **The approval prompt is quiet and unworded.** No blocked rule matches it, so before the screen
decided, the quiet-process lane would have settled it. That lane stays open only while the pane
has no readable screen.
- **Windows:** the reported bug (#23268) is Windows, where the text tail reorders rows. Only the
3.0.65 contributor capture covers it, and it reads ready from the screen. The rule depends on
rendered rows, not byte order, but no Windows turn or dialog is recorded.
## Prime Agent
Recorded 2026-09-30 on macOS, `prime-agent` 0.9.8, OpenRouter
`inclusionai/ling-3.0-flash-sante:free`, with an isolated `HOME` (the first-launch question only
reappears in a fresh one). `prime-agent-0-9-5-*.txt` are 120x35 captures from PR #22154.
| Fixture (`prime-agent-0-9-8-*.txt`) | Screen at the end | Rule |
| ----------------------------------- | ---------------------------------------------- | --------- |
| `ready`, `ready-80x24` | bare `>` over the `← manage` footer | ready |
| `ready-after-question` | trace question answered "Not now" | ready |
| `turn-ended`, `tool-turn` | a turn (one with a Python tool call) has ended | ready |
| `trace-question` | animated "Share agent traces" question | not ready |
| `slash-menu` | `> /` with the command list | not ready |
| `busy-streaming` | `⠦ Writing · 6s` status row above the composer | not ready |
| `draft` | unsent text in the composer | not ready |
- **The footer and caret stay up mid-turn.** Only the braille status row Prime keeps directly above
them says a turn is running, so the rule vetoes on it.
- **The rule reads ready for moments it should not.** Replayed in 64-byte chunks, Prime erases that
status row before redrawing it, and on first launch it paints the idle composer a few tens of
milliseconds before the trace question covers it. Both are covered by quiescence: the spinner and the question's
animation repaint continuously, and an idle Prime is silent.
- **No permission prompt exists to capture.** With default settings Prime ran the tool call without
asking.
- **Why the Prime captures are large.** Prime does no cell diffing. Each synchronized frame
(`ESC[?2026h`…`ESC[?2026l`) erases and rewrites every row it touches, so a streaming turn costs
about 4 KB per spinner tick or token (337 frames, 8,593 `ESC[2K` in 1.45 MB). The first-launch
welcome animates a full-screen dotted background with a colour code per glyph, about 10 KB a
frame at roughly ten frames a second. `busy-streaming` and `trace-question` are truncated to the
first frame that shows the screen their tests need (see each `.meta.json`).
`ready-after-question` cannot be: the animation precedes the answer, and truncation only drops
the end.
- 0.9.4's `← agents/resume` layout is not supported; the rule needs 0.9.5 or later.
-44
View File
@@ -1,44 +0,0 @@
# CodeBuddy terminal harness
CodeBuddy uses its own identity in the desktop and mobile agent catalogs, process
detection, telemetry, hooks, resume records and AI Vault. `codebuddy` and `cbc`
identify the interactive CLI; print, server, ACP and detached background modes do
not identify an interactive pane. Model and effort selections use `--model` and
`--effort`; the CLI resolves its stable model aliases for the current account.
Managed hooks use the existing Claude-compatible installer and transport, with
CodeBuddy's own `.codebuddy/settings.json` and hook source. Installation preserves
user hooks and statusline configuration. Local and SSH installers share the same
plan; Windows explicitly selects CodeBuddy's supported PowerShell hook shell.
Status flows through the execution host's canonical hook store.
## Observed lifecycle
Verified with the authenticated CodeBuddy 2.159.0 CLI on macOS:
- `UserPromptSubmit` starts work. `SessionStart` can arrive afterward and must not
settle that work.
- An unanswered `AskUserQuestion` emits a `Notification` with
`notification_type: permission_prompt` and the message
`needs your permission to use AskUserQuestion`.
- In this version, `PreToolUse(AskUserQuestion)` arrives **after** the user answers;
it resumes working, followed by `PostToolUse` and `Stop`.
- `Stop` supplies the final assistant message and the provider session identity
supplies the resume command.
The sanitized real event sequence is
`src/shared/__fixtures__/codebuddy-question-hooks.jsonl`. The regression test
replays it through the shared hook listener. Qoder keeps its existing lifecycle
semantics while sharing the common event projection.
AI Vault reads CodeBuddy's `type: message` JSONL records with top-level roles and
`input_text` / `output_text` content blocks. Local, WSL and SSH discovery use the
provider's `.codebuddy/projects` tree and the same incremental parser.
## Validation scope
Live macOS checks exercised launch through Orca's agent menu, model switching,
question waiting, answer submission, working and completion indicators, history
discovery and a resumed session recalling its earlier answer. Hidden-renderer CDP
screenshots record the working, question and completed states. Windows, Linux,
WSL and SSH runtime execution have not been exercised live on this machine.
@@ -1,28 +0,0 @@
# DeepSeek Build terminal identity
DeepSeek Build is the third-party [`innocarpe/deepseek-build`](https://github.com/innocarpe/deepseek-build) product, published as `@innocarpe/deepseek-build`. It is distinct from official DeepSeek Harness (`@deepseek-ai/dsh`), Reasonix, DSH Console and generic DeepSeek TUI wrappers.
Orca recognizes manually started Build terminals through its existing process and title observations. `TerminalAgent` includes `dsb`; the launchable `TuiAgent` registry does not. No Build launcher, hook, readiness profile, resume command or history reader is registered. Existing generic terminal input remains available.
The source and actual macOS release were checked at **v6.9.0**, source commit `74df67a56988e9a32845c4565cc62b021ea68c7d`. The darwin-arm64 release tarball SHA-256 is `a57f225a537fc5c027ac4592e3f37f7bdc28cc2d6e27a366934511ea565cb874`.
- `package.json` publishes `dsb.js` and `deepseek-build.js` npm shims; the native child is `deepseek-build-agent`.
- `crates/dsb-cli/src/main.rs` separates the full-screen entry from `run`. Global value options such as `--cwd` can precede `run`; those invocations remain excluded from interactive recognition.
- `crates/dsb-cli/src/agent_launch.rs` emits the product OSC 0 title. The vendored pager's `notifications/title.rs` composes spinner, activity and product segments with ` - ` separators.
- The committed `dsb-6-9-0-folder` PTY fixture records the released binary's welcome screen and actual title in an isolated home and plain folder. It makes no successful-authentication or completed-model-turn claim. Its runtime test feeds raw chunks through `onPtyData` with foreground inspection unavailable.
Explicit native owner markers retain their existing precedence. A Claude task merely mentioning Build is not a Build identity. Runtime publication reuses the existing optional `agentIdentity` string; no new RPC, stream opcode or status producer is added. Older hosts can omit identity, while older readers retain their existing unknown-agent handling. Execution-host process/title observations work without a Git repository; local source tests do not establish native Windows, Linux or SSH device coverage.
For rendered proof, isolate both Electron and the actual PTY. On macOS, `login(1)` can replace the shell's inherited home. Test-only `ORCA_DISABLE_MACOS_LOGIN_SHELL=1` avoids that wrapper; do not change production launch policy for a proof. Require a nonce-bound file written by a helper executed in the spawned PTY, containing its actual `HOME`, `USERPROFILE`, `DEEPSEEK_BUILD_HOME`, `GROK_HOME` and trust-RPC flag, and verify it before agent launch. A terminal-text assertion can match command echo and is not isolation evidence. Explicit provider environment at the final execution boundary protects the test even after shell startup.
The observation-type propagation and title/process recognition adapt Wooseong Kim's (`innocarpe`) [PR #23485](https://github.com/stablyai/orca/pull/23485), with source-backed corrections for the second npm shim and value options before `run`. Keep that predecessor open until a reviewed successor merges.
Independent review follow-up: upstream 6.9.0 outer `Commands::Agent` forwards native PagerArgs options. Native `-p`/`--single` (alias `--print`), `--prompt-json` and `--prompt-file` are one-shot forms and are excluded from interactive process/foreground identity, including equals/compact short forms, npm wrappers and preceding value options. Positional interactive prompt text, native option values and the native `--` terminator stay distinct. Actual release native `--help` confirms exposed flags; source alias and forwarding are pinned above.
Title follow-up confines Gemini identity/normalization and status sniffing before inspecting Build activity text. A verified Build title uses its leading own braille frame for working and leading `⚠ Action Required - ` for permission; embedded Gemini glyphs in activity/session/cwd text do not change its identity or status. Source-backed frame tests cover wrapped and alert variants, plus OSC input through the actual runtime/listing path. Native Gemini and other provider corpus contracts stay covered.
Actual released outer `dsb agent -- --help` prints the native TUI help (`outer-forwarded-help.txt`), confirming Clap consumes the outer separator before forwarding. The observer distinguishes this from the native `--`: `dsb agent -- --print task` is one-shot, whereas `dsb agent -- -- --print` and direct native `-- --print` retain literal interactive prompt text.
Native grammar follow-up: only the outer wrapper's `run` subcommand is one-shot. Native `deepseek-build-agent run`, forwarded `dsb agent run`, and `--leader-socket run` remain interactive; the last consumes `run` as a path value. Native `-c` is boolean, so Clap accepts `-cp task` and `-cptask` as continue plus single-turn prompt. Attached `-m`/`-r`/`-s`/`-w` values (including after `c`) do not expose a prompt flag. The released binary accepted the five review argument topologies with `--help` under an executed private-child environment assertion (`native-grammar-oracle.json`); this proves parsing/help, not successful model generation.
Attached prompt values can begin with hyphens: native `-p-` and `-cp--print` consume `-` and `--print` as the single-turn prompt. The observer accepts the entire remainder after `p`, while the attached m/r/s/w value shields stay covered. The released native binary and outer `agent` wrapper accepted both forms with `--help` in a nonce-asserted private child (`attached-p-oracle.json`).
-44
View File
@@ -1,44 +0,0 @@
# DeepSeek Harness integration
Orca detects the community `@deepseek-harness-tui/dsh-tui` launcher (`dsh-tui`, alias
`dst`) and requires the official `@deepseek-ai/dsh` executable too. The launcher
boots the `dsh-tui` profile; Orca passes `.` to select the current workspace and
reach its composer on the first launch. The official Harness does not bundle this community TUI.
DSH Console and DeepSeek Build are separate products and are not interchangeable
with this launch contract.
The official DSH 0.2 CLI accepts both `dsh --profile headless` and `dsh headless`.
Orca excludes the known `web`, `headless`, `sdk`, `sdk-minimal`, `acp`, and `desktop`
profiles from interactive process recognition, along with plugin management and
configuration dumps. Custom profile names remain eligible because profiles are
user configurable. Only launcher arguments are inspected; app prompts, resume IDs,
and patch filenames cannot change the selected profile's identity.
Status hooks use the official `@deepseek-ai/dsh-hooks-claude-code` plugin, installed
as an owned block in `$DSH_HOME/cordis.patch.yml`. User entries outside the block are
preserved. Local installation respects `DSH_HOME`; the existing SSH installer uses
the execution host's default `~/.dsh` because SFTP cannot read its environment.
Hooks report session start, prompt submission, tool start/end, and stopping through
Orca's host status store. Approval has no dedicated hook; it is not inferred from
an uncaptured screen. Subagent lifecycle events are ignored for parent-pane status.
DSH 0.2 still emits an empty `transcript_path` in Claude-compatible hooks. Its
session persistence defaults to compressed JSONL under `$DSH_HOME/sessions`.
Orca can resume a hook-associated session through `dsh-tui --resume <id>`, but
currently does not discover DSH logs in Agent Session History. Resume support alone
does not establish transcript-history support.
## Reproduce the official launcher check
Install `@deepseek-ai/dsh@0.2.0-rc.2` into a disposable prefix, then run:
```sh
ORCA_BACKGROUND_LAUNCH=1 ORCA_REAL_DSH_CLI=/path/to/prefix/node_modules/.bin/dsh \
pnpm test src/shared/dsh-real-cli.test.ts
```
The opt-in test checks published version, composed profile configurations, and
headless help in an isolated home and working folder without a model request.
Interactive readiness is separately pinned to the captured community TUI transcript
in `src/main/runtime/__fixtures__/dsh-tui-ready-no-key.txt`; that older capture is
not proof of current TUI compatibility or paid generation.
-94
View File
@@ -1,94 +0,0 @@
# Git Compatibility Policy
## Scope
Orca executes the user's Git binary on three kinds of execution host: native,
WSL, and SSH. Each host can have a different Git version, so compatibility
state must be scoped to the host that actually runs the command.
Git 2.25 is the core-workflow compatibility baseline for command selection. It
is the oldest line that covers Orca's baseline use of porcelain v2, `branch
--show-current`, `restore`, and sparse checkout. Optional features that need a
newer Git must degrade safely and cache the missing capability. Orca does not
currently block older Git at startup, but new command construction should not
assume features introduced after this baseline.
## Capability Rules
When a newer Git feature materially improves correctness or performance:
1. Keep a baseline-compatible command or parser as the fallback.
2. Detect rejection with a narrow predicate for that option or subcommand.
3. Run the preferred command through `GitCapabilityCache` so a rejection is
remembered for the native host, WSL distro, or SSH provider that produced it.
4. Retry after the cache interval so an in-place Git upgrade self-heals without
restarting Orca.
5. Test the first fallback, later calls that skip the rejected probe, concurrent
probe coalescing, and execution-host isolation where applicable.
Do not branch only on a parsed `git --version`. Vendor builds can backport
features, and wrappers can report a host version that differs from the binary
used inside WSL or SSH. A behavior probe plus a precise fallback is the final
authority.
## Current Capabilities
| Capability | Preferred behavior | Compatibility behavior |
| --------------------------- | ----------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `fetch-no-write-fetch-head` | Fetch a private rebase ref without changing worktree-local `FETCH_HEAD` | Serialize all Orca fetch/pull operations per worktree Git directory before Git 2.29 |
| `worktree-list-z` | NUL-delimited worktree paths with `prunable` marks | Line-block parser for Git before `worktree list -z` (2.36); the `prunable`/`locked` annotations still parse on Git 2.31–2.35, and a path-existence probe restores `prunable` detection for Git before 2.31 |
| `worktree-add-lock-reason` | Create a prepared checkout with its ownership marker already present | Before Git 2.33, add without checkout, then exclusively create the same reason marker before materializing files |
| `rev-parse-path-format` | Absolute repo metadata paths | Resolve legacy relative output against the scanned repo |
| `for-each-ref-exclude` | Exclude remote HEAD before the output limit | Request extra refs, then filter remote HEAD in Orca |
| `merge-tree-write-tree` | Derive real-merge conflicts and no-op tree proofs | Omit the conflict summary and keep conservative branch cleanup behavior before Git 2.38 |
| `merge-tree-merge-base` | Supply the already-resolved merge base | Use the older two-commit `merge-tree --write-tree` form |
Prepared creation registers and locks without checking out files while its shared
exact-base fetch runs. A cancellable in-process barrier waits for fetch settlement,
including offline failure, then resolves the current commit OID on the owning host
and materializes files once. Existing preparations queue a tip refresh on that same
barrier before a create can claim them. Finalization still resolves the latest base
and runs the post-checkout hook only when attaching the requested branch. This uses
baseline-compatible `rev-parse` and `reset --hard`; the barrier never enters Git
transport options or the remote wire.
### Placeholders That Fail Open
`GitCapabilityCache` records commands Git _rejects_. A `git log --format`
placeholder Git does not know is not rejected: Git echoes it verbatim and exits
zero, so there is no error to remember and no probe to cache. Ask for both forms
in one record and pick at parse time.
| Placeholder | Preferred behavior | Compatibility behavior |
| --------------- | ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `%(decorate:…)` | Git 2.43 separates commit decorations with `\x1f`, so ref names containing commas survive | The same record also carries `%D` (Git 2.10); an unexpanded `%(decorate` placeholder selects it, at the cost of comma-splitting |
## Why Not `simple-git`
`simple-git` is a process wrapper around the installed Git binary. Its custom
options and `raw` API pass arguments through to Git, so it cannot make a newer
flag work on an older binary or choose Orca's semantic fallback automatically.
It provides version reporting and subprocess queueing, but Orca already needs
its own WSL/SSH routing, cancellation, tracing, redaction, process cleanup, and
bounded output handling. Replacing the runner would move—not remove—the
capability problem.
## CI Contract
PR checks run the capability contract against real Git 2.25.5, 2.38.1, and
2.49.1 binaries. This spans the pre-2.29 serialized `FETCH_HEAD` fallback, the transitional
`merge-tree --write-tree` behavior before `--merge-base`, and current Git.
The three lanes run in parallel and each Git call in the container lanes costs a
container start, so their wall clock is runner contention, not Git. Build the
2.25.5 binary and pull the images before the lanes start: anything heavy left
running alongside them is charged to whichever boundary case is in flight and
surfaces as a Vitest timeout rather than as a slow setup step.
The idle maintenance contract also verifies `multi-pack-index write` and packed
object reads. The command arrived in Git 2.20 and needs no newer-Git fallback;
Orca uses only index metadata writes, respecting `core.multiPackIndex=false`.
Keep the unit tests alongside that matrix. They cover concurrent probes,
native/WSL/SSH/relay isolation, and error-stream shapes that a single real
binary invocation cannot exercise deterministically.
File diff suppressed because it is too large Load Diff
-146
View File
@@ -1,146 +0,0 @@
# IME Regression Checklist
## Follow-up issue ledger
These reports define durable acceptance contracts, not only the symptoms from
one machine.
| Issue | Root cause | Ownership invariant | Required evidence |
| ----------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| [#16911](https://github.com/stablyai/orca/issues/16911) Native Chat preedit overwritten by streaming or attachment settlement | React reconciliation or an asynchronously resolved attachment writes an application draft while the browser owns the composing textarea; duplicate settlement can then re-adopt stale DOM. | From `compositionstart` until the first `compositionend` or blur, the browser owns the textarea. Resolved paths queue at the shared semantic sink; settlement adopts the browser DOM before flushing once. | Repeated stale streaming rerenders preserve the element and preedit; idle external drafts still synchronize; concurrent SSH completions preserve completion order and duplicates; both settlement orders run once; disable discards queued work; blur does not steal focus; native composition commits once. |
| [#16949](https://github.com/stablyai/orca/issues/16949) terminal preedit has no visible cursor | The opaque composition overlay covers the renderer cursor; at the final cell an over-wide inline preedit can also place its caret beyond the clipped screen. | The existing xterm `CompositionHelper` owns a visible caret after the preedit and before any row remainder; final-cell composition end-aligns within the screen while mid-line composition stays left-anchored. | Start, update, arbitrary-width final-cell containment, mid-line remainder placement, cleanup, and update-without-start are covered; preview and normal terminals inherit the live cursor theme. |
| [#16950](https://github.com/stablyai/orca/issues/16950) typing diagnostic records no CJK samples | The probe observes echoing keydowns but not reconciled composition commits, then guesses which queued input owns opaque TUI output. | A reconciled composition is observed even when `compositionend.data` is empty; only an isolated input enters exact percentiles, while overlap or a dropped-input gap produces one aggregate ambiguous burst. | Recorded Linux IBus empty-data commit, isolated direct and IME samples, mixed-source ambiguity, timeout/cap gaps, UTF-8 output bytes, and stop/drain cleanup are covered. |
| [#17104](https://github.com/stablyai/orca/issues/17104) Korean preedit repeats the Codex placeholder | Generic xterm row-tail reproduction exposed an application-semantic Codex or Claude composer placeholder that presentation style cannot identify safely. | Xterm always preserves generic covered row text. Orca's existing structural composer classifier masks only a verified placeholder during the exact active composition session; repaint reclassification runs only while composing, and end, blur, or disposal clears ownership, class, and listeners. Arbitrary dim output and shell lookalikes remain visible. | Codex prompt/footer and Claude prompt/frame classification, arbitrary all-dim and shell-lookalike negatives, repaint entry and exit, end/blur/disposal cleanup, and rendered Electron proof at cursor column 2 preserving generic row text are covered. |
## Preedit cell advances (#19315)
Single-codepoint CJK graphemes use the active Unicode provider's cell width and
measured font advance. Ordinary inline spans preserve browser bidi and baseline
layout; equal corrections share a run. Keep glyphs unscaled and the underline,
caret, and candidate textarea aligned with the rendered preedit. Appending ASCII,
emoji, or another script must not change an existing CJK prefix's correction.
Combining sequences, emoji, other scripts, and whitespace retain native shaping.
Font loading, typography changes, and renderer metric changes must update an open
composition; row-tail repaints preserve its unchanged nodes.
Cold font measurements and styled runs share a fixed work budget. Repeated CJK
can remain one corrected run; after the budget is exhausted, the remaining text
keeps its native advance. This deliberately leaves the original spacing mismatch
in the tail of unusually varied long compositions, without switching the prefix
back to native spacing or rebuilding thousands of spans.
`terminal-ime-xterm-preedit-cell-grid.test.ts` covers text preservation, native
clusters, lifecycle, and bounded work. `terminal-ime-preedit-cell-grid.spec.ts`
checks rendered glyph origins, caret/textarea geometry, underlines, font changes,
and native shaping at DPR 1, 1.25, and 2 with WebGL on/off.
`terminal-ime-preedit-continuity.spec.ts` covers mixed suffixes and budget crossings.
These checks use Chromium composition through CDP; they do not replace native OS
IME evidence.
## Bounded-state and ownership contracts
Every transient collection and ownership tracker must have an explicit lifetime and bound:
- Native Chat uses `NATIVE_FILE_DROP_MAX_PATHS` (`256`). If a resolved completion would cross the cap, the whole batch is rejected atomically and the overflow notice remains visible through settlement; accepted paths keep order and duplicates. The queue is cleared before re-entry and on disable or pane-owner remount.
- The terminal placeholder mask tracks one scalar `activeSessionId` because xterm renders one composition view. A newer start supersedes an older one, a stale end cannot clear the latest owner, and blur or disposal clears it. The composition route keeps its per-ID reference-counted map intentionally for transport ownership; it is not replaced by the scalar.
- Typing diagnostics cap pending and ignored echo candidates at `MAX_PENDING_ECHO_CANDIDATES` (`64`), cap pending user-input signals, drain timed-out candidates, and clear all series on pane detach. Overflow becomes an explicitly ambiguous burst rather than an arbitrary attribution.
## Native Chat asynchronous attachment settlement
Attachment resolution is an external semantic write, including local file
selection, pasted-image temp saves, and SSH uploads. While composition is active,
it must not replace the browser-owned textarea value.
- Start two concurrent SSH uploads, resolve the second first, and return one path
twice. Preserve completion order and both duplicates.
- On the first settlement event, adopt the browser DOM before flushing queued
paths. Exercise `compositionend` then blur and blur then `compositionend` in
one React batch; both orders must adopt and flush exactly once.
- If the composer becomes disabled before an upload resolves or before the queue
flushes, discard that result.
- If `compositionend` is omitted, blur performs the same one-time settlement
without focusing the textarea or stealing focus back.
## Cross-platform verification
Synthetic DOM events prove Orca's event and rendering contracts, but they do
not exercise the operating system's input method. Changes must also cover:
| Environment | Native evidence |
| ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| macOS | A native Korean 2-set composition in Native Chat and a terminal; preedit survives external renders, the caret remains visible, and commit occurs once. |
| Windows | Microsoft Korean IME over an untouched Codex placeholder; the preedit is the only visible text, the placeholder returns after cancel, and ordinary mid-line content remains visible. |
| Linux / SSH | IBus Hangul with an SSH-hosted PTY; an empty-data `compositionend` still produces one diagnostic sample and one committed syllable. |
For remote evidence, `live` means the owning host reported the current
verification session or process identity. `exited` requires positive
host-owned evidence that the same identity terminated or is absent. Any
transport failure, stale identity, timeout, or inability to ask the owning host
makes the result `unverifiable`; it is never evidence that the composition or
PTY process exited.
## Code elegance gate
Each fix must pass all of these checks:
- Reuse the component that already owns the state or overlay; do not install a
second composition state machine.
- Make browser, renderer, and PTY ownership boundaries explicit. Provisional
text must not leak into committed state or PTY input.
- Route local selection, pasted-image temp saves, and SSH upload results through
one resolved-attachment sink; queue semantic paths, never whole draft snapshots.
- Keep correctness changes separate from unrelated micro-optimizations.
- Use bounded per-composition state and work. Dispose every listener, timer,
observer, and DOM node with its owner.
- Preserve ordinary Latin input, mixed styled terminal content, local and SSH
PTYs, preview terminals, and folder workspaces with paired negative tests.
- Keep platform quirks behind event contracts or runtime platform checks; do
not branch on an IME vendor, language, or terminal agent name.
- Treat the canonical xterm source patch as the only hand-edited source, then
regenerate its bundle patch and lockfile together.
- Prefer deterministic replay or state-transition tests. Native evidence is a
second layer, never a substitute for regression coverage.
## Enter in application text fields (#25035)
Use `Input`, `Textarea`, or `CommandInput` for styled fields. Existing unstyled
fields with keyboard actions use `ImeInput` / `ImeTextarea` from
`lib/ime-text-field.tsx`; those preserve the DOM element, styles, refs, and
composition callbacks. They share `useImeEnterGestureOwnership` and keep
IME-owned keys out of both field actions and bubbling form/menu shortcuts.
Overlay primitives also reject IME-marked Escape in document capture, where
field-level propagation guards cannot intercept dismissal.
Do not add a second tracker at a call site already using a guarded field.
Native Chat and the File Explorer inline name field retain their existing
trackers because they also own specialized composition or element lifetimes.
Required cases:
- `isComposing`, `keyCode: 229` without `isComposing`, and `Process/229` must
never submit, choose a suggestion, or dismiss the field.
- The unmarked Enter redispatch stays owned on either side of keyup, including
a `Process/229` release. A
subsequent ordinary typing/navigation key ends that carry immediately;
hidden renderers may defer animation frames, and typing a filename suffix
must not cause the next deliberate Enter to disappear.
- Composition callbacks, blur, refs, and keyed remount cleanup still work.
Normal Enter, modifier submits, and Shift+Enter newlines remain available.
- Test the actual shared field when a consumer delegates IME handling to it;
a mock that replaces `CommandInput` with a raw input removes the protection.
`ime-text-field.test.tsx` covers primitives, raw fields, parent handlers, and
command selection. File Explorer component tests cover all three operations
and input replacement. `file-explorer-ime-enter.spec.ts` drives Chromium
composition in New File, New Folder, and Rename in a folder workspace, then
checks the complete name in the Explorer and on disk, with both continued typing
and a redispatch followed by deliberate Enter. Overlay tests cover IME Escape
and ordinary dismissal; Markdown tests preserve an unmarked save shortcut while
composition state lingers. These are CDP event
contracts, not native OS keyboard evidence.
The audit also covers settings and title fields, issue/review creation and
pickers, comments and annotations, search fields, Native Chat questions,
notebook execution shortcuts, and Markdown menu handlers. Terminal input keeps
its existing xterm/PTY ownership; mobile native fields use `onSubmitEditing`
instead of desktop DOM keydown actions. Remote workspaces use the same renderer
fields; file-operation routing and mixed-version wire contracts are unchanged.
-182
View File
@@ -1,182 +0,0 @@
# Linux glibc Compatibility
Orca's Linux builds target **stock Ubuntu 20.04 and newer** — glibc 2.31 and
libstdc++ `GLIBCXX_3.4.28` (also Debian 11, RHEL 9), on both x64 and arm64.
Packaging enforces this floor automatically; keep it in mind when adding or
upgrading native dependencies. (The optional speech feature is the one
exception — see below.)
## Local package build prerequisites
`pnpm run build:linux` produces AppImage, deb, and RPM artifacts. The RPM target
requires `rpmbuild` on `PATH`; install `rpm` on Ubuntu/Debian, `rpm-build` on
Fedora/RHEL, or `rpm` through Homebrew on macOS, then verify it with
`rpmbuild --version` before packaging. Cross-host builds have the same
requirement.
## Why this needs attention
A native module (`.node`) links against the glibc of the machine that compiled
it. Our release CI compiles node-pty from source on GitHub's `ubuntu-latest`
runner, whose glibc rises over time as the image is bumped. A binary compiled on
a newer glibc can reference symbol versions that do not exist on an older target,
and the dynamic loader then refuses to load it:
```
/lib/x86_64-linux-gnu/libc.so.6: version `GLIBC_2.34' not found (required by .../pty.node)
```
Because the Orca main process loads node-pty at startup, that failure crashes the
whole app before a window appears — this is exactly what shipped in v1.4.150 and
broke launch on Ubuntu 20.04 ([#9902](https://github.com/stablyai/orca/issues/9902)).
The specific trap is glibc's 2.32–2.34 "libpthread/libutil merge", which moved
several long-stable functions into libc under brand-new symbol versions:
| Symbol | New version | node-pty use |
| ----------------- | ------------ | ----------------------- |
| `pthread_sigmask` | `GLIBC_2.32` | reset child signal mask |
| `openpty` | `GLIBC_2.34` | allocate the pty |
| `forkpty` | `GLIBC_2.34` | fork the shell |
Electron itself (glibc 2.25) and the other bundled native modules
(`sherpa-onnx`, `@parcel/watcher`, both prebuilt on old glibc) stay well under
the floor, so node-pty was the sole blocker.
## How we keep the floor
**1. Pin the relocated symbols (the fix).**
[`config/patches/node-pty@1.1.0.patch`](../../config/patches/node-pty@1.1.0.patch)
adds a `.symver` shim in `src/unix/pty.cc` that binds `openpty`, `forkpty`, and
`pthread_sigmask` to their pre-merge version node — `GLIBC_2.2.5` on x64,
`GLIBC_2.17` on arm64 (each architecture's baseline glibc). glibc still ships
those as compatibility aliases, so the reference resolves on both new build hosts
and old targets.
The catch: gcc defaults to `--as-needed` and, since the pinned symbols now
resolve from libc's compat aliases at build time, it drops `libutil`/`libpthread`
from `DT_NEEDED`. On the target those libraries are where the symbols actually
live, so the patch's `binding.gyp` `ldflags` force
`-Wl,--no-as-needed,-l:libutil.so.1,-l:libpthread.so.0` back into `DT_NEEDED`.
The shim is guarded by `#if defined(__linux__)`; macOS and Windows are untouched.
**2. Gate packaging (the regression guard).**
[`config/scripts/verify-linux-glibc-floor.cjs`](../../config/scripts/verify-linux-glibc-floor.cjs)
runs in the electron-builder `afterPack` hook for Linux. It reads every bundled
native binary's version needs (`objdump -p` "Version References" — the
authoritative load-time list, which also captures symbol-less markers like
`GLIBC_ABI_DT_RELR`) and fails the build if any strong `GLIBC_`/`GLIBCXX_`/
`CXXABI_` node is newer than stock Ubuntu 20.04 provides, naming the file and the
offending node. Weak needs are ignored (the loader tolerates them). It also
asserts the flip side of the `.symver` fix: any binary that imports
`openpty`/`forkpty` must keep `libutil.so.1` in `DT_NEEDED` — otherwise the
pinned `openpty@GLIBC_2.2.5` resolves from libc's compat alias at build time (so
the version check passes) yet fails to load on 20.04, where those functions live
only in libutil. A future runner bump, a new native dependency, or a dropped
ldflag therefore fails the release build instead of shipping a Linux app that
crashes on launch.
> The gate is a static invariant, not an integration test. The load path was
> verified by hand for this fix (real Ubuntu 20.04, x64 + arm64: `require`
> node-pty and spawn a shell). A CI smoke test that loads the packaged
> `pty.node` in a glibc-2.31 container and spawns a shell is the recommended
> follow-up — it would make the load path self-verifying and stay valid even if
> the build ever moves to an old-glibc sysroot.
The one carve-out is the `sherpa-onnx` speech prebuilt, which already requires
`GLIBCXX_3.4.29` (GCC 11). It loads lazily in the speech worker
(`src/main/speech/stt-worker.ts`), never at app launch, so it is exempt from the
libstdc++ floor — its glibc needs are still checked. Speech-to-text therefore
needs a host with libstdc++ from GCC 11+ (Ubuntu 21.10 / 22.04 LTS or newer); the
app itself still launches on stock 20.04.
**3. Check before loading, on hosts that ship without a compiler (`orcad`).**
The two gates above protect the packaged desktop app, where the binary is built and
verified by the same pipeline. `orcad` is deployed to hosts Orca never built on, so it
adds a runtime precondition
([`src/main/orcad/node-pty-precondition.ts`](../../src/main/orcad/node-pty-precondition.ts)),
run from `main.ts` before anything requires `node-pty`. It loads the addon in a **child
process**, so a binary the loader refuses — or one that aborts outright — is data rather
than this process's death, and the operator gets a sentence naming the host's libc, its
Node ABI, its prebuild slot and the command to run. A proven-unloadable binary exits 78
(`EX_CONFIG`) instead of reaching the `require`; a probe that never answered is reported
as unverifiable and boots anyway, because a silent probe is not evidence. Whatever it
finds is published in `status.get`'s `degradations[]` under `terminal_unavailable`.
**4. Ship the binary, built from patched sources.**
[`config/scripts/build-orcad-prebuilds.mjs`](../../config/scripts/build-orcad-prebuilds.mjs)
(`pnpm run build:orcad-prebuilds`, before `build:orcad`, which copies its target's slot into
the package's `node_modules/node-pty/build/Release`) compiles node-pty for the current
host against the pinned Node's hash-verified headers at N-API 8, and files it under
`out/orcad-prebuilds/<slot>/`, where a slot is `linux-{x64,arm64}-{glibc,musl}`,
`darwin-{x64,arm64}` or `win32-{x64,arm64}`. Its `manifest.json` records each file's
sha256, the N-API level and, for glibc slots, the highest `GLIBC_` version the binary
needs; the loader checks N-API, libc, arch and that glibc version before it installs a
slot. glibc slots are built in a `manylinux_2_28` (AlmaLinux 8) container and pass the
same gate at a **glibc 2.28 / `GLIBCXX_3.4.25`** floor instead of the desktop's 2.31, because
the pinned Node they ship beside already runs on 2.28 and a 2.31 slot would leave 2.28–2.30
hosts with a runtime but no terminal (design D6). The container's gcc-toolset supplies C++20
and links newer libstdc++ symbols statically, so the slot needs only RHEL 8's system
libstdc++. musl slots skip the gate, since they never meet glibc's libraries. libc is part
of the slot name because node-pty's own loader falls back to `prebuilds/<platform>-<arch>`
and cannot tell glibc from musl — a glibc binary parked there is loaded on Alpine and dies at `dlopen`.
The script refuses to compile a tree where `config/patches/node-pty@1.1.0.patch` is not
applied: without the patch the prebuilt is a #9902 crash shipped as an artifact rather
than a first-connect error. CI runs it once per slot inside the matching container
(`--slot=` forces the label), merges the trees, and `--require-slots` fails a release with
a hole in the matrix; `--require-slots <slot>` checks one slot's files against their
hashes and `--smoke` loads it under the pinned Node and spawns a PTY
(`.github/workflows/node-server-tests.yml` runs both on every slot's runner).
The opt-in `linux-x64-glibc217` compat slot (rung B, not part of the default matrix) is
built in `manylinux2014_x86_64` (glibc 2.17, devtoolset C++20) with `-static-libstdc++`,
gated at a glibc 2.17 floor, refused if `libstdc++.so`/`libgcc_s.so` remains in
`DT_NEEDED`, and smoked under the unofficial glibc-217 Node pinned in
`NODE_RUNTIME_COMPAT_ASSETS`. Nothing installs it yet: the loader and the SSH deploy
still choose only default slots.
## Adding or upgrading a native dependency
- Prefer packages that ship prebuilt binaries compiled against an old toolchain
(manylinux / `glibc 2.17`-class), like `@parcel/watcher`.
- For a module we compile from source, if the gate flags it, either pin the
offending symbols the way node-pty does, or build it in an old-glibc container.
- To check locally on a Linux host, list what a binary requires (skipping the
weak `0x02`-flagged needs the loader tolerates):
```bash
objdump -p path/to/module.node | sed -n '/Version References/,/^$/p'
```
No strong `GLIBC_` node may exceed `2.31`, and no `GLIBCXX_`/`CXXABI_` node may
exceed `3.4.28`/`1.3.12` — what stock Ubuntu 20.04 ships.
## Runtime floor: the `environ` race below glibc 2.41 (Electron ≥ 43.7.0)
Separate from the build floor above, one glibc runtime bug constrains which
Electron we may ship. Before glibc 2.41, `setenv`/`unsetenv` reallocate the
`environ` array and **free** the old one, so a concurrent `getenv()` on another
thread reads freed memory. Ubuntu 20.04–24.04 (2.31–2.39) are all below that
line, so every Linux target we support is exposed.
Electron 43.5.0 made that latent race reachable on every launch: it started
setting `GDK_GL=disable` around `gtk_init()` and unsetting it right after, while
in the same change moving FontConfig warm-up onto a thread-pool thread that runs
concurrently and calls `getenv()` constantly
([electron#53070](https://github.com/electron/electron/pull/53070)). The result
is a browser-process use-after-free about a second into startup — no window, no
GPU child involved, and the corruption surfaces wherever the next allocation
lands, which is why reports name unrelated frames (`gtk_widget_realize`,
libxcb-dri3, FontConfig/expat). Orca 1.4.199/1.4.200 shipped that runtime and
died on launch on Ubuntu + NVIDIA/X11
([#20081](https://github.com/stablyai/orca/issues/20081)).
Electron 43.7.0 fixes it by overriding `setenv`/`unsetenv`/`putenv`/`clearenv`
so a published `environ` is never freed, deferring to glibc on 2.41+
([electron#53491](https://github.com/electron/electron/pull/53491), backported
to 42/43/44/45). **Do not downgrade Electron below 43.7.0, or move to another
line, without confirming that backport is in the target release** —
`config/scripts/electron-runtime-floor.test.ts` fails the suite if the pin drops
below the floor. Orca itself writes `process.env` during early startup
(`patchPackagedProcessPath`, `configureOrcaUserDataPathEnv`,
`hydrate-shell-path`), so it is a first-class trigger, not just a bystander.
-58
View File
@@ -1,58 +0,0 @@
# macOS press-and-hold and key repeat
macOS opens the accent picker when a key is held unless an application opts out in its preferences
domain. That prevents held keys from repeating in terminal applications such as vim. On the first
eligible launch, Orca writes:
```sh
defaults write com.stablyai.orca ApplePressAndHoldEnabled -bool false
```
The write is scoped to Orca's packaged bundle domain. Bare Electron development bundles and
non-macOS platforms are left untouched. A fresh write is conservatively treated as taking effect
on the next launch.
## Precedence and decision record
Orca checks for an explicit domain value before writing. Either `true` or `false` is treated as a
user choice and preserved. Only an unset key receives the `false` default.
The decision is stored once in
`<userData>/macos-press-and-hold-default.json`. An `applied` or
`kept-user-preference` decision prevents future launches from touching the domain again.
Probe and write failures remain retryable so a transient failure does not permanently disable the
fix.
`defaults read <domain> <key>` is used instead of
`systemPreferences.getUserDefault`: the Electron API cannot distinguish an unset key from an
explicit `false`. Only the missing-key exit status is interpreted as unset; spawn failures,
timeouts, and other exit statuses leave the preference alone.
## Restoring the accent picker
Set the preference explicitly, then restart Orca:
```sh
defaults write com.stablyai.orca ApplePressAndHoldEnabled -bool true
```
After Orca has recorded its one-time decision, deleting the key also restores the macOS default
without Orca recreating it:
```sh
defaults delete com.stablyai.orca ApplePressAndHoldEnabled
```
Development and prerelease channels may use a channel-suffixed Orca bundle identifier; use that
domain instead when applicable.
## Reverting
Deleting the startup code is not enough. AppKit reads the persisted preference, so a code revert
must also arrange to delete the key for users who ran an affected build.
## Test coverage
Unit tests cover platform guards, explicit-value preservation, retry behavior, domain ownership,
record persistence, and subprocess exit interpretation on CI. A macOS-only test additionally pins
the real `defaults(1)` behavior, but current PR CI does not execute tests on macOS.
@@ -1,43 +0,0 @@
# Malformed worktree registration removal
Git can report a linked worktree at `<checkout>/.git` when its administrative
`gitdir` backlink incorrectly ends in `.git/.git`. That reproduces #17316's
validation error. The reproduction establishes the malformed registration, not
which program created it; current OMP uses ordinary `git worktree add`.
Orca's desktop and runtime removal entry points use registration-only recovery
when Git positively marks the row prunable, the row has a named local branch and
HEAD, it is neither main nor locked, and the execution filesystem confirms the
selected `.git` path is a regular file. Missing or unknown evidence does not
permit this recovery. A symlink or directory is not a regular-file proof.
Recovery reuses `git worktree prune` followed by a strict worktree listing that
must confirm the selected registration is gone. It does not delete the selected
file, infer a parent path for deletion, or delete the branch. Archive hooks and
checkout teardown are skipped because the selected row is not a checkout.
Two consequences are intentional:
- Git's prune also clears other stale, unlocked registrations in the repository;
it is not a path-scoped command. Live and locked registrations remain Git's
responsibility, and Orca verifies that the requested registration disappeared.
- The surviving checkout's `.git` file points at removed administrative metadata.
Files and its named branch are preserved; recovery removes the broken navigation
entry and does not repair or claim to restore that checkout.
Native and WSL checks use the existing execution-filesystem accessor. WSL prune
and verification use the same selected distro. Paired runtimes run the recovery
on their owning host. Direct SSH does not enter this local recovery: its current
provider has no registration-only removal operation, and a failed remote removal
never authorizes a local fallback.
The Git commands already exist in the 2.25-compatible cleanup path. On an older
Git that cannot positively attest this file-shaped registration as prunable, Orca
refuses this recovery. Deferred deletion independently rejects non-directory and
symlink targets, so force cannot move a `.git` file into deletion trash.
Regression coverage is in `worktree-prunable-git-file.test.ts`,
`worktrees-removal-recovery.test.ts`, and
`worktree-deferred-removal-real-git.test.ts`. The latter reproduces the exact
malformation against the installed Git binary in a disposable repository and
checks surviving file contents, branch HEAD, and removed registration.
-23
View File
@@ -1,23 +0,0 @@
# Managed OpenCode and Devin accounts
Run enrollment in a terminal on the machine running Orca:
```sh
orca account add --agent opencode --label Work
orca account add --agent opencode --integration opencode-go --label Work
orca account add --agent devin --label Work
orca account list --agent opencode --json
orca account select --agent opencode --account <id>
orca account select --agent opencode --account system
orca account rm --agent opencode --account <id>
```
OpenCode enrollment requires OpenCode 2 and runs its official `auth login --standalone` command. Devin runs `auth login --force-manual-token-flow`; obtain the enrollment token through Devin's supported login flow. These commands neither reuse a guessed token nor sign out the system account. Settings → AI Provider Accounts provides the enrollment command, refresh, selection, and removal for the selected Orca host.
Each profile belongs to the execution host. OpenCode's SQLite credentials and Devin's credential TOML stay in private Orca user-data directories. Enrollment isolates XDG data/config/cache/state, copies only authenticated credentials, and then deletes the temporary directory. OpenCode capture rejects databases containing conversations and includes SQLite WAL contents. RPC summaries contain labels, IDs, and integration names, never tokens or credential paths. Only the authenticated local runtime socket can import a credential directory; paired clients cannot ask the host to read arbitrary paths.
Selection affects newly launched explicit OpenCode/Devin commands and agent launches. It redirects XDG data and state; OpenCode inline-auth/database overrides cannot bypass the profile. Shell wrappers restore this selection after user startup files. Existing provider configuration and environment-based integrations remain available. Running terminals retain their current profile. Stop agents before removing a profile: removal also deletes conversations created in that private profile, without changing the system login.
For SSH, enroll by running the command on a headless Orca runtime on the remote machine. The remote runtime owns its profiles and selection; a desktop client's credential paths never cross SSH. Direct SSH relay launches and Windows-hosted WSL panes do not consume the desktop host's profiles. Run a headless runtime inside that execution environment instead. Folder workspaces use the same host account store as git worktrees. Older Orca hosts reject new operations before login through capability negotiation.
Validation covers OpenCode 2.0.16 on macOS and Linux arm64, including isolated official enrollment, selected and System background-terminal credential checks, reselection, and profile deletion. Linux checks used the Node headless runtime in an Ubuntu 24.04 container. Devin 3000.10.31 saved-login recognition was checked on macOS. Fresh Devin manual-token enrollment, a physical SSH host, Linux desktop UI, Windows, and Windows-hosted WSL still require verification.
@@ -1,34 +0,0 @@
# Monaco filename associations
Orca imports Monaco's full `editor.main.js` entry point, which registers the built-in
languages and loads their grammars on demand. Filename detection must not load the
editor itself: it also runs during session restoration and before the editor mounts.
`monaco-language-associations.json` contains the registration metadata from the
installed package's entry point, plus curated Ruby associations in the generator.
Those add `.rake`, `.ru`, `.jbuilder`, `.thor`, `Guardfile`, `Capfile`, `Podfile`,
`Brewfile` and `Vagrantfile` to the existing Ruby grammar. Change these in
`config/scripts/generate-monaco-associations.mjs`, not the generated JSON.
Regenerate after changing the curated associations or upgrading Monaco:
```sh
node config/scripts/generate-monaco-associations.mjs
pnpm exec oxfmt --write src/renderer/src/lib/monaco-language-associations.json
```
The generator reads syntax trees without executing contributions or grammar loaders.
Its test compares the checked-in metadata to the installed package and curated associations. The original
list was verified against a clone of `microsoft/monaco-editor`, tag `v0.55.1`, commit
`516f350bdaf7a82f6731bd128a9ec86a6e5fa47d` (`src/basic-languages` and `src/language`).
Existing Orca filename and extension choices take precedence. This preserves custom
Vue, Svelte, Astro, Nim, Typst, JSONL, notebook and preview handling, as well as the
Markdown mapping for MDX. The fallback matches exact filenames before the longest
extension, case-insensitively, and resolves duplicate associations in upstream
registration order (the last registration wins, so `.pp` selects Ruby over Pascal).
Monaco 0.55.1 has 90 registrations, 81 with filenames or extensions. Registrations
without either remain available in Monaco but cannot be inferred from a path. This
does not add VS Code extensions or language servers, guess from file contents, or
change mobile's separate lowlight grammar set. The renderer uses only the basename,
so local, Windows, SSH and folder-workspace paths share the same detection.
-51
View File
@@ -1,51 +0,0 @@
# Fresh OMP launches
Orca's new-session and draft launch plans apply a one-time `--config` overlay
containing `autoResume: false`. OMP's configured session directory, settings,
authentication and extensions remain in their usual locations. Saved launch
configuration omits the overlay so explicit resume keeps its normal semantics.
Custom commands with session selectors, unknown flags, positional arguments or
shell compounds are left unchanged.
The execution host creates the overlay. Local and WSL terminals use Orca userData
(with WSLENV path translation); SSH relays use their own managed directory. Config
creation is independent of status-hook preferences and does not require plugin
source installation. An unavailable file produces a terminal diagnostic and skips
OMP. Filesystem failures do not prevent unrelated agents or bare shells starting.
The guard invokes OMP in the current shell, preserving functions, aliases and the
managed status wrapper. A nonzero agent exit never triggers a second launch.
POSIX commands also support fish; environment presence is checked before expansion
so an old host under `set -u` reports the same missing-settings diagnostic.
## Mixed versions
The existing command and environment transport carries the launch unchanged; no
new RPC or stream opcode is introduced. A new host exports the path before shell
startup, including bare shells that receive an OMP command later. An old client
continues its existing launch behavior on a new host. A new fresh-launch command
on an old host without the managed environment fails visibly and requests a host
update and terminal restart. It must not silently fall back to OMP auto-resume.
## Verification
`src/shared/omp-fresh-launch-shell.test.ts` runs actual available bash, zsh and fish
shells, checking exact argv, a single invocation, nonzero exit, deleted settings
and an absent environment variable. `src/relay/omp-fresh-launch-environment.test.ts`
checks the guarded command through relay environment assembly and the actual OMP
shell wrapper, retaining both extension and config arguments and prefill.
Local host assembly tests recognize guarded POSIX, cmd and PowerShell commands.
Run the actual OMP storage smoke against a read-only OMP checkout:
```sh
ORCA_BACKGROUND_LAUNCH=1 bun tests/tools/omp-fresh-session-runtime-smoke.mjs /path/to/oh-my-pi
```
For Windows, bundle `tests/tools/omp-fresh-launch-windows-smoke.ts` with
`bun build --target=node --outfile=/tmp/omp-fresh-windows-smoke.mjs`, transfer the
bundle to the host and run it using Node with `ORCA_BACKGROUND_LAUNCH=1`.
The smoke uses temporary files and process-local environment only. Both cmd and
PowerShell must pass existing/missing/unset/directory settings cases, preserving
exit 17 for the single successful launch and returning exit 1 without launching
when settings are unavailable. This passed on Windows host `awin` on 2026-09-14.
-34
View File
@@ -1,34 +0,0 @@
# OMP history titles
The message-graph scanner uses persisted OMP names ahead of the first user prompt:
`session.title`, version-1 `title` slots, `title_change.title`, and legacy
`session_info.name`. Empty or unsupported metadata leaves the previous name or
prompt fallback intact. Non-OMP graph parsing keeps its existing title policy.
Explicit user names outrank automatic names. Within the same source, timestamps
prevent the current first-line title slot from being replaced by older rename
entries later in the file. Newer appended renames still update the row. Legacy
records without timestamps retain file-order handling.
The graph fold stores title authority alongside the existing accumulator. Clones
retain it without sharing mutable accumulator or preview state, while preserving
the existing identity and message-consumer contracts. Cached append parsing uses
the normal durable offset; no extra scan, process, poll or watcher is introduced.
The parser is shared by local and remote content readers and uses transcript data
from the execution host. It performs no client-side path lookup and changes no
wire shape. Folder workspaces require no git metadata.
Run actual persistence and cache validation with a read-only OMP checkout:
```sh
ORCA_BACKGROUND_LAUNCH=1 bun tests/tools/omp-history-title-smoke.mjs /path/to/oh-my-pi
```
The smoke persists a first prompt, performs a real OMP user rename, and verifies
both cold and incrementally cached scans. It checks one full parse, one append
parse and identical-object reuse on an unchanged scan. All home/config/data roots
are disposable; no model requests are made.
This is the OMP subset of the history-name behavior proposed in PR #15696 by
Brennan Benson. Pi naming and title changes in the terminal are separate concerns.
@@ -1,28 +0,0 @@
# OMP recorded transcript resume
A sleeping OMP session can retain its transcript path from a hook without an
explicit `launchConfig.ompResumeFilePath`. Both cold-restore startup and generic
sleeping-session launch already forward the provider metadata to
`getAgentResumeArgv`; that builder must keep the recorded path.
Resolution order is explicit launch path, recorded transcript path, then UUID.
The existing shell-aware builder quotes the selected argument for the execution
host. An older metadata record without a path retains UUID fallback. OMP provider
claim keys and equality remain UUID-based, so a later hook adding the path does
not create a second automatic-resume identity.
`tests/tools/omp-resume-transcript-locator-smoke.mjs` creates an actual OMP session
outside its default session store. UUID lookup fails there; the absolute path and
Orca's generated argv resume the original session. Run it with Bun and a read-only
OMP checkout as argv[2], under `ORCA_BACKGROUND_LAUNCH=1`. It uses a disposable home
and makes no model requests.
This bounded correction follows the resume-locator portion of
[PR #16276](https://github.com/stablyai/orca/pull/16276) by @CodeHourra. It does not
adopt that PR's reattach injection or title changes. The reattach proposal treats
missing snapshot/replay as permission to type a resume command, but
`daemon-pty-spawn-result.ts` explicitly permits `isReattach: true` without a
snapshot. That payload absence is not positive evidence of a newly created shell.
The proposal also adds the path to OMP claim identity, which separates UUID-only
metadata from a later path-enriched record for the same provider session. Those
changes require separate evidence and are outside this patch's review scope.
@@ -1,23 +0,0 @@
# OMP runtime session provenance
OMP computes whether a runtime session is a task child, but the released
`ExtensionContext` does not expose that value. The status extension therefore uses
the session manager's parent header and nested task transcript path only when a
root owner is already known. A nested transcript with no known owner remains
eligible because it may have been resumed directly as the pane's main session.
The remaining child-first case is inherently ambiguous to Orca: task children and
resumed child transcripts have the same public session-manager shape. A complete
child-first fence requires OMP to expose its computed `agentKind` through
`ExtensionRunner.createContext`; until then the conservative fallback avoids
silencing valid resumed sessions.
Older runtimes retain the manager-identity guard. That guard assumes the main
session reaches Orca's callback before any child. An earlier user extension can
initialize a child during session_start and violate that assumption. Keep the
ownership merge assessment conditional until the runtime API is available and the
combined flow is validated. Neither callback timeouts, UI presence, nor transcript
paths establish runtime ownership.
The guard remains scoped to one pane and launch token. It does not define how
several independent SDK/ACP roots sharing one process and pane should be attributed.
-53
View File
@@ -1,53 +0,0 @@
# OMP transcript roots
Native chat lookup, history discovery and the history path allowlist use
`src/main/ai-vault/omp-session-root.ts` on the execution host. The resolver follows
the active runtime environment rather than scanning every profile.
The behavior matches OMP's `packages/utils/src/dirs.ts`:
- On Linux and macOS, an existing `$XDG_DATA_HOME/omp` selects that app's `sessions`
directory, even when legacy transcripts coexist. The sessions directory itself
need not exist yet. There is no implicit `~/.local/share` fallback.
- A named profile uses XDG only when its own `omp/profiles/<name>` path exists;
otherwise it uses the profile's config-root `agent/sessions` directory.
- `OMP_PROFILE` takes precedence over `PI_PROFILE`, including an explicitly empty
canonical value. Named profiles ignore custom `PI_CODING_AGENT_DIR` values.
Default mode respects custom agent directories, except an inherited agent path
derived from the lower-priority profile. `PI_CONFIG_DIR` selects the config root
relative to the owning host's home, as upstream specifies.
- Orca retains its legacy `OMP_CODING_AGENT_DIR` sessions-root override and prefix
normalization. Explicit scan roots override environment discovery. Empty or
filesystem-root scan overrides and invalid profile names refuse discovery;
they never fall back to a different profile or the process working directory.
The desktop scanner child allowlist forwards only the required directory/profile
variables. SSH relay discovery still builds legacy roots from its host-owned home
and does not gain XDG/profile discovery here. Client XDG/profile values are not
applied to WSL home roots. Exact hook
paths and existing WSL attestation/refusal remain authoritative. No wire fields
or opcodes change; older clients receive the existing session record shape.
This resolves environment-visible configuration. Per-command `--profile` choices
or directory values loaded only inside the agent are not inferred by a runtime
that never received them; hook-reported transcript paths remain the exact route.
Run the read-only upstream parity smoke with:
```sh
ORCA_BACKGROUND_LAUNCH=1 bun tests/tools/omp-session-root-upstream-smoke.mjs /path/to/oh-my-pi
```
The smoke uses disposable home/data roots and compares Orca's result with OMP's
actual directory resolver. It makes no model requests. Unit tests also cover
Windows XDG exclusion, legacy override normalization and refusal paths.
For actual persistence-to-reader validation, run:
```sh
ORCA_BACKGROUND_LAUNCH=1 bun tests/tools/omp-transcript-root-reader-smoke.mjs /path/to/oh-my-pi
```
This creates default and named-profile transcripts through OMP's SessionManager,
with legacy directories still present, then resolves and decodes each by session
ID through Orca's native reader. All files use disposable roots; no model runs.
@@ -1,15 +0,0 @@
# OMP startup keyboard capability query
OMP's ProcessTerminal sends `CSI ? u` and then a DA1 sentinel before selecting its keyboard encoding. A fresh desktop terminal already advertises Kitty support, but a direct New Tab launch can query before that renderer owns replies. Previously startup ingress answered only OSC color queries.
The renderer now supplies optional `terminalKittyKeyboardProtocol: true` from its actual xterm `vtExtensions.kittyKeyboard` setting. The existing local/SSH/paired spawn route places it in startup ingress as optional `kittyKeyboardProtocol`. Missing or false capability leaves behavior unchanged, including panes that deliberately withhold Kitty on native Windows. The ingress version and stream opcodes do not change. Terminal creation accepts additive fields, but host-authoritative `terminal.createAgentSession` and `terminal.ensureAgentSession` use strict schemas. Clients send the keyboard flag on those methods only after the host advertises `agent-session.keyboard.v1`. The negotiated payload stays fixed across launch retries; old hosts receive the original payload and retain the renderer fallback. New hosts accept older clients that omit the flag. Paired background launches use the same negotiated support and default paired-terminal advertisement; legacy terminal creation receives the additive flag.
Source ingress answers only the exact first `CSI ? u` before its deadline/renderer handoff. It uses the existing mode tracker for preceding flag pushes and the existing reply-delivery echo guard. Its transformed source span consumes the query once, while the following DA1 and Kitty mode-setting bytes retain their sequence ranges and reach the renderer. Keyboard intent does not require theme colors. Color and Kitty authority end independently: answering both colors does not end Kitty handling, and ConPTY's persistent color ownership does not retain Kitty ownership after handoff.
Run the actual OMP protocol smoke with a read-only reference checkout:
```sh
ORCA_BACKGROUND_LAUNCH=1 bun tests/tools/omp-startup-keyboard-smoke.mjs /path/to/oh-my-pi > /tmp/omp-startup-keyboard.json
```
This uses OMP's real ProcessTerminal with intercepted process-local stdin/stdout, a disposable HOME, and no model call. It verifies negotiation before any renderer attaches, a single reply, preserved mode push, and contiguous raw sequence coverage. It does not constitute live Windows/SSH or rendered shortcut proof.
@@ -1,39 +0,0 @@
# Pi/OMP status tool-input redaction
Generated tool_call and tool_execution_start hooks sanitize inputs on the agent's
execution host before passing them to the existing status transport. Inputs that
reference `.ssh`, `.ssh-mcp`, `.mcp-secrets.env`, or
`.omp-backups-archive/omp-bak-keyfile` become `{ redacted: true }`. Matching includes
nested values, property names, Windows separators, case variants and shell token
boundaries. Sibling names such as `.ssh-backup` remain ordinary data.
The sanitizer copies data descriptors into objects without prototypes. It never
passes the source object's toJSON or getters to the transport. Cycles, accessors,
class instances, symbols, functions and non-JSON primitives redact the whole input.
Repeated ordinary object references are allowed. Depth, visited values, reserved
array slots (including holes) and inspected text have conservative bounds to avoid
moving unbounded work into synchronous JSON serialization.
This policy targets credential-path references in status tool inputs. It does not
scan transcript files, tool outputs, prompts or arbitrary secret values, and does
not erase previously persisted status. Proxy reflection traps can still run when
JavaScript inspects a proxy; the sanitizer is not an isolation boundary against a
malicious extension in the same process.
Ordinary question envelopes and preview input shapes remain unchanged. Older
clients already accept object-valued tool_input; no RPC fields or opcodes change.
The same generated code runs on local, WSL and SSH agent hosts; no local filesystem
lookup or substitution is introduced. Folder workspaces require no special path.
This follows PR #9554's credential-reference policy, with descriptor copying and
bounded serialization correcting its validation-then-original-object approach.
Run the actual OMP loader/native HTTP smoke with a read-only checkout:
```sh
ORCA_BACKGROUND_LAUNCH=1 bun tests/tools/omp-status-input-redaction-smoke.mjs /path/to/oh-my-pi
```
It loads Orca's generated extension through OMP, invokes synthetic tool events,
and inspects three real loopback HTTP payloads. Home/config/data roots are
disposable; it makes no model requests and does not claim an interactive tool run.
-399
View File
@@ -1,399 +0,0 @@
# Running orcad
`orcad` is the Orca runtime served from plain Node. This is the contract between it and
whatever supervises it: what it binds, what it owns on disk, who restarts what, and what its
readiness payload actually proves.
## Two long-lived processes, not one
A deployment is **orcad** plus **the terminal daemon**.
| | orcad | terminal daemon |
| ---------- | -------------------------------- | ------------------------------------- |
| Started by | the supervisor | orcad, detached |
| Owns | RPC, git, worktrees, persistence | every local PTY |
| Lifetime | one supervised run | detached from orcad, not its service |
| Endpoint | `ws://<bind>:<port>` | `<data-root>/daemon/daemon-v<N>.sock` |
orcad detaches the daemon and calls `disconnectDaemon()`, never `shutdownDaemon()`. The
built-in remote deployment path stops only the recorded orcad PID, so the daemon and its PTYs
survive. The successor adopts the current endpoint and routes supported previous protocol
versions through legacy adapters. This makes a PID-scoped update, rollback or restart
non-destructive to live work.
Process detachment is not service isolation. A daemon that orcad launches directly, and every
PTY it owns, remain in the same systemd service cgroup. `KillMode=mixed` does **not** preserve
them: it sends the graceful stop signal only to the main process, then sends `SIGKILL` to every
process remaining in the cgroup the moment that main process exits — `TimeoutStopSec` never gets
the chance to apply. `KillMode=control-group` is destructive too. `KillMode=process` leaves
service-owned processes unmanaged and is not a supported preservation mechanism.
Service-restart survival therefore requires a separately supervised cgroup, and orcad now asks
for one: on Linux it launches the daemon through `systemd-run --user --scope`, which places the
daemon and its PTYs in their own transient `orca-daemon-<launch-nonce>.scope` unit under the
user slice instead of the caller's service cgroup. A stop or restart of the service unit then
leaves that scope — and the live terminals in it — running, and the successor adopts the
endpoint as it always has.
A newly launched private daemon scope also follows the daemon's own lifetime. A small
detached shell holds an input pipe from the daemon; after that pipe closes and `/proc`
confirms the daemon PID is gone, it asks the user manager to stop that exact scope.
Systemd sends remaining processes SIGTERM and escalates after five seconds. This includes
children that double-forked or called `setsid` and can no longer be found by parent PID.
Disconnecting or restarting the runtime does not close the pipe: the daemon owns it.
The cleanup only arms on a fresh scoped launch with a matching launch nonce. Adopted
legacy scopes can contain GUI processes and are never armed retroactively. Unscoped
launches and children deliberately moved into another systemd unit remain outside this
cleanup. `nohup`, `disown`, and `tmux` alone do not move a process out of its cgroup, so
those children now end when their terminal daemon dies. Work intended to outlive that
daemon needs its own service or scope.
The scope is requested only where it can work. All of these must hold:
- **Linux with systemd as PID 1** (`/run/systemd/system` exists).
- **A reachable user bus** — a connectable `bus` socket in the per-UID runtime dir
(`/run/user/<uid>`, or whatever `XDG_RUNTIME_DIR` points at). For a service account that is
not otherwise logged in, that means `loginctl enable-linger <user>`; a unit whose
`RuntimeDirectory=` hardening moves `XDG_RUNTIME_DIR` off the per-UID path is handled, because
the real per-UID path is probed first.
- **`systemd-run` on `PATH`** and answering `--version`.
Any of those missing, or a `StartTransientUnit` call that fails anyway, falls back to the
direct launch — and in that unscoped fallback case the paragraph above still describes reality:
the daemon shares the service cgroup and a combined-unit stop ends live terminals. Read the
`cgroupUnit` field in the daemon health payload to tell the two cases apart on a running host;
it is populated from `/proc/self/cgroup`, so it reports the isolation the daemon actually has
rather than what the launcher intended.
## `orca serve` on this machine
`orca serve` runs on the local orcad slot by default. The CLI asks the app's
`out/main/orcad/orcad-local-serve-selection-entry.js` (run as plain Node on the app's executable)
which host to use. Any reason orcad cannot serve falls back to Electron serve with one
`[serve] using Electron serve: <reason>` line on stderr. Those reasons are: no slot for this host,
no template in the install, the pinned Node could not be fetched, or a failed native preflight.
- `ORCA_SERVE_RUNTIME=electron` keeps Electron serve and skips the question. `orcad` (or unset) is
the default, and any other value falls back with a reason.
- Packaged macOS stays on Electron: only Electron serve, supervised by the CLI, can take a remote
app update there, and orcad has no updater. Recipe-JSON serve has no handoff and uses orcad.
- Windows serves on orcad too. Both hosts share `<userData>\daemon`, so the daemon pipe name
(hashed from that path) is the same, and the relocated Electron daemon host changes only the
executable, not the pipe. The `orcad-serve-mode-switch-windows` e2e job checks D7 there, in the
daily run and on PRs routed to it; it does not block merges.
The slot and its pinned Node live under the desktop's `<userData>/orcad-artifacts`.
## Bind policy
`--bind <literal-ip>`, **default `127.0.0.1`**.
Only literal IPs are accepted; hostnames are refused because DNS would decide which
interface got bound. `localhost` maps to `127.0.0.1`. `0.0.0.0` / `::` are the explicit
opt-ins to network reach, and the startup log says so on every launch.
The bind is **pinned**, not defaulted. Two things widen the desktop's listener on their own —
`orca serve`'s wide default, and a startup where some device has connected before — and an
unattended host's exposure must be exactly what the operator asked for on every launch. A
mobile pairing offer, which normally rebinds to all interfaces, is refused while the bind is
pinned to loopback and reports `network_exposure_failed` rather than advertising an endpoint
nothing can reach.
Under the shipping design a client reaches a remote orcad over an SSH local port-forward, so
loopback is the correct default and the pairing credential travels over SSH.
A host whose sshd refuses forwarding (`AllowTcpForwarding no`) is reached through the stdio
bridge instead: the client keeps the same local port, and each connection to it opens one SSH
exec channel running a small script on the host's pinned Node that dials orcad's loopback port.
Windows hosts run it as the host script's `stdio-bridge` op and frame bytes as base64 lines,
because a PowerShell DefaultShell re-decodes native output. Bridges are capped below OpenSSH's
default `MaxSessions` of 10 per connection; further connections wait for a free one. The choice
is made each time the tunnel starts (`orcad-managed-tunnel-transport.ts`), so nothing is
recorded per host, and only a host where even the bridge cannot run keeps the relay, recorded as
`ssh_tunnel_unavailable`.
## Data root and the instance lock
The data root is `$ORCA_USER_DATA`, else `$XDG_DATA_HOME/Orca`, else `~/.orca`.
Before the profile index or the store is touched, orcad takes `<data-root>/orcad.lock`.
It refuses to start when:
| Code | Meaning |
| -------------------------------------- | ------------------------------------------------------------- |
| `orcad_data_root_wrong_owner` | the root is owned by another uid (POSIX) |
| `orcad_data_root_shared` | the root is group/world accessible and could not be tightened |
| `orcad_instance_lock_held` | another live orcad owns this root |
| `orcad_instance_lock_foreign_identity` | the lock belongs to a different identity |
| `orcad_data_root_unusable` | the root cannot be created, stat'd or written |
A root that is merely too permissive and that we own is tightened to `0700` rather than
refused — orcad stores credentials there unsealed (no OS keyring on this host), so the goal
is a private root, and refusing when we could just fix it helps nobody. We refuse when the
permissions are not ours to fix. Windows has no owner or mode check, because ACLs are not
expressible as a POSIX mode and `statSync().mode` there reports a synthesized one. Instead
orcad restricts the root's ACL to its own user with `icacls` (the same verified restriction
`secure-file.ts` applies to credential files) and refuses with `orcad_data_root_shared` when
that cannot be applied.
A dead holder's record is reclaimed (PID plus process start time, so a recycled PID does not
read as alive). On Windows the start time is the kernel creation time read through the
process-tree addon the slot stages; without the addon it is null and the PID alone fences,
which errs toward "held". A record belonging to a different identity is never reclaimed.
**The lock scopes one role — who is the runtime.** It deliberately says nothing about the
daemon, which lives under `<data-root>/daemon` and fences its own endpoint with its own PID
record. A lock that asked "is any process using this root" would refuse exactly the restarts
a live daemon makes worthwhile.
## Supervision
### Process-scoped and cgroup-wide stops
The built-in remote updater performs a PID-scoped stop and keeps the daemon's install version
pinned while it owns sessions. A combined-unit systemd stop or restart is different: unless the
daemon holds a durable cgroup scope of its own (see
[Two long-lived processes, not one](#two-long-lived-processes-not-one)), it reaps the daemon and
every live terminal after the graceful window. Treat a stop as destructive unless
`health.terminalDaemon.cgroupUnit` names an `orca-daemon-*.scope` on that host.
Before a cgroup-wide stop, obtain a fresh `orca-ide terminal list --json` result using the same OS
account and home as the daemon. Invoke the installer's absolute launcher path so `sudo`'s
`secure_path` cannot hide a per-user registration (for example,
`sudo -Hu orca /home/orca/.local/bin/orca-ide terminal list --json`). Replace both `orca` and
`/home/orca` with the service account and home used by the unit; an extracted deployment may use
its absolute `resources/bin/orca-ide` launcher instead. A safe empty census is untruncated, has an explicit `hostScope`, covers every
execution host affected by the stop, and lists no terminals on those hosts. Every
`omittedHostIds` entry must be explicitly accounted for outside the target service's execution
boundary. A separately paired runtime is outside that boundary; local execution and SSH hosts
reached through this runtime are not. An affected or unknown omission, missing scope,
truncation, a failed request or lost contact makes the result `unverifiable`: defer the stop. Do
not admit new work after the census. Orca does not yet provide an atomic census-and-stop fence.
### Who supervises orcad
An external supervisor (systemd, launchd, a process manager). orcad conforms to it:
- **Readiness.** One JSON line on stdout (`--json`), `type: "orca_server_ready"`, published
after the listener is bound and the daemon verdict is in. There is no separate readiness
socket; the line is the signal. Set the supervisor's start timeout generously — the daemon
launch has its own retries and can take tens of seconds on a cold host.
- **Shutdown.** `SIGTERM` or `SIGINT` starts one graceful stop. Repeated signals share
that stop because a supervisor may signal both the launcher and its child. A 15s deadline
exits with code 1 if teardown stalls. The bundled runtime also stops gracefully if its
launcher's IPC channel closes. On POSIX, both the launcher and runtime ignore `SIGHUP`,
so terminal hangups do not stop a headless host. Use `SIGTERM` or `SIGINT` to stop it.
- **Stop requests.** A file stops orcad the same way `SIGTERM` does, without a PID that may
since have been reused by another process:
- `.orcad-stop-request` beside `orcad.js` in the running slot. orcad deletes it and stops.
- An instance-bound request in the data root, named
`.orcad-managed-stop-request.<sha256 of the instance lock nonce>`. orcad stops only when it
names this orcad's version, runtime ID, PID, start time and lock nonce, and while the
instance lock still holds that record. The file is kept as evidence.
- `orcad --complete-managed-stop '<request JSON>'` writes that request, waits for the
instance to exit, and prints one JSON line whose `verdict` is `live`, `unverifiable` or
`exited`. `exited` needs proof: no process with that PID, or a PID whose start time shows
it now belongs to another process. On `exited` it writes
`<data-root>/orcad-stop-receipts/<transactionId>.json`. It exits 0 whenever it printed a
verdict, 64 for a malformed invocation, and 1 for a failure before any verdict, which is
never evidence of exit.
- A request with `retireIdleDaemon: true` asks orcad to retire the terminal daemon too. This
is best effort and never blocks or fails the stop:
- The daemon is retired only when it proves it owns no live session across every
generation.
- A busy daemon (`live`) or one whose state cannot be proven (`unverifiable`) stays up with
its terminals, and orcad reopens new-terminal admission before exiting.
- The completed-stop receipt records `retirement` as `retired`, `live` or `unverifiable`.
If orcad exits without recording an outcome, the receipt says `unverifiable`.
- `orcad --cancel-managed-stop '<request JSON>'` withdraws a request orcad has not acted on.
orcad and the canceller each try to create `<transactionId>.decision.json` exclusively,
so exactly one wins. `canceled` means orcad keeps running and the request file is removed;
`dispatched` means orcad already began stopping, and only the completion can say how it
ended.
- A build advertises all of the above with `health.stopRequests: 1` in its readiness line.
Clients stop such a build through the slot request file and older builds with `SIGTERM`,
after corroborating the PID with readiness either way. A launch clears a slot request
that the previous process never consumed.
- **Decommissioning a managed slot.** An Orca client decommissions through the same activation
journal and fence as deploy and rollback. It refuses while the terminal census reports live
or uncounted terminals, stops the instance with a managed request that also asks to retire
the daemon, and records that no version is active only after `exited` is proven. A stop
that did not finish is cancelled; if orcad already acted on it, or the host cannot answer,
the fence stays for recovery.
- **Instance lock.** `<data-root>/orcad.lock` names the running orcad. A record that is
unreadable, malformed or over 64 KiB is never reclaimed: orcad exits 78 until an operator
removes it. A shutdown whose teardown failed keeps the lock until the process exits, so a
second orcad cannot start beside a writer that may still be running.
- **Exit codes.**
| Code | Meaning | Supervisor should |
| ---- | ------------------------------------------------------------ | -------------------- |
| 0 | clean shutdown | restart per policy |
| 1 | startup or shutdown failure | restart with backoff |
| 78 | configuration fault (bind address, data root, instance lock) | **not** restart |
78 is `EX_CONFIG`. Put it in systemd's `RestartPreventExitStatus`: restarting on a data
root owned by someone else is a restart-spin, not a recovery.
- **Logs.** orcad writes human-readable diagnostics to **stderr** and its readiness contract
to **stdout**; the supervisor owns capture and rotation. The daemon, being detached, writes
its own NDJSON lifecycle log to `<data-root>/logs/daemon.log` (suppressed by
`ORCA_DIAGNOSTICS_DISABLED=1`). Rotation of that file is not implemented — see
[What is not covered](#what-is-not-covered). orcad records every trace span it emits
(git commands, worktree paths, terminal spawns, structured-chat failures and the rest) to
`<data-root>/logs/orcad.trace.ndjson`, rotated at 10 MB × 10 files, private to its user and
redacted for secret-shaped strings. It stays on the host: a desktop's diagnostics bundle does
not collect it. `ORCA_DIAGNOSTICS_DISABLED=1` turns it off, and a logs folder orcad cannot
open leaves it off with one stderr warning rather than stopping orcad.
### orcad supervising the daemon
- **Launch.** On Linux, through `systemd-run --user --scope` so the daemon gets its own
transient cgroup and survives a service-unit restart; everywhere else, and wherever that
scope is unavailable, forked detached. Either way it runs `daemon-entry.js` beside
`orcad.js` with its own PID record, token and socket under `<data-root>/daemon`.
- **Adoption before spawn.** A daemon already answering the endpoint is adopted, not
replaced, unless it is unhealthy, foreign, or built from a superseded bundle _and_ owns no
live sessions. Replacing a healthy daemon kills its PTYs, so code freshness always defers
to live work.
- **Restart.** The adapter respawns the daemon on death, transparently to callers.
- **Crash-loop containment.** At most **5 launches per 60s rolling window** per orcad run;
past that, launches are refused with `daemon_crash_loop` and terminals fail with that
message instead of the process forking forever. The window slides, so a repaired host
recovers without restarting orcad. An operator-initiated daemon restart clears it — that
is the deliberate "try again".
- **No macOS login-session watch.** That watch retires the daemon when the spawning GUI login
session dies. An orcad daemon must survive its SSH session ending.
- **Shutdown.** orcad never stops the daemon. A daemon that was never adopted retires itself
after its adoption window; an adopted one stays resident (see Decommissioning).
### Decommissioning
After a PID-scoped stop, an adopted daemon stays resident so the next orcad can reattach. A
combined-unit systemd stop also leaves a scope-isolated daemon resident, but kills one that
fell back to the service cgroup. To retire a process-scoped deployment, apply the census rule
above, stop orcad, then stop the daemon named by `health.terminalDaemon.pid`.
Only report it `exited` after verification on the execution host; loss of contact is
`unverifiable`.
### Windows hosts
What differs on a Windows SSH host, and what deliberately does not:
- **Stop path.** A signal is TerminateProcess on Windows: no flush, no lock release. The
slot's `.orcad-stop-request` file (and the managed, instance-bound request) is therefore the
only graceful stop. A detached orcad receives no console control events, so the listener
(`fs.watch` plus a one-second poll) is what stops it; the packaged-slot test proves it exits
cleanly within the 15 s shutdown deadline on every server lane, Windows included.
- **Exit proof.** `--complete-managed-stop` proves a reused PID by the addon's creation time.
Without the addon a live PID stays `live` or `unverifiable`, never `exited`.
- **Daemon endpoint.** The terminal daemon listens on a named pipe
(`\\?\pipe\orca-terminal-host-v<protocol>-<suffix>`), not a socket under the data root.
- **Leaving sshd's job.** orcad is started outside the SSH session's kill-on-close job, so the
daemon it forks inherits no such job and outlives the connection the same way.
- **Per-PTY jobs.** Each ConPTY child gets its own job (`windows-pty-job.ts`), and Git Bash /
MSYS panes follow [`windows-msys-job-breakaway.md`](./windows-msys-job-breakaway.md)
unchanged. A ConPTY smoke test runs inside a process started exactly that way (breakaway,
no window) on the Windows server lanes.
- **No daemon-host relocation.** The desktop copies its runtime to `%LOCALAPPDATA%` because
the NSIS updater deletes the install directory under a running daemon
([`windows-daemon-host-relocation.md`](./windows-daemon-host-relocation.md)). orcad slots are
versioned directories that nothing deletes while a process runs from them: Windows refuses
to delete a running image, and GC treats an in-use slot as live.
## Idle exit (client-managed orcad only)
An orcad that a desktop client launched over SSH stops itself, like the relay, once its host has
been unused for 15 minutes. The client's launch sets `ORCA_ORCAD_MANAGED_ACTIVATION_ROOT`; an
orcad started by hand, by a supervisor, or as a paired server never carries it and never idles
out.
"Unused" means every one of these held on every check for the whole period:
- no client socket open and no RPC request running;
- no terminal in the PTY provider, and the daemon answered with zero live sessions (a daemon
that does not answer keeps orcad up);
- no agent reporting `working`;
- no staged migration into this server;
- no enabled automation and no automation run still in flight (nothing on the host would start
orcad again for the next scheduled run, so a server with an enabled automation never idles out);
- no activation fence on the host (an update, rollback, decommission or recovery in flight).
The stop is the ordinary graceful shutdown, which disconnects from the daemon and never shuts it
down, so it cannot kill a terminal. It then asks the daemon to retire only if the daemon itself
proves it holds no session. Before stopping, orcad writes `<data-root>/orcad-idle-stop.json`;
the next start reports it once as `health.previousIdleStop` and removes it, so a later crash is
never read as an idle stop. A managed start with no record reports `previousIdleStop: null`.
The client starts a stopped server again, whatever stopped it (an idle stop, a kill, a host
reboot): on every connect, on every fresh tunnel (including after the client wakes from sleep),
and before a call through an environment the client restored at launch. A server that does not
answer is checked on the host; only a proven exit starts the activated slot, under the activation
fence, and the status line shows "Starting managed server…". A daemon that survived is adopted
with its terminals; after a reboot both start fresh. A process that is live or cannot be proven
gone is left alone, and a start that fails keeps the host managed with the reason and orcad.log's
tail, never as a verdict about its terminals. `ORCA_E2E_ORCAD_IDLE_TIMEOUT_MS` shortens the idle
period for tests; the client forwards it to the servers it launches.
## Health
The readiness payload carries a `health` object:
```
buildHash sha256 (16 hex) of the running orcad bundle — build identity that a version
string cannot give, so a rollback that did not replace the file is visible
buildVersion ORCA_VERSION
nodeVersion / nodeAbi process.versions.node / .modules — the ABI native addons must match
platform / arch / pid
terminalDaemon:
state live | degraded | absent
ownsFreshSessions whether NEW terminals are daemon-owned; this supports PID-scoped
restart recovery, not supervisor or service-cgroup isolation
pid the live daemon's pid, from its own PID record
buildVersion the build the LIVE daemon was forked from (may legitimately predate
this orcad after an update — reporting orcad's version for both would
hide exactly that)
entryPath / protocolVersion
selfTest { ok, coverage, verdict, durationMs }
```
### What the self-test proves
`selfTest` runs `checkDaemonHealth` against the daemon's socket. It is green only when the
daemon **opened its socket, completed the protocol handshake, and ran `ptySpawnHealth` — a
real short-lived PTY spawned inside the daemon's own process**. It therefore spans both
processes: orcad drives it, the daemon performs it, the verdict crosses the socket.
- `coverage: 'pty-spawn'` — the full round trip above.
- `coverage: 'handshake'` — **win32 only**, where `checkPtySpawnHealth` returns without
spawning anything. A green verdict there covers the handshake and nothing more. It is
reported separately rather than folded into `ok` so nobody reads it as a PTY round trip.
`state` is `live` only when the self-test passed **and** `ownsFreshSessions` is true. A
daemon that answers but has fallen back to local spawning for new terminals is `degraded`,
because those terminals die with orcad. A daemon that answered and then failed its spawn
probe is also `degraded`, not `absent`: it still holds live sessions, and calling those
exited would be the verdict `ssh-execution-boundary.md` forbids guessing.
## What is not covered
Named here so nothing reads as implemented that is not:
- **A continuous health endpoint.** `health` is published once, in the readiness payload. A
supervisor's periodic liveness/readiness probe needs an HTTP or RPC surface over the same
`collectOrcadHealth()`; that surface does not exist yet.
- **Supervision of an unscoped fallback daemon.** When the durable `systemd-run --user --scope`
launch is unavailable (see [above](#two-long-lived-processes-not-one)) orcad and its daemon
share one service cgroup, and a combined-unit stop cannot preserve live terminals. There is
no mechanism that re-isolates such a daemon after the fact.
- **libc slot.** There is no honest health value to publish until native libc detection owns
it.
- **`degradations[]`.** The readiness contract does not publish this collection yet.
- **Credential administration** (list / revoke / rotate devices, expiring pending offers,
structured security logging).
- **Pinned-port fail-closed.** A pinned `--port` still falls back to an OS-assigned port on
conflict.
- **Reconciling `webClientUrl` with reachability** under the loopback default.
- **State-schema rollback rules.**
- **Daemon log rotation.** `<data-root>/logs/daemon.log` grows unbounded.
@@ -1,22 +0,0 @@
# Configured command aliases for orchestration workers
A worker can use the name of a direct executable configured in **Settings → Agents**.
Choose the built-in agent whose command-line interface the executable implements,
then set that agent's command override to the executable name or quoted full path.
For example, configure Codex's command as `codex-fugu`, then select it with
`orca orchestration worker-start --agent codex-fugu` and the normal placement options.
The execution host resolves its own configuration. Launch receipts use the canonical
agent (`codex` in this example), and model/effort handling reuses that agent's existing
launch rules. A receipt records applied launch preferences; it does not prove provider
entitlement, successful generation, or an arbitrary vendor's model selection behavior.
Aliases require a single executable token. Commands containing interpreter arguments,
environment assignments, or shell wrappers are not aliases. Multiple built-in agents
configured with the same executable name are ambiguous and require the canonical agent
ID. Disabled launchers remain disabled. An unconfigured name is refused even if it is
on PATH; Orca cannot infer a compatible launch interface from a process name.
The same rule applies to folder workspaces and git worktrees. For a remote worker,
configure the command on its execution host. Older hosts may refuse aliases they do not
support. Configuring a command does not create new status producers or grant permissions.
-92
View File
@@ -1,92 +0,0 @@
# Native dependency install policy
Ordinary `pnpm install` installs optional native dependencies for the current OS
and CPU only. This applies to local development and root-project CI jobs,
including jobs using `.github/actions/install-node-dependencies`. Mobile and
cloud projects with their own workspace configuration are separate.
The one cross-target build in the repo is macOS: `pnpm build:mac` and the four
macOS packaging workflows produce both x64 and arm64 artifacts from an arm64
runner. Before packaging for another architecture, widen the CPU set:
```sh
pnpm install:release
```
This runs `pnpm install --frozen-lockfile --cpu=current,x64,arm64`. It never
widens the OS set: every packaging job runs on a runner whose OS matches its
target, so cross-OS installs are never needed. The macOS workflows pass
`--cpu=current,x64,arm64` directly; Windows and Linux packaging jobs use a plain
host-only `pnpm install --frozen-lockfile`. Keeping `current` in the list
preserves the host's build tools alongside the target resources. An install for
another target does not itself cross-compile native addons.
## Packaging guard
electron-builder only logs a warning for a missing `extraResources` source and
continues, so without a check a foreign-architecture slice would ship silently
broken. `beforePack` in
[`config/electron-builder.config.cjs`](../../config/electron-builder.config.cjs)
therefore calls `assertPackagedNativeVariantsInstalled` in
[`config/packaged-runtime-node-modules.cjs`](../../config/packaged-runtime-node-modules.cjs),
which fails the build when the target platform/architecture's native variants
are not installed: `sherpa-onnx-*`, `@parcel/watcher-*`, and on Windows the
node-gyp addon `@vscode/windows-process-tree`. The error names every missing
package and gives the remedy that fits: another architecture's variants come
from `pnpm install:release`, the Windows addon does not (see below).
Windows packaging requires a Windows host. `@vscode/windows-process-tree` is an
`os: win32` npm addon, so it is installed only where that matches;
`@orca/windows-registry` is a workspace package that links on every host, but
its native binary is still compiled only on Windows. Both are compiled only by
the Windows-only rebuild in `config/scripts/rebuild-native-deps.mjs`
(`allowBuilds` in `pnpm-workspace.yaml` keeps pnpm itself from running node-gyp
for them). The guard checks `@vscode/windows-process-tree` alone because the
workspace link is present everywhere and proves nothing. `pnpm install:release`
does not help on macOS or Linux because it does not widen the OS set.
Tests that inspect installed Windows addons and their packaging closure run on
Windows, where those dependencies are required. The PR Windows lane explicitly
includes them. Loading the packaging config tolerates absent Windows addons; only
`beforePack` rejects Windows packaging until they are installed. Patch-source
assertions, fixture tests, and the isolated real patch-install test continue to
run on other hosts.
## Existing checkouts
A narrowing incremental install can leave previously installed variants behind.
Stop development processes using this checkout, remove **this checkout's**
`node_modules`, then run `pnpm install --frozen-lockfile` to obtain a fresh host
install. Do not use `--force`: pnpm 12 documents that it installs optional
packages even when their OS/CPU/libc do not match. The shared pnpm download store
is separate; this change does not clear it.
## Measurement: macOS arm64, pnpm 12.0.0
Measured 2026-09-12 with the same package manifest, lockfile, patch files, and
existing download store in two fresh install directories, both with
`--frozen-lockfile --ignore-scripts --offline`. One used the earlier broad
policy (`--os=current,darwin,linux,win32 --cpu=current,x64,arm64`); the other
used `--os=current --cpu=current`. Numbers are logical file bytes in the virtual
store, **not unique disk usage**: pnpm hardlinks and APFS clones may share
storage.
| Measure | Broad install | Host install | Reduction |
| ------------------------------------------ | ------------: | ------------: | --------------------: |
| Package directories | 1,296 | 1,206 | 90 |
| Regular files | 53,972 | 52,915 | 1,057 |
| Logical package bytes | 2,485,153,773 | 1,159,396,649 | 1,325,757,124 (53.3%) |
| Logical package GiB | 2.31 | 1.08 | 1.23 |
| Install wall time, single warm-cache trial | 11.98 s | 11.25 s | 0.73 s |
| Native family | Broad MiB | Host MiB |
| ------------------- | --------: | -------: |
| Canvas | 235.7 | 25.8 |
| Sherpa speech | 235.2 | 71.6 |
| oxlint and tsgolint | 234.0 | 32.7 |
| SWC | 225.8 | 24.4 |
| TypeScript | 158.4 | 26.2 |
Electron downloads and native rebuild outputs are absent from both trials. The
single warm-cache timing pair does not establish a speedup, and Linux/Windows
results were not measured; do not present the macOS numbers as CI savings.
-90
View File
@@ -1,90 +0,0 @@
# Qoder CLI integration
Orca registers `qodercli` as `qoder`: detection, picker/settings, desktop and mobile
identity, prompt launch, permission flags, managed hooks, workspace trust, and
session resume use the existing TUI-agent and hook-status paths.
## Verified contract
Verified on macOS with Qoder CLI 1.1.64, including its versioned executable
`qodercli-1.1.64`:
- `--prompt-interactive` starts an interactive prompt; `--resume <id>` resumes a
session; `--dangerously-skip-permissions` is the permission bypass flag.
- Hooks use the Claude-shaped nested configuration in `~/.qoder/settings.json`.
Orca registers its own `/hook/qoder` source and preserves user hooks.
- Real SessionStart, UserPromptSubmit, Notification and SessionEnd events are
captured in `src/shared/__fixtures__/qoder-no-account-hooks.jsonl`, including
`source: resume` with the original session ID.
- Trust is `permissions.trustDirectories` in that settings file, using the
canonical workspace path. A workspace `.trusted` marker did not bypass the
trust dialog in this version. Remote trust uses the execution host filesystem.
- Qoder emits its `◇ … | Ready` title while the trust menu owns input. Readiness
therefore requires the live composer text, not merely this title or silence.
Raw PTY fixtures under `src/main/runtime/__fixtures__/qoder-*` cover trust,
unauthenticated prompt handling and the ready composer.
- A hidden Orca dev instance launched Qoder from New Tab in a folder workspace.
Its icon, label, terminal rendering and failed-turn indicator were inspected.
A harmless prompt reached Qoder, which reported its credit usage limit; the
canonical hook store recorded Qoder identity, session metadata and failure.
## Compatibility and limits
Windows hooks explicitly select Qoder's documented PowerShell shell, avoiding
an assumption that Qoder uses Git Bash just because it is installed. They reuse
the existing managed `.cmd` payload without adding an encoded PowerShell hop.
The configuration shape and preservation are tested; Windows, Linux and SSH
execution have not been exercised live here.
Resume requests to remote hosts require `agent-session.qoder-resume.v1`, so an
older host is not sent an agent enum it cannot accept. Remote hook installation
and trust preservation have automated coverage.
Successful model output, tool execution and live permission dialogs remain
unverified because the available account has exhausted its credits. Hook event
mapping for these paths follows the official documentation. The China executable
`qoderclicn` is not registered; it was not available for verification.
Every launch Orca starts (New Tab, workspace and draft launches, automations)
pre-trusts the workspace at spawn while the agent-wide "Trust the folder when
Orca starts an agent" setting is on. A hand-typed `qodercli` in a plain terminal
can still show Qoder's trust dialog.
## Sources
- [Official hooks documentation](https://docs.qoder.com/cli/hooks)
- [#15291](https://github.com/stablyai/orca/pull/15291): registration and hook proposal
- [#13311](https://github.com/stablyai/orca/pull/13311): icon asset (commit
`b56197025530adb1b82d97a364a8c3d4a02c0d42`) and agent integration
- [#8611](https://github.com/stablyai/orca/pull/8611) and
[#12910](https://github.com/stablyai/orca/pull/12910): contributor implementations
- [#16540](https://github.com/stablyai/orca/issues/16540): Qoder misidentified as Gemini
Contributor patches informed the integration; they were not applied wholesale.
In particular, the trust marker proposal was replaced using the installed CLI's
observed settings write, and no renderer workaround was added because the live
WebGL terminal rendered correctly.
## Cross-review of open proposals
Reviewed the current patches for all six open Qoder PRs on 2026-09-28:
| PR | Incorporated or confirmed | Deliberate differences |
| --- | --- | --- |
| #7502 (vincent-lxc) | Agent registration, mobile parity, identity before Gemini | Use structural titles and the current public permission flag. |
| #8611 (Eridanus117, building on #7502) | Qoder hook source, notification/permission mapping, session resume | 1.1.64 supports `--prompt-interactive`; preserve startup session identity and ignore compact restarts. |
| #9655 (xingqingzzp-gif) | Cross-checked minimum registration coverage | Placeholder icon and detection-only scope are superseded by the fuller integration; no code copied. |
| #12910 (jyang2004) | Reuse the Claude-compatible installer with an event subset, plus remote installation | Use observed `.qoder` path and `/hook/qoder`; Qwen and the China build remain separate work. |
| #13311 (sorrycc) | Icon, structural title disambiguation, headless detection, display glyph cleanup | No DOM renderer override: 1.1.64 renders correctly in the live WebGL terminal. No unverified executable alias. |
| #15291 (adlternative) | Versioned binary recognition, trust preflight wiring, skills mapping and resume | Trust comes from settings, not `.trusted`; include notifications, permission and failure events with Qoder-specific normalization. |
The rendered left sidebar was checked in the hidden Electron app using a temporary
folder workspace and real Qoder hooks. The workspace fixture needed a parent path
on its project group to appear in the sidebar. No agent status was injected into
the renderer. A real prompt updated the row text and returned a credit-limit error;
the row and workspace card showed failure with the Qoder icon. A DOM observer
recorded the prompt-row transition before failure. Permission/waiting and successful
completion remain automated event-mapping tests, not live model/tool verification.
Commit co-authors credit all six proposal authors for the implementation and
registration groundwork, including xingqingzzp-gif’s minimal registration proposal.
@@ -1,39 +0,0 @@
# Relay regional placement
Orca selects a Relay region in the Electron main process before requesting a new assignment. The
director publishes an allowlisted region catalog containing only HTTPS cell subdomains of that
director. Orca discards one warm-up `/health` request per probe origin — a cold request pays TCP and
TLS setup that can exceed the round trip it measures — then takes three bounded samples and compares
regions by their minimum. A wide spread still rejects a region, but only a genuinely flapping one.
The stable choice is cached for 24 hours, and a cached region changes only when the alternative is
materially faster.
A region wins only against a measured competitor. If any region in the catalog is rejected or cannot
be measured, Orca sends no hint rather than selecting the sole survivor. Sending no hint is not
neutral placement: the director assigns `preferredRegion ?? RELAY_DEFAULT_REGION`, and the default
is `us-central1`. So an `asia-east2` user whose `us-central1` probe fails or flaps once is placed in
`us-central1` for that refresh. That trade is accepted because the relay database is
`us-central1`-only, and it is bounded: the withheld hint is cached for one hour, not the 24 hours a
chosen region gets, so the next hour re-measures. An origin that fails its warm-up probe is dropped
before the sampling rounds, so an unreachable region costs one probe timeout rather than four.
After a control socket registers, Orca probes the cell it actually landed on, once per cell URL per
process. The cache is deleted only when it names a region other than the best measured one and the
assigned cell is more than three times slower than that region — a far cell under a cache that still
names the best region means the director declined the hint, and re-measuring would return the same
answer. Self-heal skips an absent, expired, or no-hint cache, and never runs under
`ORCA_RELAY_REGION_OVERRIDE`.
The assignment request sends only `preferredRegion`. It does not send latency, IP address, country,
pairing data, or credentials. Catalog, probe, and cache failures fall back to an assignment without
a region preference. A rolled-back director that rejects the new field is retried once without
only that field while preserving reconnect behavior.
The selection measures the desktop network path. Folder workspaces and SSH workspaces share the
same local broker and do not run probes on remote hosts. The phone continues to connect to the
exact cell URL in the desktop pairing payload, so its location is not measured independently and
no mobile protocol update is required.
For deterministic local diagnostics, set `ORCA_RELAY_REGION_OVERRIDE` to `us-central1` or
`asia-east2` before launching Orca. The override is not an end-user setting and is not written to
the preference cache.
-397
View File
@@ -1,397 +0,0 @@
# Remote wire compatibility
Orca's remote-server feature pairs a desktop client to a remote Orca runtime, and
users update the two independently. **Mixed versions are the normal state**, not an
edge case. This page is the contract for changing anything a paired client and host
exchange: the runtime RPC envelope, the terminal binary stream, and the content
either side publishes over them.
`src/shared/protocol-version.ts` says when to bump `RUNTIME_PROTOCOL_VERSION`. This
page covers the changes that do _not_ bump it and are therefore easy to get wrong.
## Rule 1 — a new optional JSON field on an existing frame is safe
Every JSON payload is parsed with a decoder that ignores unknown keys (zod `.strip()`
on RPC params, `JSON.parse` on stream frames). An older peer that has never heard of
the field simply does not read it.
Safe:
```ts
// host adds a field; older clients ignore it
encodeTerminalStreamJson({ kind, cols, rows, hiddenOutputReason })
```
**The field is safe only for as long as every reader treats it as optional.** The
moment a newer client _requires_ it, that client is broken against every host that
predates the field — which is the same defect as removing a field, just discovered
later. If new behavior depends on the field being present, that is Rule 2: negotiate
it, or make the reader fall back.
## Rule 2 — a new stream opcode is NOT safe; negotiate it
`decodeTerminalStreamFrame` returns `null` for an opcode it does not know, and
`runtime-rpc.ts` drops that frame without an error:
```ts
const frame = decodeTerminalStreamFrame(bytes)
if (!frame) {
return // silently dropped — the sender never learns
}
```
So a new opcode sent to an older peer does not fail loudly. It vanishes, and the
feature behind it appears to hang. Input sent under a new opcode is swallowed.
A new opcode must be announced in the subscribe handshake and sent only after the
peer confirms it. The existing pattern is `SetOutputPaused` (opcode 16):
- the client advertises support in the `Subscribe` frame's `capabilities`;
- the host echoes `capabilities: { outputPause: 1 }` on the `subscribed` event;
- the client sends opcode 16 only after that echo (`stream.supportsOutputPause`);
- the host only acts on opcode 16 when it negotiated it (`stream.supportsOutputPause`).
Reuse an existing opcode with a new optional payload field (Rule 1) whenever that
expresses the change; reach for a new opcode only when framing genuinely differs.
Opcode numbers are permanent. See the `Ack = 13` and `ClaimViewport = 14` comments
in `src/shared/terminal-stream-protocol.ts` for why a shipped number cannot be
reused even if the feature behind it is removed.
## Rule 3 — changing what the host publishes breaks old clients with no wire change
The frame shape can be untouched and the skew still real, because clients react to
frame _content_. PR #12641 is the worked example: the host stopped synthesizing a
finished agent status, and clients running older code saw different content in an
identical frame.
Treat these as wire changes even though nothing in the codec moves:
- a field the host stops populating (an old client reading it now sees `undefined`);
- a value whose meaning, units, or nullability changes;
- content the host stops synthesizing, trims, or starts deriving from a new source;
- a frame the host stops sending, or starts sending, on an existing path.
If old clients cannot interpret the new projection correctly, gate it behind a
runtime capability the same way Rule 2 gates an opcode.
## Rule 4 — an enum arm set is a wire surface; unknown arms must degrade, never reject
A closed `z.enum` in a client-side reply schema is a version claim: it asserts the host
will never send an arm this build has not heard of. A newer host that adds one arm then
has its whole reply refused, or has the row carrying it silently dropped, even though
every member the client actually reads is present and well-formed.
Declare the arm set open instead, with `openEnum` in `src/shared/zod-salvage.ts`:
```ts
// unknown arm degrades to a member the reader already handles; a non-string stays fatal
status: openEnum(GIT_BRANCH_COMPARE_STATUS, 'error')
```
Do not reach for `.catch()`. It swallows absence and the wrong type as well, which turns
a member the reader depends on into a silent default.
A fallback does not have to be an arm. `git.status`'s `area` is the worked example:
`staged`, `unstaged` and `untracked` each grant an affordance, so coercing an unknown area
to one of them offers stage, unstage or commit against a row the client cannot place. It
degrades to absent instead, which withholds all three — every reader is an equality check
against a known arm, so the row lands in no section — while keeping the row itself. That
last part is the point. Dropping the row would also drop it from the unresolved-conflict
gate, and a conflicted worktree that looks clean is granted a hosted-review create it
should not have. Withholding an affordance is a degrade; removing the evidence a gate
reads is not.
A fallback is only ever allowed to shape a _reading_. If the member is sent back to the
host — a token the client echoes into a later call's params — pass it through as
`z.string()` and let the send site keep it verbatim. `hostedReview`'s `provider` is the
case: the eligibility reply names it and the create call returns it, so an
`openEnum(..., 'unsupported')` there does not soften how the client reads a newer host's
provider, it puts `unsupported` on the wire and makes that host refuse its own. A
reply-schema fallback must never shape a param. Gate on the token instead, where the
client decides what it is willing to do with an arm it does not know.
## Worktree activation belongs to a viewer
Worktree creation and activation default to the host desktop for host/CLI requests,
and to the caller for paired desktop/web requests. Mobile creation still activates
the host renderer to provision setup and default tabs. A headless host does not
borrow a paired client's view. Catalog and session updates continue to reach every
subscriber independently of navigation.
Headless host/CLI creates provision their shells, setup, and default tabs on the
execution host in the background, without depending on an observer to open them.
The same fallback applies when an attached host renderer is unavailable or
reloading, including explicit `all` requests that target the host; availability is
checked after creation, rather than before its awaits.
`activateWorktree` client events carry an optional `navigation` field. Only explicit
`clients` or `all` requests publish these events; updated clients ignore events
without that address, including implicit broadcasts from older hosts. This is an
optional JSON field (Rule 1): older clients ignore it and still understand explicit
activation. The new host's default publication changes deliberately remove implicit
navigation for older clients too, rather than retaining the unwanted behavior.
An older host cannot express explicit follow intent to an updated client, so that
client must open the workspace itself.
Accepted navigation is also fenced through repository/worktree discovery; bridge
cleanup, reconnection, or re-pairing revokes any activation still awaiting a fetch.
Older CLIs hardcode `navigation: 'all'` for `--activate` and `--run-hooks`. The host
recognizes their `cliProvenanceRequest` and normalizes that automatic target to
`host`. An API caller without the CLI marker can still explicitly request `clients`
or `all`. Neither creation nor activation navigates remote clients by default.
Coverage: `multi-client-navigation-isolation.integration.test.ts` exercises real
paired WebSocket clients and host/headless creation, the renderer bridge tests
exercise older unaddressed publications, and
`paired-worktree-activation-isolation.spec.ts` checks a paired desktop remaining in
a local workspace while its remote catalog updates.
## Session search agent negotiation
`aiVault.searchStatus` optionally advertises `supportedAgents`; current search clients
send their own `supportedAgents` with `aiVault.searchSessions`. These are string lists,
so a future provider name does not make a peer reject the capability reply. The client
narrows explicit agent filters to the host's list before calling its request parser.
The host narrows retrieval to the client's list before publishing a page.
A peer without this field uses the frozen v1.4.211 search vocabulary. The existing
`supportsQoderHistory` flag proves CodeBuddy, ZCode, and Qoder support;
`supportsJcodeHistory` independently proves Jcode support. An explicit list takes
precedence over both flags. If only the status method is missing, the client still
searches the conservative legacy subset; other status errors propagate. Empty host
intersections keep the requested filters and use the existing no-match scope, so
consent, readiness, and unknown-scope results retain their normal precedence.
Local IPC advertises this build's full list, and every remote leg negotiates separately.
## Enforcement
`tests/e2e/cross-version-wire/cross-version-terminal-wire.unit.test.ts` runs the real
host RPC methods and the real renderer multiplexer from two builds against each
other — current working tree against the newest release tag, in both skew
directions — over one scripted terminal journey (subscribe, input, hide/reveal
snapshot, drop, reconnect).
Run it with:
```bash
pnpm exec vitest run --config config/vitest.config.ts tests/e2e/cross-version-wire/cross-version-terminal-wire.unit.test.ts
```
It fails when a frame is refused by the receiving build's decoder (Rule 2), when the
observed frame sequence changes (Rule 3), or when published snapshot content or
negotiated capabilities differ from the contract. Repeated frame shapes are compared
by corresponding journey occurrence (initial, reveal, reconnect), so a field removed
from one occurrence cannot hide behind a sibling that still publishes it. Adding an
optional field keeps the suite green (Rule 1); making a client depend on that field
turns the new-client/old-host pairing red.
### Never write down what the old side has
The baseline is whichever release tag is newest, so it moves on every cut. An
expectation of the form "the old side does not have X" — a `not.toHaveProperty`, a
`not.toContain`, a hard-coded field list — stops being true the first time a release
ships X. The suite then reddens on whatever pull request is in flight, with no code
change anywhere, and the job trains people to ignore it. That is worse than no test,
because a rolling baseline eventually contains every additive field the wire has, and
adding one is the sanctioned way to evolve it.
Derive the expectation from the baseline that was actually checked out:
- for a published frame, pair each build against a client of its own version and
compare the skewed pairing against that same-version reference, so the expectation
is whatever that build publishes today. Compare repeated frames by corresponding
occurrence with `comparePublishedFieldOccurrences` in
`tests/e2e/cross-version-wire/published-field-shape.ts`; never union keys across
initial, reveal, and reconnect frames, because a sibling can mask one occurrence's
removed field;
- for a negotiated surface, read the old build's advertised capabilities and
registered method names from its checkout, and assert they agree with each other
rather than asserting the old build lacks them;
- for a "client too old to know X", derive that client's advertised list by removing
X from the baseline's own list, so the gate stays exercised after X ships.
Name the direction in the assertion. `new client against old server` and `old client
against new server` fail for different reasons, and the host is the only side that
authors a published frame — the terminal `terminalOwner` false positive on 2026-08-29
was misread as a new client sending an unknown field when the old server was
publishing it. Two things are still safe to state literally: the current build's own
contract, and an invariant that holds for every version.
Pinning a legacy ref is the fallback when a contract genuinely needs a release from
before a feature shipped, as `cross-version-browser-placement.unit.test.ts` does with
`LEGACY_BROWSER_PLACEMENT_RELEASE_REF`. It does not rot on a cut, but it is
hand-maintained, so prefer deriving.
`tests/e2e/cross-version-wire/cross-version-agent-session-wire.unit.test.ts` pairs the
same two builds over the structured `agentSession.*` surface. Because a released build
cannot name a capability string its own source never contains, the old side's advertised
list and registered method names are read from the extracted checkout rather than
hand-written. It covers the three skews that surface can fail on:
- an old client — advertising the baseline's list minus this capability — is told the
whole surface does not exist and reaches no host method;
- a new client against the old dispatcher always gets an answer rather than silence,
and `method_not_found` for every method that release does not register, so the
absence is visible during negotiation instead of by calling;
- a cursor survives a host restart: a reattach at the client's fence is refused as stale
with the live one attached, a write still carrying that fence is delivered (writes are
named by their target, and every released client still sends a fence), and resuming from
the held cursor replays only what it missed.
Run it with:
```bash
pnpm exec vitest run --config config/vitest.config.ts tests/e2e/cross-version-wire/cross-version-agent-session-wire.unit.test.ts
```
The harness covers the terminal stream and the structured agent-session surface. It does
**not** cover the session-tab sync channel, legacy agent-session publications, file or Git
RPCs, mobile/E2EE framing, or the relay transport. A change on those paths still needs its
own reasoning against the three rules above.
## Worked example: `agentWait` on terminal and worker reads
`terminal.show`, `orchestration.workerShow` and `orchestration.federationShow` carry an
optional `agentWait` naming a pane parked on a prompt only a human can answer. It is Rule 1 —
a new optional field — but it has a second state that Rule 1 alone does not describe, and
getting that wrong turns a skew into a false "nothing is blocked".
- **present object** — this pane is waiting, with the evidence that proved it.
- **present `null`** — the host evaluated this pane and nothing proves a wait.
- **absent** — the host never evaluated it: it predates the field, the worker identity was
unverifiable, the pane was unreadable, or the agent probe did not answer in time.
A new client against an old host sees the field absent, which is why absence must read as
_unknown_ and never as _not waiting_. Collapsing absent into `null` at any hop — including a
convenience `?? null` in an RPC handler — makes an old or unreachable peer indistinguishable
from a healthy idle worker, which is the exact failure the field exists to remove.
An old client against a new host ignores the key, as Rule 1 allows. New members added to
`RuntimeTerminalWaitBlockedReason` are also Rule 1: no consumer switches exhaustively on it,
and both the CLI and worker-start interpolate it as an opaque string.
## Worked example: the `turn` journal item and its transitional downgrade
The structured chat journal records a turn as a first-class item,
`{ kind: 'turn', turnId, state, userItemId?, startedAt?, completedAt?, durationMs? }`, where it
used to write `{ kind: 'status', text, turnLifecycle }`. Nothing in the codec moves, but it is
Rule 3: a client that predates the item does not know the kind and renders it as a text bubble
with no text. So the item is gated on a client capability, `agent-session.turn-item.v1`.
The gate lives at the RPC boundary only, in
`src/main/runtime/rpc/methods/structured-agent-session-turn-item-capability.ts`, composed
around `agentSession.history` and `agentSession.subscribe` next to the background-task
projection. A client that does not advertise the capability receives every `turn` item
rewritten to the legacy status form with the full lifecycle under `turnLifecycle`; a client that
advertises it, and any in-process caller, receives the canonical body. The journal, the status
feed, and every host-side reader keep the `turn` item; `readAgentJournalTurn` in
`src/shared/agent-session-turn-record.ts` reads either form, so a new client against an old
host that still writes the status row also works.
An old client against a new host sees the status row it always did. A new client against an old
host advertises a capability the host ignores and reads the status row through the shared
reader. The downgrade is transitional: once no supported release lacks the capability, delete
the projection module and the capability check, and leave the reader.
The cross-version suite derives the old client's list by removing this capability from the
baseline's own list, per the rule above, so the downgrade stays exercised after a release ships
it.
## Known debt: JSON-RPC errors drop Node's string code
An error raised on an SSH host crosses the relay as JSON-RPC, and
`ssh-channel-multiplexer` rebuilds it with the TRANSPORT's numeric `code`. Node's
string code — `'ENOENT'`, `'EACCES'` — does not survive, so a caller on this side
cannot ask what kind of failure it was.
`isENOENT` in `src/main/ipc/filesystem-path-containment.ts` pays for that by also
matching Node's canonical message text, which is what makes remote worktree creation
work. The cost is that a host can make an unrelated failure read as "absent" by
putting that sentence in a message.
The exit is Rule 1: carry the original string code in a new optional field on the
error payload and read that instead. An old host omits it and the message match still
covers them; once hosts that send it are the floor, the message match can be deleted
rather than lived with at its ~10 call sites. Narrowing `isENOENT` back to `.code`
without doing this reinstates the bug — the transport has already overwritten it.
## Known hazard: clients ignore host-published failure fields on client-placed pages
`RuntimeMobileSessionBrowserTab` — the browser tab a host publishes on the session-tab sync
channel — permits `placement`, `loadError` and `certificateFailure` together. But for a tab
whose `placement.kind` is `'client'` the engine runs in the client's own app: the failure is
raised by the local guest webview, and the host has no view of it (`RuntimeBrowserClientPage`,
what the registry actually publishes from, carries neither field). Clients from
this version on therefore refuse host ownership of both records for client-placed pages
(`web-session-tabs-sync.ts`, the `placement?.kind !== 'client'` carve-outs) — without that,
each metadata snapshot deletes the locally recorded failure and the page's failure overlay
disappears mid-navigation.
The hazard is forward-facing and Rule 3 shaped. A host that later starts publishing
`loadError` or `certificateFailure` for a client-placed page reaches these clients as content
they silently drop, so the host would see no error and no effect. Publishing it has to be
capability-gated, with the carve-out narrowed to clients that did not negotiate the
capability. Note the cross-version harness does not exercise the session-tab sync channel, so
nothing fails if this is forgotten — this note is the only record.
A related carve-out covers `title`, `url`, `loading`, `canGoBack` and `canGoForward`
(`resolveMirroredBrowserPageContent`), and for those the hazard is already live rather than
forward-facing: the host does publish them, from a `RuntimeBrowserClientPage` it can only learn
about second-hand through the client's own `browser.clientHost.pageMetadata` calls. Its copy
therefore starts at the registry defaults (`'Browser'`, the create-time url), and while those
publishes are failing it never leaves them.
That copy is not simply behind, though, and a client must not treat it as such. When a lease
reattaches, the host refreshes the page from the client host's own inventory
(`runtime-browser-client-page-recovery.ts`), which reads the live guest — so it can be strictly
fresher than a local row whose pane is unmounted and whose metadata publisher was disposed with
it. A client that ignores the host url is relying on its own guest to re-answer on remount,
which `ClientHostedBrowserPagePane`'s mount-time `syncNavigation` is what makes true.
These five are therefore refused only by the client whose guest actually runs the page:
`placement.browserHostClientId` is compared against this client's own host id
(`readBrowserClientHostId`). Main stamps that id into the guest-hosting window's
`additionalArguments` at creation, and the preload reads it back out of its own argv — the answer
has to be there before the first snapshot is interpreted, which is earlier than any IPC handler a
renderer could wait on. Every other viewer — a second desktop, the web client, which installs no
page renderer at all, the dashboard pop-out, which is deliberately left unstamped — keeps tracking
the host, which is the only reason a mirrored viewer shows anything but its first snapshot
forever. Improving what a _second_ client sees still means fixing the publish, not the carve-out;
the carve-out no longer stands in the way of it.
The two failure fields above are deliberately left on the looser `placement?.kind !== 'client'`
predicate. It is unobservable today — the host publishes neither field for a client-placed page at
all, so a mirror has nothing to take either way. If the capability-gated publish this section
anticipates ever lands, narrow them the same way rather than by placement kind: a mirror should
take a failure it cannot otherwise see, and only the hosting client should refuse it.
## Known hazard: on the mobile surface a scope refusal is not a missing method
The agent-session harness above asserts that a peer probing an unknown method is told
`method_not_found`. That holds for the runtime-scoped surface and **not for the mobile one**.
`runtime-rpc-websocket-dispatch.ts` checks `MOBILE_RPC_METHOD_ALLOWLIST` and answers
`forbidden` _before_ it calls the dispatcher. A method a desktop predates is on neither the
allowlist nor the registry, and the gate answers first, so a phone never sees
`method_not_found` for it. `method_not_found` would reach a mobile-scoped device only for a
method that is allowlisted but unregistered, and `src/main/runtime/mobile-rpc-allowlist.test.ts`
requires every method the phone calls to be both, so that combination cannot ship. The reverse
— a registered method that no release has allowlisted yet — is the skew that does occur.
So a phone-side "is this host new enough to serve X?" probe must read **both** codes as
absence. Keying it on `method_not_found` alone compiles and passes every same-version test
while never firing. The Files and Git fallbacks have read both since they shipped
(`isMobileMethodUnavailableError`, `isMobileGitUnavailable`); the Relay pairing probes were the
outlier, and the cost was a write-once direct-relay upgrade journal, holding a pending resume
secret, that an old desktop could never retire
(`mobile/src/transport/pairing-relay-rpc-unavailable.ts`, with the gate pinned desktop-side by
`src/main/runtime/runtime-rpc-mobile-unknown-method-scope-refusal.test.ts`).
Widening is safe only where `forbidden` cannot also mean a real authorization failure. Prove
that per probe rather than globally: the pairing handlers cannot emit it (an unwired provider
answers `runtime_error`), a bad or revoked token answers `unauthorized`, and the gate is one
of only two places in `src/main` that emits the code at all. A probe whose handler _can_
refuse by authorization must not be widened.
The cross-version harness dispatches as a mobile client but calls the dispatcher directly, so
it never runs this gate and nothing reddens if any of the above is forgotten.
@@ -1,340 +0,0 @@
# Renderer agent-status performance
## Status
This design is adopted for the renderer's high-frequency agent-status path. It
keeps the existing event semantics while bounding the amount of synchronous
store fanout performed for one IPC burst.
It lands in slices. This document and the `bench:idle-cpu` harness land first so
the store changes can be reviewed against a baseline someone else measured. Until
the store slice lands, the `setAgentStatuses` / `transactAgentStatuses` actions
and the agent-status write workload described below are not yet on `main`; the
harness measures scale, listener census, and raw publication fanout only.
## Context
Orca can display a large expanded worktree lineage inside one virtualized list
row. Virtualizing the root row does not virtualize its descendants, so a
100-worktree lineage can mount 100 `WorktreeCard` instances at once.
Agent-status IPC events are bursty. The renderer already groups live events into
a 33 ms window, but the original flush applied every queued event with a
separate Zustand write. Zustand synchronously visits every listener for every
publication. The resulting work therefore grew with both the number of status
events and the number of mounted subscriptions:
```text
burst work ~= status events x store listeners x selector work
```
A production trace captured the renderer repeatedly entering
`flushLiveAgentStatusBurst -> applyAgentStatus -> setAgentStatus -> setState`
through `Set.forEach`. A deterministic 100-worktree fixture reproduces the
structural multiplier; see "Baseline on `main`" below for the currently measured
listener count and publication cost.
The production app later recovered substantially when all configured remote
hosts were removed. Host removal can stop relay/reconnect traffic, remove
mounted remote worktrees, or both, depending on host type and removal options.
That observation identifies remote presence as the production trigger but does
not by itself distinguish traffic volume from mounted-listener fanout.
A read-only reconnect audit ruled out systematic double status emission from a
full PTY replay: replay bytes bypass OSC status parsing. Reconnect still causes
a full terminal-buffer repaint for every attached remote pane, which is a
separate source of renderer work and remains a follow-up investigation.
## Goals
- Keep a large expanded lineage responsive during dense agent-status traffic.
- Preserve every ordered status transition, including repeated updates for one
pane inside the same burst.
- Publish agent-status state once for a deferred live burst.
- Preserve selector identity and child render isolation.
- Make the regression reproducible without relying on a user's production data.
## Non-goals
- Changing the client/server status payload or remote protocol.
- Deduplicating status events by pane.
- Changing agent freshness, retention, history, title, completion, or provider
session behavior.
- Changing remote reconnect, PTY replay, or terminal repaint behavior.
- Redesigning lineage presentation or collapsing worktrees automatically.
## Design
### Bound mounted subscription fanout
Sidebar components select cohesive state bundles with shallow equality instead
of registering one listener per field. Derived arrays and maps retain their
existing shallow identity behavior so unrelated store writes do not rerender a
card. Full agent-list mode keeps its child-level subscription boundary; compact
mode passes the already selected rows to avoid selecting the same inputs twice.
The deterministic 100-worktree fixture pins the resulting listener budget:
| Surface | Listener budget |
| ------------------------------ | --------------: |
| Worktree card state and caches | 2 |
| Agent-row inputs | 1 |
| Worktree activity status | 1 |
| Closed context menu | 1 |
Unmount tests require the listener count to return to its prior baseline. In the
bundled prototype, the fixture without seeded agents fell from 8,518 listeners
to 1,218; with 100 visible agent rows the candidate mounted 1,618. Compare
against the census in "Baseline on `main`", which the harness reports directly.
### Share working-spinner phase without synchronous mount queries
Working rows keep the existing compositor-driven CSS animation and shared
visual phase. `animationstart` anchors each animation to document time zero.
Deferring the animation query until that event avoids a synchronous style flush
at each mount and restores the shared phase after `display:none` or a motion
preference change. A negative mount-time delay cannot preserve that phase after
an animation restarts.
### Bound spinner animation overhead
Working rings keep compositor-driven CSS animation, but repeat the animation
once per day rather than once per second. The same 12 steps per second now
avoid recurring React animation-iteration dispatch. The existing stationary
wrapper and ring rendering stay unchanged. Offscreen containment was evaluated
and rejected after a pixel regression at low zoom on 1x displays.
The history, isolated measurements, full-app workspace/agent/subagent benchmark,
and limitations are documented in [Spinner rendering performance](./spinner-rendering-performance.md).
### Fold a burst in event order
The store exposes the single-update action and two batch forms:
- `setAgentStatus(paneKey, payload, ...)` retains the positional single-update
API for the immediate live path.
- `setAgentStatuses(updates)` applies a prebuilt ordered list, while
`transactAgentStatuses(operation)` lets IPC derive each update against the
exact staged state before the single commit.
Both entry points reuse the same single-update state transition. The batch
reducer passes each resulting state into the next update, so a sequence such as
`working -> waiting -> done` retains the same history and timestamps as three
sequential calls. Updates are never keyed or deduplicated before the fold.
The live IPC queue is spliced before it is processed. This preserves the
existing reentrancy guarantee: a synchronous subscriber can enqueue another
event without causing the current queue to be drained recursively. The first
event outside an active burst remains immediate; events accumulated within the
33 ms window are applied as one ordered transaction. Startup snapshots and
bounded pending-hydration retries use the same transaction path instead of
publishing once per restored pane.
Each transaction builds pane-routing ownership once with the same first-match
semantics as the standalone resolver. Split-layout leaf membership is indexed
once per layout root, so a large snapshot performs linear tab and leaf work
instead of rescanning every mounted worktree for every pane.
### Run effects after the transaction
Generated-title work that requires committed state is deferred until after the
transaction. Accepted updates also request freshness scheduling; the outer batch
coalesces those requests and schedules the shared freshness timer once after its
single commit. Generated-title requests are folded in event order and published
together, including first-write and forced-replacement semantics. Resolved tab
titles are projected while the transaction folds, then final title changes are
published together. Completion-triggered review refreshes remain deferred
microtasks.
Bulk title application preserves event order and duplicate-tab behavior while
indexing owners once, cloning each changed owner array once, and replacing each
top-level map once. This keeps the post-commit title phase linear in mounted
tabs plus changed titles.
This separation is important: invoking store actions from inside a Zustand
updater would re-enter the store, while running an effect before the commit
would let it observe stale state.
## Semantic invariants
Sequential and batched application must agree on:
- live and retained agent maps;
- state history, `updatedAt`, and `stateStartedAt`;
- agent identity, model, prompt, tools, assistant messages, and subagents;
- orchestration and provider-session continuity;
- sleeping-session and launch-config recovery records;
- retired/closed-pane rejection and inherited-status suppression;
- retention cleanup and live-map eviction;
- `agentStatusEpoch` and `sortEpoch`;
- automation completion observation across intermediate transitions;
- generated-title inputs, freshness scheduling, and completion refreshes.
Equivalence tests use fixed timestamps and include repeated same-pane
transitions. A publication-count test subscribes to the real store and requires
one notification for a non-empty batch and none for an empty batch.
## Benchmark contract
The benchmark launches an E2E-mode Electron build with the store exposed only
for instrumentation. It creates 100 worktrees in one expanded lineage, verifies
100 mounted cards, captures the store listener census, and then applies seeded
ordered agent-status traffic through the real store action.
The benchmark measures the synchronous store action, not the live IPC leading
edge or post-commit notification path. A real-store snapshot test covers the
end-to-end budget for 100 panes with auto-generated titles enabled: one status,
one bulk generated-title, and one bulk resolved-title publication. Disabling
generated titles removes that middle publication, independent of pane count.
The artifact records only fixed diagnostic fields needed for comparison:
- requested and completed batches and updates;
- store action calls and observed publications;
- elapsed time, throughput, and scheduling drift;
- final-state verification;
- renderer mean, p95, and maximum CPU;
- renderer timer drift and long tasks;
- mounted-card and listener counts.
Raw process inventories, temporary paths, pane identifiers, and DOM text are
diagnostic-only and must not be embedded in the shareable report.
Run baseline and candidate on the same machine and OS with the same Electron
build mode. CPU samples from macOS and Linux are comparable within that
constraint; Windows process CPU collection currently cannot support this
comparison.
## Harness
`pnpm run bench:idle-cpu` drives `config/scripts/run-idle-cpu-benchmark.mjs`,
which composes four modules:
| Module | Responsibility |
| ------------------------------------- | ------------------------------------------------------------------------------------------------- |
| `idle-cpu-renderer-scale-fixture.mjs` | Seeds the lineage, agent rows, and sidebar view state; takes the mounted-card and listener census |
| `idle-cpu-renderer-timing-probe.mjs` | In-page timer drift and long-task probe; runs the no-op publication workload |
| `idle-cpu-process-sampling.mjs` | Classifies the Electron process tree and samples per-role CPU/RSS |
| `idle-cpu-synthetic-spinners.mjs` | Measurement-only visible spinners |
The sampling window extends past `--sample-ms` until the workload settles, and
fails the run rather than reporting a truncated window if the workload overruns
the guard. That is why a 2,000-publication run reports a measured window longer
than the requested one.
`--zustand-publications` publishes an empty partial through the real store, so
each publication costs exactly one full subscriber visit and nothing else. It
isolates the `listeners x selector work` half of the burst-cost model from
agent-status payload work, and it is store-API independent — it measures the
same thing before and after the batching slice.
The agent-status write workload (`--agent-status-batches`,
`--agent-status-write-mode`) is not in this harness yet. It depends on
`setAgentStatuses`, so it lands with the store slice.
## Baseline on `main`
Measured on `main` at `077f5a11cd4` (macOS, arm64, 16 CPUs), Electron built with
`electron-vite --mode e2e`, headless, 100 worktrees at lineage depth 99 with 100
seeded agent rows, 10 s warmup and a 30 s sampling window.
Fixture scale is confirmed by the census rather than assumed: 100 store
worktrees, 100 mounted cards, 100 mounted agent rows, and **9,279 store
listeners**. That listener count is the multiplier the design targets.
2,000 no-op store publications at a 1 ms cadence, three repetitions. The
listener census was 9,279 in every run.
| Measure | Median | Runs |
| ------------------------ | ----------: | ------------------------------ |
| Wall time to complete | 12,325.7 ms | 13,148.6 / 12,325.7 / 11,870.5 |
| p50 scheduling drift | 5,150.3 ms | 5,432.3 / 5,150.3 / 4,945.2 |
| p95 scheduling drift | 9,802.9 ms | 10,570.4 / 9,802.9 / 9,390.0 |
| Renderer mean CPU | 18.25% | 20.25 / 18.25 / 15.88 |
| Renderer p95 CPU | 32.59% | 40.07 / 32.59 / 31.28 |
| Renderer timer drift p95 | 7.0 ms | 7.3 / 5.3 / 7.0 |
2,000 publications requested over 2 s take about 12 s, so the renderer sustains
roughly 160 publications per second at this scale. Each publication is
individually short - the long-task observer recorded zero entries in all three
runs - so the cost surfaces as scheduling drift and sustained CPU rather than as
discrete long tasks. Compare drift and CPU here, not long-task counts.
The idle control at the same scale with `--zustand-publications 0` reports 6.63%
renderer mean CPU, 17.11% p95, and 1.6 ms p95 timer drift. Roughly 11.6 points
of mean renderer CPU are therefore attributable to publication fanout rather than
to the mounted fixture itself. 200 spinner animations run in both cases, so the
control also bounds the animation cost out of the comparison.
Reproduce with:
```bash
pnpm run bench:idle-cpu -- --worktrees 100 --lineage-depth 99 \
--agents-per-worktree 1 --warmup-ms 10000 --sample-ms 30000 \
--zustand-publications 2000 --zustand-publication-interval-ms 1 \
--output /tmp/idle-cpu-baseline.json
```
## Results
Three repetitions used 100 mounted worktrees, lineage depth 99, 100 seeded
agent rows, and verified final state. Medians from the regenerated evidence set
are:
| Single 2,000-update burst | Sequential | Batched |
| ------------------------- | ---------: | ---------: |
| Status-state publications | 2,000 | 1 |
| Store action time | 3,692.0 ms | 188.7 ms |
| Update throughput | 541.7/s | 10,598.8/s |
| Renderer mean CPU | 36.2% | 2.9% |
| Renderer p95 CPU | 107.3% | 8.2% |
| p95 long task | 4,653 ms | 216 ms |
The direct store transaction performs 99.95% fewer status-state publications,
spends 94.9% less time in the store action, and processes updates 19.6 times
faster. Renderer mean CPU falls 92.0%, renderer p95 CPU falls 92.4%, and the p95
long task falls 95.4%.
The 60-burst × 32-update case at 33 ms is a sustained saturation stress, not a
real-time production SLO. Publications fall from 1,920 to 60 and median store
action time falls from 2,791.9 ms to 323.3 ms. Median completion time falls from
5,298.2 ms to 2,710.6 ms, p95 scheduling drift falls from 3,073.0 ms to 664.9
ms, and long-task count falls from 57 to 1. Renderer p95 CPU remains saturated
and noisy in this cadence, so it is not used as the discriminating measure.
The 20-pane artificial OpenCode regression passes with 12.4 ms median key echo,
25.2 ms worst key echo, 19.4 ms maximum timer drift, and zero dropped renderer
backlogs.
These figures come from the bundled prototype and are restated here as the
target. They are re-measured with the harness when the store slice lands.
## Acceptance criteria
- The 100-worktree fixture stays at or below the pinned listener budgets.
- A deferred transaction performs one status-state publication while preserving
ordered final state, including live-map eviction at the 500-row cap.
- A 100-pane startup snapshot performs one status and one bulk resolved-title
publication with generated titles disabled; enabling generated titles adds at
most one ordered bulk publication while preserving final statuses and titles.
- Sequential-versus-batch equivalence tests pass across same-pane transitions
and side-effect-bearing updates.
- Renderer CPU tails and scheduling drift improve in repeated candidate runs.
- The 20-pane artificial terminal test reports no dropped output backlog and no
material typing-latency regression.
- Web typecheck, focused unit tests, lint, max-lines ratchet, and E2E build pass.
## Compatibility
This is renderer-local. It adds no RPC field, stream opcode, persisted data, Git
command, or provider-specific contract. Native, WSL, SSH, relay, folder
workspace, and git-worktree status events enter the same renderer action. Mixed
client/server versions therefore need no capability negotiation.
## Failure containment
The first live event remains immediate. Startup replay and bounded pending
retries fold synchronously without waiting for the 33 ms live-burst window, but
publish their accepted updates together. Empty batches are no-ops. If an update
is stale or targets retired authority, the reducer skips only that update and
continues folding later events in order.
@@ -1,92 +0,0 @@
# Runtime file Base64 padding
Padded runtime file writes must have a length divisible by four. Empty strings and
unpadded Base64 with length modulo four equal to zero, two, or three remain valid.
The change rejects exactly the previously accepted strings containing trailing
padding whose total length modulo four is two or three. It does not enforce
canonical unused pad bits or change the alphabet.
## Boundary evidence
| Input | Before | After |
| ------------------------- | ------ | ------ |
| `A=` | Accept | Reject |
| `AA==` | Accept | Accept |
| `AAA=` | Accept | Accept |
| `AAAA` | Accept | Accept |
| `''` | Accept | Accept |
| `A` | Reject | Reject |
| `==` | Accept | Reject |
| `AA=A` (interior padding) | Reject | Reject |
| `AA=` | Accept | Reject |
| `A==` | Accept | Reject |
| `AAAA==` | Accept | Reject |
| `AA`, `AAA` (unpadded) | Accept | Accept |
`Buffer.from('A=', 'base64')` decodes to zero bytes. Rejecting malformed padding at
the RPC boundary prevents an accepted request from silently writing different bytes.
## Caller census
| Caller / surface | Reachability and compatibility verdict |
| ------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Desktop `runtime-file-import-client.ts` → `uploadRuntimeFileWithoutClobber` → `writeRuntimeBase64File` | The only production producer of `files.writeBase64` and `files.writeBase64Chunk`. Staging in `filesystem-runtime-upload-staging.ts` encodes the complete file with `buffer.toString('base64')`; newly rejected values cannot be produced. |
| Desktop single-frame uploads | Sends the staged string unchanged when its length is at most 512 × 1024 characters. Standard Node Base64 always has length divisible by four. Empty files remain accepted. |
| Desktop chunked uploads | Slices the encoded stream at 512 × 1024 = 524,288 characters, divisible by four. Every offset and every complete chunk is quartet-aligned. The final chunk is the difference of two multiples of four, including when it ends in `=` or `==`. No separately assembled final chunk or per-chunk padding is added. |
| Web implementation of `stageExternalPathsForRuntimeUpload` | Returns an empty source list; no file-write payload is produced. |
| Mobile file editor | Uses `files.writeTerminalArtifact` with text content and its separate schema. Does not reach this predicate. |
| Mobile clipboard / image attachments | Uses `clipboard.startImageUpload`, `clipboard.appendImageUploadChunk`, `clipboard.commitImageUpload`, and the `clipboard.saveImageAsTempFile` fallback. Their validator is `isValidBase64` in `clipboard-params.ts`, not `isValidRuntimeFileBase64`. Unchanged, including the mobile normalizer's existing permissive padding behavior. |
| CLI | No producer of either Base64 file-write method. File commands in `src/cli/handlers/file.ts` call `files.open` / `files.openDiff`; other CLI RPC call sites do not construct Base64 file writes. |
| Generated params catalog | References both schemas; `RpcParams` consumers use inferred types. Mobile's entry point is `export type` only, so no new client-side parsing occurs. |
| RPC dispatch | `files-mutation-methods.ts` registers both schemas. The chunk schema extends the whole-file schema; these are the only runtime consumers of the predicate. Direct runtime/provider calls do not parse these schemas. |
Repository-wide searches covered method names, schema names, the predicate and its
pattern, and all callers of the upload/staging functions. Targeted history search
on `HEAD` under `mobile/src` found no introduction/removal of the affected methods
or predicate. The local release refs `mobile-ios-v0.0.27` and `mobile-v0.0.13` also
contain no callers of either Base64 file-write method; the iOS ref uses the separate
clipboard and terminal-artifact methods above. No shipped mobile producer of a
newly rejected value was found in this source/history audit.
## Remote and workspace compatibility
Old desktop clients using the audited producer send valid quartets to a new host.
A new client still sends the same bytes to an old host. No method, field, opcode,
or host-published content changes. This follows the mixed-version requirements in
[remote-wire-compatibility.md](./remote-wire-compatibility.md).
The RPC validation runs before workspace resolution and provider selection, so the
same rule applies to folder workspaces, git worktrees, local hosts, and SSH hosts.
SSH ownership fences and provider writes are unchanged. Arbitrary external RPC
callers sending malformed padding will now receive a validation error; valid
padded and unpadded payloads remain accepted.
## Regression evidence
`src/main/runtime/rpc/methods/files-base64-padding.test.ts` exercises both actual RPC
registrations, asserts rejected input never reaches the writer, and verifies
accepted content is forwarded unchanged. With the original predicate, the test
run produced **16 failed / 16 passed**; all 16 failures were newly rejected padding
shapes accepted by the old implementation. This was run before editing the predicate.
The existing desktop external-import test now uses a final `AA==` chunk after a
524,288-character first chunk, pinning padded final-chunk forwarding in the real
upload path. No producer changes or clipboard validation changes were necessary.
## Validation results
All test/typecheck commands used `ORCA_BACKGROUND_LAUNCH=1`.
- `pnpm tc`: exit 0; all root typecheck projects passed.
- `pnpm exec vitest run src/main/runtime/rpc`: 277 files passed, one failed;
2,419 tests passed, two timed out, one skipped. Both timeouts were in the unchanged
`terminal-output-frame-chunks-equivalence.test.ts` (5s surrogate-range test and
30s 800-payload fuzz test).
- `pnpm --dir mobile typecheck`: exit 0 (`tsc --noEmit`).
- `pnpm run check:code-quality:changed`: exit 0; code quality, type-aware code
quality, and React Doctor each reported zero new findings across three source files.
- Focused run with `--config config/vitest.config.ts --maxWorkers=2`: all three
files / 58 tests passed, covering padding, desktop external imports, and the
terminal-output equivalence file that timed out in the initial run.
- Full RPC rerun with `--maxWorkers=2`: exit 0; all 278 files passed,
2,421 tests passed and one skipped (198.34s). No timeout overrides were needed.
-88
View File
@@ -1,88 +0,0 @@
# Share and install agent skills
Orca can put one skill or a bundle of skills behind one unlisted, revocable link. Shared bundles
do not appear in search, a catalog, or a public index. Anyone who has an active link can inspect
and install its contents without signing in, so treat the link like a credential.
## Share skills
Publishing and link management require an Orca account in the desktop app.
1. Open **Skills** and choose **Share skills**.
2. Select one or more skills. One link can contain a large collection, such as 30 skills.
3. Review the bundle name, included skills and files, scripts, executable files, digest, account,
and optional release notes.
4. Choose **Publish skill**, **Publish bundle**, or **Publish new version**, then copy the link.
Orca publishes an immutable version. Later changes do not silently alter a link's current bytes;
publish a new version to update the Cloud package.
Use **Settings → Share Skills** to copy or revoke active links. Revocation blocks new previews and
download grants. A grant issued immediately before revocation can remain usable for up to five
minutes, and revocation does not remove copies that recipients already installed.
## Install from a link
Opening an Orca skill link shows a preview before changing any files. You can also open **Skills**,
choose **Install from link**, and paste the URL.
1. Verify the author and organization.
2. Review the version, release notes, included skills, scripts, executable files, and digest.
3. Select all skills or only the ones you want.
4. Choose the destination machine and either global or workspace scope.
5. Review new, unchanged, updated, and conflicting skills, then choose **Install N skills**.
Supported destinations include the local machine, paired Orca runtimes, WSL, and SSH hosts. The
destination runtime resolves its own home and workspace paths, so folder workspaces and remote
filesystems do not borrow paths from the client machine.
Orca keeps one canonical installed copy and places it where supported agents can discover it.
Current provider coverage is documented in
[Agent skill provider paths](./agent-skill-provider-paths.md).
## Conflicts, updates, and rollback
**Keep local** is the default when an existing skill differs. Orca replaces modified content only
after you explicitly choose to discard it.
Open **Skills → Manage installs** to inspect managed skills and their immutable version history.
Installing the latest version performs an update; selecting an older retained version performs a
rollback. Both use the same protected install transaction. If a bundle changes between versions,
Orca updates only the selected skills that still exist in that version.
An interrupted install is recovered on restart. If Orca reports a conflict or partial result,
review the named skill and retry; completed skills do not need to be installed again.
## Remove an installed skill
Use **Skills → Manage installs → Remove**. Orca removes only copies and provider placements that it
owns and can verify. Modified or unowned files are preserved and reported. Discarding modified
content requires a separate explicit confirmation.
Removing a local install does not revoke its share or delete its Cloud package. Likewise,
revoking or deleting Cloud data does not reach into recipients' machines.
## Retention and deletion
- Upload grants expire after 15 minutes.
- Abandoned upload bytes are removed from quarantine after one day.
- Published versions have no automatic age-based deletion.
- Deleting a package revokes its links before unreferenced objects are deleted.
- Deleted GCS objects remain operator-recoverable through a seven-day soft-delete window.
- Installed copies remain until someone removes them on each destination machine.
Organization legal or retention requirements can override normal rollback and deletion timing.
## Trust and privacy
A skill is code from its author. `SKILL.md` can change agent behavior, and included scripts or
executables may run later when a person or agent uses the skill. Orca validates the package and
never executes its contents during installation, but you should install only from people you trust
and review unexpected scripts or executable files.
Orca records bounded operational identifiers and outcomes. Normal logs, telemetry, and support
bundles exclude skill contents, filenames, manifests, local paths, share URLs, upload policies,
download grants, credentials, and access lists.
If a link no longer works, ask its owner for an active link. Missing, expired, revoked, and deleted
links intentionally show the same response so Orca does not disclose private package existence.
@@ -1,203 +0,0 @@
# Spinner rendering performance
## ELI5
Imagine a wheel that tells the front desk every time it completes a lap. The
front desk is also handling your typing. CSS already turns the wheel for us,
but React still receives its once-per-second lap notifications.
We put a day's worth of laps into one animation. The wheel moves at the same
speed, while sending one lap notification a day. Drawing visible wheels still
costs something. This removes recurring bookkeeping from the input thread; it
does not make rendering or the rest of Orca free.
## How this builds on earlier changes
| Change | What it achieved | Remaining cost |
| ------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- |
| [#9380](https://github.com/stablyai/orca/pull/9380): shared JavaScript clock | Reduced frame-pipeline CPU in the original one-agent measurement | Wrote each spinner's style 12 times per second on the input thread |
| [#12359](https://github.com/stablyai/orca/pull/12359): compositor CSS rotation | Removed those recurring JavaScript style writes; fixed the reported typing regression | React still receives CSS iteration events |
| [#13987](https://github.com/stablyai/orca/pull/13987): synchronize on animationstart | Avoided a synchronous style query at every mount | Steady-state animation overhead stayed the same |
| This change | Preserves both later fixes and removes almost all iteration boundaries | Compositing, other app work, mount/reveal work, and a daily iteration boundary remain |
The historical measurements in #12359 reported 41 rings causing about 490 style
writes per second, with typing input-delay p90 of 363 ms versus 19 ms when those
writes stopped. Those are historical production measurements, not numbers from
this benchmark or a direct comparison with today's app.
## Implementation
The production change is entirely in CSS. `AgentWorkingSpinner`, its callers,
markup, border, animation-start handler, and reduced-motion behavior stay the
same. No DOM node, pseudo-element, containment boundary, timer, observer, or
JavaScript animation loop is added.
The transform travels 86,400 turns in 86,400 seconds with 1,036,800 steps: exactly
one revolution and 12 steps per second. `animationstart` sets `startTime = 0` as
before, preserving shared phase after mount and animation restart. The step
count is a timing-function parameter, not a million-entry keyframe list.
React installs delegated `animationiteration` listeners even when the component
has no iteration handler. A native 2.2-second trace of 200 isolated rings counted
400 iteration events and 800 JavaScript calls before the change, versus zero of
either with the long cycle. That trace installed no animation-event listener.
These are event dispatches, not component rerenders or 400 separate OS wakeups.
## Full-app benchmark
The opt-in Playwright benchmark launches a fresh, hidden Orca app for each
scenario. It creates real Git workspaces and seeds working statuses through the
existing renderer fixture, including in-process subagent data. It renders the
normal sidebar, virtualizer, lineage, agent rows, tabs, and terminal.
| Scenario | Git workspaces | Root agents | Subagents | Mounted / visible rings | Layout |
| ------------- | -------------: | ----------: | --------: | ----------------------: | -------------------------------------------- |
| `one-agent` | 1 | 1 | 0 | 3 / 3 | One working agent |
| `one-family` | 1 | 2 | 4 | 8 / 8 | All family rows expanded |
| `200-flat` | 200 | 400 | 800 | 162 / 15 | Normal virtualization; 23 workspaces mounted |
| `200-lineage` | 200 | 400 | 800 | 1,401 / 15 | Expanded lineage; all 200 workspaces mounted |
Measurement-only styles switch between the original one-second cycle and the
new long cycle on the same elements. The real React root, callers, status data,
and app stay the same. The reported run alternates A/B and B/A, with four
ten-second CPU samples per variant after warmup. CPU samples use cumulative
Electron process CPU and CDP main-thread task/script/style/layout metrics. No
renderer polling, screenshots, or benchmark iteration listeners run during
those CPU windows. No samples are discarded.
Typing is measured separately using the existing paced terminal-typing probe:
64 keys at 113 ms cadence, twice per variant, after two seconds of warmup with
status traffic. Status updates arrive in groups of up to eight every 200 ms.
Keys pass through the DOM, real PTY, and xterm. A sidecar timestamps arrival at
the PTY, and a bounded terminal-buffer scan observes each echo. Missing input
or echoes fail the benchmark. Echo measurements include the 10 ms scan interval;
they do not measure native display presentation. Native animation traces also
run separately from CPU and typing samples.
The statuses are deterministic test data, not hundreds of paid model sessions.
The test exercises UI cost under agent-status traffic, not the compute or network
cost of model inference, SSH traffic, or hundreds of streaming PTYs.
## Results
CPU values are medians of four samples. "CPU ms/s" means milliseconds of
processor time used in one wall-clock second: 100 ms/s is about 10% of one CPU
core. Renderer + GPU-process CPU includes their other app work and CPU used by
the graphics process; it is not GPU hardware utilization or whole-machine CPU.
The main thread handles input and is included in renderer CPU, not extra work.
Echo p90 means 90% of sampled keys were observed within that time; ranges show
the two runs, not confidence intervals. No keys or echoes were missing.
| Scenario | Renderer + GPU CPU ms/s, old → new | Main-thread ms/s, old → new | Echo p90 ms, old → new |
| ------------- | ---------------------------------: | --------------------------: | ---------------------- |
| `one-agent` | 37.0 → 38.2 | 5.4 → 3.5 | 19 → 18–19 |
| `one-family` | 46.2 → 44.6 | 7.8 → 4.4 | 17–19 → 18–19 |
| `200-flat` | 141.6 → 122.8 | 28.2 → 16.3 | 26–28 → 26–28 |
| `200-lineage` | 324.6 → 295.6 | 140.8 → 70.0 | 159–239 → 93–160 |
The consistent gain is less main-thread work: about 35%, 43%, 42%, and 50%
less in these four scenarios. Native 2.2-second traces counted 6, 16, 324, and
2,802 iteration events before, and zero in each new variant, without adding an
iteration listener. That avoided work also exists in Orca itself, independently
of the isolated fixture and CPU noise.
Total CPU was roughly unchanged in the one-worktree cases. In this run it fell
13% with normal virtualization and 9% with expanded lineage; seven of eight
paired large-case CPU samples favored the change. These percentages are not
universal: a shorter three-variant ablation measured flat-list CPU at 89.0 ms/s before and
108.0 ms/s with the long cycle, while main-thread time still fell from 26.3 to
17.4 ms/s. The repeatable main-thread reduction is stronger evidence than a
single total-CPU percentage.
Typing was similar in the small and flat-list cases. Expanded-lineage echo p90
improved in the final run, but a shorter ablation had similar before/after
latencies. No general typing speedup or statistical non-regression guarantee
is established by these short experiments.
### All CPU samples
Values are rounded to one decimal and listed by round, with no outliers removed.
The first new small-case samples were higher than their paired baselines; they
remain included. CPU and typing were sampled separately.
| Scenario | Version | Renderer + GPU CPU ms/s | Main-thread ms/s |
| ------------- | ------- | -------------------------- | -------------------------- |
| `one-agent` | Old | 37.8, 36.2, 26.2, 39.6 | 6.5, 5.2, 4.8, 5.7 |
| `one-agent` | New | 53.3, 35.8, 37.5, 39.0 | 6.7, 2.8, 3.0, 4.1 |
| `one-family` | Old | 46.3, 46.1, 47.9, 44.6 | 7.8, 7.6, 9.8, 7.7 |
| `one-family` | New | 53.8, 45.4, 42.5, 43.8 | 6.8, 4.5, 3.1, 4.3 |
| `200-flat` | Old | 142.4, 140.8, 147.0, 136.5 | 28.3, 28.1, 32.2, 27.0 |
| `200-flat` | New | 122.9, 97.2, 122.7, 126.7 | 19.2, 10.7, 16.6, 16.0 |
| `200-lineage` | Old | 317.9, 385.2, 315.9, 331.2 | 134.4, 159.6, 133.9, 147.3 |
| `200-lineage` | New | 318.6, 256.9, 296.5, 294.7 | 89.2, 60.8, 71.6, 68.3 |
## Reproduce
```sh
ORCA_BACKGROUND_LAUNCH=1 pnpm bench:spinners --sample-ms=5000
ORCA_BACKGROUND_LAUNCH=1 pnpm bench:spinners --verify-only --scale-factor=1
ORCA_BACKGROUND_LAUNCH=1 pnpm bench:spinners --verify-only --scale-factor=2
ORCA_BACKGROUND_LAUNCH=1 ORCA_SPINNER_BENCH=1 ORCA_SPINNER_KEYS=64 \
pnpm test:e2e spinner-workspace-perf.spec.ts --workers=1
```
The full-app command rebuilds in `e2e` mode. For a fresh build already made with
`pnpm exec electron-vite build --mode e2e`, `SKIP_BUILD=1` reuses it. Do not reuse
an old launch-policy build. `ORCA_SPINNER_SAMPLE_MS`, `ORCA_SPINNER_ROUNDS`,
`ORCA_SPINNER_KEYS`, `ORCA_SPINNER_KEY_CADENCE_MS`, `ORCA_SPINNER_VARIANTS`, and
`ORCA_SPINNER_OUTPUT` control the experiment. `ORCA_SPINNER_CPU=0` repeats only
typing; `--grep one-agent` selects one scenario. Reports, native traces, typing
sidecars, and CDP screenshots are written under `.bench-fixtures/`. Run one
benchmark at a time, without concurrent builds or tests.
The optional `contained` variant retains the rejected offscreen experiment for
ablation. It adds `content-visibility:auto` to the existing wrapper through
measurement-only styles. It is not enabled in production or the default
benchmark comparison.
## Visual and behavioral checks
Both 1x and 2x display-density checks passed 720 ring comparisons each: 6/8 px
rings, light/dark themes, supported zoom extremes, all 12 phases, long elapsed
times, and the daily wrap. The comparison pauses each animation and sets its
`currentTime`, so the long-elapsed and daily-wrap cases exercise the deterministic
style path rather than a running compositor animation. Against that path the
tolerance is one channel level for floating-point antialias rounding. A running
animation at multi-hour ages can differ by a few channels on the ring edge — a
fraction-of-a-pixel antialias difference at large accumulated angles, not a phase
or shape change. Checks also cover shared phase, reduced motion, initial offscreen
reveal, repeated scroll-away/reveal, and `display:none` restoration.
## Limits and rejected approaches
Adding `content-visibility:auto` to the existing stationary wrapper saved more
CPU at large mounted counts, but a 1x display check found a one-pixel shift at
the minimum UI zoom. That containment change is excluded. A previous
pseudo-element version also regressed typing latency in the virtualized list.
Neither prototype's CPU or typing numbers describe the final patch.
An initial typing run used a 100 ms key cadence, which can repeatedly align with
200 ms status bursts. Follow-up runs use 113 ms, more keys, and two seconds of
warmup under status traffic. This reduces timing bias; it does not excuse a
regression. CPU measurements run separately and do not depend on key cadence.
An early isolated test suggested a 31% process-CPU reduction that a longer audit
did not reproduce. The longer isolated audit measured original 104.04 versus
long-cycle 92.32 CPU ms/s, and main-thread 10.08 versus 0.24 ms/s. A fixture with
every ring far offscreen and containment enabled could also approach idle; that
is not representative of Orca with visible animations. Neither result justifies
claiming "free spinners" or a universal CPU percentage. Virtualized, unmounted
rows already cost nothing, and this patch does not add offscreen culling.
All local measurements use an Apple M4 (10 cores), macOS, Electron 43.4.1 /
Chromium 150.0.7871.224. Native windows stay hidden and unfocused;
benchmark-only settings disable background throttling to exercise the frame
pipeline. These are not visible-window power measurements. No battery benefit
is established. Linux/Windows need their own runtime measurements. The
renderer-only change does not alter SSH execution, wire data, status semantics,
Git operations, or folder-workspace ownership.
Animated PNGs, masks, layer promotion, CSS sprites, individual `rotate`, and
containment on the rotating element were also explored. Shared images added
raster work and regressed the single-ring case; sprites reintroduced per-frame
style work. They did not meet the appearance and responsiveness requirements.
-129
View File
@@ -1,129 +0,0 @@
# SSH Execution Boundary
How Orca splits work between your machine and an SSH host, what survives a disconnect, and how to keep `unverifiable` distinct from `exited`. Nothing under `docs/` stated this before; agents and humans were inferring it from error strings and getting it wrong.
## The rule
**The execution host owns everything that touches execution** — tools, credentials, identity, environment, processes, and artifacts. The client owns the UI, transport, and Orca control-plane state, but is not authoritative for execution state.
Two consequences, both non-negotiable:
1. **No silent substitution.** An operation on a remote `repoPath` must never fall back to running on the client. A missing SSH provider is not permission to answer locally — a local run can answer for the _wrong repository_.
2. **No asserting what you cannot observe.** Loss of contact is not evidence of `exited`. Report `unverifiable`, never `exited`.
The vocabulary is fixed: **`live` / `unverifiable` / `exited`**, taken from the incumbent `UnstoppedPtyVerdict`. Do not introduce synonyms, and never collapse `unverifiable` into either neighbour. `exited` requires positive evidence of absence from the host that owns the process; a transport failure can only ever produce `unverifiable`.
Rule 1 is stated at `src/main/source-control/repo-default-branch.ts:76-78`, `src/main/repo-worktrees.ts:45-48`, `OrcaRuntimeService.probeWorktreeDrift` in `src/main/runtime/orca-runtime.ts`, and `src/renderer/src/lib/connection-context.ts:22-24`. It is enforced throughout `src/main/runtime/orca-runtime-git.ts` by `requireRuntimeGitProvider` in `src/main/runtime/runtime-git-command-target.ts`, and throughout the runtime filesystem commands by `requireRuntimeFileProvider` in `src/main/runtime/runtime-file-command-target.ts`. Both route on the target's resolved `executionHostId` rather than on a repo row's `connectionId`: they throw the provider-unavailable message when an SSH host has no registered provider, throw `ExecutionHostNotDispatchableError` for a `runtime:` host this process does not execute, and return `null` only for `local`. Grep those names for the current call sites rather than trusting a count.
`src/main/runtime/unstopped-pty-verification.ts:12-16` is the reference implementation of rule 2: it keeps `live` / `unverifiable` / `exited` as three distinct verdicts, and treats "we could not ask" as its own answer.
## What runs where
| Concern | Executes on | Notes |
| -------------------------------------------------------------- | ------------------ | -------------------------------------------------------------------- |
| PTYs, agent CLIs | **remote** | children of the detached relay daemon, not of the ssh channel |
| git (status, diff, log, fetch, push, commit, branch, worktree) | **remote** | via `src/relay/git-handler.ts` |
| filesystem, watching, search | **remote** | |
| repo setup hooks (`--setup`) | **remote** | identical policy to local |
| commit-message / PR-field AI generation | **remote** | uses the remote agent CLI and its auth |
| `gh` / GitHub API, `glab` / GitLab | **client** | inconsistent with the rule; PRs carry the client's identity |
| the `orca` CLI inside a remote terminal | **client runtime** | control plane only — your files and processes stay remote; see below |
## Survival: what a disconnect does _not_ do
By default, remote work survives your machine going away. The relay is a detached daemon (`nohup … </dev/null &`), its handler in `src/relay/relay.ts` ignores `SIGHUP`, the PTY is its child rather than the ssh channel's, and quitting Orca is a **detach, not a dispose** (`src/main/ssh/ssh-relay-session.ts:901-915`). Sleep additionally pushes `graceTimeSeconds: 0` to un-bound any running grace window.
Two ways remote work _can_ actually stop:
- **A bounded grace period.** The shipped default is `0` = keep alive until reset. If "keep terminals alive until reset" is unchecked, the configurable range is **60s–7d** and the form defaults to **24h**. The countdown starts when the client disconnects, after which the relay SIGKILLs every PTY. Note the asymmetry: sleep protects you, but ordinary disconnect and app quit do not. No command reports which setting is in effect for a target, so at N hours since disconnect you cannot tell "unlimited" from "24h with 7 left" — treat the remote as `unverifiable`, not `exited`.
- **Host-acknowledged explicit user action** — End Remote Terminals, Reset Relay, removing the target, or closing the tab. When the host cannot acknowledge the request, closing a tab or removing a target may clear only client state; the remote verdict remains `unverifiable`.
Reconnect re-attaches to the same live PTYs and replays a bounded buffer (`REPLAY_BUFFER_MAX`, a 102,400-code-unit tail). Output beyond that while you were away is lost to the client even though the process was never interrupted: **the transcript is truncated; the work stays `live`.**
## Updating Orca strands relay-backed terminals
There is a third outcome that is neither of the two above, and the vocabulary matters: the work does not stop, it becomes unreachable through the new build's own relay.
The relay's install directory — and therefore its socket path — is namespaced by a content hash of the relay bundle (`computeRemoteRelayDir` in `src/main/ssh/ssh-relay-versioned-install.ts`, consumed by `resolveRemoteInstallState` in `src/main/ssh/ssh-relay-deploy.ts`), and the daemon refuses any client whose bundle hash differs (`handleDaemonHandshakeFrame` in `src/relay/relay-handshake.ts`, exit `EXIT_CODE_VERSION_MISMATCH` 42). Two builds whose relay protocol is byte-identical still refuse each other. So the first reconnect after an app update deploys a new relay at a path the incumbent was never listening on, and this build's own bridge cannot reach the incumbent. The incumbent's own bridge can: its version directory still holds its `relay.js` and `.version`, so `relay.js --connect` run from there presents its own hash by construction. Each deploy takes a census of this target's older endpoints (`startPreviousRelayCensus` in `src/main/ssh/ssh-previous-relay-terminals.ts`); while one is live or unverifiable, a not-found reattach keeps its lease and answers `SSH_PTY_HELD_BY_PREVIOUS_RELAY` instead of expiring. On POSIX hosts the pane is then reattached through that older bridge (`SshLegacyRelayRoute` in `src/main/ssh/ssh-legacy-relay-route.ts`), which takes the PTY owner role without output flow control, and every later operation on that PTY id is routed there (`src/main/providers/ssh-pty-legacy-relay-delegation.ts`). When the last pane it serves exits, the route hangs up and the incumbent's own grace retires it. On Windows the census probes each older version directory's pipe for the target (`src/main/ssh/ssh-previous-relay-windows-census.ts`), and a census that cannot run or does not know the host platform counts as `unverifiable`, never as "no older relay". The connect-time host census, which decides whether a host may convert to a managed server before any relay session exists, instead lists every `orca-relay-*` pipe on a Windows machine and maps each to the relay instance that owns it through that instance's credential file or pipe marker, whichever desktop launched it (`src/main/ssh/ssh-host-relay-windows-inventory.ts`); a pipe no directory accounts for counts as `unverifiable` unless the host proves it another account's. The migration terminal gate (`assessOrcadMigrationTerminals`) answers `exited` only from a complete inventory: every relay session answered, or, with no session, a host census proved the endpoints idle. A missing or failed inventory is `unverifiable` even when this desktop leases nothing. Where no route can be opened — Windows named pipes, a socket relocated under the short `/tmp` base, a bridge that fails — the PTY stays `unverifiable`: the pane says it is still running under the previous Orca version, and is never respawned. The old relay keeps its directory pinned against GC while its socket is live (`hasLiveRelaySocket` in `src/main/ssh/remote-install-gc.ts`). See #13852: on POSIX hosts that is no longer a dead end, because a terminal held by a previous relay version now resumes in its pane through that relay's own bridge.
The peer model does not have this failure, and that is the concrete reason behind "One host, one model" below. The daemon's endpoint is namespaced by a **semantic protocol version** rather than a build (`daemon-v<N>.sock`, from `getDaemonSocketPath` in `src/main/daemon/daemon-spawner.ts`), every earlier protocol version stays attachable (`PROTOCOL_VERSION` in `src/main/daemon/daemon-protocol-version.ts`), and a daemon holding live sessions is preserved across a version change instead of replaced (`shouldPreserveDaemonWithLiveSessions` in `src/main/daemon/daemon-replacement-preflight.ts`).
## Control plane
On an SSH host, `orca` is a shim (`~/.orca-relay/bin/orca`) that proxies **back to the client's runtime** over the relay socket. Your repository, processes, and files remain remote — only the control plane is on the client. This is correct for an SSH target, but it has a consequence worth stating plainly:
> When the client disconnects, every `orca …` command run on the SSH host fails with `No owning Orca client is connected to the relay`. The PTY stays `live`; its control plane does not.
Orchestration state (Runs, Tasks, Dispatches, mailboxes) is client-resident for the same reason. An agent on an SSH host should not depend on `orca` for anything it must finish while you are away. **Commit and push early** — unpushed work on a remote box is unavailable to the client until it reconnects.
## Distinguishing `unverifiable` from `exited`
A verdict needs evidence from the host that owns the process. Apply these tests in order.
**Was the signal produced by the owning host, or by the client's own bookkeeping?** Absence from a client-side set, a lookup that threw, a socket that closed, a command that timed out — none of these observe the process. They are `unverifiable` by construction, whatever the field is named.
**Did every remote PTY on that target go quiet at once?** A transport drop takes them all together. Simultaneous silence across a host indicates a lost link, not simultaneous death.
**Does the termination event match the current identity?** A host-delivered exit for the live PTY incarnation and provider generation, while its siblings still report, establishes `exited`. A stale event, an event for a superseded incarnation, or one quiet terminal with no host evidence does not.
**Did the answer carry its evidence, or only the same wording?** `pty.attach` refuses with `PTY "<id>" not found` both for a pid the relay probed and found gone and for an id its session map never had — which is every id minted before a relay restart, since ids carry a per-start mint epoch. Only the probed refusal carries `PTY_ATTACH_PROVEN_EXITED_MARKER` (`src/shared/pty-attach-absence-evidence.ts`) and reaches the client as `SshPtyProvenExitedOnRelayError`; the unmarked union arrives as `SshPtyAbsentFromRelayError`, which licenses retiring the client's own route to the PTY and nothing more. A missing marker is never evidence — an older relay omits it too.
**Is a returned status actually a claim of success?** An operation that reports failure may have succeeded, and one that reports success may not have run — check the durable state it should have changed rather than trusting the return.
Anything short of positive host evidence is `unverifiable`. Reporting it as `exited` is the error this document exists to prevent: it orphans live work and can cold-start a duplicate over the same worktree.
## Host contact is a different question from process liveness
The `live` / `unverifiable` / `exited` triple above answers one question: is this PTY running. It has
no synonyms, and nothing below adds any.
A second, narrower question — can we currently reach the host at all, and what is its last answer
worth — is answered by `RuntimeHostContact` (`src/shared/runtime-host-contact.ts`), whose arms are
`live` / `unverifiable` / `refused` / `retired`. These are **not** extra process verdicts and must
never be mapped onto one:
- `refused` is the host answering and turning us away — unauthorized, a protocol mismatch, a status
method it does not implement. That is positive evidence about the _connection_, and it says
nothing whatever about whether the host's PTYs are running. They almost certainly still are.
- `retired` is the pairing being ended by explicit user action. Same point: the client stops having
a route, the remote work is unaffected.
Both are reasons to stop _trusting a cached answer_, never reasons to report a process `exited`. A
reader that needs a process verdict must still get it from the host that owns the process, by the
tests above.
Why the extra arms exist at all: the renderer previously expressed every non-answer as one nullable
`status`, so a probe in flight, a probe that failed, a host that refused us and a retired pairing
all reached readers as the same `null` — and readers spent that `null` on decisions of very
different weight, including destructive ones. Folding `refused` and `retired` back into
`unverifiable` to match this document's triple would recreate exactly that collapse. The vocabularies
are deliberately separate because the questions are.
## Deciding a remote pane is idle
The orphan-PTY sweep is the one flow that turns an observation into a SIGKILL, so its idleness evidence has to be measured against the same thing the signal reaches. It is not the terminal.
`forceKillPosixPtyProcessGroups` (`src/main/pty/posix-pty-process-groups.ts`) collects every process group on the pane's tty and `killpg`s each one. The blast radius is therefore _(process groups on the tty) × (members of those groups, wherever they are)_, and the second factor is not bounded by the terminal at all. Two facts make that gap reachable:
- **Job control can be off.** With `set +m` a background job does not get its own process group — it keeps the shell's. `ps` then shows one process group on the tty, running a build. Nothing in a tty-shaped predicate can see it.
- **A group member can leave the terminal.** `ioctl(TIOCNOTTY)` without `setsid` drops the controlling terminal but keeps the pgid, so the process reports `tpgid == -1`, never appears in `ps -t <tty>`, and is still killed by `killpg(shellPgid)`. A double-forked grandchild similarly keeps the pgid while reparenting to pid 1, so no walk by `ppid` from the PTY root can name it either.
So `shellOwnsEveryTtyProcessGroup` (`src/main/providers/agent-foreground-process-batch.ts`) requires both measurements: every process group on the tty is the shell's own with none stopped, **and** the shell's own process group has no other member anywhere in the host's process table. The name is tty-shaped for wire-compatibility reasons only.
Two residuals remain, and neither is removable here. The capture is a snapshot, so work started between the `ps` and the signal is invisible — bounded by `RELAY_PTY_SWEEP_MAX_EVIDENCE_AGE_MS` on the reading side, not eliminated. And a process the host's own `ps` cannot enumerate (another PID namespace, `hidepid=2`, a table truncated by a permission boundary) is unobservable while `killpg` still reaches it.
The general rule this instantiates: **evidence must be measured in the unit the destructive action operates on.** Evidence in a different unit is `unverifiable` no matter how precise it looks.
## Reading artifacts instead of process state
Artifacts are stronger evidence than liveness signals, but they answer a narrower question than they appear to.
A matching commit from `git ls-remote --heads origin <branch>` or a PR head lookup proves **that commit reached the remote** — not that the current run pushed it, and not that the latest work was included. An absent result proves nothing was found, not that nothing was pushed: the ref may have been deleted, the PR closed, or the query may simply have failed.
A listing is only evidence about the hosts it actually covered. When a result does not name its scope, an empty answer is not evidence that nothing is running elsewhere. A clean **local** worktree says nothing at all about the remote one.
## One host, one model
An SSH host and a paired runtime (`orca environment`) imply opposite boundaries: the first is a dumb execution host driven by your client, the second is a peer that owns its own control plane. Registering the same machine both ways splits its worktrees across two identities, makes `terminal list` return different sets depending on `--environment`, and reliably confuses both humans and agents. Pick one per machine.
For work that must continue while you are offline, use the peer/headless-runtime model on the remote host instead of the direct-SSH model. Its control plane is host-local, and its daemon-backed PTYs can stay `live` across a PID-scoped runtime restart so the runtime can reattach. A service manager that reaps the runtime's cgroup, or an explicit daemon shutdown, makes them `exited`; see [Running orcad](./orcad-operations.md#process-scoped-and-cgroup-wide-stops). Do not register the same machine through both models. A detached agent process outside Orca can also survive a control-plane outage, but it has no stdin, so its instructions cannot be amended mid-run.
-424
View File
@@ -1,424 +0,0 @@
# SSH host key verification (STA-4319)
Revised after security and migration review. Where a first draft was wrong, the correction is kept
visible rather than quietly edited out — the reasoning matters for anyone changing this later.
## The defect
`src/main/ssh/ssh-connection.ts:1184` installs a `hostVerifier` that records a SHA-256 fingerprint
and then `return true`. Every ssh2 connection accepts every host key. There is no `known_hosts`
consult, no trust record, and no change detection anywhere in `src/main/ssh/`. There is exactly one
ssh2 `Client` construction site, so the fix has a single chokepoint.
Scope is per-connection, not per-feature: one `SshConnection` per target serves exec, SFTP, port
forwarding, the filesystem watcher and relay deploy.
### Threat model, corrected
Traffic is still encrypted, so a passive observer gets nothing. The exposure is an **active**
attacker who can redirect the connection — ARP/DNS spoofing, hostile Wi-Fi, a hijacked internal name.
Three corrections to the first draft:
- **Jump hosts are NOT the worst case; they are already safe.** `shouldUseSystemSshTransport`
(`ssh-transport-selection.ts:71-91`) returns true for exactly the conditions under which
`resolveEffectiveProxy` (`ssh-proxy-command.ts:17-38`) returns a proxy — the two branch on the same
inputs in the same order — and `attemptConnect` returns unconditionally after the system probe
(`ssh-connection.ts:670-673`). So ProxyJump/ProxyCommand go through OpenSSH and are already
verified. The ssh2 proxy-spawn at `:697` is effectively unreachable. Good news for migration, and
the first draft's motivating example was simply wrong.
- **Agent forwarding was overstated.** `agentForward` is gated on the user's `ForwardAgent yes`
(`ssh-connection-utils.ts:203-205`). `config.agent` is always set, but that is agent _auth_, whose
signatures bind the session id and cannot be replayed onward. The risk applies to users who opted
into `ForwardAgent`, not everyone.
- **Credential theft was understated, and the relay claim was backwards.** `isAgentFallbackError`
treats _any_ auth error as agent fallback (`ssh-connection-utils.ts:59-61`), so a MITM that rejects
publickey walks the user to the password prompt (`ssh-connection.ts:844`) and the private-key
**passphrase** prompt (`:834`), and `cachedPassword` is replayed without prompting on every
reconnect (`:709`). Meanwhile the relay upload matters less than assumed — the attacker already
owns their machine. The real client-side impact is the **return** direction: the attacker becomes
the host our workspace trusts, driving relay protocol frames, landing SFTP content in local
worktrees, and feeding agent-hook payloads in.
## Decisions
### D1. Read the user's `known_hosts`; write only to our own store
Consult the user's real `known_hosts` as a trust source — most developers already have their hosts
there from `ssh` and `git`, which is the entire migration story. Do **not** write to it: that file is
shared with every other SSH tool on the machine, and appending brings line-endings, permissions,
concurrent writers, and a corruption blast radius well beyond us.
Two consequences to own rather than discover:
- **Revocation does not propagate into our store.** `ssh-keygen -R host` clears `known_hosts` but not
our record. That is survivable for the ordinary rotation, because a `known_hosts` MATCH is now
decided before our store's mismatch — running the remedy we print and reconnecting works, which it
did not when the store was consulted first. What it does not cure is a host we only ever knew
ourselves, never written to `known_hosts`: there is no `ssh-keygen -R` for that one, so the
rejection names the store file directly. The "forget" action (D5) replaces that with a button; its
helper is deliberately absent until then, since an exported API nothing can reach is unverified in
production. Mismatch messaging must keep naming which source disagreed.
- **`ssh -G` on the HOME-divergent `-F` path suppresses `/etc/ssh/ssh_config`**
(`ssh-g-config-resolution.ts:44-52`), hiding site-wide `StrictHostKeyChecking yes` and
`GlobalKnownHostsFile`. On that path we must fail **strict**, never laxer than `ssh` would.
There is no `ssh`-only way out of this: `-F /dev/null` does NOT invert the exclusion, it reports
built-in defaults, so a probe built on it looks permissive on every machine. Verified against
OpenSSH 10.2p1. So the file is read directly, answering a deliberately weaker question — _could_
the site config be restricting host keys — where anything ambiguous (unreadable, an unresolvable
`Include`, the directive present at all) keeps the refusal. Only a site config that demonstrably
says nothing about host keys clears it, which is what stops the rule punishing every devcontainer,
`su` shell and Nix shell.
### D2. Ask `ssh -G`, do not reimplement config resolution
`ssh -G` reports `userknownhostsfile`, `globalknownhostsfile`, `stricthostkeychecking`,
`checkhostip`, `hostkeyalgorithms`, `fingerprinthash`, `hashknownhosts`, `updatehostkeys` and
**`hostkeyalias`** — with `Match` and `Include` already applied. `resolveWithSshG` exists and simply
does not read them yet.
`userknownhostsfile` is a space-separated list on one line, may contain `~`, and may contain
double-quoted paths with spaces. When `ssh -G` is unavailable (no `ssh`, non-zero exit, >5s timeout)
fall back to `~/.ssh/known_hosts` + `known_hosts2` — never to accept.
`HostKeyAlias` must be honoured: users tunnelling bastions through `localhost:port` depend on it and
would otherwise hit spurious mismatches. It appears nowhere in `src/main/ssh/` today.
**Lookup key.** Config resolution uses `configHost || label` (`ssh-connection.ts:660`) while ssh2
dials `effectiveHost` (`ssh-connection-utils.ts:188`). The `known_hosts` lookup must use
`HostKeyAlias` if set, else the **resolved hostname** — keying on the Orca label would miss every
existing entry.
**Two ordered lookup passes, not one candidate set.** Verified against OpenSSH 10.2p1: a non-default
port looks up `[host]:port` first, and if that finds nothing it retries the **bare** host. Crucially,
on that second pass a wrong key is downgraded to `unknown` rather than reported as changed. So the
passes are `[['[host]:port'], ['host']]`, and the fallback pass can only yield `match` or `unknown`.
Collapsing them into one set would give a spurious first-contact prompt to anyone who has a bare
line and connects on a non-default port; treating the fallback as authoritative would raise a false
change-of-key alarm.
**The entry condition to that second pass is the part that bites.** ssh runs it only when the
port-qualified lookup matched no plain entry of ANY key type — not "no match". Gating it on
"no match and no same-type mismatch" reaches the bare line when an off-port entry of another type
exists, and returns `match` where ssh prints `IDENTIFICATION HAS CHANGED`: an accept-a-changed-key
path, reproduced live. And the observations from each pass must not leak into the other, or an entry
found only on the fallback refuses a host ssh accepts as first contact.
**`HostKeyAlias` suppresses the port entirely.** ssh looks the alias up bare and never brackets it,
so an alias gets ONE pass regardless of port. Combined with the rule above, a stale `[alias]:port`
line would otherwise block the bare lookup ssh actually performs — turning the bastion case this
feature cites `HostKeyAlias` for into a hard failure.
**Hashed entries hash the candidate form, not the bare host** — `[example.com]:2222` is what gets
HMAC'd for a bracketed entry, so each candidate must be hashed separately.
**Multiple files union.** Any exact hit in any file wins; a disagreeing entry in another file does
not make it a mismatch. Confirmed live in both orderings.
**A `@cert-authority` line whose key equals the presented plain host key is not a match** — a CA line
only validates certificates. A normal line alongside it still decides. But ssh's verdict for a
CA-covered host presenting a plain key is `HOST_NEW`, not a failure: it connects. See D4.
### D3. Six outcomes, and type scoping is only safe with algorithm ordering
`match | mismatch | revoked | ca-only | unknown-type-known-host | unknown`.
Mismatch is scoped to the same key type: a host with only an RSA entry that presents ed25519 is not
"changed". Without scoping we would false-alarm nearly every RSA-era user on their first upgraded
connect, training them to dismiss the one warning that matters.
> **Corrected against a live client.** The premise above is wrong about OpenSSH, though the
> conclusion survives. `check_key_in_hostkeys` is not type-scoped at all: ANY non-marker entry for
> the host that is not byte-equal produces `HOST_CHANGED`. Verified on 127.0.0.1:2223 — `known_hosts`
> holding only `ssh-rsa` against an ed25519-only server prints `IDENTIFICATION HAS CHANGED` and
> refuses. So ssh does not avoid the false alarm by scoping; it avoids the _situation_ via
> `order_hostkeyalgs`, and hard-fails when the situation arises anyway. Our split into `mismatch`
> and `unknown-type-known-host` therefore only chooses the wording — both refuse, which is ssh's
> action. What the ordering below buys us is what it buys ssh: the situation mostly never arises.
**But scoping alone is a downgrade vector, and this is the correction that most changes the design.**
OpenSSH is safe here only because `order_hostkeyalgs()` reorders the client's proposed host-key
algorithms to put the types already in `known_hosts` first, and RFC 4253 gives the _client's_ order
priority — so a server cannot choose a type the client deprioritised. ssh2 negotiates ed25519 first
regardless. An attacker who cannot forge the RSA key on file simply presents ed25519 and receives a
friendly first-contact prompt instead of a hard failure.
Therefore: **set ssh2's `algorithms.serverHostKey` to lead with the key types already known for that
host.** Type scoping without algorithm ordering is not a safe design.
And when the presented type is unknown _while other types are known for this host_, that is
`unknown-type-known-host` — never a plain TOFU prompt. It must say we already hold a different key
for this host.
### D4. Outcomes
- **match** → connect silently.
- **unknown** → trust-on-first-use (see the phasing below for whether that is silent or prompted).
- **mismatch** → hard fail, no override in the failure surface.
- **revoked** → hard fail, always.
- **ca-only** → ~~hard fail~~ **REVERSED: treated as first contact.** See below.
- **unknown-type-known-host** → treat as suspicious, not first contact.
`StrictHostKeyChecking` is honoured: `no`/`off` accepts unknown but **never persists** and still
hard-fails changed and revoked; `accept-new` persists silently; `yes` denies unknown.
> **`ssh -G` does not report the spelling the user wrote.** StrictHostKeyChecking is rendered through
> `fmt_multistate_int`, which prints the first entry of `multistate_strict_hostkey`, and that table
> lists true/false before yes/no. So `yes` arrives as `true`, `no` and `off` both as `false`; only
> `ask` and `accept-new` pass through unchanged. Matching on `yes`/`no`/`off` matches nothing a real
> config can produce. `UpdateHostKeys` has the same shape (`true`, not `yes`).
**ca-only, reversed after review.** The rejection was stricter than ssh, and the blast radius was
mispriced. An SSH CA user holds ONE line — very often `@cert-authority *` — which matches every
candidate, so EVERY target failed, not just CA-signed ones, including on-demand runtime VMs, and
`StrictHostKeyChecking=no` did not help. `ORCA_SSH_FORCE_SYSTEM_TRANSPORT=1` is read from the
process environment, which an Electron app launched from the Dock or Start Menu does not have, so
the documented escape was unreachable for exactly the people who needed it. And OpenSSH's own
verdict for a CA-covered host presenting a plain key is `HOST_NEW`: it connects. ssh2 cannot
validate certificates at all, so refusing conceded nothing ssh was not already conceding.
The residual risk is accepted, not resolved: for a CA-protected host we take a plain key we cannot
tie to the CA. Certificate support is Phase 2 work. The `ca-only` outcome is still produced and
carried through the decision so the log shows a CA line was involved.
**An unreadable known_hosts connects but records nothing.** A file that EXISTS and will not open is
the absence of evidence, and the common trigger is not exotic — a Windows OneDrive Known Folder Move
placeholder while offline fails with a cloud-file error, not ENOENT. Refusing there broke an
ordinary corporate laptop while blaming a config file that was fine, and was asymmetric with our own
store, which degrades to "nothing trusted" and connects. ssh warns and treats the host as unknown;
so do we — but we write no record, so a first contact we could not check never becomes durable
trust. An ABSENT file is not this case: that is the normal state for a fresh profile and genuinely
means nothing is known.
### D5. Recovery must not live in the failure dialog
A "forget this host key" button _in_ the mismatch dialog is D4's rejected "trust anyway" with one
extra click. Recovery lives in target settings: a separate, deliberate surface, no auto-retry, and it
shows the stored fingerprint so the user is choosing knowingly.
Offer it only when **our** store is what disagreed; when `known_hosts` disagrees, forgetting our
record cannot unblock the connect. Messages, written to avoid naming internals:
> **Ours disagreed** — "The host key for `build-01` changed since you last connected from Orca. If you
> rebuilt or reprovisioned this machine, this is expected." → _Forget the saved key_ / _Cancel_
> **`known_hosts` disagreed** — "The host key for `build-01` does not match the entry in
> `~/.ssh/known_hosts`. `ssh` and `git` will refuse this host too. Run `ssh-keygen -R build-01`." →
> no button, because a button would not help.
### D6. Never prompt on a background reconnect
A prompt only means something when a human initiated the connect. `userInitiated` does not exist on
the connect path today and must be threaded through `connect → attemptConnect → doSsh2Connect`,
defaulting **false**.
Two traps: `useAutomationDispatchEvents.ts:203` and `pty-connection.ts:857` reach `ssh:connect`
without a human click — automation must pass `false`, but **terminal-pane focus reconnects must count
as user-initiated** or terminals die silently. And the denial string must avoid "authentication
failed"/"permission denied", or `isAgentFallbackError`/`isAuthError`
(`ssh-connection-utils.ts:46-61`) misclassifies it and the reconnect ladder retries a decision that
will never change.
### D7. Fail closed — three known fail-open shapes
1. The existing generation/disposed guard at `:1185` has the fail-open shape today: skip recording,
still `return true`. Post-fix that branch must **deny**.
2. A synchronous throw inside the verifier may not be caught by ssh2 — wrap and `verify(false)`.
3. Any non-`undefined` return accepts immediately (see Traps).
Plus: no prompt channel registered → deny (the load-bearing default lives in `doSsh2Connect`, not in
IPC, so a caller that forgets to wire it cannot accidentally accept); no window → deny; timeout →
deny; dialog dismissed → deny.
### D8. Store shape and scope
Accepted keys are scoped to **host + port + key type**, not target id — aliases point at different
machines, two targets can name one machine, and a re-created target must not lose trust.
The store is a **dedicated file**, not the main persistence blob (`persistence.ts:7088`): a settings
restore or rollback must not silently reset trust. Accept and mismatch events are logged.
`hostKeyFingerprint` is now security-relevant _and_ wire-relevant — it is an isolation namespace sent
to the host (`ssh-relay-session.ts:1298`, `managed-hook-owner-identity.ts:187`). It is `undefined` on
the system transport, so **no trust logic may key off it**, and its format must not change (see
Traps).
## Phasing — ship the defence before the dialog
Review made the case that the riskiest part of this change is not the security model but the modal.
Startup restore fires eager connects for _all_ previously-active targets in parallel (`App.tsx:1041`)
with a 15s timeout, while a prompt would live 120s — N unknown hosts means N stacked dialogs
outliving the timeout that already deferred them. Runtime-owned ephemeral VMs
(`ephemeral-vm-runtime-ssh.ts:31`) dial a freshly provisioned host with a brand-new key on every
launch. Paired-web connects run on the _host desktop_ (`runtime/rpc/methods/ssh.ts:32`), so the
dialog would open on someone else's screen while the web user watches a spinner.
**Phase 1 — no new modal.** Consult `known_hosts` + our store. `match` connects. `unknown` persists
silently with `accept-new` semantics and a passive notification naming the host and fingerprint.
`mismatch` (same type) and `revoked` hard-fail. This is the entire MITM defence with zero prompts,
zero startup storms and zero web hang.
**Phase 2** — the TOFU dialog, `StrictHostKeyChecking` honouring, `ca-only`, `userInitiated`
plumbing, and the D5 settings surface.
Carve-outs required before Phase 1 ships:
- **Runtime-owned ephemeral targets are exempt from persistence** — a new key every launch is
expected, not suspicious, and recording one would accumulate a row per launch that eventually
reads as a spurious change. Implemented via `target.owner?.type === 'on-demand-runtime'`.
- **RPC-originated connects: NOT needed in Phase 1, required in Phase 2.** The review asked for
these to fail fast rather than leave a paired-web user watching a spinner for the 120s prompt
timeout. That hang is only reachable if a prompt exists, and Phase 1 has none — the decision
function is pinned by a test asserting it never returns `prompt`. An RPC connect therefore behaves
exactly like a local one: it accepts and records on first contact, or fails immediately with the
host-key reason. Adding a fail-fast path now would introduce a failure mode for a hang that cannot
occur. It becomes load-bearing the moment the dialog lands, and is listed in Phase 2.
Worth noting for Phase 2: `runtime/rpc/methods/ssh.ts` already swallows the specific error and
rethrows `getPublicSshError(status)`, so a web client sees a generic failure rather than the
host-key reason. Pre-existing, but it means the Phase 2 message will not reach the web user
without a change there too.
## Traps
Each of these makes the fix silently do nothing. All confirmed in our tree.
1. **An `async` verifier defeats it entirely.** ssh2 does
`const ret = hashCb(key, verify); if (ret !== undefined) verify(ret)`. An async function returns a
Promise — not `undefined`, and truthy — so ssh2 accepts before our callback settles.
2. **Do not set ssh2's `hostHash`.** It hands the callback a hex digest and discards the raw blob we
must compare — and it would change `hostKeyFingerprint`'s format, which is a cross-version state
break, not a local refactor.
3. **The existing test mock calls `hostVerifier(key)` with one argument** and ignores the return
(`ssh-connection.test.ts:86-91`). Under an async verifier every connect test there breaks. The
mock must change — flagged deliberately, not rewritten silently.
4. **Validate the blob**: embedded algorithm name must match the line's key-type field; reject empty
decodes, empty salts, and hashed entries whose hash is not 20 bytes.
5. **`ssh-relay-live-connect.test.ts:59`** constructs a connection with no credential callback —
headless with no prompt channel must deny, not hang.
## Scope
**In scope, corrected:** IPv6 literals and `[host]:port` bracket parsing. Review was right that this
is a _parser_ requirement, not a scope call — getting it wrong means hosts `ssh` knows come back
`unknown`, which is the prompt-training harm D3 exists to avoid.
**Out of scope, with consequences stated:**
- **`CheckHostIP`** — OpenSSH defaults it off; we form candidates from the hostname only.
- **WSL** — `src/main/ssh/` has no WSL awareness; a distro's `known_hosts` is unreachable, so WSL
users get first-contact treatment for hosts they already verified.
- **`UpdateHostKeys`** — we read it and use nothing, so we never learn a rotated key, which makes D5
the routine path for key rotation rather than an exception.
- **Moving SFTP to the system transport** — correct direction, separate change.
## Test plan
**Parser** (against the file format, not our code's shape): plain lines, `host,host2` lists,
`[host]:port` used only when port ≠ 22, IPv6 literals, hashed `|1|salt|hash` with a real computable
vector, `@revoked`, `@cert-authority`, `*`/`?` globs, `!` negation vetoing a whole line, unrecognised
`@marker` skipping the line, malformed lines skipped not fatal, multiple keys per host, CRLF, blank
lines, comments, user file and global file disagreeing.
**Decision function**: all six outcomes; type scoping; revocation resolved before match regardless of
line order; every `StrictHostKeyChecking` value; `no`/`off` never persists.
**Algorithm ordering**: `algorithms.serverHostKey` leads with types on file — the test that makes D3
safe rather than merely scoped.
**Wiring**: unknown persists (Phase 1) without a prompt; match never notifies; mismatch fails with no
accept path; revoked fails; background reconnect denies; aborted connect settles pending verify
false; no prompt channel denies; runtime-owned targets are exempt; the denial string does not match
`isAuthError`; and — catching the worst regression — **the verifier returns nothing**, so a refactor
to `async` reddens a test rather than reaching a user.
**Checked against a live client, not just the file format.** Two assumptions the design leans on were
verified by running an OpenSSH 10.2p1 client against a real `sshd` on `127.0.0.1:2222` and recording
its verdict:
- **The bare-host fallback pass never reports a change.** With `StrictHostKeyChecking=accept-new`, a
bare line holding a _different_ key, dialed on a non-default port, made ssh connect and append a
new `[127.0.0.1]:2222` line — first contact, no `IDENTIFICATION HAS CHANGED`. Reporting `mismatch`
on that pass would refuse hosts ssh connects to happily, and would have looked like the cautious
choice.
- **`unknown-type-known-host` is ssh's own behaviour.** known_hosts holding `ssh-rsa` while the
server offers ed25519 makes ssh print `IDENTIFICATION HAS CHANGED` and refuse. So the rejection is
neither stricter nor laxer than ssh — and treating it as first contact, which a naive type-scoped
lookup does, is the laxer mistake. It also means `ssh-keygen -R` is the right remedy to name there.
## What Phase 1 shipped, and what review changed
The design above survived implementation. Every defect found afterwards was in the wiring, and the
pattern is worth recording because it repeats: **each one made us either blind or unusable, never
subtly wrong.**
Fixed after review:
1. **Our own store was type-downgradable.** The inline lookup filtered by key type first and could
only answer match/mismatch/unknown, so a record of a _different_ type read as `unknown`. D3's
downgrade, applied to the records we create ourselves. Stored types now also feed the algorithm
ordering — without that the guard is only half present.
2. **We keyed on the Orca label.** `ssh -G` echoes its own argument back as `hostname` when no Host
block matches, so for a manual target `resolved.hostname` _is_ the label — the one name D2
forbids. We consulted no entries at all.
3. **A refused key still walked the credential ladder.** ssh2 reports a denial as a generic auth
failure, so we went on to prompt for the passphrase and hand it to the host we had just refused.
Rejections are now a typed error recognised before any fallback.
4. **Fail-closed nearly became fail-always.** "No readable `known_hosts`" counted a _missing_ file the
same as an unreadable one, so a profile that had never connected — everyone's first run — would
have been refused, and the suite passed only because dev machines have a `known_hosts`.
5. **Ephemeral runtimes were refused for a policy they cannot satisfy.** The carve-out sat below the
incomplete-sources check, so a HOME-divergent environment turned on-demand runtimes off entirely.
A second review round, run against a live OpenSSH client and sshd rather than against the source,
found five more — and the pattern held: the two that mattered most were both cases where we refused
a host `ssh` connects to, and the worst single defect was that **`StrictHostKeyChecking` had never
been read correctly at all**, so a config saying `yes` was silently accepted AND persisted. See the
D2/D3/D4 corrections above. The lesson worth keeping: every one of these was invisible to unit tests
that fed the code the value a human writes, rather than the value the tool emits.
## Action items (STA-4319)
**Where the message actually lands.** Traced end to end, because a rejection the user cannot read is
a half-shipped feature. Fixed in this branch: the settings card clamped it to one line with no
tooltip, and the terminal reconnect overlay never asked for it at all. Still open:
- **The "Remote Hosts" status bar shows only `Error`.** `SshTargetStatusRow` does not receive the
error, so the status bar is a dead end for the most likely place a user notices the failure.
- **"Connect again" is the wrong advice for a decision that will never change.** The terminal overlay
now prints the reason underneath, but its call to action still invites an action that cannot
succeed. Telling a permanent rejection from a transient fault in the renderer needs a typed reason
on the wire rather than a string — a remote-wire-compatibility decision, so deliberately deferred.
- **Toasts carry Electron's `Error invoking remote method 'ssh:connect':` prefix.** The repo has
strippers for exactly this; no SSH call site uses one. Also worth noting sonner auto-dismisses in
4s, which is short for a message ending in a command the user is meant to copy.
**Before Phase 2:**
- **`UpdateHostKeys` (out of scope above, now the highest-value gap).** We read it and use nothing,
so a rotated key is a hard failure the user must resolve by hand. Combined with D5 this is the
routine path for key rotation, and it will be the most common way a legitimate user meets a
rejection. Decide whether Phase 2 honours it or D5's recovery surface absorbs it.
- **The web user never sees the reason.** `runtime/rpc/methods/ssh.ts` rethrows
`getPublicSshError(status)` on all three paths, and push events are redacted through
`getPublicSshState`, so a paired-web client always sees exactly `SSH connection unavailable`.
Pre-existing, but it makes the Phase 2 dialog message unreachable there without a change. Note the
redaction is not web-only: any target owned by a paired runtime environment is redacted, so a
_desktop_ user viewing a remote-Orca-server-owned host gets the same generic string.
- **RPC fail-fast** becomes load-bearing the moment the dialog exists (see Phasing).
**Known gaps that Phase 1 accepts, listed so they are choices and not surprises:**
- **WSL** — a distro's `known_hosts` is unreachable, so WSL users get first-contact treatment for
hosts they already verified through `ssh` inside the distro.
- **`CheckHostIP`** — candidates are formed from the hostname only.
- **Certificate validation** is still absent — a CA-covered host is now accepted on first contact
rather than refused (D4), so those users connect, but the CA itself verifies nothing for us.
- **`DEFAULT_SERVER_HOST_KEY_ALGORITHMS`** is a hand-copy of an ssh2 internal. A test pins it, so an
ssh2 upgrade that changes it fails CI rather than shipping — but the pin has to be honoured, not
deleted, because ssh2 throws `Unsupported algorithm` and every target stops connecting.
**Rollout:** the first release carrying this is the first time Orca can refuse an SSH connection at
all. Worth a staged rollout or a kill switch: the failure modes we could not find are, by the shape
of the five above, far more likely to be "a legitimate host is refused" than "a bad key is accepted".
@@ -1,159 +0,0 @@
# SSH reconnect: why the pane retry gets a byte tail, and what would actually change it
Status: investigation result. The obvious follow-up to PR #14844 was traced and **rejected**, and
tracing it turned up the actual root cause: checkpointed source recovery has never run on an SSH
reconnect. Both are recorded here — the rejected shape so nobody re-proposes it, and the verified
cause with the fix it implies.
## The shape of the problem
A reconnect remounts the pane (`tab.generation` is its React key), so the xterm is disposed with its
buffer and something must repaint it. Today that is a **byte tail**: `reattachSshPtySession` sends
`requireReplay: true` and the relay returns `RecentPtyOutputBuffer.read()` — the last 100KB, read
non-destructively, with no notion of what this client already consumed.
Two costs follow. Main's `@xterm/headless` model never sees those bytes (the tail bypasses
`onPtyData`), so it is stale by exactly the outage — which is what forces
`sshReconnectPaintsFromModel` to restrict the grid repaint to the alternate screen. And a shell loses
outage output past 100KB permanently.
## The proposal that does not work
"Make the pane-retry path request source recovery like `reattachKnownPtys` does." Mechanically this
is trivial — `sourceRecovery` is already an optional `pty.attach` param the relay parses, Path C
already calls the same `requestSshPtyAttach` helper and already parses the response field. The
required checkpoint state also survives a transport drop, in the module-level `recoveryByTarget` map
(`ssh-pty-consumer-recovery.ts:17`), reachable from `connectionId` because `connectionId === targetId`.
It still fails, three ways:
1. **The relay answers `'existing'` before it looks at the recovery argument.**
`relay-pty-source-publication.ts:99-109` short-circuits on a same-`clientId` attach, and a
reconnected client presents the same id — see the root-cause section below, where this turns out
to be the whole story rather than an obstacle specific to this proposal.
2. **A failed `reattachKnownPtys` deletes the checkpoint on purpose** (`ssh-relay-session.ts:3006-3007`)
and detaches the lease (`:3008`). The pane retry runs _after_ that, so it would present
`checkpointUnavailable`, which the relay converts to `restoreRequired`
(`relay-pty-source-publication.ts:124-130`) and the provider converts to
`SSH_SESSION_EXPIRED_ERROR` (`ssh-pty-provider.ts:103-107`). We would trade a blank-pane-with-tail
for a **killed session**.
3. **Wrong payload shape.** Recovery replays only the post-checkpoint delta
`(acceptedSourceEndSu → receivedEndSu]`. The byte tail is a screen snapshot for a _fresh, empty_
xterm. Even a successful recovery returns roughly nothing in the common case, and the pane stays
blank.
These two mechanisms answer different questions. Recovery keeps main's model whole; the tail repaints
a new terminal. Substituting one for the other is a category error.
## A correction worth recording
The motivating argument was "`requireReplay` is optional, so older relays ignore it and still show
blank panes." **That is wrong for the SSH relay.** The client deploys and launches its own relay
build into a version-scoped directory (`ssh-relay-deploy.ts:231`, `:594`), and `validateGrant`
rejects any grant whose `serverBuildId` differs from the expected one
(`ssh-pty-consumer-session.ts:58-65`, rationale in-code: _"client and relay ship in one build"_).
Client and SSH relay are version-locked; mixed versions do not occur on this channel. The
independent-update rule in `remote-wire-compatibility.md` still governs remote _runtime_ hosts — just
not this one.
So there is no old-host population to rescue, and the urgency that argument created was false.
## ANSWERED: checkpointed recovery never runs on an SSH reconnect
The question above was "does the reconnecting client present a new `clientId`?" It does not, and the
consequence is that the whole checkpoint mechanism is dead on this path. Every link verified:
1. **The client keeps its id.** A reconnect calls `Dispatcher.setWrite`
(`src/relay/dispatcher.ts:149-157`), which reuses `this.primaryClient` — including its `id` — and
replaces only the writer. The dispatcher refuses to detach the primary. This is already stated
in-repo at `src/relay/pty-handler.ts:1736-1742`.
2. **So `activate()` short-circuits.** `relay-pty-source-publication.ts:99` tests
`current?.clientId === context.clientId` and returns `'existing'` at `:108`. The `rotateDelivery`
recovery path at `:118-142` is reachable **only** when the ids differ — i.e. never, here.
3. **So the relay returns no `sourceRecovery`.**
4. **So the client abandons.** `finishSourceRecovery` (`ssh-relay-session.ts:2766-2785`) fails its
`!pendingRecovery` guard, calls `abandonPtySourceRecovery`, and returns false — which cancels the
delivery and deletes the checkpoint (`:3006-3008`).
5. **So the pane retry opens fresh and gets the byte tail**, via the `requireReplay` fix.
The byte tail is therefore not a fallback. It is the only path that has ever run for an SSH
reconnect, and the flow-control/checkpoint machinery is inert on this path.
That also explains the original blank-pane bug exactly: the relay concluded "this client already
holds the stream" because, by its own identity rule, it does.
### The fix this implies
Give a reconnected primary a distinguishable identity — a transport generation on the client record,
bumped in `setWrite` — and have `activate()` compare it alongside `clientId`, so a reconnect takes
`rotateDelivery` instead of `'existing'`.
Why this is the tractable shape:
- **No wire change.** `RequestContext`, `setWrite` and the publication are all relay-internal.
- **No compatibility exposure.** Client and relay ship in one build and are version-locked.
- **It does not disturb the invariant that broke three earlier attempts.** Deliveries still outlive
their clients; nothing retires on `onClientDetached`. The delivery is _rotated on re-attach_,
which is what the recovery design already intends and what its tests already cover.
**UNVERIFIED and to be checked before implementing:** that `rotateDelivery`'s preconditions hold at
that moment (the checkpoint's `deliveryToken`, `clientGeneration`, `ownerGeneration` and
`ptyIncarnation` must match the live identity, `:124-128`); that `outputFlowControl` is granted on
the reconnected session; and what a rotation implies for the _renderer_, which still remounts with an
empty xterm and needs a screen, not a post-checkpoint delta. Recovery keeps main's model whole — it
does not by itself repaint a fresh terminal, so the tail may still be wanted for the pane even once
the model stops going stale.
## Do not start at `onClientDetached`
Three attempts failed there, each plausible until run:
- Retiring the delivery on `dispatcher.onClientDetached` **breaks checkpoint recovery** (10 tests).
A delivery outliving its client is deliberate — it is what lets a client resume from a checkpoint.
- Retiring without `session.cancelDelivery()` orphans the credit ledger's one-upstream-owner slot;
the next open throws `PTY source delivery already has an upstream owner`. Seen live as a toast and
a blank pane.
- Comparing `record.identity.clientGeneration` to the request is impossible: that value is
client-supplied via `pty.openClient`, and `RequestContext` carries no generation of its own.
## The lead that survives
`reattachRejectedPty` (`ssh-relay-session.ts:1957-2004`) is an existing **single-PTY** entry point
into the `reattachKnownPtys` machinery, taking `(relayPtyId, mux, providerGeneration, mode)` and
driving recovery with `targetedDeliveryRecovery`. If per-pane recovery is wanted, that is the hook —
and it does not involve the pane-retry path at all. Unverified whether it is reachable at the moment
the renderer retries.
## Preconditions, unchanged
The SSH e2e lane must be green and triggering on **source** changes before any of this is attempted.
It was skipping for 15 specs; four regressions reached a user during that window.
## Resolved: a disposed pane killed its successor's new shell
A pane rebuilt during its first spawn uses the same reservation key, so main can return the
same PTY to both transports (#19386, #22578). The disposed transport must keep that shell
while its tab and layout leaf remain and its execution host's workspace is not being deleted.
A live transport refusing the id still retires it (#11003). This applies to local, WSL and SSH
IPC terminals; the remote-runtime transport has no corresponding kill.
The remount trigger in the Scan-22 user report remains unknown. A retained shell can outlive
its tab if the tab closes before a successor binds it; keeping potentially owned work follows
the SSH execution boundary.
## Open: the pane behind a preserved tab does not always rebind
The merge now keeps a local tab the host has never been told about, so the tab and its title survive
a reconnect. The reattach behind it does not, reliably — measured at three runs in four against the
Docker-SSH lane. When it misses, the store holds the tab, the tab bar renders it, and the pane never
rebinds: the "frozen tab" shape the original report described, one layer down from the deletion that
used to cause it.
Deliberately NOT asserted in `ssh-reconnect-tab-destruction.spec.ts`. A one-in-four flake in the lane
that exists to catch this class costs more than it proves — the lane stops being trusted, which is
exactly how the earlier silent-skip failure happened. Tab survival is asserted there and is
deterministic; the liveness gap is recorded here instead.
Worth checking first, since it is the same shape as everything else in this file: the tab is absent
from the host snapshot, so whatever drives the per-tab reattach after an apply may simply not know to
reattach a tab the snapshot never mentioned.
@@ -1,81 +0,0 @@
# Task provider identity RPC validation
The automation RPC identity schema follows `src/shared/task-provider-identity.ts`,
re-exported by `src/shared/task-source-context.ts`. Only GitHub requires fields beyond
`provider`: `owner` and `repo` are strings, and `host` is optional. GitLab, Linear,
and Jira fields are optional nullable strings. Requiring a GitLab project or a
Linear workspace/Jira site would contradict the domain type and account-wide scopes.
The discriminated union validates these existing field types without trimming,
coercing, or stripping identity fields. Unknown fields pass through as they did under
`z.custom`, including fields from newer clients. Absent and explicit-null identities
remain distinct. The schema does not infer providers from owner/repo or require a
git worktree, repository slug, or local execution host for a source context.
## Producer census
Paths below are relative to the repository root. Searches covered production
`providerIdentity`, `TaskProviderIdentity`, `sourceContext`, and
`linkedTaskSourceContext` uses across desktop, shared code, mobile, and CLI.
| Producer or forwarding path | Populated verdict |
| ------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `src/shared/project-host-setup-projection.ts`: `getProjectProviderIdentity` | GitHub owner/repo populated together, or no identity. Supplies project identities consumed by desktop and migration. |
| `src/renderer/src/components/task-page-source-context.tsx`: `getTaskPageRepoSourceContext` | GitHub fields populated through projection, or null. GitLab uses explicit provider and `buildGitLabProviderIdentity`; projectId/project/webUrl populated, namespace can be null. |
| Same file: `buildGitLabProviderIdentity` | GitLab fields come from project path/host; missing path components become null. No required GitLab field is invented. |
| `src/renderer/src/hooks/composer-state/source-context-state.ts` | Derived GitHub context carries a complete project identity or null. Jira folder/project-group context explicitly has null identity. Draft/linked contexts are forwarded. |
| `src/renderer/src/components/use-task-page-source-availability.ts` | Linear workspaceId/name and Jira siteId/URL can be null for account-wide selections; team/project fields are not populated. Valid under the existing optional-field contract. |
| `src/renderer/src/components/task-page-jira-item-source-context.ts` | Bound issue context populates siteId, siteUrl, projectKey. |
| `src/renderer/src/components/new-workspace/use-jira-url-source.ts` | Bound URL issue context populates siteId, siteUrl, projectKey. |
| `src/renderer/src/components/worktree-jump-palette-create-worktree.ts` | Linear teamId/key populated; workspaceId/name may be null. |
| `src/shared/task-source-context.ts`: normalize/build functions | GitHub missing owner/repo becomes null identity; other providers' missing fields become null. Provider mismatch becomes null, never an inferred provider. |
| `src/cli/handlers/automation-handler-flags.ts`, `src/cli/handlers/automations.ts` | Explicit JSON source-context input is normalized before create/update. GitHub fields populated or null identity; other fields nullable. Omitted/null context preserved by flag handling. |
| `src/main/persistence/scheduling-automations/automation-context-migration.ts` | Builds source context from projected complete GitHub identity, or null context. |
| Desktop automation save/scoped-list/host clients and web transport | Forward existing source contexts, not new identity constructors. `automation-orca-save.ts` forwards the current automation context or null. Legacy arbitrary malformed RPC input is deliberately rejected by the new schema. |
| Mobile | No task-provider identity/source-context constructor or sender found. `mobile/src/components/new-workspace-project-targets.ts` uses project identity solely for display. |
## Compatibility evidence and limits
Read `docs/reference/remote-wire-compatibility.md` before changing validation.
A source search against release tag `v1.4.199` also finds no mobile
`sourceContext`/`linkedTaskSourceContext` sender; its sole `providerIdentity` use
is the display-only project target above. The released CLI flag reader also calls
`normalizeTaskSourceContext`. This is source-level evidence for the checked release,
not a claim to have executed every historical mobile binary.
No shipped mobile producer with a newly rejected payload was found. No new required
field was added to the domain contract. GitLab, Linear, and Jira discriminant-only
identities remain valid. Folder-workspace null/absent identities remain valid on
both local and SSH hosts. No execution/status logic or client-side parsing changed.
## Regression evidence
`src/main/runtime/rpc/methods/task-provider-identity.test.ts` checks unchanged valid
identities for all four providers, required GitHub fields, every declared field's
type, optional/null non-GitHub fields, unknown-field preservation, explicit GitLab
discrimination with owner/repo present, local/SSH folder contexts, and update patches.
The focused run passed 81 tests. Temporarily replacing GitHub's field validators
with optional `z.unknown()` validators (discriminant-only acceptance) caused 17
failures and 64 passes. The mutation was restored before running the gates.
Counts re-measured at `cf4f77f` after the blank-field commit added seven tests;
the earlier 74/16/58 figures described the commit before it.
## Gate results
All commands ran with `ORCA_BACKGROUND_LAUNCH=1`.
- `pnpm tc`: exit 0; completed the repository typecheck runner.
- `pnpm exec vitest run src/main/runtime/rpc`: exit 1; 277 files passed,
one failed; 2,462 tests passed, one failed, one skipped. The only failure was
the unrelated `structured-agent-session-adoption-replay.test.ts` hitting its
5,000 ms timeout. All identity tests passed.
- `pnpm --dir mobile typecheck`: exit 0; `tsc --noEmit` passed.
- `pnpm run check:code-quality:changed`: exit 0; zero new code-quality,
type-aware, or React Doctor findings across the two changed code files.
- Isolated retry of `structured-agent-session-adoption-replay.test.ts`: exit 0;
one test passed, with the test body completing in 358 ms.
- Full RPC retry with `pnpm exec vitest run src/main/runtime/rpc --maxWorkers=4`:
exit 0; all 278 files passed, 2,470 tests passed, one skipped (97.24 seconds).
The bounded-concurrency rerun resolved the timeout without changing test code.
@@ -1,85 +0,0 @@
# Terminal artifact grant integrity
Clicking a path in terminal output mints a short-lived grant over one file in a
world-writable temp directory. Between minting the grant and using it, anything on the
host can replace that file. This page is the contract for the checks that catch such a
replacement — and, more importantly, for the windows they do **not** close.
Read this before weakening a check in
`src/main/runtime/runtime-file-commands-terminal-artifact-access.ts`, before assuming
the sequence is atomic, or before extending the guarantee to remote hosts.
## The stat identity is weaker than it looks
A grant pins the file as `dev:ino:nlink:size:mtimeMs`. Four of those five fields
survive an unlink-and-recreate at the same path, measured on Linux (Node 24) by
replaying that exact sequence:
| Field | After `rm` + recreate in the same directory |
| --------- | ------------------------------------------------------------------------ |
| `dev` | unchanged — same filesystem |
| `ino` | **reused 100% of the time** (3000/3000); ext4 hands back the freed inode |
| `nlink` | `1` before and after |
| `size` | unchanged whenever the replacement is the same length |
| `mtimeMs` | **quantised to 1 ms** — the kernel's coarse clock advances once a tick |
So for a same-length replacement the whole identity string collapses to a 1 ms
timestamp race. The full string collided in 63.7% of back-to-back iterations, 19.8%
with a 0.5 ms gap, and 0% at ≥1 ms. The same probe on macOS collided 0 times in 2000 —
no inode reuse, nanosecond mtimes — which is why this only ever showed up on Linux CI.
`ino` contributes no discriminating power against the exact case it is there to catch.
Do not add `ctimeMs` or `birthtimeMs` hoping to fix this: they come from the same
coarse clock and collide in the same window. `bigint: true` stats do not help either —
the precision loss is in the stored kernel timestamp, not in Node's `Number`.
## The content digest sits alongside the identity, and cannot replace it
Local grants also pin a sha256 of the artifact's bytes, read from the **same open
handle** as the stat so the two describe one inode with no gap between them.
The identity string stays exactly as it is because it is a wire contract, not a
host-local detail: `src/relay/fs-handler-terminal-artifact.ts` recomputes that same
`dev:ino:nlink:size:mtimeMs` format field-for-field from its own stat and compares it
against the `expectedStatIdentity` the host sends. Changing the format would have a new
host publishing a string an older relay can never reproduce, failing every remote
artifact read as `terminal_file_grant_stale` — a break that reaches old peers with no
wire-schema change at all. See [remote wire compatibility](./remote-wire-compatibility.md).
## What is still open
The digest narrows these checks. It does not make any of them atomic.
- **Read and preview — closed.** The bytes returned to the caller are the same
in-memory buffer that was digested, with no re-read in between, and the handle pins
the inode for the whole operation. A swap landing after the `open` leaves the handle
on the granted inode, so the granted content is what is served; a swap landing before
it fails the digest.
- **Write — a window survives.** Between the final pre-rename verification and the
`rename()` itself, the target path can still be swapped, and the rename clobbers
whatever is there. POSIX `rename()` has no "only if the target is still inode X"
form; Linux's `renameat2(RENAME_EXCHANGE)` would close it but is not portable and is
not exposed by Node. Narrowing this further means changing the commit strategy, not
adding another check before it.
- **Artifacts over 10 MB fall back to stat-only.** They digest to `null`, so only the
identity guards them. Nothing leaks: every read, preview, and write path rejects on
size before reading. The weakness is unreachable, not fixed — a later cap change
could expose it.
- **Remote and SSH grants are untouched.** They keep the stat-only check, with the full
1 ms weakness on whatever filesystem the relay runs. Closing it needs a negotiated
capability so an older relay is never sent a digest it cannot verify.
## What is proven, and what is inferred
The filesystem numbers above are direct measurements. That identical stats plus changed
content previously returned the swapped bytes, and now do not, is pinned by
`orca-runtime-files-terminal-artifact-swap-detection.test.ts`, which replays the first
stat seen for a path so the collision is deterministic rather than a 1 ms coin flip.
The join between the two is a chain, not a single observation. This defect surfaced as
an intermittent failure of `orca-runtime-files-terminal-artifact-io.test.ts` on
`rejects stale absolute terminal artifact previews before returning changed content`,
which swaps an 8-byte artifact for 8 different bytes. No one has instrumented a Linux
runner to prove that a specific CI failure was a same-tick mtime collision; the
conclusion rests on every ingredient being measured separately. Treat it accordingly if
a future failure does not fit.
@@ -1,106 +0,0 @@
# Terminal latency investigation (2026-09-21)
The September 21 scheduled report has five latency violations: three restores
above 1,000 ms, a worst key around 2,091 ms, and timer drift around 2,034 ms.
These remain failures under the [historically calibrated budgets](terminal-perf-report-budgets.md).
Passing the looser Electron assertions does not establish that the report passed.
## Historical boundary
Slow restores predate September: the July 20 scheduled run measured a 1,226.6 ms
Latin restore; August 1 measured 1,492 ms. Earlier June/July logs did not contain
usable summary rows and their artifacts expired. There is no established good/bad
application revision boundary, and no completed git bisect. A newly visible report
failure is not, by itself, evidence of a newly introduced application regression.
The scheduled workflow uses one Playwright worker. Parallel Electron workers do
not explain these particular failures. Repeated macOS runs did not reproduce the
Linux stalls (15 targeted samples: restore 136–236 ms, hidden worst key ≤23.6 ms).
## Controlled Linux experiments
All comparisons use the same application build within their run, one worker,
real PTYs, and the original terminal workload. Diagnostic tracing can perturb
measurements, so it identifies the mechanism rather than setting new budgets.
| Comparison | Evidence | Finding |
| --- | --- | --- |
| Default versus disabled background throttling | [35657300611](https://github.com/stablyai/orca/actions/runs/35657300611) | 14/20 restores exceed 1 second; four typing measurements approach 1 second. `setBackgroundThrottling(false)` does not remove the stalls. |
| Native browser trace | [35658595456](https://github.com/stablyai/orca/actions/runs/35658595456) | Renderer waits roughly 960–1,010 ms in `LayerTreeHost::WaitForCommitCompletion`. Some typing samples contain consecutive waits. |
| Current flags versus flags preceding `c64777d1bcd` versus SwiftShader | [35659601151](https://github.com/stablyai/orca/actions/runs/35659601151) | All three configurations still trigger undrawn-frame throttling. Reverting the May 26 flags is not a demonstrated fix. |
In the graphics comparison, current flags had one of six restores above 1 second;
the earlier flags had three and a 1,036 ms worst key. SwiftShader had six measured
restores between 301 and 396 ms, but still contained one-second native waits after
the restore measurement ended, and one 202 ms timer-drift violation. Its faster
restore numbers are insufficient evidence of a fix.
## Native mechanism
The trace shows the renderer blocked inside:
```
ProxyMain::BeginMainFrame
Commit
ProxyMain::BeginMainFrame::commit
LayerTreeHost::WaitForCommitCompletion
```
During the gap, Viz repeatedly emits `SendBeginFrameDecision` with
`reason: ThrottleUndrawnFrames` and `should_send: false`. The graphics comparison
recorded 346 such decisions with current flags, 626 with the earlier flags, and
401 with SwiftShader. Raster work was already ready before the wait ended.
The [matching Chromium source](https://github.com/chromium/chromium/blob/150.0.7871.250/components/viz/service/frame_sinks/compositor_frame_sink_support.cc)
limits begin frames to once per second when too many submitted frames remain
undrawn. Input can block the renderer waiting for the next compositor commit;
this is not a one-second xterm parse or proof of a second of CPU consumption.
The same throttle is present in Chromium 148.0.7778.218 (Electron 42.3.3) and
146.0.7680.177. This source comparison does not prove identical runtime behavior.
## Visibility control
[Run 35660847455](https://github.com/stablyai/orca/actions/runs/35660847455)
compared ten hidden-window samples with ten visible-window samples on the same
isolated Xvfb runner, using SwiftShader in both modes. Actual window visibility
was recorded. The background terminal panes remained hidden in both modes.
| Measurement | Hidden window | Visible window |
| --- | ---: | ---: |
| Undrawn-frame throttle decisions | 1,251 | 0 |
| Largest worst-key latency | 3,062.8 ms | 30.5 ms |
| Largest timer drift | 3,111.8 ms | 67.0 ms |
| Restore range | 213.8–1,862.2 ms (9 completed) | 223.4–734.3 ms (10 completed) |
| Electron tests passed | 9/10 | 10/10 |
All ten visible-window samples satisfy the existing latency limits. This isolates
the never-presented Linux test window as the trigger for the reproduced native
stalls. It does not establish a newly introduced application-code regression or
prove that every historical outlier had the same cause.
## Full scale validation
[Run 35662787327](https://github.com/stablyai/orca/actions/runs/35662787327)
passed all 21 scenarios and all 32 strict report rows, with zero skipped,
unexpected, or retried tests. All 21 scenarios recorded successful isolated-display
presentation. This run restored the original graphics flags and removed profiling.
| Metric | Largest measurement | Unchanged report limit |
| --- | ---: | ---: |
| Median typing | 15.5 ms | 25 ms |
| Worst key | 45.7 ms | 300 ms |
| Hidden-output restore | 395.7 ms | 1,000 ms |
| Worktree revisit | 50.3 ms | 300 ms |
| Scroll | 54.7 ms | 150 ms |
Timer drift and queue/drop checks also passed. Coverage includes 100-pane
same-workspace and cross-workspace redraws, 50 real PTYs under held-ACK pressure,
and the original plain/Latin/title/rich-model hidden-output scenarios. No workload
or performance limit changed.
The correction presents the benchmark window only when explicitly enabled inside
`xvfb-run` on a GitHub-hosted Linux runner. Ordinary local automation stays
windowless; production launch policy and hidden-terminal delivery remain unchanged.
For comparable Linux latency evidence, use the Terminal Perf workflow: a never-
presented local Linux window can still encounter the same compositor throttle.
Temporary profiling hooks and the comparison workflow were removed before the PR.
@@ -1,59 +0,0 @@
# Terminal performance report budgets
The saved-report gate is a performance regression gate, deliberately stricter than
some Electron test timeouts. Passing the Electron suite does not establish that a
slow sample is acceptable. CLI and HTML reports use the same policy in
`config/scripts/terminal-perf-report-budgets.mjs`.
## Historical evidence (2026-09-21)
Sample: 11 full scheduled Ubuntu runs, 32 annotation rows per run, August 1 through
September 21 (352 rows). Runs include failures, rather than selecting only green
runs. These are samples across different revisions and shared runners, not a
controlled A/B experiment or a statistical tail-latency estimate. Original June
logs returned HTTP 410 and could not establish an original baseline.
Values below are milliseconds except peak queue chars (JavaScript character
counts, not process memory bytes). Maximum median spans every typing scenario;
peak queue spans the active/revisit ACK-pressure scenarios.
| Date / run | Baseline median | Maximum typing median | Latin restore | Hidden 25-pane worst key | Peak queue chars |
| ----------------------------------------------------------------------- | --------------: | --------------------: | ------------: | -----------------------: | ---------------: |
| [2026-08-01](https://github.com/stablyai/orca/actions/runs/30693362353) | 11.7 | 12.7 | 1492.0 | 61.2 | 1867776 |
| [2026-08-03](https://github.com/stablyai/orca/actions/runs/30802935899) | 7.1 | 9.4 | 414.0 | 12.8 | 294912 |
| [2026-08-15](https://github.com/stablyai/orca/actions/runs/31875140053) | 8.9 | 13.8 | 1297.5 | 15.4 | 2523136 |
| [2026-08-24](https://github.com/stablyai/orca/actions/runs/32708128219) | 12.2 | 12.2 | 463.6 | 16.3 | 3227648 |
| [2026-09-01](https://github.com/stablyai/orca/actions/runs/33488510218) | 6.6 | 6.8 | 302.2 | 11.5 | 360448 |
| [2026-09-07](https://github.com/stablyai/orca/actions/runs/34102348134) | 7.5 | 11.3 | 328.2 | 269.6 | 2818048 |
| [2026-09-11](https://github.com/stablyai/orca/actions/runs/34580494139) | 9.5 | 10.4 | 233.5 | 186.2 | 2441216 |
| [2026-09-15](https://github.com/stablyai/orca/actions/runs/34948597774) | 10.3 | 12.6 | 282.2 | 1178.4 | 2818048 |
| [2026-09-17](https://github.com/stablyai/orca/actions/runs/35201162712) | 10.3 | 12.7 | 1182.2 | 15.7 | 2998272 |
| [2026-09-19](https://github.com/stablyai/orca/actions/runs/35432611564) | 7.6 | 10.3 | 236.7 | 1435.5 | 2588672 |
| [2026-09-21](https://github.com/stablyai/orca/actions/runs/35579708852) | 7.8 | 11.6 | 1640.7 | 2090.6 | 2523136 |
## Decisions
- Median typing: **25 ms**, tightened from 75 ms. The largest observed median was
13.8 ms, leaving about 81% headroom without accepting a sustained 5x slowdown.
- Worst key: retain **300 ms**, including stress scenarios. Typical per-scenario
worst-key samples were tens of milliseconds; isolated 1–3 second samples are
failures to investigate, not a reason to adopt the e2e 3–3.5 second ceiling.
- Revisit: retain **300 ms**. Median of the 11 revisit samples was 160.7 ms;
the 1,030.2 ms outlier remains a failure.
- Restore: retain **1,000 ms**. Median restore per scenario ranged from 123 to
493.9 ms. Observed outliers up to 1,706.4 ms do not justify a 4 second budget.
- Timer drift: retain **150 ms** and the pre-existing **3,500 ms** allowance only
for injected same/cross-workspace redraw scenarios. No new scenario receives
the broad allowance. This retains the CLI gate's existing policy in HTML too.
- Scroll: retain **150 ms**. Dropped backlogs: retain **zero**.
- Current queue: retain **2 Mi characters** everywhere. Only the transient peak
in active/revisit ACK-pressure scenarios gets **3.5 Mi characters** (3,670,016).
The maximum observed peak was 3,227,648 (3.08 Mi), leaving about 14% headroom.
The old 2 Mi peak budget rejects ordinary deliberately held-ACK bursts; the
proposed 5 Mi e2e ceiling was unnecessarily loose. Other scenarios keep 2 Mi.
For the linked September 21 report, the two peak-queue failures are corrected;
the three slow restores, worst-key stall, and timer stall still fail. This change
does not claim to fix those stalls or make that run green. Re-evaluate future
budget changes against recorded measurements; do not set limits just above a new
failure or mirror relaxed test timeouts.
-30
View File
@@ -1,30 +0,0 @@
# Terminal startup timing
For #19333, enable the renderer's opt-in recorder in its DevTools console before opening a new terminal:
```js
localStorage.setItem('orca:terminal-startup-timing', '1')
```
Remove the key to disable it. Existing sessions are unaffected. To capture the existing host spawn phases, start the host with `ORCA_PTY_SPAWN_TIMING=1`. Do not restart a host with active work just to enable diagnostics.
The renderer emits one `terminal_startup_timing` breadcrumb per transport callback generation through the existing local diagnostic channel. In `main.trace.ndjson`, find the `renderer.breadcrumb` record whose `breadcrumb.name` matches. The host's existing console timing line also becomes a `pty.spawn.timing` trace record. Correlate available PTY IDs; renderer generation distinguishes retries. Compare elapsed durations within each process, not wall clocks across hosts.
Renderer offsets are monotonic milliseconds from callback-generation creation immediately before a transport operation:
| Field | Observation |
|---|---|
| connected | Transport connection callback accepted for the current generation |
| liveData | First nonempty live delivery, including control-only output |
| submitted | First live batch sent to the renderer output scheduler |
| writeStarted | Scheduler invokes the batch's pre-write callback |
| parsed | Xterm invokes that batch's completion callback |
| renderEvent | First public xterm render event after the batch starts writing |
A render event can precede the parse callback. These observations do not establish the first printable glyph, physical screen presentation, React mount time or click-to-paint latency. Replay and synthetic reset writes do not claim the first live batch. A replay or resize can still contribute to a render event after a live write, so the event is temporal evidence rather than attribution to exact content. Hidden or restored panes may never submit a live batch; missing fields remain missing. A queue-cap warning can inherit the pre-write callback while discarding the original batch’s parse callback. In that case writeStarted/renderEvent describe incomplete pre-write activity, not successful delivery of the original batch; outcome cannot be observed without parsed.
The recorder ends after connection, parse and render observations, or on replacement, disposal, error or a ten-second diagnostic deadline. It retains only phase numbers and identifiers, with one timer and at most one render listener while enabled. It does not retain terminal text, commands, credentials or transcript buffers. Disabled recording adds no listeners or timers.
Host `phaseDurations` preserve the current phase boundaries: the timer starts after initial ownership lookups and logs before all commit/serializer work finishes. `totalMs` is that measured interval, not full IPC latency. `provider_spawn` includes provider call and surrounding reconciliation; it is not raw process creation time. The enclosing trace record is a diagnostic snapshot, not a span covering that interval.
Reliability invariant: diagnostics must not change terminal output, delivery credits, provider ownership or spawn outcome. Failure source: Windows OMP first-paint report #19333. Oracle: opt-in recorder tests distinguish queued, parsed and render milestones; existing live-delivery and synchronized-output suites preserve output behavior. No matching startup diagnostic reliability gate exists; full click-to-physical-presentation remains an explicit validation gap. Native, daemon, WSL and SSH execution remain host-owned; this adds no wire fields or remote process queries. Mobile has no recorder change. macOS/Linux/Windows renderer timing uses the same public xterm events; physical-device timing requires a separate capture.
@@ -1,77 +0,0 @@
# Resolving Windows `.cmd` shims past cmd.exe
Node refuses to spawn a `.cmd`/`.bat` target without a shell (the
CVE-2024-27980 mitigation), so `resolveSpawn` has to make `cmd.exe` the program
and hand it `/d /v:off /s /c "<caret-escaped argv>"`. For an agent CLI that
means a long `cmd.exe /c` line whose caret-escaped payload is natural-language
prompt text — which Microsoft Defender for Endpoint's command-line model scores
as obfuscation. `codex.cmd` appeared in the spawn cluster of an MDE incident
against Orca for exactly this reason.
`src/shared/child-process/windows-cmd-shim-resolution.ts` sidesteps it. npm's
`cmd-shim` and pnpm's `@zkochan/cmd-shim` generate files whose entire body is
"find a Node interpreter and run this script". Reading one lets `resolveSpawn`
spawn `node.exe <script> <args…>` directly: no cmd.exe in the tree, and no
caret escaping at all.
## What resolution changes
Only `runProcess` / `spawnProcess` callers. Two things people expect it to
cover, and it does not:
- **The interactive terminal.** `src/main/daemon/pty-subprocess/native-pty-spawn.ts`
calls `pty.spawn` directly, so typing `codex` in an Orca terminal is
completely unaffected.
- **Orca's own hook wrappers** (`codex-hook.cmd` and friends). These are batch
files Orca writes, matching none of the generator shapes, so they keep the
cmd.exe path. They are addressable — we generate them — but not by this
module.
## Adding a shape
Four shapes are recognised, each transcribed verbatim from a real install into
`src/shared/child-process/__fixtures__/windows-cmd-shim-bodies.ts`. If you add a
fifth, add its real body there too. A shape guessed from documentation is not
evidence.
The rule for the parser is all-or-nothing: the whole canonicalised body must
match end to end, and anything unrecognised returns null and keeps the cmd.exe
path. **A mis-resolution silently runs the wrong program or drops arguments,
which is far worse than an EDR alert** — when in doubt, refuse.
Resolution also refuses a captured path that is absolute, drive-relative
(`D:evil.js` — `win32.isAbsolute` says false, but `win32.resolve` leaves the
shim directory), or contains `% ^ & | < > " :` or a line break; a script or
target that is not on disk; an interpreter-less target that is not `.exe`/`.com`;
and a program path that is not absolute.
Refusing every `:` cannot cause a false refusal. Windows reserves the character
within a path segment, so a relative path cannot contain one — the only
spellings that can are drive-qualified, an alternate data stream (`a.js:zone`),
or a `\\?\` device path, and the last is already refused as absolute.
## Kill switch
Set **`ORCA_DISABLE_CMD_SHIM_RESOLUTION`** to any non-empty value in the
environment a child is spawned with, and every `.cmd` goes back through
`cmd.exe /c` unchanged. It is read from the spawn's own environment, so
exporting it before launching Orca disables resolution process-wide.
Use it to confirm a suspected mis-resolution: run the failing operation with and
without it. Identical behaviour means resolution is not the cause. If it is,
report the shim's body — the parser is only allowed to recognise shapes we have
seen for real.
## Behaviour that changes, deliberately
A resolved shim is not merely a quieter spelling of the cmd.exe path. Two limits
of `cmd.exe` disappear with it:
- An argument containing `\r`/`\n` was rejected outright, because cmd ends its
command at a raw line break whatever the quote state. Multi-line agent prompts
now work.
- A command line over 8191 characters returned `The command line is too long.`
Long prompts now work.
Both are improvements, but they are behaviour changes: an unresolved shim still
hits both limits, so a caller must not assume every `.cmd` accepts them.
@@ -1,132 +0,0 @@
# Windows daemon-host relocation
On Windows the terminal daemon does not run from the install directory. Before it forks the
daemon, Orca materializes a trimmed copy of its own runtime under
`%LOCALAPPDATA%\Orca\daemon-host\<app version>\` and forks the daemon from there
(`src/main/daemon/daemon-host-relocation.ts`). This is what keeps live terminals alive across an
auto-update and across a crash of the main process.
Read this before changing the copy plan, the host exe name, the LOCALAPPDATA layout, or
`config/nsis/orca-installer-hooks.nsh`.
## What the relocation actually escapes
The killer is **electron-builder's process sweep** — not file deletion.
Windows will not delete a running image, so `RMDir /r "$INSTDIR"` cannot end the daemon on its own.
In app-builder-lib's `allowOnlyOneInstallerInstance.nsh`, `FIND_PROCESS` / `KILL_PROCESS` have two
branches:
| Branch | Condition | Selector |
| -------- | --------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Primary | `powershell.exe` runs, `Get-CimInstance` resolves, and `Get-ExecutionPolicy -Scope Process` is not `Restricted` | `Win32_Process` where `$_.Path.StartsWith('$INSTDIR', 'CurrentCultureIgnoreCase')` — **path-scoped** |
| Fallback | otherwise | per-user: `taskkill /F /IM "<AppName>.exe" /FI "PID ne $pid" /FI "USERNAME eq %USERNAME%"`; per-machine: the same without the username filter — **image-name-scoped** |
The upstream policy check can select the image-name fallback even when the inline process query
works. `Restricted` disallows script files, not inline commands. A packaged trace for #22872
showed that fallback successfully killing the relocated `Orca.exe` during an update; the
genuine-uninstall cleanup guard was not responsible.
Orca's `customCheckAppRunning` in `config/nsis/orca-process-check.nsh` tests the actual inline
`Get-CimInstance Win32_Process` query with terminating errors. Success selects the path-scoped
branch; any other result retains the upstream fallback. The hook reuses upstream process
selection, retry, permission and installation-mode handling. It neither overrides execution
policy nor assumes PowerShell is available.
A daemon outside `$INSTDIR` survives the path-scoped sweep. A genuinely unavailable process query
still permits an image-name sweep and cold restore. Updates also run the **old installed
uninstaller**, whose old capability check cannot be repaired by the new installer: the first
upgrade can still cold-restore on an affected host. Once both binaries contain the hook, later
updates use the actual capability test on both sides.
## Why the exe is copied verbatim (and not renamed)
The host exe keeps the app exe's own file name (`daemonHostExeName()` returns
`basename(process.execPath)`), so the relocated image is a byte-for-byte copy of the app binary
under its original name.
An earlier revision copied it as `orca-terminal-daemon.exe` specifically so the fallback
`taskkill /IM Orca.exe` could not match. That bought survival on the rare no-PowerShell host and
cost a textbook defence-evasion signature: _a process copies its own image into a user-writable
directory under a different name so a kill-by-image-name cannot match it, then runs detached and
survives the installer._ Microsoft Defender for Endpoint flagged it as MITRE **T1036
(Masquerading)**, and — because it is the process every other flagged action is attributed to — it
acted as a reputation multiplier on unrelated findings. No VS Code fork does this.
Trading the fallback branch for the name is the right trade:
- On the primary branch nothing changes: the daemon still survives the update.
- On the fallback branch the daemon is killed with the app and terminals **cold-restore** on
relaunch. That is the documented pre-relocation behaviour, a first-class outcome the update
harness already asserts (`--expect cold-restore`), not a failure.
- Relocation is fail-open end to end anyway: any materialization failure returns `null` and the
caller forks the install-dir host.
One new failure mode comes with it, on the fallback branch only. The daemon now matches
`FIND_PROCESS` under the app's image name, so it enters electron-builder's retry loop
(`allowOnlyOneInstallerInstance.nsh:136-141`). If the `taskkill` there fails to end it — an elevated
or otherwise unkillable host — the loop reaches `MessageBox ... /SD IDCANCEL` and `Quit`s, aborting a
silent update rather than completing it. Under the old distinct name the daemon was invisible to
that loop. Low probability (fallback branch _and_ an unkillable daemon), but it is a real new path.
What this does **not** buy. Two things bound the win honestly:
- The strongest T1036 indicator is a PE-resource-vs-disk-name mismatch, and it was **never firing**:
the shipped binary's `OriginalFilename` is empty (only `InternalName = Orca` is set), so there was
no embedded name for the old disk name to contradict.
- The remaining behaviour — a signed app copying its own ~225 MB image into user-writable
`%LOCALAPPDATA%` and running it detached under `ELECTRON_RUN_AS_NODE=1` — is still execution from
a non-standard user-writable location, which maps to **T1036.005** and is a standard heuristic on
its own.
So this removes a real but partial signal. Expect the score to drop; do not expect the process to
stop being scored.
## Options that were rejected
| Option | Why not |
| ----------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Materialize the tree from the NSIS installer | The daemon host is ~246 MB. Writing it at install time doubles install footprint and lengthens the window in which the app is down during a silent update. Worse, on a per-machine install (`INSTALL_MODE_PER_ALL_USERS`) the installer runs as the installing admin, so `$LOCALAPPDATA` is the wrong user's — every other user still needs the runtime path, which means the runtime self-copy stays in the product and the signal is only made rarer. |
| Ship a second signed `orca-terminal-daemon.exe` in the installer | `Orca.exe` is 235,555,328 bytes (224.6 MiB). electron-builder's NSIS uses solid LZMA with a 64 MB dictionary, so a second copy 224 MB downstream does not dedupe; the compressed installer grows by roughly a whole compressed Electron binary, paid by every user on every update download. It also does not remove the runtime copy — the helper still has to reach `%LOCALAPPDATA%` to escape the sweep — so it buys the same signal reduction as the verbatim copy at a large download cost. |
| Override `customCheckAppRunning` to force a path-scoped kill on both branches | Cheap to write (~6 lines: `!include "getProcessInfo.nsh"`, `Var pid`, and a macro that pins `IsPowerShellAvailable`, reusing upstream's dialog, retry loop and elevated handling) — but wrong at any size. Forcing the PowerShell branch on a host where PowerShell is genuinely absent makes `FIND_PROCESS` and `KILL_PROCESS` silently no-op, so the installer proceeds with the **real app** still running and its files in use. That is a worse outcome than the cold restore it would prevent, so this is not worth doing ever, not merely not now. |
| Hardlink instead of copy | Avoids the 246 MB entirely and is not a "copy" at all, but is NTFS-and-same-volume-only and introduces fresh failure modes (link counts, AV interception, cross-volume installs). Worth revisiting deliberately, not as part of a signal fix. |
## Invariants to preserve
- The host exe name is **derived from `process.execPath`**, never a literal. A future
`executableName` or dev-channel rename must follow automatically; pinning a name of our own is
how the mismatch creeps back.
- The daemon is identified by **PID and command line**, never by image name — in the product
(`daemon-pid-file-parse`, `daemon-process-inspection`) and in the harness
(`tests/tools/win-update-e2e/daemon-processes.mjs`). Nothing may start matching on the exe name.
- `config/nsis/orca-installer-hooks.nsh` kills the daemon by image name. That now also matches the
app's own exe, which is correct on a genuine uninstall — the product is being removed — but its
`${isUpdated}` guard must stay: electron-builder runs the uninstaller during every update's
`uninstallOldVersion`, and killing the daemon there defeats the whole feature. The legacy
`orca-terminal-daemon.exe` name stays in the macro to reap hosts left by older builds.
- `LOCAL_HOST_ROOT_NAME` in `daemon-host-relocation.ts` and the path in the uninstall macro are the
same directory. Change both together.
- Every native module the daemon bundle `require()`s must be in the copy plan
(`daemon-host-manifest.ts`). A missing one does not fail the fork: bare `require` walks up from the
relocated bundle, finds nothing, and the caller's fallback runs forever. That is how
`@vscode/windows-process-tree` went missing and every process-table read became a
`powershell.exe` CIM scan (#16905). A host missing the addon's runtime files counts as
unmaterialized, so hosts built before it was copied get rebuilt.
## Verifying a change
Unit coverage lives in `src/main/daemon/daemon-host-relocation.test.ts` (copy plan, native-addon mirror, verbatim
naming, marker/atomic publish, fail-open, prune veto). Nothing in unit tests can prove survival, so
any change to this file or to the NSIS macro needs the packaged harnesses:
- `.github/workflows/win-update-survival-e2e.yml` — builds an installer from the branch and updates
it over itself with `--expect survival`. Verify normal and child-scoped `Restricted` policy;
also verify released-to-branch migration with the expected cold-restore limitation.
- `.github/workflows/win-crash-survival-e2e.yml` — proves the daemon survives a main-process crash.
- `.github/workflows/windows-terminal-restart-e2e.yml` — terminal restart behaviour.
- `.github/workflows/win-update-e2e.yml` — release-tag-to-release-tag update, both `survival` and
`cold-restore` profiles.
All four are `workflow_dispatch`-only (the two update workflows also carry a push trigger pinned to
one historical feature branch), so they must be dispatched by hand against this branch before
merging a change here — which requires the workflow files to already exist on `main`.
-562
View File
@@ -1,562 +0,0 @@
# Windows EDR signal surface
Orca's Windows process tree is shaped like the thing behavioural EDR is built to
find. An enterprise Windows 11 / Intune tenant opened **six Microsoft Defender
for Endpoint incidents against Orca 1.4.192 in eight days**. All six fired as
active incidents and stayed open; three closed only because a human classified
them by hand in the portal. Defender never downgraded or closed one on its own.
None were signature hits. Every one was behavioural process-tree scoring, and
two escalated to multi-stage incidents carrying ATT&CK tactic mappings
(Execution, Collection).
The framing this document keeps throughout, because both halves matter:
> **Defender is not malfunctioning. It is describing the code accurately.** Orca
> really does copy its own signed image under a different name, really did read
> every process's memory on a timer, really does run base64-encoded PowerShell
> with the execution policy bypassed, and really does take screenshots and
> synthesise input from a runtime-compiled assembly. Each of those is a
> deliberate engineering choice with issue history behind it. The problem is not
> that the capabilities are illegitimate — it is that their **behavioural
> signature overlaps with attack techniques**, and an EDR scoring behaviour
> cannot see the difference.
Do not read this as a bug report against Defender, and do not read it as a claim
that Orca is malware. It is a map of which of our behaviours are legible to an
EDR as attack-technique-shaped, why each one exists, and what engineers and
administrators can do about it.
## What the tenant actually saw
Four independent evidence clusters, from six incidents:
| Cluster | Incidents | Evidence |
| ----------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------- |
| **Update** | A, B, C | `orca-windows-setup.exe` → `old-uninstaller.exe`, `Uninstall Orca.exe` (electron-builder generates these; they are in no repo file) |
| **Spawn** | all six | `Orca.exe` → `orca-terminal-daemon.exe` → `powershell.exe` / `pwsh.exe` / `cmd.exe` / `reg.exe` → `claude.exe`, `gh.exe`, `codex.cmd` |
| **Process table** | D | "suspicious memory activity" — `OpenProcess` plus a PEB read against every process on a repeating cadence |
| **Computer use** | E, F | `runtime.ps1`, `computer-sidecar.js`, many `operation.json`, a burst of ~10 short-lived `powershell.exe` |
Incident E is the one to look at hardest: 5 alerts, 37 evidence items, ATT&CK
**Execution + Collection**, and a description reading _"Screenshots were taken
unexpectedly on this device… Screen capture code was found in a script launched
by powershell.exe."_ Incident F added _"suspicious MSIL code"_, from the
`Add-Type -TypeDefinition` that recompiles inline C# P/Invoke on every
operation.
In the update cluster the uninstaller is genuinely `NotSigned`, while `Orca.exe`
and `orca-terminal-daemon.exe` report `Valid CN=SignPath Foundation`.
## The behaviours, and why each one exists
### The daemon runs from a copy of our own image
`src/main/daemon/daemon-host-relocation.ts` copies the Electron runtime into
`%LOCALAPPDATA%\Orca\daemon-host\<version>\` and forks the terminal daemon from
there.
It exists because the NSIS installer deletes the old install directory and force-
kills every process imaged under it. Without relocation, an auto-update kills the
terminal daemon and every live terminal with it. The copy is a run-as-node
`Orca.exe` rather than `node.exe` so there is no console flash and asar still
resolves; `config/nsis/orca-installer-hooks.nsh` reaps it on a real uninstall
(guarded by `${isUpdated}` so an update's `uninstallOldVersion` never fires it).
**At the time of these incidents the copy was also renamed** to
`orca-terminal-daemon.exe`, the image name every incident here reports, and
`DAEMON_HOST_EXE_NAME`'s comment stated the reason without varnish: _"so the NSIS
updater's `taskkill /IM Orca.exe` can't match it."_ The rename has since been
removed; the copy now keeps the app exe's own file name, because the updater's
kill sweep is path-scoped on every host that has PowerShell and the rename only
ever bought the no-PowerShell fallback. See
[`windows-daemon-host-relocation.md`](./windows-daemon-host-relocation.md).
**How an EDR reads it: MITRE T1036, masquerading** — and, for what remains,
**T1036.005**. A signed executable copied out of the install directory into
`%LOCALAPPDATA%` under a different name, which then spawns shells, matches the
textbook description closely enough that no behavioural engine can be expected to
score it low. Dropping the rename removes that literal indicator but not the
underlying shape: execution from a non-standard user-writable location is scored
on its own. Note also that the strongest form of the T1036 signal was never
present here — the shipped binary's `OriginalFilename` is empty, so there was no
embedded name for the old disk name to contradict.
### Every process gets a handle, on a timer
`src/main/windows/windows-process-table.ts` takes a Toolhelp32 snapshot under one
of **two** flag sets: identity (`None | CreationTime`) for callers that read only
pid, ppid and name, and detailed (`+ CommandLine`) for callers that match on a
command line. pid, ppid and name come out of the snapshot itself and open
nothing, so an identity scan opens nothing at all. `CommandLine` is what opens a
handle: the addon calls `GetProcessCommandLine` per process, which opens
`PROCESS_QUERY_LIMITED_INFORMATION` — the same right Task Manager takes — and
asks the kernel for the string. Upstream it opened
`PROCESS_QUERY_INFORMATION | PROCESS_VM_READ` and walked the PEB with three
`ReadProcessMemory` calls (`src/process_commandline.cc:32,41-47` in the vendored
`@vscode/windows-process-tree` 0.8.0 source that `config/patches/` patches).
`Memory` is retired as of this change, and that is a real reduction: it made
`GetProcessMemoryUsage` open a **second** `PROCESS_QUERY_INFORMATION |
PROCESS_VM_READ` handle per process for a `GetProcessMemoryInfo` call whose
result no caller read (`src/process.cc:47-63`). Dropping it halves the handles
opened per snapshot. On its own it removed no memory read — both handles carried
`PROCESS_VM_READ` at the time — so it composes with the patch below rather than
substituting for it.
It exists because seven independent readers used to fork `powershell.exe` for a
`Get-CimInstance Win32_Process` scan. That cost, measured: a PowerShell
Transcription policy recorded **~289 GB across 1.4 million files** because a scan
ran every ~2 seconds (#15209); a Group Policy or AV block turned a query into
"unavailable", which callers read as "no evidence", which is how a PTY tree
survived its own teardown (#9045, #10475); and the scan cost ~700 ms per pane, so
panes multiplied it (#15036). The native snapshot answers the same question in
15.9 ms against 706 ms for CIM — p50, measured on Windows 11 at 1050 processes.
See
[`windows-process-enumeration.md`](./windows-process-enumeration.md).
Asking for fewer fields is cheaper, and each caller now asks for the smallest set
that answers it. There are exactly **two** TTL-cached snapshots, one per flag
set, never one per caller: the fan-out the cache exists to remove is one scan per
_caller_, and each reader still serves every caller wanting its flag set, so a
32-wide teardown still collapses into one scan of each. Teardown identity and the
owner probe select the identity set and therefore open no handles; the per-pane
foreground tracker genuinely needs a command line and still pays for one. A third
cache would need a third flag set, not a third caller. Measured at 492 processes,
p50: identity 6.3 ms, detailed 12.3 ms — see
[`windows-process-enumeration.md`](./windows-process-enumeration.md).
**How an EDR read it:** a cross-process handle plus a remote memory read against
every process on the box, repeating on a cadence, is the read half of the
telemetry that credential dumping and process injection produce. MDE surfaced it
as "suspicious memory activity".
**The memory read is gone.** A fourth hunk in
`config/patches/@vscode__windows-process-tree@0.8.0.patch` has
`GetProcessCommandLine` call `NtQueryInformationProcess` with
`ProcessCommandLineInformation` (class 60, Windows 8.1+; Electron's floor is
Windows 10), which returns a `UNICODE_STRING` the kernel builds and needs only
`PROCESS_QUERY_LIMITED_INFORMATION`. Measured on ~540 processes, per detailed
scan: `ReadProcessMemory` 1128 → **0**, desired access `0x0410` → `0x1000`, with
byte-identical command lines on every process both readers recovered. There is no
PEB fallback to reinstate it — a hooked `ntdll` answering
`STATUS_INVALID_INFO_CLASS` for one target would have flipped a process-wide,
one-way switch back to `PROCESS_VM_READ` on exactly the machines this exists for.
Because the property is the _absence_ of an import, it is checkable on the
artifact rather than the source: `inspectWindowsProcessTreeAddon()` answers
`clean` / `unpatched` / `missing`, and the rebuild, `ensure-native-runtime.mjs`,
the relay build and `loadWindowsProcessTree()` all key on it. That check is load-
bearing because the published tarball ships a _loadable_ prebuilt built from
unpatched source, so "it required cleanly" is not evidence.
What to declare to administrators is now one
`PROCESS_QUERY_LIMITED_INFORMATION` handle per process on a detailed snapshot and
no remote memory access at all; an identity snapshot opens nothing. What this
does not narrow is _which_ processes are asked — a detailed scan still queries
every pid, including `lsass.exe`. Restricting the command-line pass to Orca's own
subtree needs job-object membership as its source of truth (a ppid-derived
allowlist would miss the detached, reparented descendants of #9045 and #10475),
and remains unclaimed work.
### Encoded, policy-bypassing PowerShell
Three sites are named in the incident analysis:
- `src/relay/windows-port-scan.ts` ran
`-NoProfile -NonInteractive -ExecutionPolicy Bypass -EncodedCommand` over a
`Get-NetTCPConnection -State Listen` script to find dev-server ports.
Enumerating listening ports is **MITRE T1049**, network service discovery, and
doing it through an encoded policy-bypassed shell is the aggravating factor
rather than the finding itself. The ordinary scan now starts no PowerShell at
all — `netstat.exe -ano`, with the owning process name projected off the shared
native table — and that payload survives only as the last-resort fallback, as
`-Command` with no policy override.
- `src/main/daemon/shell-ready.ts` uses `-EncodedCommand` for the OSC 133
bootstrap.
- `src/main/agent-hooks/windows-powershell-hook-launcher.ts` wraps managed hooks.
**No site spells the pair any more.** `src/main/ssh/ssh-remote-powershell.ts`,
`src/shared/setup-agent-sequencing.ts`,
`src/shared/windows-cmd-runner-delayed-launch.ts` and
`src/shared/windows-interactive-login-spawn.ts` each dropped
`-ExecutionPolicy Bypass` as a measured no-op: the policy gates script _files_,
never `-EncodedCommand`. Where the bypass was load-bearing it moved in-payload as
a process-scope `Set-ExecutionPolicy` (`setup-agent-sequencing.ts`), which is the
pattern to copy rather than restoring the switch — the switch loses to a GPO
scope anyway, so it never covered the locked-down case.
What remains is `-EncodedCommand` without the bypass: the PTY bootstraps
(`src/main/daemon/shell-ready.ts`, `src/main/providers/local-pty-shell-ready.ts`,
`src/main/providers/windows-shell-args.ts`), the hook wrappers
(`src/main/agent-hooks/windows-powershell-hook-launcher.ts` and its callers
`src/main/agent-hooks/runtime-home-hook-command.ts`,
`src/main/agent-hooks/installer-utils.ts`, and `src/main/claude/hook-settings.ts`
— that last one only as a _fallback_ since #18875, see below),
`src/main/runtime/windows-default-route-interfaces.ts`,
`src/main/runtime/orchestration/setup-completion-signal.ts`,
`src/shared/hermes-startup-query.ts`, and the four ex-bypass sites above.
`src/main/runtime/windows-mobile-firewall.ts` encodes a script and launches it
_elevated_ through `Start-Process -Verb RunAs`, which is a stronger shape than
any of those; only that hop is encoded, because `-ArgumentList` re-splits an
unquoted parameter string on whitespace.
One site still spells `-ExecutionPolicy Bypass` with **no** encoding, the weaker
signal: `src/main/cli/wsl-cli-scripts.ts` (`-File`, and it is a real script
file, so the switch is not a no-op there). Every Orca-managed WSL terminal now
reaches it by default on each CLI call (`docs/reference/wsl-managed-cli.md`), not
only after the user registers the WSL CLI. `src/main/system-fonts.ts` dropped it
for plain `-Command`; `src/shared/secure-path-windows-acl.ts` no longer runs
PowerShell at all, having moved to `icacls.exe`; and computer use now asks for
`-ExecutionPolicy RemoteSigned` in
`src/main/computer/windows-powershell-execution-policy.ts`, falling back to
`Bypass` only after a policy-blocked start.
Regenerate with `rg -- '-EncodedCommand|-ExecutionPolicy' src/` rather than
trusting the lists above, and note that a raw grep under-reports: the hook sites
reach `-EncodedCommand` through `wrapWindowsPowerShellEncodedCommand` and never
spell the flag themselves.
Encoding is not gratuitous: it shields paths and switches from `cmd.exe` and MSYS
rewriting (#6078, #14815), which is a real class of corruption. But
`-EncodedCommand` is a first-class Defender alert title ("Suspicious PowerShell
command line"), and base64 raises the score rather than lowering it, because it
denies the analyser the payload it would otherwise clear.
The hook launcher is prior art worth knowing about. #16003 measured, on a
reporting Kaspersky host, that `-WindowStyle Hidden` paired with
`-EncodedCommand` was denied at `CreateProcess` with exit 126 regardless of
payload — `exit 0` was denied too. The fix was to stop _spelling_ the flags:
`WINDOWS_POWERSHELL_HOOK_SWITCHES` is now just `-NoProfile`, and separately, in
#16576, the execution policy bypass moved in-payload as a process-scope
`Set-ExecutionPolicy` — a real command-line signal reduction, though #16003's
measured denial keyed on `-WindowStyle Hidden` + `-EncodedCommand`, not on the
bypass. It is also honest that the underlying behaviour did not change.
Copy the pattern, but copy its caveat too. `windows-powershell-hook-launcher.ts`
records that dropping `-WindowStyle Hidden` was a real tradeoff whose suppression
"was never measured" and "remains unverified on a real box". Reducing spelled
flags is the right instinct; treat any specific claim about what a removed flag
was doing as unproven until someone measures it.
### `cmd.exe /c` carrying caret-escaped free text
`buildWindowsCmdShimCommandLine` in
`src/shared/child-process/windows-command-line.ts` builds `/d /v:off /s /c "…"`
for the `.cmd` and `.bat` targets Windows can only start through `cmd.exe`
(`codex.cmd` being the one that matters). Because cmd expands `%VAR%` even inside
a quoted token, each `%` is broken with `"^%"`.
The escaping is not decorative. Measured on Windows 11 against a real `.cmd`
shim, `["a b", 'c"d', "e%F%g", "h&i", "j^k"]` came back as `["a b", 'c"d',
"e^%F^%g", "h"]` — the `&` truncated the argument _and_ ran the remainder as a
command.
**How an EDR reads it:** caret escaping is the canonical obfuscation marker in
`cmd.exe` command lines, and the free text being escaped here is an agent prompt,
so the line is long, high-entropy, and attacker-shaped. It is the exact input an
obfuscated-command-line detector is tuned on.
### The spawn tree itself
`Orca.exe` → the relocated daemon host (`orca-terminal-daemon.exe` in the builds
these incidents cover, `Orca.exe` since) → a shell → an agent CLI is what a
terminal multiplexer for coding agents _is_. `reg.exe` appears from
`src/main/win32-utils.ts`,
`src/main/agent-hooks/managed-hook-owner-identity.ts` and
`src/relay/pty-shell-utils.ts` (reading the OpenSSH `DefaultShell`).
Nothing here is avoidable in principle. What is controllable is depth and
breadth: every interpreter hop between Orca and the thing the user asked for adds
a scored edge, which is why the shipped doctrine of #15520 and #15595 is to
_shorten the interpreter chain_ rather than to hide a window.
#18875 is a worked example of that doctrine. The Claude Code lifecycle hook was
registered as `powershell.exe -NoProfile -EncodedCommand <...>` whose entire
decoded payload was a `Test-Path` and a call to `~/.orca/agent-hooks/claude-hook.cmd`.
It now registers the bare script path itself, with no shell operators, so `bash ->
powershell -> cmd -> curl` became `bash -> cmd -> curl` and one
`powershell.exe -EncodedCommand` per hook event — a first-class Defender alert
title — leaves the tree. The reporting box fired ~6 900 of them in five days,
70% from Claude sessions that were not running under Orca at all and whose hook
exits at its first `ORCA_PANE_KEY` guard.
What is measured is latency and the hop count, nothing else: median 471 ms ->
213 ms per event idle, and 656 ms -> 296 ms (p95 696 ms -> 337 ms) under 10-way
concurrency, invoked as Claude Code invokes it. **No EDR verdict on either tree
was measured**, so claim the removed `-EncodedCommand` spelling and the shorter
chain, not a score. `cmd.exe` remains in the tree, spelled by MSYS's own `.cmd`
spawn rather than by us — the doc's one "unavoidable for `.cmd`/`.bat`" case,
carrying only an absolute path, with no caret escaping, no
encoding and no free text. The encoded launcher is still the shape for profile
paths the shells cannot carry bare (a space, `%`, `^`, `&`, non-ASCII, a UNC
profile).
The direct shape first shipped as `<path> || echo {}`, gated on Git Bash being
resolvable. That gate guessed at a choice Claude Code makes on its own: it runs
hooks under Windows PowerShell 5.1 when it picks PowerShell, and 5.1 rejects `||`
(parse error, exit 1), so every hook failed (STA-8913). The registered command is
now the bare path, which parses in Git Bash, cmd.exe, pwsh and PowerShell 5.1.
The neutral `{}` for a missing payload lives in the `claude-hook.cmd` entry, which
hands off to `claude-hook-impl.cmd` in the same `cmd.exe`; a deleted entry is an
ordinary non-blocking hook error (never exit 2) until the next install rewrites it.
Compat consumers such as cursor-agent and Devin import `~/.claude/settings.json`
and run `command` through their own launcher (the payload carries a
`DEVIN_PROJECT_DIR` skip for that); one that cannot start a `.cmd` at all still
gets a launch failure. Before widening the direct shape to another agent, measure
that consumer's host.
### Computer use: screen capture, synthetic input, runtime-compiled MSIL
`native/computer-use-windows/runtime.ps1` is a large PowerShell script.
`src/main/computer/desktop-script-provider-bridge.ts` launches it as
`powershell.exe -NoLogo -NoProfile -NonInteractive -ExecutionPolicy RemoteSigned
-File runtime.ps1 <operation.json>`, retrying once at `Bypass` only if the start
comes back policy-blocked — **once per operation**, with
`desktop-script-provider-client.ts` writing a fresh `operation.json` into a new
temp directory each time. On every launch the script runs `Add-Type
-TypeDefinition` over inline C# that P/Invokes `SendInput` and the window APIs,
then captures the screen through `Graphics.CopyFromScreen`.
That is four separate high-signal behaviours stacked in one process:
| Behaviour | How it is scored |
| --------------------------------------------- | ------------------------------------------------------------- |
| `Graphics.CopyFromScreen` | **MITRE T1113**, screen capture — Collection tactic |
| `SendInput` synthetic keyboard/mouse | input synthesis against other applications |
| `Add-Type -TypeDefinition` on every operation | MSIL compiled at runtime; incident F's "suspicious MSIL code" |
| One `powershell.exe` per operation | a burst of short-lived interpreters under one parent |
The bottom two rows are the two the incident text named directly, and they are
also the two a persistent runtime host would remove: a long-lived helper compiles
its P/Invoke stubs once and answers operations over a channel, so neither the
MSIL recompilation nor the interpreter burst repeats. A change doing that is in
flight and unmerged at the time of writing; check the code rather than this
paragraph for what the shipped build does. Screen capture and `SendInput` are
inherent to the feature and no refactor removes them.
### SSH hosts: upload-stage file identity and runtime-store GC
These run on the _remote_ Windows host over SSH, not on the desktop, but the
host's EDR scores them the same way.
The relay upload stage fences each slot with the directory's file ID (volume
serial plus file index). That used to come from `Add-Type -TypeDefinition` over
a P/Invoke of `GetFileInformationByHandle`, compiled in every stage command. When
the relay runs on Orca's pinned Node (design D5), node.exe is already hashed
against the pin and has run once, so the stage commands now ask it instead:
`src/main/ssh/ssh-relay-upload-stage-windows-commands.ts` runs
`node.exe -e <fixed script> -- <path>`, a fixed `fs.lstatSync(..., { bigint: true })`
with the path as an argument. libuv fills `dev` and `ino` from the same volume
serial and file index, so both readers write the same `vol:high:low` lowercase
hex, and identity files are compared after normalising hex spelling. An old
client can recover a stage a new one reserved, and the reverse.
Two alternatives were rejected:
- **PowerShell alone.** Neither .NET Framework (Windows PowerShell 5.1) nor .NET
exposes a file index without P/Invoke, which is what `Add-Type` compiles.
`fsutil file queryfileid` would spawn another binary per lookup and prints a
different format, which would break mixed-version recovery.
- **Host Node.** Relays still on the host's own Node (rung C and the legacy
path) keep the `Add-Type` helper, because Orca has not verified that binary.
That is the one remaining `Add-Type` site on SSH hosts; it goes when those
rungs do.
A lookup costs one short-lived node.exe per existing stage directory the command
inspects, usually one or two. It is not a loop over the whole pool.
Runtime-store GC (`src/main/ssh/remote-node-runtime-store-windows.ts`) reads the
store in one PowerShell invocation. It learns which runtimes are in use from a
single `Get-CimInstance Win32_Process` query, filtered on an image path under
`runtimes\`. It never matches on the image name, so another program's node.exe
holds nothing. WMI refuses a standard user's SSH logon, so a refusal falls back to
`Get-Process`, which reads the image path of the account's own processes — the
only ones running from its store. If both fail, no process check has run and the
pass keeps everything. Windows itself also refuses to delete a running image, which is a
second safeguard.
### SSH hosts: starting the relay outside the session
Win32-OpenSSH puts each session's shell in a job with
`JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE | JOB_OBJECT_LIMIT_BREAKAWAY_OK`
(`contrib/win32/win32compat/w32-doexec.c`, unchanged since 2018), so the relay
must leave that job to outlive the connection. It used to leave through WMI
`Win32_Process.Create`, which is both an EDR-scored remote-execution shape
(T1047) and refused to a standard user's network logon unless an administrator
grants Remote Enable on `root\cimv2`.
Now `relay.js --windows-breakaway-launch` runs once per launch on the same
node.exe and calls `spawnOutsideJob` in the staged process-tree addon
(`src/process_launch.cc` in the patch): one `CreateProcessW` with
`CREATE_BREAKAWAY_FROM_JOB`, and a handle list that passes only the relay's
three stdio handles, so no SSH channel pipe is inherited. libuv never passes that
flag, so Node alone cannot do this. WMI remains only as the fallback for a relay
built without the addon or a job that refuses breakaway, and a refusal there is
reported as `ORCA_RELAY_LAUNCH_REFUSED`. The Windows SSH-host lanes run with no
WMI grant and assert the breakaway route.
`orcad.js` exposes the same launcher (`src/shared/windows-breakaway-launcher.ts`)
with **no** WMI fallback: a host that cannot break away refuses the orcad launch.
Windows orcad operations start **no PowerShell**: sshd's DefaultShell runs the
pinned `node.exe` directly with plain path arguments, against one content-addressed
host script staged beside the slots (`src/main/ssh/orcad-windows-host-script.ts`).
Only a `node.exe` path that itself needs quoting (a profile name with a space)
falls back to one unencoded `powershell.exe -Command`. The launch waits for
readiness host-side, one exec per 20 s at most, rather than re-running an exec
every 500 ms.
## Signing is not the gate
The most useful calibration in the whole incident set came from the reporter's
own machine: **Antigravity IDE's main executable is `NotSigned` and was not
flagged, while Orca's is signed and was flagged six times.** Their conclusion:
_"signing is not the gate here — behaviour is."_
The mechanism is that Defender reputation is signer **plus prevalence**, and
prevalence is keyed on **file hash**. A widely installed unsigned binary clears
on install count alone. Orca's signature is a free OV certificate from SignPath
Foundation (`config/electron-builder.config.cjs` sets
`win.signtoolOptions.publisherName`; `config/scripts/verify-windows-inner-signature.mjs`
pins `CN=SignPath Foundation, O=SignPath Foundation, L=Lewes, S=Delaware, C=US`),
shared across many OSS projects, with no independent SmartScreen or MAPS
reputation of its own. Every release ships new hashes, so whatever prevalence a
build accumulates resets on the next update. Dev channels ship unsigned by
design, because SignPath's approval waits cannot fit a dev cadence
(`config/scripts/verify-dev-channel-packaging.mjs`).
Signing the uninstaller is worth doing — an unsigned `old-uninstaller.exe`
running under a signed installer is a gratuitous contribution to the update
cluster — but do not expect it to change the behavioural verdict. The three
non-update clusters contain no unsigned binary at all.
## What we do not know
Two limits the incident analysis recorded, kept here rather than smoothed over:
- **No data on Hermes.** Nothing in this document describes how Hermes behaves
under the same tenant policy — though `src/shared/hermes-startup-query.ts` does
spell `-EncodedCommand`, so the gap is telemetry, not surface.
- **Antigravity not being flagged is absence of evidence, not proof.** It is one
reporter's recollection from one machine, not a measurement. It is strong
enough to falsify "the problem is that we are not signed well enough"; it is
not strong enough to support a positive claim about how Defender scores that
product.
Add to those: this is one tenant with one policy configuration. Whether the same
build scores the same way elsewhere is unmeasured.
## Guidance for engineers
Fixes for several of the shapes above are in flight in separate changes; nothing
in this section should be read as a statement that a given site has already
changed. Check the code before relying on it.
The checklist. On Windows, do not reach for:
| Don't | Instead |
| --------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `-ExecutionPolicy Bypass` on the command line | Set the policy in-payload at process scope, as `windows-powershell-hook-launcher.ts` does, or do not run a `.ps1` at all |
| `-EncodedCommand` | A temp `.ps1` with an argument, or no PowerShell hop: prefer a native API or an existing Node path |
| `cmd.exe /c` carrying escaped free text | Spawn the real target directly. `cmd.exe` is only unavoidable for `.cmd`/`.bat`; keep free text out of the line where you can |
| Forking `powershell.exe` to read system state | The native reader — [`windows-process-enumeration.md`](./windows-process-enumeration.md) is the standing rule for the process table |
| A process per operation in a loop | One long-lived helper with a request channel. A burst of short-lived interpreters under one parent is itself the signal |
| `Add-Type -TypeDefinition` at runtime | A precompiled, signed assembly, or a native helper |
| Copying our own image under a different name | Copy it verbatim — [`windows-daemon-host-relocation.md`](./windows-daemon-host-relocation.md) (done for the daemon host) |
| Deriving a script runner from a UI preference | [`windows-setup-shell.md`](./windows-setup-shell.md) — the script declares its own interpreter |
Two framing rules that outlast the table:
- **Shorten the interpreter chain.** Each hop between Orca and the user's actual
target is a scored edge and a place for AV to deny a `CreateProcess`. This is
the shipped doctrine of #15520 and #15595.
- **Do not spell a flag you can avoid spelling.** #16003 measured a denial that
was independent of the payload and keyed purely on the switch combination on
the command line. What is on the line is itself the detection surface.
## Guidance for administrators deploying Orca
### Path exclusions alone will not silence these
This is the single most important operational point, and it is the one most
commonly got wrong. The six incidents are **MDE EDR behavioural alerts**.
Defender Antivirus path exclusions suppress _scan_ detections; they do not
suppress EDR behavioural alerts the same way. Adding
`%LOCALAPPDATA%\Programs\orca\` to the AV exclusion list and expecting the
incidents to stop will not work.
### What actually stops incidents being created
An **MDE alert suppression rule** scoped to the process tree. Build it in
Microsoft 365 Defender (Settings → Endpoints → Alert suppression), conditioned
on:
- **Alert titles** — `A suspicious file was observed` and
`Suspicious PowerShell command line`, plus any further titles your tenant
actually produced. Take the titles from your own incidents rather than from
this list.
- **File paths** — `Orca.exe` and `orca-terminal-daemon.exe` under
`%LOCALAPPDATA%\Programs\orca\` and `%LOCALAPPDATA%\Orca\daemon-host\`.
Scope it as narrowly as your tenant will tolerate, and review it when Orca
updates: the `daemon-host` path carries a `<version>` segment, so a rule pinned
to one version will silently stop matching. Two traps in that path in particular.
Materialization stages into a `<version>.staging-<hex>` sibling before renaming
it into place, so an exact-version rule misses the tree **mid-update** — which is
precisely when the update-cluster incidents fire. And the root falls back to the
Electron `userData` path when `LOCALAPPDATA` is unset, so
`%LOCALAPPDATA%\Orca\daemon-host\` is the normal location rather than a
guaranteed one. Prefer a prefix match on `…\Orca\daemon-host\` over a rule
pinned to one full path.
Add AV path exclusions for those two directories as well — they cut scan cost on
a tree that is rewritten on every update — but understand the division of
labour. The exclusions reduce scanning; **the suppression rule is what stops
incidents being created.**
### Check your ASR rules
Check whether the tenant has the Attack Surface Reduction rule **"Block
executable files from running unless they meet a prevalence, age, or trusted list
criterion"** enabled. If it is, that alone explains a freshly signed Orca build
being hit immediately after every update: each release ships new hashes, so every
build starts at zero prevalence and zero age no matter how it is signed. Either
allowlist the Orca install paths for that rule or expect a hit on each update.
### Expect the alerts to recur after each update
Prevalence is keyed on file hash. An update replaces the hashes, the reputation
starts over, and a suppression rule is the only thing carrying across.
## Computer use: decide before you deploy
Read this section before enabling computer use on a monitored endpoint, not
after.
> **On a monitored endpoint, an alert reading "Screenshots were taken
> unexpectedly on this device" is not the kind of finding a SOC dismisses on
> sight.**
Incident E is the shape to expect: 5 alerts, 37 evidence items, a multi-stage
incident mapped to ATT&CK **Execution + Collection**, and a description naming
screen capture found in a script launched by `powershell.exe`. Incident F adds
runtime-compiled MSIL to the same tree.
Every part of that is an accurate description of what the feature does. Orca's
computer use takes screenshots, synthesises keyboard and mouse input into other
applications, and compiles the P/Invoke stubs it needs at runtime. An
organisation that monitors for Collection-tactic activity — and any organisation
running MDE with default incident creation does — will see it, and will see it
as Collection.
So decide deliberately, in advance:
- **Allowlist it**, with a suppression rule covering the computer-use tree
(`powershell.exe` with `-File …\runtime.ps1`) as well as the base Orca paths,
and tell your SOC what it is before the first incident rather than during it.
- **Or leave it disabled** on monitored endpoints.
What does not work is deploying it un-triaged and handling the incidents
reactively. By the time a Collection-tactic incident is open, an analyst is
already reading a description of screenshots being taken without the user's
knowledge, and the burden of proof has moved to you.
@@ -1,165 +0,0 @@
# Why an MSYS pane's children escape the per-PTY job
Every child started from a Git Bash / MSYS2 / Cygwin pane leaves the pane's job
object unless the job is created **without** `JOB_OBJECT_LIMIT_BREAKAWAY_OK`.
`terminatePtyJob` then reports `terminated` and leaves the child running — the
orphan that holds a worktree directory open.
The denial is already in `config/patches/node-pty@1.1.0.patch`
(`usesCygwinRuntime`, added in #19068). This page records the measurement
behind it, because the failure mode it prevents is indistinguishable from a
stale native addon and the gates of the day could not tell the two apart.
## The mechanism
The MSYS/Cygwin runtime asks for `CREATE_BREAKAWAY_FROM_JOB` on the
`CreateProcessW` inside its `spawn`/`exec` path. A job that carries
`JOB_OBJECT_LIMIT_BREAKAWAY_OK` grants it, so the child is created outside the
job; a job without that limit denies it with `ERROR_ACCESS_DENIED`, and the
runtime retries without the flag rather than failing the spawn. `fork` is not
affected — forked Cygwin processes stay in the job either way.
Measured on Windows 11 `10.0.26200.9168`, Git `2.55.0.windows.3`,
bash `5.3.15(1)-release`, node `v24.18.0`, `useConptyDll: true`, for
`node-pty.spawn('C:\Program Files\Git\bin\bash.exe', ['--noprofile','--norc','-i'])`
— `+J` / `-J` is membership of the per-PTY job, read with
`QueryInformationJobObject(JobObjectBasicProcessIdList)`:
```
bin\bash.exe +J ConPTY shell (assigned by node-pty)
└ ..\usr\bin\bash.exe +J launcher hand-off, plain CreateProcess
└ usr\bin\bash.exe +J Cygwin fork for the typed command
└ node.exe -J Cygwin exec -- ESCAPES HERE
```
`bin\bash.exe` is a 47 KB launcher, not an MSYS binary: `C:\Program Files\Git\bin`
holds only `bash.exe`, `git.exe` and `sh.exe`, with no `msys-2.0.dll`. Its
hand-off to `bin\..\usr\bin\bash.exe` is an ordinary `CreateProcess` and keeps
job membership. Only the MSYS runtime's own spawn breaks away.
The shell-replacement shape (`bash -c 'exec "$BASH" --noprofile --norc -i'`)
loses membership one step earlier, at the `exec`, and everything below inherits
the loss:
```
bin\bash.exe +J
└ ..\usr\bin\bash.exe +J
└ usr\bin\bash.exe -J Cygwin exec -- ESCAPES HERE
└ usr\bin\bash -J
└ node.exe -J
```
Both shapes leak. The `exec` is not the cause; it only moves the escape earlier.
## The A/B that pins it
One source tree, one toolchain, one variable — `usesCygwinRuntime` forced to
`false` so the per-PTY job keeps `JOB_OBJECT_LIMIT_BREAKAWAY_OK`:
| per-PTY job limit | `listPtyJobProcessIds` | child reaped by `terminatePtyJob` | runs |
| ---------------------- | ---------------------- | --------------------------------- | ---- |
| `BREAKAWAY_OK` set | 2 pids, child absent | no | 0/2 |
| `BREAKAWAY_OK` cleared | 5 pids, child present | yes | 4/4 |
The job **is** the right boundary. With breakaway denied it holds the whole MSYS
tree, including the child that detached from the console, and one
`terminateJob` reaps all of it. No alternative tracking mechanism is needed.
Denying breakaway did not break ordinary launches from the pane: `git`,
`cmd //c`, an absolute-path `node`, a `&`-backgrounded job with `disown`, and
`where.exe` all returned 0 with no `Access is denied`, identically to the
breakaway-allowed control. Untested: a **non-Cygwin** program that itself passes
`CREATE_BREAKAWAY_FROM_JOB` (installers, updaters) and therefore has no runtime
to retry for it. That needs a helper that calls `CreateProcess` with the flag;
`start /b` does not exercise it (it uses `CREATE_NEW_CONSOLE`).
## A stale addon looks exactly like the bug
`config/scripts/node-pty-job-ownership.cjs` used to assert only that
`terminateJob`, `listJobProcessIds` and `assignCurrentProcessToJob` are
exported. All three predate #19068, so a `conpty.node` built before it passed
every gate: `isPtyJobOwnershipAvailable()` returned true and
`windows-pty-job.win32.test.ts` passed 6/6, while
`windows-msys-job.win32.test.ts` failed with a two-pid job list that read as a
source defect rather than a build-freshness one.
When that test fails, check the binary before the code:
```js
// UTF-16LE, because usesCygwinRuntime holds the literals
readFileSync(conptyNodePath).includes(Buffer.from('msys-2.0.dll', 'utf16le'))
```
False means the addon predates the fix; rebuild node-pty from patched source.
Note that a git worktree sharing `node_modules` with its main checkout shares
that checkout's `build/Release/conpty.node`, so pinning the _source_ to a commit
does not pin the _addon_.
The gate asserts that marker, the way `stagedRelayAddonIsUnpatched()` in
`src/main/windows/windows-process-table.ts` already sniffs a patched addon by a
binary import name. Symbol presence cannot distinguish patch revisions; a marker
can.
Because the marker is a literal in `conpty.cc` and the gate's copy of it is a
separate constant, `ensure-native-runtime-job-ownership.test.mjs` asserts the
patch still adds `L"msys-2.0.dll"` to that file. Without that, editing the patch
would turn the gate into a permanent false positive that fails every correctly
rebuilt addon and tells the developer to do the one thing that cannot help.
## A stale source tree looks exactly like a stale addon
`rebuild-native-deps.mjs` rejects a marker-less addon in its Electron probe and
again after the rebuild, so an unpatched `build/Release/conpty.node` is never
left in place silently. But the rebuild compiles whatever `node_modules/node-pty`
holds, and pnpm materializes that from the patch only at install time. On a
checkout whose `node_modules` predates the denial, `--force` compiles for
minutes and rewrites `conpty.node` byte-identical and unpatched; measured on a
Windows dev checkout, same size, new mtime, marker still absent. The
post-rebuild gate then said "rebuild from source", which was the step that had
just run.
So the script reads `src/win/conpty.cc` before it compiles: if the source does
not carry `L"msys-2.0.dll"`, it stops before the rebuild and says to run
`pnpm install`, which re-applies the current patch. If the patch itself lacks
the literal, the checkout predates the denial and a reinstall cannot help.
## Every path the loader can fall through to
`loadNativeModule` tries `build/Release`, then `build/Debug`, then
`prebuilds/win32-<arch>`, each relative to node-pty's root and then to `lib/`,
swallowing every failure in between. A require of a wrong-architecture `.node`
is one of those failures, so the candidate that runs is the first one the target
arch can actually load. The published prebuild is always the last candidate and
never carries the patch:
| package | `build/Release` | prebuild pruned? | what the app loads |
| -------------------- | ----------------------------- | ---------------- | ------------------ |
| same host, same arch | patched | yes | `build/Release` |
| cross host | absent, cannot be cross-built | no | the prebuild |
| cross arch, built | patched, target arch | yes | `build/Release` |
| cross arch, failed | the host's arch | no | the prebuild |
`beforeBuild` runs `rebuild-native-deps.mjs --platform=win32 --arch=<target>`, so
a cross-arch slice normally does get a patched `build/Release` for the target —
row three is a correct package. `prunePackagedNodePty` asks the same question the
loader does, reading the PE machine of `build/Release` rather than comparing
`electronArch` to `process.arch`, so row three's leftover prebuild goes. Keying
off the host arch kept it: unreached in the normal case, but still the binary the
loader takes if `build/Release` ever fails to load for an unrelated reason — an
AV quarantine, a missing dependency — which is the silent fall-through this whole
gate exists to close. Rows two and four keep the prebuild because it is the only
thing there the target could load. Measured on Windows 11 x64 with the VS 2022
ARM64 cross toolset: `node-gyp rebuild --arch=arm64` does emit a `conpty.node`
with machine `0xaa64`, so row three is a real package shape — but as of this
writing no release produces it, because `electron-builder --win` is run without
an arch and packages x64 only.
The verifier still does not key on presence: the prune is the step it is
checking, and `build/Debug` is never pruned, so an unmarked file beside a
correct `build/Release` cannot by itself separate row three from row four, and
failing on one would reject a correct package with advice its builder could not
act on. `verifyPackagedConptyBreakawayMarker` instead resolves the
addon the way the loader does — first candidate whose PE `IMAGE_FILE_HEADER`
machine matches the target — and checks the marker on that one. A package with
no candidate at all, or none of the target's architecture, is refused: it has no
ConPTY backend to load.
@@ -1,672 +0,0 @@
# Reading the Windows process table
Orca needs three things from the Windows process table: who a PID's parent is
(descendant walks and teardown identity), what a process is running (agent
recognition), and how much memory/CPU it uses (Resource Manager).
Node cannot answer the first one without native code. That is why seven
independent readers existed, each forking `powershell.exe` to run
`Get-CimInstance Win32_Process`, with a `wmic` fallback that Windows 11 24H2 has
since removed.
## Use the native snapshot
`src/main/windows/windows-process-table.ts` is the only module that may read the
table. It wraps a Toolhelp32 snapshot from `@vscode/windows-process-tree`.
```ts
import {
readWindowsProcessIdentityTable,
readWindowsProcessIdentityTableFresh,
readWindowsProcessTable,
readWindowsProcessTableFresh
} from '../windows/windows-process-table'
```
Each pair is a shared TTL cache plus a `Fresh` variant that starts its scan
after the call. Use `Fresh` for teardown identity, where a cached row can
predate the exit it is being asked about, and the cached one for anything
periodic.
All four **reject** when the table cannot be read. Do not convert that into an
empty array. An empty table is a claim that nothing is running, and callers act
on that claim by declaring a tree dead or a shell childless. "Unavailable" has
to stay distinguishable from "empty" — collapsing the two is how a PTY tree
survived its own teardown (#9045).
## Two flag sets: ask for a command line only if you read one
Neither flag is a wider column on the same query. Each is a separate
per-process syscall sequence, and they are not equally expensive to the EDR
watching:
- `CommandLine` (`process_commandline.cc`) —
`OpenProcess(PROCESS_QUERY_LIMITED_INFORMATION)`, then
`NtQueryInformationProcess(ProcessCommandLineInformation)` twice: once to size
the buffer, once to fill it. The kernel builds the string, so no address space
is opened or read. It used to walk the target's PEB with three chained
`ReadProcessMemory` calls; the patched addon no longer contains that primitive.
- `Memory` (`process.cc`) — retired. It took a **second** `OpenProcess`, and that
one carried `PROCESS_VM_READ`, which it acquired and never used.
Measured here (541 processes, 405 openable), per detailed scan, before → after
dropping `Memory`: `OpenProcess` 1082 → 541. That halving is all the `Memory`
drop bought on its own — both handles carried `PROCESS_VM_READ` at the time, so
it moved the PEB traffic not at all. Replacing the PEB walk with the kernel
query is what took `PROCESS_VM_READ` and `ReadProcessMemory` out of the addon
altogether; the two changes compose, and neither substitutes for the other.
So be precise about what these two flag sets buy now. A detailed scan is one
`OpenProcess(PROCESS_QUERY_LIMITED_INFORMATION)` per process and no memory
access at all. What the split buys on top of that is the handle itself: an
identity scan opens nothing.
So the module exposes two snapshots, and the row types differ so a cheap caller
cannot read what its flag set did not pay for:
| reader | row type | flags | per-process handles |
| ------------------------------------------ | --------------------------- | ---------------------- | ------------------- |
| `readWindowsProcessIdentityTable[Fresh]()` | `WindowsProcessIdentityRow` | `None \| CreationTime` | none |
| `readWindowsProcessTable[Fresh]()` | `WindowsProcessRow` | `+ CommandLine` | one `OpenProcess` |
`Memory` is requested by neither. Nothing reads a working set off this table —
`windows-process-resource-collector.ts` runs its own sweep because it needs
commit and CPU counters in the same pass, and the addon stores `WorkingSetSize`
into a `DWORD` so anything above 4 GB wraps anyway.
Measured on Windows 11 with 492 processes (p50 / p95):
| | p50 | p95 |
| -------------------------------- | ------- | ------- |
| identity (pid + ppid + name) | 6.3 ms | 7.0 ms |
| detailed (+ command line) | 12.3 ms | 13.4 ms |
| _retired_ (+ memory) | 13.1 ms | 14.1 ms |
| `Get-CimInstance` via PowerShell | 706 ms | 723 ms |
There are exactly **two** caches, never one per caller. The fan-out this module
exists to prevent is one scan per _caller_, and each reader still serves every
caller wanting its flag set, so a 32-wide teardown still collapses into one scan
of each. A third cache would need a third flag set, not a third caller.
### Only one native read may be in flight, ever
This is the price of having two flag sets, and it is not optional.
The npm wrapper **coalesces rather than queues**. `getRawProcessList` pushes the
callback onto one list and calls the addon only when no request is in progress,
so a second concurrent caller's `flags` are **discarded** and it is handed the
first caller's rows. Measured against the real addon: issue identity first, both
callers get the same array, 0 of 541 rows carry a command line. A detailed read
that overlaps an identity read therefore returns a table with **every command
line empty**, and agent recognition reads that as "no agent" — silently, and
only under concurrency.
Nothing else in this module prevents that. Each snapshot cache single-flights
only within itself (`inFlight` is a closure per reader), and the wedge set
latches only _after_ a read misses its 3 s deadline, so through the healthy
~12 ms of a scan neither excludes the other. Overlap is the normal state rather
than an edge case: other panes keep polling detailed at 750 ms while a teardown
takes identity snapshots, and `codex-structured-turn-processes.ts` issues fresh
detailed scans on turn stop.
`nativeReadGate` serializes every native read across both flag sets. It is also
what makes the relay's bare addon safe: `adaptAddon` has no queue at all, and
two simultaneous `CreateToolhelp32Snapshot` calls are the crash the vendor's
queue exists to prevent. Every link settles — a wedged read still rejects on its
deadline — so a waiter is never stranded; it re-checks the wedge and rejects.
Because only one native call is ever outstanding, the wedge gate and the 3 s
deadline stay **shared** and retention stays bounded at exactly one callback,
not one per reader. Read ids are module-global and monotonic, so a late callback
can only clear its own wedge.
`resetNativeReaderState` **chains** onto the gate rather than replacing it. A
replacement would let a waiter still holding the old chain run beside a read
queued on the new one; every link settles within the deadline, so chaining costs
a bounded wait and keeps the exclusion whole. That path is test-only, which is
exactly why it matters — it would otherwise hand a suite two concurrent calls
into its own mock, the condition these tests exist to detect.
### Testing this module: assert a positive property, on the right mock
Three defects have now shipped in this file's tests, all the same shape — a case
that passed for a reason other than the one it claimed to check:
1. A loader that built a **fresh mock per call**, so the coalescing it was meant
to reproduce could never happen.
2. An identity-side assertion of only `!('command' in row)`, which a correctly
flagged read and a coalesced one satisfy equally, so the test would go green
on the very regression it guards.
3. A concurrency assertion placed on the **coalescing** mock, whose own
`requestInProgress` latch means it can never report more than one call in
flight — so it held whether or not this module excluded anything, and passed
against a read gate that had genuinely lost exclusion.
The third arrived in the fix for the first two, which is the point: this is not a
mistake you make once.
So: assert what each flag set **did** get, not only what it lacks, and put those
assertions in the helper both orderings run through, or the reverse order keeps
the blind spot. The identity set is checked on `creationTimeMs` because that is
the field it exists to carry. Keep both that check and the flags-array check —
they catch **different** failures and neither is redundant. The flags array
catches a read served another flag set's rows (the coalescing bug); the
positional `creationTimeMs` check catches field shaping — identity dropping
`CreationTime` from its flags, or `toIdentityRow` failing to forward it — which
no flags assertion would notice.
And pick the mock to match the claim. The coalescing mock models the npm
wrapper's queue semantics and is the only place to assert those. Concurrency has
to be measured against the bare-addon mock, which has no queue and so makes
re-entry observable.
With no native binding there is only one scan to run and it is the 1.4 s
PowerShell one, so the identity view rides the detailed snapshot — projected
through `toIdentityRow`, so an identity row carries no command line on any host.
### Which callers need which
| caller | reads | flag set |
| ------------------------------------------ | --------------------------------- | -------- |
| `windows-agent-foreground-process.ts` | `command` (agent recognition) | detailed |
| `local-workspace-platform-port-scanner.ts` | `command` (port attribution) | detailed |
| `codex-structured-turn-processes.ts` | `command` (turn-process identity) | detailed |
| `structured-tui-process-identity.ts` | `command` (child match) | detailed |
| `windows-pty-root-identity.ts` | `pid` / `ppid` only | identity |
| `agent-session-process-identity-probe.ts` | `creationTimeMs` only | identity |
| `relay/windows-port-scan.ts` | `name` (port owner label) | detailed |
`windows-port-scan.ts` is the one mismatch in the table: it reads only `pid` and
`name`, which the identity set answers, but it calls the detailed reader. On a
host with a live pane that costs nothing extra — the detailed snapshot is
already cached — and on a headless relay it pays for a command line no caller
reads. Left as-is deliberately, because moving it to identity would trade that
for a second scan whenever a pane is polling; revisit if the relay ever scans
ports without one.
The per-pane foreground tracker is the hot one (750 ms / 2 s cadence) and it
genuinely needs the command line, so the repeating per-process `OpenProcess` is
not something the split removes. What the split removes is that handle from
teardown identity and from the owner probe, which now open nothing.
### `creationTimeMs` does not exist on any shipped build
Nothing in the repo supplies a `CreationTime` flag. The package enum is
`None`/`Memory`/`CommandLine`, `process_worker.cc` emits no `creationTimeMs`,
the vendored patch adds none, and `adaptAddon`'s `PROCESS_DATA_FLAG` lacks the
bit. So `creationTimeMs` is always `undefined` in production and
`isWindowsProcessStartTimeAvailable()` is always `false` — a latent product gap
that predates the split and needs its own owner.
Two consequences. `IDENTITY_PROJECTION.flags` evaluates to `0` today, so the
identity reader really does open zero handles. And
`agent-session-process-identity-probe.ts` early-returns on
`isWindowsProcessStartTimeAvailable()` rather than scanning the whole table to
produce `null`. Do not build anything on Windows start time working.
Those CIM numbers are from a 1050-process host. The scan scales with process
count: on a 1486-process Windows SSH host it measured **1.36 s** and produced
**4.8 MiB** of JSON, against the fallback's 3 s and 8 MiB limits. Both limits
match the pre-#15749 reader, so relay hosts are at parity rather than newly at
risk — but the headroom is roughly 2x on time and 1.7x on bytes, not the ~4x the
706 ms figure implies. On overflow the output is truncated, the JSON fails to
parse, and the read rejects, so a busy host loses the table rather than
receiving a wrong one.
## When a read wedges
The vendored reader pushes every callback onto a module-global queue and drains
that queue only when the request holding its `requestInProgress` latch
completes. If a Toolhelp32 snapshot never comes back — an EDR hook, a restricted
token, a worker that dies — the latch is stuck for the life of the process and
every later call parks another closure in that queue.
Two guards, and they work together:
- a **3 s deadline** on each read, so a caller gets a rejection instead of a
promise that never settles;
- a **sticky wedge**: once a read misses its deadline and has not called back,
the module refuses every further read until that read's callback fires.
The wedge used to be a 30 s cooldown that let one probe through per window. That
bounded the _rate_ of new callbacks but not the total: a permanently wedged
reader retained one more closure every 30 s for as long as the app ran, and each
probe also blocked its caller for the full 3 s deadline first. Gating on the
outstanding read instead bounds retention at exactly one callback, and gives up
nothing on recovery — a probe queued behind the latch could never have observed
recovery anyway, whereas the stuck callback firing is the drain itself. On the
relay's bare addon, which has no queue of its own, it is also what keeps Orca
from re-entering `CreateToolhelp32Snapshot` while a call is still running.
That last part is not just tidiness. The addon runs each read as a
`Napi::AsyncWorker`, so a wedged read holds a libuv threadpool slot for good. On
the relay the JS queue is not there to absorb the retries, so one probe per
window would have pinned all four default threads inside ~2 minutes — hanging
every async `fs` and DNS call in that process, not only the process table.
A wedge does **not** engage the PowerShell fallback; see the next section for
why only absence does.
## The relay has no binding, and falls back
Relay deployment installs only `node-pty` and `@parcel/watcher` on the remote
host (`RELAY_NATIVE_DEPS` in `src/main/ssh/ssh-relay-deploy.ts`), so a Windows
machine used as an SSH host has no `@vscode/windows-process-tree` at all. It is
not added there on purpose. Both ways of installing it fail, and both were
checked on a real Windows SSH host with 1486 processes:
**Installing it normally rebuilds from source, and that build fails.** The
tarball carries a `binding.gyp`, so npm runs `node-gyp rebuild` regardless of
what is already compiled inside it. On a host that _already had_ MSVC Build
Tools 2022 installed, that build still failed:
```
error MSB8040: Spectre-mitigated libraries are required for this project.
```
That is the requirement the `binding.gyp` hunk of our patch deletes, and the
patch cannot reach a remote host — pnpm patches do not cross SSH. Relay deploy
would then break outright rather than degrade: `installNativeDeps` throws on
failure, and the toolchain-skip retry is gated to Linux.
**Skipping the build and using the shipped binary returns a truncated table.**
Contrary to what this file used to claim, the published 0.8.0 tarball _does_
contain `build/Release/windows_process_tree.node` — an MSVC build directory that
looks accidentally published (`.obj` and `.tlog` files ship with it). It is
N-API, so it loads on any modern Node. But it predates our patch and still has
the `process_count < 1024` cap, so on that 1486-process host:
```
LOADED OK
rows=1024
selfPid=21964 present=false
```
Exactly 1024 rows, with the querying process itself among the missing. The
self-presence guard rejects that, so the fallback engages anyway — but only on
hosts busy enough to cross the cap. That is worse than no binding at all: it
works on a quiet machine and fails silently under load, which is precisely the
shape of bug that survives testing.
So the constraint is not that no binary exists to ship. It is that the only
binary available to ship is the broken one, and building the good one needs a
toolchain the remote does not have.
Instead, `windows-process-table.ts` falls back to
`readWindowsProcessRowsWithCim` (`windows-process-table-cim-scan.ts`), the
`Get-CimInstance` scan this module replaced. The gate is deliberately narrow:
- it engages **only** when the module cannot be required, never when a loaded
module fails, wedges, or returns an unreadable table — a present-but-failing
reader must not silently start forking a shell at the caller's poll rate;
- a fallback that also fails still rejects, so "unavailable" never degrades into
"nothing is running";
- the scan applies the same self-presence guard as the native path.
`src/main/ssh/relay-native-dependency-coverage.test.ts` asserts that every
native addon reachable from the relay entry is either installed on relay hosts
or listed there with the reason its absence is safe. That test exists because
#15749 shipped this gap: the relay tests injected a fake module through
`__setWindowsProcessTreeLoaderForTests`, so nothing exercised the real require.
## Shipping the native reader to a relay anyway
The scan is the floor, not the destination: it costs ~1.4 s and a `powershell.exe`
where the addon costs ~57 ms. Release builds therefore compile the addon and ship
it as an optional relay artifact.
`config/scripts/build-windows-process-tree-relay-addon.mjs` builds it from the
source pnpm has already patched, on a Windows runner, and refuses to run if
any patch hunk is missing — the Spectre hunk fails loudly, the 1024-process
hunk fails _silently_, and the relative gyp path dies at configure on Windows.
The source is checked rather than the install trusted. It also reads the PE
machine field of the output, because a cross-build that quietly emitted host
arch would ship a binary the target cannot load.
Windows arm64 cross-compiles from the x64 runner — verified on real hardware,
producing `IMAGE_FILE_MACHINE_ARM64` (0xaa64) against x64's 0x8664. It needs the
optional _MSVC v143 ARM64 build tools_ component; without it node-gyp fails with
`MSB8020`, which is why the addon build runs before the long packaging step.
`ORCA_REQUIRE_RELAY_NATIVE_ADDONS` is a per-arch list so a future arch can be
added best-effort before it is promoted to required.
`windows-process-table.ts` binds the bare addon directly rather than the package
wrapper. That wrapper adds only a queue over `getProcessList`, and that queue is
the wedge described above — it latches a module-global `requestInProgress` with
no try/catch. This module already holds a single-flight and a deadline, so going
straight to the addon drops the duplicate.
The artifact is optional in `RELAY_ARTIFACTS`: hashed when present, so a relay
carrying it never shares an immutable directory with one that does not, and
never probed, because requiring a file only a Windows build machine can produce
would make a correct relay read as MISSING and redeploy forever. A local build
on another OS has no addon, so its Windows relays use the scan.
Every desktop package ships relays for Windows hosts, not just the Windows
installer, and the addon also carries the relay launcher (`spawnOutsideJob`,
see `windows-edr-posture.md`). So one Windows job,
`.github/workflows/relay-windows-process-tree.yml`, compiles both arches and
uploads the `relay-windows-process-tree` artifact. The release and dev-channel
macOS and Linux packaging jobs download it into `.build/windows-process-tree`
and set `ORCA_REQUIRE_RELAY_NATIVE_ADDONS=x64,arm64`, as the Windows jobs do.
`config/scripts/relay-windows-process-tree-staging.mjs` checks each staged
binary for its PE machine, the missing `ReadProcessMemory` import, and the
`spawnOutsideJob` export. Without that last check a pre-launcher build from an
old `.build` dir or cached artifact would pass. A required arch that fails any
check fails the build. An unrequired one (a local build) is left out, and that
relay uses the scan and the WMI launch fallback.
## Why the package is patched
`config/patches/@vscode__windows-process-tree@0.8.0.patch` carries six changes.
1. **Spectre mitigation.** The upstream `binding.gyp` requires Spectre-mitigated
libraries, which Orca's Windows build agents do not install. `node-pty` is
patched the same way for the same reason.
2. **The 1024-process cap.** `GetRawProcessList` stopped after 1024 entries.
Measured on a real host with 1051 processes, the module returned exactly
1024 and the querying process was itself among the 27 missing. A truncated
snapshot silently hides the descendants a teardown is trying to reap — the
exact failure the native path exists to remove.
3. **Absolute `node-addon-api` gyp path.** `require('node-addon-api').targets`
is cwd-relative. node-gyp on Windows evaluates it from the pnpm store
realpath, then loads the relative path from the `node_modules` symlink, so
`node_addon_api.gyp` resolves outside the repo and hourly Windows builds
die at configure. `node-pty` is patched the same way for the same reason.
4. **No PEB reads, no `PROCESS_VM_READ`.** See below.
5. **The `CreationTime` flag (4).** Upstream exposes no process start time.
Structured Claude and Codex chat no longer depend on it; the owner probe and
the orcad runtime preflight still read it. `GetProcessCreationTime` opens
`PROCESS_QUERY_LIMITED_INFORMATION` and converts `GetProcessTimes`' FILETIME
to Unix ms; a process that denies the handle is emitted with the field
absent, never zero, because callers must be able to tell "cannot identify"
from a timestamp.
6. **`supportedProcessDataFlags`.** `addon.cc` exports the flag bits the
compiled binary understands, and `lib/index.js` re-exports it.
Why a separate hunk and not just the enum: unlike `node-pty`, this package
publishes a prebuilt `.node` at the same `build/Release/` path node-gyp
writes to. pnpm patches the source tree and leaves that prebuilt alone, so a
host can hold a patched `lib/index.js` — `ProcessDataFlag.CreationTime` and
all — over a binary that ignores flag 4. CI produced exactly that: the gate
read available and every row came back without `creationTimeMs`. Neither a
load check nor a path check can see the difference, so the binary has to say
so itself.
Two readers depend on it. `isWindowsProcessStartTimeAvailable()` returns
false unless this bit is set, so the owner probe never scans the whole table
for times it cannot get (the Windows Claude descendant snapshot that once
relied on it is gone). And `windows-process-tree-creation-time.cjs`
asserts it during install, which is what forces a from-source rebuild —
the same role `node-pty-job-ownership.cjs` plays for node-pty's job exports.
The typings claim `commandLine` is truncated at 512 characters. Measured, it is
not: the longest observed on a real host was 26,059.
### The command line comes from the kernel, not the target's memory
Upstream, `GetProcessCommandLine` opens every process with
`PROCESS_QUERY_INFORMATION | PROCESS_VM_READ` and issues three chained
`ReadProcessMemory` calls — PEB, `RTL_USER_PROCESS_PARAMETERS`, then the string
— to recover the command line. Walking another process's address space for
credentials-adjacent data on a repeating timer is what a credential dumper does,
so Defender for Endpoint scores it as such regardless of intent. Nothing about
the flag sets above changes that; only removing the read does.
Windows 8.1 added `NtQueryInformationProcess`'s `ProcessCommandLineInformation`
class (60), which returns the same string as a `UNICODE_STRING` the kernel
builds, needing only `PROCESS_QUERY_LIMITED_INFORMATION`. Electron's floor is
Windows 10, so every OS Orca supports has it. The entry point is resolved with
`GetProcAddress` on `ntdll.dll` — it has no import library — and the size is
probed with a null-buffer call that answers `STATUS_INFO_LENGTH_MISMATCH`.
The same hunk drops `PROCESS_VM_READ` from `GetProcessMemoryUsage` and
`GetCpuUsage`, which acquired it and never read an address space:
`GetProcessMemoryInfo` and `GetProcessTimes` are satisfied by
`PROCESS_QUERY_LIMITED_INFORMATION`. Measured, both return identical values
under the weaker right on every process that opens at all.
Measured on Windows 11, ~540 processes, counted in-process by replacing the
addon's import table entries with counting stubs:
| per `CommandLine` scan | before | after |
| ---------------------- | ----------------------------------------- | -------------------------------------- |
| `OpenProcess` calls | 543 | 543 |
| desired access | `0x0410` (`VM_READ \| QUERY_INFORMATION`) | `0x1000` (`QUERY_LIMITED_INFORMATION`) |
| `ReadProcessMemory` | 1128 | **0** |
| p50 / p95 | 13.5 / 14.5 ms | 12.3 / 13.5 ms |
Command lines were byte-identical on every process both readers recovered
(405/405, and 399/399 and 376/376 on other runs), including a 24,087-character
argv with embedded quotes, non-ASCII characters and trailing whitespace, and a
WOW64 target. The weaker right is also a strict superset in reach: three
processes that refused `PROCESS_QUERY_INFORMATION | PROCESS_VM_READ` granted
`PROCESS_QUERY_LIMITED_INFORMATION`, and none went the other way.
### There is no PEB fallback, deliberately
An earlier revision kept the PEB reader for a kernel without class 60, behind a
latch. That was wrong, and the reason is worth recording: `ClassifyQueryFailure`
mapped `STATUS_INVALID_INFO_CLASS` / `NOT_SUPPORTED` / `NOT_IMPLEMENTED` from
**any single target** onto a process-wide, one-way switch back to
`PROCESS_VM_READ` plus three `ReadProcessMemory` per pid per scan, for the life
of the process, with nothing observable from JS.
The environment this reader exists for is one where an EDR hooks `ntdll`. A hook
that returns `STATUS_INVALID_INFO_CLASS` for a class it does not recognise would
have silently reinstated the exact primitive the patch removes, on precisely the
machines it was written for — and one stray status from one process was enough.
The same applies under Wine or any instrumented `ntdll`.
So the fallback is gone rather than guarded. `GetProcessCommandLine` returns
false and leaves the command line empty, which is already a normal outcome
(`WindowsProcessRow.command` is documented as empty when a process denies a
query handle, and callers fall back to the image name). Degrading to no command
line is recoverable; silently resuming address-space reads is not.
This also makes the property checkable on the artifact rather than the source:
the patched reader never calls `ReadProcessMemory`, so the symbol is absent from
the compiled addon's import table. `inspectWindowsProcessTreeAddon()` in
`config/scripts/windows-process-tree-gyp-rebuild.mjs` is that check, and it is
the only way to tell the two binaries apart — see below. It answers
`clean` / `unpatched` / `missing` rather than a boolean, because a binary that is
not there has not been cleared, and a caller reading `false` as “verified” would
pass exactly the thing the check exists to catch.
Because the returned `UNICODE_STRING` comes from that same hookable boundary,
its `Buffer` and `Length` are bounds-checked against the allocation before the
characters are encoded, and the probed size is capped at the header plus 64 KiB
(`Length` is a `USHORT`) so a bogus size cannot turn into a `bad_alloc` that
fails an entire scan instead of one process.
### The published tarball ships a loadable unpatched prebuilt
`@vscode/windows-process-tree@0.8.0` publishes
`build/Release/windows_process_tree.node` in the tarball. It is node-addon-api,
so it is ABI-stable and loads cleanly under both Node and Electron — and it was
built from unpatched source, so it performs 1179 `ReadProcessMemory` calls and
opens every process at `0x0410` per scan.
That matters because `allowBuilds` is `false` for this package and CI installs
with `--ignore-scripts`, so nothing compiles it at install time. A `require()`
health check cannot tell the two binaries apart, and a rebuild that is skipped —
`rebuild-native-deps.mjs` soft-exits 0 on a Windows file lock during postinstall
— leaves the upstream prebuilt in place and cached.
Four checks close that, all keyed on the absent `ReadProcessMemory` import:
- `ensureWindowsProcessTreeCommandLinePatch()` deletes a binary that still has
it, so a skipped rebuild fails loudly instead of using the prebuilt;
- `ensure-native-runtime.mjs` treats such a binary as a load failure, which is
what triggers the rebuild;
- the relay build asserts it on the artifact it just produced;
- `loadWindowsProcessTree()` asserts it again on the addon staged beside a relay
bundle and refuses to bind one that still imports the symbol, falling back to
the CIM scan. The build-time assertion is not enough on its own: a bundle and
the addon beside it redeploy independently, so a host that has not taken a new
bundle keeps whatever `.node` is already there.
What none of this does is narrow _which_ processes are asked. A detailed scan
still queries every pid, including `lsass.exe`; it now asks with the same right
Task Manager uses instead of `PROCESS_VM_READ`. Restricting the command-line
pass to Orca's own subtree is the complementary change, and it belongs with the
identity/detailed reader split rather than here — a ppid-derived allowlist would
miss exactly the detached, reparented descendants the trackers exist to find
(#9045, #10475), so it needs the job-object membership as its source of truth.
## Packaging
The addon is Windows-only, so it follows the same contract as
`@orca/windows-registry` (asserted by
`config/scripts/package-electron-runtime-contract.test.mjs`):
- an `optionalDependency`, so a macOS/Linux install tolerates its absence;
- **not** enabled in `allowBuilds` in `pnpm-workspace.yaml` — pnpm installs optional dependencies
on every host, and macOS/Linux must never run `node-gyp` for it;
- listed in the win32 branch of `rebuild-native-deps.mjs` and
`ensure-native-runtime.mjs`;
- copied into the packaged `node_modules` for win32 only.
The relay's copy is a separate artifact staged beside the bundle, so a relay host
only picks up a rebuilt addon on redeploy. Until then it keeps whatever binary it
already has, which is why the addon is checked again at load.
## What the snapshot does not provide
`CreationDate` (process start time) now has an equivalent — `creationTimeMs`,
above — but only inside this module. Daemon identity, managed-hook ownership and
CPU accounting in the memory collector still read a start time through their own
queries; those callers are not migrated.
Committed private bytes have no equivalent either, and the one memory value the
addon can produce is unusable for the sizes Orca now sees: `process.cc` stores
`pmc.WorkingSetSize` into a `DWORD`, so anything above 4 GB wraps — which is why
neither flag set asks for it. That is the second reason
`windows-process-resource-collector.ts` still runs its own
`Get-CimInstance` sweep — it needs `PageFileUsage` (commit) and the CPU-time
counters in the same pass. Migrating it to the native table would cost both, and
it is why this module no longer sets the `Memory` flag at all: the field had no
reader, and asking for it opened a handle per process on every snapshot.
Start time is a proxy for identity, not identity. For the process trees Orca
itself spawns the durable answer is still an inherited handle: a job object
names the tree Orca created, so no start-time comparison is needed. The
`creationTimeMs` this snapshot now carries is for the trees Orca did **not**
create the handle for — a recovered agent session, a descendant walked out of
the table — where a bare PID is all there is to re-identify.
Do not adopt `getProcessCpuUsage()` from the package. It takes both CPU samples
inside one call with a blocking `Sleep(1000)` in the middle, which would hold a
libuv threadpool slot for a full second out of the Resource Manager's two-second
poll.
## Owning a PTY's process tree
`src/main/windows/windows-pty-job.ts` is the counterpart to reading the table:
it answers "is this tree mine, and how do I kill it?" with a handle instead of
an inference.
node-pty is patched (`config/patches/node-pty@1.1.0.patch`) to create a job
object per ConPTY and assign the shell to it under `CREATE_SUSPENDED`, before
the shell can spawn anything. Assigning after the fact leaves a window in which
a fast child escapes the job.
- `terminatePtyJob(proc)` — one `TerminateJobObject` call for the whole tree.
- `listPtyJobProcessIds(proc)` — the live pids under a tree that is still
tracked, including children that detached from the console.
Measured on Windows 11 against a shell whose grandchild was spawned `detached`:
job membership was `[shell, grandchild]` and one call killed both. Neither a
parent-pid walk nor `GetConsoleProcessList` sees that grandchild — it leaves
the console and reparents, which is what left `claude.exe`/`node.exe`/`cmd.exe`
holding worktree directories open (#9045, #10475, #10897).
The per-PTY job deliberately does **not** set
`JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE`. Measured on Windows 11: with that flag,
releasing the handle when the shell exits also kills whatever the user left
running, so typing `exit` in a pane reaped a `start /b` server that used to
survive. The job exists to make an _explicit_ teardown exact, not to redefine
what a clean exit means.
Git Bash needs one additional restriction. The Cygwin runtime — and the MSYS2
fork of it that Git for Windows ships — reads `JOB_OBJECT_LIMIT_BREAKAWAY_OK`
off its own job and then adds `CREATE_BREAKAWAY_FROM_JOB` to **every** child it
spawns when that flag is set (`spawn.cc`, there since 2011), so offering
breakaway hands the whole tree its escape. The per-PTY job therefore omits
`BREAKAWAY_OK` whenever `msys-2.0.dll` or `cygwin1.dll` sits on the shell's DLL
search path — beside the executable, or under `usr/bin` for Git's `bin`
launcher. Native shells keep explicit breakaway. Denying it costs Cygwin
nothing, because it _pre-checks_ the limit rather than retrying, so no spawn
fails; but a _native_ program that passes `CREATE_BREAKAWAY_FROM_JOB` itself
inside such a pane now gets `ERROR_ACCESS_DENIED`. `nohup` and `disown` are
unaffected — they are Cygwin signal/session concepts, unrelated to job
membership. The daemon's host job is unchanged.
Reaping a dead daemon's shells (#9195, #10415) is therefore a **second, nested
job**, not this one. The terminal daemon assigns itself to a kill-on-close job
at startup (`assignHostProcessToKillOnCloseJob`); children inherit membership,
so every pty is covered and the per-PTY jobs nest inside it. Its handle is
released only when the daemon process dies, so a crashed daemon reaps its tree
without changing what a clean shell exit means.
The split is the point. One job answers "kill exactly this pane's tree, now";
the other answers "do not strand anything if the host dies". Trying to get both
from one job is what reaped users' backgrounded work on a clean `exit`.
It belongs to the daemon and never to the app: an app-main crash must still
leave sessions alive, which `.github/workflows/win-crash-survival-e2e.yml`
asserts. The app spawns the daemon `detached` and is itself in no job, so
nothing is inherited across that boundary.
The consequence is that a PTY hosted by the app rather than the daemon gets a
per-PTY job but no crash reaping. That is deliberate — the alternative is a
kill-on-close job on the app, which is exactly what the crash-survival
guarantee forbids.
Once the shell exits, node-pty drops its handle record and closes the job, so a
terminated tree reports `null` rather than `[]`. Null means _unverifiable_ in
the sense of [`ssh-execution-boundary.md`](./ssh-execution-boundary.md) — no job
support, not a ConPTY, or no longer tracked. It is never evidence that
processes died.
Both functions report `unavailable` / `null` rather than a false success when a
pty has no job — an outer job without `JOB_OBJECT_LIMIT_BREAKAWAY_OK` (some EDR
and container hosts) can refuse the assignment, and a pty started before this
build has none. Callers must fall back, not conclude the tree is gone. That
conflation is the original bug.
### Known limitation: the baton table is not synchronised
node-pty keeps its per-terminal handles in a plain `std::vector` and erases from
it on a detached exit thread, while `get_pty_baton` is called from the main JS
thread. That race predates this change — `PtyResize`, `PtyClear` and `PtyKill`
all read the table the same way — but `terminatePtyJob` adds an instance of it:
the exit thread can close `hJob` between the lookup and `TerminateJobObject`.
Losing that race normally just returns `FALSE`, which surfaces as `unavailable`
and falls back. The case that would not be benign is a recycled `HANDLE` value,
where the call could reach a different job in the same process. Fixing it
properly means synchronising node-pty's handle table rather than adding a lock
around one accessor, so it is deliberately left alone here.
### The patch must actually be compiled
node-pty prefers its upstream prebuild and only builds from source when
`npm_config_build_from_source` is set or no prebuild exists for the platform.
The Windows prebuild does **not** contain this patch, so a plain `pnpm install`
on Windows yields a node-pty without the job-object exports — and
`terminatePtyJob` then reports `unavailable` on every call, which is
indistinguishable from a correctly degraded build.
Packaging is unaffected: `rebuild-native-deps.mjs` rebuilds node-pty from source
for Electron and restores the ConPTY runtime files that a bare `node-gyp
rebuild` skips. The gap is the **node-runtime test environment**, which is why
the Windows CI job rebuilds from source before running the win32 suites.
`isPtyJobOwnershipAvailable()` exists for exactly this: the win32 suite asserts
it is true before asserting anything else, so an unpatched binary fails loudly
instead of passing every case vacuously. That guard is what caught this.
`requiresPatchedNodePtySourceBuild()` in `ensure-native-runtime.mjs` now covers
win32 as well, and `pnpm rebuild node-pty` sets `npm_config_build_from_source`
so the patched source build actually replaces the upstream prebuild.
-86
View File
@@ -1,86 +0,0 @@
# Windows setup-runner shell
On native Windows, Orca writes the `orca.yaml` setup script (and the issue command) to a generated
runner file and types a launch command into a terminal. The runner is a **`.cmd` batch file by
default**, exactly as it has been since setup hooks shipped.
A script opts into bash by starting with a `#!` interpreter line:
```yaml
scripts:
setup: |
#!/usr/bin/env bash
[ -f .env ] || cp .env.example .env
pnpm install
```
Without that line the script keeps running under `cmd.exe`:
```yaml
scripts:
setup: |
copy .env.example .env
xcopy /E assets dist
```
## Why the script declares it, not the terminal preference
`terminalWindowsShell` says which shell _interactive terminals_ open in. It says nothing about the
language a project's setup script is written in. Deriving the runner from it had two consequences:
- Windows users with batch-syntax setup scripts silently switched to bash on upgrade, so `copy`,
`xcopy`, `set VAR=value`, and `if errorlevel 1` stopped working.
- Two people on the same repo got different interpreters for the same `orca.yaml`, so no project
could write a setup script that worked for all of its Windows contributors.
A `#!` line is per-project, explicit, and identical for everyone who checks the repo out.
The same rule applies to the per-user setup command in **Settings → repository hooks**
(`repo.hookSettings.scripts.setup`): it is merged into the same script that reaches the runner, so a
POSIX one-liner stored there needs its own `#!` line to run under bash on Windows.
## What the `#!` line does and does not select
The generated runner is always executed by bash (`bash <runner>`; Git Bash on native Windows), on
every platform. The `#!` line therefore does two things:
- It declares the script is written for a POSIX shell, which is what selects the bash runner.
- Its option flags are replayed with `set`, so `#!/usr/bin/env -S bash -euo pipefail` really does
get `pipefail`. Without that replay the flags would be silently dropped, because `bash <runner>`
never parses the interpreter line. Only the flags `set` itself accepts
(`[--abefhkmnptuvxBCHP] [-o option]`) are replayed; invocation-only ones such as `-l` are
dropped, because `set -l` exits 2 and would abort the runner before its first line.
The interpreter name itself is not honored beyond "is this a POSIX shell": `#!/bin/sh` and
`#!/bin/zsh` scripts run under bash, exactly as they already did on macOS and Linux.
## Requirements for the bash runner
A `#!` line only takes effect when Orca can actually launch bash from the configured terminal — the
terminal shell must resolve to Git Bash (`resolveWindowsGitBashShellPath`). The generated runner
uses MSYS `/c/...` paths, which Cygwin and the WSL shim do not accept, and the launch command is
typed into whatever shell the terminal opened with.
When bash is not available (a PowerShell/cmd terminal, or an SSH-to-Windows host, which always uses
the remote's `.cmd` runner) the `#!` script is **not** executed under cmd. The generated `.cmd`
runner prints why and exits 1, because running the interpreter-agnostic prefix of a bash script
(`pnpm install`, `git submodule update`) and only failing at the first bash-only line leaves a
half-set-up worktree that looks finished.
## Launching a `.cmd` runner from a Git Bash terminal
The runner format and the shell that types the launch command are independent: a Git Bash terminal
with a batch-syntax setup script gets a `.cmd` runner launched from a bash pane. `cmd.exe /c
"C:\..."` cannot be used there — MSYS rewrites the bare `/c` switch into a drive path, so cmd opens
interactively and the runner never executes (issue #6896). Those launches reuse the PowerShell
`ProcessStartInfo` launcher (`buildWindowsCmdRunnerDelayedLaunchCommand`), which carries the switch
and the runner path outside the command line. `WorktreeSetupLaunch.shell` therefore describes the
launching pane; the runner file's `.cmd`/`.sh` extension describes the format.
The `wait-for-setup` gate follows the same split. The pane types the gate and already quoted the
agent startup command for itself, so a `.cmd` runner launched from a Git Bash pane still gets the
bash gate — PowerShell's `Invoke-Expression` cannot parse POSIX `'\''` escaping. The gate wraps the
same `ProcessStartInfo` launcher, so the batch runner is never handed to bash.
WSL worktrees and non-Windows platforms are unaffected: they always use the bash runner. SSH hosts
choose their runner from the remote path format, never from local Windows preferences.
@@ -1,137 +0,0 @@
# Windows signing without occupying a runner during approval
Status: implementation proposal; production signing behavior is unchanged.
## Measured cost
In [release run 33821033674](https://github.com/stablyai/orca/actions/runs/33821033674)
(September 4, 2026), the Windows job took 21m56s. The inner-binary download step
took 13m19s and the installer download step took 40s: 13m59s, or 64% of the job,
was spent in the signing download/wait steps. These durations include the
download itself, so they are an upper bound on removable idle time, not a
prediction of net savings after transferring state between jobs.
`release-cut.yml` submits both requests with `wait-for-completion: false`, but
then invokes `Get-SignedArtifact` on the same Windows runner with one-hour and
four-hour completion timeouts. The six-hour job timeout accommodates both
waits. Changing the submission flag again, polling less often, or running the
wait inside a container does not release the runner slot.
This is runner occupancy, not a billing estimate. Standard GitHub-hosted
runners in a public repository may be free; removing the waits still releases
concurrency for other work. Check actual billing before assigning dollar savings.
The same release also occupied an Ubuntu runner for 11m38s while
`run-release-mac-build-workflow.mjs` waited on the isolated macOS workflow.
That is a separate orchestration optimization. Windows development-channel
builds deliberately ship unsigned and have no SignPath wait to remove.
## Proposed execution graph
Keep all Windows stages in the original `release-cut.yml` run to preserve the
current SignPath GitHub artifact provenance boundary:
1. `build-windows` builds and uploads the unpacked app, original installer,
updater metadata, and inner-signing manifest. It submits the inner request,
sends the existing notification, exposes the request ID, and finishes.
2. `package-windows` depends on that job and uses a protected environment named
`windows-inner-signing`. Its runner is allocated only after GitHub approval.
It restores the exact build, downloads the signed binaries with a short,
bounded completion wait, applies the existing signature restoration and
signed `elevate.exe` cache replacement, builds the NSIS installer, uploads it,
submits the second signing request, notifies approvers, and finishes.
3. `finalize-windows` depends on packaging and uses a second protected environment
named `windows-installer-signing`. After approval it downloads the signed
installer, regenerates its blockmap and `latest.yml`, runs existing outer and
inner signature checks, uploads evidence, and uploads the assets to the draft.
4. `publish-release` depends on finalization as well as the existing Linux, macOS,
and blocking release gates. It remains the only job that publishes the draft.
The approver signs in SignPath, waits for that request to finish, and then
approves the corresponding pending GitHub job. Each notification should link
to both places and explain the order. GitHub approval is an extra action;
approving in SignPath alone does not release an environment gate.
## Required configuration
The repository environments were inspected through the GitHub API on
September 5, 2026. Neither Windows environment exists. `adhoc-mac-build` has no
protection rules; it cannot be reused as an approval gate. No SignPath callback
handler was found in the repository's workflows, scripts, application, or cloud
code.
Before enabling the graph:
1. Create both environments in repository Settings → Environments.
2. Add the release approvers as required reviewers for each environment. Decide
whether a release initiator may approve their own job, and configure that
consistently with the existing SignPath policy.
3. Restrict deployment branches to the trusted refs used to dispatch release
workflows, and check that the release workflow's ref passes the restriction.
The workflow ref and the checked-out release tag are different concepts.
4. Read back both environments through the API and verify that
`required_reviewers` rules exist before changing the release graph. Merely
referring to a new environment name in YAML can create an unprotected
environment and silently leave the wait on the runner.
5. Add a preflight assertion for those rules so accidental removal fails before
any signing request is submitted. Verify the API access required for this
assertion using the release workflow's token; do not assume an administrator's
local `gh` access proves workflow-token access.
An automatic alternative requires a SignPath completion callback and an
authenticated integration that releases the corresponding deployment gate.
Confirm the Foundation plan supports the necessary callback before choosing
that architecture. Do not introduce a long-running GitHub polling job as the
callback substitute: it would continue occupying a slot.
## State and failure contracts
- Use artifacts from this exact run and attempt, with a manifest containing the
tag, tag commit SHA, workflow SHA, request IDs, artifact IDs, and SHA-256 hashes.
Artifact names alone are insufficient. Preserve the original unsigned
installer for the existing inner-signing fallback.
- Restore `dist/win-unpacked`, the staging list, the installer, and updater
metadata as one checkpoint. Use an archive to preserve the tree. Do not ship
a fresh rebuild of the app after approving a different binary tree.
- Each new Windows runner needs the pinned Node/pnpm toolchain, build
dependencies, SignPath module, and electron-builder tool cache. The second
runner must populate the NSIS cache before replacing `elevate.exe`; the old
code assumes the first installer build already populated that cache.
- Retain checkout-from-tag behavior and the existing support for release tags
that predate the composite action. Explicitly restore new orchestration code
from the workflow SHA when necessary.
- Preserve the rule that rerunning a workflow never submits a new signing
request. A resume must consume the recorded request and artifacts. Test failed
stage reruns, whole-workflow reruns, and missing/expired checkpoints separately.
- Keep installer signature checks blocking. Keep inner verification evidence
and its current warning-only policy unless changed in a separate decision.
- Resolve the current one-hour inner-signing fallback deliberately: an
environment approval can remain pending longer than one hour and rejection
skips dependent jobs. It cannot reproduce the existing automatic timeout
fallback by itself. A first migration should explicitly document the new
manual release/cancellation behavior; silently treating rejected approval as
permission to ship is not acceptable.
- Keep the release-wide concurrency lock while the graph waits, preventing
another release from overtaking this draft. This saves worker occupancy, but
does not shorten the serialized release queue's human approval time.
## Validation before production
First adapt `windows-signing-rehearsal.yml` to exercise the same staged code
using the auto-approved test-signing policy. Then run a manual rehearsal with
the protected environments and confirm that pending approval has no allocated
Windows runner. Verify signed bytes through the existing extraction-based
installer checks, not only the outer installer signature.
Cover approval before SignPath completion, rejected approval, missing signed
files, changed checkpoint hashes, lost checkpoints, expired artifacts, failed
packaging, and stage reruns without duplicate submissions. Confirm no release
becomes public until all platform and signature gates pass. Compare transferred
artifact/setup time with the original 13m59s wait sample to measure net savings.
A separate `workflow_dispatch` continuation can avoid environment provisioning,
but changes this design substantially: the original release run finishes,
workflow-level concurrency no longer protects the pending draft, and SignPath
must accept artifacts assembled from a prior run. That option needs a durable
release state machine and provenance validation before production use; it is
not a drop-in replacement for the two download steps.
@@ -1,54 +0,0 @@
# Windows terminal shell selection
Two different things can put `cmd.exe` on a Windows terminal, and only one of them makes the
terminal _be_ cmd.
- **`--shell` / `shellOverride`** names the executable the PTY is spawned as. The terminal's own
process is that shell for its whole life.
- **`--command` / `startupCommand`** is text the provider types into whatever shell it spawned.
`--command cmd.exe` therefore starts cmd as a **child** of the host's default shell.
The difference is invisible until the child exits. Leaving that cmd returns the caller's handle to
a Git Bash or PowerShell prompt it never asked for, and anything that keyed off "this terminal is
cmd" is now wrong — while `terminal list` still shows one connected, healthy terminal, because the
PTY never changed.
## Why the runtime path needed its own fix
There are two spawn preflights, and they are twins:
- `src/main/ipc/pty/ipc/spawn-preflight.ts` — renderer/IPC spawns (a terminal tab in the app).
- `src/main/ipc/pty/runtime/spawn-preflight.ts` — runtime spawns: `terminal.create` from the CLI,
headless `orca serve`, and every paired remote environment.
Only the IPC twin read the caller's requested shell. The runtime twin passed a literal `undefined`,
so a runtime-created terminal could only ever be the host's default shell. `orca terminal create
--command cmd.exe` against a Windows environment had no way to say "be cmd" — it could only type
`cmd.exe` into Git Bash. `src/main/ipc/pty/pty-spawn-shell-override-parity.test.ts` pins the pair.
## Rules
- A caller choosing a shell passes `--shell`; a caller running a program passes `--command`. Do not
route a shell choice through `command` — it looks like it worked.
- The allowlist is `isSupportedWindowsShellOverride` in `src/shared/windows-terminal-shell.ts`, and
it is the reason `--shell` cannot name an arbitrary executable. The CLI, the `terminal.create`
RPC schema, and the relay all check the same set; add a shell in one place only.
- Bare shell names only. A path or anything with arguments is refused, so `--shell` can never carry
a command line into `pty.spawn`.
- A host that predates `--shell` STRIPS it (`terminal.create` params are a zod object, which drops
unknown keys) and answers with a healthy terminal running its default shell — a reply that reads
as success. So the CLI gates on `TERMINAL_CREATE_SHELL_SELECTION_RUNTIME_CAPABILITY` and refuses
before creating anything, rather than creating the wrong shell quietly.
- `--shell` is Windows-only, and a host that cannot apply it REFUSES the create
(`terminalShellOverrideRefusal`). macOS and Linux execution hosts spawn the login shell, and a
terminal routed over SSH resolves its shell on the SSH host, whose platform and installed shells
this runtime cannot see. Refusing is the point: spawning the default shell and reporting success
is the failure `--shell` exists to remove.
- A project's execution runtime decides which MACHINE the shell runs on, so it outranks a
per-terminal pick — but it outranks it by REFUSING, not by rewriting. A `--shell` that
contradicts the project runtime (a Windows shell on a WSL project, or a WSL name on a
Windows-host project) is refused. `resolveLocalWindowsTerminalRuntimeOptions` would otherwise
rewrite the value — a WSL project forces `wsl.exe`, a Windows-host project discards a WSL name in
favour of `COMSPEC` — and hand back a terminal running something the caller never asked for. It
also splits an agent launch's quoting from the shell that receives it: POSIX-quoted args typed
into cmd, or cmd-quoted args typed into a WSL shell.
-274
View File
@@ -1,274 +0,0 @@
# Steady-state worktree rescan: Git-admin fingerprint gate
## Status
Adopted for the main-process worktree resolution cache
(`OrcaRuntimeService.listRepoWorktreesForResolution`). It keeps the existing
30-second freshness contract for externally created, removed, moved, locked, and
re-checked-out worktrees while removing the `git worktree list` subprocess that
previously ran for every registered repository every 30 seconds.
## Context
A production trace (10 registered repositories, 3 h 27 min) recorded 4,272
`git worktree` invocations and 8,663 total Git subprocesses. The invocations
arrived in full-fleet sweeps roughly every 30.5 s — one sweep per repository per
`WORKTREE_SCAN_CACHE_TTL_MS`.
### What actually expires the cache
The originating report assumed some specific terminal/status/orchestration
request was invalidating the 30-second cache. That is not what happens, and the
distinction changes the fix:
- `resolvedWorktreeCache` (the whole-fleet snapshot) has a **1-second** TTL
(`RESOLVED_WORKTREE_CACHE_TTL_MS`). Any caller polling faster than 1 Hz
recomputes the snapshot.
- `computeResolvedWorktrees` fans out over **every** registered repo
(`src/main/runtime/orca-runtime.ts`), calling
`listRepoWorktreesForResolution` per repo.
- That per-repo call is backed by `worktreeScanCache` with a **30-second** TTL.
When it expires, the next poll shells out.
So no request "expires" the 30-second cache. It expires on wall-clock time, and
whichever poller arrives first afterwards pays a full-fleet `git worktree list`
fan-out. The many high-frequency callers — `listTerminals` without a selector,
`showTerminal`, `getWorktreePs`, `listManagedWorktrees`,
`resolveWorktreeSelector`, orchestration authority refresh — only determine
_who_ pays, not _how often_. Steady-state subprocess volume is therefore
`repos / 30 s`, independent of poll rate, which is exactly the observed
~1 sweep / 30.5 s.
The in-Orca mutation surface is already event-driven: create, remove, rename,
folder-rename, sparse edits, repo add/update/remove, SSH reconnect, and
mixed-version remote invalidation all call
`invalidateWorktreeScanCacheForRepo` / `invalidateResolvedWorktreeCache`
(≈40 call sites). The 30-second TTL exists for exactly one reason: discovering
worktree changes made **outside** Orca (`git worktree add/remove/move/prune`,
`git checkout` in another worktree, `rm -rf` of a worktree directory).
That is a filesystem question, and the filesystem can answer it without a
subprocess.
## Goals
- Remove the periodic all-repository `git worktree list` fan-out in steady
state.
- Preserve the current ≈30 s discovery latency for externally created, removed,
moved, pruned, locked/unlocked, and re-checked-out worktrees.
- Keep a bounded reconciliation so anything the cheap probe cannot observe still
converges.
- Change nothing for SSH repos, WSL-routed repos, folder workspaces, bare repos,
or hosts where the probe cannot resolve Git's admin layout.
- Fail open: any probe error must behave exactly like today (run the real scan).
## Non-goals
- Changing `RESOLVED_WORKTREE_CACHE_TTL_MS` or the whole-fleet snapshot shape.
- Changing the renderer-facing `worktrees:list` / `worktrees:listAll` IPC scan
cache in `src/main/ipc/worktrees.ts` (separate 5 s cache, invalidated by
`registerWorktreeChangeInvalidator`, not a polling source in the trace).
- Scoping `resolveWorktreeSelector` to a single repository. See
"Rejected alternatives".
- Adding a filesystem watcher per repository.
- Any wire/RPC/persisted-schema change.
## Design
Introduce a cheap, subprocess-free **Git worktree admin fingerprint** for a
local repository, and consult it before re-running a scan whose TTL has expired.
### Fingerprint inputs
`readRepoWorktreeAdminFingerprint(repoPath)` in
`src/main/runtime/repo-worktree-admin-fingerprint.ts` resolves the repo's Git
common directory without a subprocess (read `.git`; if it is a `gitdir:` file,
follow it and then its `commondir`; if `.git` is absent, treat `repoPath` as a
bare gitdir), then records:
| Input | External change it catches |
| -------------------------------------------------- | ------------------------------------------------------------- |
| sorted entry names of `<commonDir>/worktrees` | `worktree add`, `worktree remove`, `worktree prune` |
| existence of `repoPath` | main checkout deleted |
| `<commonDir>/packed-refs` mtime + size | a tip moved while its loose ref is packed away |
| `<commonDir>/reftable` mtime + size | a tip moved under the reftable backend |
| per checkout: `HEAD` contents | branch switch, detach (the detached oid is in HEAD itself) |
| per checkout: contents of the ref HEAD names | a plain `git commit`, `reset`, or `fetch` that moves the tip |
| per entry: `gitdir` contents | `worktree move`, `worktree repair` |
| per entry: `locked` presence | `worktree lock` / `unlock` |
| per entry: existence of the path named by `gitdir` | a worktree directory deleted with `rm -rf` (flips `prunable`) |
"per checkout" covers the main worktree and each linked worktree. Reading the
ref HEAD names is what makes an ordinary commit visible: committing rewrites
`refs/heads/<branch>` and leaves `HEAD` untouched, but moves the oid
`git worktree list --porcelain` prints. A symref target is only followed when it
is a relative path under `refs/`, so a hand-edited `HEAD` cannot steer the probe
outside the ref store.
Every input is a `stat`, `readdir`, or small `readFile` on an already-hot inode,
and the fingerprint depends on nothing but the repo path — no prior scan result
is threaded in. Per-repo fan-out is capped at 8 concurrent linked-worktree
probes, mirroring `SPARSE_CHECKOUT_DETECTION_CONCURRENCY`. A 10-repo fleet with
10 worktrees each costs a few hundred filesystem calls per 30 s window versus 10
process spawns; the spawns dominate by orders of magnitude, and process-table
churn was the reported symptom.
The fingerprint is a NUL-delimited string; any read failure yields `null`, which
the caller treats as "cannot prove unchanged".
### Cache decision
`listRepoWorktreesForResolution` gains one branch on the expired-TTL path:
```text
cached entry exists, same generation + runtimeKey, TTL expired
└─ probe eligible? (no connectionId, no wslDistro, fingerprint recorded)
├─ no → real scan (today's behaviour)
└─ yes → read fingerprint now
├─ null or different → real scan
├─ equal, last real scan < 5 min → extend TTL, no subprocess
└─ equal, last real scan ≥ 5 min → real scan (bounded reconcile)
```
`WORKTREE_SCAN_ADMIN_RECONCILE_INTERVAL_MS` is 5 min, matching the existing
`WORKTREE_SCAN_AGENT_SCRATCH_TTL_MS` precedent. A cached result whose scan
failed (`ok: false`) is never extended, so a transient Git failure still retries
on the 30 s TTL.
### Git version compatibility
Every path the probe reads is part of Git's on-disk layout well before the 2.25
baseline in [`git-compatibility.md`](./git-compatibility.md): `.git` as a
directory or `gitdir:` file, `commondir`, `worktrees/<name>/{HEAD,gitdir,locked}`,
`packed-refs`, and loose `refs/`. `reftable` arrived in 2.45; on older Git it
simply stats as missing, which is a stable value and therefore harmless. No new
Git command is introduced — this change only skips one.
Agent-scratch repos already carry a 5-minute scan TTL
(`WORKTREE_SCAN_AGENT_SCRATCH_TTL_MS`), which equals the reconciliation
interval, so the gate never fires for them and their behaviour is unchanged.
### Ordering
The fingerprint stored with a scan result is captured **before** the scan runs.
A mutation landing while the scan is in flight therefore leaves the stored
fingerprint stale-by-construction, so the next probe sees a difference and
rescans. Capturing after the scan would let that mutation be masked forever.
### Interaction with existing invalidation
Untouched. `invalidateWorktreeScanCacheForRepo` deletes the entry (fingerprint
included) and bumps the generation, so every event-driven path still forces a
real scan on the next read. The fingerprint only ever _extends_ an entry that
the TTL alone would have refreshed.
## Freshness budget
| Change | Before | After |
| --------------------------------------------------------------------- | ----------------- | ----------------- |
| Orca-initiated create/remove/rename/sparse/repo edit | immediate (event) | immediate (event) |
| SSH reconnect / provider generation bump | immediate (event) | immediate (event) |
| External `worktree add/remove/move/prune/lock` | ≤ 30 s | ≤ 30 s |
| External `git checkout` / `commit` / `reset` in any worktree | ≤ 30 s | ≤ 30 s |
| External `rm -rf <worktree>` | ≤ 30 s | ≤ 30 s |
| External sparse-checkout pattern edit | ≤ 30 s | ≤ 5 min |
| Packed/reftable tip moved within one mtime tick at an equal file size | ≤ 30 s | ≤ 5 min |
| SSH / WSL repos, folder workspaces | unchanged | unchanged |
The two regressions are bounded by the reconciliation interval and are both
changes Orca does not make itself.
### Main-thread cost
Both the scan and the probe are asynchronous, so neither "runs on the main
thread" in the naive sense — but they are not equally free there. Measured on
macOS with a 1 ms interval sampling event-loop lag while each ran 30 times
against a repo with 20 linked worktrees:
| | wall per call | main-thread stall per call | worst single stall |
| ------------------- | ------------- | -------------------------- | ------------------ |
| `git worktree list` | 18.66 ms | 2.69 ms | 3.02 ms |
| fingerprint probe | 1.66 ms | 0.01 ms | 0.04 ms |
`fs/promises` dispatches to libuv's threadpool, so ~99 % of the probe's latency
is off-thread. Spawning Git does not: `uv_spawn`, fd and pipe setup, and stdout
collection and decoding are real synchronous main-process work. Ten repos
refreshing together therefore cost ≈27 ms of event-loop stall per sweep before
this change and ≈0.1 ms after — this reduces main-thread pressure rather than
adding to it, which is why moving either side onto a worker thread would not
help.
### Residual risk
A repo registered at a UNC path (`\\wsl$\...`) but executed by the local Windows
Git runtime is still probed, because it is not WSL-routed from Orca's point of
view. The probe is correct there and strictly cheaper than the subprocess it
replaces, but its filesystem calls cross the 9p boundary like the existing
sparse-checkout probes already do.
## Measured effect
`src/main/runtime/worktree-scan-admin-fingerprint-gate.test.ts` drives the
reported steady state — 10 idle local repos, a caller polling at 1 Hz for 30
simulated minutes — and counts `git worktree list` invocations:
| | `git worktree list` per 30 min | per hour |
| ------------------------ | ------------------------------ | -------- |
| TTL only (before) | 600 | 1,200 |
| fingerprint gate (after) | 60 | 120 |
A 90 % reduction, with the remainder being the bounded reconciliation. Repos
with genuine external activity keep rescanning at the 30 s cadence because the
fingerprint flips.
Extrapolating to the original trace's shape (10 repos, 3 h 27 min): 4,272
`git worktree` invocations would become ≈427.
## Rejected alternatives
**Raise `WORKTREE_SCAN_CACHE_TTL_MS` to 5 min.** One line, same subprocess
reduction, but it degrades _every_ external-change latency to 5 min, including
the common "I ran `git worktree add` in a terminal" case. The fingerprint buys
the same reduction without that regression.
**Per-repo `fs.watch` on `<commonDir>/worktrees`.** Lower latency, but adds
persistent watcher handles per repo, inherits recursive-watch platform
differences, and would need its own dormancy/rearm story. The existing watcher
infrastructure is scoped to workspace files; extending it to Git admin dirs is a
larger change with a worse risk profile for the same steady-state win.
**Scope `resolveWorktreeSelector` to the owning repository.** The original brief
asks for this. `listTerminals` already avoids the fan-out for explicit worktree
ids via `buildResolvedWorktreeFromId` +
`listKnownResolvedWorktreesForExplicitTarget`. Extending that to
`resolveWorktreeSelector` means splitting the highest-fan-in method in the
runtime (38 call sites) and rebuilding lineage projection for a repo subset —
material regression risk. Once the fingerprint gate lands, the remaining
fan-out cost for a targeted call is a batch of stats, not a subprocess, so the
gain no longer justifies the risk. Tracked as follow-up, not in this change.
## Test plan
`src/main/runtime/repo-worktree-admin-fingerprint.test.ts` (real temp dirs, real
`git` binary — a mocked filesystem would only restate the assumptions):
- stable across repeated reads with no change
- changes after `worktree add`, `worktree remove`, `worktree move`,
`worktree lock`, `checkout` in a linked worktree, `checkout` in the main
worktree, a commit in either, and `rm -rf` of a worktree directory
- still tracks a tip whose loose ref has been packed away by `git pack-refs`
- a linked worktree path and the main repo path produce the same fingerprint
- a bare repo resolves through its own gitdir and still tracks `worktree add`
- `null` for a non-Git directory and for a missing path
`src/main/runtime/worktree-scan-admin-fingerprint-gate.test.ts`:
- unchanged fingerprint suppresses the rescan past the 30 s TTL and re-arms it
- changed fingerprint rescans at the 30 s TTL
- the 5-minute reconciliation forces a rescan while the fingerprint is unchanged
- `notifyBranchRenamed` (event invalidation) still forces an immediate rescan
- a `null` fingerprint (probe failure) falls back to scanning
- SSH repos never consult the probe
- a failed scan is never extended
- concurrent callers share one probe and one scan
- the 1 Hz / 10-repo workload measurement above
-173
View File
@@ -1,173 +0,0 @@
# Running commands inside WSL
Two properties of `wsl.exe` decide how every guest invocation has to be written. Both are silent
when you get them wrong: the command still runs and still exits 0, it just returns the wrong bytes.
Those are sections 1 and 2. A closing section answers the question that running Orca's writes
inside a distro raises next: what happens to the distro's disk image.
## 1. Always `--exec`, never `--`
`wsl.exe -d <distro> -- <argv>` expands `$name` in **every argument** against the guest environment
before the guest runs. This is `wsl.exe` itself, not the guest shell — it happens with no shell in
the command at all:
```
$ wsl.exe -d Ubuntu-24.04 -- /usr/bin/printf %s '$HOME'
/home/you
$ wsl.exe -d Ubuntu-24.04 --exec /usr/bin/printf %s '$HOME'
$HOME
```
So under `--`, a script means something other than what it says. `awk '{print $2}'` reaches the
guest as `awk '{print }'` and prints the whole line; a positional `"$1"`, a shell local, and a
`"\$literal"` are blanked or rewritten the same way. (Expansions with no `$` are unaffected — a
`sed` backreference like `s/(a)(b)/\2\1/` survives either way.) Escaping `$` on the Windows side
cannot fix this reliably — an earlier attempt skipped every `$` preceded by a backslash, which is
exactly the case a POSIX script uses to mean a literal dollar.
Build argv with `buildWslExecArgs()` in `src/shared/wsl-login-shell-command.ts`. A test walks the
tree and fails if the `--` form reappears.
The `--` inside `sh -s -- <path>` is a _shell_ argument separator and is unrelated; leave it alone.
## 2. Machine-read output must be fenced
Orca runs guest commands through the distro user's **interactive** login shell (`-ilc` for
bash/zsh) because that is the only shell that reads `~/.bashrc`, where `nvm`, `mise` and `asdf`
install their PATH entries. Dropping `-i` would break tool detection for those users.
The cost is that an interactive shell also runs the distro's rc/motd, and that output goes to
**stdout** — the same stream the answer arrives on. Stock Ubuntu 24.04 needs no customization to
reproduce it:
```
$ wsl.exe -d Ubuntu-24.04 --exec bash -ilc 'git --version'
To run a command as administrator (user "root"), use "sudo <command>".
See "man sudo_root" for details.
git version 2.43.0
```
Any caller that parses stdout must use `buildWslCapturedLoginShellCommand()`, which fences the
payload and returns a matching `readStdout`. `.trim()` does not help: the banner is a prefix, not
surrounding whitespace, so a stat probe compared against `"directory"` simply never matches.
The fence carries a per-call nonce so that `cat`-ing a file whose contents happen to quote a marker
is not truncated, and it preserves the payload's exit status so `exit 2` → `ENOENT` mappings keep
working.
**Do not fence a command that `exec`s into a long-running program** (`codex app-server`, an
interactive terminal). It never reaches the closing fence, and there the shell's own output either
belongs to the program or is what the user wants to see.
## Prefer no shell at all
When a caller only needs a known binary with a known environment, skip the login shell entirely and
run the binary directly:
```
wsl.exe -d <distro> --exec /usr/bin/env PATH=… HOME=… /usr/bin/git -C <dir> status
```
This is what the direct-git read path does. It is immune to both problems above by construction and
avoids paying login-shell startup on every call, which also sidesteps profiles that block or print.
Resolve the PATH/HOME once through a fenced probe, cache it per distro, then use this form.
## Disk: the distro VHDX only grows
Everything above puts Orca's writes inside the distro, which raises a separate question. WSL2 keeps
the entire guest filesystem in a single dynamically-expanding `ext4.vhdx`. Deleting files inside
the distro does free the blocks — for ext4 to reuse — but the host-visible `.vhdx` does not shrink
on its own.
### Finding the file
The path depends on how the distro was installed, so do not assume one:
| Install method | `ext4.vhdx` lives under |
| ---------------------- | --------------------------------------------------------- |
| recent `wsl --install` | `%LOCALAPPDATA%\wsl\{guid}\` |
| Microsoft Store | `%LOCALAPPDATA%\Packages\<PackageFamilyName>\LocalState\` |
| `wsl --import` | wherever the operator pointed it |
The install-agnostic answer is the registry, which records every distro's directory as `BasePath`:
```powershell
Get-ChildItem HKCU:\Software\Microsoft\Windows\CurrentVersion\Lxss |
ForEach-Object { Get-ItemProperty $_.PSPath } |
Select-Object DistributionName, BasePath
```
### Measured behavior
WSL 2.7.11.0 / Ubuntu-24.04, one machine. Sizes are **size on disk** — allocated bytes, via
`GetCompressedFileSize`, not the logical file length. That distinction matters below: on this
machine the vhdx was not sparse, so the two numbers were identical, but on a sparse vhdx the
logical size stays pinned at the high-water mark while only size on disk falls when space is
reclaimed. Measure the wrong one and reclaim looks like it did nothing.
| Step | `ext4.vhdx` size on disk (bytes) |
| ---------------------------------- | -------------------------------- |
| baseline | 21,673,017,344 |
| write 1 GiB | 22,746,759,168 |
| delete it | 22,746,759,168 |
| write a fresh incompressible 1 GiB | 22,746,759,168 |
| hold 3 GiB live at once | 24,894,242,816 |
The fourth row is the point: the second gigabyte cost zero growth, because ext4 handed it the
blocks the first one freed. Only exceeding the previous peak moved the file.
### Reclaiming space
Sparse mode lets the guest hand freed blocks back to the host, so the file can shrink instead of
only growing. Two preconditions, both easy to miss: the distro has to be stopped (the vhdx cannot
be converted while it is mounted), and `wsl --manage` exists only on WSL 2.5 and newer — check with
`wsl --version`.
```
wsl --terminate <distro>
wsl --manage <distro> --set-sparse true
```
The equivalent for distros not yet created is `sparseVhd = true` under `[experimental]` in
`%UserProfile%\.wslconfig` — that file lives in the Windows user profile, **not** inside the distro
and not at `~/.wslconfig`, and does not exist until you create it.
Neither touches slack that already exists. For that, shut WSL down and compact the file by hand
from an **elevated** prompt. `compact vdisk` on its own fails because no virtual disk is selected,
so the `select` and the read-only `attach` are required, not optional:
```
wsl --shutdown
diskpart
DISKPART> select vdisk file="C:\path\to\ext4.vhdx"
DISKPART> attach vdisk readonly
DISKPART> compact vdisk
DISKPART> detach vdisk
DISKPART> exit
```
Two caveats, neither verified here: field reports say `compact vdisk` is a no-op on a vhdx that is
already sparse (convert back with `--set-sparse false` first), and sparse mode's runtime cost was
not measured. Microsoft's [disk-space guide](https://learn.microsoft.com/windows/wsl/disk-space)
carries the current locate/expand/compact procedure and the `--manage` version floor;
[`.wslconfig`](https://learn.microsoft.com/windows/wsl/wsl-config) carries `sparseVhd`. Enabling
sparse mode and compacting are both per-machine decisions; Orca does not make either.
On the measured machine the vhdx was **not** sparse: `fsutil sparse queryflag` reported "NOT set as
sparse", and no `%UserProfile%\.wslconfig` existed to opt in. Microsoft documents `sparseVhd` as
defaulting to `false`, so that is the expected state rather than a local quirk — but the flag is
per-vhdx, set when the disk is created or by an explicit conversion, so check your own distro
rather than assuming either way.
### What this means for Orca
A vhdx that grows as speculative worktree preparation and mirrored worktrees write into the distro
is expected. Its size is monotonically non-decreasing and roughly tracks peak concurrent usage —
but it can drift above peak, and the measurement above is the best case for reuse: the second
gigabyte was allocated immediately after the first was freed, out of the same block group. Under
sustained churn — many worktrees created and removed over weeks, no `fstrim`/discard, sparse off —
allocation spreads and the file settles higher than live peak.
So growth on its own is not evidence of a leak. Growth well above live peak usage is worth
investigating, and is the case the reclaim steps above address.
-37
View File
@@ -1,37 +0,0 @@
# CLI access in Orca-managed WSL shells
A WSL terminal on a Windows host gets this app's CLI (`orca-ide` packaged,
`orca-dev` in development) on its PATH with nothing installed in the guest:
`~/.local/bin`, shell profiles, and the Windows user PATH are untouched. External
WSL shells still need Settings → General registration.
1. **Host.** `buildPtyHostEnv` calls `getManagedWslCliDir` for WSL panes only. It
writes a launcher and PowerShell bridge (reusing `wsl-cli-scripts.ts`) under
`<userData>/wsl-managed-cli/<content hash>` and exports `ORCA_WSL_CLI_DIR`.
Content addressing follows `shell-wrapper-content-address.ts`: builds sharing
user data never overwrite each other, and a present file is complete because
each one lands by rename. Old directories are not collected.
2. **Crossing.** `addOrcaWslInteropEnv` adds `ORCA_WSL_CLI_DIR/p`, so both the
daemon and in-process spawn paths translate it with the distro's own mounts.
3. **Guest.** `WSL_MANAGED_CLI_PATH_RESTORE` runs after user startup files in the
bash rcfile and the local zsh first-prompt hook, which run once per Orca shell.
It leads PATH with the directory when `$ORCA_WSL_CLI_DIR/$ORCA_CLI_COMMAND` is
executable, and otherwise prints one warning. Other login shells get no CLI;
nothing blocks a shell.
The colocated launcher finds its bridge beside itself and PowerShell by Windows
path, so neither guest PATH nor the automount root matters. The bridge pins this
app's user-data directory and is written with a UTF-8 BOM so Windows PowerShell 5.1
reads non-ASCII paths correctly. It clears `ORCA_WSL_CLI_DIR`, which WSLENV maps
back to Windows, so an app the CLI starts never inherits it. In development it runs
Electron as Node on `out/cli/index.js` directly, with the environment
`buildWindowsDevLauncher` sets (`ORCA_APP_EXECUTABLE`, stashed `NODE_OPTIONS`).
Otherwise it launches its child exactly like the registered bridge.
A missing runtime (logged once) or a failed write (logged per spawn) leaves
`ORCA_WSL_CLI_DIR` unset. A terminal daemon from an older build adds no WSLENV
entry, so its WSL panes lack the CLI until the daemon restarts.
Run the opt-in end-to-end test on Windows with `ORCA_BACKGROUND_LAUNCH=1`,
`ORCA_TEST_MANAGED_WSL=1`, and optionally `ORCA_TEST_WSL_DISTRO=<distro>`. The zsh
case skips when the distro lacks `zsh` or `script`.
@@ -1,66 +0,0 @@
# WSL probe failure semantics
A WSL probe answers a question about a distro: is `git` installed, what is
`$HOME`, which distros are running. Every one of those probes can fail for a
reason that has nothing to do with the answer — the distro is booting, `wsl.exe`
is slow under load, the VM was just shut down.
The recurring bug in this subsystem is reporting that failure as a negative
answer.
## The shape
```ts
try {
await execCommandInWslOrThrow(target, `${shellQuote(command)} --version`)
return true
} catch {
return false // "not installed" and "could not ask" are now the same value
}
```
Nothing downstream can tell those two apart, because by this point they aren't
two things.
## Why it keeps shipping
Swallowing on its own is survivable. An uncached caller asks again a moment
later and the answer corrects itself, so the bug stays invisible in review and
in manual testing.
It becomes user-visible when the swallowed value is **cached** or used to
**gate discovery**. Then a distro that was busy for one second reports no git,
or no agent sessions, until the app is relaunched. The failure is sticky,
silent, and indistinguishable from the real thing.
Three instances so far:
| Where | What the user saw | Status |
| ------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- |
| Preflight CLI probes | Caching the result would have pinned "git not installed" until relaunch | Bounded entry ([#17350](https://github.com/stablyai/orca/pull/17350)) |
| `glab auth status` fallback into WSL | Idle VM woken repeatedly for users who never touch GitLab | Open ([#8941](https://github.com/stablyai/orca/issues/8941)) |
| `listRunningWslDistrosAsync` | Fails closed to `[]` with no last-known-good, polled every 2s — a persistently broken `wsl.exe` makes every WSL session vanish app-wide | Open (PR #17072 review) |
## What to do instead
Pick the cheapest option that fits the call site.
1. **Don't pin it.** If the probe is cheap and uncached, swallowing is fine —
the next call self-heals. This is what most of `src/` legitimately does.
2. **Bound the entry.** If you cache, give it a TTL so a transient failure
expires instead of lasting the session. Cheap, no signature change, and what
[#17350](https://github.com/stablyai/orca/pull/17350) does.
3. **Keep last-known-good.** If the probe gates discovery, fall back to the
previous successful answer on failure rather than to empty. `listWslDistrosAsync`
in `src/main/wsl.ts` already does this — `listRunningWslDistrosAsync`, added
beside it, does not.
4. **Propagate the third state.** The durable fix: return
`present | absent | unreachable` instead of a boolean, so a caller cannot
accidentally treat "could not ask" as "no". This reaches past WSL into shared
exec code and hasn't been done.
Whichever you pick, document why the fallback is safe for callers to cache or use for discovery.
## Reviewing failure fallbacks
Review the caller as well as the catch: whether a fallback is cached or gates discovery is a dataflow question. Returning `false`, `[]`, or `null` after a failure can be safe for some operations, but a WSL probe must preserve the distinction between absence and a distro that could not be reached. Prefer behavioral tests that exercise probe failure, discovery, caching, and recovery together.
-27
View File
@@ -1,27 +0,0 @@
# Verifying the W1–W3 Windows/WSL work
Unit and real-binary tests cover Windows and WSL behavior.
## 1. Unit — runs everywhere, every PR
| Suite | Pins |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `src/main/wsl/wsl-runner.test.ts` | Separator, lane selection, fencing, WSLENV, guest cwd, script interpreter, budget split, refusal on unresolved PATH |
| `src/main/wsl/wsl-guest-environment.test.ts` | Burst collapse, per-distro isolation, malformed-payload rejection, transient vs permanent, retry windows, joiner budget |
| `src/main/wsl/wsl-w1-w3-contract.test.ts` | The W1→W3 chain end to end: absolute `wsl.exe`, argv array, bounded call, no `--`, script byte-identical, WSLENV, no shell on probe, login PATH still applied |
| `src/shared/source-scan/source-tree-scan.test.ts` | The guard helpers. A guard that under-reports is worse than none |
## 2. Real-binary — the assertions nothing else can make
**Windows CI** (`package (windows)` job in `pr.yml`) rebuilds node-pty from patched source and runs the `win32` suites against a real ConPTY: a real detached grandchild, a real job kill, and the inverse — a clean `exit` must leave backgrounded work alone.
**Real WSL distro** — not in CI; WSL isn't available on hosted runners.
```
ORCA_REAL_WSL_RUNNER_TEST=1 ORCA_WSL_TEST_DISTRO=Ubuntu-24.04 \
pnpm vitest run src/main/wsl/wsl-runner.wsl.test.ts
```
It appends `sleep 60` to the distro's `~/.profile` and asserts the probe lane still answers inside its budget — **#14288 reproduced, not simulated** — then restores the profile. Also covers banner stripping, a script carrying quotes and `$` arriving byte-identical, WSLENV crossing, and guest cwd.
Run this before shipping a change to `src/main/wsl/`. It is the only evidence that the probe lane does what the workstream claims, and it has already gone stale once against a runner change while passing in CI, because CI skips it.
-320
View File
@@ -1,320 +0,0 @@
# xterm Patch Regeneration
## Scope
Orca ships `@xterm/xterm` with four source changes it needs and upstream has
not taken: the IME composition hooks, the `xterm-composition-*` custom events
they raise, the `ICompositionHelper` surface those hooks widen, and a `SortedList`
fix. pnpm applies them through `config/patches/@xterm__xterm@<version>.patch`.
That patch touches eight files. Four are hand-authored source
(`src/browser/CoreBrowserTerminal.ts`, `src/browser/Types.ts`,
`src/browser/input/CompositionHelper.ts`, `src/common/SortedList.ts`) and four
are the build output those sources produce (`lib/xterm.js`, `lib/xterm.mjs`,
and both sourcemaps). The bundle half is 7.3 MB of minified code. It is
generated, and this document exists so nobody edits it by hand.
The two halves are the same edits diffed two ways, so the generator requires
them to match byte for byte on every source file. A hunk the shipped patch
cannot name — upstream's `.npmignore` strips `src/**/*.test.ts` — would be
dropped by the next `--write`, so it fails the run instead.
`config/patches/xterm-src/@xterm__xterm@<version>.src.patch` is the source of
truth. Everything else is derived from it by
`config/scripts/regenerate-xterm-patches.mjs`, which is pinned to the exact
upstream commit the published tarball was built from.
`@xterm/addon-webgl`, `@xterm/addon-search`, `@xterm/addon-serialize` and `@xterm/addon-image` are
generated the same way, from their own source patches under
`config/patches/xterm-src/`. Their entries differ only in `packageDir` and build
steps; everything below applies to all five. `@xterm/addon-ligatures` is the one
patch still written by hand — see [Known Gaps](#known-gaps).
The image patch bounds pending Kitty decoders by their maximum WASM capacity
and caps transmitted image blobs by byte size. Both use the configured storage
budget; upstream's displayed-pixel budget does not cover these allocations.
Byte-budget eviction drops unplaced payloads first, so a new upload cannot erase a
visible image while abandoned blobs still hold budget; displayed images go only
when that is not enough, because the cap is a hard bound. The incoming image is
always stored, so the cap overshoots by at most one payload rather than dropping
an image the protocol already acked as `OK`. Orca uses fixed 32 MB storage and
8 MiB sequence limits, not arbitrary addon configurations.
`config/scripts/xterm-image-memory-contract.test.mjs` exercises the installed
bundle with unfinished uploads, chunk continuation, both eviction orders and
disposal.
The patch also bounds decompression before joining decoded chunks, validates PNG
dimensions before native decoding, and closes stale asynchronous image results
after reset, disable or disposal. `config/scripts/xterm-image-lifecycle-contract.test.mjs`
exercises those boundaries against the installed addon. Font zoom scales visible
tiles without creating enlarged full-image canvases;
`config/scripts/xterm-image-resize-contract.test.mjs` checks allocation and tile mapping.
Every wasm memory reserves its guard region inside V8's sandbox, which caps a
renderer at about 124 live memories no matter how much RAM is free. Upstream
gave each terminal a SIXEL decoder at activation and kept IIP decoders after the
first image, so a window with ~120+ terminals ran out. The patch borrows SIXEL
decoders from a small shared pool only while a sequence is open, keeping color
registers on the terminal, and drops IIP decoders after each image. A decoder
that cannot be allocated drops that image; before the patch it threw out of the
parser and left the terminal's write queue stuck.
`config/scripts/xterm-image-wasm-budget-contract.test.mjs` covers both.
## Rules
1. Never edit `config/patches/@xterm__*@<version>.patch`. Edit the source
patch and regenerate.
2. Never edit `lib/` inside a patched `node_modules` tree and re-run
`pnpm patch-commit`. That is how bundle hunks stop matching their sources.
3. Every source change must land together with the regenerated bundle hunks and
the `pnpm-lock.yaml` hash bump, in one commit.
4. The upstream commit lives in `config/patches/xterm-upstream.json`, not in a
comment. A version bump that leaves it stale fails the generator, it does not
silently patch the wrong tree.
5. Sourcemaps move with the bundle, and are never silently omitted. The patch
moves the code, so dropping only the map hunks would ship offsets pointing at
the wrong lines. `sourcemaps.policy` accepts `include` and nothing else: it
costs about 5.8 MB of the emitted patch and is required because
`src/renderer/src/components/terminal-pane/terminal-ime-xterm-transaction-events.test.ts`
reads `lib/*.map` and asserts the mapped `Version.ts` matches the runtime
version. Deleting the maps was once an option; the code that did it was
removed as unreachable, so re-adding the policy means re-adding that code.
6. `--check` is the authority on the lockfile, not `pnpm install`. pnpm writes the
patch hash in two places — `patchedDependencies` and every resolution key that
depends on the patched package — and on a warm store it will leave the
resolution keys at their previous value while reporting success. That installs
locally and drifts on CI's cold store. For a version bump, follow the **Version
Bumps** workflow through step 5 (the final `--check`); if it reports a stale hash
after an install, rerun `--write`. For a source-only edit, the four-step workflow
above ends at `--check`.
## Workflow
```sh
# 1. Edit the source hunks.
$EDITOR config/patches/xterm-src/@xterm__xterm@6.1.0-beta.303.src.patch
# 2. Rebuild the bundle hunks, the full patch, and the lockfile hash.
node config/scripts/regenerate-xterm-patches.mjs --write
# 3. Reinstall so node_modules picks up the new patch hash.
pnpm install
# 4. Confirm the tree is self-consistent.
node config/scripts/regenerate-xterm-patches.mjs --check
```
Editing a patch file by hand is awkward for anything larger than a one-liner.
For a substantial change, work in the generator's own checkout instead — after
any run it is left at the pinned commit with the source patch applied:
```sh
node config/scripts/regenerate-xterm-patches.mjs --check --work-dir=/tmp/xterm
$EDITOR /tmp/xterm/upstream/src/browser/input/CompositionHelper.ts
git -C /tmp/xterm/upstream diff -- src/ > config/patches/xterm-src/@xterm__xterm@6.1.0-beta.303.src.patch
node config/scripts/regenerate-xterm-patches.mjs --write --work-dir=/tmp/xterm
```
For an addon, edit under `addons/<name>/` and take the diff from that directory
with `--relative`, so the patch is rooted at the package the way the published
tarball is:
```sh
$EDITOR /tmp/xterm/upstream/addons/addon-webgl/src/TextureAtlas.ts
git -C /tmp/xterm/upstream/addons/addon-webgl diff --relative -- src/ \
> config/patches/xterm-src/@xterm__addon-webgl@0.20.0-beta.299.src.patch
```
`--write` rewrites the source patch into the canonical form it would emit on a
re-diff, so a hand-produced `git diff` gets normalized on the first run rather
than fighting `--check` forever.
Run the checkout outside this repository. A build tree underneath it makes
`tsgo` walk up into Orca's own `node_modules` and fail with `TS2300: Duplicate
identifier`, which is a symptom of where the tree sits and not of the patch.
## How the Commit Is Known
Upstream `bin/publish.js` sets `packageJson.commit` before `npm publish`, so
each published tarball names the commit that built it. The generator asserts
that stamp against `xterm-upstream.json` and then compares the tarball's `src/`
against the checkout file by file. Only `src/common/Version.ts` may differ,
because `publish.js` rewrites the version immediately before packaging; the
generator applies the same stamp.
That pair of checks is what makes the rebuild trustworthy. Without them a wrong
commit would still produce a plausible-looking 7 MB patch.
## Build Order
Upstream's publish path is `npm ci` → stamp `Version.ts` → `npm run package`.
`npm run package` runs webpack for `lib/xterm.js` and then, via `postpackage`,
`bin/esbuild_all.mjs --prod` for `lib/xterm.mjs`.
An addon needs three steps, in this order, and the first is easy to miss:
1. **root `npm run build`.** The addon's own `npm run build` is
`tsgo -p .` against a tsconfig whose `files` and `include` are both empty and
which only lists project references. In `-p` mode tsgo does not build
references, so it succeeds while emitting nothing, and the addon's webpack
then fails on a missing `./out/`. The root build is what populates it.
2. **addon `npm run package`** — the addon's own webpack, which emits the CJS
`lib/addon-*.js`. The root `package` script never builds this.
3. **root `npm run esbuild-package`** — `bin/esbuild_all.mjs --prod`, which emits
the ESM `lib/addon-*.mjs` for every addon at once.
**Do not run `npm run setup` after the packaging build.** `setup` is the
development esbuild pass with `minify: false`. Running it afterwards overwrites
`lib/xterm.mjs` with an unminified bundle and a map that no longer matches, and
the resulting patch is silently wrong — the failure mode is a `.mjs` that is
50% larger than the published one, which is easy to miss inside a 7 MB diff.
`forbiddenBuildScripts` in the manifest encodes this and the generator refuses
to run a build step that names one of those scripts.
The generator also builds the _unmodified_ commit first and asserts that it
reproduces the published `lib/` byte for byte before it emits anything. A
toolchain or build-order problem therefore surfaces as an explicit "did not
reproduce the published bundles" error rather than as 7 MB of mystery diff.
## Recovering From Hand-Edited Bundles
Between 2026-08-09 and 2026-08-17 this harness did not exist, and four fixes
landed by editing the minified bundles directly. The tell is code no minifier
emits: `const` in an otherwise `let`-only bundle, and identifiers like `$rl`,
`$hp`, `$tid`.
Recovery is not a rewrite. The hand-edits were applied to `src/` as well, so the
source hunks in the shipped patch were already correct and `--write` re-derives
the bundles from them. What changes is cosmetic and expected:
- Hand-written locals collapse back into minifier names, which shifts esbuild's
frequency-ordered allocation and can swap two short names bundle-wide (`i`↔`t`
in the `.mjs`, `w`↔`y` in the `.js`). Most differing lines are the same length.
- Hand-written equivalents normalize to what the toolchain actually emits
(`!!x` back to `Boolean(x)`, an escaped `\u200E` back to the literal
character).
To confirm a regeneration is semantically a no-op rather than a revert, compare
identifier multisets between the old and new bundle instead of reading the diff:
every name that is not a single-letter minifier local should appear the same
number of times in both. Anything else is a real change and needs explaining.
## The Lockfile Moves With the Patch
pnpm derives the `patchedDependencies` hash in `pnpm-lock.yaml` — and the
`.pnpm/@xterm+xterm@<version>_patch_hash=<hash>/` store directory name — from
the sha256 of the patch file itself. A regenerated patch without the lockfile
bump fails `pnpm install --frozen-lockfile` on every machine except the
author's. `--write` makes that edit; `--check` fails if it is missing.
`config/scripts/regenerate-xterm-patches.test.mjs` asserts the same thing
without a network or a build, so the ordinary test job catches lockfile drift
in milliseconds even though the full rebuild runs in its own CI lane.
## Toolchain Pin
`toolchain` in the manifest records what upstream's `package-lock.json` resolves
at the pinned commit, and the generator fails if `npm ci` produces something
else. The entry that matters is `@typescript/native-preview`
(`tsgo`), which upstream pins to a **dated development build** —
`7.0.0-dev.20260521.1` at the time of writing. It is a real published version
and npm does not prune old releases, but it is the one dependency of this scheme
that is not a stable release.
If that version ever becomes unresolvable the generator fails with a toolchain
error naming it. Recovery is to move the pin to the next upstream commit whose
`package-lock.json` resolves, re-verify that the rebuild still reproduces the
published bundles, and regenerate. The committed patch keeps working the whole
time — only regeneration is blocked, so this is never an outage.
## Patch Path Rooting
A published tarball is rooted at the package, so an addon's patch names
`src/TextureAtlas.ts`, not `addons/addon-webgl/src/TextureAtlas.ts`. Two places
have to agree with that, and both fail silently if they do not:
- The checkout diff passes `--relative`, which must sit **before** the `--`
separator in `CHECKOUT_DIFF_FLAGS`. After it, git reads it as a pathspec and
keeps repo-root-relative paths, and every source hunk then falls out of the
emitted patch.
- `git apply` runs from the repo root with `--directory=<packageDir>`. Run from
a subdirectory instead, git still resolves patch paths from the repo root,
skips every hunk, and **exits 0**. The generator guards this by failing when
applying a source patch leaves the checkout unchanged.
## Version Bumps
Upstream publishes each package only when its own output changes, so the four
packages carry different beta numbers while sharing one commit — at the time of
writing `@xterm/xterm@6.1.0-beta.303` and `@xterm/headless@6.1.0-beta.302` are
both built from `d3e32b3`. Match on `package.json.commit`, never on the version
string; `xterm-user-scrolling-contract.test.ts` asserts that pairing for
headless and core.
Bumping `@xterm/xterm` is:
1. Update the version in `package.json` and run `pnpm install`.
2. Rename both patch files to the new version and update `patch`,
`sourcePatch`, and `version` in `xterm-upstream.json`.
3. Update `upstream.commit` to the `commit` field of the new tarball's
`package.json`, and `toolchain` to whatever the new `package-lock.json`
resolves.
4. `node config/scripts/regenerate-xterm-patches.mjs --write`.
5. `pnpm install`, then `--check`. On a bump the lockfile has no entry under the
new key yet, so `--write` reports the gap and leaves the hash to `pnpm
install`; `--check` is what proves the two agree afterwards.
Step 4 is where a real upstream conflict shows up: `git apply` of the source
patch fails against the new tree. Resolve it in the checkout, re-diff, and
rerun. The bundle hunks need no attention at any point.
## Why Not Vendor a Fork
A vendored `@xterm/xterm` fork removes the patch entirely, but it moves Orca off
the published package, so every upstream beta becomes a merge rather than a
version bump, and Orca inherits responsibility for building and publishing a
package it does not own. The patch is four small source hunks against a commit
that reproduces byte for byte; a fork is a much larger standing cost for the
same result.
## Why Not Handle Composition at Runtime
`CompositionHelper` hooks four private call sites upstream of `onData`, and
`SortedList` has no public surface at all. There is no supported extension point
that reaches either, so a runtime shim would mean reaching into `_core`
internals that upstream renames freely between betas. The patch is the smaller
risk.
## CI Contract
`xterm_patch_sync` in `.github/workflows/pr.yml` runs
`regenerate-xterm-patches.mjs --check` on every PR and is part of the `verify`
aggregate. It clones the pinned commit, installs upstream's toolchain, builds
twice, and byte-compares the result against the committed patch. Both builds and
the diff together are about eight seconds; `npm ci` for upstream's toolchain is
what the job actually spends its minutes on, and the cache key is the manifest.
`config/scripts/regenerate-xterm-patches.test.mjs` covers the pure pieces —
pnpm's diff flags and normalization, hunk splitting, round-trip stability, the
commit and build-order assertions, and lockfile coupling — with no network and
no build, so they run in the ordinary test shards.
## Known Gaps
`@xterm/addon-ligatures` is still patched by hand, and can stay that way: the
patch is a fifteen-line `package.json` edit that repoints `module` and adds an
`exports` block, touching no bundle and no sourcemap. Nothing about it is
generated, so there is nothing for this harness to verify.
The addons were folded into this manifest on 2026-08-29. Before that they were
hand-edited minified bundles carrying a literal `/* PATCH(orca): ... */` comment
inside minified code, parser round-trip artifacts (`!0` printed back as `true`,
locals renamed `i` → `i5`), and no `.map` hunks at all — so both shipped
sourcemaps whose offsets did not match the bundle beside them. All four
`@xterm/addon-webgl` artifacts and all four `@xterm/addon-serialize` artifacts
now reproduce byte for byte from the pinned commit, which is what closed it.
The one thing still unproven is that this holds across upstream revisions rather
than at this commit. `addon-serialize.js.map` did not reproduce on the first
attempt here; the cause was a stale `out/` from a wrong build order, not
upstream nondeterminism, and it reproduced exactly once the root build ran
first. Treat a future non-reproducing artifact as a build-order bug until proven
otherwise.
@@ -57,6 +57,8 @@ async function host(remote: boolean, launchToken?: string) {
method: 'POST',
headers: {
'Content-Type': 'application/json',
// Advancing fake time can expire a pooled HTTP socket before the next post.
Connection: 'close',
'X-Orca-Agent-Hook-Token': relay.getCoordinates().token
},
body: JSON.stringify(body)
@@ -8,9 +8,6 @@
*
* One case is pinned as a KNOWN DEFECT: the shipped detector refuses a ready screen whose retained
* tail ends on the error block. That asserts what it does, not what it should.
*
* Capture protocol: docs/reference/agent-pty-transcript-capture.md
* What each transcript decides: docs/reference/antigravity-readiness-evidence.md
*/
import { existsSync, readFileSync } from 'node:fs'
import { join } from 'node:path'
@@ -26,15 +23,6 @@ vi.mock('electron', () => ({
}))
const FIXTURE_DIR = join(__dirname, '__fixtures__')
const EVIDENCE_DOC = join(
__dirname,
'..',
'..',
'..',
'docs',
'reference',
'antigravity-readiness-evidence.md'
)
// Why asymmetric: a ready verdict has to survive the settle window, while a refusal only has to
// hold for one poll. Keeping the refusal short keeps seven transcripts off the suite's clock.
const READY_TIMEOUT_MS = 2_000
@@ -48,7 +36,7 @@ const ESC = String.fromCharCode(27)
type TranscriptCase = {
/** Fixture basename; `<name>.txt` under `__fixtures__/`. */
name: string
/** Capture in docs/reference/antigravity-readiness-evidence.md. */
/** Capture group identifier. */
capture: string
what: string
/** What a correct detector must answer. Not what the shipped one answers. */
@@ -114,8 +102,7 @@ const TRANSCRIPTS: readonly TranscriptCase[] = [
expectReady: true
},
// Not captured: this machine's agy has no OAuth session and offers only Gemini models, and
// reaching the rest would mean signing the operator out or deleting their config. See
// docs/reference/antigravity-readiness-evidence.md § What could not be captured.
// reaching the rest would mean signing the operator out or deleting their config.
{
name: 'antigravity-ready-business-non-gemini',
capture: 'A',
@@ -225,13 +212,4 @@ describe('Antigravity readiness, decided by captured transcripts', () => {
expect(text).toContain(ESC)
})
}
it('documents every transcript the detector is allowed to depend on', () => {
// Why a test: the doc is the operator's checklist. A name that drifts out of it is a
// transcript nobody will capture, and a case that silently skips forever.
const doc = readFileSync(EVIDENCE_DOC, 'utf8')
for (const transcript of TRANSCRIPTS) {
expect(doc).toContain(`${transcript.name}.txt`)
}
})
})
@@ -1,30 +1,17 @@
import { readFileSync } from 'node:fs'
import { join } from 'node:path'
import { describe, expect, it } from 'vitest'
import { SINGLE_INSTANCE_ALREADY_RUNNING_EXIT_CODE } from './single-instance-lock'
// Why #11935: a pre-`ready` graceful quit is deferred, so a lock-losing headless `orca serve`
// kept booting into Linux Ozone/X11 init, died with SIGSEGV, and systemd restarted it forever
// until the leaked AppImage FUSE mounts hit the kernel's 1000-mount ceiling.
function readSystemdUnitBlocks(doc: string): Map<string, string[]> {
const blocks = new Map<string, string[]>()
// Why: key on the unit's path comment — splitting on directives mixes `[Unit]` and `[Service]` across blocks.
for (const match of doc.matchAll(/^# \/etc\/systemd\/system\/(\S+\.service)$/gm)) {
const start = match.index + match[0].length
const name = match[1]
blocks.set(name, [...(blocks.get(name) ?? []), doc.slice(start, doc.indexOf('```', start))])
}
return blocks
}
describe('headless lock-loss exit contract', () => {
const preflightSource = readFileSync(
join(process.cwd(), 'src/main/startup/main-process-preflight.ts'),
'utf8'
)
const entrySource = readFileSync(join(process.cwd(), 'src/main/index.ts'), 'utf8')
const doc = readFileSync(join(process.cwd(), 'docs/reference/headless-linux-server.md'), 'utf8')
it('exits the lock-losing launch immediately instead of scheduling a graceful quit', () => {
const gateStart = preflightSource.indexOf('if (!hasLock) {')
@@ -48,41 +35,4 @@ describe('headless lock-loss exit contract', () => {
'shouldActivateDesktopForSecondInstance(argv)'
)
})
it('makes every documented serve unit treat a duplicate owner as terminal', () => {
const serveUnits = readSystemdUnitBlocks(doc).get('orca-serve.service') ?? []
expect(serveUnits.length).toBeGreaterThan(0)
for (const unit of serveUnits) {
expect(unit).toContain(
`RestartPreventExitStatus=${SINGLE_INSTANCE_ALREADY_RUNNING_EXIT_CODE}`
)
expect(unit).toContain('StartLimitIntervalSec=')
expect(unit).toContain('StartLimitBurst=')
}
})
it('clears the start limit before every scripted start, which a tripped burst would refuse', () => {
const lines = doc.split('\n')
const startLines = lines.flatMap((line, index) =>
/^\s*sudo systemctl start orca-serve/.test(line) ? [index] : []
)
expect(startLines.length).toBeGreaterThan(0)
for (const index of startLines) {
expect(lines.slice(Math.max(0, index - 3), index).join('\n')).toContain(
'systemctl reset-failed orca-serve'
)
}
})
it('leaves the Xvfb unit free to self-heal from a transient display flap', () => {
const xvfbUnits = readSystemdUnitBlocks(doc).get('orca-xvfb.service') ?? []
expect(xvfbUnits.length).toBeGreaterThan(0)
// Why: a start limit here would down the display unit permanently and take orca-serve with it.
for (const unit of xvfbUnits) {
expect(unit).not.toContain('StartLimitBurst=')
}
})
})
+1 -1
View File
@@ -36,7 +36,7 @@ Keep `ORCA_BACKGROUND_LAUNCH=1`: the application must still suppress automatic r
`isWindowlessLaunch` describes that automatic launch policy, not the window's current visibility.
This exception belongs only to this benchmark fixture; do not generalize it to local or self-hosted
runs, paired-client helpers, native-focus tests, or production window policy. Background terminal
panes remain hidden. Evidence: `docs/reference/terminal-perf-latency-investigation.md`.
panes remain hidden.
## Isolated native IBus presentation