- AI-оркестратор для розробників рівня 100x.
- Запускайте Codex, Claude Code, OpenCode або Pi паралельно — кожен у власному worktree, усі під контролем в одному місці.
-
-
-### Супутній мобільний застосунок
-
-Стежте за агентами та керуйте ними з телефону — отримуйте сповіщення про завершення роботи агента та надсилайте подальші вказівки, де б ви не були.
-
-[App Store для iOS](https://apps.apple.com/us/app/orca-ide/id6766130217) · [TestFlight](https://testflight.apple.com/join/YjeGMQBA) · [Android APK 0.0.44](https://github.com/stablyai/orca/releases/download/mobile-android-v0.0.44/app-release.apk) · [Документація →](https://www.onorca.dev/docs/mobile)
-
-
-
-
-
-
-
-
-
-### Паралельні worktree
-
-Надішліть один промпт одразу п’ятьом агентам, кожен із яких працюватиме у власному ізольованому git worktree, — порівняйте результати та виконайте злиття найкращого з них.
-
-[Документація →](https://www.onorca.dev/docs/model/worktrees)
-
-
-
-
-
-
-
-
-
-### Розділені термінали
-
-Термінали рівня Ghostty з рендерингом на WebGL, необмеженою кількістю розділень і буфером прокручування, який зберігається після перезапуску.
-
-[Документація →](https://www.onorca.dev/docs/terminal)
-
-
-
-
-
-
-
-
-
-### Режим дизайну
-
-Клацніть на будь-якому елементі інтерфейсу у справжньому вікні Chromium, щоб надіслати його HTML, CSS і обрізаний скриншот прямо в промпт агента.
-
-[Документація →](https://www.onorca.dev/docs/browser/design-mode)
-
-
-
-
-
-
-
-
-
-### GitHub і Linear, нативно
-
-Переглядайте PR, issue та дошки проєктів прямо в застосунку — відкривайте worktree з будь-якої задачі та рев'юйте без перемикання контексту.
-
-[Документація →](https://www.onorca.dev/docs/review/linear)
-
-
-
-
-
-
-
-
-
-### SSH worktree
-
-Запускайте агентів на потужній віддаленій машині з повноцінним редагуванням файлів, git і терміналами — з автоперепідключенням і прокиданням портів.
-
-[Документація →](https://www.onorca.dev/docs/ssh)
-
-
-
-
-
-
-
-
-
-### Анотуйте diff-и агентів
-
-Залишайте коментарі на будь-якому рядку diff-у й надсилайте їх агенту — рев'юйте, редагуйте та комітьте, не виходячи з Orca.
-
-[Документація →](https://www.onorca.dev/docs/review/annotate-ai-diff)
-
-
-
-
-
-
-
-
-
-### Перетягуйте файли агентам
-
-Редактор на базі VS Code з автозбереженням усюди — перетягуйте файли чи зображення прямо в промпт агента.
-
-[Документація →](https://www.onorca.dev/docs/editing/file-explorer)
-
-
-
-
-
-
-
-
-
-### Orca CLI
-
-Агенти теж керують Orca — автоматизуйте будь-який робочий процес командами `orca worktree create`, `snapshot`, `click` і `fill`.
-
-[Документація →](https://www.onorca.dev/docs/cli/overview)
-
-
-
-
-
-
-
-
-**Також у комплекті:**
-
-- **[Швидкий пошук](https://www.onorca.dev/docs/model/quick-open)** — Шукайте серед worktree, файлів, агентів, команд і контексту репозиторію, не відриваючись від роботи.
-- **[Перемикач акаунтів і відстеження використання](https://www.onorca.dev/docs/agents/usage-tracking)** — Стежте за використанням Claude і Codex та скиданням лімітів, перемикайте акаунти на льоту без повторного входу.
-- **[Розширені перегляди репозиторію](https://www.onorca.dev/docs/editing/markdown)** — Переглядайте Markdown, зображення, PDF та документацію репозиторію прямо в робочому просторі.
-- **[Computer Use](https://www.onorca.dev/docs/cli/computer-use)** — Дозвольте агентам керувати десктопними застосунками та видимим інтерфейсом, коли робочий процес потребує реальної взаємодії.
-- **[Сповіщення та статус непрочитаного](https://www.onorca.dev/docs/notifications)** — Дізнавайтеся, коли агент завершив роботу або потребує уваги, і позначайте треди як непрочитані, щоб повернутися пізніше.
-- **І багато іншого** — ми випускаємо оновлення щодня, тож цей список завжди відстає. Справжній перелік можливостей — це [changelog](https://github.com/stablyai/orca/releases).
-
----
-
-## Підтримувані агенти
-
-Працює з **будь-яким CLI-агентом** — якщо він запускається в терміналі, він запуститься і в Orca.
-
-
-
----
-
-## Встановлення
-
-### Десктоп — macOS, Windows, Linux
-
-- **[Завантажити з onOrca.dev](https://onorca.dev/download)**
-- Або завантажте білд напряму: [macOS Apple Silicon](https://github.com/stablyai/orca/releases/latest/download/orca-macos-arm64.dmg) · [macOS Intel](https://github.com/stablyai/orca/releases/latest/download/orca-macos-x64.dmg) · [Windows (.exe)](https://github.com/stablyai/orca/releases/latest/download/orca-windows-setup.exe) · [Linux AppImage](https://github.com/stablyai/orca/releases/latest/download/orca-linux.AppImage) · [Усі білди](https://github.com/stablyai/orca/releases/latest)
-- Запускаєте `orca serve` на headless Linux-сервері? Дивіться [посібник із headless Linux-сервера](../reference/headless-linux-server.md).
-
-_Або через пакетний менеджер:_
-
-```bash
-# macOS (Homebrew)
-brew install --cask stablyai/orca/orca
-
-# Arch Linux (AUR) — або stably-orca-git для збірки з джерела
-yay -S stably-orca-bin
-```
-
-### Супутній мобільний застосунок — iOS, Android
-
-Під’єднайте мобільний застосунок до десктопного, щоб стежити за агентами та керувати ними з телефону.
-
-- **iOS:** [Завантажити з App Store](https://apps.apple.com/us/app/orca-ide/id6766130217) або [приєднатися до TestFlight](https://testflight.apple.com/join/YjeGMQBA)
-- **Android:** [Завантажити APK 0.0.44](https://github.com/stablyai/orca/releases/download/mobile-android-v0.0.44/app-release.apk) · [Інструкція зі встановлення](https://www.onorca.dev/docs/android-apk)
-
----
-
-## Спільнота та підтримка
-
-- **Discord:** Приєднуйтеся до спільноти в **[Discord](https://discord.gg/fzjDKHxv8Q)**.
-- **Twitter / X:** Стежте за **[@orca_build](https://x.com/orca_build)**, щоб бути в курсі оновлень і анонсів.
-- **WeChat:** Відскануйте QR-код, щоб приєднатися до групи № 7 спільноти Orca у WeChat. Якщо вона заповнена, приєднайтеся до групи № 8.
-
-
-
-
-- **Зворотний зв'язок та ідеї:** Ми випускаємо оновлення швидко. Чогось бракує? [Запропонуйте нову функцію](https://github.com/stablyai/orca/issues).
-- **Конфіденційність:** Перегляньте [документацію про конфіденційність і телеметрію](https://www.onorca.dev/docs/telemetry), щоб дізнатися, які анонімні дані про використання збирає Orca і як від цього відмовитися.
-- **Підтримайте нас:** Поставте [зірку](https://github.com/stablyai/orca) цьому репозиторію, щоб стежити за нашими щоденними релізами.
-
----
-
-## Розробка
-
-Хочете зробити внесок або запустити проєкт локально? Перегляньте наш посібник [CONTRIBUTING.md](../../.github/CONTRIBUTING.md).
-
-
-
-
-
-
-
-
-
-## Підписані білди
-Підписання коду для Windows надано за підтримки [SignPath.io](https://signpath.io), сертифікат надано [SignPath Foundation](https://signpath.org).
-
-## Ліцензія
-
-Orca — безкоштовний проєкт із відкритим кодом за ліцензією [MIT](../../LICENSE).
diff --git a/docs/readme/README.zh-CN.md b/docs/readme/README.zh-CN.md
index 160f5b05611..d7bae3fba9e 100644
--- a/docs/readme/README.zh-CN.md
+++ b/docs/readme/README.zh-CN.md
@@ -12,7 +12,7 @@
From f40e94d844d9c1f243c210a55b52394d2b492b4c Mon Sep 17 00:00:00 2001
From: OrcaWin
Date: Thu, 3 Sep 2026 21:37:52 -0700
Subject: [PATCH 16/49] Revert "docs: document localization workflow" (#18571)
This reverts commit 912463c2780237d7f51b254e50dd271684302631.
---
AGENTS.md | 12 ------------
1 file changed, 12 deletions(-)
diff --git a/AGENTS.md b/AGENTS.md
index 6915b246bcc..8b0156ba6b1 100644
--- a/AGENTS.md
+++ b/AGENTS.md
@@ -53,18 +53,6 @@ Orca targets macOS, Linux, and Windows. Keep all platform-dependent behavior beh
- **WSL commands**: build argv with `buildWslExecArgs` (always `--exec` — under `--`, `wsl.exe` expands `$name` in every argument and silently rewrites the script), and fence anything whose stdout you parse with `buildWslCapturedLoginShellCommand`, because the interactive login shell prints the distro banner to stdout. See [`docs/reference/wsl-command-execution.md`](./docs/reference/wsl-command-execution.md).
- **Linux native modules**: keep the glibc floor at Ubuntu 20.04 / glibc 2.31. A module compiled from source on a newer runner can reference symbol versions absent on the floor and crash the app on startup. See [`docs/reference/linux-glibc-compatibility.md`](./docs/reference/linux-glibc-compatibility.md); packaging fails if a bundled native binary needs newer glibc.
-## Localization (i18n)
-
-All user-facing copy is localized. `src/renderer/src/i18n/locales/en.json` is the source of truth; the `zh`, `ja`, `ko`, and `es` catalogs mirror its keys. Strings reach the UI through `translate('auto.', 'English fallback')` — never hardcode display text.
-
-When you touch user-facing copy, keep all five catalogs in sync:
-
-- **New strings** — wrap them in `translate(...)` with an English fallback, run `pnpm sync:localization-catalog` to register the keys in `en.json` and add placeholders to every other locale, then `pnpm bootstrap:-catalog` (e.g. `bootstrap:ja-catalog`) to translate the placeholders.
-- **Reworded strings** — changing the value of an existing key updates only `en.json`. The other locales keep the key with its old translation, and **the lint checks will not catch this**: `verify:localization-catalog` enforces key _parity_, not translation _freshness_. Update the same key in `zh/ja/ko/es` by hand, or re-translate it via `bootstrap:-catalog`.
-- **Removed strings** — delete the key from _every_ locale; the parity check rejects a key that exists in one catalog but not another.
-
-Before pushing copy changes, run `pnpm verify:localization-catalog` and `pnpm verify:localization-coverage` (both also run in `pnpm lint`).
-
## SSH Use Case
All changes must consider the SSH use case. Don't assume local-only execution. Before changing anything that reports on, stops, or lists remote work, follow [`docs/reference/ssh-execution-boundary.md`](./docs/reference/ssh-execution-boundary.md): the execution host owns everything that touches execution, and loss of contact is never evidence of process death — the verdict vocabulary is `live` / `unverifiable` / `exited`, with no synonyms.
From 36354f174224836b83e802e3ecaa25e44d9df766 Mon Sep 17 00:00:00 2001
From: Neil <4138956+nwparker@users.noreply.github.com>
Date: Thu, 3 Sep 2026 21:38:03 -0700
Subject: [PATCH 17/49] perf(remote): read the repo catalog once per publish,
not once per worktree (#18410)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
* perf(remote): read the repo catalog once per publish, not once per worktree
`remoteWorkspace:setForConnectedTargets` costs 13 ms of main-thread time per
call at 0.48 calls/sec — 0.63% of wall on a real session, the second most
expensive IPC handler in the main process.
Almost all of it is one line. `exportRemoteWorkspaceSession` asks
`isTargetWorktree(worktreeId)` once per worktree in the session, and that
callback called `targetForWorktree(store, ...)`, which called
`store.getRepos()` — and `getRepos()` maps `hydrateRepo` over every repo row.
So publishing to one SSH target re-hydrated the whole repo catalog once per
worktree, then threw a fresh `createRepoRowExecutionHostLookup` (which itself
`filter`s the catalog per lookup) away each time.
The lookup is now built once per handler invocation and shared across targets:
the rows cannot change inside one synchronous projection, and they are the same
for every target.
On the session that surfaced this — 413 worktrees, 13 repos, 1 connected target
— that is 413 catalog hydrations (5369 `hydrateRepo` calls) per publish reduced
to 1 (13 calls). The repo normaliser reached through `hydrateRepo` was the #2
self-time function in a 30 s main-process CPU profile at 0.25%.
No user-facing trade-off: identical ownership resolution, identical exported
session, identical stale-revision handling.
* perf(remote): resolve each worktree's owning target once per publish
Follow-up on the same handler: hoisting `store.getRepos()` removed the repeat
hydration, but the ownership resolution itself was still repeated once per
connected target.
`targetForWorktree` computes a connection id from the repo catalog alone — only
the final `=== targetId` differs — so exporting to N targets ran the identical
resolution N times over every worktree key, and the projection asks the question
once per key of `tabsByWorktree`, `activeTabIdByWorktree`,
`lastVisitedAtByWorktreeId` and `defaultTerminalTabsAppliedByWorktreeId`.
Resolutions are now memoised for the life of one publish, keyed on
`(worktreeId, executionHostId)` because both participate in resolution.
Test asserts 6 worktree keys resolve 6 times across 2 targets instead of 12.
* perf(remote): skip the session and repo reads when no hydrated target is connected
Hoisting the catalog read made a zero-connected-target publish pay for a full
repo hydration it never did before. Return early instead.
---
src/main/ipc/remote-workspace.test.ts | 98 ++++++++++++++++++++++++++-
src/main/ipc/remote-workspace.ts | 54 ++++++++++++---
2 files changed, 141 insertions(+), 11 deletions(-)
diff --git a/src/main/ipc/remote-workspace.test.ts b/src/main/ipc/remote-workspace.test.ts
index 56eb4804a82..3b900e175bc 100644
--- a/src/main/ipc/remote-workspace.test.ts
+++ b/src/main/ipc/remote-workspace.test.ts
@@ -7,18 +7,35 @@ import type {
RemoteWorkspaceSnapshot
} from '../../shared/remote-workspace-types'
import type { SshTarget } from '../../shared/ssh-types'
+import type * as WorktreeExecutionHostResolution from '../../shared/worktree-execution-host-resolution'
import type { WorkspaceSessionState } from '../../shared/workspace-session-state-types'
const {
getActiveMultiplexerMock,
getSshConnectionStoreMock,
- registerRemoteWorkspaceNotificationHandlerMock
+ registerRemoteWorkspaceNotificationHandlerMock,
+ resolveWorktreeExecutionHostCalls
} = vi.hoisted(() => ({
getActiveMultiplexerMock: vi.fn(),
getSshConnectionStoreMock: vi.fn(),
- registerRemoteWorkspaceNotificationHandlerMock: vi.fn(() => vi.fn())
+ registerRemoteWorkspaceNotificationHandlerMock: vi.fn(() => vi.fn()),
+ resolveWorktreeExecutionHostCalls: { count: 0 }
}))
+// Counts ownership resolutions without changing any of them.
+vi.mock('../../shared/worktree-execution-host-resolution', async (importOriginal) => {
+ const actual = (await importOriginal()) as typeof WorktreeExecutionHostResolution
+ return {
+ ...actual,
+ resolveWorktreeExecutionHost: (
+ ...args: Parameters
+ ) => {
+ resolveWorktreeExecutionHostCalls.count += 1
+ return actual.resolveWorktreeExecutionHost(...args)
+ }
+ }
+})
+
vi.mock('electron', () => ({
ipcMain: {
handle: vi.fn(),
@@ -154,9 +171,10 @@ describe('remoteWorkspace:setForConnectedTargets', () => {
const getRepoMock = vi.fn()
const getWorkspaceSessionMock = vi.fn()
// Ownership resolution reads the catalog, not one id-keyed row, so the fake has to project one.
+ const getReposMock = vi.fn(() => [getRepoMock('repo-target-1')].filter(Boolean))
const store = {
getRepo: getRepoMock,
- getRepos: () => [getRepoMock('repo-target-1')].filter(Boolean),
+ getRepos: getReposMock,
getWorkspaceSession: getWorkspaceSessionMock
} as unknown as Store
@@ -176,6 +194,7 @@ describe('remoteWorkspace:setForConnectedTargets', () => {
getTarget: (targetId: string) => targets.find((target) => target.id === targetId)
})
getRepoMock.mockReset()
+ getReposMock.mockClear()
getWorkspaceSessionMock.mockReset()
getWorkspaceSessionMock.mockReturnValue(baseSession)
getRepoMock.mockImplementation((repoId: string) =>
@@ -251,6 +270,79 @@ describe('remoteWorkspace:setForConnectedTargets', () => {
return observed as RemoteWorkspaceObservedSnapshot
}
+ it('reads the repo catalog once per publish, not once per worktree', async () => {
+ // `store.getRepos()` re-hydrates every repo row. The export asks "is this worktree mine?" once
+ // per worktree, so reading the catalog inside that callback multiplied hydration by the
+ // worktree count — 413 on the session that surfaced this.
+ const worktrees = Object.fromEntries(
+ Array.from({ length: 12 }, (_, index) => [`repo-target-1::/remote/repo-${index}`, []])
+ )
+ getWorkspaceSessionMock.mockReturnValue({
+ ...baseSession,
+ tabsByWorktree: worktrees
+ } as WorkspaceSessionState)
+ const observed = await observeTarget('target-1')
+ getReposMock.mockClear()
+
+ await callSetForConnectedTargets({
+ hydratedTargetIds: ['target-1'],
+ expectedRevisionsByTargetId: { 'target-1': observed.revision },
+ expectedHostObservationTokensByTargetId: {
+ 'target-1': observed.hostObservationToken
+ }
+ })
+
+ expect(getReposMock).toHaveBeenCalledTimes(1)
+ })
+
+ it('resolves each worktree ownership once for the whole publish, not once per target', async () => {
+ // Ownership is a function of the repo catalog alone; only the final `=== targetId` differs, so
+ // exporting to N targets used to repeat the identical resolution N times per worktree key.
+ const worktrees = Object.fromEntries(
+ Array.from({ length: 6 }, (_, index) => [`repo-target-1::/remote/repo-${index}`, []])
+ )
+ getWorkspaceSessionMock.mockReturnValue({
+ ...baseSession,
+ tabsByWorktree: worktrees
+ } as WorkspaceSessionState)
+ const observed = await Promise.all(targets.map((target) => observeTarget(target.id)))
+ getReposMock.mockClear()
+ resolveWorktreeExecutionHostCalls.count = 0
+
+ await callSetForConnectedTargets({
+ hydratedTargetIds: targets.map((target) => target.id),
+ expectedRevisionsByTargetId: Object.fromEntries(
+ targets.map((target, index) => [target.id, observed[index].revision])
+ ),
+ expectedHostObservationTokensByTargetId: Object.fromEntries(
+ targets.map((target, index) => [target.id, observed[index].hostObservationToken])
+ )
+ })
+
+ expect(getReposMock).toHaveBeenCalledTimes(1)
+ // 6 worktree keys resolved once each, regardless of how many targets are published to.
+ expect(resolveWorktreeExecutionHostCalls.count).toBe(6)
+ })
+
+ it('skips the session and repo-catalog reads when no hydrated target is connected', async () => {
+ // A hydrated but disconnected target leaves nothing to project onto, so hoisting the catalog
+ // read must not make the idle path pay for a full repo hydration it never used before.
+ getActiveMultiplexerMock.mockReturnValue(undefined)
+ getReposMock.mockClear()
+ getWorkspaceSessionMock.mockClear()
+
+ await expect(
+ callSetForConnectedTargets({
+ hydratedTargetIds: ['target-1'],
+ expectedRevisionsByTargetId: { 'target-1': 7 },
+ expectedHostObservationTokensByTargetId: { 'target-1': 'token' }
+ })
+ ).resolves.toEqual([])
+
+ expect(getReposMock).not.toHaveBeenCalled()
+ expect(getWorkspaceSessionMock).not.toHaveBeenCalled()
+ })
+
it('does not write without an explicit non-empty hydrated target set', async () => {
await expect(callSetForConnectedTargets({ session: baseSession })).resolves.toEqual([])
await expect(
diff --git a/src/main/ipc/remote-workspace.ts b/src/main/ipc/remote-workspace.ts
index f9479f15ce9..935fd1c9f72 100644
--- a/src/main/ipc/remote-workspace.ts
+++ b/src/main/ipc/remote-workspace.ts
@@ -1,5 +1,6 @@
import { ipcMain, type BrowserWindow } from 'electron'
import type { Store } from '../persistence'
+import type { Repo } from '../../shared/repo-types'
import { getActiveMultiplexer, getSshConnectionStore } from './ssh'
import { exportRemoteWorkspaceSession } from '../../shared/remote-workspace-session-projection'
import {
@@ -107,7 +108,7 @@ function getExpectedHostObservationTokens(
}
function targetForWorktree(
- store: Store,
+ repoLookup: ReturnType>,
worktreeId: string,
executionHostId?: string
): string | null {
@@ -115,21 +116,49 @@ function targetForWorktree(
// `getRepo(id)?.connectionId`, which is host-blind — the same repo id can name rows on several
// hosts, so a session could be published to a machine that never owned the worktree (#11163).
// Unresolvable ownership exports to nobody rather than guessing.
- const resolution = resolveWorktreeExecutionHost(
- createRepoRowExecutionHostLookup(store.getRepos()),
- { repoId: getRepoIdFromWorktreeId(worktreeId), hostId: executionHostId ?? null }
- )
+ const resolution = resolveWorktreeExecutionHost(repoLookup, {
+ repoId: getRepoIdFromWorktreeId(worktreeId),
+ hostId: executionHostId ?? null
+ })
return resolution.kind === 'resolved' ? resolution.connectionId : null
}
+/**
+ * Resolve each worktree's owning connection at most once for a whole publish.
+ *
+ * Why this is shared and not per target: `targetForWorktree` computes a connection id from the
+ * repo catalog alone — only the final `=== targetId` differs — so exporting to N targets used to
+ * repeat the identical resolution N times over every worktree key. `store.getRepos()` also
+ * re-hydrates every repo row on each call, and the projection asks this question once per key of
+ * `tabsByWorktree`, `activeTabIdByWorktree`, `lastVisitedAtByWorktreeId` and
+ * `defaultTerminalTabsAppliedByWorktreeId`.
+ */
+function createWorktreeTargetResolver(
+ repoLookup: ReturnType>
+): (worktreeId: string, executionHostId?: string) => string | null {
+ const resolved = new Map()
+ return (worktreeId, executionHostId) => {
+ // Host id participates in resolution, so it has to participate in the key. NUL cannot appear
+ // in either id, so it is a collision-free separator.
+ const key = `${worktreeId}\u0000${executionHostId ?? ''}`
+ const cached = resolved.get(key)
+ if (cached !== undefined) {
+ return cached
+ }
+ const connectionId = targetForWorktree(repoLookup, worktreeId, executionHostId)
+ resolved.set(key, connectionId)
+ return connectionId
+ }
+}
+
function exportSessionForTarget(
- store: Store,
+ resolveWorktreeTarget: (worktreeId: string, executionHostId?: string) => string | null,
targetId: string,
session: WorkspaceSessionState
): RemoteWorkspaceSession {
return exportRemoteWorkspaceSession(session, {
isTargetWorktree: (worktreeId, executionHostId) =>
- targetForWorktree(store, worktreeId, executionHostId) === targetId
+ resolveWorktreeTarget(worktreeId, executionHostId) === targetId
})
}
@@ -245,12 +274,21 @@ export function registerRemoteWorkspaceHandlers(
(target) => hydratedTargetIds.has(target.id) && getActiveMultiplexer(target.id)
) ?? []
+ if (targets.length === 0) {
+ // Nothing to project onto, so skip the session and repo-catalog reads entirely.
+ return []
+ }
+
const workspaceSession = args.session ?? store.getWorkspaceSession()
+ // One repo read, and ownership resolutions shared across targets: neither depends on the target.
+ const resolveWorktreeTarget = createWorktreeTargetResolver(
+ createRepoRowExecutionHostLookup(store.getRepos())
+ )
const results = await Promise.all(
targets.map(async (target) => {
// Why: each target has its own revision stream. Keep same-target
// writes queued, but do not let one slow relay block others.
- const session = exportSessionForTarget(store, target.id, workspaceSession)
+ const session = exportSessionForTarget(resolveWorktreeTarget, target.id, workspaceSession)
const result = await queueRemoteWorkspacePatch(target.id, async () => {
const current =
getCachedRemoteWorkspaceSnapshot(target.id) ?? (await getRemoteSnapshot(target))
From f36c03e84a66fc188bd1846935c245762f313a9a Mon Sep 17 00:00:00 2001
From: Neil <4138956+nwparker@users.noreply.github.com>
Date: Thu, 3 Sep 2026 21:39:34 -0700
Subject: [PATCH 18/49] fix(windows): make the install-dir ACL repair rescue
the launch it runs in (#18361)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
* fix(windows): repair the poisoned install-dir ACL before the window, not after
The install-dir LPAC ACL poison (electron/electron#51761) still costs every
affected machine at least one crash: the probe that detects it is
setImmediate-deferred and answers 0.9-3.0s in, while createMainWindow runs
synchronously in the same frame and its renderer dies at init 48-1373ms later.
- Persist the poison verdict the moment the probe reports it, and await the
repair (bounded at 20s) before any window is created on a launch that already
carries the marker.
- Do not engage the GPU safe-graphics fallback while the install-dir ACL verdict
is poisoned or still outstanding. Safe graphics does not rescue a poisoned
tree, and --in-process-gpu removes the GPU child, erasing the sibling-death
evidence that identifies the shape (4 field reports landed in 'misc' this way).
- Clear the safe-graphics marker once the repair lands, so a repaired machine
stops launching software-rendered for the rest of that build.
- Give the repair marker a bounded retry budget: it was written on failure and
matched regardless of outcome, so one transient failure pinned a machine to
'marker-hit' for the life of that version.
* test(windows): pin the install-dir ACL repair against the real icacls binary
* fix(windows): stop the install-DACL verdict from outliving the evidence
Adversarial review round 1. Five blocking findings, all addressed.
1. gpu-lifecycle guard had only a source grep (green with the polarity
inverted). The stated justification -- that gpu-lifecycle's import graph
cannot be driven in-process -- was wrong: mocking `electron` plus
`@electron-toolkit/utils` imports it fine. Replaced with
gpu-lifecycle-install-dir-acl-guard.test.ts, which drives the real
handleGpuChildCrash against a stub tracker. All four cases go red when the
guard is flipped to `if (!isInstallDirAclSuspect())`.
2. A clean probe verdict retired the on-disk marker but not the in-memory
`poison` verdict, so a machine the probe just proved healthy kept
suppressing the GPU safe-graphics fallback and kept the dialog accusing the
install folder -- permanently, since a `status:'failed'` probe deliberately
keeps the marker. A positive clean reading now latches `installDirReadClean`,
drops the verdict, and outranks a repair result that lands after it (a
'failed' from a repair with nothing left to fix must not re-accuse).
'repaired' is kept: it is not a contradiction and it is what tells the user
to reload.
3. `noteWindowsInstallDirAclProbePending()` ran on every `openMainWindow` while
the probe is once-per-process, so every tray/second-instance reopen armed a
15s window in which `recordGpuCrash` was never called at all -- on healthy
machines. `probeWindowsInstallDirAcl` now reports whether THIS call
dispatched, and only a dispatch arms the grace window.
4. The pre-window ordering guarantee was defeatable and untested.
`focusExistingMainWindow` opens a window whenever there is none and the app
is ready -- true for the whole 20s gate, which is exactly when a user
double-clicks the shortcut again. Added a `canOpenWindow` seam (same
'pending' semantics as the existing `!app.isReady()` case) wired to
`isBlockingInstallDirAclRepairInFlight()`, plus
windows-install-dir-acl-startup-wiring.test.ts pinning the await ahead of
both window-creation paths and both new call sites.
5. windows-install-dir-acl-repair.win32.test.ts was absent from the pr.yml
win32 allowlist, so it ran nowhere. Added.
Also from the non-blocking list:
- The repair no longer clears a `userConfirmed: true` safe-graphics marker;
"keep safe graphics" is a user choice, not Orca's automatic latch.
- `repairWindowsInstallDirPackageAcl` now reports its dispatch too, so a second
entry into the gate resolves immediately instead of eating the full 20s
budget waiting on an `onDone` that is never coming.
- The gate is wrapped in try/catch/finally, matching the contract the probe
documents as mandatory for anything upstream of window creation.
Rebutted, not applied:
- "Gate should be conditioned on app.isPackaged." A dev launch only carries the
poison marker if a dev launch actually probed that tree and found the
signature, in which case the dev renderer is dying the same way and the
repair is exactly what is needed. The adjacent `isPackaged` check guards a
packaged-only early-window optimisation, not a correctness boundary.
- "Fold the poison marker into the repair marker's `outcome`." They answer
different questions with different lifetimes. The repair marker is a retry
budget (`attempts >= 3` disables the repair for that version) and is never
cleared; the poison marker is cleared by a successful repair and by a clean
probe. A `'pending'` outcome written before the attempt would bump `attempts`,
so three launches killed mid-repair would permanently disable a repair that
never once ran icacls to completion.
* fix(windows): keep counting GPU crashes while the install-DACL verdict is pending
Adversarial review round 2. Both blocking findings addressed.
1. handleGpuChildCrash early-returned on isInstallDirAclSuspect() BEFORE
recordGpuCrash, so the crash left no trace in the 30s rolling window. The
suspect window is armed on every win32 non-serve launch, and the field
bundles put it at 0.8-1.7s after main_window_created on hosts whose DACL is
clean (matchesPoisonSignature=false) -- squarely inside the 2.1-6.2s
bad-driver bursts this repo already pinned in
gpu-crash-fallback-field-sessions.test.ts. A healthy machine with a failing
driver could lose an entire coalesced burst and never engage safe graphics.
The crash is now always recorded; only the engagement consults the verdict,
and it waits for the verdict rather than acting on the suspicion
(waitForInstallDirAclVerdict, resolved by the probe's onDone or by the
existing 15s grace, whichever lands first).
Deviation from the review's suggested shape, deliberately: awaiting the
verdict before persisting anything reintroduces the exact race
gpu-fallback-engagement.ts documents -- Chromium aborts the whole browser
process on the 6th GPU crash, ~1.3s after the 3rd, which is less than the
probe takes to answer. So the unconfirmed marker is written up front and
withdrawn if the verdict comes back poisoned. A machine killed mid-wait
still comes back software-rendered, and its marker is unconfirmed, which is
the state the repair's own clear already retires.
gpu-lifecycle-install-dir-acl-guard.test.ts now drives the real
GpuCrashFallbackTracker and the real engagement path (the restart prompt
firing is the signal) instead of a stub tracker, and covers the case the
previous suite could not express: a burst that lands entirely inside the
pending window still engages once the probe reports clean. Four reverts go
red -- restoring the pre-record guard (2 tests), dropping the wait, dropping
the post-wait re-check, and dropping the pre-wait marker write (2 tests).
2. The round-1 evidence block quoted commits, a test name and pass counts that
no longer exist, and its real-icacls Windows run predated the commit that
rewrote the gate. Re-run at this commit; counts and the live-Windows result
are restated in the handoff rather than carried forward.
Also from the non-blocking list:
- 'marker-hit' conflated "already repaired" with "retry budget spent", because
hasMarkerFor matches outcome === 'repaired' too. The result now carries
alreadyRepaired, and the recovery maps that to stage 'repaired' -- so a launch
killed between a successful repair and its marker clear no longer tells the
user the folder needs an administrator, no longer latches
isInstallDirAclSuspect() for the session, and does retire the poison marker.
Not applied, with reasoning:
- "clearGpuFallbackMarker narrowed to userConfirmed === false leaves the target
population software-rendered after a repair." The summary was overstated and
is corrected, but the narrowing stands: a userConfirmed marker now requires a
clean DACL verdict, because the restart prompt that writes it is exactly what
the gate above withholds while the install is a suspect. The population this
family targets can no longer reach confirmMarker while poisoned.
- "writeInstallDirAclPoisonMarker re-stamps on a budget-exhausted machine
forever." True, but on that machine the tree really is still poisoned and the
gate resolves immediately ('skipped', no icacls spawn, no 20s wait), so the
marker is telling the truth. Retiring it would be wrong; only a clean probe
reading should.
* fix(windows): register the real-icacls spec and stop its teardown racing icacls
Two ratchets were red:
- windows-lane-tree-removal-boundary: the win32 spec's afterAll used raw
rmSync on a tree two icacls.exe children had just rewritten DACLs on, which
is the EPERM race removeTreeSync exists for.
- win32-test-lane-registration: the spec was in the pr.yml argv but not in
WINDOWS_PACKAGE_TESTS, so a future diff touching only test files would not
select package_windows and the spec would self-skip on ubuntu and report
success.
* fix(windows): re-arm the GPU fallback latch when the install-DACL verdict withholds it
recordGpuCrash reports the threshold crossing exactly once and latches `engaged`.
handleGpuChildCrash consumes that report before consulting the DACL verdict, and
installDirAclClearsGpuFallback then discards it — so nothing could ever engage
safe graphics again in that process. A machine whose tree the repair fixes and
whose driver is genuinely broken stayed hardware-accelerated through an unbounded
crash loop, with no prompt and no marker.
disengage() releases only the one-shot latch; the crash window is untouched, so a
real driver burst is still never erased. Test is RED without the re-arm.
* fix(windows): keep the safe-graphics marker while an install-DACL repair is in flight
The gate dispatches a repair without arming the probe clock, so
waitForInstallDirAclVerdict() returns immediately and the withdrawal deleted the
marker inside Chromium's FATAL window (crash 6 lands ~1.3s after crash 3, well
inside the 20s gate). The process then died mid-repair, spent no attempt, and
relaunched hardware accelerated into the same gate — spawning the same GPU
children, FATALing again, forever.
Hold the marker while poison.stage is 'pending' so that launch comes back
software rendered and the next gate runs to completion. Still not engaged this
launch, so --in-process-gpu does not erase the sibling-death evidence. A
terminal verdict has no next step to rescue, so it still withdraws. Both new
tests are RED without the retention.
* fix(windows): stop a repaired marker outranking a fresh poison verdict
The probe reads the install DACL and finds it poisoned; `startRepair` dispatches;
`markerHitFor` sees a repair marker recording `outcome: 'repaired'` for the same
installDir+appVersion and reports `alreadyRepaired`, which the recovery module maps
to stage 'repaired'. So the launch that just proved the tree poisoned runs no icacls,
deletes the poison marker that arms the next launch's pre-window gate, clears the
suspect flag so `--in-process-gpu` can engage on a tree safe graphics cannot rescue,
and tells the user "Orca repaired the permissions."
Reachable whenever the tree is re-poisoned after one successful repair of the same
version, and whenever a repair reports success without clearing the tree — the silent
icacls no-op this module exists to document.
A DACL reading taken this launch now outranks the marker: `probeConfirmedPoisoned`
stops `outcome: 'repaired'` short-circuiting the repair. The attempt budget still
bounds it, so an unrepairable tree does not re-spawn icacls forever. The pre-window
gate does not set the flag — it acts on a marker from an earlier launch, not on
evidence of its own, so a recorded repair still outranks it there.
Also drives the GPU-fallback re-arm test through a repair that actually completes
'repaired', rather than a later clean probe, which is the route the review exercised.
* fix(windows): make the pre-window ACL gate act on the poison evidence it fired on
The gate fired on a poison marker — an earlier launch's DACL reading that nothing has
retired — but withheld `probeConfirmedPoisoned` from the repair, so a repair marker
recording an older success still short-circuited it. On the three-launch shape the gate
exists for (repair succeeds; tree is re-poisoned; the next launch's probe records the
poison but dies before writing its repair marker) the gate ran no icacls, deleted the
poison marker that arms every later gate, un-suspected the tree so --in-process-gpu could
engage, and told the user "Orca repaired the permissions." `applyInstallDirAclProbeVerdict`
then swallowed that launch's own reading behind `if (poison) return`.
Both callers of `startRepair` hold outstanding poison evidence, so the flag is now
unconditional (renamed `poisonEvidenceOutstanding`) and `marker-hit` means only that the
attempt budget is spent. The probe guard is narrowed to an in-flight gate repair: a reading
taken after the gate finished re-arms the poison marker and downgrades a claimed repair.
Also: withholding safe graphics now ends with the repair budget. A machine whose attempts
are spent while the signature persists was denied safe graphics on every launch for the
life of that appVersion — and had its marker deleted each time — including the healthy
installs the probe's flag-blind ACE match over-matches, where the driver really is broken.
Non-blocking, same lane: re-read `isQuitting` after the up-to-15s verdict wait, and skip
the recovered-launch prompt when the ACL gate retired the marker read before whenReady.
* fix(windows): stop a timed-out gate repair outranking a later poison reading
The gate's 20s budget expires while icacls runs on under its own 120s cap, so
the probe can read the tree poisoned while that repair is still in flight. Its
success claim then deleted the poison marker, un-suspected the tree and told the
user their permissions were fixed. The reading is now latched and outranks it.
* fix(windows): stop a gate repair claim pre-empting this launch's probe reading
Round-7 adversarial findings, both driven against the real modules:
- isInstallDirAclSuspect returned false the moment the pre-window gate set
stage 'repaired', short-circuiting ahead of the probe-pending grace check.
The GPU children die 48-1373ms after window creation while the probe
answers 0.9-3.0s in, so an icacls that silently no-opped (exit 0, tree
untouched) opened exactly that interval to --in-process-gpu on a
still-poisoned tree - and a 'keep safe graphics' answer then pinned a
userConfirmed marker no later repair may clear, with the poison marker
already deleted so no later launch gates. The claim now stays provisional
until this launch's probe corroborates it or the grace window lapses.
- A probe reading that disproves a 'repaired' claim re-armed the poison
marker but never restored the unconfirmed safe-graphics marker the claim
had cleared, so the next launch relaunched hardware-accelerated into the
re-armed gate. The clear is now captured and handed back on disproof.
* test(windows): pin the nested and update-inherited grants against real icacls
The live spec asserted the grant landed on the root-level module file only.
It now also pins that the flagless /T pass reaches a nested file carrying
its own protected DACL (the shape app.asar.unpacked and node_modules have),
and that a file written after the repair inherits the (OI)(CI) root grant -
the stated reason that grant form exists.
* fix(windows): keep the recovered-launch prompt silent while the tree is the suspect
Round-8 fresh-eyes finding, driven against the real modules: the prompt
re-read the marker the pre-window gate may have retired, but never consulted
isInstallDirAclSuspect() - so after a FAILED gate (tree still a live suspect,
window blank behind the 10s reveal fallback, Keep as both defaultId and
cancelId) a 'keep it' answer pinned a userConfirmed marker no later repair
may clear, on the exact victim class the repair cannot help. The guard now
covers both gate outcomes; staying silent leaves the marker unconfirmed,
which a successful repair still retires.
---------
Co-authored-by: Orca Worker
Co-authored-by: OrcaWin
---
.github/workflows/pr.yml | 1 +
config/scripts/pr-code-change-scope.mjs | 1 +
.../gpu-crash-fallback-decision.ts | 13 +
...pu-lifecycle-install-dir-acl-guard.test.ts | 442 ++++++++++
src/main/startup/gpu-lifecycle.ts | 89 ++
.../startup/main-process-runtime-launch.ts | 12 +
src/main/startup/main-window-actions.ts | 8 +-
src/main/startup/main-window-controller.ts | 13 +-
.../windows-install-dir-acl-poison-marker.ts | 77 ++
.../startup/windows-install-dir-acl-probe.ts | 10 +-
.../windows-install-dir-acl-recovery.test.ts | 801 +++++++++++++++++-
.../windows-install-dir-acl-recovery.ts | 282 +++++-
...ndows-install-dir-acl-repair.win32.test.ts | 124 +++
...ows-install-dir-acl-startup-wiring.test.ts | 56 ++
...ows-install-dir-package-acl-repair.test.ts | 38 +-
.../windows-install-dir-package-acl-repair.ts | 74 +-
src/main/window/focus-existing-window.test.ts | 35 +
src/main/window/focus-existing-window.ts | 4 +-
18 files changed, 2037 insertions(+), 43 deletions(-)
create mode 100644 src/main/startup/gpu-lifecycle-install-dir-acl-guard.test.ts
create mode 100644 src/main/startup/windows-install-dir-acl-poison-marker.ts
create mode 100644 src/main/startup/windows-install-dir-acl-repair.win32.test.ts
create mode 100644 src/main/startup/windows-install-dir-acl-startup-wiring.test.ts
diff --git a/.github/workflows/pr.yml b/.github/workflows/pr.yml
index ed37cefb434..ca1651325c9 100644
--- a/.github/workflows/pr.yml
+++ b/.github/workflows/pr.yml
@@ -816,6 +816,7 @@ jobs:
src/main/cli/wsl-cli-powershell-boundary.test.ts
src/main/cursor/hook-service.test.ts
src/main/orca-profiles/profile-index-store.test.ts
+ src/main/startup/windows-install-dir-acl-repair.win32.test.ts
src/main/runtime/repo-worktree-admin-fingerprint.test.ts
src/main/runtime/worktree-scan-admin-fingerprint-gate.test.ts
src/shared/secure-file-fsync-flags.test.ts
diff --git a/config/scripts/pr-code-change-scope.mjs b/config/scripts/pr-code-change-scope.mjs
index 7111531c35c..f5a73f6239f 100644
--- a/config/scripts/pr-code-change-scope.mjs
+++ b/config/scripts/pr-code-change-scope.mjs
@@ -228,6 +228,7 @@ const WINDOWS_PACKAGE_TESTS = [
'src/main/cli/wsl-cli-powershell-boundary.test.ts',
'src/main/cursor/hook-service.test.ts',
'src/main/orca-profiles/profile-index-store.test.ts',
+ 'src/main/startup/windows-install-dir-acl-repair.win32.test.ts',
'src/main/runtime/repo-worktree-admin-fingerprint.test.ts',
'src/main/runtime/worktree-scan-admin-fingerprint-gate.test.ts',
'src/shared/secure-file-fsync-flags.test.ts',
diff --git a/src/main/crash-reporting/gpu-crash-fallback-decision.ts b/src/main/crash-reporting/gpu-crash-fallback-decision.ts
index e962952c695..ed2034a62be 100644
--- a/src/main/crash-reporting/gpu-crash-fallback-decision.ts
+++ b/src/main/crash-reporting/gpu-crash-fallback-decision.ts
@@ -50,6 +50,8 @@ export class GpuCrashFallbackTracker {
crashesInWindow: number
} {
if (this.engaged || !Number.isFinite(msSinceLaunch) || msSinceLaunch < 0) {
+ // Crashes landing while engaged (e.g. during a verdict wait before `disengage`) are
+ // not recorded, so a reported crashesInWindow can understate the actual burst.
return { shouldEngageFallback: false, crashesInWindow: this.recentCrashes.length }
}
// Why: out-of-order arrivals would corrupt the sorted window, and a clock
@@ -73,6 +75,17 @@ export class GpuCrashFallbackTracker {
return this.engaged
}
+ /**
+ * Re-arm after an engagement the caller decided not to act on. `recordGpuCrash`
+ * latches `engaged` and reports the threshold crossing exactly once, so a caller
+ * that discards that one report would otherwise silence safe graphics for the
+ * rest of the process — including a later burst it would have acted on.
+ * Leaves the crash window intact; only the one-shot latch is released.
+ */
+ disengage(): void {
+ this.engaged = false
+ }
+
/** Crash times currently inside the window. Exposed to assert the pruning invariant. */
windowSnapshot(): readonly number[] {
return [...this.recentCrashes]
diff --git a/src/main/startup/gpu-lifecycle-install-dir-acl-guard.test.ts b/src/main/startup/gpu-lifecycle-install-dir-acl-guard.test.ts
new file mode 100644
index 00000000000..e35fe401cdc
--- /dev/null
+++ b/src/main/startup/gpu-lifecycle-install-dir-acl-guard.test.ts
@@ -0,0 +1,442 @@
+import { mkdtempSync, writeFileSync } from 'node:fs'
+import { tmpdir } from 'node:os'
+import { join } from 'node:path'
+import { afterAll, beforeAll, beforeEach, describe, expect, it, vi } from 'vitest'
+
+// Hoisted with the vi.mock factory below. 'Keep Running' — the prompt firing at all is the signal.
+const { showMessageBox, userData } = vi.hoisted(() => ({
+ showMessageBox: vi.fn(async () => ({ response: 1 })),
+ userData: { path: '' }
+}))
+
+// Why the mocks: gpu-lifecycle's import graph reaches electron and the toolkit's
+// electron re-export. Everything below this is the real module under test.
+vi.mock('electron', () => ({
+ app: {
+ getPath: () => userData.path,
+ getVersion: () => '1.4.184',
+ getGPUFeatureStatus: () => ({}),
+ setAboutPanelOptions: vi.fn(),
+ commandLine: { appendSwitch: vi.fn() },
+ disableHardwareAcceleration: vi.fn(),
+ isReady: () => true,
+ exit: vi.fn(),
+ on: vi.fn(),
+ name: 'Orca'
+ },
+ dialog: { showMessageBox }
+}))
+vi.mock('@electron-toolkit/utils', () => ({
+ is: { dev: false },
+ optimizer: { watchWindowShortcuts: vi.fn() },
+ electronApp: { setAppUserModelId: vi.fn() }
+}))
+
+import type { ProcessResult, ProcessSpec } from '../../shared/child-process/run-process'
+import {
+ DEFAULT_GPU_CRASH_FALLBACK_THRESHOLD,
+ DEFAULT_GPU_CRASH_FALLBACK_WINDOW_MS,
+ GpuCrashFallbackTracker
+} from '../crash-reporting/gpu-crash-fallback-decision'
+import {
+ readGpuFallbackMarker,
+ writeGpuFallbackMarker,
+ type GpuFallbackMarker
+} from './gpu-fallback-marker'
+import { handleGpuChildCrash, presentGpuFallbackRecoveredLaunchPrompt } from './gpu-lifecycle'
+import { gpuFallbackEnvironment, mainProcessState as state } from './main-process-state'
+import { writeInstallDirAclPoisonMarker } from './windows-install-dir-acl-poison-marker'
+import {
+ isInstallDirAclRepairPending,
+ noteWindowsInstallDirAclProbePending,
+ repairKnownPoisonedInstallDirBeforeWindow,
+ resetWindowsInstallDirAclRecoveryForTest,
+ startWindowsInstallDirAclRepairIfPoisoned
+} from './windows-install-dir-acl-recovery'
+import {
+ resetWindowsInstallDirAclRepairForTest,
+ WINDOWS_INSTALL_DIR_ACL_REPAIR_MARKER_FILE,
+ WINDOWS_INSTALL_DIR_ACL_REPAIR_SCHEME_VERSION
+} from './windows-install-dir-package-acl-repair'
+
+const INSTALL_DIR = 'C:\\Users\\neil\\AppData\\Local\\Programs\\orca'
+
+function recoveryOptions(userDataPath?: string): {
+ platform: 'win32'
+ installDir: string
+ appVersion: string
+ userDataPath: string
+ recordBreadcrumb: () => void
+} {
+ return {
+ platform: 'win32',
+ installDir: INSTALL_DIR,
+ appVersion: '1.4.184',
+ userDataPath: userDataPath ?? mkdtempSync(join(tmpdir(), 'orca-acl-gpu-guard-')),
+ recordBreadcrumb: () => undefined
+ }
+}
+
+/** icacls hangs until `finishRepair` — the in-flight window is when the GPU children die. */
+function reportProbePoisoned(): { finishRepair: () => Promise } {
+ let release = (): void => undefined
+ const walkingTheTree = new Promise((resolve) => {
+ release = resolve
+ })
+ startWindowsInstallDirAclRepairIfPoisoned(
+ { status: 'ok', matchesPoisonSignature: true, wellKnownNameCheckReliable: true },
+ {
+ ...recoveryOptions(),
+ runProcessFn: (async () => {
+ await walkingTheTree
+ return {
+ code: 0,
+ signal: null,
+ stdout: 'Successfully processed 3200 files; Failed processing 0 files',
+ stderr: '',
+ timedOut: false
+ }
+ }) as unknown as (spec: ProcessSpec) => Promise
+ }
+ )
+ return {
+ finishRepair: async () => {
+ release()
+ for (let i = 0; i < 200 && isInstallDirAclRepairPending(); i += 1) {
+ await new Promise((resolve) => setTimeout(resolve, 5))
+ }
+ }
+ }
+}
+
+/** A repair that settles, so `poison.stage` leaves 'pending' for a terminal verdict. */
+async function reportProbePoisonedWithSettledRepair(
+ exitCode: number,
+ userDataPath?: string
+): Promise {
+ startWindowsInstallDirAclRepairIfPoisoned(
+ { status: 'ok', matchesPoisonSignature: true, wellKnownNameCheckReliable: true },
+ {
+ ...recoveryOptions(userDataPath),
+ runProcessFn: (async () => ({
+ code: exitCode,
+ signal: null,
+ stdout: 'Successfully processed 3200 files; Failed processing 0 files',
+ stderr: exitCode === 0 ? '' : 'access denied',
+ timedOut: false
+ })) as unknown as (spec: ProcessSpec) => Promise
+ }
+ )
+ for (let i = 0; i < 200 && isInstallDirAclRepairPending(); i += 1) {
+ await new Promise((resolve) => setTimeout(resolve, 5))
+ }
+}
+
+/**
+ * The pre-window gate meeting a spent repair budget: the tree is still marked poisoned and
+ * Orca has no repair left to try. icacls must never be reached, so the runner throws.
+ */
+async function gateFindsRepairBudgetSpent(): Promise {
+ const options = recoveryOptions()
+ writeFileSync(
+ join(options.userDataPath, WINDOWS_INSTALL_DIR_ACL_REPAIR_MARKER_FILE),
+ JSON.stringify({
+ schemeVersion: WINDOWS_INSTALL_DIR_ACL_REPAIR_SCHEME_VERSION,
+ installDir: INSTALL_DIR,
+ appVersion: options.appVersion,
+ attemptedAt: Date.now(),
+ outcome: 'failed',
+ attempts: 3
+ })
+ )
+ writeInstallDirAclPoisonMarker(options.userDataPath, INSTALL_DIR, options.appVersion)
+ const mode = await repairKnownPoisonedInstallDirBeforeWindow({
+ ...options,
+ runProcessFn: (() => {
+ throw new Error('the spent budget must not spawn icacls')
+ }) as never
+ })
+ expect(mode).toBe('marker-hit')
+}
+
+function reportProbeClean(): void {
+ startWindowsInstallDirAclRepairIfPoisoned(
+ { status: 'ok', matchesPoisonSignature: false },
+ recoveryOptions()
+ )
+}
+
+/** One short of the fallback threshold, so the caller's next crash is the decisive one. */
+async function crashUpToThreshold(): Promise {
+ for (let i = 1; i < DEFAULT_GPU_CRASH_FALLBACK_THRESHOLD; i += 1) {
+ await handleGpuChildCrash('crashed', null, i * 200)
+ }
+}
+
+/**
+ * Driven end-to-end against the real tracker rather than asserted against the source:
+ * a source match is equally happy with the polarity inverted, and the property that
+ * matters is that a driver burst survives the ACL verdict either way.
+ */
+describe('handleGpuChildCrash vs the install-dir ACL verdict', () => {
+ let tracker: GpuCrashFallbackTracker
+ const realPlatform = process.platform
+
+ beforeAll(() => {
+ // The whole guard is win32-only, and so is the safe-graphics marker it writes.
+ Object.defineProperty(process, 'platform', { value: 'win32', configurable: true })
+ })
+
+ afterAll(() => {
+ Object.defineProperty(process, 'platform', { value: realPlatform, configurable: true })
+ })
+
+ beforeEach(() => {
+ userData.path = mkdtempSync(join(tmpdir(), 'orca-acl-gpu-userdata-'))
+ resetWindowsInstallDirAclRepairForTest()
+ resetWindowsInstallDirAclRecoveryForTest()
+ showMessageBox.mockClear()
+ state.isQuitting = false
+ state.isServeMode = false
+ state.gpuFallbackActiveThisLaunch = false
+ tracker = new GpuCrashFallbackTracker({
+ windowMs: DEFAULT_GPU_CRASH_FALLBACK_WINDOW_MS,
+ threshold: DEFAULT_GPU_CRASH_FALLBACK_THRESHOLD
+ })
+ state.gpuCrashFallbackTracker = tracker
+ })
+
+ it('engages safe graphics on a driver burst when nothing implicates the install DACL', async () => {
+ await crashUpToThreshold()
+ await handleGpuChildCrash('crashed', null, 600)
+ expect(showMessageBox).toHaveBeenCalledTimes(1)
+ })
+
+ // The regression this guard must never reintroduce: the probe is armed on every
+ // win32 launch, so a burst landing inside its window is the common driver case.
+ it('keeps counting crashes that land while the probe verdict is outstanding', async () => {
+ noteWindowsInstallDirAclProbePending()
+ await crashUpToThreshold()
+ const decisive = handleGpuChildCrash('crashed', null, 600)
+ expect(showMessageBox).not.toHaveBeenCalled()
+ reportProbeClean()
+ await decisive
+ expect(tracker.windowSnapshot()).toHaveLength(DEFAULT_GPU_CRASH_FALLBACK_THRESHOLD)
+ expect(showMessageBox).toHaveBeenCalledTimes(1)
+ })
+
+ it('withholds safe graphics while the install DACL is the suspect, but keeps the evidence', async () => {
+ reportProbePoisoned()
+ await crashUpToThreshold()
+ await handleGpuChildCrash('crashed', null, 600)
+ expect(tracker.windowSnapshot()).toHaveLength(DEFAULT_GPU_CRASH_FALLBACK_THRESHOLD)
+ expect(showMessageBox).not.toHaveBeenCalled()
+ })
+
+ it('withholds safe graphics when the outstanding verdict comes back poisoned', async () => {
+ noteWindowsInstallDirAclProbePending()
+ await crashUpToThreshold()
+ const decisive = handleGpuChildCrash('crashed', null, 600)
+ reportProbePoisoned()
+ await decisive
+ expect(showMessageBox).not.toHaveBeenCalled()
+ })
+
+ // The gate's 'repaired' is icacls's exit claim, not a reading of the tree, and an icacls
+ // that silently no-opped exits 0 on a tree it left poisoned. The GPU children die in the
+ // interval before this launch's probe answers, so a claim that un-suspects the tree there
+ // engages --in-process-gpu on a tree safe graphics cannot rescue — and a "keep it" answer
+ // then pins a userConfirmed marker no later repair may clear.
+ it('withholds safe graphics between a gate repair claim and this launch probe reading', async () => {
+ writeInstallDirAclPoisonMarker(userData.path, INSTALL_DIR, '1.4.184')
+ const mode = await repairKnownPoisonedInstallDirBeforeWindow({
+ ...recoveryOptions(userData.path),
+ runProcessFn: (async () => ({
+ code: 0,
+ signal: null,
+ stdout: 'Successfully processed 3200 files; Failed processing 0 files',
+ stderr: '',
+ timedOut: false
+ })) as unknown as (spec: ProcessSpec) => Promise
+ })
+ expect(mode).toBe('repaired')
+ noteWindowsInstallDirAclProbePending()
+
+ await crashUpToThreshold()
+ const decisive = handleGpuChildCrash('crashed', null, 600)
+ expect(showMessageBox).not.toHaveBeenCalled()
+
+ // The reading lands poisoned: the claim was false, and engagement stays withheld.
+ startWindowsInstallDirAclRepairIfPoisoned(
+ { status: 'ok', matchesPoisonSignature: true, wellKnownNameCheckReliable: true },
+ recoveryOptions(userData.path)
+ )
+ await decisive
+ expect(showMessageBox).not.toHaveBeenCalled()
+ })
+
+ // Chromium aborts the browser on the 6th GPU crash, sooner than the probe can answer,
+ // so the wait must not be the reason a machine comes back hardware-accelerated.
+ it('holds an unconfirmed safe-graphics marker on disk across the wait', async () => {
+ noteWindowsInstallDirAclProbePending()
+ await crashUpToThreshold()
+ const decisive = handleGpuChildCrash('crashed', null, 600)
+ expect(readGpuFallbackMarker(userData.path)?.userConfirmed).toBe(false)
+ reportProbePoisoned()
+ await decisive
+ // The verdict dispatched a repair, so the marker stays for the launch that repair rescues.
+ expect(readGpuFallbackMarker(userData.path)?.userConfirmed).toBe(false)
+ })
+
+ it('engages immediately once the probe has already reported the install clean', async () => {
+ noteWindowsInstallDirAclProbePending()
+ reportProbeClean()
+ await crashUpToThreshold()
+ await handleGpuChildCrash('crashed', null, 600)
+ expect(showMessageBox).toHaveBeenCalledTimes(1)
+ })
+
+ // Both from the re-run adversarial round. The gate dispatches a repair without arming the
+ // probe clock, so `waitForInstallDirAclVerdict` returns immediately and the withdrawal used
+ // to delete the marker inside Chromium's ~1.3s FATAL window — leaving the machine to
+ // relaunch hardware accelerated into the same 20s gate, forever.
+ it('keeps the safe-graphics marker on disk while a repair is still in flight', async () => {
+ reportProbePoisoned()
+ await crashUpToThreshold()
+ await handleGpuChildCrash('crashed', null, 600)
+
+ expect(showMessageBox).not.toHaveBeenCalled()
+ expect(readGpuFallbackMarker(userData.path)?.userConfirmed).toBe(false)
+ })
+
+ it('still withdraws the marker once the verdict is terminal rather than a pending repair', async () => {
+ await reportProbePoisonedWithSettledRepair(1)
+ await crashUpToThreshold()
+ await handleGpuChildCrash('crashed', null, 600)
+
+ expect(showMessageBox).not.toHaveBeenCalled()
+ // No repair is in flight to rescue a later launch, so the marker is not held.
+ expect(readGpuFallbackMarker(userData.path)).toBeNull()
+ })
+
+ // The verdict wait can span the probe's whole 15s grace window, and the entry guard was
+ // read before it. A quit that starts inside the wait must not be answered with a modal.
+ it('does not prompt when the user quits during the verdict wait', async () => {
+ noteWindowsInstallDirAclProbePending()
+ await crashUpToThreshold()
+ const decisive = handleGpuChildCrash('crashed', null, 600)
+ state.isQuitting = true
+ reportProbeClean()
+ await decisive
+ expect(showMessageBox).not.toHaveBeenCalled()
+ })
+
+ // Withholding is a bounded delay, not a permanent suppression. Once the repair budget is
+ // spent no repair is coming on this launch or any later one, so pinning the tree as the
+ // suspect forever denied safe graphics on EVERY launch for the life of that version — and
+ // deleted the marker each time, so the machine also relaunched hardware accelerated. The
+ // victims are a standard-user install icacls can never fix and, via the probe's flag-blind
+ // ACE match, healthy installs whose driver genuinely is broken.
+ it('offers safe graphics on every launch once the ACL repair budget is spent', async () => {
+ for (let launch = 1; launch <= 3; launch += 1) {
+ resetWindowsInstallDirAclRepairForTest()
+ resetWindowsInstallDirAclRecoveryForTest()
+ showMessageBox.mockClear()
+ state.gpuCrashFallbackTracker = new GpuCrashFallbackTracker({
+ windowMs: DEFAULT_GPU_CRASH_FALLBACK_WINDOW_MS,
+ threshold: DEFAULT_GPU_CRASH_FALLBACK_THRESHOLD
+ })
+ await gateFindsRepairBudgetSpent()
+
+ await crashUpToThreshold()
+ await handleGpuChildCrash('crashed', null, 600)
+ expect(showMessageBox).toHaveBeenCalledTimes(1)
+ }
+ })
+
+ // Still withheld while the budget has an attempt left: the repair is the better answer,
+ // and this is the launch a next one can be rescued on.
+ it('still withholds while the repair has an attempt left to spend', async () => {
+ await reportProbePoisonedWithSettledRepair(1)
+ await crashUpToThreshold()
+ await handleGpuChildCrash('crashed', null, 600)
+ expect(showMessageBox).not.toHaveBeenCalled()
+ })
+
+ // recordGpuCrash reports the threshold crossing once and latches. Withholding consumes
+ // that one report, so without a re-arm the same process could never engage again — a
+ // machine whose tree is repaired and whose driver is genuinely broken would be stuck
+ // hardware-accelerated through an unbounded crash loop.
+ it('can still engage a later burst after a withheld one, once the tree is repaired', async () => {
+ const repair = reportProbePoisoned()
+ await crashUpToThreshold()
+ await handleGpuChildCrash('crashed', null, 600)
+ expect(showMessageBox).not.toHaveBeenCalled()
+
+ // The repair itself reports 'repaired': the tree is no longer the suspect.
+ await repair.finishRepair()
+ expect(isInstallDirAclRepairPending()).toBe(false)
+
+ for (let i = 1; i <= DEFAULT_GPU_CRASH_FALLBACK_THRESHOLD; i += 1) {
+ await handleGpuChildCrash('crashed', null, 10_000 + i * 200)
+ }
+ expect(showMessageBox).toHaveBeenCalledTimes(1)
+ })
+})
+
+// The safe-graphics marker is read before whenReady, and the pre-window ACL gate runs after
+// that read. Asking "keep safe graphics?" on a machine Orca has just repaired invites a
+// `userConfirmed: true` marker that pins software rendering on healthy hardware.
+describe('presentGpuFallbackRecoveredLaunchPrompt vs a marker retired since it was read', () => {
+ const realPlatform = process.platform
+ const window = { isDestroyed: () => false } as unknown as Parameters<
+ typeof presentGpuFallbackRecoveredLaunchPrompt
+ >[0]
+
+ beforeAll(() => {
+ Object.defineProperty(process, 'platform', { value: 'win32', configurable: true })
+ })
+
+ afterAll(() => {
+ Object.defineProperty(process, 'platform', { value: realPlatform, configurable: true })
+ })
+
+ beforeEach(() => {
+ userData.path = mkdtempSync(join(tmpdir(), 'orca-acl-gpu-recovered-'))
+ resetWindowsInstallDirAclRepairForTest()
+ resetWindowsInstallDirAclRecoveryForTest()
+ showMessageBox.mockClear()
+ state.isQuitting = false
+ const info = { engagedAt: Date.now(), crashesInWindow: 3, userConfirmed: false }
+ writeGpuFallbackMarker(userData.path, info, {
+ ...gpuFallbackEnvironment(),
+ platform: 'win32'
+ })
+ state.activeGpuFallbackMarker = readGpuFallbackMarker(userData.path) as GpuFallbackMarker
+ })
+
+ it('asks while the marker is still on disk', async () => {
+ showMessageBox.mockResolvedValueOnce({ response: 0 })
+ await presentGpuFallbackRecoveredLaunchPrompt(window)
+ expect(showMessageBox).toHaveBeenCalledTimes(1)
+ })
+
+ it('stays silent once the install-DACL repair has cleared it', async () => {
+ await reportProbePoisonedWithSettledRepair(0, userData.path)
+ expect(readGpuFallbackMarker(userData.path)).toBeNull()
+
+ await presentGpuFallbackRecoveredLaunchPrompt(window)
+ expect(showMessageBox).not.toHaveBeenCalled()
+ })
+
+ // The symmetric case to the one above: a FAILED repair leaves the marker on disk and the
+ // tree a live suspect, the window the prompt lands on is blank, and Keep is both defaultId
+ // and cancelId — so asking invites a userConfirmed pin no later repair may clear.
+ it('stays silent while the install DACL is still the suspect', async () => {
+ await reportProbePoisonedWithSettledRepair(1, userData.path)
+ expect(readGpuFallbackMarker(userData.path)).not.toBeNull()
+
+ await presentGpuFallbackRecoveredLaunchPrompt(window)
+ expect(showMessageBox).not.toHaveBeenCalled()
+ })
+})
diff --git a/src/main/startup/gpu-lifecycle.ts b/src/main/startup/gpu-lifecycle.ts
index da852bcb491..89055aa2ba3 100644
--- a/src/main/startup/gpu-lifecycle.ts
+++ b/src/main/startup/gpu-lifecycle.ts
@@ -16,6 +16,12 @@ import { promptForGpuFallbackRestart } from '../crash-reporting/gpu-fallback-res
import { engageGpuFallbackAfterCrashBurst } from '../crash-reporting/gpu-fallback-engagement'
import { recordCrashBreadcrumb } from '../crash-reporting/crash-breadcrumb-store'
import { recordDurableCrashBreadcrumb } from '../crash-reporting/durable-crash-breadcrumb'
+import {
+ isInstallDirAclRepairExhausted,
+ isInstallDirAclRepairPending,
+ isInstallDirAclSuspect,
+ waitForInstallDirAclVerdict
+} from './windows-install-dir-acl-recovery'
import { mainProcessState as state, gpuFallbackEnvironment } from './main-process-state'
import { createGpuAccelerationAboutPanelOptions } from '../menu/gpu-acceleration-about-panel'
@@ -85,6 +91,18 @@ export async function presentGpuFallbackRecoveredLaunchPrompt(
// One prompt per process. A failure leaves the on-disk marker unconfirmed so the next launch retries.
state.activeGpuFallbackMarker = null
const userDataPath = app.getPath('userData')
+ // The marker was read before whenReady; the pre-window ACL gate can have retired it since.
+ // Asking then would let a "keep it" answer pin software rendering on a machine Orca just fixed.
+ if (!readActiveGpuFallbackMarker(userDataPath, gpuFallbackEnvironment())) {
+ return
+ }
+ // The symmetric case: while the tree, not the driver, is on trial (a failed gate leaves it
+ // a live suspect), a "keep it" answer would pin a userConfirmed marker no later repair may
+ // clear — on the window the poison keeps blank. Staying silent leaves the marker
+ // unconfirmed, which a successful repair still retires.
+ if (isInstallDirAclSuspect()) {
+ return
+ }
await handleGpuFallbackRecoveredLaunch({
isQuitting: () => state.isQuitting,
prompt: () => promptForGpuFallbackRecoveredLaunch(window),
@@ -114,6 +132,66 @@ export async function presentGpuFallbackRecoveredLaunchPrompt(
})
}
+/**
+ * Why withholding ends with the repair budget: withholding only buys the ACL repair the
+ * chance to land first. Once its attempts are spent no repair is coming on this launch or
+ * any later one, so holding safe graphics back forever would deny the only recovery left —
+ * on a genuinely poisoned tree Orca has already told the user the admin commands, and the
+ * probe's flag-blind ACE match also over-matches healthy installs whose driver really is
+ * the fault. It is a bounded delay, not a permanent suppression.
+ */
+function installDirAclWithholdsGpuFallback(): boolean {
+ return isInstallDirAclSuspect() && !isInstallDirAclRepairExhausted()
+}
+
+/**
+ * Why: a poisoned install DACL kills the GPU child exactly like a bad driver, but safe
+ * graphics does not rescue it and --in-process-gpu removes the GPU child, erasing the
+ * sibling deaths that identify the real cause.
+ *
+ * Why the marker is written before the wait rather than after: Chromium aborts the whole
+ * browser process on the 6th GPU crash, ~1.3s after the 3rd — less than the probe takes
+ * to answer — so a machine that dies waiting must still come back software-rendered.
+ * The withdrawal below, and the repair's own clear of an unconfirmed marker, undo it.
+ */
+async function installDirAclClearsGpuFallback(
+ userDataPath: string,
+ crashesInWindow: number
+): Promise {
+ if (!installDirAclWithholdsGpuFallback()) {
+ return true
+ }
+ const persisted = persistGpuFallbackMarker(userDataPath, {
+ engagedAt: Date.now(),
+ crashesInWindow,
+ userConfirmed: false
+ })
+ await waitForInstallDirAclVerdict()
+ if (!installDirAclWithholdsGpuFallback()) {
+ return true
+ }
+ // Why the marker survives a pending repair: withdrawing it here left a machine that
+ // Chromium FATALs mid-repair (crash 6 lands ~1.3s after crash 3, well inside the gate)
+ // relaunching hardware accelerated into the same 20s gate, spawning the same GPU children,
+ // FATALing again — with no attempt spent, so the loop never advances. Keeping it costs a
+ // healthy machine nothing: a successful repair clears an unconfirmed marker itself, and a
+ // clean probe reading means we never reach here. It is still not *engaged* this launch, so
+ // --in-process-gpu does not erase the sibling-death evidence on the launch that is running.
+ const repairPending = isInstallDirAclRepairPending()
+ if (persisted && !repairPending) {
+ clearGpuFallbackMarker(userDataPath)
+ }
+ // Why re-arm: recordGpuCrash reports the threshold crossing once and latches. Withholding
+ // consumed that one report, so without this a later burst — including one after the repair
+ // succeeds and the tree is no longer the suspect — could never engage safe graphics again.
+ state.gpuCrashFallbackTracker.disengage()
+ recordDurableCrashBreadcrumb('gpu_fallback_withheld_install_dir_acl', {
+ crashesInWindow,
+ markerHeldForPendingRepair: repairPending
+ })
+ return false
+}
+
// Why: a burst of GPU child crashes means HW acceleration is unusable — persist a build-scoped marker and offer software rendering.
export async function handleGpuChildCrash(
reason: string,
@@ -124,12 +202,23 @@ export async function handleGpuChildCrash(
if (state.gpuFallbackActiveThisLaunch || state.isQuitting || state.isServeMode) {
return
}
+ // Recorded before any install-DACL consideration: the verdict decides whether safe
+ // graphics is the right answer, never whether the crash happened. Dropping it here
+ // would erase a real driver burst from the rolling window on healthy machines too.
const result = state.gpuCrashFallbackTracker.recordGpuCrash(crashedAt)
if (!result.shouldEngageFallback) {
return
}
const fallbackData = { processReason: reason, exitCode, crashesInWindow: result.crashesInWindow }
const userDataPath = app.getPath('userData')
+ if (!(await installDirAclClearsGpuFallback(userDataPath, result.crashesInWindow))) {
+ return
+ }
+ // Re-read after that wait: it can span the probe's whole grace window, and a quit that
+ // started inside it must not be answered with a modal and a relaunch.
+ if (state.isQuitting) {
+ return
+ }
await engageGpuFallbackAfterCrashBurst(
{ reason, exitCode, crashesInWindow: result.crashesInWindow, engagedAt: Date.now() },
{
diff --git a/src/main/startup/main-process-runtime-launch.ts b/src/main/startup/main-process-runtime-launch.ts
index 9390fd23497..5f2691d6f31 100644
--- a/src/main/startup/main-process-runtime-launch.ts
+++ b/src/main/startup/main-process-runtime-launch.ts
@@ -24,6 +24,7 @@ import {
import { prepareCodexRuntimeHomeForLaunch } from './codex-launch-preparation'
import { prepareCodexSessionResumeForLaunch } from './codex-session-resume-launch'
import { startWindowsDesktopBeforeShellPathReady } from './windows-desktop-shell-path-startup'
+import { repairKnownPoisonedInstallDirBeforeWindow } from './windows-install-dir-acl-recovery'
import { registerServeSignalHandlers } from './serve-signal-handlers'
import { settleServeDesktopActivation } from './serve-desktop-activation'
import {
@@ -304,6 +305,17 @@ export async function initializeMainProcessRuntimeLaunch(
// Why published: the renderer's git-environment barrier must fence on the same
// generation the terminal startup services wait for, not a later re-read.
state.shellPathReady = shellPathReady
+ // Why before any window: the poisoned install DACL kills the renderer at init, and
+ // the probe that detects it cannot finish before createMainWindow. Bounded, and a
+ // no-op (one absent-file read) unless a previous launch already recorded the verdict.
+ const aclGate = await repairKnownPoisonedInstallDirBeforeWindow({
+ isServeMode: state.isServeMode || serveOptions !== null,
+ userDataPath: app.getPath('userData'),
+ appVersion: app.getVersion()
+ })
+ if (aclGate !== 'not-marked' && aclGate !== 'skipped') {
+ logStartupMilestone('install-dir-acl-repair-blocking-done', { mode: aclGate })
+ }
let desktopWindow: BrowserWindow | null = null
if (process.platform === 'win32' && app.isPackaged && !serveOptions) {
const desktopStartup = startWindowsDesktopBeforeShellPathReady({
diff --git a/src/main/startup/main-window-actions.ts b/src/main/startup/main-window-actions.ts
index 96751d645db..585fa0a754e 100644
--- a/src/main/startup/main-window-actions.ts
+++ b/src/main/startup/main-window-actions.ts
@@ -12,7 +12,10 @@ import { ensureAutoUpdaterConfigured } from '../window/attach-main-window-servic
import { focusExistingMainWindow, safelyRevealWindow } from '../window/focus-existing-window'
import { mainProcessState as state } from './main-process-state'
import { loadMainWindow } from '../window/createMainWindow'
-import { describeInstallDirAclPoison } from './windows-install-dir-acl-recovery'
+import {
+ describeInstallDirAclPoison,
+ isBlockingInstallDirAclRepairInFlight
+} from './windows-install-dir-acl-recovery'
import { presentRendererRecoveryPrompt } from '../window/renderer-recovery-prompt'
// The window module injects this callback to avoid a cycle between actions and lifecycle code.
@@ -28,6 +31,9 @@ export function focusExistingWindow(): void {
app,
getWindow: () => state.mainWindow,
openWindow,
+ // Why: a 20s blank launch invites a second double-click, and icacls is rewriting
+ // the per-file DACLs a fresh renderer would read. The gated launch opens the window.
+ canOpenWindow: () => !isBlockingInstallDirAclRepairInFlight(),
warn: console.warn
})
}
diff --git a/src/main/startup/main-window-controller.ts b/src/main/startup/main-window-controller.ts
index 0d935d5f84a..63d84df763c 100644
--- a/src/main/startup/main-window-controller.ts
+++ b/src/main/startup/main-window-controller.ts
@@ -10,7 +10,10 @@ import { resolveConsent } from '../telemetry/consent'
import { trackAppOpenedOnce } from '../telemetry/client'
import { ensureWindowsUserDataAclGrant } from './windows-user-data-acl'
import { probeWindowsInstallDirAcl } from './windows-install-dir-acl-probe'
-import { startWindowsInstallDirAclRepairIfPoisoned } from './windows-install-dir-acl-recovery'
+import {
+ noteWindowsInstallDirAclProbePending,
+ startWindowsInstallDirAclRepairIfPoisoned
+} from './windows-install-dir-acl-recovery'
import { logStartupMilestone } from './startup-diagnostics'
import { notifyMainWindowBecameVisible } from '../window/main-window-visibility'
import { setTrayAttention } from '../tray/system-tray'
@@ -74,7 +77,7 @@ export function openMainWindow(options: { revealOnDidFinishLoad?: boolean } = {}
})
// Why here: read-only, and the install DACL is the one thing a 0x80000003
// child death cannot tell us about itself. See electron/electron#51761.
- probeWindowsInstallDirAcl({
+ const probeDispatched = probeWindowsInstallDirAcl({
isServeMode: state.isServeMode,
onDone: (data) =>
startWindowsInstallDirAclRepairIfPoisoned(data, {
@@ -83,6 +86,12 @@ export function openMainWindow(options: { revealOnDidFinishLoad?: boolean } = {}
appVersion: app.getVersion()
})
})
+ // Why gated on the dispatch: the probe is once-per-process while openMainWindow
+ // re-runs on every reopen, so arming this again would wait on a verdict that
+ // already landed — and drop every GPU crash for the grace window.
+ if (probeDispatched) {
+ noteWindowsInstallDirAclProbePending()
+ }
}
const window = createMainWindow(store, {
getIsQuitting: () => state.isQuitting,
diff --git a/src/main/startup/windows-install-dir-acl-poison-marker.ts b/src/main/startup/windows-install-dir-acl-poison-marker.ts
new file mode 100644
index 00000000000..46035e04e74
--- /dev/null
+++ b/src/main/startup/windows-install-dir-acl-poison-marker.ts
@@ -0,0 +1,77 @@
+import { existsSync, mkdirSync, readFileSync, rmSync, writeFileSync } from 'node:fs'
+import { join } from 'node:path'
+
+/**
+ * "This install directory was found poisoned and has not been proven healthy since."
+ *
+ * Why a separate marker from `windows-install-dir-acl-repair.json`: that one is
+ * written after an attempt finishes, so a launch the poison kills mid-repair
+ * leaves no state at all and the next launch repeats the whole late-repair dance.
+ * This one is written the moment the probe's verdict lands, and it is the only
+ * thing that lets a later launch know it is poisoned *before* it creates a window
+ * — the probe itself cannot answer that early. Same tiny synchronous-JSON shape
+ * as `gpu-fallback-marker.ts`, for the same reason.
+ */
+
+export const WINDOWS_INSTALL_DIR_ACL_POISON_MARKER_FILE = 'windows-install-dir-acl-poison.json'
+export const WINDOWS_INSTALL_DIR_ACL_POISON_SCHEME_VERSION = 1
+
+type PoisonMarker = {
+ schemeVersion: number
+ installDir: string
+ appVersion: string
+ detectedAt: number
+}
+
+function markerPath(userDataPath: string): string {
+ return join(userDataPath, WINDOWS_INSTALL_DIR_ACL_POISON_MARKER_FILE)
+}
+
+/** Keyed on both: a reinstall elsewhere or an update ships files with a fresh DACL. */
+export function hasInstallDirAclPoisonMarker(
+ userDataPath: string,
+ installDir: string,
+ appVersion: string
+): boolean {
+ try {
+ const parsed = JSON.parse(readFileSync(markerPath(userDataPath), 'utf-8')) as
+ | Partial
+ | undefined
+ return (
+ parsed?.schemeVersion === WINDOWS_INSTALL_DIR_ACL_POISON_SCHEME_VERSION &&
+ parsed.installDir === installDir &&
+ parsed.appVersion === appVersion
+ )
+ } catch {
+ return false // missing or corrupt -> treat the install as healthy
+ }
+}
+
+export function writeInstallDirAclPoisonMarker(
+ userDataPath: string,
+ installDir: string,
+ appVersion: string
+): void {
+ const marker: PoisonMarker = {
+ schemeVersion: WINDOWS_INSTALL_DIR_ACL_POISON_SCHEME_VERSION,
+ installDir,
+ appVersion,
+ detectedAt: Date.now()
+ }
+ try {
+ if (!existsSync(userDataPath)) {
+ mkdirSync(userDataPath, { recursive: true })
+ }
+ writeFileSync(markerPath(userDataPath), JSON.stringify(marker))
+ } catch {
+ // Best effort: without it the next launch just falls back to today's late repair.
+ }
+}
+
+export function clearInstallDirAclPoisonMarker(userDataPath: string): void {
+ try {
+ rmSync(markerPath(userDataPath), { force: true })
+ } catch {
+ // Best effort; a stale marker only costs one redundant icacls pass.
+ }
+}
diff --git a/src/main/startup/windows-install-dir-acl-probe.ts b/src/main/startup/windows-install-dir-acl-probe.ts
index 42db3ee3d10..25d1e581c9b 100644
--- a/src/main/startup/windows-install-dir-acl-probe.ts
+++ b/src/main/startup/windows-install-dir-acl-probe.ts
@@ -211,13 +211,16 @@ export function resetWindowsInstallDirAclProbeForTest(): void {
* Fire-and-forget; returns before any spawn. win32 only — no spawn and no fs I/O
* anywhere else. Called from openMainWindow, which runs after initObservability,
* so the durable record also emits a span into the diagnostics bundle.
+ *
+ * Returns whether THIS call dispatched the probe: openMainWindow re-runs on every
+ * reopen, and only a dispatch will ever produce an `onDone`.
*/
-export function probeWindowsInstallDirAcl(options: WindowsInstallDirAclProbeOptions = {}): void {
+export function probeWindowsInstallDirAcl(options: WindowsInstallDirAclProbeOptions = {}): boolean {
if ((options.platform ?? process.platform) !== 'win32' || options.isServeMode === true) {
- return
+ return false
}
if (probeStarted) {
- return
+ return false
}
probeStarted = true
// Why the try: this runs inline in openMainWindow, so anything thrown here
@@ -229,4 +232,5 @@ export function probeWindowsInstallDirAcl(options: WindowsInstallDirAclProbeOpti
} catch {
// Nothing left to report to that would not throw again.
}
+ return true
}
diff --git a/src/main/startup/windows-install-dir-acl-recovery.test.ts b/src/main/startup/windows-install-dir-acl-recovery.test.ts
index 2b5d00ff40a..7641cdc0700 100644
--- a/src/main/startup/windows-install-dir-acl-recovery.test.ts
+++ b/src/main/startup/windows-install-dir-acl-recovery.test.ts
@@ -1,18 +1,34 @@
-import { mkdtempSync } from 'node:fs'
+import { mkdtempSync, writeFileSync } from 'node:fs'
import { tmpdir } from 'node:os'
import { join } from 'node:path'
-import { beforeEach, describe, expect, it } from 'vitest'
+import { beforeEach, describe, expect, it, vi } from 'vitest'
import type { ProcessResult, ProcessSpec } from '../../shared/child-process/run-process'
+import type { CrashReportBreadcrumbData } from '../../shared/crash-reporting'
+import { readActiveGpuFallbackMarker, writeGpuFallbackMarker } from './gpu-fallback-marker'
import {
probeWindowsInstallDirAcl,
resetWindowsInstallDirAclProbeForTest
} from './windows-install-dir-acl-probe'
+import {
+ hasInstallDirAclPoisonMarker,
+ writeInstallDirAclPoisonMarker
+} from './windows-install-dir-acl-poison-marker'
import {
describeInstallDirAclPoison,
+ isBlockingInstallDirAclRepairInFlight,
+ isInstallDirAclRepairExhausted,
+ isInstallDirAclSuspect,
+ noteWindowsInstallDirAclProbePending,
+ repairKnownPoisonedInstallDirBeforeWindow,
resetWindowsInstallDirAclRecoveryForTest,
- startWindowsInstallDirAclRepairIfPoisoned
+ startWindowsInstallDirAclRepairIfPoisoned,
+ type WindowsInstallDirAclRecoveryOptions
} from './windows-install-dir-acl-recovery'
-import { resetWindowsInstallDirAclRepairForTest } from './windows-install-dir-package-acl-repair'
+import {
+ resetWindowsInstallDirAclRepairForTest,
+ WINDOWS_INSTALL_DIR_ACL_REPAIR_MARKER_FILE,
+ WINDOWS_INSTALL_DIR_ACL_REPAIR_SCHEME_VERSION
+} from './windows-install-dir-package-acl-repair'
import {
ALL_PACKAGES_ACE,
fakeIcaclsSpawn,
@@ -26,6 +42,8 @@ import {
const INSTALL_DIR = 'C:\\Users\\neil\\AppData\\Local\\Programs\\orca'
const APP_VERSION = '1.4.184'
+type Runner = (spec: ProcessSpec) => Promise
+
/**
* Drives the production path: the real probe hands its verdict to the real gate,
* which decides whether icacls ever runs. Only the two process seams are faked.
@@ -185,3 +203,778 @@ describe('describeInstallDirAclPoison', () => {
expect(describeInstallDirAclPoison()?.detail).toContain('repairing the permissions now')
})
})
+
+const POISON_VERDICT: CrashReportBreadcrumbData = {
+ status: 'ok',
+ matchesPoisonSignature: true,
+ wellKnownNameCheckReliable: true
+}
+const GPU_ENV = { appVersion: APP_VERSION, electronVersion: '43.4.1', platform: 'win32' } as const
+
+function recoveryOptions(userDataPath: string, run: Runner): WindowsInstallDirAclRecoveryOptions {
+ return {
+ platform: 'win32',
+ installDir: INSTALL_DIR,
+ appVersion: APP_VERSION,
+ userDataPath,
+ runProcessFn: run as never,
+ recordBreadcrumb: () => undefined
+ }
+}
+
+/** icacls' real success summary, as the repair's parser expects it. */
+const okRun: Runner = async () => ({
+ code: 0,
+ signal: null,
+ stdout: 'Successfully processed 3200 files; Failed processing 0 files',
+ stderr: '',
+ timedOut: false
+})
+
+describe('install-dir ACL repair vs the GPU safe-graphics marker', () => {
+ beforeEach(() => {
+ resetWindowsInstallDirAclProbeForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ resetWindowsInstallDirAclRecoveryForTest()
+ })
+
+ it('clears the sticky safe-graphics marker once the real cause is repaired', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-gpu-'))
+ // The machine is in the reproduced state: the poisoned install DACL killed the
+ // GPU child three times, so Orca latched safe graphics for this build.
+ writeGpuFallbackMarker(
+ userDataPath,
+ { engagedAt: Date.now(), crashesInWindow: 3, userConfirmed: false },
+ GPU_ENV
+ )
+ expect(readActiveGpuFallbackMarker(userDataPath, GPU_ENV)).not.toBeNull()
+
+ await new Promise((resolve) => {
+ startWindowsInstallDirAclRepairIfPoisoned(POISON_VERDICT, {
+ ...recoveryOptions(userDataPath, okRun),
+ // Settles after the repair's own setImmediate hop and its two icacls passes.
+ recordBreadcrumb: () => {
+ setTimeout(resolve, 0)
+ return undefined
+ }
+ })
+ })
+
+ expect(describeInstallDirAclPoison()?.detail).toContain('repaired the permissions')
+ // The GPU child deaths were never a driver fault, so safe graphics — and the
+ // --in-process-gpu launch that hides the next crash's evidence — must not outlive the repair.
+ expect(readActiveGpuFallbackMarker(userDataPath, GPU_ENV)).toBeNull()
+ })
+
+ // "Keep safe graphics" is a durable user choice with its own reasons; the repair
+ // only retires the latch Orca engaged on its own.
+ it('leaves a user-confirmed safe-graphics marker alone', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-gpu-'))
+ writeGpuFallbackMarker(
+ userDataPath,
+ { engagedAt: Date.now(), crashesInWindow: 3, userConfirmed: true },
+ GPU_ENV
+ )
+
+ await new Promise((resolve) => {
+ startWindowsInstallDirAclRepairIfPoisoned(POISON_VERDICT, {
+ ...recoveryOptions(userDataPath, okRun),
+ recordBreadcrumb: () => {
+ setTimeout(resolve, 0)
+ return undefined
+ }
+ })
+ })
+
+ expect(describeInstallDirAclPoison()?.detail).toContain('repaired the permissions')
+ expect(readActiveGpuFallbackMarker(userDataPath, GPU_ENV)?.userConfirmed).toBe(true)
+ })
+})
+
+describe('isInstallDirAclSuspect', () => {
+ beforeEach(() => {
+ resetWindowsInstallDirAclProbeForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ resetWindowsInstallDirAclRecoveryForTest()
+ })
+
+ it('is false when nothing has suggested the install DACL is involved', () => {
+ expect(isInstallDirAclSuspect()).toBe(false)
+ })
+
+ // The GPU child dies ~74ms in and the probe answers 0.9-3.0s later, so "no verdict
+ // yet" is the entire window in which the misdiagnosis happens.
+ it('holds while the probe verdict is outstanding, and releases on a clean verdict', () => {
+ noteWindowsInstallDirAclProbePending()
+ expect(isInstallDirAclSuspect()).toBe(true)
+
+ startWindowsInstallDirAclRepairIfPoisoned(
+ { status: 'ok', matchesPoisonSignature: false },
+ recoveryOptions(mkdtempSync(join(tmpdir(), 'orca-acl-suspect-')), okRun)
+ )
+ expect(isInstallDirAclSuspect()).toBe(false)
+ })
+
+ it('releases once the wait exceeds the grace window, so a silent probe cannot pin it', () => {
+ noteWindowsInstallDirAclProbePending()
+ expect(isInstallDirAclSuspect(Date.now() + 14_000)).toBe(true)
+ expect(isInstallDirAclSuspect(Date.now() + 16_000)).toBe(false)
+ })
+
+ it('holds through a repair that failed, and releases once one succeeds', async () => {
+ const failing: Runner = async () => ({
+ code: 5,
+ signal: null,
+ stdout: '',
+ stderr: 'Access is denied.',
+ timedOut: false
+ })
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-suspect-'))
+ startWindowsInstallDirAclRepairIfPoisoned(
+ POISON_VERDICT,
+ recoveryOptions(userDataPath, failing)
+ )
+ expect(isInstallDirAclSuspect()).toBe(true)
+ await vi.waitFor(() => expect(describeInstallDirAclPoison()?.detail).toContain('could not'))
+ // Still suspect: the tree is proven poisoned, and safe graphics does not rescue it.
+ expect(isInstallDirAclSuspect()).toBe(true)
+
+ resetWindowsInstallDirAclRecoveryForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ startWindowsInstallDirAclRepairIfPoisoned(
+ POISON_VERDICT,
+ recoveryOptions(mkdtempSync(join(tmpdir(), 'orca-acl-suspect-')), okRun)
+ )
+ await vi.waitFor(() => expect(isInstallDirAclSuspect()).toBe(false))
+ })
+})
+
+describe('repairKnownPoisonedInstallDirBeforeWindow', () => {
+ beforeEach(() => {
+ resetWindowsInstallDirAclProbeForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ resetWindowsInstallDirAclRecoveryForTest()
+ })
+
+ it('costs a healthy machine one absent-file read and no icacls', async () => {
+ const specs: ProcessSpec[] = []
+ const run: Runner = async (spec) => {
+ specs.push(spec)
+ return okRun(spec)
+ }
+ const mode = await repairKnownPoisonedInstallDirBeforeWindow(
+ recoveryOptions(mkdtempSync(join(tmpdir(), 'orca-acl-gate-')), run)
+ )
+ expect(mode).toBe('not-marked')
+ expect(specs).toHaveLength(0)
+ })
+
+ // The crash this fixes: launch 1 detects the poison but createMainWindow already
+ // ran, so the renderer is dead before icacls is spawned. Launch 2 must not repeat it.
+ it('repairs a launch that a previous one recorded as poisoned, before returning', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-gate-'))
+ // Launch 1: the probe reports poison and the app dies mid-repair.
+ startWindowsInstallDirAclRepairIfPoisoned(
+ POISON_VERDICT,
+ recoveryOptions(userDataPath, (() => new Promise(() => undefined)) as Runner)
+ )
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+
+ // Launch 2.
+ resetWindowsInstallDirAclRecoveryForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ writeGpuFallbackMarker(
+ userDataPath,
+ { engagedAt: Date.now(), crashesInWindow: 3, userConfirmed: false },
+ GPU_ENV
+ )
+ const specs: ProcessSpec[] = []
+ const run: Runner = async (spec) => {
+ specs.push(spec)
+ return okRun(spec)
+ }
+ const mode = await repairKnownPoisonedInstallDirBeforeWindow(recoveryOptions(userDataPath, run))
+ expect(mode).toBe('repaired')
+ // Both passes have already run by the time the window may be created.
+ expect(specs.map((spec) => spec.args?.[2])).toEqual([
+ '*S-1-15-2-2:(OI)(CI)(RX)',
+ '*S-1-15-2-2:(RX)'
+ ])
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(false)
+ expect(readActiveGpuFallbackMarker(userDataPath, GPU_ENV)).toBeNull()
+ })
+
+ it('gives up on its budget rather than holding the window open forever', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-gate-'))
+ startWindowsInstallDirAclRepairIfPoisoned(
+ POISON_VERDICT,
+ recoveryOptions(userDataPath, (() => new Promise(() => undefined)) as Runner)
+ )
+ resetWindowsInstallDirAclRecoveryForTest()
+ resetWindowsInstallDirAclRepairForTest()
+
+ const mode = await repairKnownPoisonedInstallDirBeforeWindow({
+ ...recoveryOptions(userDataPath, (() => new Promise(() => undefined)) as Runner),
+ timeoutMs: 20
+ })
+ expect(mode).toBe('timeout')
+ })
+
+ it('is a no-op off win32 and in serve mode', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-gate-'))
+ startWindowsInstallDirAclRepairIfPoisoned(
+ POISON_VERDICT,
+ recoveryOptions(userDataPath, (() => new Promise(() => undefined)) as Runner)
+ )
+ resetWindowsInstallDirAclRecoveryForTest()
+ resetWindowsInstallDirAclRepairForTest()
+
+ expect(
+ await repairKnownPoisonedInstallDirBeforeWindow({
+ ...recoveryOptions(userDataPath, okRun),
+ platform: 'darwin'
+ })
+ ).toBe('skipped')
+ expect(
+ await repairKnownPoisonedInstallDirBeforeWindow({
+ ...recoveryOptions(userDataPath, okRun),
+ isServeMode: true
+ })
+ ).toBe('skipped')
+ })
+
+ it('retires the marker when a later probe reports the install clean', () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-gate-'))
+ startWindowsInstallDirAclRepairIfPoisoned(
+ POISON_VERDICT,
+ recoveryOptions(userDataPath, (() => new Promise(() => undefined)) as Runner)
+ )
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+
+ resetWindowsInstallDirAclRecoveryForTest()
+ startWindowsInstallDirAclRepairIfPoisoned(
+ { status: 'ok', matchesPoisonSignature: false },
+ recoveryOptions(userDataPath, okRun)
+ )
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(false)
+ })
+
+ // An unreadable DACL is not evidence of health; forgetting the verdict there would
+ // hand the next launch straight back to the crash it already recorded.
+ it('keeps the marker when the probe could not read the DACL', () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-gate-'))
+ startWindowsInstallDirAclRepairIfPoisoned(
+ POISON_VERDICT,
+ recoveryOptions(userDataPath, (() => new Promise(() => undefined)) as Runner)
+ )
+ resetWindowsInstallDirAclRecoveryForTest()
+ startWindowsInstallDirAclRepairIfPoisoned(
+ { status: 'failed', reason: 'all-targets-unreadable' },
+ recoveryOptions(userDataPath, okRun)
+ )
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+ })
+
+ // The gate-timed-out ordering. The gate's budget is 20s while the tree grant's own cap is
+ // 120s, so icacls routinely outlives the gate: the window opens, and this launch's probe
+ // reads the tree POISONED while that repair is still in flight. When the orphaned icacls
+ // then claims success -- exit 0, and on a localized Windows no parsable failure summary to
+ // contradict it -- the claim must not outrank a reading taken after it was dispatched.
+ // Otherwise this launch deletes the poison marker that arms every later gate, un-suspects
+ // the tree so --in-process-gpu can engage, clears the safe-graphics marker, and tells the
+ // user their permissions are fixed.
+ it('does not let a timed-out gate repair outrank a poison reading taken after it', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-gate-timeout-'))
+ writeInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)
+ writeGpuFallbackMarker(
+ userDataPath,
+ { engagedAt: Date.now(), crashesInWindow: 3, userConfirmed: false },
+ GPU_ENV
+ )
+
+ let releaseIcacls: () => void = () => undefined
+ const stalled = new Promise((resolve) => {
+ releaseIcacls = resolve
+ })
+ let repairReported: () => void = () => undefined
+ const reported = new Promise((resolve) => {
+ repairReported = resolve
+ })
+
+ const mode = await repairKnownPoisonedInstallDirBeforeWindow({
+ ...recoveryOptions(userDataPath, async (spec) => {
+ await stalled
+ return okRun(spec)
+ }),
+ recordBreadcrumb: () => {
+ setTimeout(repairReported, 0)
+ return undefined
+ },
+ timeoutMs: 20
+ })
+ expect(mode).toBe('timeout')
+
+ // The window is open now, and this launch's own probe reads the tree still poisoned.
+ startWindowsInstallDirAclRepairIfPoisoned(POISON_VERDICT, recoveryOptions(userDataPath, okRun))
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+
+ releaseIcacls()
+ await reported
+
+ expect(isInstallDirAclSuspect()).toBe(true)
+ expect(describeInstallDirAclPoison()?.detail).toContain('could not repair them')
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+ expect(readActiveGpuFallbackMarker(userDataPath, GPU_ENV)).not.toBeNull()
+ })
+
+ // The other ordering of the same two events: the orphaned icacls exits 0 and clears the
+ // safe-graphics marker BEFORE the probe reads the tree still poisoned. The disproof must
+ // give the marker back, or the two orderings disagree about the same launch.
+ it('restores the safe-graphics marker when the probe disproves a timed-out gate repair', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-gate-timeout-restore-'))
+ writeInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)
+ writeGpuFallbackMarker(
+ userDataPath,
+ { engagedAt: Date.now(), crashesInWindow: 3, userConfirmed: false },
+ GPU_ENV
+ )
+
+ let releaseIcacls: () => void = () => undefined
+ const stalled = new Promise((resolve) => {
+ releaseIcacls = resolve
+ })
+ let repairReported: () => void = () => undefined
+ const reported = new Promise((resolve) => {
+ repairReported = resolve
+ })
+
+ const mode = await repairKnownPoisonedInstallDirBeforeWindow({
+ ...recoveryOptions(userDataPath, async (spec) => {
+ await stalled
+ return okRun(spec)
+ }),
+ recordBreadcrumb: () => {
+ setTimeout(repairReported, 0)
+ return undefined
+ },
+ timeoutMs: 20
+ })
+ expect(mode).toBe('timeout')
+
+ // The orphan claims success first; the claim is believed and takes the marker with it.
+ releaseIcacls()
+ await reported
+ expect(readActiveGpuFallbackMarker(userDataPath, GPU_ENV)).toBeNull()
+
+ // Then this launch's probe reads the tree still poisoned.
+ startWindowsInstallDirAclRepairIfPoisoned(POISON_VERDICT, recoveryOptions(userDataPath, okRun))
+
+ expect(readActiveGpuFallbackMarker(userDataPath, GPU_ENV)?.userConfirmed).toBe(false)
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+ expect(isInstallDirAclSuspect()).toBe(true)
+ })
+})
+
+// The repair marker matches whatever the outcome, so on its own 'marker-hit' cannot tell a
+// finished tree from one Orca gave up on. Both callers hold outstanding poison evidence —
+// this launch's probe reading, or the persisted marker that armed the gate — so a recorded
+// success never stands in for the repair, and 'marker-hit' only ever means budget spent.
+describe('a repair marker recording a completed repair', () => {
+ beforeEach(() => {
+ resetWindowsInstallDirAclProbeForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ resetWindowsInstallDirAclRecoveryForTest()
+ })
+
+ /** One launch: fresh module latches, then the gate runs against the userData on disk. */
+ async function gateLaunch(
+ userDataPath: string,
+ run: Runner
+ ): Promise>> {
+ resetWindowsInstallDirAclProbeForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ resetWindowsInstallDirAclRecoveryForTest()
+ return repairKnownPoisonedInstallDirBeforeWindow(recoveryOptions(userDataPath, run))
+ }
+
+ // The three-launch shape the gate exists for, and the one it used to disarm itself on:
+ // launch 1 repairs; the tree is re-poisoned (an installer, AV, or an icacls run that
+ // silently no-opped); launch 2's probe records the poison but Chromium FATALs before the
+ // repair can write its marker. Launch 3's gate then meets a poison marker and a repair
+ // marker claiming success. Treating that as 'repaired' ran no icacls, deleted the poison
+ // marker so no later gate ever fires again, un-suspected the tree so --in-process-gpu
+ // could engage, and told the user their permissions were fixed.
+ it('re-runs icacls when a poison marker outlives it', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-repaired-hit-'))
+
+ // Launch 1: the gate repairs the tree and retires the poison marker.
+ writeInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)
+ expect(await gateLaunch(userDataPath, okRun)).toBe('repaired')
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(false)
+
+ // Launch 2: the probe reads the tree as poisoned again; the process dies mid-repair,
+ // so the repair marker still records launch 1's success.
+ resetWindowsInstallDirAclRecoveryForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ startWindowsInstallDirAclRepairIfPoisoned(
+ POISON_VERDICT,
+ recoveryOptions(userDataPath, (() => new Promise(() => undefined)) as Runner)
+ )
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+
+ // Launch 3: the gate must repair, not congratulate itself on launch 1's work.
+ const spent: ProcessSpec[] = []
+ const mode = await gateLaunch(userDataPath, async (spec) => {
+ spent.push(spec)
+ return { code: 5, signal: null, stdout: '', stderr: 'Access is denied.', timedOut: false }
+ })
+ expect(mode).toBe('failed')
+ expect(spent.map((spec) => spec.args?.[2])).toEqual([
+ '*S-1-15-2-2:(OI)(CI)(RX)',
+ '*S-1-15-2-2:(RX)'
+ ])
+ expect(isInstallDirAclSuspect()).toBe(true)
+ expect(describeInstallDirAclPoison()?.detail).toContain('could not repair them')
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+ })
+
+ // The budget is what stops the retry above running forever; a spent one must still read
+ // as "Orca could not fix this", never as a repair it never made.
+ it('does not let the gate report a spent budget as a repair', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-gate-budget-'))
+ writeFileSync(
+ join(userDataPath, WINDOWS_INSTALL_DIR_ACL_REPAIR_MARKER_FILE),
+ JSON.stringify({
+ schemeVersion: WINDOWS_INSTALL_DIR_ACL_REPAIR_SCHEME_VERSION,
+ installDir: INSTALL_DIR,
+ appVersion: APP_VERSION,
+ attemptedAt: Date.now(),
+ outcome: 'repaired',
+ attempts: 3
+ })
+ )
+ writeInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)
+ const spent: ProcessSpec[] = []
+ const mode = await gateLaunch(userDataPath, async (spec) => {
+ spent.push(spec)
+ return okRun(spec)
+ })
+ expect(mode).toBe('marker-hit')
+ expect(spent).toHaveLength(0)
+ expect(isInstallDirAclSuspect()).toBe(true)
+ expect(isInstallDirAclRepairExhausted()).toBe(true)
+ expect(describeInstallDirAclPoison()?.detail).toContain('could not repair them')
+ // Still armed: nothing has proven this tree healthy, so a later launch still gates.
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+ })
+
+ // The probe reads the tree AFTER the pre-window gate has finished with it, so a signature
+ // still matching means the repair never landed however icacls exited.
+ it('is overruled by a probe that reads the tree poisoned after the gate repaired it', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-noop-icacls-'))
+ writeInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)
+ expect(
+ await repairKnownPoisonedInstallDirBeforeWindow(recoveryOptions(userDataPath, okRun))
+ ).toBe('repaired')
+
+ startWindowsInstallDirAclRepairIfPoisoned(POISON_VERDICT, recoveryOptions(userDataPath, okRun))
+
+ expect(isInstallDirAclSuspect()).toBe(true)
+ expect(isInstallDirAclRepairExhausted()).toBe(false)
+ expect(describeInstallDirAclPoison()?.detail).toContain('could not repair them')
+ // Re-armed: the next launch gates before it opens a window it cannot render.
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+ })
+
+ // The gate's 'repaired' is icacls's exit claim, not a reading of the tree — and the GPU
+ // children die 48-1373ms after window creation while the probe answers 0.9-3.0s in.
+ // Un-suspecting the tree on the claim alone opens exactly that interval to
+ // --in-process-gpu on a tree safe graphics cannot rescue.
+ it('keeps a gate-repaired tree suspect until this launch probe has read it', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-provisional-'))
+ writeInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)
+ expect(
+ await repairKnownPoisonedInstallDirBeforeWindow(recoveryOptions(userDataPath, okRun))
+ ).toBe('repaired')
+
+ // openMainWindow dispatches the probe: the reading is outstanding.
+ noteWindowsInstallDirAclProbePending()
+ expect(isInstallDirAclSuspect()).toBe(true)
+ // A probe that never answers releases at the grace window, like any pending verdict.
+ expect(isInstallDirAclSuspect(Date.now() + 15_000)).toBe(false)
+
+ // A clean reading corroborates the claim and releases immediately.
+ startWindowsInstallDirAclRepairIfPoisoned(
+ { status: 'ok', matchesPoisonSignature: false },
+ recoveryOptions(userDataPath, okRun)
+ )
+ expect(isInstallDirAclSuspect()).toBe(false)
+ })
+
+ // A disproved claim owes back everything it took on the false premise — the poison
+ // marker (above) and the safe-graphics marker, or the machine relaunches hardware
+ // accelerated into the re-armed gate and FATALs before that gate can finish.
+ it('restores the safe-graphics marker a disproved repair claim cleared', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-gpu-restore-'))
+ writeInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)
+ writeGpuFallbackMarker(
+ userDataPath,
+ { engagedAt: Date.now(), crashesInWindow: 3, userConfirmed: false },
+ GPU_ENV
+ )
+ expect(
+ await repairKnownPoisonedInstallDirBeforeWindow(recoveryOptions(userDataPath, okRun))
+ ).toBe('repaired')
+ // The claim was believed, so the marker went with it.
+ expect(readActiveGpuFallbackMarker(userDataPath, GPU_ENV)).toBeNull()
+
+ startWindowsInstallDirAclRepairIfPoisoned(POISON_VERDICT, recoveryOptions(userDataPath, okRun))
+
+ expect(readActiveGpuFallbackMarker(userDataPath, GPU_ENV)?.userConfirmed).toBe(false)
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+ })
+
+ // The opposite evidence: the probe has just READ this tree and found it poisoned, so a
+ // marker claiming success describes a tree that was re-poisoned, or an icacls run that
+ // silently no-opped. Reporting 'repaired' there runs no icacls, deletes the poison marker
+ // that arms the next launch's gate, un-suspects the tree so --in-process-gpu can engage,
+ // and tells the user their permissions are fixed.
+ it('re-runs icacls when a fresh probe verdict contradicts the repaired marker', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-repoisoned-'))
+ writeInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)
+ expect(
+ await repairKnownPoisonedInstallDirBeforeWindow(recoveryOptions(userDataPath, okRun))
+ ).toBe('repaired')
+
+ // Next launch: the gate is disarmed, and the probe reads the same tree as poisoned.
+ resetWindowsInstallDirAclRecoveryForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ const spent: ProcessSpec[] = []
+ const failing: Runner = async (spec) => {
+ spent.push(spec)
+ return { code: 5, signal: null, stdout: '', stderr: 'Access is denied.', timedOut: false }
+ }
+ await new Promise((resolve) => {
+ startWindowsInstallDirAclRepairIfPoisoned(POISON_VERDICT, {
+ ...recoveryOptions(userDataPath, failing),
+ recordBreadcrumb: () => {
+ setTimeout(resolve, 0)
+ return undefined
+ }
+ })
+ })
+
+ expect(spent.map((spec) => spec.args?.[2])).toEqual([
+ '*S-1-15-2-2:(OI)(CI)(RX)',
+ '*S-1-15-2-2:(RX)'
+ ])
+ expect(isInstallDirAclSuspect()).toBe(true)
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+ expect(describeInstallDirAclPoison()?.detail).toContain('could not repair them')
+ })
+
+ // The contradiction re-opens the budget, it does not remove it: a tree that has spent
+ // every attempt must not re-spawn icacls on every launch forever.
+ it('still stops at the attempt budget when the probe keeps reporting poison', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-repoisoned-budget-'))
+ writeFileSync(
+ join(userDataPath, WINDOWS_INSTALL_DIR_ACL_REPAIR_MARKER_FILE),
+ JSON.stringify({
+ schemeVersion: WINDOWS_INSTALL_DIR_ACL_REPAIR_SCHEME_VERSION,
+ installDir: INSTALL_DIR,
+ appVersion: APP_VERSION,
+ attemptedAt: Date.now(),
+ outcome: 'repaired',
+ attempts: 3
+ })
+ )
+ const spent: ProcessSpec[] = []
+ const run: Runner = async (spec) => {
+ spent.push(spec)
+ return okRun(spec)
+ }
+ await new Promise((resolve) => {
+ startWindowsInstallDirAclRepairIfPoisoned(POISON_VERDICT, {
+ ...recoveryOptions(userDataPath, run),
+ recordBreadcrumb: () => {
+ setTimeout(resolve, 0)
+ return undefined
+ }
+ })
+ })
+
+ expect(spent).toHaveLength(0)
+ // Nothing was repaired, so the user still gets the commands and the gate stays armed.
+ expect(isInstallDirAclSuspect()).toBe(true)
+ expect(describeInstallDirAclPoison()?.detail).toContain('could not repair them')
+ expect(hasInstallDirAclPoisonMarker(userDataPath, INSTALL_DIR, APP_VERSION)).toBe(true)
+ })
+})
+
+describe('a clean probe verdict', () => {
+ beforeEach(() => {
+ resetWindowsInstallDirAclProbeForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ resetWindowsInstallDirAclRecoveryForTest()
+ })
+
+ // The launch this covers: the repair budget is spent, so the gate can only report
+ // 'marker-hit' — and then the probe reads the tree and finds it healthy.
+ it('retires a verdict the gate could no longer act on', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-clean-'))
+ writeFileSync(
+ join(userDataPath, WINDOWS_INSTALL_DIR_ACL_REPAIR_MARKER_FILE),
+ JSON.stringify({
+ schemeVersion: WINDOWS_INSTALL_DIR_ACL_REPAIR_SCHEME_VERSION,
+ installDir: INSTALL_DIR,
+ appVersion: APP_VERSION,
+ attemptedAt: Date.now(),
+ outcome: 'failed',
+ attempts: 3
+ })
+ )
+ startWindowsInstallDirAclRepairIfPoisoned(
+ POISON_VERDICT,
+ recoveryOptions(userDataPath, (() => new Promise(() => undefined)) as Runner)
+ )
+ resetWindowsInstallDirAclRecoveryForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ expect(
+ await repairKnownPoisonedInstallDirBeforeWindow(recoveryOptions(userDataPath, okRun))
+ ).toBe('marker-hit')
+ expect(isInstallDirAclSuspect()).toBe(true)
+
+ startWindowsInstallDirAclRepairIfPoisoned(
+ { status: 'ok', matchesPoisonSignature: false },
+ recoveryOptions(userDataPath, okRun)
+ )
+ // Neither the driver fallback stays suppressed nor does the dialog accuse a healthy folder.
+ expect(isInstallDirAclSuspect()).toBe(false)
+ expect(describeInstallDirAclPoison()).toBeNull()
+ })
+
+ // The probe answers while the repair is still walking the tree: 'failed' from a
+ // repair with nothing left to fix must not re-accuse an install just read clean.
+ it('outranks a repair verdict that lands after it', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-clean-'))
+ const failing: Runner = async () => ({
+ code: 5,
+ signal: null,
+ stdout: '',
+ stderr: 'Access is denied.',
+ timedOut: false
+ })
+ let repairSettled = false
+ startWindowsInstallDirAclRepairIfPoisoned(POISON_VERDICT, {
+ ...recoveryOptions(userDataPath, failing),
+ recordBreadcrumb: () => {
+ repairSettled = true
+ return undefined
+ }
+ })
+ startWindowsInstallDirAclRepairIfPoisoned(
+ { status: 'ok', matchesPoisonSignature: false },
+ recoveryOptions(userDataPath, okRun)
+ )
+ await vi.waitFor(() => expect(repairSettled).toBe(true))
+ expect(isInstallDirAclSuspect()).toBe(false)
+ expect(describeInstallDirAclPoison()).toBeNull()
+ })
+})
+
+describe('the probe-pending grace window', () => {
+ beforeEach(() => {
+ resetWindowsInstallDirAclProbeForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ resetWindowsInstallDirAclRecoveryForTest()
+ })
+
+ // openMainWindow re-runs on every tray/second-instance reopen while the probe is
+ // once-per-process, so a re-arm would wait 15s on a verdict that already landed
+ // and drop every GPU child crash in between.
+ it('is armed by a dispatched probe only, so a reopen cannot re-arm it', async () => {
+ const probeArgs = {
+ platform: 'win32' as const,
+ installDir: INSTALL_DIR,
+ fileExists: () => false,
+ spawnFn: fakeIcaclsSpawn((target) => icaclsDacl(target, [RESTRICTED_PACKAGES_ACE])).spawnFn,
+ recordBreadcrumb: () => undefined
+ }
+ let settleVerdict: () => void = () => undefined
+ const verdict = new Promise((resolve) => (settleVerdict = resolve))
+ // Launch, wired exactly as main-window-controller wires it.
+ const dispatched = probeWindowsInstallDirAcl({
+ ...probeArgs,
+ onDone: (data) => {
+ startWindowsInstallDirAclRepairIfPoisoned(
+ data,
+ recoveryOptions(mkdtempSync(join(tmpdir(), 'orca-acl-rearm-')), okRun)
+ )
+ settleVerdict()
+ }
+ })
+ if (dispatched) {
+ noteWindowsInstallDirAclProbePending()
+ }
+ expect(dispatched).toBe(true)
+ expect(isInstallDirAclSuspect()).toBe(true)
+ await verdict
+ expect(isInstallDirAclSuspect()).toBe(false)
+
+ // Reopen: the probe declines, so nothing arms the grace window again.
+ const reopened = probeWindowsInstallDirAcl({ ...probeArgs, onDone: () => undefined })
+ if (reopened) {
+ noteWindowsInstallDirAclProbePending()
+ }
+ expect(reopened).toBe(false)
+ expect(isInstallDirAclSuspect()).toBe(false)
+ expect(isInstallDirAclSuspect(Date.now() + 14_000)).toBe(false)
+ })
+})
+
+describe('isBlockingInstallDirAclRepairInFlight', () => {
+ beforeEach(() => {
+ resetWindowsInstallDirAclProbeForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ resetWindowsInstallDirAclRecoveryForTest()
+ })
+
+ it('is false on a healthy machine and clears once the gate returns', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-inflight-'))
+ expect(isBlockingInstallDirAclRepairInFlight()).toBe(false)
+
+ startWindowsInstallDirAclRepairIfPoisoned(
+ POISON_VERDICT,
+ recoveryOptions(userDataPath, (() => new Promise(() => undefined)) as Runner)
+ )
+ resetWindowsInstallDirAclRecoveryForTest()
+ resetWindowsInstallDirAclRepairForTest()
+
+ let inFlightDuringRepair = false
+ const gate = repairKnownPoisonedInstallDirBeforeWindow({
+ ...recoveryOptions(userDataPath, async (spec) => {
+ inFlightDuringRepair = isBlockingInstallDirAclRepairInFlight()
+ return okRun(spec)
+ }),
+ timeoutMs: 5_000
+ })
+ expect(await gate).toBe('repaired')
+ expect(inFlightDuringRepair).toBe(true)
+ expect(isBlockingInstallDirAclRepairInFlight()).toBe(false)
+ })
+
+ // A second entry has no `onDone` coming, so waiting out the 20s budget for it
+ // would hold the window closed for nothing.
+ it('returns immediately when the once-per-process repair already ran', async () => {
+ const userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-inflight-'))
+ startWindowsInstallDirAclRepairIfPoisoned(POISON_VERDICT, recoveryOptions(userDataPath, okRun))
+ resetWindowsInstallDirAclRecoveryForTest()
+
+ const mode = await repairKnownPoisonedInstallDirBeforeWindow({
+ ...recoveryOptions(userDataPath, okRun),
+ timeoutMs: 30_000
+ })
+ expect(mode).toBe('skipped')
+ expect(isBlockingInstallDirAclRepairInFlight()).toBe(false)
+ })
+})
diff --git a/src/main/startup/windows-install-dir-acl-recovery.ts b/src/main/startup/windows-install-dir-acl-recovery.ts
index 0aa6870192e..c251bda9267 100644
--- a/src/main/startup/windows-install-dir-acl-recovery.ts
+++ b/src/main/startup/windows-install-dir-acl-recovery.ts
@@ -1,6 +1,17 @@
import { dirname } from 'node:path'
import type { CrashReportBreadcrumbData } from '../../shared/crash-reporting'
import { logStartupMilestone } from './startup-diagnostics'
+import {
+ clearGpuFallbackMarker,
+ readGpuFallbackMarker,
+ writeGpuFallbackMarker,
+ type GpuFallbackMarker
+} from './gpu-fallback-marker'
+import {
+ clearInstallDirAclPoisonMarker,
+ hasInstallDirAclPoisonMarker,
+ writeInstallDirAclPoisonMarker
+} from './windows-install-dir-acl-poison-marker'
import {
buildInstallDirAclRepairCommands,
isInstallDirAclPoisonVerdict,
@@ -25,10 +36,165 @@ export type WindowsInstallDirAclRecoveryOptions = Omit void>()
+
+function settleVerdictWaiters(): void {
+ // `wake` deletes only itself, which is safe to do on the entry being visited.
+ for (const wake of verdictWaiters) {
+ wake()
+ }
+ verdictWaiters.clear()
+}
export function resetWindowsInstallDirAclRecoveryForTest(): void {
poison = null
+ probePendingSince = null
+ installDirReadClean = false
+ installDirReadPoisonedMidRepair = false
+ gpuMarkerClearedByRepairClaim = null
+ blockingRepairInFlight = false
+ settleVerdictWaiters()
+}
+
+/** Call when the install-DACL probe is dispatched: its verdict is not in yet. */
+export function noteWindowsInstallDirAclProbePending(): void {
+ probePendingSince = Date.now()
+}
+
+/**
+ * Resolves when the probe's verdict lands, or when its grace window runs out.
+ * For callers that must not act on a suspicion the probe is about to withdraw.
+ */
+export function waitForInstallDirAclVerdict(now: number = Date.now()): Promise {
+ const remainingMs =
+ probePendingSince === null ? 0 : PROBE_VERDICT_GRACE_MS - (now - probePendingSince)
+ if (remainingMs <= 0) {
+ return Promise.resolve()
+ }
+ return new Promise((resolve) => {
+ const wake = (): void => {
+ clearTimeout(timer)
+ verdictWaiters.delete(wake)
+ resolve()
+ }
+ const timer = setTimeout(wake, remainingMs)
+ timer.unref?.()
+ verdictWaiters.add(wake)
+ })
+}
+
+/**
+ * True while a sandboxed-child death could be the install DACL rather than the
+ * graphics driver. Safe graphics does not rescue a poisoned tree — it still kills
+ * the renderer — and it removes the GPU child, erasing the sibling-death evidence
+ * that is the only way to recognise the shape in a crash report.
+ */
+export function isInstallDirAclSuspect(now: number = Date.now()): boolean {
+ if (installDirReadClean) {
+ return false
+ }
+ if (poison && poison.stage !== 'repaired') {
+ return true
+ }
+ // A 'repaired' stage is icacls's exit claim, not a reading of the tree — and the GPU
+ // children die 48-1373ms after window creation while the probe answers 0.9-3.0s in. So
+ // the claim stays provisional while this launch's probe is still out: the grace check
+ // below keeps the suspicion until the reading corroborates it or the window lapses.
+ return probePendingSince !== null && now - probePendingSince < PROBE_VERDICT_GRACE_MS
+}
+
+/**
+ * True while a repair for this tree is dispatched and has not reported yet.
+ *
+ * Why it is not the same question as `isInstallDirAclSuspect`: a suspect tree we are
+ * actively repairing is one a *future* launch can still be rescued on, so the safe-graphics
+ * marker earns its keep there — a launch Chromium FATALs mid-repair comes back software
+ * rendered, stops spawning the GPU children that trigger the FATAL, and lets the next gate
+ * run to completion. A terminal verdict has no such next step.
+ */
+export function isInstallDirAclRepairPending(): boolean {
+ return poison?.stage === 'pending'
+}
+
+/**
+ * True once the repair has nothing left to try for this install and version: `marker-hit`
+ * is reachable only through the spent attempt budget. The suspicion itself stands — the
+ * dialog still names the cause and the admin commands — but a caller that was *withholding*
+ * a recovery to give the repair first go has nothing left to wait for.
+ */
+export function isInstallDirAclRepairExhausted(): boolean {
+ return poison?.stage === 'marker-hit'
+}
+
+/** True while the pre-window gate is rewriting the very files a new renderer would load. */
+export function isBlockingInstallDirAclRepairInFlight(): boolean {
+ return blockingRepairInFlight
+}
+
+/** False when the once-per-process repair had already been dispatched, so no `onDone` is coming. */
+function startRepair(
+ installDir: string,
+ options: WindowsInstallDirAclRecoveryOptions,
+ onDone?: (result: WindowsInstallDirAclRepairResult) => void
+): boolean {
+ writeInstallDirAclPoisonMarker(options.userDataPath, installDir, options.appVersion)
+ const started = repairWindowsInstallDirPackageAcl({
+ ...options,
+ installDir,
+ // Every caller here holds outstanding poison evidence — this launch's probe reading, or
+ // the persisted marker that armed the gate — so a marker recording a completed repair
+ // describes a re-poisoned tree, or an icacls run that silently no-opped. It must not
+ // stand in for a repair. `marker-hit` therefore only ever means the budget is spent.
+ poisonEvidenceOutstanding: true,
+ onDone: (result) => {
+ // A tree read poisoned AFTER this repair was dispatched disproves its success claim,
+ // whatever icacls exited: the gate's budget can expire while the child runs on under
+ // its own, so the probe's reading is the later evidence. A clean reading since then
+ // retires it — there was nothing left to repair.
+ const claimDisproved =
+ result.mode === 'repaired' && installDirReadPoisonedMidRepair && !installDirReadClean
+ // A clean reading of the tree outranks this: there was nothing left to repair.
+ if (!installDirReadClean) {
+ poison = { installDir, stage: claimDisproved ? 'failed' : result.mode }
+ }
+ logStartupMilestone('install-dir-acl-repair-done', { mode: result.mode })
+ if (result.mode === 'repaired' && !claimDisproved) {
+ clearInstallDirAclPoisonMarker(options.userDataPath)
+ // The GPU child deaths were never a driver fault, so safe graphics — and the
+ // --in-process-gpu launch that hides the next crash's evidence — must not outlive the repair.
+ // Never a user-confirmed marker: "keep safe graphics" is a choice, not Orca's latch.
+ const gpuMarker = readGpuFallbackMarker(options.userDataPath)
+ if (gpuMarker?.userConfirmed === false) {
+ // Kept: a probe reading that later disproves this claim restores the marker,
+ // or the next launch relaunches hardware accelerated into the re-armed gate.
+ gpuMarkerClearedByRepairClaim = gpuMarker
+ clearGpuFallbackMarker(options.userDataPath)
+ }
+ }
+ if (result.mode === 'failed') {
+ console.warn('[win32-acl] install dir package ACL repair failed:', result.reason)
+ }
+ onDone?.(result)
+ }
+ })
+ if (started) {
+ poison = { installDir, stage: 'pending' }
+ }
+ return started
}
/** The probe's `onDone`: no-op unless the machine is in the reproduced state. */
@@ -36,22 +202,112 @@ export function startWindowsInstallDirAclRepairIfPoisoned(
data: CrashReportBreadcrumbData,
options: WindowsInstallDirAclRecoveryOptions
): void {
- if (!isInstallDirAclPoisonVerdict(data)) {
- return
+ // Cleared for every verdict, including an unreadable one that proves nothing: that
+ // releases a provisional 'repaired' claim early, but holding it would only move the
+ // same release to the grace-window expiry — an unreadable probe can never corroborate.
+ probePendingSince = null
+ try {
+ applyInstallDirAclProbeVerdict(data, options)
+ } finally {
+ // Only after the verdict is applied: a waiter wakes to re-read `isInstallDirAclSuspect()`.
+ settleVerdictWaiters()
}
- const installDir = options.installDir ?? dirname(process.execPath)
- poison = { installDir, stage: 'pending' }
- repairWindowsInstallDirPackageAcl({
- ...options,
- installDir,
- onDone: (result) => {
- poison = { installDir, stage: result.mode }
- logStartupMilestone('install-dir-acl-repair-done', { mode: result.mode })
- if (result.mode === 'failed') {
- console.warn('[win32-acl] install dir package ACL repair failed:', result.reason)
+}
+
+function applyInstallDirAclProbeVerdict(
+ data: CrashReportBreadcrumbData,
+ options: WindowsInstallDirAclRecoveryOptions
+): void {
+ if (!isInstallDirAclPoisonVerdict(data)) {
+ // Only a positive clean reading retires the verdict; an unreadable DACL proves nothing.
+ if (data.matchesPoisonSignature === false) {
+ clearInstallDirAclPoisonMarker(options.userDataPath)
+ installDirReadClean = true
+ // The reading corroborates any repair claim, so its marker clear stands.
+ gpuMarkerClearedByRepairClaim = null
+ // Keeping 'repaired' costs nothing and is what tells the user to reload; anything
+ // else would go on suppressing the driver fallback and accusing a healthy folder.
+ if (poison?.stage !== 'repaired') {
+ poison = null
}
}
- })
+ return
+ }
+ // The blocking pre-window gate still owns this launch's repair; restarting it would
+ // reset the verdict to 'pending' against a repair that can no longer report. The reading
+ // is kept, not dropped: it is later evidence than the repair's own exit code.
+ if (poison?.stage === 'pending') {
+ installDirReadPoisonedMidRepair = true
+ return
+ }
+ // This reading was taken after the gate finished, so it outranks the gate's own verdict:
+ // a tree that still matches the signature was never repaired, whatever icacls exited.
+ if (poison?.stage === 'repaired') {
+ poison = { installDir: poison.installDir, stage: 'failed' }
+ // The claim also cleared the safe-graphics marker; disproved, it owes that back, or
+ // the next launch relaunches hardware accelerated and FATALs before its gate can win.
+ const cleared = gpuMarkerClearedByRepairClaim
+ gpuMarkerClearedByRepairClaim = null
+ if (cleared) {
+ try {
+ writeGpuFallbackMarker(options.userDataPath, cleared, cleared)
+ } catch {
+ // Best effort: the re-armed poison marker below still gates the next launch.
+ }
+ }
+ }
+ // Re-writes the poison marker — re-arming the next launch's gate — even when the
+ // once-per-process latch means no icacls can run again this launch.
+ startRepair(options.installDir ?? dirname(process.execPath), options)
+}
+
+/**
+ * Pre-window gate for a machine a previous launch already found poisoned.
+ *
+ * Why blocking, and why only here: the probe is `setImmediate`-deferred and takes
+ * 0.9-3.0s on the affected hosts, while the renderer it has to save is spawned
+ * synchronously by `createMainWindow` and dies at init 48-1373ms in. The
+ * persisted verdict is what buys that knowledge for free — a healthy machine
+ * reads one absent file and pays nothing.
+ */
+export async function repairKnownPoisonedInstallDirBeforeWindow(
+ options: WindowsInstallDirAclRecoveryOptions & { timeoutMs?: number }
+): Promise<'not-marked' | 'skipped' | WindowsInstallDirAclRepairResult['mode'] | 'timeout'> {
+ if ((options.platform ?? process.platform) !== 'win32' || options.isServeMode === true) {
+ return 'skipped'
+ }
+ const installDir = options.installDir ?? dirname(process.execPath)
+ if (!hasInstallDirAclPoisonMarker(options.userDataPath, installDir, options.appVersion)) {
+ return 'not-marked'
+ }
+ logStartupMilestone('install-dir-acl-repair-blocking-start')
+ blockingRepairInFlight = true
+ try {
+ return await new Promise((resolve) => {
+ const timer = setTimeout(
+ () => resolve('timeout'),
+ options.timeoutMs ?? BLOCKING_REPAIR_BUDGET_MS
+ )
+ timer.unref?.()
+ // The marker is an earlier launch's DACL reading that nothing has retired, so a
+ // repair marker claiming success cannot stand in for the repair this launch owes.
+ const started = startRepair(installDir, options, (result) => {
+ clearTimeout(timer)
+ resolve(result.mode)
+ })
+ // No dispatch means no `onDone`, so waiting out the whole budget would buy nothing.
+ if (!started) {
+ clearTimeout(timer)
+ resolve('skipped')
+ }
+ })
+ } catch (error) {
+ // This sits in the critical path ahead of window creation; it must never throw into it.
+ console.warn('[win32-acl] blocking install dir ACL repair faulted:', error)
+ return 'skipped'
+ } finally {
+ blockingRepairInFlight = false
+ }
}
const CAUSE =
diff --git a/src/main/startup/windows-install-dir-acl-repair.win32.test.ts b/src/main/startup/windows-install-dir-acl-repair.win32.test.ts
new file mode 100644
index 00000000000..9a04aa1edef
--- /dev/null
+++ b/src/main/startup/windows-install-dir-acl-repair.win32.test.ts
@@ -0,0 +1,124 @@
+import { mkdirSync, mkdtempSync, writeFileSync } from 'node:fs'
+import { tmpdir } from 'node:os'
+import { join } from 'node:path'
+import { afterAll, beforeAll, describe, expect, it } from 'vitest'
+import { runProcess } from '../../shared/child-process/run-process'
+import { getIcaclsExePath } from '../win32-utils'
+import { removeTreeSync } from '../../shared/windows-transient-lock-removal'
+import {
+ probeWindowsInstallDirAcl,
+ resetWindowsInstallDirAclProbeForTest
+} from './windows-install-dir-acl-probe'
+import { writeInstallDirAclPoisonMarker } from './windows-install-dir-acl-poison-marker'
+import {
+ repairKnownPoisonedInstallDirBeforeWindow,
+ resetWindowsInstallDirAclRecoveryForTest
+} from './windows-install-dir-acl-recovery'
+import { resetWindowsInstallDirAclRepairForTest } from './windows-install-dir-package-acl-repair'
+
+/**
+ * The other half of the ACL proof: the unit tests fake icacls, and this one runs
+ * the real binary against a real poisoned tree on a real Windows box.
+ *
+ * Both are needed. `icacls /grant "*S-1-15-2-2:(OI)(CI)(RX)"` exits 0 and
+ * prints "Failed processing 0 files" while writing no ACE at all — a model of
+ * icacls cannot catch that, and it is the exact mistake that leaves the app dead.
+ *
+ * Runs only on win32; skipped elsewhere.
+ */
+const describeOnWindows = process.platform === 'win32' ? describe : describe.skip
+
+/** An unresolvable AppContainer SID, the shape the field hosts carry. */
+const ORPHAN_SID =
+ '*S-1-15-2-1111111111-2222222222-3333333333-4444444444-5555555555-6666666666-7777777777'
+const RESTRICTED_PACKAGES_NAME = /ALL RESTRICTED APPLICATION PACKAGES/i
+
+async function icacls(...args: string[]): Promise<{ code: number | null; out: string }> {
+ const result = await runProcess({ program: getIcaclsExePath(), args, timeoutMs: 30_000 })
+ return { code: result.code, out: `${result.stdout}\n${result.stderr}` }
+}
+
+/** Explicit DACL, inheritance off: what a shipped module carries, and why a root grant alone is not enough. */
+async function createProtectedFile(path: string): Promise {
+ writeFileSync(path, 'binary')
+ await icacls(path, '/inheritance:d')
+}
+
+describeOnWindows('install-dir package ACL repair against the real icacls', () => {
+ let installDir: string
+ let userDataPath: string
+ let moduleFile: string
+ let trapFile: string
+
+ beforeAll(async () => {
+ installDir = mkdtempSync(join(tmpdir(), 'orca-acl-live-'))
+ userDataPath = mkdtempSync(join(tmpdir(), 'orca-acl-live-ud-'))
+ mkdirSync(join(installDir, 'resources'), { recursive: true })
+ moduleFile = join(installDir, 'ffmpeg.dll')
+ trapFile = join(installDir, 'resources', 'trap.dll')
+ await createProtectedFile(moduleFile)
+ await createProtectedFile(trapFile)
+ // Poison: an orphan package ACE on the tree and on the module, no well-known grant.
+ await icacls(installDir, '/grant', `${ORPHAN_SID}:(OI)(CI)(RX)`)
+ await icacls(moduleFile, '/grant', `${ORPHAN_SID}:(RX)`)
+ await icacls(trapFile, '/grant', `${ORPHAN_SID}:(RX)`)
+ })
+
+ afterAll(() => {
+ // Why removeTreeSync: two icacls.exe children just rewrote DACLs on this tree, so a
+ // raw rmSync races handles Windows has not released and throws EPERM after the
+ // assertions already passed.
+ removeTreeSync(installDir)
+ removeTreeSync(userDataPath)
+ })
+
+ function probeVerdict(): Promise> {
+ resetWindowsInstallDirAclProbeForTest()
+ return new Promise((resolve) => {
+ probeWindowsInstallDirAcl({
+ installDir,
+ recordBreadcrumb: () => undefined,
+ onDone: (data) => resolve(data as Record)
+ })
+ })
+ }
+
+ // The trap, pinned against the real binary: this is the form that looks like it worked.
+ it('confirms an inheritance-flagged grant silently writes nothing to a file', async () => {
+ const flagged = await icacls(trapFile, '/grant', '*S-1-15-2-2:(OI)(CI)(RX)')
+ expect(flagged.code).toBe(0)
+ expect(flagged.out).toMatch(/Failed processing 0 files?/i)
+ const after = await icacls(trapFile)
+ expect(after.out).not.toMatch(RESTRICTED_PACKAGES_NAME)
+ })
+
+ it('repairs the tree before the window, and the grant lands on the module file', async () => {
+ expect((await probeVerdict()).matchesPoisonSignature).toBe(true)
+
+ resetWindowsInstallDirAclRecoveryForTest()
+ resetWindowsInstallDirAclRepairForTest()
+ // The state a launch that died mid-repair leaves behind.
+ writeInstallDirAclPoisonMarker(userDataPath, installDir, '1.4.196')
+
+ const startedAt = Date.now()
+ const mode = await repairKnownPoisonedInstallDirBeforeWindow({
+ installDir,
+ userDataPath,
+ appVersion: '1.4.196',
+ recordBreadcrumb: () => undefined
+ })
+ console.log(`[live-acl] blocking repair ${mode} in ${Date.now() - startedAt}ms`)
+ expect(mode).toBe('repaired')
+
+ // A directory grant is not enough: the file carries its own DACL.
+ expect((await icacls(moduleFile)).out).toMatch(RESTRICTED_PACKAGES_NAME)
+ // The /T pass must also reach a NESTED protected file — the shape app.asar.unpacked
+ // and node_modules actually have.
+ expect((await icacls(trapFile)).out).toMatch(RESTRICTED_PACKAGES_NAME)
+ // And the (OI)(CI) root grant exists so files a later update writes inherit it.
+ const updateFile = join(installDir, 'resources', 'added-by-update.dll')
+ writeFileSync(updateFile, 'binary')
+ expect((await icacls(updateFile)).out).toMatch(RESTRICTED_PACKAGES_NAME)
+ expect((await probeVerdict()).matchesPoisonSignature).toBe(false)
+ })
+})
diff --git a/src/main/startup/windows-install-dir-acl-startup-wiring.test.ts b/src/main/startup/windows-install-dir-acl-startup-wiring.test.ts
new file mode 100644
index 00000000000..c4175e87a06
--- /dev/null
+++ b/src/main/startup/windows-install-dir-acl-startup-wiring.test.ts
@@ -0,0 +1,56 @@
+import { readFileSync } from 'node:fs'
+import { join } from 'node:path'
+import { describe, expect, it } from 'vitest'
+
+/**
+ * The three call sites that make the repair real. Each is one line of wiring in a
+ * module whose import graph makes it untestable in-process; the behaviour each
+ * line depends on is driven for real in `windows-install-dir-acl-recovery.test.ts`,
+ * `gpu-lifecycle-install-dir-acl-guard.test.ts` and `focus-existing-window.test.ts`.
+ */
+
+function readSource(relativePath: string): string {
+ return readFileSync(join(process.cwd(), relativePath), 'utf8')
+}
+
+describe('install-dir ACL repair startup wiring', () => {
+ // The entire premise: a renderer must never be spawned onto a tree a previous
+ // launch recorded as poisoned before icacls has had its bounded chance at it.
+ it('awaits the pre-window gate before any window creation', () => {
+ const source = readSource('src/main/startup/main-process-runtime-launch.ts')
+ const launchStart = source.indexOf('export async function initializeMainProcessRuntimeLaunch(')
+ expect(launchStart).toBeGreaterThanOrEqual(0)
+ const launch = source.slice(launchStart)
+
+ const gateIndex = launch.indexOf('await repairKnownPoisonedInstallDirBeforeWindow(')
+ const winEarlyWindowIndex = launch.indexOf('startWindowsDesktopBeforeShellPathReady(')
+ const desktopLaunchIndex = launch.indexOf('await launchDesktopMode(')
+ expect(gateIndex).toBeGreaterThanOrEqual(0)
+ expect(winEarlyWindowIndex).toBeGreaterThan(gateIndex)
+ expect(desktopLaunchIndex).toBeGreaterThan(gateIndex)
+ })
+
+ // A 20s blank launch invites a second double-click, and `focusExistingMainWindow`
+ // opens a window whenever there is none and the app is ready.
+ it('holds the second-instance reopen while the gate owns the launch', () => {
+ const source = readSource('src/main/startup/main-window-actions.ts')
+ const start = source.indexOf('export function focusExistingWindow(')
+ const end = source.indexOf('\nexport function showMainWindowFromTray(', start)
+ expect(start).toBeGreaterThanOrEqual(0)
+ expect(end).toBeGreaterThan(start)
+ expect(source.slice(start, end)).toContain(
+ 'canOpenWindow: () => !isBlockingInstallDirAclRepairInFlight()'
+ )
+ })
+
+ // openMainWindow re-runs on every reopen while the probe is once-per-process, so
+ // arming the grace window unconditionally would drop GPU crashes on a healthy machine.
+ it('arms the probe grace window only for a dispatched probe', () => {
+ const source = readSource('src/main/startup/main-window-controller.ts')
+ const dispatchIndex = source.indexOf('const probeDispatched = probeWindowsInstallDirAcl(')
+ const armIndex = source.indexOf('noteWindowsInstallDirAclProbePending()')
+ expect(dispatchIndex).toBeGreaterThanOrEqual(0)
+ expect(armIndex).toBeGreaterThan(dispatchIndex)
+ expect(source.slice(dispatchIndex, armIndex)).toContain('if (probeDispatched) {')
+ })
+})
diff --git a/src/main/startup/windows-install-dir-package-acl-repair.test.ts b/src/main/startup/windows-install-dir-package-acl-repair.test.ts
index 65d11e9fcde..cc41388b811 100644
--- a/src/main/startup/windows-install-dir-package-acl-repair.test.ts
+++ b/src/main/startup/windows-install-dir-package-acl-repair.test.ts
@@ -135,7 +135,7 @@ describe('repairWindowsInstallDirPackageAcl', () => {
const second = fakeRunner()
const { result, data } = await repair({ userDataPath, run: second.run })
expect(second.specs).toHaveLength(0)
- expect(result).toEqual({ mode: 'marker-hit' })
+ expect(result).toEqual({ mode: 'marker-hit', alreadyRepaired: true })
expect(data.reason).toBe('marker-hit')
})
@@ -233,6 +233,42 @@ describe('repairWindowsInstallDirPackageAcl', () => {
expect(marker.outcome).toBe('failed')
})
+ // The bricking mechanism: a marker was written on failure and matched regardless of
+ // outcome, so one Defender-locked file or one timeout pinned the machine to
+ // 'marker-hit' — repair permanently skipped — for the life of that version.
+ it('retries a failed repair on later launches, then stops once the budget is spent', async () => {
+ const userDataPath = userDataDir()
+ const failing = fakeRunner(() => ({ code: 5, stderr: 'Access is denied.' }))
+ for (let attempt = 0; attempt < 3; attempt++) {
+ resetWindowsInstallDirAclRepairForTest()
+ expect((await repair({ userDataPath, run: failing.run })).result.mode).toBe('failed')
+ }
+ expect(failing.specs).toHaveLength(6)
+
+ resetWindowsInstallDirAclRepairForTest()
+ const spent = fakeRunner()
+ const { result } = await repair({ userDataPath, run: spent.run })
+ // Not alreadyRepaired: the budget ran out, so the tree is still poisoned.
+ expect(result).toEqual({ mode: 'marker-hit', alreadyRepaired: false })
+ expect(spent.specs).toHaveLength(0)
+ })
+
+ it('stops retrying immediately once a repair has succeeded', async () => {
+ const userDataPath = userDataDir()
+ resetWindowsInstallDirAclRepairForTest()
+ await repair({ userDataPath, run: fakeRunner(() => ({ code: 5 })).run })
+ resetWindowsInstallDirAclRepairForTest()
+ expect((await repair({ userDataPath })).result).toEqual({ mode: 'repaired' })
+
+ resetWindowsInstallDirAclRepairForTest()
+ const after = fakeRunner()
+ expect((await repair({ userDataPath, run: after.run })).result).toEqual({
+ mode: 'marker-hit',
+ alreadyRepaired: true
+ })
+ expect(after.specs).toHaveLength(0)
+ })
+
it('is a no-op off win32 and in serve mode', async () => {
const off = fakeRunner()
repairWindowsInstallDirPackageAcl({
diff --git a/src/main/startup/windows-install-dir-package-acl-repair.ts b/src/main/startup/windows-install-dir-package-acl-repair.ts
index 606edae417c..6bd6505a2cb 100644
--- a/src/main/startup/windows-install-dir-package-acl-repair.ts
+++ b/src/main/startup/windows-install-dir-package-acl-repair.ts
@@ -54,7 +54,8 @@ const TREE_GRANT_TIMEOUT_MS = 120_000
const FAILED_PROCESSING = /Failed processing (\d+) files?/i
export type WindowsInstallDirAclRepairResult =
- | { mode: 'marker-hit' }
+ /** `alreadyRepaired`: the marker records a completed repair, not an exhausted retry budget. */
+ | { mode: 'marker-hit'; alreadyRepaired: boolean }
| { mode: 'repaired' }
| { mode: 'failed'; reason: string; failedFileCount: number | null }
@@ -62,6 +63,14 @@ export type WindowsInstallDirAclRepairOptions = {
installDir?: string
platform?: NodeJS.Platform
isServeMode?: boolean
+ /**
+ * A DACL reading found this tree poisoned and nothing has read it clean since — this
+ * launch's probe, or a persisted poison marker from an earlier one. A marker claiming a
+ * completed repair therefore describes a tree that has since been re-poisoned, or an
+ * icacls run that silently no-opped: it stops outranking the reading. The attempt
+ * budget still bounds retries.
+ */
+ poisonEvidenceOutstanding?: boolean
/** Test seams. */
runProcessFn?: typeof runProcess
recordBreadcrumb?: typeof recordDurableCrashBreadcrumb
@@ -81,8 +90,17 @@ type RepairMarker = {
appVersion: string
attemptedAt: number
outcome: string
+ /** Absent on schemeVersion-1 markers written before the retry budget existed. */
+ attempts?: number
}
+// Why bounded rather than one-and-done: the failure modes are not all permanent.
+// A Defender-locked file, a timeout or a contended volume fails one launch and
+// succeeds the next, and pinning on the first failure leaves the machine blank
+// forever for that version. Three is enough to stop a standard-user Program Files
+// install — which can never win — from re-spawning icacls on every launch.
+const MAX_REPAIR_ATTEMPTS = 3
+
/**
* The probe's verdict is the only trigger: an orphan package ACE with no
* well-known package grant to satisfy it. A localized icacls prints those grants
@@ -106,31 +124,46 @@ function markerPath(userDataPath: string): string {
return join(userDataPath, WINDOWS_INSTALL_DIR_ACL_REPAIR_MARKER_FILE)
}
-function hasMarkerFor(args: WindowsInstallDirAclRepairArgs): boolean {
+/** The marker for this exact install and version, or null. */
+function readMarkerFor(args: WindowsInstallDirAclRepairArgs): Partial | null {
try {
const parsed = JSON.parse(readFileSync(markerPath(args.userDataPath), 'utf-8')) as
| Partial
| undefined
- return (
- parsed?.schemeVersion === WINDOWS_INSTALL_DIR_ACL_REPAIR_SCHEME_VERSION &&
- parsed.installDir === args.installDir &&
- parsed.appVersion === args.appVersion
- )
+ if (
+ parsed?.schemeVersion !== WINDOWS_INSTALL_DIR_ACL_REPAIR_SCHEME_VERSION ||
+ parsed.installDir !== args.installDir ||
+ parsed.appVersion !== args.appVersion
+ ) {
+ return null
+ }
+ return parsed
} catch {
- return false // missing or corrupt -> attempt again
+ return null // missing or corrupt -> attempt again
}
}
-// Why write it on failure too: a standard-user Program Files install can never
-// win, and re-spawning icacls on every launch forever buys nothing. Reinstall or
-// update changes the key and retries.
+function markerHitFor(args: WindowsInstallDirAclRepairArgs): { alreadyRepaired: boolean } | null {
+ const marker = readMarkerFor(args)
+ if (!marker) {
+ return null
+ }
+ if (marker.outcome === 'repaired' && args.poisonEvidenceOutstanding !== true) {
+ return { alreadyRepaired: true }
+ }
+ return (marker.attempts ?? 0) >= MAX_REPAIR_ATTEMPTS ? { alreadyRepaired: false } : null
+}
+
+// Why write it on failure too: re-spawning icacls on every launch forever buys
+// nothing, so failures spend the retry budget. Reinstall or update changes the key.
function writeMarker(args: WindowsInstallDirAclRepairArgs, outcome: string): void {
const marker: RepairMarker = {
schemeVersion: WINDOWS_INSTALL_DIR_ACL_REPAIR_SCHEME_VERSION,
installDir: args.installDir ?? '',
appVersion: args.appVersion,
attemptedAt: Date.now(),
- outcome
+ outcome,
+ attempts: (readMarkerFor(args)?.attempts ?? 0) + 1
}
if (!existsSync(args.userDataPath)) {
mkdirSync(args.userDataPath, { recursive: true })
@@ -185,9 +218,10 @@ async function runRepair(args: WindowsInstallDirAclRepairArgs): Promise {
let result: WindowsInstallDirAclRepairResult
let data: CrashReportBreadcrumbData
try {
- if (hasMarkerFor(resolved)) {
- result = { mode: 'marker-hit' }
- data = { status: 'skipped', reason: 'marker-hit' }
+ const markerHit = markerHitFor(resolved)
+ if (markerHit) {
+ result = { mode: 'marker-hit', alreadyRepaired: markerHit.alreadyRepaired }
+ data = { status: 'skipped', reason: 'marker-hit', alreadyRepaired: markerHit.alreadyRepaired }
} else {
const runner = args.runProcessFn ?? runProcess
const root = await runGrant(
@@ -257,13 +291,16 @@ export function resetWindowsInstallDirAclRepairForTest(): void {
* Fire-and-forget; returns before any spawn. Call only when the probe reported
* `matchesPoisonSignature`. win32 only, exempt in serve mode, and it must never
* throw into window creation.
+ *
+ * Returns whether THIS call dispatched the repair. A caller that waits on `onDone`
+ * would otherwise wait forever on the once-per-process latch.
*/
-export function repairWindowsInstallDirPackageAcl(args: WindowsInstallDirAclRepairArgs): void {
+export function repairWindowsInstallDirPackageAcl(args: WindowsInstallDirAclRepairArgs): boolean {
if ((args.platform ?? process.platform) !== 'win32' || args.isServeMode === true) {
- return
+ return false
}
if (repairStarted) {
- return
+ return false
}
repairStarted = true
try {
@@ -273,4 +310,5 @@ export function repairWindowsInstallDirPackageAcl(args: WindowsInstallDirAclRepa
} catch {
// Nothing left to report to that would not throw again.
}
+ return true
}
diff --git a/src/main/window/focus-existing-window.test.ts b/src/main/window/focus-existing-window.test.ts
index 85b422e02a4..babb9a15490 100644
--- a/src/main/window/focus-existing-window.test.ts
+++ b/src/main/window/focus-existing-window.test.ts
@@ -127,6 +127,41 @@ describe('focusExistingMainWindow', () => {
expect(timer.scheduledMs()).toEqual([])
})
+ // The blocking install-DACL repair holds the first window for up to 20s of blank
+ // screen, which is exactly when a user double-clicks the shortcut again. That
+ // second instance must not spawn a renderer onto a tree icacls is rewriting.
+ it('drops a reopen while another path must own the first window', () => {
+ const openWindow = vi.fn()
+
+ const result = focusExistingMainWindow({
+ app: makeFakeApp(),
+ getWindow: () => null,
+ openWindow,
+ canOpenWindow: () => false
+ })
+
+ expect(result).toBe('pending')
+ expect(openWindow).not.toHaveBeenCalled()
+ })
+
+ it('still focuses a window that already exists while reopening is held', () => {
+ const window = makeFakeWindow()
+ const openWindow = vi.fn()
+
+ const result = focusExistingMainWindow({
+ app: makeFakeApp(),
+ getWindow: () => window,
+ openWindow,
+ canOpenWindow: () => false,
+ platform: 'darwin',
+ setTimeout: makeTimer().setTimeout
+ })
+
+ expect(result).toBe('focused')
+ expect(openWindow).not.toHaveBeenCalled()
+ expect(window.calls.focus).toHaveBeenCalledTimes(1)
+ })
+
it('waits for normal startup when no window exists before app readiness', () => {
const openWindow = vi.fn()
diff --git a/src/main/window/focus-existing-window.ts b/src/main/window/focus-existing-window.ts
index 8c8ac85ca2e..4cadb49a743 100644
--- a/src/main/window/focus-existing-window.ts
+++ b/src/main/window/focus-existing-window.ts
@@ -9,6 +9,8 @@ export type FocusExistingMainWindowOptions = {
app: Pick
getWindow: () => BrowserWindow | null
openWindow: () => BrowserWindow
+ /** False while some other path must own the first window; the reopen is dropped, not queued. */
+ canOpenWindow?: () => boolean
platform?: NodeJS.Platform
setTimeout?: FocusTimer
warn?: (message: string, error?: unknown) => void
@@ -143,7 +145,7 @@ export function focusExistingMainWindow(
let openedWindow = false
if (!window || window.isDestroyed()) {
- if (!opts.app.isReady()) {
+ if (!opts.app.isReady() || opts.canOpenWindow?.() === false) {
return 'pending'
}
window = openWindowWithRetry(opts, platform, setTimer, 1)
From 8854b5ded52bbc45b695dc88bae476de850f4bb9 Mon Sep 17 00:00:00 2001
From: Neil <4138956+nwparker@users.noreply.github.com>
Date: Thu, 3 Sep 2026 22:03:05 -0700
Subject: [PATCH 19/49] perf(ssh): coalesce concurrent git.listWorktrees reads
(#18419)
* perf(ssh): coalesce concurrent pty.inspectProcess and git.listWorktrees reads
Both were the only reads in their provider class with no in-flight dedupe while
their siblings already had it. Route them through the existing
InFlightPromiseDedupe, keyed per (relay pty id, incarnation) and per repoPath,
scoped to the provider instance so two hosts never share an entry. The worktree
listing clears from invalidateGitReads(), and a signalled read keeps its own
request so one caller's abort cannot cancel its joiners' scan.
In-flight only, no TTL: the relay does answer inspectProcess from a 500ms
TTL-cached process table, but a client TTL would compound with it rather than
match it, so it is not a free win.
* perf(ssh): drop the inspectProcess half, ratchet per-read host observations
The `git.listWorktrees` dedupe ships unchanged. The `pty.inspectProcess` dedupe
is reverted: the host mints one `observationEpoch` per request and the pane
foreground reader commits it per read, so two overlapping probes sharing one
reply make the second read a stale replay and `admitRemoteForegroundEvidence`
rejects it -- a would-be `live` identity read becomes `unverifiable`. The pane
foreground tracker overlaps its own probes by design (cancel-and-reissue after
a 350 ms settle), so that path is reachable.
Adds a ratchet that fails when the dedupe returns, driving the real reader
through the real provider operations.
* fix(ssh): move the inspect ratchet to the provider it guards
The ratchet lived under src/renderer and imported
src/main/providers/ssh-pty-provider-rpc-operations, dragging the whole
main-process graph into config/tsconfig.tc.web.json (TS6307).
Split it: the request-counter ratchet moves next to the provider it
pins, and the renderer file keeps the why -- a shared host observation
degrades the second overlapping read to unverifiable -- against the real
reader with no cross-project import.
* docs(ssh): document the worktree-list coalescing contract
---
src/main/providers/ssh-git-read-provider.ts | 3 +-
.../ssh-git-worktree-list-dedupe.test.ts | 156 ++++++++++++++++++
.../providers/ssh-git-worktree-provider.ts | 28 +++-
...h-pty-inspect-observation-identity.test.ts | 68 ++++++++
.../ssh-pty-provider-rpc-operations.ts | 4 +
...round-inspect-observation-identity.test.ts | 95 +++++++++++
6 files changed, 349 insertions(+), 5 deletions(-)
create mode 100644 src/main/providers/ssh-git-worktree-list-dedupe.test.ts
create mode 100644 src/main/providers/ssh-pty-inspect-observation-identity.test.ts
create mode 100644 src/renderer/src/components/terminal-pane/pane-foreground-inspect-observation-identity.test.ts
diff --git a/src/main/providers/ssh-git-read-provider.ts b/src/main/providers/ssh-git-read-provider.ts
index 5cbcbc17b5a..cbf20271003 100644
--- a/src/main/providers/ssh-git-read-provider.ts
+++ b/src/main/providers/ssh-git-read-provider.ts
@@ -44,7 +44,8 @@ export class SshGitReadProvider {
}
}
- private invalidateGitReads(): void {
+ /** Overridden by subclasses that own additional read caches (worktree listings). */
+ protected invalidateGitReads(): void {
this.gitDiffReadDedupe.clear()
this.statusReadLeaseOwner.invalidate()
this.upstreamStatusReadOwner.invalidate()
diff --git a/src/main/providers/ssh-git-worktree-list-dedupe.test.ts b/src/main/providers/ssh-git-worktree-list-dedupe.test.ts
new file mode 100644
index 00000000000..4030dbe87e4
--- /dev/null
+++ b/src/main/providers/ssh-git-worktree-list-dedupe.test.ts
@@ -0,0 +1,156 @@
+/**
+ * Local repos coalesce concurrent `git worktree list` scans (`shareWorktreeScan`); the SSH path
+ * branched away from that and paid one relay round trip per independent caller (`worktrees:list`,
+ * `worktrees:listAll`, the space repo scan, provisioned-root adoption). These are call counters.
+ */
+import { describe, expect, it } from 'vitest'
+import { SshGitProvider } from './ssh-git-provider'
+import { createMockMux, type MockMultiplexer } from './ssh-git-provider-test-harness'
+
+const REPO_PATH = '/home/user/repo'
+
+const WORKTREES = [
+ { path: REPO_PATH, head: 'abc123', branch: 'main', isBare: false, isMainWorktree: true }
+]
+
+type Deferred = { resolve: (value: unknown) => void; reject: (error: unknown) => void }
+
+/** Holds `git.listWorktrees` open so overlap is deterministic; answers everything else at once. */
+function createPendingListMux(): { mux: MockMultiplexer; listDeferreds: Deferred[] } {
+ const mux = createMockMux()
+ const listDeferreds: Deferred[] = []
+ mux.request.mockImplementation((method: string) => {
+ if (method !== 'git.listWorktrees') {
+ return Promise.resolve(undefined)
+ }
+ return new Promise((resolve, reject) => {
+ listDeferreds.push({ resolve, reject })
+ })
+ })
+ return { mux, listDeferreds }
+}
+
+const flush = (): Promise => new Promise((resolve) => setTimeout(resolve, 0))
+
+function countListRequests(mux: MockMultiplexer): number {
+ return mux.request.mock.calls.filter((call) => call[0] === 'git.listWorktrees').length
+}
+
+describe('SSH git.listWorktrees in-flight dedupe', () => {
+ it('collapses concurrent listings of one repo into a single relay request', async () => {
+ const { mux, listDeferreds } = createPendingListMux()
+ const provider = new SshGitProvider('conn-1', mux as never)
+
+ const listings = Array.from({ length: 6 }, () => provider.listWorktrees(REPO_PATH))
+ await flush()
+
+ expect(countListRequests(mux)).toBe(1)
+ expect(mux.request).toHaveBeenCalledWith(
+ 'git.listWorktrees',
+ { repoPath: REPO_PATH },
+ { signal: undefined }
+ )
+
+ listDeferreds[0].resolve(WORKTREES)
+ expect(await Promise.all(listings)).toEqual(Array.from({ length: 6 }, () => WORKTREES))
+ })
+
+ it('does not share across repos or connections', async () => {
+ const { mux } = createPendingListMux()
+ const provider = new SshGitProvider('conn-1', mux as never)
+
+ void provider.listWorktrees(REPO_PATH)
+ void provider.listWorktrees('/home/user/other')
+ await flush()
+ expect(countListRequests(mux)).toBe(2)
+
+ const second = createPendingListMux()
+ void new SshGitProvider('conn-2', second.mux as never).listWorktrees(REPO_PATH)
+ await flush()
+
+ expect(countListRequests(second.mux)).toBe(1)
+ expect(countListRequests(mux)).toBe(2)
+ })
+
+ it('keeps a signalled listing on its own request', async () => {
+ const { mux } = createPendingListMux()
+ const provider = new SshGitProvider('conn-1', mux as never)
+
+ void provider.listWorktrees(REPO_PATH)
+ await flush()
+ const controller = new AbortController()
+ void provider.listWorktrees(REPO_PATH, { signal: controller.signal })
+ await flush()
+
+ expect(countListRequests(mux)).toBe(2)
+ expect(mux.request).toHaveBeenCalledWith(
+ 'git.listWorktrees',
+ { repoPath: REPO_PATH },
+ { signal: controller.signal }
+ )
+ })
+
+ it('re-requests after the shared listing settles instead of caching it', async () => {
+ const { mux, listDeferreds } = createPendingListMux()
+ const provider = new SshGitProvider('conn-1', mux as never)
+
+ const first = provider.listWorktrees(REPO_PATH)
+ await flush()
+ listDeferreds[0].resolve(WORKTREES)
+ await first
+
+ void provider.listWorktrees(REPO_PATH)
+ await flush()
+
+ expect(countListRequests(mux)).toBe(2)
+ })
+
+ it.each([
+ ['addWorktree', (p: SshGitProvider) => p.addWorktree(REPO_PATH, 'feature', '/home/user/feat')],
+ ['removeWorktree', (p: SshGitProvider) => p.removeWorktree('/home/user/feat')]
+ ])('invalidates the shared listing after %s', async (_name, mutate) => {
+ const { mux } = createPendingListMux()
+ const provider = new SshGitProvider('conn-1', mux as never)
+
+ void provider.listWorktrees(REPO_PATH)
+ await flush()
+ expect(countListRequests(mux)).toBe(1)
+
+ await mutate(provider)
+
+ // The catalog moved, so a joiner must not inherit the pre-mutation scan.
+ void provider.listWorktrees(REPO_PATH)
+ await flush()
+ expect(countListRequests(mux)).toBe(2)
+ })
+
+ it('shares a failed listing with its joiners and re-requests afterwards', async () => {
+ const { mux, listDeferreds } = createPendingListMux()
+ const provider = new SshGitProvider('conn-1', mux as never)
+
+ const listings = [provider.listWorktrees(REPO_PATH), provider.listWorktrees(REPO_PATH)]
+ await flush()
+ expect(countListRequests(mux)).toBe(1)
+
+ const failure = new Error('relay request failed')
+ listDeferreds[0].reject(failure)
+ await expect(listings[0]).rejects.toBe(failure)
+ await expect(listings[1]).rejects.toBe(failure)
+
+ void provider.listWorktrees(REPO_PATH)
+ await flush()
+ expect(countListRequests(mux)).toBe(2)
+ })
+
+ it('refuses an unauthoritative relay answer for every joiner (#14004)', async () => {
+ const { mux, listDeferreds } = createPendingListMux()
+ const provider = new SshGitProvider('conn-1', mux as never)
+
+ const listings = [provider.listWorktrees(REPO_PATH), provider.listWorktrees(REPO_PATH)]
+ await flush()
+ listDeferreds[0].resolve([])
+
+ await expect(listings[0]).rejects.toThrow()
+ await expect(listings[1]).rejects.toThrow()
+ })
+})
diff --git a/src/main/providers/ssh-git-worktree-provider.ts b/src/main/providers/ssh-git-worktree-provider.ts
index 8dae1413321..17e5b273de7 100644
--- a/src/main/providers/ssh-git-worktree-provider.ts
+++ b/src/main/providers/ssh-git-worktree-provider.ts
@@ -2,6 +2,7 @@ import type { GitStatusResult } from '../../shared/git-status-types'
import type { RemoveWorktreeResult } from '../../shared/worktree/create-types'
import type { GitWorktreeInfo } from '../../shared/worktree/types'
import { CapabilityProbeCache } from '../../shared/capability-probe-cache'
+import { InFlightPromiseDedupe, stableInFlightKey } from '../../shared/in-flight-promise-dedupe'
import { assertAuthoritativeWorktreeCatalog } from '../../shared/worktree/worktree-catalog-availability'
import { isJsonRpcMethodNotFoundError } from './ssh-git-relay-errors'
import { SshGitReviewHeadProvider } from './ssh-git-review-head-provider'
@@ -29,16 +30,35 @@ export class SshGitWorktreeProvider extends SshGitReviewHeadProvider {
private readonly worktreeIsCleanCapabilityCache = new CapabilityProbeCache<
typeof WORKTREE_IS_CLEAN_CAPABILITY
>(Number.POSITIVE_INFINITY)
+ // Scoped to this provider instance, so two SSH hosts never share an entry.
+ private readonly worktreeListDedupe = new InFlightPromiseDedupe()
+ protected override invalidateGitReads(): void {
+ super.invalidateGitReads()
+ this.worktreeListDedupe.clear()
+ }
+
+ /** Un-signalled reads of one repo coalesce onto the request already in flight; nothing is cached. */
async listWorktrees(
repoPath: string,
options?: { signal?: AbortSignal }
): Promise {
- const response = await this.mux.request(
- 'git.listWorktrees',
- { repoPath },
- { signal: options?.signal }
+ // Why: same rule as shareWorktreeScan — one caller's abort must not cancel the scan its
+ // joiners are still waiting on, so a signalled read keeps its own request.
+ if (options?.signal) {
+ return this.requestWorktreeList(repoPath, options.signal)
+ }
+ return this.worktreeListDedupe.run(stableInFlightKey(['listWorktrees', repoPath]), () =>
+ this.requestWorktreeList(repoPath)
)
+ }
+
+ /** The one real relay round trip a coalesced read's joiners all wait on. */
+ private async requestWorktreeList(
+ repoPath: string,
+ signal?: AbortSignal
+ ): Promise {
+ const response = await this.mux.request('git.listWorktrees', { repoPath }, { signal })
// Why (#14004): relays before this fix answered a failed worktree scan with `[]`. Mixed versions are
// normal, so refuse the shape here too — a Git repo always lists its own checkout.
return assertAuthoritativeWorktreeCatalog(response, repoPath)
diff --git a/src/main/providers/ssh-pty-inspect-observation-identity.test.ts b/src/main/providers/ssh-pty-inspect-observation-identity.test.ts
new file mode 100644
index 00000000000..4e0dae1ead1
--- /dev/null
+++ b/src/main/providers/ssh-pty-inspect-observation-identity.test.ts
@@ -0,0 +1,68 @@
+/**
+ * Ratchet (#18419): `pty.inspectProcess` must NOT be in-flight coalesced the way the sibling git
+ * reads in `SshGitReadProvider` are. The host mints one `observationEpoch` per request and the
+ * pane foreground reader commits that epoch per read, so a shared reply reads as a stale replay to
+ * the second reader to settle — see the companion renderer proof in
+ * `src/renderer/src/components/terminal-pane/pane-foreground-inspect-observation-identity.test.ts`.
+ * These are request counters, not timings.
+ */
+import { describe, expect, it, vi } from 'vitest'
+import { createSshPtyProviderRpcOperations } from './ssh-pty-provider-rpc-operations'
+
+const RELAY_PTY_ID = 'pty-1'
+const APP_PTY_ID = `ssh:conn-1@@${RELAY_PTY_ID}`
+const INCARNATION_ID = 'inc-1'
+
+/** Answers `pty.inspectProcess` with a fresh host observation per request, held open on demand. */
+function createInspectingOperations(): {
+ operations: ReturnType
+ request: ReturnType
+ resolvers: ((value: unknown) => void)[]
+} {
+ const resolvers: ((value: unknown) => void)[] = []
+ const request = vi.fn(() => new Promise((resolve) => resolvers.push(resolve)))
+ return {
+ operations: createSshPtyProviderRpcOperations({
+ mux: { request } as never,
+ toRelayPtyId: () => RELAY_PTY_ID
+ }),
+ request,
+ resolvers
+ }
+}
+
+const flush = (): Promise => new Promise((resolve) => setTimeout(resolve, 0))
+
+describe('SSH pty.inspectProcess observation identity', () => {
+ it('gives each overlapping probe of one pane+incarnation its own host observation', async () => {
+ const { operations, request, resolvers } = createInspectingOperations()
+
+ const first = operations.inspectProcess(APP_PTY_ID, { expectedIncarnationId: INCARNATION_ID })
+ const second = operations.inspectProcess(APP_PTY_ID, { expectedIncarnationId: INCARNATION_ID })
+ await flush()
+
+ expect(request).toHaveBeenCalledTimes(2)
+ resolvers[0]({ foregroundProcess: 'claude', observationEpoch: 1 })
+ resolvers[1]({ foregroundProcess: 'claude', observationEpoch: 2 })
+ // Each read settles on the observation minted for it, never a neighbour's.
+ expect(await first).toMatchObject({ observationEpoch: 1 })
+ expect(await second).toMatchObject({ observationEpoch: 2 })
+ })
+
+ it('does not share a failed probe with an overlapping one', async () => {
+ const { operations, request, resolvers } = createInspectingOperations()
+
+ const failing = operations.inspectProcess(APP_PTY_ID, { expectedIncarnationId: INCARNATION_ID })
+ const overlapping = operations.inspectProcess(APP_PTY_ID, {
+ expectedIncarnationId: INCARNATION_ID
+ })
+ await flush()
+
+ expect(request).toHaveBeenCalledTimes(2)
+ resolvers[0](Promise.reject(new Error('relay dropped the probe')))
+ resolvers[1]({ foregroundProcess: 'claude', observationEpoch: 1 })
+
+ await expect(failing).rejects.toThrow('relay dropped the probe')
+ expect(await overlapping).toMatchObject({ observationEpoch: 1 })
+ })
+})
diff --git a/src/main/providers/ssh-pty-provider-rpc-operations.ts b/src/main/providers/ssh-pty-provider-rpc-operations.ts
index 3f9953b9d1c..71ddbefce97 100644
--- a/src/main/providers/ssh-pty-provider-rpc-operations.ts
+++ b/src/main/providers/ssh-pty-provider-rpc-operations.ts
@@ -50,6 +50,10 @@ export function createSshPtyProviderRpcOperations({ mux, toRelayPtyId }: SshPtyP
const result = await mux.request('pty.getForegroundProcess', { id: toRelayPtyId(id) })
return result as string | null
},
+ // Do NOT in-flight coalesce this the way the sibling git reads are: the host mints one
+ // `observationEpoch` per request and the pane foreground reader commits it per read, so a
+ // shared reply reads as a stale replay and degrades a `live` identity read to `unverifiable`.
+ // Guarded by ssh-pty-inspect-observation-identity.test.ts; #17525 removes the poll.
inspectProcess: async (
id: string,
options?: { expectedIncarnationId?: string }
diff --git a/src/renderer/src/components/terminal-pane/pane-foreground-inspect-observation-identity.test.ts b/src/renderer/src/components/terminal-pane/pane-foreground-inspect-observation-identity.test.ts
new file mode 100644
index 00000000000..71719b11e4f
--- /dev/null
+++ b/src/renderer/src/components/terminal-pane/pane-foreground-inspect-observation-identity.test.ts
@@ -0,0 +1,95 @@
+/**
+ * Why `pty.inspectProcess` is not in-flight coalesced (#18419). The host mints one
+ * `observationEpoch` per request and this reader commits that epoch per read, so a reply shared by
+ * two overlapping probes reads as a stale replay to the second reader to settle and its would-be
+ * `live` identity read degrades to `unverifiable`. The pane foreground tracker overlaps its own
+ * probes on purpose (`cancelPendingRead` bumps the generation but lets the in-flight probe finish,
+ * then reissues after a 350 ms settle), so that path is reachable. The provider-side ratchet that
+ * fails if the dedupe returns lives in `src/main/providers/ssh-pty-inspect-observation-identity.test.ts`.
+ */
+import { describe, expect, it } from 'vitest'
+import { createPaneForegroundProcessReader } from './pane-foreground-process-reader'
+
+const CONNECTION_ID = 'conn-1'
+const RELAY_PTY_ID = 'pty-1'
+const APP_PTY_ID = `ssh:${CONNECTION_ID}@@${RELAY_PTY_ID}`
+const INCARNATION_ID = 'inc-1'
+
+/** One host scan per request => one epoch per request. */
+const hostObservation = (observationEpoch: number): unknown => ({
+ foregroundProcess: 'claude',
+ hasChildProcesses: true,
+ foregroundProcessEvidence: {
+ verdict: 'live',
+ processName: 'claude',
+ ptyId: RELAY_PTY_ID,
+ ptyIncarnationId: INCARNATION_ID,
+ authorityGeneration: 'gen-1',
+ observationEpoch,
+ capturedAgeMs: 0,
+ fence: {
+ platform: 'posix',
+ shellPid: 100,
+ shellStartTime: '1000',
+ tty: '/dev/pts/3',
+ foregroundPgid: 200
+ }
+ }
+})
+
+/** Holds every probe open so the tracker's supersede-and-reissue pair really overlaps. */
+function createOverlappingReader(replies: { shared: boolean }): {
+ readProcess: ReturnType
+ settle: (index: number) => void
+} {
+ const resolvers: ((value: unknown) => void)[] = []
+ return {
+ // One reader instance per pane, exactly as the foreground tracker holds it.
+ readProcess: createPaneForegroundProcessReader({
+ readForegroundProcess: () => new Promise((resolve) => resolvers.push(resolve)) as never,
+ isRemotePtyId: () => true,
+ getExpectedIncarnationId: () => INCARNATION_ID
+ }),
+ // `shared` models what an in-flight dedupe would do: every joiner gets one host observation.
+ settle: (index) => resolvers[index]?.(hostObservation(replies.shared ? 1 : index + 1))
+ }
+}
+
+const flush = (): Promise => new Promise((resolve) => setTimeout(resolve, 0))
+
+describe('pane foreground inspect observation identity', () => {
+ it('keeps a reissued read `live` when it overlaps the probe it superseded', async () => {
+ const { readProcess, settle } = createOverlappingReader({ shared: false })
+
+ // The tracker cancels the first read (generation bump) but lets it run to completion, then
+ // reissues after the settle window — so both are in flight against the same pane.
+ const superseded = readProcess(APP_PTY_ID, false)
+ const reissued = readProcess(APP_PTY_ID, false)
+ await flush()
+
+ // The superseded read's continuation commits its epoch first.
+ settle(0)
+ expect((await superseded).remoteEvidenceVerdict).toBe('live')
+
+ settle(1)
+ const result = await reissued
+ expect(result.remoteEvidenceVerdict).toBe('live')
+ expect(result.processName).toBe('claude')
+ })
+
+ it('degrades the second overlapping read to `unverifiable` when one observation is shared', async () => {
+ const { readProcess, settle } = createOverlappingReader({ shared: true })
+
+ const superseded = readProcess(APP_PTY_ID, false)
+ const reissued = readProcess(APP_PTY_ID, false)
+ await flush()
+
+ settle(0)
+ expect((await superseded).remoteEvidenceVerdict).toBe('live')
+
+ settle(1)
+ const result = await reissued
+ expect(result.remoteEvidenceVerdict).toBe('unverifiable')
+ expect(result.processName).toBeNull()
+ })
+})
From 7574ee8403b30dbe3142c114e6633228986ac0a6 Mon Sep 17 00:00:00 2001
From: iverJisty
Date: Fri, 4 Sep 2026 13:19:40 +0800
Subject: [PATCH 20/49] fix(ports): route the status-bar popover scan to the
workspace's host (#17048)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
* fix(ports): route the status-bar popover scan to the workspace's host
- PortsStatusSegment resolved its runtime target from the global active
runtime, so opening the popover on a paired-remote workspace scanned the
client OS and reported zero workspace ports
- Resolve the target from the active worktree's owner host, matching
PortsPanel, PortRow, and WorktreeCardPorts
- Add publishWorkspacePortScanForHost: store the host's scan under its own
key, then republish the aggregate through setWorkspacePortScanProjection
so a single-host refresh no longer drops every other host's ports
- Publish through the projection setter instead of setWorkspacePortScan,
which wrote the synthetic all-hosts key back into workspacePortScansByKey
and made the next merge fold the aggregate into itself (duplicate rows)
- Reuse the helper for the manual panel refresh and the post-stop refresh,
and share the aggregate key constant with WorkspacePortScanner
* test(ports): cover popover host routing and aggregate preservation
- PortsStatusSegment.host-routing: popover scans the active workspace's
owner host, keeps other hosts in the projection, publishes a failed scan
under its own host, and stays local when the workspace has no owner
- workspace-port-scan-publish: single-host key vs all-hosts projection, and
repeated publishes never accumulate duplicate rows
* fix(ports): surface a host whose port scan failed instead of dropping it
- The merged projection only carries unavailableReason when every host failed,
so one unreachable server read as "this workspace has no ports"
- Add getUnavailableWorkspacePortHosts: hosts that failed while another host
still answered, with the local host distinguished by a null environment id
- Show one notice per failed host in the popover, named by its runtime
environment or the local host label, above the surviving hosts' ports
- Reuse the existing scan-unavailable string so no catalog entry is added
* test(ports): prove the popover's own failed scan reaches the host notice
- Make the mocked store setters write back, so a publish and the notice that
reads it can no longer name different scan keys with every assertion green
- Cover open popover -> remote scan rejects -> notice names the host, the seam
the store-write and render-only tests each stopped short of
- Drop an assertion comment that claimed to prove port preservation when it
only exercised the render path
* fix(ports): review nits — single-write publish, colon-safe host keys, failure port retention, docstrings
* fix(ports): keep the popover count and body in agreement, name every failed host
- A failed scan retains the host's last-good ports, and the badge/header count
them; the notice now sits above the list instead of replacing it, so the
popover no longer claims N ports over an empty body.
- getUnavailableWorkspacePortHosts reports all-hosts-failed too, so total loss
of contact names each host instead of printing raw scan keys under platform
'unknown'.
- Scan keys parse to a discriminated host ref, so an unrecognised key is
'unknown' rather than silently blamed on the local machine.
- Extract useWorktreeRuntimeTarget for the four ports surfaces that hand-rolled
the same owner-settings spread.
* fix(ports): label the local host from the failed scan's platform, not the renderer's userAgent
A paired web client's browser is not the Orca host, so deriving 'Local Mac'
from navigator.userAgent mislabels a Linux host. Carry each failed scan's
own platform through the unavailable-host list instead.
* fix(ports): keep the Ports panel list under its failure notice too
The retained-ports change gave a failed scan both ports and an
unavailableReason, and the right-sidebar panel hid every section behind the
notice — stripping the stop and open actions for ports the status bar still
counts. Gate the sections on whether anything is left to list, matching the
popover, behind a testable predicate.
* fix(ports): let a retained-port failure keep its debounce grace period
The popover publishes the host's last-good ports alongside the failure reason
the moment its own scan fails. reconcileTransientPortScanFailures treated any
published result carrying unavailableReason as a spent grace period, so the
very next background poll replaced those ports with an empty unavailable scan
— the retention never survived one poll interval. Keep the grace while the
published result still has ports; the tolerance still clears them on schedule.
* fix(ports): prune stale hosts in the poll's single map write
A manual publish (the ports popover) can resolve after the host-set change
already pruned its key, re-adding it; the poll's per-key writes only ever added,
so a removed host kept its ports in the count and held a permanent unavailable
notice until the next host-set change. Publish the poll's already-pruned map in
one replaceWorkspacePortScans instead, which also collapses N per-host
notifications into one and drops any synthetic all-hosts key that leaked in.
* fix(ports): fail closed for direct SSH workspaces
---------
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
---
.../ports/WorkspacePortScanner.test.tsx | 31 ++
.../components/ports/WorkspacePortScanner.tsx | 37 +-
.../right-sidebar/PortsPanel.test.tsx | 53 +--
.../local-workspace-port-sections.test.ts | 30 ++
.../local-workspace-port-sections.ts | 20 +
.../local-workspace-ports-panel.tsx | 63 ++-
.../WorktreeCard.compact-hover.test.tsx | 4 +-
....compact-ports-hover-independence.test.tsx | 4 +-
.../sidebar/WorktreeCardPorts.test.tsx | 2 +-
.../components/sidebar/WorktreeCardPorts.tsx | 24 +-
.../PortsStatusSegment.host-routing.test.tsx | 407 ++++++++++++++++++
.../status-bar/PortsStatusSegment.tsx | 141 ++++--
.../ports-status-popover-rows.test.tsx | 5 +-
.../status-bar/ports-status-popover-rows.tsx | 26 +-
.../src/lib/workspace-port-actions.ts | 85 ++--
.../workspace-port-host-availability.test.ts | 148 +++++++
.../lib/workspace-port-host-availability.ts | 65 +++
.../lib/workspace-port-scan-debounce.test.ts | 32 ++
.../src/lib/workspace-port-scan-debounce.ts | 10 +-
.../lib/workspace-port-scan-publish.test.ts | 153 +++++++
.../src/runtime/runtime-client-target.ts | 15 +
.../runtime/use-worktree-runtime-target.ts | 16 +
22 files changed, 1185 insertions(+), 186 deletions(-)
create mode 100644 src/renderer/src/components/right-sidebar/local-workspace-port-sections.test.ts
create mode 100644 src/renderer/src/components/status-bar/PortsStatusSegment.host-routing.test.tsx
create mode 100644 src/renderer/src/lib/workspace-port-host-availability.test.ts
create mode 100644 src/renderer/src/lib/workspace-port-host-availability.ts
create mode 100644 src/renderer/src/lib/workspace-port-scan-publish.test.ts
create mode 100644 src/renderer/src/runtime/use-worktree-runtime-target.ts
diff --git a/src/renderer/src/components/ports/WorkspacePortScanner.test.tsx b/src/renderer/src/components/ports/WorkspacePortScanner.test.tsx
index be1c69cbac7..0ae53fee931 100644
--- a/src/renderer/src/components/ports/WorkspacePortScanner.test.tsx
+++ b/src/renderer/src/components/ports/WorkspacePortScanner.test.tsx
@@ -493,6 +493,37 @@ describe('WorkspacePortScanner', () => {
expect(useAppStore.getState().workspacePortScansByKey['environment:env-3:all']).toBeUndefined()
})
+ // Why: a manual publish (the ports popover) can resolve after the host-set
+ // change already pruned its key, re-adding it. Per-key writes never delete, so
+ // that removed host would otherwise hold its ports and a permanent
+ // unavailable notice until the next host-set change.
+ it('drops a stale host re-added after pruning on the next poll', async () => {
+ await act(async () => {
+ root?.render()
+ await flushPromises()
+ })
+
+ const staleKey = 'environment:env-removed:all'
+ act(() => {
+ const state = useAppStore.getState()
+ state.replaceWorkspacePortScans(
+ {
+ ...state.workspacePortScansByKey,
+ [staleKey]: { ...emptyScan, unavailableReason: 'gone' }
+ },
+ state.workspacePortScan
+ )
+ })
+ expect(useAppStore.getState().workspacePortScansByKey[staleKey]).toBeDefined()
+
+ await act(async () => {
+ vi.advanceTimersByTime(30_000)
+ await flushPromises()
+ })
+
+ expect(useAppStore.getState().workspacePortScansByKey[staleKey]).toBeUndefined()
+ })
+
it('clears ports immediately when the final worktree is removed', async () => {
runtimeEnvironmentCall.mockImplementation(({ method }) => {
if (method === 'workspacePorts.scan') {
diff --git a/src/renderer/src/components/ports/WorkspacePortScanner.tsx b/src/renderer/src/components/ports/WorkspacePortScanner.tsx
index f2805eb88a7..2b12bd86cd8 100644
--- a/src/renderer/src/components/ports/WorkspacePortScanner.tsx
+++ b/src/renderer/src/components/ports/WorkspacePortScanner.tsx
@@ -4,10 +4,11 @@ import { getHasAnyWorktreesFromState } from '@/store/selectors'
import { getActiveRuntimeTarget, type RuntimeClientTarget } from '@/runtime/runtime-rpc-client'
import {
mergeWorkspacePortScans,
- runtimeTargetForExecutionHostId,
+ WORKSPACE_PORT_ALL_HOSTS_SCAN_KEY,
scanWorkspacePortsForTarget,
workspacePortScanKeyForTarget
} from '@/lib/workspace-port-actions'
+import { runtimeTargetForExecutionHostId } from '@/runtime/runtime-client-target'
import { installWindowVisibilityInterval, isWindowVisible } from '@/lib/window-visibility-interval'
import {
reconcileTransientPortScanFailures,
@@ -41,7 +42,6 @@ export function WorkspacePortScanner({ enabled = true }: { enabled?: boolean }):
const setWorkspacePortScan = useAppStore((s) => s.setWorkspacePortScan)
const setWorkspacePortScanProjection = useAppStore((s) => s.setWorkspacePortScanProjection)
const replaceWorkspacePortScans = useAppStore((s) => s.replaceWorkspacePortScans)
- const setWorkspacePortScanForKey = useAppStore((s) => s.setWorkspacePortScanForKey)
const setWorkspacePortScanRefreshing = useAppStore((s) => s.setWorkspacePortScanRefreshing)
const inFlightRef = useRef | null>(null)
const generationRef = useRef(0)
@@ -124,40 +124,40 @@ export function WorkspacePortScanner({ enabled = true }: { enabled?: boolean }):
const activeTargetKeys = new Set(
allTargets.map((target) => workspacePortScanKeyForTarget(target))
)
+ const publishedScans = useAppStore.getState().workspacePortScansByKey
const reconciled = reconcileTransientPortScanFailures(
results,
- useAppStore.getState().workspacePortScansByKey,
+ publishedScans,
portScanDebounceRef.current,
WORKSPACE_PORT_SCAN_FAILURE_THRESHOLD,
activeTargetKeys
)
const scansByKey = Object.fromEntries(
- Object.entries(useAppStore.getState().workspacePortScansByKey).filter(([key]) =>
- activeTargetKeys.has(key)
- )
+ Object.entries(publishedScans).filter(([key]) => activeTargetKeys.has(key))
)
- let sourceChanged = false
+ // Why: a manual publish that lands after a host is pruned re-adds its key,
+ // and per-key writes never delete. Dropping the inactive keys here is what
+ // stops a removed host from holding a permanent unavailable notice.
+ let sourceChanged =
+ Object.keys(scansByKey).length !== Object.keys(publishedScans).length
for (const { key, result } of reconciled) {
sourceChanged ||= scansByKey[key] !== result
scansByKey[key] = result
- setWorkspacePortScanForKey(key, result)
}
const activeScan = scansByKey[scanKey]
const merged = mergeWorkspacePortScans(scansByKey)
const projectionKey =
allTargets.length > 1
- ? 'all-hosts:all'
+ ? WORKSPACE_PORT_ALL_HOSTS_SCAN_KEY
: activeScan
? scanKey
: workspacePortScanKeyForTarget(allTargets[0])
if (sourceChanged || useAppStore.getState().workspacePortScan?.key !== projectionKey) {
- setWorkspacePortScanProjection(
- merged
- ? {
- key: projectionKey,
- result: merged
- }
- : null
+ // Why: one store update for the whole poll — a large host set must not
+ // fan out a notification to every subscriber per host.
+ replaceWorkspacePortScans(
+ sourceChanged ? scansByKey : publishedScans,
+ merged ? { key: projectionKey, result: merged } : null
)
}
}
@@ -177,8 +177,7 @@ export function WorkspacePortScanner({ enabled = true }: { enabled?: boolean }):
hasWorktrees,
scanKey,
setWorkspacePortScan,
- setWorkspacePortScanProjection,
- setWorkspacePortScanForKey,
+ replaceWorkspacePortScans,
setWorkspacePortScanRefreshing
]
)
@@ -215,7 +214,7 @@ export function WorkspacePortScanner({ enabled = true }: { enabled?: boolean }):
: Object.fromEntries(retainedEntries)
const retainedProjection = mergeWorkspacePortScans(retainedScans)
const retainedProjectionKey =
- targetKeys.size > 1 ? 'all-hosts:all' : Object.keys(retainedScans)[0]
+ targetKeys.size > 1 ? WORKSPACE_PORT_ALL_HOSTS_SCAN_KEY : Object.keys(retainedScans)[0]
// Why: unchanged hosts stay visible while the replacement RPC runs; removed
// hosts and the old synthetic aggregate are excluded immediately.
const nextProjection =
diff --git a/src/renderer/src/components/right-sidebar/PortsPanel.test.tsx b/src/renderer/src/components/right-sidebar/PortsPanel.test.tsx
index ea1d85f15b8..1121a05ae64 100644
--- a/src/renderer/src/components/right-sidebar/PortsPanel.test.tsx
+++ b/src/renderer/src/components/right-sidebar/PortsPanel.test.tsx
@@ -451,25 +451,26 @@ describe('PortsPanel runtime routing', () => {
})
it('returns post-stop refresh failures without throwing', async () => {
- const setWorkspacePortScan = vi.fn()
+ const replaceWorkspacePortScans = vi.fn()
const setWorkspacePortScanRefreshing = vi.fn()
localScan.mockRejectedValueOnce(new Error('scan failed'))
await expect(
refreshWorkspacePortScanAfterStop({
runtimeTarget: { kind: 'local' },
- setWorkspacePortScan: setWorkspacePortScan as never,
+ replaceWorkspacePortScans: replaceWorkspacePortScans as never,
+ getWorkspacePortScansByKey: () => ({}),
setWorkspacePortScanRefreshing: setWorkspacePortScanRefreshing as never
})
).resolves.toEqual({ ok: false, reason: 'scan failed' })
- expect(setWorkspacePortScan).not.toHaveBeenCalled()
+ expect(replaceWorkspacePortScans).not.toHaveBeenCalled()
expect(setWorkspacePortScanRefreshing).toHaveBeenNthCalledWith(1, true)
expect(setWorkspacePortScanRefreshing).toHaveBeenNthCalledWith(2, false)
})
it('ignores settled remote post-stop refresh failures after updating state', async () => {
- const setWorkspacePortScan = vi.fn()
+ const replaceWorkspacePortScans = vi.fn()
const setWorkspacePortScanRefreshing = vi.fn()
const firstScan = { ...emptyScan, scannedAt: 2 }
let scanCalls = 0
@@ -500,7 +501,8 @@ describe('PortsPanel runtime routing', () => {
await expect(
refreshWorkspacePortScanAfterStop({
runtimeTarget: { kind: 'environment', environmentId: 'env-1' },
- setWorkspacePortScan: setWorkspacePortScan as never,
+ replaceWorkspacePortScans: replaceWorkspacePortScans as never,
+ getWorkspacePortScansByKey: () => ({}),
setWorkspacePortScanRefreshing: setWorkspacePortScanRefreshing as never
})
).resolves.toEqual({ ok: true })
@@ -510,18 +512,20 @@ describe('PortsPanel runtime routing', () => {
'workspacePorts.scan',
'workspacePorts.scan'
])
- expect(setWorkspacePortScan).toHaveBeenCalledTimes(1)
- expect(setWorkspacePortScan).toHaveBeenCalledWith({
- key: 'environment:env-1:all',
- result: firstScan
- })
+ expect(replaceWorkspacePortScans).toHaveBeenCalledTimes(1)
+ expect(replaceWorkspacePortScans).toHaveBeenCalledWith(
+ { 'environment:env-1:all': firstScan },
+ {
+ key: 'environment:env-1:all',
+ result: firstScan
+ }
+ )
expect(setWorkspacePortScanRefreshing).toHaveBeenNthCalledWith(1, true)
expect(setWorkspacePortScanRefreshing).toHaveBeenNthCalledWith(2, false)
})
it('preserves an all-host projection after refreshing one host post-stop', async () => {
- const setWorkspacePortScan = vi.fn()
- const setWorkspacePortScanForKey = vi.fn()
+ const replaceWorkspacePortScans = vi.fn()
const setWorkspacePortScanRefreshing = vi.fn()
const localPort: WorkspacePort = { ...workspacePort, id: 'local-port', port: 5173 }
const refreshedRemotePort: WorkspacePort = {
@@ -571,23 +575,24 @@ describe('PortsPanel runtime routing', () => {
await expect(
refreshWorkspacePortScanAfterStop({
runtimeTarget: { kind: 'environment', environmentId: 'env-1' },
- setWorkspacePortScan: setWorkspacePortScan as never,
- setWorkspacePortScanForKey: setWorkspacePortScanForKey as never,
+ replaceWorkspacePortScans: replaceWorkspacePortScans as never,
getWorkspacePortScansByKey: () => ({ 'local:all': localHostScan }),
setWorkspacePortScanRefreshing: setWorkspacePortScanRefreshing as never
})
).resolves.toEqual({ ok: true })
- expect(setWorkspacePortScanForKey).toHaveBeenCalledWith('environment:env-1:all', remoteHostScan)
- expect(setWorkspacePortScan).toHaveBeenLastCalledWith({
- key: 'all-hosts:all',
- result: expect.objectContaining({
- ports: expect.arrayContaining([
- expect.objectContaining({ port: 5173 }),
- expect.objectContaining({ port: 3000 })
- ])
- })
- })
+ expect(replaceWorkspacePortScans).toHaveBeenLastCalledWith(
+ { 'local:all': localHostScan, 'environment:env-1:all': remoteHostScan },
+ {
+ key: 'all-hosts:all',
+ result: expect.objectContaining({
+ ports: expect.arrayContaining([
+ expect.objectContaining({ port: 5173 }),
+ expect.objectContaining({ port: 3000 })
+ ])
+ })
+ }
+ )
expect(scanCalls).toBe(2)
})
diff --git a/src/renderer/src/components/right-sidebar/local-workspace-port-sections.test.ts b/src/renderer/src/components/right-sidebar/local-workspace-port-sections.test.ts
new file mode 100644
index 00000000000..68733927ee7
--- /dev/null
+++ b/src/renderer/src/components/right-sidebar/local-workspace-port-sections.test.ts
@@ -0,0 +1,30 @@
+import { describe, expect, it } from 'vitest'
+import { shouldShowLocalWorkspacePortSections } from './local-workspace-port-sections'
+
+const empty = { activePorts: [], otherWorkspacePorts: [], externalPorts: [] }
+
+describe('shouldShowLocalWorkspacePortSections', () => {
+ it('shows the sections whenever the scan succeeded', () => {
+ expect(shouldShowLocalWorkspacePortSections(null, empty)).toBe(true)
+ expect(shouldShowLocalWorkspacePortSections({}, empty)).toBe(true)
+ })
+
+ // Why: a failed scan keeps the host's last-good ports, and the status bar
+ // still counts and lists them — hiding the sections here would strip the
+ // stop and open actions for ports the user can still see elsewhere.
+ it.each([
+ ['activePorts', { ...empty, activePorts: [{}] }],
+ ['otherWorkspacePorts', { ...empty, otherWorkspacePorts: [{}] }],
+ ['externalPorts', { ...empty, externalPorts: [{}] }]
+ ])('keeps the sections when a failed scan retained %s', (_section, sections) => {
+ expect(shouldShowLocalWorkspacePortSections({ unavailableReason: 'dropped' }, sections)).toBe(
+ true
+ )
+ })
+
+ it('lets the notice stand alone when a failed scan has nothing left to list', () => {
+ expect(shouldShowLocalWorkspacePortSections({ unavailableReason: 'dropped' }, empty)).toBe(
+ false
+ )
+ })
+})
diff --git a/src/renderer/src/components/right-sidebar/local-workspace-port-sections.ts b/src/renderer/src/components/right-sidebar/local-workspace-port-sections.ts
index d968bb85144..2a0eadf380a 100644
--- a/src/renderer/src/components/right-sidebar/local-workspace-port-sections.ts
+++ b/src/renderer/src/components/right-sidebar/local-workspace-port-sections.ts
@@ -35,6 +35,26 @@ export function getLocalWorkspacePortSections(
}
}
+/**
+ * Whether the panel still renders its port sections under a failure notice.
+ * Why: a failed scan retains the host's last-good ports, so hiding every
+ * section would drop the stop and open actions for ports the status bar still
+ * counts and lists.
+ */
+export function shouldShowLocalWorkspacePortSections(
+ scan: { unavailableReason?: string } | null | undefined,
+ sections: { activePorts: unknown[]; otherWorkspacePorts: unknown[]; externalPorts: unknown[] }
+): boolean {
+ if (!scan?.unavailableReason) {
+ return true
+ }
+ return (
+ sections.activePorts.length > 0 ||
+ sections.otherWorkspacePorts.length > 0 ||
+ sections.externalPorts.length > 0
+ )
+}
+
function workspacePortAsExternal(port: WorkspacePort & { kind: 'workspace' }): WorkspacePort {
return {
id: port.id,
diff --git a/src/renderer/src/components/right-sidebar/local-workspace-ports-panel.tsx b/src/renderer/src/components/right-sidebar/local-workspace-ports-panel.tsx
index 0955e89031b..1737e7da487 100644
--- a/src/renderer/src/components/right-sidebar/local-workspace-ports-panel.tsx
+++ b/src/renderer/src/components/right-sidebar/local-workspace-ports-panel.tsx
@@ -4,11 +4,11 @@ import { toast } from 'sonner'
import { useAppStore } from '@/store'
import { useActiveWorktree, useRepoById } from '@/store/selectors'
import { cn } from '@/lib/utils'
-import { getActiveRuntimeTarget } from '@/runtime/runtime-rpc-client'
-import { getRuntimeEnvironmentIdForWorktree } from '@/lib/worktree-runtime-owner'
+import { useWorktreeRuntimeTarget } from '@/runtime/use-worktree-runtime-target'
import {
killWorkspacePortForTarget,
openWorkspacePortInBrowser,
+ publishWorkspacePortScanForHost,
refreshWorkspacePortScanAfterStop,
resolvePortOpenInOrcaBrowser,
scanWorkspacePortsForTarget,
@@ -19,10 +19,14 @@ import { Button } from '@/components/ui/button'
import { Tooltip, TooltipContent, TooltipTrigger } from '@/components/ui/tooltip'
import type { WorkspacePort } from '../../../../shared/workspace-ports'
import { translate } from '@/i18n/i18n'
-import { getLocalWorkspacePortSections } from './local-workspace-port-sections'
+import {
+ getLocalWorkspacePortSections,
+ shouldShowLocalWorkspacePortSections
+} from './local-workspace-port-sections'
import { LocalPortSection } from './local-port-section'
import { LocalPortDetailsDialog } from './local-port-details-dialog'
+/** Right-sidebar Ports panel scoped to the active workspace's owner host. */
export function LocalWorkspacePortsPanel({ isVisible }: { isVisible: boolean }): React.JSX.Element {
const activeWorktree = useActiveWorktree()
const activeRepo = useRepoById(activeWorktree?.repoId ?? null)
@@ -31,8 +35,7 @@ export function LocalWorkspacePortsPanel({ isVisible }: { isVisible: boolean }):
const setRemoteBrowserPageHandle = useAppStore((s) => s.setRemoteBrowserPageHandle)
const scansByKey = useAppStore((s) => s.workspacePortScansByKey)
const refreshing = useAppStore((s) => s.workspacePortScanRefreshing)
- const setWorkspacePortScan = useAppStore((s) => s.setWorkspacePortScan)
- const setWorkspacePortScanForKey = useAppStore((s) => s.setWorkspacePortScanForKey)
+ const replaceWorkspacePortScans = useAppStore((s) => s.replaceWorkspacePortScans)
const setWorkspacePortScanRefreshing = useAppStore((s) => s.setWorkspacePortScanRefreshing)
const [detailsPort, setDetailsPort] = useState(null)
const [collapsedSections, setCollapsedSections] = useState>({
@@ -40,26 +43,24 @@ export function LocalWorkspacePortsPanel({ isVisible }: { isVisible: boolean }):
external: true
})
- const runtimeTarget = useMemo(() => {
- const activeRuntimeEnvironmentId = getRuntimeEnvironmentIdForWorktree(
- useAppStore.getState(),
- activeWorktree?.id
- )
- // Why: the Ports panel acts on the active workspace; use that workspace's
- // host owner even if the sidebar is focused elsewhere.
- return getActiveRuntimeTarget({ ...settings, activeRuntimeEnvironmentId })
- }, [activeWorktree?.id, settings])
- const scanKey = `${workspacePortRuntimeTargetKey(runtimeTarget)}:all`
+ // Why: the Ports panel acts on the active workspace; use that workspace's
+ // host owner even if the sidebar is focused elsewhere.
+ const runtimeTarget = useWorktreeRuntimeTarget(activeWorktree?.id)
+ const scanKey = runtimeTarget ? `${workspacePortRuntimeTargetKey(runtimeTarget)}:all` : null
const refresh = useCallback(() => {
- if (!activeRepo) {
+ if (!activeRepo || !runtimeTarget || !scanKey) {
return Promise.resolve()
}
setWorkspacePortScanRefreshing(true)
const promise = scanWorkspacePortsForTarget(runtimeTarget)
.then((nextScan) => {
- setWorkspacePortScanForKey(scanKey, nextScan)
- setWorkspacePortScan({ key: scanKey, result: nextScan })
+ publishWorkspacePortScanForHost({
+ scanKey,
+ scan: nextScan,
+ replaceWorkspacePortScans,
+ getWorkspacePortScansByKey: () => useAppStore.getState().workspacePortScansByKey
+ })
})
.catch((error) => {
const message = error instanceof Error ? error.message : String(error)
@@ -86,14 +87,13 @@ export function LocalWorkspacePortsPanel({ isVisible }: { isVisible: boolean }):
activeRepo,
runtimeTarget,
scanKey,
- setWorkspacePortScan,
- setWorkspacePortScanForKey,
+ replaceWorkspacePortScans,
setWorkspacePortScanRefreshing
])
// Why: WorkspacePortScanner already owns the 30s all-worktree poll. The
// panel scopes that shared result instead of starting a second scan loop.
- const displayScan = isVisible ? (scansByKey[scanKey] ?? null) : null
+ const displayScan = isVisible && scanKey ? (scansByKey[scanKey] ?? null) : null
const toggleSection = useCallback((sectionId: string) => {
setCollapsedSections((current) => ({ ...current, [sectionId]: !current[sectionId] }))
@@ -122,8 +122,7 @@ export function LocalWorkspacePortsPanel({ isVisible }: { isVisible: boolean }):
)
const refreshResult = await refreshWorkspacePortScanAfterStop({
runtimeTarget,
- setWorkspacePortScan,
- setWorkspacePortScanForKey,
+ replaceWorkspacePortScans,
getWorkspacePortScansByKey: () => useAppStore.getState().workspacePortScansByKey,
setWorkspacePortScanRefreshing
})
@@ -139,13 +138,7 @@ export function LocalWorkspacePortsPanel({ isVisible }: { isVisible: boolean }):
)
}
},
- [
- activeRepo,
- runtimeTarget,
- setWorkspacePortScan,
- setWorkspacePortScanForKey,
- setWorkspacePortScanRefreshing
- ]
+ [activeRepo, runtimeTarget, replaceWorkspacePortScans, setWorkspacePortScanRefreshing]
)
const handleOpenPortInBrowser = useCallback(
@@ -181,6 +174,12 @@ export function LocalWorkspacePortsPanel({ isVisible }: { isVisible: boolean }):
[activeRepo?.id, activeWorktree?.id, displayScan]
)
+ const showPortSections = shouldShowLocalWorkspacePortSections(displayScan, {
+ activePorts,
+ otherWorkspacePorts,
+ externalPorts
+ })
+
if (!activeRepo) {
return (
+ )
+}
diff --git a/src/renderer/src/components/status-bar/ports-status-popover-rows.test.tsx b/src/renderer/src/components/status-bar/ports-status-popover-rows.test.tsx
index 47c86d89c57..ba926909272 100644
--- a/src/renderer/src/components/status-bar/ports-status-popover-rows.test.tsx
+++ b/src/renderer/src/components/status-bar/ports-status-popover-rows.test.tsx
@@ -17,8 +17,7 @@ const {
settings: { openLinksInApp: true },
createBrowserTab: vi.fn(),
setRemoteBrowserPageHandle: vi.fn(),
- setWorkspacePortScan: vi.fn(),
- setWorkspacePortScanForKey: vi.fn(),
+ replaceWorkspacePortScans: vi.fn(),
setWorkspacePortScanRefreshing: vi.fn(),
recordFeatureInteraction: vi.fn(),
workspacePortScansByKey: {}
@@ -46,7 +45,7 @@ vi.mock('@/lib/worktree-activation', () => ({
}))
vi.mock('@/lib/worktree-runtime-owner', () => ({
- getRuntimeEnvironmentIdForWorktree: () => null
+ getExecutionHostIdForWorktree: () => 'local'
}))
vi.mock('@/runtime/runtime-rpc-client', () => ({
diff --git a/src/renderer/src/components/status-bar/ports-status-popover-rows.tsx b/src/renderer/src/components/status-bar/ports-status-popover-rows.tsx
index c6d7eb41625..4c07c44ed21 100644
--- a/src/renderer/src/components/status-bar/ports-status-popover-rows.tsx
+++ b/src/renderer/src/components/status-bar/ports-status-popover-rows.tsx
@@ -1,4 +1,4 @@
-import React, { useCallback, useMemo } from 'react'
+import React, { useCallback } from 'react'
import { Copy, ExternalLink, FolderOpen, Trash2 } from 'lucide-react'
import { toast } from 'sonner'
import { Button } from '@/components/ui/button'
@@ -15,9 +15,8 @@ import {
} from '@/lib/workspace-port-actions'
import type { WorkspacePortGroup } from '@/lib/workspace-port-groups'
import { useLocalhostLabelRouteForPort } from '@/lib/workspace-port-localhost-label-selector'
-import { getActiveRuntimeTarget } from '@/runtime/runtime-rpc-client'
+import { useWorktreeRuntimeTarget } from '@/runtime/use-worktree-runtime-target'
import { useAppStore } from '@/store'
-import { getRuntimeEnvironmentIdForWorktree } from '@/lib/worktree-runtime-owner'
import type { WorkspacePort } from '../../../../shared/workspace-ports'
import { translate } from '@/i18n/i18n'
@@ -67,6 +66,7 @@ function PortAction({
)
}
+/** One port row in the status-bar popover, with open/copy/stop actions on its owner host. */
export function PortRow({
port,
activeWorktreeId,
@@ -78,21 +78,13 @@ export function PortRow({
}): React.JSX.Element {
const settings = useAppStore((s) => s.settings)
const localhostLabelRoute = useLocalhostLabelRouteForPort(port)
- const runtimeEnvironmentId = useAppStore((s) =>
- getRuntimeEnvironmentIdForWorktree(
- s,
- port.kind === 'workspace' ? port.owner.worktreeId : activeWorktreeId
- )
- )
const createBrowserTab = useAppStore((s) => s.createBrowserTab)
const setRemoteBrowserPageHandle = useAppStore((s) => s.setRemoteBrowserPageHandle)
- const setWorkspacePortScan = useAppStore((s) => s.setWorkspacePortScan)
- const setWorkspacePortScanForKey = useAppStore((s) => s.setWorkspacePortScanForKey)
+ const replaceWorkspacePortScans = useAppStore((s) => s.replaceWorkspacePortScans)
const setWorkspacePortScanRefreshing = useAppStore((s) => s.setWorkspacePortScanRefreshing)
const recordFeatureInteraction = useAppStore((s) => s.recordFeatureInteraction)
- const runtimeTarget = useMemo(
- () => getActiveRuntimeTarget({ ...settings, activeRuntimeEnvironmentId: runtimeEnvironmentId }),
- [runtimeEnvironmentId, settings]
+ const runtimeTarget = useWorktreeRuntimeTarget(
+ port.kind === 'workspace' ? port.owner.worktreeId : activeWorktreeId
)
const processLabel = port.processName ?? (port.pid ? `PID ${port.pid}` : 'Unknown process')
const canStop = canStopWorkspacePort(port)
@@ -187,8 +179,7 @@ export function PortRow({
)
const refreshResult = await refreshWorkspacePortScanAfterStop({
runtimeTarget,
- setWorkspacePortScan,
- setWorkspacePortScanForKey,
+ replaceWorkspacePortScans,
getWorkspacePortScansByKey: () => useAppStore.getState().workspacePortScansByKey,
setWorkspacePortScanRefreshing
})
@@ -210,8 +201,7 @@ export function PortRow({
port,
recordFeatureInteraction,
runtimeTarget,
- setWorkspacePortScan,
- setWorkspacePortScanForKey,
+ replaceWorkspacePortScans,
setWorkspacePortScanRefreshing
]
)
diff --git a/src/renderer/src/lib/workspace-port-actions.ts b/src/renderer/src/lib/workspace-port-actions.ts
index cca201234a4..cf745d3e25f 100644
--- a/src/renderer/src/lib/workspace-port-actions.ts
+++ b/src/renderer/src/lib/workspace-port-actions.ts
@@ -7,7 +7,6 @@ import {
type RuntimeClientTarget
} from '@/runtime/runtime-rpc-client'
import { toRuntimeWorktreeSelector } from '@/runtime/runtime-worktree-selector'
-import { parseExecutionHostId, type ExecutionHostId } from '../../../shared/execution-host'
import type {
WorkspacePort,
WorkspacePortKillResult,
@@ -22,6 +21,11 @@ import { RUNTIME_BROWSER_UNAVAILABLE_MESSAGE } from './client-creation-action-po
export { addressForPort } from './workspace-port-urls'
const WORKSPACE_PORT_STOP_SETTLE_MS = 500
+const WORKSPACE_PORT_TARGET_UNAVAILABLE_REASON =
+ 'Workspace ports are unavailable for this execution host.'
+
+/** Projection key for the merged multi-host view; never a per-host scan key. */
+export const WORKSPACE_PORT_ALL_HOSTS_SCAN_KEY = 'all-hosts:all'
export function canStopWorkspacePort(
port: WorkspacePort
@@ -33,13 +37,17 @@ type BrowserTabCreator = ReturnType['createBrowserT
type RemoteBrowserPageHandleSetter = ReturnType<
typeof useAppStore.getState
>['setRemoteBrowserPageHandle']
-type WorkspacePortScanSetter = ReturnType['setWorkspacePortScan']
-type WorkspacePortScanByKeySetter = ReturnType<
- typeof useAppStore.getState
->['setWorkspacePortScanForKey']
type WorkspacePortScanRefreshingSetter = ReturnType<
typeof useAppStore.getState
>['setWorkspacePortScanRefreshing']
+type ReplaceWorkspacePortScansSetter = ReturnType<
+ typeof useAppStore.getState
+>['replaceWorkspacePortScans']
+
+export type WorkspacePortScanPublisher = {
+ replaceWorkspacePortScans: ReplaceWorkspacePortScansSetter
+ getWorkspacePortScansByKey: () => Record
+}
function delay(ms: number): Promise {
return new Promise((resolve) => window.setTimeout(resolve, ms))
@@ -94,12 +102,15 @@ export function goToWorkspacePortOwner(port: WorkspacePort): boolean {
export async function openWorkspacePortInBrowser(args: {
port: WorkspacePort
activeWorktreeId?: string | null
- runtimeTarget: RuntimeClientTarget
+ runtimeTarget: RuntimeClientTarget | null
createBrowserTab: BrowserTabCreator
setRemoteBrowserPageHandle: RemoteBrowserPageHandleSetter
openInOrcaBrowser?: boolean
localhostLabelRoute?: LocalhostWorktreeLabelRoute | null
}): Promise<{ ok: true } | { ok: false; reason: string }> {
+ if (!args.runtimeTarget) {
+ return { ok: false, reason: WORKSPACE_PORT_TARGET_UNAVAILABLE_REASON }
+ }
const rawUrl = browserUrlForPort(args.port)
let url = rawUrl
if (args.runtimeTarget.kind === 'local' && args.localhostLabelRoute) {
@@ -166,22 +177,38 @@ export async function openWorkspacePortInBrowser(args: {
}
}
-export async function refreshWorkspacePortScanAfterStop(args: {
- runtimeTarget: RuntimeClientTarget
- setWorkspacePortScan: WorkspacePortScanSetter
- setWorkspacePortScanForKey?: WorkspacePortScanByKeySetter
- setWorkspacePortScanRefreshing: WorkspacePortScanRefreshingSetter
- getWorkspacePortScansByKey?: () => Record
-}): Promise<{ ok: true } | { ok: false; reason: string }> {
+/**
+ * Stores one host's scan and republishes the aggregate the status bar reads.
+ * Why: a single-host publish used to overwrite that aggregate, so every other
+ * host's ports vanished from the count until the next background poll. One
+ * replaceWorkspacePortScans update (not setWorkspacePortScan) keeps the synthetic
+ * all-hosts key out of workspacePortScansByKey, where re-merging it would
+ * duplicate rows — and notifies subscribers once instead of twice for one scan.
+ */
+export function publishWorkspacePortScanForHost(
+ args: WorkspacePortScanPublisher & { scanKey: string; scan: WorkspacePortScanResult }
+): void {
+ const scansByKey = { ...args.getWorkspacePortScansByKey(), [args.scanKey]: args.scan }
+ const merged = mergeWorkspacePortScans(scansByKey)
+ args.replaceWorkspacePortScans(scansByKey, {
+ key: Object.keys(scansByKey).length > 1 ? WORKSPACE_PORT_ALL_HOSTS_SCAN_KEY : args.scanKey,
+ result: merged ?? args.scan
+ })
+}
+
+/** Re-scans one host after a port stop (immediately, then settled) and republishes the aggregate. */
+export async function refreshWorkspacePortScanAfterStop(
+ args: WorkspacePortScanPublisher & {
+ runtimeTarget: RuntimeClientTarget | null
+ setWorkspacePortScanRefreshing: WorkspacePortScanRefreshingSetter
+ }
+): Promise<{ ok: true } | { ok: false; reason: string }> {
+ if (!args.runtimeTarget) {
+ return { ok: false, reason: WORKSPACE_PORT_TARGET_UNAVAILABLE_REASON }
+ }
const scanKey = workspacePortScanKeyForTarget(args.runtimeTarget)
const publishScan = (scan: WorkspacePortScanResult): void => {
- args.setWorkspacePortScanForKey?.(scanKey, scan)
- const currentScans = args.getWorkspacePortScansByKey?.() ?? {}
- const merged = mergeWorkspacePortScans({ ...currentScans, [scanKey]: scan })
- args.setWorkspacePortScan({
- key: merged && Object.keys(currentScans).length > 0 ? 'all-hosts:all' : scanKey,
- result: merged ?? scan
- })
+ publishWorkspacePortScanForHost({ ...args, scanKey, scan })
}
args.setWorkspacePortScanRefreshing(true)
try {
@@ -216,19 +243,6 @@ export function workspacePortRuntimeTargetKey(target: RuntimeClientTarget): stri
return target.kind === 'local' ? 'local' : `environment:${target.environmentId}`
}
-export function runtimeTargetForExecutionHostId(
- hostId: ExecutionHostId
-): RuntimeClientTarget | null {
- const parsed = parseExecutionHostId(hostId)
- if (parsed?.kind === 'local') {
- return { kind: 'local' }
- }
- if (parsed?.kind === 'runtime') {
- return { kind: 'environment', environmentId: parsed.environmentId }
- }
- return null
-}
-
export function workspacePortScanKeyForTarget(target: RuntimeClientTarget): string {
return `${workspacePortRuntimeTargetKey(target)}:all`
}
@@ -295,9 +309,12 @@ export async function scanWorkspacePortsForTarget(
}
export async function killWorkspacePortForTarget(
- target: RuntimeClientTarget,
+ target: RuntimeClientTarget | null,
args: { repoId: string; pid: number; port: number }
): Promise {
+ if (!target) {
+ return { ok: false, reason: WORKSPACE_PORT_TARGET_UNAVAILABLE_REASON }
+ }
if (target.kind === 'local') {
return window.api.workspacePorts.kill(args)
}
diff --git a/src/renderer/src/lib/workspace-port-host-availability.test.ts b/src/renderer/src/lib/workspace-port-host-availability.test.ts
new file mode 100644
index 00000000000..7652c65a21a
--- /dev/null
+++ b/src/renderer/src/lib/workspace-port-host-availability.test.ts
@@ -0,0 +1,148 @@
+import { describe, expect, it } from 'vitest'
+import type { WorkspacePortScanResult } from '../../../shared/workspace-ports'
+import {
+ getUnavailableWorkspacePortHosts,
+ workspacePortHostForScanKey
+} from './workspace-port-host-availability'
+
+function scan(overrides: Partial = {}): WorkspacePortScanResult {
+ return { platform: 'linux', scannedAt: 1, ports: [], ...overrides }
+}
+
+describe('getUnavailableWorkspacePortHosts', () => {
+ it('reports the failed host while another host still answers', () => {
+ expect(
+ getUnavailableWorkspacePortHosts({
+ 'local:all': scan(),
+ 'environment:env-1:all': scan({ unavailableReason: 'Remote connection dropped' })
+ })
+ ).toEqual([
+ {
+ scanKey: 'environment:env-1:all',
+ host: { kind: 'environment', environmentId: 'env-1' },
+ platform: 'linux',
+ reason: 'Remote connection dropped'
+ }
+ ])
+ })
+
+ it('reports the local host as a local host ref, not an absent environment id', () => {
+ expect(
+ getUnavailableWorkspacePortHosts({
+ 'local:all': scan({ unavailableReason: 'lsof is unavailable' }),
+ 'environment:env-1:all': scan()
+ })
+ ).toEqual([
+ {
+ scanKey: 'local:all',
+ host: { kind: 'local' },
+ platform: 'linux',
+ reason: 'lsof is unavailable'
+ }
+ ])
+ })
+
+ it('keeps colons inside an environment id when parsing the scan key', () => {
+ // Why: keys are `${targetKey}:all`, so the id runs to the last `:all` —
+ // splitting on the first colon would truncate ids that contain colons.
+ expect(
+ getUnavailableWorkspacePortHosts({
+ 'local:all': scan(),
+ 'environment:weird:id:all': scan({ unavailableReason: 'Remote connection dropped' })
+ })
+ ).toEqual([
+ {
+ scanKey: 'environment:weird:id:all',
+ host: { kind: 'environment', environmentId: 'weird:id' },
+ platform: 'linux',
+ reason: 'Remote connection dropped'
+ }
+ ])
+ })
+
+ // Why: total loss of contact is where naming the host matters most — the merged
+ // projection joins raw internal scan keys, so it cannot name them itself.
+ it('names every host when all of them failed', () => {
+ expect(
+ getUnavailableWorkspacePortHosts({
+ 'local:all': scan({ unavailableReason: 'lsof is unavailable' }),
+ 'environment:env-1:all': scan({ unavailableReason: 'Remote connection dropped' })
+ })
+ ).toEqual([
+ {
+ scanKey: 'local:all',
+ host: { kind: 'local' },
+ platform: 'linux',
+ reason: 'lsof is unavailable'
+ },
+ {
+ scanKey: 'environment:env-1:all',
+ host: { kind: 'environment', environmentId: 'env-1' },
+ platform: 'linux',
+ reason: 'Remote connection dropped'
+ }
+ ])
+ })
+
+ it('names a single failed host', () => {
+ expect(
+ getUnavailableWorkspacePortHosts({
+ 'local:all': scan({ unavailableReason: 'lsof is unavailable' })
+ })
+ ).toEqual([
+ {
+ scanKey: 'local:all',
+ host: { kind: 'local' },
+ platform: 'linux',
+ reason: 'lsof is unavailable'
+ }
+ ])
+ })
+
+ // Why: the synthetic all-hosts projection key must never be labelled as the
+ // local machine — that would blame the wrong host for a remote failure.
+ it('marks an unrecognised scan key as an unknown host', () => {
+ expect(
+ getUnavailableWorkspacePortHosts({
+ 'all-hosts:all': scan({ unavailableReason: 'Remote connection dropped' })
+ })
+ ).toEqual([
+ {
+ scanKey: 'all-hosts:all',
+ host: { kind: 'unknown' },
+ platform: 'linux',
+ reason: 'Remote connection dropped'
+ }
+ ])
+ })
+
+ // Why: a paired web client's userAgent is not the Orca host's platform, so the
+ // caller labels the local host from the scan's own platform.
+ it("carries the failed scan's platform, and null when it is unknown", () => {
+ expect(
+ getUnavailableWorkspacePortHosts({
+ 'local:all': scan({ platform: 'win32', unavailableReason: 'netstat failed' }),
+ 'environment:env-1:all': scan({ platform: 'unknown', unavailableReason: 'dropped' })
+ }).map((entry) => entry.platform)
+ ).toEqual(['win32', null])
+ })
+
+ it('stays silent when nothing failed', () => {
+ expect(getUnavailableWorkspacePortHosts({ 'local:all': scan() })).toEqual([])
+ expect(getUnavailableWorkspacePortHosts({})).toEqual([])
+ })
+})
+
+describe('workspacePortHostForScanKey', () => {
+ it.each([
+ ['local:all', { kind: 'local' }],
+ ['environment:env-1:all', { kind: 'environment', environmentId: 'env-1' }],
+ ['environment:weird:id:all', { kind: 'environment', environmentId: 'weird:id' }],
+ ['all-hosts:all', { kind: 'unknown' }],
+ ['environment::all', { kind: 'unknown' }],
+ ['environment:env-1', { kind: 'unknown' }],
+ ['local', { kind: 'unknown' }]
+ ])('maps %s', (scanKey, expected) => {
+ expect(workspacePortHostForScanKey(scanKey)).toEqual(expected)
+ })
+})
diff --git a/src/renderer/src/lib/workspace-port-host-availability.ts b/src/renderer/src/lib/workspace-port-host-availability.ts
new file mode 100644
index 00000000000..4c643a5a118
--- /dev/null
+++ b/src/renderer/src/lib/workspace-port-host-availability.ts
@@ -0,0 +1,65 @@
+import type { WorkspacePortScanResult } from '../../../shared/workspace-ports'
+
+/**
+ * Host a per-host scan key points at. `unknown` is kept distinct from `local` so
+ * an unrecognised key (a synthetic projection key that leaked into the per-host
+ * map, say) is never mislabelled as a local failure.
+ */
+export type WorkspacePortHostRef =
+ | { kind: 'local' }
+ | { kind: 'environment'; environmentId: string }
+ | { kind: 'unknown' }
+
+export type UnavailableWorkspacePortHost = {
+ scanKey: string
+ host: WorkspacePortHostRef
+ /** Platform the failed scan last ran on; drives the local host's label. */
+ platform: NodeJS.Platform | null
+ reason: string
+}
+
+// Why: mirrors workspacePortScanKeyForTarget (`${targetKey}:all`, where the
+// target key is `local` or `environment:`) without importing the heavier
+// workspace-port-actions module into this pure helper. Splitting on the last
+// `:all` keeps environment ids that themselves contain colons intact.
+const ENVIRONMENT_SCAN_KEY_PREFIX = 'environment:'
+const SCAN_KEY_SUFFIX = ':all'
+const LOCAL_SCAN_KEY = `local${SCAN_KEY_SUFFIX}`
+
+/** Host a per-host scan key names; `unknown` for any other key shape. */
+export function workspacePortHostForScanKey(scanKey: string): WorkspacePortHostRef {
+ if (scanKey === LOCAL_SCAN_KEY) {
+ return { kind: 'local' }
+ }
+ if (!scanKey.endsWith(SCAN_KEY_SUFFIX) || !scanKey.startsWith(ENVIRONMENT_SCAN_KEY_PREFIX)) {
+ return { kind: 'unknown' }
+ }
+ const environmentId = scanKey.slice(
+ ENVIRONMENT_SCAN_KEY_PREFIX.length,
+ scanKey.length - SCAN_KEY_SUFFIX.length
+ )
+ return environmentId ? { kind: 'environment', environmentId } : { kind: 'unknown' }
+}
+
+/**
+ * Every host whose latest scan failed, named by host rather than by scan key.
+ * Why: on a remote host "none listening" and "could not look" are different
+ * answers, and the merged projection collapses both the partial case (no reason
+ * at all) and the total case (reasons joined with raw internal keys).
+ */
+export function getUnavailableWorkspacePortHosts(
+ scansByKey: Record
+): UnavailableWorkspacePortHost[] {
+ return Object.entries(scansByKey).flatMap(([scanKey, scan]) =>
+ scan?.unavailableReason
+ ? [
+ {
+ scanKey,
+ host: workspacePortHostForScanKey(scanKey),
+ platform: scan.platform === 'unknown' ? null : scan.platform,
+ reason: scan.unavailableReason
+ }
+ ]
+ : []
+ )
+}
diff --git a/src/renderer/src/lib/workspace-port-scan-debounce.test.ts b/src/renderer/src/lib/workspace-port-scan-debounce.test.ts
index bacac7ecacf..747de6672aa 100644
--- a/src/renderer/src/lib/workspace-port-scan-debounce.test.ts
+++ b/src/renderer/src/lib/workspace-port-scan-debounce.test.ts
@@ -25,6 +25,11 @@ function unavailable(): WorkspacePortScanResult {
return { platform: 'unknown', scannedAt: 1, ports: [], unavailableReason: 'scan failed' }
}
+/** What the ports popover publishes when its own scan fails: reason + last-good ports. */
+function unavailableWithRetainedPorts(portIds: string[]): WorkspacePortScanResult {
+ return { ...good(portIds), unavailableReason: 'scan failed' }
+}
+
const FAILURE_THRESHOLD = 2
function createHarness(): {
@@ -125,6 +130,33 @@ describe('reconcileTransientPortScanFailures', () => {
expect(state.has('flaky:all')).toBe(false)
})
+ // Why: the popover publishes reason + last-good ports the moment its own scan
+ // fails. Counting that as a spent grace period would drop those ports on the
+ // very next poll, so the retention would never survive one poll interval.
+ it('still grants the grace period after a failure that retained its ports', () => {
+ const { apply, publish } = createHarness()
+ apply([{ key: 'h:all', result: good(['tcp:3000']) }])
+ const popoverResult = unavailableWithRetainedPorts(['tcp:3000'])
+ publish('h:all', popoverResult)
+
+ const next = apply([{ key: 'h:all', result: unavailable() }])
+
+ expect(next[0].result).toBe(popoverResult)
+ expect(next[0].result.ports).toHaveLength(1)
+ })
+
+ it('still drops retained ports once failures reach the tolerance', () => {
+ const { apply, publish } = createHarness()
+ apply([{ key: 'h:all', result: good(['tcp:3000']) }])
+ publish('h:all', unavailableWithRetainedPorts(['tcp:3000']))
+ apply([{ key: 'h:all', result: unavailable() }])
+
+ const next = apply([{ key: 'h:all', result: unavailable() }])
+
+ expect(next[0].result.ports).toHaveLength(0)
+ expect(next[0].result.unavailableReason).toBe('scan failed')
+ })
+
it('uses a newer manual result instead of resurrecting stale ports', () => {
const { apply, publish } = createHarness()
apply([{ key: 'h:all', result: good(['tcp:3000']) }])
diff --git a/src/renderer/src/lib/workspace-port-scan-debounce.ts b/src/renderer/src/lib/workspace-port-scan-debounce.ts
index a6281f0ed9f..2677cbb0e52 100644
--- a/src/renderer/src/lib/workspace-port-scan-debounce.ts
+++ b/src/renderer/src/lib/workspace-port-scan-debounce.ts
@@ -34,10 +34,14 @@ export function reconcileTransientPortScanFailures(
return { key, result }
}
const failures = previousFailures + 1
+ // Why: a surface that hit the same failure first (the ports popover) republishes
+ // the host's last-good ports alongside the reason. Treating that as a spent grace
+ // period would drop those ports on the very next poll, undoing the retention.
+ const publishedIsRetainable =
+ Boolean(publishedResult) &&
+ (!publishedResult.unavailableReason || publishedResult.ports.length > 0)
const nextResult =
- failures < failureThreshold && publishedResult && !publishedResult.unavailableReason
- ? publishedResult
- : result
+ failures < failureThreshold && publishedIsRetainable ? publishedResult : result
state.set(key, { consecutiveFailures: failures, publishedResult: nextResult })
return { key, result: nextResult }
})
diff --git a/src/renderer/src/lib/workspace-port-scan-publish.test.ts b/src/renderer/src/lib/workspace-port-scan-publish.test.ts
new file mode 100644
index 00000000000..2f1386855a8
--- /dev/null
+++ b/src/renderer/src/lib/workspace-port-scan-publish.test.ts
@@ -0,0 +1,153 @@
+// @vitest-environment happy-dom
+
+import { beforeEach, describe, expect, it, vi } from 'vitest'
+import type { WorkspacePortScanResult } from '../../../shared/workspace-ports'
+
+vi.mock('@/lib/worktree-activation', () => ({
+ activateAndRevealWorktree: vi.fn()
+}))
+
+vi.mock('@/runtime/runtime-rpc-client', () => ({
+ getActiveRuntimeTarget: vi.fn(),
+ callRuntimeRpc: vi.fn(),
+ assertRuntimeEnvironmentCapability: vi.fn(),
+ RuntimeRpcCallError: class RuntimeRpcCallError extends Error {
+ code?: string
+ }
+}))
+
+vi.mock('./workspace-port-scan-client', () => ({
+ runWorkspacePortScanForTarget: vi.fn()
+}))
+
+const { publishWorkspacePortScanForHost, WORKSPACE_PORT_ALL_HOSTS_SCAN_KEY } =
+ await import('./workspace-port-actions')
+type WorkspacePortScanPublisher = Parameters[0]
+
+function scanWithPort(port: number, scannedAt: number): WorkspacePortScanResult {
+ return {
+ platform: 'linux',
+ scannedAt,
+ ports: [
+ {
+ id: `tcp:${port}`,
+ bindHost: '0.0.0.0',
+ connectHost: '127.0.0.1',
+ port,
+ pid: 100 + port,
+ processName: 'node',
+ protocol: 'http',
+ kind: 'external'
+ }
+ ]
+ }
+}
+
+/** Mirrors the store's replaceWorkspacePortScans semantics: one atomic update. */
+function makeStoreHarness(initial: Record = {}): {
+ scansByKey: Record
+ projections: { key: string; result: WorkspacePortScanResult }[]
+ publisher: Omit
+} {
+ let scansByKey: Record = { ...initial }
+ const projections: { key: string; result: WorkspacePortScanResult }[] = []
+ return {
+ get scansByKey() {
+ return scansByKey
+ },
+ projections,
+ publisher: {
+ replaceWorkspacePortScans: (
+ nextScansByKey: Record,
+ projection: { key: string; result: WorkspacePortScanResult } | null
+ ) => {
+ scansByKey = nextScansByKey
+ if (projection) {
+ projections.push(projection)
+ }
+ },
+ getWorkspacePortScansByKey: () => scansByKey
+ }
+ }
+}
+
+describe('publishWorkspacePortScanForHost', () => {
+ let localScan: WorkspacePortScanResult
+ let remoteScan: WorkspacePortScanResult
+
+ beforeEach(() => {
+ localScan = scanWithPort(5173, 10)
+ remoteScan = scanWithPort(3000, 20)
+ })
+
+ it('publishes the single tracked host under its own key', () => {
+ const harness = makeStoreHarness()
+
+ publishWorkspacePortScanForHost({
+ ...harness.publisher,
+ scanKey: 'local:all',
+ scan: localScan
+ })
+
+ expect(harness.projections).toEqual([{ key: 'local:all', result: localScan }])
+ expect(Object.keys(harness.scansByKey)).toEqual(['local:all'])
+ })
+
+ it('keeps the other host in the projection when one host refreshes', () => {
+ const harness = makeStoreHarness({ 'local:all': localScan })
+
+ publishWorkspacePortScanForHost({
+ ...harness.publisher,
+ scanKey: 'environment:env-1:all',
+ scan: remoteScan
+ })
+
+ const projection = harness.projections.at(-1)
+ expect(projection?.key).toBe(WORKSPACE_PORT_ALL_HOSTS_SCAN_KEY)
+ expect(projection?.result.ports.map((port) => port.port).sort()).toEqual([3000, 5173])
+ expect(Object.keys(harness.scansByKey).sort()).toEqual(['environment:env-1:all', 'local:all'])
+ })
+
+ it('publishes map and projection in a single store update', () => {
+ const harness = makeStoreHarness({ 'local:all': localScan })
+ const replaceSpy = vi.spyOn(harness.publisher, 'replaceWorkspacePortScans')
+
+ publishWorkspacePortScanForHost({
+ ...harness.publisher,
+ scanKey: 'environment:env-1:all',
+ scan: remoteScan
+ })
+
+ // Why: two sequential setter calls notify subscribers twice for one scan;
+ // one atomic replace keeps map and projection from ever disagreeing.
+ expect(replaceSpy).toHaveBeenCalledTimes(1)
+ const [nextScans, projection] = replaceSpy.mock.calls[0]
+ expect(Object.keys(nextScans).sort()).toEqual(['environment:env-1:all', 'local:all'])
+ expect(projection?.key).toBe(WORKSPACE_PORT_ALL_HOSTS_SCAN_KEY)
+ })
+
+ it('does not accumulate duplicate rows across repeated publishes', () => {
+ const harness = makeStoreHarness({ 'local:all': localScan })
+
+ publishWorkspacePortScanForHost({
+ ...harness.publisher,
+ scanKey: 'environment:env-1:all',
+ scan: remoteScan
+ })
+ publishWorkspacePortScanForHost({
+ ...harness.publisher,
+ scanKey: 'environment:env-1:all',
+ scan: { ...remoteScan, scannedAt: 30 }
+ })
+
+ // Why: the aggregate must never land in the per-host map, or the next merge
+ // folds the merged result back into itself and rows multiply.
+ expect(harness.scansByKey[WORKSPACE_PORT_ALL_HOSTS_SCAN_KEY]).toBeUndefined()
+ expect(
+ harness.projections
+ .at(-1)
+ ?.result.ports.map((port) => port.port)
+ .sort()
+ ).toEqual([3000, 5173])
+ })
+})
diff --git a/src/renderer/src/runtime/runtime-client-target.ts b/src/renderer/src/runtime/runtime-client-target.ts
index 1a0b8e4b8a4..fbf9af17374 100644
--- a/src/renderer/src/runtime/runtime-client-target.ts
+++ b/src/renderer/src/runtime/runtime-client-target.ts
@@ -1,4 +1,5 @@
import type { GlobalSettings } from '../../../shared/global-settings-types'
+import { parseExecutionHostId, type ExecutionHostId } from '../../../shared/execution-host'
export type RuntimeClientTarget = { kind: 'local' } | { kind: 'environment'; environmentId: string }
@@ -9,6 +10,20 @@ export function getActiveRuntimeTarget(
return environmentId ? { kind: 'environment', environmentId } : { kind: 'local' }
}
+/** RPC target for a dispatchable host; direct SSH cannot use this client path. */
+export function runtimeTargetForExecutionHostId(
+ hostId: ExecutionHostId
+): RuntimeClientTarget | null {
+ const parsed = parseExecutionHostId(hostId)
+ if (parsed?.kind === 'local') {
+ return { kind: 'local' }
+ }
+ if (parsed?.kind === 'runtime') {
+ return { kind: 'environment', environmentId: parsed.environmentId }
+ }
+ return null
+}
+
export function settingsForRuntimeOwner(
settings: Pick | null | undefined,
runtimeEnvironmentId: string | null | undefined
diff --git a/src/renderer/src/runtime/use-worktree-runtime-target.ts b/src/renderer/src/runtime/use-worktree-runtime-target.ts
new file mode 100644
index 00000000000..25bd9099e90
--- /dev/null
+++ b/src/renderer/src/runtime/use-worktree-runtime-target.ts
@@ -0,0 +1,16 @@
+import { useAppStore } from '@/store'
+import { getExecutionHostIdForWorktree } from '@/lib/worktree-runtime-owner'
+import { runtimeTargetForExecutionHostId, type RuntimeClientTarget } from './runtime-client-target'
+
+/**
+ * Runtime target that owns `worktreeId`, which is not always the globally
+ * focused runtime — acting on the focused one scans the wrong host and reports
+ * that workspace as having no ports. Direct-SSH owners return null.
+ */
+export function useWorktreeRuntimeTarget(
+ worktreeId: string | null | undefined
+): RuntimeClientTarget | null {
+ return useAppStore((state) =>
+ runtimeTargetForExecutionHostId(getExecutionHostIdForWorktree(state, worktreeId))
+ )
+}
From 3941edd4b6d474bf1c170cfbb7a0c80798d97bdc Mon Sep 17 00:00:00 2001
From: Neil <4138956+nwparker@users.noreply.github.com>
Date: Thu, 3 Sep 2026 22:20:35 -0700
Subject: [PATCH 21/49] perf(ipc): build the filesystem allowed-root list once
per authorization (#18423)
* perf(ipc): build the filesystem allowed-root list once per authorization
* perf(ipc): keep the allowed-root snapshot lazy so granted external paths build nothing
Hoisting getAllowedRoots to the top of resolveAuthorizedPath made every read of a
path covered by an external grant build the full root list, where main built none
(the grant answered before isPathAllowed reached the roots). Build on first use
instead: still one build per authorization, zero when a grant already answers.
* test(ipc): skip the allowed-root symlink escapes on Windows
Unprivileged Windows cannot create symlinks (EPERM), so both cases failed in
setup instead of exercising the escape check.
---
src/main/ipc/filesystem-allowed-roots.test.ts | 372 ++++++++++++++++++
src/main/ipc/filesystem-allowed-roots.ts | 59 ++-
src/main/ipc/filesystem-auth.ts | 65 ++-
src/main/project-runtime-git-options.ts | 7 +-
src/shared/project-groups.ts | 23 +-
5 files changed, 492 insertions(+), 34 deletions(-)
create mode 100644 src/main/ipc/filesystem-allowed-roots.test.ts
diff --git a/src/main/ipc/filesystem-allowed-roots.test.ts b/src/main/ipc/filesystem-allowed-roots.test.ts
new file mode 100644
index 00000000000..f94c99c5fdb
--- /dev/null
+++ b/src/main/ipc/filesystem-allowed-roots.test.ts
@@ -0,0 +1,372 @@
+import { mkdir, mkdtemp, realpath, rm, symlink, writeFile } from 'node:fs/promises'
+import { tmpdir } from 'node:os'
+import { join, resolve } from 'node:path'
+import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
+import type { Store } from '../persistence'
+import type * as RepoWorktrees from '../repo-worktrees'
+import { listRepoWorktreeGraph } from '../repo-worktrees'
+import type * as ProjectGroupsModule from '../../shared/project-groups'
+import { buildProjectGroupChildIndex, getProjectGroupSubtreeIds } from '../../shared/project-groups'
+import { isPathInsideOrEqual } from '../../shared/cross-platform-path'
+import { getWorktreeMirrorDistro } from '../project-runtime-git-options'
+import type { FolderWorkspace } from '../../shared/folder-workspace-types'
+import type { ProjectGroup } from '../../shared/project-group-types'
+import type { Project } from '../../shared/project-types'
+import type { Repo } from '../../shared/repo-types'
+import { getAllowedRoots } from './filesystem-allowed-roots'
+import { authorizeExternalPath, resolveAuthorizedPath } from './filesystem-auth'
+import { invalidateAuthorizedRootsCache } from './registered-worktree-roots-cache'
+import { computeWorkspaceRoot, getWorktreePathSettings } from './worktree-logic'
+
+vi.mock('../repo-worktrees', async () => {
+ const actual = await vi.importActual('../repo-worktrees')
+ return { ...actual, listRepoWorktreeGraph: vi.fn(async () => []) }
+})
+
+vi.mock('../../shared/project-groups', async () => {
+ const actual = await vi.importActual('../../shared/project-groups')
+ return {
+ ...actual,
+ buildProjectGroupChildIndex: vi.fn(actual.buildProjectGroupChildIndex),
+ getProjectGroupSubtreeIds: vi.fn(actual.getProjectGroupSubtreeIds)
+ }
+})
+
+type StoreFixture = {
+ repos: Repo[]
+ projects: Project[]
+ projectGroups: ProjectGroup[]
+ folderWorkspaces: FolderWorkspace[]
+ workspaceDir?: string
+}
+
+type StoreCallCounts = {
+ getRepos: number
+ getProjects: number
+ getProjectGroups: number
+ getFolderWorkspaces: number
+}
+
+function makeCountingStore(fixture: StoreFixture): { store: Store; counts: StoreCallCounts } {
+ const counts: StoreCallCounts = {
+ getRepos: 0,
+ getProjects: 0,
+ getProjectGroups: 0,
+ getFolderWorkspaces: 0
+ }
+ const store = {
+ getRepos: () => {
+ counts.getRepos += 1
+ // Match the real store, which rehydrates fresh repo objects on every read.
+ return fixture.repos.map((repo) => ({ ...repo }))
+ },
+ getProjects: () => {
+ counts.getProjects += 1
+ return fixture.projects.map((project) => ({ ...project }))
+ },
+ getProjectGroups: () => {
+ counts.getProjectGroups += 1
+ return fixture.projectGroups.map((group) => ({ ...group }))
+ },
+ getFolderWorkspaces: () => {
+ counts.getFolderWorkspaces += 1
+ return fixture.folderWorkspaces.map((workspace) => ({ ...workspace }))
+ },
+ getSettings: () => ({ nestWorkspaces: false, workspaceDir: fixture.workspaceDir ?? '' })
+ } as unknown as Store
+ return { store, counts }
+}
+
+/**
+ * The pre-change `getAllowedRoots` algorithm, kept verbatim so the equivalence test compares the
+ * new root list against the old one rather than against a hand-written expectation.
+ */
+function referenceAllowedRoots(store: Store): string[] {
+ const scopeStore = store as unknown as {
+ getRepos: () => Repo[]
+ getProjectGroups?: () => ProjectGroup[]
+ getFolderWorkspaces?: () => FolderWorkspace[]
+ getSettings: () => { workspaceDir?: string; nestWorkspaces?: boolean }
+ }
+ const localRepos = scopeStore.getRepos().filter((repo) => !repo.connectionId)
+ const settings = scopeStore.getSettings()
+
+ const scopeRepos = scopeStore.getRepos()
+ const projectGroups = scopeStore.getProjectGroups?.() ?? []
+ const isRemoteOnly = (
+ folderPath: string,
+ projectGroupId: string,
+ connectionId: string | null | undefined
+ ): boolean => {
+ if (connectionId) {
+ return true
+ }
+ const groupIds = getProjectGroupSubtreeIds(projectGroups, projectGroupId)
+ const candidates = scopeRepos.filter(
+ (repo) =>
+ (typeof repo.projectGroupId === 'string' && groupIds.has(repo.projectGroupId)) ||
+ isPathInsideOrEqual(folderPath, repo.path)
+ )
+ return candidates.length > 0 && candidates.every((repo) => Boolean(repo.connectionId))
+ }
+ const folderScopeRoots: string[] = []
+ for (const group of projectGroups) {
+ if (group.parentPath && !isRemoteOnly(group.parentPath, group.id, group.connectionId)) {
+ folderScopeRoots.push(resolve(group.parentPath))
+ }
+ }
+ for (const workspace of scopeStore.getFolderWorkspaces?.() ?? []) {
+ const connectionId =
+ workspace.connectionId ??
+ projectGroups.find((group) => group.id === workspace.projectGroupId)?.connectionId ??
+ null
+ if (!isRemoteOnly(workspace.folderPath, workspace.projectGroupId, connectionId)) {
+ folderScopeRoots.push(resolve(workspace.folderPath))
+ }
+ }
+
+ const roots = [...localRepos.map((repo) => resolve(repo.path)), ...folderScopeRoots]
+ if (settings.workspaceDir) {
+ if (localRepos.length === 0) {
+ roots.push(resolve(settings.workspaceDir))
+ } else {
+ for (const repo of localRepos) {
+ roots.push(
+ resolve(
+ computeWorkspaceRoot(
+ repo.path,
+ getWorktreePathSettings(repo, settings as never, getWorktreeMirrorDistro(store, repo))
+ )
+ )
+ )
+ }
+ }
+ }
+ return roots
+}
+
+function makeRepo(overrides: Partial & Pick): Repo {
+ return {
+ displayName: overrides.id,
+ badgeColor: '#000000',
+ addedAt: 1,
+ kind: 'git',
+ ...overrides
+ }
+}
+
+function makeGroup(overrides: Partial & Pick): ProjectGroup {
+ return {
+ name: overrides.id,
+ parentPath: null,
+ parentGroupId: null,
+ createdFrom: 'folder-scan',
+ tabOrder: 0,
+ isCollapsed: false,
+ color: null,
+ createdAt: 1,
+ updatedAt: 1,
+ ...overrides
+ }
+}
+
+function makeWorkspace(
+ overrides: Partial & Pick
+): FolderWorkspace {
+ return {
+ projectGroupId: 'group-root',
+ name: overrides.id,
+ comment: '',
+ linkedTask: null,
+ isArchived: false,
+ isUnread: false,
+ isPinned: false,
+ sortOrder: 1,
+ lastActivityAt: 1,
+ createdAt: 1,
+ updatedAt: 1,
+ ...overrides
+ }
+}
+
+/** Repos, nested groups, folder workspaces (one not a git worktree), and an SSH repo. */
+function makeMixedFixture(): StoreFixture {
+ const repos = [
+ makeRepo({ id: 'repo-local', path: '/repos/app', projectGroupId: 'group-root' }),
+ makeRepo({ id: 'repo-nested', path: '/repos/nested', projectGroupId: 'group-child' }),
+ makeRepo({ id: 'repo-folder', path: '/folders/plain', kind: 'folder' }),
+ makeRepo({
+ id: 'repo-ssh',
+ path: '/remote/app',
+ connectionId: 'ssh-1',
+ projectGroupId: 'group-remote'
+ })
+ ]
+ const projectGroups = [
+ makeGroup({ id: 'group-root', parentPath: '/folders/root' }),
+ makeGroup({ id: 'group-child', parentGroupId: 'group-root', parentPath: '/folders/child' }),
+ makeGroup({ id: 'group-grandchild', parentGroupId: 'group-child' }),
+ makeGroup({ id: 'group-remote', parentPath: '/remote/scope' }),
+ makeGroup({ id: 'group-connection', parentPath: '/remote/via-group', connectionId: 'ssh-1' })
+ ]
+ const folderWorkspaces = [
+ makeWorkspace({ id: 'ws-git', folderPath: '/folders/root/feature' }),
+ // Not a git worktree: a plain folder workspace under a folder-kind repo.
+ makeWorkspace({
+ id: 'ws-plain',
+ folderPath: '/folders/plain/scratch',
+ projectGroupId: 'group-child'
+ }),
+ makeWorkspace({ id: 'ws-remote', folderPath: '/remote/ws', projectGroupId: 'group-remote' }),
+ makeWorkspace({
+ id: 'ws-connection',
+ folderPath: '/remote/direct',
+ projectGroupId: 'group-connection'
+ }),
+ makeWorkspace({
+ id: 'ws-unlinked',
+ folderPath: '/folders/unlinked',
+ projectGroupId: 'group-orphan'
+ })
+ ]
+ const projects: Project[] = [
+ {
+ id: 'project-1',
+ displayName: 'App',
+ badgeColor: '#000000',
+ sourceRepoIds: ['repo-local', 'repo-nested'],
+ createdAt: 1,
+ updatedAt: 1
+ },
+ {
+ id: 'project-2',
+ displayName: 'Folder',
+ badgeColor: '#000000',
+ sourceRepoIds: ['repo-folder'],
+ createdAt: 1,
+ updatedAt: 1
+ }
+ ]
+ return { repos, projects, projectGroups, folderWorkspaces, workspaceDir: '/workspaces' }
+}
+
+beforeEach(() => {
+ invalidateAuthorizedRootsCache()
+ vi.mocked(buildProjectGroupChildIndex).mockClear()
+ vi.mocked(getProjectGroupSubtreeIds).mockClear()
+})
+
+describe('getAllowedRoots', () => {
+ it('produces the same roots as the pre-change implementation', () => {
+ const { store } = makeCountingStore(makeMixedFixture())
+
+ expect(getAllowedRoots(store)).toEqual(referenceAllowedRoots(store))
+ })
+
+ it('reads the store once and indexes project groups once per build', () => {
+ const fixture = makeMixedFixture()
+ const { store, counts } = makeCountingStore(fixture)
+
+ getAllowedRoots(store)
+
+ expect.soft(counts.getRepos).toBe(1)
+ expect.soft(counts.getProjectGroups).toBe(1)
+ expect.soft(counts.getFolderWorkspaces).toBe(1)
+ // Batched runtime resolution scans the project list once, not once per local repo.
+ expect.soft(counts.getProjects).toBe(1)
+ // The per-scope subtree walk no longer rebuilds the parent->children index.
+ expect.soft(vi.mocked(buildProjectGroupChildIndex)).toHaveBeenCalledTimes(1)
+ expect.soft(vi.mocked(getProjectGroupSubtreeIds)).not.toHaveBeenCalled()
+ })
+})
+
+describe('resolveAuthorizedPath allowed-root reuse', () => {
+ let repoRoot: string
+ let outsideRoot: string
+ let store: Store
+ let counts: StoreCallCounts
+
+ beforeEach(async () => {
+ repoRoot = await mkdtemp(join(await realpath(tmpdir()), 'orca-allowed-roots-'))
+ outsideRoot = await mkdtemp(join(await realpath(tmpdir()), 'orca-outside-'))
+ const fixture = makeMixedFixture()
+ fixture.repos = [makeRepo({ id: 'repo-local', path: repoRoot }), ...fixture.repos]
+ fixture.projects[0]!.sourceRepoIds = ['repo-local']
+ ;({ store, counts } = makeCountingStore(fixture))
+ })
+
+ afterEach(async () => {
+ await rm(repoRoot, { recursive: true, force: true })
+ await rm(outsideRoot, { recursive: true, force: true })
+ })
+
+ it('builds the allowed-root list once per call across repeated reads', async () => {
+ const dirPath = join(repoRoot, 'src')
+ await mkdir(dirPath)
+ await writeFile(join(dirPath, 'index.ts'), 'export {}\n')
+ const callCount = 5
+
+ for (let index = 0; index < callCount; index += 1) {
+ await resolveAuthorizedPath(dirPath, store)
+ await resolveAuthorizedPath(join(dirPath, 'index.ts'), store)
+ }
+
+ const buildCount = callCount * 2
+ // One build per authorization, not one per raw-path check plus one per realpath check.
+ expect.soft(counts.getFolderWorkspaces).toBe(buildCount)
+ expect.soft(counts.getRepos).toBe(buildCount)
+ expect.soft(counts.getProjects).toBe(buildCount)
+ expect.soft(vi.mocked(buildProjectGroupChildIndex)).toHaveBeenCalledTimes(buildCount)
+ expect.soft(vi.mocked(getProjectGroupSubtreeIds)).not.toHaveBeenCalled()
+ })
+
+ // Why (both symlink cases): creating a symlink on Windows needs elevation or
+ // Developer Mode, so these would fail EPERM in setup rather than exercise the
+ // escape check. Every non-symlink case still runs there.
+ it.skipIf(process.platform === 'win32')(
+ 'still refuses a symlink that escapes every allowed root',
+ async () => {
+ const secret = join(outsideRoot, 'secret.txt')
+ await writeFile(secret, 'secret\n')
+ const escape = join(repoRoot, 'escape.txt')
+ await symlink(secret, escape)
+
+ await expect(resolveAuthorizedPath(escape, store)).rejects.toThrow('Access denied')
+ expect(vi.mocked(listRepoWorktreeGraph)).toHaveBeenCalled()
+ }
+ )
+
+ it('builds no allowed-root list at all for a granted external path', async () => {
+ const external = join(outsideRoot, 'external.md')
+ await writeFile(external, 'notes\n')
+ authorizeExternalPath(external)
+ counts.getRepos = 0
+ counts.getProjects = 0
+ counts.getFolderWorkspaces = 0
+
+ for (let index = 0; index < 5; index += 1) {
+ await expect(resolveAuthorizedPath(external, store)).resolves.toBe(external)
+ }
+
+ // The grant answers on its own; hoisting the snapshot must not turn zero builds into one per read.
+ expect.soft(counts.getRepos).toBe(0)
+ expect.soft(counts.getProjects).toBe(0)
+ expect.soft(counts.getFolderWorkspaces).toBe(0)
+ expect.soft(vi.mocked(buildProjectGroupChildIndex)).not.toHaveBeenCalled()
+ })
+
+ it.skipIf(process.platform === 'win32')(
+ 'still refuses a directory symlink that escapes every allowed root',
+ async () => {
+ const outsideDir = join(outsideRoot, 'nested')
+ await mkdir(outsideDir)
+ await writeFile(join(outsideDir, 'file.txt'), 'secret\n')
+ const escape = join(repoRoot, 'escape-dir')
+ await symlink(outsideDir, escape)
+
+ await expect(resolveAuthorizedPath(join(escape, 'file.txt'), store)).rejects.toThrow(
+ 'Access denied'
+ )
+ }
+ )
+})
diff --git a/src/main/ipc/filesystem-allowed-roots.ts b/src/main/ipc/filesystem-allowed-roots.ts
index 3cb7fe4fa55..cef249430c6 100644
--- a/src/main/ipc/filesystem-allowed-roots.ts
+++ b/src/main/ipc/filesystem-allowed-roots.ts
@@ -1,9 +1,16 @@
import { resolve } from 'node:path'
import type { Store } from '../persistence'
import { computeWorkspaceRoot, getWorktreePathSettings } from './worktree-logic'
-import { getWorktreeMirrorDistro } from '../project-runtime-git-options'
+import {
+ getWorktreeMirrorDistroForRuntime,
+ resolveLocalProjectRuntimesForRepos
+} from '../project-runtime-git-options'
import { isPathInsideOrEqual } from '../../shared/cross-platform-path'
-import { getProjectGroupSubtreeIds } from '../../shared/project-groups'
+import {
+ buildProjectGroupChildIndex,
+ collectProjectGroupSubtreeIds,
+ type ProjectGroupChildIndex
+} from '../../shared/project-groups'
import type { FolderWorkspace } from '../../shared/folder-workspace-types'
import type { ProjectGroup } from '../../shared/project-group-types'
import type { Repo } from '../../shared/repo-types'
@@ -11,18 +18,22 @@ import type { Repo } from '../../shared/repo-types'
type FolderScopeStore = Pick &
Partial>
+// Why: SSH repo paths are remote-host paths; treating them as local roots could authorize unrelated local folders or probe SSH-only paths.
+function filterLocalRepos(repos: readonly Repo[]): Repo[] {
+ return repos.filter((repo) => !repo.connectionId)
+}
+
export function getLocalRepos(store: Store) {
- // Why: SSH repo paths are remote-host paths; treating them as local roots could authorize unrelated local folders or probe SSH-only paths.
- return store.getRepos().filter((repo) => !repo.connectionId)
+ return filterLocalRepos(store.getRepos())
}
function getFolderScopeCandidateRepos(
folderPath: string,
projectGroupId: string,
- projectGroups: readonly ProjectGroup[],
+ childGroupIndex: ProjectGroupChildIndex,
repos: readonly Repo[]
): Repo[] {
- const groupIds = getProjectGroupSubtreeIds(projectGroups, projectGroupId)
+ const groupIds = collectProjectGroupSubtreeIds(childGroupIndex, projectGroupId)
return repos.filter(
(repo) =>
(typeof repo.projectGroupId === 'string' && groupIds.has(repo.projectGroupId)) ||
@@ -34,13 +45,18 @@ function isRemoteOnlyFolderScope(
folderPath: string,
projectGroupId: string,
connectionId: string | null | undefined,
- projectGroups: readonly ProjectGroup[],
+ childGroupIndex: ProjectGroupChildIndex,
repos: readonly Repo[]
): boolean {
if (connectionId) {
return true
}
- const candidates = getFolderScopeCandidateRepos(folderPath, projectGroupId, projectGroups, repos)
+ const candidates = getFolderScopeCandidateRepos(
+ folderPath,
+ projectGroupId,
+ childGroupIndex,
+ repos
+ )
return candidates.length > 0 && candidates.every((repo) => Boolean(repo.connectionId))
}
@@ -55,16 +71,22 @@ function getFolderWorkspaceConnectionId(
)
}
-function getLocalFolderScopeRoots(store: Store): string[] {
+function getLocalFolderScopeRoots(store: Store, repos: readonly Repo[]): string[] {
const scopeStore = store as FolderScopeStore
- const repos = scopeStore.getRepos()
// Why: many filesystem tests use narrow Store doubles; folder scopes are additive.
const projectGroups = scopeStore.getProjectGroups?.() ?? []
+ const childGroupIndex = buildProjectGroupChildIndex(projectGroups)
const roots: string[] = []
for (const group of projectGroups) {
if (
group.parentPath &&
- !isRemoteOnlyFolderScope(group.parentPath, group.id, group.connectionId, projectGroups, repos)
+ !isRemoteOnlyFolderScope(
+ group.parentPath,
+ group.id,
+ group.connectionId,
+ childGroupIndex,
+ repos
+ )
) {
roots.push(resolve(group.parentPath))
}
@@ -75,7 +97,7 @@ function getLocalFolderScopeRoots(store: Store): string[] {
workspace.folderPath,
workspace.projectGroupId,
getFolderWorkspaceConnectionId(workspace, projectGroups),
- projectGroups,
+ childGroupIndex,
repos
)
) {
@@ -86,16 +108,19 @@ function getLocalFolderScopeRoots(store: Store): string[] {
}
export function getAllowedRoots(store: Store): string[] {
- const localRepos = getLocalRepos(store)
+ // Why one read: `getRepos` rehydrates every repo, and this runs twice per filesystem IPC.
+ const repos = store.getRepos()
+ const localRepos = filterLocalRepos(repos)
const settings = store.getSettings()
const roots = [
...localRepos.map((repo) => resolve(repo.path)),
- ...getLocalFolderScopeRoots(store)
+ ...getLocalFolderScopeRoots(store, repos)
]
if (settings.workspaceDir) {
if (localRepos.length === 0) {
roots.push(resolve(settings.workspaceDir))
} else {
+ const projectRuntimeByRepoId = resolveLocalProjectRuntimesForRepos(store, localRepos)
for (const repo of localRepos) {
roots.push(
resolve(
@@ -104,7 +129,11 @@ export function getAllowedRoots(store: Store): string[] {
// Why enriched here too: placement has to agree with the create
// flow, or renderer file access is denied for a worktree Orca
// just put on the WSL side.
- getWorktreePathSettings(repo, settings, getWorktreeMirrorDistro(store, repo))
+ getWorktreePathSettings(
+ repo,
+ settings,
+ getWorktreeMirrorDistroForRuntime(projectRuntimeByRepoId.get(repo.id))
+ )
)
)
)
diff --git a/src/main/ipc/filesystem-auth.ts b/src/main/ipc/filesystem-auth.ts
index 122617845ed..894e39945c1 100644
--- a/src/main/ipc/filesystem-auth.ts
+++ b/src/main/ipc/filesystem-auth.ts
@@ -43,7 +43,24 @@ export function authorizeExternalPath(targetPath: string): void {
} catch {}
}
-export function isPathAllowed(targetPath: string, store: Store): boolean {
+/**
+ * One allowed-root list shared by every check in a single authorization.
+ *
+ * Lazy so a path already covered by an external grant still builds nothing at all, the way it did
+ * before the list was hoisted out of the individual checks.
+ */
+type AllowedRootsSnapshot = { get: () => readonly string[] }
+
+function createAllowedRootsSnapshot(store: Store): AllowedRootsSnapshot {
+ let roots: readonly string[] | undefined
+ return { get: () => (roots ??= getAllowedRoots(store)) }
+}
+
+export function isPathAllowed(
+ targetPath: string,
+ store: Store,
+ allowedRoots?: AllowedRootsSnapshot
+): boolean {
const resolvedTarget = resolve(targetPath)
if (authorizedExternalPaths.has(resolvedTarget)) {
return true
@@ -53,7 +70,9 @@ export function isPathAllowed(targetPath: string, store: Store): boolean {
return true
}
}
- return getAllowedRoots(store).some((root) => isDescendantOrEqual(resolvedTarget, root))
+ return (allowedRoots?.get() ?? getAllowedRoots(store)).some((root) =>
+ isDescendantOrEqual(resolvedTarget, root)
+ )
}
export type ResolveAuthorizedPathOptions = {
@@ -69,7 +88,10 @@ export async function resolveAuthorizedPath(
options: ResolveAuthorizedPathOptions = {}
): Promise {
const resolvedTarget = resolve(targetPath)
- if (!(await isPathAllowedIncludingRegisteredWorktrees(resolvedTarget, store))) {
+ // Why: the roots depend only on store state, not on the candidate path, so one snapshot serves
+ // every authorization below; each candidate is still checked against it in full.
+ const allowedRoots = createAllowedRootsSnapshot(store)
+ if (!(await isPathAllowedIncludingRegisteredWorktrees(resolvedTarget, store, { allowedRoots }))) {
throw new Error(PATH_ACCESS_DENIED_MESSAGE)
}
@@ -80,14 +102,15 @@ export async function resolveAuthorizedPath(
realParent = await realpath(dirname(resolvedTarget))
} catch (error) {
if (isENOENT(error)) {
- return resolveAuthorizedMissingPath(resolvedTarget, store)
+ return resolveAuthorizedMissingPath(resolvedTarget, store, allowedRoots)
}
throw error
}
const candidateTarget = resolve(realParent, basename(resolvedTarget))
if (
!(await isPathAllowedIncludingRegisteredWorktrees(candidateTarget, store, {
- canonicalSourcePath: resolvedTarget
+ canonicalSourcePath: resolvedTarget,
+ allowedRoots
}))
) {
throw new Error(PATH_ACCESS_DENIED_MESSAGE)
@@ -100,7 +123,8 @@ export async function resolveAuthorizedPath(
const realTarget = resolve(await realpath(resolvedTarget))
if (
!(await isPathAllowedIncludingRegisteredWorktrees(realTarget, store, {
- canonicalSourcePath: resolvedTarget
+ canonicalSourcePath: resolvedTarget,
+ allowedRoots
}))
) {
throw new Error(PATH_ACCESS_DENIED_MESSAGE)
@@ -110,11 +134,15 @@ export async function resolveAuthorizedPath(
if (!isENOENT(error)) {
throw error
}
- return resolveAuthorizedMissingPath(resolvedTarget, store)
+ return resolveAuthorizedMissingPath(resolvedTarget, store, allowedRoots)
}
}
-async function resolveAuthorizedMissingPath(resolvedTarget: string, store: Store): Promise {
+async function resolveAuthorizedMissingPath(
+ resolvedTarget: string,
+ store: Store,
+ allowedRoots: AllowedRootsSnapshot
+): Promise {
let existingAncestor = resolvedTarget
const missingSegments: string[] = []
@@ -124,7 +152,8 @@ async function resolveAuthorizedMissingPath(resolvedTarget: string, store: Store
const candidateTarget = resolve(realAncestor, ...missingSegments)
if (
!(await isPathAllowedIncludingRegisteredWorktrees(candidateTarget, store, {
- canonicalSourcePath: resolvedTarget
+ canonicalSourcePath: resolvedTarget,
+ allowedRoots
}))
) {
throw new Error(PATH_ACCESS_DENIED_MESSAGE)
@@ -148,9 +177,9 @@ async function resolveAuthorizedMissingPath(resolvedTarget: string, store: Store
async function isPathAllowedIncludingRegisteredWorktrees(
targetPath: string,
store: Store,
- options: { canonicalSourcePath?: string } = {}
+ options: { canonicalSourcePath?: string; allowedRoots?: AllowedRootsSnapshot } = {}
): Promise {
- if (isPathAllowed(targetPath, store)) {
+ if (isPathAllowed(targetPath, store, options.allowedRoots)) {
return true
}
@@ -158,7 +187,14 @@ async function isPathAllowedIncludingRegisteredWorktrees(
return true
}
- if (await isPathAllowedByCanonicalAllowedRoot(targetPath, options.canonicalSourcePath, store)) {
+ if (
+ await isPathAllowedByCanonicalAllowedRoot(
+ targetPath,
+ options.canonicalSourcePath,
+ store,
+ options.allowedRoots
+ )
+ ) {
return true
}
@@ -178,12 +214,13 @@ async function isPathAllowedIncludingRegisteredWorktrees(
async function isPathAllowedByCanonicalAllowedRoot(
targetPath: string,
sourcePath: string | undefined,
- store: Store
+ store: Store,
+ allowedRoots?: AllowedRootsSnapshot
): Promise {
if (!sourcePath) {
return false
}
- for (const root of getAllowedRoots(store)) {
+ for (const root of allowedRoots?.get() ?? getAllowedRoots(store)) {
const resolvedRoot = resolve(root)
if (!isDescendantOrEqual(sourcePath, resolvedRoot)) {
continue
diff --git a/src/main/project-runtime-git-options.ts b/src/main/project-runtime-git-options.ts
index 808d31d5fcf..20aa0e9659a 100644
--- a/src/main/project-runtime-git-options.ts
+++ b/src/main/project-runtime-git-options.ts
@@ -102,7 +102,12 @@ export function getWorktreeMirrorDistro(
store: ProjectRuntimeResolutionStore,
repo: Repo
): string | undefined {
- const projectRuntime = resolveLocalProjectRuntimeForRepo(store, repo)
+ return getWorktreeMirrorDistroForRuntime(resolveLocalProjectRuntimeForRepo(store, repo))
+}
+
+export function getWorktreeMirrorDistroForRuntime(
+ projectRuntime: ProjectExecutionRuntimeResolution | undefined
+): string | undefined {
if (!projectRuntime || projectRuntime.status !== 'resolved') {
return undefined
}
diff --git a/src/shared/project-groups.ts b/src/shared/project-groups.ts
index 67c2897ead6..c4fe8badb47 100644
--- a/src/shared/project-groups.ts
+++ b/src/shared/project-groups.ts
@@ -109,10 +109,12 @@ export function clearMissingProjectGroupMemberships(repos: Repo[], groups: Proje
)
}
-export function getProjectGroupSubtreeIds(
- groups: readonly Pick[],
- rootGroupId: string
-): Set {
+export type ProjectGroupChildIndex = ReadonlyMap
+
+/** Build once and reuse when collecting subtrees for more than one root. */
+export function buildProjectGroupChildIndex(
+ groups: readonly Pick[]
+): ProjectGroupChildIndex {
const childGroupsByParentId = new Map()
for (const group of groups) {
if (!group.parentGroupId) {
@@ -122,7 +124,20 @@ export function getProjectGroupSubtreeIds(
children.push(group.id)
childGroupsByParentId.set(group.parentGroupId, children)
}
+ return childGroupsByParentId
+}
+export function getProjectGroupSubtreeIds(
+ groups: readonly Pick[],
+ rootGroupId: string
+): Set {
+ return collectProjectGroupSubtreeIds(buildProjectGroupChildIndex(groups), rootGroupId)
+}
+
+export function collectProjectGroupSubtreeIds(
+ childGroupsByParentId: ProjectGroupChildIndex,
+ rootGroupId: string
+): Set {
const subtreeIds = new Set()
const pending = [rootGroupId]
while (pending.length > 0) {
From 79d5fb469a31f0fce51fe360bb894ac2f9d7b121 Mon Sep 17 00:00:00 2001
From: Jinwoo Hong <73622457+Jinwoo-H@users.noreply.github.com>
Date: Fri, 4 Sep 2026 01:23:54 -0400
Subject: [PATCH 22/49] fix(cloud): recalibrate the relay monitor's
postgres-retry freeze to a measured bar (#18580)
The global relay_cells FOR UPDATE lock made successful retries a
steady-state rate: fleet-wide p50 430 / p90 924 / p99 1320 / max 1504
per five minutes over the last 24 h, 55% of windows over the 300 bar,
only 22% of 15-minute gates clean. Three read-only dry-runs on
2026-09-04 froze on it, blocking the same-cap roll that carries #18521
and the beginProof crash guard to the 23 cells. 2000 clears every
measured healthy gate; the exhausted-retry, director concurrency, and
pool bars keep the incident discriminator role.
---
.../relay-ops/src/incident-monitor.test.ts | 17 +++++++++------
cloud/apps/relay-ops/src/incident-monitor.ts | 21 +++++++++++++------
cloud/docs/relay-incident-monitor.md | 16 +++++++++++++-
3 files changed, 41 insertions(+), 13 deletions(-)
diff --git a/cloud/apps/relay-ops/src/incident-monitor.test.ts b/cloud/apps/relay-ops/src/incident-monitor.test.ts
index 61a73b64dbe..4e1da9fab26 100644
--- a/cloud/apps/relay-ops/src/incident-monitor.test.ts
+++ b/cloud/apps/relay-ops/src/incident-monitor.test.ts
@@ -111,14 +111,19 @@ describe('incident monitor evaluator', () => {
})
})
- it('freezes when postgres retries exceed the recalibrated ceiling', () => {
- const sample = healthySample()
- sample.sources['relay-logs']!.signals['relay.postgres_retries'] =
- signal(INCIDENT_MONITOR_THRESHOLDS.relayPostgresRetries + 1)
- expect(evaluateIncidentSample(sample, startedAt)).toMatchObject({
+ // Why: the global relay_cells lock made retries a steady-state rate (24 h p99
+ // 1320/5min on 2026-09-04); the bar fences only unbounded growth beyond that.
+ it('tolerates the measured healthy retry rate and freezes above the bar', () => {
+ const healthy = healthySample()
+ healthy.sources['relay-logs']!.signals['relay.postgres_retries'] = signal(1504)
+ expect(evaluateIncidentSample(healthy, startedAt).status).toBe('green')
+
+ const incident = healthySample()
+ incident.sources['relay-logs']!.signals['relay.postgres_retries'] = signal(2001)
+ expect(evaluateIncidentSample(incident, startedAt)).toMatchObject({
status: 'freeze',
failures: [
- expect.objectContaining({ signal: 'relay.postgres_retries', threshold: 300 })
+ expect.objectContaining({ signal: 'relay.postgres_retries', threshold: 2000 })
]
})
})
diff --git a/cloud/apps/relay-ops/src/incident-monitor.ts b/cloud/apps/relay-ops/src/incident-monitor.ts
index 868bb86fb93..a121568d918 100644
--- a/cloud/apps/relay-ops/src/incident-monitor.ts
+++ b/cloud/apps/relay-ops/src/incident-monitor.ts
@@ -32,11 +32,20 @@ export const INCIDENT_MONITOR_THRESHOLDS = {
relayPoolWaiting: 800,
relayPoolWaitMs: 2_500,
// Why: successful lock retries are the contention machinery working, not harm.
- // Healthy 2026-08-26 baseline bursts to 234/5min (26% of windows crossed the old
- // bar of 20, set unmeasured at the monitor's 2026-07-28 birth); the 2026-08-23
- // incident ran ~2,200-3,000/5min. 300 clears healthy bursts with ~10x incident
- // margin; relayPostgresRetryExhausted below bounds the terminally failed share.
- relayPostgresRetries: 300,
+ // Recalibrated 2026-09-04 from 300, which was set 2026-08-26 when healthy bursts
+ // reached 234/5min. The global relay_cells FOR UPDATE lock has since become the
+ // fleet's steady state: measured fleet-wide (director + cells, summed per five
+ // minutes) 2026-09-03T05Z..2026-09-04T05Z p50 430 / p90 924 / p99 1320 / max
+ // 1504, with 55% of windows over 300 and only 22% of 15-minute gates clean, so
+ // the bar blocked the very cell roll that carries the 500 ms lock wait (#18521)
+ // and the beginProof crash guard to the cells. The 2026-08-23 lock incident on
+ // this same metric peaked at 1510 in one window and 646 in the next, so it is
+ // not separable from today's contention by retries alone; it is caught by
+ // relayPostgresRetryExhausted (467 at the peak vs a 300 bar), director
+ // concurrency, and the pool bars. 2000 passes every healthy 15-minute window
+ // measured in the last 24 h and still fences unbounded growth. Re-tighten once
+ // the fleet is on the 500 ms lock wait and the baseline is re-measured.
+ relayPostgresRetries: 2000,
// Why: 300 per five minutes, recalibrated 2026-09-04 from a bar of zero that no
// production window has cleared since #18521 shipped to the director. That
// change cut the request-path cell-inventory wait from the 1 s pool lock_timeout
@@ -48,7 +57,7 @@ export const INCIDENT_MONITOR_THRESHOLDS = {
// quiet hours p50 2 / max 36; pre-#18521 daytime p50 10 / p90 25 / max 87;
// post-#18521 p50 42 / p90 147 / max 220. The 2026-08-23 lock incident peaked
// at 467. 300 clears every measured healthy window and still sits below the
- // incident shape; relayPostgresRetries above stays the ~10x discriminator.
+ // incident shape; retries above fence only unbounded growth.
// User-facing /v1/assign 503 share did not move with #18521 (13.9% old image
// vs 12.3% new, same evening), so exhaustion is not a proxy for user harm.
relayPostgresRetryExhausted: 300,
diff --git a/cloud/docs/relay-incident-monitor.md b/cloud/docs/relay-incident-monitor.md
index 337d3f1b20f..5c563f6a2e2 100644
--- a/cloud/docs/relay-incident-monitor.md
+++ b/cloud/docs/relay-incident-monitor.md
@@ -99,7 +99,7 @@ durably marked consumed before mutation and cannot authorize another run.
| Cloud SQL deadlocks | over 0 |
| Relay pool waiters | over 800 |
| Relay pool wait | over 2,500 ms |
-| PostgreSQL retries in five minutes | over 300 |
+| PostgreSQL retries in five minutes | over 2,000 |
| Exhausted PostgreSQL retries in five minutes | over 300 |
| Director instances | outside 5–6 |
| Director CPU or memory | over 80% |
@@ -139,6 +139,20 @@ heartbeats, and matching live admission.
logs: healthy-day bursts reach 234/5min with zero exhausted retries and 26%
of five-minute windows over 20, while the 2026-08-23 lock-contention
incident ran roughly 2,200–3,000/5min.
+- Recalibrated the PostgreSQL-retry freeze from 300 to 2,000 per five minutes
+ (2026-09-04). Basis: the global `relay_cells FOR UPDATE` lock made
+ successful retries a steady-state rate. Measured fleet-wide (director +
+ cells, summed per five minutes from the `orca_relay_postgres_retries`
+ log metric) over 2026-09-03T05Z..2026-09-04T05Z: p50 430 / p90 924 /
+ p99 1,320 / max 1,504; 55% of windows over 300; only 22% of 15-minute gates
+ clean at 300 versus 100% at 2,000. Three read-only dry-runs on 2026-09-04
+ froze on this bar (runs 33836470590, 33838698725) or on a genuine six-cell
+ crash storm (33837160275), blocking the same-cap roll that carries #18521
+ and the `beginProof` crash guard to the 23 cells. The 2026-08-23 incident
+ on this metric peaked at 1,510 then 646, so retries alone no longer
+ separate it from today's baseline; the exhausted-retry bar (incident peak
+ 467 vs bar 300), director concurrency, and the pool bars carry that role.
+ Re-tighten after the fleet is on the 500 ms lock wait.
- Recalibrated the exhausted-PostgreSQL-retry freeze from 0 to 300 per five
minutes (2026-09-04). Basis: #18521 cut the request-path cell-inventory
lock wait from the 1 s pool `lock_timeout` to 500 ms, so contended waiters
From b378101901d8062765ec81faa88addf6ad037d94 Mon Sep 17 00:00:00 2001
From: Jinwoo Hong <73622457+Jinwoo-H@users.noreply.github.com>
Date: Fri, 4 Sep 2026 01:40:55 -0400
Subject: [PATCH 23/49] docs(cloud): reconcile the 2026-08-23 retry figure with
the gate metric (#18581)
---
cloud/docs/relay-incident-monitor.md | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)
diff --git a/cloud/docs/relay-incident-monitor.md b/cloud/docs/relay-incident-monitor.md
index 5c563f6a2e2..870c95dd413 100644
--- a/cloud/docs/relay-incident-monitor.md
+++ b/cloud/docs/relay-incident-monitor.md
@@ -138,7 +138,9 @@ heartbeats, and matching live admission.
`jsonPayload.event="orca_relay_postgres_transaction_retry"` in production
logs: healthy-day bursts reach 234/5min with zero exhausted retries and 26%
of five-minute windows over 20, while the 2026-08-23 lock-contention
- incident ran roughly 2,200–3,000/5min.
+ incident ran roughly 2,200–3,000/5min by raw log-line count (the gate's
+ own `orca_relay_postgres_retries` metric read 1,510 for that window; see the
+ 2026-09-04 entry).
- Recalibrated the PostgreSQL-retry freeze from 300 to 2,000 per five minutes
(2026-09-04). Basis: the global `relay_cells FOR UPDATE` lock made
successful retries a steady-state rate. Measured fleet-wide (director +
From 561a94038c3ab550be22bbf680f15e06b6afd7dc Mon Sep 17 00:00:00 2001
From: Neil <4138956+nwparker@users.noreply.github.com>
Date: Fri, 4 Sep 2026 00:06:09 -0700
Subject: [PATCH 24/49] fix(ssh): stop the daemon's own services from blocking
the superseded-relay reap (#18586)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
`isReapableRelayHusk` required `childCount === 0`, where `childCount` came from
`pgrep -P | grep -c .`. But the relay forks service children of its own,
and `relay-ai-vault-service.js` never exits once spawned. Any relay that had
served a single AI Vault request therefore reported a non-zero child count
forever, so the sweep answered `retained-live-work` for a superseded,
disconnected relay holding no user work at all — and its version directory
stayed pinned against GC by its own live socket.
The probe now censuses each direct child instead of counting them, and the reap
gate reads the count of children it could *not* positively identify as relay
infrastructure. The asymmetry is the safety argument
(docs/reference/ssh-execution-boundary.md): subtracting a child we can name is
positive knowledge, assuming about one we cannot is not. An unrecognised argv,
an argv `ps` would not print, and a host without `pgrep` all keep the relay
unreapable. `reapEmptyRelayHuskCommand` re-runs the same census on the host
immediately before signalling.
Fixes #13614
---
src/main/ssh/relay-daemon-service-children.ts | 62 +++++++++
...dpoint-incumbent-shell.integration.test.ts | 118 ++++++++++++++++--
.../ssh/ssh-relay-endpoint-incumbent.test.ts | 64 +++++++---
src/main/ssh/ssh-relay-endpoint-incumbent.ts | 47 +++++--
.../ssh/ssh-relay-endpoint-takeover.test.ts | 29 +++--
src/main/ssh/ssh-relay-endpoint-takeover.ts | 15 ++-
.../ssh-relay-superseded-endpoints.test.ts | 10 +-
src/shared/relay-artifacts.ts | 18 ++-
8 files changed, 307 insertions(+), 56 deletions(-)
create mode 100644 src/main/ssh/relay-daemon-service-children.ts
diff --git a/src/main/ssh/relay-daemon-service-children.ts b/src/main/ssh/relay-daemon-service-children.ts
new file mode 100644
index 00000000000..35e755ea518
--- /dev/null
+++ b/src/main/ssh/relay-daemon-service-children.ts
@@ -0,0 +1,62 @@
+/**
+ * Telling a relay daemon's own service processes apart from the work it holds.
+ *
+ * The reap gate used to ask `pgrep -P | grep -c .` and demand zero. But the daemon
+ * forks service children of its own — `relay-ai-vault-service.js` is spawned lazily and then
+ * never exits — so that count is permanently non-zero on any relay that has touched the AI
+ * Vault, whether or not it holds a single PTY. A superseded, disconnected relay holding
+ * nothing therefore reported `retained-live-work` forever, its version directory stayed
+ * pinned against GC by its own live socket, and the population grew without bound (#13614).
+ *
+ * The asymmetry below is the whole safety argument, and it follows
+ * docs/reference/ssh-execution-boundary.md: *subtracting a child we can positively identify
+ * as relay infrastructure is sound; assuming anything about a child we cannot identify is
+ * not.* An argv that does not match, an argv `ps` would not print, and a host without
+ * `pgrep` all count against the relay and keep it unreapable. Losing sight of a child is
+ * never evidence that it holds nothing.
+ */
+import { RELAY_DAEMON_SERVICE_ENTRY_FILENAMES } from '../../shared/relay-artifacts'
+import { shellEscape } from './ssh-connection-utils'
+
+/** Shell variable set to the daemon's direct-child count, or `unknown`. */
+export const RELAY_CHILD_COUNT_VAR = 'kids'
+
+/** Shell variable set to the count of children not identified as relay services, or `unknown`. */
+export const RELAY_UNRECOGNIZED_CHILD_COUNT_VAR = 'unrecognized_kids'
+
+/**
+ * `case` patterns matching a service child's argv. Suffix-anchored on purpose: both entries
+ * are forked with no script arguments, so the argv ends at the filename, and the leading `/`
+ * requires the absolute path the daemon forks rather than a bare mention of the name. A
+ * future arg would stop matching and the relay would go back to being retained — the safe
+ * direction to fail in.
+ */
+function serviceChildArgvPatterns(): string {
+ return RELAY_DAEMON_SERVICE_ENTRY_FILENAMES.map(
+ (filename) => `*${shellEscape(`/${filename}`)}`
+ ).join('|')
+}
+
+/**
+ * POSIX shell that censuses the direct children of `$pid`, setting `kids` and
+ * `unrecognized_kids`. Both stay `unknown` when the host cannot enumerate children at all.
+ */
+export function relayDaemonChildCensusShell(): string[] {
+ return [
+ `${RELAY_CHILD_COUNT_VAR}=unknown`,
+ `${RELAY_UNRECOGNIZED_CHILD_COUNT_VAR}=unknown`,
+ 'if command -v pgrep >/dev/null 2>&1; then',
+ ` ${RELAY_CHILD_COUNT_VAR}=0`,
+ ` ${RELAY_UNRECOGNIZED_CHILD_COUNT_VAR}=0`,
+ ' for kid in $(pgrep -P "$pid" 2>/dev/null); do',
+ ` ${RELAY_CHILD_COUNT_VAR}=$((${RELAY_CHILD_COUNT_VAR}+1))`,
+ ' kid_args=$(ps -o args= -p "$kid" 2>/dev/null | tr -d "\\n")',
+ ' case "$kid_args" in',
+ ` ${serviceChildArgvPatterns()}) ;;`,
+ // An unreadable or unrecognised argv lands here, which is what keeps the relay retained.
+ ` *) ${RELAY_UNRECOGNIZED_CHILD_COUNT_VAR}=$((${RELAY_UNRECOGNIZED_CHILD_COUNT_VAR}+1)) ;;`,
+ ' esac',
+ ' done',
+ 'fi'
+ ]
+}
diff --git a/src/main/ssh/ssh-relay-endpoint-incumbent-shell.integration.test.ts b/src/main/ssh/ssh-relay-endpoint-incumbent-shell.integration.test.ts
index 7ece0b6532e..a8975d0520b 100644
--- a/src/main/ssh/ssh-relay-endpoint-incumbent-shell.integration.test.ts
+++ b/src/main/ssh/ssh-relay-endpoint-incumbent-shell.integration.test.ts
@@ -4,7 +4,7 @@
* generated scripts through /bin/sh against real unix sockets and real processes.
*/
import { execFile, spawn, type ChildProcess } from 'node:child_process'
-import { mkdtempSync, rmSync, writeFileSync } from 'node:fs'
+import { mkdirSync, mkdtempSync, rmSync, symlinkSync, writeFileSync } from 'node:fs'
import { tmpdir } from 'node:os'
import { join } from 'node:path'
import { afterAll, afterEach, beforeAll, describe, expect, it } from 'vitest'
@@ -15,21 +15,32 @@ import {
type RelayEndpointIncumbent
} from './ssh-relay-endpoint-incumbent'
import { reapEmptyRelayHuskCommand } from './ssh-relay-endpoint-takeover'
+import { RELAY_DAEMON_SERVICE_ENTRY_FILENAMES } from '../../shared/relay-artifacts'
const posixOnly = process.platform === 'win32' ? describe.skip : describe
const FAKE_RELAY_SOURCE = `
const net = require('net')
+const path = require('path')
const sock = process.argv[process.argv.indexOf('--sock-path') + 1]
+function spawnChild(args) {
+ require('child_process').spawn(process.execPath, args, { stdio: 'ignore' })
+}
if (process.argv.includes('--with-child')) {
- require('child_process').spawn(process.execPath, ['-e', 'setInterval(() => {}, 1000)'], {
- stdio: 'ignore'
- })
+ spawnChild(['-e', 'setTimeout(() => {}, 60000)'])
+}
+// Why forked the same way production does: the exclusion is argv-shaped, so a hand-written
+// stand-in would test the test rather than the shell that runs on someone's host.
+for (const name of process.argv.filter((arg) => arg.startsWith('--service-child='))) {
+ spawnChild([path.join(__dirname, name.slice('--service-child='.length))])
}
net.createServer(() => {}).listen(sock, () => process.stdout.write('READY\\n'))
process.on('SIGTERM', () => process.exit(0))
`
+// Self-limiting: these are orphaned when the relay under test is reaped.
+const IDLE_SERVICE_SOURCE = 'setTimeout(() => {}, 60000)\n'
+
function sh(script: string): Promise {
return new Promise((resolve, reject) => {
execFile('/bin/sh', ['-c', script], { timeout: 20_000 }, (error, stdout) => {
@@ -43,14 +54,21 @@ function sh(script: string): Promise {
}
let workDir: string
+let pgreplessBinDir: string
let hasLsof = false
const running: ChildProcess[] = []
-function startFakeRelay(sockPath: string, withChild = false): Promise {
+function startFakeRelay(
+ sockPath: string,
+ options: { withChild?: boolean; serviceChildren?: readonly string[] } = {}
+): Promise {
const args = [join(workDir, 'relay.js'), '--sock-path', sockPath]
- if (withChild) {
+ if (options.withChild) {
args.push('--with-child')
}
+ for (const name of options.serviceChildren ?? []) {
+ args.push(`--service-child=${name}`)
+ }
const child = spawn(process.execPath, args, { stdio: ['ignore', 'pipe', 'ignore'] })
running.push(child)
return new Promise((resolve, reject) => {
@@ -68,9 +86,31 @@ async function probe(sockPath: string): Promise {
return parseRelayEndpointIncumbentProbe(sockPath, output)
}
+/** The relay forks its children after it starts listening, so the probe can race them. */
+async function waitForChildCount(
+ sockPath: string,
+ expected: number
+): Promise {
+ let incumbent = await probe(sockPath)
+ for (let attempt = 0; attempt < 50 && incumbent.holders[0]?.childCount !== expected; attempt++) {
+ await new Promise((resolve) => setTimeout(resolve, 100))
+ incumbent = await probe(sockPath)
+ }
+ return incumbent
+}
+
beforeAll(async () => {
workDir = mkdtempSync(join(tmpdir(), 'orca-relay-incumbent-'))
writeFileSync(join(workDir, 'relay.js'), FAKE_RELAY_SOURCE)
+ for (const filename of RELAY_DAEMON_SERVICE_ENTRY_FILENAMES) {
+ writeFileSync(join(workDir, filename), IDLE_SERVICE_SOURCE)
+ }
+ writeFileSync(join(workDir, 'looks-like-relay-watcher.js'), IDLE_SERVICE_SOURCE)
+ pgreplessBinDir = join(workDir, 'pgrepless-bin')
+ mkdirSync(pgreplessBinDir)
+ for (const tool of ['ps', 'tr']) {
+ symlinkSync((await sh(`command -v ${tool}`)).trim(), join(pgreplessBinDir, tool))
+ }
hasLsof = await sh('command -v lsof >/dev/null 2>&1 && echo yes || echo no').then(
(out) => out.trim() === 'yes'
)
@@ -105,13 +145,51 @@ posixOnly('relay endpoint probe against a real socket', () => {
return
}
expect(incumbent.holders.map((holder) => holder.pid)).toEqual([relay.pid])
- expect(incumbent.holders[0]).toMatchObject({ matchesRelayArgv: true, childCount: 0 })
+ expect(incumbent.holders[0]).toMatchObject({
+ matchesRelayArgv: true,
+ childCount: 0,
+ unrecognizedChildCount: 0
+ })
expect(isReapableRelayHusk(incumbent)).toBe(true)
})
+ it("counts the daemon's own service children but does not hold them against it", async () => {
+ const sockPath = join(workDir, 'services.sock')
+ await startFakeRelay(sockPath, { serviceChildren: RELAY_DAEMON_SERVICE_ENTRY_FILENAMES })
+ const incumbent = await waitForChildCount(sockPath, RELAY_DAEMON_SERVICE_ENTRY_FILENAMES.length)
+
+ expect(incumbent.holders[0].childCount).toBe(RELAY_DAEMON_SERVICE_ENTRY_FILENAMES.length)
+ expect(incumbent.holders[0].unrecognizedChildCount).toBe(0)
+ expect(isReapableRelayHusk(incumbent)).toBe(true)
+ })
+
+ it('still retains a relay holding work alongside its service children', async () => {
+ const sockPath = join(workDir, 'services-and-work.sock')
+ await startFakeRelay(sockPath, {
+ withChild: true,
+ serviceChildren: RELAY_DAEMON_SERVICE_ENTRY_FILENAMES
+ })
+ const incumbent = await waitForChildCount(
+ sockPath,
+ RELAY_DAEMON_SERVICE_ENTRY_FILENAMES.length + 1
+ )
+
+ expect(incumbent.holders[0].unrecognizedChildCount).toBe(1)
+ expect(isReapableRelayHusk(incumbent)).toBe(false)
+ })
+
+ it('does not excuse a child that merely mentions a service entry name', async () => {
+ const sockPath = join(workDir, 'lookalike.sock')
+ await startFakeRelay(sockPath, { serviceChildren: ['looks-like-relay-watcher.js'] })
+ const incumbent = await waitForChildCount(sockPath, 1)
+
+ expect(incumbent.holders[0].unrecognizedChildCount).toBe(1)
+ expect(isReapableRelayHusk(incumbent)).toBe(false)
+ })
+
it('refuses to call a relay with a live child an empty husk', async () => {
const sockPath = join(workDir, 'busy.sock')
- await startFakeRelay(sockPath, true)
+ await startFakeRelay(sockPath, { withChild: true })
const incumbent = await probe(sockPath)
expect(incumbent.verdict).toBe('live')
@@ -150,12 +228,34 @@ posixOnly('empty relay husk reap against a real process', () => {
it('refuses to signal a relay that acquired a child after it was probed', async () => {
const sockPath = join(workDir, 'raced.sock')
- const relay = await startFakeRelay(sockPath, true)
+ const relay = await startFakeRelay(sockPath, { withChild: true })
const output = await sh(reapEmptyRelayHuskCommand(relay.pid!, sockPath))
expect(output.trim()).toBe('BUSY')
expect(relay.killed).toBe(false)
})
+ it('terminates a relay whose only children are its own service processes (#13614)', async () => {
+ const sockPath = join(workDir, 'service-husk.sock')
+ const relay = await startFakeRelay(sockPath, {
+ serviceChildren: RELAY_DAEMON_SERVICE_ENTRY_FILENAMES
+ })
+ await waitForChildCount(sockPath, RELAY_DAEMON_SERVICE_ENTRY_FILENAMES.length)
+ const output = await sh(reapEmptyRelayHuskCommand(relay.pid!, sockPath))
+ expect(output.trim()).toBe('GONE')
+ })
+
+ it('refuses to signal when the host cannot enumerate children at all', async () => {
+ const sockPath = join(workDir, 'no-pgrep.sock')
+ const relay = await startFakeRelay(sockPath)
+ // A PATH carrying every tool the script needs except `pgrep`: the census answers
+ // `unknown`, which must reach BUSY rather than the zero a missing tool would imply.
+ const output = await sh(
+ `PATH=${pgreplessBinDir}\n${reapEmptyRelayHuskCommand(relay.pid!, sockPath)}`
+ )
+ expect(output.trim()).toBe('BUSY')
+ expect(relay.killed).toBe(false)
+ })
+
it('refuses to signal a pid whose argv is not this relay at this socket', async () => {
const sockPath = join(workDir, 'mismatch.sock')
await startFakeRelay(sockPath)
diff --git a/src/main/ssh/ssh-relay-endpoint-incumbent.test.ts b/src/main/ssh/ssh-relay-endpoint-incumbent.test.ts
index a65cb33fc57..de4cc28d170 100644
--- a/src/main/ssh/ssh-relay-endpoint-incumbent.test.ts
+++ b/src/main/ssh/ssh-relay-endpoint-incumbent.test.ts
@@ -32,17 +32,24 @@ describe('parseRelayEndpointIncumbentProbe', () => {
it('reports live when the socket accepted a connection', () => {
const incumbent = parseRelayEndpointIncumbentProbe(
SOCK,
- probeOutput(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=4242 yes 13'])
+ probeOutput([
+ 'PRESENT=yes',
+ 'LISTEN=accepted',
+ 'HOLDERS_SOURCE=lsof',
+ 'HOLDER=4242 yes 13 11'
+ ])
)
expect(incumbent.verdict).toBe('live')
expect(incumbent.evidence).toBe('accepted-connection')
- expect(incumbent.holders).toEqual([{ pid: 4242, matchesRelayArgv: true, childCount: 13 }])
+ expect(incumbent.holders).toEqual([
+ { pid: 4242, matchesRelayArgv: true, childCount: 13, unrecognizedChildCount: 11 }
+ ])
})
it('reports live when a process still holds an inode that refuses connections', () => {
const incumbent = parseRelayEndpointIncumbentProbe(
SOCK,
- probeOutput(['PRESENT=yes', 'LISTEN=refused', 'HOLDERS_SOURCE=lsof', 'HOLDER=91 yes 2'])
+ probeOutput(['PRESENT=yes', 'LISTEN=refused', 'HOLDERS_SOURCE=lsof', 'HOLDER=91 yes 2 2'])
)
expect(incumbent.verdict).toBe('live')
expect(incumbent.evidence).toBe('holder-process')
@@ -85,7 +92,12 @@ describe('parseRelayEndpointIncumbentProbe', () => {
it('drops holder lines that do not carry a usable pid', () => {
const incumbent = parseRelayEndpointIncumbentProbe(
SOCK,
- probeOutput(['PRESENT=yes', 'LISTEN=refused', 'HOLDERS_SOURCE=lsof', 'HOLDER=- no unknown'])
+ probeOutput([
+ 'PRESENT=yes',
+ 'LISTEN=refused',
+ 'HOLDERS_SOURCE=lsof',
+ 'HOLDER=- no unknown unknown'
+ ])
)
expect(incumbent.holders).toEqual([])
expect(incumbent.verdict).toBe('exited')
@@ -94,9 +106,24 @@ describe('parseRelayEndpointIncumbentProbe', () => {
it('keeps an unreadable child count as null rather than zero', () => {
const [holder] = parseRelayEndpointIncumbentProbe(
SOCK,
- probeOutput(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=7 yes unknown'])
+ probeOutput([
+ 'PRESENT=yes',
+ 'LISTEN=accepted',
+ 'HOLDERS_SOURCE=lsof',
+ 'HOLDER=7 yes unknown unknown'
+ ])
).holders
expect(holder.childCount).toBeNull()
+ expect(holder.unrecognizedChildCount).toBeNull()
+ })
+
+ it('keeps a holder line with no unrecognized-child field unreapable', () => {
+ const incumbent = parseRelayEndpointIncumbentProbe(
+ SOCK,
+ probeOutput(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=7 yes 0'])
+ )
+ expect(incumbent.holders[0].unrecognizedChildCount).toBeNull()
+ expect(isReapableRelayHusk(incumbent)).toBe(false)
})
})
@@ -172,27 +199,36 @@ describe('mayLaunchOverRelayEndpoint', () => {
describe('isReapableRelayHusk', () => {
const husk = parseRelayEndpointIncumbentProbe(
SOCK,
- probeOutput(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=500 yes 0'])
+ probeOutput(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=500 yes 0 0'])
)
- it('accepts a single proven relay holder with zero children', () => {
+ it('accepts a single proven relay holder with no unaccounted-for children', () => {
expect(isReapableRelayHusk(husk)).toBe(true)
})
- it('refuses a relay that still holds children', () => {
+ it('accepts a relay whose only children are its own service processes (#13614)', () => {
+ const withServices = parseRelayEndpointIncumbentProbe(
+ SOCK,
+ probeOutput(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=500 yes 2 0'])
+ )
+ expect(withServices.holders[0].childCount).toBe(2)
+ expect(isReapableRelayHusk(withServices)).toBe(true)
+ })
+
+ it('refuses a relay that still holds children it could not account for', () => {
expect(
isReapableRelayHusk({
...husk,
- holders: [{ pid: 500, matchesRelayArgv: true, childCount: 1 }]
+ holders: [{ pid: 500, matchesRelayArgv: true, childCount: 3, unrecognizedChildCount: 1 }]
})
).toBe(false)
})
- it('refuses a holder whose child count could not be read', () => {
+ it('refuses a holder whose unrecognized-child count could not be read', () => {
expect(
isReapableRelayHusk({
...husk,
- holders: [{ pid: 500, matchesRelayArgv: true, childCount: null }]
+ holders: [{ pid: 500, matchesRelayArgv: true, childCount: 0, unrecognizedChildCount: null }]
})
).toBe(false)
})
@@ -201,7 +237,7 @@ describe('isReapableRelayHusk', () => {
expect(
isReapableRelayHusk({
...husk,
- holders: [{ pid: 500, matchesRelayArgv: false, childCount: 0 }]
+ holders: [{ pid: 500, matchesRelayArgv: false, childCount: 0, unrecognizedChildCount: 0 }]
})
).toBe(false)
})
@@ -211,8 +247,8 @@ describe('isReapableRelayHusk', () => {
isReapableRelayHusk({
...husk,
holders: [
- { pid: 500, matchesRelayArgv: true, childCount: 0 },
- { pid: 501, matchesRelayArgv: true, childCount: 0 }
+ { pid: 500, matchesRelayArgv: true, childCount: 0, unrecognizedChildCount: 0 },
+ { pid: 501, matchesRelayArgv: true, childCount: 0, unrecognizedChildCount: 0 }
]
})
).toBe(false)
diff --git a/src/main/ssh/ssh-relay-endpoint-incumbent.ts b/src/main/ssh/ssh-relay-endpoint-incumbent.ts
index 2688267f4c7..628a9558793 100644
--- a/src/main/ssh/ssh-relay-endpoint-incumbent.ts
+++ b/src/main/ssh/ssh-relay-endpoint-incumbent.ts
@@ -20,6 +20,11 @@
*/
import type { SshConnection } from './ssh-connection'
import { shellEscape } from './ssh-connection-utils'
+import {
+ RELAY_CHILD_COUNT_VAR,
+ RELAY_UNRECOGNIZED_CHILD_COUNT_VAR,
+ relayDaemonChildCensusShell
+} from './relay-daemon-service-children'
import { execCommand, isUnconfirmedSshCommandTermination } from './ssh-relay-deploy-helpers'
import { isWindowsRemoteHost, type RemoteHostPlatform } from './ssh-remote-platform'
@@ -38,6 +43,12 @@ export type RelayEndpointHolder = {
matchesRelayArgv: boolean
/** Direct children, or null when `pgrep` could not answer. Never guessed. */
childCount: number | null
+ /**
+ * Direct children *not* positively identified as the daemon's own service processes, or
+ * null when the host could not enumerate them. This — not `childCount` — is what says
+ * whether the relay holds anything; see relay-daemon-service-children.ts.
+ */
+ unrecognizedChildCount: number | null
}
export type RelayEndpointIncumbent = {
@@ -95,11 +106,9 @@ export function relayEndpointIncumbentProbeCommand(nodePath: string, sockPath: s
' args=$(ps -o args= -p "$pid" 2>/dev/null | tr "\\n" " ")',
' match=no',
' case "$args" in *relay.js*"$sock"*) match=yes ;; esac',
- ' kids=unknown',
- ' if command -v pgrep >/dev/null 2>&1; then',
- ' kids=$(pgrep -P "$pid" 2>/dev/null | grep -c .)',
- ' fi',
- ' printf \'HOLDER=%s %s %s\\n\' "$pid" "$match" "$kids"',
+ ...relayDaemonChildCensusShell().map((line) => ` ${line}`),
+ ' printf \'HOLDER=%s %s %s %s\\n\' "$pid" "$match" ' +
+ `"$${RELAY_CHILD_COUNT_VAR}" "$${RELAY_UNRECOGNIZED_CHILD_COUNT_VAR}"`,
' done',
'else',
" printf 'HOLDERS_SOURCE=unavailable\\n'",
@@ -159,19 +168,25 @@ export function parseRelayEndpointIncumbentProbe(
}
function parseHolder(value: string): RelayEndpointHolder | null {
- const [rawPid, rawMatch, rawKids] = value.split(/\s+/)
+ const [rawPid, rawMatch, rawKids, rawUnrecognized] = value.split(/\s+/)
const pid = Number.parseInt(rawPid ?? '', 10)
if (!Number.isInteger(pid) || pid <= 0) {
return null
}
- const childCount = Number.parseInt(rawKids ?? '', 10)
return {
pid,
matchesRelayArgv: rawMatch === 'yes',
- childCount: Number.isInteger(childCount) && childCount >= 0 ? childCount : null
+ childCount: parseChildCount(rawKids),
+ unrecognizedChildCount: parseChildCount(rawUnrecognized)
}
}
+/** `unknown`, a missing field, and anything unparseable are all "could not tell" — never 0. */
+function parseChildCount(raw: string | undefined): number | null {
+ const count = Number.parseInt(raw ?? '', 10)
+ return Number.isInteger(count) && count >= 0 ? count : null
+}
+
function unverifiableEndpoint(sockPath: string): RelayEndpointIncumbent {
return {
sockPath,
@@ -239,8 +254,12 @@ export function mayLaunchOverRelayEndpoint(incumbent: RelayEndpointIncumbent): b
/**
* A live relay that provably holds nothing: identity confirmed against its argv, exactly one
- * holder, and zero children. Reaping it destroys no user work. Anything less is retained —
- * killing the wrong pid on someone's remote host is the worst outcome available here.
+ * holder, and no child the host could not account for as one of the daemon's own service
+ * processes. Reaping it destroys no user work. Anything less is retained — killing the wrong
+ * pid on someone's remote host is the worst outcome available here.
+ *
+ * Why not `childCount === 0`: the daemon's AI Vault sidecar never exits once spawned, so that
+ * gate was unreachable for any relay that had ever served a vault request (#13614).
*/
export function isReapableRelayHusk(incumbent: RelayEndpointIncumbent): boolean {
if (incumbent.verdict !== 'live' || !incumbent.holdersEnumerable) {
@@ -250,12 +269,16 @@ export function isReapableRelayHusk(incumbent: RelayEndpointIncumbent): boolean
return false
}
const [holder] = incumbent.holders
- return holder.matchesRelayArgv && holder.childCount === 0
+ return holder.matchesRelayArgv && holder.unrecognizedChildCount === 0
}
export function describeRelayEndpointIncumbent(incumbent: RelayEndpointIncumbent): string {
const holders = incumbent.holders
- .map((holder) => `${holder.pid}(children=${holder.childCount ?? 'unknown'})`)
+ .map(
+ (holder) =>
+ `${holder.pid}(children=${holder.childCount ?? 'unknown'},` +
+ `unrecognized=${holder.unrecognizedChildCount ?? 'unknown'})`
+ )
.join(',')
return (
`${incumbent.sockPath} verdict=${incumbent.verdict} evidence=${incumbent.evidence} ` +
diff --git a/src/main/ssh/ssh-relay-endpoint-takeover.test.ts b/src/main/ssh/ssh-relay-endpoint-takeover.test.ts
index 687d633b92b..d23f065f478 100644
--- a/src/main/ssh/ssh-relay-endpoint-takeover.test.ts
+++ b/src/main/ssh/ssh-relay-endpoint-takeover.test.ts
@@ -14,6 +14,7 @@ import {
resolveRelayEndpointBeforeRelaunch
} from './ssh-relay-endpoint-takeover'
import { RelayVersionMismatchError } from './ssh-relay-version-mismatch-error'
+import { RELAY_DAEMON_SERVICE_ENTRY_FILENAMES } from '../../shared/relay-artifacts'
import type { SshConnection } from './ssh-connection'
import { getRemoteHostPlatform } from './ssh-remote-platform'
@@ -42,7 +43,7 @@ beforeEach(() => {
describe('incumbent alive and refusing', () => {
it('refuses to rebind a live relay holding PTYs, and signals nothing', async () => {
execCommand.mockResolvedValueOnce(
- probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=3669803 yes 13'])
+ probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=3669803 yes 13 11'])
)
await expect(resolve()).rejects.toSatisfy(isRelayEndpointHeldError)
// The whole point of #8585: the incumbent's socket must survive so it is not orphaned.
@@ -52,9 +53,9 @@ describe('incumbent alive and refusing', () => {
it('names the incumbent pid and the Reset Relay escape hatch in the error', async () => {
execCommand.mockResolvedValue(
- probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=3669803 yes 13'])
+ probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=3669803 yes 13 11'])
)
- await expect(resolve()).rejects.toThrow(/3669803\(children=13\)/)
+ await expect(resolve()).rejects.toThrow(/3669803\(children=13,unrecognized=11\)/)
await expect(resolve()).rejects.toThrow(/Reset Relay/)
})
@@ -70,7 +71,7 @@ describe('incumbent alive and refusing', () => {
it('reaps a live relay only when it provably holds nothing, and confirms it is gone', async () => {
execCommand
.mockResolvedValueOnce(
- probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=80583 yes 0'])
+ probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=80583 yes 2 0'])
)
.mockResolvedValueOnce('GONE\n')
await expect(resolve()).resolves.toMatchObject({ verdict: 'live' })
@@ -80,7 +81,7 @@ describe('incumbent alive and refusing', () => {
it('does not launch over an empty relay whose death could not be confirmed', async () => {
execCommand
.mockResolvedValueOnce(
- probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=80583 yes 0'])
+ probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=80583 yes 2 0'])
)
.mockResolvedValueOnce('LIVE\n')
await expect(resolve()).rejects.toSatisfy(isRelayEndpointHeldError)
@@ -89,7 +90,7 @@ describe('incumbent alive and refusing', () => {
it('does not launch over a relay the host refused to signal on its own re-check', async () => {
execCommand
.mockResolvedValueOnce(
- probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=80583 yes 0'])
+ probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=80583 yes 2 0'])
)
.mockResolvedValueOnce('BUSY\n')
await expect(resolve()).rejects.toSatisfy(isRelayEndpointHeldError)
@@ -138,9 +139,19 @@ describe('reapEmptyRelayHuskCommand', () => {
})
it('aborts without signalling when the host cannot count children', () => {
- expect(reapEmptyRelayHuskCommand(4242, SOCK)).toContain(
- "command -v pgrep >/dev/null 2>&1 || { printf 'BUSY\\n'; exit 0; }"
- )
+ const command = reapEmptyRelayHuskCommand(4242, SOCK)
+ // The census leaves both counters at `unknown` without pgrep, and the gate demands "0".
+ expect(command).toContain('unrecognized_kids=unknown')
+ expect(command).toContain('command -v pgrep >/dev/null 2>&1')
+ expect(command).toContain('[ "$unrecognized_kids" = "0" ] ||')
+ })
+
+ it('subtracts only the daemon service children it can name from the reap gate', () => {
+ const command = reapEmptyRelayHuskCommand(4242, SOCK)
+ for (const filename of RELAY_DAEMON_SERVICE_ENTRY_FILENAMES) {
+ expect(command).toContain(`*'/${filename}'`)
+ }
+ expect(command).toContain('unrecognized_kids=$((unrecognized_kids+1))')
})
})
diff --git a/src/main/ssh/ssh-relay-endpoint-takeover.ts b/src/main/ssh/ssh-relay-endpoint-takeover.ts
index f104aab5256..8f6130620cb 100644
--- a/src/main/ssh/ssh-relay-endpoint-takeover.ts
+++ b/src/main/ssh/ssh-relay-endpoint-takeover.ts
@@ -2,13 +2,17 @@
* Deciding whether a relay socket path is ours to take, and acting on the answer.
*
* The only destructive action available here is a SIGTERM to a relay that has been proven —
- * by argv, by socket-holder enumeration, and by a zero child count re-checked on the host
- * immediately before the signal — to hold nothing at all. Everything else is left running.
+ * by argv, by socket-holder enumeration, and by a child census re-run on the host immediately
+ * before the signal — to hold nothing at all. Everything else is left running.
* Per docs/reference/ssh-execution-boundary.md, a relay we merely failed to reach is
* `unverifiable`, and `unverifiable` never authorizes a kill or a rebind.
*/
import type { SshConnection } from './ssh-connection'
import { shellEscape } from './ssh-connection-utils'
+import {
+ RELAY_UNRECOGNIZED_CHILD_COUNT_VAR,
+ relayDaemonChildCensusShell
+} from './relay-daemon-service-children'
import { execCommand, isUnconfirmedSshCommandTermination } from './ssh-relay-deploy-helpers'
import {
describeRelayEndpointIncumbent,
@@ -39,9 +43,10 @@ export function reapEmptyRelayHuskCommand(pid: number, sockPath: string): string
`sock=${shellEscape(sockPath)}`,
'args=$(ps -o args= -p "$pid" 2>/dev/null | tr "\\n" " ")',
'case "$args" in *relay.js*"$sock"*) ;; *) printf \'MISMATCH\\n\'; exit 0 ;; esac',
- "command -v pgrep >/dev/null 2>&1 || { printf 'BUSY\\n'; exit 0; }",
- 'kids=$(pgrep -P "$pid" 2>/dev/null | grep -c .)',
- '[ "$kids" = "0" ] || { printf \'BUSY\\n\'; exit 0; }',
+ // Why the same census as the probe: `unknown` (no pgrep) and any child this host could
+ // not account for as a relay service both land on BUSY, so nothing is signalled.
+ ...relayDaemonChildCensusShell(),
+ `[ "$${RELAY_UNRECOGNIZED_CHILD_COUNT_VAR}" = "0" ] || { printf 'BUSY\\n'; exit 0; }`,
// SIGTERM only: the relay's own handler disposes and unlinks. SIGKILL would leave the
// socket inode behind and skip that shutdown path for no gain on an empty daemon.
'kill -TERM "$pid" 2>/dev/null || true',
diff --git a/src/main/ssh/ssh-relay-superseded-endpoints.test.ts b/src/main/ssh/ssh-relay-superseded-endpoints.test.ts
index 874d9aae3fe..9168f4688bd 100644
--- a/src/main/ssh/ssh-relay-superseded-endpoints.test.ts
+++ b/src/main/ssh/ssh-relay-superseded-endpoints.test.ts
@@ -67,7 +67,7 @@ describe('classifySupersededRelay', () => {
'PRESENT=yes',
'LISTEN=accepted',
'HOLDERS_SOURCE=lsof',
- 'HOLDER=3669803 yes 13'
+ 'HOLDER=3669803 yes 13 11'
])
)
).toBe('retained-live-work')
@@ -76,7 +76,7 @@ describe('classifySupersededRelay', () => {
it('nominates only a proven empty relay for reaping', () => {
expect(
classifySupersededRelay(
- incumbent(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=80583 yes 0'])
+ incumbent(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=80583 yes 2 0'])
)
).toBe('reap-candidate')
})
@@ -101,7 +101,7 @@ describe('sweepSupersededRelayEndpoints', () => {
execCommand
.mockResolvedValueOnce(`${OLD_SOCK}\n`)
.mockResolvedValueOnce(
- probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=3669803 yes 13'])
+ probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=3669803 yes 13 11'])
)
const findings = await sweepSupersededRelayEndpoints(CONN, HOST, SWEEP)
expect(findings).toHaveLength(1)
@@ -114,7 +114,7 @@ describe('sweepSupersededRelayEndpoints', () => {
execCommand
.mockResolvedValueOnce(`${OLD_SOCK}\n`)
.mockResolvedValueOnce(
- probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=80583 yes 0'])
+ probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=80583 yes 2 0'])
)
.mockResolvedValueOnce('GONE\n')
const findings = await sweepSupersededRelayEndpoints(CONN, HOST, SWEEP)
@@ -126,7 +126,7 @@ describe('sweepSupersededRelayEndpoints', () => {
execCommand
.mockResolvedValueOnce(`${OLD_SOCK}\n`)
.mockResolvedValueOnce(
- probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=80583 yes 0'])
+ probe(['PRESENT=yes', 'LISTEN=accepted', 'HOLDERS_SOURCE=lsof', 'HOLDER=80583 yes 2 0'])
)
.mockResolvedValueOnce('LIVE\n')
const findings = await sweepSupersededRelayEndpoints(CONN, HOST, SWEEP)
diff --git a/src/shared/relay-artifacts.ts b/src/shared/relay-artifacts.ts
index 273f6e059b8..2f9e563f839 100644
--- a/src/shared/relay-artifacts.ts
+++ b/src/shared/relay-artifacts.ts
@@ -37,6 +37,12 @@ export type RelayArtifact = {
* optional one would loop forever redeploying a relay that is already correct.
*/
optional?: boolean
+ /**
+ * Forked by the relay daemon as a long-lived child of its own. These are relay
+ * infrastructure, never user work, and the reap gate subtracts them from a daemon's
+ * child census; see src/main/ssh/relay-daemon-service-children.ts.
+ */
+ daemonServiceChild?: boolean
}
/** The bare Windows process-table addon; see docs/reference/windows-process-enumeration.md. */
@@ -44,8 +50,8 @@ export const RELAY_WINDOWS_PROCESS_TREE_FILENAME = 'windows-process-tree.node'
export const RELAY_ARTIFACTS: readonly RelayArtifact[] = [
{ filename: 'relay.js' },
- { filename: 'relay-watcher.js' },
- { filename: 'relay-ai-vault-service.js' },
+ { filename: 'relay-watcher.js', daemonServiceChild: true },
+ { filename: 'relay-ai-vault-service.js', daemonServiceChild: true },
{ filename: 'managed-hook-runtime.js' },
// Forked by the AI Vault title reader; without it a relay answers every WSL
// title request with no title and no error.
@@ -62,6 +68,14 @@ export const RELAY_ARTIFACTS: readonly RelayArtifact[] = [
{ filename: RELAY_WINDOWS_PROCESS_TREE_FILENAME, windowsOnly: true, optional: true }
]
+/**
+ * The daemon's own service children, by entry filename. Anything else under a relay pid is
+ * either user work or unidentified, and both keep the relay unreapable.
+ */
+export const RELAY_DAEMON_SERVICE_ENTRY_FILENAMES: readonly string[] = RELAY_ARTIFACTS.filter(
+ (artifact) => artifact.daemonServiceChild
+).map((artifact) => artifact.filename)
+
/** Written after the artifacts, so it is never an input to its own hash. */
export const RELAY_VERSION_FILENAME = '.version'
From b85510f3a9a3751b7320f18f6c753720c265ed86 Mon Sep 17 00:00:00 2001
From: Neil <4138956+nwparker@users.noreply.github.com>
Date: Fri, 4 Sep 2026 00:51:21 -0700
Subject: [PATCH 25/49] fix(terminal): warn about remote work when closing the
window or quitting (#18593)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
The native window-close warning was built from a local-only pty set: any
worktree with a connectionId was dropped whole, and any remote runtime pty
was filtered out. A build, test run, or agent on an SSH or Orca Remote host
was therefore structurally invisible to it, on every platform. The quit path
skipped the check entirely (#524), so remote work got no prompt at all.
Route both paths through the same probe the tab-close guard uses, so the two
cannot drift, and keep the verdict vocabulary of the SSH execution boundary:
only a host that answers "no children" suppresses the warning. An unreachable
host is `unverifiable`, never `exited`, so it warns rather than quitting
silently — with its own copy, because "could not reach the host" is a
different claim than "processes are running".
Quit still ignores local ptys, preserving #524: quitting is an unambiguous
instruction to end this machine's processes, but not to end execution on
someone else's, which a bounded relay grace period will SIGKILL once the
countdown expires.
The probe budget is 1.5s (vs the tab guard's 4s) because quit is time
sensitive; expiry raises the prompt, so an unreachable host costs a click
rather than the 15s RPC timeout or a silently orphaned build.
---
config/scripts/locale-ko-key-overrides.json | 2 +-
.../components/TerminalWorkspaceDialogs.tsx | 14 +-
.../terminal/pty-running-work-probe.ts | 88 +++++++
.../running-terminal-close-guard.test.ts | 12 +-
.../terminal/running-terminal-close-guard.ts | 38 +--
...terminal-tab-close-running-confirm.test.ts | 8 +-
.../window-close-running-work.test.ts | 231 ++++++++++++++++++
.../terminal/window-close-running-work.ts | 70 ++++++
.../use-terminal-editor-close-foundation.ts | 53 ++--
...tor-close-foundation.window-close.test.tsx | 103 ++++++++
src/renderer/src/i18n/locales/en.json | 3 +-
src/renderer/src/i18n/locales/es.json | 2 +-
src/renderer/src/i18n/locales/fr.json | 2 +-
src/renderer/src/i18n/locales/ja.json | 2 +-
src/renderer/src/i18n/locales/ko.json | 2 +-
src/renderer/src/i18n/locales/zh.json | 2 +-
src/shared/remote-execution-host-pty-id.ts | 14 ++
17 files changed, 575 insertions(+), 71 deletions(-)
create mode 100644 src/renderer/src/components/terminal/pty-running-work-probe.ts
create mode 100644 src/renderer/src/components/terminal/window-close-running-work.test.ts
create mode 100644 src/renderer/src/components/terminal/window-close-running-work.ts
create mode 100644 src/renderer/src/components/use-terminal-editor-close-foundation.window-close.test.tsx
create mode 100644 src/shared/remote-execution-host-pty-id.ts
diff --git a/config/scripts/locale-ko-key-overrides.json b/config/scripts/locale-ko-key-overrides.json
index f368ecc3cbc..bf5f62d1fa5 100644
--- a/config/scripts/locale-ko-key-overrides.json
+++ b/config/scripts/locale-ko-key-overrides.json
@@ -492,7 +492,7 @@
"ko": "agent CLI를 찾지 못했습니다. 하나를 설치하거나 설정에서 기본 agent를 선택하세요."
},
"auto.components.Terminal.7958465754": {
- "ko": "실행 중인 프로세스가 있는 로컬 terminals이 있습니다. 그래도 창을 닫으시겠습니까?"
+ "ko": "실행 중인 프로세스가 있는 terminals이 있습니다. 그래도 창을 닫으시겠습니까?"
},
"auto.components.Terminal.cdc9ac4b2d": {
"ko": "편집기"
diff --git a/src/renderer/src/components/TerminalWorkspaceDialogs.tsx b/src/renderer/src/components/TerminalWorkspaceDialogs.tsx
index 52ddbf1ba5d..bb3ffbfd621 100644
--- a/src/renderer/src/components/TerminalWorkspaceDialogs.tsx
+++ b/src/renderer/src/components/TerminalWorkspaceDialogs.tsx
@@ -24,6 +24,7 @@ export function TerminalWorkspaceDialogs({
saveDialogFile,
saveDialogFileId,
setWindowCloseDialogOpen,
+ windowCloseDialogKind,
windowCloseDialogOpen
} = controller
return (
@@ -82,10 +83,15 @@ export function TerminalWorkspaceDialogs({
{translate('auto.components.Terminal.2fa9c69ff3', 'Close Window?')}
- {translate(
- 'auto.components.Terminal.7958465754',
- 'There are local terminals with running processes. Close the window anyway?'
- )}
+ {windowCloseDialogKind === 'unverifiable'
+ ? translate(
+ 'auto.components.Terminal.b7c1f0a934',
+ 'A remote host could not be reached, so Orca cannot tell whether work is still running there. Close the window anyway?'
+ )
+ : translate(
+ 'auto.components.Terminal.7958465754',
+ 'There are terminals with running processes. Close the window anyway?'
+ )}
diff --git a/src/renderer/src/components/terminal/pty-running-work-probe.ts b/src/renderer/src/components/terminal/pty-running-work-probe.ts
new file mode 100644
index 00000000000..609b71f12ca
--- /dev/null
+++ b/src/renderer/src/components/terminal/pty-running-work-probe.ts
@@ -0,0 +1,88 @@
+import type { GlobalSettings } from '../../../../shared/global-settings-types'
+import { inspectRuntimeTerminalProcess } from '@/runtime/runtime-terminal-inspection'
+import { isRemoteExecutionHostPtyId } from '../../../../shared/remote-execution-host-pty-id'
+import { isClientOnlyUnverifiableInspection } from '../../../../shared/terminal-process-inspection'
+
+/**
+ * One probe answer in the fixed `live` / `unverifiable` / `exited` vocabulary of
+ * `docs/reference/ssh-execution-boundary.md`. `exited` is only ever produced by a host that
+ * answered; every failure to reach the owner — a rejection, a closed transport, or a deadline
+ * that expired first — stays `unverifiable`, because loss of contact is not evidence of death.
+ */
+export type PtyRunningWorkVerdict = 'live' | 'unverifiable' | 'exited'
+
+export type PtyRunningWorkProbe = {
+ ptyId: string
+ verdict: PtyRunningWorkVerdict
+ /** Why the owner could not be observed. Only set for `unverifiable`. */
+ reason?: string
+ /** The deadline expired before this pty's probe answered at all. */
+ timedOut: boolean
+ /** The pty is owned by a remote execution host (relay runtime or app SSH). */
+ remote: boolean
+}
+
+type ProbeSettings = Pick | null | undefined
+
+/**
+ * Probes every pty for running work and resolves at whichever comes first: every answer, or the
+ * deadline. Never rejects, and never reports a pty it did not hear back about as idle.
+ *
+ * Callers own the policy. This owns only the measurement, so the tab-close guard and the
+ * window-close guard cannot drift apart on what an unanswered remote host means.
+ */
+export async function probePtyRunningWork(
+ settings: ProbeSettings,
+ ptyIds: readonly string[],
+ options: { timeoutMs: number }
+): Promise {
+ if (ptyIds.length === 0) {
+ return []
+ }
+ const probes: PtyRunningWorkProbe[] = ptyIds.map((ptyId) => ({
+ ptyId,
+ verdict: 'unverifiable',
+ reason: 'probe_deadline',
+ timedOut: true,
+ remote: isRemoteExecutionHostPtyId(ptyId)
+ }))
+
+ const settle = Promise.all(
+ ptyIds.map(async (ptyId, index) => {
+ const probe = probes[index]
+ if (!probe) {
+ return
+ }
+ try {
+ const inspection = await inspectRuntimeTerminalProcess(settings, ptyId)
+ probe.timedOut = false
+ if (isClientOnlyUnverifiableInspection(inspection)) {
+ probe.verdict = 'unverifiable'
+ probe.reason = inspection.reason
+ return
+ }
+ probe.verdict = inspection.hasChildProcesses ? 'live' : 'exited'
+ delete probe.reason
+ } catch {
+ // Why: `inspectRuntimeTerminalProcess` already maps every failure it can classify onto a
+ // reason; an unclassified throw is still a failure to observe, so it stays unverifiable.
+ probe.timedOut = false
+ probe.verdict = 'unverifiable'
+ probe.reason = 'probe_failed'
+ }
+ })
+ )
+
+ let deadline: ReturnType | undefined
+ try {
+ await Promise.race([
+ settle,
+ new Promise((resolve) => {
+ deadline = setTimeout(resolve, options.timeoutMs)
+ })
+ ])
+ } finally {
+ clearTimeout(deadline)
+ }
+ return probes
+}
diff --git a/src/renderer/src/components/terminal/running-terminal-close-guard.test.ts b/src/renderer/src/components/terminal/running-terminal-close-guard.test.ts
index d63c69df7c9..165119f8b31 100644
--- a/src/renderer/src/components/terminal/running-terminal-close-guard.test.ts
+++ b/src/renderer/src/components/terminal/running-terminal-close-guard.test.ts
@@ -46,10 +46,12 @@ function visibleRequest() {
return useRunningTerminalCloseConfirmStore.getState().runningTerminalCloseConfirm
}
+// Drains pending microtasks. The probe resolves through several await points (per-pty inspect,
+// the batch join, the deadline race), so this flushes generously rather than counting ticks.
async function settleProbe(): Promise {
- await Promise.resolve()
- await Promise.resolve()
- await Promise.resolve()
+ for (let tick = 0; tick < 12; tick += 1) {
+ await Promise.resolve()
+ }
}
describe('shouldConfirmRunningTerminalClose', () => {
@@ -329,6 +331,7 @@ describe('guardRunningTerminalClose', () => {
vi.advanceTimersByTime(RUNNING_CLOSE_PROBE_TIMEOUT_MS)
vi.useRealTimers()
+ await settleProbe()
expect(onClose).not.toHaveBeenCalled()
expect(visibleRequest()).toMatchObject({ terminalTabId: 'tab-1', tabLabel: 'npm run dev' })
@@ -351,6 +354,7 @@ describe('guardRunningTerminalClose', () => {
guard()
vi.advanceTimersByTime(RUNNING_CLOSE_PROBE_TIMEOUT_MS)
vi.useRealTimers()
+ await settleProbe()
expect(visibleRequest()?.copyKind).toBe('agent')
})
@@ -368,6 +372,7 @@ describe('guardRunningTerminalClose', () => {
guard(onClose)
vi.advanceTimersByTime(RUNNING_CLOSE_PROBE_TIMEOUT_MS)
vi.useRealTimers()
+ await settleProbe()
requestSpy.mockRestore()
expect(onClose).toHaveBeenCalledTimes(1)
@@ -385,6 +390,7 @@ describe('guardRunningTerminalClose', () => {
vi.advanceTimersByTime(RUNNING_CLOSE_PROBE_TIMEOUT_MS)
vi.useRealTimers()
await settleProbe()
+ await settleProbe()
expect(onClose).not.toHaveBeenCalled()
useRunningTerminalCloseConfirmStore.getState().confirmRunningTerminalClose()
diff --git a/src/renderer/src/components/terminal/running-terminal-close-guard.ts b/src/renderer/src/components/terminal/running-terminal-close-guard.ts
index 881cd262353..6bb9ff5a582 100644
--- a/src/renderer/src/components/terminal/running-terminal-close-guard.ts
+++ b/src/renderer/src/components/terminal/running-terminal-close-guard.ts
@@ -1,10 +1,9 @@
import { useAppStore } from '@/store'
-import { inspectRuntimeTerminalProcess } from '@/runtime/runtime-terminal-inspection'
import { useRunningTerminalCloseConfirmStore } from '@/store/running-terminal-close-confirm'
import type { TerminalTabCloseReason } from '@/store/slices/terminal-tab-retirement'
import type { AppState } from '@/store/types'
import { resolveBusyPtyCloseCopyKind } from './terminal-close-copy-kind'
-import { isClientOnlyUnverifiableInspection } from '../../../../shared/terminal-process-inspection'
+import { probePtyRunningWork } from './pty-running-work-probe'
export type RunningTerminalCloseGuardOptions = {
force?: boolean
@@ -44,7 +43,7 @@ export function shouldConfirmRunningTerminalClose(
* the store's own teardown collector unions both for exactly that reason — reading only
* the map would let a close slip through the window with no prompt. A stale id costs
* nothing: its probe fails and the guard falls open. */
-function collectTabPtyIds(
+export function collectTabPtyIds(
state: Pick,
terminalTabId: string
): string[] {
@@ -112,44 +111,33 @@ export function guardRunningTerminalClose(params: {
decided = true
}
- const probeTimeout = setTimeout(() => {
- try {
+ void probePtyRunningWork(settings, ptyIds, { timeoutMs: RUNNING_CLOSE_PROBE_TIMEOUT_MS })
+ .then((probes) => {
+ if (decided) {
+ return
+ }
// Why: a probe that has not answered yet is unknown, not idle. Ask, treating every pty
// as a candidate, so a degraded relay costs a click instead of a killed remote command.
- confirmClose(ptyIds)
- } catch {
- closeNow()
- }
- }, RUNNING_CLOSE_PROBE_TIMEOUT_MS)
-
- void Promise.allSettled(ptyIds.map((ptyId) => inspectRuntimeTerminalProcess(settings, ptyId)))
- .then((results) => {
- clearTimeout(probeTimeout)
- if (decided) {
+ if (probes.some((probe) => probe.timedOut)) {
+ confirmClose(ptyIds)
return
}
// Why: fail open on an *answered* probe, matching the Cmd+W pane path — a rejection
// (wedged relay, legacy provider) or a stale remote handle is not evidence of a live
// child, and a close button that silently does nothing is worse than closing a busy tab.
- const busyPtyIds = ptyIds.filter((_, index) => {
- const result = results[index]
- return (
- result?.status === 'fulfilled' &&
- !isClientOnlyUnverifiableInspection(result.value) &&
- result.value.hasChildProcesses
- )
- })
+ const busyPtyIds = probes
+ .filter((probe) => probe.verdict === 'live')
+ .map((probe) => probe.ptyId)
if (busyPtyIds.length === 0) {
closeNow()
return
}
confirmClose(busyPtyIds)
})
- // Why: allSettled never rejects, so this only fires when the decision above throws (a
+ // Why: the probe never rejects, so this only fires when the decision above throws (a
// copy-kind lookup, a store subscriber). Without it the tab would silently never close
// and the user would get no feedback at all; the pane path it replaced had this catch.
.catch(() => {
- clearTimeout(probeTimeout)
closeNow()
})
}
diff --git a/src/renderer/src/components/terminal/terminal-tab-close-running-confirm.test.ts b/src/renderer/src/components/terminal/terminal-tab-close-running-confirm.test.ts
index 066c5795ef9..9dfe20ef501 100644
--- a/src/renderer/src/components/terminal/terminal-tab-close-running-confirm.test.ts
+++ b/src/renderer/src/components/terminal/terminal-tab-close-running-confirm.test.ts
@@ -87,10 +87,12 @@ function visibleRequest() {
return useRunningTerminalCloseConfirmStore.getState().runningTerminalCloseConfirm
}
+// Drains pending microtasks. The probe resolves through several await points (per-pty inspect,
+// the batch join, the deadline race), so this flushes generously rather than counting ticks.
async function settleProbe(): Promise {
- await Promise.resolve()
- await Promise.resolve()
- await Promise.resolve()
+ for (let tick = 0; tick < 12; tick += 1) {
+ await Promise.resolve()
+ }
}
describe('closeTerminalTab running-process confirmation', () => {
diff --git a/src/renderer/src/components/terminal/window-close-running-work.test.ts b/src/renderer/src/components/terminal/window-close-running-work.test.ts
new file mode 100644
index 00000000000..77dc4ad745b
--- /dev/null
+++ b/src/renderer/src/components/terminal/window-close-running-work.test.ts
@@ -0,0 +1,231 @@
+import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
+
+const { getStateMock, inspectRuntimeTerminalProcessMock } = vi.hoisted(() => ({
+ getStateMock: vi.fn(),
+ inspectRuntimeTerminalProcessMock: vi.fn()
+}))
+
+vi.mock('@/store', () => ({
+ useAppStore: { getState: getStateMock }
+}))
+
+vi.mock('@/runtime/runtime-terminal-inspection', () => ({
+ inspectRuntimeTerminalProcess: inspectRuntimeTerminalProcessMock
+}))
+
+import {
+ assessWindowCloseRunningWork,
+ WINDOW_CLOSE_PROBE_TIMEOUT_MS
+} from './window-close-running-work'
+
+const LOCAL_PTY = 'pty-local'
+const SSH_PTY = 'ssh:openclaw@@pty-7'
+const RUNTIME_PTY = 'remote:env-1@@handle-1'
+/** A runtime pty minted without an owner id. Still someone else's machine. */
+const OWNERLESS_RUNTIME_PTY = 'remote:handle-2'
+
+const BUSY = {
+ foregroundProcess: 'pnpm build',
+ hasChildProcesses: true,
+ foregroundProcessEvidence: {}
+}
+const IDLE = { foregroundProcess: 'bash', hasChildProcesses: false, foregroundProcessEvidence: {} }
+const UNVERIFIABLE = {
+ foregroundProcess: null,
+ hasChildProcesses: false,
+ verdict: 'unverifiable',
+ reason: 'transport_loss'
+}
+
+/** One worktree, one tab, owning `ptyIds`. */
+function setState(ptyIds: string[]): void {
+ getStateMock.mockReturnValue({
+ settings: { activeRuntimeEnvironmentId: null },
+ tabsByWorktree: { 'worktree-1': [{ id: 'tab-1' }] },
+ ptyIdsByTabId: { 'tab-1': ptyIds },
+ terminalLayoutsByTabId: {}
+ })
+}
+
+/** Answers each pty id from `byPtyId`; anything unlisted never settles. */
+function answerWith(byPtyId: Record): void {
+ inspectRuntimeTerminalProcessMock.mockImplementation((_settings: unknown, ptyId: string) =>
+ ptyId in byPtyId ? Promise.resolve(byPtyId[ptyId]) : new Promise(() => {})
+ )
+}
+
+beforeEach(() => {
+ vi.clearAllMocks()
+})
+
+afterEach(() => {
+ vi.useRealTimers()
+})
+
+describe('assessWindowCloseRunningWork', () => {
+ it('warns about a live process on an SSH host (F15: remote work was filtered out entirely)', async () => {
+ setState([SSH_PTY])
+ answerWith({ [SSH_PTY]: BUSY })
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: false })).resolves.toEqual({
+ kind: 'running'
+ })
+ })
+
+ it('warns on quit about a live process on an SSH host', async () => {
+ setState([SSH_PTY])
+ answerWith({ [SSH_PTY]: BUSY })
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: true })).resolves.toEqual({
+ kind: 'running'
+ })
+ })
+
+ it('warns on quit about a live process on a paired runtime host', async () => {
+ setState([RUNTIME_PTY])
+ answerWith({ [RUNTIME_PTY]: BUSY })
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: true })).resolves.toEqual({
+ kind: 'running'
+ })
+ })
+
+ it('counts an owner-less remote pty as remote work', async () => {
+ setState([OWNERLESS_RUNTIME_PTY])
+ answerWith({ [OWNERLESS_RUNTIME_PTY]: BUSY })
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: true })).resolves.toEqual({
+ kind: 'running'
+ })
+ })
+
+ // The crux of docs/reference/ssh-execution-boundary.md: an unreachable host is `unverifiable`,
+ // and quitting on `unverifiable` as though it were `exited` is what orphans live remote work.
+ it('warns rather than quitting silently when a remote host answers unverifiable', async () => {
+ setState([SSH_PTY])
+ answerWith({ [SSH_PTY]: UNVERIFIABLE })
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: true })).resolves.toEqual({
+ kind: 'unverifiable'
+ })
+ })
+
+ it('warns rather than quitting silently when a remote probe throws', async () => {
+ setState([SSH_PTY])
+ inspectRuntimeTerminalProcessMock.mockRejectedValue(new Error('relay wedged'))
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: true })).resolves.toEqual({
+ kind: 'unverifiable'
+ })
+ })
+
+ it('stops waiting at the budget and warns, so an unreachable host cannot hang the quit', async () => {
+ setState([SSH_PTY])
+ answerWith({})
+ vi.useFakeTimers()
+
+ const pending = assessWindowCloseRunningWork({ isQuitting: true })
+ await vi.advanceTimersByTimeAsync(WINDOW_CLOSE_PROBE_TIMEOUT_MS)
+
+ await expect(pending).resolves.toEqual({ kind: 'unverifiable' })
+ })
+
+ it('does not resolve before the budget expires', async () => {
+ setState([SSH_PTY])
+ answerWith({})
+ vi.useFakeTimers()
+ const settled = vi.fn()
+
+ void assessWindowCloseRunningWork({ isQuitting: true }).then(settled)
+ await vi.advanceTimersByTimeAsync(WINDOW_CLOSE_PROBE_TIMEOUT_MS - 1)
+
+ expect(settled).not.toHaveBeenCalled()
+ })
+
+ it('does not warn when the owning remote host reports an idle shell', async () => {
+ setState([SSH_PTY])
+ answerWith({ [SSH_PTY]: IDLE })
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: true })).resolves.toEqual({
+ kind: 'none'
+ })
+ })
+
+ it('reports a live process even when a sibling remote pane is only unverifiable', async () => {
+ setState([SSH_PTY, RUNTIME_PTY])
+ answerWith({ [SSH_PTY]: UNVERIFIABLE, [RUNTIME_PTY]: BUSY })
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: true })).resolves.toEqual({
+ kind: 'running'
+ })
+ })
+
+ it('still warns about a live local process when closing the window', async () => {
+ setState([LOCAL_PTY])
+ answerWith({ [LOCAL_PTY]: BUSY })
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: false })).resolves.toEqual({
+ kind: 'running'
+ })
+ })
+
+ // A local probe has no transport to lose, so its failure means the pty is gone — unlike a
+ // remote host going quiet, it is not a reason to hold up the close.
+ it('does not warn when only a local probe is unverifiable', async () => {
+ setState([LOCAL_PTY])
+ answerWith({ [LOCAL_PTY]: UNVERIFIABLE })
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: false })).resolves.toEqual({
+ kind: 'none'
+ })
+ })
+
+ // #524 decided quitting is an unambiguous instruction to end this machine's processes. It is
+ // not an instruction to end execution on someone else's, which is why remote still warns above.
+ it('leaves local-only quit unprompted, and never probes for it', async () => {
+ setState([LOCAL_PTY])
+ answerWith({ [LOCAL_PTY]: BUSY })
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: true })).resolves.toEqual({
+ kind: 'none'
+ })
+ expect(inspectRuntimeTerminalProcessMock).not.toHaveBeenCalled()
+ })
+
+ it('probes a pane the layout has bound before the liveness map caught up', async () => {
+ getStateMock.mockReturnValue({
+ settings: { activeRuntimeEnvironmentId: null },
+ tabsByWorktree: { 'worktree-1': [{ id: 'tab-1' }] },
+ ptyIdsByTabId: {},
+ terminalLayoutsByTabId: { 'tab-1': { ptyIdsByLeafId: { leaf: SSH_PTY } } }
+ })
+ answerWith({ [SSH_PTY]: BUSY })
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: true })).resolves.toEqual({
+ kind: 'running'
+ })
+ })
+
+ it('probes each pty once when the map and the layout name the same one', async () => {
+ getStateMock.mockReturnValue({
+ settings: { activeRuntimeEnvironmentId: null },
+ tabsByWorktree: { 'worktree-1': [{ id: 'tab-1' }] },
+ ptyIdsByTabId: { 'tab-1': [SSH_PTY] },
+ terminalLayoutsByTabId: { 'tab-1': { ptyIdsByLeafId: { leaf: SSH_PTY } } }
+ })
+ answerWith({ [SSH_PTY]: IDLE })
+
+ await assessWindowCloseRunningWork({ isQuitting: true })
+
+ expect(inspectRuntimeTerminalProcessMock).toHaveBeenCalledTimes(1)
+ })
+
+ it('closes without probing when no workspace owns a pty', async () => {
+ setState([])
+
+ await expect(assessWindowCloseRunningWork({ isQuitting: false })).resolves.toEqual({
+ kind: 'none'
+ })
+ expect(inspectRuntimeTerminalProcessMock).not.toHaveBeenCalled()
+ })
+})
diff --git a/src/renderer/src/components/terminal/window-close-running-work.ts b/src/renderer/src/components/terminal/window-close-running-work.ts
new file mode 100644
index 00000000000..427ee72dae6
--- /dev/null
+++ b/src/renderer/src/components/terminal/window-close-running-work.ts
@@ -0,0 +1,70 @@
+import { useAppStore } from '@/store'
+import { isRemoteExecutionHostPtyId } from '../../../../shared/remote-execution-host-pty-id'
+import { collectTabPtyIds } from './running-terminal-close-guard'
+import { probePtyRunningWork } from './pty-running-work-probe'
+
+/**
+ * Upper bound on how long closing the window or quitting may wait on the probes.
+ *
+ * Shorter than the tab-close guard's 4s because quit is time-sensitive in a way one tab close is
+ * not: the user has already asked to leave, and a quit that stalls on an unreachable host is its
+ * own bug. A healthy local inspect answers in single-digit milliseconds and a healthy remote one
+ * is a single RPC round-trip on an already-open mux channel, so this leaves roughly 3x headroom
+ * over a slow-but-live transcontinental host while capping the worst case — a host that is simply
+ * gone — at ~1.5s instead of the 15s RPC timeout the probe would otherwise inherit.
+ *
+ * Expiry raises the prompt rather than quitting silently: an unanswered probe is `unverifiable`,
+ * and `unverifiable` is never evidence that remote work has stopped.
+ */
+export const WINDOW_CLOSE_PROBE_TIMEOUT_MS = 1_500
+
+/** Which warning the close should raise, if any. */
+export type WindowCloseRunningWork =
+ /** Every pty that mattered answered, and none had children. */
+ | { kind: 'none' }
+ /** An owning host reported a live child process. */
+ | { kind: 'running' }
+ /** A remote execution host could not be observed, so its work may still be live. */
+ | { kind: 'unverifiable' }
+
+/**
+ * Decides whether a window close or quit should stop and ask.
+ *
+ * Two deliberate asymmetries:
+ *
+ * - **Quit only considers remote ptys.** Quitting is an unambiguous instruction to end this
+ * machine's processes (#524), but it is not an instruction to end execution on someone else's:
+ * the client detaches while the relay keeps running, and a target with a bounded grace period
+ * then SIGKILLs that work once the countdown expires.
+ * - **Only a remote `unverifiable` warns.** A local probe has no transport to lose, so its failure
+ * means the pty is gone. A remote one that cannot be reached is the case
+ * `docs/reference/ssh-execution-boundary.md` exists to protect: loss of contact is not evidence
+ * of `exited`, so it must fail toward asking rather than toward a silent quit.
+ */
+export async function assessWindowCloseRunningWork(params: {
+ isQuitting: boolean
+}): Promise {
+ const state = useAppStore.getState()
+ const ptyIds = new Set(
+ Object.values(state.tabsByWorktree)
+ .flatMap((worktreeTabs) => worktreeTabs ?? [])
+ .flatMap((tab) => collectTabPtyIds(state, tab.id))
+ )
+ const candidatePtyIds = params.isQuitting
+ ? [...ptyIds].filter(isRemoteExecutionHostPtyId)
+ : [...ptyIds]
+ if (candidatePtyIds.length === 0) {
+ return { kind: 'none' }
+ }
+
+ const probes = await probePtyRunningWork(state.settings, candidatePtyIds, {
+ timeoutMs: WINDOW_CLOSE_PROBE_TIMEOUT_MS
+ })
+ if (probes.some((probe) => probe.verdict === 'live')) {
+ return { kind: 'running' }
+ }
+ if (probes.some((probe) => probe.remote && probe.verdict === 'unverifiable')) {
+ return { kind: 'unverifiable' }
+ }
+ return { kind: 'none' }
+}
diff --git a/src/renderer/src/components/use-terminal-editor-close-foundation.ts b/src/renderer/src/components/use-terminal-editor-close-foundation.ts
index 74849e3f5a2..2aceced683d 100644
--- a/src/renderer/src/components/use-terminal-editor-close-foundation.ts
+++ b/src/renderer/src/components/use-terminal-editor-close-foundation.ts
@@ -1,8 +1,9 @@
import { useCallback, useRef, useState } from 'react'
-import { useAppStore } from '../store'
-import { getConnectionId } from '../lib/connection-context'
-import { isRemoteRuntimePtyId } from '@/runtime/runtime-terminal-inspection'
import { CLOSE_DIALOG_DEBOUNCE_MS } from './terminal-workspace-model'
+import {
+ assessWindowCloseRunningWork,
+ type WindowCloseRunningWork
+} from './terminal/window-close-running-work'
import type { TerminalWorkspaceProjectionController } from './use-terminal-workspace-projection'
import { runWithWindowCloseCheckpointScope } from './window-close-request-coordinator'
import { showShutdownCheckpointFailureToast } from '@/lib/shutdown-checkpoint-failure-toast'
@@ -27,6 +28,11 @@ export function useTerminalEditorCloseFoundation(
closeDialogDebounceTimersRef.current.add(timer)
}, [])
const [windowCloseDialogOpen, setWindowCloseDialogOpen] = useState(false)
+ // Why: "running" and "could not reach the host" are different claims, and telling the user
+ // processes are running when the truth is that a host went quiet is the fabricated certainty
+ // docs/reference/ssh-execution-boundary.md forbids.
+ const [windowCloseDialogKind, setWindowCloseDialogKind] =
+ useState>('running')
const windowCloseAfterDirtyRef = useRef<{ isQuitting: boolean } | null>(null)
const confirmNativeWindowClose = useCallback(() => {
@@ -46,33 +52,21 @@ export function useTerminalEditorCloseFoundation(
const proceedToNativeWindowClose = useCallback(
(isQuitting: boolean) => {
- if (!isQuitting) {
- const state = useAppStore.getState()
- const localPtyIds = Object.entries(state.tabsByWorktree).flatMap(
- ([worktreeId, worktreeTabs]) => {
- const connectionId = getConnectionId(worktreeId)
- if (connectionId !== null) {
- return []
- }
- return worktreeTabs
- .flatMap((tab) => state.ptyIdsByTabId[tab.id] ?? [])
- .filter((ptyId) => !isRemoteRuntimePtyId(ptyId))
+ void assessWindowCloseRunningWork({ isQuitting })
+ .then((runningWork) => {
+ if (runningWork.kind === 'none') {
+ confirmNativeWindowClose()
+ return
}
- )
- if (localPtyIds.length > 0) {
- void Promise.all(localPtyIds.map((id) => window.api.pty.hasChildProcesses(id))).then(
- (results) => {
- if (results.some(Boolean)) {
- setWindowCloseDialogOpen(true)
- } else {
- confirmNativeWindowClose()
- }
- }
- )
- return
- }
- }
- confirmNativeWindowClose()
+ setWindowCloseDialogKind(runningWork.kind)
+ setWindowCloseDialogOpen(true)
+ })
+ // Why: the assessment must never be able to trap the window. A thrown store read is
+ // not evidence either way, and a close that silently does nothing is unrecoverable
+ // without SIGKILL, so fall through to the close the user actually asked for.
+ .catch(() => {
+ confirmNativeWindowClose()
+ })
},
[confirmNativeWindowClose]
)
@@ -88,6 +82,7 @@ export function useTerminalEditorCloseFoundation(
releaseCloseDialogGuardAfterDebounce,
windowCloseDialogOpen,
setWindowCloseDialogOpen,
+ windowCloseDialogKind,
windowCloseAfterDirtyRef,
confirmNativeWindowClose,
proceedToNativeWindowClose
diff --git a/src/renderer/src/components/use-terminal-editor-close-foundation.window-close.test.tsx b/src/renderer/src/components/use-terminal-editor-close-foundation.window-close.test.tsx
new file mode 100644
index 00000000000..d7f3f69d925
--- /dev/null
+++ b/src/renderer/src/components/use-terminal-editor-close-foundation.window-close.test.tsx
@@ -0,0 +1,103 @@
+// @vitest-environment happy-dom
+
+/**
+ * Wiring for the window-close/quit running-work warning. The policy in
+ * `terminal/window-close-running-work.ts` is inert unless `proceedToNativeWindowClose` actually
+ * consults it, so pin that it does — and that a warning stops the native close rather than
+ * confirming it.
+ */
+import { act, cleanup, renderHook } from '@testing-library/react'
+import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
+
+const { assessWindowCloseRunningWorkMock, confirmWindowCloseMock } = vi.hoisted(() => ({
+ assessWindowCloseRunningWorkMock: vi.fn(),
+ confirmWindowCloseMock: vi.fn()
+}))
+
+vi.mock('./terminal/window-close-running-work', () => ({
+ assessWindowCloseRunningWork: assessWindowCloseRunningWorkMock
+}))
+vi.mock('./window-close-request-coordinator', () => ({
+ runWithWindowCloseCheckpointScope: (fn: () => unknown) => fn()
+}))
+vi.mock('@/lib/shutdown-checkpoint-failure-toast', () => ({
+ showShutdownCheckpointFailureToast: vi.fn()
+}))
+
+const { useTerminalEditorCloseFoundation } = await import('./use-terminal-editor-close-foundation')
+
+const controller = { openFiles: [] } as unknown as Parameters<
+ typeof useTerminalEditorCloseFoundation
+>[0]
+
+function mountFoundation() {
+ return renderHook(() => useTerminalEditorCloseFoundation(controller))
+}
+
+beforeEach(() => {
+ vi.clearAllMocks()
+ Object.assign(globalThis, {
+ window: Object.assign(globalThis.window, {
+ api: { ui: { confirmWindowClose: confirmWindowCloseMock } }
+ })
+ })
+})
+
+afterEach(() => {
+ cleanup()
+})
+
+describe('proceedToNativeWindowClose', () => {
+ it('asks the running-work policy about the quit rather than assuming it is safe', async () => {
+ assessWindowCloseRunningWorkMock.mockResolvedValue({ kind: 'none' })
+ const { result } = mountFoundation()
+
+ await act(async () => {
+ result.current.proceedToNativeWindowClose(true)
+ })
+
+ expect(assessWindowCloseRunningWorkMock).toHaveBeenCalledWith({ isQuitting: true })
+ expect(confirmWindowCloseMock).toHaveBeenCalledTimes(1)
+ expect(result.current.windowCloseDialogOpen).toBe(false)
+ })
+
+ it('raises the dialog and does not close when a host reports live work', async () => {
+ assessWindowCloseRunningWorkMock.mockResolvedValue({ kind: 'running' })
+ const { result } = mountFoundation()
+
+ await act(async () => {
+ result.current.proceedToNativeWindowClose(true)
+ })
+
+ expect(result.current.windowCloseDialogOpen).toBe(true)
+ expect(result.current.windowCloseDialogKind).toBe('running')
+ expect(confirmWindowCloseMock).not.toHaveBeenCalled()
+ })
+
+ it('raises the unverifiable copy when a remote host could not be reached', async () => {
+ assessWindowCloseRunningWorkMock.mockResolvedValue({ kind: 'unverifiable' })
+ const { result } = mountFoundation()
+
+ await act(async () => {
+ result.current.proceedToNativeWindowClose(true)
+ })
+
+ expect(result.current.windowCloseDialogOpen).toBe(true)
+ expect(result.current.windowCloseDialogKind).toBe('unverifiable')
+ expect(confirmWindowCloseMock).not.toHaveBeenCalled()
+ })
+
+ // Why: a thrown assessment is not evidence either way, and a close that silently does nothing
+ // leaves SIGKILL as the user's only exit.
+ it('falls through to the close when the assessment throws', async () => {
+ assessWindowCloseRunningWorkMock.mockRejectedValue(new Error('store blew up'))
+ const { result } = mountFoundation()
+
+ await act(async () => {
+ result.current.proceedToNativeWindowClose(false)
+ })
+
+ expect(confirmWindowCloseMock).toHaveBeenCalledTimes(1)
+ expect(result.current.windowCloseDialogOpen).toBe(false)
+ })
+})
diff --git a/src/renderer/src/i18n/locales/en.json b/src/renderer/src/i18n/locales/en.json
index 87fcb50a2a1..f2da58bffb3 100644
--- a/src/renderer/src/i18n/locales/en.json
+++ b/src/renderer/src/i18n/locales/en.json
@@ -2251,7 +2251,8 @@
"Terminal": {
"73768427cf": "Close",
"f82e9f02df": "Cancel",
- "7958465754": "There are local terminals with running processes. Close the window anyway?",
+ "7958465754": "There are terminals with running processes. Close the window anyway?",
+ "b7c1f0a934": "A remote host could not be reached, so Orca cannot tell whether work is still running there. Close the window anyway?",
"2fa9c69ff3": "Close Window?",
"cd51e28d8b": "Save",
"0037b21794": "Don't Save",
diff --git a/src/renderer/src/i18n/locales/es.json b/src/renderer/src/i18n/locales/es.json
index 9999634daee..98b07917880 100644
--- a/src/renderer/src/i18n/locales/es.json
+++ b/src/renderer/src/i18n/locales/es.json
@@ -1924,7 +1924,7 @@
"Terminal": {
"73768427cf": "Cerrar",
"f82e9f02df": "Cancelar",
- "7958465754": "Hay terminales locales con procesos en ejecución. ¿Cerrar la ventana de todos modos?",
+ "7958465754": "Hay terminales con procesos en ejecución. ¿Cerrar la ventana de todos modos?",
"2fa9c69ff3": "¿Cerrar ventana?",
"cd51e28d8b": "Guardar",
"0037b21794": "No guardar",
diff --git a/src/renderer/src/i18n/locales/fr.json b/src/renderer/src/i18n/locales/fr.json
index 8eff250a820..5bf316d36f7 100644
--- a/src/renderer/src/i18n/locales/fr.json
+++ b/src/renderer/src/i18n/locales/fr.json
@@ -2087,7 +2087,7 @@
"Terminal": {
"73768427cf": "Fermer",
"f82e9f02df": "Annuler",
- "7958465754": "Des terminaux locaux exécutent des processus. Fermer quand même la fenêtre ?",
+ "7958465754": "Des terminaux exécutent des processus. Fermer quand même la fenêtre ?",
"2fa9c69ff3": "Fermer la fenêtre ?",
"cd51e28d8b": "Enregistrer",
"0037b21794": "Ne pas enregistrer",
diff --git a/src/renderer/src/i18n/locales/ja.json b/src/renderer/src/i18n/locales/ja.json
index b3c30da9446..4dc81b9201d 100644
--- a/src/renderer/src/i18n/locales/ja.json
+++ b/src/renderer/src/i18n/locales/ja.json
@@ -1924,7 +1924,7 @@
"Terminal": {
"73768427cf": "閉じる",
"f82e9f02df": "キャンセル",
- "7958465754": "プロセスが実行中のローカルターミナルがあります。このままウィンドウを閉じますか?",
+ "7958465754": "プロセスが実行中のターミナルがあります。このままウィンドウを閉じますか?",
"2fa9c69ff3": "ウィンドウを閉じますか?",
"cd51e28d8b": "保存",
"0037b21794": "保存しないでください",
diff --git a/src/renderer/src/i18n/locales/ko.json b/src/renderer/src/i18n/locales/ko.json
index ac75cca2209..d9d822b8fc3 100644
--- a/src/renderer/src/i18n/locales/ko.json
+++ b/src/renderer/src/i18n/locales/ko.json
@@ -1929,7 +1929,7 @@
"Terminal": {
"73768427cf": "닫기",
"f82e9f02df": "취소",
- "7958465754": "실행 중인 프로세스가 있는 로컬 terminals이 있습니다. 그래도 창을 닫으시겠습니까?",
+ "7958465754": "실행 중인 프로세스가 있는 terminals이 있습니다. 그래도 창을 닫으시겠습니까?",
"2fa9c69ff3": "창을 닫으시겠습니까?",
"cd51e28d8b": "저장",
"0037b21794": "저장하지 않음",
diff --git a/src/renderer/src/i18n/locales/zh.json b/src/renderer/src/i18n/locales/zh.json
index 7a3b47c8f2a..46dc592dd2c 100644
--- a/src/renderer/src/i18n/locales/zh.json
+++ b/src/renderer/src/i18n/locales/zh.json
@@ -1927,7 +1927,7 @@
"Terminal": {
"73768427cf": "关闭",
"f82e9f02df": "取消",
- "7958465754": "有正在运行的进程的本地终端。还是关窗吧?",
+ "7958465754": "有正在运行的进程的终端。还是关窗吧?",
"2fa9c69ff3": "关闭窗口?",
"cd51e28d8b": "保存",
"0037b21794": "不保存",
diff --git a/src/shared/remote-execution-host-pty-id.ts b/src/shared/remote-execution-host-pty-id.ts
new file mode 100644
index 00000000000..c76ada39e03
--- /dev/null
+++ b/src/shared/remote-execution-host-pty-id.ts
@@ -0,0 +1,14 @@
+import { parseRemoteRuntimePtyId } from './remote-runtime-pty-id'
+import { parseAppSshPtyId } from './ssh-pty-id'
+
+/**
+ * Whether the process behind this pty runs on an execution host other than this machine —
+ * a paired runtime environment or an app SSH target.
+ *
+ * Deliberately broader than the inspection module's private remote check, which only counts a
+ * `remote:` id that carries an owner environment id. An owner-less `remote:` still runs
+ * somewhere else, and treating it as local is how remote work becomes invisible to a guard.
+ */
+export function isRemoteExecutionHostPtyId(ptyId: string): boolean {
+ return parseRemoteRuntimePtyId(ptyId) !== null || parseAppSshPtyId(ptyId) !== null
+}
From 11e459e9330b0712988b9c4afb063c71ee0b1507 Mon Sep 17 00:00:00 2001
From: Neil <4138956+nwparker@users.noreply.github.com>
Date: Fri, 4 Sep 2026 00:53:24 -0700
Subject: [PATCH 26/49] fix(crash-reporting): bound replay-guard wedge bursts
in the ring without losing their spans (#18441)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
`terminal_replay_guard_wedged_release` was not in
COALESCED_RENDERER_BREADCRUMB_NAMES, and its per-pane hashes give every entry a
unique ring identity. One mount/reveal/wake transition expires every in-flight
replay write at once, so a burst arrives as N distinct entries against a 30-slot
FIFO ring.
Measured, from the 09-02 corpus (121 `renderer.breadcrumb` spans across 9 of 55
diagnostic bundles):
- bundle 26461769: 26 events in 0.96s
- murlock1000: 62 events over 85s
- 8907a508 mixes two call sites in one window (2 crumbs carry `tabIdHash`, 2 do not)
Not measured: no captured report's ring actually lost slots to this crumb. All
57 reports have zero wedge crumbs in `Recent activity:`, and in all 9 bundles the
burst predates the report's ring window — for 26461769 the burst ran
13:36:03.784Z-13:36:04.742Z while the ring-owning main process started at
13:43:44.380Z, 7m40s later. So this bounds a demonstrated hazard, not an observed
loss. An earlier draft of this commit asserted "26 of 30 slots / 87% of the
pre-crash trail" as a measurement; that was a model, and it is removed.
The burst evidence lives entirely in the durable span stream, and suppressed
repeats normally emit no span (see the 1000-emissions/1-span case in
crash-reporting-renderer-breadcrumbs.test.ts), so coalescing alone would have cut
that 121-event corpus to 13 with the multiplicity recorded nowhere. Instead:
- the ring coalesces: one slot per call site, plus `suppressedSinceLast`
- every wedge event still emits its own `renderer.breadcrumb` span, via
PER_EVENT_TRACED_COALESCED_BREADCRUMB_NAMES. Span volume is unchanged at 121,
and the span deliberately carries no count so a span-stream total cannot
double-count what the ring already claims
- the coalesce key is `ptyId`/`tabIdHash` *presence*, not name alone: those
fields are absent on the restore call site (restoreScrollbackBuffers) and
present on reattach, so name-only keying would collapse 8907a508's two call
sites into whichever crumb landed last. Bounded at 4 slots per storm, matching
the webgl `kind` and duplicate-tab `resolvedToActiveWorktree` precedents in the
same file. Replaying the corpus timestamps: 121 events -> 14 ring writes.
This is a diagnostics fix, not a crash fix. It does not stop panes wedging, and
it does not explain the "can't type" reports in this round.
---
.../crash-reporting-renderer-breadcrumbs.ts | 24 +++-
...reporting-replay-guard-wedge-burst.test.ts | 128 ++++++++++++++++++
2 files changed, 149 insertions(+), 3 deletions(-)
create mode 100644 src/main/ipc/crash-reporting-replay-guard-wedge-burst.test.ts
diff --git a/src/main/ipc/crash-reporting-renderer-breadcrumbs.ts b/src/main/ipc/crash-reporting-renderer-breadcrumbs.ts
index 97e0a8f9d65..8d126f557b5 100644
--- a/src/main/ipc/crash-reporting-renderer-breadcrumbs.ts
+++ b/src/main/ipc/crash-reporting-renderer-breadcrumbs.ts
@@ -49,6 +49,7 @@ function recordRendererBreadcrumbTrace(
const DUPLICATE_TAB_OWNER_BREADCRUMB = 'terminal_tab_id_owned_by_multiple_worktrees'
const PARK_VERDICT_CHURN_BREADCRUMB = 'terminal_park_verdict_churn'
const REACT_COMMIT_CASCADE_BREADCRUMB = 'react_commit_cascade'
+const REPLAY_GUARD_WEDGED_BREADCRUMB = 'terminal_replay_guard_wedged_release'
const COALESCED_RENDERER_BREADCRUMB_NAMES = new Set([
'renderer_error',
'renderer_unhandled_rejection',
@@ -56,6 +57,7 @@ const COALESCED_RENDERER_BREADCRUMB_NAMES = new Set([
DUPLICATE_TAB_OWNER_BREADCRUMB,
PARK_VERDICT_CHURN_BREADCRUMB,
REACT_COMMIT_CASCADE_BREADCRUMB,
+ REPLAY_GUARD_WEDGED_BREADCRUMB,
TERMINAL_WEBGL_DIAGNOSTIC_BREADCRUMB
])
const RENDERER_BREADCRUMB_COALESCE_MS = 30_000
@@ -69,6 +71,11 @@ const RENDERER_BREADCRUMB_COALESCE_MS = 30_000
// 30-entry ring to two such bursts. `suppressedSinceLast` keeps the pane count
// — the only signal these carry — in one slot.
const NAME_ONLY_COALESCED_BREADCRUMB_NAMES = new Set(['terminal_safe_fit_retry_exhausted'])
+// Why: the 30-slot ring is the scarce sink; the durable span stream is not. For
+// bounded-rate pane telemetry whose multiplicity is the whole signal, spans are the
+// only place a burst survives the restart that clears the ring, so coalesce the ring
+// but keep every event's span.
+const PER_EVENT_TRACED_COALESCED_BREADCRUMB_NAMES = new Set([REPLAY_GUARD_WEDGED_BREADCRUMB])
function rendererBreadcrumbCoalesceKey(
name: string,
@@ -77,6 +84,13 @@ function rendererBreadcrumbCoalesceKey(
if (NAME_ONLY_COALESCED_BREADCRUMB_NAMES.has(name)) {
return name
}
+ // Why presence and not value: `ptyId`/`tabIdHash` are absent on the restore call
+ // site (layout-serialization restoreScrollbackBuffers) and present on reattach, so
+ // their presence is the call-site identity a mixed burst would otherwise lose. Four
+ // slots per storm at most, regardless of pane count.
+ if (name === REPLAY_GUARD_WEDGED_BREADCRUMB) {
+ return `${name}:${data?.ptyId ? 'pty' : ''}:${data?.tabIdHash ? 'tab' : ''}`
+ }
// Why trigger and not name alone: `burst` means damping engaged a commit
// short of React #185, `window` means slow benign churn. Collapsing them
// would drop the near-crash signal into a slow-churn slot. Still bounded —
@@ -191,9 +205,13 @@ export function recordRendererBreadcrumbFromRenderer(
minIntervalMs: RENDERER_BREADCRUMB_COALESCE_MS,
...(origin ? { origin } : {})
})
- // Why: tracing every suppressed duplicate would preserve the same
- // serialization and disk churn that breadcrumb coalescing removes.
- if (coalesceResult) {
+ if (PER_EVENT_TRACED_COALESCED_BREADCRUMB_NAMES.has(args.name)) {
+ // Why the raw data: every event already gets its own span, so folding the ring's
+ // running count in here would double-count in any span-stream total.
+ recordRendererBreadcrumbTrace(args.name, data)
+ } else if (coalesceResult) {
+ // Why gated: tracing every suppressed duplicate would preserve the same
+ // serialization and disk churn that breadcrumb coalescing removes.
recordRendererBreadcrumbTrace(
args.name,
coalesceResult.suppressedSinceLast > 0
diff --git a/src/main/ipc/crash-reporting-replay-guard-wedge-burst.test.ts b/src/main/ipc/crash-reporting-replay-guard-wedge-burst.test.ts
new file mode 100644
index 00000000000..e3823f0b313
--- /dev/null
+++ b/src/main/ipc/crash-reporting-replay-guard-wedge-burst.test.ts
@@ -0,0 +1,128 @@
+import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
+
+import {
+ clearCrashBreadcrumbsForTest,
+ getCrashBreadcrumbSnapshot,
+ recordCrashBreadcrumb
+} from '../crash-reporting/crash-breadcrumb-store'
+import { recordRendererBreadcrumbFromRenderer } from './crash-reporting-renderer-breadcrumbs'
+
+type SpanOptions = { attributes: Record }
+const startSpanMock = vi.fn((_name: string, _options: SpanOptions) => ({ end: () => {} }))
+vi.mock('../observability/tracer', () => ({
+ startSpan: (name: string, options: SpanOptions) => startSpanMock(name, options)
+}))
+
+const WEDGE_BREADCRUMB = 'terminal_replay_guard_wedged_release'
+
+/** Reattach-path shape: identity-bearing (`tabIdHash`, optionally `ptyId`). */
+function emitReattachWedge(pane: number, withPtyId = false): void {
+ recordRendererBreadcrumbFromRenderer({
+ name: WEDGE_BREADCRUMB,
+ data: {
+ paneId: pane,
+ leafIdHash: `leaf${String(pane).padStart(5, '0')}`,
+ tabIdHash: `tab${String(pane).padStart(6, '0')}`,
+ worktreeIdHash: 'caa15fa9',
+ ...(withPtyId ? { ptyId: `…@@pty-${pane}` } : {})
+ }
+ })
+}
+
+/** Restore-path shape (restoreScrollbackBuffers): no tabIdHash, no ptyId. */
+function emitRestoreWedge(pane: number): void {
+ recordRendererBreadcrumbFromRenderer({
+ name: WEDGE_BREADCRUMB,
+ data: { paneId: pane, leafIdHash: `leaf${String(pane).padStart(5, '0')}` }
+ })
+}
+
+function wedgeCrumbs(): ReturnType {
+ return getCrashBreadcrumbSnapshot().filter((entry) => entry.name === WEDGE_BREADCRUMB)
+}
+
+function wedgeSpanCount(): number {
+ return startSpanMock.mock.calls.filter(
+ (call) => call[1].attributes['breadcrumb.name'] === WEDGE_BREADCRUMB
+ ).length
+}
+
+beforeEach(() => {
+ startSpanMock.mockClear()
+})
+
+afterEach(() => {
+ clearCrashBreadcrumbsForTest()
+})
+
+// One mount/reveal/wake transition expires every in-flight replay write at once, so
+// the burst reaches the 30-slot ring as N distinct entries. Field span streams measure
+// bursts of 26 in 0.96s and 62 over 85s. No captured report in the 09-02 corpus shows
+// a ring that actually drained — all nine bursts predate their report's ring window —
+// so this bounds a demonstrated hazard, not an observed loss, and must not cost the
+// durable span evidence that did carry those bursts.
+describe('replay-guard wedge burst against the fixed-size breadcrumb ring', () => {
+ it('costs one ring slot per call site and preserves the pre-crash trail', () => {
+ for (let index = 0; index < 10; index += 1) {
+ recordCrashBreadcrumb(`pre_crash_evidence_${index}`, { index })
+ }
+
+ for (let pane = 0; pane < 26; pane += 1) {
+ emitReattachWedge(pane)
+ }
+
+ const snapshot = getCrashBreadcrumbSnapshot()
+ expect(snapshot.filter((entry) => entry.name.startsWith('pre_crash_evidence_'))).toHaveLength(
+ 10
+ )
+ expect(wedgeCrumbs()).toHaveLength(1)
+ })
+
+ it('carries the burst multiplicity into the ring as suppressedSinceLast', () => {
+ for (let pane = 0; pane < 26; pane += 1) {
+ emitReattachWedge(pane)
+ }
+
+ // 26 emissions: one owns the slot, 25 fold into it.
+ expect(wedgeCrumbs()[0]?.data?.suppressedSinceLast).toBe(25)
+ })
+
+ // The 121-event field corpus lives entirely in the renderer.breadcrumb span stream,
+ // and the ring is cleared by the restart that usually precedes the crash report, so
+ // ring coalescing must not suppress the per-event spans.
+ it('still emits one durable span per wedge event', () => {
+ for (let pane = 0; pane < 26; pane += 1) {
+ emitReattachWedge(pane)
+ }
+
+ expect(wedgeSpanCount()).toBe(26)
+ // Why no count on the span: one span per event already carries the multiplicity.
+ expect(
+ startSpanMock.mock.calls.some((call) =>
+ JSON.stringify(call[1]).includes('suppressedSinceLast')
+ )
+ ).toBe(false)
+ })
+
+ // Bundle 8907a508 mixes restore-path (identity-less) and reattach-path crumbs in one
+ // window; name-only keying would report only the last one's shape.
+ it('keeps restore-path and reattach-path call sites in separate slots', () => {
+ emitRestoreWedge(1)
+ emitRestoreWedge(2)
+ emitReattachWedge(3)
+ emitReattachWedge(4, true)
+
+ const crumbs = wedgeCrumbs()
+ expect(crumbs).toHaveLength(3)
+ expect(crumbs.map((crumb) => Boolean(crumb.data?.tabIdHash))).toEqual([false, true, true])
+ expect(crumbs.map((crumb) => Boolean(crumb.data?.ptyId))).toEqual([false, false, true])
+ })
+
+ it('bounds a many-pane burst to one slot within a call site', () => {
+ for (let pane = 0; pane < 40; pane += 1) {
+ emitReattachWedge(pane, pane % 2 === 0)
+ }
+
+ expect(wedgeCrumbs()).toHaveLength(2)
+ })
+})
From 9acfba401a91f6be5419950fe920447f3bf047f6 Mon Sep 17 00:00:00 2001
From: Neil <4138956+nwparker@users.noreply.github.com>
Date: Fri, 4 Sep 2026 00:53:31 -0700
Subject: [PATCH 27/49] fix(crash-reporting): stop claiming kills that never
landed, and leave proof when the own-Chromium pid set is unreadable (#18578)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
* fix(crash-reporting): stop the codex POSIX teardown claiming a group that was already gone
terminatePosixTree's default group signal swallowed every process.kill error and
then recorded a self_tree_kill unconditionally, so an ESRCH — proof the group was
already gone and this teardown killed nothing — still put a suspect in the
five-second render-process-gone attribution window.
Every sibling group-kill in the tree already records only on a proven signal:
terminateDedicatedPosixGroup in this same file, forceKillPosixPtyProcessGroups,
and the claude account-login teardown. This makes the outlier match them.
* fix(crash-reporting): leave proof when the own-Chromium pid set cannot be read
`readOrcaChromiumProcessPids` returns an empty set when `getAppMetrics()`
throws, which is the right decision — refusing every kill would orphan every
PTY, git, codex and notebook tree main tears down, and on main a refusal from
`killSourceControlAgentProcess` releases the managed-home lock with the agent
still alive. But the empty set was byte-identical to "no Chromium on this
host", so the fail-open was invisible in a field bundle.
Keeps the decision, adds a coalesced durable `own_chromium_pids_unreadable`
crumb so the two cases are distinguishable. Coalesced because the gate reads
this set on every tree kill.
* style(crash-reporting): tighten the group-signal comments to the WHY
---
.../codex-app-server-process-teardown.test.ts | 62 ++++++++++++++++++-
.../codex-app-server-process-teardown.ts | 29 +++++----
src/main/orca-chromium-process-pids.ts | 30 ++++++++-
src/main/own-chromium-tree-kill-guard.test.ts | 34 ++++++++++
4 files changed, 141 insertions(+), 14 deletions(-)
diff --git a/src/main/codex/codex-app-server-process-teardown.test.ts b/src/main/codex/codex-app-server-process-teardown.test.ts
index 1ddb672a531..cec8f91d081 100644
--- a/src/main/codex/codex-app-server-process-teardown.test.ts
+++ b/src/main/codex/codex-app-server-process-teardown.test.ts
@@ -1,7 +1,14 @@
import type { ChildProcess } from 'node:child_process'
-import { describe, expect, it, vi } from 'vitest'
+import { beforeEach, describe, expect, it, vi } from 'vitest'
+import {
+ findSelfInitiatedTreeKills,
+ resetSelfInitiatedTreeKillLogForTest
+} from '../crash-reporting/self-initiated-tree-kill-log'
import { terminateCodexAppServerProcessTree } from './codex-app-server-process-teardown'
+/** Above pid_max on every supported POSIX host, so the group signal is a real ESRCH. */
+const UNREACHABLE_PGID = 2_147_483_647
+
function child() {
return {
pid: 1234,
@@ -10,6 +17,10 @@ function child() {
}
describe('terminateCodexAppServerProcessTree', () => {
+ beforeEach(() => {
+ resetSelfInitiatedTreeKillLogForTest()
+ })
+
it('waits for the Windows tree kill before releasing the wrapper', async () => {
const target = child()
const release = Promise.withResolvers()
@@ -120,6 +131,55 @@ describe('terminateCodexAppServerProcessTree', () => {
expect(target.kill).not.toHaveBeenCalled()
})
+ /**
+ * `selfInitiatedTreeKillCount` decides whether a `render-process-gone` was
+ * ours. A group that had already exited was killed by nobody, so crediting it
+ * puts a suspect in the five-second window that Orca never issued. Exercised
+ * through the real `process.kill(-pgid)` because the swallow being tested
+ * lives in the production default, not in an injectable seam.
+ */
+ it('does not claim a snapshot group that was already gone', async () => {
+ const target = { pid: UNREACHABLE_PGID, kill: vi.fn(() => true) as ChildProcess['kill'] }
+
+ await expect(
+ terminateCodexAppServerProcessTree(target, undefined, {
+ platform: 'darwin',
+ captureDescendants: async () => ({
+ rootPgid: UNREACHABLE_PGID,
+ descendants: [],
+ capturedAtMs: 1
+ }),
+ terminateDescendants: async () => true
+ })
+ ).resolves.toBe(true)
+
+ expect(target.kill).toHaveBeenLastCalledWith('SIGKILL')
+ expect(findSelfInitiatedTreeKills(Date.now())).toEqual([])
+ })
+
+ it('claims a snapshot group the signal actually reached', async () => {
+ const target = child()
+ const signalProcessGroup = vi.fn()
+
+ await expect(
+ terminateCodexAppServerProcessTree(target, undefined, {
+ platform: 'darwin',
+ captureDescendants: async () => ({ rootPgid: 1234, descendants: [], capturedAtMs: 1 }),
+ terminateDescendants: async () => true,
+ signalProcessGroup
+ })
+ ).resolves.toBe(true)
+
+ expect(signalProcessGroup).toHaveBeenCalledWith(1234, 'SIGKILL')
+ expect(findSelfInitiatedTreeKills(Date.now())).toEqual([
+ expect.objectContaining({
+ pid: 1234,
+ site: 'codex-app-server-teardown',
+ scope: 'posix-process-group'
+ })
+ ])
+ })
+
it('tears down 40 dedicated groups without process-table scans or cross-group fanout', async () => {
const killMocks = Array.from({ length: 40 }, () => vi.fn(() => true))
const targets = killMocks.map((kill, index) => ({
diff --git a/src/main/codex/codex-app-server-process-teardown.ts b/src/main/codex/codex-app-server-process-teardown.ts
index a35ad4d3164..5a9c6e3574b 100644
--- a/src/main/codex/codex-app-server-process-teardown.ts
+++ b/src/main/codex/codex-app-server-process-teardown.ts
@@ -128,19 +128,24 @@ async function terminatePosixTree(
if (descendantsExited && snapshot.rootPgid === rootPid) {
const signalGroup =
deps.signalProcessGroup ??
- ((pgid: number, signal: NodeJS.Signals) => {
- try {
- process.kill(-pgid, signal)
- } catch {
- // Group already exited.
- }
+ ((pgid: number, signal: NodeJS.Signals) => process.kill(-pgid, signal))
+ let groupSignalled = false
+ try {
+ signalGroup(snapshot.rootPgid, 'SIGKILL')
+ groupSignalled = true
+ } catch {
+ // Already-gone is still the desired outcome, but nothing here killed it,
+ // and a crumb for a kill we never landed is a false render-process-gone suspect.
+ }
+ if (groupSignalled) {
+ // Outside the try, as in terminateDedicatedPosixGroup: that catch is the
+ // already-gone contract, not a breadcrumb handler.
+ recordSelfInitiatedTreeKill({
+ pid: snapshot.rootPgid,
+ site: 'codex-app-server-teardown',
+ scope: 'posix-process-group'
})
- signalGroup(snapshot.rootPgid, 'SIGKILL')
- recordSelfInitiatedTreeKill({
- pid: snapshot.rootPgid,
- site: 'codex-app-server-teardown',
- scope: 'posix-process-group'
- })
+ }
}
if (!descendantsExited) {
child.kill('SIGCONT')
diff --git a/src/main/orca-chromium-process-pids.ts b/src/main/orca-chromium-process-pids.ts
index f22babc6921..b2613e42b79 100644
--- a/src/main/orca-chromium-process-pids.ts
+++ b/src/main/orca-chromium-process-pids.ts
@@ -1,4 +1,5 @@
import { getAppEnvironment, hasAppEnvironment } from '../shared/app-environment'
+import { recordCoalescedDurableCrashBreadcrumb } from './crash-reporting/durable-crash-breadcrumb'
/**
* PIDs of Orca's own Chromium processes — browser, renderers, GPU, utilities.
@@ -11,6 +12,14 @@ import { getAppEnvironment, hasAppEnvironment } from '../shared/app-environment'
* Empty on a Node host and empty on failure: that is "no refusal proven", never
* "safe to kill" — callers must keep every other guard they already have.
*
+ * Why failure stays open rather than refusing everything: a refusal is not free.
+ * `terminateWindowsProcessTree` resolves without killing, and
+ * `killSourceControlAgentProcess` returns that straight to a caller that then
+ * releases the managed-home lock, so failing closed would trade one unreadable
+ * metrics table for every PTY, git, codex and notebook tree in main leaking at
+ * once. The `own_chromium_pids_unreadable` crumb is the price of that choice:
+ * without it a throw is byte-identical to "no Chromium on this host".
+ *
* Host coverage: only Electron main installs a Chromium-backed AppEnvironment
* (main-process-preflight). The standalone daemon installs none and `orcad`
* installs a Node one whose `getAppMetrics()` is `[]`, so this set is empty in
@@ -30,7 +39,26 @@ export function readOrcaChromiumProcessPids(): ReadonlySet {
.map((metric) => metric.pid)
.filter((pid) => Number.isInteger(pid) && pid > 0)
return new Set(pids)
- } catch {
+ } catch (error) {
+ recordUnreadableOwnChromiumMetrics(error)
return new Set()
}
}
+
+// Why coalesced: the gate reads this set on every tree kill, so a persistently
+// broken metrics table would otherwise flood the 30-slot ring it shares.
+const UNREADABLE_METRICS_COALESCE_MS = 60_000
+
+function recordUnreadableOwnChromiumMetrics(error: unknown): void {
+ try {
+ recordCoalescedDurableCrashBreadcrumb({
+ name: 'own_chromium_pids_unreadable',
+ data: { cause: error instanceof Error ? error.message : String(error) },
+ coalesceKey: 'own-chromium-pids-unreadable',
+ minIntervalMs: UNREADABLE_METRICS_COALESCE_MS
+ })
+ } catch {
+ // Diagnostics must never turn an admitted kill into a thrown one: callers
+ // read this set outside their own try.
+ }
+}
diff --git a/src/main/own-chromium-tree-kill-guard.test.ts b/src/main/own-chromium-tree-kill-guard.test.ts
index 7e98661aca6..bd3b1674e18 100644
--- a/src/main/own-chromium-tree-kill-guard.test.ts
+++ b/src/main/own-chromium-tree-kill-guard.test.ts
@@ -143,6 +143,40 @@ describe('refusing to tree-kill our own Chromium processes', () => {
)
})
+ /**
+ * Fail-open is the deliberate choice — see `orca-chromium-process-pids.ts` for
+ * why refusing everything is worse — so the crumb is the only thing that keeps
+ * an unreadable metrics table distinguishable from a host that has no Chromium.
+ */
+ it('leaves proof, and still admits the kill, when the Chromium metrics cannot be read', () => {
+ appMetricsMock.mockImplementation(() => {
+ throw new Error('getAppMetrics unavailable')
+ })
+
+ expect([...readOrcaChromiumProcessPids()]).toEqual([])
+ // Coalesced: the gate reads this set on every kill, so a broken table must
+ // not evict the ring it shares with the refusal crumb.
+ expect([...readOrcaChromiumProcessPids()]).toEqual([])
+ expect(
+ admitSelfInitiatedTreeKill({
+ pid: RENDERER_PID,
+ site: 'pty-descendant-sweep',
+ scope: 'win-taskkill-tree'
+ })
+ ).toBe(true)
+
+ expect(
+ getCrashBreadcrumbSnapshot().filter(
+ (breadcrumb) => breadcrumb.name === 'own_chromium_pids_unreadable'
+ )
+ ).toEqual([
+ expect.objectContaining({
+ name: 'own_chromium_pids_unreadable',
+ data: expect.objectContaining({ cause: 'getAppMetrics unavailable' })
+ })
+ ])
+ })
+
it('refuses an own-Chromium pid at the gate the account teardowns share', () => {
expect(
admitSelfInitiatedTreeKill({
From 7a714d1bd2750789cd26ad213f9ae2f129807eef Mon Sep 17 00:00:00 2001
From: Neil <4138956+nwparker@users.noreply.github.com>
Date: Fri, 4 Sep 2026 00:53:34 -0700
Subject: [PATCH 28/49] fix(terminals): add equality bailouts to the tab
pane-expansion actions (#18332)
* fix(terminals): bail out of no-op pane-expansion store writes
* test(terminals): lock the root-state identity of the bailout
A `return {}` bailout keeps the map reference but still allocates a new
root state, so zustand walks every listener. Assert root identity too.
---
.../store/terminals/terminal-layout-state.ts | 17 +-
...inal-pane-expansion-write-bailout.test.tsx | 152 ++++++++++++++++++
2 files changed, 163 insertions(+), 6 deletions(-)
create mode 100644 src/renderer/src/store/terminals/terminal-pane-expansion-write-bailout.test.tsx
diff --git a/src/renderer/src/store/terminals/terminal-layout-state.ts b/src/renderer/src/store/terminals/terminal-layout-state.ts
index d95439cf5a8..9b5cf8323c2 100644
--- a/src/renderer/src/store/terminals/terminal-layout-state.ts
+++ b/src/renderer/src/store/terminals/terminal-layout-state.ts
@@ -43,15 +43,20 @@ export function createTerminalLayoutActions(
}
})
},
+ // Why: pane mount/unmount re-asserts the same booleans; bailing like setTabLayout keeps map subscribers asleep.
setTabPaneExpanded: (tabId, expanded) => {
- set((s) => ({
- expandedPaneByTabId: { ...s.expandedPaneByTabId, [tabId]: expanded }
- }))
+ set((s) =>
+ s.expandedPaneByTabId[tabId] === expanded
+ ? s
+ : { expandedPaneByTabId: { ...s.expandedPaneByTabId, [tabId]: expanded } }
+ )
},
setTabCanExpandPane: (tabId, canExpand) => {
- set((s) => ({
- canExpandPaneByTabId: { ...s.canExpandPaneByTabId, [tabId]: canExpand }
- }))
+ set((s) =>
+ s.canExpandPaneByTabId[tabId] === canExpand
+ ? s
+ : { canExpandPaneByTabId: { ...s.canExpandPaneByTabId, [tabId]: canExpand } }
+ )
},
setTabLayout: (tabId, layout) => {
let ownershipTransfers: ReturnType = []
diff --git a/src/renderer/src/store/terminals/terminal-pane-expansion-write-bailout.test.tsx b/src/renderer/src/store/terminals/terminal-pane-expansion-write-bailout.test.tsx
new file mode 100644
index 00000000000..d0d32f7742f
--- /dev/null
+++ b/src/renderer/src/store/terminals/terminal-pane-expansion-write-bailout.test.tsx
@@ -0,0 +1,152 @@
+// @vitest-environment happy-dom
+
+import { Profiler } from 'react'
+import { act, cleanup, render } from '@testing-library/react'
+import { afterEach, describe, expect, it } from 'vitest'
+import { createTestStore } from '../slices/store-test-helpers'
+
+afterEach(cleanup)
+
+type TestStore = ReturnType
+
+const TAB_ID = 'tab-1'
+const NO_OP_WRITES = 25
+
+function recordPublishedMapKeys(store: TestStore): string[] {
+ const published: string[] = []
+ store.subscribe((next, previous) => {
+ if (next.expandedPaneByTabId !== previous.expandedPaneByTabId) {
+ published.push('expandedPaneByTabId')
+ }
+ if (next.canExpandPaneByTabId !== previous.canExpandPaneByTabId) {
+ published.push('canExpandPaneByTabId')
+ }
+ })
+ return published
+}
+
+// Mirrors use-terminal-workspace-store-bindings.ts:17, which subscribes to the raw map.
+function ExpandedPaneSubscriber({ store }: { store: TestStore }): React.JSX.Element {
+ const expandedPaneByTabId = store((s) => s.expandedPaneByTabId)
+ return {String(expandedPaneByTabId[TAB_ID] === true)}
+}
+
+function CanExpandPaneSubscriber({ store }: { store: TestStore }): React.JSX.Element {
+ const canExpandPaneByTabId = store((s) => s.canExpandPaneByTabId)
+ return {String(canExpandPaneByTabId[TAB_ID] === true)}
+}
+
+function renderCommitCounter(subscriber: React.JSX.Element): () => number {
+ let commits = 0
+ render(
+ {
+ commits += 1
+ }}
+ >
+ {subscriber}
+
+ )
+ const mountCommits = commits
+ return () => commits - mountCommits
+}
+
+describe('setTabPaneExpanded', () => {
+ it('publishes the first write for an unseen tab and a real toggle', () => {
+ const store = createTestStore()
+ const published = recordPublishedMapKeys(store)
+
+ store.getState().setTabPaneExpanded(TAB_ID, false)
+ expect(published).toEqual(['expandedPaneByTabId'])
+ expect(store.getState().expandedPaneByTabId[TAB_ID]).toBe(false)
+
+ store.getState().setTabPaneExpanded(TAB_ID, true)
+ expect(published).toEqual(['expandedPaneByTabId', 'expandedPaneByTabId'])
+ expect(store.getState().expandedPaneByTabId[TAB_ID]).toBe(true)
+ })
+
+ it('bails out when the value is unchanged', () => {
+ const store = createTestStore()
+ store.getState().setTabPaneExpanded(TAB_ID, false)
+ const before = store.getState().expandedPaneByTabId
+ // Root identity too: returning `{}` keeps the map but allocates a new root, so zustand still walks every listener.
+ const rootBefore = store.getState()
+ const published = recordPublishedMapKeys(store)
+
+ for (let i = 0; i < NO_OP_WRITES; i += 1) {
+ store.getState().setTabPaneExpanded(TAB_ID, false)
+ }
+
+ expect(published).toEqual([])
+ expect(store.getState().expandedPaneByTabId).toBe(before)
+ expect(store.getState()).toBe(rootBefore)
+ })
+
+ it('costs no React commit in a map subscriber when the value is unchanged', () => {
+ const store = createTestStore()
+ store.getState().setTabPaneExpanded(TAB_ID, false)
+ const commitsSinceMount = renderCommitCounter()
+
+ for (let i = 0; i < NO_OP_WRITES; i += 1) {
+ act(() => {
+ store.getState().setTabPaneExpanded(TAB_ID, false)
+ })
+ }
+ expect(commitsSinceMount()).toBe(0)
+
+ act(() => {
+ store.getState().setTabPaneExpanded(TAB_ID, true)
+ })
+ expect(commitsSinceMount()).toBe(1)
+ })
+})
+
+describe('setTabCanExpandPane', () => {
+ it('publishes the first write for an unseen tab and a real toggle', () => {
+ const store = createTestStore()
+ const published = recordPublishedMapKeys(store)
+
+ store.getState().setTabCanExpandPane(TAB_ID, false)
+ expect(published).toEqual(['canExpandPaneByTabId'])
+ expect(store.getState().canExpandPaneByTabId[TAB_ID]).toBe(false)
+
+ store.getState().setTabCanExpandPane(TAB_ID, true)
+ expect(published).toEqual(['canExpandPaneByTabId', 'canExpandPaneByTabId'])
+ expect(store.getState().canExpandPaneByTabId[TAB_ID]).toBe(true)
+ })
+
+ it('bails out when the value is unchanged', () => {
+ const store = createTestStore()
+ store.getState().setTabCanExpandPane(TAB_ID, false)
+ const before = store.getState().canExpandPaneByTabId
+ const rootBefore = store.getState()
+ const published = recordPublishedMapKeys(store)
+
+ for (let i = 0; i < NO_OP_WRITES; i += 1) {
+ store.getState().setTabCanExpandPane(TAB_ID, false)
+ }
+
+ expect(published).toEqual([])
+ expect(store.getState().canExpandPaneByTabId).toBe(before)
+ expect(store.getState()).toBe(rootBefore)
+ })
+
+ it('costs no React commit in a map subscriber when the value is unchanged', () => {
+ const store = createTestStore()
+ store.getState().setTabCanExpandPane(TAB_ID, false)
+ const commitsSinceMount = renderCommitCounter()
+
+ for (let i = 0; i < NO_OP_WRITES; i += 1) {
+ act(() => {
+ store.getState().setTabCanExpandPane(TAB_ID, false)
+ })
+ }
+ expect(commitsSinceMount()).toBe(0)
+
+ act(() => {
+ store.getState().setTabCanExpandPane(TAB_ID, true)
+ })
+ expect(commitsSinceMount()).toBe(1)
+ })
+})
From cc9e9ed65fc3d4f69d79224eb66c437fb819528d Mon Sep 17 00:00:00 2001
From: Neil <4138956+nwparker@users.noreply.github.com>
Date: Fri, 4 Sep 2026 00:53:37 -0700
Subject: [PATCH 29/49] fix(crash-reporting): sample system memory before the
process is gone (#18356)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
* fix(crash-reporting): sample system memory before the renderer dies
* fix(crash-reporting): make the pre-gone host sample decisive, not just present
Round-1 review said the shipped field set could not decide G4-oom. Fixed.
Decisive field (blocking #1). The investigation's own win-lowspec repro
falsified "low available commit kills": at a 127 MB commit floor Windows grew
the pagefile to 2029 MB and nothing died, and it named the missing datum —
pagefile-growth headroom / system-drive free space. `getSystemMemoryInfo()`
gives neither. Added `swap-volume-free-space.ts`: one `fs.statfs` on the volume
backing the pagefile (SystemRoot on Windows, the root fs elsewhere, resolved
via `path.parse().root`), published as `systemMemoryPreGoneSwapVolumeFreeMB`.
Together with the already-emitted commit limit that separates "commit was low"
from "commit was refused". Pagefile *max* size needs a registry read; skipped
deliberately — per-operation interpreter spawning is exactly what
docs/reference/windows-edr-posture.md says not to add for telemetry.
Darwin honesty (blocking #2). Every reading now carries
`systemMemoryPressureSignal`: `available-commit` on Windows (swapFree is
ullAvailPageFile), `mem-available` on Linux when MemAvailable is present,
`none` otherwise — which is always on darwin. A future analyst cannot now table
`freeMB: 272` from a healthy Mac as evidence of exhaustion, because the same
record says the platform gave no pressure verdict. Partial rebuttal on the
suggested reuse: `host-memory.ts:86` was considered and rejected as a periodic
source. It spawns `/usr/bin/memory_pressure` per call, and the sampler this PR
needs runs every 10 s for the app's lifetime; a subprocess at that cadence is
worse than the gap it closes. The reviewer conceded this tradeoff is arguable —
what was not acceptable was shipping the darwin gap silently, so it is now in
the data, not only in a comment.
Staleness (blocking #3). Confirmed the measurement: four of five G4 reports
carried a ~37 s-old sample (4872/36796/37332/38017/39715 ms). Host memory no
longer rides the 60 s process-metrics sweep; `pre-gone-host-memory.ts` samples
it on its own 10 s timer with its own `systemMemoryPreGoneSampleAgeMs`. One
GlobalMemoryStatusEx-class call plus one statfs is cheap enough at that rate. A
refusal shorter than the interval stays invisible and the module comment says
so — no polling cadence fixes that.
Non-blocking, all taken: renamed `gone-time-system-memory.ts` ->
`system-memory-details.ts` with the now-false "reads AFTER the crash" framing
scoped to the gone-time caller; pre-gone host keys moved out of the
`processMetrics` namespace to `systemMemoryPreGone*`, so the string-surgery
`preGoneDetailKey` helper is gone and a `systemMemory` prefix scan sees both
reads; the bare catch no longer spans both halves of the sample, and a test
pins that a throwing host read leaves the process-metric sample intact; the
inert second test is replaced by three that go red without this change
(verified: swap-volume, pressure-signal and cadence assertions all fail when
the production hunks are reverted).
Rebuttal, non-blocking #5 (duplicated electron mock across two test files):
declined. `vi.mock` is hoisted per file, so the mock cannot be shared without a
setup module, and this directory already has 26 focused test files that each
re-declare it. Splitting by concern is the local convention.
`startPreGoneProcessMetricsSampling` is renamed `startPreGoneCrashSampling`
since it now starts two samplers.
* fix(crash-reporting): test the arming, gate the swap volume, unblock the host read
Round-2 review blocked on four items. All four addressed.
WHAT THIS BRANCH ACTUALLY DOES, AT HEAD (blocking #4). The commit-1 message
("13 lines, 1 production file, no new module, new optional numeric fields
only", `processMetricsPreGoneSystemMemory*` keys, a `preGoneDetailKey` helper,
a `pre-gone-system-memory.test.ts`) describes a superseded revision; every one
of those claims is false now, so it must not be used as the PR description.
The change against origin/main is: 3 new production modules
(`pre-gone-host-memory.ts`, `system-memory-details.ts`,
`swap-volume-free-space.ts`), 1 deleted (`gone-time-system-memory.ts`), plus
edits to `process-gone-diagnostics.ts` and `main-process-ready-runtime.ts` and
2 test files. It adds a second main-process interval timer that runs for the
life of the app: every 10 s one synchronous GlobalMemoryStatusEx-class read,
and on win32/darwin one `fs.statfs` on the swap-backing volume. Details are
`systemMemoryPreGone*`, and two of them are STRINGS, not numbers:
`systemMemoryPreGonePressureSignal` (enum) and `systemMemoryPreGoneSwapVolume`
(a drive label, separator-trimmed so it is not a path). Both are assigned after
`sanitizeCrashReportDetails`; neither carries user content.
Arming is now tested (blocking #1). The reviewer deleted
`startPreGoneSystemMemorySampling(...)` from `startPreGoneCrashSampling` and
all 264 tests stayed green — confirmed and fixed. `pre-gone-host-memory.test.ts`
now calls `startPreGoneCrashSampling()` with production defaults and asserts
both `setInterval` calls, their literal periods `[60_000, 10_000]`, that both
timers are unref'd, and that advancing 10 s takes a fresh host sample that
reaches `buildProcessGoneCrashDetails` with `SampleAgeMs: 0`. Verified red on
revert: deleting the arming line -> 1 failure; changing the interval constant
to 30_000 -> 1 failure (the old assertion compared the constant to itself and
caught neither). The tautological `10_000 < 60_000 / 2` test is gone,
superseded by this one.
Swap volume is win32/darwin only (blocking #2). On Linux swap is a fixed
partition, a fixed-size swapfile, or zram; none grow into root-fs free space,
so `SwapVolumeFreeMB: 380000` beside `SwapFreeMB: 0` would have invited exactly
the wrong verdict on the two Linux cluster members. `swapVolumeAnchor` returns
undefined off win32/darwin, so no field and no statfs at all. The comment
claiming "elsewhere swap is on the root fs" was wrong and is gone. The Windows
anchor is still the DEFAULT pagefile volume, so the measured volume now ships
with the number (`systemMemoryPreGoneSwapVolume: 'C:'`) instead of being
implied. The honesty label covers it: win32 reads `available-commit` only when
the volume datum is present, and `available-commit-unqualified` otherwise —
which also fixes non-blocking #5, where the synchronous gone-time read claimed
a verdict its own fields could not support.
Host read no longer waits on statfs (blocking #3). `samplePreGoneSystemMemory`
now commits the synchronous memory reading first and merges volume free space
in afterwards, so the cadence is 10 s regardless of disk-metadata latency and a
hung volume can no longer stop host sampling — precisely the paging-storm case
this exists for. The in-flight latch now guards only the statfs. Verified red
on revert to the serialized shape (2 failures). A stale-but-slow-moving volume
value merging into a newer memory sample is deliberate and commented.
Non-blocking #3 (reset does not invalidate an in-flight sample): fixed with a
generation counter bumped by `resetPreGoneSystemMemorySamplingForTest`, so a
late statfs cannot repopulate a reset sample. Separately, the volume read now
only runs after a host sample committed, which removes the real `statfs('/')`
side effect from `process-gone-diagnostics.test.ts` entirely.
REBUTTAL, darwin `memory_pressure` reuse (non-blocking #2): declined, with
evidence. `readDarwinAvailableMemory` at src/main/memory/host-memory.ts:87 is
reached only via `collectHostMemory` <- `runSnapshot` <- `collectMemorySnapshot`,
whose only callers are the `memory:getSnapshot` IPC handler and
orca-runtime-pty-foreground-process-reads.ts:170 — both on demand. There is no
periodic snapshot, so there is no cached reading to reuse for free; adopting it
means spawning `/usr/bin/memory_pressure` on a main-process timer for the life
of the app, and its module-global `darwinAvailabilitySupported` latch is shared
with the memory UI. The gap is not hidden: darwin ships
`PressureSignal: 'none'` in the data, and the module comment now cites the
existing reader and why it is not used here rather than claiming Orca lacks
one.
Verified: `vitest src/main/crash-reporting src/main/startup src/main/memory` =
769 passed / 6 skipped (crash-reporting re-run 5x, no flake);
`tsc --noEmit -p config/tsconfig.node.json` 0; `oxlint` 0; `oxfmt --check` 0.
* fix(crash-reporting): stop a stale statfs qualifying the commit verdict
Round-3 adversarial review, 2 blocking. Both fixed with mutation-verified
tests.
1. `mergeSwapVolumeFreeSpace` merged the volume reading into whatever sample
was current at RESOLUTION time, and `pressureSignal` then upgraded win32
from `available-commit-unqualified` to the decisive `available-commit` on
the strength of it. The `swapVolumeReadInFlight` latch makes every
intervening tick skip the merge, so the lag is as old as the last STARTED
statfs, not the last tick — and no age field exposed it, because
`systemMemoryPreGoneSampleAgeMs` describes only the synchronous memory read.
Reviewer's executed scenario: a statfs issued at t=0 on a healthy host
(40 GB free) resolving at t=20 s of commit pressure emitted
`SwapFreeMB: 200` beside `SwapVolumeFreeMB: 40000`, labelled
`available-commit`, with `SampleAgeMs: 0`. That reads as "the pagefile had
room, so this was not a commit refusal" — the opposite conclusion, wearing
the branch's highest-confidence label, on exactly the win32 G4-oom reports
this exists to decide.
The datum still ships (it is the only pagefile-expandability signal there
is), but now:
- the sample carries `swapVolumeSampledAtMs` — the tick that ISSUED the
statfs, never the one it resolved on — surfaced as
`systemMemoryPreGoneSwapVolumeAgeMs`;
- only a statfs that answers on its own tick may qualify the verdict.
`withSwapVolumeFreeSpace` takes `coTimed`; false keeps
`available-commit-unqualified`.
The next tick issues a fresh statfs, so the verdict recovers on its own.
2. The branch's sole production entry point — `startPreGoneCrashSampling()` at
main-process-ready-runtime.ts:128 — was untested. Deleting it left 691
tests across crash-reporting/ and startup/ green, while a comment in the
new test file claimed that gap was why the test was written. This is pure
instrumentation, so that one line is the whole of its value in the shipped
app. Added a source-level wiring test (the pattern this repo already uses
for arm-once ready-phase lines) that pins the import, exactly one call, the
call at statement indent, and that `main-process-ready.ts` awaits the
function it lives in. The misleading comment is gone.
Mutation-verified — each goes red alone:
coTimed -> always true 1 failed (verdict)
drop swapVolumeSampledAtMs age 1 failed (verdict test)
delete startPreGoneCrashSampling() 1 failed (wiring)
wrap it in `if (!is.dev) { ... }` 1 failed (wiring)
Verified: tsc -p config/tsconfig.node.json exit 0; oxlint
src/main/crash-reporting src/main/startup exit 0; 268 tests in
crash-reporting/ pass. Across crash-reporting/ + startup/ + memory/: 769
passed, 2 failed — both environment-dependent and failing identically on the
unmodified tree (Xvfb rebind, and a whole-repo glob census that times out).
* fix(crash-reporting): stop free disk standing in for pagefile growability
The win32 reading was promoted to the decisive `available-commit` whenever a
co-timed volume number merely existed, which the data cannot support: a fixed
or disabled pagefile grows into no amount of empty disk, its maximum is
unreadable here, and the measured volume is only the DEFAULT pagefile drive. A
host with 180 MB of available commit, a commit limit at RAM and 812 GB free
read as "the pagefile had room" — the opposite conclusion, under the branch's
most confident label.
The volume datum is now named for what it is (`available-commit-volume-cotimed`,
context beside the commit number), and the one decisive win32 case — a commit
limit at or below RAM, i.e. no pagefile behind it — gets its own label.
Also: carry the last volume reading onto the sample that replaces it, aged and
non-qualifying, so a statfs slower than one tick no longer makes the field
vanish from the reports it exists for; don't commit a reading whose every
memory field failed, which shipped an age and a disk-free number with no host
memory beside them; and move the startup wiring test beside the file it pins,
scoped to the ready-phase entry's own body so the call cannot satisfy it from
a sibling export nothing calls.
* fix(crash-reporting): co-time the statfs by tick, not sample identity
A tick whose host read fails leaves the pre-gone sample object in place, so
the identity check still read a 25 s-late statfs as co-timed.
---
.../gone-time-system-memory.ts | 77 ----
.../pre-gone-host-memory.test.ts | 379 ++++++++++++++++++
.../crash-reporting/pre-gone-host-memory.ts | 164 ++++++++
.../process-gone-diagnostics.test.ts | 22 +-
.../process-gone-diagnostics.ts | 23 +-
.../crash-reporting/swap-volume-free-space.ts | 67 ++++
.../crash-reporting/system-memory-details.ts | 161 ++++++++
.../startup/main-process-ready-runtime.ts | 9 +-
.../pre-gone-crash-sampling-wiring.test.ts | 49 +++
9 files changed, 854 insertions(+), 97 deletions(-)
delete mode 100644 src/main/crash-reporting/gone-time-system-memory.ts
create mode 100644 src/main/crash-reporting/pre-gone-host-memory.test.ts
create mode 100644 src/main/crash-reporting/pre-gone-host-memory.ts
create mode 100644 src/main/crash-reporting/swap-volume-free-space.ts
create mode 100644 src/main/crash-reporting/system-memory-details.ts
create mode 100644 src/main/startup/pre-gone-crash-sampling-wiring.test.ts
diff --git a/src/main/crash-reporting/gone-time-system-memory.ts b/src/main/crash-reporting/gone-time-system-memory.ts
deleted file mode 100644
index 7cca89d2d4b..00000000000
--- a/src/main/crash-reporting/gone-time-system-memory.ts
+++ /dev/null
@@ -1,77 +0,0 @@
-import type { CrashReportDetailValue } from '../../shared/crash-reporting'
-
-// ─── System memory at gone time ─────────────────────────────────────
-// Why: the system outlives the crashed process, so this IS sampleable at
-// process-gone — it separates "renderer grew huge" from "machine out of
-// memory/commit", which the per-process buckets alone cannot.
-// Timing honesty: this reads AFTER the crashed process's memory returned to
-// the OS, so free/swapFree can look healthier than they were at kill time.
-// Platform honesty: swap* exist on Windows/Linux only. On Linux `free` is
-// /proc/meminfo MemFree and is NOT the pressure signal — it excludes page cache
-// and other reclaimable memory; `available` (MemAvailable, Linux-only) is. On
-// macOS `free` is near-meaningless (file cache and compression keep it low on
-// healthy machines); fileBacked/purgeable are the only reclaimability proxy this
-// API gives there, and none of these fields answers "was the machine under
-// pressure" on macOS — that needs a signal Electron does not expose.
-
-type CrashReportDetails = Record
-
-export function memoryKBFieldMB(value: unknown): number | undefined {
- const kb = typeof value === 'number' && Number.isFinite(value) ? value : undefined
- return kb === undefined ? undefined : Math.round(Math.max(0, kb) / 1024)
-}
-
-type SystemMemoryInfoLike = {
- total?: unknown
- free?: unknown
- available?: unknown
- swapTotal?: unknown
- swapFree?: unknown
- fileBacked?: unknown
- purgeable?: unknown
-}
-
-type SystemMemoryInfoReader = () => SystemMemoryInfoLike | null
-
-function readElectronSystemMemoryInfo(): SystemMemoryInfoLike | null {
- const read = (process as NodeJS.Process & { getSystemMemoryInfo?: () => SystemMemoryInfoLike })
- .getSystemMemoryInfo
- if (typeof read !== 'function') {
- return null
- }
- try {
- return read.call(process)
- } catch {
- return null
- }
-}
-
-let systemMemoryInfoReader: SystemMemoryInfoReader = readElectronSystemMemoryInfo
-
-export function setSystemMemoryInfoReaderForTest(reader: SystemMemoryInfoReader | null): void {
- systemMemoryInfoReader = reader ?? readElectronSystemMemoryInfo
-}
-
-export function getSystemMemoryAtGoneDetails(): CrashReportDetails {
- const info = systemMemoryInfoReader()
- if (!info) {
- return {}
- }
- const details: CrashReportDetails = {}
- const fields: readonly [keyof SystemMemoryInfoLike, string][] = [
- ['total', 'systemMemoryTotalMB'],
- ['free', 'systemMemoryFreeMB'],
- ['available', 'systemMemoryAvailableMB'],
- ['swapTotal', 'systemMemorySwapTotalMB'],
- ['swapFree', 'systemMemorySwapFreeMB'],
- ['fileBacked', 'systemMemoryFileBackedMB'],
- ['purgeable', 'systemMemoryPurgeableMB']
- ]
- for (const [field, key] of fields) {
- const mb = memoryKBFieldMB(info[field])
- if (mb !== undefined) {
- details[key] = mb
- }
- }
- return details
-}
diff --git a/src/main/crash-reporting/pre-gone-host-memory.test.ts b/src/main/crash-reporting/pre-gone-host-memory.test.ts
new file mode 100644
index 00000000000..0df13d4fee5
--- /dev/null
+++ b/src/main/crash-reporting/pre-gone-host-memory.test.ts
@@ -0,0 +1,379 @@
+import { beforeEach, describe, expect, it, vi } from 'vitest'
+import {
+ getSystemMemoryDetails,
+ setSystemMemoryInfoReaderForTest,
+ withSwapVolumeFreeSpace
+} from './system-memory-details'
+import {
+ readSwapVolumeFreeSpace,
+ setSwapVolumeFreeSpaceReaderForTest,
+ type SwapVolumeFreeSpace
+} from './swap-volume-free-space'
+import { samplePreGoneSystemMemory } from './pre-gone-host-memory'
+import {
+ buildProcessGoneCrashDetails,
+ resetPreGoneCrashSamplingForTest,
+ samplePreGoneProcessMetrics,
+ startPreGoneCrashSampling
+} from './process-gone-diagnostics'
+
+type MetricFixture = {
+ pid: number
+ creationTime: number
+ type: string
+ memory: { workingSetSize: number; peakWorkingSetSize?: number; privateBytes?: number }
+}
+
+const { appMetricsMock } = vi.hoisted(() => ({
+ appMetricsMock: vi.fn<() => MetricFixture[]>(() => [])
+}))
+
+vi.mock('electron', () => ({ app: { getAppMetrics: appMetricsMock } }))
+
+const BROWSER_AND_RENDERER: MetricFixture[] = [
+ { pid: 10, creationTime: 1, type: 'Browser', memory: { workingSetSize: 1024 * 250 } },
+ {
+ pid: 11,
+ creationTime: 2,
+ type: 'Tab',
+ memory: { workingSetSize: 1024 * 400, peakWorkingSetSize: 1024 * 420, privateBytes: 1024 * 260 }
+ }
+]
+
+const BROWSER_ONLY: MetricFixture[] = [BROWSER_AND_RENDERER[0]]
+
+const UNDER_COMMIT_PRESSURE = {
+ total: 16_000 * 1024,
+ free: 400 * 1024,
+ swapTotal: 48_000 * 1024,
+ swapFree: 200 * 1024
+}
+
+const AFTER_THE_CORPSE_RELEASED = {
+ total: 16_000 * 1024,
+ free: 3_000 * 1024,
+ swapTotal: 48_000 * 1024,
+ swapFree: 2_900 * 1024
+}
+
+// Commit limit ~= RAM: a disabled or fixed pagefile, which no amount of empty
+// disk can grow into. `swapTotal > total` is all this API can say about that.
+const FIXED_PAGEFILE_UNDER_PRESSURE = {
+ total: 16_000 * 1024,
+ free: 300 * 1024,
+ swapTotal: 16_100 * 1024,
+ swapFree: 180 * 1024
+}
+
+const NO_PAGEFILE_UNDER_PRESSURE = {
+ ...FIXED_PAGEFILE_UNDER_PRESSURE,
+ swapTotal: 15_900 * 1024
+}
+
+const BEFORE_THE_STORM = {
+ total: 16_000 * 1024,
+ free: 9_000 * 1024,
+ swapTotal: 48_000 * 1024,
+ swapFree: 30_000 * 1024
+}
+
+describe('pre-gone host memory', () => {
+ beforeEach(() => {
+ resetPreGoneCrashSamplingForTest()
+ setSystemMemoryInfoReaderForTest(null)
+ setSwapVolumeFreeSpaceReaderForTest(null)
+ appMetricsMock.mockClear()
+ appMetricsMock.mockReturnValue(BROWSER_AND_RENDERER)
+ })
+
+ it('carries a pre-gone host reading, not only the post-mortem one', async () => {
+ setSystemMemoryInfoReaderForTest(() => UNDER_COMMIT_PRESSURE)
+ setSwapVolumeFreeSpaceReaderForTest(() => Promise.resolve({ freeMB: 120, volume: 'C:' }))
+ await samplePreGoneSystemMemory(Date.now() - 5_000)
+
+ // The renderer dies; its ~400 MB returns to the OS, so the gone-time read
+ // now shows a much healthier machine than the one that refused the alloc.
+ setSystemMemoryInfoReaderForTest(() => AFTER_THE_CORPSE_RELEASED)
+ appMetricsMock.mockReturnValue(BROWSER_ONLY)
+
+ const details = buildProcessGoneCrashDetails({ processType: 'renderer' }, 'renderer')
+
+ expect(details.systemMemorySwapFreeMB).toBe(2_900)
+ expect(details.systemMemoryPreGoneSwapFreeMB).toBe(200)
+ expect(details.systemMemoryPreGoneFreeMB).toBe(400)
+ expect(details.systemMemoryPreGoneTotalMB).toBe(16_000)
+ // Why: host memory keeps its own key family, so a `systemMemory` prefix scan sees both reads.
+ expect(
+ Object.keys(details).filter((key) => key.startsWith('processMetricsPreGoneSystem'))
+ ).toEqual([])
+ })
+
+ // Why this decides the cluster: 200 MB available commit is only a REFUSAL when
+ // the pagefile cannot grow, which is what the volume's free space says.
+ it('reports swap-volume free space so low commit can be told from refused commit', async () => {
+ setSystemMemoryInfoReaderForTest(() => UNDER_COMMIT_PRESSURE)
+ setSwapVolumeFreeSpaceReaderForTest(() => Promise.resolve({ freeMB: 120, volume: 'C:' }))
+ await samplePreGoneSystemMemory(Date.now() - 5_000)
+
+ const details = buildProcessGoneCrashDetails({ processType: 'renderer' }, 'renderer')
+
+ expect(details.systemMemoryPreGoneSwapVolumeFreeMB).toBe(120)
+ // Which volume was measured: Windows only names the DEFAULT pagefile drive.
+ expect(details.systemMemoryPreGoneSwapVolume).toBe('C:')
+ })
+
+ it('omits swap-volume free space on Linux, where swap cannot grow into free disk', async () => {
+ // Linux swap is a fixed partition, a fixed-size swapfile, or zram; reporting
+ // root-fs free space next to SwapFreeMB 0 would read as headroom that is not there.
+ setSwapVolumeFreeSpaceReaderForTest(null)
+
+ await expect(readSwapVolumeFreeSpace('linux')).resolves.toBeUndefined()
+ })
+
+ it('labels the reading with the pressure verdict the platform can actually give', () => {
+ // Windows available commit is only a REFUSAL when the pagefile cannot grow,
+ // which nothing here proves, so no label may read as that verdict.
+ setSystemMemoryInfoReaderForTest(() => UNDER_COMMIT_PRESSURE)
+ const windowsCommit = getSystemMemoryDetails('win32')
+ expect(windowsCommit.systemMemoryPressureSignal).toBe('available-commit-unqualified')
+ expect(
+ withSwapVolumeFreeSpace(windowsCommit, { freeMB: 120, volume: 'C:' }, 'win32')
+ .systemMemoryPressureSignal
+ ).toBe('available-commit-volume-cotimed')
+ // A volume number from a different moment describes a different machine.
+ expect(
+ withSwapVolumeFreeSpace(windowsCommit, { freeMB: 120, volume: 'C:' }, 'win32', false)
+ .systemMemoryPressureSignal
+ ).toBe('available-commit-unqualified')
+
+ setSystemMemoryInfoReaderForTest(() => ({ total: 16_000 * 1024, free: 400 * 1024 }))
+ expect(getSystemMemoryDetails('linux').systemMemoryPressureSignal).toBe('none')
+
+ setSystemMemoryInfoReaderForTest(() => ({ total: 16_000 * 1024, available: 900 * 1024 }))
+ expect(getSystemMemoryDetails('linux').systemMemoryPressureSignal).toBe('mem-available')
+
+ // darwin free/fileBacked/purgeable answer reclaimability, never pressure.
+ setSystemMemoryInfoReaderForTest(() => ({
+ total: 16_000 * 1024,
+ free: 272 * 1024,
+ fileBacked: 2_694 * 1024,
+ purgeable: 0
+ }))
+ expect(getSystemMemoryDetails('darwin').systemMemoryPressureSignal).toBe('none')
+ })
+
+ // Why this and not the volume number: the branch's own repro needed a pagefile
+ // that CANNOT grow to kill anything, and neither the pagefile maximum nor its
+ // drive is readable here — `swapVolumeAnchor` measures SystemRoot's volume,
+ // which a relocated pagefile does not live on.
+ it('never reads free disk as proof the pagefile could have grown', () => {
+ setSystemMemoryInfoReaderForTest(() => FIXED_PAGEFILE_UNDER_PRESSURE)
+ const fixedPagefile = withSwapVolumeFreeSpace(
+ getSystemMemoryDetails('win32'),
+ { freeMB: 812_000, volume: 'C:' },
+ 'win32'
+ )
+ // 180 MB of commit beside 812 GB of free disk: co-timed, and still not a
+ // verdict — reading it as "the pagefile had room" is the opposite conclusion.
+ expect(fixedPagefile.systemMemoryPressureSignal).toBe('available-commit-volume-cotimed')
+
+ // The one decisive win32 case: commit limit at or below RAM means there is
+ // no pagefile behind it, so the floor cannot heal however empty the disk is.
+ setSystemMemoryInfoReaderForTest(() => NO_PAGEFILE_UNDER_PRESSURE)
+ expect(
+ withSwapVolumeFreeSpace(
+ getSystemMemoryDetails('win32'),
+ { freeMB: 812_000, volume: 'C:' },
+ 'win32'
+ ).systemMemoryPressureSignal
+ ).toBe('available-commit-hard-capped')
+ })
+
+ // Why the verdict and not just the field: a statfs issued on a healthy host at
+ // t=0 that resolves 20 s into a commit storm prints "200 MB commit, 40 GB of
+ // pagefile headroom" — which reads as NOT a commit refusal, the opposite
+ // conclusion, under the branch's most confident label.
+ it('will not let a statfs that outlived its tick qualify the win32 commit verdict', async () => {
+ const platform = Object.getOwnPropertyDescriptor(process, 'platform')!
+ Object.defineProperty(process, 'platform', { configurable: true, value: 'win32' })
+ vi.useFakeTimers()
+ let resolveVolume: (value: SwapVolumeFreeSpace) => void = () => {}
+ try {
+ setSystemMemoryInfoReaderForTest(() => BEFORE_THE_STORM)
+ setSwapVolumeFreeSpaceReaderForTest(
+ () =>
+ new Promise((resolve) => {
+ resolveVolume = resolve
+ })
+ )
+ void samplePreGoneSystemMemory(0)
+
+ // The storm arrives; the in-flight latch makes every tick skip the merge,
+ // so the pending statfs is as old as the tick that STARTED it.
+ setSystemMemoryInfoReaderForTest(() => UNDER_COMMIT_PRESSURE)
+ await samplePreGoneSystemMemory(10_000)
+ await samplePreGoneSystemMemory(20_000)
+
+ resolveVolume({ freeMB: 40_000, volume: 'C:' })
+ await vi.advanceTimersByTimeAsync(0)
+
+ vi.setSystemTime(20_000)
+ const stale = buildProcessGoneCrashDetails({}, 'renderer')
+ expect(stale.systemMemoryPreGoneSwapFreeMB).toBe(200)
+ // The pre-storm volume number still ships — but carrying its own age, and
+ // without promoting the verdict the analyst reads.
+ expect(stale.systemMemoryPreGoneSwapVolumeFreeMB).toBe(40_000)
+ expect(stale.systemMemoryPreGoneSampleAgeMs).toBe(0)
+ expect(stale.systemMemoryPreGoneSwapVolumeAgeMs).toBe(20_000)
+ expect(stale.systemMemoryPreGonePressureSignal).toBe('available-commit-unqualified')
+
+ // The next tick's statfs answers on its own tick, so it qualifies again.
+ setSwapVolumeFreeSpaceReaderForTest(() => Promise.resolve({ freeMB: 900, volume: 'C:' }))
+ await samplePreGoneSystemMemory(30_000)
+ vi.setSystemTime(30_000)
+ const fresh = buildProcessGoneCrashDetails({}, 'renderer')
+ expect(fresh.systemMemoryPreGoneSwapVolumeFreeMB).toBe(900)
+ expect(fresh.systemMemoryPreGoneSwapVolumeAgeMs).toBe(0)
+ expect(fresh.systemMemoryPreGonePressureSignal).toBe('available-commit-volume-cotimed')
+ } finally {
+ vi.useRealTimers()
+ Object.defineProperty(process, 'platform', platform)
+ }
+ })
+
+ // Round 5: sample identity alone could not see these ticks. A host read that
+ // returns nothing leaves the sample object in place, so `sample === issuedFor`
+ // still held 25 s and two ticks later and the statfs re-qualified the verdict.
+ it('will not let ticks with a failed host read pass a stale statfs off as co-timed', async () => {
+ const platform = Object.getOwnPropertyDescriptor(process, 'platform')!
+ Object.defineProperty(process, 'platform', { configurable: true, value: 'win32' })
+ vi.useFakeTimers()
+ let resolveVolume: (value: SwapVolumeFreeSpace) => void = () => {}
+ try {
+ setSystemMemoryInfoReaderForTest(() => BEFORE_THE_STORM)
+ setSwapVolumeFreeSpaceReaderForTest(
+ () =>
+ new Promise((resolve) => {
+ resolveVolume = resolve
+ })
+ )
+ void samplePreGoneSystemMemory(0)
+
+ // GlobalMemoryStatusEx starts failing: the sample is neither replaced nor erased.
+ setSystemMemoryInfoReaderForTest(() => null)
+ await samplePreGoneSystemMemory(10_000)
+ await samplePreGoneSystemMemory(20_000)
+
+ resolveVolume({ freeMB: 40_000, volume: 'C:' })
+ await vi.advanceTimersByTimeAsync(0)
+
+ vi.setSystemTime(25_000)
+ const details = buildProcessGoneCrashDetails({}, 'renderer')
+ // 25 s of lag: the label must not say co-timed beside that age.
+ expect(details.systemMemoryPreGoneSwapVolumeAgeMs).toBe(25_000)
+ expect(details.systemMemoryPreGonePressureSignal).toBe('available-commit-unqualified')
+ } finally {
+ vi.useRealTimers()
+ Object.defineProperty(process, 'platform', platform)
+ }
+ })
+
+ it("arms the host sampler on its own unref'd 10 s timer, not the metric sweep's", async () => {
+ vi.useFakeTimers()
+ vi.setSystemTime(0)
+ const readHostMemory = vi.fn(() => UNDER_COMMIT_PRESSURE)
+ setSystemMemoryInfoReaderForTest(readHostMemory)
+ setSwapVolumeFreeSpaceReaderForTest(() => Promise.resolve({ freeMB: 120, volume: 'C:' }))
+ const setIntervalSpy = vi.spyOn(globalThis, 'setInterval')
+ try {
+ startPreGoneCrashSampling()
+
+ // Literal millisecond values: asserting the constants against themselves
+ // would let a cadence regression through, and 37 s of staleness is the bug.
+ expect(setIntervalSpy.mock.calls.map(([, ms]) => ms)).toEqual([60_000, 10_000])
+ for (const { value } of setIntervalSpy.mock.results) {
+ expect((value as NodeJS.Timeout).hasRef()).toBe(false)
+ }
+ expect(readHostMemory).toHaveBeenCalledTimes(1)
+
+ readHostMemory.mockReturnValue(AFTER_THE_CORPSE_RELEASED)
+ await vi.advanceTimersByTimeAsync(10_000)
+ // One host tick, no extra metric sweep: the two samplers run independently.
+ expect(readHostMemory).toHaveBeenCalledTimes(2)
+ expect(appMetricsMock).toHaveBeenCalledTimes(1)
+
+ const details = buildProcessGoneCrashDetails({}, 'renderer')
+ expect(details.systemMemoryPreGoneSampleAgeMs).toBe(0)
+ expect(details.systemMemoryPreGoneSwapFreeMB).toBe(2_900)
+ } finally {
+ setIntervalSpy.mockRestore()
+ vi.useRealTimers()
+ }
+ })
+
+ it('commits the host reading without waiting on the swap-volume statfs', async () => {
+ // Why: statfs is slowest during the paging storm this sampler targets, and
+ // a hung volume must not stall or silently skip host sampling.
+ setSystemMemoryInfoReaderForTest(() => UNDER_COMMIT_PRESSURE)
+ setSwapVolumeFreeSpaceReaderForTest(() => new Promise(() => {}))
+
+ void samplePreGoneSystemMemory(Date.now() - 5_000)
+ expect(buildProcessGoneCrashDetails({}, 'renderer').systemMemoryPreGoneSwapFreeMB).toBe(200)
+
+ // A second tick still refreshes the reading while that statfs hangs.
+ setSystemMemoryInfoReaderForTest(() => AFTER_THE_CORPSE_RELEASED)
+ void samplePreGoneSystemMemory(Date.now())
+ expect(buildProcessGoneCrashDetails({}, 'renderer').systemMemoryPreGoneSwapFreeMB).toBe(2_900)
+ })
+
+ it('publishes no pre-gone host keys when every memory field failed to read', async () => {
+ // Why not "no keys at all": the reading always carries its signal label, so a
+ // committed empty one would ship an age and a volume number with no memory
+ // numbers beside them — a disk-free figure standing in for a host reading.
+ setSystemMemoryInfoReaderForTest(() => ({ total: Number.NaN, free: undefined }))
+ await samplePreGoneSystemMemory(Date.now())
+
+ const details = buildProcessGoneCrashDetails({}, 'renderer')
+
+ expect(Object.keys(details).filter((key) => key.startsWith('systemMemoryPreGone'))).toEqual([])
+ })
+
+ it('carries the last volume reading forward, aged, instead of dropping it', async () => {
+ vi.useFakeTimers()
+ try {
+ setSystemMemoryInfoReaderForTest(() => UNDER_COMMIT_PRESSURE)
+ setSwapVolumeFreeSpaceReaderForTest(() => Promise.resolve({ freeMB: 42, volume: 'C:' }))
+ await samplePreGoneSystemMemory(0)
+
+ // The next tick's statfs hangs — during the paging storm this targets, that
+ // is the normal case — so the tick has no volume reading of its own, and
+ // the sample that replaces the last one would otherwise drop the field.
+ setSwapVolumeFreeSpaceReaderForTest(() => new Promise(() => {}))
+ void samplePreGoneSystemMemory(10_000)
+ vi.setSystemTime(10_000)
+
+ const details = buildProcessGoneCrashDetails({}, 'renderer')
+ expect(details.systemMemoryPreGoneSwapVolumeFreeMB).toBe(42)
+ expect(details.systemMemoryPreGoneSwapVolume).toBe('C:')
+ expect(details.systemMemoryPreGoneSampleAgeMs).toBe(0)
+ // Carried, not re-read: it ships at its real age, never as a fresh number.
+ expect(details.systemMemoryPreGoneSwapVolumeAgeMs).toBe(10_000)
+ } finally {
+ vi.useRealTimers()
+ }
+ })
+
+ it('keeps a failed host read from erasing the process-metric sample', async () => {
+ samplePreGoneProcessMetrics(Date.now() - 5_000)
+ setSystemMemoryInfoReaderForTest(() => {
+ throw new Error('getSystemMemoryInfo unavailable')
+ })
+ await samplePreGoneSystemMemory(Date.now() - 5_000)
+ setSystemMemoryInfoReaderForTest(null)
+
+ const details = buildProcessGoneCrashDetails({ processType: 'renderer' }, 'renderer')
+
+ expect(details.processMetricsPreGoneRendererWorkingSetMB).toBe(400)
+ expect(Object.keys(details).filter((key) => key.startsWith('systemMemoryPreGone'))).toEqual([])
+ })
+})
diff --git a/src/main/crash-reporting/pre-gone-host-memory.ts b/src/main/crash-reporting/pre-gone-host-memory.ts
new file mode 100644
index 00000000000..0db56796750
--- /dev/null
+++ b/src/main/crash-reporting/pre-gone-host-memory.ts
@@ -0,0 +1,164 @@
+import type { CrashReportDetailValue } from '../../shared/crash-reporting'
+import { readSwapVolumeFreeSpace } from './swap-volume-free-space'
+import {
+ getSystemMemoryDetails,
+ SYSTEM_MEMORY_KEY_PREFIX,
+ withSwapVolumeFreeSpace
+} from './system-memory-details'
+
+// ─── Pre-gone host memory sampling ──────────────────────────────────
+// Why sample at all: the gone-time host read lands after the corpse released
+// its pages, so it reports a healthier machine than the one that refused the
+// allocation.
+// Why 10 s and not the 60 s process-metrics cadence: at 60 s, four of five
+// field OOMs carried a ~37 s old host reading — far too stale to see a
+// transient commit refusal. A refusal shorter than the interval stays
+// invisible; no cadence fixes that.
+
+export const PRE_GONE_SYSTEM_MEMORY_SAMPLE_INTERVAL_MS = 10_000
+
+type CrashReportDetails = Record
+
+type PreGoneSystemMemorySample = {
+ details: CrashReportDetails
+ sampledAtMs: number
+ /** Tick that ISSUED the statfs now merged in — never the tick it resolved on. */
+ swapVolumeSampledAtMs?: number
+}
+
+let preGoneSample: PreGoneSystemMemorySample | null = null
+let preGoneTimer: ReturnType | null = null
+let swapVolumeReadInFlight = false
+let samplingGeneration = 0
+let sampleTick = 0
+
+const PRESSURE_SIGNAL_KEY = `${SYSTEM_MEMORY_KEY_PREFIX}PressureSignal`
+
+/**
+ * Carries the last volume reading onto the sample that replaces its own.
+ *
+ * Why: a statfs slower than one tick would otherwise make the field vanish from
+ * the reports it exists for — the next tick replaces the sample wholesale, and
+ * the in-flight latch keeps intervening ticks from merging anything. It ships
+ * with its own (now larger) age and, not being co-timed, never names the label.
+ */
+function withCarriedSwapVolume(sample: PreGoneSystemMemorySample): PreGoneSystemMemorySample {
+ const previous = preGoneSample
+ if (!previous || previous.swapVolumeSampledAtMs === undefined) {
+ return sample
+ }
+ const freeMB = previous.details[`${SYSTEM_MEMORY_KEY_PREFIX}SwapVolumeFreeMB`]
+ const volume = previous.details[`${SYSTEM_MEMORY_KEY_PREFIX}SwapVolume`]
+ if (typeof freeMB !== 'number' || typeof volume !== 'string') {
+ return sample
+ }
+ return {
+ ...sample,
+ details: withSwapVolumeFreeSpace(sample.details, { freeMB, volume }, process.platform, false),
+ swapVolumeSampledAtMs: previous.swapVolumeSampledAtMs
+ }
+}
+
+function commitHostMemorySample(nowMs: number): boolean {
+ try {
+ const details = getSystemMemoryDetails()
+ // Why not `length === 0`: the signal label is appended unconditionally, so a
+ // reading that resolved no memory field at all still arrives with one key.
+ if (!Object.keys(details).some((key) => key !== PRESSURE_SIGNAL_KEY)) {
+ return false
+ }
+ preGoneSample = withCarriedSwapVolume({ details, sampledAtMs: nowMs })
+ return true
+ } catch {
+ // Why: a failed read must not erase the previous good sample.
+ return false
+ }
+}
+
+async function mergeSwapVolumeFreeSpace(issuedOnTick: number): Promise {
+ if (swapVolumeReadInFlight) {
+ return
+ }
+ swapVolumeReadInFlight = true
+ const generation = samplingGeneration
+ const issuedFor = preGoneSample
+ try {
+ const volume = await readSwapVolumeFreeSpace()
+ if (volume && preGoneSample && generation === samplingGeneration) {
+ // Why only its own tick qualifies: a statfs that outlived its tick carries a
+ // pre-storm volume number, and the latch makes that lag unbounded. It still
+ // ships beside its age, but it may not decide the verdict.
+ // Why the tick counter and not sample identity: a tick whose host read fails
+ // leaves the sample object in place, so identity alone reads as co-timed.
+ const coTimed = issuedOnTick === sampleTick
+ preGoneSample = {
+ ...preGoneSample,
+ details: withSwapVolumeFreeSpace(preGoneSample.details, volume, process.platform, coTimed),
+ swapVolumeSampledAtMs: issuedFor?.sampledAtMs
+ }
+ }
+ } catch {
+ // Why: the memory reading is already committed and stands on its own.
+ } finally {
+ swapVolumeReadInFlight = false
+ }
+}
+
+export async function samplePreGoneSystemMemory(nowMs: number = Date.now()): Promise {
+ // Why commit before awaiting: the volume read is a statfs, and under the very
+ // paging storm this targets it is slowest — it must never delay, or (via an
+ // in-flight latch) skip, the cheap synchronous host reading.
+ const tick = ++sampleTick
+ if (!commitHostMemorySample(nowMs)) {
+ return
+ }
+ await mergeSwapVolumeFreeSpace(tick)
+}
+
+export function startPreGoneSystemMemorySampling(
+ intervalMs: number = PRE_GONE_SYSTEM_MEMORY_SAMPLE_INTERVAL_MS
+): void {
+ if (preGoneTimer) {
+ return
+ }
+ void samplePreGoneSystemMemory()
+ preGoneTimer = setInterval(() => void samplePreGoneSystemMemory(), intervalMs)
+ preGoneTimer.unref?.()
+}
+
+export function resetPreGoneSystemMemorySamplingForTest(): void {
+ if (preGoneTimer) {
+ clearInterval(preGoneTimer)
+ }
+ preGoneTimer = null
+ preGoneSample = null
+ swapVolumeReadInFlight = false
+ // Why bump: an already-awaited volume read must not repopulate a reset sample.
+ samplingGeneration += 1
+}
+
+/** Keyed as `systemMemoryPreGone*` so a scan over the `systemMemory` family sees both reads. */
+export function preGoneSystemMemoryDetails(nowMs: number): CrashReportDetails {
+ if (!preGoneSample) {
+ return {}
+ }
+ const details: CrashReportDetails = {
+ [`${SYSTEM_MEMORY_KEY_PREFIX}PreGoneSampleAgeMs`]: Math.max(
+ 0,
+ nowMs - preGoneSample.sampledAtMs
+ )
+ }
+ // Why its own age: the volume read resolves out of band, so it can be older
+ // than the memory reading printed beside it, and that gap must be readable.
+ if (preGoneSample.swapVolumeSampledAtMs !== undefined) {
+ details[`${SYSTEM_MEMORY_KEY_PREFIX}PreGoneSwapVolumeAgeMs`] = Math.max(
+ 0,
+ nowMs - preGoneSample.swapVolumeSampledAtMs
+ )
+ }
+ for (const [key, value] of Object.entries(preGoneSample.details)) {
+ details[`${SYSTEM_MEMORY_KEY_PREFIX}PreGone${key.slice(SYSTEM_MEMORY_KEY_PREFIX.length)}`] =
+ value
+ }
+ return details
+}
diff --git a/src/main/crash-reporting/process-gone-diagnostics.test.ts b/src/main/crash-reporting/process-gone-diagnostics.test.ts
index e31645865d8..6a6ec410733 100644
--- a/src/main/crash-reporting/process-gone-diagnostics.test.ts
+++ b/src/main/crash-reporting/process-gone-diagnostics.test.ts
@@ -2,11 +2,11 @@ import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import {
buildProcessGoneCrashDetails,
collectProcessGoneMetricDetails,
- resetPreGoneProcessMetricsSamplingForTest,
+ resetPreGoneCrashSamplingForTest,
samplePreGoneProcessMetrics,
- startPreGoneProcessMetricsSampling
+ startPreGoneCrashSampling
} from './process-gone-diagnostics'
-import { setSystemMemoryInfoReaderForTest } from './gone-time-system-memory'
+import { setSystemMemoryInfoReaderForTest } from './system-memory-details'
type MetricFixture = {
pid?: number
@@ -27,7 +27,7 @@ vi.mock('electron', () => ({
describe('process gone diagnostics', () => {
beforeEach(() => {
- resetPreGoneProcessMetricsSamplingForTest()
+ resetPreGoneCrashSamplingForTest()
setSystemMemoryInfoReaderForTest(null)
})
@@ -141,8 +141,8 @@ describe('process gone diagnostics', () => {
appMetricsMock.mockReturnValue([
{ pid: 30, type: 'Tab', memory: { workingSetSize: 1024 * 100 } }
])
- startPreGoneProcessMetricsSampling(1_000)
- startPreGoneProcessMetricsSampling(1_000)
+ startPreGoneCrashSampling(1_000)
+ startPreGoneCrashSampling(1_000)
// A crash inside the first interval already has a sample to draw from.
expect(buildProcessGoneCrashDetails({}, 'renderer')).toMatchObject({
@@ -582,12 +582,12 @@ describe('process gone diagnostics', () => {
it("arms an unref'd interval so sampling never holds the event loop open", () => {
const setIntervalSpy = vi.spyOn(globalThis, 'setInterval')
try {
- startPreGoneProcessMetricsSampling(60_000)
+ startPreGoneCrashSampling(60_000)
const timer = setIntervalSpy.mock.results[0]?.value as NodeJS.Timeout
expect(timer.hasRef()).toBe(false)
} finally {
setIntervalSpy.mockRestore()
- resetPreGoneProcessMetricsSamplingForTest()
+ resetPreGoneCrashSamplingForTest()
}
})
@@ -641,7 +641,7 @@ describe('process gone diagnostics', () => {
expect(details.systemMemoryTotalMB).toBe(16_384)
})
- it('samples system memory at gone time but never into the pre-gone snapshot', () => {
+ it('samples system memory at gone time but never into the processMetrics family', () => {
appMetricsMock.mockReturnValue([{ pid: 1, type: 'Browser', memory: { workingSetSize: 0 } }])
samplePreGoneProcessMetrics()
setSystemMemoryInfoReaderForTest(() => ({
@@ -658,7 +658,9 @@ describe('process gone diagnostics', () => {
systemMemorySwapTotalMB: 8_192,
systemMemorySwapFreeMB: 40
})
- expect(details.processMetricsPreGoneSystemMemoryTotalMB).toBeUndefined()
+ expect(
+ Object.keys(details).filter((key) => key.startsWith('processMetricsPreGoneSystem'))
+ ).toEqual([])
})
it('leaves records unflagged when the crashed bucket is still populated', () => {
diff --git a/src/main/crash-reporting/process-gone-diagnostics.ts b/src/main/crash-reporting/process-gone-diagnostics.ts
index d0bb380a2b6..bf0d735a2c7 100644
--- a/src/main/crash-reporting/process-gone-diagnostics.ts
+++ b/src/main/crash-reporting/process-gone-diagnostics.ts
@@ -3,7 +3,13 @@ import {
sanitizeCrashReportDetails,
type CrashReportDetailValue
} from '../../shared/crash-reporting'
-import { getSystemMemoryAtGoneDetails, memoryKBFieldMB } from './gone-time-system-memory'
+import { getSystemMemoryDetails, memoryKBFieldMB } from './system-memory-details'
+import {
+ PRE_GONE_SYSTEM_MEMORY_SAMPLE_INTERVAL_MS,
+ preGoneSystemMemoryDetails,
+ resetPreGoneSystemMemorySamplingForTest,
+ startPreGoneSystemMemorySampling
+} from './pre-gone-host-memory'
type ProcessMetricLike = {
pid?: unknown
@@ -204,8 +210,9 @@ export function samplePreGoneProcessMetrics(nowMs: number = Date.now()): void {
}
}
-export function startPreGoneProcessMetricsSampling(
- intervalMs: number = PROCESS_METRICS_PRE_GONE_SAMPLE_INTERVAL_MS
+export function startPreGoneCrashSampling(
+ intervalMs: number = PROCESS_METRICS_PRE_GONE_SAMPLE_INTERVAL_MS,
+ systemMemoryIntervalMs: number = PRE_GONE_SYSTEM_MEMORY_SAMPLE_INTERVAL_MS
): void {
if (preGoneSampleTimer) {
return
@@ -213,14 +220,16 @@ export function startPreGoneProcessMetricsSampling(
samplePreGoneProcessMetrics()
preGoneSampleTimer = setInterval(() => samplePreGoneProcessMetrics(), intervalMs)
preGoneSampleTimer.unref?.()
+ startPreGoneSystemMemorySampling(systemMemoryIntervalMs)
}
-export function resetPreGoneProcessMetricsSamplingForTest(): void {
+export function resetPreGoneCrashSamplingForTest(): void {
if (preGoneSampleTimer) {
clearInterval(preGoneSampleTimer)
}
preGoneSampleTimer = null
preGoneSample = null
+ resetPreGoneSystemMemorySamplingForTest()
}
const PROCESS_METRICS_KEY_PREFIX = 'processMetrics'
@@ -271,7 +280,7 @@ export function buildProcessGoneCrashDetails(
const crashDetails: CrashReportDetails = {
...sanitizedDetails,
...liveMetricDetails,
- ...getSystemMemoryAtGoneDetails()
+ ...getSystemMemoryDetails()
}
// Why: with the crasher gone, Largest names a survivor — flag that so the
// live buckets are read as "everyone else", not as the crashed process.
@@ -290,8 +299,10 @@ export function buildProcessGoneCrashDetails(
if (liveMetricDetails[crashedBucketCountKey] === 0 || sampledSameBucketProcessVanished) {
crashDetails.processMetricsCrashedProcessAbsent = true
}
+ const nowMs = Date.now()
if (preGoneSample) {
- Object.assign(crashDetails, preGoneSampleDetails(preGoneSample, Date.now()))
+ Object.assign(crashDetails, preGoneSampleDetails(preGoneSample, nowMs))
}
+ Object.assign(crashDetails, preGoneSystemMemoryDetails(nowMs))
return crashDetails
}
diff --git a/src/main/crash-reporting/swap-volume-free-space.ts b/src/main/crash-reporting/swap-volume-free-space.ts
new file mode 100644
index 00000000000..3ad40b7629b
--- /dev/null
+++ b/src/main/crash-reporting/swap-volume-free-space.ts
@@ -0,0 +1,67 @@
+import { statfs } from 'node:fs/promises'
+import path from 'node:path'
+
+// Why: a system-managed Windows pagefile — and a macOS swapfile — only grows
+// into free space on its own volume, so low available commit is a REFUSED
+// allocation only when that volume is full too. Linux is excluded on purpose:
+// its swap is a fixed partition, a fixed-size swapfile, or zram, none of which
+// grow into root-fs free space, so the number would read as headroom that
+// cannot exist. The measured volume ships alongside because Windows only names
+// the DEFAULT pagefile drive; a relocated pagefile lives elsewhere.
+
+const BYTES_PER_MB = 1024 * 1024
+
+export type SwapVolumeFreeSpace = {
+ freeMB: number
+ /** Which volume was measured, separator-trimmed so redaction sees no path. */
+ volume: string
+}
+
+type SwapVolumeFreeSpaceReader = (
+ platform: NodeJS.Platform
+) => Promise
+
+function swapVolumeAnchor(platform: NodeJS.Platform): string | undefined {
+ if (platform === 'win32') {
+ const anchor = process.env.SystemRoot || process.env.SystemDrive
+ return anchor ? path.parse(anchor).root || anchor : undefined
+ }
+ return platform === 'darwin' ? path.sep : undefined
+}
+
+function volumeLabel(root: string): string {
+ const trimmed = root.replace(/[\\/]+$/, '')
+ return trimmed.length > 0 ? trimmed : root
+}
+
+async function statfsSwapVolumeFreeSpace(
+ platform: NodeJS.Platform
+): Promise {
+ const root = swapVolumeAnchor(platform)
+ if (!root) {
+ return undefined
+ }
+ try {
+ const stats = await statfs(root)
+ const bytes = Number(stats.bsize) * Number(stats.bavail)
+ return Number.isFinite(bytes)
+ ? { freeMB: Math.round(Math.max(0, bytes) / BYTES_PER_MB), volume: volumeLabel(root) }
+ : undefined
+ } catch {
+ return undefined
+ }
+}
+
+let swapVolumeFreeSpaceReader: SwapVolumeFreeSpaceReader = statfsSwapVolumeFreeSpace
+
+export function setSwapVolumeFreeSpaceReaderForTest(
+ reader: SwapVolumeFreeSpaceReader | null
+): void {
+ swapVolumeFreeSpaceReader = reader ?? statfsSwapVolumeFreeSpace
+}
+
+export function readSwapVolumeFreeSpace(
+ platform: NodeJS.Platform = process.platform
+): Promise {
+ return swapVolumeFreeSpaceReader(platform)
+}
diff --git a/src/main/crash-reporting/system-memory-details.ts b/src/main/crash-reporting/system-memory-details.ts
new file mode 100644
index 00000000000..1f2cf556faa
--- /dev/null
+++ b/src/main/crash-reporting/system-memory-details.ts
@@ -0,0 +1,161 @@
+import type { CrashReportDetailValue } from '../../shared/crash-reporting'
+import type { SwapVolumeFreeSpace } from './swap-volume-free-space'
+
+// ─── Host system memory for crash reports ───────────────────────────
+// Why: the system outlives the crashed process, so this IS sampleable at
+// process-gone — it separates "renderer grew huge" from "machine out of
+// memory/commit", which the per-process buckets alone cannot. The gone-time
+// caller reads AFTER the corpse returned its pages, so free/swapFree read
+// healthier than at kill time; the pre-gone sampler carries a live reading past
+// that.
+// Every reading is labelled `systemMemoryPressureSignal` so no report can be
+// read as a pressure verdict the platform never gave:
+// win32 — swapFree is MEMORYSTATUSEX.ullAvailPageFile, i.e. available
+// COMMIT, which pagefile growth can heal (a 127 MB commit floor healed to
+// 2029 MB mid-hold on the win-lowspec repro, killing nothing). Free space on
+// the swap volume does NOT establish that it could: a fixed-size or disabled
+// pagefile grows into no amount of empty disk, its maximum is unreadable
+// here (needs a registry read), and the measured volume is only the DEFAULT
+// pagefile drive. So a co-timed volume reading is context beside the commit
+// number — `available-commit-volume-cotimed` — never a verdict. The one
+// decisive win32 case is a commit limit at or below RAM: no pagefile exists
+// to grow, so the floor cannot heal (`available-commit-hard-capped`).
+// linux — MemAvailable is the real signal; MemFree is not (it excludes page
+// cache and other reclaimable memory).
+// darwin — none. `free` stays low on healthy machines and
+// fileBacked/purgeable are only a reclaimability proxy. The real signal
+// needs `memory_pressure -Q`; Orca's reader for it
+// (src/main/memory/host-memory.ts) is on-demand, and spawning a subprocess
+// on a 10 s app-lifetime timer costs more than the gap it closes.
+
+type CrashReportDetails = Record
+
+export const SYSTEM_MEMORY_KEY_PREFIX = 'systemMemory'
+
+export function memoryKBFieldMB(value: unknown): number | undefined {
+ const kb = typeof value === 'number' && Number.isFinite(value) ? value : undefined
+ return kb === undefined ? undefined : Math.round(Math.max(0, kb) / 1024)
+}
+
+type SystemMemoryInfoLike = {
+ total?: unknown
+ free?: unknown
+ available?: unknown
+ swapTotal?: unknown
+ swapFree?: unknown
+ fileBacked?: unknown
+ purgeable?: unknown
+}
+
+type SystemMemoryInfoReader = () => SystemMemoryInfoLike | null
+
+/** How far this reading may be read as a "was the host under pressure" verdict. */
+export type SystemMemoryPressureSignal =
+ | 'available-commit-hard-capped'
+ | 'available-commit-volume-cotimed'
+ | 'available-commit-unqualified'
+ | 'mem-available'
+ | 'none'
+
+function readElectronSystemMemoryInfo(): SystemMemoryInfoLike | null {
+ const read = (process as NodeJS.Process & { getSystemMemoryInfo?: () => SystemMemoryInfoLike })
+ .getSystemMemoryInfo
+ if (typeof read !== 'function') {
+ return null
+ }
+ try {
+ return read.call(process)
+ } catch {
+ return null
+ }
+}
+
+let systemMemoryInfoReader: SystemMemoryInfoReader = readElectronSystemMemoryInfo
+
+export function setSystemMemoryInfoReaderForTest(reader: SystemMemoryInfoReader | null): void {
+ systemMemoryInfoReader = reader ?? readElectronSystemMemoryInfo
+}
+
+function numericDetail(details: CrashReportDetails, suffix: string): number | undefined {
+ const value = details[`${SYSTEM_MEMORY_KEY_PREFIX}${suffix}`]
+ return typeof value === 'number' ? value : undefined
+}
+
+/** Windows commit limit = RAM + pagefile, so a limit at or below RAM has no pagefile behind it. */
+function pagefileBacksCommit(details: CrashReportDetails): boolean | undefined {
+ const total = numericDetail(details, 'TotalMB')
+ const swapTotal = numericDetail(details, 'SwapTotalMB')
+ return total === undefined || swapTotal === undefined ? undefined : swapTotal > total
+}
+
+function pressureSignal(
+ platform: NodeJS.Platform,
+ details: CrashReportDetails,
+ volumeCoTimed = true
+): SystemMemoryPressureSignal {
+ if (platform === 'win32' && `${SYSTEM_MEMORY_KEY_PREFIX}SwapFreeMB` in details) {
+ if (pagefileBacksCommit(details) === false) {
+ return 'available-commit-hard-capped'
+ }
+ return volumeCoTimed && `${SYSTEM_MEMORY_KEY_PREFIX}SwapVolumeFreeMB` in details
+ ? 'available-commit-volume-cotimed'
+ : 'available-commit-unqualified'
+ }
+ if (platform === 'linux' && `${SYSTEM_MEMORY_KEY_PREFIX}AvailableMB` in details) {
+ return 'mem-available'
+ }
+ return 'none'
+}
+
+export function getSystemMemoryDetails(
+ platform: NodeJS.Platform = process.platform
+): CrashReportDetails {
+ const info = systemMemoryInfoReader()
+ if (!info) {
+ return {}
+ }
+ const details: CrashReportDetails = {}
+ const fields: readonly [keyof SystemMemoryInfoLike, string][] = [
+ ['total', 'TotalMB'],
+ ['free', 'FreeMB'],
+ ['available', 'AvailableMB'],
+ ['swapTotal', 'SwapTotalMB'],
+ ['swapFree', 'SwapFreeMB'],
+ ['fileBacked', 'FileBackedMB'],
+ ['purgeable', 'PurgeableMB']
+ ]
+ for (const [field, suffix] of fields) {
+ const mb = memoryKBFieldMB(info[field])
+ if (mb !== undefined) {
+ details[`${SYSTEM_MEMORY_KEY_PREFIX}${suffix}`] = mb
+ }
+ }
+ details[`${SYSTEM_MEMORY_KEY_PREFIX}PressureSignal`] = pressureSignal(platform, details)
+ return details
+}
+
+/**
+ * Merges the statfs-derived volume datum, which needs an await and so is only
+ * reachable from the periodic sampler, and relabels the reading it sits beside.
+ *
+ * `coTimed` false means the statfs outlived the tick that issued it, so this
+ * volume number and the commit number beside it describe different moments —
+ * during a pagefile-growth storm that is exactly when they diverge, and a
+ * pre-storm 40 GB printed next to 200 MB of commit reads as "the pagefile had
+ * room", the opposite conclusion. The datum still ships (with its own age), but
+ * only a co-timed one is named in the label.
+ */
+export function withSwapVolumeFreeSpace(
+ details: CrashReportDetails,
+ volume: SwapVolumeFreeSpace,
+ platform: NodeJS.Platform = process.platform,
+ coTimed = true
+): CrashReportDetails {
+ const merged: CrashReportDetails = {
+ ...details,
+ [`${SYSTEM_MEMORY_KEY_PREFIX}SwapVolumeFreeMB`]: volume.freeMB,
+ [`${SYSTEM_MEMORY_KEY_PREFIX}SwapVolume`]: volume.volume
+ }
+ merged[`${SYSTEM_MEMORY_KEY_PREFIX}PressureSignal`] = pressureSignal(platform, merged, coTimed)
+ return merged
+}
diff --git a/src/main/startup/main-process-ready-runtime.ts b/src/main/startup/main-process-ready-runtime.ts
index 75784a4c196..26b920652df 100644
--- a/src/main/startup/main-process-ready-runtime.ts
+++ b/src/main/startup/main-process-ready-runtime.ts
@@ -9,7 +9,7 @@ import { RpcDispatcher } from '../runtime/rpc/dispatcher'
import { browserManager } from '../browser/browser-manager'
import { configureBrowserClientPageAutomationRuntime } from '../browser/browser-client-page-automation-runtime'
import { BrowserClientPageCommandError } from '../browser/browser-client-page-command-failure'
-import { startPreGoneProcessMetricsSampling } from '../crash-reporting/process-gone-diagnostics'
+import { startPreGoneCrashSampling } from '../crash-reporting/process-gone-diagnostics'
import { recordProcessGoneCrash } from './main-window-lifecycle-flags'
import { handleGpuChildCrash } from './gpu-lifecycle'
import { isGpuFallbackCrashCandidate } from '../crash-reporting/gpu-crash-fallback-decision'
@@ -130,9 +130,10 @@ export async function initializeReadyRuntimeServices(): Promise {
console.warn('[agent-hooks] failed to reconcile managed hooks on startup:', error)
)
}
- // Why: process-gone metrics only see survivors; retain a recent whole-app
- // snapshot for comparison in crash reports.
- startPreGoneProcessMetricsSampling()
+ // Why: process-gone metrics only see survivors, and the gone-time host memory
+ // read lands after the corpse released its pages; both need a live pre-gone
+ // sample to compare against in crash reports.
+ startPreGoneCrashSampling()
app.on('child-process-gone', (_event, details) => {
recordProcessGoneCrash('child', details.type, details.reason, details.exitCode ?? null, {
name: details.name,
diff --git a/src/main/startup/pre-gone-crash-sampling-wiring.test.ts b/src/main/startup/pre-gone-crash-sampling-wiring.test.ts
new file mode 100644
index 00000000000..2a8e008c0b3
--- /dev/null
+++ b/src/main/startup/pre-gone-crash-sampling-wiring.test.ts
@@ -0,0 +1,49 @@
+import { readFileSync } from 'node:fs'
+import { join } from 'node:path'
+import { describe, expect, it } from 'vitest'
+
+/**
+ * Guards the one line that arms pre-gone crash sampling.
+ *
+ * That branch is pure instrumentation, so this line is the whole of its value in
+ * the shipped app: deleting it left all 691 tests across `src/main/crash-reporting/`
+ * and `src/main/startup/` green while every crash report silently lost its only
+ * host reading taken before the dying process returned its pages.
+ *
+ * Source-level because that is the property: the sampler is armed once inside the
+ * ready-phase composition, which has no runtime seam to assert against.
+ */
+describe('pre-gone crash sampling startup wiring', () => {
+ // Why normalize: the indent anchors below are `\n`-prefixed, and nothing pins
+ // src/**/*.ts to LF, so a CRLF Windows checkout would fail them spuriously.
+ const readSource = (name: string): string =>
+ readFileSync(join(process.cwd(), 'src/main/startup', name), 'utf8').replace(/\r\n/g, '\n')
+
+ const readyRuntimeSource = readSource('main-process-ready-runtime.ts')
+ const readySource = readSource('main-process-ready.ts')
+
+ const READY_ENTRY = 'export async function initializeReadyRuntimeServices('
+ // Why the entry's body and not the file: the call satisfies a whole-file grep
+ // just as well from a sibling export nothing calls, which arms nothing.
+ const readyRuntimeEntryBody = readyRuntimeSource
+ .slice(readyRuntimeSource.indexOf(READY_ENTRY) + READY_ENTRY.length)
+ .split('\nexport ')[0]
+
+ it('arms the sampler unconditionally inside the function app readiness runs', () => {
+ expect(readyRuntimeSource).toContain(
+ "import { startPreGoneCrashSampling } from '../crash-reporting/process-gone-diagnostics'"
+ )
+ expect(readyRuntimeSource).toContain(READY_ENTRY)
+ expect(readyRuntimeEntryBody.split('startPreGoneCrashSampling()').length - 1).toBe(1)
+ // Why pin the indent: the call also matches as the body of an added
+ // `if (...)` guard, which keeps every other assertion here true while the
+ // sampler silently stops arming on most startups.
+ expect(readyRuntimeEntryBody).toContain('\n startPreGoneCrashSampling()')
+
+ // ...and that this really is the function app readiness runs.
+ expect(readySource).toContain(
+ "import { initializeReadyRuntimeServices } from './main-process-ready-runtime'"
+ )
+ expect(readySource).toContain('\n await initializeReadyRuntimeServices()')
+ })
+})
From 2ee507d744b8f8abc61f91bd0c77563ee3bb5c79 Mon Sep 17 00:00:00 2001
From: Neil <4138956+nwparker@users.noreply.github.com>
Date: Fri, 4 Sep 2026 01:22:09 -0700
Subject: [PATCH 30/49] fix(ssh): move Windows file writes off PowerShell 5.1
stdin onto sftp (#18596)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
* fix(ssh): move Windows file writes off PowerShell 5.1 stdin onto sftp
#16432 was fixed by chunking writes to 32KB, on the belief that a
`DefaultShell=cmd.exe` host caps one stdin at roughly 50KB. Re-measured on
Windows 11 26200.9168 / OpenSSH_for_Windows_10.0p2, that premise is wrong in
both directions, and the chunking does not fix the hang.
The real constraint: a read on Windows PowerShell 5.1's redirected-stdin handle
over a non-pty ssh exec can die permanently when it finds the stream momentarily
empty, taking both the remaining data and the EOF with it. It is probabilistic
per such read — not a size threshold, and not certain on the first one. Measured
by swapping the copy loop for a counting reader:
a 1.5s gap before any byte -> 0 bytes received, 6 of 6
1 byte, 1.5s gap, then 32767 -> exactly 1 byte
32768, 1.5s gap, then 32768 -> exactly 32768
a continuous 2MB -> 167936 / 270336 / 372736
Those three 2MB figures are one payload run three times under the same
conditions, which is what rules out a threshold. Independently reproduced by a
second harness where one 1.9MB counted read completed through 39 reads and
another died after 11.
A payload that fits one burst usually presents only one read that can find the
stream empty, which is why 32KB mostly works — and it still failed 15 times in
120 under load, and 1 in 40 on a quiet host. Neither rate survives the 62 execs
a 1.9MB file needs: even 2.5% compounds to about four uploads in five failing.
No chunk size helps, because the defect is per blocking read, not per byte.
Three controls on the same host, same DefaultShell, rule out both a size limit
and cmd.exe: `findstr` took 2,016,000 bytes through one exec's stdin, sftp moved
1.9MB 5/5, and PowerShell 7 took 2MB in one exec.
Windows writes now go over the sftp subsystem, whose batch script is read by
the *local* client, so no remote process reads a pipe at all. PowerShell 7 is
the fallback where sftp is unavailable, and Windows PowerShell 5.1 is last,
still bounded, and now reports the host limitation and its remedy instead of a
bare timeout.
Measured on the same host, through this code: 1.9MB x20 all succeeded,
hash-verified, median 315ms, against 0/6 before. 32KB x120 zero hangs, against
15/120.
Also:
- Stage under a unique name per attempt. An abandoned write leaves a remote
process that may still hold the staging file, and losing contact is not
evidence it died (docs/reference/ssh-execution-boundary.md), so a retry must
not reuse a name its predecessor may own. Sweep is best-effort and never
treated as proof of anything.
- Create upload directories over sftp too; the JSON mkdir batch rode the same
defective read.
- Cover makeWindowsWriteFileCommand and the publish command against the
8000-char budget, which F11 flagged as untested.
* fix(ssh): replace the staged Windows write atomically, and translate ssh -l
Three review findings, all on the failure path that the success-path
measurements say nothing about.
CodeRabbit, Critical: the publish deleted the destination before moving the
staged file onto it, so a failed move destroyed the user's existing file and
left a window where a reader saw no file at all. That is worse than the
truncated partial the staging discipline exists to prevent. Now File.Replace
(Win32 ReplaceFile, atomic), falling back to a plain Move only when the
destination is absent — and that race is safe, because a destination appearing
in between makes Move throw with the staged file preserved. The exclusive
branch already had it right: Move throwing on an existing destination is the
exclusive contract. Append stays non-atomic and now says why.
buildSshArgs can emit '-l ' for a config alias no Host block claims,
and the translator threw on it. isSftpUnavailableError read that throw as 'this
host cannot do sftp', so those hosts fell back to the defective PowerShell 5.1
path and had the refusal cached against them for 30 minutes, silently. '-l' now
maps to '-o User=', with a test for the exact argument shape buildSshArgs
produces in that case.
CodeRabbit, minor: two assertions passed on an absent observation — an
unmatched regex yields '' and every() is true of an empty list. Both now assert
the positive form first, and the same audit was applied to the three other
some()/every() assertions in the file. The temp-file test now asserts mode 0600
rather than only that the file is cleaned up.
* fix(ssh): keep a path sftp cannot spell from becoming a verdict about the host
Audit of isSftpUnavailableError, prompted by the '-l' gap having the same
shape: a per-operation condition being written into a per-host cache that
holds for 30 minutes.
It had a second instance, and this one was mine. UnsupportedSftpPathError was
classified as 'this host cannot do sftp', but it is thrown for a UNC or
relative destination and for any path sftp's batch lexer cannot quote --
including a *local* filename containing a newline, which POSIX clients allow.
One such file would have routed every later Windows write to that host down
the defective PowerShell 5.1 path for the rest of the cache window.
The host verdict is now only the errors that really are host-scoped: a refused
subsystem, a client that will not start, and an untranslatable argument list.
A path refusal falls back for that one write and leaves the cache alone, in
both the file-write and directory-creation paths.
Revert-tested. Removing the operation-scoped catch fails all three new tests,
whether or not the predicate is also widened. Widening the predicate alone
does not fail them, correctly: with the catch in place the predicate no longer
gates that path, so keeping it narrow is defence-in-depth rather than the live
mechanism.
Flag audit at the same time: -F, -o, -T, -S, -p, -i, -J, -l and -- are now the
complete set buildSshArgs can emit, and all are handled.
* fix(ssh): make the atomic publish actually run, and unroll the mkdir batch
Two runtime defects that only a real host could surface. Both were invisible
to unit tests that assert the shape of the generated command string, because
both are PowerShell rejecting an argument at execution time.
File.Replace was passed a bare $null for destinationBackupFileName. PowerShell
coerces $null to an empty string when binding a .NET string parameter, and
Replace rejects that with 'The path is not of a legal form' -- so every
create-mode publish failed. The Critical fix was inert as shipped. Now
[NullString]::Value, which is the construct that exists for this.
Measured on awin, same staging-file lock, opposite outcomes:
old publish rc=1 destination MISSING <- prior contents destroyed
new publish rc=1 destination PRESENT, sha 7f06b7e0... unchanged
control, destination present, no lock rc=0 replaced exactly
control, destination absent, no lock rc=0 Move fallback created it
End-to-end through the real uploader afterwards: 1.9MB x15 all hashes exact,
median 303ms; overwrite of an existing destination exact both times.
Separately, the PowerShell mkdir fallback could not create a tree of more than
one directory. '@($json | ConvertFrom-Json)' wraps the parsed array in another
array, so the loop variable binds to the whole thing and [string] of it is the
paths joined by spaces. It only ever worked for a one-element batch, where
stringifying a single-element array happens to yield the element -- which is
why no existing test caught it. Pre-existing on main; fixed here because this
PR puts that command on the fallback tier and claims the ladder works.
Both tiers now verified live against a three-directory tree.
---
src/main/ssh/ssh-remote-powershell.ts | 20 +-
...-remote-windows-command-line-limit.test.ts | 44 ++
src/main/ssh/ssh-system-fallback.test.ts | 85 ++-
.../ssh/system-ssh-file-binary-transfer.ts | 236 ++-----
src/main/ssh/system-ssh-file-transfer.ts | 80 ++-
src/main/ssh/system-ssh-sftp-args.test.ts | 141 ++++
src/main/ssh/system-ssh-sftp-args.ts | 95 +++
src/main/ssh/system-ssh-sftp-path.test.ts | 59 ++
src/main/ssh/system-ssh-sftp-path.ts | 46 ++
src/main/ssh/system-ssh-sftp-transfer.ts | 191 +++++
src/main/ssh/system-ssh-windows-file-write.ts | 138 ++++
.../ssh/system-ssh-windows-upload.test.ts | 659 ++++++++++++++----
...tem-ssh-windows-write-capabilities.test.ts | 79 +++
.../system-ssh-windows-write-capabilities.ts | 52 ++
.../ssh/system-ssh-windows-write-strategy.ts | 329 +++++++++
15 files changed, 1900 insertions(+), 354 deletions(-)
create mode 100644 src/main/ssh/system-ssh-sftp-args.test.ts
create mode 100644 src/main/ssh/system-ssh-sftp-args.ts
create mode 100644 src/main/ssh/system-ssh-sftp-path.test.ts
create mode 100644 src/main/ssh/system-ssh-sftp-path.ts
create mode 100644 src/main/ssh/system-ssh-sftp-transfer.ts
create mode 100644 src/main/ssh/system-ssh-windows-file-write.ts
create mode 100644 src/main/ssh/system-ssh-windows-write-capabilities.test.ts
create mode 100644 src/main/ssh/system-ssh-windows-write-capabilities.ts
create mode 100644 src/main/ssh/system-ssh-windows-write-strategy.ts
diff --git a/src/main/ssh/ssh-remote-powershell.ts b/src/main/ssh/ssh-remote-powershell.ts
index 8c94fd3c483..420223ced29 100644
--- a/src/main/ssh/ssh-remote-powershell.ts
+++ b/src/main/ssh/ssh-remote-powershell.ts
@@ -11,14 +11,24 @@ export {
// to leave room for the `/c` wrapper sshd adds before cmd.exe counts the line.
const WINDOWS_REMOTE_COMMAND_LINE_BUDGET_CHARS = 8_000
-export function powerShellCommand(script: string): string {
- const inline = encodedPowerShellCommand(script)
+/**
+ * `pwsh.exe` is PowerShell 7. It is not present on a stock Windows install, so it is only ever
+ * chosen after a probe — but where it exists it reads a redirected stdin correctly, which Windows
+ * PowerShell 5.1 does not (see `system-ssh-file-binary-transfer.ts`).
+ */
+export type WindowsPowerShellExecutable = 'powershell.exe' | 'pwsh.exe'
+
+export function powerShellCommand(
+ script: string,
+ executable: WindowsPowerShellExecutable = 'powershell.exe'
+): string {
+ const inline = encodedPowerShellCommand(script, executable)
if (inline.length <= WINDOWS_REMOTE_COMMAND_LINE_BUDGET_CHARS) {
return inline
}
// Why: these scripts are repetitive enough that gzip beats the UTF-16LE tax by
// ~4x, which is the difference between a line cmd.exe runs and one it refuses.
- const compressed = encodedPowerShellCommand(selfExtractingPowerShellScript(script))
+ const compressed = encodedPowerShellCommand(selfExtractingPowerShellScript(script), executable)
if (compressed.length > WINDOWS_REMOTE_COMMAND_LINE_BUDGET_CHARS) {
throw new Error(
`Remote Windows command needs ${compressed.length} characters; Orca budgets ${WINDOWS_REMOTE_COMMAND_LINE_BUDGET_CHARS} for a line sshd hands to cmd.exe, which itself refuses more than ${CMD_EXE_COMMAND_LINE_MAX_CHARS}.`
@@ -27,8 +37,8 @@ export function powerShellCommand(script: string): string {
return compressed
}
-function encodedPowerShellCommand(script: string): string {
- return `powershell.exe -NoProfile -NonInteractive -ExecutionPolicy Bypass -EncodedCommand ${encodePowerShellCommand(script)}`
+function encodedPowerShellCommand(script: string, executable: WindowsPowerShellExecutable): string {
+ return `${executable} -NoProfile -NonInteractive -ExecutionPolicy Bypass -EncodedCommand ${encodePowerShellCommand(script)}`
}
/** Orca-prefixed names so the payload can never shadow the bootstrap's own state. */
diff --git a/src/main/ssh/ssh-remote-windows-command-line-limit.test.ts b/src/main/ssh/ssh-remote-windows-command-line-limit.test.ts
index c104078238e..fa34a8fbda7 100644
--- a/src/main/ssh/ssh-remote-windows-command-line-limit.test.ts
+++ b/src/main/ssh/ssh-remote-windows-command-line-limit.test.ts
@@ -4,6 +4,10 @@ import { CMD_EXE_COMMAND_LINE_MAX_CHARS } from '../providers/windows-shell-args'
import { getRemoteHostPlatform } from './ssh-remote-platform'
import { tryStealInstallLockCommand } from './ssh-relay-install-lock-commands'
import { decodeRemotePowerShellScript, powerShellCommand } from './ssh-remote-powershell'
+import {
+ makeWindowsPublishStagedFileCommand,
+ makeWindowsWriteFileCommand
+} from './system-ssh-windows-file-write'
import {
cleanupOwnedRelayUploadStageCommand,
promoteOwnedRelayUploadStageCommand,
@@ -38,6 +42,17 @@ describe('Windows remote command line limit', () => {
[
'steal stale install lock',
tryStealInstallLockCommand(windows, 'C:\\Users\\orca\\.orca-remote\\relay', 1_200)
+ ],
+ // F11 flagged these two as uncovered. They carry one path literal each, so they are the file
+ // commands whose length a caller can actually move.
+ ['write file', makeWindowsWriteFileCommand('C:\\Users\\orca\\.orca-remote\\relay.js')],
+ [
+ 'publish staged file',
+ makeWindowsPublishStagedFileCommand(
+ 'C:\\Users\\orca\\.orca-remote\\relay.js.orca-partial-0123456789ab',
+ 'C:\\Users\\orca\\.orca-remote\\relay.js',
+ 'create'
+ )
]
])('keeps the %s command inside what sshd\u2019s cmd.exe accepts', (_name, command) => {
expect(command.length).toBeLessThanOrEqual(CMD_EXE_COMMAND_LINE_MAX_CHARS)
@@ -76,3 +91,32 @@ describe('Windows remote command line limit', () => {
)
})
})
+
+/**
+ * F11 asked whether a pathological path could reach the budget, and what happens if it does.
+ * Measured: the inline encoding crosses 8000 at roughly 2500 high-entropy path characters — an
+ * order of magnitude past what Windows itself accepts — and the failure is a throw before any ssh
+ * is spawned, never a hang.
+ */
+describe('Windows file command budget headroom', () => {
+ it('absorbs a path far longer than Windows will accept', () => {
+ const deep = `C:\\Users\\orca\\${'segment\\'.repeat(30)}relay.js`
+
+ expect(deep.length).toBeGreaterThan(260)
+ expect(makeWindowsWriteFileCommand(deep).length).toBeLessThanOrEqual(
+ CMD_EXE_COMMAND_LINE_MAX_CHARS
+ )
+ })
+
+ it('throws rather than spawning a line cmd.exe would refuse', () => {
+ // Random segments so gzip cannot rescue it, which is the only way to reach the ceiling at all.
+ const incompressible = Array.from(
+ { length: 400 },
+ (_unused, index) => `${index}-${Math.random().toString(36).slice(2)}`
+ ).join('\\')
+
+ expect(() => makeWindowsWriteFileCommand(`C:\\${incompressible}\\f.bin`)).toThrow(
+ /Orca budgets 8000/
+ )
+ })
+})
diff --git a/src/main/ssh/ssh-system-fallback.test.ts b/src/main/ssh/ssh-system-fallback.test.ts
index c366e899bf6..b477ad682ef 100644
--- a/src/main/ssh/ssh-system-fallback.test.ts
+++ b/src/main/ssh/ssh-system-fallback.test.ts
@@ -709,9 +709,13 @@ describe('spawnSystemSsh', () => {
expect(args[standaloneControlIdx + 1]).toBe('none')
})
- it('writes files to Windows system SSH targets with PowerShell stdin bytes', async () => {
- const proc = createEventedProcess()
- spawnMock.mockImplementation(() => closeOnceSpawned(proc))
+ it('sends Windows file writes over sftp, not through a remote PowerShell stdin', async () => {
+ const spawned: EventedProcess[] = []
+ spawnMock.mockImplementation(() => {
+ const proc = createEventedProcess()
+ spawned.push(proc)
+ return closeOnceSpawned(proc)
+ })
const hostPlatform = getRemoteHostPlatform('win32-x64')
const promise = writeFileViaSystemSsh(
@@ -722,16 +726,24 @@ describe('spawnSystemSsh', () => {
)
await expect(promise).resolves.toBeUndefined()
- const args = spawnMock.mock.calls[0][1] as string[]
- const remoteCommand = args.at(-1) ?? ''
- expect(remoteCommand).toContain('powershell.exe')
- expect(remoteCommand).not.toContain('/bin/sh')
- expect(proc.stdin.end).toHaveBeenCalledWith(Buffer.from('0.1.0', 'utf-8'))
+ // #16432, re-measured: Windows PowerShell 5.1 can lose a redirected stdin for good when a read
+ // finds it momentarily empty, so the bytes must not travel that way at all.
+ const batch = String(spawned[0]!.stdin.end.mock.calls[0]?.[0] ?? '')
+ expect(batch).toContain('put ')
+ expect(batch).toContain('/C:/Users/me/.orca-remote/relay/.version.orca-partial-')
+ const sftpArgs = spawnMock.mock.calls[0][1] as string[]
+ expect(sftpArgs).toContain('-b')
+ // The rename that publishes it reads the staged file, never a pipe.
+ const publish = (spawnMock.mock.calls[1][1] as string[]).at(-1) ?? ''
+ expect(publish).toContain('powershell.exe')
+ expect(decodePowerShellCommand(publish)).toContain(
+ '[System.IO.File]::Replace($staging, $path, [NullString]::Value)'
+ )
+ expect(publish).not.toContain('/bin/sh')
})
- it('writes binary buffers to Windows system SSH targets with CreateNew mode', async () => {
- const proc = createEventedProcess()
- spawnMock.mockImplementation(() => closeOnceSpawned(proc))
+ it('enforces an exclusive Windows buffer write at the rename, where it is atomic', async () => {
+ spawnMock.mockImplementation(() => closeOnceSpawned(createEventedProcess()))
const hostPlatform = getRemoteHostPlatform('win32-x64')
const promise = writeBufferViaSystemSsh(
@@ -742,12 +754,11 @@ describe('spawnSystemSsh', () => {
)
await expect(promise).resolves.toBeUndefined()
- const args = spawnMock.mock.calls[0][1] as string[]
- const remoteCommand = args.at(-1) ?? ''
- expect(remoteCommand).toContain('powershell.exe')
- expect(decodePowerShellCommand(remoteCommand)).toContain('CreateNew')
- expect(remoteCommand).not.toContain('/bin/sh')
- expect(proc.stdin.end).toHaveBeenCalledWith(Buffer.from('png'))
+ const publish = decodePowerShellCommand((spawnMock.mock.calls[1][1] as string[]).at(-1) ?? '')
+ // `File::Move` raising on an existing destination is what carries the exclusive contract now;
+ // a `CreateNew` on the staged file would only refuse a leftover of our own.
+ expect(publish).toContain('[System.IO.File]::Move($staging, $path)')
+ expect(publish).not.toContain('[System.IO.File]::Delete($path)')
})
it('downloads files from Windows system SSH targets with PowerShell stdout bytes', async () => {
@@ -779,8 +790,7 @@ describe('spawnSystemSsh', () => {
})
it('forces standalone SSH for Windows file writes when requested', async () => {
- const proc = createEventedProcess()
- spawnMock.mockImplementation(() => closeOnceSpawned(proc))
+ spawnMock.mockImplementation(() => closeOnceSpawned(createEventedProcess()))
const hostPlatform = getRemoteHostPlatform('win32-x64')
const promise = writeFileViaSystemSsh(
@@ -791,10 +801,14 @@ describe('spawnSystemSsh', () => {
)
await expect(promise).resolves.toBeUndefined()
- const args = spawnMock.mock.calls[0][1] as string[]
- const standaloneControlIdx = args.indexOf('-S')
+ const sftpArgs = spawnMock.mock.calls[0][1] as string[]
+ // sftp's own `-S` names a program to run, so the same request has to be spelled as an option.
+ expect(sftpArgs).not.toContain('-S')
+ expect(sftpArgs).toContain('ControlPath=none')
+ const publishArgs = spawnMock.mock.calls[1][1] as string[]
+ const standaloneControlIdx = publishArgs.indexOf('-S')
expect(standaloneControlIdx).toBeGreaterThan(-1)
- expect(args[standaloneControlIdx + 1]).toBe('none')
+ expect(publishArgs[standaloneControlIdx + 1]).toBe('none')
})
it('uploads a Windows directory as a mkdir batch plus per-file writes, never one blob', async () => {
@@ -819,19 +833,20 @@ describe('spawnSystemSsh', () => {
rmSync(localDir, { recursive: true, force: true })
}
+ // #16432: directories first, then the file — but both over sftp now, so the only PowerShell
+ // left is the rename that publishes the staged file, which reads a file rather than a pipe.
+ const mkdirBatch = String(spawned[0]!.stdin.end.mock.calls[0]?.[0] ?? '')
+ expect(mkdirBatch).toBe('-mkdir "/C:/Users/me/.orca-remote/relay"\n')
+ const putBatch = String(spawned[1]!.stdin.end.mock.calls[0]?.[0] ?? '')
+ expect(putBatch).toContain('put ')
+ expect(putBatch).toContain('/C:/Users/me/.orca-remote/relay/relay.js.orca-partial-')
const commands = spawnMock.mock.calls.map((call) => (call[1] as string[]).at(-1) ?? '')
- // #16432: directories first (metadata only), then the file bytes on their own stdin. One batch
- // meant base64-ing the whole bundle into a single PowerShell string, which the remote never read.
- expect(commands).toHaveLength(2)
- expect(commands.every((command) => command.includes('powershell.exe'))).toBe(true)
expect(commands.every((command) => !command.includes('/bin/sh'))).toBe(true)
expect(commands.join('\n')).not.toContain('tar -xzf')
- expect(JSON.parse(spawned[0].stdin.end.mock.calls[0]?.[0] as string)).toEqual([
- 'C:/Users/me/.orca-remote/relay'
- ])
- expect(Buffer.from(spawned[1].stdin.end.mock.calls[0]?.[0] as Buffer).toString('utf-8')).toBe(
- 'console.log("relay")'
- )
+ // Nothing base64s the bundle into one PowerShell string any more, and nothing reads one.
+ expect(
+ commands.some((command) => decodePowerShellCommand(command).includes('OpenStandardInput'))
+ ).toBe(false)
})
it('forces standalone SSH for Windows upload packages when requested', async () => {
@@ -855,9 +870,9 @@ describe('spawnSystemSsh', () => {
}
const args = spawnMock.mock.calls[0][1] as string[]
- const standaloneControlIdx = args.indexOf('-S')
- expect(standaloneControlIdx).toBeGreaterThan(-1)
- expect(args[standaloneControlIdx + 1]).toBe('none')
+ // The first spawn is the sftp client, whose own `-S` names a program to run.
+ expect(args).not.toContain('-S')
+ expect(args).toContain('ControlPath=none')
})
it('throws when no system ssh is found', () => {
diff --git a/src/main/ssh/system-ssh-file-binary-transfer.ts b/src/main/ssh/system-ssh-file-binary-transfer.ts
index b0c5b662ed1..149d10dfbd3 100644
--- a/src/main/ssh/system-ssh-file-binary-transfer.ts
+++ b/src/main/ssh/system-ssh-file-binary-transfer.ts
@@ -1,5 +1,7 @@
import { constants, createWriteStream } from 'node:fs'
-import { lstat, open } from 'node:fs/promises'
+import { lstat, mkdtemp, open, rm, writeFile } from 'node:fs/promises'
+import { tmpdir } from 'node:os'
+import { join } from 'node:path'
import type { Writable } from 'node:stream'
import { pipeline } from 'node:stream/promises'
import type { SshTarget } from '../../shared/ssh-types'
@@ -16,6 +18,16 @@ import {
throwIfAborted,
waitForChannelClose
} from './system-ssh-operation-lifecycle'
+import {
+ writeWindowsRemoteFile,
+ type WindowsWriteSource
+} from './system-ssh-windows-write-strategy'
+
+export {
+ WINDOWS_STDIN_WRITE_CHUNK_BYTES,
+ WINDOWS_STDIN_WRITE_TIMEOUT_MS
+} from './system-ssh-windows-write-strategy'
+export { WINDOWS_STAGED_WRITE_SUFFIX } from './system-ssh-windows-file-write'
type SystemSshOperationOptions = SystemSshBuildArgsOptions & {
signal?: AbortSignal
@@ -74,13 +86,16 @@ export async function writeBufferViaSystemSsh(
): Promise {
throwIfAborted(options?.signal)
if (options?.hostPlatform && isWindowsRemoteHost(options.hostPlatform)) {
- await writeWindowsBytesViaSystemSsh(
+ await writeWindowsRemoteFile(
target,
remotePath,
- contents.length,
- (offset, maxBytes) =>
- Promise.resolve(contents.subarray(offset, Math.min(offset + maxBytes, contents.length))),
- options
+ {
+ totalBytes: contents.length,
+ readChunk: (offset, maxBytes) =>
+ Promise.resolve(contents.subarray(offset, Math.min(offset + maxBytes, contents.length))),
+ withLocalFile: (send) => withTemporaryLocalFile(contents, send)
+ },
+ options ?? {}
)
return
}
@@ -127,20 +142,19 @@ export async function uploadFileViaSystemSsh(
throwIfAborted(options?.signal)
if (options?.hostPlatform && isWindowsRemoteHost(options.hostPlatform)) {
- // #16432: a Windows host cannot take a whole file through one stdin, however the local side
- // paces it — see WINDOWS_STDIN_WRITE_CHUNK_BYTES. This is the path that carries the large
- // files, so it is the one that has to be chunked and bounded.
- await writeWindowsBytesViaSystemSsh(
- target,
- remotePath,
- openedStat.size,
- async (offset, maxBytes) => {
+ // This is the path that carries the large files, so it is the one the transport choice is
+ // made for; see the #16432 note below.
+ const source: WindowsWriteSource = {
+ totalBytes: openedStat.size,
+ readChunk: async (offset, maxBytes) => {
const buffer = Buffer.allocUnsafe(Math.min(maxBytes, openedStat.size - offset))
const { bytesRead } = await handle.read(buffer, 0, buffer.length, offset)
return buffer.subarray(0, bytesRead)
},
- options
- )
+ // The verified local file is already exactly the payload, so sftp sends it as is.
+ withLocalFile: (send) => send(localPath)
+ }
+ await writeWindowsRemoteFile(target, remotePath, source, options ?? {})
return
}
@@ -173,158 +187,56 @@ export async function uploadFileViaSystemSsh(
}
/**
- * #16432: Windows PowerShell 5.1 stops draining a redirected stdin over a non-pty ssh exec
- * somewhere between 50KB and 1MB, depending on the host's `DefaultShell`, and it hangs rather than
- * failing. The reporter measured that on both constructs he tried — `[Console]::In.ReadToEnd()` and
- * `new IO.StreamReader([Console]::OpenStandardInput())`, the latter reading incrementally, which is
- * why the limit cannot be attributed to materializing the payload. `Stream.CopyTo` reads the same
- * `[Console]::OpenStandardInput()` object with the same incremental `Read` loop, so nothing in it
- * escapes that limit either: no single write may exceed what one stdin is known to carry.
+ * #16432, re-measured: the constraint is not a size limit, and it is not cmd.exe's.
*
- * 32KB is an order of magnitude under the low end of the measured range, and under 50KB, which the
- * reporter measured succeeding against a stream reader on the worse of the two `DefaultShell`
- * settings.
- */
-export const WINDOWS_STDIN_WRITE_CHUNK_BYTES = 32 * 1024
-
-/** No Windows stdin write should ever outlive this; a wedged PowerShell never closes on its own. */
-export const WINDOWS_STDIN_WRITE_TIMEOUT_MS = 60_000
-
-/** Suffix for the path a multi-exec Windows write lands on before it is published by rename. */
-export const WINDOWS_STAGED_WRITE_SUFFIX = '.orca-partial'
-
-/**
- * Splits one logical Windows write into stdin-sized execs.
+ * A read on Windows PowerShell 5.1's redirected-stdin handle over a non-pty ssh exec can die
+ * permanently when it finds the stream momentarily empty: no further bytes arrive, and no EOF ever
+ * does. It is probabilistic per such read — not a size threshold, and not certain on the first one.
+ * Measured on Windows 11 26200.9168 / OpenSSH_for_Windows_10.0p2 with `DefaultShell = cmd.exe`, by
+ * replacing the copy loop with a counting reader:
*
- * A write that needs more than one exec cannot land on the destination directly: a chunk failing
- * mid-file would leave a truncated artifact under the real name with nothing marking it incomplete,
- * and the retry would then meet its own leftovers — under `exclusive` the retry's `CreateNew` fails
- * on them. Multi-exec creates therefore land on a staging path and are published by a rename, which
- * is also where `exclusive` is enforced: once, at the destination, instead of smeared across the
- * first chunk. A caller-requested append cannot be staged without reading the remote file back, so
- * it keeps writing straight through, as its own protocol already implies.
+ * - a 1.5s gap before any byte, which forces the first read to find nothing -> 0 bytes, 6 of 6
+ * - one byte, a 1.5s gap, then 32767 more -> exactly 1 byte, then nothing
+ * - 32768, a 1.5s gap, then 32768 more -> exactly 32768, then nothing
+ * - a continuous 2MB -> 167936 / 270336 / 372736, then nothing
+ *
+ * Those three 2MB death points are one payload run three times under the same conditions, which is
+ * what rules out a threshold: a stream that died at a fixed point would not vary by 2x. Independently reproduced by
+ * a second harness, where one 1.9MB counted read survived 39 reads to completion and another died
+ * after 11 — same construct, same payload.
+ *
+ * A payload small enough to arrive in one burst usually presents only one read that can find the
+ * stream empty (the one waiting for EOF), which is why 32KB mostly works: it still failed 15 times
+ * in 120 with the host under load, and 1 in 40 on a quiet one. Neither rate is survivable across
+ * the 62 execs a 1.9MB file needs — even 2.5% compounds to roughly four uploads in five failing —
+ * and no chunk size helps, because the client does not control whether its bytes arrive together.
+ *
+ * The same host, same `DefaultShell`, same connection pattern contradicts every size-limit reading:
+ * `findstr` took 2,016,000 bytes through one exec's stdin, and PowerShell 7 took 2MB. So cmd.exe is
+ * not the ceiling and neither is ~50KB. Writes now go over sftp, which moves the whole payload
+ * without any remote process reading a pipe; see `system-ssh-windows-write-strategy.ts` for the
+ * fallback order.
+ *
+ * Successes are never partial. Across every run in both harnesses a failed write hung; not one
+ * produced a short file, so this defect cannot silently truncate an upload.
*/
-async function writeWindowsBytesViaSystemSsh(
- target: SshTarget,
- remotePath: string,
- totalBytes: number,
- readChunk: (offset: number, maxBytes: number) => Promise,
- options: SystemSshWriteBufferOptions
-): Promise {
- throwIfAborted(options.signal)
- const staged = !options.append && totalBytes > WINDOWS_STDIN_WRITE_CHUNK_BYTES
- const writePath = staged ? `${remotePath}${WINDOWS_STAGED_WRITE_SUFFIX}` : remotePath
- let offset = 0
- // An empty write still has to run: it is what creates (or truncates) the file.
- do {
- const chunk = await readChunk(offset, WINDOWS_STDIN_WRITE_CHUNK_BYTES)
- if (chunk.length === 0 && offset < totalBytes) {
- throw new Error(`Source ran short during upload of ${remotePath}`)
- }
- await writeWindowsChunkViaSystemSsh(
- target,
- writePath,
- chunk,
- {
- ...options,
- append: staged ? offset > 0 : options.append === true || offset > 0,
- exclusive: staged ? false : options.exclusive === true && offset === 0
- },
- offset
- )
- offset += chunk.length
- } while (offset < totalBytes)
- if (staged) {
- await publishWindowsStagedWrite(target, writePath, remotePath, options)
+
+/** A staged write is materialized locally first when the source is a buffer rather than a file. */
+async function withTemporaryLocalFile(
+ contents: Buffer,
+ send: (localPath: string) => Promise
+): Promise {
+ const directory = await mkdtemp(join(tmpdir(), 'orca-win-upload-'))
+ const localPath = join(directory, 'payload.bin')
+ try {
+ // 0600: the payload can be repository content, and tmpdir is shared on every platform.
+ await writeFile(localPath, contents, { mode: 0o600 })
+ return await send(localPath)
+ } finally {
+ await rm(directory, { recursive: true, force: true }).catch(() => {})
}
}
-async function writeWindowsChunkViaSystemSsh(
- target: SshTarget,
- remotePath: string,
- chunk: Buffer,
- options: SystemSshWriteBufferOptions,
- offset: number
-): Promise {
- throwIfAborted(options.signal)
- const channel = spawnSystemSshCommand(target, makeWindowsWriteFileCommand(remotePath, options), {
- wrapCommand: false,
- ...getSystemSshBuildArgsFromOperationOptions(options)
- })
- const closePromise = awaitWithSystemSshAbort(
- options.signal,
- () => channel.close(),
- waitForChannelClose(
- channel,
- `write ${remotePath} at offset ${offset}`,
- WINDOWS_STDIN_WRITE_TIMEOUT_MS
- )
- )
- if (!options.signal?.aborted) {
- channel.stdin.end(chunk)
- }
- await closePromise
-}
-
-async function publishWindowsStagedWrite(
- target: SshTarget,
- stagingPath: string,
- remotePath: string,
- options: SystemSshWriteBufferOptions
-): Promise {
- throwIfAborted(options.signal)
- const channel = spawnSystemSshCommand(
- target,
- makeWindowsPublishStagedFileCommand(stagingPath, remotePath, options.exclusive === true),
- { wrapCommand: false, ...getSystemSshBuildArgsFromOperationOptions(options) }
- )
- const closePromise = awaitWithSystemSshAbort(
- options.signal,
- () => channel.close(),
- waitForChannelClose(channel, `publish ${remotePath}`, WINDOWS_STDIN_WRITE_TIMEOUT_MS)
- )
- if (!options.signal?.aborted) {
- channel.stdin.end()
- }
- await closePromise
-}
-
-function makeWindowsWriteFileCommand(
- remotePath: string,
- options?: { append?: boolean; exclusive?: boolean }
-): string {
- const fileMode = options?.append ? 'Append' : options?.exclusive ? 'CreateNew' : 'Create'
- return powerShellCommand(
- [
- '$ErrorActionPreference = "Stop"',
- `$path = ${powerShellLiteral(remotePath)}`,
- '$parent = [System.IO.Path]::GetDirectoryName($path)',
- 'if ($parent) { $null = [System.IO.Directory]::CreateDirectory($parent) }',
- '$inputStream = [Console]::OpenStandardInput()',
- `$outputStream = [System.IO.File]::Open($path, [System.IO.FileMode]::${fileMode}, [System.IO.FileAccess]::Write, [System.IO.FileShare]::None)`,
- 'try { $inputStream.CopyTo($outputStream) } finally { $outputStream.Dispose() }'
- ].join('; ')
- )
-}
-
-// `File::Move` throws when the destination exists, which is exactly the exclusive contract; the
-// non-exclusive caller asked to replace, so it deletes first (a no-op on an absent path).
-function makeWindowsPublishStagedFileCommand(
- stagingPath: string,
- remotePath: string,
- exclusive: boolean
-): string {
- return powerShellCommand(
- [
- '$ErrorActionPreference = "Stop"',
- `$staging = ${powerShellLiteral(stagingPath)}`,
- `$path = ${powerShellLiteral(remotePath)}`,
- ...(exclusive ? [] : ['[System.IO.File]::Delete($path)']),
- '[System.IO.File]::Move($staging, $path)'
- ].join('; ')
- )
-}
-
function makePosixWriteFileCommand(
remotePath: string,
options?: { append?: boolean; exclusive?: boolean }
diff --git a/src/main/ssh/system-ssh-file-transfer.ts b/src/main/ssh/system-ssh-file-transfer.ts
index f728c0eab02..d5757904356 100644
--- a/src/main/ssh/system-ssh-file-transfer.ts
+++ b/src/main/ssh/system-ssh-file-transfer.ts
@@ -27,6 +27,12 @@ import {
WINDOWS_STDIN_WRITE_TIMEOUT_MS,
writeBufferViaSystemSsh
} from './system-ssh-file-binary-transfer'
+import {
+ isSftpPathUnsupportedError,
+ isSftpUnavailableError,
+ makeDirectoriesViaSftp
+} from './system-ssh-sftp-transfer'
+import { getWindowsRemoteWriteCapabilities } from './system-ssh-windows-write-capabilities'
type SystemSshOperationOptions = SystemSshBuildArgsOptions & {
signal?: AbortSignal
@@ -161,9 +167,14 @@ async function collectWindowsUploadPlan(
return plan
}
-// Why the JSON envelope survives here: a path list is metadata, so this payload stays in the
-// hundreds of bytes even for a deep tree. Batched anyway, so a pathological tree cannot walk back
-// into the same stdin size that wedges PowerShell.
+/**
+ * Creates the upload's directories, preferring sftp's own `mkdir`.
+ *
+ * The PowerShell fallback keeps the JSON envelope, batched under one stdin's worth: a path list is
+ * metadata, so it stays in the hundreds of bytes even for a deep tree. It is still a redirected
+ * stdin read though, so on Windows PowerShell 5.1 it carries the same defect as any other — which
+ * is why sftp is tried first even for a payload this small.
+ */
async function createWindowsUploadDirectories(
target: SshTarget,
directories: readonly string[],
@@ -175,23 +186,27 @@ async function createWindowsUploadDirectories(
if (batch.length === 0) {
return
}
+ const pending = batch
const payload = JSON.stringify(batch)
batch = []
batchBytes = 0
throwIfAborted(options.signal)
- const channel = spawnSystemSshCommand(target, makeWindowsCreateDirectoriesCommand(), {
- wrapCommand: false,
- ...getSystemSshBuildArgsFromOperationOptions(options)
- })
- const closePromise = awaitWithSystemSshAbort(
- options.signal,
- () => channel.close(),
- waitForChannelClose(channel, 'windows relay upload mkdir', WINDOWS_STDIN_WRITE_TIMEOUT_MS)
+ await getWindowsRemoteWriteCapabilities(target).runWithFallback(
+ 'sftp-subsystem',
+ async () => {
+ try {
+ await makeDirectoriesViaSftp(target, pending, options)
+ } catch (error) {
+ // A directory sftp cannot address is this batch's problem, not the host's verdict.
+ if (!isSftpPathUnsupportedError(error)) {
+ throw error
+ }
+ await createWindowsUploadDirectoriesViaPowerShell(target, payload, options)
+ }
+ },
+ () => createWindowsUploadDirectoriesViaPowerShell(target, payload, options),
+ isSftpUnavailableError
)
- if (!options.signal?.aborted) {
- channel.stdin.end(payload)
- }
- await closePromise
}
for (const directory of directories) {
const entryBytes = Buffer.byteLength(directory) + 4
@@ -204,17 +219,44 @@ async function createWindowsUploadDirectories(
await flush()
}
+async function createWindowsUploadDirectoriesViaPowerShell(
+ target: SshTarget,
+ payload: string,
+ options: SystemSshOperationOptions
+): Promise {
+ const channel = spawnSystemSshCommand(target, makeWindowsCreateDirectoriesCommand(), {
+ wrapCommand: false,
+ ...getSystemSshBuildArgsFromOperationOptions(options)
+ })
+ const closePromise = awaitWithSystemSshAbort(
+ options.signal,
+ () => channel.close(),
+ waitForChannelClose(channel, 'windows relay upload mkdir', WINDOWS_STDIN_WRITE_TIMEOUT_MS)
+ )
+ if (!options.signal?.aborted) {
+ channel.stdin.end(payload)
+ }
+ await closePromise
+}
+
function makeWindowsCreateDirectoriesCommand(): string {
return powerShellCommand(
[
'$ErrorActionPreference = "Stop"',
- // The reporter measured this reader surviving 50KB where `[Console]::In` wedged at the same
- // size (#16432); the batch above stays under that.
+ // Reached only where the host has no sftp subsystem. Windows PowerShell 5.1 can lose a
+ // redirected stdin for good when a read finds it empty (#16432); a batch this small usually
+ // arrives in one piece, and "usually" is exactly why sftp is preferred.
'$reader = New-Object System.IO.StreamReader([Console]::OpenStandardInput())',
'try { $json = $reader.ReadToEnd() } finally { $reader.Dispose() }',
'if ([string]::IsNullOrWhiteSpace($json)) { return }',
- 'foreach ($path in @($json | ConvertFrom-Json)) {',
- ' $null = [System.IO.Directory]::CreateDirectory([string]$path)',
+ // `[string[]]`, not `@(...)`: ConvertFrom-Json emits the parsed array as a single pipeline
+ // object, so `@(...)` wraps it in *another* array and the loop variable binds to the whole
+ // thing. `[string]` of that is the paths joined by spaces, which CreateDirectory rejects with
+ // "The given path's format is not supported". It only ever worked for a one-element batch,
+ // where stringifying a single-element array happens to yield the element. Measured on
+ // WindowsPowerShell 5.1.26100 against a three-directory tree.
+ 'foreach ($path in [string[]]($json | ConvertFrom-Json)) {',
+ ' $null = [System.IO.Directory]::CreateDirectory($path)',
'}'
].join('; ')
)
diff --git a/src/main/ssh/system-ssh-sftp-args.test.ts b/src/main/ssh/system-ssh-sftp-args.test.ts
new file mode 100644
index 00000000000..d971390f8dd
--- /dev/null
+++ b/src/main/ssh/system-ssh-sftp-args.test.ts
@@ -0,0 +1,141 @@
+/**
+ * `buildSshArgs` is shared with the sftp client, and three of its flags mean something else there.
+ * Every case below is a silent wrong-target rather than an error if the translation is skipped,
+ * which is why the fallback is "refuse and use another transport", never "pass it through".
+ */
+import { describe, expect, it } from 'vitest'
+import {
+ SftpArgTranslationError,
+ translateSshArgsToSftpArgs,
+ withSftpKeepalive
+} from './system-ssh-sftp-args'
+
+describe('translateSshArgsToSftpArgs', () => {
+ it('sends the port as an option, since sftp -p preserves mtimes instead', () => {
+ const args = translateSshArgsToSftpArgs(['-p', '2222', '--', 'dev@win.example'])
+
+ expect(args).toEqual(['-o', 'Port=2222', '--', 'dev@win.example'])
+ })
+
+ it('sends the login name as an option, since sftp has no -l', () => {
+ // `buildSshArgs` emits `-l` for a config alias no Host block claims. Throwing here would send
+ // exactly those hosts to the transport this PR exists to stop using, silently.
+ const args = translateSshArgsToSftpArgs(['-l', 'neil', '--', 'awin'])
+
+ expect(args).toEqual(['-o', 'User=neil', '--', 'awin'])
+ })
+
+ it('translates the whole unclaimed-alias shape buildSshArgs emits', () => {
+ const args = translateSshArgsToSftpArgs([
+ '-o',
+ 'BatchMode=no',
+ '-T',
+ '-S',
+ 'none',
+ '-o',
+ 'Hostname=192.168.0.186',
+ '-p',
+ '2222',
+ '-l',
+ 'neil',
+ '--',
+ 'awin'
+ ])
+
+ expect(args).toEqual([
+ '-o',
+ 'BatchMode=no',
+ '-o',
+ 'ControlPath=none',
+ '-o',
+ 'Hostname=192.168.0.186',
+ '-o',
+ 'Port=2222',
+ '-o',
+ 'User=neil',
+ '--',
+ 'awin'
+ ])
+ })
+
+ it('spells ControlPath=none out, since sftp -S names a program to run', () => {
+ // `sftp -S none` would try to exec a binary called `none`.
+ const args = translateSshArgsToSftpArgs(['-S', 'none', '--', 'dev@win.example'])
+
+ expect(args).toEqual(['-o', 'ControlPath=none', '--', 'dev@win.example'])
+ })
+
+ it('refuses any other -S, which would hand sftp an ssh binary Orca did not choose', () => {
+ expect(() => translateSshArgsToSftpArgs(['-S', '/tmp/ctl.sock'])).toThrow(
+ SftpArgTranslationError
+ )
+ })
+
+ it('drops -T, which sftp does not have', () => {
+ expect(translateSshArgsToSftpArgs(['-T', '--', 'host'])).toEqual(['--', 'host'])
+ })
+
+ it('passes through the flags both clients spell the same way', () => {
+ const args = translateSshArgsToSftpArgs([
+ '-F',
+ '/tmp/config',
+ '-o',
+ 'BatchMode=yes',
+ '-i',
+ '/tmp/key',
+ '-J',
+ 'jump.example',
+ '--',
+ 'dev@win.example'
+ ])
+
+ expect(args).toEqual([
+ '-F',
+ '/tmp/config',
+ '-o',
+ 'BatchMode=yes',
+ '-i',
+ '/tmp/key',
+ '-J',
+ 'jump.example',
+ '--',
+ 'dev@win.example'
+ ])
+ })
+
+ it('takes everything after -- as the destination without reinterpreting it', () => {
+ // A host literally named `-p` is not a flag once `--` has been seen.
+ expect(translateSshArgsToSftpArgs(['--', '-p'])).toEqual(['--', '-p'])
+ })
+
+ it('refuses an unknown flag rather than guessing what sftp would do with it', () => {
+ // The point of the throw: a flag added to buildSshArgs later must degrade to another
+ // transport, not reach sftp carrying a different meaning.
+ expect(() => translateSshArgsToSftpArgs(['-A', '--', 'host'])).toThrow(SftpArgTranslationError)
+ })
+
+ it('refuses a value flag with no value', () => {
+ expect(() => translateSshArgsToSftpArgs(['-o'])).toThrow(SftpArgTranslationError)
+ })
+})
+
+describe('withSftpKeepalive', () => {
+ it('asks OpenSSH to notice a dead peer, since the transfer itself has no wall-clock bound', () => {
+ expect(withSftpKeepalive(['--', 'host'])).toEqual([
+ '-o',
+ 'ServerAliveInterval=15',
+ '-o',
+ 'ServerAliveCountMax=3',
+ '--',
+ 'host'
+ ])
+ })
+
+ it('leaves a caller-stated keepalive policy alone', () => {
+ const args = withSftpKeepalive(['-o', 'ServerAliveInterval=60', '--', 'host'])
+
+ expect(args.filter((arg) => arg.startsWith('ServerAliveInterval'))).toEqual([
+ 'ServerAliveInterval=60'
+ ])
+ })
+})
diff --git a/src/main/ssh/system-ssh-sftp-args.ts b/src/main/ssh/system-ssh-sftp-args.ts
new file mode 100644
index 00000000000..17fa37e78bb
--- /dev/null
+++ b/src/main/ssh/system-ssh-sftp-args.ts
@@ -0,0 +1,95 @@
+/**
+ * Rewrites `buildSshArgs` output for the sftp(1) client.
+ *
+ * Three flags ssh and sftp share spell different things: sftp's `-p` is "preserve mtime", its `-S`
+ * names the ssh binary to run, and it has no `-T` at all. Passing ssh's list through unchanged
+ * would silently connect to the wrong port and try to exec a program called `none`.
+ *
+ * Anything this table does not recognize throws. A flag added to `buildSshArgs` later must degrade
+ * to the non-sftp transfer path, never reach sftp carrying a different meaning.
+ */
+
+/** `buildSshArgs` emitted a flag with no sftp equivalent; the caller should use another transport. */
+export class SftpArgTranslationError extends Error {
+ constructor(flag: string) {
+ super(`No sftp equivalent for system ssh argument ${JSON.stringify(flag)}`)
+ this.name = 'SftpArgTranslationError'
+ }
+}
+
+/** Flags whose spelling and meaning are identical in both clients. */
+const PASSTHROUGH_VALUE_FLAGS = new Set(['-F', '-o', '-i', '-J'])
+
+export function translateSshArgsToSftpArgs(sshArgs: readonly string[]): string[] {
+ const sftpArgs: string[] = []
+ let index = 0
+ while (index < sshArgs.length) {
+ const flag = sshArgs[index]!
+ if (flag === '--') {
+ // Everything after `--` is the destination, which both clients spell the same way.
+ sftpArgs.push(...sshArgs.slice(index))
+ return sftpArgs
+ }
+ const value = sshArgs[index + 1]
+ if (PASSTHROUGH_VALUE_FLAGS.has(flag)) {
+ if (value === undefined) {
+ throw new SftpArgTranslationError(flag)
+ }
+ sftpArgs.push(flag, value)
+ index += 2
+ continue
+ }
+ if (flag === '-T') {
+ // sftp never allocates a tty, so ssh's "no tty" request has nothing to translate to.
+ index += 1
+ continue
+ }
+ if (flag === '-p') {
+ if (value === undefined) {
+ throw new SftpArgTranslationError(flag)
+ }
+ sftpArgs.push('-o', `Port=${value}`)
+ index += 2
+ continue
+ }
+ if (flag === '-l') {
+ // sftp has no `-l`; the login name is an option there. `buildSshArgs` emits this for an
+ // unclaimed config alias, so throwing would route those hosts down the defective path and
+ // then cache the refusal against them for half an hour.
+ if (value === undefined) {
+ throw new SftpArgTranslationError(flag)
+ }
+ sftpArgs.push('-o', `User=${value}`)
+ index += 2
+ continue
+ }
+ if (flag === '-S') {
+ // ssh's `-S none` is ControlPath=none; sftp's `-S` would run a binary called `none`.
+ if (value !== 'none') {
+ throw new SftpArgTranslationError(flag)
+ }
+ sftpArgs.push('-o', 'ControlPath=none')
+ index += 2
+ continue
+ }
+ throw new SftpArgTranslationError(flag)
+ }
+ return sftpArgs
+}
+
+/**
+ * A transfer that stalls mid-stream has no per-write bound to catch it, so ask OpenSSH to notice a
+ * dead peer itself. Only added when the caller has not already stated a keepalive policy.
+ */
+export function withSftpKeepalive(sftpArgs: readonly string[]): string[] {
+ const hasOption = (name: string): boolean =>
+ sftpArgs.some((arg, position) => sftpArgs[position - 1] === '-o' && arg.startsWith(`${name}=`))
+ const keepalive: string[] = []
+ if (!hasOption('ServerAliveInterval')) {
+ keepalive.push('-o', 'ServerAliveInterval=15')
+ }
+ if (!hasOption('ServerAliveCountMax')) {
+ keepalive.push('-o', 'ServerAliveCountMax=3')
+ }
+ return [...keepalive, ...sftpArgs]
+}
diff --git a/src/main/ssh/system-ssh-sftp-path.test.ts b/src/main/ssh/system-ssh-sftp-path.test.ts
new file mode 100644
index 00000000000..e196c2e3e8f
--- /dev/null
+++ b/src/main/ssh/system-ssh-sftp-path.test.ts
@@ -0,0 +1,59 @@
+/**
+ * Both functions here guard against the same measured failure: sftp's batch lexer treats `\` as an
+ * escape, so a Windows path handed over raw is silently mis-targeted *and the client still exits
+ * 0*. On Windows 11 / OpenSSH 10.0p2, `put src C:\Users\neil\qt\a.bin` created a file literally
+ * named `C` in the start directory and reported success.
+ */
+import { describe, expect, it } from 'vitest'
+import {
+ quoteSftpBatchArgument,
+ toSftpRemotePath,
+ UnsupportedSftpPathError
+} from './system-ssh-sftp-path'
+
+describe('toSftpRemotePath', () => {
+ it('roots a drive path under /, which is the namespace the Windows sftp-server exposes', () => {
+ // `pwd` in that session reports `/C:/Users/dev`.
+ expect(toSftpRemotePath('C:/Users/dev/f.bin')).toBe('/C:/Users/dev/f.bin')
+ })
+
+ it('accepts a path already in that namespace unchanged', () => {
+ expect(toSftpRemotePath('/C:/Users/dev/f.bin')).toBe('/C:/Users/dev/f.bin')
+ })
+
+ it('converts the separators Orca stores paths with', () => {
+ expect(toSftpRemotePath('C:\\Users\\dev\\f.bin')).toBe('/C:/Users/dev/f.bin')
+ })
+
+ it('declines a UNC path rather than guessing where it lands', () => {
+ // A guess here writes real bytes to the wrong place; declining falls back to another transport.
+ expect(() => toSftpRemotePath('//server/share/f.bin')).toThrow(UnsupportedSftpPathError)
+ })
+
+ it('declines a relative path, which would resolve against the session start directory', () => {
+ expect(() => toSftpRemotePath('Users/dev/f.bin')).toThrow(UnsupportedSftpPathError)
+ })
+})
+
+describe('quoteSftpBatchArgument', () => {
+ it('escapes the backslashes in a Windows client local path', () => {
+ // Unescaped, sftp reads this as C:srcf.bin and fails to find the source.
+ expect(quoteSftpBatchArgument('C:\\src\\f.bin')).toBe('"C:\\\\src\\\\f.bin"')
+ })
+
+ it('keeps a path with spaces as one argument', () => {
+ expect(quoteSftpBatchArgument('/tmp/two words.bin')).toBe('"/tmp/two words.bin"')
+ })
+
+ it('escapes an embedded quote, which would otherwise end the argument early', () => {
+ expect(quoteSftpBatchArgument('/tmp/dq".bin')).toBe('"/tmp/dq\\".bin"')
+ })
+
+ it('refuses a line break, which would split one batch command into two', () => {
+ expect(() => quoteSftpBatchArgument('/tmp/a\nrm -rf b')).toThrow(UnsupportedSftpPathError)
+ })
+
+ it('refuses a NUL, which truncates the argument', () => {
+ expect(() => quoteSftpBatchArgument('/tmp/a\0b')).toThrow(UnsupportedSftpPathError)
+ })
+})
diff --git a/src/main/ssh/system-ssh-sftp-path.ts b/src/main/ssh/system-ssh-sftp-path.ts
new file mode 100644
index 00000000000..2b5bfe53f02
--- /dev/null
+++ b/src/main/ssh/system-ssh-sftp-path.ts
@@ -0,0 +1,46 @@
+import { normalizeWindowsRemotePath } from './ssh-remote-platform'
+
+/**
+ * A path this transfer cannot express to sftp. Callers treat it as "use another transport", never
+ * as a transfer failure.
+ */
+export class UnsupportedSftpPathError extends Error {
+ constructor(path: string) {
+ super(`Path cannot be addressed over sftp: ${JSON.stringify(path)}`)
+ this.name = 'UnsupportedSftpPathError'
+ }
+}
+
+/**
+ * Converts a Windows remote path to the namespace OpenSSH's Windows sftp-server exposes, which
+ * roots every drive under `/`: `C:/Users/dev/f` is `/C:/Users/dev/f`, and `pwd` there reports
+ * `/C:/Users/dev`.
+ */
+export function toSftpRemotePath(remotePath: string): string {
+ const normalized = normalizeWindowsRemotePath(remotePath)
+ if (/^\/[a-zA-Z]:\//.test(normalized)) {
+ return normalized
+ }
+ if (/^[a-zA-Z]:\//.test(normalized)) {
+ return `/${normalized}`
+ }
+ // UNC (`//server/share`) and relative paths have no settled mapping in this namespace, and a
+ // guess here writes real bytes to the wrong place. Decline instead.
+ throw new UnsupportedSftpPathError(remotePath)
+}
+
+/**
+ * Quotes one argument of an sftp batch line.
+ *
+ * Escaping is load-bearing, not cosmetic: sftp's batch lexer treats `\` as an escape even inside
+ * double quotes, so an unescaped Windows local path `C:\src\f.bin` is read as `C:srcf.bin`, and an
+ * unescaped destination `C:\Users\dev\f.bin` writes a file literally named `C` in the start
+ * directory — while sftp still exits 0. Both measured on Windows 11 / OpenSSH 10.0p2.
+ */
+export function quoteSftpBatchArgument(value: string): string {
+ if (/[\n\r\0]/.test(value)) {
+ // A line break would split one batch command into two; NUL truncates the argument.
+ throw new UnsupportedSftpPathError(value)
+ }
+ return `"${value.replace(/([\\"])/g, '\\$1')}"`
+}
diff --git a/src/main/ssh/system-ssh-sftp-transfer.ts b/src/main/ssh/system-ssh-sftp-transfer.ts
new file mode 100644
index 00000000000..c50375fa8eb
--- /dev/null
+++ b/src/main/ssh/system-ssh-sftp-transfer.ts
@@ -0,0 +1,191 @@
+import { accessSync, constants, existsSync, statSync } from 'node:fs'
+import { posix, win32 } from 'node:path'
+import type { SshTarget } from '../../shared/ssh-types'
+import { buildSshArgs, type SystemSshBuildArgsOptions } from './system-ssh-args'
+import { findSystemSsh } from './system-ssh-binary'
+import {
+ SftpArgTranslationError,
+ translateSshArgsToSftpArgs,
+ withSftpKeepalive
+} from './system-ssh-sftp-args'
+import {
+ quoteSftpBatchArgument,
+ toSftpRemotePath,
+ UnsupportedSftpPathError
+} from './system-ssh-sftp-path'
+import { throwIfAborted } from './system-ssh-operation-lifecycle'
+import { runProcess } from '../../shared/child-process/run-process'
+
+/** The host answered, but not with an sftp subsystem. The caller must fall back, not fail. */
+export class SftpSubsystemUnavailableError extends Error {
+ constructor(detail: string) {
+ super(`Remote host has no usable sftp subsystem: ${detail}`)
+ this.name = 'SftpSubsystemUnavailableError'
+ }
+}
+
+/**
+ * True for the errors that mean "this host cannot serve sftp at all".
+ *
+ * Host-scoped, and therefore the only errors safe to remember: a capability cache keyed by host
+ * turns anything it accepts into a verdict about every later write to that host. Deliberately
+ * narrow — a permission denial or a missing directory is a real failure that must surface, not a
+ * reason to retry the whole upload down a slower path.
+ */
+export function isSftpUnavailableError(error: unknown): boolean {
+ return error instanceof SftpSubsystemUnavailableError || error instanceof SftpArgTranslationError
+}
+
+/**
+ * True when *this path* cannot be spelled for sftp, which says nothing about the host.
+ *
+ * Kept apart from the host verdict on purpose. A UNC destination, or a local file whose name
+ * contains a newline, is a property of one operation; caching it would degrade every subsequent
+ * write to that host for the cache's whole retry window on the strength of one odd filename.
+ */
+export function isSftpPathUnsupportedError(error: unknown): boolean {
+ return error instanceof UnsupportedSftpPathError
+}
+
+/** Neither kind of refusal moves a byte, so a staged file cannot exist to sweep. */
+export function isSftpRefusalBeforeStaging(error: unknown): boolean {
+ return isSftpUnavailableError(error) || isSftpPathUnsupportedError(error)
+}
+
+function systemSftpCandidates(sshPath: string | null, platform: NodeJS.Platform): string[] {
+ const pathApi = platform === 'win32' ? win32 : posix
+ const executable = platform === 'win32' ? 'sftp.exe' : 'sftp'
+ const candidates: string[] = []
+ // Why the ssh binary's own directory first: a host with two OpenSSH installs must pair the sftp
+ // client with the ssh that `buildSshArgs` was built for, not whichever one PATH happens to reach.
+ if (sshPath) {
+ candidates.push(pathApi.join(pathApi.dirname(sshPath), executable))
+ }
+ if (platform === 'win32') {
+ const systemRoot = process.env.SystemRoot || process.env.WINDIR
+ if (systemRoot) {
+ candidates.push(win32.join(systemRoot, 'System32', 'OpenSSH', executable))
+ }
+ } else {
+ candidates.push('/usr/bin/sftp', '/usr/local/bin/sftp', '/opt/homebrew/bin/sftp')
+ }
+ return candidates
+}
+
+/** Locate the sftp client paired with the system ssh binary. Returns null when there is none. */
+export function findSystemSftp(): string | null {
+ if (process.env.ORCA_SYSTEM_SFTP_PATH) {
+ return process.env.ORCA_SYSTEM_SFTP_PATH
+ }
+ const sshPath = findSystemSsh()
+ for (const candidate of systemSftpCandidates(sshPath, process.platform)) {
+ try {
+ if (!statSync(candidate).isFile()) {
+ continue
+ }
+ if (process.platform !== 'win32') {
+ accessSync(candidate, constants.X_OK)
+ }
+ return candidate
+ } catch {
+ continue
+ }
+ }
+ return findSftpOnPath()
+}
+
+function findSftpOnPath(): string | null {
+ const pathValue = process.env.PATH
+ if (!pathValue) {
+ return null
+ }
+ const pathApi = process.platform === 'win32' ? win32 : posix
+ const executable = process.platform === 'win32' ? 'sftp.exe' : 'sftp'
+ for (const entry of pathValue.split(pathApi.delimiter)) {
+ const directory = entry.trim().replace(/^"|"$/g, '')
+ if (!directory) {
+ continue
+ }
+ const candidate = pathApi.join(directory, executable)
+ if (existsSync(candidate)) {
+ return candidate
+ }
+ }
+ return null
+}
+
+/**
+ * OpenSSH prints this when the server refuses the subsystem — a host with `Subsystem sftp`
+ * commented out, or an internal-sftp block that does not apply to this user.
+ */
+const SUBSYSTEM_REFUSED_PATTERN = /subsystem request failed|no such file or directory.*sftp-server/i
+
+export type SftpBatchOptions = SystemSshBuildArgsOptions & { signal?: AbortSignal }
+
+/**
+ * Runs one sftp batch script.
+ *
+ * The script goes to the *local* sftp client's stdin, which is the point: no remote process ever
+ * reads a redirected stdin, so none of this rides the Windows PowerShell stdin defect.
+ */
+export async function runSftpBatch(
+ target: SshTarget,
+ commands: readonly string[],
+ options?: SftpBatchOptions
+): Promise {
+ throwIfAborted(options?.signal)
+ const sftpPath = findSystemSftp()
+ if (!sftpPath) {
+ throw new SftpSubsystemUnavailableError('no sftp client binary found alongside ssh')
+ }
+ const args = withSftpKeepalive(translateSshArgsToSftpArgs(buildSshArgs(target, options)))
+ let result
+ try {
+ result = await runProcess({
+ program: sftpPath,
+ args: ['-b', '-', ...args],
+ // `-b -` takes the script on stdin, and that stdin is the *local* client's — no remote
+ // process reads a pipe anywhere in this transfer, which is the whole point of preferring it.
+ input: `${commands.join('\n')}\n`,
+ // Why no timeout: a large upload is legitimately slow, and a wall-clock cap would fail a
+ // healthy transfer on a slow link. A dead peer is caught by the ServerAlive options instead.
+ timeoutMs: null,
+ signal: options?.signal
+ })
+ } catch (error) {
+ // A client that will not start is "this host cannot do sftp" from the caller's side, not a
+ // transfer failure: the payload never left. Falling back is the only useful answer.
+ throw new SftpSubsystemUnavailableError(
+ `sftp client at ${sftpPath} could not be started: ${error instanceof Error ? error.message : String(error)}`
+ )
+ }
+ if (result.code === 0) {
+ return
+ }
+ throwIfAborted(options?.signal)
+ const detail = result.stderr.trim()
+ if (SUBSYSTEM_REFUSED_PATTERN.test(detail)) {
+ throw new SftpSubsystemUnavailableError(detail)
+ }
+ throw new Error(`sftp batch failed (exit ${result.code}): ${detail}`)
+}
+
+/**
+ * Creates remote directories, parents first.
+ *
+ * `-mkdir` keeps sftp going when a directory is already there; batch mode otherwise aborts the
+ * whole script on the first non-zero status, which for an idempotent tree walk is not a failure.
+ */
+export function makeDirectoriesViaSftp(
+ target: SshTarget,
+ remoteDirectories: readonly string[],
+ options?: SftpBatchOptions
+): Promise {
+ const commands = remoteDirectories.map(
+ (directory) => `-mkdir ${quoteSftpBatchArgument(toSftpRemotePath(directory))}`
+ )
+ if (commands.length === 0) {
+ return Promise.resolve()
+ }
+ return runSftpBatch(target, commands, options)
+}
diff --git a/src/main/ssh/system-ssh-windows-file-write.ts b/src/main/ssh/system-ssh-windows-file-write.ts
new file mode 100644
index 00000000000..8b98d4aec50
--- /dev/null
+++ b/src/main/ssh/system-ssh-windows-file-write.ts
@@ -0,0 +1,138 @@
+import { randomBytes } from 'node:crypto'
+import { powerShellCommand, powerShellLiteral } from './ssh-remote-powershell'
+import { normalizeWindowsRemotePath } from './ssh-remote-platform'
+
+/**
+ * Suffix marking the path a Windows write lands on before it is published by rename.
+ *
+ * The random tail is the fix for a measured harm, not decoration. A write that loses contact with
+ * the host leaves a remote process that may still hold the staging file open exclusively, and
+ * `docs/reference/ssh-execution-boundary.md` is explicit that losing contact is not evidence that
+ * process died — so the retry must not reuse the name it may still own. A fresh name per attempt
+ * means a retry never meets its predecessor's lock; the abandoned file is cleaned up best-effort
+ * and never treated as proof of anything.
+ */
+export const WINDOWS_STAGED_WRITE_SUFFIX = '.orca-partial'
+
+export function makeWindowsStagingPath(remotePath: string): string {
+ return `${remotePath}${WINDOWS_STAGED_WRITE_SUFFIX}-${randomBytes(6).toString('hex')}`
+}
+
+export type WindowsPublishMode = 'create' | 'exclusive' | 'append'
+
+/**
+ * Publishes a staged upload onto its real name.
+ *
+ * Every branch reads the staged *file*, never a redirected stdin, which is what makes this safe on
+ * a host whose Windows PowerShell 5.1 cannot drain a piped stdin.
+ *
+ * The replacing branch must never delete the destination first. Deleting and then moving loses the
+ * user's existing file outright if the move fails, and exposes a window where a reader sees no file
+ * at all — a worse outcome than the truncated-partial this staging discipline exists to prevent.
+ * `File.Replace` is the atomic swap (Win32 `ReplaceFile`), and it requires the destination to
+ * exist, so an absent one falls back to a plain `Move`. That fallback is raced deliberately: if the
+ * destination appears in between, `Move` throws, the staged file survives, and the destination is
+ * left exactly as whoever created it left it.
+ *
+ * `File::Move` throwing on an existing destination is also precisely the exclusive contract, which
+ * is why that branch needs nothing else.
+ */
+export function makeWindowsPublishStagedFileCommand(
+ stagingPath: string,
+ remotePath: string,
+ mode: WindowsPublishMode
+): string {
+ const preamble = [
+ '$ErrorActionPreference = "Stop"',
+ `$staging = ${powerShellLiteral(stagingPath)}`,
+ `$path = ${powerShellLiteral(remotePath)}`,
+ '$parent = [System.IO.Path]::GetDirectoryName($path)',
+ 'if ($parent) { $null = [System.IO.Directory]::CreateDirectory($parent) }'
+ ]
+ if (mode === 'append') {
+ return powerShellCommand(
+ [
+ ...preamble,
+ // Not atomic, and cannot cheaply be: appending is defined as extending the destination, so
+ // a failure part-way leaves it longer than it was rather than destroyed. The caller's
+ // chunked-append protocol already restarts from its own offset.
+ '$in = [System.IO.File]::OpenRead($staging)',
+ '$out = [System.IO.File]::Open($path, [System.IO.FileMode]::Append, [System.IO.FileAccess]::Write, [System.IO.FileShare]::None)',
+ 'try { $in.CopyTo($out) } finally { $out.Dispose(); $in.Dispose() }',
+ '[System.IO.File]::Delete($staging)'
+ ].join('; ')
+ )
+ }
+ if (mode === 'exclusive') {
+ return powerShellCommand([...preamble, '[System.IO.File]::Move($staging, $path)'].join('; '))
+ }
+ return powerShellCommand(
+ [
+ ...preamble,
+ // `[NullString]::Value`, not `$null`: PowerShell coerces a bare `$null` to an empty string
+ // when binding a .NET `string` parameter, and `Replace` rejects that with "The path is not
+ // of a legal form" — so every publish would fail. Measured on WindowsPowerShell 5.1.26100.
+ 'try { [System.IO.File]::Replace($staging, $path, [NullString]::Value) } catch [System.IO.FileNotFoundException] { [System.IO.File]::Move($staging, $path) }'
+ ].join('; ')
+ )
+}
+
+/** Best-effort removal of a staged file whose write was abandoned. Never asserts the writer died. */
+export function makeWindowsDiscardStagedFileCommand(stagingPath: string): string {
+ return powerShellCommand(
+ [
+ // Deliberately not `Stop`: the previous writer may still hold this file, and that is a
+ // possibility to tolerate, not an error to report. The unique staging name means a leftover
+ // blocks nothing; sweeping it is housekeeping.
+ '$ErrorActionPreference = "SilentlyContinue"',
+ `$staging = ${powerShellLiteral(stagingPath)}`,
+ '[System.IO.File]::Delete($staging)'
+ ].join('; ')
+ )
+}
+
+/**
+ * The ancestor directories of a Windows remote path, drive root first.
+ *
+ * sftp's `mkdir` creates one level, so a batch has to name each level itself. The drive root is
+ * excluded: `-mkdir "/C:/"` is not a directory anyone creates.
+ */
+export function windowsRemoteAncestorDirectories(remotePath: string): string[] {
+ const normalized = normalizeWindowsRemotePath(remotePath)
+ const segments = normalized.split('/')
+ segments.pop()
+ const ancestors: string[] = []
+ // Start past the drive (`C:`) or the UNC host, which are never created.
+ for (let depth = 2; depth <= segments.length; depth += 1) {
+ const directory = segments.slice(0, depth).join('/')
+ if (directory) {
+ ancestors.push(directory)
+ }
+ }
+ return ancestors
+}
+
+/**
+ * `[Console]::OpenStandardInput()` into a `FileStream`, used only by the two stdin fallbacks.
+ *
+ * On Windows PowerShell 5.1 this is the defective read; see the strategy comment in
+ * `system-ssh-file-binary-transfer.ts`. It is correct under PowerShell 7.
+ */
+export function makeWindowsWriteFileCommand(
+ remotePath: string,
+ options?: { append?: boolean; exclusive?: boolean; executable?: 'powershell.exe' | 'pwsh.exe' }
+): string {
+ const fileMode = options?.append ? 'Append' : options?.exclusive ? 'CreateNew' : 'Create'
+ return powerShellCommand(
+ [
+ '$ErrorActionPreference = "Stop"',
+ `$path = ${powerShellLiteral(remotePath)}`,
+ '$parent = [System.IO.Path]::GetDirectoryName($path)',
+ 'if ($parent) { $null = [System.IO.Directory]::CreateDirectory($parent) }',
+ '$inputStream = [Console]::OpenStandardInput()',
+ `$outputStream = [System.IO.File]::Open($path, [System.IO.FileMode]::${fileMode}, [System.IO.FileAccess]::Write, [System.IO.FileShare]::None)`,
+ 'try { $inputStream.CopyTo($outputStream) } finally { $outputStream.Dispose() }'
+ ].join('; '),
+ options?.executable ?? 'powershell.exe'
+ )
+}
diff --git a/src/main/ssh/system-ssh-windows-upload.test.ts b/src/main/ssh/system-ssh-windows-upload.test.ts
index 207c3e2df8e..c3a5ff80276 100644
--- a/src/main/ssh/system-ssh-windows-upload.test.ts
+++ b/src/main/ssh/system-ssh-windows-upload.test.ts
@@ -1,29 +1,40 @@
/**
- * #16432: the Windows relay upload pushed the whole bundle into one PowerShell stdin, which
- * Windows PowerShell 5.1 cannot drain over a non-pty ssh exec — the remote blocks forever, and
- * `waitForChannelClose()` had no timeout, so the UI sat at "Connecting…" with no error. Covered
- * here: no write exceeds one stdin's worth on any Windows path (bundle upload *and* single-file
- * upload, which is the one that carries large files), a partial write never lands under the real
- * name, and a remote that never closes fails instead of hanging.
+ * #16432. The original fix chunked the payload because the constraint was believed to be a ~50KB
+ * cmd.exe stdin ceiling. Re-measured on Windows 11 26200.9168 / OpenSSH_for_Windows_10.0p2, it is
+ * not a size limit and not cmd.exe's: a read on Windows PowerShell 5.1's redirected-stdin handle
+ * over a non-pty ssh exec can die permanently when it finds the stream momentarily empty, taking
+ * both the remaining data and the EOF with it. It is probabilistic per such read — identical 2MB
+ * payloads died at 167936, 270336 and 372736 — so a 32KB chunk still failed 15 times in 120 under
+ * load, while `findstr` took 2,016,000 bytes through one exec on the same host.
+ *
+ * So the covering property is no longer "every write is small". It is "the bytes do not cross a
+ * remote process's stdin at all": sftp first, PowerShell 7 next, and Windows PowerShell 5.1 last,
+ * bounded and loud. The staging-and-rename discipline is kept on every path, with a unique staging
+ * name per attempt so a retry never meets a predecessor's lock.
*/
import { EventEmitter } from 'node:events'
import { mkdirSync, mkdtempSync, writeFileSync } from 'node:fs'
-import { rm } from 'node:fs/promises'
+import { readFile, rm, stat } from 'node:fs/promises'
import { tmpdir } from 'node:os'
import { join } from 'node:path'
import { PassThrough, Writable } from 'node:stream'
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import type * as SystemSshOperationLifecycle from './system-ssh-operation-lifecycle'
-const { spawnSystemSshCommandMock, waitForChannelCloseSpy } = vi.hoisted(() => ({
+const { spawnSystemSshCommandMock, waitForChannelCloseSpy, runProcessMock } = vi.hoisted(() => ({
spawnSystemSshCommandMock: vi.fn(),
- waitForChannelCloseSpy: vi.fn()
+ waitForChannelCloseSpy: vi.fn(),
+ runProcessMock: vi.fn()
}))
vi.mock('./system-ssh-command', () => ({
spawnSystemSshCommand: spawnSystemSshCommandMock
}))
+vi.mock('../../shared/child-process/run-process', () => ({
+ runProcess: runProcessMock
+}))
+
// Delegates to the real implementation; the spy only records whether each wait was given a bound.
vi.mock('./system-ssh-operation-lifecycle', async (importActual) => {
const actual = (await importActual()) as typeof SystemSshOperationLifecycle
@@ -41,6 +52,11 @@ import {
} from './system-ssh-file-binary-transfer'
import { waitForChannelClose } from './system-ssh-operation-lifecycle'
import { getRemoteHostPlatform } from './ssh-remote-platform'
+import {
+ clearWindowsRemoteWriteCapabilitiesForTests,
+ getWindowsRemoteWriteCapabilities
+} from './system-ssh-windows-write-capabilities'
+import { explainWindowsPowerShellStdinFailure } from './system-ssh-windows-write-strategy'
import type { SshTarget } from '../../shared/ssh-types'
type FakeChannel = EventEmitter & {
@@ -50,7 +66,12 @@ type FakeChannel = EventEmitter & {
written: Buffer
}
-const target = { id: 'win-1', host: 'win.example', username: 'dev' } as unknown as SshTarget
+const target = {
+ id: 'win-1',
+ host: 'win.example',
+ username: 'dev',
+ port: 22
+} as unknown as SshTarget
const hostPlatform = getRemoteHostPlatform('win32-x64')
const remoteRoot = 'C:/Users/dev/.orca-remote'
@@ -78,96 +99,238 @@ function createFakeChannel(onEnd: (channel: FakeChannel) => void): FakeChannel {
return channel
}
-type RecordedCommand = { script: string; stdin: Buffer }
+type RecordedCommand = { script: string; executable: string; stdin: Buffer }
+type RecordedSftpBatch = { args: string[]; script: string }
-describe('Windows upload stdin framing', () => {
- let localDir: string
- const commands: RecordedCommand[] = []
- /** Index of the spawn that should report a non-zero exit, to model a chunk failing mid-file. */
- let failAtSpawn = -1
+const sftpBatches: RecordedSftpBatch[] = []
+const commands: RecordedCommand[] = []
+/** Index of the exec that should report a non-zero exit, to model a chunk failing mid-file. */
+let failAtSpawn = -1
+let localDir: string
- const fileWrites = (): RecordedCommand[] =>
- commands.filter((command) => command.script.includes('FileMode]::'))
- const writtenPath = (command: RecordedCommand): string =>
- /\$path = '((?:[^']|'')*)'/.exec(command.script)?.[1].replace(/''/g, "'") ?? ''
- const fileMode = (command: RecordedCommand): string | undefined =>
- /FileMode\]::(\w+)/.exec(command.script)?.[1]
+const fileWrites = (): RecordedCommand[] =>
+ commands.filter((command) => command.script.includes('OpenStandardInput'))
+const writtenPath = (command: RecordedCommand): string =>
+ /\$path = '((?:[^']|'')*)'/.exec(command.script)?.[1]?.replace(/''/g, "'") ?? ''
+const fileMode = (command: RecordedCommand): string | undefined =>
+ /FileMode\]::(\w+)/.exec(command.script)?.[1]
+const putLines = (): string[] =>
+ sftpBatches.flatMap((batch) => batch.script.split('\n').filter((line) => line.startsWith('put ')))
+const putDestination = (line: string): string => /put "(?:[^"]*)" "([^"]*)"/.exec(line)?.[1] ?? ''
+const putSource = (line: string): string => /put "([^"]*)"/.exec(line)?.[1] ?? ''
- beforeEach(() => {
- commands.length = 0
- failAtSpawn = -1
- waitForChannelCloseSpy.mockClear()
- localDir = mkdtempSync(join(tmpdir(), 'orca-win-upload-'))
- spawnSystemSshCommandMock.mockReset()
- spawnSystemSshCommandMock.mockImplementation((_target: SshTarget, command: string) => {
- const spawnIndex = spawnSystemSshCommandMock.mock.calls.length - 1
- return createFakeChannel((channel) => {
- commands.push({ script: decodePowerShellCommand(command), stdin: channel.written })
- setImmediate(() =>
- spawnIndex === failAtSpawn
- ? channel.emit('close', 1, null)
- : channel.emit('close', 0, null)
- )
+/** Makes every sftp batch succeed, recording what it was asked to do. */
+function acceptSftp(): void {
+ runProcessMock.mockImplementation(
+ async (spec: { args: string[]; input: string; program: string }) => {
+ const script = spec.input
+ sftpBatches.push({ args: spec.args, script })
+ // Model the real client: `put` copies the local file, so read it while it still exists.
+ for (const line of script.split('\n').filter((entry) => entry.startsWith('put '))) {
+ await readFile(putSource(line))
+ }
+ return { code: 0, signal: null, stdout: '', stderr: '', timedOut: false }
+ }
+ )
+}
+
+/** Models a host whose sshd has no `Subsystem sftp` line. */
+function refuseSftp(): void {
+ runProcessMock.mockImplementation(async (spec: { args: string[]; input: string }) => {
+ sftpBatches.push({ args: spec.args, script: spec.input })
+ return {
+ code: 255,
+ signal: null,
+ stdout: '',
+ stderr: 'subsystem request failed on channel 0\nConnection closed',
+ timedOut: false
+ }
+ })
+}
+
+/** Models a host with no PowerShell 7, which cmd.exe reports as an unrecognized command. */
+function refusePwsh(): void {
+ spawnSystemSshCommandMock.mockImplementation((_target: SshTarget, command: string) => {
+ const spawnIndex = spawnSystemSshCommandMock.mock.calls.length - 1
+ const executable = command.split(' ')[0] ?? ''
+ return createFakeChannel((channel) => {
+ commands.push({
+ script: decodePowerShellCommand(command),
+ executable,
+ stdin: channel.written
+ })
+ setImmediate(() => {
+ if (executable === 'pwsh.exe') {
+ channel.stderr.write(
+ "'pwsh.exe' is not recognized as an internal or external command,\noperable program or batch file."
+ )
+ channel.emit('close', 9009, null)
+ return
+ }
+ channel.emit('close', spawnIndex === failAtSpawn ? 1 : 0, null)
})
})
})
+}
- afterEach(async () => {
- await rm(localDir, { recursive: true, force: true })
+beforeEach(() => {
+ commands.length = 0
+ sftpBatches.length = 0
+ failAtSpawn = -1
+ clearWindowsRemoteWriteCapabilitiesForTests()
+ waitForChannelCloseSpy.mockClear()
+ localDir = mkdtempSync(join(tmpdir(), 'orca-win-upload-'))
+ process.env.ORCA_SYSTEM_SFTP_PATH = '/usr/bin/sftp'
+ runProcessMock.mockReset()
+ acceptSftp()
+ spawnSystemSshCommandMock.mockReset()
+ spawnSystemSshCommandMock.mockImplementation((_target: SshTarget, command: string) => {
+ const spawnIndex = spawnSystemSshCommandMock.mock.calls.length - 1
+ return createFakeChannel((channel) => {
+ commands.push({
+ script: decodePowerShellCommand(command),
+ executable: command.split(' ')[0] ?? '',
+ stdin: channel.written
+ })
+ setImmediate(() =>
+ spawnIndex === failAtSpawn ? channel.emit('close', 1, null) : channel.emit('close', 0, null)
+ )
+ })
})
+})
- it('never pushes a whole artifact bundle into one PowerShell stdin', async () => {
- mkdirSync(join(localDir, 'node'), { recursive: true })
- // Comfortably past the ~50KB point at which the reporter measured PowerShell 5.1 wedging.
- writeFileSync(join(localDir, 'node', 'relay.js'), Buffer.alloc(600 * 1024, 0x61))
- writeFileSync(join(localDir, 'index.js'), Buffer.alloc(300 * 1024, 0x62))
+afterEach(async () => {
+ delete process.env.ORCA_SYSTEM_SFTP_PATH
+ await rm(localDir, { recursive: true, force: true })
+})
- await uploadDirectoryViaSystemSsh(target, localDir, remoteRoot, { hostPlatform })
-
- const largest = Math.max(...commands.map((command) => command.stdin.length))
- expect(largest).toBeLessThanOrEqual(WINDOWS_STDIN_WRITE_CHUNK_BYTES)
- // The base64 + JSON envelope is gone entirely: nothing reads the bundle as one string.
- expect(commands.some((command) => command.script.includes('FromBase64String'))).toBe(false)
- // `[Console]::In` wedged at 50KB where the stream reader did not, so the mkdir batch — the one
- // payload still read as a string — must use the reader the reporter measured surviving.
- expect(commands.some((command) => command.script.includes('[Console]::In.ReadToEnd()'))).toBe(
- false
- )
- expect(
- commands.filter((command) => command.script.includes('StreamReader([Console]::'))
- ).toHaveLength(1)
- })
-
- it('bounds the single-file upload too, which is the path large files take', async () => {
- const contents = Buffer.alloc(WINDOWS_STDIN_WRITE_CHUNK_BYTES * 3 + 11, 0x64)
+describe('Windows upload over sftp', () => {
+ it('moves the payload without any remote process reading a stdin', async () => {
+ const contents = Buffer.alloc(WINDOWS_STDIN_WRITE_CHUNK_BYTES * 60 + 11, 0x64)
const localPath = join(localDir, 'big.node')
writeFileSync(localPath, contents)
await uploadFileViaSystemSsh(target, localPath, `${remoteRoot}/big.node`, { hostPlatform })
- const writes = fileWrites()
- expect(writes).toHaveLength(4)
- expect(Math.max(...writes.map((write) => write.stdin.length))).toBe(
- WINDOWS_STDIN_WRITE_CHUNK_BYTES
- )
- expect(Buffer.concat(writes.map((write) => write.stdin)).equals(contents)).toBe(true)
- // A wedged PowerShell never closes on its own, so no wait on this path may be unbounded.
- expect(
- waitForChannelCloseSpy.mock.calls.every((call) => call[2] === WINDOWS_STDIN_WRITE_TIMEOUT_MS)
- ).toBe(true)
+ // The defect is a remote stdin read; the fix is that there is not one.
+ expect(fileWrites()).toHaveLength(0)
+ expect(putLines()).toHaveLength(1)
+ // One transfer, not 61 execs: the whole point of the change.
+ expect(sftpBatches).toHaveLength(1)
})
- it('writes every byte of every artifact across the chunked writes', async () => {
- const contents = Buffer.alloc(WINDOWS_STDIN_WRITE_CHUNK_BYTES * 2 + 17, 0x63)
- writeFileSync(join(localDir, 'relay.js'), contents)
+ it('creates the parent chain and sends the payload in one round trip', async () => {
+ writeFileSync(join(localDir, 'relay.js'), 'x')
- await uploadDirectoryViaSystemSsh(target, localDir, remoteRoot, { hostPlatform })
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), `${remoteRoot}/a/b/relay.js`, {
+ hostPlatform
+ })
- const writes = fileWrites()
- expect(writes).toHaveLength(3)
- expect(Buffer.concat(writes.map((write) => write.stdin)).equals(contents)).toBe(true)
- // Only the first write creates the staging file; the rest must extend it or it is truncated.
- expect(writes.map(fileMode)).toEqual(['Create', 'Append', 'Append'])
+ expect(sftpBatches).toHaveLength(1)
+ expect(sftpBatches[0]!.script.split('\n').filter(Boolean)).toEqual([
+ '-mkdir "/C:/Users"',
+ '-mkdir "/C:/Users/dev"',
+ '-mkdir "/C:/Users/dev/.orca-remote"',
+ '-mkdir "/C:/Users/dev/.orca-remote/a"',
+ '-mkdir "/C:/Users/dev/.orca-remote/a/b"',
+ expect.stringContaining('put ') as unknown as string
+ ])
+ })
+
+ it('addresses the destination in the drive-rooted namespace sftp exposes', async () => {
+ writeFileSync(join(localDir, 'relay.js'), 'x')
+
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), `${remoteRoot}/relay.js`, {
+ hostPlatform
+ })
+
+ // A backslash destination silently writes a file named `C` and still exits 0, so the leading
+ // slash and forward separators are correctness, not style.
+ expect(putDestination(putLines()[0]!)).toMatch(
+ /^\/C:\/Users\/dev\/\.orca-remote\/relay\.js\.orca-partial-[0-9a-f]{12}$/
+ )
+ })
+
+ it('never lands a partial under the real name, and publishes by rename', async () => {
+ writeFileSync(join(localDir, 'relay.js'), 'x')
+ const remotePath = `${remoteRoot}/relay.js`
+
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), remotePath, { hostPlatform })
+
+ const destination = putDestination(putLines()[0]!)
+ // Assert the positive first: an unmatched regex yields '', which would satisfy the `not.toBe`
+ // below without this test ever having seen a destination.
+ expect(destination).toContain(WINDOWS_STAGED_WRITE_SUFFIX)
+ expect(destination).not.toBe(`/C:${remotePath.slice(2)}`)
+ const publish = commands.at(-1)!
+ expect(publish.script).toContain(
+ '[System.IO.File]::Replace($staging, $path, [NullString]::Value)'
+ )
+ // The publish reads the staged file, never a pipe, so it is safe on PowerShell 5.1.
+ expect(publish.script).not.toContain('OpenStandardInput')
+ })
+
+ it('never deletes the destination it is replacing', async () => {
+ writeFileSync(join(localDir, 'relay.js'), 'x')
+
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), `${remoteRoot}/relay.js`, {
+ hostPlatform
+ })
+
+ const publish = commands.at(-1)!
+ // Delete-then-move destroys the user's existing file outright if the move then fails, and
+ // exposes a window where a reader sees no file at all — worse than the truncated partial the
+ // staging discipline exists to prevent. `File.Replace` is the atomic swap.
+ expect(publish.script).not.toContain('[System.IO.File]::Delete($path)')
+ expect(publish.script).toContain(
+ '[System.IO.File]::Replace($staging, $path, [NullString]::Value)'
+ )
+ // An absent destination cannot be Replaced, so that case falls back to a plain Move.
+ expect(publish.script).toContain(
+ 'catch [System.IO.FileNotFoundException] { [System.IO.File]::Move($staging, $path) }'
+ )
+ })
+
+ it('gives every attempt its own staging name, so a retry cannot meet a predecessor lock', async () => {
+ writeFileSync(join(localDir, 'relay.js'), 'x')
+
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), `${remoteRoot}/relay.js`, {
+ hostPlatform
+ })
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), `${remoteRoot}/relay.js`, {
+ hostPlatform
+ })
+
+ const [first, second] = putLines().map(putDestination)
+ expect(first).toContain(WINDOWS_STAGED_WRITE_SUFFIX)
+ // Losing contact is not evidence the previous writer died, so the name must not be reused.
+ expect(second).not.toBe(first)
+ })
+
+ it('enforces exclusive at the rename, where it is atomic', async () => {
+ writeFileSync(join(localDir, 'import.bin'), 'x')
+
+ await uploadFileViaSystemSsh(target, join(localDir, 'import.bin'), `${remoteRoot}/import.bin`, {
+ hostPlatform,
+ exclusive: true
+ })
+
+ const publish = commands.at(-1)!
+ expect(publish.script).toContain('[System.IO.File]::Move($staging, $path)')
+ expect(publish.script).not.toContain('[System.IO.File]::Delete($path)')
+ })
+
+ it('appends by concatenating the staged file, not by piping bytes to the remote', async () => {
+ await writeBufferViaSystemSsh(target, `${remoteRoot}/log.bin`, Buffer.from('tail'), {
+ hostPlatform,
+ append: true
+ })
+
+ expect(fileWrites()).toHaveLength(0)
+ const publish = commands.at(-1)!
+ expect(publish.script).toContain('FileMode]::Append')
+ expect(publish.script).toContain('$in.CopyTo($out)')
+ expect(publish.script).toContain('[System.IO.File]::Delete($staging)')
})
it('still creates an empty artifact on the host', async () => {
@@ -175,80 +338,310 @@ describe('Windows upload stdin framing', () => {
await uploadDirectoryViaSystemSsh(target, localDir, remoteRoot, { hostPlatform })
- expect(fileWrites().map(writtenPath)).toEqual([`${remoteRoot}/empty.txt`])
- expect(fileWrites()[0].stdin).toHaveLength(0)
- expect(fileMode(fileWrites()[0])).toBe('Create')
+ expect(putLines()).toHaveLength(1)
+ expect(commands.at(-1)!.script).toContain('[System.IO.File]::Move($staging, $path)')
})
- it('lands a multi-chunk write on a staging path and publishes it by rename', async () => {
- const remotePath = `${remoteRoot}/relay.js`
- writeFileSync(join(localDir, 'relay.js'), Buffer.alloc(WINDOWS_STDIN_WRITE_CHUNK_BYTES + 1))
+ it('writes a buffer through a 0600 temp file that does not outlive the transfer', async () => {
+ const seen: { path: string; contents: Buffer; mode: number }[] = []
+ runProcessMock.mockImplementation(async (spec: { args: string[]; input: string }) => {
+ sftpBatches.push({ args: spec.args, script: spec.input })
+ for (const line of spec.input.split('\n').filter((entry) => entry.startsWith('put '))) {
+ const path = putSource(line)
+ seen.push({
+ path,
+ contents: await readFile(path),
+ mode: (await stat(path)).mode & 0o777
+ })
+ }
+ return { code: 0, signal: null, stdout: '', stderr: '', timedOut: false }
+ })
+
+ await writeBufferViaSystemSsh(target, `${remoteRoot}/version`, Buffer.from('1.2.3'), {
+ hostPlatform
+ })
+
+ expect(seen).toHaveLength(1)
+ expect(seen[0]!.contents.toString()).toBe('1.2.3')
+ // The payload can be repository content and tmpdir is world-readable on every platform, so the
+ // window between write and upload must not be group- or world-readable.
+ expect(seen[0]!.mode).toBe(0o600)
+ await expect(readFile(seen[0]!.path)).rejects.toThrow()
+ })
+
+ it('creates upload directories over sftp rather than a PowerShell stdin batch', async () => {
+ mkdirSync(join(localDir, 'node'), { recursive: true })
+ writeFileSync(join(localDir, 'node', 'relay.js'), 'x')
await uploadDirectoryViaSystemSsh(target, localDir, remoteRoot, { hostPlatform })
- // Nothing touches the real name until every byte is on the host.
- expect(fileWrites().map(writtenPath)).toEqual([
- `${remotePath}${WINDOWS_STAGED_WRITE_SUFFIX}`,
- `${remotePath}${WINDOWS_STAGED_WRITE_SUFFIX}`
- ])
- const publish = commands.at(-1)!
- expect(publish.script).toContain('[System.IO.File]::Move($staging, $path)')
- expect(publish.script).toContain('[System.IO.File]::Delete($path)')
+ // Anchor on a non-empty observation: `some` is false of an empty list, so this would pass even
+ // if no command had been recorded at all.
+ expect(commands.length).toBeGreaterThan(0)
+ expect(commands.some((command) => command.script.includes('StreamReader([Console]::'))).toBe(
+ false
+ )
+ expect(sftpBatches[0]!.script).toContain('-mkdir "/C:/Users/dev/.orca-remote"')
+ })
+
+ it('sweeps the staged bytes when the publish is the thing that fails', async () => {
+ writeFileSync(join(localDir, 'import.bin'), 'x')
+ // An exclusive conflict is the ordinary way to get here: the payload is on the host, and the
+ // rename that would have given it a name refuses.
+ spawnSystemSshCommandMock.mockImplementation((_target: SshTarget, command: string) => {
+ const script = decodePowerShellCommand(command)
+ return createFakeChannel((channel) => {
+ commands.push({ script, executable: command.split(' ')[0] ?? '', stdin: channel.written })
+ const failed = script.includes('::Move($staging, $path)')
+ setImmediate(() => channel.emit('close', failed ? 1 : 0, null))
+ })
+ })
+
+ await expect(
+ uploadFileViaSystemSsh(target, join(localDir, 'import.bin'), `${remoteRoot}/import.bin`, {
+ hostPlatform,
+ exclusive: true
+ })
+ ).rejects.toThrow()
+
+ const sweep = commands.at(-1)!
+ expect(sweep.script).toContain('[System.IO.File]::Delete($staging)')
+ // Tolerated, not asserted: the previous writer may still hold the file, and losing contact is
+ // not evidence it died.
+ expect(sweep.script).toContain('$ErrorActionPreference = "SilentlyContinue"')
+ })
+
+ it('reports a cancelled transfer as an abort, not as a failed one', async () => {
+ writeFileSync(join(localDir, 'relay.js'), 'x')
+ const controller = new AbortController()
+ // runProcess reports the kill as a non-zero exit rather than throwing, so without checking the
+ // signal first a user pressing cancel is indistinguishable from the transfer genuinely failing.
+ runProcessMock.mockImplementation(async (spec: { args: string[]; input: string }) => {
+ sftpBatches.push({ args: spec.args, script: spec.input })
+ controller.abort()
+ return { code: 255, signal: 'SIGTERM', stdout: '', stderr: '', timedOut: false }
+ })
+
+ let error: Error | undefined
+ try {
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), `${remoteRoot}/relay.js`, {
+ hostPlatform,
+ signal: controller.signal
+ })
+ } catch (thrown) {
+ error = thrown as Error
+ }
+
+ expect(error?.name).toBe('AbortError')
+ expect(error?.message).not.toContain('sftp batch failed')
+ // A cancel is also not evidence about the host, so it must not send later writes to the slow
+ // path, and must not fall through to the defective reader now.
+ expect(getWindowsRemoteWriteCapabilities(target).shouldTry('sftp-subsystem')).toBe(true)
+ expect(fileWrites()).toHaveLength(0)
+ })
+
+ it('does not let one unaddressable path become a verdict about the host', async () => {
+ writeFileSync(join(localDir, 'relay.js'), 'x')
+
+ // A UNC destination has no settled mapping in sftp's drive-rooted namespace, so this write
+ // falls back — but the host still serves sftp perfectly well for every other path.
+ await uploadFileViaSystemSsh(
+ target,
+ join(localDir, 'relay.js'),
+ '//fileserver/share/relay.js',
+ { hostPlatform }
+ )
+
+ expect(fileWrites().length).toBeGreaterThan(0)
+ expect(sftpBatches).toHaveLength(0)
+ // The 30-minute capability cache is keyed by host; caching this would send every later write
+ // to the same machine down the defective path on the strength of one odd destination.
+ expect(getWindowsRemoteWriteCapabilities(target).shouldTry('sftp-subsystem')).toBe(true)
+ })
+
+ it('keeps using sftp for the next file after one path it could not spell', async () => {
+ writeFileSync(join(localDir, 'relay.js'), 'x')
+
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), '//fileserver/share/a.js', {
+ hostPlatform
+ })
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), `${remoteRoot}/b.js`, {
+ hostPlatform
+ })
+
+ expect(putLines()).toHaveLength(1)
+ expect(putDestination(putLines()[0]!)).toContain('/C:/Users/dev/.orca-remote/b.js')
+ })
+
+ it('does not let a local filename sftp cannot quote become a verdict either', async () => {
+ // POSIX clients allow a newline in a filename, and sftp's batch lexer would read it as the end
+ // of one command and the start of another.
+ const awkward = join(localDir, 'two\nlines.js')
+ writeFileSync(awkward, 'x')
+
+ await uploadFileViaSystemSsh(target, awkward, `${remoteRoot}/relay.js`, { hostPlatform })
+
+ expect(fileWrites().length).toBeGreaterThan(0)
+ expect(getWindowsRemoteWriteCapabilities(target).shouldTry('sftp-subsystem')).toBe(true)
+ })
+
+ it('translates the ssh argument list rather than passing it to a client that reads it differently', async () => {
+ writeFileSync(join(localDir, 'relay.js'), 'x')
+
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), `${remoteRoot}/relay.js`, {
+ hostPlatform,
+ disableControlMaster: true
+ })
+
+ const args = sftpBatches[0]!.args
+ // sftp's `-T` does not exist, its `-p` preserves mtime, and its `-S` names a program to run.
+ expect(args).not.toContain('-T')
+ expect(args).not.toContain('-p')
+ expect(args).not.toContain('-S')
+ expect(args).toContain('ControlPath=none')
+ expect(args).toContain('ServerAliveInterval=15')
+ })
+})
+
+describe('Windows upload on a host with no sftp subsystem', () => {
+ beforeEach(() => {
+ refuseSftp()
+ })
+
+ it('creates a multi-directory tree, which the one-element case never exercised', async () => {
+ mkdirSync(join(localDir, 'node', 'deep'), { recursive: true })
+ writeFileSync(join(localDir, 'index.js'), 'a')
+ writeFileSync(join(localDir, 'node', 'deep', 'x.js'), 'b')
+
+ await uploadDirectoryViaSystemSsh(target, localDir, remoteRoot, { hostPlatform })
+
+ const mkdir = commands.find((command) => command.script.includes('ConvertFrom-Json'))!
+ // `@($json | ConvertFrom-Json)` wraps the parsed array in another array, so the loop variable
+ // binds to the whole thing and `[string]` of it is the paths joined by spaces — which
+ // CreateDirectory rejects. It only ever worked for a single directory, where stringifying a
+ // one-element array happens to yield the element, so no batch of one can catch this.
+ expect(mkdir.script).toContain('[string[]]($json | ConvertFrom-Json)')
+ expect(mkdir.script).not.toContain('@($json | ConvertFrom-Json)')
+ const batch = JSON.parse(mkdir.stdin.toString('utf-8')) as string[]
+ expect(batch.length).toBeGreaterThan(1)
+ })
+
+ it('falls back rather than failing the transfer', async () => {
+ const contents = Buffer.alloc(WINDOWS_STDIN_WRITE_CHUNK_BYTES + 5, 0x61)
+ writeFileSync(join(localDir, 'relay.js'), contents)
+
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), `${remoteRoot}/relay.js`, {
+ hostPlatform
+ })
+
+ expect(Buffer.concat(fileWrites().map((write) => write.stdin)).equals(contents)).toBe(true)
+ })
+
+ it('remembers the refusal, so a multi-file upload probes once', async () => {
+ writeFileSync(join(localDir, 'a.js'), 'a')
+ writeFileSync(join(localDir, 'b.js'), 'b')
+ writeFileSync(join(localDir, 'c.js'), 'c')
+
+ await uploadDirectoryViaSystemSsh(target, localDir, remoteRoot, { hostPlatform })
+
+ // One refusal is enough; re-probing per file is a wasted round trip on every file.
+ expect(sftpBatches).toHaveLength(1)
+ })
+
+ it('does not spend a sweep round trip when sftp declined before moving any bytes', async () => {
+ writeFileSync(join(localDir, 'relay.js'), 'x')
+
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), `${remoteRoot}/relay.js`, {
+ hostPlatform
+ })
+
+ // A refused subsystem staged nothing, so there is nothing to delete — and on a host without
+ // sftp that sweep would otherwise be paid on every single write.
+ expect(commands.length).toBeGreaterThan(0)
+ expect(commands.some((command) => command.script.includes('Delete($staging)'))).toBe(false)
+ })
+
+ it('prefers PowerShell 7, which reads a redirected stdin correctly', async () => {
+ writeFileSync(join(localDir, 'relay.js'), Buffer.alloc(WINDOWS_STDIN_WRITE_CHUNK_BYTES * 3))
+
+ await uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), `${remoteRoot}/relay.js`, {
+ hostPlatform
+ })
+
+ expect(fileWrites().map((write) => write.executable)).toEqual(['pwsh.exe'])
+ // PowerShell 7 took 2MB through one exec when measured, so chunking it buys nothing.
+ expect(fileWrites()[0]!.stdin).toHaveLength(WINDOWS_STDIN_WRITE_CHUNK_BYTES * 3)
+ })
+
+ it('bounds every write when only Windows PowerShell 5.1 is available', async () => {
+ refusePwsh()
+ const contents = Buffer.alloc(WINDOWS_STDIN_WRITE_CHUNK_BYTES * 3 + 11, 0x64)
+ writeFileSync(join(localDir, 'big.node'), contents)
+
+ await uploadFileViaSystemSsh(target, join(localDir, 'big.node'), `${remoteRoot}/big.node`, {
+ hostPlatform
+ })
+
+ const writes = fileWrites().filter((write) => write.executable === 'powershell.exe')
+ expect(writes).toHaveLength(4)
+ expect(Math.max(...writes.map((write) => write.stdin.length))).toBe(
+ WINDOWS_STDIN_WRITE_CHUNK_BYTES
+ )
+ expect(Buffer.concat(writes.map((write) => write.stdin)).equals(contents)).toBe(true)
+ expect(writes.map(fileMode)).toEqual(['Create', 'Append', 'Append', 'Append'])
+ // A wedged PowerShell never closes on its own, so no wait on this path may be unbounded.
+ // Count first: `every` is true of zero calls, so a wait that moved to a different helper would
+ // pass this silently.
+ expect(waitForChannelCloseSpy.mock.calls.length).toBeGreaterThan(0)
+ expect(
+ waitForChannelCloseSpy.mock.calls.every((call) => call[2] === WINDOWS_STDIN_WRITE_TIMEOUT_MS)
+ ).toBe(true)
+ })
+
+ it('remembers that PowerShell 7 is absent instead of re-probing per chunk', async () => {
+ refusePwsh()
+ writeFileSync(join(localDir, 'big.node'), Buffer.alloc(WINDOWS_STDIN_WRITE_CHUNK_BYTES * 3))
+
+ await uploadFileViaSystemSsh(target, join(localDir, 'big.node'), `${remoteRoot}/big.node`, {
+ hostPlatform
+ })
+
+ expect(fileWrites().filter((write) => write.executable === 'pwsh.exe')).toHaveLength(1)
})
it('leaves no truncated file under the real name when a chunk fails mid-file', async () => {
writeFileSync(join(localDir, 'relay.js'), Buffer.alloc(WINDOWS_STDIN_WRITE_CHUNK_BYTES * 3))
- // Spawns: 0 = mkdir batch, 1..3 = chunk writes. Fail the second chunk.
- failAtSpawn = 2
+ // Spawn 0 is the pwsh write; fail it and every retry beneath it.
+ failAtSpawn = 0
await expect(
- uploadDirectoryViaSystemSsh(target, localDir, remoteRoot, { hostPlatform })
+ uploadFileViaSystemSsh(target, join(localDir, 'relay.js'), `${remoteRoot}/relay.js`, {
+ hostPlatform
+ })
).rejects.toThrow()
+ expect(fileWrites().length).toBeGreaterThan(0)
expect(fileWrites().map(writtenPath)).not.toContain(`${remoteRoot}/relay.js`)
expect(commands.some((command) => command.script.includes('::Move('))).toBe(false)
})
+})
- it('enforces exclusive once at the rename, so a retry is not blocked by its own leftovers', async () => {
- const localPath = join(localDir, 'import.bin')
- writeFileSync(localPath, Buffer.alloc(WINDOWS_STDIN_WRITE_CHUNK_BYTES + 1))
+describe('last-resort Windows PowerShell failure reporting', () => {
+ it('names the host limitation and its remedy, not just the timeout', () => {
+ const timeout = new Error('write C:/x at offset 0 timed out after 60000ms with no response')
- await uploadFileViaSystemSsh(target, localPath, `${remoteRoot}/import.bin`, {
- hostPlatform,
- exclusive: true
- })
+ const explained = explainWindowsPowerShellStdinFailure(timeout) as Error
- // CreateNew on chunk one would fail against a leftover staging file from a failed attempt;
- // `File::Move` raising on an existing destination is what carries the exclusive contract.
- expect(fileWrites().map(fileMode)).toEqual(['Create', 'Append'])
- const publish = commands.at(-1)!
- expect(publish.script).toContain('[System.IO.File]::Move($staging, $path)')
- expect(publish.script).not.toContain('[System.IO.File]::Delete($path)')
+ // "timed out" alone sends the user to retry a network they cannot fix; the fix is host-side.
+ expect(explained.message).toContain('Windows PowerShell 5.1')
+ expect(explained.message).toContain('Subsystem sftp sftp-server.exe')
+ expect(explained.cause).toBe(timeout)
})
- it('keeps a single-chunk write on the destination, with the caller mode intact', async () => {
- await writeBufferViaSystemSsh(target, `${remoteRoot}/version`, Buffer.from('1.2.3'), {
- hostPlatform,
- exclusive: true
- })
+ it('leaves a real failure alone, so a permission error is not reported as a host limitation', () => {
+ const denied = new Error('write C:/x at offset 0 failed (exit 1): Access to the path is denied')
- expect(fileWrites()).toHaveLength(1)
- expect(writtenPath(fileWrites()[0])).toBe(`${remoteRoot}/version`)
- expect(fileMode(fileWrites()[0])).toBe('CreateNew')
- expect(commands.some((command) => command.script.includes('::Move('))).toBe(false)
- })
-
- it('appends onto the destination rather than staging, since append cannot be staged', async () => {
- const remotePath = `${remoteRoot}/log.bin`
- await writeBufferViaSystemSsh(
- target,
- remotePath,
- Buffer.alloc(WINDOWS_STDIN_WRITE_CHUNK_BYTES + 1),
- { hostPlatform, append: true }
- )
-
- expect(fileWrites().map(writtenPath)).toEqual([remotePath, remotePath])
- expect(fileWrites().map(fileMode)).toEqual(['Append', 'Append'])
+ expect(explainWindowsPowerShellStdinFailure(denied)).toBe(denied)
})
})
diff --git a/src/main/ssh/system-ssh-windows-write-capabilities.test.ts b/src/main/ssh/system-ssh-windows-write-capabilities.test.ts
new file mode 100644
index 00000000000..ad723d592b0
--- /dev/null
+++ b/src/main/ssh/system-ssh-windows-write-capabilities.test.ts
@@ -0,0 +1,79 @@
+/**
+ * Whether a Windows host has an sftp subsystem is a fact about that host, so the cache is keyed by
+ * the endpoint that executes rather than by Orca's target id — otherwise a hardened host is
+ * re-probed once per file, and two targets pointing at one machine learn the same fact twice.
+ */
+import { afterEach, describe, expect, it } from 'vitest'
+import type { SshTarget } from '../../shared/ssh-types'
+import {
+ clearWindowsRemoteWriteCapabilitiesForTests,
+ getWindowsRemoteWriteCapabilities,
+ getWindowsRemoteWriteExecutionHostKey
+} from './system-ssh-windows-write-capabilities'
+
+const asTarget = (fields: Partial): SshTarget => fields as SshTarget
+
+afterEach(() => {
+ clearWindowsRemoteWriteCapabilitiesForTests()
+})
+
+describe('getWindowsRemoteWriteExecutionHostKey', () => {
+ it('gives two targets on one endpoint the same key', () => {
+ const first = asTarget({ id: 'a', host: 'win.example', username: 'dev', port: 22 })
+ const second = asTarget({ id: 'b', host: 'win.example', username: 'dev', port: 22 })
+
+ // A target re-created under a new id has not changed what the host supports.
+ expect(getWindowsRemoteWriteExecutionHostKey(first)).toBe(
+ getWindowsRemoteWriteExecutionHostKey(second)
+ )
+ })
+
+ it('separates hosts, ports and users', () => {
+ const base = { id: 'a', host: 'win.example', username: 'dev', port: 22 }
+ const keys = [
+ asTarget(base),
+ asTarget({ ...base, host: 'other.example' }),
+ asTarget({ ...base, port: 2222 }),
+ asTarget({ ...base, username: 'ops' })
+ ].map(getWindowsRemoteWriteExecutionHostKey)
+
+ expect(new Set(keys).size).toBe(4)
+ })
+
+ it('keys a config alias by the alias, since ssh_config decides where it lands', () => {
+ const alias = asTarget({ id: 'a', host: 'stale.example', configHost: 'winbox' })
+
+ expect(getWindowsRemoteWriteExecutionHostKey(alias)).toBe('config:winbox')
+ })
+})
+
+describe('getWindowsRemoteWriteCapabilities', () => {
+ it('shares one cache across targets that reach the same host', () => {
+ const first = asTarget({ id: 'a', host: 'win.example', username: 'dev', port: 22 })
+ const second = asTarget({ id: 'b', host: 'win.example', username: 'dev', port: 22 })
+
+ getWindowsRemoteWriteCapabilities(first).rememberUnsupported('sftp-subsystem')
+
+ expect(getWindowsRemoteWriteCapabilities(second).shouldTry('sftp-subsystem')).toBe(false)
+ })
+
+ it('does not let one host answer for another', () => {
+ const hardened = asTarget({ id: 'a', host: 'hardened.example', username: 'dev', port: 22 })
+ const ordinary = asTarget({ id: 'b', host: 'ordinary.example', username: 'dev', port: 22 })
+
+ getWindowsRemoteWriteCapabilities(hardened).rememberUnsupported('sftp-subsystem')
+
+ expect(getWindowsRemoteWriteCapabilities(ordinary).shouldTry('sftp-subsystem')).toBe(true)
+ })
+
+ it('keeps the two capabilities independent', () => {
+ const target = asTarget({ id: 'a', host: 'win.example', username: 'dev', port: 22 })
+ const capabilities = getWindowsRemoteWriteCapabilities(target)
+
+ capabilities.rememberUnsupported('pwsh')
+
+ // No PowerShell 7 says nothing about whether the host will serve sftp.
+ expect(capabilities.shouldTry('sftp-subsystem')).toBe(true)
+ expect(capabilities.shouldTry('pwsh')).toBe(false)
+ })
+})
diff --git a/src/main/ssh/system-ssh-windows-write-capabilities.ts b/src/main/ssh/system-ssh-windows-write-capabilities.ts
new file mode 100644
index 00000000000..dcd03f19807
--- /dev/null
+++ b/src/main/ssh/system-ssh-windows-write-capabilities.ts
@@ -0,0 +1,52 @@
+import type { SshTarget } from '../../shared/ssh-types'
+import { CapabilityProbeCache } from '../../shared/capability-probe-cache'
+
+/**
+ * Whether a Windows host can take a file write over the sftp subsystem, and whether it has a
+ * PowerShell 7 to fall back to. Both are host facts, so they are cached per execution host rather
+ * than per transfer — a hardened host with `Subsystem sftp` removed must not be re-probed on every
+ * file of a multi-file upload.
+ */
+export type WindowsRemoteWriteCapability = 'sftp-subsystem' | 'pwsh'
+
+// Why re-probe at all: an admin can enable the subsystem, or install PowerShell 7, without the
+// user restarting Orca. Long enough that a hardened host costs one failed probe per half hour.
+export const WINDOWS_WRITE_CAPABILITY_RETRY_INTERVAL_MS = 30 * 60_000
+
+const capabilitiesByExecutionHost = new Map<
+ string,
+ CapabilityProbeCache
+>()
+
+/**
+ * Keyed by the endpoint that executes, not by target id: two Orca targets pointing at one host
+ * describe the same sshd, and a target re-created under a new id has not changed what that host
+ * supports. A config alias is its own key because ssh_config, not Orca, resolves where it lands.
+ */
+export function getWindowsRemoteWriteExecutionHostKey(target: SshTarget): string {
+ if (target.configHost) {
+ return `config:${target.configHost}`
+ }
+ const port = target.port ?? 22
+ return target.username
+ ? `host:${target.username}@${target.host}:${port}`
+ : `host:${target.host}:${port}`
+}
+
+export function getWindowsRemoteWriteCapabilities(
+ target: SshTarget
+): CapabilityProbeCache {
+ const key = getWindowsRemoteWriteExecutionHostKey(target)
+ let cache = capabilitiesByExecutionHost.get(key)
+ if (!cache) {
+ cache = new CapabilityProbeCache(
+ WINDOWS_WRITE_CAPABILITY_RETRY_INTERVAL_MS
+ )
+ capabilitiesByExecutionHost.set(key, cache)
+ }
+ return cache
+}
+
+export function clearWindowsRemoteWriteCapabilitiesForTests(): void {
+ capabilitiesByExecutionHost.clear()
+}
diff --git a/src/main/ssh/system-ssh-windows-write-strategy.ts b/src/main/ssh/system-ssh-windows-write-strategy.ts
new file mode 100644
index 00000000000..f2cdca12516
--- /dev/null
+++ b/src/main/ssh/system-ssh-windows-write-strategy.ts
@@ -0,0 +1,329 @@
+import type { SshTarget } from '../../shared/ssh-types'
+import { getSystemSshBuildArgsFromOperationOptions } from './system-ssh-args'
+import { spawnSystemSshCommand } from './system-ssh-command'
+import {
+ awaitWithSystemSshAbort,
+ throwIfAborted,
+ waitForChannelClose
+} from './system-ssh-operation-lifecycle'
+import {
+ isSftpPathUnsupportedError,
+ isSftpRefusalBeforeStaging,
+ isSftpUnavailableError,
+ runSftpBatch
+} from './system-ssh-sftp-transfer'
+import { quoteSftpBatchArgument, toSftpRemotePath } from './system-ssh-sftp-path'
+import { getWindowsRemoteWriteCapabilities } from './system-ssh-windows-write-capabilities'
+import {
+ makeWindowsDiscardStagedFileCommand,
+ makeWindowsPublishStagedFileCommand,
+ makeWindowsStagingPath,
+ makeWindowsWriteFileCommand,
+ windowsRemoteAncestorDirectories,
+ type WindowsPublishMode
+} from './system-ssh-windows-file-write'
+
+/** No Windows stdin write should ever outlive this; a wedged PowerShell never closes on its own. */
+export const WINDOWS_STDIN_WRITE_TIMEOUT_MS = 60_000
+
+/**
+ * Bound on one stdin write for the last-resort Windows PowerShell 5.1 path.
+ *
+ * Measured on Windows 11 26200 / OpenSSH 10.0p2: a 32KB write still hangs 15 times in 120 under
+ * load, and no smaller value removes the risk. The defect is per blocking read, not per byte, so
+ * shrinking the chunk trades one risky read for more execs that each carry their own. This is a
+ * damage bound on a path known to be unreliable, not a safe size.
+ */
+export const WINDOWS_STDIN_WRITE_CHUNK_BYTES = 32 * 1024
+
+export type WindowsWriteOptions = Parameters<
+ typeof getSystemSshBuildArgsFromOperationOptions
+>[0] & {
+ signal?: AbortSignal
+ append?: boolean
+ exclusive?: boolean
+}
+
+/** Bytes to write, plus a way to present them to sftp, which can only send a local file. */
+export type WindowsWriteSource = {
+ totalBytes: number
+ readChunk: (offset: number, maxBytes: number) => Promise
+ withLocalFile: (send: (localPath: string) => Promise) => Promise
+}
+
+function publishMode(options: WindowsWriteOptions): WindowsPublishMode {
+ return options.append ? 'append' : options.exclusive === true ? 'exclusive' : 'create'
+}
+
+/**
+ * Writes one file to a Windows host, preferring transports that do not push bytes through a remote
+ * PowerShell's stdin.
+ *
+ * Order, and why: sftp carries the whole payload in one transfer and never has a remote process
+ * read a pipe. Measured on Windows 11 / OpenSSH 10.0p2: 1.9MB in a median 315ms over sftp against
+ * 0 of 6 completions on the chunked path, whose best case was ~62 execs at ~350ms each. PowerShell
+ * 7 reads a redirected stdin correctly but is not installed by default. Windows PowerShell 5.1 is
+ * always present and is the defective reader, so it is last and it is bounded.
+ *
+ * Every transport stages under a unique name and publishes by rename, so no partial write is ever
+ * visible under the real name and no retry inherits a predecessor's lock.
+ */
+export async function writeWindowsRemoteFile(
+ target: SshTarget,
+ remotePath: string,
+ source: WindowsWriteSource,
+ options: WindowsWriteOptions
+): Promise {
+ throwIfAborted(options.signal)
+ const capabilities = getWindowsRemoteWriteCapabilities(target)
+ await capabilities.runWithFallback(
+ 'sftp-subsystem',
+ () => writeViaSftp(target, remotePath, source, options),
+ () => writeViaRemoteStdin(target, remotePath, source, options),
+ isSftpUnavailableError
+ )
+}
+
+/**
+ * Stages under a name nothing else can own, publishes it, and sweeps the staging file if either
+ * step fails.
+ *
+ * Shared by both transports so the cleanup contract cannot drift between them: a failed publish —
+ * an exclusive conflict is the ordinary case — leaves bytes on the host that no longer have a
+ * purpose, and the sweep is what stops them accumulating.
+ */
+async function stageThenPublish(
+ target: SshTarget,
+ remotePath: string,
+ options: WindowsWriteOptions,
+ stage: (stagingPath: string) => Promise,
+ nothingStaged: (error: unknown) => boolean = () => false
+): Promise {
+ const stagingPath = makeWindowsStagingPath(remotePath)
+ try {
+ await stage(stagingPath)
+ await publishStagedWrite(target, stagingPath, remotePath, options)
+ } catch (error) {
+ // A transport that declined before it moved any bytes has nothing to sweep, and sweeping
+ // anyway would spend a round trip on every write to a host that has no sftp subsystem.
+ if (!nothingStaged(error)) {
+ await discardStagedWrite(target, stagingPath, options)
+ }
+ throw error
+ }
+}
+
+/**
+ * A path sftp cannot address falls back for this write alone, without touching the host verdict.
+ *
+ * The distinction matters because the capability cache is keyed by host and holds for half an hour:
+ * routing one UNC destination, or one local filename containing a newline, into
+ * `rememberUnsupported` would send every later write to that host down the defective path too.
+ */
+async function writeViaSftp(
+ target: SshTarget,
+ remotePath: string,
+ source: WindowsWriteSource,
+ options: WindowsWriteOptions
+): Promise {
+ try {
+ await attemptSftpWrite(target, remotePath, source, options)
+ } catch (error) {
+ if (!isSftpPathUnsupportedError(error)) {
+ throw error
+ }
+ await writeViaRemoteStdin(target, remotePath, source, options)
+ }
+}
+
+function attemptSftpWrite(
+ target: SshTarget,
+ remotePath: string,
+ source: WindowsWriteSource,
+ options: WindowsWriteOptions
+): Promise {
+ const mkdirs = windowsRemoteAncestorDirectories(remotePath).map(
+ (directory) => `-mkdir ${quoteSftpBatchArgument(toSftpRemotePath(directory))}`
+ )
+ return stageThenPublish(
+ target,
+ remotePath,
+ options,
+ (stagingPath) =>
+ source.withLocalFile((localPath) =>
+ // One round trip: the parent chain and the payload travel in the same batch.
+ runSftpBatch(
+ target,
+ [
+ ...mkdirs,
+ `put ${quoteSftpBatchArgument(localPath)} ${quoteSftpBatchArgument(toSftpRemotePath(stagingPath))}`
+ ],
+ options
+ )
+ ),
+ isSftpRefusalBeforeStaging
+ )
+}
+
+function writeViaRemoteStdin(
+ target: SshTarget,
+ remotePath: string,
+ source: WindowsWriteSource,
+ options: WindowsWriteOptions
+): Promise {
+ const capabilities = getWindowsRemoteWriteCapabilities(target)
+ return stageThenPublish(target, remotePath, options, (stagingPath) =>
+ capabilities.runWithFallback(
+ 'pwsh',
+ () => writeStdinChunks(target, stagingPath, source, options, 'pwsh.exe'),
+ () => writeStdinChunks(target, stagingPath, source, options, 'powershell.exe'),
+ isPwshUnavailableError
+ )
+ )
+}
+
+/**
+ * PowerShell 7 takes the whole payload in one exec — measured at 2MB — so only the 5.1 path pays
+ * for chunking, and only because a bounded write is the most that path can be trusted with.
+ */
+async function writeStdinChunks(
+ target: SshTarget,
+ stagingPath: string,
+ source: WindowsWriteSource,
+ options: WindowsWriteOptions,
+ executable: 'powershell.exe' | 'pwsh.exe'
+): Promise {
+ const chunkBytes =
+ executable === 'pwsh.exe' ? Math.max(source.totalBytes, 1) : WINDOWS_STDIN_WRITE_CHUNK_BYTES
+ let offset = 0
+ // An empty write still has to run: it is what creates the staged file.
+ do {
+ const chunk = await source.readChunk(offset, chunkBytes)
+ if (chunk.length === 0 && offset < source.totalBytes) {
+ throw new Error(`Source ran short during upload of ${stagingPath}`)
+ }
+ await writeOneStdinChunk(
+ target,
+ stagingPath,
+ chunk,
+ { ...options, append: offset > 0, exclusive: false },
+ offset,
+ executable
+ )
+ offset += chunk.length
+ } while (offset < source.totalBytes)
+}
+
+async function writeOneStdinChunk(
+ target: SshTarget,
+ stagingPath: string,
+ chunk: Buffer,
+ options: WindowsWriteOptions,
+ offset: number,
+ executable: 'powershell.exe' | 'pwsh.exe'
+): Promise {
+ throwIfAborted(options.signal)
+ const channel = spawnSystemSshCommand(
+ target,
+ makeWindowsWriteFileCommand(stagingPath, {
+ append: options.append,
+ exclusive: options.exclusive,
+ executable
+ }),
+ { wrapCommand: false, ...getSystemSshBuildArgsFromOperationOptions(options) }
+ )
+ const closePromise = awaitWithSystemSshAbort(
+ options.signal,
+ () => channel.close(),
+ waitForChannelClose(
+ channel,
+ `write ${stagingPath} at offset ${offset}`,
+ WINDOWS_STDIN_WRITE_TIMEOUT_MS
+ )
+ ).catch((error: unknown) => {
+ throw executable === 'powershell.exe' ? explainWindowsPowerShellStdinFailure(error) : error
+ })
+ if (!options.signal?.aborted) {
+ channel.stdin.end(chunk)
+ }
+ await closePromise
+}
+
+/**
+ * Names the cause on the one path that can hang, so the failure is not just "timed out".
+ *
+ * A user seeing this needs to know it is a host limitation with a host-side remedy, not a network
+ * fault they should retry into.
+ */
+export function explainWindowsPowerShellStdinFailure(error: unknown): unknown {
+ const message = error instanceof Error ? error.message : String(error)
+ if (!/timed out/i.test(message)) {
+ return error
+ }
+ return new Error(
+ `${message}\nWindows PowerShell 5.1 can lose a redirected stdin permanently when a read finds it momentarily empty, so this write cannot be made reliable from the client. Enable the sftp subsystem on the host (sshd_config: "Subsystem sftp sftp-server.exe"), or install PowerShell 7, and Orca will use it automatically.`,
+ { cause: error instanceof Error ? error : undefined }
+ )
+}
+
+function isPwshUnavailableError(error: unknown): boolean {
+ const message = error instanceof Error ? error.message : String(error)
+ // cmd.exe's "not recognized" and sshd's exit 9009 both mean "no pwsh here". A timeout does not:
+ // that is the stdin defect, and PowerShell 7 does not have it, so it must not be cached as absent.
+ return /is not recognized as an internal or external command|9009|CommandNotFoundException/i.test(
+ message
+ )
+}
+
+async function publishStagedWrite(
+ target: SshTarget,
+ stagingPath: string,
+ remotePath: string,
+ options: WindowsWriteOptions
+): Promise {
+ await runWindowsCommandWithoutStdin(
+ target,
+ makeWindowsPublishStagedFileCommand(stagingPath, remotePath, publishMode(options)),
+ `publish ${remotePath}`,
+ options
+ )
+}
+
+async function discardStagedWrite(
+ target: SshTarget,
+ stagingPath: string,
+ options: WindowsWriteOptions
+): Promise {
+ try {
+ await runWindowsCommandWithoutStdin(
+ target,
+ makeWindowsDiscardStagedFileCommand(stagingPath),
+ `discard ${stagingPath}`,
+ { ...options, signal: undefined }
+ )
+ } catch {
+ // Housekeeping only. The staging name is unique, so a leftover blocks nothing, and a failure
+ // here says nothing about whether the abandoned writer is still alive.
+ }
+}
+
+function runWindowsCommandWithoutStdin(
+ target: SshTarget,
+ command: string,
+ label: string,
+ options: WindowsWriteOptions
+): Promise {
+ const channel = spawnSystemSshCommand(target, command, {
+ wrapCommand: false,
+ ...getSystemSshBuildArgsFromOperationOptions(options)
+ })
+ const closePromise = awaitWithSystemSshAbort(
+ options.signal,
+ () => channel.close(),
+ waitForChannelClose(channel, label, WINDOWS_STDIN_WRITE_TIMEOUT_MS)
+ )
+ if (!options.signal?.aborted) {
+ channel.stdin.end()
+ }
+ return closePromise
+}
From 7b108abf710d21069691554da8bb4e7d3e2c805a Mon Sep 17 00:00:00 2001
From: Jinwoo Hong <73622457+Jinwoo-H@users.noreply.github.com>
Date: Fri, 4 Sep 2026 04:34:22 -0400
Subject: [PATCH 31/49] fix(relay): stop taking the fleet-wide cell inventory
lock on per-connection paths (#18606)
* fix(relay): stop taking the fleet-wide cell inventory lock on per-connection paths
activateControl, acquireActivity, changeActivity and
removeSupersededSameCellControls each adjust exactly one cell's
reservation, yet took SELECT * FROM relay_cells FOR UPDATE, so every
desktop rebind and phone reconnect in the fleet queued behind every
other one and behind placement. They now use the single-row atomic
update (or lock only their own cell row), leaving the inventory lock to
placement and sweeps.
Fleet-wide 55P03 retries ran p50 430 / p99 1320 per five minutes on
2026-09-03, every cell pinned sqlLatencyMsMax at the lock timeout, and
the old cell image crashed on the resulting pool timeouts ~every 15
minutes. A real-Postgres test holds another cell's row and asserts a
rebind proceeds; re-adding the inventory lock fails it.
* fix(relay): lock the touched cell rows in order on cross-cell activity moves
Review found that acquireActivity's existing-lease branch could lock the
old lease's cell row (via removeActivityLease) before the new cell's row,
which cycles with placement's ascending inventory lock; reproduced on
real Postgres as paired 55P03 retries. lockCellRows now takes the one or
two rows a per-connection path touches in cell_id order with the 500 ms
request bound, and the census fails on any inline relay_cells FOR UPDATE
outside the named lock helpers. A three-cell Postgres test moves an
activity from the highest cell to a lower one while the target row is
held and asserts the mover holds nothing else; five revert-mutants
(inventory lock on each path, dropped ordering, dropped ORDER BY) fail it.
* test(relay): make the inline relay_cells lock census scan whole statements
Review showed two evasions: a FOR UPDATE inside query() and a queryLocked
whose FROM relay_cells sat past a fixed line window. The guard now matches
every query()/queryLocked() template statement in full; both evasions
fail it. Also clears relay_cell_connection_snapshots in the connection-
headroom Postgres suite so an aborted run does not poison the next.
---
...nment-connection-headroom-postgres.test.ts | 6 +
...ment-control-supersession-postgres.test.ts | 4 +
cloud/apps/relay/src/assignment-store.ts | 32 ++-
.../src/cell-inventory-lock-census.test.ts | 73 ++++-
...rol-rebind-inventory-lock-postgres.test.ts | 260 ++++++++++++++++++
5 files changed, 361 insertions(+), 14 deletions(-)
create mode 100644 cloud/apps/relay/src/control-rebind-inventory-lock-postgres.test.ts
diff --git a/cloud/apps/relay/src/assignment-connection-headroom-postgres.test.ts b/cloud/apps/relay/src/assignment-connection-headroom-postgres.test.ts
index 6ac9521c3d6..80a74a47eeb 100644
--- a/cloud/apps/relay/src/assignment-connection-headroom-postgres.test.ts
+++ b/cloud/apps/relay/src/assignment-connection-headroom-postgres.test.ts
@@ -44,6 +44,12 @@ describePostgres('PostgreSQL assignment connection headroom', () => {
`DELETE FROM relay_assignments
WHERE user_id LIKE 'connection-headroom-postgres-%'`
)
+ // A snapshot left by an aborted run rejects the replayed watermark
+ // with stale_connection_snapshot.
+ await database.query(
+ `DELETE FROM relay_cell_connection_snapshots WHERE cell_id = ?`,
+ [cell.id]
+ )
await database.query(
`DELETE FROM relay_cell_connection_runtime WHERE cell_id = ?`,
[cell.id]
diff --git a/cloud/apps/relay/src/assignment-control-supersession-postgres.test.ts b/cloud/apps/relay/src/assignment-control-supersession-postgres.test.ts
index 10193b78cc6..cf8819686b5 100644
--- a/cloud/apps/relay/src/assignment-control-supersession-postgres.test.ts
+++ b/cloud/apps/relay/src/assignment-control-supersession-postgres.test.ts
@@ -38,6 +38,10 @@ describePostgres('PostgreSQL control supersession', () => {
[identity.userId]
)
await database.query(`DELETE FROM relay_assignments WHERE user_id = ?`, [identity.userId])
+ // A snapshot left by an aborted run rejects the replayed watermark with stale_connection_snapshot.
+ await database.query(`DELETE FROM relay_cell_connection_snapshots WHERE cell_id = ?`, [
+ cell.id
+ ])
await database.query(`DELETE FROM relay_cell_connection_runtime WHERE cell_id = ?`, [cell.id])
await database.query(`DELETE FROM relay_cell_connection_limits WHERE cell_id = ?`, [cell.id])
await database.query(`DELETE FROM relay_cell_runtime WHERE cell_id = ?`, [cell.id])
diff --git a/cloud/apps/relay/src/assignment-store.ts b/cloud/apps/relay/src/assignment-store.ts
index d0517d46746..9d240e304a7 100644
--- a/cloud/apps/relay/src/assignment-store.ts
+++ b/cloud/apps/relay/src/assignment-store.ts
@@ -3202,8 +3202,7 @@ export class RelayAssignmentStore {
)
const requestDelta = ACTIVITY_REQUEST_UNITS[kind] * (after - before)
if (requestDelta !== 0) {
- await this.lockCellInventory(transaction, 'request')
- await this.adjustCellReservation(transaction, text(row, 'cell_id'), requestDelta)
+ await this.adjustCellReservationAtomically(transaction, text(row, 'cell_id'), requestDelta)
}
})
})
@@ -3263,9 +3262,12 @@ export class RelayAssignmentStore {
}
const units = ACTIVITY_REQUEST_UNITS[input.kind]
if (existing) {
- await this.lockCellInventory(transaction, 'request')
+ // Why: a client-chosen activity id can move between cells, so lock the
+ // one or two rows this path touches in cell_id order, the same order
+ // placement takes the inventory in, and no cycle can form.
+ await this.lockCellRows(transaction, [text(existing, 'cell_id'), input.cellId])
await this.removeActivityLease(transaction, identity, existing, now)
- await this.adjustCellReservation(transaction, input.cellId, units)
+ await this.adjustCellReservationAtomically(transaction, input.cellId, units)
}
await this.adjustActivityCount(transaction, identity, input.kind, 1, expiresAt, now)
await transaction.query(
@@ -3580,8 +3582,7 @@ export class RelayAssignmentStore {
)
await this.touchAssignment(transaction, identity, expiresAt, now)
} else {
- await this.lockCellInventory(transaction, 'request')
- await this.adjustCellReservation(transaction, input.cellId, 1)
+ await this.adjustCellReservationAtomically(transaction, input.cellId, 1)
await this.adjustActivityCount(transaction, identity, 'control', 1, expiresAt, now)
await transaction.query(
`INSERT INTO relay_assignment_activity_leases
@@ -6954,6 +6955,19 @@ export class RelayAssignmentStore {
return rows
}
+ // Per-connection paths touch one or two cells. Locking exactly those rows,
+ // in the same ascending order the inventory lock uses (ORDER BY fixes the
+ // row-lock order), keeps them off the fleet-wide lock without a cycle.
+ private async lockCellRows(database: RelayDatabase, cellIds: string[]): Promise {
+ const distinct = [...new Set(cellIds)]
+ return await database.queryLocked(
+ `SELECT * FROM relay_cells WHERE cell_id IN (${distinct.map(() => '?').join(', ')})
+ ORDER BY cell_id ASC`,
+ distinct,
+ { lockTimeoutMs: CELL_INVENTORY_LOCK_TIMEOUT_MS }
+ )
+ }
+
private async lockGeneralCellInventory(
database: RelayDatabase,
mode: CellInventoryLockMode
@@ -7590,7 +7604,10 @@ export class RelayAssignmentStore {
) {
throw new Error('activity_lease_shape_mismatch')
}
- const cells = await this.lockCellInventory(database, 'request')
+ // Why: this recomputes one cell's reservation from its leases, so only that
+ // row needs to be held; the 23-row inventory lock here serialised every
+ // desktop control rebind in the fleet behind every other one.
+ const cellRow = (await this.lockCellRows(database, [cellId]))[0]
await database.query(
`DELETE FROM relay_assignment_activity_leases
WHERE user_id = ? AND relay_host_id = ? AND activity_kind = 'control'
@@ -7611,7 +7628,6 @@ export class RelayAssignmentStore {
[cellId]
)
)[0]!
- const cellRow = cells.find((cell) => text(cell, 'cell_id') === cellId)
const cellUnits = integer(cellUnitsRow, 'request_units')
if (!cellRow) throw new Error('assigned_cell_missing')
if (cellUnits > integer(cellRow, 'capacity_requests')) {
diff --git a/cloud/apps/relay/src/cell-inventory-lock-census.test.ts b/cloud/apps/relay/src/cell-inventory-lock-census.test.ts
index 8ca7cee55f5..a26e15f8e1d 100644
--- a/cloud/apps/relay/src/cell-inventory-lock-census.test.ts
+++ b/cloud/apps/relay/src/cell-inventory-lock-census.test.ts
@@ -25,10 +25,12 @@ const CENSUS: CensusEntry[] = [
{ method: 'assignOnce', mode: 'nowait', reach: 'both' },
{ method: 'assignOnce', mode: 'nowait', reach: 'both' },
{ method: 'refreshDrainMigrationLeasesOnce', mode: 'request', reach: 'request' },
- // Reachable from neither: changeActivity has no production callers, only tests.
- { method: 'changeActivity', mode: 'request', reach: 'orphan' },
- { method: 'acquireActivity', mode: 'request', reach: 'request' },
- { method: 'activateControl', mode: 'request', reach: 'request' },
+ // changeActivity, acquireActivity, activateControl and
+ // removeSupersededSameCellControls no longer take the inventory: they lock
+ // only the one or two cell rows they touch, in cell_id order (lockCellRows),
+ // so they cannot cycle with placement's ordered inventory lock, and the
+ // 23-row lock there had serialised every reconnect in the fleet behind every
+ // other one.
{ method: 'startEvacuation', mode: 'request', reach: 'request' },
{ method: 'completeEvacuationFromDeadSourceOnce', mode: 'request', reach: 'request' },
{ method: 'completeEvacuationFromDeadSourceOnce', mode: 'nowait', reach: 'request' },
@@ -48,8 +50,31 @@ const CENSUS: CensusEntry[] = [
{ method: 'releaseExpiredActivityLeases', mode: 'nowait', reach: 'sweep' },
{ method: 'releaseExpiredActivity', mode: 'nowait', reach: 'sweep' },
{ method: 'reconcileReservationAccounting', mode: 'pool-default', reach: 'both' },
- { method: 'leastLoadedCell', mode: 'pool-default', reach: 'both' },
- { method: 'removeSupersededSameCellControls', mode: 'request', reach: 'request' }
+ { method: 'leastLoadedCell', mode: 'pool-default', reach: 'both' }
+]
+
+// Every inline `FROM relay_cells ... FOR UPDATE` outside the named lock helpers,
+// in source order: whole-table locks in reconciliation and sticky placement,
+// and single-row locks for a cell the method is already scoped to (heartbeat,
+// fence, drain generation, configuration, or a reservation adjust that runs
+// under a lock its caller already holds). A new inline lock fails the census
+// below until it is listed here; per-connection paths that touch more than one
+// cell go through lockCellRows so the order is fixed.
+const NAMED_LOCK_HELPERS = ['lockCellInventory', 'lockGeneralCellInventory', 'lockCellRows']
+
+const INLINE_CELL_LOCK_SITES = [
+ 'reconcileCellsWithOptions',
+ 'assignStickyOnce',
+ 'recordCellHeartbeat',
+ 'attestCellFence',
+ 'adoptLegacyCellFence',
+ 'commitLegacyCellFenceAdoption',
+ 'prepareCellFenceAttempt',
+ 'attestCellFenceAttempt',
+ 'attestCellFenceAttempt',
+ 'configureCell',
+ 'assertDrainCellGeneration',
+ 'adjustCellReservation'
]
// The background sweeps, and nothing else. A method reachable from one of these
@@ -151,6 +176,42 @@ describe('cell inventory lock call-site census', () => {
)
})
+ // Why: the census only sees lockCellInventory calls, so a hand-written
+ // `relay_cells ... FOR UPDATE` would escape classification entirely.
+ it('routes every relay_cells row lock through a named lock helper', () => {
+ const lines = storeSource()
+ const rawSites: string[] = []
+ // Whole statements, not a fixed window: a wide column list or a raw
+ // FOR UPDATE inside query() must not slip past.
+ const source = lines.join('\n')
+ const bounds: { name: string; start: number }[] = []
+ lines.forEach((line, index) => {
+ const declaration = DECLARATION.exec(line)
+ if (declaration) bounds.push({ name: declaration[1]!, start: index })
+ })
+ const methodAt = (offset: number): string => {
+ const lineIndex = source.slice(0, offset).split('\n').length - 1
+ let name = ''
+ for (const bound of bounds) if (bound.start <= lineIndex) name = bound.name
+ return name
+ }
+ const tick = String.fromCharCode(96)
+ const statementCall = new RegExp(
+ '\\.(queryLocked|query)\\(\\s*' + tick + '([^' + tick + ']*)' + tick,
+ 'g'
+ )
+ for (const call of source.matchAll(statementCall)) {
+ const statement = call[2]!
+ if (!/\bFROM\s+relay_cells\b/.test(statement)) continue
+ const locks = call[1] === 'queryLocked' || /\bFOR\s+UPDATE\b/.test(statement)
+ if (!locks) continue
+ const method = methodAt(call.index)
+ if (NAMED_LOCK_HELPERS.includes(method)) continue
+ rawSites.push(method)
+ }
+ expect(rawSites).toEqual(INLINE_CELL_LOCK_SITES)
+ })
+
it('leaves no call site taking the inventory without naming a mode', () => {
const source = readFileSync(new URL('./assignment-store.ts', import.meta.url), 'utf8')
const unclassified = source
diff --git a/cloud/apps/relay/src/control-rebind-inventory-lock-postgres.test.ts b/cloud/apps/relay/src/control-rebind-inventory-lock-postgres.test.ts
new file mode 100644
index 00000000000..e990ac1ed1a
--- /dev/null
+++ b/cloud/apps/relay/src/control-rebind-inventory-lock-postgres.test.ts
@@ -0,0 +1,260 @@
+import { afterAll, beforeAll, describe, expect, it } from 'vitest'
+import { RelayAssignmentStore } from './assignment-store.js'
+import { openRelayDatabase, type RelayDatabase } from './database.js'
+
+const databaseUrl = process.env.ORCA_RELAY_TEST_POSTGRES_URL
+const describePostgres = databaseUrl ? describe : describe.skip
+
+// Three cells: the inventory lock covers more than the rows a move touches, and
+// a high-to-low move exposes any lock taken out of cell_id order.
+const cells = [
+ {
+ id: 'rebind-inventory-postgres-a',
+ url: 'https://rebind-inventory-postgres-a.example.com',
+ capacityRequests: 1_000,
+ connectionHardCap: 600 as const,
+ connectionUnobservedBound: 50
+ },
+ {
+ id: 'rebind-inventory-postgres-b',
+ url: 'https://rebind-inventory-postgres-b.example.com',
+ capacityRequests: 1_000,
+ connectionHardCap: 600 as const,
+ connectionUnobservedBound: 50
+ },
+ {
+ id: 'rebind-inventory-postgres-c',
+ url: 'https://rebind-inventory-postgres-c.example.com',
+ capacityRequests: 1_000,
+ connectionHardCap: 600 as const,
+ connectionUnobservedBound: 50
+ }
+]
+const identity = { userId: 'rebind-inventory-postgres-user', relayHostId: 'rebindinvhost001' }
+
+function heartbeat(cell: (typeof cells)[number]) {
+ return {
+ cellId: cell.id,
+ cellUrl: cell.url,
+ cellIncarnation: '11111111-1111-4111-8111-111111111111',
+ startedAt: 50,
+ ready: true,
+ observedRequests: 0,
+ totalConnections: 0,
+ inFlightConnections: 0,
+ reservedConnectionUnits: 0,
+ enforcedConnectionUnits: 0,
+ connectionInclusionWatermark: 1,
+ connectionHardCap: 600 as const,
+ connectionUnobservedBound: 50
+ }
+}
+
+// Why: every desktop control rebind used to take the fleet-wide relay_cells
+// FOR UPDATE lock, so a rebind on one cell queued behind whatever held any
+// other cell's row, until COMMIT (55P03 at the request bound). A rebind only
+// touches its own cell row, so it must proceed while another cell's row is
+// held elsewhere.
+describePostgres('PostgreSQL control rebind under a held cell row', () => {
+ const databases: RelayDatabase[] = []
+
+ beforeAll(async () => {
+ databases.push(
+ await openRelayDatabase({ databaseUrl, dataDir: '' }),
+ await openRelayDatabase({ databaseUrl, dataDir: '' })
+ )
+ })
+
+ async function removeTestRows(database: RelayDatabase): Promise {
+ await database.query(
+ `DELETE FROM relay_control_connection_reservations WHERE user_id = ?`,
+ [identity.userId]
+ )
+ for (const table of [
+ 'relay_assignment_activity_leases',
+ 'relay_post_drain_migration_pins',
+ 'relay_assignment_migration_incarnations',
+ 'relay_assignment_migrations',
+ 'relay_assignments'
+ ]) {
+ await database.query(`DELETE FROM ${table} WHERE user_id = ?`, [identity.userId])
+ }
+ for (const cell of cells) {
+ for (const table of [
+ 'relay_cell_connection_snapshots',
+ 'relay_cell_connection_runtime',
+ 'relay_cell_connection_limits',
+ 'relay_cell_runtime',
+ 'relay_cells'
+ ]) {
+ await database.query(`DELETE FROM ${table} WHERE cell_id = ?`, [cell.id])
+ }
+ }
+ }
+
+ afterAll(async () => {
+ if (databases[0]) await removeTestRows(databases[0])
+ for (const connection of databases) await connection.close()
+ })
+
+ it("rebinds and supersedes a control while another cell's row is held", async () => {
+ // A prior aborted run leaves connection snapshots that reject a replayed watermark.
+ await removeTestRows(databases[0]!)
+ const store = new RelayAssignmentStore(databases[0]!, () => 100)
+ await store.reconcileCells(cells)
+ for (const cell of cells) await store.recordCellHeartbeat(heartbeat(cell))
+ // Pin the host to cell A so placement is deterministic.
+ await store.setCellEnabled(cells[1]!.id, false)
+ await store.setCellEnabled(cells[2]!.id, false)
+ const assignment = await store.assign(identity)
+ expect(assignment.cellId).toBe(cells[0]!.id)
+ await store.setCellEnabled(cells[1]!.id, true)
+ await store.setCellEnabled(cells[2]!.id, true)
+ await store.activateControl(identity, {
+ cellId: cells[0]!.id,
+ assignmentEpoch: assignment.assignmentEpoch,
+ generation: 1,
+ connectionInclusionWatermark: 10
+ })
+
+ // Hold only cell B's row on a second connection, the way a rebind on B
+ // does, for longer than the request-path lock bound.
+ let releaseInventory!: () => void
+ const inventoryReleased = new Promise((resolve) => {
+ releaseInventory = resolve
+ })
+ let inventoryHeld!: () => void
+ const inventoryHeldPromise = new Promise((resolve) => {
+ inventoryHeld = resolve
+ })
+ const holder = databases[1]!.transaction(async (transaction) => {
+ await transaction.queryLocked(`SELECT * FROM relay_cells WHERE cell_id = ?`, [cells[1]!.id])
+ inventoryHeld()
+ await inventoryReleased
+ })
+ await inventoryHeldPromise
+
+ // A generation-2 rebind on cell A supersedes generation 1. It must not
+ // wait on cell B's row.
+ const startedAt = Date.now()
+ const blockedStatement = async (): Promise => {
+ const rows = await databases[1]!.query(
+ `SELECT left(query, 160) AS q FROM pg_stat_activity
+ WHERE datname = current_database() AND wait_event_type = 'Lock'`
+ )
+ return rows.map((row) => String(row.q)).join(' | ')
+ }
+ const timeout = new Promise((_, reject) =>
+ setTimeout(
+ () =>
+ void blockedStatement().then((statement) =>
+ reject(new Error(`rebind on cell A blocked behind cell B's row: ${statement}`))
+ ),
+ 2_000
+ )
+ )
+ const rebound = await Promise.race([
+ store.activateControl(identity, {
+ cellId: cells[0]!.id,
+ assignmentEpoch: assignment.assignmentEpoch,
+ generation: 2,
+ connectionInclusionWatermark: 11
+ }),
+ timeout
+ ])
+ const elapsedMs = Date.now() - startedAt
+ releaseInventory()
+ await holder
+
+ expect(rebound).toBe(`control:${cells[0]!.id}:2`)
+ expect(elapsedMs).toBeLessThan(2_000)
+ const controls = await databases[0]!.query(
+ `SELECT activity_id FROM relay_assignment_activity_leases
+ WHERE user_id = ? AND activity_kind = 'control' ORDER BY activity_id`,
+ [identity.userId]
+ )
+ expect(controls).toEqual([{ activity_id: `control:${cells[0]!.id}:2` }])
+ const reserved = await databases[0]!.query(
+ `SELECT reserved_requests FROM relay_cells WHERE cell_id = ?`,
+ [cells[0]!.id]
+ )
+ expect(Number(reserved[0]!.reserved_requests)).toBe(1)
+ }, 15_000)
+
+ // Why: a phone's activity id is client-chosen and can follow the host across
+ // a migration, so acquireActivity may touch two cell rows. Moving from the
+ // higher cell to the lower one is where an unordered lock cycles with
+ // placement's ascending inventory lock (reproduced live before this fix).
+ it('moves an activity from a higher cell to a lower one in cell_id order', async () => {
+ await removeTestRows(databases[0]!)
+ const [cellA, cellB, cellC] = cells as [typeof cells[0], typeof cells[0], typeof cells[0]]
+ const store = new RelayAssignmentStore(databases[0]!, () => 100)
+ await store.reconcileCells(cells)
+ for (const cell of cells) await store.recordCellHeartbeat(heartbeat(cell))
+ await store.setCellEnabled(cellA.id, false)
+ await store.setCellEnabled(cellB.id, false)
+ const assignment = await store.assign(identity)
+ expect(assignment.cellId).toBe(cellC.id)
+ await store.setCellEnabled(cellA.id, true)
+ await store.setCellEnabled(cellB.id, true)
+ const activityId = 'splice:rebind-inventory-postgres'
+ await store.acquireActivity(identity, { activityId, kind: 'splice', cellId: cellC.id })
+ // The migration makes B authoritative; the lease still sits on C.
+ const migration = await store.startEvacuation(identity, cellB.id)
+ expect(migration.targetCellId).toBe(cellB.id)
+
+ // Hold B elsewhere. An ordered move locks B first and queues here holding
+ // nothing else. Locking C first (the old lease's row, as an unordered move
+ // does) or the whole inventory (which takes A) shows up as a held row.
+ let releaseRow!: () => void
+ const rowReleased = new Promise((resolve) => {
+ releaseRow = resolve
+ })
+ let rowHeld!: () => void
+ const rowHeldPromise = new Promise((resolve) => {
+ rowHeld = resolve
+ })
+ const heldWhileMoverWaits: string[] = []
+ const holder = databases[1]!.transaction(async (transaction) => {
+ await transaction.queryLocked(`SELECT * FROM relay_cells WHERE cell_id = ?`, [cellB.id])
+ rowHeld()
+ await rowReleased
+ for (const cell of [cellA, cellC]) {
+ try {
+ await transaction.queryLocked(`SELECT * FROM relay_cells WHERE cell_id = ?`, [cell.id], {
+ failIfUnavailable: true
+ })
+ } catch {
+ heldWhileMoverWaits.push(cell.id)
+ }
+ }
+ })
+ await rowHeldPromise
+ const move = store.acquireActivity(identity, { activityId, kind: 'splice', cellId: cellB.id })
+ let moved = false
+ void move.then(() => {
+ moved = true
+ })
+ await new Promise((resolve) => setTimeout(resolve, 250))
+ expect(moved).toBe(false)
+ releaseRow()
+ await holder
+ await move
+ expect(heldWhileMoverWaits).toEqual([])
+
+ const reservations = await databases[0]!.query(
+ `SELECT cell_id, reserved_requests FROM relay_cells
+ WHERE cell_id IN (?, ?, ?) ORDER BY cell_id ASC`,
+ [cellA.id, cellB.id, cellC.id]
+ )
+ const reserved = reservations.map((row) => [String(row.cell_id), Number(row.reserved_requests)])
+ expect(reserved).toEqual([
+ [cellA.id, 0],
+ // Migration grant plus the moved splice, as in the SQLite origin-scoped
+ // reservation case: the lock change did not alter accounting.
+ [cellB.id, 6],
+ // The sticky grant stays on the source until the migration completes.
+ [cellC.id, 1]
+ ])
+ }, 15_000)
+})
From fb69f00b65bb3096ae58010a780fcbf771b7ea17 Mon Sep 17 00:00:00 2001
From: Neil <4138956+nwparker@users.noreply.github.com>
Date: Fri, 4 Sep 2026 01:34:47 -0700
Subject: [PATCH 32/49] fix(hosts): resolve a folder workspace's SSH host from
the repo's host, not its raw connectionId (#18598)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
* fix(hosts): resolve a folder workspace's SSH host from the repo's host, not its raw connectionId
`resolveFolderWorkspaceHost` inferred a workspace's host by reading
`repo.connectionId` directly. SSH ownership has two spellings on a repo row, and
a row carrying only `executionHostId: 'ssh:'` has no `connectionId` to
read — so it counted as a local repo and the workspace resolved `{ kind: 'local' }`.
That is an execute-here answer for a workspace whose files are on an SSH host,
the #11163 class, and it fires on a well-formed row.
Resolve the host first, then read the target off it. Every other row keeps its
existing contribution, including a `runtime:` row's nested SSH target: that
target is not this client's to dial, but narrowing it here would be a second
behaviour change riding on this one. The runtime branch above still answers
`local`, and now says so — `FolderWorkspaceHost` has no runtime variant, and
widening the type is its own change, not an oversight to be silently corrected.
Three smaller items that stand on their own:
- `resolveWorktreeExecutionHost` gains a `malformed` reason distinct from
`unknown`. `unknown` (nothing carries the id) is a verdict the launch path may
legitimately dispose of as a plain local folder; `malformed` (the row named a
host that cannot be parsed) must fail closed. One word for two situations is
the shape that lost the distinction in #18006. The strict read is private to
that module: `getRepoExecutionHostId` stays the answer everywhere else, since
its fall-through to `local` is harmless for the grouping, label and index
callers that are nearly all of its ~340 call sites.
- `readAllWorktreeMetaForRepo` / `readWorktreeMetaForRepo` replace four
open-coded copies of the same host-qualified read (the F7/F8 lockstep shape).
- `getExecutionHostLabel` answers 'Unknown host' rather than 'All hosts' for an
id that names no host. Showing one unroutable row as though it were on every
host is wrong on its own terms. Plain English like every other label in that
module, none of which resolve through the renderer's i18n catalog.
* fix(hosts): resolve the host in candidate selection too, not just in resolution
The first pass fixed how a repo row is classified once it reaches
`resolveFolderWorkspaceHost`. The candidate filter decides which rows reach it at
all, and it read `repo.connectionId` raw as well — so an SSH-only row outside the
project-group subtree was dropped before the new logic could see it, and the
execute-here bug survived for the population the fix was for, via a different
path. Found in review by CodeRabbit.
Three repo-row reads had the same root cause, not one:
- the scope-connection filter, comparing a path repo's raw field against the
workspace/group connection;
- the group-connection set, built from group repos' raw fields;
- that set's membership test against path repos' raw fields.
The last two are one comparison with the mismatch on either side, so resolving
only the path side would have reintroduced it from the other direction.
All three, plus the resolution loop, now go through one `getRepoScopeConnectionId`
helper. Non-SSH hosts still fall back to the raw field, so a `runtime:` row keeps
contributing its nested target exactly as before.
The new tests use a repo matched only by path, outside the subtree — the
population every existing test missed, which is why four passing revert-tests
did not catch this. One of them is labelled as pinning the resolver rather than
the filter: under the old raw read both rows came back connectionless and matched
each other by accident, so it survives a filter revert and must not be counted as
coverage for it.
---
.../listing/detected-provider-listing.ts | 7 +-
.../register-worktree-catalog-handlers.ts | 11 +-
.../listing/ssh-worktree-fallback.ts | 4 +-
.../listing/worktree-discovery-metadata.ts | 4 +-
.../host-qualified-worktree-meta.ts | 23 ++-
src/main/runtime/worktree-launch-host-repo.ts | 10 +-
src/shared/execution-host.test.ts | 13 ++
src/shared/execution-host.ts | 9 +-
.../folder-workspace-execution-host.test.ts | 188 ++++++++++++++++++
src/shared/folder-workspace-execution-host.ts | 38 +++-
...worktree-execution-host-resolution.test.ts | 34 ++++
.../worktree-execution-host-resolution.ts | 38 +++-
12 files changed, 348 insertions(+), 31 deletions(-)
diff --git a/src/main/ipc/worktrees/listing/detected-provider-listing.ts b/src/main/ipc/worktrees/listing/detected-provider-listing.ts
index 25c08d262fb..388c825530f 100644
--- a/src/main/ipc/worktrees/listing/detected-provider-listing.ts
+++ b/src/main/ipc/worktrees/listing/detected-provider-listing.ts
@@ -26,8 +26,7 @@ import {
type DetectedWorktreeSideEffectToken
} from './detected-worktree-scan-cache'
import { loggedWorktreeListFailures, warnOnce } from './worktree-listing-diagnostics'
-import { readAllWorktreeMetaForHost } from '../../../persistence/host-qualified-worktree-meta'
-import { getRepoExecutionHostId } from '../../../../shared/execution-host'
+import { readAllWorktreeMetaForRepo } from '../../../persistence/host-qualified-worktree-meta'
export async function listDetectedWorktreesForCapturedRepo(
store: Store,
@@ -40,9 +39,7 @@ export async function listDetectedWorktreesForCapturedRepo(
providerAbort?.signal.aborted
? ({ providerAbortStatus: providerAbort.status() } as const)
: undefined
- const allMeta = isFolderRepo(repo)
- ? undefined
- : readAllWorktreeMetaForHost(store, getRepoExecutionHostId(repo))
+ const allMeta = isFolderRepo(repo) ? undefined : readAllWorktreeMetaForRepo(store, repo)
// Why: only the disconnected fallbacks read this, so keep parseWorktreeId over the whole host snapshot
// off the connected path entirely.
let cachedSshWorktreeMetaIndex: SshWorktreeMetaIndex | undefined
diff --git a/src/main/ipc/worktrees/listing/register-worktree-catalog-handlers.ts b/src/main/ipc/worktrees/listing/register-worktree-catalog-handlers.ts
index d684461a381..4f3055c63a2 100644
--- a/src/main/ipc/worktrees/listing/register-worktree-catalog-handlers.ts
+++ b/src/main/ipc/worktrees/listing/register-worktree-catalog-handlers.ts
@@ -23,7 +23,10 @@ import {
warnOnce
} from './worktree-listing-diagnostics'
import type { WorktreeIpcContext } from '../worktree-ipc-context'
-import { readAllWorktreeMetaForHost } from '../../../persistence/host-qualified-worktree-meta'
+import {
+ readAllWorktreeMetaForHost,
+ readAllWorktreeMetaForRepo
+} from '../../../persistence/host-qualified-worktree-meta'
import type { WorktreeMeta } from '../../../../shared/worktree/meta-types'
const WORKTREE_LIST_ALL_CONCURRENCY = 8
@@ -174,9 +177,7 @@ export function registerWorktreeCatalogHandlers(context: WorktreeIpcContext): vo
if (!repo) {
return []
}
- const allMeta = repo.connectionId
- ? readAllWorktreeMetaForHost(store, getRepoExecutionHostId(repo))
- : undefined
+ const allMeta = repo.connectionId ? readAllWorktreeMetaForRepo(store, repo) : undefined
const sshWorktreeMetaIndex = repo.connectionId
? createSshWorktreeMetaIndex(Object.entries(allMeta ?? {}))
: new Map()
@@ -226,7 +227,7 @@ export function registerWorktreeCatalogHandlers(context: WorktreeIpcContext): vo
})
}
loggedWorktreeListFailures.delete(`${repo.id}:${repo.path}`)
- const metadata = allMeta ?? readAllWorktreeMetaForHost(store, getRepoExecutionHostId(repo))
+ const metadata = allMeta ?? readAllWorktreeMetaForRepo(store, repo)
return buildDetectedGitWorktrees(store, repo, gitWorktrees, metadata)
.filter((worktree) => worktree.visible)
.map((worktree) => stampAndMergeVisibleDetectedWorktree(store, repo, worktree, metadata))
diff --git a/src/main/ipc/worktrees/listing/ssh-worktree-fallback.ts b/src/main/ipc/worktrees/listing/ssh-worktree-fallback.ts
index 5c734d8bcd9..ecf8ab3abc1 100644
--- a/src/main/ipc/worktrees/listing/ssh-worktree-fallback.ts
+++ b/src/main/ipc/worktrees/listing/ssh-worktree-fallback.ts
@@ -9,7 +9,7 @@ import type { GitWorktreeInfo, DetectedWorktree, Worktree } from '../../../../sh
import type { Store } from '../../../persistence/loading-store/store'
import { getRepoExecutionHostId } from '../../../../shared/execution-host'
import {
- readWorktreeMetaForHost,
+ readWorktreeMetaForRepo,
writeWorktreeMetaForHost
} from '../../../persistence/host-qualified-worktree-meta'
import { getRepoOwnedWorktreeMeta } from '../../../worktree-metadata-ownership'
@@ -159,7 +159,7 @@ export function buildDetectedGitWorktrees(
const legacyMeta = allMeta === undefined ? store.getWorktreeMeta?.(worktreeId) : undefined
const metaById = allMeta ?? (legacyMeta ? { [worktreeId]: legacyMeta } : {})
const meta =
- readWorktreeMetaForHost(store, worktreeId, getRepoExecutionHostId(repo)) ??
+ readWorktreeMetaForRepo(store, worktreeId, repo) ??
getRepoOwnedWorktreeMeta(repo, worktreeId, metaById, repoOwnerCount)
const worktree = mergeWorktree(repo.id, gitWorktree, meta, repo.displayName)
const detected = toDetectedWorktree({
diff --git a/src/main/ipc/worktrees/listing/worktree-discovery-metadata.ts b/src/main/ipc/worktrees/listing/worktree-discovery-metadata.ts
index b17677ddfcc..3edb8efc1d7 100644
--- a/src/main/ipc/worktrees/listing/worktree-discovery-metadata.ts
+++ b/src/main/ipc/worktrees/listing/worktree-discovery-metadata.ts
@@ -4,7 +4,7 @@ import type { WorktreeMeta } from '../../../../shared/worktree/meta-types'
import { getProjectHostSetupWorktreeMeta } from '../../../../shared/project-host-setup-lookup'
import { getRepoExecutionHostId } from '../../../../shared/execution-host'
import {
- readWorktreeMetaForHost,
+ readWorktreeMetaForRepo,
writeWorktreeMetaForHost
} from '../../../persistence/host-qualified-worktree-meta'
import { getRepoOwnedWorktreeMeta } from '../../../worktree-metadata-ownership'
@@ -44,7 +44,7 @@ export function resolveWorktreeMetaWithDiscoveryBackfill(
// Why: the locator-keyed row is only a stand-in for a missing snapshot, so don't read it when we have one.
const legacyMeta = allMeta === undefined ? store.getWorktreeMeta?.(worktreeId) : undefined
const existing =
- readWorktreeMetaForHost(store, worktreeId, executionHostId) ??
+ readWorktreeMetaForRepo(store, worktreeId, repo) ??
getRepoOwnedWorktreeMeta(
repo,
worktreeId,
diff --git a/src/main/persistence/host-qualified-worktree-meta.ts b/src/main/persistence/host-qualified-worktree-meta.ts
index a9d1c0e8fb1..6267311f471 100644
--- a/src/main/persistence/host-qualified-worktree-meta.ts
+++ b/src/main/persistence/host-qualified-worktree-meta.ts
@@ -1,4 +1,5 @@
-import type { ExecutionHostId } from '../../shared/execution-host'
+import { getRepoExecutionHostId, type ExecutionHostId } from '../../shared/execution-host'
+import type { Repo } from '../../shared/repo-types'
import type { WorktreeMeta } from '../../shared/worktree/meta-types'
/**
@@ -52,6 +53,26 @@ export function readWorktreeMetaForHost(
return store.getWorktreeMetaForHost?.(worktreeId, executionHostId)
}
+/**
+ * The same two reads keyed off a repo row, so the resolve-then-read pair lives in one place. Four
+ * call sites had open-coded it identically, which is the shape that lets one copy drift from the
+ * rest (F7/F8).
+ */
+export function readAllWorktreeMetaForRepo(
+ store: Pick,
+ repo: Pick
+): Record {
+ return readAllWorktreeMetaForHost(store, getRepoExecutionHostId(repo))
+}
+
+export function readWorktreeMetaForRepo(
+ store: Pick,
+ worktreeId: string,
+ repo: Pick
+): WorktreeMeta | undefined {
+ return readWorktreeMetaForHost(store, worktreeId, getRepoExecutionHostId(repo))
+}
+
export function writeWorktreeMetaForHost(
store: Pick,
worktreeId: string,
diff --git a/src/main/runtime/worktree-launch-host-repo.ts b/src/main/runtime/worktree-launch-host-repo.ts
index 4decb7acb56..7db9f4dae18 100644
--- a/src/main/runtime/worktree-launch-host-repo.ts
+++ b/src/main/runtime/worktree-launch-host-repo.ts
@@ -17,7 +17,10 @@ export type WorktreeHostRouting =
| { kind: 'resolved'; hostId: ExecutionHostId; repo: T | null }
/** No row carries this repo id and the worktree names no host — nothing ever named a host. */
| { kind: 'unowned' }
- /** Rival rows disagree about the host; guessing one is the cross-host leak. */
+ /**
+ * No single trustworthy host: rival rows disagree, or the resolved row named one that cannot be
+ * parsed. Guessing is the cross-host leak in both cases.
+ */
| { kind: 'ambiguous' }
/**
@@ -33,7 +36,10 @@ export function resolveWorktreeHostRouting {
const resolution = resolveWorktreeExecutionHost(createRepoRowExecutionHostLookup(repos), worktree)
if (resolution.kind === 'unresolved') {
- return resolution.reason === 'ambiguous' ? { kind: 'ambiguous' } : { kind: 'unowned' }
+ // Only `unknown` — nothing anywhere carries the id — becomes `unowned`, which callers dispose of
+ // as a plain local folder. `malformed` is a row that declared a host and named an unparseable
+ // one, so it joins `ambiguous`: guessing is the cross-host leak either way.
+ return resolution.reason === 'unknown' ? { kind: 'unowned' } : { kind: 'ambiguous' }
}
return { kind: 'resolved', hostId: resolution.hostId, repo: resolution.owner }
}
diff --git a/src/shared/execution-host.test.ts b/src/shared/execution-host.test.ts
index 9905fc5fe2a..fba1d1fee8d 100644
--- a/src/shared/execution-host.test.ts
+++ b/src/shared/execution-host.test.ts
@@ -2,6 +2,7 @@ import { afterEach, describe, expect, it, vi } from 'vitest'
import {
ALL_EXECUTION_HOSTS_SCOPE,
LOCAL_EXECUTION_HOST_ID,
+ getExecutionHostLabel,
getLocalExecutionHostLabel,
getRepoExecutionHostId,
getRepoSshConnectionId,
@@ -169,4 +170,16 @@ describe('execution host id delimiter invariant', () => {
targetId: 'a|b'
})
})
+
+ // "All hosts" is the everything-scope. Answering with it for an id that names no host shows one
+ // unroutable row as though it were on every host, which is the opposite of what it is.
+ it('labels an id that names no host as one unknown host, not as every host', () => {
+ for (const id of ['ssh:', 'ssh:a|b', 'ssh:%zz', 'runtime:', 'quantum:box'] as const) {
+ expect(getExecutionHostLabel(id as never)).toBe('Unknown host')
+ }
+ expect(getExecutionHostLabel(null)).toBe('Unknown host')
+ expect(getExecutionHostLabel(ALL_EXECUTION_HOSTS_SCOPE)).toBe('All hosts')
+ expect(getExecutionHostLabel('ssh:box')).toBe('box')
+ expect(getExecutionHostLabel('runtime:env-1')).toBe('env-1')
+ })
})
diff --git a/src/shared/execution-host.ts b/src/shared/execution-host.ts
index a77d02b3882..bbe55aea1a0 100644
--- a/src/shared/execution-host.ts
+++ b/src/shared/execution-host.ts
@@ -226,13 +226,18 @@ export function getSettingsFocusedExecutionHostId(
: LOCAL_EXECUTION_HOST_ID
}
-export function getExecutionHostLabel(id: ExecutionHostScope): string {
+export function getExecutionHostLabel(id: ExecutionHostScope | null | undefined): string {
if (id === ALL_EXECUTION_HOSTS_SCOPE) {
return 'All hosts'
}
const parsed = parseExecutionHostId(id)
if (!parsed) {
- return 'All hosts'
+ // Not "All hosts": an id that names no host is one *unknown* host, and answering with the
+ // everything-scope label shows an unroutable row as though it were on every host.
+ // Plain English like every other label in this module (`Local Mac`, `This computer`,
+ // `All hosts`) — none of them resolve through the renderer's i18n catalog, so a lone
+ // translated string here would read inconsistently.
+ return 'Unknown host'
}
switch (parsed.kind) {
case 'local':
diff --git a/src/shared/folder-workspace-execution-host.test.ts b/src/shared/folder-workspace-execution-host.test.ts
index 2ccced49907..e1af0a7920e 100644
--- a/src/shared/folder-workspace-execution-host.test.ts
+++ b/src/shared/folder-workspace-execution-host.test.ts
@@ -174,6 +174,194 @@ describe('folder workspace execution host', () => {
expect(resolveFolderWorkspaceHost(state({ repos: [] }), 'fw-1')).toEqual({ kind: 'local' })
})
+ // SSH ownership has two spellings on a repo row. A row carrying only `executionHostId: 'ssh:*'`
+ // has no `connectionId`, and reading the raw field counted it as a local repo — so a workspace
+ // whose files live on an SSH host resolved `local`, which is an execute-here answer for a remote
+ // path. These fire on well-formed rows; nothing malformed is involved.
+ it('resolves a repo that names its SSH host only through executionHostId', () => {
+ const resolved = resolveFolderWorkspaceHost(
+ state({
+ repos: [
+ repo({
+ id: 'repo-1',
+ path: '/work/app/a',
+ projectGroupId: 'group-1',
+ executionHostId: 'ssh:box'
+ })
+ ]
+ }),
+ 'fw-1'
+ )
+
+ expect(resolved).toEqual({ kind: 'ssh', targetId: 'box' })
+ })
+
+ it('mixes such a repo with a local one as ambiguous rather than local', () => {
+ const resolved = resolveFolderWorkspaceHost(
+ state({
+ repos: [
+ repo({ id: 'repo-1', path: '/work/app/a', projectGroupId: 'group-1' }),
+ repo({
+ id: 'repo-2',
+ path: '/work/app/b',
+ projectGroupId: 'group-1',
+ executionHostId: 'ssh:box'
+ })
+ ]
+ }),
+ 'fw-1'
+ )
+
+ expect(resolved).toEqual({ kind: 'ambiguous' })
+ })
+
+ it('matches a scope connection against such a repo instead of calling it ambiguous', () => {
+ const resolved = resolveFolderWorkspaceHost(
+ state({
+ folderWorkspaces: [workspace({ connectionId: 'box' })],
+ repos: [
+ repo({
+ id: 'repo-1',
+ path: '/work/app/a',
+ projectGroupId: 'group-1',
+ executionHostId: 'ssh:box'
+ })
+ ]
+ }),
+ 'fw-1'
+ )
+
+ expect(resolved).toEqual({ kind: 'ssh', targetId: 'box' })
+ })
+
+ it('reads the target off the host, so a percent-encoded id decodes', () => {
+ const resolved = resolveFolderWorkspaceHost(
+ state({
+ repos: [
+ repo({
+ id: 'repo-1',
+ path: '/work/app/a',
+ projectGroupId: 'group-1',
+ executionHostId: `ssh:${encodeURIComponent('box 1')}`
+ })
+ ]
+ }),
+ 'fw-1'
+ )
+
+ expect(resolved).toEqual({ kind: 'ssh', targetId: 'box 1' })
+ })
+
+ // Deliberately unchanged: a `runtime:` row's nested SSH target is not this client's to dial, but
+ // narrowing that here would be a second behaviour change riding on the SSH fix.
+ it('leaves a runtime row contributing its nested connection exactly as before', () => {
+ const resolved = resolveFolderWorkspaceHost(
+ state({
+ repos: [
+ repo({
+ id: 'repo-1',
+ path: '/work/app/a',
+ projectGroupId: 'group-1',
+ executionHostId: 'runtime:env-1',
+ connectionId: 'nested-box'
+ })
+ ]
+ }),
+ 'fw-1'
+ )
+
+ expect(resolved).toEqual({ kind: 'ssh', targetId: 'nested-box' })
+ })
+
+ it('still answers local for a runtime pin, which the type cannot express otherwise', () => {
+ const resolved = resolveFolderWorkspaceHost(
+ state({
+ folderWorkspaces: [workspace({ executionHostId: 'runtime:env-1' })]
+ }),
+ 'fw-1'
+ )
+
+ expect(resolved).toEqual({ kind: 'local' })
+ })
+
+ // The candidate FILTER decides which rows reach the resolver, and it read `repo.connectionId` raw
+ // too — so an SSH-only repo outside the project-group subtree was dropped before any of the above
+ // could classify it. Every test before this one uses a repo inside the subtree, which is never
+ // filtered, so none of them could have caught it (found in review by CodeRabbit).
+ describe('a repo matched only by path, outside the project-group subtree', () => {
+ const sshOnlyPathRepo = repo({
+ id: 'repo-path',
+ path: '/work/app/nested',
+ executionHostId: 'ssh:box'
+ })
+
+ it('survives the scope-connection filter instead of being dropped as connectionless', () => {
+ const scoped = state({
+ folderWorkspaces: [workspace({ connectionId: 'box' })],
+ repos: [sshOnlyPathRepo]
+ })
+
+ expect(findFolderWorkspaceCandidateRepos(scoped, 'fw-1')).toEqual([sshOnlyPathRepo])
+ expect(resolveFolderWorkspaceHost(scoped, 'fw-1')).toEqual({ kind: 'ssh', targetId: 'box' })
+ })
+
+ // Pins the resolver, not the filter: under the old raw read BOTH rows came back connectionless,
+ // so they matched each other by accident and this case survived the filter either way. The
+ // legacy-vs-unified pairing below is the one that discriminates.
+ it('survives the group-connection filter when the group is on that same SSH host', () => {
+ const scoped = state({
+ repos: [
+ repo({
+ id: 'repo-group',
+ path: '/work/app/group',
+ projectGroupId: 'group-1',
+ executionHostId: 'ssh:box'
+ }),
+ sshOnlyPathRepo
+ ]
+ })
+
+ expect(findFolderWorkspaceCandidateRepos(scoped, 'fw-1')).toHaveLength(2)
+ expect(resolveFolderWorkspaceHost(scoped, 'fw-1')).toEqual({ kind: 'ssh', targetId: 'box' })
+ })
+
+ // Both sides of the group comparison are resolved, so the legacy spelling on one side and the
+ // unified spelling on the other still match.
+ it('matches a legacy-spelled group repo against a unified-spelled path repo', () => {
+ const scoped = state({
+ repos: [
+ repo({
+ id: 'repo-group',
+ path: '/work/app/group',
+ projectGroupId: 'group-1',
+ connectionId: 'box'
+ }),
+ sshOnlyPathRepo
+ ]
+ })
+
+ expect(findFolderWorkspaceCandidateRepos(scoped, 'fw-1')).toHaveLength(2)
+ expect(resolveFolderWorkspaceHost(scoped, 'fw-1')).toEqual({ kind: 'ssh', targetId: 'box' })
+ })
+
+ // A `runtime:` row's nested target is still read from the raw field, so it matches a scope
+ // connection exactly as it does today. Pinned so the carve-out stays a decision.
+ it('leaves a runtime row matching the scope connection through its nested target', () => {
+ const runtimePathRepo = repo({
+ id: 'repo-path',
+ path: '/work/app/nested',
+ executionHostId: 'runtime:env-1',
+ connectionId: 'box'
+ })
+ const scoped = state({
+ folderWorkspaces: [workspace({ connectionId: 'box' })],
+ repos: [runtimePathRepo]
+ })
+
+ expect(findFolderWorkspaceCandidateRepos(scoped, 'fw-1')).toEqual([runtimePathRepo])
+ })
+ })
+
it('reads each repository membership once while collecting candidates', () => {
let membershipReads = 0
const repos = Array.from({ length: 32 }, (_, index) => {
diff --git a/src/shared/folder-workspace-execution-host.ts b/src/shared/folder-workspace-execution-host.ts
index dc4dc80c51f..0aee75de543 100644
--- a/src/shared/folder-workspace-execution-host.ts
+++ b/src/shared/folder-workspace-execution-host.ts
@@ -17,7 +17,7 @@ import type { ProjectGroup } from './project-group-types'
import type { Repo } from './repo-types'
import { isPathInsideOrEqual } from './cross-platform-path'
import { getProjectGroupSubtreeIds } from './project-groups'
-import { parseExecutionHostId } from './execution-host'
+import { getRepoExecutionHostId, parseExecutionHostId } from './execution-host'
export type FolderWorkspaceHostState = {
folderWorkspaces: readonly FolderWorkspace[]
@@ -36,6 +36,24 @@ export function normalizeConnectionId(value: string | null | undefined): string
return value?.trim() || null
}
+/**
+ * The SSH target whose filesystem holds this repo's files, or `null` for anything else.
+ *
+ * SSH ownership has two spellings on a repo row — the legacy `connectionId` field and the unified
+ * `executionHostId` — so reading the raw field sees only one of them and a row carrying only
+ * `executionHostId: 'ssh:'` reads as if it had no connection at all. Every comparison in
+ * this file goes through here: the candidate filters decide which rows reach the resolver, so
+ * reading raw in either place drops the row before the resolver can classify it.
+ *
+ * A non-SSH host falls back to the raw field so a `runtime:` row keeps contributing its nested
+ * target exactly as it does today. That target is not this client's to dial, but changing it is a
+ * separate defect with its own reasoning — see the note in `resolveFolderWorkspaceHost`.
+ */
+function getRepoScopeConnectionId(repo: Repo): string | null {
+ const host = parseExecutionHostId(getRepoExecutionHostId(repo))
+ return host?.kind === 'ssh' ? host.targetId : normalizeConnectionId(repo.connectionId)
+}
+
function getFolderScopeCandidateRepos(args: {
folderPath: string
projectGroupId: string
@@ -59,18 +77,18 @@ function getFolderScopeCandidateRepos(args: {
if (args.connectionId) {
return [
...groupRepos,
- ...pathRepos.filter((repo) => normalizeConnectionId(repo.connectionId) === args.connectionId)
+ ...pathRepos.filter((repo) => getRepoScopeConnectionId(repo) === args.connectionId)
]
}
if (groupRepos.length === 0) {
return pathRepos
}
- const groupConnectionIds = new Set(
- groupRepos.map((repo) => normalizeConnectionId(repo.connectionId))
- )
+ // Both sides resolved: comparing a resolved path repo against a raw group read would reintroduce
+ // the same mismatch from the other direction.
+ const groupConnectionIds = new Set(groupRepos.map(getRepoScopeConnectionId))
return [
...groupRepos,
- ...pathRepos.filter((repo) => groupConnectionIds.has(normalizeConnectionId(repo.connectionId)))
+ ...pathRepos.filter((repo) => groupConnectionIds.has(getRepoScopeConnectionId(repo)))
]
}
@@ -102,6 +120,12 @@ export function resolveFolderWorkspaceHost(
}
const explicitHost = parseExecutionHostId(workspace.executionHostId)
if (explicitHost) {
+ // A `runtime:` workspace deliberately answers `local`, and `FolderWorkspaceHost` has no runtime
+ // variant to answer with instead. That omission is known: a runtime environment's own server
+ // normalizes its work to `local`, and the nested SSH target on such a row is addressable only as
+ // the pair (environmentId, targetId) — handing it to this client's SSH table would dial a
+ // same-named box in the wrong namespace. Widening the type is its own change, not an oversight
+ // here.
return explicitHost.kind === 'ssh'
? { kind: 'ssh', targetId: explicitHost.targetId }
: { kind: 'local' }
@@ -114,7 +138,7 @@ export function resolveFolderWorkspaceHost(
let hasLocalRepo = false
const connectionIds = new Set()
for (const repo of candidateRepos) {
- const connectionId = normalizeConnectionId(repo.connectionId)
+ const connectionId = getRepoScopeConnectionId(repo)
if (connectionId) {
connectionIds.add(connectionId)
} else {
diff --git a/src/shared/worktree-execution-host-resolution.test.ts b/src/shared/worktree-execution-host-resolution.test.ts
index b82cea476eb..b31281fe925 100644
--- a/src/shared/worktree-execution-host-resolution.test.ts
+++ b/src/shared/worktree-execution-host-resolution.test.ts
@@ -156,6 +156,40 @@ describe('resolveWorktreeExecutionHost', () => {
it('reports an unknown owner distinctly from a conflicting one', () => {
expect(resolve([], { repoId: 'r' })).toEqual({ kind: 'unresolved', reason: 'unknown' })
})
+
+ // `unknown` is a verdict the launch path disposes of as a plain local folder, so a row that
+ // declared a host and named an unparseable one must not share the word — it has to fail closed.
+ it('reports a row naming an unparseable host distinctly from an unknown one', () => {
+ for (const executionHostId of ['ssh:', 'ssh:a|b', 'ssh:%zz', 'runtime:', 'quantum:box']) {
+ expect(resolve([{ id: 'r', executionHostId }], { repoId: 'r' })).toEqual({
+ kind: 'unresolved',
+ reason: 'malformed'
+ })
+ }
+ })
+
+ it('does not recover a host from the connectionId such a row overrode', () => {
+ expect(
+ resolve([{ id: 'r', executionHostId: 'ssh:a|b', connectionId: 'openclaw' }], {
+ repoId: 'r'
+ })
+ ).toEqual({ kind: 'unresolved', reason: 'malformed' })
+ })
+
+ it('still resolves every row that names a parseable host', () => {
+ expect(resolve([{ id: 'r', executionHostId: 'ssh:box' }], { repoId: 'r' })).toMatchObject({
+ kind: 'resolved',
+ hostId: 'ssh:box'
+ })
+ expect(resolve([{ id: 'r', connectionId: 'box' }], { repoId: 'r' })).toMatchObject({
+ kind: 'resolved',
+ hostId: 'ssh:box'
+ })
+ expect(resolve([{ id: 'r' }], { repoId: 'r' })).toMatchObject({
+ kind: 'resolved',
+ hostId: 'local'
+ })
+ })
})
it('ignores an unparseable host id rather than treating it as a host', () => {
diff --git a/src/shared/worktree-execution-host-resolution.ts b/src/shared/worktree-execution-host-resolution.ts
index 00b66b7f11c..0172106aee0 100644
--- a/src/shared/worktree-execution-host-resolution.ts
+++ b/src/shared/worktree-execution-host-resolution.ts
@@ -54,7 +54,29 @@ export type WorktreeExecutionHostResolution =
/** Display metadata only. The decisions are `hostId` / `connectionId`. */
owner: T | null
}
- | { kind: 'unresolved'; reason: 'ambiguous' | 'unknown' }
+ /**
+ * Three reasons, not two, and deliberately not collapsed. `unknown` (nothing carries the id) is a
+ * verdict the launch path may legitimately dispose of as a plain local folder; `malformed` (the
+ * row named a host that cannot be parsed) must fail closed. A vocabulary that cannot express the
+ * difference guarantees it is lost at the first caller that switches on it — the same shape as
+ * #18006, where one word had to stand for two liveness situations.
+ */
+ | { kind: 'unresolved'; reason: 'ambiguous' | 'unknown' | 'malformed' }
+
+/**
+ * The owner row's host, or `null` when the row names one that cannot be parsed.
+ *
+ * Module-private and deliberately not a second exported reading of a repo row: only this resolution
+ * needs the distinction, because only this resolution is routing. `getRepoExecutionHostId` stays the
+ * answer everywhere else — its fall-through to `local` is harmless for the grouping, label and index
+ * callers that make up nearly all of its ~340 call sites, and is wrong only when the value decides
+ * where work runs.
+ */
+function resolveOwnerRowHostId(row: ExecutionHostOwnerRow): ExecutionHostId | null {
+ return row.executionHostId?.trim()
+ ? normalizeExecutionHostId(row.executionHostId)
+ : getRepoExecutionHostId(row)
+}
export function resolveWorktreeExecutionHost(
lookup: ExecutionHostOwnerLookup,
@@ -80,9 +102,13 @@ export function resolveWorktreeExecutionHost(
if (match.kind !== 'resolved') {
return { kind: 'unresolved', reason: match.kind === 'ambiguous' ? 'ambiguous' : 'unknown' }
}
+ const hostId = resolveOwnerRowHostId(match.owner)
+ if (!hostId) {
+ return { kind: 'unresolved', reason: 'malformed' }
+ }
return {
kind: 'resolved',
- hostId: getRepoExecutionHostId(match.owner),
+ hostId,
connectionId: getRepoSshConnectionId(match.owner),
owner: match.owner
}
@@ -115,12 +141,14 @@ export function createRepoRowExecutionHostLookup getRepoExecutionHostId(repo) !== ownerHostId)
+ const ownerHostId = resolveOwnerRowHostId(owner)
+ return rows.some((repo) => resolveOwnerRowHostId(repo) !== ownerHostId)
? { kind: 'ambiguous' }
: { kind: 'resolved', owner }
},
+ // A row naming an unparseable host matches no host, which is what stops a worktree on a real
+ // host from adopting it.
byHost: (repoId, hostId) =>
- rowsFor(repoId).find((repo) => getRepoExecutionHostId(repo) === hostId) ?? null
+ rowsFor(repoId).find((repo) => resolveOwnerRowHostId(repo) === hostId) ?? null
}
}
From 01d7228b7e971749b1f3da06f8e9a6be18e76b94 Mon Sep 17 00:00:00 2001
From: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Date: Fri, 4 Sep 2026 01:44:52 -0700
Subject: [PATCH 33/49] docs: add WeChat group 9 QR code
Adds group 9 QR fallback QR codes and capacity guidance to the README variants.
---
README.md | 4 ++--
docs/assets/wechat-qr-group9.jpg | Bin 0 -> 395815 bytes
docs/readme/README.fr.md | 4 ++--
docs/readme/README.ko.md | 4 ++--
docs/readme/README.zh-CN.md | 4 ++--
5 files changed, 8 insertions(+), 8 deletions(-)
create mode 100644 docs/assets/wechat-qr-group9.jpg
diff --git a/README.md b/README.md
index 7a3cbe2360c..2ae59035da8 100644
--- a/README.md
+++ b/README.md
@@ -238,9 +238,9 @@ Pair with your desktop app to monitor and steer your agents from your phone.
- **Discord:** Join the community on **[Discord](https://discord.gg/fzjDKHxv8Q)**.
- **Twitter / X:** Follow **[@orca_build](https://x.com/orca_build)** for updates and announcements.
-- **WeChat:** Scan to join the Orca community WeChat group 8.
+- **WeChat:** Scan to join the Orca community WeChat group 8. Group 8 may be full; if so, scan the Group 9 QR code instead.
-
+
- **Feedback & Ideas:** We ship fast. Missing something? [Request a new feature](https://github.com/stablyai/orca/issues).
- **Privacy:** See the [privacy & telemetry docs](https://www.onorca.dev/docs/telemetry) for what anonymous usage data Orca collects and how to opt out.
diff --git a/docs/assets/wechat-qr-group9.jpg b/docs/assets/wechat-qr-group9.jpg
new file mode 100644
index 0000000000000000000000000000000000000000..2bf46a28c3d464682461ffe47a0bde9a797fd1e8
GIT binary patch
literal 395815
zcmce;2~<Gw4frE!k&CD%+IBsco%HF~8^qI4lE?;qVbNBG{yK(cDe*iu(
zBsAGW-v!^gcHQnH>$jb_sN@^0w&&>28`Muecvju2tbfb}+-Z^^RECQ0Zd(Iuhz)2K?Qsw?wy_j5;+
zDzHl6niXCch~CGU3z{o;^(J(p^`T%vg-@?*{MT0ADv+GRzQU{lWWTQhi749aUe>o_
zn*8!2Kx@}zEj`3Z2Ux4Xp-2EDRN=4uiIU^_tH5I$04Z9xAiI&V3Yade0&jBeBA}RM
zd5zO5kQu%Te8Ac&Nl!$r0v`KTfgY<>;7f6?qI_5LDp3A#75JFF3VgfNr?3ia@mU3i
z&(OY}Uj?M4Q^5by0I1hjfmS=r#5=^Olh_oz^GZaTs7hYMjA~A-0&%8sLsDGW1WRQc
zQkg`=l@pDFf|&C^EeqCtZl2u0?0!mRp4nPAc{6*>+R!G?C=6(
zH8$^Kt*}w%DADU;Te?1?m5}3?U61~@J_Dcc>4h!58n{0!fZCy2_Ws**zc2MM425|l
z_$P+xFViuYS6&6;{e+IIz%j*d^2-X7R1J`xObM=XCwsN}Xp5BYTVtOy+oq{I!MfBE
zDAs2%lf#LU6u#h?W)YfCiYm&9L8PE{&5R-)r-;utR5l&8yc_v}KUXTbh!WUp&(I`E~(wDLiRHshgX-jiH}0hkd>
z!A;KJ;kL=wg2hSJ1fkd8uby~Goz`D$8X}gp;T+j
zHAGIPh5D&uwl;QdGFYqdG{m%JPIDFY?JA&a>Wnxwun-pQ?=d9b2CV|loX}JQ=ej3f
zFAgr*IDBRQ;_(hI*7&g}JSr;O+I+aK-^uz**EXvXrQacxU=>J5@0~@!_XMi4C_J4B
z%@qw6^hq@Z&Hf#hZTMl%3DpkCbi?+?vKbGr?hv9n*O~OU{l50Cf6NT*pFr4hS+~-8
zm^Pkyp%A4e@8#<>9T&*7hrMJtZ*Ge^^!0Q
z-7nP@gr>j+rHBb!99>O*;A)@-YmFSMCq}lkHQVKIi~Ksm8J#Ihw3!_ez9GlPgJapj&)-L0$6KcC;}>
zn8a=lxbtxQsw2UF``e20DOXF%Ht#W~?aUgdA1iFyAF#LJb+_L=+@u6P64rNM=5dfK
zW)pQc#ZlfSL5TU%T~rm3LCPvn&0@$$WLdJZRbYcSqm<@P3s?pGXOO-U+ygP|#dtLR
zFNvuHFta9r29pVF;vUQvialf|@_KE#ss6=uHGygdMF>{`75uh9iz8T@7P?+xRvEG+$;8fw+b*=H7vpI0aN}wa&8rvSH;ZpZ>U8-ny;e;
zjEH};xEnDKr~?-iY}Wzm3ZHqnxmS{NvzfQDD*{-x#a%?4!_XYj%-O(|Bi|~u9W?E^
zbZ*divhl@W=&S^*+Czx)iz1l1nP=SLbdy3zMj>|kol}bnAAAmh>B&;%Djek2Dc2Gq
zVks93Wf>AIn|i1}422-@!bi*Jmvm6|Q!C9CPxA5N9`t99oI9`4FUxy7Gd`mWI`2$V
z=!J)cJV|5kl@HMuZ_^kC@*U|(NNIPUQpz~w*WM7G_}Lm+LXNunLKIf?VqC8QWPW}`
zBM7aC(I)cOE4rOYHcgSgTejh7&&;#-Js+>&uN=91&isNB>n>R%$xDq&1IrqjJyxDU|s89O!4YD`E@QcI(tmU9honb*ko693^w}4
zxLrVaw8$Q_@Uo+CzPb#uZovk^lbT)3dym8M`IKEy0{*7oO}tufSVTk^CO$ugH_Y4c
zccxH?*H6}Z@b`nE?N4p3tJ=^TMH%g?Kk?aUMKB|&i6nG-S!uuom+yt@y0vG1_)Fu*
z=8pHnE9X=3(KJ)nx~Z@4&*(PpJ)GZt@UrVoU~vp5#>on$n1Y!|akIoq8q#lSqlRrB
zTFNH4K~9r0i!&V43Ui)UkK?7mB_00e(UI>a?w6jhsHm|2Vd4r-f6mp5n4l-LprOZQ
z-MN(GAm+X`t#b68^K)z-*?{fIO=glqyCW(pD$RC;j&1C^(3IAulG1GzQ8}(a+R1h8
zKeMN)-w+7T*(J+l1@tOUHKZq5UjY>-{o^TC`J-kZtE$Mc0$kc6QPYWWa{Fhd?xl;c
ziAznbId%a5IcQIk6L9s`UF-AT_c4%*umh28nB-MJJ^ndWOL`be5NQS6`h(#A`XA<{
zGc63o%>E&;n7!gtVEVO0MP;E{g{41_yWaT4y0^CHg`#68e~sLNV#@L5%Jro?p;B-a
zIQf#5aJGK9i3l!tjBn@o51H(G)lfg+I~X>mgROlck0Ka$C*J1VdHF=e>FLA$Y0A2Q
zA*3@=)rn0KTQtzwxO*%$^zP{fBIn(VWx~%*O2R4)_jfq!gFn`V*GH~@P|`{W?K}Ob
zqUnp{kvNWRmew%<(}i5ya!h(;pu#LY3A3TaKx)GwEjVQPJtt4TMjL`T*-<47HE-j%
zLNP0UdYG$O#N!DkJlRH%e%qDli_B1~xqD_J{b!aAdOzMmdznp(<*UngP2I`7&$Az{
zok&lR8;g7%5z9^x@e$~!!6j3c=K21c^HsBDsqcqf&SxAM_2Hct8Mz&9yz3(3x2^&y
zG?lyoe<%yA7h14O!k(Js^?wxcV<^r7EvCn9SEc53=SS_@u9@{5kHVf>?|r3DbG;Hb
zwXUsx0t|uSuTg3IEsK{$2JBTJhOQwE6O06z7)^4ey5M9ddVl|oM#i;z*V5+!nfqiZ
z@4b#4dT(35-reKw1N8^CS}j(PljxFwfG3Eo1oU=ff^nrvXTMt)b(08rPsiQ+2VSb}
zv|qdSt;{fmYozix0YgpYQbeG^h&VC)Zm7?8SvD&GsXJw^cL;QRU3H&&6l^IKSIzZq
zBfAMQQk(2@XE44^F9)iqt>t*_?VsrpVw;nk&uz1F%SWSgz6=Y8y=zY#+NESLNsp&(
zLu=tY@bqmhNHSmWv9N@vBg32*^~RYkDD!de%6o&jb*`E6?Bc85p2qDz*#_-OJAJCZ
z$$q))(B0sFS@x)wY>6>-AhVP1TQ-fX&FG4DJ2PZd5vF_W(m=~r%88bNC1do;hR?UF
z=1@e6g&Fh1;}4ALtDa+C-A0;Q&g^{ghp3?$h`=T{=?htFX1Vh1B38V$-2`79kGm-!
z0Qt!bZzncSoF$&MigxIpUqM$9{7XD!RoEs5iz%^Z-~wtJ>Kjr*kG0RiR$YkzeiC%0
zQ&HsEHr6JpCfTlC4-LglrUv%1$G3vvag6S4Kb&YilPdRS#%C3ebnX56hHv^1-M_uL
z>8JPa|J=?v`X(1S%%AIrgMuh3hqXnH^n&umIkFO|lEjQW;NqoIhK#3fhJJz_pk7(N
z$NQ0j`a3=rBD0KkR=l3tfc2ST+zz7;FF@B~ktVx-OWwFA)ZqPb#1@_SzXZqs|BnYM
zB4C-X`8fLDS)obQ@9<`1#>G)Ogx(QtFBPklj-(!jW_TZU!1LE(*MW1}j(-*mRDGyV
z$_bdO%+^o>y(r+w`^maUSPgjC^h5N=?lTO|aYBrU%6IF54Z6bq?UM(`=%@>n=
zLb^Icc->H(V3gX^Bt97R+^boNy-}}Xi=5l;>+4(S?tzc=eq5AQgs*!ttbHfjm#CT@
z7Z>;YK@$GvO?=RUQ6A58v}>ovaG&GO!zYu}PnSG@t`PK33;hqk1^g3tDQsA~#dcTB
z2K`;Jt!wo*=$BxI?_%E6%QqjFr8gav=)utyR=y8E4x~5?SSv|v$rv_{Q6C38TKBGE
zFY=M5s4{XsTkmk6YGFlTS~h7Df2&k^60Wq2P885{Ao0;pRrC@tE-
ztM!T;Wy^)7eO8xlIGlgz&jCf%Mke_1RK*X^&mD48XkJ;rRIH@hGR4
zyBnAd@m^6`gv5u0sFa8Y`I*Xs2cE}$pO^=HNVdQ6V8`RZ(%pA4zobLvDz95M^kD*M
zkP)#nZen5MR4sH#oZT-)2=I$^I3x=l*9Kk4{K-GD8v?d@jGSD-#Pv-z9QRhl61)$d
ze^m6~7Me}?0MM?kh-05Y#ixd($Ar;3XGO?(sa+8+paG7;mc}f5-TKX%k(cV8^ty1u
zq75&^3%gtf!c>kIFF_^S6rscHy#(vD|R43
z?82!3X_0wC0Xg9;riR`%e`-4OW0rpz{?mniss@zr(Z7(IE!EUm1#WY~qEvLs=mh@yZXUw^pnnB8$
zQoP)>vCqg%HFAZwp)xFGR#)XLXVJ^4;-pv7JR{qZeLntFEu;Qt0OLTLsPfSOdsA2w
z8Myk-*+SOlDIbG@#Z_SB*d#sqwh_gy0KYPnou1UB3oiSJu=gnT{=`zR_pQWA$g6$D
z;gJxtIS|YX)ZX7*IK9&HG?-B0;r0HMg52g`UXoT#l5V55$+f|BshW;ldASI>#BM;w
zht947??%NzBKxI;1V)ys_M7GQmZOl}cTCP|ce~R-Y-VU`FC;fzEDx0L73ozRwUNK7
znrV1aUHaZT(1Po>?Jo5qt(BXj*Jp--_JwsemP|!DJga~M*j1n{!uoi*&vae5XXVfP
z%Xwn@^Olqn?AnY-ubDfWhAm^=dxA(tf7ZSF>T7jCQQAHAFEge%Q}yl=y9Ftt%{NPS
zlCd3`dg(&iM%V{Fm*q5B3300AgmX2+48&G4%@cYhP_^*rTk1iC$qG#EAAf|LWC?
z!NqO~A;nN)m^2@h`$XAWOHr09Og^wDP0R1JZ``&6c}W^SCu1
zB{1cwt{~k28;8H;J3$$}3;8ak!#q+Wl^`V21mPVreQ(kBx~DhSH&)=$IwEXLrL|4r
z46o^e$+MDao@!TO(+PCy&4^&Lx(xGICLj~$t&O|_A0Q0oWkb)Rx`NkWu10>oKV&995Z*6X^JVBd7ZkyRN>B+$eRKb#gg$kBHkAwb0K^A*K>6A*lXI7?fWL(M}CB5lYlfx|-Bk
zpfm4v_QfE{+E{p;ijuAY87;fO4?Q262BtFK=V*y6lB|#KFDp7jpsr{pnlpb-<7vObS+My3~5e+NFJdXYwC^9Z>wHx(dX4Qq(4T5JHl4
zd#1!5OpiBB;UnY&EB-MIogUN)PPqvKyMC0g3TW)B2cJ9`|w-2#+WJU0wZ9bz{#2
zM;K-0me%h%z9$#SMI<&^!91Gglr}fc+T&S(+yG=HnU={^+I@6Ho1fmS2(XkVJ?C}%
zJr7tJPA@$XQaqFCT6oiXIAbMxi6PE-QtLzt-t^$~)N7=m|2xkDQ)4U1wARWT9q=BuN9U_QVlHIXQ#+bx{W0kzYw8+&LgLFVP(Iy7?Pi~
z_4;ti!=k)UTVb=N*`g1&u<)dk6T#c->IJMV)0sidxq&b8Z~-fC={HNKzrxX2t0h3G
zf>m=EqMhhxO{8LKeDd?BG94b5kc%cKi(Uz%<}I9Ss+&Z5Or`m)A{@JqpuN;negs=K
z^UIXBM$)s6l0ZaoxL$hxbmaB3b&=5pM}SZmdj?x1FjPb28$^=pS7#&j9+Y1thn$3h
ztH_a)mj`YyDKm_JLfyw87N?f7wY^^Mk3Qy_jvZv`Xf5gNj}aLyED)dgs{S=Kp$Fra
zwI^X8Bej4Za%6E8O;*pQR@Su?gF0-zJ%bgC-bcI7#$FI9C3?=fa+kWN%CE&(HTzv1
z3WCXO+@K$Ug@S_1iP7Av3P78Z;@HuRaL@Hdh(+NoNNgN&y{R1n2Rs7kWW^dAWG
zlHgyXS73y5kxXK@!2UAE4>lcW#wxIB$`X}egrfOjMylE~o>*#R_!W4JK`9+9CFyP<
zZ99Q$&Wmnl9GmsoFEXim4kff}7nCgX-o7dg)UL}94xvSb1RqfP8za@thJE%0xMiPd
zyH)M?!t?S93
zdQSJND=s&RPlDo~1=tZMaUi%H$FM=NBiCnu`T}e1TM4=NI!|0cD+q4E5~L@d%{*E5yj7b=u)H}moc(;-Gqs{8KHoRQrf{rj`0E*+yR_QS;I2)<
z>`i#}i+j4%FFr2_?o8gUpQ^s@l)b{0%j>Sdo^Q=vOz*55u?yC?IaQh*`8
z{A!CyjRfQ~n2BD}yXCg<7F}5ZD~Pq#E^iinpjvbvQ|rdoGN?|2U7kBfY9?SM;?$a8
z#(u1zjEG4wY9dUiDxGcQRL)mbl!xLp-|@?bJ^xm=MZUNHH=Fw}Quq4}xm!L?tARP8
zAk+wQGGnewI7);Oep@$aUR?}1c?{28Z()K7?N&~BRFmT5UR6SP*lqbvk#oyaT5xdD
zUeA2|pV5Sy!J9Iy1{I~(VGx2??U%xFLhNnY2Qy4{g6zGUyoT203CAg}XMt}~;=g2-
z=xuT~rZ%CS&Ld1-S?XZLP<8RYyw^oJXPnpDE%IqB^=21Av^F12=HNo8B-#$g4YT^q
zxWc%FlXH2^jGS2a#LeU`@{h_=SgUfT`vW8Id!nRjIpS5oAzA*}X8$U1Syt}LF|!&w
zVzSpurP=WeLzcDF=xrkZIB0J+o=q&D5sTdr`xurv&vu>(Kd=q~rLJR5>%bN_DTXgo
z`4vv}1sJaar{CkG=ir?=brnw5>7O92IKWiC@LD}g!hL}AecCECj)mqbptg>+qIm!o
zFtN`|{@^7G`VD@!?-iW!bBN(Drgp3X?X4tv+#5J!XbQJh5QLKNSD`JMTu@jAw&vE$
z^n+=l!(OYvv!Cu_o~J_wSgE?<1YN3LE51`Ldq&-8Ulvez1WIrP6H=tA!l+bjtR|${
z;b~j<^3fUY{Gzwr#ri}XK7)6sw(#n=2PV&oa3@04Z8bJ3EI$TlSzS`y^od5!
zF;FQQHoK6gWeMD`WOmg5~(324funoj`Vwd{;-c4Zb2B+?=Cj0z1Gta
z9hz2h5wvu=|0(g+Qcgw
zauv8qgAO=$E{ra8;5JjXiU@A9LZN71RTm=OlK6fVP+hEW?Y@%3ztTs9$QYOc?^BFRjPL&>iM$U{z3yqcBMN0P=y?1>!4pJ`W;WP6+$7fn}
z2c$OzhGRi(%8R2|wSZ&qi^>8Oz60I0-+jvq>Ikl;YhF>}jFaw`{nF$wALh5-yU1vr6i)oteM3ZH~GMkia;prGR_t`qUYC3Z`#URG(bY^t;sl$N{
zr|PzzCx0kX&x!G@nf?9
zxw7;sG%IX|O;3B~r-YWx6f;qJdvo~Ne&;3T{f67sY@AAWB!`2pL
zjCh6XR7>{~VUnzo#Q~X@Eh4j*o@5jE$(9XE?iJ`#R2UMNi;SSY!+>{K?Fb%5902_ZstLO0yq`(6xbnjg%#wyg
zi8QZ5O9mZL(JE~%sM?H^qK6*jV{&ym`(5QyG`4NIq4IITjD?=QQ}ni#FS_1~vU4Ma
zzvd$M(pvXpCMI!7Nbnqe>$4w35JsE8hFesXAhi+{LnpxTZYOS&
zW6@aZevCSQ#`Dco!{wFIfzVUj%Fy&JQvmYMEB|zh$oGivp}h)ht=DOS7IPk;p8yw6
zoA`y!K+H1*$XK+_KnsG5nq;-=Ad^>+aZAYc6iW_+uL@mi%VeN8%d)5@-L}#|IOHzaT!KVLY_Le~o*AnCFY!+O##>QE_l0+i{*G(tnLUz<~SP1Wmri
z)zUgKS}?!vdN9~5e|*YkAGjP#S|ob*n^b%@&YZ+Hzlsamm8rJtp@BW{1qR(fixst3
zZ}&jcJ`t>$RUK77w-^pB1~l0hK>m{mcDk1Im}qmU;M)R;@h)u4%3f5Io;PFBV9HB-
zT8bA1-WleMoQW#tdlr0*@=6o}Up#?%`3|X_2-}8CLLyp>$l796t4{tVYmxw)NJR{q
z+^z^SBp}owL!TESgM@`3)7hz_8deO~b->u*$(t`bh(Y-dF+zj5F)c-j1Ncd*7bdhK
zAp%^-VY?VZ4#?yu&o2^>BGd>zJ7uLPsj*-l*S^)`gkEjwl+O;>j`%NV1)_EPtCqIE
zY8Wszw(Dp)(9AI^jKwyJH0o^{XWVR&y>O7omv0;r8OS&Gis-E{p1TP4vgzwv^rVNt
z4t_jZi(=l->`vDX~ie|!X22oFU
z&$357mvTjT4e{q$U`11WQC39VOYMZS5pHcHUWi~D`3^0Su0*+rOtwDmFO-=?6WVxL
z&8eRsl}?TC8o0f%n0<2c^e}_ybdgi_Hlu86rK7p|?+iMt#Eu&Ms7coS`xp1COVIiD0dk@gC{&K_#%1&CV#f6>5y#_P?RWvG|5@Fft>nPV+WF)~;
zp~??jYxKCH=mMuHBy8f5fVPcYcsBN8zxDNo-uU??8?_5a-byMxdGQ+JeYBq(VLj~6
z2|FE%fDFz*Eh5z7x%~kRY1IPVWH%!d+$nEUruHc_>PLUk{l;s%9&s)j2@`ogl+
z6~(n3EHI)zxk_p%Bp}z)TG6K1P0-tA_DWCjJ<1s&hDZI{sfPq%JpI{p_P0<~k~jD`
z9WFANm|aj0;GKUiu3w=i+?IL?2o*tZloz{cTc<*<78M66K!3upA`jlcvlI{bdP{0A
z%WJw)E=9r|JZ>YEv+`2JN^3G2H^Saq8LDcRMEhMBG?IGcLp9y4R7+-EKX0HNZBR2P
z-8-%xVDhnxy$%5~I;7rDQ@9@_wYEct_!%aH