rm bugs.md

This commit is contained in:
Jinjing
2026-09-01 03:53:50 -07:00
parent 20c6d5d304
commit 32ca206e9c
164 changed files with 28408 additions and 1879 deletions
+59 -9
View File
@@ -821,6 +821,51 @@ jobs:
- name: Custom agent platform suites (resolver, tokenizer, startup, runtime)
run: pnpm test:custom-agent-platform
# Why required rather than folded into the advisory `e2e` job: this is the only
# lane that observes a real remote custom-agent launch (argv+env read from
# /proc inside a throwaway Docker SSH container). The path-filtered `e2e` job
# skips whenever a PR changes launch code without touching tests/e2e/**, and
# the release workflow only runs full E2E after publication — so remote launch
# would otherwise reach users unverified. Ubuntu-only: macOS/Windows runners
# cannot host a Linux Docker container.
ssh_custom_agent_e2e:
name: e2e ssh custom agent
runs-on: ubuntu-latest
timeout-minutes: 40
steps:
- name: Checkout
uses: actions/checkout@v6
with:
persist-credentials: false
# Why explicit: headless Electron still needs an X display, and the e2e
# workflow's equivalent lane installs it rather than trusting the image.
- name: Install xvfb
run: sudo apt-get update && sudo apt-get install -y xvfb
- uses: ./.github/actions/install-node-dependencies
with:
native-runtime: electron
# Why no SKIP_BUILD: this lane owns its build (no shared e2e artifact in
# PR checks), and globalSetup adds the SSH relay bundle because the runner
# script exports ORCA_E2E_SSH_DOCKER=1. Docker is preinstalled on
# ubuntu-latest.
- name: Run SSH custom-agent e2e
env:
ORCA_E2E_FORWARD_APP_LOGS: '1'
run: xvfb-run --auto-servernum pnpm run test:e2e:ssh-custom-agent
- name: Upload Playwright traces
if: failure()
uses: actions/upload-artifact@v7
with:
name: playwright-traces-pr-ssh-custom-agent
path: test-results/
retention-days: 7
if-no-files-found: ignore
verify:
if: always()
needs:
@@ -837,17 +882,20 @@ jobs:
- package
- package_windows
- custom_agent_platform
- ssh_custom_agent_e2e
runs-on: ubuntu-latest
steps:
# Why: e2e is deliberately absent from needs. The suite is currently red on
# main (every scheduled run), so gating merges on it would block any PR that
# touches tests/e2e/** — including the ones fixing the suite. Until it is
# green the job runs and reports for E2E-path PRs without blocking. To flip
# it on: add `e2e` to needs, add E2E to the env below, and require
# `"$E2E" = success || skipped` after the loop — skipped is the normal
# result for a path-filtered job and must keep passing, so it has to be
# checked outside the loop or it would excuse the jobs above.
# Why: the path-filtered e2e job is deliberately absent from needs. That
# suite is currently red on main (every scheduled run), so gating merges on
# it would block any PR that touches tests/e2e/** — including the ones
# fixing the suite. Until it is green the job runs and reports for E2E-path
# PRs without blocking. To flip it on: add `e2e` to needs, add E2E to the
# env below, and require `"$E2E" = success || skipped` after the loop —
# skipped is the normal result for a path-filtered job and must keep
# passing, so it has to be checked outside the loop or it would excuse the
# jobs above. ssh_custom_agent_e2e is a different case: it is unfiltered,
# so `skipped` is never a legitimate result and it belongs in the loop.
- name: Require successful checks
env:
CODE_PATHS: ${{ needs.code_paths.result }}
@@ -864,6 +912,7 @@ jobs:
PACKAGE: ${{ needs.package.result }}
PACKAGE_WINDOWS: ${{ needs.package_windows.result }}
CUSTOM_AGENT_PLATFORM: ${{ needs.custom_agent_platform.result }}
SSH_CUSTOM_AGENT_E2E: ${{ needs.ssh_custom_agent_e2e.result }}
run: |
if [ "$CODE_PATHS" != "success" ]; then
exit 1
@@ -904,7 +953,8 @@ jobs:
"$MANAGED_HOOK_NODE18" \
"$PACKAGE" \
"$PACKAGE_WINDOWS" \
"$CUSTOM_AGENT_PLATFORM"; do
"$CUSTOM_AGENT_PLATFORM" \
"$SSH_CUSTOM_AGENT_E2E"; do
if [ "$result" != "success" ]; then
exit 1
fi
+385
View File
@@ -0,0 +1,385 @@
The zero-regression bar is not met. Review of af702b1c54 against
merge-base 9772da844a found 4 P0 release blockers, 27 P1 regressions,
and 9 P2 issues. The branch also fails typecheck, desktop build, lint,
and its new SSH E2E coverage.
## P0 — Release blockers
1. Mobile can execute AI text as shell commands against older hosts.
Identity-only launch is sent unconditionally at mobile/src/session/
identity-create-terminal-params.ts:15. Old hosts strip it and
create a bare shell; PR triage and review then submit provider/
user-controlled text with Enter at mobile/src/session/pr-ai-triage-
launch.ts:42 and mobile/src/session/use-mobile-diff-review-send-
actions.ts:109. Filenames, titles, comments, or check output
containing shell metacharacters can execute locally or over SSH.
2. Catalog/reference authoring acknowledges changes before they are
durable. Mutations return success immediately at src/main/agent-
launch/agent-catalog-service.ts:247, while persistence waits 1–5
seconds and merely logs failures at src/main/persistence.ts:3946.
Crash, power loss, or disk failure can lose acknowledged agents/
references or resurrect deletions.
3. Migration accepts an unreadable pre-v1 backup as valid. Read
failure returns “usable” at src/main/agent-launch/agent-catalog-
pre-v1-backup.ts:16. Reproduced with a directory at the backup
path: migration returned success without a readable rollback point.
4. New “durable” launch stores silently swallow every write failure
while callers report success:
- src/main/agent-launch/agent-launch-operation-store-
persistence.ts:269
- src/main/agent-launch/agent-session-record-store-
persistence.ts:171
- src/main/agent-launch/background-agent-launch-store-
persistence.ts:85
These stores also use the non-fsync writer. Restart can lose
idempotency, recovery, capacity, session-resume, or background-
launch state.
## P1 — Functionality, reliability, security, and performance
1. Paired web never consumes the custom-agent catalog. The host
publishes it at src/main/runtime/rpc/methods/client-ui.ts:38, but
paired web discards it at src/renderer/src/web/web-preload-
api.ts:3926, rejects catalog access at src/renderer/src/web/web-
preload-api.ts:756, and no-ops revision events at src/renderer/src/
hooks/useIpcEvents.ts:1007. Custom agents cannot be selected or
refreshed; a custom default becomes Blank.
2. Paired-web identity launches are not capability-gated. src/
renderer/src/runtime/web-runtime-session.ts:261 sends agentLaunch
to old hosts, which create a bare shell and report success. The
explicit gate already exists in src/renderer/src/runtime/agent-
launch-identity-negotiation.ts:1.
3. Folder-workspace quick-create has the same older-host regression:
it sends an empty command plus only agentLaunch at src/renderer/
src/components/sidebar/folder-workspace-composer-submit.ts:53. Old
hosts open a blank shell and drop the linked draft/note.
4. launchToken changed the daemon and SSH relay wire without version/
capability negotiation. The daemon remains v33 at src/main/daemon/
daemon-protocol-version.ts:1, while new main sends the field at
src/main/daemon/daemon-pty-adapter.ts:512. Old surviving daemons/
relays ignore it. After a main crash, the live agent appears absent
and Retry can create a duplicate.
5. A successful empty scoped SSH relist is never authoritative because
connection IDs are recorded only while iterating returned sessions
at src/main/runtime/orca-runtime.ts:32363. Pending launches remain
unknown, retain capacity, and cannot retry.
6. Catalog repair deadlocks with more than 256 corrupt rows. Snapshot
generation requests tokens for every row, but src/main/agent-
launch/agent-catalog-repair-mutations.ts:35 evicts them at 256.
Reproduced with 257 rows: every repair token becomes stale.
7. Reference payload budgeting measures the pre-persistence
representation at src/main/agent-launch/agent-catalog-
service.ts:229, not the expanded projection. A reproduced 180 KB
input was accepted and persisted as a 541 KB payload over the 512
KiB limit.
8. Quick-command RPC updates bypass reference authority at src/main/
runtime/orca-runtime.ts:4016: no live/enabled validation, reference
revision, tombstone handling, or change event.
9. Tombstone pruning is quadratic. src/main/agent-launch/agent-
catalog-tombstone-gc.ts:23 rescans every reference owner for every
tombstone. A minimal 3,500-row benchmark took about 591 ms before
real scanner overhead.
10. Mobile replay authorization changes are omitted from the admission
fingerprint. Replay environment can change after an env rotation
while the fingerprint remains identical: src/main/agent-launch/
resolve-agent-launch-result.ts:107.
11. WSL descriptors can be classified as native/local, leaving Windows
or UNC paths inside Linux argv: src/main/agent-launch/agent-launch-
host-state.ts:124, src/main/runtime/orca-runtime.ts:25930.
12. The command-length cap runs before the final prompt is appended.
RPC permits 100,000 characters at src/main/runtime/rpc/methods/
agent-launch-spawn-schema.ts:44, the cap checks earlier at src/
main/agent-launch/resolve-agent-launch.ts:296, and prompt suffixes
are added later at src/shared/resolved-agent-startup-plan.ts:181.
Windows can receive an oversized command instead of a typed
failure.
13. Stage-two host-state rejection can leak reservations because the
read occurs before guarded transaction cleanup: src/main/agent-
launch/agent-launch-worktree-resolution.ts:141, src/main/agent-
launch/agent-launch-worktree-transaction.ts:115. Repeated SSH
disconnect races can exhaust capacity until restart.
14. Admission principals identify only mobile or runtime, not devices,
by design at src/main/agent-launch/agent-launch-admission-
store.ts:25. Multiple phones share limits and recovery rows;
worktree/background Forget does not verify the stored principal,
allowing one paired device to forget another’s launch.
15. orchestration.dispatchForget claims owner authorization but
performs no Run/caller ownership check at src/main/runtime/rpc/
methods/orchestration.ts:1465. Stronger mutations use src/main/
runtime/rpc/methods/orchestration-run-scope.ts:37.
16. Environment-size validation checks only resolved custom env, not
the actual inherited/default spawn environment. Native Windows can
exceed CreateProcess’s environment-block limit and fail opaquely.
17. Settled operation history is globally unbounded. Only each
individual scope is capped at src/main/agent-launch/agent-launch-
operation-store.ts:237; unique scopes live forever and every write
serializes all scopes. At 5,000 scopes: 5,000 retained,
approximately 2.8 GB cumulative serialization, 8.2 seconds before
encryption/disk.
18. Session records are permanently unbounded and synchronously
rewrite encrypted full history at src/main/agent-launch/agent-
session-record-store.ts:101. Valid-cap benchmark: 1,000 records,
16.8 MB final payload, 8.4 GB cumulative writes, 11 seconds before
real encryption/disk.
19. Unbound launch staging also leaks. Runtime/mobile registrations
provide terminalId but no paneKey, while the sole cleanup only
removes by paneKey: src/main/agent-launch/agent-session-launch-
registration.ts:32, src/main/agent-launch/agent-session-record-
store.ts:170.
20. Background launch attempts are globally unbounded at src/main/
agent-launch/background-agent-launch-store.ts:43, with full-ledger
writes on creation and settlement. At 5,000 attempts: 4.7 GB
cumulative serialization and 8.1 seconds before disk.
21. Production launch rebuilds the full normalized catalog 2–3 times
per launch at src/main/agent-launch/agent-launch-spawn.ts:155 and
src/main/agent-launch/agent-launch-boundary.ts:343. At a valid
2,500-agent catalog, the three-pass p95 was about 32 ms of main-
thread CPU before preflight/spawn. The registered performance test
injects a prebuilt catalog and misses this production work.
22. useLocalAgentCatalog creates a separate full IPC fetch, cache, and
settings listener for every split group and closed quick-command
dialog: src/renderer/src/hooks/useLocalAgentCatalog.ts:28, src/
renderer/src/components/tab-bar/use-tab-bar-runtime-model.ts:151.
Sixteen 2,500-agent builds cost about 161 ms and clone roughly 9.7
MB.
23. Direct work-item launch rejects custom overrides/defaults because
it validates them against built-in-only detection at src/renderer/
src/lib/launch-work-item-direct.ts:142.
24. Source-control “Don’t save” launches stale persisted CLI args.
Current edited agentArgs exist at src/renderer/src/components/
right-sidebar/runSourceControlAgentActionStart.ts:123, but the
launch sends only a recipe locator at line 136.
25. The source-control action picker omits named custom agents because
options are built from the built-in catalog/detection only: src/
renderer/src/components/right-sidebar/
useSourceControlAgentActionDialog.ts:146.
26. Automation terminal cleanup is discarded. src/main/automations/
service.ts:143 early-returns for already-final runs, but the
renderer intentionally sends a second update clearing retired
terminal pointers at src/renderer/src/hooks/
useAutomationDispatchEvents.ts:377. “View run” can target a dead
terminal.
27. Remote capacity-recovery sheets do not live-refresh. src/renderer/
src/components/agent/AgentLaunchCapacityRecoverySheet.tsx:135
subscribes only to local preload events, not remote runtime
worktreesChanged.
28. All four new SSH custom-agent E2Es fail before exercising the
feature. tests/e2e/helpers/docker-ssh-custom-agent-remote.ts:112
treats SshTargetAddResult as the target instead of
destructuring .target, producing SSH target "undefined" not found.
29. The broken SSH suite is non-blocking in PR verification, and the
release workflow runs full E2E only after publication: .github/
workflows/pr.yml:541.
## P2 — Narrow but concrete regressions
- Repaired null defaults map back to legacy Auto and can unexpectedly
launch an installed agent: src/shared/tui-agent-selection.ts:82.
- Legacy Windows prefix parsing does not understand PowerShell
backticks or cmd carets: src/shared/legacy-agent-prefix-
tokenizer.ts:115.
- Duplicating a built-in flattens grouped args with .join(' '), losing
quoted boundaries: src/main/agent-launch/agent-catalog-lifecycle-
mutations.ts:117.
- Future catalog schema versions remain writable by older app builds:
src/shared/agent-catalog-schema-migration.ts:54.
- Quick-create can submit before the catalog loads and silently choose
a built-in: src/renderer/src/components/
NewWorkspaceComposerModal.tsx:159.
- Required Linux packaging bypasses the CLI require-graph verifier:
package.json:72, config/scripts/run-electron-vite-targets-in-
parallel.mjs:35. This reopens a previously shipped dead-CLI failure
class.
- Recovery UI advertises any existing backup as restorable even if it
cannot be read.
- settings.get now carries up to 512 KiB of catalog data on unrelated
mobile/paired-web settings reads; paired web then discards it: src/
main/runtime/rpc/methods/client-ui.ts:38.
- Mobile revision caches clear data but retain every historical host
key for process lifetime.
@@ -102,6 +102,20 @@ describe('PR E2E gate contract', () => {
expect(verifyStep.run).not.toContain('$E2E')
})
it('blocks merges on the SSH custom-agent E2E lane', () => {
// Why unfiltered and required: it is the only live evidence that the host
// resolves and spawns a custom agent on a remote, and a launch-code change
// can break it without touching tests/e2e/**. Leaving it advisory let a
// broken suite reach a release, where full E2E only runs after publication.
const sshJob = prWorkflow.jobs.ssh_custom_agent_e2e
expect(sshJob.if).toBeUndefined()
expect(sshJob['continue-on-error']).toBeUndefined()
expect(sshJob.steps.find((step) => step.name === 'Run SSH custom-agent e2e').run).toContain(
'pnpm run test:e2e:ssh-custom-agent'
)
expect(prWorkflow.jobs.verify.needs).toContain('ssh_custom_agent_e2e')
})
it('passes only changed specs to the reusable E2E workflow', () => {
// Why: without this the job could lose its filter and run on every PR — the
// cost the path filter exists to avoid — while the gate assertions above
@@ -102,8 +102,12 @@ describe('PR workflow parallelism', () => {
.split(/\s+/)
.filter((token) => !['apt-get', 'install', 'sudo', ''].includes(token))
.filter((token) => !token.startsWith('-'))
const jobsInstallingPackages = Object.entries(workflow.jobs)
.filter(([, job]) => (job.steps ?? []).some((step) => aptPackages(step).length > 0))
// Scoped to shells: other lanes legitimately apt-install unrelated tooling.
const shells = ['zsh', 'fish', 'bash', 'dash', 'ksh']
const jobsInstallingShells = Object.entries(workflow.jobs)
.filter(([, job]) =>
(job.steps ?? []).some((step) => aptPackages(step).some((pkg) => shells.includes(pkg)))
)
.map(([name]) => name)
expect(shellStep).toBeDefined()
@@ -111,7 +115,7 @@ describe('PR workflow parallelism', () => {
expect(shellStep.run.split(/\s+/)).toContain('--maxWorkers=1')
// Why the whole workflow, not just the general shards: any other lane installing
// these shells would silently start running the real-shell tests twice.
expect(jobsInstallingPackages).toEqual(['shell_contracts'])
expect(jobsInstallingShells).toEqual(['shell_contracts'])
// Why each shell is asserted: the live tests skip themselves when the binary is
// missing, so a dropped package silently empties this lane instead of failing it.
const shellPackages = workflow.jobs.shell_contracts.steps.flatMap(aptPackages)
@@ -463,7 +467,8 @@ describe('PR workflow parallelism', () => {
'managed_hook_node18',
'package',
'package_windows',
'custom_agent_platform'
'custom_agent_platform',
'ssh_custom_agent_e2e'
])
const verifyStep = workflow.jobs.verify.steps.find(
(step) => step.name === 'Require successful checks'
@@ -2,31 +2,28 @@ import { spawn } from 'node:child_process'
import { fileURLToPath } from 'node:url'
const buildScript = fileURLToPath(new URL('./run-electron-vite-build.mjs', import.meta.url))
// Keep this wrapper CommonJS (the `.cts` extension) so electron-vite can load
// each parallel target without sharing its timestamp-named ESM temp file.
const targetConfig = fileURLToPath(new URL('../electron-vite-target.config.cts', import.meta.url))
const verifyRequiresScript = fileURLToPath(
new URL('./verify-cli-require-resolution.mjs', import.meta.url)
)
const targetConfig = fileURLToPath(new URL('../electron-vite-target.config.ts', import.meta.url))
const targets = ['main', 'preload', 'renderer']
function buildTarget(target) {
function runNodeScript(args, env, label) {
return new Promise((resolve, reject) => {
const child = spawn(
process.execPath,
[buildScript, '--config', targetConfig, '--ignoreConfigWarning'],
{
stdio: 'inherit',
env: {
...process.env,
ORCA_ELECTRON_VITE_TARGET: target
}
const child = spawn(process.execPath, args, {
stdio: 'inherit',
env: {
...process.env,
...env
}
)
})
child.on('error', reject)
child.on('exit', (code, signal) => {
if (signal) {
reject(new Error(`Electron Vite ${target} build exited with signal ${signal}`))
reject(new Error(`${label} exited with signal ${signal}`))
} else if (code !== 0) {
reject(new Error(`Electron Vite ${target} build exited with code ${code}`))
reject(new Error(`${label} exited with code ${code}`))
} else {
resolve()
}
@@ -34,6 +31,14 @@ function buildTarget(target) {
})
}
function buildTarget(target) {
return runNodeScript(
[buildScript, '--config', targetConfig, '--ignoreConfigWarning'],
{ ORCA_ELECTRON_VITE_TARGET: target },
`Electron Vite ${target} build`
)
}
const results = await Promise.allSettled(targets.map(buildTarget))
const failures = results.filter((result) => result.status === 'rejected')
@@ -43,3 +48,12 @@ if (failures.length > 0) {
}
process.exit(1)
}
// Why: build:electron-vite chains this verifier, and Linux packaging takes the
// parallel path — skipping it here reopens the dead-CLI require graph (1.4.150-rc.1.perf).
try {
await runNodeScript([verifyRequiresScript], {}, 'CLI require resolution check')
} catch (error) {
console.error(error instanceof Error ? error.message : String(error))
process.exit(1)
}
@@ -0,0 +1,38 @@
import { readFileSync } from 'node:fs'
import { join, resolve } from 'node:path'
import { describe, expect, it } from 'vitest'
const projectDir = resolve(import.meta.dirname, '../..')
const packageJson = JSON.parse(readFileSync(join(projectDir, 'package.json'), 'utf8'))
const runnerSource = readFileSync(
join(projectDir, 'config/scripts/run-electron-vite-targets-in-parallel.mjs'),
'utf8'
)
// Why: the serial script is the source of truth for what a completed
// electron-vite build owes, so renaming its verifier fails here instead of
// silently leaving the parallel path — the one Linux packaging takes — unverified.
const serialVerifier = packageJson.scripts['build:electron-vite'].match(
/config\/scripts\/(verify-[\w-]+\.mjs)/
)?.[1]
describe('parallel electron-vite target runner', () => {
it('verifies the compiled CLI require graph like the serial build does', () => {
expect(serialVerifier).toBe('verify-cli-require-resolution.mjs')
expect(runnerSource).toContain(`./${serialVerifier}`)
})
it('runs the verifier only after every target build succeeded', () => {
const failureGate = runnerSource.indexOf('failures.length > 0')
const verifierRun = runnerSource.indexOf('runNodeScript([verifyRequiresScript]')
expect(failureGate).toBeGreaterThan(-1)
expect(verifierRun).toBeGreaterThan(failureGate)
})
it('keeps packaging on the parallel entry point this runner backs', () => {
expect(packageJson.scripts['build:electron-vite:parallel']).toContain(
'run-electron-vite-targets-in-parallel.mjs'
)
})
})
@@ -12,6 +12,12 @@ import { pathToFileURL } from 'node:url'
const RELATIVE_REQUIRE_PATTERN = /\brequire\(\s*(["'])(\.{1,2}\/[^"']+)\1\s*\)/g
// Why: build:cli emits compiled specs into out/, but the shipped CLI never
// requires them — they run from src under vitest. Seeding the walk with them
// would fail the build on imports (e.g. main-side parity checks) that no user
// command can reach. They are still walked if a real entry pulls one in.
const COMPILED_SPEC_PATTERN = /\.(test|spec)\.js$/
function listJsFilesRecursively(dir) {
const files = []
for (const entry of readdirSync(dir, { withFileTypes: true })) {
@@ -51,7 +57,7 @@ export function verifyCliRequireResolution({
return { checkedFiles: 0, missing: [], skipped: true }
}
const queue = listJsFilesRecursively(cliDir)
const queue = listJsFilesRecursively(cliDir).filter((file) => !COMPILED_SPEC_PATTERN.test(file))
const visited = new Set(queue)
const missing = []
while (queue.length > 0) {
@@ -75,6 +75,29 @@ describe('verifyCliRequireResolution', () => {
expect(result.missing).toEqual([])
})
it('ignores compiled specs that import main modules the shipped CLI never loads', () => {
// serve-electron-flag-parity.test.ts imports a main module listed only for
// typecheck parity; no CLI command can reach it.
const projectDir = makeProject({
'out/cli/index.js': 'module.exports = {}',
'out/cli/serve-electron-flag-parity.test.js': `require("../main/startup/serve-mode-argv");`
})
const result = verifyCliRequireResolution({ projectDir })
expect(result.missing).toEqual([])
expect(result.checkedFiles).toBe(1)
})
it('still reports a spec pulled in by a real CLI entry', () => {
const projectDir = makeProject({
'out/cli/index.js': `require("./fixtures.test");`,
'out/cli/fixtures.test.js': `require("../main/deleted-by-clean");`
})
const result = verifyCliRequireResolution({ projectDir })
expect(result.missing).toEqual([
{ from: path.join('out', 'cli', 'fixtures.test.js'), specifier: '../main/deleted-by-clean' }
])
})
it('skips when out/cli has not been built', () => {
const projectDir = makeProject({ 'out/main/index.js': '' })
const result = verifyCliRequireResolution({ projectDir })
File diff suppressed because it is too large Load Diff
@@ -41,7 +41,12 @@ export async function loadMobileResumeMetadata(
.sendRequest('projectGroup.list', undefined, { timeoutMs: RESUME_RPC_TIMEOUT_MS })
.catch(() => null),
client
.sendRequest('settings.get', undefined, { timeoutMs: RESUME_RPC_TIMEOUT_MS })
.sendRequest(
'settings.get',
// Resume reads only `settings`; the catalog comes from settings.agentCatalog.get.
{ includeAgentCatalog: false },
{ timeoutMs: RESUME_RPC_TIMEOUT_MS }
)
.catch(() => null),
client
.sendRequest('worktree.ps', { limit: 10000 }, { timeoutMs: RESUME_RPC_TIMEOUT_MS })
@@ -93,7 +93,9 @@ export function useNewWorkspaceCreateSubmit(args: {
}
let latestRuntimeSettings = args.runtimeSettings
try {
const settingsResponse = await client.sendRequest('settings.get')
const settingsResponse = await client.sendRequest('settings.get', {
includeAgentCatalog: false
})
if (settingsResponse.ok) {
const result = (settingsResponse as RpcSuccess).result as {
settings: NewWorktreeRuntimeSettings
@@ -39,7 +39,7 @@ export function useNewWorkspaceRuntimeContext(
client.sendRequest('linear.status')
])
const [settingsRes, uiRes] = await Promise.allSettled([
client.sendRequest('settings.get'),
client.sendRequest('settings.get', { includeAgentCatalog: false }),
client.sendRequest('ui.get')
])
if (stale) {
+3 -1
View File
@@ -73,7 +73,9 @@ export function fetchMobileHomeTaskProviders(
disposed: () => boolean
): void {
Promise.all([
sendSingleFlightRequest(client, hostId, 'settings.get'),
sendSingleFlightRequest(client, hostId, 'settings.get', {
includeAgentCatalog: false
}),
sendSingleFlightRequest(client, hostId, 'preflight.check'),
sendSingleFlightRequest(client, hostId, 'linear.status')
])
@@ -40,6 +40,10 @@ describe('mobile new-tab agent loading', () => {
'preflight.detectAgents',
'settings.get'
])
// The 512 KiB agent catalog must not ride this hot read.
expect(
client.sendRequest.mock.calls.find(([method]) => method === 'settings.get')?.[1]
).toEqual({ includeAgentCatalog: false })
})
it('detects agents through the worktree repo connection for SSH sessions', async () => {
@@ -23,7 +23,8 @@ export async function loadMobileNewTabAgentOptions(args: {
? client.sendRequest('preflight.detectAgents')
: loadWorkspaceDetectedAgents(client, worktreeId)
const [settingsResponse, detectedResponse] = await Promise.all([
client.sendRequest('settings.get'),
// Only `settings` is read here; opt out of the piggybacked agent catalog (old hosts ignore the param).
client.sendRequest('settings.get', { includeAgentCatalog: false }),
detectedAgentsRequest
])
if (!settingsResponse.ok) {
@@ -33,7 +33,8 @@ export function usePRBotAuthorOverrides(
}
let stale = false
void client
.sendRequest('settings.get')
// Only prBotAuthorOverrides is read; opt out of the piggybacked agent catalog.
.sendRequest('settings.get', { includeAgentCatalog: false })
.then((response) => {
if (stale || !response.ok) {
return
+92 -18
View File
@@ -6,45 +6,119 @@ import type { RpcClient } from './rpc-client'
const tick = () => new Promise((resolve) => setTimeout(resolve, 0))
function fakeClient(response: unknown): { client: RpcClient; methods: string[] } {
return routingClient(() => response)
}
/** Lets a test answer the dedicated read and the legacy piggyback differently. */
function routingClient(route: (method: string) => unknown): {
client: RpcClient
methods: string[]
} {
const methods: string[] = []
const client = {
sendRequest: async (method: string) => {
methods.push(method)
return response
return route(method)
},
getState: () => 'connected'
} as unknown as RpcClient
return { client, methods }
}
function catalogResult(revision: number): unknown {
return {
ok: true,
id: '1',
result: {
agentCatalog: {
version: 1,
revision,
defaultAgent: 'auto',
disabledAgents: [],
customAgents: [],
deletedCustomAgents: []
}
},
_meta: { runtimeId: 'runtime-1' }
}
}
function failure(code: string, message = 'nope'): unknown {
return { ok: false, id: '1', error: { code, message }, _meta: { runtimeId: 'runtime-1' } }
}
describe('agentCatalogSync', () => {
it('hydrates the catalog snapshot from settings.get', async () => {
const { client, methods } = fakeClient({
ok: true,
id: '1',
result: {
agentCatalog: {
version: 1,
revision: 7,
defaultAgent: 'auto',
disabledAgents: [],
customAgents: [],
deletedCustomAgents: []
}
},
_meta: { runtimeId: 'runtime-1' }
})
it('hydrates the catalog snapshot from the dedicated settings.agentCatalog.get read', async () => {
const { client, methods } = fakeClient(catalogResult(7))
const conn = agentCatalogSync.openConnection('cat-a', client)
conn.hydrate()
await tick()
expect(methods).toEqual(['settings.get'])
expect(methods).toEqual(['settings.agentCatalog.get'])
expect(agentCatalogSync.getSnapshot('cat-a')).toMatchObject({ version: 1, revision: 7 })
conn.dispose()
agentCatalogSync.clear('cat-a')
})
it.each(['forbidden', 'method_not_found'])(
'falls back to the settings.get piggyback when an old host answers %s',
async (code) => {
const { client, methods } = routingClient((method) =>
method === 'settings.agentCatalog.get' ? failure(code) : catalogResult(3)
)
const conn = agentCatalogSync.openConnection('cat-old', client)
conn.hydrate()
await tick()
expect(methods).toEqual(['settings.agentCatalog.get', 'settings.get'])
expect(agentCatalogSync.getSnapshot('cat-old')).toMatchObject({ revision: 3 })
conn.dispose()
agentCatalogSync.clear('cat-old')
}
)
it('keeps the cache and skips the fallback when the dedicated read fails transiently', async () => {
let dedicated: unknown = catalogResult(5)
const { client, methods } = routingClient((method) =>
method === 'settings.agentCatalog.get' ? dedicated : catalogResult(9)
)
const conn = agentCatalogSync.openConnection('cat-transient', client)
conn.hydrate()
await tick()
expect(agentCatalogSync.getSnapshot('cat-transient')).toMatchObject({ revision: 5 })
dedicated = failure('internal_error', 'boom')
methods.length = 0
conn.announce(6)
await tick()
expect(methods).toEqual(['settings.agentCatalog.get'])
expect(agentCatalogSync.getSnapshot('cat-transient')).toMatchObject({ revision: 5 })
conn.dispose()
agentCatalogSync.clear('cat-transient')
})
it('drops the cache when a downgraded host answers the dedicated read without a catalog', async () => {
let result: unknown = catalogResult(5)
const { client } = routingClient(() => result)
const conn = agentCatalogSync.openConnection('cat-absent', client)
conn.hydrate()
await tick()
expect(agentCatalogSync.getSnapshot('cat-absent')).toMatchObject({ revision: 5 })
result = { ok: true, id: '1', result: {}, _meta: { runtimeId: 'runtime-1' } }
conn.hydrate()
await tick()
expect(agentCatalogSync.getSnapshot('cat-absent')).toBeNull()
conn.dispose()
agentCatalogSync.clear('cat-absent')
})
it('stores the projection error variant so repair copy can render', async () => {
const { client } = fakeClient({
ok: true,
+38 -13
View File
@@ -1,9 +1,11 @@
// Why: mobile's per-host cache of the runtime agent catalog. The host owns the
// catalog; mobile only mirrors the env-free revisioned snapshot it publishes on
// `settings.get`, refetching on `agentCatalogChanged` events. Never merges with
// local state and never carries a custom env key or value (env-free by DTO
// construction — no fields are added here).
// `settings.agentCatalog.get` (falling back to the legacy `settings.get`
// piggyback against old hosts), refetching on `agentCatalogChanged` events.
// Never merges with local state and never carries a custom env key or value
// (env-free by DTO construction — no fields are added here).
import type { RpcClient } from './rpc-client'
import type { RpcResponse, RpcSuccess } from './types'
import type {
AgentCatalogProjectionError,
AgentCatalogSnapshot
@@ -29,28 +31,51 @@ function parseCatalogValue(raw: unknown): AgentCatalogValue | null {
return raw as AgentCatalogValue
}
async function fetchCatalog(client: RpcClient): Promise<SnapshotFetchOutcome<AgentCatalogValue>> {
let response
async function sendCatalogRead(client: RpcClient, method: string): Promise<RpcResponse | null> {
try {
response = await client.sendRequest('settings.get')
return (await client.sendRequest(method)) ?? null
} catch {
return { kind: 'unavailable' }
}
if (!response || !response.ok) {
return { kind: 'unavailable' }
return null
}
}
// Mobile's method allowlist is checked before dispatch, so a host that predates
// the dedicated read rejects it with 'forbidden' rather than 'method_not_found'.
function isMethodUnavailable(response: RpcResponse | null): boolean {
const error = response && !response.ok ? response.error : undefined
return (
error?.code === 'forbidden' ||
error?.code === 'method_not_found' ||
error?.message.includes('not available to mobile clients') === true
)
}
function readCatalog(response: RpcSuccess): SnapshotFetchOutcome<AgentCatalogValue> {
const result = response.result as { agentCatalog?: unknown } | null
const runtimeId = (response as { _meta?: { runtimeId?: string } })._meta?.runtimeId ?? ''
const value = parseCatalogValue(result?.agentCatalog)
if (!value) {
// settings.get answered without a catalog: this host publishes none (it was
// downgraded, or predates the catalog). Keeping the cached one would leave the
// picker offering custom agents the host can no longer launch.
// The host answered without a catalog: it publishes none (it was downgraded,
// or predates the catalog). Keeping the cached one would leave the picker
// offering custom agents the host can no longer launch.
return { kind: 'absent', runtimeId }
}
return { kind: 'value', runtimeId, value }
}
async function fetchCatalog(client: RpcClient): Promise<SnapshotFetchOutcome<AgentCatalogValue>> {
const dedicated = await sendCatalogRead(client, 'settings.agentCatalog.get')
if (dedicated?.ok) {
return readCatalog(dedicated)
}
if (!isMethodUnavailable(dedicated)) {
return { kind: 'unavailable' }
}
// Old host: the catalog ships only piggybacked on settings.get.
const legacy = await sendCatalogRead(client, 'settings.get')
return legacy?.ok ? readCatalog(legacy) : { kind: 'unavailable' }
}
const sync = createRevisionedSnapshotSync<AgentCatalogValue>()
export const agentCatalogSync = {
@@ -151,6 +151,39 @@ describe('createRevisionedSnapshotSync', () => {
expect(sync.getSnapshot('b')?.revision).toBe(20)
})
// Every host disconnect clears its cache, so retaining the key would grow this
// map by one entry per host ever connected to, for the process lifetime.
it('drops the host key on clear, not just its cached value', async () => {
const sync = createRevisionedSnapshotSync<Snap>()
const { fetch, resolvers } = deferredFetch()
for (let host = 0; host < 5; host++) {
const conn = sync.openConnection(`host-${host}`, fetch)
conn.hydrate()
resolvers[host]!({ kind: 'value', runtimeId: 'r', value: { revision: 1 } })
await tick()
conn.dispose()
sync.clear(`host-${host}`)
}
expect(sync.trackedHostCountForTests()).toBe(0)
})
it('still fences an in-flight fetch whose host was cleared', async () => {
const sync = createRevisionedSnapshotSync<Snap>()
const { fetch, resolvers } = deferredFetch()
const conn = sync.openConnection('host', fetch)
conn.hydrate()
sync.clear('host')
// Resolves against the dropped state; must not resurrect the host entry.
resolvers[0]!({ kind: 'value', runtimeId: 'r', value: { revision: 7 } })
await tick()
expect(sync.getSnapshot('host')).toBeNull()
expect(sync.trackedHostCountForTests()).toBe(0)
})
it('stores a projection error as the cached value so the UI can render repair copy', async () => {
const sync = createRevisionedSnapshotSync<Snap>()
const { fetch, resolvers } = deferredFetch()
@@ -41,6 +41,8 @@ export type RevisionedSnapshotSync<V extends RevisionedSnapshot> = {
clear: (hostId: string) => void
getSnapshot: (hostId: string) => V | null
subscribe: (hostId: string, listener: () => void) => () => void
// Hosts still holding a state entry; proves clear() drops the key, not just the value.
trackedHostCountForTests: () => number
}
type HostState<V> = {
@@ -207,6 +209,9 @@ export function createRevisionedSnapshotSync<
if (!state) {
return
}
// Fence in-flight fetches and stale handles on the object they captured,
// THEN drop the key — keeping every host ever connected to would grow this
// map for the process lifetime.
state.epoch += 1
state.value = null
state.runtimeId = null
@@ -214,11 +219,15 @@ export function createRevisionedSnapshotSync<
state.fetching = false
state.hydratePending = false
state.highestAnnounced = null
states.delete(hostId)
notify(hostId)
},
getSnapshot(hostId) {
return states.get(hostId)?.value ?? null
},
trackedHostCountForTests() {
return states.size
},
subscribe(hostId, listener) {
let set = listeners.get(hostId)
if (!set) {
@@ -0,0 +1,87 @@
// Why: the agent catalog projection is budgeted at 512 KiB and only rides
// settings.get as the legacy piggyback for hosts without settings.agentCatalog.get.
// A caller that reads settings without opting out drags that payload onto
// every unrelated read, so the opt-out is asserted rather than left to review.
import { readdirSync, readFileSync, statSync } from 'node:fs'
import { join } from 'node:path'
import { describe, expect, it } from 'vitest'
const MOBILE_ROOT = join(__dirname, '..', '..')
// The only caller allowed to piggyback: it falls back to settings.get against
// hosts that predate settings.agentCatalog.get.
const CATALOG_PIGGYBACK_OWNER = join(MOBILE_ROOT, 'src', 'transport', 'agent-catalog-sync.ts')
function listSourceFiles(root: string): string[] {
const files: string[] = []
for (const entry of readdirSync(root)) {
const path = join(root, entry)
if (statSync(path).isDirectory()) {
files.push(...listSourceFiles(path))
continue
}
if (/\.[cm]?[jt]sx?$/.test(entry) && !/\.test\.[cm]?[jt]sx?$/.test(entry)) {
files.push(path)
}
}
return files
}
function scannedSources(): { path: string; source: string }[] {
return [join(MOBILE_ROOT, 'app'), join(MOBILE_ROOT, 'src')]
.flatMap((root) => listSourceFiles(root))
.filter((path) => path !== CATALOG_PIGGYBACK_OWNER)
.map((path) => ({
path,
// Drop line comments so an explanatory comment between the method name and
// its params can't push the opt-out outside the match window.
source: readFileSync(path, 'utf8').replace(/^\s*\/\/.*$/gm, '')
}))
}
function describeSite(path: string, match: string): string {
return `${path.slice(MOBILE_ROOT.length + 1)}: ${match.split('\n')[0]}`
}
function callSitesNotOptingOut(): string[] {
const sites: string[] = []
for (const { path, source } of scannedSources()) {
for (const match of source.matchAll(/['"]settings\.get['"](.{0,80})/gs)) {
if (!match[1]!.includes('includeAgentCatalog')) {
sites.push(describeSite(path, match[0]!))
}
}
}
return sites
}
// sendSingleFlightRequest coalesces by (client, host, requestKind) and lets the
// latest queued params win, so two settings.get callers disagreeing about
// includeAgentCatalog could hand one of them the other's response. Keeping every
// single-flight settings.get caller opted out makes that mismatch unreachable.
function singleFlightSettingsSitesNotOptingOut(): string[] {
const sites: string[] = []
for (const { path, source } of scannedSources()) {
for (const match of source.matchAll(
/sendSingleFlightRequest\([^)]*?['"]settings\.get['"](.{0,80})/gs
)) {
if (!match[1]!.includes('includeAgentCatalog: false')) {
sites.push(describeSite(path, match[0]!))
}
}
}
return sites
}
describe('mobile settings.get catalog opt-out', () => {
it('opts every mobile settings.get caller out of the piggybacked agent catalog', () => {
expect(callSitesNotOptingOut()).toEqual([])
})
it('keeps every single-flight settings.get caller on the same opted-out params', () => {
expect(singleFlightSettingsSitesNotOptingOut()).toEqual([])
})
it('keeps the catalog piggyback off the params-blind single-flight path', () => {
expect(readFileSync(CATALOG_PIGGYBACK_OWNER, 'utf8')).not.toContain('sendSingleFlightRequest')
})
})
@@ -0,0 +1,295 @@
// Timeout-flag validation and task-list brief output for the orchestration CLI
// handlers; payload/handle mapping lives in orchestration.test.ts.
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import {
callMock,
getTerminalHandleMock,
handlerInvoker,
restoreTerminalIdentityEnv
} from './orchestration-handler-test-harness'
// Why: isolate the handler's flag-to-param mapping; printResult only writes output.
vi.mock('../format', () => ({ printResult: vi.fn() }))
vi.mock('../selectors', () => ({ getTerminalHandle: getTerminalHandleMock }))
import { ORCHESTRATION_HANDLERS } from './orchestration'
import { printResult } from '../format'
afterEach(restoreTerminalIdentityEnv)
describe('orchestration timeout flag validation', () => {
const invalidTimeoutValues: [string, string | boolean][] = [
['missing', true],
['empty', ''],
['non-numeric', 'not-a-number'],
['zero', '0'],
['negative', '-1']
]
beforeEach(() => {
callMock.mockReset()
delete process.env.ORCA_TERMINAL_HANDLE
delete process.env.ORCA_PANE_KEY
})
const invokeCheck = handlerInvoker(ORCHESTRATION_HANDLERS['orchestration check'])
const invokeAsk = handlerInvoker(ORCHESTRATION_HANDLERS['orchestration ask'])
it.each(invalidTimeoutValues)('rejects invalid check --timeout-ms: %s', async (_label, value) => {
const flags = new Map<string, string | boolean>([
['wait', true],
['timeout-ms', value]
])
await expect(invokeCheck(flags)).rejects.toThrow(/--timeout-ms/)
expect(callMock).not.toHaveBeenCalled()
})
it('passes a parsed check timeout and peek mode into the RPC payload', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
callMock.mockResolvedValue({ result: { messages: [], count: 0 } })
await invokeCheck(
new Map<string, string | boolean>([
['wait', true],
['peek', true],
['timeout-ms', '250']
])
)
// Why: --peek rides with unread:false so pre-peek runtimes fall back to
// the non-consuming all mode instead of the destructive mark-read default.
expect(callMock).toHaveBeenCalledWith('orchestration.check', {
terminal: 'term_worker',
terminalPaneKey: undefined,
unread: false,
peek: true,
all: undefined,
types: undefined,
format: undefined,
compatibilityCliCommand: expect.stringMatching(/^orca(?:-ide)?$/),
run: undefined,
ack: undefined,
wait: true,
timeoutMs: 250
})
})
it('filters already-read rows from a peek response for pre-peek runtimes', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
callMock.mockResolvedValue({
result: {
messages: [
{ id: 'msg_old', from_handle: 'a', subject: 'seen', read: 1 },
{ id: 'msg_new', from_handle: 'a', subject: 'fresh', read: 0 }
],
count: 2,
formatted: 'banners built from all rows'
}
})
vi.mocked(printResult).mockClear()
await invokeCheck(new Map<string, string | boolean>([['peek', true]]))
const response = vi.mocked(printResult).mock.calls[0]?.[0] as {
result: { messages: { id: string }[]; count: number; formatted?: string }
}
expect(response.result.messages.map((m) => m.id)).toEqual(['msg_new'])
expect(response.result.count).toBe(1)
// Why: the pre-peek runtime built `formatted` from all rows, including
// the read one the filter just removed.
expect(response.result.formatted).toBeUndefined()
})
it('rejects combined read modes before calling the runtime', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
callMock.mockClear()
await expect(
invokeCheck(
new Map<string, string | boolean>([
['unread', true],
['peek', true]
])
)
).rejects.toMatchObject({
code: 'invalid_argument',
message: expect.stringContaining('read mode')
})
expect(callMock).not.toHaveBeenCalled()
})
it('warns when a pre-peek runtime returned a full 100-row page', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
const rows = Array.from({ length: 100 }, (_, i) => ({
id: `msg_${i}`,
from_handle: 'a',
subject: `s${i}`,
read: i === 0 ? 0 : 1
}))
callMock.mockResolvedValue({ result: { messages: rows, count: 100 } })
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
await invokeCheck(new Map<string, string | boolean>([['peek', true]]))
expect(errorSpy).toHaveBeenCalledWith(expect.stringContaining('newest 100 messages'))
errorSpy.mockRestore()
})
it('fails --peek --wait against a runtime that returned only read rows', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
callMock.mockResolvedValue({
result: {
messages: [{ id: 'msg_old', from_handle: 'a', subject: 'seen', read: 1 }],
count: 1
}
})
await expect(
invokeCheck(
new Map<string, string | boolean>([
['peek', true],
['wait', true]
])
)
).rejects.toMatchObject({ code: 'peek_wait_unsupported' })
})
it.each(invalidTimeoutValues)('rejects invalid ask --timeout-ms: %s', async (_label, value) => {
const flags = new Map<string, string | boolean>([
['to', 'term_coord'],
['question', 'Proceed?'],
['timeout-ms', value]
])
await expect(invokeAsk(flags)).rejects.toThrow(/--timeout-ms/)
expect(callMock).not.toHaveBeenCalled()
})
it('uses the parsed ask timeout for both runtime wait and client timeout', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
callMock.mockResolvedValue({
result: {
answer: 'yes',
messageId: 'msg_1',
threadId: 'thread_1',
timedOut: false
}
})
vi.spyOn(console, 'log').mockImplementation(() => {})
await invokeAsk(
new Map<string, string | boolean>([
['to', 'term_coord'],
['question', 'Proceed?'],
['timeout-ms', '123']
])
)
expect(callMock).toHaveBeenCalledWith(
'orchestration.ask',
{
to: 'term_coord',
run: undefined,
question: 'Proceed?',
resume: undefined,
options: undefined,
timeoutMs: 123,
from: 'term_worker',
compatibilityCliCommand: expect.stringMatching(/^orca(?:-ide)?$/),
compatibilityWindowsCommand: undefined
},
{ timeoutMs: 5_123, orchestrationCapability: undefined }
)
})
it('passes an ask resume without creating a new question payload', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
callMock.mockResolvedValue({
result: {
answer: 'yes',
messageId: 'msg_question',
threadId: 'msg_question',
timedOut: false
}
})
vi.spyOn(console, 'log').mockImplementation(() => {})
await invokeAsk(new Map<string, string | boolean>([['resume', 'msg_question']]))
expect(callMock).toHaveBeenCalledWith(
'orchestration.ask',
{
to: undefined,
run: undefined,
question: undefined,
resume: 'msg_question',
options: undefined,
timeoutMs: undefined,
from: 'term_worker',
compatibilityCliCommand: expect.stringMatching(/^orca(?:-ide)?$/),
compatibilityWindowsCommand: undefined
},
{ timeoutMs: 605_000, orchestrationCapability: undefined }
)
})
it('rejects ambiguous ask create/resume input before RPC', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
await expect(
invokeAsk(
new Map<string, string | boolean>([
['question', 'new'],
['resume', 'msg_old']
])
)
).rejects.toMatchObject({ code: 'invalid_argument' })
expect(callMock).not.toHaveBeenCalled()
})
})
describe('orchestration task-list brief output', () => {
it('requests server-side brief and falls back client-side for older runtimes', async () => {
callMock.mockReset().mockResolvedValue({
result: {
// No spec_truncated field — the pre-brief-runtime signature.
tasks: [{ id: 'task_1', spec: `First line\n${'detail '.repeat(40)}`, status: 'ready' }],
count: 1
}
})
vi.mocked(printResult).mockClear()
await handlerInvoker(ORCHESTRATION_HANDLERS['orchestration task-list'])(
new Map([['brief', true]])
)
expect(callMock).toHaveBeenCalledWith(
'orchestration.taskList',
expect.objectContaining({ brief: true })
)
const response = vi.mocked(printResult).mock.calls[0]?.[0] as {
result: { tasks: { spec: string; spec_truncated: boolean }[] }
}
expect(response.result.tasks[0].spec).toHaveLength(160)
expect(response.result.tasks[0].spec_truncated).toBe(true)
})
it('passes server-abbreviated rows through untouched', async () => {
const serverTasks = [
{ id: 'task_1', spec: 'already brief…', status: 'ready', spec_truncated: true }
]
callMock.mockReset().mockResolvedValue({ result: { tasks: serverTasks, count: 1 } })
vi.mocked(printResult).mockClear()
await handlerInvoker(ORCHESTRATION_HANDLERS['orchestration task-list'])(
new Map([['brief', true]])
)
const response = vi.mocked(printResult).mock.calls[0]?.[0] as {
result: { tasks: { spec: string; spec_truncated: boolean }[] }
}
// Why: re-abbreviating a server-truncated spec would flip spec_truncated
// back to false (the truncated text fits the cap).
expect(response.result.tasks).toBe(serverTasks)
})
})
+37 -281
View File
@@ -1,5 +1,6 @@
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import {
type CliFlagMap,
callMock,
getTerminalHandleMock,
handlerInvoker,
@@ -7,8 +8,7 @@ import {
restoreTerminalIdentityEnv,
staleHandleError,
stubStaleHandleRemint,
stubStaleHandleRemintFailure,
type CliFlagMap
stubStaleHandleRemintFailure
} from './orchestration-handler-test-harness'
// Why: isolate the handler's flag-to-param mapping; printResult only writes output.
@@ -17,7 +17,6 @@ vi.mock('../selectors', () => ({ getTerminalHandle: getTerminalHandleMock }))
import { ORCHESTRATION_HANDLERS } from './orchestration'
import { RuntimeClientError } from '../runtime-client'
import { printResult } from '../format'
afterEach(restoreTerminalIdentityEnv)
@@ -452,12 +451,16 @@ describe('orchestration dispatch Forget + raw read CLI handlers (W-T2)', () => {
'orchestration dispatch-forget',
new Map<string, string | boolean>([
['task', 'task_1'],
['expected-failure-id', 'fail-1']
['expected-failure-id', 'fail-1'],
['from', 'term_coord'],
['run', 'run_1']
])
)
expect(callMock).toHaveBeenCalledWith('orchestration.dispatchForget', {
task: 'task_1',
run: 'run_1',
from: 'term_coord',
expectedFailureId: 'fail-1'
})
})
@@ -469,11 +472,40 @@ describe('orchestration dispatch Forget + raw read CLI handlers (W-T2)', () => {
await invoke(
'orchestration dispatch-forget',
new Map<string, string | boolean>([['task', 'task_1']])
new Map<string, string | boolean>([
['task', 'task_1'],
['from', 'term_coord']
])
)
expect(callMock).toHaveBeenCalledWith('orchestration.dispatchForget', {
task: 'task_1',
run: undefined,
from: 'term_coord',
expectedFailureId: undefined
})
})
// Owner authorization is Run-scoped, so an unattested coordinator terminal must
// still send a resolved handle rather than relying on the envelope fallback.
it('dispatch-forget sends the coordinator handle resolved from the live terminal', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_env'
callMock
.mockResolvedValueOnce({ result: { terminal: { handle: 'term_env' } } })
.mockResolvedValueOnce({
dispatch: { id: 'ctx_1', task_id: 'task_1', status: 'forgotten' }
})
await invoke(
'orchestration dispatch-forget',
new Map<string, string | boolean>([['task', 'task_1']])
)
expect(callMock).toHaveBeenCalledWith('terminal.show', { terminal: 'term_env' })
expect(callMock).toHaveBeenCalledWith('orchestration.dispatchForget', {
task: 'task_1',
run: undefined,
from: 'term_env',
expectedFailureId: undefined
})
})
@@ -634,279 +666,3 @@ describe('orchestration task-create caller handle', () => {
})
})
})
describe('orchestration timeout flag validation', () => {
const invalidTimeoutValues: [string, string | boolean][] = [
['missing', true],
['empty', ''],
['non-numeric', 'not-a-number'],
['zero', '0'],
['negative', '-1']
]
beforeEach(() => {
callMock.mockReset()
delete process.env.ORCA_TERMINAL_HANDLE
delete process.env.ORCA_PANE_KEY
})
const invokeCheck = handlerInvoker(ORCHESTRATION_HANDLERS['orchestration check'])
const invokeAsk = handlerInvoker(ORCHESTRATION_HANDLERS['orchestration ask'])
it.each(invalidTimeoutValues)('rejects invalid check --timeout-ms: %s', async (_label, value) => {
const flags = new Map<string, string | boolean>([
['wait', true],
['timeout-ms', value]
])
await expect(invokeCheck(flags)).rejects.toThrow(/--timeout-ms/)
expect(callMock).not.toHaveBeenCalled()
})
it('passes a parsed check timeout and peek mode into the RPC payload', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
callMock.mockResolvedValue({ result: { messages: [], count: 0 } })
await invokeCheck(
new Map<string, string | boolean>([
['wait', true],
['peek', true],
['timeout-ms', '250']
])
)
// Why: --peek rides with unread:false so pre-peek runtimes fall back to
// the non-consuming all mode instead of the destructive mark-read default.
expect(callMock).toHaveBeenCalledWith('orchestration.check', {
terminal: 'term_worker',
terminalPaneKey: undefined,
unread: false,
peek: true,
all: undefined,
types: undefined,
format: undefined,
compatibilityCliCommand: expect.stringMatching(/^orca(?:-ide)?$/),
run: undefined,
ack: undefined,
wait: true,
timeoutMs: 250
})
})
it('filters already-read rows from a peek response for pre-peek runtimes', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
callMock.mockResolvedValue({
result: {
messages: [
{ id: 'msg_old', from_handle: 'a', subject: 'seen', read: 1 },
{ id: 'msg_new', from_handle: 'a', subject: 'fresh', read: 0 }
],
count: 2,
formatted: 'banners built from all rows'
}
})
vi.mocked(printResult).mockClear()
await invokeCheck(new Map<string, string | boolean>([['peek', true]]))
const response = vi.mocked(printResult).mock.calls[0]?.[0] as {
result: { messages: { id: string }[]; count: number; formatted?: string }
}
expect(response.result.messages.map((m) => m.id)).toEqual(['msg_new'])
expect(response.result.count).toBe(1)
// Why: the pre-peek runtime built `formatted` from all rows, including
// the read one the filter just removed.
expect(response.result.formatted).toBeUndefined()
})
it('rejects combined read modes before calling the runtime', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
callMock.mockClear()
await expect(
invokeCheck(
new Map<string, string | boolean>([
['unread', true],
['peek', true]
])
)
).rejects.toMatchObject({
code: 'invalid_argument',
message: expect.stringContaining('read mode')
})
expect(callMock).not.toHaveBeenCalled()
})
it('warns when a pre-peek runtime returned a full 100-row page', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
const rows = Array.from({ length: 100 }, (_, i) => ({
id: `msg_${i}`,
from_handle: 'a',
subject: `s${i}`,
read: i === 0 ? 0 : 1
}))
callMock.mockResolvedValue({ result: { messages: rows, count: 100 } })
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
await invokeCheck(new Map<string, string | boolean>([['peek', true]]))
expect(errorSpy).toHaveBeenCalledWith(expect.stringContaining('newest 100 messages'))
errorSpy.mockRestore()
})
it('fails --peek --wait against a runtime that returned only read rows', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
callMock.mockResolvedValue({
result: {
messages: [{ id: 'msg_old', from_handle: 'a', subject: 'seen', read: 1 }],
count: 1
}
})
await expect(
invokeCheck(
new Map<string, string | boolean>([
['peek', true],
['wait', true]
])
)
).rejects.toMatchObject({ code: 'peek_wait_unsupported' })
})
it.each(invalidTimeoutValues)('rejects invalid ask --timeout-ms: %s', async (_label, value) => {
const flags = new Map<string, string | boolean>([
['to', 'term_coord'],
['question', 'Proceed?'],
['timeout-ms', value]
])
await expect(invokeAsk(flags)).rejects.toThrow(/--timeout-ms/)
expect(callMock).not.toHaveBeenCalled()
})
it('uses the parsed ask timeout for both runtime wait and client timeout', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
callMock.mockResolvedValue({
result: {
answer: 'yes',
messageId: 'msg_1',
threadId: 'thread_1',
timedOut: false
}
})
vi.spyOn(console, 'log').mockImplementation(() => {})
await invokeAsk(
new Map<string, string | boolean>([
['to', 'term_coord'],
['question', 'Proceed?'],
['timeout-ms', '123']
])
)
expect(callMock).toHaveBeenCalledWith(
'orchestration.ask',
{
to: 'term_coord',
run: undefined,
question: 'Proceed?',
resume: undefined,
options: undefined,
timeoutMs: 123,
from: 'term_worker',
compatibilityCliCommand: expect.stringMatching(/^orca(?:-ide)?$/),
compatibilityWindowsCommand: undefined
},
{ timeoutMs: 5_123, orchestrationCapability: undefined }
)
})
it('passes an ask resume without creating a new question payload', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
callMock.mockResolvedValue({
result: {
answer: 'yes',
messageId: 'msg_question',
threadId: 'msg_question',
timedOut: false
}
})
vi.spyOn(console, 'log').mockImplementation(() => {})
await invokeAsk(new Map<string, string | boolean>([['resume', 'msg_question']]))
expect(callMock).toHaveBeenCalledWith(
'orchestration.ask',
{
to: undefined,
run: undefined,
question: undefined,
resume: 'msg_question',
options: undefined,
timeoutMs: undefined,
from: 'term_worker',
compatibilityCliCommand: expect.stringMatching(/^orca(?:-ide)?$/),
compatibilityWindowsCommand: undefined
},
{ timeoutMs: 605_000, orchestrationCapability: undefined }
)
})
it('rejects ambiguous ask create/resume input before RPC', async () => {
process.env.ORCA_TERMINAL_HANDLE = 'term_worker'
await expect(
invokeAsk(
new Map<string, string | boolean>([
['question', 'new'],
['resume', 'msg_old']
])
)
).rejects.toMatchObject({ code: 'invalid_argument' })
expect(callMock).not.toHaveBeenCalled()
})
})
describe('orchestration task-list brief output', () => {
it('requests server-side brief and falls back client-side for older runtimes', async () => {
callMock.mockReset().mockResolvedValue({
result: {
// No spec_truncated field — the pre-brief-runtime signature.
tasks: [{ id: 'task_1', spec: `First line\n${'detail '.repeat(40)}`, status: 'ready' }],
count: 1
}
})
vi.mocked(printResult).mockClear()
await handlerInvoker(ORCHESTRATION_HANDLERS['orchestration task-list'])(
new Map([['brief', true]])
)
expect(callMock).toHaveBeenCalledWith(
'orchestration.taskList',
expect.objectContaining({ brief: true })
)
const response = vi.mocked(printResult).mock.calls[0]?.[0] as {
result: { tasks: { spec: string; spec_truncated: boolean }[] }
}
expect(response.result.tasks[0].spec).toHaveLength(160)
expect(response.result.tasks[0].spec_truncated).toBe(true)
})
it('passes server-abbreviated rows through untouched', async () => {
const serverTasks = [
{ id: 'task_1', spec: 'already brief…', status: 'ready', spec_truncated: true }
]
callMock.mockReset().mockResolvedValue({ result: { tasks: serverTasks, count: 1 } })
vi.mocked(printResult).mockClear()
await handlerInvoker(ORCHESTRATION_HANDLERS['orchestration task-list'])(
new Map([['brief', true]])
)
const response = vi.mocked(printResult).mock.calls[0]?.[0] as {
result: { tasks: { spec: string; spec_truncated: boolean }[] }
}
// Why: re-abbreviating a server-truncated spec would flip spec_truncated
// back to false (the truncated text fits the cap).
expect(response.result.tasks).toBe(serverTasks)
})
})
@@ -88,12 +88,14 @@ export const ORCHESTRATION_DISPATCH_INSPECTION_HANDLERS: Record<string, CommandH
})
},
'orchestration dispatch-forget': async ({ flags, client, json }) => {
'orchestration dispatch-forget': async ({ flags, client, cwd, json }) => {
const result = await client.call<{
dispatch: { id: string; task_id: string; status: string } | null
}>('orchestration.dispatchForget', {
task: getRequiredStringFlag(flags, 'task'),
expectedFailureId: getOptionalStringFlag(flags, 'expected-failure-id')
expectedFailureId: getOptionalStringFlag(flags, 'expected-failure-id'),
run: getOptionalStringFlag(flags, 'run'),
from: await resolveCoordinatorTerminalHandle(flags, cwd, client)
})
printResult(result, json, (value) => {
if (!value.dispatch) {
+2 -2
View File
@@ -200,8 +200,8 @@ export const ORCHESTRATION_COMMAND_SPECS: CommandSpec[] = [
path: ['orchestration', 'dispatch-forget'],
summary: 'Forget a dispatch stranded in an unknown launch state',
usage:
'orca orchestration dispatch-forget --task <task_id> [--expected-failure-id <id>] [--json]',
allowedFlags: [...GLOBAL_FLAGS, 'task', 'expected-failure-id'],
'orca orchestration dispatch-forget --task <task_id> [--expected-failure-id <id>] [--run <run_id>] [--from <handle>] [--json]',
allowedFlags: [...GLOBAL_FLAGS, 'task', 'expected-failure-id', 'run', 'from'],
notes: [
'The task returns to blocked; retry with: orca orchestration task-update --id <task_id> --status ready.'
]
@@ -0,0 +1,159 @@
// Catalog/reference authoring against the real Store: an ack means the bytes are on
// disk, a refused write is a typed failure, and payload budgets measure the projection
// persistence actually produces (which expands the caller's patch).
import { describe, it, expect, vi, beforeEach, afterEach } from 'vitest'
import type { GlobalSettings } from '../../shared/types'
import { mkdtempSync, readFileSync, rmSync } from 'node:fs'
import { join } from 'node:path'
import { tmpdir } from 'node:os'
const testState = { dir: '' }
vi.mock('electron', () => ({
app: { getPath: () => testState.dir },
safeStorage: {
isEncryptionAvailable: () => false,
encryptString: (plaintext: string) => Buffer.from(plaintext, 'utf-8'),
decryptString: (ciphertext: Buffer) => ciphertext.toString('utf-8')
}
}))
vi.mock('../telemetry/client', () => ({ track: vi.fn() }))
vi.mock('../telemetry/cohort-classifier', () => ({ getCohortAtEmit: vi.fn() }))
vi.mock('../ssh/ssh-config-parser', () => ({
loadUserSshConfig: vi.fn(() => null),
sshConfigHostsToTargets: vi.fn(() => [])
}))
async function createService(dataFile: string) {
vi.resetModules()
const { Store } = await import('../persistence')
const { AgentCatalogService } = await import('./agent-catalog-service')
const store = new Store({ dataFile })
return { store, service: new AgentCatalogService(store) }
}
function readPersisted(dataFile: string): { settings: GlobalSettings } {
return JSON.parse(readFileSync(dataFile, 'utf-8')) as { settings: GlobalSettings }
}
const CREATE_CODEX = {
expectedRevision: 1,
mutation: {
kind: 'create' as const,
baseAgent: 'codex' as const,
draft: {
label: 'Durable',
commandOverride: null,
args: '',
env: {},
syncEnv: false
}
}
}
describe('durable catalog authoring through the real Store', () => {
let dir = ''
let dataFile = ''
beforeEach(() => {
dir = mkdtempSync(join(tmpdir(), 'orca-agent-catalog-durable-'))
testState.dir = dir
dataFile = join(dir, 'orca-data.json')
})
afterEach(() => {
rmSync(dir, { recursive: true, force: true })
})
it('has the created agent on disk before the mutation returns ok', async () => {
const { service } = await createService(dataFile)
const result = service.mutate(CREATE_CODEX)
expect(result.ok).toBe(true)
// No debounce wait, no flush: the ack itself is the durability barrier.
expect(readPersisted(dataFile).settings.customTuiAgents).toHaveLength(1)
})
it('has the tombstoned deletion on disk before the mutation returns ok', async () => {
const { service, store } = await createService(dataFile)
const created = service.mutate(CREATE_CODEX)
expect(created.ok).toBe(true)
const [persistedAgent] = store.getSettings().customTuiAgents ?? []
const deleted = service.mutate({
expectedRevision: service.getRevision(),
mutation: { kind: 'delete-custom', id: persistedAgent.id }
})
expect(deleted.ok).toBe(true)
const persisted = readPersisted(dataFile).settings
expect(persisted.customTuiAgents).toHaveLength(0)
expect(persisted.deletedCustomTuiAgents).toHaveLength(1)
})
it('reports a typed failure instead of acking when the write cannot land', async () => {
const { service, store } = await createService(dataFile)
expect(service.mutate(CREATE_CODEX).ok).toBe(true)
store.freezeWrites()
const result = service.mutate({
expectedRevision: service.getRevision(),
mutation: {
kind: 'create',
baseAgent: 'claude',
draft: {
label: 'Lost',
commandOverride: null,
args: '',
env: {},
syncEnv: false
}
}
})
expect(result).toMatchObject({
ok: false,
code: 'agent_catalog_write_failed'
})
// Rolled back: an unacknowledged agent must not linger in memory either.
expect(store.getSettings().customTuiAgents).toHaveLength(1)
expect(readPersisted(dataFile).settings.customTuiAgents).toHaveLength(1)
})
it('rejects a reference write whose persisted projection expands past 512 KiB', async () => {
const { service, store } = await createService(dataFile)
// Under the cap as written; persistence derives legacy commitMessageAi from it,
// so the projection that ships carries the instructions twice.
const instructions = 'x'.repeat(300_000)
const result = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: {
kind: 'source-control-update',
changes: { instructionsByOperation: { commitMessage: instructions } }
}
})
expect(result).toMatchObject({
ok: false,
code: 'agent_reference_payload_too_large'
})
expect(store.getSettings().sourceControlAi?.instructionsByOperation?.commitMessage).not.toBe(
instructions
)
})
it('keeps a reference write whose persisted projection stays under 512 KiB', async () => {
const { service, store } = await createService(dataFile)
const result = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: {
kind: 'source-control-update',
changes: {
instructionsByOperation: { commitMessage: 'x'.repeat(1_000) }
}
}
})
expect(result.ok).toBe(true)
expect(
store.getSettings().sourceControlAi?.instructionsByOperation?.commitMessage
).toHaveLength(1_000)
expect(
readPersisted(dataFile).settings.sourceControlAi?.instructionsByOperation?.commitMessage
).toHaveLength(1_000)
})
})
@@ -6,6 +6,7 @@ import type {
GlobalSettings
} from '../../shared/types'
import type { CustomAgentDraft } from '../../shared/agent-catalog-snapshot'
import { tokenizeStartupCommand } from '../../shared/tui-agent-startup-shell'
import {
AgentCatalogRepairTokenRegistry,
applyAgentCatalogMutation,
@@ -323,6 +324,28 @@ describe('duplicate', () => {
expect(copy?.args).toBe('codex-real --fast --user-arg')
})
it('keeps a grouped prefix argument grouped instead of flattening it into two args', () => {
const result = apply({
settings: settingsWith({
agentCmdOverrides: { codex: '/opt/wrap --prompt "hello world"' },
agentDefaultArgs: { codex: '--user-arg' }
}),
mutation: { kind: 'duplicate', sourceAgent: 'codex', label: 'Wrapped' }
})
expect(result.ok).toBe(true)
if (!result.ok) {
return
}
const copy = result.patch.customTuiAgents?.[0]
expect(copy?.commandOverride).toBe('/opt/wrap')
// A bare join would store `--prompt hello world`, which re-splits into three.
expect(copy?.args).toBe('--prompt "hello world" --user-arg')
for (const shell of ['posix', 'powershell', 'cmd'] as const) {
const tokenized = tokenizeStartupCommand(copy?.args ?? '', shell)
expect(tokenized.ok && tokenized.tokens).toEqual(['--prompt', 'hello world', '--user-arg'])
}
})
it('rejects a platform-ambiguous built-in prefix instead of guessing a grammar', () => {
const result = apply({
settings: settingsWith({
@@ -13,10 +13,8 @@ import type {
} from '../../shared/types'
import type { AgentCatalogMutationRequest } from '../../shared/agent-catalog-snapshot'
import { normalizeAgentCatalog, type AgentCatalog } from '../../shared/custom-tui-agents'
import type {
AgentCatalogMutationApplication,
TombstoneReferenceCount
} from './agent-catalog-draft-validation'
import type { AgentCatalogMutationApplication } from './agent-catalog-draft-validation'
import type { TombstoneReferenceCounter } from './agent-catalog-tombstone-gc'
import {
applyCreate,
applyDelete,
@@ -41,8 +39,9 @@ export type ApplyAgentCatalogMutationArgs = {
currentRevision: number
repairTokens: AgentCatalogRepairTokenRegistry
/** Authoritative reference count per tombstone id; 'unknown' means an owner
* store could not be checked and the tombstone must be retained. */
countTombstoneReferences: (id: CustomTuiAgentId) => TombstoneReferenceCount
* store could not be checked and the tombstone must be retained. Accepts a
* batch counter so a prune indexes the owners once instead of per tombstone. */
countTombstoneReferences: TombstoneReferenceCounter
}
export type MutationContext = {
@@ -0,0 +1,69 @@
import { describe, expect, it } from 'vitest'
import type { CustomTuiAgentId } from '../../shared/types'
import { AgentTombstoneReferenceIndex } from './agent-tombstone-reference-index'
import { createBatchTombstoneReferenceCounter } from './agent-catalog-owner-scanners'
const idA = 'custom-agent:codex:fedcba98-7654-4321-8fed-cba987654321' as CustomTuiAgentId
const idB = 'custom-agent:claude:01234567-89ab-4cde-8f01-23456789abcd' as CustomTuiAgentId
const idC = 'custom-agent:gemini:01234567-89ab-4cde-8f01-23456789abce' as CustomTuiAgentId
function countForIds(
index: AgentTombstoneReferenceIndex,
ids: readonly CustomTuiAgentId[]
): ReadonlyMap<CustomTuiAgentId, number | 'unknown'> {
const counter = createBatchTombstoneReferenceCounter(index)
if (typeof counter === 'function') {
throw new Error('expected a batch counter')
}
return counter.countForIds(ids)
}
describe('createBatchTombstoneReferenceCounter', () => {
it('matches per-id countReferences for every id in one owner pass', () => {
const index = new AgentTombstoneReferenceIndex()
let scans = 0
index.register({
owner: 'automation',
scan: () => {
scans += 1
return { ok: true, referencedIds: [idA, idA, 'auto', null] }
}
})
index.register({
owner: 'workspace',
scan: () => {
scans += 1
return { ok: true, referencedIds: [idB] }
}
})
const batch = countForIds(index, [idA, idB, idC])
expect(scans).toBe(2)
expect([...batch]).toEqual([
[idA, index.countReferences(idA)],
[idB, index.countReferences(idB)],
[idC, index.countReferences(idC)]
])
expect(batch.get(idA)).toBe(2)
expect(batch.get(idB)).toBe(1)
expect(batch.get(idC)).toBe(0)
})
it('reports unknown for the whole batch when any owner cannot be read', () => {
const index = new AgentTombstoneReferenceIndex()
index.register({ owner: 'automation', scan: () => ({ ok: true, referencedIds: [idA] }) })
index.register({ owner: 'session', scan: () => ({ ok: false }) })
const batch = countForIds(index, [idA, idB])
// Same conservative retain the per-id path produces.
expect(index.countReferences(idB)).toBe('unknown')
expect(batch.get(idA)).toBe('unknown')
expect(batch.get(idB)).toBe('unknown')
})
it('answers an empty batch without claiming references', () => {
const index = new AgentTombstoneReferenceIndex()
index.register({ owner: 'automation', scan: () => ({ ok: true, referencedIds: [idA] }) })
expect([...countForIds(index, [])]).toEqual([])
})
})
@@ -0,0 +1,124 @@
// P1-9: the service's prune must index the owners ONCE per prune. A per-id
// counter re-scans every owner store per tombstone, which is quadratic on real
// catalogs, so these tests count owner scans rather than prune outcomes.
import { describe, expect, it } from 'vitest'
import type {
CustomTuiAgentId,
DeletedCustomTuiAgent,
GlobalSettings,
TerminalAgentQuickCommand
} from '../../shared/types'
import type { Store } from '../persistence'
import { AgentCatalogService } from './agent-catalog-service'
const UUID_TAIL = '-89ab-4cde-8f01-23456789abcd'
const TOMBSTONE_COUNT = 25
function customId(index: number): CustomTuiAgentId {
return `custom-agent:codex:${`${index}`.padStart(8, '0')}${UUID_TAIL}` as CustomTuiAgentId
}
function unreferencedTombstones(count: number): DeletedCustomTuiAgent[] {
return Array.from({ length: count }, (_, index) => ({
id: customId(index),
baseAgent: 'codex' as const,
label: `Gone ${index}`,
deletedAt: index
}))
}
type ScanCounts = { worktreeMeta: number; automations: number }
function makeCountingStore(settings: GlobalSettings): { store: Store; scans: ScanCounts } {
const state = { settings }
const scans: ScanCounts = { worktreeMeta: 0, automations: 0 }
const preview = (updates: Partial<GlobalSettings>): GlobalSettings => ({
...state.settings,
...updates
})
const stub = {
getSettings: () => state.settings,
getAgentCatalogMigrationError: () => null,
getAgentCatalogSchemaTooNew: () => null,
previewSettingsUpdate: preview,
updateSettingsDurable: (updates: Partial<GlobalSettings>) => {
state.settings = preview(updates)
return state.settings
},
updateSettings: (updates: Partial<GlobalSettings>) => {
state.settings = preview(updates)
return state.settings
},
getRepos: () => [],
listAutomations: () => {
scans.automations += 1
return []
},
listAutomationRuns: () => [],
getAllWorktreeMeta: () => {
scans.worktreeMeta += 1
return {}
}
}
return { store: stub as unknown as Store, scans }
}
function baseSettings(overrides: Partial<GlobalSettings> = {}): GlobalSettings {
return {
defaultTuiAgent: 'auto',
disabledTuiAgents: [],
customTuiAgents: [],
deletedCustomTuiAgents: [],
agentCatalogRevision: 1,
agentReferenceRevision: 1,
terminalQuickCommands: [],
agentCmdOverrides: {},
...overrides
} as GlobalSettings
}
function agentQuickCommand(agent: CustomTuiAgentId): TerminalAgentQuickCommand {
return { id: 'qc-1', label: 'Q', action: 'agent-prompt', agent, prompt: 'p' }
}
describe('agent catalog service tombstone prune batching (P1-9)', () => {
it('scans each owner once for a mutation prune, not once per tombstone', () => {
const { store, scans } = makeCountingStore(
baseSettings({ deletedCustomTuiAgents: unreferencedTombstones(TOMBSTONE_COUNT) })
)
const service = new AgentCatalogService(store)
const created = service.mutate({
expectedRevision: 1,
mutation: {
kind: 'create',
baseAgent: 'claude',
draft: { label: 'New One', commandOverride: null, args: '', env: {}, syncEnv: false }
}
})
expect(created.ok).toBe(true)
expect(store.getSettings().deletedCustomTuiAgents).toHaveLength(0)
expect(scans).toEqual({ worktreeMeta: 1, automations: 1 })
})
it('scans each owner once for the post-reference-removal prune', () => {
const { store, scans } = makeCountingStore(
baseSettings({
terminalQuickCommands: [agentQuickCommand(customId(0))],
deletedCustomTuiAgents: unreferencedTombstones(TOMBSTONE_COUNT)
})
)
const service = new AgentCatalogService(store)
const result = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: { kind: 'quick-command-delete', id: 'qc-1' }
})
expect(result.ok).toBe(true)
expect(store.getSettings().deletedCustomTuiAgents).toHaveLength(0)
expect(scans).toEqual({ worktreeMeta: 1, automations: 1 })
})
})
@@ -0,0 +1,108 @@
// Shared store stub and catalog fixtures for the AgentCatalogService suites.
import type {
CustomTuiAgent,
CustomTuiAgentId,
GlobalSettings,
Repo,
TerminalAgentQuickCommand,
WorktreeMeta
} from '../../shared/types'
import type { Automation, AutomationRun } from '../../shared/automations-types'
import type { Store } from '../persistence'
export const UUID_A = '01234567-89ab-4cde-8f01-23456789abcd'
export const UUID_B = 'fedcba98-7654-4321-8fed-cba987654321'
export function customId(base: string, uuid = UUID_A): CustomTuiAgentId {
return `custom-agent:${base}:${uuid}` as CustomTuiAgentId
}
export function liveAgent(overrides: Partial<CustomTuiAgent> = {}): CustomTuiAgent {
return {
id: customId('codex'),
baseAgent: 'codex',
label: 'My Codex',
args: '',
env: {},
syncEnv: false,
...overrides
}
}
export type StoreStubState = {
settings: GlobalSettings
repos: Repo[]
automations: Automation[]
automationRuns?: AutomationRun[]
worktreeMeta?: Record<string, WorktreeMeta>
failAutomationScan?: boolean
failWorktreeScan?: boolean
agentCatalogMigrationError?: string | null
agentCatalogSchemaTooNew?: { persistedVersion: number; supportedVersion: number } | null
failDurableWrite?: boolean
// Stands in for persistence-side normalization/derivation of a patch.
expandOnPersist?: (updates: Partial<GlobalSettings>) => Partial<GlobalSettings>
}
export function makeStoreStub(state: StoreStubState): Store {
const preview = (updates: Partial<GlobalSettings>): GlobalSettings => ({
...state.settings,
...updates,
...state.expandOnPersist?.(updates)
})
const stub = {
getSettings: () => state.settings,
getAgentCatalogMigrationError: () => state.agentCatalogMigrationError ?? null,
getAgentCatalogSchemaTooNew: () => state.agentCatalogSchemaTooNew ?? null,
previewSettingsUpdate: preview,
updateSettingsDurable: (updates: Partial<GlobalSettings>) => {
if (state.failDurableWrite) {
throw new Error('disk failure')
}
state.settings = preview(updates)
return state.settings
},
updateSettings: (updates: Partial<GlobalSettings>) => {
state.settings = preview(updates)
return state.settings
},
getRepos: () => state.repos,
listAutomations: () => {
if (state.failAutomationScan) {
throw new Error('store unavailable')
}
return state.automations
},
listAutomationRuns: () => state.automationRuns ?? [],
getAllWorktreeMeta: () => {
if (state.failWorktreeScan) {
throw new Error('store unavailable')
}
return state.worktreeMeta ?? {}
}
}
return stub as unknown as Store
}
export function baseSettings(overrides: Partial<GlobalSettings> = {}): GlobalSettings {
return {
defaultTuiAgent: 'auto',
disabledTuiAgents: [],
customTuiAgents: [],
deletedCustomTuiAgents: [],
agentCatalogRevision: 1,
agentReferenceRevision: 1,
terminalQuickCommands: [],
agentCmdOverrides: {},
...overrides
} as GlobalSettings
}
export function tombstoneFor(id: CustomTuiAgentId) {
return { id, baseAgent: 'codex' as const, label: 'Gone', deletedAt: 1 }
}
export function agentQuickCommand(agent: CustomTuiAgentId): TerminalAgentQuickCommand {
return { id: 'qc-1', label: 'Q', action: 'agent-prompt', agent, prompt: 'p' }
}
@@ -0,0 +1,332 @@
// Write-policy suites for AgentCatalogService: the pre-v1 migration gate, the
// read-only newer-schema gate, the reference payload budget, and the
// durable-before-ack contract.
import { describe, expect, it } from 'vitest'
import type { GlobalSettings } from '../../shared/types'
import { AgentCatalogService } from './agent-catalog-service'
import {
baseSettings,
makeStoreStub,
type StoreStubState
} from './agent-catalog-service-test-fixtures'
describe('pre-v1 migration gate', () => {
it('blocks catalog and reference mutations while the pinned pre-v1 backup has failed', () => {
const state: StoreStubState = {
settings: baseSettings(),
repos: [],
automations: [],
agentCatalogMigrationError: 'disk full'
}
const service = new AgentCatalogService(makeStoreStub(state))
const before = state.settings
const catalogResult = service.mutate({
expectedRevision: 1,
mutation: {
kind: 'create',
baseAgent: 'codex',
draft: { label: 'Blocked', commandOverride: null, args: '', env: {}, syncEnv: false }
}
})
expect(catalogResult).toEqual({
ok: false,
code: 'agent_catalog_migration_blocked',
migrationError: 'disk full'
})
const referenceResult = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: { kind: 'quick-command-delete', id: 'qc-1' }
})
expect(referenceResult).toEqual({
ok: false,
code: 'agent_catalog_migration_blocked',
migrationError: 'disk full'
})
// No v1 write of any kind may land on the unbacked-up profile.
expect(state.settings).toBe(before)
})
it('carries the block into the local snapshot so Settings surfaces it on load', () => {
const state: StoreStubState = {
settings: baseSettings(),
repos: [],
automations: [],
agentCatalogMigrationError: 'disk full'
}
const service = new AgentCatalogService(makeStoreStub(state))
expect(service.getLocalSnapshot().migrationBlockedError).toBe('disk full')
state.agentCatalogMigrationError = null
expect(service.getLocalSnapshot().migrationBlockedError).toBeUndefined()
})
it('projects the block to remote clients as a boolean only — never the error text', () => {
const state: StoreStubState = {
settings: baseSettings(),
repos: [],
automations: [],
agentCatalogMigrationError: 'disk full at /Users/someone/Library'
}
const service = new AgentCatalogService(makeStoreStub(state))
const remote = service.getRemoteSnapshot()
expect('customAgents' in remote && remote.migrationBlocked).toBe(true)
expect(JSON.stringify(remote)).not.toContain('disk full')
expect(JSON.stringify(remote)).not.toContain('/Users/')
state.agentCatalogMigrationError = null
const healthy = service.getRemoteSnapshot()
expect('migrationBlocked' in healthy).toBe(false)
})
})
describe('schema-newer-than-supported read-only gate', () => {
function readOnlyState(): StoreStubState {
return {
settings: baseSettings(),
repos: [],
automations: [],
agentCatalogSchemaTooNew: { persistedVersion: 2, supportedVersion: 1 }
}
}
it('rejects catalog and reference mutations up front instead of via the write refusal', () => {
const state = readOnlyState()
// A durable write here would be persistence's refusal, reported as a generic
// failure; the gate must answer before any write is attempted.
state.failDurableWrite = true
const service = new AgentCatalogService(makeStoreStub(state))
const before = state.settings
expect(
service.mutate({
expectedRevision: 1,
mutation: {
kind: 'create',
baseAgent: 'codex',
draft: { label: 'Blocked', commandOverride: null, args: '', env: {}, syncEnv: false }
}
})
).toEqual({
ok: false,
code: 'agent_catalog_schema_too_new',
persistedVersion: 2,
supportedVersion: 1
})
expect(
service.mutateReferences({
expectedReferenceRevision: 1,
mutation: { kind: 'quick-command-delete', id: 'qc-1' }
})
).toEqual({
ok: false,
code: 'agent_catalog_schema_too_new',
persistedVersion: 2,
supportedVersion: 1
})
expect(state.settings).toBe(before)
})
it('carries the read-only state into the local snapshot so Settings surfaces it on load', () => {
const state = readOnlyState()
const service = new AgentCatalogService(makeStoreStub(state))
expect(service.getLocalSnapshot().schemaTooNew).toEqual({
persistedVersion: 2,
supportedVersion: 1
})
state.agentCatalogSchemaTooNew = null
expect(service.getLocalSnapshot().schemaTooNew).toBeUndefined()
})
it('leaves the remote projection unchanged: no new wire field for old clients', () => {
const service = new AgentCatalogService(makeStoreStub(readOnlyState()))
const remote = service.getRemoteSnapshot()
expect('migrationBlocked' in remote).toBe(false)
expect(JSON.stringify(remote)).not.toContain('schemaTooNew')
})
})
describe('reference payload budget (L1-#2)', () => {
function sourceControlAiWith(
instructions: Partial<Record<'commitMessage' | 'pullRequest' | 'branchName', string>>
): GlobalSettings['sourceControlAi'] {
return {
enabled: true,
agentId: null,
actions: {},
selectedModelByAgent: {},
selectedThinkingByModel: {},
customAgentCommand: '',
instructionsByOperation: instructions
} as GlobalSettings['sourceControlAi']
}
function serviceWith(settings: GlobalSettings): {
service: AgentCatalogService
state: StoreStubState
} {
const state: StoreStubState = { settings, repos: [], automations: [] }
return { service: new AgentCatalogService(makeStoreStub(state)), state }
}
it('rejects a growing write that pushes the reference snapshot past 512 KiB', () => {
const { service, state } = serviceWith(
baseSettings({
sourceControlAi: sourceControlAiWith({ commitMessage: 'x'.repeat(600_000) })
})
)
const before = state.settings
const result = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: {
kind: 'source-control-update',
changes: { customAgentCommand: 'generate-with-huge-profile' }
}
})
expect(result).toMatchObject({
ok: false,
code: 'agent_reference_payload_too_large',
referenceRevision: 1
})
// A rejected budget check performs no write.
expect(state.settings).toBe(before)
})
it('still commits a shrinking write while the snapshot is over budget', () => {
const { service, state } = serviceWith(
baseSettings({
sourceControlAi: sourceControlAiWith({
commitMessage: 'x'.repeat(600_000),
pullRequest: 'y'.repeat(600_000)
})
})
)
const result = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: {
kind: 'source-control-update',
changes: {
instructionsByOperation: { commitMessage: '', pullRequest: 'y'.repeat(600_000) }
}
}
})
// Still > 512 KiB after the write, but smaller than before: the user must be
// able to edit an over-budget profile back under budget.
expect(result.ok).toBe(true)
expect(state.settings.sourceControlAi?.instructionsByOperation?.commitMessage).toBe('')
})
it('keeps a small write under budget flowing normally', () => {
const { service } = serviceWith(baseSettings())
const result = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: { kind: 'source-control-update', changes: { customAgentCommand: 'ok' } }
})
expect(result.ok).toBe(true)
})
it('budgets the persisted projection, not the pre-persistence patch', () => {
// Persistence derives legacy commitMessageAi from sourceControlAi, so an
// under-budget patch lands as a projection carrying the instructions twice.
const state: StoreStubState = {
settings: baseSettings(),
repos: [],
automations: [],
expandOnPersist: (updates) =>
updates.sourceControlAi
? {
commitMessageAi: {
enabled: true,
agentId: null,
selectedModelByAgent: {},
selectedThinkingByModel: {},
customPrompt: updates.sourceControlAi.instructionsByOperation?.commitMessage ?? '',
customAgentCommand: ''
} as GlobalSettings['commitMessageAi']
}
: {}
}
const service = new AgentCatalogService(makeStoreStub(state))
const before = state.settings
const result = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: {
kind: 'source-control-update',
changes: { instructionsByOperation: { commitMessage: 'x'.repeat(300_000) } }
}
})
expect(result).toMatchObject({ ok: false, code: 'agent_reference_payload_too_large' })
expect(state.settings).toBe(before)
})
})
describe('durable authoring acknowledgement (P0-2)', () => {
function serviceWith(state: StoreStubState): AgentCatalogService {
const store = makeStoreStub(state)
// Any debounced write here would ack before the bytes are durable.
Object.assign(store, {
updateSettings: () => {
throw new Error('authoring must not use the debounced settings write')
}
})
return new AgentCatalogService(store)
}
const createCodex = {
expectedRevision: 1,
mutation: {
kind: 'create' as const,
baseAgent: 'codex' as const,
draft: { label: 'Durable', commandOverride: null, args: '', env: {}, syncEnv: false }
}
}
it('commits catalog and reference mutations through the durable write path', () => {
const state: StoreStubState = { settings: baseSettings(), repos: [], automations: [] }
const service = serviceWith(state)
expect(service.mutate(createCodex).ok).toBe(true)
expect(state.settings.customTuiAgents).toHaveLength(1)
const referenceResult = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: { kind: 'source-control-update', changes: { customAgentCommand: 'ok' } }
})
expect(referenceResult.ok).toBe(true)
})
it('reports a typed failure and keeps state unchanged when the durable write fails', () => {
const state: StoreStubState = {
settings: baseSettings(),
repos: [],
automations: [],
failDurableWrite: true
}
const service = serviceWith(state)
const revisions: number[] = []
service.onDidChange((revision) => revisions.push(revision))
const before = state.settings
expect(service.mutate(createCodex)).toEqual({
ok: false,
code: 'agent_catalog_write_failed',
revision: 1
})
expect(
service.mutateReferences({
expectedReferenceRevision: 1,
mutation: { kind: 'source-control-update', changes: { customAgentCommand: 'ok' } }
})
).toEqual({
ok: false,
code: 'agent_reference_write_failed',
referenceRevision: 1,
catalogRevision: 1
})
expect(state.settings).toBe(before)
expect(revisions).toEqual([])
})
})
@@ -1,113 +1,21 @@
import { afterEach, describe, expect, it, vi } from 'vitest'
import type {
CustomTuiAgent,
CustomTuiAgentId,
GlobalSettings,
Repo,
TerminalAgentQuickCommand,
TuiAgent,
WorktreeMeta
} from '../../shared/types'
import type { GlobalSettings, Repo, TuiAgent, WorktreeMeta } from '../../shared/types'
import type { Automation, AutomationRun } from '../../shared/automations-types'
import type { Store } from '../persistence'
import { AgentCatalogService } from './agent-catalog-service'
import { getHostAgentSessionRecordStore } from './agent-session-record-store-host'
import type { HostSessionLaunchRecord } from './agent-session-record-store'
import { getHostBackgroundAgentLaunchStore } from './background-agent-launch-store-host'
const UUID_A = '01234567-89ab-4cde-8f01-23456789abcd'
const UUID_B = 'fedcba98-7654-4321-8fed-cba987654321'
function customId(base: string, uuid = UUID_A): CustomTuiAgentId {
return `custom-agent:${base}:${uuid}` as CustomTuiAgentId
}
function liveAgent(overrides: Partial<CustomTuiAgent> = {}): CustomTuiAgent {
return {
id: customId('codex'),
baseAgent: 'codex',
label: 'My Codex',
args: '',
env: {},
syncEnv: false,
...overrides
}
}
type StoreStubState = {
settings: GlobalSettings
repos: Repo[]
automations: Automation[]
automationRuns?: AutomationRun[]
worktreeMeta?: Record<string, WorktreeMeta>
failAutomationScan?: boolean
failWorktreeScan?: boolean
agentCatalogMigrationError?: string | null
failDurableWrite?: boolean
// Stands in for persistence-side normalization/derivation of a patch.
expandOnPersist?: (updates: Partial<GlobalSettings>) => Partial<GlobalSettings>
}
function makeStoreStub(state: StoreStubState): Store {
const preview = (updates: Partial<GlobalSettings>): GlobalSettings => ({
...state.settings,
...updates,
...state.expandOnPersist?.(updates)
})
const stub = {
getSettings: () => state.settings,
getAgentCatalogMigrationError: () => state.agentCatalogMigrationError ?? null,
previewSettingsUpdate: preview,
updateSettingsDurable: (updates: Partial<GlobalSettings>) => {
if (state.failDurableWrite) {
throw new Error('disk failure')
}
state.settings = preview(updates)
return state.settings
},
updateSettings: (updates: Partial<GlobalSettings>) => {
state.settings = preview(updates)
return state.settings
},
getRepos: () => state.repos,
listAutomations: () => {
if (state.failAutomationScan) {
throw new Error('store unavailable')
}
return state.automations
},
listAutomationRuns: () => state.automationRuns ?? [],
getAllWorktreeMeta: () => {
if (state.failWorktreeScan) {
throw new Error('store unavailable')
}
return state.worktreeMeta ?? {}
}
}
return stub as unknown as Store
}
function baseSettings(overrides: Partial<GlobalSettings> = {}): GlobalSettings {
return {
defaultTuiAgent: 'auto',
disabledTuiAgents: [],
customTuiAgents: [],
deletedCustomTuiAgents: [],
agentCatalogRevision: 1,
agentReferenceRevision: 1,
terminalQuickCommands: [],
agentCmdOverrides: {},
...overrides
} as GlobalSettings
}
function tombstoneFor(id: CustomTuiAgentId) {
return { id, baseAgent: 'codex' as const, label: 'Gone', deletedAt: 1 }
}
function agentQuickCommand(agent: CustomTuiAgentId): TerminalAgentQuickCommand {
return { id: 'qc-1', label: 'Q', action: 'agent-prompt', agent, prompt: 'p' }
}
import {
UUID_A,
UUID_B,
agentQuickCommand,
baseSettings,
customId,
liveAgent,
makeStoreStub,
tombstoneFor,
type StoreStubState
} from './agent-catalog-service-test-fixtures'
describe('tombstone reference GC across owners', () => {
const deadId = customId('codex', UUID_B)
@@ -596,255 +504,3 @@ describe('local draft endpoint', () => {
}
})
})
describe('pre-v1 migration gate', () => {
it('blocks catalog and reference mutations while the pinned pre-v1 backup has failed', () => {
const state: StoreStubState = {
settings: baseSettings(),
repos: [],
automations: [],
agentCatalogMigrationError: 'disk full'
}
const service = new AgentCatalogService(makeStoreStub(state))
const before = state.settings
const catalogResult = service.mutate({
expectedRevision: 1,
mutation: {
kind: 'create',
baseAgent: 'codex',
draft: { label: 'Blocked', commandOverride: null, args: '', env: {}, syncEnv: false }
}
})
expect(catalogResult).toEqual({
ok: false,
code: 'agent_catalog_migration_blocked',
migrationError: 'disk full'
})
const referenceResult = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: { kind: 'quick-command-delete', id: 'qc-1' }
})
expect(referenceResult).toEqual({
ok: false,
code: 'agent_catalog_migration_blocked',
migrationError: 'disk full'
})
// No v1 write of any kind may land on the unbacked-up profile.
expect(state.settings).toBe(before)
})
it('carries the block into the local snapshot so Settings surfaces it on load', () => {
const state: StoreStubState = {
settings: baseSettings(),
repos: [],
automations: [],
agentCatalogMigrationError: 'disk full'
}
const service = new AgentCatalogService(makeStoreStub(state))
expect(service.getLocalSnapshot().migrationBlockedError).toBe('disk full')
state.agentCatalogMigrationError = null
expect(service.getLocalSnapshot().migrationBlockedError).toBeUndefined()
})
it('projects the block to remote clients as a boolean only — never the error text', () => {
const state: StoreStubState = {
settings: baseSettings(),
repos: [],
automations: [],
agentCatalogMigrationError: 'disk full at /Users/someone/Library'
}
const service = new AgentCatalogService(makeStoreStub(state))
const remote = service.getRemoteSnapshot()
expect('customAgents' in remote && remote.migrationBlocked).toBe(true)
expect(JSON.stringify(remote)).not.toContain('disk full')
expect(JSON.stringify(remote)).not.toContain('/Users/')
state.agentCatalogMigrationError = null
const healthy = service.getRemoteSnapshot()
expect('migrationBlocked' in healthy).toBe(false)
})
})
describe('reference payload budget (L1-#2)', () => {
function sourceControlAiWith(
instructions: Partial<Record<'commitMessage' | 'pullRequest' | 'branchName', string>>
): GlobalSettings['sourceControlAi'] {
return {
enabled: true,
agentId: null,
actions: {},
selectedModelByAgent: {},
selectedThinkingByModel: {},
customAgentCommand: '',
instructionsByOperation: instructions
} as GlobalSettings['sourceControlAi']
}
function serviceWith(settings: GlobalSettings): {
service: AgentCatalogService
state: StoreStubState
} {
const state: StoreStubState = { settings, repos: [], automations: [] }
return { service: new AgentCatalogService(makeStoreStub(state)), state }
}
it('rejects a growing write that pushes the reference snapshot past 512 KiB', () => {
const { service, state } = serviceWith(
baseSettings({
sourceControlAi: sourceControlAiWith({ commitMessage: 'x'.repeat(600_000) })
})
)
const before = state.settings
const result = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: {
kind: 'source-control-update',
changes: { customAgentCommand: 'generate-with-huge-profile' }
}
})
expect(result).toMatchObject({
ok: false,
code: 'agent_reference_payload_too_large',
referenceRevision: 1
})
// A rejected budget check performs no write.
expect(state.settings).toBe(before)
})
it('still commits a shrinking write while the snapshot is over budget', () => {
const { service, state } = serviceWith(
baseSettings({
sourceControlAi: sourceControlAiWith({
commitMessage: 'x'.repeat(600_000),
pullRequest: 'y'.repeat(600_000)
})
})
)
const result = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: {
kind: 'source-control-update',
changes: {
instructionsByOperation: { commitMessage: '', pullRequest: 'y'.repeat(600_000) }
}
}
})
// Still > 512 KiB after the write, but smaller than before: the user must be
// able to edit an over-budget profile back under budget.
expect(result.ok).toBe(true)
expect(state.settings.sourceControlAi?.instructionsByOperation?.commitMessage).toBe('')
})
it('keeps a small write under budget flowing normally', () => {
const { service } = serviceWith(baseSettings())
const result = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: { kind: 'source-control-update', changes: { customAgentCommand: 'ok' } }
})
expect(result.ok).toBe(true)
})
it('budgets the persisted projection, not the pre-persistence patch', () => {
// Persistence derives legacy commitMessageAi from sourceControlAi, so an
// under-budget patch lands as a projection carrying the instructions twice.
const state: StoreStubState = {
settings: baseSettings(),
repos: [],
automations: [],
expandOnPersist: (updates) =>
updates.sourceControlAi
? {
commitMessageAi: {
enabled: true,
agentId: null,
selectedModelByAgent: {},
selectedThinkingByModel: {},
customPrompt: updates.sourceControlAi.instructionsByOperation?.commitMessage ?? '',
customAgentCommand: ''
} as GlobalSettings['commitMessageAi']
}
: {}
}
const service = new AgentCatalogService(makeStoreStub(state))
const before = state.settings
const result = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: {
kind: 'source-control-update',
changes: { instructionsByOperation: { commitMessage: 'x'.repeat(300_000) } }
}
})
expect(result).toMatchObject({ ok: false, code: 'agent_reference_payload_too_large' })
expect(state.settings).toBe(before)
})
})
describe('durable authoring acknowledgement (P0-2)', () => {
function serviceWith(state: StoreStubState): AgentCatalogService {
const store = makeStoreStub(state)
// Any debounced write here would ack before the bytes are durable.
Object.assign(store, {
updateSettings: () => {
throw new Error('authoring must not use the debounced settings write')
}
})
return new AgentCatalogService(store)
}
const createCodex = {
expectedRevision: 1,
mutation: {
kind: 'create' as const,
baseAgent: 'codex' as const,
draft: { label: 'Durable', commandOverride: null, args: '', env: {}, syncEnv: false }
}
}
it('commits catalog and reference mutations through the durable write path', () => {
const state: StoreStubState = { settings: baseSettings(), repos: [], automations: [] }
const service = serviceWith(state)
expect(service.mutate(createCodex).ok).toBe(true)
expect(state.settings.customTuiAgents).toHaveLength(1)
const referenceResult = service.mutateReferences({
expectedReferenceRevision: 1,
mutation: { kind: 'source-control-update', changes: { customAgentCommand: 'ok' } }
})
expect(referenceResult.ok).toBe(true)
})
it('reports a typed failure and keeps state unchanged when the durable write fails', () => {
const state: StoreStubState = {
settings: baseSettings(),
repos: [],
automations: [],
failDurableWrite: true
}
const service = serviceWith(state)
const revisions: number[] = []
service.onDidChange((revision) => revisions.push(revision))
const before = state.settings
expect(service.mutate(createCodex)).toEqual({
ok: false,
code: 'agent_catalog_write_failed',
revision: 1
})
expect(
service.mutateReferences({
expectedReferenceRevision: 1,
mutation: { kind: 'source-control-update', changes: { customAgentCommand: 'ok' } }
})
).toEqual({
ok: false,
code: 'agent_reference_write_failed',
referenceRevision: 1,
catalogRevision: 1
})
expect(state.settings).toBe(before)
expect(revisions).toEqual([])
})
})
+52 -125
View File
@@ -17,7 +17,6 @@ import {
AgentCatalogRepairTokenRegistry,
applyAgentCatalogMutation
} from './agent-catalog-mutations'
import { buildUnreferencedTombstonePrunePatch } from './agent-catalog-tombstone-gc'
import {
buildAgentCatalogSnapshot,
buildLocalAgentCatalogSnapshot,
@@ -29,15 +28,16 @@ import {
AgentTombstoneReferenceIndex,
type AgentReferenceSummary
} from './agent-tombstone-reference-index'
import { registerBuiltInOwnerScanners } from './agent-catalog-owner-scanners'
import {
buildAgentReferenceSnapshot,
measureAgentReferenceProjection
} from './agent-reference-snapshot-projection'
createBatchTombstoneReferenceCounter,
registerBuiltInOwnerScanners
} from './agent-catalog-owner-scanners'
import {
agentCatalogMigrationBlockedError,
agentCatalogSchemaTooNewError,
isSecurityReducingMutation,
type AgentCatalogMigrationBlockedError
type AgentCatalogMigrationBlockedError,
type AgentCatalogSchemaTooNewError
} from './agent-catalog-write-policy'
import { computeBaseDisableImpact } from './agent-catalog-base-disable-impact'
import {
@@ -45,7 +45,7 @@ import {
type AgentCatalogWriteFailedError,
type AgentReferenceWriteFailedError
} from './agent-catalog-durable-write'
import { applyAgentReferenceMutation } from './agent-reference-mutations'
import { AgentReferenceAuthoringService } from './agent-reference-authoring-service'
import type {
AgentReferenceMutationRequest,
AgentReferenceMutationResult,
@@ -72,10 +72,25 @@ export function getOrCreateAgentCatalogService(store: Store): AgentCatalogServic
export class AgentCatalogService {
private readonly repairTokens = new AgentCatalogRepairTokenRegistry()
private readonly referenceIndex = new AgentTombstoneReferenceIndex()
// Reads the index live, so scanners registered after construction still count.
private readonly tombstoneReferenceCounter = createBatchTombstoneReferenceCounter(
this.referenceIndex
)
private readonly changeListeners = new Set<(revision: number) => void>()
private readonly references: AgentReferenceAuthoringService
constructor(private readonly store: Store) {
registerBuiltInOwnerScanners(this.referenceIndex, this.store)
this.references = new AgentReferenceAuthoringService({
store,
countTombstoneReferences: this.tombstoneReferenceCounter,
catalogRevision: () => this.getRevision(),
publishCatalogRevision: (revision) => {
for (const listener of this.changeListeners) {
listener(revision)
}
}
})
}
/** Later units (worktree pending launches, background attempts, orchestration,
@@ -96,9 +111,20 @@ export class AgentCatalogService {
}
getLocalSnapshot(): LocalAgentCatalogSnapshot {
const snapshot = buildLocalAgentCatalogSnapshot(this.store.getSettings(), this.repairTokens)
const base = buildLocalAgentCatalogSnapshot(this.store.getSettings(), this.repairTokens)
const blocked = agentCatalogMigrationBlockedError(this.store)
return blocked ? { ...snapshot, migrationBlockedError: blocked.migrationError } : snapshot
const snapshot = blocked ? { ...base, migrationBlockedError: blocked.migrationError } : base
const tooNew = agentCatalogSchemaTooNewError(this.store)
if (!tooNew) {
return snapshot
}
return {
...snapshot,
schemaTooNew: {
persistedVersion: tooNew.persistedVersion,
supportedVersion: tooNew.supportedVersion
}
}
}
getRemoteSnapshot(): ReturnType<typeof buildAgentCatalogSnapshot> {
@@ -167,37 +193,15 @@ export class AgentCatalogService {
}
getReferenceRevision(): number {
return this.store.getSettings().agentReferenceRevision ?? 1
return this.references.getRevision()
}
/** Remote (runtime RPC) reference snapshot; typed projection error when over
* the 512 KiB frame budget. */
getRemoteReferenceSnapshot(): AgentReferenceSnapshot | AgentReferenceProjectionError {
const settings = this.store.getSettings()
const snapshot = buildAgentReferenceSnapshot(settings)
const { tooLarge } = measureAgentReferenceProjection(settings)
if (tooLarge) {
return {
version: 1,
revision: snapshot.revision,
code: 'agent_reference_payload_too_large',
maxBytes: 524_288
}
}
return snapshot
return this.references.getRemoteSnapshot()
}
/** Uncapped authoring/repair view over local preload IPC only. */
getLocalReferenceSnapshot(): LocalAgentReferenceSnapshot {
const settings = this.store.getSettings()
const snapshot = buildAgentReferenceSnapshot(settings)
const { bytes, tooLarge } = measureAgentReferenceProjection(settings)
return {
...snapshot,
projection: tooLarge
? { status: 'too-large', bytes, maxBytes: 524_288 }
: { status: 'ready', bytes, maxBytes: 524_288 }
}
return this.references.getLocalSnapshot()
}
mutateReferences(
@@ -205,107 +209,30 @@ export class AgentCatalogService {
):
| AgentReferenceMutationResult<LocalAgentReferenceSnapshot>
| AgentCatalogMigrationBlockedError
| AgentCatalogSchemaTooNewError
| AgentReferenceWriteFailedError {
// The failed-backup invariant is "no v1 write": reference mutations stamp
// agentReferenceRevision, so they are blocked alongside catalog mutations.
const blocked = agentCatalogMigrationBlockedError(this.store)
if (blocked) {
return blocked
}
const settings = this.store.getSettings()
const currentReferenceRevision = settings.agentReferenceRevision ?? 1
const application = applyAgentReferenceMutation({
settings,
request,
currentReferenceRevision,
catalog: normalizeCatalogFromSettings(settings)
})
if (!application.ok) {
return {
ok: false,
code: application.code,
referenceRevision: currentReferenceRevision,
catalogRevision: this.getRevision(),
...(application.code === 'reference_revision_conflict'
? { snapshot: this.getLocalReferenceSnapshot() }
: {}),
...(application.owner ? { owner: application.owner } : {}),
...(application.field ? { field: application.field } : {}),
...(application.reason ? { reason: application.reason } : {})
}
}
// The 512 KiB remote-projection budget is checked on the post-mutation
// snapshot; a non-growing write still commits so an over-budget profile can
// always be edited back under budget (mirrors mutate()'s security-reducing
// allowlist). Why preview: persistence normalizes the patch and derives
// legacy commitMessageAi from sourceControlAi, so the patch as written is
// smaller than the projection that actually ships.
const projected = measureAgentReferenceProjection(
this.store.previewSettingsUpdate(application.patch)
)
if (projected.tooLarge && projected.bytes > measureAgentReferenceProjection(settings).bytes) {
return {
ok: false,
code: 'agent_reference_payload_too_large',
referenceRevision: currentReferenceRevision,
catalogRevision: this.getRevision()
}
}
// Owner change commits before any prune; a failure between the two leaves
// the tombstone conservatively retained for the next indexed recheck.
if (!commitAuthoringPatchDurable(this.store, application.patch)) {
return {
ok: false,
code: 'agent_reference_write_failed',
referenceRevision: currentReferenceRevision,
catalogRevision: this.getRevision()
}
}
this.pruneUnreferencedTombstonesAfterReferenceRemoval()
return {
ok: true,
referenceRevision: application.newReferenceRevision,
catalogRevision: this.getRevision(),
snapshot: this.getLocalReferenceSnapshot()
}
}
/** Reference-aware prune run after a reference removal; a prune advances and
* publishes the catalog revision so receivers replace their snapshot. */
private pruneUnreferencedTombstonesAfterReferenceRemoval(): void {
const settings = this.store.getSettings()
// The patch also strips any persisted row the pruned tombstone suppressed,
// so the prune can never resurrect it.
const prunePatch = buildUnreferencedTombstonePrunePatch(
settings,
normalizeCatalogFromSettings(settings),
(id) => this.referenceIndex.countReferences(id)
)
if (!prunePatch) {
return
}
const newRevision = (settings.agentCatalogRevision ?? 1) + 1
// GC only: a rolled-back prune keeps the tombstone retained, and the next
// indexed recheck retries it — the reference mutation is already durable.
if (
!commitAuthoringPatchDurable(this.store, { ...prunePatch, agentCatalogRevision: newRevision })
) {
return
}
for (const listener of this.changeListeners) {
listener(newRevision)
}
return this.references.mutate(request)
}
mutate(
request: AgentCatalogMutationRequest
): AgentCatalogMutationResult | AgentCatalogMigrationBlockedError | AgentCatalogWriteFailedError {
):
| AgentCatalogMutationResult
| AgentCatalogMigrationBlockedError
| AgentCatalogSchemaTooNewError
| AgentCatalogWriteFailedError {
// A failed pinned pre-v1 backup means no v1 write may land on the profile;
// fail closed here so authoring cannot bypass the migration invariant.
const blocked = agentCatalogMigrationBlockedError(this.store)
if (blocked) {
return blocked
}
// Persistence refuses the write anyway; reporting it up front keeps a
// non-retryable read-only profile from looking like a disk failure.
const tooNew = agentCatalogSchemaTooNewError(this.store)
if (tooNew) {
return tooNew
}
const settings = this.store.getSettings()
const currentRevision = settings.agentCatalogRevision ?? 1
const application = applyAgentCatalogMutation({
@@ -313,7 +240,7 @@ export class AgentCatalogService {
request,
currentRevision,
repairTokens: this.repairTokens,
countTombstoneReferences: (id) => this.referenceIndex.countReferences(id)
countTombstoneReferences: this.tombstoneReferenceCounter
})
if (!application.ok) {
return {
@@ -0,0 +1,128 @@
import { describe, expect, it } from 'vitest'
import type { CustomTuiAgentId, DeletedCustomTuiAgent } from '../../shared/types'
import { pruneTombstones, type TombstoneReferenceCounter } from './agent-catalog-tombstone-gc'
const UUID_A = '01234567-89ab-4cde-8f01-23456789abcd'
const UUID_B = 'fedcba98-7654-4321-8fed-cba987654321'
function customId(base: string, uuid = UUID_A): CustomTuiAgentId {
return `custom-agent:${base}:${uuid}` as CustomTuiAgentId
}
function sequentialId(index: number): CustomTuiAgentId {
return customId('codex', `${index}`.padStart(8, '0') + UUID_A.slice(8))
}
function tombstonesFor(ids: readonly CustomTuiAgentId[]): DeletedCustomTuiAgent[] {
return ids.map((id, index) => ({ id, baseAgent: 'codex', label: `T${index}`, deletedAt: index }))
}
type OwnerFixture = { referencedIds: readonly string[]; readable: boolean }
/** Pre-fix shape: one full owner sweep per tombstone. */
function perIdCounter(owners: readonly OwnerFixture[], sweeps: { count: number }) {
return (id: CustomTuiAgentId): number | 'unknown' => {
let total = 0
for (const owner of owners) {
sweeps.count += 1
if (!owner.readable) {
return 'unknown'
}
total += owner.referencedIds.filter((value) => value === id).length
}
return total
}
}
/** Post-fix shape: owners indexed once, then a single pass over tombstones. */
function batchCounter(
owners: readonly OwnerFixture[],
sweeps: { count: number }
): TombstoneReferenceCounter {
return {
countForIds: (ids) => {
const wanted = new Set<string>(ids)
const matched = new Map<string, number>()
let complete = true
for (const owner of owners) {
sweeps.count += 1
if (!owner.readable) {
complete = false
continue
}
for (const value of owner.referencedIds) {
if (wanted.has(value)) {
matched.set(value, (matched.get(value) ?? 0) + 1)
}
}
}
const counts = new Map<CustomTuiAgentId, number | 'unknown'>()
for (const id of ids) {
counts.set(id, complete ? (matched.get(id) ?? 0) : 'unknown')
}
return counts
}
}
}
describe('pruneTombstones batch counting (P1-9)', () => {
const unreferenced = customId('codex', UUID_A)
const referencedOnce = customId('claude', UUID_B)
const referencedTwice = customId('gemini', UUID_A)
// Mixed: unreferenced (twice over, so a duplicate id is covered), single and
// multi reference, plus owner values that match no tombstone.
const mixed = [unreferenced, referencedOnce, referencedTwice, unreferenced]
const owners: OwnerFixture[] = [
{ referencedIds: [referencedOnce, 'auto'], readable: true },
{ referencedIds: [], readable: true },
{ referencedIds: [referencedTwice, referencedTwice, 'claude'], readable: true }
]
it('produces identical retain/prune decisions for both counter shapes', () => {
const tombstones = tombstonesFor(mixed)
const perId = pruneTombstones(tombstones, perIdCounter(owners, { count: 0 }))
const batch = pruneTombstones(tombstones, batchCounter(owners, { count: 0 }))
expect(batch).toEqual(perId)
expect(perId.prunedIds).toEqual([unreferenced, unreferenced])
expect(perId.retained.map((entry) => entry.id)).toEqual([referencedOnce, referencedTwice])
})
it('retains everything when an owner is unreadable, under both counter shapes', () => {
const tombstones = tombstonesFor(mixed)
const withFailure = [...owners, { referencedIds: [], readable: false }]
const perId = pruneTombstones(tombstones, perIdCounter(withFailure, { count: 0 }))
const batch = pruneTombstones(tombstones, batchCounter(withFailure, { count: 0 }))
expect(batch).toEqual(perId)
expect(perId.prunedIds).toEqual([])
expect(perId.retained).toHaveLength(mixed.length)
})
it('retains a tombstone the counter omits rather than treating it as unreferenced', () => {
const tombstones = tombstonesFor([unreferenced])
const result = pruneTombstones(tombstones, { countForIds: () => new Map() })
expect(result.prunedIds).toEqual([])
expect(result.retained).toEqual(tombstones)
})
it('sweeps every owner once for the whole batch instead of once per tombstone', () => {
const tombstones = tombstonesFor(Array.from({ length: 200 }, (_, index) => sequentialId(index)))
const batchSweeps = { count: 0 }
pruneTombstones(tombstones, batchCounter(owners, batchSweeps))
expect(batchSweeps.count).toBe(owners.length)
const perIdSweeps = { count: 0 }
pruneTombstones(tombstones, perIdCounter(owners, perIdSweeps))
expect(perIdSweeps.count).toBe(owners.length * tombstones.length)
})
it('stays linear on a 3,500-tombstone catalog', () => {
const ids = Array.from({ length: 3500 }, (_, index) => sequentialId(index))
const bigOwner: OwnerFixture[] = [{ referencedIds: ids.slice(0, 1750), readable: true }]
const started = performance.now()
const result = pruneTombstones(tombstonesFor(ids), batchCounter(bigOwner, { count: 0 }))
expect(performance.now() - started).toBeLessThan(150)
expect(result.prunedIds).toHaveLength(1750)
expect(result.retained).toHaveLength(1750)
})
})
@@ -3,6 +3,7 @@
// that blocks every v1 write while the pinned pre-v1 backup is failing.
import type { AgentCatalogMutationRequest } from '../../shared/agent-catalog-snapshot'
import type { AgentCatalogSchemaTooNew } from '../../shared/data-recovery'
/** Mutations that reduce risk/size and stay allowed while a payload budget is
* already exceeded; they must never add arbitrary user text or a reference. */
@@ -44,3 +45,28 @@ export function agentCatalogMigrationBlockedError(store: {
}
return { ok: false, code: 'agent_catalog_migration_blocked', migrationError }
}
/** Returned while the persisted catalog schema is newer than this build: authoring
* would clobber fields this build cannot represent, and nothing here is retryable
* — only a newer Orca clears it. */
export type AgentCatalogSchemaTooNewError = {
ok: false
code: 'agent_catalog_schema_too_new'
persistedVersion: number
supportedVersion: number
}
export function agentCatalogSchemaTooNewError(store: {
getAgentCatalogSchemaTooNew(): AgentCatalogSchemaTooNew | null
}): AgentCatalogSchemaTooNewError | null {
const schemaTooNew = store.getAgentCatalogSchemaTooNew()
if (schemaTooNew === null) {
return null
}
return {
ok: false,
code: 'agent_catalog_schema_too_new',
persistedVersion: schemaTooNew.persistedVersion,
supportedVersion: schemaTooNew.supportedVersion
}
}
@@ -220,14 +220,23 @@ describe('describeSpawnExecutionHost', () => {
})
// A WSL UNC cwd runs a Linux userland; win32 here picks the Windows executable
// variant and PowerShell-quotes a bash command line.
it('describes a WSL UNC cwd as a linux local target with no Windows shell', () => {
// variant and PowerShell-quotes a bash command line. Classifying it 'local'
// would also host-id it `local`, leaving UNC/drive paths in the Linux argv.
it('describes a WSL UNC cwd as its own wsl host with no Windows shell', () => {
const descriptor = describeSpawnExecutionHost({
connectionId: null,
cwd: '\\\\wsl.localhost\\Ubuntu\\home\\me\\repo',
terminalWindowsShell: 'powershell'
})
expect(descriptor).toEqual({ kind: 'local', platform: 'linux' })
expect(descriptor).toEqual({ kind: 'wsl', distro: 'Ubuntu' })
expect(platformForDescriptor(descriptor)).toBe('linux')
expect(executionHostIdForDescriptor(descriptor)).toBe('wsl:Ubuntu')
})
it('describes a legacy \\\\wsl$ UNC cwd as the same wsl host', () => {
expect(
describeSpawnExecutionHost({ connectionId: null, cwd: '\\\\wsl$\\Debian\\srv\\app' })
).toEqual({ kind: 'wsl', distro: 'Debian' })
})
})
@@ -7,6 +7,7 @@ import type { AgentLaunchSnapshot } from '../../shared/agent-launch-host-contrac
import {
AgentLaunchOperationStore,
MAX_SETTLED_OPERATIONS_PER_SCOPE,
MAX_SETTLED_OPERATION_SCOPES,
agentLaunchIdempotencyKey,
canonicalPayloadDigest,
mintAgentLaunchOperationId,
@@ -225,3 +226,57 @@ describe('settled ledger', () => {
expect(bucket.at(-1)?.operationId).toBe(`op-${MAX_SETTLED_OPERATIONS_PER_SCOPE + 2}`)
})
})
describe('settled scope bound', () => {
it('evicts the least-recently-settled scopes past the global bound', () => {
const store = new AgentLaunchOperationStore()
const overflow = 5
for (let index = 0; index < MAX_SETTLED_OPERATION_SCOPES + overflow; index += 1) {
store.recordSettled(
settled({ scope: `wt-${index}`, operationId: `op-${index}`, idempotencyKey: `k-${index}` })
)
}
for (let index = 0; index < overflow; index += 1) {
expect(store.settledForScope(`wt-${index}`)).toHaveLength(0)
}
expect(store.settledForScope(`wt-${overflow}`)).toHaveLength(1)
expect(store.settledForScope(`wt-${MAX_SETTLED_OPERATION_SCOPES + overflow - 1}`)).toHaveLength(
1
)
expect(store.durableState().settled).toHaveLength(MAX_SETTLED_OPERATION_SCOPES)
})
it('re-settling an old scope keeps it alive as newest', () => {
const store = new AgentLaunchOperationStore()
store.recordSettled(
settled({ scope: 'wt-old', operationId: 'op-old', idempotencyKey: 'k-old' })
)
for (let index = 0; index < MAX_SETTLED_OPERATION_SCOPES; index += 1) {
store.recordSettled(
settled({ scope: `wt-${index}`, operationId: `op-${index}`, idempotencyKey: `k-${index}` })
)
// Touching wt-old on every append keeps it out of the eviction window.
store.recordSettled(
settled({ scope: 'wt-old', operationId: `op-old-${index}`, idempotencyKey: `k-o-${index}` })
)
}
expect(store.settledForScope('wt-old').length).toBeGreaterThan(0)
})
it('never evicts a scope an in-flight launch still needs', () => {
const store = new AgentLaunchOperationStore()
// The pending launch lands FIRST, so its scope is the oldest by insertion —
// exactly the entry a naive oldest-first eviction would drop.
store.recordSettled(
settled({ scope: 'wt-live', operationId: 'op-live', idempotencyKey: 'k-live' })
)
store.beginPending(pending({ scope: 'wt-live', launchToken: 'token-live' }))
for (let index = 0; index < MAX_SETTLED_OPERATION_SCOPES + 10; index += 1) {
store.recordSettled(
settled({ scope: `wt-${index}`, operationId: `op-${index}`, idempotencyKey: `k-${index}` })
)
}
expect(store.settledForScope('wt-live')).toHaveLength(1)
expect(store.findSettledByIdempotencyKey('wt-live', 'k-live')?.operationId).toBe('op-live')
})
})
@@ -80,6 +80,10 @@ function buildDeps(
operationStore: store,
liveTerminalByToken: overrides.liveTerminalByToken ?? (() => null),
isHostAuthoritative: overrides.isHostAuthoritative ?? ((id) => id === 'local'),
isHostTokenAuthoritative: overrides.isHostTokenAuthoritative ?? (() => true),
...(overrides.identifyLaunchWithoutTokenEcho
? { identifyLaunchWithoutTokenEcho: overrides.identifyLaunchWithoutTokenEcho }
: {}),
expectedWorktreeId: overrides.expectedWorktreeId ?? ((p) => p.scope),
arms: overrides.arms ?? {
worktree: noopArm,
@@ -244,6 +248,114 @@ describe('buildReconcileAgentLaunchDeps liveness', () => {
expect(arm.calls).toEqual(['failed'])
})
it('settles absent on a missing token echo when the host echoes launch tokens', () => {
const arm = spyArm()
const { store, deps } = buildDeps({
isHostAuthoritative: () => true,
isHostTokenAuthoritative: () => true,
identifyLaunchWithoutTokenEcho: () => 'inconclusive',
arms: {
worktree: () => arm,
automation: () => arm,
orchestration: () => arm,
background: () => arm
}
})
const entry = pending({ launchToken: 'token-v34' }, 'ssh:host-a')
store.beginPending(entry)
const outcome = reconcileOnePendingAgentLaunch(deps, entry)
expect(outcome).toEqual({ kind: 'spawn_failed' })
expect(arm.calls).toEqual(['failed'])
})
it('never settles absent on a missing echo from a host that cannot echo tokens', () => {
// Regression (P1-4): a pre-v34 daemon / old relay accepts launchToken and
// drops it, so its listing can never echo one. Settling spawn_failed here
// would let Retry spawn a DUPLICATE beside the agent it is still running.
const arm = spyArm()
const { store, deps } = buildDeps({
isHostAuthoritative: () => true,
isHostTokenAuthoritative: () => false,
identifyLaunchWithoutTokenEcho: () => 'inconclusive',
arms: {
worktree: () => arm,
automation: () => arm,
orchestration: () => arm,
background: () => arm
}
})
const entry = pending({ launchToken: 'token-legacy' }, 'ssh:host-a')
store.beginPending(entry)
const outcome = reconcileOnePendingAgentLaunch(deps, entry)
expect(outcome).toEqual({ kind: 'launch_state_unknown' })
expect(arm.calls).toEqual(['unknown'])
// Held, not settled: the reservation survives and Retry stays gated.
expect(store.getPending('token-legacy')).not.toBeNull()
})
it('holds a non-token-authoritative launch pending when no fallback is wired', () => {
const arm = spyArm()
const { store, deps } = buildDeps({
isHostAuthoritative: () => true,
isHostTokenAuthoritative: () => false,
arms: {
worktree: () => arm,
automation: () => arm,
orchestration: () => arm,
background: () => arm
}
})
const entry = pending({ launchToken: 'token-nofallback' }, 'ssh:host-a')
store.beginPending(entry)
expect(reconcileOnePendingAgentLaunch(deps, entry)).toEqual({ kind: 'launch_state_unknown' })
expect(store.getPending('token-nofallback')).not.toBeNull()
})
it('settles absent on a non-echoing host once the fallback proves the terminal gone', () => {
const arm = spyArm()
const { store, deps } = buildDeps({
isHostAuthoritative: () => true,
isHostTokenAuthoritative: () => false,
identifyLaunchWithoutTokenEcho: () => 'absent',
arms: {
worktree: () => arm,
automation: () => arm,
orchestration: () => arm,
background: () => arm
}
})
const entry = pending({ launchToken: 'token-proven' }, 'ssh:host-a')
store.beginPending(entry)
expect(reconcileOnePendingAgentLaunch(deps, entry)).toEqual({ kind: 'spawn_failed' })
expect(arm.calls).toEqual(['failed'])
})
it('keeps a live token match launched even on a non-token-authoritative host', () => {
const arm = spyArm()
const { store, deps } = buildDeps({
liveTerminalByToken: () => ({ ptyId: 'term-live', worktreeId: 'wt-1' }),
isHostAuthoritative: () => true,
isHostTokenAuthoritative: () => false,
identifyLaunchWithoutTokenEcho: () => 'absent',
arms: {
worktree: () => arm,
automation: () => arm,
orchestration: () => arm,
background: () => arm
}
})
const entry = pending({ launchToken: 'token-live' }, 'ssh:host-a')
store.beginPending(entry)
expect(reconcileOnePendingAgentLaunch(deps, entry)).toEqual({ kind: 'launched' })
})
it('routes a background pending to the background store keyed by attempt id', () => {
const background = new BackgroundAgentLaunchStore({ now: () => 1000 })
background.create({
@@ -15,6 +15,13 @@
// until its own terminal-list/reconnect event re-probes. `isHostAuthoritative`
// encodes which hosts a given reconcile pass can speak for, so a daemon/SSH
// survivor is never falsely settled `absent` before its provider reconnects.
// - Listing authority is ANDed with TOKEN authority: a peer that predates the
// launch-token echo (pre-v34 daemon, old SSH relay) accepts the token on
// create and silently drops it, so its listing can NEVER carry one. Reading
// that missing echo as absence settles spawn_failed for a live agent and the
// user's Retry then spawns a duplicate beside it. Such hosts fall back to the
// pre-launchToken identification, and hold the launch pending when that is
// inconclusive.
import type { AgentLaunchExecutionHostId } from '../../shared/agent-launch-host-contract'
import {
@@ -41,6 +48,11 @@ import type {
* the launch's own terminal is alive, so it must never resolve `absent`. */
export type LiveTerminalForToken = { ptyId: string; worktreeId: string | null }
/** Verdict of the pre-launchToken identification used for hosts that cannot echo
* tokens: `absent` only when the re-listed terminals PROVE the launch's terminal
* is gone. Anything less is `inconclusive` and holds the launch pending. */
export type TokenlessLaunchLiveness = 'absent' | 'inconclusive'
export type ReconcileRuntimeDeps = {
operationStore: AgentLaunchOperationStore
/** The live terminal holding a launch token, or null if none is live. */
@@ -48,6 +60,12 @@ export type ReconcileRuntimeDeps = {
/** Whether a non-live pending's host can be spoken for authoritatively in this
* reconcile pass (→ `absent`); false leaves it `unknown`. */
isHostAuthoritative: (executionHostId: AgentLaunchExecutionHostId) => boolean
/** Whether a MISSING launchToken in this host's listing is absence proof. False
* for peers that predate the echo and drop the token they were handed. */
isHostTokenAuthoritative: (executionHostId: AgentLaunchExecutionHostId) => boolean
/** Pre-launchToken identification, consulted only for non-token-authoritative
* hosts. Omitted (or `inconclusive`) holds the launch pending. */
identifyLaunchWithoutTokenEcho?: (pending: PendingAgentLaunchSnapshot) => TokenlessLaunchLiveness
/** The worktree a live token must belong to for attribution: the scope for a
* worktree launch, the attempt's worktree for a background launch, or null when
* the intent has no worktree to compare (attribution then trusts the token). */
@@ -105,7 +123,17 @@ function resolveLiveness(
}
}
const host = pending.snapshot.target.executionHostId
return deps.isHostAuthoritative(host) ? { kind: 'absent' } : { kind: 'unknown' }
if (!deps.isHostAuthoritative(host)) {
return { kind: 'unknown' }
}
if (!deps.isHostTokenAuthoritative(host)) {
// A peer that never echoes tokens can only be identified the pre-token way;
// absent needs positive proof there, or Retry duplicates a running agent.
return deps.identifyLaunchWithoutTokenEcho?.(pending) === 'absent'
? { kind: 'absent' }
: { kind: 'unknown' }
}
return { kind: 'absent' }
}
export function buildReconcileAgentLaunchDeps(
@@ -182,6 +182,94 @@ describe('resolveAgentLaunchSpawn', () => {
})
})
// P1-24: a "Don't save" launch sends the edited args with the locator; the
// stored recipe snapshot is stale and must not win.
it('prefers client unsaved args over the recipe stored agentArgs (P1-24)', async () => {
const resolve = vi.fn((_request: ResolveAgentLaunchRequest) => ({
ok: true as const,
launch: makeLaunch()
}))
const deps = makeDeps(resolve)
await resolveAgentLaunchSpawn(
deps,
baseInput({
request: {
selection: { kind: 'agent', agent: 'claude' },
prompt: 'x',
sourceRecord: { owner: 'source-control-recipe', id: 'fixChecks' },
unsavedAgentArgs: '--edited two'
},
recipeRepo: {
sourceControlAi: { actionOverrides: { fixChecks: { agentArgs: '--recipe one' } } }
}
})
)
expect(resolve.mock.calls[0]![0].perLaunchArgs).toBe('--edited two')
})
it('treats empty unsaved args as a real "no args" edit (P1-24)', async () => {
const resolve = vi.fn((_request: ResolveAgentLaunchRequest) => ({
ok: true as const,
launch: makeLaunch()
}))
const deps = makeDeps(resolve)
await resolveAgentLaunchSpawn(
deps,
baseInput({
request: {
selection: { kind: 'agent', agent: 'claude' },
prompt: 'x',
sourceRecord: { owner: 'source-control-recipe', id: 'fixChecks' },
unsavedAgentArgs: ''
},
recipeRepo: {
sourceControlAi: { actionOverrides: { fixChecks: { agentArgs: '--recipe one' } } }
}
})
)
expect(resolve.mock.calls[0]![0].perLaunchArgs).toBe('')
})
it('ignores unsaved args without a source-control-recipe owner (P1-24)', async () => {
const resolve = vi.fn((_request: ResolveAgentLaunchRequest) => ({
ok: true as const,
launch: makeLaunch()
}))
const deps = makeDeps(resolve)
await resolveAgentLaunchSpawn(
deps,
baseInput({
request: {
selection: { kind: 'agent', agent: 'claude' },
prompt: 'x',
unsavedAgentArgs: '--sneaky'
}
})
)
expect('perLaunchArgs' in resolve.mock.calls[0]![0]).toBe(false)
})
it('still rejects an unknown recipe id even with unsaved args (P1-24)', async () => {
const resolve = vi.fn((_request: ResolveAgentLaunchRequest) => ({
ok: true as const,
launch: makeLaunch()
}))
const deps = makeDeps(resolve)
const result = await resolveAgentLaunchSpawn(
deps,
baseInput({
request: {
selection: { kind: 'agent', agent: 'claude' },
prompt: 'x',
sourceRecord: { owner: 'source-control-recipe', id: 'not-a-real-action' },
unsavedAgentArgs: '--edited'
}
})
)
expect(result).toEqual({ ok: false, requestError: { code: 'untrusted_reference' } })
expect(resolve).not.toHaveBeenCalled()
})
it('rejects an unknown recipe action id with untrusted_reference and never resolves (U7)', async () => {
const resolve = vi.fn((_request: ResolveAgentLaunchRequest) => ({
ok: true as const,
+9 -1
View File
@@ -130,7 +130,10 @@ function referenceFor(request: AgentLaunchSpawnRequest): AgentReferenceAuthority
* a source-control-recipe owner contributes args today: the host validates the id
* is a real action id (unknown/mismatched → untrusted_reference, no PTY), then
* reads the recipe's stored agentArgs from repo-scoped settings (global fallback
* when the repo id is absent). Clients never send args — only the recipe id. */
* when the repo id is absent). A client may substitute UNSAVED edits of those
* args (`unsavedAgentArgs`) for this one launch — still bounded text the resolver
* tokenizes itself, never a resolved argv — and the stored recipe stays the
* fallback. Clients never send anything else about the launch. */
function resolvePerLaunchArgs(
request: AgentLaunchSpawnRequest,
recipeRepo: Pick<Repo, 'sourceControlAi'> | null | undefined,
@@ -143,6 +146,11 @@ function resolvePerLaunchArgs(
if (!sourceRecord.id || !isSourceControlActionId(sourceRecord.id)) {
return { ok: false, requestError: { code: 'untrusted_reference' } }
}
// Why: the unsaved edit IS the args the user launched with; the stored recipe
// is a stale snapshot of them. An empty string is a real "no args" edit.
if (request.unsavedAgentArgs !== undefined) {
return { ok: true, perLaunchArgs: request.unsavedAgentArgs }
}
const recipe = resolveSourceControlActionRecipe({
settings,
repo: recipeRepo,
@@ -231,6 +231,44 @@ describe('two-stage worktree agent-launch resolution', () => {
expect(store.pendingForWorktree('attempt-1')).toBe(0)
})
it('frees the hold when a stage-2 host read rejects, so repeated races keep capacity', async () => {
const resolve = vi
.fn<(request: ResolveAgentLaunchRequest) => ResolveAgentLaunchOutcome>()
.mockReturnValue({ ok: true, launch: makeLaunch('fp-real', 'sd-1', '/wt-real') })
const { deps, store } = makeSetup(resolve)
// More rounds than MAX_PENDING_LAUNCHES_PER_PRINCIPAL: a leaked hold per round
// exhausts the principal's capacity until restart.
for (let round = 0; round < 70; round++) {
deps.resolveTargetHomePath = async () => '/home/dev'
const prepared = await prepareWorktreeAgentLaunch(deps, CONTEXT, {
repoPath: '/repo',
worktreePath: '/wt-provisional'
})
expect(prepared.ok).toBe(true)
if (!prepared.ok) {
return
}
// The SSH connection drops between git creating the worktree and stage 2.
deps.resolveTargetHomePath = async () => {
throw new Error('ssh channel closed')
}
const executed = await executeWorktreeAgentLaunch(
deps,
CONTEXT,
{ repoPath: '/repo', worktreePath: '/wt-real' },
{ reservationId: prepared.reservationId, expectedStableInputDigest: 'sd-1' }
)
expect(executed.ok).toBe(false)
if (executed.ok) {
return
}
expect('failure' in executed && executed.failure.code).toBe('spawn_failed')
expect(store.pendingForPrincipal(LOCAL)).toBe(0)
expect(store.pendingCount()).toBe(0)
}
})
it('takes no reservation when pre-git resolution fails', async () => {
const resolve = vi
.fn<(request: ResolveAgentLaunchRequest) => ResolveAgentLaunchOutcome>()
@@ -136,7 +136,7 @@ export async function prepareWorktreeAgentLaunch(
/** Stage 2: with the authoritative worktree path and the pinned reservation,
* re-resolve, recheck the config-only digest, and convert the reservation into
* a startup plan + receipt (or release it on any failure). Creates no PTY: the
* a startup plan + receipt (or release it on any rejection). Creates no PTY: the
* caller persists the pending record, then spawns and settles. */
export async function executeWorktreeAgentLaunch(
deps: WorktreeAgentLaunchDeps,
@@ -144,50 +144,60 @@ export async function executeWorktreeAgentLaunch(
authoritativePaths: { repoPath: string | null; worktreePath: string | null },
reservation: { reservationId: string; expectedStableInputDigest: string }
): Promise<ExecuteAgentLaunchResult> {
const hostState = await deriveAgentLaunchHostState(
{
getSettings: deps.getSettings,
getCatalogRevision: deps.getCatalogRevision,
detectStockBaseAgents: deps.detectStockBaseAgents,
resolveTargetHomePath: deps.resolveTargetHomePath,
...(deps.resolveTransportConfidentiality
? { resolveTransportConfidentiality: deps.resolveTransportConfidentiality }
: {})
},
context.descriptor,
authoritativePaths
)
const resolve = buildHostStateResolve(toSpawnDeps(deps), {
request: context.request,
intent: context.intent,
target: hostState.target,
variables: hostState.variables,
scope: context.scope,
principal: context.principal
})
return deps.boundary.executeReservedAgentLaunch({
scope: context.scope,
// Stage 2 holds the authoritative worktree; a background intent names it
// even when the caller did not thread context.worktreeId explicitly.
worktreeId:
context.worktreeId !== undefined
? context.worktreeId
: context.intent.kind === 'background'
? context.intent.worktreeId
: null,
principal: context.principal,
resolve,
prompt: context.request.prompt ?? '',
...(context.request.allowEmptyPromptLaunch !== undefined
? { allowEmptyPromptLaunch: context.request.allowEmptyPromptLaunch }
: {}),
...(context.request.promptDelivery !== undefined
? { promptDelivery: context.request.promptDelivery }
: {}),
maxInlineDraftChars: STARTUP_COMMAND_TEXT_MAX_CHARS,
...(deps.markWorkspaceTrusted ? { preflight: deps.markWorkspaceTrusted } : {}),
...(deps.prepareEnv ? { prepareEnv: deps.prepareEnv } : {}),
reservationId: reservation.reservationId,
expectedStableInputDigest: reservation.expectedStableInputDigest
})
try {
const hostState = await deriveAgentLaunchHostState(
{
getSettings: deps.getSettings,
getCatalogRevision: deps.getCatalogRevision,
detectStockBaseAgents: deps.detectStockBaseAgents,
resolveTargetHomePath: deps.resolveTargetHomePath,
...(deps.resolveTransportConfidentiality
? { resolveTransportConfidentiality: deps.resolveTransportConfidentiality }
: {})
},
context.descriptor,
authoritativePaths
)
const resolve = buildHostStateResolve(toSpawnDeps(deps), {
request: context.request,
intent: context.intent,
target: hostState.target,
variables: hostState.variables,
scope: context.scope,
principal: context.principal
})
return await deps.boundary.executeReservedAgentLaunch({
scope: context.scope,
// Stage 2 holds the authoritative worktree; a background intent names it
// even when the caller did not thread context.worktreeId explicitly.
worktreeId:
context.worktreeId !== undefined
? context.worktreeId
: context.intent.kind === 'background'
? context.intent.worktreeId
: null,
principal: context.principal,
resolve,
prompt: context.request.prompt ?? '',
...(context.request.allowEmptyPromptLaunch !== undefined
? { allowEmptyPromptLaunch: context.request.allowEmptyPromptLaunch }
: {}),
...(context.request.promptDelivery !== undefined
? { promptDelivery: context.request.promptDelivery }
: {}),
maxInlineDraftChars: STARTUP_COMMAND_TEXT_MAX_CHARS,
...(deps.markWorkspaceTrusted ? { preflight: deps.markWorkspaceTrusted } : {}),
...(deps.prepareEnv ? { prepareEnv: deps.prepareEnv } : {}),
reservationId: reservation.reservationId,
expectedStableInputDigest: reservation.expectedStableInputDigest
})
} catch {
// The stage-2 host reads (SSH/WSL detection + home probe) happen before the
// boundary's guarded release path, so a disconnect rejecting here would
// strand the pre-git hold until restart and repeated races would exhaust
// capacity. Releasing is idempotent, so a hold the boundary already freed is
// unaffected.
deps.boundary.releaseReservedAgentLaunch(reservation.reservationId)
return { ok: false, failure: { code: 'spawn_failed' } }
}
}
@@ -0,0 +1,182 @@
// Reference-side authoring authority (owner references over agent tombstones):
// snapshots, the reference mutation path, and the reference-driven tombstone GC.
// Split out of AgentCatalogService, which composes it and delegates.
import type { Store } from '../persistence'
import { buildUnreferencedTombstonePrunePatch } from './agent-catalog-tombstone-gc'
import { normalizeCatalogFromSettings } from './agent-catalog-projections'
import {
buildAgentReferenceSnapshot,
measureAgentReferenceProjection
} from './agent-reference-snapshot-projection'
import {
agentCatalogMigrationBlockedError,
agentCatalogSchemaTooNewError,
type AgentCatalogMigrationBlockedError,
type AgentCatalogSchemaTooNewError
} from './agent-catalog-write-policy'
import {
commitAuthoringPatchDurable,
type AgentReferenceWriteFailedError
} from './agent-catalog-durable-write'
import { applyAgentReferenceMutation } from './agent-reference-mutations'
import type {
AgentReferenceMutationRequest,
AgentReferenceMutationResult,
AgentReferenceProjectionError,
AgentReferenceSnapshot,
LocalAgentReferenceSnapshot
} from '../../shared/agent-reference-snapshot'
const MAX_REFERENCE_PROJECTION_BYTES = 524_288
export type AgentReferenceAuthoringDeps = {
store: Store
countTombstoneReferences: Parameters<typeof buildUnreferencedTombstonePrunePatch>[2]
catalogRevision: () => number
publishCatalogRevision: (revision: number) => void
}
export class AgentReferenceAuthoringService {
constructor(private readonly deps: AgentReferenceAuthoringDeps) {}
private get store(): Store {
return this.deps.store
}
getRevision(): number {
return this.store.getSettings().agentReferenceRevision ?? 1
}
/** Remote (runtime RPC) reference snapshot; typed projection error when over
* the 512 KiB frame budget. */
getRemoteSnapshot(): AgentReferenceSnapshot | AgentReferenceProjectionError {
const settings = this.store.getSettings()
const snapshot = buildAgentReferenceSnapshot(settings)
const { tooLarge } = measureAgentReferenceProjection(settings)
if (tooLarge) {
return {
version: 1,
revision: snapshot.revision,
code: 'agent_reference_payload_too_large',
maxBytes: MAX_REFERENCE_PROJECTION_BYTES
}
}
return snapshot
}
/** Uncapped authoring/repair view over local preload IPC only. */
getLocalSnapshot(): LocalAgentReferenceSnapshot {
const settings = this.store.getSettings()
const snapshot = buildAgentReferenceSnapshot(settings)
const { bytes, tooLarge } = measureAgentReferenceProjection(settings)
return {
...snapshot,
projection: tooLarge
? { status: 'too-large', bytes, maxBytes: MAX_REFERENCE_PROJECTION_BYTES }
: { status: 'ready', bytes, maxBytes: MAX_REFERENCE_PROJECTION_BYTES }
}
}
mutate(
request: AgentReferenceMutationRequest
):
| AgentReferenceMutationResult<LocalAgentReferenceSnapshot>
| AgentCatalogMigrationBlockedError
| AgentCatalogSchemaTooNewError
| AgentReferenceWriteFailedError {
// The failed-backup invariant is "no v1 write": reference mutations stamp
// agentReferenceRevision, so they are blocked alongside catalog mutations.
const blocked = agentCatalogMigrationBlockedError(this.store)
if (blocked) {
return blocked
}
// Persistence refuses the write anyway; reporting it up front keeps a
// non-retryable read-only profile from looking like a disk failure.
const tooNew = agentCatalogSchemaTooNewError(this.store)
if (tooNew) {
return tooNew
}
const settings = this.store.getSettings()
const currentReferenceRevision = settings.agentReferenceRevision ?? 1
const application = applyAgentReferenceMutation({
settings,
request,
currentReferenceRevision,
catalog: normalizeCatalogFromSettings(settings)
})
if (!application.ok) {
return {
ok: false,
code: application.code,
referenceRevision: currentReferenceRevision,
catalogRevision: this.deps.catalogRevision(),
...(application.code === 'reference_revision_conflict'
? { snapshot: this.getLocalSnapshot() }
: {}),
...(application.owner ? { owner: application.owner } : {}),
...(application.field ? { field: application.field } : {}),
...(application.reason ? { reason: application.reason } : {})
}
}
// The 512 KiB remote-projection budget is checked on the post-mutation
// snapshot; a non-growing write still commits so an over-budget profile can
// always be edited back under budget (mirrors mutate()'s security-reducing
// allowlist). Why preview: persistence normalizes the patch and derives
// legacy commitMessageAi from sourceControlAi, so the patch as written is
// smaller than the projection that actually ships.
const projected = measureAgentReferenceProjection(
this.store.previewSettingsUpdate(application.patch)
)
if (projected.tooLarge && projected.bytes > measureAgentReferenceProjection(settings).bytes) {
return {
ok: false,
code: 'agent_reference_payload_too_large',
referenceRevision: currentReferenceRevision,
catalogRevision: this.deps.catalogRevision()
}
}
// Owner change commits before any prune; a failure between the two leaves
// the tombstone conservatively retained for the next indexed recheck.
if (!commitAuthoringPatchDurable(this.store, application.patch)) {
return {
ok: false,
code: 'agent_reference_write_failed',
referenceRevision: currentReferenceRevision,
catalogRevision: this.deps.catalogRevision()
}
}
this.pruneUnreferencedTombstones()
return {
ok: true,
referenceRevision: application.newReferenceRevision,
catalogRevision: this.deps.catalogRevision(),
snapshot: this.getLocalSnapshot()
}
}
/** Reference-aware prune run after a reference removal; a prune advances and
* publishes the catalog revision so receivers replace their snapshot. */
private pruneUnreferencedTombstones(): void {
const settings = this.store.getSettings()
// The patch also strips any persisted row the pruned tombstone suppressed,
// so the prune can never resurrect it.
const prunePatch = buildUnreferencedTombstonePrunePatch(
settings,
normalizeCatalogFromSettings(settings),
this.deps.countTombstoneReferences
)
if (!prunePatch) {
return
}
const newRevision = (settings.agentCatalogRevision ?? 1) + 1
// GC only: a rolled-back prune keeps the tombstone retained, and the next
// indexed recheck retries it — the reference mutation is already durable.
if (
!commitAuthoringPatchDurable(this.store, { ...prunePatch, agentCatalogRevision: newRevision })
) {
return
}
this.deps.publishCatalogRevision(newRevision)
}
}
@@ -2,7 +2,7 @@
// staging, provider-session bind (by launch token) → durable resume record,
// ownership-key resolution, incompatible/non-resumable bind rejection, spawn-
// failure rollback, dispose-keeps-record, and the one-time legacy handoff.
import { describe, expect, it } from 'vitest'
import { describe, expect, it, vi } from 'vitest'
import type { AgentLaunchSnapshot } from '../../shared/agent-launch-host-contract'
import type { TuiAgent } from '../../shared/types'
import {
@@ -16,6 +16,7 @@ import {
type AgentSessionRecordStoreDurableState,
type StagedLaunchRegistration
} from './agent-session-record-store'
import { MAX_SESSION_RECORDS } from './agent-session-record-retention'
function snapshot(overrides: Partial<AgentLaunchSnapshot> = {}): AgentLaunchSnapshot {
return {
@@ -191,6 +192,51 @@ describe('AgentSessionRecordStore lifecycle', () => {
expect(store.bindProviderSessionByToken('token-a', SESSION)).toBeNull()
})
it('dispose by terminal clears staging a pane-less surface registered', () => {
const store = new AgentSessionRecordStore()
// Runtime/mobile launches carry a terminal id but never a stable pane key,
// so the pane teardown can never reach them.
store.register(registration({ paneKey: undefined, terminalId: 'term-a' }))
store.disposeStagingForPane('pane-a')
store.disposeStagingForTerminal('term-a')
expect(store.bindProviderSessionByToken('token-a', SESSION)).toBeNull()
})
it('bounds retained durable records, least-recently-updated first', () => {
let clock = 0
const store = new AgentSessionRecordStore({ now: () => (clock += 1) })
const total = MAX_SESSION_RECORDS + 4
for (let index = 0; index < total; index += 1) {
store.register(registration({ launchToken: `token-${index}`, paneKey: `pane-${index}` }))
store.bindProviderSessionByToken(`token-${index}`, { key: 'session_id', id: `sess-${index}` })
}
expect(store.durableState().records).toHaveLength(MAX_SESSION_RECORDS)
for (let index = 0; index < 4; index += 1) {
expect(
store.resolveByOwnershipKey({ ...OWNERSHIP, providerSessionId: `sess-${index}` })
).toBeNull()
}
// The record the last bind just created is never the eviction candidate.
expect(
store.resolveByOwnershipKey({ ...OWNERSHIP, providerSessionId: `sess-${total - 1}` })
).not.toBeNull()
})
it('trims a pre-bound record file at rehydrate without a write', () => {
let clock = 0
const seed = new AgentSessionRecordStore({ now: () => (clock += 1) })
for (let index = 0; index < MAX_SESSION_RECORDS + 3; index += 1) {
seed.register(registration({ launchToken: `token-${index}` }))
seed.bindProviderSessionByToken(`token-${index}`, { key: 'session_id', id: `sess-${index}` })
}
const store = new AgentSessionRecordStore()
const sink = vi.fn()
store.setDurablePersistence(sink)
store.rebuildRecordsFrom(seed.durableState().records)
expect(store.durableState().records).toHaveLength(MAX_SESSION_RECORDS)
expect(sink).not.toHaveBeenCalled()
})
it('two custom ids on one base/provider session resolve to one owner record', () => {
const store = new AgentSessionRecordStore()
store.register(
@@ -1,5 +1,8 @@
import { describe, expect, it, vi } from 'vitest'
import { BackgroundAgentLaunchStore } from './background-agent-launch-store'
import {
BackgroundAgentLaunchStore,
MAX_SETTLED_BACKGROUND_ATTEMPTS
} from './background-agent-launch-store'
import type { BackgroundAgentLaunchCreateInput } from './background-agent-launch-store'
import type { PersistedAgentLaunchFailure } from '../../shared/agent-launch-contract'
import { parsePersistedAgentLaunchFailure } from '../../shared/agent-launch-failure-schema'
@@ -203,6 +206,54 @@ describe('BackgroundAgentLaunchStore', () => {
)
})
it('bounds retained settled attempts, oldest-settled first', () => {
let clock = 0
const store = new BackgroundAgentLaunchStore({ now: () => (clock += 1) })
const total = MAX_SETTLED_BACKGROUND_ATTEMPTS + 5
for (let index = 0; index < total; index += 1) {
store.create(createInput({ attemptId: `attempt-${index}` }))
store.settleLaunched(`attempt-${index}`)
}
expect(store.all()).toHaveLength(MAX_SETTLED_BACKGROUND_ATTEMPTS)
for (let index = 0; index < 5; index += 1) {
expect(store.get(`attempt-${index}`)).toBeNull()
}
expect(store.get(`attempt-${total - 1}`)?.state).toBe('launched')
})
it('never evicts a pending attempt, however old', () => {
let clock = 0
const store = new BackgroundAgentLaunchStore({ now: () => (clock += 1) })
// Created first, so it is the oldest record in the ledger and would be the
// first casualty of an age-only bound — but its reservation and private
// snapshot are still live.
store.create(createInput({ attemptId: 'attempt-live' }))
for (let index = 0; index < MAX_SETTLED_BACKGROUND_ATTEMPTS + 20; index += 1) {
store.create(createInput({ attemptId: `attempt-${index}` }))
store.settleFailed(`attempt-${index}`, failure('spawn_failed'))
}
expect(store.get('attempt-live')?.state).toBe('pending')
expect(store.all()).toHaveLength(MAX_SETTLED_BACKGROUND_ATTEMPTS + 1)
})
it('trims a pre-bound ledger at rehydrate without a write', () => {
let clock = 0
const seed = new BackgroundAgentLaunchStore({ now: () => (clock += 1) })
for (let index = 0; index < MAX_SETTLED_BACKGROUND_ATTEMPTS + 7; index += 1) {
seed.create(createInput({ attemptId: `attempt-${index}` }))
seed.settleLaunched(`attempt-${index}`)
}
const persisted = seed.durableState().attempts
const overSized = [...persisted, { ...persisted[0], attemptId: 'legacy-0', updatedAt: 0 }]
const store = new BackgroundAgentLaunchStore()
const sink = vi.fn()
store.setDurablePersistence(sink)
store.rebuildFrom(overSized)
expect(store.all()).toHaveLength(MAX_SETTLED_BACKGROUND_ATTEMPTS)
expect(store.get('legacy-0')).toBeNull()
expect(sink).not.toHaveBeenCalled()
})
it('persistenceForAttempt binds the reconcile slice to one attempt', () => {
const store = new BackgroundAgentLaunchStore()
store.create(createInput())
@@ -130,6 +130,19 @@ export function composeAgentLaunchEnv(input: ComposeAgentLaunchEnvInput): Record
return env
}
/** The layer a spawned child actually inherits: `process.env` values are typed
* `string | undefined` and an undefined key is absent from the child's block. */
export function inheritedEnvLayer(source: NodeJS.ProcessEnv): EnvLayer {
const layer = nullProtoEnv()
for (const key of Object.keys(source)) {
const value = source[key]
if (value !== undefined) {
layer[key] = value
}
}
return layer
}
/** Native-Windows environment-block size in UTF-16 code units: each entry is
* `key=value\0`, with one extra terminating NUL after the final entry. */
export function measureWindowsEnvironmentBlockCodeUnits(env: EnvLayer): number {
@@ -0,0 +1,187 @@
// One root cause, three durable launch stores: they wrote through the NON-fsync
// writer and their sinks swallowed every write error, so a launch could report
// success while its idempotency/recovery/resume row never reached the platter
// (or vanished with the page cache on power loss). These tests hold both halves
// of the fix — the fsync'd writer, and a failed write that reaches the caller.
import type * as SecureFileModule from '../../shared/secure-file'
import { mkdtempSync, rmSync } from 'node:fs'
import { tmpdir } from 'node:os'
import { join } from 'node:path'
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
const durableWrites: string[] = []
const plainWrites: string[] = []
const failingWritePaths = new Set<string>()
vi.mock('../../shared/secure-file', async (importOriginal) => {
const actual = await importOriginal<typeof SecureFileModule>()
const failIfMarked = (path: string): void => {
if (failingWritePaths.has(path)) {
const error = new Error(
`ENOSPC: no space left on device, write '${path}'`
) as NodeJS.ErrnoException
error.code = 'ENOSPC'
throw error
}
}
return {
...actual,
writeSecureJsonFile: (path: string, value: unknown) => {
plainWrites.push(path)
failIfMarked(path)
actual.writeSecureJsonFile(path, value)
},
writeDurableSecureJsonFile: (path: string, value: unknown) => {
durableWrites.push(path)
failIfMarked(path)
actual.writeDurableSecureJsonFile(path, value)
}
}
})
vi.mock('electron', () => ({
safeStorage: {
isEncryptionAvailable: () => false,
encryptString: (value: string) => Buffer.from(value, 'utf-8'),
decryptString: (value: Buffer) => value.toString('utf-8')
}
}))
import type { AgentLaunchSnapshot } from '../../shared/agent-launch-host-contract'
import {
AgentLaunchOperationStore,
mintAgentLaunchOperationId
} from './agent-launch-operation-store'
import {
agentLaunchOperationStorePath,
initAgentLaunchOperationStorePersistence
} from './agent-launch-operation-store-persistence'
import { AgentSessionRecordStore } from './agent-session-record-store'
import {
agentSessionRecordStorePath,
initAgentSessionRecordStorePersistence,
type AgentSessionRecordCipher
} from './agent-session-record-store-persistence'
import { BackgroundAgentLaunchStore } from './background-agent-launch-store'
import {
backgroundAgentLaunchStorePath,
initBackgroundAgentLaunchStorePersistence
} from './background-agent-launch-store-persistence'
const plaintextCipher: AgentSessionRecordCipher = {
available: () => false,
encrypt: (plaintext) => Buffer.from(plaintext, 'utf-8'),
decrypt: (ciphertext) => ciphertext.toString('utf-8')
}
const snapshot: AgentLaunchSnapshot = {
version: 1,
requestedAgent: 'claude',
baseAgent: 'claude',
displayLabel: 'Claude',
mode: 'built-in',
argv: ['claude'],
agentEnv: {},
capturedEnvPolicy: 'none',
target: {
platform: 'linux',
execution: 'native',
shell: 'posix',
isRemote: false,
executionHostId: 'local'
}
}
function mutateOperationStore(store: AgentLaunchOperationStore, launchToken: string): void {
store.beginPending({
operationId: mintAgentLaunchOperationId(),
idempotencyKey: `key-${launchToken}`,
scope: 'wt-1',
clientMutationId: null,
payloadDigest: 'digest-a',
launchToken,
intent: 'interactive',
snapshot
})
}
function mutateSessionRecordStore(store: AgentSessionRecordStore, launchToken: string): void {
store.register({
worktreeId: 'wt-1',
requestedAgent: 'claude',
baseAgent: 'claude',
launchSnapshot: snapshot,
launchToken
})
store.bindProviderSessionByToken(launchToken, { key: 'session_id', id: `sess-${launchToken}` })
}
function mutateBackgroundStore(store: BackgroundAgentLaunchStore, attemptId: string): void {
store.create({
attemptId,
worktreeId: 'r1::/wt',
operationId: `op-${attemptId}`,
requestedAgent: 'claude',
baseAgent: null
})
}
describe('durable launch store writes', () => {
let dir: string
beforeEach(() => {
dir = mkdtempSync(join(tmpdir(), 'launch-store-durable-'))
durableWrites.length = 0
plainWrites.length = 0
failingWritePaths.clear()
})
afterEach(() => {
rmSync(dir, { recursive: true, force: true })
})
it('persists the operation store through the fsync writer and surfaces failures', () => {
const path = agentLaunchOperationStorePath(dir)
const store = new AgentLaunchOperationStore()
initAgentLaunchOperationStorePersistence(store, path, plaintextCipher, {
rebuildAdmission: () => {},
worktreeIdForBackgroundScope: () => null
})
mutateOperationStore(store, 'tok-ok')
expect(durableWrites).toContain(path)
expect(plainWrites).not.toContain(path)
failingWritePaths.add(path)
expect(() => mutateOperationStore(store, 'tok-fail')).toThrow(/ENOSPC/)
})
it('persists session records through the fsync writer and surfaces failures', () => {
const path = agentSessionRecordStorePath(dir)
const store = new AgentSessionRecordStore()
initAgentSessionRecordStorePersistence(store, path, plaintextCipher)
mutateSessionRecordStore(store, 'tok-ok')
expect(durableWrites).toContain(path)
expect(plainWrites).not.toContain(path)
failingWritePaths.add(path)
expect(() => mutateSessionRecordStore(store, 'tok-fail')).toThrow(/ENOSPC/)
})
it('persists background attempts through the fsync writer and surfaces failures', () => {
const path = backgroundAgentLaunchStorePath(dir)
const store = new BackgroundAgentLaunchStore()
initBackgroundAgentLaunchStorePersistence(store, path)
mutateBackgroundStore(store, '0199f7a1-0000-7000-8000-0000000000d1')
expect(durableWrites).toContain(path)
expect(plainWrites).not.toContain(path)
failingWritePaths.add(path)
expect(() => mutateBackgroundStore(store, '0199f7a1-0000-7000-8000-0000000000d2')).toThrow(
/ENOSPC/
)
})
})
@@ -0,0 +1,137 @@
// The admission fingerprint must move whenever anything that changes THIS launch
// changes — including the env a mobile/paired remove-only replay actually ships,
// which the coarse capture policy alone cannot distinguish. The config-only
// stable digest must stay blind to the path variables (and to the argv/env values
// they are substituted into) so the two-stage worktree recheck still passes.
import { describe, expect, it } from 'vitest'
import { resolveAgentLaunch } from './resolve-agent-launch'
import {
catalogOf,
customAgent,
customId,
requestOf,
settingsOf
} from './agent-launch-test-catalog'
import type { AgentLaunchSnapshot } from '../../shared/agent-launch-host-contract'
import type { ResolveAgentLaunchOutcome } from './resolve-agent-launch'
const AGENT_ID = customId('claude', '00000000-0000-4000-8000-0000000f0001')
const CAPTURED_ENV = { API_KEY: 'v1', REGION: 'eu' }
function snapshotOf(): AgentLaunchSnapshot {
return {
version: 1,
requestedAgent: AGENT_ID,
baseAgent: 'claude',
displayLabel: 'My Agent',
mode: 'custom',
argv: ['claude'] as unknown as AgentLaunchSnapshot['argv'],
agentEnv: CAPTURED_ENV,
capturedEnvPolicy: 'full',
target: {
platform: 'linux',
execution: 'native',
shell: 'posix',
isRemote: false,
executionHostId: 'local'
}
}
}
/** A mobile replay of the same snapshot against a live definition whose env may
* have rotated since capture. */
function mobileReplay(definitionEnv: Record<string, string>): ResolveAgentLaunchOutcome {
return resolveAgentLaunch(
{
...requestOf({ selection: { kind: 'agent', agent: AGENT_ID } }),
intent: { kind: 'resume', operation: 'resume', client: 'mobile' },
reference: { kind: 'persisted', owner: 'session' },
persistedSnapshot: snapshotOf()
},
catalogOf({
customTuiAgents: [
customAgent({
id: AGENT_ID,
label: 'My Agent',
env: definitionEnv,
syncEnv: true
})
]
}),
settingsOf()
)
}
function launchOf(outcome: ResolveAgentLaunchOutcome) {
if (!outcome.ok || !('launch' in outcome)) {
throw new Error('expected the launch to resolve')
}
return outcome.launch
}
describe('admission fingerprint over replay env authorization', () => {
it('differs when a rotated key is withheld, though every coarse input matches', () => {
const authorized = launchOf(mobileReplay({ API_KEY: 'v1', REGION: 'eu' }))
const rotated = launchOf(mobileReplay({ API_KEY: 'v2', REGION: 'eu' }))
expect(authorized.agentEnv).toEqual(CAPTURED_ENV)
expect(rotated.agentEnv).toEqual({ REGION: 'eu' })
// Same argv, same capture policy ('full' — one entry survived), same replay
// definition digest: only the shipped env moved, so only it can carry the
// difference. A stale admission would otherwise ship the pre-rotation value.
expect(authorized.policy.env).toBe(rotated.policy.env)
expect(authorized.admissionGuard.fingerprint).not.toBe(rotated.admissionGuard.fingerprint)
})
it('stays identical for two replays that ship the same env', () => {
const first = launchOf(mobileReplay({ API_KEY: 'v1', REGION: 'eu' }))
const second = launchOf(mobileReplay({ API_KEY: 'v1', REGION: 'eu' }))
expect(first.admissionGuard.fingerprint).toBe(second.admissionGuard.fingerprint)
})
it('differs when a per-launch arg changes the resolved argv', () => {
const argvFingerprint = (args: string): string => {
const outcome = resolveAgentLaunch(
requestOf({ selection: { kind: 'agent', agent: AGENT_ID } }),
catalogOf({ customTuiAgents: [customAgent({ id: AGENT_ID, args })] }),
settingsOf()
)
return launchOf(outcome).admissionGuard.fingerprint
}
expect(argvFingerprint('--model a')).not.toBe(argvFingerprint('--model b'))
})
})
describe('config-only stable digest', () => {
function resolveAtWorktree(worktreePath: string): ResolveAgentLaunchOutcome {
return resolveAgentLaunch(
requestOf({
selection: { kind: 'agent', agent: AGENT_ID },
variables: { repoPath: '/repo', worktreePath }
}),
catalogOf({
customTuiAgents: [
customAgent({
id: AGENT_ID,
args: '{worktreePath}',
env: { WT: '{worktreePath}' }
})
]
}),
settingsOf()
)
}
it('ignores the paths substituted into argv and env so stage 2 rechecks cleanly', () => {
const provisional = launchOf(resolveAtWorktree('/wt-provisional'))
const authoritative = launchOf(resolveAtWorktree('/wt-real'))
expect(provisional.argv).toContain('/wt-provisional')
expect(authoritative.agentEnv.WT).toBe('/wt-real')
expect(provisional.admissionGuard.fingerprint).not.toBe(
authoritative.admissionGuard.fingerprint
)
expect(provisional.admissionGuard.stableInputDigest).toBe(
authoritative.admissionGuard.stableInputDigest
)
})
})
@@ -0,0 +1,88 @@
// A native-Windows spawn inherits this process' env, so CreateProcess sizes the
// composed inherited+custom block, not the custom layer the resolver admits.
// Measuring only the custom layer admits a launch that then dies opaquely at
// spawn, so the cap composes the real block for that target and no other.
import { describe, expect, it } from 'vitest'
import { resolveAgentLaunch } from './resolve-agent-launch'
import { catalogOf, requestOf, settingsOf } from './agent-launch-test-catalog'
import type { AgentLaunchExecutionHostId } from '../../shared/agent-launch-host-contract'
import type { ResolveAgentLaunchOutcome } from './resolve-agent-launch'
const WINDOWS_TARGET = {
platform: 'win32' as NodeJS.Platform,
shell: 'cmd' as const,
targetHomePath: 'C:\\Users\\me'
}
/** Just past WINDOWS_ENVIRONMENT_BLOCK_MAX_CODE_UNITS once composed. */
const OVERSIZED_INHERITED: NodeJS.ProcessEnv = { HUGE: 'x'.repeat(33_000) }
function resolveWith(
overrides: Partial<Parameters<typeof requestOf>[0]>,
inheritedEnv: NodeJS.ProcessEnv
): ResolveAgentLaunchOutcome {
return resolveAgentLaunch(
requestOf({ selection: { kind: 'agent', agent: 'claude' }, ...overrides }),
catalogOf({}),
settingsOf(),
inheritedEnv
)
}
function failureCodeOf(outcome: ResolveAgentLaunchOutcome): string | null {
return !outcome.ok && 'failure' in outcome ? outcome.failure.code : null
}
describe('native-Windows environment-block cap', () => {
it('rejects a launch whose inherited block already exceeds the ceiling', () => {
const outcome = resolveWith({ ...WINDOWS_TARGET }, OVERSIZED_INHERITED)
expect(failureCodeOf(outcome)).toBe('invalid_agent_env')
if (!outcome.ok && 'failure' in outcome && outcome.failure.code === 'invalid_agent_env') {
expect(outcome.failure.reason).toBe('environment_block_too_large')
expect(outcome.failure.field).toBe('env')
}
})
it('admits the same launch under a normal inherited block', () => {
expect(resolveWith({ ...WINDOWS_TARGET }, { PATH: 'C:\\bin' }).ok).toBe(true)
})
it('ignores the local inherited block for a remote target', () => {
const outcome = resolveWith(
{
isRemote: true,
targetHomePath: null,
executionHostId: 'ssh:box' as AgentLaunchExecutionHostId
},
OVERSIZED_INHERITED
)
expect(outcome.ok).toBe(true)
})
it('ignores it for a WSL target, which spawns inside the distro', () => {
const outcome = resolveWith(
{ ...WINDOWS_TARGET, executionHostId: 'wsl:Ubuntu' as AgentLaunchExecutionHostId },
OVERSIZED_INHERITED
)
expect(outcome.ok).toBe(true)
})
it('applies the same ceiling to a snapshot replay', () => {
const admitted = resolveWith({ ...WINDOWS_TARGET }, { PATH: 'C:\\bin' })
if (!admitted.ok || !('launch' in admitted)) {
throw new Error('expected the baseline launch to resolve')
}
const replay = resolveAgentLaunch(
{
...requestOf({ selection: { kind: 'agent', agent: 'claude' }, ...WINDOWS_TARGET }),
intent: { kind: 'resume', operation: 'resume', client: 'desktop' },
reference: { kind: 'persisted', owner: 'session' },
persistedSnapshot: admitted.launch.snapshot
},
catalogOf({}),
settingsOf(),
OVERSIZED_INHERITED
)
expect(failureCodeOf(replay)).toBe('invalid_agent_env')
})
})
@@ -5,7 +5,7 @@
import os from 'node:os'
import { performance } from 'node:perf_hooks'
import { describe, expect, it, vi } from 'vitest'
import { beforeEach, describe, expect, it, vi } from 'vitest'
import type { CustomTuiAgent, CustomTuiAgentId, GlobalSettings } from '../../shared/types'
import type { AgentCatalog } from '../../shared/agent-catalog-normalization'
import { resolveAgentLaunch, type ResolveAgentLaunchOutcome } from './resolve-agent-launch'
@@ -16,7 +16,17 @@ import {
requestOf,
settingsOf
} from './agent-launch-test-catalog'
import {
resolveAgentLaunchSpawn,
type AgentLaunchSpawnDeps,
type AgentLaunchSpawnTarget
} from './agent-launch-spawn'
import { AgentLaunchBoundary } from './agent-launch-boundary'
import type * as AgentCatalogProjectionsModule from './agent-catalog-projections'
import {
AgentLaunchAdmissionStore,
LaunchAdmissionCoordinator
} from './agent-launch-admission-store'
// The pure-resolver gates below inject a PREBUILT catalog, so they never see the
// normalize pass production pays per launch. This counter lets the production-
@@ -82,6 +92,27 @@ function buildFixture(
return { catalog: catalogOf({ customTuiAgents: agents }), selectedId }
}
/** RAW persisted settings — what production actually hands the launch path. The
* host normalizes these per launch, so a prebuilt-catalog fixture measures a
* strictly cheaper pipeline than the one users pay for. */
function buildSettingsFixture(
size: number,
envEntries: number
): { settings: GlobalSettings; selectedId: CustomTuiAgentId } {
const { agents, selectedId } = buildAgents(size, envEntries)
return {
settings: {
...settingsOf(),
customTuiAgents: agents,
deletedCustomTuiAgents: [],
disabledTuiAgents: [],
defaultTuiAgent: 'auto',
agentCatalogRevision: 1
} as unknown as GlobalSettings,
selectedId
}
}
const ARRAY_SCAN_PROPS = new Set([
'forEach',
'map',
@@ -239,3 +270,123 @@ describe('resolve-agent-launch performance budget', () => {
expect(okAll).toBe(true)
})
})
// P1-21: the gates above resolve a PREBUILT catalog, so they never exercise the
// normalize pass the host runs per launch — the production surface (settings →
// normalized catalog → resolver → boundary) is measured here instead.
describe('production launch path performance budget (P1-21)', () => {
const PRODUCTION_TARGET: AgentLaunchSpawnTarget = {
platform: 'linux',
shell: 'posix',
isRemote: false,
executionHostId: 'local',
targetHomePath: '/home/dev',
// null = detection unavailable, the honest host default; never a fabricated set.
detectedStockBaseAgents: null
}
function productionDeps(getSettings: () => GlobalSettings): AgentLaunchSpawnDeps {
return {
getSettings,
getCatalogRevision: () => getSettings().agentCatalogRevision ?? 1,
boundary: new AgentLaunchBoundary({
admissionStore: new AgentLaunchAdmissionStore(),
coordinator: new LaunchAdmissionCoordinator()
})
}
}
async function launchOnce(
deps: AgentLaunchSpawnDeps,
selectedId: CustomTuiAgentId
): Promise<boolean> {
const result = await resolveAgentLaunchSpawn(deps, {
request: { selection: { kind: 'agent', agent: selectedId }, prompt: 'go' },
intent: { kind: 'interactive', client: 'desktop' },
target: PRODUCTION_TARGET,
variables: { repoPath: '/repo', worktreePath: '/repo/wt' },
scope: 'perf-launch',
principal: { kind: 'local' }
})
if (result.ok) {
// Release the admitted slot so a repeated-launch loop never hits the cap.
deps.boundary.settleAgentLaunch(result.receipt.launchToken, 'failed')
}
return result.ok
}
beforeEach(() => {
normalizeCatalogCalls.count = 0
})
it('normalizes a 2,500-agent catalog once per launch, not once per resolve pass', async () => {
const { settings, selectedId } = buildSettingsFixture(2500, 8)
const deps = productionDeps(() => settings)
expect(await launchOnce(deps, selectedId)).toBe(true)
// The boundary resolves twice (initial view + coordinator re-resolve); both
// passes must share one normalized catalog.
expect(normalizeCatalogCalls.count).toBe(1)
})
it('reuses one normalized catalog across launches on an unchanged settings revision', async () => {
const { settings, selectedId } = buildSettingsFixture(2500, 8)
const deps = productionDeps(() => settings)
for (let i = 0; i < 8; i += 1) {
expect(await launchOnce(deps, selectedId)).toBe(true)
}
expect(normalizeCatalogCalls.count).toBe(1)
})
it('re-normalizes after the settings object is replaced (a real catalog edit)', async () => {
const { settings, selectedId } = buildSettingsFixture(2500, 8)
let current = settings
const deps = productionDeps(() => current)
expect(await launchOnce(deps, selectedId)).toBe(true)
current = { ...settings, agentCatalogRevision: 2 } as GlobalSettings
expect(await launchOnce(deps, selectedId)).toBe(true)
expect(normalizeCatalogCalls.count).toBe(2)
})
it('records production launch throughput as PR evidence (never a CI wall-clock assertion)', async () => {
const { settings, selectedId } = buildSettingsFixture(2500, 8)
let current = settings
const deps = productionDeps(() => current)
let okAll = true
for (let i = 0; i < 5; i += 1) {
okAll = (await launchOnce(deps, selectedId)) && okAll
}
const runs = process.env.ORCA_PERF_EVIDENCE ? 50 : 5
const coldMs: number[] = []
const warmMs: number[] = []
for (let r = 0; r < runs; r += 1) {
// Cold: a fresh settings revision, so this launch pays the normalize.
current = { ...settings, agentCatalogRevision: r + 10 } as GlobalSettings
let start = performance.now()
okAll = (await launchOnce(deps, selectedId)) && okAll
coldMs.push(performance.now() - start)
// Warm: same revision, so the normalized catalog is reused.
start = performance.now()
okAll = (await launchOnce(deps, selectedId)) && okAll
warmMs.push(performance.now() - start)
}
const p95 = (samples: number[]): number => {
const sorted = [...samples].sort((a, b) => a - b)
return sorted[Math.min(sorted.length - 1, Math.floor(sorted.length * 0.95))] ?? 0
}
// Evidence only — asserted nowhere (plan §1371: no wall-clock CI threshold).
console.log(
`[agent-launch-spawn.perf] node=${process.version} cpu=${os.cpus()[0]?.model ?? 'unknown'} ` +
`agents=2500 runs=${runs} p95ColdLaunchMs=${p95(coldMs).toFixed(3)} ` +
`p95WarmLaunchMs=${p95(warmMs).toFixed(3)}`
)
expect(okAll).toBe(true)
})
})
+45 -6
View File
@@ -21,6 +21,7 @@ import { assembleCommand } from './resolve-agent-command'
import { interpolateVariables, prepareVariableValues } from './resolve-agent-variables'
import { clientOfIntent } from './resolve-agent-env-admission'
import { checkCommandTooLong, checkEnvPayloadTooLarge } from './agent-launch-payload-caps'
import { composeAgentLaunchEnv, inheritedEnvLayer } from './compose-agent-launch-env'
import { buildResolvedLaunch, type LaunchTarget } from './resolve-agent-launch-result'
import {
resolveMobileRemoveOnlyReplayEnv,
@@ -35,6 +36,31 @@ function hasUserPathOverride(env: Record<string, string>): boolean {
return Object.keys(env).some((key) => key.toLowerCase() === 'path')
}
/** The env a native-Windows spawn actually ships. CreateProcess rejects the whole
* inherited+custom block over its ceiling, so measuring the custom layer alone
* admits a launch that then fails opaquely at spawn. Every other target measures
* argv+env bytes instead and composes nothing here. */
function spawnEnvForPayloadCap(
env: Record<string, string>,
target: LaunchTarget,
inheritedEnv: NodeJS.ProcessEnv | undefined
): Record<string, string> {
if (target.platform !== 'win32' || target.execution !== 'native' || target.isRemote) {
return env
}
// Only the machine that will spawn owns the inherited layer; a win32 target
// resolved from anywhere else has no local block to measure.
const source = inheritedEnv ?? (process.platform === 'win32' ? process.env : undefined)
if (!source) {
return env
}
return composeAgentLaunchEnv({
platform: 'win32',
inherited: inheritedEnvLayer(source),
agentEnv: env
})
}
function deriveTarget(request: ResolveAgentLaunchRequest): LaunchTarget {
return {
platform: request.platform,
@@ -79,7 +105,8 @@ function replayFromSnapshot(
request: ResolveAgentLaunchRequest,
target: LaunchTarget,
catalog: AgentCatalog,
settings: GlobalSettings
settings: GlobalSettings,
inheritedEnv: NodeJS.ProcessEnv | undefined
): ResolveAgentLaunchOutcome {
const snapshot = request.persistedSnapshot
if (!snapshot) {
@@ -169,7 +196,11 @@ function replayFromSnapshot(
if (replayCommandCap) {
return { ok: false, failure: replayCommandCap }
}
const replayEnvCap = checkEnvPayloadTooLarge(replayArgv, env, target)
const replayEnvCap = checkEnvPayloadTooLarge(
replayArgv,
spawnEnvForPayloadCap(env, target, inheritedEnv),
target
)
if (replayEnvCap) {
return { ok: false, failure: replayEnvCap }
}
@@ -209,11 +240,15 @@ function replayFromSnapshot(
}
}
/** Resolve a launch request against the normalized catalog and current settings. */
/** Resolve a launch request against the normalized catalog and current settings.
* `inheritedEnv` is the layer the spawn will inherit, needed only to size a
* native-Windows environment block; it defaults to this process' env, which is
* that layer whenever this host is the one spawning. */
export function resolveAgentLaunch(
request: ResolveAgentLaunchRequest,
catalog: AgentCatalog,
settings: GlobalSettings
settings: GlobalSettings,
inheritedEnv?: NodeJS.ProcessEnv
): ResolveAgentLaunchOutcome {
const selection = resolveSelection(request, catalog)
if (selection.kind === 'failure') {
@@ -226,7 +261,7 @@ export function resolveAgentLaunch(
const target = deriveTarget(request)
if (selection.decision.launch === 'replay-snapshot') {
return replayFromSnapshot(request, target, catalog, settings)
return replayFromSnapshot(request, target, catalog, settings, inheritedEnv)
}
const client = clientOfIntent(request.intent)
@@ -297,7 +332,11 @@ export function resolveAgentLaunch(
if (commandCap) {
return { ok: false, failure: commandCap }
}
const envCap = checkEnvPayloadTooLarge(command.argv, resolvedEnv, target)
const envCap = checkEnvPayloadTooLarge(
command.argv,
spawnEnvForPayloadCap(resolvedEnv, target, inheritedEnv),
target
)
if (envCap) {
return { ok: false, failure: envCap }
}
@@ -332,4 +332,85 @@ describe('buildAgentStartupPlanFromResolvedLaunch', () => {
expect(submitted?.draftPrompt).toBeUndefined()
})
})
describe('final assembled command ceiling', () => {
// The resolver caps only the RESOLVED argv; the RPC admits a 100k-char prompt
// that is appended here, so the ceiling has to hold for the final command.
it('moves an argv prompt off a cmd command line over the 8191-char ceiling', () => {
const launch = resolvedLaunch({
request: { platform: 'win32', shell: 'cmd', targetHomePath: 'C:\\Users\\me' }
})
const bigPrompt = 'x'.repeat(20_000)
const plan = buildAgentStartupPlanFromResolvedLaunch({ launch, prompt: bigPrompt })
expect(plan?.launchCommand).toBe(`"codex"`)
expect(plan?.launchCommand.length).toBeLessThanOrEqual(8191)
// The FULL prompt is retained — never truncated or dropped.
expect(plan?.followupPrompt).toBe(bigPrompt)
})
it('moves an oversized draft flag prompt to the paste path with no inline ceiling given', () => {
const launch = resolvedLaunch({
agent: 'claude',
request: { platform: 'win32', shell: 'cmd', targetHomePath: 'C:\\Users\\me' }
})
const bigDraft = 'x'.repeat(20_000)
const plan = buildAgentStartupPlanFromResolvedLaunch({
launch,
prompt: bigDraft,
promptDelivery: 'draft'
})
expect(plan?.launchCommand).toBe(`"claude"`)
expect(plan?.draftPrompt).toBe(bigDraft)
})
it('moves an env-var draft off argv when its cleanup clause crosses the cmd ceiling', () => {
const agent = customId('pi', '00000000-0000-4000-8000-0000000f0002')
// Base command sits just under 8191; the ` & set "ORCA_PI_PREFILL="` clause
// the draft path appends is what pushes the final command over.
const launch = resolvedLaunch({
agent,
catalog: catalogOf({
customTuiAgents: [customAgent({ id: agent, baseAgent: 'pi', args: 'y'.repeat(8_173) })]
}),
request: { platform: 'win32', shell: 'cmd', targetHomePath: 'C:\\Users\\me' }
})
const plan = buildAgentStartupPlanFromResolvedLaunch({
launch,
prompt: 'do it',
promptDelivery: 'draft'
})
expect(plan?.launchCommand.length).toBeLessThanOrEqual(8191)
expect(plan?.launchCommand).not.toContain('ORCA_PI_PREFILL')
expect(plan?.env).toBeUndefined()
expect(plan?.draftPrompt).toBe('do it')
})
it('applies the encoded-length ceiling on a powershell target', () => {
const launch = resolvedLaunch({
request: { platform: 'win32', shell: 'powershell', targetHomePath: 'C:\\Users\\me' }
})
const bigPrompt = 'x'.repeat(30_000)
const plan = buildAgentStartupPlanFromResolvedLaunch({ launch, prompt: bigPrompt })
expect(plan?.launchCommand).toBe(`& 'codex'`)
expect(plan?.followupPrompt).toBe(bigPrompt)
})
it('applies the byte ceiling on a posix target', () => {
const launch = resolvedLaunch()
const bigPrompt = 'x'.repeat(200_000)
const plan = buildAgentStartupPlanFromResolvedLaunch({ launch, prompt: bigPrompt })
expect(plan?.launchCommand).toBe(`'codex'`)
expect(plan?.followupPrompt).toBe(bigPrompt)
})
it('still inlines a prompt that fits the target ceiling', () => {
const launch = resolvedLaunch({
request: { platform: 'win32', shell: 'cmd', targetHomePath: 'C:\\Users\\me' }
})
const prompt = 'x'.repeat(4_000)
const plan = buildAgentStartupPlanFromResolvedLaunch({ launch, prompt })
expect(plan?.launchCommand).toBe(`"codex" "${prompt}"`)
expect(plan?.followupPrompt).toBeNull()
})
})
})
@@ -342,6 +342,38 @@ describe('DegradedDaemonPtyProvider', () => {
expect(provider.providesAgentSessionOwnerListings('unknown-session')).toBe(false)
})
it('withholds launch-token listing authority unless every daemon route echoes tokens', () => {
const echoing = (label: string, echoes: boolean): DaemonPtyAdapter & ProviderMock => {
const adapter = createDaemonAdapter(label)
adapter.providesLaunchTokenListings = vi.fn(() => echoes)
return adapter
}
const fallback = (): ProviderMock => createProvider('fallback')
expect(
new DegradedDaemonPtyProvider({
current: echoing('daemon', true),
legacy: [echoing('legacy', true)],
fallback: fallback()
}).providesLaunchTokenListings()
).toBe(true)
// A pre-v34 legacy daemon still owning the launched agent poisons the whole host.
expect(
new DegradedDaemonPtyProvider({
current: echoing('daemon', true),
legacy: [echoing('legacy', false)],
fallback: fallback()
}).providesLaunchTokenListings()
).toBe(false)
expect(
new DegradedDaemonPtyProvider({
current: createDaemonAdapter('unanswering'),
legacy: [],
fallback: fallback()
}).providesLaunchTokenListings()
).toBe(false)
})
it('routes fresh foreground confirmation to the session owner', async () => {
const current = createDaemonAdapter('daemon', ['daemon-session'])
const fallback = createProvider('fallback')
@@ -118,6 +118,13 @@ export class DegradedDaemonPtyProvider implements IPtyProvider {
? await provider.writeWithSettlement(id, data)
: provider.write(id, data) !== false
}
// Why every daemon route (and why an adapter that cannot answer withholds): one
// preserved pre-v34 daemon may own the very agent a pending launch is looking for and
// will never echo its token, so a missing echo must not settle the launch failed.
// Daemon adapters only — the in-process fallback has main's own token registry.
providesLaunchTokenListings = (): boolean =>
this.allDaemonAdapters().every((adapter) => adapter.providesLaunchTokenListings?.() === true)
}
resize(id: string, cols: number, rows: number): void {
this.providerFor(id).resize(id, cols, rows)
@@ -2,6 +2,7 @@ import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest'
import {
chmodSync,
existsSync,
mkdirSync,
mkdtempSync,
readFileSync,
rmSync,
@@ -61,6 +62,21 @@ describe('data recovery points', () => {
expect(points[0].createdAtMs).toBeGreaterThan(0)
expect(JSON.stringify(points)).not.toContain(dir)
expect(JSON.stringify(points)).not.toContain('settings')
expect(points[0].restorable).toBe(true)
})
it('lists an unreadable pinned backup as not restorable', () => {
// A directory at the backup path reproduces EISDIR everywhere; existence alone
// would otherwise advertise a rollback the restore path always rejects.
mkdirSync(`${dataFile}${PINNED_SUFFIX}`)
const points = listRecoveryPoints(dataFile)
expect(points).toHaveLength(1)
expect(points[0].restorable).toBe(false)
})
it('lists a torn pinned backup as not restorable', () => {
writeFileSync(`${dataFile}${PINNED_SUFFIX}`, '{"settings":')
expect(listRecoveryPoints(dataFile)[0].restorable).toBe(false)
})
it('restores atomically: freeze before replace, safety copy kept, point preserved', async () => {
+8 -3
View File
@@ -3,7 +3,10 @@
// filesystem path or raw backup contents; restore/retry are main-owned.
import { existsSync, readFileSync, renameSync, statSync } from 'node:fs'
import { pinnedPreV1BackupPath } from '../agent-launch/agent-catalog-pre-v1-backup'
import {
classifyPinnedPreV1Backup,
pinnedPreV1BackupPath
} from '../agent-launch/agent-catalog-pre-v1-backup'
import { durableWriteTempPath, writeFileDurableSync } from '../durable-file-write'
import type { RecoveryPointDto, RecoveryPointId } from '../../shared/data-recovery'
@@ -38,13 +41,15 @@ export function listRecoveryPoints(dataFile: string): RecoveryPointDto[] {
createdAtMs = stat.birthtimeMs > 0 ? stat.birthtimeMs : stat.mtimeMs
sizeBytes = stat.size
} catch {
// Metadata is best-effort; the point is still restorable.
// Metadata is best-effort; readability is decided below, not here.
}
points.push({
id: 'agent-catalog-pre-v1',
compatibility: 'previous-binary',
createdAtMs,
sizeBytes
sizeBytes,
// Existence alone would advertise a rollback nobody can actually read.
restorable: classifyPinnedPreV1Backup(pinned).state === 'usable'
})
}
return points
+54
View File
@@ -0,0 +1,54 @@
import { beforeEach, describe, expect, it, vi } from 'vitest'
const { handleMock } = vi.hoisted(() => ({ handleMock: vi.fn() }))
vi.mock('electron', () => ({
app: { quit: vi.fn() },
ipcMain: { handle: handleMock }
}))
vi.mock('../data-recovery/recovery-points', () => ({
listRecoveryPoints: vi.fn(() => []),
restoreRecoveryPoint: vi.fn(async () => ({ ok: true }))
}))
import { registerDataRecoveryHandlers } from './data-recovery'
import type { DataRecoveryMigrationStatus } from '../../shared/data-recovery'
function migrationStatusFor(store: {
getAgentCatalogMigrationError(): string | null
getAgentCatalogSchemaTooNew(): { persistedVersion: number; supportedVersion: number } | null
}): DataRecoveryMigrationStatus {
registerDataRecoveryHandlers({ ...store, getDataFilePath: () => '/tmp/orca.json' } as never)
const handler = handleMock.mock.calls.find(
(call) => call[0] === 'dataRecovery:migrationStatus'
)?.[1] as () => DataRecoveryMigrationStatus
return handler()
}
describe('dataRecovery:migrationStatus', () => {
beforeEach(() => {
handleMock.mockClear()
})
it('reports a read-only profile stamped by a newer build', () => {
expect(
migrationStatusFor({
getAgentCatalogMigrationError: () => null,
getAgentCatalogSchemaTooNew: () => ({ persistedVersion: 2, supportedVersion: 1 })
})
).toEqual({
agentCatalogMigrationError: null,
agentCatalogSchemaTooNew: { persistedVersion: 2, supportedVersion: 1 }
})
})
it('reports null for a healthy profile so the renderer shows no notice', () => {
expect(
migrationStatusFor({
getAgentCatalogMigrationError: () => null,
getAgentCatalogSchemaTooNew: () => null
})
).toEqual({ agentCatalogMigrationError: null, agentCatalogSchemaTooNew: null })
})
})
+2 -1
View File
@@ -11,7 +11,8 @@ export function registerDataRecoveryHandlers(store: Store): void {
ipcMain.handle(
'dataRecovery:migrationStatus',
(): DataRecoveryMigrationStatus => ({
agentCatalogMigrationError: store.getAgentCatalogMigrationError()
agentCatalogMigrationError: store.getAgentCatalogMigrationError(),
agentCatalogSchemaTooNew: store.getAgentCatalogSchemaTooNew()
})
)
@@ -0,0 +1,128 @@
// A profile stamped by a newer build must stay read-only: the Store latches the
// too-new state at load, refuses durable authoring writes, and never turns it
// into a retryable pinned-backup failure.
import { describe, it, expect, vi, beforeEach, afterEach } from 'vitest'
import { chmodSync, existsSync, mkdtempSync, rmSync, writeFileSync } from 'node:fs'
import { join } from 'node:path'
import { tmpdir } from 'node:os'
const testState = { dir: '' }
vi.mock('electron', () => ({
app: {
getPath: () => testState.dir
},
safeStorage: {
isEncryptionAvailable: () => false,
encryptString: (plaintext: string) => Buffer.from(plaintext, 'utf-8'),
decryptString: (ciphertext: Buffer) => ciphertext.toString('utf-8')
}
}))
vi.mock('./telemetry/client', () => ({ track: vi.fn() }))
vi.mock('./telemetry/cohort-classifier', () => ({ getCohortAtEmit: vi.fn() }))
vi.mock('./ssh/ssh-config-parser', () => ({
loadUserSshConfig: vi.fn(() => null),
sshConfigHostsToTargets: vi.fn(() => [])
}))
async function createStore(dataFile: string) {
vi.resetModules()
const { Store } = await import('./persistence')
return new Store({ dataFile })
}
function writeProfile(dataFile: string, settings: Record<string, unknown>): void {
writeFileSync(dataFile, JSON.stringify({ settings }), { mode: 0o600 })
}
const PINNED_BACKUP_SUFFIX = '.pre-agent-catalog-v1.backup'
describe('agent-catalog schema newer than this build', () => {
let dir = ''
let dataFile = ''
beforeEach(() => {
dir = mkdtempSync(join(tmpdir(), 'orca-agent-catalog-too-new-'))
testState.dir = dir
dataFile = join(dir, 'orca-data.json')
vi.spyOn(console, 'error').mockImplementation(() => {})
})
afterEach(() => {
chmodSync(dir, 0o755)
rmSync(dir, { recursive: true, force: true })
vi.restoreAllMocks()
})
it('latches the too-new state at load without a retryable migration error', async () => {
writeProfile(dataFile, {
defaultTuiAgent: 'codex',
agentCatalogSchemaVersion: 2,
agentCatalogRevision: 7
})
const store = await createStore(dataFile)
expect(store.getAgentCatalogSchemaTooNew()).toEqual({
persistedVersion: 2,
supportedVersion: 1
})
// A too-new profile is not a blocked migration: retrying a backup fixes nothing.
expect(store.getAgentCatalogMigrationError()).toBeNull()
expect(store.retryAgentCatalogMigration()).toEqual({ ok: true })
expect(store.getSettings().agentCatalogSchemaVersion).toBe(2)
expect(store.getSettings().agentCatalogRevision).toBe(7)
expect(existsSync(`${dataFile}${PINNED_BACKUP_SUFFIX}`)).toBe(false)
})
it('refuses durable authoring writes while the profile is too new', async () => {
writeProfile(dataFile, {
agentCatalogSchemaVersion: 3,
defaultTuiAgent: 'codex'
})
const store = await createStore(dataFile)
expect(() => store.updateSettingsDurable({ defaultTuiAgent: 'auto' })).toThrow(
/newer than this build supports/
)
expect(store.getSettings().defaultTuiAgent).toBe('codex')
})
it('allows durable authoring writes on a supported profile', async () => {
writeProfile(dataFile, {
defaultTuiAgent: 'codex',
agentCatalogSchemaVersion: 1,
agentCatalogRevision: 1,
agentReferenceRevision: 1
})
const store = await createStore(dataFile)
expect(store.getAgentCatalogSchemaTooNew()).toBeNull()
expect(store.updateSettingsDurable({ defaultTuiAgent: 'auto' }).defaultTuiAgent).toBe('auto')
})
// Why skip on win32: chmod cannot make a directory read-only on Windows.
it.skipIf(process.platform === 'win32')(
'converts a blocked migration retry into read-only when the profile turned out newer',
async () => {
writeFileSync(dataFile, JSON.stringify({ settings: { defaultTuiAgent: null } }), {
mode: 0o600
})
chmodSync(dir, 0o500)
const store = await createStore(dataFile)
chmodSync(dir, 0o755)
expect(store.getAgentCatalogMigrationError()).not.toBeNull()
// A newer build's stamp reached this session (e.g. a synced profile) while
// the pinned backup was still owed.
store.updateSettings({ agentCatalogSchemaVersion: 4 })
const retry = store.retryAgentCatalogMigration()
expect(retry.ok).toBe(false)
expect(retry.ok === false && retry.error).toMatch(/read-only/)
expect(store.getAgentCatalogSchemaTooNew()).toEqual({
persistedVersion: 4,
supportedVersion: 1
})
// The retryable error is dropped: nothing about a backup can make it writable.
expect(store.getAgentCatalogMigrationError()).toBeNull()
expect(existsSync(`${dataFile}${PINNED_BACKUP_SUFFIX}`)).toBe(false)
}
)
})
@@ -0,0 +1,71 @@
import { beforeEach, describe, expect, it, vi } from 'vitest'
import { LAUNCH_TOKEN_ECHO_PROTOCOL_VERSION } from '../../shared/agent-launch-token-echo-protocol'
import { SshPtyProvider } from './ssh-pty-provider'
describe('SSH launch-token echo negotiation', () => {
const request = vi.fn()
let provider: SshPtyProvider
beforeEach(() => {
request.mockReset()
provider = new SshPtyProvider('conn-1', {
request,
notify: vi.fn(),
onNotification: vi.fn(),
dispose: vi.fn(),
isDisposed: vi.fn(() => false)
} as never)
})
function spawnParams(): Record<string, unknown> {
const call = request.mock.calls.find(([method]) => method === 'pty.spawn')
return (call?.[1] ?? {}) as Record<string, unknown>
}
// New main + old relay: the relay accepts launchToken and never re-lists it, so a
// crash-recovery re-list would read the live agent as absent and Retry would duplicate it.
it('withholds the token from a relay that cannot echo it', async () => {
request.mockImplementation(async (method: string) =>
method === 'pty.getCapabilities' ? {} : { id: 'pty-old', incarnationId: 'inc-old' }
)
await provider.spawn({ cols: 80, rows: 24, command: 'claude', launchToken: 'tok-1' })
expect('launchToken' in spawnParams()).toBe(false)
expect(provider.providesLaunchTokenListings()).toBe(false)
})
it('sends the token once the relay advertises the echo', async () => {
request.mockImplementation(async (method: string) =>
method === 'pty.getCapabilities'
? { launchTokenEchoVersion: LAUNCH_TOKEN_ECHO_PROTOCOL_VERSION }
: { id: 'pty-new', incarnationId: 'inc-new' }
)
await provider.spawn({ cols: 80, rows: 24, command: 'claude', launchToken: 'tok-2' })
expect(spawnParams().launchToken).toBe('tok-2')
expect(provider.providesLaunchTokenListings()).toBe(true)
})
// Old main + new relay: the advertisement is purely additive, so a tokenless spawn
// (all an old main ever sends) stays byte-for-byte what it was and probes nothing.
it('leaves a tokenless spawn unprobed and unchanged', async () => {
request.mockResolvedValue({ id: 'pty-plain', incarnationId: 'inc-plain' })
await provider.spawn({ cols: 80, rows: 24, command: 'claude' })
expect(request.mock.calls.map(([method]) => method)).toEqual(['pty.spawn'])
expect('launchToken' in spawnParams()).toBe(false)
})
it('re-probes a negative echo capability after an in-place relay upgrade', async () => {
request.mockResolvedValueOnce({}).mockResolvedValueOnce({
launchTokenEchoVersion: LAUNCH_TOKEN_ECHO_PROTOCOL_VERSION
})
await expect(provider.supportsLaunchTokenEcho()).resolves.toBe(false)
await expect(provider.supportsLaunchTokenEcho()).resolves.toBe(true)
expect(request).toHaveBeenCalledTimes(2)
})
})
@@ -11,6 +11,9 @@ const MOBILE_DYNAMIC_RPC_METHODS = [
'accounts.selectCodexForTarget',
'terminal.createAgentSession',
'terminal.ensureAgentSession',
// agent-catalog-sync passes the dedicated read and its legacy settings.get
// fallback through one sender, so neither is a literal sendRequest argument.
'settings.agentCatalog.get',
'github.updateIssue',
'github.updatePRState',
'gitlab.updateIssue',
@@ -7,6 +7,11 @@ import { OrcaRuntimeService } from './orca-runtime'
import { resolveTerminalAgentLaunch } from './terminal-agent-launch-resolution'
import { buildClaudeAgentTeamsLaunchPlan } from './claude-agent-teams-shim-env'
import { getHostAgentLaunchBoundary } from '../agent-launch/agent-launch-boundary-host'
import {
executionHostIdForDescriptor,
platformForDescriptor,
type AgentLaunchHostDescriptor
} from '../agent-launch/agent-launch-host-state'
vi.mock('electron', () => ({
BrowserWindow: { fromId: vi.fn(() => null) },
@@ -158,3 +163,53 @@ describe('createTerminal host-resolved agentLaunch', () => {
})
})
})
// A WSL surface classified native/local host-ids as `local`, which resolves
// execution 'native' and leaves Windows/UNC {repoPath} values in the Linux argv.
describe('WSL execution-host classification', () => {
type DescriptorInternals = {
buildTerminalAgentLaunchDescriptor: (
workspace: Record<string, unknown>
) => AgentLaunchHostDescriptor
buildRepoAgentLaunchDescriptor: (repo: Record<string, unknown>) => AgentLaunchHostDescriptor
}
function makeDescriptorRuntime(): DescriptorInternals {
const runtime = new OrcaRuntimeService({
getSettings: () => ({}),
getProjects: () => []
} as never)
return runtime as unknown as DescriptorInternals
}
function workspace(path: string): Record<string, unknown> {
return { id: 'wt-1', path, connectionId: null, repo: null, folderWorkspace: null }
}
it('describes a WSL UNC workspace as its own wsl host, never local', () => {
const descriptor = makeDescriptorRuntime().buildTerminalAgentLaunchDescriptor(
workspace('\\\\wsl.localhost\\Ubuntu\\home\\me\\app')
)
expect(descriptor).toEqual({ kind: 'wsl', distro: 'Ubuntu' })
expect(executionHostIdForDescriptor(descriptor)).toBe('wsl:Ubuntu')
expect(platformForDescriptor(descriptor)).toBe('linux')
})
it('describes a WSL UNC repo as a wsl host at worktree-create time', () => {
const descriptor = makeDescriptorRuntime().buildRepoAgentLaunchDescriptor({
id: 'r1',
path: '\\\\wsl$\\Debian\\srv\\app',
connectionId: null
})
expect(descriptor).toEqual({ kind: 'wsl', distro: 'Debian' })
expect(executionHostIdForDescriptor(descriptor)).toBe('wsl:Debian')
})
it('keeps an ordinary local workspace on the local host', () => {
expect(
makeDescriptorRuntime().buildTerminalAgentLaunchDescriptor(workspace('/repo/app'))
).toMatchObject({ kind: 'local' })
})
})
@@ -62,9 +62,19 @@ type Internals = {
filter?: (pending: PendingAgentLaunchSnapshot) => boolean,
relistedTokenPtyIds?: ReadonlyMap<string, string>
) => void
refreshPtyWorktreeRecordsWithControllerInventory: (
resolvedWorktrees: unknown[],
targetWorktreeId: string | null,
deadline?: number,
connectionId?: string | null
) => Promise<unknown>
}
function makeRuntime(): { internals: Internals; metaWrites: Record<string, unknown>[] } {
function makeRuntime(): {
runtime: OrcaRuntimeService
internals: Internals
metaWrites: Record<string, unknown>[]
} {
const runtime = new OrcaRuntimeService()
const internals = runtime as unknown as Internals
const metaWrites: Record<string, unknown>[] = []
@@ -77,7 +87,7 @@ function makeRuntime(): { internals: Internals; metaWrites: Record<string, unkno
},
getSettings: () => ({})
}
return { internals, metaWrites }
return { runtime, internals, metaWrites }
}
afterEach(() => {
@@ -141,3 +151,57 @@ describe('reconcilePendingAgentLaunches liveness composition', () => {
expect(opStore.isSpawnInFlight('tok-spawning')).toBe(false)
})
})
// A scoped SSH re-list that SUCCEEDS speaks for its own connection even when it
// returns zero sessions. Deriving the authoritative scope from the response rows
// instead left such a host "unknown" forever: its pending held capacity and could
// never be retried.
describe('scoped SSH re-list authority', () => {
function makeSshRuntime(listProcesses: () => Promise<unknown[]>) {
const made = makeRuntime()
made.runtime.setPtyController({
listProcesses,
hasPty: () => false
} as never)
return made
}
it('settles a pending on a successful but empty scoped re-list', async () => {
const opStore = getHostAgentLaunchOperationStore()
opStore.rebuildPendingFrom([pending('tok-ssh-empty')])
const { internals } = makeSshRuntime(async () => [])
await internals.refreshPtyWorktreeRecordsWithControllerInventory([], null, undefined, 'host-a')
expect(opStore.getPending('tok-ssh-empty')).toBeNull()
expect(opStore.findSettledByIdempotencyKey(WORKTREE_ID, 'key-tok-ssh-empty')).toMatchObject({
status: 'failed'
})
})
it('leaves the pending untouched when the scoped re-list fails', async () => {
const opStore = getHostAgentLaunchOperationStore()
opStore.rebuildPendingFrom([pending('tok-ssh-failed')])
const { internals, metaWrites } = makeSshRuntime(async () => {
throw new Error('relay down')
})
await internals.refreshPtyWorktreeRecordsWithControllerInventory([], null, undefined, 'host-a')
// "Re-list failed" is not evidence of absence: retryable, still holding its slot.
expect(opStore.getPending('tok-ssh-failed')).not.toBeNull()
expect(opStore.findSettledByIdempotencyKey(WORKTREE_ID, 'key-tok-ssh-failed')).toBeNull()
expect(metaWrites).toEqual([])
})
it('does not let one connection speak for another SSH host', async () => {
const opStore = getHostAgentLaunchOperationStore()
opStore.rebuildPendingFrom([pending('tok-other-host')])
const { internals } = makeSshRuntime(async () => [])
await internals.refreshPtyWorktreeRecordsWithControllerInventory([], null, undefined, 'host-b')
// The pending targets ssh:host-a; host-b's empty list says nothing about it.
expect(opStore.getPending('tok-other-host')).not.toBeNull()
})
})
@@ -0,0 +1,205 @@
// P1-4: a peer that predates the launch-token echo (pre-v34 daemon, old SSH
// relay) accepts `launchToken` on create and silently drops it, so its re-list
// can NEVER carry one. Reading that missing echo as absence settled spawn_failed
// for a launch the peer was still running, and the user's Retry then spawned a
// DUPLICATE agent beside it. The runtime must therefore AND its re-list
// authority with the provider's token-echo authority.
import { afterEach, describe, expect, it, vi } from 'vitest'
import { OrcaRuntimeService } from './orca-runtime'
import { getHostAgentLaunchOperationStore } from '../agent-launch/agent-launch-operation-store-host'
import { retryRecoveryGateForFailureCode } from '../agent-launch/agent-launch-reconciliation'
import type { PendingAgentLaunchSnapshot } from '../agent-launch/agent-launch-operation-store'
import type { AgentLaunchSnapshot } from '../../shared/agent-launch-host-contract'
import type { IPtyProvider } from '../providers/types'
vi.mock('electron', () => ({
BrowserWindow: { fromId: vi.fn(() => null) },
webContents: { fromId: vi.fn(() => null) },
ipcMain: { on: vi.fn(), removeListener: vi.fn() },
app: { getPath: vi.fn(() => '/tmp') }
}))
const WORKTREE_ID = 'r1::/wt-a'
const CONNECTION_ID = 'host-a'
const LIVE_SESSION_ID = `ssh:${CONNECTION_ID}@@pty-1`
function snapshot(): AgentLaunchSnapshot {
return {
version: 1,
requestedAgent: 'claude',
baseAgent: 'claude',
displayLabel: 'Claude',
mode: 'built-in',
argv: ['claude'],
agentEnv: {},
capturedEnvPolicy: 'none',
target: {
platform: 'linux',
execution: 'native',
shell: 'posix',
isRemote: true,
executionHostId: `ssh:${CONNECTION_ID}`
}
}
}
function pending(token: string): PendingAgentLaunchSnapshot {
return {
operationId: `op-${token}`,
idempotencyKey: `key-${token}`,
scope: WORKTREE_ID,
clientMutationId: null,
payloadDigest: 'digest-1',
launchToken: token,
intent: 'interactive',
principal: { kind: 'local' },
snapshot: snapshot()
}
}
type Internals = {
store: unknown
refreshPtyWorktreeRecordsWithControllerInventory: (
resolvedWorktrees: unknown[],
targetWorktreeId: string | null,
deadline?: number,
connectionId?: string | null
) => Promise<unknown>
}
/** A relay re-list that returns `sessions`, fronted by an SSH provider that does
* or does not echo launch tokens back in its listings. */
function makeRuntime(args: { echoesLaunchTokens: boolean; sessions: unknown[] }): {
internals: Internals
metaWrites: Record<string, unknown>[]
} {
const provider = {
providesLaunchTokenListings: () => args.echoesLaunchTokens
} as unknown as IPtyProvider
const runtime = new OrcaRuntimeService(null, undefined, {
getSshProvider: () => provider
})
const internals = runtime as unknown as Internals
const metaWrites: Record<string, unknown>[] = []
internals.store = {
getWorktreeMeta: () => ({}),
setWorkspaceSession: () => {},
setWorktreeMeta: (_id: string, meta: Record<string, unknown>) => {
metaWrites.push(meta)
return meta
},
getSettings: () => ({})
}
runtime.setPtyController({
listProcesses: async () => args.sessions,
hasPty: () => false,
getForegroundProcess: async () => null
} as never)
return { internals, metaWrites }
}
/** One live session the peer lists WITHOUT a launchToken (it dropped the one it
* was handed) — the surviving agent the pending launch belongs to. */
function tokenlessLiveSession(worktreeId: string | null): Record<string, unknown> {
return {
id: LIVE_SESSION_ID,
cwd: '/remote/wt-a',
title: 'claude',
...(worktreeId ? { worktreeId } : {})
}
}
afterEach(() => {
// The host operation store is a singleton; leave nothing behind for other suites.
getHostAgentLaunchOperationStore().rebuildPendingFrom([])
getHostAgentLaunchOperationStore().rebuildSettledFrom([])
})
describe('launch-token echo authority in reconcile', () => {
it('holds the launch pending when the host cannot echo tokens (no duplicate on Retry)', async () => {
const opStore = getHostAgentLaunchOperationStore()
opStore.rebuildPendingFrom([pending('tok-legacy')])
const { internals, metaWrites } = makeRuntime({
echoesLaunchTokens: false,
sessions: [tokenlessLiveSession(WORKTREE_ID)]
})
await internals.refreshPtyWorktreeRecordsWithControllerInventory(
[],
null,
undefined,
CONNECTION_ID
)
// Never settled: the reservation and private snapshot survive for a real proof.
expect(opStore.getPending('tok-legacy')).not.toBeNull()
expect(opStore.findSettledByIdempotencyKey(WORKTREE_ID, 'key-tok-legacy')).toBeNull()
const failure = metaWrites.at(-1)?.agentLaunchFailure as { code?: string } | undefined
expect(failure?.code).toBe('launch_state_unknown')
// The Retry that duplicated the live agent is gated behind an explicit Forget.
expect(retryRecoveryGateForFailureCode('launch_state_unknown')).toEqual({
kind: 'launch_state_unknown'
})
})
it('settles absent on the same re-list when the host DOES echo tokens', async () => {
const opStore = getHostAgentLaunchOperationStore()
opStore.rebuildPendingFrom([pending('tok-v34')])
const { internals } = makeRuntime({
echoesLaunchTokens: true,
sessions: [tokenlessLiveSession(WORKTREE_ID)]
})
await internals.refreshPtyWorktreeRecordsWithControllerInventory(
[],
null,
undefined,
CONNECTION_ID
)
// A v34 peer lists every token it holds, so a missing echo IS absence proof.
expect(opStore.getPending('tok-v34')).toBeNull()
expect(opStore.findSettledByIdempotencyKey(WORKTREE_ID, 'key-tok-v34')).toMatchObject({
status: 'failed'
})
})
it('still settles a non-echoing host absent when nothing live remains in the worktree', async () => {
const opStore = getHostAgentLaunchOperationStore()
opStore.rebuildPendingFrom([pending('tok-gone')])
const { internals } = makeRuntime({ echoesLaunchTokens: false, sessions: [] })
await internals.refreshPtyWorktreeRecordsWithControllerInventory(
[],
null,
undefined,
CONNECTION_ID
)
// Pre-token identification still proves absence, so recovery is not stranded.
expect(opStore.findSettledByIdempotencyKey(WORKTREE_ID, 'key-tok-gone')).toMatchObject({
status: 'failed'
})
})
it('holds the launch when a non-echoing host lists a session it cannot attribute', async () => {
const opStore = getHostAgentLaunchOperationStore()
opStore.rebuildPendingFrom([pending('tok-unattributed')])
const { internals } = makeRuntime({
echoesLaunchTokens: false,
sessions: [tokenlessLiveSession(null)]
})
await internals.refreshPtyWorktreeRecordsWithControllerInventory(
[],
null,
undefined,
CONNECTION_ID
)
// The session never enters ptysById and carries no token: it may BE this launch.
expect(opStore.getPending('tok-unattributed')).not.toBeNull()
expect(opStore.findSettledByIdempotencyKey(WORKTREE_ID, 'key-tok-unattributed')).toBeNull()
})
})
@@ -0,0 +1,158 @@
// P1-14 (ninth principal site): a terminal-create agentLaunch resolves against
// the PAIRED DEVICE's admission principal, so two phones cannot spend one
// launch-capacity bucket. An unpaired/local caller keeps the coarse principal —
// the bucket pre-device persisted rows were counted in.
import { describe, expect, it, vi } from 'vitest'
import { OrcaRuntimeService } from './orca-runtime'
import { resolveTerminalAgentLaunch } from './terminal-agent-launch-resolution'
import {
AgentLaunchAdmissionStore,
MAX_PENDING_LAUNCHES_PER_PRINCIPAL,
admissionPrincipalOwns,
type AdmissionPrincipal
} from '../agent-launch/agent-launch-admission-store'
import type { AgentLaunchSnapshot } from '../../shared/agent-launch-host-contract'
vi.mock('electron', () => ({
BrowserWindow: { fromId: vi.fn(() => null) },
webContents: { fromId: vi.fn(() => null) },
ipcMain: { on: vi.fn(), removeListener: vi.fn() },
app: { getPath: vi.fn(() => '/tmp') }
}))
vi.mock('./terminal-agent-launch-resolution', async (importOriginal) => {
const actual = await importOriginal<Record<string, unknown>>()
return {
...actual,
// A typed pre-spawn failure short-circuits both create paths after the
// principal has already been built — all this test cares about.
resolveTerminalAgentLaunch: vi.fn(async () => ({
kind: 'failed',
outcome: {
status: 'failed',
failure: { code: 'base_agent_disabled', baseAgent: 'claude' }
}
}))
}
})
const resolveMock = vi.mocked(resolveTerminalAgentLaunch)
type Internals = {
resolveAgentTerminalCreateOptions: (workspace: unknown, opts: unknown) => Promise<unknown>
resolveMobileTerminalStartup: (workspace: unknown, opts: unknown) => Promise<unknown>
}
const WORKSPACE = { id: 'wt-1', path: '/wt', connectionId: null, repo: null }
const AGENT_LAUNCH = { selection: { kind: 'agent', agent: 'claude' } }
function makeRuntime(): Internals {
const runtime = new OrcaRuntimeService()
;(runtime as unknown as { store: unknown }).store = {
getSettings: () => ({})
}
return runtime as unknown as Internals
}
async function principalFor(
path: 'terminal' | 'mobile',
opts: { clientKind?: 'mobile' | 'runtime'; deviceId?: string }
): Promise<AdmissionPrincipal> {
resolveMock.mockClear()
const internals = makeRuntime()
const createOpts = { agentLaunch: AGENT_LAUNCH, ...opts }
await (path === 'terminal'
? internals.resolveAgentTerminalCreateOptions(WORKSPACE, createOpts)
: internals.resolveMobileTerminalStartup(WORKSPACE, createOpts))
return (resolveMock.mock.calls[0]![1] as { principal: AdmissionPrincipal }).principal
}
function snapshot(): AgentLaunchSnapshot {
return {
baseAgent: 'claude',
target: { executionHostId: 'local' }
} as unknown as AgentLaunchSnapshot
}
function admit(
store: AgentLaunchAdmissionStore,
principal: AdmissionPrincipal,
scope: string
): boolean {
return store.admit({
principal,
intent: 'interactive',
scope,
worktreeId: null,
fingerprint: 'fp',
snapshot: snapshot(),
admittedAt: 1
}).ok
}
describe.each(['terminal', 'mobile'] as const)(
'%s terminal-create agentLaunch admission principal',
(path) => {
it('scopes the principal to the authenticated paired device', async () => {
expect(
await principalFor(path, {
clientKind: 'mobile',
deviceId: 'device-a'
})
).toEqual({
kind: 'remote',
id: 'mobile',
deviceId: 'device-a'
})
})
it('keeps the coarse principal when the transport carries no paired device', async () => {
// Legacy/pre-device rows were persisted under this exact key; an
// `undefined`-valued deviceId would fork it.
expect(await principalFor(path, { clientKind: 'mobile' })).toEqual({
kind: 'remote',
id: 'mobile'
})
})
it('is local for an in-process desktop caller', async () => {
expect(await principalFor(path, {})).toEqual({ kind: 'local' })
})
it('gives two paired devices isolated capacity buckets', async () => {
const deviceA = await principalFor(path, {
clientKind: 'mobile',
deviceId: 'device-a'
})
const deviceB = await principalFor(path, {
clientKind: 'mobile',
deviceId: 'device-b'
})
const store = new AgentLaunchAdmissionStore()
for (let i = 0; i < MAX_PENDING_LAUNCHES_PER_PRINCIPAL; i += 1) {
expect(admit(store, deviceA, `wt-a-${i}`)).toBe(true)
}
expect(admit(store, deviceA, 'wt-a-over')).toBe(false)
// Before the device id was threaded, both phones shared `remote:mobile`
// and the second device was rejected here.
expect(admit(store, deviceB, 'wt-b-1')).toBe(true)
})
it('still owns a legacy coarse-principal row admitted before device scoping', async () => {
const device = await principalFor(path, {
clientKind: 'mobile',
deviceId: 'device-a'
})
const store = new AgentLaunchAdmissionStore()
const legacy: AdmissionPrincipal = { kind: 'remote', id: 'mobile' }
expect(admit(store, legacy, 'wt-legacy')).toBe(true)
// Forget/recovery surfaces join through ownership, so a pre-upgrade row
// must stay visible and claimable by any device of its kind.
expect(admissionPrincipalOwns(device, legacy)).toBe(true)
expect(store.capacitySummaryFor(device).map((row) => row.scope)).toContain('wt-legacy')
})
}
)
@@ -1,6 +1,7 @@
import { describe, expect, it } from 'vitest'
import { CreateTerminalTab } from './session-tabs-schemas'
import { AgentLaunchInputSchema, AgentLaunchSpawnRequestSchema } from './agent-launch-spawn-schema'
import { MAX_AGENT_ARGS_CODE_UNITS } from '../../../../shared/custom-tui-agent-fields'
const CUSTOM_ID = 'custom-agent:claude:11111111-2222-4333-8444-555555555555'
@@ -88,6 +89,50 @@ describe('agentLaunch unattended declaration', () => {
})
})
// P1-24: unsaved recipe arg edits ride the launch as client text, so they must
// survive parsing AND be length-bounded like the stored recipe args.
describe('agentLaunch unsaved recipe args', () => {
it('preserves client-supplied unsaved args instead of stripping them', () => {
const parsed = AgentLaunchSpawnRequestSchema.safeParse({
selection: { kind: 'agent', agent: 'claude' },
sourceRecord: { owner: 'source-control-recipe', id: 'fixChecks' },
unsavedAgentArgs: '--model sonnet'
})
expect(parsed.success).toBe(true)
expect(parsed.success && parsed.data.unsavedAgentArgs).toBe('--model sonnet')
})
it('accepts an empty string (the user cleared the args)', () => {
const parsed = AgentLaunchSpawnRequestSchema.safeParse({
selection: { kind: 'agent', agent: 'claude' },
sourceRecord: { owner: 'source-control-recipe', id: 'fixChecks' },
unsavedAgentArgs: ''
})
expect(parsed.success).toBe(true)
expect(parsed.success && parsed.data.unsavedAgentArgs).toBe('')
})
it('rejects oversized unsaved args', () => {
const parsed = AgentLaunchSpawnRequestSchema.safeParse({
selection: { kind: 'agent', agent: 'claude' },
sourceRecord: { owner: 'source-control-recipe', id: 'fixChecks' },
unsavedAgentArgs: 'a'.repeat(MAX_AGENT_ARGS_CODE_UNITS + 1)
})
expect(parsed.success).toBe(false)
})
it('rejects a non-string on the CreateTerminalTab agentLaunch field', () => {
const parsed = CreateTerminalTab.safeParse({
worktree: 'w1',
agentLaunch: {
selection: { kind: 'agent', agent: 'claude' },
unsavedAgentArgs: ['--model', 'sonnet']
}
})
expect(parsed.success).toBe(false)
})
})
describe('agentLaunch resume/fork variant', () => {
const validKey = { worktreeId: 'wt-1', baseAgent: 'claude', providerSessionId: 'sess-1' }
@@ -19,6 +19,7 @@ import {
type ResumableTuiAgent
} from '../../../../shared/agent-session-resume'
import { AI_VAULT_AGENTS, type AiVaultAgent } from '../../../../shared/ai-vault-types'
import { MAX_AGENT_ARGS_CODE_UNITS } from '../../../../shared/custom-tui-agent-fields'
import { parseExecutionHostId } from '../../../../shared/execution-host'
const AgentLaunchSelection = z.union([
@@ -47,6 +48,9 @@ export const AgentLaunchSpawnRequestSchema: z.ZodType<AgentLaunchSpawnRequest> =
allowEmptyPromptLaunch: z.boolean().optional(),
promptDelivery: z.enum(['submit', 'draft']).optional(),
sourceRecord: AgentLaunchSourceRecord.optional(),
// Untrusted client text that substitutes for the stored recipe args, so it gets
// the same length ceiling those args are save-validated against.
unsavedAgentArgs: z.string().max(MAX_AGENT_ARGS_CODE_UNITS).optional(),
// Why: without this, an unrecognized key would be silently stripped by Zod's
// default object parsing, downgrading a caller's declared unattended/
// background launch to interactive instead of failing closed.
@@ -101,6 +101,8 @@ describe('client UI settings agent-authoring boundary', () => {
const names = CLIENT_UI_METHODS.map((method) => method.name)
const settingsMethods = names.filter((name) => name.startsWith('settings.'))
expect(settingsMethods.sort()).toEqual([
// Read-only catalog fetch; authoring stays desktop preload IPC.
'settings.agentCatalog.get',
'settings.agentReferences.get',
'settings.get',
// Terminal quick commands are plain settings state, not agent authoring.
@@ -174,4 +176,56 @@ describe('client UI settings agent-authoring boundary', () => {
expect(strings).not.toContain('SECRET_TOKEN')
expect(strings).not.toContain('super-secret-value')
})
it('omits the catalog from settings.get when the client opts out, and serves it standalone', async () => {
const catalogSettings = {
customTuiAgents: [
{
id: 'custom-agent:codex:01234567-89ab-4cde-8f01-23456789abcd',
baseAgent: 'codex',
label: 'Secret Codex',
args: '',
env: { SECRET_TOKEN: 'super-secret-value' },
syncEnv: true
}
],
agentCatalogRevision: 7
} as unknown as GlobalSettings
const runtime = {
getRuntimeId: () => 'test-runtime',
getClientSettings: vi.fn(() => ({ defaultTaskSource: 'github' })),
getAgentCatalogSnapshot: vi.fn(() => buildAgentCatalogSnapshot(catalogSettings)),
getAgentReferenceRevision: vi.fn(() => 4)
} as unknown as OrcaRuntimeService
const dispatcher = new RpcDispatcher({ runtime, methods: CLIENT_UI_METHODS })
const opted = await dispatcher.dispatch(
makeRequest('settings.get', { includeAgentCatalog: false })
)
expect(opted.ok).toBe(true)
const optedResult = (opted as { result: Record<string, unknown> }).result
expect(optedResult).not.toHaveProperty('agentCatalog')
expect(optedResult.agentReferences).toEqual({ version: 1, revision: 4 })
expect(runtime.getAgentCatalogSnapshot).not.toHaveBeenCalled()
// Explicit true and omitted both keep the legacy piggyback for old clients.
const included = await dispatcher.dispatch(
makeRequest('settings.get', { includeAgentCatalog: true })
)
expect((included as { result: Record<string, unknown> }).result.agentCatalog).toMatchObject({
version: 1,
revision: 7
})
const standalone = await dispatcher.dispatch(makeRequest('settings.agentCatalog.get'))
expect(standalone.ok).toBe(true)
const standaloneResult = (standalone as { result: Record<string, unknown> }).result
expect(standaloneResult.agentCatalog).toMatchObject({ version: 1, revision: 7 })
expect(standaloneResult).not.toHaveProperty('settings')
const strings: string[] = []
collectStringsAndKeys(standaloneResult, strings)
expect(strings).not.toContain('SECRET_TOKEN')
expect(strings).not.toContain('super-secret-value')
})
})
+26 -5
View File
@@ -1,8 +1,14 @@
import { omitPairingLocalUiFields } from '../../../../shared/pairing-local-ui-fields'
import { z } from 'zod'
import type { PersistedUIState } from '../../../../shared/persisted-ui-state-types'
import { defineMethod, type RpcMethod } from '../core'
import { PRBotAuthorOverrideUpdate, SettingsUpdate } from './client-settings-schemas'
import { FeatureInteractionIdParam, UiUpdate } from './client-ui-schemas'
import { OptionalBoolean } from '../schemas'
import {
FeatureInteractionIdParam,
PRBotAuthorOverrideUpdate,
SettingsUpdate,
UiUpdate
} from './client-ui-schemas'
// Type-only side effect: keeps the schema/PersistedUIState parity assertions in
// the typecheck graph so drift fails the build instead of a paired client.
@@ -30,18 +36,33 @@ const AGENT_REJECTED_SETTINGS_UPDATE_KEYS = [
'sourceControlAi'
] as const
// Why: the catalog projection is budgeted at 512 KiB, so a client that only wants
// settings opts out and fetches it from settings.agentCatalog.get instead. Omitted
// keeps the legacy piggyback: an old client has no other way to read the catalog.
const SettingsGet = z.object({ includeAgentCatalog: OptionalBoolean }).strict()
export const CLIENT_UI_METHODS: RpcMethod[] = [
defineMethod({
name: 'settings.get',
params: null,
handler: (_params, { runtime }) => ({
params: SettingsGet,
handler: (params, { runtime }) => ({
settings: runtime.getClientSettings(),
agentCatalog: runtime.getAgentCatalogSnapshot(),
...(params.includeAgentCatalog === false
? {}
: { agentCatalog: runtime.getAgentCatalogSnapshot() }),
// Small capability descriptor; the full snapshot ships from
// settings.agentReferences.get so the two never compete under one frame.
agentReferences: { version: 1 as const, revision: runtime.getAgentReferenceRevision() }
})
}),
defineMethod({
name: 'settings.agentCatalog.get',
params: null,
// Why: dedicated read so the catalog rides its own frame instead of every
// unrelated settings.get; a client that gets method_not_found here is talking
// to an old host and falls back to the settings.get piggyback.
handler: (_params, { runtime }) => ({ agentCatalog: runtime.getAgentCatalogSnapshot() })
}),
defineMethod({
name: 'settings.agentReferences.get',
params: null,
@@ -71,7 +71,9 @@ export const SESSION_TAB_METHODS: RpcAnyMethod[] = [
...(params.launchAgent ? { launchAgent: params.launchAgent } : {}),
...(params.viewMode ? { viewMode: params.viewMode } : {}),
// Authenticated RPC scope for host-resolved launches; never client JSON.
// The paired device narrows the admission principal to this device.
clientKind,
...(pairedDeviceId ? { deviceId: pairedDeviceId } : {}),
activate: params.activate,
select: params.select,
clientNavigationId: pairedDeviceId,
@@ -0,0 +1,97 @@
// P1-14: both terminal-create surfaces must carry the AUTHENTICATED paired
// device into the runtime create options, or two phones share one launch
// admission principal (and one capacity bucket). The device id comes from the
// envelope only — never from params.
import { describe, expect, it, vi } from 'vitest'
import { isStreamingMethod, type RpcAnyMethod, type RpcContext, type RpcMethod } from '../core'
import { TERMINAL_METHODS } from './terminal'
import { SESSION_TAB_METHODS } from './session-tabs'
const AGENT_LAUNCH = {
selection: { kind: 'agent', agent: 'claude' },
prompt: 'hi'
}
function method(methods: readonly RpcAnyMethod[], name: string): RpcMethod {
const found = methods.find((candidate) => candidate.name === name)
if (!found || isStreamingMethod(found)) {
throw new Error(`${name} method missing`)
}
return found
}
async function createTerminalOpts(
ctx: Partial<RpcContext>,
params: Record<string, unknown> = {}
): Promise<Record<string, unknown>> {
const createTerminal = vi.fn(async (_worktree: unknown, opts: Record<string, unknown>) => ({
handle: 'terminal-1',
worktreeId: 'wt-1',
opts
}))
await method(TERMINAL_METHODS, 'terminal.create').handler(
{ worktree: 'id:wt-1', agentLaunch: AGENT_LAUNCH, ...params },
{ runtime: { createTerminal }, ...ctx } as unknown as RpcContext
)
return createTerminal.mock.calls[0]![1]
}
async function createMobileTabOpts(
ctx: Partial<RpcContext>,
params: Record<string, unknown> = {}
): Promise<Record<string, unknown>> {
const createMobileSessionTerminal = vi.fn(
async (_worktree: unknown, opts: Record<string, unknown>) => ({
tab: { id: 'tab-1::leaf-1' },
opts
})
)
await method(SESSION_TAB_METHODS, 'session.tabs.createTerminal').handler(
{ worktree: 'id:wt-1', agentLaunch: AGENT_LAUNCH, ...params },
{
runtime: { createMobileSessionTerminal },
...ctx
} as unknown as RpcContext
)
return createMobileSessionTerminal.mock.calls[0]![1]
}
const PATHS: [
string,
(ctx: Partial<RpcContext>, params?: Record<string, unknown>) => Promise<Record<string, unknown>>
][] = [
['terminal.create', createTerminalOpts],
['session.tabs.createTerminal', createMobileTabOpts]
]
describe.each(PATHS)('%s device-scoped agent-launch admission', (_name, optsFor) => {
it('forwards the authenticated paired device alongside clientKind', async () => {
const opts = await optsFor({
clientKind: 'mobile',
pairedDeviceId: 'device-a'
})
expect(opts).toMatchObject({
clientKind: 'mobile',
deviceId: 'device-a'
})
})
it('omits deviceId (never undefined-valued) when no device is paired', async () => {
// An omitted id keeps the coarse principal pre-device rows were admitted
// under, so they stay forgettable after the upgrade.
const opts = await optsFor({ clientKind: 'mobile' })
expect(opts.clientKind).toBe('mobile')
expect('deviceId' in opts).toBe(false)
})
it('takes the device id from the envelope, never from client params', async () => {
const opts = await optsFor(
{ clientKind: 'mobile', pairedDeviceId: 'device-a' },
{ deviceId: 'device-victim' }
)
expect(opts.deviceId).toBe('device-a')
})
})
@@ -49,6 +49,7 @@ export const TERMINAL_LIFECYCLE_METHODS: RpcAnyMethod[] = [
? { terminalColorQueryReplies: params.terminalColorQueryReplies }
: {}),
clientKind,
...(pairedDeviceId ? { deviceId: pairedDeviceId } : {}),
title: params.title,
focus: params.focus === true,
rendererBacked: params.rendererBacked === true,
@@ -0,0 +1,103 @@
// P1-14: every agent-launch recovery method forwards the AUTHENTICATED
// pairedDeviceId alongside clientKind, so the runtime can scope caps, recovery
// rows, and idempotency to the calling device rather than to its client kind.
// The device id comes from the authenticated envelope only — never from params.
import { describe, expect, it, vi } from 'vitest'
import type { RpcContext } from '../core'
import { WORKTREE_AGENT_LAUNCH_RECOVERY_METHODS } from './worktree-agent-launch-recovery-methods'
const CANONICAL_UUID = 'aaaaaaaa-bbbb-4ccc-8ddd-eeeeeeeeeeee'
function handler(name: string): (params: unknown, ctx: RpcContext) => Promise<unknown> {
const method = WORKTREE_AGENT_LAUNCH_RECOVERY_METHODS.find((m) => m.name === name)
if (!method) {
throw new Error(`${name} not registered`)
}
return method.handler as (params: unknown, ctx: RpcContext) => Promise<unknown>
}
// [rpc method, runtime method, params, trailing args before (clientKind, deviceId)]
const CASES: [string, string, Record<string, unknown>][] = [
[
'worktree.retryAgentLaunch',
'retryWorktreeAgentLaunch',
{
worktree: 'id:wt-1',
expectedFailureId: 'f1',
clientMutationId: CANONICAL_UUID,
action: { kind: 'same' }
}
],
[
'worktree.forgetAgentLaunch',
'forgetUnknownWorktreeAgentLaunch',
{
worktree: 'id:wt-1',
expectedOperationId: 'op-1',
clientMutationId: CANONICAL_UUID
}
],
[
'worktree.retryBackgroundAgentLaunch',
'retryBackgroundAgentLaunch',
{
attemptId: 'att-1',
expectedFailureId: 'f1',
clientMutationId: CANONICAL_UUID,
action: { kind: 'same' }
}
],
[
'worktree.forgetBackgroundAgentLaunch',
'forgetBackgroundAgentLaunch',
{
attemptId: 'att-1',
expectedOperationId: 'op-1',
clientMutationId: CANONICAL_UUID
}
],
['worktree.pendingAgentLaunchSummary', 'pendingAgentLaunchSummary', {}],
[
'worktree.unknownAgentLaunchSiblingCount',
'unknownWorktreeAgentLaunchSiblingCount',
{ worktree: 'id:wt-1' }
],
[
'worktree.forgetUnknownAgentLaunchSiblings',
'forgetUnknownWorktreeAgentLaunchSiblings',
{ worktree: 'id:wt-1' }
]
]
describe('worktree agent-launch recovery methods forward the paired device', () => {
it.each(CASES)('%s passes pairedDeviceId to %s', async (rpcName, runtimeName, params) => {
const spy = vi.fn().mockResolvedValue({})
const runtime = {
[runtimeName]: spy
} as unknown as RpcContext['runtime']
await handler(rpcName)(params, {
runtime,
clientKind: 'mobile',
pairedDeviceId: 'device-a'
})
const args = spy.mock.calls[0]
expect(args.slice(-2)).toEqual(['mobile', 'device-a'])
})
it.each(CASES)(
'%s falls back to the coarse principal when no device is paired',
async (rpcName, runtimeName, params) => {
const spy = vi.fn().mockResolvedValue({})
const runtime = {
[runtimeName]: spy
} as unknown as RpcContext['runtime']
// In-process/local dispatch: no clientKind and no paired device id.
await handler(rpcName)(params, { runtime })
expect(spy.mock.calls[0].slice(-2)).toEqual([undefined, undefined])
}
)
})
@@ -2,8 +2,11 @@
// worktree launch failure, retry/forget a generic background attempt, and the
// redacted capacity-recovery summary. Split out of worktree.ts to keep that file
// under the max-lines limit. Every method scopes admission/idempotency from the
// authenticated clientKind — never from client JSON — and treats the
// expectedFailureId/expectedOperationId fields as anti-race guards, not secrets.
// authenticated clientKind AND the authenticated pairedDeviceId — never from
// client JSON — so two paired devices get isolated caps/recovery rows and neither
// can forget the other's launch. A connection with no paired device id (in-process
// runtime callers) falls back to the pre-device coarse principal. The
// expectedFailureId/expectedOperationId fields stay anti-race guards, not secrets.
import { defineMethod, type RpcMethod } from '../core'
import {
@@ -22,7 +25,7 @@ export const WORKTREE_AGENT_LAUNCH_RECOVERY_METHODS: RpcMethod[] = [
params: WorktreeRetryAgentLaunch,
// Authorization is authenticated worktree access, the same boundary as every
// other worktree mutation.
handler: async (params, { runtime, clientKind }) =>
handler: async (params, { runtime, clientKind, pairedDeviceId }) =>
runtime.retryWorktreeAgentLaunch(
params.worktree,
{
@@ -30,27 +33,29 @@ export const WORKTREE_AGENT_LAUNCH_RECOVERY_METHODS: RpcMethod[] = [
clientMutationId: params.clientMutationId,
action: params.action
},
clientKind
clientKind,
pairedDeviceId
)
}),
defineMethod({
name: 'worktree.forgetAgentLaunch',
params: WorktreeForgetAgentLaunch,
handler: async (params, { runtime, clientKind }) =>
handler: async (params, { runtime, clientKind, pairedDeviceId }) =>
runtime.forgetUnknownWorktreeAgentLaunch(
params.worktree,
{
expectedOperationId: params.expectedOperationId,
clientMutationId: params.clientMutationId
},
clientKind
clientKind,
pairedDeviceId
)
}),
defineMethod({
name: 'worktree.retryBackgroundAgentLaunch',
params: WorktreeRetryBackgroundAgentLaunch,
// Authorization is authenticated access to the attempt's worktree.
handler: async (params, { runtime, clientKind }) =>
handler: async (params, { runtime, clientKind, pairedDeviceId }) =>
runtime.retryBackgroundAgentLaunch(
{
attemptId: params.attemptId,
@@ -58,37 +63,43 @@ export const WORKTREE_AGENT_LAUNCH_RECOVERY_METHODS: RpcMethod[] = [
clientMutationId: params.clientMutationId,
action: params.action
},
clientKind
clientKind,
pairedDeviceId
)
}),
defineMethod({
name: 'worktree.forgetBackgroundAgentLaunch',
params: WorktreeForgetBackgroundAgentLaunch,
handler: async (params, { runtime, clientKind }) =>
handler: async (params, { runtime, clientKind, pairedDeviceId }) =>
runtime.forgetBackgroundAgentLaunch(
{
attemptId: params.attemptId,
expectedOperationId: params.expectedOperationId,
clientMutationId: params.clientMutationId
},
clientKind
clientKind,
pairedDeviceId
)
}),
defineMethod({
name: 'worktree.pendingAgentLaunchSummary',
params: WorktreePendingAgentLaunchSummary,
// clientKind scopes the admission principal (own rows only). The redacted rows
// are secret-free and carry no token.
handler: async (_params, { runtime, clientKind }) =>
runtime.pendingAgentLaunchSummary(clientKind)
// clientKind + pairedDeviceId scope the admission principal (own rows only).
// The redacted rows are secret-free and carry no token.
handler: async (_params, { runtime, clientKind, pairedDeviceId }) =>
runtime.pendingAgentLaunchSummary(clientKind, pairedDeviceId)
}),
defineMethod({
name: 'worktree.unknownAgentLaunchSiblingCount',
params: WorktreeUnknownAgentLaunchSiblingCount,
// Lazy preflight for the ":498 Also forget N other stranded launches" affordance;
// clientKind scopes the principal, siblings are host-derived, no secrets cross.
handler: async (params, { runtime, clientKind }) => ({
count: await runtime.unknownWorktreeAgentLaunchSiblingCount(params.worktree, clientKind)
// the principal is device-scoped, siblings are host-derived, no secrets cross.
handler: async (params, { runtime, clientKind, pairedDeviceId }) => ({
count: await runtime.unknownWorktreeAgentLaunchSiblingCount(
params.worktree,
clientKind,
pairedDeviceId
)
})
}),
defineMethod({
@@ -96,7 +107,7 @@ export const WORKTREE_AGENT_LAUNCH_RECOVERY_METHODS: RpcMethod[] = [
params: WorktreeForgetUnknownAgentLaunchSiblings,
// Same-principal bulk forget on the anchor's disconnected host. Never kills or
// spawns; each sibling settles only its own reservation and self-guards.
handler: async (params, { runtime, clientKind }) =>
runtime.forgetUnknownWorktreeAgentLaunchSiblings(params.worktree, clientKind)
handler: async (params, { runtime, clientKind, pairedDeviceId }) =>
runtime.forgetUnknownWorktreeAgentLaunchSiblings(params.worktree, clientKind, pairedDeviceId)
})
]
@@ -57,8 +57,8 @@ describe('worktree.retryBackgroundAgentLaunch RPC', () => {
)
expect(response).toMatchObject({ ok: true, result: { status: 'launched' } })
// clientKind is undefined for a local dispatch; it scopes the idempotency
// principal and is never derived from the client JSON.
// clientKind and pairedDeviceId are both undefined for a local dispatch;
// together they scope the idempotency principal, never client JSON.
expect(retryBackgroundAgentLaunch).toHaveBeenCalledWith(
{
attemptId: 'attempt-1',
@@ -66,6 +66,7 @@ describe('worktree.retryBackgroundAgentLaunch RPC', () => {
clientMutationId: CANONICAL_UUID,
action: { kind: 'change-agent', agent: 'codex' }
},
undefined,
undefined
)
})
@@ -115,6 +116,7 @@ describe('worktree.forgetBackgroundAgentLaunch RPC', () => {
expectedOperationId: 'op-1',
clientMutationId: CANONICAL_UUID
},
undefined,
undefined
)
})
@@ -0,0 +1,85 @@
// P1-14: worktree.create's host-atomic agentLaunch must carry the AUTHENTICATED
// paired device into createManagedWorktree, so its pre-git capacity reservation
// and idempotency key are per device rather than per client kind.
import { describe, expect, it, vi } from 'vitest'
import type { RpcContext } from '../core'
import { WORKTREE_METHODS } from './worktree'
const AGENT_LAUNCH = {
selection: { kind: 'default' as const },
allowEmptyPromptLaunch: true
}
function worktreeCreateHandler(): (params: unknown, ctx: RpcContext) => Promise<unknown> {
const found = WORKTREE_METHODS.find((candidate) => candidate.name === 'worktree.create')
if (!found) {
throw new Error('worktree.create method missing')
}
return found.handler as (params: unknown, ctx: RpcContext) => Promise<unknown>
}
async function createArgs(
ctx: Partial<RpcContext>,
params: Record<string, unknown> = {}
): Promise<Record<string, unknown>> {
const createManagedWorktree = vi.fn(async (_args: Record<string, unknown>) => ({
created: true
}))
const runtime = {
getRuntimeId: () => 'test-runtime',
dedupeWorktreeCreate: <T>(
_repo: string,
_mutationId: string | undefined,
run: () => Promise<T>
) => run(),
showRepo: vi.fn().mockResolvedValue({ id: 'repo-1', path: '/repo', kind: 'git' }),
createManagedWorktree
}
await worktreeCreateHandler()(
{ repo: 'repo-1', name: 'wt', agentLaunch: AGENT_LAUNCH, ...params },
{ runtime, ...ctx } as unknown as RpcContext
)
return createManagedWorktree.mock.calls[0]![0]
}
describe('worktree.create device-scoped agent launch', () => {
it('forwards the authenticated paired device alongside the client kind', async () => {
const args = await createArgs({
clientKind: 'mobile',
pairedDeviceId: 'device-a'
})
expect(args).toMatchObject({
agentLaunch: AGENT_LAUNCH,
agentLaunchClientKind: 'mobile',
agentLaunchDeviceId: 'device-a'
})
})
it('omits the device id (never undefined-valued) for an in-process caller', async () => {
// Keeps the coarse/local principal that pre-device persisted rows were
// reserved under, so they stay forgettable after the upgrade.
const args = await createArgs({})
expect(args.agentLaunchClientKind).toBeUndefined()
expect('agentLaunchDeviceId' in args).toBe(false)
})
it('takes the device id from the envelope, never from client params', async () => {
const args = await createArgs(
{ clientKind: 'mobile', pairedDeviceId: 'device-a' },
{ agentLaunchDeviceId: 'device-victim' }
)
expect(args.agentLaunchDeviceId).toBe('device-a')
})
it('sends no device id on the legacy (no agentLaunch) create path', async () => {
const args = await createArgs(
{ clientKind: 'mobile', pairedDeviceId: 'device-a' },
{ agentLaunch: undefined }
)
expect('agentLaunchDeviceId' in args).toBe(false)
})
})
@@ -15,7 +15,7 @@ export const WORKTREE_CREATE_METHOD: RpcMethod = defineMethod({
name: 'worktree.create',
params: WorktreeCreate,
handler: async (params, context) => {
const { runtime, clientKind } = context
const { runtime, clientKind, pairedDeviceId } = context
// U7: a remote client (authenticated clientKind) may not name a custom id on
// the legacy built-in create path — it cannot be host-resolved without the
// host-atomic agentLaunch request. Reject at the boundary (no worktree),
@@ -97,9 +97,15 @@ export const WORKTREE_CREATE_METHOD: RpcMethod = defineMethod({
startupDraft: params.startupDraft,
// The host-atomic launch request; when present the host ignores the
// client startup/createdWithAgent for the agent terminal. clientKind
// scopes admission/intent and is never derived from client JSON.
// scopes admission/intent and is never derived from client JSON; the
// paired device narrows that principal so one phone's capacity and
// recovery rows are its own.
...(params.agentLaunch
? { agentLaunch: params.agentLaunch, agentLaunchClientKind: clientKind }
? {
agentLaunch: params.agentLaunch,
agentLaunchClientKind: clientKind,
...(pairedDeviceId ? { agentLaunchDeviceId: pairedDeviceId } : {})
}
: {}),
...(params.agentLaunchTelemetry
? { agentLaunchTelemetry: params.agentLaunchTelemetry }
@@ -49,11 +49,12 @@ describe('worktree.forgetAgentLaunch RPC', () => {
)
expect(response).toMatchObject({ ok: true, result: { status: 'forgotten' } })
// clientKind is undefined for an in-process/local dispatch; it scopes the
// idempotency principal and is never derived from the client JSON.
// clientKind and pairedDeviceId are both undefined for an in-process/local
// dispatch; together they scope the idempotency principal, never client JSON.
expect(forgetUnknownWorktreeAgentLaunch).toHaveBeenCalledWith(
'id:wt-1',
{ expectedOperationId: 'op-1', clientMutationId: CANONICAL_UUID },
undefined,
undefined
)
})
@@ -32,8 +32,8 @@ describe('worktree.pendingAgentLaunchSummary RPC', () => {
)
expect(response).toMatchObject({ ok: true, result: { rows: [{ liveness: 'live' }] } })
// clientKind is undefined for an in-process/local dispatch; it scopes the
// admission principal and is never derived from the client JSON.
expect(pendingAgentLaunchSummary).toHaveBeenCalledWith(undefined)
// clientKind and pairedDeviceId are both undefined for an in-process/local
// dispatch; together they scope the admission principal, never client JSON.
expect(pendingAgentLaunchSummary).toHaveBeenCalledWith(undefined, undefined)
})
})
@@ -58,8 +58,8 @@ describe('worktree.retryAgentLaunch RPC', () => {
)
expect(response).toMatchObject({ ok: true, result: { status: 'launched' } })
// clientKind is undefined for an in-process/local dispatch; it scopes the
// idempotency principal and is never derived from the client JSON.
// clientKind and pairedDeviceId are both undefined for an in-process/local
// dispatch; together they scope the idempotency principal, never client JSON.
expect(retryWorktreeAgentLaunch).toHaveBeenCalledWith(
'id:wt-1',
{
@@ -67,6 +67,7 @@ describe('worktree.retryAgentLaunch RPC', () => {
clientMutationId: CANONICAL_UUID,
action: { kind: 'change-agent', agent: 'codex' }
},
undefined,
undefined
)
})
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,179 @@
// @vitest-environment happy-dom
import React, { act } from 'react'
import { createRoot, type Root } from 'react-dom/client'
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import type { LocalAgentCatalogSnapshot } from '../../../shared/agent-catalog-snapshot'
import type { CustomTuiAgentId, TuiAgent } from '../../../shared/types'
globalThis.IS_REACT_ACT_ENVIRONMENT = true
const CUSTOM_CODEX = 'custom-agent:codex:22222222-2222-4222-8222-222222222222' as CustomTuiAgentId
const mocks = vi.hoisted(() => ({
submitQuick: vi.fn(),
createDisabled: false,
catalog: {
snapshot: null as unknown,
loading: true,
unavailable: false
},
cardProps: [] as { createDisabled?: boolean; onCreate?: () => void }[]
}))
vi.mock('@/store', () => ({
useAppStore: (selector: (state: unknown) => unknown) =>
selector({
activeModal: 'new-workspace-composer',
modalData: {},
closeModal: vi.fn(),
settings: { defaultTuiAgent: CUSTOM_CODEX, disabledTuiAgents: [] }
})
}))
vi.mock('@/hooks/useComposerState', () => ({
useComposerState: () => ({
cardProps: {
detectedAgentIds: new Set<TuiAgent>(['codex']),
projectOptions: [],
selectedProjectId: null,
selectedRepoIsGit: true
},
composerRef: { current: null },
onComposerNodeChange: vi.fn(),
nameInputRef: { current: null },
submitQuick: mocks.submitQuick,
createDisabled: mocks.createDisabled,
selectAddedProjectRepo: vi.fn()
})
}))
vi.mock('@/hooks/useLocalAgentCatalog', () => ({
useLocalAgentCatalog: () => ({
snapshot: mocks.catalog.snapshot,
loading: mocks.catalog.loading,
unavailable: mocks.catalog.unavailable,
refetch: vi.fn(),
applySnapshot: vi.fn()
})
}))
vi.mock('@/components/NewWorkspaceComposerCard', () => ({
default: (props: { createDisabled?: boolean; onCreate?: () => void }) => {
mocks.cardProps.push(props)
return (
<button
type="button"
data-testid="create"
disabled={props.createDisabled}
onClick={props.onCreate}
>
create
</button>
)
}
}))
vi.mock('@/components/agent/AgentSettingsDialog', () => ({ default: () => null }))
vi.mock('@/lib/lazy-with-retry', () => ({ lazyWithRetry: () => () => null }))
vi.mock('@/components/ui/dialog', () => ({
Dialog: ({ children }: { children: React.ReactNode }) => <div>{children}</div>,
DialogContent: ({ children }: { children: React.ReactNode }) => <div>{children}</div>,
DialogDescription: ({ children }: { children: React.ReactNode }) => <p>{children}</p>,
DialogHeader: ({ children }: { children: React.ReactNode }) => <header>{children}</header>,
DialogTitle: ({ children }: { children: React.ReactNode }) => <h1>{children}</h1>
}))
const READY_CUSTOM_CATALOG = {
customAgents: [
{
status: 'ready',
definition: {
id: CUSTOM_CODEX,
baseAgent: 'codex',
label: 'Team Codex',
args: '--model team',
syncEnv: false,
commandOverride: '/opt/bin/codex'
},
envSummary: { entryCount: 0, bytes: 0 },
availabilityReason: 'configured-executable'
}
]
} as unknown as LocalAgentCatalogSnapshot
let container: HTMLDivElement
let root: Root
async function render(): Promise<void> {
const { default: NewWorkspaceComposerModal } = await import('./NewWorkspaceComposerModal')
await act(async () => {
root.render(<NewWorkspaceComposerModal />)
})
}
function clickCreate(): void {
const button = container.querySelector<HTMLButtonElement>('[data-testid="create"]')
expect(button).not.toBeNull()
act(() => {
button?.click()
})
}
beforeEach(() => {
mocks.submitQuick.mockReset()
mocks.createDisabled = false
mocks.catalog.snapshot = null
mocks.catalog.loading = true
mocks.catalog.unavailable = false
mocks.cardProps = []
container = document.createElement('div')
document.body.appendChild(container)
root = createRoot(container)
})
afterEach(() => {
act(() => root.unmount())
container.remove()
})
describe('NewWorkspaceComposerModal quick create', () => {
it('blocks create while the local agent catalog is still loading', async () => {
await render()
expect(mocks.cardProps.at(-1)?.createDisabled).toBe(true)
clickCreate()
expect(mocks.submitQuick).not.toHaveBeenCalled()
})
it('submits the custom default once the catalog has loaded', async () => {
mocks.catalog.snapshot = READY_CUSTOM_CATALOG
mocks.catalog.loading = false
await render()
expect(mocks.cardProps.at(-1)?.createDisabled).toBe(false)
clickCreate()
expect(mocks.submitQuick).toHaveBeenCalledWith(CUSTOM_CODEX)
})
it('still allows create where the local catalog surface does not exist', async () => {
mocks.catalog.loading = false
mocks.catalog.unavailable = true
await render()
expect(mocks.cardProps.at(-1)?.createDisabled).toBe(false)
clickCreate()
expect(mocks.submitQuick).toHaveBeenCalledWith('codex')
})
it('keeps the composer create gate disabled independently of the catalog', async () => {
mocks.createDisabled = true
mocks.catalog.snapshot = READY_CUSTOM_CATALOG
mocks.catalog.loading = false
await render()
expect(mocks.cardProps.at(-1)?.createDisabled).toBe(true)
})
})
@@ -151,7 +151,10 @@ function QuickTabBody({
enableIssueAutomation: modalData.enableIssueAutomation === true,
createGateMode: 'quick'
})
const { snapshot: localAgentCatalog } = useLocalAgentCatalog()
// Why: quick-create resolves the agent from this catalog, so submitting before
// it lands would silently swap a custom default for a built-in. `unavailable`
// (paired web) never resolves a snapshot, so it must not block creation.
const { snapshot: localAgentCatalog, loading: localAgentCatalogLoading } = useLocalAgentCatalog()
const quickAgentOptions = useMemo(
() =>
buildWorkspaceAgentOptions({
@@ -202,9 +205,14 @@ function QuickTabBody({
setQuickAgentOverride(agent)
}, [])
const createDisabledForQuickAgent = createDisabled || localAgentCatalogLoading
const handleCreate = useCallback(async (): Promise<void> => {
if (localAgentCatalogLoading) {
return
}
await submitQuick(quickAgent)
}, [quickAgent, submitQuick])
}, [localAgentCatalogLoading, quickAgent, submitQuick])
// Why: Add Project layers over the composer as a nested dialog instead of
// replacing it in the activeModal slot — closing the composer mid-flow (and
// losing the typed name/prompt) was the old, abrupt behavior. Once opened it
@@ -280,7 +288,7 @@ function QuickTabBody({
if (!shouldAllowComposerEnterSubmitTarget(target, composerRef.current)) {
return
}
if (createDisabled) {
if (createDisabledForQuickAgent) {
return
}
event.preventDefault()
@@ -288,7 +296,7 @@ function QuickTabBody({
}
window.addEventListener('keydown', onKeyDown, { capture: true })
return () => window.removeEventListener('keydown', onKeyDown, { capture: true })
}, [active, composerRef, createDisabled, handleCreate, nestedDialogOpen])
}, [active, composerRef, createDisabledForQuickAgent, handleCreate, nestedDialogOpen])
return (
<>
@@ -322,6 +330,7 @@ function QuickTabBody({
onQuickAgentChange={handleQuickAgentChange}
quickAgentOptions={quickAgentOptions}
{...cardProps}
createDisabled={createDisabledForQuickAgent}
primaryActionLabel={primaryActionLabel}
onOpenAgentSettings={() => setAgentSettingsOpen(true)}
onCreate={() => void handleCreate()}
@@ -0,0 +1,54 @@
// @vitest-environment happy-dom
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import { cleanup, render, screen } from '@testing-library/react'
import { DataRecoveryDialog } from './DataRecoveryDialog'
const PIN = {
id: 'agent-catalog-pre-v1' as const,
compatibility: 'previous-binary' as const,
createdAtMs: 1_752_000_000_000,
sizeBytes: 1024
}
const UNREADABLE_COPY = /cannot be read, so it cannot be restored/
function installApi(points: unknown[]): void {
;(window as unknown as { api: unknown }).api = {
dataRecovery: {
listPoints: vi.fn().mockResolvedValue(points),
restore: vi.fn().mockResolvedValue({ ok: true })
}
}
}
describe('DataRecoveryDialog', () => {
beforeEach(() => {
vi.clearAllMocks()
})
afterEach(() => {
cleanup()
})
it('offers the restore affordance for a readable point', async () => {
installApi([{ ...PIN, restorable: true }])
render(<DataRecoveryDialog open onOpenChange={() => {}} />)
expect(await screen.findByText('Restore these settings…')).toBeTruthy()
expect(screen.queryByText(UNREADABLE_COPY)).toBeNull()
})
it('withdraws the restore affordance when the point cannot be read', async () => {
installApi([{ ...PIN, restorable: false }])
render(<DataRecoveryDialog open onOpenChange={() => {}} />)
expect(await screen.findByText(UNREADABLE_COPY)).toBeTruthy()
expect(screen.queryByText('Restore these settings…')).toBeNull()
// The loss warning describes a restore that can no longer happen.
expect(screen.queryByText(/discards settings and custom agents/)).toBeNull()
})
it('stays restorable for a host that reports no readability at all', async () => {
installApi([PIN])
render(<DataRecoveryDialog open onOpenChange={() => {}} />)
expect(await screen.findByText('Restore these settings…')).toBeTruthy()
})
})
@@ -18,9 +18,7 @@ export type DataRecoveryDialogProps = {
/** `restorable` is optional so a host that predates the readability probe still
* lists its points; only an explicit `false` withdraws the restore affordance. */
type ListedRecoveryPoint = RecoveryPointDto & { restorable?: boolean }
function isRestorable(point: ListedRecoveryPoint): boolean {
function isRestorable(point: RecoveryPointDto): boolean {
return point.restorable !== false
}
@@ -29,7 +27,7 @@ function pointTitle(point: RecoveryPointDto): string {
case 'agent-catalog-pre-v1':
return translate(
'auto.components.dataRecovery.pointAgentCatalogPreV1Title',
'Before the custom-agents update (agent catalog v1)'
'Before custom agents'
)
}
}
@@ -39,7 +37,7 @@ function pointLossSummary(point: RecoveryPointDto): string {
case 'agent-catalog-pre-v1':
return translate(
'auto.components.dataRecovery.pointAgentCatalogPreV1Loss',
'Restoring discards all settings and workspace metadata saved after this point, including any custom agents.'
'This discards settings and custom agents saved after this update.'
)
}
}
@@ -48,8 +46,8 @@ function pointLossSummary(point: RecoveryPointDto): string {
* points by metadata only; the pinned pre-v1 point restores via Prepare
* downgrade — Orca restores atomically and quits without relaunching. */
export function DataRecoveryDialog({ open, onOpenChange }: DataRecoveryDialogProps) {
const [points, setPoints] = useState<ListedRecoveryPoint[] | null>(null)
const [confirming, setConfirming] = useState<ListedRecoveryPoint | null>(null)
const [points, setPoints] = useState<RecoveryPointDto[] | null>(null)
const [confirming, setConfirming] = useState<RecoveryPointDto | null>(null)
const [restoring, setRestoring] = useState(false)
const [restored, setRestored] = useState(false)
const [error, setError] = useState<string | null>(null)
@@ -133,7 +131,7 @@ export function DataRecoveryDialog({ open, onOpenChange }: DataRecoveryDialogPro
<DialogDescription>
{translate(
'auto.components.dataRecovery.description',
'Restore this profile from a recovery point created before a data migration. Restoring never deletes the recovery point.'
'Restore the settings Orca saved before this update. The saved copy is kept.'
)}
</DialogDescription>
</DialogHeader>
@@ -218,7 +216,7 @@ export function DataRecoveryDialog({ open, onOpenChange }: DataRecoveryDialogPro
>
{translate(
'auto.components.dataRecovery.prepareDowngrade',
'Prepare downgrade…'
'Restore these settings…'
)}
</Button>
)}
@@ -81,14 +81,54 @@ describe('DataRecoveryMigrationNotice', () => {
await vi.waitFor(() => expect(screen.queryByRole('alert')).toBeNull())
})
it('opens Data recovery listing the pinned point with Prepare downgrade', async () => {
it('shows a non-retryable read-only notice when the profile schema is too new', async () => {
installApi({
migrationStatus: vi.fn().mockResolvedValue({
agentCatalogMigrationError: null,
agentCatalogSchemaTooNew: { persistedVersion: 2, supportedVersion: 1 }
})
})
render(<DataRecoveryMigrationNotice />)
expect(
await screen.findByText('Custom agents are read-only in this version of Orca')
).toBeTruthy()
expect(screen.getByRole('alert').textContent).toContain('saved by a newer version of Orca')
expect(
screen.getByText('Profile agent catalog v2; this version of Orca supports v1.')
).toBeTruthy()
// Retry re-runs the pinned backup, which cannot clear a newer schema.
expect(screen.queryByText('Retry migration')).toBeNull()
})
it('prefers the read-only notice over a retry offer when both states are reported', async () => {
installApi({
migrationStatus: vi.fn().mockResolvedValue({
agentCatalogMigrationError: 'disk full',
agentCatalogSchemaTooNew: { persistedVersion: 3, supportedVersion: 1 }
})
})
render(<DataRecoveryMigrationNotice />)
expect(
await screen.findByText('Custom agents are read-only in this version of Orca')
).toBeTruthy()
expect(screen.queryByText('Retry migration')).toBeNull()
})
it('ignores an older host that never sends the schema field', async () => {
installApi({
migrationStatus: vi.fn().mockResolvedValue({ agentCatalogMigrationError: 'disk full' })
})
render(<DataRecoveryMigrationNotice />)
expect(await screen.findByText('Retry migration')).toBeTruthy()
expect(screen.queryByText('Custom agents are read-only in this version of Orca')).toBeNull()
})
it('opens Data recovery listing the pinned point with restore', async () => {
const api = installApi()
render(<DataRecoveryMigrationNotice />)
fireEvent.click(await screen.findByText('Open Data recovery'))
expect(
await screen.findByText('Before the custom-agents update (agent catalog v1)')
).toBeTruthy()
fireEvent.click(screen.getByText('Prepare downgrade…'))
expect(await screen.findByText('Before custom agents')).toBeTruthy()
fireEvent.click(screen.getByText('Restore these settings…'))
fireEvent.click(screen.getByText('Restore and quit'))
await vi.waitFor(() =>
expect(api.restore).toHaveBeenCalledWith({
@@ -2,6 +2,12 @@ import { useCallback, useEffect, useState } from 'react'
import { AlertTriangle } from 'lucide-react'
import { Button } from '@/components/ui/button'
import { translate } from '@/i18n/i18n'
import {
agentCatalogSchemaTooNewMessage,
agentCatalogSchemaTooNewTitle,
agentCatalogSchemaTooNewVersions
} from '../settings/agent-catalog-schema-too-new'
import type { AgentCatalogSchemaTooNew } from '../../../../shared/data-recovery'
import { DataRecoveryDialog } from './DataRecoveryDialog'
/** App-level persistent notice for a blocked agent-catalog migration (runbook:
@@ -10,6 +16,8 @@ import { DataRecoveryDialog } from './DataRecoveryDialog'
* nothing on paired web (no dataRecovery preload surface) or when healthy. */
export function DataRecoveryMigrationNotice() {
const [migrationError, setMigrationError] = useState<string | null>(null)
// Absent on an older host, so a missing field is simply "not read-only".
const [schemaTooNew, setSchemaTooNew] = useState<AgentCatalogSchemaTooNew | null>(null)
const [retrying, setRetrying] = useState(false)
const [recoveryOpen, setRecoveryOpen] = useState(false)
@@ -17,8 +25,10 @@ export function DataRecoveryMigrationNotice() {
try {
const status = await window.api.dataRecovery?.migrationStatus()
setMigrationError(status?.agentCatalogMigrationError ?? null)
setSchemaTooNew(status?.agentCatalogSchemaTooNew ?? null)
} catch {
setMigrationError(null)
setSchemaTooNew(null)
}
}, [])
@@ -26,6 +36,31 @@ export function DataRecoveryMigrationNotice() {
void refresh()
}, [refresh])
// Read-only wins over the retryable block: no retry can clear a newer schema,
// so offering one would only fail repeatedly.
if (schemaTooNew) {
const versions = agentCatalogSchemaTooNewVersions(schemaTooNew)
return (
<div
role="alert"
className="flex flex-wrap items-start gap-2 border-b border-destructive/40 bg-destructive/5 px-3 py-2 text-sm"
>
<AlertTriangle className="mt-0.5 size-3.5 shrink-0 text-destructive" />
<div className="min-w-0 flex-1">
<p className="font-medium text-destructive">{agentCatalogSchemaTooNewTitle()}</p>
<p className="text-muted-foreground">{agentCatalogSchemaTooNewMessage()}</p>
{versions ? <p className="text-xs text-muted-foreground">{versions}</p> : null}
</div>
<div className="flex shrink-0 items-center gap-2">
<Button type="button" size="xs" variant="outline" onClick={() => setRecoveryOpen(true)}>
{translate('auto.components.dataRecovery.openDataRecovery', 'Open Data recovery')}
</Button>
</div>
<DataRecoveryDialog open={recoveryOpen} onOpenChange={setRecoveryOpen} />
</div>
)
}
if (migrationError === null) {
return null
}
@@ -0,0 +1,68 @@
import { AgentIcon } from '@/lib/agent-catalog'
import { Badge } from '@/components/ui/badge'
import { cn } from '@/lib/utils'
import { translate } from '@/i18n/i18n'
import type { BuiltInTuiAgent } from '../../../../shared/types'
export const PIN_EXIT_CUSTOM_AGENT_EXAMPLE_NAME = 'gpt-5.6-luna-low'
export const PIN_EXIT_CUSTOM_AGENT_EXAMPLE_COMMAND =
'codex --model gpt-5.6-luna -c model_reasoning_effort="low"'
/** Static picker mock: shows a custom agent as a named row among built-ins. */
export function DataRecoveryPinExitCustomAgentExample() {
return (
<div aria-hidden className="overflow-hidden rounded-md border border-border bg-card text-left">
<ExamplePickerRow agent="claude" label="Claude" />
<ExamplePickerRow
agent="codex"
label={PIN_EXIT_CUSTOM_AGENT_EXAMPLE_NAME}
command={PIN_EXIT_CUSTOM_AGENT_EXAMPLE_COMMAND}
selected
/>
<ExamplePickerRow agent="grok" label="Grok" />
</div>
)
}
function ExamplePickerRow({
agent,
label,
command,
selected = false
}: {
agent: BuiltInTuiAgent
label: string
command?: string
selected?: boolean
}) {
return (
<div
className={cn(
'flex items-start gap-2.5 px-3 py-2',
selected &&
'bg-[color-mix(in_srgb,var(--foreground)_10%,var(--background))] shadow-[inset_0_0_0_1px_color-mix(in_srgb,var(--foreground)_12%,transparent)]',
!selected && 'border-t border-border/60 first:border-t-0'
)}
data-current={selected ? 'true' : undefined}
>
<span className="mt-0.5 inline-flex size-4 shrink-0 items-center justify-center">
<AgentIcon agent={agent} size={16} />
</span>
<div className="min-w-0 flex-1">
<div className="flex items-center gap-2">
<span className="truncate text-sm font-medium text-foreground">{label}</span>
{selected ? (
<Badge variant="outline" className="font-normal">
{translate('auto.components.dataRecovery.pinExitExampleBadge', 'Custom')}
</Badge>
) : null}
</div>
{command ? (
<p className="mt-0.5 font-mono text-xs leading-relaxed text-muted-foreground">
{command}
</p>
) : null}
</div>
</div>
)
}
@@ -2,6 +2,10 @@
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import { cleanup, fireEvent, render, screen } from '@testing-library/react'
import { DataRecoveryPinExitNotice } from './DataRecoveryPinExitNotice'
import {
PIN_EXIT_CUSTOM_AGENT_EXAMPLE_COMMAND,
PIN_EXIT_CUSTOM_AGENT_EXAMPLE_NAME
} from './DataRecoveryPinExitCustomAgentExample'
import { dismissPinExitNotice } from './data-recovery-pin-exit-notice-dismissal'
const PIN = {
@@ -44,7 +48,7 @@ describe('DataRecoveryPinExitNotice', () => {
it('renders nothing when the dataRecovery surface is absent', async () => {
;(window as unknown as { api: unknown }).api = {}
render(<DataRecoveryPinExitNotice />)
await vi.waitFor(() => expect(screen.queryByRole('status')).toBeNull())
await vi.waitFor(() => expect(screen.queryByRole('dialog')).toBeNull())
})
it('renders nothing when migration is blocked', async () => {
@@ -52,31 +56,55 @@ describe('DataRecoveryPinExitNotice', () => {
migrationStatus: vi.fn().mockResolvedValue({ agentCatalogMigrationError: 'disk full' })
})
render(<DataRecoveryPinExitNotice />)
await vi.waitFor(() => expect(screen.queryByRole('status')).toBeNull())
await vi.waitFor(() => expect(screen.queryByRole('dialog')).toBeNull())
})
it('renders nothing when no recovery point exists', async () => {
installApi({ listPoints: vi.fn().mockResolvedValue([]) })
render(<DataRecoveryPinExitNotice />)
await vi.waitFor(() => expect(screen.queryByRole('status')).toBeNull())
await vi.waitFor(() => expect(screen.queryByRole('dialog')).toBeNull())
})
it('renders nothing when the only pin cannot be restored', async () => {
installApi({
listPoints: vi.fn().mockResolvedValue([{ ...PIN, restorable: false }])
})
render(<DataRecoveryPinExitNotice />)
await vi.waitFor(() => expect(screen.queryByRole('dialog')).toBeNull())
})
it('shows exit guidance for a host that omits restorable', async () => {
installApi({
listPoints: vi.fn().mockResolvedValue([{ ...PIN, restorable: undefined }])
})
render(<DataRecoveryPinExitNotice />)
expect(await screen.findByRole('dialog')).toBeTruthy()
})
it('shows exit guidance when a pin exists and migration is healthy', async () => {
installApi()
render(<DataRecoveryPinExitNotice />)
expect(await screen.findByRole('status')).toBeTruthy()
expect(screen.getByText('This profile was upgraded for custom agents')).toBeTruthy()
expect(screen.getByText('Open Data recovery')).toBeTruthy()
expect(await screen.findByRole('dialog')).toBeTruthy()
expect(screen.getByText('Custom agents are now available')).toBeTruthy()
expect(screen.getByText(PIN_EXIT_CUSTOM_AGENT_EXAMPLE_NAME)).toBeTruthy()
expect(screen.getByText(PIN_EXIT_CUSTOM_AGENT_EXAMPLE_COMMAND)).toBeTruthy()
expect(screen.getByText('Before you install an older Orca')).toBeTruthy()
expect(screen.getByText('Restore the data backup. Orca will quit.')).toBeTruthy()
expect(screen.getByText('Then install the older Orca.')).toBeTruthy()
expect(screen.getByText('Restore data backup…')).toBeTruthy()
expect(screen.getByText('Continue')).toBeTruthy()
expect(screen.queryByText('Open Data recovery')).toBeNull()
expect(screen.queryByText(/profile/i)).toBeNull()
})
it('dismisses for this pin and stays hidden after remount', async () => {
installApi()
const { unmount } = render(<DataRecoveryPinExitNotice />)
fireEvent.click(await screen.findByText('Got it'))
await vi.waitFor(() => expect(screen.queryByRole('status')).toBeNull())
fireEvent.click(await screen.findByText('Continue'))
await vi.waitFor(() => expect(screen.queryByRole('dialog')).toBeNull())
unmount()
render(<DataRecoveryPinExitNotice />)
await vi.waitFor(() => expect(screen.queryByRole('status')).toBeNull())
await vi.waitFor(() => expect(screen.queryByRole('dialog')).toBeNull())
})
it('reappears for a newer pin after a prior dismiss', async () => {
@@ -85,15 +113,13 @@ describe('DataRecoveryPinExitNotice', () => {
listPoints: vi.fn().mockResolvedValue([{ ...PIN, createdAtMs: PIN.createdAtMs + 1 }])
})
render(<DataRecoveryPinExitNotice />)
expect(await screen.findByRole('status')).toBeTruthy()
expect(await screen.findByRole('dialog')).toBeTruthy()
})
it('opens Data recovery from the banner', async () => {
it('opens the restore dialog from the data-backup action', async () => {
installApi()
render(<DataRecoveryPinExitNotice />)
fireEvent.click(await screen.findByText('Open Data recovery'))
expect(
await screen.findByText('Before the custom-agents update (agent catalog v1)')
).toBeTruthy()
fireEvent.click(await screen.findByText('Restore data backup…'))
expect(await screen.findByText('Before custom agents')).toBeTruthy()
})
})
@@ -1,25 +1,43 @@
import { useCallback, useEffect, useState } from 'react'
import { Info } from 'lucide-react'
import { useCallback, useEffect, useRef, useState } from 'react'
import { Button } from '@/components/ui/button'
import {
Dialog,
DialogContent,
DialogDescription,
DialogFooter,
DialogHeader,
DialogTitle
} from '@/components/ui/dialog'
import { Separator } from '@/components/ui/separator'
import { translate } from '@/i18n/i18n'
import type { RecoveryPointDto } from '../../../../shared/data-recovery'
import { DataRecoveryDialog } from './DataRecoveryDialog'
import { DataRecoveryPinExitCustomAgentExample } from './DataRecoveryPinExitCustomAgentExample'
import {
dismissPinExitNotice,
isPinExitNoticeDismissed
} from './data-recovery-pin-exit-notice-dismissal'
function preV1Point(points: RecoveryPointDto[]): RecoveryPointDto | null {
return points.find((point) => point.id === 'agent-catalog-pre-v1') ?? null
/** Only a pin that can actually be restored: an unreadable one would send people
* to a downgrade they cannot perform. `restorable` is optional (older hosts omit
* it), so only an explicit false withdraws the dialog. */
function restorablePreV1Point(points: RecoveryPointDto[]): RecoveryPointDto | null {
return (
points.find((point) => point.id === 'agent-catalog-pre-v1' && point.restorable !== false) ??
null
)
}
/** One-shot info banner after a successful agent-catalog pin: tells people how
* to leave for stable. Hidden when migration is blocked (red notice owns that
* state), when no pin exists, on paired web, or after dismiss for this pin. */
/** One-shot dialog after a successful agent-catalog pin. Most people should
* continue; the optional path is how to return to the previous Orca without
* reinstalling over live data. Hidden when migration is blocked (red notice
* owns that state), when no restorable pin exists, on paired web, or after
* dismiss for this pin. */
export function DataRecoveryPinExitNotice() {
const [pin, setPin] = useState<RecoveryPointDto | null>(null)
const [dismissed, setDismissed] = useState(false)
const [recoveryOpen, setRecoveryOpen] = useState(false)
const continueRef = useRef<HTMLButtonElement>(null)
const refresh = useCallback(async () => {
try {
@@ -29,7 +47,7 @@ export function DataRecoveryPinExitNotice() {
return
}
const points = (await window.api.dataRecovery?.listPoints()) ?? []
const next = preV1Point(points)
const next = restorablePreV1Point(points)
setPin(next)
setDismissed(next != null && isPinExitNoticeDismissed(next.createdAtMs))
} catch {
@@ -41,50 +59,114 @@ export function DataRecoveryPinExitNotice() {
void refresh()
}, [refresh])
if (pin === null || dismissed) {
return null
}
const handleDismiss = (): void => {
dismissPinExitNotice(pin.createdAtMs)
if (pin !== null) {
dismissPinExitNotice(pin.createdAtMs)
}
setDismissed(true)
}
const pinDialogOpen = pin !== null && !dismissed && !recoveryOpen
if (!pinDialogOpen && !recoveryOpen) {
return null
}
return (
<div role="status" className="border-b border-border bg-muted/40 py-2 text-sm">
{/* Why: this strip sits under native macOS traffic lights; center the
content block and reserve the same pad on both sides so the title is
not clipped and reads as middle-of-window, not left-edge. */}
<div className="flex items-start justify-center gap-3 px-3">
<div className="titlebar-traffic-light-pad shrink-0" aria-hidden />
<div className="flex min-w-0 max-w-3xl flex-1 flex-wrap items-start justify-center gap-2 sm:flex-nowrap">
<Info className="mt-0.5 size-3.5 shrink-0 text-muted-foreground" />
<div className="min-w-0 flex-1 text-left">
<p className="font-medium">
<>
<Dialog
open={pinDialogOpen}
onOpenChange={(open) => {
// Closing because the restore dialog is on top is temporary.
if (!open && !recoveryOpen) {
handleDismiss()
}
}}
>
<DialogContent
className="max-w-lg gap-5"
onOpenAutoFocus={(event) => {
event.preventDefault()
continueRef.current?.focus()
}}
>
<DialogHeader className="gap-3">
<DialogTitle className="leading-snug">
{translate(
'auto.components.dataRecovery.pinExitTitle',
'This profile was upgraded for custom agents'
'Custom agents are now available'
)}
</p>
<p className="text-muted-foreground">
</DialogTitle>
<DialogDescription className="text-sm leading-relaxed text-foreground">
{translate(
'auto.components.dataRecovery.pinExitBody',
'Orca kept a recovery point so you can return to the previous Orca. Use Data recovery → Prepare downgrade, then install that older version. Do not only reinstall the older app over this profile. Going back discards settings and custom agents saved after the recovery point.'
'auto.components.dataRecovery.pinExitLead',
'A custom agent is saved arguments for a harness like Codex or Claude, picked by name.'
)}
</DialogDescription>
</DialogHeader>
<DataRecoveryPinExitCustomAgentExample />
<p className="text-sm leading-relaxed text-muted-foreground">
{translate(
'auto.components.dataRecovery.pinExitExampleHint',
'Create them in Settings → Agents. Keep working as usual — nothing else is required.'
)}
</p>
<Separator />
<div className="flex flex-col gap-3">
<p className="text-sm font-medium text-foreground">
{translate(
'auto.components.dataRecovery.pinExitRollbackTitle',
'Before you install an older Orca'
)}
</p>
<p className="text-sm leading-relaxed text-muted-foreground">
{translate(
'auto.components.dataRecovery.pinExitRollbackReinstall',
'This version changed the local data format. Restore the previous backup first, or the older app can break.'
)}
</p>
<ol className="list-decimal space-y-1 pl-4 text-sm leading-relaxed text-muted-foreground">
<li>
{translate(
'auto.components.dataRecovery.pinExitRollbackStepRestore',
'Restore the data backup. Orca will quit.'
)}
</li>
<li>
{translate(
'auto.components.dataRecovery.pinExitRollbackStepInstall',
'Then install the older Orca.'
)}
</li>
</ol>
<p className="text-sm leading-relaxed text-muted-foreground">
{translate(
'auto.components.dataRecovery.pinExitRollbackLoss',
'The restore discards changes since this update, including custom agents.'
)}
</p>
<div>
<Button
type="button"
size="sm"
variant="outline"
onClick={() => setRecoveryOpen(true)}
>
{translate('auto.components.dataRecovery.pinExitGoBack', 'Restore data backup…')}
</Button>
</div>
</div>
<div className="flex shrink-0 items-center gap-2">
<Button type="button" size="xs" variant="outline" onClick={() => setRecoveryOpen(true)}>
{translate('auto.components.dataRecovery.openDataRecovery', 'Open Data recovery')}
<DialogFooter>
<Button ref={continueRef} type="button" onClick={handleDismiss}>
{translate('auto.components.dataRecovery.dismissPinExit', 'Continue')}
</Button>
<Button type="button" size="xs" variant="ghost" onClick={handleDismiss}>
{translate('auto.components.dataRecovery.dismissPinExit', 'Got it')}
</Button>
</div>
</div>
<div className="titlebar-traffic-light-pad shrink-0" aria-hidden />
</div>
</DialogFooter>
</DialogContent>
</Dialog>
<DataRecoveryDialog open={recoveryOpen} onOpenChange={setRecoveryOpen} />
</div>
</>
)
}
@@ -34,7 +34,7 @@ export function DataRecoverySettingsRow() {
<span>
{translate(
'auto.components.dataRecovery.settingsRowHint',
'A recovery point can return this profile to the previous Orca (Prepare downgrade, then install that version).'
'You can restore the settings from before this update, then install the previous Orca.'
)}
</span>
<Button type="button" variant="outline" size="xs" onClick={() => setOpen(true)}>
@@ -21,6 +21,7 @@ import {
type SourceControlAiWriteTarget
} from '../../../../shared/source-control-ai-recipe-save'
import { saveSourceControlAiSettings } from '@/lib/agent-catalog-authoring'
import { notifyAgentAuthoringWriteFailure } from '@/lib/agent-authoring-write-failure-toast'
import type {
SourceControlActionRecipe,
SourceControlLaunchActionId
@@ -183,7 +184,7 @@ export function useCheckRunDetailsFixWithAI(args: {
recipe
})
if ('sourceControlAi' in result) {
await saveSourceControlAiSettings(result.sourceControlAi)
notifyAgentAuthoringWriteFailure(await saveSourceControlAiSettings(result.sourceControlAi))
return
}
await updateRepo(result.target.repoId, result.update)
@@ -9,6 +9,7 @@ import {
} from '../../../../shared/commit-message-agent-spec'
import { getAgentCatalog } from '@/lib/agent-catalog'
import { saveCommitMessageAiSettings } from '@/lib/agent-catalog-authoring'
import { notifyAgentAuthoringWriteFailure } from '@/lib/agent-authoring-write-failure-toast'
import { useAppStore } from '@/store'
import {
EMPTY_COMMIT_MESSAGE_AI_SETTINGS,
@@ -89,7 +90,9 @@ export function useAiCommitPrSettings(): AiCommitPrSettingsViewModel {
if (!settings) {
return
}
void saveCommitMessageAiSettings({ ...config, ...patch })
// Why: these cards render in the feature wall and onboarding overlay, neither
// of which has an error slot — toast so a rejected write is not read as saved.
void saveCommitMessageAiSettings({ ...config, ...patch }).then(notifyAgentAuthoringWriteFailure)
}
const toggleAi = (): void => {
@@ -305,6 +305,7 @@ import {
type SourceControlAiWriteTarget
} from '../../../../shared/source-control-ai-recipe-save'
import { saveSourceControlAiSettings } from '@/lib/agent-catalog-authoring'
import { notifyAgentAuthoringWriteFailure } from '@/lib/agent-authoring-write-failure-toast'
import { resolveSourceControlLaunchPlatform } from '@/lib/source-control-launch-platform'
import { getLocalProjectExecutionRuntimeContext } from '@/lib/local-preflight-context'
import { getRuntimeEnvironmentIdForWorktree } from '@/lib/worktree-runtime-owner'
@@ -932,7 +933,7 @@ export default function ChecksPanel(): React.JSX.Element {
recipe
})
if ('sourceControlAi' in result) {
await saveSourceControlAiSettings(result.sourceControlAi)
notifyAgentAuthoringWriteFailure(await saveSourceControlAiSettings(result.sourceControlAi))
return
}
await updateRepo(result.target.repoId, result.update)
@@ -313,6 +313,100 @@ describe('runSourceControlAgentActionStart', () => {
expect(mocks.launchAgentInNewTab).toHaveBeenCalledTimes(1)
})
// Why: the host reads agentArgs from the stored recipe the locator names, so a
// "Don't save" launch must send the edited args in place of that stale snapshot.
it('sends the edited args with the locator when they diverge from the stored args', async () => {
mocks.launchAgentInNewTab.mockReturnValue({ tabId: 'tab-1', pasteDraftAfterLaunch: true })
const settings = {
sourceControlAi: {
actions: { resolveComments: { agentId: 'codex', agentArgs: '--model gpt-4' } }
}
} as Parameters<typeof runSourceControlAgentActionStart>[0]['settings']
await expect(
runSourceControlAgentActionStart(buildArgs({ saveTargetValue: 'none', settings }))
).resolves.toBe(true)
expect(mocks.onSaveAgentDefault).not.toHaveBeenCalled()
expect(mocks.launchAgentInNewTab.mock.calls[0]?.[0]).toMatchObject({
sourceRecord: { owner: 'source-control-recipe', id: 'resolveComments' },
unsavedAgentArgs: '--model gpt-5'
})
})
it('keeps the recipe locator when the stored args already match the launch', async () => {
mocks.launchAgentInNewTab.mockReturnValue({ tabId: 'tab-1', pasteDraftAfterLaunch: true })
const settings = {
sourceControlAi: {
actions: { resolveComments: { agentId: 'codex', agentArgs: '--model gpt-5' } }
}
} as Parameters<typeof runSourceControlAgentActionStart>[0]['settings']
await expect(
runSourceControlAgentActionStart(buildArgs({ saveTargetValue: 'none', settings }))
).resolves.toBe(true)
expect(mocks.launchAgentInNewTab.mock.calls[0]?.[0]).toMatchObject({
sourceRecord: { owner: 'source-control-recipe', id: 'resolveComments' }
})
expect(mocks.launchAgentInNewTab.mock.calls[0]?.[0]).not.toHaveProperty('unsavedAgentArgs')
})
it('keeps the recipe locator once the edited args are persisted', async () => {
mocks.launchAgentInNewTab.mockReturnValue({ tabId: 'tab-1', pasteDraftAfterLaunch: true })
mocks.onSaveAgentDefault.mockResolvedValue(undefined)
await expect(
runSourceControlAgentActionStart(buildArgs({ saveTargetValue: 'global' }))
).resolves.toBe(true)
expect(mocks.launchAgentInNewTab.mock.calls[0]?.[0]).toMatchObject({
sourceRecord: { owner: 'source-control-recipe', id: 'resolveComments' }
})
})
// Why: a repo agentArgs override shadows the global recipe, so saving globally
// leaves the host resolving the repo's stale args from the locator.
it('sends the edited args when a global save is shadowed by a repo args override', async () => {
mocks.launchAgentInNewTab.mockReturnValue({ tabId: 'tab-1', pasteDraftAfterLaunch: true })
mocks.onSaveAgentDefault.mockResolvedValue(undefined)
const repo = {
id: 'repo-1',
sourceControlAi: { actionOverrides: { resolveComments: { agentArgs: '--model gpt-4' } } }
} as unknown as Parameters<typeof runSourceControlAgentActionStart>[0]['repo']
await expect(
runSourceControlAgentActionStart(
buildArgs({ saveTargetValue: 'global', repoId: 'repo-1', repo })
)
).resolves.toBe(true)
expect(mocks.onSaveAgentDefault).toHaveBeenCalledTimes(1)
expect(mocks.launchAgentInNewTab.mock.calls[0]?.[0]).toMatchObject({
sourceRecord: { owner: 'source-control-recipe', id: 'resolveComments' },
unsavedAgentArgs: '--model gpt-5'
})
})
it('keeps the recipe locator when the save lands on the shadowing repo override', async () => {
mocks.launchAgentInNewTab.mockReturnValue({ tabId: 'tab-1', pasteDraftAfterLaunch: true })
mocks.onSaveAgentDefault.mockResolvedValue(undefined)
const repo = {
id: 'repo-1',
sourceControlAi: { actionOverrides: { resolveComments: { agentArgs: '--model gpt-4' } } }
} as unknown as Parameters<typeof runSourceControlAgentActionStart>[0]['repo']
await expect(
runSourceControlAgentActionStart(
buildArgs({ saveTargetValue: 'repo', repoId: 'repo-1', repo })
)
).resolves.toBe(true)
expect(mocks.launchAgentInNewTab.mock.calls[0]?.[0]).toMatchObject({
sourceRecord: { owner: 'source-control-recipe', id: 'resolveComments' }
})
})
it('keeps injected onStart successes immediate', async () => {
const onStart = vi.fn().mockResolvedValue(true)
@@ -11,7 +11,10 @@ import type {
import type { SourceControlAiWriteTarget } from '../../../../shared/source-control-ai-recipe-save'
import { toast } from 'sonner'
import { translate } from '@/i18n/i18n'
import { resolveSourceControlActionRecipe } from '../../../../shared/source-control-ai'
import {
normalizeRepoSourceControlAiOverrides,
resolveSourceControlActionRecipe
} from '../../../../shared/source-control-ai'
import { sourceControlActionRecipeMatchesTarget } from './source-control-action-recipe-match'
import { resolveSourceControlAgentSaveTarget } from './source-control-agent-action-dialog-support'
@@ -94,17 +97,22 @@ export async function runSourceControlAgentActionStart({
repo
})
)
// Why: the host resolves agentArgs from the owner locator, so claiming it is
// only correct while the persisted recipe still carries the edited args — a
// "Don't save" launch of edited args would otherwise run the stale snapshot.
// Why: the host resolves agentArgs from the owner locator, so the launch needs
// the edited args sent alongside it unless the EFFECTIVE recipe (repo override
// shadowing global) already carries them.
let launchArgsPersisted =
launchRecipeAlreadySaved ||
(resolveSourceControlActionRecipe({ settings, repo, actionId }).agentArgs?.trim() ?? '') ===
agentArgs.trim()
agentArgs.trim()
// Why: a repo-level agentArgs override shadows the global recipe per field, so
// a global save never becomes what the host resolves for this repo.
const repoOverridesActionArgs =
normalizeRepoSourceControlAiOverrides(repo?.sourceControlAi)?.actionOverrides?.[actionId]
?.agentArgs !== undefined
if (saveTarget && onSaveAgentDefault && !launchRecipeAlreadySaved) {
try {
await onSaveAgentDefault(saveTarget, actionId, launchRecipe)
launchArgsPersisted = true
launchArgsPersisted =
launchArgsPersisted || saveTarget.type === 'repo' || !repoOverridesActionArgs
} catch (error) {
console.error('saving the agent recipe before launch failed', error)
toast.error(
@@ -142,13 +150,12 @@ export async function runSourceControlAgentActionStart({
worktreeId,
groupId: groupId ?? worktreeId,
prompt: trimmedCommandInput,
// Why: the host resolves this recipe's stored agentArgs from the owner
// locator; the client no longer sends assembled args on the launch path.
// Unsaved edits are not in that stored recipe, so the locator is omitted
// rather than launching the stale persisted args.
...(launchArgsPersisted
? { sourceRecord: { owner: 'source-control-recipe' as const, id: actionId } }
: {}),
sourceRecord: { owner: 'source-control-recipe' as const, id: actionId },
// Why: the host resolves the recipe's stored agentArgs from the locator, so
// a "Don't save" launch must send the edited args for the host to use in
// their place — otherwise it runs the stale persisted snapshot. Best-effort:
// a host predating the field still launches the stored args.
...(launchArgsPersisted ? {} : { unsavedAgentArgs: agentArgs.trim() }),
promptDelivery,
launchPlatform,
launchSource
@@ -0,0 +1,77 @@
import { describe, expect, it } from 'vitest'
import type {
LocalAgentCatalogSnapshot,
LocalCustomTuiAgent
} from '../../../../shared/agent-catalog-snapshot'
import type { CustomTuiAgentId, TuiAgent } from '../../../../shared/types'
import { buildSourceControlAgentActionOptions } from './source-control-agent-action-agent-options'
const CUSTOM_CODEX = 'custom-agent:codex:11111111-1111-4111-8111-111111111111' as CustomTuiAgentId
function readyCustom(): LocalCustomTuiAgent {
return {
status: 'ready',
definition: {
id: CUSTOM_CODEX,
baseAgent: 'codex',
label: 'Model-specific Codex',
args: '--model custom-model',
syncEnv: false,
commandOverride: '/opt/bin/codex'
},
envSummary: { entryCount: 0, bytes: 0 },
availabilityReason: 'configured-executable'
}
}
function snapshot(): LocalAgentCatalogSnapshot {
return { customAgents: [readyCustom()] } as LocalAgentCatalogSnapshot
}
describe('buildSourceControlAgentActionOptions', () => {
it('offers named custom agents alongside the detected built-ins', () => {
const options = buildSourceControlAgentActionOptions({
enabledDetectedAgents: ['claude', 'codex'],
disabledAgents: [],
localAgentCatalog: snapshot(),
selectedAgent: null
})
expect(options.map((option) => option.id)).toEqual(['claude', 'codex', CUSTOM_CODEX])
expect(options.at(-1)).toMatchObject({ label: 'Model-specific Codex', baseAgent: 'codex' })
})
it('drops a disabled custom agent from the picker', () => {
const options = buildSourceControlAgentActionOptions({
enabledDetectedAgents: ['codex'],
disabledAgents: [CUSTOM_CODEX],
localAgentCatalog: snapshot(),
selectedAgent: null
})
expect(options.map((option) => option.id)).toEqual(['codex'])
})
it('keeps a saved custom selection listed with its real label when it is unavailable', () => {
const options = buildSourceControlAgentActionOptions({
enabledDetectedAgents: ['claude'],
disabledAgents: [CUSTOM_CODEX],
localAgentCatalog: snapshot(),
selectedAgent: CUSTOM_CODEX
})
expect(options.map((option) => option.id)).toEqual(['claude', CUSTOM_CODEX])
expect(options.at(-1)?.label).toBe('Model-specific Codex')
})
it('keeps listing an undetected built-in selection', () => {
const options = buildSourceControlAgentActionOptions({
enabledDetectedAgents: ['claude'],
disabledAgents: undefined,
localAgentCatalog: null,
selectedAgent: 'cursor' as TuiAgent
})
expect(options.map((option) => option.id)).toEqual(['claude', 'cursor'])
})
})

Some files were not shown because too many files have changed in this diff Show More