Files
orca/src/cli/skills.test.ts
T
Jinwoo Hong 06a607a1d7 feat(orchestration): make multi-agent workflows durable (#16904)
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->

| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 225 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$​21666 | $\color{#cf222e}{\Huge{\mathbf{−}}}$​2820 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$​18846 |
| Prod | 348 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$​17107 | $\color{#cf222e}{\Huge{\mathbf{−}}}$​4706 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$​12401 |

<!-- /orca-pr-loc -->

## ELI5

Orca now treats orchestration like a durable control plane instead of inferring success from terminal keystrokes. Agents can tell whether a prompt was accepted or a turn started, replay an ambiguous request without sending twice, and recover coordinator mail after a crash. Completed workers can be inspected, released, or retained, and their panes no longer auto-resume as if the work were still running.

## What changed

- **Run receipts** from `run-create/use/current/show/list` are the row without routing plumbing (`home_database`, `coordinator_pane_key`) and without the duplicate `binding` object.
- **`terminal send` receipts are honest and idempotent.** `input_accepted` and `turn_started` are the only stages; `--wait-submit` observes without resending; `--retry-request <uuid>` replays the exact request against the same process incarnation. A transport timeout keeps the retry ID; only a different runtime answering strips it. Value-less or non-UUID `--retry-request` is rejected on the CLI and the SSH shim.
- **Mailbox delivery is committed before wakeup.** Pointer writes are staged in the DB before any PTY byte, replayed once after restart, and never emit a naked Enter. The watermark that parks concurrent deliveries is released with the DB reservation. Restart rescans pointer-pending and `dispatch:` mailboxes.
- **Lifecycle is a guarded transition graph** (`lifecycle-transition.ts`) with a table-driven test over every caller edge. Task reopen/overturn stays in the public contract. A PTY exit during `worker-stop` is the stop succeeding, not a failure.
- **Worker lifecycle CLI:** `worker-start` (`--spec` creates Task + attempt in one call), `worker-show`, `worker-read` (provider transcript first, bounded terminal fallback with a typed reason, local/WSL/SSH), `worker-stop`, `worker-abandon`, `worker-release`, `worker-retain`, `worker-list` (rowid-fenced pagination, fleet liveness, `attention`, literal `nextAction`).
- **Release is an explicit ownership table** (`decideWorkerTerminalRelease`): only an `owned` resource can be settled, the archive is mandatory where reachable, and an owner whose process is proven exited can always get out of `retained` via `archive_status: unavailable`. User-taken-over, external, and transferred panes stay retained.
- **Settled-worker resume fence** (folds in #17651): a settled dispatch whose pane is still open is fenced at settlement, on stop/abandon/exit, and at startup; lifted on release, retain, takeover, and pane reuse.
- **Liveness is `live` / `unverifiable` / `exited` only**, from execution-host evidence. Fleet projection reads the evidence clock, not the relay delivery clock. A host-certified exit outranks the worker's settled state. `unverifiable` never authorizes stop, abandon, retry, or release, in code or in the guide.
- **Federation:** structured reads negotiate by `method_not_found` so every shipped host keeps transcript-first output; exited remote workers are closed before being reported closed; epoch fencing holds across peer restart, downgrade, and pairing rotation; no per-second forced capability probe.
- **Schema v35:** repairs databases stamped v34 by the pre-fix branch (mailbox_handle default, index predicates), drops the write-only `lifecycle_transition_receipts` ledger and five never-read v31 identity columns.
- **Schema v36:** `dispatch:<id>` mailboxes get a real consumer generation on `dispatch_contexts` and `remote_dispatch_attachments`, bumped and fenced in the same transaction on every re-attach (manual inject, worker-start, federated attach). A stale worker whose Dispatch moved to another process now gets `consumer_fenced` instead of silently acking the new worker's Delivery. Run mailboxes already worked this way.
- **Schema v37:** `dispatch_contexts` records its creator (`creator_handle`, `creator_pane_key`), so a coordinator's context-only self-dispatch is bookkeeping rather than a nesting parent; before this, one self-dispatch made every later `worker-start` from that coordinator fail the depth cap. Pre-v37 rows keep counting (fails closed).
- **Dispatch-mailbox ownership is checked, not inferred.** A `check` from a process whose pane no longer holds the Dispatch, or whose last Attempt was abandoned/failed and moved to another terminal, gets `consumer_fenced` instead of an empty inbox that reads as "no mail yet". `--peek`/`--all` stay readable. A paneless caller still gets `stable_pane_required` with the rebind recovery.
- **Liveness certification is stricter:** a `process_exited` stage whose termination reason is `unknown` (a stop that was issued but never observed) projects `unverifiable`, not `exited`. Federated `worker-show` carries the execution host's verdict and host kind instead of a local guess. A live, ready worker with nothing pending has `nextAction: none` rather than pointing at the `worker-show` that produced it.
- **Wire:** `workerShow` keeps `dispatch.task_id` next to `taskId` for shipped CLIs. `ask --json` uses the standard `{ok, result}` envelope like every sibling verb.
- **Migration start-version detection** treats the two v32 recovery columns as versioned. Before this, every shipped database stamped below 32 resolved to the v6 floor and replayed the whole chain (the v23 backfill synthesized 68 phantom retained workers on a real v30 profile). Verified on a copy of a real 62 MB v30 profile: starts at 30, no row delta, integrity ok, 11 ms.
- **Skill guide** rewritten as a ≤200-line kernel plus seven references, to the outcome-first standard (Result / Done / Safe failure first, conditions not case lists, one done bar, references loaded at the point of use). The canonical loop uses `worker-start --spec`, names `worker-list` for completion accounting, documents `--retry-request` / `request-show` / `--wait-submit`, and requires positive evidence before any stall action. The other seven guides get the same treatment in #18724, split out so this PR stays orchestration-only.
- **`rpc/methods/orchestration-*`** (126 flat files) regrouped into `orchestration/{worker,federation,messaging,runs,gates}/`.

## Why

User reports showed the same boundary failures: false `agent_prompt_stalled` causing duplicate sends (#15180), coordinators unable to trust screen scrapes, cold-parked terminals receiving a pointer without the submit, settled workers accumulating as live tabs and auto-resuming after restart, and no way to tell a stalled worker from a working one.

## Linked issues

Fixes #15180. Fixes #17935 (orchestration skill description is 866 characters; a guard now caps every bundled skill at 1,024). Supersedes #17651 (fence folded in). Advances #16660, #16522, #14907, #13047.

## Review record

This PR was reviewed adversarially after revival: eight independent lenses (lifecycle, mailbox, send, worker, federation, transcript, complexity, live ergonomics), each required to prove findings with a failing test. That produced 16 proven blockers, all fixed with red-then-green regression tests, followed by two re-review rounds and a third fix wave that caught 3 regressions introduced by the fixes and 7 fixes that missed their target; all closed. A final pass (five lenses incl. a live built-runtime smoke, then a re-review of the fix wave) found and fixed seven more, chiefly the stale-worker mailbox steal, the self-dispatch depth wedge, and the unproven-exit certification. Three independent Codex (gpt-6-astra) passes followed: the first found nothing new, the second found and fixed 3 defects (task-status reachability, WSL-local host classification, peer-capability epoch), the third found and fixed 6 (production PTY controller never installed settled writes, ambiguous in-flight pointer failures allowed duplicate replay, SSH/relay deadlines cut off a valid `--wait-submit`, stop-vs-exit race during inspection, and two release-recovery paths for vanished or exited terminals). The full record (findings, proof tests, triage, declines with reasons) is archived outside the repo.

**Rework after the live smoke.** A first live cross-host run on the shipped adhoc build (this Mac, a paired Windows host on the same build, a paired Mac on 1.4.195, and an SSH host) found a P1: a running local worker read `unverifiable`/`missing_status` because the fleet snapshot rows lacked the terminal handle the matcher keyed on. A 59-row failure table over every bug fixed during review showed the same two classes recurring: a fact dropped in transit through optional fields, and two authorities for one fact. Two blind designs (Opus, Codex) converged on the same mechanisms, and the scoped tranches landed here with red-then-green seam tests from the real producer to the real consumer, faults injected only at the transport or hook-ingest boundary:

- **Settlement (data-loss class):** one three-valued `WriteSettlement` (`accepted | refused{reason} | unverifiable{reason, bytesHandedToTransport}`) from the SSH multiplexer through daemon client, providers, controller, to pointer staging. No boolean, no rejection-as-third-state. The two silent degrades that fabricated a handoff are deleted; a provider that cannot settle refuses before any effect. Pointer text and Enter share the contract; a partial flush is `unverifiable`, never `refused`.
- **Evidence identity (false-liveness class):** fleet agent-status evidence is a tagged union (`binding: worker | pane | unresolved{reason}`, `clock: observed | delivery`) minted once at ingest, so a hook row captured on one process incarnation can never bind to a later dispatch on the same pane. The matcher's `!worker.paneKey ||` defaults are gone. One host-scope parser replaces two.
- **Small pre-merge items:** `capability_unsupported` from an old peer is no longer relabelled `host_unavailable`; a producer census test asserts every agent-status consumer path projects a pane-only hook row as `live`.

Two ergonomics defects the second live run surfaced on a real database are fixed here too: a pre-v3 dispatch already marked `completed` projected as `outcome_unknown` / `requiresAction: true` forever (three copies of the outcome ladder disagreed on legacy rows; now one resolver, legacy `completed` reads `succeeded` with nothing to act on, legacy `failed` stays actionable on the failure), and an unscoped `worker-list` enumerated the entire database (now defaults to the Run bound to the calling terminal, `--run` overrides, and the receipt's additive `scope` field says which).

A third live round on the shipped adhoc build of `b082443e1f` (same four hosts) plus an unscripted run in the user's own prompt style (a plain Claude Code shell, `/orchestration`, three workers, zero errors, bound-Run default confirmed) found two more branch defects, fixed with red-then-green tests: a worker freshly started on a paired server projected `unverifiable`/`host_indeterminate` with `requiresAction` for ~3 minutes, including after its own `worker_done`, because the host's federation observation returned `missing_liveness_verdict` for any PTY the liveness register had not yet swept (the host now reads a connected pane it owns locally as `live`; disconnected or SSH-scoped panes stay `unverifiable`); and six pre-v3 completed rows still carried an `input` category because settling through the task-status path or `failDispatch` never closed the Dispatch's pending question threads (both paths close them now, and schema v38 closes threads already pending on settled rows). The guide's `worker-start` examples now show `--model sonnet`, since an omitted model inherits the launcher's default.

A Codex adversarial pass on the tranche diff found one real design hole (identity minted at read time instead of ingest, now closed) and two daemon settlement paths that threw instead of settling (fixed). Two `@ts-nocheck` runtime mixins on these paths were extracted into checked modules; the repo-wide `@ts-nocheck` count is unchanged at 171.

Deletions during review: ~1,900 lines (write-only ledger, unread columns, dead v1 archive path, test harnesses shipped in prod, duplicated liveness and state-machine copies, self-capability checks that were compile-time true).

## Testing

- `pnpm typecheck:tsc:node|cli|web` clean
- `pnpm run check:code-quality:changed` 0 findings; `check:react-doctor:changed` 0
- `pnpm verify:bundled-skill-guides`, `verify:skill-bundle-manifest`
- full `pnpm test` on the integrated head: 72,332 pass / 292 skipped; the only failures were three non-PR files (two zsh live-shell suites hit a node-pty spawn-helper ENOENT while a concurrent native rebuild ran, 44/44 in isolation; `release-checkout.unit.test.ts` is a known 30 s load timeout that passes in isolation on `origin/main` too).
- CI on 70b4811267 (rerun, pre-Codex): the only reds are five SSH e2e specs plus `terminal-send-agent-prompt-submit:198`, each shown failing identically on main (main's E2E workflow is red on its last 40 runs). The terminal-send spec is root-caused and fixed separately in #18707. The Windows hook-service flake (#17721) and the federation load flake did not recur.
- Skills: `pnpm exec vitest run` over the skill gate files plus `src/cli`, `config/scripts`, `src/main/skills` pass; live smoke on the built CLI of `skills get orchestration` and `--full` (7 references).
- live headless runtime (`orca-dev serve`, isolated profile): canonical loop, stop, release, archive read, retry rejection, stale-handle check, SIGKILL-and-replay all verified with receipts
- Live cross-host smoke on the shipped adhoc build of `0d465e7931` (this Mac and a paired Windows host on the build, a paired Mac left on 1.4.195, an SSH host): local, paired-new, paired-old and SSH loops all settle; running workers read `live` on every host and `exited` after release; the old peer reads `capability_unsupported` and refuses release honestly. Injected 10 s relay stall with a send in flight: delivered exactly once after recovery, zero duplicates. Every liveness field across 104 receipts is only `live` / `unverifiable` / `exited`.
- Final live cross-host smoke on the shipped adhoc build of `b082443e1f` (same hosts): every loop settles; 942 of 948 legacy completed rows read settled with `requiresAction: false` before the question-thread fix and all of them after; `worker-list` scope reads `bound` / `flag` / `all` correctly; 122 JSON receipts carry only `live` / `unverifiable` / `exited`. Unscripted prompt-style run: clean.
- Confirmation smoke on the shipped adhoc build of `2da076d4e9` (this Mac and the paired Windows host, both updated): a freshly started Windows worker reads `live` on the first fleet poll and on all 20 that follow, with no `host_indeterminate` at any point, and `exited` after release; all 948 legacy completed rows read `requiresAction: false` with `nextAction: none` after schema v38; every verdict across 60 receipts is `live` / `unverifiable` / `exited`.
- Not physically exercised: WSL hosts, the renderer notification bell (headless has no renderer), same-session fence via a real pane close (renderer-only state), restart mid-delivery on a real app (covered by e2e only).

## Notes

- Remote-wire additions are optional fields or `method_not_found`-negotiated methods; one new Electron-only IPC channel (`agentStatus:legacyWorkerTerminalResumeFence`) never crosses the wire.
- SSH contact loss remains `unverifiable`; the execution host stays authoritative.
- Intentional wire projection change: an SSH host scope with an empty `targetId` now projects host id `ssh` instead of an empty string (remote-wire-compatibility rule 3, old clients decode the same field). A fleet pane key without a terminal handle is now `unidentifiable` rather than matched by pane key alone.
- Found live but pre-existing on main, filed separately: a relay daemon-start collision during transport loss rewrites the endpoint credential and wedges the surviving relay (host needs a manual kill); `terminal create` on a reconnecting SSH host reports an opaque `No PTY provider for connection`; `terminal list` reports `orphaned:false` and `terminal close` reports `ptyKilled:true` for a pane whose relay is gone (orchestration's own projection reads `unverifiable` correctly at the same moment).
- Downgrade after this PR is not a supported path: main opens a v37 database and early-returns (its inserts still work against the v36/v37 defaulted columns), but its one-outstanding-Delivery-per-Run index is a no-op against the branch's mailbox-scoped index of the same name.
- Known follow-ups (not blockers): `worker-list` materializes every dispatch row per call; a positive "agent absent" signal distinct from PTY liveness is a product decision left open (a headless fake agent never reaches `live`, so its `nextAction` stays `inspect`); a context-only self-dispatch still lists as `role: worker` in `worker-list`; `dispatch` task-not-found / task-not-ready / inject-rejected still surface as `runtime_error`; task and inbox receipts still expose raw row columns. Deferred skill product decisions live on #18724.
2026-09-06 14:34:03 -04:00

999 lines
36 KiB
TypeScript

import { chmodSync, mkdtempSync, writeFileSync } from 'node:fs'
import { tmpdir } from 'node:os'
import { EventEmitter } from 'node:events'
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import { delimiter, join } from 'node:path'
import type * as CodexCliCommandModule from '../shared/node-cli-command-resolution'
import { WINDOWS_BATCH_UNSAFE_CHARACTERS_LABEL } from '../shared/windows-batch-spawn'
const {
detectCommandsMock,
guideModuleLoadMock,
resolveCliCommandMock,
runtimeClientConstructorMock,
spawnMock
} = vi.hoisted(() => ({
detectCommandsMock: vi.fn(() => new Set<string>(['claude'])),
guideModuleLoadMock: vi.fn(),
resolveCliCommandMock: vi.fn(() => 'npx'),
runtimeClientConstructorMock: vi.fn(),
spawnMock: vi.fn()
}))
// Why: agent detection probes the real machine, so pin it or every install
// assertion depends on what the test runner happens to have installed.
vi.mock('../shared/local-agent-install-dir-detection', () => ({
detectCommandsInInstallDirs: detectCommandsMock
}))
// Why: override only the npx lookup so the real Windows .cmd rail still runs.
vi.mock('../shared/node-cli-command-resolution', async (importOriginal) => ({
...(await importOriginal<typeof CodexCliCommandModule>()),
resolveCliCommand: resolveCliCommandMock
}))
vi.mock('node:child_process', () => ({
spawn: spawnMock
}))
vi.mock('./bundled-skill-guides.js', () => {
guideModuleLoadMock()
return {
BUNDLED_SKILL_GUIDES: [
{
name: 'zeta',
description: 'Use when zeta work\nspans lines.',
markdown: '# Zeta\n',
fullMarkdown: '# Zeta\n\n## References\n\nZeta reference.\n',
aliases: []
},
{
name: 'alpha',
description: 'Use when alpha work is needed.',
markdown: '# Alpha\n\nShort.\n',
fullMarkdown: '# Alpha\n\nShort.\n\n## References\n\nFull.\n',
aliases: ['legacy-alpha']
},
{
name: 'gamma',
description:
'Use when gamma work spans several sentences describing exactly how a ' +
'coding agent should decide whether gamma applies to the current task at hand.',
markdown: '# Gamma\n',
fullMarkdown: '# Gamma\n\n## References\n\nGamma reference.\n',
aliases: []
}
]
}
})
vi.mock('./runtime-client', async () => {
// Why: re-export the REAL error classes rather than redefining them. format.ts
// narrows with `instanceof` against ./runtime/types, so a look-alike class
// here would make every CLI error fall through to the generic `runtime_error`
// shape — mirroring the barrel keeps the mock faithful to production.
const { RuntimeClientError, RuntimeRpcFailureError } = await import('./runtime/types.js')
class RuntimeClient {
constructor() {
runtimeClientConstructorMock()
}
}
return {
RuntimeClient,
RuntimeClientError,
RuntimeRpcFailureError,
serveOrcaApp: vi.fn(),
getDefaultUserDataPath: vi.fn(() => '/tmp/orca-user-data')
}
})
import { dispatch } from './dispatch'
import { main } from './index'
describe('orca skills CLI', () => {
beforeEach(() => {
vi.restoreAllMocks()
runtimeClientConstructorMock.mockClear()
resolveCliCommandMock.mockReset()
resolveCliCommandMock.mockReturnValue('npx')
detectCommandsMock.mockReset()
detectCommandsMock.mockReturnValue(new Set<string>(['claude']))
spawnMock.mockReset()
process.exitCode = undefined
})
afterEach(() => {
vi.restoreAllMocks()
vi.unstubAllEnvs()
})
it('keeps the bundled table off the eager command-registry path', async () => {
vi.spyOn(console, 'log').mockImplementation(() => {})
expect(guideModuleLoadMock).not.toHaveBeenCalled()
await main(['status', '--help'], '/tmp/repo')
expect(guideModuleLoadMock).not.toHaveBeenCalled()
})
it('dispatches an alias locally and emits the exact Markdown', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await dispatch(['skills', 'get'], {
flags: new Map([['topic', 'legacy-alpha']]),
get client(): never {
throw new Error('skills get accessed RuntimeClient')
},
cwd: '/tmp/repo',
json: false
})
expect(stdoutText(stdoutSpy)).toBe('# Alpha\n\nShort.\n')
})
it('lists canonical topics deterministically without constructing RuntimeClient', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(['skills', 'list'], '/tmp/repo')
expect(stdoutText(stdoutSpy)).toBe(
'alpha: Use when alpha work is needed.\n' +
'gamma: Use when gamma work spans several sentences describing exactly how a ' +
'coding agent should decide whether gamma applies to the current task at hand.\n' +
'zeta: Use when zeta work spans lines.\n'
)
expect(runtimeClientConstructorMock).not.toHaveBeenCalled()
})
it('emits full Markdown for --full', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(['skills', 'get', 'alpha', '--full'], '/tmp/repo')
expect(stdoutText(stdoutSpy)).toBe('# Alpha\n\nShort.\n\n## References\n\nFull.\n')
})
it('supports the canonical single-item show verb as an alias', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(['skills', 'show', 'alpha'], '/tmp/repo')
expect(stdoutText(stdoutSpy)).toBe('# Alpha\n\nShort.\n')
})
it('gives list --json a stable canonical schema', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(['skills', 'list', '--json'], '/tmp/repo')
expect(stdoutText(stdoutSpy)).toBe(
`${JSON.stringify(
{
topics: [
{ name: 'alpha', description: 'Use when alpha work is needed.' },
{
name: 'gamma',
description:
'Use when gamma work spans several sentences describing exactly how a ' +
'coding agent should decide whether gamma applies to the current task at hand.'
},
{ name: 'zeta', description: 'Use when zeta work spans lines.' }
]
},
null,
2
)}\n`
)
})
it('gives alias get --json the canonical name, selection, and Markdown', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(['skills', 'get', 'legacy-alpha', '--full', '--json'], '/tmp/repo')
expect(stdoutText(stdoutSpy)).toBe(
`${JSON.stringify(
{
name: 'alpha',
full: true,
markdown: '# Alpha\n\nShort.\n\n## References\n\nFull.\n'
},
null,
2
)}\n`
)
})
it('shows leaf, group, and root help for skills', async () => {
const logSpy = vi.spyOn(console, 'log').mockImplementation(() => {})
await main(['skills', 'get', '--help'], '/tmp/repo')
await main(['skills', '--help'], '/tmp/repo')
await main(['--help'], '/tmp/repo')
expect(String(logSpy.mock.calls[0]?.[0])).toContain(
'Usage: orca skills get <topic> [--full | --reference <name>] [--json]'
)
expect(String(logSpy.mock.calls[1]?.[0])).toContain(
'Commands:\n installed List installed skill selectors'
)
expect(String(logSpy.mock.calls[1]?.[0])).toContain(
'get Print a version-matched skill guide'
)
expect(String(logSpy.mock.calls[1]?.[0])).toContain(
'install Install bundled Orca skills'
)
expect(String(logSpy.mock.calls[1]?.[0])).toContain(
'update Update already-installed Orca skills'
)
expect(String(logSpy.mock.calls[2]?.[0])).toContain('Skills:\n skills installed')
expect(String(logSpy.mock.calls[2]?.[0])).toContain('skills update')
expect(runtimeClientConstructorMock).not.toHaveBeenCalled()
})
it('returns a nonzero error with all canonical topics for an unknown topic', async () => {
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
await main(['skills', 'get', 'missing'], '/tmp/repo')
expect(process.exitCode).toBe(1)
expect(errorSpy).toHaveBeenCalledWith(
'Unknown skill topic "missing". Available topics: alpha, gamma, zeta'
)
expect(runtimeClientConstructorMock).not.toHaveBeenCalled()
})
it('lists installable skills when no --skill/--all is given', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(['skills', 'install'], '/tmp/repo')
expect(stdoutText(stdoutSpy)).toBe(
[
'Choose one or more skills to install:',
' alpha',
' gamma',
' zeta',
'',
'Usage: orca skills install --skill <name> [--skill <name> ...]',
' or: orca skills install --all',
''
].join('\n')
)
expect(spawnMock).not.toHaveBeenCalled()
expect(runtimeClientConstructorMock).not.toHaveBeenCalled()
})
it('gives install --json (no selection) a stable schema', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(['skills', 'install', '--json'], '/tmp/repo')
expect(stdoutText(stdoutSpy)).toBe(
`${JSON.stringify({ availableSkills: ['alpha', 'gamma', 'zeta'] }, null, 2)}\n`
)
})
it('rejects combining --all with --skill', async () => {
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
await main(['skills', 'install', '--all', '--skill', 'alpha'], '/tmp/repo')
expect(process.exitCode).toBe(1)
expect(errorSpy).toHaveBeenCalledWith('Use either --all or --skill, not both.')
expect(spawnMock).not.toHaveBeenCalled()
})
it('rejects an unknown --skill name', async () => {
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
await main(['skills', 'install', '--skill', 'missing'], '/tmp/repo')
expect(process.exitCode).toBe(1)
expect(errorSpy).toHaveBeenCalledWith(
'Unknown skill "missing". Available skills: alpha, gamma, zeta'
)
expect(spawnMock).not.toHaveBeenCalled()
})
it('rejects --skill without a value', async () => {
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
await main(['skills', 'install', '--skill'], '/tmp/repo')
expect(process.exitCode).toBe(1)
expect(errorSpy).toHaveBeenCalledWith('--skill requires a value; it was passed with none.')
expect(spawnMock).not.toHaveBeenCalled()
})
it('rejects --json for a real (non-dry-run) install', async () => {
const logSpy = vi.spyOn(console, 'log').mockImplementation(() => {})
await main(['skills', 'install', '--skill', 'alpha', '--json'], '/tmp/repo')
expect(process.exitCode).toBe(1)
expect(logSpy).toHaveBeenCalledWith(
JSON.stringify(
{
id: 'local',
ok: false,
error: {
code: 'invalid_argument',
message:
"orca skills install --json only supports --dry-run. Real installs stream npx's " +
"own output, which isn't JSON."
},
_meta: { runtimeId: null }
},
null,
2
)
)
expect(spawnMock).not.toHaveBeenCalled()
})
it('prints the resolved install command without running it for --dry-run', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(['skills', 'install', '--skill', 'alpha', '--dry-run'], '/tmp/repo')
expect(stdoutText(stdoutSpy)).toBe(
'npx --yes skills add https://github.com/stablyai/orca --skill alpha --global --agent claude-code --agent universal -y\n\n' +
'Rerun without --dry-run to install now.\n'
)
expect(spawnMock).not.toHaveBeenCalled()
})
it('gives dry-run --json a stable schema', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(['skills', 'install', '--skill', 'legacy-alpha', '--dry-run', '--json'], '/tmp/repo')
expect(stdoutText(stdoutSpy)).toBe(
`${JSON.stringify(
{
command:
'npx --yes skills add https://github.com/stablyai/orca --skill alpha --global --agent claude-code --agent universal -y',
skills: ['alpha'],
global: true,
executed: false
},
null,
2
)}\n`
)
})
it('drops --global for --local in the dry-run command and JSON', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(['skills', 'install', '--skill', 'alpha', '--local', '--dry-run'], '/tmp/repo')
expect(stdoutText(stdoutSpy)).toBe(
'npx --yes skills add https://github.com/stablyai/orca --skill alpha --agent claude-code --agent universal -y\n\n' +
'Rerun without --dry-run to install now.\n'
)
stdoutSpy.mockClear()
await main(
['skills', 'install', '--skill', 'alpha', '--local', '--dry-run', '--json'],
'/tmp/repo'
)
expect(stdoutText(stdoutSpy)).toBe(
`${JSON.stringify(
{
command:
'npx --yes skills add https://github.com/stablyai/orca --skill alpha --agent claude-code --agent universal -y',
skills: ['alpha'],
global: false,
executed: false
},
null,
2
)}\n`
)
})
it('runs npx without --global for --local', async () => {
const child = createFakeChild()
spawnMock.mockReturnValue(child)
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'install', '--skill', 'alpha', '--local'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
expect(spawnMock).toHaveBeenCalledWith(
'npx',
[
'--yes',
'skills',
'add',
'https://github.com/stablyai/orca',
'--skill',
'alpha',
'--agent',
'claude-code',
'--agent',
'universal',
'-y'
],
expect.objectContaining({ stdio: 'inherit' })
)
})
it('routes a resolved Windows .cmd shim through cmd.exe', async () => {
vi.spyOn(process, 'platform', 'get').mockReturnValue('win32')
vi.stubEnv('ComSpec', 'C:\\Windows\\System32\\cmd.exe')
resolveCliCommandMock.mockReturnValue('C:\\Program Files\\nodejs\\npx.cmd')
const child = createFakeChild()
spawnMock.mockReturnValue(child)
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'install', '--skill', 'alpha'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
expect(spawnMock).toHaveBeenCalledWith(
'C:\\Windows\\System32\\cmd.exe',
[
'/d',
'/c',
'C:\\Program Files\\nodejs\\npx.cmd',
'--yes',
'skills',
'add',
'https://github.com/stablyai/orca',
'--skill',
'alpha',
'--global',
'--agent',
'claude-code',
'--agent',
'universal',
'-y'
],
expect.objectContaining({ stdio: 'inherit' })
)
})
it('spawns a resolved npx path directly when it is not a .cmd shim', async () => {
vi.spyOn(process, 'platform', 'get').mockReturnValue('win32')
resolveCliCommandMock.mockReturnValue('C:\\Program Files\\nodejs\\npx.exe')
const child = createFakeChild()
spawnMock.mockReturnValue(child)
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'install', '--skill', 'alpha'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
// Why: an .exe shim must stay a direct spawn so a missing npx still raises
// ENOENT on the child instead of hiding inside cmd.exe's own exit code.
expect(spawnMock.mock.calls[0]?.[0]).toBe('C:\\Program Files\\nodejs\\npx.exe')
expect(spawnMock.mock.calls[0]?.[1]?.[0]).toBe('--yes')
})
it('resolves a legacy topic alias to the canonical skill name for install', async () => {
const child = createFakeChild()
spawnMock.mockReturnValue(child)
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'install', '--skill', 'legacy-alpha'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
expect(spawnMock).toHaveBeenCalledWith(
'npx',
[
'--yes',
'skills',
'add',
'https://github.com/stablyai/orca',
'--skill',
'alpha',
'--global',
'--agent',
'claude-code',
'--agent',
'universal',
'-y'
],
expect.objectContaining({ stdio: 'inherit' })
)
})
it('runs npx for --all and forwards its exit code', async () => {
const child = createFakeChild()
spawnMock.mockReturnValue(child)
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'install', '--all'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 1, null)
await resultPromise
expect(spawnMock).toHaveBeenCalledWith(
'npx',
[
'--yes',
'skills',
'add',
'https://github.com/stablyai/orca',
'--skill',
'alpha',
'--skill',
'gamma',
'--skill',
'zeta',
'--global',
'--agent',
'claude-code',
'--agent',
'universal',
'-y'
],
expect.objectContaining({ stdio: 'inherit' })
)
expect(process.exitCode).toBe(1)
})
it('propagates a spawn error as a nonzero exit', async () => {
const child = createFakeChild()
spawnMock.mockReturnValue(child)
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'install', '--skill', 'alpha'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('error', new Error('spawn npx ENOENT'))
await resultPromise
expect(process.exitCode).toBe(1)
expect(errorSpy).toHaveBeenCalledWith(
'Could not run npx: spawn npx ENOENT. Install Node.js and ensure npx is on PATH.'
)
})
it('lists updatable skills when no --skill/--all is given', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(['skills', 'update'], '/tmp/repo')
expect(stdoutText(stdoutSpy)).toBe(
[
'Choose one or more skills to update:',
' alpha',
' gamma',
' zeta',
'',
'Usage: orca skills update --skill <name> [--skill <name> ...]',
' or: orca skills update --all',
''
].join('\n')
)
expect(spawnMock).not.toHaveBeenCalled()
})
it('prints the resolved update command without running it for --dry-run', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(['skills', 'update', '--skill', 'legacy-alpha', '--dry-run'], '/tmp/repo')
expect(stdoutText(stdoutSpy)).toBe(
'npx --yes skills update alpha --global -y\n\nRerun without --dry-run to update now.\n'
)
expect(spawnMock).not.toHaveBeenCalled()
})
it('selects project scope for --local on update', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(
['skills', 'update', '--skill', 'alpha', '--local', '--dry-run', '--json'],
'/tmp/repo'
)
expect(stdoutText(stdoutSpy)).toBe(
`${JSON.stringify(
{
command: 'npx --yes skills update alpha --project -y',
skills: ['alpha'],
global: false,
executed: false
},
null,
2
)}\n`
)
})
it('runs local updates with explicit project scope', async () => {
const child = createFakeChild()
spawnMock.mockReturnValue(child)
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'update', '--skill', 'alpha', '--local'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
expect(spawnMock).toHaveBeenCalledWith(
'npx',
['--yes', 'skills', 'update', 'alpha', '--project', '-y'],
expect.objectContaining({ stdio: 'inherit' })
)
})
it('refuses a real run when the shell forwards orca to the Orca host', async () => {
vi.stubEnv('ORCA_CLI_CWD', '/home/alice/wt')
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
await main(['skills', 'install', '--skill', 'alpha'], '/tmp/repo')
// Why: the SSH relay and WSL bridge run argv on the Orca host, so a real
// install there would silently skip the machine the user is sitting on.
expect(spawnMock).not.toHaveBeenCalled()
expect(process.exitCode).toBe(1)
expect(String(errorSpy.mock.calls[0]?.[0])).toContain('writes to the machine that runs it')
})
it('refuses --dry-run through the host-forwarding shim too', async () => {
vi.stubEnv('ORCA_CLI_CWD', '/home/alice/wt')
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
await main(['skills', 'install', '--skill', 'alpha', '--dry-run'], '/tmp/repo')
// Why: the targets are resolved from THIS host's agents, so a command printed
// here would name the wrong machine's agents. Point at the target instead.
expect(spawnMock).not.toHaveBeenCalled()
expect(process.exitCode).toBe(1)
expect(String(errorSpy.mock.calls[0]?.[0])).toContain('writes to the machine that runs it')
})
it('puts the resolved npx directory on the child PATH', async () => {
// Why a real directory with a real sibling node: pairing only fires when the
// node it would add actually exists, so a fictional path proves nothing.
const npxBin = mkdtempSync(join(tmpdir(), 'orca-npx-'))
for (const name of ['node', 'npx']) {
writeFileSync(join(npxBin, name), '')
chmodSync(join(npxBin, name), 0o755)
}
const child = createFakeChild()
spawnMock.mockReturnValue(child)
resolveCliCommandMock.mockReturnValue(join(npxBin, 'npx'))
vi.stubEnv('PATH', `/usr/bin${delimiter}/bin`)
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'install', '--skill', 'alpha'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
// Why: npx is an `env node` script, so an off-PATH npx exits 127 with no
// 'error' event unless node ships alongside it on the child's PATH.
const env = spawnMock.mock.calls[0]?.[2]?.env
// Why: the child still needs the inherited PATH and the rest of the parent
// environment; replacing it outright breaks git, node, HOME and npm config.
expect(env?.PATH).toBe(`${npxBin}${delimiter}/usr/bin${delimiter}/bin`)
expect(env?.HOME ?? env?.USERPROFILE).toBe(process.env.HOME ?? process.env.USERPROFILE)
})
it('leaves PATH untouched when no node ships beside the resolved npx', async () => {
// Why: prepending a directory that has no node buys nothing and would shadow
// the caller's own ordering for every other binary the child resolves.
const npxBin = mkdtempSync(join(tmpdir(), 'orca-npx-bare-'))
writeFileSync(join(npxBin, 'npx'), '')
chmodSync(join(npxBin, 'npx'), 0o755)
const child = createFakeChild()
spawnMock.mockReturnValue(child)
resolveCliCommandMock.mockReturnValue(join(npxBin, 'npx'))
vi.stubEnv('PATH', `/usr/bin${delimiter}/bin`)
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'install', '--skill', 'alpha'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
expect(spawnMock.mock.calls[0]?.[2]?.env?.PATH).toBe(`/usr/bin${delimiter}/bin`)
})
it('reports a Windows npx path cmd.exe would reinterpret', async () => {
vi.spyOn(process, 'platform', 'get').mockReturnValue('win32')
vi.stubEnv('ComSpec', 'C:\\Windows\\System32\\cmd.exe')
resolveCliCommandMock.mockReturnValue('C:\\Users\\A&B\\npx.cmd')
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
await main(['skills', 'install', '--skill', 'alpha'], '/tmp/repo')
expect(spawnMock).not.toHaveBeenCalled()
expect(process.exitCode).toBe(1)
expect(String(errorSpy.mock.calls[0]?.[0])).toContain('cmd.exe would reinterpret')
// Why: remediation advice that names the wrong characters is unactionable.
expect(String(errorSpy.mock.calls[0]?.[0])).toContain(WINDOWS_BATCH_UNSAFE_CHARACTERS_LABEL)
})
it('runs npx from a Program Files (x86) install instead of refusing it', async () => {
const child = createFakeChild()
spawnMock.mockReturnValue(child)
vi.spyOn(process, 'platform', 'get').mockReturnValue('win32')
vi.stubEnv('ComSpec', 'C:\\Windows\\System32\\cmd.exe')
const npx = 'C:\\Program Files (x86)\\nodejs\\npx.cmd'
resolveCliCommandMock.mockReturnValue(npx)
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'install', '--skill', 'alpha'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
expect(spawnMock.mock.calls[0]?.[0]).toBe('C:\\Windows\\System32\\cmd.exe')
expect(spawnMock.mock.calls[0]?.[1]?.slice(0, 3)).toEqual(['/d', '/c', npx])
})
it('never puts the current directory on the child PATH when npx is unresolvable', async () => {
const child = createFakeChild()
spawnMock.mockReturnValue(child)
// Why: resolveCliCommand returns the bare name when it finds nothing, and
// dirname('npx') is '.', which would run ./npx out of the caller's checkout.
resolveCliCommandMock.mockReturnValue('npx')
vi.stubEnv('PATH', `/usr/bin${delimiter}/bin`)
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'install', '--skill', 'alpha'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
// Why: assert the constructed value, not the ambient one — a dev PATH with a
// trailing separator carries its own '' entry and would fake a failure here.
expect(spawnMock.mock.calls[0]?.[2]?.env?.PATH).toBe(`/usr/bin${delimiter}/bin`)
})
it('refuses to install when Orca detects no agent, instead of targeting them all', async () => {
detectCommandsMock.mockReturnValue(new Set<string>())
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
await main(['skills', 'install', '--skill', 'alpha'], '/tmp/repo')
// Why: `skills add -y` with nothing detected installs into every agent it
// knows (~75), creating config dirs for agents the host does not have.
expect(spawnMock).not.toHaveBeenCalled()
expect(process.exitCode).toBe(1)
expect(String(errorSpy.mock.calls[0]?.[0])).toContain('No coding agent detected')
})
it('honours an explicit --agent list without probing the host', async () => {
const child = createFakeChild()
spawnMock.mockReturnValue(child)
detectCommandsMock.mockReturnValue(new Set<string>())
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(
['skills', 'install', '--skill', 'alpha', '--agent', 'codex, claude-code ,codex'],
'/tmp/repo'
)
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
const argv = spawnMock.mock.calls[0]?.[1] ?? []
const agents = argv.filter((_: string, i: number) => argv[i - 1] === '--agent')
// Why: trimmed and de-duplicated, and detection is not consulted at all.
expect(agents).toEqual(['codex', 'claude-code'])
expect(detectCommandsMock).not.toHaveBeenCalled()
})
it('maps detected agents onto the skills CLI namespace, not Orca ids', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
detectCommandsMock.mockReturnValue(new Set<string>(['claude', 'cursor-agent', 'rovo']))
await main(['skills', 'install', '--skill', 'alpha', '--dry-run'], '/tmp/repo')
// Why: `skills add` exits 1 on an unknown --agent, and the ids differ —
// Orca's `claude` is `claude-code` and its `rovo` is `rovodev`.
expect(stdoutText(stdoutSpy)).toContain(
'--agent claude-code --agent cursor --agent rovodev --agent universal'
)
})
it('never sends --agent for an update, and never refuses on a bare host', async () => {
const child = createFakeChild()
spawnMock.mockReturnValue(child)
// Why: update refreshes what is already placed, so the no-agent refusal that
// guards install must not reach it.
detectCommandsMock.mockReturnValue(new Set<string>())
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'update', '--skill', 'alpha'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
// Why: update refreshes what is already placed; it chooses no new targets.
expect(spawnMock.mock.calls[0]?.[1]).not.toContain('--agent')
})
it.each([
['a bare --agent', ['skills', 'install', '--skill', 'alpha', '--agent']],
['an empty --agent', ['skills', 'install', '--skill', 'alpha', '--agent', '']],
['a separator-only --agent', ['skills', 'install', '--skill', 'alpha', '--agent', ' , ,']]
])('rejects %s instead of installing to every agent', async (_label, argv) => {
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
await main(argv, '/tmp/repo')
// Why: an --agent that resolves to nothing must not fall back to detection or
// emit no --agent at all — the latter restores the ~75-agent install.
expect(spawnMock).not.toHaveBeenCalled()
expect(process.exitCode).toBe(1)
expect(String(errorSpy.mock.calls[0]?.[0])).toContain('Missing required --agent')
})
it.each([
['a dash-leading value', ['skills', 'install', '--skill', 'alpha', '--agent', '-y']],
['an inline dash value', ['skills', 'install', '--skill', 'alpha', '--agent=--copy']],
['a value with a space', ['skills', 'install', '--skill', 'alpha', '--agent', 'a b']]
])('rejects %s the skills CLI would silently drop', async (_label, argv) => {
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
await main(argv, '/tmp/repo')
// Why: the skills CLI drops such a value, leaving it with no target — the same
// all-agents install as omitting --agent entirely.
expect(spawnMock).not.toHaveBeenCalled()
expect(process.exitCode).toBe(1)
expect(String(errorSpy.mock.calls[0]?.[0])).toContain('Invalid --agent value')
})
it('reports forwarding, not missing agents, when a forwarded host detects none', async () => {
vi.stubEnv('ORCA_CLI_CWD', '/home/alice/wt')
detectCommandsMock.mockReturnValue(new Set<string>())
const errorSpy = vi.spyOn(console, 'error').mockImplementation(() => {})
await main(['skills', 'install', '--skill', 'alpha'], '/tmp/repo')
// Why: resolving targets first would hide the forwarding problem behind a
// no-agent error about the wrong machine.
expect(String(errorSpy.mock.calls[0]?.[0])).toContain('writes to the machine that runs it')
})
it('documents --agent for skills install rather than the terminal-launch flag', async () => {
const logSpy = vi.spyOn(console, 'log').mockImplementation(() => {})
await main(['skills', 'install', '--help'], '/tmp/repo')
const help = String(logSpy.mock.calls[0]?.[0])
expect(help).toContain('--agent <names>')
expect(help).not.toContain('Launch a known TUI agent')
})
it('accumulates a repeated --skill instead of keeping only the last one', async () => {
const child = createFakeChild()
spawnMock.mockReturnValue(child)
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(
['skills', 'install', '--skill', 'zeta', '--skill', 'alpha'],
'/tmp/repo'
)
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
// Why: the documented primary invocation. Dropping 'skill' from the
// repeatable-flag set silently installs one skill instead of two.
expect(spawnMock).toHaveBeenCalledWith(
'npx',
[
'--yes',
'skills',
'add',
'https://github.com/stablyai/orca',
'--skill',
'alpha',
'--skill',
'zeta',
'--global',
'--agent',
'claude-code',
'--agent',
'universal',
'-y'
],
expect.objectContaining({ stdio: 'inherit' })
)
})
it('collapses an alias and its canonical name into one --skill', async () => {
const stdoutSpy = vi.spyOn(process.stdout, 'write').mockImplementation(() => true)
await main(
['skills', 'install', '--skill', 'alpha', '--skill', 'legacy-alpha', '--dry-run'],
'/tmp/repo'
)
expect(stdoutText(stdoutSpy)).toBe(
'npx --yes skills add https://github.com/stablyai/orca --skill alpha --global --agent claude-code --agent universal -y\n\n' +
'Rerun without --dry-run to install now.\n'
)
expect(spawnMock).not.toHaveBeenCalled()
})
it('reports the command it is about to run on stderr', async () => {
const child = createFakeChild()
spawnMock.mockReturnValue(child)
const stderrSpy = vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'install', '--skill', 'alpha'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 0, null)
await resultPromise
// Why: stdout belongs to the child, so this record has to go to stderr.
expect(stderrSpy).toHaveBeenCalledWith(
'Running: npx --yes skills add https://github.com/stablyai/orca --skill alpha --global --agent claude-code --agent universal -y\n'
)
})
it('runs npx skills update for --all and forwards its exit code', async () => {
const child = createFakeChild()
spawnMock.mockReturnValue(child)
vi.spyOn(process.stderr, 'write').mockImplementation(() => true)
const resultPromise = main(['skills', 'update', '--all'], '/tmp/repo')
await vi.waitFor(() => expect(spawnMock).toHaveBeenCalled())
child.emit('exit', 2, null)
await resultPromise
expect(spawnMock).toHaveBeenCalledWith(
'npx',
['--yes', 'skills', 'update', 'alpha', 'gamma', 'zeta', '--global', '-y'],
expect.objectContaining({ stdio: 'inherit' })
)
expect(process.exitCode).toBe(2)
})
it('rejects --json for a real (non-dry-run) update', async () => {
const logSpy = vi.spyOn(console, 'log').mockImplementation(() => {})
await main(['skills', 'update', '--skill', 'alpha', '--json'], '/tmp/repo')
expect(process.exitCode).toBe(1)
expect(logSpy).toHaveBeenCalledWith(
JSON.stringify(
{
id: 'local',
ok: false,
error: {
code: 'invalid_argument',
message:
"orca skills update --json only supports --dry-run. Real updates stream npx's " +
"own output, which isn't JSON."
},
_meta: { runtimeId: null }
},
null,
2
)
)
expect(spawnMock).not.toHaveBeenCalled()
})
})
function stdoutText(spy: ReturnType<typeof vi.spyOn>): string {
return spy.mock.calls.map((call) => String(call[0])).join('')
}
function createFakeChild(): EventEmitter {
return new EventEmitter()
}