mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 08:02:21 +00:00
* feat(agent-launch): host-side prompt delivery for agent.launch The host's agent.launch typed any launch prompt into the shell as part of the launch command. A long or multi-line prompt then ran line by line in the shell, and an agent that never showed readiness or crashed at startup had nothing guarding where its text went. agent.launch now carries a prompt on the typed line only when the line stays one line, control-free and at most 512 bytes; otherwise the agent starts clean and the host pastes the prompt once the agent's own ready signal fires (bracketed paste plus its composer marker or a quiet render, read only after the shell's last hand-off, never while the pane's own shell is proven in front), with main's draft-paste bytes and an Enter 50 ms later. Orchestration worker starts wait on tui-idle as before. A replay-safe launch admits and claims its ledger row in one write, Qwen Code gets a second Enter, the desktop and phone share one launch-refusal classifier, and hosts advertise agent.launch.prompt-carry.v1. Split out of #23748, which moves the desktop source-control buttons onto this path. * fix(agent-launch): keep a short-lined multi-line prompt on a local zsh launch line, as main did #24257 moved every multi-line or over-512-byte prompt off the typed launch line and pasted it after readiness. The phone's AI buttons and review notes, whose multi-line prompts main typed whole into zsh, then reached Claude 0.5-3 s later and their RPC reply waited for the paste. The host now names the shell a local macOS or Linux line is typed into, the way the spawn picks it, and a multi-line prompt rides a zsh line when every line is at most 512 bytes and the whole line at most 8 KB. A real-zsh test types such a line through Orca's own ready barrier and startup write, including when a slow user config makes the write land early. Elsewhere the measured unsafe cases keep the paste: bash 3.2 runs multi-line lines piecemeal, fish drops an early multi-line write, and any shell loses a line over 1 KB written early. * test(agent-launch): keep the real-zsh launch-line test out of the Windows lane's gate scan The Windows lane registration check read `const ZSH_PATH = process.platform === 'win32'` (the head of a multi-line ternary) as a Windows-true flag, so `describe.skipIf(!ZSH_PATH)` looked like a Windows-only suite. The file is POSIX-only; the zsh lookup is now a function. * refactor(protocol): move the agent.launch capabilities into their own module Main's protocol-version.ts sits at the 300-line cap, so the prompt-carry capability pushed it over. The four agent.launch capabilities and their doc move to agent-launch-runtime-capability.ts, re-exported by name and spread into RUNTIME_CAPABILITIES at the same position; the advertised lists and every export are unchanged. * refactor(protocol): import the agent.launch capabilities from their own module `export *` from protocol-version left the four names undefined under the mobile recording loader, which resolves a relative import through a Proxy with no own keys, so 37 phone recordings lost agent.launchReplay. Importers now name agent-launch-runtime-capability directly; protocol-version only spreads its list. * refactor(agent-launch): drop the unshipped viewMode field and trusted local caller id Both were inert in step 1 and existed only for step 2. agent.launch will become a public plugin API, so every wire field is permanent once shipped; a top-level viewMode reads as "choose terminal vs chat", which the host decides. Step 2 introduces placement and view intent under a placement object instead. * fix(agent-launch): read Codex's provisional startup from the rule files' hold anchor Main (#24375) moved Codex's provisional-header check into codex.json's provisional_startup hold anchor and deleted codex-terminal-readiness.ts, so the launch readiness hold now asks showsHoldAnchor, as main's own settled check does. * fix(agent-launch): hold rule-file name titles to quiet for a launch, and census the zsh fixture Main (#24375) answers a name-only title from each agent's rule file ahead of the sustained-title lane, so gemini.json's name_title settled a launch readiness wait on the shell's auto-title while Gemini was still booting. A launch now asks quiet of every weak idle verdict, as that lane did. Main's readiness census requires a recorder for every runtime fixture; the zsh prompt recording is a non-agent control. Gemini's synthetic baseline is regenerated for this PR's stated change: a bare gemini title is no longer its rest mark, so name-only rows settle weak, and a fresh working or blocked status is no longer overridden. * fix(agent-launch): paste a launch prompt only when the launched agent is proven in front A launch pasted its prompt unless a shell was proven in the terminal's foreground, so any read that could not prove one let the prompt through. After an agent exited at startup, its shell turned bracketed paste on at the next prompt, readiness fired on it, and the prompt was typed into the shell: - macOS: a pane runs its shell under login, so the process-group fence's root was never the shell's group and never proved it; the cached foreground name could also still name the exited process. - Windows Git Bash and WSL: the shell-alone-in-its-job check never answers. Now one fresh read of the terminal's foreground decides: agent, shell or unknown. Only 'agent' lets a write through (paste, Enter, second Enter, reused panes too); 'shell' still drops a ready signal. A Windows host never proves the agent, so there the launch line carries the prompt at any size, as on main. * test(agent-launch): cover the Windows QA stub, a grok override that exits at once * fix(agent-launch): keep the local socket alive while a prompted launch waits for its agent A launch with a prompt now waits up to 60 s for the terminal agent to be ready before it writes the prompt, and reports not-delivered when the agent never is. The local runtime socket closes a connection idle for 30 s unless the request is a long poll, so a launch whose agent exited at startup lost its reply and the caller saw 'runtime closed the connection' instead of not-delivered. Classify a prompted agent.launch and agent.launchReplay as a long poll, as orchestration.workerStart already is for the same wait. * refactor(agent-launch): narrow the launch params by 'in' instead of a cast * fix(agent-launch): find a launched agent behind a wrapper that leads its process group A tcsh or nu launch line runs the agent from /bin/sh '<script>', and a wrapper script that does not exec its agent does the same: the wrapper leads the terminal's foreground process group and the agent is a member of it. The fresh foreground read names the group's leader, sh, so a prompted launch was refused or pasted late (M4Air tcsh: 2 of 4 not delivered, 2 pasted ~9 s late). Before that read, take the host's process-group observation as positive proof when it names the launched agent among the foreground group's members and is younger than a ready signal's quiet window. It never proves a shell. * fix(agent-launch): judge the foreground-group proof by when its capture began, not how long ps took The age the host stamps on a process-group observation runs from the start of its whole-machine ps, so on a loaded Mac a capture begun after the read was asked for still read as older than 1 s and the proof was dropped. Count an observation whose capture began after the read was asked for, less the window a shared capture is reused across. * test(agent-launch): keep the crash-guard live test out of the Windows lane's gate scan The Windows-lane registration scan read the const assigned from a platform check as a Windows-only gate, though the suite runs everywhere but Windows; find zsh in a function instead, as the real-zsh typed-line test does. Under load the fresh foreground scan can fail to answer, which lets the shell's prompt settle readiness (2 of 4 paired runs). The guard still refuses that write, so assert the refused write, the property that must always hold. * perf(agent-launch): read a local pane's foreground from its own terminal, not the whole process table The foreground read that gates every launch paste ran the daemon's inspectProcess capture and then a fresh scan, each a whole-machine ps; the fresh one also waits for any capture already running before it starts its own. Measured here at load 5: 1.2 s a read (M4Air QA: 3.4-5.0 s, and worker starts 17.6-32 s against main's 9-12 s at load 25-84). On a local macOS or Linux host, take the pane's root pid from the provider's session inventory and run one ps limited to that pane's terminal. Its foreground process group decides: the launched agent or any non-shell member is the agent (a wrapper that did not exec its agent leads the group), a group of shells alone is the shell. Same pane, same verdict: 2.7 ms a read. SSH hosts keep the relay's observation and name. * test(mobile): re-measure the web app's script sweep after agent.launch's capabilities moved out of protocol-version The mobile web bundle check failed at 124 assets against a ceiling of 123. Main already sat exactly on that ceiling: its sweep table read 69 scripts at 16 routes while the tree builds 73, the whole margin of 4. This branch imports the agent.launch capabilities from their own module, so protocol-version is no longer pulled into the root layout and four other routes. That moves which routes share which modules, and the Qoder capability module, imported by protocol-version and the AI-vault resume path, no longer shares an importer set with anything, so it gets a chunk of its own: 74 scripts. The fence says to re-derive the bound rather than raise it, so the sweep is re-measured on this head (every prefix of the sorted route list). The worst route now adds 10 scripts (session), not 9, which moves the pinned shell crossing from 32 to 30 routes; main re-measured on its own lands on the same crossing. * fix(agent-launch): a worker's brief needs its agent found in front, and Grok's start answers on its composer A paired-server worker start whose agent exited at startup typed its brief into the server's shell, which ran it: the idle wait can settle on a shell back at its prompt, and the brief was written with no foreground read. Both worker-start paths now check before each brief write, as a launch prompt is checked: on a host that can find the agent in front it must be there; on one that cannot (Windows) a shell proven in front still refuses, and anything else writes as before. A Grok worker start waited ~10 s more than main: its only rest signal is its bare name, which a launch holds to quiet output, and Grok animates its logo for ten seconds after its composer glyph. A worker start for an agent whose rest signal is its bare name and whose composer draws a marker (Grok, DSH, mimo-code) now also answers on that marker, whichever comes first.
596 lines
23 KiB
TypeScript
596 lines
23 KiB
TypeScript
/**
|
||
* The executor's ordering contract, which is the defect this module exists to remove.
|
||
*
|
||
* The old shape created a new worktree agent-first, so its startup terminal WAS the agent and the
|
||
* structured branch below it could not be reached for any new worktree. The assertions that matter
|
||
* here are therefore about *order and arguments*, not just the returned mode: a structured launch
|
||
* must create the worktree with `startupAgent: undefined`, and it must ask the host only after the
|
||
* workspace exists.
|
||
*/
|
||
|
||
import { describe, expect, it, vi } from 'vitest'
|
||
import { executeAgentLaunch, type AgentLaunchExecution } from './agent-launch-executor'
|
||
import { AgentLaunchStructuredSessionRefusedError } from './agent-launch-surface-factories'
|
||
import type { AgentLaunchIntent } from '../../shared/agent-launch-intent'
|
||
import { FLOATING_TERMINAL_WORKTREE_ID } from '../../shared/constants'
|
||
|
||
const STRUCTURED_PREFERENCE = {
|
||
experimentalNativeChat: true,
|
||
experimentalStructuredNativeChat: true,
|
||
openAgentTabsInChatByDefault: true
|
||
}
|
||
|
||
function harness(options: {
|
||
settings?: Record<string, unknown> | null
|
||
createSupport?: { supported: boolean; reason?: 'agent' | 'remote' | 'wsl' }
|
||
createSupportThrows?: boolean
|
||
structuredCreateError?: Error
|
||
deliveredMessageId?: string | null
|
||
terminalPromptDelivered?: boolean
|
||
/** Whether the surface reports that its typed line took the offered prompt. */
|
||
lineCarriesPrompt?: boolean
|
||
}) {
|
||
const calls: string[] = []
|
||
const carried = (startupPrompt: string | undefined) =>
|
||
startupPrompt && (options.lineCarriesPrompt ?? true) ? { promptRodeLaunchCommand: true } : {}
|
||
const createWorktree = vi.fn(
|
||
async (args: {
|
||
create: Record<string, unknown>
|
||
startupAgent: string | undefined
|
||
startupPrompt?: string
|
||
}) => {
|
||
calls.push(`createWorktree(startupAgent=${String(args.startupAgent)})`)
|
||
return {
|
||
worktreeId: 'wt-new',
|
||
startupTerminalHandle: args.startupAgent ? 'term_agent_first' : undefined,
|
||
...carried(args.startupPrompt)
|
||
}
|
||
}
|
||
)
|
||
const getStructuredAgentSessionCreateSupport = vi.fn(async () => {
|
||
calls.push('createSupport')
|
||
if (options.createSupportThrows) {
|
||
throw new Error('host unreachable')
|
||
}
|
||
return options.createSupport ?? { supported: true }
|
||
})
|
||
const createStructuredSession = vi.fn(async () => {
|
||
calls.push('createStructuredSession')
|
||
if (options.structuredCreateError) {
|
||
throw options.structuredCreateError
|
||
}
|
||
return { sessionId: 'sess-1', handle: 'handle_structured', fence: 4 }
|
||
})
|
||
const createTerminalAgent = vi.fn(async (args: { startupPrompt?: string }) => {
|
||
calls.push('createTerminalAgent')
|
||
return { handle: 'term_1', ...carried(args.startupPrompt) }
|
||
})
|
||
const deliverStructuredPrompt = vi.fn(async () => {
|
||
calls.push('deliverStructuredPrompt')
|
||
return options.deliveredMessageId === undefined ? 'msg-1' : options.deliveredMessageId
|
||
})
|
||
const deliverTerminalPrompt = vi.fn(async () => {
|
||
calls.push('deliverTerminalPrompt')
|
||
return options.terminalPromptDelivered ?? true
|
||
})
|
||
const runtime = {
|
||
getClientSettings: () =>
|
||
options.settings === undefined ? STRUCTURED_PREFERENCE : options.settings,
|
||
getStructuredAgentSessionCreateSupport
|
||
}
|
||
return {
|
||
calls,
|
||
createWorktree,
|
||
createStructuredSession,
|
||
createTerminalAgent,
|
||
deliverStructuredPrompt,
|
||
deliverTerminalPrompt,
|
||
run: (intent: AgentLaunchIntent) =>
|
||
executeAgentLaunch({
|
||
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the stub implements only the two runtime methods the executor reaches, and each test asserts the calls made, so an omitted method throws rather than reading a wrong value.
|
||
runtime: runtime as unknown as AgentLaunchExecution['runtime'],
|
||
intent,
|
||
surfaces: {
|
||
createStructuredSession,
|
||
createTerminalAgent,
|
||
deliverStructuredPrompt,
|
||
deliverTerminalPrompt
|
||
},
|
||
workspaces: { createWorktree }
|
||
})
|
||
}
|
||
}
|
||
|
||
const CREATE_INTENT: AgentLaunchIntent = {
|
||
agent: 'claude',
|
||
target: { kind: 'create-worktree', create: { repo: 'id:repo-1', name: 'task' } }
|
||
}
|
||
|
||
describe('a structured launch that creates its own worktree', () => {
|
||
it('creates the worktree with no startup agent, then asks the host, then opens a session', async () => {
|
||
const h = harness({})
|
||
const result = await h.run(CREATE_INTENT)
|
||
|
||
// The whole defect in one assertion: the worktree must not be created agent-first.
|
||
expect(h.calls).toEqual([
|
||
'createWorktree(startupAgent=undefined)',
|
||
'createSupport',
|
||
'createStructuredSession'
|
||
])
|
||
expect(result.outcome).toEqual({
|
||
kind: 'structured',
|
||
sessionId: 'sess-1',
|
||
handle: 'handle_structured'
|
||
})
|
||
expect(result.worktreeId).toBe('wt-new')
|
||
expect(result.receipt.mode).toBe('structured')
|
||
})
|
||
|
||
it('asks the host only after the workspace exists, never before', async () => {
|
||
const h = harness({})
|
||
await h.run(CREATE_INTENT)
|
||
expect(h.calls.indexOf('createSupport')).toBeGreaterThan(
|
||
h.calls.indexOf('createWorktree(startupAgent=undefined)')
|
||
)
|
||
})
|
||
|
||
it('falls back to a terminal in the worktree it just created when the host refuses', async () => {
|
||
const h = harness({ createSupport: { supported: false, reason: 'wsl' } })
|
||
const result = await h.run(CREATE_INTENT)
|
||
|
||
expect(h.calls).toEqual([
|
||
'createWorktree(startupAgent=undefined)',
|
||
'createSupport',
|
||
'createTerminalAgent'
|
||
])
|
||
expect(result.outcome).toEqual({ kind: 'terminal', handle: 'term_1' })
|
||
// Not a failed launch, and the workspace is the one just created.
|
||
expect(result.worktreeId).toBe('wt-new')
|
||
expect(result.receipt).toMatchObject({ mode: 'terminal', reason: 'wsl_execution_runtime' })
|
||
})
|
||
|
||
it('falls back to a terminal when the host cannot be reached at all', async () => {
|
||
const h = harness({ createSupportThrows: true })
|
||
const result = await h.run(CREATE_INTENT)
|
||
expect(result.outcome.kind).toBe('terminal')
|
||
expect(result.receipt).toMatchObject({ reason: 'structured_support_unknown' })
|
||
})
|
||
|
||
it('falls back only for a definitive structured refusal after the worktree exists', async () => {
|
||
const h = harness({
|
||
structuredCreateError: new AgentLaunchStructuredSessionRefusedError(
|
||
'structured_agent_session_unsupported',
|
||
'unsupported'
|
||
)
|
||
})
|
||
const result = await h.run(CREATE_INTENT)
|
||
|
||
expect(h.calls).toEqual([
|
||
'createWorktree(startupAgent=undefined)',
|
||
'createSupport',
|
||
'createStructuredSession',
|
||
'createTerminalAgent'
|
||
])
|
||
expect(result.outcome).toEqual({ kind: 'terminal', handle: 'term_1' })
|
||
expect(result.receipt).toMatchObject({
|
||
mode: 'terminal',
|
||
reason: 'structured_unsupported_on_host'
|
||
})
|
||
})
|
||
|
||
it('does not create a duplicate terminal when structured creation is unknown', async () => {
|
||
const h = harness({
|
||
structuredCreateError: new AgentLaunchStructuredSessionRefusedError(
|
||
'agent_session_operation_unknown',
|
||
'unknown'
|
||
)
|
||
})
|
||
|
||
await expect(h.run(CREATE_INTENT)).rejects.toThrow('unknown')
|
||
expect(h.calls).toEqual([
|
||
'createWorktree(startupAgent=undefined)',
|
||
'createSupport',
|
||
'createStructuredSession'
|
||
])
|
||
})
|
||
|
||
it('strips a stale startupAgent out of a migrated create payload', async () => {
|
||
const h = harness({})
|
||
await h.run({
|
||
agent: 'claude',
|
||
target: {
|
||
kind: 'create-worktree',
|
||
// Exactly what mobile sends `worktree.create` today.
|
||
create: { repo: 'id:repo-1', name: 'task', startupAgent: 'claude', startupDraft: 'url' }
|
||
}
|
||
})
|
||
const passed = h.createWorktree.mock.calls[0]?.[0]
|
||
expect(passed?.create).not.toHaveProperty('startupAgent')
|
||
expect(passed?.create).not.toHaveProperty('startupDraft')
|
||
expect(passed?.create).toMatchObject({ repo: 'id:repo-1', name: 'task' })
|
||
})
|
||
})
|
||
|
||
describe('a launch the user did not ask to be structured', () => {
|
||
it('creates the worktree agent-first and never asks the host', async () => {
|
||
const h = harness({ settings: null })
|
||
const result = await h.run(CREATE_INTENT)
|
||
|
||
// Agent-first is preserved for PTY launches: it is what sequences the agent's startup command
|
||
// behind the setup runner, so the wait-for-setup gate comes for free.
|
||
expect(h.calls).toEqual(['createWorktree(startupAgent=claude)'])
|
||
expect(result.outcome).toEqual({ kind: 'terminal', handle: 'term_agent_first' })
|
||
expect(result.receipt).toMatchObject({ mode: 'terminal', reason: 'user_default' })
|
||
})
|
||
})
|
||
|
||
describe('a launch into a workspace that already exists', () => {
|
||
it('opens a session without creating anything', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({ agent: 'codex', target: { kind: 'existing', worktree: 'wt-7' } })
|
||
expect(h.calls).toEqual(['createSupport', 'createStructuredSession'])
|
||
expect(h.createWorktree).not.toHaveBeenCalled()
|
||
expect(result.worktreeId).toBe('wt-7')
|
||
})
|
||
|
||
it('reuses a running terminal without creating or asking', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({
|
||
agent: 'claude',
|
||
target: { kind: 'existing', worktree: 'wt-7' },
|
||
reuseTerminal: { handle: 'term_live' }
|
||
})
|
||
expect(h.calls).toEqual([])
|
||
expect(result.outcome).toEqual({ kind: 'terminal', handle: 'term_live' })
|
||
expect(result.receipt).toMatchObject({ mode: 'terminal', reason: 'reused_terminal' })
|
||
})
|
||
})
|
||
|
||
describe('an agent with no structured session', () => {
|
||
it('stays a terminal without asking the host', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({ agent: 'grok', target: { kind: 'existing', worktree: 'wt-7' } })
|
||
expect(h.calls).toEqual(['createTerminalAgent'])
|
||
expect(result.receipt).toMatchObject({ reason: 'agent_without_structured_session' })
|
||
})
|
||
})
|
||
|
||
describe('the prompt receipt', () => {
|
||
const SUBMIT = { text: 'do the thing', delivery: 'submit' } as const
|
||
|
||
it('commits a submitted prompt to the session the launch created and names the row', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({ ...CREATE_INTENT, prompt: SUBMIT })
|
||
|
||
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'journaled', messageId: 'msg-1' })
|
||
// Delivery is sequenced after the surface exists; there is nothing to send into before that.
|
||
expect(h.calls).toEqual([
|
||
'createWorktree(startupAgent=undefined)',
|
||
'createSupport',
|
||
'createStructuredSession',
|
||
'deliverStructuredPrompt'
|
||
])
|
||
// The send carries the create's own fence; nothing re-reads the session for it.
|
||
expect(h.deliverStructuredPrompt).toHaveBeenCalledWith({
|
||
sessionId: 'sess-1',
|
||
fence: 4,
|
||
prompt: SUBMIT
|
||
})
|
||
})
|
||
|
||
it('under-claims as not delivered when nothing was committed', async () => {
|
||
const h = harness({ deliveredMessageId: null })
|
||
const result = await h.run({ ...CREATE_INTENT, prompt: SUBMIT })
|
||
// A resend costs a duplicate; claiming a row that does not exist loses the text silently.
|
||
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'not-delivered' })
|
||
})
|
||
|
||
it('leaves a draft with the caller, because the host has no composer to hold one', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({
|
||
...CREATE_INTENT,
|
||
prompt: { text: 'do the thing', delivery: 'draft' }
|
||
})
|
||
expect(result.prompt).toEqual({ delivery: 'draft', outcome: 'not-delivered' })
|
||
expect(h.deliverStructuredPrompt).not.toHaveBeenCalled()
|
||
})
|
||
|
||
it('omits the receipt when no prompt was requested', async () => {
|
||
const h = harness({})
|
||
expect((await h.run(CREATE_INTENT)).prompt).toBeUndefined()
|
||
})
|
||
})
|
||
|
||
/**
|
||
* A terminal takes its prompt one of two ways. An agent whose CLI accepts a prompt argument is
|
||
* offered it on the launch command, and the surface that types that line reports whether it rode;
|
||
* everything else is written as keystrokes once the agent is ready. `claude` is argv-mode, `aider`
|
||
* is `stdin-after-start` — the two halves of the table.
|
||
*/
|
||
describe('delivering a launch prompt to a terminal agent', () => {
|
||
const SUBMIT = { text: 'do the thing', delivery: 'submit' } as const
|
||
|
||
it('folds an argv agent’s prompt into the command that starts it, never a paste', async () => {
|
||
const h = harness({ createSupport: { supported: false, reason: 'wsl' } })
|
||
const result = await h.run({ ...CREATE_INTENT, prompt: SUBMIT })
|
||
|
||
expect(result.outcome.kind).toBe('terminal')
|
||
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
|
||
expect(h.createTerminalAgent.mock.calls[0]?.[0]).toMatchObject({
|
||
startupPrompt: 'do the thing'
|
||
})
|
||
// The text was in the process's argv at exec time; a paste on top would be a second copy.
|
||
expect(h.deliverTerminalPrompt).not.toHaveBeenCalled()
|
||
})
|
||
|
||
it('writes a stdin-after-start agent’s prompt into its PTY, because its CLI takes none', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({
|
||
agent: 'aider',
|
||
target: { kind: 'existing', worktree: 'wt-7' },
|
||
prompt: SUBMIT
|
||
})
|
||
|
||
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
|
||
expect(h.deliverTerminalPrompt).toHaveBeenCalledWith({
|
||
handle: 'term_1',
|
||
agent: 'aider',
|
||
freshLaunch: true,
|
||
prompt: SUBMIT
|
||
})
|
||
// Folding it into argv would have appended it as an argument the CLI does not accept.
|
||
expect(h.createTerminalAgent.mock.calls[0]?.[0]).not.toHaveProperty('startupPrompt')
|
||
})
|
||
|
||
it('carries an argv prompt through an agent-first create, which builds the startup command', async () => {
|
||
const h = harness({ settings: null })
|
||
const result = await h.run({ ...CREATE_INTENT, prompt: SUBMIT })
|
||
|
||
expect(result.outcome).toEqual({ kind: 'terminal', handle: 'term_agent_first' })
|
||
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
|
||
expect(h.createWorktree.mock.calls[0]?.[0]).toMatchObject({
|
||
startupAgent: 'claude',
|
||
startupPrompt: 'do the thing'
|
||
})
|
||
expect(h.deliverTerminalPrompt).not.toHaveBeenCalled()
|
||
})
|
||
|
||
it('pastes an argv agent’s prompt after start when the surface reports its typed line could not carry it', async () => {
|
||
const h = harness({
|
||
createSupport: { supported: false, reason: 'wsl' },
|
||
lineCarriesPrompt: false
|
||
})
|
||
const result = await h.run({ ...CREATE_INTENT, prompt: SUBMIT })
|
||
|
||
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
|
||
// Offered to the launch command; the surface, not the executor, decided it did not ride.
|
||
expect(h.createTerminalAgent.mock.calls[0]?.[0]).toMatchObject({
|
||
startupPrompt: 'do the thing'
|
||
})
|
||
expect(h.deliverTerminalPrompt).toHaveBeenCalledWith({
|
||
handle: 'term_1',
|
||
agent: 'claude',
|
||
freshLaunch: true,
|
||
prompt: SUBMIT
|
||
})
|
||
})
|
||
|
||
it('pastes into an agent-first create’s startup terminal when its typed line could not carry the prompt', async () => {
|
||
const h = harness({ settings: null, lineCarriesPrompt: false })
|
||
const result = await h.run({ ...CREATE_INTENT, prompt: SUBMIT })
|
||
|
||
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
|
||
expect(h.deliverTerminalPrompt).toHaveBeenCalledWith({
|
||
handle: 'term_agent_first',
|
||
agent: 'claude',
|
||
freshLaunch: true,
|
||
prompt: SUBMIT
|
||
})
|
||
})
|
||
|
||
it('writes into a reused terminal, whose process started before the launch existed', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({
|
||
agent: 'claude',
|
||
target: { kind: 'existing', worktree: 'wt-7' },
|
||
reuseTerminal: { handle: 'term_existing' },
|
||
prompt: SUBMIT
|
||
})
|
||
|
||
// Argv is unreachable here however argv-friendly the agent is: the process already exists.
|
||
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
|
||
expect(h.deliverTerminalPrompt).toHaveBeenCalledWith({
|
||
handle: 'term_existing',
|
||
agent: 'claude',
|
||
freshLaunch: false,
|
||
prompt: SUBMIT
|
||
})
|
||
})
|
||
|
||
it('under-claims as not delivered when the write did not land', async () => {
|
||
const h = harness({ terminalPromptDelivered: false })
|
||
const result = await h.run({
|
||
agent: 'aider',
|
||
target: { kind: 'existing', worktree: 'wt-7' },
|
||
prompt: SUBMIT
|
||
})
|
||
// A launch whose agent is running must not fail because its text did not; the caller resends.
|
||
expect(result.outcome.kind).toBe('terminal')
|
||
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'not-delivered' })
|
||
})
|
||
|
||
it('leaves a terminal draft with the caller, because the TUI composer is not the host’s to fill', async () => {
|
||
const h = harness({ settings: null })
|
||
const result = await h.run({
|
||
...CREATE_INTENT,
|
||
prompt: { text: 'do the thing', delivery: 'draft' }
|
||
})
|
||
expect(result.prompt).toEqual({ delivery: 'draft', outcome: 'not-delivered' })
|
||
expect(h.deliverTerminalPrompt).not.toHaveBeenCalled()
|
||
// A draft must not be submitted as a turn by riding the launch command either.
|
||
expect(h.createWorktree.mock.calls[0]?.[0]).not.toHaveProperty('startupPrompt')
|
||
})
|
||
})
|
||
|
||
/**
|
||
* The kind is read off the resolved workspace id, so a workspace with nowhere to keep a session is
|
||
* decided here rather than offered to a host probe that cannot answer for it.
|
||
*/
|
||
describe('a launch into an existing workspace, by workspace kind', () => {
|
||
it('runs the floating workspace as a terminal, never a structured session', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({
|
||
agent: 'claude',
|
||
target: { kind: 'existing', worktree: FLOATING_TERMINAL_WORKTREE_ID }
|
||
})
|
||
|
||
// The invariant, not the call order: the floating sentinel has no session store to open into.
|
||
expect(h.createStructuredSession).not.toHaveBeenCalled()
|
||
expect(result.outcome).toEqual({ kind: 'terminal', handle: 'term_1' })
|
||
expect(result.receipt).toMatchObject({
|
||
mode: 'terminal',
|
||
reason: 'structured_unsupported_on_host'
|
||
})
|
||
})
|
||
|
||
it('still opens a structured session in a folder workspace', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({
|
||
agent: 'claude',
|
||
target: { kind: 'existing', worktree: 'folder:fw-1' }
|
||
})
|
||
|
||
// A folder workspace has no git worktree either; it must not be swept up with the sentinel.
|
||
expect(h.createTerminalAgent).not.toHaveBeenCalled()
|
||
expect(result.outcome).toEqual({
|
||
kind: 'structured',
|
||
sessionId: 'sess-1',
|
||
handle: 'handle_structured'
|
||
})
|
||
expect(result.receipt).toMatchObject({ mode: 'structured' })
|
||
})
|
||
|
||
it('still opens a structured session in a git worktree', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({ agent: 'claude', target: { kind: 'existing', worktree: 'wt-7' } })
|
||
|
||
expect(result.outcome.kind).toBe('structured')
|
||
expect(result.receipt).toMatchObject({ mode: 'structured' })
|
||
})
|
||
})
|
||
|
||
/**
|
||
* The launch inputs the host cannot derive for itself.
|
||
*
|
||
* The pair is deliberately asymmetric and the asymmetry is the contract: a requested `cwd` is
|
||
* something only a terminal can apply, so it decides the route; launch arguments are a TUI concern
|
||
* the structured providers version independently, so they do NOT decide the route and a structured
|
||
* launch that received some has to admit it ignored them.
|
||
*/
|
||
describe('caller-supplied launch inputs', () => {
|
||
const EXISTING = { kind: 'existing' as const, worktree: 'wt-7' }
|
||
|
||
it('downgrades a structured preference to a terminal when the launch names a cwd', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({ agent: 'claude', target: EXISTING, cwd: '/repo/packages/api' })
|
||
|
||
// A structured session runs in its workspace, so honouring the cwd and honouring the
|
||
// preference are mutually exclusive; the receipt has to say which one lost.
|
||
expect(result.outcome).toEqual({ kind: 'terminal', handle: 'term_1' })
|
||
expect(result.receipt).toMatchObject({
|
||
mode: 'terminal',
|
||
preferred: 'structured',
|
||
reason: 'tui_launch_command'
|
||
})
|
||
expect(h.createStructuredSession).not.toHaveBeenCalled()
|
||
})
|
||
|
||
it('still opens a structured session when the cwd names the workspace root', async () => {
|
||
// The root the RPC layer resolved rides on the target, so a cwd spelled as the root is not a
|
||
// custom directory and does not decide the route.
|
||
const h = harness({})
|
||
const result = await h.run({
|
||
agent: 'claude',
|
||
target: { kind: 'existing', worktree: 'wt-7', workspacePath: '/repo' },
|
||
cwd: '/repo/'
|
||
})
|
||
expect(result.outcome.kind).toBe('structured')
|
||
expect(result.receipt).toMatchObject({ mode: 'structured' })
|
||
})
|
||
|
||
it('still downgrades for a subdirectory of a resolved root', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({
|
||
agent: 'claude',
|
||
target: { kind: 'existing', worktree: 'wt-7', workspacePath: '/repo' },
|
||
cwd: '/repo/packages/api'
|
||
})
|
||
expect(result.receipt).toMatchObject({ mode: 'terminal', reason: 'tui_launch_command' })
|
||
})
|
||
|
||
it('still opens a structured session when the cwd is only whitespace', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({ agent: 'claude', target: EXISTING, cwd: ' ' })
|
||
|
||
expect(result.outcome.kind).toBe('structured')
|
||
})
|
||
|
||
it('hands cwd, agentArgs and launchSource to the terminal it creates', async () => {
|
||
const h = harness({ settings: null })
|
||
await h.run({
|
||
agent: 'claude',
|
||
target: EXISTING,
|
||
cwd: '/repo/packages/api',
|
||
agentArgs: '--model opus',
|
||
launchSource: 'source_control_recovery'
|
||
})
|
||
|
||
expect(h.createTerminalAgent.mock.calls[0]?.[0]).toMatchObject({
|
||
cwd: '/repo/packages/api',
|
||
agentArgs: '--model opus',
|
||
launchSource: 'source_control_recovery'
|
||
})
|
||
})
|
||
|
||
it('forwards an explicit "no arguments" rather than dropping it as falsy', async () => {
|
||
const h = harness({ settings: null })
|
||
await h.run({ agent: 'claude', target: EXISTING, agentArgs: null })
|
||
|
||
// `null` means the caller wants none; dropping it here would silently restore the user's
|
||
// configured default, which is the opposite of what was asked.
|
||
expect(h.createTerminalAgent.mock.calls[0]?.[0]).toHaveProperty('agentArgs', null)
|
||
})
|
||
|
||
it('omits agentArgs entirely when the caller sent none, so the settings default still applies', async () => {
|
||
const h = harness({ settings: null })
|
||
await h.run({ agent: 'claude', target: EXISTING })
|
||
|
||
expect(h.createTerminalAgent.mock.calls[0]?.[0]).not.toHaveProperty('agentArgs')
|
||
})
|
||
|
||
it('warns that a structured session ignored the launch arguments, without changing the route', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({ agent: 'claude', target: EXISTING, agentArgs: '--model opus' })
|
||
|
||
expect(result.outcome.kind).toBe('structured')
|
||
expect(result.warning).toContain('does not apply launch arguments')
|
||
})
|
||
|
||
it('warns when a structured session ignored an explicit "no arguments" too', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({ agent: 'claude', target: EXISTING, agentArgs: null })
|
||
|
||
// The structured path reads the bypass-permissions bit from the user's SETTINGS default, so a
|
||
// caller that asked for no arguments can still get a session with more permission than it asked
|
||
// for. Staying silent about that is the failure mode worth a test.
|
||
expect(result.warning).toContain('does not apply launch arguments')
|
||
})
|
||
|
||
it('leaves a structured launch unwarned when it carried no arguments at all', async () => {
|
||
const h = harness({})
|
||
const result = await h.run({ agent: 'claude', target: EXISTING })
|
||
|
||
expect(result.warning).toBeUndefined()
|
||
})
|
||
})
|