feat(agent-launch): keep long prompts off the launch line and paste them after readiness (step 1 of 7) (#24257)

* feat(agent-launch): host-side prompt delivery for agent.launch

The host's agent.launch typed any launch prompt into the shell as part of the
launch command. A long or multi-line prompt then ran line by line in the
shell, and an agent that never showed readiness or crashed at startup had
nothing guarding where its text went.

agent.launch now carries a prompt on the typed line only when the line stays
one line, control-free and at most 512 bytes; otherwise the agent starts
clean and the host pastes the prompt once the agent's own ready signal fires
(bracketed paste plus its composer marker or a quiet render, read only after
the shell's last hand-off, never while the pane's own shell is proven in
front), with main's draft-paste bytes and an Enter 50 ms later. Orchestration
worker starts wait on tui-idle as before. A replay-safe launch admits and
claims its ledger row in one write, Qwen Code gets a second Enter, the
desktop and phone share one launch-refusal classifier, and hosts advertise
agent.launch.prompt-carry.v1.

Split out of #23748, which moves the desktop source-control buttons onto
this path.

* fix(agent-launch): keep a short-lined multi-line prompt on a local zsh launch line, as main did

#24257 moved every multi-line or over-512-byte prompt off the typed launch line and pasted it after readiness. The phone's AI buttons and review notes, whose multi-line prompts main typed whole into zsh, then reached Claude 0.5-3 s later and their RPC reply waited for the paste.

The host now names the shell a local macOS or Linux line is typed into, the way the spawn picks it, and a multi-line prompt rides a zsh line when every line is at most 512 bytes and the whole line at most 8 KB. A real-zsh test types such a line through Orca's own ready barrier and startup write, including when a slow user config makes the write land early. Elsewhere the measured unsafe cases keep the paste: bash 3.2 runs multi-line lines piecemeal, fish drops an early multi-line write, and any shell loses a line over 1 KB written early.

* test(agent-launch): keep the real-zsh launch-line test out of the Windows lane's gate scan

The Windows lane registration check read `const ZSH_PATH = process.platform === 'win32'` (the head of a multi-line ternary) as a Windows-true flag, so `describe.skipIf(!ZSH_PATH)` looked like a Windows-only suite. The file is POSIX-only; the zsh lookup is now a function.

* refactor(protocol): move the agent.launch capabilities into their own module

Main's protocol-version.ts sits at the 300-line cap, so the prompt-carry capability pushed it over. The four agent.launch capabilities and their doc move to agent-launch-runtime-capability.ts, re-exported by name and spread into RUNTIME_CAPABILITIES at the same position; the advertised lists and every export are unchanged.

* refactor(protocol): import the agent.launch capabilities from their own module

`export *` from protocol-version left the four names undefined under the mobile recording loader, which resolves a relative import through a Proxy with no own keys, so 37 phone recordings lost agent.launchReplay. Importers now name agent-launch-runtime-capability directly; protocol-version only spreads its list.

* refactor(agent-launch): drop the unshipped viewMode field and trusted local caller id

Both were inert in step 1 and existed only for step 2. agent.launch will become a
public plugin API, so every wire field is permanent once shipped; a top-level
viewMode reads as "choose terminal vs chat", which the host decides. Step 2
introduces placement and view intent under a placement object instead.

* fix(agent-launch): read Codex's provisional startup from the rule files' hold anchor

Main (#24375) moved Codex's provisional-header check into codex.json's
provisional_startup hold anchor and deleted codex-terminal-readiness.ts, so the
launch readiness hold now asks showsHoldAnchor, as main's own settled check does.

* fix(agent-launch): hold rule-file name titles to quiet for a launch, and census the zsh fixture

Main (#24375) answers a name-only title from each agent's rule file ahead of the
sustained-title lane, so gemini.json's name_title settled a launch readiness wait
on the shell's auto-title while Gemini was still booting. A launch now asks quiet
of every weak idle verdict, as that lane did.

Main's readiness census requires a recorder for every runtime fixture; the zsh
prompt recording is a non-agent control. Gemini's synthetic baseline is
regenerated for this PR's stated change: a bare gemini title is no longer its
rest mark, so name-only rows settle weak, and a fresh working or blocked status
is no longer overridden.

* fix(agent-launch): paste a launch prompt only when the launched agent is proven in front

A launch pasted its prompt unless a shell was proven in the terminal's
foreground, so any read that could not prove one let the prompt through. After
an agent exited at startup, its shell turned bracketed paste on at the next
prompt, readiness fired on it, and the prompt was typed into the shell:

- macOS: a pane runs its shell under login, so the process-group fence's root
  was never the shell's group and never proved it; the cached foreground name
  could also still name the exited process.
- Windows Git Bash and WSL: the shell-alone-in-its-job check never answers.

Now one fresh read of the terminal's foreground decides: agent, shell or
unknown. Only 'agent' lets a write through (paste, Enter, second Enter, reused
panes too); 'shell' still drops a ready signal. A Windows host never proves the
agent, so there the launch line carries the prompt at any size, as on main.

* test(agent-launch): cover the Windows QA stub, a grok override that exits at once

* fix(agent-launch): keep the local socket alive while a prompted launch waits for its agent

A launch with a prompt now waits up to 60 s for the terminal agent to be
ready before it writes the prompt, and reports not-delivered when the agent
never is. The local runtime socket closes a connection idle for 30 s unless
the request is a long poll, so a launch whose agent exited at startup lost its
reply and the caller saw 'runtime closed the connection' instead of
not-delivered. Classify a prompted agent.launch and agent.launchReplay as a
long poll, as orchestration.workerStart already is for the same wait.

* refactor(agent-launch): narrow the launch params by 'in' instead of a cast

* fix(agent-launch): find a launched agent behind a wrapper that leads its process group

A tcsh or nu launch line runs the agent from /bin/sh '<script>', and a
wrapper script that does not exec its agent does the same: the wrapper leads
the terminal's foreground process group and the agent is a member of it. The
fresh foreground read names the group's leader, sh, so a prompted launch was
refused or pasted late (M4Air tcsh: 2 of 4 not delivered, 2 pasted ~9 s late).

Before that read, take the host's process-group observation as positive proof
when it names the launched agent among the foreground group's members and is
younger than a ready signal's quiet window. It never proves a shell.

* fix(agent-launch): judge the foreground-group proof by when its capture began, not how long ps took

The age the host stamps on a process-group observation runs from the start
of its whole-machine ps, so on a loaded Mac a capture begun after the read
was asked for still read as older than 1 s and the proof was dropped. Count
an observation whose capture began after the read was asked for, less the
window a shared capture is reused across.

* test(agent-launch): keep the crash-guard live test out of the Windows lane's gate scan

The Windows-lane registration scan read the const assigned from a platform
check as a Windows-only gate, though the suite runs everywhere but Windows;
find zsh in a function instead, as the real-zsh typed-line test does.

Under load the fresh foreground scan can fail to answer, which lets the
shell's prompt settle readiness (2 of 4 paired runs). The guard still refuses
that write, so assert the refused write, the property that must always hold.

* perf(agent-launch): read a local pane's foreground from its own terminal, not the whole process table

The foreground read that gates every launch paste ran the daemon's
inspectProcess capture and then a fresh scan, each a whole-machine ps; the
fresh one also waits for any capture already running before it starts its own.
Measured here at load 5: 1.2 s a read (M4Air QA: 3.4-5.0 s, and worker starts
17.6-32 s against main's 9-12 s at load 25-84).

On a local macOS or Linux host, take the pane's root pid from the provider's
session inventory and run one ps limited to that pane's terminal. Its
foreground process group decides: the launched agent or any non-shell member
is the agent (a wrapper that did not exec its agent leads the group), a group
of shells alone is the shell. Same pane, same verdict: 2.7 ms a read. SSH hosts
keep the relay's observation and name.

* test(mobile): re-measure the web app's script sweep after agent.launch's capabilities moved out of protocol-version

The mobile web bundle check failed at 124 assets against a ceiling of 123. Main already sat
exactly on that ceiling: its sweep table read 69 scripts at 16 routes while the tree builds 73,
the whole margin of 4. This branch imports the agent.launch capabilities from their own module,
so protocol-version is no longer pulled into the root layout and four other routes. That moves
which routes share which modules, and the Qoder capability module, imported by protocol-version
and the AI-vault resume path, no longer shares an importer set with anything, so it gets a chunk
of its own: 74 scripts.

The fence says to re-derive the bound rather than raise it, so the sweep is re-measured on this
head (every prefix of the sorted route list). The worst route now adds 10 scripts (session), not
9, which moves the pinned shell crossing from 32 to 30 routes; main re-measured on its own lands
on the same crossing.

* fix(agent-launch): a worker's brief needs its agent found in front, and Grok's start answers on its composer

A paired-server worker start whose agent exited at startup typed its brief into the server's
shell, which ran it: the idle wait can settle on a shell back at its prompt, and the brief was
written with no foreground read. Both worker-start paths now check before each brief write, as a
launch prompt is checked: on a host that can find the agent in front it must be there; on one that
cannot (Windows) a shell proven in front still refuses, and anything else writes as before.

A Grok worker start waited ~10 s more than main: its only rest signal is its bare name, which a
launch holds to quiet output, and Grok animates its logo for ten seconds after its composer glyph.
A worker start for an agent whose rest signal is its bare name and whose composer draws a marker
(Grok, DSH, mimo-code) now also answers on that marker, whichever comes first.
This commit is contained in:
Brennan Benson
2026-10-05 11:26:52 -07:00
committed by GitHub
parent f250db55b7
commit 761d63a4e5
102 changed files with 4767 additions and 371 deletions
@@ -539,10 +539,10 @@ describe('the Phase C budget', () => {
key
).toBeGreaterThanOrEqual(measured)
}
// 1 to 9 per route, which is why four per route was a bound rather than a fit and why the
// 1 to 10 per route, which is why four per route was a bound rather than a fit and why the
// envelope cannot be a line through the measurement either.
expect(Math.min(...MOBILE_WEB_APP_BUNDLE_ROUTE_SCRIPT_SPREAD)).toBe(1)
expect(Math.max(...MOBILE_WEB_APP_BUNDLE_ROUTE_SCRIPT_SPREAD)).toBe(9)
expect(Math.max(...MOBILE_WEB_APP_BUNDLE_ROUTE_SCRIPT_SPREAD)).toBe(10)
})
it('sits exactly one margin over the swept tree and grants the worst route beyond it', () => {
@@ -614,13 +614,13 @@ describe('the Phase C budget', () => {
it('fails the build when the derived ceiling passes what the phone will accept', async () => {
// The shell hands back null for a manifest over its own ceiling, so a derived ceiling above
// that ships a green build no device can open. At the 42 images the tree carries, the envelope
// plus 42 plus the document crosses 256 at 32 routes, which Phase C reaches. The crossing came
// plus 42 plus the document crosses 256 at 30 routes, which Phase C reaches. The crossing came
// in from 50 with the envelope: it grants the worst swept route to each one past the sweep,
// where `4r + 16` granted four, so re-measuring a tree whose routes share more moves it out.
expect(await readMobileWebBundleMaxAssets()).toBe(MOBILE_WEB_BUNDLE_MAX_ASSETS)
expect(assertAssetCeilingFitsShell(31, 42, MOBILE_WEB_BUNDLE_MAX_ASSETS)).toBe(251)
expect(() => assertAssetCeilingFitsShell(32, 42, MOBILE_WEB_BUNDLE_MAX_ASSETS)).toThrow(
/260 .*256/
expect(assertAssetCeilingFitsShell(29, 42, MOBILE_WEB_BUNDLE_MAX_ASSETS)).toBe(251)
expect(() => assertAssetCeilingFitsShell(30, 42, MOBILE_WEB_BUNDLE_MAX_ASSETS)).toThrow(
/261 .*256/
)
})
})
+16 -16
View File
@@ -50,9 +50,9 @@ export const MOBILE_WEB_APP_BUNDLE_MAX_TOTAL_BYTES = 9 * 1024 * 1024
* A chunk is emitted per distinct set of importers, not per route, so a route's marginal cost is
* what it fails to share rather than what it weighs. Re-measured on this head by building
* `routes.slice(0, n)` for every n, which is what the fence below is derived from rather than
* fitted to. The spread it shows is 1 to 9: `pr` and `web` add one script each, `review` adds nine.
* fitted to. The spread it shows is 1 to 10: `pr` and `web` add one script each, `session` adds ten.
* The root `./_layout.tsx` (the page's web sibling of the native root) sorts first; with it the
* swept tree reads 69 scripts at 16 routes, the old 15 read 67 on the same head.
* swept tree reads 74 scripts at 16 routes.
*
* This table is the fence's only input, so a route added to the tree stales it and the pins beside
* the fence fail until it is re-measured. That is the point: the bound is re-derived, never bumped.
@@ -61,19 +61,19 @@ export const MOBILE_WEB_APP_BUNDLE_SCRIPT_SWEEP = [
['./_layout.tsx', 3],
['./h/[hostId]/[...page].tsx', 7],
['./h/[hostId]/accounts.tsx', 9],
['./h/[hostId]/agent-history/[worktreeId].tsx', 13],
['./h/[hostId]/edit.tsx', 18],
['./h/[hostId]/files/[worktreeId].tsx', 21],
['./h/[hostId]/files/preview/[worktreeId].tsx', 28],
['./h/[hostId]/history/[worktreeId].tsx', 30],
['./h/[hostId]/index.tsx', 35],
['./h/[hostId]/pr/[worktreeId].tsx', 36],
['./h/[hostId]/review/[worktreeId].tsx', 45],
['./h/[hostId]/session/[worktreeId].tsx', 54],
['./h/[hostId]/source-control/[worktreeId].tsx', 59],
['./h/[hostId]/tasks.tsx', 66],
['./h/[hostId]/web.tsx', 67],
['./h/_layout.tsx', 69]
['./h/[hostId]/agent-history/[worktreeId].tsx', 14],
['./h/[hostId]/edit.tsx', 19],
['./h/[hostId]/files/[worktreeId].tsx', 24],
['./h/[hostId]/files/preview/[worktreeId].tsx', 32],
['./h/[hostId]/history/[worktreeId].tsx', 34],
['./h/[hostId]/index.tsx', 39],
['./h/[hostId]/pr/[worktreeId].tsx', 40],
['./h/[hostId]/review/[worktreeId].tsx', 48],
['./h/[hostId]/session/[worktreeId].tsx', 58],
['./h/[hostId]/source-control/[worktreeId].tsx', 63],
['./h/[hostId]/tasks.tsx', 71],
['./h/[hostId]/web.tsx', 72],
['./h/_layout.tsx', 74]
]
const sweptScripts = MOBILE_WEB_APP_BUNDLE_SCRIPT_SWEEP.map(([, scripts]) => scripts)
@@ -191,7 +191,7 @@ export async function readMobileWebBundleMaxAssets() {
* shells return null for a manifest over MOBILE_WEB_BUNDLE_MAX_ASSETS rather than dropping the
* extra assets, so a route count that pushes the chunk envelope plus images plus the document past
* it would pass this build and fail on the device with nothing to read. At today's 42 images that
* is 31 routes, inside what Phase C adds, which is why this is a build failure and not a comment.
* is 29 routes, inside what Phase C adds, which is why this is a build failure and not a comment.
* The envelope grants the worst swept route to each one past the sweep, so re-measuring a tree
* whose routes share more moves that crossing out again.
*/
@@ -7,19 +7,17 @@
*/
import type { AgentLaunchPrompt, AgentLaunchResult } from '../../../src/shared/agent-launch-intent'
import { AGENT_LAUNCH_PANE_ALREADY_LIVE_CODE } from '../../../src/shared/agent-launch-pane-already-live'
import { AGENT_LAUNCH_SESSION_ALREADY_EXISTS_CODE } from '../../../src/shared/agent-launch-session-already-exists'
import { isAgentSessionHandleProvider } from '../../../src/shared/agent-session-provider-handle'
import { makePaneKey } from '../../../src/shared/stable-pane-id'
import { createStructuredAgentSessionId } from '../../../src/shared/structured-agent-session-create'
import type { TuiAgent } from '../../../src/shared/tui-agent'
import type { RpcClient } from '../transport/rpc-client'
import { agentLaunchReplayRun } from '../tasks/mobile-workspace-create-operations'
import { agentLaunchExistingParams, readAgentLaunchSupport } from '../tasks/agent-launch-request'
import {
agentLaunchExistingParams,
isAgentLaunchReplayUnsupportedRefusal,
readAgentLaunchSupport
} from '../tasks/agent-launch-request'
classifyAgentLaunchReplayRefusal,
isAgentLaunchReservationTakenRefusal
} from '../../../src/shared/agent-launch-replay-refusal'
import { sendReplayingAmbiguousDelivery } from '../tasks/replay-on-ambiguous-delivery'
import {
structuredSessionOperationId,
@@ -152,29 +150,17 @@ function classifyLaunchRefusal(
error: { code?: string; message?: string },
replayed: boolean
): MobileExistingAgentLaunch {
if (isAgentLaunchReplayUnsupportedRefusal(error)) {
// Only a refusal of the first send proves nothing ran; after a replay it may be a replacement
// connection whose capability list hasn't landed, answering for an attempt that did start.
return replayed
? { kind: 'unknown', message: AGENT_LAUNCH_UNCONFIRMED_MESSAGE }
: { kind: 'unsupported' }
switch (classifyAgentLaunchReplayRefusal(error, replayed)) {
case 'unsupported':
return { kind: 'unsupported' }
case 'unknown':
return { kind: 'unknown', message: AGENT_LAUNCH_UNCONFIRMED_MESSAGE }
case 'failed': {
if (isAgentLaunchReservationTakenRefusal(error)) {
return { kind: 'failed', message: AGENT_LAUNCH_RESERVATION_TAKEN_MESSAGE }
}
const message = error.message?.trim()
return { kind: 'failed', message: message || "Couldn't start the agent." }
}
}
if (
error.code === 'agent_session_operation_unknown' ||
error.code === 'agent_session_operation_expired'
) {
return { kind: 'unknown', message: AGENT_LAUNCH_UNCONFIRMED_MESSAGE }
}
if (
error.code === AGENT_LAUNCH_PANE_ALREADY_LIVE_CODE ||
error.code === AGENT_LAUNCH_SESSION_ALREADY_EXISTS_CODE
) {
// A taken reservation proves nothing started only on the first send; after a replay the pane
// or chat holding it may be this launch's own.
return replayed
? { kind: 'unknown', message: AGENT_LAUNCH_UNCONFIRMED_MESSAGE }
: { kind: 'failed', message: AGENT_LAUNCH_RESERVATION_TAKEN_MESSAGE }
}
const message = error.message?.trim()
return { kind: 'failed', message: message || "Couldn't start the agent." }
}
@@ -20,7 +20,7 @@ import { runtimeStub } from '../../../src/main/runtime/rpc/methods/agent-launch.
import {
AGENT_LAUNCH_REPLAY_REQUIRED_RUNTIME_CAPABILITY,
AGENT_LAUNCH_RUNTIME_CAPABILITY
} from '../../../src/shared/protocol-version'
} from '../../../src/shared/agent-launch-runtime-capability'
import {
launchAgentInExistingWorkspace,
reserveMobileAgentLaunch
+3 -9
View File
@@ -24,7 +24,7 @@ import {
import {
AGENT_LAUNCH_REPLAY_REQUIRED_RUNTIME_CAPABILITY,
AGENT_LAUNCH_RUNTIME_CAPABILITY
} from '../../../src/shared/protocol-version'
} from '../../../src/shared/agent-launch-runtime-capability'
import type { TuiAgent } from '../../../src/shared/tui-agent'
import type { RpcSendParams } from '../transport/rpc-params-contract'
import type { WorkspaceCreateParams } from './workspace-create-params'
@@ -142,11 +142,5 @@ export function isAgentLaunchUnsupportedRefusal(error: {
return (error.message ?? '').includes('agent_launch_unsupported')
}
/** The `agent.launchReplay` twin: an older host rejects the method rather than a field. */
export function isAgentLaunchReplayUnsupportedRefusal(error: { code?: string }): boolean {
return (
error.code === 'method_not_found' ||
error.code === 'forbidden' ||
error.code === 'agent_launch_replay_unsupported'
)
}
// The `agent.launchReplay` twin lives in shared so the desktop classifies refusals the same way.
export { isAgentLaunchReplayUnsupportedRefusal } from '../../../src/shared/agent-launch-replay-refusal'
@@ -23,7 +23,7 @@ import { readNewWorktreeRuntimeCapabilities } from './worktree-create-capability
import {
AGENT_LAUNCH_RUNTIME_CAPABILITY,
AGENT_LAUNCH_REPLAY_REQUIRED_RUNTIME_CAPABILITY
} from '../../../src/shared/protocol-version'
} from '../../../src/shared/agent-launch-runtime-capability'
const createStructuredSession = vi.fn()
vi.mock('../../../src/main/runtime/rpc/methods/structured-agent-session-create', () => ({
@@ -3,7 +3,7 @@ import {
AGENT_LAUNCH_REPLAY_REQUIRED_RUNTIME_CAPABILITY,
AGENT_LAUNCH_REPLAY_RUNTIME_CAPABILITY,
AGENT_LAUNCH_RUNTIME_CAPABILITY
} from '../../../src/shared/protocol-version'
} from '../../../src/shared/agent-launch-runtime-capability'
import type { RpcClient } from '../transport/rpc-client'
import { LogicalClientCutoverError } from '../transport/stable-logical-rpc-client'
import { readNewWorktreeRuntimeCapabilities } from './worktree-create-capability'
@@ -1,12 +1,12 @@
import { describe, expect, it } from 'vitest'
import {
AGENT_LAUNCH_RUNTIME_CAPABILITY,
AGENT_SESSION_TURN_ITEM_CAPABILITY,
CLAUDE_STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY,
SESSION_TABS_SPLIT_GROUP_PLACEMENT_RUNTIME_CAPABILITY,
STRUCTURED_AGENT_SESSION_HOLD_RUNTIME_CAPABILITY,
STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY
} from '../../../src/shared/protocol-version'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../../src/shared/agent-launch-runtime-capability'
import { MOBILE_RUNTIME_CLIENT_CAPABILITIES } from './mobile-runtime-client-capabilities'
/** Mirrors the host's `parseRuntimeClientCapabilities`, which returns an EMPTY list — silently
@@ -1,5 +1,4 @@
import {
AGENT_LAUNCH_RUNTIME_CAPABILITY,
AGENT_SESSION_PENDING_SEND_RESULT_RUNTIME_CAPABILITY,
AGENT_SESSION_TURN_ITEM_CAPABILITY,
CLAUDE_STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY,
@@ -7,6 +6,7 @@ import {
STRUCTURED_AGENT_SESSION_HOLD_RUNTIME_CAPABILITY,
STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY
} from '../../../src/shared/protocol-version'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../../src/shared/agent-launch-runtime-capability'
import { remoteRuntimeClientCapabilities } from '../../../src/shared/remote-runtime-client-capabilities'
export const MOBILE_RUNTIME_CLIENT_CAPABILITIES = remoteRuntimeClientCapabilities([
@@ -2,8 +2,8 @@
# attributes a fresh agent the host builds; carries agentStartedTelemetry
# resume continues an existing agent session; not a new start
# mobile-followup the phone's session-tab launch; attribution is a separate follow-up
# The shared execution-host builder also serves bare commands; only fresh-agent callers supply attribution.
src/main/opencode/opencode-model-startup-plan.ts 2 attributes
src/main/runtime/runtime-worktree-agent-startup.ts 3 attributes
# The shared execution-host builders also serve bare commands; only fresh-agent callers supply attribution.
src/main/opencode/opencode-model-startup-plan.ts 3 attributes
src/main/runtime/runtime-worktree-agent-startup.ts 4 attributes
src/main/runtime/orca-runtime-get-agent-session-execution-namespace.ts 1 resume
src/main/runtime/orca-runtime-resolve-mobile-session-terminal-command.ts 1 mobile-followup
@@ -27,8 +27,12 @@ function harness(options: {
structuredCreateError?: Error
deliveredMessageId?: string | null
terminalPromptDelivered?: boolean
/** Whether the surface reports that its typed line took the offered prompt. */
lineCarriesPrompt?: boolean
}) {
const calls: string[] = []
const carried = (startupPrompt: string | undefined) =>
startupPrompt && (options.lineCarriesPrompt ?? true) ? { promptRodeLaunchCommand: true } : {}
const createWorktree = vi.fn(
async (args: {
create: Record<string, unknown>
@@ -38,7 +42,8 @@ function harness(options: {
calls.push(`createWorktree(startupAgent=${String(args.startupAgent)})`)
return {
worktreeId: 'wt-new',
startupTerminalHandle: args.startupAgent ? 'term_agent_first' : undefined
startupTerminalHandle: args.startupAgent ? 'term_agent_first' : undefined,
...carried(args.startupPrompt)
}
}
)
@@ -56,9 +61,9 @@ function harness(options: {
}
return { sessionId: 'sess-1', handle: 'handle_structured', fence: 4 }
})
const createTerminalAgent = vi.fn(async (_args: { startupPrompt?: string }) => {
const createTerminalAgent = vi.fn(async (args: { startupPrompt?: string }) => {
calls.push('createTerminalAgent')
return { handle: 'term_1' }
return { handle: 'term_1', ...carried(args.startupPrompt) }
})
const deliverStructuredPrompt = vi.fn(async () => {
calls.push('deliverStructuredPrompt')
@@ -297,10 +302,10 @@ describe('the prompt receipt', () => {
})
/**
* A terminal takes its prompt one of two ways, and which one is not a preference: an agent whose
* CLI accepts a prompt argument must get it on argv, because that is the transport that survives
* multi-line and special-character text. Only an agent with no such argument is written to as
* keystrokes. `claude` is argv-mode, `aider` is `stdin-after-start` — the two halves of the table.
* A terminal takes its prompt one of two ways. An agent whose CLI accepts a prompt argument is
* offered it on the launch command, and the surface that types that line reports whether it rode;
* everything else is written as keystrokes once the agent is ready. `claude` is argv-mode, `aider`
* is `stdin-after-start` — the two halves of the table.
*/
describe('delivering a launch prompt to a terminal agent', () => {
const SUBMIT = { text: 'do the thing', delivery: 'submit' } as const
@@ -329,6 +334,8 @@ describe('delivering a launch prompt to a terminal agent', () => {
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
expect(h.deliverTerminalPrompt).toHaveBeenCalledWith({
handle: 'term_1',
agent: 'aider',
freshLaunch: true,
prompt: SUBMIT
})
// Folding it into argv would have appended it as an argument the CLI does not accept.
@@ -348,6 +355,39 @@ describe('delivering a launch prompt to a terminal agent', () => {
expect(h.deliverTerminalPrompt).not.toHaveBeenCalled()
})
it('pastes an argv agent’s prompt after start when the surface reports its typed line could not carry it', async () => {
const h = harness({
createSupport: { supported: false, reason: 'wsl' },
lineCarriesPrompt: false
})
const result = await h.run({ ...CREATE_INTENT, prompt: SUBMIT })
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
// Offered to the launch command; the surface, not the executor, decided it did not ride.
expect(h.createTerminalAgent.mock.calls[0]?.[0]).toMatchObject({
startupPrompt: 'do the thing'
})
expect(h.deliverTerminalPrompt).toHaveBeenCalledWith({
handle: 'term_1',
agent: 'claude',
freshLaunch: true,
prompt: SUBMIT
})
})
it('pastes into an agent-first create’s startup terminal when its typed line could not carry the prompt', async () => {
const h = harness({ settings: null, lineCarriesPrompt: false })
const result = await h.run({ ...CREATE_INTENT, prompt: SUBMIT })
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
expect(h.deliverTerminalPrompt).toHaveBeenCalledWith({
handle: 'term_agent_first',
agent: 'claude',
freshLaunch: true,
prompt: SUBMIT
})
})
it('writes into a reused terminal, whose process started before the launch existed', async () => {
const h = harness({})
const result = await h.run({
@@ -361,6 +401,8 @@ describe('delivering a launch prompt to a terminal agent', () => {
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
expect(h.deliverTerminalPrompt).toHaveBeenCalledWith({
handle: 'term_existing',
agent: 'claude',
freshLaunch: false,
prompt: SUBMIT
})
})
+35 -12
View File
@@ -71,8 +71,12 @@ export type AgentLaunchExecution = {
vocabulary?: AgentLaunchModeVocabulary
/** Attributes a throw to the step that was running, the way a dispatch's own stages do. */
onStage?: (stage: 'worktree_create' | 'mode_settle' | 'surface_create') => void
/** The surface exists and its tab is published; runs before any prompt delivery. Must not throw. */
onSurfacePublished?: (surface: AgentLaunchPublishedSurface) => void
}
export type AgentLaunchPublishedSurface = Pick<AgentLaunchResult, 'outcome' | 'worktreeId'>
export async function executeAgentLaunch(
execution: AgentLaunchExecution
): Promise<AgentLaunchResult> {
@@ -99,13 +103,18 @@ export async function executeAgentLaunch(
// A reused terminal already downgraded in the pre-flight; there is nothing to create. Its agent
// was running before this launch existed, so argv is unreachable and the PTY is the only way in.
if (intent.reuseTerminal) {
return {
const reused = published(execution, {
outcome: { kind: 'terminal', handle: intent.reuseTerminal.handle },
worktreeId: existingWorktreeId(intent.target),
worktreeId: existingWorktreeId(intent.target)
})
return {
...reused,
receipt: preflight,
...promptReceipt(
intent,
await deliverTerminalLaunchPrompt(execution, intent.reuseTerminal.handle)
await deliverTerminalLaunchPrompt(execution, intent.reuseTerminal.handle, {
freshLaunch: false
})
)
}
}
@@ -113,20 +122,25 @@ export async function executeAgentLaunch(
const placed = await resolveWorkspace(execution, preflight)
// Agent-first creation already produced the agent, so the pre-flight verdict is final.
if (placed.startupTerminalHandle) {
return {
const startup = published(execution, {
outcome: {
kind: 'terminal',
handle: placed.startupTerminalHandle,
...(placed.startupTerminalPaneKey ? { paneKey: placed.startupTerminalPaneKey } : {})
},
worktreeId: placed.worktreeId,
worktreeId: placed.worktreeId
})
return {
...startup,
receipt: preflight,
...(placed.warning ? { warning: placed.warning } : {}),
...promptReceipt(
intent,
placed.promptRodeLaunchCommand
? HANDED_TO_TERMINAL
: await deliverTerminalLaunchPrompt(execution, placed.startupTerminalHandle)
: await deliverTerminalLaunchPrompt(execution, placed.startupTerminalHandle, {
freshLaunch: true
})
)
}
}
@@ -169,15 +183,23 @@ export async function executeAgentLaunch(
// not start while looking at it. Telling those apart needs `createManagedWorktree` to stop
// multiplexing "couldn't copy untracked files" and "startup terminal failed" into one string.
const warning = combineLaunchWarnings(placed.warning, created.warning)
const surface = published(execution, { outcome: created.outcome, worktreeId: placed.worktreeId })
return {
outcome: created.outcome,
worktreeId: placed.worktreeId,
...surface,
receipt: settled,
...(warning ? { warning } : {}),
...promptReceipt(intent, await settleLaunchPromptDisposal(execution, created))
}
}
function published(
execution: AgentLaunchExecution,
surface: AgentLaunchPublishedSurface
): AgentLaunchPublishedSurface {
execution.onSurfacePublished?.(surface)
return surface
}
function downgradeAgentLaunchModeForStructuredRefusal(
receipt: AgentLaunchModeReceipt,
vocabulary: AgentLaunchModeVocabulary
@@ -223,9 +245,10 @@ async function resolveWorkspace(
})
// Only when a startup terminal actually came back: a create that produced none ran no command,
// so nothing carried the prompt and the launch still owes it to whatever surface it builds next.
return created.startupTerminalHandle && startupPrompt
? { ...created, promptRodeLaunchCommand: true }
: created
const { promptRodeLaunchCommand, ...rest } = created
return rest.startupTerminalHandle && promptRodeLaunchCommand
? { ...rest, promptRodeLaunchCommand: true }
: rest
}
/** `structured` is the same surface `outcome` names, kept typed so prompt delivery reads the create's
@@ -327,7 +350,7 @@ async function createTerminalSurface(
...(terminal.paneKey ? { paneKey: terminal.paneKey } : {})
},
...(terminal.warning ? { warning: terminal.warning } : {}),
...(startupPrompt ? { promptRodeLaunchCommand: true } : {})
...(startupPrompt && terminal.promptRodeLaunchCommand ? { promptRodeLaunchCommand: true } : {})
}
}
@@ -5,16 +5,15 @@
* surface exists and in what order; everything here decides how the text reaches whichever surface
* that turned out to be, and each surface takes it differently:
*
* structured session -> committed to the transcript, named by a message id -> journaled
* terminal, argv CLI -> folded into the command that execs the agent -> handed-to-terminal
* terminal, no argv -> bracketed paste into the live PTY -> handed-to-terminal
* anything unproven -> -> not-delivered
* structured session -> committed to the transcript, named by a message id -> journaled
* terminal, line fits -> folded into the command that execs the agent -> handed-to-terminal
* terminal, otherwise -> bracketed paste into the live PTY once it is ready -> handed-to-terminal
* anything unproven -> -> not-delivered
*
* The argv/PTY fork is not a preference. `argv` exists so multi-line and special-character text
* reaches a CLI as one argument instead of keystrokes, and it has no readiness race because the
* text is in the process's arguments at exec time. So it is preferred wherever the agent's CLI
* takes a prompt argument, and `agentPromptRidesLaunchCommand` — derived from the same injection
* table `buildAgentStartupPlan` branches on — is the one place that question is asked.
* argv has no readiness race, so it is offered wherever the agent's CLI takes a prompt argument
* (`agentPromptRidesLaunchCommand`). But that command is TYPED into the user's shell, and a long or
* multi-line typed line fails in ways argv itself does not, so whether the offer was taken is decided
* where the line is built (`startup-line-prompt-carry`) and reported back, never predicted here.
*/
import type {
@@ -39,11 +38,11 @@ export async function settleLaunchPromptDisposal(
const messageId = await deliverStructuredLaunchPrompt(execution, created.structured)
return messageId ? { outcome: 'journaled', messageId } : NOT_DELIVERED
}
// The startup command already carries an argv agent's prompt; there is nothing left to write.
// The surface reported that the startup command carried the prompt; there is nothing left to write.
if (created.promptRodeLaunchCommand) {
return HANDED_TO_TERMINAL
}
return deliverTerminalLaunchPrompt(execution, created.outcome.handle)
return deliverTerminalLaunchPrompt(execution, created.outcome.handle, { freshLaunch: true })
}
/**
@@ -82,13 +81,19 @@ async function deliverStructuredLaunchPrompt(
*/
export async function deliverTerminalLaunchPrompt(
execution: AgentLaunchExecution,
handle: string
handle: string,
{ freshLaunch }: { freshLaunch: boolean }
): Promise<AgentLaunchPromptDisposal> {
const { intent, surfaces } = execution
if (!intent.prompt || intent.prompt.delivery !== 'submit') {
return NOT_DELIVERED
}
const delivered = await surfaces.deliverTerminalPrompt?.({ handle, prompt: intent.prompt })
const delivered = await surfaces.deliverTerminalPrompt?.({
handle,
agent: intent.agent,
freshLaunch,
prompt: intent.prompt
})
return delivered ? HANDED_TO_TERMINAL : NOT_DELIVERED
}
@@ -98,8 +103,8 @@ function launchSubmitText(intent: AgentLaunchIntent): string | undefined {
}
/**
* The prompt a terminal's launch command should carry, which is an argv-mode agent's and only an
* argv-mode agent's.
* The prompt offered to a terminal's launch command, which is an argv-mode agent's and only an
* argv-mode agent's. An offer, not a decision: the surface reports whether the typed line took it.
*/
export function argvLaunchPrompt(intent: AgentLaunchIntent): string | undefined {
const text = launchSubmitText(intent)
@@ -25,8 +25,8 @@ export type AgentLaunchSurfaceFactory = {
worktreeId: string
agent: TuiAgent
options?: Readonly<Record<string, unknown>>
/** Set only for an agent whose CLI takes the prompt on argv, so the text is in the process's
* arguments at exec time rather than raced into its composer afterwards. */
/** Offered only for an agent whose CLI takes the prompt on argv. It rides the launch command
* only when the typed line can carry it; `promptRodeLaunchCommand` reports which happened. */
startupPrompt?: string
/** Replaces the settings default for this launch only; `null` means no arguments at all. */
agentArgs?: string | null
@@ -40,6 +40,8 @@ export type AgentLaunchSurfaceFactory = {
/** The pane this create minted; a factory whose runtime reports none omits it, never invents. */
paneKey?: string
warning?: string
/** Reported by the surface that built the typed line, never predicted by the executor. */
promptRodeLaunchCommand?: boolean
}>
/**
* Commits the launch text as the session's first turn, answering with the transcript row's id.
@@ -56,12 +58,19 @@ export type AgentLaunchSurfaceFactory = {
/**
* Writes the launch text into a terminal agent's live PTY, answering whether it landed.
*
* The other half of `startupPrompt`, for the two cases argv cannot serve: a `stdin-after-start`
* agent, whose CLI takes no prompt argument, and a reused terminal, whose process was already
* running before this launch existed. `false` for every failure, on the same rule the structured
* The other half of `startupPrompt`, for the cases the launch command cannot serve: a
* `stdin-after-start` agent, whose CLI takes no prompt argument; a prompt the typed line cannot
* carry; and a reused terminal, whose process was already running before this launch existed. `false` for every failure, on the same rule the structured
* twin follows — a launch whose agent is running must not fail because its text did not land.
*/
deliverTerminalPrompt?(args: { handle: string; prompt: AgentLaunchPrompt }): Promise<boolean>
deliverTerminalPrompt?(args: {
handle: string
/** The launched agent, whose own readiness signal the write waits for. */
agent: TuiAgent
/** False for a reused terminal, which has no fresh launch readiness to wait for. */
freshLaunch: boolean
prompt: AgentLaunchPrompt
}): Promise<boolean>
}
/** `fence` is the lease the create was admitted at, carried so the launch prompt's send can fill its
@@ -95,8 +104,8 @@ export type AgentLaunchWorkspaceFactory = {
* wait-for-setup gate for free. A structured launch has no startup command to sequence and
* must await that gate explicitly instead. */
startupAgent: TuiAgent | undefined
/** Set only alongside a `startupAgent` whose CLI takes the prompt on argv: agent-first creation
* builds the startup command, so that is where an argv prompt belongs. */
/** Offered only alongside a `startupAgent` whose CLI takes the prompt on argv: agent-first
* creation builds the startup command, so that is where the typed line is measured. */
startupPrompt?: string
/** Inputs needed when this terminal is created as the worktree's startup surface. */
agentArgs?: string | null
@@ -112,5 +121,7 @@ export type AgentLaunchWorkspaceFactory = {
startupTerminalPaneKey?: string
/** Created, but incomplete — surfaced on the launch result rather than dropped. */
warning?: string
/** Reported by the create that built the startup command's typed line. */
promptRodeLaunchCommand?: boolean
}>
}
@@ -22,8 +22,9 @@ const LISTED: ReadonlyMap<string, { calls: number; decision: string }> = new Map
})
)
// planStartupWithPromptCandidate wraps buildAgentStartupPlan in shared/, which this scan does not reach.
const BUILDER_CALL =
/\b(?:buildAgentStartupPlan|buildAgentDraftLaunchPlan|buildAgentResumeStartupPlan)\s*\(/g
/\b(?:buildAgentStartupPlan|buildAgentDraftLaunchPlan|buildAgentResumeStartupPlan|planStartupWithPromptCandidate)\s*\(/g
const DECIDE =
'Decide attribution: a fresh agent the host builds must carry ' +
@@ -7,7 +7,6 @@
import { describe, expect, it } from 'vitest'
import {
AGENT_LAUNCH_RUNTIME_CAPABILITY,
AGENT_SESSION_BOUNDARY_RUNTIME_CAPABILITY,
AUTOMATION_CREATE_IDEMPOTENCY_RUNTIME_CAPABILITY,
AUTOMATION_OWNER_FENCING_RUNTIME_CAPABILITY,
@@ -23,6 +22,7 @@ import {
WORKTREE_VISIBILITY_SOURCE_DEFAULTS_RUNTIME_CAPABILITY,
type RuntimeCapability
} from '../../shared/protocol-version'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../shared/agent-launch-runtime-capability'
import { ELECTRON_REMOTE_RUNTIME_CLIENT_CAPABILITIES } from '../../shared/electron-remote-runtime-client-capabilities'
import { remoteRuntimeClientCapabilities } from '../../shared/remote-runtime-client-capabilities'
import { supportsAgentLaunch } from '../runtime/rpc/methods/agent-launch'
@@ -1,5 +1,4 @@
import {
AGENT_LAUNCH_RUNTIME_CAPABILITY,
AGENT_SESSION_ACCEPTED_SEND_RUNTIME_CAPABILITY,
AGENT_SESSION_BACKGROUND_TASK_ROW_STOP_CAPABILITY,
AGENT_SESSION_BACKGROUND_TASK_STOP_CAPABILITY,
@@ -10,6 +9,7 @@ import {
STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY,
type RuntimeCapability
} from '../../shared/protocol-version'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../shared/agent-launch-runtime-capability'
import { AGENT_SESSION_BACKGROUND_TASK_CHILD_VIEWS_CAPABILITY } from '../../shared/agent-session-background-task-child-views-capability'
/**
@@ -6,10 +6,8 @@
*/
import { beforeEach, describe, expect, it, vi } from 'vitest'
import {
AGENT_LAUNCH_RUNTIME_CAPABILITY,
type RuntimeCapability
} from '../../shared/protocol-version'
import type { RuntimeCapability } from '../../shared/protocol-version'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../shared/agent-launch-runtime-capability'
type AdvertisedClient = {
clientKind?: 'mobile' | 'runtime'
@@ -1,5 +1,10 @@
import type { AgentStartupPlanInputs } from '../../shared/agent-startup-plan-inputs'
import { buildAgentDraftLaunchPlan, buildAgentStartupPlan } from '../../shared/tui-agent-startup'
import {
buildAgentDraftLaunchPlan,
buildAgentStartupPlan,
type AgentStartupPlan
} from '../../shared/tui-agent-startup'
import { planStartupWithPromptCandidate } from '../../shared/startup-line-prompt-carry'
import { buildSleepingAgentLaunchConfig } from '../../shared/sleeping-agent-launch-config'
import type { SleepingAgentLaunchConfig } from '../../shared/agent-session-resume'
import type { TuiAgent } from '../../shared/tui-agent'
@@ -162,6 +167,22 @@ export async function buildExecutionHostAgentStartupPlan(
return plan
}
/** `buildExecutionHostAgentStartupPlan` for a caller that delivers an uncarried prompt itself. */
export async function planExecutionHostStartupWithPromptCandidate(
options: StartupScope & {
prompt: string
host: { shellName?: string; provesAgentInFront: boolean }
}
): Promise<{ plan: AgentStartupPlan | null; promptCarried: boolean }> {
const prepared = await prepareOpenCodeModelStartupInputs(options)
const offered = planStartupWithPromptCandidate(prepared.inputs, options.prompt, options.host)
if (offered.plan && prepared.launchConfig) {
offered.plan.launchConfig = prepared.launchConfig
offered.plan.sessionOptions = { ...options.inputs.sessionOptions }
}
return offered
}
export function assertOpenCodeModelLaunchPreferencesAbsent(
agent: TuiAgent | undefined,
preferences: Readonly<Record<string, unknown>> | undefined
@@ -0,0 +1,152 @@
/**
* Real-zsh proof for the typed launch lines `startup-line-prompt-carry.ts` lets a prompt ride: a
* multi-line line whose every line fits the per-line budget, up to the whole-line budget, reaches
* zsh intact through Orca's own wrapper, ready marker and startup write.
*
* Why a slow user config too: the write is released 1.5 s after spawn even without the marker, and
* then lands while the terminal is still line-buffered, where macOS keeps at most MAX_CANON bytes of
* one line. The per-line budget is what survives that; a single 1.1 KB line did not.
*/
import { spawnSync } from 'node:child_process'
import { existsSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs'
import { tmpdir } from 'node:os'
import { join } from 'node:path'
import * as pty from 'node-pty'
import { afterEach, beforeEach, describe, expect, it } from 'vitest'
import {
createShellStartupOutputScanState,
scanShellStartupOutput
} from '../shell-startup-output-scanner'
import { selectShellStartupFeatures } from '../shell-startup-features'
import { isBracketedPasteSafeShell } from '../../shared/startup-command-submission'
import {
TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES,
ZSH_MULTI_LINE_STARTUP_LINE_BUDGET_BYTES
} from '../../shared/startup-line-prompt-carry'
import {
restoreUserDataPathAfterEach,
setTestUserDataPath
} from './local-pty-shell-ready-test-harness'
function findZsh(): string {
if (process.platform === 'win32') {
return ''
}
return (spawnSync('sh', ['-c', 'command -v zsh'], { encoding: 'utf8' }).stdout ?? '').trim()
}
const ZSH_PATH = findZsh()
/** Lines of prompt text, each `perLine` bytes. */
function promptLines(lineCount: number, perLine: number): string {
const words = 'Fix the failing checks on this branch, then explain what changed and why.'
return Array.from({ length: lineCount }, (_unused, index) => {
let line = `L${index}:`
while (line.length < perLine) {
line += ` ${words}`
}
return line.slice(0, perLine)
}).join('\n')
}
let home = ''
beforeEach(() => {
home = mkdtempSync(join(tmpdir(), 'orca-typed-line-home-'))
setTestUserDataPath(mkdtempSync(join(tmpdir(), 'orca-typed-line-ud-')))
})
afterEach(() => {
rmSync(home, { recursive: true, force: true, maxRetries: 5, retryDelay: 100 })
})
restoreUserDataPathAfterEach()
/** Types `printf '%s' '<text>' > out` the way a local pane types a launch line; returns what ran. */
async function typeIntoZsh(text: string, userConfigSeconds: number): Promise<string | null> {
if (userConfigSeconds > 0) {
writeFileSync(join(home, '.zshrc'), `sleep ${userConfigSeconds}\n`)
}
const out = join(home, 'out.txt')
const done = join(home, 'done.txt')
const { ensureShellReadyWrappers } = await import('./local-pty-shell-ready-wrapper-generation')
const { getShellLaunchConfig } = await import('./local-pty-shell-ready')
const { writeStartupCommandWhenShellReady, STARTUP_COMMAND_READY_MAX_WAIT_MS } =
await import('./local-pty-shell-ready-startup-command')
ensureShellReadyWrappers()
const env: Record<string, string> = { PATH: '/usr/bin:/bin', HOME: home }
const launch = getShellLaunchConfig(
ZSH_PATH,
selectShellStartupFeatures({
shellPath: ZSH_PATH,
env,
hasStartupCommand: true,
waitsForShellReady: true,
emitsStartupIdentity: false
})
)
expect(launch.supportsReadyMarker).toBe(true)
const proc = pty.spawn(ZSH_PATH, launch.args ?? [], {
name: 'xterm-256color',
cols: 200,
rows: 40,
cwd: home,
// Why ORCA_ORIG_ZDOTDIR: the user's config is read from the sandbox home, not the real one.
env: { ...env, ...launch.env, ORCA_ORIG_ZDOTDIR: home, TERM: 'xterm-256color' }
})
const scan = createShellStartupOutputScanState()
let resolveReady: ((signal: { postMarkerBytesObserved: boolean }) => void) | null = null
const ready = new Promise<{ postMarkerBytesObserved: boolean }>((resolve) => {
resolveReady = resolve
})
// The provider's cap (local-pty-shell-readiness-session.ts): a later marker releases the write.
const cap = setTimeout(() => {
resolveReady?.({ postMarkerBytesObserved: false })
resolveReady = null
}, STARTUP_COMMAND_READY_MAX_WAIT_MS)
proc.onData((data) => {
if (resolveReady && scanShellStartupOutput(scan, data).ready) {
resolveReady({ postMarkerBytesObserved: true })
resolveReady = null
}
})
writeStartupCommandWhenShellReady(
ready,
proc,
`printf '%s' '${text}' > '${out}'; printf ok > '${done}'`,
() => {},
{
bracketedPasteSafe: isBracketedPasteSafeShell({
shellName: 'zsh',
waitsForShellReady: launch.supportsReadyMarker === true
})
}
)
const deadline = Date.now() + 20_000 + userConfigSeconds * 1000
while (!existsSync(done) && Date.now() < deadline) {
await new Promise((resolve) => setTimeout(resolve, 50))
}
clearTimeout(cap)
const ran = existsSync(out) ? readFileSync(out, 'utf8') : null
const exited = new Promise<void>((resolve) => proc.onExit(() => resolve()))
proc.kill()
await Promise.race([exited, new Promise((resolve) => setTimeout(resolve, 2000))])
return ran
}
describe.skipIf(!ZSH_PATH)('a multi-line launch line typed into zsh', () => {
// Lines at the per-line budget, as many as the whole-line budget leaves room for.
const perLine = TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES - 1
const text = promptLines(
Math.floor((ZSH_MULTI_LINE_STARTUP_LINE_BUDGET_BYTES - 200) / (perLine + 1)),
perLine
)
it('arrives whole through the ready barrier', async () => {
expect(await typeIntoZsh(text, 0)).toBe(text)
}, 60_000)
it('arrives whole when the user config outlasts the barrier and the write lands early', async () => {
expect(await typeIntoZsh(text, 3)).toBe(text)
}, 60_000)
})
@@ -29,20 +29,20 @@
"title=working-spinner status=blocked-fresh screen=present fg=agent clock=clockless": "verdict=pending:closed wait=pending",
"title=working-spinner status=blocked-stale screen=present fg=agent clock=clocked": "now=working edge=working quiet=working wait=pending",
"title=working-spinner status=blocked-stale screen=present fg=agent clock=clockless": "verdict=working wait=pending",
"title=name-only status=none screen=present fg=agent clock=clocked": "now=ready-strong edge=ready-strong quiet=ready-strong wait=ready@start",
"title=name-only status=none screen=present fg=agent clock=clockless": "verdict=ready-strong wait=ready@start",
"title=name-only status=done-fresh screen=present fg=agent clock=clocked": "now=ready-strong edge=ready-strong quiet=ready-strong wait=ready@start",
"title=name-only status=done-fresh screen=present fg=agent clock=clockless": "verdict=ready-strong wait=ready@start",
"title=name-only status=done-stale screen=present fg=agent clock=clocked": "now=ready-strong edge=ready-strong quiet=ready-strong wait=ready@start",
"title=name-only status=done-stale screen=present fg=agent clock=clockless": "verdict=ready-strong wait=ready@start",
"title=name-only status=working-fresh screen=present fg=agent clock=clocked": "now=ready-strong edge=ready-strong quiet=ready-strong wait=ready@start",
"title=name-only status=working-fresh screen=present fg=agent clock=clockless": "verdict=ready-strong wait=ready@start",
"title=name-only status=working-stale screen=present fg=agent clock=clocked": "now=ready-strong edge=ready-strong quiet=ready-strong wait=ready@start",
"title=name-only status=working-stale screen=present fg=agent clock=clockless": "verdict=ready-strong wait=ready@start",
"title=name-only status=blocked-fresh screen=present fg=agent clock=clocked": "now=ready-strong edge=ready-strong quiet=ready-strong wait=ready@start",
"title=name-only status=blocked-fresh screen=present fg=agent clock=clockless": "verdict=ready-strong wait=ready@start",
"title=name-only status=blocked-stale screen=present fg=agent clock=clocked": "now=ready-strong edge=ready-strong quiet=ready-strong wait=ready@start",
"title=name-only status=blocked-stale screen=present fg=agent clock=clockless": "verdict=ready-strong wait=ready@start",
"title=name-only status=none screen=present fg=agent clock=clocked": "now=ready-weak edge=ready-weak quiet=ready-weak wait=ready@start",
"title=name-only status=none screen=present fg=agent clock=clockless": "verdict=ready-weak wait=ready@start",
"title=name-only status=done-fresh screen=present fg=agent clock=clocked": "now=ready-weak edge=ready-weak quiet=ready-weak wait=ready@start",
"title=name-only status=done-fresh screen=present fg=agent clock=clockless": "verdict=ready-weak wait=ready@start",
"title=name-only status=done-stale screen=present fg=agent clock=clocked": "now=ready-weak edge=ready-weak quiet=ready-weak wait=ready@start",
"title=name-only status=done-stale screen=present fg=agent clock=clockless": "verdict=ready-weak wait=ready@start",
"title=name-only status=working-fresh screen=present fg=agent clock=clocked": "now=working edge=working quiet=working wait=pending",
"title=name-only status=working-fresh screen=present fg=agent clock=clockless": "verdict=working wait=pending",
"title=name-only status=working-stale screen=present fg=agent clock=clocked": "now=ready-weak edge=ready-weak quiet=ready-weak wait=ready@start",
"title=name-only status=working-stale screen=present fg=agent clock=clockless": "verdict=ready-weak wait=ready@start",
"title=name-only status=blocked-fresh screen=present fg=agent clock=clocked": "now=pending:closed edge=pending:closed quiet=pending:closed wait=pending",
"title=name-only status=blocked-fresh screen=present fg=agent clock=clockless": "verdict=pending:closed wait=pending",
"title=name-only status=blocked-stale screen=present fg=agent clock=clocked": "now=ready-weak edge=ready-weak quiet=ready-weak wait=ready@start",
"title=name-only status=blocked-stale screen=present fg=agent clock=clockless": "verdict=ready-weak wait=ready@start",
"title=none status=none screen=present fg=agent clock=clocked": "now=pending:closed edge=pending:closed quiet=pending:closed wait=pending",
"title=none status=none screen=present fg=agent clock=clockless": "verdict=pending:closed wait=pending",
"title=none status=done-fresh screen=present fg=agent clock=clocked": "now=pending:closed edge=pending:closed quiet=pending:closed wait=pending",
@@ -57,16 +57,16 @@
"title=none status=blocked-fresh screen=present fg=agent clock=clockless": "verdict=pending:closed wait=pending",
"title=none status=blocked-stale screen=present fg=agent clock=clocked": "now=pending:closed edge=pending:closed quiet=pending:closed wait=pending",
"title=none status=blocked-stale screen=present fg=agent clock=clockless": "verdict=pending:closed wait=pending",
"title=name-only status=none screen=present fg=shell clock=clocked": "now=ready-strong edge=ready-strong quiet=ready-strong wait=ready@start",
"title=name-only status=none screen=present fg=shell clock=clockless": "verdict=ready-strong wait=ready@start",
"title=name-only status=none screen=untrusted fg=agent clock=clocked": "now=ready-strong edge=ready-strong quiet=ready-strong wait=ready@start",
"title=name-only status=none screen=untrusted fg=agent clock=clockless": "verdict=ready-strong wait=ready@start",
"title=name-only status=none screen=untrusted fg=shell clock=clocked": "now=ready-strong edge=ready-strong quiet=ready-strong wait=ready@start",
"title=name-only status=none screen=untrusted fg=shell clock=clockless": "verdict=ready-strong wait=ready@start",
"title=name-only status=none screen=absent fg=agent clock=clocked": "now=ready-strong edge=ready-strong quiet=ready-strong wait=ready@start",
"title=name-only status=none screen=absent fg=agent clock=clockless": "verdict=ready-strong wait=ready@start",
"title=name-only status=none screen=absent fg=shell clock=clocked": "now=ready-strong edge=ready-strong quiet=ready-strong wait=ready@start",
"title=name-only status=none screen=absent fg=shell clock=clockless": "verdict=ready-strong wait=ready@start",
"title=name-only status=none screen=present fg=shell clock=clocked": "now=ready-weak edge=ready-weak quiet=ready-weak wait=ready@start",
"title=name-only status=none screen=present fg=shell clock=clockless": "verdict=ready-weak wait=ready@start",
"title=name-only status=none screen=untrusted fg=agent clock=clocked": "now=ready-weak edge=ready-weak quiet=ready-weak wait=ready@start",
"title=name-only status=none screen=untrusted fg=agent clock=clockless": "verdict=ready-weak wait=ready@start",
"title=name-only status=none screen=untrusted fg=shell clock=clocked": "now=ready-weak edge=ready-weak quiet=ready-weak wait=ready@start",
"title=name-only status=none screen=untrusted fg=shell clock=clockless": "verdict=ready-weak wait=ready@start",
"title=name-only status=none screen=absent fg=agent clock=clocked": "now=ready-weak edge=ready-weak quiet=ready-weak wait=ready@start",
"title=name-only status=none screen=absent fg=agent clock=clockless": "verdict=ready-weak wait=ready@start",
"title=name-only status=none screen=absent fg=shell clock=clocked": "now=ready-weak edge=ready-weak quiet=ready-weak wait=ready@start",
"title=name-only status=none screen=absent fg=shell clock=clockless": "verdict=ready-weak wait=ready@start",
"title=none status=none screen=present fg=shell clock=clocked": "now=pending:closed edge=pending:closed quiet=pending:closed wait=pending",
"title=none status=none screen=present fg=shell clock=clockless": "verdict=pending:closed wait=pending",
"title=none status=none screen=untrusted fg=agent clock=clocked": "now=pending:closed edge=pending:closed quiet=pending:closed wait=pending",
@@ -0,0 +1,7 @@
{
"description": "non-agent recording at 80x24 replayed on the unknown pane, one entry per chunk",
"observations": {
"clocked": ["0-4: now=pending:open edge=pending:open quiet=pending:open wait=pending"],
"clockless": ["0-4: verdict=pending:open wait=pending"]
}
}
@@ -0,0 +1,9 @@
{
"capturedAt": "2026-09-30T12:00:00.000Z",
"platform": "darwin",
"command": ["/bin/zsh", "-f", "-i"],
"cols": 80,
"rows": 24,
"note": "zsh 5.9, macOS arm64, clean environment, PS1='%# '. The prompt turns bracketed paste on (ESC[?2004h), `sleep 1` is typed and run, which turns it off (ESC[?2004l), and the next prompt turns it on again: the bytes a launch sees from the shell before its agent starts, and after an agent exits",
"exitCode": null
}
@@ -0,0 +1,2 @@
% ␍ ␍␍% [?2004hssleep 1[?2004l␍
% ␍ ␍␍% [?2004h
@@ -0,0 +1,47 @@
import { describe, expect, it } from 'vitest'
import {
launchHostProvesAgentInFront,
nameLocalTypedLineShell
} from './agent-launch-typed-line-shell'
describe('naming the shell a local launch line is typed into', () => {
it('takes the request’s shell first, then the setting, then SHELL, as the spawn does', () => {
const base = { isRemote: false, platform: 'darwin' as const, envShell: '/bin/bash' }
expect(
nameLocalTypedLineShell({
...base,
shellOverride: '/opt/homebrew/bin/fish',
defaultShellSetting: '/bin/zsh'
})
).toBe('fish')
expect(nameLocalTypedLineShell({ ...base, defaultShellSetting: ' /bin/zsh ' })).toBe('zsh')
expect(nameLocalTypedLineShell(base)).toBe('bash')
expect(nameLocalTypedLineShell({ ...base, envShell: '' })).toBe('zsh')
})
it('names none for a remote host, whose relay picks its own login shell', () => {
expect(
nameLocalTypedLineShell({ isRemote: true, platform: 'darwin', envShell: '/bin/zsh' })
).toBeUndefined()
})
it('names none on Windows, where the pane may be cmd, PowerShell, Git Bash or WSL', () => {
expect(
nameLocalTypedLineShell({ isRemote: false, platform: 'win32', envShell: '/bin/zsh' })
).toBeUndefined()
})
})
describe('whether a launch host can prove its agent is in front before a paste', () => {
it.each([
['a local macOS host', false, 'darwin', 'darwin', true],
['a local Linux host', false, 'linux', 'linux', true],
['a local Windows host', false, 'win32', 'win32', false],
// The pane runs in the distro, but the reads run on the Windows host.
['a local WSL pane', false, 'linux', 'win32', false],
['an SSH Linux host from Windows', true, 'linux', 'win32', true],
['an SSH Windows host', true, 'win32', 'darwin', false]
] as const)('%s: %s', (_label, isRemote, launchPlatform, hostPlatform, proves) => {
expect(launchHostProvesAgentInFront({ isRemote, launchPlatform, hostPlatform })).toBe(proves)
})
})
@@ -0,0 +1,42 @@
import { basename } from 'node:path'
/**
* The shell a local pane's startup line gets typed into, named before the spawn the way the spawn
* picks it: the request's shell, the default-shell setting, then `SHELL`
* (`ipc/pty/runtime/spawn-preflight.ts`, then the local or daemon launch plan).
*
* Undefined where the host cannot name it: a remote host, whose relay picks its own login shell,
* and Windows, whose pane may be cmd, PowerShell, Git Bash or a WSL distro.
*/
export function nameLocalTypedLineShell(args: {
isRemote: boolean
shellOverride?: string
defaultShellSetting?: string
platform?: NodeJS.Platform
envShell?: string
}): string | undefined {
if (args.isRemote || (args.platform ?? process.platform) === 'win32') {
return undefined
}
const shellPath =
args.shellOverride?.trim() ||
args.defaultShellSetting?.trim() ||
(args.envShell ?? process.env.SHELL) ||
'/bin/zsh'
return basename(shellPath).toLowerCase()
}
/**
* Whether the execution host can prove a launched agent holds its terminal before a paste
* (`launched-agent-foreground`). A Windows host cannot, and a local WSL pane runs on one; an SSH
* host is judged by its own platform.
*/
export function launchHostProvesAgentInFront(args: {
isRemote: boolean
launchPlatform: NodeJS.Platform
hostPlatform?: NodeJS.Platform
}): boolean {
return args.isRemote
? args.launchPlatform !== 'win32'
: (args.hostPlatform ?? process.platform) !== 'win32'
}
@@ -1,3 +1,4 @@
import { vi } from 'vitest'
import type { TuiAgent } from '../../shared/tui-agent'
import { OrcaRuntimeService } from './orca-runtime'
import { makeStore } from './runtime-rpc-worktree-store-fixtures'
@@ -24,5 +25,7 @@ export async function createAgentPromptSubmissionRuntime(
const terminal = await runtime.createTerminal(`path:${AGENT_PROMPT_TEST_WORKTREE_PATH}`, {
launchAgent
})
// The pane holds a live agent, which its host would find in front of its terminal.
vi.spyOn(runtime, 'readLaunchedAgentForeground').mockResolvedValue('agent')
return { runtime, handle: terminal.handle, writes }
}
@@ -193,6 +193,23 @@ export function claimAgentSessionOperationInto(
return claimed.claim
}
/** Whether an admission left the right to run open, so the same transaction should claim it. */
export type ClaimAfterAdmission = (decision: AgentSessionOperationDecision) => boolean
/** Admission and, when `claimAfter` says so, the claim, in one transaction: the same swap as
* `claimAgentSessionOperationInto`, with one durable write instead of two. */
export function admitAndClaimAgentSessionOperationInto(
state: { operations: Map<string, AgentSessionOperationRow> },
args: AgentSessionOperationAdmission,
claimAfter: ClaimAfterAdmission
): { decision: AgentSessionOperationDecision; claim: AgentSessionOperationClaim | null } {
const decision = admitAgentSessionOperationInto(state, args)
return {
decision,
claim: claimAfter(decision) ? claimAgentSessionOperationInto(state, args) : null
}
}
export function settleAgentSessionOperationInto(
state: { operations: Map<string, AgentSessionOperationRow> },
args: { callerKey?: string; operationId: string; outcome: AgentSessionOperationOutcome }
@@ -21,6 +21,8 @@ import {
evaluateAgentSessionMutationOperation,
admitAgentSessionOperationInto,
claimAgentSessionOperationInto,
admitAndClaimAgentSessionOperationInto,
type ClaimAfterAdmission,
settleAgentSessionOperationInto,
type AgentSessionMutationOperationAdmission,
type AgentSessionOperationAdmission
@@ -298,6 +300,12 @@ export class AgentSessionRecordStore {
}): Promise<AgentSessionOperationClaim> =>
this.transact((draft) => claimAgentSessionOperationInto(draft, args))
/** Admission and, when `claimAfter` allows, the claim: one durable write before the effect. */
admitAndClaimOperation = (
args: AgentSessionOperationAdmission,
claimAfter: ClaimAfterAdmission
) => this.transact((draft) => admitAndClaimAgentSessionOperationInto(draft, args, claimAfter))
async recordOperationOutcome(args: AgentSessionOperationSettlement): Promise<void> {
await this.transact((draft) => settleAgentSessionOperationInto(draft, args))
}
@@ -17,7 +17,10 @@ vi.mock('electron', () => ({
app: { getPath: vi.fn(() => '/tmp') }
}))
function runtimeWithAgentLaunch(): {
// The shell the host names for a local line; bash takes no multi-line line, zsh does.
function runtimeWithAgentLaunch(
options: { terminalDefaultShell?: string; connectionId?: string } = {}
): {
runtime: OrcaRuntimeService
spawn: ReturnType<typeof vi.fn>
} {
@@ -27,11 +30,13 @@ function runtimeWithAgentLaunch(): {
store: { getSettings: () => Record<string, unknown> }
resolveTerminalWorkspaceLaunchScope: (selector: string) => Promise<unknown>
}
internal.store = { getSettings: () => ({}) }
internal.store = {
getSettings: () => ({ terminalDefaultShell: options.terminalDefaultShell ?? '/bin/bash' })
}
vi.spyOn(internal, 'resolveTerminalWorkspaceLaunchScope').mockResolvedValue({
id: 'wt-1',
path: '/repo/app',
connectionId: null,
connectionId: options.connectionId ?? null,
repo: null,
folderWorkspace: null
})
@@ -52,16 +57,80 @@ function spawnedCommand(spawn: ReturnType<typeof vi.fn>): string {
}
describe('a terminal create that is handed a launch prompt', () => {
it('folds an argv agent’s prompt into the command it spawns', async () => {
it('folds an argv agent’s prompt into the command it spawns, and says it did', async () => {
const { runtime, spawn } = runtimeWithAgentLaunch()
const onStartupPromptCarry = vi.fn()
await runtime.createTerminal('id:wt-1', {
startupAgent: 'claude',
startupPrompt: 'summarize the diff'
startupPrompt: 'summarize the diff',
onStartupPromptCarry
})
expect(spawnedCommand(spawn)).toContain('summarize the diff')
expect(spawn).toHaveBeenCalledWith(expect.objectContaining({ launchAgent: 'claude' }))
expect(onStartupPromptCarry).toHaveBeenCalledWith(true)
})
it('starts clean, and says so, when the typed line cannot carry a multi-line prompt', async () => {
const { runtime, spawn } = runtimeWithAgentLaunch()
const onStartupPromptCarry = vi.fn()
await runtime.createTerminal('id:wt-1', {
startupAgent: 'claude',
startupPrompt: 'summarize the diff\nthen list the risks',
onStartupPromptCarry
})
// Typed into a shell, each newline would be Enter; the caller pastes it once the agent is up.
expect(spawnedCommand(spawn)).toContain('claude')
expect(spawnedCommand(spawn)).not.toContain('summarize')
expect(onStartupPromptCarry).toHaveBeenCalledWith(false)
})
it('carries a short-lined multi-line prompt on a local zsh line, so the agent starts with it', async () => {
const { runtime, spawn } = runtimeWithAgentLaunch({ terminalDefaultShell: '/bin/zsh' })
const onStartupPromptCarry = vi.fn()
await runtime.createTerminal('id:wt-1', {
startupAgent: 'claude',
startupPrompt: 'summarize the diff\nthen list the risks',
onStartupPromptCarry
})
expect(spawnedCommand(spawn)).toContain('summarize the diff\nthen list the risks')
expect(onStartupPromptCarry).toHaveBeenCalledWith(true)
})
it('starts clean on a remote host, whose shell this host cannot name', async () => {
const { runtime, spawn } = runtimeWithAgentLaunch({
terminalDefaultShell: '/bin/zsh',
connectionId: 'ssh-1'
})
const onStartupPromptCarry = vi.fn()
await runtime.createTerminal('id:wt-1', {
startupAgent: 'claude',
startupPrompt: 'summarize the diff\nthen list the risks',
onStartupPromptCarry
})
expect(spawnedCommand(spawn)).not.toContain('summarize')
expect(onStartupPromptCarry).toHaveBeenCalledWith(false)
})
it('starts Hermes clean instead of refusing when its env budget cannot hold the prompt', async () => {
const { runtime, spawn } = runtimeWithAgentLaunch()
const onStartupPromptCarry = vi.fn()
await runtime.createTerminal('id:wt-1', {
startupAgent: 'hermes',
startupPrompt: 'x'.repeat(30_000),
onStartupPromptCarry
})
expect(spawn).toHaveBeenCalledTimes(1)
expect(onStartupPromptCarry).toHaveBeenCalledWith(false)
})
it('still builds a bare agent launch when no prompt is handed to it', async () => {
@@ -2,6 +2,7 @@
import { vi } from 'vitest'
import { OrcaRuntimeService } from './orca-runtime'
import type { TuiAgent } from '../../shared/tui-agent'
import type { PtyProcessInspection } from '../providers/pty-process-inspection'
const TRANSCRIPT_PANE_LEAF_ID = '11111111-1111-4111-8111-111111111111'
const TRANSCRIPT_PANE_TAB_ID = 'tab-1'
@@ -15,11 +16,24 @@ export type TranscriptPaneOptions = {
launchAgent?: TuiAgent
/** Set for a pane whose PTY lives on an SSH host or WSL distro rather than locally. */
connectionId?: string
/** The remote host of a `connectionId` pane is Windows. */
remoteWindowsHost?: boolean
/** Simulates a PTY controller whose foreground probe never settles. */
foregroundProbeHangs?: boolean
onForegroundProbe?: () => void
/** PTY grid the controller reports; the runtime's emulator otherwise defaults to 80x24. */
size?: { cols: number; rows: number }
/** What a fresh foreground scan finds, where it differs from the cached foreground read. */
confirmedForegroundProcess?: string | null
onForegroundScan?: () => void
/** What the host's process inspection answers; absent for a host without one. */
processInspection?: PtyProcessInspection
onProcessInspection?: () => void
/** What the host's shell-foreground check answers: the spawned shell holds the foreground. */
shellForegroundProven?: boolean
/** The pane's root process the provider reports; absent for a provider without an inventory. */
paneRootPid?: number
onShellForegroundProof?: () => void
}
export async function createTranscriptPane(
@@ -27,16 +41,21 @@ export async function createTranscriptPane(
runtimeDeps?: ConstructorParameters<typeof OrcaRuntimeService>[2]
): Promise<{ runtime: OrcaRuntimeService; handle: string }> {
const runtime = new OrcaRuntimeService(null, undefined, runtimeDeps)
// The runtime reads a remote pane's host OS from its worktree path.
const worktreeId = options.remoteWindowsHost
? 'repo-1::C:\\repo\\app'
: TRANSCRIPT_PANE_WORKTREE_ID
const internals = runtime as unknown as {
resolveTerminalWorkspaceLaunchScope: (selector: string) => Promise<unknown>
}
vi.spyOn(internals, 'resolveTerminalWorkspaceLaunchScope').mockResolvedValue({
id: TRANSCRIPT_PANE_WORKTREE_ID,
id: worktreeId,
path: '/repo/app',
connectionId: options.connectionId ?? null,
repo: null,
folderWorkspace: null
})
const processInspection = options.processInspection
runtime.setPtyController({
spawn: vi.fn().mockResolvedValue({ id: TRANSCRIPT_PANE_PTY_ID, incarnationId: 'inc-1' }),
write: () => true,
@@ -47,9 +66,45 @@ export async function createTranscriptPane(
return options.foregroundProbeHangs === true
? new Promise<string | null>(() => {})
: Promise.resolve(options.foregroundProcess)
}
},
...(options.confirmedForegroundProcess !== undefined
? {
confirmForegroundProcess: async () => {
options.onForegroundScan?.()
return options.confirmedForegroundProcess ?? null
}
}
: {}),
...(processInspection
? {
inspectProcess: async () => {
options.onProcessInspection?.()
return processInspection
}
}
: {}),
...(options.paneRootPid !== undefined
? {
listProcesses: async () => [
{
id: TRANSCRIPT_PANE_PTY_ID,
rootProcessId: options.paneRootPid,
cwd: '/repo/app',
title: 'Terminal'
}
]
}
: {}),
...(options.shellForegroundProven !== undefined
? {
confirmShellForeground: async () => {
options.onShellForegroundProof?.()
return options.shellForegroundProven === true
}
}
: {})
})
const terminal = await runtime.createTerminal(`id:${TRANSCRIPT_PANE_WORKTREE_ID}`, {
const terminal = await runtime.createTerminal(`id:${worktreeId}`, {
tabId: TRANSCRIPT_PANE_TAB_ID,
leafId: TRANSCRIPT_PANE_LEAF_ID,
title: 'Terminal'
@@ -59,7 +114,7 @@ export async function createTranscriptPane(
tabs: [
{
tabId: TRANSCRIPT_PANE_TAB_ID,
worktreeId: TRANSCRIPT_PANE_WORKTREE_ID,
worktreeId,
title: 'Terminal',
activeLeafId: TRANSCRIPT_PANE_LEAF_ID,
layout: null
@@ -68,7 +123,7 @@ export async function createTranscriptPane(
leaves: [
{
tabId: TRANSCRIPT_PANE_TAB_ID,
worktreeId: TRANSCRIPT_PANE_WORKTREE_ID,
worktreeId,
leafId: TRANSCRIPT_PANE_LEAF_ID,
paneRuntimeId: 1,
ptyId: TRANSCRIPT_PANE_PTY_ID,
@@ -77,7 +132,7 @@ export async function createTranscriptPane(
]
})
if (options.launchAgent) {
runtime.registerPty(TRANSCRIPT_PANE_PTY_ID, TRANSCRIPT_PANE_WORKTREE_ID, null, {
runtime.registerPty(TRANSCRIPT_PANE_PTY_ID, worktreeId, options.connectionId ?? null, {
tabId: TRANSCRIPT_PANE_TAB_ID,
leafId: TRANSCRIPT_PANE_LEAF_ID,
incarnationId: 'inc-1',
@@ -0,0 +1,386 @@
/**
* A freshly launched Claude's first input, replayed from captured transcripts
* (`__fixtures__/claude-dialog-trust-workspace*.txt`).
*
* The launch pastes on the signal the desktop's own paste used: bracketed paste turned on, then a
* quiet render. Claude's first-launch trust dialog renders in that same mode, so the quiet window
* also settles over it, and only the screen check keeps the prompt out of the dialog.
*/
import { readFileSync } from 'node:fs'
import { join } from 'node:path'
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import { createTranscriptPane, TRANSCRIPT_PANE_PTY_ID } from './agent-transcript-pane-test-harness'
import {
waitForLaunchedAgentComposer,
waitForWorkerStartComposer
} from './launched-agent-composer-readiness'
import { resolveRemoteForegroundEvidence } from '../providers/agent-foreground-process'
import type { ProcessTableRow } from '../../shared/process-table-snapshot'
import type * as TerminalForegroundGroup from './terminal-foreground-group'
// What `ps` limited to the pane's terminal answers; the verdict over it stays the real one.
const paneTerminal = vi.hoisted(() => {
const state: { rows: ProcessTableRow[] | null } = { rows: null }
return state
})
vi.mock('./terminal-foreground-group', async (importOriginal) => ({
...(await importOriginal<typeof TerminalForegroundGroup>()),
readTerminalProcessRows: vi.fn(async () => paneTerminal.rows)
}))
vi.mock('electron', () => ({
BrowserWindow: { fromId: vi.fn(() => null) },
webContents: { fromId: vi.fn(() => null) },
ipcMain: { on: vi.fn(), removeListener: vi.fn() },
app: { getPath: vi.fn(() => '/tmp') }
}))
/** The desktop paste's quiet window after bracketed paste, which the launch now shares. */
const QUIET_WINDOW_MS = 1_500
function readCapture(name: string): { data: string; size: { cols: number; rows: number } } {
const base = join(__dirname, '__fixtures__', name)
const meta: { cols: number; rows: number } = JSON.parse(readFileSync(`${base}.meta.json`, 'utf8'))
return { data: readFileSync(`${base}.txt`, 'utf8'), size: { cols: meta.cols, rows: meta.rows } }
}
/** Starts the launch wait on an empty pane, then streams the capture in, as a live launch does. */
async function launchAndStream(
name: string,
timeoutMs: number,
pane: Partial<Parameters<typeof createTranscriptPane>[0]> = {}
) {
const { data, size } = readCapture(name)
const { runtime, handle } = await createTranscriptPane({
paneTitle: 'Claude Code',
foregroundProcess: 'claude',
launchAgent: 'claude',
size,
data: '',
...pane
})
// Pane creation awaits real timers; the wait and its quiet window use the virtual clock.
vi.useFakeTimers()
const ready = waitForLaunchedAgentComposer(runtime, handle, 'claude', timeoutMs)
const settled = vi.fn()
ready.then(settled, settled)
runtime.onPtyData(TRANSCRIPT_PANE_PTY_ID, data, Date.now())
return { ready, settled }
}
describe('launch readiness for a freshly launched Claude', () => {
afterEach(() => {
vi.useRealTimers()
})
it('claude-dialog-trust-workspace-answered: reads the composer after the answered dialog as ready', async () => {
const { data } = readCapture('claude-dialog-trust-workspace-answered')
// Presence precondition: the capture turns bracketed paste on and ends on the idle composer.
expect(data).toContain('\x1b[?2004h')
const { ready, settled } = await launchAndStream(
'claude-dialog-trust-workspace-answered',
60_000
)
await vi.advanceTimersByTimeAsync(QUIET_WINDOW_MS + 100)
expect(settled).toHaveBeenCalled()
await expect(ready).resolves.toMatchObject({ satisfied: true })
})
it('claude-dialog-trust-workspace-answered: over SSH, settles on the quiet window, not the 8 s fallback', async () => {
const { ready, settled } = await launchAndStream(
'claude-dialog-trust-workspace-answered',
60_000,
// The relay offers no scan or shell check; either claiming a shell would refuse this signal.
{ connectionId: 'ssh-1', confirmedForegroundProcess: 'zsh', shellForegroundProven: true }
)
await vi.advanceTimersByTimeAsync(QUIET_WINDOW_MS + 300)
expect(settled).toHaveBeenCalled()
await expect(ready).resolves.toMatchObject({ satisfied: true })
})
it('claude-dialog-trust-workspace: never reads the trust dialog as the composer', async () => {
const { ready, settled } = await launchAndStream('claude-dialog-trust-workspace', 60_000)
// Reported inside the desktop paste's budget: the quiet window settles over the dialog, the
// screen check refuses it, and the idle wait names it.
await vi.advanceTimersByTimeAsync(4_000)
expect(settled).toHaveBeenCalled()
await expect(ready).resolves.toMatchObject({
satisfied: false,
blockedReason: 'agent-trust-workspace'
})
})
})
describe('a fresh orchestration worker start for Claude', () => {
afterEach(() => {
vi.useRealTimers()
})
// Main's worker start pasted on this cue; the launch's quiet window put the brief 1.6 s later.
it('settles on Claude’s ready title, before the launch paste’s quiet window', async () => {
const { data, size } = readCapture('claude-dialog-trust-workspace-answered')
const { runtime, handle } = await createTranscriptPane({
paneTitle: 'Claude Code',
foregroundProcess: 'claude',
launchAgent: 'claude',
size,
data: ''
})
vi.useFakeTimers()
const ready = waitForWorkerStartComposer(runtime, handle, 'claude', 60_000)
const settled = vi.fn()
ready.then(settled, settled)
runtime.onPtyData(TRANSCRIPT_PANE_PTY_ID, data, Date.now())
await vi.advanceTimersByTimeAsync(QUIET_WINDOW_MS / 2)
expect(settled).toHaveBeenCalled()
await expect(ready).resolves.toMatchObject({ satisfied: true })
})
})
describe('what holds a launched agent’s terminal', () => {
const originalPlatform = Object.getOwnPropertyDescriptor(process, 'platform')!
const setPlatform = (value: NodeJS.Platform): void => {
Object.defineProperty(process, 'platform', { configurable: true, value })
}
afterEach(() => {
Object.defineProperty(process, 'platform', originalPlatform)
})
describe('macOS and Linux: the pane terminal’s own foreground group decides', () => {
beforeEach(() => {
setPlatform('darwin')
paneTerminal.rows = null
})
// Captured with `ps -o pid=,ppid=,pgid=,tpgid=,stat=,command=` on a pane spawned the way a macOS
// pane is, under `login`: zsh holds the terminal in its own group, never the root's.
const login = (tpgid: number): ProcessTableRow => ({
pid: 60404,
ppid: 60394,
pgid: 60404,
tpgid,
stat: 'Ss',
command: '/usr/bin/login -flpq user /bin/bash --noprofile --norc -p -c'
})
const zshAt = (tpgid: number): ProcessTableRow => ({
pid: 60406,
ppid: 60404,
pgid: 60406,
tpgid,
stat: tpgid === 60406 ? 'S+' : 'S',
command: '-/bin/zsh -f'
})
it.each([
// A crashed stub's name can outlive it in the cached read; the rows show zsh back in front.
['zsh back at its prompt', [login(60406), zshAt(60406)], 'shell'],
[
'claude in front',
[
login(60500),
zshAt(60500),
{ pid: 60500, ppid: 60406, pgid: 60500, tpgid: 60500, stat: 'S+', command: 'claude' }
],
'agent'
]
] as const)(
'%s: answers from the rows, never the cached name or a whole-machine scan',
async (_label, rows, found) => {
paneTerminal.rows = [...rows]
const scan = vi.fn()
const inspection = vi.fn()
const { runtime } = await createTranscriptPane({
paneTitle: 'Claude Code',
foregroundProcess: 'python3',
confirmedForegroundProcess: 'claude',
onForegroundScan: scan,
processInspection: { foregroundProcess: 'claude', hasChildProcesses: true },
onProcessInspection: inspection,
paneRootPid: 60404,
launchAgent: 'claude',
data: ''
})
await expect(
runtime.readLaunchedAgentForeground(TRANSCRIPT_PANE_PTY_ID, 'claude')
).resolves.toBe(found)
expect(scan).not.toHaveBeenCalled()
expect(inspection).not.toHaveBeenCalled()
}
)
it.each([
['without the pane’s root process', undefined, [login(60406), zshAt(60406)]],
['when ps cannot read the pane’s terminal', 60404, null]
] as const)('proves nothing %s', async (_label, paneRootPid, rows) => {
paneTerminal.rows = rows ? [...rows] : null
const { runtime } = await createTranscriptPane({
paneTitle: 'Claude Code',
foregroundProcess: 'claude',
...(paneRootPid ? { paneRootPid } : {}),
launchAgent: 'claude',
data: ''
})
await expect(
runtime.readLaunchedAgentForeground(TRANSCRIPT_PANE_PTY_ID, 'claude')
).resolves.toBe('unknown')
})
})
// A tcsh or nu launch line runs the agent from `/bin/sh '<script>'`, which leads the terminal's
// foreground group with the agent a member of it, so the relay's name is `sh`.
describe('SSH: the relay’s process-group observation, then its name', () => {
beforeEach(() => setPlatform('darwin'))
const shLeadsClaude = (capturedAgeMs: number, agentCommand: string) => {
const row = (pid: number, ppid: number, stat: string, command: string): ProcessTableRow => ({
pid,
ppid,
pgid: 40210,
tpgid: 40210,
stat,
tty: 'pts/7',
startTime: '1790950000',
command
})
return {
foregroundProcess: 'sh',
hasChildProcesses: true,
foregroundProcessEvidence: resolveRemoteForegroundEvidence(
{ rootPid: 40100, fallbackProcess: 'sh' },
{
ptyId: TRANSCRIPT_PANE_PTY_ID,
ptyIncarnationId: 'inc-1',
authorityGeneration: 'gen-1',
observationEpoch: 1,
capturedAgeMs,
platform: 'linux'
},
[
{ ...row(40100, 40090, 'Ss', '-tcsh'), pgid: 40100 },
row(40210, 40100, 'S+', '/bin/sh /tmp/orca-launch/run.sh'),
row(40211, 40210, 'S+', agentCommand)
]
)
}
}
it.each([
[
'finds the launched agent behind the sh that leads its group',
0,
'/opt/bin/claude',
'agent'
],
// Reused from before the read was asked for, it may predate the agent's exit.
[
'takes the relay’s name over an observation too old to trust',
1_500,
'/opt/bin/claude',
'shell'
],
['takes the relay’s name when the group holds another agent', 0, '/opt/bin/codex', 'shell']
] as const)('%s', async (_label, capturedAgeMs, agentCommand, found) => {
const inspection = shLeadsClaude(capturedAgeMs, agentCommand)
// Presence precondition: the observation is live and names the agent in the group.
expect(inspection.foregroundProcessEvidence).toMatchObject({
verdict: 'live',
processName: expect.any(String)
})
const { runtime } = await createTranscriptPane({
paneTitle: 'Terminal',
foregroundProcess: 'sh',
processInspection: inspection,
connectionId: 'ssh-1',
launchAgent: 'claude',
data: ''
})
await expect(
runtime.readLaunchedAgentForeground(TRANSCRIPT_PANE_PTY_ID, 'claude')
).resolves.toBe(found)
})
})
// The scan names the pane's shell for an agent it cannot recognize (an npm agent as `node.exe`),
// and Git Bash and WSL keep other processes in the shell's job, so nothing proves the agent.
describe('Windows: only the shell-foreground check, and never the agent', () => {
beforeEach(() => setPlatform('win32'))
it.each([
['powershell.exe', false, 'unknown'],
['powershell.exe', true, 'shell'],
['claude', true, 'shell'],
['claude', false, 'unknown'],
['node', false, 'unknown']
] as const)('cached %s, shell check %s: %s', async (foregroundProcess, proven, found) => {
const scan = vi.fn()
const { runtime } = await createTranscriptPane({
paneTitle: 'Claude Code',
foregroundProcess,
confirmedForegroundProcess: 'claude',
onForegroundScan: scan,
shellForegroundProven: proven,
launchAgent: 'claude',
data: ''
})
await expect(
runtime.readLaunchedAgentForeground(TRANSCRIPT_PANE_PTY_ID, 'claude')
).resolves.toBe(found)
expect(scan).not.toHaveBeenCalled()
})
})
// The stubs always disagree with the relay's name, so asking either would flip the answer.
it.each([
['claude', 'agent'],
['bash', 'shell']
] as const)(
'SSH: takes the relay’s own read %s (%s), without a scan or check',
async (relayRead, found) => {
setPlatform('darwin')
const scan = vi.fn()
const proof = vi.fn()
const { runtime } = await createTranscriptPane({
paneTitle: 'Claude Code',
foregroundProcess: relayRead,
confirmedForegroundProcess: found === 'shell' ? 'claude' : 'zsh',
onForegroundScan: scan,
shellForegroundProven: found !== 'shell',
onShellForegroundProof: proof,
connectionId: 'ssh-1',
launchAgent: 'claude',
data: ''
})
await expect(
runtime.readLaunchedAgentForeground(TRANSCRIPT_PANE_PTY_ID, 'claude')
).resolves.toBe(found)
expect(scan).not.toHaveBeenCalled()
expect(proof).not.toHaveBeenCalled()
}
)
// Why: a Windows relay names the pane's shell for an agent its scan cannot recognize (node.exe).
it('SSH to a Windows host: never the agent', async () => {
const { runtime } = await createTranscriptPane({
paneTitle: 'Copilot',
foregroundProcess: 'node',
connectionId: 'ssh-1',
remoteWindowsHost: true,
launchAgent: 'copilot',
data: ''
})
await expect(
runtime.readLaunchedAgentForeground(TRANSCRIPT_PANE_PTY_ID, 'copilot')
).resolves.toBe('unknown')
})
})
@@ -0,0 +1,107 @@
import { describe, expect, it, vi } from 'vitest'
import { createTranscriptPane, TRANSCRIPT_PANE_PTY_ID } from './agent-transcript-pane-test-harness'
import { readRuntimeFixture } from './agent-transcript-replay-test-harness'
import { waitForLaunchedAgentComposer } from './launched-agent-composer-readiness'
vi.mock('electron', () => ({
BrowserWindow: { fromId: vi.fn(() => null) },
webContents: { fromId: vi.fn(() => null) },
ipcMain: { on: vi.fn(), removeListener: vi.fn() },
app: { getPath: vi.fn(() => '/tmp') }
}))
// A source-control button's multi-line prompt reaches Codex only through this wait, so a Codex
// release that the launch wait cannot read ready reports "prompt wasn't sent".
async function launchedCodexPane(data: string) {
return createTranscriptPane({
paneTitle: 'Terminal',
foregroundProcess: 'codex',
launchAgent: 'codex',
data,
size: { cols: 120, rows: 40 }
})
}
// Why cut: on an Astra model 0.158 sparkles braille stars in the empty composer, redrawing every
// 150 ms, then repaints it clean after 15 s. The capture ends on the first star frame, so replayed
// whole it holds a starred composer still, which the live stream never does.
function greetingBeforeStarfield(): string {
const data = readRuntimeFixture('codex-0158-fresh-home-greeting')
const firstStar = data.search(/38;2;\d+;\d+;\d+;48;2;30;30;30m[⠀-⣿]/)
return data.slice(0, data.lastIndexOf('\x1b[?2026h', firstStar))
}
describe('launch readiness for a freshly launched Codex', () => {
it('codex-0158-fresh-home-greeting: reads the live chat as ready, so the prompt is pasted', async () => {
const data = greetingBeforeStarfield()
// Presence precondition: the cut keeps the live status row and drops every star.
expect(data).toContain('GPT-6-Astra default')
expect(data).not.toMatch(/48;2;30;30;30m[⠀-⣿]/)
const { runtime, handle } = await launchedCodexPane(data)
await expect(
waitForLaunchedAgentComposer(runtime, handle, 'codex', 8_000)
).resolves.toMatchObject({ condition: 'tui-idle', satisfied: true })
}, 15_000)
it('codex-0157-plain-ready: reads the live chat as ready, so the prompt is pasted', async () => {
const { runtime, handle } = await launchedCodexPane(
readRuntimeFixture('codex-0157-plain-ready')
)
await expect(
waitForLaunchedAgentComposer(runtime, handle, 'codex', 8_000)
).resolves.toMatchObject({ condition: 'tui-idle', satisfied: true })
}, 15_000)
it.each([
['codex-0158-update-available-dialog', 'agent-update-prompt'],
['codex-0158-hooks-review-dialog', 'agent-hooks-review-prompt'],
['codex-0158-model-retired-dialog', 'codex-model-migration-prompt'],
['codex-0158-model-announcement-dialog', 'codex-model-migration-prompt']
])(
'%s: stops at the dialog as %s instead of pasting into it',
async (fixture, reason) => {
const { runtime, handle } = await launchedCodexPane(readRuntimeFixture(fixture))
await expect(
waitForLaunchedAgentComposer(runtime, handle, 'codex', 2_500)
).resolves.toMatchObject({ satisfied: false, blockedReason: reason })
},
15_000
)
// Why streamed: live, the dialog's `›` selector arrives after bracketed paste and fires the
// composer signal, so only the screen check keeps it from pasting; it hands the dialog to
// `tui-idle` to report, well inside the desktop paste's budget.
it.each([
['codex-0157-update-available-dialog', 'agent-update-prompt'],
['codex-0158-update-available-dialog', 'agent-update-prompt'],
['codex-0158-hooks-review-dialog', 'agent-hooks-review-prompt'],
['codex-0158-model-retired-dialog', 'codex-model-migration-prompt'],
['codex-0158-model-announcement-dialog', 'codex-model-migration-prompt']
])(
'%s streamed in after the launch: reports the dialog as %s, never as the composer',
async (fixture, reason) => {
const data = readRuntimeFixture(fixture)
// Presence precondition: the dialog draws Codex's composer glyph after bracketed paste.
expect(data.slice(data.indexOf('\x1b[?2004h'))).toContain('›')
const { runtime, handle } = await launchedCodexPane('')
const ready = waitForLaunchedAgentComposer(runtime, handle, 'codex', 2_500)
runtime.onPtyData(TRANSCRIPT_PANE_PTY_ID, data, Date.now())
await expect(ready).resolves.toMatchObject({ satisfied: false, blockedReason: reason })
},
15_000
)
// Why: 0.157 discards input typed behind its provisional `model: loading` screen.
it('codex-0157-fresh-home-daemon-install: does not read the provisional screen as ready', async () => {
const full = readRuntimeFixture('codex-0157-fresh-home-daemon-install')
const install = full.indexOf('Installing daemon')
// Presence precondition: the cut keeps the provisional header and stops before the live chat.
expect(install).toBeGreaterThan(0)
const provisional = full.slice(0, full.indexOf('\n', install) + 1)
expect(provisional).toMatch(/model:.*loading/)
const { runtime, handle } = await launchedCodexPane(provisional)
await expect(waitForLaunchedAgentComposer(runtime, handle, 'codex', 5_000)).rejects.toThrow(
/timeout/
)
}, 15_000)
})
@@ -1,9 +1,13 @@
import { existsSync, readFileSync } from 'node:fs'
import { join } from 'node:path'
import { describe, expect, it } from 'vitest'
import { afterEach, describe, expect, it, vi } from 'vitest'
import { createDraftPasteReadyScanner } from '../../shared/draft-paste-ready-scanner'
import type { TuiAgent } from '../../shared/tui-agent'
import { TUI_AGENT_CONFIG } from '../../shared/tui-agent-config'
import type { RuntimeTerminalWait } from '../../shared/runtime-terminal-contracts'
import { GROK_STARTUP_PTY_TRACE } from '../../shared/__fixtures__/grok-startup-pty-trace'
import type { GrokStartupTraceChunk } from '../../shared/__fixtures__/grok-startup-pty-trace'
import { GROK_INLINE_STARTUP_PTY_TRACE } from '../../shared/__fixtures__/grok-inline-startup-pty-trace'
import {
readRuntimeFixture,
readTimedRuntimeFixture,
@@ -11,8 +15,12 @@ import {
} from './agent-transcript-replay-test-harness'
import {
getLaunchedAgentReadinessLane,
type LaunchedAgentReadinessLane
waitForLaunchedAgentComposer,
workerStartReadsComposerMarker,
type LaunchedAgentReadinessLane,
type LaunchedAgentReadinessRuntime
} from './launched-agent-composer-readiness'
import { waitForWorktreeStartupDraft } from './runtime-worktree-startup-readiness'
// Why a full Record: adding a TuiAgent fails to compile here until someone decides whether its
// worker start waits for a captured input-box marker, and a row that gains or loses evidence
@@ -117,6 +125,23 @@ describe('which lane a freshly launched worker waits on', () => {
})
})
describe('which idle-lane worker starts also answer on their composer marker', () => {
// Their only rest signal is their bare name, which a launch holds to quiet output; Grok's logo
// animates for ten seconds after its composer glyph. Gemini's title needs corroboration: absent.
it('is exactly the bare-name agents whose composer draws a marker', () => {
expect(
Object.keys(EXPECTED_LANES)
.filter(isTuiAgent)
.filter(
(agent) =>
getLaunchedAgentReadinessLane(agent) === 'tui-idle' &&
workerStartReadsComposerMarker(agent)
)
.sort()
).toEqual(['dsh', 'grok', 'mimo-code'])
})
})
describe('every capture a row cites proves its input-box marker', () => {
it('cites at least the rows worker start moved', () => {
expect(CITED_CAPTURES.map(([agent]) => agent)).toEqual(
@@ -170,3 +195,147 @@ describe('every capture a row cites proves its input-box marker', () => {
}
)
})
const READY: RuntimeTerminalWait = {
handle: 'term-1',
condition: 'tui-idle',
satisfied: true,
status: 'running',
exitCode: null
}
/** The runtime's composer wait over a replayed PTY stream, with its defaults. */
function replayRuntime() {
let listener = (_data: string): void => {}
const waitForFreshWorkerComposer = vi.fn(
async (
handle: string,
agent: TuiAgent,
timeoutMs: number,
{ requireComposerMarker = true }: { requireComposerMarker?: boolean } = {}
): Promise<RuntimeTerminalWait> => {
const ptyId = await waitForWorktreeStartupDraft(
{
getPtyId: () => 'pty-1',
getForegroundProcess: async () => agent,
subscribeToData: (_ptyId, onData) => {
listener = onData
return () => {
listener = () => {}
}
},
readRecentOutput: () => undefined,
write: vi.fn()
},
handle,
agent,
{ timeoutMs, requireComposerMarker }
)
if (!ptyId) {
throw new Error('timeout')
}
return READY
}
)
/** The `tui-idle` fallback: pending until a test settles it. */
let settleIdle = (_wait: RuntimeTerminalWait): void => {}
const waitForTerminal = vi.fn(
(): Promise<RuntimeTerminalWait> =>
new Promise((resolve) => {
settleIdle = resolve
})
)
const runtime: LaunchedAgentReadinessRuntime = {
waitForTerminal,
waitForFreshWorkerComposer
}
const play = async (trace: GrokStartupTraceChunk[]): Promise<void> => {
let now = 0
for (const chunk of trace) {
await vi.advanceTimersByTimeAsync(chunk.t - now)
now = chunk.t
listener(chunk.data ?? 'x'.repeat(chunk.bytes ?? 0))
}
}
return {
runtime,
waitForTerminal,
play,
feed: (data: string) => listener(data),
settleIdle: (wait: RuntimeTerminalWait) => settleIdle(wait)
}
}
describe('launched grok composer readiness', () => {
afterEach(() => vi.useRealTimers())
it('opens on the composer frame in the default full-screen mode', async () => {
vi.useFakeTimers()
const h = replayRuntime()
const ready = waitForLaunchedAgentComposer(h.runtime, 'term-1', 'grok', 60_000)
await h.play(GROK_STARTUP_PTY_TRACE)
await expect(ready).resolves.toEqual(READY)
})
it('still opens in inline mode, which never switches to the alternate screen', async () => {
// `grok --no-alt-screen` / `screen_mode = "minimal"` paints its `❯` without the anchor the
// marker needs, so the quiet window after bracketed paste is its only readiness.
vi.useFakeTimers()
const h = replayRuntime()
const ready = waitForLaunchedAgentComposer(h.runtime, 'term-1', 'grok', 60_000)
const settled = vi.fn()
void ready.then(settled, settled)
await h.play(GROK_INLINE_STARTUP_PTY_TRACE)
await vi.advanceTimersByTimeAsync(5_000)
expect(settled).toHaveBeenCalledWith(READY)
})
it('keeps ZCode on its composer marker alone', async () => {
vi.useFakeTimers()
const h = replayRuntime()
const ready = waitForLaunchedAgentComposer(h.runtime, 'term-1', 'zcode', 1_000)
void ready.catch(() => {})
expect(h.runtime.waitForFreshWorkerComposer).toHaveBeenCalledWith('term-1', 'zcode', 1_000)
await vi.advanceTimersByTimeAsync(1_000)
await expect(ready).rejects.toThrow('timeout')
})
})
describe('launched composer readiness for agents the idle evidence can also read', () => {
afterEach(() => vi.useRealTimers())
it('pastes on the quiet window after bracketed paste, as the desktop did, without the idle evidence', async () => {
vi.useFakeTimers()
const h = replayRuntime()
const ready = waitForLaunchedAgentComposer(h.runtime, 'term-1', 'claude', 60_000)
const settled = vi.fn()
void ready.then(settled, settled)
h.feed('\x1b[?2004h\x1b[?25l welcome to claude code \x1b[?25h')
await vi.advanceTimersByTimeAsync(1_400)
expect(settled).not.toHaveBeenCalled()
await vi.advanceTimersByTimeAsync(200)
expect(settled).toHaveBeenCalledWith(READY)
expect(h.waitForTerminal).not.toHaveBeenCalled()
})
it('checks the idle evidence only once the desktop paste’s budget ran out, where it pasted blind', async () => {
vi.useFakeTimers()
const h = replayRuntime()
const ready = waitForLaunchedAgentComposer(h.runtime, 'term-1', 'claude', 60_000)
// Output, but bracketed paste never turns on, so the composer signal cannot fire.
h.feed('\x1b[?25l welcome to claude code \x1b[?25h')
await vi.advanceTimersByTimeAsync(7_900)
expect(h.waitForTerminal).not.toHaveBeenCalled()
await vi.advanceTimersByTimeAsync(200)
expect(h.waitForTerminal).toHaveBeenCalledWith('term-1', {
condition: 'tui-idle',
timeoutMs: 52_000,
launchReadiness: true
})
h.settleIdle(READY)
await expect(ready).resolves.toEqual(READY)
})
})
@@ -1,28 +1,135 @@
import type { RuntimeTerminalWait } from '../../shared/runtime-terminal-contracts'
/**
* The one answer to "has the agent Orca just launched opened its composer?", shared by every host
* path that writes a first input into a fresh agent: `agent.launch`'s terminal prompt and, through
* `waitForWorkerStartComposer`, an orchestration worker's first dispatch.
*
* The signal is the one the desktop's own paste used: bracketed paste turned on (DECSET 2004) plus
* the agent's `draftPasteReadySignal` (its composer marker, or a quiet render after 2004), read by
* the shared `draft-paste-ready-scanner`, within the same per-agent budget. That signal cannot tell
* a composer from a startup dialog drawn in the same mode, so it counts only while the pane shows no
* startup dialog and no Codex provisional header (`readFreshComposerHold`). Unlike the desktop's
* paste it reads only output after the shell's last `?2004l`, and drops a signal while a shell is
* proven in front (`readLaunchedAgentForeground`). A signal is never proof by itself: a shell back
* at its prompt turns bracketed paste on too, so the write still needs the agent found in front.
*
* Where the desktop pasted blind once its budget ran out, the host falls back to the `tui-idle`
* evidence ranking (idle titles, known ready screens), which also reports a dialog left up. Agents
* whose composer marker a captured boot proves (`composerReadyCaptures`) wait for their marker alone.
*/
import type { TuiAgent } from '../../shared/tui-agent'
import type { RuntimeTerminalWait } from '../../shared/runtime-terminal-contracts'
import { resolveDraftPasteReadyTimeoutMs } from '../../shared/draft-paste-ready-timeout'
import { TUI_AGENT_CONFIG } from '../../shared/tui-agent-config'
import { draftPasteReadySignalHasMarker } from '../../shared/draft-paste-ready-scanner'
import { nameOnlyIdleNeedsCorroboration } from './tui-idle-evidence'
import type { OrcaRuntimeService } from './orca-runtime'
import { showsHoldAnchor } from './agent-state-rules/agent-state-text-anchors'
import { detectTerminalWaitBlockedReason } from './terminal-wait-detection'
export type LaunchedAgentReadinessLane = 'composer-marker' | 'tui-idle'
export function getLaunchedAgentReadinessLane(agent: TuiAgent): LaunchedAgentReadinessLane {
return TUI_AGENT_CONFIG[agent].composerReadyCaptures?.length ? 'composer-marker' : 'tui-idle'
}
export type LaunchedAgentReadinessRuntime = Pick<
OrcaRuntimeService,
'waitForTerminal' | 'waitForFreshWorkerComposer'
>
export function getLaunchedAgentReadinessLane(agent: TuiAgent): LaunchedAgentReadinessLane {
return TUI_AGENT_CONFIG[agent].composerReadyCaptures?.length ? 'composer-marker' : 'tui-idle'
/**
* What keeps a ready signal from counting: a startup dialog in the pane's text or on its screen, or
* a rule file's hold anchor (Codex 0.157's provisional `model: loading` header, which discards input).
*/
export function readFreshComposerHold(
waitText: string,
screenLines: readonly string[] | null
): 'dialog' | 'starting' | null {
if (
detectTerminalWaitBlockedReason(waitText) !== null ||
(screenLines !== null && detectTerminalWaitBlockedReason(screenLines.join('\n')) !== null)
) {
return 'dialog'
}
return showsHoldAnchor(waitText.toLowerCase()) ? 'starting' : null
}
export function waitForLaunchedAgentComposer(
/**
* A fresh orchestration worker's first dispatch. Marker agents wait for their marker, as a launch
* does; every other agent takes main's cue for this path, `tui-idle`, which settles on the agent's
* own ready title instead of a quiet window after it. That dispatch waits for the render to settle
* before Enter, so it never needed the desktop paste's later cue.
*/
export async function waitForWorkerStartComposer(
runtime: LaunchedAgentReadinessRuntime,
handle: string,
agent: TuiAgent,
timeoutMs: number
): Promise<RuntimeTerminalWait> {
return getLaunchedAgentReadinessLane(agent) === 'composer-marker'
? runtime.waitForFreshWorkerComposer(handle, agent, timeoutMs)
: runtime.waitForTerminal(handle, { condition: 'tui-idle', timeoutMs })
if (getLaunchedAgentReadinessLane(agent) === 'composer-marker') {
return runtime.waitForFreshWorkerComposer(handle, agent, timeoutMs)
}
if (!workerStartReadsComposerMarker(agent)) {
return runtime.waitForTerminal(handle, {
condition: 'tui-idle',
timeoutMs,
launchReadiness: true
})
}
const stop = new AbortController()
try {
return await firstAnswer(
runtime.waitForTerminal(handle, {
condition: 'tui-idle',
timeoutMs,
launchReadiness: true,
signal: stop.signal
}),
runtime.waitForFreshWorkerComposer(handle, agent, timeoutMs, {
requireComposerMarker: true,
signal: stop.signal
})
)
} finally {
stop.abort()
}
}
/**
* An agent whose only rest signal is its bare name, which a launch holds to quiet output, and whose
* composer draws a marker: that marker answers first. Grok draws its glyph at 0.6 s, then animates
* its logo for ten.
*/
export function workerStartReadsComposerMarker(agent: TuiAgent): boolean {
const signal = TUI_AGENT_CONFIG[agent].draftPasteReadySignal
return (
signal !== undefined &&
draftPasteReadySignalHasMarker(signal) &&
!nameOnlyIdleNeedsCorroboration(agent)
)
}
/** The first wait to answer; a failure counts only once both failed, and then as the idle wait's. */
function firstAnswer(
idle: Promise<RuntimeTerminalWait>,
marker: Promise<RuntimeTerminalWait>
): Promise<RuntimeTerminalWait> {
return new Promise((resolve, reject) => {
let failures = 0
let idleError: unknown
const fail = (error: unknown, fromIdle: boolean): void => {
if (fromIdle) {
idleError = error
}
failures += 1
if (failures === 2) {
reject(idleError ?? error)
}
}
idle.then(resolve, (error: unknown) => fail(error, true))
marker.then(resolve, (error: unknown) => fail(error, false))
})
}
export function waitForWorkerAgentReady(
@@ -32,6 +139,40 @@ export function waitForWorkerAgentReady(
): Promise<RuntimeTerminalWait> {
// A caller-supplied terminal was not freshly launched, so its composer marker may be long gone.
return args.agent && !args.reusesTerminal
? waitForLaunchedAgentComposer(runtime, handle, args.agent, args.timeoutMs)
? waitForWorkerStartComposer(runtime, handle, args.agent, args.timeoutMs)
: runtime.waitForTerminal(handle, { condition: 'tui-idle', timeoutMs: args.timeoutMs })
}
/**
* The composer signal's wait, then — if it did not settle within the desktop paste's budget, or a
* startup dialog is up — the `tui-idle` wait for what is left of `timeoutMs`, whose result says
* ready, blocked by a dialog, or not ready. Throws when that runs out too.
*/
export async function waitForLaunchedAgentComposer(
runtime: LaunchedAgentReadinessRuntime,
handle: string,
agent: TuiAgent,
timeoutMs: number
): Promise<RuntimeTerminalWait> {
if (getLaunchedAgentReadinessLane(agent) === 'composer-marker') {
return runtime.waitForFreshWorkerComposer(handle, agent, timeoutMs)
}
const startedAt = Date.now()
try {
return await runtime.waitForFreshWorkerComposer(
handle,
agent,
Math.min(timeoutMs, resolveDraftPasteReadyTimeoutMs(agent)),
{ requireComposerMarker: false, stopOnDialog: true }
)
} catch {
// Out of budget, a dialog up, or a pane it could not read: the idle wait answers each, and
// throws for a handle that is gone.
}
// Checked, where the desktop pasted blind: an agent that shows no readiness keeps its text.
return runtime.waitForTerminal(handle, {
condition: 'tui-idle',
timeoutMs: Math.max(1, timeoutMs - (Date.now() - startedAt)),
launchReadiness: true
})
}
@@ -0,0 +1,183 @@
/**
* A launched agent that exits at startup hands the terminal back to its shell, which turns bracketed
* paste on at its next prompt just as an agent's composer does. Real zsh with a slow user config,
* spawned the way a pane is (under `login` on macOS), read through the terminal daemon's own
* foreground tracker: the launch must find the shell and write nothing, and must still find a live
* agent.
*/
import { spawnSync } from 'node:child_process'
import { chmodSync, mkdirSync, mkdtempSync, rmSync, writeFileSync } from 'node:fs'
import { tmpdir } from 'node:os'
import { join } from 'node:path'
import * as pty from 'node-pty'
import { afterEach, beforeEach, describe, expect, it } from 'vitest'
import { recognizeAgentProcess } from '../../shared/agent-process-recognition'
import { getStrictProcessTableSnapshotWithAge } from '../../shared/process-table-snapshot-reader'
import { createPtyForegroundProcessTracker } from '../daemon/pty-subprocess/foreground-process-tracker'
import { resolveRemoteForegroundEvidence } from '../providers/agent-foreground-process'
import {
prepareMacosTccLoginShell,
wrapShellSpawnForMacosTccAttribution
} from '../providers/macos-tcc-login-shell'
import { readLaunchedAgentForeground } from './launched-agent-foreground'
import { createLaunchedAgentWriteGuard } from './launched-agent-write-guard'
import { waitForWorktreeStartupDraft } from './runtime-worktree-startup-readiness'
const PTY_ID = 'pty-1'
// A function, not a ternary: the Windows-lane registration scan reads a const assigned from a
// platform check as a Windows-only gate, and this suite runs everywhere but Windows.
function findZsh(): string {
if (process.platform === 'win32') {
return ''
}
return (spawnSync('sh', ['-c', 'command -v zsh'], { encoding: 'utf8' }).stdout ?? '').trim()
}
const ZSH_PATH = findZsh()
const describeWithZsh = ZSH_PATH ? describe : describe.skip
const STUBS = {
// Exits at startup, as an agent with a bad config or a missing dependency does.
crash: "#!/bin/sh\nprintf 'stub: crashing at startup\\r\\n'\nexit 1\n",
// Opens its composer the way an agent does, then stays up.
ready: "#!/bin/sh\nprintf '\\033[?2004hstub agent\\r\\n> '\nexec sleep 30\n"
} as const
let home = ''
let savedLoginOptOut: string | undefined
beforeEach(() => {
home = mkdtempSync(join(tmpdir(), 'orca-crash-guard-'))
// The production spawn: a pane on macOS runs under `login` unless the user opted out.
savedLoginOptOut = process.env.ORCA_DISABLE_MACOS_LOGIN_SHELL
delete process.env.ORCA_DISABLE_MACOS_LOGIN_SHELL
})
afterEach(() => {
if (savedLoginOptOut !== undefined) {
process.env.ORCA_DISABLE_MACOS_LOGIN_SHELL = savedLoginOptOut
}
rmSync(home, { recursive: true, force: true, maxRetries: 5, retryDelay: 100 })
})
/** Launches the `claude` stub in a fresh zsh pane and runs the launch's own readiness wait. */
async function launchStub(stub: keyof typeof STUBS, readyTimeoutMs: number) {
const bin = join(home, 'bin')
mkdirSync(bin)
// Typed by absolute path: `login` resets HOME and PATH, so a bare name would find a real agent.
const stubPath = join(bin, 'claude')
writeFileSync(stubPath, STUBS[stub])
chmodSync(stubPath, 0o755)
writeFileSync(join(home, '.zshrc'), 'sleep 2\n')
const env = { PATH: '/usr/bin:/bin', HOME: home, ZDOTDIR: home, TERM: 'xterm-256color' }
await prepareMacosTccLoginShell()
const spawn = wrapShellSpawnForMacosTccAttribution(ZSH_PATH, ['-i'], env)
const proc = pty.spawn(spawn.file, spawn.args, { cols: 80, rows: 24, cwd: home, env })
let dead = false
proc.onExit(() => {
dead = true
})
const listeners = new Set<(data: string) => void>()
let output = ''
const tracker = createPtyForegroundProcessTracker({
process: proc,
shellPath: ZSH_PATH,
cwd: home,
sessionId: 'wt-crash-guard@@pane',
startupAgentRecognition: recognizeAgentProcess('claude'),
isDead: () => dead
})
proc.onData((data) => {
tracker.recordOutput(data)
output += data
for (const listener of listeners) {
listener(data)
}
})
// The daemon's answers, each from the same pieces the daemon uses.
const controller = {
getForegroundProcess: async () => tracker.getForegroundProcess(),
confirmForegroundProcess: () => tracker.confirmForegroundProcess(),
confirmShellForeground: () => tracker.confirmShellForeground(),
listProcesses: async () => [{ id: PTY_ID, rootProcessId: proc.pid, cwd: home, title: 'zsh' }],
inspectProcess: async () => {
const snapshot = await getStrictProcessTableSnapshotWithAge()
return {
foregroundProcess: tracker.getForegroundProcess(),
hasChildProcesses: true,
foregroundProcessEvidence: resolveRemoteForegroundEvidence(
{ rootPid: proc.pid, fallbackProcess: tracker.getForegroundProcess() },
{
ptyId: PTY_ID,
ptyIncarnationId: 'inc-1',
authorityGeneration: 'gen-1',
observationEpoch: 1,
capturedAgeMs: snapshot.capturedAgeMs,
platform: process.platform
},
snapshot.rows
)
}
}
}
const readForeground = (ptyId: string) =>
readLaunchedAgentForeground(controller, { remote: false, windows: false }, ptyId, 'claude')
const subscribe = (_ptyId: string, listener: (data: string) => void) => {
listeners.add(listener)
return () => listeners.delete(listener)
}
const ready = waitForWorktreeStartupDraft(
{
getPtyId: () => PTY_ID,
getForegroundProcess: controller.getForegroundProcess,
subscribeToData: subscribe,
readRecentOutput: () => undefined,
write: () => {}
},
'term-1',
'claude',
{
timeoutMs: readyTimeoutMs,
isShellInFront: async (ptyId) => (await readForeground(ptyId)) === 'shell'
}
)
proc.write(`'${stubPath}'\r`)
const guard = createLaunchedAgentWriteGuard(
{ readLaunchedAgentForeground: readForeground, subscribeToTerminalData: subscribe },
'claude'
)
const close = async (): Promise<void> => {
guard.dispose()
proc.kill()
await new Promise((resolve) => setTimeout(resolve, 200))
}
return { ready, guard, close, output: () => output }
}
describeWithZsh('a launch prompt after the launched agent exits at startup', () => {
it('refuses the write after the shell’s prompt turns bracketed paste on', async () => {
const launch = await launchStub('crash', 9_000)
try {
// The shell's prompt after the crash turned bracketed paste on and went quiet. A read that
// finds the shell drops that signal; one that cannot answer on a loaded host lets it settle.
// Either way only a read that finds the agent may let the text through.
await launch.ready
// Presence precondition: the stub, not anything else on the host, is what ran and exited.
expect(launch.output()).toContain('stub: crashing at startup')
await expect(launch.guard.beforeWrite(PTY_ID)).rejects.toThrow('agent_not_in_foreground')
} finally {
await launch.close()
}
}, 30_000)
it('still finds an agent that stays up, and lets the write through', async () => {
const launch = await launchStub('ready', 15_000)
try {
await expect(launch.ready).resolves.toBe(PTY_ID)
expect(launch.output()).toContain('stub agent')
await expect(launch.guard.beforeWrite(PTY_ID)).resolves.toBeUndefined()
} finally {
await launch.close()
}
}, 30_000)
})
@@ -0,0 +1,253 @@
/**
* An agent that exits at startup leaves its shell at a prompt that turns bracketed paste on, which
* is the same signal an agent's composer gives. Driven through the launch's own readiness wait,
* foreground read and write guard, with each host answering the way it does after the exit.
*/
import { readFileSync } from 'node:fs'
import { join } from 'node:path'
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import { deliverTerminalAgentLaunchPrompt } from './rpc/methods/agent-launch-terminal-prompt'
import type { ProcessTableRow } from '../../shared/process-table-snapshot'
import { readLaunchedAgentForeground } from './launched-agent-foreground'
import type * as TerminalForegroundGroup from './terminal-foreground-group'
import { waitForWorktreeStartupDraft } from './runtime-worktree-startup-readiness'
const PTY_ID = 'pty-1'
// What `ps` limited to the pane's terminal answers; the verdict over it stays the real one.
const paneTerminal = vi.hoisted(() => {
const state: { rows: ProcessTableRow[] | null } = { rows: null }
return state
})
vi.mock('./terminal-foreground-group', async (importOriginal) => ({
...(await importOriginal<typeof TerminalForegroundGroup>()),
readTerminalProcessRows: vi.fn(async () => paneTerminal.rows)
}))
/** A macOS pane under `login` whose terminal `group` holds, with `jobs` launched from its zsh. */
function loginPane(
group: number,
jobs: { pid: number; command: string }[] = []
): ProcessTableRow[] {
return [
{
pid: 100,
ppid: 1,
pgid: 100,
tpgid: group,
stat: 'Ss',
command: '/usr/bin/login -flpq user'
},
{ pid: 101, ppid: 100, pgid: 101, tpgid: group, stat: 'S', command: '-zsh' },
...jobs.map(({ pid, command }) => ({
pid,
ppid: 101,
pgid: pid,
tpgid: group,
stat: 'S+',
command
}))
]
}
const SHELL_HANDOFF = '\x1b[?2004l'
/** zsh's prompt, the launch line it runs, then its prompt again once that command exited. */
function zshRunsLaunchLineThatExits(): string[] {
const zsh = readFileSync(join(__dirname, '__fixtures__', 'zsh-prompt-runs-command.txt'), 'utf8')
const split = zsh.indexOf('\n', zsh.indexOf(SHELL_HANDOFF)) + 1
return [zsh.slice(0, split), 'stub: crashing at startup\r\n', zsh.slice(split)]
}
/** Git Bash and WSL's bash: readline turns bracketed paste on at each prompt. */
function bashRunsLaunchLineThatExits(): string[] {
return [
'\x1b[?2004hqa@host:~$ claude\r\n\x1b[?2004l\r',
'stub: crashing at startup\r\n',
'\x1b[?2004hqa@host:~$ '
]
}
type HostAnswers = {
agent?: 'claude' | 'grok'
host: { remote: boolean; windows: boolean }
cached: string | null
scanned: string | null
/** The pane terminal's processes, for a local macOS or Linux host. */
rows?: ProcessTableRow[]
shellAlone: boolean
}
const HOSTS: [string, HostAnswers, () => string[]][] = [
// The shell keeps other processes in its job, so the shell-alone check never answers.
[
'Windows Git Bash',
{
host: { remote: false, windows: true },
cached: 'claude',
scanned: 'bash',
shellAlone: false
},
bashRunsLaunchLineThatExits
],
// The QA stub: a `grok` override that exits at once.
[
'Windows Git Bash, grok',
{
agent: 'grok',
host: { remote: false, windows: true },
cached: 'node',
scanned: 'bash',
shellAlone: false
},
bashRunsLaunchLineThatExits
],
[
'Windows WSL',
{ host: { remote: false, windows: true }, cached: 'claude', scanned: 'wsl', shellAlone: false },
bashRunsLaunchLineThatExits
],
// A pane under `login`, whose cached name is still the exited stub's.
[
'macOS zsh',
{
host: { remote: false, windows: false },
cached: 'python3',
scanned: 'zsh',
rows: loginPane(101),
shellAlone: false
},
zshRunsLaunchLineThatExits
]
]
function launchedPane(answers: HostAnswers) {
paneTerminal.rows = answers.rows ?? null
const listeners = new Set<(data: string) => void>()
const subscribe = (_ptyId: string, listener: (data: string) => void) => {
listeners.add(listener)
return () => listeners.delete(listener)
}
const controller = {
getForegroundProcess: async () => answers.cached,
confirmForegroundProcess: async () => answers.scanned,
confirmShellForeground: async () => answers.shellAlone,
listProcesses: async () => [{ id: PTY_ID, rootProcessId: 100, cwd: '/repo', title: 'zsh' }]
}
const readForeground = (ptyId: string) =>
readLaunchedAgentForeground(controller, answers.host, ptyId, answers.agent ?? 'claude')
const writes: string[] = []
const runtime = {
waitForFreshWorkerComposer: async (
_handle: string,
agent: 'claude' | 'grok',
timeoutMs: number
) => {
const ptyId = await waitForWorktreeStartupDraft(
{
getPtyId: () => PTY_ID,
getForegroundProcess: controller.getForegroundProcess,
subscribeToData: subscribe,
readRecentOutput: () => undefined,
write: () => {}
},
'term-1',
agent,
{
timeoutMs,
requireComposerMarker: false,
isShellInFront: async (id) => (await readForeground(id)) === 'shell'
}
)
if (!ptyId) {
throw new Error('timeout')
}
return { handle: 'term-1', satisfied: true, status: 'running' }
},
// The idle evidence after the budget: the shell's prompt can look settled too.
waitForTerminal: vi.fn(async () => ({ handle: 'term-1', satisfied: true, status: 'idle' })),
readLaunchedAgentForeground: (ptyId: string) => readForeground(ptyId),
subscribeToTerminalData: subscribe,
sendTerminalAgentPrompt: vi.fn(
async (
handle: string,
text: string,
options: { beforeWrite?: (id: string) => Promise<void> }
) => {
await options.beforeWrite?.(PTY_ID)
writes.push(text)
return { handle, accepted: true, bytesWritten: text.length }
}
)
}
return {
runtime,
writes,
emit: (data: string) => {
for (const listener of listeners) {
listener(data)
}
}
}
}
describe('a launch prompt after the launched agent exits at startup', () => {
beforeEach(() => {
vi.useFakeTimers()
vi.spyOn(console, 'warn').mockImplementation(() => {})
})
afterEach(() => {
vi.useRealTimers()
vi.restoreAllMocks()
})
it.each(HOSTS)(
'%s: types nothing and reports it undelivered',
async (_label, answers, transcript) => {
const pane = launchedPane(answers)
const delivery = deliverTerminalAgentLaunchPrompt({
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the deliverer reaches only the methods built above.
runtime: pane.runtime as unknown as Parameters<
typeof deliverTerminalAgentLaunchPrompt
>[0]['runtime'],
handle: 'term-1',
agent: answers.agent ?? 'claude',
freshLaunch: true,
text: 'QA prompt that must never run as a shell command'
})
for (const chunk of transcript()) {
pane.emit(chunk)
}
await vi.advanceTimersByTimeAsync(61_000)
await expect(delivery).resolves.toBe(false)
expect(pane.writes).toEqual([])
}
)
it('macOS zsh: still writes into an agent that stays up', async () => {
const pane = launchedPane({
host: { remote: false, windows: false },
cached: 'claude',
scanned: 'claude',
rows: loginPane(200, [{ pid: 200, command: '/opt/bin/claude' }]),
shellAlone: false
})
const delivery = deliverTerminalAgentLaunchPrompt({
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the deliverer reaches only the methods built above.
runtime: pane.runtime as unknown as Parameters<
typeof deliverTerminalAgentLaunchPrompt
>[0]['runtime'],
handle: 'term-1',
agent: 'claude',
freshLaunch: true,
text: 'QA prompt'
})
pane.emit(zshRunsLaunchLineThatExits()[0])
pane.emit('\x1b[?2004hClaude Code\r\n> ')
await vi.advanceTimersByTimeAsync(2_000)
await expect(delivery).resolves.toBe(true)
expect(pane.writes).toEqual(['QA prompt'])
})
})
@@ -0,0 +1,114 @@
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import type { RemoteForegroundEvidence } from '../../shared/foreground-process-evidence'
import type { ProcessTableRow } from '../../shared/process-table-snapshot'
import type * as TerminalForegroundGroup from './terminal-foreground-group'
import { readLaunchedAgentForeground } from './launched-agent-foreground'
const paneTerminal = vi.hoisted(() => {
const state: { rows: ProcessTableRow[] | null } = { rows: null }
return state
})
vi.mock('./terminal-foreground-group', async (importOriginal) => ({
...(await importOriginal<typeof TerminalForegroundGroup>()),
readTerminalProcessRows: vi.fn(async () => paneTerminal.rows)
}))
/** A live observation naming `claude` in the terminal's foreground group. */
function claudeInGroup(capturedAgeMs: number): RemoteForegroundEvidence {
return {
verdict: 'live',
processName: 'claude',
fence: {
platform: 'posix',
shellPid: 40100,
shellStartTime: 'Fri Oct 2 09:30:01 2026',
tty: 'ttys007',
foregroundPgid: 40210
},
authorityGeneration: 'gen-1',
observationEpoch: 1,
capturedAgeMs,
ptyId: 'pty-1',
ptyIncarnationId: 'inc-1'
}
}
/** The relay names the `sh` that leads the group, as it does behind a wrapper. */
function controllerAnswering(inspect: () => Promise<RemoteForegroundEvidence>) {
return {
getForegroundProcess: async () => 'sh',
confirmForegroundProcess: async () => 'sh',
confirmShellForeground: async () => false,
inspectProcess: async () => ({
foregroundProcess: 'sh',
hasChildProcesses: true,
foregroundProcessEvidence: await inspect()
})
}
}
const POSIX_SSH = { remote: true, windows: false }
describe('the relay’s foreground group as proof of a launched agent', () => {
beforeEach(() => {
vi.useFakeTimers()
})
afterEach(() => {
vi.useRealTimers()
})
// On a loaded host the whole-machine capture can take seconds; it still describes the moment
// after the read was asked for.
it('counts a capture begun after the read was asked for, however long `ps` took', async () => {
const controller = controllerAnswering(async () => {
vi.advanceTimersByTime(1_800)
return claudeInGroup(1_800)
})
await expect(
readLaunchedAgentForeground(controller, POSIX_SSH, 'pty-1', 'claude')
).resolves.toBe('agent')
})
it('takes the relay’s name over a capture reused from before the read was asked for', async () => {
const controller = controllerAnswering(async () => claudeInGroup(1_500))
await expect(
readLaunchedAgentForeground(controller, POSIX_SSH, 'pty-1', 'claude')
).resolves.toBe('shell')
})
})
// Why: this read gates every guarded paste, so it answers from the pane's own terminal and never
// waits on a whole-machine capture, which took seconds a read on a loaded host.
describe('a local pane’s foreground read', () => {
it('answers from the pane’s own terminal while every whole-machine read is still pending', async () => {
paneTerminal.rows = [
{ pid: 100, ppid: 1, pgid: 100, tpgid: 200, stat: 'Ss', command: '-zsh' },
{ pid: 200, ppid: 100, pgid: 200, tpgid: 200, stat: 'S+', command: '/opt/bin/claude' }
]
const never = <T>(): Promise<T> => new Promise<T>(() => {})
const controller = {
getForegroundProcess: never<string | null>,
confirmForegroundProcess: never<string | null>,
confirmShellForeground: never<boolean>,
inspectProcess: never<never>,
listProcesses: async () => [{ id: 'pty-1', rootProcessId: 100, cwd: '/repo', title: 'zsh' }]
}
const settled = new Promise<string>((resolve) =>
setTimeout(() => resolve('still waiting'), 1_000)
)
await expect(
Promise.race([
readLaunchedAgentForeground(
controller,
{ remote: false, windows: false },
'pty-1',
'claude'
),
settled
])
).resolves.toBe('agent')
})
})
@@ -0,0 +1,92 @@
import {
isExpectedAgentProcess,
recognizeAgentProcess
} from '../../shared/agent-process-recognition'
import { PROCESS_TABLE_SNAPSHOT_MAX_STALENESS_MS } from '../../shared/process-table-snapshot'
import { isShellProcess } from '../../shared/shell-process-detection'
import type { TuiAgent } from '../../shared/tui-agent'
import { TUI_AGENT_CONFIG } from '../../shared/tui-agent-config'
import type { RuntimePtyController } from './runtime-pty-controller-contract'
import { judgeTerminalForeground, readTerminalProcessRows } from './terminal-foreground-group'
/**
* What holds a launched agent's terminal: the agent (anything but the pane's shell), the shell
* (the launch line has not run yet, or the agent exited), or `unknown` when the host cannot tell.
* Only `agent` lets a launch write its prompt; only `shell` drops a ready signal.
*/
export type LaunchedAgentForeground = 'agent' | 'shell' | 'unknown'
/** A login shell is reported as `-zsh`. */
function isLaunchShell(processName: string): boolean {
return isShellProcess(processName.replace(/^-/, ''))
}
function isLaunchedAgent(processName: string, agent: TuiAgent): boolean {
return (
recognizeAgentProcess(processName)?.agent === agent ||
isExpectedAgentProcess(processName, TUI_AGENT_CONFIG[agent].expectedProcess)
)
}
/**
* On a local macOS or Linux host, one `ps` limited to the pane's own terminal
* (`terminal-foreground-group`): its foreground process group decides, read fresh in a few
* milliseconds. Never the whole-machine capture behind a fresh scan or `inspectProcess`, which
* took seconds a read on a loaded host and gated every paste behind it, and never the cached name a
* tab icon uses, which can still name a process that already exited.
*
* On an SSH host, the relay's process-group observation when it names the launched agent in the
* foreground group (a wrapper that did not `exec` it leads the group), then the relay's own name,
* which it reads from the terminal when asked. The observation never proves a shell.
*
* Windows has no foreground process group, and its scan names the pane's shell for an agent it
* cannot recognize, while Git Bash and WSL keep other processes in the shell's job, so nothing
* there proves the agent. A Windows host can still prove its shell alone (`confirmShellForeground`).
*/
export async function readLaunchedAgentForeground(
controller: Pick<
RuntimePtyController,
'getForegroundProcess' | 'confirmShellForeground' | 'inspectProcess' | 'listProcesses'
> | null,
host: { remote: boolean; windows: boolean },
ptyId: string,
agent: TuiAgent
): Promise<LaunchedAgentForeground> {
if (!controller) {
return 'unknown'
}
try {
if (host.windows) {
// An SSH pane has no such check, so on a Windows relay nothing proves a shell.
return (await controller.confirmShellForeground?.(ptyId)) ? 'shell' : 'unknown'
}
if (!host.remote) {
const rootPid = (await controller.listProcesses?.(null))?.find(
(pane) => pane.id === ptyId
)?.rootProcessId
const rows = rootPid ? await readTerminalProcessRows(rootPid) : null
return rows && rootPid ? judgeTerminalForeground(rows, rootPid, agent) : 'unknown'
}
const askedAt = Date.now()
const evidence = (await controller.inspectProcess?.(ptyId))?.foregroundProcessEvidence
if (
evidence?.verdict === 'live' &&
evidence.fence.platform === 'posix' &&
Date.now() - evidence.capturedAgeMs >= askedAt - PROCESS_TABLE_SNAPSHOT_MAX_STALENESS_MS &&
evidence.processName &&
isLaunchedAgent(evidence.processName, agent)
) {
return 'agent'
}
const foreground = await controller.getForegroundProcess(ptyId)
if (!foreground) {
return 'unknown'
}
return isLaunchShell(foreground) &&
!isExpectedAgentProcess(foreground, TUI_AGENT_CONFIG[agent].expectedProcess)
? 'shell'
: 'agent'
} catch {
return 'unknown'
}
}
@@ -0,0 +1,249 @@
/**
* A launched agent's ready signal must come from the agent, not from the shell that ran it.
*
* Replayed from captures: zsh turns bracketed paste on at its prompt and off when it runs the
* typed command, then on again at its next prompt (`zsh-prompt-runs-command.txt`); Claude turns it
* on when it draws (`claude-dialog-trust-workspace-answered.txt`). Read as the agent's, the shell's
* prompt settled the quiet window before Claude had drawn anything, and after an agent that exited.
*/
import { readFileSync } from 'node:fs'
import { join } from 'node:path'
import { afterEach, describe, expect, it, vi } from 'vitest'
import type { TuiAgent } from '../../shared/tui-agent'
import { waitForWorktreeStartupDraft } from './runtime-worktree-startup-readiness'
const QUIET_WINDOW_MS = 1_500
const SHELL_HANDOFF = '\x1b[?2004l'
function readFixture(name: string): string {
return readFileSync(join(__dirname, '__fixtures__', `${name}.txt`), 'utf8')
}
/** The shell's prompt and the launch line it runs, up to where it hands the terminal over. */
function shellRunsLaunchLine(): string {
const zsh = readFixture('zsh-prompt-runs-command')
const handoff = zsh.indexOf(SHELL_HANDOFF)
return zsh.slice(0, zsh.indexOf('\n', handoff) + 1)
}
/** The prompt the shell draws again once the command it ran has exited. */
function shellPromptAfterCommandExits(): string {
const zsh = readFixture('zsh-prompt-runs-command')
return zsh.slice(zsh.indexOf('\n', zsh.indexOf(SHELL_HANDOFF)) + 1)
}
function launchedPane(
options: {
agent?: TuiAgent
/** Whether the check proves a shell for this foreground name; default: only `zsh` does. */
provesShell?: (name: string) => boolean
checkDelayMs?: number
} = {}
) {
let listener = (_data: string): void => {}
let foreground = 'zsh'
let reads = 0
const host = {
getPtyId: () => 'pty-1',
getForegroundProcess: async () => foreground,
subscribeToData: (_ptyId: string, onData: (data: string) => void) => {
listener = onData
return () => {
listener = () => {}
}
},
readRecentOutput: () => undefined,
write: vi.fn()
}
const provesShell = options.provesShell ?? ((name: string) => name === 'zsh')
const ready = waitForWorktreeStartupDraft(host, 'term-1', options.agent ?? 'claude', {
timeoutMs: 8_000,
isShellInFront: async () => {
reads += 1
const name = foreground
if (options.checkDelayMs) {
await new Promise((resolve) => setTimeout(resolve, options.checkDelayMs))
}
return provesShell(name)
}
})
const settled = vi.fn()
void ready.then(settled)
return {
settled,
get reads() {
return reads
},
emit: (data: string) => listener(data),
setForeground: (name: string) => {
foreground = name
}
}
}
describe('a launched agent’s ready signal, after the shell that ran it', () => {
afterEach(() => vi.useRealTimers())
it('waits for the agent’s own bracketed paste when the agent starts well after the shell’s prompt', async () => {
vi.useFakeTimers()
const pane = launchedPane()
const launchLine = shellRunsLaunchLine()
// Presence precondition: the shell enabled bracketed paste before it handed the terminal over.
expect(launchLine).toContain('\x1b[?2004h')
pane.emit(launchLine)
// The agent's process takes the terminal, then draws nothing for 3 s while it starts.
pane.setForeground('claude')
await vi.advanceTimersByTimeAsync(3_000)
expect(pane.settled).not.toHaveBeenCalled()
pane.emit(readFixture('claude-dialog-trust-workspace-answered'))
await vi.advanceTimersByTimeAsync(QUIET_WINDOW_MS - 100)
expect(pane.settled).not.toHaveBeenCalled()
await vi.advanceTimersByTimeAsync(200)
expect(pane.settled).toHaveBeenCalledWith('pty-1')
})
it('never reads the shell’s next prompt as the agent’s composer after the agent exits', async () => {
vi.useFakeTimers()
const pane = launchedPane()
pane.emit(shellRunsLaunchLine())
// The agent exits at startup: its shell is back in the foreground and draws its prompt.
pane.emit('claude: failed to start\r\n')
pane.emit(shellPromptAfterCommandExits())
await vi.advanceTimersByTimeAsync(8_000)
expect(pane.settled).toHaveBeenCalledWith(null)
// The refused signal is dropped, not retried until the budget ends.
expect(pane.reads).toBe(1)
})
// Why: a quiet agent never signals again, so only a proven shell may drop its signal. Requiring
// proof of the agent left Claude idle until the 8 s fallback wherever the read could not answer.
it('settles on the signal when the check cannot prove a shell, with one read', async () => {
vi.useFakeTimers()
const pane = launchedPane({ provesShell: () => false })
pane.emit(shellRunsLaunchLine())
pane.setForeground('2.1.285')
pane.emit(readFixture('claude-dialog-trust-workspace-answered'))
await vi.advanceTimersByTimeAsync(QUIET_WINDOW_MS)
expect(pane.settled).toHaveBeenCalledWith('pty-1')
expect(pane.reads).toBe(1)
})
it('settles a Claude that is in front the moment its signal fires, with one check', async () => {
vi.useFakeTimers()
const pane = launchedPane()
pane.emit(shellRunsLaunchLine())
pane.setForeground('2.1.285')
pane.emit(readFixture('claude-dialog-trust-workspace-answered'))
await vi.advanceTimersByTimeAsync(QUIET_WINDOW_MS)
expect(pane.settled).toHaveBeenCalledWith('pty-1')
expect(pane.reads).toBe(1)
})
it('reads nothing before the signal when the shell never enables bracketed paste', async () => {
vi.useFakeTimers()
const pane = launchedPane()
pane.emit('$ claude\r\n')
pane.setForeground('2.1.285')
await vi.advanceTimersByTimeAsync(300)
expect(pane.reads).toBe(0)
pane.emit(readFixture('claude-dialog-trust-workspace-answered'))
await vi.advanceTimersByTimeAsync(QUIET_WINDOW_MS)
expect(pane.settled).toHaveBeenCalledWith('pty-1')
expect(pane.reads).toBe(1)
})
// A startup file's subprocess (a conda or pyenv hook) in front before the prompt proves nothing
// about the agent, so the shell's prompt after it must not settle the wait.
it('does not settle while Claude is still silent after a startup-file subprocess ran', async () => {
vi.useFakeTimers()
const pane = launchedPane()
pane.setForeground('python3.11')
await vi.advanceTimersByTimeAsync(300)
pane.setForeground('zsh')
pane.emit(shellRunsLaunchLine())
pane.setForeground('2.1.285')
await vi.advanceTimersByTimeAsync(3_000)
expect(pane.settled).not.toHaveBeenCalled()
})
it('counts only what follows the last hand-off when the shell draws two prompts first', async () => {
vi.useFakeTimers()
const pane = launchedPane({ provesShell: () => false })
// A first command runs and the prompt comes back, as separate chunks, before the launch line.
pane.emit(shellRunsLaunchLine())
pane.emit(shellPromptAfterCommandExits())
await vi.advanceTimersByTimeAsync(100)
pane.emit(`claude${SHELL_HANDOFF}\r\r\n`)
pane.setForeground('2.1.285')
await vi.advanceTimersByTimeAsync(3_000)
expect(pane.settled).not.toHaveBeenCalled()
pane.emit(readFixture('claude-dialog-trust-workspace-answered'))
await vi.advanceTimersByTimeAsync(QUIET_WINDOW_MS)
expect(pane.settled).toHaveBeenCalledWith('pty-1')
})
it('sees a hand-off split across two chunks', async () => {
vi.useFakeTimers()
const pane = launchedPane({ provesShell: () => false })
const launchLine = shellRunsLaunchLine()
const cut = launchLine.indexOf(SHELL_HANDOFF) + 4
pane.emit(launchLine.slice(0, cut))
pane.emit(launchLine.slice(cut))
pane.setForeground('2.1.285')
await vi.advanceTimersByTimeAsync(3_000)
expect(pane.settled).not.toHaveBeenCalled()
pane.emit(readFixture('claude-dialog-trust-workspace-answered'))
await vi.advanceTimersByTimeAsync(QUIET_WINDOW_MS)
expect(pane.settled).toHaveBeenCalledWith('pty-1')
})
it('never settles a shell signal whose check was still running when the shell handed over', async () => {
vi.useFakeTimers()
// The check answers after the launch line ran, when the agent is in front.
const pane = launchedPane({ provesShell: () => false, checkDelayMs: 400 })
const zsh = readFixture('zsh-prompt-runs-command')
const handoff = zsh.indexOf(SHELL_HANDOFF)
// The prompt with its launch line typed, then quiet long enough to fire on the shell's 2004.
pane.emit(zsh.slice(0, handoff))
await vi.advanceTimersByTimeAsync(QUIET_WINDOW_MS + 100)
expect(pane.reads).toBe(1)
pane.emit(zsh.slice(handoff, zsh.indexOf('\n', handoff) + 1))
pane.setForeground('2.1.285')
await vi.advanceTimersByTimeAsync(3_000)
expect(pane.settled).not.toHaveBeenCalled()
})
it('settles Codex 0.157 on its marker although it turns bracketed paste off and on as it starts', async () => {
vi.useFakeTimers()
const codex = readFixture('codex-0157-plain-ready')
// Presence precondition: the capture hands bracketed paste off mid-startup.
expect(codex).toContain(SHELL_HANDOFF)
const pane = launchedPane({ agent: 'codex' })
pane.emit(shellRunsLaunchLine())
pane.setForeground('codex')
pane.emit(codex)
await vi.advanceTimersByTimeAsync(50)
expect(pane.settled).toHaveBeenCalledWith('pty-1')
})
})
@@ -0,0 +1,63 @@
/**
* A Grok worker start replayed through the runtime at its recorded read times. Grok draws its
* composer glyph at 0.6 s and then shimmers its logo until 9.9 s, and its only title is its bare
* name, so a wait that holds that title to quiet output answers ten seconds after the composer.
*/
import { afterEach, describe, expect, it, vi } from 'vitest'
import { GROK_STARTUP_PTY_TRACE } from '../../shared/__fixtures__/grok-startup-pty-trace'
import { createTranscriptPane, TRANSCRIPT_PANE_PTY_ID } from './agent-transcript-pane-test-harness'
import { waitForWorkerStartComposer } from './launched-agent-composer-readiness'
vi.mock('electron', () => ({
BrowserWindow: { fromId: vi.fn(() => null) },
webContents: { fromId: vi.fn(() => null) },
ipcMain: { on: vi.fn(), removeListener: vi.fn() },
app: { getPath: vi.fn(() => '/tmp') }
}))
const COMPOSER_FRAME_MS = GROK_STARTUP_PTY_TRACE.find((chunk) => chunk.data?.includes('❯'))?.t
const LAST_FRAME_MS = GROK_STARTUP_PTY_TRACE.at(-1)?.t ?? 0
async function replayGrokWorkerStart(): Promise<number | null> {
const { runtime, handle } = await createTranscriptPane({
paneTitle: 'Terminal',
foregroundProcess: 'grok',
launchAgent: 'grok',
size: { cols: 120, rows: 30 },
data: ''
})
vi.useFakeTimers()
const startedAt = Date.now()
let settledAt: number | null = null
void waitForWorkerStartComposer(runtime, handle, 'grok', 60_000).then(
(wait) => {
settledAt = wait.satisfied ? Date.now() - startedAt : null
},
() => {}
)
for (const chunk of GROK_STARTUP_PTY_TRACE) {
await vi.advanceTimersByTimeAsync(Math.max(0, startedAt + chunk.t - Date.now()))
runtime.onPtyData(
TRANSCRIPT_PANE_PTY_ID,
chunk.data ?? 'x'.repeat(chunk.bytes ?? 0),
Date.now()
)
}
await vi.advanceTimersByTimeAsync(5_000)
return settledAt
}
describe('a Grok worker start', () => {
afterEach(() => vi.useRealTimers())
it('is ready on its composer glyph, not once its logo stops animating', async () => {
expect(COMPOSER_FRAME_MS).toBeLessThan(1_000)
const settledAt = await replayGrokWorkerStart()
expect(settledAt).not.toBeNull()
expect(settledAt).toBeGreaterThanOrEqual(COMPOSER_FRAME_MS ?? 0)
// Within a second of the glyph: main's own wait answered on the title at once (2.2-2.8 s live).
expect(settledAt).toBeLessThan((COMPOSER_FRAME_MS ?? 0) + 1_000)
expect(settledAt).toBeLessThan(LAST_FRAME_MS)
})
})
@@ -0,0 +1,91 @@
import { describe, expect, it, vi } from 'vitest'
import type { LaunchedAgentForeground } from './launched-agent-foreground'
import {
createLaunchedAgentWriteGuard,
type LaunchedAgentWriteGuardRuntime
} from './launched-agent-write-guard'
function guardHarness(foregrounds: LaunchedAgentForeground[]) {
const answers = [...foregrounds]
let listener: ((data: string) => void) | null = null
const unsubscribe = vi.fn(() => {
listener = null
})
const runtime: LaunchedAgentWriteGuardRuntime = {
readLaunchedAgentForeground: vi.fn(
async (): Promise<LaunchedAgentForeground> => answers.shift() ?? 'agent'
),
subscribeToTerminalData: vi.fn((_ptyId: string, next: (data: string) => void) => {
listener = next
return unsubscribe
})
}
const guard = createLaunchedAgentWriteGuard(runtime, 'claude')
return { runtime, guard, unsubscribe, emit: (data: string) => listener?.(data) }
}
describe('the check before each write of a launch prompt', () => {
it('reuses a read that found the agent for the writes after it', async () => {
const { runtime, guard } = guardHarness(['agent'])
await guard.beforeWrite('pty-1')
// The agent's own output after the paste says nothing about a shell.
await guard.beforeWrite('pty-1')
await guard.beforeWrite('pty-1')
expect(runtime.readLaunchedAgentForeground).toHaveBeenCalledTimes(1)
})
// Why: a ready signal can come from a shell whose agent exited, so only a read that finds the
// agent may let the text through; an unanswered read is not that.
it.each(['shell', 'unknown'] as const)(
'refuses a write when the read finds %s',
async (found) => {
const { guard, unsubscribe } = guardHarness([found])
await expect(guard.beforeWrite('pty-1')).rejects.toThrow('agent_not_in_foreground')
expect(unsubscribe).toHaveBeenCalledTimes(1)
}
)
it.each([
['a shell turning bracketed paste on', '$ \x1b[?2004h'],
['a shell integration prompt mark', '\x1b]133;D;1\x07\x1b]133;A\x07$ '],
['a mark split across two chunks', ['\x1b[?20', '04h']]
] as const)('reads again after %s, and refuses a shell', async (_label, output) => {
const { runtime, guard, emit } = guardHarness(['agent', 'shell'])
await guard.beforeWrite('pty-1')
for (const chunk of typeof output === 'string' ? [output] : output) {
emit(chunk)
}
await expect(guard.beforeWrite('pty-1')).rejects.toThrow('agent_not_in_foreground')
expect(runtime.readLaunchedAgentForeground).toHaveBeenCalledTimes(2)
})
it('catches a shell that returns while the first read is still running', async () => {
const { runtime, guard, emit } = guardHarness([])
vi.mocked(runtime.readLaunchedAgentForeground)
.mockImplementationOnce(async () => {
emit('\x1b[?2004h')
return 'agent'
})
.mockResolvedValueOnce('shell')
await guard.beforeWrite('pty-1')
await expect(guard.beforeWrite('pty-1')).rejects.toThrow('agent_not_in_foreground')
})
it('stops watching when it refuses and when it is disposed', async () => {
const refused = guardHarness(['shell'])
await expect(refused.guard.beforeWrite('pty-1')).rejects.toThrow('agent_not_in_foreground')
expect(refused.unsubscribe).toHaveBeenCalledTimes(1)
const cleared = guardHarness(['agent'])
await cleared.guard.beforeWrite('pty-1')
cleared.guard.dispose()
expect(cleared.unsubscribe).toHaveBeenCalledTimes(1)
})
})
@@ -0,0 +1,91 @@
import type { TuiAgent } from '../../shared/tui-agent'
import { isTuiAgent } from '../../shared/tui-agent-config'
import type { OrcaRuntimeService } from './orca-runtime'
import type { LaunchedAgentForeground } from './launched-agent-foreground'
/** A shell back at its prompt turns bracketed paste on, and Orca's shell integration marks it. */
const SHELL_RETURN_MARKERS = ['\x1b[?2004', '\x1b]133;'] as const
const MARKER_CARRY_CHARS = Math.max(...SHELL_RETURN_MARKERS.map((marker) => marker.length)) - 1
export type LaunchedAgentWriteGuardRuntime = Pick<
OrcaRuntimeService,
'readLaunchedAgentForeground' | 'subscribeToTerminalData'
> &
Partial<Pick<OrcaRuntimeService, 'launchedAgentHostProvesAgent'>>
export type LaunchedAgentWriteGuard = {
beforeWrite: (ptyId: string) => Promise<void>
dispose: () => void
}
/**
* The check before each write of a launch prompt: the paste, its Enter, and Codex's second Enter.
* A write needs a fresh read that finds the agent in front; a shell, or a host that cannot tell,
* refuses it, since a ready signal alone can come from a shell back at its prompt. Once a read finds
* the agent, later writes reuse it until the terminal shows a shell coming back to its prompt, so
* Enter follows the paste on the desktop's timing instead of waiting out another process read.
*/
export function createLaunchedAgentWriteGuard(
runtime: LaunchedAgentWriteGuardRuntime,
agent: TuiAgent,
/** `write-unless-shell`: on a host that cannot find the agent in front (Windows), only a shell
* proven in front refuses, so a caller that wrote there before keeps doing so. */
{ unprovableHost = 'refuse' }: { unprovableHost?: 'refuse' | 'write-unless-shell' } = {}
): LaunchedAgentWriteGuard {
let cleared: { ptyId: string; shellMayHaveReturned: boolean; unsubscribe: () => void } | null =
null
const dispose = (): void => {
cleared?.unsubscribe()
cleared = null
}
const beforeWrite = async (ptyId: string): Promise<void> => {
if (cleared?.ptyId === ptyId && !cleared.shellMayHaveReturned) {
return
}
dispose()
let carry = ''
const watch = { ptyId, shellMayHaveReturned: false, unsubscribe: (): void => {} }
// Subscribed before the read, so a shell that returns while it runs is not missed.
watch.unsubscribe = runtime.subscribeToTerminalData(ptyId, (data) => {
const window = carry + data
carry = window.slice(-MARKER_CARRY_CHARS)
if (SHELL_RETURN_MARKERS.some((marker) => window.includes(marker))) {
watch.shellMayHaveReturned = true
}
})
let foreground: LaunchedAgentForeground
try {
foreground = await runtime.readLaunchedAgentForeground(ptyId, agent)
} catch (error) {
watch.unsubscribe()
throw error
}
if (foreground === 'agent') {
cleared = watch
return
}
watch.unsubscribe()
if (
foreground === 'shell' ||
unprovableHost === 'refuse' ||
runtime.launchedAgentHostProvesAgent?.(ptyId) !== false
) {
throw new Error('agent_not_in_foreground')
}
}
return { beforeWrite, dispose }
}
/**
* The check before a worker start writes its brief into the agent it just launched, as a launch
* prompt's write is checked; null for a terminal the caller supplied, which no launch put an agent in.
*/
export function createWorkerBriefWriteGuard(
runtime: LaunchedAgentWriteGuardRuntime,
agent: string | null | undefined,
freshLaunch: boolean
): LaunchedAgentWriteGuard | null {
return freshLaunch && isTuiAgent(agent)
? createLaunchedAgentWriteGuard(runtime, agent, { unprovableHost: 'write-unless-shell' })
: null
}
@@ -31,6 +31,12 @@ import {
} from './runtime-worktree-startup-readiness'
import type { CreateWorktreeResult } from '../../shared/worktree/create-types'
import { provisionWorktreeTerminals } from './runtime-worktree-terminal-provisioning'
import { readFreshComposerHold } from './launched-agent-composer-readiness'
import { buildTerminalWaitText } from './terminal-wait-tail-state'
import {
readLaunchedAgentForeground,
type LaunchedAgentForeground
} from './launched-agent-foreground'
export class OrcaRuntimeWithActivateManagedWorktree extends OrcaRuntimeWithListManagedWorktrees {
async activateManagedWorktree(
@@ -145,6 +151,7 @@ export class OrcaRuntimeWithActivateManagedWorktree extends OrcaRuntimeWithListM
launchInputs?: {
agentArgs?: string | null
launchSource?: string
onPromptCarry?: (carried: boolean) => void
}
): { agent: TuiAgent; startup: WorktreeStartupLaunch; followup?: WorktreeStartupFollowup } {
if (!this.store) {
@@ -157,6 +164,7 @@ export class OrcaRuntimeWithActivateManagedWorktree extends OrcaRuntimeWithListM
...(launchPreferences ? { launchPreferences } : {}),
...(launchInputs?.agentArgs !== undefined ? { agentArgs: launchInputs.agentArgs } : {}),
...(launchInputs?.launchSource ? { launchSource: launchInputs.launchSource } : {}),
...(launchInputs?.onPromptCarry ? { onPromptCarry: launchInputs.onPromptCarry } : {}),
settings: this.store.getSettings(),
getLaunchPlatform: () => this.getAgentLaunchPlatformForRepo(repo),
toSessionOptions: (preferences) => this.toAgentSessionOptions(preferences)
@@ -178,28 +186,91 @@ export class OrcaRuntimeWithActivateManagedWorktree extends OrcaRuntimeWithListM
pasteWorktreeStartupDraftWhenReady(this.getWorktreeStartupReadinessHost(), handle, draft)
}
/** Only for a newly launched worker, before its first dispatch input. */
/**
* Only for a newly launched agent, before its first input. Settles when the agent's composer
* signal fires on a screen with no startup dialog and no Codex provisional header; with
* `stopOnDialog`, a dialog ends the wait so the caller's idle wait can report it.
*/
async waitForFreshWorkerComposer(
handle: string,
agent: TuiAgent,
timeoutMs: number
timeoutMs: number,
{
requireComposerMarker = true,
stopOnDialog = false,
signal
}: { requireComposerMarker?: boolean; stopOnDialog?: boolean; signal?: AbortSignal } = {}
): Promise<RuntimeTerminalWait> {
const initialPtyId =
this.getLivePtyForHandle(handle)?.pty.ptyId ?? this.getLiveLeafForHandle(handle).leaf.ptyId
const stop = new AbortController()
const onAbort = (): void => stop.abort()
signal?.addEventListener('abort', onAbort, { once: true })
const ptyId = await waitForWorktreeStartupDraft(
{ ...this.getWorktreeStartupReadinessHost(), getPtyId: () => initialPtyId },
handle,
agent,
{ timeoutMs, requireComposerMarker: true }
{
timeoutMs,
requireComposerMarker,
signal: stop.signal,
isShellInFront: async (ownerPtyId) =>
(await this.readLaunchedAgentForeground(ownerPtyId, agent)) === 'shell',
accept: (readyPtyId) => {
const pty = this.ptysById.get(readyPtyId)
const hold = pty
? readFreshComposerHold(
buildTerminalWaitText(pty.tailBuffer, pty.tailPartialLine, pty.preview),
this.readLiveTerminalScreenLines(readyPtyId)
)
: null
if (hold === 'dialog' && stopOnDialog) {
stop.abort()
}
return hold === null
}
}
)
signal?.removeEventListener('abort', onAbort)
if (!ptyId) {
throw new Error('timeout')
throw new Error(
signal?.aborted
? 'request_aborted'
: stop.signal.aborted
? 'agent_startup_dialog'
: 'timeout'
)
}
this.assertLiveTerminalHandleTargetsPty(handle, ptyId)
if (!this.ptysById.get(ptyId)?.connected) {
throw new Error('terminal_handle_stale')
}
return { handle, condition: 'tui-idle', satisfied: true, status: 'running', exitCode: null }
return this.buildTuiIdleProbeResult(handle, null)
}
/** What holds the terminal a launch started its agent in, read fresh from the execution host. */
readLaunchedAgentForeground(ptyId: string, agent: TuiAgent): Promise<LaunchedAgentForeground> {
return readLaunchedAgentForeground(
this.ptyController,
this.launchedAgentHost(ptyId),
ptyId,
agent
)
}
/** Whether the pane's execution host can find a launched agent in front: a Windows one cannot. */
launchedAgentHostProvesAgent(ptyId: string): boolean {
return !this.launchedAgentHost(ptyId).windows
}
private launchedAgentHost(ptyId: string): { remote: boolean; windows: boolean } {
const pty = this.ptysById.get(ptyId)
const remote = !!pty?.connectionId
// A local WSL pane still runs on a Windows host, whose process reads cannot see into it.
return {
remote,
windows: remote ? this.pathFlavorForPty(pty) === 'win32' : process.platform === 'win32'
}
}
protected sendStartupFollowupWhenReady(handle: string, followup: WorktreeStartupFollowup): void {
@@ -1,5 +1,6 @@
// @ts-nocheck -- mechanically split from OrcaRuntimeService; behavior is covered by AST equivalence and characterization tests.
import { OrcaRuntimeWithResolveTerminalPane } from './orca-runtime-resolve-terminal-pane'
import { wrapTerminalBracketedPasteText } from '../../shared/terminal-bracketed-paste-text'
import { PROVEN_ABSENT_LEAF_PTY_TTL_MS } from './orca-runtime-core'
import { pruneExpiredProvenAbsentLeafPtyVerdicts } from './proven-absent-leaf-pty-verdicts'
import type { RuntimeTerminalSend } from '../../shared/runtime-types'
@@ -158,6 +159,10 @@ export class OrcaRuntimeWithControllerKnowsPtyIsLive extends OrcaRuntimeWithReso
): Promise<RuntimeTerminalSend> {
// Why the consuming agent: the foreground process reads the bytes; launchAgent covers startup.
const payloadFor = (ptyId: string): string => {
// Why: a launch prompt replaced the desktop's draft paste, so it sends that paste's bytes.
if (options.inputKind === 'launch') {
return wrapTerminalBracketedPasteText(prompt)
}
const pty = this.ptysById.get(ptyId)
const agent = pty?.foregroundAgent ?? pty?.launchAgent
return buildAgentPromptPasteBytes(
@@ -134,6 +134,7 @@ export class OrcaRuntimeWithCreateAgentPromptRenderGate extends OrcaRuntimeWithW
condition?: RuntimeTerminalWaitCondition
timeoutMs?: number
signal?: AbortSignal
launchReadiness?: boolean
}
): Promise<RuntimeTerminalWait> {
return this.terminalWait.wait(handle, options)
@@ -0,0 +1,126 @@
import { describe, expect, it, vi } from 'vitest'
import {
AGENT_PROMPT_BRACKETED_PASTE_END,
OrcaRuntimeService,
acknowledgeAgentPromptSubmit
} from '../orca-runtime-test-mocks.spec'
import { TEST_WORKTREE_PATH, store } from '../orca-runtime-test-fixtures.spec'
import { wrapTerminalBracketedPasteText } from '../../../shared/terminal-bracketed-paste-text'
// A launch's first prompt goes in right after the caller saw the agent's composer accept input, as
// the desktop's own draft paste did: Enter one turn after the paste, not after the render settles.
async function launchedAgentPane(agent: 'claude' | 'codex' | 'qwen-code') {
const writes: { data: string; at: number }[] = []
const runtime = new OrcaRuntimeService(store)
const startedAt = Date.now()
runtime.setPtyController({
spawn: vi.fn().mockResolvedValue({ id: 'pty-bg' }),
write: (_ptyId, data) => {
writes.push({ data, at: Date.now() - startedAt })
if (data.includes(AGENT_PROMPT_BRACKETED_PASTE_END)) {
// A live Claude keeps repainting after a paste, so its render never settles.
runtime.onPtyData('pty-bg', '\x1b[?25h', Date.now())
for (let delay = 500; delay <= 9_000; delay += 500) {
setTimeout(() => runtime.onPtyData('pty-bg', `frame ${delay}`, Date.now()), delay)
}
}
acknowledgeAgentPromptSubmit(runtime, 'pty-bg', data)
return true
},
kill: () => true,
getForegroundProcess: async () => null
})
const { handle } = await runtime.createTerminal(`path:${TEST_WORKTREE_PATH}`, {
launchAgent: agent
})
return { runtime, handle, writes }
}
describe('a launch prompt into a composer the caller saw ready', () => {
it('pastes the same bytes the desktop draft paste sent for the prompt', async () => {
vi.useFakeTimers()
try {
const { runtime, handle, writes } = await launchedAgentPane('claude')
const prompt = 'Explain this commit:\n- first line\r\n- second line with \x1b[31m colour'
const send = runtime.sendTerminalAgentPrompt(handle, prompt, {
inputKind: 'launch',
composerReady: true
})
await vi.advanceTimersByTimeAsync(100)
await send
const pasted = writes
.map((write) => write.data)
.filter((data) => data !== '\r')
.join('')
expect(pasted).toBe(wrapTerminalBracketedPasteText(prompt))
// Main's draft paste: every line break as CR, and an embedded ESC made inert.
expect(pasted).toBe(
'\x1b[200~Explain this commit:\r- first line\r- second line with \u241b[31m colour\x1b[201~'
)
} finally {
vi.useRealTimers()
}
})
it('submits one turn after the paste instead of waiting out the render cap', async () => {
vi.useFakeTimers()
try {
const { runtime, handle, writes } = await launchedAgentPane('claude')
const send = runtime.sendTerminalAgentPrompt(handle, 'fix the checks\nlog tail', {
inputKind: 'launch',
composerReady: true
})
await vi.advanceTimersByTimeAsync(49)
expect(writes.map((write) => write.data)).not.toContain('\r')
await vi.advanceTimersByTimeAsync(2)
await send
// One Enter: Claude takes no second one.
expect(writes.filter((write) => write.data === '\r')).toHaveLength(1)
} finally {
vi.useRealTimers()
}
})
it.each(['codex', 'qwen-code'] as const)(
'sends %s its second Enter after its retry delay, as the desktop paste did',
async (agent) => {
vi.useFakeTimers()
try {
const { runtime, handle, writes } = await launchedAgentPane(agent)
const send = runtime.sendTerminalAgentPrompt(handle, 'explain this commit', {
inputKind: 'launch',
composerReady: true
})
await vi.advanceTimersByTimeAsync(51)
expect(writes.filter((write) => write.data === '\r')).toHaveLength(1)
await vi.advanceTimersByTimeAsync(1_200)
await send
expect(writes.filter((write) => write.data === '\r')).toHaveLength(2)
} finally {
vi.useRealTimers()
}
}
)
it('keeps a prompt into a running agent behind the render settle, as before', async () => {
vi.useFakeTimers()
try {
const { runtime, handle, writes } = await launchedAgentPane('claude')
const send = runtime.sendTerminalAgentPrompt(handle, 'fix the checks', {
inputKind: 'driving'
})
await vi.advanceTimersByTimeAsync(7_000)
expect(writes.map((write) => write.data)).not.toContain('\r')
await vi.advanceTimersByTimeAsync(2_000)
await send
expect(writes.filter((write) => write.data === '\r')).toHaveLength(1)
} finally {
vi.useRealTimers()
}
})
})
@@ -8,6 +8,7 @@ import {
waitForAgentPromptPromise
} from './orca-runtime-core'
import {
AGENT_PROMPT_POST_PASTE_SUBMIT_DELAY_MS,
AGENT_PROMPT_SUBMIT,
agentPromptSubmitJoinsPasteFrame,
getAgentPromptSubmitDelayMs,
@@ -20,6 +21,7 @@ import {
resolveAgentPromptEffectTimeoutMs,
verifyAgentPromptSubmission
} from './agent-prompt-submission-verification'
import { TUI_AGENT_CONFIG } from '../../shared/tui-agent-config'
export class OrcaRuntimeWithWriteTerminalAgentPrompt extends OrcaRuntimeWithResolveAuthoritativeTerminalWaitPermission {
protected async writeTerminalAgentPrompt(
@@ -43,7 +45,11 @@ export class OrcaRuntimeWithWriteTerminalAgentPrompt extends OrcaRuntimeWithReso
)
const pasteByteLength = Buffer.byteLength(pastePayload, 'utf8')
const pasteIngestMs = getTerminalPasteIngestMs(writeHostPlatform, pasteByteLength)
const renderGate = this.createAgentPromptRenderGate(ptyId, pasteIngestMs)
// Why no gate for a ready composer: a live Claude never settles it, so Enter always waited out
// its 8 s cap, where the desktop's own paste submitted in about 2 s.
const renderGate = options.composerReady
? null
: this.createAgentPromptRenderGate(ptyId, pasteIngestMs)
const waitTextCache: AgentPromptWaitTextCache = {}
const preSubmitBaseline = submitWithPaste
? this.getAgentPromptActivity(handle, ptyId, waitTextCache)
@@ -79,6 +85,11 @@ export class OrcaRuntimeWithWriteTerminalAgentPrompt extends OrcaRuntimeWithReso
} finally {
renderGate.dispose()
}
} else if (options.composerReady) {
await waitForAgentPromptDelay(
AGENT_PROMPT_POST_PASTE_SUBMIT_DELAY_MS + pasteIngestMs,
options.signal
)
} else {
const agent = this.getPtyAgent(ptyId)
const submitDelayMs = options.promptForSchedule
@@ -107,6 +118,10 @@ export class OrcaRuntimeWithWriteTerminalAgentPrompt extends OrcaRuntimeWithReso
throw new Error(options.suffixFailureError ?? 'terminal_not_writable')
}
}
const submits =
options.composerReady && !submitWithPaste
? 1 + (await this.resubmitAgentPromptAfterRetryDelay(ptyId, generation, options))
: 1
const effectTimeoutMs = resolveAgentPromptEffectTimeoutMs(this.getPtyAgent(ptyId))
if (!options.acceptQueued || !options.requestId) {
await verifyAgentPromptSubmission({
@@ -115,7 +130,7 @@ export class OrcaRuntimeWithWriteTerminalAgentPrompt extends OrcaRuntimeWithReso
timeoutMs: effectTimeoutMs,
signal: options.signal
})
return { submits: 1 }
return { submits }
}
const binding = this.getTerminalPromptRequestBinding(handle)
const foregroundAgent = this.ptysById.get(ptyId)?.foregroundAgent
@@ -147,7 +162,7 @@ export class OrcaRuntimeWithWriteTerminalAgentPrompt extends OrcaRuntimeWithReso
// receipt; they must not fail a Dispatch merely because Orca cannot prove
// submission through hooks.
if (!settlementAgent) {
return { submits: 1, prompt: inputAccepted }
return { submits, prompt: inputAccepted }
}
this.registerAgentPromptRequest(
ptyId,
@@ -175,7 +190,7 @@ export class OrcaRuntimeWithWriteTerminalAgentPrompt extends OrcaRuntimeWithReso
})
this.forgetAgentPromptRequest(ptyId, generation, options.requestId)
return {
submits: 1,
submits,
prompt: {
...inputAccepted,
stages: ['input_accepted', 'turn_started']
@@ -183,16 +198,40 @@ export class OrcaRuntimeWithWriteTerminalAgentPrompt extends OrcaRuntimeWithReso
}
} catch (error) {
if (error instanceof Error && error.message === 'agent_prompt_stalled') {
return { submits: 1, prompt: inputAccepted }
return { submits, prompt: inputAccepted }
}
if (error instanceof Error && error.message === 'agent_prompt_blocked') {
this.forgetAgentPromptRequest(ptyId, generation, options.requestId)
return {
submits: 1,
submits,
prompt: { ...inputAccepted, observation: 'permission' }
}
}
throw error
}
}
/**
* The desktop draft paste's second Enter, for agents whose composer can render before Enter is
* live (`submitRetryDelayMs`). Best effort: it never fails a prompt the first Enter submitted.
*/
private async resubmitAgentPromptAfterRetryDelay(
ptyId: string,
generation: number,
options: RuntimeAgentPromptWriteOptions
): Promise<number> {
const agent = this.getPtyAgent(ptyId)
const retryDelayMs = agent ? TUI_AGENT_CONFIG[agent]?.submitRetryDelayMs : undefined
if (retryDelayMs === undefined) {
return 0
}
try {
await waitForAgentPromptDelay(retryDelayMs, options.signal)
this.assertAgentPromptGeneration(ptyId, generation)
await options.beforeWrite?.(ptyId)
return this.ptyController?.write(ptyId, AGENT_PROMPT_SUBMIT, options.inputKind) ? 1 : 0
} catch {
return 0
}
}
}
+1
View File
@@ -47,6 +47,7 @@ await import('./orca-runtime-tests/terminal-creation-and-readiness-part-05.spec'
await import('./orca-runtime-tests/terminal-creation-and-readiness-part-06.spec')
await import('./orca-runtime-tests/terminal-creation-and-readiness-part-07.spec')
await import('./orca-runtime-tests/terminal-creation-and-readiness-part-08.spec')
await import('./orca-runtime-tests/agent-launch-prompt-submit.spec')
await import('./orca-runtime-tests/terminal-creation-and-readiness-part-09.spec')
await import('./orca-runtime-tests/terminal-creation-and-readiness-part-10.spec')
await import('./orca-runtime-tests/terminal-creation-and-readiness-part-11.spec')
+1 -1
View File
@@ -156,7 +156,7 @@ ${params.taskSpec}`
export type DispatchPreambleSendOptions = Pick<
RuntimeAgentPromptWriteOptions,
'leadLine' | 'acceptQueued' | 'observationTimeoutMs' | 'requestId' | 'inputKind'
'leadLine' | 'acceptQueued' | 'observationTimeoutMs' | 'requestId' | 'inputKind' | 'beforeWrite'
>
export function dispatchPreambleSendOptions(requestId: string): DispatchPreambleSendOptions {
@@ -79,6 +79,8 @@ export type TerminalTurnSend<TReceipt> = {
runtime: TerminalAgentTurnRuntime<TReceipt>
handle: string
turn: TerminalTurn
/** Runs before each PTY write; throwing refuses it (`createWorkerBriefWriteGuard`). */
beforeWrite?: DispatchPreambleSendOptions['beforeWrite']
}
/**
@@ -103,11 +105,10 @@ export function sendAgentTurn<TReceipt>(
return sendStructuredSessionTurn(send)
case 'terminal':
// Not async: the caller awaits the runtime's own promise.
return send.runtime.sendTerminalAgentPrompt(
send.handle,
send.turn.body,
terminalTurnOptions(send.turn)
)
return send.runtime.sendTerminalAgentPrompt(send.handle, send.turn.body, {
...terminalTurnOptions(send.turn),
...(send.beforeWrite ? { beforeWrite: send.beforeWrite } : {})
})
}
}
@@ -50,7 +50,9 @@ const RUNTIME_RECORDERS: readonly (readonly [string, Recorder])[] = [
['prime-agent-', { agent: 'prime-agent', foregroundProcess: 'prime-agent' }],
['qoder-cn-', { agent: 'qoder-cn', foregroundProcess: 'qoderclicn' }],
['qoder-', { agent: 'qoder', foregroundProcess: 'qodercli' }],
['zcode-', { agent: 'zcode', foregroundProcess: 'zcode' }]
['zcode-', { agent: 'zcode', foregroundProcess: 'zcode' }],
// A bare shell's prompt, a non-agent control like the daemon's less/nano/vim.
['zsh-', { agent: null, foregroundProcess: 'zsh' }]
]
// less, nano and vim are non-agent controls for the agent-unknown pane.
@@ -25,6 +25,14 @@ const createStructuredSession = vi.hoisted(() => vi.fn())
vi.mock('./structured-agent-session-create', () => ({
createStructuredAgentSessionForWorktree: createStructuredSession
}))
const deliverTerminalPrompt = vi.hoisted(() => vi.fn(async () => true))
vi.mock('./agent-launch-terminal-prompt', () => ({
deliverTerminalAgentLaunchPrompt: deliverTerminalPrompt
}))
const commitChatPrompt = vi.hoisted(() => vi.fn(async () => 'message-1'))
vi.mock('./agent-launch-structured-prompt', () => ({
commitStructuredAgentSessionLaunchPrompt: commitChatPrompt
}))
const { AGENT_LAUNCH_METHODS } = await import('./agent-launch')
const AGENT_LAUNCH = methodNamed(AGENT_LAUNCH_METHODS, 'agent.launch')
@@ -39,6 +47,8 @@ const CREATE_LAUNCH = {
target: { kind: 'create-worktree', create: { repo: 'id:repo-1', name: 'task' } }
}
const CALLER = 'device-1'
const REVEAL_WARNING =
'Terminal term_1 is running, but Orca could not make it discoverable. Run `orca terminal focus --terminal term_1` to reveal and focus it.'
function selectionRuntime(options: Parameters<typeof runtimeStub>[0]) {
return Object.assign(runtimeStub(options), {
@@ -59,6 +69,8 @@ function chatActivation(): unknown {
}
beforeEach(() => {
deliverTerminalPrompt.mockClear()
commitChatPrompt.mockClear()
createStructuredSession
.mockReset()
.mockResolvedValue({ ok: true, value: { sessionId: 'sess-1' } })
@@ -91,6 +103,27 @@ describe('a paired client launching into an existing workspace', () => {
)
})
// Why: a pasted prompt waits up to a minute for the agent; the caller's view must not wait with it.
it.each([
['terminal', {}, deliverTerminalPrompt],
['chat', STRUCTURED_PREFERENCE, commitChatPrompt]
] as const)(
'selects the new %s for the caller before its prompt is delivered',
async (_surface, settings, deliver) => {
const runtime = selectionRuntime({ settings, terminalPaneKey: PANE_KEY })
await launch(
{ ...EXISTING_LAUNCH, prompt: { text: 'Fix it.\nLog:', delivery: 'submit' } },
runtime
)
const selectedAt = runtime.selectCreatedMobileSessionTabForClient.mock.invocationCallOrder[0]
const deliveredAt = deliver.mock.invocationCallOrder[0]
expect(deliveredAt).toBeDefined()
expect(selectedAt).toBeLessThan(deliveredAt!)
}
)
it('still reports the launch when selecting its tab fails', async () => {
const runtime = selectionRuntime({ settings: {}, terminalPaneKey: PANE_KEY })
runtime.selectCreatedMobileSessionTabForClient.mockImplementationOnce(() => {
@@ -104,6 +137,20 @@ describe('a paired client launching into an existing workspace', () => {
warn.mockRestore()
})
// Why: a headless `--serve` host has no window to reveal into; the caller mirrors the tab anyway.
it('does not pass on the host’s own reveal warning, since the caller shows the tab itself', async () => {
const runtime = selectionRuntime({
settings: {},
terminalPaneKey: PANE_KEY,
terminalWarning: REVEAL_WARNING
})
const result = await launch(EXISTING_LAUNCH, runtime)
expect(result.outcome).toMatchObject({ kind: 'terminal', handle: 'term_1' })
expect(result.warning).toBeUndefined()
})
it('selects nothing when the runtime reported no pane for the terminal', async () => {
const runtime = selectionRuntime({ settings: {} })
@@ -123,6 +170,14 @@ describe('launches that keep the host-wide behaviour', () => {
expect(runtime.selectCreatedMobileSessionTabForClient).not.toHaveBeenCalled()
})
it('an in-process caller still hears that the host could not reveal the tab', async () => {
const runtime = selectionRuntime({ settings: {}, terminalWarning: REVEAL_WARNING })
const result = await launch(EXISTING_LAUNCH, runtime, {})
expect(result.warning).toBe(REVEAL_WARNING)
})
it("the host's own desktop window, a runtime client with no paired device, still activates the chat", async () => {
const runtime = selectionRuntime({ settings: STRUCTURED_PREFERENCE })
@@ -8,7 +8,8 @@
* what runs the new workspace's setup.
*/
import type { AgentLaunchResult, AgentLaunchTarget } from '../../../../shared/agent-launch-intent'
import type { AgentLaunchTarget } from '../../../../shared/agent-launch-intent'
import type { AgentLaunchPublishedSurface } from '../../../agent-launch/agent-launch-executor'
import { parsePaneKey } from '../../../../shared/stable-pane-id'
import type { OrcaRuntimeService } from '../../orca-runtime'
import type { RpcContext } from '../core'
@@ -27,25 +28,25 @@ export function agentLaunchCallerNavigationId(
/** Bookkeeping, never a gate: the agent already runs, so a failure here only leaves the view as it was. */
export function selectAgentLaunchTabForCaller(
runtime: Pick<OrcaRuntimeService, 'selectCreatedMobileSessionTabForClient'>,
result: AgentLaunchResult,
surface: AgentLaunchPublishedSurface,
clientNavigationId: string
): void {
const { outcome } = result
const { outcome } = surface
// Found by pane or session, never by a predicted tab id.
const surface =
const selector =
outcome.kind === 'terminal'
? outcome.paneKey
? parsePaneKey(outcome.paneKey)
: null
: { sessionId: outcome.sessionId }
if (!surface) {
if (!selector) {
return
}
try {
if (
!runtime.selectCreatedMobileSessionTabForClient(
result.worktreeId,
surface,
surface.worktreeId,
selector,
clientNavigationId
)
) {
@@ -0,0 +1,113 @@
/**
* Whether `agent.launch` pastes an argv agent's prompt is the runtime's report about the line it
* typed, not the executor's guess: a carried prompt must not be pasted a second time, and an
* uncarried one must not be left undelivered.
*/
import { describe, expect, it, vi } from 'vitest'
import {
CAPABLE_CLIENT,
methodNamed,
rpcContext,
runtimeStub,
type AgentLaunchRuntimeStub
} from './agent-launch.test-fixture'
const { AGENT_LAUNCH_METHODS } = await import('./agent-launch')
const AGENT_LAUNCH = methodNamed(AGENT_LAUNCH_METHODS, 'agent.launch')
const SUBMIT = { text: 'fix the failing checks\nlog tail follows', delivery: 'submit' }
function withPromptWriter(runtime: AgentLaunchRuntimeStub) {
const waitForTerminal = vi.fn(async () => ({ satisfied: true, status: 'idle' }))
// The idle evidence settles these launches; the composer signal never fires.
const waitForFreshWorkerComposer = vi.fn(async () => {
throw new Error('timeout')
})
const sendTerminalAgentPrompt = vi.fn(async () => ({
handle: 'term_1',
accepted: true,
bytesWritten: 1
}))
return {
runtime: Object.assign(runtime, {
waitForTerminal,
waitForFreshWorkerComposer,
sendTerminalAgentPrompt
}),
sendTerminalAgentPrompt
}
}
async function launch(params: unknown, runtime: AgentLaunchRuntimeStub) {
const parsed = AGENT_LAUNCH.params.safeParse(params)
if (!parsed.success) {
throw new Error(parsed.error.issues[0]?.message ?? 'invalid')
}
return AGENT_LAUNCH.handler(parsed.data, rpcContext(runtime, CAPABLE_CLIENT))
}
describe('an argv agent’s launch prompt, by what the runtime reports about its typed line', () => {
const EXISTING = { agent: 'claude', target: { kind: 'existing', worktree: 'id:wt-7' } }
it('pastes it once the agent is ready when the line could not carry it', async () => {
const { runtime, sendTerminalAgentPrompt } = withPromptWriter(
runtimeStub({ settings: {}, lineCarriesPrompt: false })
)
const result = await launch({ ...EXISTING, prompt: SUBMIT }, runtime)
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
expect(sendTerminalAgentPrompt).toHaveBeenCalledWith('term_1', SUBMIT.text, expect.anything())
})
it('does not paste a prompt the line already carried', async () => {
const { runtime, sendTerminalAgentPrompt } = withPromptWriter(
runtimeStub({ settings: {}, lineCarriesPrompt: true })
)
const result = await launch({ ...EXISTING, prompt: SUBMIT }, runtime)
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
expect(sendTerminalAgentPrompt).not.toHaveBeenCalled()
})
it('pastes into an agent-first create’s startup terminal when its line could not carry it', async () => {
const { runtime, sendTerminalAgentPrompt } = withPromptWriter(
runtimeStub({ settings: {}, lineCarriesPrompt: false })
)
const result = await launch(
{
agent: 'claude',
target: { kind: 'create-worktree', create: { repo: 'id:repo-1', name: 'task' } },
prompt: SUBMIT
},
runtime
)
expect(result.prompt).toEqual({ delivery: 'submit', outcome: 'handed-to-terminal' })
expect(sendTerminalAgentPrompt).toHaveBeenCalledWith(
'term_agent_first',
SUBMIT.text,
expect.anything()
)
})
it('does not paste into an agent-first create’s startup terminal whose line carried it', async () => {
const { runtime, sendTerminalAgentPrompt } = withPromptWriter(
runtimeStub({ settings: {}, lineCarriesPrompt: true })
)
await launch(
{
agent: 'claude',
target: { kind: 'create-worktree', create: { repo: 'id:repo-1', name: 'task' } },
prompt: SUBMIT
},
runtime
)
expect(sendTerminalAgentPrompt).not.toHaveBeenCalled()
})
})
@@ -12,7 +12,7 @@ import { mkdtemp, rm } from 'node:fs/promises'
import { tmpdir } from 'node:os'
import { join } from 'node:path'
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../../../shared/protocol-version'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../../../shared/agent-launch-runtime-capability'
import {
computeAgentLaunchFingerprint,
deriveAgentLaunchChildOperationId,
@@ -185,6 +185,49 @@ describe('exactly one execution per launch operation', () => {
})
})
describe('a fresh launch writes the ledger once before its effect', () => {
it('admits and claims in one durable transaction, claimed on disk when the effect starts', async () => {
const admitAndClaim = vi.spyOn(store, 'admitAndClaimOperation')
const admit = vi.spyOn(store, 'admitOperation')
const claimOnly = vi.spyOn(store, 'claimOperation')
const runtime = runtimeStub()
let statusAtEffect: string | undefined
runtime.createManagedWorktree.mockImplementationOnce(async () => {
const persisted = await readPersistedTestAgentSessionStore(directory)
statusAtEffect =
persisted.operations[agentSessionOperationKey('device-1', OPERATION_ID)]?.outcome.status
return { worktree: { id: 'wt-new' }, startupTerminal: undefined }
})
await launch(createLaunch({ operationId: OPERATION_ID }), runtime)
expect(admitAndClaim).toHaveBeenCalledTimes(1)
expect(admit).not.toHaveBeenCalled()
expect(claimOnly).not.toHaveBeenCalled()
// A claim moves the row from `pending` to `unknown`: taken, not yet settled.
expect(statusAtEffect).toBe('unknown')
})
it('still lets exactly one of two concurrent admissions run the effect', async () => {
const admission = {
callerKey: 'device-1',
operationId: OPERATION_ID,
fingerprint: 'fp-1',
now: NOW
}
const claimAfter = (decision: { decision: string; row?: AgentSessionOperationRow }) =>
decision.decision === 'admit' ||
(decision.decision === 'replay' && decision.row?.outcome.status === 'pending')
const results = await Promise.all([
store.admitAndClaimOperation(admission, claimAfter),
store.admitAndClaimOperation(admission, claimAfter)
])
expect(results.filter((result) => result.claim?.claim === 'won')).toHaveLength(1)
})
})
describe('stable replay identity', () => {
it('replays across a bearer-credential change under the paired device subject', async () => {
const params = createLaunch({ operationId: OPERATION_ID })
@@ -116,7 +116,7 @@ function answerFromRecordedRow(
}
/**
* Admit, then claim.
* Admit, then claim, in one durable transaction so a launch writes the ledger once before its effect.
*
* Two steps because they answer different questions — "is this id known and consistent?" and "may
* *I* run it?" — and the second cannot be folded into the first. Admission hands two concurrent
@@ -136,12 +136,14 @@ export async function admitAgentLaunchOperation(
}
const store = await requireLaunchOperationStore(context)
const callerKey = agentLaunchOperationCallerKey(context)
const admitted = await store.admitOperation({
callerKey,
operationId,
fingerprint,
now
})
const { decision: admitted, claim } = await store.admitAndClaimOperation(
{ callerKey, operationId, fingerprint, now },
// A fresh row, or a replayed one no one has answered yet, leaves the right to run open.
(decision) =>
decision.decision === 'admit' ||
(decision.decision === 'replay' &&
answerFromRecordedRow(operationId, decision.row.outcome) === null)
)
if (admitted.decision === 'refused') {
return refusal(operationId, admitted.code, `was refused: ${admitted.code}`)
}
@@ -151,7 +153,14 @@ export async function admitAgentLaunchOperation(
return answer
}
}
const claim = await store.claimOperation({ callerKey, operationId })
// Unreachable with both steps in one transaction; answered as uncertain rather than run twice.
if (!claim || claim.claim === 'absent') {
return refusal(
operationId,
'agent_session_operation_unknown',
'has no claim; its outcome is unknown'
)
}
if (claim.claim === 'lost') {
// The handler joins same-process retries before admission. Reaching a claimed row here means
// this runtime did not start it, so treating it as restart uncertainty is the safe answer.
@@ -160,15 +169,6 @@ export async function admitAgentLaunchOperation(
refusal(operationId, 'agent_session_operation_unknown', 'is claimed but unsettled')
)
}
if (claim.claim === 'absent') {
// Admitted a moment ago and gone already: the row cannot be re-admitted without reopening the
// duplicate-spawn window it exists to close, so this stays uncertain.
return refusal(
operationId,
'agent_session_operation_unknown',
'was pruned between admission and its claim; its outcome is unknown'
)
}
return {
decision: 'execute',
attachOperationId,
@@ -42,8 +42,8 @@ export function agentLaunchSurfaceFactory(
context: RpcContext,
attachOperationId?: string,
operationCallerKey?: string,
// False when the launch selects the chat for its paired caller instead of for everyone.
activateChat = true,
// True when the launch shows its surface to the paired caller itself rather than to everyone.
callerPresentsSurface = false,
terminalSpawn: TerminalSpawnDispatch = trackTerminalSpawnDispatch()
): AgentLaunchSurfaceFactory {
return {
@@ -80,7 +80,7 @@ export function agentLaunchSurfaceFactory(
...(seeded ? { options: seeded } : {}),
...(tabId ? { tabId } : {}),
// The user asked for this chat, so it takes the surface — unlike a dispatched worker.
activate: activateChat
activate: !callerPresentsSurface
})
if (!created.ok) {
// The caller named this session, so a taken id is its answer, not an opaque refusal; and not
@@ -121,13 +121,21 @@ export function agentLaunchSurfaceFactory(
options
}) => {
const launchPreferences = toAgentLaunchPreferences(options)
let promptRodeLaunchCommand = false
const created = context.runtime.createTerminal(`id:${worktreeId}`, {
// The agent id is not a shell command — `cursor` is the desktop app, its CLI is
// `cursor-agent` — so the runtime builds the configured launcher.
startupAgent: agent,
// Folded into that launcher by the same startup plan a new agent tab is built from, so an
// argv agent's prompt is in its argv at exec time rather than typed in afterwards.
...(startupPrompt ? { startupPrompt } : {}),
// Offered to that launcher's startup plan; it rides only when the typed line can carry it,
// and the runtime reports which so an uncarried prompt is pasted once the agent is ready.
...(startupPrompt
? {
startupPrompt,
onStartupPromptCarry: (carried: boolean) => {
promptRodeLaunchCommand = carried
}
}
: {}),
...(agentArgs !== undefined ? { agentArgs } : {}),
...(cwd ? { cwd } : {}),
// The model the user picked outranks configured args here too, as it does on a chat.
@@ -143,13 +151,17 @@ export function agentLaunchSurfaceFactory(
// The runtime already minted this pane and baked it into the PTY's env and its own reveal;
// dropping it here was what left a client with no way to name the tab it just asked for.
...(terminal.paneKey ? { paneKey: terminal.paneKey } : {}),
...(terminal.warning ? { warning: terminal.warning } : {})
// Its only warning is that the host could not reveal the tab, which the caller shows itself.
...(terminal.warning && !callerPresentsSurface ? { warning: terminal.warning } : {}),
...(promptRodeLaunchCommand ? { promptRodeLaunchCommand } : {})
}
},
deliverTerminalPrompt: async ({ handle, prompt }) =>
deliverTerminalPrompt: async ({ handle, agent, freshLaunch, prompt }) =>
deliverTerminalAgentLaunchPrompt({
runtime: context.runtime,
handle,
agent,
freshLaunch,
text: prompt.text
})
}
@@ -18,18 +18,46 @@ type SendFn = (
options: Record<string, unknown>
) => Promise<SendResult>
function runtimeStub(overrides: { wait?: unknown; send?: SendFn }) {
const waitForTerminal = vi.fn(async () => overrides.wait ?? { satisfied: true, status: 'idle' })
const COMPOSER_READY = { satisfied: true, status: 'running' }
function runtimeStub(overrides: {
wait?: unknown
waits?: unknown[]
send?: SendFn
/** Whether the agent's composer signal fires within its budget; else the idle evidence decides. */
composerSignal?: boolean
/** What a fresh read finds in the terminal's foreground; default: the agent. */
foreground?: 'agent' | 'shell' | 'unknown'
}) {
const queued = [...(overrides.waits ?? [])]
const waitForTerminal = vi.fn(
async (_handle: string, _options?: { condition?: string; timeoutMs?: number }) =>
queued.shift() ?? overrides.wait ?? { satisfied: true, status: 'idle' }
)
const waitForFreshWorkerComposer = vi.fn(async () => {
if (!overrides.composerSignal) {
throw new Error('timeout')
}
return COMPOSER_READY
})
const readLaunchedAgentForeground = vi.fn(async () => overrides.foreground ?? 'agent')
const subscribeToTerminalData = vi.fn(() => () => {})
const sendTerminalAgentPrompt = vi.fn<SendFn>(
overrides.send ?? (async () => ({ handle: 'term_1', accepted: true, bytesWritten: 12 }))
)
return {
waitForTerminal,
waitForFreshWorkerComposer,
sendTerminalAgentPrompt,
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the deliverer reaches exactly these two runtime methods; anything else would throw rather than read a wrong value.
runtime: { waitForTerminal, sendTerminalAgentPrompt } as unknown as Parameters<
typeof deliverTerminalAgentLaunchPrompt
>[0]['runtime']
readLaunchedAgentForeground,
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: the deliverer reaches exactly these five runtime methods; anything else would throw rather than read a wrong value.
runtime: {
waitForTerminal,
waitForFreshWorkerComposer,
sendTerminalAgentPrompt,
readLaunchedAgentForeground,
subscribeToTerminalData
} as unknown as Parameters<typeof deliverTerminalAgentLaunchPrompt>[0]['runtime']
}
}
@@ -48,13 +76,22 @@ describe('writing a launch prompt into a terminal agent', () => {
const delivered = await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'claude',
freshLaunch: true,
text: 'do the thing'
})
expect(delivered).toBe(true)
// The composer signal did not fire within its budget, so the idle evidence decided.
expect(stub.waitForFreshWorkerComposer).toHaveBeenCalledWith('term_1', 'claude', 8_000, {
requireComposerMarker: false,
stopOnDialog: true
})
expect(stub.waitForTerminal).toHaveBeenCalledWith('term_1', {
condition: 'tui-idle',
timeoutMs: 60_000
timeoutMs: expect.any(Number),
// A name-only title proves nothing about a just-launched agent until its stream is quiet.
launchReadiness: true
})
const [handle, text, options] = stub.sendTerminalAgentPrompt.mock.calls[0]!
expect(handle).toBe('term_1')
@@ -63,6 +100,8 @@ describe('writing a launch prompt into a terminal agent', () => {
// agent would be reported as undelivered while its prompt sat in the pane.
expect(options.acceptQueued).toBe(true)
expect(options.requestId).toEqual(expect.any(String))
// Enter follows the paste on the desktop draft paste's timing: this composer was just seen ready.
expect(options.composerReady).toBe(true)
})
it('does not write when the composer never opened', async () => {
@@ -70,6 +109,8 @@ describe('writing a launch prompt into a terminal agent', () => {
const delivered = await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'claude',
freshLaunch: true,
text: 'do the thing'
})
@@ -87,6 +128,8 @@ describe('writing a launch prompt into a terminal agent', () => {
const delivered = await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'claude',
freshLaunch: true,
text: 'do the thing'
})
@@ -103,6 +146,8 @@ describe('writing a launch prompt into a terminal agent', () => {
const delivered = await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'claude',
freshLaunch: true,
text: 'do the thing'
})
@@ -115,6 +160,8 @@ describe('writing a launch prompt into a terminal agent', () => {
const delivered = await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'claude',
freshLaunch: true,
text: 'do the thing'
})
@@ -128,10 +175,240 @@ describe('writing a launch prompt into a terminal agent', () => {
await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'claude',
freshLaunch: true,
text: ' '
})
).toBe(false)
expect(stub.waitForTerminal).not.toHaveBeenCalled()
expect(stub.sendTerminalAgentPrompt).not.toHaveBeenCalled()
})
it('waits out a blocking prompt the user dismisses, then writes', async () => {
const blocked = { satisfied: false, status: 'running', blockedReason: 'trust-prompt' }
const stub = runtimeStub({ waits: [blocked, blocked, { satisfied: true, status: 'running' }] })
const clock = fakeClock()
const delivered = await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'claude',
freshLaunch: true,
text: 'do the thing',
clock
})
expect(delivered).toBe(true)
expect(stub.waitForTerminal).toHaveBeenCalledTimes(3)
// Each re-wait spends only what is left of the one launch budget.
const lastTimeoutMs = stub.waitForTerminal.mock.calls.at(-1)?.[1]?.timeoutMs ?? 0
expect(lastTimeoutMs).toBeLessThanOrEqual(58_000)
expect(lastTimeoutMs).toBeGreaterThan(57_000)
expect(stub.sendTerminalAgentPrompt).toHaveBeenCalledTimes(1)
})
it('writes nothing into a blocking prompt still up when the budget ends', async () => {
const stub = runtimeStub({
wait: { satisfied: false, status: 'running', blockedReason: 'trust-prompt' }
})
const delivered = await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'claude',
freshLaunch: true,
text: 'do the thing',
clock: fakeClock()
})
expect(delivered).toBe(false)
expect(stub.sendTerminalAgentPrompt).not.toHaveBeenCalled()
// One check per second of a 60 s budget, then it stops rather than spinning.
expect(stub.waitForTerminal.mock.calls.length).toBeLessThanOrEqual(60)
})
it.each([
['writes when a fresh read finds the agent in front', 'agent', true],
['refuses the write when the shell is back in front', 'shell', false],
// A ready signal can be a shell's own prompt, so a read that cannot tell proves nothing.
['refuses the write when the host cannot tell what is in front', 'unknown', false]
] as const)('%s', async (_label, foreground, delivered) => {
// An agent that exits at startup hands its shell back, which must never run the prompt.
const stub = runtimeStub({ composerSignal: true, foreground })
stub.sendTerminalAgentPrompt.mockImplementation(async (handle, _text, options) => {
const { beforeWrite } = options
if (typeof beforeWrite === 'function') {
await beforeWrite('pty-1')
}
return { handle, accepted: true, bytesWritten: 12 }
})
await expect(
deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'claude',
freshLaunch: true,
text: 'do the thing'
})
).resolves.toBe(delivered)
expect(stub.readLaunchedAgentForeground).toHaveBeenCalledWith('pty-1', 'claude')
})
it('checks the foreground once for the paste, its Enter and the second Enter', async () => {
const stub = runtimeStub({ composerSignal: true })
stub.sendTerminalAgentPrompt.mockImplementation(async (handle, _text, options) => {
const { beforeWrite } = options
if (typeof beforeWrite === 'function') {
await beforeWrite('pty-1')
await beforeWrite('pty-1')
await beforeWrite('pty-1')
}
return { handle, accepted: true, bytesWritten: 12 }
})
await expect(
deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'codex',
freshLaunch: true,
text: 'do the thing'
})
).resolves.toBe(true)
expect(stub.readLaunchedAgentForeground).toHaveBeenCalledTimes(1)
})
it('writes as soon as the composer signal fires, without consulting the idle evidence', async () => {
const stub = runtimeStub({ composerSignal: true })
const delivered = await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'claude',
freshLaunch: true,
text: 'do the thing'
})
expect(delivered).toBe(true)
expect(stub.waitForFreshWorkerComposer).toHaveBeenCalledWith('term_1', 'claude', 8_000, {
requireComposerMarker: false,
stopOnDialog: true
})
expect(stub.waitForTerminal).not.toHaveBeenCalled()
})
it.each(['zcode', 'opencode', 'opencode2'] as const)(
'waits for %s’s composer readiness',
async (agent) => {
const stub = runtimeStub({ composerSignal: true })
const delivered = await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent,
freshLaunch: true,
text: 'do the thing'
})
expect(delivered).toBe(true)
expect(stub.waitForFreshWorkerComposer).toHaveBeenCalledWith('term_1', agent, 60_000)
expect(stub.waitForTerminal).not.toHaveBeenCalled()
}
)
it('keeps the text when an agent shows no readiness evidence at all', async () => {
// Nothing is pasted blind.
const stub = runtimeStub({ wait: { satisfied: false, status: 'running' } })
const delivered = await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'goose',
freshLaunch: true,
text: 'do the thing'
})
expect(delivered).toBe(false)
expect(stub.sendTerminalAgentPrompt).not.toHaveBeenCalled()
})
it('never lets a re-wait fall back to the terminal wait’s 5-minute default', async () => {
// A late 1 s re-check sleep can land past the deadline; 0 would read as "use the default".
const stub = runtimeStub({
wait: { satisfied: false, status: 'running', blockedReason: 'trust-prompt' }
})
let now = 0
const lateClock = {
now: () => now,
sleep: async (ms: number) => {
now += ms + 1_000
}
}
await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_1',
agent: 'claude',
freshLaunch: true,
text: 'do the thing',
clock: lateClock
})
for (const [, options] of stub.waitForTerminal.mock.calls) {
expect(options?.timeoutMs).toBeGreaterThan(0)
}
})
it('waits on a reused terminal’s idle state, not a fresh launch’s composer marker', async () => {
// A long-running grok pane may no longer show its composer marker in recent output.
const stub = runtimeStub({})
const delivered = await deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_existing',
agent: 'grok',
freshLaunch: false,
text: 'do the thing'
})
expect(delivered).toBe(true)
expect(stub.waitForFreshWorkerComposer).not.toHaveBeenCalled()
expect(stub.waitForTerminal).toHaveBeenCalledWith('term_existing', {
condition: 'tui-idle',
timeoutMs: 60_000
})
// A reused pane's render state is only inferred, so its Enter still waits for the render.
expect(stub.sendTerminalAgentPrompt.mock.calls[0]?.[2]?.composerReady).toBe(false)
})
it('writes nothing into a reused terminal whose agent is no longer in front', async () => {
const stub = runtimeStub({ foreground: 'shell' })
stub.sendTerminalAgentPrompt.mockImplementation(async (handle, _text, options) => {
const { beforeWrite } = options
if (typeof beforeWrite === 'function') {
await beforeWrite('pty-1')
}
return { handle, accepted: true, bytesWritten: 12 }
})
await expect(
deliverTerminalAgentLaunchPrompt({
runtime: stub.runtime,
handle: 'term_existing',
agent: 'claude',
freshLaunch: false,
text: 'do the thing'
})
).resolves.toBe(false)
})
})
function fakeClock() {
let now = 0
return {
now: () => now,
sleep: async (ms: number) => {
now += ms
}
}
}
@@ -7,11 +7,11 @@
* Mobile, the CLI and orchestration got an agent and no prompt. The host owns the PTY, so it can
* write into one whether or not any window is open on it.
*
* This is the half argv cannot serve. An agent whose CLI takes the prompt as an argument gets it on
* the launch command instead (`agentPromptRidesLaunchCommand`), where it is in the process's argv
* at exec time and no readiness race exists. What reaches here is a `stdin-after-start` agent,
* whose CLI accepts no such argument, and a reused terminal, whose process was already running
* before this launch existed.
* This is the half the launch command cannot serve. An argv agent's prompt rides that command when
* its typed line can carry it (`startup-line-prompt-carry`), and then no readiness race exists. What
* reaches here is a `stdin-after-start` agent, whose CLI accepts no such argument; a prompt too long
* or multi-line for the typed line; and a reused terminal, whose process was already running before
* this launch existed. Readiness is `waitForLaunchedAgentComposer`, the one the worker start uses.
*
* Nothing here writes to a PTY itself. `sendTerminalAgentPrompt` is the runtime's one agent-prompt
* writer: it frames the text as a bracketed paste so multi-line and special-character content is
@@ -22,13 +22,64 @@
*/
import { randomUUID } from 'node:crypto'
import type { TuiAgent } from '../../../../shared/tui-agent'
import type { RuntimeTerminalWait } from '../../../../shared/runtime-terminal-contracts'
import { isAgentPromptStalledError } from '../../agent-prompt-submission-verification'
import {
waitForLaunchedAgentComposer,
type LaunchedAgentReadinessRuntime
} from '../../launched-agent-composer-readiness'
import type { OrcaRuntimeService } from '../../orca-runtime'
import {
createLaunchedAgentWriteGuard,
type LaunchedAgentWriteGuardRuntime
} from '../../launched-agent-write-guard'
/** The same budget orchestration gives a worker to reach its composer before dispatching to it. */
const AGENT_READY_TIMEOUT_MS = 60_000
/** How often a launch re-checks a blocking prompt the user may still dismiss. */
const BLOCKED_RECHECK_MS = 1_000
type TerminalPromptRuntime = Pick<OrcaRuntimeService, 'waitForTerminal' | 'sendTerminalAgentPrompt'>
type TerminalPromptRuntime = LaunchedAgentReadinessRuntime &
LaunchedAgentWriteGuardRuntime &
Pick<OrcaRuntimeService, 'sendTerminalAgentPrompt'>
type ReadinessClock = { now: () => number; sleep: (ms: number) => Promise<void> }
const REAL_CLOCK: ReadinessClock = {
now: () => Date.now(),
sleep: (ms) => new Promise((resolve) => setTimeout(resolve, ms))
}
/**
* The launched agent's readiness, waiting out a blocking prompt until the budget ends.
*
* A trust or update dialog is the user's to answer, and they often do within seconds; giving up on
* the first sight of it dropped a prompt the pane was about to accept. A dialog still up when the
* budget ends is reported as it is, and nothing is written into it.
*/
async function waitThroughBlockingPrompts(
runtime: TerminalPromptRuntime,
handle: string,
agent: TuiAgent,
freshLaunch: boolean,
clock: ReadinessClock
): Promise<RuntimeTerminalWait | undefined> {
const deadline = clock.now() + AGENT_READY_TIMEOUT_MS
for (;;) {
// At least 1 ms: the terminal wait reads 0 as "use the 5-minute default", and a late sleep can
// land past the deadline.
const timeoutMs = Math.max(1, deadline - clock.now())
const wait = freshLaunch
? await waitForLaunchedAgentComposer(runtime, handle, agent, timeoutMs)
: // A reused pane was not freshly launched: its composer marker may be long gone.
await runtime.waitForTerminal(handle, { condition: 'tui-idle', timeoutMs })
if (!wait?.blockedReason || wait.satisfied || deadline - clock.now() <= BLOCKED_RECHECK_MS) {
return wait
}
await clock.sleep(BLOCKED_RECHECK_MS)
}
}
/**
* Whether the text reached the pane.
@@ -47,18 +98,28 @@ type TerminalPromptRuntime = Pick<OrcaRuntimeService, 'waitForTerminal' | 'sendT
export async function deliverTerminalAgentLaunchPrompt(args: {
runtime: TerminalPromptRuntime
handle: string
agent: TuiAgent
/** False for a reused terminal, whose agent was already running before this launch. */
freshLaunch: boolean
text: string
clock?: ReadinessClock
}): Promise<boolean> {
if (args.text.trim().length === 0) {
return false
}
// Before the paste and again before Enter, for a reused pane too: a ready signal can come from a
// shell whose agent exited, so only a read that finds the agent in front lets the text through.
const guard = createLaunchedAgentWriteGuard(args.runtime, args.agent)
try {
const wait = await args.runtime.waitForTerminal(args.handle, {
condition: 'tui-idle',
timeoutMs: AGENT_READY_TIMEOUT_MS
})
// An unsatisfied wait is a composer that never opened — a trust prompt, an update prompt, a
// dead process. Pasting anyway would answer whatever question is on screen with the prompt.
const wait = await waitThroughBlockingPrompts(
args.runtime,
args.handle,
args.agent,
args.freshLaunch,
args.clock ?? REAL_CLOCK
)
// An unsatisfied wait is a composer that never opened — a dialog left up, a dead process, an
// agent that showed no readiness. Pasting anyway would answer whatever is on screen with it.
if (wait && !wait.satisfied) {
console.warn(
`[agent-launch] the terminal agent did not become ready (${wait.status}); its launch prompt was not delivered`
@@ -67,6 +128,9 @@ export async function deliverTerminalAgentLaunchPrompt(args: {
}
const sent = await args.runtime.sendTerminalAgentPrompt(args.handle, args.text, {
inputKind: 'launch',
// A fresh launch's composer was just seen ready; a reused pane's state is only inferred.
composerReady: args.freshLaunch,
beforeWrite: guard.beforeWrite,
// Paired: together these take the queued path, which settles an unobserved turn start into
// an `input_accepted` receipt rather than raising it. Without the id the write is verified
// strictly and a slow first turn throws.
@@ -86,5 +150,7 @@ export async function deliverTerminalAgentLaunchPrompt(args: {
error
)
return false
} finally {
guard.dispose()
}
}
@@ -47,6 +47,7 @@ export function agentLaunchWorkspaceFactory(
options
}) => {
const startupLaunchPreferences = toAgentLaunchPreferences(options)
let promptRodeLaunchCommand = false
// oxlint-disable-next-line typescript/consistent-type-assertions -- SAFETY: already validated by `AgentLaunch`; the executor only removed the reserved agent fields, so the rest of the payload is the parsed shape.
const params = create as WorktreeCreateParams
const { runtime } = context
@@ -65,8 +66,8 @@ export function agentLaunchWorkspaceFactory(
...params,
...(startupAgent ? { startupAgent } : {}),
// Only ever set alongside `startupAgent`, which is what the create requires; the
// executor sends it exclusively for an agent that takes its prompt on argv, so this
// is the startup command carrying the text rather than a second delivery path.
// executor offers it only to an agent that takes its prompt on argv, and it rides only
// when the typed line can carry it.
...(startupPrompt ? { startupPrompt } : {})
},
{
@@ -80,6 +81,13 @@ export function agentLaunchWorkspaceFactory(
context.clientKind ? { clientKind: context.clientKind } : {}
),
...(agentArgs !== undefined ? { startupAgentArgs: agentArgs } : {}),
...(startupPrompt
? {
onStartupPromptCarry: (carried: boolean) => {
promptRodeLaunchCommand = carried
}
}
: {}),
...(cwd ? { startupCwd: cwd } : {}),
...(launchSource ? { startupLaunchSource: launchSource } : {}),
...(paneKey ? { startupPaneKey: paneKey } : {}),
@@ -100,6 +108,7 @@ export function agentLaunchWorkspaceFactory(
return {
worktreeId: result.worktree.id,
startupTerminalHandle: result.startupTerminal?.handle,
...(promptRodeLaunchCommand ? { promptRodeLaunchCommand } : {}),
...(result.startupTerminal?.paneKey
? { startupTerminalPaneKey: result.startupTerminal.paneKey }
: {}),
@@ -6,7 +6,7 @@
*/
import { vi } from 'vitest'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../../../shared/protocol-version'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../../../shared/agent-launch-runtime-capability'
import { AgentLaunchPaneAlreadyLiveError } from '../../../../shared/agent-launch-pane-already-live'
import type { RpcContext } from '../core'
@@ -35,6 +35,18 @@ export type AgentLaunchRuntimeStubOptions = {
startupTerminalPaneKey?: string
/** The reserved pane is already live, so a create that requires a fresh pane is refused. */
terminalPaneAlreadyLive?: boolean
/** What the runtime reports about an offered prompt's typed line; unset reports nothing. */
lineCarriesPrompt?: boolean
}
function reportPromptCarry(
options: AgentLaunchRuntimeStubOptions,
report: unknown,
offered: unknown
): void {
if (typeof report === 'function' && offered && options.lineCarriesPrompt !== undefined) {
report(options.lineCarriesPrompt)
}
}
export function runtimeStub(options: AgentLaunchRuntimeStubOptions = {}) {
@@ -66,22 +78,26 @@ export function runtimeStub(options: AgentLaunchRuntimeStubOptions = {}) {
}
),
showRepo: vi.fn(async () => ({ id: 'repo-1' })),
createManagedWorktree: vi.fn(async (args: Record<string, unknown>) => ({
worktree: { id: 'wt-new' },
startupTerminal: args.startupAgent
? {
handle: 'term_agent_first',
...(options.startupTerminalPaneKey ? { paneKey: options.startupTerminalPaneKey } : {})
}
: undefined,
...(options.setupReceipt ? { setupReceipt: options.setupReceipt } : {}),
...(options.createWarning ? { warning: options.createWarning } : {})
})),
createManagedWorktree: vi.fn(async (args: Record<string, unknown>) => {
reportPromptCarry(options, args.onStartupPromptCarry, args.startupPrompt)
return {
worktree: { id: 'wt-new' },
startupTerminal: args.startupAgent
? {
handle: 'term_agent_first',
...(options.startupTerminalPaneKey ? { paneKey: options.startupTerminalPaneKey } : {})
}
: undefined,
...(options.setupReceipt ? { setupReceipt: options.setupReceipt } : {}),
...(options.createWarning ? { warning: options.createWarning } : {})
}
}),
// Args are declared so a test can assert what the launch asked for, not merely that it asked.
createTerminal: vi.fn(async (_selector: string, createOptions?: Record<string, unknown>) => {
if (options.terminalPaneAlreadyLive && createOptions?.requireFreshPane === true) {
throw new AgentLaunchPaneAlreadyLiveError()
}
reportPromptCarry(options, createOptions?.onStartupPromptCarry, createOptions?.startupPrompt)
return {
handle: 'term_1',
...(options.terminalPaneKey ? { paneKey: options.terminalPaneKey } : {}),
@@ -9,7 +9,7 @@
*/
import { beforeEach, describe, expect, it, vi } from 'vitest'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../../../shared/protocol-version'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../../../shared/agent-launch-runtime-capability'
import type { RpcContext } from '../core'
import {
CAPABLE_CLIENT,
+11 -8
View File
@@ -19,7 +19,7 @@
* harmless, and does nothing to reunite a caller with a surface a dead attempt left behind.
*/
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../../../shared/protocol-version'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../../../shared/agent-launch-runtime-capability'
import { computeAgentLaunchFingerprint } from '../../../../shared/agent-launch-operation'
import type {
AgentLaunchIntent,
@@ -154,22 +154,25 @@ async function runAgentLaunch(
terminalSpawn?: TerminalSpawnDispatch
): Promise<AgentLaunchResult> {
const callerNavigationId = agentLaunchCallerNavigationId(intent.target, context)
const result = await executeAgentLaunch({
return executeAgentLaunch({
runtime: context.runtime,
intent,
surfaces: agentLaunchSurfaceFactory(
context,
attachOperationId,
operationCallerKey,
callerNavigationId === null,
callerNavigationId !== null,
terminalSpawn
),
workspaces: agentLaunchWorkspaceFactory(context, intent.agent)
workspaces: agentLaunchWorkspaceFactory(context, intent.agent),
// The tab is shown as it is published, not after a prompt that can take a minute to land.
...(callerNavigationId !== null
? {
onSurfacePublished: (surface) =>
selectAgentLaunchTabForCaller(context.runtime, surface, callerNavigationId)
}
: {})
})
if (callerNavigationId !== null) {
selectAgentLaunchTabForCaller(context.runtime, result, callerNavigationId)
}
return result
}
/**
@@ -0,0 +1,105 @@
import { afterEach, describe, expect, it, vi } from 'vitest'
import { OrcaRuntimeService } from '../../../../orca-runtime'
import { OrchestrationDb } from '../../../../orchestration/db'
import type { LaunchedAgentForeground } from '../../../../launched-agent-foreground'
import { ORCHESTRATION_METHODS } from '../../orchestration'
import { configureFederationWorkerRuntime } from './federation-runtime.test-support'
const PTY_ID = 'pty_worker'
// Why: a shell back at its prompt after the agent exits reads as ready too, so the worker host
// must find the launched agent in front before it types the brief, or the shell runs the brief.
describe('a paired-server worker start writes its brief only into the agent it launched', () => {
const databases: OrchestrationDb[] = []
afterEach(() => {
for (const db of databases.splice(0)) {
db.close()
}
vi.restoreAllMocks()
})
function workerHost(foreground: LaunchedAgentForeground, provesAgent = true) {
const db = new OrchestrationDb(':memory:')
databases.push(db)
const runtime = new OrcaRuntimeService()
runtime.setOrchestrationDb(db)
configureFederationWorkerRuntime(runtime)
const writes: string[] = []
vi.spyOn(runtime, 'readLaunchedAgentForeground').mockResolvedValue(foreground)
vi.spyOn(runtime, 'launchedAgentHostProvesAgent').mockReturnValue(provesAgent)
vi.spyOn(runtime, 'subscribeToTerminalData').mockReturnValue(() => {})
vi.mocked(runtime.sendTerminalAgentPrompt).mockImplementation(async (handle, text, options) => {
await options?.beforeWrite?.(PTY_ID)
writes.push(text)
return { handle, accepted: true, bytesWritten: text.length }
})
return { runtime, writes }
}
async function attach(runtime: OrcaRuntimeService): Promise<unknown> {
const method = ORCHESTRATION_METHODS.find(
(candidate) => candidate.name === 'orchestration.federationAttachStart'
)
if (!method) {
throw new Error('federationAttachStart method is not registered')
}
return await method.handler(
method.params!.parse({
runId: 'run-home',
dispatchId: 'ctx_remote',
taskId: 'task_remote',
taskSpec: 'remote worker',
protocolVersion: 3,
worktree: 'new-top-level',
repo: 'windows-repo',
name: 'remote-worker',
agent: 'claude'
}),
{
runtime,
orchestrationMutation: {
callerFingerprint: 'home_peer',
requestId: 'request_remote',
method: 'orchestration.federationAttachStart',
payloadHash: 'remote_payload'
}
}
)
}
it('types nothing when the agent exited and its shell is in front', async () => {
const host = workerHost('shell')
await expect(attach(host.runtime)).resolves.toMatchObject({
state: 'failed',
failedStage: 'dispatch_input',
lastError: 'agent_not_in_foreground'
})
expect(host.writes).toEqual([])
})
it('types nothing where the host cannot tell what is in front', async () => {
const host = workerHost('unknown')
await expect(attach(host.runtime)).resolves.toMatchObject({ state: 'failed' })
expect(host.writes).toEqual([])
})
it('writes the brief once into the agent found in front', async () => {
const host = workerHost('agent')
await expect(attach(host.runtime)).resolves.toMatchObject({ state: 'ready' })
expect(host.writes).toHaveLength(1)
})
it('on a host that cannot prove the agent (Windows), writes unless a shell is proven', async () => {
const unknown = workerHost('unknown', false)
await expect(attach(unknown.runtime)).resolves.toMatchObject({ state: 'ready' })
expect(unknown.writes).toHaveLength(1)
const shell = workerHost('shell', false)
await expect(attach(shell.runtime)).resolves.toMatchObject({ state: 'failed' })
expect(shell.writes).toEqual([])
})
})
@@ -3,6 +3,7 @@ import type { TuiAgent } from '../../../../../../shared/tui-agent'
import { describeTerminalWaitBlockedReason } from '../../../../../../shared/terminal-wait-blocked-reason-legacy-alias'
import { buildDispatchPreamble } from '../../../../orchestration/preamble'
import { sendAgentTurn } from '../../../../orchestration/send-agent-turn'
import { createWorkerBriefWriteGuard } from '../../../../launched-agent-write-guard'
import { OrchestrationError } from '../../../../orchestration/orchestration-error'
import { defineMethod } from '../../../core'
import { assertOrchestrationWorktreeCreationSupported } from '../worker/folder-worktree-placement'
@@ -244,10 +245,13 @@ export const ORCHESTRATION_FEDERATION_ATTACH_METHODS = [
reusesTerminal: Boolean(params.terminal)
})
failedStage = 'dispatch_input'
// A shell back at its prompt also reads as ready, so the brief needs the agent found in front.
const briefGuard = createWorkerBriefWriteGuard(runtime, agent, !params.terminal)
const prompt = await sendAgentTurn({
kind: 'terminal',
runtime,
handle: terminalHandle,
...(briefGuard ? { beforeWrite: briefGuard.beforeWrite } : {}),
turn: {
purpose: 'dispatch-preamble',
operationId: orchestrationMutation.requestId,
@@ -264,7 +268,7 @@ export const ORCHESTRATION_FEDERATION_ATTACH_METHODS = [
cliCommand: runtime.getTerminalOrchestrationCliCommand(terminalHandle)
})
}
})
}).finally(() => briefGuard?.dispose())
effects.push({
kind: 'dispatch_input',
role: 'agent',
@@ -9,6 +9,7 @@ import { canonicalOrcaSessionId } from '../../../../orchestration/canonical-orca
import { orcaSessionIdOrHandle } from '../../../../orchestration/orchestration-party'
import { buildDispatchPreamble } from '../../../../orchestration/preamble'
import { sendAgentTurn } from '../../../../orchestration/send-agent-turn'
import { createWorkerBriefWriteGuard } from '../../../../launched-agent-write-guard'
import { sendStructuredWorkerPreamble } from '../../orchestration-structured-worker-session'
import type { WorkerTurnStartObservation } from './worker-start-turn-observation'
import type { createStructuredWorkerSessionForWorktree } from './worker-topology'
@@ -36,6 +37,8 @@ export async function deliverWorkerDispatchPreamble(args: {
coordinatorHandle: string
devMode: boolean | undefined
requestId: string
/** The agent this worker start launched into `terminalHandle`; absent for a caller's terminal. */
launchedAgent?: string | null
}): Promise<{
prompt?: RuntimeTerminalSend['prompt']
structuredTurnStart?: WorkerTurnStartObservation
@@ -78,14 +81,18 @@ export async function deliverWorkerDispatchPreamble(args: {
}
}
}
return {
prompt: (
await sendAgentTurn({
kind: 'terminal',
runtime,
handle: terminalHandle,
turn: { purpose: 'dispatch-preamble', body: preamble, operationId: args.requestId }
})
).prompt
// A shell back at its prompt also reads as ready, so the brief needs the agent found in front.
const briefGuard = createWorkerBriefWriteGuard(runtime, args.launchedAgent, !!args.launchedAgent)
try {
const sent = await sendAgentTurn({
kind: 'terminal',
runtime,
handle: terminalHandle,
...(briefGuard ? { beforeWrite: briefGuard.beforeWrite } : {}),
turn: { purpose: 'dispatch-preamble', body: preamble, operationId: args.requestId }
})
return { prompt: sent.prompt }
} finally {
briefGuard?.dispose()
}
}
@@ -269,6 +269,7 @@ export async function startLocalWorker(args: {
devMode: params.devMode,
requestId: orchestrationMutation?.requestId ?? started.dispatch.id,
agent: agent ?? null,
launchedAgent: params.terminal ? null : (agent ?? null),
setupReceipt,
launchReceipt: launch.receipt,
mode,
@@ -0,0 +1,50 @@
import { afterEach, describe, expect, it, vi } from 'vitest'
import type { LaunchedAgentForeground } from '../../../../launched-agent-foreground'
import { createOrchestrationWorkerReleaseHarness } from './worker-release.test-support'
const PTY_ID = 'pty_worker'
// Why: a shell back at its prompt after the agent exits reads as ready too, so the brief needs the
// launched agent found in front, or the shell runs it.
describe('a worker start writes its brief only into the agent it launched', () => {
const h = createOrchestrationWorkerReleaseHarness()
afterEach(() => h.cleanup())
function launchedPane(foreground: LaunchedAgentForeground): string[] {
h.setup()
const writes: string[] = []
vi.spyOn(h.runtime, 'readLaunchedAgentForeground').mockResolvedValue(foreground)
vi.spyOn(h.runtime, 'launchedAgentHostProvesAgent').mockReturnValue(true)
vi.spyOn(h.runtime, 'subscribeToTerminalData').mockReturnValue(() => {})
vi.mocked(h.runtime.sendTerminalAgentPrompt).mockImplementation(
async (handle, text, options) => {
await options?.beforeWrite?.(PTY_ID)
writes.push(text)
return { handle, accepted: true, bytesWritten: text.length }
}
)
return writes
}
it('types nothing when the agent exited and its shell is in front', async () => {
const writes = launchedPane('shell')
await expect(h.startWorker({ agent: 'claude' })).rejects.toThrow()
expect(writes).toEqual([])
})
it('writes the brief once into the agent found in front', async () => {
const writes = launchedPane('agent')
await h.startWorker({ agent: 'claude' })
expect(writes).toHaveLength(1)
})
it('leaves a terminal the caller supplied to its own idle wait', async () => {
const writes = launchedPane('shell')
await h.startWorker({ terminal: 'term_worker' })
expect(h.runtime.readLaunchedAgentForeground).not.toHaveBeenCalled()
expect(writes).toHaveLength(1)
})
})
@@ -34,6 +34,8 @@ export async function deliverAndSettleWorkerStartReadiness(args: {
devMode: boolean | undefined
requestId: string
agent: string | null
/** The agent this start launched into `terminalHandle`; null when the caller supplied it. */
launchedAgent: string | null
setupReceipt: WorkerSetupReceipt
launchReceipt: OrchestrationWorkerLaunchReceipt
mode: WorkerStartModeReceipt
@@ -57,7 +59,8 @@ export async function deliverAndSettleWorkerStartReadiness(args: {
taskSpec: task.spec,
coordinatorHandle: args.coordinatorHandle,
devMode: args.devMode,
requestId: args.requestId
requestId: args.requestId,
launchedAgent: args.launchedAgent
})
effects.push({
kind: 'dispatch_input',
@@ -454,6 +454,25 @@ describe('orchestration new-worktree workers', () => {
)
})
// Why: a first dispatch the composer signal settles must record setup exactly as one the idle
// wait settles; the composer lane once returned nothing and skipped this record.
it('settles a fresh worker start on main’s idle wait, not the launch paste’s signal', async () => {
mockCreatedWorktree({ startupPolicy: 'wait-for-setup', state: 'running' })
const composerSignal = vi.spyOn(runtime, 'waitForFreshWorkerComposer')
const { result } = await startWorker()
expect(composerSignal).not.toHaveBeenCalled()
expect(runtime.waitForTerminal).toHaveBeenCalledWith(
'term_worker',
expect.objectContaining({ condition: 'tui-idle', launchReadiness: true })
)
expect(result).toMatchObject({
state: 'ready',
setup: { startupPolicy: 'wait-for-setup', state: 'succeeded' }
})
})
it('does not inject task input when the gated setup terminal fails to start', async () => {
mockCreatedWorktree({ startupPolicy: 'wait-for-setup', state: 'spawn_failed' })
vi.mocked(runtime.waitForTerminal).mockResolvedValue({
@@ -2,7 +2,7 @@ import { afterEach, describe, expect, it, vi } from 'vitest'
import type { RuntimeTerminalWait } from '../../../../../../shared/runtime-terminal-contracts'
import { createOrchestrationWorkerReleaseHarness } from './worker-release.test-support'
describe('ZCode first dispatch readiness', () => {
describe('composer-marker first dispatch readiness', () => {
const h = createOrchestrationWorkerReleaseHarness()
afterEach(() => h.cleanup())
@@ -3,11 +3,11 @@
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import {
AGENT_LAUNCH_RUNTIME_CAPABILITY,
CLAUDE_STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY,
STRUCTURED_AGENT_SESSION_CLIENT_LAUNCH_MODE_CAPABILITY,
STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY
} from '../../../../shared/protocol-version'
import { AGENT_LAUNCH_RUNTIME_CAPABILITY } from '../../../../shared/agent-launch-runtime-capability'
import {
CLEANUP_METHODS,
WORK_METHODS
@@ -4,7 +4,12 @@ import type { TerminalWorkspaceLaunchScope } from './runtime-legacy-worker-termi
import type { TerminalCreateOptions } from './runtime-terminal-contracts'
import { isTuiAgentEnabled } from '../../shared/tui-agent-selection'
import { resolveBareAgentLaunchCommand } from './runtime-agent-launch-resolution'
import { buildExecutionHostAgentStartupPlan } from '../opencode/opencode-model-startup-plan'
import { planExecutionHostStartupWithPromptCandidate } from '../opencode/opencode-model-startup-plan'
import { agentPromptRidesLaunchCommand } from '../../shared/tui-agent-startup'
import {
launchHostProvesAgentInFront,
nameLocalTypedLineShell
} from './agent-launch-typed-line-shell'
import { resolveTerminalStartupCwd } from '../../shared/terminal-startup-cwd'
import { resolveAgentStartupPlanInputs } from '../../shared/agent-startup-plan-inputs'
import { agentStartedTelemetry } from '../agent-launch/agent-started-telemetry'
@@ -35,7 +40,12 @@ export async function buildRuntimeAgentTerminalStartupOptions(
return opts
}
const startupPlan = await buildExecutionHostAgentStartupPlan({
// A prompt this launch command cannot carry has nowhere to go from here — the create returns
// options, not a live PTY — so refuse rather than spawn the agent and drop the text.
if (opts.startupPrompt && !agentPromptRidesLaunchCommand(agent)) {
throw new Error(`Agent ${agent} does not take a startup prompt on its launch command.`)
}
const { plan: startupPlan, promptCarried } = await planExecutionHostStartupWithPromptCandidate({
inputs: resolveAgentStartupPlanInputs({
agent,
settings,
@@ -48,7 +58,17 @@ export async function buildRuntimeAgentTerminalStartupOptions(
}),
prompt: opts.startupPrompt ?? '',
cwd: resolveTerminalStartupCwd(workspace.path, opts.cwd) ?? workspace.path,
hostIdentity
hostIdentity,
host: {
shellName: nameLocalTypedLineShell({
isRemote,
...(opts.shellOverride ? { shellOverride: opts.shellOverride } : {}),
...(settings.terminalDefaultShell
? { defaultShellSetting: settings.terminalDefaultShell }
: {})
}),
provesAgentInFront: launchHostProvesAgentInFront({ isRemote, launchPlatform: platform })
}
})
if (!startupPlan) {
// Why: an explicit agent that yields no plan would otherwise spawn a bare
@@ -58,10 +78,8 @@ export async function buildRuntimeAgentTerminalStartupOptions(
}
return opts
}
// A prompt this launch command cannot carry has nowhere to go from here — the create returns
// options, not a live PTY — so refuse rather than spawn the agent and drop the text.
if (opts.startupPrompt && 'followupPrompt' in startupPlan && startupPlan.followupPrompt) {
throw new Error(`Agent ${agent} does not take a startup prompt on its launch command.`)
if (opts.startupPrompt) {
opts.onStartupPromptCarry?.(promptCarried)
}
return {
@@ -51,6 +51,9 @@ export type RuntimeManagedWorktreeCreateArgs = {
startupAgent?: TuiAgent
startupLaunchPreferences?: AgentLaunchPreferences
startupPrompt?: string
/** Main-internal: set by a caller that delivers an uncarried `startupPrompt` itself, so the text
* rides only a typed line that can carry it; reports whether it did. */
onStartupPromptCarry?: (carried: boolean) => void
/** Per-launch inputs used when `startupAgent` is the created terminal surface. */
startupAgentArgs?: string | null
startupCwd?: string
@@ -44,6 +44,25 @@ describe('OrcaRuntimeRpcServer', () => {
).toBe('wait')
})
it.each(['agent.launch', 'agent.launchReplay'])(
'keeps a prompted %s alive while it waits for the agent to be ready',
(method) => {
const prompt = { text: 'fix the build', delivery: 'submit' }
expect(
classifyRuntimeLongPoll({
id: 'req_launch',
authToken: 'token',
method,
params: { prompt }
})
).toBe('wait')
// Without a prompt nothing waits on the agent, so the launch stays a short call.
expect(
classifyRuntimeLongPoll({ id: 'req_launch', authToken: 'token', method, params: {} })
).toBeNull()
}
)
it('keeps agent-prompt submission sockets alive during verification', () => {
expect(
classifyRuntimeLongPoll({
@@ -26,6 +26,17 @@ export function classifyRuntimeLongPoll(request: RpcRequest): RuntimeLongPollCla
if (request.method === 'orchestration.workerStart') {
return 'wait'
}
// A launch with a prompt waits for the agent's readiness before writing it, up to 60 s, and a
// reply lost to the 30 s idle timer reads as a dead runtime instead of an undelivered prompt.
if (
(request.method === 'agent.launch' || request.method === 'agent.launchReplay') &&
typeof request.params === 'object' &&
request.params !== null &&
'prompt' in request.params &&
request.params.prompt !== undefined
) {
return 'wait'
}
if (request.method === 'browser.clientHost.attach') {
return 'browser-host'
}
@@ -86,6 +86,8 @@ export type RuntimeStore = {
agentDefaultArgs?: GlobalSettings['agentDefaultArgs']
agentDefaultEnv?: GlobalSettings['agentDefaultEnv']
terminalWindowsShell?: GlobalSettings['terminalWindowsShell']
// Read by the launch-line carry rule to name the shell a local line is typed into.
terminalDefaultShell?: GlobalSettings['terminalDefaultShell']
floatingTerminalEnabled?: GlobalSettings['floatingTerminalEnabled']
agentStatusHooksEnabled?: GlobalSettings['agentStatusHooksEnabled']
terminalCopyTrimsGutter?: GlobalSettings['terminalCopyTrimsGutter']
+12 -2
View File
@@ -45,12 +45,16 @@ export type TerminalCreateOptions = {
launchAgent?: TuiAgent
startupAgent?: TuiAgent
/**
* Initial text folded into `startupAgent`'s launch command, for an agent whose CLI takes a prompt
* Initial text offered to `startupAgent`'s launch command, for an agent whose CLI takes a prompt
* argument. Not a general prompt channel: an agent that takes its text only after start has no
* launch command to carry it, and a caller that sets this for one is refused rather than having
* the prompt silently dropped. Post-start delivery belongs to whoever owns the live PTY.
* the prompt silently dropped. The text rides only when the typed line can carry it
* (`startup-line-prompt-carry`); otherwise the agent starts clean, `onStartupPromptCarry` says so,
* and post-start delivery belongs to whoever owns the live PTY.
*/
startupPrompt?: string
/** Main-internal: whether `startupPrompt` rode the launch command. Called once the plan is built. */
onStartupPromptCarry?: (carried: boolean) => void
/**
* Replaces the Settings launch arguments for this `startupAgent` only; `null` means none at all.
*
@@ -206,6 +210,9 @@ export type TerminalWaiter = {
/** Retires this waiter from the shared idle-poll sweep; null when not polling. */
cancelIdlePoll: (() => void) | null
abortCleanup: (() => void) | null
/** Main-internal: waiting for a just-launched agent, so a name-only title must be held to a
* quiet stream (`TuiIdleEvaluationInput.launchReadiness`). */
launchReadiness?: boolean
}
/** How a provider-held screen should be fetched when runtime bytes are absent. */
@@ -221,6 +228,9 @@ export type RuntimeAgentPromptWriteOptions = Omit<RuntimeTerminalWriteOptions, '
inputKind: Exclude<TerminalInputKind, 'query-reply'>
/** Raw prompt text for submit scheduling; not written, only used for line-aware delays. */
promptForSchedule?: string
/** The caller just saw this agent's composer accept input, so Enter follows the paste on the
* desktop draft paste's timing instead of waiting for the render to settle. */
composerReady?: boolean
/** See buildAgentPromptPasteBytes. */
leadLine?: string
/** Return an accepted receipt as soon as input lands, instead of waiting for the turn. */
@@ -131,7 +131,10 @@ export class RuntimeTerminalIdlePolls {
const readWaitText = () =>
buildTerminalWaitText(pty.tailBuffer, pty.tailPartialLine, pty.preview)
return {
verdict: evaluateTuiIdle(ptyTuiIdleEvidence(this.deps, pty, readWaitText)),
verdict: evaluateTuiIdle({
...ptyTuiIdleEvidence(this.deps, pty, readWaitText),
launchReadiness: entry.waiter.launchReadiness
}),
ptyId: pty.ptyId,
ready: () => buildPtyTerminalWaitResult(handle, 'tui-idle', pty),
blocked: (reason) => buildPtyTerminalWaitBlockedResult(handle, 'tui-idle', pty, reason),
@@ -150,7 +153,10 @@ export class RuntimeTerminalIdlePolls {
buildTerminalWaitText(leaf.tailBuffer, leaf.tailPartialLine, leaf.preview)
const live = () => this.deps.getLiveLeaf(entry.leaf)
return {
verdict: evaluateTuiIdle(leafTuiIdleEvidence(this.deps, leaf, readWaitText)),
verdict: evaluateTuiIdle({
...leafTuiIdleEvidence(this.deps, leaf, readWaitText),
launchReadiness: entry.waiter.launchReadiness
}),
ptyId: leaf.ptyId,
ready: () => buildTerminalWaitResult(handle, 'tui-idle', live()),
blocked: (reason) => buildTerminalWaitBlockedResult(handle, 'tui-idle', live(), reason),
+6 -2
View File
@@ -76,6 +76,8 @@ export class RuntimeTerminalWait {
condition?: RuntimeTerminalWaitCondition
timeoutMs?: number
signal?: AbortSignal
/** Main-internal, never on the wire: see `TerminalWaiter.launchReadiness`. */
launchReadiness?: boolean
}
): Promise<RuntimeTerminalWaitResult> {
const condition = options?.condition ?? 'exit'
@@ -112,7 +114,8 @@ export class RuntimeTerminalWait {
reject,
timeout: null,
cancelIdlePoll: null,
abortCleanup: null
abortCleanup: null,
...(options?.launchReadiness ? { launchReadiness: true } : {})
}
if (!this.waiters.bindAbort(waiter, options?.signal)) {
reject(new Error('request_aborted'))
@@ -191,7 +194,8 @@ export class RuntimeTerminalWait {
reject,
timeout: null,
cancelIdlePoll: null,
abortCleanup: null
abortCleanup: null,
...(options?.launchReadiness ? { launchReadiness: true } : {})
}
if (!this.waiters.bindAbort(waiter, options?.signal)) {
@@ -108,6 +108,41 @@ describe('buildWorktreeStartupForAgent host resolution', () => {
})
})
describe('buildWorktreeStartupForAgent prompt carry', () => {
const build = (onPromptCarry?: (carried: boolean) => void, terminalDefaultShell = '/bin/bash') =>
buildWorktreeStartupForAgent({
repo: makeRepo({}),
settings: Object.assign({}, settings, { terminalDefaultShell }),
agent: 'claude',
prompt: 'summarize the diff\nthen list the risks',
getLaunchPlatform: () => 'linux',
toSessionOptions: () => undefined,
...(onPromptCarry ? { onPromptCarry } : {})
})
it('starts clean and reports it when a caller that pastes offers a prompt the line cannot carry', () => {
const onPromptCarry = vi.fn()
const result = build(onPromptCarry)
expect(result.startup.command).not.toContain('summarize')
expect(result.followup).toBeUndefined()
expect(onPromptCarry).toHaveBeenCalledWith(false)
})
it('carries a short-lined multi-line prompt on a local zsh line, as main typed it', () => {
const onPromptCarry = vi.fn()
const result = build(onPromptCarry, '/bin/zsh')
expect(result.startup.command).toContain('summarize the diff\nthen list the risks')
expect(onPromptCarry).toHaveBeenCalledWith(true)
})
it('keeps folding the prompt for a caller that delivers nothing afterwards', () => {
// `orca worktree create --prompt` has no post-start paste of its own for an argv agent.
expect(build().startup.command).toContain('summarize the diff')
})
})
describe('buildWorktreeStartupForDraft agent detection', () => {
it('probes the SSH host named only by executionHostId instead of this client', async () => {
mocks.detectRemoteAgents.mockResolvedValueOnce(['claude'])
@@ -9,6 +9,11 @@ import { isTuiAgent } from '../../shared/tui-agent-config'
import { isTuiAgentEnabled, pickTuiAgent } from '../../shared/tui-agent-selection'
import { resolveAgentStartupPlanInputs } from '../../shared/agent-startup-plan-inputs'
import { buildAgentDraftLaunchPlan, buildAgentStartupPlan } from '../../shared/tui-agent-startup'
import { planStartupWithPromptCandidate } from '../../shared/startup-line-prompt-carry'
import {
launchHostProvesAgentInFront,
nameLocalTypedLineShell
} from './agent-launch-typed-line-shell'
import {
detectInstalledAgentsWithShellPathHydration,
detectRemoteAgents
@@ -128,6 +133,9 @@ export function buildWorktreeStartupForAgent(
toSessionOptions: (
preferences?: AgentLaunchPreferences
) => Parameters<typeof buildAgentStartupPlan>[0]['sessionOptions'] | undefined
/** Set by a caller that delivers an uncarried prompt itself: the prompt then rides only a typed
* line that can carry it, and this reports whether it did. Absent keeps the CLI's fold. */
onPromptCarry?: (carried: boolean) => void
}
): {
agent: TuiAgent
@@ -138,18 +146,36 @@ export function buildWorktreeStartupForAgent(
if (!isTuiAgentEnabled(agent, settings.disabledTuiAgents)) {
throw new Error('Selected agent is disabled. Choose an enabled agent before creating.')
}
const startupPlan = buildAgentStartupPlan({
...resolveAgentStartupPlanInputs({
agent,
settings,
platform: environment.getLaunchPlatform(),
isRemote: repoIsRemote(repo),
...(environment.agentArgs !== undefined ? { agentArgs: environment.agentArgs } : {}),
sessionOptions: environment.toSessionOptions(environment.launchPreferences)
}),
prompt: environment.prompt ?? '',
allowEmptyPromptLaunch: true
const planInputs = resolveAgentStartupPlanInputs({
agent,
settings,
platform: environment.getLaunchPlatform(),
isRemote: repoIsRemote(repo),
...(environment.agentArgs !== undefined ? { agentArgs: environment.agentArgs } : {}),
sessionOptions: environment.toSessionOptions(environment.launchPreferences)
})
const prompt = environment.prompt ?? ''
let startupPlan: ReturnType<typeof buildAgentStartupPlan>
if (environment.onPromptCarry) {
const offered = planStartupWithPromptCandidate(planInputs, prompt, {
shellName: nameLocalTypedLineShell({
isRemote: repoIsRemote(repo),
...(settings.terminalDefaultShell
? { defaultShellSetting: settings.terminalDefaultShell }
: {})
}),
provesAgentInFront: launchHostProvesAgentInFront({
isRemote: repoIsRemote(repo),
launchPlatform: environment.getLaunchPlatform()
})
})
startupPlan = offered.plan
if (startupPlan && prompt.trim()) {
environment.onPromptCarry(offered.promptCarried)
}
} else {
startupPlan = buildAgentStartupPlan({ ...planInputs, prompt, allowEmptyPromptLaunch: true })
}
if (!startupPlan) {
throw new Error(`Could not build launch command for ${agent}.`)
}
@@ -181,7 +207,11 @@ export function resolveWorktreeCreateAgentStartup(
agent: TuiAgent,
prompt: string | undefined,
preferences: AgentLaunchPreferences | undefined,
inputs: { agentArgs?: string | null; launchSource?: string }
inputs: {
agentArgs?: string | null
launchSource?: string
onPromptCarry?: (carried: boolean) => void
}
) => { agent: TuiAgent; startup: WorktreeStartupLaunch; followup?: WorktreeStartupFollowup }
) {
if (args.startup || !args.startupAgent) {
@@ -189,6 +219,7 @@ export function resolveWorktreeCreateAgentStartup(
}
return build(args.startupAgent, args.startupPrompt, args.startupLaunchPreferences, {
...(args.startupAgentArgs !== undefined ? { agentArgs: args.startupAgentArgs } : {}),
...(args.startupLaunchSource ? { launchSource: args.startupLaunchSource } : {})
...(args.startupLaunchSource ? { launchSource: args.startupLaunchSource } : {}),
...(args.onStartupPromptCarry ? { onPromptCarry: args.onStartupPromptCarry } : {})
})
}
@@ -13,6 +13,9 @@ import type {
const BRACKETED_PASTE_BEGIN = '\x1b[200~'
const BRACKETED_PASTE_END = '\x1b[201~'
const BRACKETED_PASTE_QUIET_MS = 1500
// Why: an interactive shell turns bracketed paste on at its prompt and off when it runs the typed
// command (`zsh-prompt-runs-command.txt`), so a 2004 before the last `?2004l` is the shell's.
const DECRST_BRACKETED_PASTE = '\x1b[?2004l'
export type WorktreeStartupReadinessHost = {
getPtyId: (handle: string) => string | null
@@ -87,24 +90,47 @@ export async function waitForWorktreeStartupFollowup(
return null
}
export type StartupDraftReadinessOptions = {
timeoutMs?: number
requireComposerMarker?: boolean
signal?: AbortSignal
/** Vetoes a ready signal whose screen still holds something the input must not answer; the
* scan continues, so the agent's next marker or quiet window asks again. */
accept?: (ptyId: string) => boolean | Promise<boolean>
/**
* For a launched agent. With it, only output after the shell's last `?2004l` counts, since the
* shell's own prompt enables bracketed paste too, and a ready signal is dropped while a shell is
* proven in front: the launch line has not run yet, or the agent exited.
*/
isShellInFront?: (ptyId: string) => Promise<boolean>
}
export function waitForWorktreeStartupDraft(
host: WorktreeStartupReadinessHost,
handle: string,
agent: TuiAgent,
options: { timeoutMs?: number; requireComposerMarker?: boolean } = {}
options: StartupDraftReadinessOptions = {}
): Promise<string | null> {
const ptyId = host.getPtyId(handle)
if (!ptyId) {
if (!ptyId || options.signal?.aborted) {
return Promise.resolve(null)
}
const signal =
TUI_AGENT_CONFIG[agent].draftPasteReadySignal ?? 'render-quiet-after-bracketed-paste'
const isShellInFront = options.isShellInFront
return new Promise((resolve) => {
let settled = false
const scanner = createDraftPasteReadyScanner(signal)
let scanner = createDraftPasteReadyScanner(signal)
let quietTimer: NodeJS.Timeout | null = null
let hardTimer: NodeJS.Timeout | null = null
let unsubscribe: (() => void) | null = null
// Bumped at each shell hand-off: a check begun before one must not settle the wait after it.
let handoffs = 0
let handoffCarry = ''
let checking = false
/** The hand-off count at a signal that fired while a check ran. */
let recheckAt: number | null = null
const onAbort = (): void => finish(null)
const finish = (value: string | null): void => {
if (settled) {
return
@@ -117,23 +143,74 @@ export function waitForWorktreeStartupDraft(
clearTimeout(hardTimer)
}
unsubscribe?.()
options.signal?.removeEventListener('abort', onAbort)
resolve(value)
}
const observe = (data: string): void => {
/** Settles a fired signal once its screen is clear and no shell is proven in front. */
const settleSignal = async (signalHandoffs: number): Promise<void> => {
checking = true
try {
if (options.accept && !(await options.accept(ptyId))) {
return
}
if (isShellInFront && (await isShellInFront(ptyId))) {
return
}
if (signalHandoffs === handoffs) {
finish(ptyId)
}
} catch {
// A check that failed is no settle; the agent's next signal asks again.
} finally {
checking = false
const next = recheckAt
recheckAt = null
if (next === handoffs && !settled) {
void settleSignal(next)
}
}
}
const onSignal = (): void => {
if (checking) {
recheckAt = handoffs
return
}
void settleSignal(handoffs)
}
/** The part of a chunk after the shell's last hand-off in it, resetting the scan at one. */
const sinceShellHandoff = (chunk: string): string => {
const window = handoffCarry + chunk
// Why 7: one short of the sequence, so a split one is rejoined and never counted twice.
handoffCarry = window.slice(-(DECRST_BRACKETED_PASTE.length - 1))
const handoff = window.lastIndexOf(DECRST_BRACKETED_PASTE)
if (handoff === -1) {
return chunk
}
handoffs += 1
scanner = createDraftPasteReadyScanner(signal)
if (quietTimer) {
clearTimeout(quietTimer)
quietTimer = null
}
return window.slice(handoff + DECRST_BRACKETED_PASTE.length)
}
const observe = (chunk: string): void => {
if (settled) {
return
}
const data = isShellInFront ? sinceShellHandoff(chunk) : chunk
const result = scanner.observe(data)
if (result.ready) {
return finish(ptyId)
return onSignal()
}
if (result.armQuietTimer && !options.requireComposerMarker) {
if (quietTimer) {
clearTimeout(quietTimer)
}
quietTimer = setTimeout(() => finish(ptyId), BRACKETED_PASTE_QUIET_MS)
quietTimer = setTimeout(onSignal, BRACKETED_PASTE_QUIET_MS)
}
}
options.signal?.addEventListener('abort', onAbort)
unsubscribe = host.subscribeToData(ptyId, observe)
hardTimer = setTimeout(
() => finish(null),
@@ -0,0 +1,71 @@
import { describe, expect, it } from 'vitest'
import type { ProcessTableRow } from '../../shared/process-table-snapshot'
import { judgeTerminalForeground } from './terminal-foreground-group'
type Row = Omit<ProcessTableRow, 'tty' | 'startTime'>
/** A macOS pane under `login`, as `ps -o pid=,ppid=,pgid=,tpgid=,stat=,command= -t` reads it. */
function pane(tpgid: number, jobs: Omit<Row, 'tpgid'>[]): Row[] {
return [
{
pid: 100,
ppid: 1,
pgid: 100,
tpgid,
stat: 'Ss',
command: '/usr/bin/login -flpq user /bin/bash --noprofile --norc -p -c'
},
{ pid: 101, ppid: 100, pgid: 101, tpgid, stat: tpgid === 101 ? 'S+' : 'S', command: '-zsh' },
...jobs.map((job) => ({ ...job, tpgid }))
]
}
describe('who holds a pane’s terminal, by its foreground process group', () => {
it.each([
// The launch line has not run, or the agent it ran exited and handed the terminal back.
['the pane’s shell at its prompt', pane(101, []), 'shell'],
[
'the agent the launch line ran',
pane(200, [{ pid: 200, ppid: 101, pgid: 200, stat: 'S+', command: '/opt/bin/claude' }]),
'agent'
],
// A tcsh or nu launch line, or a wrapper script that does not `exec` its agent.
[
'the agent behind the sh that leads its group',
pane(200, [
{ pid: 200, ppid: 101, pgid: 200, stat: 'S+', command: '/bin/sh /tmp/orca-launch/run.sh' },
{ pid: 201, ppid: 200, pgid: 200, stat: 'S+', command: '/opt/bin/claude' }
]),
'agent'
],
[
'an agent it cannot recognize behind a bash wrapper',
pane(200, [
{ pid: 200, ppid: 101, pgid: 200, stat: 'S+', command: 'bash ./start-agent.sh' },
{ pid: 201, ppid: 200, pgid: 200, stat: 'S+', command: 'node /opt/agent/cli.js' }
]),
'agent'
],
// Between its commands a wrapper is only shells, and typing into it could run the text.
[
'a bash wrapper alone',
pane(200, [{ pid: 200, ppid: 101, pgid: 200, stat: 'S+', command: 'bash ./start-agent.sh' }]),
'shell'
],
[
'an agent stopped in the background',
pane(101, [{ pid: 200, ppid: 101, pgid: 200, stat: 'T', command: '/opt/bin/claude' }]),
'shell'
]
] as const)('%s: %s', (_label, rows, found) => {
expect(judgeTerminalForeground(rows, 100, 'claude')).toBe(found)
})
it.each([
['the root is not on the terminal', pane(101, []).slice(1)],
['the terminal has no foreground group', pane(0, [])],
['no process is in the foreground group', pane(300, [])]
])('proves nothing when %s', (_label, rows) => {
expect(judgeTerminalForeground(rows, 100, 'claude')).toBe('unknown')
})
})
@@ -0,0 +1,103 @@
/**
* Who holds one pane's terminal, from a `ps` limited to that terminal's processes: a few
* milliseconds, where the whole-machine capture behind a fresh scan or an `inspectProcess`
* observation takes seconds on a loaded host and gated every launch paste behind it.
*/
import { recognizeAgentProcessFromCommandLine } from '../../shared/agent-process-recognition'
import { runProcess } from '../../shared/child-process/run-process'
import { parseShellForegroundRows, type ProcessTableRow } from '../../shared/process-table-snapshot'
import { isShellProcess } from '../../shared/shell-process-detection'
import type { TuiAgent } from '../../shared/tui-agent'
const PS_TIMEOUT_MS = 3_000
const PS_MAX_OUTPUT_BYTES = 256 * 1024
/** Root pid to its terminal: a live process keeps its controlling terminal. */
const ttyByRootPid = new Map<number, string>()
const TTY_CACHE_LIMIT = 256
async function ps(args: readonly string[]): Promise<string | null> {
try {
const result = await runProcess({
program: 'ps',
args,
timeoutMs: PS_TIMEOUT_MS,
maxOutputBytes: PS_MAX_OUTPUT_BYTES
})
return result.code === 0 ? result.stdout : null
} catch {
return null
}
}
async function terminalOf(rootPid: number): Promise<string | null> {
const cached = ttyByRootPid.get(rootPid)
if (cached) {
return cached
}
const tty = (await ps(['-o', 'tty=', '-p', String(rootPid)]))?.trim()
// `??` (macOS) and `?` (Linux): no controlling terminal.
if (!tty || tty.startsWith('?')) {
return null
}
if (ttyByRootPid.size >= TTY_CACHE_LIMIT) {
ttyByRootPid.clear()
}
ttyByRootPid.set(rootPid, tty)
return tty
}
/** Every process on the terminal a pane's root process holds, or null when `ps` cannot say. */
export async function readTerminalProcessRows(rootPid: number): Promise<ProcessTableRow[] | null> {
const tty = await terminalOf(rootPid)
if (!tty) {
return null
}
const stdout = await ps(['-o', 'pid=,ppid=,pgid=,tpgid=,stat=,command=', '-t', tty])
try {
const rows = stdout === null ? null : parseShellForegroundRows(stdout)
// A pid reused by another process on another terminal reads as this pane's root missing.
if (!rows?.some((row) => row.pid === rootPid)) {
ttyByRootPid.delete(rootPid)
return null
}
return rows
} catch {
return null
}
}
function isShellCommand(command: string): boolean {
const executable = command.trim().split(/\s+/, 1)[0]?.replace(/^-/, '') ?? ''
return isShellProcess(executable.split('/').pop() ?? executable)
}
/**
* The terminal's foreground process group decides. The launched agent among its members, or any
* member that is not a shell, is the agent: that finds it behind a wrapper that did not `exec` it,
* whose own shell name leads the group. A group of shells alone is the shell, whether the pane's
* own (the launch line has not run, or the agent exited) or a wrapper between commands. The pane's
* root is never assumed to be the shell: a macOS pane runs its shell under `login`.
*/
export function judgeTerminalForeground(
rows: readonly ProcessTableRow[],
rootPid: number,
agent: TuiAgent
): 'agent' | 'shell' | 'unknown' {
const root = rows.find((row) => row.pid === rootPid)
const foregroundGroup = root?.tpgid
if (foregroundGroup === undefined || foregroundGroup <= 0) {
return 'unknown'
}
const front = rows.filter((row) => row.pgid === foregroundGroup)
if (front.length === 0) {
return 'unknown'
}
return front.some(
(row) =>
recognizeAgentProcessFromCommandLine(row.command)?.agent === agent ||
!isShellCommand(row.command)
)
? 'agent'
: 'shell'
}
+15 -6
View File
@@ -148,12 +148,15 @@ export function nameOnlyIdleNeedsCorroboration(
export function hasSustainedTitleIdle(
record: TuiIdleEvidenceRecord,
agent: TuiAgent | null | undefined,
quiescenceMs: number
quiescenceMs: number,
launchReadiness = false
): boolean {
if (record.lastAgentStatus !== 'idle') {
return false
}
if (!nameOnlyIdleNeedsCorroboration(agent, record.lastOscTitle)) {
// Why launch readiness always corroborates: a shell auto-title (`grok`, `gemini`) names the
// agent before its TUI mounts, and a paste then lands in a booting TUI or the shell itself.
if (!launchReadiness && !nameOnlyIdleNeedsCorroboration(agent, record.lastOscTitle)) {
// The title is the only rest signal this agent emits, so there is nothing to wait for.
return true
}
@@ -214,6 +217,9 @@ export type TuiIdleEvaluationInput = {
/** Tier 0: the hook server's fresh row for the pane, read only for an authoritative agent. */
readHookTurn?: () => TuiIdleHookTurn | null
quiescenceMs: number
/** Waiting for a just-launched agent to open its composer, where a name-only title proves
* nothing until the stream goes quiet. */
launchReadiness?: boolean
}
export type TuiIdleVerdict =
@@ -337,11 +343,11 @@ function rankTuiIdleEvidence(input: TuiIdleEvaluationInput): TuiIdleVerdict {
// refused, and an agent's own idle-title rule replaces the sustained-title lane below.
const ruled = input.readAgentRuleVerdict()
if (ruled !== null) {
return isSettledWeakIdle(ruled, input.record, input.quiescenceMs)
return isSettledWeakIdle(ruled, input.record, input.quiescenceMs, input.launchReadiness)
? READY_WEAK
: { kind: 'pending', quietForeground: 'closed' }
}
if (hasSustainedTitleIdle(input.record, input.agent, input.quiescenceMs)) {
if (hasSustainedTitleIdle(input.record, input.agent, input.quiescenceMs, input.launchReadiness)) {
return READY_WEAK
}
return {
@@ -353,15 +359,18 @@ function rankTuiIdleEvidence(input: TuiIdleEvaluationInput): TuiIdleVerdict {
}
}
// Why launch readiness asks quiet of every weak idle: those rules read a name-only title, which a
// shell auto-title writes before the TUI mounts, as in the sustained-title lane.
function isSettledWeakIdle(
verdict: AgentStateVerdict,
record: TuiIdleEvidenceRecord,
quiescenceMs: number
quiescenceMs: number,
launchReadiness = false
): boolean {
return (
verdict.state === 'idle' &&
verdict.strength === 'weak' &&
(!verdict.requiresQuiet || hasQuietOutput(record, quiescenceMs))
((!verdict.requiresQuiet && !launchReadiness) || hasQuietOutput(record, quiescenceMs))
)
}
@@ -0,0 +1,95 @@
/**
* A shell that titles the command it is about to run (zsh preexec auto-title) writes the agent's
* bare name before the agent's TUI has mounted. For an agent whose only rest signal is that name,
* the title alone used to settle `tui-idle`, so a launch pasted its prompt into a TUI still
* booting, or into the shell. A launch readiness wait holds that title to a quiet stream.
*/
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import { createTranscriptPane, TRANSCRIPT_PANE_PTY_ID } from './agent-transcript-pane-test-harness'
vi.mock('electron', () => ({
BrowserWindow: { fromId: vi.fn(() => null) },
webContents: { fromId: vi.fn(() => null) },
ipcMain: { on: vi.fn(), removeListener: vi.fn() },
app: { getPath: vi.fn(() => '/tmp') }
}))
const ESC = String.fromCharCode(27)
const BEL = String.fromCharCode(7)
const BOOT_FRAME = 'Loading…\r\n'
/** The shell echoes the launch line, then its preexec hook titles the pane with the command. */
async function launchedPaneTitledByShell(agent: 'gemini' | 'copilot') {
return createTranscriptPane({
paneTitle: 'Terminal',
foregroundProcess: agent,
launchAgent: agent,
data: `$ ${agent}\r\n${ESC}]2;${agent}${BEL}`
})
}
/** The agent is still painting its boot screen: output keeps arriving each second. */
async function keepBooting(
runtime: Awaited<ReturnType<typeof launchedPaneTitledByShell>>['runtime'],
seconds: number
): Promise<void> {
for (let second = 0; second < seconds; second += 1) {
runtime.onPtyData(TRANSCRIPT_PANE_PTY_ID, BOOT_FRAME, Date.now())
await vi.advanceTimersByTimeAsync(1000)
}
}
describe('a launch waiting on an agent whose shell titled the pane with its name', () => {
beforeEach(() => {
vi.useFakeTimers()
})
afterEach(() => {
vi.useRealTimers()
})
// Gemini's bare name was once normalized into its rest glyph; Copilot's name is its only signal.
it.each(['gemini', 'copilot'] as const)(
'does not treat the shell’s auto-title as readiness while %s is still booting',
async (agent) => {
const { runtime, handle } = await launchedPaneTitledByShell(agent)
const settled = vi.fn()
runtime
.waitForTerminal(handle, {
condition: 'tui-idle',
timeoutMs: 60_000,
launchReadiness: true
})
.then(settled, settled)
await keepBooting(runtime, 10)
expect(settled).not.toHaveBeenCalled()
}
)
it('settles once the agent’s stream goes quiet under that title', async () => {
const { runtime, handle } = await launchedPaneTitledByShell('gemini')
const settled = vi.fn()
runtime
.waitForTerminal(handle, { condition: 'tui-idle', timeoutMs: 60_000, launchReadiness: true })
.then(settled, settled)
await keepBooting(runtime, 4)
await vi.advanceTimersByTimeAsync(8_000)
expect(settled).toHaveBeenCalledWith(expect.objectContaining({ satisfied: true }))
})
it('keeps settling a later, non-launch wait on the name alone, as agents that never quiet need', async () => {
const { runtime, handle } = await launchedPaneTitledByShell('gemini')
const settled = vi.fn()
runtime
.waitForTerminal(handle, { condition: 'tui-idle', timeoutMs: 60_000 })
.then(settled, settled)
await keepBooting(runtime, 10)
expect(settled).toHaveBeenCalledWith(expect.objectContaining({ satisfied: true }))
})
})
@@ -3,8 +3,10 @@ import {
markTerminalBracketedPasteInterrupted,
observeTerminalBracketedPasteModeOutput,
pasteTerminalText,
sanitizeTerminalPasteText
sanitizeTerminalPasteText,
wrapTerminalBracketedPasteText
} from './terminal-bracketed-paste'
import { wrapTerminalBracketedPasteText as hostPasteFrame } from '../../../../shared/terminal-bracketed-paste-text'
function createTerminal(bracketedPasteMode = true) {
const terminal = {
@@ -21,6 +23,11 @@ function createTerminal(bracketedPasteMode = true) {
}
describe('terminal bracketed paste policy', () => {
// The host's launch prompt replaced the desktop's draft paste; one function keeps their bytes equal.
it('frames pasted text with the same function as the host launch prompt', () => {
expect(wrapTerminalBracketedPasteText).toBe(hostPasteFrame)
})
it('temporarily ignores bracketed paste wrappers for single-line paste after Ctrl+C', () => {
const terminal = createTerminal(true)
const observedIgnoreValues: (boolean | undefined)[] = []
@@ -1,5 +1,20 @@
import type { Terminal } from '@xterm/xterm'
import type { WindowsInputRecordNewline } from './terminal-paste-model'
import {
BRACKETED_PASTE_END,
BRACKETED_PASTE_START,
normalizeTerminalPasteLineEndings,
sanitizeBracketedPasteText,
wrapTerminalBracketedPasteText
} from '../../../../shared/terminal-bracketed-paste-text'
export {
BRACKETED_PASTE_END,
BRACKETED_PASTE_START,
normalizeTerminalPasteLineEndings,
sanitizeBracketedPasteText,
wrapTerminalBracketedPasteText
}
type BracketedPasteTerminal = {
modes: {
@@ -21,8 +36,6 @@ type PasteTerminalTextOptions = {
const interruptedBracketedPasteTerminals = new WeakSet<object>()
const bracketedPasteModeOutputTail = new WeakMap<object, string>()
const ESCAPE = '\u001b'
export const BRACKETED_PASTE_START = `${ESCAPE}[200~`
export const BRACKETED_PASTE_END = `${ESCAPE}[201~`
const BRACKETED_PASTE_MODE_SEQUENCE_RE = /^\[\?(?:\d+;)*2004(?:;\d+)*[hl]/
const BRACKETED_PASTE_MODE_TAIL_MAX = 128
const BRACKETED_PASTE_MODE_SEQUENCE_SCAN_MAX = BRACKETED_PASTE_MODE_TAIL_MAX
@@ -45,40 +58,10 @@ function hasBracketedPasteModeSequence(data: string): boolean {
return false
}
// Why: an embedded ESC (e.g. a pasted `\x1b[201~` from scrollback) would close
// the bracketed-paste frame early and run the tail as keystrokes. Replacing ESC
// with its printable substitute (\u241b, U+241B) neutralizes every framing escape.
export function sanitizeBracketedPasteText(text: string): string {
let escapeIndex = text.indexOf(ESCAPE)
if (escapeIndex === -1) {
return text
}
let sanitized = ''
let start = 0
while (escapeIndex !== -1) {
sanitized += `${text.slice(start, escapeIndex)}\u241b`
start = escapeIndex + ESCAPE.length
escapeIndex = text.indexOf(ESCAPE, start)
}
return sanitized + text.slice(start)
}
export function sanitizeTerminalPasteText(text: string): string {
return sanitizeBracketedPasteText(text)
}
export function normalizeTerminalPasteLineEndings(text: string): string {
// Why: xterm's native paste path converts every clipboard newline to CR.
// Direct frames must match it or ConPTY TUIs can treat raw LF as submit.
return text.replace(/\r?\n/g, '\r')
}
export function wrapTerminalBracketedPasteText(text: string): string {
const normalizedText = normalizeTerminalPasteLineEndings(text)
return `${BRACKETED_PASTE_START}${sanitizeBracketedPasteText(normalizedText)}${BRACKETED_PASTE_END}`
}
export function encodeWindowsInputRecordPasteText(
text: string,
newline: WindowsInputRecordNewline
+2 -1
View File
@@ -2,6 +2,7 @@ import type { GlobalSettings } from '../../../shared/global-settings-types'
import type { TuiAgent } from '../../../shared/tui-agent'
import type { TerminalInputKind } from '../../../shared/terminal-input-kind'
import { TUI_AGENT_CONFIG } from '../../../shared/tui-agent-config'
import { AGENT_PROMPT_POST_PASTE_SUBMIT_DELAY_MS } from '../../../shared/agent-prompt-injection'
import { resolveDraftPasteReadyTimeoutMs } from '../../../shared/draft-paste-ready-timeout'
import { useAppStore } from '@/store'
import {
@@ -34,7 +35,7 @@ export {
// line-edit shortcuts. Callers choose whether to append Enter after the paste.
export const BRACKETED_PASTE_BEGIN = BRACKETED_PASTE_START
export { BRACKETED_PASTE_END }
export const POST_PASTE_SUBMIT_DELAY_MS = 50
export const POST_PASTE_SUBMIT_DELAY_MS = AGENT_PROMPT_POST_PASTE_SUBMIT_DELAY_MS
// Why: "the tab has a PTY" and "the agent's composer accepts input" are separate
// states with separate failure modes, so they get separate budgets. A PTY that
+52
View File
@@ -0,0 +1,52 @@
/**
* What an `agent.launchReplay` refusal proves about the launch, for every client that replays one.
*
* Kept in one place because the three answers call for opposite actions: `unsupported` means nothing
* ran and the client may use its older launch path; `unknown` means the agent may be running and the
* client must not launch again on its own; `failed` means the host refused before starting anything.
*/
import { AGENT_LAUNCH_PANE_ALREADY_LIVE_CODE } from './agent-launch-pane-already-live'
import { AGENT_LAUNCH_SESSION_ALREADY_EXISTS_CODE } from './agent-launch-session-already-exists'
export type AgentLaunchReplayRefusal = 'unsupported' | 'unknown' | 'failed'
/** The pane or chat the launch reserved is already held. */
export function isAgentLaunchReservationTakenRefusal(error: { code?: string }): boolean {
return (
error.code === AGENT_LAUNCH_PANE_ALREADY_LIVE_CODE ||
error.code === AGENT_LAUNCH_SESSION_ALREADY_EXISTS_CODE
)
}
/** An older host rejects the method rather than a field. */
export function isAgentLaunchReplayUnsupportedRefusal(error: { code?: string }): boolean {
return (
error.code === 'method_not_found' ||
error.code === 'forbidden' ||
error.code === 'agent_launch_replay_unsupported'
)
}
export function classifyAgentLaunchReplayRefusal(
error: { code?: string },
replayed: boolean
): AgentLaunchReplayRefusal {
if (isAgentLaunchReplayUnsupportedRefusal(error)) {
// Only a refusal of the first send proves nothing ran; after a replay it may be a replacement
// connection whose capability list hasn't landed, answering for an attempt that did start.
return replayed ? 'unknown' : 'unsupported'
}
if (
error.code === 'agent_session_operation_unknown' ||
error.code === 'agent_session_operation_expired'
) {
return 'unknown'
}
if (isAgentLaunchReservationTakenRefusal(error)) {
// A taken reservation proves nothing started only on the first send; after a replay the pane
// or chat holding it may be this launch's own.
return replayed ? 'unknown' : 'failed'
}
return 'failed'
}
@@ -0,0 +1,37 @@
// Split out of protocol-version.ts, which spreads the list below into RUNTIME_CAPABILITIES in this
// order. Import these names from here: the mobile recording loader cannot follow `export *`.
/**
* `agent.launch` exists: one host-side method that decides structured-vs-terminal and creates the
* surface, instead of each client routing for itself.
*
* Negotiated rather than assumed because a client that cannot see it must keep using
* `worktree.create` + `startupAgent`, which stays supported verbatim. The reverse skew is the
* dangerous one: `worktree.create` returns `agentTerminalHandle` only when a startup agent was
* requested, so a host that quietly routed that call to a structured session would hand an old
* client a response with no handle and no error.
*
* Advertising it is a statement that the client understands EITHER outcome, since the host is what
* picks: a structured session it can open, or a terminal agent. A client that renders only one of
* the two keeps using the surface-specific methods.
*/
// v2 makes prompt delivery an outcome union and top-level warnings the only supported shape.
export const AGENT_LAUNCH_RUNTIME_CAPABILITY = 'agent.launch.v2' as const
// Optional identity support on agent.launch; mobile replay across replacement hosts requires the new method.
export const AGENT_LAUNCH_REPLAY_RUNTIME_CAPABILITY = 'agent.launch.replay.v1' as const
// A host that sends a launch prompt its typed startup line cannot carry to a paste after
// readiness; an older host folds any prompt into that line, so clients gate prompted launches on it.
export const AGENT_LAUNCH_PROMPT_CARRY_RUNTIME_CAPABILITY = 'agent.launch.prompt-carry.v1' as const
// agent.launchReplay requires the ledger; older replacement hosts must reject the method.
export const AGENT_LAUNCH_REPLAY_REQUIRED_RUNTIME_CAPABILITY =
'agent.launch.replay-required.v1' as const
export const AGENT_LAUNCH_RUNTIME_CAPABILITIES = [
AGENT_LAUNCH_RUNTIME_CAPABILITY,
AGENT_LAUNCH_REPLAY_RUNTIME_CAPABILITY,
AGENT_LAUNCH_REPLAY_REQUIRED_RUNTIME_CAPABILITY,
AGENT_LAUNCH_PROMPT_CARRY_RUNTIME_CAPABILITY
] as const
+3
View File
@@ -6,6 +6,9 @@ import { isTuiAgent, TUI_AGENT_CONFIG } from './tui-agent-config'
export const AGENT_PROMPT_BRACKETED_PASTE_START = '\x1b[200~'
export const AGENT_PROMPT_BRACKETED_PASTE_END = '\x1b[201~'
export const AGENT_PROMPT_SUBMIT = '\r'
/** Why: Claude Code can leave a prompt as editable text when paste-end and Enter arrive in the
* same PTY write, so the desktop's draft paste sends Enter on the next turn, this long after. */
export const AGENT_PROMPT_POST_PASTE_SUBMIT_DELAY_MS = 50
/** Why unknown agents keep the lead: an unidentified Claude still needs it, while known non-Claude
* TUIs get pre-lead bytes because Codex drops typed text that shares the paste's write (STA-8200). */
+3 -1
View File
@@ -153,7 +153,9 @@ export function normalizeTerminalTitle(title: string): string {
if (status === 'working') {
return `${GEMINI_WORKING} Gemini CLI`
}
if (status === 'idle') {
// Why only with the glyph: a bare `gemini` title (a shell auto-title) reads idle by default,
// and stamping Gemini's rest glyph on it turned a name into explicit readiness.
if (status === 'idle' && title.includes(GEMINI_IDLE)) {
return `${GEMINI_IDLE} Gemini CLI`
}
}
+5
View File
@@ -101,6 +101,11 @@ const DRAFT_PASTE_READY_SIGNALS: Record<DraftPasteReadySignal, DraftPasteReadySi
/** Longest anchor sequence minus one — the carry needed to rejoin one split across chunks. */
const ANCHOR_CARRY_CHARS = 7
/** Whether the signal has a composer marker, rather than only a quiet window after its anchor. */
export function draftPasteReadySignalHasMarker(readySignal: DraftPasteReadySignal): boolean {
return DRAFT_PASTE_READY_SIGNALS[readySignal].marker !== null
}
export type DraftPasteReadyScanResult = {
/** The agent-specific ready signal fired — caller should deliver the paste now. */
ready: boolean
+5 -27
View File
@@ -27,6 +27,10 @@ import {
SKILL_UPLOAD_CAPABILITY
} from './skill-install-capability'
export { SKILL_INSTALL_RESULT_V2_CAPABILITY } from './skill-install-capability'
import {
AGENT_LAUNCH_RUNTIME_CAPABILITIES,
AGENT_LAUNCH_RUNTIME_CAPABILITY
} from './agent-launch-runtime-capability'
// Why: declares the Orca runtime RPC compatibility contract. Desktop,
// headless server, CLI, and mobile builds may drift in app version, but
@@ -309,30 +313,6 @@ export const AUTOMATION_CREATE_IDEMPOTENCY_RUNTIME_CAPABILITY =
// Hosts without this capability have no notifications.registerPush RPC.
export const NOTIFICATIONS_REMOTE_PUSH_RUNTIME_CAPABILITY = 'notifications.remote-push.v1' as const
/**
* `agent.launch` exists: one host-side method that decides structured-vs-terminal and creates the
* surface, instead of each client routing for itself.
*
* Negotiated rather than assumed because a client that cannot see it must keep using
* `worktree.create` + `startupAgent`, which stays supported verbatim. The reverse skew is the
* dangerous one: `worktree.create` returns `agentTerminalHandle` only when a startup agent was
* requested, so a host that quietly routed that call to a structured session would hand an old
* client a response with no handle and no error.
*
* Advertising it is a statement that the client understands EITHER outcome, since the host is what
* picks: a structured session it can open, or a terminal agent. A client that renders only one of
* the two keeps using the surface-specific methods.
*/
// v2 makes prompt delivery an outcome union and top-level warnings the only supported shape.
export const AGENT_LAUNCH_RUNTIME_CAPABILITY = 'agent.launch.v2' as const
// Optional identity support on agent.launch; mobile replay across replacement hosts requires the new method.
export const AGENT_LAUNCH_REPLAY_RUNTIME_CAPABILITY = 'agent.launch.replay.v1' as const
// agent.launchReplay requires the ledger; older replacement hosts must reject the method.
export const AGENT_LAUNCH_REPLAY_REQUIRED_RUNTIME_CAPABILITY =
'agent.launch.replay-required.v1' as const
// Generic native clients include the CLI and must not claim Electron-only page
// placement support.
export const NATIVE_REMOTE_RUNTIME_CLIENT_CAPABILITIES = [
@@ -452,9 +432,7 @@ export const RUNTIME_CAPABILITIES = [
AUTOMATION_OWNER_FENCING_RUNTIME_CAPABILITY,
AUTOMATION_CREATE_IDEMPOTENCY_RUNTIME_CAPABILITY,
NOTIFICATIONS_REMOTE_PUSH_RUNTIME_CAPABILITY,
AGENT_LAUNCH_RUNTIME_CAPABILITY,
AGENT_LAUNCH_REPLAY_RUNTIME_CAPABILITY,
AGENT_LAUNCH_REPLAY_REQUIRED_RUNTIME_CAPABILITY
...AGENT_LAUNCH_RUNTIME_CAPABILITIES
] as const
export type RuntimeCapability = (typeof RUNTIME_CAPABILITIES)[number] | (string & {})
@@ -0,0 +1,207 @@
import { describe, expect, it } from 'vitest'
import {
hasControlByte,
planStartupWithPromptCandidate,
TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES,
ZSH_MULTI_LINE_STARTUP_LINE_BUDGET_BYTES
} from './startup-line-prompt-carry'
import type { TuiAgent } from './tui-agent'
import { RUNTIME_CAPABILITIES } from './protocol-version'
import { AGENT_LAUNCH_PROMPT_CARRY_RUNTIME_CAPABILITY } from './agent-launch-runtime-capability'
function offer(
agent: TuiAgent,
prompt: string,
extra: {
cmdOverride?: string
shellName?: string | undefined
platform?: NodeJS.Platform
provesAgentInFront?: boolean
} = {}
) {
return planStartupWithPromptCandidate(
{
agent,
cmdOverrides: extra.cmdOverride ? { [agent]: extra.cmdOverride } : {},
platform: extra.platform ?? 'darwin'
},
prompt,
{
...(extra.shellName ? { shellName: extra.shellName } : {}),
provesAgentInFront: extra.provesAgentInFront ?? true
}
)
}
/** `count` prompt lines of `bytes` each, as a multi-line source-control prompt is. */
function linesOf(count: number, bytes: number): string {
return Array.from({ length: count }, (_unused, index) => `${index}:`.padEnd(bytes, 'x')).join(
'\n'
)
}
/** A prompt that brings the quoted `claude '<prompt>'` line to exactly `bytes`. */
function promptForClaudeLineOf(bytes: number): string {
const base = offer('claude', '').plan?.launchCommand ?? ''
// `<base> '<prompt>'`: one space and two quotes around the text.
return 'x'.repeat(bytes - base.length - 3)
}
describe('whether a launch prompt rides the typed startup line', () => {
it('carries a short single-line prompt on the launch command', () => {
const { plan, promptCarried } = offer('claude', 'explain this repo')
expect(promptCarried).toBe(true)
expect(plan?.launchCommand).toContain('explain this repo')
})
it('carries a line of exactly the budget and refuses one byte past it', () => {
const atBudget = offer('claude', promptForClaudeLineOf(TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES))
expect(new TextEncoder().encode(atBudget.plan?.launchCommand ?? '').byteLength).toBe(
TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES
)
expect(atBudget.promptCarried).toBe(true)
const pastBudget = offer(
'claude',
promptForClaudeLineOf(TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES + 1)
)
expect(pastBudget.promptCarried).toBe(false)
})
it('measures the quoted line, so quote-heavy text under the budget raw can exceed it typed', () => {
// 200 quotes are 200 raw bytes but 600 once portable quoting expands each to `"'"`.
const quotes = "'".repeat(200)
expect(quotes.length).toBeLessThan(TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES)
const { plan, promptCarried } = offer('claude', quotes)
expect(promptCarried).toBe(false)
expect(plan?.launchCommand).not.toContain(`"'"`)
})
it.each([
['LF', 'first line\nsecond line'],
['CRLF', 'first line\r\nsecond line'],
['CR', 'first line\rsecond line']
])('never types a %s-bearing prompt, which a shell would read as Enter', (_label, prompt) => {
const { plan, promptCarried } = offer('codex', prompt)
expect(promptCarried).toBe(false)
// The clean launch: the prompt is left for the paste after start.
expect(plan?.launchCommand).not.toContain('first line')
expect(plan?.followupPrompt).toBeNull()
})
it.each([
['TAB', 'see\tthis'],
['ESC', 'red \x1b[31mtext'],
['^C', 'stop\x03here'],
['^U', 'kill\x15line'],
['DEL', 'erase\x7fme']
])(
'never types a %s-bearing prompt, which a line editor would read as a key',
(_label, prompt) => {
const { plan, promptCarried } = offer('claude', prompt)
expect(promptCarried).toBe(false)
expect(hasControlByte(plan?.launchCommand ?? '')).toBe(false)
}
)
it('counts the launcher toward the line, so a long configured one leaves no room for the prompt', () => {
const launcher = `claude ${'--add-dir /very/long/path '.repeat(30)}`.trim()
expect(launcher.length).toBeGreaterThan(TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES)
const { plan, promptCarried } = offer('claude', 'hi', { cmdOverride: launcher })
expect(promptCarried).toBe(false)
expect(plan?.launchCommand).toBe(launcher)
})
it('carries a Hermes prompt through its env transport, whose typed line never holds the text', () => {
const multiLine = 'line one\nline two'
const { plan, promptCarried } = offer('hermes', multiLine)
expect(promptCarried).toBe(true)
expect(plan?.launchCommand).not.toContain('line one')
expect(Object.values(plan?.env ?? {})).toContain(multiLine)
})
it('launches Hermes clean instead of refusing when its env budget cannot hold the prompt', () => {
const { plan, promptCarried } = offer('hermes', 'x'.repeat(30_000))
expect(promptCarried).toBe(false)
expect(plan).not.toBeNull()
expect(Object.values(plan?.env ?? {}).join('')).not.toContain('xxxx')
})
it('never carries a stdin-after-start agent’s prompt, whose CLI takes none', () => {
const { plan, promptCarried } = offer('aider', 'fix it')
expect(promptCarried).toBe(false)
expect(plan?.followupPrompt).toBeNull()
})
it('reports nothing carried for an empty prompt', () => {
expect(offer('claude', ' ').promptCarried).toBe(false)
})
})
// Pinned live by startup-line-typed-length.live-shell.test.ts; the refusals below were measured on
// macOS through the same write: a 1.1 KB single line lost to zsh's late write, every multi-line line
// lost to bash 3.2, and to fish whenever its config outlasted the ready barrier.
describe('a multi-line prompt typed into a shell the host names', () => {
it('rides a zsh line as main typed it when each line is short, so the agent starts with it', () => {
const prompt = linesOf(21, 100)
const { plan, promptCarried } = offer('claude', prompt, { shellName: 'zsh' })
expect(promptCarried).toBe(true)
expect(plan?.launchCommand).toContain(prompt)
})
it('keeps a zsh line off the launch command when any one line exceeds the per-line budget', () => {
const prompt = `${linesOf(3, 100)}\n${'y'.repeat(TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES + 1)}`
expect(offer('claude', prompt, { shellName: 'zsh' }).promptCarried).toBe(false)
})
it('keeps a zsh line off the launch command past the whole-line budget', () => {
const prompt = linesOf(Math.ceil(ZSH_MULTI_LINE_STARTUP_LINE_BUDGET_BYTES / 400) + 1, 400)
expect(offer('claude', prompt, { shellName: 'zsh' }).promptCarried).toBe(false)
})
it('keeps a long single line off a zsh launch command, which a late write truncates', () => {
const prompt = promptForClaudeLineOf(TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES + 1)
expect(offer('claude', prompt, { shellName: 'zsh' }).promptCarried).toBe(false)
})
it.each([['bash'], ['fish'], ['pwsh'], [undefined]])(
'never types a multi-line line into %s',
(shellName) => {
const { plan, promptCarried } = offer('claude', linesOf(3, 50), { shellName })
expect(promptCarried).toBe(false)
expect(plan?.launchCommand).not.toContain('0:')
}
)
it.each([
['TAB', 'see\tthis\nnext line'],
['CR', 'first\r\nsecond']
])('never types a %s-bearing multi-line prompt into zsh', (_label, prompt) => {
expect(offer('claude', prompt, { shellName: 'zsh' }).promptCarried).toBe(false)
})
})
// Why: a host that cannot prove the agent took the terminal could paste into the shell of one that
// exited, so there the line carries the prompt whatever its size, as it did before.
describe('on a host that cannot prove the launched agent is in front', () => {
it.each([
['a long single line', 'x'.repeat(TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES * 4)],
['a multi-line prompt', linesOf(5, 40)]
])('carries %s on the launch line', (_label, prompt) => {
const { plan, promptCarried } = offer('claude', prompt, {
platform: 'win32',
provesAgentInFront: false
})
expect(promptCarried).toBe(true)
expect(plan?.launchCommand).toContain(prompt.split('\n')[0])
// Control: the same prompt is pasted where the host can prove the agent.
expect(offer('claude', prompt, { platform: 'win32' }).promptCarried).toBe(false)
})
})
describe('the capability clients gate a prompted launch on', () => {
it('is advertised by every host that applies the typed-line rule', () => {
// An older host folds any prompt into the typed line, so its absence is the gate.
expect(RUNTIME_CAPABILITIES).toContain(AGENT_LAUNCH_PROMPT_CARRY_RUNTIME_CAPABILITY)
})
})
+123
View File
@@ -0,0 +1,123 @@
/**
* Whether a launch prompt rides the command line that gets TYPED into the user's shell, or the agent
* starts clean and the prompt is pasted once it is ready. A paste needs a host that can prove the
* agent holds its terminal (`launched-agent-foreground`); elsewhere the line carries it, as on main.
*
* Measured on the built line, not the raw prompt: quoting, the launcher, its arguments and session
* options all land on that line, and every failure a long or multi-line typed line has is a property
* of the line — macOS bash 3.2 reads each newline as Enter, a line editor reads any other control
* byte as a key, a canonical-mode write truncates a line
* past MAX_CANON (1024 on macOS), and cmd caps a line at 8191. Decided here, where the line exists,
* so the answer is a fact about what was built rather than a prediction of it.
*
* `startup-line-typed-length.live-shell.test.ts` types lines through Orca's own ready barrier and
* submission into real shells, including one whose config outlasts the barrier's 1.5 s cap so the
* line lands while the terminal is still line-buffered. Only zsh read a multi-line line whole in
* both cases, and only while each of its lines stayed short; a long single line was lost there.
*/
import { TUI_AGENT_CONFIG } from './tui-agent-config'
import { buildAgentStartupPlan, type AgentStartupPlan } from './tui-agent-startup'
import type { TuiAgent } from './tui-agent'
/** Half of macOS MAX_CANON: a single typed line this long survived every canonical-mode write
* measured, and 1 KiB did not. Also the cap on each line of a multi-line one. */
export const TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES = 512
/** A multi-line line typed into zsh: the largest measured intact behind a late write (8 KiB). */
export const ZSH_MULTI_LINE_STARTUP_LINE_BUDGET_BYTES = 8192
const encoder = new TextEncoder()
function typedLineBytes(line: string): number {
return encoder.encode(line).byteLength
}
export function startupLineCarriesPrompt(args: {
agent: TuiAgent
withPrompt: AgentStartupPlan | null
/** The shell the line is typed into, when the host can name it before the spawn. */
shellName?: string
/** Whether the host can prove the launched agent holds its terminal before it pastes. */
hostProvesAgentInFront: boolean
}): boolean {
const { withPrompt } = args
if (!withPrompt || withPrompt.followupPrompt !== null) {
return false
}
// Where nothing can prove the agent took the terminal, a paste could land in the shell of an
// agent that exited, so the line carries the prompt at any size, as it always did.
if (!args.hostProvesAgentInFront) {
return true
}
// Hermes types a fixed line that reads the prompt from the spawn env, so the line never grows
// with the text; its own env budget already returned null above when the text did not fit.
if (TUI_AGENT_CONFIG[args.agent].promptInjectionMode === 'hermes-query') {
return true
}
const line = withPrompt.launchCommand
if (!hasControlByte(line)) {
return typedLineBytes(line) <= TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES
}
return args.shellName === 'zsh' && zshReadsMultiLineWhole(line)
}
/**
* zsh takes a multi-line line as one bracketed paste, and a write that beats its line editor still
* reaches it line by line, so each line must fit the terminal's line buffer. bash 3.2 has no
* bracketed paste, and fish drops a paste written before its reader is up.
*/
function zshReadsMultiLineWhole(line: string): boolean {
const lines = line.split('\n')
return (
lines.every(
(part) =>
!hasControlByte(part) && typedLineBytes(part) <= TYPED_STARTUP_LINE_PROMPT_BUDGET_BYTES
) && typedLineBytes(line) <= ZSH_MULTI_LINE_STARTUP_LINE_BUDGET_BYTES
)
}
/** Any C0 byte or DEL, not just CR/LF: no quoter escapes them, and a single-line command is written
* raw, so a TAB completes, ESC starts a key sequence, and ^C/^U/^W kill or edit the line. */
export function hasControlByte(line: string): boolean {
for (let i = 0; i < line.length; i += 1) {
const code = line.charCodeAt(i)
if (code < 0x20 || code === 0x7f) {
return true
}
}
return false
}
type StartupPlanInputs = Omit<
Parameters<typeof buildAgentStartupPlan>[0],
'prompt' | 'allowEmptyPromptLaunch'
>
/**
* The startup plan for a launch that offers a prompt: the prompted plan when its typed line can
* carry the text, else the clean plan, with which one it was.
*/
export function planStartupWithPromptCandidate(
inputs: StartupPlanInputs,
prompt: string,
host: { shellName?: string; provesAgentInFront: boolean }
): { plan: AgentStartupPlan | null; promptCarried: boolean } {
if (prompt.trim()) {
const withPrompt = buildAgentStartupPlan({ ...inputs, prompt, allowEmptyPromptLaunch: true })
if (
startupLineCarriesPrompt({
agent: inputs.agent,
withPrompt,
...(host.shellName ? { shellName: host.shellName } : {}),
hostProvesAgentInFront: host.provesAgentInFront
})
) {
return { plan: withPrompt, promptCarried: true }
}
}
return {
plan: buildAgentStartupPlan({ ...inputs, prompt: '', allowEmptyPromptLaunch: true }),
promptCarried: false
}
}
+1
View File
@@ -143,6 +143,7 @@ export const launchSourceSchema = z.enum([
'conflict_resolution',
'source_control_recovery',
'terminal_context_menu',
'explain_commit',
// Launches the host performs for a caller outside the desktop app.
'cli',
'mobile',

Some files were not shown because too many files have changed in this diff Show More