Files
orca/src/main/runtime/structured-claude-scripted-runtime-test-support.ts
T
Brennan Benson c6cfcc034e refactor(native-chat): structured chat failures always reach the diagnostics log (#24312)
* refactor(native-chat): give the structured chat host one required logger

The structured chat runtime took an optional onError callback that the
desktop never passed, so a late dispatch settlement, an unanswered-dispatch
release, a journal event-sink write and a provider lifecycle delivery that
failed were dropped with no trace. Other host failures went to scattered
console.warn calls, which reach nothing in a packaged desktop build.

The runtime and host now take one required logger (warn/error with a scope
and fields). The production logger writes each entry as a failed span to
<userData>/logs/main.trace.ndjson, which the diagnostic bundle collects, and
to the console (stderr under a supervised headless host). The runtime and the
host wrap it so a logger that throws never fails what it reports, and the
install refuses without one. Sites that deliberately kept a recovery-capsule
error out of the log still log no error object.

* refactor(native-chat): hand the chat host's collaborators the logger, and give orcad its trace file

The delivery loop, idle sweep, queued-message drain, lease renewer, event
sink, conversation map and provider start/exit settlement each took an
internal error callback that the host mapped onto the logger. They now take
the logger itself and log under their own scope. The event sink keeps one
onFailed hook, which decides whether to stop the provider, not whether to
report. The dead-generation settlement returns its failure so each caller
logs it under its own scope.

orcad now installs the desktop's local trace sink under its own data root, so
a headless host's chat failures reach <data-root>/logs/main.trace.ndjson as
well as stderr.

Also passes the logger in the test fixtures the first commit missed, which
tc:node caught.

* fix(native-chat): keep repeated chat failures from flooding the trace file, and record their causes

- The production structured-chat logger writes a repeated failure (same level, scope, session,
  message and error text) once per 5 minutes, carrying how many repeats it swallowed; the
  tracked set is capped at 256.
- Trace entries now carry the error's code (and SQLite errcode) and up to three causes by name and
  message.
- A chat read whose conversation will not open is logged through the host's logger
  (open-for-read), and so are the runtime's chat-tab bookkeeping failures that already hold the
  host.
- orcad writes its own orcad.trace.ndjson, closes it after every quit handler, and flushes it on
  process exit; a trace file that cannot be opened leaves tracing off instead of stopping the app
  or orcad.
- Tests: the desktop wiring test proves the logger reaches the trace sink, and the privacy tests
  read every level the logger received.

* fix(native-chat): log a created chat's tab-publication and launch-prompt failures through the host's logger

* fix(native-chat): key a repeated chat failure on everything its entry writes

The repeat suppression keyed on the message and the error's text, so two refusals with the same
code but different causes, a plain error and a refusal of one code, or two object-valued errors
shared a key and the second was swallowed for five minutes. The key is now the entry's whole
written content (fields, code, errcode, refusal reason, cause chain, a stable rendering of a
non-error value) plus the error's name and message; a refusal's reason is also written.

* test(native-chat): pin that an error's name keeps two repeated failures apart

* test(native-chat): build the refusal in the repeat-key test as the wire does

* fix(native-chat): read an error's code and a refusal's reason by narrowing, not Reflect.get
2026-10-01 11:21:50 -07:00

271 lines
11 KiB
TypeScript

// A structured-session runtime whose Claude children are scripted: the production runtime,
// adapter, record store and host, with only the CLI process replaced.
import { providerDiagnostic, withProviderDiagnostic } from '../../shared/agent-session-failure'
import { mkdir, mkdtemp, rm } from 'node:fs/promises'
import { tmpdir } from 'node:os'
import { join } from 'node:path'
import type {
ClaudeStreamJsonConnection,
ClaudeStreamJsonConnectionHandlers,
ClaudeStreamJsonLaunch,
openClaudeStreamJsonConnection
} from '../claude/claude-stream-json-connection'
import { runClaudeControl } from '../claude/claude-agent-sdk-control-requests'
import { claudeSessionIdForOrcaSession } from '../claude/claude-structured-launch-resolution'
import type { AgentSessionAttachParams } from '../native-chat/agent-session-wire/structured-agent-session-attach'
import { hostTestAttachParams } from '../native-chat/agent-session-wire/structured-agent-session-host-test-data'
import type { StructuredAgentSessionHost } from '../native-chat/agent-session-wire/structured-agent-session-host'
import {
ensureStructuredAgentSessionHost,
stopStructuredAgentSessionRuntime
} from './structured-agent-session-runtime'
import { createStructuredAgentSessionLogger } from '../native-chat/agent-session-wire/structured-agent-session-logger'
/** The error a real CLI's exit reaches the adapter as: Orca's message, the stderr as a log detail. */
export function scriptedClaudeExitError(diagnostic: string): Error {
return withProviderDiagnostic(new Error(diagnostic), providerDiagnostic(diagnostic, 'log'))
}
export type ScriptedClaudeBehavior = {
/** Initialize never answers; only the child's exit settles it. */
initHangs?: boolean
/** Every control read after startup's own settings read waits for `releaseStalls`. */
stallsControlReads?: boolean
/** The spawn itself waits for `releaseStalls`, holding its acquisition open. */
spawnHangs?: boolean
/** The CLI exits with this diagnostic before its spawn returns, or while its start time is read. */
exitsDuringSpawn?: { diagnostic: string; at: 'spawn' | 'start-time-read' }
/** Closing cannot prove the descendant tree gone, as when it was never snapshottable. */
closeUnproven?: boolean
/** Control requests run under this deadline, the way production's run under the default one. */
controlTimeoutMs?: number
/** Every option write (set_model, set_permission_mode, apply_flag_settings) goes unanswered. */
optionWritesHang?: boolean
/** Startup's own settings read goes unanswered. */
startupSettingsReadHangs?: boolean
/** Every option write loses its answer while the CLI keeps running (not a refusal), so a
* start that restores one faults. */
optionWritesFail?: boolean
/** The init frame names another provider session than the one launched. */
initNamesForeignSession?: boolean
}
export type ScriptedClaudeChild = {
sessionId: string
launch: ClaudeStreamJsonLaunch
calls: string[]
handlers: ClaudeStreamJsonConnectionHandlers
connection: Omit<ClaudeStreamJsonConnection, 'closed' | 'exitVerdict'> & {
closed: boolean
exitVerdict: ClaudeStreamJsonConnection['exitVerdict']
}
/** The CLI exits on its own: its root is gone, its tree unverifiable. */
exit: (error: Error) => void
/** The CLI answers initialize now; only meaningful under `initHangs`. */
answerInit: () => void
}
export function createScriptedClaudeRuntime(sessionIds: readonly string[]) {
const children: ScriptedClaudeChild[] = []
const behaviors = new Map<string, ScriptedClaudeBehavior>()
let releaseStalls = (): void => {}
const stall = new Promise<void>((resolve) => {
releaseStalls = resolve
})
let root: string | null = null
let operations = 0
const openConnection: typeof openClaudeStreamJsonConnection = async (launch, handlers = {}) => {
const providerSessionId = String(launch.options.sessionId ?? launch.options.resume)
const sessionId = sessionIds.find(
(candidate) => claudeSessionIdForOrcaSession(candidate) === providerSessionId
)
if (!sessionId) {
throw new Error(`no scripted Claude session for ${providerSessionId}`)
}
const behavior = behaviors.get(sessionId) ?? {}
if (behavior.spawnHangs) {
await stall
}
let failInit = (_error: Error): void => {}
let answerInit = (): void => {}
const answer = <T>(value: T, startup: boolean): Promise<T> =>
behavior.stallsControlReads && !startup ? stall.then(() => value) : Promise.resolve(value)
// Untimed unless the behavior sets a deadline; a hang then settles only on the child's exit.
const control = <T>(subtype: string, run: () => Promise<T>): Promise<T> =>
runClaudeControl(subtype, run, behavior.controlTimeoutMs ?? null)
const never = <T>(): Promise<T> => new Promise<T>(() => {})
const optionWrite = (subtype: string): Promise<void> => {
child.calls.push(subtype)
return control(subtype, () =>
behavior.optionWritesFail
? Promise.reject(new Error('Query closed before response received'))
: behavior.optionWritesHang
? never()
: Promise.resolve()
)
}
let settingsReads = 0
const child: ScriptedClaudeChild = {
sessionId,
launch,
calls: [],
handlers,
exit: (error) => {
child.connection.exitVerdict = { root: 'exited', tree: 'unverifiable' }
failInit(error)
handlers.onExit?.(error)
},
answerInit: () => answerInit(),
connection: {
pid: 5000 + children.length,
closed: false,
exitVerdict: { root: 'live', tree: 'unverifiable' },
initializationResult: () => {
const initialized = { models: [{ value: 'sonnet', displayName: 'Sonnet' }] }
const announce = (): void =>
handlers.onMessage?.({
type: 'system',
subtype: 'init',
session_id: behavior.initNamesForeignSession ? 'foreign-session' : providerSessionId,
model: 'claude-sonnet-5',
apiKeySource: 'none'
})
if (behavior.initHangs) {
return new Promise((resolve, reject) => {
failInit = reject
answerInit = () => {
announce()
resolve(initialized)
}
})
}
announce()
return Promise.resolve(initialized)
},
getContextUsage: () => {
child.calls.push('get_context_usage')
return Promise.resolve({})
},
getSettings: () => {
child.calls.push('get_settings')
settingsReads += 1
const startup = settingsReads === 1
return control('get_settings', () =>
behavior.startupSettingsReadHangs && startup
? never()
: answer({ effective: { effortLevel: 'high' } }, startup)
)
},
supportedModels: () => {
child.calls.push('list_models')
return control('list_models', () =>
answer(
[
{ value: 'sonnet', displayName: 'Sonnet' },
{ value: 'opus', displayName: 'Opus' }
],
false
)
)
},
setModel: () => optionWrite('set_model'),
setPermissionMode: () => optionWrite('set_permission_mode'),
applyFlagSettings: () => optionWrite('apply_flag_settings'),
interrupt: async () => undefined,
cancelAsyncMessage: async () => false,
stopTask: async () => {},
send: async () => {
child.calls.push('send')
},
close: async () => {
child.connection.closed = true
return !behavior.closeUnproven
}
}
}
children.push(child)
if (behavior.exitsDuringSpawn?.at === 'spawn') {
exitAtSpawn(child, behavior.exitsDuringSpawn.diagnostic)
}
return child.connection
}
const exitAtSpawn = (child: ScriptedClaudeChild, diagnostic: string): void => {
child.connection.closed = true
child.exit(scriptedClaudeExitError(diagnostic))
}
/** A pid whose process is gone has no start time to read. */
const readProcessStartTime = async (pid: number): Promise<number | null> => {
const child = children.find((entry) => entry.connection.pid === pid)
const exits = child ? behaviors.get(child.sessionId)?.exitsDuringSpawn : undefined
if (child && exits?.at === 'start-time-read' && !child.connection.closed) {
exitAtSpawn(child, exits.diagnostic)
return pid * 10
}
return child?.connection.exitVerdict.root === 'exited' ? null : pid * 10
}
return {
behave: (sessionId: string, behavior: ScriptedClaudeBehavior): void => {
behaviors.set(sessionId, behavior)
},
/** The latest child spawned for `sessionId`. */
child: (sessionId: string): ScriptedClaudeChild => {
const found = children.findLast((entry) => entry.sessionId === sessionId)
if (!found) {
throw new Error(`no Claude child for ${sessionId}`)
}
return found
},
children: (sessionId: string): ScriptedClaudeChild[] =>
children.filter((entry) => entry.sessionId === sessionId),
install: async (): Promise<StructuredAgentSessionHost> => {
root = await mkdtemp(join(tmpdir(), 'orca-scripted-claude-runtime-'))
await mkdir(join(root, 'claude-home'), { recursive: true })
const directory = root
return ensureStructuredAgentSessionHost({
logger: createStructuredAgentSessionLogger(),
stateDirectory: directory,
hostId: 'local',
claimKeyId: 'key-1',
resolveWorkspacePath: async () => directory,
resolveClaudeCommand: () => '/usr/local/bin/claude',
resolveClaudeAuthPolicy: () => ({ stripAuthEnv: false }),
openClaudeConnection: openConnection,
readProcessStartTime
})
},
attachParams: (
sessionId: string,
expectedRuntimeFence: number | null,
overrides: Partial<AgentSessionAttachParams> = {}
) =>
hostTestAttachParams(expectedRuntimeFence, {
envelope: {
sessionId,
// The runtime's own clock admits operation ids, so these are minted against it.
clientOperationId: `${Date.now()}-${(++operations).toString(16).padStart(32, '0')}`,
expectedRuntimeFence,
payloadFingerprint: ''
},
provider: 'claude',
agent: 'claude',
accountHome: { variable: 'CLAUDE_CONFIG_DIR', path: join(root ?? '', 'claude-home') },
providerHandle: {
kind: 'claude',
sessionId: claudeSessionIdForOrcaSession(sessionId),
leafUuid: null
},
...overrides
}),
dispose: async (): Promise<void> => {
releaseStalls()
await stopStructuredAgentSessionRuntime()
if (root) {
await rm(root, { recursive: true, force: true })
root = null
}
}
}
}