mirror of
https://github.com/stablyai/orca.git
synced 2026-10-08 16:02:37 +00:00
* fix(native-chat): a child's root exit is reported even during its close, and bookkeeping after it never reads as unproven - Both connections report the root process's exit once, with `expected` set when a close had begun. A close that came back unproven and whose root exits later is finished by the adapter, and its end reaches the host like any other. - A Claude close whose resume-point write fails after the exit was proven, and a Codex close whose terminal row is refused, now end the session and report the failure, instead of keeping a dead child indexed as if its exit were unproven. - A Codex close whose forced tree kill can't prove the descendants gone but saw the root exit reports the descendants and counts the root exit. - Every child exit with an identity, expected or not, is forwarded to the host. * fix(native-chat): the exit ends the child's record; an unfinished stop is the child's own close, which everyone joins - The host keeps no stored "stop still owed" record any more. A stop begins the child's close (`child.close`), which lives on the child and ends with it. A second Stop, the idle reaper, quit, a send and an option/answer/goal/rewind all join that close instead of retrying a separate obligation. - A caller waits on the close only as long as the step deadline; the close itself is never abandoned. A proof that lands after every caller stopped waiting reaches the host as the adapter's report of that exit, which ends the record through the same handler. - Once the exit is proven, draining, settling, the lease release and the adapter's acknowledgement are each attempted and reported on failure; none keeps the child on record. A start, and the handle's close, write a release that failed from this host's proof of that exit, so a failed write never refuses a send. - A start that meets a close still unverifiable is refused with `previousExitUnverifiable`, so the queued message is rejected with a send-again reason; nothing is held and nothing starts beside the old process. - The idle sweep goes back to idle reaping only. - Removes #24333's retry entry points, the wait row and its hold rule, the ask/failure cursors on the stored record, and the stop's own wake. Tests replace the #24333 unproven-stop test: a send joining an unproven close and an in-flight one, a late proof past the caller's bound, a root exiting after its close gave up, a proven exit whose resume-point write and lease release both failed, an unverifiable close rejecting the send and refusing an option change, a surviving descendant, quit and the idle reaper; and Codex's unverifiable, late-exit and joined-close cases. * fix(native-chat): a message refused because the old process's exit is unverifiable says so, and to send again The start failure for a refusal with reason `previousExitUnverifiable` reads "Orca couldn't confirm Claude's previous process ended. Send your message to try again." instead of "Claude couldn't restart." The status-row kind and the refusal reason stay in the shared lists for rows and hosts that still carry them; the catalogs keep one sentence for both. * fix(native-chat): a close's verdict is the root's exit alone, and what follows it is logged - A Claude close resolves as soon as the root's exit is proven: the session ends and its `ended` report goes out then. Saving the resume point runs afterwards and a failure is logged, so a slow or hung write never reads as an unproven exit or keeps a dead child on record. - A root that exits after its close came back unproven finishes that close through the same path as any close, so the session's child work is published as ended (background tasks and subagents no longer stay shown running for a dead agent), and a failure there is logged. - Codex logs a refused final row, and reports a root exit whose forced tree kill could not prove the rest of the tree gone the way Claude does, so the host logs it and blocks nothing. - Both adapters take the host's logger for this bookkeeping. * fix(native-chat): one handler ends every child's exit, and a join waits on the adapter's own close - One exit handler (`structured-agent-session-child-exit`) ends a child's record for an exit expected or not. `expected` only changes what the chat is told: the stop's cause, its end at the stop's ask, the settlement id, and no crash outcome row. The lease release keeps the exit's evidence; the handoff guard, lifecycle barrier, sink release and adapter acknowledgement apply to both. A Claude journal-sink failure ends in the same step as its stop, as Orca's own fault. - Joining a close is asking the adapter, whose close is memoized while it runs and bounded by its own kill escalation; the host keeps no attempt of its own and no 10 s caller bound. An ask after a close came back unproven runs the stop again. - A close's end is stamped where its stop was asked for (a repeated ask moves it), so the closed chat and failed start checks order a message accepted meanwhile after it. - A start refused because the old exit is unverifiable rejects what was queued in the same step. - The end of a close the host asked for no longer waits on the cross-session recovery chain. - The kill no longer waits for the stop event's write; the journal writes rows in order. * fix(native-chat): an exit's lease release lands whatever the length of its reason A crash's reason can carry kilobytes of the provider's stderr, and a lease whose death detail is over 512 characters fails the store's own check. The exit handler cut it, but the release a start or the chat handle's close re-derives did not, so after a crash whose own release failed every message was refused as not resumable until restart. The record's builder now cuts the detail to the record's bound, so no writer can hand it one too long. * fix(claude): a proven close waits at most 2 s for the output it already wrote Once the root's exit is proven, the close still waited for the SDK's output reader to end. Something outside the process tree that holds the output open would keep that close, and every send, Stop and quit joining it, waiting with no bound. The wait is now bounded; past it the close resolves as proven and the open output is logged. * fix(codex): an exit reported inside Orca's close keeps the reason Orca closed it for The connection reports the app-server's exit inside the close that ends it, so that report ended every Codex close and replaced the close's own reason (for example, a provider frame that could not be recorded) with the connection's stderr text in the ended record and the lease's exit evidence. The session now records Orca's close with its reason, and the exit it ends keeps that reason. The test connection reports its exit inside close the way the real one does. * fix(native-chat): quit stops delivery before it drains exit recovery Every exit now wakes delivery, and teardown drained exit recovery before it stopped delivery, so an exit settled in that window could start a fresh agent that teardown then killed. Teardown stops delivery first; queued messages wait for the next launch. * docs(native-chat): the unverifiable-exit refusal no longer names a caller's wait The caller's bounded wait was removed; the comment describes the close as it is now. * fix(native-chat): a stop whose kill did not take is logged, and the next ask kills again When a close's kill leaves the agent's root running, the host now logs it. Tests pin what a later ask does: each connection runs its whole stop again (Codex sends SIGKILL a second time), refuses input meanwhile, and proves the exit once the kill takes. * fix(native-chat): a start refused over the old process says Orca couldn't stop it The host reaches an unverifiable verdict only after its own kill left the agent's root running, on the machine that runs the agent, so the sentence now says that: "Orca couldn't stop {agent}'s previous process." The refusal reason, failure kind and wire shapes are unchanged. The host test also checks the failed kill is logged. * docs(native-chat): an unverifiable close verdict is a root that survived the kill The host's close runs where the agent runs, so lost contact never yields this verdict; the comment no longer says it does. * fix(native-chat): a kill that did not take is reported once, by whoever met it The log added at the close fired beside a Stop's own failure report for the same event. A stop still reports it through its failure; a send or option change refused over it now logs it at the refusal, the only place it is otherwise invisible. * test(native-chat): a second Stop joins a close the first could not prove and retries its kill * fix(native-chat): say a start refused beside an unstopped process plainly The rejection now reads "Couldn't stop {{agent}} from before. Send your message again to try once more." This kind has its own send-again step; every other failure keeps "Send your message to try again."
285 lines
10 KiB
TypeScript
285 lines
10 KiB
TypeScript
import type { StructuredAgentSessionLogger } from '../native-chat/agent-session-wire/structured-agent-session-logger'
|
|
import type {
|
|
AgentJournalItemIdentity,
|
|
AgentSessionJournalIdentity
|
|
} from '../../shared/agent-session-journal-types'
|
|
import { randomUUID } from 'node:crypto'
|
|
import type { AgentJournalDispatchRejection } from '../../shared/agent-session-failure-words'
|
|
import { cancelProcessAcquisition } from '../../shared/child-process/cancel-process-acquisition'
|
|
import type {
|
|
CodexAppServerConnection,
|
|
openCodexAppServerConnection
|
|
} from './codex-app-server-connection'
|
|
import { CodexAcquisitionWindow } from './codex-structured-acquisition-window'
|
|
import {
|
|
createCodexTurnOpenWaits,
|
|
type CodexTurnOpenWaits
|
|
} from './codex-structured-turn-open-wait'
|
|
import type { CodexDispatchEchoes } from './codex-structured-dispatch-echo'
|
|
import type { AgentChildWorkEvidence } from '../../shared/agent-status-child-work-evidence'
|
|
import type { CodexBackgroundTaskTracker } from './codex-background-task-tracker'
|
|
import type { CodexJournalTranslator } from './codex-structured-journal-translation'
|
|
import type { StructuredAgentSessionEndedEvent } from '../native-chat/agent-session-wire/structured-agent-session-adapter'
|
|
import type { CodexStructuredPermissionPolicy } from './codex-structured-permission-policy'
|
|
import type {
|
|
AgentModelCatalogSessionAccess,
|
|
AgentModelCatalogStore
|
|
} from '../native-chat/agent-model-catalog/agent-model-catalog-store'
|
|
|
|
export type CodexSessionCatalogAccess = AgentModelCatalogSessionAccess
|
|
|
|
export type CodexStructuredLaunch = {
|
|
command: string
|
|
args: string[]
|
|
cwd: string
|
|
codexHome: string | null
|
|
resumeThreadId: string | null
|
|
resumePath?: string | null
|
|
/** The resumed thread is this session's own creation: when Codex answers that it holds no
|
|
* rollout for it, start a new thread in its place. Never set for a thread a resume proved. */
|
|
supersedeIfUnsaved?: boolean
|
|
permissionPolicy?: CodexStructuredPermissionPolicy
|
|
/** The model the session chose; the thread opens on it so its first turn is not a switch. */
|
|
model?: string
|
|
env?: Record<string, string>
|
|
}
|
|
|
|
export type CodexStructuredSessionEvent =
|
|
| {
|
|
type: 'notification'
|
|
sessionId: string
|
|
threadId: string
|
|
method: string
|
|
params: unknown
|
|
/** Host receipt time of a turn boundary; survives retry and deferral so a replay is not re-stamped. */
|
|
observedAt?: number
|
|
/** Highest dispatch sequence armed when this turn-start was first received. */
|
|
dispatchSequenceAtReceipt?: number
|
|
}
|
|
| { type: 'server-request'; sessionId: string; threadId: string; method: string; params: unknown }
|
|
| { type: 'provider-frame'; sessionId: string; threadId: string; kind: string; payload: unknown }
|
|
| {
|
|
type: 'prompt'
|
|
sessionId: string
|
|
threadId: string
|
|
method: string
|
|
params: unknown
|
|
codexItemId: string
|
|
promptKey: string
|
|
}
|
|
| StructuredAgentSessionEndedEvent
|
|
/** Translator-only compatibility for callers that do not participate in host recovery. */
|
|
| { type: 'ended'; sessionId: string; reason: string; observedAt?: number }
|
|
|
|
export type CodexStructuredSessionAdapterDeps = {
|
|
resolveLaunch: (input: {
|
|
identity: AgentSessionJournalIdentity
|
|
}) => Promise<CodexStructuredLaunch>
|
|
/** Host capability seam; production uses the native Windows process table. */
|
|
isWindowsProcessStartTimeAvailable?: () => boolean
|
|
onEvent?: (event: CodexStructuredSessionEvent) => void
|
|
/** Where bookkeeping a close or exit does after the child is gone reports a failure. */
|
|
logger?: StructuredAgentSessionLogger
|
|
/** What the session's child work did, delivered after the journal handled the frame. */
|
|
onChildWorkEvidence?: (sessionId: string, evidence: AgentChildWorkEvidence[]) => void
|
|
/** A send admitted earlier: its identity once Codex echoes it, or its rejection when the turn
|
|
* Codex answered it into ended without taking it. */
|
|
onDispatchSettledLate?: (
|
|
input: { sessionId: string; clientMessageId: string } & (
|
|
| { providerIdentity: AgentJournalItemIdentity }
|
|
| ({ state: 'rejected' } & AgentJournalDispatchRejection)
|
|
)
|
|
) => void
|
|
/** Codex reported its thread not running with no turn open: a send whose
|
|
* dispatch was never answered is owed nothing after this. */
|
|
onPrimaryThreadStoppedRunning?: (input: { sessionId: string }) => void
|
|
openConnection?: typeof openCodexAppServerConnection
|
|
readProcessStartTime?: (pid: number) => Promise<number | null>
|
|
mintLinkId?: () => string
|
|
mintAcquisitionGeneration?: () => string
|
|
now?: () => number
|
|
requestTimeoutMs?: number
|
|
/** Host model catalog; sessions write their listings through and read back. */
|
|
modelCatalog?: AgentModelCatalogStore
|
|
}
|
|
|
|
export type CodexSession = {
|
|
connection: CodexAppServerConnection
|
|
ended: boolean
|
|
/** First observed child exit survives rejected settlement admission. */
|
|
exitObservedAt?: number
|
|
/** The close Orca began for this child: asked for, or forced as a death, and why. Whatever ends
|
|
* the child after it (that close, or the exit the connection reports meanwhile) keeps this. */
|
|
orcaClose?: { requested: boolean; reason: Error }
|
|
fence: number
|
|
acquisitionGeneration: string
|
|
threadId: string
|
|
historyMode?: 'legacy' | 'paginated'
|
|
/** Primary-thread turns Codex reported started and not yet ended, as read off the wire: what
|
|
* rewind waits out and what a Stop naming no turn interrupts when the journal shows none. */
|
|
activeTurnIds?: Set<string>
|
|
/** Stops waiting for the turn Codex answered a send into to open. */
|
|
turnOpenWaits: CodexTurnOpenWaits
|
|
dispatchPending?: boolean
|
|
prompts: CodexAcquisitionWindow['prompts']
|
|
options: Map<string, string>
|
|
reportedOptions: {
|
|
model?: string
|
|
effort?: string
|
|
serviceTier?: string | null
|
|
serviceTierKnown?: true
|
|
}
|
|
/** Exact provider-advertised Fast request value for each discovered model. */
|
|
fastModeTierByModel: Map<string, string>
|
|
/** Absent when the adapter runs without a host catalog store (tests). */
|
|
catalogAccess?: CodexSessionCatalogAccess
|
|
/** Sends whose identity is still to be settled by the provider echo. */
|
|
dispatchEchoes: CodexDispatchEchoes
|
|
translator: CodexJournalTranslator | null
|
|
/** Ephemeral roster behind the background-tasks strip; never durable state. */
|
|
backgroundTasks: CodexBackgroundTaskTracker
|
|
unbindReadingControl?: () => void
|
|
/** Terminates this exact child as an unexpected death and enters host recovery. */
|
|
forceCloseUnexpected?: (reason: Error) => Promise<boolean>
|
|
}
|
|
|
|
export function mintCodexAcquisitionGeneration(deps: CodexStructuredSessionAdapterDeps): string {
|
|
return deps.mintAcquisitionGeneration?.() ?? randomUUID()
|
|
}
|
|
|
|
export function codexSessionLifecycle(
|
|
fence: number,
|
|
acquisitionGeneration: string
|
|
): Pick<CodexSession, 'ended' | 'fence' | 'acquisitionGeneration' | 'turnOpenWaits'> {
|
|
return {
|
|
ended: false,
|
|
fence,
|
|
acquisitionGeneration,
|
|
turnOpenWaits: createCodexTurnOpenWaits()
|
|
}
|
|
}
|
|
|
|
/** A child that exited while being acquired never becomes the session's. */
|
|
export function assertCodexConnectionOpen(
|
|
connection: Pick<CodexAppServerConnection, 'closed'>,
|
|
sessionId: string
|
|
): void {
|
|
if (connection.closed) {
|
|
throw new Error(`codex app-server for session ${sessionId} exited while being acquired`)
|
|
}
|
|
}
|
|
|
|
export function requireLiveCodexSession(
|
|
sessions: Map<string, CodexSession>,
|
|
sessionId: string
|
|
): CodexSession {
|
|
const session = sessions.get(sessionId)
|
|
if (!session || session.ended) {
|
|
throw new Error(`no live codex app-server for session ${sessionId}`)
|
|
}
|
|
return session
|
|
}
|
|
|
|
export type CodexAcquisitionAttempt = {
|
|
window: CodexAcquisitionWindow
|
|
cancelled: boolean
|
|
exitProven: boolean
|
|
finished: Promise<void>
|
|
finish: () => void
|
|
}
|
|
|
|
export function createCodexAcquisitionAttempt(): CodexAcquisitionAttempt {
|
|
let finish = (): void => {}
|
|
const finished = new Promise<void>((resolve) => {
|
|
finish = resolve
|
|
})
|
|
return {
|
|
window: new CodexAcquisitionWindow(),
|
|
cancelled: false,
|
|
exitProven: false,
|
|
finished,
|
|
finish
|
|
}
|
|
}
|
|
|
|
export class CodexAcquisitionRegistry {
|
|
private readonly attempts = new Map<string, CodexAcquisitionAttempt>()
|
|
private closing = false
|
|
|
|
get size(): number {
|
|
return this.attempts.size
|
|
}
|
|
|
|
start(sessionId: string): {
|
|
previousAttempt: CodexAcquisitionAttempt | undefined
|
|
attempt: CodexAcquisitionAttempt
|
|
} {
|
|
if (this.closing) {
|
|
throw new Error('codex structured session adapter is closing')
|
|
}
|
|
const previousAttempt = this.attempts.get(sessionId)
|
|
const attempt = createCodexAcquisitionAttempt()
|
|
this.attempts.set(sessionId, attempt)
|
|
return { previousAttempt, attempt }
|
|
}
|
|
|
|
assertCurrent(sessionId: string, attempt: CodexAcquisitionAttempt): void {
|
|
if (this.closing || attempt.cancelled || this.attempts.get(sessionId) !== attempt) {
|
|
throw new Error(`codex session ${sessionId} was superseded while being acquired`)
|
|
}
|
|
}
|
|
|
|
get(sessionId: string): CodexAcquisitionAttempt | undefined {
|
|
return this.attempts.get(sessionId)
|
|
}
|
|
|
|
deleteIfCurrent(sessionId: string, attempt: CodexAcquisitionAttempt): void {
|
|
if (this.attempts.get(sessionId) === attempt) {
|
|
this.attempts.delete(sessionId)
|
|
}
|
|
}
|
|
|
|
restoreIfCurrent(
|
|
sessionId: string,
|
|
replacement: CodexAcquisitionAttempt,
|
|
previous: CodexAcquisitionAttempt
|
|
): void {
|
|
if (this.attempts.get(sessionId) === replacement) {
|
|
this.attempts.set(sessionId, previous)
|
|
}
|
|
}
|
|
|
|
async closeFailedAttempt(sessionId: string, attempt: CodexAcquisitionAttempt): Promise<boolean> {
|
|
const stopped = (await attempt.window.connection?.close()) ?? true
|
|
if (stopped) {
|
|
attempt.exitProven = true
|
|
this.deleteIfCurrent(sessionId, attempt)
|
|
}
|
|
return stopped
|
|
}
|
|
|
|
sessionIds(): IterableIterator<string> {
|
|
return this.attempts.keys()
|
|
}
|
|
|
|
close(): void {
|
|
this.closing = true
|
|
}
|
|
}
|
|
|
|
export async function cancelCodexAcquisitionAttempt(
|
|
attempt: CodexAcquisitionAttempt | undefined
|
|
): Promise<boolean> {
|
|
if (!attempt) {
|
|
return true
|
|
}
|
|
return cancelProcessAcquisition({
|
|
cancel: () => {
|
|
attempt.cancelled = true
|
|
},
|
|
connection: () => attempt.window.connection,
|
|
exitProven: () => attempt.exitProven,
|
|
finished: attempt.finished
|
|
})
|
|
}
|