mirror of
https://github.com/stablyai/orca.git
synced 2026-10-08 00:02:38 +00:00
* fix(native-chat): a provider's own retry progress is quoted in its retry row A retry row whose fact carries a detail the provider wrote for a person now quotes it, the same way a rejected message or failed compaction does, so the row says how the retry is going. A log detail still stays out of the sentence. * fix(native-chat): a Codex stream retry is one warning row that updates in place An error Codex says it will retry used to fall through to the generic frame row: red, and a new row for every attempt. It now writes one providerRetrying row per retry run, warning-toned, revised by each attempt with Codex's own progress sentence. A run is the retry frames of one turn with nothing else the thread journals between them; every attempt still publishes, so the idle sweep keeps seeing activity. Errors Codex will not retry are unchanged. * fix(native-chat): a Codex retry row says it is retrying and keeps the frame behind Details The quoted retry sentence leads with "is retrying", which holds for any provider's progress text. The Codex retry row also keeps the whole bounded frame behind the row's Details, as the generic row did, so Codex's additionalDetails stays available to diagnose a retry. * test(native-chat): a Codex retry re-handled after backpressure keeps its one row Pins the run being opened before the write: a first attempt whose publish is refused and is handed back must revise the row it already wrote, not open a second run. Also stops the fixture claiming Codex sends an idle thread status beside each retry, which the app server does not do. * perf(native-chat): a Codex frame with no retry run open is not classified Ending a retry run classified every non-retry frame, and classifying walks the whole payload: every streaming delta, and every large item/completed, paid a walk about as costly as parsing the frame. Only a thread with a run open needs the answer, so the classification now runs only there. * fix(native-chat): each Codex retry attempt is its own warning row, with what failed on its second line A stream error Codex says it will retry is written as its own warning row with a providerRetrying fact, under the same per-frame identity every Codex frame row gets. The host no longer tracks retry runs or rewrites one row in place, so there is no run state to open, end or clear, and no frame has to be classified to end a run. Every attempt publishes, which keeps renewing the idle clock while Codex retries. Codex's additionalDetails, which its own UI shows under the progress message, is kept on the fact as the retry's cause and printed on the row's second line. Errors Codex will not retry are unchanged. * fix(native-chat): a transcript draws only the latest row of a provider retry run The shared structured message projection, which both the desktop and the mobile transcript read, collapses a run of retry rows into its latest row. A run is retry rows from the same agent with no other drawn row between them; a row that draws nothing, or a queued send drawn after the conversation, does not split it. The earlier attempts stay in the journal. * test(native-chat): the retry-run render test uses the message list's current props * fix(native-chat): a Codex frame row is named for its connection, so a later one never revises it Frame rows were named provider-frame:codex:<n> from a counter that starts over with every connection, so the first frame row after a reconnect in the same session revised an earlier connection's row in place, at its old spot. Each connection's frame rows now carry the acquisition generation, minted once before the translator is built: provider-frame:codex:<generation>:<n>. Rows already written keep their identities. * fix(native-chat): agents retrying at once each keep one row, read from the row's own agent * test(native-chat): a reconnect's Codex rows are named for the acquisition that received them * fix(native-chat): a retry run is one agent's, so another agent's row never splits it Each agent's rows are drawn apart: the session's own rows are the conversation, and a subagent's rows open in that subagent's section. Splitting a run on any other row in the flat list left two adjacent retry rows on screen whenever another agent wrote between two attempts: a subagent finishing a command while the session reconnected, or the session working while a subagent reconnected. A run is now per agent: an agent's retry rows with none of its own other rows between them, drawn as its latest. * test(native-chat): a subagent's retry run is checked in its own section, and the run rule's words say same agent * test(mobile): the retry-rows test typechecks, so the test ratchet keeps checking it
152 lines
5.9 KiB
TypeScript
152 lines
5.9 KiB
TypeScript
import { describe, expect, it } from 'vitest'
|
|
import {
|
|
MAX_PROVIDER_DIAGNOSTIC_CHARS,
|
|
providerDiagnostic,
|
|
providerDiagnosticOf,
|
|
agentSessionFailureFact,
|
|
readAgentSessionFailureFact,
|
|
readProviderRetry,
|
|
readWholeAgentSessionFailureFact,
|
|
withProviderDiagnostic
|
|
} from './agent-session-failure'
|
|
import { AGENT_SESSION_REFUSAL_REASONS } from './agent-session-refusal-details'
|
|
|
|
describe('provider diagnostics', () => {
|
|
it('bounds the text and records nothing for an empty one', () => {
|
|
expect(providerDiagnostic(' ', 'log')).toBeUndefined()
|
|
expect(
|
|
providerDiagnostic('x'.repeat(MAX_PROVIDER_DIAGNOSTIC_CHARS + 50), 'log')?.text
|
|
).toHaveLength(MAX_PROVIDER_DIAGNOSTIC_CHARS)
|
|
})
|
|
|
|
it('is read from the error that carries it, or from what it wraps', () => {
|
|
const diagnostic = { text: 'bad request', audience: 'person' as const }
|
|
const carried = withProviderDiagnostic(
|
|
new Error('codex turn/start failed: bad request'),
|
|
diagnostic
|
|
)
|
|
expect(providerDiagnosticOf(carried)).toEqual(diagnostic)
|
|
expect(providerDiagnosticOf(new Error('wrapped', { cause: carried }))).toEqual(diagnostic)
|
|
expect(providerDiagnosticOf(new AggregateError([new Error('other'), carried], 'both'))).toEqual(
|
|
diagnostic
|
|
)
|
|
})
|
|
|
|
it('ends on an aggregate error that contains itself', () => {
|
|
const loop = new AggregateError([], 'loop')
|
|
loop.errors.push(loop)
|
|
expect(providerDiagnosticOf(loop)).toBeUndefined()
|
|
})
|
|
|
|
it('is never inferred from an error that did not carry one', () => {
|
|
// Orca's own wording, even when it quotes something that looks like a provider message.
|
|
expect(
|
|
providerDiagnosticOf(new Error('claude stream-json exited (code 1): not signed in'))
|
|
).toBe(undefined)
|
|
expect(providerDiagnosticOf('a string')).toBeUndefined()
|
|
})
|
|
})
|
|
|
|
describe('reading a failure fact', () => {
|
|
it('keeps what this build can place', () => {
|
|
expect(
|
|
readAgentSessionFailureFact({
|
|
kind: 'restartFailed',
|
|
refusal: { code: 'agent_session_conflict', details: { reason: 'claimConflicted' } },
|
|
detail: { text: 'x', audience: 'person' }
|
|
})
|
|
).toEqual({
|
|
kind: 'restartFailed',
|
|
refusal: { code: 'agent_session_conflict', details: { reason: 'claimConflicted' } },
|
|
detail: { text: 'x', audience: 'person' }
|
|
})
|
|
})
|
|
|
|
it('keeps what a provider said it is retrying, and only that', () => {
|
|
expect(
|
|
readAgentSessionFailureFact({
|
|
kind: 'providerRetrying',
|
|
retry: { error: 'rate_limit', status: 429, attempt: 3 }
|
|
})
|
|
).toEqual({ kind: 'providerRetrying', retry: { error: 'rate_limit', status: 429 } })
|
|
expect(
|
|
readAgentSessionFailureFact({ kind: 'providerRetrying', retry: { error: '', status: 'x' } })
|
|
).toEqual({ kind: 'providerRetrying' })
|
|
// The provider's own account of what failed survives a read, bounded like any detail.
|
|
expect(
|
|
readAgentSessionFailureFact({
|
|
kind: 'providerRetrying',
|
|
retry: { cause: ` stream disconnected${' x'.repeat(400)}` }
|
|
})?.retry?.cause
|
|
).toBe(`stream disconnected${' x'.repeat(400)}`.slice(0, 512).trim())
|
|
expect(
|
|
readAgentSessionFailureFact({ kind: 'providerRetrying', retry: { cause: ' ' } })
|
|
).toEqual({ kind: 'providerRetrying' })
|
|
})
|
|
|
|
it('reads a row an unreleased build wrote with a cause as a refusal with no details', () => {
|
|
expect(
|
|
readAgentSessionFailureFact({
|
|
kind: 'restartFailed',
|
|
refusal: { code: 'agent_session_conflict', cause: 'claimConflicted' }
|
|
})
|
|
).toEqual({ kind: 'restartFailed', refusal: { code: 'agent_session_conflict' } })
|
|
})
|
|
|
|
it('drops what a newer host wrote that this build cannot place', () => {
|
|
expect(readAgentSessionFailureFact({ kind: 'futureKind' })).toBeUndefined()
|
|
expect(readAgentSessionFailureFact(undefined)).toBeUndefined()
|
|
expect(
|
|
readAgentSessionFailureFact({
|
|
kind: 'restartFailed',
|
|
refusal: { code: 'agent_session_conflict', details: { reason: 'futureReason' } },
|
|
detail: { text: 'x', audience: 'future' }
|
|
})
|
|
).toEqual({ kind: 'restartFailed', refusal: { code: 'agent_session_conflict' } })
|
|
})
|
|
})
|
|
|
|
describe('reading all of a failure fact', () => {
|
|
it('reads every part of a fact the host built, as it arrives off the wire', () => {
|
|
for (const fact of [
|
|
agentSessionFailureFact('startFailed', {
|
|
refusal: { code: 'agent_session_conflict', details: { reason: 'claimConflicted' } }
|
|
}),
|
|
agentSessionFailureFact('providerRejected', {
|
|
detail: providerDiagnostic('Image type .bmp', 'person')
|
|
}),
|
|
agentSessionFailureFact('attachmentInvalid', {
|
|
attachment: { reason: 'tooLarge', limit: 5 * 1024 * 1024 }
|
|
}),
|
|
agentSessionFailureFact('providerRetrying', {
|
|
retry: readProviderRetry({ error: 'rate_limit', status: 429 })
|
|
})
|
|
]) {
|
|
expect(readWholeAgentSessionFailureFact(JSON.parse(JSON.stringify(fact)))).toEqual(fact)
|
|
}
|
|
// Every reason a refusal can name, on the code that names it.
|
|
for (const [code, reasons] of Object.entries(AGENT_SESSION_REFUSAL_REASONS)) {
|
|
for (const reason of reasons) {
|
|
const fact = { kind: 'startFailed', refusal: { code, details: { reason } } }
|
|
expect(readWholeAgentSessionFailureFact(fact)).toEqual(fact)
|
|
}
|
|
}
|
|
})
|
|
|
|
it('reads nothing when this build would drop any part, however deep', () => {
|
|
for (const value of [
|
|
{ kind: 'futureKind' },
|
|
{ kind: 'startFailed', refusal: { code: 'agent_session_future_code' } },
|
|
{
|
|
kind: 'startFailed',
|
|
refusal: { code: 'agent_session_conflict', details: { reason: 'futureReason' } }
|
|
},
|
|
{ kind: 'attachmentInvalid', attachment: { reason: 'futureReason' } },
|
|
{ kind: 'providerRejected', detail: { text: 'x', audience: 'future' } },
|
|
{ kind: 'startFailed', futurePart: {} }
|
|
]) {
|
|
expect(readWholeAgentSessionFailureFact(value)).toBeUndefined()
|
|
}
|
|
})
|
|
})
|