Files
orca/src/shared/structured-agent-session-projection.ts
T
Brennan Benson 5e6fcca0b3 fix(native-chat): date a session by its own lifecycle, not its subagents' work (#22520)
* fix(native-chat): date a session by its own agent's rows, not its subagents'

The journal reducer's lastActivityAt is the structured status summary's
updatedAt, which the status row uses as its completion stamp and
acknowledgement clock. It took the max over every journal row, and a
session's subagents write into the same journal after its own agent has
settled, so an idle parent was re-dated and marked unread on child work.

A row now dates the session only when the session's own agent produced it:
not a row whose producer linkage names a subagent, and not a subagent
roster row (a subagent-group block), which the session writes but revises
on every child transition. The roster rule is derived from the row body;
no new persisted field. Replay folds through the same rule, so existing
journals are re-dated to their own last row on reopen.

Claude: a backgrounded subagent emits no child frames, so its re-dating
came entirely from roster revisions (task_updated, task_notification) and
from the stale-roster revision written when a journal reopens. Codex: the
roster is revised on every child token-usage report; child-thread rows
carry no producer linkage yet, and read as the session's own until they do.

* fix(native-chat): a reopened journal's verdict on stale work does not date the session

Reopening a journal settles rows the previous host left live (a working
subagent roster, a live background task) to unverifiable. Those revisions
were appended at the reopen moment, and a background-task row is the
session's own non-roster row, so a crash-restarted session with a live
shell was re-dated to the restart although no agent acted.

The reconciler now writes each verdict revision with the row's own observed
time. That is one rule at the one writer, covering both settle shapes; the
render item's observedAt was already pinned to the row's first write, so
nothing the transcript shows changes. The live-transition roster exclusion
stays: live roster revisions are written by the providers, not here.

The `recovered` row flag is not used as the discriminator: the live
unexpected-exit settlement also writes recovered rows, and a clock rule
keyed on it would stop dating a provider crash the host just observed.

* fix(native-chat): date a session by the reducer's attribution of what a row wrote

The clock read producer linkage off the raw row. A lifecycle batch names no
row-level producer, a tombstone names none, and a revision may name none while
the reducer still attributes the item to a subagent, so each of those dated an
idle parent. The clock now asks the reducer: after a row applies, whether any
item it wrote is the session's own work; before a removal, whether the item it
removes was.

* fix(native-chat): date a session's status by its own lifecycle edges

A subagent writes into its parent's journal and keeps going after the parent
settles. Every one of its rows advanced the summary's updatedAt, and the status
row re-dated a done parent to it, so an idle parent read as newly finished and
unread on each child step.

The host now publishes statusStartedAt beside updatedAt: when the session's own
agent entered its status, read off edges only it writes. Idle is when its newest
turn ended; working is when the running turn was requested, or the earliest
send still unanswered; attention is its own oldest pending ask, or a subagent's
when that alone holds it. A turn that recovery settled after its host went away
ended when that settle was written, so it reads as a completion the user has not
seen; it carries no outcome, so no completion event or notification calls it a
success. Render items carry recoveredAt, the recovered row's own write time, so
nothing new is persisted.

The sidebar bridge and the host ingest date the row and the main agent's clock
from it whenever the row shows the main agent's own state, and keep their
existing rules for a row child work holds open or a summary from an older host.
The status feed republishes when the clock moves instead of on every idle row.

* revert(native-chat): keep the journal clock over every row

The row filter this branch put on the reducer's lastActivityAt decided which
rows could date a session: a list of exclusions that each new row kind could
slip past. The session's state is now dated by its own lifecycle edges, so the
filter, its attribution helper and the backdated reopen verdicts go back to
main. lastActivityAt, and the summary's updatedAt it feeds, is again the
evidence clock over every row, including a subagent's.

* fix(native-chat): keep republishing an idle session its live child work holds open

A row held open by live child work is dated by when each reader saw the
publish, and mobile decays a working row whose evidence is older than the
staleness window. Suppressing row-activity republishes for every dated idle
session froze that evidence, so a subagent running more than 30 minutes past
its parent's turn made the row read idle on mobile. Only a session nothing
holds open stays quiet on row activity now; its state clock is unchanged.

* fix(activity): order an agent's timeline by when each state was seen

An answered ask returns a settled parent to its own turn's end, so its done
repeats the time of the done before the ask. Activity keyed and ordered
events by that time: the new done collided with the old one and was
dropped, and the row took its state from the newest-dated event, the
blocked ask, so a done parent read Blocked and needed attention.

Each state switch now records the `updatedAt` it was seen at. Events are
keyed and ordered by that, while unread and "Clear completed" still compare
the state's own time, so the answer neither re-lights unread nor revives a
cleared done. The row's state comes from the pane's own status entry, so a
clear that hid the done cannot leave it reading Blocked either.

* test(activity): pin the timeline across repeated asks, a clear, and a stale turn

Three parts of ordering the timeline by when each state was seen had no test
that failed without them:

- A second ask moves the answered done into history. Both dones share the
  turn's end, so only the history entry's own seen time keeps them apart;
  without it one done collided with the other and the timeline showed two
  Blocked events in a row. Three asks also exceed the per-pane cap, which must
  keep the most recently seen events, not the most recently started.
- "Clear completed" on an answered row must cut off past the ask, which is
  dated after the done, or the cleared row stays listed. A done that the user
  cleared must also stay hidden once a later ask moves it into history.
- A stale working row must not read as running just because the pane's own
  status says working.
2026-09-24 16:17:08 -07:00

376 lines
14 KiB
TypeScript

import {
AGENT_STATUS_MAX_FIELD_LENGTH,
normalizeOptionalField,
normalizePromptField
} from './agent-status-field-normalization'
import {
AGENT_JOURNAL_MESSAGE_SEND_MODES,
type AgentJournalMessageSendMode,
type AgentJournalRenderItem,
type AgentJournalSubmission,
type AgentJournalTurnOutcome
} from './agent-session-journal-types'
import { isRootAgentJournalItem } from './agent-session-journal-producer'
import { readAgentJournalTurnOutcome } from './agent-session-turn-record'
import {
AGENT_STATUS_TOOL_INPUT_MAX_LENGTH,
AGENT_STATUS_TOOL_NAME_MAX_LENGTH
} from './agent-status-types'
import { describeToolInput } from './native-chat-tool-summary'
import {
activeStructuredAgentSessionTurnId,
newestStructuredAgentSessionTurn,
statusStructuredAgentSessionToolCall
} from './structured-agent-session-live-turn'
import {
isStructuredAgentSessionToolAction,
structuredAgentSessionToolCallBlock
} from './structured-agent-session-tool-call-block'
import type { NativeChatBlock, NativeChatMessage } from './native-chat-types'
import { sha256 } from './sha256'
import { structuredAgentSessionStatusStartedAt } from './structured-agent-session-status-started-at'
import { isUnansweredStructuredAgentSessionDispatch } from './structured-agent-session-unanswered-dispatch'
// Re-exported so the live-turn readers' existing consumers keep one import site.
export {
activeStructuredAgentSessionTurnId,
newestStructuredAgentSessionTurn
} from './structured-agent-session-live-turn'
function boundedText(payload: { head: string; truncated: boolean; byteLength: number }): string {
return payload.truncated ? `${payload.head}\n… (${payload.byteLength} bytes)` : payload.head
}
/** The markers a clipped payload carries in its own text, anchored to the end
* so nothing that merely looks like one inside the body can match. */
const BOUNDED_TEXT_MARKERS = [
/\n… \(\d+ bytes\)$/,
/\n\[Orca: output truncated — \d+ bytes total, digest [0-9a-f]+\]$/
]
/** Recovers the clipped body from a bounded payload's text, and says whether a
* marker was there. A reader that treats the text as content renders the
* marker as a line of it — with a line number, which reads as a real position
* in the file — and reports the body as complete. */
export function stripBoundedTextMarker(text: string): { text: string; truncated: boolean } {
const stripped = BOUNDED_TEXT_MARKERS.reduce((value, marker) => value.replace(marker, ''), text)
return { text: stripped, truncated: stripped.length !== text.length }
}
function itemBlocks(item: AgentJournalRenderItem): {
role: NativeChatMessage['role']
blocks: NativeChatBlock[]
} | null {
const body = item.body
if (body.kind === 'message') {
return { role: body.role, blocks: body.blocks }
}
if (isStructuredAgentSessionToolAction(body)) {
const call = structuredAgentSessionToolCallBlock(body)
if (body.kind === 'diff') {
return {
role: 'assistant',
blocks: [call, { type: 'tool-result', output: boundedText(body.patch) }]
}
}
return {
role: 'assistant',
blocks: [
call,
...(body.output
? [
{
type: 'tool-result' as const,
output: boundedText(body.output),
isError: body.state === 'failed'
}
]
: [])
]
}
}
if (body.kind === 'approval') {
if (body.resolution.state === 'pending') {
return null
}
return {
role: 'system',
blocks: [
{
type: 'text',
text: `${body.title}\n${body.detail ?? ''}\n${body.resolution.state}`.trim()
}
]
}
}
if (body.kind === 'question') {
if (body.resolution.state === 'pending') {
return null
}
const choices = body.options.map((option) => option.label).join(' · ')
return {
role: 'system',
blocks: [{ type: 'text', text: `${body.question}\n${choices}`.trim() }]
}
}
// A turn record is timing, not content; a kind this build does not know is
// never painted as text either, so a newer host can add kinds freely.
if (body.kind !== 'status' || body.turnLifecycle) {
return null
}
return {
role: 'system',
blocks: [
{
type: 'text',
text: body.text,
...(body.presentation !== undefined ? { presentation: body.presentation } : {}),
...(body.tone !== undefined ? { tone: body.tone } : {}),
...(body.providerFrame ? { providerFrame: body.providerFrame } : {})
}
]
}
}
function isAgentJournalMessageSendMode(value: string): value is AgentJournalMessageSendMode {
return AGENT_JOURNAL_MESSAGE_SEND_MODES.some((mode) => mode === value)
}
const projectedItems = new WeakMap<AgentJournalRenderItem, NativeChatMessage | null>()
/** Deliberately NOT scoped by producer: the transcript shows every agent's
* output. The line this module draws is that the transcript renders every item,
* while every "what is this agent doing right now" scan renders only the
* session's own agent's. */
export function projectStructuredItemsToNativeChat(
items: readonly AgentJournalRenderItem[]
): NativeChatMessage[] {
const messages: NativeChatMessage[] = []
items.forEach((item) => {
const projected = projectStructuredItemToNativeChat(item)
if (projected) {
messages.push(projected)
}
})
return messages
}
export function projectStructuredItemToNativeChat(
item: AgentJournalRenderItem
): NativeChatMessage | null {
const cached = projectedItems.get(item)
if (cached !== undefined) {
return cached
}
// Reducer updates replace journal items, so unchanged rows keep their render caches.
const projected = itemBlocks(item)
const sentAs = item.body.kind === 'message' ? item.body.sentAs : undefined
const message: NativeChatMessage | null = projected
? {
id: item.itemId,
role: projected.role,
blocks: projected.blocks,
timestamp: item.observedAt,
source: 'transcript',
// A send mode this build cannot name renders as an ordinary message.
...(sentAs !== undefined && isAgentJournalMessageSendMode(sentAs) ? { sentAs } : {})
}
: null
projectedItems.set(item, message)
return message
}
/** Deliberately NOT scoped by producer: this is an existence test ("is this
* session listable at all"), not an attribution one. A session whose only
* content came from a subagent still has content. */
export function hasPersistedStructuredAgentSessionTurn(
items: readonly AgentJournalRenderItem[]
): boolean {
return items.some(
(item) =>
item.body.kind === 'message' && (item.body.role === 'user' || item.body.role === 'assistant')
)
}
/**
* A send the host has journaled that the provider has neither opened a turn for nor refused.
*
* Codex declares `turn/started` within ~150ms, but Claude's running row can only be written once
* the SDK echoes the user message back — a 3.4s median and 18s at p90 on real journals. Waiting
* on that echo to call a session working leaves the whole gap reading idle in the chat and in
* every session list, so the send itself is the evidence.
*
* A live `unknown` still counts because an ambiguous adapter reply does not prove the provider
* stopped. A recovered `unknown` does not — it outlived the host generation that sent it, so
* there is nothing still running to report.
*/
export function hasUnansweredStructuredAgentSessionDispatch(
submissions: readonly AgentJournalSubmission[],
currentFence?: number | null
): boolean {
return submissions.some((submission) =>
isUnansweredStructuredAgentSessionDispatch(submission, currentFence)
)
}
export type StructuredAgentSessionProjectedStatus = 'working' | 'attention' | 'idle'
export function structuredAgentSessionTabId(sessionId: string): string {
return `structured-agent-session-${sessionId}`
}
export function projectStructuredAgentSessionStatus(
items: readonly AgentJournalRenderItem[],
submissions: readonly AgentJournalSubmission[] = [],
currentFence?: number | null
): StructuredAgentSessionProjectedStatus {
if (
items.some(
(item) =>
(item.body.kind === 'approval' || item.body.kind === 'question') &&
item.body.resolution.state === 'pending'
)
) {
return 'attention'
}
return activeStructuredAgentSessionTurnId(items) ||
hasUnansweredStructuredAgentSessionDispatch(submissions, currentFence)
? 'working'
: 'idle'
}
function messageProse(blocks: readonly NativeChatBlock[]): string {
return blocks.flatMap((block) => (block.type === 'text' ? [block.text] : [])).join('\n')
}
/** The newest prompt the session's own user turn carries, as the sidebar quotes
* it. Scoped to root rows for the same reason the assistant line is: a provider
* that journals a subagent's own prompt would otherwise requote it as the
* session's. */
export function latestStructuredAgentSessionPrompt(
items: readonly AgentJournalRenderItem[]
): string {
const body = latestStructuredAgentSessionUserItem(items)?.body
return body?.kind === 'message' ? messageProse(body.blocks) : ''
}
export function latestStructuredAgentSessionUserItem(
items: readonly AgentJournalRenderItem[]
): AgentJournalRenderItem | null {
for (let index = items.length - 1; index >= 0; index -= 1) {
const item = items[index]
if (
item?.body.kind === 'message' &&
item.body.role === 'user' &&
isRootAgentJournalItem(item)
) {
return item
}
}
return null
}
/** The newest prose THE SESSION'S OWN AGENT wrote in the latest user turn — not a
* subagent's, whose rows share this journal and are usually the newer ones while
* a child runs. Tool-only assistant items are skipped; the user boundary clears
* prose from the preceding turn. */
export function latestStructuredAgentSessionAssistantMessage(
items: readonly AgentJournalRenderItem[]
): string {
for (let index = items.length - 1; index >= 0; index -= 1) {
const item = items[index]
const body = item?.body
if (!isRootAgentJournalItem(item)) {
continue
}
if (body?.kind === 'message' && body.role === 'user') {
return ''
}
if (body?.kind === 'message' && body.role === 'assistant') {
const prose = messageProse(body.blocks)
if (prose.trim()) {
return prose
}
}
}
return ''
}
/** The activity fields a sidebar row shows beside the prompt, named as the agent-status
* entry names them so the client can hand them straight to a row. */
export type StructuredAgentSessionStatusProjection = {
status: StructuredAgentSessionProjectedStatus | null
latestPrompt: string
/** Present only while a turn is running — see showsAgentToolPreview, which reads
* these on any state that carries them. */
toolName?: string
toolInput?: string
lastAssistantMessage?: string
/** The newest settled turn's provider verdict; present only while `status` is idle. */
turnOutcome?: AgentJournalTurnOutcome
statusStartedAt?: number
}
/** One projection shared by host and client: null status means "no turn yet", not idle.
* Every text field is bounded to the same preview an agent-status row carries — a send
* admits 256 KB, and one status frame carries every retained session at once. The
* assistant line is bounded harder than the hook field it stands in for (a preview, not
* the 8 KB body): a streamed reply re-projects on every journal checkpoint, so the frame
* has to stay small even though the row only ever renders one line of it. */
export function projectStructuredAgentSessionStatusSummary(
items: readonly AgentJournalRenderItem[],
submissions: readonly AgentJournalSubmission[] = [],
currentFence?: number | null
): StructuredAgentSessionStatusProjection {
// A first send has no journalled message until the provider replays it, so the pending
// dispatch is also what makes a brand-new session listable at all.
if (
!hasPersistedStructuredAgentSessionTurn(items) &&
!hasUnansweredStructuredAgentSessionDispatch(submissions, currentFence)
) {
return { status: null, latestPrompt: '' }
}
const status = projectStructuredAgentSessionStatus(items, submissions, currentFence)
const statusToolCall = status === 'working' ? statusStructuredAgentSessionToolCall(items) : null
const toolName = statusToolCall
? normalizeOptionalField(statusToolCall.name, AGENT_STATUS_TOOL_NAME_MAX_LENGTH)
: undefined
const toolInput = statusToolCall
? normalizeOptionalField(
describeToolInput(statusToolCall.input),
AGENT_STATUS_TOOL_INPUT_MAX_LENGTH
)
: undefined
const lastAssistantMessage = normalizeOptionalField(
latestStructuredAgentSessionAssistantMessage(items),
AGENT_STATUS_MAX_FIELD_LENGTH
)
// A verdict is a fact about a finished turn: only an idle session has one to report, and
// `readAgentJournalTurnOutcome` already answers null for anything it cannot place.
const turnOutcome =
status === 'idle' ? readAgentJournalTurnOutcome(newestStructuredAgentSessionTurn(items)) : null
const statusStartedAt = structuredAgentSessionStatusStartedAt(
status,
items,
submissions,
currentFence
)
return {
status,
latestPrompt: normalizePromptField(latestStructuredAgentSessionPrompt(items)),
...(toolName ? { toolName } : {}),
...(toolInput ? { toolInput } : {}),
...(lastAssistantMessage ? { lastAssistantMessage } : {}),
...(turnOutcome ? { turnOutcome } : {}),
...(statusStartedAt !== undefined ? { statusStartedAt } : {})
}
}
export function structuredAgentSessionPaneKey(tabId: string, sessionId: string): string {
const bytes = sha256(new TextEncoder().encode(sessionId))
const hex = Array.from(bytes.slice(0, 16), (byte) => byte.toString(16).padStart(2, '0')).join('')
const leaf = `${hex.slice(0, 8)}-${hex.slice(8, 12)}-4${hex.slice(13, 16)}-a${hex.slice(17, 20)}-${hex.slice(20, 32)}`
return `${tabId}:${leaf}`
}