Files
orca/src/shared/native-chat-transcript-projection.test.ts
T
Brennan Benson 21124db4d5 refactor(native-chat): a subagent's rows live in its own section, not in the parent's conversation (#23752)
* refactor(native-chat): a subagent's rows live with that subagent, not in the conversation

A subagent's rows were drawn in its parent's conversation, each captioned with
the subagent's name. They now belong to the subagent: the transcript projection
keeps the session's own rows as the conversation and each subagent's rows apart,
keyed by the agent id its roster entry already carries, folded on their own.

Desktop: a subagent's rows open in a section under the roster row that names it,
from that agent's roster entry, and are windowed like any other rows. A subagent
no loaded roster names opens where its first row happened, inside the section of
the agent that spawned it or in the conversation. Its edits still count in the
turn they were made, and revealing one opens the sections around it.

Mobile shows the conversation, with each spawn's roster line. Worker reads and
structured terminal reads serve the worker's own rows.

Removes what the move makes redundant: the per-row caption and its copy, the
producer check in the tool fold and the turn answer, the per-agent frontier
interleaved in the conversation, worker-text subagent tags, and the agent id on
worker-read messages.

* refactor(native-chat): a diff target names the sections its row sits in

Revealing a subagent's edit opens the sections around it from the target the
rollup already holds, instead of looking the row up at click time. The section
head keeps to the agent's name and dot; its state in words stays on the roster
entry. The worker page test stubs the host through its module rather than a cast.

* fix(native-chat): a working subagent's section is open; a worker page windows its own rows

A subagent's section is open while its agent works and closes once it settles,
the way the turn's own live run does; a section the reader opened or closed by
hand keeps that choice. A subagent another subagent spawned opens inside that
one's section, so a working grandchild shows inside its working parent. Openness
is derived from the roster's state and the reader's choices; nothing stores an
automatic open.

A worker page is now the newest page of the worker's own rows. The host windows
the read over them before the limit, so a subagent's burst can no longer crowd
the worker's rows off the page, and "older" still means older worker rows. The
scope is an in-process argument of the host's history read; no wire request
carries it.

* fix(native-chat): a subagent section head names the turn it sits in, for the outline rail

* fix(native-chat): a subagent section's rows sit in the turn the section is shown in, for the outline rail

A background subagent's rows written during a later turn carried that later
turn onto their slots, so scrolling through its section lit the later turn's
rail tick and then snapped back. The rollup still counts each edit in the turn
it was made; only the slot, which the rail reads, takes the shown turn.

* fix(mobile): Load earlier reads past pages that hold only a subagent's rows

Mobile draws only the session's own rows, so an older page made entirely of a
subagent's rows landed as nothing: the reader tapped Load earlier, saw the
spinner, and got the same transcript back. One load now reads on (up to 8 pages)
until a page holds a row of the session's own, then applies the pages in order.

* test(mobile): stub the RPC client the way the other structured-session hook tests do

* perf(native-chat): order subagent rows for the changed-files rollup once per change to them

The rollup flattened and re-sorted every subagent row on each update, including
every token the parent streamed. The ordering now keys on the projection's
subagent rows, which keep their identity while only the conversation changes.

* refactor(native-chat): order subagent rows in the sections hook, keeping the list under its line limit

* fix(mobile): a transcript whose newest page is only a subagent's rows reads back on its own

Opened while a subagent is busy, the newest page can hold nothing but that
subagent's rows. Mobile draws none of them, so the reader saw an empty chat with
a Load earlier button, and an empty list cannot be scrolled to page. The hook now
reads back once from each such head, and the read runs on to the session's own rows.

* fix(native-chat): count the live window in the session's own rows, so a subagent's burst keeps its roster

The live window kept the newest 1,024 rows of every agent. A subagent writing
more than that trimmed its own spawn's roster row and the prompt, and its
section fell back to a closed, unnamed header. The window now keeps the newest
1,024 of the session's own rows and everything after, with an 8,192-row cap on
every agent's rows as the memory backstop. A transcript with no subagent rows
trims exactly as before.

* fix(agent-session): window history pages by the session's own rows, with a subagent's rows riding along

A history page held the newest 200 rows of every agent, so a subagent's burst
could fill a page on its own: the phone opened on an empty chat and "Load
earlier" landed nothing. A page now starts at the oldest of the newest `limit`
rows of the session's own and serves every row from there, so the subagent's
rows come with the conversation they happened in. The page stays contiguous,
the cursor still names its first row, and the byte bound still applies. A
transcript with no subagent rows gets the same pages as before.

Clients already take a page larger than its limit: both reducers raise their
retained window to the page's size. The mobile read-on and read-back stay for
older hosts.

* test(agent-session): a page reaches back to the start rather than leaving a subagent-only page

* fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it

The live window trimmed to just after the own row it dropped, so a subagent
whose roster row went kept its rows at the top as an unnamed section until
the parent wrote again. Trim to the oldest own row kept instead; it still
fires only once an own row passes the limit, so a paged-in run of subagent
rows at the head stays until then. With no subagent rows nothing changes.

* perf(native-chat): cap the live window at 4,096 rows, bounding each delta's re-derivation

Every live batch re-derives the transcript over every retained row. On the
largest real window (7,374 rows) that cost 7-8 ms a delta on desktop against
0.6 ms at the old 1,024-row window, and held about 26 MB of row content.
4,096 halves both. The most rows any local journal puts between a roster and
its subagent's last row, with the parent inside its own-row limit, is 3,005,
so no observed subagent loses its roster to the lower cap.

* fix(native-chat): a subagent section opens only while its roster is the running scope's live frontier

A section used to open whenever its roster said the subagent was working, anywhere
in the transcript and whether or not the session was running, so a background
subagent's section stayed open and grew mid-transcript while the parent moved on.

It now opens by default only while the session runs and the roster row naming the
subagent is the newest thing the parent produced, user rows aside. Newer parent
output closes it even while the subagent still works; the roster row keeps
showing that live state. A subagent still working is a running scope of its own
for the sections it spawned; a settled one closes its scope. Derived every
render, no latch; the reader's own open or close still wins.

* fix(native-chat): name a subagent's section from a client roster the window never trims

A section took its name and state from a roster row in the loaded window. Once a
burst trimmed that row, or the row sat on an older page, the section fell back to
an unnamed, closed "Subagent" header.

The shared reducer now keeps a roster keyed by agent id, folded from every roster
row and revision the client receives: pages, older pages and live batches,
including revisions of roster rows outside the window, which live batches already
carry. The first roster naming an agent wins and its revisions update it; a
removed roster row drops its entries; it is rebuilt on every page that replaces
the window and bounded to 512 agents. Sections take their name, state and
live-frontier place from it; placement stays under the loaded roster row, else
at the section's first loaded row. Only a subagent no roster ever named stays
unnamed.

* feat(agent-session): a history page names the subagents whose roster row is older than it

A page is a contiguous run of the journal whose older-page cursor is its first
item, so it cannot pull an older roster row in without skipping the rows between.
When a page held a subagent's rows but not the roster row naming it (about 11% of
the moments a reader could open a session on local journals), that subagent drew
as an unnamed "Subagent" header.

History and hydration pages now carry an optional `subagentRoster`: the first
roster entry naming each subagent whose rows are on the page and whose roster row
is not, with the row's id, sequence and revision; bounded to 64 entries and
16 KB. Items and cursor are unchanged. The client seeds its roster from it.

Rule 1 in docs/reference/remote-wire-compatibility.md: an optional field on an
existing frame, no capability gate. An older client ignores it (the released
reducer reads a page with it exactly as one without); against an older host the
field is absent and the section falls back to an unnamed header.

* Revert "fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it"

This reverts commit 22078656b2.

Its only purpose was to stop a subagent whose roster row an own-row trim had
dropped from showing at the head of the window as an unnamed section. The client
roster now names that section whatever the window holds, so the cut is back at
just after the own row the limit passes. The retention test that pinned the
unnamed-section case now asserts the section at the head keeps its name.

* chore(native-chat): state the retention limits' own reasons, now that no name depends on the window

Own-row retention keeps the conversation a reader sees from being crowded out by
rows drawn as a one-row section on desktop and not at all on mobile; the
every-agent cap bounds memory and each live delta's re-derivation. Neither is
about keeping a roster row loaded any more.

* fix(native-chat): hold the roster fold's draft map where type narrowing can see closure writes

* fix(native-chat): a roster row's newer revision replaces it in the client roster too

A revision that stops naming an agent (the host drops an entry it learns is not a
subagent, or re-keys a provisional one) left the client roster holding the old
entry, often still "working", with nothing to re-derive it. The section then read
as working forever and could auto-open, while a fresh read of the same journal
left it unnamed. The fold now drops an entry when a newer revision of the row that
named it no longer does, before any roster takes it over.

* fix(native-chat): a parent's spawn and wait calls keep the subagent they name open

A subagent section auto-opened only while its roster row was the running session's
newest row, so any later row closed it: a Codex wait on the agent, or the parent's
text before its next spawn call. Now a row that is part of delegating to a subagent
keeps that subagent open:

- a Codex collab call (spawn, wait, resume, message, close) opens each agent its
  receiver thread ids name; one naming none is ordinary output;
- a Claude spawn call names no agent, so it counts toward the roster announcing it;
- a roster row at the frontier opens its most recently added agent, not all of them.

A roster or call naming only agents one subagent spawned is that subagent's output,
so a grandchild's roster, which the host journals as the session's row, no longer
closes the spawner's section.

* fix(native-chat): a parent's call right after the roster closes its subagent's section

A parent's tool calls after a roster row fold into the tool run drawn above
the roster, so the roster stayed the newest drawn row and its section stayed
open while the parent was already reading or running commands. The fold now
records the newest journal position among the rows it merged, and the live
frontier orders rows by that newest part. The layout is unchanged. A spawn
call folded there still counts as part of the roster announcing it.

* fix(native-chat): a Codex call naming several subagents delegates to the first

A Codex collab call that names several agents opened every one of their
sections. It now counts as delegating to the first agent it names, so one
section opens, the same as a call naming one agent.

* fix(native-chat): closing a roster's list closes the sections under it

Collapsing a roster row's list of subagents left their open sections drawn,
so the section's own head became the only way to close them. And the list's
open state lived in the row, so a row the window unmounted came back
collapsed.

The transcript now holds each roster list's open state beside the section
choices. A closed list hides every section it anchors; each section keeps
its own open or closed choice for when the list reopens. With no choice from
the reader, a list is open while a section under it is open. Closing a
section from its entry keeps the list open, and revealing a subagent's edit
opens the list it sits under.

* perf(native-chat): a reveal finds the roster lists it opens with one set lookup per entry

* fix(native-chat): a subagent's roster entry heads its own rows

An open section drew the agent's name twice: its entry in the roster's list,
then a separate section head above its rows. The entry is now the head. The
roster row draws its entries through the first open one, that agent's rows
follow, then the entries after it, each run in its own windowed slot. A
section no loaded roster row holds (an older page, a grandchild, an unnamed
agent) keeps its own head.

A roster list is open while the live frontier or a reader's choice is on
one of its agents, unless the reader closed the list, so closing an agent
from its entry no longer needs to pin the list open.

The section emitter moves to its own module, and the trailing-run
predicates it shares with the slot builder to theirs, to keep the slot
builder under its line limit.

* fix(native-chat): the entries after an open subagent's rows set in its roster's type

The roster row's list inherits the system row's small muted type; the entries that
follow an open section sit outside that row, so they now carry the same type.

* docs(native-chat): a current host can also serve a page of only a subagent's rows

A page is bounded by bytes after it is windowed by the session's own rows, so a
burst that fills the bound yields a page, or an opening page, with none of the
session's own rows. Mobile's read-on and read-back therefore serve current hosts
too, not only older ones; the comments said otherwise. The retention comment
still described a closed section as a row of its own; it now sits behind its
roster entry.

* fix(native-chat): a section's prose keeps its copy/timestamp controls inside the section

An assistant row's hover controls (copy, scroll-to-top, timestamp) hang 20px
below the row into the gap before the next one (`-mb-5`). Inside a subagent's
section that put them below the section's left border, and on the section's
last row they touched the parent's next row with no gap.

Inside a section the controls now stay in flow, so the border covers them and
the next row sits the normal gap below. The row-height estimate reserves the
same 20px for a section's prose so windowing does not jump on measure.

* test(agent-session): state each appended row's turn scope, as the journal now requires

* refactor(native-chat): the client's journal retention policy lives in its own module
2026-09-29 23:47:02 -07:00

180 lines
6.5 KiB
TypeScript

import { describe, expect, it } from 'vitest'
import type {
AgentJournalItemBody,
AgentJournalRenderItem,
AgentJournalTurnScope
} from './agent-session-journal-types'
import type { NativeChatMessage } from './native-chat-types'
import {
projectNativeChatTranscript,
projectNativeChatTranscriptMessages
} from './native-chat-transcript-projection'
let sequence = 0
function row(
id: string,
blocks: NativeChatMessage['blocks'],
overrides: Partial<NativeChatMessage> = {}
): NativeChatMessage {
sequence += 1
return {
id,
role: 'assistant',
blocks,
timestamp: 1,
source: 'transcript',
journalPosition: { sequence, index: 0 },
...overrides
}
}
const say = (value: string) => [{ type: 'text' as const, text: value }]
const call = (name: string) => [{ type: 'tool-call' as const, name, input: {} }]
const child = { agentId: 'task-1', producerKind: 'agent' as const }
/** A journal item as a host that states each row's turn writes it; `turnItemId` absent: none. */
function item(
itemId: string,
body: AgentJournalItemBody,
turnItemId?: string,
agentId?: string
): AgentJournalRenderItem {
const turnScope: AgentJournalTurnScope =
turnItemId === undefined ? { kind: 'thread' } : { kind: 'turn', turnItemId }
return {
itemId,
body,
sequence: 0,
observedAt: 0,
revision: 1,
turnScope,
...(agentId === undefined ? {} : { agentId })
}
}
const message = (role: 'user' | 'assistant'): AgentJournalItemBody => ({
kind: 'message',
role,
blocks: []
})
const turn = (turnId: string, userItemId: string): AgentJournalItemBody => ({
kind: 'turn',
turnId,
state: 'completed',
userItemId,
startedAt: 0
})
describe("a subagent's rows are not the conversation's", () => {
// The shape a background subagent leaves: its rows interleave with its parent's.
const transcript = [
row('ask', say('review the PR'), { role: 'user' }),
row('delegate', say('Delegating the review.')),
row('child-look', say('Looking at the diff.'), child),
row('parent-read', call('Read')),
row('child-grep', call('Grep'), child),
row('child-verdict', say('The PR is CLEAN.'), child),
row('answer', say('The review found nothing.'))
]
it("keeps the session's own rows as the conversation, and none of the subagent's", () => {
const { conversation, subagentRows } = projectNativeChatTranscript(transcript)
expect(conversation.map((message) => message.id)).toEqual(['ask', 'delegate', 'answer'])
expect(conversation[1]?.blocks).toEqual([...say('Delegating the review.'), ...call('Read')])
expect(
subagentRows.get('task-1')?.map(({ message, turnKey }) => [message.id, turnKey])
).toEqual([
['child-look', 'ask'],
['child-verdict', 'ask']
])
expect(projectNativeChatTranscriptMessages(transcript)).toEqual(conversation)
})
it("folds the subagent's calls into its own run however its parent's interleave", () => {
const [look] = projectNativeChatTranscript(transcript).subagentRows.get('task-1') ?? []
expect(look?.message.blocks).toEqual([...say('Looking at the diff.'), ...call('Grep')])
})
it('splits a subagent run at the turn it crossed, so each call stays in its own turn', () => {
const rows = projectNativeChatTranscript([
row('first', say('go'), { role: 'user' }),
row('child-start', say('Starting.'), child),
row('second', say('and then'), { role: 'user' }),
row('child-edit', call('Edit'), child)
]).subagentRows.get('task-1')
expect(rows?.map(({ message, turnKey }) => [message.id, turnKey])).toEqual([
['child-start', 'first'],
['child-edit', 'second']
])
})
// The conversation's grouping rule, not position: a send made mid-turn does not own the
// rows written before its own turn opened, and a turn keyed to its record owns its rows.
it("takes each row's turn from the journal's turn records, as the conversation's rows do", () => {
const rows = [
row('first', say('go'), { role: 'user' }),
row('child-start', say('Starting.'), child),
row('second', say('and then'), { role: 'user' }),
row('child-edit', call('Edit'), child),
row('child-woke', call('Write'), child)
]
const journal = {
items: [
item('first', message('user')),
item('turn-1', turn('1', 'first')),
item('child-start', message('assistant'), 'turn-1', 'task-1'),
item('second', message('user')),
item('child-edit', message('assistant'), 'turn-1', 'task-1'),
item('turn-2', turn('2', 'second')),
// Its opener is outside the loaded window, so the turn keys to its own record.
item('turn-3', turn('3', 'older-send')),
item('child-woke', message('assistant'), 'turn-3', 'task-1')
],
submissions: []
}
const projected = projectNativeChatTranscript(rows, undefined, journal).subagentRows.get(
'task-1'
)
expect(projected?.map(({ message, turnKey }) => [message.id, turnKey])).toEqual([
['child-start', 'first'],
['child-woke', 'turn-3']
])
expect(projected?.[0]?.message.blocks).toEqual([...say('Starting.'), ...call('Edit')])
})
it("never lets a subagent's own prompt open a conversation turn", () => {
const transcript = [
row('ask', say('go'), { role: 'user' }),
row('child-prompt', say('Review the diff.'), { ...child, role: 'user' }),
row('child-look', say('Looking.'), child)
]
const rows = projectNativeChatTranscript(transcript).subagentRows.get('task-1')
expect(rows?.every(({ turnKey }) => turnKey === 'ask')).toBe(true)
// A host that states turns scopes a prompt written after the turn ended to none.
const journal = {
items: [
item('ask', message('user')),
item('turn-1', turn('1', 'ask')),
item('child-prompt', message('user'), undefined, 'task-1'),
item('child-look', message('assistant'), 'turn-1', 'task-1')
],
submissions: []
}
const scoped = projectNativeChatTranscript(transcript, undefined, journal).subagentRows
expect(scoped.get('task-1')?.map(({ turnKey }) => turnKey)).toEqual([undefined, 'ask'])
})
it('projects a transcript that names no producer exactly as before', () => {
const plain = transcript.map(({ agentId: _agentId, producerKind: _kind, ...rest }) => rest)
const { conversation, subagentRows } = projectNativeChatTranscript(plain)
expect(subagentRows.size).toBe(0)
expect(conversation.map((message) => message.id)).toEqual([
'ask',
'delegate',
'child-look',
'child-verdict',
'answer'
])
})
})