Files
orca/src/shared/agent-session-context-usage-schema.ts
T
Brennan Benson 25d7c21fcb feat(native-chat): show context window usage in the composer (#22301)
* refactor(native-chat): move the composer's stop action into its own hook

The composer is at its line budget; lifting the stop action out makes room
for the context usage ring without changing what Stop does.

* feat(native-chat): record the Claude CLI's context window facts on the structured journal

A structured Claude session now keeps what the CLI says about its context
window on journal rows, so every client reads the same answer and a restart
replays it:

- Each main-thread assistant response keeps its API usage. A subagent's
  response measures its own window, so it carries none.
- The turn a result settles records the session's window: the largest
  contextWindow across the result's per-model usage, since side calls to a
  smaller model report their own smaller window.
- After a result and after a compaction boundary the host asks the CLI for
  its /context breakdown (5s bound) and records the answer on the current or
  last turn. An answer is dropped when the main conversation moved, a send was
  accepted, a newer request was issued, or the session was released while it
  was in flight; a failure or an older CLI leaves the row unchanged.
- A compaction boundary or conversation reset records that the used count is
  unknown until the next response or report, so the pre-compaction size is
  never shown as current.

Every part carries its own host clock, and the reader takes the newest, since
a revised turn row keeps its place in the transcript. The persisted validator
admits every value the writer can write, including a zero auto-compact
threshold: a row replay rejects truncates the journal from that row.

* feat(native-chat): show context window usage in the composer

A structured Claude chat shows a ring beside send once the journal can state
the session's context usage. Hovering shows used/window with a bar and, when
the CLI has reported its breakdown, one row per CLI category as a share of the
window, largest first. Between reports the ring shows the newest response's
usage against the newest window the CLI reported, marked as estimated. Before
the CLI has reported any window, and after a compaction or reset until the next
response, there is no ring. A terminal-backed chat shows none.

* fix(native-chat): measure the context ring against the main thread's model window

The result's per-model usage is cumulative across the session and includes
subagents and side calls, so the largest window was often not the one the
main conversation runs in: after switching from a 1M model to a 200k one, or
when a subagent ran on a larger-window model, the ring read against the
wrong window after every turn. Pick the entry named by the model that served
the newest main-thread response, and among its [1m]/non-[1m] entries the one
the result moved; fall back to the largest only when nothing names it.

* fix(native-chat): keep the context ring moving through tool-only responses

The live estimate lived on assistant message rows, and a response with only
tool calls or thinking writes no message row, so the ring froze through long
tool loops and stayed hidden after a mid-turn auto-compaction until the next
text reply. Record every main-thread response's usage on its turn row
instead, once per response, so the selector sees each one.

* fix(native-chat): read the ring's window from the model the turn's init names

The CLI keys per-model usage by the main loop's model string, [1m] included,
and every turn's system/init frame carries that exact string, while a
response drops the suffix. Match the init's key first, so a session that
switched between the 1M and 200k variants of one model reads the right
window; fall back to the newest response's model, then the largest entry.

* fix(native-chat): correct the composer's control-order note for the context ring

* fix(native-chat): keep a running turn's context facts when the host settles it

A turn row now carries the live context estimate while it runs. When the host
settles a running row itself (a crashed or stale generation, a close the
translator never saw), it rebuilt the record field by field and dropped those
facts, so after a crash the ring fell back to an older turn's size, or to a
pre-compaction size the dropped reset had superseded.

* refactor(native-chat): revise Claude turn rows from the journal so the ring survives a restart

The context ring's facts were written to turn rows through an in-memory list
of recent turns. A new translator is built on every acquisition, so after a
restart or reattach that list was empty and every fact for a turn that was
not open was dropped: a /compact as the first action after a restart never
cleared the ring and never showed the fresh breakdown.

Every Claude turn-row write is now a revision of the row as the bound journal
holds it when the write runs. The sink gains a resolved revision that reads
the target row and its body at execution; the queue runs one operation at a
time, so the read-modify-write cannot interleave, and revisions are never
coalesced. Lifecycle writes own the lifecycle fields and context writes own
contextUsage; each keeps every other field. Only the open turn is kept in
memory. A fact with no open turn lands on the newest turn row, and a report
lands on the turn it was requested for.

The persisted facts are simplified to a window, which now names the model it
was measured for, and a single used part (report, estimate or unknown) that
each write replaces. The ring reads the newest turn row carrying each part,
and hides an estimate whose model the window was not measured for instead of
dividing by another model's window. A turn opening, and a reset, count as
activity, so a late report can never land behind a newer turn.

Host settlement of a stale running turn now drops only the fields its verdict
owns, so context facts and any field a newer build wrote survive it.

* fix(native-chat): keep a turn row whose context facts this build cannot read

Context facts are validated deeply, so one malformed or future-shaped fact
made the whole turn row malformed, and replay truncates the journal from that
row on. Replay now drops unreadable facts from a turn row, in item rows and in
settlement batches, and keeps the row, the same way it already drops producer
linkage it cannot trust. The ring shows nothing for that turn instead of the
session losing its history.

* test(native-chat): pin that a child exit mid-turn keeps the ring's last size

A lifecycle-only revision, the end a turn gets when its child exits without a
result, must keep the context facts the row already carries.

* perf(native-chat): revise a named Claude turn row by key instead of scanning the journal

Every Claude turn-row write walked every reduced journal item to find its
row, even when it already knew the row's identity, so a long session paid
O(items) per write on the main process. The journal now answers a keyed read,
and a context report names its turn by row identity rather than turn id, so
only a write made while no turn is open still scans.

* fix(native-chat): tell a 1M window from a 200k one of the same model

Responses drop the [1m] suffix, so after a switch between the 1M and 200k
windows of one model the running turn was measured against the previous
turn's window until its result arrived. An estimate now records the turn's
init model, which keys the window exactly, and the reader requires the full
model id to match.

* fix(native-chat): show no ring for a context kind a newer host writes

A paired client reads turn rows from the host unvalidated, so a used-count
kind this build does not know fell through to the estimate branch and threw
reading its missing usage. Only the kinds this build can measure now produce
a ring.

* fix(native-chat): keep the context ring through plan-mode turns on another model

Plan mode can run a turn on a model the turn's init does not name (opusplan
upgrades to Opus's 1M window). The estimate then carried only the response's
id, which drops [1m], and the exact comparison against the window hid the ring
for every plan-mode turn. The estimate now records the response's id beside
the init's exact key, and the reader matches the base model only when no exact
key was recorded.

* refactor(native-chat): pair the context ring's window by model change, not by model id

The ring divided the newest response's size by the newest window only when
their model ids matched, which meant comparing ids from the init frame, the
response, per-model usage keys and canonical ids. Those disagree in plan mode
and across 1M and 200k windows of one model.

The writer now knows when the model may have changed: after a model or
permission-mode write that changes the value, when a restore cannot put the
stored model back, and when a main-thread response comes from a different
model than the one the window serves (an approved plan). It then marks the
size unknown, holds estimates, and asks the CLI for its context report, which
states the new model's window. Any new window, from a report or a turn
result, releases the hold. The reader compares nothing: a report, or the
newest estimate over the newest window.

Turn rows no longer store window.model, window.canonicalModel,
estimate.model or estimate.responseModel.

* fix(native-chat): keep a late context report's window when only its count went stale

* test(native-chat): pin that each turn's init lets its result restate the context window

* fix(native-chat): open the context card on click and tap

* fix(native-chat): wait a beat before a mouse hover opens the context card

* fix(native-chat): publish each context write in the operation that makes it

A context report answers after the turn's last frame, so a revision that waited
for the next frame's publish reached live clients only on the next turn. Context
writes now queue their revision and its publication as one operation.

* fix(native-chat): write million-token counts with a capital M

A lowercase m read as minutes on the context card.

* fix(native-chat): keep the context card open while the pointer crosses into it

The card closed the moment a mouse left the ring, so the pointer could not
cross the gap into the card. Leaving now waits a beat, and entering the card
cancels the close.

* fix(native-chat): let Escape close the context card without stopping the agent

The card keeps focus in the composer, so the Escape that closed it also
reached the composer and interrupted the running turn. The composer now
skips an Escape an open layer already handled.

* fix(native-chat): show the context ring when the chat has not loaded the turn it belongs to

The ring read context facts only from the rows the chat had loaded, so a
reopened chat whose recent page started after the last turn row, or a live
turn longer than the retained window, showed no ring until the next turn.

The host now derives the newest context facts from its whole journal with
the same selector the chat uses, and returns them on agentSession.options
for sessions that write them. The chat prefers each fact its loaded rows
carry and takes the host answer for a fact they lack. When a live batch
revises a turn row older than the loaded window, the chat asks for options
again so that answer stays current.

* fix(native-chat): bound context refresh reads and refresh when the turn row is trimmed

Each turn-row revision the loaded window missed started its own options
read. Those reads share the session's host queue with sends and interrupts,
and each asks the CLI for its settings, so a burst could pile reads in front
of a user action and discard every answer before it landed. The chat now
keeps one options read in flight and at most one behind it.

A live turn longer than the retained window also lost its turn row to the
trim without asking for a fresh host answer, so the ring fell back to the
answer read at turn start until the next response. Trimming a turn row now
asks again, like a dropped revision does.

* fix(native-chat): show the context ring from the first response, sized from the session's model

A new session has no measured window until its first result, so the ring
stayed hidden for the whole first turn. The host now keeps the window the
applied model's name implies (1M for a [1m] name, unknown for default, 200k
otherwise) and writes it beside an estimate when the journal holds no window,
or after a model write, until the result or the CLI's report replaces it.

* fix(native-chat): imply a context window only from a [1m] model name

A bare model name does not fix the window: first-party runs today's opus,
sonnet and fable models natively at 1M while a gateway or cloud provider runs
them at 200k, and opusplan and haiku run another model in plan mode. Sizing
their first response at 200k read the ring about five times too full, so only
a [1m] name implies a window now; any other name waits for the result.

* fix(native-chat): size the first response from a report taken before any turn

A model picked in a chat with no turn yet asks the CLI for its context
report, but with no turn row the report's write lands nowhere. Recording it
still marked the journal as holding a window, so the first response wrote
none and the ring stayed hidden until the turn's result.

The report's window now serves as the fallback a response writes while the
journal holds no window, and recording a report no longer assumes its write
landed.

* test(native-chat): move the fake Claude connection out of the structured integration suite

The context-report delivery case pushed the suite past the 800-line limit,
failing repo-wide lint. The fake child now lives in its own fixture.
2026-09-24 00:12:08 -07:00

62 lines
1.9 KiB
TypeScript

// Admission for the context facts a turn row carries. Bounded to what the
// writer can produce, so a replayed row can never hold more than a live one.
import { z } from 'zod'
import {
MAX_CONTEXT_CATEGORIES,
MAX_CONTEXT_CATEGORY_NAME_CHARS,
MAX_CONTEXT_MODEL_ID_CHARS,
type AgentSessionContextUsage
} from './agent-session-context-usage'
const TokenCount = z.number().finite().nonnegative()
const WindowTokens = z.number().finite().positive()
const CapturedAt = z.number().finite()
const ModelId = z.string().min(1).max(MAX_CONTEXT_MODEL_ID_CHARS)
const TokenUsage = z.object({
inputTokens: TokenCount,
cacheCreationInputTokens: TokenCount,
cacheReadInputTokens: TokenCount,
outputTokens: TokenCount
})
const Used = z.discriminatedUnion('kind', [
z.object({
kind: z.literal('report'),
model: ModelId,
usedTokens: TokenCount,
windowTokens: WindowTokens,
percentage: z.number().finite(),
autoCompactAtTokens: TokenCount.optional(),
categories: z
.array(
z.object({
name: z.string().min(1).max(MAX_CONTEXT_CATEGORY_NAME_CHARS),
tokens: TokenCount,
deferred: z.literal(true).optional()
})
)
.max(MAX_CONTEXT_CATEGORIES),
capturedAt: CapturedAt
}),
z.object({ kind: z.literal('estimate'), usage: TokenUsage, capturedAt: CapturedAt }),
z.object({ kind: z.literal('unknown'), capturedAt: CapturedAt })
])
export const AgentSessionContextUsageSchema = z.object({
window: z.object({ tokens: WindowTokens, capturedAt: CapturedAt }).optional(),
used: Used.optional()
})
export function isAdmissibleAgentSessionContextUsage(
value: unknown
): value is AgentSessionContextUsage {
return AgentSessionContextUsageSchema.safeParse(value).success
}
type Admits<T extends true> = T
export type CanonicalContextUsageIsAdmissible = Admits<
AgentSessionContextUsage extends z.input<typeof AgentSessionContextUsageSchema> ? true : false
>