mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 16:02:29 +00:00
* refactor(native-chat): the host hands out client delivery's status subscriptions as they are subscribeStatus and subscribeTurnCompletions wrapped client delivery's bound methods in forwarding lambdas; they are now the same members, the way waitForSendSettlement already is. The host is at its size limit, and the next channel it hands out needs the line. * feat(native-chat): Claude sessions write their subagents into the host status store The Claude background-task tracker queues child-work evidence at each decision it already makes (start, update, progress, terminal frame, roster replacement, turn end, session end), plus the two facts its legacy row ignores: a foreground child's progress and a foreground spawn call's result. The adapter drains that evidence after the journal handled the frame and the parent row was republished, and the host folds it into one record per child in its canonical store. Nothing reads the records yet; the strip and sidebar keep their current sources. * test(native-chat): pin the Claude child-work evidence and the host reduction of it * test(native-chat): prove every hop from a Claude frame to the host's child record The adapter delivers evidence after the frame's journal rows and the parent's republished row; the frame script keeps the parent state today reads while the records add outcome and activity; the runtime hands the evidence to the status sink under the session's own address; both entry points wire the sink to the ingest. * test(native-chat): read an optional task list as optional in the producer script * test(native-chat): an address whose publish threw carries no child work * test(agent-status): a foreign record differs from ours by producer alone * feat(native-chat): a foreground Claude child's own tool call is what its record says it is doing A child's tool traffic reaches the parent stream only for a foreground child. Read after the journal handled the frame, the child's newest call still awaiting a result becomes its open operation, previewed the way a hook-reported row previews its own tool; the result closes it. The open call is derived from the journal's own bookkeeping, not held a second time. * fix(native-chat): a Claude child restarted under a new spawn call keeps reporting to its record A task that ended and starts again stays hidden from the legacy row until a roster lists it, so the tracker held no run for it: the new run's progress reached nothing and a foreground re-run's own spawn result settled nothing. The run is now held beside the live map, where the legacy row never reads it, until a roster hands it back or it ends. A parity test pins the record's run count to the journal roster's attempt on a new spawn call, the one event both count. * refactor(native-chat): the Claude child-tool queries and translator contract get their own homes The translator's child-tool queries move into claude-child-tool-queries.ts and its contract type into claude-journal-translator-contract.ts. Brings the translator back under the size limit. * refactor(native-chat): Claude child evidence carries only its own edge's facts Admission now keeps what a child's record already knows: labels, model, owner, residency, the last message within an invocation, and a token count that never shrinks. The evidence side copied all of those forward itself, a second owner of the same rule. It now sends only what this edge observed, and a task's token count comes from the frame that reported it. * refactor(native-chat): Claude child evidence hands admission its raw labels Admission now folds provider text to one line and drops a malformed fact instead of refusing the record, so the evidence side no longer folds labels itself. The description keeps admission's longer bound. * fix(agent-status): admission alone decides a settled child's second ending The reconciliation returned before admission whenever a record had already settled with a definite outcome. That dropped the evidence an `unknown` ending carries (its last message and tokens), which admission's refine-only rule keeps, so that rule never ran for the structured producers. The latch goes. Admission keeps the definite outcome, lands the late evidence, and refuses a conflicting definite ending as `stale-invocation`, which the host ingest already counts as the fence doing its job, not a fault. Pinned through the real Claude producer and the host's own ingest. * perf(agent-status): keep child records off the status hot paths Child records made every store write and every status notification scale with the whole store. Each Claude child progress frame cost about 2 ms with 5 chats holding ~200 child records (about 14 ms at ~1,400), and every status change on any lane re-parsed every child record just to list parent rows. - The store derives each frozen record's key once instead of re-parsing it on every mutation's validation and every alias lookup. - Settled history is trimmed only when a batch settles something. - Parent listing and the structured row's revision stamp read the parents and the revision directly instead of building a full snapshot. A progress frame now costs about 0.3 ms at the same size, and listing parent rows no longer depends on how many child records the store holds. * fix(native-chat): an errored Claude spawn result no longer decides how the child ended Interrupting a foreground Claude agent while it runs a tool delivers the spawn call's errored result before the child's own killed/stopped frames. The spawn result settled the record `failed` first, and admission then refused the later `cancelled` as a conflicting ending, so an interrupted child read as a failure. An errored spawn result now settles the child `unknown`; the child's own terminal frame refines it to `cancelled` or `failed`. A successful spawn result still settles `succeeded`. The test replays both frame orders the real CLI produced when interrupted. * test(native-chat): pin a Claude foreground child's real finishing order The real CLI ends a foreground agent with its own completed update, then a notification carrying the final summary and usage, and only then the spawn call's result. Existing tests modeled the spawn result arriving first, so nothing checked that the notification's summary and tokens still land on a record the update already settled. * perf(agent-status): a store write costs what it touches, not the whole store With child records on the host, every mutation copied all five store maps and re-validated every record, and reads scanned every child and alias. A parent status publish cost about 10 ms with 4,000 child records in the store, and a child update about 13 ms. - A mutation writes into drafts over the committed maps and lands in place; a refused one is dropped with nothing to undo. The drafts keep the exact map order a copy would have. - Only what a mutation touched is re-validated: touched parents, children, aliases, facts and tombstones, plus every alias of a touched child and whatever a removed parent owned. The full validation stays for snapshot restore. - The snapshot byte budget is a running total instead of a re-measure. - Children by parent, facts by parent, aliases by child, aliases by identity and retired aliases are indexed, so reads return stored records without scanning or re-parsing. - The memoized alias identity and tombstone-key checks are gone: indexes derive them once. A parent publish now costs about 0.015 ms and a child update about 0.06 ms at 40, 1,000 and 4,000 children alike. A seeded fuzz holds the store to the copy-and-validate-everything path decision for decision, snapshot for snapshot and read for read, and a replica fed the envelopes ends identical. * fix(native-chat): a Claude child ends only on its own terminal frame The child records were fed from the legacy background-task tracker's display decisions, so they inherited rules that are not truth: a turn ending swept foreground children, a roster omitting a background child settled it, a foreground spawn call's result ended the child, and a new background start after any roster produced no record. Captured from the real CLI, an agent moved to the background keeps its own shell running for 40 s after the parent's turn ends, and that shell was settled `unknown` at the parent's `result`. Replayed with the spawn result ahead of the roster, the same agent settled as a false success and was then revived as a spurious second run. A new decoder reads the task frames directly. `task_started` opens a child (a start for an ended task id is a restart, the way messaging a finished agent resumes it), progress and a live `task_updated` update it, and a terminal `task_updated` or `task_notification` ends it. Rosters, turn ends and spawn results say nothing about a child. Every child in every capture gets its own terminal frame, so no evidenced ending is lost. The notification's `tool_use_id` names the run that ended (captured on a resumed agent's second run), so an ending from a run that is already over no longer ends the current one; a run id the record never saw still ends it, so nothing strands. The tracker, its settled-task retention and the frame readers are back to exactly what main has: the aggregate-roster split and the restart holding map are deleted, and the legacy row is unchanged by construction. * fix(agent-status): a session's end settles its live children instead of erasing them When a structured session ended, the reducer removed every child record it held, finished or not, so a reader could no longer tell how the session's work had ended. Now a child still live when its session ends settles `unknown` (nothing reported how it ended), and a child that had already ended keeps its outcome. The records still die with their parent: closing or releasing the session drops the parent row, and the store drops its children with it. A child's own outcome arriving after the session ended still refines the `unknown`. The `inventory` and `turn-ended` edges, and the rules that settled children on a roster omission or at a turn boundary, are deleted: no producer sends them any more. A restart is now its own flag on a live edge, which is what a producer reports when a finished child starts again under the same run handle. * test(native-chat): replay the real Claude CLI's frame orders into a real host Scrubbed cuts of five Claude CLI 2.1.280 stream-json captures (ids, paths and prompts replaced, frame order and relative clock kept), replayed through the adapter into a hook server: - an agent moved to the background keeps its own shell live past the parent's turn, and the shell settles at its own notification's time; - the same capture with the spawn result ahead of the move ends the agent once, from its own notification, with no second run; - a roster that omits a background child without its own ending leaves it live; - a session that ends settles what still runs `unknown` and keeps every record; - messaging a finished background agent opens its second run, which ends from its own frame; - interrupts in both captured orders end `cancelled`, and a finished foreground agent keeps its summary and usage. * test(agent-status): hold the store's running indexes and byte total to a rebuild The copying-store fuzz never reaches the snapshot byte budget, so a drift in the running byte total (or any index the public reads do not surface) passed it. After every fuzzed step, including refusals, compare every index with one rebuilt from the committed maps.
200 lines
6.3 KiB
TypeScript
200 lines
6.3 KiB
TypeScript
// A mutation applied in place: its steps write into drafts over the committed maps, only what they
|
|
// touched is re-validated, and the drafts land together or not at all. The committed store was
|
|
// valid, so a record the mutation did not touch, and whose dependencies it did not touch, still is.
|
|
|
|
import type { AgentStatusStoreMutation } from './agent-status-store-contract'
|
|
import { agentStatusStoreHeaderBytes } from './agent-status-store-byte-budget'
|
|
import { AGENT_STATUS_STORE_LIMITS } from './agent-status-store-contract'
|
|
import {
|
|
commitAgentStatusStoreIndexes,
|
|
draftedRecordBytes,
|
|
inMapOrder,
|
|
type AgentStatusStoreDrafts,
|
|
type AgentStatusStoreIndexes
|
|
} from './agent-status-store-indexes'
|
|
import {
|
|
applyAgentStatusStoreMutationSteps,
|
|
type AgentStatusStoreMutationTables
|
|
} from './agent-status-store-mutation'
|
|
import {
|
|
agentStatusStoreSizesFit,
|
|
storedAliasIsValid,
|
|
storedChildIsValid,
|
|
storedFactIsValid,
|
|
storedParentIsValid,
|
|
storedTombstoneIsValid,
|
|
type AgentStatusStoreState
|
|
} from './agent-status-store-state'
|
|
import { AgentStatusStoreDraftTable } from './agent-status-store-table'
|
|
import { serializeAgentStatusSubject } from './agent-status-subject'
|
|
|
|
/** Present keys of one draft that an index (or this mutation) places under a query, in map order. */
|
|
function draftQuery<V>(
|
|
indexes: AgentStatusStoreIndexes,
|
|
name: 'children' | 'aliases' | 'facts',
|
|
draft: AgentStatusStoreDraftTable<V>,
|
|
indexed: Iterable<string>,
|
|
matches: (record: V) => boolean
|
|
): string[] {
|
|
const inPlace = new Set<string>()
|
|
for (const key of indexed) {
|
|
if (!draft.removedKeys.has(key)) {
|
|
inPlace.add(key)
|
|
}
|
|
}
|
|
for (const { key, next } of draft.touched()) {
|
|
if (next !== undefined && !draft.appendedKeys.has(key)) {
|
|
inPlace.add(key)
|
|
}
|
|
}
|
|
const present = (key: string) => {
|
|
const record = draft.get(key)
|
|
return record !== undefined && matches(record)
|
|
}
|
|
const keys = inMapOrder(indexes, name, [...inPlace].filter(present))
|
|
for (const key of draft.appendedKeys) {
|
|
if (present(key)) {
|
|
keys.push(key)
|
|
}
|
|
}
|
|
return keys
|
|
}
|
|
|
|
function draftTables(
|
|
state: AgentStatusStoreState,
|
|
indexes: AgentStatusStoreIndexes,
|
|
revision: number
|
|
): { drafts: AgentStatusStoreDrafts; tables: AgentStatusStoreMutationTables } {
|
|
const drafts: AgentStatusStoreDrafts = {
|
|
parents: new AgentStatusStoreDraftTable(state.parents),
|
|
children: new AgentStatusStoreDraftTable(state.children),
|
|
aliases: new AgentStatusStoreDraftTable(state.aliases),
|
|
facts: new AgentStatusStoreDraftTable(state.facts),
|
|
tombstones: new AgentStatusStoreDraftTable(state.tombstones)
|
|
}
|
|
const tables: AgentStatusStoreMutationTables = {
|
|
revision,
|
|
...drafts,
|
|
childrenOf: (parentKey) =>
|
|
draftQuery(
|
|
indexes,
|
|
'children',
|
|
drafts.children,
|
|
indexes.childrenByParent.get(parentKey) ?? [],
|
|
(child) => serializeAgentStatusSubject(child.parent) === parentKey
|
|
),
|
|
factsOf: (parentKey) =>
|
|
draftQuery(
|
|
indexes,
|
|
'facts',
|
|
drafts.facts,
|
|
indexes.factsByParent.get(parentKey) ?? [],
|
|
(fact) => serializeAgentStatusSubject(fact.subject) === parentKey
|
|
),
|
|
aliasesOfChildren: (childWorkIds) =>
|
|
draftQuery(
|
|
indexes,
|
|
'aliases',
|
|
drafts.aliases,
|
|
[...childWorkIds].flatMap((id) => [...(indexes.aliasesByChild.get(id) ?? [])]),
|
|
(alias) => childWorkIds.has(alias.childWorkId)
|
|
)
|
|
}
|
|
return { drafts, tables }
|
|
}
|
|
|
|
function touchedAreValid(
|
|
drafts: AgentStatusStoreDrafts,
|
|
tables: AgentStatusStoreMutationTables
|
|
): boolean {
|
|
const removedParents: string[] = []
|
|
for (const { key, next } of drafts.parents.touched()) {
|
|
if (!next) {
|
|
removedParents.push(key)
|
|
} else if (!storedParentIsValid(tables, key, next)) {
|
|
return false
|
|
}
|
|
}
|
|
// A parent's removal must have taken its children and facts with it.
|
|
if (
|
|
removedParents.some(
|
|
(key) => tables.childrenOf(key).length > 0 || tables.factsOf(key).length > 0
|
|
)
|
|
) {
|
|
return false
|
|
}
|
|
const touchedChildren = new Set<string>()
|
|
for (const { key, next } of drafts.children.touched()) {
|
|
touchedChildren.add(key)
|
|
if (next && !storedChildIsValid(tables, key, next)) {
|
|
return false
|
|
}
|
|
}
|
|
// An alias is valid against its child, so a touched child re-checks every alias naming it.
|
|
const aliases = new Set(tables.aliasesOfChildren(touchedChildren))
|
|
for (const { key, next } of drafts.aliases.touched()) {
|
|
if (next) {
|
|
aliases.add(key)
|
|
}
|
|
}
|
|
for (const key of aliases) {
|
|
const alias = drafts.aliases.get(key)
|
|
if (alias && !storedAliasIsValid(tables, key, alias)) {
|
|
return false
|
|
}
|
|
}
|
|
for (const { key, next } of drafts.facts.touched()) {
|
|
if (next && !storedFactIsValid(tables, key, next)) {
|
|
return false
|
|
}
|
|
}
|
|
for (const { next } of drafts.tombstones.touched()) {
|
|
if (next && !storedTombstoneIsValid(tables, next)) {
|
|
return false
|
|
}
|
|
}
|
|
return true
|
|
}
|
|
|
|
function fitsByteBudget(
|
|
epoch: string,
|
|
revision: number,
|
|
drafts: AgentStatusStoreDrafts,
|
|
recordBytes: Record<keyof AgentStatusStoreDrafts, number>
|
|
): boolean {
|
|
let bytes = agentStatusStoreHeaderBytes(epoch, revision)
|
|
for (const name of ['parents', 'children', 'aliases', 'facts', 'tombstones'] as const) {
|
|
bytes += Math.max(0, drafts[name].size - 1) + recordBytes[name]
|
|
}
|
|
return bytes <= AGENT_STATUS_STORE_LIMITS.serializedBytes
|
|
}
|
|
|
|
/** Apply one mutation to the committed store in place; false (and nothing changed) on refusal. */
|
|
export function commitAgentStatusStoreMutation(
|
|
state: AgentStatusStoreState,
|
|
indexes: AgentStatusStoreIndexes,
|
|
mutation: AgentStatusStoreMutation,
|
|
revision: number
|
|
): boolean {
|
|
const { drafts, tables } = draftTables(state, indexes, revision)
|
|
if (
|
|
!applyAgentStatusStoreMutationSteps(tables, mutation, revision) ||
|
|
!agentStatusStoreSizesFit(drafts) ||
|
|
!touchedAreValid(drafts, tables)
|
|
) {
|
|
return false
|
|
}
|
|
const recordBytes = draftedRecordBytes(indexes, drafts)
|
|
if (!fitsByteBudget(state.epoch, revision, drafts, recordBytes)) {
|
|
return false
|
|
}
|
|
commitAgentStatusStoreIndexes(indexes, drafts, recordBytes)
|
|
drafts.parents.commitInto(state.parents)
|
|
drafts.children.commitInto(state.children)
|
|
drafts.aliases.commitInto(state.aliases)
|
|
drafts.facts.commitInto(state.facts)
|
|
drafts.tombstones.commitInto(state.tombstones)
|
|
state.revision = revision
|
|
return true
|
|
}
|