mirror of
https://github.com/stablyai/orca.git
synced 2026-10-08 08:02:32 +00:00
* refactor(native-chat): the host hands out client delivery's status subscriptions as they are subscribeStatus and subscribeTurnCompletions wrapped client delivery's bound methods in forwarding lambdas; they are now the same members, the way waitForSendSettlement already is. The host is at its size limit, and the next channel it hands out needs the line. * feat(native-chat): Claude sessions write their subagents into the host status store The Claude background-task tracker queues child-work evidence at each decision it already makes (start, update, progress, terminal frame, roster replacement, turn end, session end), plus the two facts its legacy row ignores: a foreground child's progress and a foreground spawn call's result. The adapter drains that evidence after the journal handled the frame and the parent row was republished, and the host folds it into one record per child in its canonical store. Nothing reads the records yet; the strip and sidebar keep their current sources. * test(native-chat): pin the Claude child-work evidence and the host reduction of it * test(native-chat): prove every hop from a Claude frame to the host's child record The adapter delivers evidence after the frame's journal rows and the parent's republished row; the frame script keeps the parent state today reads while the records add outcome and activity; the runtime hands the evidence to the status sink under the session's own address; both entry points wire the sink to the ingest. * test(native-chat): read an optional task list as optional in the producer script * test(native-chat): an address whose publish threw carries no child work * test(agent-status): a foreign record differs from ours by producer alone * feat(native-chat): a foreground Claude child's own tool call is what its record says it is doing A child's tool traffic reaches the parent stream only for a foreground child. Read after the journal handled the frame, the child's newest call still awaiting a result becomes its open operation, previewed the way a hook-reported row previews its own tool; the result closes it. The open call is derived from the journal's own bookkeeping, not held a second time. * fix(native-chat): a Claude child restarted under a new spawn call keeps reporting to its record A task that ended and starts again stays hidden from the legacy row until a roster lists it, so the tracker held no run for it: the new run's progress reached nothing and a foreground re-run's own spawn result settled nothing. The run is now held beside the live map, where the legacy row never reads it, until a roster hands it back or it ends. A parity test pins the record's run count to the journal roster's attempt on a new spawn call, the one event both count. * refactor(native-chat): the Claude child-tool queries and translator contract get their own homes The translator's child-tool queries move into claude-child-tool-queries.ts and its contract type into claude-journal-translator-contract.ts. Brings the translator back under the size limit. * refactor(native-chat): Claude child evidence carries only its own edge's facts Admission now keeps what a child's record already knows: labels, model, owner, residency, the last message within an invocation, and a token count that never shrinks. The evidence side copied all of those forward itself, a second owner of the same rule. It now sends only what this edge observed, and a task's token count comes from the frame that reported it. * refactor(native-chat): Claude child evidence hands admission its raw labels Admission now folds provider text to one line and drops a malformed fact instead of refusing the record, so the evidence side no longer folds labels itself. The description keeps admission's longer bound. * fix(agent-status): admission alone decides a settled child's second ending The reconciliation returned before admission whenever a record had already settled with a definite outcome. That dropped the evidence an `unknown` ending carries (its last message and tokens), which admission's refine-only rule keeps, so that rule never ran for the structured producers. The latch goes. Admission keeps the definite outcome, lands the late evidence, and refuses a conflicting definite ending as `stale-invocation`, which the host ingest already counts as the fence doing its job, not a fault. Pinned through the real Claude producer and the host's own ingest. * perf(agent-status): keep child records off the status hot paths Child records made every store write and every status notification scale with the whole store. Each Claude child progress frame cost about 2 ms with 5 chats holding ~200 child records (about 14 ms at ~1,400), and every status change on any lane re-parsed every child record just to list parent rows. - The store derives each frozen record's key once instead of re-parsing it on every mutation's validation and every alias lookup. - Settled history is trimmed only when a batch settles something. - Parent listing and the structured row's revision stamp read the parents and the revision directly instead of building a full snapshot. A progress frame now costs about 0.3 ms at the same size, and listing parent rows no longer depends on how many child records the store holds. * fix(native-chat): an errored Claude spawn result no longer decides how the child ended Interrupting a foreground Claude agent while it runs a tool delivers the spawn call's errored result before the child's own killed/stopped frames. The spawn result settled the record `failed` first, and admission then refused the later `cancelled` as a conflicting ending, so an interrupted child read as a failure. An errored spawn result now settles the child `unknown`; the child's own terminal frame refines it to `cancelled` or `failed`. A successful spawn result still settles `succeeded`. The test replays both frame orders the real CLI produced when interrupted. * test(native-chat): pin a Claude foreground child's real finishing order The real CLI ends a foreground agent with its own completed update, then a notification carrying the final summary and usage, and only then the spawn call's result. Existing tests modeled the spawn result arriving first, so nothing checked that the notification's summary and tokens still land on a record the update already settled. * perf(agent-status): a store write costs what it touches, not the whole store With child records on the host, every mutation copied all five store maps and re-validated every record, and reads scanned every child and alias. A parent status publish cost about 10 ms with 4,000 child records in the store, and a child update about 13 ms. - A mutation writes into drafts over the committed maps and lands in place; a refused one is dropped with nothing to undo. The drafts keep the exact map order a copy would have. - Only what a mutation touched is re-validated: touched parents, children, aliases, facts and tombstones, plus every alias of a touched child and whatever a removed parent owned. The full validation stays for snapshot restore. - The snapshot byte budget is a running total instead of a re-measure. - Children by parent, facts by parent, aliases by child, aliases by identity and retired aliases are indexed, so reads return stored records without scanning or re-parsing. - The memoized alias identity and tombstone-key checks are gone: indexes derive them once. A parent publish now costs about 0.015 ms and a child update about 0.06 ms at 40, 1,000 and 4,000 children alike. A seeded fuzz holds the store to the copy-and-validate-everything path decision for decision, snapshot for snapshot and read for read, and a replica fed the envelopes ends identical. * fix(native-chat): a Claude child ends only on its own terminal frame The child records were fed from the legacy background-task tracker's display decisions, so they inherited rules that are not truth: a turn ending swept foreground children, a roster omitting a background child settled it, a foreground spawn call's result ended the child, and a new background start after any roster produced no record. Captured from the real CLI, an agent moved to the background keeps its own shell running for 40 s after the parent's turn ends, and that shell was settled `unknown` at the parent's `result`. Replayed with the spawn result ahead of the roster, the same agent settled as a false success and was then revived as a spurious second run. A new decoder reads the task frames directly. `task_started` opens a child (a start for an ended task id is a restart, the way messaging a finished agent resumes it), progress and a live `task_updated` update it, and a terminal `task_updated` or `task_notification` ends it. Rosters, turn ends and spawn results say nothing about a child. Every child in every capture gets its own terminal frame, so no evidenced ending is lost. The notification's `tool_use_id` names the run that ended (captured on a resumed agent's second run), so an ending from a run that is already over no longer ends the current one; a run id the record never saw still ends it, so nothing strands. The tracker, its settled-task retention and the frame readers are back to exactly what main has: the aggregate-roster split and the restart holding map are deleted, and the legacy row is unchanged by construction. * fix(agent-status): a session's end settles its live children instead of erasing them When a structured session ended, the reducer removed every child record it held, finished or not, so a reader could no longer tell how the session's work had ended. Now a child still live when its session ends settles `unknown` (nothing reported how it ended), and a child that had already ended keeps its outcome. The records still die with their parent: closing or releasing the session drops the parent row, and the store drops its children with it. A child's own outcome arriving after the session ended still refines the `unknown`. The `inventory` and `turn-ended` edges, and the rules that settled children on a roster omission or at a turn boundary, are deleted: no producer sends them any more. A restart is now its own flag on a live edge, which is what a producer reports when a finished child starts again under the same run handle. * test(native-chat): replay the real Claude CLI's frame orders into a real host Scrubbed cuts of five Claude CLI 2.1.280 stream-json captures (ids, paths and prompts replaced, frame order and relative clock kept), replayed through the adapter into a hook server: - an agent moved to the background keeps its own shell live past the parent's turn, and the shell settles at its own notification's time; - the same capture with the spawn result ahead of the move ends the agent once, from its own notification, with no second run; - a roster that omits a background child without its own ending leaves it live; - a session that ends settles what still runs `unknown` and keeps every record; - messaging a finished background agent opens its second run, which ends from its own frame; - interrupts in both captured orders end `cancelled`, and a finished foreground agent keeps its summary and usage. * test(agent-status): hold the store's running indexes and byte total to a rebuild The copying-store fuzz never reaches the snapshot byte budget, so a drift in the running byte total (or any index the public reads do not surface) passed it. After every fuzzed step, including refusals, compare every index with one rebuilt from the committed maps.
117 lines
3.7 KiB
TypeScript
117 lines
3.7 KiB
TypeScript
// One keyed table of the status store, as a mutation sees it.
|
|
//
|
|
// A mutation writes into a draft over the committed map instead of a copy of it: the draft holds
|
|
// only the keys the mutation touched, so a refused mutation is dropped without undoing anything,
|
|
// and an accepted one lands in place. Map order is part of the store's observable state (snapshot
|
|
// order, tombstone compaction), so the draft reproduces exactly the order a copied map would have.
|
|
|
|
export type AgentStatusStoreTable<V> = {
|
|
readonly size: number
|
|
get(key: string): V | undefined
|
|
has(key: string): boolean
|
|
set(key: string, value: V): void
|
|
delete(key: string): void
|
|
/** In map order, tolerating deletes of the entry being visited. */
|
|
entries(): Iterable<[string, V]>
|
|
}
|
|
|
|
const REMOVED: unique symbol = Symbol('removed')
|
|
type Change<V> = V | typeof REMOVED
|
|
|
|
export class AgentStatusStoreDraftTable<V> implements AgentStatusStoreTable<V> {
|
|
/** Every key this mutation touched, with its latest value. */
|
|
private readonly changes = new Map<string, Change<V>>()
|
|
/** Committed keys this mutation deleted (a later set re-appends them). */
|
|
private readonly removedFromBase = new Set<string>()
|
|
/** Keys present now that the committed map does not hold in place, in append order. */
|
|
private readonly appended = new Set<string>()
|
|
|
|
constructor(private readonly base: ReadonlyMap<string, V>) {}
|
|
|
|
get size(): number {
|
|
return this.base.size - this.removedFromBase.size + this.appended.size
|
|
}
|
|
|
|
get(key: string): V | undefined {
|
|
const change = this.changes.get(key)
|
|
if (change === undefined) {
|
|
return this.base.get(key)
|
|
}
|
|
return change === REMOVED ? undefined : change
|
|
}
|
|
|
|
has(key: string): boolean {
|
|
return this.get(key) !== undefined
|
|
}
|
|
|
|
set(key: string, value: V): void {
|
|
if (!this.has(key)) {
|
|
// A Map appends a key it does not hold, including one deleted earlier.
|
|
this.appended.delete(key)
|
|
this.appended.add(key)
|
|
}
|
|
this.changes.set(key, value)
|
|
}
|
|
|
|
delete(key: string): void {
|
|
if (!this.has(key)) {
|
|
return
|
|
}
|
|
this.changes.set(key, REMOVED)
|
|
if (this.base.has(key)) {
|
|
this.removedFromBase.add(key)
|
|
}
|
|
this.appended.delete(key)
|
|
}
|
|
|
|
*entries(): Iterable<[string, V]> {
|
|
for (const [key, value] of this.base) {
|
|
if (this.removedFromBase.has(key)) {
|
|
continue
|
|
}
|
|
const change = this.changes.get(key)
|
|
yield [key, change === undefined || change === REMOVED ? value : change]
|
|
}
|
|
for (const key of this.appended) {
|
|
const value = this.get(key)
|
|
if (value !== undefined) {
|
|
yield [key, value]
|
|
}
|
|
}
|
|
}
|
|
|
|
/** Each touched key with its committed value and its value after this mutation. */
|
|
*touched(): Iterable<{ key: string; previous: V | undefined; next: V | undefined }> {
|
|
for (const [key, change] of this.changes) {
|
|
yield { key, previous: this.base.get(key), next: change === REMOVED ? undefined : change }
|
|
}
|
|
}
|
|
|
|
/** Committed keys this mutation deleted, then keys it appended, in the order they now sit. */
|
|
get removedKeys(): ReadonlySet<string> {
|
|
return this.removedFromBase
|
|
}
|
|
|
|
get appendedKeys(): ReadonlySet<string> {
|
|
return this.appended
|
|
}
|
|
|
|
/** Land the draft in the committed map, leaving it in the order a copied map would have. */
|
|
commitInto(base: Map<string, V>): void {
|
|
for (const key of this.removedFromBase) {
|
|
base.delete(key)
|
|
}
|
|
for (const [key, change] of this.changes) {
|
|
if (change !== REMOVED && !this.appended.has(key)) {
|
|
base.set(key, change)
|
|
}
|
|
}
|
|
for (const key of this.appended) {
|
|
const value = this.get(key)
|
|
if (value !== undefined) {
|
|
base.set(key, value)
|
|
}
|
|
}
|
|
}
|
|
}
|