mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 08:02:21 +00:00
7c6a37eeba08172f0aae3029fbad9a0e106a70aa
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
21124db4d5 |
refactor(native-chat): a subagent's rows live in its own section, not in the parent's conversation (#23752)
* refactor(native-chat): a subagent's rows live with that subagent, not in the conversation
A subagent's rows were drawn in its parent's conversation, each captioned with
the subagent's name. They now belong to the subagent: the transcript projection
keeps the session's own rows as the conversation and each subagent's rows apart,
keyed by the agent id its roster entry already carries, folded on their own.
Desktop: a subagent's rows open in a section under the roster row that names it,
from that agent's roster entry, and are windowed like any other rows. A subagent
no loaded roster names opens where its first row happened, inside the section of
the agent that spawned it or in the conversation. Its edits still count in the
turn they were made, and revealing one opens the sections around it.
Mobile shows the conversation, with each spawn's roster line. Worker reads and
structured terminal reads serve the worker's own rows.
Removes what the move makes redundant: the per-row caption and its copy, the
producer check in the tool fold and the turn answer, the per-agent frontier
interleaved in the conversation, worker-text subagent tags, and the agent id on
worker-read messages.
* refactor(native-chat): a diff target names the sections its row sits in
Revealing a subagent's edit opens the sections around it from the target the
rollup already holds, instead of looking the row up at click time. The section
head keeps to the agent's name and dot; its state in words stays on the roster
entry. The worker page test stubs the host through its module rather than a cast.
* fix(native-chat): a working subagent's section is open; a worker page windows its own rows
A subagent's section is open while its agent works and closes once it settles,
the way the turn's own live run does; a section the reader opened or closed by
hand keeps that choice. A subagent another subagent spawned opens inside that
one's section, so a working grandchild shows inside its working parent. Openness
is derived from the roster's state and the reader's choices; nothing stores an
automatic open.
A worker page is now the newest page of the worker's own rows. The host windows
the read over them before the limit, so a subagent's burst can no longer crowd
the worker's rows off the page, and "older" still means older worker rows. The
scope is an in-process argument of the host's history read; no wire request
carries it.
* fix(native-chat): a subagent section head names the turn it sits in, for the outline rail
* fix(native-chat): a subagent section's rows sit in the turn the section is shown in, for the outline rail
A background subagent's rows written during a later turn carried that later
turn onto their slots, so scrolling through its section lit the later turn's
rail tick and then snapped back. The rollup still counts each edit in the turn
it was made; only the slot, which the rail reads, takes the shown turn.
* fix(mobile): Load earlier reads past pages that hold only a subagent's rows
Mobile draws only the session's own rows, so an older page made entirely of a
subagent's rows landed as nothing: the reader tapped Load earlier, saw the
spinner, and got the same transcript back. One load now reads on (up to 8 pages)
until a page holds a row of the session's own, then applies the pages in order.
* test(mobile): stub the RPC client the way the other structured-session hook tests do
* perf(native-chat): order subagent rows for the changed-files rollup once per change to them
The rollup flattened and re-sorted every subagent row on each update, including
every token the parent streamed. The ordering now keys on the projection's
subagent rows, which keep their identity while only the conversation changes.
* refactor(native-chat): order subagent rows in the sections hook, keeping the list under its line limit
* fix(mobile): a transcript whose newest page is only a subagent's rows reads back on its own
Opened while a subagent is busy, the newest page can hold nothing but that
subagent's rows. Mobile draws none of them, so the reader saw an empty chat with
a Load earlier button, and an empty list cannot be scrolled to page. The hook now
reads back once from each such head, and the read runs on to the session's own rows.
* fix(native-chat): count the live window in the session's own rows, so a subagent's burst keeps its roster
The live window kept the newest 1,024 rows of every agent. A subagent writing
more than that trimmed its own spawn's roster row and the prompt, and its
section fell back to a closed, unnamed header. The window now keeps the newest
1,024 of the session's own rows and everything after, with an 8,192-row cap on
every agent's rows as the memory backstop. A transcript with no subagent rows
trims exactly as before.
* fix(agent-session): window history pages by the session's own rows, with a subagent's rows riding along
A history page held the newest 200 rows of every agent, so a subagent's burst
could fill a page on its own: the phone opened on an empty chat and "Load
earlier" landed nothing. A page now starts at the oldest of the newest `limit`
rows of the session's own and serves every row from there, so the subagent's
rows come with the conversation they happened in. The page stays contiguous,
the cursor still names its first row, and the byte bound still applies. A
transcript with no subagent rows gets the same pages as before.
Clients already take a page larger than its limit: both reducers raise their
retained window to the page's size. The mobile read-on and read-back stay for
older hosts.
* test(agent-session): a page reaches back to the start rather than leaving a subagent-only page
* fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it
The live window trimmed to just after the own row it dropped, so a subagent
whose roster row went kept its rows at the top as an unnamed section until
the parent wrote again. Trim to the oldest own row kept instead; it still
fires only once an own row passes the limit, so a paged-in run of subagent
rows at the head stays until then. With no subagent rows nothing changes.
* perf(native-chat): cap the live window at 4,096 rows, bounding each delta's re-derivation
Every live batch re-derives the transcript over every retained row. On the
largest real window (7,374 rows) that cost 7-8 ms a delta on desktop against
0.6 ms at the old 1,024-row window, and held about 26 MB of row content.
4,096 halves both. The most rows any local journal puts between a roster and
its subagent's last row, with the parent inside its own-row limit, is 3,005,
so no observed subagent loses its roster to the lower cap.
* fix(native-chat): a subagent section opens only while its roster is the running scope's live frontier
A section used to open whenever its roster said the subagent was working, anywhere
in the transcript and whether or not the session was running, so a background
subagent's section stayed open and grew mid-transcript while the parent moved on.
It now opens by default only while the session runs and the roster row naming the
subagent is the newest thing the parent produced, user rows aside. Newer parent
output closes it even while the subagent still works; the roster row keeps
showing that live state. A subagent still working is a running scope of its own
for the sections it spawned; a settled one closes its scope. Derived every
render, no latch; the reader's own open or close still wins.
* fix(native-chat): name a subagent's section from a client roster the window never trims
A section took its name and state from a roster row in the loaded window. Once a
burst trimmed that row, or the row sat on an older page, the section fell back to
an unnamed, closed "Subagent" header.
The shared reducer now keeps a roster keyed by agent id, folded from every roster
row and revision the client receives: pages, older pages and live batches,
including revisions of roster rows outside the window, which live batches already
carry. The first roster naming an agent wins and its revisions update it; a
removed roster row drops its entries; it is rebuilt on every page that replaces
the window and bounded to 512 agents. Sections take their name, state and
live-frontier place from it; placement stays under the loaded roster row, else
at the section's first loaded row. Only a subagent no roster ever named stays
unnamed.
* feat(agent-session): a history page names the subagents whose roster row is older than it
A page is a contiguous run of the journal whose older-page cursor is its first
item, so it cannot pull an older roster row in without skipping the rows between.
When a page held a subagent's rows but not the roster row naming it (about 11% of
the moments a reader could open a session on local journals), that subagent drew
as an unnamed "Subagent" header.
History and hydration pages now carry an optional `subagentRoster`: the first
roster entry naming each subagent whose rows are on the page and whose roster row
is not, with the row's id, sequence and revision; bounded to 64 entries and
16 KB. Items and cursor are unchanged. The client seeds its roster from it.
Rule 1 in docs/reference/remote-wire-compatibility.md: an optional field on an
existing frame, no capability gate. An older client ignores it (the released
reducer reads a page with it exactly as one without); against an older host the
field is absent and the section falls back to an unnamed header.
* Revert "fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it"
This reverts commit
|
||
|
|
e8e144bf3c |
fix(native-chat): a subagent's words are presented as that subagent's, never the parent's (#23605)
* fix(native-chat): a subagent's words are presented as that subagent's, never the parent's The journal already names the agent that produced every row, but the transcript projection dropped it, so a subagent's prose rendered as the parent's reply, its tool calls folded into the parent's runs, and a settled turn could fold down to a subagent's words as its only visible answer. The transcript message now keeps the row's producer. The fold keeps each agent's calls in that agent's own run, a turn's answer is the session's own agent's last prose, and a subagent's row names the subagent on desktop, mobile and a worker's transcript text. * test(native-chat): give the window fixture's slot the attribution field it now carries * fix(mobile): read the subagent label the row is given, and pin the caption * fix(native-chat): keep interleaved agents in order and each agent's own run live Review follow-ups: - the fold is main's adjacency fold plus one condition: a row never folds into another agent's run, so an agent's later call stays below its subagent's work instead of jumping back into its earlier row - each agent has its own live frontier, so a parent still inside its spawn call reads as running while its subagent works below it - mobile names no one on a row whose only content is hidden behind its settled turn - a pending question from a subagent keeps its producer - worker reads serve only the producing agent's id, bounded like the roster key that names it, and drop the provenance fields - the single-message worker formatter is private, so no caller can drop names |
||
|
|
7a1f55c52a |
fix(native-chat): give a failed Claude background task a typed row instead of an opcode (#20519)
* fix(native-chat): give a failed Claude background task a typed row instead of an opcode
A failed backgrounded command printed red rows whose visible text was the wire
opcode, and printed one failure twice. All five task lifecycle kinds are
catalogued status-chrome, but the payload sniffer in classifyProviderFrame runs
first and promotes any frame reporting a failure to the generic unknown-frame
fallback, whose sentence lookup has no key for the field Claude puts its own
sentence in. Two frames for one task therefore produced two rows, both of them
the method name.
Suppressing those frames is not the fix: when the last background task settles
the tracker flushes it and the strip unmounts, local_bash is excluded from the
subagent roster, and the status feed publishes only live tasks, so for a lone
backgrounded command the transcript row is the only report of the failure that
exists anywhere.
So the catalogue now binds: kinds a dedicated typed translator owns are named
as covered, and the generic fallback refuses to emit for them in either
direction. A new row owner keeps one durable row per task id, opened by the
announcement, revised in place by the lifecycle frames and closed by the
notification, carrying the provider's summary, error, output path, usage and a
run state. The row is written on the same dual carrier the subagent roster
uses: a frozen text twin plus a typed block, so a client without the block type
reads the sentence rather than nothing.
hasProviderError keeps its authority everywhere else, unchanged.
* fix(native-chat): keep tool attribution across a background-task row
A background task's row is a system message landing mid-turn between the
assistant's tool calls, exactly where the spawn-group roster row lands. Without
the same exemption it ended the run the following tool messages fold into, so a
tool result arriving after one stopped folding into its own assistant turn.
Also syncs the catalog with the row's one new translate key.
* fix(native-chat): harden background task rows
* fix(native-chat): settle background rows on provider end
* fix(native-chat): scope malformed task fallback text
* test(native-chat): assert only eligibility at the disposition layer
The malformed-task-frame test asserted the generic fallback resolves Claude's
`summary` field itself, which was true only while that key sat in the shared
key list. Eligibility is what this layer decides; the sentence the row leads
with is Claude's, supplied through the display-text seam and proven in the
translation test.
* fix(native-chat): gate background-task admission and scope rows per run
Admission now matches the reference on all three gates. Type is the whole gate
and MONITORS ARE NOT ADMITTED: a monitor runs for the life of the session and
has no outcome a row could report, so it never reaches the timeline. On first
admission only, the task's tool_use_id must name a tool call this session
forwarded at the TOP level — a Task spawned inside a subagent's sidechain names
an id that never reached the transcript, and a top-level row for it would claim
an invocation the user never saw. And a task that already exists and has not
finished is not re-opened: a duplicate announcement is a redelivery, not a
second run.
Rows are now keyed per RUN. A provider may reuse a task id for a distinct later
invocation, and a row keyed by the id alone overwrote the first run's transcript
history instead of leaving it standing. Generation 1 keeps the bare key, so
every row already written is unaffected.
The spawning tool call is carried on the row as parentToolUseId. Orca's journal
has no structural parent link for an item — AgentJournalItemIdentity has four
arms and none carries one — so the relationship is data on the item rather than
nesting.
A terminal frame that names NO tool still opens a row. That is a named
deviation, recorded at its call site, and the measurement behind it is in the PR.
* fix(native-chat): read the aggregate roster by membership, not a phantom status
The background-tasks payload types every entry as exactly
{task_id, task_type, description, ambient?}. It has no per-entry status, so the
state this owner derived from one was always undefined and the reopen branch it
guarded was unreachable on every real payload — proven by deriving the state
from an SDK-shaped entry and getting null.
Membership is the only liveness the payload carries: it is the whole live set
after a change, so presence means live and absence means merely "no longer
listed", never an outcome. Presence does not revive a settled row either — the
level's ordering against the start/stop edges is unspecified and it carries no
evidence of a new run, so the task's own frames stay the only thing that opens
or settles one. Only the identity fields it really sends are read, and ambient
housekeeping entries are excluded as the payload asks.
The two helpers that served the dead branch are removed, along with the test
that exercised it through a synthetic status the CLI cannot send.
* fix(native-chat): mirror reference task admission and drop the synthesis path
The forwarded-parent gate is conditional on the field being PRESENT. An
announcement naming a tool this session never forwarded is a nested child and is
refused; one naming no tool at all is admitted, because absence of the field is
not evidence of an unforwarded parent. The previous rule required the field and
so refused every tool-less task.
Terminal frames now match on task_id alone. The forwarded-parent question is
settled once, at admission, and is never re-asked on a notification or a patch.
A frame for a task that was never admitted yields no row, and a patch is folded
into the row it names rather than opening one.
That removes the synthesized-row path entirely, and with it the named deviation
it carried: the captured tool-less failure lands on a row that already exists,
because its own tool-less announcement is admitted. The dead builders go with
it.
Left deliberately stricter than the reference, and flagged rather than changed:
a terminal frame still records its task id as terminal even for a task never
admitted, so a late announcement cannot open a row for work already reported
finished. Two existing tests pin that.
* fix(native-chat): preserve background task ownership across restarts
* chore: restore pnpm-lock.yaml to origin/main
A local pnpm run rewrote the lockfile and the merge commit swept it in. The
branch changes no dependencies, so it must carry no lockfile delta at all.
* fix(claude): harden background task lifecycle
* fix(claude): bound task generation history
* fix(claude): preserve task identity after history eviction
* fix(claude): resolve background task identity after rebind
* fix(claude): isolate queued task runs
* refactor(claude): give the background-task ledgers one bounded owner
The bounded collections behind a background-task row were read out of the
class with `Reflect.get` to prove they stay capped, which the anti-slop
gate rejects. Move them into `ClaudeBackgroundTaskLedgers`, which owns
the caps beside the eviction helpers and reports a typed readonly size
view the tests assert against.
Also replace a `Reflect.get` in the mobile recording proxy with typed
property access.
* fix(native-chat): report a background task failure the transcript never admitted
A terminal `task_notification` for a task no announcement ever admitted rendered
nothing at all. The typed row owner declined the row because its map held no
entry for the id, and reported the frame as handled — which is exactly what
tells the generic provider-frame fallback to stay quiet. Both surfaces declined
the same frame, so a real failed background task was dropped on the floor.
A terminal frame is self-sufficient: it states an outcome, and it carries the
summary, status, error, output path and usage that outcome needs. It now opens
its own row from those fields, with the summary as the label and `unknown` as
the kind when the frame names no task type. The row map enriches a terminal
frame; it never gates one. Every terminal status writes one, not failures alone,
so there is one rule here rather than a third behaviour for failures.
The deliberate hand-offs still win, because they are recorded rather than
implied: ambient, subagent and foreground tasks are claimed in the foreign-owner
ledger the notification path already checks first. A Task spawned inside a
subagent's sidechain now records its refusal there too, under `sidechain`,
instead of leaving no trace and reading as a task nothing ever decided about. A
capacity-refused task whose outcome the generic fallback already printed records
`fallback` the same way, so a redelivery neither prints twice nor mints the row
capacity refused.
The anti-resurrection guard stays and stays scoped to announcements: a late
`task_started` cannot reopen work already reported finished. The restart rule is
now stated once instead of twice — a different parent alias is the provider's
restart signal only when BOTH runs name their parent, which is what the terminal
ledger already required of an evicted row and what the live row now requires too.
* test(native-chat): pin that an orphan task row reopens the provider's turn
* fix(native-chat): stop an orphan task row drawing its sentence twice
An orphan row took its header label from the notification's summary, which
also renders as the row's sentence, so the same string appeared in both slots
of the same collapsed row. The label now stays empty and the header falls back
to the task kind, leaving the sentence to carry the provider's words.
* fix(native-chat): keep Claude task outcomes owned through capacity and redelivery
* fix(native-chat): keep settled overflow notifications from reopening a turn
* fix(native-chat): retain Claude task rows through journal pressure
|