mirror of
https://github.com/stablyai/orca.git
synced 2026-10-08 16:02:37 +00:00
7faa9f7cd3bda68189646b1e9def6cc8aae3b1cb
39
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7faa9f7cd3 |
fix(native-chat): reasoning rows with a real open/finished state, and readable Claude thinking (#19221)
* fix(native-chat): render structured reasoning as collapsible messages
* fix(native-chat): align expanded reasoning with summary
* fix(native-chat): place reasoning chevron after summary
* De-emphasize reasoning headlines with observed timing labels
* Exclude later turn work from observed reasoning duration
* fix(native-chat): give reasoning rows a host-owned open/closed lifecycle
A reasoning row now says whether its block is still streaming (`state`)
and, when the host saw it end, when (`completedAt`). The row's own
start is its first-write time, so "Thought for N s" is measured on the
execution host instead of by whichever window happened to be watching.
Rows open with their first non-empty text and close on every path that
ends them: the block's final frame, a new message in the same stream,
every Claude turn end through one hook on the open turn, Codex
item/completed, turn and session settlement, active-item eviction, and
both host sweeps for a dead generation. Closing writes are lifecycle
writes so backpressure cannot leave a row open. Rows without the field
(older hosts, older journals) never read as live.
The reasoning row's headline reads the row's lifecycle and its own
turn's liveness: Thinking while open in a running turn, then Thought
for N s, Thought when no span was seen, Reasoning when the host kept no
lifecycle.
* feat(native-chat): ask Claude for readable thinking summaries
Under Orca's launch the Claude CLI streams thinking blocks with empty
text, so no reasoning row ever had anything to show. Pass only
`--thinking-display summarized`: it fills thinking blocks with the
API's summaries without turning thinking on, so a user who disabled
thinking keeps it off.
Orca runs the user's own binary, and a CLI older than the flag exits on
it before the session starts. The launch probes the binary it is about
to run, overlapping the rest of launch resolution, and passes the flag
only when that probe has already answered with 2.1.94 or newer. A slow
or failed probe never delays a launch and is never remembered; a
successful one is kept per binary until the binary changes.
* test(native-chat): type the superseding send in the reasoning lifecycle test
* fix(native-chat): stop reading an ended reasoning row as thinking
The spinner line infers "Thinking" from the newest root row being a
reasoning message. With summaries on, a closed reasoning row stays the
newest row while Claude streams a tool's input, so the line read
Thinking for the whole Write. A reasoning row now counts only while its
own state is running; a row from a host that keeps no state reads as
before.
* fix(native-chat): measure reasoning from the block's start, not its first text
A row opens with its first summary text, which trails the thinking
block's start by seconds, and the journal stamps a row with its queued
append time. So "Thought for N s" read 1 s for blocks that ran 4.65 s
and 12.95 s. The block registry now records the host time at each
block's content_block_start (else its first delta), and every write of
that block's row carries it as the row's observed time; the close keeps
the frame's own receipt time. Codex reasoning rows likewise carry their
item/started time.
* perf(native-chat): stop rewriting the reasoning row for every thinking token
Claude sends a thinking_tokens frame after every thinking delta. Every
frame the stream path did not consume forced the streamed text out
first, bypassing the checkpoint widening and the coalescing window, so
a single thinking block was rewritten and republished once per token:
212 full-row writes and 258 KB for one captured block. A frame that can
write no row (token tallies, stream deltas no registry carries, pings)
no longer forces that flush; every frame that can write still does.
The same block now takes 14 writes.
* fix(native-chat): time Codex reasoning by receipt and close evicted rows
Codex summary text lands 13-50 ms before item/completed, inside one
coalescing window, so a reasoning row's first write can be its
completion. Item boundaries are now stamped with their host receipt
time the way turn boundaries already were, so a retried or buffered
delivery keeps it, and the item translator falls back to the host
clock rather than Date.now(). The row carries its item/started time on
every first write, including the completion, so its span is
item/started to item/completed.
An evicted active item is now closed from the text streamed so far,
like both settle paths, and the eviction runs before the incoming item
is tracked: tracking first let the stream bound drop the evictee's text
before the eviction could close its row, stranding it running.
completedAt now has one meaning everywhere: the host time the message
was seen to end, or the end of the turn or stream that cut it off;
absent only when no end was seen live.
* refactor(native-chat): keep the Codex streaming body translation pure
The streaming translation preserves a reasoning body as-is again; the
stream writer, which is what knows the item has not completed, stamps
it running.
* test(native-chat): keep a re-collapsed reasoning row collapsed through a revision
* fix(native-chat): estimate a collapsed reasoning row as its trigger
A reasoning row renders collapsed, as one small button, but its height
was estimated from its full text: a 4,129-character summary reserved
about 950 px for a 24 px row, so long chats jumped as rows were
measured. It is now estimated at the trigger's height; opening the row
remeasures it.
* fix(native-chat): probe the CLI the launch will run, and learn from a refusal
The thinking-display gate probed `claude --version` with Orca's own cwd
and env, while the launch spawns with the workspace's cwd and the shell
env. Behind a version manager's shim those can pick different CLIs, so
the probe could approve a CLI the launch never ran, and an older CLI
exits on the unknown flag before the session starts. The probe also
only counted if it had already finished when resolution did, so a
first launch, or the first after a CLI update, usually went without
the flag.
The probe now runs with the launch's own resolved cwd and env, is keyed
by the binary and the workspace, and a launch waits up to 200 ms for it
(about 3x the probe's measured p95) before going without the flag. Only
answers are kept, so a slow or failed probe is asked again next launch.
A child that exits with commander's "unknown option '--thinking-display'"
marks that binary in that workspace so the next launch skips the flag;
that one start fails exactly as any CLI startup failure does today.
* test(native-chat): type the thinking-display probe mock with both of its parameters
* fix(native-chat): recheck the account switch after the probe, last as before
Moving the invocation ahead of the probe, the transcript check and the
permission mode put its account-switch recheck before those awaits, so a
switch that began during them launched unchecked. The invocation is the
last await again. The probe gets its own env from the same sources the
launch uses, the inherited env and the overlay with the CLI's runtime on
PATH, built by the same code, minus every credential: asking a CLI its
version needs none.
* fix(native-chat): wait up to 1.5 s for a cold probe, and remember every outcome
200 ms only covered a warm CLI; a cold disk, a node install or an
antivirus scan exceeds it, and that chat's child then ran its whole
life without summaries. A launch now waits up to 1.5 s, once per binary
per workspace, measured from when that binary's probe began, so a later
launch never waits again on a probe already past it. The probe gets its
own 10 s kill timeout, and every outcome, including no version printed,
a failure or a kill, is kept for the binary's life, so a probe that
hangs costs one launch rather than every one. A refusal seen while a
probe still runs wins over its late answer.
* fix(native-chat): leave Thinking to the activity line while reasoning runs
A reasoning row still being written drew a pulsing "Thinking…" header
right under the turn's activity line, which already says Thinking: two
live indicators for one fact. In a running turn an open reasoning row
now draws nothing and reserves no height; it appears when it closes, as
"Thought for N s". A row from a host that keeps no state, a closed row,
and a row left open by a turn that ended draw as before. The row's
Thinking headline is gone with its catalog key.
* fix(native-chat): end every unfinished Codex item through one rule
A reasoning completion with no text of its own left the row its stream
wrote running for good: the completion translated to nothing and the
item left the active set, so no settle could reach it. It now closes
from the text streamed so far.
Settlement, eviction and that completion now build an unfinished item's
row through one choice (the streamed text when there is any, else the
item as it started) and end it through one rule. An evicted file change
with streamed tool output no longer keeps that output as its patch; it
reads as interrupted, as a settled one does. A completion whose start
was never recorded claims no span, so it reads "Thought".
* fix(native-chat): catalogue Claude's stream keep-alive as benign
An uncatalogued `ping` stream frame classified as substantive, so the
fallback wrote a visible "claude · message:stream_event:ping" row, and
since such frames no longer force streamed text out first, a ping
inside a coalescing window landed above the open reasoning row. A ping
is now benign: it writes no row.
* fix(native-chat): bound finding the binary by the probe budget, and keep it LRU
Resolving the command's real path and its mtime was awaited before the
budgeted wait, so a slow filesystem could hold a launch indefinitely;
it now counts against the same budget, and running out caches nothing.
The cache is least recently used rather than first written, and holds
32 binary-and-workspace entries rather than 16.
* feat(mobile): collapse reasoning rows the way desktop does
With summaries on, every Claude turn now carries reasoning text, and the
phone drew all of it inline, dimmed, between the prompt and the answer.
Mobile now draws a reasoning row as desktop does: collapsed to "Thought
for N s", "Thought" or "Reasoning", its text mounted only once opened,
and nothing at all while the row is still being written in the live
turn or has no text. The headline and the visibility rule live in one
shared module both clients read, so they cannot drift.
* fix(native-chat): let a failed CLI probe heal instead of latching
A probe killed at its timeout, failing to spawn under a loaded boot, or
printing no version was cached as "no flag" for the binary's life, so
that workspace never got summaries again in that run. Only a version
(either side of the floor) or the CLI's own refusal is kept for good
now; a probe that gave no version is kept for 10 minutes, so a hung CLI
still costs one wait per stretch and a boot-time failure heals.
* docs(native-chat): say exactly what the version probe's env leaves out
* fix(native-chat): record a Codex item's start whatever its first frame carried
The start was recorded only for an item/started that wrote no row, so a
reasoning item that started with text lost it and its completion
claimed no span. Every tracked item/started now records its receipt
time, and the started write carries it too.
* fix(mobile): label the reasoning toggle and give it a full touch target
The toggle now tells a screen reader what it is, "Reasoning: Thought
for 3s", as desktop's prefix does, and reaches a 44 pt target. The
shared English copy stays private to the module that formats it.
* refactor(native-chat): build reasoning rows from one provider-neutral helper
Claude, Codex and the terminal sweeps each built the reasoning row body and
its running/ended stamp themselves. They now share journal-reasoning-row:
blank text journals no row, text is bounded the same way, and an end carries
completedAt only when the host saw it.
* refactor(codex): move the active journal item type into the contracts file
codex-unfinished-item-body imported the type from the settlement module,
which imports values from it.
* test(claude): read the launch PATH the way Windows spells it
* test(claude): compare the probe's PATH to the launch's without Orca's CLI dir
When the CLI's directory also holds node (Linux CI's /usr/local/bin), the
runtime pairing puts that directory first, ahead of the Orca CLI directory the
launch adds, so the launch PATH no longer ends with the probe's. Both still
resolve the same claude and shims. The test now checks that exactly, for a CLI
with and without a sibling node.
* feat(native-chat): lead the reasoning row with a brain glyph in the tool-row column
* fix(native-chat): route the reasoning glyph through the shared icon names, keep its chevron findable, and match it on mobile
* fix(mobile): keep the long-press actions sheet on reasoning rows for Android
* fix(native-chat): forward every exit argument through the thinking-display connection wrapper
* refactor(codex): keep the receipt-timed notification methods with the event they stamp
* feat(native-chat): read an open reasoning block through the one live "Thinking" line
While the agent's open reasoning block has text, the turn's live activity line is its
disclosure: collapsed by default, expandable to the live text (capped and scrollable), and
the block's row draws nothing meanwhile. When the block ends, its row appears in place,
open if the reader opened it live, because the line and the row read one disclosure key.
Which block the line discloses is derived from the line's own render condition, so a row
is never hidden while nothing on screen shows it; any other open block (a subagent's, or
one a prompt pushed off the line) draws as "Reasoning". Desktop and mobile alike; no host
or wire change.
* fix(native-chat): a slot kept for its turn bar or diff rollup no longer draws its message
The transcript row drew the message of every message slot, while the slot builder pushes a slot
for a row it declined to draw whenever that row also carries its turn's bar or diff rollup. So the
open reasoning block the live line discloses still drew as a "Reasoning" row when it was a
provider-opened turn's first row or the last row of a turn that changed files, and one click
opened both. The builder's decision now travels on the slot (`drawsMessage`) and the row draws
only the bar and rollup when it is false; the row-level `folded` guard it made redundant is gone.
* refactor(native-chat): draw the live line from one shared value, with one live region
Desktop and mobile now render the live activity line from one pure function,
`nativeChatLiveLine`: whether it draws, what it says, and the open reasoning block it
discloses with the text it has so far. The open block's row is hidden from that same value,
so desktop no longer restates the line's render condition beside it, and the lines no longer
re-derive the text.
The line keeps one element, and so one live region, through every state; only its trigger
and body come and go, so a screen reader hears "Thinking" and the label after it. On mobile
the live text gets the finished row's Android long press (copy or select through the message
actions sheet), the line's touch target is the row's 44 pt, and its label and body sit in
the finished row's column so nothing moves when the row takes over.
* fix(mobile): keep the live text's actions sheet on the block it was opened for
On Android the sheet opened by a long press on the live reasoning was a flag gated on a live
block: it vanished when the block ended, mid Select text, and the stale flag reopened it
unprompted on the next block. The sheet now holds the message it was opened for.
* test(native-chat): the reasoning body owns its tone, live and once landed
Rendered QA on a pre-merge build showed the live line's open reasoning in full foreground and the
landed row's in muted text, so it dimmed as the block landed: the body set no colour of its own and
inherited one from wherever it was mounted. Since the main merge (
|
||
|
|
9905765e3b |
feat(native-chat): open structured chat's wire and stored records to registered agents, behind a negotiated capability (#25159)
* refactor(native-chat): keep the provider resume handle opaque to shared code
Shared structured-chat code parsed each provider's resume handle: Claude's
session id and branch leaf, Codex's thread id, through a 'claude' | 'codex'
union every new agent had to widen. The in-memory handle is now
{ transport, agent, nativeId, providerData? }: shared readers use nativeId,
lease and handle-chain checks compare transport and agent, and only the
Claude adapter reads its leaf (providerData).
Stored and wire forms are unchanged for Claude and Codex. One encoding
module writes their typed shapes and decodes both those and the neutral
shape a new transport uses, which an older build refuses as unreadable
rather than reading as Codex. Key and root strings, which fork seeds,
superseded creations and resume offers persist, stay byte-identical.
The journal's own handle type becomes the journal-row and attach-wire
encoding of the same handle, and the journal identity carries the
neutral handle (null before the provider proves one).
No user-visible change.
* fix(native-chat): derive journal-row provider handles from the journal identity
The journal row converter now takes the identity every caller already holds,
so a row's handle has one obvious constructor. Tests that wrote the in-memory
handle straight into journal rows now build it through that converter, and the
processless Claude fixture names a not-yet-proved handle as null.
* fix(native-chat): refuse a stored provider handle written in both forms
A typed Claude or Codex handle that also carries the neutral form's
transport, agent, native id or provider data named two identities; it
was read as Claude or Codex and the next write dropped the other one.
Such a row now stays unreadable and is set aside untouched.
* refactor(native-chat): route structured agents through registered definitions
The structured-chat shared layer named Claude and Codex in a closed router
type, in option reads, in frame classification, and in renderer label and
validation branches, so adding an agent meant editing shared types. Each
agent now has a definition (capabilities, resting option rules) owned by its
own module; the router is a registry of definition + adapter pairs built at
runtime setup, and shared code reads the definition instead of the name.
No behavior change for Claude or Codex; no wire or stored shape change.
* refactor(native-chat): name the structured agent list once in host types
The host types, conversation commands, provider-session ownership and the
runtime's create path spelled the structured agent list inline; they now use
the one shared type the record and wire already validate against.
* fix(native-chat): narrow the record before reading its agent's option rules
* refactor(native-chat): make the router's registrations the only agent definition lookup
The at-rest option reads looked definitions up in a second, separately populated map, so a
registered definition could route and declare capabilities while the static map decided which
options a chat at rest accepted. The router now answers `definition(agent)` from its
registrations, the at-rest readers take it through the host's adapter, and the built-in
definitions are listed only where the runtime composes the router.
* feat(native-chat): open the structured-chat wire and stored records to registered agents
A host's structured agents are the ones its runtime registered. Records, RPC
params, persisted tabs and the model catalog accept any registered agent instead
of naming Claude and Codex; each agent's definition declares the transport its
handles live in and the variable its account home pins. A new runtime
capability, agent-session.structured.registered-agents.v1, advertises that a
host accepts and lists its agents (agentSession.agents, with each agent's
capability record), and the host withholds any other agent's tabs and restart
offers from clients that do not advertise it.
* test(native-chat): cover registered agents on the wire, in storage and across versions
* refactor(native-chat): let the record store decide which agents' tabs exist
* test(native-chat): declare the pilot test agent's storage
* test(native-chat): read the old build's saved tabs through a parsed shape
* fix(native-chat): act on restart offers only for agents the calling client can show
A paired client too old to show an agent's chat was listed only the offers it could show, but
dismissing or continuing all reached every offer on the host, and a named continuation answered
with the host's whole remaining inventory. The client's audience now goes to the host with every
restart operation: only offers it sees are reserved, dismissed or returned. Without an audience
(this host's own process, or a client that shows every agent) nothing changes.
* test(native-chat): read an agent-registering baseline's storage on its own terms
The registered-agents downgrade test assumed its baseline release predates registered agents: it
expected the saved-tab parser to erase an unknown agent and called the record reader without the
agents list. Once a release with this change becomes the baseline, both break. The expectations now
follow what the baseline host advertises, and an agent-registering baseline is handed its own
Claude and Codex storage.
* refactor(native-chat): derive record-store admission from the runtime's agent registrations
Which agents a stored record may name and which agents the router drives came from two lists in
the runtime, so a newly registered agent could be routed while its records were set aside. One
list of registrations now holds each agent's definition and the factory for its adapter: the
store's admitted agents are derived from it before the store opens, and the adapters are built
from it once it has.
* fix(native-chat): hand the exit drain a promise for every registered agent
* fix(native-chat): let the adoption conflict check read any agent's ownership
Ownership rows name any registered agent since the stored records opened to them; the adoption
check compares by agent, so it takes the same open id. Only Claude and Codex still adopt.
* test(native-chat): use opaque handle in queued rejection fixture
* test(native-chat): share one Codex journal identity in the integration suite
Main grew the suite to the 800-line limit; the opaque-handle import pushed it
over. The two tests built the same identity inline.
* refactor(agent-session): name the handle's adapter state resumeCursor
Rename the neutral provider handle's providerData to resumeCursor before any
row persists the neutral form: it is an adapter-owned resume position (Claude's
transcript leaf), never identity. Claude/Codex stored and wire bytes are
unchanged; their typed shapes never carried the field.
State the stored-form contract (a handle's field set is closed; later per-link
data goes on the chain link, which every build preserves) and pin it with a
record round-trip test. Document that transport records the id space the
native id was minted in, which can differ from the agent's current transport.
* refactor(agent-session): one required agent registry; declarations admit what they claim
A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.
/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.
Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).
* refactor(agent-session): the router applies the declared rewind itself
The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.
* test(agent-session): register the agents the merged-in tests now need
The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.
* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record
* fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop
The desktop renderer calls the restart methods as a paired runtime client
without the registered-agents capability, so it was handed a Claude/Codex
audience and its "dismiss all" took the scoped path: no dismissal fence and
no clearing of unwritten teardown witnesses, so a late teardown write could
bring a dismissed offer back.
The audience is now derived against this host's registry: a client that can
show every agent the host registered gets none, and dismisses exactly as
before (one fence for every offer). A client that truly cannot show some
agent still gets a final dismissal for what it sees: each cleared session is
fenced on its own (bounded, superseded by a dismiss-all fence) and only the
witnesses of agents it sees are dropped.
One predicate now answers which agents a client renders, for both tab
projection and restart offers; the Claude capability rule is folded in.
* fix(agent-session): a changed agent definition never hides that agent's chats
A record was readable only if every handle used the transport its agent's
definition declares today, and its account variable was declared by some
registered agent. A later build that drives the same agent over another
protocol, or under a renamed variable, would have set every existing chat of
that agent aside: no tab, no history, no error.
Readability now asks only that the record agree with itself: its agent is
registered, every handle shares one namespace owned by that agent, and the
account variable is a well-formed name. Whether this build can drive it (the
chain's transport is the one the agent speaks, and the account variable is
that agent's own, since it becomes the child's environment) is decided in
performAttach, the one admission every start of every agent passes through,
and refused there as hostUnsupported. A record pinning another agent's
variable is refused the same way rather than passing because some agent
declares it.
* refactor(agent-session): each agent's registration says where it runs and which account it pins
createSupport, create and the model catalog's account read still chose a
location rule and an account resolver by name (Claude, Codex, else none),
so a new agent would have needed a third branch beside the registration
list that already decides storage, routing and publication.
Each runtime registration now carries supportsLocation(location) and
resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }),
and Claude's and Codex's rules move into their entries unchanged: a read
still syncs no home and starts no bridge, a Codex launch still trusts the
workspace first, and WSL locations resolve as before. The list exists
before the host is built, so none of these reads installs the host, and an
agent the list does not hold is answered no without opening the journal.
* test(agent-session): drop a duplicate registry key; keep fences out of the capsule state
A spread input already carries the registry, so the explicit key was
overwritten (TS2783). The session-fence helper now returns only the fence,
not the capsule state it was given.
* fix(agent-session): a scoped dismiss-all persists no per-session fence
The per-session dismissal fence added a forever-persisted capsule field
with an arbitrary cap that served no caller: only the local desktop calls
the restart methods, and on a host whose agents it all shows it gets no
audience and takes the unscoped, fenced dismissal. A scoped dismiss-all
(a caller that cannot show some registered agent) keeps that agent's
offers, clears only its audience's unwritten witnesses, and writes no
fence; a late write from this process is serialized behind it. The field
never left this unreleased branch.
* fix(agent-session): refuse an attach whose agent is not the session's own
The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal.
* fix(agent-session): offer to start a chat only when the start would accept it
The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired.
* fix(agent-session): the model-catalog probe uses a record's account only if this build drives it
A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record.
* refactor(agent-session): the record store admits agent ids; comments say where transport is checked
The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept.
* docs(agent-session): the record store admits the registered agents' ids
* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer
Uses an audience production sends (one that cannot show every agent), per review.
* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge
* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way
The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.
* test(ratchet): require src/main/provider-process now that it has landed
* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test
Main's #25706 and this branch both added the import at different lines; the merge kept both.
* Keep saved chats readable without provider registration
* Keep stored-record compatibility checks independent of registration
* Keep saved providers in restart client audiences
* test(native-chat): type reveal fixtures without assertions
* test(wire): expose known agents in structured host fixture
* Supply startability dependency in the new Codex catalog fixture
|
||
|
|
3a03441580 |
refactor(native-chat): structured agents declare their capabilities instead of shared code naming Claude and Codex (#25076)
* refactor(native-chat): keep the provider resume handle opaque to shared code
Shared structured-chat code parsed each provider's resume handle: Claude's
session id and branch leaf, Codex's thread id, through a 'claude' | 'codex'
union every new agent had to widen. The in-memory handle is now
{ transport, agent, nativeId, providerData? }: shared readers use nativeId,
lease and handle-chain checks compare transport and agent, and only the
Claude adapter reads its leaf (providerData).
Stored and wire forms are unchanged for Claude and Codex. One encoding
module writes their typed shapes and decodes both those and the neutral
shape a new transport uses, which an older build refuses as unreadable
rather than reading as Codex. Key and root strings, which fork seeds,
superseded creations and resume offers persist, stay byte-identical.
The journal's own handle type becomes the journal-row and attach-wire
encoding of the same handle, and the journal identity carries the
neutral handle (null before the provider proves one).
No user-visible change.
* fix(native-chat): derive journal-row provider handles from the journal identity
The journal row converter now takes the identity every caller already holds,
so a row's handle has one obvious constructor. Tests that wrote the in-memory
handle straight into journal rows now build it through that converter, and the
processless Claude fixture names a not-yet-proved handle as null.
* fix(native-chat): refuse a stored provider handle written in both forms
A typed Claude or Codex handle that also carries the neutral form's
transport, agent, native id or provider data named two identities; it
was read as Claude or Codex and the next write dropped the other one.
Such a row now stays unreadable and is set aside untouched.
* refactor(native-chat): route structured agents through registered definitions
The structured-chat shared layer named Claude and Codex in a closed router
type, in option reads, in frame classification, and in renderer label and
validation branches, so adding an agent meant editing shared types. Each
agent now has a definition (capabilities, resting option rules) owned by its
own module; the router is a registry of definition + adapter pairs built at
runtime setup, and shared code reads the definition instead of the name.
No behavior change for Claude or Codex; no wire or stored shape change.
* refactor(native-chat): name the structured agent list once in host types
The host types, conversation commands, provider-session ownership and the
runtime's create path spelled the structured agent list inline; they now use
the one shared type the record and wire already validate against.
* fix(native-chat): narrow the record before reading its agent's option rules
* refactor(native-chat): make the router's registrations the only agent definition lookup
The at-rest option reads looked definitions up in a second, separately populated map, so a
registered definition could route and declare capabilities while the static map decided which
options a chat at rest accepted. The router now answers `definition(agent)` from its
registrations, the at-rest readers take it through the host's adapter, and the built-in
definitions are listed only where the runtime composes the router.
* test(native-chat): use opaque handle in queued rejection fixture
* test(native-chat): share one Codex journal identity in the integration suite
Main grew the suite to the 800-line limit; the opaque-handle import pushed it
over. The two tests built the same identity inline.
* refactor(agent-session): name the handle's adapter state resumeCursor
Rename the neutral provider handle's providerData to resumeCursor before any
row persists the neutral form: it is an adapter-owned resume position (Claude's
transcript leaf), never identity. Claude/Codex stored and wire bytes are
unchanged; their typed shapes never carried the field.
State the stored-form contract (a handle's field set is closed; later per-link
data goes on the chain link, which every build preserves) and pin it with a
record round-trip test. Document that transport records the id space the
native id was minted in, which can differ from the agent's current transport.
* refactor(agent-session): one required agent registry; declarations admit what they claim
A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.
/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.
Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).
* refactor(agent-session): the router applies the declared rewind itself
The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.
* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record
* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way
The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.
* test(ratchet): require src/main/provider-process now that it has landed
* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test
Main's #25706 and this branch both added the import at different lines; the merge kept both.
|
||
|
|
57fedeed79 |
fix(native-chat): the agent's exit ends its record, and an unconfirmed stop is joined instead of held (#24862)
* fix(native-chat): a child's root exit is reported even during its close, and bookkeeping after it never reads as unproven - Both connections report the root process's exit once, with `expected` set when a close had begun. A close that came back unproven and whose root exits later is finished by the adapter, and its end reaches the host like any other. - A Claude close whose resume-point write fails after the exit was proven, and a Codex close whose terminal row is refused, now end the session and report the failure, instead of keeping a dead child indexed as if its exit were unproven. - A Codex close whose forced tree kill can't prove the descendants gone but saw the root exit reports the descendants and counts the root exit. - Every child exit with an identity, expected or not, is forwarded to the host. * fix(native-chat): the exit ends the child's record; an unfinished stop is the child's own close, which everyone joins - The host keeps no stored "stop still owed" record any more. A stop begins the child's close (`child.close`), which lives on the child and ends with it. A second Stop, the idle reaper, quit, a send and an option/answer/goal/rewind all join that close instead of retrying a separate obligation. - A caller waits on the close only as long as the step deadline; the close itself is never abandoned. A proof that lands after every caller stopped waiting reaches the host as the adapter's report of that exit, which ends the record through the same handler. - Once the exit is proven, draining, settling, the lease release and the adapter's acknowledgement are each attempted and reported on failure; none keeps the child on record. A start, and the handle's close, write a release that failed from this host's proof of that exit, so a failed write never refuses a send. - A start that meets a close still unverifiable is refused with `previousExitUnverifiable`, so the queued message is rejected with a send-again reason; nothing is held and nothing starts beside the old process. - The idle sweep goes back to idle reaping only. - Removes #24333's retry entry points, the wait row and its hold rule, the ask/failure cursors on the stored record, and the stop's own wake. Tests replace the #24333 unproven-stop test: a send joining an unproven close and an in-flight one, a late proof past the caller's bound, a root exiting after its close gave up, a proven exit whose resume-point write and lease release both failed, an unverifiable close rejecting the send and refusing an option change, a surviving descendant, quit and the idle reaper; and Codex's unverifiable, late-exit and joined-close cases. * fix(native-chat): a message refused because the old process's exit is unverifiable says so, and to send again The start failure for a refusal with reason `previousExitUnverifiable` reads "Orca couldn't confirm Claude's previous process ended. Send your message to try again." instead of "Claude couldn't restart." The status-row kind and the refusal reason stay in the shared lists for rows and hosts that still carry them; the catalogs keep one sentence for both. * fix(native-chat): a close's verdict is the root's exit alone, and what follows it is logged - A Claude close resolves as soon as the root's exit is proven: the session ends and its `ended` report goes out then. Saving the resume point runs afterwards and a failure is logged, so a slow or hung write never reads as an unproven exit or keeps a dead child on record. - A root that exits after its close came back unproven finishes that close through the same path as any close, so the session's child work is published as ended (background tasks and subagents no longer stay shown running for a dead agent), and a failure there is logged. - Codex logs a refused final row, and reports a root exit whose forced tree kill could not prove the rest of the tree gone the way Claude does, so the host logs it and blocks nothing. - Both adapters take the host's logger for this bookkeeping. * fix(native-chat): one handler ends every child's exit, and a join waits on the adapter's own close - One exit handler (`structured-agent-session-child-exit`) ends a child's record for an exit expected or not. `expected` only changes what the chat is told: the stop's cause, its end at the stop's ask, the settlement id, and no crash outcome row. The lease release keeps the exit's evidence; the handoff guard, lifecycle barrier, sink release and adapter acknowledgement apply to both. A Claude journal-sink failure ends in the same step as its stop, as Orca's own fault. - Joining a close is asking the adapter, whose close is memoized while it runs and bounded by its own kill escalation; the host keeps no attempt of its own and no 10 s caller bound. An ask after a close came back unproven runs the stop again. - A close's end is stamped where its stop was asked for (a repeated ask moves it), so the closed chat and failed start checks order a message accepted meanwhile after it. - A start refused because the old exit is unverifiable rejects what was queued in the same step. - The end of a close the host asked for no longer waits on the cross-session recovery chain. - The kill no longer waits for the stop event's write; the journal writes rows in order. * fix(native-chat): an exit's lease release lands whatever the length of its reason A crash's reason can carry kilobytes of the provider's stderr, and a lease whose death detail is over 512 characters fails the store's own check. The exit handler cut it, but the release a start or the chat handle's close re-derives did not, so after a crash whose own release failed every message was refused as not resumable until restart. The record's builder now cuts the detail to the record's bound, so no writer can hand it one too long. * fix(claude): a proven close waits at most 2 s for the output it already wrote Once the root's exit is proven, the close still waited for the SDK's output reader to end. Something outside the process tree that holds the output open would keep that close, and every send, Stop and quit joining it, waiting with no bound. The wait is now bounded; past it the close resolves as proven and the open output is logged. * fix(codex): an exit reported inside Orca's close keeps the reason Orca closed it for The connection reports the app-server's exit inside the close that ends it, so that report ended every Codex close and replaced the close's own reason (for example, a provider frame that could not be recorded) with the connection's stderr text in the ended record and the lease's exit evidence. The session now records Orca's close with its reason, and the exit it ends keeps that reason. The test connection reports its exit inside close the way the real one does. * fix(native-chat): quit stops delivery before it drains exit recovery Every exit now wakes delivery, and teardown drained exit recovery before it stopped delivery, so an exit settled in that window could start a fresh agent that teardown then killed. Teardown stops delivery first; queued messages wait for the next launch. * docs(native-chat): the unverifiable-exit refusal no longer names a caller's wait The caller's bounded wait was removed; the comment describes the close as it is now. * fix(native-chat): a stop whose kill did not take is logged, and the next ask kills again When a close's kill leaves the agent's root running, the host now logs it. Tests pin what a later ask does: each connection runs its whole stop again (Codex sends SIGKILL a second time), refuses input meanwhile, and proves the exit once the kill takes. * fix(native-chat): a start refused over the old process says Orca couldn't stop it The host reaches an unverifiable verdict only after its own kill left the agent's root running, on the machine that runs the agent, so the sentence now says that: "Orca couldn't stop {agent}'s previous process." The refusal reason, failure kind and wire shapes are unchanged. The host test also checks the failed kill is logged. * docs(native-chat): an unverifiable close verdict is a root that survived the kill The host's close runs where the agent runs, so lost contact never yields this verdict; the comment no longer says it does. * fix(native-chat): a kill that did not take is reported once, by whoever met it The log added at the close fired beside a Stop's own failure report for the same event. A stop still reports it through its failure; a send or option change refused over it now logs it at the refusal, the only place it is otherwise invisible. * test(native-chat): a second Stop joins a close the first could not prove and retries its kill * fix(native-chat): say a start refused beside an unstopped process plainly The rejection now reads "Couldn't stop {{agent}} from before. Send your message again to try once more." This kind has its own send-again step; every other failure keeps "Send your message to try again." |
||
|
|
c6cfcc034e |
refactor(native-chat): structured chat failures always reach the diagnostics log (#24312)
* refactor(native-chat): give the structured chat host one required logger The structured chat runtime took an optional onError callback that the desktop never passed, so a late dispatch settlement, an unanswered-dispatch release, a journal event-sink write and a provider lifecycle delivery that failed were dropped with no trace. Other host failures went to scattered console.warn calls, which reach nothing in a packaged desktop build. The runtime and host now take one required logger (warn/error with a scope and fields). The production logger writes each entry as a failed span to <userData>/logs/main.trace.ndjson, which the diagnostic bundle collects, and to the console (stderr under a supervised headless host). The runtime and the host wrap it so a logger that throws never fails what it reports, and the install refuses without one. Sites that deliberately kept a recovery-capsule error out of the log still log no error object. * refactor(native-chat): hand the chat host's collaborators the logger, and give orcad its trace file The delivery loop, idle sweep, queued-message drain, lease renewer, event sink, conversation map and provider start/exit settlement each took an internal error callback that the host mapped onto the logger. They now take the logger itself and log under their own scope. The event sink keeps one onFailed hook, which decides whether to stop the provider, not whether to report. The dead-generation settlement returns its failure so each caller logs it under its own scope. orcad now installs the desktop's local trace sink under its own data root, so a headless host's chat failures reach <data-root>/logs/main.trace.ndjson as well as stderr. Also passes the logger in the test fixtures the first commit missed, which tc:node caught. * fix(native-chat): keep repeated chat failures from flooding the trace file, and record their causes - The production structured-chat logger writes a repeated failure (same level, scope, session, message and error text) once per 5 minutes, carrying how many repeats it swallowed; the tracked set is capped at 256. - Trace entries now carry the error's code (and SQLite errcode) and up to three causes by name and message. - A chat read whose conversation will not open is logged through the host's logger (open-for-read), and so are the runtime's chat-tab bookkeeping failures that already hold the host. - orcad writes its own orcad.trace.ndjson, closes it after every quit handler, and flushes it on process exit; a trace file that cannot be opened leaves tracing off instead of stopping the app or orcad. - Tests: the desktop wiring test proves the logger reaches the trace sink, and the privacy tests read every level the logger received. * fix(native-chat): log a created chat's tab-publication and launch-prompt failures through the host's logger * fix(native-chat): key a repeated chat failure on everything its entry writes The repeat suppression keyed on the message and the error's text, so two refusals with the same code but different causes, a plain error and a refusal of one code, or two object-valued errors shared a key and the second was swallowed for five minutes. The key is now the entry's whole written content (fields, code, errcode, refusal reason, cause chain, a stable rendering of a non-error value) plus the error's name and message; a refusal's reason is also written. * test(native-chat): pin that an error's name keeps two repeated failures apart * test(native-chat): build the refusal in the repeat-key test as the wire does * fix(native-chat): read an error's code and a refusal's reason by narrowing, not Reflect.get |
||
|
|
0b79720c2e |
feat(native-chat): the chat strip and the sidebar read the host's child records (#22614)
* feat(native-chat): publish the host's child records to the status summary and the chat strip
The status summary and the background-task channel now read a session's child
records from the host's canonical store, through the status sink its row
landed in, and derive the legacy task and subagent shapes from the same views.
The parent row folds its child-work liveness from those records at ingest,
not from the summary's task list. The adapters no longer push their task DTO
to clients: the onBackgroundTasksChanged path is gone, and a child-work ingest
is what republishes both the summary and the strip. Finished children stay
listed until the session's own next turn starts. A reader that predates child
views never receives a roster whose rows are all settled.
* feat(native-chat): the chat strip reads the host's child records with its parent's verdict
The strip's roster now renders from the child views its channel carries, and
passes the verdict the session's own status row gives its children, built the
way the sidebar builds it (the row's freshness and the status feed's
observation). So one child reads the same in the strip and the sidebar, live,
after the transport drops, and once the row goes stale. A roster of finished
children stays shown until the next turn but no longer animates the monitoring
indicator or blocks conversation commands.
Tests: an end-to-end run on a host with no renderer (a real hook server as the
status sink) shows the summary and the strip channel carrying the same records
at every step, the parent row folded from them, retention, and an older
client's task list holding live work only; a wired renderer test shows both
surfaces agree when live, lost and stale.
* test(native-chat): a newer host's view degrades, an older reader keeps its live roster, a finished roster holds nothing open
- The view decoder ignores unknown keys, degrades unknown kinds, states,
outcomes and memberships, and drops only rows it cannot identify.
- At the RPC boundary a reader that predates child views gets no strip for a
roster of finished children and never the views themselves; a stop-only
reader keeps rows whose host offers no targeted stop.
- The strip shows finished children without reading them as live work.
- The row keeps its child list's identity when a summary repeats it.
* refactor(native-chat): the summary's task list is the legacy projection's live rows, unfiltered
* test(native-chat): type the switch tests' mocks instead of asserting them
* test: remote clients advertise reading child views
* docs(agent-status): the structured row folds the store's child records
* fix(agent-status): keep the view reader's header from reading as a value import to the renderer boundary
The renderer node-builtin boundary test scans raw text, so a header comment
that said "imports" ahead of the import block turned the type-only import of
agent-status-child-work into a value edge that reaches node:crypto.
* refactor(native-chat): the status summary's broadcast equality gets its own module
The status feed crossed the file-size limit once the summary gained the main agent's turn
outcome beside the child views. Which summary changes reach every session list now lives in
structured-agent-session-status-summary-equality.ts.
* fix(native-chat): command admission reads the strip's child records
A conversation command was refused on the provider tracker's own roster
while the strip read the host's child records, so a drift between the two
rule sets could refuse /clear with a stop instruction the strip had no
button for. Admission now reads the same records through the same read as
the strip, uses the strip's liveness fold, and asks for a stop only when
the strip renders one. The adapter contract no longer exposes the tracker
roster, so no host decision can read it.
Also records when the legacy child shapes die, every earlier death of a
settled child, and the display-precision invariant behind the summary's
clock tolerance.
* refactor(native-chat): command admission takes only what it reads of a turn
* fix(native-chat): the session list drops a session's children when the store does
A session's end no longer removes its child records: a child still running
settles with an outcome nobody reported, and a finished one stays listed.
Records now leave only at the session's own next turn, at the cap on settled
records, or when the host lets go of the session and its row leaves the store.
The summary kept after the host lets go used to strip its children on close,
a rule of its own. It now re-reads them from the store when the row leaves,
through the same read every live summary uses, so the session list and the
chat strip list the same children at each step, including a forget with no
close. Closing only revokes ownership, as before the child records existed.
* test(native-chat): write the Codex frame script's parent row out step by step
Once every surface reads the child records, the provider tracker's roster is
no oracle: it and the records read the same child executions, so agreeing
with it cannot catch a defect in either. Each frame now states the child
liveness and the parent row it must fold to.
* fix(native-chat): the idle sweep and the restart snapshot read the host's child records
The idle sweep (keep an agent running while its subagents or commands run) and
the restart-resume snapshot (what a chat was doing when Orca stopped it) both
read the provider tracker's roster through the adapter interface, which no
longer carries it. Both now take the host's one child-record read, the same one
the status summary, the chat strip and command admission use.
The snapshot's working test also folded that roster through the shared fold's
old `backgroundTasks` input, which the fold no longer reads, so a settled lead
whose subagent was still running would have been offered nothing. It now hands
the fold the records.
* test(native-chat): the child-record tests follow the merged command lifecycle
A command is a live child record from its start and is removed, not settled,
when it stops, whatever Codex tagged it. The end-to-end switch now shows the
child's `npm test` as a live row beside its dev server, and both are gone once
they exit; only the finished subagent stays listed until the next turn. Letting
go of the session is its tab closing, since a closed conversation whose tab
remains keeps its row.
Command admission's finished row is a subagent, the one kind that settles, and
the failed-verdict row test admits its live subagent as a host record, the only
thing the row folds.
* refactor(native-chat): the status feed's journal projection cache gets its own module
The status feed crossed the file-size limit once the child records joined the
agent-start signal and the completion feed's status read. The per-journal
projection, cached per commit, now lives in
structured-agent-session-status-journal-projection.ts.
* test(native-chat): the admission test's compaction resolves with a real outcome
Main's compaction result is a tagged outcome; the host-level admission test
resolved its mock compaction with an empty object.
* feat(native-chat): the sidebar lists running subagents; the strip, running then the newest finished
The host keeps every child record; what each surface lists is picked from them on every read, so
nothing is stored twice. The status summary, which every session list reads, now carries only
running children (and a finished one whose shell still runs, which reads monitoring): a finished
or failed subagent leaves the sidebar and stays in the chat's strip. The strip lists every running
child, then the newest finished ones, 100 rows in all; more than 100 running all show.
This matches common practice: sidebars show live subagents, and finished ones stay in the chat's
panel, newest first. No wire field is added. An older client reads fewer rows: its legacy task
lists were already live-only in the summary, and the strip's settled tasks come from the same
bounded roster.
* fix(sidebar): one rule for what the worktree sidebar lists: running children, from every source
`worktreeSidebarListsChild` is the one definition: a child that runs, counting a finished one
whose own shell still runs (it reads monitoring). The sidebar's row builder applies it to every
child source it reads, a terminal agent's hook roster and a chat session's records alike, and the
host's status summary applies the same predicate, so the sidebar's payload stays small. The chat's
strip keeps finished children, newest first.
A terminal agent's hook roster already drops a child on its own stop, so nothing changes there:
a teammate between turns and a child gone quiet still run, and still show. The selection module
moves to `agent-child-work-listing.ts`, since it now covers every source, not only chat sessions.
* test(native-chat): the switch test passes the startup child key main's status bar takes
* fix(native-chat): a finished child stays until the user's next send, not a turn Claude opens on its own
Claude wakes the agent on its own when a background task ends, and that wake is a
new root turn. Keying retention on the newest root turn retired every finished
child about two seconds after a background agent or shell finished, so its
outcome never showed in the strip.
Retention now keys on the user's newest send the provider accepted (a message, a
steer or a command), with the journal epoch so a rewind still retires. A turn the
provider opens itself and a subagent's turn carry no send. Replayed captured wake
orders through the real adapter, hook server and status feed.
* fix(native-chat): a background Stop reaches the tasks the child records show
The strip draws a row's Stop, and /clear, /compact and rewind wait for background
work, from the host's child records, but the Claude adapter still resolved which
tasks a Stop reached from its own tracker's roster, and refused to stop at all
once that roster was empty. A task the records kept live after the roster dropped
it showed a Stop that sent nothing and blocked those commands until the chat tab
closed.
The host now resolves the provider ids a Stop sends from the records (the same
per-row rule the strip and admission use; every such row for stop-all), and the
adapter stops exactly those, with no tracker guard. An acknowledged stop ends the
record: a running task sends its own stopped frame first, and the CLI answers
success with no frame for a task it no longer knows. A refused stop leaves the
record live. No production code reads the tracker's roster any more.
* fix(native-chat): one rule for a finished child that still owns live work, at any depth
The listing kept a finished child whose work ran through any depth of ownership,
but retention at the user's next turn protected only the direct owner, so a
finished agent whose finished subagent still ran a shell was removed and that
subagent jumped to the top level. Both now read settledOwnersOfLiveWork.
* fix(native-chat): an older client sees a Codex child's shell as it did before views
Clients that predate child views read a flat task roster derived from the views.
It listed a Codex child agent's shell as an extra row beside the running agent,
then as a bare command once the agent finished. The derivation now hides a
running agent's commands and names a finished agent's as "<agent> — <command>",
as the Codex tracker did; the label rule moves to a shared module both use.
* fix(native-chat): the chat decodes a roster's child rows once, as the frame arrives
The client reducer compared raw wire rows, so a row shaped by a newer host could
throw there, and the strip decoded a new array on every render, which defeated its
grouped-rows memo while a turn streamed. Rows are now decoded where the frame
enters the reducer, an unchanged roster keeps its identity, and the strip's parent
context is rebuilt only when one of its values changes.
* fix(native-chat): the strip channel forgets a closed conversation's roster
It kept the last roster fingerprint of every conversation for the host's lifetime.
The conversation map now tells observers when one leaves it. Also corrects the
summary's children comment: it carries running children only.
* docs(native-chat): rewrap the retention comment
* fix(native-chat): a task's own ending replaces a Stop's, and a child finished after the user wrote stays
Two lifecycle gaps from the round-1 fixes.
A Stop acknowledged ahead of the task's own ending relabelled it. The SDK hands
Orca a control answer as soon as it reads it and queues other frames, so when a
task finished just as the user pressed Stop, the acknowledgement arrived before
the task's completion the CLI wrote first, and the task read "Stopped" with its
result lost. An acknowledged Stop now ends a record provisionally (outcome basis
`stop-acknowledged`); the task's own terminal frame replaces it, and nothing
replaces an ending the task reported itself.
A finished child still vanished with no new action from the user when the send
the provider took was written before the child finished: a steer Claude takes at
its next boundary, or a queued draft handed over at the end of the turn. The
user's next turn now carries when they acted (the send's written time, or the
draft's queued time, both on the host clock), and only children that finished at
or before it retire; a later one stays until the user's next send. A rewind
still retires every finished child.
Also: a Stop-all keeps stopping the remaining tasks after one request fails, then
reports the failure.
* fix(native-chat): the strip keeps one empty list for a roster that omits one
A roster with only running or only finished rows made a new empty array on every
render, so the strip regrouped its rows each time. One shared empty list keeps
its memo.
* fix(native-chat): a strip row whose owner the 100-row budget cut renders under the main agent
The budget can keep a finished child and cut its finished owner. The child still
named that owner, so it rendered nowhere. The selection now clears an owner it
did not keep, as the view contract says for an owner outside the projection.
* fix(native-chat): a stop-all that times out stops asking, and "no children" is sent once
A stop-all kept asking after a request timed out, so a Claude CLI that stopped
answering control requests cost one full deadline per task, while the chat's
sends, Stop and /compact waited behind it. A timeout now ends the loop; other
request failures still let the remaining tasks be stopped.
The chat strip channel never remembered that it had sent "no children", so
every change in a chat with none re-sent that frame to each subscriber. It now
remembers it, and forgets only when the conversation closes.
Also renames agent-child-work-stop.ts to agent-child-work-stop-targets.ts, which
says what it answers: the provider ids a background Stop reaches.
* fix(native-chat): the status feed reads a journal snapshot with no submissions, and e2e tests use the child-work reader
CI on
|
||
|
|
5cda0f4508 |
refactor(native-chat): keep agent-session records in the chat journal database (#24006)
* fix(native-chat): report a failed startup chat reconcile instead of failing app startup At startup the chat host re-checks every saved chat's lease and writes the result to agent-sessions.json. If that write failed (the file lock gave up, the file could not be written, or the file was written by a newer Orca and is read-only here), reconcileRestartLeases rejected, the startup IPC call rejected, and the renderer fell into its degraded "Session restore failed. Changes won't be saved until restart" mode. The reconcile is bookkeeping: a lease left unreconciled grants no writer, and every attach, send and read of a chat reconciles its own lease again. So the startup reconcile now reports its failure through a new optional host dependency, onStartupReconcileFailure, and resolves. The runtime routes it to its onError sink under the scope structured-agent-session-startup-reconcile, or logs it when no sink is installed (the desktop installs none). * fix(native-chat): read restored chats without waiting on lease bookkeeping With native chat on and a chat tab open at quit, the renderer's startup also awaits the chat tab restore (session.tabs.listAll). That restore re-ran the lease reconcile before reading each chat and rethrew its store failure, then recorded each restored tab as visible through a store transaction that throws on a held lock or a read-only store. Either one failed the restore, so startup still fell into "Session restore failed". Reading a chat grants no writer, so the reconcile startup and the restore run is now a reader's: createReaderReconcile never throws, answers whether every lease is settled (recovery is resolved only then; the journal opens either way), and reports each distinct failure once until a reconcile settles. Attach and agent start keep the strict reconcile. The restore's tab republish logs a failed visibility write and still publishes the tab, since a client drops every unpublished chat tab; user-driven publishes still refuse. The host dependency is renamed onLeaseReconcileFailure (scope structured-agent-session-lease-reconcile), since it now also reports for reads. * fix(native-chat): keep every record-store write off the startup chat read path Round-2 review found two more writes on the startup chat restore that could still fail it and put the app into "Session restore failed": republishing a /clear replacement recorded its tab visibility strictly, and resolving a chat's recovery rethrew its store error. The restore also paid one lock wait per tab and per batch of chats while the lock stayed held. The restore now derives tabs from state it already holds: - publishStructuredAgentSessionTab splits into the strict write and projectStructuredAgentSessionTab, which only updates the runtime's snapshot. The restore and /clear replacements only project: a saved tab index already lists every restored chat, and a /clear moves the tab in the same write that commits it. visibilityWriteMayFail is gone. - Chats a legacy profile restores that the index does not list are recorded in one best-effort transaction (store.showSessionTabs), so a failure leaves the index absent to seed again rather than partial. - The read restore's recovery resolution is caught and reported through onLeaseReconcileFailure, deduplicated with the reconcile's reports. - Once lease bookkeeping fails in a restore pass, the rest of that pass skips it, so a held lock costs one wait for the startup reconcile and one for the restore, however many chats are open. User actions (create, reveal, attach, send, the /clear commit) keep their strict writes. * test: open, seed and read the agent-session record store through one harness Tests that open the durable agent-session record store, seed it, or read back what it persisted now go through agent-session-record-store-test-harness.ts instead of calling AgentSessionRecordStore.open or touching agent-sessions.json themselves. A later change that moves the store into the chat database then changes the harness instead of every test. No production code changes. Tests whose subject is the JSON file itself (its .bak recovery, salvage, schema versions, permissions, and what older builds read back) keep reading and writing the file directly; the storage move rewrites or deletes them. * fix(native-chat): start each restore pass from one lease check and stop its bookkeeping at the first failure The restore now runs one reader lease check for the pass and lets each chat re-check and resolve recovery only while the pass is still settled. The first refusal or failed write clears it for the rest of the pass, and every chat is still opened for reading. With another process holding the lock, startup waits on it once in prepare and once in the restore, however many chats are open; a legacy profile waits once more for its tab-index seed. * docs(native-chat): correct restore comments and a test name to match the final design * test: address the record-store harness by the host's state directory The harness took the store's own folder, so each caller picked one (join(root, 'store'), or 'agent-sessions' where a test read the store the runtime owns). A later change that moves the store into the state directory's journal database could not tell those apart, and would have had to edit every caller again. Every harness function now takes the state directory, the one the test's journal database and recovery capsule already live in, and keeps the store in the same subfolder the runtime uses. Callers pass that directory; store-only tests pass their temp directory unchanged. Format tests that share a directory with harness calls take the file path from testAgentSessionStoreFilePath. The folder name moves from a private constant in the runtime to AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness shares it without importing the runtime. Its value and every path built from it are unchanged. * refactor(native-chat): keep agent-session records in the chat journal database The record store's records, operation ledger, retired claim keys and chat tab index become tables in agent-session-journal.db (user_version 4). The version-4 migration copies agent-sessions.json in its own transaction and never writes, renames or deletes that file or its .bak. Each store write is one journal transaction over exactly the rows it changed, checked with the load rules; the file lock, the external-change refresh and its hash, the .bak rotation, salvage and the hot-path recovery fence are gone from the store. * wip: importer tests * test(native-chat): cover the records migration, the import, row writes and read-only records * docs(native-chat): retire comments that describe the records file as the live store * test(native-chat): drop the record-store harness's leftover file path and type the import fixture * test(native-chat): let the host harness cleanup wait out a recovery-offer read's lock * fix(native-chat): let Stop reach the agent when its ledger row cannot be written Stop's operation-ledger row now shares the database with the chat history, so damage, a full disk or a stranded transaction on that write refused the Stop before the interrupt. A cancel plan now takes its decision from the committed ledger in memory, runs without settling, and warns that the row was skipped. Other mutations answer proven damage with the typed "Unable to load this chat." refusal instead of the raw SQLite error. * fix(native-chat): answer whether a profile holds chats from the database's rows Every host install creates agent-session-journal.db, chats or not, and the version probe created it too, so its mere existence made every profile that ever installed the host wait on host install and reconcile at startup. The check now opens the database read-only and looks for a record or tab row, lets the records file answer while its import is still owed, and counts an unreadable database as present. The version probe no longer creates the file. * fix(native-chat): open a chat from history when its tab index cannot be written Over records a newer Orca wrote, every write is refused, so opening a closed chat from Agent Session History failed on the tab-visibility write and the chat read as unreachable. Like closing a tab, opening one now reports a failed restore-index write and still publishes the tab. * fix(native-chat): keep the records import owed when the backup read fails transiently A torn records file whose .bak could not be read (EACCES, EIO) was reported as unusable, so the migration completed with nothing copied and never retried. A non-ENOENT read failure of either copy now carries its cause, which the importer classifies as a read that can clear. * test(native-chat): pin that an unreadable records file never falls back to its backup * fix(native-chat): restore imported chats' tabs when the records file had no tab index A chat created while the import was owed recorded a tab index holding only itself. When the file it later imported had no index, that index still read as recorded, so the imported chats' tabs never came back. The import now clears the recorded marker in that case, and restore falls back to the profile's tabs. * refactor(native-chat): drop the unused in-transaction store write Nothing called it, and it bypassed the write queue and the read-only refusal. * docs(native-chat): say that an unusable records file is left untouched but never re-imported * refactor(native-chat): keep the provider handle chain check as main has it The chain-validation refactor has no measured need in this change. * docs(native-chat): retire lease-renewer comments that describe the records file as the live store * fix(native-chat): keep a throwing failure sink from failing the startup chat read The lease bookkeeping failure reporter called the host's failure sink directly, so a sink that threw turned a reported, recoverable store failure back into a rejected startup reconcile or read restore. The reporter now catches a sink throw and logs both the original failure and the sink error with console.warn. * test(native-chat): wait for a replaced host's restart-offer writes before cleanup A restart test replaces the host without tearing the old one down, so the old host's fire-and-forget restart-offer withdrawal could still hold the recovery capsule's lock directory when cleanup removed the test directory (ENOTEMPTY). The harness now hands hosts a capsule that tracks running operations and waits for them before removing the directory, replacing the rm retries. * docs(native-chat): retire the abandon helper's note that the store re-creates its directory * fix(native-chat): restore a chat opened while the import was owed beside the profile's chats When the imported records file had no tab index, restore fell back to the profile's saved tabs, which never list a Claude chat, and the seed then rewrote the tab table without the chat opened while the import was owed. The tab rows that chat left are now loaded as unrecorded, restore takes them together with the profile's chats, and the seed keeps their tab ids. * test(native-chat): pin that a create whose tab index write fails still opens the chat * docs(native-chat): say why restore puts chats opened while the import was owed first * test(native-chat): replace a ledger row rather than change it in place in the Send-now rerun test The record store freezes published rows in tests, so setting a row's outcome in place threw; the test now swaps in a changed copy, as its sibling cases do. |
||
|
|
007b7c0d32 |
fix(claude): end a message Claude started but never confirmed, and keep Claude running while it holds one (#23898)
* fix(claude): settle a queued send the CLI withdrew from its own cancelled frame Claude reports each uuid-stamped command's lifecycle (queued, started, completed, cancelled). A send it withdraws from its queue gets `cancelled` before the interrupt or cancel_async_message answer, so a lost or failed answer no longer leaves that send pending: it settles as withdrawn, with the same reason and words as the receipt path. A command the CLI already started also ends `cancelled` when its turn is interrupted or fails, so `cancelled` after `started` is not a withdrawal; an echoed send has left the waiter lists and is never reached. Tests replay real 2.1.280 captures, scrubbed. * fix(claude): release a doubted send when the CLI reports its session idle A Claude send whose write ended in doubt is recorded `unknown`, and a live `unknown` reads as work still owed, so the chat showed Working until the child exited. Claude sends `session_state_changed idle` only once its whole queue has drained, so it can no longer be holding that send. The runtime now routes that report to the host's existing release, the same one Codex's thread-stopped report uses; it retires `unknown` only, never `pending`. * fix(claude): keep a command's started mark when a redelivery re-emits queued; fixtures name msg_lifecycle_v1 * fix(claude): settle every terminal lifecycle state of a send the CLI never echoed A send the CLI started, then cancelled before any echo, stayed pending: it may already be in the conversation, so it is released as doubt (unknown, recovered), never withdrawn and never re-sent. A late echo still accepts it. The 2.1.280 schema has two more terminal states. `discarded` (the CLI ended its session with the send still queued) settles as not delivered; `refused` (declined before it queued) settles as not accepted by the provider. After `started`, either one is doubt, as `cancelled` is. The late-settlement path gains an `unknown` outcome, which the host records as released doubt. * fix(claude): release a send the CLI took but left unanswered when it goes idle `session_state_changed idle` comes only once the CLI's queue has drained, so a send it took that is still unanswered there got no echo and never will: a turn that throws can leave `started` with no terminal state. Idle releases it as doubt. What proves the CLI took a send is its lifecycle frame. On a CLI that reports no lifecycle, it is the send's place on stdin: one whose write finished before an interrupt went out was read before the interrupt was, so the first idle after that interrupt releases it too. A send armed ahead of the interrupt but written after it is left alone, since the CLI may still run it. * fix(native-chat): keep the idle sweep off a Claude child that holds a send A Claude retrying a rate-limited request has taken the send but echoes nothing, so no turn row exists yet and the sweep rested the child after the idle window, turning the send into doubt. The adapter now reports whether the CLI holds a send (lifecycle `queued` or `started`, not yet echoed or ended), derived from the live waiters, and owed work counts it. Nothing is stored: every held send leaves the live set on its echo, its terminal lifecycle state, the CLI's idle, or the child's exit, so the hold ends with the send. * fix(claude): count only a started send at idle and as a held send 2.1.280's end-of-turn cleanup can report idle before it re-reads its queue, so a send read in that window goes queued, idle, started. Releasing every taken send at idle doubted that live send and dropped Working. Only a `started` send is released at idle or keeps the child from the idle sweep; a `queued` one ends by starting and echoing, by a terminal lifecycle frame, or with the child. The stdin-order path for CLIs without lifecycle frames is removed: a doubted send retired there disables content matching on CLIs that mint their own echo ids, and no Orca failure called for it. Those CLIs keep the earlier behaviour. Comments that said only a failed write or child exit ends a waiter, or that idle comes only once the queue has drained, now say what ends one. * docs(claude): say only what the CLI's lifecycle frames and idle actually prove * fix(claude): hold the idle sweep while Claude has a send queued, not only started The sweep rested a child whose CLI had queued a follow-up behind a turn, dropping the send it had already taken. The hold now spans the CLI reporting it took the send until its echo, a terminal lifecycle state, or the child's exit. The idle release still covers only started sends: 2.1.280 can report idle before it re-reads its queue. * refactor(native-chat): give provider-proven late dispatch settlement its own module * test(claude): pin a steer a Stop interrupts after it started as doubt, not withdrawn |
||
|
|
ba39c6d5fd |
fix(native-chat): a failed startup chat-lease save no longer puts the app into "Session restore failed" (#23964)
* fix(native-chat): report a failed startup chat reconcile instead of failing app startup At startup the chat host re-checks every saved chat's lease and writes the result to agent-sessions.json. If that write failed (the file lock gave up, the file could not be written, or the file was written by a newer Orca and is read-only here), reconcileRestartLeases rejected, the startup IPC call rejected, and the renderer fell into its degraded "Session restore failed. Changes won't be saved until restart" mode. The reconcile is bookkeeping: a lease left unreconciled grants no writer, and every attach, send and read of a chat reconciles its own lease again. So the startup reconcile now reports its failure through a new optional host dependency, onStartupReconcileFailure, and resolves. The runtime routes it to its onError sink under the scope structured-agent-session-startup-reconcile, or logs it when no sink is installed (the desktop installs none). * fix(native-chat): read restored chats without waiting on lease bookkeeping With native chat on and a chat tab open at quit, the renderer's startup also awaits the chat tab restore (session.tabs.listAll). That restore re-ran the lease reconcile before reading each chat and rethrew its store failure, then recorded each restored tab as visible through a store transaction that throws on a held lock or a read-only store. Either one failed the restore, so startup still fell into "Session restore failed". Reading a chat grants no writer, so the reconcile startup and the restore run is now a reader's: createReaderReconcile never throws, answers whether every lease is settled (recovery is resolved only then; the journal opens either way), and reports each distinct failure once until a reconcile settles. Attach and agent start keep the strict reconcile. The restore's tab republish logs a failed visibility write and still publishes the tab, since a client drops every unpublished chat tab; user-driven publishes still refuse. The host dependency is renamed onLeaseReconcileFailure (scope structured-agent-session-lease-reconcile), since it now also reports for reads. * fix(native-chat): keep every record-store write off the startup chat read path Round-2 review found two more writes on the startup chat restore that could still fail it and put the app into "Session restore failed": republishing a /clear replacement recorded its tab visibility strictly, and resolving a chat's recovery rethrew its store error. The restore also paid one lock wait per tab and per batch of chats while the lock stayed held. The restore now derives tabs from state it already holds: - publishStructuredAgentSessionTab splits into the strict write and projectStructuredAgentSessionTab, which only updates the runtime's snapshot. The restore and /clear replacements only project: a saved tab index already lists every restored chat, and a /clear moves the tab in the same write that commits it. visibilityWriteMayFail is gone. - Chats a legacy profile restores that the index does not list are recorded in one best-effort transaction (store.showSessionTabs), so a failure leaves the index absent to seed again rather than partial. - The read restore's recovery resolution is caught and reported through onLeaseReconcileFailure, deduplicated with the reconcile's reports. - Once lease bookkeeping fails in a restore pass, the rest of that pass skips it, so a held lock costs one wait for the startup reconcile and one for the restore, however many chats are open. User actions (create, reveal, attach, send, the /clear commit) keep their strict writes. * fix(native-chat): start each restore pass from one lease check and stop its bookkeeping at the first failure The restore now runs one reader lease check for the pass and lets each chat re-check and resolve recovery only while the pass is still settled. The first refusal or failed write clears it for the rest of the pass, and every chat is still opened for reading. With another process holding the lock, startup waits on it once in prepare and once in the restore, however many chats are open; a legacy profile waits once more for its tab-index seed. * docs(native-chat): correct restore comments and a test name to match the final design * fix(native-chat): keep a throwing failure sink from failing the startup chat read The lease bookkeeping failure reporter called the host's failure sink directly, so a sink that threw turned a reported, recoverable store failure back into a rejected startup reconcile or read restore. The reporter now catches a sink throw and logs both the original failure and the sink error with console.warn. |
||
|
|
75040eba5a |
test: open, seed and read the agent-session record store through one test harness (#23986)
* test: open, seed and read the agent-session record store through one harness Tests that open the durable agent-session record store, seed it, or read back what it persisted now go through agent-session-record-store-test-harness.ts instead of calling AgentSessionRecordStore.open or touching agent-sessions.json themselves. A later change that moves the store into the chat database then changes the harness instead of every test. No production code changes. Tests whose subject is the JSON file itself (its .bak recovery, salvage, schema versions, permissions, and what older builds read back) keep reading and writing the file directly; the storage move rewrites or deletes them. * test: address the record-store harness by the host's state directory The harness took the store's own folder, so each caller picked one (join(root, 'store'), or 'agent-sessions' where a test read the store the runtime owns). A later change that moves the store into the state directory's journal database could not tell those apart, and would have had to edit every caller again. Every harness function now takes the state directory, the one the test's journal database and recovery capsule already live in, and keeps the store in the same subfolder the runtime uses. Callers pass that directory; store-only tests pass their temp directory unchanged. Format tests that share a directory with harness calls take the file path from testAgentSessionStoreFilePath. The folder name moves from a private constant in the runtime to AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness shares it without importing the runtime. Its value and every path built from it are unchanged. |
||
|
|
59b746ff3c |
feat(native-chat): one structured-chat journal database per host, owned by one process (#23613)
* feat(native-chat): one structured-chat journal database per host, owned by one process
Every structured chat on a state directory now lives in one SQLite file,
agent-session-journal.db, opened once by the process holding
agent-session-journal.owner: an empty SQLite file whose held BEGIN EXCLUSIVE is a
kernel byte-range lock, refused while another process holds it and released when
the holder dies.
- Stores own no connection: the per-chat handle, its close contract and the
close-retry registry are gone; closing a conversation drains its writes, and the
one connection closes last at teardown.
- The owner lock is taken at runtime start, before orca-runtime.json is written;
a process that does not own the chats is not published and refuses every
structured request with journalUnavailable and words that say what to do. It
retries the lock with backoff and runs the full install once it holds it.
- A journal that will not open fails the host install: every chat says "Unable
to load this chat." (journalCorrupt), and nothing is renamed, deleted or
rebuilt. A newer build's database is refused and left byte-identical.
- An append is one INSERT. The listing status is a column, written after the
rows it describes and keyed by (epoch, sequence).
- A per-chat journal from an earlier build is copied in verbatim (epoch UUID and
every sequence) on that chat's first open, and its directory is retired only
after the copy commits.
- auto_vacuum = INCREMENTAL, with freed pages handed back in bounded steps after
every delete.
* perf(native-chat): key journal rows by block so one chat's rows sit together
Each chat's live epoch owns a block of row ids, block * 2^32 + seq, so a chat's
rows share leaf pages with nobody else's, a replay is one range scan, and
replacing or rewinding a chat deletes one contiguous range. Measured on the
largest real chat (61 MB) beside 19 interleaved peers: 39 ms and 7.5 MB of WAL,
against 214 ms and 102 MB for a (session_id, epoch, seq) key.
- Ids are computed in Number arithmetic, never bitwise. A sequence is refused
outside [1, 2^32) and a block at 2^21, which keeps every id below 2^53.
- A replace, roll or import allocates a fresh block, moves the chat's pointer,
and deletes the old block in the same transaction, so no orphan block exists.
- The listing status write moves into its own writer beside the column.
* feat(native-chat): copy a chat's per-chat journal again when an older Orca wrote it after a downgrade
A per-chat journal.db that reappears after its chat was copied in is the newer
history: an older build, run after a downgrade, attached the chat and wrote it.
- journal_imports records the (epoch, tip) each chat was copied from, in the
copy's own transaction. A file already copied is never copied again, across
any number of restarts after a failed rename; a file that differs always is.
- Newest writer wins, per chat, with a row saying the chat was continued in an
older version of Orca. When both builds wrote past the recorded tip under one
epoch, the copy takes a fresh epoch, so readers reset instead of skipping rows.
- Each copied directory retires to its own .imported-<epoch8>-<ms> name, so a
second downgrade and re-upgrade never collides with the first.
* test(native-chat): fixture deps match the host journal database shape
Attach-flow and reconcile-attach fixtures stop passing a journal database those inputs do not take, and host and restore fixtures pass the one they now require instead of the removed journal root.
* test(native-chat): state why the runtime-state fixtures' existing casts are safe
* fix(native-chat): start up normally when this process cannot open the chat journal
A process refused the chat journal, because another Orca owns it or because its own journal will not open, failed startup restoration: the window booted in degraded no-save mode and a paired phone could not list any tabs. Startup restoration now treats the refusal structured requests are getting as having no structured host; terminals, tabs and saving go on, structured requests are still refused by the gate, and the install is retried on the next one. Any other install error fails startup as before.
* test(native-chat): name the owner-lock sweep test after the two sweeps it runs
* fix(native-chat): session history and terminal resume work while chats are refused
Session history (listing and preparing a resume) and a terminal typing a resume command only check whether a structured chat owns a provider session. In a process refused the chat journal they failed outright. They now take the refusal chats are getting as having no structured host, the same treatment startup restoration gets, through one shared helper; any other install failure still fails them. Chat requests keep the gate's refusal.
* test(native-chat): the first-work rename's fake journal saves the listing status
* fix(native-chat): open a chat whose per-chat journal file never got its schema
A crash between creating a chat's journal.db and creating its tables left an empty or schema-less file. Each chat used to open that file as an empty chat; the importer instead refused the open as "try again" forever. A file with no journal_sessions table is now read as never written, the same as one with no rows. A file that is not a database, or whose read fails, is still refused.
* fix(native-chat): let the event loop run between chats during startup restore
Opening a chat's journal is synchronous SQLite now that no per-chat directory
is created first, so the restore of every visible chat ran as one main-thread
task. Each chat now waits for a macrotask before it opens.
* fix(native-chat): import a per-chat journal in bounded batches
The one-time copy of an earlier build's per-chat journal ran as one
transaction, which blocked the main thread for 650 ms on the largest chat.
Rows now copy 512 at a time, each batch its own transaction, yielding to the
event loop between batches. The rows go into a block journal_import_blocks
reserves, which no reader follows and no other chat is allocated; the last
batch publishes the chat's pointer, repair marker and import marker together
and releases the reservation. A copy that stops midway leaves only that
block, which the next open clears and copies again. Two opens of one chat
import one after the other.
* fix(native-chat): refuse chats when the owner lock file cannot be opened
A lock file that is not a database, or cannot be opened, made the claim throw
before any refusal was recorded, so startup restoration failed on every
launch. The claim now sits in the same try as the database open and records
the same typed refusal.
* fix(native-chat): finish reclaiming pages a delete frees during a running pass
A reclaim pass ended as soon as the freelist stopped shrinking between steps,
so a second delete that freed more than one step's worth mid-pass ended it
early and left those pages on the freelist. A pass now ends only when a step
itself frees nothing, or the freelist is empty.
* test(native-chat): desktop session history is served while chats are refused
* fix(native-chat): a send to a chat holding a newer Orca's rows says to update
A chat opened read-only because a newer Orca wrote rows to it answered a send
with the generic write failure. It now refuses the way a database a newer Orca
wrote does, with the same reason and words.
* fix(native-chat): keep chat tabs while this process cannot list its chats
A process whose chats another Orca owns, or whose chat journal will not
open, has no structured host. Its session-tabs inventory still answered,
with no chat rows, and the renderer read that as "every chat was closed":
it removed the restored chat tabs and the next session save persisted
their placement away.
The inventory now says `agentSessionsUnverifiable` when the last tab
restore ran with chats on disk but no host to list them. The flag is set
and cleared at the per-client projection point beside the client-hosted
page hold, and the restore is memoised only once a host answered, so a
later lock takeover or journal open republishes the chats and clears it.
The renderer keeps agent-session tabs, and keeps cancellation tombstones,
against an inventory that does not affirm its chat set.
* fix(native-chat): say chats are open in another Orca, with this process's way past it
A process refused because another Orca owns the profile's chats sent the
generic `journalUnavailable` reason, so current desktop and phone surfaces
said "couldn't open this chat's history right now. Try again." — a step
that never helps while the other Orca runs.
The refusal now names its own reason, `journalOwnedElsewhere`, with the
refused process's kind (dev desktop, packaged, orcad) as a fact. Each kind
gets its own step: quit the other Orca, or give this one its own profile
(ORCA_DEV_USER_DATA_PATH) or data folder (ORCA_USER_DATA). The sentences
are added to the shared notice copy, the desktop catalogs in all six
locales, and the boot catalog.
A client that predates the reason reads it as none and keeps the code's
words; an unknown kind reads as the packaged app's step. The `message`
released clients print is unchanged. Which requests refuse does not change.
* fix(native-chat): restore chats on taking ownership, without a list to ask
A refused startup kept its hostless result, so after the owner quit this
process never installed a host, never reconciled restart leases, and kept
telling clients it could not list its chats until a desktop chat request.
Taking the lock now reruns startup restoration once and pushes the chats.
* test(native-chat): a navigation reply says chats are unverifiable while refused
* fix(native-chat): retry a refused owner lock at most every 5 seconds
The lock frees as its holder exits, but a refused process only learns that
on its next retry, and the 30 s cap left a second Orca refusing chats for up
to half a minute after the owner quit. One retry is an open and BEGIN
EXCLUSIVE on an empty file.
* fix(native-chat): install before deciding whether a takeover must republish chats
A list that landed on the refused startup after the lock was taken finished
after the takeover had already checked, so nothing republished. The takeover
now installs first, waits for any restore in flight, and restores only then;
the restore that clears "cannot tell" pushes the frames itself, so a list
that heals the inventory first reaches subscribers too.
* fix(native-chat): no takeover lands a host after the runtime stop
Quitting cancelled a refused claim's retry only at its end, so a retry firing
during the stop's awaits took the lock and installed a host the stop never
tore down, and the lock was then released under an open journal. The stop
now cancels the retry first, keeping the refusal, and repeats its teardown
while an install that began during it (a takeover already under way) is
pending, so no journal connection outlives the lock.
* fix(native-chat): show a thrown refusal in its own words, not its code
A refusal the host throws reaches the client as an RPC error whose message is
the bare code; its reason and facts ride only in the error's data, which no
client read. The chat pane's status line therefore printed
agent_session_journal_unreadable, a send took the bare "not sent" path, and
other writes said the outcome was unconfirmed.
One shared reader, agentSessionThrownRefusal, now reads the refusal from the
error data. A failed history read shows the refusal's read-history words, a
send keeps the refusal behind its Retry exactly as a returned refusal does, and
the other writes (desktop and phone) name the refusal instead of doubting the
outcome. The phone's read failure goes through the same reader.
* fix(native-chat): log a failed journal open once per distinct failure
Every chat request retries a journal open that failed, which is intended, but
each retry also logged the failure with its full stack: a junk database file
logged the same "file is not a database" error 189 times in a minute. The open
now logs a failure only when its code and message differ from the last one
logged, and forgets it once an open succeeds. The retry is unchanged.
The open moves to its own module beside the runtime, which had no room left.
* fix(native-chat): restore lists a chat from its per-chat file and copies it on first use
Startup restore opened every restored chat, and that open copied the chat's
per-chat file into the host database, so the first boot after an upgrade paid
the whole one-time copy before the chat list appeared.
A restore open now reads a chat that is still in its per-chat file straight
from that file, read-only, with the importer's own reader, and closes the file
before moving on. That read drives the listing, the status row and the
restart offer, as it did when every chat had its own file. The copy becomes
owed work on the chat's write queue: it runs before the chat's first write,
and a reader that reaches the chat awaits it. A chat the host already holds,
or that was copied before, still opens through the import and its reimport
rules, and so does a file whose read needs a repair written.
* fix(native-chat): no host stays registered after a stop an install spanned
Each teardown pass clears the registered host before it awaits an install in
flight, and that install registers its host when it finishes. The pass then
tore the host down but left it registered, so a request after the stop was
served by a host whose journal was closed. The stop now clears the slot once
its passes are done.
* fix(native-chat): checkpoint the journal with a full flush on macOS
synchronous = FULL fsyncs each commit, but macOS fsync leaves the drive cache
unflushed, so FULL alone does not survive a power loss there. With
checkpoint_fullfsync, each checkpoint uses F_FULLFSYNC; elsewhere it is a no-op.
The comment that said FULL alone was enough is corrected.
* fix(native-chat): delete a per-chat journal once its copy verifies
An imported chat's per-chat file was kept under an `.imported-*` name, which
doubled the disk its history takes. The copy now reads back from the host
database before it is published: its items, submissions, epoch and tip must
match the file's. Only then does one transaction publish the chat with its
import marker, and the file and its WAL files are deleted, the directory too
when nothing else is in it (a pre-SQLite transcript there is kept).
A copy that does not match is never published: the file stays, the chat is
refused as unreadable ("Unable to load this chat."), and the mismatch is logged
once. A file left behind by a failed delete or a crash matches the marker, so
the next open deletes it rather than copying it again; a file an older build
wrote after a downgrade still differs, and is still copied again.
* fix(native-chat): verify an imported chat a batch at a time
The check that a copied chat reads back as its per-chat file folded both whole,
each in one synchronous task: over half a second on the largest chat. Both
reads now go a batch at a time between turns of the event loop, like the copy
itself, and count rows as well, so a copy that lost a row with no item in it
is caught too.
* fix(native-chat): restore reads a chat's per-chat file a batch at a time
Restore folded a chat still in its per-chat file in one task, so the largest
chat's file held the main thread for about half a second at startup. The fold
now takes the file a batch of rows per turn of the event loop, into the same
fold a replay uses, and nothing reads it before it is done. The file is still
closed before restore moves on.
* fix(native-chat): end a per-chat copy on a turn of its own
A chat's first open ran the copy's last steps (the verified publish and the
per-chat file delete) and the replay of what was copied in one task. The copy
now yields before it returns, so the replay, which every open runs, is a task
of its own.
* fix(native-chat): commit a per-chat copy's batches without an fsync each
Each 512-row batch of a chat's first-use copy committed under synchronous =
FULL, so a large chat paid one fsync per batch, about a quarter of its first
open. The batches now commit under NORMAL, set and restored in the batch's own
task so no other chat's commit runs under it. The publish that makes the copy
visible still commits under FULL, and under WAL that sync makes every earlier
batch durable with it. A crash before it leaves only the unpublished block,
which the next open clears and copies again.
* fix(native-chat): roll back a chat journal transaction whose COMMIT fails
The shared connection's transaction rolled back only when its body threw. A
COMMIT that failed left the transaction open, so every later write, for any
chat, failed with "cannot start a transaction within a transaction", and reads
saw rows that never committed. Under the unsynced copy the failure also tried
to restore the sync level inside the open transaction, which SQLite refuses,
so the caller got that error instead of the COMMIT's.
One transaction helper now covers the body and the COMMIT, rolls back whatever
transaction survives, and rethrows the original error. Schema creation uses it
too. If that ROLLBACK fails as well, the connection is marked stranded: each
later use retries the ROLLBACK, and until one goes through every chat gets the
same "history unavailable, try again" refusal a journal that will not open
gives. The rollback that frees it also restores the FULL sync level.
* fix(native-chat): keep the chat journal connection until its close succeeds
Closing the journal dropped its connection handle before closing it. A close
that failed left the database reporting itself closed with the connection still
open, so the stop that retried the teardown found nothing to close and released
the owner lock over a live connection.
The handle is now dropped only once the close succeeds. A failed close keeps
the runtime pending and the lock held, and the next stop closes that same
connection before it releases the lock.
* fix(native-chat): publish the runtime only once it holds the chat journal lock
When this process could not open the owner lock file at all (a permission
error, or a file that is not a database), the runtime counted that as owning
the chats and wrote orca-runtime.json. That overwrote the real owner's entry,
so the CLI was sent to a process that cannot serve its chats.
A claim that throws is now refused like one another process holds: the runtime
starts but does not publish, the claim's existing retry keeps asking for the
lock, and discovery publishes once the retry takes it. Chats still get the
refusal for the failure itself, and startup restoration reruns on the takeover
the same way it does after another owner quits. A sole process whose lock file
never opens is not found by the CLI until it does.
* fix(native-chat): keep a chat's history when an older build started it over
The first copy deletes a chat's per-chat file, so an older build run after a
downgrade finds no file and starts the chat from nothing. On the re-upgrade that
fresh file was copied in as the newer history, replacing everything the shared
database held for the chat, and then deleted.
A file whose epoch is not the one last copied and that opens with
`session_created` is now kept: neither copied nor deleted, and the chat keeps
the history it has. A file that carried the copied epoch on is still copied
again, as before.
* test(native-chat): pin which chats startup restore copies
Restore copies a chat still in its per-chat file only when restore itself has
to write to it: settling what the last run left open, here a running tool call
or a send handed over and never answered. Every other restored chat stays in
its file until its first use.
* test(native-chat): pin the copy wait on a read that opens a chat restore opened
A read queued behind restore's open of the same chat reaches the conversation
through its own open rather than the listing. It must still wait for the
owed copy, or it reads the chat before its history is in the one database.
* fix(native-chat): record a set-aside per-chat file so no later open reads it
Setting aside a file an older build started over is decided once and kept in
the new `journal_set_aside` table (schema 2, additive), with the file's epoch
and tip as they were. Every later open of the chat skips the file without
opening it, across restarts and after the older build writes more to it:
anything written there grows from that build's own start, never from this
build's history.
The best-effort delete moves beside the per-chat file reader.
* fix(native-chat): set aside any per-chat file at an epoch this build never copied
A chat's per-chat file is deleted once its copy verifies, so a file that
reappears at another epoch was never this build's history, whatever its first
row says: an older build started the chat over, possibly rewinding it after
(`handle_forked`), or rolled the epoch of a file whose delete had failed.
Copying any of them would replace everything the chat holds, so each is set
aside. Only a file still at the copied epoch is copied again (it grew) or
deleted (it did not). The first-row check is gone.
* fix(native-chat): copy a reappearing per-chat file again only while this build has not written past the copy
A per-chat file that an older build carried on under the copied epoch was
copied again even when this build had also written to the chat since the
copy, or had rolled its epoch. The second copy replaced the chat's block,
so what was sent in this build after the copy was gone for good.
Now the file is copied again only when the chat still stands exactly as it
was copied: the same epoch and tip the import marker recorded. Otherwise it
is set aside like any other file that is not this build's history, left on
disk untouched and recorded so no later open reads it. A second copy
therefore never replaces rows this build wrote, keeps the file's own epoch,
and the fresh-epoch rewrite goes away. The row it adds now says the history
includes what the older version recorded, not that anything was replaced.
* test(native-chat): pin that a chat founded here keeps its history, and the v1 schema upgrade
A chat this build founded has a pointer and no import marker, so a per-chat
file an older build later starts for it is set aside. Nothing pinned that
half of the rule: letting such a chat be copied again replaced its history
and every test still passed. A second test pins that a database written at
schema version 1 upgrades in place, gaining the set-aside table and keeping
its import markers.
* test(native-chat): drop a lost copied row by patching the source, not wrapping it
* chore(mobile): restore the mobile lockfile to main's
* fix(native-chat): pass a classified journal refusal through a send or Stop unchanged
* fix(native-chat): refuse a read whose owed copy fails as a failed open does
* test(native-chat): measure only the replace's WAL in the block-key case
Opening the chats starts a free-page pass that waits one event-loop turn,
and the seed never yields one, so that pass was still pending when the
replace committed. It woke during the async stat and reclaimed the pages
the replace freed, adding ~500 KB of WAL whenever the stat lost the race
(Linux CI). Drain that pass before measuring and stub the replace's own.
* test(native-chat): the RPC fixture's status journal can save its listing status
The status feed now hands every projection to the journal, which decides whether it is worth saving.
* test(native-chat): state why the RPC fixture's status journal cast is safe
* fix(native-chat): refuse a per-chat copy whose rows differ from the file, not only its counts
* fix(bench): build the replay benchmark's baseline arm from the base tree and release its handles on failure
* fix(native-chat): retry a failed listing status save on the next read of a cached status
* refactor(native-chat): drop the chat journal owner lock; the process instance lock already guards the profile
The journal carried its own exclusive lock, with a retry loop, an in-process
takeover, lock-gated runtime discovery and a "chats are open in another Orca"
refusal. Every shipped process kind (packaged desktop, serve mode, orcad)
already refuses a second instance on one profile before the journal opens, so
the lock only ever mattered for dev desktops, which the next commit covers at
the process level instead.
The host now opens its one journal connection at install with no lock. What a
sole process whose journal will not open needs stays: the install refusal
recorded for the gate, the no-host startup path, and the unverifiable chat
inventory, now in structured-agent-session-host-refusal.ts. The unreleased
journalOwnedElsewhere reason, its processKind fact and their copy are removed.
* fix(startup): dev desktops take the single-instance lock, and a second one says why it quit
Dev skipped Electron's single-instance lock so parallel `pnpm dev` runs from
several worktrees would not quit silently, but two dev processes on the
default orca-dev profile then write the same stores at once. Dev now takes
the lock like packaged builds: a second launch on the same profile focuses
the first window and exits with code 3, printing one stderr line that names
the taken profile and how to run another copy (ORCA_DEV_USER_DATA_PATH).
Serve mode, the macOS diagnostic bypass and the E2E harness are unchanged:
an E2E launch still skips the lock unless it sets
ORCA_E2E_ENFORCE_SINGLE_INSTANCE_LOCK=1.
* refactor(native-chat): key journal rows by chat, epoch and sequence
Rows in the host's journal database are now addressed by the chat's own
identity, with `(session_id, epoch, seq)` as the primary key, the same
shape each per-chat file already used. The block-keyed layout goes with
everything built on it: the block column and its allocator, the 2^21
block ceiling, and the import's reserved block table.
A first-use copy writes its rows under the file's epoch, which the chat's
pointer does not name until the verified copy publishes it, so no reader
sees a half-copied chat. A try that stopped midway leaves only rows no
pointer names, and the next try deletes them before it copies again.
Replace, rollover and repair delete by (chat, epoch).
This build's history always wins: once a chat was copied or founded here,
any per-chat file that reappears is set aside, and the same-epoch copy
again after a downgrade is removed.
The bounded free-page reclaim after every delete is dropped;
`auto_vacuum = INCREMENTAL` stays at file creation, so a later periodic
reclaim can still be added. Session search keeps its own step.
The schema moves to version 3. Versions 1 and 2 were written only by
unreleased builds of this change and are refused as found, not migrated.
* fix(native-chat): open a chat journal a newer Orca wrote read-only instead of refusing it
After a downgrade, the host's journal database carries a newer user_version. It was refused
outright, so every chat's history disappeared. It now opens on a read-only connection, as the
per-chat journals did: each chat shows what this build can read, from the database or a per-chat
file never copied in, and every write is refused with "Chats were saved by a newer Orca. Update
Orca to keep using them." Nothing is written, copied, repaired or founded, and the file stays
byte-identical. A table the newer schema changed reads as the same read-only refusal, not damage.
* refactor(native-chat): leave the saved listing status to the change that reads it
Nothing in this change reads the per-chat listing status column: it was a stored copy of a fact
the status feed derives, written after every turn end and cleared on every epoch change. The
status_json / status_seq columns, their writer, the saved-status type, the status feed's save and
its retry on a cached projection all go, with their tests. The change that lists chats from a
saved status adds the column back beside its reader.
* fix(native-chat): a chat saved by a newer Orca says to update Orca, not to try again
When a newer Orca wrote the chat journal, this build opens it read-only. A send or a Stop was
refused with the reason `journalUnavailable`, so today's desktop and phone clients chose the
words for an open that can clear: "Orca couldn't open this chat's history right now. Try again."
Retrying never cleared it; only updating Orca does.
The refusal now names its own reason, `journalWrittenByNewerOrca`, whose words are "Chats were
saved by a newer Orca. Update Orca to keep using them." A read refused the same way names it
too. An older client does not know the reason, drops it, and falls back to the code's words
("Orca couldn't read this chat's saved history."), and released clients still print the message.
* fix(native-chat): a chat journal from an unreleased build reads as unusable, not as retryable
A chat journal database stamped with schema 1 or 2 was written only by unreleased development
builds of this change. Opening it threw a plain error, which every chat reported as "Orca couldn't
open this chat's history right now. Try again." Retrying never cleared it.
It now throws a named error that is classified as unusable, so every chat says "Unable to load
this chat." The one log line names the file, says an unreleased development build wrote it, and
says to move it aside. Nothing migrates or renames it.
* docs(native-chat): drop the second-Orca-owns-the-chats case from three comments
The chat-only owner lock is gone, so only a chat journal that will not open leaves a runtime
unable to list its chats.
* docs(native-chat): correct three chat-journal comments the redesign left behind
A per-chat file left without its WAL is set aside, not copied again; nothing runs an incremental
vacuum yet, so the auto_vacuum mode is kept for a later pass; and the idle sweep drops a chat's
in-memory fold, since a chat holds no journal connection.
* refactor(native-chat): stop exporting chat-journal names nothing imports
Each is used only inside its own module now; the teardown's export served a deleted test.
* test(native-chat): name the version-0 test for what it covers, and check every journal table
The test named 'migrates an older user_version forward' covers only a version-0 file that already
has its tables; versions 1 and 2 are refused. The table test now also checks journal_imports and
journal_set_aside.
* fix(startup): a second dev launch's exit line no longer claims it focused a window
The running dev instance may be a background launch or a server, which show no window. The line
now says only that this launch passed its request to that instance.
* fix(native-chat): a failed structured-chat install closes the journal connection it opened
The install opened the chat journal database and closed it only if the record store then failed
to open. A later failure, such as the model catalog wiring or the host constructor, left the
connection open, and the next install opened a second one in the same process. Every failure
after the open now closes it.
|
||
|
|
18327d9665 |
fix(claude): a queued message Claude withdrew is settled from Claude's own cancelled event (#23862)
* fix(claude): settle a queued send the CLI withdrew from its own cancelled frame Claude reports each uuid-stamped command's lifecycle (queued, started, completed, cancelled). A send it withdraws from its queue gets `cancelled` before the interrupt or cancel_async_message answer, so a lost or failed answer no longer leaves that send pending: it settles as withdrawn, with the same reason and words as the receipt path. A command the CLI already started also ends `cancelled` when its turn is interrupted or fails, so `cancelled` after `started` is not a withdrawal; an echoed send has left the waiter lists and is never reached. Tests replay real 2.1.280 captures, scrubbed. * fix(claude): release a doubted send when the CLI reports its session idle A Claude send whose write ended in doubt is recorded `unknown`, and a live `unknown` reads as work still owed, so the chat showed Working until the child exited. Claude sends `session_state_changed idle` only once its whole queue has drained, so it can no longer be holding that send. The runtime now routes that report to the host's existing release, the same one Codex's thread-stopped report uses; it retires `unknown` only, never `pending`. * fix(claude): keep a command's started mark when a redelivery re-emits queued; fixtures name msg_lifecycle_v1 |
||
|
|
c5330d0d52 |
fix(native-chat): stop killing processes that only inherited a chat's spawn tag (#23460)
* fix(native-chat): stop signalling processes that only inherited a spawn token A spawn token is an environment variable, so every descendant of a provider child carries it. The Linux-only startup scan treated any carrier no lease claimed as a lost provider child and sent it SIGTERM, which also hit editors, tmux servers and nested Orca processes the agent had started. Remove that scan's killing consumer; the token scan stays for the reservation probe, and recorded owners are still stopped by identity during recovery. * fix(codex): remove the token-scan kill path from app-server teardown Every descendant inherits the spawn token, so killing each pid that carries it can reach processes the agent started that are not the provider. Production never injected this path; teardown always uses the process-group and descendant-snapshot proof. Drop it, its deps, and the now-unused spawn-token argument. |
||
|
|
153d3fd3fa |
feat(native-chat): Codex sessions write their subagents into the host status store (#22553)
* refactor(native-chat): the Codex acquire names its turn-boundary methods as a set Behavior-neutral: the same two methods stamp receipt time. Keeps the file under the size limit once the child-work sink lands. * feat(native-chat): Codex sessions write their subagents into the host status store A Codex child thread and each persistent command become host child records, fed through the same delivery, ingest and reducer the Claude lane uses. The child's own turn decides it: turn start is live, turn completion settles it with the outcome Codex reports, and a follow-up turn reopens the same record as a new run. Its open tool call, last message, usage and waiting-on-user flag come from its own thread's frames. A parent turn ending settles nothing. * fix(native-chat): close a Codex child's tool call by its item id alone A completion frame need not restate the tool it ran, so reading the tool name before closing left the call open and the record naming a finished tool. * test(native-chat): pin the Codex child-work evidence and every hop to the host's records Child turn start/end/follow-up, open tool call, last message, usage, waiting, the persistent command a child owns and its monitoring display, a primary turn end settling nothing, and session end. End to end through the real adapter: evidence after the journal and the legacy republish, and the parent state the records imply equals today's at every frame of a scripted session. Through the production runtime: a Codex session's child work reaches the status sink under its own address, and a provider exit ends it there. * test(native-chat): a Codex child's new run never inherits the last run's open call * test(native-chat): a Codex session with no child-work sink holds no evidence * test(native-chat): deliver a Codex child's announcement twice, as Codex does, before counting edges * refactor(native-chat): hand the Codex producer's pending edge over directly * fix(native-chat): name every Codex turn state in the outcome map; type the runtime test's fake opener * fix(native-chat): a Codex child's turn ends on the error that ends it, or on its thread closing Codex can end a child's turn with no turn/completed: an error it will not retry is that turn's own end (the verdict the transcript already settles the same turn on), and a closed thread ran its last turn. The executions, the one owner of child turn state, now end the turn on both, so the strip drops the child and its record settles (failed, or unknown for a close) together, instead of reading working for the life of the session. A systemError status is not an ending: Codex raises it for errors that leave the turn running. A child fact whose frame names no turn now belongs to the turn the child is running, instead of counting for every run. * test(native-chat): a Codex child's turn ending by fatal error or thread close settles strip and record together * test(native-chat): the Codex parity script reads a waiting child through the shared fold's waiting arm * test(native-chat): a Codex child row's journal attempt is its record's generation The journal numbers a Codex child's runs by the turns it observed on the child's thread; the host record numbers them by the runs its evidence opened. Both are keyed by the child's own turn id, so they must agree run for run, including when Codex reports the child's first turn before the spawn that announces it. * test(native-chat): a Codex session's end settles its live children and keeps the ended ones The host no longer erases a session's children when its provider goes away: a child still running settles with an outcome nobody reported, and a child that had already ended keeps what it said. The producer tests now expect exactly that, from the close path and from an unexpected exit. * fix(native-chat): a Codex subagent's shell is its open tool until the process exits Codex runs every agent shell through unified exec, so every subagent shell arrives with the source the persistent-command tracker keys on. The producer skipped those items, so a working subagent never named its shell, and an approved command (started on the approval path, completed from unified exec) stayed its open tool until the turn ended. The tracker still records the process separately, so a command that outlives the turn reads as monitoring. * fix(native-chat): a Codex shell becomes a subagent's own work only once it outlives its turn Codex runs every agent shell through unified exec and never says when one is left running, so the producer turned every shell, even a millisecond `rg`, into a command record the moment it started. Each settled into the session's pool of 32 settled records, so a busy turn evicted a finished subagent's record (its outcome row would vanish) and listed dozens of finished shells beside it. A command now becomes a record at the first turn boundary of the thread that launched it while its process still runs: until then it is the agent's open call. A shell that exits within its turn never becomes a record. * refactor(native-chat): child records keep every settled child and can be removed outright Settled child records now stay until the host drops the session's row; the 32-record trim is gone. A producer can say work stopped with nothing to report, and its record (and the handles it answered to) goes instead of settling. Evidence stays host-internal: the producer and the store share one process. * fix(native-chat): a Codex command is live work from its start until its process stops The command tracker is now the one owner of a Codex command's lifetime. It admits every command whatever `source` Codex tags it with (the approval path starts one as `agent`), and ends it when its process exits, when its thread closes (Codex stops the processes first, so no exit ever arrives), or when the session ends. The producer mirrors that one-to-one: a live record from the start, removed when the command stops, never settled. This removes the turn-boundary rule: a command that was only recorded at its turn's end left the parent reading done for one publish when the main agent's turn ended with a shell still running. The parity script now checks the parent at every journal write, not only at frame end. * fix(native-chat): a Codex command whose approval its turn abandoned never ran Codex starts an approval's command item before it asks, and when the turn ends with the question unanswered (the user stops at the approval), it drops the question and never completes the item. The command tracker admitted that start as a running process, so the strip kept a phantom command row and the session row read working until the session ended. The prompt registry, which owns which approvals are still unanswered, reports the command approvals a turn ended without; the tracker ends those commands with the frame that ended the turn. An answered approval keeps its command. * test(native-chat): start the Codex child-work runtime test without the removed hold Main no longer has host.hold: creating the session starts its child, and nothing a viewer does keeps it running. The test attaches and asserts the one child that attach started, then drives it as before. |
||
|
|
7a24d3d335 |
fix(native-chat): the conversation outlives its agent (#22835)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * fix(native-chat): name a chat write by its target, not the owner generation A write carried the fence of the last frame the pane read, and the host refused it unless that fence was still current. An idle release and the restart after it each move the fence, and the release publishes nothing, so a send after a release was refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a cold start was refused as stale. Every write already names what it acts on: a send its conversation, a cancel its turn, a prompt answer its item revision, a rewind its epoch; an option is last-writer-wins. So admission stops comparing the client's fence, and the rebase that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it. The writer-lease check stays, and so does the attach's compare-and-swap. Frames now stamp the fence read when each frame is sent instead of a copy each subscriber kept, which went stale on the same release. * fix(native-chat): every journal append reaches the chats that are open A journal write and its delivery to open readers were two calls, and some writers made only the first. A failed start whose lease could not be handed back, a provider revision with no frame behind it, and eviction's settlement were all journaled without reaching an open chat. A journal handle now reports every durable change, and the host's session map binds that report to the session's readers when the handle is set. Writers no longer publish what they append; the per-writer publish calls are deleted. * test(native-chat): an epoch replacement reaches the open chat * test(native-chat): each row reaches an open chat once, and a live handle enters only through the map * test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite The seeded record had no surface tab id, so the next open backfilled one and that rewrite alone made the no-op transaction write. The test passed with the legacy-lease rewrite signal removed. * test(worktree-activation): restore the OMP surfaced-agent resume test The handoff removal deleted it alongside the terminal-owner tests, but it covers the surfaced-PTY block that still guards resume, including an agent whose ownership is unknown. * perf(native-chat): a publish behind a delivered commit reads nothing Each commit now delivers itself, so the publish a provider frame still sends afterwards found every reader caught up but still read rows and rebuilt the timeline for each one. A caught-up reader now skips the read. * test(native-chat): state why the teardown test's fake journal is safe to cast * docs(native-chat): say mutation admission checks only the writer lease * docs(native-chat): drop the send rebase from comments that still described it * fix(native-chat): a message is accepted, then delivered A send to a chat with no running agent restarted the agent inside the send call, before the message was recorded, so the client waited for the whole start and a failed restart refused the message. Claude held prompts sent during startup, and those could settle as "unconfirmed". A send is now accepted inside the session's serialized queue: one ledger row and one submission row marked handoverRecorded, published, answered pending. A per-session delivery loop exists while a message is queued. It starts the agent through the same serialized attach a hold uses, waits outside the queue for a Claude child to prove its start, and hands the oldest queued message over as its own serialized step, writing dispatch{pending} before the adapter call. A start it needed and did not get writes one error-tone row and rejects every queued message with the same words; a start Stop cancelled writes none. Settlement follows from the rows. A queued message is provably unwritten, so a close, an eviction or an exit rejects it. A handed-over message stays in doubt. A queued row at or below the sequence a handle found when it opened was left by an earlier process and is rejected at open, with no latch. Stop withdraws queued messages with no writer lease and no fence. An attach failure keeps the conversation open, and the attach adopts its journal. Owed work counts the loop and queued rows. A compaction or rewind found prepared when a conversation opens was started under a child this process no longer has, so the open settles it rather than leaving it to refuse every send until a view attaches. The open cursor is scoped to its epoch, because sequences restart when an epoch is replaced. Deleted: restart-before-admission, recordFailedRestart, the fence rebase, Claude's startup gate, the attach's forget on failure and its own crash boundary. Clients without agent-session.accepted-send.v1 get their reply held until the handover; the desktop and paired desktop lists advertise it. * fix(native-chat): settle queued messages only for the child that ended A child that proved its start and then exited before its message was handed over left the message queued: the exit settlement returned early when nothing else was in flight. Delivery then started another child for it, and a child that died the same way started another, without end and without a row. A retried settlement for an earlier generation, run by the attach that delivery started, did the opposite: with that generation's turn unfinished it rejected the message queued for the child being attached. The settlement now takes the rejection for queued messages from its caller. The unexpected exit and the eviction pass one, and it applies even with no other work in flight; the retry for an earlier generation passes none. * fix(native-chat): an adoption that fails to import keeps the conversation open The attach now writes into the conversation's own open journal, but a failed transcript import still closed it as if it were the attach's provisional one. The conversation stayed indexed with a closed journal, so every later send answered "could not be recorded" and every attach failed again until the app restarted. The import now closes only a journal the attach opened for itself. * perf(native-chat): the recovering open reads the journal once Every conversation open now goes through the recovering open, including the read restore of every chat at startup, which used to replay its journal once. The recovering open replayed it twice: once to probe it and again inside the open. The probe is now handed to the open as its load. * fix(native-chat): an attach that fails after indexing its child leaves no child behind A failed attach now keeps the conversation open, but a failure after `onAttached` indexed the child (the rewind or compaction recovery, or the attach's own success record) left that entry claiming a child the failure path had already released. The next send found the phantom, skipped the start, and wrote at a fence the journal had moved past, so the message stayed queued for good. The entry now drops the released child and its event sink, and follows the record's fence, as a failure before indexing already did. * fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer The error strip for a message the host accepted and then did not deliver matched the entry before the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not send your message" with nothing to retry. It now reads the reconciled entry. A rejection the journal records before the send's own pending answer lands is final as well: that answer no longer puts the entry back to dispatching with no Retry. * fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down The preamble waits for its submission to be delivered while the worker's agent starts. When that wait ran out it threw operation_unknown, and the failed-start teardown then closed the session, which rejected the very preamble the host was about to deliver. It now reports a turn start nobody observed yet: the worker is start-unknown with its session kept, the host delivers the preamble when the agent starts, and the worker's report settles the dispatch as for any unobserved start. The receipt no longer suggests reading a screen a structured worker lacks. * fix(native-chat): a message rejected while its chat was closed reads as not sent A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked every later message behind a Retry and no reason, and the delivery probe, seeing the journal already answered, never ran. The reconcile now settles it as rejected like a dispatching one. * test(orchestration): name why the readiness settlement fakes are cast * fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent * docs(native-chat): drop the fence from the admission the send effects run behind * docs(native-chat): give the fence move on release the reason that still holds * docs(native-chat): stop citing a write fence check in launch and mailbox comments Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease. * refactor(native-chat): the provider child is its own record A conversation now outlives any number of provider children, so the child is one record on the conversation's entry instead of five loose fields beside its journal. It is written in one place: indexed only once an attach has fully succeeded, and ended through one function that an exit, a failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence. - A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence patch after it are gone. - Conversation writes read the record's fence, the way mutation admission already does; a child's own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of the conversation's fence, are gone. - The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer dropped when an attach replaced the whole entry. - Stop on a child still proving its start stops only the child: its lease goes back and the chat is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus the conversation's close. - The settlement retry uses the conversation's own journal, opened through the host's one open. * fix(native-chat): the delivery loop alone settles a message its start or child failed A queued message was settled by whichever path happened to end the child first: the loop, the unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that rejected every pending row. That gave two failure rows with different tones for one start, a loop that could hand over to a different child than the one it waited on, and a Claude start that died while starting reading unlike every other failed start. - The loop remembers the child it waited on. At handover, if that child is gone or replaced, it reads how it ended: a Stop continues; anything else writes one failure row and rejects every queued message with the same words, then stops. A child still starting whose start the adapter says did not land fails the same way. The exit, eviction and the settlement retry only settle the handed-over and legacy rows of the child that ended. - One failure row, always an error, keyed by the start. A start a view began that dies with nothing queued writes the same row through the same builder, so a second report revises it. - The open no longer rejects leftovers; the loop's first step does, and the open wakes it. - `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the failure before the exit is processed. - Quit closes every conversation the way closing a chat does: what is still queued is rejected as closed, with or without a child, and a start the loop already has in flight is waited for so the child it produces is stopped rather than left behind. * refactor(native-chat): a stopped child ends on the one reading of its stop The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that verdict to the child's ending, so the host never forms a second view of whether the root is gone. Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition, and a failed re-attach passes what its release saw. The end-of-child record can therefore also carry a stop whose root was not seen to go, which nothing ends on yet. * feat(native-chat): the host says it accepts a send before any agent has it The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same string capable clients already send. A client can then tell a host that answers a send at acceptance, and admits a Stop with no writer before a turn starts, from an older one that still restarts the agent inside the send. Additive: an older client ignores a capability it does not know. * refactor(native-chat): an attach never opens a journal of its own The attach adopts the conversation's open journal, which outlives it, so it no longer opens one for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag that told the two cases apart is gone. Tests that attach without a host open the conversation the way a host does. * fix(native-chat): a moved fence resends nothing on a host that accepts first The outbox treated any fence change as a new owner: it dropped the answer of a send in flight, queued that send to go out again under the same id, and unblocked a refused head. On an older host that is how a send the restart refused, unrecorded, gets another try. On a host that records every send before it starts an agent, a fence moves because that start ran, so the same rule resent into every failed start. With a fence stamped on every frame, that became a loop. The outbox now reacts to a fence change only when the host has not advertised that it accepts a send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed start reaches the client as a rejected message it keeps with its Retry. Against an older host, or before one has answered, the outbox behaves as it did. Desktop and paired web share this hook. * refactor(native-chat): a child's end says whether the user or the host stopped it The end-of-child record's cause now tells a user's Stop from the host stopping the child for a cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a user's Stop, as before, and fails the start it was waiting on after a host stop, with the one error row and every queued message rejected, in the stop's reason when it gave one. The reason stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet. * fix(native-chat): a chat whose only work is a queued message is not offered for resume A message accepted while the agent was starting counts as working in the chat, and quit rejects it as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a chat whose agent never had the message. The snapshot now reads only what was handed over. * fix(native-chat): the conversation outlives its agent Opening a chat no longer starts its agent. A conversation is reached through one host accessor that opens its journal at rest, and a send is what starts the agent, through the delivery loop. One idle sweep, every five minutes, stops an agent that has been quiet for thirty minutes and owes no work, then drops an open journal handle that is only a cache. Its record, tab, status row and readers stay. - hold and release are no-ops; hold still builds the host for shipped mobile builds. - The holders, the holds, the release clock and the exit respawn are deleted. - Options, the model list, the goal and the context meter answer at rest; a model pick at rest is recorded as intent for the next start. - Compact, rewind, clear and goal changes start the agent first. A send does too when a rewind is still in doubt after the conversation opens. - Orchestration routes mail and group addresses on ownership (the record plus the chat tab), not on whether the process runs. An open dispatch keeps its worker running. - The restart continuation is a send; Resume all holds each slot until the message is handed over or rejected. - A read error never replaces a loaded transcript, and shows the host's own words. * test(native-chat): type the queued-message fixtures in the resume-offer tests * fix(native-chat): a start that dies while a message waits on it is that message's failed start Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When that start died, its exit wrote the start's error row and left the message queued, so the delivery loop started a second agent into the same failure and wrote a second row. A child's end now records where the conversation's journal stood, and the loop settles a message accepted before a failed start ended with that start: one row, under its key, and no second start. A message sent after the failure still gets a fresh start. * docs(native-chat): say what an attach's open conversation and unconfirmed ids are now * test(native-chat): pin what a failed start settles, and what a resume offer names A view's child that dies while a sent message waits settles that message only when it died starting and no child has taken its place: a proven child's crash, or a second start since, gets the message delivered. The resume offer names the handed-over message, never a newer one still queued. * test(native-chat): the failed-start pins fail on what the message became, not on a timeout * fix(native-chat): a restart offer ends when the chat's agent starts again The offer used to end only when the chat's newest user message changed, because opening a chat started its agent and that start could not be told apart from real activity. Opening a chat starts nothing now, so the host reads the fact it already publishes: a chat's status row goes from not host-owned to host-owned exactly when its agent is started. At that edge the offer and any failure record for the chat are withdrawn, unless the start is a resume action's own (its continuation is the oldest undelivered message). A continuation and a message racing to be first are decided at acceptance: the continuation is refused, quietly and with nothing filed, when any other message was accepted since the restart. A failed continuation start leaves the offer retryable, and each resume action sends its own message id. Deleted: the newest-user-message comparison, its journal reader, the continuation filter, and the failure ledger's own "answered by the chat" check. The marker still carries its message id for one release, so the previous build can read it. * fix(runtime): end a transcript stream when its client unsubscribes Desktop: the IPC subscription controller was dropped as soon as the streaming handler returned, which for most streams is right after it binds. A later runtime:unsubscribe then found nothing to abort, so the host kept the subscriber and derived and sent every publish to a channel no one listened to. The controller now lives until the renderer unsubscribes, resubscribes the same id, or goes away. Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe with the stream's frame id, so the host ends that subscriber and leaves a sibling stream on the same socket running. The direct path now passes the frame id the relay path already passed. * test(orchestration): the preamble's host stub is typed, not cast The preamble send now takes only what it reads of the host, the send, the settlement wait and the record's fence, so its test builds that host with real types instead of `as never`. * fix(native-chat): one fact ends a restart offer: the chat moved on since the restart The offer is live while no other message has been accepted in the chat since the restart and its agent has not proved a start since. The offer list, the resume's reservation check and the continuation's acceptance check all read that one fact, so a message whose start then failed withdraws the offer too, and a stale click finds nothing to act on. The fact is read off the conversation's open handle, which the restart closed, so it is retired durably whenever it may have changed: a message accepted, a start proven. A close and reopen within the same run therefore cannot bring the offer back. A continuation rejected before it reached the agent does not count, so a retry after a failed start still runs. Deleted: the quit-time gate on withdrawal, which changed nothing because the withdrawal and the quit's own offer write share one queue; the per-action "withdrawn" flag and the separate acceptance check it paired with. * test(native-chat): an older build reads the restart offer this build records The offer lives in a file the previous release reads after a downgrade. Pin that against the pinned release's own capsule, and run the lane when the marker or the capsule changes. * fix(native-chat): read a restart offer against where the journal stood when it was taken "Since the restart" was read off the conversation's open handle, which the idle sweep closes: after a reopen, a message the user had already sent looked older than the handle and the withdrawn offer came back. The offer now records the journal position (epoch and sequence) at the moment it is taken, and a message accepted after that position, or a journal on another epoch, means the chat moved on. That is derived from the journal, so it holds across any number of closes and reopens. An older build's offer has no position; only a start withdraws it. Because the message half is now durable, the offer is no longer rewritten in the recovery file on every accepted message; a proven start still writes it, since only the host that saw the start knows of it. * test(native-chat): wait for the listing's retire write before reading the recovery file * fix(native-chat): keep the terminal-backed chat's read error over its local echoes Messages winning over a read error is right for the structured chat, whose read retries and whose messages came from the transcript. The terminal-backed view assembles its list from local echoes too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no error. Only the structured pane now keeps messages over an error. * fix(native-chat): a start retries the exit settlement a failed journal write left owed An agent exit whose journal settlement write failed releases the lease latched until a retry lands. Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried it before the next app launch, and every send was refused. The start the send needs now runs the retry first, where the attach would. * perf(native-chat): answer the owner check without opening the chat Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the answer comes from the session record alone. Reaching it through the accessor opened each resting chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it open for the idle window. It now checks the record and the adapter's support, as before this series, and opens nothing. * fix(native-chat): a read waiting on the session lock opens nothing once quit began The accessor checked for quit before queueing the open, so a read queued behind a session task ran its open after teardown had begun and indexed a journal no teardown step would close. The check now runs at the open itself. * fix(native-chat): read a failed resume's chat before calling it retryable Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the restart, read from its journal. The failure list read it only for a chat already open, so once the idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did nothing. The list now opens the failed chats first, as the offer list does. * test(native-chat): type the provider event sink the settlement test reaches for * fix(native-chat): say the structured read keeps trying only where it does The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an untranslated fallback whenever the read error had no text, and the empty state prefers any message. The view state now leaves the message out, so the structured pane shows that line and the terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an error frame, so it no longer makes the claim. * test(native-chat): await the send's settlement instead of polling for the start The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a loaded machine outran. They now await the host's own settlement of the message. * fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now dated by the resume action. Telling a rejected continuation from the user's own message read the operation ledger, whose rows expire after about a day; after that a failed resume stopped being retryable. The offer now records the continuation each action sends on its own capsule entry, bounded to the newest 16, so the ids end with the offer. The ledger read is deleted. * fix(orchestration): route no mail to a structured worker its orchestration released A structured worker is routed on ownership, and a resting worker's lease is released, so ownership held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one restarted its agent. Routing now also reads the orchestration's own resource row: once it is released, direct mail, group addressing and worker-show's addressable answer drop the worker, as they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored. * fix(native-chat): a failed retry names the user's prompt, not Orca's continuation A resume's continuation is written to the chat before its start, so after a failed attempt the chat's newest user message is that rejected continuation. A second failure then showed Orca's own restart text as the chat's prompt. A retry now keeps the prompt its first failure named. * fix(orchestration): read the released row optionally, as the authority does worker-show's observation called the row lookup directly, which a runtime double without it threw on and failed the structured tab-retirement release. * fix(native-chat): the status bar drops a restart offer the chat moved on from The renderer re-read the host's restart offer only when a failed chat showed activity, so after a message withdrew a pending offer the host answered no chats while the status bar kept counting one, and clicking it opened nothing. The same watch now covers pending offers: a status change in an offered chat asks the host again, once. * test(native-chat): a roster of idle or finished children does not keep an agent awake The sweep reads owed background work through the shared child-work liveness that upstream's release clock adopted; a child that went idle or finished is not work the agent still owes. * fix(orchestration): a task dispatched into a resting structured worker keeps it running The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's process incarnation now counts, derived from the existing rows. * docs(native-chat): comments stop describing the hold this PR removed Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it. Comment-only. * fix(native-chat): a restart offer keeps the start its own continuation made Whose start ended an offer was decided at read time, from whether the offer's continuation was still the queued message. Once the provider refused that continuation, the child it had started read as someone else's start, so the offer ended and its failure showed no Retry. The delivery loop now records which queued message a start is for on the in-memory child, and the child's end carries it; the offer counts a start as its own when that message is one of its continuations. * fix(native-chat): an agent gets a full idle window after its owed work ends The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can read done before the lead's wake-up turn writes anything, and stopping in that gap loses the wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full window afterwards, as the release clock it replaced did. * test(claude): the options-read fixture runs a live child The fixture marked its conversation running with a hasProviderChild field the session type does not have, so the read took the at-rest path and refused a session with no record. It now carries a child, which is what the read checks. * test(native-chat): host tests reach its collaborators through a typed seam The rest-test rig and three test files read the host's private members with Reflect.get and cast the result. The host now exposes one test-only accessor, collaboratorsForTests(), and the subscribers class a subscriberCountForTests() beside its existing retainedActivityCountForTests(), so the tests are checked against the real types and the casts are gone. * refactor(orchestration): one owner answers a structured worker's custody Routing, group addressing, worker-show and the idle sweep each composed their own reading of whether orchestration still holds a structured worker, so each new obligation or retirement state had to be added to every reader. structured-worker-custody now derives both answers from the worker-terminal list state coordinators see in worker-list: addressable is owned and not released, and owed work is an active custody or an unsettled task dispatched to the same incarnation. The owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests. * refactor(orchestration): owed work is an open dispatch on the worker's incarnation A supervised worker's own dispatch context stays open exactly while the worker is active, so the separate active-custody branch only repeated it. Owed work is now one fact, which also states the policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are written once at the top of the module. * fix(native-chat): a restart offer knows its continuations by a tag in their id The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running action's id in memory. Both could disagree with the journal: past the cap an old rejected continuation read as the chat moving on, and a crash during a retry restored the failure's older entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer (its teardown and chat), then the action's own part, so any continuation of this offer, queued or rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own continuation the start was for, read against the stored marker. * test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait The tui-idle probe reads through readTerminal, which now awaits the structured worker check before the PTY read, so the probe's snapshot request starts a microtask later. vi.waitFor missed it on its first check and polled again at 50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot then resolved after the wait had already timed out, so the test passed without judging it, and the rejection landed before any handler was attached. Vitest reported that as an unhandled error and failed the shard. Polling every 1 ms sees the request within a few ms, so the snapshot is judged while the wait is still pending. * fix(native-chat): a view never restarts a chat whose last start failed A Claude chat whose CLI exits during startup left one red row per start, and every time a view bound to it (the chat opening right after its create died, or the user switching back to it) the hold started the CLI again, so the same launch-failure row repeated. Only a send retries a failed start now, the same rule provider-exit recovery already applied; the rule lives in one predicate the hold, exit recovery and the delivery loop share. * fix(native-chat): the idle sweep reads owed work every tick Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it refreshed the clock at most once a window. Work that ended just before the next read left the agent to be stopped at that read, moments after the work ended, which is the gap the refresh was meant to cover. The sweep now reads owed work on every tick for a started agent, so the window always runs from the last tick that saw work owed. * fix(native-chat): a continuation handed to the agent stays sent The offer read its own continuation as not reaching the agent while its dispatch was pending, which also covered one already handed over and still unanswered. When the wait for that answer ended first, the failure it filed read as retryable, and a retry sent a second continuation to an agent that may have acted on the first. Only a continuation still queued, or rejected, is now read as unsent. * test(native-chat): start the child the loop waits on with an attach, not a second view A view no longer starts a child whose last start failed, so the R2 case that waits on a child started since the failure now gets that child from a client attach, the one non-send starter left. * fix(native-chat): settle a gone generation's turn wherever a conversation opens A send that opens a chat this process had not read yet (after a crash, from a phone or the CLI) went through the delivery open, which never settled what the dead generation left running; only the read restore and a successful acquire did. When the send's start then failed, the turn stayed running for every reader. The settlement now runs in the one journal open, at the crash boundary, for every opener except an acquisition, which settles from the evidence it read before its reserve; the read restore's separate step is gone. * test(native-chat): prove the next child's start settles the turn an earlier child left The R1 case lost its only settlement assertion when the latch it checked was deleted. It now seeds the running turn the earlier child left and asserts it ends at the exit's receipt, with the exit's row, before the message is handed to the new child. * test(native-chat): count a failed start's rows by row, not by text Comparing the set of texts passed when two different rows carried the same words, which is the duplicate the test exists to catch. * test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget * test(native-chat): the interrupted create's own retry continues again The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its own operation id, with a fresh start whose result nothing read. That fresh start passes with the released-reservation continuation deleted, so the case the fix exists for went untested. The retry and its assertion are main's again. * docs(native-chat): three comments that still had views starting agents A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction left alone would refuse every send, so no agent would ever start to finish it; and a current host raises the unattached read refusal only once quit began, with the attach window belonging to an older host. * test(native-chat): pin the open's and the send's start and row counts, however the view binds Opening a fresh chat whose starts fail makes one start and one row, with two views bound before or after the create's child died; one send makes one more of each. * test(native-chat): a reader's open settles the turn a failed exit settlement left running An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle sweep closed, and a read that opens the chat before the restart restore reaches it. * test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner A subscription reads the conversation before it returns, so under load the two views took longer than the create child's 300 ms start, which then exited before the test checked that it had not. The child now takes a second to fail. * fix(native-chat): settle a gone generation's turn at every open but an acquisition's The journal open skipped the settlement whenever the lease read reserved or live, to leave an acquisition's own open to the acquisition. But a lease a crashed process left in recovery also reads live, until the next acquire resolves it. A send that opened such a chat, from a phone or the CLI after a crash on a host that could not prove the old owner gone, skipped the settlement; when its start then failed, the dead turn stayed running for every reader. The acquisition now says it is the opener, and every other open settles, whatever the lease still claims. * test(native-chat): hold the create's start open until the views bind The "view binds while the create is still starting" case gave the create a 300 ms head start and asserted the views bound before it died. On a loaded runner the holds took longer, the create's exit landed first, and the case failed its own precondition. The create's initialize now waits on a gate the test releases once the views are bound. * test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the dead process's lease still reads live. The open settles the turn it left running anyway, and the restore that follows finds it settled. * refactor(native-chat): drop the composer's second error formatter After the merge with main, every chat write in the composer path reports its failure as a typed outcome worded by the refusal-notice table, so the send's catch sees only a local throw. The {code, message} formatter this branch added for it has no payload left to format, and its claim to be the one way a chat words a failure is no longer true. The composer send is main's again. * test(native-chat): pin the reason on a message rejected while its chat was closed The reopen test checked only that the message reads as not sent; it now also checks the Retry row carries the host's reason. * docs(native-chat): drop the removed dispatch hold from six comments A worker's session no longer takes a dispatch hold, and no release clock rests a chat by visibility; the agent-launch comments, the abandon test, the teardown test and the refusal census still said so. * test(native-chat): rest the owner-status chat through the idle sweep, not a hold The activation-gate test from #22808 put its chat at rest by holding and releasing it, and passed the release-clock grace. This branch deleted both, so the case threw before it reached its assertions. It now moves the host's clock past the idle window and lets the sweep stop the agent and close the conversation, then asserts the same owner answer and activation gate. * fix(native-chat): show the structured pane's retrying line when a read fails The read transport always hands the pane the host's words, so the error state's "Orca keeps trying to load it" line, which showed only when there were none, was never seen: the pane showed the host's text twice, as its subtitle and on the status line under it. The structured pane now always says its read keeps retrying, and the host's text stays on the status line. The terminal-backed chat is unchanged. * test(native-chat): wait for a send's background start before the refusal oracle removes its store An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals. * fix(native-chat): a start a message waited on gets one failure row, the delivery loop's When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice. The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
acf8e679ea |
feat(native-chat): Claude sessions write their subagents into the host status store (#22536)
* refactor(native-chat): the host hands out client delivery's status subscriptions as they are subscribeStatus and subscribeTurnCompletions wrapped client delivery's bound methods in forwarding lambdas; they are now the same members, the way waitForSendSettlement already is. The host is at its size limit, and the next channel it hands out needs the line. * feat(native-chat): Claude sessions write their subagents into the host status store The Claude background-task tracker queues child-work evidence at each decision it already makes (start, update, progress, terminal frame, roster replacement, turn end, session end), plus the two facts its legacy row ignores: a foreground child's progress and a foreground spawn call's result. The adapter drains that evidence after the journal handled the frame and the parent row was republished, and the host folds it into one record per child in its canonical store. Nothing reads the records yet; the strip and sidebar keep their current sources. * test(native-chat): pin the Claude child-work evidence and the host reduction of it * test(native-chat): prove every hop from a Claude frame to the host's child record The adapter delivers evidence after the frame's journal rows and the parent's republished row; the frame script keeps the parent state today reads while the records add outcome and activity; the runtime hands the evidence to the status sink under the session's own address; both entry points wire the sink to the ingest. * test(native-chat): read an optional task list as optional in the producer script * test(native-chat): an address whose publish threw carries no child work * test(agent-status): a foreign record differs from ours by producer alone * feat(native-chat): a foreground Claude child's own tool call is what its record says it is doing A child's tool traffic reaches the parent stream only for a foreground child. Read after the journal handled the frame, the child's newest call still awaiting a result becomes its open operation, previewed the way a hook-reported row previews its own tool; the result closes it. The open call is derived from the journal's own bookkeeping, not held a second time. * fix(native-chat): a Claude child restarted under a new spawn call keeps reporting to its record A task that ended and starts again stays hidden from the legacy row until a roster lists it, so the tracker held no run for it: the new run's progress reached nothing and a foreground re-run's own spawn result settled nothing. The run is now held beside the live map, where the legacy row never reads it, until a roster hands it back or it ends. A parity test pins the record's run count to the journal roster's attempt on a new spawn call, the one event both count. * refactor(native-chat): the Claude child-tool queries and translator contract get their own homes The translator's child-tool queries move into claude-child-tool-queries.ts and its contract type into claude-journal-translator-contract.ts. Brings the translator back under the size limit. * refactor(native-chat): Claude child evidence carries only its own edge's facts Admission now keeps what a child's record already knows: labels, model, owner, residency, the last message within an invocation, and a token count that never shrinks. The evidence side copied all of those forward itself, a second owner of the same rule. It now sends only what this edge observed, and a task's token count comes from the frame that reported it. * refactor(native-chat): Claude child evidence hands admission its raw labels Admission now folds provider text to one line and drops a malformed fact instead of refusing the record, so the evidence side no longer folds labels itself. The description keeps admission's longer bound. * fix(agent-status): admission alone decides a settled child's second ending The reconciliation returned before admission whenever a record had already settled with a definite outcome. That dropped the evidence an `unknown` ending carries (its last message and tokens), which admission's refine-only rule keeps, so that rule never ran for the structured producers. The latch goes. Admission keeps the definite outcome, lands the late evidence, and refuses a conflicting definite ending as `stale-invocation`, which the host ingest already counts as the fence doing its job, not a fault. Pinned through the real Claude producer and the host's own ingest. * perf(agent-status): keep child records off the status hot paths Child records made every store write and every status notification scale with the whole store. Each Claude child progress frame cost about 2 ms with 5 chats holding ~200 child records (about 14 ms at ~1,400), and every status change on any lane re-parsed every child record just to list parent rows. - The store derives each frozen record's key once instead of re-parsing it on every mutation's validation and every alias lookup. - Settled history is trimmed only when a batch settles something. - Parent listing and the structured row's revision stamp read the parents and the revision directly instead of building a full snapshot. A progress frame now costs about 0.3 ms at the same size, and listing parent rows no longer depends on how many child records the store holds. * fix(native-chat): an errored Claude spawn result no longer decides how the child ended Interrupting a foreground Claude agent while it runs a tool delivers the spawn call's errored result before the child's own killed/stopped frames. The spawn result settled the record `failed` first, and admission then refused the later `cancelled` as a conflicting ending, so an interrupted child read as a failure. An errored spawn result now settles the child `unknown`; the child's own terminal frame refines it to `cancelled` or `failed`. A successful spawn result still settles `succeeded`. The test replays both frame orders the real CLI produced when interrupted. * test(native-chat): pin a Claude foreground child's real finishing order The real CLI ends a foreground agent with its own completed update, then a notification carrying the final summary and usage, and only then the spawn call's result. Existing tests modeled the spawn result arriving first, so nothing checked that the notification's summary and tokens still land on a record the update already settled. * perf(agent-status): a store write costs what it touches, not the whole store With child records on the host, every mutation copied all five store maps and re-validated every record, and reads scanned every child and alias. A parent status publish cost about 10 ms with 4,000 child records in the store, and a child update about 13 ms. - A mutation writes into drafts over the committed maps and lands in place; a refused one is dropped with nothing to undo. The drafts keep the exact map order a copy would have. - Only what a mutation touched is re-validated: touched parents, children, aliases, facts and tombstones, plus every alias of a touched child and whatever a removed parent owned. The full validation stays for snapshot restore. - The snapshot byte budget is a running total instead of a re-measure. - Children by parent, facts by parent, aliases by child, aliases by identity and retired aliases are indexed, so reads return stored records without scanning or re-parsing. - The memoized alias identity and tombstone-key checks are gone: indexes derive them once. A parent publish now costs about 0.015 ms and a child update about 0.06 ms at 40, 1,000 and 4,000 children alike. A seeded fuzz holds the store to the copy-and-validate-everything path decision for decision, snapshot for snapshot and read for read, and a replica fed the envelopes ends identical. * fix(native-chat): a Claude child ends only on its own terminal frame The child records were fed from the legacy background-task tracker's display decisions, so they inherited rules that are not truth: a turn ending swept foreground children, a roster omitting a background child settled it, a foreground spawn call's result ended the child, and a new background start after any roster produced no record. Captured from the real CLI, an agent moved to the background keeps its own shell running for 40 s after the parent's turn ends, and that shell was settled `unknown` at the parent's `result`. Replayed with the spawn result ahead of the roster, the same agent settled as a false success and was then revived as a spurious second run. A new decoder reads the task frames directly. `task_started` opens a child (a start for an ended task id is a restart, the way messaging a finished agent resumes it), progress and a live `task_updated` update it, and a terminal `task_updated` or `task_notification` ends it. Rosters, turn ends and spawn results say nothing about a child. Every child in every capture gets its own terminal frame, so no evidenced ending is lost. The notification's `tool_use_id` names the run that ended (captured on a resumed agent's second run), so an ending from a run that is already over no longer ends the current one; a run id the record never saw still ends it, so nothing strands. The tracker, its settled-task retention and the frame readers are back to exactly what main has: the aggregate-roster split and the restart holding map are deleted, and the legacy row is unchanged by construction. * fix(agent-status): a session's end settles its live children instead of erasing them When a structured session ended, the reducer removed every child record it held, finished or not, so a reader could no longer tell how the session's work had ended. Now a child still live when its session ends settles `unknown` (nothing reported how it ended), and a child that had already ended keeps its outcome. The records still die with their parent: closing or releasing the session drops the parent row, and the store drops its children with it. A child's own outcome arriving after the session ended still refines the `unknown`. The `inventory` and `turn-ended` edges, and the rules that settled children on a roster omission or at a turn boundary, are deleted: no producer sends them any more. A restart is now its own flag on a live edge, which is what a producer reports when a finished child starts again under the same run handle. * test(native-chat): replay the real Claude CLI's frame orders into a real host Scrubbed cuts of five Claude CLI 2.1.280 stream-json captures (ids, paths and prompts replaced, frame order and relative clock kept), replayed through the adapter into a hook server: - an agent moved to the background keeps its own shell live past the parent's turn, and the shell settles at its own notification's time; - the same capture with the spawn result ahead of the move ends the agent once, from its own notification, with no second run; - a roster that omits a background child without its own ending leaves it live; - a session that ends settles what still runs `unknown` and keeps every record; - messaging a finished background agent opens its second run, which ends from its own frame; - interrupts in both captured orders end `cancelled`, and a finished foreground agent keeps its summary and usage. * test(agent-status): hold the store's running indexes and byte total to a rebuild The copying-store fuzz never reaches the snapshot byte budget, so a drift in the running byte total (or any index the public reads do not surface) passed it. After every fuzzed step, including refusals, compare every index with one rebuilt from the committed maps. |
||
|
|
d443320af2 |
refactor(native-chat): remove the unused terminal handoff (#22783)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite The seeded record had no surface tab id, so the next open backfilled one and that rewrite alone made the no-op transaction write. The test passed with the legacy-lease rewrite signal removed. * test(worktree-activation): restore the OMP surfaced-agent resume test The handoff removal deleted it alongside the terminal-owner tests, but it covers the surfaced-PTY block that still guards resume, including an agent whose ownership is unknown. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
e0144a9eb6 |
fix(native-chat): show the Codex and Claude model picker the moment a chat opens (#22756)
* fix(native-chat): show the Codex and Claude model picker the moment a chat opens A new structured chat showed no model picker until its session had been created, spawned, initialized and had answered a model listing — and the picker then listed the models a second time. Codex's listing often goes to the network, so the picker took 0.6-2 s to appear. - Keep a host-owned model catalog per agent and account home, persisted on success only and refreshed in the background once it ages out. A new read-only agentSession.modelCatalog RPC answers from it without a live session; sessions reuse it instead of listing again. - Render the picker while the launch is still provisional, showing the saved default. A pick made before the session exists is held and applied once it publishes; only the host's acceptance saves it as the default. - Mark the model and effort set in the user's Codex config as the listing's default (config/read), so the first frame names what the chat will run. - Resolve the account a record-less read would use without running launch preparation, which writes and syncs account state. * fix(codex): disable plugins in the model catalog probe app-server * fix(native-chat): read the host model catalog only for panes on this machine * fix(native-chat): name a pre-report model only for a chat this view launched * test(native-chat): pin the launch latch across publish * test(native-chat): pin a held pick reaching the host before the first send * fix(claude): pin the catalog probe's config dir by the session spawn's rule * fix(native-chat): read the host model catalog only for a visible chat * fix(native-chat): rewrite the model catalog file only when a listing changes * test(native-chat): type-check the first-send order fixture * refactor(native-chat): keep the structured options hook under the line cap * fix(native-chat): send the first turn only after every pick held during launch settles * fix(claude): name no default effort from the catalog probe listing * fix(native-chat): name no listed default model for a chat resumed from history * refactor(native-chat): let the launch own picks made before it publishes A pick made while a chat launches had no fence to go to, so the pane held it and flushed it after publish; every other sender (the outbox, the launch prompt) then needed its own gate to wait for that flush. The launch now keeps those picks in its own state, applies them against the create receipt's fence before it counts as published, and every sender follows publish by construction. The pane flush, the outbox gate and the module-wide held-pick registry are gone. The launch also snapshots the saved selection its create seeds when the intent is built, so a pick in another chat no longer relabels one still launching, and a pick the host refuses is reported the way a refused mid-session pick is. * fix(native-chat): name no default model a workspace's own config can replace The catalog's default is the account's, read without a working directory, but a chat runs in its worktree, where a project config (Codex's .codex/config.toml between the project root and the worktree, or a Claude .claude settings file that sets a model) picks the model instead. The picker named the account default there while the chat ran the project's model. A new chat's catalog read now names its worktree. The host checks that workspace for such config (existence only for Codex, the model key for Claude) and, when any is present or the workspace is not a local directory, serves the listing with no default, so the picker names nothing until the chat reports its model. * fix(native-chat): name the listed default model before the report only for Codex * fix(claude): let an option pick made while Claude starts wait for it instead of being refused * fix(codex): name no listed default when the configured model is not in the listing * fix(native-chat): show the picker as unavailable until a published chat attaches * fix(codex): keep the catalog probe's listing when config/read stalls * fix(native-chat): write the pending model catalog save before quit * chore: drop an unrelated lockfile rewrite * fix(native-chat): rename the catalog store's listing parameter off the global fetch name * chore: drop an unrelated lockfile rewrite * fix(native-chat): name the model Claude will run before its first turn * fix(native-chat): keep Claude's pre-turn applied effort out of the saved session options |
||
|
|
6ae6ed08bb |
fix(claude): open structured chat without a startup deadline, and make Retry start fresh (#22364)
* fix(claude): open structured chat without a startup deadline, and make Retry start fresh
Publish the Claude session as soon as its process is spawned instead of racing
initialize against a fixed 10s deadline. Prompts sent before startup lands are
held and written in order once it does. An exit or sign-in failure before startup
ends the session with the reason and the CLI's stderr.
A create that failed because the process provably exited now carries
ownerVerdict 'exited', so the client marks the launch failed and Retry mints a
new operation instead of replaying the stored failure.
* fix(native-chat): sending into a chat that failed to start restarts it
* fix(native-chat): a send with no live owner restarts it once
A provider child that timed out or exited hands its lease back, and every
later send was refused agent_session_ownership_unknown. Clients read that
code as "not admitted yet" and resend forever, while only a surface hold
could make a new child, once per mount, with its failure swallowed.
The send now routes to a live owner, otherwise restarts one from the
persisted resume state where resume eligibility allows it (single-flight
per session), otherwise refuses with the new settled
agent_session_owner_unrecoverable. Unverifiable, reserved and handed-off
leases are left alone. The desktop hold now logs its failure.
* test(native-chat): pin the unrecoverable refusal as settled in the outbox
* test(native-chat): pin the release clock after a send restarts an unheld owner
* test: read the sent operation id without a cast
* fix(native-chat): type the send-recovery record lookup as the store returns it
* fix(native-chat): a send ensures its owner before admission, and an unheld owner idles for 30 minutes
* fix(native-chat): a create that throws releases its event sink
A child that dies between spawn and journal attach can still write through
the host's event sink, which attach unbound in onAcquiring and never re-bound
because onAttached never ran. The orchestration released that sink only when
performAttach returned a refusal; a thrown failure (the root-exit path) kept
the sink cached with its queued write, so the next attach's drain barrier and
runtime shutdown's flush waited forever.
Also pins the publish-on-root-exit clause for a start that never proved:
deleting it reddened nothing before.
* fix(native-chat): a resend the journal answers restarts nothing, and a send joining a restart rebases from the fence it replaced
* fix(native-chat): the host learns a Claude start positively, and persists only proven options
A publish-first create used to read the session's options before Claude had
answered initialize. With startup pending that read fell back to the built-in
catalog's default, so `record.options.model` was persisted as `sonnet` for
every user whose CLI default is something else; an owner handoff or a reopen
then replayed `set_model('sonnet')` and silently switched their model.
The adapter now reports `started` once startup facts are applied and saved
options restored. The host keeps a `providerChildPhase` on the session it
owns: a starting child hands over nothing but the saved options as intent,
and the `started` event re-reads the options as fact and persists them through
the same record write a user's option change takes. The status summary carries
`hostExecutionPhase` (optional, wire-safe), and the chat pane says the agent is
still starting instead of showing nothing.
A child whose exit already reached the adapter before acquire returns is no
longer handed over as live; the create fails with the CLI's diagnostic.
* fix(native-chat): a hold and a send that find the owner gone share one restart, and a send the ledger already holds restarts nothing
* fix(native-chat): a failed create answers one refusal shape, stamped once at the boundary
A create whose Claude process was seen to exit answered twice in two shapes:
the first call threw a generic runtime error, and only the replay of the same
operation carried the `ownerVerdict: 'exited'` refusal that lets a client
retry under a new operation. Three sites stamped the verdict and the store
failure path stamped nothing.
The first-hand root exit is now returned as the refusal on the first call,
with the provider's own diagnostic as its message. The verdict is stamped in
one place, at the boundary of the attach, from the durable row the operation
settled to, so every refusal shape answers the same fact and no site can
forget it. The per-site stamps are gone.
* fix(native-chat): a send into a session whose child ended restarts it before admission
A session that published and then lost its Claude child before startup (not
signed in, for one) keeps a released lease and a chat the user can still type
into. The send was refused as ownership-unknown, the outbox parked it as
pending admission, and nothing ever restarted the child: the message sat
there until the user closed and reopened the tab.
A send reaching a session with no provider child now runs the same resume a
surface's first hold runs, before the write is admitted. The resume reserves
a new fence, so that send is answered stale with the published fence and the
client's outbox re-drives under it, as after any fence change. A resume that
fails is not this send's answer; admission reports the lease as it stands.
* chore: restore pnpm-lock.yaml to origin/main (local pnpm rewrote it)
* test(native-chat): pin the pre-handover exit as a failed acquire; stub the status feed in the delivery test
An exit the adapter observes before acquire returns now fails the acquire
with the CLI's diagnostic instead of handing over a dead child; the
published-then-ended path stays pinned by the slow-init startup case. The
delivery test renders the pane, which now activates the host status feed.
* test(native-chat): a same-ID re-hold over the wire joins the one resume, and a replay reopen goes on the idle clock
* test(native-chat): a re-hold that joins a failing resume proves one resume ran
* fix(native-chat): a create whose child was proven gone answers the refusal on the first call
The previous change answered a first-hand root exit as the exited refusal on the
first call, but the common failed start never took that path: when the close
ladder proves the whole tree dead the acquisition error is a plain one, the
store-failure classifier rethrows it, and the client still saw a runtime error
first and the refusal only on replay.
The cleanup that proves the child gone now names such a failure
`AgentSessionAcquisitionExitProvenError`, carrying the provider's diagnostic,
unless it already names its own verdict (a refusal, a typed exit proof, a host
store code). The attach answers both proven-exit kinds as the refusal its replay
gives. How a failed acquisition settles and how it is first answered now live
beside the verdict stamp, in the failed-create module.
* test(native-chat): pin the outbox re-driving a stale-refused send under the resumed fence
A send into a session whose child ended is answered stale once the host has
restarted the child. The outbox keeps that operation queued and blocked, and the
fence change the resume publishes re-drives the same operation under the new
fence; the host admits it.
* fix(native-chat): a child restarted for a send nobody holds is still released
The restart a send runs for a childless session takes no holder, on the premise
that the sending surface already holds one. A one-shot writer holds nothing, so
the child it restarted had no release clock and lived until the app quit. The
write resume now arms the clock when no holder is present, as the first-hold
resume already does. The send-after-failed-start cases also pin that the stale
answer's operation is admitted when re-sent under the new fence, and that two
racing sends restart the child once.
* test(native-chat): pin the picked Claude model across a resume whose child starts on its own default
The started event re-reads and persists what the child reports. A resumed child
answers initialize with its CLI default before the saved pick is restored over
it; the record must hold the pick while starting and after started.
* Revert "fix(native-chat): a child restarted for a send nobody holds is still released"
This reverts commit
|
||
|
|
60bd1dfdea |
feat(native-chat): one shell-environment setting for every structured chat (#22387)
* feat(native-chat): one shell-environment setting for every structured chat Structured Codex chats started from the login-shell environment, while structured Claude chats started from Orca's own process environment, so a variable exported in .zshrc reached one and not the other. Both now start from the same base, chosen by a new setting: - on (default): the whole login-shell environment, as a terminal gets - off: Orca's environment plus PATH, locale, SSH_AUTH_SOCK, and the variable names the user lists The setting is re-read each time a chat starts or resumes. It is shown only when Chat UI, the Chat UI default view, and structured native chat are all on. Terminal-backed chat is unchanged. * fix(native-chat): normalize the shell-environment settings when a profile loads A hand-edited settings file could store the variable list as something other than an array, and the structured runtime called `.filter` on it per launch, so a malformed value failed every structured chat create and resume, and the settings pane render. Normalize both keys where the profile loads, the same way the other array settings are, through one shared normalizer the runtime policy also uses. Also pin that an uncommitted name draft survives an unrelated settings re-render. * fix(native-chat): keep the pinned account as the only source of a structured chat's Claude home The session record owns which Claude home a structured chat uses, and the acquisition pin (claudeConfigDirEnvPatch) is the only emitter of CLAUDE_CONFIG_DIR, compared against what the child would otherwise inherit. With the login-shell snapshot as the inherited base, a CLAUDE_CONFIG_DIR exported only in a shell rc flipped that comparison and produced an explicit pin to the CLI default home, which moves the CLI off its default Keychain item. Drop the inherited CLAUDE_CONFIG_DIR in the Claude launch resolver before the pin runs, as Codex already does for an inherited CODEX_HOME. A configured per-agent overlay still passes through, since the record already honors it. * fix(native-chat): drop Orca's own CLAUDE_CONFIG_DIR from a structured Claude child too The process spawner merges Orca's process env under the launch env, so a CLAUDE_CONFIG_DIR exported to Orca itself reached the child around the launch resolver's drop and unseen by the account pin. One helper now strips it from both inherited bases. Also declare the two shell-environment settings on the runtime store contract and add the six new strings to every locale catalog. * feat(native-chat): add shell variables one at a time with a removable list * fix(native-chat): return focus to the name input after removing a shell variable * fix(native-chat): use a neutral placeholder for the shell variable input The empty input showed a grey HTTPS_PROXY as its placeholder, which reads as a saved value, especially right after that exact entry is removed from the list. Use "Variable name" instead, in every locale catalog. |
||
|
|
dc8cf30554 |
fix(native-chat): end a structured turn when the agent reports it failed (#22047)
* fix(native-chat): end a structured turn when the provider reports it failed (#22044) A turn reads as working while its durable turn row says `running`, and only two events could write a terminal row: the provider's turn-completed notification and the provider process going away. A provider error that ends a turn is neither, so the row stayed `running` with nothing re-deriving it, and the chat counted "Working for N" for the life of the session. Codex reports such a failure as an `error` notification naming the turn it ended, with `willRetry` distinguishing it from a stream error it is about to retry. That frame now settles the turn it names. Claude's CLI reports the same through its session-state frame, whose `idle` arm the SDK documents as the authoritative turn-over signal; that now settles the open turn too. Codex's `thread/status/changed` deliberately settles no open turn: the app server clears `running` on every error, including ones it reports as not affecting turn status, so a turn still open there is still running. What it does settle is a send whose dispatch was never answered — a timed-out dispatch is recorded as unverified delivery, reads as work still owed, and nothing in a live session retired it. Retiring it never makes the send re-deliverable. Splits the codex notification translator so the file stays inside its line budget. * fix(codex): defer idle dispatch release until turn end * fix(claude): enable session state lifecycle events |
||
|
|
434365d2de |
Offer to reconnect native chats that were working when Orca restarted (#21096)
* feat(native-chat): resume structured chats that were working at restart
Teardown records a marker for every session this host was genuinely running a
turn for, derived from the LIVE runtime rather than a persisted status row, so
a stale `running` row left by an older crash can never trigger a resume. On the
next launch a modal lists exactly which chats would resume and resumes them via
native continuation (Claude resume/resumeSessionAt, Codex thread id) — never by
re-sending the prompt, which is what makes an agent redo finished work.
A session resumes only when all of these hold: a teardown marker exists and has
not expired, the record's lease is released and reconciled, a provider resume
cursor exists and still matches the marker, the journal's own turn record names
the same turn, and the marker has not already been spent. Markers are consumed
before the resume is submitted, so a crash mid-resume cannot double-fire, and an
admission gate refuses a second concurrent resume for one session. Resumes are
staggered three at a time rather than spawning every provider at once.
The modal's "Don't ask again" checkbox writes the nativeChatResumeWorkOnRestart
setting, which Settings can turn back off; automatic mode runs the identical
predicate and staggering and reports what it did. Declining consumes the markers
so the prompt cannot return every launch — nothing is lost, because opening a
chat still re-acquires it at the same cursor.
* fix(native-chat): compare handle ROOT and turn state when offering a resume
Four defects QA found in the restart-resume offer, fixed together because the
first two interact: shipping the root fix without the state fix would convert a
silent no-op into actively offering finished chats.
1. Claude was never offered (0/4). The marker recorded agentSessionProviderHandleKey,
which embeds Claude's leaf uuid — a branch cursor. The adapter's own close path
appends a `resumed` link with an advanced leaf during the SAME teardown, so the
marker went stale seconds after it was written and the drift guard refused every
Claude session forever. Record and compare agentSessionProviderHandleRoot instead:
the root is the part a resume must preserve, and changing it is a fork, which is
exactly what this guard is for. Codex is unaffected (its thread id is the whole
key) but uses the root too, so the rule is uniform.
2. The predicate compared turn IDENTITY but discarded turn STATE, so a `completed`
turn satisfied it as readily as an interrupted one. Eviction rewrites `running`
to `interrupted` and never to `completed`, so the state is what separates work
that was cut off from work that finished. Require `interrupted` or `unverifiable`.
3. A chat blocked on a pending approval or question was marked as working, because
the teardown reader accepted any `running` turn while the product's own projection
calls that state `attention`. Teardown now defers to that projection: an agent
waiting on the USER is not interrupted work.
4. "Resume all" could silently no-op. The modal fetched candidates at mount; by click
time the chat's own pane may have bound and taken the hold, moving the lease to
`live` so the predicate dropped it and the call returned no results, leaving the
dialog open behind a dead button. Re-derive at click time and settle an
already-live session as resumed — it is running, which is what the user asked for.
Test fakes now model the Claude close path that advances the leaf, which is why no
unit test could previously exhibit defect 1. Ablation covers all eleven guards.
* fix(native-chat): gate the already-live settlement on the full resume predicate
Two follow-ups from re-QA, both cases of a rule stated by intent rather than by
discriminator.
1. The already-live path bypassed the predicate. "Resume all" sends no session
ids, so the fallback's target set was every marker, and it was gated only on
the session having a live provider child. A chat the predicate had refused --
a completed turn, say -- whose pane happened to own the lease was therefore
settled as `already_live` and had its marker spent, inflating the "Resumed N"
count with chats that were never eligible. No provider spawned and no tokens
were spent, but a marker the predicate rejected must never be consumed.
The resumable set now takes an explicit `leaseState`. The already-live path
derives a second set with ONLY the released-lease clause relaxed, and settles
a session just when it is in that set. Every other clause still applies.
2. The `attention` rule was one-sided. Teardown refuses to mint a marker for a
chat blocked on the user, but the set predicate had no equivalent, so a marker
arriving by any other route was offered once eviction rewrote its turn to
`interrupted` -- the same asymmetry the completed-turn case had.
Gated on projectStructuredAgentSessionStatus === 'attention'. That projection
tests for a pending approval or question BEFORE it looks at turn state, so it
still reports `attention` after the turn is settled, which makes it the durable
signal and keeps one source of truth with teardown.
Ablation now covers thirteen guards, including one for each of the above.
* fix(native-chat): capture awaits-user on the marker instead of re-deriving it
The awaits-user clause could never fire. It asked the live projection for
`attention`, which needs a prompt whose resolution is still `pending` -- but
teardown CANCELS that prompt a few phases after it writes the marker. By the next
launch the evidence is gone, for precisely the sessions the clause was written
for. QA measured the injection still being offered and then resumed.
This is the same shape as the leaf-drift bug: state read after teardown is not the
state that justified the marker. The discriminator, now applied across the whole
predicate:
- a fact teardown itself destroys or mutates must be CAPTURED on the marker
while it is still true;
- a fact that evolves on its own must be RE-DERIVED at read time, never
snapshotted.
So `awaitsUser` is now recorded at teardown and the predicate reads the recorded
value. Teardown still declines to mint a marker for such a session, so the
recorded flag is the second line rather than the only one.
Audit of every other clause against the same test:
- turn id (captured) -- teardown rewrites turn STATE but never the id. Correct.
- provider handle root (captured) -- the close path appends a resumed link, and
appendAgentSessionProviderHandleLink refuses one that changes the root, so the
root is invariant under exactly the mutation that broke the key. Correct.
- turn state (re-derived) -- DELIBERATE exception, stated here rather than left
implicit: we are not reading the state that justified the marker, we are
reading teardown's receipt that it settled the turn. A turn still `running`
means eviction never finished, and we refuse. Correct, and intentionally so.
- lease reconciled / released / handoff stage (re-derived) -- these answer a
different, launch-time question: may this host take the lease NOW. The
teardown-time value would be meaningless, and `unreconciled` is cleared by
this launch's own reconciliation. Correct.
- adapter support, marker TTL, marker consumption (re-derived) -- all evolve
independently of teardown. Correct.
Only awaitsUser was on the wrong side.
* fix(native-chat): drop the unreachable awaits-user marker flag
The captured flag was dead code. `awaitsUser` could only be true when the
projected status was `attention`, and `attention` hits the `continue` above the
push -- so every marker teardown can ever write carries `false` (QA measured
22 of 22 across two real teardowns). The predicate clause reading it was
unreachable by any production path.
A flag that is structurally always false is worse than no flag: it reads as a
safeguard, so the next person to touch this trusts it. The asymmetry it was
added to close was only ever reachable by fault injection, because teardown is
the sole writer of markers and already refuses attention sessions.
Removing it also drops an upgrade discontinuity: as a required field it made a
marker written by the previous build fail validation and be silently discarded,
costing a resume offer on precisely the upgrade where the user was mid-turn.
Markers predating the providerHandleRoot rename still will not parse, but those
carry a leaf-sensitive key the predicate would refuse anyway, so nothing usable
is lost.
In its place the teardown gate now states that `status !== 'working'` is the
SINGLE gate for awaiting-user sessions, why a predicate-side mirror would be
unreachable, and why it could not even re-derive the fact -- so the reasoning is
inherited rather than rediscovered.
Ablation is back to twelve guards; every other clause is unchanged.
* fix(native-chat): say reconnect, not resume, and show each offer's age
Two changes, both independent of the parked continuation decision.
1. The copy claimed something QA disproved. "Resuming continues each agent where
it left off" is false: reconnection restores the session at the point it
stopped, with full context and without re-sending the prompt, but the
interrupted reply does not continue on its own. The toast's "Resumed N chats"
implied work had restarted.
Audited every user-facing string against the rule that none may claim work
continues or that a reply resumes -- which caught more than the three strings
the fix started from. The title, the row button, "Resume all", "Resuming...",
the not-now hint ("picks it up where it left off"), the checkbox and its hint
("resume on their own"), the list's aria-label and the Settings row all made
the same claim. The user-facing verb is now reconnect throughout; the body and
update variant state outright that the interrupted reply will not continue.
en.json synced, runtime boot catalog regenerated.
If we later decide to send a continuation instruction, this is one commit to
change back. Shipping text we know to be false was the worse option.
2. Rows now show each offer's age. The TTL is 24 hours and a stale offer looked
identical to a fresh one. The marker already carried `recordedAt`, so this is
a render change plus one field on the renderer's candidate type, formatted
with the existing formatUiRelativeTime helper rather than a new one.
The clock is stamped once when the list arrives rather than read during render:
ages then stay stable across re-renders, and the render stays pure, which the
react(purity) rule requires.
Guards, predicate and RPC are untouched; ablation still covers twelve.
* feat(native-chat): show the workspace name on each reconnect row
A row read `codex · folder:8f3a1c22-… · 8 hours ago`. Recognising which chats
would reconnect is the entire point of the list, and at twenty rows a UUID
identifies nothing.
No RPC or host change was needed: the renderer can already resolve this id.
Resolved the way automation dispatch resolves the same id space
(resolveAutomationDispatchWorkspace) -- a folder workspace by its full
`folder:<uuid>` key via getKnownWorktreeById, a git worktree by its bare
`repoId::path` id via allWorktrees. Both return a Worktree, whose displayName is
a required field, and DetectedWorktree extends Worktree so either shape answers.
Falls back to the id when nothing resolves, which is what the row showed before
and also covers the window before the worktree store has hydrated.
The lookup lives in a per-row subcomponent because a hook cannot run inside
`map`, and its selector returns a primitive string so repeated selector runs
cannot churn referential equality.
* feat(native-chat): group the reconnect modal by worktree and add opt-in continuation
Grouping. Rows are now grouped under a worktree heading with the repo glyph and
an agent count, using the sidebar's own collapse mechanics. Only presentational
pieces are reused -- RepoIconGlyph, CompactAgentExpansion, AgentIcon and
formatShortTimeAgo. The sidebar's agent row cannot be: worktree-card-compact-agent-row
imports DashboardAgentRow, the dashboard's own type, so both surfaces render one
live-agent model requiring a pane, tab and status entry. Every chat offered here
is by definition stopped, so supplying that would mean inventing live state.
Two things I had assumed were reusable and were not:
- DashboardHostBadge returns null unless hostKind is ssh or remote. Structured
chat is local-only, so it would always render nothing. The host line is
omitted rather than faked; the badge is the right element to add if and when
structured chat gains remote support.
- No state dot. Every AgentDotState misleads here: idle and unverifiable both
presuppose a live pane, interrupted renders red like an error, done green,
working a spinner. A missing dot beats one saying these agents are running.
One worktree renders flat with no heading -- a name, count and chevron around a
single group says nothing the dialog has not already said.
The age column now uses formatShortTimeAgo for sidebar consistency. It takes
(timestamp, now) and subtracts internally rather than taking a delta, so the call
is (recordedAt, listedAt); passing the old delta would have rendered plausible
nonsense. The clock is still stamped once into state, so ages stay stable and the
render stays pure.
Continuation. A secondary "Reconnect and continue" action sends one message, from
a single shared constant, identical for both providers. Reconnect is unchanged and
still sends nothing. An info popover quotes the literal message read from that
same constant, so what is shown cannot drift from what is sent.
Ablation now covers fourteen guards. Two are new: continuation only follows a
reconnect that actually happened, and -- inversely -- a send injected into the
reconnect path must turn the test red, since "don't ask again" rests on reconnect
never sending.
* feat(native-chat): say terminal sessions kept running, and clear the quality gate
The modal lists stopped chats with no way to tell that CLI agents are fine, and
the true state of the world is counterintuitive: the terminal sessions survived
the restart and the chats did not. One line now says so, next to the heading
where it frames the list rather than as a footnote at the bottom.
Wording follows the app's own vocabulary rather than inventing a term: the
catalog settles on "terminal sessions" (terminalSessionCount, "Terminal sessions
are grouped by workspace", "No terminal sessions yet"), and UpdateCard already
reassures with "Your terminal sessions won't be interrupted during the update" in
the same text-xs text-muted-foreground treatment. "kept running" rather than
"were restored" -- nothing reconnected them, they never stopped, and the line
says nothing about why.
Also clears check:code-quality:changed, which I had not been running -- oxlint
alone covers neither the design-system nor the casting audit, so 18 findings had
accumulated across the branch.
- design system (4): Button spacing hand-rolled as gap-1/px-2 is just size="xs";
PopoverContent and DialogTitle own their typography and spacing, so the
text-xs moved to the popover's own children and the title's icon gap moved to
a plain wrapper.
- casting (14): production code loses its assertions outright via Reflect.get,
the idiom already used in managed-hook-detection-commands and
worktree-name-retirement. The marker validator reads each field through
Reflect.get and now checks recordedAt is a number rather than asserting it;
the store-file parse uses the existing `file` shape instead of a second
assertion; the runner narrows the admission error's owner with typeof.
Test fixtures keep their assertions behind the line-specific SAFETY:
rationale the repo mandates for exactly this case.
One trap worth recording: the audit reports an assertion at the line its
EXPRESSION OPENS, not where `as` appears, so a disable-next-line above the
closing brace of a multi-line literal is inert and silently changes nothing.
Guards unchanged; ablation re-proved 14/14 at this head.
* fix(native-chat): give the reconnect row's provider icon an accessible name
Every row rendered the provider as a bare AgentIcon, whose svg carries no
aria-label, title or alt. With a Claude chat and a Codex chat in one worktree the
two rows were identical to any non-visual consumer, and the dialog offered
several identically-named "Reconnect" buttons with nothing to tell them apart.
A regression from
|
||
|
|
7f5141ae2d |
Make the Agent Permissions toggle apply to Codex chat (#20977)
* fix(structured-chat): deliver the permission posture through each transport's own contract Codex posture moves off app-server argv onto typed `thread/start` and `thread/resume` params. Manual states `on-request` / `workspace-write` explicitly instead of omitting the fields, which app-server resolved through the mirrored config.toml — a Manual thread on a home carrying `approval_policy = "never"` never prompted. Claude keeps its owned `--dangerously-skip-permissions` flag through SDK `extraArgs`; the SDK's typed bypass option emits a newer allow flag that older user-installed binaries reject. Posture is re-derived from current settings on every session acquisition. * fix(structured-chat): parse permission arguments as argv * fix(structured-chat): keep permission policy authoritative |
||
|
|
c702e77bc7 |
Stop reading the terminal arguments field on the structured chat route (#20944)
* fix(native-chat): stop reading the terminal arguments field on the structured chat route Setting Claude's Arguments to "--dangerously-skip-permissions --model Opus" made every new Claude tab open in the old terminal-backed chat instead of the new structured one, with nothing on screen to explain why. Removing "--model Opus" fixed it. The cause was a whole-string comparison: the configured arguments were checked against a single blessed value per agent, so any added token at all — including one the agent supports — stopped the string matching and the launch was demoted. Structured chat does not run the interactive CLI. It drives Claude through the Agent SDK and Codex through app-server, and those take narrower option sets that are versioned separately from the CLI's, so one free-text field cannot have a guaranteed meaning for all three. The structured route now reads only what it can actually honour: a replaced launch command, or a launch that names its own working directory. Terminal launches still apply the field exactly as before. Permission posture no longer travels as a raw flag. It is derived from the resolved launch arguments, which is the same fact a terminal launch acts on and which falls back to the default Orca ships when the field was never touched, so bypass stays on by default and Manual is still honoured. Claude gets the SDK's typed permissionMode and allowDangerouslySkipPermissions at query start; Codex gets its bypass flag placed before the app-server subcommand. Both are re-derived per acquisition beside the auth policy and environment overlay rather than stored in the session record, so nothing can disagree with the setting. Codex also loses the --profile, --add-dir and -c passthrough that reached app-server through that field. Only the permission posture comes back. * test(native-chat): pin routing authority on the narrowed feasibility input The routing-authority pin still named the old bundled blocker and built its "customized" fixture out of the arguments field, which is no longer a feasibility input. Both are now the launch command, and arguments and environment are customized on both passes of the loop, so the flag handed to the shared resolver tracks the command alone — a caller that resumed reading either one fails here. No case is dropped and no assertion is relaxed: the blocker list is still exhaustive and every caller must still honour a refusal from the shared resolver. |
||
|
|
f55b7ba680 |
fix(native-chat): cancel pending prompts precisely (#20601)
* fix(native-chat): hide activity while awaiting input * fix(native-chat): keep approval turns cancellable * test(native-chat): satisfy split PR quality gate * fix(native-chat): catalog approval cancellation label * fix(native-chat): include approval cancellation runtime label * fix(codex): settle prompts when cancelled turns complete * fix(codex): settle prompt registry fallbacks * test(native-chat): cover pending interaction fallbacks * test(native-chat): split prompt state coverage * test(native-chat): keep prompt state isolated * fix(native-chat): bound prompt turn backfill * refactor(codex): centralize prompt registry bounds * fix(native-chat): cancel pending prompts precisely * fix(native-chat): consolidate capability imports * fix(native-chat): harden precise prompt cancellation * fix claude cancellation teardown races * retry claude prompt lifecycle admission * bound claude prompt cancellation retry work * fix(codex): bound prompt turn identity on registration * fix(native-chat): route rejected late dispatch settlements * fix(codex): retain exact cancellable prompt turn ids --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
955051ded0 |
fix(codex): settle a structured send on admission, and stop minting a colliding identity (#20138)
* fix(codex): settle a structured send on admission, and stop minting a colliding identity Two sends could be written into the journal under one durable identity. Codex coalesces a mid-turn `turn/start` into the running turn rather than refusing it -- measured against real `codex app-server` builds 0.147.0, 0.150.1 and 0.153.4, none of which refuse and none of which fire a second `turn/started`. The dispatch path read the turn id from the turn/start response and stamped every accepted send `ordinal: 0`. Since a coalesced send gets the running turn's id back, two submissions persisted the same `providerItemId`. That string is durable, and it is the key a restore uses to match a submission against provider history, so the second message's real history row matched nothing and rendered as an extra bubble on replay. On 0.147.0 it is worse than a collision: the coalesced response returns a turn id that never starts and never completes, so the persisted key named a turn absent from history and NEITHER message could match. Identity is now minted from the echoed user message at `identityFor` -- the single point that mints the journal row's own identity -- so the settled key is by construction the one replay computes, rather than a parallel calculation that can drift. Dispatch returns `admitted` when the transport write completes; identity settles on the echo through a channel that did not previously exist for Codex. Waiters are keyed by client message id instead of being shifted off the front of an array by arrival order, and they are cleared on session close and child exit -- previously a timeout was the only thing that ever ended one. `TURN_ID_WAIT_MS` is deleted. It was never reachable on any build measured: `readCodexTurnId` returns non-null on all three, so the 10s wait never fired. The comment justifying it claimed older builds acknowledge before the id exists, which no tested build does. Three comments asserting Codex answers a mid-turn send with `turn already running` are corrected. Their only backing was a test fixture inventing that error string. The correction is factual only -- every changed line in `src/main/runtime/orchestration/` is a comment, and mid-turn delivery is still refused for both providers. Whether that policy is right is a separate question; it was resting on a false premise. Known gap, stated rather than implied: this prevents new collisions and does not repair journals already written with a colliding or phantom key. Those conversations keep duplicating on restore. Repairing them means re-matching persisted submissions against provider history and rewriting `providerItemId` -- which is what `journal-submission-reconciler.ts` is written for, and it still has no production caller. * test(codex): drop the synchronous-accept contract and the colliding `:0` from the integration fakes Three tests in the structured-session integration suites encoded the dispatch contract this branch replaces, and two of them pinned the defect it fixes. They asserted `agentSession.send` answers `dispatchState: 'accepted'` carrying `providerItemId: codex:<thread>:<turn>:0` at send time. That ordinal was never observed; it was stamped on every accepted send, which is exactly the collision this branch removes -- a send coalesced into a running turn is answered with the running turn's id, so two submissions persisted one durable key. The visible failure was a 30s timeout rather than a failed assertion. The fake client advertised no `agent-session.pending-send-result.v1`, and without it the host holds the reply until the send settles: a shim for clients too old to render a pending bubble. The fake provider then echoed the user message with no `clientId`, so nothing could correlate that echo back to the submission, and the wait ran to its own 30s ceiling. Real Codex sends `clientId` on that echo, and the fake now does too, which is what makes it a model of the provider rather than a sketch of one. The identity assertion is kept rather than dropped. Each send now asserts `pending` with no identity at admission, then asserts the submission settles `accepted` at `codex:<thread>:<turn>:0` once the echo lands. Same ordinal, but earned from `identityFor` on the echo -- the key a replay recomputes -- instead of guessed from the turn/start response. Ablated: removing `clientId` from the two echoes leaves both submissions `pending` and fails both assertions, so the assertion is load-bearing and not satisfied by something incidental. Both suites' client fixtures now advertise the capability set the desktop renderer sends in `src/main/ipc/runtime.ts`, which is what these suites mean by a client. The older-client settlement wait keeps its own coverage in `src/main/runtime/rpc/methods/structured-agent-session.test.ts`. `structured-agent-session-runtime-exit.test.ts` asserts `pending` for the same reason; it drives the host directly, so it never took the compatibility path, and what proves delivery there is still the turn the reacquired provider starts. The replay suite's "without dispatching it twice" property is untouched: one `turn/start` call, one replayed ledger row. * fix(codex): preserve unsettled dispatch correlations * test(codex): type the dispatch fixtures instead of asserting over them main's new casting gate (#20367 base) flags type assertions on changed lines. Replace them with checked types: the recording sink already satisfies its interface, both CodexSession fixtures are now annotated and carry real collaborators, the settlement assertion compares whole identities, and the integration helper reads submissions through the host's public journalSnapshot instead of its private session map. * fix(test): merge the duplicate doubt-reasons import the merge left behind Both sides added an import from journal-dispatch-doubt-reasons and the merge kept both statements, which the whole-repo native plugin gate refuses under --deny-warnings. * test(codex): a Fast mode turn is admitted, not accepted #20506 landed its Fast mode tests against the dispatch contract this branch replaces: a Codex send now returns admitted and settles its identity on the provider echo. The tier assertions the test exists for are untouched. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
539d4d1f32 |
fix(native-chat): resume structured chats cleanly after restart (#20509)
* fix(native-chat): retire provider ownership on restart * fix(native-chat): stop showing a restart eviction as a provider death Restarting Orca turned a resumable structured chat into a user-visible `Provider exited: recorded pid absent on host`. Quit never released the durable lease, so restart probed the recorded pid, adjudicated the session evicted, and wrote a synthetic status row against a chat that was perfectly resumable. The fix is the missing teardown phase plus the missing fence check: quit now evicts every provider child this host owns — stopping it, settling its journal and handing the lease back — and the release compare-and-swaps on the fence it expected. Restart then finds a released lease and reopens the chat silently. What the user sees is decided by the typed death evidence rather than the shape of a settlement id: only an `exit-observed` death writes copy, and that copy now carries its cause so an auth failure and an OOM kill do not read alike. The reassuring wording stays. Historical synthetic rows are filtered out of the render projection, which needs no schema change and leaves every real provider-exit row alone. Also: - Bound the new eviction phase well below the quit deadline; a quit that dies mid-eviction leaves the lease unreleased, which is the original bug. - Scope the interruption verdict to work that was mid-response. A provider that died while waiting on an approval interrupted nothing. - Keep host bookkeeping in step with the adapter: the provider-child flag clears when the child is proven stopped, not seven steps later. - Drop the router's duplicate shutdown gate and acquisition drain — both adapters already own theirs — and latch the router closed so a late acquire cannot fan a session back out to closed adapters. - Attach the real cause to the settlement failure a quit reports, and remove a recovery-ticket field that was hardcoded at its only construction site. * fix(native-chat): scope the legacy status filter to the copy it retires The read-time filter hid every status row carrying a `restart-eviction:` identity. That identity is still minted, so a genuine provider death settled under it would have been dropped from every rendered page. Match the retired `Provider exited` copy as well, so only the legacy rows are hidden. Three smaller corrections alongside it: - The settlement retry path now applies the same unfinished-work check the live exit path uses, so a provider that died waiting on an approval no longer gets told a response was in progress. - Bound the exit reason before composing the outcome copy, so a stderr dump in the reason cannot push the "you can continue" sentence past the row's byte cap. - Correct the teardown comment: tail rows are protected by eviction's own per-session ordering, and `closeAll` is a backstop for children eviction never took, including one whose eviction was refused. * fix(native-chat): retire legacy status rows at the projection source The read-time filter that hides the retired `Provider exited …` rows ran on the way OUT of the page builder, after the paging math had already measured the unfiltered timeline. A backward window landing entirely on those rows returned an empty page that still reported `hasOlder: true` with a null `window.oldest`, so the renderer's backfill loop re-asked from the same anchor forever. Its only no-progress guard compares `window.oldest?.sequence` to the anchor, and `undefined === n` never breaks. The live subscription opens behind that loop, so the transcript never finished loading either. Filter where items ENTER the page pipeline instead: the reduced snapshot gets one renderable timeline, the forward path gets one renderable batch, and the window bound, effective limit, `hasOlder`, `window.oldest` and `nextCursor` are all computed over that single array. A window with nothing left behind it now reports end-of-history. Also restore the eviction retry contract. Clearing `hasProviderChild` as soon as the adapter proves the child gone is honest, but it is a different fact from the wind-down this host still owes. A retry after a step aborted between the two was reading "no child here" and skipping both the dead-generation settlement and the lease release the aborted attempt had promised to repeat. The obligation is now tracked separately and cleared only by a release that actually landed. And rename the filter to the copy it retires: it drops only rows carrying the retired `Provider exited` text, not restart-eviction status rows in general. * fix(native-chat): read the wind-down a close owes from the live child An eviction recorded "nothing owed" whenever it ran over a session with no provider child of its own, and the retry then read that record in preference to the child in front of it. A session suspended to an agent terminal is exactly that shape, and the trip back to native re-acquires into the SAME session object rather than replacing it, so the next close skipped both the dead-generation settlement and the lease release — leaving the record claiming a live owner this host had just stopped, and a pending send unsettled. The obligation is now derived the way the quit sweep already derived it, from one shared predicate: a live child always owes a wind-down, and a remembered `false` only carries the obligation forward, never cancels it. Also drops a memoization in the history page that could never hit. Its key was the snapshot's items array, which the reducer rebuilds on every `snapshot()` call, so each backward page allocated a fresh key; the one reader that does share a snapshot across pages reads forward and never calls it. The comment claimed a multi-page read filtered once, which was not true of either path. Tests: the handoff round trip that strands the lease, and the quit sweep picking up an eviction whose close retry never came. * chore(native-chat): scope three helpers to their file and pin the teardown order retryUnexpectedExitSettlement, hasUnfinishedStructuredAgentSessionWork and isRetiredProviderExitStatusItem each have no consumer outside the file that defines them, so they no longer advertise an external contract. The quit-path phase list documents its order as load-bearing, but nothing asserted it. Pin the phase names so evict-owned-sessions cannot drift out of its slot between drain-attaches and flush-event-sinks. * fix(native-chat): stop the router reporting a stop it never observed `closeAll` cleared the route table and set one boolean, after which that boolean was the only surviving evidence about any session. Two call sites then spent it: `releaseAcquisition` and the stop path each turned a route-lookup MISS into reported success. Eviction reads a `true` from the stop path as proof the provider child is gone and releases the durable lease on it, so a session the router never routed could have its lease handed back on the strength of "I have no record, but everything is closed." Loss of contact is not evidence of process death. The fix keeps the evidence instead of the inference: adapter shutdown only resolves once every child is proven stopped, so `closeAll` now marks each routed session `stopped` rather than forgetting it. A routed session still answers `true` from its own retained proof; a session with no route answers `false`, which leaves it indexed for a real retry. `releaseAcquisition` drops its short-circuit and asks the adapters, which answer from their own session maps. The acquire-side latch is unchanged: once closed, the router stays closed and refuses new work. Behaviour that changed: a post-`closeAll` stop for a session the router never routed, or one the host already acknowledged as released, now reports unproven instead of proven. That matches what the same call already answered before `closeAll`, and no real flow reaches it — quit evicts every owned session before `closeAll` runs, and eviction only asks the adapter for sessions whose provider child this host acquired through the router. * test(native-chat): ratchet the retired provider-exit copy out of production The retirement filter hides a status row on two facts: a restart-eviction item id and copy that opens with the retired prefix. The identity half is still minted today, so the filter cannot tell a new producer's row from the legacy row it exists to hide — any future writer of that copy would be dropped from every transcript with no trace. Until now that safety property lived only in a doc comment. Scan the shipped tree for a string literal that OPENS with the retired prefix, which is exactly what the filter's `startsWith` reads. Comments are stripped first, so prose about the retirement is not a producer, and the filter's own constant is exempt. Tests are excluded: writing the copy is how the filter is exercised. * revert(native-chat): drop the read-time retired provider-exit filter Fix forward instead. The lifecycle change in this branch stops any new `Provider exited: <reason>` row from being written; rows a previous build already persisted stay in those transcripts and age out with them. A permanent read-time filter for a cosmetic, shrinking set was not worth its maintenance cost, and its paging seam was the only place a backward window could land entirely on hidden rows. Removes the filter module and its test, restores agent-session-history-page.ts to its pre-branch form, and drops the tests that only existed to prove the filter did not over-match or wedge the backfill loop. The copy ratchet stays and now carries the whole guarantee: with no filter in front of it, any production writer that resurrects the retired prefix reaches the user's transcript directly. |
||
|
|
ebb1acfa37 |
refactor(agent-status): publish structured sessions into the hook server store (#19683)
* refactor(agent-status): publish structured sessions into the hook server store Structured (native chat) sessions have no PTY and no hook script, so their status never reached the hook server's store; #19217 gave `worktree ps` its own adapter over the structured feed instead. The feed now writes every projection into that store through a status sink the runtime wires, drops the row when the host closes the session, and `worktree ps` reads the one snapshot like every other agent. Rows carry a `structuredHost` marker and the journal clock; they are never persisted to last-status.json, and the main process does not forward them to the renderer yet, whose feed bridge still owns them until it is retired. Design and the two follow-ups: docs/reference/agent-status-store.md. * chore: drop stray @pnpm/exe lockfile entry An unrelated local pnpm run added @pnpm/exe as a packageManagerDependency with no package.json change, so CI's --frozen-lockfile install failed before any job ran. * docs(agent-status): describe the step that actually landed The design record claimed PR 1 deletes RuntimeAgentRowStore, drops the retained-versus-hook reconciliation, stamps terminalHandle on OSC rows, and tags rows with a source field of 'structured-host'. None of that is true of the shipped code: the retained store and its reconciliation are still in place, and the row field is structuredHost: 'held' | 'owned'. AGENTS.md points every future contributor here before they touch agent status, so split the roadmap into the 1a that landed and the 1b that has not, and name the fields the code actually writes. * fix(agent-status): pair session removal with the status-row forget A session dropped from the host's map without an explicit forget left its row in the store forever: `structuredHostOwned` bypasses the staleness check, so a failed re-attach (the Claude rewind path reaches one) stranded a permanently working agent in `worktree ps` and on mobile with no UI able to clear it. Deletion and forget are now one operation both callers route through. * fix(agent-status): give orcad the store worktree ps reads from `orcad` constructed its runtime with neither `getAgentStatusSnapshot` nor `structuredAgentStatusSink`, so once `worktree ps` sourced rows only from that snapshot the headless host published nowhere and listed nothing. The hook server's store is a module singleton whose import tree never reaches Electron, and its file paths come from `start()`, which orcad never calls. * fix(agent-status): drop a structured row without a renderer clear `dropStructuredStatus` went through `clearPaneState`, which fans a pane clear out to the renderer for a pane key the renderer's own feed bridge still writes - so 'exactly one writer per pane key' held for writes and not for deletes. `dropStatusEntry` routes through the status-drop tap instead, and skips the resume-identity remnant: a structured session has no pane to resume into, and every null-status publish would otherwise re-mint one. * test(agent-status): pin both half-migration structured-row filters Neither the `agentStatus:getSnapshot` filter nor the main-window listener's had a single assertion, so deleting either — the first step of PR 2 — was green everywhere. Also covers the perf skip and the drop's lack of a renderer clear. * docs(agent-status): correct three statements this PR made false The sink JSDoc claimed only tests construct a host without one; `orcad` did. The doc argued a structured row needs no tab mirror 'because headless serve has no renderer', reasoning about exactly the topology the wiring had not reached. The deleted runtime adapter's warning that the pane key must be the DERIVED one - never a bearer handle or minted worker key - was lost with it. * test(agent-status): declare orcad in the hook-row producer census Wiring the hook store into the orcad runtime added a production site that hands hook rows to a consumer, which the census ratchet pins deliberately. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
5868fdc9e3 |
feat(native-chat): report Codex background tasks in the chat strip (#19346)
* feat(native-chat): report Codex background tasks in the chat strip The background-tasks strip works for Claude only; a structured Codex session shows nothing in it. Feed it from the Codex app-server stream. The strip stands for work that OUTLIVED a turn, which is what the monitoring header, Claude's foreground suppression, and the conversation command gate all already assume. Codex has no `is_backgrounded` flag, so that fact is derived from the turn boundary: a `subAgentActivity` child or a primary-thread `commandExecution` becomes visible once the turn it belongs to completes and it is still unsettled. `turn/completed` only reveals a task here, never settles one — measured on `codex app-server` 0.153.4, a spawn_agent child reported `completed` 95.8s after its parent turn ended. Only a child's own activity kind settles it. Codex exposes no honest stop: `turn/interrupt` on a child ends its turn without emitting a terminal activity item and leaves its shell running. So the state carries a new optional `supportsStopAll: false`, the strip hides a control that could not act, and the blocked-command message asks the user to wait rather than to press a button that does not exist. * refactor(codex): move session teardown out of the structured adapter Merging main crossed the 300-line cap on `codex-structured-session-adapter.ts`: the rewind backend (#19235) and this branch's close-time strip clear both landed in it. The four close paths move verbatim into `codex-structured-session-teardown.ts`, where they funnel through one `settled` helper instead of repeating the notification-retry and background-task cleanup at each call site. No ratchet bump. Also normalize a background task's description once at receipt rather than on every projection; the roster is re-projected on each observed frame. * fix(codex): drop the shell row the journal already settles A `commandExecution` still `inProgress` when its turn ends was reported as a `command` task. But `settleCodexJournalTurn` writes exactly those items to the journal as `state: 'failed'` on `turn/completed` and forgets them, so the strip row would have claimed a shell was still running at the same instant Orca recorded that it was not — two surfaces contradicting each other about the same process. A subagent is the opposite case and stays: the roster pointedly does not sweep at a turn boundary, because children measurably outlive it. That leaves the producer making exactly one claim — these spawn_agent children are still live after their turn — which the durable roster row corroborates. * fix(native-chat): track Codex background execution lifetimes * fix(native-chat): keep running tool groups from claiming completion * Fix runtime catalog and capability expectation * fix(codex): keep a child's name on the command row that outlives it A child agent's commands stay hidden behind its agent row while the child works. Once the child's turn settles with a command still running, that command surfaces as its own row labelled from the raw command string, so 'long_probe' became "/bin/zsh -lc 'ping -c 300 127.0.0.1 > /dev/null'" at the moment that row was the only remaining signal for the work. Qualify a child's command row with the child's label. Resolved on read, so a label registered after the command still lands, and bounded by the existing description cap so admission accounting stays valid. Primary- thread commands are left unqualified: they have no child to name. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
b7b6ea3942 |
fix(native-chat): auto-rename the workspace on a structured chat's first turn (#19138)
* fix(native-chat): auto-rename the workspace on a structured chat's first turn Structured native chat (Claude and Codex) never reached the first-work workspace rename. The orchestrator has a single production caller, the agent-hook server listener, and structured sessions never set ORCA_PANE_KEY, so no hook event could ever be attributed to one. The renderer knew this and suppressed pendingFirstAgentMessageRename for structured launches at three sites, which also closed the gate the folder-workspace title rename depends on. The host's status feed already computes the exact edge: status 'working' with a latestPrompt normalized the same way the hook payload is, and a workspaceId that IS the worktree id. Publish that projection to the host, thread it out to the runtime, and hand it to the same orchestrator the hook path uses. Re-projections of state the host already knew (restore, an arriving subscriber) are flagged as replays and map to the orchestrator's existing isReplay gate, so a host restart cannot rename off a stale journal. One host and one journal serve both providers, so this covers Claude and Codex together. Verified in a live Electron instance, worktrees created through the real composer and prompts sent through the real chat composer: Codex langouste -> retry-helper-exponential-backoff Claude prowfish -> parse-csv-headers * fix(native-chat): preserve first-work rename across runtime and queued turns * fix(native-chat): skip branch rename for folder projects --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
ad4dc353f3 |
fix(native-chat): settle a structured send the provider proves it received after the ack window (#19140)
* fix(native-chat): settle a structured send the provider proves it received after the ack window A send waits a bounded window for the provider to echo the message it was given. On timeout the dispatch resolves `unknown`. The echo that arrives later IS matched — `recoverLateIdentity` uses it to repair the session's turn identity — but nothing tells the journal, and `unknown` is terminal there. The submission stays unknown for the life of the session. Two consequences, both reachable on any ordinary session: - The composer renders "Message delivery is unconfirmed." with a Retry, forever, for a message that was delivered and answered. - Retry redispatches, because the host only replays a recorded outcome unless `retryUnknown` is set, which that button is the only thing that sets. So the banner is a duplicate delivery armed and waiting for a click — and a user who believes the banner and resends is doing exactly that by hand. Every send made while a turn is already running takes this path: the provider does not echo a queued message until the running turn ends, which is far past the 10s ack window. Sends made while idle are unaffected, which is why this reads as intermittent. Carry the `clientMessageId` on the dispatch waiter and settle the journal submission `accepted` when the late echo proves delivery. Deliberately unfenced against the dispatch sequence: that fence decides which turn owns the identity, while delivery is settled either way. Already-terminal rows are untouched. * fix(native-chat): persist late dispatch receipts before session close --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
af82126058 |
fix(native-chat): give the Claude exit barrier a handle on unpublished exits (#18826)
A first-hand Claude exit is not published where it is observed. `handleExit` re-enters the close ladder and persists the transcript cursor before it emits `ended`, and only that emission reaches the runtime's recovery chain. So the runtime's `waitForRecovery` — whose whole job is to drain an in-flight recovery before teardown stops children — returns immediately for an exit that is still climbing the ladder, and nothing outside the adapter can tell an observed exit from a published one. The integration test for fenced host reconciliation had no handle on that barrier, so it bounded-polled the lease for 100ms instead. Measured under 16x local concurrency, publication alone takes 77-204ms: 19/24 runs failed. Retain the ladder-then-settle tail on the exit record and expose `drainObservedExits`, fold it into `waitForRecovery`, and export the barrier so a caller that needs the settled lease can await it. Codex publishes inside its own exit callback and needs nothing. The test now awaits the barrier: 0/24 under the same load, and it fails on an idle machine without the drain. |
||
|
|
e89deb63c9 |
Show Claude background task status in Native Chat (#18757)
* feat(native-chat): show Claude background task status * fix(native-chat): carry background task fence forward * fix(claude): bound background task stop requests * Show running Claude background task details * Harden Claude background task status updates --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
a65332a8bd |
feat(claude): move structured native chat onto the Claude Agent SDK and enable it on macOS and Linux (#18560)
* Join structured attach teardown through journal bind * fix: restore structured chat parity * feat: add Claude structured session adapter * fix: harden Claude structured adapter * fix: close Claude adapter edge cases * fix: start Claude init deadline after launch * feat: wire Claude structured sessions * fix: harden Claude structured runtime * fix: fence Claude structured compatibility * fix: preserve Claude free-text prompt answers * fix: decode addressed Claude prompt text * feat: enable Claude structured chat on mobile * fix(mobile): keep structured chat provider-aware * fix(mobile): negotiate Claude structured tabs * fix: keep scoped RPC tests native-free * fix: secure mobile structured image delivery * fix: close structured session data-loss gaps * fix: prove real Claude structured startup * fix: consume pre-spawn proof before retry * feat(native-chat): add desktop structured sessions * fix(native-chat): satisfy structured session cleanup gates * fix(native-chat): keep structured renders pure * fix(native-chat): open composer pickers upward * fix(native-chat): use existing view for structured sessions * fix: harden structured desktop status projection * fix: close structured desktop lifecycle gaps * fix: fence structured AI Vault resumes * fix: fence structured AI Vault resumes * fix: preserve structured tabs during activation * feat: toggle structured sessions between chat and TUI * fix: harden structured session handoffs * fix: bind structured TUI before rollout proof * fix: complete structured chat round trips * fix: align structured TUI return readiness * fix(native-chat): make reverse handoff transactional * Add Claude structured TUI handoff seams * fix(native-chat): clear sticky handoff recovery * fix(native-chat): complete mobile reverse after TUI exit * fix(native-chat): keep TUI transcripts readable * fix(native-chat): recover TUI transcript gaps * fix(native-chat): recover claimed TUI owners * fix(native-chat): retain cold TUI proof authority * fix(native-chat): preserve Claude handoff authority * fix(native-chat): recover TUI transcripts read-only * fix(native-chat): harden Claude handoff recovery * fix(native-chat): serialize structured handoff recovery * fix(native-chat): close handoff admission races * fix(native-chat): validate pinned launch environment * fix(native-chat): revalidate restored and retried owners * fix(native-chat): gate restart recovery publications * fix(i18n): catalog Claude session controls * fix(native-chat): wait for structured TUI process proof * fix(native-chat): queue stale idle TUI handoffs * fix(native-chat): route structured Codex options directly * fix(native-chat): persist structured session options * fix(native-chat): hydrate resumed structured options * fix(native-chat): preserve options across structured handoffs * fix(native-chat): replay pending option mutations * fix(native-chat): rotate settled handoff operations * fix(native-chat): rotate refused send operations * test(native-chat): derive refusal retry state from host * test(native-chat): give the host-oracle matrix test an explicit timeout * fix(native-chat): keep Claude option controls idle * fix mobile structured first-send hydration race * fix(native-chat): preserve handoff launch authority * fix(native-chat): harden shared handoff recovery * fix(native-chat): serialize structured handoff recovery * fix(native-chat): close handoff admission races * fix(native-chat): validate pinned launch environment * fix(native-chat): revalidate restored and retried owners * fix(native-chat): gate restart recovery publications * fix(i18n): catalog structured session recovery control * fix(native-chat): wait for structured TUI process proof * fix(native-chat): queue stale idle TUI handoffs * fix(native-chat): keep structured recovery provider-neutral * fix(native-chat): drop local terminal topology from structured sync * fix structured outbox and tab restore races * fix(native-chat): preserve Claude question groups * fix structured provider visibility and request handling * fix structured session TUI handoff recovery * fix reverse structured session handoff * fix(native-chat): recover Claude outbox and resume state * chore(mobile): preserve the working-tree lockfile state before the main merge Carries the pre-existing uncommitted mobile/pnpm-lock.yaml modification into history so the main merge cannot overwrite it. Verified benign pnpm drift (babel 7.29.7->7.29.8 transitives plus deprecation metadata); drops no patchedDependencies (the mobile lockfile declares none). * test(native-chat): drop orphaned Claude handoff-auth test left by the main merge 'pins Claude handoff auth through the terminal provider boundary' is absent from main and its production counterpart preserveClaudeAuthEnv no longer exists outside this test - orphaned residue of the terminal/native handoff work this PR excludes by scope. Removed rather than repaired: the failure was a renamed field (providerHome -> providerRoot), and renaming it would have carried out-of-scope handoff code into the merge. Body preserved as evidence and logged in CLAUDE-STRUCTURED-DISPOSITION-TABLE.md. * Fix mobile structured turn state * fix Claude structured session blockers * fix claude structured lane blockers * fix Claude acquisition exit proof * fix(claude): route stream-json launch through process wrapper * fix(claude): gate structured chat support * Fix Claude structured launch gating * fix(claude): split session acquisition and prune mobile scope * test(claude): align structured session fixtures * fix(agent-session): preserve handoff launch arguments * fix(claude): open journals through the factory after origin/main split The journal opener moved to journal-store-factory on main; retarget the Claude structured tests that still imported the old path. * fix(claude): resolve Claude structured launch args, auth, and win32 proof The origin/main merge re-expressed the lane's Claude wiring onto main's split orca-runtime facade and dropped three wires past green typecheck and lint. - resolveLaunchArgs discarded its provider parameter, so structured Claude sessions were launched with Codex app-server flags; Claude exits on --dangerously-bypass-approvals-and-sandbox, and a Codex arg-parse throw could block Claude session creation outright. - resolveClaudeLaunchEnv was no longer supplied, so the launch resolver fell back to the whole process env as configuredEnv and buildClaudeChildProcessEnv re-applied every auth var it had just stripped. The resolver now merges the Claude overlay onto a strip-applied copy of the inherited env, which also keeps PATH intact for withCliRuntimeOnPath. - The windowsProcessStartTimeAvailable producer was gone while the contract field and both consumers survived, so the renderer gate fail-closed and structured native chat was unreachable on every win32 host. Separately, structured Claude pinned CLAUDE_CONFIG_DIR unconditionally. An explicit pin makes the CLI abandon the macOS Keychain even when it names the CLI's own default, so a default claude.ai account could not authenticate where the legacy Claude terminal could. Pin only a home the CLI would not resolve on its own, matching ClaudeRuntimePathResolver, and compare against the env the child would otherwise inherit so a diverging overlay cannot outrank the record's account home. Also await the now-async revealNativeSession in its regression test, and set the native status before revealing so a rejecting reveal cannot leave a session released but never marked native. Claude-Session: https://claude.ai/code/session_013UqKCRB6k5e8UaYhXUHeWY * fix(claude): scrub case-insensitive Windows auth env * fix(native-chat): settle handoff outcome-write failures instead of leaking them A store write failure while recording a handoff outcome escaped the flow runner's catch handler, so the client never received the failure and the flow surfaced as an unhandled rejection (seen as an intermittent agent_session_store_corrupt error in the proven-dead-retry suite, whose teardown raced the flow's trailing outcome write). Record the failed outcome best-effort, and drain the coordinator before that test's teardown removes the store root. Claude-Session: https://claude.ai/code/session_011aXkcHyeiRJuezupQdjZaM * fix(native-chat): make the structured close-failure toast provider-neutral The structuredSessionCloseFailed toast fires for any structured session, but its copy said 'Codex chat', so a Claude structured session that fails to close showed the wrong provider name. The launch-failure toast is only reachable behind the agent === 'codex' gate, so its copy stays as is. Claude-Session: https://claude.ai/code/session_013ugSpCx4AWkySaJb69BQax * fix(native-chat): wire structured handoff proof recovery * fix(native-chat): wire structured handoff proof recovery * fix(native-chat): correct the structured chat opt-in copy The one `experimentalStructuredNativeChat` toggle gates both providers — `useStructuredAgentSessionCreate` runs `canUseStructuredNativeChat` for `'claude'` as well as `'codex'` — but its description named only Codex. Its scope line also said Windows keeps using terminal chat, while the gate refuses win32 only until the host proves it can read a process start time. `structured-native-chat-availability.test.ts` already pins that Windows is allowed once the proof is cached, so the two contradicted each other. Claude-Session: https://claude.ai/code/session_01RJFsidQWmKYFmeoUuVu4Tp * test(claude): pin @anthropic-ai/claude-agent-sdk 0.3.251 contracts against a scripted CLI PR 1 of the SDK migration: dependency + test-only harness, no product wiring. - Pin @anthropic-ai/claude-agent-sdk to exactly 0.3.251 — not the newest release — because 0.3.251 (published 2026-08-28) clears the repo's 3-day minimumReleaseAge supply-chain gate with no exclusion, while the newest release was minutes old and would have required excluding a brand-new publish from the exact control built to catch brand-new malicious publishes. Every contract this design depends on was verified identical on 0.3.251: the full option surface, no pid on SpawnedProcess (custom spawner stays mandatory), env defaulting to process.env when omitted, and --replay-user-messages appearing only via extraArgs. - Exclude all eight bundled CLI platform binaries via ignoredOptionalDependencies. The setting lives in pnpm-workspace.yaml because pnpm 12 no longer reads the package.json "pnpm" field (it warns and ignores it; verified by install ablation). Excluding the binaries is what makes Orca's pathToClaudeCodeExecutable override mandatory rather than merely preferred. Note: pnpm 12.0.0 honors the ignore list when reconciling an existing lockfile but not on fresh resolution of a new dependency, so the lockfile's SDK entry was pinned surgically; both 'pnpm install' and 'pnpm install --frozen-lockfile' verify clean and stable against the committed lockfile. - Contract-pin suite drives the real SDK against a scripted fake CLI and pins: unknown type/field/content-block pass-through (and keep_alive interception), spawner env fidelity plus the omitted-env process.env inheritance sharp edge, extraArgs producing --replay-user-messages, argument parity for every CLAUDE_STRUCTURED_BASE_ARGS entry plus --session-id/--resume/ --resume-session-at, canUseTool wire request_id stability and abort on control_cancel_request, one spawn per query, pathToClaudeCodeExecutable honored by the default spawner, the exact SDK version, and the eight platform binaries staying uninstalled. Claude-Session: https://claude.ai/code/session_01FGCRfYUnb4hbvfTAHGtJKQ * feat(claude): drive the structured transport through the agent SDK Replaces the hand-rolled `claude -p --input-format stream-json` transport with @anthropic-ai/claude-agent-sdk 0.3.251, keeping the existing connection interface for this commit so the acquisition path changes minimally. The control-plane rewrite is a separate change. Orca still supplies the process. `spawnClaudeCodeProcess` routes through `spawnProcess`, retains the child and its pid — the triple the durable lease adjudicates on — drains stderr so exit errors keep their tail, and hands `.cmd` shims to Orca's Windows argument encoder rather than the SDK's plain spawn. `close()` keeps Orca's own bounded tree-kill and exit deadline, so it still resolves true only after an observed exit. Launch resolution emits an SDK options object instead of argv; durable `launchArgs` translate to a typed option where one exists and to `extraArgs` otherwise, refusing a token neither can carry rather than dropping it. The child env is always passed explicitly — omitting it would let the SDK inherit `process.env` and reintroduce the ambient `ANTHROPIC_*` leak. The stdout line parser is deleted; the SDK owns framing, and unknown frames still reach the translator verbatim. Claude-Session: https://claude.ai/code/session_01JMhFjh9HEnkcJ5YTfCdgD3 * fix(claude): settle the frame the SDK pulled but never wrote The SDK's input pump is `for await (frame of prompt) { await transport.write(frame) }`. When that write rejects — the child dies between Orca's liveness guard and the write — the for-await ends abruptly and calls the generator's `return()`, so the code after `yield` never runs. The frame was already shift()ed out of `queued`, so the later `fail()` from the exit path could not reach it and `send()` never settled: `dispatchClaudeTurn` awaits that send before it can return `unknown`, wedging the caller and the durable outbox. The pre-SDK transport rejected on the stdin write callback instead. Retain the in-flight entry and settle it from the generator's cleanup, and let fail() reach it too for the pump that never resumes at all. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * fix(claude): keep the agent SDK behind the structured-Claude boundary The ordinary OrcaRuntimeService graph statically reaches the Claude adapter and so the transport module, whose first line imported @anthropic-ai/claude-agent-sdk. The SDK is evaluated whenever the regular runtime loads, before any structured Claude session is chosen: it sets process.env.NoDefaultCurrentDirectoryInExePath, changing Windows executable resolution for later subprocesses, and a missing or incompatible install would break normal runtime startup — for a user who never leaves the terminal/TUI path. Defer the SDK to the connection, memoized so it loads once per process, and add the import-graph ratchet: a walk from the Electron main entry that fails on any static import of the package, plus a clean-fork check that loading the runtime leaves the Windows search variable untouched and a child-process pin that the side effect is still real. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * fix(claude): answer list_models so the picker stops serving the seed sendControlRequest had no list_models case, so every request hit the default reject; readClaudeStructuredSessionOptions swallows that with .catch(() => null) and falls back to the static catalog. Every structured session therefore served a hardcoded model list with no per-model effort levels, no resolvedModel and no default detection, and nothing surfaced the failure. The pre-SDK transport got the live catalog from the CLI. Route it through the SDK's supportedModels(), wrapped in the { models } envelope the existing parser reads. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * fix(claude): reap the child's descendants before killing it The forced step of the exit ladder went through the Codex helper, which spawns `pkill -KILL -P <pid>` and SIGKILLs the parent in the same tick: the parent usually dies first, the descendants reparent to pid 1, and `-P` matches nothing. An MCP or launcher descendant of a stubborn Claude child was left running. The test named for that requirement declined to assert it and killed the survivor by hand instead, so it could not fail for the thing it was named after. Route the Claude reap through Orca's existing sweep, which snapshots descendants while their parent link still exists and signals them before the root goes, and on Windows uses the identity-gated `taskkill /T /F`. The test now asserts the descendant is dead; the manual kill stays only as a failure-safe. close() still returns true only on an observed exit. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * fix(native-chat): merge the duplicated handoff type import CI's static-analysis lint (`oxlint --config config/oxlint-code-quality-native-plugins.json src config tests mobile --deny-warnings`) exits 1 on the two separate `import type` statements from the same module. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * fix(claude): answer a permission callback whose signal already aborted settleFrom registered the abort listener and then delivered the request. A callback that arrives already aborted never fires that event, so the promise stayed pending behind a durable prompt with no cancel path. Check the signal first, emit the cancel, and resolve the SDK's null sentinel without registering. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * test(claude): wait for the child to record the frame, not just for its report The scripted CLI writes its report at startup, so `until(readReport)` returned a report with no user messages whenever the child had not yet read the line. The assertion then failed under parallel load. Poll for the frame instead of for the file. Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum * fix(claude): coalesce partial deltas onto one assistant item and stop painting result frames Under --include-partial-messages every stream_event frame carries its own uuid, and the final assistant frame for a block carries yet another; only message.id ties them. The translator keyed each delta by its frame uuid, so a reply painted as one bubble per delta chunk followed by a complete duplicate under the final frame's uuid. The block's first stream frame now mints the claude:(sessionId, uuid) identity, deltas coalesce onto it through the shared 60ms seam, and the final frame reconciles onto that same item. Known SDK bookkeeping no longer reaches the provider-fallback row: result subtypes are catalogued and settled by the turn lifecycle, an empty thinking block (redacted thinking) is a modeled kind, a string-content user replay is a text block, and an empty user frame paints nothing. An unmodeled result subtype or content kind still lands on the bounded fallback row. Claude-Session: https://claude.ai/code/session_01GaP5HpYQbvy2hYehVhwfEW * fix(claude): prove descendant exit at the close boundary instead of on an unref'd timer close() reported proven=true as soon as the direct child exited while the descendant sweep's SIGKILL sat on an unref'd 2 s timer, so a SIGTERM-resistant MCP server outlived the lease release. The reaper now composes the same shared primitives the Codex structured provider uses: snapshot, verified bounded descendant termination on POSIX, taskkill /T /F on Windows. The proof is false whenever descendants outlive the deadline, a retried close re-verifies the retained snapshot rather than trusting the dead root, and the raw pipe child no longer goes through the PTY job sweep it never owned a job for. Measured on macOS: a killed child of a SIGSTOPped parent stays a matching zombie row in ps, so the root is killed while verification runs rather than stopped first as the Codex non-group path does. Claude-Session: https://claude.ai/code/session_0161QFm3KVRNJKfdzWVGVNWk * feat(claude): replace the hand-rolled control plane with the SDK's native surface PR 3 of the Claude structured SDK migration removes the wire-frame scaffolding PR 2 kept, so Orca drives the SDK's typed control surface directly. Inbound permissions move from a rebuilt control_request dispatch to the SDK's canUseTool / onUserDialog callbacks. The prompt registry now carries the callback's own resolver: a decodable can_use_tool becomes a durable prompt whose answer settles the callback; a malformed one is denied without registering; the SDK's abort signal (fired on control_cancel_request, which the SDK matches and dedups itself) forgets the prompt and settles it null, and a late answer after abort finds no prompt and is refused. Closing settles every in-flight callback so no promise dangles. The claude-agent-sdk-control-bridge that rebuilt the wire frame is deleted. Outbound control maps to Query methods: interrupt() for cancel, setModel / setPermissionMode / applyFlagSettings for options, supportedModels for the model list, initializationResult() for init proof, each under Orca's own request deadline and error classification. Cancel is interrupt-receipt aware: a CLI advertising interrupt_cancel_queued_v1 gets cancel_queued in one round trip, otherwise the receipt's still_queued uuids are swept with cancel_async_message so a cancelled turn cannot spawn a later unexpected turn; older CLIs resolve no receipt. Init keeps the 10s deadline and the unauthenticated-startup guidance. Every behavior is failing-first and ablation-proven; the toggle-off import boundary and the accepted loss of unknown-control visibility rows are unchanged. Claude-Session: https://claude.ai/code/session_01Pqjduxt5G4rr9aYvtp7rNm * fix(claude): arm the descendant snapshot before stdin closes and make the tree verdict unproven by default A healthy Claude root leaves within the graceful window, and the close ladder only snapshotted descendants when the root was still alive after that window. So the common close never looked at the tree: `treeExited` stayed null, `!== false` passed it, and close() reported a proven exit with an MCP child still running. A root that died before the walk made the snapshot vacuous too. The proof is now unproven by default. The reaper holds one verdict in Orca's vocabulary (exited / live / unverifiable), assigned in exactly one place from the bounded verification, and close() returns true only on `exited`. The snapshot is armed before stdin closes, while the root can still be walked, and is verified after the root exits; a root that left before any snapshot could be armed stays unverifiable rather than vouching for descendants it never showed us. The shared verifier gains the three-way verdict behind its boolean face, and the connection reports the root and tree verdicts separately along with the child's exit status. One verification per close attempt: the retried close re-verifies, so the intra-attempt re-reap is gone from the teardown budget. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * fix(claude): verify the Windows tree after taskkill instead of trusting that it ran `terminateWindowsProcessTree` resolves from taskkill's callback whatever the error says, so a timeout, an access denial, a recycled root and a surviving descendant all looked identical to the reaper — which then returned a proven exit unconditionally. close() reported true and the lease was released with an MCP descendant potentially still live. The Windows branch now snapshots the root's descendants while it is alive and, after taskkill, polls a fresh process table to a bounded deadline: a row still matching by pid AND creation time is `live`, an unreadable table is `unverifiable`, and only a table with no match is `exited`. Creation time is the PID-reuse guard the POSIX path gets from ps lstart, so a descendant that denied a creation-time query is omitted rather than signalled on a bare pid. A root already observed exited is never taskkilled: `/T /F` on a recycled pid would take an unrelated tree down with it. The captured tree is tagged by platform so neither verifier can be handed the other's rows. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * fix(claude): release a reservation on a first-hand root exit instead of latching it into manual recovery Making close() strict about the descendant tree exposed a second defect at the same boundary. A create-time acquisition has no ownerProcess until publication, so an unproven cleanup mapped to handoffStage `manual-recovery`, and adjudication then refuses every later attach with agent_session_ownership_unknown. A user who was merely signed out, or whose --resume the CLI rejected, wedged the session id permanently. Each question now answers from its own evidence. close() is unchanged and stays strict about the tree. Separately, the lease is keyed on the root's pid and start time, so when Orca's own child handle observed that root exit and no descendant snapshot was ever admissible, the reservation is released and the CLI's exit code and stderr reach the user. A descendant observed still alive, or a root Orca never saw leave, stays unproven and keeps the reservation. The settlement records only what was observed: the released lease says the provider process exited and its descendants were not verifiable, rather than reusing the wording that claims cleanup proved no child remains. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * fix(claude): surface an API error a result frame reports instead of settling the turn on it The SDK models an API failure as a SUCCESS-subtype result whose `result` string is the user-facing error text, with no assistant frame behind it. The translator suppressed every catalogued result subtype as turn bookkeeping, so that turn tombstoned its lifecycle and showed the user a completed, empty reply with no sign anything had failed. Suppression is now by meaning. A result reporting a failure routes to the bounded provider-error surface, leading with the provider's own sentence and keeping the raw frame behind the row's disclosure; ordinary successful results stay off the timeline as before. A turn the user aborted also stays suppressed: its interrupt frame already says so, and its execution diagnostic would only be noise on every stop. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * fix(claude): drop the stream state of turns that never received their final frame Every streamed delta recorded its block's identity, latest text and checkpoint length. Only the final assistant frame removed them, so an interrupted turn left its whole accumulated reply reachable until the session was disposed, and a long session with repeated interruptions grew those maps without bound. The partial text was already journaled by the flush that precedes settlement, so the live copy was pure retention. That state now lives in its own module, named for what it does — grow a streamed block's journal row between its deltas and its final frame — and turn settlement drops every block still awaiting a final. The translator reports how many remain, which is the invariant: a settled turn leaves none. Also makes a timed-out process-table read retryable while the root is still alive. A loaded host can miss the table's one-second deadline, and latching that as "no descendants" both lost the descendant sweep and, on a busy machine, made the close ladder report unproven for a tree it never actually looked at. Only the root's death still makes a missing snapshot final. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * perf(claude): capture the Windows descendant tree from one process-table read The capture walked the descendant tree and then read the table again for the creation times the walk's projection drops. Each read is bounded in seconds and both run inside the close ladder's budget, so the second one cost the worst-case teardown three seconds for data the first read already held. The walk is now exported from the module that owns it and runs over rows the caller has already read, which is also what lets the snapshot keep the PID-reuse guard the projection cannot carry. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * fix(pty): spend the descendant verification window instead of surrendering on one slow table read The verification abandoned the whole check the first time a process-table read missed its own one-second deadline, with seconds of its window still unspent. On a loaded host that reported a tree unverifiable without ever having looked at it, which the Claude close ladder then turned into an unproven close and a retried teardown. It also made the descendant-exit tests flake under a parallel suite run, for the same reason and with the same honest-but-premature verdict. A read that missed its deadline is now simply not an answer: the loop waits and reads again until its own deadline, and only a window that ends without a readable table reports unverifiable. This can only turn a premature verdict into one backed by evidence; it never manufactures a proof. Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9 * fix(claude): never let a later failed look collapse an observed live descendant into unverifiable The reaper's single assignment site latched only 'exited', so a second reap whose table reads all missed their deadline overwrote an earlier completed verification's 'live' with 'unverifiable'. The acquisition release gate discriminates on exactly that pair, so a root exit after such a decay released the lease over a descendant that had been observed alive. The latch is now monotone in trust order: exited is final, and live is only ever raised to exited. Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP * fix(claude): never prove a Windows tree gone while a descendant denied identification The Windows snapshot dropped rows that denied the creation-time query, and an emptied snapshot was judged exited without any table read: a descendant Orca was refused information about was treated as one that had left. The snapshot now counts the unidentified rows it saw, and verification caps its verdict at unverifiable while any exist. Nothing is ever signalled on a bare pid, as before. Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP * fix(claude): classify cleanup after a first-hand exit as a root exit instead of a proven tree When the CLI died between a successful acquire and the host's commit or proof of the lease, handleExit had already removed the session, so releaseAcquisition found nothing and reported true. The attach flow then settled exit-proven with deathEvidence claiming cleanup proved no provider child remains, though the tree was never verified. The adapter now keeps the exit that removed a published session until the session is acquired again; acquisition cleanup runs that connection's close ladder and classifies its verdict exactly as a start-time failure would be, so the record reads root-exit-observed. The wire helper keeps that typed classification and its provider diagnostic instead of wrapping it as unproven, and the router gives up its owner even when the release throws. Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP * fix(claude): integrate SDK teardown and picker lifecycle fixes * fix(claude): preserve resume leaf and settle processless spawns * fix(claude): reacquire from persisted resume leaf * fix(native-chat): restore Claude grouped question handling * fix(claude): persist only resumable transcript leaves * fix(claude): recover structured session exits safely * fix(claude): close remaining structured session P1s * fix(claude): harden transcript branch proof * Remove superseded root fix reports * fix(windows): restore indexed descendant row walk * fix(router): forward force-close lifecycle * fix(claude): fence stale turn cancellations * fix(claude): fence cancellation after unknown dispatch * fix(claude): fence replay and option recovery races * fix(claude): block replay fallback after waiter eviction * fix(claude): fence evicted slash results * fix(claude): fence ambiguous results and restore options safely * fix(claude): scrub SDK child env and localize pending launch * fix(claude): pin transcript roots and exit recovery proofs * fix(claude): retain unproven SDK exits * fix(claude): settle retained exit before reacquire * fix(claude): resume from settled retained cursor * chore: remove tracked review artifact * fix: harden Claude SDK transport session cleanup * fix: close Claude sessions safely * fix(claude): close races with fresh child snapshots * fix(claude): fail closed on recycled child identities * fix(claude): gate root cleanup on process identity * fix(claude): fence same-second root identity reuse * fix(claude): restore the root SIGKILL fallback the identity gate took away The direct root kill goes through the handle Node owns, not through a pid: libuv drops that handle in the same turn it reaps, so the signal either reaches the process Orca spawned or reaches nothing at all. Gating it on a process-table probe therefore bought no safety and cost the tree its only fallback whenever the probe declined -- a first capture landing in the fork's own second, a recycled descendant pid voiding the snapshot, or a process table that could not be read on either platform. Identity verification stays where a bare pid is genuinely addressed: Windows `taskkill /T /F`, and the descendant sweep's own revalidation before it signals. Also stops a declined root probe from collapsing an observed `live` or `exited` descendant verdict into `unverifiable`, and stops a successful taskkill from reporting `unverifiable` because a later probe found the root correctly dead. * docs(claude): rewrap the root-kill ordering comment * Match the Claude structured launch to the terminal path's managed-account auth rules The SDK path stripped ambient Anthropic auth unconditionally, let an explicit agentDefaultEnv override beat a pinned managed account, and had no account-switch guard. Reuse the terminal preflight's own predicate and messages so both transports strip, refuse, and report identically, and cover the CLI transcript location that mobile native chat depends on. * Reach the Claude structured chat lane from the desktop UI The main process has had a complete, correctly gated Claude Agent SDK lane for a while, but no renderer ever asked for it: the launch route accepted only `codex`, and the create path was typed `agent: 'codex'` end to end. Widen both to the structured provider union that already exists (`AgentSessionHandleProvider`), and generalize the codex-named create path instead of adding a Claude twin beside it. The pending-launch registry is now keyed by agent as well as workspace — a shared key handed a second caller the first agent's intent, so a Claude and a Codex launch in one worktree collided. Windows, per agent. Codex's client-side win32 refusal is deliberate and settled elsewhere, so it stays exactly as it was. Claude's answer is no longer guessed from the client's platform: a structured session fences its provider child on that child's process start time, and only the executing host knows whether it can read one. `agentSession.createSupport` already answers precisely that, per agent, and had no renderer caller — so the Claude create path asks it before creating and turns a "no", or a probe it cannot get answered, into the definitive refusal the launch fallback already handles. Fail closed either way. That refusal mapping also closes a real gap: the host reports an unsupported location by throwing `structured_agent_session_unsupported`, which reaches the client as a transport rejection rather than a refusal envelope, so `StructuredAgentSessionCreateRefusalError` never fired. The launch would retry the create, strand itself in `visibilityUnknown`, run no legacy fallback, and show an error toast. Close a fail-open hole while Claude and win32 become reachable: `create` with a client-supplied location, and `ensure`, both skip the worktree-resolving support check. They now ask the executing host the same question directly, so a host that cannot fence a provider child no longer creates one on a client's say-so. Also deletes `structured-agent-session-provider-routing.ts`, a duplicate of `structured-agent-session-provider-support.ts` with no importers. WSL, SSH and paired hosts, floating workspaces, draft prompt delivery, explicit TUI customization and initial session options all keep refusing; folder workspaces keep working. * P1-1: make the structured Claude auth policy required and testable The optional dep plus a {stripAuthEnv:false} fallback meant a dropped wiring under-stripped silently. Required at all three hops, asserted at install time for the @ts-nocheck caller, and the settings-to-policy mapping is now a named tested function. * P2-3: mobile's default Claude transcript root must follow CLAUDE_CONFIG_DIR session-file-resolver's default ignored the variable the pinned account home follows, so a CLAUDE_CONFIG_DIR launch wrote one tree and mobile read another. The Task-4 test now resolves with no root override (mobile's own call) and checks the answer against the root the CLI itself reports, instead of mirroring the code under test's own expression. * P2-1/P2-2/P3: close the teardown window, join the live-auth gate, align the refusal P2-1: a switch beginning inside the acquire teardown left a dead chat and no replacement. Past that point the launch waits the swap out and refuses only if it never settles; the entry guard still refuses outright, because nothing is torn down there yet. P2-2: structured children now hold the same OAuth-refresh gate a Claude PTY does, so a managed refresh cannot rotate the token out from under a live turn. P3: the refusal now matches the strip it guards (case-folded on win32, presence not truthiness), and the dead structured-to-TUI builder states its auth policy instead of silently signing a system-auth user out. * Make the live-auth gate tests independent of sibling connection teardown order * Do not offer structured Claude under a WSL-only managed account Structured Claude launches against the ambient Claude config, which the account service keeps in sync with the selected HOST account. A WSL-bound managed account lives inside the distro and is never synced there, so on Windows a structured session would authenticate as whatever the ambient identity happens to be while the UI names the WSL account — the user is told one identity and given another. That was unreachable only because nothing offered structured Claude on win32. Enabling it makes it reachable, so gate it here rather than patching the auth layer: refuse the structured path when the active managed Claude account is WSL-bound, and let the terminal-backed path — which resolves the account per runtime — handle that account shape. The answer rides the agentSession.createSupport seam the renderer already consumes, so no new capability and no renderer knowledge of account internals. A create the host declines becomes the definitive refusal the launch fallback already turns into a legacy native chat tab, with no error toast. Unknown answers refuse. An install with no managed accounts claims no identity and is fine, but an active selection that cannot be resolved — or account state that cannot be read at all — is not evidence that the ambient identity is right. Claude only. Codex resolves its account through a different path and its createSupport answer is untouched, as is every Codex routing decision. * Read the structured Claude account gate through the auth policy's accessor The gate resolved the active account from the account-service snapshot's runtime map; the auth policy resolves it with getSelectedClaudeAccountIdForTarget(settings, { runtime: 'host' }). Those are two sources and two resolution rules, and they disagree on a legacy settings blob that carries the selection only in the flat activeClaudeManagedAccountId: the accessor falls through to it, a direct read of the runtime map does not. The gate would then refuse a launch the policy would have run under host-1 — and in the mirror case a session could be admitted under a policy computed from a different account than the gate approved. Read the same settings through the same accessor so agreement is structural rather than coincidental, and drop the controller accessor that existed only to reach the snapshot. No behaviour change for any state both already agreed on; Codex is untouched. * Round-3 review fixes: N-1 empty-value regression, N-2 gate leak window, N-4 lost history N-1: my presence-based conflict predicate refused a terminal launch that works today. 'ANTHROPIC_API_KEY=' is how a user blanks a variable and the settings pipeline preserves that empty value; an empty override cannot beat the pinned account and the strip removes the name anyway. Back to truthiness for the value, keeping the win32 case folding. N-2: enter the live-auth gate only after the exit/close handlers that release it, so no throw in between can leave an entry nothing reconciles. N-4: the Claude transcript resolver searches config-dir-then-default and de-dupes, matching the Codex sibling in the same file, so adopting CLAUDE_CONFIG_DIR no longer hides history written before it. * Run the managed-account gate on every Claude acquisition, not just create createSupport gates the create path, but a session's account state can change while it lives. A reacquire after an unexpected child exit re-resolves the launch and re-derives auth, with nothing re-checking the gate — so a session created while supported could come back up in the refused shape. With the strip predicate keyed on there being an active non-WSL account, the WSL-only user's normalized steady state (accounts exist, none active) does not strip, and that reacquire reaches the child with ambient auth while the UI names the account. Gate at resolveLaunch, the one choke point every acquisition passes through, refusing with the pre-spawn error the caller already handles. Same predicate as create-time, now sharing one settings reader so the two cannot drift. Claude only; Codex resolves its account on a different path and is untouched. The runtime class that wires this does not typecheck its own `this` calls — a missing hookup compiles clean — so the wiring is pinned behaviourally rather than trusted to the compiler. * Move the structured Claude gate out of the @ts-nocheck runtime files Both call sites of the managed-account gate sat in files whose first line is `// @ts-nocheck`, so neither was typechecked: three arguments to a one-argument function plus an undeclared identifier compiled clean. New auth-identity decision logic had no compiler behind it. Move the verdict into a checked module that takes the two facts the runtime owns — the adapter's answer and a settings getter — and decides. The runtime class now only forwards. Move the gate reader's construction into the checked installer too, so the nocheck file passes a plain settings closure and never names a gate symbol. Every reference to the gate predicate and its reader now lives in a checked file, so the ablation that used to pass silently is a compile error at both the create-support and reacquire sites. Removing the file-level @ts-nocheck is a separate, larger job and is not attempted here. * Derive the gate test's auth policy from the settings under test A hardcoded stripAuthEnv asserts a gate/policy pairing production cannot produce, and false additionally lets launch.env inherit the runner's real process.env. Derive via claudeStructuredAuthPolicyForSettings instead: the gate settings type is the same Pick the policy takes, and both resolve the account through getSelectedClaudeAccountIdForTarget. * Pin the absent-vs-empty distinction in the managed-account gate An empty claudeManagedAccounts array is a real answer: the user has no managed accounts, nothing claims an identity, and the ambient path is legitimate. A readable settings object with no such field is settings we failed to parse — the same unknown as unreadable — so it refuses. The two are one character apart in the code and the difference is invisible without the reasoning, so record it at the branch and pin both sides. The test fails under the obvious "consistency fix" of treating a missing field as empty. * fix(claude): keep command queue bookkeeping out of the transcript Claude Code 2.1.258 emits a `command_lifecycle` frame for every uuid-stamped command it starts, completes or cancels. The frame carries a command uuid and a state and no content, and the CLI keeps it out of its own transcript -- but it is absent from the SDK's SDKMessage union and so from Orca's frame catalogue, where an uncatalogued kind defaults to a substantive row. Every structured turn therefore painted raw JSON rows into the user-visible transcript. Catalogue it and disposition it as status chrome. The unknown-kind default stays `timeline-substantive`: a kind we have never seen is likelier to carry content than to be chrome, and a visible row we can catalogue later beats content we silently dropped. A lifecycle state that reads as a failure still surfaces, because the payload error check in `classifyProviderFrame` outranks the catalogue. * fix(claude): let a re-walked descendant become eligible for the forced sweep A descendant first observed by a capture inside its own birth second could never be SIGKILLed: `ps lstart` is second-resolution, so that capture cannot rule out a pid recycled later in the same second, and the merge pinned each retained row to the boundary of the walk that first saw it. SIGTERM-resistant children forked in that window were signalled and then never escalated -- they survived close, quit and restart, reparented to init, and had to be killed by hand. Advancing that boundary on any later capture would be unsound: a later capture matching pid, pgid and start-second is exactly what an impostor would also show. But a capture is not a match -- it is a fresh ppid walk from a root Node pins through its own handle, so a row it re-derives is proved ours at that instant without appealing to its start time. Chain the fence from there instead, and take that walk at the close boundary while the root certainly still lives: the root may leave inside the grace window, and the post-timeout refresh never runs. A row absent from the later walk still keeps its earlier boundary, and a row no walk has ever re-derived in a later second is still never escalated. * Treat an absent managed-account list as empty, not as unreadable An empty claudeManagedAccounts array and a missing one are the same answer: this user has no managed Claude accounts, so nothing claims an identity and ambient auth is the truth. Refusing on absence strands any profile that simply never wrote the key, and it disagrees with the auth policy, whose own predicate takes `(accounts ?? [])` for exactly this reason. Only settings that cannot be READ stay unknown, and those still refuse — as do a WSL-bound active account and a selection naming an account the list does not explain. The earlier reasoning treated a missing field as settings we failed to parse. That conflated "not present" with "not readable"; only the second is unknown. * Support structured Claude when accounts are registered but none is selected Registered-but-deselected Claude accounts were refused, which is behaviourally identical to having no accounts at all: the auth policy does not strip, ambient auth is the truth, and the UI names no host identity. A user who deselected their accounts silently got legacy chat with nothing explaining why. Nothing selected for the host runtime is two states the settings cannot tell apart after the fact, because pruneInvalidClaudeRuntimeSelection empties the host slot and persists null in the second one: honest deselection -> ambient auth, UI names nothing -> SUPPORTED the WSL-only steady state -> ambient auth, UI names the WSL account -> REFUSED The presence of any WSL-bound account in the list decides. Simplifying this to "none active -> supported" re-opens the auth-identity misrepresentation, so the tests fail loudly on exactly that: five of them, across the unit rule and the createSupport path. * Stop treating an unanswerable create-support probe as a refusal A worktree is not resolvable for a beat after createWorktree resolves, so a probe fired immediately after creation fails the RPC with selector_not_found instead of answering. The catch collapsed that into `supported = false`, so the composer refused and quietly built a terminal session — the gate never said no, it was never asked successfully. Elapsed time was the only input that decided whether a Claude launch went structured. "Could not answer" and "answered no" are different states and only the second is a verdict. Retry while the host cannot yet resolve the selector, with a bounded backoff that covers the measured window with margin, and keep refusing on the first ask for everything else. Fail-closed is unchanged: a probe that still cannot be answered when the budget is spent refuses. The retry is narrowed with the shared error-code matcher, which classifies a token that transports re-wrap into a longer message without matching prose that merely mentions it. Codex never probes, so this race has never been able to refuse a Codex launch — the race itself is identical for it. Recorded at the early return, because whoever gives Codex a probe inherits the bug. * fix(claude): fence the forced sweep on re-derivation, not on lstart's second A descendant forked in the same wall-clock second as every walk that sees it was signalled with SIGTERM and then never escalated, so a SIGTERM-resistant child survived tab close, app quit and a full relaunch. Two children of one parent 96ms apart across a second boundary took opposite paths. The leak predates this branch: it reproduces with the change reverted. `ps lstart` has one-second resolution, so a walk landing inside a row's birth second can never rule out a pid recycled later in that same second. But a walk is not a match: a ppid walk only reaches what the root actually parents, and the root is pinned by Node's own handle, so a row the walk re-derived is ours whatever second it was born in -- a stranger would have to have been forked into our tree, and then it is not a stranger. Fence the escalation on that. Rows a merge retained from an earlier walk are not re-derived and still answer to the start-time fence, which remains correct for them. Scoped to callers that revalidate identity before signalling, which is the Claude close path. Codex teardown reaches this same verifier and is unchanged; the argument holds there too, but widening it is its own deliberate change. Also reverts two changes from the previous attempt at this leak. Advancing the capture boundary on a later walk is inert once the sweep fences on re-derivation -- both key on the same set of rows, so the new term short-circuits for exactly the rows whose boundary it advanced. The extra ladder refresh was a duplicate full process-table read: close() already awaits tree.refresh() immediately before proveClaudeChildExit, on the only path that reaches it. Known property: the kill lands roughly a grace window after the walk that proved membership, so a pid recycled inside that gap could in principle be signalled. It is bounded -- matchingSnapshotRows already requires the live row to carry the same start-second and pgid, so an impostor must be born in the remainder of that one second, land on that exact pid, and sit in the same process group, and it has already received the unfenced SIGTERM from the same loop. * Run the Claude structured integration suite as a runtime client The suite exercises agentSession.* for Claude, not the mobile surface: nothing in it asserts anything mobile-specific and its sibling integration suites use 'runtime'. Mobile now additionally requires the experimental structured-chat setting, which structured-agent-session.test.ts pins in both states, so the stale 'mobile' fixture was claiming coverage it never had. * fix(claude): report effort from get_settings, which is the only frame that has it The composer's Effort pill rendered blank in every structured session. This is not a missing source: the publication reads `effortLevel` off the `system/init` frame, and that frame has never carried an effort of any kind, while the correct value is already fetched at acquisition and thrown away on the auth diagnostic. Verified two ways -- a live get_settings probe against Claude Code 2.1.258, and the shipped binary's own init frame construction, which lists `model` and no effort. So `reportedOptions.effort` was always empty, the options reader dropped the key, and the pill had no value. Model survived only because `currentModelId()` has a fallback chain. The get_settings call acquisition already makes reports the session's current effort as `effective.effortLevel`; pass that into the publication instead. Selecting an effort already worked, so this is the arrival value only. The legacy PTY path is unaffected and must not be "fixed" to match: it reads its effort by parsing the startup banner (`CLAUDE_MODEL_EFFORT` in src/renderer/src/components/native-chat/claude-terminal-session-options.ts), which is why it shows a value where the structured path does not. Also removes the fixture that hid this: the fake init frame invented `effortLevel: 'high'`, a field the CLI does not send, which is why every gate stayed green over a value that is always empty in production. The fixture's get_settings now returns the real {applied, effective, sources} shape instead of a bare `{env: {}}`, so the two adapter tests that asserted an effort keep asserting it through the path production actually uses. The reader returns null rather than defaulting: an effort nothing measured would repeat the fixture's mistake, and a blank pill is the honest degradation if the provider ever renames the key. * fix(claude): only record an effort the child confirms it adopted apply_flag_settings answers `success` for an effort it then ignores. Measured against Claude Code 2.1.258: applying `bogus-effort-xyz` returns subtype "success" with no error while `applied.effort` stays at its previous value, and a valid `low` moves it. The option write treated the absence of a throw as adoption and recorded the requested value unconditionally, so Orca would show and persist an effort the child was not using, with nothing anywhere reporting a problem. Read the effort back after applying it, through the same reader the arrival value uses, and reject when the child reports a different one. A readback that could not be taken is not evidence of a refusal -- the apply itself succeeded -- so it still records; only a readback that disagrees rejects. Not reachable from today's picker, which offers catalog values only, but the CLI's effort catalog is server-delivered and has changed before, so a retired id would otherwise become a pill confidently displaying a setting that never took. * test(claude): assert the effort contract against the real binary The blank pill survived every gate because the only tests that touched it were fixture-backed, and the fixture invented the field. A test that pins the shape we read cannot catch the provider renaming the key, which is the failure mode that produced this defect. Asserts both halves against a live authenticated CLI: that no frame it publishes carries an effort at all, and that the session's current effort arrives through get_settings. Which frame proves the session varies by host -- this machine proves it with a SessionStart hook rather than a system/init frame -- so the negative half asserts over every published frame rather than picking one. Skips with the rest of the file when no authenticated CLI is present. * fix(claude): stop the synthesised content-part kinds leaking into the transcript Sending an image put a bare `claude · message:user:content:image` row between the user's bubble and the answer. Two causes, and only the second is a family. An image part counted as modelled only when `source.type === 'url'`, but claudeDispatchMessageContent sends a local attachment as a base64 source and the CLI replays that shape back, so every attached image was classified unmodelled. Accept the base64 and file sources Orca itself sends. The family is the real defect. `message:<role>:content:<type>` kinds are synthesised at runtime from whatever `part.type` arrives, so unlike the top-level frame catalogue they can never be enumerated ahead of time -- the `?? 'timeline-substantive'` default then prints the synthesised name at a user who cannot act on it. That default is right for top-level frames, where "substantive" means show the frame; here it meant show our own vocabulary, which drops the content AND leaks the opcode. So an unrenderable part now renders a sentence saying exactly that, with the kind and payload still on the row's disclosure. A part that carries its own readable sentence keeps it -- the placeholder is a fallback, not an override. An unknown future part type is therefore visible, never silently dropped and never printed as a kind: the same principle as the effort readback, which records only what the provider confirms. * Declare agentSession.requestHandoff on the cross-version wire surface The manifest is a ratchet for cross-version reachability, so the method is declared with real HandoffParams rather than counted. requestHandoff is capability-gated through requireStructuredHost and has no client caller, so declaring it is the whole of the change. Also model two host capabilities the harness omitted: the stub host's supportsCreate, and the fake adapter's, without which adapterSupportsCreate falls through to a supportsLocation the fake also lacks. Every ensure was refused for the harness's silence rather than for its location. * Gate structured Claude session tabs on the client capability that names them The Claude structured lane deleted the projection's `agent !== 'codex'` filter and added CLAUDE_STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY in the same commit, but never wired the constant to anything. Paired clients then received agent-session tabs for Claude, which no shipped client renders -- mobile's resolveMobileNativeChat returns null for every agent but codex, so the row listed and selected into a pane with neither chat nor terminal. Restore the filter behind the declared capability instead of the bare agent name. No client advertises it yet, so this matches main's behaviour today and becomes a negotiation a future client can opt into. * Confirm the structured Claude model against the model the CLI reports set_model answers success for any string, including a model it cannot resolve — the failure only surfaces when the turn runs — and get_settings reports the settings-file model, not the session's. The init frame that opens each turn is the only channel carrying the adopted model, so keep the session's reported model current from it instead of reading it once at acquisition. Also stop rejecting an effort the readback cannot represent: max is session-scoped and excluded from the persisted effortLevel, so a readback reporting the level underneath it is an absence of evidence, not a refusal. * Clear the session-option hedge when the provider confirms the value The pill claimed every option was unconfirmed for the life of the session: the renderer recorded each write as dispatched and nothing ever moved it, so a model the CLI had already reported back still read as unconfirmed. Carry the provider's own confirmation to the surface. Main reports which option ids the provider named rather than merely accepted, and the client re-reads options as a turn changes, because the frame that opens a turn is where the adopted model arrives. A value the provider has not reported stays hedged, including an effort whose readback could not be taken. The confirmed list is optional on the wire: a host that predates it sends nothing and the client keeps hedging, which is the behaviour it had. * Keep the model report current across an acquisition fence bump * Show the picked session-option value and let the provider report correct it The pill showed a "not confirmed" second tooltip line for any value we had sent but not yet seen reported back. Nothing acts on it, and for the PTY lane it was permanent — that transport has no report channel. The pill now shows the picked value immediately and the provider's per-turn report corrects it when the two disagree; a newer local write still outranks a report that precedes it. `dispatched` stays as a provenance member rather than collapsing into `applied`: it is produced independently by the PTY lane, and it is where the `confirmed` wire field lands, which would otherwise be unobservable. Effort keeps its readback and its rejection path. That matters more now, not less: with the hedge gone the rejection is the only user-visible failure signal on this surface, so a spurious one would be the loudest bug here. Skipping the readback for an effort the settings response structurally cannot echo is what prevents it — the response carries the persisted level, so reading it back for a session-scoped value would report the level underneath and fail a valid write. * Hedge a session-option value only when the terminal transport sent it Both lanes emit `dispatched`, so it could never say which one produced a value. The descriptor now carries the transport that built it, set once in the shared snapshot builder from a parameter that is required rather than defaulted — the builder is the only place a descriptor is constructed, so a new producer has to name its lane or fail to compile. The structured lane confirms every value from the provider's own per-turn report, which makes the hedge transient noise there. The terminal lane can only learn an outcome by parsing the screen back, and only for Claude: every other agent's `dispatched` value stays unconfirmed for the life of the session, so the line is the only signal that we sent something we never saw land. * Refuse an effort the session's model advertises no control for * Refuse tab mutations on a Claude row the client never negotiated The branch added a case asserting a client advertising only agent-session.structured.v1 may mutate a claude row. That is the same ungated behaviour the projection gate removes, encoded a second time — mutation authorization reads the projection, so hiding the row refuses the write. Assert that contract instead, and add the positive case for a client that does negotiate Claude rows. * Resolve the Claude session's current model in one place so the effort guard and the pill agree * Record an effort the child did not adopt instead of refusing the write apply_flag_settings answers success for an effort it then ignores, so the readback exists to detect that. Refusing on it made the detection a veto, and a veto is only correct if the readback can never be wrong about which model is current -- which it was, twice. The pre-flight guard already refuses a level the model advertises no control for, so the veto guarded a door that is now locked upstream. Keep the detection, drop the refusal: a disagreement records the child's own answer and omits the option from confirmed, so main stops vouching for a value the provider rejected without blocking the user's write. * Stop a slow whole-machine ps from being read as an absent process `ps -axo ...command=` pays a per-pid argv read: measured 1.15s for 1,948 processes (0.03s without `command=`), and CPU contention stretched the same capture to 6.0s. Two budgets sized for a cheap look then misreport a readable machine. The reader's 3s ceiling killed 6 of 20 consecutive captures at load 27, so every consumer answered "unverifiable" about a table it could read. Raise it to 15s, and stamp the capture instant at ps START so `capturedAgeMs` is the upper bound its contract promises -- a 6s capture used to report itself as freshly taken, understating staleness against a 5s kill gate. The TTL keys on completion so a slow capture still coalesces instead of forking ps per caller. `readStructuredTuiProcessIdentity` then spent its whole 5s wait inside one capture and concluded "no exact child" after a single look taken before the child existed (observed landing at ~3.5s). Absence needs a look that did not race the spawn, so require two captures before the deadline can end the loop. Both surfaced by the real-binary Claude TUI resume test, which failed ~1 in 5 under load; 14/14 now, 8 of those runs containing a capture the old 3s budget would have killed. * Let the desktop renderer negotiate Claude structured tabs The paired-client gate hides agent-session rows an agent the client cannot render. The desktop renderer's own IPC dispatches as clientKind 'runtime' advertising only agent-session.structured.v1, so the gate hid Claude rows from the surface this feature ships on. It renders them; it should say so. * Stop a slow process table from silently blinding every freshness gate Stamping `capturedAgeMs` at ps START made the number honest, and honest broke both consumers that read it. `ps -axo ...command=` measured 2.5-9.0s on an idle 2,002-process laptop and 4.0-18.6s at load 46, so the age it now reports lands past every budget: `planRelayPtySweep` refuses the stop as "too old", and the renderer's `admitRemoteForegroundEvidence` refuses the record outright. That second one is the expensive half and was outside the diff -- a refusal bumps `consecutiveInspectionErrors`, the poll scheduler backs off to its 10s floor, and agent-completion detection stops for the pane. The subsystem went blind on exactly the loaded hosts the honest stamp was meant to serve. The evidence-publishing read now gives up at 1,200ms instead of waiting out `PS_TIMEOUT_MS`. It is one budget for one question: these consumers ask whether an observation describes NOW, and past this it does not -- a late answer is refused by the age gate anyway, having first blocked a polled path for the whole capture, so a prompt `unverifiable` is both the truthful verdict and the cheap one. Both relay call sites already produce it from a rejection, and an admitted `unverifiable` costs a poll where a refusal costs the cadence. Identity proof keeps the full 15s through `getFreshProcessTableSnapshot`, because it asks whether a process EXISTS and must never read slow as absent. The budget bounds the wait, never the capture: the reader coalesces, so an abandoned wait leaves its capture running to fill the cache rather than forking a second whole-machine `ps` on the host that can least afford one. 1,200ms is bracketed rather than picked. The floor is the capture's own cost -- `command=` measured 1.15s for 1,948 processes on an idle host, and a budget under that answers `unverifiable` about a machine nobody is straining. The ceiling is the consumer's: 2,000ms, less the 500ms a TTL-shared capture may already have aged, leaves 1,500ms, and transit takes the rest. That ceiling only fits once the capture stops being charged twice. `ps` runs inside the RPC round trip, so its duration is already in `receiveDelay`, and `capturedAgeMs` is that same duration on the host's clock; summing them halved the budget this gate grants a host from ~2.0s of `ps` to ~1.0s, which is why a 1.2s capture arriving at 1.3s read as 2.5s old and was refused. Admission now takes the larger of the two. The sweep's gate keeps its sum, which is correct there: `evidenceAgeSinceListingMs` is stamped after the listing ARRIVES, so it measures planning time and overlaps nothing. A stated limit rather than an assumed one: 15s is not proven sufficient for identity proof. The same capture reached 18.6s at load 46, so that path can still time out and answer "no exact child" about a host it simply could not read in time. Narrowing it needs a cheaper question than a whole-machine argv read, not a larger number. The one test guarding this field could not fail. `beginPtyHandlerTest` installs fake timers, so `Date.now()` is frozen, the real reader reports exactly +0, and `0 <= 500` held identically for a hardcoded zero, for completion-stamping and for start-stamping -- while the real reader on that host returns thousands of ms. It now drives a measured age in and asserts the handler publishes it rather than restamping; that the reader MEASURES it correctly stays pinned separately, against a controllable clock. Both consumers get boundary coverage either side, and each new gate was ablated red before it went green. * Keep the compatibility fields off the capture the budget just abandoned inspectProcess falls back to processHasChildren and listProcesses to getForegroundProcessName, and both read the same TTL-shared capture with no budget of their own. On a slow host they joined the in-flight capture the budgeted evidence read had just given up on, so the call still blocked for the full 6-18s and the budget bought nothing -- once for inspectProcess and once per managed PTY for listProcesses. Use the degraded answers those helpers already give for an unreadable table, reached promptly. pty.hasChildProcesses keeps its unbudgeted fresh probe: it is a one-shot destructive gate that can afford to wait. --------- Co-authored-by: Merge Sim <merge-sim@local> Co-authored-by: Merge Sim <sim@local> |
||
|
|
f4c2821167 |
refactor(agent-session-journal): move the session journal onto SQLite (#18652)
* refactor(agent-session-journal): move the session journal onto SQLite The agent-session journal kept its state in three hand-rolled file formats: an append-only `log.jsonl` with torn-tail repair, a `snapshot.json` holding folded state plus a retained tail, and byte-quarantine files for anything unreadable. This replaces all of it with one SQLite database per session — `journal.db` beside the existing `blobs/` store — using the in-house adapter and the open/pragma/migrate/harden pattern the orchestration database already follows. Two tables: `journal_rows` (the append-only log, keyed by `(session_id, epoch, seq)`) and `journal_sessions` (the derived projection, upserted in the SAME transaction as every insert). Rows stay JSON in one column, so the row schema, the version upcast chain, and the reducer survive byte for byte — `journal-reducer.test.ts` and four other suites pass unchanged and are the regression proof. Deleted: `journal-log-file.ts`, `journal-compaction.ts`, `journal-corruption-quarantine.ts`, and the public `compact()` / `compactionBoundary` / `autoCompact` members, none of which had a non-test caller. Existing `log.jsonl` / `snapshot.json` journals are deliberately abandoned. No importer: a session created on the old path stops working, which is acceptable because the feature is off by default. ## The physical quota is repriced, because SQLite does not charge like a file The 256 MiB per-session bound is unchanged, but the arithmetic under it could not survive: SQLite grows the database in pages and the WAL in frames, and the checkpoint that copies the WAL forward holds the same pages in both files at once, so a transaction's peak is about twice its content. Admission now charges the candidate transaction's own measured page cost, validated against a sweep that runs as a regression test (`journal-database-space.test.ts`) rather than derived from reasoning about the allocator. Four things are load-bearing rather than tuning, each measured: - `auto_vacuum = INCREMENTAL` must be set BEFORE `journal_mode = WAL`. Set it after and it is ignored with no error, reclamation silently becomes a no-op, and the file never shrinks again. Both halves are asserted. - `wal_autocheckpoint = 0` plus an explicit `wal_checkpoint(TRUNCATE)` at the end of every write path, so the one moment the same pages live in two files is a moment the charge accounts for. - Reclamation runs in bounded chunks. A single unbounded `incremental_vacuum` took a 252 MB directory to 504 MB — the reclamation added to defend the bound would have breached it. `PRAGMA incremental_vacuum(N)` also frees exactly one page unless it is stepped to completion, which no size assertion catches, so the freed page count is asserted directly. - A blocked checkpoint leaves the WAL on disk together with the database growth it already copied, so admission charges that deferred copy explicitly. The term is zero whenever the last checkpoint succeeded, so the uncontended path admits and refuses an identical set. The epoch discard is `DELETE FROM journal_rows` with no WHERE clause, which takes SQLite's truncate optimization: measured at ~0.26% of the database in WAL bytes where the `WHERE session_id = ?` form rewrote every emptied leaf at up to 99%. One database per session is what makes the unqualified form correct. An open, empty journal costs 57,344 bytes before a single row exists, so a configured quota below `JOURNAL_MIN_SESSION_BYTES` now fails loudly at open with the existing `journal_bound_exceeded` instead of as a run of identical append failures. No production caller configures one; the affected surface is test fixtures, rescaled to the smallest value that restores what each case proves. ## One deliberate behaviour change Compaction was the only mechanism that shed bytes inside an epoch, and the write path called it precisely so an append at the bound was not refused. The SQLite-shaped replacement — a bounded prefix delete — cannot be used: with the snapshot gone the surviving rows ARE the state, so dropping the oldest of them loses the oldest transcript silently at the next reopen. So no row is ever shed inside an epoch, and a session whose row bytes alone reach the bound now refuses every append where it previously compacted and continued. A loud typed refusal beats silent data loss. What still sheds is unreferenced BLOB bytes — the dominant and unbounded byte source — on the same write-path hook. The escape from the hard stop is the fold that already exists, `replaceEpochItems`, which now actually returns bytes to the filesystem instead of leaving them on the freelist. The prune's protected set is a union of live reducer digests AND the candidate row's own digests, including those cited only by a nested lifecycle-batch mutation. Content addressing never rewrites a digest already on disk, so protecting live state alone deletes the blob the append is about to cite — a dangling reference that surfaces one reopen later as an empty expansion on an item the user can see. `journal-store-blob-budget.test.ts` pins it, and it goes red when the set is narrowed back. ## Handle ownership A file handle used to be opened and closed per append; a SQLite handle is held for the session's lifetime. Every path that can open a connection now has one owner: the open function owns its raw connection until it returns, the store owns its retained one and releases it in a new `close()`, and every other connection is closed by the call that opened it. The attach, recovery, eviction, map-overwrite and host-teardown paths close what they drop, and host teardown is failure-complete — the sink-barrier flush throws by design, so a trailing close statement would be skipped on exactly the path that leaks. `close()` has a stated contract: admission at enqueue and permanent, the close step on the same queue past that gate, one shared in-flight attempt, fulfilment terminal, and the release last and deliberately unguarded so a retry re-enters it. Guarding the release would skip it on retry, guaranteeing a permanent leak in exactly the case where it did not release. `journal_closed` joins the error union for a write after `close()`; no file outside the directory references any of these codes. * fix(agent-session-journal): make a COMMIT final, stop repairs deleting valid rows, and keep rejected closes retryable Six review findings on the SQLite journal migration. 1. A successful COMMIT is now the point of no return. The ordinary append, the epoch roll and the epoch replacement each adopt the committed row or epoch BEFORE any post-commit filesystem work; checkpoint, reclaim, blob prune and directory measurement run through `runJournalPostCommit`, which is best-effort by design and falls back to the transaction's own charge as a conservative footprint. Previously a post-COMMIT scan failure rejected a durable append and the next one reused its sequence, and a failed epoch housekeeping step left the store writing into a prefix already deleted. 2. Corruption repair preserves instead of destroying. A rejected suffix is copied into a new `journal_quarantine` table and removed from the live epoch in ONE transaction per chunk, charged against the session bound before a byte is written; a journal that cannot afford the copy refuses to open rather than falling back to deletion. The repair state is exposed as `journal.repair` and the rows are readable through `recoverQuarantinedRows()`, so Orca-owned submission, receipt and lifecycle identity survives a gap or a malformed row. 3. The physical charge covers the B-tree key payload. `session_id` and `epoch` are stored in both tables and both primary-key indexes and appear nowhere in `row_json`, so the journal boundary now bounds them and `journalTxnPhysicalCost` charges those bounds plus the projection upsert. The charge sweep runs the exact production transaction at maximum admitted key sizes. 4. A rejected `close()` no longer orphans its handle. Callers hand the journal to `agentSessionJournalCloseRetries` instead of swallowing the rejection, the attach map replacement is ABORTED when the previous journal will not close, host teardown retries what the registry holds, and a failed runtime teardown is retained so the next stop is a real retry. 5. `journalWalBytes()` returns zero only for ENOENT and propagates every other stat error, so admission and reclamation fail closed. 6. The WAL contention test closes the writer before removing its temp root and asserts the directory is removable once handles close. Regression coverage: post-commit divergence (4), corruption repair (5), key bounds (5), WAL stat (8), close retry (5), plus a runtime stop-retry case. Each fix was ablated on this head and the matching tests go red. * fix(agent-session-journal): anchor replay at sequence 1, make quarantine append-only, and charge it in bytes Three ways the corruption quarantine still lost rows it was written to keep. Replay validated contiguity from the first row that HAPPENED to remain, so an epoch missing only its sequence-1 row declared the leftovers contiguous and set no `truncateFrom`. The load was still corrupt, so recovery imported provider history and `replaceEpochItems` deleted every live row — including Orca-minted submission, receipt and lifecycle identity that no transcript can reconstruct, and that nothing had quarantined. Replay now anchors at sequence 1, so a missing epoch row rejects the whole surviving range before any replacement runs. `journal_quarantine` was keyed on `(session_id, epoch, seq)` and copied with `INSERT OR REPLACE`. A repair frees the sequences it removed and the live epoch reuses them, so a second repair in the same epoch silently deleted what the first preserved. The table is now keyed on a surrogate `quarantine_id`, the copy is a plain append, and `(epoch, seq)` is metadata; existing v1 databases are rekeyed in the migration that already bumps `user_version`. The admission charge read `length(row_json)`, which counts CHARACTERS for a TEXT value where `journalTxnPhysicalCost` expects physical UTF-8 bytes. A multibyte suffix was charged at up to a third of what it writes, which defeats the pre-write physical bound — over a megabyte on a maximum-size lifecycle batch. * fix(agent-session-journal): keep a repaired epoch anchored and stop the v1 quarantine migration doubling the file Replay validated numeric contiguity from sequence 1 but never that sequence 1 IS the epoch row. When the anchor was missing the repair set aside every surviving row, and if provider-history import then failed — a transcript that is temporarily gone is enough — the journal reopened as a clean, row-less epoch: an ordinary append took sequence 1, replay accepted it, read-restore published it as history, and automatic recovery never ran again while the user's real messages sat in quarantine. Replay now rejects an unanchored prefix, the open publishes an `unreconcilable_prefix` anchor for an epoch its repair emptied, and that anchor keeps reporting corrupt — so provider history is retried on every attach — until the timeline is rebuilt or the session writes content of its own. A repair also discloses rows it set aside when no line was unreadable at all, which is the case that removes the most. The v1 quarantine rekey copied every legacy row into the new table inside one transaction and dropped the old one. A quarantine holds whole rejected rows: a single 8 MiB row nearly doubled the database past the physical bound the open had already checked, the dropped pages only reached the freelist, and the next open refused the session it had just migrated. The v1 table is renamed and frozen instead, and reads take both generations. Table creation also moves inside the migration transaction, so a crash can no longer leave a v2-shaped database still reporting version 0 for an older build to write into. * fix(agent-session-journal): stop an empty provider transcript retiring the repair marker A transcript that exists but decodes to zero messages was imported as a success: the import published an empty `legacy_import` replacement that deleted the `unreconcilable_prefix` anchor and its disclosure, so the next probe read the session as clean and every later attach skipped provider recovery while the user's rows sat in quarantine for good. The import now leaves the epoch untouched when nothing decodes, reporting `replaced: false`, and recovery treats that like a transcript it could not read — the marker stands and a later attach with real history rebuilds the timeline. * style(agent-session-journal): merge the duplicate journal-database-space import * refactor(agent-session-journal): drop quarantine, byte bound, blob spill and rate limit Match what comparable implementations do: the journal is an unbounded append-only SQLite log with no side tables and no admission control. Corruption: the rejected suffix is DELETED rather than copied into a quarantine table. The load still reports `corrupt` and recovery still rebuilds the epoch from provider history, so the observable outcome is unchanged — only the preservation half is gone. The schema is back to one version with two tables; no v1 database exists outside unmerged commits of this branch, so the rekey migration and the two-generation read path go with it. Sequence-1 epoch anchoring and the empty-provider-transcript retry are kept: both are about the corrupt signal being correct. Size: no `maxSessionBytes`, so no page-cost arithmetic, reclaim band, incremental vacuum, lifecycle byte reservations or `journal_bound_exceeded`. `auto_vacuum` and `wal_autocheckpoint = 0` existed only to make a transaction's physical cost predictable for that charge; with the charge gone SQLite's default checkpointing is what the journal wants, and the explicit pre-close checkpoint is redundant with the one `db.close()` performs. WAL, `synchronous = FULL` and `busy_timeout` stay. Payloads: an oversized body is truncated at the existing inline cap with the existing marker and the remainder is discarded, bounded at the translation layer that already calls these helpers. The truncation point and message do not change; the content-addressed blob directory and all digest tracking do. Rate: no `maxAppendsPerWindow` and no `journal_rate_exceeded`. `JournalPayloadLimits` is now just the inline cap. * fix(agent-session-journal): mark a partial repair pending and bound multi-block tool input A repair that keeps its prefix had nothing durable to show for the suffix it deleted: a sequence gap costs no malformed row, so no disclosure is appended, and the surviving rows keep their epoch anchor. The next probe read a contiguous anchored prefix, called it clean, and the deleted stretch of timeline was never asked for again — silent loss, with the deletion already committed. The deletion now writes a `journal_repairs` marker in the SAME transaction, and replay keeps reporting corrupt while it stands. It retires under exactly the rule the emptied-epoch anchor takes: a fresh epoch carries the rebuild, or the session writes content of its own past the sequence the repair left free. The repair's own disclosure is not that content. Legacy import bounded a tool call's input only when it was the message's sole block; the multi-block path returned `tool-call` unchanged, so a mixed message from Claude, Grok or an omp execution cell persisted the whole input despite `inlineHeadBytes`. `boundBlock` now routes it through `boundToolInput`. Also drops canonical comments describing quarantine, snapshot files, blob storage and blob compaction — none of which exist any more. * fix(agent-session-wire): stop awaiting the synchronous journal probe loadJournal runs on a sync-database connection and returns JournalLoad | null, so both wire call sites were awaiting a non-Promise. The type-aware code-quality gate flags it; the native gate does not. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
872bd51d47 |
fix(native-chat): reland large structured command results (#17720)
* fix(native-chat): preserve large structured command results (#17707) * fix(native-chat): preserve large structured command results * chore: place native chat validation artifacts under docs * chore: drop stale root package config * fix(native-chat): enforce rebuilt lifecycle append slots --------- Co-authored-by: Merge Sim <sim@local> * chore: omit native-chat reland planning docs * fix(native-chat): remove journal store import cycle * fix(native-chat): keep journal factory acyclic --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
894ed75abb |
Revert "fix(native-chat): preserve large structured command results (#17707)" (#17719)
This reverts commit
|
||
|
|
5fe37729ea |
fix(native-chat): preserve large structured command results (#17707)
* fix(native-chat): preserve large structured command results * chore: place native chat validation artifacts under docs * chore: drop stale root package config * fix(native-chat): enforce rebuilt lifecycle append slots --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
fd9125ea8c |
feat(native-chat): Codex structured native chat restructure (#16729)
* feat(native-chat): port structured Codex sessions from restructure-recovery Rebuilds the desktop structured native-chat implementation from brennanb2025/native-chat-restructure-recovery (tip 4e31c08db3) on top of current main as a single commit, scoped to the local Codex path. Ported: - Structured agent-session core: durable record store + single-writer lease, canonical journal, agent-session wire host/attach/eviction/subscribers, `agentSession.*` RPC surface (registered via ALL_RPC_METHODS; host-side mobile allowlist included for wire compat), pty write gate, transcript additions, and the Codex app-server adapter/launch resolution. - Renderer: NativeChatStructuredSession view/composer stack, structured launch path with the single-flight guard, local structured session tabs sync, activation gate + structured inventory (read-only `agentSession.handoffStatus` probe), agent-session tabs in the tab strip, AI-vault structured session activation, and the settings pane with the parent Experimental Chat UI toggle plus the nested "Use updated structured native chat" toggle. New sessions require both flags, agent codex, no prompt, and a local non-WSL, non-Windows-host execution host (structured-native-chat-availability). - Fixes 72c013cea6 (verified Codex launch recovery), 8ddbaf5e3d (defer native terminal view switching affordances), and 4e31c08db3 (release the launch gate after a visibility retry) with their regression tests, including the third-launch-after-retry guard case. - Cross-version agent-session wire test + CI lane, packaging entries (proper-lockfile, agent-tooling asar excludes), and the wire-compat doc section. Deliberately not ported: mobile/ changes, the Claude structured runtime (only the claude-transcript-branch-proof and claude-structured-owner-identity leaf modules remain, backing the kept TUI-recovery arms), the terminal↔chat adoption/handoff flow (`agentSession.adoptTerminal`/`requestHandoff`, the handoff request engine, TUI adoption machinery, orca-runtime adoption methods), renderer switching affordances and their dead leftovers, the hook/subagent-status refactor cluster, and unrelated branch changes. The crash-during-acquisition recovery path (restart handoff adjudication, restore/reverse re-acquire, lease schema handoff keys) is kept because every plain direct launch depends on it; a trimmed handoff coordinator exposes only status/restore/close. Branch edits that targeted files main has since split (ipc/pty.ts, worktrees.ts, rpc/methods/terminal.ts, useIpcEvents, pty-connection, store/slices/terminals.ts, runtime-types, web preload) were re-applied to the split modules, preserving main's newer logic (Windows CIM fallback, browser tab close rework, cold-restore resume flow, dispatcher threading). Known seam: the mobile clipboard image-provenance CONSUMER gate ships (agentSession.send refuses unproven mobile image refs with agent_session_image_untrusted) but the producer hunk in rpc/methods/clipboard.ts stays with the unported mobile cluster, so mobile image sends into structured chat fail closed until that side ports. * fix(native-chat): trust only authenticated local image uploads * fix(build): preserve Windows process-tree patch application * test(windows): include process creation time in addon fixture * fix(build): run windows-process-tree node-gyp from the physical package dir gyp expands the node-addon-api dependency by probing node, whose cwd resolves to the package's physical directory in the store, so the emitted target is a store-relative ../../../../node-addon-api@... hop. gyp then resolves that hop against the rebuild cwd; from the node_modules symlink/junction it escapes the store and configure fails with "node_addon_api.gyp not found" (run 32999886072). Rebuild from realpath(package dir) so both bases agree, matching how the package manager itself runs native install scripts. The regression test replays gyp's expansion+resolution against the planned cwd and fails without the fix. * fix(native-chat): keep chat tabs visible through terminal closes and empty-worktree launches Two proven blockers in the native Codex tab contract: closeTerminalTab pre-empted the canonical unified close. With one terminal left it deactivated the worktree on a terminal/editor/browser-only check, blanking a workspace that still held a renderable agent-session tab; with two or more it pre-picked a successor from terminal entities only, re-stamping the group active before closeUnifiedTab's MRU/neighbor repair could land on the chat tab. Successor choice now defers to the unified contract whenever the terminal has a unified row, and deactivation is gated on the unified renderable count (matching leaveWorktreeIfEmpty), with the legacy pre-pick kept only for terminals without a unified row. A structured session created on an empty worktree was published into the host's headless group while preserveLocalLayout froze the local layout, leaving the tab in store but permanently off screen. A preserveLocalLayout owner now always takes client-owned placement — repairing a rendered leaf whose group record is missing, or materializing a rendered group on a truly empty worktree — and applies the client-derived layout repair while still rejecting host-authored layout. Regression tests drive the real store through closeTerminalTab (git worktree and folder workspace) and the real snapshot applier for the empty-worktree adoption states; all fail without the fixes. * fix(native-chat): close stale turns and retry rejected sends * fix(native-chat): retire hosted rows on structured tab activation * fix(native-chat): preserve rpc defaults across main merge * chore: format remote wire compatibility guide * test(native-chat): cover retry after unconfirmed send * fix(native-chat): reload outbox on session switch * docs(settings): disclose structured chat platform limits * fix(native-chat): await Codex launch-home preparation * fix(codex): align child-process allowlist with async trust bridge * test(identity): update inventory for tab surface refactor * fix(windows): preserve process-tree CRLF patch sources * fix(native-chat): anchor an unmatched chat echo where it was sent (#16117) * fix(native-chat): anchor an unmatched chat echo where it was sent The reported symptom was old user messages replaying below every new turn, so the conversation read as scrambled. The cause was not that the echo failed to match a transcript row. Claude consumes a mid-turn send through a `queued_command` attachment and writes no `type:"user"` record for it, so some echoes can never match, and no amount of matching will change that. The cause was WHERE an unmatched echo rendered: buildMobileNativeChatTransientData appended every pending item after the entire transcript, so it re-read below each turn that landed afterwards. Render each echo directly after the transcript row it was sent against, using the baseline the send already captures. An unmatched echo is then at worst a duplicate in the right position rather than a scrambled one, and it stays visible. Echoes sharing an anchor keep send order; a send with no baseline, or one whose anchor folding dropped, still falls back to the tail. Deliberately NOT fixed by deleting the echo. Inferring from send ordering that an echo can never match, then removing it, loses the user's own text for a message the agent did receive, and it cannot fire in the common case anyway - measured drain groups are 1,017 of size 1 against 55 larger. It also escalates an existing gap: the count pass has no baseline-tail guard, unlike the glue pass, while `messages` is a 40-row window that head-trims, resets on reconnect and grows at the front on loadEarlier, so a false landing there would license deleting a DIFFERENT outstanding message. That count-pass gap is real and left for a separate change; anchoring makes its worst case a duplicate in place rather than a scrambled conversation. * fix(native-chat): preserve folded echo anchors * fix(native-chat): preserve forward-folded echo anchors * fix(native-chat): keep leading folded echoes in place * fix(workspace-cleanup): show git status for every row (#16690) * fix(native-chat): refuse structured chat on every Windows execution path canUseStructuredNativeChat only refused win32 when a project runtime resolved, so folder-workspace keys (and other keys with no project runtime) failed open into structured chat on Windows. Fail closed on win32 unconditionally after the host check, matching the settings copy: local macOS/Linux only; Windows/WSL/SSH stay on terminal chat. * fix(native-chat): restore runtime refusals behind the win32 gate |