mirror of
https://github.com/stablyai/orca.git
synced 2026-10-04 08:02:09 +00:00
ff59e2cd7b0f2e73d558bbffdd888d07aa83a437
232
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ff59e2cd7b |
fix(codex): log a late reply to a timed-out request instead of showing it in the chat (#23893)
* fix(codex): log a late reply to a timed-out request instead of showing it in the chat * fix(codex): say a reply had no waiting request rather than an unknown id |
||
|
|
d60f999f94 |
fix(native-chat): turn facts come from the turn record, and /compact is a message the chat sends (#23059)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * fix(native-chat): name a chat write by its target, not the owner generation A write carried the fence of the last frame the pane read, and the host refused it unless that fence was still current. An idle release and the restart after it each move the fence, and the release publishes nothing, so a send after a release was refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a cold start was refused as stale. Every write already names what it acts on: a send its conversation, a cancel its turn, a prompt answer its item revision, a rewind its epoch; an option is last-writer-wins. So admission stops comparing the client's fence, and the rebase that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it. The writer-lease check stays, and so does the attach's compare-and-swap. Frames now stamp the fence read when each frame is sent instead of a copy each subscriber kept, which went stale on the same release. * fix(native-chat): every journal append reaches the chats that are open A journal write and its delivery to open readers were two calls, and some writers made only the first. A failed start whose lease could not be handed back, a provider revision with no frame behind it, and eviction's settlement were all journaled without reaching an open chat. A journal handle now reports every durable change, and the host's session map binds that report to the session's readers when the handle is set. Writers no longer publish what they append; the per-writer publish calls are deleted. * test(native-chat): an epoch replacement reaches the open chat * test(native-chat): each row reaches an open chat once, and a live handle enters only through the map * test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite The seeded record had no surface tab id, so the next open backfilled one and that rewrite alone made the no-op transaction write. The test passed with the legacy-lease rewrite signal removed. * test(worktree-activation): restore the OMP surfaced-agent resume test The handoff removal deleted it alongside the terminal-owner tests, but it covers the surfaced-PTY block that still guards resume, including an agent whose ownership is unknown. * perf(native-chat): a publish behind a delivered commit reads nothing Each commit now delivers itself, so the publish a provider frame still sends afterwards found every reader caught up but still read rows and rebuilt the timeline for each one. A caught-up reader now skips the read. * test(native-chat): state why the teardown test's fake journal is safe to cast * docs(native-chat): say mutation admission checks only the writer lease * docs(native-chat): drop the send rebase from comments that still described it * fix(native-chat): a message is accepted, then delivered A send to a chat with no running agent restarted the agent inside the send call, before the message was recorded, so the client waited for the whole start and a failed restart refused the message. Claude held prompts sent during startup, and those could settle as "unconfirmed". A send is now accepted inside the session's serialized queue: one ledger row and one submission row marked handoverRecorded, published, answered pending. A per-session delivery loop exists while a message is queued. It starts the agent through the same serialized attach a hold uses, waits outside the queue for a Claude child to prove its start, and hands the oldest queued message over as its own serialized step, writing dispatch{pending} before the adapter call. A start it needed and did not get writes one error-tone row and rejects every queued message with the same words; a start Stop cancelled writes none. Settlement follows from the rows. A queued message is provably unwritten, so a close, an eviction or an exit rejects it. A handed-over message stays in doubt. A queued row at or below the sequence a handle found when it opened was left by an earlier process and is rejected at open, with no latch. Stop withdraws queued messages with no writer lease and no fence. An attach failure keeps the conversation open, and the attach adopts its journal. Owed work counts the loop and queued rows. A compaction or rewind found prepared when a conversation opens was started under a child this process no longer has, so the open settles it rather than leaving it to refuse every send until a view attaches. The open cursor is scoped to its epoch, because sequences restart when an epoch is replaced. Deleted: restart-before-admission, recordFailedRestart, the fence rebase, Claude's startup gate, the attach's forget on failure and its own crash boundary. Clients without agent-session.accepted-send.v1 get their reply held until the handover; the desktop and paired desktop lists advertise it. * fix(native-chat): settle queued messages only for the child that ended A child that proved its start and then exited before its message was handed over left the message queued: the exit settlement returned early when nothing else was in flight. Delivery then started another child for it, and a child that died the same way started another, without end and without a row. A retried settlement for an earlier generation, run by the attach that delivery started, did the opposite: with that generation's turn unfinished it rejected the message queued for the child being attached. The settlement now takes the rejection for queued messages from its caller. The unexpected exit and the eviction pass one, and it applies even with no other work in flight; the retry for an earlier generation passes none. * fix(native-chat): an adoption that fails to import keeps the conversation open The attach now writes into the conversation's own open journal, but a failed transcript import still closed it as if it were the attach's provisional one. The conversation stayed indexed with a closed journal, so every later send answered "could not be recorded" and every attach failed again until the app restarted. The import now closes only a journal the attach opened for itself. * perf(native-chat): the recovering open reads the journal once Every conversation open now goes through the recovering open, including the read restore of every chat at startup, which used to replay its journal once. The recovering open replayed it twice: once to probe it and again inside the open. The probe is now handed to the open as its load. * fix(native-chat): an attach that fails after indexing its child leaves no child behind A failed attach now keeps the conversation open, but a failure after `onAttached` indexed the child (the rewind or compaction recovery, or the attach's own success record) left that entry claiming a child the failure path had already released. The next send found the phantom, skipped the start, and wrote at a fence the journal had moved past, so the message stayed queued for good. The entry now drops the released child and its event sink, and follows the record's fence, as a failure before indexing already did. * fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer The error strip for a message the host accepted and then did not deliver matched the entry before the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not send your message" with nothing to retry. It now reads the reconciled entry. A rejection the journal records before the send's own pending answer lands is final as well: that answer no longer puts the entry back to dispatching with no Retry. * fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down The preamble waits for its submission to be delivered while the worker's agent starts. When that wait ran out it threw operation_unknown, and the failed-start teardown then closed the session, which rejected the very preamble the host was about to deliver. It now reports a turn start nobody observed yet: the worker is start-unknown with its session kept, the host delivers the preamble when the agent starts, and the worker's report settles the dispatch as for any unobserved start. The receipt no longer suggests reading a screen a structured worker lacks. * fix(native-chat): a message rejected while its chat was closed reads as not sent A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked every later message behind a Retry and no reason, and the delivery probe, seeing the journal already answered, never ran. The reconcile now settles it as rejected like a dispatching one. * test(orchestration): name why the readiness settlement fakes are cast * fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent * docs(native-chat): drop the fence from the admission the send effects run behind * docs(native-chat): give the fence move on release the reason that still holds * docs(native-chat): stop citing a write fence check in launch and mailbox comments Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease. * refactor(native-chat): the provider child is its own record A conversation now outlives any number of provider children, so the child is one record on the conversation's entry instead of five loose fields beside its journal. It is written in one place: indexed only once an attach has fully succeeded, and ended through one function that an exit, a failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence. - A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence patch after it are gone. - Conversation writes read the record's fence, the way mutation admission already does; a child's own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of the conversation's fence, are gone. - The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer dropped when an attach replaced the whole entry. - Stop on a child still proving its start stops only the child: its lease goes back and the chat is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus the conversation's close. - The settlement retry uses the conversation's own journal, opened through the host's one open. * fix(native-chat): the delivery loop alone settles a message its start or child failed A queued message was settled by whichever path happened to end the child first: the loop, the unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that rejected every pending row. That gave two failure rows with different tones for one start, a loop that could hand over to a different child than the one it waited on, and a Claude start that died while starting reading unlike every other failed start. - The loop remembers the child it waited on. At handover, if that child is gone or replaced, it reads how it ended: a Stop continues; anything else writes one failure row and rejects every queued message with the same words, then stops. A child still starting whose start the adapter says did not land fails the same way. The exit, eviction and the settlement retry only settle the handed-over and legacy rows of the child that ended. - One failure row, always an error, keyed by the start. A start a view began that dies with nothing queued writes the same row through the same builder, so a second report revises it. - The open no longer rejects leftovers; the loop's first step does, and the open wakes it. - `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the failure before the exit is processed. - Quit closes every conversation the way closing a chat does: what is still queued is rejected as closed, with or without a child, and a start the loop already has in flight is waited for so the child it produces is stopped rather than left behind. * refactor(native-chat): a stopped child ends on the one reading of its stop The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that verdict to the child's ending, so the host never forms a second view of whether the root is gone. Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition, and a failed re-attach passes what its release saw. The end-of-child record can therefore also carry a stop whose root was not seen to go, which nothing ends on yet. * feat(native-chat): the host says it accepts a send before any agent has it The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same string capable clients already send. A client can then tell a host that answers a send at acceptance, and admits a Stop with no writer before a turn starts, from an older one that still restarts the agent inside the send. Additive: an older client ignores a capability it does not know. * refactor(native-chat): an attach never opens a journal of its own The attach adopts the conversation's open journal, which outlives it, so it no longer opens one for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag that told the two cases apart is gone. Tests that attach without a host open the conversation the way a host does. * fix(native-chat): a moved fence resends nothing on a host that accepts first The outbox treated any fence change as a new owner: it dropped the answer of a send in flight, queued that send to go out again under the same id, and unblocked a refused head. On an older host that is how a send the restart refused, unrecorded, gets another try. On a host that records every send before it starts an agent, a fence moves because that start ran, so the same rule resent into every failed start. With a fence stamped on every frame, that became a loop. The outbox now reacts to a fence change only when the host has not advertised that it accepts a send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed start reaches the client as a rejected message it keeps with its Retry. Against an older host, or before one has answered, the outbox behaves as it did. Desktop and paired web share this hook. * refactor(native-chat): a child's end says whether the user or the host stopped it The end-of-child record's cause now tells a user's Stop from the host stopping the child for a cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a user's Stop, as before, and fails the start it was waiting on after a host stop, with the one error row and every queued message rejected, in the stop's reason when it gave one. The reason stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet. * fix(native-chat): a chat whose only work is a queued message is not offered for resume A message accepted while the agent was starting counts as working in the chat, and quit rejects it as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a chat whose agent never had the message. The snapshot now reads only what was handed over. * fix(native-chat): the conversation outlives its agent Opening a chat no longer starts its agent. A conversation is reached through one host accessor that opens its journal at rest, and a send is what starts the agent, through the delivery loop. One idle sweep, every five minutes, stops an agent that has been quiet for thirty minutes and owes no work, then drops an open journal handle that is only a cache. Its record, tab, status row and readers stay. - hold and release are no-ops; hold still builds the host for shipped mobile builds. - The holders, the holds, the release clock and the exit respawn are deleted. - Options, the model list, the goal and the context meter answer at rest; a model pick at rest is recorded as intent for the next start. - Compact, rewind, clear and goal changes start the agent first. A send does too when a rewind is still in doubt after the conversation opens. - Orchestration routes mail and group addresses on ownership (the record plus the chat tab), not on whether the process runs. An open dispatch keeps its worker running. - The restart continuation is a send; Resume all holds each slot until the message is handed over or rejected. - A read error never replaces a loaded transcript, and shows the host's own words. * test(native-chat): type the queued-message fixtures in the resume-offer tests * fix(native-chat): a start that dies while a message waits on it is that message's failed start Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When that start died, its exit wrote the start's error row and left the message queued, so the delivery loop started a second agent into the same failure and wrote a second row. A child's end now records where the conversation's journal stood, and the loop settles a message accepted before a failed start ended with that start: one row, under its key, and no second start. A message sent after the failure still gets a fresh start. * fix(native-chat): a request that failed reads as failed A structured chat whose only message the agent's start refused read as a green finish, and a cancelled structured turn did too: the host published a verdict only for turn records, and structured rows carried no `interrupted`. The host projection now reads the session's latest request: its turn's outcome, or `failure` for a send the agent or its start refused. A send that was withdrawn, or left undelivered by a restart or a close, fails nobody and makes nothing listable. The ingest publishes `interrupted` as the hook lanes do, and every reader decodes the verdict through one accessor, so a failure reads Failed on the dot, the rollups, history and `worktree ps`, behaves like a cancellation in every clean-finish policy, and notifies as "failed". * docs(native-chat): say what an attach's open conversation and unconfirmed ids are now * test(native-chat): a verdict change republishes the mobile status projection * refactor(native-chat): the store's retention trigger keeps its flag compare A verdict change always moves the completion clock the same check already reads, so a second verdict compare there caught nothing new. * test(native-chat): a user message the provider journaled keeps its session listed * test(native-chat): pin what a failed start settles, and what a resume offer names A view's child that dies while a sent message waits settles that message only when it died starting and no child has taken its place: a proven child's crash, or a second start since, gets the message delivered. The resume offer names the handed-over message, never a newer one still queued. * test(native-chat): the failed-start pins fail on what the message became, not on a timeout * fix(native-chat): a restart offer ends when the chat's agent starts again The offer used to end only when the chat's newest user message changed, because opening a chat started its agent and that start could not be told apart from real activity. Opening a chat starts nothing now, so the host reads the fact it already publishes: a chat's status row goes from not host-owned to host-owned exactly when its agent is started. At that edge the offer and any failure record for the chat are withdrawn, unless the start is a resume action's own (its continuation is the oldest undelivered message). A continuation and a message racing to be first are decided at acceptance: the continuation is refused, quietly and with nothing filed, when any other message was accepted since the restart. A failed continuation start leaves the offer retryable, and each resume action sends its own message id. Deleted: the newest-user-message comparison, its journal reader, the continuation filter, and the failure ledger's own "answered by the chat" check. The marker still carries its message id for one release, so the previous build can read it. * fix(runtime): end a transcript stream when its client unsubscribes Desktop: the IPC subscription controller was dropped as soon as the streaming handler returned, which for most streams is right after it binds. A later runtime:unsubscribe then found nothing to abort, so the host kept the subscriber and derived and sent every publish to a channel no one listened to. The controller now lives until the renderer unsubscribes, resubscribes the same id, or goes away. Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe with the stream's frame id, so the host ends that subscriber and leaves a sibling stream on the same socket running. The direct path now passes the frame id the relay path already passed. * fix(native-chat): a late provider-session update keeps a failed recovery record failed A provider-session heartbeat that rewrites a completed recovery record kept its interrupted flag but dropped the outcome it was copied with, so a live failed checkpoint read as a clean finish until the next status write. * test(orchestration): the preamble's host stub is typed, not cast The preamble send now takes only what it reads of the host, the send, the settlement wait and the record's fence, so its test builds that host with real types instead of `as never`. * test(native-chat): the terminal-bell check asserts the renamed verdict field The bell notification test still checked for agentInterrupted, which no longer exists, so it could not catch a verdict leaking into a bell dispatch. * fix(native-chat): a failed turn ranks like a completion for attention Attention readers (completion time, Smart Sort, sticky retention, Cmd+J Recent) now demote only a turn the user stopped. A failure is news the user has not seen, so it keeps its completion time, ranks in the Done class, stays retained after its pane goes away, and a retained failure reads failed in the worktree rollup instead of done. Clean-finish policy (hibernation, pane ownership, the value moment) still treats a failure like a stop. The retention trigger compares verdicts again: success -> failure no longer moves the completion clock. * fix(native-chat): one fact ends a restart offer: the chat moved on since the restart The offer is live while no other message has been accepted in the chat since the restart and its agent has not proved a start since. The offer list, the resume's reservation check and the continuation's acceptance check all read that one fact, so a message whose start then failed withdraws the offer too, and a stale click finds nothing to act on. The fact is read off the conversation's open handle, which the restart closed, so it is retired durably whenever it may have changed: a message accepted, a start proven. A close and reopen within the same run therefore cannot bring the offer back. A continuation rejected before it reached the agent does not count, so a retry after a failed start still runs. Deleted: the quit-time gate on withdrawal, which changed nothing because the withdrawal and the quit's own offer write share one queue; the per-action "withdrawn" flag and the separate acceptance check it paired with. * test(native-chat): an older build reads the restart offer this build records The offer lives in a file the previous release reads after a downgrade. Pin that against the pinned release's own capsule, and run the lane when the marker or the capsule changes. * fix(native-chat): read a restart offer against where the journal stood when it was taken "Since the restart" was read off the conversation's open handle, which the idle sweep closes: after a reopen, a message the user had already sent looked older than the handle and the withdrawn offer came back. The offer now records the journal position (epoch and sequence) at the moment it is taken, and a message accepted after that position, or a journal on another epoch, means the chat moved on. That is derived from the journal, so it holds across any number of closes and reopens. An older build's offer has no position; only a start withdraws it. Because the message half is now durable, the offer is no longer rewritten in the recovery file on every accepted message; a proven start still writes it, since only the host that saw the start knows of it. * test(native-chat): wait for the listing's retire write before reading the recovery file * refactor(native-chat): every journal row states which turn it belongs to Rows gain a turn scope stated by the write that creates them: the open root turn, or the conversation. A queued message takes its scope from its handover. Rows stored before scopes existed are placed on replay by the root turn open when they were created, so no persisted state is needed for them. Rewind keeps each retained row's scope and producer, so a subagent's row stays its own. * fix(native-chat): keep the terminal-backed chat's read error over its local echoes Messages winning over a read error is right for the structured chat, whose read retries and whose messages came from the transcript. The terminal-backed view assembles its list from local echoes too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no error. Only the structured pane now keeps messages over an error. * fix(native-chat): a start retries the exit settlement a failed journal write left owed An agent exit whose journal settlement write failed releases the lease latched until a retry lands. Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried it before the next app launch, and every send was refused. The start the send needs now runs the retry first, where the attach would. * fix(native-chat): a failed main agent reads failed while its subagents still work The verdict is now read from the main agent's own state, not the folded row: a main agent that is done and failed has a verdict even while its subagents keep the row working. Without mainAgent (history, worktree ps, older hosts) the old combined-done rule stands. Display marks the verdict through agentVerdictDisplayMark: a failure outranks every combined state on the agent's dot, label, tab badge, dashboard and activity rows; a stop marks only a done row, so a successful or stopped main agent with live subagents still reads working. Subagent rows keep their own state. The worktree card, terminal tab and Cmd+J rollups share one pane fold and rank a pending question, then failed, then working, monitoring, interrupted and done. worktree ps publishes the main agent's outcome on a working row, and the mobile mirror reads it. The store's change check, the paired-client mirror's equality and its epoch now see a verdict change on a working row, which otherwise moves no state or clock and left the worktree card reading working. Clean-finish policy is unchanged: a working row is never hibernated and has no completion time. * perf(native-chat): answer the owner check without opening the chat Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the answer comes from the session record alone. Reaching it through the accessor opened each resting chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it open for the idle window. It now checks the record and the adapter's support, as before this series, and opens nothing. * fix(native-chat): a read waiting on the session lock opens nothing once quit began The accessor checked for quit before queueing the open, so a read queued behind a session task ran its open after teardown had begun and indexed a journal no teardown step would close. The check now runs at the open itself. * test(native-chat): pin stated turn scopes, the upcast of unscoped rows, and rewind attribution * fix(native-chat): /compact is a message the chat sends, run as a turn of its own The conversation command RPC now accepts /compact into the queue like any send and answers once it is handed over. The delivery loop opens the command's own turn, starts the provider on it, and waits for the provider's end off the session's queue, so messages typed meanwhile are held and delivered after it, even when it fails. It settles by re-reading the journal: a child that died meanwhile already wrote the verdict. Stop ends the command at once. The 180 s completion window, the unconfirmed row and the recovery of an older build's compaction record are gone; that record no longer gates anything. On Codex the provider turn the command opens is claimed into the command's turn. * fix(native-chat): read a failed resume's chat before calling it retryable Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the restart, read from its journal. The failure list read it only for a chat already open, so once the idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did nothing. The list now opens the failed chats first, as the offer list does. * test(native-chat): type the provider event sink the settlement test reaches for * docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it The field is new: an old host sends no outcome at all, so a reader falls back to interrupted. The removed clause said old hosts send it on done rows, which never shipped. * fix(native-chat): say the structured read keeps trying only where it does The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an untranslated fallback whenever the read error had no text, and the empty state prefers any message. The view state now leaves the message out, so the structured pane shows that line and the terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an error frame, so it no longer makes the claim. * fix(native-chat): rows group under the turn their record names, not the one above them Each row's turn is the turn its stated scope names, anchored on the entry that opened it, or on the turn itself when the provider opened it unasked. So /compact groups its own rows and the previous turn is untouched, a message typed into a running turn joins it, and a provider-resumed turn folds under its own Worked-for. A row reporting how a turn ended, an error or the compaction separator, never folds. Desktop and mobile read the same keys; a host that states no scope keeps today's positional grouping. * test(native-chat): await the send's settlement instead of polling for the start The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a loaded machine outran. They now await the host's own settlement of the message. * docs(native-chat): the status-store listing rule names provider-journaled user messages * fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now dated by the resume action. Telling a rejected continuation from the user's own message read the operation ledger, whose rows expire after about a day; after that a failed resume stopped being retryable. The offer now records the continuation each action sends on its own capsule entry, bounded to the newest 16, so the ids end with the offer. The ledger read is deleted. * fix(native-chat): a /compact is not a request the sidebar, notifications or restart resume report The sidebar's prompt, preview, verdict and instant, the turn-completion feed, and the restart-resume marker read past a conversation command and its turn to the last real request, so a /compact neither notifies nor re-dates the row, and a command in flight is never offered as work to resume. An older client shown a command's turn in the legacy form names the session's own agent. * fix(orchestration): route no mail to a structured worker its orchestration released A structured worker is routed on ownership, and a resting worker's lease is released, so ownership held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one restarted its agent. Routing now also reads the orchestration's own resource row: once it is released, direct mail, group addressing and worker-show's addressable answer drop the worker, as they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored. * fix(native-chat): a failed retry names the user's prompt, not Orca's continuation A resume's continuation is written to the chat before its start, so after a failed attempt the chat's newest user message is that rejected continuation. A second failure then showed Orca's own restart text as the chat's prompt. A retry now keeps the prompt its first failure named. * test(native-chat): pin what a conversation command's admission refuses at rest and at handover * test(native-chat): tests merged from the base state which turn their rows belong to * fix(native-chat): a refused send notifies failed through the completion feed The host's completion feed followed only the newest turn, so a send the agent or its start refused, which creates no turn, read Failed on its row but sent no notification. The feed now follows the session's latest request, read from the projection the status feed already makes for the commit: a turn keeps its id, a refused send is named by its journal item key. It announces only while the session is idle, as the row reports a verdict, so queued sends refused one commit at a time notify once, and a withdrawn send falls back to a request already announced. * fix(orchestration): read the released row optionally, as the authority does worker-show's observation called the row lookup directly, which a runtime double without it threw on and failed the structured tab-retirement release. * chore(native-chat): one import per module and no unexplained casts in the turn-scope changes * test(claude): pin which turn a Claude row joins, including a subagent's after the turn ends * fix(native-chat): the status bar drops a restart offer the chat moved on from The renderer re-read the host's restart offer only when a failed chat showed activity, so after a message withdrew a pending offer the host answered no chats while the status bar kept counting one, and clicking it opened nothing. The same watch now covers pending offers: a status change in an offered chat asks the host again, once. * fix(native-chat): a refused steer is read from the turn its handover named The latest-request reader decided whether a refused send had joined a running turn by comparing host clocks: its handover time against the previous turn's end. The handover row now states the turn it delivered into, so the reader reads that instead and the clock comparison goes. A journal written before handover rows stated a turn is scoped on replay from the turn open when each row was written, which can differ from the clock reading only when a send and a turn's end share a millisecond. * fix(mobile): the native-chat controller contract carries the turn journal The controller and overlay already pass nativeChatTurnJournal, but the contract type never declared it, so mobile failed to typecheck. * fix(native-chat): the live turn is the running turn, not the newest user row A turn the provider opened on its own (a background wake, a resumed turn) anchors on its own record, but the list still treated the newest user row as the live turn. While such a turn ran, the settled user turn before it lost its duration and the running turn's own rows were drawn as settled, so its tool calls lost their live state. nativeChatTurnMembership now answers both questions from the turn record: each row's turn, and the live turn (the running root turn's anchor, else the newest user row, which is also all an unscoped host has). Desktop and mobile key liveness, the timing clock and the live status's row on it. * test(native-chat): a turn the provider opened keeps its own clock Pins that the local turn clock follows the live turn, so a wake after a settled turn does not restart that turn's clock when no host durations are recorded. * fix(native-chat): a running turn no message opened draws its status on no row Its live status belongs to the transcript-tail indicator alone. Once it settles, its duration draws at its first row as before; a running turn a message opened still draws on that message. * fix(native-chat): every copy of a row carries the main agent's own status History entries, sleep records and `worktree ps` rows carried a flattened top-level `outcome`, copied under different gates and without the main agent's clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type the live row already persists and sends, and every copy site takes it with `interrupted` through one function, `agentVerdictFields`. - The accessor reads `mainAgent` then the legacy flag; the mobile mirror matches it line for line. - Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a malformed value drops the field, never the record. - Mobile dates a main agent that failed under live subagents by its own clock, as desktop does, and its row equality compares `mainAgent`. - The activity feed reads a history entry's own `mainAgent` instead of rebuilding one; the sync key and history equality compare it. * test(native-chat): pin the worktree ps verdict across host and phone versions Pairs the real v1.4.212 host and phone row reader with this build: an old phone reads a new host's rows by `interrupted`, a new phone reads an old host's rows (no `mainAgent`) the same way, and a new phone reads a failure under live subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout now carries the phone's self-contained row reader, and the lane runs when the `worktree ps` row producers change. * test(mobile): name the parity table's row for its role * test(native-chat): a roster of idle or finished children does not keep an agent awake The sweep reads owed background work through the shared child-work liveness that upstream's release clock adopted; a child that went idle or finished is not work the agent still owes. * fix(native-chat): a request that settles while the user is asked something notifies once The completion edge waited for an idle session, and a pending prompt (including a subagent's approval) is not idle. Structured chat has no other attention producer, so a main turn that finished while a subagent waited on the user sent nothing until the prompt was answered. The edge now waits only on owed work (a running turn or an unanswered send), which the projection reports even beneath a pending prompt. A request that settles with a prompt pending announces once; the renderer words it "needs input" from the host status mirror's `attention`, and answering the prompt keeps the same request identity, so it does not announce again. The wire shape is unchanged. * fix(orchestration): a task dispatched into a resting structured worker keeps it running The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's process incarnation now counts, derived from the existing rows. * fix(native-chat): a command's wait ends when its child does The delivery loop waited for a /compact only on the adapter's compaction tracker, which learns of the child's end only on some exit paths: a Codex exit or close, and a Claude close, never reach it. The wait then never ended, so nothing queued behind the command was delivered again, Stop had no child to answer through, and the tracker's leftover entry refused the next /compact. Every way a child ends passes endProviderChild, so the host now offers a per-child end signal there. The loop races the tracker against it (the dead-generation settlement has already written the command's verdict), and on that end asks every adapter to release the command, so a later command runs and no later provider turn is claimed into the dead one. The adapters' own exit-time releases were unreachable (Codex) or covered one path of several (Claude), and are removed. The Codex RPC test harness moves to its own module so the exit can be driven through the real adapter's connection callback. * fix(native-chat): keep refusing sends during a command on an older host An older host's controller still refuses a send while a conversation command runs, so dropping the client's block turned every message typed during /compact into a 'not sent' row with Retry there. The block stays for hosts that do not run the command as a send-path turn, and goes only for those that do. The signal is one the client already holds: a host that runs /compact on the send path states a turn scope on every journal row it writes, the same fact turn membership uses to tell it from an older host. Both now read it from one predicate. On an empty conversation, or one whose rows all predate the upgrade, the signal is absent until the command's own entry streams in, so that brief window keeps the old local refusal; no capability or wire field is added. * docs(native-chat): comments stop describing the hold this PR removed Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it. Comment-only. * fix(native-chat): the completion says when the user is being asked A request that settles while a prompt waits on the user was worded "needs input" from the renderer's status-feed mirror. Remote clients receive the status and completion streams over separate sockets, so they can arrive in either order and the wording could be wrong both ways. The host already knows at emit time, so the completion now carries an optional `awaitingUser: true` in that case and omits it otherwise. The renderer words the notification from that field alone and no longer reads the status mirror. Old clients ignore the field and word by outcome; old hosts never send it. * fix(native-chat): a restart offer keeps the start its own continuation made Whose start ended an offer was decided at read time, from whether the offer's continuation was still the queued message. Once the provider refused that continuation, the child it had started read as someone else's start, so the offer ended and its failure showed no Retry. The delivery loop now records which queued message a start is for on the in-memory child, and the child's end carries it; the offer counts a start as its own when that message is one of its continuations. * fix(native-chat): a rewound turn still names the message that opened it A Codex rewind rebuilds the epoch without submissions, so each sent message survives only under its provider key. The kept turn records still named the submission key, so each turn anchored on itself and its rows grouped apart from the message that opened it. The rewind now renames the turn's opener along with the message. * fix(native-chat): Stop ends only the command it names Stop on a command turn abandoned whatever compaction the session had pending, so a late Stop for an earlier /compact cancelled the one running now. The tracker now ends a command only when the Stop names its turn, and the cancel reply reports whether it did. * fix(native-chat): an agent gets a full idle window after its owed work ends The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can read done before the lead's wake-up turn writes anything, and stopping in that gap loses the wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full window afterwards, as the release clock it replaced did. * test(claude): the options-read fixture runs a live child The fixture marked its conversation running with a hasProviderChild field the session type does not have, so the read took the at-rest path and refused a session with no record. It now carries a child, which is what the read checks. * test(native-chat): host tests reach its collaborators through a typed seam The rest-test rig and three test files read the host's private members with Reflect.get and cast the result. The host now exposes one test-only accessor, collaboratorsForTests(), and the subscribers class a subscriberCountForTests() beside its existing retainedActivityCountForTests(), so the tests are checked against the real types and the casts are gone. * fix(worktree-status): a departed agent's failure yields to live work on the worktree card A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome. * refactor(orchestration): one owner answers a structured worker's custody Routing, group addressing, worker-show and the idle sweep each composed their own reading of whether orchestration still holds a structured worker, so each new obligation or retirement state had to be added to every reader. structured-worker-custody now derives both answers from the worker-terminal list state coordinators see in worker-list: addressable is owned and not released, and owed work is an active custody or an unsettled task dispatched to the same incarnation. The owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests. * refactor(orchestration): owed work is an open dispatch on the worker's incarnation A supervised worker's own dispatch context stays open exactly while the worker is active, so the separate active-custody branch only repeated it. Owed work is now one fact, which also states the policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are written once at the top of the module. * docs(agent-status): a departed agent's failure ranks below live work on the worktree card * fix(native-chat): a restart offer knows its continuations by a tag in their id The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running action's id in memory. Both could disagree with the journal: past the cap an old rejected continuation read as the chat moving on, and a crash during a retry restored the failure's older entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer (its teardown and chat), then the action's own part, so any continuation of this offer, queued or rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own continuation the start was for, read against the stored marker. * test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait The tui-idle probe reads through readTerminal, which now awaits the structured worker check before the PTY read, so the probe's snapshot request starts a microtask later. vi.waitFor missed it on its first check and polled again at 50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot then resolved after the wait had already timed out, so the test passed without judging it, and the rejection landed before any handler was attached. Vitest reported that as an unhandled error and failed the shard. Polling every 1 ms sees the request within a few ms, so the snapshot is judged while the wait is still pending. * fix(native-chat): a message held behind /compact is drawn where it was handed over A message typed while /compact runs was drawn above the compaction's result, between itself and its own answer. The reducer kept every item at the sequence and timestamp of the row that created it, and a queued message is created at acceptance, long before the command it waits behind writes its result. The phone orders by that sequence and the desktop by that timestamp, so both put the message first. A queued message now takes its position from its handover row, the same row that already states its turn scope. Everything the agent did before the handover, a command it waited behind included, draws above it. This holds for every held message, not only /compact's, and needs no client change: every client, older builds included, reads the position the host publishes. A live batch already carries the item when its dispatch row lands, and history pages cut the reduced timeline by sequence, so paging stays contiguous. * fix(native-chat): a phone's send during /compact answers without waiting out the compaction A client that predates accepted-send replies, which is every phone build, has its send reply held until the host hands the message over. A message sent during /compact is not handed over until the compaction ends, so the phone's 15 s request timeout fired first and showed the message as unconfirmed. That wait now also ends once the message is queued behind a running command. This is read from the journal's running turn and needs no new state. Every other wait still ends at the handover: behind a starting child or an ordinary turn, and for restart resume, the command front door and orchestration, which keep the plain handover point. * perf(native-chat): a rewind places provider items with one pass over the merged rows A Codex rewind gives each provider item the old epoch never held the turn record for its provider turn. It found that record by scanning every merged row, restoring each row's body, once per provider item. That is quadratic, and it runs on the host's main thread up to the journal's 10,000-row cap, twice per rewind. A rewind record written before rows carried their scope holds no scope for any provider item, so it paid the full cost. The merge now indexes turn records by provider turn id once, keeping the first match as the scan did, and each provider item looks its record up. * fix(native-chat): a view never restarts a chat whose last start failed A Claude chat whose CLI exits during startup left one red row per start, and every time a view bound to it (the chat opening right after its create died, or the user switching back to it) the hold started the CLI again, so the same launch-failure row repeated. Only a send retries a failed start now, the same rule provider-exit recovery already applied; the rule lives in one predicate the hold, exit recovery and the delivery loop share. * fix(native-chat): a message waiting behind /compact is drawn after it until it is sent A message sent while /compact runs is placed where it was handed over. It was still drawn where it was accepted until then. /compact writes its result one step before the handover, so for that step the waiting message sat above the compaction's separator. A message the host accepted but has not handed over is not part of the conversation yet, so both clients now draw it after everything the agent has done. The shared projection moves it to the end, which is the order the phone draws. The desktop ranks it with the other not-yet-sent rows, after the streaming preview. At handover it takes its place from its handover row, which is also after the separator, so it never appears above the compaction it waited for. * fix(native-chat): the idle sweep reads owed work every tick Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it refreshed the clock at most once a window. Work that ended just before the next read left the agent to be stopped at that read, moments after the work ended, which is the gap the refresh was meant to cover. The sweep now reads owed work on every tick for a started agent, so the window always runs from the last tick that saw work owed. * fix(native-chat): a continuation handed to the agent stays sent The offer read its own continuation as not reaching the agent while its dispatch was pending, which also covered one already handed over and still unanswered. When the wait for that answer ended first, the failure it filed read as retryable, and a retry sent a second continuation to an agent that may have acted on the first. Only a continuation still queued, or rejected, is now read as unsent. * test(native-chat): start the child the loop waits on with an attach, not a second view A view no longer starts a child whose last start failed, so the R2 case that waits on a child started since the failure now gets that child from a client attach, the one non-send starter left. * fix(native-chat): settle a gone generation's turn wherever a conversation opens A send that opens a chat this process had not read yet (after a crash, from a phone or the CLI) went through the delivery open, which never settled what the dead generation left running; only the read restore and a successful acquire did. When the send's start then failed, the turn stayed running for every reader. The settlement now runs in the one journal open, at the crash boundary, for every opener except an acquisition, which settles from the evidence it read before its reserve; the read restore's separate step is gone. * test(native-chat): prove the next child's start settles the turn an earlier child left The R1 case lost its only settlement assertion when the latch it checked was deleted. It now seeds the running turn the earlier child left and asserts it ends at the exit's receipt, with the exit's row, before the message is handed to the new child. * test(native-chat): count a failed start's rows by row, not by text Comparing the set of texts passed when two different rows carried the same words, which is the duplicate the test exists to catch. * test(cross-version): load the phone row readers without mobile's toolchain Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json extends expo/tsconfig.base.json, which the root-only cross-version lane never installs. The worktree ps verdict suite imported the current phone row reader from mobile/ directly, so CI failed with TSConfckParseError before any test ran. The harness now imports a copy of the working-tree reader placed under the checkout cache, where the root tsconfig applies, as it already does for the release checkout's copy. Both readers are still the real files. * test(cross-version): keep the checkout path-guard message and justify the copy import's cast * fix(native-chat): a command ends only by its own provider answer or its child's end Stop no longer settles a conversation command. It interrupts it like any turn, and when the provider cannot take that (Codex has not opened the command's turn yet, or Claude refuses the interrupt) it stops the child, whose dead-generation settlement writes the verdict. The pending command now lives on the provider child's own session instead of an adapter-wide map keyed by session, so it dies with the child and nothing has to release it. Claude's /compact is sent under a uuid the slot records, and only a root result naming that input (or naming none) ends it; its outcome is read with the ordinary result reading, so a stopped /compact is a cancellation. * fix(native-chat): a command's settle answers its message before ending its turn The two writes are not one batch. Writing the message's answer first means a crash between them leaves a running command turn, which the stale-turn sweep already settles, instead of an ended turn whose message reads as in flight forever. The settle now writes only while the command turn is still running. * fix(native-chat): "Worked for" counts from the handover, not the send A message held behind /compact, or behind a cold start, used to count the wait as the agent's work, although its row is drawn at the handover. Every handed-over submission's turn, the command's own included, now starts at the handover row's instant, falling back to the send time for a host that recorded none. * test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget * test(native-chat): the interrupted create's own retry continues again The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its own operation id, with a fresh start whose result nothing read. That fresh start passes with the released-reservation continuation deleted, so the case the fix exists for went untested. The retry and its assertion are main's again. * docs(native-chat): three comments that still had views starting agents A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction left alone would refuse every send, so no agent would ever start to finish it; and a current host raises the unattached read refusal only once quit began, with the attach window belonging to an older host. * test(native-chat): pin the open's and the send's start and row counts, however the view binds Opening a fresh chat whose starts fail makes one start and one row, with two views bound before or after the create's child died; one send makes one more of each. * fix(native-chat): a second Stop on a command ends its child; one compaction verdict for every provider A Stop's note now names itself in its key, so a later Stop on a command still running reads, from the journal, that the provider was already asked and never answered, and stops the child instead of interrupting again. Nothing is held in memory for it. Adds the rule both translators will read a compaction's end by: only a compaction the provider reported is a success; none after Orca's interrupt is a cancellation; anything else is a failure. A real Claude capture, pinned as a fixture, is why: a stopped /compact ends in the same success result as a finished one. * test(native-chat): a reader's open settles the turn a failed exit settlement left running An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle sweep closed, and a read that opens the chat before the restart restore reaches it. * test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner A subscription reads the conversation before it returns, so under load the two views took longer than the create child's 300 ms start, which then exited before the test checked that it had not. The child now takes a second to fail. * fix(native-chat): settle a gone generation's turn at every open but an acquisition's The journal open skipped the settlement whenever the lease read reserved or live, to leave an acquisition's own open to the acquisition. But a lease a crashed process left in recovery also reads live, until the next acquire resolves it. A send that opened such a chat, from a phone or the CLI after a crash on a host that could not prove the old owner gone, skipped the settlement; when its start then failed, the dead turn stayed running for every reader. The acquisition now says it is the opener, and every other open settles, whatever the lease still claims. * test(native-chat): hold the create's start open until the views bind The "view binds while the create is still starting" case gave the create a 300 ms head start and asserted the views bound before it died. On a loaded runner the holds took longer, the create's exit landed first, and the case failed its own precondition. The create's initialize now waits on a gate the test releases once the views are bound. * refactor(native-chat): the provider's translator ends a command's turn; the loop holds no command state A conversation command is now a turn of the provider child's own journal pipeline. The adapter-wide tracker, its promise and the loop's settle step are gone. - Codex: the translator claims the provider turn that carries the command, scopes its rows to the command's turn, and writes the command's end in the same batch that settles that turn. Codex's own compaction marker is the success row. - Claude: the command's turn is the translator's open turn until the result that answers the /compact input ends it. The command's own frames, such as the continuation summary, its echo and "Compaction canceled.", draw nothing. - Both read the end with the one compaction rule: success needs the provider's report of the compaction; none after Orca's interrupt is a cancellation. - The message resolves at the provider's receipt, as any send does: the Codex ack, or the Claude slash-command waiter on its result. The host writes a command's end only when the provider never took it. - The delivery loop stops while a command's turn runs, and every journal commit re-wakes it through the session's serialize, so an end that lands while a step decides to stop is never lost. A child that ends first is settled with it. * test(native-chat): pin a command's end to real /compact frames and to each path it threads The captured /compact frames drive the Claude translator's command turn: a finished compaction ends as a success with only the separator drawn; a stopped one ends as a cancellation with no failure row, and the next send answers in its own turn; a result naming another input ends nothing. The command's end is checked at each point the ordinary result path threads through: the reopen latch after a failure, the settling of a child still working, the context facts the result reports, and the provider's own error row. On the host: a message held behind a command is handed over when the command ends just as the loop stops for it, a refused command settles as a failure and the loop moves on, and a Claude child that exits mid-command settles the command and hands what waited to a fresh child. * test(native-chat): tests merged from the base state which turn their rows belong to * refactor(native-chat): drop the child-end waiter nothing waits on A command no longer waits for its child here: its turn ends from the provider's frames or from that child's settlement, and the delivery loop is woken by the commit. The waiter and its test were left from the earlier shape. * fix(native-chat): a command holds the queue only while its child runs it The delivery loop stopped whenever the journal showed a command's turn running. When the command's child ended and its settlement could not be written, that turn stayed running with no child to end it, and the loop's gate kept it from ever starting the next child, which is what settles a gone generation's leftovers. Every later send was held for good, and Stop had no child to end. The gate now holds only while the conversation has a child: with none, the command belongs to a gone generation, and the loop's start settles it like any turn a dead child left running. * fix(native-chat): a Claude /compact succeeds only on its compaction boundary The command's evidence counted Claude's `compact_result: 'success'` status as the compaction done. That status comes before the boundary that replaces the history, so a Stop landing between the two read as a finished compaction even though no boundary was ever written. Only the boundary now counts, as the rule for both providers states; the capture's finished compaction carries one, so it still reads as a success. * fix(native-chat): a Claude child's exit says why the turn it ended stopped When a Claude child exited mid-/compact, the command showed "Worked for 0s" and no reason. The child's translator ends its open turn the moment the exit is reported, stamped with the exit's instant, so by the time the exit settlement ran nothing was running. The settlement recognises a turn the exit already ended by that same instant, but the Claude lifecycle event dropped it on the way to the host, which then used its own clock, matched nothing, and wrote no row. When the clocks did agree, the row was scoped to the running turn, of which there was none, so it landed outside the turn it explained. The exit's instant now reaches the host, and the exit row belongs to the turn the exit ended: still running, or ended by the translator at that instant. * fix(native-chat): a message waiting behind /compact draws below its live activity A message sent while /compact runs waits on the host until the command ends. Both clients moved it to the end of the transcript rows, but the running turn's live activity line ("Compacting the conversation") draws after every row, so the waiting message sat between the command and its own live status. A row that is queued, and not what the live turn is for, now draws after that live activity: on desktop outside the transcript window, below the activity line; on the phone in the list footer, below the live status. A message whose own start is pending still draws above the activity that start reports. * fix(native-chat): only a running command holds a message below its live activity A message is accepted, then handed over a moment later, and in between it reads as waiting. Every message waiting behind a live turn drew below that turn's activity line, so an ordinary message sent while the agent was working crossed below "Thinking" and jumped back up once it was handed over, on desktop and phone. Only a conversation command's turn holds the queue on the host. A message now waits below the live activity only while the running turn is one a command opened, read from the entry that opened it. The phone test also typechecks, which the mobile test ratchet requires. * test(codex): the claim test names its notification params as a record * test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the dead process's lease still reads live. The open settles the turn it left running anyway, and the restore that follows finds it settled. * refactor(native-chat): drop the composer's second error formatter After the merge with main, every chat write in the composer path reports its failure as a typed outcome worded by the refusal-notice table, so the send's catch sees only a local throw. The {code, message} formatter this branch added for it has no payload left to format, and its claim to be the one way a chat words a failure is no longer true. The composer send is main's again. * test(native-chat): pin the reason on a message rejected while its chat was closed The reopen test checked only that the message reads as not sent; it now also checks the Retry row carries the host's reason. * docs(native-chat): drop the removed dispatch hold from six comments A worker's session no longer takes a dispatch hold, and no release clock rests a chat by visibility; the agent-launch comments, the abandon test, the teardown test and the refusal census still said so. * test(native-chat): rest the owner-status chat through the idle sweep, not a hold The activation-gate test from #22808 put its chat at rest by holding and releasing it, and passed the release-clock grace. This branch deleted both, so the case threw before it reached its assertions. It now moves the host's clock past the idle window and lets the sweep stop the agent and close the conversation, then asserts the same owner answer and activation gate. * fix(native-chat): show the structured pane's retrying line when a read fails The read transport always hands the pane the host's words, so the error state's "Orca keeps trying to load it" line, which showed only when there were none, was never seen: the pane showed the host's text twice, as its subtitle and on the status line under it. The structured pane now always says its read keeps retrying, and the host's text stays on the status line. The terminal-backed chat is unchanged. * fix(native-chat): a send the provider never received after a restart has no verdict Restart reconciliation rejects a crash-stranded send that is absent from a trustworthy provider history with reason 'not_delivered'. Nobody failed that send, but the verdict allowlist did not name it, so after a crash the chat read Failed, was listed, and could notify "failed". Give the reason a shared constant (persisted value unchanged), add it to the no-verdict set, and treat it as an internal marker so the Retry row no longer shows the raw string. * fix(native-chat): a failed Codex compaction's late completion writes no turn of its own Codex ends a failed turn with an error and then still completes it as failed. The error settled the compaction and released its claim on the provider turn, so the completion read that turn as an ordinary one and wrote a stray record. The claim now lasts until the completion, which adds nothing to a command the error already ended. * test(native-chat): the mid-command exit case resumes its next child as a real one does The case's fake started every child as a newly created thread with the same generation. The store refuses a created link once the conversation has a thread, so the next child's start failed and wrote its own error row, which landed before or after the case read the journal. The next child now resumes the thread under its own generation, and the case reads the journal once the waiting message is delivered, which also proves the loop moved on. * test(native-chat): wait for a send's background start before the refusal oracle removes its store An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals. * fix(native-chat): a start a message waited on gets one failure row, the delivery loop's When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice. The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit. * fix(native-chat): a /compact whose start failed says to run /compact again The failure-words context named only /clear as a command to retry, so a /compact whose agent failed to start read "Send your message to try again." on its row, its rejected message and the command reply. The context now carries any conversation command; the host derives it from the oldest message still waiting on the provider, which is the one a failed start fails first, and the /compact reply names it directly. * fix(native-chat): a Codex /compact ends only on its turn's completion, below Codex's own error row Since only turn/completed ends a Codex turn, Codex's turn-ending `error` is a row inside the still-open command turn, and the failed completion that follows it is the command's end: completed, outcome failure, at the completion's receipt time. The command's own "Compaction failed" row was written on that completion too, so a failed /compact read its reason twice. The command turn now notes when Codex's turn-ending error for the turn it carries was written as a row, and its end then adds no second row. A retried stream error ends nothing and is not counted. The flag that let the error end the command and kept the claim until the completion is gone with the error-driven end. A test replays the captured failed compaction from the real app-server through a claimed command turn. * test(native-chat): main's crash-turn test states its row's turn, and a dead /compact settles on its recorded exit Two tests the main merge brought together: - The crash-turn test from #23456 writes a turn record through the event sink without options; every row here states its turn scope, and a turn record's is the thread. - The /compact whose exit settlement could not be written no longer stays running until the next start: main now settles an open chat from the exit it recorded, so the command reads interrupted before the next message, which is then delivered. * test(native-chat): main's new journal tests state each row's turn The crash-turn, stale-turn and sink-queue tests main added wrote rows without a turn scope, which every item write now states. Rows written inside a running turn name that turn; the sink-queue batch and a send handed over with no live turn name the thread. * fix(native-chat): draw a queued turn's message after the earlier turn's rows A message sent while A runs is written to the journal when it is sent. When the provider queues it (Claude answers it after A), A's remaining rows - its last tool run and its answer - are written after that message, and the message's own turn opens only after them. Grouping put those rows in A's turn, but the transcript still drew them in journal order, below B's bubble and bar, where A's answer read as B's reply. This is the residual #23671 left open. A message that opened a turn now draws after the earlier turns' rows the journal wrote after it, just before its own turn's rows (nativeChatTurnDrawOrder, returned by nativeChatTurnMembership as drawOrder). Desktop and mobile both draw in that order. A steer, and a message that has opened no turn yet, stay where they were written. It applies on hosts that state turn scopes and, through journal order, on older ones. * test(native-chat): run #23026's Stop tests against #23059's command turns Two of #23026's tests call APIs #23059 changed, and failed after the merge: - codex-structured-conversation-stop: a compaction now goes through adapter.compact with the command run the host wrote (#23059), not a bare turn id, and answers with the provider's receipt. With the command claimed, a Stop that names no turn while the compaction's provider turn has not opened still interrupts nothing. - main-agent-working-agreement: a provider row states its turn scope (#23059's appendItem contract); the retry and subagent rows are conversation-scoped. * fix(native-chat): typecheck main's Stop and restore-grouping code against #23059 A Stop's compaction interrupt reads the narrowed requested turn, and the restore-grouping test states whether each row reports its turn's outcome. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
62e3a5a9a2 |
fix(codex): turning Codex off per agent keeps its hook entry off (#23667)
* refactor(agent-hooks): one predicate for whether an agent's status hooks are on "Global switch on and this agent not turned off" was spelled out separately in the startup controls, the settings reconcile, the retained-home reconcile, the WSL preflight RPC, the CLI preflight and the OpenCode plugin selection. They now share one function, in a module light enough for the CLI's per-launch Codex preflight to load. The PTY spawn env derives the Codex flag from the switch and opt-out list it already carries, the same way it does for OpenCode and Pi, instead of receiving a second copy. * fix(codex): launch and resume prep honour Codex's per-agent hook opt-out Turning Codex off in the per-agent hook settings removes Orca's Codex hook entry, but launch prep and session resume read only the global hooks switch, so the next Codex launch or resume wrote the entry straight back into the real ~/.codex or the account's home. Both now read the per-agent predicate, which the PTY spawn env and startup already honoured. * fix(codex): turning Codex off per agent clears the real ~/.codex entry While the real-home lane owns ~/.codex/hooks.json, the legacy system-home sweep stands down. That gate read only the global switch, so turning Codex off per agent ran remove() with the sweep still suppressed and left Orca's entry in the real ~/.codex. The gate now reads the per-agent predicate, the same as turning every hook off. * test(codex): cover the system ~/.codex sweep gate for Codex turned off The gate that lets the legacy system-home sweep run was an inline closure in startup, so reverting it to the global switch left CI green. It is now a pure function beside the gate it feeds, with a table test and a remove() test on a seeded ~/.codex: turning Codex off strips Orca's entry and keeps user hooks; with Codex on the entry stays. * fix(cli): keep the agent-status hooks predicate loadable by the packaged CLI The CLI's prepare-codex handler imported the predicate from src/main, but the Electron build rebuilds out/main from its declared entries only, so the packaged `orca agent hooks` commands could not load it (package jobs and the CLI bundle-parity test were red). The predicate reads only settings, so it now lives in src/shared, which the CLI compiles itself. |
||
|
|
31012aeb09 |
test: remove assertion-free probes, copied inventories and export-shape checks (#23816)
Second audit wave, targeting three more junk patterns: - assertion-free cases that run code and assert nothing, so they pass no matter what the code does; - inventory literals re-typed from a production declaration, where the only way the assertion can fail is someone editing one of the two copies; - export key-set and export-shape loops (`typeof x === 'function'` over every export) that restate what TypeScript already enforces. Yield is much smaller than wave 1 on purpose: the assertion-free scanner has a high false-positive rate, because many flagged blocks assert through a shared helper or their oracle is "this must not throw". Those were kept. `mobileWebCheckArgs` in `config/scripts/run-mobile-web-app-checks.mjs` is de-exported — after the inventory comparison went away, nothing outside the module read it. |
||
|
|
c49388cd33 |
fix(native-chat): Stop is there from the moment a message is sent (#23026)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * fix(native-chat): name a chat write by its target, not the owner generation A write carried the fence of the last frame the pane read, and the host refused it unless that fence was still current. An idle release and the restart after it each move the fence, and the release publishes nothing, so a send after a release was refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a cold start was refused as stale. Every write already names what it acts on: a send its conversation, a cancel its turn, a prompt answer its item revision, a rewind its epoch; an option is last-writer-wins. So admission stops comparing the client's fence, and the rebase that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it. The writer-lease check stays, and so does the attach's compare-and-swap. Frames now stamp the fence read when each frame is sent instead of a copy each subscriber kept, which went stale on the same release. * fix(native-chat): every journal append reaches the chats that are open A journal write and its delivery to open readers were two calls, and some writers made only the first. A failed start whose lease could not be handed back, a provider revision with no frame behind it, and eviction's settlement were all journaled without reaching an open chat. A journal handle now reports every durable change, and the host's session map binds that report to the session's readers when the handle is set. Writers no longer publish what they append; the per-writer publish calls are deleted. * test(native-chat): an epoch replacement reaches the open chat * test(native-chat): each row reaches an open chat once, and a live handle enters only through the map * test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite The seeded record had no surface tab id, so the next open backfilled one and that rewrite alone made the no-op transaction write. The test passed with the legacy-lease rewrite signal removed. * test(worktree-activation): restore the OMP surfaced-agent resume test The handoff removal deleted it alongside the terminal-owner tests, but it covers the surfaced-PTY block that still guards resume, including an agent whose ownership is unknown. * perf(native-chat): a publish behind a delivered commit reads nothing Each commit now delivers itself, so the publish a provider frame still sends afterwards found every reader caught up but still read rows and rebuilt the timeline for each one. A caught-up reader now skips the read. * test(native-chat): state why the teardown test's fake journal is safe to cast * docs(native-chat): say mutation admission checks only the writer lease * docs(native-chat): drop the send rebase from comments that still described it * fix(native-chat): a message is accepted, then delivered A send to a chat with no running agent restarted the agent inside the send call, before the message was recorded, so the client waited for the whole start and a failed restart refused the message. Claude held prompts sent during startup, and those could settle as "unconfirmed". A send is now accepted inside the session's serialized queue: one ledger row and one submission row marked handoverRecorded, published, answered pending. A per-session delivery loop exists while a message is queued. It starts the agent through the same serialized attach a hold uses, waits outside the queue for a Claude child to prove its start, and hands the oldest queued message over as its own serialized step, writing dispatch{pending} before the adapter call. A start it needed and did not get writes one error-tone row and rejects every queued message with the same words; a start Stop cancelled writes none. Settlement follows from the rows. A queued message is provably unwritten, so a close, an eviction or an exit rejects it. A handed-over message stays in doubt. A queued row at or below the sequence a handle found when it opened was left by an earlier process and is rejected at open, with no latch. Stop withdraws queued messages with no writer lease and no fence. An attach failure keeps the conversation open, and the attach adopts its journal. Owed work counts the loop and queued rows. A compaction or rewind found prepared when a conversation opens was started under a child this process no longer has, so the open settles it rather than leaving it to refuse every send until a view attaches. The open cursor is scoped to its epoch, because sequences restart when an epoch is replaced. Deleted: restart-before-admission, recordFailedRestart, the fence rebase, Claude's startup gate, the attach's forget on failure and its own crash boundary. Clients without agent-session.accepted-send.v1 get their reply held until the handover; the desktop and paired desktop lists advertise it. * fix(native-chat): settle queued messages only for the child that ended A child that proved its start and then exited before its message was handed over left the message queued: the exit settlement returned early when nothing else was in flight. Delivery then started another child for it, and a child that died the same way started another, without end and without a row. A retried settlement for an earlier generation, run by the attach that delivery started, did the opposite: with that generation's turn unfinished it rejected the message queued for the child being attached. The settlement now takes the rejection for queued messages from its caller. The unexpected exit and the eviction pass one, and it applies even with no other work in flight; the retry for an earlier generation passes none. * fix(native-chat): an adoption that fails to import keeps the conversation open The attach now writes into the conversation's own open journal, but a failed transcript import still closed it as if it were the attach's provisional one. The conversation stayed indexed with a closed journal, so every later send answered "could not be recorded" and every attach failed again until the app restarted. The import now closes only a journal the attach opened for itself. * perf(native-chat): the recovering open reads the journal once Every conversation open now goes through the recovering open, including the read restore of every chat at startup, which used to replay its journal once. The recovering open replayed it twice: once to probe it and again inside the open. The probe is now handed to the open as its load. * fix(native-chat): an attach that fails after indexing its child leaves no child behind A failed attach now keeps the conversation open, but a failure after `onAttached` indexed the child (the rewind or compaction recovery, or the attach's own success record) left that entry claiming a child the failure path had already released. The next send found the phantom, skipped the start, and wrote at a fence the journal had moved past, so the message stayed queued for good. The entry now drops the released child and its event sink, and follows the record's fence, as a failure before indexing already did. * fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer The error strip for a message the host accepted and then did not deliver matched the entry before the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not send your message" with nothing to retry. It now reads the reconciled entry. A rejection the journal records before the send's own pending answer lands is final as well: that answer no longer puts the entry back to dispatching with no Retry. * fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down The preamble waits for its submission to be delivered while the worker's agent starts. When that wait ran out it threw operation_unknown, and the failed-start teardown then closed the session, which rejected the very preamble the host was about to deliver. It now reports a turn start nobody observed yet: the worker is start-unknown with its session kept, the host delivers the preamble when the agent starts, and the worker's report settles the dispatch as for any unobserved start. The receipt no longer suggests reading a screen a structured worker lacks. * fix(native-chat): a message rejected while its chat was closed reads as not sent A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked every later message behind a Retry and no reason, and the delivery probe, seeing the journal already answered, never ran. The reconcile now settles it as rejected like a dispatching one. * test(orchestration): name why the readiness settlement fakes are cast * fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent * docs(native-chat): drop the fence from the admission the send effects run behind * docs(native-chat): give the fence move on release the reason that still holds * docs(native-chat): stop citing a write fence check in launch and mailbox comments Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease. * refactor(native-chat): the provider child is its own record A conversation now outlives any number of provider children, so the child is one record on the conversation's entry instead of five loose fields beside its journal. It is written in one place: indexed only once an attach has fully succeeded, and ended through one function that an exit, a failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence. - A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence patch after it are gone. - Conversation writes read the record's fence, the way mutation admission already does; a child's own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of the conversation's fence, are gone. - The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer dropped when an attach replaced the whole entry. - Stop on a child still proving its start stops only the child: its lease goes back and the chat is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus the conversation's close. - The settlement retry uses the conversation's own journal, opened through the host's one open. * fix(native-chat): the delivery loop alone settles a message its start or child failed A queued message was settled by whichever path happened to end the child first: the loop, the unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that rejected every pending row. That gave two failure rows with different tones for one start, a loop that could hand over to a different child than the one it waited on, and a Claude start that died while starting reading unlike every other failed start. - The loop remembers the child it waited on. At handover, if that child is gone or replaced, it reads how it ended: a Stop continues; anything else writes one failure row and rejects every queued message with the same words, then stops. A child still starting whose start the adapter says did not land fails the same way. The exit, eviction and the settlement retry only settle the handed-over and legacy rows of the child that ended. - One failure row, always an error, keyed by the start. A start a view began that dies with nothing queued writes the same row through the same builder, so a second report revises it. - The open no longer rejects leftovers; the loop's first step does, and the open wakes it. - `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the failure before the exit is processed. - Quit closes every conversation the way closing a chat does: what is still queued is rejected as closed, with or without a child, and a start the loop already has in flight is waited for so the child it produces is stopped rather than left behind. * refactor(native-chat): a stopped child ends on the one reading of its stop The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that verdict to the child's ending, so the host never forms a second view of whether the root is gone. Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition, and a failed re-attach passes what its release saw. The end-of-child record can therefore also carry a stop whose root was not seen to go, which nothing ends on yet. * feat(native-chat): the host says it accepts a send before any agent has it The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same string capable clients already send. A client can then tell a host that answers a send at acceptance, and admits a Stop with no writer before a turn starts, from an older one that still restarts the agent inside the send. Additive: an older client ignores a capability it does not know. * refactor(native-chat): an attach never opens a journal of its own The attach adopts the conversation's open journal, which outlives it, so it no longer opens one for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag that told the two cases apart is gone. Tests that attach without a host open the conversation the way a host does. * fix(native-chat): a moved fence resends nothing on a host that accepts first The outbox treated any fence change as a new owner: it dropped the answer of a send in flight, queued that send to go out again under the same id, and unblocked a refused head. On an older host that is how a send the restart refused, unrecorded, gets another try. On a host that records every send before it starts an agent, a fence moves because that start ran, so the same rule resent into every failed start. With a fence stamped on every frame, that became a loop. The outbox now reacts to a fence change only when the host has not advertised that it accepts a send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed start reaches the client as a rejected message it keeps with its Retry. Against an older host, or before one has answered, the outbox behaves as it did. Desktop and paired web share this hook. * refactor(native-chat): a child's end says whether the user or the host stopped it The end-of-child record's cause now tells a user's Stop from the host stopping the child for a cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a user's Stop, as before, and fails the start it was waiting on after a host stop, with the one error row and every queued message rejected, in the stop's reason when it gave one. The reason stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet. * fix(native-chat): a chat whose only work is a queued message is not offered for resume A message accepted while the agent was starting counts as working in the chat, and quit rejects it as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a chat whose agent never had the message. The snapshot now reads only what was handed over. * test(native-chat): type the queued-message fixtures in the resume-offer tests * fix(native-chat): a start that dies while a message waits on it is that message's failed start Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When that start died, its exit wrote the start's error row and left the message queued, so the delivery loop started a second agent into the same failure and wrote a second row. A child's end now records where the conversation's journal stood, and the loop settles a message accepted before a failed start ended with that start: one row, under its key, and no second start. A message sent after the failure still gets a fresh start. * docs(native-chat): say what an attach's open conversation and unconfirmed ids are now * test(native-chat): pin what a failed start settles, and what a resume offer names A view's child that dies while a sent message waits settles that message only when it died starting and no child has taken its place: a proven child's crash, or a second start since, gets the message delivered. The resume offer names the handed-over message, never a newer one still queued. * test(native-chat): the failed-start pins fail on what the message became, not on a timeout * test(orchestration): the preamble's host stub is typed, not cast The preamble send now takes only what it reads of the host, the send, the settlement wait and the record's fence, so its test builds that host with real types instead of `as never`. * fix(native-chat): a Stop that names no turn stops what the conversation has in flight Between handing a message to the agent and the agent opening its turn, there is no turn id a client could name, so a Stop in that gap was refused as "already finished" while the agent went on to answer. A cancel's turn id is now an optional precondition instead of its target: with none, the host withdraws what is queued and, when the journal still reads working, asks the adapter to stop whatever the child has in flight. Claude's interrupt is session-scoped, so it is guarded by fence and acquisition generation rather than a turn identity. Codex interrupts the turn its latest turn/start answered with until the journal shows one. A cancel that names its turn behaves exactly as before. * fix(native-chat): Stop is there from the moment a message is sent The composer showed Stop only once the agent had opened a turn, so for the second or two after a send the chat read "thinking" with no way to stop it. Against a host that takes a Stop naming no turn, Stop now shows whenever the chat reads working (a turn, a queued message, or a handed-over one still unanswered) or this client still has a message on its way. Pressing it, or Escape, first drops every outbox entry the journal does not hold yet, so nothing goes out after the Stop, then sends the conversation-wide cancel. A send already on its way reaches the host ahead of the cancel, which withdraws it there. Against an older host Stop still needs a running turn. The unconfirmed-send probe moves into its own hook so the outbox hook stays in budget. * fix(native-chat): Stop before a turn is gated on its own host capability A host that accepts sends first (agent-session.accepted-send.v1) can still predate the cancel that names no turn and would refuse it as invalid, since clients and hosts ship independently. Hosts that take that cancel now advertise agent-session.conversation-stop.v1, and the renderer shows Stop before a turn opens, and sends the no-turn cancel, only to a host advertising it. Every other host keeps a Stop that needs, and names, a running turn. The host capability probe the accepted-send hook used is generalized so both read one path. * test(native-chat): a build advertises conversation stop exactly where its cancel may name no turn * fix(native-chat): a view never restarts a chat whose last start failed A Claude chat whose CLI exits during startup left one red row per start, and every time a view bound to it (the chat opening right after its create died, or the user switching back to it) the hold started the CLI again, so the same launch-failure row repeated. Only a send retries a failed start now, the same rule provider-exit recovery already applied; the rule lives in one predicate the hold, exit recovery and the delivery loop share. * test(native-chat): start the child the loop waits on with an attach, not a second view A view no longer starts a child whose last start failed, so the R2 case that waits on a child started since the failure now gets that child from a client attach, the one non-send starter left. * fix(native-chat): settle a gone generation's turn wherever a conversation opens A send that opens a chat this process had not read yet (after a crash, from a phone or the CLI) went through the delivery open, which never settled what the dead generation left running; only the read restore and a successful acquire did. When the send's start then failed, the turn stayed running for every reader. The settlement now runs in the one journal open, at the crash boundary, for every opener except an acquisition, which settles from the evidence it read before its reserve; the read restore's separate step is gone. * test(native-chat): prove the next child's start settles the turn an earlier child left The R1 case lost its only settlement assertion when the latch it checked was deleted. It now seeds the running turn the earlier child left and asserts it ends at the exit's receipt, with the exit's row, before the message is handed to the new child. * test(native-chat): count a failed start's rows by row, not by text Comparing the set of texts passed when two different rows carried the same words, which is the duplicate the test exists to catch. * test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget * test(native-chat): pin the open's and the send's start and row counts, however the view binds Opening a fresh chat whose starts fail makes one start and one row, with two views bound before or after the create's child died; one send makes one more of each. * fix(native-chat): settle a gone generation's turn at every open but an acquisition's The journal open skipped the settlement whenever the lease read reserved or live, to leave an acquisition's own open to the acquisition. But a lease a crashed process left in recovery also reads live, until the next acquire resolves it. A send that opened such a chat, from a phone or the CLI after a crash on a host that could not prove the old owner gone, skipped the settlement; when its start then failed, the dead turn stayed running for every reader. The acquisition now says it is the opener, and every other open settles, whatever the lease still claims. * test(native-chat): hold the create's start open until the views bind The "view binds while the create is still starting" case gave the create a 300 ms head start and asserted the views bound before it died. On a loaded runner the holds took longer, the create's exit landed first, and the case failed its own precondition. The create's initialize now waits on a gate the test releases once the views are bound. * fix(native-chat): Stop reads the one working rule every session list reads While Claude retries a rate-limited request it never echoes the message, so no turn opens: the sidebar read Working from the unanswered send while the composer showed Send. The chat's working state, the host's session-list status and the host's no-turn Stop check now call one shared rule instead of three copies. * test(native-chat): a rate-limit retry pins only that no turn opens, not how its rows are kept * fix(native-chat): Stop leaves a message waiting on its Retry, and does not show for one A send that failed holds the queue until the user retries it, and one the host restarted under is parked the same way. Stop counted both as still on their way, so it showed in an idle chat and could never go away, and pressing it dropped the failed message along with its Retry. * test(native-chat): the chat's Stop and a session list read the main agent alike over their own copies The chat reduces its stream and a list reads the status feed. Driven through the real host for a rate-limit retry with no turn, a subagent still running after the main turn, and the handed-over child exiting. * refactor(mobile): the chat reads the main agent's working state through the shared rule Behaviour is unchanged: the same two terms, now from the one function the host projection and the desktop chat read. * fix(codex): a Stop naming no turn never interrupts an earlier turn It fell back to the id an earlier turn/start answered with when the latest start went unanswered, or when the journal showed a compaction Codex had not started, and reported that as stopped. * fix(native-chat): a Stop naming no turn never says a turn had already finished When the provider found nothing left to stop, for instance a turn that ended between the host's check and the interrupt, the chat got "The provider had already finished this turn." for a turn the Stop never named. It now ends quietly, as a Stop with nothing in flight does. * fix(native-chat): one Stop the host could not settle no longer refuses every later one A Stop naming no turn has one operation key per session. When the host could not settle one, it answered every later Stop under the same id as unknown until the id expired. Once the host says so, the next press is a new Stop; transport doubt still replays the same id. * refactor(native-chat): drop the composer's second error formatter After the merge with main, every chat write in the composer path reports its failure as a typed outcome worded by the refusal-notice table, so the send's catch sees only a local throw. The {code, message} formatter this branch added for it has no payload left to format, and its claim to be the one way a chat words a failure is no longer true. The composer send is main's again. * test(native-chat): pin the reason on a message rejected while its chat was closed The reopen test checked only that the message reads as not sent; it now also checks the Retry row carries the host's reason. * test(native-chat): read Stop operation ids without a cast * fix(native-chat): a Stop whose answer was lost no longer swallows the next one A Stop that names no turn has one operation key per chat. When its answer was lost in transit, the chat kept the id, so every later Stop replayed it; the host answers a replay as already handled, so for up to a day Stop stopped nothing. The id is now dropped once the call settles, however it settles. A second press while the first is still on its way still shares its id. * refactor(native-chat): a Stop naming no target keeps its operation id only for its own call The chat kept each write's operation id per payload across calls, and dropped it only on some settle paths. That is right for a write naming what it acts on, but a Stop naming no turn, and a stop of every background task, share one payload with every later one, so any path that kept the id made the next Stop replay as already handled and stop nothing. One path was still open: an answer that arrived after the chat moved to a new fence. Whether a write names its target is now decided once, before its id is picked. One that names none keeps its id only while its call is in flight, so a press made meanwhile joins it, and releases it when the call settles, however it settles. The release runs only while the key still holds that call's id, so a joined call settling late cannot drop a newer one's. This replaces the per-path exceptions for a thrown call. * test(native-chat): read the Stop fences without a cast * test(native-chat): pin the new id for a named cancel the host could not settle After the Stop naming no turn moved to a per-call id, the only test of the unknown-refusal release was gone, and the half that stays, for a cancel naming its turn, could be removed with every test green. * fix(native-chat): a Stop pressed after a new message stops it, even while the last Stop is unanswered A Stop naming no turn shared its operation id with any press made while it was still in flight. The host runs a chat's writes in order, so a message sent between two presses was accepted after the first Stop ran, and the second press replayed that Stop as already handled and left the message running, although the chat had already withdrawn it from the outbox. A write naming no target now gets a new id on every press and is never kept, so each Stop acts on whatever is running when the host reaches it. A write naming its target keeps its id exactly as before. A double press can ask the provider to stop the same turn twice, which it tolerates. * fix(native-chat): Stop no longer blinks off as Claude opens the turn for a message Claude's echo of a sent message both answers the send and opens its turn. The echo settled the send first, so the host published the message as answered one frame before the turn it opened, and for that frame the chat read nothing running: Stop turned back into Send, and Working blinked off in every session list, for tens of milliseconds on each turn. The echo now settles the send after the turn it opens has been emitted, so the running turn is published first. * fix(native-chat): a message a Stop withdrew comes back to its sender's composer A Stop withdraws every message the host holds but has not run, and S also drops the ones this client had not handed over yet. Either way the message left the chat and its text survived only in a hidden journal row and the in-memory ArrowUp history. The sending client now puts the withdrawn text and images back in that pane's composer, after whatever is typed there. Withdrawn is read from the rejection reason through one shared check, which the outbox reconcile now uses too. The composer is written before the entry leaves storage, so a failure between the two repeats the text instead of losing it, and an entry storage no longer holds is never given back again, so a replay, a second view or a remount restores it once. Only this client's outbox holds the entry, so other viewers still see the message disappear. A failed Stop withdraws nothing on the host and gives nothing back. * fix(native-chat): withdrawn text put back during an IME composition is not lost While the IME owns the field, the composer ignores a programmatic draft, and the next composed keystroke wrote the draft without the restored text, after its outbox entry had already been dropped. The composer now holds text appended mid-composition, keeps it in the cache after each composed write, and shows it once the composition settles, the way attachments that land mid-composition already wait for it. * test(native-chat): pin that only a withdrawn message comes back to the composer * test(native-chat): set up the composer's window API for every describe in the composition-race file * docs(native-chat): note that the withdrawn check reads the legacy reason until a typed category lands * test(native-chat): pin that text put back mid-composition shows once, even beside a mid-composition clear * fix(native-chat): land a late settlement from a streamed turn's end after that turn's rows A settlement that says a streamed turn ended waits for the session's event sink to drain before writing its dispatch row. The journal reducer still refuses to overwrite an accepted or rejected send. * fix(codex): settle a send from the end of the turn Codex answered it into The turn/start answer names the turn that holds a send. The send's echo entry now keeps that binding, in memory only. If the bound turn is interrupted without echoing the send, the send is withdrawn: Codex clears a turn's pending input on interrupt, so the model never saw it. If the turn fails first, the send is rejected in Codex's words. A completed turn settles nothing, since Codex records pending input when it finishes and the echo is still due. An answer read after its turn already ended is settled by that end. The echo is still the acceptance and carries the item key. * test(codex): a send settles from the end of the turn Codex answered it into The fake Codex keeps 0.157's turn bookkeeping, and can deliver the turn/start answer after turn/started or after turn/completed. The tests cover: - a Stop before any echo withdraws the send, and the working rule reads idle; - a steered follow-up is withdrawn when the turn is interrupted; - a failed turn rejects the send in Codex's words; - a completed turn leaves the send to its echo; - a normal echo and a late echo; - two steered sends in one turn; - an answer read after the turn ended; - a timed-out answer; - child-thread turns; - how a binding dies. * refactor(native-chat): drop the stream flush before a late turn-end settlement Nothing reads the order of a dispatch row against the turn's terminal row: the reducer keeps a settled send terminal and the working state is derived from both. The echo acceptance on the same path never waited either, and the wait could drop the settlement on a failed sink barrier. * fix(codex): settle a failed turn's sends at its end, not at its error Codex keeps a failed turn's pending input and records it after the error frame, before turn/completed. Settling at the error rejected a steered follow-up the model had in fact received, so a Retry would send it twice. * docs(codex): say a completed turn echoes what it took before it ends Codex records a completed turn's pending input before `turn/completed`, so a bound send that turn never echoed is left for recovery, not awaiting an echo. The comments and one test title said the echo was still due. * test(codex): settle a send whose answer is read after its turn failed or completed A failed turn that ended before the answer rejects the send in Codex's words, once; a completed one leaves it admitted and still armed for its echo. * refactor(codex): read a failed turn's reason with the typed thread-fact reader * fix(native-chat): a Stop that names no turn and ends nothing says why The host sends a Stop naming no turn to the agent only while the chat reads working. When the agent ended nothing, the Stop wrote no row, so it looked ignored. It now writes one: Stop could not reach the agent, in the agent's own words when it refused the interrupt. * fix(codex): hold a cold send until Codex opens its turn, and stop the turn it opened Codex answers turn/start before it opens the turn, and refuses an interrupt until then. A Stop queued behind a cold send ran in that gap, named the answered turn, and was refused. The send's handover now lasts until Codex opens that turn, or provably will not: the turn ended, the primary thread stopped running, or the child ended, bounded by the turn/start deadline. A steered send, whose turn is already open, does not wait. A Stop naming no turn now interrupts the turn the journal shows, else the primary-thread turn Codex reported started and not yet ended. The per-start answered id is gone: it was never cleared at a turn's end. * fix(native-chat): give back a send already on its way only when the host withdraws it Stop took the in-flight send out of the outbox and put its text back in the composer at once. The send still landed ahead of the Stop, so the chat read working for a moment before the host withdrew it, and a send the agent had already taken came back as well. The in-flight send now stays until the host answers it, and the withdrawn-message restore gives it back from that answer. * refactor(native-chat): read a Stop's withdrawal through the rejection classifier dispatchWasWithdrawn matched the legacy reason string. It now asks the classifier, which reads the typed fact first and keeps that string only as its own fallback, so a withdrawal written as a fact with a sentence is still given back to the composer. * fix(native-chat): a Stop Codex took but Orca could not confirm is not reported as reaching nothing When Codex acknowledged the interrupt but Orca could not verify the turn's processes ended, a Stop naming no turn wrote "it had no turn running to stop", though the turn then ended as interrupted. The adapter now says the Stop was taken but unconfirmed, and the row says Cancellation was not confirmed. * fix(codex): a held cold send never delays closing the chat or quitting A cold send's handover waits for Codex to open its turn, and that wait sat in the session queue ahead of the close and quit eviction, so either could wait out the 30 s request deadline. The host now releases those waits before it queues a close or starts quit teardown, and the adapter releases them when it is asked to close the child and on every exit, including one whose end publication is backpressured. * fix(native-chat): a refused Stop says the agent declined, and names it "Stop could not reach the agent" was wrong: the agent was reached and declined. The row now reads "Codex had no turn running to stop." or "Codex didn't stop: <Codex's words>.", naming the chat's agent. * fix(codex): a Stop waits for the turn Codex answered to open; the send no longer does Codex answers turn/start before it opens the turn and refuses turn/interrupt until then, so a Stop naming no turn in that window was lost. The send's handover used to wait for the turn to open, and a close or quit needed its own release to get past that wait. Now only the Stop waits. A Stop naming no turn, finding no journal turn and no open one, reads the turn Codex answered the latest pending send into and has neither opened nor ended, and waits for it: bounded at 5 s, under the quit eviction budget, and ended by that turn opening or ending, the thread going idle or failing, or the child exiting. If the turn opened it is stopped; otherwise the Stop reports that Codex had no turn running. The send returns at Codex's answer, so a close or quit with no Stop pending is never delayed, and the release plumbing through the adapter, router, close and quit is gone. * fix(codex): the Stop's wait reads a stopped thread from the thread-facts reader main kept --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
6e1b7e7fa3 |
test: remove junk tests that assert source text instead of behavior (#23815)
Deletes 101 test files and trims 112 more, all matching documented junk patterns: exact source/import/string greps, copied inventories and export lists, duplicate invocations of a contract another test already owns, typeof-shape checks TypeScript already enforces, and self-comparisons. The largest group read a production `.ts` file and asserted on its text — for example a TaskPage test that required the source to contain `selectedRepos.find((r) => r.id === newIssueRepoId) ?? selectedRepos[0] ?? null`. Any behavior-preserving rename broke it; no behavior change ever did. Production-side follow-through: exports that only these tests imported are de-exported or deleted, stale comments pointing at removed censuses are dropped, and the reliability-gate registry, `cloud/package.json` test lists, and orphaned source-reading helpers are updated so nothing references a deleted file. Two files kept their real coverage and lost only the census scaffolding: `agent-status-producer-census.test.ts` now drives all five producers end to end instead of grepping the source tree, and `config-toml-trust-stale-writes` replaces an export-list parity check. |
||
|
|
bb91e39eb9 |
test(native-chat): main's Codex child tests expect a turn to end on its own turn/completed (#23783)
* fix(native-chat): a Codex child's turn ends on its own turn/completed #22553 ended a child thread's turn on an `error` Codex will not retry, reading the verdict module #23682 deleted when it made turn/completed the only end of a Codex turn. The two landed minutes apart with no textual conflict, so main no longer typechecks. Codex runs every thread's events, spawned children included, through the same per-thread handler: a turn-ending error is recorded as the turn's last error and the turn then completes as failed. So a child's failed turn/completed is its end, as on the primary thread, and a closed thread stays the one child ending with no completion. * refactor(native-chat): a closed Codex child thread is its own frame With a child's turn now ending only on its own turn/completed, the turn-ended frame carried a fixed turnId (null) and state (unverifiable), and endTurn took parameters nothing passed. Name the one remaining case: a thread-closed frame, and closeThread ending the running turn as unverifiable. Drop a test step that no longer exercised anything. |
||
|
|
b6c2de9f92 |
fix(codex): index a new account home before bridging history into it (#22971)
* fix(codex): index a new account home before bridging history into it Codex indexes every rollout present when it first creates its state DB, and the TUI gives up after 30s, so large histories broke a new account's launch. Fixes #20669 * test(codex): keep the account migration test from starting the real Codex binary Selecting an account starts the history bridge, which spawned codex app-server on the fixture homes and raced the test's temp-dir cleanup. * fix(codex): index a new account's most recent bridged history first Indexing a large history takes minutes, and Codex's /resume hides unindexed threads once a directory has any indexed one, so recent conversations stayed missing until the heal reached them in directory-listing order. * test(codex): cover the history bridge's quit, no-history and failed-link guards Drop the stop check before creating the state DB: it ran in the same tick as the check at bridge start, so it could never observe a quit. * test(codex): make the account heal tests fail closed instead of reaching a real codex Fake invocations now point at a nonexistent binary and a throwaway CODEX_HOME, and cover per-home failure memory and a session that fails during quit. --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
bd133058a9 |
refactor(sidebar): one subagent row for CLI and structured children, shared with the chat strip (#22565)
* test(sidebar): pin today's subagent rows and chat strip rows from legacy shapes Captured on the unmodified renderer: the compact and full sidebar child rows built from a legacy subagents snapshot, and the expanded chat strip built from a legacy background-task roster. A host that sends only these shapes must keep rendering exactly these rows. * refactor(sidebar): one subagent row for CLI and structured children, shared with the chat strip A child row's dot, name and detail are now decided once, by a shared row model (src/shared/agent-child-row-model.ts), and rendered by one piece (AgentChildRowContent) that both the sidebar's child rows and the chat strip use. The model reads the host's child views when a session publishes them and today's legacy shapes otherwise (the subagents snapshot for the sidebar, the task roster for the strip), so old hosts and CLI panes render exactly as before. From views it keeps what the legacy path lost: a finished subagent whose shell still runs reads monitoring through the shared child fold, a settled child reads by its outcome, the tool it runs shows as the CLI row shows it, and each row keeps its own clock. In the strip, work a child owns renders nested beneath it. * i18n: add the child row's Ended and ended-ago strings to every catalog * test(sidebar): a child reads the same in the sidebar and the chat strip A row-model table for every display state, and a parity table that renders the same child views through the compact sidebar row and the chat strip and asserts both show the same dot, name and detail. Covers a finished subagent whose shell still runs (monitoring on both, no stale tool text), sibling rows with their own clocks, a settled row timed from when it ended, and owned shells nested under their owner in the strip. The strip's view input is named childViews so it cannot be mistaken for React children. * fix(sidebar): keep the child display derivation loadable in the renderer The row model imported deriveAgentChildDisplayState from the view module, whose owner resolution reaches status subjects and, through them, agent-hook-relay's node:crypto. The renderer cannot evaluate that chunk, so the app booted blank (the renderer node-builtin boundary test fails on the previous commit). The display derivation (agentChildWorkOwnedLiveness, deriveAgentChildDisplayState, AgentChildDisplayState) now lives in agent-status-child-work-display.ts, which imports only the fold, liveness and the one-pass grouping both modules share (grouped-by.ts); the view module keeps the host-side projection. * fix(sidebar): the strip reads its parent's verdict, and the full row says Ended Review fixes: - The strip's view path took no freshness input, so once views are wired the same lost child would read unverifiable in the sidebar and working in the strip. agentChildRowContextForParent builds the one context a parent row gives its children; the sidebar uses it, and buildBackgroundTaskGroupsFromViews / the strip's new optional childRowContext prop accept it (absent: claims stand as reported, as before). - The full sidebar row showed no word for a child that ended with no outcome; its message line now reads the row model's (the message, or Ended). An unlabeled child falls back to its display state there too. - The summary order now includes failed, so a failed row is never dropped from the counts. - Parity now covers the full row for every state, a live parked child, an unlabeled child, and a lost or stale parent on both surfaces. * refactor(agent-status): the subagent snapshot, its normalization and equality get their own module agent-status-types.ts crossed the file-size limit once the row gained child views beside the main agent fact. The subagent snapshot shape, its admission normalization and the array equality move into agent-status-subagent-snapshot.ts (re-exported, so importers are unchanged); the three field caps it shares with the row move to the field normalization. * refactor(sidebar): one legacy builder family, one reader clock, frozen settled rows - The chat strip's task-roster rows are built in the shared row model beside the other two builders, with one placeholder set and the one detail rule. The module header states when each legacy builder is deleted. - A child's "No update" duration reads the parent's reader clock (receipt time for a mirrored parent); a view child's own stamp is moved onto that clock. The full sidebar row reads the same value as the compact row. - The failed->blocked / interrupted->idle collapse has one owner, shared by the sidebar row state and the strip header. - A settled strip row shows its run frozen at settledAt instead of a growing "ended N ago", so a strip of only finished work never wakes the 1 Hz tick. * test(sidebar): the sidebar row state and the strip header share one lifecycle word * test(sidebar): a child running a shell in its turn shows its Bash line with the shell nested beneath it While a child's turn runs its shell, the host records the shell both as the child's operation and as a live command the child owns. The strip shows the child working its Bash line with the shell row nested beneath it, the header counts only the agent, and the sidebar shows the child alone with the same text. * test(sidebar): re-pin the legacy chat strip golden to main's scoped goal-dock selector Main's #23725 rewrote the strip's goal-dock variants from `group-has-[...]/tasks:` to an ancestor-scoped `[[data-native-chat-background-tasks]:has(+[data-native-chat-thread-goal])_&]:` selector. The merged head renders byte-identically to main for this fixture; only the golden was captured before that change. |
||
|
|
357a2fed08 |
fix(native-chat): group chat rows by the turn that produced them (#23671)
* fix(native-chat): keep a turn's bar on the prompt that opened it A message sent while a structured turn runs appears in the transcript at once, so "the newest user message" is not the running turn's owner. The live "Working for" bar moved to the mid-turn message and counted from the earlier prompt's start, and a send queued behind the running turn counted its wait twice: once in the previous turn and again from its own send. Derive both from the host's turn records in one ordered pass: - The running turn's bar belongs to the user message its lifecycle row names (resolved exactly as settled timing resolves it). A message sent mid-turn gets no bar until its own turn opens; a send folded into the running turn never gets one. Surfaces fall back to the latest user message only when the host names no opener. - A turn counts from its send, but never before the previous turn in the journal ended (its recorded end, else its row's last host revision), capped at the turn's own start. The same origin feeds the live counter and the settled duration. Desktop and mobile share the derivation; no wire, host, or storage change. * feat(native-chat): derive each transcript row's owning turn from the journal Rows between a turn record and the next belong to that record's turn, so a message the provider folds into a running turn no longer captures the rows produced after it. A turn whose opener the host cannot name in the loaded window anchors to its own record instead of a bystander prompt, and shared nativeChatRowTurnKeys keeps positional preceding-user grouping for anything the host does not attribute (older hosts stay pixel-identical). * fix(native-chat): fold and time transcript rows by their owning turn on desktop A settled turn's bar now folds every row the turn produced, including rows after a mid-turn send; the steered bubble stays visible and carries no bar. A provider-opened turn renders its bar above its first row instead of borrowing the newest prompt, and row liveness follows the owning turn rather than the newest user message, so a running turn's rows stay live while a send waits. * fix(mobile): group phone transcript rows and bars by their owning turn Same shared derivation as desktop: the opener's bar owns every row of its turn across a mid-turn send, a steered bubble never grows a bar, a provider-opened turn's bar sits above its first row, and a running turn's tool rows stay live while a newer message waits behind it. * fix(mobile): declare the turn ownership map on the chat controller contract * fix(native-chat): one diff rollup per turn, and no wake-turn clock on a later prompt A turn's rows are no longer contiguous once rows are grouped by owning turn: a prompt sent before the running turn's last row lands among its rows. The diff rollup was drawn at every run boundary, so such a turn showed its rollup twice and the later prompt's rollup appeared under its bubble. It now renders once, under the turn's last row. On the phone, a turn keyed to its own record is not a user message, so when it ended its clock was treated as a replaced optimistic echo and handed to the newest prompt - a message sent during a wake turn got a bar with the wake turn's duration. Host-attributed turn keys now count as live turn keys. * fix(native-chat): keep a Codex turn's bar on its send until the echo lands Codex reports turn/started before it runs hooks and prewarm and before it echoes the send, so for that gap the turn names a provider key no alias resolves yet. Anchoring it to its own record left the running turn with no bar at all; treat the send still in flight ahead of the record as its opener. * fix(codex): restore each turn's record ahead of that turn's items Rows are grouped by the nearest turn record before them in journal order, which holds on the live path because a turn's record is written when the turn opens. Full-history restore (the fallback for Codex app-servers that reject excludeTurns) wrote each turn's items first and its record after, so every restored turn's rows were credited to the previous turn and turn 1's answer folded away. The restore now appends the record before the turn's items. The record itself is unchanged: same identity, state, outcome, opener key, endpoints and duration, one append each. The restore-order test now expects the record first, because that order is what keeps grouping correct; its old order was incidental, not a contract any reader relied on. Readers that key turns by id or opener key, or that scan for the newest record (all restored records are settled), read the same result in either order. Journals already imported in the old order stay as written; no migration. |
||
|
|
25d9c57e3a |
fix(native-chat): keep a child turn settling on an error Codex will not retry (#23808)
#23801 removed the child-path reading of Codex's turn-ending `error` along with the import of the module #23682 deleted. That was the wrong half to remove: the primary journal path can rely on Codex's failed `turn/completed` arriving within ~32 ms, which is what #23682 established, but a child turn has no such guarantee, and without the error as its end the child's lifecycle row latches on `working` for the life of the session. Three tests assert exactly that and could not run, because the unresolved import had been skipping the unit matrix since #23682 merged. The reading is restored inline against `readCodexErrorWillRetry`, itself restored to `codex-structured-thread-facts.ts`, rather than by reviving the deleted module: its `thread-stopped-running` arm lost its only consumer when #23682 rewrote the primary path, so restoring the file would re-add dead code. Also drops `pr-workflow-parallelism.test.mjs`'s read of `.github/workflows/track-community-prs.yaml`, which #23796 deleted while leaving the assertion behind. Same failure class, and it fails the same shard. |
||
|
|
643f93f009 |
fix(native-chat): drop the Codex error path #23682 removed (#23801)
#23682 established that only `turn/completed` ends a Codex turn and deleted `codex-structured-journal-provider-verdicts.ts`, but the child-turn reader kept importing `readCodexProviderVerdict` and branching on its `turn-failed` verdict. The module is gone, so the import resolves to nothing: typecheck and static analysis fail on every pull request, which skips the whole unit matrix behind them, and the esbuild pass behind Bun profile scope detection throws, so every run falls back to all six platforms. A closed thread is now the only child-turn end without `turn/completed`, which is what #23682 intended: Codex follows a turn-ending `error` with a failed `turn/completed` for the same turn 0-32 ms later, and that completion carries the duration and receipt time the error does not. |
||
|
|
b4c19f12c4 |
fix(claude): run structured Claude under the POSIX provider supervisor (#23476)
* fix(codex): the provider supervisor outlives its provider group when stopped A signalled supervisor forwards the signal to the provider group, escalates to SIGKILL after the grace, and exits only once the group is gone, so recovery's proof that the recorded pid is dead also proves the provider is. It refuses to spawn when its parent is already not the owner named in its spec, and watches that owner rather than whichever parent it first saw. The grace is a spec field. Recovery's SIGTERM stage now outlasts the supervisor's own stop, since a SIGKILL that lands first cannot be handled and leaves the group running. * fix(codex): a closed owner pipe no longer ends the supervisor before its provider group When Orca dies, the supervisor's stdout pipe has no reader. Provider output in the window before the parent-death watch fired raised an unhandled EPIPE that exited the supervisor with the provider group still running. * fix(codex): bound the supervisor grace so recovery's SIGTERM stage always covers it Recovery sized its SIGTERM stage from the default grace, so a launch with a longer grace would be SIGKILLed mid-stop and orphan its group with no test noticing. The spec now refuses any grace above one exported maximum, and recovery derives its SIGTERM stage from that maximum. * fix(codex): every supervisor stop asks the provider with SIGTERM first Owner death, stdin end after the grace, and a signal to the supervisor now all take one path: SIGTERM the provider group, SIGKILL it after the grace, and exit only once it is gone. The signal handlers are registered before the provider is spawned, so a stop that lands in the spawn window still reaps it. The longest stop grows to two graces plus the reap wait, and both recovery's SIGTERM stage and the connection's graceful close now wait that long before forcing, since forcing the supervisor sooner can orphan its group. * fix(claude): run the structured Claude child under the POSIX provider supervisor A close now stops Claude with a SIGTERM through the supervisor instead of letting stdin end finish the turn, and Orca's death stops it through the supervisor. * test(claude): pin the supervised stop against a real Claude CLI, opt-in * test(claude): a requested stop reads interrupted through the frames the supervised SIGTERM makes Claude emit * test(claude): show what the real CLI did when it never ran the tool * test(claude): Orca's death now reaps Claude's own tool through its SIGTERM * refactor(claude): take supervision from the spawn spec so the close ladder cannot disagree with the spawn createProviderSpawnSpec now reports whether it wrapped the provider in the supervisor, and the Claude spawner reads that instead of repeating the platform check. The close ladder's SIGTERM follows the process actually spawned. * fix(native-chat): derive quit's chat-eviction bound from the longest supervised provider close Quit's child-eviction phase was a hand-picked 8 s. It is now the sink drain plus the longest supervised close over Claude and Codex plus a named 1 s margin, so a provider close that grows widens it instead of silently outrunning it. A close's tree-kill fallback stays outside the bound: once main exits, the supervisor stops its provider group on owner death, which a new test now proves for a clean owner exit, and next launch's recovery settles the lease. |
||
|
|
94b5a7d256 |
fix(native-chat): only Codex's turn completion ends a Codex turn (#23682)
* fix(native-chat): the sink queue keeps a settlement's first batch, as the journal does The journal applies a lifecycle batch's settlement id once and skips any later batch with the same id. The deferred sink queue coalesced the same key the other way: a second batch replaced the first while it was still queued. So which record survived depended on whether the first had drained yet. A lifecycle batch now keeps the queued operation with its key, and a later one is accepted and dropped, which is what the journal does once the first is written. * fix(native-chat): only turn/completed ends a Codex turn Codex follows every turn-ending `error` (willRetry=false) with a failed `turn/completed` for the same turn, 0-32 ms later. That was captured from the real app-server on 0.141.0 and 0.158.0 across eight failure scenarios, and it is how Codex builds a failed turn: it records the error as the turn's last error, records any pending input, and then derives `failed` from that error when it completes the turn. The translator ended the turn twice: once on the error, and again on the completion, with a guard to make the first end final. Ending on the error threw away what only the completion carries: Codex's duration, and the completion's receipt time. It also forgot the turn before Codex recorded the turn's pending input. Now the error is only the row the user reads, inside the still-open turn, and `turn/completed` is the turn's only live end. A process exit between the two is the existing exit sweep's observed end, recorded as interrupted. A failed completion is stored as completed with outcome failure, live and on restore alike. Only `interrupted` maps to the interrupted state. The first-end-final guard is gone. Codex sends one completion per turn, the only redelivery Orca has is the retry of a refused frame (which changes nothing), and the settlement id already keeps the first record in the queue and the journal. * refactor(codex): delete the unreachable oversized-notification settlement The translator settled a streamed item when the transport rejected its notification as oversized. Nothing can produce that frame. The Codex stdio reader frames with `maxLineBytes: Number.POSITIVE_INFINITY` (codex-app-server-record-reader.ts), which it has done since the app-server records were uncapped. With an infinite limit the framer never reports `line-too-long`: no line, pending suffix or paused queue can exceed it. So the dispatcher never emits `frame:oversized-notification`, and the arm that settles it never runs. The arm, its helper module and its test go. In place of the test, the connection test now proves the reason: a notification past the old 16 MiB wire limit arrives whole, and no oversized frame is reported. * test(codex): replace the captured ids in the turn-endings fixture with synthetic ones The replay reads ids only to group frames, so the real thread, turn and response ids from the capture account carry nothing the test needs. The fixture moves beside the Codex tests that read it. * test(codex): use a neutral made-up status as the unknown-status example 'cancelled' read as a stop being recorded as a completion. * test(codex): a restored turn with a status Orca cannot place ends with no verdict Codex's history carries the same status field as the live completion, so the restore path is pinned to the same mapping: completed, and no outcome. |
||
|
|
2af897d7ea |
fix(codex): the provider supervisor outlives its provider group when stopped (#23466)
* fix(codex): the provider supervisor outlives its provider group when stopped A signalled supervisor forwards the signal to the provider group, escalates to SIGKILL after the grace, and exits only once the group is gone, so recovery's proof that the recorded pid is dead also proves the provider is. It refuses to spawn when its parent is already not the owner named in its spec, and watches that owner rather than whichever parent it first saw. The grace is a spec field. Recovery's SIGTERM stage now outlasts the supervisor's own stop, since a SIGKILL that lands first cannot be handled and leaves the group running. * fix(codex): a closed owner pipe no longer ends the supervisor before its provider group When Orca dies, the supervisor's stdout pipe has no reader. Provider output in the window before the parent-death watch fired raised an unhandled EPIPE that exited the supervisor with the provider group still running. * fix(codex): bound the supervisor grace so recovery's SIGTERM stage always covers it Recovery sized its SIGTERM stage from the default grace, so a launch with a longer grace would be SIGKILLed mid-stop and orphan its group with no test noticing. The spec now refuses any grace above one exported maximum, and recovery derives its SIGTERM stage from that maximum. * fix(codex): every supervisor stop asks the provider with SIGTERM first Owner death, stdin end after the grace, and a signal to the supervisor now all take one path: SIGTERM the provider group, SIGKILL it after the grace, and exit only once it is gone. The signal handlers are registered before the provider is spawned, so a stop that lands in the spawn window still reaps it. The longest stop grows to two graces plus the reap wait, and both recovery's SIGTERM stage and the connection's graceful close now wait that long before forcing, since forcing the supervisor sooner can orphan its group. * fix(codex): give the provider 1 s after stdin end and 3 s after SIGTERM to flush before SIGKILL The supervisor's stop was stdin end, 1.25 s, SIGTERM, 1.25 s, SIGKILL. Codex now gets 3 s after SIGTERM to flush its state. The two graces are separate constants, the longest stop they derive becomes 5.5 s, and a test keeps it inside quit's 8 s child-eviction bound. * test(codex): count eviction's pre-stop drain in the quit budget test Eviction drains the sink for up to 1 s before it stops the child, inside the same 8 s bound. |
||
|
|
c5330d0d52 |
fix(native-chat): stop killing processes that only inherited a chat's spawn tag (#23460)
* fix(native-chat): stop signalling processes that only inherited a spawn token A spawn token is an environment variable, so every descendant of a provider child carries it. The Linux-only startup scan treated any carrier no lease claimed as a lost provider child and sent it SIGTERM, which also hit editors, tmux servers and nested Orca processes the agent had started. Remove that scan's killing consumer; the token scan stays for the reservation probe, and recorded owners are still stopped by identity during recovery. * fix(codex): remove the token-scan kill path from app-server teardown Every descendant inherits the spawn token, so killing each pid that carries it can reach processes the agent started that are not the provider. Production never injected this path; teardown always uses the process-group and descendant-snapshot proof. Drop it, its deps, and the now-unused spawn-token argument. |
||
|
|
2ca4ecbc61 |
feat(orchestration): let a structured chat run orchestration as itself (#22568)
* feat(orchestration): inject the Orca session id into structured children and let the CLI act as it Every structured session's child (native Claude, native Codex, and the terminal view) carries ORCA_AGENT_SESSION_ID and reaches the Orca CLI. The CLI sends the id in the orchestration envelope; when present it is the caller, and a caller flag naming anyone else is refused before any request. The id is stripped from inherited PTY env and from the SSH host-CLI passthrough, and crosses into WSL so the host can refuse the cross-host claim. * test(orchestration): pin session id injection for native Claude, native Codex, the terminal view, WSL, PTY inheritance and SSH * test(orchestration): pin one caller precedence rule across every CLI verb that names its caller Adds the per-verb table (flagless acts as the session; a conflicting --from or --terminal is refused before any request; the session's own spellings are accepted), the enumerated guess population with its positive control, the structured worker's own handle, the identity-less refusal for an older child, the unchanged terminal agent, and the envelope. dispatch-show's --from only fills preview text, so it passes through unfenced and a session's flagless preview names the address the real dispatch writes. * refactor(orchestration): keep the identity-less marker reader to the marker; the id is checked first * test(orchestration): pin that a host refusal of the session surfaces verbatim from the CLI * fix(orchestration): keep the identity-less marker beside the id for CLIs that predate it A CLI older than the id, reached through a global install when a shell rc resets PATH, would otherwise guess a sibling's terminal in a chat that no longer carries the marker. It refuses on the marker instead; a current CLI checks the id first, so the marker never makes a session with an id identity-less. * fix(orchestration): refuse a conflicting --from on gate-list and task-list scoped by --run A --run listing needs no caller, so both handlers skipped the resolver and a --from naming another actor was dropped silently under a session. The conflict check now runs on that branch too; terminal callers are unchanged. * fix(orchestration): name this app's CLI by absolute path for a structured session's login shells A provider can run each command in a login shell: Codex runs zsh -lc, and the profile rebuilds PATH, putting a global install (possibly an older Orca) ahead of the directory Orca prepended. ORCA_CLI_COMMAND, which an agent resolves the CLI from first, is now the absolute launcher in that directory (the native launcher on Windows), so no shell's startup files can swap it. The PATH prepend stays for shells that read no profile. Found by the live coordinator run of the next PR. * test(orchestration): pin a structured worker's CLI command as this app's absolute launcher * test(orchestration): run the zsh login-shell arm in the real-shell lane that installs zsh The ordinary Linux unit lane has no /bin/zsh, so the zsh arm failed there with ENOENT. It moves to a live-shell file registered in the shell-contracts lane; the bash arm keeps running in every lane. The lane guard's detector now also sees a zsh spawned through the ProcessSpec program field, which is how this test escaped it. * fix(orchestration): omit a structured child's CLI command when no launcher resolves, and pin its instance A bare `orca` fallback named GNOME's screen reader on packaged Linux, and an inherited value named another app's CLI. The builder now deletes any inherited value, sets the absolute launcher only when one resolved, and pins ORCA_USER_DATA_PATH so a current CLI dials the instance that minted the id. Renames the marker reader to hasStructuredSessionMarker and records why the terminal view carries the id without the marker. * fix(terminal): name this app's CLI launcher by absolute path in every local terminal ORCA_CLI_COMMAND meant three things by lane: an absolute launcher for a structured session, a bare name for WSL, and nothing for any other terminal, so a structured session's terminal view lost it. Local terminals now get the same absolute launcher the structured lane gets; WSL keeps its guest command name, and a terminal whose launcher does not resolve still gets none. * feat(cli): hand a command to the session's own CLI when another Orca CLI was invoked A login shell can reorder PATH behind a global install, and an agent or its helper script can run bare `orca`, so the binary that answered depended on the agent following instructions. Orca's packaged launchers and bare-orca shims now export ORCA_CLI_SELF (outermost wins). At the CLI entry, when it names a different launcher than ORCA_CLI_COMMAND, the command re-runs once through the named launcher with ORCA_CLI_REEXEC=1 and exits with its status; both variables are consumed so no child inherits them. Dev launchers export no self on purpose, WSL and SSH names never qualify, and a launcher that cannot start leaves the command to run here. The Windows launcher no longer rewrites ORCA_CLI_COMMAND; the legacy ask protocol normalizes its resume command itself. * refactor(orchestration): declare which flag names the caller on each spec and refuse at the CLI entry Each handler hand-classified its --from/--terminal as the caller or a target, and the refusal of a conflicting caller flag ran inside the caller resolver plus two standalone calls for --run listings, so a new verb that read its flag raw would pass a sibling's handle to a pre-session host. Specs now declare identityFlagRoles, the CLI entry refuses a conflicting caller flag once from the spec, the resolver only applies the id-wins rule, and a test fails any orchestration verb that accepts --from or --terminal without classifying it. * perf(cli): keep the session caller check off the actor codec's module graph The check runs at the CLI entry for every command, and the actor codec pulls zod through the session record. Compare the session's own spellings as plain strings instead. * refactor(cli): spell a session's address from the one prefix constant, off the codec's module graph The Orca session address prefix moves to a leaf module with no imports, re-exported by the address codec, so the CLI entry check derives `session:<id>` from that constant instead of re-typing it and still stays off the codec's zod graph. Prose and test names say caller or Orca session id, not actor. * refactor(orchestration): drop the session id's terminal-view spawn now that the handoff is gone The terminal handoff was removed, so no terminal is ever a structured session: - delete the terminal-view identity env and its WSL passthrough, and their tests; - strip the session caller keys from every terminal's env unconditionally; - the CLI's own-address spelling moves beside the injected id in src/shared, with a test pinning it to the address the host's party resolver gives that session. * fix(terminal): run the Codex launch preflight through the CLI the terminal names Packaged Linux names the userData shim in ORCA_CLI_COMMAND, while the preflight ran the bundled launcher behind it. The CLI saw a different launcher and handed the preflight off to the shim, booting Electron twice before every codex launch. * revert(terminal): keep terminals on main's ORCA_CLI_COMMAND and Codex preflight Only a structured session needs an absolute ORCA_CLI_COMMAND; local terminals go back to naming none (WSL keeps its guest command), and the Codex launch preflight goes back to the bundled launcher. The CLI handoff is scoped to sessions, so a terminal's preflight can no longer be handed off and start Electron twice. This reverts commit |
||
|
|
153d3fd3fa |
feat(native-chat): Codex sessions write their subagents into the host status store (#22553)
* refactor(native-chat): the Codex acquire names its turn-boundary methods as a set Behavior-neutral: the same two methods stamp receipt time. Keeps the file under the size limit once the child-work sink lands. * feat(native-chat): Codex sessions write their subagents into the host status store A Codex child thread and each persistent command become host child records, fed through the same delivery, ingest and reducer the Claude lane uses. The child's own turn decides it: turn start is live, turn completion settles it with the outcome Codex reports, and a follow-up turn reopens the same record as a new run. Its open tool call, last message, usage and waiting-on-user flag come from its own thread's frames. A parent turn ending settles nothing. * fix(native-chat): close a Codex child's tool call by its item id alone A completion frame need not restate the tool it ran, so reading the tool name before closing left the call open and the record naming a finished tool. * test(native-chat): pin the Codex child-work evidence and every hop to the host's records Child turn start/end/follow-up, open tool call, last message, usage, waiting, the persistent command a child owns and its monitoring display, a primary turn end settling nothing, and session end. End to end through the real adapter: evidence after the journal and the legacy republish, and the parent state the records imply equals today's at every frame of a scripted session. Through the production runtime: a Codex session's child work reaches the status sink under its own address, and a provider exit ends it there. * test(native-chat): a Codex child's new run never inherits the last run's open call * test(native-chat): a Codex session with no child-work sink holds no evidence * test(native-chat): deliver a Codex child's announcement twice, as Codex does, before counting edges * refactor(native-chat): hand the Codex producer's pending edge over directly * fix(native-chat): name every Codex turn state in the outcome map; type the runtime test's fake opener * fix(native-chat): a Codex child's turn ends on the error that ends it, or on its thread closing Codex can end a child's turn with no turn/completed: an error it will not retry is that turn's own end (the verdict the transcript already settles the same turn on), and a closed thread ran its last turn. The executions, the one owner of child turn state, now end the turn on both, so the strip drops the child and its record settles (failed, or unknown for a close) together, instead of reading working for the life of the session. A systemError status is not an ending: Codex raises it for errors that leave the turn running. A child fact whose frame names no turn now belongs to the turn the child is running, instead of counting for every run. * test(native-chat): a Codex child's turn ending by fatal error or thread close settles strip and record together * test(native-chat): the Codex parity script reads a waiting child through the shared fold's waiting arm * test(native-chat): a Codex child row's journal attempt is its record's generation The journal numbers a Codex child's runs by the turns it observed on the child's thread; the host record numbers them by the runs its evidence opened. Both are keyed by the child's own turn id, so they must agree run for run, including when Codex reports the child's first turn before the spawn that announces it. * test(native-chat): a Codex session's end settles its live children and keeps the ended ones The host no longer erases a session's children when its provider goes away: a child still running settles with an outcome nobody reported, and a child that had already ended keeps what it said. The producer tests now expect exactly that, from the close path and from an unexpected exit. * fix(native-chat): a Codex subagent's shell is its open tool until the process exits Codex runs every agent shell through unified exec, so every subagent shell arrives with the source the persistent-command tracker keys on. The producer skipped those items, so a working subagent never named its shell, and an approved command (started on the approval path, completed from unified exec) stayed its open tool until the turn ended. The tracker still records the process separately, so a command that outlives the turn reads as monitoring. * fix(native-chat): a Codex shell becomes a subagent's own work only once it outlives its turn Codex runs every agent shell through unified exec and never says when one is left running, so the producer turned every shell, even a millisecond `rg`, into a command record the moment it started. Each settled into the session's pool of 32 settled records, so a busy turn evicted a finished subagent's record (its outcome row would vanish) and listed dozens of finished shells beside it. A command now becomes a record at the first turn boundary of the thread that launched it while its process still runs: until then it is the agent's open call. A shell that exits within its turn never becomes a record. * refactor(native-chat): child records keep every settled child and can be removed outright Settled child records now stay until the host drops the session's row; the 32-record trim is gone. A producer can say work stopped with nothing to report, and its record (and the handles it answered to) goes instead of settling. Evidence stays host-internal: the producer and the store share one process. * fix(native-chat): a Codex command is live work from its start until its process stops The command tracker is now the one owner of a Codex command's lifetime. It admits every command whatever `source` Codex tags it with (the approval path starts one as `agent`), and ends it when its process exits, when its thread closes (Codex stops the processes first, so no exit ever arrives), or when the session ends. The producer mirrors that one-to-one: a live record from the start, removed when the command stops, never settled. This removes the turn-boundary rule: a command that was only recorded at its turn's end left the parent reading done for one publish when the main agent's turn ended with a shell still running. The parity script now checks the parent at every journal write, not only at frame end. * fix(native-chat): a Codex command whose approval its turn abandoned never ran Codex starts an approval's command item before it asks, and when the turn ends with the question unanswered (the user stops at the approval), it drops the question and never completes the item. The command tracker admitted that start as a running process, so the strip kept a phantom command row and the session row read working until the session ended. The prompt registry, which owns which approvals are still unanswered, reports the command approvals a turn ended without; the tracker ends those commands with the frame that ended the turn. An answered approval keeps its command. * test(native-chat): start the Codex child-work runtime test without the removed hold Main no longer has host.hold: creating the session starts its child, and nothing a viewer does keeps it running. The test attaches and asserts the one child that attach started, then drives it as before. |
||
|
|
2c92e5e7da |
fix(native-chat): a failed Codex turn keeps its failure and "Worked for" (#23514)
* fix(native-chat): a Codex turn's first terminal settlement is final Codex follows a turn-ending `error` with a failed `turn/completed` for the same turn. The error settled the turn and dropped its start and attributed send; the completion then re-settled it from nothing, so the record lost its start time and flipped from completed/failure to interrupted/failure. A failed turn lost its "Worked for" and read as interrupted. Both ends now go through one settlement that refuses a turn already in the recent-turns window, so every later end (error then completion, a duplicate completion) adds nothing and overwrites nothing. * test(native-chat): the settles-once fixture is an ordinary failed turn The fixture was labelled and worded as a failed compaction; this change covers the ordinary-turn path, so its frames now read as one. * test(native-chat): a failed Codex turn keeps its record when its two ends coalesce in the queue |
||
|
|
d691486b3f |
fix(codex): a message whose turn was stopped before Codex took it is withdrawn, not stuck (#23618)
* fix(native-chat): land a late settlement from a streamed turn's end after that turn's rows A settlement that says a streamed turn ended waits for the session's event sink to drain before writing its dispatch row. The journal reducer still refuses to overwrite an accepted or rejected send. * fix(codex): settle a send from the end of the turn Codex answered it into The turn/start answer names the turn that holds a send. The send's echo entry now keeps that binding, in memory only. If the bound turn is interrupted without echoing the send, the send is withdrawn: Codex clears a turn's pending input on interrupt, so the model never saw it. If the turn fails first, the send is rejected in Codex's words. A completed turn settles nothing, since Codex records pending input when it finishes and the echo is still due. An answer read after its turn already ended is settled by that end. The echo is still the acceptance and carries the item key. * test(codex): a send settles from the end of the turn Codex answered it into The fake Codex keeps 0.157's turn bookkeeping, and can deliver the turn/start answer after turn/started or after turn/completed. The tests cover: - a Stop before any echo withdraws the send, and the working rule reads idle; - a steered follow-up is withdrawn when the turn is interrupted; - a failed turn rejects the send in Codex's words; - a completed turn leaves the send to its echo; - a normal echo and a late echo; - two steered sends in one turn; - an answer read after the turn ended; - a timed-out answer; - child-thread turns; - how a binding dies. * refactor(native-chat): drop the stream flush before a late turn-end settlement Nothing reads the order of a dispatch row against the turn's terminal row: the reducer keeps a settled send terminal and the working state is derived from both. The echo acceptance on the same path never waited either, and the wait could drop the settlement on a failed sink barrier. * fix(codex): settle a failed turn's sends at its end, not at its error Codex keeps a failed turn's pending input and records it after the error frame, before turn/completed. Settling at the error rejected a steered follow-up the model had in fact received, so a Retry would send it twice. * docs(codex): say a completed turn echoes what it took before it ends Codex records a completed turn's pending input before `turn/completed`, so a bound send that turn never echoed is left for recovery, not awaiting an echo. The comments and one test title said the echo was still due. * test(codex): settle a send whose answer is read after its turn failed or completed A failed turn that ended before the answer rejects the send in Codex's words, once; a completed one leaves it admitted and still armed for its echo. * refactor(codex): read a failed turn's reason with the typed thread-fact reader |
||
|
|
a68d62911e |
fix(native-chat): the host writes chat failures for a person, with a typed fact beside them (#23116)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * fix(native-chat): name a chat write by its target, not the owner generation A write carried the fence of the last frame the pane read, and the host refused it unless that fence was still current. An idle release and the restart after it each move the fence, and the release publishes nothing, so a send after a release was refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a cold start was refused as stale. Every write already names what it acts on: a send its conversation, a cancel its turn, a prompt answer its item revision, a rewind its epoch; an option is last-writer-wins. So admission stops comparing the client's fence, and the rebase that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it. The writer-lease check stays, and so does the attach's compare-and-swap. Frames now stamp the fence read when each frame is sent instead of a copy each subscriber kept, which went stale on the same release. * fix(native-chat): every journal append reaches the chats that are open A journal write and its delivery to open readers were two calls, and some writers made only the first. A failed start whose lease could not be handed back, a provider revision with no frame behind it, and eviction's settlement were all journaled without reaching an open chat. A journal handle now reports every durable change, and the host's session map binds that report to the session's readers when the handle is set. Writers no longer publish what they append; the per-writer publish calls are deleted. * test(native-chat): an epoch replacement reaches the open chat * test(native-chat): each row reaches an open chat once, and a live handle enters only through the map * test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite The seeded record had no surface tab id, so the next open backfilled one and that rewrite alone made the no-op transaction write. The test passed with the legacy-lease rewrite signal removed. * test(worktree-activation): restore the OMP surfaced-agent resume test The handoff removal deleted it alongside the terminal-owner tests, but it covers the surfaced-PTY block that still guards resume, including an agent whose ownership is unknown. * perf(native-chat): a publish behind a delivered commit reads nothing Each commit now delivers itself, so the publish a provider frame still sends afterwards found every reader caught up but still read rows and rebuilt the timeline for each one. A caught-up reader now skips the read. * test(native-chat): state why the teardown test's fake journal is safe to cast * docs(native-chat): say mutation admission checks only the writer lease * docs(native-chat): drop the send rebase from comments that still described it * fix(native-chat): a message is accepted, then delivered A send to a chat with no running agent restarted the agent inside the send call, before the message was recorded, so the client waited for the whole start and a failed restart refused the message. Claude held prompts sent during startup, and those could settle as "unconfirmed". A send is now accepted inside the session's serialized queue: one ledger row and one submission row marked handoverRecorded, published, answered pending. A per-session delivery loop exists while a message is queued. It starts the agent through the same serialized attach a hold uses, waits outside the queue for a Claude child to prove its start, and hands the oldest queued message over as its own serialized step, writing dispatch{pending} before the adapter call. A start it needed and did not get writes one error-tone row and rejects every queued message with the same words; a start Stop cancelled writes none. Settlement follows from the rows. A queued message is provably unwritten, so a close, an eviction or an exit rejects it. A handed-over message stays in doubt. A queued row at or below the sequence a handle found when it opened was left by an earlier process and is rejected at open, with no latch. Stop withdraws queued messages with no writer lease and no fence. An attach failure keeps the conversation open, and the attach adopts its journal. Owed work counts the loop and queued rows. A compaction or rewind found prepared when a conversation opens was started under a child this process no longer has, so the open settles it rather than leaving it to refuse every send until a view attaches. The open cursor is scoped to its epoch, because sequences restart when an epoch is replaced. Deleted: restart-before-admission, recordFailedRestart, the fence rebase, Claude's startup gate, the attach's forget on failure and its own crash boundary. Clients without agent-session.accepted-send.v1 get their reply held until the handover; the desktop and paired desktop lists advertise it. * fix(native-chat): settle queued messages only for the child that ended A child that proved its start and then exited before its message was handed over left the message queued: the exit settlement returned early when nothing else was in flight. Delivery then started another child for it, and a child that died the same way started another, without end and without a row. A retried settlement for an earlier generation, run by the attach that delivery started, did the opposite: with that generation's turn unfinished it rejected the message queued for the child being attached. The settlement now takes the rejection for queued messages from its caller. The unexpected exit and the eviction pass one, and it applies even with no other work in flight; the retry for an earlier generation passes none. * fix(native-chat): an adoption that fails to import keeps the conversation open The attach now writes into the conversation's own open journal, but a failed transcript import still closed it as if it were the attach's provisional one. The conversation stayed indexed with a closed journal, so every later send answered "could not be recorded" and every attach failed again until the app restarted. The import now closes only a journal the attach opened for itself. * perf(native-chat): the recovering open reads the journal once Every conversation open now goes through the recovering open, including the read restore of every chat at startup, which used to replay its journal once. The recovering open replayed it twice: once to probe it and again inside the open. The probe is now handed to the open as its load. * fix(native-chat): an attach that fails after indexing its child leaves no child behind A failed attach now keeps the conversation open, but a failure after `onAttached` indexed the child (the rewind or compaction recovery, or the attach's own success record) left that entry claiming a child the failure path had already released. The next send found the phantom, skipped the start, and wrote at a fence the journal had moved past, so the message stayed queued for good. The entry now drops the released child and its event sink, and follows the record's fence, as a failure before indexing already did. * fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer The error strip for a message the host accepted and then did not deliver matched the entry before the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not send your message" with nothing to retry. It now reads the reconciled entry. A rejection the journal records before the send's own pending answer lands is final as well: that answer no longer puts the entry back to dispatching with no Retry. * fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down The preamble waits for its submission to be delivered while the worker's agent starts. When that wait ran out it threw operation_unknown, and the failed-start teardown then closed the session, which rejected the very preamble the host was about to deliver. It now reports a turn start nobody observed yet: the worker is start-unknown with its session kept, the host delivers the preamble when the agent starts, and the worker's report settles the dispatch as for any unobserved start. The receipt no longer suggests reading a screen a structured worker lacks. * fix(native-chat): a message rejected while its chat was closed reads as not sent A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked every later message behind a Retry and no reason, and the delivery probe, seeing the journal already answered, never ran. The reconcile now settles it as rejected like a dispatching one. * test(orchestration): name why the readiness settlement fakes are cast * fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent * docs(native-chat): drop the fence from the admission the send effects run behind * docs(native-chat): give the fence move on release the reason that still holds * docs(native-chat): stop citing a write fence check in launch and mailbox comments Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease. * refactor(native-chat): the provider child is its own record A conversation now outlives any number of provider children, so the child is one record on the conversation's entry instead of five loose fields beside its journal. It is written in one place: indexed only once an attach has fully succeeded, and ended through one function that an exit, a failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence. - A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence patch after it are gone. - Conversation writes read the record's fence, the way mutation admission already does; a child's own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of the conversation's fence, are gone. - The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer dropped when an attach replaced the whole entry. - Stop on a child still proving its start stops only the child: its lease goes back and the chat is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus the conversation's close. - The settlement retry uses the conversation's own journal, opened through the host's one open. * fix(native-chat): the delivery loop alone settles a message its start or child failed A queued message was settled by whichever path happened to end the child first: the loop, the unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that rejected every pending row. That gave two failure rows with different tones for one start, a loop that could hand over to a different child than the one it waited on, and a Claude start that died while starting reading unlike every other failed start. - The loop remembers the child it waited on. At handover, if that child is gone or replaced, it reads how it ended: a Stop continues; anything else writes one failure row and rejects every queued message with the same words, then stops. A child still starting whose start the adapter says did not land fails the same way. The exit, eviction and the settlement retry only settle the handed-over and legacy rows of the child that ended. - One failure row, always an error, keyed by the start. A start a view began that dies with nothing queued writes the same row through the same builder, so a second report revises it. - The open no longer rejects leftovers; the loop's first step does, and the open wakes it. - `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the failure before the exit is processed. - Quit closes every conversation the way closing a chat does: what is still queued is rejected as closed, with or without a child, and a start the loop already has in flight is waited for so the child it produces is stopped rather than left behind. * refactor(native-chat): a stopped child ends on the one reading of its stop The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that verdict to the child's ending, so the host never forms a second view of whether the root is gone. Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition, and a failed re-attach passes what its release saw. The end-of-child record can therefore also carry a stop whose root was not seen to go, which nothing ends on yet. * feat(native-chat): the host says it accepts a send before any agent has it The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same string capable clients already send. A client can then tell a host that answers a send at acceptance, and admits a Stop with no writer before a turn starts, from an older one that still restarts the agent inside the send. Additive: an older client ignores a capability it does not know. * refactor(native-chat): an attach never opens a journal of its own The attach adopts the conversation's open journal, which outlives it, so it no longer opens one for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag that told the two cases apart is gone. Tests that attach without a host open the conversation the way a host does. * fix(native-chat): a moved fence resends nothing on a host that accepts first The outbox treated any fence change as a new owner: it dropped the answer of a send in flight, queued that send to go out again under the same id, and unblocked a refused head. On an older host that is how a send the restart refused, unrecorded, gets another try. On a host that records every send before it starts an agent, a fence moves because that start ran, so the same rule resent into every failed start. With a fence stamped on every frame, that became a loop. The outbox now reacts to a fence change only when the host has not advertised that it accepts a send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed start reaches the client as a rejected message it keeps with its Retry. Against an older host, or before one has answered, the outbox behaves as it did. Desktop and paired web share this hook. * refactor(native-chat): a child's end says whether the user or the host stopped it The end-of-child record's cause now tells a user's Stop from the host stopping the child for a cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a user's Stop, as before, and fails the start it was waiting on after a host stop, with the one error row and every queued message rejected, in the stop's reason when it gave one. The reason stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet. * fix(native-chat): a chat whose only work is a queued message is not offered for resume A message accepted while the agent was starting counts as working in the chat, and quit rejects it as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a chat whose agent never had the message. The snapshot now reads only what was handed over. * test(native-chat): type the queued-message fixtures in the resume-offer tests * fix(native-chat): a start that dies while a message waits on it is that message's failed start Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When that start died, its exit wrote the start's error row and left the message queued, so the delivery loop started a second agent into the same failure and wrote a second row. A child's end now records where the conversation's journal stood, and the loop settles a message accepted before a failed start ended with that start: one row, under its key, and no second start. A message sent after the failure still gets a fresh start. * docs(native-chat): say what an attach's open conversation and unconfirmed ids are now * test(native-chat): pin what a failed start settles, and what a resume offer names A view's child that dies while a sent message waits settles that message only when it died starting and no child has taken its place: a proven child's crash, or a second start since, gets the message delivered. The resume offer names the handed-over message, never a newer one still queued. * test(native-chat): the failed-start pins fail on what the message became, not on a timeout * test(orchestration): the preamble's host stub is typed, not cast The preamble send now takes only what it reads of the host, the send, the settlement wait and the record's fence, so its test builds that host with real types instead of `as never`. * feat(native-chat): a typed failure fact beside every failure sentence Adds the shared vocabulary the host writes a failure with: a closed failure kind, a provider diagnostic that says who it is for (a person, or a log), and a refusal cause beside the refusal code. Status rows gain an optional failure fact and rejected submissions an optional rejection fact; the dispatch row carries it, the reducer reads it field by field, and the projection forwards it. Older rows and older readers are untouched: every field is optional and the schemas stay open. * fix(native-chat): durable failure rows and rejection reasons are written for a person Every host writer that records a failure now writes a sentence for a person beside a typed fact, instead of embedding a refusal's message, an exception or a composed exit string. A provider's own words travel as a separate diagnostic from the places Orca composes them - the Claude and Codex exit stderr (a log), Codex's JSON-RPC message, Claude's compact_error and Codex's turn error (for a person) - and are never inferred from a string afterwards. Not signed in and oversized history are typed at the adapter that detects them. Covers start and restart failures, the delivery loop, dispatch rejections (content, queue-full, write failures, provider refusals), cancel and answer confirmation rows, compaction, the rewind fallback, and not_delivered, which released clients printed as it was. Two leaks close on the way: a settlement retry no longer writes Orca's probe evidence into the exit row, and an attach or journal-sink failure is recorded as Orca's fault rather than as the provider stopping. The legacy rejection markers and the reasons on sends in doubt stay byte-identical. * feat(native-chat): refusals name their cause, and a failed start is worded in one place A refusal now carries an optional cause beside its code: one closed enum of the situations a chat write can meet, set at every emitter a structured-chat write reaches. Returned refusals build it with refuse(code, cause, message). Store and host paths that raised a bare Error(code) now throw AgentSessionRefusalError, whose message is still the code and which has no code property; the RPC error mapper handles it before any other passthrough, keeps today's wire code and message byte-identical, and adds { refusal: { code, cause } } to the error's data. The hold throws it, and restart-resume files the cause beside the unchanged reason. The operation ledger stores the cause beside the code, so a replay names the same situation as the first answer. The store fallback copy picks its words and cause by situation, so a stale replay or a moved lease no longer reads as a latched owner. Every failed start is worded by structuredAgentSessionStartFailure(cause, context), which returns the row sentence and the typed fact together; the delivery loop, the exit settlement and the dispatch that met a starting child all call it. Provider diagnostics are capped at the lease record's 512 characters wherever a fact is built. * refactor(native-chat): one reader of why a submission was rejected classifyDispatchRejection(submission) returns { category, verdict, kind? }. It reads the typed rejection fact when the row carries one this build can place, and the legacy markers otherwise - all six, including not_delivered, which released clients printed as it was. The verdict is null only for a withdrawal, a host restart and a closed chat; write failures, a full queue and not-delivered stay failures. It replaces dispatchRejectionReasonIsInternal and every string comparison against the markers: the outbox reconcile, the rejection notice, and the send disposition, where a replayed Stop-withdrawn send no longer surfaces as a failed send. The journal reducer's echo-aliasing guard reads the narrow isWriteFailureSubmission, which matches the typed kind or the legacy prefix in any dispatch state, so its behaviour on legacy unknown rows is unchanged. * test(native-chat): pin provider diagnostics where they are composed The Claude exit status and stderr, the Codex stderr tail and Codex's JSON-RPC message are each checked at the place Orca composes its own error around them, so the typed detail is proven to come from the provider's value and never from Orca's wording. * fix(native-chat): an attachment Orca refuses says which limit it broke The content check's refusals (20 images, 5 MB per image, 20 MB in total, supported types) are written for a person, but the rejection writer replaced them all with one generic sentence. Each refusal now carries its own sentence, in MB rather than bytes, and the writer records it beside kind attachmentInvalid. Only an attachment that could not be read keeps the generic sentence. * fix(native-chat): the chat tab table refuses with a typed cause Showing a chat tab refused with bare Error('agent_session_conflict') and Error('agent_session_identity_required'), the only chat-reachable refusals still thrown without a cause (opening a chat from history can reach the second when the chat is removed mid-open). Both now throw the typed refusal; wire code and message are unchanged. * fix(native-chat): a compaction Codex refuses up front keeps Codex's words When Codex refused thread/compact/start, the adapter passed on only Orca's wrapped error text and dropped Codex's own message, so the chat's row read just "Compaction failed." The refusal now carries Codex's message as the failure detail, as a compaction that fails later already did. * fix(native-chat): an unreadable chat record no longer promises an update fixes it The recordUnreadable copy said a newer Orca saved the chat, but the store marks a record unreadable for damage and key mismatches too, where updating does nothing. The sentence now says Orca can't read it and gives both next steps. * fix(native-chat): a restart or a close leaves released clients a sentence, not a marker host_restarted_before_delivery and provider_closed_before_delivery are not in the markers released desktop and mobile builds hide, so they printed raw on every rejected message a restart or a chat close left. Neither marker has shipped. New rows carry a sentence plus kind hostRestarted or chatClosed, as not_delivered already did; the classifier still reads both markers, and the verdict for both stays no-failure. * fix(native-chat): log why an attachment could not be read The rejected message now says only that the attachment couldn't be read, so the error that said why (a missing file, a permission, or an unexpected throw) went nowhere. It is logged instead. The content rejection moves beside the content check that owns its errors. * refactor(native-chat): drop an unused thrown-refusal cause reader It had no callers, and its comment claimed it looked through wrappers, which it did not. * fix(native-chat): a compaction the provider never confirmed is recorded as unconfirmed, not failed * fix(native-chat): an empty or non-user message is not recorded as a bad attachment * fix(native-chat): an undelivered preamble's error ends in one period * refactor(native-chat): a compaction ends as compacted, failed, or unconfirmed, never an unlabelled error * refactor(native-chat): a failure's sentence is written only from its fact A writer could choose its sentence and its kind separately, so five writers put hand-written words beside a fact that said something else. One shared constructor, agentSessionFailureWords(fact, { surface, agentName }), now makes both, and the journal types refuse anything else: a status row, a rejected message or a conversation command that carries a fact must carry the sentence that constructor branded. Persisted rows keep their shape. - The sentence table and the restart table move to src/shared, as the English default a client copy table can reuse. - A rejection's legacy markers come from the same constructor. A write failure is now the bare `provider_write_failed` marker, which released clients already hide; its error goes to the log. - An image Orca refuses carries which check it failed (and the limit) in the fact instead of a sentence; an empty message gets its own kind, and a non-user message is Orca's fault. - The exit row, the interrupted compaction, the /clear failure and the rewind placeholder no longer carry their own words beside a fact: the exit row says what its fact says, and the placeholder carries no fact. * fix(native-chat): a start that failed without an observed exit no longer blames the provider Every untyped start error was recorded as "The provider stopped before it finished starting.", so an Orca fault, a failed spawn or a close that ended a start was blamed on the provider. Only an exit the adapter observed says so now: the Claude adapter marks the error it saw the child exit with, and anything else is a new `startFailed` kind, "<Agent> couldn't start.", keeping the provider's diagnostic when the error carried one. A child gone with no end observed, and a /clear whose new conversation was refused, read the same way. * fix(native-chat): a view never restarts a chat whose last start failed A Claude chat whose CLI exits during startup left one red row per start, and every time a view bound to it (the chat opening right after its create died, or the user switching back to it) the hold started the CLI again, so the same launch-failure row repeated. Only a send retries a failed start now, the same rule provider-exit recovery already applied; the rule lives in one predicate the hold, exit recovery and the delivery loop share. * fix(native-chat): a chat a terminal agent still holds says to quit that agent The merged base words a restart a terminal agent's claim refused with the refusal's own message, which names the process. That message is Orca's text and never reaches a durable row here, so the row read only "<Agent> couldn't restart." and lost the one step that frees the chat. The restart sentence now derives it from the refusal's cause: a claimConflicted refusal adds "This chat is still open in a terminal agent. Quit that agent to continue the chat here." The live refusal still names the process. The restart-resume ledger test now sets up a claim the base still refuses: a terminal owner that is proven running. * refactor(native-chat): keep the changed files inside their lint limits The legacy-marker lookup is a table, not a non-exhaustive switch; the preamble tests read the error without a cast; and the Codex history refusal goes through a named constructor so its file stays under the line limit. * test(native-chat): start the child the loop waits on with an attach, not a second view A view no longer starts a child whose last start failed, so the R2 case that waits on a child started since the failure now gets that child from a client attach, the one non-send starter left. * fix(native-chat): settle a gone generation's turn wherever a conversation opens A send that opens a chat this process had not read yet (after a crash, from a phone or the CLI) went through the delivery open, which never settled what the dead generation left running; only the read restore and a successful acquire did. When the send's start then failed, the turn stayed running for every reader. The settlement now runs in the one journal open, at the crash boundary, for every opener except an acquisition, which settles from the evidence it read before its reserve; the read restore's separate step is gone. * test(native-chat): prove the next child's start settles the turn an earlier child left The R1 case lost its only settlement assertion when the latch it checked was deleted. It now seeds the running turn the earlier child left and asserts it ends at the exit's receipt, with the exit's row, before the message is handed to the new child. * test(native-chat): count a failed start's rows by row, not by text Comparing the set of texts passed when two different rows carried the same words, which is the duplicate the test exists to catch. * test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget * refactor(native-chat): refusals carry details keyed by their code A refusal named its situation with one flat cause list shared by every code, so nothing stopped a site pairing a code with a situation that code never means, and the loose fields released clients read (fence, revision, resolution, verdict, rewind reason) were written by hand at each emitter. A refusal is now one variant per code with optional details: a reason that code lists plus that code's own facts. refuse(code, details, message) rejects a reason the code does not list at compile time, and it is the one place the loose top-level fields are copied from details, so released clients read exactly what they read before. A site that cannot name its situation uses refuseUnclassified, which carries facts but no reason, the same as an older host; there is no catch-all reason. Thrown refusals put { refusal: { code, details } } in the RPC error's data (wire code and message unchanged). The operation ledger, the restart-resume record and a restart failure's embedded refusal keep details beside the code and read them back against it; a row an unreleased build wrote with a cause parses and reads as naming none. * test(native-chat): pin the open's and the send's start and row counts, however the view binds Opening a fresh chat whose starts fail makes one start and one row, with two views bound before or after the create's child died; one send makes one more of each. * test(native-chat): a restart-resume failure keeps its refusal details The recovery capsule reads a failure's details back against the refusal code in its reason: facts the code does not list and a reason another code owns are dropped, and a record an unreleased build wrote with a cause still parses, naming nothing. * test(native-chat): import the failure words once in the provider child test The merge left two imports of the same modules. * fix(native-chat): settle a gone generation's turn at every open but an acquisition's The journal open skipped the settlement whenever the lease read reserved or live, to leave an acquisition's own open to the acquisition. But a lease a crashed process left in recovery also reads live, until the next acquire resolves it. A send that opened such a chat, from a phone or the CLI after a crash on a host that could not prove the old owner gone, skipped the settlement; when its start then failed, the dead turn stayed running for every reader. The acquisition now says it is the opener, and every other open settles, whatever the lease still claims. * test(native-chat): hold the create's start open until the views bind The "view binds while the create is still starting" case gave the create a 300 ms head start and asserted the views bound before it died. On a loaded runner the holds took longer, the create's exit landed first, and the case failed its own precondition. The create's initialize now waits on a gate the test releases once the views are bound. * fix(native-chat): a start the provider refused or Orca broke no longer says the provider stopped A failed start whose cleanup proved the child gone was typed as `providerStartFailed` whatever failed it: Codex refusing to resume a thread, a timeout, or Orca's own store fault. The chat then read "The provider stopped before it finished starting.", which was untrue, and the provider's own words were dropped. A refused restart also inferred the same from an `exited` verdict, which only says nothing runs now. Only an exit the adapter observed names that situation now; anything else is refused with no reason, keeping its verdict, and the chat reads "<Agent> couldn't restart.". The provider's words travel host-side from where the acquisition failed into the start-failure fact's detail, never onto the refusal, and the sentence does not quote them. What failed is logged once where the start failed. * fix(native-chat): a refused /clear start keeps the situation it named The replacement start that /clear makes built its own start-failure fact, so a refusal that named its situation, such as not being signed in or a history too large to restore, was recorded as a bare "couldn't start". It now takes its fact from the same start-failure function as every other start, as a new session that failed to start. * fix(native-chat): the journal schema and comments describe a refusal's details, not its cause The persisted failure fact's schema still described `refusal.cause`, which this branch replaced with `details`. It now describes `details` as an optional open object; a row an earlier build wrote with a `cause` still parses. A comment and three test descriptions that still named the cause now name the details. * fix(native-chat): a Claude child that exits while being acquired still reads as the provider stopping Now that only an exit the adapter observed says the provider stopped, the exit the Claude adapter saw during acquisition has to be marked where it is seen, as the exit after acquisition already is. Without the mark, a Claude CLI that exited at spawn read "Claude couldn't restart." instead of "The provider stopped before it finished starting." * fix(native-chat): word a failed chat start's refusal as its start failure A chat whose agent failed to start answered the create with the raw error: the launch strip read "Chat could not be started. claude stream-json exited (code 1): claude: not signed in", and the ledger replayed the same text. The first answer and the replay now carry the sentence the chat's start-failure row reads as ("The provider stopped before it finished starting.", "Claude couldn't start.", the not-signed-in and history-too-large sentences), and the raw error goes to the log. A store refusal's code and the unproven-exit marker are unchanged. * fix(native-chat): show Claude's API retries as one sentence row While Claude retried a refused request (a 429, say), the chat gained one red row per attempt reading "rate_limit", with the raw retry frame behind Details. Each retry run now writes one warning row that later attempts revise in place: "Claude is rate-limited and retrying." for a rate limit (error `rate_limit` or status 429), and "Claude hit a temporary problem and is retrying." otherwise. The row carries a `providerRetrying` fact with the provider's error type and status, and the frame as a log detail capped at 512 characters. * fix(native-chat): tell the user to run /clear again when its new conversation can't start When /clear's replacement conversation failed to start, the result told the user to "send your message again", which would go into the old conversation. The failure words now take the command the start was for, so a failed /clear reads "Codex is not signed in for the selected account. Sign in, then run /clear again.", "Codex couldn't start. Run /clear again." or "The provider stopped before it finished starting. Run /clear again." A message send keeps its wording. * test(native-chat): expect a failed Claude create to be refused in a sentence The runtime suites asserted the CLI's stderr reached the create refusal; it now goes to the log and the refusal reads as the chat's start failure. * fix(native-chat): name the agent that stopped starting and say how to retry a failed start A start the provider ended now reads "Claude stopped before it finished starting." (or "The agent ..." when the chat's agent is unknown) instead of naming "the provider". A start or restart that failed with a chat left to retry now ends in "Send your message to try again.", or "Run /clear again." for /clear; released clients print only this sentence. A message rejected at dispatch because its child died while starting names the same agent as the start's row. * refactor(native-chat): drop the composer's second error formatter After the merge with main, every chat write in the composer path reports its failure as a typed outcome worded by the refusal-notice table, so the send's catch sees only a local throw. The {code, message} formatter this branch added for it has no payload left to format, and its claim to be the one way a chat words a failure is no longer true. The composer send is main's again. * test(native-chat): pin the reason on a message rejected while its chat was closed The reopen test checked only that the message reads as not sent; it now also checks the Retry row carries the host's reason. * fix(native-chat): answer a create refused before spawn in the words its replay reads A create that failed before any process started threw Orca's own error text as its first answer, while its replay from the operation ledger read the generic start sentence. The two refusals a person can act on, a launch whose Anthropic sign-in variables override the managed Claude account and a Claude account switch in progress, are now typed where they are thrown and worded by the shared failure constructor, so the first answer and the replay say the same thing. Every other pre-spawn failure reads the generic start sentence, with its own text in the log. The first answer keeps its wire code; a message that is itself a code is unchanged. * fix(native-chat): say how to retry after an agent stopped before it finished starting "<Agent> stopped before it finished starting." gave no next step outside /clear, unlike every other failed start. It now ends "Send your message to try again.", the same step a start or restart that could not run gives; after /clear it still says "Run /clear again." * fix(native-chat): say how to reach a Claude chat when a WSL Claude account blocks it A Claude chat that restarts while a Claude account is added in WSL and no Windows Claude account is selected was refused before spawn with Orca's own text as its first answer, and the generic start sentence on replay. The refusal is now typed where the account gate throws, and both answers read the same sentence: choose or add a Windows Claude account in Claude Accounts settings, then send the message again. Account settings that cannot be read name no situation and keep the generic sentence. * fix(native-chat): say the reason a host names for a refused chat write A refused Stop, answer, setting, goal, command or queued message now reads the refusal's reason as well as its code. A code stands for several situations, so the code alone could only say what did not happen; with the reason, the notice says why and, where the person has a step to take, what it is: "The agent is still responding. The command didn't run. Wait for the agent to finish responding, or stop it." The phone uses the same words. The notice table keys on code, then reason, then the kind of write. Every reason of every code has an entry, so a reason the host adds does not compile until it has words; a reason whose honest words are its code's keeps the code's row. A refusal with no reason, or one this build does not know, reads exactly as before, which is what an older host gets. A start that failed reuses the failure row's own sentence rather than a second one. A queued message keeps the reason, the rewind reason and the owner's verdict with its saved failure, and a rejected one keeps the host's typed fact without its provider detail. Nothing that moves with the owner or comes from the provider is saved; entries saved before this load as they were. The saved failure moves to its own module beside the words chosen from it. * fix(native-chat): retry a refused send under a new id only once its agent is proven gone A send refused because Orca could not tell who owns the chat keeps its operation id, since the first attempt may still land. When the refusal also says the agent process has exited, nothing can run that attempt, so the message moves to its Retry row under a new id instead of holding the queue. The owner's verdict is read as a floor: a saved `exited` is final, and any other saved verdict never changes the id on its own. Only a verdict re-derived from the current lease can raise it to `exited`, and nothing lowers a saved `exited`. Today no host sends a verdict on a send refusal, so nothing a person sees changes; the rule is in place for the saved verdict a reload reads back. * fix(native-chat): say a /clear that never finished did not finish, instead of that it cleared the chat A send into a chat whose /clear started but never committed was refused as if the conversation had been cleared: "This conversation has been cleared. Your message was not sent. Open the current conversation to continue." The clear never finished, so its new conversation may not exist and there is nothing to open. That refusal now has its own reason and reads "The last /clear didn't finish. Your message was not sent. Start a new chat to continue.", which is the only way on today. Its message for released clients says the same: "The last /clear didn't finish. Start a new chat to continue." Only a committed /clear still says the conversation was cleared. * fix(native-chat): a send the provider never received after a restart has no verdict Restart reconciliation rejects a crash-stranded send the provider's history proves it never received. Nobody failed that send, but the verdict table treated it as a failure. Each rejection kind now has its verdict in one exhaustive table, so a new kind does not compile until its verdict is chosen; no verdict for a withdrawal, a host restart, a chat close, or this lost send. A rejection whose kind this build cannot place, such as one a newer host added, now reads as undelivered with no verdict instead of falling back to the reason beside it: all it proves is that the message did not happen. The host keeps such a fact's kind when it reads the row back, rather than dropping it and letting the reason decide. Only kinds that can be why a message was not sent may reject one, by type: compaction, cancel/answer confirmation and provider-retry kinds stay on status rows. The one dispatch-row builder takes its input from the type that makes a rejected row carry its fact. * fix(native-chat): a send refused after its agent exited keeps its place in the queue When Orca cannot tell who owns a chat but the refusal says the agent process has exited, the next attempt may use a new id, since nothing can run the old one. It no longer marks the message as rejected: nothing recorded it, so it still holds the head of the queue, and later messages wait behind it instead of being sent ahead of it. * fix(native-chat): stop reading a provider's words from an error that contains itself A cleanup that aggregates errors restarted the depth count for each one, so an aggregate error that contains itself recursed until the host ran out of stack. One depth bound now covers both the cause chain and the aggregated errors. * fix(native-chat): say a chat whose history can't be read can't continue, and to start a new one A read of a chat's history is refused with `agent_session_journal_unreadable` only when the chat's journal file is corrupt or not a database, which no retry can change. The notice table had no words for a read at all, so a pane had nothing to show but the raw code. Reading a chat's history is now its own request kind, `read-history`, and that refusal on it reads "This chat's history couldn't be read, so it can't continue here. Start a new chat to continue.", whether the host names the reason or raises the bare code. A write refused under the same code keeps its words, because its cause is any failed open, which can clear. Any other read refusal says only "This chat's history couldn't be loaded." `agentSessionReadHistoryRefusalParts(code, details)` gives a pane those words from a read error. The notice sentences move to their own module so the table stays within its size limit. * fix(native-chat): decide what every way a child ends means for queued messages in one table What a child's end means for the messages queued behind it was an if-chain: a user's Stop was checked in one place, a host stop in another, and every other cause, including one added later, fell through to "the provider exited". It is now one table over every end cause, so a new cause does not compile until someone says whether it fails what is queued and how. A user's Stop still fails nothing, a host stop is still Orca's fault, and an exit, a failed attach or an eviction still carry the failure the end recorded. Nothing a person sees changes. * test(native-chat): a close that stops the child and then fails rejects what is queued by how the child ended When a chat closes, stops its agent, and then fails a later step, the chat stays open with its messages still queued. The delivery loop then rejects them by the way the child ended: an eviction during startup reads as a failed start, otherwise as the provider having stopped, and either counts as a failure. The cases are rows over the end cause, so another way a chat closes is one more row. * test(native-chat): type the refusal a persisted-schema test admits * fix(native-chat): say whether a chat's history is damaged or just couldn't open A write refused because the chat's journal would not open named one reason, `journalUnreadable`, for every failed open, and a read of the history took that same reason as final. So the words depended on what was asked, not on what happened: a busy or permission-denied open could tell a person to start a new chat, and a damaged one could read as something that clears. The host now decides at the refusal which it was. `journalCorrupt` is set only when SQLite itself reports the journal damaged or not a database (SQLITE_CORRUPT or SQLITE_NOTADB, extended codes included, read from the driver's result code and never from message text, through any `cause` chain). Every other failed open is `journalUnavailable`. A corrupt history reads "This chat's history couldn't be read, so it can't continue here. Start a new chat to continue.", with "Your message was not sent." before the step on a send. One that couldn't open reads "Orca couldn't open this chat's history right now. Try again.", likewise on a send. A host that names no reason gets "Orca couldn't read this chat's saved history.", which promises neither, because damage can't be proven from the code alone. The refusal's message, which released clients print for a send, is now that person sentence instead of the open error's own text; the error is logged instead. `journalUnreadable` is replaced outright: no released build wrote it, and a stored one reads as a refusal with no reason. * docs(native-chat): say why a history that couldn't open names its retry step * fix(native-chat): a failed start's row keeps the words its rejected messages carry When the delivery loop settles a failed start before the child's exit is published, it writes the start's error row and rejects every queued message with the adapter's startup answer. The exit settlement then rewrote the same row from the exit event, so the row could say one thing while the rejected messages, which are terminal, said another. The exit now leaves a start's row it finds already written. * fix(native-chat): a Claude start Orca itself failed no longer says Claude stopped Every error that ended a Claude session was marked as an exit the adapter observed, including a start Orca failed while the CLI was still running: a saved option whose restore lost its answer, an init frame naming another session, or a journal write fault. Those read "Claude stopped before it finished starting." although Claude never stopped on its own. The mark now stays where the child's exit is seen (the connection's exit callback), so those starts read "Claude couldn't start." with any diagnostic beside it, and a real exit before the start lands still says Claude stopped. * fix(native-chat): name the agent that stopped, and blame Orca for its own closes A chat that lost its agent mid-response said "The provider stopped…", and a started Claude session that Orca itself closed after a journal fault said the same, as if Claude had exited on its own. The exit row and the rejected-message reason now name the chat's agent ("Claude stopped while this response was in progress…", "Codex stopped before this message was sent."), or "The agent" when the name is unknown; the stale-state settlement now passes the agent name too. After a Claude start has landed, the ended event reports providerExited only when the child's own exit was observed; any other close is Orca's fault and reads as one. * test(native-chat): pin the sidebar verdict to the rejection classifier for every kind and legacy marker * fix(native-chat): blame Orca, not Codex, when Orca closes the Codex child A Codex chat that Orca itself closed (a journal sink that could not take a frame, or a forced close) said "Codex stopped while this response was in progress", as if Codex had exited on its own. Orca's own close path now reports hostFault. providerExited is left to the app-server connection's exit callback, which the connection withholds while Orca is closing the child, so it only ever reports the child's own exit. * fix(native-chat): a chat whose history is damaged reads "Unable to load this chat." * fix(native-chat): route the conversation-outlives-agent writers through the typed refusals Three writers that arrived with the merge wrote refusals the old way: - An operation that starts the agent itself, such as a goal change, turned any error the start threw into a refusal whose message was Orca's own error text, which released clients print. It now logs the error and says only that the agent couldn't restart. - An option picked while the chat is at rest, for a key the provider would not accept, is refused with the rejected-option reason, like the same pick on a running agent. - A restart continuation whose agent was refused a start filed the rejected message's sentence as the failure's reason. It files the refusal's code with its details again, which is what the restart-failure guidance keys on. * fix(native-chat): a start Orca stopped because it never finished reads as that The idle sweep now stops an agent whose start never finished and rejects the messages waiting on it. The chat read "Orca ran into a problem, so this didn't go through. Try again." for that, because every host stop was worded as Orca's own fault. It now reads "Codex never finished starting, so Orca stopped it." in the chat's row and on each rejected message, carried as its own failure kind so newer clients can tell it apart. The message counts as failed, like any start that did not land. * test(native-chat): pin the merged close and host-stop rows to their typed facts The merge left two expectations on the old words: the close tests looked for the marker a close used to write, and the host-stop test for the host-fault sentence. A close now writes "The chat closed before this message was sent." with its fact, and a host stop the hostStopped sentence the constructor gives, whatever reason the stop carried. Also folds the conversation command's two imports from send preparation into one. * fix(native-chat): a read of a chat this host cannot open says why Reads now reach a chat through one accessor, which refused a missing record and a provider this host does not run as a bare code with nothing beside it. Revealing the same chat already names those reasons, so a client could tell "this chat no longer exists" and "update Orca" apart there but not on the history or subscribe read that follows. The accessor now throws the same typed refusals. The wire code and message are unchanged; the reason rides only in the error's data, which released clients ignore. * test(native-chat): a Claude retrying past the idle window keeps its conversation open Every api_retry frame publishes the journal, and that publish is the activity the idle sweep reads, so a retry run revised into one row still renews the clock on each attempt. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
1b4587956a |
test(native-chat): await the async history and journal snapshot in three tests (#23560)
#22835 made history() and journalSnapshot() async; tests from #23502 and #22944 still call them synchronously, so the typecheck job is red on every PR while main pushes do not run it. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
7a24d3d335 |
fix(native-chat): the conversation outlives its agent (#22835)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * fix(native-chat): name a chat write by its target, not the owner generation A write carried the fence of the last frame the pane read, and the host refused it unless that fence was still current. An idle release and the restart after it each move the fence, and the release publishes nothing, so a send after a release was refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a cold start was refused as stale. Every write already names what it acts on: a send its conversation, a cancel its turn, a prompt answer its item revision, a rewind its epoch; an option is last-writer-wins. So admission stops comparing the client's fence, and the rebase that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it. The writer-lease check stays, and so does the attach's compare-and-swap. Frames now stamp the fence read when each frame is sent instead of a copy each subscriber kept, which went stale on the same release. * fix(native-chat): every journal append reaches the chats that are open A journal write and its delivery to open readers were two calls, and some writers made only the first. A failed start whose lease could not be handed back, a provider revision with no frame behind it, and eviction's settlement were all journaled without reaching an open chat. A journal handle now reports every durable change, and the host's session map binds that report to the session's readers when the handle is set. Writers no longer publish what they append; the per-writer publish calls are deleted. * test(native-chat): an epoch replacement reaches the open chat * test(native-chat): each row reaches an open chat once, and a live handle enters only through the map * test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite The seeded record had no surface tab id, so the next open backfilled one and that rewrite alone made the no-op transaction write. The test passed with the legacy-lease rewrite signal removed. * test(worktree-activation): restore the OMP surfaced-agent resume test The handoff removal deleted it alongside the terminal-owner tests, but it covers the surfaced-PTY block that still guards resume, including an agent whose ownership is unknown. * perf(native-chat): a publish behind a delivered commit reads nothing Each commit now delivers itself, so the publish a provider frame still sends afterwards found every reader caught up but still read rows and rebuilt the timeline for each one. A caught-up reader now skips the read. * test(native-chat): state why the teardown test's fake journal is safe to cast * docs(native-chat): say mutation admission checks only the writer lease * docs(native-chat): drop the send rebase from comments that still described it * fix(native-chat): a message is accepted, then delivered A send to a chat with no running agent restarted the agent inside the send call, before the message was recorded, so the client waited for the whole start and a failed restart refused the message. Claude held prompts sent during startup, and those could settle as "unconfirmed". A send is now accepted inside the session's serialized queue: one ledger row and one submission row marked handoverRecorded, published, answered pending. A per-session delivery loop exists while a message is queued. It starts the agent through the same serialized attach a hold uses, waits outside the queue for a Claude child to prove its start, and hands the oldest queued message over as its own serialized step, writing dispatch{pending} before the adapter call. A start it needed and did not get writes one error-tone row and rejects every queued message with the same words; a start Stop cancelled writes none. Settlement follows from the rows. A queued message is provably unwritten, so a close, an eviction or an exit rejects it. A handed-over message stays in doubt. A queued row at or below the sequence a handle found when it opened was left by an earlier process and is rejected at open, with no latch. Stop withdraws queued messages with no writer lease and no fence. An attach failure keeps the conversation open, and the attach adopts its journal. Owed work counts the loop and queued rows. A compaction or rewind found prepared when a conversation opens was started under a child this process no longer has, so the open settles it rather than leaving it to refuse every send until a view attaches. The open cursor is scoped to its epoch, because sequences restart when an epoch is replaced. Deleted: restart-before-admission, recordFailedRestart, the fence rebase, Claude's startup gate, the attach's forget on failure and its own crash boundary. Clients without agent-session.accepted-send.v1 get their reply held until the handover; the desktop and paired desktop lists advertise it. * fix(native-chat): settle queued messages only for the child that ended A child that proved its start and then exited before its message was handed over left the message queued: the exit settlement returned early when nothing else was in flight. Delivery then started another child for it, and a child that died the same way started another, without end and without a row. A retried settlement for an earlier generation, run by the attach that delivery started, did the opposite: with that generation's turn unfinished it rejected the message queued for the child being attached. The settlement now takes the rejection for queued messages from its caller. The unexpected exit and the eviction pass one, and it applies even with no other work in flight; the retry for an earlier generation passes none. * fix(native-chat): an adoption that fails to import keeps the conversation open The attach now writes into the conversation's own open journal, but a failed transcript import still closed it as if it were the attach's provisional one. The conversation stayed indexed with a closed journal, so every later send answered "could not be recorded" and every attach failed again until the app restarted. The import now closes only a journal the attach opened for itself. * perf(native-chat): the recovering open reads the journal once Every conversation open now goes through the recovering open, including the read restore of every chat at startup, which used to replay its journal once. The recovering open replayed it twice: once to probe it and again inside the open. The probe is now handed to the open as its load. * fix(native-chat): an attach that fails after indexing its child leaves no child behind A failed attach now keeps the conversation open, but a failure after `onAttached` indexed the child (the rewind or compaction recovery, or the attach's own success record) left that entry claiming a child the failure path had already released. The next send found the phantom, skipped the start, and wrote at a fence the journal had moved past, so the message stayed queued for good. The entry now drops the released child and its event sink, and follows the record's fence, as a failure before indexing already did. * fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer The error strip for a message the host accepted and then did not deliver matched the entry before the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not send your message" with nothing to retry. It now reads the reconciled entry. A rejection the journal records before the send's own pending answer lands is final as well: that answer no longer puts the entry back to dispatching with no Retry. * fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down The preamble waits for its submission to be delivered while the worker's agent starts. When that wait ran out it threw operation_unknown, and the failed-start teardown then closed the session, which rejected the very preamble the host was about to deliver. It now reports a turn start nobody observed yet: the worker is start-unknown with its session kept, the host delivers the preamble when the agent starts, and the worker's report settles the dispatch as for any unobserved start. The receipt no longer suggests reading a screen a structured worker lacks. * fix(native-chat): a message rejected while its chat was closed reads as not sent A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked every later message behind a Retry and no reason, and the delivery probe, seeing the journal already answered, never ran. The reconcile now settles it as rejected like a dispatching one. * test(orchestration): name why the readiness settlement fakes are cast * fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent * docs(native-chat): drop the fence from the admission the send effects run behind * docs(native-chat): give the fence move on release the reason that still holds * docs(native-chat): stop citing a write fence check in launch and mailbox comments Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease. * refactor(native-chat): the provider child is its own record A conversation now outlives any number of provider children, so the child is one record on the conversation's entry instead of five loose fields beside its journal. It is written in one place: indexed only once an attach has fully succeeded, and ended through one function that an exit, a failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence. - A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence patch after it are gone. - Conversation writes read the record's fence, the way mutation admission already does; a child's own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of the conversation's fence, are gone. - The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer dropped when an attach replaced the whole entry. - Stop on a child still proving its start stops only the child: its lease goes back and the chat is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus the conversation's close. - The settlement retry uses the conversation's own journal, opened through the host's one open. * fix(native-chat): the delivery loop alone settles a message its start or child failed A queued message was settled by whichever path happened to end the child first: the loop, the unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that rejected every pending row. That gave two failure rows with different tones for one start, a loop that could hand over to a different child than the one it waited on, and a Claude start that died while starting reading unlike every other failed start. - The loop remembers the child it waited on. At handover, if that child is gone or replaced, it reads how it ended: a Stop continues; anything else writes one failure row and rejects every queued message with the same words, then stops. A child still starting whose start the adapter says did not land fails the same way. The exit, eviction and the settlement retry only settle the handed-over and legacy rows of the child that ended. - One failure row, always an error, keyed by the start. A start a view began that dies with nothing queued writes the same row through the same builder, so a second report revises it. - The open no longer rejects leftovers; the loop's first step does, and the open wakes it. - `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the failure before the exit is processed. - Quit closes every conversation the way closing a chat does: what is still queued is rejected as closed, with or without a child, and a start the loop already has in flight is waited for so the child it produces is stopped rather than left behind. * refactor(native-chat): a stopped child ends on the one reading of its stop The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that verdict to the child's ending, so the host never forms a second view of whether the root is gone. Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition, and a failed re-attach passes what its release saw. The end-of-child record can therefore also carry a stop whose root was not seen to go, which nothing ends on yet. * feat(native-chat): the host says it accepts a send before any agent has it The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same string capable clients already send. A client can then tell a host that answers a send at acceptance, and admits a Stop with no writer before a turn starts, from an older one that still restarts the agent inside the send. Additive: an older client ignores a capability it does not know. * refactor(native-chat): an attach never opens a journal of its own The attach adopts the conversation's open journal, which outlives it, so it no longer opens one for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag that told the two cases apart is gone. Tests that attach without a host open the conversation the way a host does. * fix(native-chat): a moved fence resends nothing on a host that accepts first The outbox treated any fence change as a new owner: it dropped the answer of a send in flight, queued that send to go out again under the same id, and unblocked a refused head. On an older host that is how a send the restart refused, unrecorded, gets another try. On a host that records every send before it starts an agent, a fence moves because that start ran, so the same rule resent into every failed start. With a fence stamped on every frame, that became a loop. The outbox now reacts to a fence change only when the host has not advertised that it accepts a send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed start reaches the client as a rejected message it keeps with its Retry. Against an older host, or before one has answered, the outbox behaves as it did. Desktop and paired web share this hook. * refactor(native-chat): a child's end says whether the user or the host stopped it The end-of-child record's cause now tells a user's Stop from the host stopping the child for a cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a user's Stop, as before, and fails the start it was waiting on after a host stop, with the one error row and every queued message rejected, in the stop's reason when it gave one. The reason stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet. * fix(native-chat): a chat whose only work is a queued message is not offered for resume A message accepted while the agent was starting counts as working in the chat, and quit rejects it as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a chat whose agent never had the message. The snapshot now reads only what was handed over. * fix(native-chat): the conversation outlives its agent Opening a chat no longer starts its agent. A conversation is reached through one host accessor that opens its journal at rest, and a send is what starts the agent, through the delivery loop. One idle sweep, every five minutes, stops an agent that has been quiet for thirty minutes and owes no work, then drops an open journal handle that is only a cache. Its record, tab, status row and readers stay. - hold and release are no-ops; hold still builds the host for shipped mobile builds. - The holders, the holds, the release clock and the exit respawn are deleted. - Options, the model list, the goal and the context meter answer at rest; a model pick at rest is recorded as intent for the next start. - Compact, rewind, clear and goal changes start the agent first. A send does too when a rewind is still in doubt after the conversation opens. - Orchestration routes mail and group addresses on ownership (the record plus the chat tab), not on whether the process runs. An open dispatch keeps its worker running. - The restart continuation is a send; Resume all holds each slot until the message is handed over or rejected. - A read error never replaces a loaded transcript, and shows the host's own words. * test(native-chat): type the queued-message fixtures in the resume-offer tests * fix(native-chat): a start that dies while a message waits on it is that message's failed start Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When that start died, its exit wrote the start's error row and left the message queued, so the delivery loop started a second agent into the same failure and wrote a second row. A child's end now records where the conversation's journal stood, and the loop settles a message accepted before a failed start ended with that start: one row, under its key, and no second start. A message sent after the failure still gets a fresh start. * docs(native-chat): say what an attach's open conversation and unconfirmed ids are now * test(native-chat): pin what a failed start settles, and what a resume offer names A view's child that dies while a sent message waits settles that message only when it died starting and no child has taken its place: a proven child's crash, or a second start since, gets the message delivered. The resume offer names the handed-over message, never a newer one still queued. * test(native-chat): the failed-start pins fail on what the message became, not on a timeout * fix(native-chat): a restart offer ends when the chat's agent starts again The offer used to end only when the chat's newest user message changed, because opening a chat started its agent and that start could not be told apart from real activity. Opening a chat starts nothing now, so the host reads the fact it already publishes: a chat's status row goes from not host-owned to host-owned exactly when its agent is started. At that edge the offer and any failure record for the chat are withdrawn, unless the start is a resume action's own (its continuation is the oldest undelivered message). A continuation and a message racing to be first are decided at acceptance: the continuation is refused, quietly and with nothing filed, when any other message was accepted since the restart. A failed continuation start leaves the offer retryable, and each resume action sends its own message id. Deleted: the newest-user-message comparison, its journal reader, the continuation filter, and the failure ledger's own "answered by the chat" check. The marker still carries its message id for one release, so the previous build can read it. * fix(runtime): end a transcript stream when its client unsubscribes Desktop: the IPC subscription controller was dropped as soon as the streaming handler returned, which for most streams is right after it binds. A later runtime:unsubscribe then found nothing to abort, so the host kept the subscriber and derived and sent every publish to a channel no one listened to. The controller now lives until the renderer unsubscribes, resubscribes the same id, or goes away. Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe with the stream's frame id, so the host ends that subscriber and leaves a sibling stream on the same socket running. The direct path now passes the frame id the relay path already passed. * test(orchestration): the preamble's host stub is typed, not cast The preamble send now takes only what it reads of the host, the send, the settlement wait and the record's fence, so its test builds that host with real types instead of `as never`. * fix(native-chat): one fact ends a restart offer: the chat moved on since the restart The offer is live while no other message has been accepted in the chat since the restart and its agent has not proved a start since. The offer list, the resume's reservation check and the continuation's acceptance check all read that one fact, so a message whose start then failed withdraws the offer too, and a stale click finds nothing to act on. The fact is read off the conversation's open handle, which the restart closed, so it is retired durably whenever it may have changed: a message accepted, a start proven. A close and reopen within the same run therefore cannot bring the offer back. A continuation rejected before it reached the agent does not count, so a retry after a failed start still runs. Deleted: the quit-time gate on withdrawal, which changed nothing because the withdrawal and the quit's own offer write share one queue; the per-action "withdrawn" flag and the separate acceptance check it paired with. * test(native-chat): an older build reads the restart offer this build records The offer lives in a file the previous release reads after a downgrade. Pin that against the pinned release's own capsule, and run the lane when the marker or the capsule changes. * fix(native-chat): read a restart offer against where the journal stood when it was taken "Since the restart" was read off the conversation's open handle, which the idle sweep closes: after a reopen, a message the user had already sent looked older than the handle and the withdrawn offer came back. The offer now records the journal position (epoch and sequence) at the moment it is taken, and a message accepted after that position, or a journal on another epoch, means the chat moved on. That is derived from the journal, so it holds across any number of closes and reopens. An older build's offer has no position; only a start withdraws it. Because the message half is now durable, the offer is no longer rewritten in the recovery file on every accepted message; a proven start still writes it, since only the host that saw the start knows of it. * test(native-chat): wait for the listing's retire write before reading the recovery file * fix(native-chat): keep the terminal-backed chat's read error over its local echoes Messages winning over a read error is right for the structured chat, whose read retries and whose messages came from the transcript. The terminal-backed view assembles its list from local echoes too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no error. Only the structured pane now keeps messages over an error. * fix(native-chat): a start retries the exit settlement a failed journal write left owed An agent exit whose journal settlement write failed releases the lease latched until a retry lands. Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried it before the next app launch, and every send was refused. The start the send needs now runs the retry first, where the attach would. * perf(native-chat): answer the owner check without opening the chat Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the answer comes from the session record alone. Reaching it through the accessor opened each resting chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it open for the idle window. It now checks the record and the adapter's support, as before this series, and opens nothing. * fix(native-chat): a read waiting on the session lock opens nothing once quit began The accessor checked for quit before queueing the open, so a read queued behind a session task ran its open after teardown had begun and indexed a journal no teardown step would close. The check now runs at the open itself. * fix(native-chat): read a failed resume's chat before calling it retryable Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the restart, read from its journal. The failure list read it only for a chat already open, so once the idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did nothing. The list now opens the failed chats first, as the offer list does. * test(native-chat): type the provider event sink the settlement test reaches for * fix(native-chat): say the structured read keeps trying only where it does The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an untranslated fallback whenever the read error had no text, and the empty state prefers any message. The view state now leaves the message out, so the structured pane shows that line and the terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an error frame, so it no longer makes the claim. * test(native-chat): await the send's settlement instead of polling for the start The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a loaded machine outran. They now await the host's own settlement of the message. * fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now dated by the resume action. Telling a rejected continuation from the user's own message read the operation ledger, whose rows expire after about a day; after that a failed resume stopped being retryable. The offer now records the continuation each action sends on its own capsule entry, bounded to the newest 16, so the ids end with the offer. The ledger read is deleted. * fix(orchestration): route no mail to a structured worker its orchestration released A structured worker is routed on ownership, and a resting worker's lease is released, so ownership held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one restarted its agent. Routing now also reads the orchestration's own resource row: once it is released, direct mail, group addressing and worker-show's addressable answer drop the worker, as they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored. * fix(native-chat): a failed retry names the user's prompt, not Orca's continuation A resume's continuation is written to the chat before its start, so after a failed attempt the chat's newest user message is that rejected continuation. A second failure then showed Orca's own restart text as the chat's prompt. A retry now keeps the prompt its first failure named. * fix(orchestration): read the released row optionally, as the authority does worker-show's observation called the row lookup directly, which a runtime double without it threw on and failed the structured tab-retirement release. * fix(native-chat): the status bar drops a restart offer the chat moved on from The renderer re-read the host's restart offer only when a failed chat showed activity, so after a message withdrew a pending offer the host answered no chats while the status bar kept counting one, and clicking it opened nothing. The same watch now covers pending offers: a status change in an offered chat asks the host again, once. * test(native-chat): a roster of idle or finished children does not keep an agent awake The sweep reads owed background work through the shared child-work liveness that upstream's release clock adopted; a child that went idle or finished is not work the agent still owes. * fix(orchestration): a task dispatched into a resting structured worker keeps it running The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's process incarnation now counts, derived from the existing rows. * docs(native-chat): comments stop describing the hold this PR removed Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it. Comment-only. * fix(native-chat): a restart offer keeps the start its own continuation made Whose start ended an offer was decided at read time, from whether the offer's continuation was still the queued message. Once the provider refused that continuation, the child it had started read as someone else's start, so the offer ended and its failure showed no Retry. The delivery loop now records which queued message a start is for on the in-memory child, and the child's end carries it; the offer counts a start as its own when that message is one of its continuations. * fix(native-chat): an agent gets a full idle window after its owed work ends The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can read done before the lead's wake-up turn writes anything, and stopping in that gap loses the wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full window afterwards, as the release clock it replaced did. * test(claude): the options-read fixture runs a live child The fixture marked its conversation running with a hasProviderChild field the session type does not have, so the read took the at-rest path and refused a session with no record. It now carries a child, which is what the read checks. * test(native-chat): host tests reach its collaborators through a typed seam The rest-test rig and three test files read the host's private members with Reflect.get and cast the result. The host now exposes one test-only accessor, collaboratorsForTests(), and the subscribers class a subscriberCountForTests() beside its existing retainedActivityCountForTests(), so the tests are checked against the real types and the casts are gone. * refactor(orchestration): one owner answers a structured worker's custody Routing, group addressing, worker-show and the idle sweep each composed their own reading of whether orchestration still holds a structured worker, so each new obligation or retirement state had to be added to every reader. structured-worker-custody now derives both answers from the worker-terminal list state coordinators see in worker-list: addressable is owned and not released, and owed work is an active custody or an unsettled task dispatched to the same incarnation. The owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests. * refactor(orchestration): owed work is an open dispatch on the worker's incarnation A supervised worker's own dispatch context stays open exactly while the worker is active, so the separate active-custody branch only repeated it. Owed work is now one fact, which also states the policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are written once at the top of the module. * fix(native-chat): a restart offer knows its continuations by a tag in their id The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running action's id in memory. Both could disagree with the journal: past the cap an old rejected continuation read as the chat moving on, and a crash during a retry restored the failure's older entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer (its teardown and chat), then the action's own part, so any continuation of this offer, queued or rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own continuation the start was for, read against the stored marker. * test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait The tui-idle probe reads through readTerminal, which now awaits the structured worker check before the PTY read, so the probe's snapshot request starts a microtask later. vi.waitFor missed it on its first check and polled again at 50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot then resolved after the wait had already timed out, so the test passed without judging it, and the rejection landed before any handler was attached. Vitest reported that as an unhandled error and failed the shard. Polling every 1 ms sees the request within a few ms, so the snapshot is judged while the wait is still pending. * fix(native-chat): a view never restarts a chat whose last start failed A Claude chat whose CLI exits during startup left one red row per start, and every time a view bound to it (the chat opening right after its create died, or the user switching back to it) the hold started the CLI again, so the same launch-failure row repeated. Only a send retries a failed start now, the same rule provider-exit recovery already applied; the rule lives in one predicate the hold, exit recovery and the delivery loop share. * fix(native-chat): the idle sweep reads owed work every tick Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it refreshed the clock at most once a window. Work that ended just before the next read left the agent to be stopped at that read, moments after the work ended, which is the gap the refresh was meant to cover. The sweep now reads owed work on every tick for a started agent, so the window always runs from the last tick that saw work owed. * fix(native-chat): a continuation handed to the agent stays sent The offer read its own continuation as not reaching the agent while its dispatch was pending, which also covered one already handed over and still unanswered. When the wait for that answer ended first, the failure it filed read as retryable, and a retry sent a second continuation to an agent that may have acted on the first. Only a continuation still queued, or rejected, is now read as unsent. * test(native-chat): start the child the loop waits on with an attach, not a second view A view no longer starts a child whose last start failed, so the R2 case that waits on a child started since the failure now gets that child from a client attach, the one non-send starter left. * fix(native-chat): settle a gone generation's turn wherever a conversation opens A send that opens a chat this process had not read yet (after a crash, from a phone or the CLI) went through the delivery open, which never settled what the dead generation left running; only the read restore and a successful acquire did. When the send's start then failed, the turn stayed running for every reader. The settlement now runs in the one journal open, at the crash boundary, for every opener except an acquisition, which settles from the evidence it read before its reserve; the read restore's separate step is gone. * test(native-chat): prove the next child's start settles the turn an earlier child left The R1 case lost its only settlement assertion when the latch it checked was deleted. It now seeds the running turn the earlier child left and asserts it ends at the exit's receipt, with the exit's row, before the message is handed to the new child. * test(native-chat): count a failed start's rows by row, not by text Comparing the set of texts passed when two different rows carried the same words, which is the duplicate the test exists to catch. * test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget * test(native-chat): the interrupted create's own retry continues again The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its own operation id, with a fresh start whose result nothing read. That fresh start passes with the released-reservation continuation deleted, so the case the fix exists for went untested. The retry and its assertion are main's again. * docs(native-chat): three comments that still had views starting agents A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction left alone would refuse every send, so no agent would ever start to finish it; and a current host raises the unattached read refusal only once quit began, with the attach window belonging to an older host. * test(native-chat): pin the open's and the send's start and row counts, however the view binds Opening a fresh chat whose starts fail makes one start and one row, with two views bound before or after the create's child died; one send makes one more of each. * test(native-chat): a reader's open settles the turn a failed exit settlement left running An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle sweep closed, and a read that opens the chat before the restart restore reaches it. * test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner A subscription reads the conversation before it returns, so under load the two views took longer than the create child's 300 ms start, which then exited before the test checked that it had not. The child now takes a second to fail. * fix(native-chat): settle a gone generation's turn at every open but an acquisition's The journal open skipped the settlement whenever the lease read reserved or live, to leave an acquisition's own open to the acquisition. But a lease a crashed process left in recovery also reads live, until the next acquire resolves it. A send that opened such a chat, from a phone or the CLI after a crash on a host that could not prove the old owner gone, skipped the settlement; when its start then failed, the dead turn stayed running for every reader. The acquisition now says it is the opener, and every other open settles, whatever the lease still claims. * test(native-chat): hold the create's start open until the views bind The "view binds while the create is still starting" case gave the create a 300 ms head start and asserted the views bound before it died. On a loaded runner the holds took longer, the create's exit landed first, and the case failed its own precondition. The create's initialize now waits on a gate the test releases once the views are bound. * test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the dead process's lease still reads live. The open settles the turn it left running anyway, and the restore that follows finds it settled. * refactor(native-chat): drop the composer's second error formatter After the merge with main, every chat write in the composer path reports its failure as a typed outcome worded by the refusal-notice table, so the send's catch sees only a local throw. The {code, message} formatter this branch added for it has no payload left to format, and its claim to be the one way a chat words a failure is no longer true. The composer send is main's again. * test(native-chat): pin the reason on a message rejected while its chat was closed The reopen test checked only that the message reads as not sent; it now also checks the Retry row carries the host's reason. * docs(native-chat): drop the removed dispatch hold from six comments A worker's session no longer takes a dispatch hold, and no release clock rests a chat by visibility; the agent-launch comments, the abandon test, the teardown test and the refusal census still said so. * test(native-chat): rest the owner-status chat through the idle sweep, not a hold The activation-gate test from #22808 put its chat at rest by holding and releasing it, and passed the release-clock grace. This branch deleted both, so the case threw before it reached its assertions. It now moves the host's clock past the idle window and lets the sweep stop the agent and close the conversation, then asserts the same owner answer and activation gate. * fix(native-chat): show the structured pane's retrying line when a read fails The read transport always hands the pane the host's words, so the error state's "Orca keeps trying to load it" line, which showed only when there were none, was never seen: the pane showed the host's text twice, as its subtitle and on the status line under it. The structured pane now always says its read keeps retrying, and the host's text stays on the status line. The terminal-backed chat is unchanged. * test(native-chat): wait for a send's background start before the refusal oracle removes its store An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals. * fix(native-chat): a start a message waited on gets one failure row, the delivery loop's When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice. The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
5e5f4f6603 |
fix(native-chat): keep a Codex ask's questions in the order it asked them (#23502)
* fix(native-chat): keep a Codex ask's questions in the order it asked them Codex journals every question of one ask in a single write, so the questions share a timestamp. The transcript list sorted rows by timestamp and broke ties by row id, and a question's id ends in the question id the model chose, so answered, cancelled and still-pending rows of one ask came out in the alphabetical order of those ids. The transcript projection now breaks timestamp ties by the order its source gave: the journal's order for structured sessions. The session assembler keeps its id tie-break, since the sources it merges share no order of their own. * refactor(native-chat): name the id tie-break comparator for what it does Two shared comparators differed only in whether they break timestamp ties by id, under near-identical names; spell the id tie-break in the name. * fix(native-chat): order structured chat rows by their journal position The desktop list and the host's conversation outline sorted structured rows by timestamp. The journal's contract is that the sequence orders the timeline and the timestamp is the provider's clock: a Codex ask writes all of its questions in one write (one sequence, one timestamp), and a row recovered after a crash carries an earlier clock at a later sequence. The host now records each item's place within the write that created it, keeps it across revisions like the sequence, and sends it as an optional field. Rows projected from the journal carry that position, and both the desktop list and the outline order journal rows by it. Rows outside the journal keep their rules: rank for the streaming and pending tail, outbox sends after every journal row, and terminal-backed chats keep time then id. * fix(native-chat): keep a refused send at its journal place, and keep list positions off worker reads A send the host journalled before the provider refused it is shown from the outbox, and it sorted after every journal row, so it dropped below whatever the agent wrote after it. It now takes the journal position of the submission the host recorded. Structured worker reads and their archives projected journal rows through the same projection, so they returned the list-only journal position on every message. The worker payload bound now drops it. * test(native-chat): a failed restart's row draws below the message it failed Since a message is accepted before its delivery starts the agent, a restart that fails is journalled after the message, and the refused message keeps that journal place in the chat. Both host paths now pin it through the chat's own projection: a start the host could not make, and a restarted child that exits before proving its start. Folds the journal reducer's batch item write onto fewer lines, which the merge of main pushed past the file's line limit. |
||
|
|
934a2d44a0 |
fix(codex): a native chat's thread opens on the model the chat chose (#23532)
* fix(codex): a native chat's thread opens on the model the chat chose * fix(codex): a resumed thread keeps its own saved model, provider and effort |
||
|
|
17690e6b9a |
style: settle oxfmt 0.70 drift and stop formatting vendored licences (#23377)
The oxfmt 0.65 -> 0.70 bump landed without a repo-wide reformat, so 36 files already in the tree no longer matched what the new version emits. Anyone running `pnpm format` picked all of them up alongside their own change. Also excludes `resources/licenses/**`: `oxfmt --write .` was rewriting the vendored PCRE2 licence, turning its `*` redistribution bullets into `-`. Third party licence text has to be reproduced verbatim, so formatting must not touch it. |
||
|
|
85ac14e9c2 |
fix(codex): retain runtime MCP entries without losing revocation (#22426)
* Retain runtime-only MCP entries Adapted from the investigation and proposal by @mmarabel. Co-authored-by: mmarabel <mmarabel@users.noreply.github.com> * fix(codex): respect inline and dotted canonical MCP ownership * fix(codex): retain canonical MCP removal across upgrades * fix(types): include MCP ownership in CLI project * Keep unrelated main test formatting unchanged --------- Co-authored-by: mmarabel <mmarabel@users.noreply.github.com> |
||
|
|
067975bfd1 |
fix(native-chat): every lease latch has a way to die (#22820)
* fix(native-chat): every lease latch has a way to die A failed exit settlement no longer leaves the lease in recovery: the release writes no stage and keeps the exit in its death evidence, and whatever the dead generation left running is settled from that evidence at the next acquire or read restore. The settlement retry flag, its disposition and every branch that read it are gone. A reservation that recorded no process is released at startup and after a failed start, the never-written conflicted status and the processless proof are deleted, recovery resolution always concludes, and Codex records its child's identity at spawn, before the handshake. * test(native-chat): a re-create needs a release proven by death evidence * test(codex): the child's pid is reported before the handshake * test(native-chat): type the crash and exit fixtures without casts * fix(native-chat): wait out a terminal owner an older build recorded, in recovery rather than manual recovery * test(native-chat): a chat mid-turn at quit reopens idle, and an older build reads an unproven release * test(native-chat): explain the baseline store cast * fix(native-chat): a terminal owner's refusal names the process instead of recursing Opening a chat whose terminal owner an older build recorded threw a stack overflow instead of the refusal that names the process to quit. * fix(native-chat): wait out a terminal owner recovery cannot verify instead of releasing it A terminal agent an older build recorded keeps its PTY across an Orca restart, so a probe that cannot answer (a start-time read that fails on a loaded host) is not evidence its transport is gone. Releasing it let a native child resume the same conversation beside the live terminal agent. Only proof of its exit now ends the claim. * ci(cross-version): run the unproven-release downgrade test The sharded unit job excludes tests/e2e/cross-version-wire, and the cross-version job runs an explicit list that did not name the new test, so it never ran in CI. A change to the record validator now also starts the job. * refactor(native-chat): map the retired manual-recovery stage to recovering at decode Nothing in this build writes manual-recovery, and restart reconciliation already rewrites it. Mapping it where the other retired handoff stages are mapped removes it from the in-memory lease type and deletes the branches that could only see it: the acquisition refusal, the renewer skip, the unproven-release stage check, and the handoff-status 'manual recovery is required' answer. Older builds accept recovering, so a record written back still loads after a downgrade. * docs(native-chat): say what happens to a live child an ownerless reservation leaves The reaper runs once at store open, while the unreconciled lease still claims the child's token, so it does not stop that child on this launch. The comment claimed it did. * test(native-chat): name the each-case label for its role * fix(native-chat): continue a create retried after recovery released its reservation The client retries a create it never heard back from under the same operation id. Recovery had released that create's reservation, so the retry was refused agent_session_ownership_unknown while its row was pending, and agent_session_operation_expired once the row aged out, and the chat never started. A retry whose lease nothing holds now continues as a fresh reservation at the next fence, which also stops the old reservation's spawn from committing. * test(native-chat): name the refusal a replayed create used to get * fix(native-chat): one quit-the-terminal-agent message for a chat a terminal agent holds A chat held by a terminal agent an older build recorded frees only when that agent exits. Sending said to reopen the chat and opening it said two runtimes claimed it; both now say the chat is open in a terminal agent, name its process, and say to quit it. Error codes are unchanged. * ci: run PR checks on the rebased head * fix(native-chat): name a terminal owner's process only when its start time can tell it from a reused pid * test(native-chat): relaunch from the dying host's durable state, so its still-pending attach cannot race the new host |
||
|
|
add99c908b |
fix(native-chat): send typed question answers as structured answers, not an option id (#22793)
* fix(native-chat): send typed question answers as structured answers, not an option id A typed "Other" answer was packed into the `optionId` of agentSession.respondToQuestion, a field capped at 1024 characters, so a long answer failed with "Invalid option id" and never reached the agent. respondToQuestion now carries per-question `answers` in their own field, bounded like a typed answer, and a host advertises agent-session.question-answers.v1 when it takes them. Clients fall back to the packed option id for older hosts. The host reads either form once into a typed response, records the structured answers on the resolution (and keeps the packed form older clients read), and the Claude and Codex adapters build their reply from the typed answers before the journal commits, so an answer the agent cannot take is refused rather than recorded unanswered. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): hold one answer per single-select question card Typing an answer deselects a picked option, and picking an option leaves the typed text in the field without sending it, so the card never shows two answers while sending one. Multi-select still sends picked options and typed text together. * fix(native-chat): keep keyboard tabbing from re-choosing a typed answer; accept untrimmed question ids Clicking or typing in the answer field chooses the typed answer; focus alone no longer does, so tabbing to Submit keeps the option the user picked. A question id is matched exactly by the host, so the wire no longer rejects agent-written ids with edge spaces, which older builds accepted. * fix(native-chat): choose the typed answer on click so a disabled or scrolled field cannot * test(native-chat): cover pointer events on a disabled answer field * refactor(native-chat): record the typed answer as a choice in the question card Choosing the typed answer is now an entry in the question's selection, set by typing or clicking the field and replaced by picking an option, instead of being inferred from an empty selection. Unpicking an option no longer silently chooses kept text. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
c220d92c03 |
fix(codex): Codex 0.157+ starts in Orca-managed homes instead of failing with SUN_LEN (#22878)
* fix(codex): turn off Codex daemon auto-start in homes whose socket path exceeds sun_path
Codex >= 0.157 auto-starts a background app-server daemon and connects to
<CODEX_HOME>/app-server-control/app-server-control.sock. Orca's managed homes
under userData make that path longer than sun_path (104 bytes on macOS, 108 on
Linux/Windows), so every interactive codex in an Orca terminal failed with
'path must be shorter than SUN_LEN'. The config mirror now writes a marked
[features] daemon_auto_start = false into only those homes, removes it when the
home fits, and never promotes it into ~/.codex.
* fix(codex): address review of the daemon socket guard
- A runtime config.toml holding only Orca's daemon override no longer reads as a
config-sync stall, so users without ~/.codex/config.toml get no false
"missing" warning in the accounts pane.
- The legacy shared-home refresh re-applies the guard, so retained pre-rollout
panes keep daemon auto-start off after a system-default launch.
- Warn once when an inline `features = {...}` or `[[features]]` blocks the
override instead of failing silently.
- Rename the upsert's TUI-specific internals now that it serves any table.
* fix(codex): apply the daemon socket guard even when the settings mirror stalls
When the settings write-back or mirror refused (unreadable baseline, failed
write to ~/.codex, unreadable source), the whole pass returned before the
daemon guard was applied. A home whose config.toml predates the guard then
kept failing with SUN_LEN on every launch for as long as the stall lasted.
The guard now lands on those paths too; the mirror itself is unchanged.
* fix(codex): guard managed account homes when ~/.codex/config.toml is missing
* test(codex): keep reset-credit ownership checks scoped to the retry, not service construction
* test(codex): build the account mirror test without a type cast
* fix(codex): keep blocking WSL ownership checks off the no-config guard pass
Guarding account homes with no ~/.codex/config.toml ran the WSL ownership
check, a synchronous wsl.exe call per account, at startup before the window
opens and on every account switch. WSL homes are guarded by WSL launch prep,
so that pass now covers host homes only.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
|
||
|
|
d443320af2 |
refactor(native-chat): remove the unused terminal handoff (#22783)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite The seeded record had no surface tab id, so the next open backfilled one and that rewrite alone made the no-op transaction write. The test passed with the legacy-lease rewrite signal removed. * test(worktree-activation): restore the OMP surfaced-agent resume test The handoff removal deleted it alongside the terminal-owner tests, but it covers the surfaced-PTY block that still guards resume, including an agent whose ownership is unknown. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
e0144a9eb6 |
fix(native-chat): show the Codex and Claude model picker the moment a chat opens (#22756)
* fix(native-chat): show the Codex and Claude model picker the moment a chat opens A new structured chat showed no model picker until its session had been created, spawned, initialized and had answered a model listing — and the picker then listed the models a second time. Codex's listing often goes to the network, so the picker took 0.6-2 s to appear. - Keep a host-owned model catalog per agent and account home, persisted on success only and refreshed in the background once it ages out. A new read-only agentSession.modelCatalog RPC answers from it without a live session; sessions reuse it instead of listing again. - Render the picker while the launch is still provisional, showing the saved default. A pick made before the session exists is held and applied once it publishes; only the host's acceptance saves it as the default. - Mark the model and effort set in the user's Codex config as the listing's default (config/read), so the first frame names what the chat will run. - Resolve the account a record-less read would use without running launch preparation, which writes and syncs account state. * fix(codex): disable plugins in the model catalog probe app-server * fix(native-chat): read the host model catalog only for panes on this machine * fix(native-chat): name a pre-report model only for a chat this view launched * test(native-chat): pin the launch latch across publish * test(native-chat): pin a held pick reaching the host before the first send * fix(claude): pin the catalog probe's config dir by the session spawn's rule * fix(native-chat): read the host model catalog only for a visible chat * fix(native-chat): rewrite the model catalog file only when a listing changes * test(native-chat): type-check the first-send order fixture * refactor(native-chat): keep the structured options hook under the line cap * fix(native-chat): send the first turn only after every pick held during launch settles * fix(claude): name no default effort from the catalog probe listing * fix(native-chat): name no listed default model for a chat resumed from history * refactor(native-chat): let the launch own picks made before it publishes A pick made while a chat launches had no fence to go to, so the pane held it and flushed it after publish; every other sender (the outbox, the launch prompt) then needed its own gate to wait for that flush. The launch now keeps those picks in its own state, applies them against the create receipt's fence before it counts as published, and every sender follows publish by construction. The pane flush, the outbox gate and the module-wide held-pick registry are gone. The launch also snapshots the saved selection its create seeds when the intent is built, so a pick in another chat no longer relabels one still launching, and a pick the host refuses is reported the way a refused mid-session pick is. * fix(native-chat): name no default model a workspace's own config can replace The catalog's default is the account's, read without a working directory, but a chat runs in its worktree, where a project config (Codex's .codex/config.toml between the project root and the worktree, or a Claude .claude settings file that sets a model) picks the model instead. The picker named the account default there while the chat ran the project's model. A new chat's catalog read now names its worktree. The host checks that workspace for such config (existence only for Codex, the model key for Claude) and, when any is present or the workspace is not a local directory, serves the listing with no default, so the picker names nothing until the chat reports its model. * fix(native-chat): name the listed default model before the report only for Codex * fix(claude): let an option pick made while Claude starts wait for it instead of being refused * fix(codex): name no listed default when the configured model is not in the listing * fix(native-chat): show the picker as unavailable until a published chat attaches * fix(codex): keep the catalog probe's listing when config/read stalls * fix(native-chat): write the pending model catalog save before quit * chore: drop an unrelated lockfile rewrite * fix(native-chat): rename the catalog store's listing parameter off the global fetch name * chore: drop an unrelated lockfile rewrite * fix(native-chat): name the model Claude will run before its first turn * fix(native-chat): keep Claude's pre-turn applied effort out of the saved session options |
||
|
|
eb746a6d32 |
fix(codex): a Codex native chat that never sent a message reopens after restart (#22639)
* fix(codex): start a new thread when a chat's thread was never saved A structured Codex chat records its thread at create time, but Codex writes no rollout until the first input. After a restart, launch resumed that thread, Codex answered "no rollout found for thread id", and the chat could never run again. When the head of the handle chain is the session's own creation and Codex answers that exact error for that exact thread, start a new thread instead. The new link supersedes the unsaved creation in place and names it, so the chain keeps one live identity and does not grow across restarts. A thread a resume, fork or adoption proved is never superseded, and no other resume error starts fresh. * test(codex): build launch-resolution chains without a type assertion * fix(codex): match only Codex's own no-rollout text, pinned through the real connection The fallback matched Orca's own error-wrapper prefix too, and every test built that string itself, so rewording the wrapper would have disabled the fallback with the suite green. Match the method, code -32600 and Codex's exact detail as the message suffix, and drive Codex's raw error frame through the real connection in a test. The link builder now refuses, at the type level, a supersession on an adopted or resumed link, which the chain would reject downstream anyway. * docs(codex): note why the no-rollout text is safe on the resume path |
||
|
|
6ae6ed08bb |
fix(claude): open structured chat without a startup deadline, and make Retry start fresh (#22364)
* fix(claude): open structured chat without a startup deadline, and make Retry start fresh
Publish the Claude session as soon as its process is spawned instead of racing
initialize against a fixed 10s deadline. Prompts sent before startup lands are
held and written in order once it does. An exit or sign-in failure before startup
ends the session with the reason and the CLI's stderr.
A create that failed because the process provably exited now carries
ownerVerdict 'exited', so the client marks the launch failed and Retry mints a
new operation instead of replaying the stored failure.
* fix(native-chat): sending into a chat that failed to start restarts it
* fix(native-chat): a send with no live owner restarts it once
A provider child that timed out or exited hands its lease back, and every
later send was refused agent_session_ownership_unknown. Clients read that
code as "not admitted yet" and resend forever, while only a surface hold
could make a new child, once per mount, with its failure swallowed.
The send now routes to a live owner, otherwise restarts one from the
persisted resume state where resume eligibility allows it (single-flight
per session), otherwise refuses with the new settled
agent_session_owner_unrecoverable. Unverifiable, reserved and handed-off
leases are left alone. The desktop hold now logs its failure.
* test(native-chat): pin the unrecoverable refusal as settled in the outbox
* test(native-chat): pin the release clock after a send restarts an unheld owner
* test: read the sent operation id without a cast
* fix(native-chat): type the send-recovery record lookup as the store returns it
* fix(native-chat): a send ensures its owner before admission, and an unheld owner idles for 30 minutes
* fix(native-chat): a create that throws releases its event sink
A child that dies between spawn and journal attach can still write through
the host's event sink, which attach unbound in onAcquiring and never re-bound
because onAttached never ran. The orchestration released that sink only when
performAttach returned a refusal; a thrown failure (the root-exit path) kept
the sink cached with its queued write, so the next attach's drain barrier and
runtime shutdown's flush waited forever.
Also pins the publish-on-root-exit clause for a start that never proved:
deleting it reddened nothing before.
* fix(native-chat): a resend the journal answers restarts nothing, and a send joining a restart rebases from the fence it replaced
* fix(native-chat): the host learns a Claude start positively, and persists only proven options
A publish-first create used to read the session's options before Claude had
answered initialize. With startup pending that read fell back to the built-in
catalog's default, so `record.options.model` was persisted as `sonnet` for
every user whose CLI default is something else; an owner handoff or a reopen
then replayed `set_model('sonnet')` and silently switched their model.
The adapter now reports `started` once startup facts are applied and saved
options restored. The host keeps a `providerChildPhase` on the session it
owns: a starting child hands over nothing but the saved options as intent,
and the `started` event re-reads the options as fact and persists them through
the same record write a user's option change takes. The status summary carries
`hostExecutionPhase` (optional, wire-safe), and the chat pane says the agent is
still starting instead of showing nothing.
A child whose exit already reached the adapter before acquire returns is no
longer handed over as live; the create fails with the CLI's diagnostic.
* fix(native-chat): a hold and a send that find the owner gone share one restart, and a send the ledger already holds restarts nothing
* fix(native-chat): a failed create answers one refusal shape, stamped once at the boundary
A create whose Claude process was seen to exit answered twice in two shapes:
the first call threw a generic runtime error, and only the replay of the same
operation carried the `ownerVerdict: 'exited'` refusal that lets a client
retry under a new operation. Three sites stamped the verdict and the store
failure path stamped nothing.
The first-hand root exit is now returned as the refusal on the first call,
with the provider's own diagnostic as its message. The verdict is stamped in
one place, at the boundary of the attach, from the durable row the operation
settled to, so every refusal shape answers the same fact and no site can
forget it. The per-site stamps are gone.
* fix(native-chat): a send into a session whose child ended restarts it before admission
A session that published and then lost its Claude child before startup (not
signed in, for one) keeps a released lease and a chat the user can still type
into. The send was refused as ownership-unknown, the outbox parked it as
pending admission, and nothing ever restarted the child: the message sat
there until the user closed and reopened the tab.
A send reaching a session with no provider child now runs the same resume a
surface's first hold runs, before the write is admitted. The resume reserves
a new fence, so that send is answered stale with the published fence and the
client's outbox re-drives under it, as after any fence change. A resume that
fails is not this send's answer; admission reports the lease as it stands.
* chore: restore pnpm-lock.yaml to origin/main (local pnpm rewrote it)
* test(native-chat): pin the pre-handover exit as a failed acquire; stub the status feed in the delivery test
An exit the adapter observes before acquire returns now fails the acquire
with the CLI's diagnostic instead of handing over a dead child; the
published-then-ended path stays pinned by the slow-init startup case. The
delivery test renders the pane, which now activates the host status feed.
* test(native-chat): a same-ID re-hold over the wire joins the one resume, and a replay reopen goes on the idle clock
* test(native-chat): a re-hold that joins a failing resume proves one resume ran
* fix(native-chat): a create whose child was proven gone answers the refusal on the first call
The previous change answered a first-hand root exit as the exited refusal on the
first call, but the common failed start never took that path: when the close
ladder proves the whole tree dead the acquisition error is a plain one, the
store-failure classifier rethrows it, and the client still saw a runtime error
first and the refusal only on replay.
The cleanup that proves the child gone now names such a failure
`AgentSessionAcquisitionExitProvenError`, carrying the provider's diagnostic,
unless it already names its own verdict (a refusal, a typed exit proof, a host
store code). The attach answers both proven-exit kinds as the refusal its replay
gives. How a failed acquisition settles and how it is first answered now live
beside the verdict stamp, in the failed-create module.
* test(native-chat): pin the outbox re-driving a stale-refused send under the resumed fence
A send into a session whose child ended is answered stale once the host has
restarted the child. The outbox keeps that operation queued and blocked, and the
fence change the resume publishes re-drives the same operation under the new
fence; the host admits it.
* fix(native-chat): a child restarted for a send nobody holds is still released
The restart a send runs for a childless session takes no holder, on the premise
that the sending surface already holds one. A one-shot writer holds nothing, so
the child it restarted had no release clock and lived until the app quit. The
write resume now arms the clock when no holder is present, as the first-hold
resume already does. The send-after-failed-start cases also pin that the stale
answer's operation is admitted when re-sent under the new fence, and that two
racing sends restart the child once.
* test(native-chat): pin the picked Claude model across a resume whose child starts on its own default
The started event re-reads and persists what the child reports. A resumed child
answers initialize with its CLI default before the saved pick is restored over
it; the record must hold the pick while starting and after started.
* Revert "fix(native-chat): a child restarted for a send nobody holds is still released"
This reverts commit
|
||
|
|
98584332a3 |
fix(native-chat): record which Codex agent produced each journal row (#22532)
* fix(journal): a batch revision restates the producer of each row it revises The reducer rebuilds a row's producer linkage from its NEWEST revision, and absence is a positive claim: no agent id means the session's own agent wrote the row. So any revision written without the stamp hands a subagent's row back to its parent, permanently. Three host paths revise rows they did not write, from the render item they already hold, and all three dropped the stamp: - answering a prompt re-appended the asker's row with the fence only; - dead-generation settlement failed running tool calls and cancelled pending prompts through a lifecycle batch; - stale-session settlement on acquire cancelled lost prompts the same way. The batch path could not carry a producer at all: linkage was removed from the batch row because one row covers N mutations, with a note that a mixed batch would have to stamp per mutation. Dead-generation settlement is such a batch already, and Codex settlement is about to become one. So each item mutation now names its own producer, inline like the row base. A mutation that names none falls back to the row's linkage, which is what a batch read before. Parse sanitizes a bad per-mutation id the same way it does a row's: the field is dropped and the mutation kept. No schema version bump. An older host's mutation validator ignores unknown keys, so it reads a stamped mutation as the session's own, which is exactly what it shows today. Old journals carry no stamp and read as before. Turn revisions still carry nothing: a turn is the session's unit of work, and the live-turn scans rely on a turn row never carrying linkage. The note recording that invariant is updated to the new write sites. * fix(native-chat): attribute a Codex subagent's journal rows to the subagent Codex journals every thread on its app-server connection into the session's journal, and a spawned subagent's items arrive on the child's own thread. None of those rows carried producer linkage, so under the journal's rule that absence means the session's own agent wrote a row, every child's command, message, reasoning, prompt and status row read as the PARENT's: the parent could show its child's running command, its child's reasoning as "thinking", and its child's prose as its own latest line. The Claude lane's model is reused, not reinvented: the same fields and the same absence rule. What differs is how the producer is known. Orca opens exactly one thread per app-server, so any other thread is one Codex spawned. That decides WHETHER a row is a child's from its first frame, announced or not, and the thread id is final at once: it is never re-minted the way a tool-call reference is, so no correction ledger is needed for identity. - agentId: the child thread id, the same id the status side keys a Codex child on. - parentAgentId: the thread whose stream carried the child's `started` activity. Codex emits that item on the spawning agent's own session, so a child that spawned a grandchild is named; the session's own thread is not. Other activity kinds ride whichever agent acted and are not used. - producerKind: 'agent'. - attempt: which run of the child the row's own turn was, counted from the child turns the roster already observes; absent on the first run. Taken from the row's turn rather than the child's latest, so a persistent shell that outlives its turn keeps its run across revisions. - providerParentRef is omitted: a Codex child's frames carry no parent reference of their own beyond the thread id, which is already agentId. One resolver, owned by the roster (which already owns what is known about each child thread), is handed to every writer: items, streams, generic and summary rows, prompts, compactions, goals, and the three settlement batches. The session-end settlement mixes every thread's rows in one batch, so each mutation names its own producer. Turn rows stay unstamped: Codex writes them only for the primary thread. The spawn-group roster row stays unstamped on purpose: a child's frame can trigger its write, but it is the parent's list of its children. The translator's construction moves to a parts module so the translator stays a router under the line cap, and the item streams reuse one append-and-publish helper instead of two copies. Children are never swept at turn end; nothing here changes that. * test(native-chat): pin Codex subagent attribution at every writer and every parent reader Two layers, so a stamp that is correct in the store and never read, or read and never persisted, cannot pass. The readers, through the real path: translator, deferred sink, on-disk journal, snapshot. Each is a defect on main: the parent named its child's running command as its own tool, read its child's reasoning as itself thinking, showed its child's compaction as its activity line, and quoted its child's prose as its latest line (checked after closing and reopening the journal, so the stamp is read back from disk). The transcript still renders the child's rows. The writers, through a sink that records the linkage of every plain append, batch mutation and lifecycle transition: start, streamed checkpoint and completion of one command all restate the child; a row that beats the spawn announcement is still the child's; a grandchild names the child that announced it, while an `interacted` activity names no parent; a follow-up turn is the child's second run, and a shell that outlives its turn keeps its own; the exit batch settles each thread's rows under its own producer and the turn row under none; a child's provider frames, approval and goal rows are its own; nothing is stamped while the session thread is still opening; and the spawn-group row stays the parent's. * test(native-chat): pin linkage forwarding on the sink's lifecycle-transition path A Codex child's goal row is written through a lifecycle transition, so a sink that forwarded only the fence there would file the child's goal as the session's own. * test(native-chat): type the Codex item fixtures as thread items * refactor(journal): keep a row's producer across revisions that name none The reducer took a row's producer linkage from its newest revision, so every writer that revised a row it did not write - a prompt answer, a dead-generation or stale-session settlement, the reopen sweep of stale subagent rosters - had to restate the producer or silently hand a subagent's row to the session's own agent. Three of those writers had been patched to restate it; the next one to forget would reintroduce the bug. Attribution is now fixed by a row's first write. A revision that names no producer keeps the row's existing linkage; one that names any replaces the whole bundle, which is how a provisional stamp is still corrected in place. A row re-created after a tombstone starts with nothing. The reducer runs the same fold on replay, so the kept producer survives a reopen. The three restatements are removed. Per-mutation linkage on lifecycle batches stays: a batch can create a row (a Codex child's prompt, or a child's item settled before any checkpoint landed) and one batch can mix producers. * test(journal): pin producer inheritance in the reducer and across a reopen A revision naming no producer keeps the row's, on the plain item path and in a batch settling a child's row beside the session's own; one naming any replaces the bundle wholesale; a tombstone clears it; a stale revision cannot touch it; and a reopened journal replays it exactly as it was folded live. * refactor(codex): name the translator's writer factory for what it builds * docs(codex): say why a settled row names its producer * test(journal): drop a producer test the stale-revision guards make unreachable The stale revision is dropped whole by two independent guards before the inheritance rule runs, so its producer assertion could never fail; the reducer's own stale-revision tests already cover the drop. Also say what the batch sink does forward: each mutation's own producer. |
||
|
|
563dd5487f |
feat(native-chat): show a Codex chat's goal above the composer, and set it from goal mode (#22377)
* feat(native-chat): show a Codex chat's goal above the composer and set it from goal mode Structured Codex chat now treats the thread goal as session state: a banner above the composer shows the current goal (pursuing / paused) with clear, pause/resume and expand; /goal enters a goal mode whose send calls thread/goal/set; the objective is journaled as a user message marked as sent as a goal. The banner is derived from the journaled goal rows, which Codex's resume snapshot refreshes, so a reopened or adopted chat shows its goal. Fixes STA-8159 * fix(native-chat): replace a recorded goal by clearing first, and recover a lost goal-change response - A set while the journal records a goal (any status) clears it before setting, so the new goal starts with its own time and token counters instead of rewriting the old goal's objective in place. - The threadGoal plan answers an unknown outcome from the goal the journal records and reruns otherwise, so one request timeout no longer refuses every later Clear/Pause/Resume as unknown for the mounted session. - The goal-mode chip says "Exit goal mode"; "Clear goal" stays the banner's action on the provider goal. - A typed bare /goal on Enter enters goal mode, the same as picking it. - The renderer reads the goal off the tail of its ordered snapshot; the host's unordered map keeps the by-sequence reader. - Drop the composer's duplicate in-flight guard; the goal controller already serializes changes. - Pin that a counter-only revision reaches a subscriber's live page under its original sequence. * fix(native-chat): keep a bare /goal inside goal mode as the entrance, and pin goal delivery and serialization - A bare `/goal` submitted while already in goal mode re-enters the mode instead of setting a goal whose objective is the literal text "/goal". - The counter-only revision pin now drives the host's own event sink bound to a real journal, so it goes red when the publish after a lifecycle transition is dropped; the previous fake sink never published. - Pin that a set which threw after journaling its objective puts that objective back exactly once when the ledger reruns the same operation id. - Cover the goal controller hook: absent without host support, the loaded window wins over the host's answer, a second change while one is unsettled answers false without a request, and a refused change frees the next one. * fix(native-chat): resume a blocked or usage-limited goal, and keep goal-mode drafts honest - The goal bar offers Resume on a blocked or usage-limited goal, which the provider resumes exactly as it resumes a paused one; a goal whose token budget is spent still offers only Clear. The rule lives beside the other goal facts in shared code so every reader answers it the same way. - A `/goal <text>` typed inside goal mode sets the objective `<text>`, as it does outside goal mode, instead of a goal whose objective is the literal command. - Setting a goal is a host round trip; a draft edited while it was in flight is no longer wiped when the goal lands, matching every other host command. - Pin that a lost status-change response is read as applied only when the recorded goal is in that status, that a cleared row in the loaded window outranks the host's earlier answer, and that the PTY lane is untouched. * fix(native-chat): keep the load-older anchor on the loaded window when a live revision lands below it A live revision of a row keeps that row's original sequence. When the row is older than the client's loaded window, the shared reducer merged it in and it became the load-older anchor, so paging `before` it skipped every row between. A goal's counter-only revisions during a long goal turn reach any client that attached after the goal row left its window, so a reopened chat lost rows on scroll-back. The reducer now admits live rows only at or above the window's oldest row while older rows remain on the host; the journal keeps the revision and the page reader serves it once the window reaches the row. With nothing older on the host the window is the whole journal, so a row below the head is admitted as before. Also drain accepted provider events before a goal set reads the journal to decide whether it replaces a recorded goal. |
||
|
|
dc8cf30554 |
fix(native-chat): end a structured turn when the agent reports it failed (#22047)
* fix(native-chat): end a structured turn when the provider reports it failed (#22044) A turn reads as working while its durable turn row says `running`, and only two events could write a terminal row: the provider's turn-completed notification and the provider process going away. A provider error that ends a turn is neither, so the row stayed `running` with nothing re-deriving it, and the chat counted "Working for N" for the life of the session. Codex reports such a failure as an `error` notification naming the turn it ended, with `willRetry` distinguishing it from a stream error it is about to retry. That frame now settles the turn it names. Claude's CLI reports the same through its session-state frame, whose `idle` arm the SDK documents as the authoritative turn-over signal; that now settles the open turn too. Codex's `thread/status/changed` deliberately settles no open turn: the app server clears `running` on every error, including ones it reports as not affecting turn status, so a turn still open there is still running. What it does settle is a send whose dispatch was never answered — a timed-out dispatch is recorded as unverified delivery, reads as work still owed, and nothing in a live session retired it. Retiring it never makes the send re-deliverable. Splits the codex notification translator so the file stays inside its line budget. * fix(codex): defer idle dispatch release until turn end * fix(claude): enable session state lifecycle events |
||
|
|
4085e1cf60 |
fix(memory): release stale session registries (#21734)
* fix(memory): bound session and lifecycle registries * fix(memory): bound transient filesystem registries * fix(memory): cap path and locale caches * fix(memory): bound runtime recovery registries * fix(memory): bound host mirror gap verdicts * fix(memory): bound shell startup env cache * fix(memory): bound gitlab host context cache * fix(memory): release removed ssh generations * fix(memory): expire cloud refresh replay guards * fix(memory): release retired plugin generations * fix(memory): bound plugin log key retention * fix(memory): bound automation authority generations * fix(memory): bound native chat enrichment cache * fix(memory): bound web session tracking generations * fix(memory): bound codex credential absence paths * fix(memory): bound WSL canonical path cache * fix(memory): bound sparse checkout cache * fix(memory): bound shared directory cache * fix(memory): bound advertised URL scan snapshots * fix(memory): bound automation manager cache * fix(memory): bound web session reorder intents * fix(memory): bound web session focus intents * fix(memory): bound web session handoffs * fix(memory): bound automation dispatch tokens * fix(memory): bound host mirror waiters * fix(memory): bound retained session activity * fix(memory): bound retained session activity * fix(memory): bound web session close intents * fix(memory): bound cloud session cache * fix(memory): bound WSL home cache * fix(memory): bound SSH capability cache * fix(memory): bound trust grant cooldowns * fix(memory): bound WSL auth drain state * fix(memory): bound Linear workspace credential cache * fix(memory): bound local Git capability cache * fix(memory): bound WSL Git environment cache * fix(memory): bound WSL Git environment cache * fix(memory): bound WSL preflight cache * fix(memory): keep hot cache entries warm * fix(memory): preserve generation fences across eviction * fix(memory): close remaining eviction fences * fix(memory): align evicted upstream generations * fix(memory): trim successful capability probes * fix(auth): retain expired refresh replay evidence --------- Co-authored-by: m4air <m4air@Mac.localdomain> |
||
|
|
9907117569 |
feat(native-chat): record an explicit provider outcome on every structured turn (#21278)
A structured turn that FAILED was recorded as `completed`, identically to one that succeeded, so nothing downstream could tell them apart. Claude mapped only its two abort reasons to `interrupted` and let an API error fall through to `completed`; Codex collapsed every non-`completed` status to `interrupted` and read a missing status as a clean finish. Add `outcome` — success / failure / cancellation — to the turn record, emitted by both providers. The four-arm lifecycle union is deliberately untouched: it stays a report on what the HOST observed, and its readers are unaffected by construction. Absent means UNKNOWN and never success. Historical rows, older hosts, and any end the host inferred rather than heard (the child going away, a turn superseded before its result) all carry no outcome, so a newer client cannot mistake an old host's `completed` API error for a clean turn. Claude's abort-reason list had a second copy in the provider-fallback reader; both now classify through one `claudeResultOutcome`, so the durable verdict and the visible error row cannot drift. |
||
|
|
fbfe3a2e74 |
fix: release Codex prompt claims when their turns complete (#21138)
Co-authored-by: m4air <m4air@Mac.localdomain> |
||
|
|
434365d2de |
Offer to reconnect native chats that were working when Orca restarted (#21096)
* feat(native-chat): resume structured chats that were working at restart
Teardown records a marker for every session this host was genuinely running a
turn for, derived from the LIVE runtime rather than a persisted status row, so
a stale `running` row left by an older crash can never trigger a resume. On the
next launch a modal lists exactly which chats would resume and resumes them via
native continuation (Claude resume/resumeSessionAt, Codex thread id) — never by
re-sending the prompt, which is what makes an agent redo finished work.
A session resumes only when all of these hold: a teardown marker exists and has
not expired, the record's lease is released and reconciled, a provider resume
cursor exists and still matches the marker, the journal's own turn record names
the same turn, and the marker has not already been spent. Markers are consumed
before the resume is submitted, so a crash mid-resume cannot double-fire, and an
admission gate refuses a second concurrent resume for one session. Resumes are
staggered three at a time rather than spawning every provider at once.
The modal's "Don't ask again" checkbox writes the nativeChatResumeWorkOnRestart
setting, which Settings can turn back off; automatic mode runs the identical
predicate and staggering and reports what it did. Declining consumes the markers
so the prompt cannot return every launch — nothing is lost, because opening a
chat still re-acquires it at the same cursor.
* fix(native-chat): compare handle ROOT and turn state when offering a resume
Four defects QA found in the restart-resume offer, fixed together because the
first two interact: shipping the root fix without the state fix would convert a
silent no-op into actively offering finished chats.
1. Claude was never offered (0/4). The marker recorded agentSessionProviderHandleKey,
which embeds Claude's leaf uuid — a branch cursor. The adapter's own close path
appends a `resumed` link with an advanced leaf during the SAME teardown, so the
marker went stale seconds after it was written and the drift guard refused every
Claude session forever. Record and compare agentSessionProviderHandleRoot instead:
the root is the part a resume must preserve, and changing it is a fork, which is
exactly what this guard is for. Codex is unaffected (its thread id is the whole
key) but uses the root too, so the rule is uniform.
2. The predicate compared turn IDENTITY but discarded turn STATE, so a `completed`
turn satisfied it as readily as an interrupted one. Eviction rewrites `running`
to `interrupted` and never to `completed`, so the state is what separates work
that was cut off from work that finished. Require `interrupted` or `unverifiable`.
3. A chat blocked on a pending approval or question was marked as working, because
the teardown reader accepted any `running` turn while the product's own projection
calls that state `attention`. Teardown now defers to that projection: an agent
waiting on the USER is not interrupted work.
4. "Resume all" could silently no-op. The modal fetched candidates at mount; by click
time the chat's own pane may have bound and taken the hold, moving the lease to
`live` so the predicate dropped it and the call returned no results, leaving the
dialog open behind a dead button. Re-derive at click time and settle an
already-live session as resumed — it is running, which is what the user asked for.
Test fakes now model the Claude close path that advances the leaf, which is why no
unit test could previously exhibit defect 1. Ablation covers all eleven guards.
* fix(native-chat): gate the already-live settlement on the full resume predicate
Two follow-ups from re-QA, both cases of a rule stated by intent rather than by
discriminator.
1. The already-live path bypassed the predicate. "Resume all" sends no session
ids, so the fallback's target set was every marker, and it was gated only on
the session having a live provider child. A chat the predicate had refused --
a completed turn, say -- whose pane happened to own the lease was therefore
settled as `already_live` and had its marker spent, inflating the "Resumed N"
count with chats that were never eligible. No provider spawned and no tokens
were spent, but a marker the predicate rejected must never be consumed.
The resumable set now takes an explicit `leaseState`. The already-live path
derives a second set with ONLY the released-lease clause relaxed, and settles
a session just when it is in that set. Every other clause still applies.
2. The `attention` rule was one-sided. Teardown refuses to mint a marker for a
chat blocked on the user, but the set predicate had no equivalent, so a marker
arriving by any other route was offered once eviction rewrote its turn to
`interrupted` -- the same asymmetry the completed-turn case had.
Gated on projectStructuredAgentSessionStatus === 'attention'. That projection
tests for a pending approval or question BEFORE it looks at turn state, so it
still reports `attention` after the turn is settled, which makes it the durable
signal and keeps one source of truth with teardown.
Ablation now covers thirteen guards, including one for each of the above.
* fix(native-chat): capture awaits-user on the marker instead of re-deriving it
The awaits-user clause could never fire. It asked the live projection for
`attention`, which needs a prompt whose resolution is still `pending` -- but
teardown CANCELS that prompt a few phases after it writes the marker. By the next
launch the evidence is gone, for precisely the sessions the clause was written
for. QA measured the injection still being offered and then resumed.
This is the same shape as the leaf-drift bug: state read after teardown is not the
state that justified the marker. The discriminator, now applied across the whole
predicate:
- a fact teardown itself destroys or mutates must be CAPTURED on the marker
while it is still true;
- a fact that evolves on its own must be RE-DERIVED at read time, never
snapshotted.
So `awaitsUser` is now recorded at teardown and the predicate reads the recorded
value. Teardown still declines to mint a marker for such a session, so the
recorded flag is the second line rather than the only one.
Audit of every other clause against the same test:
- turn id (captured) -- teardown rewrites turn STATE but never the id. Correct.
- provider handle root (captured) -- the close path appends a resumed link, and
appendAgentSessionProviderHandleLink refuses one that changes the root, so the
root is invariant under exactly the mutation that broke the key. Correct.
- turn state (re-derived) -- DELIBERATE exception, stated here rather than left
implicit: we are not reading the state that justified the marker, we are
reading teardown's receipt that it settled the turn. A turn still `running`
means eviction never finished, and we refuse. Correct, and intentionally so.
- lease reconciled / released / handoff stage (re-derived) -- these answer a
different, launch-time question: may this host take the lease NOW. The
teardown-time value would be meaningless, and `unreconciled` is cleared by
this launch's own reconciliation. Correct.
- adapter support, marker TTL, marker consumption (re-derived) -- all evolve
independently of teardown. Correct.
Only awaitsUser was on the wrong side.
* fix(native-chat): drop the unreachable awaits-user marker flag
The captured flag was dead code. `awaitsUser` could only be true when the
projected status was `attention`, and `attention` hits the `continue` above the
push -- so every marker teardown can ever write carries `false` (QA measured
22 of 22 across two real teardowns). The predicate clause reading it was
unreachable by any production path.
A flag that is structurally always false is worse than no flag: it reads as a
safeguard, so the next person to touch this trusts it. The asymmetry it was
added to close was only ever reachable by fault injection, because teardown is
the sole writer of markers and already refuses attention sessions.
Removing it also drops an upgrade discontinuity: as a required field it made a
marker written by the previous build fail validation and be silently discarded,
costing a resume offer on precisely the upgrade where the user was mid-turn.
Markers predating the providerHandleRoot rename still will not parse, but those
carry a leaf-sensitive key the predicate would refuse anyway, so nothing usable
is lost.
In its place the teardown gate now states that `status !== 'working'` is the
SINGLE gate for awaiting-user sessions, why a predicate-side mirror would be
unreachable, and why it could not even re-derive the fact -- so the reasoning is
inherited rather than rediscovered.
Ablation is back to twelve guards; every other clause is unchanged.
* fix(native-chat): say reconnect, not resume, and show each offer's age
Two changes, both independent of the parked continuation decision.
1. The copy claimed something QA disproved. "Resuming continues each agent where
it left off" is false: reconnection restores the session at the point it
stopped, with full context and without re-sending the prompt, but the
interrupted reply does not continue on its own. The toast's "Resumed N chats"
implied work had restarted.
Audited every user-facing string against the rule that none may claim work
continues or that a reply resumes -- which caught more than the three strings
the fix started from. The title, the row button, "Resume all", "Resuming...",
the not-now hint ("picks it up where it left off"), the checkbox and its hint
("resume on their own"), the list's aria-label and the Settings row all made
the same claim. The user-facing verb is now reconnect throughout; the body and
update variant state outright that the interrupted reply will not continue.
en.json synced, runtime boot catalog regenerated.
If we later decide to send a continuation instruction, this is one commit to
change back. Shipping text we know to be false was the worse option.
2. Rows now show each offer's age. The TTL is 24 hours and a stale offer looked
identical to a fresh one. The marker already carried `recordedAt`, so this is
a render change plus one field on the renderer's candidate type, formatted
with the existing formatUiRelativeTime helper rather than a new one.
The clock is stamped once when the list arrives rather than read during render:
ages then stay stable across re-renders, and the render stays pure, which the
react(purity) rule requires.
Guards, predicate and RPC are untouched; ablation still covers twelve.
* feat(native-chat): show the workspace name on each reconnect row
A row read `codex · folder:8f3a1c22-… · 8 hours ago`. Recognising which chats
would reconnect is the entire point of the list, and at twenty rows a UUID
identifies nothing.
No RPC or host change was needed: the renderer can already resolve this id.
Resolved the way automation dispatch resolves the same id space
(resolveAutomationDispatchWorkspace) -- a folder workspace by its full
`folder:<uuid>` key via getKnownWorktreeById, a git worktree by its bare
`repoId::path` id via allWorktrees. Both return a Worktree, whose displayName is
a required field, and DetectedWorktree extends Worktree so either shape answers.
Falls back to the id when nothing resolves, which is what the row showed before
and also covers the window before the worktree store has hydrated.
The lookup lives in a per-row subcomponent because a hook cannot run inside
`map`, and its selector returns a primitive string so repeated selector runs
cannot churn referential equality.
* feat(native-chat): group the reconnect modal by worktree and add opt-in continuation
Grouping. Rows are now grouped under a worktree heading with the repo glyph and
an agent count, using the sidebar's own collapse mechanics. Only presentational
pieces are reused -- RepoIconGlyph, CompactAgentExpansion, AgentIcon and
formatShortTimeAgo. The sidebar's agent row cannot be: worktree-card-compact-agent-row
imports DashboardAgentRow, the dashboard's own type, so both surfaces render one
live-agent model requiring a pane, tab and status entry. Every chat offered here
is by definition stopped, so supplying that would mean inventing live state.
Two things I had assumed were reusable and were not:
- DashboardHostBadge returns null unless hostKind is ssh or remote. Structured
chat is local-only, so it would always render nothing. The host line is
omitted rather than faked; the badge is the right element to add if and when
structured chat gains remote support.
- No state dot. Every AgentDotState misleads here: idle and unverifiable both
presuppose a live pane, interrupted renders red like an error, done green,
working a spinner. A missing dot beats one saying these agents are running.
One worktree renders flat with no heading -- a name, count and chevron around a
single group says nothing the dialog has not already said.
The age column now uses formatShortTimeAgo for sidebar consistency. It takes
(timestamp, now) and subtracts internally rather than taking a delta, so the call
is (recordedAt, listedAt); passing the old delta would have rendered plausible
nonsense. The clock is still stamped once into state, so ages stay stable and the
render stays pure.
Continuation. A secondary "Reconnect and continue" action sends one message, from
a single shared constant, identical for both providers. Reconnect is unchanged and
still sends nothing. An info popover quotes the literal message read from that
same constant, so what is shown cannot drift from what is sent.
Ablation now covers fourteen guards. Two are new: continuation only follows a
reconnect that actually happened, and -- inversely -- a send injected into the
reconnect path must turn the test red, since "don't ask again" rests on reconnect
never sending.
* feat(native-chat): say terminal sessions kept running, and clear the quality gate
The modal lists stopped chats with no way to tell that CLI agents are fine, and
the true state of the world is counterintuitive: the terminal sessions survived
the restart and the chats did not. One line now says so, next to the heading
where it frames the list rather than as a footnote at the bottom.
Wording follows the app's own vocabulary rather than inventing a term: the
catalog settles on "terminal sessions" (terminalSessionCount, "Terminal sessions
are grouped by workspace", "No terminal sessions yet"), and UpdateCard already
reassures with "Your terminal sessions won't be interrupted during the update" in
the same text-xs text-muted-foreground treatment. "kept running" rather than
"were restored" -- nothing reconnected them, they never stopped, and the line
says nothing about why.
Also clears check:code-quality:changed, which I had not been running -- oxlint
alone covers neither the design-system nor the casting audit, so 18 findings had
accumulated across the branch.
- design system (4): Button spacing hand-rolled as gap-1/px-2 is just size="xs";
PopoverContent and DialogTitle own their typography and spacing, so the
text-xs moved to the popover's own children and the title's icon gap moved to
a plain wrapper.
- casting (14): production code loses its assertions outright via Reflect.get,
the idiom already used in managed-hook-detection-commands and
worktree-name-retirement. The marker validator reads each field through
Reflect.get and now checks recordedAt is a number rather than asserting it;
the store-file parse uses the existing `file` shape instead of a second
assertion; the runner narrows the admission error's owner with typeof.
Test fixtures keep their assertions behind the line-specific SAFETY:
rationale the repo mandates for exactly this case.
One trap worth recording: the audit reports an assertion at the line its
EXPRESSION OPENS, not where `as` appears, so a disable-next-line above the
closing brace of a multi-line literal is inert and silently changes nothing.
Guards unchanged; ablation re-proved 14/14 at this head.
* fix(native-chat): give the reconnect row's provider icon an accessible name
Every row rendered the provider as a bare AgentIcon, whose svg carries no
aria-label, title or alt. With a Claude chat and a Codex chat in one worktree the
two rows were identical to any non-visual consumer, and the dialog offered
several identically-named "Reconnect" buttons with nothing to tell them apart.
A regression from
|
||
|
|
7f5141ae2d |
Make the Agent Permissions toggle apply to Codex chat (#20977)
* fix(structured-chat): deliver the permission posture through each transport's own contract Codex posture moves off app-server argv onto typed `thread/start` and `thread/resume` params. Manual states `on-request` / `workspace-write` explicitly instead of omitting the fields, which app-server resolved through the mirrored config.toml — a Manual thread on a home carrying `approval_policy = "never"` never prompted. Claude keeps its owned `--dangerously-skip-permissions` flag through SDK `extraArgs`; the SDK's typed bypass option emits a newer allow flag that older user-installed binaries reject. Posture is re-derived from current settings on every session acquisition. * fix(structured-chat): parse permission arguments as argv * fix(structured-chat): keep permission policy authoritative |
||
|
|
533b0bd02e |
fix(native-chat): count a turn from the send that opened it (#21086)
* fix(native-chat): count a turn from the send that opened it The live turn indicator switched on at the submission but anchored its clock at the provider turn-open, so it jumped back by exactly the dispatch latency the moment the turn opened. Measured on a real Claude session: the counter climbed to "Working for 25s", reset to "Working for 0s", then settled "Worked for 26s" — three readings of one turn, from two different instants. The host now resolves the send that opened a turn and publishes it as an additive optional `requestedAt` on the turn lifecycle row. `startedAt` keeps its exact meaning, the provider turn-open, and is never rewritten, so clients that cannot be upgraded see no change to any value they already read. Both providers write it; it is omitted when no send can be named (provider-resumed turns, replayed history). Readers take one origin, `requestedAt ?? startedAt`, for both the live counter and the settled host interval, so the two cannot disagree. The provider's own reported duration keeps outranking the host interval, unchanged. The host-to-local clock conversion is now latched once per turn rather than re-derived per render. `receivedAt - hostNow` carries that sample's one-way delivery latency as well as skew, and the reducer replaces the sample on every frame, so re-deriving imported fresh jitter and could move the anchor later — the same class of backwards jump this change removes. With the conversion fixed, an origin that improves moves the anchor earlier by exactly that much, so displayed elapsed only grows. No monotonicity guard is added; the ordering is structural. Desktop and mobile drove byte-identical copies of the timing hook, so both are collapsed onto one React-free helper in shared. Regression tests drive the origin resolution rather than an already-resolved anchor, assert in milliseconds because second-flooring hides the sub-second case, and include a deliberate host/client skew so a raw timestamp assignment cannot pass on a machine where the two clocks agree. * fix(native-chat): correlate Codex turn origins by echo * fix(native-chat): preserve causal turn timing ownership * fix(native-chat): keep settled turn timing continuous |
||
|
|
d62328aa4d |
fix(codex): remove redundant Windows hook launcher for Unicode profiles (#20952)
* fix(codex): reuse the Windows hook shell for Unicode profile paths * test(codex): register Unicode hook tests in Windows CI * test(codex): pin trust hash replacement during Windows upgrade * test(codex): retry transient Windows teardown locks |
||
|
|
c702e77bc7 |
Stop reading the terminal arguments field on the structured chat route (#20944)
* fix(native-chat): stop reading the terminal arguments field on the structured chat route Setting Claude's Arguments to "--dangerously-skip-permissions --model Opus" made every new Claude tab open in the old terminal-backed chat instead of the new structured one, with nothing on screen to explain why. Removing "--model Opus" fixed it. The cause was a whole-string comparison: the configured arguments were checked against a single blessed value per agent, so any added token at all — including one the agent supports — stopped the string matching and the launch was demoted. Structured chat does not run the interactive CLI. It drives Claude through the Agent SDK and Codex through app-server, and those take narrower option sets that are versioned separately from the CLI's, so one free-text field cannot have a guaranteed meaning for all three. The structured route now reads only what it can actually honour: a replaced launch command, or a launch that names its own working directory. Terminal launches still apply the field exactly as before. Permission posture no longer travels as a raw flag. It is derived from the resolved launch arguments, which is the same fact a terminal launch acts on and which falls back to the default Orca ships when the field was never touched, so bypass stays on by default and Manual is still honoured. Claude gets the SDK's typed permissionMode and allowDangerouslySkipPermissions at query start; Codex gets its bypass flag placed before the app-server subcommand. Both are re-derived per acquisition beside the auth policy and environment overlay rather than stored in the session record, so nothing can disagree with the setting. Codex also loses the --profile, --add-dir and -c passthrough that reached app-server through that field. Only the permission posture comes back. * test(native-chat): pin routing authority on the narrowed feasibility input The routing-authority pin still named the old bundled blocker and built its "customized" fixture out of the arguments field, which is no longer a feasibility input. Both are now the launch command, and arguments and environment are customized on both passes of the loop, so the flag handed to the shared resolver tracks the command alone — a caller that resumed reading either one fails here. No case is dropped and no assertion is relaxed: the blocker list is still exhaustive and every caller must still honour a refusal from the shared resolver. |
||
|
|
231e805b1e |
fix(lint): enable anti-slop/no-shape-in-symbol-names (#20785)
Flip `anti-slop/no-shape-in-symbol-names` from "off" to "error" and clear
every violation under src, config, tests and mobile.
What the rule bans
------------------
The case-insensitive substring "shape" in any JS/TS identifier: variables,
functions, parameters, types, type parameters, class members, private names,
object-literal keys and JSX identifiers. The one exemption is a statically
accessed member read owned by another value (`zodObject.shape` is fine), so
third-party APIs stay readable without a suppression.
"Shape" names a value's structure rather than its domain role. `UserShape`,
`validateArgShape` and `errorShape` all tell you the symbol is "an object
with some fields" -- which is already what a type says -- while saying
nothing about what the value is for or who owns it. The rule forces the
name to carry the domain instead.
Violations fixed
----------------
689 violations across 109 files at baseline (verified by re-running the
audit against the pre-change tree with the rule set to "error").
Fix pattern
-----------
Rename for the domain role, not the structure:
-type FieldShape = 'list' | 'map' | 'whole'
-const FIELD_SHAPES = { ... } satisfies Record<keyof Observation, FieldShape>
+type FieldEncoding = 'list' | 'map' | 'whole'
+const FIELD_ENCODINGS = { ... } satisfies Record<keyof Observation, FieldEncoding>
-function assertGitPushTargetShape(target: unknown): void
+function assertValidGitPushTarget(target: unknown): void
-function describeReadDirPathShape(p: string): ReadDirPathKind
+function classifyReadDirPath(p: string): ReadDirPathKind
Predicates became statements about the value (`isDeltaShapedProviderFrameKind`
-> `isDeltaProviderFrameKind`, `isDeleteShapedDiscardEntry` ->
`discardDeletesEntryFile`, `isSkillsCliAgentKeyShaped` ->
`isUsableSkillsCliAgentKey`). Type aliases dropped the suffix where the
remaining name was already unambiguous (`GhGraphqlErrorShape` ->
`GhGraphqlError`).
No wire-visible name was renamed: no IPC or RPC channel, stream opcode,
request/response param, persisted field, or i18n key. The `--shape=symlink|copy`
CLI flag read by .github/workflows/skill-update-roundtrip.yml is unchanged --
only the local variable holding it was renamed.
Exemptions
----------
They are file-scoped entries in config/oxlint-anti-slop.json, not inline
`oxlint-disable` comments. An inline directive naming an anti-slop rule reads
back as an UNUSED directive under the root lint scan, which does not load this
plugin -- the changed-code quality gate counts that warning, so the comment form
cannot be used for a rule that lives only in this config.
* src/renderer/src/components/browser-pane/annotate/**:
in the screenshot annotator a "shape" is the drawn geometry -- pen, arrow,
rect, ellipse, highlight. That is a genuine domain noun, and it pervades
every symbol in the module.
* repo-icon.tsx, repo-header-project-actions.tsx, mobile MobileRepoIcon.tsx:
lucide exports the icon component as `Shapes`. The name is theirs, and the
matching REPO_LUCIDE_ICONS key is the persisted icon name shared with the
desktop picker -- renaming it would orphan saved repo icons.
* src/shared/onboarding-state-types.ts, src/shared/constants.ts:
`shapedSidebar` is a persisted onboarding-checklist field and a telemetry
enum member; renaming it would orphan saved state.
* src/shared/rpc-contract/rpc-send-params.ts: matching zod's own literal `shape`
property is what selects the ZodObject branch of the conditional type.
No exemption was added merely to avoid a rename. Eight symbols initially
suppressed as "a cross-module refactor outside this change" were proven to have
zero non-TypeScript references repo-wide and renamed instead.
Zod's `ZodRawShape` needed no exemption at all: `Readonly<Record<string,
z.ZodType>>` is its definition, so repo-update-params.ts and
ui-update-value-tolerance-params.ts spell it out instead. Likewise
telemetry-event-classification.ts now reads `.shape` through an `in` narrowing,
which also retires two pre-existing type assertions; three more assertions the
rename had dragged onto changed lines (two `JSON.parse` sites, one node:sqlite
row read) became annotations and an explicit row mapping.
Verified
--------
* Audit reports zero violations; confirmed the rule genuinely fires by
planting a probe violation.
* node config/scripts/run-typecheck-projects-in-parallel.mjs exits 0.
* Vitest over src/shared, src/main/github/project-view, the annotate module,
the repo-icon components and the Chromium SameSite electron spec: all green.
* All 66 removed "shape" identifiers grepped repo-wide across every file type;
none survive.
* node config/scripts/generate-rpc-params-catalog.mjs --check exits 0.
* node --check on every changed .mjs; oxfmt clean on all changed files.
* `pnpm run check:code-quality:changed` reports 0 findings.
Not machine-verified: the 3 mobile/ files (its Vitest run cannot resolve
`expo/tsconfig.base.json` in this worktree), and the WSL- and Playwright-gated
specs. All are rename- or comment-only hunks, read in full.
|
||
|
|
bfdec26352 |
fix(lint): enable anti-slop/no-object-parameters (#20781)
The rule rejects the broad `object` type on any function input (declarations, expressions, arrows, methods, call/construct signatures, function types), plus local aliases and unions that resolve to `object`. `object` accepts every non-primitive while exposing no properties, so it documents nothing and pushes callers into assertions at the boundary. Fixes all 185 violations across src, config, tests and mobile, and flips the rule from "off" to "error" in config/oxlint-anti-slop.json. Approach: replace each `object` input with the type its owner already has. Most sites took an existing domain type or a type-only import (36 added); 40 new aliases name shapes that had none. Where a value is genuinely only compared by reference, it gets a named identity token instead of a shape -- `Record<string, never>`, the built-in `WeakKey`, or a `unique symbol` brand, matching the branding already used in src/shared. Same treatment for WeakMap and Map key parameters. Two `as unknown as` casts became unnecessary once the parameter carried a real type and were removed; no new casts were added. Suppressions added: none. No `oxlint-disable` for this rule anywhere, and no max-lines disable or per-file bump. Three files sat exactly at their max-lines cap, so the added type imports were made line-neutral rather than suppressed: - src/main/ipc/browser.ts exports the existing guest-registration args type (renamed BrowserGuestArgs) so browser.test.ts reuses it on one line. - pane-scroll.ts takes TerminalScrollIntentTarget through the existing pane-manager-types import via a type-only re-export. - direct-rpc-client.ts drops the identity parameter entirely: the session check moved into the sendProbe callback that owns the token. Verified: anti-slop config reports zero violations over src config tests mobile; run-typecheck-projects-in-parallel exits 0; 144 affected test files pass (1749 tests); oxlint and oxfmt clean on all changed files. Mobile has no runnable test/typecheck target in this worktree (expo is not installed), so its 6 files were typechecked against a standalone config and diffed against the base branch -- error sets are byte-identical, including test files. |
||
|
|
f107499e44 |
fix(lint): enable anti-slop/no-reflect-get (#20786)
`anti-slop/no-reflect-get` rejects every call to `Reflect.get`. The
reflective read bypasses ordinary property access and throws away the
type evidence the compiler would otherwise give you: the result is
`any`/`unknown` with no narrowing, so a typo in the key or a shape drift
in the source object is invisible until runtime. The rule's remedy is to
parse dynamic input into a named domain type (or narrow it with `in`)
and then read the field normally.
Baseline: 86 violations across 67 files. Now zero unsuppressed
violations under
`npx oxlint --config config/oxlint-anti-slop.json --ignore-pattern 'config/oxlint-plugins/anti-slop/**' src config tests mobile`.
Fix pattern
-----------
44 of the 86 were rewritten. The dominant shape was an `unknown` value
read through `Reflect.get` right after a `typeof === 'object'` guard;
those became `in`-narrowed property access, which TypeScript checks:
- Reflect.get(value, 'agents')
+ 'agents' in value ? value.agents : null
Two further shapes:
- `Reflect.get(Object(x), 'k')` on a possibly-primitive envelope became a
small named reader that boxes once and indexes a
`Record<string, unknown>` (`settingsField` in
mobile/src/transport/settings-read-operations.ts).
- Tests reaching into private state moved to TypeScript's checked
bracket-index escape hatch (`runtime['layoutQueues']`), or to a
documented read-only accessor on the owning class
(`SearchSubprocessLineAccumulator.retainedCapacityBytes()`,
`CodexSubagentExecutions.retentionSizes()`).
No type assertion was added anywhere: the diff contains zero net-new
`as` casts, `as any`, `as unknown as`, `@ts-ignore`, or
`@ts-expect-error`, so nothing was laundered into the sibling
assertion rules.
Suppressions
------------
42x `// oxlint-disable-next-line anti-slop/no-reflect-get` across 38
files. Every one is the default-forward branch of a `Proxy` `get` trap:
get(target, property, receiver) {
...
return Reflect.get(target, property, receiver)
}
`Reflect.get(target, property, receiver)` is the only construct that
forwards with correct `receiver` semantics; `target[property]` invokes
an accessor with the wrong `this` and silently breaks getters that read
sibling state. There is no typed alternative, so these are suppressed
rather than rewritten.
3x `// oxlint-disable-next-line typescript-eslint/consistent-type-definitions
-- declaration merging requires interface` in
tests/e2e/github-url-smart-input-transition.spec.ts,
tests/e2e/linear-url-workspace-entry.spec.ts, and
tests/e2e/worktree-active-delete-scroll-position.spec.ts. Replacing
`Reflect.get(window, 'x')` with typed `window.x` requires a
`declare global { interface Window }` block, and `interface` is
mandatory for declaration merging. Matches the existing convention at
tests/e2e/helpers/runtime-types.ts:63.
1x `// eslint-disable-next-line no-var -- main-process gate handle for
this spec` in tests/e2e/project-group-creation-visibility.spec.ts, for
the same reason a `var` global is needed to type the handle. Matches
tests/e2e/agent-session-log-tail-stability.spec.ts:24.
Also updates two source-text anchors in mobile's rpc-recording mutation
harness (mobile/src/test-support/rpc-recording/operation-mutations.ts
and recording-runner.test.ts), which pin the exact text of the rewritten
line in settings-read-operations.ts and would otherwise fail with
"Mutant anchor matched 0 sites, expected 1".
|
||
|
|
f55b7ba680 |
fix(native-chat): cancel pending prompts precisely (#20601)
* fix(native-chat): hide activity while awaiting input * fix(native-chat): keep approval turns cancellable * test(native-chat): satisfy split PR quality gate * fix(native-chat): catalog approval cancellation label * fix(native-chat): include approval cancellation runtime label * fix(codex): settle prompts when cancelled turns complete * fix(codex): settle prompt registry fallbacks * test(native-chat): cover pending interaction fallbacks * test(native-chat): split prompt state coverage * test(native-chat): keep prompt state isolated * fix(native-chat): bound prompt turn backfill * refactor(codex): centralize prompt registry bounds * fix(native-chat): cancel pending prompts precisely * fix(native-chat): consolidate capability imports * fix(native-chat): harden precise prompt cancellation * fix claude cancellation teardown races * retry claude prompt lifecycle admission * bound claude prompt cancellation retry work * fix(codex): bound prompt turn identity on registration * fix(native-chat): route rejected late dispatch settlements * fix(codex): retain exact cancellable prompt turn ids --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
955051ded0 |
fix(codex): settle a structured send on admission, and stop minting a colliding identity (#20138)
* fix(codex): settle a structured send on admission, and stop minting a colliding identity Two sends could be written into the journal under one durable identity. Codex coalesces a mid-turn `turn/start` into the running turn rather than refusing it -- measured against real `codex app-server` builds 0.147.0, 0.150.1 and 0.153.4, none of which refuse and none of which fire a second `turn/started`. The dispatch path read the turn id from the turn/start response and stamped every accepted send `ordinal: 0`. Since a coalesced send gets the running turn's id back, two submissions persisted the same `providerItemId`. That string is durable, and it is the key a restore uses to match a submission against provider history, so the second message's real history row matched nothing and rendered as an extra bubble on replay. On 0.147.0 it is worse than a collision: the coalesced response returns a turn id that never starts and never completes, so the persisted key named a turn absent from history and NEITHER message could match. Identity is now minted from the echoed user message at `identityFor` -- the single point that mints the journal row's own identity -- so the settled key is by construction the one replay computes, rather than a parallel calculation that can drift. Dispatch returns `admitted` when the transport write completes; identity settles on the echo through a channel that did not previously exist for Codex. Waiters are keyed by client message id instead of being shifted off the front of an array by arrival order, and they are cleared on session close and child exit -- previously a timeout was the only thing that ever ended one. `TURN_ID_WAIT_MS` is deleted. It was never reachable on any build measured: `readCodexTurnId` returns non-null on all three, so the 10s wait never fired. The comment justifying it claimed older builds acknowledge before the id exists, which no tested build does. Three comments asserting Codex answers a mid-turn send with `turn already running` are corrected. Their only backing was a test fixture inventing that error string. The correction is factual only -- every changed line in `src/main/runtime/orchestration/` is a comment, and mid-turn delivery is still refused for both providers. Whether that policy is right is a separate question; it was resting on a false premise. Known gap, stated rather than implied: this prevents new collisions and does not repair journals already written with a colliding or phantom key. Those conversations keep duplicating on restore. Repairing them means re-matching persisted submissions against provider history and rewriting `providerItemId` -- which is what `journal-submission-reconciler.ts` is written for, and it still has no production caller. * test(codex): drop the synchronous-accept contract and the colliding `:0` from the integration fakes Three tests in the structured-session integration suites encoded the dispatch contract this branch replaces, and two of them pinned the defect it fixes. They asserted `agentSession.send` answers `dispatchState: 'accepted'` carrying `providerItemId: codex:<thread>:<turn>:0` at send time. That ordinal was never observed; it was stamped on every accepted send, which is exactly the collision this branch removes -- a send coalesced into a running turn is answered with the running turn's id, so two submissions persisted one durable key. The visible failure was a 30s timeout rather than a failed assertion. The fake client advertised no `agent-session.pending-send-result.v1`, and without it the host holds the reply until the send settles: a shim for clients too old to render a pending bubble. The fake provider then echoed the user message with no `clientId`, so nothing could correlate that echo back to the submission, and the wait ran to its own 30s ceiling. Real Codex sends `clientId` on that echo, and the fake now does too, which is what makes it a model of the provider rather than a sketch of one. The identity assertion is kept rather than dropped. Each send now asserts `pending` with no identity at admission, then asserts the submission settles `accepted` at `codex:<thread>:<turn>:0` once the echo lands. Same ordinal, but earned from `identityFor` on the echo -- the key a replay recomputes -- instead of guessed from the turn/start response. Ablated: removing `clientId` from the two echoes leaves both submissions `pending` and fails both assertions, so the assertion is load-bearing and not satisfied by something incidental. Both suites' client fixtures now advertise the capability set the desktop renderer sends in `src/main/ipc/runtime.ts`, which is what these suites mean by a client. The older-client settlement wait keeps its own coverage in `src/main/runtime/rpc/methods/structured-agent-session.test.ts`. `structured-agent-session-runtime-exit.test.ts` asserts `pending` for the same reason; it drives the host directly, so it never took the compatibility path, and what proves delivery there is still the turn the reacquired provider starts. The replay suite's "without dispatching it twice" property is untouched: one `turn/start` call, one replayed ledger row. * fix(codex): preserve unsettled dispatch correlations * test(codex): type the dispatch fixtures instead of asserting over them main's new casting gate (#20367 base) flags type assertions on changed lines. Replace them with checked types: the recording sink already satisfies its interface, both CodexSession fixtures are now annotated and carry real collaborators, the settlement assertion compares whole identities, and the integration helper reads submissions through the host's public journalSnapshot instead of its private session map. * fix(test): merge the duplicate doubt-reasons import the merge left behind Both sides added an import from journal-dispatch-doubt-reasons and the merge kept both statements, which the whole-repo native plugin gate refuses under --deny-warnings. * test(codex): a Fast mode turn is admitted, not accepted #20506 landed its Fast mode tests against the dispatch contract this branch replaces: a Codex send now returns admitted and settles its identity on the provider echo. The tier assertions the test exists for are untouched. --------- Co-authored-by: Merge Sim <sim@local> |