mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 16:02:29 +00:00
4e64fa9940f5a4a9bfb7b021a3c405fd76643ca5
973
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ca4e239861 | Remove low-value test inventories and duplicate fuzz oracles (#25791) | ||
|
|
ca2cdc0ad9 |
feat(native-chat): answer /context in chat view for OpenClaude and OMP (#22294)
* feat(native-chat): answer /context in chat view for Claude Over a terminal-backed chat, Claude's /context paints a grid only the hidden terminal sees and records nothing in the transcript, so the chat showed nothing. The chat host now answers it itself: the transcript decoder keeps each assistant response's API usage and model, and the composer replies with the same estimate the CLI statusline shows (input + cache tokens against the model's window) instead of typing the command into the PTY. The command is also listed in the curated Claude catalog so the slash menu offers it. * fix(native-chat): only state a /context window the host can establish The /context answer sized the window from the transcript's model id, which never carries the [1m] suffix, so a 1M session read as 200k (e.g. 450k / 200k, 225%). The window now comes from the session's tracked model resolved through the host CLI's own model listing (the resolvedModel field the list_models probe already returned), and only a [1m] id that matches the model the transcript last answered with is sized; anything else reports the used figure alone. Also: - Offer /context only in the desktop terminal-backed composer; the shared catalog is mobile's too, and mobile cannot answer it. - Answer a typed /context even with images attached, as the picker does, and keep the attachments armed instead of sending them to the terminal. - Report nothing until a response lands after a /compact or /clear sent from the chat, instead of the pre-compaction size. * test(native-chat): pin resolvedModel through Claude model discovery * fix(native-chat): say /context is unavailable when the host never reports usage A host that predates usage decoding, or a scraped view, decodes Claude's replies without a model or usage, so the chat promised an answer after the next response that would never come. Once the agent has answered and no answer names its model, say usage is not available for this session. * feat(native-chat): answer /context for OMP instead of Claude Claude and Codex answer /context themselves in the structured chat lane, so the terminal-backed Claude answer, its [1m] window rule, and the resolved-model plumbing through the model listing go away. OMP's /context only paints a panel inside its TUI. The OMP decoder now keeps the provider, model and usage of each reply (none for aborted or errored replies, which OMP's own gauge skips) and turns compaction rows into the existing compaction boundary. The OMP model listing keeps each model's context window, and the chat answers /context from the last reply's prompt size against the window of the provider/model that served it. * feat(native-chat): show OpenClaude's own /context report in chat OpenClaude writes its /context grid to the transcript as a <local-command-stdout> row, which the chat hid as harness noise. The live message preparation now turns the stdout row that answers an opted-in command (/context for OpenClaude) into command output with the ANSI colors stripped, and the notice row renders command output in monospace so the grid keeps its columns. Replies to other local commands stay hidden, and Claude is not opted in. OpenClaude's slash menu now lists /context; Claude's does not. * test(native-chat): pin OpenClaude /context output through live message preparation * test(native-chat): pin the OMP context window through model discovery * test(native-chat): pin monospace rendering of command output * test(native-chat): restore the notice row suite the command-output test replaced The previous commit rewrote NativeChatNoticeRow.test.tsx wholesale, deleting main's compaction, tone, plan, provider-notice and old-reader coverage. Restore it and pin monospace command output through the message row instead. * fix(mobile): keep OpenClaude /context off mobile, which cannot show its report The shared slash catalog now lists /context for OpenClaude because the desktop chat surfaces its transcript reply. Mobile's transcript fold does not, so the command did nothing there. Mobile now offers the catalog minus commands answered only by a surfaced reply, which returns its OpenClaude menu and send classification to main's behavior. * refactor(native-chat): drop the unread estimated flag from context usage Left from the earlier design; the answer text states the estimate itself. * fix(native-chat): keep /model replies hidden after an OpenClaude /context Skill surfacing ran first and rewrote non-catalog envelopes such as /model into plain user text, so command-output pairing no longer saw them. A later /model reply then paired with the earlier /context envelope and appeared in the chat. Pair outputs against the raw envelopes before skill surfacing. * fix(native-chat): word the OMP /context answer for a compacted session and show it as an aside After a compaction the agent has already answered, so the unknown-usage line now promises the next response instead. The host answer is one sentence, so it draws as the muted line it replaces rather than a monospace grid. The provider/model selector now comes from one helper shared by the OMP listing parser and the answer, and new tests drive real OpenClaude and OMP transcript lines through the decoders. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): pair command replies by row link and declare /context replies on the catalog A `<local-command-stdout>` row now answers the command row its `parentUuid` names, carried on decoded messages as the optional `parentId`, instead of the newest command row at or before its timestamp. A reply with no loaded linked row (older host, command outside the tail window) stays hidden. How each agent's `/context` is answered is declared once, as `reply` on its catalog row (`composer` for OMP, `transcript` for OpenClaude). The slash menu, the composer's answer, the transcript surfacing and the mobile catalog all derive from it, and both send paths share one intercept. Mobile no longer lists OMP `/context`, which did nothing there. * perf(native-chat): skip command-output surfacing for agents that declare no transcript reply Every Claude-format row now carries its parent link, so once a session holds any local-command reply (/model, /compact) the surfacing pass built a map of every message on each live update and changed nothing for agents like Claude. Agents whose catalog declares no transcript reply now return before the scan. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
b354aab18e |
fix(mobile-chat): a failed message's text is added to the draft, and an expired resend no longer blocks that text (#25149)
* fix(mobile-chat): a returned message is appended to the draft, and an expired resend id is spent * test(mobile-chat): leave the idempotent append to the shared rule's own test * docs(native-chat): say a whitespace-only draft counts as empty * docs(mobile-chat): don't overclaim where returned text survives |
||
|
|
761d63a4e5 |
feat(agent-launch): keep long prompts off the launch line and paste them after readiness (step 1 of 7) (#24257)
* feat(agent-launch): host-side prompt delivery for agent.launch The host's agent.launch typed any launch prompt into the shell as part of the launch command. A long or multi-line prompt then ran line by line in the shell, and an agent that never showed readiness or crashed at startup had nothing guarding where its text went. agent.launch now carries a prompt on the typed line only when the line stays one line, control-free and at most 512 bytes; otherwise the agent starts clean and the host pastes the prompt once the agent's own ready signal fires (bracketed paste plus its composer marker or a quiet render, read only after the shell's last hand-off, never while the pane's own shell is proven in front), with main's draft-paste bytes and an Enter 50 ms later. Orchestration worker starts wait on tui-idle as before. A replay-safe launch admits and claims its ledger row in one write, Qwen Code gets a second Enter, the desktop and phone share one launch-refusal classifier, and hosts advertise agent.launch.prompt-carry.v1. Split out of #23748, which moves the desktop source-control buttons onto this path. * fix(agent-launch): keep a short-lined multi-line prompt on a local zsh launch line, as main did #24257 moved every multi-line or over-512-byte prompt off the typed launch line and pasted it after readiness. The phone's AI buttons and review notes, whose multi-line prompts main typed whole into zsh, then reached Claude 0.5-3 s later and their RPC reply waited for the paste. The host now names the shell a local macOS or Linux line is typed into, the way the spawn picks it, and a multi-line prompt rides a zsh line when every line is at most 512 bytes and the whole line at most 8 KB. A real-zsh test types such a line through Orca's own ready barrier and startup write, including when a slow user config makes the write land early. Elsewhere the measured unsafe cases keep the paste: bash 3.2 runs multi-line lines piecemeal, fish drops an early multi-line write, and any shell loses a line over 1 KB written early. * test(agent-launch): keep the real-zsh launch-line test out of the Windows lane's gate scan The Windows lane registration check read `const ZSH_PATH = process.platform === 'win32'` (the head of a multi-line ternary) as a Windows-true flag, so `describe.skipIf(!ZSH_PATH)` looked like a Windows-only suite. The file is POSIX-only; the zsh lookup is now a function. * refactor(protocol): move the agent.launch capabilities into their own module Main's protocol-version.ts sits at the 300-line cap, so the prompt-carry capability pushed it over. The four agent.launch capabilities and their doc move to agent-launch-runtime-capability.ts, re-exported by name and spread into RUNTIME_CAPABILITIES at the same position; the advertised lists and every export are unchanged. * refactor(protocol): import the agent.launch capabilities from their own module `export *` from protocol-version left the four names undefined under the mobile recording loader, which resolves a relative import through a Proxy with no own keys, so 37 phone recordings lost agent.launchReplay. Importers now name agent-launch-runtime-capability directly; protocol-version only spreads its list. * refactor(agent-launch): drop the unshipped viewMode field and trusted local caller id Both were inert in step 1 and existed only for step 2. agent.launch will become a public plugin API, so every wire field is permanent once shipped; a top-level viewMode reads as "choose terminal vs chat", which the host decides. Step 2 introduces placement and view intent under a placement object instead. * fix(agent-launch): read Codex's provisional startup from the rule files' hold anchor Main (#24375) moved Codex's provisional-header check into codex.json's provisional_startup hold anchor and deleted codex-terminal-readiness.ts, so the launch readiness hold now asks showsHoldAnchor, as main's own settled check does. * fix(agent-launch): hold rule-file name titles to quiet for a launch, and census the zsh fixture Main (#24375) answers a name-only title from each agent's rule file ahead of the sustained-title lane, so gemini.json's name_title settled a launch readiness wait on the shell's auto-title while Gemini was still booting. A launch now asks quiet of every weak idle verdict, as that lane did. Main's readiness census requires a recorder for every runtime fixture; the zsh prompt recording is a non-agent control. Gemini's synthetic baseline is regenerated for this PR's stated change: a bare gemini title is no longer its rest mark, so name-only rows settle weak, and a fresh working or blocked status is no longer overridden. * fix(agent-launch): paste a launch prompt only when the launched agent is proven in front A launch pasted its prompt unless a shell was proven in the terminal's foreground, so any read that could not prove one let the prompt through. After an agent exited at startup, its shell turned bracketed paste on at the next prompt, readiness fired on it, and the prompt was typed into the shell: - macOS: a pane runs its shell under login, so the process-group fence's root was never the shell's group and never proved it; the cached foreground name could also still name the exited process. - Windows Git Bash and WSL: the shell-alone-in-its-job check never answers. Now one fresh read of the terminal's foreground decides: agent, shell or unknown. Only 'agent' lets a write through (paste, Enter, second Enter, reused panes too); 'shell' still drops a ready signal. A Windows host never proves the agent, so there the launch line carries the prompt at any size, as on main. * test(agent-launch): cover the Windows QA stub, a grok override that exits at once * fix(agent-launch): keep the local socket alive while a prompted launch waits for its agent A launch with a prompt now waits up to 60 s for the terminal agent to be ready before it writes the prompt, and reports not-delivered when the agent never is. The local runtime socket closes a connection idle for 30 s unless the request is a long poll, so a launch whose agent exited at startup lost its reply and the caller saw 'runtime closed the connection' instead of not-delivered. Classify a prompted agent.launch and agent.launchReplay as a long poll, as orchestration.workerStart already is for the same wait. * refactor(agent-launch): narrow the launch params by 'in' instead of a cast * fix(agent-launch): find a launched agent behind a wrapper that leads its process group A tcsh or nu launch line runs the agent from /bin/sh '<script>', and a wrapper script that does not exec its agent does the same: the wrapper leads the terminal's foreground process group and the agent is a member of it. The fresh foreground read names the group's leader, sh, so a prompted launch was refused or pasted late (M4Air tcsh: 2 of 4 not delivered, 2 pasted ~9 s late). Before that read, take the host's process-group observation as positive proof when it names the launched agent among the foreground group's members and is younger than a ready signal's quiet window. It never proves a shell. * fix(agent-launch): judge the foreground-group proof by when its capture began, not how long ps took The age the host stamps on a process-group observation runs from the start of its whole-machine ps, so on a loaded Mac a capture begun after the read was asked for still read as older than 1 s and the proof was dropped. Count an observation whose capture began after the read was asked for, less the window a shared capture is reused across. * test(agent-launch): keep the crash-guard live test out of the Windows lane's gate scan The Windows-lane registration scan read the const assigned from a platform check as a Windows-only gate, though the suite runs everywhere but Windows; find zsh in a function instead, as the real-zsh typed-line test does. Under load the fresh foreground scan can fail to answer, which lets the shell's prompt settle readiness (2 of 4 paired runs). The guard still refuses that write, so assert the refused write, the property that must always hold. * perf(agent-launch): read a local pane's foreground from its own terminal, not the whole process table The foreground read that gates every launch paste ran the daemon's inspectProcess capture and then a fresh scan, each a whole-machine ps; the fresh one also waits for any capture already running before it starts its own. Measured here at load 5: 1.2 s a read (M4Air QA: 3.4-5.0 s, and worker starts 17.6-32 s against main's 9-12 s at load 25-84). On a local macOS or Linux host, take the pane's root pid from the provider's session inventory and run one ps limited to that pane's terminal. Its foreground process group decides: the launched agent or any non-shell member is the agent (a wrapper that did not exec its agent leads the group), a group of shells alone is the shell. Same pane, same verdict: 2.7 ms a read. SSH hosts keep the relay's observation and name. * test(mobile): re-measure the web app's script sweep after agent.launch's capabilities moved out of protocol-version The mobile web bundle check failed at 124 assets against a ceiling of 123. Main already sat exactly on that ceiling: its sweep table read 69 scripts at 16 routes while the tree builds 73, the whole margin of 4. This branch imports the agent.launch capabilities from their own module, so protocol-version is no longer pulled into the root layout and four other routes. That moves which routes share which modules, and the Qoder capability module, imported by protocol-version and the AI-vault resume path, no longer shares an importer set with anything, so it gets a chunk of its own: 74 scripts. The fence says to re-derive the bound rather than raise it, so the sweep is re-measured on this head (every prefix of the sorted route list). The worst route now adds 10 scripts (session), not 9, which moves the pinned shell crossing from 32 to 30 routes; main re-measured on its own lands on the same crossing. * fix(agent-launch): a worker's brief needs its agent found in front, and Grok's start answers on its composer A paired-server worker start whose agent exited at startup typed its brief into the server's shell, which ran it: the idle wait can settle on a shell back at its prompt, and the brief was written with no foreground read. Both worker-start paths now check before each brief write, as a launch prompt is checked: on a host that can find the agent in front it must be there; on one that cannot (Windows) a shell proven in front still refuses, and anything else writes as before. A Grok worker start waited ~10 s more than main: its only rest signal is its bare name, which a launch holds to quiet output, and Grok animates its logo for ten seconds after its composer glyph. A worker start for an agent whose rest signal is its bare name and whose composer draws a marker (Grok, DSH, mimo-code) now also answers on that marker, whichever comes first. |
||
|
|
98171a8934 |
fix(native-chat): never delete chat history on read; a chat Orca can't load says why once (#24576)
* fix(native-chat): an older Orca keeps a chat with newer content read-only, and an unreadable annotation costs only itself A body or lifecycle mutation of a kind this build does not know, inside a row of a known kind, now latches the chat read-only with every row kept, as a newer row kind or version already does. It used to read as damage, and the open deleted the journal from that row on. Each optional field of a body schema now declares what an unparseable value means: - droppable (an annotation): removed from the row in memory; the chat stays writable; - must-understand (a prompt's questions, a plan approval's plan, a turn's lifecycle, a goal change): the row is unreadable and the chat latches read-only. Only a required field that fails is damage, repaired as before. A test holds every optional field reachable from a body schema to a policy; the known body kinds are read off the schema. Prompt options and questions no longer reject a key this build does not know, which deleted the journal from that row on. The context-usage drop is now one case of the droppable rule. * feat(native-chat): a chat an older Orca keeps read-only says so before a send is refused A chat a newer Orca wrote opens read-only, and until now nothing said so: the person found out when a send failed with "Chats were saved by a newer Orca". The host now names the reason on every whole history page it serves (`page.readOnly: 'written-by-newer-orca'`: hydration, history and catch-up resets), the shared client reducer keeps it from the latest whole page, and the desktop status strip and the phone's composer area say "This chat was saved by a newer Orca, so it's read-only here. Update Orca to continue this chat." Every read-only latch the host has is a newer Orca's (its database, a row version, a row kind, a body kind or must-understand fact), so the reason is derived from the journal's latch, not stored. The field is optional: older clients ignore it (cross-version test against the newest release's reducer), and older hosts never send it, so a new client against one says nothing, as today. A reason this client does not know shows nothing rather than words that may be wrong. The window-merge helpers move out of the client reducer into their own module. * fix(native-chat): only a newer build's value goes read-only; damage repairs as before An older Orca keeps a chat read-only only on evidence a newer build wrote it: a value outside a closed set this build knows (a body kind, a nested discriminant or literal, a mutation kind), read off zod's issues. That evidence wins over damage beside it. Anything else that fails (a wrong type, a missing field, a bad optional value) is damage, repaired as before, so damage no longer freezes a chat behind an "update Orca" it cannot fix. Drops the per-field droppable/must-understand policy. A lifecycle mutation's kind is decided before its item id, so a newer kind without one is kept. The context-usage drop is unchanged from main. * fix(native-chat): a read-only chat locks its composer and says why there Removes the separate read-only line (desktop status strip, phone notice component). The host's `readOnly` on whole pages now feeds the existing composer lock instead: desktop disables the composer with "Saved by a newer Orca. Update Orca to continue this chat." as its placeholder, and the phone adds a 'read-only' input-lock reason with the same words. A read-only chat shows no prompt card it would refuse; the draft stays in the composer. The phone's send-failure line is back exactly as on main. * fix(native-chat): a read-only chat pages back through its history A read-only journal answered every history request with a reset, so a chat this PR now keeps read-only more often could show only its newest page. Backward reads now serve the snapshot the open folded, as the first page already does; a forward read of rows still resets. * test(native-chat): an unknown key in a question entry is read and kept * test(mobile): the read-only overlay test typechecks * chore(native-chat): satisfy the changed-code gate in admission and the read-only overlay test * fix(native-chat): a new value inside a known block or goal arm reads as a newer build's The open fallback arm of the block and goal unions now aborts, so zod reports the known arm's own failure instead of only the fallback's; a future closed set there reads unreadable, damage there still repairs. The schema headers name the context-usage exception and say a new open-string value must be safe for every older build (rewind keeps only known states; pages carry no row version). * fix(native-chat): a read-only chat reconnects for a whole page, so an updated host unlocks it The client keeps `readOnly` until a whole page replaces it, but reconnected at its cursor and got only batches, so a desktop attached to a host updated in place stayed locked with "Update Orca". While it holds the latch it now subscribes without a cursor; the snapshot re-derives the state on any host version. The phone already subscribes without one. * fix(native-chat): every desktop write control follows the one read-only fact The structured session's transport state derives one fact from the page's `readOnly`: the words for why the host refuses writes. Every write control reads it. While it holds, the chat has no turn and no work, as the host projects it, so the working row and Stop go away; a pending prompt stays shown above the composer with its card disabled; queued cards' Send now, Delete and Edit, the queue's Resume, the goal banner's actions, the model and option pickers, dictation and background-task Stop are disabled or withheld. The composer placeholder stays the one place the reason is said. `lockReason` now locks the composer by itself. * fix(mobile): every phone write control follows the one read-only fact The phone session derives one fact from the page's `readOnly`, the words for why the host refuses writes, and every write control reads it: no turn or work while it holds, so no Stop; the pending prompt stays shown with its card disabled; queued cards and Resume are inert; the model and option pickers stay shut; the composer field is not editable, so its placeholder keeps saying why, and the lock applies at once instead of after the transport lock's settle. * test(native-chat): the newest release hydrates a read-only chat and scrolls back through it * test(native-chat): a read-only chat reads idle even with a send its provider never answered * test(native-chat): composer lock test without type assertions * fix(native-chat): round-3 read-only fixes - An old /model or option request no longer reopens its picker when a read-only chat unlocks: the menu is keyed by the request alone, and a lock only keeps it from mounting open. - A disabled question card still steps through its questions; only answering is locked. - A failed send's notice in a read-only chat says only that the message was not sent, with no Retry: the composer already says why, and no retry can land. Its wiring moves into useStructuredAgentSessionDeliveryNotices. - The host's background-task stop flags stand; the client no longer overrides them. - The view reads the read-only fact through one local flag. - The schema header says what an older build's rewind does with each open string. * test(native-chat): count picker mounts without an effect * fix(native-chat): read-only keeps other failures' words and shuts open option menus - In a read-only chat, only a send the newer-Orca refusal stopped is shortened to "Your message was not sent."; any other saved failure keeps its own words, and none offers a Retry. - A model or option menu open when the lock lands has its items disabled, so it cannot send a change the host would refuse. * test(native-chat): lock-lands picker test uses a typed surface * fix(native-chat): a locked composer draws its reason The editor's placeholder used the default that draws nothing while the editor is not editable, so a read-only chat's composer was empty and its reason lived only in the aria-label. The placeholder now draws while disabled, and its words are empty unless the composer is locked with a reason, so every other disabled composer (a pending prompt, no terminal) looks as before. * fix(native-chat): sync the locked-placeholder flag in an effect, not during render * fix(native-chat): a read-only chat is not offered for resume or counted as failed to resume After a relaunch, a chat whose journal this build keeps read-only (a newer Orca's) was listed in "Resume interrupted chats?", failed with "Orca couldn't resume this chat. Open it to continue manually.", and left a status-bar "N chats failed to resume" that never cleared, since nothing can continue a chat this build cannot write. The resume set now skips such a chat after the checks that end an offer (fork, moved on), so the offer is kept, not spent, and an updated Orca offers it again; the failure ledger keeps an already-filed failure for it but does not show it. Both read the chat's open journal, which listing and acting open first. * fix(native-chat): a refused resume continuation keeps the refusal's reason A send refused while continuing a chat after a restart was filed with its code alone, so a newer Orca's refusal lost its `journalWrittenByNewerOrca` reason. The continuation now carries the whole refusal, as a refusal from the agent's start already did, so the filed failure keeps it. * fix(native-chat): a newer Orca's chat keeps its resume offer however it is met - A newer Orca's whole database counts as read-only for resume, so a chat whose journal cannot be opened at all (a table the newer schema changed) is not offered, run or counted either. - The checks made right before sending no longer skip read-only chats: turning one away there was read as the user having moved on and deleted the offer. Its send is refused instead. - Settling a resume files nothing for a newer Orca's refusal; the existing rollback reopens the offer for an updated Orca. - The continuation records the refusal as a reference, without the wire prose. * chore(native-chat): one import of the session wire types in the continuation * docs(native-chat): resume and read-only comments match the current rule * test(native-chat): a closed set of a journal row cannot change without the row version The test walks the body schema and the row and mutation kind tables, collects every closed set (a row kind, a mutation kind, each enum, literal or discriminant), and compares them with a snapshot recorded beside the row version. A new value makes older builds go read-only, so it fails until the reader ships first or `v` is bumped and the snapshot updated. Tags a catch-all arm accepts as any string (block types, goal states) are listed apart: older builds read a new one as-is. * feat(native-chat): a body or plan subject of a newer kind is kept and the chat stays writable An item body kind and an approval subject kind this build does not know now read through the same catch-all blocks and goal states use (`openDiscriminatedUnion`), instead of making the chat read-only. The item is kept, drawn by nobody, and ignored by everything that reads items (turns, prompts, status); a rewind carries it as it was, kept by its place like every other row. An approval whose subject this build cannot draw shows its `detail`, which Orca's writers fill with the subject's text. A sent message of a kind this build does not know stays a newer build's. * fix(native-chat): rewind narrows a retained body by its kind before normalizing it * fix(native-chat): one chat that cannot open no longer stops the startup restore of the rest * fix(native-chat): a damaged chat fails to load and nothing is deleted A row this build cannot parse, a gap, or an epoch without its first row used to be repaired on open: every row from the damage on was deleted and a rebuild from the provider transcript was attempted. The open now fails through one refusal (failLoadOnJournalDamage, journalCorrupt), every row stays, and a damaged per-chat file is kept whole and never copied. A newer build's row still wins and keeps the chat read-only. The body classifier for closed-set values, the repair marker and disclosure, and journal recovery from the transcript go. * chore(native-chat): no reset or adapter left on the open's fixtures * feat(native-chat): an approval of a newer Orca's subject kind can only be cancelled Desktop and phone show its detail and the line "This request needs a newer version of Orca.", disable every answer, and keep the card's cancel, which ends the turn. The chat stays writable. * test(native-chat): the closed-set guard says a missed version bump fails older builds' load * test(native-chat): the closed-set guard names what a missed version bump does * fix(native-chat): a row the reader rejects is never written, and a chat that cannot load keeps its tab Every journal insert reads its serialized row back with the reader inside the write's transaction and refuses one it rejects, so Orca's own writer can no longer leave a chat that fails to load. A restored chat whose open fails keeps its tab and says why when opened. An epoch named with no rows is founded afresh again, deleting nothing. The at-rest test now expects a damaged chat to be refused; the replay gate keeps its 2026-09-11 evidence as it was run. * fix(native-chat): a newer Orca's approval subject is carried as it was, and its card cancels with or without a turn The approval subject's type is open, so code that reads a plan narrows first; the host's prompt bounding carries an unknown subject instead of rebuilding it as a plan, which threw in every start's stale settlement. A Codex rewind keeps a newer Orca's item in its place. A whitespace-only tag is damage. The card's cancel reaches the host without a running turn on hosts that take that. A provider row the reader rejects ends its turn as interrupted and the next send works. A guard fails if an approval subject kind is added before clients can be gated (STA-9262). * fix(native-chat): a card's cancel sends nothing without a turn again, and rewind keeps a newer row after any held row The host's cancel request has to name a turn on every version, so the turnless prompt cancel is reverted on desktop and phone. The rewind merge moves its anchor for every row the provider still holds, so a newer Orca's row stays after one whose kind this build does not know. * fix(native-chat): a card this build cannot answer leaves the composer open, and a send starts a turn An approval of a subject kind this build cannot draw, raised by a newer host's background agent while no turn runs, left nothing to press: its answers are disabled, its cancel needs a turn, and the desktop composer gave the card its slot. A send the host queues waits behind any pending prompt, so even the phone's open composer only queued. While every pending prompt is one this build cannot answer, the desktop composer stays open, and desktop and phone send without asking to queue: the send starts a turn, and the card's cancel then settles it. A chat a dead run left is unaffected: opening it already cancels those prompts. * feat(native-chat): a chat a newer Orca saved fails to load and says to update, with nothing to send into Opening a chat a newer Orca saved (a row kind, batch-change kind or sent-message kind this build does not know, a newer row version, or a newer history store) used to open it read-only: the history shown, the composer locked with a reason, every control greyed. Now it fails the load, like a damaged chat, with its own words: "This chat was saved by a newer Orca. Update Orca to open it." Every row is kept. The read is final, so the pane stops retrying; reopening the tab reads again. A chat whose load failed for good (damaged, or a newer Orca's) offers no composer under the error, on desktop and on the phone, so the failure is said once and no send can be refused a second time. A restart offer for a newer Orca's chat is spent without a word: no failure filed, nothing counted. Removed with the read-only mode: the journal's read-only latch and every check of it, the history page's readOnly field (never released), the cursorless reconnect, read-only paging, the composer lock and placeholder, the greyed controls, and the resume exclusion. * test(native-chat): a newer Orca's records open no chat, and a status fake carries its submissions * test(native-chat): the status fake's cast says what the feed reads * fix(native-chat): a chat's read that failed for good takes the whole pane, over a loaded transcript too * test(native-chat): the final-read-failure test names its reasons * fix(native-chat): a newer Orca's chat keeps its restart offer, and its refused load is logged once * refactor(native-chat): the newer-Orca chats a restart listing skips live beside the candidate reader * fix(native-chat): a refused chat load keeps where it failed, so the log names it and a new reason logs again * chore(native-chat): the reducer's window merge back where main has it, and comments say a newer Orca's chat fails to load * test(agent-hooks): the rename test's journal fakes carry submissions, as a journal snapshot does * test(agent-hooks): the rename test's journal fakes say what the status feed reads * fix(native-chat): Dismiss all ends the restart offers this Orca lists, and keeps a newer Orca's hidden ones * test(native-chat): a resend against a newer Orca's store answers unknown, and the phone says nothing beside the read's words This build never opens a chat a newer Orca's database holds, so a resent send id there cannot be answered from its journal: it answers "outcome unknown", never a refusal and never a made-up record, with no second delivery and no write. On the phone, a send error left beside a read that failed for good is not shown; the pane keeps only the read's words. |
||
|
|
c2c7649849 |
fix(mobile): stop Android from selecting words while the chat transcript scrolls (#22871)
* fix(mobile): stop Android from selecting words while the chat transcript scrolls On Android every paragraph, heading, quote, code block and table cell in the native chat transcript was a selectable TextView. Android starts a word selection, with the magnifier, on a double tap or a long press, and two flicks in the same spot while scrolling a FlatList register as a double tap, so scrolling the chat kept selecting words. iOS is unaffected: its UITextView path arbitrates scroll against selection itself. Android now renders transcript text without inline selection: one gate in MarkdownText covers every selectable span, and the user bubble follows it. A long press on a message opens a sheet with "Copy message" and "Select text", the latter a screen whose only content is one selectable Text, so a selection can only start where the user asked for it. The message row owns that sheet and mounts it only while open. iOS and web keep their inline selection and get no long-press handler. Verified on a Pixel 10 Pro Fold (Android 17): an adb double tap on the transcript selects nothing, the same double tap inside "Select text" selects a word, tool rows inside the bubble still expand on tap, and the long press opens the sheet. * fix(mobile): scope the Android selection gate to the transcript and route span long presses Review follow-ups on #22871: - The gate now applies only where `rangeSelectable` is passed (the chat transcript). Task comments and file previews keep their selectable text on Android as before. - On the Android transcript, spans that take taps (links, file paths) also take the row's long press, so a link under the finger no longer swallows the copy/select sheet. - Copied text keeps its whitespace; only whitespace-only blocks are dropped. - The Android markdown test compares `String(node.type)` instead of a type assertion, which the changed-code quality gate rejects. * fix(mobile): route long presses on Markdown images to the row on Android An image block is a Pressable of its own, so on the Android transcript it now carries the row's onLongPress like tappable spans do; a long press on an image opens the copy/select sheet instead of being swallowed. Test extended with an image block. * test(mobile): pin the Android long press from a chat row to its actions sheet The existing row suite runs as iOS, where the bubble has no long press. This one runs as Android: the bubble's long press mounts the actions sheet with the message, the markdown receives the same handler, closing unmounts the sheet, and the user bubble carries no inline selection. * fix(mobile): preserve Android message selection while replies stream --------- Co-authored-by: Neil <neil@stably.ai> |
||
|
|
a154562a89 |
fix(native-chat): show a reply cut off by an Orca crash or quit like a finished turn, with one explanation (#25043)
* fix(native-chat): read a crash-cut reply like a finished turn, with one explanation
A reply that an Orca crash or restart cut off said it failed three times: the
turn bar read "Failed after N", the chat's notice row said Claude stopped, and
the sidebar dot stayed red after the chat was read.
The turn bar now reads "Worked for N" for a proven crash/restart cut, the
notice row stays the one explanation, and the sidebar card's dot and agent
rows read failed only until the user has visited the chat (the same
acknowledgement that un-bolds the row), then read done. A failure, a user's
Stop, a replaced turn and an unconfirmed end keep their labels. The stored
outcome is unchanged.
* fix(native-chat): explain a turn a quit or eviction cut, once
A turn cut off by quitting Orca, an idle eviction or a teardown recorded the
same outcome as a crash-cut turn but wrote no notice row, so with the turn bar
now reading "Worked for N" nothing in the chat said it stopped.
The host stop's settle now writes the existing providerExited notice for the
latest turn when it ends as news (no person's Stop decided it): in the same
write as the host's own turn end, or right after the adapter's. It is keyed by
the turn, and skipped when an error row already explains that turn, so a retry
or an earlier exit row never leaves two.
* test(native-chat): narrow the turn scope and add the seen map in the crash-cut tests
* fix(native-chat): derive a cut turn's one notice on read instead of writing it at stop time
A turn cut short when the agent stopped without anyone asking now reads
"Worked for N", so it needs a row saying it stopped. Writing that row only on
the quit and eviction paths missed journals written before this change, a quit
that died between its drain and its settle, and any future stop cause.
The transcript now derives it from the journal on desktop and phone alike: a
root turn that ended interrupted with no verdict and no row about the stop
gets one notice in the provider-exit row's words. A stored exit row, matched by
its exit fact or its writer's identity rather than its tone, stays the
explanation, and so does the restart continuation's own outcome note, so a
refused resume says it once. The host's stop path is back to what it was.
* fix(sidebar): read a cut-short turn's red mark from its acknowledgement on every surface
A turn cut short with nobody asking reads failed only until the user has seen
it. The previous revision passed an optional "seen" flag to each caller, so any
surface that did not pass it, the chat's own tab dot among them, stayed red
beside a sidebar that read Done.
The verdict's display now takes the acknowledgement as part of what it reads:
the entry's stateStartedAt beside the time the user last acknowledged it, both
required, judged by one shared rule. Each surface joins it once where it builds
its rows: the sidebar's agent rows (and so the notes send menu), the workspace
card summary, the terminal tab bar, Cmd-J recent rows, and Activity threads.
Failures, Stops, replaced turns and unproven ends keep their marks as before;
the phone, which has no acknowledgement record, keeps the unseen reading.
* refactor(attention): use the one acknowledgement rule where it was copied
Auto-acknowledgement, the Activity unread count, dashboard row buckets and
notification acknowledgement each spelled the same "acknowledged at or after
the current state began" comparison; they now call the shared rule.
* test(mobile): type the cut-turn notice test's client and hook holder
* test: pin an older host's exit row by its writer and the sidebar rows' acknowledgement join
* test(sidebar): re-read the card and the tab when only an acknowledgement changes
* fix(native-chat): keep a cut turn's notice through a resume that carries on
A resume after a quit writes its "asked this agent to continue" note after its
own message, so counting that note as the cut's explanation removed the notice
once the resume went on, and made it flash away under automatic resume. Only the
notes that say the chat was not carried on (refused, not connected, not
confirmed) stand in for the notice now.
A row about the whole conversation now explains a cut turn only when no other
turn lies between them, and a failed start's row, which is about a start, never
does. The host writers and the rule share the row identity prefixes, with the
contract that a new row explaining a stop carries the provider-exited fact or
one of them.
* fix(sidebar): judge a cut as seen by the main agent's clock when a subagent holds the row
A subagent can keep a row working after its main agent was cut, and the row's
clock then predates the cut, so a look at the working chat counted as having
seen the cut and no red mark showed. The mark now compares the acknowledgement
with the later of the row's and the main agent's clocks; auto-acknowledgement of
the chat on screen reads the same clock, and its stamp covers it, so looking at
the chat still clears the mark. Bold rows, dashboard buckets, notifications and
unread counts keep the row's clock.
* fix(native-chat): tie a conversation-wide exit row to a cut only with no message sent since
A send after a quit's cut whose new agent died before its turn opened leaves a
provider-exit row about the whole conversation. That row is about the send,
not the earlier cut, so the cut kept no explanation. An exit row now explains
a preceding cut only when no message was sent in between; the restart notes,
which follow the continuation's own message, still need only no turn between.
The notes that say a resume did not carry the chat on are now told apart from
the continued note by their tone ('error' or 'warning'), which every host that
wrote them has set, instead of by their words.
* fix(native-chat): keep an owner's proven death the explanation of its cut after a send
A chat read before the startup reconcile settles its cut turn unverifiable;
if the user then sends a message, the reconcile proves the old agent dead and
writes its row about the conversation after that message. The row names the
old owner's death, never the send, so a message since no longer detaches it
from the cut. Only a provider-exit row, which a later start can write, still
needs no message sent since.
* test(sidebar): give the activity-status store mock the acknowledgement map
The card summary now reads acknowledgedAgentsByPaneKey; this test's hand-built
store state lacked it, so every summary read threw.
* refactor: move the seen-gated red mark out of this change
This change now keeps only the cut turn's label and its one derived notice.
The red mark that clears once the user has seen a cut turn moves to its own
change, so each can land alone; until it lands, the sidebar, tab bar, Cmd-J,
Activity and the notes send menu read a cut turn as failed, as before.
* test(native-chat): pin that a restart note after a continuation's turn leaves the cut's notice
* test(native-chat): keep a cut turn's notice beside a later message drawn as not sent
|
||
|
|
e347aa4e67 |
fix(mobile): honour the desktop pinned-worktree placement setting (#25301)
* fix(mobile): honour the desktop pinned-worktree placement setting Mobile always kept pinned rows in their groups as well as in Pinned. Desktop's showPinnedWorktreesInGroups (default off) shows them only in Pinned. Read that setting with its own settings.get operation, refreshed with the list's view settings, and leave host and device-local pins out of their groups unless it is on. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): keep lineage children with a pinned parent Under single-location a pinned parent left its group while its unpinned child stayed behind as a root. Visible descendants now follow a pinned ancestor into Pinned, nested under it, as on desktop. Drop the inert host-id half of the pinned-policy stale-reply guard; a host switch replaces the client. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): fold pinned placement into the desktop view-settings sync One "sync from desktop" callback now also reads settings.get for pinned placement, unawaited so a slow read never holds the ui.get merge; the separate callback and its plumbing are gone. Re-record the three host.view-settings goldens, which gain only that settings.get send. The pin expansion walks the lineage children index the renderer now shares, over visible rows only. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): one stale-reply rule for the desktop view-settings sync Both reads in syncViewSettingsFromDesktop now drop a reply only when the client was replaced; the host-id half compared a captured value with itself. Pin the no-wait invariant: the ui.get merge applies while settings.get never answers. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): follow a pin through rows hidden by search Desktop's Pinned section walks lineage over every worktree and keeps the visible descendants, so a grandchild still follows a pinned root when search hides the middle row. Mobile walked visible rows only and left it in its group. Build the children index from the unfiltered list; membership stays visible-only. Drop the hostId arg the view-settings sync no longer reads. Pinned-lineage tests move to their own file to stay under max-lines. * refactor(mobile): read pinned placement in its own hook The placement read was folded into the desktop view-settings sync, so that callback juggled an awaited and an unawaited read and three goldens recorded a settings.get that never answers. useHostShowPinnedInGroups now owns the setting and the catalog refreshes it beside the view-settings sync on connect, focus and mount; the sync, screen state and goldens are back to main. The reader returns desktop's boolean, as its siblings do, so transport no longer imports a worktree type and the two-literal policy union is gone. makeSection always applies lineage, the pinned walk is a Set worklist, and the lineage children index drops a set that duplicated its map's keys. * fix(mobile): key pinned placement to the client that reported it The host screen is reused across hosts, so the previous host's "show in groups" value survived a switch until the new read landed, or for good if it failed. The setting now carries the client that reported it and counts only for that client, which also makes a late reply from a replaced client inert and drops the clientRef plumbing. * fix(mobile): drop a replaced client's late placement reply Keying the value to its client hid a previous host's value but still let that host's late reply overwrite the current host's, reverting the list to the default until the next refresh. Restore the clientRef write guard; the client key still covers the effect-long window before clientRef follows a switch. * refactor(mobile): treat pinned placement like the screen's other host state The placement hook kept its own host-switch handling (a client-keyed value plus a clientRef guard) beside the screen's existing one. showPinnedInGroups now lives in HostScreenState, resets with the other host-scoped values in useHostScreenIdentity, and is read in the catalog refresh behind the same clientRef check as the catalog fetch. The hook and the catalog argument go. |
||
|
|
ebeea319b2 |
Support modern Qoder commands and verify authenticated resume
Start and resume Qoder through the existing execution-host selector when only the documented qoder command exists. Preserve legacy qodercli preference, explicit commands, quoting and session identity; disconnected SSH execution refuses without local fallback. Mobile history uses command-at-create only with the optional owned-create capability and an existing stable mutation identity. Reconcile authoritative execution-host inventory before retrying creation, adopt the same surviving operation, preserve WSL/incarnation metadata and restore only missing original launch metadata. Changed retry settings cannot replace original capture. Bounded evidence expires after 15 minutes; unavailable inventory or original evidence refuses recovery and leaves surviving execution untouched. Authoritative inventory proving absence preserves the existing recreation behavior; this is not a durable exactly-once ledger for completed one-shot side effects. Older hosts retain the acknowledged create-then-send path. Scope mobile launch authority to the current committed host/client generation, and recheck operation ownership after asynchronous preparation before later sends. Retire the mutation once the host acknowledges the resume; later ownership changes stop navigation without reporting a completed resume as failed or reusing a cached legacy pane. Preserve genuinely interrupted mutation identity. Preserve current-main OpenCode validation and merged Pi/Cursor behavior. Correct unchanged-main editor test fixtures to their production insertion-range and store contracts while retaining original assertions and production behavior. Credit: Neil Parker (@nwparker); Soperf and jyang for Qoder integration/history groundwork in #24614; actual Pullfrog and CodeRabbit reviews and the independent Source reviewer for the mobile retry, launch capture, evidence lifetime and connection ownership findings. Conservative positional process recognition and existing folder-history matching-worktree limitations remain documented. |
||
|
|
1978469fd2 |
fix(mobile): keep the working rings turning on the OTA page (#25299)
* fix(mobile): keep the working rings turning on the OTA page Animated.loop starts a native loop whenever the timing asks for the native driver; the web has none, so the JS fallback ran one turn and froze at 360deg. Ask for the native driver only off the web. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile): pin the native driver on native spinners Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): share the working ring rotation between both rings Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
baa56fd10d |
Simplify the phone-control and phone-size terminal dialogs (#25307)
* Redesign the phone-control and phone-size terminal dialogs Drop the eyebrow label and circled icon, shorten the copy so it no longer restates the buttons, and give each state one primary action with a quieter "all" action. Collapse moves out of the button row into a Minimize icon in the corner. Behavior is unchanged. * Point the phone settings copy at the renamed Restore button; drop dead ko overrides The phone app and desktop update independently, so name only "Restore", which matches both the old and new desktop banner labels. |
||
|
|
0971479866 |
Add shared workspace settings note and prevent filter modal expansion (#25300)
* fix(mobile): say that workspace sort, grouping, and filters are shared The Manual sort option was subtitled 'Server order', but it orders by the desktop's drag ranks. Sort, grouping, and filters on the phone all write the host's shared view settings, so changing them also changes every other device on that host, which the screen never said. Relabel Manual as 'Desktop drag order' and add 'Shared with other devices on this host' under the Sort By, Group By, and Filter titles. The note avoids naming a desktop sidebar because headless hosts have none. * fix(mobile): prevent filter modal heading expansion Add flexShrink: 1 to allow the heading container to shrink when space is constrained. Update comment to clarify why workspace view is shared across devices. * update wording |
||
|
|
97fa6aee74 |
fix(native-chat): show a message Orca accepted and then failed to deliver as "Not sent" in the chat (#24710)
* fix(native-chat): keep a message the host accepted then rejected in the desktop chat as not sent Draw it in place from the host's history, so a crash that loses the outbox no longer makes it vanish. A later copy of the same body supersedes it; the outbox row wins while it holds the message; the phone is unchanged. * test(native-chat): pin the same-id rule apart from the body match * test(native-chat): type the rejected-in-place fixture body as a text block * fix(native-chat): let the host's row own a message it recorded and then rejected Once the host's journal records a send as rejected, the desktop outbox lets it go, as it already does for delivered and Stop-withdrawn sends: the host's row shows it as not sent, with the host's reason and no Retry. The outbox keeps only sends the host refused before recording them, which keep their Retry. A send whose own reply says it was rejected is drawn by its outbox entry, with no Retry, until the journal carries the row; a copy left by an earlier session is dropped when the chat opens. - the transcript no longer hides a host row behind an outbox entry with the same id or the same text; those rules and their cache are gone - a rejected message the queue holds (a draft's hand-off, or a live card under its id) is drawn as its card, not as a row - a later copy of the same text hides a rejected row only when it was sent once the rejection was known, so a deliberate repeat stays - delivery notices read the same visibility rule as the transcript; a chat whose only rejection a Stop withdrew no longer rebuilds them per batch - the body fingerprint helper goes back to the host, its only user * test(native-chat): keep one row when copies of a rejected message share an instant * refactor(native-chat): let the host's notice replace the outbox's under the same id * test(native-chat): pass the queued card ids in the tool-stream cost transcript * fix(native-chat): keep the host's record as what lets a rejected message go - the outbox no longer drops a host-rejected message when a chat opens; the reconcile lets it go once the journal's submissions say it was rejected, and that drop is written to storage, so nothing reads as still owed - a message the host rejected while the chat watched waits for its journal row with no Retry; one read back from storage with no row loaded keeps its Retry under a new id, since the host may have lost it - the delivery notices keep the same map and notice objects across a batch that words every row the same, so a submission batch re-renders no row - a rejected command such as /compact stays hidden: its own reply reports it - the desktop transcript requires the queued card ids, with a controller-level test that a card holding a rejected message keeps its row hidden * fix(native-chat): draw a queued message where the host rejected it A message accepted to hand over later and rejected before any handover now sits at its rejection, as a handover places one: what the agent did while it waited happened before it, and the newest history page holds it. One handed over, or dispatched as it was recorded, keeps its place. An older host does not move it, so it stays at its submission, still drawn. A failed start now rejects the queued messages and writes its row in ONE journal append, the messages first: no reader ever meets one without the other, and the messages still draw above the row that says why. * test(native-chat): pin that rows written together roll back together * fix(native-chat): draw every rejected message where it was rejected Not only a queued message: one handed over into a turn and then rejected, or sent directly and rejected, also sits at its rejection, in no turn. A message in doubt stays where it was, a plain bubble: it may have reached the agent. * fix(native-chat): decide a rejected message's Retry from the host's stored fact - a message the host recorded and then rejected has no Retry on any mount, however that mount learned of it, and a Dismiss that clears it from storage; a send refused before the host recorded it keeps its Retry - the rule that keeps a rejected command such as /compact out of the transcript moves into the one visibility function rows and notices share - the outbox state docs say what lets a recorded message go: the client holding its rejected submission, whose row the host places at the rejection * fix(native-chat): write no start-failure row when a Stop withdrew every queued message first * test(native-chat): pass the Dismiss action in the delivery-notice hook tests * fix(native-chat): keep a rejected message's outbox copy until its row loads An older host leaves a rejected message where it was sent, which may be older than the loaded window: the chat then holds the rejected submission but not the row that draws it. The outbox copy now stays until that row loads, marked as the host recorded it (Dismiss, no Retry, in the host's words), and leaves once the page holding the row is loaded. Derived from the loaded rows each time. Tests that label their projection as the phone's now pass the phone's own setting. * test(native-chat): type the outbox hook props that carry loaded rows * fix(native-chat): write nothing when a journal batch settles nothing in the outbox The outbox re-reads the journal on every batch since it waits for a rejected message's row to load. Its reconcile now returns each unchanged entry, and the list, as themselves (a message left in doubt included), so a batch that changes nothing writes nothing to storage. The reconcile moves to its own module. A copy the host recorded and rejected owes no delivery, so it no longer keeps a hidden pane reading the journal. * test(native-chat): count storage writes on the outbox's own storage object * fix(native-chat): let a recorded rejected message's outbox copy leave on its own, with no Dismiss The outbox copy of a message the host recorded and then rejected draws it only while the host's row is not loaded, and leaves on the batch or page that loads that row. It owes no delivery and offers no control: sending it again is a new message. The Dismiss that let the user clear it is gone, from the outbox, the notices and the session controller. |
||
|
|
f199a20c3a |
Preserve OpenCode reasoning and recorded patches in native history (#24790)
* Use bounded OpenCode context for vault session continuation OpenCode database and synthetic row paths are not text transcripts. Use the vault preview or captured pane context, preserving actual transcript paths containing a hash and supporting both OpenCode lanes and Windows paths. Adapted the intent of #11859 and extended it to actual installed v2 vault rows. Co-authored-by: mrcha033 <mrcha033@users.noreply.github.com> * Read real OpenCode sessions in terminal-backed native Chat Reuse the bounded AI Vault SQLite worker for v1 and v2 session pages and live updates. Keep terminal input as the real execution path and pace OpenCode Stop through its two-Escape interrupt. Co-authored-by: xodmd45-ctrl <xodmd45-ctrl@users.noreply.github.com> * fix(opencode): publish approval cards for permission requests * Send OpenCode native approval through its Enter selector * Resolve mobile Chat readability for folder workspaces * Bound OpenCode part batches and preserve v2 image attachments * Prefer live migrated OpenCode sessions over legacy copies * Consolidate mobile Chat eligibility test imports * Consolidate OpenCode SQLite protocol type imports * fix(native-chat): preserve OpenCode reasoning and patch parts Separate genuine reasoning from answer blocks in both native SQLite schemas and retain recorded patches as completed patch tools. Keep each database row together at page boundaries so the existing raw-row cursors cannot drop half of a mixed row. Adapted from @akhan157's OpenCode native history work in #13287 at bb661d10d716764fb472d824cd434678875b1947; retains the current bounded reader and account discovery instead of restoring the older capture and cursor implementation. Verified against genuine private installed 2.0.16 and official 1.18.30 CLI ingestion. Co-authored-by: Adnan Khan <adnank11427@gmail.com> * fix(native-chat): keep split OpenCode rows intact on desktop and mobile Preserve the native reader's bounded OpenCode row groups in paired reads and snapshot/replacement frames so a second presentation-count slice cannot drop reasoning while advancing the database cursor. Sort derived reasoning before its answer under the same provider timestamp while retaining journal order. These two boundaries were reproduced with genuine installed 2.0.16 and official 1.18.30 sessions in a hidden desktop renderer and the current mobile view over an actual authenticated encrypted pairing. Completes the semantic presentation from @akhan157's #13287 without importing its older clipping or cursor implementation. Co-authored-by: Adnan Khan <adnank11427@gmail.com> * fix(native-chat): keep reasoning and answers together in live windows * Bound OpenCode transcript RPC pages and present omission notices * Bound OpenCode transcript RPC pages and present omission notices * Bound OpenCode transcript RPC pages and present omission notices * Update native worker oversized-history notice contract --------- Co-authored-by: mrcha033 <mrcha033@users.noreply.github.com> Co-authored-by: xodmd45-ctrl <xodmd45-ctrl@users.noreply.github.com> Co-authored-by: nwparker <nwparker@users.noreply.github.com> |
||
|
|
e1b046a1bd |
Skip new legacy file inventories after mobile search cleanup (#24792)
Reuse the existing mobile mounted-ref pattern only at legacy fallback entry after completed passive cleanup; preserve admitted work and all live search/authority/cache paths. Correct only the strict fully-unmounted inventory scenario and its sole golden. |
||
|
|
06194626ba |
Delete unreturned clipboard cache files after a failed write (#24599)
Reuse existing provider-copy best-effort deletion for a newly created clipboard cache file whose write fails before its URI reaches the caller. |
||
|
|
831710c380 |
Avoid restarting error timers after mobile relay pairing closes (#24568)
Check the existing pairing owner closed flag before allocating its missing-close fallback alarm. |
||
|
|
4498e099f6 |
Avoid restarting error timers on closed mobile relay links (#24567)
Check the existing irreversible closed flag before scheduling the existing missing-close fallback timer. |
||
|
|
1de8396f5b |
Skip new mobile toast work after feedback owner cleanup (#24759)
Reuse the existing mobile mounted-ref lifecycle pattern at toast presentation entry, preserving admitted clipboard outcomes and every live animation/sequence/timer operation. |
||
|
|
8b1dc63459 |
Release waiting terminal output when a mobile subscription fails to start (#24547)
Dispose the failed subscription record’s existing terminal backlog before rethrowing its original start error. |
||
|
|
4fdf6df25b |
Reuse the ancestor path while building mobile agent rows (#24539)
Preserve traversal order and cycle guards using one call-local path Set rather than a copy at every depth. |
||
|
|
84d246b9f0 |
fix(jcode): register in main's remote-installer guard, drop our duplicate
Rebasing onto 853 commits of main surfaced two things the earlier branch had hidden. main already owns a guard for the issue-#7253 bug class (`remote-hook-service-registry-coverage.test.ts`). This branch had added a second, near-identical one — a parallel implementation of a test that already existed, which is what AGENTS.md's reuse rule is about. Deleted ours and registered jcode in main's, which is the one that has kept pace with every agent added since. Also fixes a missing separator in the mobile icon map. `pnpm tc` does not cover `mobile/`, so only the session-route closure suite caught it. Co-authored-by: czzczz <chanzrz_zbf@foxmail.com> |
||
|
|
4049e63714 |
feat(jcode): add Jcode as a supported TUI agent with managed hooks
Ports PR #10521 onto current main: agent catalog, managed hook service, agent-status listener, session resume, AI Vault parser, per-pane daemon isolation, and Source Control AI support. Co-authored-by: Neil <neil@stably.ai> |
||
|
|
a2896f5470 |
Support real OpenCode sessions in native Chat (#24647)
* Use bounded OpenCode context for vault session continuation OpenCode database and synthetic row paths are not text transcripts. Use the vault preview or captured pane context, preserving actual transcript paths containing a hash and supporting both OpenCode lanes and Windows paths. Adapted the intent of #11859 and extended it to actual installed v2 vault rows. Co-authored-by: mrcha033 <mrcha033@users.noreply.github.com> * Read real OpenCode sessions in terminal-backed native Chat Reuse the bounded AI Vault SQLite worker for v1 and v2 session pages and live updates. Keep terminal input as the real execution path and pace OpenCode Stop through its two-Escape interrupt. Co-authored-by: xodmd45-ctrl <xodmd45-ctrl@users.noreply.github.com> * fix(opencode): publish approval cards for permission requests * Send OpenCode native approval through its Enter selector * Resolve mobile Chat readability for folder workspaces * Bound OpenCode part batches and preserve v2 image attachments * Prefer live migrated OpenCode sessions over legacy copies * Consolidate mobile Chat eligibility test imports * Consolidate OpenCode SQLite protocol type imports * Update native chat settings contract for both OpenCode agents * fix(native-chat): reconcile bounded OpenCode transcript reads * fix(native-chat): dispatch OpenCode questions safely * fix(native-chat): keep native discovery and transcript windows current * feat(accounts): link standalone GLM Coding Plans (#24618) * feat(accounts): link standalone GLM Coding Plans Adapt the reviewed GLM accounts contribution to current main, retain Antigravity behavior, guard late credential results, expose storage protection, and redact quota errors. Co-authored-by: Luchong <lu740528977@gmail.com> * fix(accounts): retain GLM credential results during quota refresh * fix(accounts): make GLM credential editing desktop-only * fix(accounts): mirror the host GLM site in paired clients * fix(accounts): report unknown GLM host details and split web settings tests Apply the independently reviewed Accounts correction from697284a without the v2 adapter commits. Preserve the saved-key store and serialized write behavior. * fix(zcode): ship required GLM account translation entries * chore: record GLM reconciliation hook validation * chore: validate installed GLM commit hooks * test: complete GLM account fixtures and web API inventory --------- Co-authored-by: Luchong <lu740528977@gmail.com> * fix(native-chat): route transcript requests through shared SQLite worker * fix(ci): prevent concurrent pnpm refresh during mobile typechecks (#24776) * fix(ci): run mobile typechecks without concurrent dependency refresh * test(ci): check effective Linux E2E package list * test(ci): preserve the mobile production compiler barrier --------- Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example> * test(terminal): restore the live fish fixture prerequisites (#24947) A restored pane waits for the initial status replay before subscribing to PTY output. This fixture never settled that replay, so fish printed its mode-2031 arm before the renderer connected. Its PTY API also omitted the reset-input listener required by the serializer, aborting attachment. Settle and dispose the existing startup-snapshot registration and provide the same reset-listener mock used by the other PTY tests. The real fish child-stdin assertions and timeouts remain unchanged. No production change. * fix(shortcuts): defer TUI editing chords in terminal-first mode (#24640) Restack the original focused change onto current main, preserving every owned source and test blob and the merged CI contract and journal cleanup fixes. Original-commit: |
||
|
|
ebe77028bd |
Check each project repository once while filtering mobile cards (#24540)
Reuse exact raw source/slug matching decisions within one project filter call, preserving the matcher, membership, negative matches, output order and identity. |
||
|
|
14ab734b0f |
Skip unused image-size calculations in mobile web browser requests (#24941)
Use the existing mobile density budget only for mobile view; pass the existing constant to the existing assembler for web/default mode, where that argument is discarded. No cache, policy, request or native path changes. |
||
|
|
843607b1bc |
Register supervised Qoder China and Qwen Code (#24616)
* Add Qoder session history and search with real CLI coverage * Allow the real Qoder marker file to end with a newline * Keep Qoder tool output out of history previews and search * Keep Qoder search pages readable by older clients * Verify persisted Qoder history after a real generated and resumed task * Negotiate Qoder filters before searching an older execution host * Combine search client imports for the CI plugin gate * Keep the relay search oracle aligned with legacy agent filtering * Register supervised Qoder China and Qwen lifecycle integration * Cover Qoder China mobile assets and mixed-host resume gates * Verify Qoder provider tags against the older released wire parser * Verify China and Qwen keep independent Windows hook scripts * Verify Qoder registrations against the installed older Windows release * test(qoder): align search capability contracts and pin old-host fencing * fix(qoder): rank exact picker identities and command aliases first * test(qoder): preserve the regional CLI shared icon expectation Keep the full bundled-asset and no-remote-image checks, with an explicit shared-logo basename for Qoder China. The map also works with older catalog type unions. * fix(qoder): align China catalog entry with fallback order --------- Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example> |
||
|
|
ac46d9efda |
fix(native-chat): plain wording for chat errors and status rows (#24594)
* fix(native-chat): plain wording for chat errors and status rows
Replaces Orca-internal words (host, journal, transcript, outbox, process,
unverifiable, "no contact", byte budgets) in native chat refusals, status
rows, the skills menu, subagent and background-task state labels and the
history-load error with plain language, and routes the two hard-coded
English composer errors through translate(). Copy only; no behaviour change.
* fix(native-chat): match chat copy to what happens and to the sidebar's words
- The held-message row says Orca keeps checking, which it does: the idle
sweep retries the unproven stop and a landed retry sends what waited.
- An unsettled earlier message reads as unconfirmed, not undelivered.
- The skills-unavailable line names SSH chats, its only cause.
- Subagent and background-task rows say "no recent update", the sidebar's
words for the same state, and "status unavailable" after a count; the
row no longer repeats the state as a reason.
* test(native-chat): find the repair row by its own text, not the old wording
* fix(native-chat): word the history-repair and too-large rows in the reader's language
Both rows were finished English the host wrote into the chat, so nothing could
translate them. Each now names itself with a presentation, the way the
compaction row does, and the chat says it through translate(). The English text
stays on the row for clients that predate these presentations and for the phone.
* fix(native-chat): say composer send errors with the chat's notice sentences
The composer's two errors had their own wording file beside the sentence table
every other chat notice uses. The send outcome now carries notice parts, worded
by the same function as the rest. A refused redelivery says "Orca couldn't
confirm your message reached the agent. Check the chat, then send it again if
needed." (the same sentence the failed-send rework uses), and a message this
client couldn't store says "Couldn't save your message. Try again."
* fix(native-chat): drop the retry line, keep one name for a lost task, and say only true causes
- The history error pane no longer adds "Orca keeps trying to load this chat."
under its title: the read still retries on its own, but the pane says only
that the chat didn't load.
- A write refused as unsupported asks for an Orca update only when no reason
came back, which means the host is older. A named reason (a location or agent
that can't run there, no chat host, a client missing the capability) now reads
"This isn't available in this chat.", since updating doesn't fix it.
- A task Orca lost track of is "no recent update" everywhere; after a count it
reads "2 agents with no recent update" instead of a second name.
- The skills menu announces the same sentence it shows, including in SSH chats.
- The row for messages held behind a previous agent reads "{{agent}} from before
may still be running. Your messages will send once it stops."
* fix(native-chat): a chat whose host can't run it says its history didn't load
A history read refused as unsupported with a named reason left the error pane
saying only "This isn't available in this chat.", which never said the chat
failed to load and named nothing the reader asked for. A read now says "This
chat's history couldn't be loaded."; other writes keep the shorter sentence.
|
||
|
|
3ad26486d5 |
fix(mobile): reduce base64 allocations for encrypted text frames (#24665)
* fix(mobile): reduce base64 allocations for encrypted text frames * Correct screencast encoder comment after byte-codec extraction --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> Co-authored-by: m4air <m4air@Mac.localdomain> |
||
|
|
3c9c2abe33 |
fix(mobile): release callbacks retained by canceled streams (#24625)
* fix(mobile): release callbacks retained by canceled streams * test(mobile): cover delayed frames after stream cancellation Preserve late-ready cleanup while rejecting canceled scrollback and terminal routing; cover unsent cancellation across browser, client events and session tabs. * test(mobile): model unavailable queued stream transport The encrypted sender rejects writes until connection setup completes. Keep existing session-tab unsubscribe attempts and assert canceled queued openers are never replayed. --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> Co-authored-by: Neil <neil@stably.ai> Co-authored-by: m4air <m4air@Mac.localdomain> |
||
|
|
1061dedcda |
fix(mobile): merge terminal backlogs without rescanning growing strings (#24680)
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
5c52dcee8f |
fix(mobile): release file previews after their tabs close (#24655)
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
7f388b70c6 |
fix(mobile): avoid backtick match arrays when editing Markdown (#24778)
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
3793c58abd |
fix(mobile): release notification and account streams from the transport (#23032)
* fix(mobile): release notification and account streams from the transport * test(mobile): drop the deleted notification-unsubscribe matrix from the bridged-parity tally * test(mobile): count the bridged-parity corpus at 786 goldens * chore(mobile): state why the notification stream frame cast is safe * test(mobile): re-record goldens after the transport took over notification release Repins the recording baseline to |
||
|
|
ddcc2796ba | Preserve Mermaid exports through mobile bundling (#24661) | ||
|
|
76b1a90ff6 |
chore(deps): update reviewed dependencies across Orca (#24561)
* chore(deps): update reviewed desktop dependencies and tooling * chore(deps): update compatible mobile packages and Fastlane * chore(deps): update cloud transports and enforce release age * chore(deps): patch documentation dependencies and record review * chore: remove dependency review reports * test(linear): smoke-load resolved SDK through CommonJS loader * fix(deps): keep native rebuilds from reinstalling addon dependencies * fix(native): invoke installed node-gyp directly for Node rebuilds * test(cloud): exclude observer probes from row-lock timing budget * test(mobile): preserve the CSS writer receiver in viewport spy * test(native): remove obsolete batch-shim fixture exception * Stream native rebuild output through the process wrapper |
||
|
|
8186ded0bd | fix(mobile): preserve terminal mode query variables in WebView builds (#24626) | ||
|
|
e3621295e6 |
Remove empty passing sentinels from opt-in socket tests (#24621)
* test: align source-control fixtures with current store contracts * Bound E2E package setup and retain cancelled-job traces * Remove empty passing sentinels from opt-in socket tests |
||
|
|
7176648759 |
fix(native-chat): every chat action press is its own action (re-land #23916 on main) (#24301)
* refactor(native-chat): a stopped child ends on the one reading of its stop
The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.
* feat(native-chat): the host says it accepts a send before any agent has it
The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.
* refactor(native-chat): an attach never opens a journal of its own
The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.
* fix(native-chat): a moved fence resends nothing on a host that accepts first
The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.
The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.
* refactor(native-chat): a child's end says whether the user or the host stopped it
The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.
* fix(native-chat): a chat whose only work is a queued message is not offered for resume
A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.
* fix(native-chat): the conversation outlives its agent
Opening a chat no longer starts its agent. A conversation is reached through one host
accessor that opens its journal at rest, and a send is what starts the agent, through
the delivery loop. One idle sweep, every five minutes, stops an agent that has been
quiet for thirty minutes and owes no work, then drops an open journal handle that is
only a cache. Its record, tab, status row and readers stay.
- hold and release are no-ops; hold still builds the host for shipped mobile builds.
- The holders, the holds, the release clock and the exit respawn are deleted.
- Options, the model list, the goal and the context meter answer at rest; a model pick
at rest is recorded as intent for the next start.
- Compact, rewind, clear and goal changes start the agent first. A send does too when
a rewind is still in doubt after the conversation opens.
- Orchestration routes mail and group addresses on ownership (the record plus the chat
tab), not on whether the process runs. An open dispatch keeps its worker running.
- The restart continuation is a send; Resume all holds each slot until the message is
handed over or rejected.
- A read error never replaces a loaded transcript, and shows the host's own words.
* test(native-chat): type the queued-message fixtures in the resume-offer tests
* fix(native-chat): a start that dies while a message waits on it is that message's failed start
Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.
* fix(native-chat): a request that failed reads as failed
A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.
The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".
* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now
* test(native-chat): a verdict change republishes the mobile status projection
* refactor(native-chat): the store's retention trigger keeps its flag compare
A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.
* test(native-chat): a user message the provider journaled keeps its session listed
* test(native-chat): pin what a failed start settles, and what a resume offer names
A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.
* test(native-chat): the failed-start pins fail on what the message became, not on a timeout
* fix(native-chat): a restart offer ends when the chat's agent starts again
The offer used to end only when the chat's newest user message changed,
because opening a chat started its agent and that start could not be told
apart from real activity. Opening a chat starts nothing now, so the host
reads the fact it already publishes: a chat's status row goes from not
host-owned to host-owned exactly when its agent is started. At that edge the
offer and any failure record for the chat are withdrawn, unless the start is
a resume action's own (its continuation is the oldest undelivered message).
A continuation and a message racing to be first are decided at acceptance:
the continuation is refused, quietly and with nothing filed, when any other
message was accepted since the restart. A failed continuation start leaves
the offer retryable, and each resume action sends its own message id.
Deleted: the newest-user-message comparison, its journal reader, the
continuation filter, and the failure ledger's own "answered by the chat"
check. The marker still carries its message id for one release, so the
previous build can read it.
* fix(runtime): end a transcript stream when its client unsubscribes
Desktop: the IPC subscription controller was dropped as soon as the streaming
handler returned, which for most streams is right after it binds. A later
runtime:unsubscribe then found nothing to abort, so the host kept the subscriber
and derived and sent every publish to a channel no one listened to. The controller
now lives until the renderer unsubscribes, resubscribes the same id, or goes away.
Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe
with the stream's frame id, so the host ends that subscriber and leaves a sibling
stream on the same socket running. The direct path now passes the frame id the relay
path already passed.
* fix(native-chat): a late provider-session update keeps a failed recovery record failed
A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.
* test(orchestration): the preamble's host stub is typed, not cast
The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.
* test(native-chat): the terminal-bell check asserts the renamed verdict field
The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.
* fix(native-chat): a failed turn ranks like a completion for attention
Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.
The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.
* fix(native-chat): one fact ends a restart offer: the chat moved on since the restart
The offer is live while no other message has been accepted in the chat since the
restart and its agent has not proved a start since. The offer list, the resume's
reservation check and the continuation's acceptance check all read that one fact,
so a message whose start then failed withdraws the offer too, and a stale click
finds nothing to act on.
The fact is read off the conversation's open handle, which the restart closed, so
it is retired durably whenever it may have changed: a message accepted, a start
proven. A close and reopen within the same run therefore cannot bring the offer
back. A continuation rejected before it reached the agent does not count, so a
retry after a failed start still runs.
Deleted: the quit-time gate on withdrawal, which changed nothing because the
withdrawal and the quit's own offer write share one queue; the per-action
"withdrawn" flag and the separate acceptance check it paired with.
* test(native-chat): an older build reads the restart offer this build records
The offer lives in a file the previous release reads after a downgrade. Pin that
against the pinned release's own capsule, and run the lane when the marker or the
capsule changes.
* fix(native-chat): read a restart offer against where the journal stood when it was taken
"Since the restart" was read off the conversation's open handle, which the idle
sweep closes: after a reopen, a message the user had already sent looked older
than the handle and the withdrawn offer came back.
The offer now records the journal position (epoch and sequence) at the moment
it is taken, and a message accepted after that position, or a journal on another
epoch, means the chat moved on. That is derived from the journal, so it holds
across any number of closes and reopens. An older build's offer has no position;
only a start withdraws it. Because the message half is now durable, the offer is
no longer rewritten in the recovery file on every accepted message; a proven
start still writes it, since only the host that saw the start knows of it.
* test(native-chat): wait for the listing's retire write before reading the recovery file
* refactor(native-chat): every journal row states which turn it belongs to
Rows gain a turn scope stated by the write that creates them: the open root
turn, or the conversation. A queued message takes its scope from its handover.
Rows stored before scopes existed are placed on replay by the root turn open
when they were created, so no persisted state is needed for them. Rewind keeps
each retained row's scope and producer, so a subagent's row stays its own.
* fix(native-chat): keep the terminal-backed chat's read error over its local echoes
Messages winning over a read error is right for the structured chat, whose read retries and whose
messages came from the transcript. The terminal-backed view assembles its list from local echoes
too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no
error. Only the structured pane now keeps messages over an error.
* fix(native-chat): a start retries the exit settlement a failed journal write left owed
An agent exit whose journal settlement write failed releases the lease latched until a retry lands.
Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried
it before the next app launch, and every send was refused. The start the send needs now runs the
retry first, where the attach would.
* fix(native-chat): a failed main agent reads failed while its subagents still work
The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.
Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.
worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.
* perf(native-chat): answer the owner check without opening the chat
Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the
answer comes from the session record alone. Reaching it through the accessor opened each resting
chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it
open for the idle window. It now checks the record and the adapter's support, as before this series,
and opens nothing.
* fix(native-chat): a read waiting on the session lock opens nothing once quit began
The accessor checked for quit before queueing the open, so a read queued behind a session task ran
its open after teardown had begun and indexed a journal no teardown step would close. The check now
runs at the open itself.
* test(native-chat): pin stated turn scopes, the upcast of unscoped rows, and rewind attribution
* fix(native-chat): /compact is a message the chat sends, run as a turn of its own
The conversation command RPC now accepts /compact into the queue like any
send and answers once it is handed over. The delivery loop opens the command's
own turn, starts the provider on it, and waits for the provider's end off the
session's queue, so messages typed meanwhile are held and delivered after it,
even when it fails. It settles by re-reading the journal: a child that died
meanwhile already wrote the verdict. Stop ends the command at once. The 180 s
completion window, the unconfirmed row and the recovery of an older build's
compaction record are gone; that record no longer gates anything. On Codex the
provider turn the command opens is claimed into the command's turn.
* fix(native-chat): read a failed resume's chat before calling it retryable
Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the
restart, read from its journal. The failure list read it only for a chat already open, so once the
idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did
nothing. The list now opens the failed chats first, as the offer list does.
* test(native-chat): type the provider event sink the settlement test reaches for
* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it
The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.
* fix(native-chat): say the structured read keeps trying only where it does
The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an
untranslated fallback whenever the read error had no text, and the empty state prefers any message.
The view state now leaves the message out, so the structured pane shows that line and the
terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an
error frame, so it no longer makes the claim.
* fix(native-chat): rows group under the turn their record names, not the one above them
Each row's turn is the turn its stated scope names, anchored on the entry
that opened it, or on the turn itself when the provider opened it unasked.
So /compact groups its own rows and the previous turn is untouched, a message
typed into a running turn joins it, and a provider-resumed turn folds under
its own Worked-for. A row reporting how a turn ended, an error or the
compaction separator, never folds. Desktop and mobile read the same keys; a
host that states no scope keeps today's positional grouping.
* test(native-chat): await the send's settlement instead of polling for the start
The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a
loaded machine outran. They now await the host's own settlement of the message.
* docs(native-chat): the status-store listing rule names provider-journaled user messages
* fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations
The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than
a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now
dated by the resume action.
Telling a rejected continuation from the user's own message read the operation ledger, whose rows
expire after about a day; after that a failed resume stopped being retryable. The offer now
records the continuation each action sends on its own capsule entry, bounded to the newest 16, so
the ids end with the offer. The ledger read is deleted.
* fix(native-chat): a /compact is not a request the sidebar, notifications or restart resume report
The sidebar's prompt, preview, verdict and instant, the turn-completion feed,
and the restart-resume marker read past a conversation command and its turn to
the last real request, so a /compact neither notifies nor re-dates the row,
and a command in flight is never offered as work to resume. An older client
shown a command's turn in the legacy form names the session's own agent.
* fix(orchestration): route no mail to a structured worker its orchestration released
A structured worker is routed on ownership, and a resting worker's lease is released, so ownership
held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found
at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one
restarted its agent. Routing now also reads the orchestration's own resource row: once it is
released, direct mail, group addressing and worker-show's addressable answer drop the worker, as
they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored.
* fix(native-chat): a failed retry names the user's prompt, not Orca's continuation
A resume's continuation is written to the chat before its start, so after a failed attempt the chat's
newest user message is that rejected continuation. A second failure then showed Orca's own restart
text as the chat's prompt. A retry now keeps the prompt its first failure named.
* test(native-chat): pin what a conversation command's admission refuses at rest and at handover
* test(native-chat): tests merged from the base state which turn their rows belong to
* fix(native-chat): a refused send notifies failed through the completion feed
The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.
* fix(orchestration): read the released row optionally, as the authority does
worker-show's observation called the row lookup directly, which a runtime double without it threw on
and failed the structured tab-retirement release.
* chore(native-chat): one import per module and no unexplained casts in the turn-scope changes
* test(claude): pin which turn a Claude row joins, including a subagent's after the turn ends
* fix(native-chat): the status bar drops a restart offer the chat moved on from
The renderer re-read the host's restart offer only when a failed chat showed activity, so after a
message withdrew a pending offer the host answered no chats while the status bar kept counting one,
and clicking it opened nothing. The same watch now covers pending offers: a status change in an
offered chat asks the host again, once.
* fix(native-chat): a refused steer is read from the turn its handover named
The latest-request reader decided whether a refused send had joined a running turn by comparing
host clocks: its handover time against the previous turn's end. The handover row now states the
turn it delivered into, so the reader reads that instead and the clock comparison goes. A journal
written before handover rows stated a turn is scoped on replay from the turn open when each row
was written, which can differ from the clock reading only when a send and a turn's end share a
millisecond.
* fix(mobile): the native-chat controller contract carries the turn journal
The controller and overlay already pass nativeChatTurnJournal, but the
contract type never declared it, so mobile failed to typecheck.
* fix(native-chat): the live turn is the running turn, not the newest user row
A turn the provider opened on its own (a background wake, a resumed turn)
anchors on its own record, but the list still treated the newest user row
as the live turn. While such a turn ran, the settled user turn before it
lost its duration and the running turn's own rows were drawn as settled,
so its tool calls lost their live state.
nativeChatTurnMembership now answers both questions from the turn record:
each row's turn, and the live turn (the running root turn's anchor, else
the newest user row, which is also all an unscoped host has). Desktop and
mobile key liveness, the timing clock and the live status's row on it.
* test(native-chat): a turn the provider opened keeps its own clock
Pins that the local turn clock follows the live turn, so a wake after a
settled turn does not restart that turn's clock when no host durations
are recorded.
* fix(native-chat): a running turn no message opened draws its status on no row
Its live status belongs to the transcript-tail indicator alone. Once it
settles, its duration draws at its first row as before; a running turn a
message opened still draws on that message.
* fix(native-chat): every copy of a row carries the main agent's own status
History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.
- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
rebuilding one; the sync key and history equality compare it.
* test(native-chat): pin the worktree ps verdict across host and phone versions
Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.
* test(mobile): name the parity table's row for its role
* test(native-chat): a roster of idle or finished children does not keep an agent awake
The sweep reads owed background work through the shared child-work liveness that upstream's
release clock adopted; a child that went idle or finished is not work the agent still owes.
* fix(native-chat): a request that settles while the user is asked something notifies once
The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.
The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.
* fix(orchestration): a task dispatched into a resting structured worker keeps it running
The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal
resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a
dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while
that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's
process incarnation now counts, derived from the existing rows.
* fix(native-chat): a command's wait ends when its child does
The delivery loop waited for a /compact only on the adapter's compaction
tracker, which learns of the child's end only on some exit paths: a Codex
exit or close, and a Claude close, never reach it. The wait then never
ended, so nothing queued behind the command was delivered again, Stop had
no child to answer through, and the tracker's leftover entry refused the
next /compact.
Every way a child ends passes endProviderChild, so the host now offers a
per-child end signal there. The loop races the tracker against it (the
dead-generation settlement has already written the command's verdict),
and on that end asks every adapter to release the command, so a later
command runs and no later provider turn is claimed into the dead one.
The adapters' own exit-time releases were unreachable (Codex) or covered
one path of several (Claude), and are removed.
The Codex RPC test harness moves to its own module so the exit can be
driven through the real adapter's connection callback.
* fix(native-chat): keep refusing sends during a command on an older host
An older host's controller still refuses a send while a conversation
command runs, so dropping the client's block turned every message typed
during /compact into a 'not sent' row with Retry there. The block stays
for hosts that do not run the command as a send-path turn, and goes only
for those that do.
The signal is one the client already holds: a host that runs /compact on
the send path states a turn scope on every journal row it writes, the
same fact turn membership uses to tell it from an older host. Both now
read it from one predicate. On an empty conversation, or one whose rows
all predate the upgrade, the signal is absent until the command's own
entry streams in, so that brief window keeps the old local refusal; no
capability or wire field is added.
* docs(native-chat): comments stop describing the hold this PR removed
Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that
pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive
subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it.
Comment-only.
* fix(native-chat): the completion says when the user is being asked
A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.
The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.
* fix(native-chat): a restart offer keeps the start its own continuation made
Whose start ended an offer was decided at read time, from whether the offer's continuation was
still the queued message. Once the provider refused that continuation, the child it had started
read as someone else's start, so the offer ended and its failure showed no Retry. The delivery
loop now records which queued message a start is for on the in-memory child, and the child's end
carries it; the offer counts a start as its own when that message is one of its continuations.
* fix(native-chat): a rewound turn still names the message that opened it
A Codex rewind rebuilds the epoch without submissions, so each sent message survives only under
its provider key. The kept turn records still named the submission key, so each turn anchored on
itself and its rows grouped apart from the message that opened it. The rewind now renames the
turn's opener along with the message.
* fix(native-chat): Stop ends only the command it names
Stop on a command turn abandoned whatever compaction the session had pending, so a late Stop for
an earlier /compact cancelled the one running now. The tracker now ends a command only when the
Stop names its turn, and the cancel reply reports whether it did.
* fix(native-chat): an agent gets a full idle window after its owed work ends
The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or
dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can
read done before the lead's wake-up turn writes anything, and stopping in that gap loses the
wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full
window afterwards, as the release clock it replaced did.
* test(claude): the options-read fixture runs a live child
The fixture marked its conversation running with a hasProviderChild field the
session type does not have, so the read took the at-rest path and refused a
session with no record. It now carries a child, which is what the read checks.
* test(native-chat): host tests reach its collaborators through a typed seam
The rest-test rig and three test files read the host's private members with
Reflect.get and cast the result. The host now exposes one test-only accessor,
collaboratorsForTests(), and the subscribers class a subscriberCountForTests()
beside its existing retainedActivityCountForTests(), so the tests are checked
against the real types and the casts are gone.
* fix(worktree-status): a departed agent's failure yields to live work on the worktree card
A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.
* refactor(orchestration): one owner answers a structured worker's custody
Routing, group addressing, worker-show and the idle sweep each composed their own reading of
whether orchestration still holds a structured worker, so each new obligation or retirement state
had to be added to every reader. structured-worker-custody now derives both answers from the
worker-terminal list state coordinators see in worker-list: addressable is owned and not released,
and owed work is an active custody or an unsettled task dispatched to the same incarnation. The
owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup
already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests.
* refactor(orchestration): owed work is an open dispatch on the worker's incarnation
A supervised worker's own dispatch context stays open exactly while the worker is active, so the
separate active-custody branch only repeated it. Owed work is now one fact, which also states the
policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are
written once at the top of the module.
* docs(agent-status): a departed agent's failure ranks below live work on the worktree card
* fix(native-chat): a restart offer knows its continuations by a tag in their id
The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running
action's id in memory. Both could disagree with the journal: past the cap an old rejected
continuation read as the chat moving on, and a crash during a retry restored the failure's older
entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer
(its teardown and chat), then the action's own part, so any continuation of this offer, queued or
rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and
the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own
continuation the start was for, read against the stored marker.
* test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait
The tui-idle probe reads through readTerminal, which now awaits the structured
worker check before the PTY read, so the probe's snapshot request starts a
microtask later. vi.waitFor missed it on its first check and polled again at
50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot
then resolved after the wait had already timed out, so the test passed without
judging it, and the rejection landed before any handler was attached. Vitest
reported that as an unhandled error and failed the shard.
Polling every 1 ms sees the request within a few ms, so the snapshot is judged
while the wait is still pending.
* fix(native-chat): a message held behind /compact is drawn where it was handed over
A message typed while /compact runs was drawn above the compaction's result, between
itself and its own answer. The reducer kept every item at the sequence and timestamp of
the row that created it, and a queued message is created at acceptance, long before the
command it waits behind writes its result. The phone orders by that sequence and the
desktop by that timestamp, so both put the message first.
A queued message now takes its position from its handover row, the same row that already
states its turn scope. Everything the agent did before the handover, a command it waited
behind included, draws above it. This holds for every held message, not only /compact's,
and needs no client change: every client, older builds included, reads the position the
host publishes. A live batch already carries the item when its dispatch row lands, and
history pages cut the reduced timeline by sequence, so paging stays contiguous.
* fix(native-chat): a phone's send during /compact answers without waiting out the compaction
A client that predates accepted-send replies, which is every phone build, has its send
reply held until the host hands the message over. A message sent during /compact is not
handed over until the compaction ends, so the phone's 15 s request timeout fired first
and showed the message as unconfirmed.
That wait now also ends once the message is queued behind a running command. This is
read from the journal's running turn and needs no new state. Every other wait still
ends at the handover: behind a starting child or an ordinary turn, and for restart
resume, the command front door and orchestration, which keep the plain handover point.
* perf(native-chat): a rewind places provider items with one pass over the merged rows
A Codex rewind gives each provider item the old epoch never held the turn record for its
provider turn. It found that record by scanning every merged row, restoring each row's
body, once per provider item. That is quadratic, and it runs on the host's main thread
up to the journal's 10,000-row cap, twice per rewind. A rewind record written before
rows carried their scope holds no scope for any provider item, so it paid the full cost.
The merge now indexes turn records by provider turn id once, keeping the first match as
the scan did, and each provider item looks its record up.
* fix(native-chat): a view never restarts a chat whose last start failed
A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.
* fix(native-chat): a message waiting behind /compact is drawn after it until it is sent
A message sent while /compact runs is placed where it was handed over. It was still
drawn where it was accepted until then. /compact writes its result one step before the
handover, so for that step the waiting message sat above the compaction's separator.
A message the host accepted but has not handed over is not part of the conversation
yet, so both clients now draw it after everything the agent has done. The shared
projection moves it to the end, which is the order the phone draws. The desktop ranks
it with the other not-yet-sent rows, after the streaming preview. At handover it takes
its place from its handover row, which is also after the separator, so it never
appears above the compaction it waited for.
* fix(native-chat): the idle sweep reads owed work every tick
Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it
refreshed the clock at most once a window. Work that ended just before the next read left the
agent to be stopped at that read, moments after the work ended, which is the gap the refresh was
meant to cover. The sweep now reads owed work on every tick for a started agent, so the window
always runs from the last tick that saw work owed.
* fix(native-chat): a continuation handed to the agent stays sent
The offer read its own continuation as not reaching the agent while its dispatch was pending, which
also covered one already handed over and still unanswered. When the wait for that answer ended first,
the failure it filed read as retryable, and a retry sent a second continuation to an agent that may
have acted on the first. Only a continuation still queued, or rejected, is now read as unsent.
* test(native-chat): start the child the loop waits on with an attach, not a second view
A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.
* fix(native-chat): settle a gone generation's turn wherever a conversation opens
A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.
* test(native-chat): prove the next child's start settles the turn an earlier child left
The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.
* test(native-chat): count a failed start's rows by row, not by text
Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.
* test(cross-version): load the phone row readers without mobile's toolchain
Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.
The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.
* test(cross-version): keep the checkout path-guard message and justify the copy import's cast
* fix(native-chat): a command ends only by its own provider answer or its child's end
Stop no longer settles a conversation command. It interrupts it like any turn,
and when the provider cannot take that (Codex has not opened the command's turn
yet, or Claude refuses the interrupt) it stops the child, whose dead-generation
settlement writes the verdict.
The pending command now lives on the provider child's own session instead of an
adapter-wide map keyed by session, so it dies with the child and nothing has to
release it. Claude's /compact is sent under a uuid the slot records, and only a
root result naming that input (or naming none) ends it; its outcome is read with
the ordinary result reading, so a stopped /compact is a cancellation.
* fix(native-chat): a command's settle answers its message before ending its turn
The two writes are not one batch. Writing the message's answer first means a
crash between them leaves a running command turn, which the stale-turn sweep
already settles, instead of an ended turn whose message reads as in flight
forever. The settle now writes only while the command turn is still running.
* fix(native-chat): "Worked for" counts from the handover, not the send
A message held behind /compact, or behind a cold start, used to count the wait
as the agent's work, although its row is drawn at the handover. Every handed-over
submission's turn, the command's own included, now starts at the handover row's
instant, falling back to the send time for a host that recorded none.
* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget
* test(native-chat): the interrupted create's own retry continues again
The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its
own operation id, with a fresh start whose result nothing read. That fresh start passes with the
released-reservation continuation deleted, so the case the fix exists for went untested. The retry
and its assertion are main's again.
* docs(native-chat): three comments that still had views starting agents
A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction
left alone would refuse every send, so no agent would ever start to finish it; and a current host
raises the unattached read refusal only once quit began, with the attach window belonging to an older
host.
* test(native-chat): pin the open's and the send's start and row counts, however the view binds
Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.
* fix(native-chat): a second Stop on a command ends its child; one compaction verdict for every provider
A Stop's note now names itself in its key, so a later Stop on a command still
running reads, from the journal, that the provider was already asked and never
answered, and stops the child instead of interrupting again. Nothing is held in
memory for it.
Adds the rule both translators will read a compaction's end by: only a
compaction the provider reported is a success; none after Orca's interrupt is a
cancellation; anything else is a failure. A real Claude capture, pinned as a
fixture, is why: a stopped /compact ends in the same success result as a
finished one.
* test(native-chat): a reader's open settles the turn a failed exit settlement left running
An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now
settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle
sweep closed, and a read that opens the chat before the restart restore reaches it.
* test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner
A subscription reads the conversation before it returns, so under load the two views took longer
than the create child's 300 ms start, which then exited before the test checked that it had not.
The child now takes a second to fail.
* fix(native-chat): settle a gone generation's turn at every open but an acquisition's
The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.
* test(native-chat): hold the create's start open until the views bind
The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.
* refactor(native-chat): the provider's translator ends a command's turn; the loop holds no command state
A conversation command is now a turn of the provider child's own journal
pipeline. The adapter-wide tracker, its promise and the loop's settle step are
gone.
- Codex: the translator claims the provider turn that carries the command, scopes
its rows to the command's turn, and writes the command's end in the same batch
that settles that turn. Codex's own compaction marker is the success row.
- Claude: the command's turn is the translator's open turn until the result that
answers the /compact input ends it. The command's own frames, such as the
continuation summary, its echo and "Compaction canceled.", draw nothing.
- Both read the end with the one compaction rule: success needs the provider's
report of the compaction; none after Orca's interrupt is a cancellation.
- The message resolves at the provider's receipt, as any send does: the Codex
ack, or the Claude slash-command waiter on its result. The host writes a
command's end only when the provider never took it.
- The delivery loop stops while a command's turn runs, and every journal commit
re-wakes it through the session's serialize, so an end that lands while a step
decides to stop is never lost. A child that ends first is settled with it.
* test(native-chat): pin a command's end to real /compact frames and to each path it threads
The captured /compact frames drive the Claude translator's command turn: a
finished compaction ends as a success with only the separator drawn; a stopped
one ends as a cancellation with no failure row, and the next send answers in its
own turn; a result naming another input ends nothing. The command's end is
checked at each point the ordinary result path threads through: the reopen latch
after a failure, the settling of a child still working, the context facts the
result reports, and the provider's own error row.
On the host: a message held behind a command is handed over when the command
ends just as the loop stops for it, a refused command settles as a failure and
the loop moves on, and a Claude child that exits mid-command settles the command
and hands what waited to a fresh child.
* test(native-chat): tests merged from the base state which turn their rows belong to
* refactor(native-chat): drop the child-end waiter nothing waits on
A command no longer waits for its child here: its turn ends from the provider's frames or from
that child's settlement, and the delivery loop is woken by the commit. The waiter and its test
were left from the earlier shape.
* fix(native-chat): a command holds the queue only while its child runs it
The delivery loop stopped whenever the journal showed a command's turn running. When the
command's child ended and its settlement could not be written, that turn stayed running with
no child to end it, and the loop's gate kept it from ever starting the next child, which is
what settles a gone generation's leftovers. Every later send was held for good, and Stop had
no child to end.
The gate now holds only while the conversation has a child: with none, the command belongs to
a gone generation, and the loop's start settles it like any turn a dead child left running.
* fix(native-chat): a Claude /compact succeeds only on its compaction boundary
The command's evidence counted Claude's `compact_result: 'success'` status as the compaction
done. That status comes before the boundary that replaces the history, so a Stop landing
between the two read as a finished compaction even though no boundary was ever written. Only
the boundary now counts, as the rule for both providers states; the capture's finished
compaction carries one, so it still reads as a success.
* fix(native-chat): a Claude child's exit says why the turn it ended stopped
When a Claude child exited mid-/compact, the command showed "Worked for 0s" and no reason. The
child's translator ends its open turn the moment the exit is reported, stamped with the exit's
instant, so by the time the exit settlement ran nothing was running. The settlement recognises a
turn the exit already ended by that same instant, but the Claude lifecycle event dropped it on the
way to the host, which then used its own clock, matched nothing, and wrote no row. When the clocks
did agree, the row was scoped to the running turn, of which there was none, so it landed outside
the turn it explained.
The exit's instant now reaches the host, and the exit row belongs to the turn the exit ended:
still running, or ended by the translator at that instant.
* fix(native-chat): a message waiting behind /compact draws below its live activity
A message sent while /compact runs waits on the host until the command ends. Both clients moved
it to the end of the transcript rows, but the running turn's live activity line ("Compacting the
conversation") draws after every row, so the waiting message sat between the command and its own
live status.
A row that is queued, and not what the live turn is for, now draws after that live activity: on
desktop outside the transcript window, below the activity line; on the phone in the list footer,
below the live status. A message whose own start is pending still draws above the activity that
start reports.
* fix(native-chat): only a running command holds a message below its live activity
A message is accepted, then handed over a moment later, and in between it reads as waiting. Every
message waiting behind a live turn drew below that turn's activity line, so an ordinary message
sent while the agent was working crossed below "Thinking" and jumped back up once it was handed
over, on desktop and phone. Only a conversation command's turn holds the queue on the host.
A message now waits below the live activity only while the running turn is one a command opened,
read from the entry that opened it. The phone test also typechecks, which the mobile test ratchet
requires.
* test(codex): the claim test names its notification params as a record
* test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn
On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the
dead process's lease still reads live. The open settles the turn it left running anyway, and the
restore that follows finds it settled.
* refactor(native-chat): drop the composer's second error formatter
After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.
* test(native-chat): pin the reason on a message rejected while its chat was closed
The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.
* docs(native-chat): drop the removed dispatch hold from six comments
A worker's session no longer takes a dispatch hold, and no release clock
rests a chat by visibility; the agent-launch comments, the abandon test,
the teardown test and the refusal census still said so.
* test(native-chat): rest the owner-status chat through the idle sweep, not a hold
The activation-gate test from #22808 put its chat at rest by holding and
releasing it, and passed the release-clock grace. This branch deleted both,
so the case threw before it reached its assertions. It now moves the host's
clock past the idle window and lets the sweep stop the agent and close the
conversation, then asserts the same owner answer and activation gate.
* fix(native-chat): show the structured pane's retrying line when a read fails
The read transport always hands the pane the host's words, so the error
state's "Orca keeps trying to load it" line, which showed only when there
were none, was never seen: the pane showed the host's text twice, as its
subtitle and on the status line under it. The structured pane now always
says its read keeps retrying, and the host's text stays on the status line.
The terminal-backed chat is unchanged.
* fix(native-chat): a send the provider never received after a restart has no verdict
Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.
* fix(native-chat): a failed Codex compaction's late completion writes no turn of its own
Codex ends a failed turn with an error and then still completes it as failed.
The error settled the compaction and released its claim on the provider turn,
so the completion read that turn as an ordinary one and wrote a stray record.
The claim now lasts until the completion, which adds nothing to a command the
error already ended.
* test(native-chat): the mid-command exit case resumes its next child as a real one does
The case's fake started every child as a newly created thread with the same generation. The
store refuses a created link once the conversation has a thread, so the next child's start
failed and wrote its own error row, which landed before or after the case read the journal.
The next child now resumes the thread under its own generation, and the case reads the
journal once the waiting message is delivered, which also proves the loop moved on.
* fix(native-chat): a /clear that never committed no longer locks the chat
A /clear wrote a durable "prepared, outcome unknown" record before starting
the replacement conversation. When that start was refused without a definite
answer (or Orca died), the record stayed forever, and while it did the chat
refused every send, /compact, a new /clear and rewind. Its only exit was a
rerun under the same operation id, which only the renderer held.
The record guarded nothing the process does not already know: a clear in
flight holds the session's serialize for its whole run and the command
controller refuses sends meanwhile, and the replacement's id and start
operation are pure functions of the clear's operation id. So the clear now
writes nothing durable before its commit, the gates refuse only a committed
clear (an older build's prepared record is inert), and a clear with no
committed answer reruns: a same-op retry re-attaches the same replacement,
a new op id runs a fresh clear.
A crash between the replacement's start and the commit leaves a replacement
record nothing points at. Verified: it has no tab, is not in the
replacement list, and a restart opens and starts nothing for it (restore
reads only the visible tab index); restart reconciliation releases its lease
like any dead owner's. In a live process its agent is stopped by the idle
sweep like any quiet agent. Session History lists provider transcripts and
only annotates them with an owner, so it can list this only if the provider
wrote a transcript for a thread that never got a message. Its record stays
on disk, as every closed chat's does; the store deletes none.
* fix(native-chat): a Codex rewind the provider did not keep no longer fails every attach
When Codex acknowledged a revert and Orca stopped before proving it, the
rewind stayed prepared with providerApplied set. On the next attach,
recovery read the provider's history, found the target turn still there
(provider-refused), and threw, because that settlement was limited to
reverts never sent. The throw ran inside the attach, so every attach, and
every send that needs one, failed for good.
The journal is replaced only once the provider proves the revert, so both
the provider and the journal still hold the target turn: settling the
rewind refused is consistent whether or not the provider acknowledged it.
* test(native-chat): a clear retried after a crash starts no second replacement
The replacement's id is the only thing that keeps a retried clear from leaving a second one, and no test held it across a restart.
* chore(native-chat): the clear rerun comment claims only the stable replacement id
* test(native-chat): wait for a send's background start before the refusal oracle removes its store
An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals.
* fix(native-chat): a start a message waited on gets one failure row, the delivery loop's
When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice.
The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit.
* fix(native-chat): a /compact whose start failed says to run /compact again
The failure-words context named only /clear as a command to retry, so a
/compact whose agent failed to start read "Send your message to try again."
on its row, its rejected message and the command reply. The context now
carries any conversation command; the host derives it from the oldest
message still waiting on the provider, which is the one a failed start
fails first, and the /compact reply names it directly.
* fix(native-chat): a Codex /compact ends only on its turn's completion, below Codex's own error row
Since only turn/completed ends a Codex turn, Codex's turn-ending `error` is a row
inside the still-open command turn, and the failed completion that follows it is
the command's end: completed, outcome failure, at the completion's receipt time.
The command's own "Compaction failed" row was written on that completion too, so a
failed /compact read its reason twice.
The command turn now notes when Codex's turn-ending error for the turn it carries
was written as a row, and its end then adds no second row. A retried stream error
ends nothing and is not counted. The flag that let the error end the command and
kept the claim until the completion is gone with the error-driven end.
A test replays the captured failed compaction from the real app-server through a
claimed command turn.
* test(native-chat): main's crash-turn test states its row's turn, and a dead /compact settles on its recorded exit
Two tests the main merge brought together:
- The crash-turn test from #23456 writes a turn record through the event sink
without options; every row here states its turn scope, and a turn record's is
the thread.
- The /compact whose exit settlement could not be written no longer stays running
until the next start: main now settles an open chat from the exit it recorded, so
the command reads interrupted before the next message, which is then delivered.
* test(native-chat): main's new journal tests state each row's turn
The crash-turn, stale-turn and sink-queue tests main added wrote rows without a
turn scope, which every item write now states. Rows written inside a running
turn name that turn; the sink-queue batch and a send handed over with no live
turn name the thread.
* fix(native-chat): draw a queued turn's message after the earlier turn's rows
A message sent while A runs is written to the journal when it is sent.
When the provider queues it (Claude answers it after A), A's remaining
rows - its last tool run and its answer - are written after that
message, and the message's own turn opens only after them. Grouping put
those rows in A's turn, but the transcript still drew them in journal
order, below B's bubble and bar, where A's answer read as B's reply. This
is the residual #23671 left open.
A message that opened a turn now draws after the earlier turns' rows the
journal wrote after it, just before its own turn's rows
(nativeChatTurnDrawOrder, returned by nativeChatTurnMembership as
drawOrder). Desktop and mobile both draw in that order. A steer, and a
message that has opened no turn yet, stay where they were written. It
applies on hosts that state turn scopes and, through journal order, on
older ones.
* test(native-chat): run #23026's Stop tests against #23059's command turns
Two of #23026's tests call APIs #23059 changed, and failed after the
merge:
- codex-structured-conversation-stop: a compaction now goes through
adapter.compact with the command run the host wrote (#23059), not a
bare turn id, and answers with the provider's receipt. With the command
claimed, a Stop that names no turn while the compaction's provider turn
has not opened still interrupts nothing.
- main-agent-working-agreement: a provider row states its turn scope
(#23059's appendItem contract); the retry and subagent rows are
conversation-scoped.
* fix(native-chat): typecheck main's Stop and restore-grouping code against #23059
A Stop's compaction interrupt reads the narrowed requested turn, and the
restore-grouping test states whether each row reports its turn's outcome.
* fix(native-chat): say a /clear cut off by a restart left the chat unchanged
A /clear retried under the same operation after Orca restarted could not reuse the new conversation its first try started, and its row said "Codex couldn't start. Run /clear again." The agent did not fail to start: the earlier try was cut off. The row now reads "This /clear didn't finish, so the chat is unchanged. Run /clear again to start fresh.", from a new clearUnfinished failure fact written through agentSessionFailureWords.
The clearUnconfirmed and conversationCommandUnconfirmed reasons stay, with their words, for older hosts that still send them.
* fix(native-chat): a retried /clear finishes onto the conversation its earlier try started
When an earlier try of the same /clear started its replacement conversation and a restart or the
idle sweep has since stopped it, the retry could not replay that settled start and reported the
chat unchanged. That replacement is a fresh conversation at rest, so the retry now commits onto it
and its first message starts its agent. A replacement whose start definitely failed still reads
that failure, and one Orca can't prove stopped still commits nothing. The clearUnfinished failure
kind this made unnecessary is removed.
* refactor(native-chat): stop recording that Codex acknowledged a rewind
A refused rewind recovery now settles as refused whether or not Codex acknowledged the revert,
so nothing reads providerApplied any more. Stop writing it and drop the hook that wrote it.
Records that still carry the field load as before; the schema ignores the extra key.
* fix(native-chat): a /clear retried under a new operation id finishes the same replacement
A /clear's replacement id came from the client's operation id, so a retry the client sent
under a fresh id started a second replacement and orphaned the first. The host now derives
it from this caller's oldest /clear since its last commit whose replacement start reached
the operation ledger, so any retry from that caller finishes the same replacement, including
after a restart. A /clear after a committed one starts a new replacement. Another caller's
/clear is refused only while such a replacement is running or not proven stopped. An older
client that resends the same operation id still lands on the same replacement.
* fix(native-chat): a /clear retry never repeats a failed start or waits on an unproven stop
A retry under a new operation id could pick an earlier try whose replacement start had already
failed, replay that failure and commit it again, so a user who had since signed in was told
they were still signed out. Such a try is now skipped, and the retry starts afresh.
Another window's /clear was refused while the first window's leftover replacement was merely
not proven stopped. Nothing but the first window's own retry would settle that, so the refusal
could last until its ledger row expired a day later. It now waits only on a replacement whose
agent is running.
* fix(native-chat): a /clear retry finishes only a replacement that started
A retry picked an earlier try whose replacement start never answered, because a crash left
that start unsettled. Replaying it could only repeat "couldn't start" or, with the old agent
unproven, refuse every /clear from that window. Only a start that succeeded left a
conversation to finish; any other try is skipped and the retry starts afresh.
* refactor(native-chat): a record's identity fields are built in one place
A created record and a founded one (a conversation no agent has run yet, at
rest) share who and where the agent is and how it launches. The founding
builder is used by the /clear commit that follows.
* feat(native-chat): the store commits a /clear and its new conversation in one write
commitConversationClear founds the at-rest replacement from the cleared
record's identity and writes the committed marker and tab move in the same
transaction, so neither can land without the other. It refuses to overwrite
an existing record under the replacement id.
* fix(native-chat): /clear starts nothing; the new chat's first message starts its agent
/clear used to start the new conversation's agent before it committed, so it
could fail on that start ("Run /clear again"), and a crash between the start
and the commit left a running conversation nothing pointed at. #23524 then
needed a ledger scan to find an earlier try's replacement, a nonce half of the
derived ids, a refusal of another window's /clear while a leftover agent ran,
and a check for a start that had already finished.
Now /clear opens the chat for writing (it no longer starts an at-rest chat's
agent either) and makes one store write: the at-rest replacement under a
random id, the committed marker, and the tab move. The first message in the
new chat starts its agent through the existing send and delivery path, fresh
because its handle chain is empty. A failed start shows on that message with
the typed failure and a Retry, and a conversation no agent ever ran now reads
"couldn't start" rather than "couldn't restart".
Deletes clearTryToFinish, otherCallersClearIsLive, the attach block and the
committed start-failure branch, and the tests of that retry machinery.
* test(native-chat): drop the /clear retry wording test; no start runs for a /clear now
* test(native-chat): another window and a phone read a /clear's replacement from the host
Both list the replacement the committed marker names, under the chat's tab,
and each one's session list shows it with nothing unread until its first
message runs. A reader that recomputed the id from the operation turns this
red.
* test(native-chat): a never-started replacement closes as settled
A worktree delete closes every chat in it and asks the user to force any it
cannot prove stopped. A replacement no agent has run is released, so its
close settles like any at-rest chat's.
* fix(native-chat): /clear settles an interrupted Codex rewind the way a send does
/clear moved from starting the chat's agent to only opening the
conversation. A Codex rewind cut off mid-way on a chat at rest can only be
settled by its agent, so /clear was refused as "rewind unconfirmed" every
time until the user happened to send a message. It now prepares like a send
or /compact: the agent starts only when such a rewind is in doubt.
* fix(native-chat): a chat whose agent is not running keeps its `/` commands
Claude reports its skills and project commands only from a running process,
and the host served the `/` menu only from the running agent. Now that
/clear starts nothing, the new chat's menu lost those entries until its
first message; a chat stopped by the idle sweep already did.
The host now keeps, in memory, the list a running agent last reported for
its launch (provider, host, workspace, account and launch arguments) and
serves it to a chat of the same launch whose agent is not running. A new
report replaces it; nothing is stored on disk, so a relaunch still shows
the short menu until the agent reports again, and no list is ever served
across accounts, workspaces or hosts.
* test(native-chat): queued drafts around /clear follow what a /clear now is
Three queue tests from #23726 are red on main
|
||
|
|
1762a138f7 |
feat(mobile): slide the page's host stack on push and Back (#24268)
expo-router's Stack on web renders native-stack's web view, which flips display and ignores animation. The page's host stack now keeps expo-router's StackRouter under its public Navigator and draws the slide with the Web Animations API; a popped screen stays mounted until it has slid out. Native is a pure move. |
||
|
|
069eaf5668 |
fix(mobile): iOS shell keeps WebKit text interaction on so page fields take text (#24270)
With isTextInteractionEnabled = false (#21589) WebKit delivered keydown to a focused page field but never inserted text. The page's user-select rule (#24277) now keeps WebKit's selection off long presses. Needs an iOS shell release. |
||
|
|
e9ec63168f |
fix(mobile): hold-to-dictate, repeat keys and the browser long-press survive the page's long-press (#24277)
On the OTA page a held press died ~500 ms in: the WebView's long-press selected nearby text and that selection's selectionchange/touchcancel ended the press. Page text is now unselectable unless it opts in (as native), hold surfaces declare onLongPress, the browser pane refuses contextmenu termination, and the chat mic's swapped icons no longer steal the touch target. |
||
|
|
3ab3c9239f |
fix(native-chat): the working line shows only what the agent is doing now (#24218)
* fix(native-chat): the working line shows only what the agent is doing now A chat's live "Working…" line could show an old notice, such as "Claude hit a temporary problem and is retrying.", long after the agent had moved on and was running new commands. When the host had no live activity for the turn, the line fell back to the newest status row in the turn, and any status row qualified: retry warnings, "Context compacted", "Cancellation requested.", and notification summaries. Those rows record the past and already appear in the transcript. The line now reads only the host's live, per-turn activity, which is never saved and is cleared at turn boundaries. Without it, the line says Thinking or Working…. Desktop and mobile share the selector, so both change. * test(codex): guard that a subagent's compaction never becomes the parent's live activity |
||
|
|
b4b708c2c4 |
fix(native-chat): a Codex stream retry is one warning row that updates in place (#23684)
* fix(native-chat): a provider's own retry progress is quoted in its retry row A retry row whose fact carries a detail the provider wrote for a person now quotes it, the same way a rejected message or failed compaction does, so the row says how the retry is going. A log detail still stays out of the sentence. * fix(native-chat): a Codex stream retry is one warning row that updates in place An error Codex says it will retry used to fall through to the generic frame row: red, and a new row for every attempt. It now writes one providerRetrying row per retry run, warning-toned, revised by each attempt with Codex's own progress sentence. A run is the retry frames of one turn with nothing else the thread journals between them; every attempt still publishes, so the idle sweep keeps seeing activity. Errors Codex will not retry are unchanged. * fix(native-chat): a Codex retry row says it is retrying and keeps the frame behind Details The quoted retry sentence leads with "is retrying", which holds for any provider's progress text. The Codex retry row also keeps the whole bounded frame behind the row's Details, as the generic row did, so Codex's additionalDetails stays available to diagnose a retry. * test(native-chat): a Codex retry re-handled after backpressure keeps its one row Pins the run being opened before the write: a first attempt whose publish is refused and is handed back must revise the row it already wrote, not open a second run. Also stops the fixture claiming Codex sends an idle thread status beside each retry, which the app server does not do. * perf(native-chat): a Codex frame with no retry run open is not classified Ending a retry run classified every non-retry frame, and classifying walks the whole payload: every streaming delta, and every large item/completed, paid a walk about as costly as parsing the frame. Only a thread with a run open needs the answer, so the classification now runs only there. * fix(native-chat): each Codex retry attempt is its own warning row, with what failed on its second line A stream error Codex says it will retry is written as its own warning row with a providerRetrying fact, under the same per-frame identity every Codex frame row gets. The host no longer tracks retry runs or rewrites one row in place, so there is no run state to open, end or clear, and no frame has to be classified to end a run. Every attempt publishes, which keeps renewing the idle clock while Codex retries. Codex's additionalDetails, which its own UI shows under the progress message, is kept on the fact as the retry's cause and printed on the row's second line. Errors Codex will not retry are unchanged. * fix(native-chat): a transcript draws only the latest row of a provider retry run The shared structured message projection, which both the desktop and the mobile transcript read, collapses a run of retry rows into its latest row. A run is retry rows from the same agent with no other drawn row between them; a row that draws nothing, or a queued send drawn after the conversation, does not split it. The earlier attempts stay in the journal. * test(native-chat): the retry-run render test uses the message list's current props * fix(native-chat): a Codex frame row is named for its connection, so a later one never revises it Frame rows were named provider-frame:codex:<n> from a counter that starts over with every connection, so the first frame row after a reconnect in the same session revised an earlier connection's row in place, at its old spot. Each connection's frame rows now carry the acquisition generation, minted once before the translator is built: provider-frame:codex:<generation>:<n>. Rows already written keep their identities. * fix(native-chat): agents retrying at once each keep one row, read from the row's own agent * test(native-chat): a reconnect's Codex rows are named for the acquisition that received them * fix(native-chat): a retry run is one agent's, so another agent's row never splits it Each agent's rows are drawn apart: the session's own rows are the conversation, and a subagent's rows open in that subagent's section. Splitting a run on any other row in the flat list left two adjacent retry rows on screen whenever another agent wrote between two attempts: a subagent finishing a command while the session reconnected, or the session working while a subagent reconnected. A run is now per agent: an agent's retry rows with none of its own other rows between them, drawn as its latest. * test(native-chat): a subagent's retry run is checked in its own section, and the run rule's words say same agent * test(mobile): the retry-rows test typechecks, so the test ratchet keeps checking it |
||
|
|
24edf0f64b |
fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm (#23467)
* refactor(native-chat): remove the unused terminal handoff
No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.
* fix(native-chat): never let the pre-stop snapshot hold a chat's stop
Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(native-chat): drop helpers only the terminal handoff called
`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.
Co-Authored-By: Claude <noreply@anthropic.com>
* docs(native-chat): stop citing the removed handoff in lifecycle comments
Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): type the stalled snapshot drain without a cast
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): pin that a start dead before proving owes no settlement
The removed restart handoff test pinned this branch; nothing else did.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(native-chat): keep the owner-status read behind an in-flight attach
The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(terminal): remove the agent-session PTY write gate
The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(native-chat): drop the transcript helpers only the handoff called
appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(native-chat): stop calling a starting chat "mid-handoff"
A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): type the stand-in roster decoder without a cast
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(codex): name the pinned rollout lookup for what it does
With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.
* refactor(native-chat): type the owner-status reply as the host sends it
The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.
* refactor(native-chat): normalize terminal-handoff lease values once at decode
Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.
The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:
- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
`conflicted`, the claim every build probes but never stops. A plain native
owner would be stopped by restart recovery, here and in older builds.
Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.
The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.
* refactor(native-chat): stop threading the owner kind through a reservation
A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.
* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else
Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.
* fix(native-chat): name a chat write by its target, not the owner generation
A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.
Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.
Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.
* fix(native-chat): every journal append reaches the chats that are open
A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.
A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.
* test(native-chat): an epoch replacement reaches the open chat
* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map
* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite
The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.
* test(worktree-activation): restore the OMP surfaced-agent resume test
The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.
* perf(native-chat): a publish behind a delivered commit reads nothing
Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.
* test(native-chat): state why the teardown test's fake journal is safe to cast
* docs(native-chat): say mutation admission checks only the writer lease
* docs(native-chat): drop the send rebase from comments that still described it
* fix(native-chat): a message is accepted, then delivered
A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".
A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.
Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.
A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.
Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.
* fix(native-chat): settle queued messages only for the child that ended
A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.
A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.
The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.
* fix(native-chat): an adoption that fails to import keeps the conversation open
The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.
* perf(native-chat): the recovering open reads the journal once
Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.
* fix(native-chat): an attach that fails after indexing its child leaves no child behind
A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.
* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer
The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.
A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.
* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down
The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.
* fix(native-chat): a message rejected while its chat was closed reads as not sent
A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.
* test(orchestration): name why the readiness settlement fakes are cast
* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent
* docs(native-chat): drop the fence from the admission the send effects run behind
* docs(native-chat): give the fence move on release the reason that still holds
* docs(native-chat): stop citing a write fence check in launch and mailbox comments
Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.
* refactor(native-chat): the provider child is its own record
A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.
- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.
* fix(native-chat): the delivery loop alone settles a message its start or child failed
A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.
- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
reads how it ended: a Stop continues; anything else writes one failure row and rejects every
queued message with the same words, then stops. A child still starting whose start the adapter
says did not land fails the same way. The exit, eviction and the settlement retry only settle
the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
closed, with or without a child, and a start the loop already has in flight is waited for so the
child it produces is stopped rather than left behind.
* refactor(native-chat): a stopped child ends on the one reading of its stop
The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.
* feat(native-chat): the host says it accepts a send before any agent has it
The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.
* refactor(native-chat): an attach never opens a journal of its own
The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.
* fix(native-chat): a moved fence resends nothing on a host that accepts first
The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.
The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.
* refactor(native-chat): a child's end says whether the user or the host stopped it
The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.
* fix(native-chat): a chat whose only work is a queued message is not offered for resume
A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.
* test(native-chat): type the queued-message fixtures in the resume-offer tests
* fix(native-chat): a start that dies while a message waits on it is that message's failed start
Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.
* fix(native-chat): a request that failed reads as failed
A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.
The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".
* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now
* test(native-chat): a verdict change republishes the mobile status projection
* refactor(native-chat): the store's retention trigger keeps its flag compare
A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.
* test(native-chat): a user message the provider journaled keeps its session listed
* test(native-chat): pin what a failed start settles, and what a resume offer names
A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.
* test(native-chat): the failed-start pins fail on what the message became, not on a timeout
* fix(native-chat): a late provider-session update keeps a failed recovery record failed
A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.
* test(orchestration): the preamble's host stub is typed, not cast
The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.
* test(native-chat): the terminal-bell check asserts the renamed verdict field
The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.
* fix(native-chat): a failed turn ranks like a completion for attention
Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.
The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.
* fix(native-chat): a failed main agent reads failed while its subagents still work
The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.
Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.
worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.
* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it
The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.
* docs(native-chat): the status-store listing rule names provider-journaled user messages
* fix(native-chat): a refused send notifies failed through the completion feed
The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.
* fix(native-chat): every copy of a row carries the main agent's own status
History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.
- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
rebuilding one; the sync key and history equality compare it.
* test(native-chat): pin the worktree ps verdict across host and phone versions
Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.
* test(mobile): name the parity table's row for its role
* fix(native-chat): a request that settles while the user is asked something notifies once
The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.
The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.
* fix(native-chat): the completion says when the user is being asked
A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.
The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.
* fix(worktree-status): a departed agent's failure yields to live work on the worktree card
A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.
* docs(agent-status): a departed agent's failure ranks below live work on the worktree card
* fix(native-chat): a view never restarts a chat whose last start failed
A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.
* test(native-chat): start the child the loop waits on with an attach, not a second view
A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.
* fix(native-chat): settle a gone generation's turn wherever a conversation opens
A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.
* test(native-chat): prove the next child's start settles the turn an earlier child left
The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.
* test(native-chat): count a failed start's rows by row, not by text
Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.
* test(cross-version): load the phone row readers without mobile's toolchain
Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.
The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.
* test(cross-version): keep the checkout path-guard message and justify the copy import's cast
* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget
* test(native-chat): pin the open's and the send's start and row counts, however the view binds
Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.
* fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm
When the provider gave no verdict, the structured status projection now derives
one from the newest turn's lifecycle: interrupted -> interruption, unverifiable ->
unconfirmed. Nothing new is journaled, the completion feed stays provider-only, and
the legacy interrupted flag stays a user stop only. Every verdict reader handles
both arms explicitly.
* fix(native-chat): settle a gone generation's turn at every open but an acquisition's
The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.
* fix(native-chat): a folded turn a crash cut off reads Interrupted after N
The settled-turn timing now carries the turn's verdict, derived by the same
agentTurnVerdict the status row uses. A turn that ended interrupted with no
provider verdict heads its fold 'Interrupted after N'; a user's stop keeps
'Worked for N'.
* test(native-chat): hold the create's start open until the views bind
The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.
* fix(native-chat): the user's close of a chat records the turn it cuts short as their cancellation
The expected-close settle writes outcome cancellation when the user aimed the
stop at this chat: a Stop while the agent starts, the chat's tab closed (the
agentSession.close RPC, or session.tabs.close with a user reason), or /clear.
A quit, an idle eviction, a worktree teardown or an orchestration stop leaves the
turn with no verdict, so it still reads Interrupted.
* test(native-chat): a Claude turn a newer send superseded reads Interrupted
The supersede fires for any send Orca dispatched, the user's or another agent's,
and nothing at that site records the sender, so the turn keeps no verdict and
folds as Interrupted after N.
* fix(native-chat): the user's close records cancellation on the turn the provider settled on its way out
The Codex and Claude adapters settle their open turn as interrupted, with no
verdict, while the host stops them. The expected-close settle then found no
running turn, so a user's close of a mid-turn chat read Interrupted. The stop
now reads the running turns before it reaches the provider and records the
user's cancellation on each one it cut short, unless the provider gave a
verdict of its own. The close-verdict test's adapter now settles its turn on
close the way the real adapters do.
* test(native-chat): update the close and settled-turn expectations for the host-observed verdict
agentSession.close now passes the user's word to the host, and a settled
interrupted turn with no provider verdict carries `interruption`. Also merge a
duplicate import the code-quality gate rejects.
* refactor(native-chat): drop the composer's second error formatter
After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.
* test(native-chat): pin the reason on a message rejected while its chat was closed
The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.
* fix(native-chat): the user's close cancels a turn whose start landed as the provider stopped
The close read which turns were running before the stop. A turn whose start was
still in flight (a send echo not yet journaled) was absent from that read, so the
provider's verdict-less settle on the way out left it Interrupted. The close now
reads which turns were already over instead, and records the user's cancellation
on every other turn the stop left running or interrupted with no verdict. A turn
cut off earlier, or finished during the stop, keeps its end.
* fix(native-chat): a send the provider never received after a restart has no verdict
Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.
* refactor(native-chat): the adapter settles the turn a stop cuts with the stop's typed cause
The host hands its stop's cause to the adapter's close. Each adapter settles its own
open turn on 'ended' through one mapping, turnVerdictForChildEnd, and the host's
dead-generation fallback uses the same mapping for any turn no adapter settled. A
user's close or stop of this chat is their cancellation; a quit, eviction, teardown
or an exit the adapter saw first is news.
Deletes the snapshot-and-diff reconstruction (endedTurnItemIds,
userStoppedTurnRevisions, settleUserStoppedTurns) and the requestedByUser flag.
host.close now takes a required cause.
* test(native-chat): a user's close drops the chat's status row like an eviction
* test(native-chat): expect the eviction cause on the host closes of idle release, worker stop and worker discard
The stop's typed cause now travels into host.close and the adapter's close, so these
three non-user closes assert the 'evict' they pass.
* refactor(native-chat): every stop names its cause, so none defaults to the user's cancellation
stopStructuredAgentSessionAgentUnderSerialize defaulted its ending to 'user-stop', which now
settles the cut turn as the user's cancellation. Every caller already passes a cause; the
parameter is now required, and a type-level test fails to compile if the default returns.
* test(native-chat): pin who a chat's session.tabs.close speaks for, older clients' reasonless close included
The mapping lived inline in a type-unchecked file, and only the explicit user reason had a test: an older client's reasonless close, or a lifecycle echo read as the user's, stayed green. It is now one exhaustive, type-checked function with a case per reason.
* style(mobile): draw the unconfirmed dot in the theme's status amber, not an inline hex
* docs(agent-status): the main agent's outcome also carries the host-observed end, interruption or unconfirmed
* fix(native-chat): a chat the user closed while its agent started is not a failed start
A still-starting child the user's close cut counted as a failed start, since only 'user-stop' was excluded: a start-failure row, and queued messages rejected as a provider failure. Whether an ending fails its start is now one exhaustive switch, shared by the failed-start read and the delivery loop's handover: a user's stop or close never does; an exit, a failed attach and the host's own stops still do.
* fix(native-chat): a chat the user closed closes its queued messages, and starts no agent for them
After
|
||
|
|
afa81dc3ad |
fix(native-chat): chat failure messages appear in the app's language (#23674)
* fix(native-chat): a read whose history will not open is refused with its reason
History, subscribe, snapshot and options reads reach a chat through one accessor, whose open had no
catch: a journal that would not open reached every client as a runtime error carrying the storage's
own text (a path, "file is not a database"). The accessor, and the options read's own open, now throw
the classified journal refusal: journalCorrupt when SQLite reports damage, journalUnavailable
otherwise. The storage text goes to the log only.
The wire code stays runtime_error and the message becomes the bare code, as for every thrown
refusal; the reason rides in the error's data.
* refactor(native-chat): the idle sweep's stop of a hung start carries no hand-written reason
The sweep passed an English sentence as the stop's reason. It lived only in memory and nothing read
it: the delivery loop words the error row and the rejection from the hostStopped fact. Dropped, with
the display-name lookup that built it.
* fix(native-chat): a refusal the host throws is worded from its data, never its message
A thrown agent-session refusal reaches the client as runtime_error with the bare code as its message
and the typed refusal in error.data. Stop, answers, options and goals, a launch's held option pick,
the option picker's failure toast, and the Retry line of a chat that could not start now word it
from that refusal through the shared notice table. What each caller decides about the outcome is
unchanged: only the words move.
The Retry line of a failed start no longer prints the host's message or a thrown error's text; it
keeps the refusal as a fact and says the cause and step its reason names, or only that the chat
could not be started. The option toast keeps a local option surface's own sentence.
* fix(native-chat): an unreadable history is worded from its refusal, and damage stops the retry
The structured chat's read failure showed the host's text on the status line, and the pane always
said Orca keeps trying. The read transport now takes the refusal from the error's data (a stream
payload or a thrown RPC error), the reducer keeps it beside the failure text, and the pane and the
status line word it through the notice table, once: on the pane when the failure took it, else
beside the transcript that stays.
A damaged journal says "Unable to load this chat." and the read stops reconnecting for that run;
reopening the chat reads again. An open that can clear names its cause without "Try again", since
the pane retries on its own. A failure that names no reason keeps today's generic line. Finality
comes from the refusal's reason, never its message, which is the bare code for both.
* fix(native-chat): a rejected message is worded from its stored fact
A message the host recorded and then rejected keeps the host's typed fact beside its reason, but
the Retry words re-read the reason alone. Now the fact decides: a hand-over failure says Orca
couldn't reach the agent, a kind whose reason may be a legacy marker gets its fact's own sentence
(a full queue now says so instead of only "not sent"), and any other kind shows the sentence the
host wrote for it, which carries the agent's name and any words the provider wrote for a person. A
row with no fact reads as before.
* fix(native-chat): each message that did not go through says why on its own row
The structured chat showed one Retry strip under the transcript for whichever single entry it
picked, so a second failed message had no reason and no Retry of its own. The terminal-backed
chat's existing per-row delivery marker now carries a notice and an optional Retry, and the
structured pane derives one per message from the outbox on each render: every rejected message,
and the one the queue stopped on (read through the drain's own rule, so a Retry never names a
message waiting behind it). Each is worded from that message's stored failure. The single strip is
deleted. Nothing new is stored, and the shared message projection is untouched.
* fix(native-chat): say each chat failure's words where its own control already acts
Three wording rules for the desktop chat:
- A chat that could not start shows Retry beside its reason, so the reason stops at its cause
where the Retry is the step: a reason whose action is to retry, and a start failure's "send
your message again". Any other step stays (quit the terminal agent, start a new chat). The
start-failure sentences take a retryControl context for this; what the host writes is unchanged.
- A history that couldn't open right now still reconnects, so the pane keeps "Orca keeps trying
to load it" under its cause. Only a damaged history, which no retry reads past, drops it.
- A read failure that names no reason while the transcript is shown is only the pane
reconnecting: the status line says "Reconnecting to this chat…" in muted text, not an error.
New key components.native-chat.state.reconnecting, hand-translated for es/fr/ja/ko/zh.
* fix(native-chat): a rejected message offers Retry only once the queue is moving
Each rejected message's row offered its own Retry even while the queue was stopped on another
message. Any Retry clears the stopped queue, so pressing a rejected message's Retry also sent the
message the queue was holding, which the user had not retried; behind a message whose delivery is
unconfirmed, the retried one instead went back into the queue with no notice and waited there.
While the queue is stopped, only the message it stopped on offers Retry, as the single Retry strip
this replaced did. A rejected message keeps its words on its row and gets its Retry back once the
queue moves.
* fix(native-chat): a chat whose history will not open logs once, not on every reconnect
A reader reconnects every 750 ms while a journal open can clear, and each attempt logged the
failure with its full stack. The read door now logs a session's failure once until that session
opens, closes, or fails differently; every attempt is still refused with its reason.
* fix(native-chat): a message's own Retry is its resend step, so its notice stops at the cause
A rejected message offering Retry read "Claude stopped before it finished starting. Send your
message to try again." beside that button. Its row now takes the rule the launch strip already
follows: beside its own Retry the words leave out sending or trying again, worded from the stored
fact with the chat's agent name. The stored fact keeps less than the host wrote from (a refusal,
the provider's words), so a reason it cannot rebuild exactly is kept as written. A rejected
message without a Retry, while the queue is held, keeps the step. What the host writes and the
phone's notice are unchanged.
* fix(native-chat): a message's Retry sends only that message, never the one the queue is held on
Retry released the queue's refusal hold whichever message it was pressed on. While a queued message waited ahead of a held one, every rejected message offered Retry, and pressing it also sent the held message the person had not retried. Retrying an unconfirmed message ahead of a held one did the same. Retry now releases the hold only for its own message.
* fix(native-chat): a not-signed-in failure beside Retry still says to sign in first
Beside a Retry the notice dropped the whole next step, so a chat that could not start because the agent was not signed in read only the cause. Pressing Retry without signing in fails the same way again. The words now keep the sign-in step and leave out only the resend, which the Retry button is.
* test(native-chat): a rejected message's hidden Retry only avoids waiting unseen
* fix(native-chat): every Retry beside a notice leaves out the retry step the same way
A message the queue stopped on worded its refusal with no agent name and with its retry step, beside its own Retry, while a rejected message next to it named the agent and left the step to the button. The launch strip and the history pane each had their own copy of the same rule. One wording context now goes through the one notice table for every surface: a Retry beside the words, or a pane that reconnects on its own, is the step for a reason whose action is to retry, and every other step stays. What the phone and the host write is unchanged.
* test(native-chat): read the sent message id without a type assertion
* fix(native-chat): a rejected message is worded from the journal's own fact, never by comparing sentences
A message the host recorded and then rejected kept only the rejection's kind on the message, so its notice was reworded from that smaller copy only when it rebuilt the host's sentence word for word. A different agent name, an older host's wording, or anything the copy dropped (why a start failed, the provider's own words) left the host's sentence in place, beside a Retry that repeated its resend step. The notice now reads the journal's own rejection for that message, found by id, with the pane's agent name and Retry, and shows the provider's words only when they were written for a person. The message's smaller copy words it only when that journal row is not loaded. Nothing new is stored.
* fix(native-chat): a message rejected before a restart retries under a new id the first time
Whether a Retry needed a new message id was remembered in memory for one message, or read from the journal row when it was loaded. After a restart, or for an older message whose row was not loaded, the first Retry resent under the old id, the host answered with the same settled rejection, and nothing visibly happened. The message now says so itself: one the host recorded and rejected always retries under a new id, including after a restart. A refusal that already gave the message a fresh id, and a message whose delivery is unconfirmed or in flight, keep their id as before.
* fix(native-chat): a rejected message older than the loaded history keeps the provider's words
When the journal row that rejected a message is not loaded, the message's own copy of the
fact has no provider detail or start refusal. For the kinds worded from those, the row now
shows the sentence the host wrote for the person instead of a thinner rebuilt one.
* test(native-chat): the chat pane words a rejected message from its loaded journal row
Nothing covered the pane handing the journal's rows to the per-message notices, so a pane that stopped passing them would quietly fall back to the message's smaller copy of the rejection and show the host's sentence, resend step and all. The new case renders the pane with a rejected message whose journal row is loaded and checks that it reads that row's refusal in the chat's own agent name.
* fix(native-chat): a chat whose history won't load says why in one line
A read the host refused for a named reason put its sentence under the generic
"Could not load conversation" title, so a damaged history read as two lines
saying the same thing. The pane's own sentence now takes the title's place; a
history that can come back keeps its line saying Orca keeps trying. A failure
that names nothing keeps the generic title.
* fix(native-chat): a message a failed start rejected says only that it was not sent
When an agent stopped before it finished starting, the chat showed the start's
red row ("Claude stopped before it finished starting. Send your message to try
again.") and then repeated that cause under every message the start rejected.
Each of those messages now reads "Your message was not sent." beside its Retry.
The match is made on typed facts, not on the words: the host writes the start's
row and the rejection of its queued messages from the same failure fact, and the
row is keyed by the start. The pane finds the loaded start-failure rows by that
key and shortens a message's notice only when its loaded journal submission was
rejected with the same fact. Any other rejection, or one whose submission or row
is not loaded, keeps its full notice. The row key moves to a shared module so the
host that writes it and the pane that reads it use one definition.
* fix(native-chat): a chat whose history keeps failing to open retries less often
A read the host kept refusing (its history store could not be opened right now)
reopened every 750 ms for as long as the chat stayed open, about 40 opens every
30 seconds. Each reconnect now waits twice as long as the last, from 750 ms up
to 30 seconds, and never gives up; the first read that delivers anything starts
the wait over at 750 ms. A damaged history still stops reconnecting at once.
Reset happens on a delivered read, not on connect: a local subscribe resolves
before the host's open refuses, so resetting there would keep the 750 ms loop.
* test(native-chat): the pane harness types its journal rows without a cast
* test(native-chat): the admission test passes no start-failure rows to the notices
* test(native-chat): import the journal types once
* fix(native-chat): a remote chat reads again as soon as its host is back
The read retry doubles its wait up to 30 s during an outage, and nothing
reset it when the remote runtime reconnected, so the transcript could
lag the reconnect by up to 30 s. The read now watches the runtime
status store's contact-regained edges (hostContactEpoch for a
same-runtime return, connectionGeneration for a new runtime session)
and, when one lands, runs a waiting retry immediately with the wait
reset to its base.
* refactor(native-chat): build each failure sentence from whole pieces
Every sentence agentSessionFailureWords writes is now assembled from a
table of whole English pieces, so a reader can supply its own words for
each piece. The host still fills them in English, byte for byte as
before.
* fix(native-chat): translate the failure sentences desktop notices show
A refused start and a rejected message now carry their failure fact to
the notice instead of its English sentence, and desktop words that fact
through translate keys whose English defaults are the host's own
pieces. The host keeps writing English into rows and reasons, and a
host sentence with no fact beside it still shows as written.
* fix(native-chat): the history pane says only that Orca keeps trying
When the pane's title already says Orca couldn't open this chat's
history right now, the line under it no longer repeats that the
transcript could not be read; it says only that Orca keeps trying to
load it. The pane with no named reason keeps its two-part line.
* test(native-chat): type the failure pieces a refusal notice shares
* test(native-chat): the Chinese failure words use no Japanese-only characters
* fix(native-chat): every history pane that says it didn't load says only that Orca keeps trying
A pane whose title is a code's own words ("This chat's history couldn't
be loaded.") now gets the short retrying line too. Only the pane with no
named reason keeps the two-part line.
* test(native-chat): a provider's words with nesting and markup stay as written in a translated notice
* fix(native-chat): French and Spanish say a withdrawn message was withdrawn before the agent began working on it
* fix(native-chat): Japanese and Chinese notices run their sentences on without a space
A notice joined its sentences with a space in every language, so Japanese and
Chinese read "Claude 无法启动。 请重新发送消息。" with a stray gap after the full
stop. The failure-sentence builder now takes the joiner alongside its words, and
desktop joins in the UI language: no space in Japanese and Chinese (including a
plugin pack that declares either), one space elsewhere. The host and the phone
keep English, joined with a space as before.
* fix(native-chat): a failed /clear or /compact says why in the app's language
The line under the composer printed the host's English sentence although the
result carries the typed failure beside it. It now words that failure the way
the host does (the chat's agent and /clear for a failed /clear, nothing for
/compact), in the app's language; an older host that sends no failure keeps
its sentence.
* fix(native-chat): Spanish says a rate limit, and Korean says a withdrawn message was never processed
The Spanish retry notice said the agent hit a usage limit, a different thing
from the rate limit the English names. The Korean withdrawn-message notice
said the agent had not started, which reads as the agent not launching; it now
says the agent had not begun processing the message, as the other languages do.
* refactor(native-chat): one rule says which words already say the history didn't load
* fix(native-chat): an image size limit says its unit the way the reader's language does
* test(native-chat): the option picker's i18n stand-in knows the reader's locale, which a refusal notice now reads
* fix(native-chat): a sentence a language pack left in English keeps its space
Sentences were joined by the UI language: no space in Japanese and Chinese,
one elsewhere. A plugin pack for a Chinese or Japanese variant that predates
the failure words falls back to English for them, so a notice read
"您的訊息未傳送。Claude couldn't start.Send your message to try again."
Each gap now follows the sentence before it: none after a full-width 。!?,
one space after anything else. The built-in Japanese and Chinese catalogs end
every sentence in 。, so they read as before, and English is unchanged. Since
the rule no longer needs the language, one joiner serves every surface and
the failure-sentence builder no longer takes one alongside its words.
* fix(native-chat): a failed /clear this build only partly understands shows the host's own sentence
The line under the composer words a failed /clear or /compact from the fact
the host sends beside its sentence. The reader drops any part this build
cannot place, such as a refusal code a newer host added, and the rest of the
fact can then give different advice: "Run /clear again." where the host said
"Start a new chat to continue."
When any part the host sent did not survive the read, the line now shows the
host's sentence as written, the same as for a host that sends no fact. A fact
this build reads whole is still worded in the app's language.
* fix(native-chat): a failure fact this build reads only in part shows the host's sentence everywhere
The previous check compared only a fact's top-level parts, so a known refusal
code carrying a reason a newer host added still counted as read: the reader
dropped the reason and the notice re-worded what was left, which can advise
differently from the host ("Run /clear again." against "Start a new chat to
continue.").
One shared reader now answers whether this build read the whole fact: it reads
the fact and keeps it only when the read equals what arrived, at every depth.
Every place that chooses between wording a fact and showing the host's text
uses it: the line under the composer after /clear or /compact, a rejected
message's notice from the journal's fact, and the smaller copy a rejected
message keeps for when its journal row is not loaded. Matching a rejected
message to the start row that already says why still uses what this build can
read, since that is identity, not wording. Facts this build reads whole are
worded as before.
* fix(native-chat): a failed /compact names /compact as its next step on desktop too
The host names the command a failed start was waiting on, for /clear and /compact alike.
The line under the composer re-worded only /clear with it, so a /compact whose start failed
read "The agent couldn't restart. Send your message to try again." in the reader's
language. It now words every command the host answers with the agent and the command the
host used, so desktop English matches the host and French or Japanese keep /compact.
* fix(native-chat): a /compact whose start failed no longer says the operation was not confirmed
A /compact on a chat whose agent is not running starts it first, and that start takes a new
lease, so the chat's fence moves before the command's reply arrives. The write settles as one
for a fence this pane no longer shows, and the composer line read that as "Conversation
operation was not confirmed." Such a write now says nothing there, as every other write
already does: the chat's own start-failure row says why, and the command's message is
rejected in the journal. The sentence it printed is gone from the catalogs.
* fix(native-chat): keep the message outbox within its line limit after main's growth
* fix(native-chat): a command's own reply is kept when its start moved the fence
A /clear or /compact on a chat whose agent is at rest starts the agent first, and that start
moves the chat's fence before the command's reply arrives. Every reply from an earlier fence was
discarded, so a /clear whose new chat failed to start said nothing at all, and a /compact that
started left "/compact" in the composer.
A conversation command's reply is now kept while the pane still shows the chat it was sent for;
a closed pane or another chat still drops it, and every other write keeps the fence rule. The
line under the composer says nothing only when the failure is the chat's own start and that
start's loaded row already says why, the rule a message that start rejected already follows.
* fix(native-chat): a returned queued message shows the host's sentence for a fact read in part
The caption under a returned queued message re-worded its failure from whatever this build could
read of the fact. A newer host's fact with a refusal code or reason this build drops read as a
shorter sentence with different advice. It now re-words only a fact read whole, and otherwise
shows the host's own sentence, as every other surface that re-words a fact does.
* test(native-chat): a command reply after any fence move, for /clear and /compact
The fence-move tests now state the rule as the code has it: a conversation command's reply is kept
whenever the pane still shows its chat, whatever moved the fence. They cover a /clear that
completed, and a failed start for each command in French with the fact that command really meets
(a /clear's new chat fails to start; a /compact's chat fails to restart).
* fix(native-chat): only the reply a pane still waits on outlives a fence move
A conversation command's reply was kept across a fence move whenever the pane still showed the
same chat. A reply the pane had stopped waiting on, because it left the chat and came back or a
newer command replaced it, was applied as if it answered the current one.
The pane now remembers the one command request it waits on; only that request's reply is kept
after the fence moves, and closing the pane or showing another chat forgets it. Tests now drive the
fence move during the request itself rather than through /clear starting an agent, and cover a
/clear that stops a running agent.
* fix(native-chat): a newer-Orca history error keeps a whole retry line, and the phone shows the host's words for a fact it reads in part
When a chat's history was saved by a newer Orca, the pane's title reads "Chats were saved by a newer
Orca. Update Orca to keep using them." and the line under it said only "Orca keeps trying to load
it.", with nothing for "it" to mean (in French and Spanish the pronoun also disagreed with "chats").
That title names every chat rather than this one, so the pane keeps the full line: "The transcript
could not be read. Orca keeps trying to load it."
On the phone, a returned queued message whose failure fact this build reads only in part was
re-worded from what it could read, dropping advice the host gave; it now shows the host's own
sentence, as the desktop card already does.
The comment on the reply the pane waits on now says what forgets it: a newer command, or disabling
the pane.
* fix(native-chat): a command reply this build can't place shows the host's own words
A newer host can answer a conversation command with a command name this build doesn't know. This
build re-worded that reply from its failure fact as if it were a command it knew, naming the
command in its own words, or said nothing when a loaded start row matched the fact. It now shows
the host's sentence as written, as it already does for a fact it can read only in part.
|
||
|
|
85f8d6b5f5 |
test: retire long-tail cases whose assertion is decided by the test itself (#24132)
Resumes the backlog sweep at a chunk size that actually gets read. Six auditors, 84 files
each, and all six read their full scope case-by-case against production — the first wave
where every chunk closed with no gap. 33 case declarations removed across 22 files, 1 test
file deleted, 826 lines gone. No production code touched.
This wave exists because a conclusion of mine was wrong. I had recorded that yield collapsed
~36x and that deletion was no longer the high-value work. I was dividing cases removed by
files IN SCOPE while the fraction auditors actually READ fell from 100% to about 4%, because
I kept handing them 300-800 files. Recomputed against files read, yield has been flat at 4-7
per 100 with no downward trend. This wave came in at 8.2.
The most instructive removal looked like the most valuable test in scope.
`orchestration-worker-release-reap-fixed.func.test.ts` cites a production bug by two
identifiers, describes orphaned PTYs accumulating until `TasksMax=4096` aborts processes on
EAGAIN, and advertises itself as the functional tier wiring the real orchestration RPC
surface, the real `OrchestrationDb` and the real release modules. Deleting it leaves no
reference to that bug anywhere in `src`.
It still had to go: its fake runtime performed the fence it asserted —
if (pty.incarnationId !== inc) { return null }
handleTable.set('term_reminted', { ptyId, epoch: rendererGraphEpoch })
— so the case checking that a reused ptyId with a mismatched incarnation does not resolve was
checking a decision its own spy made twenty lines earlier. The real fence is owned by
`orca-runtime-terminal-handle-incarnation.test.ts:257`, and the other two cases replay
`orchestration-worker-release-incarnation-fallback.test.ts` (which uses a plain
`mockReturnValue` rather than reimplementing the remint) and `worker/worker-release.test.ts:23`.
"Integration test" and "wires real modules" describe the scaffolding, not the asserted step.
Other removals: a self-comparison disguised by an alias, where
`export const getIssueOwnerRepo = getOwnerRepo` makes a case asserting the two "agree" into
`f(x) === f(x)`; four cases whose `vi.mock` of `resolveIssueSource` made both the preference
value and the topology inert; five verdict-precedence cases owned by a verdict-agnostic block;
three call-shape probes on one-line store pass-throughs whose real contracts are driven by
behavioural neighbours; and a `export type _Ref = [...]` declaration whose own comment admits
it exists only to preserve test-only module-surface references.
Kept after checking production rather than shape. An auditor found two near-identical
ten-reconnect loops and kept both: one uses a test-local live-lease filter, the other the
shipped `sshRemotePtyLeaseAllowsReattach` predicate, and the file's own comment explains the
duality is deliberate "so the two cannot drift". Another kept a paths-alignment case that
looks like a validator tested against its own list, because adding a generated file without
registering its path does fail it — and `shellReadyWrappersExist` uses that registered list to
decide whether a partial tree needs regeneration.
Production duplication is now confirmed four times over, and it is why mirrored tests exist:
`createUpdateWorktreeLineage`/`createAssignWorktreeParent` differ by one `console.error`
string; `terminal-path-tap.ts` and `document/path-tap.ts` carry hand-maintained copies of
`matchFilePathAtColumn` under a docblock reading "keep the two in sync". In those cases both
test sides are load-bearing and the duplication belongs on a refactor list.
`mobile/tests-typecheck-baseline.txt` loses one entry. Trimming
`relay-host-signed-out-verdict.test.ts` made it typecheck clean, so the ratchet required
pruning its grandfathered entry — the file graduates from exempt to enforced. Baseline is now
124 entries, down from 125.
Verified: 690 test files / 7,560 cases pass across the touched desktop areas; the modified
mobile files pass (162 cases); `check-tests-typecheck-ratchet.mjs` OK (898 files in program,
124 grandfathered); `check-reliability-gates.mjs` 140 gates; the deleted file is absent from
the gate manifest, `cloud/package.json` and the mobile baseline; nothing under
`mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` touched.
|
||
|
|
a781a602a8 |
test: retire duplicate cases that replay an owner across a re-export or provider shim (#24114)
Resolves 208 candidate pairs where the same case title appears verbatim in two or more
files, produced by a repo-wide scan calibrated against a known positive. 46 case
declarations removed across 32 files, 798 lines gone. No file deleted whole, no
production code touched.
The headline result is the measurement, not the deletions: across the three buckets that
reported in detail, the signal ran roughly 86% false-positive (3/42, 9/42, and the rest).
It has good recall and poor precision, and it reorders a reading queue rather than
replacing one. Calibrating a detector against a known positive proves recall, not
precision.
What the deletions were:
- Duplicate invocation through a re-export shim. `native-chat-tool-summary.ts` is a
ten-line `export {...} from '../../../../shared/native-chat-tool-summary'`, and
`agent-status.ts:161` is `export { isExplicitAgentStatusFresh } from
'./pane-agent-evidence'`. Cases on the shim side were byte-equivalent to the owner's
with no rendering or transport hop.
- Provider-local replays of a shared helper: three `repository-ref` providers that are
each `createRemoteRefProbeCache(parseXRef)` and contribute nothing to transient
handling; two `local-pty` and `daemon/session` tables replaying
`shell-startup-output-scanner`, whose owner additionally checks every split point.
- A reader-side replay of store policy. `runtime-worktree-agent-rows-structured.test.ts`
asserted an attention-to-blocked mapping; the reader contains zero `attention` or
`blocked` tokens and copies `state` through. The mapping lives in
`structuredAgentSessionAgentStatus`. Consistent with
`docs/reference/agent-status-store.md`: readers keep only presentation policy.
- Constructor-only subclass duplication: the shared capability-cache case is covered by
`codex-app-server-capability-cache.test.ts`, whose ten cases include the identical
title plus all four risks `docs/reference/git-compatibility.md` names — first fallback,
later cached call, concurrent probes, per-host isolation.
- A private predicate duplicated at a real boundary, varying only a path passed straight
into the shared predicate.
Why most pairs were KEPT, because the false positives are principled rather than noise:
- Two independent execution hosts. `src/relay/git-handler-*` and `src/main/git/*` are
separate Git implementations that cannot import each other and hold separate capability
caches, exactly as the compatibility doc requires; the repo already ships
`status-branch-line-total-relay-parity.test.ts` to pin the duality deliberately. Neither
side's argv, timeout or cache regression is visible to the other.
- Deliberately duplicated production siblings: Codex vs Claude (different account fields,
different CLIs, different wire protocols), gitea vs bitbucket (`/pulls/42` vs
`/pullrequests/42`), gl-utils vs gh-utils (separate in-flight maps). Same contract
shape, different implementations — an identical title is the correct naming.
- Shared-predicate consumers: one side tests the predicate, the other tests a caller's
wiring to it. A caller that forgot to call the predicate passes the shared test.
In a codebase with intentional provider and host symmetry, identical test titles are
expected, and the signal cannot distinguish "copied" from "parallel by design" because
both produce the same prose. Only reading both bodies separates them.
Verified: 6,968 desktop test files pass; the three modified mobile files pass (39 cases);
`check-reliability-gates.mjs` 140 gates; nothing under
`mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` touched.
62 local failures across 12 files were each accounted for and none is caused by this
change: `browser-manager-tab-identity`, `browser-manager-viewport-ownership`,
`session-scanner-codex-workers` and `managed-hook-script-refresh` all fail identically in
a pristine `origin/main` worktree; five `mobile-web-app-*-render` tests need Playwright
browsers this machine lacks; `structured-agent-session-restart-ownership` and
`ssh-remote-commands` pass in isolation and fail only under concurrent load.
|
||
|
|
cef66fbab8 |
test: retire long-tail cases whose input cannot reach the behavior they name (#24101)
Sweeps the triage-only backlog: 2,269 files that earlier waves saw and skipped for size, reconstructed from the unread lists five waves of auditors disclosed. 35 case declarations removed across 15 files, 2 test files deleted, 487 lines gone. No production file touched. These are large integration suites, so the junk here is individual cases buried among real coverage rather than whole bad files. The dominant defect was again a case whose input cannot reach the behavior its title names: - `resume-sleeping-agent-session-remote-compat.test.ts` (deleted) — two cases titled for "transport-level host authority on a capable host" and "host authority is not known". `resume-sleeping-agent-session.ts` has no host-authority or capability concept at all, and its only read of `origin` is `if (!record.origin && record.state === 'done')`, unreachable for both rows. Both executed one identical path. The surviving contract is owned by `resume-sleeping-agent-session-execution-host-scope.test.ts`, which drives a real host catalog. - `project-group-header-drag.test.ts` (deleted) — four cases setting `data-project-group-header-id`, which the predicate never reads. Its subject, `isProjectGroupHeaderActionTarget`, is byte-identical to `isRepoHeaderActionTarget` apart from the function name and imports the same `REPO_HEADER_ACTION_SELECTOR`, so all four cases were a strict subset of `project-header-drag.test.ts` using identical `data-repo-header-*` fixtures. - `remote-worktree-history-cleanup.test.ts` — "repeats idempotent cleanup through the PTY owner" against a six-line best-effort forward with zero dedupe state. The case called it twice and asserted the mock recorded two calls, which is arithmetic over the test's own loop; nothing about idempotence was established. Also removed: - Runtime assertions of type-level facts, where production already makes the check at a stronger boundary: `const adapterSatisfiesPort: AdapterIsPort = true` followed by `expect(...).toBe(true)` — unconditionally true, while `createExpoGenerationFileSystem(): GenerationFileSystem` is explicitly annotated and passed into `createGenerationStore` at a typed call site. And a case named "does not typecheck" whose runtime assertion is a length check on its own literal, declaring its own local annotation so it could never notice the production annotation weakening. - Private predicate tests duplicated at a real boundary: four `repo-slug-cache` cases delivered by `repo-slug-index.test.ts`, which drives the same resolution through the hook, the real store and the preload bridge, while the cache-level versions hand-seed the internal map and break on a cache-key format change. - Duplicate invocations owned at the shared boundary, including commit and push recovery cases owned by `src/shared/source-control-recovery-agent-command.test.ts`. Kept deliberately, verified rather than assumed: the production duplication behind the deleted drag test was left alone, because `REPO_HEADER_ACTION_SELECTOR` ends in generic `button, a, input, textarea, select`, so genuine action targets inside a group header still match — it is an unspecialised copy-paste, not a live bug, and collapsing two functions is a refactor. Reported instead. Auditors' probes produced 20, 11 and 13 candidate hits for the signature-versus-title shape across their chunks; every one was inspected and every one was genuine coverage. No deletion in this wave rests on a probe alone. Coverage is partial and stated as such: of 2,269 files, roughly 100 were read case-by-case and the remainder reviewed at title-plus-import level. Each auditor listed its own unread set. The largest remaining surfaces are `src/main/agent-hooks` (95), `src/main/claude` (100), `src/renderer/src/lib/pane-manager` (62) and the 20 largest sidebar suites. Verified: 2,583 desktop test files / 25,636 cases pass, plus one pre-existing `it.fails` marker; the two modified mobile files pass (57 cases); `check-reliability-gates.mjs` 140 gates; `check:code-quality:changed` 0 new findings. Both deleted files confirmed absent from the gate manifest, `cloud/package.json` and `mobile/tests-typecheck-baseline.txt`. |
||
|
|
b99462ac1c |
test: retire mobile, cloud, config and e2e cases their input cannot reach (#24077)
Completes the first pass over every test area in the repository. Sweep over `mobile/src`, `config/scripts`, `cloud/`, and `tests/` (1,494 files in scope, with the 24 files under `mobile/src/test-support/rpc-recording/` deliberately excluded). 31 case declarations removed across 17 files, 2 test files deleted, 356 lines gone. What went, by pattern: - Cross-boundary replays of a shared helper. A whole mobile file re-ran `extractPendingAsk`/`parseAskFromStatus`/`formatAskAnswer`, all owned by `src/shared/native-chat-ask.test.ts`, `native-chat-ask-fifo.test.ts` and the renderer's interactive-prompt suite — one case title was verbatim identical to the owner's, and the owners' inputs are supersets. The mobile file imported the shared module directly and exercised no mobile transport, lifecycle or rendering. - A case whose input cannot reach the behavior its title names: "arms it on Android while the drawer is open", where `use-back-claim.ts` has zero Platform/OS references, so flipping the mocked OS changes only shadow styles. - Identity copiers, including one asserting `prSidebarRenderBranch(state) === state.kind` against a production body that is `return state.kind`. The function stays; it has three live callers. - A test of the runtime rather than the product: a case asserting Node's own `EventEmitter` crash contract on a bare emitter, with zero production code in the path. The guard it documents is exercised behaviourally by the case after it. - Duplicate invocations, one of them provable rather than eyeballed: with `MODULE_SCOPE_ENV_WRITER_PIN = 0`, `files.size <= 0` is strictly implied by the sibling's `expect(offenders).toEqual([])`, since a non-empty `offenders` forces `files.size >= 1`. The pin's own doc says it may only ever be decreased from 0, so it could never become a meaningful bound either. Its policy guidance survives as a comment; the file's real ratchet and its regex self-test both stay. - Expected values produced by the test's own arithmetic, and a p95 case strictly implied by a sibling that already pins exact p95 and exact max over a wider range. One production line goes: the `export` keyword on `assignmentCleanupSteps` in `cloud/apps/relay/src/assignment-cleanup-steps.ts`. The function itself stays and is still called internally; only the test-only export was orphaned. Kept deliberately: everything a gate cites, checked by case title and not only by file path; a gate-cited case that does not deliver its claim (reported instead — see below); a cross-version wire cell whose ledger is never invoked, left under the raised bar for wire coverage; and every limit, bound, quota and provenance guard. Nothing under `mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` was touched — those bytes feed a `recorderSha256` digest pinning 398 golden recordings. Verified: `mobile` vitest over the modified mobile files (8 files, 50 cases); `mobile/scripts/check-tests-typecheck-ratchet.mjs` OK (898 files in program, 125 grandfathered, none @ts-nocheck); relay suite 799 passed; `check-reliability-gates.mjs` 140 gates; both deleted files confirmed absent from the gate manifest, `cloud/package.json` and `mobile/tests-typecheck-baseline.txt`. Seven local failures were investigated and none is caused by this change: five `mobile-web-app-*-render` tests drive `playwright-core` chromium/webkit and need browsers this machine lacks, `release-checkout.unit.test.ts` needs cross-version git refs, and `e2e-worker-env-isolation.unit.test.ts` fails identically with its HEAD content restored — it recurses `tests/e2e` with symlink-following `statSync` and no depth guard. |