mirror of
https://github.com/stablyai/orca.git
synced 2026-10-01 08:01:56 +00:00
d2dbe2c385c30fb02be12dceb71fc5f19a560ecc
12192
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d2dbe2c385 |
fix(windows): replace the managed CLI launcher with a native one (#24094)
* docs(security): add the antivirus clearance path for future releases Every AV false positive here has been handled one vendor and one shipped version at a time. Document the programs that clear future releases instead -- signer and product enrollment rather than per-build sample submission -- and add a script that reports an RC's current detection state by hash, so a verdict is found before users meet it in an issue report. Hash lookup only by default; --upload transmits the artifact and stays manual. * fix(windows): replace the managed CLI launcher with a native one resources\bin\orca.exe was a csc-compiled MSIL assembly: a small, freshly compiled .NET image in a user-writable directory that mutates environment variables and proxies a child process. That is the shape .NET dropper heuristics are trained on, and every verdict against it named the family -- MSILHeracles from two vendors, Wacatac!ml from a third. Signing the file does not change its shape, so signing never cleared it. Rebuild it in Rust. Same resolution, same environment contract, same argv passthrough that keeps newline-bearing orchestration bodies intact (#8374), and the child still inherits our environment block rather than an explicit map, so a block carrying both PATH and Path survives (#12046). The PE now carries publisher, version, icon and an asInvoker manifest from build.rs. Refs #23383 * ci(windows): install the Rust toolchain before building the CLI launcher The hosted runners happen to ship cargo, but a real Windows dev box does not -- verified on our own Windows QA host, where cargo and rustc were both absent. Relying on the image means a future image change fails deep inside electron-builder's native hook instead of at an obvious step. |
||
|
|
cef66fbab8 |
test: retire long-tail cases whose input cannot reach the behavior they name (#24101)
Sweeps the triage-only backlog: 2,269 files that earlier waves saw and skipped for size, reconstructed from the unread lists five waves of auditors disclosed. 35 case declarations removed across 15 files, 2 test files deleted, 487 lines gone. No production file touched. These are large integration suites, so the junk here is individual cases buried among real coverage rather than whole bad files. The dominant defect was again a case whose input cannot reach the behavior its title names: - `resume-sleeping-agent-session-remote-compat.test.ts` (deleted) — two cases titled for "transport-level host authority on a capable host" and "host authority is not known". `resume-sleeping-agent-session.ts` has no host-authority or capability concept at all, and its only read of `origin` is `if (!record.origin && record.state === 'done')`, unreachable for both rows. Both executed one identical path. The surviving contract is owned by `resume-sleeping-agent-session-execution-host-scope.test.ts`, which drives a real host catalog. - `project-group-header-drag.test.ts` (deleted) — four cases setting `data-project-group-header-id`, which the predicate never reads. Its subject, `isProjectGroupHeaderActionTarget`, is byte-identical to `isRepoHeaderActionTarget` apart from the function name and imports the same `REPO_HEADER_ACTION_SELECTOR`, so all four cases were a strict subset of `project-header-drag.test.ts` using identical `data-repo-header-*` fixtures. - `remote-worktree-history-cleanup.test.ts` — "repeats idempotent cleanup through the PTY owner" against a six-line best-effort forward with zero dedupe state. The case called it twice and asserted the mock recorded two calls, which is arithmetic over the test's own loop; nothing about idempotence was established. Also removed: - Runtime assertions of type-level facts, where production already makes the check at a stronger boundary: `const adapterSatisfiesPort: AdapterIsPort = true` followed by `expect(...).toBe(true)` — unconditionally true, while `createExpoGenerationFileSystem(): GenerationFileSystem` is explicitly annotated and passed into `createGenerationStore` at a typed call site. And a case named "does not typecheck" whose runtime assertion is a length check on its own literal, declaring its own local annotation so it could never notice the production annotation weakening. - Private predicate tests duplicated at a real boundary: four `repo-slug-cache` cases delivered by `repo-slug-index.test.ts`, which drives the same resolution through the hook, the real store and the preload bridge, while the cache-level versions hand-seed the internal map and break on a cache-key format change. - Duplicate invocations owned at the shared boundary, including commit and push recovery cases owned by `src/shared/source-control-recovery-agent-command.test.ts`. Kept deliberately, verified rather than assumed: the production duplication behind the deleted drag test was left alone, because `REPO_HEADER_ACTION_SELECTOR` ends in generic `button, a, input, textarea, select`, so genuine action targets inside a group header still match — it is an unspecialised copy-paste, not a live bug, and collapsing two functions is a refactor. Reported instead. Auditors' probes produced 20, 11 and 13 candidate hits for the signature-versus-title shape across their chunks; every one was inspected and every one was genuine coverage. No deletion in this wave rests on a probe alone. Coverage is partial and stated as such: of 2,269 files, roughly 100 were read case-by-case and the remainder reviewed at title-plus-import level. Each auditor listed its own unread set. The largest remaining surfaces are `src/main/agent-hooks` (95), `src/main/claude` (100), `src/renderer/src/lib/pane-manager` (62) and the 20 largest sidebar suites. Verified: 2,583 desktop test files / 25,636 cases pass, plus one pre-existing `it.fails` marker; the two modified mobile files pass (57 cases); `check-reliability-gates.mjs` 140 gates; `check:code-quality:changed` 0 new findings. Both deleted files confirmed absent from the gate manifest, `cloud/package.json` and `mobile/tests-typecheck-baseline.txt`. |
||
|
|
23a874a2a1 |
fix(secrets): seal credentials on Linux desktops Chromium cannot detect (#24035)
* fix(secrets): stop telling Linux users to install a keyring they already run On a desktop Chromium does not recognise (Hyprland, sway, river, niri) the selected backend is basic_text and sealing is unavailable, so the at-rest protection report told the user to install and unlock gnome-keyring — which is usually already running and serving org.freedesktop.secrets. Chromium simply never looked, because it picks the backend from XDG_CURRENT_DESKTOP. Split the Linux unavailable path on the selected backend: an unrecognised desktop now says so, names XDG_CURRENT_DESKTOP, and points at --password-store. A backend that did resolve but cannot seal keeps the install-and-unlock text, which is correct there. The backend read stays after isEncryptionAvailable(), the call that performs the D-Bus probe, so nothing new blocks (STA-5765 timing unchanged). * fix(secrets): seal credentials on Linux desktops Chromium cannot detect Chromium picks its os_crypt backend from the desktop-environment env vars and recognises none of the tiling compositors (Hyprland, sway, river, niri). On those it selects basic_text, whose key Electron only exposes after an explicit setUsePlainTextEncryption() this app never calls — so isEncryptionAvailable() is false and every credential store takes its plaintext fallback, while gnome-keyring sits on the session bus unasked (#21827). Name gnome-libsecret ourselves, but only where it provably cannot hurt: - Never for a desktop Chromium does resolve. Overriding a working selection is the one change that could strand already-sealed credentials, and a KDE session identified only by KDE_FULL_SESSION is the case that matters. - Only when the secret service's default collection is present AND unlocked, measured out of process with a killable 1.5s deadline. A locked collection with no unlock prompter is what made isEncryptionAvailable() block for 76s to first window (STA-5765); selecting libsecret there would trade silent plaintext for a frozen app. The probe reads the Locked property, which cannot itself trigger a prompt. Anything unexpected — no session bus, no gdbus, no name owner, timeout — leaves Chromium's own choice alone, so the worst case is today's behaviour. The desktop-presence env vars come from a review finding by @Raajik on #21831, confirmed there against base/nix/xdg_util.cc. |
||
|
|
5fa290fa50 |
feat(secrets): warn in Settings when a credential is stored unencrypted (#24048)
* feat(secrets): warn in Settings when a credential is stored unencrypted When no OS keyring is usable, the MiniMax stores write the credential as a plaintext envelope and say so with a console.warn nobody reads. The users this affects are exactly the ones who never see a main-process log, so in practice they were told nothing (#21827). Report it where the credential is managed instead. Each store gains a protection reader, the status IPC carries it, and Settings renders a warning next to the credential it applies to. Keyed on the stored bytes, not isEncryptionAvailable(): a credential saved before a keyring existed stays plaintext until it is saved again, so reporting current capability would call it protected while the file says otherwise. The readers parse the envelope kind without decrypting, so opening Settings cannot provoke a keychain prompt. The console.warn stays. It carries no secret material, and it is still the only signal on a headless host with no Settings window. * feat(secrets): extend the unsealed-credential warning to every affected store The speech key, Linear tokens, Jira tokens and the Bitbucket credential have the same plaintext fallback the MiniMax stores do, and the same console-only warning nobody reads. Add a shared `readCredentialFileProtection` for the four stores that write bare ciphertext with no envelope, classifying with the same printable-UTF-8 test `readStoredCredentialToken` already uses — so the reporter cannot drift into disagreeing with the reader about the same bytes. Linear and Jira report across every stored workspace/site rather than the active one: sealing is a host-wide property, so a second workspace stored while the keyring was missing is exposed even when the active one is sealed. Both fields are optional, so an older remote host that omits them reads as unknown rather than as sealed. Bitbucket reports null for env-supplied auth, where Orca stores nothing and has no claim to make. Also fixes the credential-connection test double, whose identity-function `encryptString` wrote a readable token — faithful enough for a round-trip assertion, but it made the suite assert that a sealed credential was exposed. * chore(i18n): extract the unsealed-credential notice strings CI's localization-extraction gate requires every translate() key to exist in the primary catalog. Inserted in place rather than re-sorting the file, which is not fully sorted and would have produced a 17k-line diff. * test(web): pin the null protection fields on the desktop-only MiniMax bridge The web bridge reports no protection because it stores nothing; the shape assertions had to move with it. |
||
|
|
b99462ac1c |
test: retire mobile, cloud, config and e2e cases their input cannot reach (#24077)
Completes the first pass over every test area in the repository. Sweep over `mobile/src`, `config/scripts`, `cloud/`, and `tests/` (1,494 files in scope, with the 24 files under `mobile/src/test-support/rpc-recording/` deliberately excluded). 31 case declarations removed across 17 files, 2 test files deleted, 356 lines gone. What went, by pattern: - Cross-boundary replays of a shared helper. A whole mobile file re-ran `extractPendingAsk`/`parseAskFromStatus`/`formatAskAnswer`, all owned by `src/shared/native-chat-ask.test.ts`, `native-chat-ask-fifo.test.ts` and the renderer's interactive-prompt suite — one case title was verbatim identical to the owner's, and the owners' inputs are supersets. The mobile file imported the shared module directly and exercised no mobile transport, lifecycle or rendering. - A case whose input cannot reach the behavior its title names: "arms it on Android while the drawer is open", where `use-back-claim.ts` has zero Platform/OS references, so flipping the mocked OS changes only shadow styles. - Identity copiers, including one asserting `prSidebarRenderBranch(state) === state.kind` against a production body that is `return state.kind`. The function stays; it has three live callers. - A test of the runtime rather than the product: a case asserting Node's own `EventEmitter` crash contract on a bare emitter, with zero production code in the path. The guard it documents is exercised behaviourally by the case after it. - Duplicate invocations, one of them provable rather than eyeballed: with `MODULE_SCOPE_ENV_WRITER_PIN = 0`, `files.size <= 0` is strictly implied by the sibling's `expect(offenders).toEqual([])`, since a non-empty `offenders` forces `files.size >= 1`. The pin's own doc says it may only ever be decreased from 0, so it could never become a meaningful bound either. Its policy guidance survives as a comment; the file's real ratchet and its regex self-test both stay. - Expected values produced by the test's own arithmetic, and a p95 case strictly implied by a sibling that already pins exact p95 and exact max over a wider range. One production line goes: the `export` keyword on `assignmentCleanupSteps` in `cloud/apps/relay/src/assignment-cleanup-steps.ts`. The function itself stays and is still called internally; only the test-only export was orphaned. Kept deliberately: everything a gate cites, checked by case title and not only by file path; a gate-cited case that does not deliver its claim (reported instead — see below); a cross-version wire cell whose ledger is never invoked, left under the raised bar for wire coverage; and every limit, bound, quota and provenance guard. Nothing under `mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` was touched — those bytes feed a `recorderSha256` digest pinning 398 golden recordings. Verified: `mobile` vitest over the modified mobile files (8 files, 50 cases); `mobile/scripts/check-tests-typecheck-ratchet.mjs` OK (898 files in program, 125 grandfathered, none @ts-nocheck); relay suite 799 passed; `check-reliability-gates.mjs` 140 gates; both deleted files confirmed absent from the gate manifest, `cloud/package.json` and `mobile/tests-typecheck-baseline.txt`. Seven local failures were investigated and none is caused by this change: five `mobile-web-app-*-render` tests drive `playwright-core` chromium/webkit and need browsers this machine lacks, `release-checkout.unit.test.ts` needs cross-version git refs, and `e2e-worker-env-isolation.unit.test.ts` fails identically with its HEAD content restored — it recurses `tests/e2e` with symlink-following `statSync` and no depth guard. |
||
|
|
f4092c06d6 |
fix(runtime): treat Hermes session start as idle, not a running turn (#24064)
Hermes fires `on_session_start` when a session is opened, switched or reset; its payload carries no turn. Orca mapped it to `working`, so a freshly launched Hermes held a fresh first-party `working` row for the whole 30-minute staleness window: `terminal wait --for tui-idle` never settled against an idle composer, and the sidebar span a phantom spinner. Map it to `done` with `sessionBoundary`, matching what Claude and the compatible-lifecycle providers already do for SessionStart, and let the tui-idle evidence lane settle on a fresh boundary row. A boundary row claims a new session owns the pane and awaits its first input, which cannot arrive mid-turn, so it carries none of the #6011 risk that scoped that lane to DSH; a turn-end `done` still does not settle. Fixes #13653. Co-authored-by: Brian Grablin <5216789+bgrablin@users.noreply.github.com> |
||
|
|
f37a84993e |
test: retire renderer cases whose input cannot reach the behavior they name (#24065)
Audit sweep over `src/renderer` (3,206 test files in scope). 58 case declarations
removed across 37 files, 3 test files deleted, 1,033 lines gone. Nine dead
production symbols removed with them.
The dominant pattern this wave was a case whose input cannot reach the behavior its
title claims:
- `getRemoteBrowserFrameStyle` takes `_metadata` unused and returns a constant
object. Four cases fed it 958x609, a uniform high-DPI 1998x1218, an uneven
high-DPI frame, and malformed metadata, all asserting the same constant. The
function cannot branch on any of them; the runtime rejects server-sized frames, so
the renderer ignores bitmap dimensions by design. Kept the one case whose name
admits that.
- `shouldIgnoreTerminalMenuPointerDownOutside` takes `{openedAtMs, nowMs}` and has no
button or modifier parameter. Two cases named for secondary-button and macOS
control-click passed byte-identical timestamps. Their titles came from the
docblock's prose, which describes behavior the CALLER delivers.
- A retry case titled "then Retry re-arms" proved that half by deleting a key from
its own ref object and asserting it was undefined. The `#6648` budget assertion
stays; the title now matches what the case proves.
Also removed:
- Cases that cannot fail: three asserting `toBeInstanceOf(Map)` and `size === 0`
immediately after a `beforeEach` that clears those maps.
- Assertion-free coverage probes, one of which writes a JSON report to tmpdir and
asserts nothing while its own comment says the sibling case is the gate.
- Copied inventories and export lists, each checked per entry: a 35-name renderer
git-client export list (every name has 7-40 production references) and a 16-path
caller census that never verified the callers pass the argument it exists to track.
- Provider-local replays of one factory (`createUsageProviderSlice`), keeping a
single representative; and a case re-running a shared window-shortcut policy the
shared suite owns.
- Duplicate invocations, including a `buildNativeChatSendBytes` block whose function
is `buildNativeChatPasteBytes(text) + '\r'`, with three titles verbatim from the
paste-bytes cases.
- A slice case whose only assertion reads back `antigravity: null`, a literal in
`createEmptyRateLimitState()`, under a title claiming a "stable pending key".
Nine production symbols went, each verified to have exactly its own declaration plus
one test reference and zero production callers: `buildNativeChatSendBytes` (its
docblock said "kept for callers/tests" and had none), `isCommandMarkerId`,
`buildNativeChatRenderItems`, `collectToolResults`, `scrapeNativeChatSession`, three
now-orphaned types, and an eagerly-evaluated `WORKTREE_CARD_PROPERTY_OPTIONS`.
One deletion was reverted. An i18n JSX-spacing guard reads `.tsx` source text and
would break under a behavior-preserving rewrite, both true. But all three components
it guards have zero referencing test files, so nothing else catches
`translate('…','Detected from')}` followed by `<code>` rendering as "Detected
fromorca.yaml". The junk patterns describe a test's shape; the ratchet rules describe
consequence, and consequence wins.
Kept deliberately: a randomized parity test measured at 483 of 517 comparisons being
`identity === identity`, because 32 seeds and 2 named cases are genuine differentials
and the fix is a rewrite, not a deletion; a title-tracker parity test whose two sides
are separately-authored implementations that a sibling case proves can diverge; a
Monaco upstream-drift detector that reads the installed package, which is the shipped
contract; and nine bound guards, all confirmed to have live production callers.
Coverage is partial and stated as such: roughly 810 of 3,206 files read case-by-case,
the rest triaged by mechanical scan. Every auditor disclosed its own unread set and
those paths are tracked rather than assumed clean.
Verified: `pnpm test` over all nine renderer chunks — 2,834 files, 25,540 cases, 0
failures (a first run showed 5 failures that did not reproduce and were concurrent
load); `pnpm tc` after clearing `.tsbuildinfo`; `check-reliability-gates.mjs` 140
gates; `check:code-quality:changed` 0 new findings. All three deleted files confirmed
absent from the gate manifest, `cloud/package.json` and
`mobile/tests-typecheck-baseline.txt`.
|
||
|
|
d68a5be13a |
fix(claude): run Windows hooks without shell operators (#23944)
* fix(claude): run Windows hooks without shell operators Keep neutral replies inside the managed entry and payload scripts, repair missing files from managed registrations, and stop using Git Bash discovery to guess Claude's hook shell. Co-authored-by: latte271 <junghyeyun27@gmail.com> Co-authored-by: Bing.Z <zzb@gxsmjx.com> * fix(claude): keep the Windows hook refresh async and its scripts after uninstall - List windows-hook-files.ts in the CLI project so typecheck passes. - Refresh the entry/payload pair only from a surviving entry, with an async existence check, so startup refresh stays off the main thread on Windows. - Keep both scripts on uninstall like every other agent; a Claude session still holding old settings keeps answering instead of erroring per event. - A payload that exists but cannot start falls through to the neutral reply, and the missing-payload branch exits early for background jobs. - Update the EDR posture reference for the operator-free command. * test(claude): run the Windows hook host legs for real The live Windows host legs never ran: runProcessSync cannot take a string stdin (it forces encoding 'buffer'), so every leg threw before starting a host. Use async runProcess, pass PATHEXT (without it Windows PowerShell 5.1 prints nothing and exits 0 for a .cmd path), and name the host in each assertion. Drop the POSIX pwsh leg: its drive-mapping shim proved nothing about Windows, and the Windows legs cover both PowerShell hosts. --------- Co-authored-by: latte271 <junghyeyun27@gmail.com> Co-authored-by: Bing.Z <zzb@gxsmjx.com> |
||
|
|
75719bd38b |
fix(terminal): keep the mouse report format in desktop pane snapshots so phone swipes never type escape text (#23943)
* fix(mobile): preserve host mouse modes for terminal scrolling * revert(terminal): drop the out-of-band mouse-modes channel The earlier commit sent the host's mouse tracking and encoding beside each snapshot and stream frame, and made the phone trust that over what its replayed bytes say. Where the snapshot text already carries the encoding this is redundant, and where the text is wrong (a desktop pane snapshot never writes ?1006h) the host's own mirror is seeded from that same text, so the side channel asserts the wrong encoding too. Remove the shared type, the host publication, the wire fields and the phone's host-modes authority. The phone keeps its guard: while mouse tracking is on but no replayed byte proved the encoding, a wheel scrolls locally instead of sending a guessed legacy report. * fix(terminal): carry the mouse encoding in desktop pane snapshots Swiping a phone terminal running Codex typed `[M`-style mouse bytes into Codex's prompt instead of scrolling (#23818). Codex turns on mouse tracking with the SGR encoding (?1006h). A desktop pane snapshot is made by xterm's serialize addon, which writes the tracking modes but never the encoding, so the phone replayed "tracking on, legacy encoding" and sent legacy `ESC[M` reports that Codex does not parse. The host's own terminal model is seeded from the same snapshot, so it lost the encoding too. The pane now mirrors its xterm's mouse encoding from the parser (1006, 1016 and a full reset, with xterm's own set/reset rules) and its snapshot ends with the matching DECSET, after the alternate-screen switch that readers keep. Programs on the default encoding get nothing appended. The phone keeps a guard for hosts without this fix: while a wheel- reporting tracking mode is on but no replayed byte proved the encoding, a wheel or swipe scrolls locally and a tap or drag sends no mouse report (the tap still focuses the keyboard). Click-only tracking (?9h) keeps its arrow-key scrolling on the alternate screen, and an explicit ?1006l still sends legacy reports. * fix(terminal): state the default mouse encoding so legacy mouse apps keep phone input Snapshots now say ?1006l while a program tracks the mouse with the default encoding (daemon rehydrate and desktop pane serializer), and the phone treats a tracking enable seen in live output as proof of its encoding. Legacy-encoding programs such as vim with mouse=a keep phone scrolling, taps and drags; replay-only tracking with no stated encoding still sends nothing. An unproven drag falls back to local selection, and an empty pane snapshot stays empty. * chore(terminal): type the mouse-encoding tracker inputs for the low-evidence audit |
||
|
|
59c05d32b0 |
fix(rate-limits): stop driving a hidden Codex TUI to read usage (#23806)
* fix(rate-limits): stop driving a hidden Codex TUI to read usage When the headless Codex usage call failed, Orca opened a hidden interactive Codex, typed /status and pressed Enter without reading the screen, then killed it after 15 s. If Codex showed its "Update available" prompt, that Enter picked "Update now", Codex started its installer, and the 15 s kill interrupted it, leaving the global install broken. Drop the hidden-terminal fallback. When the headless call fails with a non-sign-in error, read the same usage from the HTTP endpoint Orca already calls on WSL and for the 5-hour window, so a usage check can no longer answer any Codex startup screen. Fixes #17415 * test(rate-limits): pin the RPC error when the Codex HTTP fallback fails Also drop comments that still described the removed hidden-terminal fallback. * test(rate-limits): use real Response objects in Codex fetcher tests The changed-code gate rejects the new type assertions this PR added. |
||
|
|
707d3dc96b |
fix(chat): decode Claude pastes and report terminal delivery uncertainty (#23788)
* fix(chat): decode Claude pastes and track terminal delivery uncertainty Keep queued prompts pending while the existing agent status reports work, and check fresh history after a later idle fact. Preserve draft text and distinguish write rejection from unconfirmed delivery. Co-authored-by: Harshul Rathod <harshulrathod1640@gmail.com> * fix(native-chat): break the observed-send import cycle and keep renderer tests out of main The observed-send path imported the clear helpers from native-chat-runtime-send, which imports it back. Move the input-clear layer into its own module both use. The Claude paste decoder test imported the renderer pending module from src/main, which the node typecheck project cannot see; the echo-retirement assertions now live in a renderer test. * fix(native-chat): still submit a Claude chat send whose write acknowledgment was lost A remote write whose acknowledgment is lost (timeout, dropped link) is not a refusal, but the observed path stopped there and never sent Enter, leaving a body that did land sitting unsubmitted in Claude's input line until the next send's clear wiped it. Continue to the next write without re-sending the bytes, as the unobserved path always did. * perf(native-chat): keep terminal Chat pending delivery from re-rendering every row The delivery notices were merged into a new Map on every render, which invalidated the transcript row context and re-rendered every memoized row on each stream update. The pending hook also wrote a fresh array on every status ping and prune pass even when nothing changed, and the phone mapped its pending list on every render, rebuilding the chat list data. Memoize the merged notices, skip no-op pending writes, and memoize the phone's rendered pending list. The phone also skips a transcript read when no send is due. * fix(native-chat): never flag a queued Claude send, and flag one an idle Claude never starts Two gaps in when terminal Chat calls a Claude send "Delivery unconfirmed": A prompt sent while Claude is mid-turn is queued, and Claude folds it into the running turn as a queued-command record. The transcript reader drops those records, so once the turn ended the prompt Claude did run read as unconfirmed, inviting a duplicate resend. A send made while the agent is busy is now never checked; it keeps the pending behaviour it had before. A prompt sent to an idle Claude that never starts a turn (Claude exited to the shell, or the paste went nowhere) left the status at the same idle fact forever, so the check never ran and the bubble stayed pending. An idle agent starts a turn on a delivered prompt at once, so a send whose idle status is unchanged after the existing 20 s bound is now checked against a fresh transcript read. * fix(native-chat): add the delivery notice strings to the English catalog The Dismiss action's translate key was missing from en.json, which fails the localization catalog and extraction gates. The desktop "Message not sent" and "Delivery unconfirmed" notices were hard-coded English; route them through translate with the same wording. * fix(mobile): sync the held-send refs after commit instead of during render Moving the acknowledgment-loss hold into its own hook made its render-time ref writes new lines, which the React Doctor changed-lines gate blocks. Held sends report after commit, so syncing those refs in a layout effect keeps them current where they are read. * fix(native-chat): report only definite terminal Chat send outcomes The delivery rule inferred "Delivery unconfirmed" from "the turn ended and the transcript has no matching row". Claude records a prompt sent mid-turn only as a queued-command attachment, which the transcript reader drops, so that rule flagged prompts Claude had answered. It also never fired for an idle Claude that lost the write, because no newer turn arrives. Keep only facts the transport reports: - a refused write reads "Message not sent", keeps its text, and can be dismissed; - a lost write acknowledgment holds the echo for 20 s, the phone's existing rule, then reads "Delivery unconfirmed" unless its row has landed. An ordinary send, including one Claude queues mid-turn, stays pending as before. Remove the agent-status subscription, the status-epoch origin, the fresh 500-row transcript read, the confirmed state and the no-status clock. The phone already implements this rule, so its changes revert to main; only a test for old-host paste envelopes remains. * fix(i18n): translate the terminal Chat delivery notices Add the Dismiss, "Message not sent" and "Delivery unconfirmed" strings to the es, fr, ja, ko and zh catalogs, reusing each catalog's existing Dismiss wording. * fix(native-chat): let a resend replace its failed terminal Chat echo A "Message not sent" or "Delivery unconfirmed" echo kept its transcript occurrence, so resending the same text numbered the resend as the second copy: the one landed row retired the failed echo and pinned the resend below the reply forever. Appending a send now drops a failed echo with the same content first. * fix(native-chat): unwrap a Claude paste that quotes pasted_content tags The envelope parser refused any body containing a pasted_content tag, so a pasted prompt that itself quotes one (a transcript excerpt, or code that handles these tags) kept its wrapper and its echo stayed pinned below the reply. Claude's per-paste id exists to disambiguate exactly that; only a same-id tag inside the body is now ambiguous. Wrappers without an id keep the strict rule. * test(native-chat): pin which terminal Chat sends observe write outcomes Only a Claude chat send (text or images) reports a refused or unacknowledged write to its pending echo; other agents and slash commands keep the unobserved write path exactly as before. * fix(native-chat): keep failed terminal Chat sends through Stop Stop cleared every optimistic echo, including a "Message not sent" or "Delivery unconfirmed" bubble whose send had already settled. Stop cannot affect that send, and the bubble is the only place its text stays copyable, so it now survives until the user dismisses or resends it. Also moves the observed-send import below the file header comment. --------- Co-authored-by: Harshul Rathod <harshulrathod1640@gmail.com> |
||
|
|
5b93c6216a |
Fix Chat UI paste intake and pane routing (#23784)
* fix(chat): separate text paste from attachments and route by pane Keep composer text independent of image checks and saving, and route pastes caught underneath chat to the originating pane's mounted input. Preserve native event data, selection replacement, undo, and target lifetime checks. Co-authored-by: Wooseong Kim <innocarpe@gmail.com> Co-authored-by: lurunzi <lurunzi@gmail.com> * fix(chat): keep focus and quiet text paste after routing it to chat - A paste inserted into the composer now moves focus there, as the old menu-paste insert did; otherwise a paste routed from the hidden terminal left the next keystrokes going to that terminal. - With a remote-server or not-ready workspace, pasted text no longer shows the "Local attachments are not available" refusal because the clipboard also held an image rendition (common for Office copies). The menu path probes for an image only when the text read is empty, so a paired browser does one permission-gated clipboard read for a text paste, not two. - Latest-value refs update in a layout effect instead of during render. * perf(clipboard): answer "is there an image?" from the format list The chat composer asks the main process whether the clipboard holds an image before explaining an image-only paste on a remote-server workspace. That probe decoded the whole image (readImage().isEmpty()) on the main thread just to return a boolean. Read clipboard.availableFormats() instead, and share the MIME check with the paired-web probe. * fix(chat): a chat cover owns focus, so input never reaches the hidden terminal When a Chat UI tab opened over its terminal, the terminal's xterm kept keyboard focus until the composer claimed it a frame later, and forever if the composer never became ready (still starting, a question card, a phone holding input). The previous commits rerouted paste from that hidden terminal to the chat, but typing and Enter still went to the terminal, an image-only or refused paste left focus there, and about twenty terminal.focus() call sites could put it back. Make "a covered terminal cannot hold focus" structural instead: - The chat cover takes focus in the commit that mounts it and marks the covered xterm inert, so every terminal.focus() path is refused by the browser. Split siblings are untouched. When the chat goes away the xterm is un-inerted, and gets focus back only if focus was inside that chat. - Terminal paste listeners skip anything inside a chat cover (previously only inside a mounted chat root). The reroute from terminal to chat is gone; terminal-only paste is back to main's code. - A paste that finds no chat input (before the chat mounts, or an approval card with no text field) gets a visible refusal from the cover. A disabled composer shows the same notice inline instead of dropping the paste. New copy: "Can't paste — this chat isn't accepting input right now." (the old "Worktree not ready" toast was wrong for a chat that is still starting). - The terminal context menu, which names its pane, keeps a small request event to that pane's chat, now without a clipboard payload and using the existing covered-pane check. - Cmd/Ctrl+V or Shift+Insert on a non-input part of the chat focuses the composer (or question answer) first, so the paste lands there. - The composer-scope check used to decide whether a text field inside the chat keeps its own paste matched the whole pane (the file-drop surface carries the same attribute). It now asks the composer whether the target is inside its input. * fix(chat): don't paste a copied file's name next to the file Copying a file in Finder or another file manager puts its name on the clipboard as text/plain beside the file itself. Since text and images are now pasted independently, pasting such a copy into a local or SSH chat inserted the file name into the prompt as well as attaching the image. On the paste-event path, text/plain that is exactly the names of the pasted files (one per line) is the file's label, not prompt text, so it is dropped when an image from that paste is being attached. Rich-text copies (text plus an image rendition) still insert their text, and a copied non-image file, which is not attached, still pastes its name as before. * fix(chat): don't type a Finder file's name on Cmd+V either On macOS, Cmd+V in the chat goes through the app-menu paste, which reads the clipboard text and saves the clipboard image separately. A file copied in Finder also puts its name on the clipboard as text, so the composer typed the name next to the attachment. On main the menu path never read text once an image saved. The main process now reports the paths of the files a file manager copied (macOS filenames plist or file URL, Explorer's FileNameW, a Linux uri-list). Text that only labels those files waits for the image outcome: dropped when an image is attached (or refused on a remote owner), typed when none came. The same label rule now also accepts a path or file URL per line, which is how Linux file managers label copied files on the paste-event path. * fix(chat): pane focus aimed at a chat lands on the chat Since the covered terminal became inert, focusing a pane that shows a chat (keyboard pane navigation, focus-follows-mouse, split activation) was refused and focus stayed on the pane the user left, so typing went to that visible sibling terminal. The one place a pane's focus is requested now puts it on the pane's chat cover, which hands it to the composer when the pane is revealed. Focus already inside the chat is left alone. * test(terminal): give fake panes the container pane focus now reads Pane focus checks the pane's container for a chat cover, and these two fixtures built panes with only a terminal, so four tests threw. * refactor(native-chat): move composer paste handle and chat-root key routing into their own modules Brings NativeChatComposer.tsx and NativeChatResolvedView.tsx back under the 400-line limit after merging main. No behavior change. --------- Co-authored-by: Wooseong Kim <innocarpe@gmail.com> Co-authored-by: lurunzi <lurunzi@gmail.com> |
||
|
|
7afa4ee3dc |
test(e2e): read the tab strip's dock samples through a typed window field (#24052)
#24010's spec read them with Reflect.get, which the low-evidence lint rejects, so every PR's static analysis now fails on main. |
||
|
|
ca7c14db08 |
fix(mobile): start + menu, quick command and diff-note agents through agent.launch (#22954)
* fix(mobile): start + menu, quick command and diff-note agents through agent.launch
The session screen's + menu, agent quick commands and diff notes' New agent
session now ask the host to start the agent with agent.launchReplay, so the
host picks chat or terminal from the desktop's default and delivers any
prompt. Hosts without the launch capabilities keep today's paths.
The phone's pending tab choice is one value (a tab, a terminal by handle, or a
launched surface) instead of two refs, and a launched chat is found by its
session id in the next snapshot rather than a predicted tab id. A launched
surface waits a bounded number of snapshots for its tab.
* test(mobile): add the + menu and diff-note launch scenarios to the recording corpus
* test(mobile): repin bridged-parity tallies for the four launch goldens; drop test casts
The corpus grows from 790 to 794 goldens; all four new ones replay identically.
* fix(mobile): show a refused agent launch as a toast beside open tabs
The inline create error renders only in an empty session, so a host refusal
(for example a disabled agent) from the + menu in a session with tabs showed
nothing. Always toast the failure: the caller's own copy when it gave one,
otherwise the host's reason.
* test(mobile): type the launch reply helper with the shared launch outcome types
* fix(mobile): record a launched agent's tab as this device's pick on the host
A launch carries no navigation, so the phone selected the new tab only
locally while the host kept this device on the tab it had before. Leaving
the session and coming back, or a reconnect that reset the screen, reopened
that old tab. The "+" terminal path this replaced asked the host to select
the tab for the caller.
When a launched surface's tab lands in a snapshot, activate it for the
caller exactly as a tap does. The resolver now names the landed tab in
place of the unused `missed` flag. Route parity re-pinned for the new
activation body, identity payload and strings.
* fix(mobile): land on a launched agent's tab without a 500 ms wait or a blank pane
The host publishes a launched tab before it replies, so the tab list the
phone already holds usually has it by the time the reply arrives. The
launch paths still waited for a refetch 500 ms later, leaving the phone on
the old tab for that long after every launch. Read the tab list at once.
On hosts without agent.launch, the chat path also unsubscribed the open
terminal and cleared its handle before the chat's tab landed, while the old
terminal tab stayed selected: a blank pane until the next tab list. Leave the
open tab live until the chat lands, as the launch path does; applying that
tab list tears the old terminal down.
Route parity re-pinned for the two bodies.
* fix(mobile): keep a tab the user picked while a prompted launch was still replying
A quick command or review-notes launch now waits for the host to deliver the prompt, which can take up to a minute. The launched tab shows up in the tab row well before that, so a user who tapped another tab meanwhile was pulled back onto the launched one when the reply arrived, and that pick was recorded on the host.
The launch now remembers which tab the phone was on when it started and only takes focus if the phone is still there when the reply lands. Any move made in between, by a tap or by the computer navigating this phone, wins. Session route parity re-pinned for the handleCreateTerminal body only.
* fix(mobile): name a launched agent's tab before asking, and land on it when it is listed
A "+" menu, quick-command or review-notes launch now reserves its tab before it asks the host: a fresh pane key (tab and leaf UUIDs) and, for an agent the host may start as a chat, a session id. Both are minted once per launch and sent unchanged on every replay, since the host's replay fingerprint covers them.
The phone arms its pending selection with that reservation before sending, so it lands on the terminal (matched by pane halves) or chat (matched by session id) as soon as the tab is listed. For an agent whose prompt is pasted after start, that is long before the reply, which waits for delivery. Landing also frees the "+" lock; the lock holds the create's id, so an older launch's reply cannot free a newer one's. The reply now only adds its own handle or session id (an older host ignores the reservation), starts the fallback countdown, and reports prompt delivery.
A tab the user picks mid-launch replaces the pending selection, so the launch-start tab check is gone. A reservation the host refuses as already taken reads "Couldn't start the agent. Try again." on the first send, and as unconfirmed after a replay. The mobile UUID fallback now yields a v4 UUID, because a pane key's leaf must be one. The host launch path moved to new-tab-agent-host-launch.ts; session route parity re-pinned for that move and the landing's lock release.
* test(mobile): expect the launch reservation in the four launch scenarios
The four launch scenarios now expect the pane key and session id the phone sends (the scripted ids come first, so the operation id moves from ...001 to ...004).
* fix(mobile): don't say an agent may not have started while the user is looking at it
When a launch's reply was lost after its tab had already landed, the phone said "Couldn't confirm the agent started", although the listed tab proves it did. Now a listed tab narrows the doubt to the prompt or notes ("The agent started, but couldn't confirm the notes were sent."), the notes stay unsent, and a bare launch says nothing. Only the nested-function parity pin moves, for handleCreateTerminal passing the tab list to the launch.
* test(mobile): read the launch's sent reservation through the host's params schema
The anti-slop audit rejects Reflect.get; parsing with AgentLaunchReplay also
asserts the host accepts the params the phone sent.
* test(mobile): check the launch reservation against the host without importing its schema
Mobile code may import the params contract only as types. The phone's tests
now read the sent reservation by narrowing, a chat reservation is checked
through the real host dispatcher, and the older-host drop is pinned host-side.
* fix(mobile): don't send the same review notes to a second new agent
The "+" lock is now freed when the launched tab lands, but review notes are
only cleared when the launch's reply confirms delivery, which for a prompted
launch can take up to a minute. In that window "Send review notes to AI" still
offered the same notes, and choosing a new agent session started a second
agent with them.
The notes a new agent session is being started with are now held from the tap
until that launch settles: the Send button no longer counts them, the sheet no
longer offers them, and a stale tap on the old sheet starts nothing. Notes the
host did not deliver become sendable again once the reply arrives.
* test(mobile): record the + menu and diff-note launch goldens
4 added (+ as a terminal, + as a chat, notes delivered, notes not delivered). 10 existing create-terminal goldens move only because the recorded state now shows one pending selection instead of two refs; their requests are unchanged.
|
||
|
|
fceca5cece |
fix(sidebar): an agent's row stays while it runs, whatever its tab title (#23948)
* fix(sidebar): keep hook-less agent rows while the agent runs, whatever its title Codex retitles its pane to the project name, so the sidebar's title-derived row (which required the title to name an agent) vanished while Codex kept running (#23767). Rows now take identity from the canonical pane resolver over the pane's foreground-process read and launch record, then the title; the title only decides idle/working/needs-input. The row still goes away when the PTY exits, the process tracker proves the shell is back, or the title is a shell or default title. * test(dashboard): justify the partial store fixture's type assertion * fix(sidebar): only a live process read keeps a plain-title agent row Review of the previous commit found ghost rows: the tab launch record is a latch nothing clears on WSL, after an SSH exit, or for a launch that never started, and a parked pane's process read went stale because only the mounted tracker re-derives it. - The launch record returns to main's role: a fallback only for titles that show activity, ranked below a title naming another agent (pane reuse), matching the tab icon's order. - A parked pane's command boundary retires its unconfirmable process read, like the mounted ladder's unavailable path; reveal re-reads it. * fix(sidebar): confirm before a parked marker retires an agent; read Git Bash prompt titles as the shell - A parked pane's end-of-command marker can be a nested shell's leak under a still-running full-screen agent, so confirm the foreground first (as the mounted ladder does) and retire the process read only on a shell or no answer. SSH/remote parked panes hold no incarnation to fence a host read with, so they still retire. - Git Bash emits no command marks; its `$MSYSTEM:$PWD` prompt title (MINGW64:/c/repo) is now shell evidence, so a stale Codex read there no longer keeps a ghost row after Codex exits. * fix(sidebar): trust only process-read agents for plain-title rows; per-worktree foreground selector groups by tab A daemon reattach seeds the pane's foreground entry with its launch agent, which can outlive the process while Orca is closed. The entry now records where its agent came from (agentEvidence), and the sidebar/dashboard title-derived rows only keep a plain-title row on an actual process read. Routing and the tab icon are unchanged. selectPaneForegroundAgentsForWorktree grouped every pane key per worktree; it now groups by tab once per map identity and skips worktrees with no tabs. * fix(sidebar): a parked pane's reattach keeps its own process read of the same agent The reattach seed marked a returning parked Codex pane as launch-record evidence, over the process read this session already took, so its row blinked out on reveal and stayed hidden if the user left the tab before the visible read landed. Keep the read when it names the same agent; the seed still drops byte-routing trust. * test(terminal): foreground confirmation publishes process-read evidence * fix(sidebar): a cleared pane title retires the agent's process read Codex clears its title when it exits, and the tab then shows its default title. A pane without shell command marks never re-reads its foreground process, so the retained read kept a "Codex · Idle" row after /quit (permanently for a hand-typed Codex; about 15 s while the marked-pane confirm ladder ran). Treat a blank title like the default title it shows. * fix(sidebar): the pane's process monitor retires an exited agent's process read A hook-less pane keeps its sidebar row from the tracker's foreground-process read, but nothing re-derived that read in a pane without OSC 133 command marks. After Codex exited there, a "Codex · Idle" row stayed: permanently when the shell titles its prompt, or when a killed Codex leaves its last title. The pane's agent-completion process monitor already confirms an agent's exit (no agent and no child processes, held past its settle window). It now reports that exit to the tracker, which retires its own process read and runs the confirmed-shell path the visible-pty read uses. A tracker read that names an agent seeds the monitor, so hidden panes and panes the monitor had not polled yet are watched too. A command read in flight still decides the pane, and launch records or other agents' reads are left alone. * fix(sidebar): a monitor-confirmed exit leaves the next agent in an unmarked pane identifiable The process-exit retire published shellForeground:true and left the one-shot visible sample settled; a pane without command marks has no command start to lift either, so a Codex typed again after quitting was never read and lost its row on retitle. Publish shellForeground:false and reopen the sample. * test(terminal): justify the pane binding cast in the process-exit relaunch test |
||
|
|
f33f3093cb |
docs(wechat): point community QR code at group 11 (#24050)
Group 10 is full; swap the README QR code and copy (all locales) to the new group 11 invite. |
||
|
|
444e0b1cf9 |
fix(codex): recognise Codex's quoted spellings in config.toml, and repair Orca's duplicates (#22592) (#23958)
* fix(codex): recognise Codex's quoted project-trust spellings in config.toml (#22592) Codex's settings screen writes project trust as ["projects"."/p"] and "trust_level" = "trusted". Orca's matchers only knew the bare spelling, so a trust write appended a second [projects."/p"] table (or a second trust_level line) and every codex command then failed with "duplicate key". The config mirror kept both spellings in Orca-managed homes for the same reason. - Project table headers are now read through the existing TOML key-path parser, so bare, quoted, literal-quoted, mixed and spaced spellings are the same table for trust writes and the managed-home mirror/dedupe. - trust_level is found by decoded key, in both the trust writer and the mirror's trust reader, and an existing key is rewritten, never duplicated. - On the next trust write, a table older Orca appended (exactly [projects."<p>"] holding only trust_level = "trusted") that duplicates the user's table, or the bare line it inserted under a quoted "trust_level", is removed; the user's table wins and the atomic writer keeps config.toml.bak. Any other duplicate, or a repair that would still leave one, leaves the file untouched and logs once. * build(cli): list the new Codex trust modules in the CLI project * fix(codex): recognise Codex's quoted hooks.state spellings and repair Orca's copies (#22592) Codex writes hook trust as ["hooks"."state"."<key>"] (and the parent as ["hooks"."state"]). Orca's hook-trust writer, parent-table check and mirror only knew the bare spelling, so a hook-trust write appended a bare copy and the file failed to parse with "Cannot declare ... twice". - The hooks.state header, parent-table and mirror checks now use the TOML key-path parser, like project tables. - The duplicate repair now also removes Orca's own hooks.state tables (an exact [hooks.state."<k>"] with only enabled + trusted_hash, or an empty [hooks.state]) that repeat a table in another spelling, and runs on hook trust writes too, so a file with both project and hook duplicates is fully repaired. The Orca-shaped copy is removed whichever order the two tables are in, only when exactly one other table (the user's) remains; anything else is left untouched and logged once. * fix(codex): carry plain-Codex plugin and project hook trust into Orca's Codex homes (#22592) Codex keeps hook trust in $CODEX_HOME/config.toml under hooks.state, keyed by the hook's source. Plugin keys (`id@mkt:path`) and project keys (`<repo>/.codex/...`) are the same in every home, but the mirror dropped every hooks.state table from ~/.codex, so Codex inside Orca asked users to re-trust plugin and project hooks they had already trusted in plain Codex. - classifyHookTrustKey splits keys into home-scoped (the home's own hooks.json/config.toml, re-keyed by install as before) and shared. - The mirror now carries shared hook trust from ~/.codex in every spelling. A key the managed home already holds keeps the managed copy, a key repeated in ~/.codex is carried once, and the parent [hooks.state] table is never copied, so the result never declares a table twice. - mergeSystemCodexConfigIntoRuntime moves to codex-config-mirror-merge.ts to keep codex-config-mirror.ts under the line limit. - Tests cover plugin/project carry in each spelling, user-hook keys staying out, repeated launches, managed-copy precedence, Windows key spellings, parent tables, and user-hook trust re-keying (trusted_hash and enabled) from every ~/.codex spelling. * fix(codex): carry session_end and interrupt hook trust into Orca's Codex homes (#22592) The shared-trust classifier parsed hook keys with Orca's own trust-key parser, which only knows the ten events Orca installs hooks for. Keys for Codex's session_end and interrupt events did not parse, so their plugin and project trust was treated as home-scoped and left out of the managed home. The classifier now reads the source path from Codex's key shape `{source}:{event}:{group}:{handler}` for any event label. A key without that shape is still never carried. Tests cover both events for plugin and project keys in both spellings, user-layer keys for both events, and five unattributable key shapes. |
||
|
|
c7218f20dc |
test: retire src/main/runtime cases that re-prove an owned contract (#24043)
Completes the `src/main` audit. Sweep over `src/main/runtime` (flat, rpc,
orchestration, relay, push) and flat `src/main`. 42 case declarations removed
across 24 files, 783 lines gone. No file deleted whole, no production code touched.
Highest-yield area by far was `runtime/orchestration` (21 cases from 151 files);
the rest of runtime measured under 1%.
What went, by pattern:
- Tests whose subject is the test itself: a source-scanning boundary test with no
production import at all, three of whose cases checked its own regex against
strings it declares; and a benchmark whose own simulation contains the
short-circuit it asserts, with `expect(wouldHaveBeen).toBe(60000)` comparing the
test's own arithmetic.
- Identity copiers: seven db cases where every asserted value is the literal input
(`type: 'question'` in, `type === 'question'` out).
- Permanently skipped tests for behavior that does not exist — two cases carrying
`// TODO: inline restore on re-subscribe not yet implemented`. A skipped test for
an unimplemented feature can never fail; it is a note in test syntax.
- Duplicate invocations of a contract owned at a stronger boundary, including three
reset scopes and two dependency-promotion cases owned by dedicated suites.
- Registration manifests whose every entry is referenced by other tests.
- Table rows and cases varying a field production never reads: a mobile tab-restore
case varying `clientCapabilities`, which the mobile branch does not consult, and
a Windows worker case in a module with no platform input at all.
- Names promising more than the input exercises: a case titled for forged AppImage
variables whose body only removes `AppRun`, byte-identical to a row of the
`it.each` table twenty lines above.
- Negative controls passing for an unrelated reason: a foreign-pane rejection whose
fixture also differs in sender handle, so the handle guard can reject it.
- Self-comparisons, including `format(m, { authority: 'current' }) === format(m)`.
Two deletions were justified by the wrong argument and kept only after checking a
better one. "A generated-catalog check gates registration" is false: that script
prevents the catalog and dispatcher from drifting apart, so a developer who removes
a method regenerates the catalog and the check passes. Like a type annotation over
an interface and its implementing class, it verifies internal consistency and
cannot see a declaration and its use removed together. Both inventories go on
per-entry evidence instead — every name is referenced by other tests.
The same rule kept a 37-entry terminal-method inventory in the same wave, because
some of its entries are pinned solely by it. An inventory is a ratchet if and only
if at least one entry is pinned solely by it; that is evidence per entry, not a
verdict by shape.
Kept deliberately: everything a gate cites, checked by case title and not only by
file path; source-scanning ratchets that pair their negative assertion with a
positive one against real source (the surviving boundary case asserts the pattern
still matches the writer module, so a silently-broken regex goes red); a
destructive-delete PTY waiver guard; `it.fails` markers, which go red if the bug is
fixed; and `it.skipIf(platform)` cases, which do run on other hosts.
Coverage is partial and stated as such: 1,195 files in scope, roughly 660 read
case-by-case, the remainder title-scanned and mechanically triaged. Unread paths are
recorded for a later sweep rather than assumed clean.
Verified: `check-reliability-gates.mjs` (140 gates), `check:code-quality:changed`
(0 new findings). No `pnpm tc` needed since no production file changed. A local run
of `src/main` surfaced 42 failures in three files this change does not touch
(`browser-manager-tab-identity`, `browser-manager-viewport-ownership`,
`session-scanner-codex-workers`); all 42 reproduce on a pristine `origin/main`
worktree, so they predate this wave and are environment-dependent locally — CI was
green on main at `2d85fdc753e2`.
|
||
|
|
5b3366f78e |
test: unit tests can no longer write a developer's real agent or Orca settings (#23979)
* fix(agent-trust): write per-user trust under the home the launched agent reads The Cursor, Copilot, Qoder and Antigravity writers and the local Codex config list resolved ~ with os.homedir() at write time, so any test that reached them wrote into the developer's real ~/.codex, ~/.cursor, ~/.copilot or ~/.gemini. Each writer now takes the home, derived once from the launch env (HOME, or USERPROFILE on Windows, else this host's home) by launchedAgentHome, which the SSH relay already used. * test: give tests that wrote the real agent or Orca home a temp one The structured Codex adoption replay pre-trusted /repos/workspace-1 in the real ~/.codex/config.toml; it now runs with a temp HOME and userData. The Codex session-resume and WSL hook tests created Orca's managed Codex home under the live userData, and the Claude Agent Teams tests wrote their tmux shim into ~/.orca; each now runs against a temp userData or HOME. * test: fail any unit test that writes the real agent or Orca home A vitest setup file wraps the node:fs mutating calls and refuses a target under the account's real ~/.codex, ~/.claude(.json), ~/.orca, ~/.cursor, ~/.copilot, ~/.gemini, ~/.qoder or Orca userData, found through os.userInfo() so a test that swaps HOME cannot hide it. The refusal is recorded and rethrown after the test, since trust writers swallow errors. Reads are untouched. It stands down only while an opted-in real-agent suite's own switch is set. It also unsets what an Orca terminal exports toward the live app (userData, Codex and Claude homes, and the Codex launch preflight CLI, which a shell test would otherwise run), so a local run matches CI. * test: type the guarded fs call from its narrowed original |
||
|
|
e8e09eed46 |
fix(push): size the claim-attempt budget from the drain count (#24040)
* fix(push): size the claim-attempt budget from the drain count With twelve drains, up to eleven peers can hold device heads, so a four-attempt claim budget can run out while claimable rows remain and the drain exits idle for a tick. Move the drain count into one module and derive the attempt budget from it. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * style(push): keep the worker and store in repo formatting Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * style(push): drop the stray semicolon in the concurrency constant Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
9cdbeba06d |
fix(tab-bar): keep the active tab visible when the tab strip scrolls (#24010)
* fix: keep active tab visible by docking to viewport edges Makes the current tab easier to locate in many-tab scenarios. The active tab now sticks to a viewport edge via sticky positioning when it would scroll out of view, with a full-foreground indicator bar for better visibility and arrow animation when a background tab opens off-screen. * fix(tab-bar): reveal offscreen tabs instead of nudge animation When a background tab opens beyond the visible area, automatically scroll to reveal it (unless hovering the tab strip). This replaces the previous arrow-nudge animation with direct visibility. revealTabStripElement now handles keeping the active tab visible alongside the revealed tab when both fit, or docks the active tab when needed. * fix(tab-bar): reveal tabs by identity, not count increase alone Detect opened tabs by comparing tab identities independently of count changes. Newly opened tabs are now revealed even when the total tab count stays the same—e.g., when a tab closes as another opens. * fix(tab-bar): track tabs by identity for reliable reveal on open/close Replace count-based tab detection with identity tracking so the strip correctly reveals tabs when they're added, replaced, or when the active tab closes and switches to a far-back history tab. Removes the tabCount parameter and simplifies overflow navigation by using identity sets. * fix(tab-bar): defer revealing tabs until pointer leaves When a background tab opens while the pointer hovers the tab strip, defer its reveal until the pointer leaves. This prevents the active tab from sliding away mid-interaction. Also support client-hosted rows taking active state while maintaining tab dock positioning. |
||
|
|
8b01590847 |
fix(editor): keep saved untitled notes when their tab closes (#23768)
* fix(editor): keep saved untitled notes when their tab closes Saving cleared the draft and dirty flag, so close-time cleanup took a saved note for an untouched placeholder and deleted it. Fixes #23688 * fix(editor): decide untitled-note cleanup from disk state, not a per-save flag Drop setUntitledFileHasSavedContent and its save-queue hook so deleteUntouchedOnClose keeps its creation-time template meaning. A clean untitled tab now counts as an empty placeholder only while its last-known disk content is empty or not yet loaded; the on-disk size check still gates every actual delete. Notes an agent filled while their tab was showing now stay reopenable with Cmd+Shift+T. Also test both Don't Save dialogs end to end, mirror the shared discard in the floating panel test mock, and remove markFileDirty selectors left without consumers. * test(editor): wait for untitled cleanup before asserting disk state Drain the stat-then-delete chain before the untitled-note tests assert the file survived, so the in-flight Don't Save case fails if the size check is removed. Type the main-window close dialog hook by the controller fields it reads, which lets its test drop a type assertion. --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
f4068747ac |
fix(claude): a restarted provider continues the subagent roster earlier runs journaled (#23758)
* fix(claude): a subagent resumed after a restart keeps its canonical id
A backgrounded Claude subagent resumed by a message after its session's provider
restarted parents its frames to its ORIGINAL spawn call, while the announcement the
new provider sees names only the resuming call. The spawn call's alias lived only
in the old provider's memory, so every row the resumed child wrote fell through to
its raw call id: one child shown as two, a named roster entry with no rows and an
unnamed section holding them.
The alias table now also recalls what an earlier run of the session resolved,
re-derived from the agent rows it journaled (canonical id beside the call its frames
arrived under), read once per bound journal epoch. Nothing new is persisted.
* fix(claude): a restarted provider continues the subagent roster earlier runs journaled
The roster's state lived in one provider process while everything it writes is
the session's. After a restart a resumed child was re-rostered in a second group
row with its attempts restarted, its resumed frames lost the alias they still
carry, and the one group row no turn owns was rewritten from empty, erasing the
children an earlier run had listed there.
A run now reads what earlier runs journaled, once per bound journal epoch: group
rows give each child's entry and group, agent rows give the spawn-call aliases
and the latest attempt. A group is inherited only when this run's own events
reach it, and inheriting writes nothing. An inherited child is reopened by an
announcement exactly as an in-process resume reopens it, and by nothing else.
The alias recall is one facet of that read. Nothing new is persisted.
* fix(claude): an earlier run's subagent takes Claude's restart verdict in one row
At restart Claude reports each agent the previous session left running as
stopped ("didn't finish before the previous session ended"). The roster ignored
it, leaving the entry unverifiable, while the background-task lane, which never
saw that agent announced, opened a second, unnamed row for the same agent.
The roster now records the verdict on the inherited entry, keeping the time the
earlier run lost contact, and still reopens the entry on an announcement. The
background-task lane leaves any agent an earlier run rostered to the roster,
from the same journal-derived reading.
* fix(claude): a restart verdict on a child whose host died invents no stop time
A journal reopened after its host died leaves a child unverifiable with no stop
time. Stamping Claude's later restart verdict with the current time would show
the whole outage as how long the child ran, so the earlier run's stamp, or its
absence, is kept. The journal-liveness comment no longer describes the roster as
unable to continue from the journal.
* fix(claude): a child two journaled rows list stays live in one of them
An older build could re-roster a resumed child in a later turn's row, so two
group rows list it (67 children in 39 local sessions). Inheriting the second
row re-pointed the child to that row's copy, so the copy the resume had
reopened was left at working and swept to unverifiable, while the outcome
landed on the stale copy. The first row reached keeps the child.
* fix(native-chat): no run length for a settled child with no stop time
A subagent whose host died without sweeping it has no stop time. After a
restart, Claude's own verdict on it ("stopped") is now recorded, but the rule
that hides the group's duration only covered `unverifiable`, so a mixed group
showed the sibling's duration as the whole group's. The rule now keys on the
missing stop time, whatever the settled state.
* fix(claude): a twice-listed child resumes in the row the journal reading chose
A child an older build listed in two rows was placed in whichever row this
run's frames reached first, so a sibling's frame reaching the older row made
the resumed child reopen there. Placement now follows the reading's one
tie-break, and evicting a row drops only placements that point at it.
* fix(native-chat): a resumed subagent's clock never counts the idle gap
A reopened Claude child now starts a new run: its startedAt is reset to the
reopening frame's time on every reopen path (resume after a restart and a
same-process reactivation). The group clock reads the union of the children's
latest runs, so an idle gap between runs is never counted; an ordinary
overlapping fan-out reads the same as before.
|
||
|
|
21124db4d5 |
refactor(native-chat): a subagent's rows live in its own section, not in the parent's conversation (#23752)
* refactor(native-chat): a subagent's rows live with that subagent, not in the conversation
A subagent's rows were drawn in its parent's conversation, each captioned with
the subagent's name. They now belong to the subagent: the transcript projection
keeps the session's own rows as the conversation and each subagent's rows apart,
keyed by the agent id its roster entry already carries, folded on their own.
Desktop: a subagent's rows open in a section under the roster row that names it,
from that agent's roster entry, and are windowed like any other rows. A subagent
no loaded roster names opens where its first row happened, inside the section of
the agent that spawned it or in the conversation. Its edits still count in the
turn they were made, and revealing one opens the sections around it.
Mobile shows the conversation, with each spawn's roster line. Worker reads and
structured terminal reads serve the worker's own rows.
Removes what the move makes redundant: the per-row caption and its copy, the
producer check in the tool fold and the turn answer, the per-agent frontier
interleaved in the conversation, worker-text subagent tags, and the agent id on
worker-read messages.
* refactor(native-chat): a diff target names the sections its row sits in
Revealing a subagent's edit opens the sections around it from the target the
rollup already holds, instead of looking the row up at click time. The section
head keeps to the agent's name and dot; its state in words stays on the roster
entry. The worker page test stubs the host through its module rather than a cast.
* fix(native-chat): a working subagent's section is open; a worker page windows its own rows
A subagent's section is open while its agent works and closes once it settles,
the way the turn's own live run does; a section the reader opened or closed by
hand keeps that choice. A subagent another subagent spawned opens inside that
one's section, so a working grandchild shows inside its working parent. Openness
is derived from the roster's state and the reader's choices; nothing stores an
automatic open.
A worker page is now the newest page of the worker's own rows. The host windows
the read over them before the limit, so a subagent's burst can no longer crowd
the worker's rows off the page, and "older" still means older worker rows. The
scope is an in-process argument of the host's history read; no wire request
carries it.
* fix(native-chat): a subagent section head names the turn it sits in, for the outline rail
* fix(native-chat): a subagent section's rows sit in the turn the section is shown in, for the outline rail
A background subagent's rows written during a later turn carried that later
turn onto their slots, so scrolling through its section lit the later turn's
rail tick and then snapped back. The rollup still counts each edit in the turn
it was made; only the slot, which the rail reads, takes the shown turn.
* fix(mobile): Load earlier reads past pages that hold only a subagent's rows
Mobile draws only the session's own rows, so an older page made entirely of a
subagent's rows landed as nothing: the reader tapped Load earlier, saw the
spinner, and got the same transcript back. One load now reads on (up to 8 pages)
until a page holds a row of the session's own, then applies the pages in order.
* test(mobile): stub the RPC client the way the other structured-session hook tests do
* perf(native-chat): order subagent rows for the changed-files rollup once per change to them
The rollup flattened and re-sorted every subagent row on each update, including
every token the parent streamed. The ordering now keys on the projection's
subagent rows, which keep their identity while only the conversation changes.
* refactor(native-chat): order subagent rows in the sections hook, keeping the list under its line limit
* fix(mobile): a transcript whose newest page is only a subagent's rows reads back on its own
Opened while a subagent is busy, the newest page can hold nothing but that
subagent's rows. Mobile draws none of them, so the reader saw an empty chat with
a Load earlier button, and an empty list cannot be scrolled to page. The hook now
reads back once from each such head, and the read runs on to the session's own rows.
* fix(native-chat): count the live window in the session's own rows, so a subagent's burst keeps its roster
The live window kept the newest 1,024 rows of every agent. A subagent writing
more than that trimmed its own spawn's roster row and the prompt, and its
section fell back to a closed, unnamed header. The window now keeps the newest
1,024 of the session's own rows and everything after, with an 8,192-row cap on
every agent's rows as the memory backstop. A transcript with no subagent rows
trims exactly as before.
* fix(agent-session): window history pages by the session's own rows, with a subagent's rows riding along
A history page held the newest 200 rows of every agent, so a subagent's burst
could fill a page on its own: the phone opened on an empty chat and "Load
earlier" landed nothing. A page now starts at the oldest of the newest `limit`
rows of the session's own and serves every row from there, so the subagent's
rows come with the conversation they happened in. The page stays contiguous,
the cursor still names its first row, and the byte bound still applies. A
transcript with no subagent rows gets the same pages as before.
Clients already take a page larger than its limit: both reducers raise their
retained window to the page's size. The mobile read-on and read-back stay for
older hosts.
* test(agent-session): a page reaches back to the start rather than leaving a subagent-only page
* fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it
The live window trimmed to just after the own row it dropped, so a subagent
whose roster row went kept its rows at the top as an unnamed section until
the parent wrote again. Trim to the oldest own row kept instead; it still
fires only once an own row passes the limit, so a paged-in run of subagent
rows at the head stays until then. With no subagent rows nothing changes.
* perf(native-chat): cap the live window at 4,096 rows, bounding each delta's re-derivation
Every live batch re-derives the transcript over every retained row. On the
largest real window (7,374 rows) that cost 7-8 ms a delta on desktop against
0.6 ms at the old 1,024-row window, and held about 26 MB of row content.
4,096 halves both. The most rows any local journal puts between a roster and
its subagent's last row, with the parent inside its own-row limit, is 3,005,
so no observed subagent loses its roster to the lower cap.
* fix(native-chat): a subagent section opens only while its roster is the running scope's live frontier
A section used to open whenever its roster said the subagent was working, anywhere
in the transcript and whether or not the session was running, so a background
subagent's section stayed open and grew mid-transcript while the parent moved on.
It now opens by default only while the session runs and the roster row naming the
subagent is the newest thing the parent produced, user rows aside. Newer parent
output closes it even while the subagent still works; the roster row keeps
showing that live state. A subagent still working is a running scope of its own
for the sections it spawned; a settled one closes its scope. Derived every
render, no latch; the reader's own open or close still wins.
* fix(native-chat): name a subagent's section from a client roster the window never trims
A section took its name and state from a roster row in the loaded window. Once a
burst trimmed that row, or the row sat on an older page, the section fell back to
an unnamed, closed "Subagent" header.
The shared reducer now keeps a roster keyed by agent id, folded from every roster
row and revision the client receives: pages, older pages and live batches,
including revisions of roster rows outside the window, which live batches already
carry. The first roster naming an agent wins and its revisions update it; a
removed roster row drops its entries; it is rebuilt on every page that replaces
the window and bounded to 512 agents. Sections take their name, state and
live-frontier place from it; placement stays under the loaded roster row, else
at the section's first loaded row. Only a subagent no roster ever named stays
unnamed.
* feat(agent-session): a history page names the subagents whose roster row is older than it
A page is a contiguous run of the journal whose older-page cursor is its first
item, so it cannot pull an older roster row in without skipping the rows between.
When a page held a subagent's rows but not the roster row naming it (about 11% of
the moments a reader could open a session on local journals), that subagent drew
as an unnamed "Subagent" header.
History and hydration pages now carry an optional `subagentRoster`: the first
roster entry naming each subagent whose rows are on the page and whose roster row
is not, with the row's id, sequence and revision; bounded to 64 entries and
16 KB. Items and cursor are unchanged. The client seeds its roster from it.
Rule 1 in docs/reference/remote-wire-compatibility.md: an optional field on an
existing frame, no capability gate. An older client ignores it (the released
reducer reads a page with it exactly as one without); against an older host the
field is absent and the section falls back to an unnamed header.
* Revert "fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it"
This reverts commit
|
||
|
|
71ac7cb66b |
Fix typo in PR template about issue linking
Correct typo in PR template regarding issue linking for outside contributors. |
||
|
|
52110982ca |
fix(push): run twelve delivery drains instead of four (#24038)
Four drains, each holding one provider round trip of ~100 ms plus its database statements, capped delivery near 30/s. Production inflow reached 35/s on 2026-09-30, so the backlog aged past the five-minute TTL and notifications expired. Twelve drains lift the ceiling to roughly 90/s; the pool is now six per instance, so the extra drains queue on connections instead of starving the request path. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
75040eba5a |
test: open, seed and read the agent-session record store through one test harness (#23986)
* test: open, seed and read the agent-session record store through one harness Tests that open the durable agent-session record store, seed it, or read back what it persisted now go through agent-session-record-store-test-harness.ts instead of calling AgentSessionRecordStore.open or touching agent-sessions.json themselves. A later change that moves the store into the chat database then changes the harness instead of every test. No production code changes. Tests whose subject is the JSON file itself (its .bak recovery, salvage, schema versions, permissions, and what older builds read back) keep reading and writing the file directly; the storage move rewrites or deletes them. * test: address the record-store harness by the host's state directory The harness took the store's own folder, so each caller picked one (join(root, 'store'), or 'agent-sessions' where a test read the store the runtime owns). A later change that moves the store into the state directory's journal database could not tell those apart, and would have had to edit every caller again. Every harness function now takes the state directory, the one the test's journal database and recovery capsule already live in, and keeps the store in the same subfolder the runtime uses. Callers pass that directory; store-only tests pass their temp directory unchanged. Format tests that share a directory with harness calls take the file path from testAgentSessionStoreFilePath. The folder name moves from a private constant in the runtime to AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness shares it without importing the runtime. Its value and every path built from it are unchanged. |
||
|
|
c3183a4556 |
test(e2e): let the completed-worker fake Codex answer the --help probe (#24033)
#23900 probes codex --help before each launch; the fake counted it as a worker spawn, breaking two specs. |
||
|
|
d414033400 |
fix(packaging): stop shipping the relay bundles inside app.asar (#24027)
resources/relay is the only relay copy a packaged build resolves, but out/relay was also packed into app.asar — 14.2MB of unreachable duplicate. Kaspersky flagged app.asar as a compound object precisely because relay.js was inside it, so one script-heuristic verdict on relay.js gutted the whole install. Excluding it decouples app.asar from that verdict and drops the duplicate bytes. |
||
|
|
e594cb06af |
test(mobile): record RPC goldens without a pinned commit, and check recorded requests against the desktop's params rules (#23732)
* test(mobile): add rpc:diff to decode what a recording change moved
The RPC recording goldens are content-addressed JSON, so their raw git diff is
pool hashes. `pnpm --dir mobile rpc:diff [<base>]` decodes both sides and prints,
per golden, the checkpoint, field and JSON path that moved with both values,
grouped across checkpoints, plus added and removed goldens. `--summary <file>`
appends a Markdown report capped for GitHub's step-summary limit.
It reads any pooled format, so it can prove the next commit's format change
moves no recorded value. Checkpoints are matched by occurrence because an id can
repeat within one golden.
This commit adds files under the recorder directory, which moves the header
digest every golden pins; the next commit removes that header.
* test(mobile): record RPC goldens without a pinned commit or input digests
Every golden carried a pinned `baseline` commit plus digests of the recorder,
its mount adapter and its scenario, and the record script refused to run unless
the product tree matched the pin. So every behaviour change repinned to its own
branch commit and rewrote all ~790 files, the squash made that commit
unreachable, and main's pin job stayed red until a hand-made repin pull request
landed (22 of them in 12 days). The digests could only fail when an input moved
and the recording did not, which is exactly the change that carries no
information; every run already re-derives each golden from the current tree and
compares it.
Format 6 keeps the format version, operation, family, named deltas, the value
pool and the recording. Removed: the pin and fence, the three digest modules and
their test, the pin guard and its CI job, and the dead scenario `version` field
(the manifest reader now refuses `baseline` and `version` with a message).
- `pnpm --dir mobile rpc:record [<golden-id>...] [--prune]` records all or some
goldens; orphans are listed, and deleted only with `--prune`. Every derived
test title now starts with its golden id so an id selects it.
- `compareGolden` reports every difference in one failure (identity fields by
name, the checkpoint list, each checkpoint/field/path grouped), keeps the
final byte compare, and ends with the command to re-record that golden.
- `unhandled-recording.test.ts` now drives a detached rejection through
`runRecording` into a checkpoint and the cleanup checkpoint; no golden carries
one, and disconnecting the capture passed every suite before.
- Seam rules that existed only to keep a digest honest are gone; the
mutant-reachability, register-completeness and one-exposure rules stay.
- CI: `Mobile tests on main` runs the whole mobile suite on every merge that
touches mobile/, src/shared/, the root lockfile or the host RPC paths, since
`verify` never runs on main. A new `Mobile RPC Recording Replay` workflow
replays the recordings on pull requests that touch src/shared/ or the root
lockfile without touching mobile/. `verify` writes the `rpc:diff` report to
the job summary.
Proof: `rpc:diff` against the parent reports no recorded behaviour moved; each
golden only loses its ten header lines.
* test(mobile): check every recorded request against the host's params contract
The goldens script the host's replies, so a scenario could record a success
for a request the real host would refuse, and a desktop change that tightens a
params schema moved no golden at all.
`recorded-request-params.test.ts` parses every distinct request the corpus puts
on the wire with the host dispatcher's own `parseRpcRequestParams` and the
schema `rpc-params-catalog.generated.ts` binds to that method. It fails on a
method the host lacks, params it refuses, params sent to a method that takes
none (the dispatcher never reads them), and keys the schema silently strips
unless an inventory entry gives the reason; a stale entry fails too. Each rule
is also shown firing on a made-up request, since the corpus has no instance of
three of them. It imports the desktop dispatcher, so it sits beside the other
Node-side tests outside the RN test program, and the params-contract boundary
now exempts test files, which are never bundled.
It found twelve requests the host would refuse, all from invented fixture
values, not product code, fixed at their source:
- git.branchDiff sent `base-oid`/`head-oid`/`merge-base` where the host needs
full object ids (diff-review and source-control adapters, and the branch
compare replies in the manifest that feed them);
- an iOS push registration without `apnsEnvironment`, which a real iOS token
always carries (`push-token.ts`); the adapter now defaults to `production`;
- `settings.update` given Linear's `assigned` filter as a GitHub preset, which
the product type forbids; the scenario now picks `my-issues`;
- GitLab `projectRef` as a string where the host and the product type take
`{ host, path }` (7 methods, 5 adapters and the manifest).
46 goldens move, and a decoded comparison of every one of them shows no change
other than those substitutions; `rpc:diff` lists them.
* ci(mobile): detect a mobile change without a SIGPIPE-prone grep pipe
Under the runner's pipefail, grep -q exiting on its first match SIGPIPEs git
diff on a long file list, so a large pull request touching mobile/ read as
uncovered and replayed the recordings a second time.
* test(mobile): drop comments that still describe the golden header and digests
Eleven adapters justified an import rule by the header a golden no longer
carries, and that rule's test is gone. The census failure now names the
rpc:record and --prune commands.
* test(mobile): refuse a golden that keeps a key no recording writes
Decoding dropped unknown top-level keys, so an old header left behind by a
hand-resolved merge conflict passed every compare unseen.
* ci(mobile): summarize RPC recording changes after a failed test step too
* test(mobile): stream rpc:record output instead of capturing it
A captured run stayed silent for its whole duration and clipped its tail,
where the failure summary sits, past 8 MB.
* test(ci): let the Ruby-gate contract skip the always-run RPC summary step
|
||
|
|
f7025d88be |
fix(push): give the push gateway six database connections per instance (#24026)
Delivery collapsed on 2026-09-29 once send volume doubled: the worker, the retention pruner and the request path share a two-connection pool, and the database transaction rate pinned at ~140/s regardless of how many notifications were delivered. Raise the pool to six so worker and pruner stop serialising on one connection. The budget precondition stays satisfied (2 x 6 x 3 = 36 <= 64). Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
2d85fdc753 |
test: retire src/main cases that replay a contract their owner already proves (#24025)
Audit sweep over `src/main/{native-chat,startup,daemon,skills,ssh,providers,git,
persistence,claude,agent-hooks,github}`. 35 case declarations removed across 23
files, 1 test file deleted, 655 lines gone. No production file touched.
What went, by pattern:
- Duplicate invocations of a contract owned exhaustively elsewhere: three
`publishDaemonEndpoint` cases that `daemon-endpoint-publish.test.ts` already
covers in 20, and three daemon health classifications (`HEALTHY`, `DEGRADED`,
`UNREACHABLE`) that `daemon-health.test.ts` owns. `WEDGED` and `WEDGED-HELLO`
stayed — the never-resolving-RPC and never-answers-hello paths have no other
owner.
- Provider-local replays of a shared helper: five `GitStatusReadLeaseOwner` cases
re-run per provider, owned by `src/main/git/git-status-read-lease-owner.test.ts`,
and `returns the connectionId` replayed in three provider suites against an
identity getter.
- Assertion-free coverage probes, including one whose comment says "no writes
should happen" while nothing checks that.
- Copied inventories that restate a type: `PROVIDER_FRAME_CLASSIFICATIONS` is
declared `as const satisfies Record<...>`, so a missing key is already a type
error and an extra key fails the excess-property check. Those cases also pinned
key order, which is not a contract.
- A negative control that cannot fail: asserting a profile-state filename is not
an unrelated literal, in a file whose first case already pins that filename
positively.
- Byte-identical duplicates across files, and a second case re-asserting the
`unverifiable -> true` mapping the case above it already proves.
`src/main/providers/ssh-git-provider-api.test.ts` goes: 52 method names asserted
`toBeTypeOf('function')` plus `toHaveLength(52)` over its own literal. Note the
reason, because the obvious one is wrong. "The `IGitProvider & SshGitProvider`
annotation enforces this at compile time" does NOT hold — removing an operation
from the interface and its implementing class in one commit still compiles. What
makes the file redundant is that all 51 extractable names are referenced by some
other test under `src`, so dropping an operation breaks a behavioral test anyway.
The same check kept the three `registers all expected handlers` manifests in
`src/relay` during the previous wave, where eleven methods had no behavioral
caller at all. An inventory test is a ratchet if and only if at least one entry is
pinned solely by it; that is verified per entry, not per file.
Kept deliberately: everything a reliability gate cites, checked by case TITLE and
not only by file path, because the gate script resolves paths only; bound, quota
and provenance guards; the Windows MSYS job-breakaway and daemon-host relocation
tests, which guard failures that pass every existing gate; SSH execution-boundary
verdict vocabulary; and Git capability tests covering first fallback, cached call,
concurrent probes and per-host isolation as four distinct risks.
Coverage is partial and stated as such: of 1,449 files in scope, roughly 990 were
read case-by-case and 452 received title-and-grep triage only. The unread paths
are recorded for a later sweep rather than assumed clean.
Verified: per-area suites green (`daemon`+`skills` 294 files/3016 cases;
`git`+`persistence` 414 files/4476 cases; and the rest), gate manifest 140 gates,
`check:code-quality:changed` 0 new findings. A combined 11-path local run put
1,563 files through one machine and surfaced three timing-sensitive failures in
files this change does not touch (`history-manager`,
`structured-agent-session-refusal-retry`, `ssh-remote-commands`); all three pass
in isolation, and no production code changed, so CI's sharded run is the arbiter.
|
||
|
|
b29947d585 | fix(secrets): stop telling Linux users to install a keyring they already run (#24013) | ||
|
|
dcaef9dee5 |
test: retire relay, preload and shared cases that re-prove an owned contract (#24007)
Audit sweep over `src/relay`, `src/preload` and `src/shared` (1,087 test files reviewed). 101 case declarations removed across 40 files, 6 test files deleted outright, 1,143 lines gone. Executed-case count falls further, since several removals were `it.each` tables. Dominant patterns, by frequency: - Self-comparisons that cannot fail: `expect(f(x)).toBe(f(x))`, `JSON.parse(JSON.stringify(literal))` deep-equalling the literal for a type with no codec, and `normalizeKeyToken(t) === normalizeKeyToken(t)` presented as proof of memoization. - Object literals asserting their own fields back, where the guarantee comes from the type annotation and the runtime assertion cannot fail. - Copied inventories: constants compared to their own initializers, and a function returning a copy of an exported constant checked against that constant's literal contents. - Duplicate invocations of a contract owned at a stronger boundary, including provider-local replays of a shared helper. - Table rows varying a field production never reads, so every row runs one path. - Names promising more than the input exercises: a "Windows launch" case in a module with no platform input, and a case whose named branch is never entered. Two production symbols go with them, each a test-only export whose sole caller was a deleted case: - `getGitHubProjectRefInputByteLength` — a one-line forward to `getClipboardTextByteLength`. The real bound (`GITHUB_PROJECT_REF_INPUT_MAX_BYTES`) and its guard stay. - `GRAB_STYLE_PROPERTIES` — an intended shared source of truth that nothing ever consulted; the property set is hand-enumerated at three independent sites. One case was deliberately restored and strengthened rather than dropped. The relay integration suite is the only place the real `SshChannelMultiplexer` is wired to `RelayDispatcher`, so it reaches transport behavior the handler suites cannot (they use `createMockDispatcher`). Its `fs.writeFile` roundtrip is the one case producing a void result, and `JSON.stringify` drops an absent `result` member — a shape no other surviving case exercises. Restored with an assertion pinning what the client actually observes: `null`, not `undefined`. That assertion failed on first run, so the fact was previously unasserted anywhere. One deletion was reverted mid-audit. A case asserting that optional fields stay invisible to "old attach and ready decoders" builds those decoders from `z.object` schemas declared in the test file, so it demonstrates zod's unknown-key stripping rather than anything shipped. It is nonetheless the only forward-compatibility coverage these envelopes have, and `reliability-gates.jsonc:6232` names it as evidence verbatim, so it stays. Note that `check-reliability-gates.mjs` passed both with and without it: the script resolves manifest paths and commands, and does not check that a named assertion still corresponds to a live case. Kept deliberately: everything a reliability gate cites as evidence; the three `registers all expected handlers` RPC manifests (a dropped registration is a silent wire break no type checker catches, and one carries the STA-4571 `pty.ackData` ratchet); the `child-process` direct-import ratchet; and prototype-spy cases paired with a `.repeat(10_000)` input, which assert a real memory bound rather than merely forbidding a technique. Verified: `pnpm test src/shared src/relay src/preload` (1073 files, 11996 passed, 1 pre-existing `it.fails`, 131 skipped), `pnpm tc` after clearing `.tsbuildinfo`, `check-reliability-gates.mjs` (140 gates), `check:code-quality:changed` (0 new findings). |
||
|
|
b7209b5ae9 |
perf(git): relist only the repo whose worktrees changed, and stop blocking main on sync git (#23998)
* perf(git): stop blocking main on the open-on-remote git cascade
`getRemoteFileUrl` ran up to 6 sequential `gitExecFileSync` calls on the Electron
main thread — `remote get-url`, then `getDefaultBaseRef`'s `symbolic-ref` plus up
to four `rev-parse --verify` probes — each with its own 15s timeout and no yield
between them.
A complete async twin already existed (`getDefaultBaseRefAsync` ->
`resolveDefaultBaseRefViaExec`, sharing DEFAULT_BASE_REF_PROBES), so the sync
cascade is deleted rather than converted. `getRemoteUrl`, `getRemoteFileUrl` and
`getRemoteCommitUrl` become async; all four downstream callers were already async
(`filesystem-git-url-handlers` inside `ipcMain.handle`, `runtime-git-diff-commands`
async methods) and the provider contract already typed both wrappers
`Promise<string | null>`, so no new async plumbing was needed.
Removes 3 of the 10 `gitExecFileSync` sites and the confusing name collision with
the unrelated async `getDefaultBaseRef` in hosted-review-creation-git-state.
The base-ref regression tests keep their coverage, repointed at the public async
`getBaseRefDefault`.
* perf(git): resolve the repo root in one sync spawn instead of two
getGitRepoRoot ran `rev-parse --is-inside-work-tree` and then `rev-parse
--show-toplevel` as separate blocking spawns. Each sync git call holds the main
thread for up to its whole 15s timeout, so the spawn count is the cost — and this
function is called twice per "Add Project" on a linked worktree, once directly and
once through getLinkedWorktreeMainRepoRoot's self-recursion.
Combined into one invocation. Safe only here: in a bare repo the combined form
exits non-zero, and both that throw and the plain `false` already land on the same
marker-scan fallback. probeGitRepo deliberately does NOT combine — it has to read
`false` cleanly to go on and detect a bare repo, which the combined form's exit 128
would misread as indeterminate.
* perf(git): rebuild only the repos whose authorized roots actually changed
One worktree create called `invalidateAuthorizedRootsCache()`, which dirties every
registered owner. The next authorization-requiring IPC then rebuilt by listing EVERY
repo — and the rebuild never consulted `dirty` when choosing what to list, so `dirty`
gated only whether a rebuild ran, not its scope. At 58 repos that is 58
`git worktree list` spawns, roughly ten seconds of git wall-clock through an
admission budget of four, to rediscover roots one repo changed.
Both halves were needed; scoping the invalidation alone changed nothing.
- `markAuthorizedRootsOwnerDirty` dirties a single owner, reusing the per-owner
primitives `registerWorktreeRootsForRepo` already used. It leaves `baseRevision`
and the per-repo revision map alone — that pair is the global side-effect-token
fence, and bumping it would retire in-flight tokens for untouched repos.
- `rebuildAuthorizedRootsCache(store, onlyDirty)` re-lists only owners that are
dirty, have no listing yet, or still hold recovered roots (those are retired by
comparison against a fresh listing, so skipping them would strand them as
authorized). Only `ensureAuthorizedRootsCache` passes `onlyDirty`; an explicit
rebuild keeps re-listing everything because callers use it to force a refresh —
`filesystem-auth.test.ts` pins that contract.
`invalidateAuthorizedRootsCacheForRepo` wraps the primitive and falls back to the
global form for an unknown owner or a missing store, rather than silently skipping an
invalidation and leaving a stale allowlist. Applied to the worktree-create path.
Changes that can alter the owner SET (store swap, host/WSL re-routing, nested-repo
import, folder->git upgrade) stay global. Removal paths are not converted yet.
The allowlist contents are unchanged and the failure direction is a false denial
rather than a false allow. The relist predicate is split into its own module so it is
testable alone and the cache file stays inside its line budget without a suppression.
* test(perf): measure what git orchestration actually costs the main thread
The existing churn probe (ORCA_MAIN_THREAD_DIAGNOSTICS=1) reported spawn-initiation
cost for git/gh/glab only — its 7 call sites all sit inside git/command-runner — so
it was blind to `spawnProcess`/`runProcess`, the repo's own mandated wrapper, and to
the blocking `execFileSync('ps')` per PTY resize. That understated total churn across
115 main call sites.
- `spawn-observer.ts`: a settable seam, since shared code cannot import src/main.
Unregistered in the daemon/relay/CLI, where it costs one boolean check.
- `spawnProcess` brackets `nodeSpawn` and reports; exec-file-capture's own report is
removed because it routes through runProcess and would double-count.
- `posix-pty-foreground-group` now reports its full blocking duration. Note this
lands on the daemon, not main, whenever the daemon hosts the PTY.
- `ORCA_UNMINIFIED_MAIN=1` build flag, because a minified main bundle cannot
attribute CPU-profile self time to real function names. Defaults unchanged.
- `main-thread-git-cost.spec.ts` + `analyze-main-cpuprofile.mjs`: sweeps concurrency
against real registered repos, captures the churn lines and a V8 CPU profile of
main per phase.
What it found, which is why this is worth keeping: at the width-4 admission ceiling
(~90 git:status/s) main sees ZERO event-loop gaps over 50ms and a worst gap of 23ms,
and is 85% idle. Git orchestration does not stall the main thread. Of the cost it
does incur, spawn-init is 58%, parse 5%, stdout drain 4%.
* test(perf): name the inspector params type the anti-slop gate requires
The broad `object` parameter trips anti-slop(no-object-parameters); the only
Profiler call that passes params sends `{ interval }`.
|
||
|
|
bb874f6bb3 |
test: retire cli cases that re-run a contract the sibling already owns (#24000)
Audit sweep over `src/cli`. 23 cases retired and 2 `it.each` tables collapsed
to the rows their parameter actually reaches.
What went, by pattern:
- Table rows whose varied parameter production never reads, so every row ran
one identical path.
- Second and third invocations of a contract already proven by the case above
them, differing only in a field the assertion ignores.
- Argument-shape and private-predicate checks duplicated at the real CLI
boundary, where the same input is already driven end to end.
- Assertions whose expected value came from the same helper under test.
`src/cli/command-suggestion.ts` loses `export { levenshtein }`, a re-export no
production caller used. The one test that stubs edit distance spies on
`../shared/edit-distance` directly, which is the module `command-suggestion`
imports, so the seam it needs is unaffected.
Kept deliberately: `orchestration-lifecycle-json-rejection.test.ts` and
`orchestration-migration.test.ts`, both named in `config/reliability-gates.jsonc`
as sole evidence for a gate.
While auditing the latter, its replay dimension turned out to be inert --
`it.each([false, true])` varies `lifecycle.duplicate`, and `hasLifecycleVerdict`
(`orchestration-worker-settlement.ts:112-132`) reads only `action`, `authority`
and `outcome`. The gate at `reliability-gates.jsonc:15861` nonetheless records
"first and replayed legacy worker_done settlements are accepted". Left exactly
as found and reported rather than collapsed, because correcting a gate's claim
or adding real replay coverage is the owner's call.
Verified: `pnpm test src/cli` (131 files, 1474 passed), `pnpm tc`,
`check-reliability-gates.mjs` (140 gates), `check:code-quality:changed`.
|
||
|
|
91ab51b9fa |
fix(terminal): attach dropped images whose filenames need escaping (#23707)
Every image drop is now sent to the terminal as a bracketed paste, so agent TUIs (Claude Code, Codex, Pi) attach it. Previously, names that needed shell escaping, such as `download (1).png`, and names with spaces, which Codex's shlex splits, were typed as keystrokes or pasted raw and stayed as text. Safe names are still pasted raw. Names with spaces or shell metacharacters are backslash-escaped inside the paste on POSIX shells, which Claude Code, Codex and pi-image-paste all unescape, apostrophes included. Windows shells keep double quotes. Non-ASCII characters such as the U+202F in macOS screenshot names stay bare. Names with control bytes are still typed. Fixes #23703 Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> |
||
|
|
98039676f3 |
test: retire codex cases whose setup is inert or whose name outruns its input (#23996)
Semantic sweep of src/main/codex, daemon and git. 16 cases and 4 it.each rows
removed across 11 files; no production code touched, no files deleted.
Inert setup — the case flips the verdict by hand and the ceremony changes nothing:
- three `codex-stale-pane-accounts` cases varied `environmentHomeOverride`, which
`codex-stale-pane-accounts.ts:38-42` never reads (it reads `selectionKey`,
`homeRoute` and `accountId`); one also rewrote `.zshrc` and called
`__resetShellStartupEnvCache()` while both verdicts came from the
`activeHostHomeRoute` argument the test sets directly.
Names outrunning their input:
- an `it.each` row named `'leap century'` stepped `['1999','12','31']` to
`['2000','01','02']` — it never touches February, so it is the `'year rollover'`
row under a name promising a leap rule;
- `it.each([1, 2, 3])('keeps pre-ownership baseline version %s canonical…')`
collapsed to version 1: `config-settings-baseline.ts:185` only validates the
version is one of the three, and the policy is driven by
`parsed.mcpServers === undefined`. Version 3 without `mcpServers` is an
impossible shape, since v3 is what the writer emits *with* it.
A self-comparison: `codex-session-index-heal.test.ts:802` looped
`CODEX_SHORT_LIVED_PROBE_APP_SERVER_ARGS` asserting `args` contains each entry,
but `buildNativeHealInvocation` sets `args: [...CODEX_SHORT_LIVED_PROBE_APP_SERVER_ARGS]`
(`codex-session-index-heal.ts:298`) — the constant against itself. The full
`toEqual` with the shim is owned by `codex-short-lived-app-server-spawn.test.ts:33`;
the `CODEX_HOME` pin stays and the case is retitled to match.
Five exact source greps over `POSIX_PROVIDER_SUPERVISOR_SCRIPT` went because
`codex-app-server-posix-supervisor.integration.test.ts` executes that same script
and asserts the behavior for real — owner-PID refusal, group leadership,
group-SIGKILL escalation after a SIGTERM-ignoring provider, and exit-code relay.
Three greps in that same file are KEPT, one wave after ~84 files of that shape
were deleted, because nothing else can reach them: no integration test inspects
the provider's env, so a leaked `ELECTRON_RUN_AS_NODE` would silently change how
codex runs; and `stdin.once('close')` matters because the integration test's
`stdin.end()` would still pass if only `'end'` were registered.
Also collapsed four `it.each` blocks whose callback took no parameter, so every
row ran the identical body while the name advertised per-scenario coverage:
`['local IPC', 'SSH remote runtime']` built one transport, and
`['visible blur', 'terminal tab switch', 'split pane switch']` plus
`['terminal close', 'tab unmount']` each ran one harness. Those scenarios were
never constructed; the false claim is removed rather than the coverage, because
there was none. A single-row `it.each` whose name renders truthfully, and one
whose callback is a named function that does take the parameter, are untouched.
|
||
|
|
80786ddccb |
fix(onboarding): build the agent step around skills, not CLI registration (#22720)
* fix(onboarding): build the agent step around skills, not CLI registration The checklist step "Enable Orca CLI" was marked done once the agent skills were installed, while Settings -> Browser still showed CLI registration as a pending step. Orca terminals already put the bundled CLI on PATH, so registration only matters for shells Orca did not launch, plus WSL, where `orca-ide` exists only once registered. - Rename the step to "Give agents Orca skills"; setup registers the CLI only for WSL (isOrcaCliRegistrationRequired), via onboarding-cli-registration.ts. - Settings -> Browser drops the CLI step outside WSL (2 steps instead of 3). - Skills panel: "All skills installed" + "Update skills" replaces the disabled button; status pills sit top-right; no "Installed" beside "Unavailable". - Full Disk Access moves to the "Start work in multiple repos" step. Fixes STA-8306 / #22524. * refactor(onboarding): simplify agent-skill step state after review - One done rule: isAgentCapabilitiesDone in feature-wall-setup-progress.ts, reused by the skills panel instead of a mirrored copy. - 'unavailable' is an install-status tone instead of a second boolean; pill/note rendering moves to AgentCapabilityStatusBadges.tsx. - "Update skills" skips Computer Use when it can't run (no warning toast); setup takes an explicit selection. - The WSL gate lives in registerOnboardingCliIfRequired; onboarding deps drop the now-unreachable host CLI branches. - BrowserUsePane: one cliRequired/cliReady pair, no host CLI status fetch, WSL-only enable path, single "Finish the steps below." string. - Full Disk Access placement goes through a SelectedStepFooter switch. - Prune orphaned locale keys (and their boot-bundle entries); fix a stale comment. * fix(skills): require CLI registration only for WSL setup * fix(skills): retain CLI install labels for WSL * docs(skills): clarify remaining WSL registration fallback * fix(skills): skip WSL registration when the host confirms managed CLI access * fix(wsl): prepare managed shell wrappers before onboarding probes * test: update daemon capability and terminal hook expectations * refactor(onboarding): remove CLI registration checks from skill setup * refactor(setup): remove redundant state and obsolete registration scaffolding * fix(settings): stop registering the CLI before installing the CLI skill The General > Orca CLI skill panel still registered `orca` on PATH before opening the install terminal, contradicting the rest of skill setup. Orca terminals already provide the CLI, so the shell command toggle now says it is only for terminals outside Orca. * style(onboarding): polish the setup checklist and first-run steps Make onboarding monochrome: completion is a neutral check, selection a neutral outline, and color only flags real problems. Tighten the checklist rail and header, single-line agent cards with a grid that scrolls only when it runs out of room, a labeled permission switch under the grid, calmer notification and skill cards, sentence-case copy, and a labeled "Hide checklist from sidebar" action. Workspace setup leads with "Add project" when no git project exists. * fix(emulator): drop the Enable Orca CLI step from agent control setup Agents that drive the emulator run in Orca terminals, which already provide the `orca` command. Agent control setup in the emulator card and Settings is now a single step: install the Orca CLI skill. * fix(onboarding): hide the Full Disk Access card once access is granted A granted card has no remaining action and only takes space on the add projects step. It also no longer flashes a "Checking" state before the first status arrives. * fix(onboarding): address review on permission warning, hide button, and translations - Name the permission switch "Yolo mode" (matching Settings > Agents) and state the risk: agents act without asking and some bypass their sandbox. - Hide the modal's "Hide checklist from sidebar" button below sm, where the header centers its title under it; the sidebar entry keeps its own control. - Translate every string this PR adds into es, fr, ja, ko, and zh. |
||
|
|
34d6041c83 |
test: retire ai-vault cases whose inputs the scanner never reads (#23992)
Semantic sweep of src/main/ai-vault. 48 cases removed across 11 files; no production code touched, no files deleted. The largest single removal is a 36-case block (6 agents x 6 env values) in `session-scanner-agent-root-overrides.test.ts` whose assertion was a self-comparison: `root === join(root)`. The sibling `falls back to the default root for %j` pins `roots[0]` to the exact absolute default, and the extra roots the block also scanned (`agent_logs`, `.clawdbot/agents`) are homedir-derived and unaffected by the env var it varied. The #13082 rationale comment is kept on the surviving case. Inputs the production path never reads: - `session-scanner-codex-tool-records.ts:89` reads only `change.unified_diff ?? change.content` and never `change.type`, so the `{ type: 'delete' }` row was the identical path as `add`; the `add`/`update` rows remain as the two real disjuncts. - `codexSpawnDepth` accepts any positive integer, so depth 2 was the same branch as depth 1. - `agentPath` is an independent `??` fallback with no cross-field logic, so "keeps the rest of the spawn when the naming path is null" passes either way. - `getAiVaultWslHomeDirs` reads only `platform` and a `hasCachedWslDistros()` gate, so a case varying which distros are "currently running" took the default branch; the filtering lives entirely inside a mocked async call. Cases that cannot fail for the reason they name: an all-unknown-agent response whose throw requires `malformedSessionCount > 0` when it is 0; a symlink rejection byte-identical to the directory case above it (`isFile: () => false`); a runtime restamp whose fixture already carries the `executionHostId` and `id` it asserts. Also removed: duplicates of a stronger sibling in `session-list-results`, `session-parse-cache-persistence` (same `schemaVersion !==` gate), `session-scanner-claude-title` (owned by the subagent-prune test, which also asserts `subagentTranscriptCount`), and a session-scanner listing case whose count is N-independent because `fixedChildFileSegments` does one readDir plus a direct stat per child — so a per-session-readDir regression fails at N=1. Three keeps worth recording. Spy-counting tests were kept where real code runs: `session-scan-cutoff` and `session-scanner-dedup-batches` count `Array.prototype.sort` / `RegExp.prototype.test`, but drive the real scanner over 128 fixtures and assert the limit-ordered result too, so a re-sort-per-candidate regression fails them for the right reason — unlike a bench whose assertion was arithmetic over its own constants. `session-scanner-claude-unicode-scope` keeps its locally re-spelled dir-name encoder deliberately: importing the production one would hide Orca drifting from Claude's actual naming. And `session-parse-cache-persistence.test.ts:175` stays although `keys.length > 0` cannot fail — it is the only reference to the `satisfies Record<keyof AiVaultSession, true>` table, so deleting the case would make that type-level ratchet dead. |
||
|
|
09784740bd |
Set iceCandidatePoolSize to 0 in WebRTC egress probe (#23991)
Disables pre-gathering of ICE candidates to avoid timeout or flakiness during test probe initialization. |
||
|
|
fb52c0602a |
fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout (#23920)
* fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout
xterm paints nothing while DEC mode 2026 (synchronized output) is open and only
force-flushes after 1000ms. Codex wraps every draw in mode 2026, so any byte gap
or chunk split that loses the closing \x1b[?2026l freezes the pane for a full
second and then repaints in one burst.
Orca never emitted \x1b[?2026l anywhere, and three paths could destroy a TUI's:
the per-PTY pending cap drops buffered output wholesale (mode 2031 was already
salvaged there, 2026 was not), main sliced pending data at a blind 16KB offset
that can land inside an open frame or sever the 8-byte marker, and the renderer's
backlog warnings replace a queued tail that may hold the close.
- salvage the 2026 latch across dropped output, mirroring the existing 2031
salvage, and append the release on both delivery sites
- ground 2026 in RESET_AFTER_BYTE_GAP and the replay baseline, and in both
backlog warnings, so every drop path is self-healing
- make main's 16KB flush split frame-aware instead of a blind byte offset
- lift the synchronized-output scanner into shared/ so main and the renderer
use one implementation
Closing a frame early costs one premature repaint; leaving it open costs a
second of blank screen, so the asymmetry favours always closing.
Also adds the reproduction this needed: the pre-existing typing bench observes
the xterm BUFFER, which the parser fills while rendering is held, so it scored
these freezes as fast echoes.
* fix(terminal): stop the renderer's queue drain cutting inside an open DEC 2026 frame
takeQueuedChunk sliced a queued chunk at a blind byte offset to fit the 16KB
coalescing budget, which can strand a frame's closing \x1b[?2026l in the residual
until a later drain. Same defect as main's flush split, same fix: reuse the
frame-aware split helper.
Usually masked because the drain coalesces adjacent chunks and reassembles what
main split, but not when the budget boundary falls inside a frame.
* fix(relay): keep the SSH path's bounded slice outside an open DEC 2026 frame
pty-handler split pending output at a byte offset with a surrogate-pair guard but
no synchronized-output awareness, so a frame straddling the 16KB wire slice had
its closing \x1b[?2026l stranded in the remainder — the same defect just fixed on
the local path, on the path AGENTS.md requires us to consider.
Placed before the surrogate guard so that guard keeps the final say, and floored
at 2 so frame alignment can never walk a healthy slice into the guard's
decrement and then into the chunkChars <= 0 pause-and-retry path.
Also drops a dead `splitAt === 0` branch in takeQueuedChunk: both callers pass a
positive limit and the helper never returns 0 for one.
The two new split tests were each confirmed to fail without their fix.
* test(terminal): sweep the DEC 2026 split helper over escape-sequence shapes and every limit
Covers OSC 52, DCS, repeated open/close markers and limits 1..len+3, asserting the
result never exceeds the limit, never reaches 0, and stays byte-exact. Also pins
that a buffer beginning inside an open frame degrades to the blind offset rather
than doing something worse, and documents that callers do not thread latch state.
* fix(terminal): ground DEC 2026 on the daemon slice, the recovery replays, and the process boundary
Four more sites could strand the latch, found by sweeping every path that drops,
splits, or replays terminal bytes.
- daemon-stream-data-batcher: the 64KB bulk-write slice used a surrogate-only
clamp, and its remainder is HELD until 'drain' — "seconds for multi-MB
backlogs" per the file's own note. A frame straddling that boundary parked its
\x1b[?2026l behind the hold, blanking the pane past xterm's 1s timeout once per
frame for as long as the backlog lasted. This is the default daemon-backed pane
path, so it is the one users actually hit. The new
clampToSafeBulkWriteSplitIndex frame-aligns first and surrogate-clamps last,
and lives in daemon-stream-data-split alongside the policy it belongs to.
- replay-data-drain and remote-runtime-terminal-binary-snapshots wrote a bare
\x1b[2J\x1b[3J\x1b[H, which does not clear mode 2026 — so on the SSH/remote
reconnect path, the very event most likely to sever a frame, the whole replay
could paint nothing.
- ipc-pty-attach: trimIncompleteTerminalControlTail can cut a half-written
\x1b[?2026l while its opening marker survives in the replayed prefix.
- PROCESS_BOUNDARY_GROUND: the "process that armed these modes is gone" ground
omitted 2026, the last unexplained gap in that file. A disable, so it still
satisfies the recovery barrier's ownership scan (only ?25h may be an enable).
Recovery-path expectations updated where they pin the emitted bytes. Deliberately
NOT touched: apply-reattach-payload and ssh-snapshot-prepaint already ground via
buildSnapshotReplayPrologue.
Still unfixed, deferred with reason: terminal-output-frame-chunks.ts splits the
remote wire on accumulated UTF-8 byte width and needs a different shape than the
char-index helper; desktop clients reassemble in main's pending buffer, so the
exposure is mobile/web only.
* fix(terminal): emit the DEC 2026 release before the mode-2031 tail, and stop claiming the drop path writes it
Two corrections from adversarial review of the earlier commits.
1. Ordering bug I introduced. getDroppedMode2031RendererData ends with
`state.tail`, which extractPrivateModeScanTail deliberately retains as an
INCOMPLETE private-mode sequence so the next chunk can resolve it. Appending the
2026 release after it put an ESC behind a dangling CSI, aborting it and silently
losing whatever mode spanned the drop boundary. The release now goes first.
2. The drop-path release does not reach xterm in the dominant case, and the comment
now says so instead of implying otherwise. live-data-callback's droppedOutput
branch discards `data` and salvages only queries
(salvageRendererQueriesFromDiscardedRestoreData handles CPR/DA1/OSC colour;
\x1b[?2026l is not a query), so for hidden panes and visible panes outside
foreground-restore backpressure the synthesized release was dropped. The grounded
snapshot replay releases the latch instead.
I tried writing it through writePtyOutputToXterm there and reverted: it consumes
the pending hidden-output snapshot and broke
pty-connection-hidden-snapshot-resize-signals ("re-restores a skipped alt frame"),
so the release rides the restore rather than perturbing that state machine.
Residual gap, documented: a cap-dropped pane whose restore never arrives.
The salvage is still load-bearing on the fall-through path, so it stays.
* fix(terminal): release DEC 2026 on the reattach clears, floor the split, and correct the freeze framing
Remaining findings from adversarial review.
- apply-reattach-payload's three bare-clear branches (:63 daemon snapshot, :229
relay replay, :269 cold restore) had no release anywhere in their sequence: I
checked all seven POST_REPLAY_* profiles reachable via chooseReattachReplayReset
and none contains \x1b[?2026l. Only the buildMainModelSnapshotReplayWrites branch
was grounded, so covering the streamed replay path and not the main reattach path
was inconsistent. Verified no production code matches these clear strings — the
three test updates are mock equality, and each was confirmed to fail without the
source change.
- clampToSafeBulkWriteSplitIndex could return 0 (('\u{1F600}aaaa', 1) — alignment
returns 1, the surrogate clamp decrements to 0), which would leave a zero-length
slice that never shifts the batcher's queue entry and spin its drain loop.
Unreachable from today's only caller, but it is exported with an unstated
precondition. Floored at 1.
- Frame alignment could halve per-PTY flush throughput: main re-queues the
remainder with eligibleRound = round + 1, so the shortfall cannot be refilled in
the same round, and aligned size is floor(W/F)*F — 50% worst case in the 8-16KB
band, which is exactly the full-screen redraw burst that reaches the pending cap.
Alignment is now rejected below half the window, preferring throughput and
letting the reset profiles release the latch.
Framing corrected throughout: bufferRows records a row range and clears nothing, so
the pane freezes on its last painted frame — it does not go blank. The real trade is
"stale but coherent for <=1s" versus "immediate partial frame", and
RESET_AFTER_BYTE_GAP (written alone, with no repaint behind it in the same write) is
the one site that can newly flash a partial frame. Said so at the constant instead
of implying the release is free.
* fix(terminal): rename the shape-flagged symbols the anti-slop audit rejects
CI's anti-slop gate rejects "shape" in symbol names as structural rather than
domain language: `shapes` -> `outputSamples`, and
`writeCodexShapedEchoProbeScript`/`codexShapedEchoProbeScript` ->
`writeCodexEchoProbeScript`/`codexEchoProbeScript`.
|
||
|
|
33351b0085 |
test: retire hook and installer cases whose guards or arithmetic cannot fail (#23981)
Semantic sweep of src/main persistence, skills, agent-hooks and providers. 22 cases removed across 14 files, plus one file; no production code touched. `opencode-message-part-flood-bench.test.ts` is deleted (137 lines). Its headline assertion is `THROTTLED_POSTS < LEGACY_PART_UPDATES / 3 + 1`, i.e. `120 < 134.33` over two constants declared in the test file itself, and the second is arithmetic over two more. The throttle it is named for lives in the OpenCode plugin and is never invoked: the test hand-simulates both client behaviours and posts them to a server that only asserts `status === 204`. Timings are logged, not asserted. Negative controls whose refusal comes from a different guard than the name claims: - "keeps permission visible for unpreviewable tool input with another tool use id" is blocked by `!hasConflictingToolUseId` AND the whole input clause, making it strictly weaker than the sibling that shares the guard with that clause satisfied; - "keeps permission visible when unknown tool previews collide" refuses on `nextToolUseId !== undefined` (`server-claude-status-rules.ts:133`), never on the collision; - "pane key present but no endpoint" hits the same `-z PORT || -z TOKEN || -z PANE_KEY` guard as its sibling, since PORT and TOKEN stay empty either way. Cases that cannot fail: "skips a no-op write when contents already match" — writing identical content yields identical content whether or not a skip fired, and the comment concedes nothing is observable. "Keeps the event loop responsive" asserted `settled === false`, restating promise pendingness, while the sibling "no synchronous HOME filesystem calls" catches the regression harder. Inputs no production path reads: `reconcileEndedProcessForPaneKeys` gates on `paneHasStateClaims`, not `state`; `applyAgentStatusHooksEnabled(false, …)` returns before reading `settings`; nothing branches on spaces in a script path, since POSIX always single-quotes and Windows base64-encodes the payload (the caret/percent case is the stronger #6078 guard); `TaskUpdate` is absent from `TOOL_INPUT_KEYS_BY_TOOL`, so it takes the same `if (!keys) return undefined` arm as the unknown-tool sibling. A provider loop drops `prime-agent`, a pure fall-through of every `pi` arm with no server-side special case; `omp` stays because `source === 'omp'` genuinely diverges in the retirement path. Kept on history rather than structure: `remote-hook-service-registry-coverage.test.ts` reads like a copied inventory but is the ratchet for #7253, where Droid and Copilot shipped `installRemote` unregistered and SSH status silently vanished. A HOME-ordering source grep also stays — its behavioural sibling only reproduces the 6.6s Xcode-stub stall on macOS, so the grep is the only cross-platform guard. |
||
|
|
45c63a66e9 |
test: delete the source-grep tests an earlier detector's regex missed (#23976)
A rebuilt detector found 195 source-grep candidates where the original found 111.
The gap was one over-specific regex: the first scanner required a literal `.ts`
path inside `readFileSync(...)`, so every test that built its path from variables
(`join(dirname, '..', 'foo.tsx')`) was invisible to it. Roughly 84 files of a
pattern an earlier wave reported as cleared had in fact survived.
Deleted whole, every case asserting on production source text:
- `app-startup-routing.test.ts` (27 cases) — exact import statements
(`"import('../components/UpdateCard').then"`), relative-path spelling, and
`indexOf` source ordering. A file move or a `lazy()` refactor breaks it.
- `pull-request-page-host-boundary.test.ts` (13) — `toContain` on whole argument
expressions concatenated across 20+ component files.
- `SmartWorkspaceNameField-source-boundaries.test.ts` (7) — placeholder copy, a
Tailwind class string, and `not.toContain` on an already-deleted symbol.
- `github-project-repo-list-load.test.ts` (9) — `indexOf` statement ordering
inside `loadTasks`.
- `github-enterprise-slug-routing-boundary.test.ts` (4) —
`toContain('host: githubProjectHost(parsed?.slug.host)')`.
- `web-viewport-shell.test.ts` (3) — a regex demanding exact CSS selector-list
ordering and whitespace.
- `agent-catalog-links.test.ts` (1) — restates two `homepageUrl` literals straight
out of `agent-catalog.ts` with nothing in between.
Trimmed, keeping only what nothing else can reach:
- `desktop-startup-ordering.test.ts` 549 -> 66 lines, retaining the three cases
named as `assertionRefs` by the `ssh-filesystem.stream-inactivity-lifecycle` and
`agent-browser.owner-boundary-cleanup` gates; 15 source-order greps went.
- `ResourceUsageStatusSegment.session-polling.test.ts` keeps its census that no
`setInterval` exists and `listSessions()` is called exactly once — an added poll
multiplies a global daemon scan and no behavioral test sees it. The
`indexOf('if (!open)')` ordering pair and four `not.toContain` lines went.
- `agent-skill-installed-command-callers.test.ts` 231 -> 86, keeping the
`readdirSync` census that discovers every `<AgentSkillSetupPanel` caller and
asserts set equality against the allowlist, so a new panel host cannot silently
show a default Update action.
Also in this wave, from the renderer lib/runtime sweep: 22 cases whose routing
signal the production path never reads — verified by mutation, stripping
`connectionId`, the WSL preference and the UNC path from four of them left all 29
tests passing — plus braille-spinner rows collapsed onto one regex range, copied
`WELL_KNOWN_LABELS` rows, and a whole `resolveAiVaultResumeStartupShell` describe
whose four darwin/linux fixtures all return before the login shell is read.
`config/reliability-gates.jsonc` drops the two `app-startup-routing.test.ts`
references; the manifest still validates for 140 gates.
|
||
|
|
3b4583694f |
Show skip reasons in terminal drop upload reports (#23951)
* Show skip reasons in terminal drop upload reports Extract skip-reason copy to shared module and extend terminal drop reports to show descriptions (e.g. 'Permission denied') when all dropped files share the same known skip reason. * Move drop-skip-reason copy to dedicated i18n namespace Reorganize file skip reason strings from hooks.useComposerState to lib.dropSkipReason. Simplifies keys by removing the attachSkip prefix and improves code organization. |
||
|
|
fb67d5d7c3 |
fix(runtime): stop a busy Codex 0.150-0.157 pane reading as tui-idle (#23805)
* fix(runtime): stop a busy Codex 0.150-0.157 pane reading as tui-idle The startup header box (OpenAI Codex / model: / directory:) stays on screen and in the tail for the whole session, so as tier-1 evidence it settled tui-idle mid-turn. For a codex pane it now counts only in the quiet lane, held to the same quiescence as the composer. * fix(runtime): keep a restored Codex pane's header as tier-1 readiness A restored or reattached pane has no lastOutputAt, so the quiet lane that now holds a Codex header can never fire and the wait sat pending until timeout, where main settled it. Gate the Codex tier-1 veto on the output clock rather than the agent name, and share one settled-prompt helper. * test(runtime): read no screen in the restored Codex pane test * test(runtime): pin the clock in the clockless Codex header cases * test(runtime): name the screen-readiness comparison for what it proves |
||
|
|
2aad275ab6 |
test: delete a suite that tested only itself, and trim browser-pane duplicates (#23971)
Semantic sweep of renderer browser-pane, tab-bar, settings and dashboard-popout.
The headline deletion is `assemble-chrome/context-menu-positioning.test.ts` — 266
lines, 21 cases, and it tested nothing. Its only import was
`{ describe, expect, it } from 'vitest'`; all three functions under test were
declared inside the test file itself (`computeViewportCoords:17`,
`computeCorrection:30`, `computeEdgeFlip:116`), and a single-path rg finds those
names nowhere else in the repo. Positioning logic was presumably prototyped in a
test and never extracted, leaving 21 cases asserting their own arithmetic.
No mechanical detector in this audit would have caught it: it has real assertions,
no source greps, no copied inventories, no literals shared with a mock, and 21
well-named cases. Only the import list gives it away. Running that check repo-wide
afterwards — no production import, no file reads, declares its own subject — now
returns zero, so the class is cleared rather than sampled.
Other removals, each naming what owns it:
- "lists supported import sources before detection runs": `formatBrowserImportSummary`
gates on `detectedBrowsersLoaded && detectedBrowsers.length > 0`, so with an empty
list the `loaded: false` this case varies is inert; the sibling empty-detection
case hits the identical branch.
- the `'document'` row of `it.each(['url','document'])`: the `page.docLocation`
ternary sits inside `DeferredBrowserContent` and `localBrowserPages` does not
filter doc pages, so retention is the same code for both rows — and both panes
are mocked to identical markup.
- `it.each(['automation','mobile','viewer'])`: mount is the plain disjunction
`isBrowserPagePanePaintable`, owned at the pure boundary by
`browser-page-paintability.test.ts` across all four disjuncts.
- a popup-notice case whose `toast` is fully mocked, so no collapsing occurs and
three identical events necessarily yield identical template-derived ids.
- a subscription probe whose only unique content was `listenerCount() === 1` in a
test that never re-renders, so a missing-cleanup regression could not reach it.
`client-hosted-browser-pane-test-rig.ts` drops the now-orphaned `listenerCount`
field with it — after that case went, nothing read it.
Kept where the assertions repeat but the coverage does not: two refusal cases that
are the distinct conjuncts of `focusedGroupId !== undefined && groups.some(...)`
(missing key vs failed lookup); a `describe.each(['plain','StrictMode'])` where
StrictMode is what the renderer actually runs under and double-invokes the focus
effect; four address-bar dismissal cases mapping to four separate listeners; and a
`MARKUP_DOWNSCALE_STEPS[0] === 1` plus descending-order assertion, since reordering
that literal would ship the smallest composite first.
|
||
|
|
ad2e1b5efa |
fix(terminal): restore the mouse format with mouse tracking, so phone swipes don't type into Codex (#23946)
* fix(terminal): restore the mouse encoding with mouse tracking in every snapshot Swiping to scroll Codex from the phone on a Windows host typed legacy `ESC [ M` mouse reports into the Codex composer (#23818). SerializeAddon re-arms mouse tracking (?1000h/?1002h/?1003h) but never the SGR encoding (?1006h/?1016h). Any snapshot taken from a desktop pane's xterm (the runtime seeds its headless model from it after a reattach, and serves it to remote viewers when no model exists) therefore restored "tracking on, legacy encoding", and the phone encoded wheel events as X10 bytes, which ConPTY hands to Codex as keystrokes. serializeWithAbsoluteCursor, the one wrapper every Orca snapshot producer uses, now appends the encoding xterm itself parsed, read from xterm's mouse state service. The daemon/runtime headless model reads tracking and encoding from xterm too, so its regex mirror of the DECSET stream is deleted (one source of truth; one less regex pass per PTY chunk). Mixed versions: no wire field changes. A new host's snapshot carries an extra DECSET that old desktop and phone clients already parse; an old host's snapshot restores exactly as before. With tracking off the encoding alone sends no reports, so the wheel still scrolls scrollback. * test(terminal): pin the mouse-encoding read against the renderer xterm build * fix(terminal): type the xterm mouse-state read behind named shapes |
||
|
|
6b36a2c3fb |
feat(native-chat): mid-turn messages wait as editable cards above the composer (#23731)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * fix(native-chat): name a chat write by its target, not the owner generation A write carried the fence of the last frame the pane read, and the host refused it unless that fence was still current. An idle release and the restart after it each move the fence, and the release publishes nothing, so a send after a release was refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a cold start was refused as stale. Every write already names what it acts on: a send its conversation, a cancel its turn, a prompt answer its item revision, a rewind its epoch; an option is last-writer-wins. So admission stops comparing the client's fence, and the rebase that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it. The writer-lease check stays, and so does the attach's compare-and-swap. Frames now stamp the fence read when each frame is sent instead of a copy each subscriber kept, which went stale on the same release. * fix(native-chat): every journal append reaches the chats that are open A journal write and its delivery to open readers were two calls, and some writers made only the first. A failed start whose lease could not be handed back, a provider revision with no frame behind it, and eviction's settlement were all journaled without reaching an open chat. A journal handle now reports every durable change, and the host's session map binds that report to the session's readers when the handle is set. Writers no longer publish what they append; the per-writer publish calls are deleted. * test(native-chat): an epoch replacement reaches the open chat * test(native-chat): each row reaches an open chat once, and a live handle enters only through the map * test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite The seeded record had no surface tab id, so the next open backfilled one and that rewrite alone made the no-op transaction write. The test passed with the legacy-lease rewrite signal removed. * test(worktree-activation): restore the OMP surfaced-agent resume test The handoff removal deleted it alongside the terminal-owner tests, but it covers the surfaced-PTY block that still guards resume, including an agent whose ownership is unknown. * perf(native-chat): a publish behind a delivered commit reads nothing Each commit now delivers itself, so the publish a provider frame still sends afterwards found every reader caught up but still read rows and rebuilt the timeline for each one. A caught-up reader now skips the read. * test(native-chat): state why the teardown test's fake journal is safe to cast * docs(native-chat): say mutation admission checks only the writer lease * docs(native-chat): drop the send rebase from comments that still described it * fix(native-chat): a message is accepted, then delivered A send to a chat with no running agent restarted the agent inside the send call, before the message was recorded, so the client waited for the whole start and a failed restart refused the message. Claude held prompts sent during startup, and those could settle as "unconfirmed". A send is now accepted inside the session's serialized queue: one ledger row and one submission row marked handoverRecorded, published, answered pending. A per-session delivery loop exists while a message is queued. It starts the agent through the same serialized attach a hold uses, waits outside the queue for a Claude child to prove its start, and hands the oldest queued message over as its own serialized step, writing dispatch{pending} before the adapter call. A start it needed and did not get writes one error-tone row and rejects every queued message with the same words; a start Stop cancelled writes none. Settlement follows from the rows. A queued message is provably unwritten, so a close, an eviction or an exit rejects it. A handed-over message stays in doubt. A queued row at or below the sequence a handle found when it opened was left by an earlier process and is rejected at open, with no latch. Stop withdraws queued messages with no writer lease and no fence. An attach failure keeps the conversation open, and the attach adopts its journal. Owed work counts the loop and queued rows. A compaction or rewind found prepared when a conversation opens was started under a child this process no longer has, so the open settles it rather than leaving it to refuse every send until a view attaches. The open cursor is scoped to its epoch, because sequences restart when an epoch is replaced. Deleted: restart-before-admission, recordFailedRestart, the fence rebase, Claude's startup gate, the attach's forget on failure and its own crash boundary. Clients without agent-session.accepted-send.v1 get their reply held until the handover; the desktop and paired desktop lists advertise it. * fix(native-chat): settle queued messages only for the child that ended A child that proved its start and then exited before its message was handed over left the message queued: the exit settlement returned early when nothing else was in flight. Delivery then started another child for it, and a child that died the same way started another, without end and without a row. A retried settlement for an earlier generation, run by the attach that delivery started, did the opposite: with that generation's turn unfinished it rejected the message queued for the child being attached. The settlement now takes the rejection for queued messages from its caller. The unexpected exit and the eviction pass one, and it applies even with no other work in flight; the retry for an earlier generation passes none. * fix(native-chat): an adoption that fails to import keeps the conversation open The attach now writes into the conversation's own open journal, but a failed transcript import still closed it as if it were the attach's provisional one. The conversation stayed indexed with a closed journal, so every later send answered "could not be recorded" and every attach failed again until the app restarted. The import now closes only a journal the attach opened for itself. * perf(native-chat): the recovering open reads the journal once Every conversation open now goes through the recovering open, including the read restore of every chat at startup, which used to replay its journal once. The recovering open replayed it twice: once to probe it and again inside the open. The probe is now handed to the open as its load. * fix(native-chat): an attach that fails after indexing its child leaves no child behind A failed attach now keeps the conversation open, but a failure after `onAttached` indexed the child (the rewind or compaction recovery, or the attach's own success record) left that entry claiming a child the failure path had already released. The next send found the phantom, skipped the start, and wrote at a fence the journal had moved past, so the message stayed queued for good. The entry now drops the released child and its event sink, and follows the record's fence, as a failure before indexing already did. * fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer The error strip for a message the host accepted and then did not deliver matched the entry before the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not send your message" with nothing to retry. It now reads the reconciled entry. A rejection the journal records before the send's own pending answer lands is final as well: that answer no longer puts the entry back to dispatching with no Retry. * fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down The preamble waits for its submission to be delivered while the worker's agent starts. When that wait ran out it threw operation_unknown, and the failed-start teardown then closed the session, which rejected the very preamble the host was about to deliver. It now reports a turn start nobody observed yet: the worker is start-unknown with its session kept, the host delivers the preamble when the agent starts, and the worker's report settles the dispatch as for any unobserved start. The receipt no longer suggests reading a screen a structured worker lacks. * fix(native-chat): a message rejected while its chat was closed reads as not sent A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked every later message behind a Retry and no reason, and the delivery probe, seeing the journal already answered, never ran. The reconcile now settles it as rejected like a dispatching one. * test(orchestration): name why the readiness settlement fakes are cast * fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent * docs(native-chat): drop the fence from the admission the send effects run behind * docs(native-chat): give the fence move on release the reason that still holds * docs(native-chat): stop citing a write fence check in launch and mailbox comments Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease. * refactor(native-chat): the provider child is its own record A conversation now outlives any number of provider children, so the child is one record on the conversation's entry instead of five loose fields beside its journal. It is written in one place: indexed only once an attach has fully succeeded, and ended through one function that an exit, a failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence. - A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence patch after it are gone. - Conversation writes read the record's fence, the way mutation admission already does; a child's own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of the conversation's fence, are gone. - The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer dropped when an attach replaced the whole entry. - Stop on a child still proving its start stops only the child: its lease goes back and the chat is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus the conversation's close. - The settlement retry uses the conversation's own journal, opened through the host's one open. * fix(native-chat): the delivery loop alone settles a message its start or child failed A queued message was settled by whichever path happened to end the child first: the loop, the unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that rejected every pending row. That gave two failure rows with different tones for one start, a loop that could hand over to a different child than the one it waited on, and a Claude start that died while starting reading unlike every other failed start. - The loop remembers the child it waited on. At handover, if that child is gone or replaced, it reads how it ended: a Stop continues; anything else writes one failure row and rejects every queued message with the same words, then stops. A child still starting whose start the adapter says did not land fails the same way. The exit, eviction and the settlement retry only settle the handed-over and legacy rows of the child that ended. - One failure row, always an error, keyed by the start. A start a view began that dies with nothing queued writes the same row through the same builder, so a second report revises it. - The open no longer rejects leftovers; the loop's first step does, and the open wakes it. - `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the failure before the exit is processed. - Quit closes every conversation the way closing a chat does: what is still queued is rejected as closed, with or without a child, and a start the loop already has in flight is waited for so the child it produces is stopped rather than left behind. * refactor(native-chat): a stopped child ends on the one reading of its stop The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that verdict to the child's ending, so the host never forms a second view of whether the root is gone. Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition, and a failed re-attach passes what its release saw. The end-of-child record can therefore also carry a stop whose root was not seen to go, which nothing ends on yet. * feat(native-chat): the host says it accepts a send before any agent has it The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same string capable clients already send. A client can then tell a host that answers a send at acceptance, and admits a Stop with no writer before a turn starts, from an older one that still restarts the agent inside the send. Additive: an older client ignores a capability it does not know. * refactor(native-chat): an attach never opens a journal of its own The attach adopts the conversation's open journal, which outlives it, so it no longer opens one for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag that told the two cases apart is gone. Tests that attach without a host open the conversation the way a host does. * fix(native-chat): a moved fence resends nothing on a host that accepts first The outbox treated any fence change as a new owner: it dropped the answer of a send in flight, queued that send to go out again under the same id, and unblocked a refused head. On an older host that is how a send the restart refused, unrecorded, gets another try. On a host that records every send before it starts an agent, a fence moves because that start ran, so the same rule resent into every failed start. With a fence stamped on every frame, that became a loop. The outbox now reacts to a fence change only when the host has not advertised that it accepts a send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed start reaches the client as a rejected message it keeps with its Retry. Against an older host, or before one has answered, the outbox behaves as it did. Desktop and paired web share this hook. * refactor(native-chat): a child's end says whether the user or the host stopped it The end-of-child record's cause now tells a user's Stop from the host stopping the child for a cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a user's Stop, as before, and fails the start it was waiting on after a host stop, with the one error row and every queued message rejected, in the stop's reason when it gave one. The reason stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet. * fix(native-chat): a chat whose only work is a queued message is not offered for resume A message accepted while the agent was starting counts as working in the chat, and quit rejects it as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a chat whose agent never had the message. The snapshot now reads only what was handed over. * test(native-chat): type the queued-message fixtures in the resume-offer tests * fix(native-chat): a start that dies while a message waits on it is that message's failed start Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When that start died, its exit wrote the start's error row and left the message queued, so the delivery loop started a second agent into the same failure and wrote a second row. A child's end now records where the conversation's journal stood, and the loop settles a message accepted before a failed start ended with that start: one row, under its key, and no second start. A message sent after the failure still gets a fresh start. * docs(native-chat): say what an attach's open conversation and unconfirmed ids are now * test(native-chat): pin what a failed start settles, and what a resume offer names A view's child that dies while a sent message waits settles that message only when it died starting and no child has taken its place: a proven child's crash, or a second start since, gets the message delivered. The resume offer names the handed-over message, never a newer one still queued. * test(native-chat): the failed-start pins fail on what the message became, not on a timeout * test(orchestration): the preamble's host stub is typed, not cast The preamble send now takes only what it reads of the host, the send, the settlement wait and the record's fence, so its test builds that host with real types instead of `as never`. * fix(native-chat): a Stop that names no turn stops what the conversation has in flight Between handing a message to the agent and the agent opening its turn, there is no turn id a client could name, so a Stop in that gap was refused as "already finished" while the agent went on to answer. A cancel's turn id is now an optional precondition instead of its target: with none, the host withdraws what is queued and, when the journal still reads working, asks the adapter to stop whatever the child has in flight. Claude's interrupt is session-scoped, so it is guarded by fence and acquisition generation rather than a turn identity. Codex interrupts the turn its latest turn/start answered with until the journal shows one. A cancel that names its turn behaves exactly as before. * fix(native-chat): Stop is there from the moment a message is sent The composer showed Stop only once the agent had opened a turn, so for the second or two after a send the chat read "thinking" with no way to stop it. Against a host that takes a Stop naming no turn, Stop now shows whenever the chat reads working (a turn, a queued message, or a handed-over one still unanswered) or this client still has a message on its way. Pressing it, or Escape, first drops every outbox entry the journal does not hold yet, so nothing goes out after the Stop, then sends the conversation-wide cancel. A send already on its way reaches the host ahead of the cancel, which withdraws it there. Against an older host Stop still needs a running turn. The unconfirmed-send probe moves into its own hook so the outbox hook stays in budget. * fix(native-chat): Stop before a turn is gated on its own host capability A host that accepts sends first (agent-session.accepted-send.v1) can still predate the cancel that names no turn and would refuse it as invalid, since clients and hosts ship independently. Hosts that take that cancel now advertise agent-session.conversation-stop.v1, and the renderer shows Stop before a turn opens, and sends the no-turn cancel, only to a host advertising it. Every other host keeps a Stop that needs, and names, a running turn. The host capability probe the accepted-send hook used is generalized so both read one path. * test(native-chat): a build advertises conversation stop exactly where its cancel may name no turn * fix(native-chat): a view never restarts a chat whose last start failed A Claude chat whose CLI exits during startup left one red row per start, and every time a view bound to it (the chat opening right after its create died, or the user switching back to it) the hold started the CLI again, so the same launch-failure row repeated. Only a send retries a failed start now, the same rule provider-exit recovery already applied; the rule lives in one predicate the hold, exit recovery and the delivery loop share. * test(native-chat): start the child the loop waits on with an attach, not a second view A view no longer starts a child whose last start failed, so the R2 case that waits on a child started since the failure now gets that child from a client attach, the one non-send starter left. * fix(native-chat): settle a gone generation's turn wherever a conversation opens A send that opens a chat this process had not read yet (after a crash, from a phone or the CLI) went through the delivery open, which never settled what the dead generation left running; only the read restore and a successful acquire did. When the send's start then failed, the turn stayed running for every reader. The settlement now runs in the one journal open, at the crash boundary, for every opener except an acquisition, which settles from the evidence it read before its reserve; the read restore's separate step is gone. * test(native-chat): prove the next child's start settles the turn an earlier child left The R1 case lost its only settlement assertion when the latch it checked was deleted. It now seeds the running turn the earlier child left and asserts it ends at the exit's receipt, with the exit's row, before the message is handed to the new child. * test(native-chat): count a failed start's rows by row, not by text Comparing the set of texts passed when two different rows carried the same words, which is the duplicate the test exists to catch. * test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget * test(native-chat): pin the open's and the send's start and row counts, however the view binds Opening a fresh chat whose starts fail makes one start and one row, with two views bound before or after the create's child died; one send makes one more of each. * fix(native-chat): settle a gone generation's turn at every open but an acquisition's The journal open skipped the settlement whenever the lease read reserved or live, to leave an acquisition's own open to the acquisition. But a lease a crashed process left in recovery also reads live, until the next acquire resolves it. A send that opened such a chat, from a phone or the CLI after a crash on a host that could not prove the old owner gone, skipped the settlement; when its start then failed, the dead turn stayed running for every reader. The acquisition now says it is the opener, and every other open settles, whatever the lease still claims. * test(native-chat): hold the create's start open until the views bind The "view binds while the create is still starting" case gave the create a 300 ms head start and asserted the views bound before it died. On a loaded runner the holds took longer, the create's exit landed first, and the case failed its own precondition. The create's initialize now waits on a gate the test releases once the views are bound. * fix(native-chat): Stop reads the one working rule every session list reads While Claude retries a rate-limited request it never echoes the message, so no turn opens: the sidebar read Working from the unanswered send while the composer showed Send. The chat's working state, the host's session-list status and the host's no-turn Stop check now call one shared rule instead of three copies. * test(native-chat): a rate-limit retry pins only that no turn opens, not how its rows are kept * fix(native-chat): Stop leaves a message waiting on its Retry, and does not show for one A send that failed holds the queue until the user retries it, and one the host restarted under is parked the same way. Stop counted both as still on their way, so it showed in an idle chat and could never go away, and pressing it dropped the failed message along with its Retry. * test(native-chat): the chat's Stop and a session list read the main agent alike over their own copies The chat reduces its stream and a list reads the status feed. Driven through the real host for a rate-limit retry with no turn, a subagent still running after the main turn, and the handed-over child exiting. * refactor(mobile): the chat reads the main agent's working state through the shared rule Behaviour is unchanged: the same two terms, now from the one function the host projection and the desktop chat read. * fix(codex): a Stop naming no turn never interrupts an earlier turn It fell back to the id an earlier turn/start answered with when the latest start went unanswered, or when the journal showed a compaction Codex had not started, and reported that as stopped. * fix(native-chat): a Stop naming no turn never says a turn had already finished When the provider found nothing left to stop, for instance a turn that ended between the host's check and the interrupt, the chat got "The provider had already finished this turn." for a turn the Stop never named. It now ends quietly, as a Stop with nothing in flight does. * fix(native-chat): one Stop the host could not settle no longer refuses every later one A Stop naming no turn has one operation key per session. When the host could not settle one, it answered every later Stop under the same id as unknown until the id expired. Once the host says so, the next press is a new Stop; transport doubt still replays the same id. * refactor(native-chat): drop the composer's second error formatter After the merge with main, every chat write in the composer path reports its failure as a typed outcome worded by the refusal-notice table, so the send's catch sees only a local throw. The {code, message} formatter this branch added for it has no payload left to format, and its claim to be the one way a chat words a failure is no longer true. The composer send is main's again. * test(native-chat): pin the reason on a message rejected while its chat was closed The reopen test checked only that the message reads as not sent; it now also checks the Retry row carries the host's reason. * test(native-chat): read Stop operation ids without a cast * fix(native-chat): a Stop whose answer was lost no longer swallows the next one A Stop that names no turn has one operation key per chat. When its answer was lost in transit, the chat kept the id, so every later Stop replayed it; the host answers a replay as already handled, so for up to a day Stop stopped nothing. The id is now dropped once the call settles, however it settles. A second press while the first is still on its way still shares its id. * refactor(native-chat): a Stop naming no target keeps its operation id only for its own call The chat kept each write's operation id per payload across calls, and dropped it only on some settle paths. That is right for a write naming what it acts on, but a Stop naming no turn, and a stop of every background task, share one payload with every later one, so any path that kept the id made the next Stop replay as already handled and stop nothing. One path was still open: an answer that arrived after the chat moved to a new fence. Whether a write names its target is now decided once, before its id is picked. One that names none keeps its id only while its call is in flight, so a press made meanwhile joins it, and releases it when the call settles, however it settles. The release runs only while the key still holds that call's id, so a joined call settling late cannot drop a newer one's. This replaces the per-path exceptions for a thrown call. * test(native-chat): read the Stop fences without a cast * test(native-chat): pin the new id for a named cancel the host could not settle After the Stop naming no turn moved to a per-call id, the only test of the unknown-refusal release was gone, and the half that stays, for a cancel naming its turn, could be removed with every test green. * fix(native-chat): a Stop pressed after a new message stops it, even while the last Stop is unanswered A Stop naming no turn shared its operation id with any press made while it was still in flight. The host runs a chat's writes in order, so a message sent between two presses was accepted after the first Stop ran, and the second press replayed that Stop as already handled and left the message running, although the chat had already withdrawn it from the outbox. A write naming no target now gets a new id on every press and is never kept, so each Stop acts on whatever is running when the host reaches it. A write naming its target keeps its id exactly as before. A double press can ask the provider to stop the same turn twice, which it tolerates. * fix(native-chat): Stop no longer blinks off as Claude opens the turn for a message Claude's echo of a sent message both answers the send and opens its turn. The echo settled the send first, so the host published the message as answered one frame before the turn it opened, and for that frame the chat read nothing running: Stop turned back into Send, and Working blinked off in every session list, for tens of milliseconds on each turn. The echo now settles the send after the turn it opens has been emitted, so the running turn is published first. * fix(native-chat): a message a Stop withdrew comes back to its sender's composer A Stop withdraws every message the host holds but has not run, and S also drops the ones this client had not handed over yet. Either way the message left the chat and its text survived only in a hidden journal row and the in-memory ArrowUp history. The sending client now puts the withdrawn text and images back in that pane's composer, after whatever is typed there. Withdrawn is read from the rejection reason through one shared check, which the outbox reconcile now uses too. The composer is written before the entry leaves storage, so a failure between the two repeats the text instead of losing it, and an entry storage no longer holds is never given back again, so a replay, a second view or a remount restores it once. Only this client's outbox holds the entry, so other viewers still see the message disappear. A failed Stop withdraws nothing on the host and gives nothing back. * fix(native-chat): withdrawn text put back during an IME composition is not lost While the IME owns the field, the composer ignores a programmatic draft, and the next composed keystroke wrote the draft without the restored text, after its outbox entry had already been dropped. The composer now holds text appended mid-composition, keeps it in the cache after each composed write, and shows it once the composition settles, the way attachments that land mid-composition already wait for it. * test(native-chat): pin that only a withdrawn message comes back to the composer * test(native-chat): set up the composer's window API for every describe in the composition-race file * docs(native-chat): note that the withdrawn check reads the legacy reason until a typed category lands * test(native-chat): pin that text put back mid-composition shows once, even beside a mid-composition clear * feat(native-chat): host-owned queued-message draft store in the session journal A queued mid-turn message is a draft row in the session's journal.db, created idempotently at every writable open with no user_version bump so a downgrade stays writable. Consume converts one draft into an ordinary submission inside the journal writer's own transaction (exactly-once), and a standing writer hook returns a consumed draft only when a committed row newly settles its current consumed submission to a non-withdrawn rejection — the same decision the reducer folds rows through. Open-time repair re-derives returned state behind the stored fact; retention never prunes a row whose refusal could still return it. * feat(native-chat): queued-messages wire contract, dark capability, and send classifiers The send result becomes a union: today's submission arm unchanged, plus a capability-gated queued arm only clients that sent delivery:'queue-if-active' ever receive. Whole-list queuedMessages fields ride the subscribe events and history pages; Stop gains withdrawQueued with the withdrawn bodies in its result; clear's result carries withdrawn drafts too. Both classifiers treat queued as accepted/spent. agent-session.queued-messages.v1 is defined but deliberately NOT advertised: the rollout prerequisites (Claude fold receipt, integrated Codex steer matrix) are not in this host. * feat(native-chat): queue a capable mid-turn send as a draft, drain it at turn end, and let Stop and clear return its text A send carrying delivery:'queue-if-active' while the session owes work — or behind an actionable backlog — becomes a host-held draft instead of a submission. A serialized drain woken by journal commits, draft mutations and conversation opens re-derives its gates from live facts (streamed-event barrier first, backlog never a gate) and converts the oldest actionable draft through the exactly-once consume; from that instant today's delivery pipeline runs unchanged. Stop pauses the withdrawable frontier at the stop step (a process-level pause set that survives handle eviction and, via the per-process host instance, restarts), then withdraws it with the text in the result for capable clients; /clear does the same for the superseded source. The draft list publishes whole per emit with identity dedup, rides only the final catch-up page, and attaches to history pages. queuedMessageSend overrides queue policy only; queuedMessageDelete hands the body back. Replays for all of it answer from op-stamped tombstone receipts. * test(native-chat): pin mid-turn queueing against the real host Accept (working/backlog/text-only/budget/replay), the one-per-settle drain, returned cards with N1 overtake and the N4 re-send loop, Stop withdraw with tombstone replays, the process-level pause across evict/reopen, Delete receipts, /clear returning the withdrawn text, and publication (hydration, unchanged-cursor insert, same-frame consume, identity dedup). * test(native-chat): read the queued receipt ids before the wait closures * chore(native-chat): SAFETY rationales on the sqlite row casts and a cast-free mobile narrowing * fix(native-chat): queued-draft bookkeeping never costs a publish, an open, a clear or a history read - Cache the draft list per draft-table revision. The drain re-checks on every journal publish, so each streamed delta was running a SELECT and parsing every draft body the handle had ever written (tombstones included). - Open-time repair/prune failures are reported and skipped; they no longer fail opening the chat. - /clear on a source with no drafts answers exactly as before: no empty `withdrawnQueued`, no empty write transaction, no extra publish. A draft read failure after the committed clear no longer turns it into a refusal. - History pages read drafts through the same guarded reader as subscribers. - Publication moves to its own module; the held-draft rule lives with the pause state; one pending-prompt check; drop an export nothing calls. - Tests: restart-held drafts, pre-consume failure pause + Send retry, failed open repair, clear with no drafts. * fix(native-chat): a Stop that withdraws a consumed draft's send gives its text back A queued draft converted into a submission leaves the sender's outbox, so when a Stop withdrew that submission before the agent received it, the text had no holder: the draft stayed `dispatched` forever and nothing restored it. - The returned-card rule now follows every effective `rejected` settlement of a consumed draft's submission, a Stop's withdrawal included, with the withdrawal reason stored as the fact (`dispatchWasWithdrawn`). The writer hook and the open-time repair share the rule, so no rejected submission can leave its draft `dispatched`. - A capable Stop withdraws the cards it returned itself along with its frontier, stamped with its caller-scoped key: the text comes back once in `withdrawnQueued` and replays from the tombstone. An old client's Stop leaves a returned card. - Stop's draft steps move to structured-agent-session-queued-stop.ts. - Tests: Stop between consume and the agent's receipt for both client kinds, its replay, a crash after the withdrawal, restart in the window, and the repair of a hookless withdrawal. * perf(native-chat): the queued-draft drain takes no serialized step while the agent works The drain was woken by every journal publish and, with a draft waiting, queued a serialized step (streamed-event flush included) per publish, only to find the session still working. During a streamed turn that is one step per delta, contending with Stop and every other mutation for the session's queue. The pre-check now also skips while the session is working. Whatever ends the work is itself a commit that schedules again, and the step still re-reads every gate after its flush, so no wake is lost. - Test: queued sends during a turn take no drain step; settling the turn drains. * fix(native-chat): a clear withdraws queued text only for a caller that can take it back; paused reasons are markers An older client running /clear had its source's waiting and returned drafts withdrawn and their text returned in a `withdrawnQueued` field it does not read, so the text was lost. Clear now mirrors Stop: `withdrawQueued: true` on `agentSession.conversationCommand` (strict params, sent only when the queued-messages capability is advertised) withdraws the drafts and returns their text once, replaying from the tombstones. Without it the source keeps its cards: the supersession fence already blocks the drain, and Delete still hands the text back. A paused card's reason was host-authored English on the wire. It is now a typed marker (`send_failed`) the client localizes, like `returnedReason`; a client treats an unknown marker as a plain pause. - Tests: an old client's clear leaves the cards and its replay stays field-free, then Delete returns the text; a capable clear returns the text once and replays it; the paused marker. * fix(native-chat): a draft pause that commits no journal row still reaches live subscribers A pause writes no journal row, so it reaches subscribers only on the next publish. Two pauses had none behind them: the drain's pre-consume failure (the session is idle by then, so nothing else commits) and an old client's Stop that interrupted nothing. A live card kept reading as waiting, with no failure marker, until some unrelated commit arrived. The drain now publishes after pausing a draft it failed to convert, and an old client's Stop publishes when it paused a frontier. - Tests: a failed conversion and an idle old-client Stop each reach a live subscriber as a paused card; both fail without the fix. * fix(native-chat): a failed clear wakes the queued drain, a failed Stop withdrawal still publishes its pause A conversation command can settle on the record alone (a retried clear that fails), so drafts held behind its prepared phase waited for an unrelated journal commit; the command controller now re-derives the drain when any command finishes. A capable Stop whose withdrawal write failed never published the pause it set, and a publish failure after a committed withdrawal (Stop or clear) dropped the bodies from the answer; publishing now happens outside the withdrawal and can no longer discard its result. Tests reset the process-level pause set between cases: operation ids repeat per test, so a shuffled order held later tests' drafts. * refactor(native-chat): the draft store notifies through the journal's commit listener, the hold is a stored row fact, and one typed gate decides every queue hold R1: every standalone draft-table transaction that changed rows (insert, withdraw, hold, open-time repair) fires the journal's own commit listener after COMMIT, so a draft or hold change publishes and wakes the drain through the same path a journal row does — no call site can forget. All hand-written publish/wake plumbing for draft changes is deleted; wakeQueuedDrain survives only as the record-input wake (a conversation command can settle on the record alone). R2: the process-level pause set becomes a hold_reason column on the draft row (pre-ship, so no migration): holds survive eviction and restart, keep their send-failed marker across restarts, die with the session's journal, and are cleared by consume and withdraw in their own UPDATE. The host-instance derivation stays the one restart mechanism. R3: one typed structuredQueueHold (blocked | command | prompt | working) consumed by admission, the drain step and Send-now, with each caller's override set written beside it. A capable send during a late-result /compact now queues instead of being refused (PLAN §3.1); the dead prepared-command branches and the drain's duplicated gate list are gone. prompt outranks working so Send-now's one override cannot swallow it. R4: one isUnsettledQueuedMessage predicate for the withdrawable/budget filters. Loop 4: a replayed send whose draft was refused answers with the returned card, never the rejected submission, so the text cannot render twice. Rewind completion was verified to publish after the record clears (the rewind path's own publish; the open path's recovery precedes the open snapshot). * fix(native-chat): a Stop with no drafts writes nothing, and a failed hold still lets a capable Stop withdraw The stored hold turned Stop's in-memory pause into a draft-table write, so every Stop (drafts or not, capability advertised or not) opened a BEGIN IMMEDIATE/COMMIT. An empty hold now returns before the serialized write. A hold that threw also emptied the frontier, so a capable Stop withdrew only returned cards and left the waiting drafts unheld to auto-send after the interrupt. The frontier is read once and survives a failed hold. * fix(native-chat): a capable Stop with no drafts writes nothing The empty-hold guard from the previous fix did not reach withdraw, so every capable Stop still opened a write transaction after the interrupt, and a closed handle turned its empty answer into a missing field. The draft store now answers an empty withdraw without a transaction, for every caller. * refactor(native-chat): Stop and /clear never withdraw queued drafts; no text rides the wire back Adopt the host-owned-queue model end to end: a Stop holds the waiting frontier ('stopped') for EVERY client and interrupts — the cards stay published as paused, Send-now overrides per card, and the pause dies when the user next starts a turn (an ordinary dispatched send lifts 'stopped' holds in the same serialized step; 'send_failed' holds still need their explicit Send). /clear carries the source's unsettled drafts to the replacement session as born-held rows — identical for every client version — then tombstones the source. Delete answers with no body: the card leaving the published list is the outcome. Removed (never shipped; the capability was dark and unadvertised, so no wire compatibility is affected): CancelParams.withdrawQueued and its refine, ConversationCommandParams.withdrawQueued, CancelResult.withdrawnQueued, ConversationCommandResult.withdrawnQueued, AgentSessionWithdrawnQueuedMessage, the Delete result body, settleStopQueuedWithdrawal and the cancel finisher, withdrawClearedSourceQueuedMessages, replayWithdrawnQueuedMessages, and cancelPlan's tombstone replay. This also removes the defect where a withdrawal took every row regardless of which client sent it (a phone Stop pulled desktop-typed text): nothing moves text anymore, so a Stop from one client can never relocate another client's drafts. Hold and carry writes are bookkeeping: a failure is logged and never gates the interrupt or the clear. * feat(native-chat): a restart hold lifts like a Stop's, and paused cards say why The user's next dispatched send lifts every stop-shaped hold in one UPDATE: stored 'stopped' rows, and restart-held rows (host_instance mismatch), which are adopted into the running instance — the same fact the derivation reads, so no second copy of the hold exists. 'send_failed' still requires its explicit Send. Publication now marks stop/restart holds with pausedReason 'stopped' (an additive optional value on a dark capability), so clients can caption them "sends after your next message" and keep "couldn't send" for 'send_failed'. * fix(native-chat): only a client's own send lifts a Stop's queue pause The lift ran for every accepted host send, so orchestration mail, a restart continuation and a launch prompt released drafts the user had stopped (and adopted restart-held rows into the running instance). The client-facing agentSession.send RPC now marks its sends as the user's own; host-internal senders leave the pause alone. Also drops comments still describing the withdrawn return-text rule. * fix(native-chat): a Stop's queue pause lifts when the user's send starts its turn The pause lifted as soon as the host accepted a user send, so a send the provider then refused (a failed child start, a refused turn/start) had already released the stopped drafts into the same failure. The host now remembers a client's own send, in memory, until the provider answers it: acceptance lifts the stop-shaped holds, a refusal forgets it with the holds intact, and a later Stop supersedes it. Nothing is persisted, so a restart between the send and its turn start leaves the cards held for the user's next send rather than sending them unasked. * fix(native-chat): a consumed draft's turn starting lifts a Stop's queue pause Drafts are only ever a client's own sends, so a drained draft or a Send-now is a user send for the pause: its submission joins the same in-memory set a direct send uses, and the provider accepting it lifts the stop-shaped holds. Before, a message typed while a stopped turn wound down drained as a draft and left the older stopped cards held, so their "sends after your next message" caption was false. A refused consumption lifts nothing, a later Stop still clears the set, and orchestration mail and restart continuations still never lift. * fix(native-chat): queue a capable send behind a /compact and re-scope /clear's carried drafts - A text send with queue-if-active during a /compact in flight is admitted on the compact's side lane as a held draft instead of being refused; it may only become a draft, so one the gate no longer holds is refused rather than dispatched. - Drafts /clear carries to the replacement are fingerprinted for the replacement session, so the provider's echo folds into the sent bubble. - The in-memory set of user sends awaiting their turn is capped; sends settling unknown no longer grow it without bound. - Correct the userSend comment: the renderer's launch prompt goes through the client RPC and does set it. * feat(native-chat): queued mid-turn drafts become editable cards above the composer Against a host advertising agent-session.queued-messages.v1, Enter stamps the send 'delivery: queue-if-active' (chat-wide 'Queue follow-ups' setting, on by default) and the host's published drafts render as compact cards between the transcript and the composer — never as transcript bubbles — with Steer (send-now, Cmd/Ctrl+Enter for the newest), Delete, and a menu with Edit message and Turn off queueing. Returned cards show the stored effective rejection with the same words a rejected submission gets (a Stop-withdrawn one says so); paused cards localize the host's typed marker, and an unknown marker reads as a plain pause. Hold captions are derived client-side; the wire carries none. Restore is write-ahead: Stop, Edit and a capable /clear (withdrawQueued on conversationCommand, fingerprint-matched to the host's digest) persist their operation identity before the RPC and append the withdrawn bodies to the composer draft exactly once — replays answer from the durable restored record, and a /clear's text lands in the replacement session's pane. A marker left by a crash is RELEASED, never replayed: an unadmitted operation-id replay would execute the command, so a reopened chat can never be cleared, nor new work stopped, by a press from before a crash; unwithdrawn drafts stay visible as cards. Text never duplicates: an outbox entry the host visibly holds as a draft (same id) or answers for in withdrawnQueued retires without a local restore, and Stop's host-side restore skips ids the outbox withdrawal already put back. The renderer carries the list everywhere frames flow: reducer (live over stale history, omitted means unchanged) and the frame coalescer (latest wins, like commands). Everything is capability-gated: an older host sees byte-for-byte today's requests — no delivery key, no withdrawQueued, no queuedMessage RPCs. The capability stays dark; nothing here advertises it. * fix(native-chat): queued-draft restore survives a lost answer and an unconfirmed /clear - A capable /clear reuses the operation id write keeps for an unconfirmed clear, so the next press replays it; a fresh id each press was refused by the host for as long as the first stayed unconfirmed. - Edit, Stop and a capable /clear replay a lost answer (the call threw) under the same operation id, bounded and in-session, so withdrawn text still comes back after the card has gone. A refusal or fence move stays final; a crash marker is still only released on remount. - Restored-id bookkeeping lives in memory beside the draft cache it guards; storage holds only in-flight markers, validated per element, removed when empty. The /clear marker is written only when the clear actually sends. - A mid-turn queue send awaiting its answer, or already held as a draft, no longer paints as a transcript bubble next to its card. - One action per card at a time; Edit/Delete hand focus to the composer. - Revert unrelated en.json reflow. * fix(native-chat): a lost Stop never lands on newer work; the steer chord never skips typed text - A Stop whose answer was lost is replayed only while the turn and sends it was aimed at are still what is in flight; once another turn opens or a newer send lands (e.g. a queued message drained), the Stop is reported unconfirmed instead of interrupting work begun after the press. - Cmd/Ctrl+Enter steers the newest queued card only from an empty composer; with text or an image in the composer it stays a plain send. - A mid-turn queue send hides from the transcript only while it is on its way: from the entry the drain is stopped on (read through the drain's own rule), sends stay visible as bubbles beside the Retry row. A rejected entry holds nothing up, so what follows it still becomes a card. * fix(native-chat): a send the host visibly holds as a draft frees the outbox's single flight The published draft list is the host answering the send, exactly as a journal row is: retiring the in-flight entry now also releases single-flight and voids the unsettled reply. Before, a slow or lost reply kept the next mid-turn message waiting, hidden (neither card nor bubble), until the RPC timed out. * fix(native-chat): Steer hands focus to the composer like Edit and Delete A steered card leaves the list once the host sends it; focus on its Steer button fell to the document body, so the next keystroke went nowhere. * fix(native-chat): a lost /clear stops replaying within seconds, so sends never wait on bookkeeping Sends are refused while a clear settles. Each clear call can run for its full 195 s timeout, so three lost-answer replays could hold the composer for about 13 minutes. Replays now start only within 10 s of the press: a slow first call is never followed by more, and at most one replay can outlast the window. * refactor(native-chat): queued drafts stay paused cards; no draft text ever rides a wire answer Stop and /clear go back to main's plain writes: the host pauses its drafts and carries them across a clear, so nothing needs restoring and cards stay visible on every device. Edit copies the text the card already shows into the composer before a plain Delete, so no RPC outcome can lose it. The write-ahead restore journal, replay loops, the Stop wrong-turn guard, and the clear replay window are deleted with the contract that needed them. Stop's local outbox step keeps an issued queue send whose answer is still out — the host may already hold it as a card, and its answer settles it — so the same text can never appear twice. * test(native-chat): drop the removed tabId option from the queued gating test * fix(native-chat): paused cards caption per published reason; first card reaches the live region A Stop's hold ('stopped') says it sends after your next message, a failed consume ('send_failed') asks for Send, and an absent or unknown marker reads as a plain "Paused" instead of promising a resume the host may not do. The live region now stays mounted while empty so the first queued card is announced. * fix(native-chat): a paused or returned card's Send tooltip no longer promises to skip a turn * fix(native-chat): show the queue follow-ups switch only when the host queues messages The switch rendered whenever structured chat was on, even though a host that does not advertise agent-session.queued-messages.v1 ignores the preference. It now reads the local host's capability through the existing structured host-capability hook and stays hidden until the host says it queues. The copy now also says that messages with images send right away, since image messages never queue. Updated in all six catalogs. * fix(settings): find the Queue follow-ups switch when searching "queue" The switch renders inside the Chat UI settings entry, whose search keywords never included "queue", so settings search hid it. Add a localized "queue" keyword to that entry in every locale catalog. * fix(native-chat): a returned queued card carries the typed rejection fact, like a rejected submission A consumed draft the agent never ran comes back as a returned card. The card kept only the rejection's sentence, while its submission now also records the typed fact a client classifies from. A host-restart rejection's sentence carries no legacy marker, so such a card could not be told apart from a provider's refusal. The draft table stores the submission's fact next to its reason (`returned_rejection`, written by the same settlement that sets the reason, and read back with the reducer's own fact reader), and the card publishes it as `returnedRejection`. Both are overwritten on every return, so a re-sent card never keeps an earlier refusal's fact, and a /clear carry inserts a plain held draft with neither. Retention moves to queued-message-retention.ts to keep the table module within max-lines. * fix(native-chat): say why Stop keeps a dispatching queue send that is not the in-flight one A pending answer frees single-flight but leaves the entry dispatching until its journal row lands. * fix(native-chat): word a returned queued card from its typed rejection fact A returned card is classified and worded exactly as a rejected submission: returnedRejection decides, returnedReason is the fallback. A host-restart card now says Orca restarted instead of the generic not-sent line. * fix(native-chat): fit the queue to main's typed rejections and compaction result Main (#23026) dropped the disposition's fresh-id retry field, gives a rejected dispatch a typed sentence plus fact, and types /compact's result. The queued-draft disposition and the queue tests now use those shapes. * fix(native-chat): a returned queued card's words leave out sending again The card offers its own Send, so its caption is worded with the retry control present, as the delivery notices are. * fix(native-chat): a queued send in doubt that survives a Stop waits for the user's Retry The unconfirmed probe resent it onto the session the user had just stopped, starting a new turn when the host never got the first attempt. A Stop now parks it the way a recovered unknown is parked. * fix(native-chat): a withdrawn send the host returns as a card is not also put back in the composer When the withdrawn submission and the returned card arrived in one frame, the journal reconcile restored the text before the card retired the entry, so it showed twice. * fix(native-chat): a send stops asking the host to queue it once the host no longer can delivery was fixed at enqueue, so after a host rollback every Retry of a queued send was refused on the same strict field. It is now decided per attempt: an id already sent keeps it while the host can read it, an id never sent takes the current choice, and a host without the capability never sees it. * fix(native-chat): draft bookkeeping can never roll back the journal row it rides The queued-draft returned transition runs inside every journal append's transaction. A throw there (a draft table an earlier build created without the returned_rejection column) rolled back the journal's own rejection row, so a Stop, a failed start or a provider refusal could not be recorded. The standing hook now runs in its own savepoint: its failure is logged and rolls back alone, and the open-time repair re-derives the missed transition from the committed row. The draft table also gains any missing nullable column at open. * fix(native-chat): a draft a Stop or restart took back waits again instead of blocking the queue Cards A, B and C wait; the turn ends and the drain consumes A, but the agent has not taken it yet. A Stop then pauses B and C and withdraws A's submission, which made A a returned card. The user's next send lifted B and C, yet a returned card blocks everything behind it, so B and C never sent although they read "sends after your next message". A restart or close before hand-over did the same. Nobody failed the user there, so the draft now goes back to waiting at its own position, under the hold that same event put on the drafts behind it: a Stop's 'stopped', or no stored hold after a restart, whose hold derives from the host instance. It carries no refusal, and records its spent submission id in consumed_as, so its next consume (the drain, or Send on the card) mints a fresh id through the same path a returned card's re-send uses. Provider refusals and other failures still return the card. The live settlement hook and the open-time repair share one decision. After a Stop and the user's next turn, A drains first, then B, then C, one per turn. * fix(native-chat): Delete and Send on a queued card answer at once during a /compact A /compact holds the chat's serialized lane for its whole provider call, and the queued-card Delete and Send ran on that lane, so both hung until the compaction finished. They now run on the side lane a draft-only send already uses while a compaction is in flight: Delete completes at once, and Send reaches its readable "wait for the conversation operation" refusal at once. The drain stays on the main lane and keeps its command hold, so nothing sends until the compaction settles. * fix(native-chat): a re-sent returned card drops the refusal it came back with Re-consuming a returned card left returned_reason and returned_rejection on the now-dispatched row, so the row described a refusal that no longer applied. The consume clears both in the same update that moves the card to dispatched. * perf(native-chat): the queue gate reads pending prompts without rendering the journal The prompt check ran on every send admission and drain step, and read journal.snapshot(), which copies and sorts every item in the chat. It now walks the reduced items in place with journal.visitItems; the answer is the same, since the snapshot only sorts those items. * fix(native-chat): a Stop that fails leaves the queued cards as it found them Stop holds the waiting cards before it withdraws queued sends and interrupts the agent. When a later step threw or the Stop was refused, the cards stayed paused ("sends after your next message") although a failed Stop is meant to change nothing. A failed Stop now undoes exactly what it added: each card it held gets back the hold it replaced, a consumed card its withdrawal sent back to waiting is released, and the user sends it had set aside can again lift the pause. Holds an earlier Stop or a restart put on the cards stay. The hold SQL moves to its own module, and the draft store's standalone transactions share one helper. * docs(native-chat): confirmed cancellation is no longer a queue rollout prerequisite Stop withdrawing queued sends with a typed cancellation landed on main with #23026. The comment gating the queued-messages capability now lists only what remains: the Codex steer matrix (#21062), the Claude fold receipt, turn-owner bars, and the desktop and phone clients. * docs(native-chat): the Claude fold receipt and turn-owner bars have landed; Codex steer and the clients remain * test(native-chat): type the returned-card restore test's hook props * fix(native-chat): a draft a Stop put back stays visible as a card The card list hid a waiting draft whose id already had a submission. A Stop that withdraws a consumed draft requeues it under the same id while the first submission stays rejected, so the draft vanished from both the cards and the transcript. Only a submission that was not rejected now hides its card. * fix(native-chat): Send on a queued card during a /compact is refused before it takes a lane Send-now chose its lane once, at entry. During a /compact it took the side lane, where it could wait behind a Stop, then run after the compaction had settled and append a real submission unserialized against the main lane. While a compaction is in flight, Send-now is now answered with the "wait for the conversation operation" refusal before entering any lane, and otherwise it runs on the main lane. Only Delete keeps the side lane, whose compare-and-set withdrawal is safe on either. * fix(native-chat): a Stop that fails after reaching the agent keeps the queue paused A failed Stop undid its queue holds whenever it threw, including after the interrupt had already gone to the provider (a status-note write failing after cancelTurn, or after stopping a starting agent). The turn could be stopped while the cards drained as if no Stop was pressed. The Stop now marks the step that reaches the provider, and undoes its holds only when it failed before that. A Stop the agent refused answers ok and keeps its holds; the comment no longer claims otherwise. * fix(native-chat): an unanswered capability probe no longer rewrites a queued send The per-attempt delivery decision was stored on the entry, so a replay during the window before the host's queued-messages probe answered, or after it failed, was saved without delivery; the host's ledger then refused every later replay of that id. The entry now keeps the user's intent, the wire field is decided per request, a queue send waits while the capability is unknown, and a failed probe is asked again when contact with a remote host is regained. * fix(native-chat): a skipped draft settlement heals on the next drain step, not only at reopen The draft settlement rides each journal append as bookkeeping, and a failure there is logged and skipped. Only the open-time repair re-derived it, so a consumed draft whose submission was rejected stayed dispatched (invisible, and blocking nothing it should) until the chat reopened. The re-derivation is now its own function, shared by the open-time repair and the drain: whenever a dispatched draft's submission is already rejected, the drain step applies the owed settlement first. * fix(native-chat): a queued send in flight when Stop lands is never resent by the probe Stop parked only sends already unconfirmed; one still dispatching whose answer later came back unknown was left to the unconfirmed probe, which resent it onto the stopped session. The entry now records that a Stop outlived it, the probe skips it, and only the user's Retry, which clears the mark, sends it again. This replaces the retryAfterUnknownSubmittedAt parking for the unconfirmed case. * fix(native-chat): one id is never recorded as a submission twice A second submission row under an id the journal already holds replaces the submission with a fresh pending one, so a rejected message could be handed over again under its own id. Send on a queued card could do exactly that: if the host died after it consumed the card under the operation's id but before its answer settled, the rerun consumed again under the same id. The journal now refuses a submission under an id it already records, so no id is delivered twice whatever the caller does. And a Send-now rerun that finds the card consumed under its own operation id answers with that submission instead of consuming again. * fix(native-chat): a waiting draft whose first send the agent echoed is withdrawn, never resent A consumed draft goes back to waiting when its submission is rejected as never delivered (a Stop's withdrawal, a restart, a close), and then sends again automatically. That rests on the "never delivered" claim. If the provider then echoes that message, the first delivery happened, and the automatic resend would give the agent the same message twice. The reducer already keeps such an echo apart, since a rejected submission may not claim it, so the draft store reads it from the appended row itself: a provider echo of a user message that no live submission claims, matching a waiting draft whose spent submission is rejected, withdraws that draft the way a Delete would. The echo-claiming rule is split out of the reducer's aliasing so both read the same decision, and the per-row draft hook moves beside the settlement re-derivation. * feat(native-chat): a submission names the queued draft it hands off Clients told a queued card's hand-off apart from other sends by comparing the draft's id with the submission's id. That holds only for a draft's first hand-off: a re-send, or a draft that goes back to waiting and drains again, goes out under a fresh id, and the clients showed the card and the sent message together, or restored text the host still held. Every submission the host creates by handing off a draft now carries queuedMessageId, the draft's id. It is written on the submission's journal row as an optional key (older readers keep it and ignore it), carried by the reducer, listed in the published submission schema (which otherwise strips it), and stamped where the row is built from the consume itself, so no hand-off path can leave it off; a caller naming a different draft is refused. A direct send names none. The queued-messages capability comment makes the link part of v1. * refactor(native-chat): every queued draft goes out under a fresh submission id A draft's first hand-off reused the draft's own id as the submission id, so comparing a draft id with a submission id looked right in every first-send test and failed only on a re-send or a requeued draft. Every hand-off now uses a fresh id (the drain mints one; Send on a card uses its operation's id), so id equality is never true and a reader must use the submission's queuedMessageId. The host gets simpler: queuedMessageNeedsFreshSubmissionId is gone, consumed_as is set on every dispatched row and cleared when a withdrawal sends the draft back to waiting (its spent submissions stay findable by their link), the consume refuses the draft's own id, and the consumedAs ?? messageId fallbacks collapse. The delivered-echo check finds spent hand-offs by link. A send this host queued, asked again (a lost answer's replay, or a rerun the operation ledger no longer covers), answers from its draft and then from the hand-off that names it, through one function. The rerun path used to be kept from sending twice only because a submission sat under the send's own id; with fresh ids that guard is now explicit. A Send-now rerun recognises its own consume by the link instead of consumed_as. * refactor(native-chat): a queued send is matched to its hand-off by queuedMessageId, never by id The host now hands every queued draft off under a fresh submission id and names the draft on the submission. The card list hides a waiting card only for a live hand-off linked to it; an outbox entry a submission links to belongs to the host in any state (one rule, in a queue-aware reconcile both readers use); a withdrawn hand-off is never restored to the composer; a send answered with the hand-off settles as held; and the Stop and in-flight checks read the same link. This replaces the rejected-submission filter and the published-id restore skip. * test(native-chat): read the outbox only after the replayed send's answer lands * fix(native-chat): an echo withdraws a draft only if its rejected hand-off reached the agent The delivered-echo rule withdrew a waiting draft when a provider echo matched any rejected hand-off of it, including one a Stop rejected before it was ever handed over. That hand-off is provably unwritten, so a matching unclaimed echo is some other message, and the rule silently deleted the card. Only a hand-off that was handed over and then rejected as never delivered can be disproved by an echo now. * fix(native-chat): a skipped echo withdrawal is re-derived before the draft can send again The delivered-echo withdrawal rides each journal append as bookkeeping, and a skipped hook left the draft waiting, so it later sent the same message a second time. Nothing re-derived it. The draft store now also withdraws, in its owed-settlement pass, each waiting draft that an echo already in the journal proves delivered: an unclaimed provider user message (still stored under its own id), carrying the draft's payload, appended after a hand-off that was handed over and rejected. The live hook and the re-derivation share one predicate. The pass runs at open and in the drain step, right before a draft would send; it reads every item, so it never runs per streamed row. * fix(native-chat): a rolled-back journal append leaves no draft state cached The draft store caches its row list by revision. The per-row hook read that list eagerly inside the append's transaction, after the consume in the same transaction had already written and bumped the revision, so a failed COMMIT left the cache showing a hand-off that never happened. The hook now reads the drafts only once a row holds an unclaimed echo, and any rollback of a journal append or of its bookkeeping savepoint invalidates the cache, so no other read inside the transaction can leave it stale either. * fix(native-chat): a replay of a deleted queued card answers withdrawn, not refused Once a deleted card's tombstone is pruned, a replay of the send that queued it found the draft through its last hand-off. When that hand-off had been rejected (the card came back, and the user then deleted it), the replay answered with the rejected submission, which clients show as a failed send with a Retry. Only a withdrawn row is pruned while its last hand-off stands rejected, so the replay now answers queued, withdrawn. * fix(native-chat): a queued send records what it sent instead of waiting on the capability The round-two hold kept a queue send back while the host's capability was unknown, which hid its text, wedged every later send, made Retry a no-op and let Stop restore text the host held. There is no hold now: the entry records what its first attempt sent and every replay sends exactly that (dropping it only for a host known not to read it); a first attempt asks to be queued only of a host known to queue with the setting on, and otherwise goes out plain as before; the transcript hides only a send whose request asks to be queued; Stop keeps any queue send that has gone out and parks it, unconfirmed, for the user's Retry, which the drain and an owner change now respect too. * test(native-chat): a replayed send of a deleted card is spent, with no restore and no Retry * refactor(native-chat): name the queue's pause-lift for what it releases * chore(native-chat): one import of the mutation helpers * fix(native-chat): a Stop marks a queue send without rewriting its state; the entry stores only what it sent Stop set an in-flight queue send to unconfirmed, so a settled refusal answering its first attempt kept the old id, and every Retry replayed into the same recorded refusal. Stop now only marks the entry, and the drain never admits a marked queued entry, so its state and refusal notice stay what its answer made them. The stored delivery intent is gone: an entry keeps only sentDelivery, recorded by its first attempt; a never-attempted send decides at attempt time. * docs(native-chat): no send waits on an unknown queued-messages capability * test(native-chat): one import of the outbox module in the owner-change test * test(native-chat): read a stale outbox entry without a JSON round-trip * fix(native-chat): keep a first attempt's recorded delivery a literal * fix(native-chat): an interrupted send attempted before a Stop reads as unconfirmed, never as not sent * feat(native-chat): a Stop pauses the whole queue, derived from the journal, with an explicit Resume After a Stop, each waiting card was held on its own row ('stopped'), lifted when the host saw, in memory, that a user send made after the Stop had its turn accepted. The cards read "sends after your next message" one by one, there was no way to resume the queue without sending something, and the in-memory record of user sends was lost on a restart or eviction. The pause is now the queue's, and derived rather than stored as a flag: - 'stopped': the user's last Stop took effect at a recorded journal position and no turn a person asked for has started since. "A person asked for it" is the new `origin: 'client'` on the submission row (a send over the client send RPC, or a card they sent now); orchestration mail, a restart continuation, a host-sent launch prompt and the queue's own drain record `host` and never lift it. - 'restarted': a waiting card was written by another host process and no person's turn has started since this conversation opened. Resume (`agentSession.queuedMessagesResume`) lifts either. Send-now sends one card; the rest stay paused until that card's turn starts, which is a person's turn like any other. The journal's row kinds are closed (an older build truncates a journal at a row kind it does not know), so the one event the journal cannot carry, where the Stop took effect, is recorded beside the drafts in `queued_message_pauses`; everything after it is read from the journal. A Stop records it only once it takes effect (after withdrawing queued sends, as it reaches the agent), so a Stop that fails first leaves nothing to undo, and the per-row hold, its undo and `userSendsAwaitingTurn` are gone. A card keeps a hold of its own only when its conversion failed ('send_failed'). The pause is published once, as `queuePause` beside `queuedMessages`, on live frames, catch-up and history. A /clear starts its replacement paused, as after a Stop, since the carried cards were written for the context it discarded. * feat(native-chat): a /clear's replacement queue reads paused because of the clear, not an interrupt The replacement's pause was recorded as 'stopped', which clients show as "Queue paused because you interrupted" although the user cleared the chat. It is now its own reason, 'cleared', on queuePause.reason ('stopped' | 'restarted' | 'cleared'). It lifts and resumes exactly like a Stop's: through Resume, or the user's next turn starting on the replacement. * feat(native-chat): a paused queue shows one header row with Resume; cards keep Steer The host now publishes the queue's pause once (queuePause: stopped, restarted or cleared) beside the list, and per-card holds mean only a failed send. The card list shows a header row above the cards naming why the queue is paused, with a Resume button that calls agentSession.queuedMessagesResume (a failure is the usual toast). Cards keep Steer, Delete and More actions while the queue is paused; the old per-card paused caption is gone. Steer's tooltip now reads Submit without interrupting the model. * test(native-chat): the coalescer keeps the pause of the latest list * test(native-chat): fit the queue tests to the queue-level pause types * fix(native-chat): a queue pause covers only the cards it paused A Stop recorded its pause fact even when the queue had no cards, and the fact outlived the cards it did pause. The published list hid a pause over no cards, but the drain still treated the queue as paused, so a card typed much later — during an orchestration-mail turn, or a correction typed before the stopped turn ended — sat under "paused because you interrupted" with no Stop of its own. A Stop now records its pause only if the queue holds a card when the Stop takes effect (the hand-offs its withdrawal sent back included). The fact is retired in the same transaction as the Delete, consume or withdrawal that empties the queue, never from an async publish. A /clear's carry now lands each card with its 'cleared' pause in one transaction, so a failed insert leaves no pause over an empty replacement. * perf(native-chat): the queue's pause reads the latest person's turn in O(1) The pause is derived on every publish, per subscriber, and each derivation copied and scanned every submission to find a person's accepted turn after the Stop. The reducer now keeps that fact as it folds rows: the submission row of the latest accepted turn whose origin is `client`. The Stop's and the restart's lift both read it directly. * fix(native-chat): each Resume press is its own operation, and a pause-only frame updates state Resume names no target, so reusing its operation id after a failed press replayed a stale answer or the same refusal. The reducer's no-change check also ignored the queue pause, dropping a frame that changed only the pause. * fix(native-chat): a card handed off after a restart belongs to the process that sent it A draft's host_instance was only ever the process that first wrote it (or adopted it while waiting). A returned card from before a restart, sent again in this process and then withdrawn back to waiting, still carried the old process, so it raised a 'restarted' pause although no restart happened since it was sent. Every hand-off (the drain, Send on a card) now stamps the handing-off process on the draft in the consume's own update. * fix(native-chat): the paused-queue row shows only over cards Resume can send, and matches the queue's icons The header row appears only when a card waits on nothing but the queue's pause; Resume shows it is pending, hands focus back to the composer like the card actions, and the truncated line keeps its full text as a title. A paused queue outranks a pending prompt in the card's hold, as on the phone. Steer carries the corner-down-right arrow, a queued card leads with the list-end glyph, and a card whose send failed leads with the alert, as a returned one does. * test(native-chat): Resume reports itself in flight until it settles * fix(native-chat): a queue pause shows only while Resume would send something After a Stop whose only remaining card was a returned one, or after a restart with only a card held by its own failed send, the queue published a pause with a Resume that could send nothing: a returned card waits for the user anyway, and a held one for its own Send. The pause is now published, recorded by a Stop, and kept only over a card it can hold back — waiting, with no hold of its own. The fact is retired in the same transaction as the write that removes the last such card, a hold or a refusal included. The publication's dedup also compared only the pause's reason, so a pause appearing or clearing with no readable reason could read as unchanged; it now compares presence first. * fix(native-chat): a queue pause counts only cards Resume would actually send A waiting card behind a returned one is blocked until the user acts on the returned card — the drain never sends past it — so a pause over only such cards still offered a Resume that sent nothing. The rule for "a card Resume would send" is now one function: waiting, no hold of its own, and not behind a returned card. The publication, a Stop's record and the fact's retirement all read it; retirement reads the rows in position order inside the same transaction as the write that took the last such card. * fix(native-chat): a returned card that blocks the paused cards hides the pause but keeps it The last change retired a Stop's pause as soon as a returned card blocked every paused card. Deleting that returned card then sent the cards behind it at once, with no Resume — not what the user asked for. The two rules are now separate. The pause is KEPT (recorded by a Stop, retired in the same transaction as the write that takes the last one) while any waiting card with no hold of its own exists, wherever it sits. It is PUBLISHED only while such a card is not behind a returned one, so the header never offers a Resume that sends nothing. Deleting the blocking card shows the pause again, and the cards behind it wait for Resume or the user's next turn. * test(native-chat): build the pause-only batch through the typed helper * fix(native-chat): a Stop pauses a card its withdrawal sent back even when that settlement was skipped The Stop checked the draft table for a card to pause. When the per-row hook that settles a withdrawn hand-off was skipped, that card was still 'dispatched', so the Stop recorded no pause; the drain later healed it back to waiting and sent it, although the user had pressed Stop. Retirement had the same blind spot and could drop a pause while such a card was owed. What a pause holds back is now one predicate, judged inside the transaction that records or retires it: a waiting card with no hold of its own (one SQL EXISTS), or a dispatched card whose consumed submission was rejected with a settlement back to waiting (read against the journal's submissions). The Stop first runs the owed settlement, as the drain does; if that fails, the owed card still counts, so the pause is recorded rather than skipped. recordPause now checks inside its own transaction and returns whether it recorded, and any draft-table write (and the per-row hook, the consume and the open-time repair) retires a pause that no longer holds anything back. * test(native-chat): pin the per-row hook's pause retirement; skip the judgement when no pause exists The retirement test recorded its second pause over a queue with nothing to hold back, so the recording returned false and the "retired" assertion proved nothing; ablating the per-row hook's retirement passed every test. The hold case now asserts the pause was recorded, and a new test has a delivered echo, through the per-row hook, withdraw the last card a recorded pause holds back. Retirement runs on every appended journal row, so it now checks the pause row by key first and judges nothing when no pause is recorded. Two comments were brought in line with the owed-hand-off rule and rewrapped. * fix(native-chat): the queued area is one bordered box, and the Steer tooltip spaces its shortcut The pause row, when shown, is the box's first row and each card a row below it, divided rather than individually bordered. The Steer tooltip groups its hint and the shortcut chips with the house gap, so the chips no longer touch the text. * test(native-chat): match main's append and dispatch shapes in the queue tests * fix(native-chat): read a compaction's settled submission through the send-result union * test(native-chat): a queued card Steered into a turn joins it under the opener's bar A Steer hands the draft over under a fresh submission id, linked by queuedMessageId, and the host scopes its row to the running turn; the transcript keeps it inside that turn with no bar of its own, settled or running. * fix(native-chat): queued messages sit above running shells and agents, which stay next to the composer --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
9b52661f92 |
fix(claude): a Stop naming the running turn acts at once, and a Stop in the gap is no longer lost (#23861)
* fix(claude): stop a named running turn at once when a follow-up is queued A Stop that named the running Claude turn waited up to 3 s whenever a later send's handover was still unresolved, for example a follow-up Claude had queued. The phone app and older desktop clients always name the turn. Claude's interrupt is session-scoped, so a Stop naming the live turn now takes the same path as a Stop naming no turn: it interrupts at once, and the queue sweep settles the follow-up as withdrawn. A Stop naming a turn that is no longer live is still refused. With that case gone, the admission wait could only delay a refusal, so it is deleted rather than left as a second gate. * fix(claude): a Stop naming a turn that just ended withdraws the follow-up behind it The phone names the turn it shows, and its copy of the chat can trail the host's. When that turn ends just before its Stop lands, with a follow-up already written to Claude but not yet started, the Stop was refused and the follow-up ran anyway. With nothing live, a written follow-up whose turn has not opened is work the naming client has not seen start, so the Stop takes the conversation path: it interrupts and Claude's queued follow-up settles as withdrawn. A Stop naming an older turn while a newer one runs is still refused. * fix(native-chat): a named Stop that withdrew a queued message says nothing finished When a Stop named a turn that had already ended, the host still withdrew the message the user had sent after it, yet the Stop reported nothing cancelled and the chat read "The provider had already finished this turn." That withdrawal is the Stop taking effect, so it now reports cancelled and writes no row, as a Stop naming no turn already does in the same case. * fix(native-chat): a named Stop the provider left unconfirmed is not reported as a success A named Stop that withdrew a host-queued message was reported cancelled, with no row, whatever the provider answered. When the provider took the interrupt without confirming it (or refused it), the turn may still be running, so the Stop now keeps the provider's answer instead of claiming success. * fix(native-chat): judge a named Stop's withdrawal by what the journal still runs A named Stop that withdrew a host-queued message was reported as a success unless the provider's answer carried a refusal or was unconfirmed. Providers report a refusal differently: Claude never sets one, so a Stop naming an older turn while a newer one ran came back cancelled with no row; Codex refuses any turn that has ended, so the ended-turn case still wrote "already finished". The success is now decided by the rule the no-turn Stop uses: nothing still reads working on the journal. * refactor(native-chat): one rule for what every conversation Stop reports A conversation Stop that withdrew what was queued and left nothing working on the journal is a success, whether or not it named a turn. The rule was scoped to named Stops, so a Stop naming no turn whose turn ended under it reported that the agent had no turn to stop after it withdrew a message. An unconfirmed interrupt is now reported as unconfirmed before the named-turn row, which said the provider had already finished a turn it had just taken an interrupt for. * fix(native-chat): judge what a Stop left working once its streamed rows land Claude's send echo accepts the send one sink write before the turn row it opens lands, so a Stop that withdrew a queued message could read the journal idle in that gap and report success with no row while the turn then ran. Drain the streamed events before the read, as the prompt Stop already does. * fix(native-chat): a Stop naming no turn reads what is working once its streamed rows land The no-turn Stop decides whether to interrupt from the journal. Claude's send echo accepts the send one sink write before its turn row lands, so in that gap the journal read idle and the Stop skipped the interrupt while the turn ran. Drain the streamed events first, through the same read the Stop's report uses. A failed drain reads working, so bookkeeping never keeps a Stop from interrupting. * test(native-chat): scope the Stop tests' journal rows to the thread after the main merge * test(native-chat): open the Stop-report journal on the host database after the main merge |