mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 16:02:29 +00:00
536b433bc718dbde4c92bd8ff944ee9089b0a5df
1020
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ddefd523e0 |
Keep selected text navigation from rewriting document links (#25175)
* Keep selected text navigation from rewriting document links * Check model selection before document link arrow coverage |
||
|
|
cbe64383dc |
Reuse source-line calculations for Markdown review selections (#25173)
* Reuse source-line calculations for Markdown review selections * Use typed editor probes in review selection performance coverage |
||
|
|
41cc77509f | Keep active notebook cells current after external reloads (#25172) | ||
|
|
8267b578b2 |
Fix Markdown Find editing without moving the caret or viewport (#25144)
* Fix Markdown Find editing without changing the caret or viewport * Preserve Markdown selections across search focus and stale updates |
||
|
|
9207aef01d |
Add an optional shortcut to toggle child workspaces (#24165)
Adds a user-assignable shortcut for the existing child-workspace chip action. It stays unassigned by default on macOS, Linux and Windows. Final-source rendered checks cover Settings recording/reset, guards, scroll preservation and restart. Fixes #24163 Continues mmarabel’s original contribution in this PR. Issue #24163 has no sibling implementation PR. Existing bot findings are fixed, withdrawn or addressed in the PR review. Co-authored-by: mmarabel <166927047+mmarabel@users.noreply.github.com> |
||
|
|
1dc6157614 |
ci: run the cross-version tests when the chat send builders or the RPC request path change (#25061)
* ci: run cross-version wire suites when agent sends or orchestration change The selector skipped the cross-version wire job for #24901, which changed the shared agent-send path, the structured send envelope builders, and orchestration RPC code. Route the send payload/fingerprint builders, src/main/runtime/orchestration/, and the orchestration RPC methods to the job, and pin each rule in a test. * ci: run cross-version wire suites only for code they execute The agent-session suites now build their send through the shared outbox builder and check it against the release host's schema and fingerprint, so the send builder and outbox are selected exactly. Load-only orchestration rules are dropped; the dispatcher-path files every suite request runs through are selected instead. * ci: run the release's own send admission, cover queued sends and the RPC reply builder * test(cross-version): a wrong send fingerprint must be refused, so an agreed one is not vacuous |
||
|
|
136b990a05 |
fix(search): negotiate supported agents across mixed host versions (#25009)
Keep older request and reply parsers usable while current peers retain all supported history. Co-authored-by: nwparker <nwparker@users.noreply.github.com> |
||
|
|
62451920ed |
fix(git): preserve SSH review context and worktree ownership (#24945)
* fix(git): preserve SSH arguments and guard background writes * test(terminal): settle fish fixture startup readiness * fix(git): preserve bare UNC SSH paths * fix(git): unescape shell operators in Windows SSH paths * test(runtime): settle removal writes before fixture cleanup * test(git): skip optional OpenSSH probe when unavailable * perf(git): skip equal-tip reads and bound relay discovery * test(processes): ratchet the removed relay Git spawn * test(shells): wait for initial zsh output before sending input * Fix SSH review context and recover Git maintenance cleanup safely * Keep relay Git compatibility fixtures outside shared client projects * Fence superseded maintenance and preserve mixed-version session search * test: model Git child termination and search catalogs * test: retain catalog authority over history flags |
||
|
|
aca2d51e0e |
fix(jcode): harden Windows hooks and negotiate remote history (#24998)
Redirect the managed Windows payload file into curl instead of starting pipeline shells, register native Windows delivery coverage, and document Jcode v0.89.0+ as the upstream launcher requirement for invisible hooks. Negotiate Jcode history in both directions with mixed-version Orca hosts, preserving supported search filters and old-client response compatibility. Co-authored-by: czzczz <chanzrz_zbf@foxmail.com> Co-authored-by: JianJia2018 <39438074+JianJia2018@users.noreply.github.com> |
||
|
|
8dc16a7df0 |
feat(native-chat): open structured chats on the paired Orca server that owns the workspace (#24205)
* fix(native-chat): a host admits structured sessions by client capability, not its own chat setting A host's experimentalStructuredNativeChat decided whether any paired client could reach agentSession.* at all, and whether session.tabs.* showed it structured tabs. That setting is the host user's own launch preference: whether a new agent opens as a chat or a terminal is decided by whoever launches it. Using it as admission control meant a client whose own preference was "structured chat" was refused on a host whose preference was "terminal", and chats opened while the setting was on were withheld from mobile once it was turned off. The gate now asks one thing: did the client advertise agent-session.structured.v1 (in-process callers negotiate nothing and are always admitted). Tab projection and restore follow the same rule. With the setting no longer gating anything, the separate cleanup gate (close, cancel, unsubscribe, release), which existed only so those kept working after the setting was switched off, is identical to the main gate and is folded into it. The settings listener that republished tabs when the setting changed is removed, since projection no longer depends on it. The host setting still picks the default for launches that start on the host itself (agent.launch from mobile, orchestration worker-start). * fix(native-chat): the desktop declares structured chat support to paired hosts The desktop renderer advertised agent-session.structured.v1 (and the Claude, turn-item and background-task capabilities that go with it) to its own main process but not to a paired Orca server. The server therefore refused every agentSession.* call from the desktop and stripped structured chat tabs out of the tab list it published to it, so a structured chat running on a paired server never appeared on the desktop, even though the renderer already mirrors a host's agent-session tabs and drives each one against the server that owns its workspace. The same renderer reads structured chats on either host, so the remote Electron list now carries the same structured-session capabilities as the local one, and the capability test pins that nothing is advertised only locally. * feat(native-chat): open structured chats on the paired server that owns the workspace With the structured-chat default on, an agent launched in a workspace that lives on a paired Orca server always opened as a terminal (or the terminal-backed chat view). Three things kept it off the structured path: the launch check refused every host but this machine, a remote workspace was handed to the host-published terminal path before the structured route was even considered, and the structured launch pipeline sent create and every follow-up call to this machine's runtime. A workspace's owning runtime is fixed, so the pipeline now derives it from the workspace instead of assuming this machine (structured-agent-session-owner.ts, the same derivation the chat pane already uses to read a session). The launch intent carries that target; the pre-create support check, create, the publication check and fence read, the launch prompt send, held option picks, the "focus this chat" marker, the placeholder tab's host, and tab close/purge all use it. The launch check now accepts a paired server and asks that server's own capabilities (read from the status the client already cached for it) rather than this machine's. An SSH workspace stays terminal-backed: no Orca runtime runs there. The host still answers createSupport before anything is created, so an older server that refuses shows the failure in the chat tab. A chat the user closed before its create landed is now also retired on the paired server when it publishes, as the local sync already does. Orchestration workers placed on another runtime are unchanged: federation creates terminal agents only. * fix(native-chat): negotiate client-chosen launch mode so released phones and old servers keep terminals Hosts advertise agent-session.structured.client-launch-mode.v1: they admit structured sessions by client capability alone. A remote client that does not advertise it (phones released before agent.launch) asks createSupport to pick the launch mode, so the host keeps answering that with its own setting, exactly as before. Cleanup methods keep their own named gate so a future admission condition cannot make close or cancel refusable. * refactor(runtime): keep the Electron client capability list in its own module protocol-version.ts is at its line budget; the list is what the desktop advertises to paired hosts, not the host's own contract. * fix(native-chat): the desktop declares it picks each launch mode itself Paired hosts and the desktop's own main process then answer createSupport by the workspace rather than by their own chat setting. * fix(native-chat): pin each structured chat to the host it was launched on - Route: a paired server opens a chat only when it advertises the client-chosen launch mode; an older server keeps its terminal. Its capabilities come from the store's host status, not the compatibility cache that is empty after boot or reconnect. - A launch command override is this machine's: the route applies it only locally, and a host's createSupport refuses on its own override. - The owning host is resolved once, from the same value the route used, and carried on the launch intent, its persisted record (legacy records load as local), the provisional tab and every mirrored chat tab. Close, purge, retry and reload read it instead of re-deriving it from a worktree id two hosts can share; an owner that cannot be named refuses. - Cancellation tombstones record their host: only that host's authoritative inventory retires one, restored cleanup closes it there, and a paired host's tombstone expires after 30 days if it never answers. - A paired server's frame settles launches it published, as the local inventory already does for this machine. * fix(native-chat): a paired server that declines a chat opens its terminal instead createSupport only reads, so both of its non-answers are settled before anything is created: - A paired server that answers it cannot run the chat (a WSL repo, a Claude account mismatch, its own launch command override) closes the chat tab and opens the terminal the route would have chosen, with a notice saying why. This machine's own decline stays a failed chat. - A host that could not be asked closes the chat tab and leaves one failure toast, instead of a lingering "could not confirm" chat. * test(native-chat): a provisional chat carries its launch's host and hands pre-create failures on * chore(native-chat): justify the two type assertions this change's lines touch * test(native-chat): state why each staged test fixture is cast * fix(native-chat): chats that already exist keep showing whatever the chat setting says The structured chat setting decides only what new agents open as. With it off, this machine's structured chats used to be hidden while the host, which no longer reads the setting, still reported them to the workspace activation gate, so a workspace holding only a chat opened empty. The local chat mirror and its startup restore now run whatever the setting says, the continue-after-restart offer follows the chats that exist, and the setting's copy says it applies to new agents. * fix(native-chat): the browser client keeps its host terminal on paired servers A browser client whose own preferences turn structured chat on took the structured route for every paired-server workspace, but its handshake never says it reads structured sessions, so the server refused the chat and the user got a failed chat tab where a host terminal used to open. The route for a paired host now also asks what this client advertises to it: the desktop's list does, the browser client's does not. Its handshake list is now a named constant the route reads, so the two cannot drift. The chat setting's copy now says it runs on paired Orca servers too; WSL and SSH hosts still use terminal chat. * fix(native-chat): a retried launch a paired server declines opens its terminal too A launch restored after a reload settles only through its Retry, so a declining paired server left a failed chat there while a first launch got the server's terminal and a notice. The chat's Retry now hands the same pre-create failures to the same replacement, carrying the prompt the launch had staged. * refactor(native-chat): a paired host's cancelled-chat record ends on its 30-day TTL The paired census re-read a host's whole inventory after every authoritative frame to retire tombstones, and a tombstone restored after a reload needed a second such frame, so in practice it retired nothing. A tombstone guards a random session id and is inert once stale; the chat is already closed on its host whenever a frame shows it. The census, its trigger in the mirror layer and its cleanup are removed; the owner-scoped tombstones, close-on-sight, the TTL and publication marking from frames stay. * fix(native-chat): a chat's pane and status read from the host recorded on its tab The chat pane and its sidebar status still derived the host from the workspace id, which two hosts can share; a paired chat in a non-active same-id workspace was read from this machine. Both now read the owner stamped on the tab, as close, purge and publication already do. * fix(native-chat): "Resume in chat" follows the terminal resume's host rule Agent Session History offered "Resume in chat" for a conversation recorded on this machine into a paired server's workspace, where its transcript does not exist. A chat now resumes a conversation only on the host that recorded it, as the terminal resume does, and that host is the one asked whether it can resume history. * fix(native-chat): the chat setting says older paired servers keep terminal chat * test(native-chat): pin that a host advertises the client-chosen launch mode * fix(native-chat): mirror this machine's chats only where it holds them Round 1 ran the local chat mirror for everyone so existing chats show whatever the setting says. That gave every desktop a permanent session-tabs listener, which turns on the runtime's phone replication paths, plus two full session-tab censuses at startup, and made the browser client mirror its remote host a second time. The runtime now says whether it holds structured chats: its structured host is built only when saved chats were restored at startup or a client created one here, and it announces the moment one is built. The mirror, the startup restore and the continue-after-restart offer run only when the setting launches chats or the host holds some, and never in the browser client. A chat a paired client creates here with the setting off still appears at once. The chat behaviour settings show wherever chats exist, and the setting's copy says it picks what new agents open as. The toggle-off teardown this made dead is removed. * test(native-chat): route a paired-server launch over the capability lists both sides really advertise * test(native-chat): record install listeners without a cast * fix(native-chat): a paired server admits a chat before any of it exists here The desktop opened a paired server's chat tab, launch record, queued prompt and focus intent before asking the server, so a "no" needed a replacement that undid and redid all of it, and every piece it missed was a bug: the workspace deselected, the caller told "failed" while a terminal ran its prompt, the caller's arguments and other queued prompts lost, and a create whose reply was lost treated as never sent. A paired launch now asks the server first and commits nothing until it answers. Admitted opens the chat as before. Declined runs the caller's own launch as the server's terminal, with the existing notice (a resume fails instead, having no terminal equivalent). Unreachable opens nothing and names the server in one toast. The new-tab launcher reports the host's surface for paired workspaces, as it did before paired chats, with the prompt delivery of whichever surface got the prompt. The replacement and its error classes are gone, and the probe inside a launch is back to its old meaning: a "no" is a failed chat with Retry, and no answer leaves "Could not confirm" with Retry and the prompt kept, here as on this machine. * fix(native-chat): mirror this machine's chats only once it holds one, not once its host is built Session history, resume preparation, terminal resume commands and replay-safe phone launches all build the structured host for users who never had a chat, which turned on the chat mirror and the structured-only settings rows until the next restart. The signal is now derived from the host's records (or a records file still owed its import) and pushed when the first chat is restored or created. A throwing listener no longer fails the install that fired it. * fix(native-chat): a fork's reveal never seeds a terminal beside the surface the launcher opens Forking into a paired-server workspace revealed it as if nothing would open there, so the reveal created a blank host terminal beside the forked chat (and beside a forked agent terminal on main). The launcher always opens the fork's surface itself, so the reveal now says so for every surface, as the fix-checks launch already does. * fix(native-chat): a declined direct launch keeps the caller's CLI args; an unreachable resume toasts once When a paired server declines a "Fix checks" chat in a new workspace, the terminal that opens instead now carries the recipe's saved CLI arguments, launch platform and launch source, as the terminal route did. "Resume in chat" to a server that cannot be reached showed the admission's "Could not reach" toast and the vault's generic one; the admission marks its failure notified and the vault adds nothing. * fix(native-chat): a declined background create opens its terminal without switching workspaces Since #23974 a worktree create the user moved away from must not pull them onto the new workspace. When a paired server declined that create's chat, the fallback terminal opened as a new agent tab, whose host create selects the workspace. The create now opens its own agent terminal the way main's background branch does: in place from the request's startup plan (so its CLI args carry), without selecting the workspace. A create the user is still watching keeps the new-tab fallback. * test(native-chat): name the launch's host in main's new outbox fence test Main's new staging-failure test calls settleStructuredAgentLaunchPrompt without the target this PR made required; it is a local launch, as in the sibling tests. * fix(native-chat): a paired server's new chat shows no model until the server reports the one it started A chat on a paired server starts with the server's saved model and options, but the picker showed this desktop's saved selection (or the catalog default) until the server reported a model, and a pick made in that window was remembered on the server under that guessed model. A paired launch now carries no desktop seed, and until the server reports its model the picker names no model and takes no picks. Local chats are unchanged. * test(native-chat): seed the paired repo without a cast The repo literal already satisfies Repo, so the changed-lines cast gate has nothing to excuse. * feat(native-chat): createSupport reports the saved selection a new chat on this host starts with A chat on a paired server starts with the server's saved model and options, which the desktop could not read, so its picker showed a guess. createSupport's answer, which the desktop already waits for before a paired launch, now also carries that seed as a new optional field (older clients ignore it). Create and createSupport read it through one resolver so they cannot drift. * fix(native-chat): a paired server's new chat shows the selection the server will start it with The paired server now names its saved model and options in the admission answer the desktop already waits for. That seed goes into the launch intent and its persisted record, so the picker shows the server's model at once, stays pickable like a local chat, and remembers picks on the server under that model; a reload shows the same. The locked picker remains only for a server too old to name a seed. Also moves host admission and launch-outcome tracking into their own modules: the latest main merge left structured-agent-session-launch.ts over the max-lines limit. * test(native-chat): expect the launch intent's new seed argument in exact-call assertions * refactor(protocol): move the Electron remote client capability list into its own module Merging main left protocol-version.ts one line over the max-lines limit on this branch. The list of capabilities the desktop advertises to a paired host moves, unchanged, into electron-remote-runtime-client-capabilities.ts, the module the next PR in the stack already uses for it; importers point there. * fix(native-chat): a paired chat with no saved server model is pickable; Retry shows the server's current seed A server whose user never saved a chat model sends no seed, and the desktop showed a locked, model-only picker for it, although that is the common case: no server that can admit a paired chat predates the seed field. Such a chat now behaves like a local chat with no saved model: the CLI default, pickable. The lock and its snapshot helper are gone. Retry kept the first admission's seed while the create probe, which already runs on every attempt, reported the server's current one and dropped it. The probe's seed now replaces a paired launch's seed and the picker's, so a retried chat shows what its create will run. * test(cross-version): stub the launch seed resolver createSupport now reads * test(protocol): pin the desktop capability divergence against what a paired server receives Every paired transport sends the shared remote base plus the Electron list, so the divergence test now compares that union with the renderer's local list instead of the declared Electron list. A capability added only to the shared base can no longer slip past it. The two base-only capabilities it surfaced are recorded: skills.install-result.v2 has no local caller; the authoritative-inventory label is read by the local tabs sync but dropped by main, and is marked unsettled. The turn-item and both background-task-stop capabilities were already sent through the shared base, so the Electron list no longer repeats them. The wire set is unchanged; this PR's real change on the wire is structured.v1, the Claude structured capability and the client launch-mode capability. * fix(native-chat): the desktop tells its own host it picks each launch mode, so retrying an existing chat works with the setting off * docs(native-chat): name the real exit for the released-phone createSupport rule * fix(native-chat): the route reads the capabilities a paired host actually receives The renderer decided whether a paired host would admit a chat from the desktop's Electron list, but every desktop transport sends that list plus the shared remote base. They agreed only because the route's checks happened to sit in both. The route input is now built with the same remoteRuntimeClientCapabilities the transports use (the browser client already sends its list as is), and a test pins each against the real handshake. * test(cross-version): a released client still gets the host-setting createSupport answer; a launch-mode client gets supported plus the seed * test(native-chat): let main's child-records test resolve each chat's owner Main's new test mocks worktree-runtime-owner with only the runtime environment id, but the status projection in this PR also resolves each structured chat's owner from the worktree. The mock keeps the module's real exports and overrides only what the test pins. * fix(deps): take #24204's lockfile that the merge reverted * test(native-chat): let the Codex child-approval e2e unit test resolve each chat's owner Its worktree-runtime-owner mock exported only the runtime environment id, but this PR's status projection also resolves each structured chat's owner from the worktree. The mock now keeps the module's real exports and overrides only that id, as structured-child-records-switch does. * test(native-chat): move the close-race launch cases into their own file Merging main added launch tests on both sides and took structured-agent-session-launch.test.ts past the 800-line limit. The three cases where a tab close races a launch move to structured-agent-session-launch-close-race.test.ts, with the same setup the other split launch suites copy. |
||
|
|
a5f28f265c |
Fix word wrap for both panes in side-by-side diffs
Forward wrapping to both diff panes through the existing editor option path and clean up listeners. Co-authored-by: Wooseong Kim <innocarpe@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
7b8498b406 |
feat(sidebar): include folder workspaces in keyboard navigation
Use the rendered sidebar row order and host identity when cycling through folder and Git workspaces. Related: https://github.com/stablyai/orca/pull/10555 Co-authored-by: Neil <neil@stably.ai> Co-authored-by: JeongUk Park <jeongph.dev@gmail.com> |
||
|
|
130b2ef425 |
feat(documents): open CSV and TSV files from the OS
Extend existing OS document associations and delivery to CSV/TSV, preserving restoration and authorization. Co-authored-by: Neil <neil@stably.ai> |
||
|
|
b332742b89 |
Speed up large Markdown Find and render oversized tables (#24948)
* Speed up large Markdown Find and render oversized tables * Poll table preview geometry outside hidden renderers * Preserve Markdown navigation through refresh and tab restoration * Confirm Markdown restoration when refreshed content is ready * Trace table refresh positions and update Unicode search reference * Recognize queued measurement scrolls before restoring Markdown anchors * Rebuild Markdown Find ranges after renderer components change |
||
|
|
843607b1bc |
Register supervised Qoder China and Qwen Code (#24616)
* Add Qoder session history and search with real CLI coverage * Allow the real Qoder marker file to end with a newline * Keep Qoder tool output out of history previews and search * Keep Qoder search pages readable by older clients * Verify persisted Qoder history after a real generated and resumed task * Negotiate Qoder filters before searching an older execution host * Combine search client imports for the CI plugin gate * Keep the relay search oracle aligned with legacy agent filtering * Register supervised Qoder China and Qwen lifecycle integration * Cover Qoder China mobile assets and mixed-host resume gates * Verify Qoder provider tags against the older released wire parser * Verify China and Qwen keep independent Windows hook scripts * Verify Qoder registrations against the installed older Windows release * test(qoder): align search capability contracts and pin old-host fencing * fix(qoder): rank exact picker identities and command aliases first * test(qoder): preserve the regional CLI shared icon expectation Keep the full bundled-asset and no-remote-image checks, with an explicit shared-logo basename for Qoder China. The map also works with older catalog type unions. * fix(qoder): align China catalog entry with fallback order --------- Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example> |
||
|
|
9a1bef48e4 |
fix(cursor): resume exact conversation after startup status replay (#24670)
* fix(cursor): restore exact conversation after startup snapshot * fix(terminal): preserve ready snapshot reattach and isolate bridge fixtures * test(cursor): seed the real startup bridge after module resets * test(terminal): settle startup snapshots in remote restore fixtures Signed-off-by: Neil <neil@stably.ai> --------- Signed-off-by: Neil <neil@stably.ai> |
||
|
|
2c2dfd028a |
test: run committed OpenCode redraw capture replay (#24811)
Port the synthetic fixture and replay coverage from Neil/nwparker original PR #19195 ( |
||
|
|
f97ca2a49d |
Add Qoder session history and search (#24614)
* Add Qoder session history and search with real CLI coverage * Allow the real Qoder marker file to end with a newline * Keep Qoder tool output out of history previews and search * Keep Qoder search pages readable by older clients * Verify persisted Qoder history after a real generated and resumed task * Negotiate Qoder filters before searching an older execution host * Combine search client imports for the CI plugin gate * Keep the relay search oracle aligned with legacy agent filtering * test(qoder): align search capability contracts and pin old-host fencing |
||
|
|
533446dde6 |
Stop mocked renderer imports from qualifying headless CI (#24902)
* Decouple headless running-work tests from the renderer * Keep the shared running-work probe contract documented |
||
|
|
3fba1952c8 |
perf(tab-bar): a change to one tab no longer re-renders every tab (#24261)
With many tabs open, a change to any one tab (a retitle, an agent finishing, a tab switch, a git status write, or a browser tab update on SSH and web clients) re-rendered every tab in the strip, so the strip stuttered. Each tab is now a memoized row that re-renders only when its own values change, with stable handlers, a stable drag id list and stable drag sensor options. Editor tabs get their own git status, and mirrored browser tabs keep their page-id list while the ids don't change. Part of #24241: opening, closing or reordering a tab still re-renders every tab once. |
||
|
|
1aa0860f7e |
Keep large Markdown previews responsive (#24880)
* Keep large Markdown previews responsive * Fix large preview review navigation and Find budgets * Initialize preview scroll caches once and check viewport visibility * Restore large previews after loaded rows are measured * Refresh loaded Markdown rows after viewport changes * Keep Markdown revisions visible and reuse bounded search text |
||
|
|
b4ff7fe6d4 |
fix(tests): run watcher crash harness against current code (#24705)
* fix(tests): resolve watcher crash harness from repository root * fix(tests): rebuild watcher crash harness from current sources * fix(tests): register watcher interruption callback as an owner hook |
||
|
|
f2257ffa69 | fix: dismiss Codex account prompt and return focus to terminal (#24683) | ||
|
|
de8bffe240 |
Fix terminal width cutoff on wide panes (#24687)
* fix(terminal): let wide panes use up to 1024 columns Adapt the wider viewport limit proposed in #16578 to the current runtime, shared RPC schemas, and preview sizing. Co-authored-by: innocarpe <innocarpe@users.noreply.github.com> * test(terminal): wait for probe output after command echo * test(terminal): align RPC boundary with wider viewport limit --------- Co-authored-by: innocarpe <innocarpe@users.noreply.github.com> |
||
|
|
0f167ac659 | test: check plugin fixture worktree cleanup (#24779) | ||
|
|
adf447d958 | test: classify acknowledged remount input as driving (#24750) | ||
|
|
99b2628aef | test: update sidebar setup and remove obsolete permission sentinel (#24734) | ||
|
|
d9fbb4eecf | Keep SSH typing replies inside narrow split terminal panes (#24682) | ||
|
|
66799f7e8f |
Keep paired browser terminal insertion in the host's requested position (#24676)
* test: align source-control fixtures with current store contracts * Bound E2E package setup and retain cancelled-job traces * Remove empty passing sentinels from opt-in socket tests * Make SSH typing pressure fixture readiness and replies observable * Advertise browser support for anchored terminal placement |
||
|
|
b49abdb1f4 |
fix: recover renderer launch failures in the running app (#24250)
* fix(recovery): back off a launch-failed renderer instead of tripping the crash breaker
A renderer that the OS refused to spawn (macOS exit 1003 = LAUNCH_RESULT_FAILURE; field
cause: per-user process limit, posix_spawn EAGAIN) burned the 3-reload crash-loop budget
in ~750ms and raised a "graphics driver" prompt, while the condition lasted minutes.
- launch-failed retries in place on a 250ms..60s backoff (~2 min), outside the breaker;
a loaded document resets it. Other crash reasons keep the breaker.
- Each launch failure records renderer_launch_failed_probe {spawnError} from a cheap
spawn probe, so bundles name EAGAIN/EACCES/ENOENT directly.
- The exhausted prompt says the process limit was hit (probe EAGAIN), drops the
graphics-driver wording, keeps Try Again as default, and offers no Restart:
app.relaunch also needs a free process slot and silently fails without one.
* fix(recovery): skip the launch probe on Windows and probe the prompt once
- Re-check quitting after the prompt's probe; don't re-probe on Copy Commands.
- recordRendererLaunchFailureProbe never rejects (breadcrumb write guarded).
- Windows: no spawn probe; a child per failed launch is the per-operation burst EDR scores.
* test: cover quitting and duplicate renderer launch failures
* test: use typed access in PTY delay regression fixture
* fix: scope extended launch retries to POSIX hosts
* test: cover launch probe behavior on native Windows
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
|
||
|
|
53930a161b |
Keep SSH typing replies visible during background pressure (#24629)
* test: align source-control fixtures with current store contracts * Bound E2E package setup and retain cancelled-job traces * Remove empty passing sentinels from opt-in socket tests * Make SSH typing pressure fixture readiness and replies observable |
||
|
|
306b4578aa |
Enable Option shortcuts for ABC keyboards in Auto mode (#24528)
* Clarify Option shortcut settings and cover punctuation input * Enable Auto Option shortcuts on ABC keyboards safely |
||
|
|
3f37fcc423 | test: align source-control fixtures with current store contracts (#24571) | ||
|
|
ba9af21d75 | test: classify terminal driver input with the current PTY contract (#24560) | ||
|
|
e2c5414f76 |
fix(native-chat): an older Orca keeps a chat with a newer row kind read-only instead of deleting the rest of its history (#24477)
* fix(native-chat): an older Orca skips and keeps a journal row of a kind it does not know
* test(native-chat): a newer build's journal row kind survives reads, writes, rewinds and reopens
* fix(native-chat): an older Orca keeps an unknown journal row kind read-only unless its writer declared it skippable
A row of a kind this build does not know, in a well-formed envelope, now latches the chat
read-only with every row kept, the same way a newer row version does. It is read past only
when its writer declared `ifUnknown` on the row: `skip` (a rewind drops it) or `carry` (a
rewind carries it after the rebuilt history, epoch, seq and fence restamped). Every existing
kind changes queue or turn state, so skipping by default would let an older build write from
a wrong fold.
- journal-row-kind-compatibility.ts: each kind states how older builds read it, typed over
every row kind, so a new kind cannot be added without a declaration.
- Rewind restates the Resume and Stop as before, then carries `carry` rows in source order;
the restatement goes back to { lifted, liveStop }.
- Replay treats a row whose body names another sequence than its stored key as malformed at
the key, so the next write never collides with it; catch-up reads stop there too.
* refactor(native-chat): drop the writer opt-in; an unknown journal row kind only latches read-only
An older Orca now treats a row of a kind it does not know exactly like a row from a newer
schema version: every row stays on disk and the chat opens read-only until an update. The
writer-declared skip/carry opt-in, its in-memory placeholder, the carry through rewinds and
the per-kind registry are removed: no current or planned kind could use them, and they can
come with the first kind that may safely be read past.
Kept: an unknown kind needs the envelope every row keeps (epoch, sequence, fence, timestamp),
else it is damage as before; a row whose body names another sequence than its stored key is
malformed at the key; the epoch row's validator names its kind. The schema header states the
rule for adding a kind: keep the envelope, and either ship the reader first or bump `v`.
* refactor(native-chat): derive the journal's known row kinds from the row union
Each kind's own-field check now lives in one table keyed by every kind JournalRow holds, and
the set of kinds this build knows is derived from that table. A kind added to the union without
a check fails to compile, rather than latching this build's own chats read-only as a newer
build's kind. A test reads one valid row of every kind.
|
||
|
|
c9a9b8d109 | test: restore delayed PTY writes in large-paste coverage (#24556) | ||
|
|
b666d07117 | test: update worktree setup and enforce cleanup results (#24552) | ||
|
|
1fbfb13e0f | test: isolate seeded Git repositories per Playwright worker (#24550) | ||
|
|
0b7b9a9af5 | test: isolate session fixtures and wait for completed indexing (#24544) | ||
|
|
026b8378a4 | test(wire): make release compatibility probes deterministic (#24538) | ||
|
|
11b1c8f353 | test(e2e): dismiss the browser tour before starting screenshot markup (#24534) | ||
|
|
444f1952c7 |
ci: run every cross-version wire test, picked up by folder so new ones can't be skipped (#24499)
* ci(cross-version-wire): run the whole directory so no compatibility test is left out Three cross-version tests ran in no CI job because the job named its files by hand. Run the directory instead, ratchet that every file kept out of the unit shards runs in some PR job, and re-run the job when the modules the newly running tests guard change. * test(cross-version): give the orchestration downgrade test its siblings' 120 s budget * ci(unit-exclusion): count only merge-gating jobs, and require each excluded file's job to fire on it The coverage check counted any pr.yml job, including e2e, terminal IME and Windows WSL, which are left out of verify.needs and so cannot block a merge. It now reads verify.needs and the reusable workflows those jobs call. It also only proved that some step names each excluded file, not that the job runs when the file changes. The structured-session zsh login-shell test runs only in shell_contracts, whose path trigger matched neither it, its harness nor its subject, so a PR touching only those ran it nowhere. The check now asserts a change to each excluded file fires a gating job that names it, and the shell trigger gains those three paths. * ci(cross-version-wire): trigger on the turn-outcome vocabulary and the schema version-skew resolver A change confined to src/shared/agent-turn-outcome (the arms a newer host publishes) or to orchestration-schema-version-skew (how current code reopens a downgraded database) skipped the job whose tests guard exactly those contracts. Also corrects the publish/read direction in the turn-end comment. * test(cross-version): state why the orchestration downgrade test needs 120 s * test(ci): glob the unit tree once for the unit-exclusion coverage checks |
||
|
|
757736628f |
fix(native-chat): a paired server admits structured chat by client capability, not its own chat setting (#24203)
* fix(native-chat): a host admits structured sessions by client capability, not its own chat setting A host's experimentalStructuredNativeChat decided whether any paired client could reach agentSession.* at all, and whether session.tabs.* showed it structured tabs. That setting is the host user's own launch preference: whether a new agent opens as a chat or a terminal is decided by whoever launches it. Using it as admission control meant a client whose own preference was "structured chat" was refused on a host whose preference was "terminal", and chats opened while the setting was on were withheld from mobile once it was turned off. The gate now asks one thing: did the client advertise agent-session.structured.v1 (in-process callers negotiate nothing and are always admitted). Tab projection and restore follow the same rule. With the setting no longer gating anything, the separate cleanup gate (close, cancel, unsubscribe, release), which existed only so those kept working after the setting was switched off, is identical to the main gate and is folded into it. The settings listener that republished tabs when the setting changed is removed, since projection no longer depends on it. The host setting still picks the default for launches that start on the host itself (agent.launch from mobile, orchestration worker-start). * fix(native-chat): negotiate client-chosen launch mode so released phones and old servers keep terminals Hosts advertise agent-session.structured.client-launch-mode.v1: they admit structured sessions by client capability alone. A remote client that does not advertise it (phones released before agent.launch) asks createSupport to pick the launch mode, so the host keeps answering that with its own setting, exactly as before. Cleanup methods keep their own named gate so a future admission condition cannot make close or cancel refusable. * chore(native-chat): justify the two type assertions this change's lines touch * fix(native-chat): chats that already exist keep showing whatever the chat setting says The structured chat setting decides only what new agents open as. With it off, this machine's structured chats used to be hidden while the host, which no longer reads the setting, still reported them to the workspace activation gate, so a workspace holding only a chat opened empty. The local chat mirror and its startup restore now run whatever the setting says, the continue-after-restart offer follows the chats that exist, and the setting's copy says it applies to new agents. * test(native-chat): pin that a host advertises the client-chosen launch mode * fix(native-chat): mirror this machine's chats only where it holds them Round 1 ran the local chat mirror for everyone so existing chats show whatever the setting says. That gave every desktop a permanent session-tabs listener, which turns on the runtime's phone replication paths, plus two full session-tab censuses at startup, and made the browser client mirror its remote host a second time. The runtime now says whether it holds structured chats: its structured host is built only when saved chats were restored at startup or a client created one here, and it announces the moment one is built. The mirror, the startup restore and the continue-after-restart offer run only when the setting launches chats or the host holds some, and never in the browser client. A chat a paired client creates here with the setting off still appears at once. The chat behaviour settings show wherever chats exist, and the setting's copy says it picks what new agents open as. The toggle-off teardown this made dead is removed. * test(native-chat): record install listeners without a cast * fix(native-chat): mirror this machine's chats only once it holds one, not once its host is built Session history, resume preparation, terminal resume commands and replay-safe phone launches all build the structured host for users who never had a chat, which turned on the chat mirror and the structured-only settings rows until the next restart. The signal is now derived from the host's records (or a records file still owed its import) and pushed when the first chat is restored or created. A throwing listener no longer fails the install that fired it. * feat(native-chat): createSupport reports the saved selection a new chat on this host starts with A chat on a paired server starts with the server's saved model and options, which the desktop could not read, so its picker showed a guess. createSupport's answer, which the desktop already waits for before a paired launch, now also carries that seed as a new optional field (older clients ignore it). Create and createSupport read it through one resolver so they cannot drift. * refactor(protocol): move the Electron remote client capability list into its own module Merging main left protocol-version.ts one line over the max-lines limit on this branch. The list of capabilities the desktop advertises to a paired host moves, unchanged, into electron-remote-runtime-client-capabilities.ts, the module the next PR in the stack already uses for it; importers point there. * test(cross-version): stub the launch seed resolver createSupport now reads * fix(native-chat): the desktop tells its own host it picks each launch mode, so retrying an existing chat works with the setting off * docs(native-chat): name the real exit for the released-phone createSupport rule * test(cross-version): a released client still gets the host-setting createSupport answer; a launch-mode client gets supported plus the seed |
||
|
|
0d2300ca8e |
test(e2e): retry the crash probe's main-process reads through the transient-evaluate helper (#24479)
expect.poll does not retry a thrown read, so one spurious 'Resulting promise was garbage collected' from Electron's main evaluate failed the crash-recovery test. |
||
|
|
4e919f3b5a | test: send Codex Ctrl+C to a live raw terminal fixture (#24480) | ||
|
|
02790c53e8 |
Fix Codex status after Ctrl+C copy and side-chat navigation (#24339)
* fix: preserve Codex status on ambiguous Ctrl+C input * fix: confirm Codex turn cancellations from host rollout records * fix: keep ephemeral side hooks separate from the Codex main turn * perf: watch active Codex rollouts and skip unrelated records * fix: retain confirmed Codex cancellation across late relay events |
||
|
|
6e7e964705 |
feat(orchestration): tell each agent its own orchestration address (#22636)
* feat(orchestration): report the caller's host-resolved orchestration address in orca status orca status --json gains a caller block: the calling agent's address as the host resolved it from the identity its environment carries. A structured session is session:<id>; a terminal agent is its handle, with whether the host still knows it. A session the host refuses reports that refusal instead. The host answers through a new read-only orchestration.callerShow, so the session claim runs through the same dispatch-entry resolver every verb uses. An older host leaves caller unresolved. The help footer and the run/check specs stop describing identity only in terminal terms. * docs(orchestration): tell agents their address and give chat coordinators a non-waiting loop The orchestration guide now states that a chat session's address is session:<id> (never the provider's id), that orca status --json reports it, and that no caller flag should name another agent. A consuming check no longer tells every caller to name itself with --terminal. A chat coordinator starts its wave, ends the turn, and on each turn Orca starts for new mail runs a non-waiting check and ack; it never blocks in check --wait. The guide also names ORCA_CLI_COMMAND as the executable in chat sessions. * feat(native-chat): add Copy Orchestration Address to a structured chat's context menu Copies session:<id>, the Orca-minted address other agents message the chat by. The existing Copy Session ID still copies the provider's id and is left as is; the new action is labelled so the two cannot be confused. Strings are added to every locale catalog. * feat(orchestration): tell every dispatched worker its own orchestration address The worker preamble names the coordinator's address rather than a terminal handle, and states the worker's own address. A structured worker is told it is session:<id>, that its coordinator reaches it there or at its dispatch mailbox, and that mail arriving while it is idle starts a new turn. Its commands invoke the CLI through ORCA_CLI_COMMAND in its own shell's form, the same rendering the pointer turn uses, because a bare orca in a login shell can reach a different Orca. * docs(orchestration): give the ORCA_CLI_COMMAND form for POSIX shells and PowerShell A chat session's shell reads the variable as "$ORCA_CLI_COMMAND" in a POSIX shell (Git Bash included) and as & $env:ORCA_CLI_COMMAND in PowerShell, the same two forms the pointer turn and worker preamble render. The chat coordinator loop now runs the check its pointer turn names. * docs(orchestration): say that /clear gives a chat a new address and Orca moves its Runs * fix(orchestration): keep CLI resolution in the shared skill stub and the orchestration kernel in budget The guide-contract tests own two rules this PR broke: only the shared skill stub may describe how to resolve the CLI, and the always-loaded orchestration kernel stays within 202 lines. The ORCA_CLI_COMMAND text moves to the stub's resolver block, which now covers chat sessions and login shells beside WSL and gives the POSIX and PowerShell forms; every skill projection and the bundle manifest are regenerated. The kernel keeps one line each for the caller's address, the environment-resolved check caller and the chat coordinator's non-waiting loop; the loop steps and the address details move to the coordinator-loop and messaging references. The two kernel pins now assert the new check contract and refuse the old --terminal <your_handle> shape. * fix(orchestration): refuse a blocking check --wait from a native chat session A chat runs turn by turn through a shell tool with its own timeout, so a blocking wait is killed mid-wait and retried. The host now refuses it with wait_requires_terminal and the turn-loop recovery, keyed on the session's lease: a session a terminal view holds still runs in a PTY and may block. * fix(orchestration): resolve orca status's caller with the verbs' ladder, host-side callerShow now answers a terminal caller the way the coordinator verbs act: the carried handle while it is live, else the handle its pane was reminted as. The CLI always asks, so the host decides that a process has no identity from the same envelope every verb sends; a pane key alone now resolves. * fix(orchestration): show a structured worker as session:<id> wherever agents read mail A structured worker was session:<id> in orca status and its preamble, but structworker_<uuid> in check rows, banners, reply hints, its own check label and a sub-worker's coordinator line. The minted handle is now only the mailbox key: mailbox reads, the check label and preamble coordinator lines spell the worker session:<id>, which the host binds back to that mailbox. Send receipts still echo the stored row, whose sender key worker_done settlement matches. * fix(orchestration): teach a chat worker the turn loop and pin preamble parity at the contract The worker preamble was byte-identical across modes except its address, so a chat worker was taught a 600s blocking ask its shell tool kills before the message ID for --resume prints, heartbeat exemptions for check --wait, and to keep a shell open. Parity now pins the contract (sections, verbs, flags, lifecycle ids); interaction discipline follows the mode: a chat asks with a 5s wait and ends its turn, owns sub-workers through the turn loop, and names itself session:<id> in every command. The guide says a chat's address survives /clear and that Orca refuses a chat's check --wait. * test(orchestration): pass the db to preamble delivery and fence a terminal-view waiter The coordinator line maps a structured coordinator's handle through the orchestration db, so delivery takes it from its caller. The consumer-fencing waiter test now waits as a terminal-view session, the only session kind that may still block in check --wait. * test(orchestration): read the Run id with the fixture's checked accessor * feat(orchestration): copy a chat's conversation address, which /clear keeps Copy Orchestration Address copied session:<live id>. A chat's address is its conversation's, derived by the host from the session records, so the menu now asks the host for it at copy time through orchestration.sessionAddress, the same derivation a verb acting as that session binds to. A host that predates the method has no /clear lineage, so there the live id is the address. The guide's /clear text says the address survives and nothing moves. * test(orchestration): pin that a cleared chat's successor copies its conversation's root address * test(orchestration): give the mode-opacity fixture's record store the listing a lineage lookup reads A structured worker's agent-visible address now resolves through its conversation's lineage, which lists the session records; the fixture's partial store lacked that listing, so the sub-worker start failed at dispatch input. * refactor(orchestration): format a chat's copied and reported address from its root Orca session id Carries the Orca session id rename into the self-address surfaces. orchestration.sessionAddress, the copy action's fallback, callerShow and the address a structured worker is shown now format `session:<id>` from the conversation's bare root Orca session id with formatOrcaSessionAddress, and ids arriving as strings are checked with isOrcaSessionId first. The CLI status line, the check caller label and the dispatch preamble spell the prefix from the one exported constant. * refactor(orchestration): resolve a session's reported address through the party resolver, and refuse every session's check --wait - orchestration.sessionAddress, and the agent-visible spelling of a structured worker, resolve through the party resolver, so they format the lineage root the one id hook derives; sessionAddress.sessionId is classified as a target. - With the terminal handoff gone every structured session runs turn by turn, so check --wait is refused for any session caller, a worker included, and the session caller no longer carries its lease's runtime kind. - The coordinator loop no longer mentions a terminal view, and the messaging reference says a chat takes messages but is refused as a Dispatch assignee. * fix(orchestration): cap a session caller's blocking wait below its shell tool instead of refusing it A chat or structured worker runs each command under its provider's shell-tool timeout, so check --wait was refused for every session caller and chats were taught a separate loop. The host now caps check --wait and ask for a session caller below that timeout (Codex 10s one-shot exec default, Claude Code Bash 120s) and answers the normal timed-out result, so the terminal coordinator loop runs unchanged in a chat. A terminal caller's wait is untouched. * refactor(orchestration): teach a chat worker the terminal worker's preamble, byte for byte but the address One preamble for both modes: the chat variant (short ask, end your turn, this chat stays available, ORCA_CLI_COMMAND invocation) is deleted. A structured worker's only difference is its address, session:<id>; the byte-parity test between modes is restored with just that substituted. * docs(orchestration): drop every chat-specific instruction; name the address once, generically The guide, its references, the shared CLI-resolution stub and the help return to main's text, with one kernel line saying `orca status --json` shows your address (the kernel stays at main's length). The status caller block reports only the opaque address, the same shape for a chat and a terminal agent. Guides regenerated. * test(orchestration): pin that a chat and a terminal agent see the same preamble, pointer and guide * test(orchestration): key the wait-cap fixture's records by plain session id strings * chore(i18n): add the copy-address strings at the head of native-chat, clear of main's catalog edits * test(orchestration): fail the capped-wait test on the settle, not on the test timeout * test(orchestration): keep main's takeover assertions on a session coordinator's waiting check With the session wait capped rather than refused, the test main extended runs as it is: the restack re-added the shorter pre-main version over it. * fix(orchestration): show a /clear-ed chat its lineage root's address everywhere it reads its own check labelled a session caller with session:<live id>, while orca status and the preamble show the conversation's root. The CLI cannot read the lineage, so the label now comes from the same host answer orca status prints (orchestration.callerShow), asked only when there are messages to render, and falling back to the live id only when the host cannot say. The host also spells a session address it shows an agent with the lineage root: a dispatch preview filled in from the chat's own address, and the provider-id refusal that names a session's address. * test(orchestration): the parity test's gate facts resolve like the host's * test(orchestration): the parity test's gate facts carry the submissions main's pointer lane reads * fix(orchestration): wait a chat's check --wait and ask exactly as long as a terminal's The host capped a session caller's blocking wait (Codex 6s, Claude 100s) so the provider's shell tool would not kill it. Neither provider kills a long shell call: default Codex's exec tool yields and keeps the command running, and Claude Code moves a timed-out Bash call to the background. Terminal agents run the same tools uncapped, so the cap only made a chat coordinator re-poll every few seconds. A session caller's check --wait and ask now wait the budget asked for. * fix(orchestration): show every agent one address, the mailbox address its mail is keyed by A structured worker was told `session:<id>` in orca status and its preamble, but its own send receipts, inbox, worker-list, dispatch previews and task rows still showed the `structworker_` handle its mail is stored under; only some reads were re-spelled. Instead of re-spelling reads, callerShow, sessionAddress and the preamble now report the caller's stored mailbox address (mailboxAddressOf): a terminal's handle, a structured worker's handle, and a chat's `session:<lineage root>`. The read-side re-spelling layer (withAgentVisibleAddresses and its check/banner/preamble call sites) is gone. Dispatch previews spell the coordinator by its party's mailbox address, so a `/clear`ed chat's dispatch-show still names its root. * refactor(orchestration): label check output from what the CLI already knows check asked the host for orchestration.callerShow after every non-empty check by a session, only to fill a label used when a legacy row lacks to_handle, which host rows never do. The label is again the caller's handle or its injected mailbox address, with no second round trip after mail is consumed. * refactor(native-chat): offer Copy Orchestration Address on chat tabs only No mount passes both terminal-pane actions and an orchestration address: a chat shown inside a terminal pane is that terminal's agent, copied by its terminal ID. Drop the unreachable terminal-pane placement and its tests. * fix(orchestration): have orca status report the handle the agent's own check reads After a window reload a terminal agent keeps ORCA_TERMINAL_HANDLE=term_old while its pane is reminted as term_new. callerShow reminted and advertised term_new, but check, send and ask act as the carried handle and never remint, so mail sent to the advertised address was never read by that agent. callerShow now answers the carried handle with its liveness, and null for a pane key alone, from which the mailbox verbs have no identity. Resolving terminal callers once on the host for every verb is a separate follow-up. * fix(orchestration): read the renamed coordinator line in the long-prompt repro, and trim round-one leftovers The reliability repro's fake worker parsed "Your coordinator's terminal handle is:", which the preamble now spells "Your coordinator's address is:", so it silently skipped worker_done; it accepts both. Dispatch and its dry-run go back to main's coordinator line (their `from` is already bound at the entry); only dispatch-show, whose `from` is unbound, resolves it. Also drops a stale status-caller comment, trims the wait test to its one uncapped-wait case, and reverts comment-only churn in the worker opacity test. * docs(orchestration): keep worker obligation 1 as main words it The guide grows by the one caller.address line; the parity test bounds the kernel at main's length plus that line instead of forcing a reword. * test(native-chat): prove a structured chat tab offers Copy Orchestration Address Renders the pane-commands hook as a structured chat tab and selects the item: it asks orchestration.sessionAddress with the tab's target and session id. Also corrects the menu item's comment to what it copies. * test(orchestration): D5's tests expect the orca_session_id prefix and 'Orca session ID' wording * fix(orchestration): name a session by its Orca session ID, and leave terminal agents as main has them Terminal agents keep main's exact wording: a terminal worker's preamble is byte-identical to main's, and orca status prints nothing new for them. A session is named by its Orca session ID (orca_session_id:<id>, its /clear root's): a structured worker's preamble says "Your Orca session ID is: …" and its commands use that ID, and a session coordinator is "Your coordinator's Orca session ID is: …". orca status shows a session caller's `caller.orcaSessionId`; callerShow answers null for anyone else. The chat tab menu item becomes "Copy Orca Session ID" with a tooltip saying what the ID is, and its toasts match. No agent-read text calls this ID an address. A structured worker's mail is still keyed by its minted handle. * test(orchestration): check CLI help and status for "address" wording from a CLI test The node project cannot compile src/cli, so the guard over CLI help, specs and status text moves to src/cli; both halves share one pattern. Also brings two comments and the long-prompt repro's coordinator-line regex to the Orca session ID wording. * fix(native-chat): keep the Orca session ID tooltip within the tooltip primitive's typography Drops a restyle the design-system gate refuses on TooltipContent, keeps "Agent" untranslated in the Japanese tooltip as that catalog does, and types the test's tooltip mock without an assertion. * fix(native-chat): the Orca session ID tooltip names the agent CLI's own session ID in the singular |
||
|
|
8c1670f2c4 |
test(e2e): fake Codex answers the --no-daemon --help probe without a spawn (#24440)
* test(e2e): fake Codex answers the --no-daemon --help probe without a spawn Since #23933 Orca runs `codex --help` before a path-named Codex launch. The fakes logged it as an agent spawn and held the probe for its 5 s timeout. Move the app-server refusal and --help answer into one shared FAKE_CODEX_LAUNCH_PROBES_SOURCE used by every fake Codex. * test(e2e): stop asserting the dispatch capability column a current worker no longer has #23994 stopped minting the per-dispatch capability, so capability_hash is null. The --help spawn failure used to stop this test before it got here. |
||
|
|
ccb63afc06 |
fix(native-chat): a Stop still reads as yours after Orca restarts, because the turn's end reads the Stop's event (#24311)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused
Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.
* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered
A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.
* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget
* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it
The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.
* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card
* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows
Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.
One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.
The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.
Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.
* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones
A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.
* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction
The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.
* fix(native-chat): stop creating the unused queue pause table
The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.
* fix(native-chat): a Stop's pause never hides the restart pause
A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.
Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.
* test(native-chat): pin the Stop's no-resend, lift and held-card rules
- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
again" at one instant, before a queue ignoring the pause re-sends. They
now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
whether or not a person's turn lifts it; it now reads the Stop's pause
before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
queued before a rewind.
* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller
The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.
* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event
* test(native-chat): pin that Stop and Resume rows never reach apps or count as history
* test(native-chat): only a person's Stop event pauses the queue
* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop
Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.
* test(native-chat): a card held at a starting agent is checked before the Stop's timing
Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.
* test(native-chat): a released build keeps and folds a journal holding Stop events
Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.
* style(native-chat): format the Stop event changes
* test(native-chat): type the released build's exports through one checked helper
* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only
* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade
The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.
Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.
* fix(native-chat): a Stop that stops nothing new writes no Stop event
A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.
It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.
* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop
* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled
* fix(native-chat): any later Stop event ends a person's Stop pause
A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.
An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.
* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed
A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.
A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.
Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.
* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes
A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.
The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.
* fix(native-chat): a Stop still reads as yours after Orca restarts before the turn ends
Every stop that ends work now writes the Stop's event before it ends the child: a
person's close of the chat, an eviction (worktree teardown, orchestration stop, tab
cleanup) and the idle sweep's stop of a start that never landed. A stop that ends
nothing writes nothing, and quit writes none: its resume marker records why.
The turn-end write reads the latest Stop event where every turn row is built, so the
adapter's settle, the host's fallback and the relaunch's settle all agree: a turn a
person's Stop or close named, ending with no verdict of its own after that Stop, ends
as their cancellation. A relaunch's probe-bounded end is no earlier than a Stop that
found the turn running. When the provider refuses the interrupt and the turn runs on,
a refusal row answers the Stop, so a later crash still reads Failed; pressing Stop
again after a refusal is a new Stop.
* refactor(native-chat): a stop no longer carries its cause; the turn's end reads the Stop event
The cause of a stop was threaded in memory from each entry through the host's stop
step, the adapter router and each adapter's close onto the `ended` it settled with,
and Claude kept a per-turn copy of a Stop it sent. All of that is gone: adapters
settle a turn they cut as interrupted with no verdict, the host's fallback does the
same, and the one rule where a turn row is built (`turnEndAfterStop`) reads the
journal's latest Stop event to say whether it was a person's.
- `closeSession` / `disposeSession` take no cause; `ended` has no `stopCause`.
- Claude reads an error result after a person's Stop as their cancellation from the
journal's Stop event (through the event sink), not from a per-turn slot, and a
refused interrupt is the host's refusal row, not `withdrawTurnStop`.
- An owed wind-down keeps no cause: its retry's fallback reads the Stop event.
- The mutation context's Stop passes no cause: its step already wrote the event, and
the delivery loop's child-end reason is read back from it.
- A Stop pressed before its turn showed applies to the turn that opens under it,
unless a send a person made since was accepted.
* test(native-chat): a turn a later send opened is no Stop's that named no turn
* test(native-chat): the restart test's death proof carries its detail
* refactor(native-chat): a refused Stop leaves no record; a Stop only ever ends the turn it names
The stop-refused mark is gone: its tombstone kind, its fold, the clock-keyed match that tied it to
a Stop, and the exception that let a second press after a refusal write a new Stop. A Stop that
stops nothing writes nothing. A Codex refusal names a turn that is no longer its active one, and
the Stop names that turn, so the turn running instead never reads as the person's by its id alone.
* fix(native-chat): a Stop pressed before any turn showed stops only the turn opened next
A Stop that named no turn read as the person's cancellation for every later turn that opened
after it, until a send a person made was accepted. The queue's drain, orchestration mail and a
restart continuation send as the host, so a turn they opened long after, cut by a crash, read
"Interrupted" as if the person had stopped it. The Stop now applies only to the first turn
opened after it.
* fix(native-chat): an older Claude's error end after a Stop pressed before its echo reads Interrupted
Claude CLIs before 2.1.91 end an interrupted turn with an error result that names no reason. The
translator judged whether a person's Stop explained it by its own copy of the Stop rule, which
ignored a Stop that named no turn, so a Stop pressed before Claude echoed the send read "Failed".
The translator now writes such an end as interrupted with no verdict and no error row whenever a
person's Stop may name the turn, and the journal's one rule decides as it writes the end.
* fix(native-chat): a person's Stop and /clear each name why they end the agent
The host's mutation path ended the agent with one "recorded" ending for every caller, which read
back the reason of whatever Stop event the journal held last, however old. /clear writes no Stop
event, so its end took an unrelated earlier reason. Each caller now names its own: the chat's Stop
`user-stop`, whose event its own step wrote, and /clear `user-close`, the user replacing this chat.
* fix(native-chat): a host stop judges whether it ends work after the provider's rows land
A close, eviction or host stop decided whether it ended a running turn from the journal as it
stood, while the provider's own rows (the turn its echo opened) could still be in the session's
event sink. A close landing in that gap wrote no Stop event, so the turn it cut read as news. It
now reads after the sink drains, as a person's Stop does, through the same check; a drain that
fails or takes over a second reads working.
* fix(native-chat): a Claude Stop naming a turn that just ended still marks the follow-up it cuts
A phone names the turn it last saw. When that turn had ended and a follow-up was still unechoed,
Claude's Stop interrupted the follow-up and ended the child, but the Stop's event named the ended
turn, so the follow-up's turn the child's end cut read "Failed" under "Cancellation requested.".
A Stop that ends the provider's session ends whatever is in flight, so its event now names the
live turn or none, and a Stop that names none binds the turn opened next. Codex keeps naming only
the turn the Stop names.
The Claude Stop turn-end tests move to their own file, since the session-ending Stop suite is at
its line budget.
* fix(native-chat): the idle sweep reads working by the same rule as a stop's event
The sweep judged a chat resting while a send whose reply was lost was still unanswered, but the
stop's event writer counts that send as work. So the sweep evicted it and wrote an evict event,
which ends a person's Stop pause and let the cards behind it drain on their own. The sweep's owed
work now reads the main agent working the way every session list and the event writer do.
* test(native-chat): an aborted eviction's injected drain failure lands on the eviction's own drain
A host stop now drains the session's sink once to judge whether it ends work, so the tests that
fail the eviction's drain-published step skip that first drain.
* fix(native-chat): the idle sweep's rest writes no Stop event; it evicts a send that never echoes
The previous commit made the sweep count an unanswered send as owed work, which pins a chat whose
admitted send Codex never echoes forever, and the sweep exists to retire exactly that. That rule
returns. The sweep stops only an agent it judged resting, so its eviction now writes no Stop
event, whatever send it retires: a person's Stop pause holds through it.
* fix(native-chat): stopping a start that carries no send writes no Stop event
A host stop, eviction or close of a starting child wrote a Stop event whatever the start carried.
A start with a send already reads working, so the clause only mattered for a start with none,
which ends no turn and no send: its event only lifted a person's Stop pause and bumped the idle
clock, which is why the idle sweep had been changed to close the conversation in the same pass.
The clause goes and the sweep is #24072's again. The child's end still reads host-stop, as before.
* test(native-chat): a Stop's pause across a restart is tested with a restart that writes no event
The rig's restart closes the chat with an eviction, which now writes a Stop event when work runs
and so ends a person's Stop pause. "A Stop never hides a restart's pause" then passed with no Stop
pause left to hide anything. Those tests, and the pause-lift test whose dropped assertion returns,
restart as a process that dies with no close, which like a quit writes no Stop event, and assert
that both the Stop's and the restart's pauses are in force first.
* fix(native-chat): a host stop of a turn a person's Stop is still ending keeps that Stop's reason
An eviction or host stop that landed while a person's Stop or close was already ending the same
turn wrote a newer Stop event, and the turn's end reads only the latest, so the person's Stop of
that turn read as news. A host reason now writes nothing while a person's Stop still decides what
runs: the live turn it names or bound, or, with none, the turn a send opens next. The person's
own close still writes. The E2 tests now open and end the stopped send's own turn, as Codex does,
so the mail turn after it is not the turnless Stop's.
* fix(native-chat): an older Claude's error on a later turn keeps its error text after a Stop
The translator left an error result that names no reason to the journal's Stop rule whenever a
person's Stop named the turn or none, but the rule binds a Stop naming no turn only to the turn
opened next. So a real error on a later turn read "Failed" with its error text dropped. The
translator now asks the journal's rule itself (`personStopDecidesTurn`, the one core
`turnEndAfterStop` and a host stop's in-force check share), so the two cannot disagree.
* fix(native-chat): a Stop of a start that never landed binds no later turn, whatever sent it
A person's Stop pressed while the agent starts names no turn, and the send it stopped is
cancelled before it opens one. The Stop then bound the next turn anything opened (orchestration
mail, a restart continuation, the queue's drain, all of which send as the host), so a host
eviction of that turn wrote nothing and its crash or close read as the person's cancellation. A
Stop that named no turn now binds only a turn no send journaled after it opened: any send since,
of any origin and not refused, opens its own. The E2 test's mail send is accepted as Codex
accepts it, instead of opening the stopped send's own turn first.
* test(native-chat): a rewind's restated turnless Stop binds no turn opened after the rewind
A Codex rewind restates a person's Stop still in force after the turns it keeps, at a new
sequence, so by sequence alone it would bind the next turn opened after the rewind. A send
journaled after the restated row voids that binding (the previous commit), which this pins.
* fix(native-chat): a relaunch settles a person's stopped turn with no "stopped while in progress" row
After a restart, a turn a person's Stop ended reads "Interrupted after N" with the muted mark, but
the relaunch still added the error row saying the provider stopped mid-response, which a live Stop
never writes. The settle now skips that row when every turn it interrupts is the person's Stop's
by the journal's one rule; a crash nobody stopped keeps it.
* test(native-chat): the unexpected-exit settle's journal fake answers whether a person's Stop decides a turn
* fix(native-chat): a host stop whose sink drain fails reads the journal as it stands
A host stop drains the session's sink before judging whether it ends work, and a failed or slow
drain read as working. So an eviction of an agent at rest wrote a Stop event that ended nothing,
which lifts a person's Stop pause, and a close wrote a person's event naming no turn. The drain is
now best effort: the stop goes ahead either way and only its record is at stake, so a failed or
slow drain leaves the journal's read as it stands. A person's Stop keeps its own rule.
* fix(native-chat): a Stop that named no turn applies only to a turn a send it stopped opened
A person's Stop pressed before any turn showed names no turn. It bound the first turn opened
after it, then (
|
||
|
|
976dc00337 |
fix(native-chat): Stop's pause is worked out from the chat's history, so a steered message is never re-sent (#24072)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused
Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.
* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered
A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.
* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget
* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it
The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.
* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card
* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows
Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.
One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.
The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.
Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.
* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones
A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.
* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction
The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.
* fix(native-chat): stop creating the unused queue pause table
The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.
* fix(native-chat): a Stop's pause never hides the restart pause
A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.
Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.
* test(native-chat): pin the Stop's no-resend, lift and held-card rules
- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
again" at one instant, before a queue ignoring the pause re-sends. They
now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
whether or not a person's turn lifts it; it now reads the Stop's pause
before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
queued before a rewind.
* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller
The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.
* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event
* test(native-chat): pin that Stop and Resume rows never reach apps or count as history
* test(native-chat): only a person's Stop event pauses the queue
* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop
Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.
* test(native-chat): a card held at a starting agent is checked before the Stop's timing
Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.
* test(native-chat): a released build keeps and folds a journal holding Stop events
Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.
* style(native-chat): format the Stop event changes
* test(native-chat): type the released build's exports through one checked helper
* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only
* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade
The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.
Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.
* fix(native-chat): a Stop that stops nothing new writes no Stop event
A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.
It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.
* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop
* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled
* fix(native-chat): any later Stop event ends a person's Stop pause
A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.
An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.
* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed
A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.
A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.
Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.
* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes
A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.
The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.
* refactor(native-chat): one reading of a Stop's turn for its event and its note
A Stop's event and its note each worked out the same two facts on their own:
which turn the Stop is about (the one it named, else the one running), and
whether a named turn is the one the journal shows running. The event decides
before the interrupt; the note and whether the session ends decide after the
provider's answer, so those decisions stay separate, but the facts they read
are now one helper each in structured-agent-session-turn-stop-notes.ts:
structuredAgentSessionStoppedTurnId and
structuredAgentSessionStopNamesTurnNotLive. The event's turn, the note's key,
the session-ending condition, the running-command check and the repeat check
all read them. No behavior change.
Tests: a Stop naming no turn records the running turn on its event, and
rewrites that turn's note as a Stop naming it does.
* refactor(native-chat): a failed-interrupt Stop reads its turn through the shared helper
The new branch that ends a Codex child after a failed interrupt asked
whether the Stop's turn still runs with `turnId ?? liveTurnId`, a third
copy of "the turn a Stop is about". It now reads
structuredAgentSessionStoppedTurnId, the value the note key already uses,
read at the same point before the cancel. No behavior change.
Test: a Codex Stop whose interrupt failed ends the child, holds the card
queued before it with the queue paused, and writes its Stop event before
the turn's end.
|