Takes #24917's request ids everywhere a launch is started: a launch joins another only on an equal
id during its first attempt, and every launch call site passes one. C2's notes hold stays the only
one: the staged message carries the notes' keys and the hold derives from the saved outbox, so
#24917's in-memory hold and new-agent-prompt-outcome stay deleted. Its new test case (a chat
still starting is closed without waiting on its create) is ported into notes-carried-by-chat.
The restack kept this branch's adapted probe tests in the delivery file and the split's older
versions (a head parked for Retry, a force-retried head) in the new probe file, so each case ran
twice in two forms, one of them for behaviour this branch removed. The probe file now holds this
branch's versions only, using the shared harness's clock and seeding helpers, and the delivery file
keeps the sending-notice cases. The tests read requests through checked helpers instead of casts,
and the harness's seeded entry no longer carries the removed Retry marker.
"The same request" was recognised by its content (agent, workspace, trimmed
text, sent or drafted), which split one action whose text changes between
deliveries ("Fix with AI" re-fetches logs) and merged two different actions
that happen to carry the same text.
Each user action now mints one request id where it is handled (the click,
menu pick, notes send, shortcut, Fix with AI press, quick command; a worktree
create uses its creation id; programmatic starts mint once at their entry)
and passes it through the launch plan, its verdict and the structured launch
options, where it is required. A new start joins a chat only when its first
attempt carries the same id: a double click or a caller retrying its own call
makes one chat and sends the first delivery's text once; any other action
opens its own chat, whatever its text. An empty chat that is still starting
is still claimed once by the first action with text, and then belongs to that
action's id. Resume launches keep joining their conversation's launch.
The + menu and new-tab search no longer grey out an agent while one of its
chats starts: every pick is a new action and opens its own chat. The id lives
in memory only; nothing reads it after a reload (a restored launch is reached
by its session id through Retry, never joined).
NativeChatStructuredSessionDelivery.test.tsx went over the 800-line lint
limit after main was merged. The eight tests for the automatic probe of an
unconfirmed message move unchanged to
NativeChatStructuredSessionDelivery.probe.test.tsx, which uses the shared
structured-session test harness for its mocks. The outbox seeding and probe
clock helpers move into that harness so both files share them.
The at-most-once reliability gate now lists the new file, with a fresh
evidence run.
* fix(native-chat): a resent send id gets its recorded answer, never an early refusal or a made-up record
The host now looks a resent send id up before preparing the session. A row
that settled refused answers with its refusal before the chat is opened. A
resend whose chat cannot be opened or made ready answers unknown instead of a
refusal. A /clear in flight refuses only ids the ledger does not hold.
A send row now records the journal epoch it was admitted into. An unsettled
row with nothing written in that same epoch runs for the first time; under a
later epoch the host answers unknown instead of reconstructing a submission
it never had. The host advertises agent-session.send-answers-proof.v1.
* test(native-chat): pass the ledger row to the thread-goal rerun check
* fix(native-chat): a send's answer commits with its write, and a resend is answered before any write
A send (and /compact) settles its ledger row `succeeded` in the same SQLite transaction as the
submission or queued draft that accepts it, so a row still `pending` proves nothing was written and
a resend runs it for the first time. The unknown-before-run mark and the per-row journal epoch go.
A resent id is answered from its row and the journal before preparation starts an agent and before
any write transaction: a recorded refusal with nothing opened; otherwise the conversation is opened
(no agent start for a send) and replayed, and a conversation that will not open answers unknown.
* fix(native-chat): a ledger refusal is answered first, and a /clear refuses only a send's first run
An id the ledger refuses (expired, conflict, invalid, capacity) is answered as admission would,
with no journal read, preparation or write, so a closed chat or a read-only store answers it too.
A re-read after the replay open that comes back refused returns that refusal.
Whether a /clear is in flight is read when a send arrives and applied in the send's preparation
for a first run only: an id the ledger holds by the send's turn, including one whose earlier
attempt was queued ahead of the clear, is answered from its record. MutationPlan makes
settlesWithWrite and settledOutcome exclusive; the capability text no longer promises a refused
id never sends.
* docs(native-chat): say what a replay's preparation does for every plan
* fix(codex): keep Orca-only MCP servers when refreshing the retained shared home
The refresh for panes that outlive an update treated the old shared
home's whole MCP root as owned by ~/.codex, so it deleted servers the
user had added from an Orca terminal, which existed only there. Read the
home's settings baseline instead, as the normal mirror does: drop only
servers the last mirror copied from ~/.codex. No baseline keeps the old
behaviour; an unreadable one skips the refresh. The baseline is not
advanced, keeping the refresh one-way.
STA-9109
* test(codex): type the MCP ownership baseline fixture
---------
Co-authored-by: Orca Worker <orca-worker@localhost>
* feat(codex): tell Windows users once what stays behind when Codex moves onto ~/.codex
When Windows' system-default Codex first runs on ~/.codex (launch prep or
the usage poll), main decides once whether Orca's managed home was ever
used and which MCP servers lived only there, and persists that in UI
state. The renderer shows one dismissible toast when a Codex terminal
exists, after the server-isolation notice rather than on top of it, and
clears the notice when shown.
The "kept only in the managed home" MCP rule is extracted into
isRuntimeOnlyMcpServer, which the config mirror merge now uses too, so
the notice names exactly the servers the mirror would have kept.
* fix(codex): stop counting Orca's own config.toml as use of the old Codex home
Orca's hook install writes that home's config.toml on every startup, so its
presence was true for nearly every Windows user with Codex. The home now
counts as used only with recorded sessions or an MCP server of its own.
Resolver tests keep one case per input source.
* refactor(codex): ask main for the shared-settings notice instead of persisting it
The persisted missing/object/null field, written from launch prep and the
usage poll, becomes a plain codexSharedSettingsNoticeSeen flag mirroring
codexTerminalServerIsolationNoticeSeen. When a Codex terminal first appears
and the flag is unset, the renderer asks codexConfigSync:sharedSettingsNotice
once; main answers read-only (Windows, system default on ~/.codex, managed
home path without mkdir) and maps any read error to null.
Runtime-home routing, launch and the test harness return to main's code.
The notice no longer waits for the server-isolation toast; they may stack.
The Codex-terminal watch moves to codex-terminal-presence.ts.
* refactor(codex): watch for the first Codex terminal in one place for both notices
The server-isolation notice now passes its due check to
whenCodexTerminalAppears instead of keeping its own copy of the presence
scan, input filter and subscription loop. Its behaviour and tests are
unchanged.
* docs(codex): trim isRuntimeOnlyMcpServer's comment to why it is shared
* refactor(codex): keep McpServerTomlOwnership private to its module
* test(codex): cover the shared-settings notice channel without type assertions
Handlers are looked up by channel now that two are registered, so the
status tests no longer depend on registration order.
* refactor(codex): show the Windows shared-settings notice without asking main
Every way of detecting who relied on Orca's old Codex folder had false
positives, so the renderer now shows one static toast on Windows the first
time a Codex terminal exists. This drops the main-process resolver, its IPC
channel, preload line, web stub and shared type, and the MCP-names variant
of the description.
* refactor(codex): restore the MCP server ownership helpers to main's shape
The static notice no longer reads MCP servers, so the shared
isRuntimeOnlyMcpServer extraction has no second caller.
* refactor(codex): let each notice decide when it is due, so the Codex watcher only watches
The isolation notice now selects its due predicate and starts the watcher only while due, so whenCodexTerminalAppears no longer takes an isDue or re-checks hydration and settings. The shared-settings notice uses isLocalWindowsDesktopClient, its test stubs the user agent instead of mocking pane-helpers, and the hydration safeguard it relies on is now tested on the UI slice itself.
* test(codex): drive the Codex notices through a reactive store, and drop a redundant hydration gate
The server-isolation notice now reads "is it due" through a store selector, but its test
mocked the store without re-rendering, so a due change after mount (persisted UI loading,
the setting turning off) was never exercised. The notice tests now share one harness backed
by a real zustand store, the shared watcher gets its own test, and both notices cover the
seen flag loading after mount.
persistedUIReady is dropped from isNoticeDue: the seen flag defaults to true and only
hydration clears it, in the same update that sets persistedUIReady. Both notices now gate
the same way.
* fix(codex): keep the shared-settings toast until dismissed, and shorten it
It is marked seen before it shows, so a 15s auto-close could lose it for good while the user
is typing in the Codex terminal that triggered it. Every other one-shot notice that marks
itself seen on show stays until dismissed; this now does too.
The text drops the sentence that repeated the title and keeps only what to expect and do.
---------
Co-authored-by: Orca Worker <orca-worker@localhost>
The listing purge is reconciliation, not a delete: re-pairing a server under
the same id purges its worktrees' tabs, and the host then restores the same
chats under the same draft keys, so deleting drafts there erased them.
Removing a worktree, a project or a folder workspace now reads the draft
keys of the workspace's open chats before the host round trip, and deletes
them only once the host confirms. A listing refresh that drops the tabs
during the round trip no longer hides them, and a refused delete keeps them.
A worktree's chat drafts were deleted only by the removal's own renderer
teardown, which looked them up from the worktree's tab lists after the
removal round trip. The host announces the change before it replies, and
the listing refresh that starts can purge those tab lists first, so the
teardown found no tabs and the drafts stayed. The bulk purge now deletes
the drafts of the tabs it drops, and the teardown deletes them before it
closes any tab.
A message's ending can fire inside a store update (a removed workspace's unpublished chats are
settled there), and clearing the notes from inside it wrote the store re-entrantly, so the update's
own result could put a note back. The clear now runs in a microtask, after that work returns.
Closing a launch that never published handed its other unsent messages to a draft no chat shows,
and ended them as "returned", which cleared the notes they carried. They now end as discarded, so
their notes go back on the shelf as before; the text is still kept in that draft.
A host that refuses a send before looking up its id (a host with structured chats turned off)
refuses every resend the same way. Since a message past the host's window is no longer handed back
on wake, such a message was resent forever and held up everything queued behind it. Now, past the
host's window for the id, an answer the host itself gave that settles nothing (a refusal returned
or thrown, or a call turned away) hands the message back with "couldn't confirm" words, unless the
journal shows its row. Only a lost connection keeps it resending, since that says nothing about
the host. Derived from that answer and the id's own time; nothing new is stored.
Also corrects the host-window docstring, which still said an answer that never comes ends it.
The merge with main kept both versions of two delivery tests. Each still sent its first message
outside act() and waited with waitFor, whose polling uses the setTimeout that main's fake clock
freezes, so both hung until the test timed out. The target-switch test also sent its message twice
and looked for "Message delivery is unconfirmed.", a notice this branch replaced with "Sending…".
Both now send inside act() as main's do, keep main's exact probe timing, and read this branch's
notice.
Some messages are handed back because the loaded journal shows no row for them: an older build's
held message, a send a Stop outran, or one past the host's window. That journal could be one kept
through an outage, so losing contact read as "the host has nothing": a Stop followed by a new send
while offline handed the outrun message back at once, and a laptop asleep for over two days handed
a message back on wake without asking the host.
The read side now knows whether its journal is live: a frame of the current subscription makes it
so, and a closed or failed stream ends that. The chat lets the journal decide alone only while it
is live and attached, and a Stop is only given up for a newer send then. A message that slept past
the host's window is sent again first, and the host's answer (expired, read with the journal)
decides.
A message that came back to the composer returned its images by path only, losing the SSH
connection they were uploaded to, so on a remote workspace the preview broke and, after a reload,
the image asked to be attached again. The outbox entry now keeps each image's connection (saved
with it, never sent to the host, which reads paths) and the hand-back gives it back with the image.
Closing a structured chat's tab cleared the chat's whole outbox: a message still being sent
("Orca will keep trying…") or one queued behind it was deleted, its text kept nowhere. Closing now
throws away only a cancelled launch's own prompt that never went out (its notes return to the
shelf). Any other message that never went out comes back to the conversation's draft, with its
notes following the text. A message that went out may be the host's, so it stays in the outbox
and settles when the chat is reopened. The same rule now serves a launch's cancel and a removed
workspace's unpublished launches.
The image hand-back that this needs moved out of the composer hook into its own module, so the
close path doesn't load the composer (and through it the app store) from inside a store slice.
Drafts of a structured chat were saved under the pane (tab id plus a hash
of the session), so the follow-up that keys them by conversation would
have left every draft saved by this build invisible after an upgrade,
never shown and never deleted. They are now keyed by the conversation
(`agent-session:<sessionId>`) from the start: every composer showing the
conversation shares one draft, Stop and a queued card's Edit give text
back to it, closing the tab keeps it, and removing the worktree deletes
it. Terminal-backed chats keep the pane key and lose their draft when the
tab is closed.
A send now leaves whatever was added since it was sent, text typed or
composed and images attached meanwhile, and clears only what was sent;
a draft replaced in the meantime is left alone.
Lifted from #25207 (62ef8e7804): the conversation key, the composer's
draftScopeKey, keep-on-close and delete-on-worktree-removal, and the
leave-what-was-added send settle.
When a message carrying notes came back to the composer, the notes were cleared first and the
text written to the draft after. A crash between the two could lose both. The text is now saved
to the draft first, on every path that hands it back (the host's answer, the journal, a Stop).
Keys whose notes were not in the store yet stayed pending forever when the note never showed up:
a second "delivered" ending for notes already cleared, a page closed or a workspace removed, or a
note edited while its send was on the way. Each kept a store listener that scanned on every store
write for the rest of the run. A key now waits only while its workspace has not loaded; once it
has, a key with no note is dropped, and the listener detaches as soon as nothing is waiting.
A message an older build left waiting on its Retry is never sent again; once the chat's journal
loads it either comes back to the composer or the host's row shows it. Until then it was drawn as
a message on its way and read "Sending…", so QA saw it flash for a moment before its text went back
to the composer. It is now never drawn as one of this client's sends: no bubble and no notice. The
host's row, when there is one, or the composer is where it shows.
Notes sent to a chat had two owners when the message came back: its text went to the composer
and the notes returned to the shelf, so sending both delivered them twice. Notes now follow their
text: once the host has the message, or its text goes back to the composer (turned away, or taken
back by a Stop), the notes are used and leave the shelf. Only a message thrown away with nothing
handed back (a cancelled launch, or this window closing a failed chat) puts them back.
A send that ended before its workspace's or browser page's notes were in the store cleared
nothing, so those notes showed as unsent once they loaded. The keys are now kept until their
owner's notes load, then cleared.
The test that only called a mocked tab refresh now drives the real host-frame entry points (a
paired host's frame, and a local snapshot through the cancelled-launch filter). Stale "Retry"
wording is gone from a test name and a comment.
Notes handed to a chat were kept out of the next "Send notes" only in memory. After a reload,
while the chat still held the unsent message built from them, the notes were offered again and a
second send delivered them twice. A chat send also cleared its notes the moment it was queued,
so a message that came back to the composer left its notes gone from the shelf.
Now the chat's saved message carries the notes' keys (each naming its workspace or browser page),
for "Send notes to > New agent" and for a chat already open. A note stays off the shelf while a
message this client still holds carries its key and the host can still deliver that message; past
the host's window for it the hold lapses on its own. The message's own ending decides the rest:
once the host has it the notes are cleared as sent, from any send or resend and after a reload;
if it comes back, is withdrawn or is discarded, the notes return to the shelf (a returned
message's text is also in the composer). Nothing is released or deleted because a tab list or a
host sync no longer shows the chat.
One mechanism instead of two: the per-message watch and outcome promise are removed; the outbox's
entry endings drive both the launch prompt's result and the notes.
A press whose answer was lost stays quiet because Orca will send the Stop again. If a newer
message then makes Orca give the Stop up before that resend (or a resend is dropped), nothing was
ever said. The press's own notice ("The agent wasn't stopped.") is now shown when the Stop is
given up; it is honest, since the outcome is unknown.
Since the line under a message Orca keeps resending dropped every step, a refusal whose reason is
a failure fact (signed out, history too large, the agent stopped while starting, a Claude account
problem) or an older host lost its cause too, and said only "Orca will keep trying to send it."
for as long as a day. Each now keeps a cause-only sentence ("The agent is not signed in for the
selected account.", "This needs a newer Orca on the computer running this chat.") with no step,
in all six languages.