mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 08:02:21 +00:00
b5869eeaae72dbf2b7a4bd409f2dd66860bfcf08
966
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f97ca2a49d |
Add Qoder session history and search (#24614)
* Add Qoder session history and search with real CLI coverage * Allow the real Qoder marker file to end with a newline * Keep Qoder tool output out of history previews and search * Keep Qoder search pages readable by older clients * Verify persisted Qoder history after a real generated and resumed task * Negotiate Qoder filters before searching an older execution host * Combine search client imports for the CI plugin gate * Keep the relay search oracle aligned with legacy agent filtering * test(qoder): align search capability contracts and pin old-host fencing |
||
|
|
533446dde6 |
Stop mocked renderer imports from qualifying headless CI (#24902)
* Decouple headless running-work tests from the renderer * Keep the shared running-work probe contract documented |
||
|
|
3fba1952c8 |
perf(tab-bar): a change to one tab no longer re-renders every tab (#24261)
With many tabs open, a change to any one tab (a retitle, an agent finishing, a tab switch, a git status write, or a browser tab update on SSH and web clients) re-rendered every tab in the strip, so the strip stuttered. Each tab is now a memoized row that re-renders only when its own values change, with stable handlers, a stable drag id list and stable drag sensor options. Editor tabs get their own git status, and mirrored browser tabs keep their page-id list while the ids don't change. Part of #24241: opening, closing or reordering a tab still re-renders every tab once. |
||
|
|
1aa0860f7e |
Keep large Markdown previews responsive (#24880)
* Keep large Markdown previews responsive * Fix large preview review navigation and Find budgets * Initialize preview scroll caches once and check viewport visibility * Restore large previews after loaded rows are measured * Refresh loaded Markdown rows after viewport changes * Keep Markdown revisions visible and reuse bounded search text |
||
|
|
f2257ffa69 | fix: dismiss Codex account prompt and return focus to terminal (#24683) | ||
|
|
de8bffe240 |
Fix terminal width cutoff on wide panes (#24687)
* fix(terminal): let wide panes use up to 1024 columns Adapt the wider viewport limit proposed in #16578 to the current runtime, shared RPC schemas, and preview sizing. Co-authored-by: innocarpe <innocarpe@users.noreply.github.com> * test(terminal): wait for probe output after command echo * test(terminal): align RPC boundary with wider viewport limit --------- Co-authored-by: innocarpe <innocarpe@users.noreply.github.com> |
||
|
|
0f167ac659 | test: check plugin fixture worktree cleanup (#24779) | ||
|
|
adf447d958 | test: classify acknowledged remount input as driving (#24750) | ||
|
|
99b2628aef | test: update sidebar setup and remove obsolete permission sentinel (#24734) | ||
|
|
d9fbb4eecf | Keep SSH typing replies inside narrow split terminal panes (#24682) | ||
|
|
66799f7e8f |
Keep paired browser terminal insertion in the host's requested position (#24676)
* test: align source-control fixtures with current store contracts * Bound E2E package setup and retain cancelled-job traces * Remove empty passing sentinels from opt-in socket tests * Make SSH typing pressure fixture readiness and replies observable * Advertise browser support for anchored terminal placement |
||
|
|
b49abdb1f4 |
fix: recover renderer launch failures in the running app (#24250)
* fix(recovery): back off a launch-failed renderer instead of tripping the crash breaker
A renderer that the OS refused to spawn (macOS exit 1003 = LAUNCH_RESULT_FAILURE; field
cause: per-user process limit, posix_spawn EAGAIN) burned the 3-reload crash-loop budget
in ~750ms and raised a "graphics driver" prompt, while the condition lasted minutes.
- launch-failed retries in place on a 250ms..60s backoff (~2 min), outside the breaker;
a loaded document resets it. Other crash reasons keep the breaker.
- Each launch failure records renderer_launch_failed_probe {spawnError} from a cheap
spawn probe, so bundles name EAGAIN/EACCES/ENOENT directly.
- The exhausted prompt says the process limit was hit (probe EAGAIN), drops the
graphics-driver wording, keeps Try Again as default, and offers no Restart:
app.relaunch also needs a free process slot and silently fails without one.
* fix(recovery): skip the launch probe on Windows and probe the prompt once
- Re-check quitting after the prompt's probe; don't re-probe on Copy Commands.
- recordRendererLaunchFailureProbe never rejects (breadcrumb write guarded).
- Windows: no spawn probe; a child per failed launch is the per-operation burst EDR scores.
* test: cover quitting and duplicate renderer launch failures
* test: use typed access in PTY delay regression fixture
* fix: scope extended launch retries to POSIX hosts
* test: cover launch probe behavior on native Windows
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
|
||
|
|
53930a161b |
Keep SSH typing replies visible during background pressure (#24629)
* test: align source-control fixtures with current store contracts * Bound E2E package setup and retain cancelled-job traces * Remove empty passing sentinels from opt-in socket tests * Make SSH typing pressure fixture readiness and replies observable |
||
|
|
306b4578aa |
Enable Option shortcuts for ABC keyboards in Auto mode (#24528)
* Clarify Option shortcut settings and cover punctuation input * Enable Auto Option shortcuts on ABC keyboards safely |
||
|
|
3f37fcc423 | test: align source-control fixtures with current store contracts (#24571) | ||
|
|
ba9af21d75 | test: classify terminal driver input with the current PTY contract (#24560) | ||
|
|
e2c5414f76 |
fix(native-chat): an older Orca keeps a chat with a newer row kind read-only instead of deleting the rest of its history (#24477)
* fix(native-chat): an older Orca skips and keeps a journal row of a kind it does not know
* test(native-chat): a newer build's journal row kind survives reads, writes, rewinds and reopens
* fix(native-chat): an older Orca keeps an unknown journal row kind read-only unless its writer declared it skippable
A row of a kind this build does not know, in a well-formed envelope, now latches the chat
read-only with every row kept, the same way a newer row version does. It is read past only
when its writer declared `ifUnknown` on the row: `skip` (a rewind drops it) or `carry` (a
rewind carries it after the rebuilt history, epoch, seq and fence restamped). Every existing
kind changes queue or turn state, so skipping by default would let an older build write from
a wrong fold.
- journal-row-kind-compatibility.ts: each kind states how older builds read it, typed over
every row kind, so a new kind cannot be added without a declaration.
- Rewind restates the Resume and Stop as before, then carries `carry` rows in source order;
the restatement goes back to { lifted, liveStop }.
- Replay treats a row whose body names another sequence than its stored key as malformed at
the key, so the next write never collides with it; catch-up reads stop there too.
* refactor(native-chat): drop the writer opt-in; an unknown journal row kind only latches read-only
An older Orca now treats a row of a kind it does not know exactly like a row from a newer
schema version: every row stays on disk and the chat opens read-only until an update. The
writer-declared skip/carry opt-in, its in-memory placeholder, the carry through rewinds and
the per-kind registry are removed: no current or planned kind could use them, and they can
come with the first kind that may safely be read past.
Kept: an unknown kind needs the envelope every row keeps (epoch, sequence, fence, timestamp),
else it is damage as before; a row whose body names another sequence than its stored key is
malformed at the key; the epoch row's validator names its kind. The schema header states the
rule for adding a kind: keep the envelope, and either ship the reader first or bump `v`.
* refactor(native-chat): derive the journal's known row kinds from the row union
Each kind's own-field check now lives in one table keyed by every kind JournalRow holds, and
the set of kinds this build knows is derived from that table. A kind added to the union without
a check fails to compile, rather than latching this build's own chats read-only as a newer
build's kind. A test reads one valid row of every kind.
|
||
|
|
c9a9b8d109 | test: restore delayed PTY writes in large-paste coverage (#24556) | ||
|
|
b666d07117 | test: update worktree setup and enforce cleanup results (#24552) | ||
|
|
1fbfb13e0f | test: isolate seeded Git repositories per Playwright worker (#24550) | ||
|
|
0b7b9a9af5 | test: isolate session fixtures and wait for completed indexing (#24544) | ||
|
|
026b8378a4 | test(wire): make release compatibility probes deterministic (#24538) | ||
|
|
11b1c8f353 | test(e2e): dismiss the browser tour before starting screenshot markup (#24534) | ||
|
|
444f1952c7 |
ci: run every cross-version wire test, picked up by folder so new ones can't be skipped (#24499)
* ci(cross-version-wire): run the whole directory so no compatibility test is left out Three cross-version tests ran in no CI job because the job named its files by hand. Run the directory instead, ratchet that every file kept out of the unit shards runs in some PR job, and re-run the job when the modules the newly running tests guard change. * test(cross-version): give the orchestration downgrade test its siblings' 120 s budget * ci(unit-exclusion): count only merge-gating jobs, and require each excluded file's job to fire on it The coverage check counted any pr.yml job, including e2e, terminal IME and Windows WSL, which are left out of verify.needs and so cannot block a merge. It now reads verify.needs and the reusable workflows those jobs call. It also only proved that some step names each excluded file, not that the job runs when the file changes. The structured-session zsh login-shell test runs only in shell_contracts, whose path trigger matched neither it, its harness nor its subject, so a PR touching only those ran it nowhere. The check now asserts a change to each excluded file fires a gating job that names it, and the shell trigger gains those three paths. * ci(cross-version-wire): trigger on the turn-outcome vocabulary and the schema version-skew resolver A change confined to src/shared/agent-turn-outcome (the arms a newer host publishes) or to orchestration-schema-version-skew (how current code reopens a downgraded database) skipped the job whose tests guard exactly those contracts. Also corrects the publish/read direction in the turn-end comment. * test(cross-version): state why the orchestration downgrade test needs 120 s * test(ci): glob the unit tree once for the unit-exclusion coverage checks |
||
|
|
757736628f |
fix(native-chat): a paired server admits structured chat by client capability, not its own chat setting (#24203)
* fix(native-chat): a host admits structured sessions by client capability, not its own chat setting A host's experimentalStructuredNativeChat decided whether any paired client could reach agentSession.* at all, and whether session.tabs.* showed it structured tabs. That setting is the host user's own launch preference: whether a new agent opens as a chat or a terminal is decided by whoever launches it. Using it as admission control meant a client whose own preference was "structured chat" was refused on a host whose preference was "terminal", and chats opened while the setting was on were withheld from mobile once it was turned off. The gate now asks one thing: did the client advertise agent-session.structured.v1 (in-process callers negotiate nothing and are always admitted). Tab projection and restore follow the same rule. With the setting no longer gating anything, the separate cleanup gate (close, cancel, unsubscribe, release), which existed only so those kept working after the setting was switched off, is identical to the main gate and is folded into it. The settings listener that republished tabs when the setting changed is removed, since projection no longer depends on it. The host setting still picks the default for launches that start on the host itself (agent.launch from mobile, orchestration worker-start). * fix(native-chat): negotiate client-chosen launch mode so released phones and old servers keep terminals Hosts advertise agent-session.structured.client-launch-mode.v1: they admit structured sessions by client capability alone. A remote client that does not advertise it (phones released before agent.launch) asks createSupport to pick the launch mode, so the host keeps answering that with its own setting, exactly as before. Cleanup methods keep their own named gate so a future admission condition cannot make close or cancel refusable. * chore(native-chat): justify the two type assertions this change's lines touch * fix(native-chat): chats that already exist keep showing whatever the chat setting says The structured chat setting decides only what new agents open as. With it off, this machine's structured chats used to be hidden while the host, which no longer reads the setting, still reported them to the workspace activation gate, so a workspace holding only a chat opened empty. The local chat mirror and its startup restore now run whatever the setting says, the continue-after-restart offer follows the chats that exist, and the setting's copy says it applies to new agents. * test(native-chat): pin that a host advertises the client-chosen launch mode * fix(native-chat): mirror this machine's chats only where it holds them Round 1 ran the local chat mirror for everyone so existing chats show whatever the setting says. That gave every desktop a permanent session-tabs listener, which turns on the runtime's phone replication paths, plus two full session-tab censuses at startup, and made the browser client mirror its remote host a second time. The runtime now says whether it holds structured chats: its structured host is built only when saved chats were restored at startup or a client created one here, and it announces the moment one is built. The mirror, the startup restore and the continue-after-restart offer run only when the setting launches chats or the host holds some, and never in the browser client. A chat a paired client creates here with the setting off still appears at once. The chat behaviour settings show wherever chats exist, and the setting's copy says it picks what new agents open as. The toggle-off teardown this made dead is removed. * test(native-chat): record install listeners without a cast * fix(native-chat): mirror this machine's chats only once it holds one, not once its host is built Session history, resume preparation, terminal resume commands and replay-safe phone launches all build the structured host for users who never had a chat, which turned on the chat mirror and the structured-only settings rows until the next restart. The signal is now derived from the host's records (or a records file still owed its import) and pushed when the first chat is restored or created. A throwing listener no longer fails the install that fired it. * feat(native-chat): createSupport reports the saved selection a new chat on this host starts with A chat on a paired server starts with the server's saved model and options, which the desktop could not read, so its picker showed a guess. createSupport's answer, which the desktop already waits for before a paired launch, now also carries that seed as a new optional field (older clients ignore it). Create and createSupport read it through one resolver so they cannot drift. * refactor(protocol): move the Electron remote client capability list into its own module Merging main left protocol-version.ts one line over the max-lines limit on this branch. The list of capabilities the desktop advertises to a paired host moves, unchanged, into electron-remote-runtime-client-capabilities.ts, the module the next PR in the stack already uses for it; importers point there. * test(cross-version): stub the launch seed resolver createSupport now reads * fix(native-chat): the desktop tells its own host it picks each launch mode, so retrying an existing chat works with the setting off * docs(native-chat): name the real exit for the released-phone createSupport rule * test(cross-version): a released client still gets the host-setting createSupport answer; a launch-mode client gets supported plus the seed |
||
|
|
0d2300ca8e |
test(e2e): retry the crash probe's main-process reads through the transient-evaluate helper (#24479)
expect.poll does not retry a thrown read, so one spurious 'Resulting promise was garbage collected' from Electron's main evaluate failed the crash-recovery test. |
||
|
|
4e919f3b5a | test: send Codex Ctrl+C to a live raw terminal fixture (#24480) | ||
|
|
02790c53e8 |
Fix Codex status after Ctrl+C copy and side-chat navigation (#24339)
* fix: preserve Codex status on ambiguous Ctrl+C input * fix: confirm Codex turn cancellations from host rollout records * fix: keep ephemeral side hooks separate from the Codex main turn * perf: watch active Codex rollouts and skip unrelated records * fix: retain confirmed Codex cancellation across late relay events |
||
|
|
8c1670f2c4 |
test(e2e): fake Codex answers the --no-daemon --help probe without a spawn (#24440)
* test(e2e): fake Codex answers the --no-daemon --help probe without a spawn Since #23933 Orca runs `codex --help` before a path-named Codex launch. The fakes logged it as an agent spawn and held the probe for its 5 s timeout. Move the app-server refusal and --help answer into one shared FAKE_CODEX_LAUNCH_PROBES_SOURCE used by every fake Codex. * test(e2e): stop asserting the dispatch capability column a current worker no longer has #23994 stopped minting the per-dispatch capability, so capability_hash is null. The --help spawn failure used to stop this test before it got here. |
||
|
|
ccb63afc06 |
fix(native-chat): a Stop still reads as yours after Orca restarts, because the turn's end reads the Stop's event (#24311)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused
Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.
* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered
A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.
* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget
* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it
The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.
* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card
* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows
Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.
One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.
The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.
Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.
* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones
A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.
* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction
The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.
* fix(native-chat): stop creating the unused queue pause table
The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.
* fix(native-chat): a Stop's pause never hides the restart pause
A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.
Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.
* test(native-chat): pin the Stop's no-resend, lift and held-card rules
- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
again" at one instant, before a queue ignoring the pause re-sends. They
now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
whether or not a person's turn lifts it; it now reads the Stop's pause
before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
queued before a rewind.
* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller
The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.
* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event
* test(native-chat): pin that Stop and Resume rows never reach apps or count as history
* test(native-chat): only a person's Stop event pauses the queue
* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop
Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.
* test(native-chat): a card held at a starting agent is checked before the Stop's timing
Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.
* test(native-chat): a released build keeps and folds a journal holding Stop events
Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.
* style(native-chat): format the Stop event changes
* test(native-chat): type the released build's exports through one checked helper
* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only
* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade
The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.
Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.
* fix(native-chat): a Stop that stops nothing new writes no Stop event
A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.
It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.
* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop
* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled
* fix(native-chat): any later Stop event ends a person's Stop pause
A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.
An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.
* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed
A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.
A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.
Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.
* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes
A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.
The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.
* fix(native-chat): a Stop still reads as yours after Orca restarts before the turn ends
Every stop that ends work now writes the Stop's event before it ends the child: a
person's close of the chat, an eviction (worktree teardown, orchestration stop, tab
cleanup) and the idle sweep's stop of a start that never landed. A stop that ends
nothing writes nothing, and quit writes none: its resume marker records why.
The turn-end write reads the latest Stop event where every turn row is built, so the
adapter's settle, the host's fallback and the relaunch's settle all agree: a turn a
person's Stop or close named, ending with no verdict of its own after that Stop, ends
as their cancellation. A relaunch's probe-bounded end is no earlier than a Stop that
found the turn running. When the provider refuses the interrupt and the turn runs on,
a refusal row answers the Stop, so a later crash still reads Failed; pressing Stop
again after a refusal is a new Stop.
* refactor(native-chat): a stop no longer carries its cause; the turn's end reads the Stop event
The cause of a stop was threaded in memory from each entry through the host's stop
step, the adapter router and each adapter's close onto the `ended` it settled with,
and Claude kept a per-turn copy of a Stop it sent. All of that is gone: adapters
settle a turn they cut as interrupted with no verdict, the host's fallback does the
same, and the one rule where a turn row is built (`turnEndAfterStop`) reads the
journal's latest Stop event to say whether it was a person's.
- `closeSession` / `disposeSession` take no cause; `ended` has no `stopCause`.
- Claude reads an error result after a person's Stop as their cancellation from the
journal's Stop event (through the event sink), not from a per-turn slot, and a
refused interrupt is the host's refusal row, not `withdrawTurnStop`.
- An owed wind-down keeps no cause: its retry's fallback reads the Stop event.
- The mutation context's Stop passes no cause: its step already wrote the event, and
the delivery loop's child-end reason is read back from it.
- A Stop pressed before its turn showed applies to the turn that opens under it,
unless a send a person made since was accepted.
* test(native-chat): a turn a later send opened is no Stop's that named no turn
* test(native-chat): the restart test's death proof carries its detail
* refactor(native-chat): a refused Stop leaves no record; a Stop only ever ends the turn it names
The stop-refused mark is gone: its tombstone kind, its fold, the clock-keyed match that tied it to
a Stop, and the exception that let a second press after a refusal write a new Stop. A Stop that
stops nothing writes nothing. A Codex refusal names a turn that is no longer its active one, and
the Stop names that turn, so the turn running instead never reads as the person's by its id alone.
* fix(native-chat): a Stop pressed before any turn showed stops only the turn opened next
A Stop that named no turn read as the person's cancellation for every later turn that opened
after it, until a send a person made was accepted. The queue's drain, orchestration mail and a
restart continuation send as the host, so a turn they opened long after, cut by a crash, read
"Interrupted" as if the person had stopped it. The Stop now applies only to the first turn
opened after it.
* fix(native-chat): an older Claude's error end after a Stop pressed before its echo reads Interrupted
Claude CLIs before 2.1.91 end an interrupted turn with an error result that names no reason. The
translator judged whether a person's Stop explained it by its own copy of the Stop rule, which
ignored a Stop that named no turn, so a Stop pressed before Claude echoed the send read "Failed".
The translator now writes such an end as interrupted with no verdict and no error row whenever a
person's Stop may name the turn, and the journal's one rule decides as it writes the end.
* fix(native-chat): a person's Stop and /clear each name why they end the agent
The host's mutation path ended the agent with one "recorded" ending for every caller, which read
back the reason of whatever Stop event the journal held last, however old. /clear writes no Stop
event, so its end took an unrelated earlier reason. Each caller now names its own: the chat's Stop
`user-stop`, whose event its own step wrote, and /clear `user-close`, the user replacing this chat.
* fix(native-chat): a host stop judges whether it ends work after the provider's rows land
A close, eviction or host stop decided whether it ended a running turn from the journal as it
stood, while the provider's own rows (the turn its echo opened) could still be in the session's
event sink. A close landing in that gap wrote no Stop event, so the turn it cut read as news. It
now reads after the sink drains, as a person's Stop does, through the same check; a drain that
fails or takes over a second reads working.
* fix(native-chat): a Claude Stop naming a turn that just ended still marks the follow-up it cuts
A phone names the turn it last saw. When that turn had ended and a follow-up was still unechoed,
Claude's Stop interrupted the follow-up and ended the child, but the Stop's event named the ended
turn, so the follow-up's turn the child's end cut read "Failed" under "Cancellation requested.".
A Stop that ends the provider's session ends whatever is in flight, so its event now names the
live turn or none, and a Stop that names none binds the turn opened next. Codex keeps naming only
the turn the Stop names.
The Claude Stop turn-end tests move to their own file, since the session-ending Stop suite is at
its line budget.
* fix(native-chat): the idle sweep reads working by the same rule as a stop's event
The sweep judged a chat resting while a send whose reply was lost was still unanswered, but the
stop's event writer counts that send as work. So the sweep evicted it and wrote an evict event,
which ends a person's Stop pause and let the cards behind it drain on their own. The sweep's owed
work now reads the main agent working the way every session list and the event writer do.
* test(native-chat): an aborted eviction's injected drain failure lands on the eviction's own drain
A host stop now drains the session's sink once to judge whether it ends work, so the tests that
fail the eviction's drain-published step skip that first drain.
* fix(native-chat): the idle sweep's rest writes no Stop event; it evicts a send that never echoes
The previous commit made the sweep count an unanswered send as owed work, which pins a chat whose
admitted send Codex never echoes forever, and the sweep exists to retire exactly that. That rule
returns. The sweep stops only an agent it judged resting, so its eviction now writes no Stop
event, whatever send it retires: a person's Stop pause holds through it.
* fix(native-chat): stopping a start that carries no send writes no Stop event
A host stop, eviction or close of a starting child wrote a Stop event whatever the start carried.
A start with a send already reads working, so the clause only mattered for a start with none,
which ends no turn and no send: its event only lifted a person's Stop pause and bumped the idle
clock, which is why the idle sweep had been changed to close the conversation in the same pass.
The clause goes and the sweep is #24072's again. The child's end still reads host-stop, as before.
* test(native-chat): a Stop's pause across a restart is tested with a restart that writes no event
The rig's restart closes the chat with an eviction, which now writes a Stop event when work runs
and so ends a person's Stop pause. "A Stop never hides a restart's pause" then passed with no Stop
pause left to hide anything. Those tests, and the pause-lift test whose dropped assertion returns,
restart as a process that dies with no close, which like a quit writes no Stop event, and assert
that both the Stop's and the restart's pauses are in force first.
* fix(native-chat): a host stop of a turn a person's Stop is still ending keeps that Stop's reason
An eviction or host stop that landed while a person's Stop or close was already ending the same
turn wrote a newer Stop event, and the turn's end reads only the latest, so the person's Stop of
that turn read as news. A host reason now writes nothing while a person's Stop still decides what
runs: the live turn it names or bound, or, with none, the turn a send opens next. The person's
own close still writes. The E2 tests now open and end the stopped send's own turn, as Codex does,
so the mail turn after it is not the turnless Stop's.
* fix(native-chat): an older Claude's error on a later turn keeps its error text after a Stop
The translator left an error result that names no reason to the journal's Stop rule whenever a
person's Stop named the turn or none, but the rule binds a Stop naming no turn only to the turn
opened next. So a real error on a later turn read "Failed" with its error text dropped. The
translator now asks the journal's rule itself (`personStopDecidesTurn`, the one core
`turnEndAfterStop` and a host stop's in-force check share), so the two cannot disagree.
* fix(native-chat): a Stop of a start that never landed binds no later turn, whatever sent it
A person's Stop pressed while the agent starts names no turn, and the send it stopped is
cancelled before it opens one. The Stop then bound the next turn anything opened (orchestration
mail, a restart continuation, the queue's drain, all of which send as the host), so a host
eviction of that turn wrote nothing and its crash or close read as the person's cancellation. A
Stop that named no turn now binds only a turn no send journaled after it opened: any send since,
of any origin and not refused, opens its own. The E2 test's mail send is accepted as Codex
accepts it, instead of opening the stopped send's own turn first.
* test(native-chat): a rewind's restated turnless Stop binds no turn opened after the rewind
A Codex rewind restates a person's Stop still in force after the turns it keeps, at a new
sequence, so by sequence alone it would bind the next turn opened after the rewind. A send
journaled after the restated row voids that binding (the previous commit), which this pins.
* fix(native-chat): a relaunch settles a person's stopped turn with no "stopped while in progress" row
After a restart, a turn a person's Stop ended reads "Interrupted after N" with the muted mark, but
the relaunch still added the error row saying the provider stopped mid-response, which a live Stop
never writes. The settle now skips that row when every turn it interrupts is the person's Stop's
by the journal's one rule; a crash nobody stopped keeps it.
* test(native-chat): the unexpected-exit settle's journal fake answers whether a person's Stop decides a turn
* fix(native-chat): a host stop whose sink drain fails reads the journal as it stands
A host stop drains the session's sink before judging whether it ends work, and a failed or slow
drain read as working. So an eviction of an agent at rest wrote a Stop event that ended nothing,
which lifts a person's Stop pause, and a close wrote a person's event naming no turn. The drain is
now best effort: the stop goes ahead either way and only its record is at stake, so a failed or
slow drain leaves the journal's read as it stands. A person's Stop keeps its own rule.
* fix(native-chat): a Stop that named no turn applies only to a turn a send it stopped opened
A person's Stop pressed before any turn showed names no turn. It bound the first turn opened
after it, then (
|
||
|
|
976dc00337 |
fix(native-chat): Stop's pause is worked out from the chat's history, so a steered message is never re-sent (#24072)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused
Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.
* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered
A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.
* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget
* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it
The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.
* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card
* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows
Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.
One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.
The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.
Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.
* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones
A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.
* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction
The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.
* fix(native-chat): stop creating the unused queue pause table
The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.
* fix(native-chat): a Stop's pause never hides the restart pause
A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.
Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.
* test(native-chat): pin the Stop's no-resend, lift and held-card rules
- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
again" at one instant, before a queue ignoring the pause re-sends. They
now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
whether or not a person's turn lifts it; it now reads the Stop's pause
before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
queued before a rewind.
* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller
The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.
* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event
* test(native-chat): pin that Stop and Resume rows never reach apps or count as history
* test(native-chat): only a person's Stop event pauses the queue
* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop
Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.
* test(native-chat): a card held at a starting agent is checked before the Stop's timing
Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.
* test(native-chat): a released build keeps and folds a journal holding Stop events
Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.
* style(native-chat): format the Stop event changes
* test(native-chat): type the released build's exports through one checked helper
* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only
* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade
The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.
Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.
* fix(native-chat): a Stop that stops nothing new writes no Stop event
A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.
It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.
* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop
* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled
* fix(native-chat): any later Stop event ends a person's Stop pause
A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.
An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.
* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed
A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.
A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.
Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.
* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes
A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.
The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.
* refactor(native-chat): one reading of a Stop's turn for its event and its note
A Stop's event and its note each worked out the same two facts on their own:
which turn the Stop is about (the one it named, else the one running), and
whether a named turn is the one the journal shows running. The event decides
before the interrupt; the note and whether the session ends decide after the
provider's answer, so those decisions stay separate, but the facts they read
are now one helper each in structured-agent-session-turn-stop-notes.ts:
structuredAgentSessionStoppedTurnId and
structuredAgentSessionStopNamesTurnNotLive. The event's turn, the note's key,
the session-ending condition, the running-command check and the repeat check
all read them. No behavior change.
Tests: a Stop naming no turn records the running turn on its event, and
rewrites that turn's note as a Stop naming it does.
* refactor(native-chat): a failed-interrupt Stop reads its turn through the shared helper
The new branch that ends a Codex child after a failed interrupt asked
whether the Stop's turn still runs with `turnId ?? liveTurnId`, a third
copy of "the turn a Stop is about". It now reads
structuredAgentSessionStoppedTurnId, the value the note key already uses,
read at the same point before the cancel. No behavior change.
Test: a Codex Stop whose interrupt failed ends the child, holds the card
queued before it with the queue paused, and writes its Stop event before
the turn's end.
|
||
|
|
8bf90ce2c2 |
test(e2e): give the sparse preset proof room to finish on CI (#24442)
The ~60-step screenshot proof takes 2-4 s per step on CI runners and hit the 120 s default on both failing nightlies, at different steps. |
||
|
|
9912042812 |
test(e2e): select the onboarding Codex card by its exact name (#24439)
#22720 dropped the command subtitle from onboarding agent cards, so the Codex card's accessible name is now "Codex" and /^Codex\s/ never matches. |
||
|
|
b3577b7c2a |
test(e2e): drag the manual-order worktree to a slot that changes the order (#24441)
Smart sort lists the new worktrees newest-first, so dropping the source before the row that already follows it is a no-op and correctly keeps Smart sort. |
||
|
|
c6cfcc034e |
refactor(native-chat): structured chat failures always reach the diagnostics log (#24312)
* refactor(native-chat): give the structured chat host one required logger The structured chat runtime took an optional onError callback that the desktop never passed, so a late dispatch settlement, an unanswered-dispatch release, a journal event-sink write and a provider lifecycle delivery that failed were dropped with no trace. Other host failures went to scattered console.warn calls, which reach nothing in a packaged desktop build. The runtime and host now take one required logger (warn/error with a scope and fields). The production logger writes each entry as a failed span to <userData>/logs/main.trace.ndjson, which the diagnostic bundle collects, and to the console (stderr under a supervised headless host). The runtime and the host wrap it so a logger that throws never fails what it reports, and the install refuses without one. Sites that deliberately kept a recovery-capsule error out of the log still log no error object. * refactor(native-chat): hand the chat host's collaborators the logger, and give orcad its trace file The delivery loop, idle sweep, queued-message drain, lease renewer, event sink, conversation map and provider start/exit settlement each took an internal error callback that the host mapped onto the logger. They now take the logger itself and log under their own scope. The event sink keeps one onFailed hook, which decides whether to stop the provider, not whether to report. The dead-generation settlement returns its failure so each caller logs it under its own scope. orcad now installs the desktop's local trace sink under its own data root, so a headless host's chat failures reach <data-root>/logs/main.trace.ndjson as well as stderr. Also passes the logger in the test fixtures the first commit missed, which tc:node caught. * fix(native-chat): keep repeated chat failures from flooding the trace file, and record their causes - The production structured-chat logger writes a repeated failure (same level, scope, session, message and error text) once per 5 minutes, carrying how many repeats it swallowed; the tracked set is capped at 256. - Trace entries now carry the error's code (and SQLite errcode) and up to three causes by name and message. - A chat read whose conversation will not open is logged through the host's logger (open-for-read), and so are the runtime's chat-tab bookkeeping failures that already hold the host. - orcad writes its own orcad.trace.ndjson, closes it after every quit handler, and flushes it on process exit; a trace file that cannot be opened leaves tracing off instead of stopping the app or orcad. - Tests: the desktop wiring test proves the logger reaches the trace sink, and the privacy tests read every level the logger received. * fix(native-chat): log a created chat's tab-publication and launch-prompt failures through the host's logger * fix(native-chat): key a repeated chat failure on everything its entry writes The repeat suppression keyed on the message and the error's text, so two refusals with the same code but different causes, a plain error and a refusal of one code, or two object-valued errors shared a key and the second was swallowed for five minutes. The key is now the entry's whole written content (fields, code, errcode, refusal reason, cause chain, a stable rendering of a non-error value) plus the error's name and message; a refusal's reason is also written. * test(native-chat): pin that an error's name keeps two repeated failures apart * test(native-chat): build the refusal in the repeat-key test as the wire does * fix(native-chat): read an error's code and a refusal's reason by narrowing, not Reflect.get |
||
|
|
2a83c9536f |
ci(daemon): gate PRs on daemon protocol crossing from the newest release (#24089)
Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working tree must attach the newest release tag's daemon. Rollback crossing is reported only. Runs in the cross-version-wire job, which already has full tags; tag selection moves to config/scripts/stable-release-tags.mjs so both use one rule. Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
c0ac2b1fcc |
feat(browser): add a rebindable shortcut for Annotate page element (#23879)
Fixes #23470 Co-authored-by: fruit <200041037+guozi-lab@users.noreply.github.com> |
||
|
|
0b79720c2e |
feat(native-chat): the chat strip and the sidebar read the host's child records (#22614)
* feat(native-chat): publish the host's child records to the status summary and the chat strip
The status summary and the background-task channel now read a session's child
records from the host's canonical store, through the status sink its row
landed in, and derive the legacy task and subagent shapes from the same views.
The parent row folds its child-work liveness from those records at ingest,
not from the summary's task list. The adapters no longer push their task DTO
to clients: the onBackgroundTasksChanged path is gone, and a child-work ingest
is what republishes both the summary and the strip. Finished children stay
listed until the session's own next turn starts. A reader that predates child
views never receives a roster whose rows are all settled.
* feat(native-chat): the chat strip reads the host's child records with its parent's verdict
The strip's roster now renders from the child views its channel carries, and
passes the verdict the session's own status row gives its children, built the
way the sidebar builds it (the row's freshness and the status feed's
observation). So one child reads the same in the strip and the sidebar, live,
after the transport drops, and once the row goes stale. A roster of finished
children stays shown until the next turn but no longer animates the monitoring
indicator or blocks conversation commands.
Tests: an end-to-end run on a host with no renderer (a real hook server as the
status sink) shows the summary and the strip channel carrying the same records
at every step, the parent row folded from them, retention, and an older
client's task list holding live work only; a wired renderer test shows both
surfaces agree when live, lost and stale.
* test(native-chat): a newer host's view degrades, an older reader keeps its live roster, a finished roster holds nothing open
- The view decoder ignores unknown keys, degrades unknown kinds, states,
outcomes and memberships, and drops only rows it cannot identify.
- At the RPC boundary a reader that predates child views gets no strip for a
roster of finished children and never the views themselves; a stop-only
reader keeps rows whose host offers no targeted stop.
- The strip shows finished children without reading them as live work.
- The row keeps its child list's identity when a summary repeats it.
* refactor(native-chat): the summary's task list is the legacy projection's live rows, unfiltered
* test(native-chat): type the switch tests' mocks instead of asserting them
* test: remote clients advertise reading child views
* docs(agent-status): the structured row folds the store's child records
* fix(agent-status): keep the view reader's header from reading as a value import to the renderer boundary
The renderer node-builtin boundary test scans raw text, so a header comment
that said "imports" ahead of the import block turned the type-only import of
agent-status-child-work into a value edge that reaches node:crypto.
* refactor(native-chat): the status summary's broadcast equality gets its own module
The status feed crossed the file-size limit once the summary gained the main agent's turn
outcome beside the child views. Which summary changes reach every session list now lives in
structured-agent-session-status-summary-equality.ts.
* fix(native-chat): command admission reads the strip's child records
A conversation command was refused on the provider tracker's own roster
while the strip read the host's child records, so a drift between the two
rule sets could refuse /clear with a stop instruction the strip had no
button for. Admission now reads the same records through the same read as
the strip, uses the strip's liveness fold, and asks for a stop only when
the strip renders one. The adapter contract no longer exposes the tracker
roster, so no host decision can read it.
Also records when the legacy child shapes die, every earlier death of a
settled child, and the display-precision invariant behind the summary's
clock tolerance.
* refactor(native-chat): command admission takes only what it reads of a turn
* fix(native-chat): the session list drops a session's children when the store does
A session's end no longer removes its child records: a child still running
settles with an outcome nobody reported, and a finished one stays listed.
Records now leave only at the session's own next turn, at the cap on settled
records, or when the host lets go of the session and its row leaves the store.
The summary kept after the host lets go used to strip its children on close,
a rule of its own. It now re-reads them from the store when the row leaves,
through the same read every live summary uses, so the session list and the
chat strip list the same children at each step, including a forget with no
close. Closing only revokes ownership, as before the child records existed.
* test(native-chat): write the Codex frame script's parent row out step by step
Once every surface reads the child records, the provider tracker's roster is
no oracle: it and the records read the same child executions, so agreeing
with it cannot catch a defect in either. Each frame now states the child
liveness and the parent row it must fold to.
* fix(native-chat): the idle sweep and the restart snapshot read the host's child records
The idle sweep (keep an agent running while its subagents or commands run) and
the restart-resume snapshot (what a chat was doing when Orca stopped it) both
read the provider tracker's roster through the adapter interface, which no
longer carries it. Both now take the host's one child-record read, the same one
the status summary, the chat strip and command admission use.
The snapshot's working test also folded that roster through the shared fold's
old `backgroundTasks` input, which the fold no longer reads, so a settled lead
whose subagent was still running would have been offered nothing. It now hands
the fold the records.
* test(native-chat): the child-record tests follow the merged command lifecycle
A command is a live child record from its start and is removed, not settled,
when it stops, whatever Codex tagged it. The end-to-end switch now shows the
child's `npm test` as a live row beside its dev server, and both are gone once
they exit; only the finished subagent stays listed until the next turn. Letting
go of the session is its tab closing, since a closed conversation whose tab
remains keeps its row.
Command admission's finished row is a subagent, the one kind that settles, and
the failed-verdict row test admits its live subagent as a host record, the only
thing the row folds.
* refactor(native-chat): the status feed's journal projection cache gets its own module
The status feed crossed the file-size limit once the child records joined the
agent-start signal and the completion feed's status read. The per-journal
projection, cached per commit, now lives in
structured-agent-session-status-journal-projection.ts.
* test(native-chat): the admission test's compaction resolves with a real outcome
Main's compaction result is a tagged outcome; the host-level admission test
resolved its mock compaction with an empty object.
* feat(native-chat): the sidebar lists running subagents; the strip, running then the newest finished
The host keeps every child record; what each surface lists is picked from them on every read, so
nothing is stored twice. The status summary, which every session list reads, now carries only
running children (and a finished one whose shell still runs, which reads monitoring): a finished
or failed subagent leaves the sidebar and stays in the chat's strip. The strip lists every running
child, then the newest finished ones, 100 rows in all; more than 100 running all show.
This matches common practice: sidebars show live subagents, and finished ones stay in the chat's
panel, newest first. No wire field is added. An older client reads fewer rows: its legacy task
lists were already live-only in the summary, and the strip's settled tasks come from the same
bounded roster.
* fix(sidebar): one rule for what the worktree sidebar lists: running children, from every source
`worktreeSidebarListsChild` is the one definition: a child that runs, counting a finished one
whose own shell still runs (it reads monitoring). The sidebar's row builder applies it to every
child source it reads, a terminal agent's hook roster and a chat session's records alike, and the
host's status summary applies the same predicate, so the sidebar's payload stays small. The chat's
strip keeps finished children, newest first.
A terminal agent's hook roster already drops a child on its own stop, so nothing changes there:
a teammate between turns and a child gone quiet still run, and still show. The selection module
moves to `agent-child-work-listing.ts`, since it now covers every source, not only chat sessions.
* test(native-chat): the switch test passes the startup child key main's status bar takes
* fix(native-chat): a finished child stays until the user's next send, not a turn Claude opens on its own
Claude wakes the agent on its own when a background task ends, and that wake is a
new root turn. Keying retention on the newest root turn retired every finished
child about two seconds after a background agent or shell finished, so its
outcome never showed in the strip.
Retention now keys on the user's newest send the provider accepted (a message, a
steer or a command), with the journal epoch so a rewind still retires. A turn the
provider opens itself and a subagent's turn carry no send. Replayed captured wake
orders through the real adapter, hook server and status feed.
* fix(native-chat): a background Stop reaches the tasks the child records show
The strip draws a row's Stop, and /clear, /compact and rewind wait for background
work, from the host's child records, but the Claude adapter still resolved which
tasks a Stop reached from its own tracker's roster, and refused to stop at all
once that roster was empty. A task the records kept live after the roster dropped
it showed a Stop that sent nothing and blocked those commands until the chat tab
closed.
The host now resolves the provider ids a Stop sends from the records (the same
per-row rule the strip and admission use; every such row for stop-all), and the
adapter stops exactly those, with no tracker guard. An acknowledged stop ends the
record: a running task sends its own stopped frame first, and the CLI answers
success with no frame for a task it no longer knows. A refused stop leaves the
record live. No production code reads the tracker's roster any more.
* fix(native-chat): one rule for a finished child that still owns live work, at any depth
The listing kept a finished child whose work ran through any depth of ownership,
but retention at the user's next turn protected only the direct owner, so a
finished agent whose finished subagent still ran a shell was removed and that
subagent jumped to the top level. Both now read settledOwnersOfLiveWork.
* fix(native-chat): an older client sees a Codex child's shell as it did before views
Clients that predate child views read a flat task roster derived from the views.
It listed a Codex child agent's shell as an extra row beside the running agent,
then as a bare command once the agent finished. The derivation now hides a
running agent's commands and names a finished agent's as "<agent> — <command>",
as the Codex tracker did; the label rule moves to a shared module both use.
* fix(native-chat): the chat decodes a roster's child rows once, as the frame arrives
The client reducer compared raw wire rows, so a row shaped by a newer host could
throw there, and the strip decoded a new array on every render, which defeated its
grouped-rows memo while a turn streamed. Rows are now decoded where the frame
enters the reducer, an unchanged roster keeps its identity, and the strip's parent
context is rebuilt only when one of its values changes.
* fix(native-chat): the strip channel forgets a closed conversation's roster
It kept the last roster fingerprint of every conversation for the host's lifetime.
The conversation map now tells observers when one leaves it. Also corrects the
summary's children comment: it carries running children only.
* docs(native-chat): rewrap the retention comment
* fix(native-chat): a task's own ending replaces a Stop's, and a child finished after the user wrote stays
Two lifecycle gaps from the round-1 fixes.
A Stop acknowledged ahead of the task's own ending relabelled it. The SDK hands
Orca a control answer as soon as it reads it and queues other frames, so when a
task finished just as the user pressed Stop, the acknowledgement arrived before
the task's completion the CLI wrote first, and the task read "Stopped" with its
result lost. An acknowledged Stop now ends a record provisionally (outcome basis
`stop-acknowledged`); the task's own terminal frame replaces it, and nothing
replaces an ending the task reported itself.
A finished child still vanished with no new action from the user when the send
the provider took was written before the child finished: a steer Claude takes at
its next boundary, or a queued draft handed over at the end of the turn. The
user's next turn now carries when they acted (the send's written time, or the
draft's queued time, both on the host clock), and only children that finished at
or before it retire; a later one stays until the user's next send. A rewind
still retires every finished child.
Also: a Stop-all keeps stopping the remaining tasks after one request fails, then
reports the failure.
* fix(native-chat): the strip keeps one empty list for a roster that omits one
A roster with only running or only finished rows made a new empty array on every
render, so the strip regrouped its rows each time. One shared empty list keeps
its memo.
* fix(native-chat): a strip row whose owner the 100-row budget cut renders under the main agent
The budget can keep a finished child and cut its finished owner. The child still
named that owner, so it rendered nowhere. The selection now clears an owner it
did not keep, as the view contract says for an owner outside the projection.
* fix(native-chat): a stop-all that times out stops asking, and "no children" is sent once
A stop-all kept asking after a request timed out, so a Claude CLI that stopped
answering control requests cost one full deadline per task, while the chat's
sends, Stop and /compact waited behind it. A timeout now ends the loop; other
request failures still let the remaining tasks be stopped.
The chat strip channel never remembered that it had sent "no children", so
every change in a chat with none re-sent that frame to each subscriber. It now
remembers it, and forgets only when the conversation closes.
Also renames agent-child-work-stop.ts to agent-child-work-stop-targets.ts, which
says what it answers: the provider ids a background Stop reaches.
* fix(native-chat): the status feed reads a journal snapshot with no submissions, and e2e tests use the child-work reader
CI on
|
||
|
|
24edf0f64b |
fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm (#23467)
* refactor(native-chat): remove the unused terminal handoff
No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.
* fix(native-chat): never let the pre-stop snapshot hold a chat's stop
Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(native-chat): drop helpers only the terminal handoff called
`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.
Co-Authored-By: Claude <noreply@anthropic.com>
* docs(native-chat): stop citing the removed handoff in lifecycle comments
Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): type the stalled snapshot drain without a cast
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): pin that a start dead before proving owes no settlement
The removed restart handoff test pinned this branch; nothing else did.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(native-chat): keep the owner-status read behind an in-flight attach
The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(terminal): remove the agent-session PTY write gate
The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(native-chat): drop the transcript helpers only the handoff called
appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(native-chat): stop calling a starting chat "mid-handoff"
A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): type the stand-in roster decoder without a cast
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(codex): name the pinned rollout lookup for what it does
With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.
* refactor(native-chat): type the owner-status reply as the host sends it
The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.
* refactor(native-chat): normalize terminal-handoff lease values once at decode
Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.
The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:
- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
`conflicted`, the claim every build probes but never stops. A plain native
owner would be stopped by restart recovery, here and in older builds.
Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.
The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.
* refactor(native-chat): stop threading the owner kind through a reservation
A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.
* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else
Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.
* fix(native-chat): name a chat write by its target, not the owner generation
A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.
Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.
Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.
* fix(native-chat): every journal append reaches the chats that are open
A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.
A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.
* test(native-chat): an epoch replacement reaches the open chat
* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map
* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite
The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.
* test(worktree-activation): restore the OMP surfaced-agent resume test
The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.
* perf(native-chat): a publish behind a delivered commit reads nothing
Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.
* test(native-chat): state why the teardown test's fake journal is safe to cast
* docs(native-chat): say mutation admission checks only the writer lease
* docs(native-chat): drop the send rebase from comments that still described it
* fix(native-chat): a message is accepted, then delivered
A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".
A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.
Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.
A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.
Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.
* fix(native-chat): settle queued messages only for the child that ended
A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.
A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.
The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.
* fix(native-chat): an adoption that fails to import keeps the conversation open
The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.
* perf(native-chat): the recovering open reads the journal once
Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.
* fix(native-chat): an attach that fails after indexing its child leaves no child behind
A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.
* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer
The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.
A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.
* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down
The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.
* fix(native-chat): a message rejected while its chat was closed reads as not sent
A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.
* test(orchestration): name why the readiness settlement fakes are cast
* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent
* docs(native-chat): drop the fence from the admission the send effects run behind
* docs(native-chat): give the fence move on release the reason that still holds
* docs(native-chat): stop citing a write fence check in launch and mailbox comments
Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.
* refactor(native-chat): the provider child is its own record
A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.
- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.
* fix(native-chat): the delivery loop alone settles a message its start or child failed
A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.
- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
reads how it ended: a Stop continues; anything else writes one failure row and rejects every
queued message with the same words, then stops. A child still starting whose start the adapter
says did not land fails the same way. The exit, eviction and the settlement retry only settle
the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
closed, with or without a child, and a start the loop already has in flight is waited for so the
child it produces is stopped rather than left behind.
* refactor(native-chat): a stopped child ends on the one reading of its stop
The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.
* feat(native-chat): the host says it accepts a send before any agent has it
The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.
* refactor(native-chat): an attach never opens a journal of its own
The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.
* fix(native-chat): a moved fence resends nothing on a host that accepts first
The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.
The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.
* refactor(native-chat): a child's end says whether the user or the host stopped it
The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.
* fix(native-chat): a chat whose only work is a queued message is not offered for resume
A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.
* test(native-chat): type the queued-message fixtures in the resume-offer tests
* fix(native-chat): a start that dies while a message waits on it is that message's failed start
Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.
* fix(native-chat): a request that failed reads as failed
A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.
The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".
* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now
* test(native-chat): a verdict change republishes the mobile status projection
* refactor(native-chat): the store's retention trigger keeps its flag compare
A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.
* test(native-chat): a user message the provider journaled keeps its session listed
* test(native-chat): pin what a failed start settles, and what a resume offer names
A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.
* test(native-chat): the failed-start pins fail on what the message became, not on a timeout
* fix(native-chat): a late provider-session update keeps a failed recovery record failed
A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.
* test(orchestration): the preamble's host stub is typed, not cast
The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.
* test(native-chat): the terminal-bell check asserts the renamed verdict field
The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.
* fix(native-chat): a failed turn ranks like a completion for attention
Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.
The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.
* fix(native-chat): a failed main agent reads failed while its subagents still work
The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.
Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.
worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.
* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it
The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.
* docs(native-chat): the status-store listing rule names provider-journaled user messages
* fix(native-chat): a refused send notifies failed through the completion feed
The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.
* fix(native-chat): every copy of a row carries the main agent's own status
History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.
- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
rebuilding one; the sync key and history equality compare it.
* test(native-chat): pin the worktree ps verdict across host and phone versions
Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.
* test(mobile): name the parity table's row for its role
* fix(native-chat): a request that settles while the user is asked something notifies once
The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.
The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.
* fix(native-chat): the completion says when the user is being asked
A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.
The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.
* fix(worktree-status): a departed agent's failure yields to live work on the worktree card
A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.
* docs(agent-status): a departed agent's failure ranks below live work on the worktree card
* fix(native-chat): a view never restarts a chat whose last start failed
A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.
* test(native-chat): start the child the loop waits on with an attach, not a second view
A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.
* fix(native-chat): settle a gone generation's turn wherever a conversation opens
A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.
* test(native-chat): prove the next child's start settles the turn an earlier child left
The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.
* test(native-chat): count a failed start's rows by row, not by text
Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.
* test(cross-version): load the phone row readers without mobile's toolchain
Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.
The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.
* test(cross-version): keep the checkout path-guard message and justify the copy import's cast
* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget
* test(native-chat): pin the open's and the send's start and row counts, however the view binds
Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.
* fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm
When the provider gave no verdict, the structured status projection now derives
one from the newest turn's lifecycle: interrupted -> interruption, unverifiable ->
unconfirmed. Nothing new is journaled, the completion feed stays provider-only, and
the legacy interrupted flag stays a user stop only. Every verdict reader handles
both arms explicitly.
* fix(native-chat): settle a gone generation's turn at every open but an acquisition's
The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.
* fix(native-chat): a folded turn a crash cut off reads Interrupted after N
The settled-turn timing now carries the turn's verdict, derived by the same
agentTurnVerdict the status row uses. A turn that ended interrupted with no
provider verdict heads its fold 'Interrupted after N'; a user's stop keeps
'Worked for N'.
* test(native-chat): hold the create's start open until the views bind
The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.
* fix(native-chat): the user's close of a chat records the turn it cuts short as their cancellation
The expected-close settle writes outcome cancellation when the user aimed the
stop at this chat: a Stop while the agent starts, the chat's tab closed (the
agentSession.close RPC, or session.tabs.close with a user reason), or /clear.
A quit, an idle eviction, a worktree teardown or an orchestration stop leaves the
turn with no verdict, so it still reads Interrupted.
* test(native-chat): a Claude turn a newer send superseded reads Interrupted
The supersede fires for any send Orca dispatched, the user's or another agent's,
and nothing at that site records the sender, so the turn keeps no verdict and
folds as Interrupted after N.
* fix(native-chat): the user's close records cancellation on the turn the provider settled on its way out
The Codex and Claude adapters settle their open turn as interrupted, with no
verdict, while the host stops them. The expected-close settle then found no
running turn, so a user's close of a mid-turn chat read Interrupted. The stop
now reads the running turns before it reaches the provider and records the
user's cancellation on each one it cut short, unless the provider gave a
verdict of its own. The close-verdict test's adapter now settles its turn on
close the way the real adapters do.
* test(native-chat): update the close and settled-turn expectations for the host-observed verdict
agentSession.close now passes the user's word to the host, and a settled
interrupted turn with no provider verdict carries `interruption`. Also merge a
duplicate import the code-quality gate rejects.
* refactor(native-chat): drop the composer's second error formatter
After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.
* test(native-chat): pin the reason on a message rejected while its chat was closed
The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.
* fix(native-chat): the user's close cancels a turn whose start landed as the provider stopped
The close read which turns were running before the stop. A turn whose start was
still in flight (a send echo not yet journaled) was absent from that read, so the
provider's verdict-less settle on the way out left it Interrupted. The close now
reads which turns were already over instead, and records the user's cancellation
on every other turn the stop left running or interrupted with no verdict. A turn
cut off earlier, or finished during the stop, keeps its end.
* fix(native-chat): a send the provider never received after a restart has no verdict
Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.
* refactor(native-chat): the adapter settles the turn a stop cuts with the stop's typed cause
The host hands its stop's cause to the adapter's close. Each adapter settles its own
open turn on 'ended' through one mapping, turnVerdictForChildEnd, and the host's
dead-generation fallback uses the same mapping for any turn no adapter settled. A
user's close or stop of this chat is their cancellation; a quit, eviction, teardown
or an exit the adapter saw first is news.
Deletes the snapshot-and-diff reconstruction (endedTurnItemIds,
userStoppedTurnRevisions, settleUserStoppedTurns) and the requestedByUser flag.
host.close now takes a required cause.
* test(native-chat): a user's close drops the chat's status row like an eviction
* test(native-chat): expect the eviction cause on the host closes of idle release, worker stop and worker discard
The stop's typed cause now travels into host.close and the adapter's close, so these
three non-user closes assert the 'evict' they pass.
* refactor(native-chat): every stop names its cause, so none defaults to the user's cancellation
stopStructuredAgentSessionAgentUnderSerialize defaulted its ending to 'user-stop', which now
settles the cut turn as the user's cancellation. Every caller already passes a cause; the
parameter is now required, and a type-level test fails to compile if the default returns.
* test(native-chat): pin who a chat's session.tabs.close speaks for, older clients' reasonless close included
The mapping lived inline in a type-unchecked file, and only the explicit user reason had a test: an older client's reasonless close, or a lifecycle echo read as the user's, stayed green. It is now one exhaustive, type-checked function with a case per reason.
* style(mobile): draw the unconfirmed dot in the theme's status amber, not an inline hex
* docs(agent-status): the main agent's outcome also carries the host-observed end, interruption or unconfirmed
* fix(native-chat): a chat the user closed while its agent started is not a failed start
A still-starting child the user's close cut counted as a failed start, since only 'user-stop' was excluded: a start-failure row, and queued messages rejected as a provider failure. Whether an ending fails its start is now one exhaustive switch, shared by the failed-start read and the delivery loop's handover: a user's stop or close never does; an exit, a failed attach and the host's own stops still do.
* fix(native-chat): a chat the user closed closes its queued messages, and starts no agent for them
After
|
||
|
|
5cda0f4508 |
refactor(native-chat): keep agent-session records in the chat journal database (#24006)
* fix(native-chat): report a failed startup chat reconcile instead of failing app startup At startup the chat host re-checks every saved chat's lease and writes the result to agent-sessions.json. If that write failed (the file lock gave up, the file could not be written, or the file was written by a newer Orca and is read-only here), reconcileRestartLeases rejected, the startup IPC call rejected, and the renderer fell into its degraded "Session restore failed. Changes won't be saved until restart" mode. The reconcile is bookkeeping: a lease left unreconciled grants no writer, and every attach, send and read of a chat reconciles its own lease again. So the startup reconcile now reports its failure through a new optional host dependency, onStartupReconcileFailure, and resolves. The runtime routes it to its onError sink under the scope structured-agent-session-startup-reconcile, or logs it when no sink is installed (the desktop installs none). * fix(native-chat): read restored chats without waiting on lease bookkeeping With native chat on and a chat tab open at quit, the renderer's startup also awaits the chat tab restore (session.tabs.listAll). That restore re-ran the lease reconcile before reading each chat and rethrew its store failure, then recorded each restored tab as visible through a store transaction that throws on a held lock or a read-only store. Either one failed the restore, so startup still fell into "Session restore failed". Reading a chat grants no writer, so the reconcile startup and the restore run is now a reader's: createReaderReconcile never throws, answers whether every lease is settled (recovery is resolved only then; the journal opens either way), and reports each distinct failure once until a reconcile settles. Attach and agent start keep the strict reconcile. The restore's tab republish logs a failed visibility write and still publishes the tab, since a client drops every unpublished chat tab; user-driven publishes still refuse. The host dependency is renamed onLeaseReconcileFailure (scope structured-agent-session-lease-reconcile), since it now also reports for reads. * fix(native-chat): keep every record-store write off the startup chat read path Round-2 review found two more writes on the startup chat restore that could still fail it and put the app into "Session restore failed": republishing a /clear replacement recorded its tab visibility strictly, and resolving a chat's recovery rethrew its store error. The restore also paid one lock wait per tab and per batch of chats while the lock stayed held. The restore now derives tabs from state it already holds: - publishStructuredAgentSessionTab splits into the strict write and projectStructuredAgentSessionTab, which only updates the runtime's snapshot. The restore and /clear replacements only project: a saved tab index already lists every restored chat, and a /clear moves the tab in the same write that commits it. visibilityWriteMayFail is gone. - Chats a legacy profile restores that the index does not list are recorded in one best-effort transaction (store.showSessionTabs), so a failure leaves the index absent to seed again rather than partial. - The read restore's recovery resolution is caught and reported through onLeaseReconcileFailure, deduplicated with the reconcile's reports. - Once lease bookkeeping fails in a restore pass, the rest of that pass skips it, so a held lock costs one wait for the startup reconcile and one for the restore, however many chats are open. User actions (create, reveal, attach, send, the /clear commit) keep their strict writes. * test: open, seed and read the agent-session record store through one harness Tests that open the durable agent-session record store, seed it, or read back what it persisted now go through agent-session-record-store-test-harness.ts instead of calling AgentSessionRecordStore.open or touching agent-sessions.json themselves. A later change that moves the store into the chat database then changes the harness instead of every test. No production code changes. Tests whose subject is the JSON file itself (its .bak recovery, salvage, schema versions, permissions, and what older builds read back) keep reading and writing the file directly; the storage move rewrites or deletes them. * fix(native-chat): start each restore pass from one lease check and stop its bookkeeping at the first failure The restore now runs one reader lease check for the pass and lets each chat re-check and resolve recovery only while the pass is still settled. The first refusal or failed write clears it for the rest of the pass, and every chat is still opened for reading. With another process holding the lock, startup waits on it once in prepare and once in the restore, however many chats are open; a legacy profile waits once more for its tab-index seed. * docs(native-chat): correct restore comments and a test name to match the final design * test: address the record-store harness by the host's state directory The harness took the store's own folder, so each caller picked one (join(root, 'store'), or 'agent-sessions' where a test read the store the runtime owns). A later change that moves the store into the state directory's journal database could not tell those apart, and would have had to edit every caller again. Every harness function now takes the state directory, the one the test's journal database and recovery capsule already live in, and keeps the store in the same subfolder the runtime uses. Callers pass that directory; store-only tests pass their temp directory unchanged. Format tests that share a directory with harness calls take the file path from testAgentSessionStoreFilePath. The folder name moves from a private constant in the runtime to AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness shares it without importing the runtime. Its value and every path built from it are unchanged. * refactor(native-chat): keep agent-session records in the chat journal database The record store's records, operation ledger, retired claim keys and chat tab index become tables in agent-session-journal.db (user_version 4). The version-4 migration copies agent-sessions.json in its own transaction and never writes, renames or deletes that file or its .bak. Each store write is one journal transaction over exactly the rows it changed, checked with the load rules; the file lock, the external-change refresh and its hash, the .bak rotation, salvage and the hot-path recovery fence are gone from the store. * wip: importer tests * test(native-chat): cover the records migration, the import, row writes and read-only records * docs(native-chat): retire comments that describe the records file as the live store * test(native-chat): drop the record-store harness's leftover file path and type the import fixture * test(native-chat): let the host harness cleanup wait out a recovery-offer read's lock * fix(native-chat): let Stop reach the agent when its ledger row cannot be written Stop's operation-ledger row now shares the database with the chat history, so damage, a full disk or a stranded transaction on that write refused the Stop before the interrupt. A cancel plan now takes its decision from the committed ledger in memory, runs without settling, and warns that the row was skipped. Other mutations answer proven damage with the typed "Unable to load this chat." refusal instead of the raw SQLite error. * fix(native-chat): answer whether a profile holds chats from the database's rows Every host install creates agent-session-journal.db, chats or not, and the version probe created it too, so its mere existence made every profile that ever installed the host wait on host install and reconcile at startup. The check now opens the database read-only and looks for a record or tab row, lets the records file answer while its import is still owed, and counts an unreadable database as present. The version probe no longer creates the file. * fix(native-chat): open a chat from history when its tab index cannot be written Over records a newer Orca wrote, every write is refused, so opening a closed chat from Agent Session History failed on the tab-visibility write and the chat read as unreachable. Like closing a tab, opening one now reports a failed restore-index write and still publishes the tab. * fix(native-chat): keep the records import owed when the backup read fails transiently A torn records file whose .bak could not be read (EACCES, EIO) was reported as unusable, so the migration completed with nothing copied and never retried. A non-ENOENT read failure of either copy now carries its cause, which the importer classifies as a read that can clear. * test(native-chat): pin that an unreadable records file never falls back to its backup * fix(native-chat): restore imported chats' tabs when the records file had no tab index A chat created while the import was owed recorded a tab index holding only itself. When the file it later imported had no index, that index still read as recorded, so the imported chats' tabs never came back. The import now clears the recorded marker in that case, and restore falls back to the profile's tabs. * refactor(native-chat): drop the unused in-transaction store write Nothing called it, and it bypassed the write queue and the read-only refusal. * docs(native-chat): say that an unusable records file is left untouched but never re-imported * refactor(native-chat): keep the provider handle chain check as main has it The chain-validation refactor has no measured need in this change. * docs(native-chat): retire lease-renewer comments that describe the records file as the live store * fix(native-chat): keep a throwing failure sink from failing the startup chat read The lease bookkeeping failure reporter called the host's failure sink directly, so a sink that threw turned a reported, recoverable store failure back into a rejected startup reconcile or read restore. The reporter now catches a sink throw and logs both the original failure and the sink error with console.warn. * test(native-chat): wait for a replaced host's restart-offer writes before cleanup A restart test replaces the host without tearing the old one down, so the old host's fire-and-forget restart-offer withdrawal could still hold the recovery capsule's lock directory when cleanup removed the test directory (ENOTEMPTY). The harness now hands hosts a capsule that tracks running operations and waits for them before removing the directory, replacing the rm retries. * docs(native-chat): retire the abandon helper's note that the store re-creates its directory * fix(native-chat): restore a chat opened while the import was owed beside the profile's chats When the imported records file had no tab index, restore fell back to the profile's saved tabs, which never list a Claude chat, and the seed then rewrote the tab table without the chat opened while the import was owed. The tab rows that chat left are now loaded as unrecorded, restore takes them together with the profile's chats, and the seed keeps their tab ids. * test(native-chat): pin that a create whose tab index write fails still opens the chat * docs(native-chat): say why restore puts chats opened while the import was owed first * test(native-chat): replace a ledger row rather than change it in place in the Send-now rerun test The record store freezes published rows in tests, so setting a row's outcome in place threw; the test now swaps in a changed copy, as its sibling cases do. |
||
|
|
9afd1101ff |
fix(orchestration): stop minting and printing the dispatch capability (#23994)
* fix(orchestration): authorize worker reports without the dispatch capability Worker lifecycle reports and questions no longer depend on the per-dispatch capability token that lives only in the agent's conversation. The host now: - ignores capability_hash/capability_revoked_at for authorization on every row and checks the exact worker process instead (ask gains that check); - refuses a report whose calling terminal is provably another orchestration party (a Run coordinator or another Dispatch's worker), treating env that names no live pane here as absent; - applies one worker-state rule locally and remotely: a stop in flight refuses, while stop_unknown and start_unknown accept and settle. Minting and printing the flag are unchanged, so an older host and older preambles keep working. * fix(orchestration): stop minting the dispatch capability Dispatches no longer mint a per-Dispatch token, and preambles, the bundled skill guide and the ask resume hint stop printing --dispatch-capability. The consumer-generation bump and delivery fence that minting carried stay, now as setDispatchConsumer. Readers that inferred meaning from capability_hash read what they meant instead: worker-show's injected stage comes from the attached consumer, and a failed start copies custody identity only when no authority was ever attached. The CLI keeps accepting and forwarding the flag for older hosts. Cancelling a Task is recorded as failed with a reason; task-list now shows that reason and the guide and task-update notes document the recipe. * fix(orchestration): name the fenced party without implying which Dispatch it owns * refactor(orchestration): one worker report rule, fence only a different party - One module owns the unproven/settleable worker states and the refusal rule; local send records it, ask and remote throw it. A stale process is worker_identity_changed on every path. - The caller fence passes the worker's own terminal when its --from handle went stale. - Document the shared-tmux-server limit; drop the dead dispatch_capability_invalid rejection member; tests assert dispatch state, not the capability column. * refactor(orchestration): drop setDispatchConsumer and the dead capability retention - dispatch --inject no longer re-points the row createDispatchContext just wrote; worker-show reports every worker-less Dispatch as context_only, since Orca keeps no record of the paste. Tests re-point through a fixture. - failWorkerStart always records when the lifecycle closed; nothing authorizes on it. - Restore the ask resume hint's echo of a passed --dispatch-capability: an old host checks it before --resume. - Move the cancellation convention to #23983. * test(orchestration): drop capability-era assertions other tests already cover * refactor(orchestration): drop the host-side capability field and no-op test fixtures - RpcRequest and the SSH bridge stop carrying orchestrationCapability; the CLI's wire field stays for older hosts. - Fixtures pass identity to createRootDispatch instead of re-pointing to the same values; drop absence checks for a flag that can no longer be produced. * test(orchestration): cover a current process whose terminal moved to another pane * chore(orchestration): finish the capability cleanup in test stubs and skill wording * test(orchestration): drop needless response casts; mark the db stub cast safe |
||
|
|
b99462ac1c |
test: retire mobile, cloud, config and e2e cases their input cannot reach (#24077)
Completes the first pass over every test area in the repository. Sweep over `mobile/src`, `config/scripts`, `cloud/`, and `tests/` (1,494 files in scope, with the 24 files under `mobile/src/test-support/rpc-recording/` deliberately excluded). 31 case declarations removed across 17 files, 2 test files deleted, 356 lines gone. What went, by pattern: - Cross-boundary replays of a shared helper. A whole mobile file re-ran `extractPendingAsk`/`parseAskFromStatus`/`formatAskAnswer`, all owned by `src/shared/native-chat-ask.test.ts`, `native-chat-ask-fifo.test.ts` and the renderer's interactive-prompt suite — one case title was verbatim identical to the owner's, and the owners' inputs are supersets. The mobile file imported the shared module directly and exercised no mobile transport, lifecycle or rendering. - A case whose input cannot reach the behavior its title names: "arms it on Android while the drawer is open", where `use-back-claim.ts` has zero Platform/OS references, so flipping the mocked OS changes only shadow styles. - Identity copiers, including one asserting `prSidebarRenderBranch(state) === state.kind` against a production body that is `return state.kind`. The function stays; it has three live callers. - A test of the runtime rather than the product: a case asserting Node's own `EventEmitter` crash contract on a bare emitter, with zero production code in the path. The guard it documents is exercised behaviourally by the case after it. - Duplicate invocations, one of them provable rather than eyeballed: with `MODULE_SCOPE_ENV_WRITER_PIN = 0`, `files.size <= 0` is strictly implied by the sibling's `expect(offenders).toEqual([])`, since a non-empty `offenders` forces `files.size >= 1`. The pin's own doc says it may only ever be decreased from 0, so it could never become a meaningful bound either. Its policy guidance survives as a comment; the file's real ratchet and its regex self-test both stay. - Expected values produced by the test's own arithmetic, and a p95 case strictly implied by a sibling that already pins exact p95 and exact max over a wider range. One production line goes: the `export` keyword on `assignmentCleanupSteps` in `cloud/apps/relay/src/assignment-cleanup-steps.ts`. The function itself stays and is still called internally; only the test-only export was orphaned. Kept deliberately: everything a gate cites, checked by case title and not only by file path; a gate-cited case that does not deliver its claim (reported instead — see below); a cross-version wire cell whose ledger is never invoked, left under the raised bar for wire coverage; and every limit, bound, quota and provenance guard. Nothing under `mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` was touched — those bytes feed a `recorderSha256` digest pinning 398 golden recordings. Verified: `mobile` vitest over the modified mobile files (8 files, 50 cases); `mobile/scripts/check-tests-typecheck-ratchet.mjs` OK (898 files in program, 125 grandfathered, none @ts-nocheck); relay suite 799 passed; `check-reliability-gates.mjs` 140 gates; both deleted files confirmed absent from the gate manifest, `cloud/package.json` and `mobile/tests-typecheck-baseline.txt`. Seven local failures were investigated and none is caused by this change: five `mobile-web-app-*-render` tests drive `playwright-core` chromium/webkit and need browsers this machine lacks, `release-checkout.unit.test.ts` needs cross-version git refs, and `e2e-worker-env-isolation.unit.test.ts` fails identically with its HEAD content restored — it recurses `tests/e2e` with symlink-following `statSync` and no depth guard. |
||
|
|
7afa4ee3dc |
test(e2e): read the tab strip's dock samples through a typed window field (#24052)
#24010's spec read them with Reflect.get, which the low-evidence lint rejects, so every PR's static analysis now fails on main. |
||
|
|
fceca5cece |
fix(sidebar): an agent's row stays while it runs, whatever its tab title (#23948)
* fix(sidebar): keep hook-less agent rows while the agent runs, whatever its title Codex retitles its pane to the project name, so the sidebar's title-derived row (which required the title to name an agent) vanished while Codex kept running (#23767). Rows now take identity from the canonical pane resolver over the pane's foreground-process read and launch record, then the title; the title only decides idle/working/needs-input. The row still goes away when the PTY exits, the process tracker proves the shell is back, or the title is a shell or default title. * test(dashboard): justify the partial store fixture's type assertion * fix(sidebar): only a live process read keeps a plain-title agent row Review of the previous commit found ghost rows: the tab launch record is a latch nothing clears on WSL, after an SSH exit, or for a launch that never started, and a parked pane's process read went stale because only the mounted tracker re-derives it. - The launch record returns to main's role: a fallback only for titles that show activity, ranked below a title naming another agent (pane reuse), matching the tab icon's order. - A parked pane's command boundary retires its unconfirmable process read, like the mounted ladder's unavailable path; reveal re-reads it. * fix(sidebar): confirm before a parked marker retires an agent; read Git Bash prompt titles as the shell - A parked pane's end-of-command marker can be a nested shell's leak under a still-running full-screen agent, so confirm the foreground first (as the mounted ladder does) and retire the process read only on a shell or no answer. SSH/remote parked panes hold no incarnation to fence a host read with, so they still retire. - Git Bash emits no command marks; its `$MSYSTEM:$PWD` prompt title (MINGW64:/c/repo) is now shell evidence, so a stale Codex read there no longer keeps a ghost row after Codex exits. * fix(sidebar): trust only process-read agents for plain-title rows; per-worktree foreground selector groups by tab A daemon reattach seeds the pane's foreground entry with its launch agent, which can outlive the process while Orca is closed. The entry now records where its agent came from (agentEvidence), and the sidebar/dashboard title-derived rows only keep a plain-title row on an actual process read. Routing and the tab icon are unchanged. selectPaneForegroundAgentsForWorktree grouped every pane key per worktree; it now groups by tab once per map identity and skips worktrees with no tabs. * fix(sidebar): a parked pane's reattach keeps its own process read of the same agent The reattach seed marked a returning parked Codex pane as launch-record evidence, over the process read this session already took, so its row blinked out on reveal and stayed hidden if the user left the tab before the visible read landed. Keep the read when it names the same agent; the seed still drops byte-routing trust. * test(terminal): foreground confirmation publishes process-read evidence * fix(sidebar): a cleared pane title retires the agent's process read Codex clears its title when it exits, and the tab then shows its default title. A pane without shell command marks never re-reads its foreground process, so the retained read kept a "Codex · Idle" row after /quit (permanently for a hand-typed Codex; about 15 s while the marked-pane confirm ladder ran). Treat a blank title like the default title it shows. * fix(sidebar): the pane's process monitor retires an exited agent's process read A hook-less pane keeps its sidebar row from the tracker's foreground-process read, but nothing re-derived that read in a pane without OSC 133 command marks. After Codex exited there, a "Codex · Idle" row stayed: permanently when the shell titles its prompt, or when a killed Codex leaves its last title. The pane's agent-completion process monitor already confirms an agent's exit (no agent and no child processes, held past its settle window). It now reports that exit to the tracker, which retires its own process read and runs the confirmed-shell path the visible-pty read uses. A tracker read that names an agent seeds the monitor, so hidden panes and panes the monitor had not polled yet are watched too. A command read in flight still decides the pane, and launch records or other agents' reads are left alone. * fix(sidebar): a monitor-confirmed exit leaves the next agent in an unmarked pane identifiable The process-exit retire published shellForeground:true and left the one-shot visible sample settled; a pane without command marks has no command start to lift either, so a Codex typed again after quitting was never read and lost its row on retitle. Publish shellForeground:false and reopen the sample. * test(terminal): justify the pane binding cast in the process-exit relaunch test |
||
|
|
9cdbeba06d |
fix(tab-bar): keep the active tab visible when the tab strip scrolls (#24010)
* fix: keep active tab visible by docking to viewport edges Makes the current tab easier to locate in many-tab scenarios. The active tab now sticks to a viewport edge via sticky positioning when it would scroll out of view, with a full-foreground indicator bar for better visibility and arrow animation when a background tab opens off-screen. * fix(tab-bar): reveal offscreen tabs instead of nudge animation When a background tab opens beyond the visible area, automatically scroll to reveal it (unless hovering the tab strip). This replaces the previous arrow-nudge animation with direct visibility. revealTabStripElement now handles keeping the active tab visible alongside the revealed tab when both fit, or docks the active tab when needed. * fix(tab-bar): reveal tabs by identity, not count increase alone Detect opened tabs by comparing tab identities independently of count changes. Newly opened tabs are now revealed even when the total tab count stays the same—e.g., when a tab closes as another opens. * fix(tab-bar): track tabs by identity for reliable reveal on open/close Replace count-based tab detection with identity tracking so the strip correctly reveals tabs when they're added, replaced, or when the active tab closes and switches to a far-back history tab. Removes the tabCount parameter and simplifies overflow navigation by using identity sets. * fix(tab-bar): defer revealing tabs until pointer leaves When a background tab opens while the pointer hovers the tab strip, defer its reveal until the pointer leaves. This prevents the active tab from sliding away mid-interaction. Also support client-hosted rows taking active state while maintaining tab dock positioning. |
||
|
|
75040eba5a |
test: open, seed and read the agent-session record store through one test harness (#23986)
* test: open, seed and read the agent-session record store through one harness Tests that open the durable agent-session record store, seed it, or read back what it persisted now go through agent-session-record-store-test-harness.ts instead of calling AgentSessionRecordStore.open or touching agent-sessions.json themselves. A later change that moves the store into the chat database then changes the harness instead of every test. No production code changes. Tests whose subject is the JSON file itself (its .bak recovery, salvage, schema versions, permissions, and what older builds read back) keep reading and writing the file directly; the storage move rewrites or deletes them. * test: address the record-store harness by the host's state directory The harness took the store's own folder, so each caller picked one (join(root, 'store'), or 'agent-sessions' where a test read the store the runtime owns). A later change that moves the store into the state directory's journal database could not tell those apart, and would have had to edit every caller again. Every harness function now takes the state directory, the one the test's journal database and recovery capsule already live in, and keeps the store in the same subfolder the runtime uses. Callers pass that directory; store-only tests pass their temp directory unchanged. Format tests that share a directory with harness calls take the file path from testAgentSessionStoreFilePath. The folder name moves from a private constant in the runtime to AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness shares it without importing the runtime. Its value and every path built from it are unchanged. |
||
|
|
c3183a4556 |
test(e2e): let the completed-worker fake Codex answer the --help probe (#24033)
#23900 probes codex --help before each launch; the fake counted it as a worker spawn, breaking two specs. |
||
|
|
b7209b5ae9 |
perf(git): relist only the repo whose worktrees changed, and stop blocking main on sync git (#23998)
* perf(git): stop blocking main on the open-on-remote git cascade
`getRemoteFileUrl` ran up to 6 sequential `gitExecFileSync` calls on the Electron
main thread — `remote get-url`, then `getDefaultBaseRef`'s `symbolic-ref` plus up
to four `rev-parse --verify` probes — each with its own 15s timeout and no yield
between them.
A complete async twin already existed (`getDefaultBaseRefAsync` ->
`resolveDefaultBaseRefViaExec`, sharing DEFAULT_BASE_REF_PROBES), so the sync
cascade is deleted rather than converted. `getRemoteUrl`, `getRemoteFileUrl` and
`getRemoteCommitUrl` become async; all four downstream callers were already async
(`filesystem-git-url-handlers` inside `ipcMain.handle`, `runtime-git-diff-commands`
async methods) and the provider contract already typed both wrappers
`Promise<string | null>`, so no new async plumbing was needed.
Removes 3 of the 10 `gitExecFileSync` sites and the confusing name collision with
the unrelated async `getDefaultBaseRef` in hosted-review-creation-git-state.
The base-ref regression tests keep their coverage, repointed at the public async
`getBaseRefDefault`.
* perf(git): resolve the repo root in one sync spawn instead of two
getGitRepoRoot ran `rev-parse --is-inside-work-tree` and then `rev-parse
--show-toplevel` as separate blocking spawns. Each sync git call holds the main
thread for up to its whole 15s timeout, so the spawn count is the cost — and this
function is called twice per "Add Project" on a linked worktree, once directly and
once through getLinkedWorktreeMainRepoRoot's self-recursion.
Combined into one invocation. Safe only here: in a bare repo the combined form
exits non-zero, and both that throw and the plain `false` already land on the same
marker-scan fallback. probeGitRepo deliberately does NOT combine — it has to read
`false` cleanly to go on and detect a bare repo, which the combined form's exit 128
would misread as indeterminate.
* perf(git): rebuild only the repos whose authorized roots actually changed
One worktree create called `invalidateAuthorizedRootsCache()`, which dirties every
registered owner. The next authorization-requiring IPC then rebuilt by listing EVERY
repo — and the rebuild never consulted `dirty` when choosing what to list, so `dirty`
gated only whether a rebuild ran, not its scope. At 58 repos that is 58
`git worktree list` spawns, roughly ten seconds of git wall-clock through an
admission budget of four, to rediscover roots one repo changed.
Both halves were needed; scoping the invalidation alone changed nothing.
- `markAuthorizedRootsOwnerDirty` dirties a single owner, reusing the per-owner
primitives `registerWorktreeRootsForRepo` already used. It leaves `baseRevision`
and the per-repo revision map alone — that pair is the global side-effect-token
fence, and bumping it would retire in-flight tokens for untouched repos.
- `rebuildAuthorizedRootsCache(store, onlyDirty)` re-lists only owners that are
dirty, have no listing yet, or still hold recovered roots (those are retired by
comparison against a fresh listing, so skipping them would strand them as
authorized). Only `ensureAuthorizedRootsCache` passes `onlyDirty`; an explicit
rebuild keeps re-listing everything because callers use it to force a refresh —
`filesystem-auth.test.ts` pins that contract.
`invalidateAuthorizedRootsCacheForRepo` wraps the primitive and falls back to the
global form for an unknown owner or a missing store, rather than silently skipping an
invalidation and leaving a stale allowlist. Applied to the worktree-create path.
Changes that can alter the owner SET (store swap, host/WSL re-routing, nested-repo
import, folder->git upgrade) stay global. Removal paths are not converted yet.
The allowlist contents are unchanged and the failure direction is a false denial
rather than a false allow. The relist predicate is split into its own module so it is
testable alone and the cache file stays inside its line budget without a suppression.
* test(perf): measure what git orchestration actually costs the main thread
The existing churn probe (ORCA_MAIN_THREAD_DIAGNOSTICS=1) reported spawn-initiation
cost for git/gh/glab only — its 7 call sites all sit inside git/command-runner — so
it was blind to `spawnProcess`/`runProcess`, the repo's own mandated wrapper, and to
the blocking `execFileSync('ps')` per PTY resize. That understated total churn across
115 main call sites.
- `spawn-observer.ts`: a settable seam, since shared code cannot import src/main.
Unregistered in the daemon/relay/CLI, where it costs one boolean check.
- `spawnProcess` brackets `nodeSpawn` and reports; exec-file-capture's own report is
removed because it routes through runProcess and would double-count.
- `posix-pty-foreground-group` now reports its full blocking duration. Note this
lands on the daemon, not main, whenever the daemon hosts the PTY.
- `ORCA_UNMINIFIED_MAIN=1` build flag, because a minified main bundle cannot
attribute CPU-profile self time to real function names. Defaults unchanged.
- `main-thread-git-cost.spec.ts` + `analyze-main-cpuprofile.mjs`: sweeps concurrency
against real registered repos, captures the churn lines and a V8 CPU profile of
main per phase.
What it found, which is why this is worth keeping: at the width-4 admission ceiling
(~90 git:status/s) main sees ZERO event-loop gaps over 50ms and a worst gap of 23ms,
and is 85% idle. Git orchestration does not stall the main thread. Of the cost it
does incur, spawn-init is 58%, parse 5%, stdout drain 4%.
* test(perf): name the inspector params type the anti-slop gate requires
The broad `object` parameter trips anti-slop(no-object-parameters); the only
Profiler call that passes params sends `{ interval }`.
|
||
|
|
fb52c0602a |
fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout (#23920)
* fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout
xterm paints nothing while DEC mode 2026 (synchronized output) is open and only
force-flushes after 1000ms. Codex wraps every draw in mode 2026, so any byte gap
or chunk split that loses the closing \x1b[?2026l freezes the pane for a full
second and then repaints in one burst.
Orca never emitted \x1b[?2026l anywhere, and three paths could destroy a TUI's:
the per-PTY pending cap drops buffered output wholesale (mode 2031 was already
salvaged there, 2026 was not), main sliced pending data at a blind 16KB offset
that can land inside an open frame or sever the 8-byte marker, and the renderer's
backlog warnings replace a queued tail that may hold the close.
- salvage the 2026 latch across dropped output, mirroring the existing 2031
salvage, and append the release on both delivery sites
- ground 2026 in RESET_AFTER_BYTE_GAP and the replay baseline, and in both
backlog warnings, so every drop path is self-healing
- make main's 16KB flush split frame-aware instead of a blind byte offset
- lift the synchronized-output scanner into shared/ so main and the renderer
use one implementation
Closing a frame early costs one premature repaint; leaving it open costs a
second of blank screen, so the asymmetry favours always closing.
Also adds the reproduction this needed: the pre-existing typing bench observes
the xterm BUFFER, which the parser fills while rendering is held, so it scored
these freezes as fast echoes.
* fix(terminal): stop the renderer's queue drain cutting inside an open DEC 2026 frame
takeQueuedChunk sliced a queued chunk at a blind byte offset to fit the 16KB
coalescing budget, which can strand a frame's closing \x1b[?2026l in the residual
until a later drain. Same defect as main's flush split, same fix: reuse the
frame-aware split helper.
Usually masked because the drain coalesces adjacent chunks and reassembles what
main split, but not when the budget boundary falls inside a frame.
* fix(relay): keep the SSH path's bounded slice outside an open DEC 2026 frame
pty-handler split pending output at a byte offset with a surrogate-pair guard but
no synchronized-output awareness, so a frame straddling the 16KB wire slice had
its closing \x1b[?2026l stranded in the remainder — the same defect just fixed on
the local path, on the path AGENTS.md requires us to consider.
Placed before the surrogate guard so that guard keeps the final say, and floored
at 2 so frame alignment can never walk a healthy slice into the guard's
decrement and then into the chunkChars <= 0 pause-and-retry path.
Also drops a dead `splitAt === 0` branch in takeQueuedChunk: both callers pass a
positive limit and the helper never returns 0 for one.
The two new split tests were each confirmed to fail without their fix.
* test(terminal): sweep the DEC 2026 split helper over escape-sequence shapes and every limit
Covers OSC 52, DCS, repeated open/close markers and limits 1..len+3, asserting the
result never exceeds the limit, never reaches 0, and stays byte-exact. Also pins
that a buffer beginning inside an open frame degrades to the blind offset rather
than doing something worse, and documents that callers do not thread latch state.
* fix(terminal): ground DEC 2026 on the daemon slice, the recovery replays, and the process boundary
Four more sites could strand the latch, found by sweeping every path that drops,
splits, or replays terminal bytes.
- daemon-stream-data-batcher: the 64KB bulk-write slice used a surrogate-only
clamp, and its remainder is HELD until 'drain' — "seconds for multi-MB
backlogs" per the file's own note. A frame straddling that boundary parked its
\x1b[?2026l behind the hold, blanking the pane past xterm's 1s timeout once per
frame for as long as the backlog lasted. This is the default daemon-backed pane
path, so it is the one users actually hit. The new
clampToSafeBulkWriteSplitIndex frame-aligns first and surrogate-clamps last,
and lives in daemon-stream-data-split alongside the policy it belongs to.
- replay-data-drain and remote-runtime-terminal-binary-snapshots wrote a bare
\x1b[2J\x1b[3J\x1b[H, which does not clear mode 2026 — so on the SSH/remote
reconnect path, the very event most likely to sever a frame, the whole replay
could paint nothing.
- ipc-pty-attach: trimIncompleteTerminalControlTail can cut a half-written
\x1b[?2026l while its opening marker survives in the replayed prefix.
- PROCESS_BOUNDARY_GROUND: the "process that armed these modes is gone" ground
omitted 2026, the last unexplained gap in that file. A disable, so it still
satisfies the recovery barrier's ownership scan (only ?25h may be an enable).
Recovery-path expectations updated where they pin the emitted bytes. Deliberately
NOT touched: apply-reattach-payload and ssh-snapshot-prepaint already ground via
buildSnapshotReplayPrologue.
Still unfixed, deferred with reason: terminal-output-frame-chunks.ts splits the
remote wire on accumulated UTF-8 byte width and needs a different shape than the
char-index helper; desktop clients reassemble in main's pending buffer, so the
exposure is mobile/web only.
* fix(terminal): emit the DEC 2026 release before the mode-2031 tail, and stop claiming the drop path writes it
Two corrections from adversarial review of the earlier commits.
1. Ordering bug I introduced. getDroppedMode2031RendererData ends with
`state.tail`, which extractPrivateModeScanTail deliberately retains as an
INCOMPLETE private-mode sequence so the next chunk can resolve it. Appending the
2026 release after it put an ESC behind a dangling CSI, aborting it and silently
losing whatever mode spanned the drop boundary. The release now goes first.
2. The drop-path release does not reach xterm in the dominant case, and the comment
now says so instead of implying otherwise. live-data-callback's droppedOutput
branch discards `data` and salvages only queries
(salvageRendererQueriesFromDiscardedRestoreData handles CPR/DA1/OSC colour;
\x1b[?2026l is not a query), so for hidden panes and visible panes outside
foreground-restore backpressure the synthesized release was dropped. The grounded
snapshot replay releases the latch instead.
I tried writing it through writePtyOutputToXterm there and reverted: it consumes
the pending hidden-output snapshot and broke
pty-connection-hidden-snapshot-resize-signals ("re-restores a skipped alt frame"),
so the release rides the restore rather than perturbing that state machine.
Residual gap, documented: a cap-dropped pane whose restore never arrives.
The salvage is still load-bearing on the fall-through path, so it stays.
* fix(terminal): release DEC 2026 on the reattach clears, floor the split, and correct the freeze framing
Remaining findings from adversarial review.
- apply-reattach-payload's three bare-clear branches (:63 daemon snapshot, :229
relay replay, :269 cold restore) had no release anywhere in their sequence: I
checked all seven POST_REPLAY_* profiles reachable via chooseReattachReplayReset
and none contains \x1b[?2026l. Only the buildMainModelSnapshotReplayWrites branch
was grounded, so covering the streamed replay path and not the main reattach path
was inconsistent. Verified no production code matches these clear strings — the
three test updates are mock equality, and each was confirmed to fail without the
source change.
- clampToSafeBulkWriteSplitIndex could return 0 (('\u{1F600}aaaa', 1) — alignment
returns 1, the surrogate clamp decrements to 0), which would leave a zero-length
slice that never shifts the batcher's queue entry and spin its drain loop.
Unreachable from today's only caller, but it is exported with an unstated
precondition. Floored at 1.
- Frame alignment could halve per-PTY flush throughput: main re-queues the
remainder with eligibleRound = round + 1, so the shortfall cannot be refilled in
the same round, and aligned size is floor(W/F)*F — 50% worst case in the 8-16KB
band, which is exactly the full-screen redraw burst that reaches the pending cap.
Alignment is now rejected below half the window, preferring throughput and
letting the reset profiles release the latch.
Framing corrected throughout: bufferRows records a row range and clears nothing, so
the pane freezes on its last painted frame — it does not go blank. The real trade is
"stale but coherent for <=1s" versus "immediate partial frame", and
RESET_AFTER_BYTE_GAP (written alone, with no repaint behind it in the same write) is
the one site that can newly flash a partial frame. Said so at the constant instead
of implying the release is free.
* fix(terminal): rename the shape-flagged symbols the anti-slop audit rejects
CI's anti-slop gate rejects "shape" in symbol names as structural rather than
domain language: `shapes` -> `outputSamples`, and
`writeCodexShapedEchoProbeScript`/`codexShapedEchoProbeScript` ->
`writeCodexEchoProbeScript`/`codexEchoProbeScript`.
|
||
|
|
bb667a33bd |
test: retire the last private-predicate duplicates in the leaked-internals sweep (#23949)
Sixth and final wave over the modules that export symbols only tests import.
Deletes private-predicate cases whose behavior is already asserted through the
module's real entry point, then makes the symbol private again.
Also removes three distinct junk shapes the earlier detectors missed:
- a self-comparison whose expected empty row was produced by the helper under
test (`worktree-palette-search`), now a literal;
- expected values computed by a sibling helper rather than asserted
(`terminal-theme`), now read through the production `getBuiltinTheme`;
- a negative control that cannot fail — `expect('json' in jsonlMonarchLanguage)
.toBe(false)`, where `IMonarchLanguage` has no such key, so it guarded nothing
while appearing to guard "does not attach the JSON language service".
Dead production code removed where tests were its only callers:
`refreshWindowsTerminalCapabilities` (a one-line alias for
`loadWindowsTerminalCapabilities({force: true})`), `readSpoolRecords`,
`buildAgentPromptSubmitBytes`, and `getCommitMessageModelCapability`.
About 70% of everything this detector flagged across the whole vein was a false
positive, so most modules were left untouched. Bounds consumed as test input,
`*ForTests` seams, non-hook cores of `useSyncExternalStore` hooks, and
value-position registrations all look identical to a leaked internal from the
outside and are not.
|