mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 08:02:21 +00:00
5cf3585b78da2dc629e77b3787f7744cb4cfdf0f
12851
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5cf3585b78 |
fix(native-chat): keep a message accepted before a quit or crash as a held card (#24660)
* fix(native-chat): keep a message accepted before a quit or crash as a held card A send the host accepted while the agent was still starting, and never handed over, was rejected unseen at quit or at the next open after a crash. The next open now keeps a person's message (typed, or a launch's first prompt) as a waiting card at the head of the queue, held until Resume, Send now, Edit or Delete; quit no longer rejects it. Each submission records its source so a restart knows which leftovers to keep. A direct send's replay answers from its own record, never the queued arm. Cards shown without the queue capability hide Turn off queueing and the steer chord. * fix(native-chat): keep an unsent message as a held card at every close, not only a restart The host now keeps a person's message it accepted and never handed over with one rule wherever it can no longer hand it over: a quit or crash (settled at the next open) and a close of the chat (tab close, worktree teardown, orchestration stop). - The hold is card state: a new per-card hold_reason 'kept', published as the existing pausedReason, instead of a fake host_instance value. host_instance means the owner again, and the pause clause, adoption filter and /clear carry special cases are gone. A kept card holds the cards behind it until the person sends, edits or deletes it; a send_failed card still does not. - A card hand-off rejected by a restart or a close returns as a kept card (rejectedDraftSettlement), so a person's next message can never release it. - One hold function, parameterized by cause (hostRestarted / chatClosed), replaces the close path's plain rejection. - The phone shows published cards and per-card holds whatever the queue capability says; only queueing a new send stays gated. - source gains 'dispatch' for the orchestration preamble (still rejected); an unknown source is kept as written and never takes the legacy rule. - Kept cards from earlier settlements stay ahead of a batch's new ones. * test(native-chat): the dispatch preamble records its source * fix(native-chat): skip a kept card like a failed one, and make a quit leave the queue as a crash does - A kept card is held on its own, as a send_failed one is: the queue sends the cards behind it, and the "a message ahead needs attention" caption no longer appears behind it (desktop and phone). - Quit disposes the queue's drain together with delivery, so it mints no hand-off that only the next process could settle. - Only a Send the person asked for (origin client) that a restart or close cut short returns kept; the queue's own hand-off returns where it stood, under the restart's pause, as on main. - Comments that said only a capable host gets cards or the card actions now say the capability gates only queueing a new send. * refactor(native-chat): one dispose gate in the queue drain's step * test(native-chat): the downgrade test names the older build's /clear exception * fix(native-chat): re-check the drain's quit gate right before it appends a hand-off A drain step already past its first check when quit begins no longer makes a hand-off. Adds regression tests for that, for Send now on an ordinary card cut short by a quit, and for the open-time repair deriving kept from the hand-off's origin; fixes the pause and settlement comments that called a kept card one the queue never passes. * fix(native-chat): a card held on its own starts no restart pause for the others After a second restart a kept (or send_failed) card is another process's, but it waits for its own Send, so it no longer pauses every other card under "Queue paused because Orca restarted". * fix(native-chat): record on a kept send which card holds it, so no trace survives an Edit or Delete The rejection row of a send the host kept as a card now names that card (`keptAsQueuedMessageId`), in the same transaction that writes the card, and the fold publishes it on the submission. The shared projection draws such a send only as its card: once the card is sent, edited or deleted, neither the send nor the sending desktop's local copy of it shows. * chore(native-chat): leave the unused submission schema as main has it No client parses published submissions with it (they arrive as typed frames, and the host's history pages carry the field, as the Edit/Delete test reads); listing the field put the file over its line budget. * fix(native-chat): retire a kept send's local copy instead of only hiding it The outbox reconcile and the send disposition drop an entry whose submission the host rejected as kept as a card, as they already do for a Stop's withdrawal, so the copy never comes back as "Not sent / Retry" once the submission falls out of the loaded page. The projection reads the reconciled outbox, so its separate filter goes. * test(native-chat): move the queued-message rig's scripted provider into its own fixture The rig fixture grew past the 300-line limit once main's changes merged in. * fix(native-chat): hide a kept send by its own record, not by its card still existing The transcript hid a rejected send while a queued card held it under its id, so the card's Edit or Delete brought back a "Not sent" row. It now reads the send's own keptAsQueuedMessageId, which the host records with the rejection, and the live card list is no longer threaded to the transcript or the delivery notices. * refactor(native-chat): move a sent message's row writes into their own journal collaborator The journal store went past its line limit once main's ledger receipt joined this branch's transaction hook. The submission and dispatch-transition writes, and what commits in their transaction, now live in JournalSubmissionWriter; the store's methods delegate to it unchanged. * fix(native-chat): list a kept send's card id in the submission schema The schema drops keys it does not list, so a reader that kept a parsed submission would lose keptAsQueuedMessageId and source, both persisted with the row. The submission schema and the failure fact it shares with item bodies move to their own modules, with room for both fields. * test(native-chat): match the transcript and outbox hook signatures main and the swap changed * test(native-chat): import the journal types once in the queue-delivery test * fix(mobile): a resend the host kept as a card shows no error and returns no text The host answers a resend of a message it kept as a card with that message's rejected submission, marked keptAsQueuedMessageId. The phone read it as any rejection: "Message not sent" and the text back in the composer, while the card showed the same text. It now answers like a queued send, as the desktop's send disposition does: the id is spent, no error, and the card holds the text. * test(native-chat): pin that quit's first step stops the queue's hand-off Quit now stops delivery, the queue drain included, at its first step (stopDelivery), before teardown drains recovery. The drain-step quit test runs from that step as well as from the flush. * test(native-chat): give cards their source and store unknown sources as another build would #25078 made a card's source required, so the tests that insert a card pass the person's. The hold's unknown-kind and unreadable-source cases now rewrite the stored row the way a newer build would leave it, instead of casting a type. * refactor(native-chat): stop delivery and the queue drain in one line at quit Main's #25159 left the host at its line limit; quit's stop now disposes both in one expression instead of a block. * test(native-chat): hold the drain step without reading the call stack Bun formats a method's stack frame without its class ("at step"), so the quit test's caller check never matched, the step was never held, and both cases timed out once CI ran Vitest on Bun (#25840). Only the drain step heals owed queue bookkeeping, so the hold needs no caller check. |
||
|
|
b75213100b |
Add a Chat settings page for structured native chat (#25685)
* Add a Chat settings page for structured native chat * Remove unused chat appearance summary translations * Fix chat settings search entries and preview anchor * Keep shortcut formatting available in settings sidebar tests * Clarify the Chat settings page description * Update Chat settings test and remove unused summary alias |
||
|
|
3ec38b8c6d |
Run Vitest on Bun with Node runtime contracts (#25840)
* Run Vitest on Bun while preserving Node runtime contracts * Preserve runtime timing provenance and keep the Bun pin in config * Scope builtin compatibility mocks to test-only lint exceptions * Give capture retention fixtures distinct filesystem timestamps * Await the copy button success state in the React fixture * Bound Node test worker shutdown and tighten migration fixtures |
||
|
|
5b8a982f8f |
Fix private Linear images in task descriptions (#25849)
* Fix private Linear images in task descriptions * Strengthen Linear image signing validation and typed access |
||
|
|
768c1ea967 |
Add confirmed conversation rewind from structured chat messages (#19338)
* feat(native-chat): confirm destructive conversation rewind from user messages * fix(native-chat): retain rewind uncertainty from host status * fix(native-chat): await the authoritative rewind epoch reset * Pause restored chat outbox during rewind recovery * Verify rewind targets after message acceptance * test(native-chat): verify rewind copy against host reason contract * fix: type rewind hook test props explicitly * Show rewind explanations on keyboard focus * fix(native-chat): read renamed background stop-all capability * fix(native-chat): consolidate epoch test vitest imports * fix(native-chat): restore rewind refusal delivery after merge * fix(native-chat): satisfy main's design-system and cast gates in rewind Match the rewind button to the copy button beside it instead of restyling the shared Button, read refusal reasons without a cast, and merge a duplicate import. * fix(native-chat): let the host govern an in-doubt rewind; offer Edit from here only on turn openers - Pending starts after Confirm and ends with the request; a confirmed rewind waits up to 120 s for its new conversation, then lets go with one toast. - An unknown outcome no longer blocks sending: the host recovers it on the next send. Its latch only disables the action. The message goes back to the composer on success and on an unknown outcome; refusals and timeouts are toasted once instead of holding the error line. - The action is offered only on sent prompts that opened their own turn (no steers, unsent, queued, /compact or goal rows), and not at all where the provider cannot rewind. - Rewind is held behind host-queued cards, gets the command's 195 s remote timeout, focuses the composer after returning the message, drops the message count, and uses plain copy. * fix(native-chat): say why a send waits while a rewind is in flight * fix(native-chat): settle an in-doubt rewind on the next send, even with the agent running - Host: a send (or /clear, /compact) that finds a rewind in doubt with the agent already running now runs the same provider recovery an attach would. Proven: the conversation is replaced and the send proceeds; proven not done: the record is refused and the send proceeds; still unknown: the send is refused as before, but before the ledger records it, so a Retry of the same id is decided afresh instead of replaying the refusal for a day. - Client: a confirmed rewind whose pane was hidden during the wait lets go silently (a hidden pane reads nothing); the timeout and busy copy say what happened and what is needed. * fix(native-chat): keep /clear's settled refusal while a rewind is in doubt; type the live-send test fixture * test(native-chat): give the live-send rewind host the declared agents main now requires * fix(native-chat): offer Edit from here only where the next send settles a rewind, and never on URL images - New runtime capability agent-session.rewind-send-recovery.v1, advertised by this host: the rewind UI hands an in-doubt prompt back for the next send to settle, which hosts with only the rewind RPC refuse while the agent runs. The action is hidden until the host (local or remote) advertises it, and while that is unknown. - A prompt with an image that has no local file is not offered: the image could not go back to the composer after the rewind discards it. - Row and provider eligibility move to native-chat-rewind-eligibility.ts. * fix(native-chat): keep when each kept message was first seen across a rewind A rewind rebuilt the provider's kept items with the rewind's own time, so every kept message showed the rewind time. The merge already carried a held row's scope and producer onto the provider item; it now carries its first-seen time too, for the rewind and its recovery alike. * fix(native-chat): hand rows the rewind through context so the list and pane stay within their line budgets after main's appearance changes * fix(native-chat): fit the rewind recovery capability within protocol-version's line budget; hoist a test context value * fix(native-chat): label the action "Rewind to here" --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
d9cf07d3f4 | Fix pinned workspace reveal expanding other hosts (#25836) | ||
|
|
4de9f9f85a | Move orchestration capabilities out of protocol version registry (#25839) | ||
|
|
ee917205bd |
fix(native-chat): a chat you return to drops background tasks that finished while it was hidden (#24305)
* fix(native-chat): a chat you return to drops background tasks that finished while it was hidden A background command that finished while its chat pane was hidden stayed in the strip as running, its timer counting for hours, and Stop failed with "The background task wasn't stopped." The pane stops listening while hidden, so it missed the "no tasks" update; once the idle sweep stopped the agent, the host reopened the pane's subscription without any roster at all, and the client reads a missing roster as "unchanged". The roster now rides each subscriber's frames the way the slash-command list and the message queue already do. The subscriber registry reads it from the host's child records on a subscriber's first frame and on every snapshot and reset, and re-sends it to each subscriber whose last copy differs when the records change. The channel always answers: a conversation with no running child is "none" whether or not an agent holds it. Ordinary journal batches never read it. This also fixes the reverse case: a snapshot or reset sent to a live pane carried no roster, which cleared a still-running task from the strip until the roster next changed. The client treats a resumed subscription's first batch as stating the roster, so a pane reconnecting to an older host that still omits it drops its stale copy too. Fixes #24227 * refactor(native-chat): require the background-task roster read and drop the unused close observers The host's delivery layer now requires every hook it is built with, so production wiring cannot leave out the roster read (an opening frame without it reads as "no tasks" to current clients). The conversations map's close observers lost their only caller in this PR and are removed. The reducer test comment states the old-host omission rule precisely. * refactor(native-chat): keep structured-agent-session-host.ts within the line limit Main brought the host file to the 300-line lint limit, and this PR's roster wiring adds one line. Name the lease reconciler's type by its factory, as the neighbouring fields do. * test(native-chat): build the roster test's Claude provider handle with claudeProviderHandle Main made the provider handle opaque (#24991); build it with the existing helper, as main's own tests now do. |
||
|
|
670c59d23a |
Translate ACP traffic into shared timeline events (#25090)
* Add standalone ACP protocol client and session runtime
* Protect ACP transport teardown from late stream errors
* Retire incoming ACP request ids before publishing responses
* Narrow ACP configuration requests and transport message types
* Remove redundant ACP request handler return unions
* Keep ACP waits caller-owned and preserve protocol extensions
* Preserve open ACP decisions through prompt completion
* Move the turn message ordinals and the turn-row revision to the neutral timeline folder
Pure moves so a shared timeline assembler can use them: Codex's message
ordinal counter becomes ProviderTurnMessageOrdinals and Claude's turn-row
revision becomes the provider-neutral agent-journal turn-row revision. Only
names and import paths change.
* Admit one provider event's writes as one transition, and let rows be found again after a restart
- A sink transition is admitted whole or not at all; its steps run back to
back at their turn in the journal's write queue, and each resolver reads
the fold with every earlier write landed. A resolver may also say where the
row belongs (turn scope, provider reference), and the writer always hears
how the transition landed. A resolved lifecycle batch chooses its
settlement mutations from the fold at execution.
- New optional row field providerItemRef: the provider's own reference for
the item a row is, written only where the row's identity cannot spell it
(Codex keys messages by their place in the turn and renumbers its item ids
on resume). Set by the creating write, kept by revisions, indexed by the
journal fold, never read by clients. A downgrade test shows an older host
and client render such rows unchanged.
- Provider timeline identity schemes (shared legacy arm, Codex) and the join
index that resolves a provider item to its row from memory or the fold:
ordinals and request incarnations are read back from the rows, so a
restart or an evicted entry finds the original row instead of placing a
new one.
* Recover message ordinals from the journal's highest place, and forget joins read from a replaced epoch
A fresh join index continued a turn's messages at the first free place, so a journal holding only a
later ordinal (an imported or removed earlier row) had its sequence back-filled. The place is now
one past the highest ordinal any row or echoed send holds there, read through a pure scheme reader.
The join caches also drop what they read when the journal's epoch is replaced.
* Spell the subagent thread's message slot without spreading an identity union
* Add a provider timeline grammar and a shared assembler that decides at its turn in the journal
Adapters translate their provider's dialect into a small grammar (turns,
items, streamed text, requests, context facts, session end/reset); one shared
assembler turns it into the journal rows every structured lane writes.
Each event is planned as one sink transition. Which row a write lands on,
whether a replay writes anything, and every change to what the assembler
knows (its ledger) are decided by the transition's resolvers at the event's
turn in the journal's write queue, against the fold as it stands then. A
forecast (the ledger plus admitted events still queued) only answers apply()
at once. So a refused event allocates nothing, a write the journal rejects
leaves no trace in memory, and a restart or evicted cache finds the same rows
again. Text and full snapshots of one provider item share one row and one
lifecycle; reset always flushes text and settles the old session from the
journal; the open-work budget is derived from what is actually open.
Codex migration contracts compare against the existing Codex translator,
including a restart mid-stream and a repeat that outlives the join cache.
* Fix the types and the exhaustive event switch CI reported for the assembler
* Let the journal decide stream lifetimes, request reuse, named sends and background work
A third review found two blockers with the earlier rounds' cause, a remembered interpretation
trusted after the journal moved on:
- A reused request id was judged by its earlier prompt's settled turn before asking which turn the
new one lands in, so a real approval in a later turn was dropped. The target turn now decides:
the old turn again is a replay; a different live turn opens the next prompt beside it.
- A text stream checked its row's turn only on its first write, and a turn's end released streams
by the turn planning expected. Every write now checks the row, a turn's end stops the streams
whose rows are in it, and turn status reads the journal first, so another writer's Stop wins.
Also: a message boundary drawn by an event the journal held as a replay no longer splits an
anonymous message; a send naming a turn not yet open waits for that turn; the budget charges a
stream's thread and turn strings and the turn caches are byte-bounded; the open turn ends when the
journal shows it settled; a turn's opener is read from the journal's row.
Background work is now Orca's existing background-task row instead of a tool call flagged
`outlivesTurn` (a flag remembered only in memory, so a restart failed the task). A turn's end
never settles that row, so it survives restarts; session end leaves one in flight unverifiable.
Three tests that opened a background tool call with `outlivesTurn` now open a background-task row
and keep their original expectations about which turn the row stays in.
* Type the unbound assembler helper's drain as the void it reports
* Bound the rows kept for a continued anonymous message and the stopped streams
A row kept for the anonymous stream that may continue it, and the marker that a stopped stream's
queued writes write nothing, lived in the live-stream map and were never removed when no stream
followed. They now live in their own bounded maps, so the live map holds open streams only.
* Translate ACP session traffic into shared timeline events
* Preserve fixture answer linkage and handle all typed ACP updates
* Keep replay chunks and long-turn tool snapshots bounded
* Forward ACP joins to the journal-derived timeline assembler
* Use exact ACP prompt identity and recover incomplete timeline replay
* Detect short and Unicode home paths in encoded ACP fixtures
* Use generic privacy patterns for ACP fixtures
* Translate Grok background work into durable task rows
* Recover background task origin from its persisted launch tool
* Restate background task fallback from its terminal evidence
* Require session identity on Grok task notifications
* Honor explicit Grok task completion without an exit code
* Match the background reset test to the final timeline grammar
* Give task completion fixtures a specific evidence type
* Settle ACP background tasks from live and replayed evidence
* Honor background task outcomes carried by replayed tool results
* Keep the journal store under its line limit after the main merge
* Drop the provider item reference, join index and identity schemes from the transition PR
Nothing in production reaches the state they defended (an assembler that lost its memory while
its child keeps streaming the same turn), and the stored Codex id was positional. The legacy
identity scheme moves to the assembler PR with its first caller; the Codex scheme and any
persisted reference wait for Codex to move onto the assembler. The Codex ordinal counter goes
back to codex/, since no neutral code imports it.
* Write a resolved settlement in one transaction through enqueueRows
A settlement too large for one row now commits all its rows or none, through the journal's
existing all-or-nothing write, instead of a row-by-row writer. Every row is built before any
commits, so a settlement naming one item twice is refused before anything is written.
* Drop the transition's landing report; keep the turn-row write fire-and-forget
Nothing reads which steps of a transition wrote. A failed step fails the sink, leaving the
steps before it written; the header says so, and tests cover it plus a settlement whose second
row fails inside the transaction. writeAgentJournalTurnRow returns nothing again, as on main.
* Generate open ACP enums and check the generated schema offline
A newer or vendor enum value (tool kind, tool status, option kind, stop
reason) no longer fails the whole message: generated enums accept the known
literals plus any other string, typed so callers can still narrow on the
known ones. The generated header now records the pinned input digests, the
generator digest and a body hash, so `verify:acp-protocol` catches a stale or
hand-edited file without network access; it runs in lint and the PR workflow.
* Land the ACP runtime contract the agent adapters use
- Deliver notifications other than session/update through
onExtensionNotification, in arrival order with session updates.
- Accept _meta on prompt, setMode, setModel, setConfigOption and cancel.
- cancel() always sends session/cancel once the session runs, since the
agent can be in a turn it began itself; only a successful send is shared,
so a failed write is retried.
- Cancel aborts each open agent request's signal and lets its handler send
its own answer; -32800 only when the handler rejects.
- Permission requests validate only the session, tool call id and options;
unreadable fields are dropped with a diagnostic, and any answer Orca
cannot send is `cancelled` instead of a JSON-RPC error. Agent-started
turns may ask; whether to show it is the caller's decision.
- AcpAgentError marks the agent's own errors; AcpInvalidResponseError keeps
the raw answer and validation issues for answers Orca could not read.
- Lines over the size limit are classified by prefix (shared with the Codex
reader): the owed request fails, an oversized agent request is answered
with an error, and an unattributable response closes the connection.
* Run a transition's steps as a prefix; drop the paced flag and resolved options
A failed step no longer lets the steps after it write: each step checks, at its own turn in the
journal's queue, whether the write handed over just ahead of it completed, using the queue's count
of completed write bodies (a promise would report the failure only after the next step ran). The
sink fails only once every step has had its turn. The item step's `paced` size bypass and the
resolver's replacement `options` are removed; nothing planned uses them.
* Rebuild the timeline assembler on one admission-order state
The assembler kept a second copy of its state (a forecast beside a ledger), hydrated
the open turn from the journal, and recognised replays, all to recover from losing its
memory while the provider child kept streaming. That never happens: one assembler lives
exactly as long as one child, and a new child is a new assembler in a new generation
whose events land after the dead-generation sweep.
- One state, changed only when the sink admits an event (minted keys included), so a
refused event takes nothing.
- Every journal-dependent choice is made when the write runs, by keyed reads: a turn row
is written only where none is, a stream checks its row's turn on every write, a running
snapshot never lands in a settled turn or relights a settled tool, a request takes the
first incarnation the journal holds no row for.
- Rows are found by spelling their ids (provider-timeline-rows.ts); no join cache.
- The identity scheme (legacy arm only) lives here with its first caller; requests are
spelled in their acquisition generation, since JSON-RPC ids restart per process.
- Saved history goes in as `input.history` plus ordinary events with the provider's ids,
into an empty journal; `session.reset` and every replay rule are gone.
- One terminal-body function (`terminalAgentJournalBody`) is shared with the
dead-generation settlement.
- The test rig's restart now sweeps and starts a new generation, as production does.
* Answer every agent request after an ACP cancel
A cancel that lands before a permission handler starts now still runs the
permission path, so the agent gets the `cancelled` outcome rather than a
request-cancelled error. A handler that ignores the abort no longer leaves
the agent waiting: once the abort has run through, any request still
unanswered gets request-cancelled. Handlers that answer on abort keep their
own reply.
Also renames a lint-rejected helper parameter, replaces a Reflect.apply in a
test, and stops the permission diagnostic from firing with an empty list.
* Cover new running work in a turn the sweep ended
* Run a transition's steps in one queued write that loops over them
The steps of one event now share one turn in the journal's write queue: a
loop writes each in its own transaction through the row writer's
synchronous writeRows (split out of enqueueRows) and stops at the first
throw. Prefix semantics and "nothing lands between the steps" now hold by
construction, so the completed-write counter on the queue, the step gate
and the allSettled barrier are gone; the queue is back to main's bytes.
* Let each ACP request handler own its answer after a cancel
Removes the next-event-loop-turn fallback that answered request-cancelled
for any handler still silent after a cancel. It raced answers that were
still being saved (an approval mid-journal-write reached the agent as an
error) and made the outcome depend on event-loop timing. The handler that
owns an agent request now always sends its answer, or throws for
request-cancelled; a request it never answers ends when the connection
closes. A permission whose handler had not started still answers
`cancelled`.
* End a turn another writer settled the way the provider's end does
A person's Stop settled the open turn's row without a word to the assembler.
The assembler then forgot the turn: its running tools and pending prompts were
never settled, the provider's own end and withdrawal were dropped, and the
turn's text streams stayed counted against the open budget for the life of the
process. Text the provider kept streaming afterwards could land as a message
outside the stopped turn.
- The open turn the journal shows settled ends first, as one transition, through
the same settlement the provider's turn.end plans; its streams stop and their
keys drop later text until that turn's end or the next turn opens.
- turn.end and request.withdrawn are admitted for a turn or request the journal
holds; their settlement writes nothing for rows already settled.
- The budget's re-check frees streams whose turn settled.
- A settled tool keeps its terminal body against any differing write.
- Session-end settlement of lost background work uses the journal's own
lostLiveWorkJournalBody instead of a copy.
- The rig's window elapses before every read, and restart swaps and disposes
the old assembler.
* Adopt ACP load history only into an empty journal, chosen by the translator's creator
The translator now takes `adopt` from whoever creates it instead of guessing from the
journal (which reads empty before the sink binds). Without adoption, history a provider
replays during session/load is dropped except its context usage. With adoption, history
becomes ordinary events with provider or position ids and the user's saved messages become
`input.history`, so re-running an interrupted adoption lands the same rows.
The translator no longer reads the journal: it lives exactly as long as its provider child.
Its tool snapshots are bounded by the assembler's open budget, and a small record of each
tool's turn keeps a late background-task notice beside the tool that started it.
* Leave a stopped turn's running tools to the agent's own end
When another writer settles the open turn (a person's Stop), the assembler
now only stops that turn's text and cancels its pending prompts. Running tool
calls stay the agent's: a progress update or completion it reports after the
Stop lands as reported, and whatever is still running settles at the agent's
turn end for that turn, the next turn's open, or the session's end.
An agent's end for an earlier turn while a newer one is open no longer clears
the open turn's activity line or ends its anonymous reply. An unnamed end right
after a Stop ends the stopped turn instead of being dropped. The test rig's
restart no longer writes the dead assembler's window text, matching dispose.
* Pin that a stopped turn's running tools hold budget until the agent's end
* Type the stopped turn's tool progress update as a tool body
* List every event the assembler hands to the decision step
The type-aware lint requires an exhaustive switch with no default case.
Also retitle a Stop test to say what it asserts.
* Say why a Grok turn failed, and keep task rows in Grok's own words
A failed Grok turn ended with no reason on screen: the translator dropped
every copy of Grok's message. The failed turn now gets one status row in
Orca's existing "provider did not accept this message" words with Grok's
reason, read from whichever copy arrives first (the given-up retry, the
turn's end, the prompt's completion notice, or the prompt's error answer);
later copies only fill a reason the row still lacks.
A running background command no longer reads "Background task <id>
started": a task's summary is mapped only once it has settled. A monitor
stays a monitor when the agent reads its output: a frame that names no
kind keeps the known one, and a "[monitor" command is a monitor.
A prompt's turn is marked started, so a late frame for an ended prompt
neither reopens it nor becomes the active turn. A tool's turn is held in
one place at a time.
* Read a monitor from Grok's exact output prefix
* Word a failed Grok turn in Grok's own text, not as a refused message
A turn that started and then failed was told "The provider did not accept
this message", Orca's sentence for a message refused before its turn. The
row now reads as a Codex turn-ending error does: an error status row with the
provider's own words. With no words, the dialect names the failure ("Grok
ended this turn with an error." / "Grok usage limit reached."), else the
agent's display name does.
* Register the ACP schema verify step in the PR preflight phase test
* Read ACP permissions, session events and prompt errors through the protocol client's own types
The translator now reads a permission request with the client's lenient reader, a session update
with its session-event reader, and takes only the agent's own error answer as a failed prompt's
reason, so an Orca-side error never reads as the provider's words. Tests cover protocol values
newer than this build.
* refactor(native-chat): drop saved-history adoption from the timeline assembler
The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.
* refactor(acp): drop session/load history adoption from the translator
The common pattern discards the history an agent replays during session/load,
keeping only what it says about the context window. Remove the adoption path
(acp-history-adoption.ts, the adopt option, and the historical background-task
liveness rewrite it fed) so load replay is always dropped except usage.
* refactor(native-chat): a pending input is only Orca's send now
Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.
* test(acp): keep the task-result status table on live frames
Review follow-up to the adoption removal: the result-status mapping was only
tested through adopted history, so run the same table on live frames, and
cover an unmarked task notice during a load being dropped.
* Use current provider handles in transition tests
* Use current provider handles in timeline fixtures
* Require the ACP directory in the runtime import check
|
||
|
|
13ea35973c |
Stop expensive checks when an unmerged PR closes (#25829)
* Cancel active checks when an unmerged PR closes * Register owned-branch cancellation qualification * Keep temporary cancellation qualification outside the review diff |
||
|
|
f6f96db6be | Build SSH hostile-host Linux slots independently (#25821) | ||
|
|
634c43787c |
Fix premature Claude automation completion and inherited CI failures (#24878)
* fix(agent-status): keep a Claude pane working until owed task wake-ups arrive Claude wakes the main agent for each background task that ends, after the task stops running. "Nothing running" was read as done, so automations closed the terminal before the final turn. Fixes #23942 * fix(agent-status): keep owed task wake-ups through failed turns, the cap and nested exits A failed turn's task list now marks a vanished shell owed like a normal turn end does. Past the cap only an already-announced sub-agent is forgotten. A process exit clears what is owed only where the server admits it from the pane's owner. * refactor(agent-status): move pane-scoped cache entry helpers out of listener-state listener-state.ts passed the 300-line limit once main's and this branch's additions met. * fix(agent-status): verify Claude task wake-up completion on its execution host * fix(agent-status): canonicalize Claude background task identifiers * Preserve out-of-order Claude task notification evidence * Verify retained released parser in cross-version checkout test * Reuse watcher directory-cache enumeration without losing fresh keys --------- Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> |
||
|
|
0bcd49c04c |
Let ACP connections own their agent processes (#25810)
* Let ACP connections own their supervised agent process * Preserve ACP cleanup evidence and isolate exit observers * Expose ACP cleanup observations and type the permission fixture |
||
|
|
a2f197fde7 |
Refresh visible reviews automatically and stop settled merged polling (#25788)
Refresh only actual on-screen review cards and the selected visible review panel, with a single metadata owner per execution host. Open reviews refresh every 60 seconds when selected and 120 seconds otherwise. Settled merged reviews stop automatic refresh; pending merged checks continue, hidden rows stop, and work changes or stale re-exposure discover fresh state. Remove the polling setting. Preserve provider/SSH/runtime ownership, mixed-version cache behavior, failure backoff, bounded foreground admission, request coalescing, and mutation invalidation. Cover older persisted caches and unknown HEADs. Validated with 1,628 focused tests, typechecking, lint/code-quality gates, hidden Electron viewport checks, and parallel Codex/Claude Opus 5.5 adversarial reviews. Measurements and policy details are in PR #25788. Fixes #25746 Co-authored-by: ggbdpq <ggbdpq@gmail.com> |
||
|
|
9905765e3b |
feat(native-chat): open structured chat's wire and stored records to registered agents, behind a negotiated capability (#25159)
* refactor(native-chat): keep the provider resume handle opaque to shared code
Shared structured-chat code parsed each provider's resume handle: Claude's
session id and branch leaf, Codex's thread id, through a 'claude' | 'codex'
union every new agent had to widen. The in-memory handle is now
{ transport, agent, nativeId, providerData? }: shared readers use nativeId,
lease and handle-chain checks compare transport and agent, and only the
Claude adapter reads its leaf (providerData).
Stored and wire forms are unchanged for Claude and Codex. One encoding
module writes their typed shapes and decodes both those and the neutral
shape a new transport uses, which an older build refuses as unreadable
rather than reading as Codex. Key and root strings, which fork seeds,
superseded creations and resume offers persist, stay byte-identical.
The journal's own handle type becomes the journal-row and attach-wire
encoding of the same handle, and the journal identity carries the
neutral handle (null before the provider proves one).
No user-visible change.
* fix(native-chat): derive journal-row provider handles from the journal identity
The journal row converter now takes the identity every caller already holds,
so a row's handle has one obvious constructor. Tests that wrote the in-memory
handle straight into journal rows now build it through that converter, and the
processless Claude fixture names a not-yet-proved handle as null.
* fix(native-chat): refuse a stored provider handle written in both forms
A typed Claude or Codex handle that also carries the neutral form's
transport, agent, native id or provider data named two identities; it
was read as Claude or Codex and the next write dropped the other one.
Such a row now stays unreadable and is set aside untouched.
* refactor(native-chat): route structured agents through registered definitions
The structured-chat shared layer named Claude and Codex in a closed router
type, in option reads, in frame classification, and in renderer label and
validation branches, so adding an agent meant editing shared types. Each
agent now has a definition (capabilities, resting option rules) owned by its
own module; the router is a registry of definition + adapter pairs built at
runtime setup, and shared code reads the definition instead of the name.
No behavior change for Claude or Codex; no wire or stored shape change.
* refactor(native-chat): name the structured agent list once in host types
The host types, conversation commands, provider-session ownership and the
runtime's create path spelled the structured agent list inline; they now use
the one shared type the record and wire already validate against.
* fix(native-chat): narrow the record before reading its agent's option rules
* refactor(native-chat): make the router's registrations the only agent definition lookup
The at-rest option reads looked definitions up in a second, separately populated map, so a
registered definition could route and declare capabilities while the static map decided which
options a chat at rest accepted. The router now answers `definition(agent)` from its
registrations, the at-rest readers take it through the host's adapter, and the built-in
definitions are listed only where the runtime composes the router.
* feat(native-chat): open the structured-chat wire and stored records to registered agents
A host's structured agents are the ones its runtime registered. Records, RPC
params, persisted tabs and the model catalog accept any registered agent instead
of naming Claude and Codex; each agent's definition declares the transport its
handles live in and the variable its account home pins. A new runtime
capability, agent-session.structured.registered-agents.v1, advertises that a
host accepts and lists its agents (agentSession.agents, with each agent's
capability record), and the host withholds any other agent's tabs and restart
offers from clients that do not advertise it.
* test(native-chat): cover registered agents on the wire, in storage and across versions
* refactor(native-chat): let the record store decide which agents' tabs exist
* test(native-chat): declare the pilot test agent's storage
* test(native-chat): read the old build's saved tabs through a parsed shape
* fix(native-chat): act on restart offers only for agents the calling client can show
A paired client too old to show an agent's chat was listed only the offers it could show, but
dismissing or continuing all reached every offer on the host, and a named continuation answered
with the host's whole remaining inventory. The client's audience now goes to the host with every
restart operation: only offers it sees are reserved, dismissed or returned. Without an audience
(this host's own process, or a client that shows every agent) nothing changes.
* test(native-chat): read an agent-registering baseline's storage on its own terms
The registered-agents downgrade test assumed its baseline release predates registered agents: it
expected the saved-tab parser to erase an unknown agent and called the record reader without the
agents list. Once a release with this change becomes the baseline, both break. The expectations now
follow what the baseline host advertises, and an agent-registering baseline is handed its own
Claude and Codex storage.
* refactor(native-chat): derive record-store admission from the runtime's agent registrations
Which agents a stored record may name and which agents the router drives came from two lists in
the runtime, so a newly registered agent could be routed while its records were set aside. One
list of registrations now holds each agent's definition and the factory for its adapter: the
store's admitted agents are derived from it before the store opens, and the adapters are built
from it once it has.
* fix(native-chat): hand the exit drain a promise for every registered agent
* fix(native-chat): let the adoption conflict check read any agent's ownership
Ownership rows name any registered agent since the stored records opened to them; the adoption
check compares by agent, so it takes the same open id. Only Claude and Codex still adopt.
* test(native-chat): use opaque handle in queued rejection fixture
* test(native-chat): share one Codex journal identity in the integration suite
Main grew the suite to the 800-line limit; the opaque-handle import pushed it
over. The two tests built the same identity inline.
* refactor(agent-session): name the handle's adapter state resumeCursor
Rename the neutral provider handle's providerData to resumeCursor before any
row persists the neutral form: it is an adapter-owned resume position (Claude's
transcript leaf), never identity. Claude/Codex stored and wire bytes are
unchanged; their typed shapes never carried the field.
State the stored-form contract (a handle's field set is closed; later per-link
data goes on the chain link, which every build preserves) and pin it with a
record round-trip test. Document that transport records the id space the
native id was minted in, which can differ from the agent's current transport.
* refactor(agent-session): one required agent registry; declarations admit what they claim
A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.
/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.
Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).
* refactor(agent-session): the router applies the declared rewind itself
The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.
* test(agent-session): register the agents the merged-in tests now need
The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.
* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record
* fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop
The desktop renderer calls the restart methods as a paired runtime client
without the registered-agents capability, so it was handed a Claude/Codex
audience and its "dismiss all" took the scoped path: no dismissal fence and
no clearing of unwritten teardown witnesses, so a late teardown write could
bring a dismissed offer back.
The audience is now derived against this host's registry: a client that can
show every agent the host registered gets none, and dismisses exactly as
before (one fence for every offer). A client that truly cannot show some
agent still gets a final dismissal for what it sees: each cleared session is
fenced on its own (bounded, superseded by a dismiss-all fence) and only the
witnesses of agents it sees are dropped.
One predicate now answers which agents a client renders, for both tab
projection and restart offers; the Claude capability rule is folded in.
* fix(agent-session): a changed agent definition never hides that agent's chats
A record was readable only if every handle used the transport its agent's
definition declares today, and its account variable was declared by some
registered agent. A later build that drives the same agent over another
protocol, or under a renamed variable, would have set every existing chat of
that agent aside: no tab, no history, no error.
Readability now asks only that the record agree with itself: its agent is
registered, every handle shares one namespace owned by that agent, and the
account variable is a well-formed name. Whether this build can drive it (the
chain's transport is the one the agent speaks, and the account variable is
that agent's own, since it becomes the child's environment) is decided in
performAttach, the one admission every start of every agent passes through,
and refused there as hostUnsupported. A record pinning another agent's
variable is refused the same way rather than passing because some agent
declares it.
* refactor(agent-session): each agent's registration says where it runs and which account it pins
createSupport, create and the model catalog's account read still chose a
location rule and an account resolver by name (Claude, Codex, else none),
so a new agent would have needed a third branch beside the registration
list that already decides storage, routing and publication.
Each runtime registration now carries supportsLocation(location) and
resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }),
and Claude's and Codex's rules move into their entries unchanged: a read
still syncs no home and starts no bridge, a Codex launch still trusts the
workspace first, and WSL locations resolve as before. The list exists
before the host is built, so none of these reads installs the host, and an
agent the list does not hold is answered no without opening the journal.
* test(agent-session): drop a duplicate registry key; keep fences out of the capsule state
A spread input already carries the registry, so the explicit key was
overwritten (TS2783). The session-fence helper now returns only the fence,
not the capsule state it was given.
* fix(agent-session): a scoped dismiss-all persists no per-session fence
The per-session dismissal fence added a forever-persisted capsule field
with an arbitrary cap that served no caller: only the local desktop calls
the restart methods, and on a host whose agents it all shows it gets no
audience and takes the unscoped, fenced dismissal. A scoped dismiss-all
(a caller that cannot show some registered agent) keeps that agent's
offers, clears only its audience's unwritten witnesses, and writes no
fence; a late write from this process is serialized behind it. The field
never left this unreleased branch.
* fix(agent-session): refuse an attach whose agent is not the session's own
The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal.
* fix(agent-session): offer to start a chat only when the start would accept it
The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired.
* fix(agent-session): the model-catalog probe uses a record's account only if this build drives it
A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record.
* refactor(agent-session): the record store admits agent ids; comments say where transport is checked
The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept.
* docs(agent-session): the record store admits the registered agents' ids
* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer
Uses an audience production sends (one that cannot show every agent), per review.
* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge
* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way
The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.
* test(ratchet): require src/main/provider-process now that it has landed
* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test
Main's #25706 and this branch both added the import at different lines; the merge kept both.
* Keep saved chats readable without provider registration
* Keep stored-record compatibility checks independent of registration
* Keep saved providers in restart client audiences
* test(native-chat): type reveal fixtures without assertions
* test(wire): expose known agents in structured host fixture
* Supply startability dependency in the new Codex catalog fixture
|
||
|
|
c4ea14cd9f |
Show a tool call you stopped as interrupted, not failed (#25181)
* Move the turn message ordinals and the turn-row revision to the neutral timeline folder Pure moves so a shared timeline assembler can use them: Codex's message ordinal counter becomes ProviderTurnMessageOrdinals and Claude's turn-row revision becomes the provider-neutral agent-journal turn-row revision. Only names and import paths change. * Admit one provider event's writes as one transition, and let rows be found again after a restart - A sink transition is admitted whole or not at all; its steps run back to back at their turn in the journal's write queue, and each resolver reads the fold with every earlier write landed. A resolver may also say where the row belongs (turn scope, provider reference), and the writer always hears how the transition landed. A resolved lifecycle batch chooses its settlement mutations from the fold at execution. - New optional row field providerItemRef: the provider's own reference for the item a row is, written only where the row's identity cannot spell it (Codex keys messages by their place in the turn and renumbers its item ids on resume). Set by the creating write, kept by revisions, indexed by the journal fold, never read by clients. A downgrade test shows an older host and client render such rows unchanged. - Provider timeline identity schemes (shared legacy arm, Codex) and the join index that resolves a provider item to its row from memory or the fold: ordinals and request incarnations are read back from the rows, so a restart or an evicted entry finds the original row instead of placing a new one. * Recover message ordinals from the journal's highest place, and forget joins read from a replaced epoch A fresh join index continued a turn's messages at the first free place, so a journal holding only a later ordinal (an imported or removed earlier row) had its sequence back-filled. The place is now one past the highest ordinal any row or echoed send holds there, read through a pure scheme reader. The join caches also drop what they read when the journal's epoch is replaced. * Spell the subagent thread's message slot without spreading an identity union * Add a provider timeline grammar and a shared assembler that decides at its turn in the journal Adapters translate their provider's dialect into a small grammar (turns, items, streamed text, requests, context facts, session end/reset); one shared assembler turns it into the journal rows every structured lane writes. Each event is planned as one sink transition. Which row a write lands on, whether a replay writes anything, and every change to what the assembler knows (its ledger) are decided by the transition's resolvers at the event's turn in the journal's write queue, against the fold as it stands then. A forecast (the ledger plus admitted events still queued) only answers apply() at once. So a refused event allocates nothing, a write the journal rejects leaves no trace in memory, and a restart or evicted cache finds the same rows again. Text and full snapshots of one provider item share one row and one lifecycle; reset always flushes text and settles the old session from the journal; the open-work budget is derived from what is actually open. Codex migration contracts compare against the existing Codex translator, including a restart mid-stream and a repeat that outlives the join cache. * Fix the types and the exhaustive event switch CI reported for the assembler * Let the journal decide stream lifetimes, request reuse, named sends and background work A third review found two blockers with the earlier rounds' cause, a remembered interpretation trusted after the journal moved on: - A reused request id was judged by its earlier prompt's settled turn before asking which turn the new one lands in, so a real approval in a later turn was dropped. The target turn now decides: the old turn again is a replay; a different live turn opens the next prompt beside it. - A text stream checked its row's turn only on its first write, and a turn's end released streams by the turn planning expected. Every write now checks the row, a turn's end stops the streams whose rows are in it, and turn status reads the journal first, so another writer's Stop wins. Also: a message boundary drawn by an event the journal held as a replay no longer splits an anonymous message; a send naming a turn not yet open waits for that turn; the budget charges a stream's thread and turn strings and the turn caches are byte-bounded; the open turn ends when the journal shows it settled; a turn's opener is read from the journal's row. Background work is now Orca's existing background-task row instead of a tool call flagged `outlivesTurn` (a flag remembered only in memory, so a restart failed the task). A turn's end never settles that row, so it survives restarts; session end leaves one in flight unverifiable. Three tests that opened a background tool call with `outlivesTurn` now open a background-task row and keep their original expectations about which turn the row stays in. * Type the unbound assembler helper's drain as the void it reports * Bound the rows kept for a continued anonymous message and the stopped streams A row kept for the anonymous stream that may continue it, and the marker that a stopped stream's queued writes write nothing, lived in the live-stream map and were never removed when no stream followed. They now live in their own bounded maps, so the live map holds open streams only. * Read a tool call a stop or the session's end cut short as interrupted, not failed A call still running when its turn was interrupted, when the provider child was seen to exit, or that the provider cancelled now carries `endedAs: 'interrupted'` beside `state: 'failed'`. Every build that predates the field keeps reading the call as the failure it always showed; this build reads both through one shared reader and counts the call apart from failures in the run header, without the error tint on its partial output. The shared assembler, the restart settlement and the Codex lane settle a running call through one rule: only a proven interruption cuts it short, so a turn the provider completed around an unclosed call, or work the host lost track of, still reads failed. * Assert a completed Codex turn leaves its unclosed call plainly failed * Assert an unverifiable restart settlement leaves its running call plainly failed * Let a later death proof correct the calls an unverifiable settle closed A settle with no proof of the child's death closed its running calls as plain failed, so a proof written afterwards corrected the turn to interrupted but could not find the calls. The call now keeps endedAs: 'unverifiable' beside its failed state, and the stale settle revises exactly those calls whose owner the proof names. Older builds still read them as failed. * Render a cut-short call's output on mobile, with and without proof * Find the mobile result box by walking to its View, and type the unverified-ending block * Walk to the mobile result box without a type assertion * Keep the journal store under its line limit after the main merge * Drop the provider item reference, join index and identity schemes from the transition PR Nothing in production reaches the state they defended (an assembler that lost its memory while its child keeps streaming the same turn), and the stored Codex id was positional. The legacy identity scheme moves to the assembler PR with its first caller; the Codex scheme and any persisted reference wait for Codex to move onto the assembler. The Codex ordinal counter goes back to codex/, since no neutral code imports it. * Write a resolved settlement in one transaction through enqueueRows A settlement too large for one row now commits all its rows or none, through the journal's existing all-or-nothing write, instead of a row-by-row writer. Every row is built before any commits, so a settlement naming one item twice is refused before anything is written. * Drop the transition's landing report; keep the turn-row write fire-and-forget Nothing reads which steps of a transition wrote. A failed step fails the sink, leaving the steps before it written; the header says so, and tests cover it plus a settlement whose second row fails inside the transaction. writeAgentJournalTurnRow returns nothing again, as on main. * Run a transition's steps as a prefix; drop the paced flag and resolved options A failed step no longer lets the steps after it write: each step checks, at its own turn in the journal's queue, whether the write handed over just ahead of it completed, using the queue's count of completed write bodies (a promise would report the failure only after the next step ran). The sink fails only once every step has had its turn. The item step's `paced` size bypass and the resolver's replacement `options` are removed; nothing planned uses them. * Rebuild the timeline assembler on one admission-order state The assembler kept a second copy of its state (a forecast beside a ledger), hydrated the open turn from the journal, and recognised replays, all to recover from losing its memory while the provider child kept streaming. That never happens: one assembler lives exactly as long as one child, and a new child is a new assembler in a new generation whose events land after the dead-generation sweep. - One state, changed only when the sink admits an event (minted keys included), so a refused event takes nothing. - Every journal-dependent choice is made when the write runs, by keyed reads: a turn row is written only where none is, a stream checks its row's turn on every write, a running snapshot never lands in a settled turn or relights a settled tool, a request takes the first incarnation the journal holds no row for. - Rows are found by spelling their ids (provider-timeline-rows.ts); no join cache. - The identity scheme (legacy arm only) lives here with its first caller; requests are spelled in their acquisition generation, since JSON-RPC ids restart per process. - Saved history goes in as `input.history` plus ordinary events with the provider's ids, into an empty journal; `session.reset` and every replay rule are gone. - One terminal-body function (`terminalAgentJournalBody`) is shared with the dead-generation settlement. - The test rig's restart now sweeps and starts a new generation, as production does. * Cover new running work in a turn the sweep ended * Run a transition's steps in one queued write that loops over them The steps of one event now share one turn in the journal's write queue: a loop writes each in its own transaction through the row writer's synchronous writeRows (split out of enqueueRows) and stops at the first throw. Prefix semantics and "nothing lands between the steps" now hold by construction, so the completed-write counter on the queue, the step gate and the allSettled barrier are gone; the queue is back to main's bytes. * End a turn another writer settled the way the provider's end does A person's Stop settled the open turn's row without a word to the assembler. The assembler then forgot the turn: its running tools and pending prompts were never settled, the provider's own end and withdrawal were dropped, and the turn's text streams stayed counted against the open budget for the life of the process. Text the provider kept streaming afterwards could land as a message outside the stopped turn. - The open turn the journal shows settled ends first, as one transition, through the same settlement the provider's turn.end plans; its streams stop and their keys drop later text until that turn's end or the next turn opens. - turn.end and request.withdrawn are admitted for a turn or request the journal holds; their settlement writes nothing for rows already settled. - The budget's re-check frees streams whose turn settled. - A settled tool keeps its terminal body against any differing write. - Session-end settlement of lost background work uses the journal's own lostLiveWorkJournalBody instead of a copy. - The rig's window elapses before every read, and restart swaps and disposes the old assembler. * Leave a stopped turn's running tools to the agent's own end When another writer settles the open turn (a person's Stop), the assembler now only stops that turn's text and cancels its pending prompts. Running tool calls stay the agent's: a progress update or completion it reports after the Stop lands as reported, and whatever is still running settles at the agent's turn end for that turn, the next turn's open, or the session's end. An agent's end for an earlier turn while a newer one is open no longer clears the open turn's activity line or ends its anonymous reply. An unnamed end right after a Stop ends the stopped turn instead of being dropped. The test rig's restart no longer writes the dead assembler's window text, matching dispose. * Pin that a stopped turn's running tools hold budget until the agent's end * Type the stopped turn's tool progress update as a tool body * List every event the assembler hands to the decision step The type-aware lint requires an exhaustive switch with no default case. Also retitle a Stop test to say what it asserts. * End a running call as its turn's journal row ends A call still running when its turn ends takes the state of that turn's row: a row another writer settled first (a person's Stop) stands, so its calls read interrupted whatever the provider's later end reports. The no-ending path that settled calls from the Stop row is gone, since a Stop now leaves running calls to the provider. Adds the two Spanish strings. * Settle a stopped turn's running call as its turn row ended after a restart too The restart sweep ended every running call by the death evidence alone, so after a person's Stop with no proof the child died the call read failed under a turn that read interrupted. The sweep and the live dead-generation settlement now ask the same rule the assembler does: a call in a turn already settled ends as that row ended; only a turn still running leaves its calls to the evidence. * Keep the dead-generation settlement under the line cap * refactor(native-chat): drop saved-history adoption from the timeline assembler The common pattern discards the history a provider replays while loading a session, so the assembler has no use for an input.history event. * refactor(native-chat): a pending input is only Orca's send now Review follow-up to the adoption removal: drop the comment naming the provider's saved message, and make requestedAt required since every pending input comes from input.accepted. * test(ratchet): require src/main/provider-process now that it has landed * test(native-chat): build this stack's journal identities with main's opaque provider handle Main's #24991 replaced the {kind, ...} handle with {transport, agent, nativeId}; three test files from this stack still wrote the old shape. Same lines the downstream ACP branch uses. |
||
|
|
41f293c264 |
fix(notifications): replay completion sounds reliably (#15940)
* fix(notifications): replay completion sounds reliably The in-flight `isNotificationSoundPlaying` gate dropped every notification sound that arrived while the previous one was still ringing, so bursts played once. Drop the gate and restart the cached Audio instead, and reuse an entry a concurrent call already cached for the same path so parallel first playbacks share one Audio rather than revoking each other's blob URL. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016S8SZpkXGcnU5UWwA8UiF2 * test(notifications): verify real sound replay and remove unused playback listeners --------- Co-authored-by: Laku <laku@LakudeMacBook-Pro.local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
6a9ba9d733 |
Keep remote worktree creation from navigating paired clients (#25793)
* fix(workspaces): keep remote creation from navigating paired viewers * fix(workspaces): revoke pending navigation after connection replacement * fix(workspaces): provision when the host renderer is unavailable * test: align navigation and cache-scan expectations with current behavior * fix(workspaces): provision broadcast creates without a host renderer |
||
|
|
8507111d87 |
Document Apple Git ref recovery retirement condition
Name the affected Apple Git-154 (2.39.5) build and document when corruption recovery can be retired while retaining current-Git disambiguation. Comments only; lint, formatting, and review pass. |
||
|
|
29e669680d |
fix(codex): never write a Codex config.toml that Codex can't load, and approve both symlink spellings (#25741)
* fix(codex): never write a hook approval Codex cannot load, and approve both symlink spellings - Refuse any hooks.state write into a Codex config.toml (upsert, move, remove, mirrored enabled state, SSH installer) that would turn a loadable file into one Codex cannot load; write nothing and surface the reason. - Read approvals written as dotted keys or inline tables. - Approve and remove Orca's ~/.codex hook under both the spelled and the resolved key when ~/.codex or HOME is a symlink; move user approvals under both keys. - Stale runtime trust cleanup no longer keeps an unexpected key whose conflicting duplicate tables read as no hash. * refactor(codex): pass every hooks.json spelling as one sourcePaths list * fix(codex): sweep retired-hook approvals under every key spelling; read literal-string trusted_hash |
||
|
|
1de3aa405f |
Fix truncated Unicode branch names in base ref search
Preserve complete branch selectors when Git truncates or disambiguates short names, and keep slash-named local branches separate from remote-tracking refs through worktree creation and reuse. Keep native/SSH namespace boundaries and mixed-version client capability gates intact. Accept Git-valid dotted components and cover collisions with configured and orphan tracking refs. Adapted from @nishino-tsukasa's #19540. Independently reviewed in seven adversarial rounds; 235 focused tests and 77 affected-Apple-Git tests pass, along with typecheck, quality checks, and the desktop build. Correct the inherited file-explorer cache-scan expectation while retaining its exact constant-work bound. Correct the inherited release-parser compatibility assertion to match the pinned test dependency; all 121 cross-version tests pass locally. Fixes #19515 Co-authored-by: kino <nishinotsukasavirgo@gmail.com> |
||
|
|
3ec030e80a | Update README downloads badge | ||
|
|
ac3a25ca37 |
Wait for Codex's live composer before pasting linked issue drafts (#25779)
* Wait for Codex's live composer before pasting launch drafts * Keep timed-out Codex draft delivery consumed across remounts * Document the checked terminal fixtures used by the remount test * Distinguish Codex's reserved footer row from early multiline input * Complete transcript baselines and suppress hidden terminal strings * Align watcher scan budget with linked-directory reconciliation * Drop unused snapshots after main removed the readiness census |
||
|
|
eaaae0196f |
feat(native-chat): record fresh sessions after failed restoration (#25747)
* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore A chat whose saved conversation the agent cannot reopen can now continue in a fresh one: the handle chain records the new conversation as a creation that replaces the lost one (which, why, and when), keeping every earlier link. Rows keep a shape older builds read: the stored chain starts at the latest replacement and carries the earlier links inside it. * refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row Older builds only read Claude and Codex records, so the nested stored form protected rows no replacement can reach while adding a cap mismatch after a downgrade. Store the chain as held, refuse a replacement in a Claude or Codex chain until one has a stored shape older builds read, and refuse a supersession key on a replacement that names no creation in the chain. * test(native-chat): prove replacement rows survive downgrade and re-upgrade |
||
|
|
d7a95782d1 |
Add a shared timeline assembler for structured agent chats (not wired yet) (#25064)
* Move the turn message ordinals and the turn-row revision to the neutral timeline folder Pure moves so a shared timeline assembler can use them: Codex's message ordinal counter becomes ProviderTurnMessageOrdinals and Claude's turn-row revision becomes the provider-neutral agent-journal turn-row revision. Only names and import paths change. * Admit one provider event's writes as one transition, and let rows be found again after a restart - A sink transition is admitted whole or not at all; its steps run back to back at their turn in the journal's write queue, and each resolver reads the fold with every earlier write landed. A resolver may also say where the row belongs (turn scope, provider reference), and the writer always hears how the transition landed. A resolved lifecycle batch chooses its settlement mutations from the fold at execution. - New optional row field providerItemRef: the provider's own reference for the item a row is, written only where the row's identity cannot spell it (Codex keys messages by their place in the turn and renumbers its item ids on resume). Set by the creating write, kept by revisions, indexed by the journal fold, never read by clients. A downgrade test shows an older host and client render such rows unchanged. - Provider timeline identity schemes (shared legacy arm, Codex) and the join index that resolves a provider item to its row from memory or the fold: ordinals and request incarnations are read back from the rows, so a restart or an evicted entry finds the original row instead of placing a new one. * Recover message ordinals from the journal's highest place, and forget joins read from a replaced epoch A fresh join index continued a turn's messages at the first free place, so a journal holding only a later ordinal (an imported or removed earlier row) had its sequence back-filled. The place is now one past the highest ordinal any row or echoed send holds there, read through a pure scheme reader. The join caches also drop what they read when the journal's epoch is replaced. * Spell the subagent thread's message slot without spreading an identity union * Add a provider timeline grammar and a shared assembler that decides at its turn in the journal Adapters translate their provider's dialect into a small grammar (turns, items, streamed text, requests, context facts, session end/reset); one shared assembler turns it into the journal rows every structured lane writes. Each event is planned as one sink transition. Which row a write lands on, whether a replay writes anything, and every change to what the assembler knows (its ledger) are decided by the transition's resolvers at the event's turn in the journal's write queue, against the fold as it stands then. A forecast (the ledger plus admitted events still queued) only answers apply() at once. So a refused event allocates nothing, a write the journal rejects leaves no trace in memory, and a restart or evicted cache finds the same rows again. Text and full snapshots of one provider item share one row and one lifecycle; reset always flushes text and settles the old session from the journal; the open-work budget is derived from what is actually open. Codex migration contracts compare against the existing Codex translator, including a restart mid-stream and a repeat that outlives the join cache. * Fix the types and the exhaustive event switch CI reported for the assembler * Let the journal decide stream lifetimes, request reuse, named sends and background work A third review found two blockers with the earlier rounds' cause, a remembered interpretation trusted after the journal moved on: - A reused request id was judged by its earlier prompt's settled turn before asking which turn the new one lands in, so a real approval in a later turn was dropped. The target turn now decides: the old turn again is a replay; a different live turn opens the next prompt beside it. - A text stream checked its row's turn only on its first write, and a turn's end released streams by the turn planning expected. Every write now checks the row, a turn's end stops the streams whose rows are in it, and turn status reads the journal first, so another writer's Stop wins. Also: a message boundary drawn by an event the journal held as a replay no longer splits an anonymous message; a send naming a turn not yet open waits for that turn; the budget charges a stream's thread and turn strings and the turn caches are byte-bounded; the open turn ends when the journal shows it settled; a turn's opener is read from the journal's row. Background work is now Orca's existing background-task row instead of a tool call flagged `outlivesTurn` (a flag remembered only in memory, so a restart failed the task). A turn's end never settles that row, so it survives restarts; session end leaves one in flight unverifiable. Three tests that opened a background tool call with `outlivesTurn` now open a background-task row and keep their original expectations about which turn the row stays in. * Type the unbound assembler helper's drain as the void it reports * Bound the rows kept for a continued anonymous message and the stopped streams A row kept for the anonymous stream that may continue it, and the marker that a stopped stream's queued writes write nothing, lived in the live-stream map and were never removed when no stream followed. They now live in their own bounded maps, so the live map holds open streams only. * Keep the journal store under its line limit after the main merge * Drop the provider item reference, join index and identity schemes from the transition PR Nothing in production reaches the state they defended (an assembler that lost its memory while its child keeps streaming the same turn), and the stored Codex id was positional. The legacy identity scheme moves to the assembler PR with its first caller; the Codex scheme and any persisted reference wait for Codex to move onto the assembler. The Codex ordinal counter goes back to codex/, since no neutral code imports it. * Write a resolved settlement in one transaction through enqueueRows A settlement too large for one row now commits all its rows or none, through the journal's existing all-or-nothing write, instead of a row-by-row writer. Every row is built before any commits, so a settlement naming one item twice is refused before anything is written. * Drop the transition's landing report; keep the turn-row write fire-and-forget Nothing reads which steps of a transition wrote. A failed step fails the sink, leaving the steps before it written; the header says so, and tests cover it plus a settlement whose second row fails inside the transaction. writeAgentJournalTurnRow returns nothing again, as on main. * Run a transition's steps as a prefix; drop the paced flag and resolved options A failed step no longer lets the steps after it write: each step checks, at its own turn in the journal's queue, whether the write handed over just ahead of it completed, using the queue's count of completed write bodies (a promise would report the failure only after the next step ran). The sink fails only once every step has had its turn. The item step's `paced` size bypass and the resolver's replacement `options` are removed; nothing planned uses them. * Rebuild the timeline assembler on one admission-order state The assembler kept a second copy of its state (a forecast beside a ledger), hydrated the open turn from the journal, and recognised replays, all to recover from losing its memory while the provider child kept streaming. That never happens: one assembler lives exactly as long as one child, and a new child is a new assembler in a new generation whose events land after the dead-generation sweep. - One state, changed only when the sink admits an event (minted keys included), so a refused event takes nothing. - Every journal-dependent choice is made when the write runs, by keyed reads: a turn row is written only where none is, a stream checks its row's turn on every write, a running snapshot never lands in a settled turn or relights a settled tool, a request takes the first incarnation the journal holds no row for. - Rows are found by spelling their ids (provider-timeline-rows.ts); no join cache. - The identity scheme (legacy arm only) lives here with its first caller; requests are spelled in their acquisition generation, since JSON-RPC ids restart per process. - Saved history goes in as `input.history` plus ordinary events with the provider's ids, into an empty journal; `session.reset` and every replay rule are gone. - One terminal-body function (`terminalAgentJournalBody`) is shared with the dead-generation settlement. - The test rig's restart now sweeps and starts a new generation, as production does. * Cover new running work in a turn the sweep ended * Run a transition's steps in one queued write that loops over them The steps of one event now share one turn in the journal's write queue: a loop writes each in its own transaction through the row writer's synchronous writeRows (split out of enqueueRows) and stops at the first throw. Prefix semantics and "nothing lands between the steps" now hold by construction, so the completed-write counter on the queue, the step gate and the allSettled barrier are gone; the queue is back to main's bytes. * End a turn another writer settled the way the provider's end does A person's Stop settled the open turn's row without a word to the assembler. The assembler then forgot the turn: its running tools and pending prompts were never settled, the provider's own end and withdrawal were dropped, and the turn's text streams stayed counted against the open budget for the life of the process. Text the provider kept streaming afterwards could land as a message outside the stopped turn. - The open turn the journal shows settled ends first, as one transition, through the same settlement the provider's turn.end plans; its streams stop and their keys drop later text until that turn's end or the next turn opens. - turn.end and request.withdrawn are admitted for a turn or request the journal holds; their settlement writes nothing for rows already settled. - The budget's re-check frees streams whose turn settled. - A settled tool keeps its terminal body against any differing write. - Session-end settlement of lost background work uses the journal's own lostLiveWorkJournalBody instead of a copy. - The rig's window elapses before every read, and restart swaps and disposes the old assembler. * Leave a stopped turn's running tools to the agent's own end When another writer settles the open turn (a person's Stop), the assembler now only stops that turn's text and cancels its pending prompts. Running tool calls stay the agent's: a progress update or completion it reports after the Stop lands as reported, and whatever is still running settles at the agent's turn end for that turn, the next turn's open, or the session's end. An agent's end for an earlier turn while a newer one is open no longer clears the open turn's activity line or ends its anonymous reply. An unnamed end right after a Stop ends the stopped turn instead of being dropped. The test rig's restart no longer writes the dead assembler's window text, matching dispose. * Pin that a stopped turn's running tools hold budget until the agent's end * Type the stopped turn's tool progress update as a tool body * List every event the assembler hands to the decision step The type-aware lint requires an exhaustive switch with no default case. Also retitle a Stop test to say what it asserts. * refactor(native-chat): drop saved-history adoption from the timeline assembler The common pattern discards the history a provider replays while loading a session, so the assembler has no use for an input.history event. * refactor(native-chat): a pending input is only Orca's send now Review follow-up to the adoption removal: drop the comment naming the provider's saved message, and make requestedAt required since every pending input comes from input.accepted. * Use current provider handles in transition tests * Use current provider handles in timeline fixtures |
||
|
|
f9c8cd4fc3 | Reduce redundant test coverage and unnecessary CI waits (#25806) | ||
|
|
4e64fa9940 |
fix(native-chat): start a Codex chat without waiting on the background model probe (#23831)
* fix(native-chat): start a Codex chat without waiting on the background model probe Opening a Codex chat kicks a session-less model-catalog probe (a throwaway read-only `codex app-server`) for the account, and the new chat's own options read then joined that probe's in-flight refresh through the catalog store's single-flight. When the probe's Codex hung, chat start waited out the probe's 15s deadline and failed with "codex app-server session exceeded 15000ms", even though the chat's own app-server was up. Single-flight now joins only a refresh by the same kind of lister. A live session lists over the connection it already holds; the probe still records its own failure in the store, and the picker keeps serving whichever listing succeeded. * fix(native-chat): never let a chat's model listing join another chat's listing Single-flight in the model catalog store was split by lister kind, so a new chat's acquire-time options read no longer joined the session-less probe, but it still joined any other live chat's in-flight listing for the same account. When that other chat's Codex was wedged or closing, the new chat waited out that chat's request timeout and failed to start with its error. Key in-flight listings by the lister's identity instead: a live session by its own connection, the probe by the probe itself. A lister still joins its own in-flight listing, and shouldRefresh still holds back a probe while any listing for the account is in flight. * test(native-chat): pin that the catalog store releases an account once every listing settles shouldRefresh reads 'any listing in flight for this account' from whether the per-account in-flight map exists, so that map must be dropped exactly when its last listing settles. Nothing covered this: removing the cleanup, or dropping the map on the first settle, passed every catalog test. A leftover map would stop every later background refresh and probe for the account. * fix(native-chat): name the catalog lister type instead of a bare object A live session is keyed by its per-spawn catalog handle (minted in the same acquire as its connection), the probe by itself. * test(codex): pin that a chat's own concurrent option reads share one listing * fix(native-chat): keep model picker current across parallel listings |
||
|
|
020cebeff6 |
Add standalone Agent Client Protocol client layer (#24990)
* Add standalone ACP protocol client and session runtime * Protect ACP transport teardown from late stream errors * Retire incoming ACP request ids before publishing responses * Narrow ACP configuration requests and transport message types * Remove redundant ACP request handler return unions * Keep ACP waits caller-owned and preserve protocol extensions * Preserve open ACP decisions through prompt completion * Generate open ACP enums and check the generated schema offline A newer or vendor enum value (tool kind, tool status, option kind, stop reason) no longer fails the whole message: generated enums accept the known literals plus any other string, typed so callers can still narrow on the known ones. The generated header now records the pinned input digests, the generator digest and a body hash, so `verify:acp-protocol` catches a stale or hand-edited file without network access; it runs in lint and the PR workflow. * Land the ACP runtime contract the agent adapters use - Deliver notifications other than session/update through onExtensionNotification, in arrival order with session updates. - Accept _meta on prompt, setMode, setModel, setConfigOption and cancel. - cancel() always sends session/cancel once the session runs, since the agent can be in a turn it began itself; only a successful send is shared, so a failed write is retried. - Cancel aborts each open agent request's signal and lets its handler send its own answer; -32800 only when the handler rejects. - Permission requests validate only the session, tool call id and options; unreadable fields are dropped with a diagnostic, and any answer Orca cannot send is `cancelled` instead of a JSON-RPC error. Agent-started turns may ask; whether to show it is the caller's decision. - AcpAgentError marks the agent's own errors; AcpInvalidResponseError keeps the raw answer and validation issues for answers Orca could not read. - Lines over the size limit are classified by prefix (shared with the Codex reader): the owed request fails, an oversized agent request is answered with an error, and an unattributable response closes the connection. * Answer every agent request after an ACP cancel A cancel that lands before a permission handler starts now still runs the permission path, so the agent gets the `cancelled` outcome rather than a request-cancelled error. A handler that ignores the abort no longer leaves the agent waiting: once the abort has run through, any request still unanswered gets request-cancelled. Handlers that answer on abort keep their own reply. Also renames a lint-rejected helper parameter, replaces a Reflect.apply in a test, and stops the permission diagnostic from firing with an empty list. * Let each ACP request handler own its answer after a cancel Removes the next-event-loop-turn fallback that answered request-cancelled for any handler still silent after a cancel. It raced answers that were still being saved (an approval mid-journal-write reached the agent as an error) and made the outcome depend on event-loop timing. The handler that owns an agent request now always sends its answer, or throws for request-cancelled; a request it never answers ends when the connection closes. A permission whose handler had not started still answers `cancelled`. * Register the ACP schema verify step in the PR preflight phase test * feat(acp): a steer's cancel asks once and never ends the agent The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle, then close the connection, which ends the agent. A steer used it too, so a slow agent lost its process just because the person added a message. requestSteerCancel() now sends session/cancel once per prompt, cancels the agent's open requests and answers later permissions cancelled, and never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel() stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel paths move into acp-prompt-cancel.ts over one cancel channel. * fix(acp): a repeated steer shares the cancel in flight; say what the caller owns Per review: a second steer before the first write lands returns that write instead of resolving early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that fails instead must not take the steer until the caller rebuilds the session; the Stop's says a prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives the runtime a handler that would allow: the open permission's signal aborts and the late one never reaches it. * test(ratchet): require src/main/acp now that this PR lands it |
||
|
|
635e7f53cb |
feat(github-projects): render Board project views as a kanban with drag-and-drop (#19074)
* Add board layout support for GitHub project views Board-layout views render as a kanban. Columns come from the view's verticalGroupByFields (the host retries without the selection on older GHES schemas and the renderer falls back to the Status field), with one column per single-select option in option order — empty ones included — plus a trailing no-value column whose drop clears the field. Card drops reuse the table's field mutation path, committed from a document-level capture listener because the preload's native-drop bridge stops drop events before React's root ever sees them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(github-projects): harden board capability probe and drop lifecycle Review follow-up. The verticalGroupByFields capability probe now matches parsed GraphQL error messages instead of substring-scanning the whole response body — partial-error responses echo the field name as a data key on healthy schemas, so one SAML/FORBIDDEN partial error could permanently degrade github.com boards for the session. Covered by new project-view-config tests per the capability-cache testing contract. Also: only single-select/iteration vertical fields shape columns (a drifted field kind no longer yields a clear-on-drop no-value column), optimistic patches resolve the column field through the board/group config, drag cleanup moved to a document-level dragend listener, the column dot follows dark mode via the chip CSS variables, the supported- layout allowlist is a single shared predicate, and the board's edit handler is stable so its drop listener stops re-registering per render. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(github-projects): serialize board edits and verify rendered drops * fix(github-projects): preserve refresh baselines and order edits across views * fix(github-projects): accept source settings projection in cache scope * fix(github-projects): distinguish view switches from refreshed field baselines * test(github-projects): type board IPC recordings and verify clear request --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
ca4e239861 | Remove low-value test inventories and duplicate fuzz oracles (#25791) | ||
|
|
c0b07d3a71 |
Admit one provider event's journal writes as one queued operation, decided when it runs (#25141)
* Move the turn message ordinals and the turn-row revision to the neutral timeline folder Pure moves so a shared timeline assembler can use them: Codex's message ordinal counter becomes ProviderTurnMessageOrdinals and Claude's turn-row revision becomes the provider-neutral agent-journal turn-row revision. Only names and import paths change. * Admit one provider event's writes as one transition, and let rows be found again after a restart - A sink transition is admitted whole or not at all; its steps run back to back at their turn in the journal's write queue, and each resolver reads the fold with every earlier write landed. A resolver may also say where the row belongs (turn scope, provider reference), and the writer always hears how the transition landed. A resolved lifecycle batch chooses its settlement mutations from the fold at execution. - New optional row field providerItemRef: the provider's own reference for the item a row is, written only where the row's identity cannot spell it (Codex keys messages by their place in the turn and renumbers its item ids on resume). Set by the creating write, kept by revisions, indexed by the journal fold, never read by clients. A downgrade test shows an older host and client render such rows unchanged. - Provider timeline identity schemes (shared legacy arm, Codex) and the join index that resolves a provider item to its row from memory or the fold: ordinals and request incarnations are read back from the rows, so a restart or an evicted entry finds the original row instead of placing a new one. * Recover message ordinals from the journal's highest place, and forget joins read from a replaced epoch A fresh join index continued a turn's messages at the first free place, so a journal holding only a later ordinal (an imported or removed earlier row) had its sequence back-filled. The place is now one past the highest ordinal any row or echoed send holds there, read through a pure scheme reader. The join caches also drop what they read when the journal's epoch is replaced. * Spell the subagent thread's message slot without spreading an identity union * Keep the journal store under its line limit after the main merge * Drop the provider item reference, join index and identity schemes from the transition PR Nothing in production reaches the state they defended (an assembler that lost its memory while its child keeps streaming the same turn), and the stored Codex id was positional. The legacy identity scheme moves to the assembler PR with its first caller; the Codex scheme and any persisted reference wait for Codex to move onto the assembler. The Codex ordinal counter goes back to codex/, since no neutral code imports it. * Write a resolved settlement in one transaction through enqueueRows A settlement too large for one row now commits all its rows or none, through the journal's existing all-or-nothing write, instead of a row-by-row writer. Every row is built before any commits, so a settlement naming one item twice is refused before anything is written. * Drop the transition's landing report; keep the turn-row write fire-and-forget Nothing reads which steps of a transition wrote. A failed step fails the sink, leaving the steps before it written; the header says so, and tests cover it plus a settlement whose second row fails inside the transaction. writeAgentJournalTurnRow returns nothing again, as on main. * Run a transition's steps as a prefix; drop the paced flag and resolved options A failed step no longer lets the steps after it write: each step checks, at its own turn in the journal's queue, whether the write handed over just ahead of it completed, using the queue's count of completed write bodies (a promise would report the failure only after the next step ran). The sink fails only once every step has had its turn. The item step's `paced` size bypass and the resolver's replacement `options` are removed; nothing planned uses them. * Run a transition's steps in one queued write that loops over them The steps of one event now share one turn in the journal's write queue: a loop writes each in its own transaction through the row writer's synchronous writeRows (split out of enqueueRows) and stops at the first throw. Prefix semantics and "nothing lands between the steps" now hold by construction, so the completed-write counter on the queue, the step gate and the allSettled barrier are gone; the queue is back to main's bytes. * Use current provider handles in transition tests |
||
|
|
dc8b3934b2 |
Add chat appearance settings: text size, code size, width (#25657)
* feat(native-chat): add chat-scoped color tokens * feat(native-chat): soften transcript and composer appearance * fix(native-chat): refine code spacing and faint text styling * fix(native-chat): wrap prose links at word boundaries * test(native-chat): refresh background task strip snapshots * Add native chat appearance settings and card scaffold * Persist chat text and code sizes with appearance width controls * Fix chat appearance typing, shortcut routing, and scoped typography * Use chat call identities in typography regression fixture * Keep markdown metadata typography independent of chat text size * Preserve terminal undo chords and stabilize chat resize estimates * Show only primary shortcuts in chat appearance settings * fix(native-chat): preserve appearance edits and configured zoom shortcuts * Remove unused chat appearance summary translation * Preserve command classification when applying chat code size |
||
|
|
6c693edf40 |
Show native chat tool calls as plain sentences (#25654)
* feat(native-chat): add chat-scoped color tokens * feat(native-chat): soften transcript and composer appearance * fix(native-chat): refine code spacing and faint text styling * fix(native-chat): wrap prose links at word boundaries * test(native-chat): refresh background task strip snapshots * feat(native-chat): show tool calls as plain sentences * fix(native-chat): make tool sentences reflect call state * fix(native-chat): clarify failed commands and subagent sentences * test(native-chat): exercise command disclosure with real results * fix(native-chat): preserve command inputs and localize failure rows * fix(native-chat): retain complete padded command input * fix(native-chat): read named tool input fields safely * Refresh native chat tool rows when the UI language changes |
||
|
|
8d2e9264bc |
fix(native-chat): ignore the terminal launch command when starting a native chat (#25720)
A custom launch command in settings made structured native chat unavailable: the renderer route, the host launch-mode decision, and the host's create-support all treated it as terminal-only, so a user with the chat default on silently got a terminal. The command names the CLI binary and now applies to terminal launches only; native chat ignores it. A start directory outside the workspace still forces a terminal. The internal route field and blocker are renamed to that concept; the wire reason `tui_launch_command` is kept and its receipt text now names the start directory. |
||
|
|
54ded3bc18 |
fix(relay): commit the cell counter in one round trip; cells boot without the database (#25765)
* fix(relay): commit the cell counter in one round trip; cells boot without the DB Step 2 (option B) cell image: - One-round-trip counter commit at acquireActivity, releaseActivity and activateControl: the final counter UPDATE and COMMIT go as one simple-query message. Server errors mean COMMIT never ran (retry as today; 22012 = no row, rolled back and disambiguated outside the transaction); a lost connection is never retried. - Cells skip the schema apply and region backfill, so they listen while the database is down and turn ready on their first successful query. - G13: rehome target connection headroom folded into the existing NOWAIT UPDATE, excluding the host's own reservation by key. - fixLevel on every runtime metrics line, plus declared (not applied) cell fix-level metrics and alert. - Per-desktop drain disconnect-gap measurement from existing log lines. - Census test that fails on floating database promises; fixes two shutdown sites. Lock-wait sample keeps the combined role. * fix(relay): make the outdated-image alert creatable: one PromQL condition, 1 h lookback, fixed floor A PromQL condition must be the only condition in its policy, and alerts on log-based metrics may look back at most 25 h. Replace the 6-day/7-day design with relay_cell_min_fix_level (tfvars, raised by a targeted apply after each wave) and one query: a serving cell below the floor or reporting no level, sustained 6 h. Drops the separate without-level metric. * fix(relay): review fixes: gap-script ordering, wider promise census, fused-path guard, row-busy as scheduled - Drain gap script: sort closes by time (gcloud exports newest first) and refuse an invalid drain start. - Census: any floating promise in relay src, including callback-discarded and never-read ones, with a reviewed never-rejects list. - Test the fused counter commit through the store the server builds, so a wrapper that stops forwarding commitWithFinal fails CI. - Same-cap shadow gate: a row-busy refusal (the host's own release still holds its row) is a scheduled 503, like an own early retry. No client change. * test(relay): judge drain redials by no host refused twice, not a refusal count The row-busy count tracks how many releases are still in flight at the dial (80 of 180 every run at 1 s, against a bar of 90). What matters is that the release has finished by the next dial: assert no host is refused twice, keep the time-to-placed p95 bound. * fix(relay): cap row-busy as scheduled at the drain-return admissions; bound the gap script's window Shadow gate: a row-busy refusal of a drained host follows its drain-return lane admission, so per minute only that many (plus a rounding margin of 2) are scheduled; the rest stay non-drain, so row contention the drain does not explain still fails the budget. Gap script: --drain-ended-at excludes the new container's closes after the roll; later grants still close a gap. |
||
|
|
fff8718c76 |
Improve Quick Open matching, file locations and recent history (#25371)
* Match Quick Open queries across path terms and identifier separators * Honor ignored-file and symlink preferences on file inventory hosts * Open pasted file locations and remember successful workspace file visits * Verify bounded directory listing preferences across execution routes * Preserve symlink preferences on legacy directory fallback * Preserve bounded Quick Open terms and separator alternatives * Revoke stale Quick Open selections and preserve literal file locations * Validate recent Quick Open candidates independently on their host * Use complete old-host inventory fixtures for Quick Open compatibility * Keep Quick Open validation current across palette and workspace changes * Route recent candidates through relay file listing dispatch * Record Quick Open hook renders as test snapshots * Localize Quick Open eligibility errors and search aliases * Validate file inventory with bundled search and runtime preferences * test: use directory junction fixtures on Windows * fix: negotiate Quick Open policies with nested SSH hosts * fix: refresh cached linked folders after target changes * Retry stale expanded links after reads settle and targets recover * Scope stale refresh request lifetime to the visible workspace * test: seed linked-folder fixture before watcher startup * Keep Quick Open responsive and compatible with older hosts * Exercise directory discovery with unprivileged Windows junctions |
||
|
|
c8ff7e8875 |
Stop repeated recovery of finished workers (#25679)
* Settle completed worker assignments and bound recovery retries * Reset queued explicit recovery before starting its retry budget * Use current provider identity in the inherited journal fixture * Keep late worker recovery automatic without repeating persistence work * Repair inherited validation fixtures after updating main * Keep the released parser available to compatibility tests * Limit recovery queries and background process inspection to pending work * Verify historical repair safety and live SSH process retention * Preserve the same provider import as main in the merge result * Reduce historical repair and duplicate workspace recovery work * Align startup prompt fixture with upstream Windows coverage * Keep healthy worker repair indexed and read-only * Align startup prompt tests with current main * Keep pending terminal releases progressing during recovery retries |
||
|
|
e2ccf52dc3 |
fix(remote-runtime): a paired terminal accepts input again after its host app relaunches (#25736)
* test(remote-runtime): reproduce dead input after a paired host relaunch A relaunched desktop host accepts RPC before its renderer publishes a window graph. During that gap session.tabs.list answers with an unpublished empty graph (publicationEpoch "none", snapshotVersion 0, tabs []). The client's reconnect inventory treats that as removal and retires the pane (retireRemoteTerminalId(-1)): the pane goes to "ended", the reconnect overlay disappears, and keystrokes never reach the surviving daemon PTY. Both tests are red on main by design; they are the repro for the fix. - e2e: holds the relaunched host's runtime:syncWindowGraph so the client's reconnect deterministically meets the unpublished host (4/4 red). - unit: transport-level repro of the same retirement (red in ~2s). - restart helper gains a beforeFirstWindow hook; LaunchOptions moves to its own module to stay under the max-lines budget. * fix(remote-runtime): a paired terminal survives its host app relaunching After the host app quit and relaunched (daemon still running), a paired client's terminal cleared its reconnect overlay and then ignored all input. The relaunched host answers session.tabs.list before its renderer publishes, with a synthesized empty frame. A paired client sees that frame through its navigation projection as epoch "none:client-navigation", which the shared "does this frame answer for the worktree" check did not recognise, so the pane read it as removal and retired itself. - hostSnapshotAffirmsWorktreeContents treats the client-projected placeholder as no answer at any version (the projection adds its navigation revision). - The reconnect inventory keeps polling on such a frame instead of retiring; a published frame lacking the surface still retires the pane. - Pushed frames that affirm nothing no longer report the surface absent. - When the bounded inventory wait ends without evidence, the first published snapshot carrying the same handle now reattaches the pane, rather than waiting out the ~3 minute auto-recovery deadline. Client-only: the host already labels the frame, old hosts send the same frame, and the wire is unchanged. The e2e helper drops its type assertions. * refactor(remote-runtime): simplify the host-relaunch reconnect fix - One epoch rule: the placeholder epoch, bare or with the shared client-navigation suffix, is no answer. Hosts only ever send it at version 0, so the version check is dropped. - Unpublished frames are dropped where pushed snapshots enter, instead of a nullable per-subscriber update. - A published same-handle snapshot fires the parked retry through retryNow(), which now also fires a retry parked while still 'recovering' (nothing is in flight then). This replaces state-reading wiring in the listener and lets online/resume fire it too. - The Windows e2e runs the fixture through PowerShell's call operator. * test(remote-runtime): prove the e2e hold beat the relaunched host's first publish The spec now asks the held host for its tab list and requires the unpublished placeholder before releasing, so a hold installed too late fails instead of passing against unfixed code. The helper keeps its gate in a typed global instead of Reflect lookups. * fix(remote-runtime): keep a fenced handle's parked retry revivable A retry parked for a handle that needs a replacement was consumed by an early external trigger and then refused by the epoch's snapshot-wait guard, leaving nothing to revive the pane. An external trigger now takes over that snapshot wait as a fresh attempt, and a republished fenced handle no longer fires the retry. Also refreshes comments that still paired the placeholder epoch with version 0. * fix(remote-runtime): a host answers an empty worktree for real once it has published The host sent its unpublished placeholder for any worktree it had no entry for, including one its window had never opened, long after startup. With the placeholder now read as "ask me later", a paired client opening such a worktree never got its first terminal. The host now sends the placeholder only until its graph first publishes. Afterwards a worktree with no tabs gets a real empty answer, and a client that was told "ask me later" during startup is sent that answer once on publication. Against an older host the client withholds the automatic first terminal for such a worktree, which the user can still create. A host push now fires only a parked retry, never cutting a scheduled backoff short. * chore(reliability-gates): reference the empty-worktree tests in gate commands * test(remote-runtime): find the never-opened worktree by repo id on every platform |
||
|
|
661372d514 |
Cancel abandoned file searches and prevent stale results (#25370)
* Bound relay git-grep records and clean up capacity failures Credits @OrcaWin for the original bounded-record proposal. * Bound filesystem listing and transfer metadata at the execution host Apply mobile limits before transport, list Markdown through its semantic producer, and retain complete directory results only within explicit capacity budgets. Stream SFTP directory packets and reject oversized transfer plans before reporting success; preserve narrow older-peer fallbacks and validate streamed response retention. * Release inactive Markdown candidates and support folder scopes Keep one current document snapshot per consumer and attach completion candidates to the actual editor model lifetime. Resolve folder workspace roots through existing workspace identities so their previews and completions receive the same authorized listing as worktrees. * Preserve full runtime inventories under aggregate byte budgets Leave unqualified inventories complete beyond 20,000 files. Charge retained paths and serialized bytes at local and relay producers, reject oversized full inventories explicitly, and bound old-peer response reassembly; keep caller-specific mobile and explicit result limits. * Bound SFTP directory handle cleanup waits Send CLOSE even after cancellation, stop waiting after five seconds without an acknowledgement, and ignore late callbacks. Preserve the original capacity or cancellation error and close handles returned after an aborted OPENDIR. * Preserve registered root spelling in Markdown document paths * Reject incomplete SFTP directory cleanup and final cancellation * Bound pending response bytes before stream ownership arrives * Preserve bounded directory reads on legacy SSH hosts * Allow bounded per-chunk padding at response byte boundaries * Bound response payload retention after stream metadata * Bound complete legacy Quick Open inventories and cache retention by bytes * Localize Markdown document listing fallback * Exercise relay search decoding with real byte streams * Preserve concrete relay stream fixture types * test: compare distinct bundled ripgrep platforms on Windows * Preserve Markdown discovery across Windows and legacy SSH hosts * Make WSL process termination available to the bounds layer * Cancel abandoned file searches and preserve search keyboard navigation * Invalidate searches when workspace roots or owners change * Cover root replacement and legacy search cancellation together * Exercise search cancellation through the current desktop access seam * Keep search ownership and cancellation contracts current * Share local text-search execution and close canceled subscriptions |
||
|
|
f59194d481 |
Bound file-search inventories and release abandoned Markdown data (#25369)
* Bound relay git-grep records and clean up capacity failures Credits @OrcaWin for the original bounded-record proposal. * Bound filesystem listing and transfer metadata at the execution host Apply mobile limits before transport, list Markdown through its semantic producer, and retain complete directory results only within explicit capacity budgets. Stream SFTP directory packets and reject oversized transfer plans before reporting success; preserve narrow older-peer fallbacks and validate streamed response retention. * Release inactive Markdown candidates and support folder scopes Keep one current document snapshot per consumer and attach completion candidates to the actual editor model lifetime. Resolve folder workspace roots through existing workspace identities so their previews and completions receive the same authorized listing as worktrees. * Preserve full runtime inventories under aggregate byte budgets Leave unqualified inventories complete beyond 20,000 files. Charge retained paths and serialized bytes at local and relay producers, reject oversized full inventories explicitly, and bound old-peer response reassembly; keep caller-specific mobile and explicit result limits. * Bound SFTP directory handle cleanup waits Send CLOSE even after cancellation, stop waiting after five seconds without an acknowledgement, and ignore late callbacks. Preserve the original capacity or cancellation error and close handles returned after an aborted OPENDIR. * Preserve registered root spelling in Markdown document paths * Reject incomplete SFTP directory cleanup and final cancellation * Bound pending response bytes before stream ownership arrives * Preserve bounded directory reads on legacy SSH hosts * Allow bounded per-chunk padding at response byte boundaries * Bound response payload retention after stream metadata * Bound complete legacy Quick Open inventories and cache retention by bytes * Localize Markdown document listing fallback * Exercise relay search decoding with real byte streams * Preserve concrete relay stream fixture types * test: compare distinct bundled ripgrep platforms on Windows * Preserve Markdown discovery across Windows and legacy SSH hosts * Make WSL process termination available to the bounds layer |
||
|
|
b58f8197dc |
Soften native chat colors and code surfaces (#25651)
* feat(native-chat): add chat-scoped color tokens * feat(native-chat): soften transcript and composer appearance * fix(native-chat): refine code spacing and faint text styling * fix(native-chat): wrap prose links at word boundaries * test(native-chat): refresh background task strip snapshots |
||
|
|
bfd1e9c579 |
fix(codex): keep working status when Escape closes search or permissions (#25769)
* fix(codex): preserve working status when Escape dismisses a view * test(codex): cover navigation Escape during terminal exit cleanup |
||
|
|
3ee3a41b6f |
fix(markdown): strengthen dark table grid lines (#25653)
Strengthen only dark-mode table cell borders in Markdown Preview and the rich editor by mixing foreground into the existing border token. Keep light mode and table geometry unchanged. Fixes #24982. Consolidates #24984, #25050, and #25653. Co-authored-by: Paramon <andrii.paramonov@gmail.com> Co-authored-by: kana001-bit <288527232+kana001-bit@users.noreply.github.com> Co-authored-by: Neil <neil@stably.ai> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
5f7c8483e9 |
test(cross-version): load releases that import a package main dropped (#25731)
* test(cross-version): stand in for packages a release declares but main dropped The newest release (v1.4.221) imports @streamparser/json, which main removed. The harness runs release source against the current install, and the release's RPC dispatcher imports its whole method table, so every wire suite failed at import time. Extraction now gives each package the release declares and the current tree no longer declares a stand-in in the checkout's own node_modules that loads but throws, naming the package, on any use. Co-Authored-By: Claude <noreply@anthropic.com> * test(cross-version): decide stand-ins by package resolution, not manifest diff A package gets a stand-in only when release source imports it at runtime, the current tree does not declare it, and it does not resolve from the checkout. A package still reachable from the checkout keeps the real install, and one the current tree declares but cannot resolve stays a loud failure. Type-only imports, comment prose and scripts in strings are ignored. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
3285f8214c |
Share the managed process lifecycle for structured providers (#25204)
* fix(native-chat): a child's root exit is reported even during its close, and bookkeeping after it never reads as unproven - Both connections report the root process's exit once, with `expected` set when a close had begun. A close that came back unproven and whose root exits later is finished by the adapter, and its end reaches the host like any other. - A Claude close whose resume-point write fails after the exit was proven, and a Codex close whose terminal row is refused, now end the session and report the failure, instead of keeping a dead child indexed as if its exit were unproven. - A Codex close whose forced tree kill can't prove the descendants gone but saw the root exit reports the descendants and counts the root exit. - Every child exit with an identity, expected or not, is forwarded to the host. * fix(native-chat): the exit ends the child's record; an unfinished stop is the child's own close, which everyone joins - The host keeps no stored "stop still owed" record any more. A stop begins the child's close (`child.close`), which lives on the child and ends with it. A second Stop, the idle reaper, quit, a send and an option/answer/goal/rewind all join that close instead of retrying a separate obligation. - A caller waits on the close only as long as the step deadline; the close itself is never abandoned. A proof that lands after every caller stopped waiting reaches the host as the adapter's report of that exit, which ends the record through the same handler. - Once the exit is proven, draining, settling, the lease release and the adapter's acknowledgement are each attempted and reported on failure; none keeps the child on record. A start, and the handle's close, write a release that failed from this host's proof of that exit, so a failed write never refuses a send. - A start that meets a close still unverifiable is refused with `previousExitUnverifiable`, so the queued message is rejected with a send-again reason; nothing is held and nothing starts beside the old process. - The idle sweep goes back to idle reaping only. - Removes #24333's retry entry points, the wait row and its hold rule, the ask/failure cursors on the stored record, and the stop's own wake. Tests replace the #24333 unproven-stop test: a send joining an unproven close and an in-flight one, a late proof past the caller's bound, a root exiting after its close gave up, a proven exit whose resume-point write and lease release both failed, an unverifiable close rejecting the send and refusing an option change, a surviving descendant, quit and the idle reaper; and Codex's unverifiable, late-exit and joined-close cases. * fix(native-chat): a message refused because the old process's exit is unverifiable says so, and to send again The start failure for a refusal with reason `previousExitUnverifiable` reads "Orca couldn't confirm Claude's previous process ended. Send your message to try again." instead of "Claude couldn't restart." The status-row kind and the refusal reason stay in the shared lists for rows and hosts that still carry them; the catalogs keep one sentence for both. * fix(native-chat): a close's verdict is the root's exit alone, and what follows it is logged - A Claude close resolves as soon as the root's exit is proven: the session ends and its `ended` report goes out then. Saving the resume point runs afterwards and a failure is logged, so a slow or hung write never reads as an unproven exit or keeps a dead child on record. - A root that exits after its close came back unproven finishes that close through the same path as any close, so the session's child work is published as ended (background tasks and subagents no longer stay shown running for a dead agent), and a failure there is logged. - Codex logs a refused final row, and reports a root exit whose forced tree kill could not prove the rest of the tree gone the way Claude does, so the host logs it and blocks nothing. - Both adapters take the host's logger for this bookkeeping. * fix(native-chat): one handler ends every child's exit, and a join waits on the adapter's own close - One exit handler (`structured-agent-session-child-exit`) ends a child's record for an exit expected or not. `expected` only changes what the chat is told: the stop's cause, its end at the stop's ask, the settlement id, and no crash outcome row. The lease release keeps the exit's evidence; the handoff guard, lifecycle barrier, sink release and adapter acknowledgement apply to both. A Claude journal-sink failure ends in the same step as its stop, as Orca's own fault. - Joining a close is asking the adapter, whose close is memoized while it runs and bounded by its own kill escalation; the host keeps no attempt of its own and no 10 s caller bound. An ask after a close came back unproven runs the stop again. - A close's end is stamped where its stop was asked for (a repeated ask moves it), so the closed chat and failed start checks order a message accepted meanwhile after it. - A start refused because the old exit is unverifiable rejects what was queued in the same step. - The end of a close the host asked for no longer waits on the cross-session recovery chain. - The kill no longer waits for the stop event's write; the journal writes rows in order. * fix(native-chat): an exit's lease release lands whatever the length of its reason A crash's reason can carry kilobytes of the provider's stderr, and a lease whose death detail is over 512 characters fails the store's own check. The exit handler cut it, but the release a start or the chat handle's close re-derives did not, so after a crash whose own release failed every message was refused as not resumable until restart. The record's builder now cuts the detail to the record's bound, so no writer can hand it one too long. * fix(claude): a proven close waits at most 2 s for the output it already wrote Once the root's exit is proven, the close still waited for the SDK's output reader to end. Something outside the process tree that holds the output open would keep that close, and every send, Stop and quit joining it, waiting with no bound. The wait is now bounded; past it the close resolves as proven and the open output is logged. * fix(codex): an exit reported inside Orca's close keeps the reason Orca closed it for The connection reports the app-server's exit inside the close that ends it, so that report ended every Codex close and replaced the close's own reason (for example, a provider frame that could not be recorded) with the connection's stderr text in the ended record and the lease's exit evidence. The session now records Orca's close with its reason, and the exit it ends keeps that reason. The test connection reports its exit inside close the way the real one does. * fix(native-chat): quit stops delivery before it drains exit recovery Every exit now wakes delivery, and teardown drained exit recovery before it stopped delivery, so an exit settled in that window could start a fresh agent that teardown then killed. Teardown stops delivery first; queued messages wait for the next launch. * docs(native-chat): the unverifiable-exit refusal no longer names a caller's wait The caller's bounded wait was removed; the comment describes the close as it is now. * fix(native-chat): a stop whose kill did not take is logged, and the next ask kills again When a close's kill leaves the agent's root running, the host now logs it. Tests pin what a later ask does: each connection runs its whole stop again (Codex sends SIGKILL a second time), refuses input meanwhile, and proves the exit once the kill takes. * fix(native-chat): a start refused over the old process says Orca couldn't stop it The host reaches an unverifiable verdict only after its own kill left the agent's root running, on the machine that runs the agent, so the sentence now says that: "Orca couldn't stop {agent}'s previous process." The refusal reason, failure kind and wire shapes are unchanged. The host test also checks the failed kill is logged. * docs(native-chat): an unverifiable close verdict is a root that survived the kill The host's close runs where the agent runs, so lost contact never yields this verdict; the comment no longer says it does. * fix(native-chat): a kill that did not take is reported once, by whoever met it The log added at the close fired beside a Stop's own failure report for the same event. A stop still reports it through its failure; a send or option change refused over it now logs it at the refusal, the only place it is otherwise invisible. * test(native-chat): a second Stop joins a close the first could not prove and retries its kill * fix(native-chat): say a start refused beside an unstopped process plainly The rejection now reads "Couldn't stop {{agent}} from before. Send your message again to try once more." This kind has its own send-again step; every other failure keeps "Send your message to try again." * Move provider process supervision and stream reading out of Codex * Preserve teardown behavior with checked mock types after move * Apply provider launch environment and caller teardown labels * Share managed provider launch, exit observation, and close * Preserve synchronous stdin closure before the grace wait * Keep the stacked Claude adapter within the module size limit * Handle nullable provider stdin and processless fixture identity * Preserve close-time cleanup diagnostics after a late provider exit * Separate provider root and descendant exit observations * Give provider child env one owner and gate Codex contract on the shared reader resolveProviderChildEnv is now the only place that overlays and strips a provider's environment; the spawn spec and the request-scoped Codex session both call it. supervisedPosixLaunch only accepts a launch without env fields, so an override can no longer be silently ignored there. Edits to the shared stream reader or the env rule now run the real-binary Codex contract job. * Give every provider one close result, stderr tail and root-only default - A close reports the root and the descendants with the one verdict vocabulary (live / unverifiable / exited); `tree` is null when the close made no descendant observation, and the fallback teardown's outcome is written into it instead of a side flag a caller could miss. - The managed process drains stderr and keeps the 8 KiB tail, so no provider can forget to drain the pipe. - Root-only providers take the default close policy and completion rule; only Claude overrides them. The one already-exited guard lives in the managed close and is recorded when no earlier close ran. - A failed spawn is never read as an observed root exit. * Report only observed descendant exits from the fallback teardown The fallback teardown answered "accepted", and the close turned that into `tree: 'exited'`, though a Windows tree kill's outcome is never read and an unreadable process table observes nothing. It now returns what it observed: `exited` only when the captured descendants were verified gone (or the kernel reported the group empty), `live` / `unverifiable` as verification found them, and null when it signalled without observing. Codex's process-tree diagnostic fires in exactly the cases it did before. * Give two test fixtures' casts a SAFETY rationale for the changed-code gate * Claim no observation from an empty process group ESRCH from the dedicated-group signal says only that the group is empty: a descendant that left it, or a root that never led it, may still be running. It now reports no observation instead of `exited`. The unreadable-table comment says why that case stays "no observation" for root-only providers. * fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts On main every close of the chat writes its own Stop and settle. Here a later close joins the first and writes no row, and the first's settle closed when its kill failed, so a turn that opened in between and was cut by the next close read as failed. A person's close joining a person's close whose Stop opened a settle now reopens that settle until its attempt is done. Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and the next reads as the person's cancellation (each fails without its half of the fix). * test(native-chat): build the Stop-opened-turn test's identity with the opaque handle The test (#25056) landed before the opaque provider handle (#24991), so main still built the old {kind, threadId} handle. * test(ratchet): require src/main/provider-process now that it has landed * test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test Main's #25706 made the same fix as this branch at a different line; the merge kept both imports. |
||
|
|
3a03441580 |
refactor(native-chat): structured agents declare their capabilities instead of shared code naming Claude and Codex (#25076)
* refactor(native-chat): keep the provider resume handle opaque to shared code
Shared structured-chat code parsed each provider's resume handle: Claude's
session id and branch leaf, Codex's thread id, through a 'claude' | 'codex'
union every new agent had to widen. The in-memory handle is now
{ transport, agent, nativeId, providerData? }: shared readers use nativeId,
lease and handle-chain checks compare transport and agent, and only the
Claude adapter reads its leaf (providerData).
Stored and wire forms are unchanged for Claude and Codex. One encoding
module writes their typed shapes and decodes both those and the neutral
shape a new transport uses, which an older build refuses as unreadable
rather than reading as Codex. Key and root strings, which fork seeds,
superseded creations and resume offers persist, stay byte-identical.
The journal's own handle type becomes the journal-row and attach-wire
encoding of the same handle, and the journal identity carries the
neutral handle (null before the provider proves one).
No user-visible change.
* fix(native-chat): derive journal-row provider handles from the journal identity
The journal row converter now takes the identity every caller already holds,
so a row's handle has one obvious constructor. Tests that wrote the in-memory
handle straight into journal rows now build it through that converter, and the
processless Claude fixture names a not-yet-proved handle as null.
* fix(native-chat): refuse a stored provider handle written in both forms
A typed Claude or Codex handle that also carries the neutral form's
transport, agent, native id or provider data named two identities; it
was read as Claude or Codex and the next write dropped the other one.
Such a row now stays unreadable and is set aside untouched.
* refactor(native-chat): route structured agents through registered definitions
The structured-chat shared layer named Claude and Codex in a closed router
type, in option reads, in frame classification, and in renderer label and
validation branches, so adding an agent meant editing shared types. Each
agent now has a definition (capabilities, resting option rules) owned by its
own module; the router is a registry of definition + adapter pairs built at
runtime setup, and shared code reads the definition instead of the name.
No behavior change for Claude or Codex; no wire or stored shape change.
* refactor(native-chat): name the structured agent list once in host types
The host types, conversation commands, provider-session ownership and the
runtime's create path spelled the structured agent list inline; they now use
the one shared type the record and wire already validate against.
* fix(native-chat): narrow the record before reading its agent's option rules
* refactor(native-chat): make the router's registrations the only agent definition lookup
The at-rest option reads looked definitions up in a second, separately populated map, so a
registered definition could route and declare capabilities while the static map decided which
options a chat at rest accepted. The router now answers `definition(agent)` from its
registrations, the at-rest readers take it through the host's adapter, and the built-in
definitions are listed only where the runtime composes the router.
* test(native-chat): use opaque handle in queued rejection fixture
* test(native-chat): share one Codex journal identity in the integration suite
Main grew the suite to the 800-line limit; the opaque-handle import pushed it
over. The two tests built the same identity inline.
* refactor(agent-session): name the handle's adapter state resumeCursor
Rename the neutral provider handle's providerData to resumeCursor before any
row persists the neutral form: it is an adapter-owned resume position (Claude's
transcript leaf), never identity. Claude/Codex stored and wire bytes are
unchanged; their typed shapes never carried the field.
State the stored-form contract (a handle's field set is closed; later per-link
data goes on the chain link, which every build preserves) and pin it with a
record round-trip test. Document that transport records the id space the
native id was minted in, which can differ from the agent's current transport.
* refactor(agent-session): one required agent registry; declarations admit what they claim
A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.
/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.
Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).
* refactor(agent-session): the router applies the declared rewind itself
The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.
* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record
* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way
The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.
* test(ratchet): require src/main/provider-process now that it has landed
* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test
Main's #25706 and this branch both added the import at different lines; the merge kept both.
|
||
|
|
db078b6549 |
fix(claude): API-key Claude users are told they're not signed in (#25163)
* fix(claude): open native chat for API-key users instead of saying "not signed in" Claude reports tokenSource "none" beside apiKeySource "ANTHROPIC_API_KEY" when it runs on an API key (environment or settings env). The startup check read only tokenSource, so it refused those chats as signed out. Refuse only when Claude reports no token source and no API-key source. Co-Authored-By: Claude <noreply@anthropic.com> * test(claude): cover a Console /login key at chat start --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
7f2541373a |
Stage long agent launch lines where they are typed, so they arrive whole (step 2 of 7) (#23962)
* feat(agent-launch): host-side prompt delivery for agent.launch The host's agent.launch typed any launch prompt into the shell as part of the launch command. A long or multi-line prompt then ran line by line in the shell, and an agent that never showed readiness or crashed at startup had nothing guarding where its text went. agent.launch now carries a prompt on the typed line only when the line stays one line, control-free and at most 512 bytes; otherwise the agent starts clean and the host pastes the prompt once the agent's own ready signal fires (bracketed paste plus its composer marker or a quiet render, read only after the shell's last hand-off, never while the pane's own shell is proven in front), with main's draft-paste bytes and an Enter 50 ms later. Orchestration worker starts wait on tui-idle as before. A replay-safe launch admits and claims its ledger row in one write, Qwen Code gets a second Enter, the desktop and phone share one launch-refusal classifier, and hosts advertise agent.launch.prompt-carry.v1. Split out of #23748, which moves the desktop source-control buttons onto this path. * fix(agent-launch): keep a short-lined multi-line prompt on a local zsh launch line, as main did #24257 moved every multi-line or over-512-byte prompt off the typed launch line and pasted it after readiness. The phone's AI buttons and review notes, whose multi-line prompts main typed whole into zsh, then reached Claude 0.5-3 s later and their RPC reply waited for the paste. The host now names the shell a local macOS or Linux line is typed into, the way the spawn picks it, and a multi-line prompt rides a zsh line when every line is at most 512 bytes and the whole line at most 8 KB. A real-zsh test types such a line through Orca's own ready barrier and startup write, including when a slow user config makes the write land early. Elsewhere the measured unsafe cases keep the paste: bash 3.2 runs multi-line lines piecemeal, fish drops an early multi-line write, and any shell loses a line over 1 KB written early. * test(agent-launch): keep the real-zsh launch-line test out of the Windows lane's gate scan The Windows lane registration check read `const ZSH_PATH = process.platform === 'win32'` (the head of a multi-line ternary) as a Windows-true flag, so `describe.skipIf(!ZSH_PATH)` looked like a Windows-only suite. The file is POSIX-only; the zsh lookup is now a function. * refactor(protocol): move the agent.launch capabilities into their own module Main's protocol-version.ts sits at the 300-line cap, so the prompt-carry capability pushed it over. The four agent.launch capabilities and their doc move to agent-launch-runtime-capability.ts, re-exported by name and spread into RUNTIME_CAPABILITIES at the same position; the advertised lists and every export are unchanged. * refactor(protocol): import the agent.launch capabilities from their own module `export *` from protocol-version left the four names undefined under the mobile recording loader, which resolves a relative import through a Proxy with no own keys, so 37 phone recordings lost agent.launchReplay. Importers now name agent-launch-runtime-capability directly; protocol-version only spreads its list. * refactor(agent-launch): drop the unshipped viewMode field and trusted local caller id Both were inert in step 1 and existed only for step 2. agent.launch will become a public plugin API, so every wire field is permanent once shipped; a top-level viewMode reads as "choose terminal vs chat", which the host decides. Step 2 introduces placement and view intent under a placement object instead. * fix(agent-launch): read Codex's provisional startup from the rule files' hold anchor Main (#24375) moved Codex's provisional-header check into codex.json's provisional_startup hold anchor and deleted codex-terminal-readiness.ts, so the launch readiness hold now asks showsHoldAnchor, as main's own settled check does. * fix(agent-launch): hold rule-file name titles to quiet for a launch, and census the zsh fixture Main (#24375) answers a name-only title from each agent's rule file ahead of the sustained-title lane, so gemini.json's name_title settled a launch readiness wait on the shell's auto-title while Gemini was still booting. A launch now asks quiet of every weak idle verdict, as that lane did. Main's readiness census requires a recorder for every runtime fixture; the zsh prompt recording is a non-agent control. Gemini's synthetic baseline is regenerated for this PR's stated change: a bare gemini title is no longer its rest mark, so name-only rows settle weak, and a fresh working or blocked status is no longer overridden. * fix(launch): stage long or multi-line launch lines so they arrive whole A launch line that is long or has several lines could lose lines or wedge a shell that is still starting, because it was typed into the terminal before the shell could read it whole. The terminal host (local daemon, and the SSH relay) now writes such a line to a script file and types one short line that runs it, for every POSIX shell; a shell that cannot read Orca's quoting runs Orca's own agent lines through /bin/sh. If the script cannot be written, the terminal says so. The SSH background launch path no longer types the line from the renderer; the relay stages it like any other terminal. Bumps the terminal daemon protocol to 40 (main took 39); a v39 daemon keeps its sessions and gets no staged line. Restacked onto #24257 from the version reviewed on top of #23748; carries none of #23748's changes. * fix(launch): run a staged launch line as its own job in fish and ksh fish and ksh run the commands of a sourced file in the shell's own process group, so a staged agent was not a job: Ctrl-Z was dropped (a typed line stops), and the shell-in-front check saw the shell's group in the foreground while the agent ran. fish now runs the script with `eval (string collect < '<path>')`, which runs it as if typed and leaves the shell's job-control mode alone. ksh and mksh run a long Orca-built line through `/bin/sh '<path>'`, a real job; a long command the user wrote is typed as before. The real-shell test now asserts the agent's process group is its own and is the terminal's foreground group. The staging rule also reuses hasControlByte and the 512-byte budget from startup-line-prompt-carry so the carry and staging limits cannot drift. * fix(agent-launch): paste a launch prompt only when the launched agent is proven in front A launch pasted its prompt unless a shell was proven in the terminal's foreground, so any read that could not prove one let the prompt through. After an agent exited at startup, its shell turned bracketed paste on at the next prompt, readiness fired on it, and the prompt was typed into the shell: - macOS: a pane runs its shell under login, so the process-group fence's root was never the shell's group and never proved it; the cached foreground name could also still name the exited process. - Windows Git Bash and WSL: the shell-alone-in-its-job check never answers. Now one fresh read of the terminal's foreground decides: agent, shell or unknown. Only 'agent' lets a write through (paste, Enter, second Enter, reused panes too); 'shell' still drops a ready signal. A Windows host never proves the agent, so there the launch line carries the prompt at any size, as on main. * test(agent-launch): cover the Windows QA stub, a grok override that exits at once * fix(agent-launch): keep the local socket alive while a prompted launch waits for its agent A launch with a prompt now waits up to 60 s for the terminal agent to be ready before it writes the prompt, and reports not-delivered when the agent never is. The local runtime socket closes a connection idle for 30 s unless the request is a long poll, so a launch whose agent exited at startup lost its reply and the caller saw 'runtime closed the connection' instead of not-delivered. Classify a prompted agent.launch and agent.launchReplay as a long poll, as orchestration.workerStart already is for the same wait. * refactor(agent-launch): narrow the launch params by 'in' instead of a cast * fix(agent-launch): find a launched agent behind a wrapper that leads its process group A tcsh or nu launch line runs the agent from /bin/sh '<script>', and a wrapper script that does not exec its agent does the same: the wrapper leads the terminal's foreground process group and the agent is a member of it. The fresh foreground read names the group's leader, sh, so a prompted launch was refused or pasted late (M4Air tcsh: 2 of 4 not delivered, 2 pasted ~9 s late). Before that read, take the host's process-group observation as positive proof when it names the launched agent among the foreground group's members and is younger than a ready signal's quiet window. It never proves a shell. * fix(launch): type a plain Orca agent line as is in tcsh and nu In tcsh and nu, every Orca-built agent line ran through `/bin/sh '<script>'`, so the agent was a child of sh in sh's process group. When the prompt was pasted after the agent was ready, the paste guard read sh in front and, on a loaded Mac, reported the prompt not delivered (4/4 live runs in tcsh). Now only a line that needs it goes through `/bin/sh`: one over 512 bytes, with a control byte, or with a character those shells would not read literally (`!`, backslash, `"`, `$`, backtick). A plain line such as `claude '--dangerously-skip-permissions'` is typed as on main, so the agent is the shell's own job. * fix(agent-launch): judge the foreground-group proof by when its capture began, not how long ps took The age the host stamps on a process-group observation runs from the start of its whole-machine ps, so on a loaded Mac a capture begun after the read was asked for still read as older than 1 s and the proof was dropped. Count an observation whose capture began after the read was asked for, less the window a shared capture is reused across. * test(agent-launch): keep the crash-guard live test out of the Windows lane's gate scan The Windows-lane registration scan read the const assigned from a platform check as a Windows-only gate, though the suite runs everywhere but Windows; find zsh in a function instead, as the real-zsh typed-line test does. Under load the fresh foreground scan can fail to answer, which lets the shell's prompt settle readiness (2 of 4 paired runs). The guard still refuses that write, so assert the refused write, the property that must always hold. * perf(agent-launch): read a local pane's foreground from its own terminal, not the whole process table The foreground read that gates every launch paste ran the daemon's inspectProcess capture and then a fresh scan, each a whole-machine ps; the fresh one also waits for any capture already running before it starts its own. Measured here at load 5: 1.2 s a read (M4Air QA: 3.4-5.0 s, and worker starts 17.6-32 s against main's 9-12 s at load 25-84). On a local macOS or Linux host, take the pane's root pid from the provider's session inventory and run one ps limited to that pane's terminal. Its foreground process group decides: the launched agent or any non-shell member is the agent (a wrapper that did not exec its agent leads the group), a group of shells alone is the shell. Same pane, same verdict: 2.7 ms a read. SSH hosts keep the relay's observation and name. * test(mobile): re-measure the web app's script sweep after agent.launch's capabilities moved out of protocol-version The mobile web bundle check failed at 124 assets against a ceiling of 123. Main already sat exactly on that ceiling: its sweep table read 69 scripts at 16 routes while the tree builds 73, the whole margin of 4. This branch imports the agent.launch capabilities from their own module, so protocol-version is no longer pulled into the root layout and four other routes. That moves which routes share which modules, and the Qoder capability module, imported by protocol-version and the AI-vault resume path, no longer shares an importer set with anything, so it gets a chunk of its own: 74 scripts. The fence says to re-derive the bound rather than raise it, so the sweep is re-measured on this head (every prefix of the sorted route list). The worst route now adds 10 scripts (session), not 9, which moves the pinned shell crossing from 32 to 30 routes; main re-measured on its own lands on the same crossing. * fix(agent-launch): a worker's brief needs its agent found in front, and Grok's start answers on its composer A paired-server worker start whose agent exited at startup typed its brief into the server's shell, which ran it: the idle wait can settle on a shell back at its prompt, and the brief was written with no foreground read. Both worker-start paths now check before each brief write, as a launch prompt is checked: on a host that can find the agent in front it must be there; on one that cannot (Windows) a shell proven in front still refuses, and anything else writes as before. A Grok worker start waited ~10 s more than main: its only rest signal is its bare name, which a launch holds to quiet output, and Grok animates its logo for ten seconds after its composer glyph. A worker start for an agent whose rest signal is its bare name and whose composer draws a marker (Grok, DSH, mimo-code) now also answers on that marker, whichever comes first. * test(startup-staging): expect the CR that submits a typed launch line since #23672 |
||
|
|
ccc0bf70e4 | Pause Pullfrog reviews while the CI runner queue recovers (#25760) | ||
|
|
059e81a106 |
chore(i18n): use 智能体 for Chinese Agent copy (#25767)
Simplified Chinese rendered the Agent concept as 代理, which collides with 代理 = proxy. Standardize on 智能体 for Agent (and 子智能体 for subagent), while keeping 代理 for proxy senses: HTTP/network proxy, SSH Proxy Command, reverse proxy, and browser user agent. - Converted 131 zh catalog values (incl. 子代理 -> 子智能体); 26 proxy / user-agent values left as 代理. - locale-phrase-fixes.mjs: 客服人员/代理商/座席 -> 智能体, 代理 -> 智能体 when the English names an agent (guard excludes "user agent"); removed the old 智能体 -> 代理 rule so the pipeline no longer reverts it. - Updated value/key/search/macos-tcc overrides to 智能体; proxy keyword and proxy override entries unchanged. - Updated the two policy tests that pinned the old 代理 output. Gates: catalog verify, coverage --check, extraction, runtime-catalog, and the locale vitest suites (296 tests) pass. |