Commit Graph
9997 Commits
Author SHA1 Message Date
Neil d9cf07d3f4 Fix pinned workspace reveal expanding other hosts (#25836) 2026-10-06 02:25:36 -07:00
Neil 4de9f9f85a Move orchestration capabilities out of protocol version registry (#25839) 2026-10-06 01:27:11 -07:00
Brennan Benson ee917205bd fix(native-chat): a chat you return to drops background tasks that finished while it was hidden (#24305)
* fix(native-chat): a chat you return to drops background tasks that finished while it was hidden

A background command that finished while its chat pane was hidden stayed in
the strip as running, its timer counting for hours, and Stop failed with "The
background task wasn't stopped." The pane stops listening while hidden, so it
missed the "no tasks" update; once the idle sweep stopped the agent, the host
reopened the pane's subscription without any roster at all, and the client
reads a missing roster as "unchanged".

The roster now rides each subscriber's frames the way the slash-command list
and the message queue already do. The subscriber registry reads it from the
host's child records on a subscriber's first frame and on every snapshot and
reset, and re-sends it to each subscriber whose last copy differs when the
records change. The channel always answers: a conversation with no running
child is "none" whether or not an agent holds it. Ordinary journal batches
never read it.

This also fixes the reverse case: a snapshot or reset sent to a live pane
carried no roster, which cleared a still-running task from the strip until the
roster next changed.

The client treats a resumed subscription's first batch as stating the roster,
so a pane reconnecting to an older host that still omits it drops its stale
copy too.

Fixes #24227

* refactor(native-chat): require the background-task roster read and drop the unused close observers

The host's delivery layer now requires every hook it is built with, so
production wiring cannot leave out the roster read (an opening frame
without it reads as "no tasks" to current clients). The conversations
map's close observers lost their only caller in this PR and are removed.
The reducer test comment states the old-host omission rule precisely.

* refactor(native-chat): keep structured-agent-session-host.ts within the line limit

Main brought the host file to the 300-line lint limit, and this PR's
roster wiring adds one line. Name the lease reconciler's type by its
factory, as the neighbouring fields do.

* test(native-chat): build the roster test's Claude provider handle with claudeProviderHandle

Main made the provider handle opaque (#24991); build it with the existing
helper, as main's own tests now do.
2026-10-06 01:20:21 -07:00
Brennan Benson 670c59d23a Translate ACP traffic into shared timeline events (#25090)
* Add standalone ACP protocol client and session runtime

* Protect ACP transport teardown from late stream errors

* Retire incoming ACP request ids before publishing responses

* Narrow ACP configuration requests and transport message types

* Remove redundant ACP request handler return unions

* Keep ACP waits caller-owned and preserve protocol extensions

* Preserve open ACP decisions through prompt completion

* Move the turn message ordinals and the turn-row revision to the neutral timeline folder

Pure moves so a shared timeline assembler can use them: Codex's message
ordinal counter becomes ProviderTurnMessageOrdinals and Claude's turn-row
revision becomes the provider-neutral agent-journal turn-row revision. Only
names and import paths change.

* Admit one provider event's writes as one transition, and let rows be found again after a restart

- A sink transition is admitted whole or not at all; its steps run back to
  back at their turn in the journal's write queue, and each resolver reads
  the fold with every earlier write landed. A resolver may also say where the
  row belongs (turn scope, provider reference), and the writer always hears
  how the transition landed. A resolved lifecycle batch chooses its
  settlement mutations from the fold at execution.
- New optional row field providerItemRef: the provider's own reference for
  the item a row is, written only where the row's identity cannot spell it
  (Codex keys messages by their place in the turn and renumbers its item ids
  on resume). Set by the creating write, kept by revisions, indexed by the
  journal fold, never read by clients. A downgrade test shows an older host
  and client render such rows unchanged.
- Provider timeline identity schemes (shared legacy arm, Codex) and the join
  index that resolves a provider item to its row from memory or the fold:
  ordinals and request incarnations are read back from the rows, so a
  restart or an evicted entry finds the original row instead of placing a
  new one.

* Recover message ordinals from the journal's highest place, and forget joins read from a replaced epoch

A fresh join index continued a turn's messages at the first free place, so a journal holding only a
later ordinal (an imported or removed earlier row) had its sequence back-filled. The place is now
one past the highest ordinal any row or echoed send holds there, read through a pure scheme reader.
The join caches also drop what they read when the journal's epoch is replaced.

* Spell the subagent thread's message slot without spreading an identity union

* Add a provider timeline grammar and a shared assembler that decides at its turn in the journal

Adapters translate their provider's dialect into a small grammar (turns,
items, streamed text, requests, context facts, session end/reset); one shared
assembler turns it into the journal rows every structured lane writes.

Each event is planned as one sink transition. Which row a write lands on,
whether a replay writes anything, and every change to what the assembler
knows (its ledger) are decided by the transition's resolvers at the event's
turn in the journal's write queue, against the fold as it stands then. A
forecast (the ledger plus admitted events still queued) only answers apply()
at once. So a refused event allocates nothing, a write the journal rejects
leaves no trace in memory, and a restart or evicted cache finds the same rows
again. Text and full snapshots of one provider item share one row and one
lifecycle; reset always flushes text and settles the old session from the
journal; the open-work budget is derived from what is actually open.

Codex migration contracts compare against the existing Codex translator,
including a restart mid-stream and a repeat that outlives the join cache.

* Fix the types and the exhaustive event switch CI reported for the assembler

* Let the journal decide stream lifetimes, request reuse, named sends and background work

A third review found two blockers with the earlier rounds' cause, a remembered interpretation
trusted after the journal moved on:

- A reused request id was judged by its earlier prompt's settled turn before asking which turn the
  new one lands in, so a real approval in a later turn was dropped. The target turn now decides:
  the old turn again is a replay; a different live turn opens the next prompt beside it.
- A text stream checked its row's turn only on its first write, and a turn's end released streams
  by the turn planning expected. Every write now checks the row, a turn's end stops the streams
  whose rows are in it, and turn status reads the journal first, so another writer's Stop wins.

Also: a message boundary drawn by an event the journal held as a replay no longer splits an
anonymous message; a send naming a turn not yet open waits for that turn; the budget charges a
stream's thread and turn strings and the turn caches are byte-bounded; the open turn ends when the
journal shows it settled; a turn's opener is read from the journal's row.

Background work is now Orca's existing background-task row instead of a tool call flagged
`outlivesTurn` (a flag remembered only in memory, so a restart failed the task). A turn's end
never settles that row, so it survives restarts; session end leaves one in flight unverifiable.
Three tests that opened a background tool call with `outlivesTurn` now open a background-task row
and keep their original expectations about which turn the row stays in.

* Type the unbound assembler helper's drain as the void it reports

* Bound the rows kept for a continued anonymous message and the stopped streams

A row kept for the anonymous stream that may continue it, and the marker that a stopped stream's
queued writes write nothing, lived in the live-stream map and were never removed when no stream
followed. They now live in their own bounded maps, so the live map holds open streams only.

* Translate ACP session traffic into shared timeline events

* Preserve fixture answer linkage and handle all typed ACP updates

* Keep replay chunks and long-turn tool snapshots bounded

* Forward ACP joins to the journal-derived timeline assembler

* Use exact ACP prompt identity and recover incomplete timeline replay

* Detect short and Unicode home paths in encoded ACP fixtures

* Use generic privacy patterns for ACP fixtures

* Translate Grok background work into durable task rows

* Recover background task origin from its persisted launch tool

* Restate background task fallback from its terminal evidence

* Require session identity on Grok task notifications

* Honor explicit Grok task completion without an exit code

* Match the background reset test to the final timeline grammar

* Give task completion fixtures a specific evidence type

* Settle ACP background tasks from live and replayed evidence

* Honor background task outcomes carried by replayed tool results

* Keep the journal store under its line limit after the main merge

* Drop the provider item reference, join index and identity schemes from the transition PR

Nothing in production reaches the state they defended (an assembler that lost its memory while
its child keeps streaming the same turn), and the stored Codex id was positional. The legacy
identity scheme moves to the assembler PR with its first caller; the Codex scheme and any
persisted reference wait for Codex to move onto the assembler. The Codex ordinal counter goes
back to codex/, since no neutral code imports it.

* Write a resolved settlement in one transaction through enqueueRows

A settlement too large for one row now commits all its rows or none, through the journal's
existing all-or-nothing write, instead of a row-by-row writer. Every row is built before any
commits, so a settlement naming one item twice is refused before anything is written.

* Drop the transition's landing report; keep the turn-row write fire-and-forget

Nothing reads which steps of a transition wrote. A failed step fails the sink, leaving the
steps before it written; the header says so, and tests cover it plus a settlement whose second
row fails inside the transaction. writeAgentJournalTurnRow returns nothing again, as on main.

* Generate open ACP enums and check the generated schema offline

A newer or vendor enum value (tool kind, tool status, option kind, stop
reason) no longer fails the whole message: generated enums accept the known
literals plus any other string, typed so callers can still narrow on the
known ones. The generated header now records the pinned input digests, the
generator digest and a body hash, so `verify:acp-protocol` catches a stale or
hand-edited file without network access; it runs in lint and the PR workflow.

* Land the ACP runtime contract the agent adapters use

- Deliver notifications other than session/update through
  onExtensionNotification, in arrival order with session updates.
- Accept _meta on prompt, setMode, setModel, setConfigOption and cancel.
- cancel() always sends session/cancel once the session runs, since the
  agent can be in a turn it began itself; only a successful send is shared,
  so a failed write is retried.
- Cancel aborts each open agent request's signal and lets its handler send
  its own answer; -32800 only when the handler rejects.
- Permission requests validate only the session, tool call id and options;
  unreadable fields are dropped with a diagnostic, and any answer Orca
  cannot send is `cancelled` instead of a JSON-RPC error. Agent-started
  turns may ask; whether to show it is the caller's decision.
- AcpAgentError marks the agent's own errors; AcpInvalidResponseError keeps
  the raw answer and validation issues for answers Orca could not read.
- Lines over the size limit are classified by prefix (shared with the Codex
  reader): the owed request fails, an oversized agent request is answered
  with an error, and an unattributable response closes the connection.

* Run a transition's steps as a prefix; drop the paced flag and resolved options

A failed step no longer lets the steps after it write: each step checks, at its own turn in the
journal's queue, whether the write handed over just ahead of it completed, using the queue's count
of completed write bodies (a promise would report the failure only after the next step ran). The
sink fails only once every step has had its turn. The item step's `paced` size bypass and the
resolver's replacement `options` are removed; nothing planned uses them.

* Rebuild the timeline assembler on one admission-order state

The assembler kept a second copy of its state (a forecast beside a ledger), hydrated
the open turn from the journal, and recognised replays, all to recover from losing its
memory while the provider child kept streaming. That never happens: one assembler lives
exactly as long as one child, and a new child is a new assembler in a new generation
whose events land after the dead-generation sweep.

- One state, changed only when the sink admits an event (minted keys included), so a
  refused event takes nothing.
- Every journal-dependent choice is made when the write runs, by keyed reads: a turn row
  is written only where none is, a stream checks its row's turn on every write, a running
  snapshot never lands in a settled turn or relights a settled tool, a request takes the
  first incarnation the journal holds no row for.
- Rows are found by spelling their ids (provider-timeline-rows.ts); no join cache.
- The identity scheme (legacy arm only) lives here with its first caller; requests are
  spelled in their acquisition generation, since JSON-RPC ids restart per process.
- Saved history goes in as `input.history` plus ordinary events with the provider's ids,
  into an empty journal; `session.reset` and every replay rule are gone.
- One terminal-body function (`terminalAgentJournalBody`) is shared with the
  dead-generation settlement.
- The test rig's restart now sweeps and starts a new generation, as production does.

* Answer every agent request after an ACP cancel

A cancel that lands before a permission handler starts now still runs the
permission path, so the agent gets the `cancelled` outcome rather than a
request-cancelled error. A handler that ignores the abort no longer leaves
the agent waiting: once the abort has run through, any request still
unanswered gets request-cancelled. Handlers that answer on abort keep their
own reply.

Also renames a lint-rejected helper parameter, replaces a Reflect.apply in a
test, and stops the permission diagnostic from firing with an empty list.

* Cover new running work in a turn the sweep ended

* Run a transition's steps in one queued write that loops over them

The steps of one event now share one turn in the journal's write queue: a
loop writes each in its own transaction through the row writer's
synchronous writeRows (split out of enqueueRows) and stops at the first
throw. Prefix semantics and "nothing lands between the steps" now hold by
construction, so the completed-write counter on the queue, the step gate
and the allSettled barrier are gone; the queue is back to main's bytes.

* Let each ACP request handler own its answer after a cancel

Removes the next-event-loop-turn fallback that answered request-cancelled
for any handler still silent after a cancel. It raced answers that were
still being saved (an approval mid-journal-write reached the agent as an
error) and made the outcome depend on event-loop timing. The handler that
owns an agent request now always sends its answer, or throws for
request-cancelled; a request it never answers ends when the connection
closes. A permission whose handler had not started still answers
`cancelled`.

* End a turn another writer settled the way the provider's end does

A person's Stop settled the open turn's row without a word to the assembler.
The assembler then forgot the turn: its running tools and pending prompts were
never settled, the provider's own end and withdrawal were dropped, and the
turn's text streams stayed counted against the open budget for the life of the
process. Text the provider kept streaming afterwards could land as a message
outside the stopped turn.

- The open turn the journal shows settled ends first, as one transition, through
  the same settlement the provider's turn.end plans; its streams stop and their
  keys drop later text until that turn's end or the next turn opens.
- turn.end and request.withdrawn are admitted for a turn or request the journal
  holds; their settlement writes nothing for rows already settled.
- The budget's re-check frees streams whose turn settled.
- A settled tool keeps its terminal body against any differing write.
- Session-end settlement of lost background work uses the journal's own
  lostLiveWorkJournalBody instead of a copy.
- The rig's window elapses before every read, and restart swaps and disposes
  the old assembler.

* Adopt ACP load history only into an empty journal, chosen by the translator's creator

The translator now takes `adopt` from whoever creates it instead of guessing from the
journal (which reads empty before the sink binds). Without adoption, history a provider
replays during session/load is dropped except its context usage. With adoption, history
becomes ordinary events with provider or position ids and the user's saved messages become
`input.history`, so re-running an interrupted adoption lands the same rows.

The translator no longer reads the journal: it lives exactly as long as its provider child.
Its tool snapshots are bounded by the assembler's open budget, and a small record of each
tool's turn keeps a late background-task notice beside the tool that started it.

* Leave a stopped turn's running tools to the agent's own end

When another writer settles the open turn (a person's Stop), the assembler
now only stops that turn's text and cancels its pending prompts. Running tool
calls stay the agent's: a progress update or completion it reports after the
Stop lands as reported, and whatever is still running settles at the agent's
turn end for that turn, the next turn's open, or the session's end.

An agent's end for an earlier turn while a newer one is open no longer clears
the open turn's activity line or ends its anonymous reply. An unnamed end right
after a Stop ends the stopped turn instead of being dropped. The test rig's
restart no longer writes the dead assembler's window text, matching dispose.

* Pin that a stopped turn's running tools hold budget until the agent's end

* Type the stopped turn's tool progress update as a tool body

* List every event the assembler hands to the decision step

The type-aware lint requires an exhaustive switch with no default case.
Also retitle a Stop test to say what it asserts.

* Say why a Grok turn failed, and keep task rows in Grok's own words

A failed Grok turn ended with no reason on screen: the translator dropped
every copy of Grok's message. The failed turn now gets one status row in
Orca's existing "provider did not accept this message" words with Grok's
reason, read from whichever copy arrives first (the given-up retry, the
turn's end, the prompt's completion notice, or the prompt's error answer);
later copies only fill a reason the row still lacks.

A running background command no longer reads "Background task <id>
started": a task's summary is mapped only once it has settled. A monitor
stays a monitor when the agent reads its output: a frame that names no
kind keeps the known one, and a "[monitor" command is a monitor.

A prompt's turn is marked started, so a late frame for an ended prompt
neither reopens it nor becomes the active turn. A tool's turn is held in
one place at a time.

* Read a monitor from Grok's exact output prefix

* Word a failed Grok turn in Grok's own text, not as a refused message

A turn that started and then failed was told "The provider did not accept
this message", Orca's sentence for a message refused before its turn. The
row now reads as a Codex turn-ending error does: an error status row with the
provider's own words. With no words, the dialect names the failure ("Grok
ended this turn with an error." / "Grok usage limit reached."), else the
agent's display name does.

* Register the ACP schema verify step in the PR preflight phase test

* Read ACP permissions, session events and prompt errors through the protocol client's own types

The translator now reads a permission request with the client's lenient reader, a session update
with its session-event reader, and takes only the agent's own error answer as a failed prompt's
reason, so an Orca-side error never reads as the provider's words. Tests cover protocol values
newer than this build.

* refactor(native-chat): drop saved-history adoption from the timeline assembler

The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.

* refactor(acp): drop session/load history adoption from the translator

The common pattern discards the history an agent replays during session/load,
keeping only what it says about the context window. Remove the adoption path
(acp-history-adoption.ts, the adopt option, and the historical background-task
liveness rewrite it fed) so load replay is always dropped except usage.

* refactor(native-chat): a pending input is only Orca's send now

Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.

* test(acp): keep the task-result status table on live frames

Review follow-up to the adoption removal: the result-status mapping was only
tested through adopted history, so run the same table on live frames, and
cover an unmarked task notice during a load being dropped.

* Use current provider handles in transition tests

* Use current provider handles in timeline fixtures

* Require the ACP directory in the runtime import check
2026-10-06 01:07:22 -07:00
Kelvin AmoabaandNeil 634c43787c Fix premature Claude automation completion and inherited CI failures (#24878)
* fix(agent-status): keep a Claude pane working until owed task wake-ups arrive

Claude wakes the main agent for each background task that ends, after the
task stops running. "Nothing running" was read as done, so automations
closed the terminal before the final turn.

Fixes #23942

* fix(agent-status): keep owed task wake-ups through failed turns, the cap and nested exits

A failed turn's task list now marks a vanished shell owed like a normal turn end does. Past
the cap only an already-announced sub-agent is forgotten. A process exit clears what is owed
only where the server admits it from the pane's owner.

* refactor(agent-status): move pane-scoped cache entry helpers out of listener-state

listener-state.ts passed the 300-line limit once main's and this branch's additions met.

* fix(agent-status): verify Claude task wake-up completion on its execution host

* fix(agent-status): canonicalize Claude background task identifiers

* Preserve out-of-order Claude task notification evidence

* Verify retained released parser in cross-version checkout test

* Reuse watcher directory-cache enumeration without losing fresh keys

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-10-06 01:02:18 -07:00
Brennan Benson 0bcd49c04c Let ACP connections own their agent processes (#25810)
* Let ACP connections own their supervised agent process

* Preserve ACP cleanup evidence and isolate exit observers

* Expose ACP cleanup observations and type the permission fixture
2026-10-06 01:01:13 -07:00
Neilandggbdpq a2f197fde7 Refresh visible reviews automatically and stop settled merged polling (#25788)
Refresh only actual on-screen review cards and the selected visible review panel, with a single metadata owner per execution host. Open reviews refresh every 60 seconds when selected and 120 seconds otherwise. Settled merged reviews stop automatic refresh; pending merged checks continue, hidden rows stop, and work changes or stale re-exposure discover fresh state.

Remove the polling setting. Preserve provider/SSH/runtime ownership, mixed-version cache behavior, failure backoff, bounded foreground admission, request coalescing, and mutation invalidation. Cover older persisted caches and unknown HEADs.

Validated with 1,628 focused tests, typechecking, lint/code-quality gates, hidden Electron viewport checks, and parallel Codex/Claude Opus 5.5 adversarial reviews. Measurements and policy details are in PR #25788.

Fixes #25746

Co-authored-by: ggbdpq <ggbdpq@gmail.com>
2026-10-06 00:53:49 -07:00
Brennan Benson 9905765e3b feat(native-chat): open structured chat's wire and stored records to registered agents, behind a negotiated capability (#25159)
* refactor(native-chat): keep the provider resume handle opaque to shared code

Shared structured-chat code parsed each provider's resume handle: Claude's
session id and branch leaf, Codex's thread id, through a 'claude' | 'codex'
union every new agent had to widen. The in-memory handle is now
{ transport, agent, nativeId, providerData? }: shared readers use nativeId,
lease and handle-chain checks compare transport and agent, and only the
Claude adapter reads its leaf (providerData).

Stored and wire forms are unchanged for Claude and Codex. One encoding
module writes their typed shapes and decodes both those and the neutral
shape a new transport uses, which an older build refuses as unreadable
rather than reading as Codex. Key and root strings, which fork seeds,
superseded creations and resume offers persist, stay byte-identical.

The journal's own handle type becomes the journal-row and attach-wire
encoding of the same handle, and the journal identity carries the
neutral handle (null before the provider proves one).

No user-visible change.

* fix(native-chat): derive journal-row provider handles from the journal identity

The journal row converter now takes the identity every caller already holds,
so a row's handle has one obvious constructor. Tests that wrote the in-memory
handle straight into journal rows now build it through that converter, and the
processless Claude fixture names a not-yet-proved handle as null.

* fix(native-chat): refuse a stored provider handle written in both forms

A typed Claude or Codex handle that also carries the neutral form's
transport, agent, native id or provider data named two identities; it
was read as Claude or Codex and the next write dropped the other one.
Such a row now stays unreadable and is set aside untouched.

* refactor(native-chat): route structured agents through registered definitions

The structured-chat shared layer named Claude and Codex in a closed router
type, in option reads, in frame classification, and in renderer label and
validation branches, so adding an agent meant editing shared types. Each
agent now has a definition (capabilities, resting option rules) owned by its
own module; the router is a registry of definition + adapter pairs built at
runtime setup, and shared code reads the definition instead of the name.

No behavior change for Claude or Codex; no wire or stored shape change.

* refactor(native-chat): name the structured agent list once in host types

The host types, conversation commands, provider-session ownership and the
runtime's create path spelled the structured agent list inline; they now use
the one shared type the record and wire already validate against.

* fix(native-chat): narrow the record before reading its agent's option rules

* refactor(native-chat): make the router's registrations the only agent definition lookup

The at-rest option reads looked definitions up in a second, separately populated map, so a
registered definition could route and declare capabilities while the static map decided which
options a chat at rest accepted. The router now answers `definition(agent)` from its
registrations, the at-rest readers take it through the host's adapter, and the built-in
definitions are listed only where the runtime composes the router.

* feat(native-chat): open the structured-chat wire and stored records to registered agents

A host's structured agents are the ones its runtime registered. Records, RPC
params, persisted tabs and the model catalog accept any registered agent instead
of naming Claude and Codex; each agent's definition declares the transport its
handles live in and the variable its account home pins. A new runtime
capability, agent-session.structured.registered-agents.v1, advertises that a
host accepts and lists its agents (agentSession.agents, with each agent's
capability record), and the host withholds any other agent's tabs and restart
offers from clients that do not advertise it.

* test(native-chat): cover registered agents on the wire, in storage and across versions

* refactor(native-chat): let the record store decide which agents' tabs exist

* test(native-chat): declare the pilot test agent's storage

* test(native-chat): read the old build's saved tabs through a parsed shape

* fix(native-chat): act on restart offers only for agents the calling client can show

A paired client too old to show an agent's chat was listed only the offers it could show, but
dismissing or continuing all reached every offer on the host, and a named continuation answered
with the host's whole remaining inventory. The client's audience now goes to the host with every
restart operation: only offers it sees are reserved, dismissed or returned. Without an audience
(this host's own process, or a client that shows every agent) nothing changes.

* test(native-chat): read an agent-registering baseline's storage on its own terms

The registered-agents downgrade test assumed its baseline release predates registered agents: it
expected the saved-tab parser to erase an unknown agent and called the record reader without the
agents list. Once a release with this change becomes the baseline, both break. The expectations now
follow what the baseline host advertises, and an agent-registering baseline is handed its own
Claude and Codex storage.

* refactor(native-chat): derive record-store admission from the runtime's agent registrations

Which agents a stored record may name and which agents the router drives came from two lists in
the runtime, so a newly registered agent could be routed while its records were set aside. One
list of registrations now holds each agent's definition and the factory for its adapter: the
store's admitted agents are derived from it before the store opens, and the adapters are built
from it once it has.

* fix(native-chat): hand the exit drain a promise for every registered agent

* fix(native-chat): let the adoption conflict check read any agent's ownership

Ownership rows name any registered agent since the stored records opened to them; the adoption
check compares by agent, so it takes the same open id. Only Claude and Codex still adopt.

* test(native-chat): use opaque handle in queued rejection fixture

* test(native-chat): share one Codex journal identity in the integration suite

Main grew the suite to the 800-line limit; the opaque-handle import pushed it
over. The two tests built the same identity inline.

* refactor(agent-session): name the handle's adapter state resumeCursor

Rename the neutral provider handle's providerData to resumeCursor before any
row persists the neutral form: it is an adapter-owned resume position (Claude's
transcript leaf), never identity. Claude/Codex stored and wire bytes are
unchanged; their typed shapes never carried the field.

State the stored-form contract (a handle's field set is closed; later per-link
data goes on the chain link, which every build preserves) and pin it with a
record round-trip test. Document that transport records the id space the
native id was minted in, which can differ from the agent's current transport.

* refactor(agent-session): one required agent registry; declarations admit what they claim

A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.

/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.

Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).

* refactor(agent-session): the router applies the declared rewind itself

The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.

* test(agent-session): register the agents the merged-in tests now need

The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.

* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record

* fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop

The desktop renderer calls the restart methods as a paired runtime client
without the registered-agents capability, so it was handed a Claude/Codex
audience and its "dismiss all" took the scoped path: no dismissal fence and
no clearing of unwritten teardown witnesses, so a late teardown write could
bring a dismissed offer back.

The audience is now derived against this host's registry: a client that can
show every agent the host registered gets none, and dismisses exactly as
before (one fence for every offer). A client that truly cannot show some
agent still gets a final dismissal for what it sees: each cleared session is
fenced on its own (bounded, superseded by a dismiss-all fence) and only the
witnesses of agents it sees are dropped.

One predicate now answers which agents a client renders, for both tab
projection and restart offers; the Claude capability rule is folded in.

* fix(agent-session): a changed agent definition never hides that agent's chats

A record was readable only if every handle used the transport its agent's
definition declares today, and its account variable was declared by some
registered agent. A later build that drives the same agent over another
protocol, or under a renamed variable, would have set every existing chat of
that agent aside: no tab, no history, no error.

Readability now asks only that the record agree with itself: its agent is
registered, every handle shares one namespace owned by that agent, and the
account variable is a well-formed name. Whether this build can drive it (the
chain's transport is the one the agent speaks, and the account variable is
that agent's own, since it becomes the child's environment) is decided in
performAttach, the one admission every start of every agent passes through,
and refused there as hostUnsupported. A record pinning another agent's
variable is refused the same way rather than passing because some agent
declares it.

* refactor(agent-session): each agent's registration says where it runs and which account it pins

createSupport, create and the model catalog's account read still chose a
location rule and an account resolver by name (Claude, Codex, else none),
so a new agent would have needed a third branch beside the registration
list that already decides storage, routing and publication.

Each runtime registration now carries supportsLocation(location) and
resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }),
and Claude's and Codex's rules move into their entries unchanged: a read
still syncs no home and starts no bridge, a Codex launch still trusts the
workspace first, and WSL locations resolve as before. The list exists
before the host is built, so none of these reads installs the host, and an
agent the list does not hold is answered no without opening the journal.

* test(agent-session): drop a duplicate registry key; keep fences out of the capsule state

A spread input already carries the registry, so the explicit key was
overwritten (TS2783). The session-fence helper now returns only the fence,
not the capsule state it was given.

* fix(agent-session): a scoped dismiss-all persists no per-session fence

The per-session dismissal fence added a forever-persisted capsule field
with an arbitrary cap that served no caller: only the local desktop calls
the restart methods, and on a host whose agents it all shows it gets no
audience and takes the unscoped, fenced dismissal. A scoped dismiss-all
(a caller that cannot show some registered agent) keeps that agent's
offers, clears only its audience's unwritten witnesses, and writes no
fence; a late write from this process is serialized behind it. The field
never left this unreleased branch.

* fix(agent-session): refuse an attach whose agent is not the session's own

The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal.

* fix(agent-session): offer to start a chat only when the start would accept it

The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired.

* fix(agent-session): the model-catalog probe uses a record's account only if this build drives it

A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record.

* refactor(agent-session): the record store admits agent ids; comments say where transport is checked

The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept.

* docs(agent-session): the record store admits the registered agents' ids

* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer

Uses an audience production sends (one that cannot show every agent), per review.

* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge

* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.

* test(ratchet): require src/main/provider-process now that it has landed

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 and this branch both added the import at different lines; the merge kept both.

* Keep saved chats readable without provider registration

* Keep stored-record compatibility checks independent of registration

* Keep saved providers in restart client audiences

* test(native-chat): type reveal fixtures without assertions

* test(wire): expose known agents in structured host fixture

* Supply startability dependency in the new Codex catalog fixture
2026-10-06 00:41:14 -07:00
Brennan Benson c4ea14cd9f Show a tool call you stopped as interrupted, not failed (#25181)
* Move the turn message ordinals and the turn-row revision to the neutral timeline folder

Pure moves so a shared timeline assembler can use them: Codex's message
ordinal counter becomes ProviderTurnMessageOrdinals and Claude's turn-row
revision becomes the provider-neutral agent-journal turn-row revision. Only
names and import paths change.

* Admit one provider event's writes as one transition, and let rows be found again after a restart

- A sink transition is admitted whole or not at all; its steps run back to
  back at their turn in the journal's write queue, and each resolver reads
  the fold with every earlier write landed. A resolver may also say where the
  row belongs (turn scope, provider reference), and the writer always hears
  how the transition landed. A resolved lifecycle batch chooses its
  settlement mutations from the fold at execution.
- New optional row field providerItemRef: the provider's own reference for
  the item a row is, written only where the row's identity cannot spell it
  (Codex keys messages by their place in the turn and renumbers its item ids
  on resume). Set by the creating write, kept by revisions, indexed by the
  journal fold, never read by clients. A downgrade test shows an older host
  and client render such rows unchanged.
- Provider timeline identity schemes (shared legacy arm, Codex) and the join
  index that resolves a provider item to its row from memory or the fold:
  ordinals and request incarnations are read back from the rows, so a
  restart or an evicted entry finds the original row instead of placing a
  new one.

* Recover message ordinals from the journal's highest place, and forget joins read from a replaced epoch

A fresh join index continued a turn's messages at the first free place, so a journal holding only a
later ordinal (an imported or removed earlier row) had its sequence back-filled. The place is now
one past the highest ordinal any row or echoed send holds there, read through a pure scheme reader.
The join caches also drop what they read when the journal's epoch is replaced.

* Spell the subagent thread's message slot without spreading an identity union

* Add a provider timeline grammar and a shared assembler that decides at its turn in the journal

Adapters translate their provider's dialect into a small grammar (turns,
items, streamed text, requests, context facts, session end/reset); one shared
assembler turns it into the journal rows every structured lane writes.

Each event is planned as one sink transition. Which row a write lands on,
whether a replay writes anything, and every change to what the assembler
knows (its ledger) are decided by the transition's resolvers at the event's
turn in the journal's write queue, against the fold as it stands then. A
forecast (the ledger plus admitted events still queued) only answers apply()
at once. So a refused event allocates nothing, a write the journal rejects
leaves no trace in memory, and a restart or evicted cache finds the same rows
again. Text and full snapshots of one provider item share one row and one
lifecycle; reset always flushes text and settles the old session from the
journal; the open-work budget is derived from what is actually open.

Codex migration contracts compare against the existing Codex translator,
including a restart mid-stream and a repeat that outlives the join cache.

* Fix the types and the exhaustive event switch CI reported for the assembler

* Let the journal decide stream lifetimes, request reuse, named sends and background work

A third review found two blockers with the earlier rounds' cause, a remembered interpretation
trusted after the journal moved on:

- A reused request id was judged by its earlier prompt's settled turn before asking which turn the
  new one lands in, so a real approval in a later turn was dropped. The target turn now decides:
  the old turn again is a replay; a different live turn opens the next prompt beside it.
- A text stream checked its row's turn only on its first write, and a turn's end released streams
  by the turn planning expected. Every write now checks the row, a turn's end stops the streams
  whose rows are in it, and turn status reads the journal first, so another writer's Stop wins.

Also: a message boundary drawn by an event the journal held as a replay no longer splits an
anonymous message; a send naming a turn not yet open waits for that turn; the budget charges a
stream's thread and turn strings and the turn caches are byte-bounded; the open turn ends when the
journal shows it settled; a turn's opener is read from the journal's row.

Background work is now Orca's existing background-task row instead of a tool call flagged
`outlivesTurn` (a flag remembered only in memory, so a restart failed the task). A turn's end
never settles that row, so it survives restarts; session end leaves one in flight unverifiable.
Three tests that opened a background tool call with `outlivesTurn` now open a background-task row
and keep their original expectations about which turn the row stays in.

* Type the unbound assembler helper's drain as the void it reports

* Bound the rows kept for a continued anonymous message and the stopped streams

A row kept for the anonymous stream that may continue it, and the marker that a stopped stream's
queued writes write nothing, lived in the live-stream map and were never removed when no stream
followed. They now live in their own bounded maps, so the live map holds open streams only.

* Read a tool call a stop or the session's end cut short as interrupted, not failed

A call still running when its turn was interrupted, when the provider child
was seen to exit, or that the provider cancelled now carries `endedAs:
'interrupted'` beside `state: 'failed'`. Every build that predates the field
keeps reading the call as the failure it always showed; this build reads both
through one shared reader and counts the call apart from failures in the run
header, without the error tint on its partial output.

The shared assembler, the restart settlement and the Codex lane settle a
running call through one rule: only a proven interruption cuts it short, so a
turn the provider completed around an unclosed call, or work the host lost
track of, still reads failed.

* Assert a completed Codex turn leaves its unclosed call plainly failed

* Assert an unverifiable restart settlement leaves its running call plainly failed

* Let a later death proof correct the calls an unverifiable settle closed

A settle with no proof of the child's death closed its running calls as plain failed, so a
proof written afterwards corrected the turn to interrupted but could not find the calls. The
call now keeps endedAs: 'unverifiable' beside its failed state, and the stale settle revises
exactly those calls whose owner the proof names. Older builds still read them as failed.

* Render a cut-short call's output on mobile, with and without proof

* Find the mobile result box by walking to its View, and type the unverified-ending block

* Walk to the mobile result box without a type assertion

* Keep the journal store under its line limit after the main merge

* Drop the provider item reference, join index and identity schemes from the transition PR

Nothing in production reaches the state they defended (an assembler that lost its memory while
its child keeps streaming the same turn), and the stored Codex id was positional. The legacy
identity scheme moves to the assembler PR with its first caller; the Codex scheme and any
persisted reference wait for Codex to move onto the assembler. The Codex ordinal counter goes
back to codex/, since no neutral code imports it.

* Write a resolved settlement in one transaction through enqueueRows

A settlement too large for one row now commits all its rows or none, through the journal's
existing all-or-nothing write, instead of a row-by-row writer. Every row is built before any
commits, so a settlement naming one item twice is refused before anything is written.

* Drop the transition's landing report; keep the turn-row write fire-and-forget

Nothing reads which steps of a transition wrote. A failed step fails the sink, leaving the
steps before it written; the header says so, and tests cover it plus a settlement whose second
row fails inside the transaction. writeAgentJournalTurnRow returns nothing again, as on main.

* Run a transition's steps as a prefix; drop the paced flag and resolved options

A failed step no longer lets the steps after it write: each step checks, at its own turn in the
journal's queue, whether the write handed over just ahead of it completed, using the queue's count
of completed write bodies (a promise would report the failure only after the next step ran). The
sink fails only once every step has had its turn. The item step's `paced` size bypass and the
resolver's replacement `options` are removed; nothing planned uses them.

* Rebuild the timeline assembler on one admission-order state

The assembler kept a second copy of its state (a forecast beside a ledger), hydrated
the open turn from the journal, and recognised replays, all to recover from losing its
memory while the provider child kept streaming. That never happens: one assembler lives
exactly as long as one child, and a new child is a new assembler in a new generation
whose events land after the dead-generation sweep.

- One state, changed only when the sink admits an event (minted keys included), so a
  refused event takes nothing.
- Every journal-dependent choice is made when the write runs, by keyed reads: a turn row
  is written only where none is, a stream checks its row's turn on every write, a running
  snapshot never lands in a settled turn or relights a settled tool, a request takes the
  first incarnation the journal holds no row for.
- Rows are found by spelling their ids (provider-timeline-rows.ts); no join cache.
- The identity scheme (legacy arm only) lives here with its first caller; requests are
  spelled in their acquisition generation, since JSON-RPC ids restart per process.
- Saved history goes in as `input.history` plus ordinary events with the provider's ids,
  into an empty journal; `session.reset` and every replay rule are gone.
- One terminal-body function (`terminalAgentJournalBody`) is shared with the
  dead-generation settlement.
- The test rig's restart now sweeps and starts a new generation, as production does.

* Cover new running work in a turn the sweep ended

* Run a transition's steps in one queued write that loops over them

The steps of one event now share one turn in the journal's write queue: a
loop writes each in its own transaction through the row writer's
synchronous writeRows (split out of enqueueRows) and stops at the first
throw. Prefix semantics and "nothing lands between the steps" now hold by
construction, so the completed-write counter on the queue, the step gate
and the allSettled barrier are gone; the queue is back to main's bytes.

* End a turn another writer settled the way the provider's end does

A person's Stop settled the open turn's row without a word to the assembler.
The assembler then forgot the turn: its running tools and pending prompts were
never settled, the provider's own end and withdrawal were dropped, and the
turn's text streams stayed counted against the open budget for the life of the
process. Text the provider kept streaming afterwards could land as a message
outside the stopped turn.

- The open turn the journal shows settled ends first, as one transition, through
  the same settlement the provider's turn.end plans; its streams stop and their
  keys drop later text until that turn's end or the next turn opens.
- turn.end and request.withdrawn are admitted for a turn or request the journal
  holds; their settlement writes nothing for rows already settled.
- The budget's re-check frees streams whose turn settled.
- A settled tool keeps its terminal body against any differing write.
- Session-end settlement of lost background work uses the journal's own
  lostLiveWorkJournalBody instead of a copy.
- The rig's window elapses before every read, and restart swaps and disposes
  the old assembler.

* Leave a stopped turn's running tools to the agent's own end

When another writer settles the open turn (a person's Stop), the assembler
now only stops that turn's text and cancels its pending prompts. Running tool
calls stay the agent's: a progress update or completion it reports after the
Stop lands as reported, and whatever is still running settles at the agent's
turn end for that turn, the next turn's open, or the session's end.

An agent's end for an earlier turn while a newer one is open no longer clears
the open turn's activity line or ends its anonymous reply. An unnamed end right
after a Stop ends the stopped turn instead of being dropped. The test rig's
restart no longer writes the dead assembler's window text, matching dispose.

* Pin that a stopped turn's running tools hold budget until the agent's end

* Type the stopped turn's tool progress update as a tool body

* List every event the assembler hands to the decision step

The type-aware lint requires an exhaustive switch with no default case.
Also retitle a Stop test to say what it asserts.

* End a running call as its turn's journal row ends

A call still running when its turn ends takes the state of that turn's
row: a row another writer settled first (a person's Stop) stands, so its
calls read interrupted whatever the provider's later end reports. The
no-ending path that settled calls from the Stop row is gone, since a Stop
now leaves running calls to the provider. Adds the two Spanish strings.

* Settle a stopped turn's running call as its turn row ended after a restart too

The restart sweep ended every running call by the death evidence alone, so after
a person's Stop with no proof the child died the call read failed under a turn
that read interrupted. The sweep and the live dead-generation settlement now ask
the same rule the assembler does: a call in a turn already settled ends as that
row ended; only a turn still running leaves its calls to the evidence.

* Keep the dead-generation settlement under the line cap

* refactor(native-chat): drop saved-history adoption from the timeline assembler

The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.

* refactor(native-chat): a pending input is only Orca's send now

Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.

* test(ratchet): require src/main/provider-process now that it has landed

* test(native-chat): build this stack's journal identities with main's opaque provider handle

Main's #24991 replaced the {kind, ...} handle with {transport, agent, nativeId}; three test
files from this stack still wrote the old shape. Same lines the downstream ACP branch uses.
2026-10-06 00:40:47 -07:00
41f293c264 fix(notifications): replay completion sounds reliably (#15940)
* fix(notifications): replay completion sounds reliably

The in-flight `isNotificationSoundPlaying` gate dropped every notification
sound that arrived while the previous one was still ringing, so bursts played
once. Drop the gate and restart the cached Audio instead, and reuse an entry a
concurrent call already cached for the same path so parallel first playbacks
share one Audio rather than revoking each other's blob URL.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016S8SZpkXGcnU5UWwA8UiF2

* test(notifications): verify real sound replay and remove unused playback listeners

---------

Co-authored-by: Laku <laku@LakudeMacBook-Pro.local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Neil <neil@stably.ai>
2026-10-06 00:39:33 -07:00
Neil 6a9ba9d733 Keep remote worktree creation from navigating paired clients (#25793)
* fix(workspaces): keep remote creation from navigating paired viewers

* fix(workspaces): revoke pending navigation after connection replacement

* fix(workspaces): provision when the host renderer is unavailable

* test: align navigation and cache-scan expectations with current behavior

* fix(workspaces): provision broadcast creates without a host renderer
2026-10-06 00:38:53 -07:00
Neil 8507111d87 Document Apple Git ref recovery retirement condition
Name the affected Apple Git-154 (2.39.5) build and document when corruption recovery can be retired while retaining current-Git disambiguation. Comments only; lint, formatting, and review pass.
2026-10-06 00:29:44 -07:00
Jinwoo Hong 29e669680d fix(codex): never write a Codex config.toml that Codex can't load, and approve both symlink spellings (#25741)
* fix(codex): never write a hook approval Codex cannot load, and approve both symlink spellings

- Refuse any hooks.state write into a Codex config.toml (upsert, move,
  remove, mirrored enabled state, SSH installer) that would turn a loadable
  file into one Codex cannot load; write nothing and surface the reason.
- Read approvals written as dotted keys or inline tables.
- Approve and remove Orca's ~/.codex hook under both the spelled and the
  resolved key when ~/.codex or HOME is a symlink; move user approvals
  under both keys.
- Stale runtime trust cleanup no longer keeps an unexpected key whose
  conflicting duplicate tables read as no hash.

* refactor(codex): pass every hooks.json spelling as one sourcePaths list

* fix(codex): sweep retired-hook approvals under every key spelling; read literal-string trusted_hash
2026-10-06 02:57:25 -04:00
Neilandkino 1de3aa405f Fix truncated Unicode branch names in base ref search
Preserve complete branch selectors when Git truncates or disambiguates short names, and keep slash-named local branches separate from remote-tracking refs through worktree creation and reuse.

Keep native/SSH namespace boundaries and mixed-version client capability gates intact. Accept Git-valid dotted components and cover collisions with configured and orphan tracking refs.

Adapted from @nishino-tsukasa's #19540. Independently reviewed in seven adversarial rounds; 235 focused tests and 77 affected-Apple-Git tests pass, along with typecheck, quality checks, and the desktop build.

Correct the inherited file-explorer cache-scan expectation while retaining its exact constant-work bound.

Correct the inherited release-parser compatibility assertion to match the pinned test dependency; all 121 cross-version tests pass locally.

Fixes #19515

Co-authored-by: kino <nishinotsukasavirgo@gmail.com>
2026-10-05 23:47:01 -07:00
Neil ac3a25ca37 Wait for Codex's live composer before pasting linked issue drafts (#25779)
* Wait for Codex's live composer before pasting launch drafts

* Keep timed-out Codex draft delivery consumed across remounts

* Document the checked terminal fixtures used by the remount test

* Distinguish Codex's reserved footer row from early multiline input

* Complete transcript baselines and suppress hidden terminal strings

* Align watcher scan budget with linked-directory reconciliation

* Drop unused snapshots after main removed the readiness census
2026-10-05 23:38:28 -07:00
Brennan Benson eaaae0196f feat(native-chat): record fresh sessions after failed restoration (#25747)
* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore

A chat whose saved conversation the agent cannot reopen can now continue in a
fresh one: the handle chain records the new conversation as a creation that
replaces the lost one (which, why, and when), keeping every earlier link.

Rows keep a shape older builds read: the stored chain starts at the latest
replacement and carries the earlier links inside it.

* refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row

Older builds only read Claude and Codex records, so the nested stored form
protected rows no replacement can reach while adding a cap mismatch after a
downgrade. Store the chain as held, refuse a replacement in a Claude or Codex
chain until one has a stored shape older builds read, and refuse a supersession
key on a replacement that names no creation in the chain.

* test(native-chat): prove replacement rows survive downgrade and re-upgrade
2026-10-05 23:22:24 -07:00
Brennan Benson d7a95782d1 Add a shared timeline assembler for structured agent chats (not wired yet) (#25064)
* Move the turn message ordinals and the turn-row revision to the neutral timeline folder

Pure moves so a shared timeline assembler can use them: Codex's message
ordinal counter becomes ProviderTurnMessageOrdinals and Claude's turn-row
revision becomes the provider-neutral agent-journal turn-row revision. Only
names and import paths change.

* Admit one provider event's writes as one transition, and let rows be found again after a restart

- A sink transition is admitted whole or not at all; its steps run back to
  back at their turn in the journal's write queue, and each resolver reads
  the fold with every earlier write landed. A resolver may also say where the
  row belongs (turn scope, provider reference), and the writer always hears
  how the transition landed. A resolved lifecycle batch chooses its
  settlement mutations from the fold at execution.
- New optional row field providerItemRef: the provider's own reference for
  the item a row is, written only where the row's identity cannot spell it
  (Codex keys messages by their place in the turn and renumbers its item ids
  on resume). Set by the creating write, kept by revisions, indexed by the
  journal fold, never read by clients. A downgrade test shows an older host
  and client render such rows unchanged.
- Provider timeline identity schemes (shared legacy arm, Codex) and the join
  index that resolves a provider item to its row from memory or the fold:
  ordinals and request incarnations are read back from the rows, so a
  restart or an evicted entry finds the original row instead of placing a
  new one.

* Recover message ordinals from the journal's highest place, and forget joins read from a replaced epoch

A fresh join index continued a turn's messages at the first free place, so a journal holding only a
later ordinal (an imported or removed earlier row) had its sequence back-filled. The place is now
one past the highest ordinal any row or echoed send holds there, read through a pure scheme reader.
The join caches also drop what they read when the journal's epoch is replaced.

* Spell the subagent thread's message slot without spreading an identity union

* Add a provider timeline grammar and a shared assembler that decides at its turn in the journal

Adapters translate their provider's dialect into a small grammar (turns,
items, streamed text, requests, context facts, session end/reset); one shared
assembler turns it into the journal rows every structured lane writes.

Each event is planned as one sink transition. Which row a write lands on,
whether a replay writes anything, and every change to what the assembler
knows (its ledger) are decided by the transition's resolvers at the event's
turn in the journal's write queue, against the fold as it stands then. A
forecast (the ledger plus admitted events still queued) only answers apply()
at once. So a refused event allocates nothing, a write the journal rejects
leaves no trace in memory, and a restart or evicted cache finds the same rows
again. Text and full snapshots of one provider item share one row and one
lifecycle; reset always flushes text and settles the old session from the
journal; the open-work budget is derived from what is actually open.

Codex migration contracts compare against the existing Codex translator,
including a restart mid-stream and a repeat that outlives the join cache.

* Fix the types and the exhaustive event switch CI reported for the assembler

* Let the journal decide stream lifetimes, request reuse, named sends and background work

A third review found two blockers with the earlier rounds' cause, a remembered interpretation
trusted after the journal moved on:

- A reused request id was judged by its earlier prompt's settled turn before asking which turn the
  new one lands in, so a real approval in a later turn was dropped. The target turn now decides:
  the old turn again is a replay; a different live turn opens the next prompt beside it.
- A text stream checked its row's turn only on its first write, and a turn's end released streams
  by the turn planning expected. Every write now checks the row, a turn's end stops the streams
  whose rows are in it, and turn status reads the journal first, so another writer's Stop wins.

Also: a message boundary drawn by an event the journal held as a replay no longer splits an
anonymous message; a send naming a turn not yet open waits for that turn; the budget charges a
stream's thread and turn strings and the turn caches are byte-bounded; the open turn ends when the
journal shows it settled; a turn's opener is read from the journal's row.

Background work is now Orca's existing background-task row instead of a tool call flagged
`outlivesTurn` (a flag remembered only in memory, so a restart failed the task). A turn's end
never settles that row, so it survives restarts; session end leaves one in flight unverifiable.
Three tests that opened a background tool call with `outlivesTurn` now open a background-task row
and keep their original expectations about which turn the row stays in.

* Type the unbound assembler helper's drain as the void it reports

* Bound the rows kept for a continued anonymous message and the stopped streams

A row kept for the anonymous stream that may continue it, and the marker that a stopped stream's
queued writes write nothing, lived in the live-stream map and were never removed when no stream
followed. They now live in their own bounded maps, so the live map holds open streams only.

* Keep the journal store under its line limit after the main merge

* Drop the provider item reference, join index and identity schemes from the transition PR

Nothing in production reaches the state they defended (an assembler that lost its memory while
its child keeps streaming the same turn), and the stored Codex id was positional. The legacy
identity scheme moves to the assembler PR with its first caller; the Codex scheme and any
persisted reference wait for Codex to move onto the assembler. The Codex ordinal counter goes
back to codex/, since no neutral code imports it.

* Write a resolved settlement in one transaction through enqueueRows

A settlement too large for one row now commits all its rows or none, through the journal's
existing all-or-nothing write, instead of a row-by-row writer. Every row is built before any
commits, so a settlement naming one item twice is refused before anything is written.

* Drop the transition's landing report; keep the turn-row write fire-and-forget

Nothing reads which steps of a transition wrote. A failed step fails the sink, leaving the
steps before it written; the header says so, and tests cover it plus a settlement whose second
row fails inside the transaction. writeAgentJournalTurnRow returns nothing again, as on main.

* Run a transition's steps as a prefix; drop the paced flag and resolved options

A failed step no longer lets the steps after it write: each step checks, at its own turn in the
journal's queue, whether the write handed over just ahead of it completed, using the queue's count
of completed write bodies (a promise would report the failure only after the next step ran). The
sink fails only once every step has had its turn. The item step's `paced` size bypass and the
resolver's replacement `options` are removed; nothing planned uses them.

* Rebuild the timeline assembler on one admission-order state

The assembler kept a second copy of its state (a forecast beside a ledger), hydrated
the open turn from the journal, and recognised replays, all to recover from losing its
memory while the provider child kept streaming. That never happens: one assembler lives
exactly as long as one child, and a new child is a new assembler in a new generation
whose events land after the dead-generation sweep.

- One state, changed only when the sink admits an event (minted keys included), so a
  refused event takes nothing.
- Every journal-dependent choice is made when the write runs, by keyed reads: a turn row
  is written only where none is, a stream checks its row's turn on every write, a running
  snapshot never lands in a settled turn or relights a settled tool, a request takes the
  first incarnation the journal holds no row for.
- Rows are found by spelling their ids (provider-timeline-rows.ts); no join cache.
- The identity scheme (legacy arm only) lives here with its first caller; requests are
  spelled in their acquisition generation, since JSON-RPC ids restart per process.
- Saved history goes in as `input.history` plus ordinary events with the provider's ids,
  into an empty journal; `session.reset` and every replay rule are gone.
- One terminal-body function (`terminalAgentJournalBody`) is shared with the
  dead-generation settlement.
- The test rig's restart now sweeps and starts a new generation, as production does.

* Cover new running work in a turn the sweep ended

* Run a transition's steps in one queued write that loops over them

The steps of one event now share one turn in the journal's write queue: a
loop writes each in its own transaction through the row writer's
synchronous writeRows (split out of enqueueRows) and stops at the first
throw. Prefix semantics and "nothing lands between the steps" now hold by
construction, so the completed-write counter on the queue, the step gate
and the allSettled barrier are gone; the queue is back to main's bytes.

* End a turn another writer settled the way the provider's end does

A person's Stop settled the open turn's row without a word to the assembler.
The assembler then forgot the turn: its running tools and pending prompts were
never settled, the provider's own end and withdrawal were dropped, and the
turn's text streams stayed counted against the open budget for the life of the
process. Text the provider kept streaming afterwards could land as a message
outside the stopped turn.

- The open turn the journal shows settled ends first, as one transition, through
  the same settlement the provider's turn.end plans; its streams stop and their
  keys drop later text until that turn's end or the next turn opens.
- turn.end and request.withdrawn are admitted for a turn or request the journal
  holds; their settlement writes nothing for rows already settled.
- The budget's re-check frees streams whose turn settled.
- A settled tool keeps its terminal body against any differing write.
- Session-end settlement of lost background work uses the journal's own
  lostLiveWorkJournalBody instead of a copy.
- The rig's window elapses before every read, and restart swaps and disposes
  the old assembler.

* Leave a stopped turn's running tools to the agent's own end

When another writer settles the open turn (a person's Stop), the assembler
now only stops that turn's text and cancels its pending prompts. Running tool
calls stay the agent's: a progress update or completion it reports after the
Stop lands as reported, and whatever is still running settles at the agent's
turn end for that turn, the next turn's open, or the session's end.

An agent's end for an earlier turn while a newer one is open no longer clears
the open turn's activity line or ends its anonymous reply. An unnamed end right
after a Stop ends the stopped turn instead of being dropped. The test rig's
restart no longer writes the dead assembler's window text, matching dispose.

* Pin that a stopped turn's running tools hold budget until the agent's end

* Type the stopped turn's tool progress update as a tool body

* List every event the assembler hands to the decision step

The type-aware lint requires an exhaustive switch with no default case.
Also retitle a Stop test to say what it asserts.

* refactor(native-chat): drop saved-history adoption from the timeline assembler

The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.

* refactor(native-chat): a pending input is only Orca's send now

Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.

* Use current provider handles in transition tests

* Use current provider handles in timeline fixtures
2026-10-05 23:21:22 -07:00
Neil f9c8cd4fc3 Reduce redundant test coverage and unnecessary CI waits (#25806) 2026-10-05 23:19:18 -07:00
Brennan Benson 4e64fa9940 fix(native-chat): start a Codex chat without waiting on the background model probe (#23831)
* fix(native-chat): start a Codex chat without waiting on the background model probe

Opening a Codex chat kicks a session-less model-catalog probe (a throwaway
read-only `codex app-server`) for the account, and the new chat's own
options read then joined that probe's in-flight refresh through the catalog
store's single-flight. When the probe's Codex hung, chat start waited out the
probe's 15s deadline and failed with "codex app-server session exceeded
15000ms", even though the chat's own app-server was up.

Single-flight now joins only a refresh by the same kind of lister. A live
session lists over the connection it already holds; the probe still records
its own failure in the store, and the picker keeps serving whichever listing
succeeded.

* fix(native-chat): never let a chat's model listing join another chat's listing

Single-flight in the model catalog store was split by lister kind, so a new
chat's acquire-time options read no longer joined the session-less probe,
but it still joined any other live chat's in-flight listing for the same
account. When that other chat's Codex was wedged or closing, the new chat
waited out that chat's request timeout and failed to start with its error.

Key in-flight listings by the lister's identity instead: a live session by
its own connection, the probe by the probe itself. A lister still joins its
own in-flight listing, and shouldRefresh still holds back a probe while any
listing for the account is in flight.

* test(native-chat): pin that the catalog store releases an account once every listing settles

shouldRefresh reads 'any listing in flight for this account' from whether the
per-account in-flight map exists, so that map must be dropped exactly when its
last listing settles. Nothing covered this: removing the cleanup, or dropping
the map on the first settle, passed every catalog test. A leftover map would
stop every later background refresh and probe for the account.

* fix(native-chat): name the catalog lister type instead of a bare object

A live session is keyed by its per-spawn catalog handle (minted in the same
acquire as its connection), the probe by itself.

* test(codex): pin that a chat's own concurrent option reads share one listing

* fix(native-chat): keep model picker current across parallel listings
2026-10-05 22:27:09 -07:00
Brennan Benson 020cebeff6 Add standalone Agent Client Protocol client layer (#24990)
* Add standalone ACP protocol client and session runtime

* Protect ACP transport teardown from late stream errors

* Retire incoming ACP request ids before publishing responses

* Narrow ACP configuration requests and transport message types

* Remove redundant ACP request handler return unions

* Keep ACP waits caller-owned and preserve protocol extensions

* Preserve open ACP decisions through prompt completion

* Generate open ACP enums and check the generated schema offline

A newer or vendor enum value (tool kind, tool status, option kind, stop
reason) no longer fails the whole message: generated enums accept the known
literals plus any other string, typed so callers can still narrow on the
known ones. The generated header now records the pinned input digests, the
generator digest and a body hash, so `verify:acp-protocol` catches a stale or
hand-edited file without network access; it runs in lint and the PR workflow.

* Land the ACP runtime contract the agent adapters use

- Deliver notifications other than session/update through
  onExtensionNotification, in arrival order with session updates.
- Accept _meta on prompt, setMode, setModel, setConfigOption and cancel.
- cancel() always sends session/cancel once the session runs, since the
  agent can be in a turn it began itself; only a successful send is shared,
  so a failed write is retried.
- Cancel aborts each open agent request's signal and lets its handler send
  its own answer; -32800 only when the handler rejects.
- Permission requests validate only the session, tool call id and options;
  unreadable fields are dropped with a diagnostic, and any answer Orca
  cannot send is `cancelled` instead of a JSON-RPC error. Agent-started
  turns may ask; whether to show it is the caller's decision.
- AcpAgentError marks the agent's own errors; AcpInvalidResponseError keeps
  the raw answer and validation issues for answers Orca could not read.
- Lines over the size limit are classified by prefix (shared with the Codex
  reader): the owed request fails, an oversized agent request is answered
  with an error, and an unattributable response closes the connection.

* Answer every agent request after an ACP cancel

A cancel that lands before a permission handler starts now still runs the
permission path, so the agent gets the `cancelled` outcome rather than a
request-cancelled error. A handler that ignores the abort no longer leaves
the agent waiting: once the abort has run through, any request still
unanswered gets request-cancelled. Handlers that answer on abort keep their
own reply.

Also renames a lint-rejected helper parameter, replaces a Reflect.apply in a
test, and stops the permission diagnostic from firing with an empty list.

* Let each ACP request handler own its answer after a cancel

Removes the next-event-loop-turn fallback that answered request-cancelled
for any handler still silent after a cancel. It raced answers that were
still being saved (an approval mid-journal-write reached the agent as an
error) and made the outcome depend on event-loop timing. The handler that
owns an agent request now always sends its answer, or throws for
request-cancelled; a request it never answers ends when the connection
closes. A permission whose handler had not started still answers
`cancelled`.

* Register the ACP schema verify step in the PR preflight phase test

* feat(acp): a steer's cancel asks once and never ends the agent

The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle,
then close the connection, which ends the agent. A steer used it too, so a slow agent lost its
process just because the person added a message. requestSteerCancel() now sends session/cancel
once per prompt, cancels the agent's open requests and answers later permissions cancelled, and
never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel()
stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel
paths move into acp-prompt-cancel.ts over one cancel channel.

* fix(acp): a repeated steer shares the cancel in flight; say what the caller owns

Per review: a second steer before the first write lands returns that write instead of resolving
early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that
fails instead must not take the steer until the caller rebuilds the session; the Stop's says a
prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives
the runtime a handler that would allow: the open permission's signal aborts and the late one never
reaches it.

* test(ratchet): require src/main/acp now that this PR lands it
2026-10-05 22:25:40 -07:00
635e7f53cb feat(github-projects): render Board project views as a kanban with drag-and-drop (#19074)
* Add board layout support for GitHub project views

Board-layout views render as a kanban. Columns come from the view's
verticalGroupByFields (the host retries without the selection on older
GHES schemas and the renderer falls back to the Status field), with one
column per single-select option in option order — empty ones included —
plus a trailing no-value column whose drop clears the field. Card drops
reuse the table's field mutation path, committed from a document-level
capture listener because the preload's native-drop bridge stops drop
events before React's root ever sees them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(github-projects): harden board capability probe and drop lifecycle

Review follow-up. The verticalGroupByFields capability probe now matches
parsed GraphQL error messages instead of substring-scanning the whole
response body — partial-error responses echo the field name as a data key
on healthy schemas, so one SAML/FORBIDDEN partial error could permanently
degrade github.com boards for the session. Covered by new
project-view-config tests per the capability-cache testing contract.

Also: only single-select/iteration vertical fields shape columns (a
drifted field kind no longer yields a clear-on-drop no-value column),
optimistic patches resolve the column field through the board/group
config, drag cleanup moved to a document-level dragend listener, the
column dot follows dark mode via the chip CSS variables, the supported-
layout allowlist is a single shared predicate, and the board's edit
handler is stable so its drop listener stops re-registering per render.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(github-projects): serialize board edits and verify rendered drops

* fix(github-projects): preserve refresh baselines and order edits across views

* fix(github-projects): accept source settings projection in cache scope

* fix(github-projects): distinguish view switches from refreshed field baselines

* test(github-projects): type board IPC recordings and verify clear request

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Neil <neil@stably.ai>
2026-10-05 22:23:57 -07:00
Neil ca4e239861 Remove low-value test inventories and duplicate fuzz oracles (#25791) 2026-10-05 22:23:48 -07:00
Brennan Benson c0b07d3a71 Admit one provider event's journal writes as one queued operation, decided when it runs (#25141)
* Move the turn message ordinals and the turn-row revision to the neutral timeline folder

Pure moves so a shared timeline assembler can use them: Codex's message
ordinal counter becomes ProviderTurnMessageOrdinals and Claude's turn-row
revision becomes the provider-neutral agent-journal turn-row revision. Only
names and import paths change.

* Admit one provider event's writes as one transition, and let rows be found again after a restart

- A sink transition is admitted whole or not at all; its steps run back to
  back at their turn in the journal's write queue, and each resolver reads
  the fold with every earlier write landed. A resolver may also say where the
  row belongs (turn scope, provider reference), and the writer always hears
  how the transition landed. A resolved lifecycle batch chooses its
  settlement mutations from the fold at execution.
- New optional row field providerItemRef: the provider's own reference for
  the item a row is, written only where the row's identity cannot spell it
  (Codex keys messages by their place in the turn and renumbers its item ids
  on resume). Set by the creating write, kept by revisions, indexed by the
  journal fold, never read by clients. A downgrade test shows an older host
  and client render such rows unchanged.
- Provider timeline identity schemes (shared legacy arm, Codex) and the join
  index that resolves a provider item to its row from memory or the fold:
  ordinals and request incarnations are read back from the rows, so a
  restart or an evicted entry finds the original row instead of placing a
  new one.

* Recover message ordinals from the journal's highest place, and forget joins read from a replaced epoch

A fresh join index continued a turn's messages at the first free place, so a journal holding only a
later ordinal (an imported or removed earlier row) had its sequence back-filled. The place is now
one past the highest ordinal any row or echoed send holds there, read through a pure scheme reader.
The join caches also drop what they read when the journal's epoch is replaced.

* Spell the subagent thread's message slot without spreading an identity union

* Keep the journal store under its line limit after the main merge

* Drop the provider item reference, join index and identity schemes from the transition PR

Nothing in production reaches the state they defended (an assembler that lost its memory while
its child keeps streaming the same turn), and the stored Codex id was positional. The legacy
identity scheme moves to the assembler PR with its first caller; the Codex scheme and any
persisted reference wait for Codex to move onto the assembler. The Codex ordinal counter goes
back to codex/, since no neutral code imports it.

* Write a resolved settlement in one transaction through enqueueRows

A settlement too large for one row now commits all its rows or none, through the journal's
existing all-or-nothing write, instead of a row-by-row writer. Every row is built before any
commits, so a settlement naming one item twice is refused before anything is written.

* Drop the transition's landing report; keep the turn-row write fire-and-forget

Nothing reads which steps of a transition wrote. A failed step fails the sink, leaving the
steps before it written; the header says so, and tests cover it plus a settlement whose second
row fails inside the transaction. writeAgentJournalTurnRow returns nothing again, as on main.

* Run a transition's steps as a prefix; drop the paced flag and resolved options

A failed step no longer lets the steps after it write: each step checks, at its own turn in the
journal's queue, whether the write handed over just ahead of it completed, using the queue's count
of completed write bodies (a promise would report the failure only after the next step ran). The
sink fails only once every step has had its turn. The item step's `paced` size bypass and the
resolver's replacement `options` are removed; nothing planned uses them.

* Run a transition's steps in one queued write that loops over them

The steps of one event now share one turn in the journal's write queue: a
loop writes each in its own transaction through the row writer's
synchronous writeRows (split out of enqueueRows) and stops at the first
throw. Prefix semantics and "nothing lands between the steps" now hold by
construction, so the completed-write counter on the queue, the step gate
and the allSettled barrier are gone; the queue is back to main's bytes.

* Use current provider handles in transition tests
2026-10-05 22:20:59 -07:00
Brennan Benson dc8b3934b2 Add chat appearance settings: text size, code size, width (#25657)
* feat(native-chat): add chat-scoped color tokens

* feat(native-chat): soften transcript and composer appearance

* fix(native-chat): refine code spacing and faint text styling

* fix(native-chat): wrap prose links at word boundaries

* test(native-chat): refresh background task strip snapshots

* Add native chat appearance settings and card scaffold

* Persist chat text and code sizes with appearance width controls

* Fix chat appearance typing, shortcut routing, and scoped typography

* Use chat call identities in typography regression fixture

* Keep markdown metadata typography independent of chat text size

* Preserve terminal undo chords and stabilize chat resize estimates

* Show only primary shortcuts in chat appearance settings

* fix(native-chat): preserve appearance edits and configured zoom shortcuts

* Remove unused chat appearance summary translation

* Preserve command classification when applying chat code size
2026-10-05 22:12:51 -07:00
Brennan Benson 6c693edf40 Show native chat tool calls as plain sentences (#25654)
* feat(native-chat): add chat-scoped color tokens

* feat(native-chat): soften transcript and composer appearance

* fix(native-chat): refine code spacing and faint text styling

* fix(native-chat): wrap prose links at word boundaries

* test(native-chat): refresh background task strip snapshots

* feat(native-chat): show tool calls as plain sentences

* fix(native-chat): make tool sentences reflect call state

* fix(native-chat): clarify failed commands and subagent sentences

* test(native-chat): exercise command disclosure with real results

* fix(native-chat): preserve command inputs and localize failure rows

* fix(native-chat): retain complete padded command input

* fix(native-chat): read named tool input fields safely

* Refresh native chat tool rows when the UI language changes
2026-10-05 22:05:13 -07:00
Brennan Benson 8d2e9264bc fix(native-chat): ignore the terminal launch command when starting a native chat (#25720)
A custom launch command in settings made structured native chat unavailable:
the renderer route, the host launch-mode decision, and the host's
create-support all treated it as terminal-only, so a user with the chat
default on silently got a terminal. The command names the CLI binary and now
applies to terminal launches only; native chat ignores it.

A start directory outside the workspace still forces a terminal. The internal
route field and blocker are renamed to that concept; the wire reason
`tui_launch_command` is kept and its receipt text now names the start
directory.
2026-10-05 22:01:56 -07:00
Neil fff8718c76 Improve Quick Open matching, file locations and recent history (#25371)
* Match Quick Open queries across path terms and identifier separators

* Honor ignored-file and symlink preferences on file inventory hosts

* Open pasted file locations and remember successful workspace file visits

* Verify bounded directory listing preferences across execution routes

* Preserve symlink preferences on legacy directory fallback

* Preserve bounded Quick Open terms and separator alternatives

* Revoke stale Quick Open selections and preserve literal file locations

* Validate recent Quick Open candidates independently on their host

* Use complete old-host inventory fixtures for Quick Open compatibility

* Keep Quick Open validation current across palette and workspace changes

* Route recent candidates through relay file listing dispatch

* Record Quick Open hook renders as test snapshots

* Localize Quick Open eligibility errors and search aliases

* Validate file inventory with bundled search and runtime preferences

* test: use directory junction fixtures on Windows

* fix: negotiate Quick Open policies with nested SSH hosts

* fix: refresh cached linked folders after target changes

* Retry stale expanded links after reads settle and targets recover

* Scope stale refresh request lifetime to the visible workspace

* test: seed linked-folder fixture before watcher startup

* Keep Quick Open responsive and compatible with older hosts

* Exercise directory discovery with unprivileged Windows junctions
2026-10-05 21:52:28 -07:00
Neil c8ff7e8875 Stop repeated recovery of finished workers (#25679)
* Settle completed worker assignments and bound recovery retries

* Reset queued explicit recovery before starting its retry budget

* Use current provider identity in the inherited journal fixture

* Keep late worker recovery automatic without repeating persistence work

* Repair inherited validation fixtures after updating main

* Keep the released parser available to compatibility tests

* Limit recovery queries and background process inspection to pending work

* Verify historical repair safety and live SSH process retention

* Preserve the same provider import as main in the merge result

* Reduce historical repair and duplicate workspace recovery work

* Align startup prompt fixture with upstream Windows coverage

* Keep healthy worker repair indexed and read-only

* Align startup prompt tests with current main

* Keep pending terminal releases progressing during recovery retries
2026-10-05 21:47:30 -07:00
Jinwoo Hong e2ccf52dc3 fix(remote-runtime): a paired terminal accepts input again after its host app relaunches (#25736)
* test(remote-runtime): reproduce dead input after a paired host relaunch

A relaunched desktop host accepts RPC before its renderer publishes a window
graph. During that gap session.tabs.list answers with an unpublished empty
graph (publicationEpoch "none", snapshotVersion 0, tabs []). The client's
reconnect inventory treats that as removal and retires the pane
(retireRemoteTerminalId(-1)): the pane goes to "ended", the reconnect overlay
disappears, and keystrokes never reach the surviving daemon PTY.

Both tests are red on main by design; they are the repro for the fix.

- e2e: holds the relaunched host's runtime:syncWindowGraph so the client's
  reconnect deterministically meets the unpublished host (4/4 red).
- unit: transport-level repro of the same retirement (red in ~2s).
- restart helper gains a beforeFirstWindow hook; LaunchOptions moves to its
  own module to stay under the max-lines budget.

* fix(remote-runtime): a paired terminal survives its host app relaunching

After the host app quit and relaunched (daemon still running), a paired
client's terminal cleared its reconnect overlay and then ignored all input.
The relaunched host answers session.tabs.list before its renderer publishes,
with a synthesized empty frame. A paired client sees that frame through its
navigation projection as epoch "none:client-navigation", which the shared
"does this frame answer for the worktree" check did not recognise, so the
pane read it as removal and retired itself.

- hostSnapshotAffirmsWorktreeContents treats the client-projected placeholder
  as no answer at any version (the projection adds its navigation revision).
- The reconnect inventory keeps polling on such a frame instead of retiring;
  a published frame lacking the surface still retires the pane.
- Pushed frames that affirm nothing no longer report the surface absent.
- When the bounded inventory wait ends without evidence, the first published
  snapshot carrying the same handle now reattaches the pane, rather than
  waiting out the ~3 minute auto-recovery deadline.

Client-only: the host already labels the frame, old hosts send the same
frame, and the wire is unchanged. The e2e helper drops its type assertions.

* refactor(remote-runtime): simplify the host-relaunch reconnect fix

- One epoch rule: the placeholder epoch, bare or with the shared
  client-navigation suffix, is no answer. Hosts only ever send it at
  version 0, so the version check is dropped.
- Unpublished frames are dropped where pushed snapshots enter, instead of
  a nullable per-subscriber update.
- A published same-handle snapshot fires the parked retry through
  retryNow(), which now also fires a retry parked while still
  'recovering' (nothing is in flight then). This replaces state-reading
  wiring in the listener and lets online/resume fire it too.
- The Windows e2e runs the fixture through PowerShell's call operator.

* test(remote-runtime): prove the e2e hold beat the relaunched host's first publish

The spec now asks the held host for its tab list and requires the
unpublished placeholder before releasing, so a hold installed too late
fails instead of passing against unfixed code. The helper keeps its gate
in a typed global instead of Reflect lookups.

* fix(remote-runtime): keep a fenced handle's parked retry revivable

A retry parked for a handle that needs a replacement was consumed by an
early external trigger and then refused by the epoch's snapshot-wait
guard, leaving nothing to revive the pane. An external trigger now takes
over that snapshot wait as a fresh attempt, and a republished fenced
handle no longer fires the retry. Also refreshes comments that still
paired the placeholder epoch with version 0.

* fix(remote-runtime): a host answers an empty worktree for real once it has published

The host sent its unpublished placeholder for any worktree it had no
entry for, including one its window had never opened, long after
startup. With the placeholder now read as "ask me later", a paired client
opening such a worktree never got its first terminal.

The host now sends the placeholder only until its graph first publishes.
Afterwards a worktree with no tabs gets a real empty answer, and a client
that was told "ask me later" during startup is sent that answer once on
publication. Against an older host the client withholds the automatic
first terminal for such a worktree, which the user can still create.

A host push now fires only a parked retry, never cutting a scheduled
backoff short.

* chore(reliability-gates): reference the empty-worktree tests in gate commands

* test(remote-runtime): find the never-opened worktree by repo id on every platform
2026-10-06 00:47:27 -04:00
Neil 661372d514 Cancel abandoned file searches and prevent stale results (#25370)
* Bound relay git-grep records and clean up capacity failures

Credits @OrcaWin for the original bounded-record proposal.

* Bound filesystem listing and transfer metadata at the execution host

Apply mobile limits before transport, list Markdown through its semantic producer, and retain complete directory results only within explicit capacity budgets. Stream SFTP directory packets and reject oversized transfer plans before reporting success; preserve narrow older-peer fallbacks and validate streamed response retention.

* Release inactive Markdown candidates and support folder scopes

Keep one current document snapshot per consumer and attach completion candidates to the actual editor model lifetime. Resolve folder workspace roots through existing workspace identities so their previews and completions receive the same authorized listing as worktrees.

* Preserve full runtime inventories under aggregate byte budgets

Leave unqualified inventories complete beyond 20,000 files. Charge retained paths and serialized bytes at local and relay producers, reject oversized full inventories explicitly, and bound old-peer response reassembly; keep caller-specific mobile and explicit result limits.

* Bound SFTP directory handle cleanup waits

Send CLOSE even after cancellation, stop waiting after five seconds without an acknowledgement, and ignore late callbacks. Preserve the original capacity or cancellation error and close handles returned after an aborted OPENDIR.

* Preserve registered root spelling in Markdown document paths

* Reject incomplete SFTP directory cleanup and final cancellation

* Bound pending response bytes before stream ownership arrives

* Preserve bounded directory reads on legacy SSH hosts

* Allow bounded per-chunk padding at response byte boundaries

* Bound response payload retention after stream metadata

* Bound complete legacy Quick Open inventories and cache retention by bytes

* Localize Markdown document listing fallback

* Exercise relay search decoding with real byte streams

* Preserve concrete relay stream fixture types

* test: compare distinct bundled ripgrep platforms on Windows

* Preserve Markdown discovery across Windows and legacy SSH hosts

* Make WSL process termination available to the bounds layer

* Cancel abandoned file searches and preserve search keyboard navigation

* Invalidate searches when workspace roots or owners change

* Cover root replacement and legacy search cancellation together

* Exercise search cancellation through the current desktop access seam

* Keep search ownership and cancellation contracts current

* Share local text-search execution and close canceled subscriptions
2026-10-05 21:46:37 -07:00
Neil f59194d481 Bound file-search inventories and release abandoned Markdown data (#25369)
* Bound relay git-grep records and clean up capacity failures

Credits @OrcaWin for the original bounded-record proposal.

* Bound filesystem listing and transfer metadata at the execution host

Apply mobile limits before transport, list Markdown through its semantic producer, and retain complete directory results only within explicit capacity budgets. Stream SFTP directory packets and reject oversized transfer plans before reporting success; preserve narrow older-peer fallbacks and validate streamed response retention.

* Release inactive Markdown candidates and support folder scopes

Keep one current document snapshot per consumer and attach completion candidates to the actual editor model lifetime. Resolve folder workspace roots through existing workspace identities so their previews and completions receive the same authorized listing as worktrees.

* Preserve full runtime inventories under aggregate byte budgets

Leave unqualified inventories complete beyond 20,000 files. Charge retained paths and serialized bytes at local and relay producers, reject oversized full inventories explicitly, and bound old-peer response reassembly; keep caller-specific mobile and explicit result limits.

* Bound SFTP directory handle cleanup waits

Send CLOSE even after cancellation, stop waiting after five seconds without an acknowledgement, and ignore late callbacks. Preserve the original capacity or cancellation error and close handles returned after an aborted OPENDIR.

* Preserve registered root spelling in Markdown document paths

* Reject incomplete SFTP directory cleanup and final cancellation

* Bound pending response bytes before stream ownership arrives

* Preserve bounded directory reads on legacy SSH hosts

* Allow bounded per-chunk padding at response byte boundaries

* Bound response payload retention after stream metadata

* Bound complete legacy Quick Open inventories and cache retention by bytes

* Localize Markdown document listing fallback

* Exercise relay search decoding with real byte streams

* Preserve concrete relay stream fixture types

* test: compare distinct bundled ripgrep platforms on Windows

* Preserve Markdown discovery across Windows and legacy SSH hosts

* Make WSL process termination available to the bounds layer
2026-10-05 21:45:52 -07:00
Brennan Benson b58f8197dc Soften native chat colors and code surfaces (#25651)
* feat(native-chat): add chat-scoped color tokens

* feat(native-chat): soften transcript and composer appearance

* fix(native-chat): refine code spacing and faint text styling

* fix(native-chat): wrap prose links at word boundaries

* test(native-chat): refresh background task strip snapshots
2026-10-05 21:40:55 -07:00
Neil bfd1e9c579 fix(codex): keep working status when Escape closes search or permissions (#25769)
* fix(codex): preserve working status when Escape dismisses a view

* test(codex): cover navigation Escape during terminal exit cleanup
2026-10-05 21:34:27 -07:00
3ee3a41b6f fix(markdown): strengthen dark table grid lines (#25653)
Strengthen only dark-mode table cell borders in Markdown Preview and the rich editor by mixing foreground into the existing border token. Keep light mode and table geometry unchanged.

Fixes #24982. Consolidates #24984, #25050, and #25653.

Co-authored-by: Paramon <andrii.paramonov@gmail.com>
Co-authored-by: kana001-bit <288527232+kana001-bit@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-05 21:25:56 -07:00
Brennan Benson 3285f8214c Share the managed process lifecycle for structured providers (#25204)
* fix(native-chat): a child's root exit is reported even during its close, and bookkeeping after it never reads as unproven

- Both connections report the root process's exit once, with `expected` set when a close had
  begun. A close that came back unproven and whose root exits later is finished by the adapter,
  and its end reaches the host like any other.
- A Claude close whose resume-point write fails after the exit was proven, and a Codex close whose
  terminal row is refused, now end the session and report the failure, instead of keeping a dead
  child indexed as if its exit were unproven.
- A Codex close whose forced tree kill can't prove the descendants gone but saw the root exit
  reports the descendants and counts the root exit.
- Every child exit with an identity, expected or not, is forwarded to the host.

* fix(native-chat): the exit ends the child's record; an unfinished stop is the child's own close, which everyone joins

- The host keeps no stored "stop still owed" record any more. A stop begins the child's close
  (`child.close`), which lives on the child and ends with it. A second Stop, the idle reaper,
  quit, a send and an option/answer/goal/rewind all join that close instead of retrying a
  separate obligation.
- A caller waits on the close only as long as the step deadline; the close itself is never
  abandoned. A proof that lands after every caller stopped waiting reaches the host as the
  adapter's report of that exit, which ends the record through the same handler.
- Once the exit is proven, draining, settling, the lease release and the adapter's
  acknowledgement are each attempted and reported on failure; none keeps the child on record.
  A start, and the handle's close, write a release that failed from this host's proof of that
  exit, so a failed write never refuses a send.
- A start that meets a close still unverifiable is refused with `previousExitUnverifiable`, so the
  queued message is rejected with a send-again reason; nothing is held and nothing starts beside
  the old process.
- The idle sweep goes back to idle reaping only.
- Removes #24333's retry entry points, the wait row and its hold rule, the ask/failure cursors on
  the stored record, and the stop's own wake.

Tests replace the #24333 unproven-stop test: a send joining an unproven close and an in-flight
one, a late proof past the caller's bound, a root exiting after its close gave up, a proven exit
whose resume-point write and lease release both failed, an unverifiable close rejecting the send
and refusing an option change, a surviving descendant, quit and the idle reaper; and Codex's
unverifiable, late-exit and joined-close cases.

* fix(native-chat): a message refused because the old process's exit is unverifiable says so, and to send again

The start failure for a refusal with reason `previousExitUnverifiable` reads "Orca couldn't
confirm Claude's previous process ended. Send your message to try again." instead of "Claude
couldn't restart." The status-row kind and the refusal reason stay in the shared lists for rows
and hosts that still carry them; the catalogs keep one sentence for both.

* fix(native-chat): a close's verdict is the root's exit alone, and what follows it is logged

- A Claude close resolves as soon as the root's exit is proven: the session ends and its `ended`
  report goes out then. Saving the resume point runs afterwards and a failure is logged, so a slow
  or hung write never reads as an unproven exit or keeps a dead child on record.
- A root that exits after its close came back unproven finishes that close through the same path
  as any close, so the session's child work is published as ended (background tasks and subagents
  no longer stay shown running for a dead agent), and a failure there is logged.
- Codex logs a refused final row, and reports a root exit whose forced tree kill could not prove
  the rest of the tree gone the way Claude does, so the host logs it and blocks nothing.
- Both adapters take the host's logger for this bookkeeping.

* fix(native-chat): one handler ends every child's exit, and a join waits on the adapter's own close

- One exit handler (`structured-agent-session-child-exit`) ends a child's record for an exit
  expected or not. `expected` only changes what the chat is told: the stop's cause, its end at the
  stop's ask, the settlement id, and no crash outcome row. The lease release keeps the exit's
  evidence; the handoff guard, lifecycle barrier, sink release and adapter acknowledgement apply to
  both. A Claude journal-sink failure ends in the same step as its stop, as Orca's own fault.
- Joining a close is asking the adapter, whose close is memoized while it runs and bounded by its
  own kill escalation; the host keeps no attempt of its own and no 10 s caller bound. An ask after
  a close came back unproven runs the stop again.
- A close's end is stamped where its stop was asked for (a repeated ask moves it), so the closed
  chat and failed start checks order a message accepted meanwhile after it.
- A start refused because the old exit is unverifiable rejects what was queued in the same step.
- The end of a close the host asked for no longer waits on the cross-session recovery chain.
- The kill no longer waits for the stop event's write; the journal writes rows in order.

* fix(native-chat): an exit's lease release lands whatever the length of its reason

A crash's reason can carry kilobytes of the provider's stderr, and a lease whose death detail is
over 512 characters fails the store's own check. The exit handler cut it, but the release a start
or the chat handle's close re-derives did not, so after a crash whose own release failed every
message was refused as not resumable until restart. The record's builder now cuts the detail to
the record's bound, so no writer can hand it one too long.

* fix(claude): a proven close waits at most 2 s for the output it already wrote

Once the root's exit is proven, the close still waited for the SDK's output reader to end. Something
outside the process tree that holds the output open would keep that close, and every send, Stop
and quit joining it, waiting with no bound. The wait is now bounded; past it the close resolves as
proven and the open output is logged.

* fix(codex): an exit reported inside Orca's close keeps the reason Orca closed it for

The connection reports the app-server's exit inside the close that ends it, so that report ended
every Codex close and replaced the close's own reason (for example, a provider frame that could not
be recorded) with the connection's stderr text in the ended record and the lease's exit evidence.
The session now records Orca's close with its reason, and the exit it ends keeps that reason. The
test connection reports its exit inside close the way the real one does.

* fix(native-chat): quit stops delivery before it drains exit recovery

Every exit now wakes delivery, and teardown drained exit recovery before it stopped delivery, so an
exit settled in that window could start a fresh agent that teardown then killed. Teardown stops
delivery first; queued messages wait for the next launch.

* docs(native-chat): the unverifiable-exit refusal no longer names a caller's wait

The caller's bounded wait was removed; the comment describes the close as it is now.

* fix(native-chat): a stop whose kill did not take is logged, and the next ask kills again

When a close's kill leaves the agent's root running, the host now logs it. Tests pin what a later
ask does: each connection runs its whole stop again (Codex sends SIGKILL a second time), refuses
input meanwhile, and proves the exit once the kill takes.

* fix(native-chat): a start refused over the old process says Orca couldn't stop it

The host reaches an unverifiable verdict only after its own kill left the agent's root running, on
the machine that runs the agent, so the sentence now says that: "Orca couldn't stop {agent}'s
previous process." The refusal reason, failure kind and wire shapes are unchanged. The host test
also checks the failed kill is logged.

* docs(native-chat): an unverifiable close verdict is a root that survived the kill

The host's close runs where the agent runs, so lost contact never yields this verdict; the comment no longer says it does.

* fix(native-chat): a kill that did not take is reported once, by whoever met it

The log added at the close fired beside a Stop's own failure report for the same event. A stop
still reports it through its failure; a send or option change refused over it now logs it at the
refusal, the only place it is otherwise invisible.

* test(native-chat): a second Stop joins a close the first could not prove and retries its kill

* fix(native-chat): say a start refused beside an unstopped process plainly

The rejection now reads "Couldn't stop {{agent}} from before. Send your message again to try once more."
This kind has its own send-again step; every other failure keeps "Send your message to try again."

* Move provider process supervision and stream reading out of Codex

* Preserve teardown behavior with checked mock types after move

* Apply provider launch environment and caller teardown labels

* Share managed provider launch, exit observation, and close

* Preserve synchronous stdin closure before the grace wait

* Keep the stacked Claude adapter within the module size limit

* Handle nullable provider stdin and processless fixture identity

* Preserve close-time cleanup diagnostics after a late provider exit

* Separate provider root and descendant exit observations

* Give provider child env one owner and gate Codex contract on the shared reader

resolveProviderChildEnv is now the only place that overlays and strips a
provider's environment; the spawn spec and the request-scoped Codex session
both call it. supervisedPosixLaunch only accepts a launch without env fields,
so an override can no longer be silently ignored there. Edits to the shared
stream reader or the env rule now run the real-binary Codex contract job.

* Give every provider one close result, stderr tail and root-only default

- A close reports the root and the descendants with the one verdict vocabulary
  (live / unverifiable / exited); `tree` is null when the close made no
  descendant observation, and the fallback teardown's outcome is written into
  it instead of a side flag a caller could miss.
- The managed process drains stderr and keeps the 8 KiB tail, so no provider
  can forget to drain the pipe.
- Root-only providers take the default close policy and completion rule; only
  Claude overrides them. The one already-exited guard lives in the managed
  close and is recorded when no earlier close ran.
- A failed spawn is never read as an observed root exit.

* Report only observed descendant exits from the fallback teardown

The fallback teardown answered "accepted", and the close turned that into
`tree: 'exited'`, though a Windows tree kill's outcome is never read and an
unreadable process table observes nothing. It now returns what it observed:
`exited` only when the captured descendants were verified gone (or the kernel
reported the group empty), `live` / `unverifiable` as verification found them,
and null when it signalled without observing. Codex's process-tree diagnostic
fires in exactly the cases it did before.

* Give two test fixtures' casts a SAFETY rationale for the changed-code gate

* Claim no observation from an empty process group

ESRCH from the dedicated-group signal says only that the group is empty: a
descendant that left it, or a root that never led it, may still be running.
It now reports no observation instead of `exited`. The unreadable-table
comment says why that case stays "no observation" for root-only providers.

* fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts

On main every close of the chat writes its own Stop and settle. Here a later close joins the
first and writes no row, and the first's settle closed when its kill failed, so a turn that opened
in between and was cut by the next close read as failed. A person's close joining a person's close
whose Stop opened a settle now reopens that settle until its attempt is done.

Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and
the next reads as the person's cancellation (each fails without its half of the fix).

* test(native-chat): build the Stop-opened-turn test's identity with the opaque handle

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle.

* test(ratchet): require src/main/provider-process now that it has landed

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 made the same fix as this branch at a different line; the merge kept both imports.
2026-10-05 21:10:07 -07:00
Brennan Benson 3a03441580 refactor(native-chat): structured agents declare their capabilities instead of shared code naming Claude and Codex (#25076)
* refactor(native-chat): keep the provider resume handle opaque to shared code

Shared structured-chat code parsed each provider's resume handle: Claude's
session id and branch leaf, Codex's thread id, through a 'claude' | 'codex'
union every new agent had to widen. The in-memory handle is now
{ transport, agent, nativeId, providerData? }: shared readers use nativeId,
lease and handle-chain checks compare transport and agent, and only the
Claude adapter reads its leaf (providerData).

Stored and wire forms are unchanged for Claude and Codex. One encoding
module writes their typed shapes and decodes both those and the neutral
shape a new transport uses, which an older build refuses as unreadable
rather than reading as Codex. Key and root strings, which fork seeds,
superseded creations and resume offers persist, stay byte-identical.

The journal's own handle type becomes the journal-row and attach-wire
encoding of the same handle, and the journal identity carries the
neutral handle (null before the provider proves one).

No user-visible change.

* fix(native-chat): derive journal-row provider handles from the journal identity

The journal row converter now takes the identity every caller already holds,
so a row's handle has one obvious constructor. Tests that wrote the in-memory
handle straight into journal rows now build it through that converter, and the
processless Claude fixture names a not-yet-proved handle as null.

* fix(native-chat): refuse a stored provider handle written in both forms

A typed Claude or Codex handle that also carries the neutral form's
transport, agent, native id or provider data named two identities; it
was read as Claude or Codex and the next write dropped the other one.
Such a row now stays unreadable and is set aside untouched.

* refactor(native-chat): route structured agents through registered definitions

The structured-chat shared layer named Claude and Codex in a closed router
type, in option reads, in frame classification, and in renderer label and
validation branches, so adding an agent meant editing shared types. Each
agent now has a definition (capabilities, resting option rules) owned by its
own module; the router is a registry of definition + adapter pairs built at
runtime setup, and shared code reads the definition instead of the name.

No behavior change for Claude or Codex; no wire or stored shape change.

* refactor(native-chat): name the structured agent list once in host types

The host types, conversation commands, provider-session ownership and the
runtime's create path spelled the structured agent list inline; they now use
the one shared type the record and wire already validate against.

* fix(native-chat): narrow the record before reading its agent's option rules

* refactor(native-chat): make the router's registrations the only agent definition lookup

The at-rest option reads looked definitions up in a second, separately populated map, so a
registered definition could route and declare capabilities while the static map decided which
options a chat at rest accepted. The router now answers `definition(agent)` from its
registrations, the at-rest readers take it through the host's adapter, and the built-in
definitions are listed only where the runtime composes the router.

* test(native-chat): use opaque handle in queued rejection fixture

* test(native-chat): share one Codex journal identity in the integration suite

Main grew the suite to the 800-line limit; the opaque-handle import pushed it
over. The two tests built the same identity inline.

* refactor(agent-session): name the handle's adapter state resumeCursor

Rename the neutral provider handle's providerData to resumeCursor before any
row persists the neutral form: it is an adapter-owned resume position (Claude's
transcript leaf), never identity. Claude/Codex stored and wire bytes are
unchanged; their typed shapes never carried the field.

State the stored-form contract (a handle's field set is closed; later per-link
data goes on the chain link, which every build preserves) and pin it with a
record round-trip test. Document that transport records the id space the
native id was minted in, which can differ from the agent's current transport.

* refactor(agent-session): one required agent registry; declarations admit what they claim

A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.

/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.

Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).

* refactor(agent-session): the router applies the declared rewind itself

The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.

* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record

* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.

* test(ratchet): require src/main/provider-process now that it has landed

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 and this branch both added the import at different lines; the merge kept both.
2026-10-05 21:09:26 -07:00
Brennan BensonandClaude db078b6549 fix(claude): API-key Claude users are told they're not signed in (#25163)
* fix(claude): open native chat for API-key users instead of saying "not signed in"

Claude reports tokenSource "none" beside apiKeySource "ANTHROPIC_API_KEY" when it runs on
an API key (environment or settings env). The startup check read only tokenSource, so it
refused those chats as signed out. Refuse only when Claude reports no token source and no
API-key source.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(claude): cover a Console /login key at chat start

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-05 21:00:34 -07:00
Brennan Benson 7f2541373a Stage long agent launch lines where they are typed, so they arrive whole (step 2 of 7) (#23962)
* feat(agent-launch): host-side prompt delivery for agent.launch

The host's agent.launch typed any launch prompt into the shell as part of the
launch command. A long or multi-line prompt then ran line by line in the
shell, and an agent that never showed readiness or crashed at startup had
nothing guarding where its text went.

agent.launch now carries a prompt on the typed line only when the line stays
one line, control-free and at most 512 bytes; otherwise the agent starts
clean and the host pastes the prompt once the agent's own ready signal fires
(bracketed paste plus its composer marker or a quiet render, read only after
the shell's last hand-off, never while the pane's own shell is proven in
front), with main's draft-paste bytes and an Enter 50 ms later. Orchestration
worker starts wait on tui-idle as before. A replay-safe launch admits and
claims its ledger row in one write, Qwen Code gets a second Enter, the
desktop and phone share one launch-refusal classifier, and hosts advertise
agent.launch.prompt-carry.v1.

Split out of #23748, which moves the desktop source-control buttons onto
this path.

* fix(agent-launch): keep a short-lined multi-line prompt on a local zsh launch line, as main did

#24257 moved every multi-line or over-512-byte prompt off the typed launch line and pasted it after readiness. The phone's AI buttons and review notes, whose multi-line prompts main typed whole into zsh, then reached Claude 0.5-3 s later and their RPC reply waited for the paste.

The host now names the shell a local macOS or Linux line is typed into, the way the spawn picks it, and a multi-line prompt rides a zsh line when every line is at most 512 bytes and the whole line at most 8 KB. A real-zsh test types such a line through Orca's own ready barrier and startup write, including when a slow user config makes the write land early. Elsewhere the measured unsafe cases keep the paste: bash 3.2 runs multi-line lines piecemeal, fish drops an early multi-line write, and any shell loses a line over 1 KB written early.

* test(agent-launch): keep the real-zsh launch-line test out of the Windows lane's gate scan

The Windows lane registration check read `const ZSH_PATH = process.platform === 'win32'` (the head of a multi-line ternary) as a Windows-true flag, so `describe.skipIf(!ZSH_PATH)` looked like a Windows-only suite. The file is POSIX-only; the zsh lookup is now a function.

* refactor(protocol): move the agent.launch capabilities into their own module

Main's protocol-version.ts sits at the 300-line cap, so the prompt-carry capability pushed it over. The four agent.launch capabilities and their doc move to agent-launch-runtime-capability.ts, re-exported by name and spread into RUNTIME_CAPABILITIES at the same position; the advertised lists and every export are unchanged.

* refactor(protocol): import the agent.launch capabilities from their own module

`export *` from protocol-version left the four names undefined under the mobile recording loader, which resolves a relative import through a Proxy with no own keys, so 37 phone recordings lost agent.launchReplay. Importers now name agent-launch-runtime-capability directly; protocol-version only spreads its list.

* refactor(agent-launch): drop the unshipped viewMode field and trusted local caller id

Both were inert in step 1 and existed only for step 2. agent.launch will become a
public plugin API, so every wire field is permanent once shipped; a top-level
viewMode reads as "choose terminal vs chat", which the host decides. Step 2
introduces placement and view intent under a placement object instead.

* fix(agent-launch): read Codex's provisional startup from the rule files' hold anchor

Main (#24375) moved Codex's provisional-header check into codex.json's
provisional_startup hold anchor and deleted codex-terminal-readiness.ts, so the
launch readiness hold now asks showsHoldAnchor, as main's own settled check does.

* fix(agent-launch): hold rule-file name titles to quiet for a launch, and census the zsh fixture

Main (#24375) answers a name-only title from each agent's rule file ahead of the
sustained-title lane, so gemini.json's name_title settled a launch readiness wait
on the shell's auto-title while Gemini was still booting. A launch now asks quiet
of every weak idle verdict, as that lane did.

Main's readiness census requires a recorder for every runtime fixture; the zsh
prompt recording is a non-agent control. Gemini's synthetic baseline is
regenerated for this PR's stated change: a bare gemini title is no longer its
rest mark, so name-only rows settle weak, and a fresh working or blocked status
is no longer overridden.

* fix(launch): stage long or multi-line launch lines so they arrive whole

A launch line that is long or has several lines could lose lines or wedge a
shell that is still starting, because it was typed into the terminal before
the shell could read it whole. The terminal host (local daemon, and the SSH
relay) now writes such a line to a script file and types one short line that
runs it, for every POSIX shell; a shell that cannot read Orca's quoting runs
Orca's own agent lines through /bin/sh. If the script cannot be written, the
terminal says so.

The SSH background launch path no longer types the line from the renderer;
the relay stages it like any other terminal.

Bumps the terminal daemon protocol to 40 (main took 39); a v39 daemon keeps
its sessions and gets no staged line.

Restacked onto #24257 from the version reviewed on top of #23748; carries
none of #23748's changes.

* fix(launch): run a staged launch line as its own job in fish and ksh

fish and ksh run the commands of a sourced file in the shell's own process
group, so a staged agent was not a job: Ctrl-Z was dropped (a typed line
stops), and the shell-in-front check saw the shell's group in the foreground
while the agent ran. fish now runs the script with
`eval (string collect < '<path>')`, which runs it as if typed and leaves the
shell's job-control mode alone. ksh and mksh run a long Orca-built line through
`/bin/sh '<path>'`, a real job; a long command the user wrote is typed as
before. The real-shell test now asserts the agent's process group is its own
and is the terminal's foreground group.

The staging rule also reuses hasControlByte and the 512-byte budget from
startup-line-prompt-carry so the carry and staging limits cannot drift.

* fix(agent-launch): paste a launch prompt only when the launched agent is proven in front

A launch pasted its prompt unless a shell was proven in the terminal's
foreground, so any read that could not prove one let the prompt through. After
an agent exited at startup, its shell turned bracketed paste on at the next
prompt, readiness fired on it, and the prompt was typed into the shell:

- macOS: a pane runs its shell under login, so the process-group fence's root
  was never the shell's group and never proved it; the cached foreground name
  could also still name the exited process.
- Windows Git Bash and WSL: the shell-alone-in-its-job check never answers.

Now one fresh read of the terminal's foreground decides: agent, shell or
unknown. Only 'agent' lets a write through (paste, Enter, second Enter, reused
panes too); 'shell' still drops a ready signal. A Windows host never proves the
agent, so there the launch line carries the prompt at any size, as on main.

* test(agent-launch): cover the Windows QA stub, a grok override that exits at once

* fix(agent-launch): keep the local socket alive while a prompted launch waits for its agent

A launch with a prompt now waits up to 60 s for the terminal agent to be
ready before it writes the prompt, and reports not-delivered when the agent
never is. The local runtime socket closes a connection idle for 30 s unless
the request is a long poll, so a launch whose agent exited at startup lost its
reply and the caller saw 'runtime closed the connection' instead of
not-delivered. Classify a prompted agent.launch and agent.launchReplay as a
long poll, as orchestration.workerStart already is for the same wait.

* refactor(agent-launch): narrow the launch params by 'in' instead of a cast

* fix(agent-launch): find a launched agent behind a wrapper that leads its process group

A tcsh or nu launch line runs the agent from /bin/sh '<script>', and a
wrapper script that does not exec its agent does the same: the wrapper leads
the terminal's foreground process group and the agent is a member of it. The
fresh foreground read names the group's leader, sh, so a prompted launch was
refused or pasted late (M4Air tcsh: 2 of 4 not delivered, 2 pasted ~9 s late).

Before that read, take the host's process-group observation as positive proof
when it names the launched agent among the foreground group's members and is
younger than a ready signal's quiet window. It never proves a shell.

* fix(launch): type a plain Orca agent line as is in tcsh and nu

In tcsh and nu, every Orca-built agent line ran through `/bin/sh '<script>'`,
so the agent was a child of sh in sh's process group. When the prompt was
pasted after the agent was ready, the paste guard read sh in front and, on a
loaded Mac, reported the prompt not delivered (4/4 live runs in tcsh).

Now only a line that needs it goes through `/bin/sh`: one over 512 bytes, with
a control byte, or with a character those shells would not read literally
(`!`, backslash, `"`, `$`, backtick). A plain line such as
`claude '--dangerously-skip-permissions'` is typed as on main, so the agent is
the shell's own job.

* fix(agent-launch): judge the foreground-group proof by when its capture began, not how long ps took

The age the host stamps on a process-group observation runs from the start
of its whole-machine ps, so on a loaded Mac a capture begun after the read
was asked for still read as older than 1 s and the proof was dropped. Count
an observation whose capture began after the read was asked for, less the
window a shared capture is reused across.

* test(agent-launch): keep the crash-guard live test out of the Windows lane's gate scan

The Windows-lane registration scan read the const assigned from a platform
check as a Windows-only gate, though the suite runs everywhere but Windows;
find zsh in a function instead, as the real-zsh typed-line test does.

Under load the fresh foreground scan can fail to answer, which lets the
shell's prompt settle readiness (2 of 4 paired runs). The guard still refuses
that write, so assert the refused write, the property that must always hold.

* perf(agent-launch): read a local pane's foreground from its own terminal, not the whole process table

The foreground read that gates every launch paste ran the daemon's
inspectProcess capture and then a fresh scan, each a whole-machine ps; the
fresh one also waits for any capture already running before it starts its own.
Measured here at load 5: 1.2 s a read (M4Air QA: 3.4-5.0 s, and worker starts
17.6-32 s against main's 9-12 s at load 25-84).

On a local macOS or Linux host, take the pane's root pid from the provider's
session inventory and run one ps limited to that pane's terminal. Its
foreground process group decides: the launched agent or any non-shell member
is the agent (a wrapper that did not exec its agent leads the group), a group
of shells alone is the shell. Same pane, same verdict: 2.7 ms a read. SSH hosts
keep the relay's observation and name.

* test(mobile): re-measure the web app's script sweep after agent.launch's capabilities moved out of protocol-version

The mobile web bundle check failed at 124 assets against a ceiling of 123. Main already sat
exactly on that ceiling: its sweep table read 69 scripts at 16 routes while the tree builds 73,
the whole margin of 4. This branch imports the agent.launch capabilities from their own module,
so protocol-version is no longer pulled into the root layout and four other routes. That moves
which routes share which modules, and the Qoder capability module, imported by protocol-version
and the AI-vault resume path, no longer shares an importer set with anything, so it gets a chunk
of its own: 74 scripts.

The fence says to re-derive the bound rather than raise it, so the sweep is re-measured on this
head (every prefix of the sorted route list). The worst route now adds 10 scripts (session), not
9, which moves the pinned shell crossing from 32 to 30 routes; main re-measured on its own lands
on the same crossing.

* fix(agent-launch): a worker's brief needs its agent found in front, and Grok's start answers on its composer

A paired-server worker start whose agent exited at startup typed its brief into the server's
shell, which ran it: the idle wait can settle on a shell back at its prompt, and the brief was
written with no foreground read. Both worker-start paths now check before each brief write, as a
launch prompt is checked: on a host that can find the agent in front it must be there; on one that
cannot (Windows) a shell proven in front still refuses, and anything else writes as before.

A Grok worker start waited ~10 s more than main: its only rest signal is its bare name, which a
launch holds to quiet output, and Grok animates its logo for ten seconds after its composer glyph.
A worker start for an agent whose rest signal is its bare name and whose composer draws a marker
(Grok, DSH, mimo-code) now also answers on that marker, whichever comes first.

* test(startup-staging): expect the CR that submits a typed launch line since #23672
2026-10-05 20:53:07 -07:00
Jinjing 059e81a106 chore(i18n): use 智能体 for Chinese Agent copy (#25767)
Simplified Chinese rendered the Agent concept as 代理, which collides with
代理 = proxy. Standardize on 智能体 for Agent (and 子智能体 for subagent),
while keeping 代理 for proxy senses: HTTP/network proxy, SSH Proxy Command,
reverse proxy, and browser user agent.

- Converted 131 zh catalog values (incl. 子代理 -> 子智能体); 26 proxy /
  user-agent values left as 代理.
- locale-phrase-fixes.mjs: 客服人员/代理商/座席 -> 智能体, 代理 -> 智能体 when
  the English names an agent (guard excludes "user agent"); removed the old
  智能体 -> 代理 rule so the pipeline no longer reverts it.
- Updated value/key/search/macos-tcc overrides to 智能体; proxy keyword and
  proxy override entries unchanged.
- Updated the two policy tests that pinned the old 代理 output.

Gates: catalog verify, coverage --check, extraction, runtime-catalog, and
the locale vitest suites (296 tests) pass.
2026-10-05 20:19:43 -07:00
Kelvin AmoabaandNeil 9b7c936929 fix(dock): clear stale workspace unread counts (#24877)
* fix(dock): count only the unread the app shows

Tab markers outlive the workspace unread flag, so the badge could
sit at 1 with no sidebar dot to find or clear.

Fixes #23363

* fix(dock): follow the sidebar's host filter in the unread count

Counting folder workspaces and per-host rows by flag put rows on the
Dock that a host-scoped sidebar does not show.

* fix(dock): preserve owned folder alerts without expanding unread counts

* fix(dock): honor the existing other-device worktree filter

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-10-05 20:15:53 -07:00
420447e063 fix(agent-status): retire a removed worktree's hook-status rows (#23881)
* fix(agent-status): retire a removed worktree's hook-status rows

Rows whose terminal was never reattached had no teardown path, so they
outlived the worktree in last-status.json for up to 7 days.

Fixes #23068

* fix(agent-status): skip panes another owner has since reclaimed

* fix(agent-status): clear only the removed owner's claims on a shared pane

* fix(agent-status): retire hook-status rows a host scan proved removed

`worktrees:forgetRemovedForExecutionHost` is the only path that ever retires an
off-host WorktreeMeta row: gcStaleWorktreeMeta skips any row whose repo or
hostId is not local. It already prunes the cleanup and space-analysis snapshots
for a worktree a remote scan proved gone, but left that worktree's hook-status
rows behind in `last-status.json` — the same stranding this branch fixes for the
in-Orca delete, reached through the other trigger.

A scan the host answered is positive evidence of removal rather than loss of
contact, so it is the host evidence `ssh-execution-boundary.md` requires, and it
publishes no verdict. The drop is scoped to the scanned host's id, so a same-id
worktree on another connection keeps its rows, and a runtime host falls through
the method's own early return because a paired server owns its own store.

Also names the shared-pane condition in `dropStatusEntriesForRemovedWorktree`,
which was the one place the two-owner logic was hard to read.

* fix(agent-status): stop a removed worktree's row returning on a shared pane

The mixed-owner branch deleted the removed worktree's row but left the pane
unfenced, so the stale row came straight back. `getAgentStatusDisposition`
returns `accept` for an unfenced pane, and the removed worktree's agent can
still post a late turn on a pane it shares with another owner — I reproduced the
`working` row being rewritten to memory and to `last-status.json`, which is the
symptom #23068 is about.

A launch-token fence cannot close this: `ingestTerminalStatus` calls the
disposition gate with no event, so the token check never runs and an OSC report
carries no token to check. The pane fence the sole-owner path already uses does
suppress it, and it lifts on PTY reattach or a new agent's turn.

Fencing the pane would otherwise discard the surviving owner's claims, which is
what the branch existed to protect, so both of its records are captured first
and restored after: its persisted authority commitment (its resume identity) and
its current authority observation. The observation also fixes a second defect on
this path — `deleteStatusEntry` drops it whatever `preserveAuthority` says, so
the surviving owner silently lost `current_runtime` attestation until its next
hook event.

Both are covered by tests that fail against the previous commit.

* fix(agent-status): decide a reused pane by its occupant, and retire rows for local scan-proved removals

A pane has one terminal, so its newest row names who occupies it. A saved
commitment naming a different owner is what an earlier occupant left behind,
and serialization already drops it while the row disagrees. Removal now retires
a pane the removed worktree occupies, and on a pane another owner occupies it
clears only the removed worktree's outlived commitment. This replaces the
capture-and-restore handling of two-owner panes.

The local authoritative-scan prune is the local twin of the SSH scan forget: it
drops a worktree's metadata once the scan proves it gone, and now retires its
status rows too.

* fix(agent-status): preserve foreign authority during worktree removal

* fix(agent-status): preserve terminal connection validation

* fix(agent-status): keep ordinary OSC admission unchanged

* refactor(agent-status): colocate the remote envelope type

* fix(agent-status): revoke removed startup authority evidence

* fix(agent-status): retain foreign startup claims during removal

---------

Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-10-05 20:12:38 -07:00
6012570996 feat(files): reveal Source Control files and terminal file links in the file manager (#24502)
* feat(reveal): reveal changed files and terminal file links

Reveal only existed on file explorer rows. Every reveal menu now shares
one action, so a missing file or a remote workspace behaves the same
in all of them.

Fixes #24003

* fix(reveal): address review findings on missing links and SSH files

A hovered hyperlink to a missing file saved "missing" in the cache plain
path links trust, hiding the path once the file existed. A tab holding an
SSH file from outside its workspace still offered reveal.

* fix(reveal): drop reveal from check-details tabs

Their path is a synthetic id, so reveal could only fail.

* fix(reveal): use scoped file access for directory checks

* fix(reveal): retire superseded Spanish menu labels

* test: preserve explicit file access and untranslated-locale contracts

* refactor(terminal): reveal a file link from its click popover, not the right-click menu

The terminal "Reveal in Finder" action now lives in the file-link popover
(the one with "Open file" and "Open with default app"), as a row with no
click shortcut. It goes through the shared reveal, so a folder link (including
a macOS .app or .xcodeproj bundle) is selected in its parent and never opened.

The row is omitted, not disabled, when the file is not on this machine: an
SSH or runtime-owned file, a pane whose shell runs on another host, or while a
remote runtime is focused. Workspace-root links keep their own "Open in
Finder"/"Open folder" row and get no reveal row.

Removes the right-click path: the context-menu item, the hovered-link tracking
and its OSC 8 / path-existence probe refactors, the directory stat (and its
user-named file-access grant, which only served the open-a-folder behaviour),
and the right-click e2e case. Terminal files the PR touched are back to main
apart from the popover row.

Co-authored-by: Kelvin Amoaba <97001695+AmoabaKelvin@users.noreply.github.com>

* test(terminal): type the reveal settings fixture without assertion

* test(startup): account for encoded Windows prompt lines

* test(compatibility): retain the released journal parser dependency

* test(journal): retain one canonical provider-handle import

* fix(terminal): give the popover's Reveal in Finder row the reveal icon

Every other Reveal in Finder item shows the external-link icon; the link
popover draws it for rows that open outside Orca, so mark the row that way.

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-10-05 20:10:31 -07:00
a259616f7c feat(pi): show Pi and OMP background children as rows (#23201)
* feat(pi): show Pi and OMP background children as rows

Refs #22868

* fix(pi): derive the extension's roster cap from the host's

The generated extension re-typed the 32-row limit as its own literal, so
a change to AGENT_STATUS_MAX_SUBAGENTS would leave the extension posting
rosters the host silently truncates on arrival. Interpolate the host
constant instead, and cover the cap with a test that posts past it.

* test(pi): pin the Pi cancel verdict beside a live child row

Publishing a subagent roster for Pi puts its panes behind the same
child-work guard Claude and Codex sit behind, so Ctrl+C beside a live
child no longer settles a stopped row. The extension reports no
main-agent state, so Orca cannot tell a cancelled turn from Ctrl+C at
the idle prompt of a lead children alone hold open, and keeps the live
row. Both halves of that are now pinned, as is the resume-placeholder
guard on the model_select arm of the same code path.

* fix(pi): end a run with its session and keep children across reload

Pi re-runs the extension on a fresh event bus for /new, resume, fork and
/reload, so the roster kept on that bus was lost and nothing ended the
run the old session's children held open.

* fix(pi): preserve pending completion and OMP pane ownership

* test(pi): use the hook owner context type

* fix(pi): retain completion timers and delivery across session boundaries

* test(omp): preserve unsent completion at session boundaries

* test(journal): retain provider handle import across main integration

* test(startup): account for encoded Windows prompt lines

* fix(pi): serialize status delivery across module reloads

* test(compatibility): retain the released journal parser dependency

---------

Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-10-05 20:04:48 -07:00
1b52be6255 fix(claude-accounts): preserve shared MCP OAuth credentials across an account switch (#21931)
* fix(claude-accounts): preserve shared MCP OAuth credentials across an account switch

Orca's account switch writes the target managed account's own stored credential
verbatim to the global "Claude Code-credentials" Keychain item. That credential
never carries mcpOAuth/mcpOAuthClientConfig/mcpXaaIdp/mcpXaaIdpConfig/pluginSecrets
(they are machine-shared MCP connector state, not per-account), so every switch
silently drops any MCP connections the previous session had.

Merge the live credential's copy of these shared fields into the target credential
right before the Keychain write, live-wins (absence included), mirroring how the
third-party claude-swap tool already treats this exact shared Keychain item.

Fixes #16098

* fix(claude-accounts): skip reformatting shared credential merge when nothing changed

mergeSharedClaudeCredentialFields always re-serialized via JSON.stringify, even when
the shared-key set was identical on both sides. The reformatted-but-semantically-equal
JSON (e.g. missing the original trailing newline) then read as an external Claude Code
refresh to the read-back byte-equality check in runtime-auth-sync, causing every
subsequent sync to wrongly adopt it into managed credential storage.

Return the original string when the merge changes nothing.

* fix(claude-accounts): validate OAuth shape, preserve key order, and degrade Keychain read failures

Address CodeRabbit/pullfrog follow-up review on the shared-credential merge:

- parseCredentialObject accepted claudeAiOauth as null or a primitive; the merge would
  then serialize the malformed target instead of passing it through unchanged. Validate
  it is a non-null object (hasClaudeOauthObject) before merging.
- The no-op guard compared whole-object JSON.stringify output, so an existing shared key
  interleaved among target-only keys got moved to the end of the rebuilt object even when
  its value did not change, producing a formatting-only rewrite. Compare per shared key
  and keep the target's own key order via object spread instead.
- runtime-auth-sync read the live Keychain credential for the merge with the raw
  (non-best-effort) function, so a transient Keychain read error aborted the whole
  account switch instead of just skipping the merge. Use the inherited
  readAggregateClaudeKeychainCredentialsBestEffort, matching every other Keychain read
  in this subsystem.
- Fixed a dead citation URL (scaryghost/claude-swap 404s; the real repo is
  realiti4/claude-swap, verified to contain the cited SHARED_CREDENTIAL_KEYS /
  merge_shared_credential_fields).
- Added docstrings to every function touched by this change.

Tests: 5 new cases (null/primitive claudeAiOauth, interleaved shared key no-op and
update, Keychain read failure does not abort the switch).

* fix(claude-accounts): preserve connector state across all runtime credential writes

Keep live MCP grants in both runtime files and keychain items, excluding shared secrets from managed account capture and refresh read-back. Preserve rotations and revocations through default restoration and reject unreadable live state before destructive writes.

* fix(claude-accounts): retain grants from all live stores on startup

Before Orca has a last-written baseline, a missing shared field in one store cannot prove revocation. Carry present fields from the other live stores, prioritizing a proved newer refresh candidate when available.

* fix(claude-accounts): reconcile connector updates without guessing token freshness

* test(preflight): freeze clock for Windows PATH probe assertion

---------

Co-authored-by: tomarai85 <tomarai85@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>
2026-10-05 19:49:32 -07:00
Brennan BensonandClaude 08a970a3f1 fix(native-chat): show the message rail from the first user message (#25707)
* fix(native-chat): show the message rail from the first user message

The rail on the right of native chat stayed hidden until a conversation had
three user messages, so short chats had no rail at all. Show it whenever there
is at least one user message (still hidden in panes too narrow for it).

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(test): remove duplicate journal fixture handle import

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-05 19:02:40 -07:00
a47e0f5657 fix(feedback): optimize oversized screenshots without losing detail (#23664)
* fix(feedback): shrink oversized screenshots instead of refusing them

Fixes #22776

* fix(feedback): release each shrink canvas and skip the PNG ladder for a JPEG

Two defects in the attach-time shrink ladder:

- Every step allocated a fresh full-size canvas and left it for the garbage
  collector, so a 32-megapixel source could keep six of them alive at once.
  Past the renderer's canvas memory budget Chromium hands back a blank canvas,
  which would upload as a blank screenshot with nothing to notice it. Each
  step's canvas is now released once toBlob has answered.
- A JPEG source ran two full-size PNG encodes first. A PNG of decoded JPEG
  noise is many times larger than the source it has to undercut, so those
  steps could only ever burn time and a large buffer before the first JPEG
  attempt. A lossy source now starts at the JPEG steps.

* fix(feedback): keep the image read queue usable after a settle callback throws

The queue tail is what the next batch awaits. A throw in either settle callback
left it rejected, so every later attach short-circuited into "Could not read the
attached images" without reading anything — and the rejection went unhandled.

* feat(feedback): say when an attachment is still being prepared

Shrinking a full-screen capture takes long enough that the gap between picking
the file and its thumbnail appearing reads as a dropped attachment: the only
other signal was the Send button quietly staying disabled. The attachment hint
now reads "Preparing attachments…" while a batch is in flight.

* fix(feedback): report an undecodable oversized image as invalid, not too large

* fix(feedback): detect an animated PNG by walking its chunks, not a 64 KB text scan

* docs(feedback): stop claiming toBlob encodes off the renderer thread

* refactor(feedback): seed the image read queue ref directly

* test(feedback): cover the real decode, downscale and release path of the shrink

* perf(feedback): skip image batches still queued when the dialog unmounts

* fix(feedback): show the preparing hint only once a read outlasts a short delay

* test(feedback): pin the animated-PNG walk on corrupt lengths and a refused image's budget

* fix(feedback): blame the shared budget, not the file size, when a shrinkable image has no room

* fix(feedback): cap each shrunk screenshot at half the attachment budget

A shrink kept the largest step that fit everything left, so the first
full-screen screenshot took a 3.8 MB PNG, the second fell to a barely
readable JPEG and the third was refused. Re-encodes now target at most
2 MB; an image that already fits is still attached untouched. The
"larger than 4.0 MB" refusal now applies whenever the capped target an
empty budget would give was the one that refused it.

* test(feedback): pin the half-budget shrink cap independently of its constant

* fix(feedback): reserve a shrinking screenshot's capped size, not its file size, for the paste gate

A 6 MB screenshot still shrinking reserved 6 MB, more than the whole budget,
so a second paste during the shrink fell through to the textarea even though
it attached. Each pending file now reserves what the reader can commit under
the same fit and shrink-target rules. Also corrects the minimum shrink target
comment to the measured smallest step.

* docs(feedback): correct the canvas-release and shrink-queue test comments; mock toast.info in the dialog test

* Revert "refactor(feedback): seed the image read queue ref directly"

This reverts commit e2272b825e.

* fix(feedback): keep preparing hint compatible with current main

* fix(feedback): preserve mixed paste text and seed read queue on attach

* fix(feedback): preserve screenshot detail and pending paste text

Optimize oversized attachments only as full-resolution lossless PNG; refuse them with existing warnings when that cannot fit. Bound each pending batch independently so earlier optimization slack, rejected files, or removed drafts cannot under-reserve bytes and swallow co-pasted text. Regression tests demonstrate the old mixed-batch text loss.

* test(journal): retain provider handle import across main integration

* test(startup): account for encoded Windows prompt lines

* test(compatibility): retain the released journal parser dependency

---------

Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-10-05 18:38:06 -07:00
Brennan Benson 8127d94782 test(agent-launch): keep the carry control on a prompt that still cannot ride the line after #23672 (#25734)
#23672 quotes a multi-line PowerShell argument as one physical line, so the control's short
multi-line prompt now rides the line on a Windows host that can prove the agent, and the unit shard
went red. The control now uses a multi-line prompt over the typed budget, which still pastes, and a
new test pins that a short multi-line prompt rides the PowerShell line as one physical line.
2026-10-05 18:20:17 -07:00
Neil 4ccfc166d5 fix(sidebar): preserve host filters across stale UI updates (#25737) 2026-10-05 18:17:51 -07:00
Jinjing 9756daeced chore(i18n): translate 76 new keys to es/fr/ja/ko/zh (#25719)
Delta since last scan (a15a5c8c1a): 87 new en keys plus 18 changed en
values. 11 of the 87 are provider credit-balance keys that
provider-credit-balance-locales.test.ts intentionally keeps absent from
target catalogs (sparse English fallback), so they were left untranslated;
the remaining 76 were translated for all five locales. 15 of the 18
changed values already matched the new English upstream; 3 stale ones
(Gemini CLI legacy copy/label, onboarding arrow spacing) were corrected.

Applied via apply-translation-delta.mjs and the repairTranslatedValue
policy gate (ja open->オープン, zh 智能体->代理, ko spacing). Strictly
additive: +380 locale keys, 0 removed.

Three subagent review rounds (max 4): R1 fixed 11, R2 fixed 3 and
rejected the ja half-width-colon and ko pinned-override findings as repo
conventions, R3 signed off clean on all five locales.

Gates: catalog verify, coverage --check, extraction, runtime-catalog, and
the locale vitest suites (34 files / 288 tests) all pass.
2026-10-05 17:52:44 -07:00
Brennan Benson ea6c9d6f63 fix(test): restore the codexProviderHandle import the #25078 squash removed again (#25728) 2026-10-05 17:46:42 -07:00