Keep older request and reply parsers usable while current peers retain all supported history.
Co-authored-by: nwparker <nwparker@users.noreply.github.com>
* feat(native-chat): a Claude subagent waiting on a permission prompt reads as waiting
A subagent's permission request reaches the parent session's callback naming the
subagent that asked (agent_id) and the tool call it gates (tool_use_id). The
pending request is recorded with the asking agent. On every drain the child-work
producer re-derives which children a pending request blocks and hands that set
to the Claude child decoder, the one owner of each child's live edges: a blocked
child reads waiting on every live edge it reports, and a child that starts or
stops waiting is a live edge of its own. Answering, denying or cancelling the
request returns the child to its prior live state; nothing is stored beyond the
pending requests.
A live task_updated carrying an error now reaches the record as the child's last
message, without an ending or a new state.
The replay test drives a scrubbed capture of the real CLI (foreground allow,
deny, interrupt, background allow, and the main agent's own request) through the
real adapter into the host's child records.
* test(native-chat): a subagent's request names it before its tool call is read
* docs(agent-status): a subagent asking for approval waits in every lane; the parent row keeps the session's own attention
* test(native-chat): hand canUseTool the asking agent without widening the helper's cast
* test(native-chat): an interrupted Claude subagent settles cancelled, not failed
A captured interrupt shows the spawn call's error result ("The user doesn't want
to proceed…") arriving before the subagent's own `task_updated {status: killed}`.
The spawn result ends nothing (the child ends only on its own terminal frame), so
the child stays live until its `killed` status settles it cancelled. A genuine
failure, captured with the subagent on a model that does not exist, sends its
`failed` status before the error result and still ends failed. Both captures now
replay through the real adapter into the host's records.
* fix(native-chat): a Claude subagent's prompt makes the parent row wait, not block
A subagent's pending prompt made the whole session `attention`, which reads as
the main agent's own `blocked` and outranks the fold's waiting arm, so the
parent row read blocked where a CLI Claude parent reads waiting. The main
agent's state now reads only its own pending prompts.
- A Claude prompt row carries the linkage of the agent that raised it: the one
the permission request names, or the owner of the tool call it gates. The
same join decides which child reads waiting, so the two cannot disagree.
- The status summary projects the session's own status from root prompts only;
every other reader (delivery gates, teardown, restart) still asks whether
anyone is waiting on a human.
- An answer keeps the prompt row's linkage by the journal's own rule: a
revision that names no producer keeps the row's existing one.
- The child-tool queries gain the prompt's producer, so a prompt row and a
child record answer "which agent" from the same join.
* test(native-chat): say which ids the permission capture scrubs and which are its own
* fix(native-chat): the status clock dates attention by the session's own asks only
The session's status is now `attention` only for its own pending prompt, so the
clock's fallback to a subagent's ask could no longer be reached, and it read the
journal by a different rule than the status it dates. Both now read root prompts.
The journal also stamps a Codex subagent's prompt with its thread (#22532), so a
Codex child's approval is that child's wait in the Codex lane too. Two tests
written for the earlier rule are updated: a subagent's ask leaves a running
session `working` on its turn's clock, and a Codex child's answered approval
leaves the settled parent's Activity row done with nothing unread.
* fix(native-chat): a completion still says the user is asked when a subagent asks
The turn-completion feed marked a completion `awaitingUser` from the status
summary's `attention`. The status now means the session's own agent is waiting,
so a subagent's pending approval stopped reaching the completion. The projection
now also says whether anyone is waiting on the user, as the delivery gates,
teardown and restart ask it, and the completion reads that.
The waiting-subagent replay answers its prompt with the adapter's current
response shape.
* revert(native-chat): a live Claude task's error stays out of the child's last message
No capture shows a live task_updated carrying an error, and it is unrelated to
a subagent waiting on a permission request; it leaves this PR.
* test(native-chat): settle the Claude session's startup before replaying a subagent's request
A startup frame drained child work during the first await, so answering a
request freed the child even with the answer's own republish removed.
* fix(native-chat): a Claude subagent's prompt row names it as its other rows do
The prompt row stamped only the asking agent's id, so a nested subagent's
request lost the agent that spawned it, its spawn call and its run. It now takes
the linkage the asker's own rows take: the gated tool call's, when that names the
same agent, else the one resolved through the agent's spawn call. The provider's
agent id stays the asker's id.
* fix(native-chat): a subagent's request makes the parent row wait without a child record
The parent row learned that a subagent needed the user only from that
subagent's child record, so a request no record carried (a Codex child the host
never registered, a Claude task past the live cap) left the row working or done
while the approval card sat in the chat.
"Someone in this session must answer" is now one derived session fact. The
projection names two facts instead of a mode flag: the main agent's own status
(attention only for its own request) and structuredAgentSessionAwaitsUser (any
pending prompt). The status summary publishes the second as an optional
awaitsUser, and the shared fold reads it: the main agent's own ask is blocked,
otherwise awaitsUser or a waiting child record makes the row wait. Every caller
picks the fact it means: the completion edge's awaitingUser and the delivery
gate read awaitsUser; the quit snapshot folds the same two inputs the sidebar
does.
* fix(native-chat): a client that predates awaitsUser still reads a subagent's request as attention
A status summary's status is now the main agent's own, so a client built before
the split would read a subagent's request as working (or idle) and fold it with
code that has no awaitsUser input. Clients advertise
agent-session.status-awaits-user.v1; at agentSession.subscribeStatus the host
sends any client that does not the pre-split summary: attention whenever
awaitsUser is set, without the main agent's own tool line, verdict and clock.
The feed and every in-process reader keep the canonical summary. Transitional,
like the turn-item downgrade.
* test(native-chat): a Codex subagent's approval makes its settled parent's Activity row wait
The test pinned the parent row done while a Codex child asked, through a harness
that fed no child records, so it proved nothing about the ask. It now drives the
ask twice through the real host status store: with no child record (the
session's awaitsUser alone) and with the child's own record waiting from
thread/status/changed. Both read waiting with needsAttention while the ask is
open, then done with nothing unread.
* docs(agent-status): a subagent's request reaches the parent row through awaitsUser in every structured lane
The store reference said a Codex child's request still read as the main agent's
blocked and that only the Codex hook lane fed a waiting child. Both structured
lanes stamp the asking child and feed child records, and awaitsUser carries the
request when no record does. The liveness comment goes back to main's: a child's
blocked is a failed task on an older host's legacy rows.
* fix(native-chat): the restart dialog still headlines a subagent's pending approval
The quit snapshot now records the main agent's own state, so a subagent asking
while the main agent worked recorded `working` and the dialog said "Was
mid-reply" where it used to say "Waiting for your approval". The headline now
comes from the snapshot's pending prompt, whoever raised it, with the existing
copy; `state` stays the main agent's own.
* test(orchestration): a subagent's pending approval holds structured mail delivery
Scoping the delivery gate to the main agent's own request left every gate test
green; a subagent's request now has its own case.
* fix(native-chat): a subagent's request is dated by when it was raised, on every client
Since the summary's clock became the main agent's own, nothing dated a wait
that only a subagent's request held: a pre-split client was sent attention with
no clock, where the old host dated it by the subagent's prompt, and a new
client's waiting row fell back to the time it first saw it, so after a reload a
request the user had already read could read unread again.
The session fact is now when someone started being asked: awaitsUserSince, the
oldest pending prompt whoever raised it, and its presence is what awaitsUser
meant. A row waiting on someone else's request takes that as its clock; the
downgrade for a client without the capability dates its attention by it, which
is what the old host published. A cross-version test pinned to the last
pre-split release runs the same journals through that release's projection and
through this one plus the downgrade, and compares the whole summary. The Codex
end-to-end test also reads the host's own status row, and keeps a read ask read
through a later row and a reload.
* test(native-chat): the pre-split parity check compares only the fields the split owns
An additive summary field is safe for old clients, so comparing whole summaries
against the pinned release would redden on one. The wire comment now says how
the downgrade dates attention: the main agent's own oldest ask, else
awaitsUserSince.
* test(runtime): an aged host-held working summary states that nobody is asked
The test built its working summary by overriding the status of a published
approval summary, which still carried awaitsUserSince, so the row correctly
read waiting. It now drops the request as its scenario says.
* fix(native-chat): the chat's subagent block says waiting when the strip does
While a Claude subagent's request was open, the sidebar and the composer strip
read waiting but the subagent block in the chat history a few pixels above
still read "Kicked off 1 subagent working": it shows the journal's roster
state, and the journal records no wait.
The structured chat now hands its transcript the subagents the strip shows
waiting, read from the host's child records through the strip's own row model
and matched by the provider id the roster names each one by. A running entry
the host says is waiting reads waiting in the group row, its entry and its
section head, with the strip's word and the question colour; it reads the
journal's state again as soon as the host stops reporting the wait.
* fix(native-chat): a collapsed subagent group shows a wait beside a failed sibling
A failed sibling took the group row's one alert slot, so a group with a waiting,
a working and a failed child read "1 working +1 failed" and hid the wait; it
now reads "1 working +1 waiting +1 failed". The waiting set keeps its identity
while a child frame changes no wait, so the transcript's subagent rows do not
re-render on every frame, and the test of a wait ending now updates one mounted
row instead of remounting it.
* refactor(claude): one needs-input state on the parent; the asking subagent alone reads waiting
Drop the split of the main agent's own status from a session-wide "someone must
answer" fact: awaitsUserSince, the agent-session.status-awaits-user.v1
capability and its old-client downgrade, and every reader change that only
consumed them (fold, equality, ingest, delivery gate, turn-completion feed,
quit snapshot, resume headline, status clock, status bridge, attention
dispatch) go back to main. The parent row again reads one needs-input state
for a pending request whoever asked, dated as before.
Kept: a request's owner recorded once on its prompt row with full producer
linkage; the asking subagent's own record reads waiting, re-derived on every
update; the chat history's subagent block reads that same state; an answered
subagent request stays in its subagent's group.
A subagent now waits only on a request the user can still answer (its card
open, no answer underway), and the adapter frees it before the host records an
answer or dismissal. So a waiting child record always sits beside the pending
card, and main's fold never reads the parent as waiting on it: no window after
an answer, and no ~3 s wait after a card dismissed by Stop.
* fix(claude): a subagent waits only beside its committed card
A subagent's wait was pushed to the host as soon as its request arrived,
while the request's card row reached the journal at least a microtask later.
So every subagent request published the parent row as waiting before
blocked (the main agent's own fold reads a waiting child that way), and
Activity got an extra unread "waiting" event that main never shows.
The card is now the one record of an open request. The translator records
the asker on the card once (its row's linkage) and counts the card open only
after the sink confirms its rows landed, then publishes the wait; anything
that closes the card (an answer underway, a dismissal handed to the host,
Claude's own withdrawal, the session's end) frees the subagent first. So
every publish that shows a subagent waiting also shows its pending card, and
the parent reads one needs-input state, exactly as on main.
This retires the registry's view of pending requests (unclaimed(), the
asking-child join) and the translator's holdsOpen. The prompt row's linkage
takes one rule: the agent the provider names, else the gated call's owner.
The parent-row proof now runs through the real deferred sink, durable
journal and status feed, publishing as production does, and checks at every
publish that waiting subagents have pending cards and that the parent row
matches a host fed no waits.
* fix(native-chat): a closed sink's dropped writes never read as landed
The sink's written() resolved ok when the sink was closed with writes still
queued, so a subagent's prompt card could count as open with no row in the
journal. written() now reports a close that dropped writes admitted so far
as not landed; drained() and lifecycleBarrier() keep reading a closed sink as
settled.
Tests: a card never opens when its sink closes first; a card Claude withdraws
while the sink holds the cancelled row back closes at once; two subagents
asking at once, and the main agent asking beside a subagent, keep the parent
row as before with each waiting subagent beside its own card; a process that
dies mid-request leaves no subagent waiting.
* fix(claude): a withdrawn subagent request frees its child before its card closes
Main's sink now hands each write to the journal as it is submitted, and an idle journal commits it
and runs the publication at once. Claude's own withdrawal of a subagent's request therefore closed
the card and published the parent row before the child's wait was freed, so one publish showed the
subagent waiting beside no pending card (fg-interrupt replay). The child's wait now also requires the
request to still be open in the registry, and a withdrawal republishes child work before the
journal takes the close.
* refactor(claude): trim subagent request waiting to the common pattern and its essential tests
The chat history no longer marks a subagent block as waiting: the approval card itself carries the
request, and the asking subagent's row in the sidebar and composer strip reads waiting, as before.
NativeChatWaitingSubagentsProvider, native-chat-waiting-subagents.ts and their renderer changes go.
A subagent waits while its request is still open and unanswered in the prompt registry and its card
has landed in the journal. The registry check also covers a withdrawal under backpressure, so the
card list no longer filters pending cancellations itself.
Tests: one integration file replays the captured CLI frames through the real adapter, sink, journal
and status feed (renamed claude-subagent-permission-request.test.ts), with the asking subagent's
state timeline, attribution, nested linkage and a card write that waits for the journal. The
producer-harness waiting test, the redundant prompt-card cases, the harness reducer swap and three
unused captures (deny, interrupt, main agent, failed subagent) are removed.
* test(claude): pin the parent row's dating when a subagent asked first, and narrow the oracle's claim
* Add host-owned OpenCode and Devin account profiles
* Manage OpenCode and Devin profiles in account Settings
* Expose registered account roots to host transcript readers
* Clarify managed profile provider flags
* Retain isolated Electron home in browser sidecars
* Restore inherited account environment and preserve cleanup retries
* Check relay environment values before merging
* Consolidate managed account type imports
* fix(accounts): use existing localized provider names
Align the new Japanese account copy with the existing catalog repair policy.
* Keep managed account baselines private to the execution host
* Align account enrollment help with accepted providers and flags
* Show the active System account in managed profile lists
* Document the validated Linux managed account scope
* Fix managed account removal, credential audits, and runtime bundling
* Quarantine removed accounts and audit captured credential snapshots
* Continue Antigravity IDE and 2.0 history in new CLI conversations (#24692)
* feat(antigravity): bridge IDE history into new CLI conversations
* fix(antigravity): preserve fresh-launch model and environment for IDE references
* fix(antigravity): forward IDE history opt-in through desktop IPC
* fix(antigravity): rebuild remote IDE reference startup on its host
* fix(antigravity): register IDE continuation action labels
* fix(antigravity): confine IDE references and bound metadata reads
* fix(antigravity): localize IDE continuation badges
* Preserve scanner service cache assertions and refresh Antigravity opening metadata
* Preserve Antigravity opening joins and target folder runtime authority
* fix(opencode): retry timed-out SSH plugin updates (#24666)
Preserve bounded retry behavior and the current-main status-envelope fields.
Original-PR: #24124
Reviewed-source: 103144f9c5
Co-authored-by: Justas Brazauskas <brazauskasjustas@gmail.com>
* fix(opencode): keep Go credentials private and resolve backend keys (#24615)
Preserve the complete credential storage, migration, IPC, Settings and rate-limit refresh change alongside standalone GLM plans, current-main database diagnostics and the reviewed unknown-backend environment correction. Keep native discovery cancellation third and selected environment fourth.
Original-topic-commit: 7903f1cddb
Original-topic-commit: 588117b5cf
Original-topic-commit: dedd4f8c86
Original-topic-commit: 6645dae104
Original-topic-commit: a25b80c02c1af7830b0e6a65e72d965b3ad98276
Restacked-from: a25b80c02c1af7830b0e6a65e72d965b3ad98276
Restacked-onto: b032867021
Co-authored-by: kespineira <kespineira@users.noreply.github.com>
Co-authored-by: kevimux <kevimux@users.noreply.github.com>
Reported-by: pullfrog
Reviewed-full-source: 849fe093073f4c1606bd65d79a0c725d955d0d1f
Native-helper-source: 80dbe23237
Reviewed-full-current-source: 36acb57d44adb3d378c0289c8c15f7da0fda214c
* Skip store notifications when refreshed work items are unchanged (#24530)
An identical forced refresh previously returned an empty Zustand patch, creating a new root state and notifying every subscriber. It now returns the current state after renewing the existing freshness timestamp. Real row or metadata changes still publish once.
* Read project activity once while sorting the sidebar (#24549)
Reuse the existing pure recent rank once per actual entry tuple within each synchronous sort; discard the lookup after sorting.
* Skip impossible link and tag matches in docs search results (#24646)
* Skip impossible HTML matches when formatting docs search results
After the existing link-removal phase, skip the unchanged HTML regex only when the current string has no closing delimiter.
* Skip impossible link and tag matches in docs search results
Skip the five unchanged link regexes when their protected-code input lacks ]; skip the unchanged HTML regex when its post-link input lacks >.
* Decode complete transcript lines without copying their bytes (#24546)
Decode each single owned Buffer synchronously; preserve concatenation for multipart records and remove a redundant tail array copy.
* Clear unused speech worker timers after errors and exit (#24924)
Call the existing idle timer cleanup when private current-worker state is cleared; keep current-worker guards, transcript/error order, deliberate warmth and stop deadline unchanged.
* Reuse documentation search excerpts during result navigation (#24574)
Memoize the existing pure excerpt renderer by complete raw text for the current result array, retaining original rendering as a miss fallback.
* Append subagent handoffs without recopying their message list (#24672)
Append each message reference to its private call-local scope list; retain the final merge and existing ordering.
* Cancel plugin retry timers when their fetch is replaced (#24700)
Reuse the existing timer clear block immediately after the fetch generation advances; preserve live retries and all state publications.
* Avoid rebuilding visible file paths for single-row selection (#24743)
Skip the visible path array only for keyboard replacement selection; the existing replacement branch never reads that order.
* Reuse prepared reads while searching saved agent conversations (#24535)
Use schema-complete session columns to enable the existing bounded statement cache without caching result rows.
* Reuse discovered mobile chunks during static traversal (#24774)
Iterate the existing live visited Set in FIFO discovery order instead of maintaining a second shifted queue; both actual callers own untouched esbuild JSON metadata.
* Skip unused image-size calculations in mobile web browser requests (#24941)
Use the existing mobile density budget only for mobile view; pass the existing constant to the existing assembler for web/default mode, where that argument is discarded. No cache, policy, request or native path changes.
* Skip diff analysis after a result has been discarded (#24711)
Move the unchanged pure render-limit calculation after both existing generation and section-token rejection fences.
* Skip late browser grab toasts after their surface closes (#24727)
Reuse the existing mounted-owner ref at the toast presentation entry while preserving all admitted extraction/screenshot and clipboard completion.
* Classify native chat waiting messages once (#24771)
Use the existing native filter traversal to populate a fresh waiting array, keeping the original predicate, policy, projection and output ordering while classifying each actual private slot once.
* Skip discarded Kanban pruning for an empty selection (#24782)
Guard only the existing open pruning branch when both its private selected Set and nullable anchor are empty; retain all original pre-guard projections and nonempty/anchor-only/public-helper behavior.
* Index discovered test files once during shard selection validation (#24532)
Verified selection plans previously validated each selected filename with a scan of all discovered files. One per-call Set now handles membership checks. Exact source/discovery checks, nonempty selection, fallback to all tests, shard balancing and manifests remain intact.
* Clear the watchdog benchmark heartbeat when CPU sampling fails (#24802)
Move the existing initial CPU observation and histogram enable inside the existing try/finally so their failures clear the sampler's owned heartbeat, preserving successful operation order and original errors.
* Reuse the parsed notification when leasing a push delivery (#24645)
* Reuse the parsed notification when leasing a push delivery
Pass the notification already parsed for the dismissal check into the existing private delivery builder.
* Reuse the parsed notification when leasing a push delivery
Pass the notification already parsed for the dismissal check into the existing private delivery builder.
* fix(accounts): canonicalize account removal to rm
Use account rm as the canonical removal command and keep account remove as an alias. Preserve the existing accounts.removeData RPC and the full original 63-path account component.
Original-Account-Source: cf71ae4cb6
Frozen-Account-Base: 08ee7ba9ef
Frozen-Integrated-Main: 53f9ea7839
Private validation source only; no ref or publication.
* Update README downloads badge
* Let retired-cache GC observations settle across the existing six-turn budget (#24967)
* Collect test-selection evidence when full unit tests fail (#24955)
* Collect advisory unit-selection evidence from failed full runs
* Trigger checks after retargeting the evidence fix to main
* Check each project repository once while filtering mobile cards (#24540)
Reuse exact raw source/slug matching decisions within one project filter call, preserving the matcher, membership, negative matches, output order and identity.
* Skip renderer callbacks after their request has ended (#24631)
Reuse the existing pending-request identity check before starting its deferred renderer callback.
* Decode single-piece saved-session tails without copying them (#24632)
Decode the existing owned Buffer directly when the EOF tail has one piece; preserve the existing concatenation for multiple pieces.
* Avoid irrelevant scroll-action searches in Mac snapshots (#24731)
Check the current pure action name before searching for vertical scroll actions in the same immutable array.
* Stop copying every retained browser page before attachment (#24553)
Find the first page/generation match directly in the existing insertion-ordered Map instead of copying all page references into an array.
* Reuse event byte sizes when relay watchers notify several clients (#24558)
Store the existing grouped numeric byte-size array on the already call-local watcher batch sizing object, for reuse by later chunking clients.
* Stop building unused import candidates in the test planner (#24611)
Replace the existing ordered extension map/find with a loop that forms each candidate only when its existing lookup is reached.
* docs(terminal): explain local macOS/Linux shell startup files
Document the actual login/interactive shell startup-file order and existing shell setup behavior.
Co-authored-by: brynnclaw <261708852+brynnclaw@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(editor): recognize Ruby task and configuration files
Add Ruby task/configuration filenames to the existing generated language associations.
Co-authored-by: ggbdpq <ggbdpq@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* Speed up large Markdown Find and render oversized tables (#24948)
* Speed up large Markdown Find and render oversized tables
* Poll table preview geometry outside hidden renderers
* Preserve Markdown navigation through refresh and tab restoration
* Confirm Markdown restoration when refreshed content is ready
* Trace table refresh positions and update Unicode search reference
* Recognize queued measurement scrolls before restoring Markdown anchors
* Rebuild Markdown Find ranges after renderer components change
* fix(source-control): generate clean OpenCode messages locally and over SSH (#24613)
* fix(source-control): generate clean OpenCode answers on local and SSH hosts
Restack the original focused change onto current main, preserving every owned source and test blob and the merged CI contract and journal cleanup fixes.
Original-commit: 64bb15e3f2
fix(source-control): generate clean OpenCode answers on local and SSH hosts
Use configured models and JSON answer/error events, preserve run-first arguments, and handle the precise v2 variant rejection. Hydrate SSH execution-host PATH through the existing bounded login environment resolver before direct spawning.
Credits: andy-murr (PR #5197 SSH environment intent) and coelho-doti (PR #13065 argument-order intent).
Original-commit: 1b60ec5d11
fix(source-control): retry inline OpenCode model and variant options
Original-commit: 98fdecc7a3
Preserve OpenCode named errors without a data message
Restacked-from: 98fdecc7a3
Restacked-onto: f7b1f9d8be
* fix(ci): prevent concurrent pnpm refresh during mobile typechecks (#24776)
* fix(ci): run mobile typechecks without concurrent dependency refresh
* test(ci): check effective Linux E2E package list
* test(ci): preserve the mobile production compiler barrier
---------
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
* test(terminal): restore the live fish fixture prerequisites (#24947)
A restored pane waits for the initial status replay before subscribing to
PTY output. This fixture never settled that replay, so fish printed its
mode-2031 arm before the renderer connected. Its PTY API also omitted the
reset-input listener required by the serializer, aborting attachment.
Settle and dispose the existing startup-snapshot registration and provide
the same reset-listener mock used by the other PTY tests. The real fish
child-stdin assertions and timeouts remain unchanged. No production change.
* fix(shortcuts): defer TUI editing chords in terminal-first mode (#24640)
Restack the original focused change onto current main, preserving every owned source and test blob and the merged CI contract and journal cleanup fixes.
Original-commit: f707cde14a
fix(shortcuts): defer TUI editing chords in terminal-first mode
Original-commit: 0c6348e49e
docs(shortcuts): describe deferred preview terminal chords
Original-commit: be62c1b6c6
Align worktree history shortcut metadata with terminal conflict policy
Restacked-from: be62c1b6c6
Restacked-onto: f7b1f9d8be
* Register supervised Qoder China and Qwen Code (#24616)
* Add Qoder session history and search with real CLI coverage
* Allow the real Qoder marker file to end with a newline
* Keep Qoder tool output out of history previews and search
* Keep Qoder search pages readable by older clients
* Verify persisted Qoder history after a real generated and resumed task
* Negotiate Qoder filters before searching an older execution host
* Combine search client imports for the CI plugin gate
* Keep the relay search oracle aligned with legacy agent filtering
* Register supervised Qoder China and Qwen lifecycle integration
* Cover Qoder China mobile assets and mixed-host resume gates
* Verify Qoder provider tags against the older released wire parser
* Verify China and Qwen keep independent Windows hook scripts
* Verify Qoder registrations against the installed older Windows release
* test(qoder): align search capability contracts and pin old-host fencing
* fix(qoder): rank exact picker identities and command aliases first
* test(qoder): preserve the regional CLI shared icon expectation
Keep the full bundled-asset and no-remote-image checks, with an explicit
shared-logo basename for Qoder China. The map also works with older
catalog type unions.
* fix(qoder): align China catalog entry with fallback order
---------
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
* Use the measured pnpm lookup policy automatically in hosted root CI (#24951)
* Select lookup automatically for the measured hosted root-install profile
* Record hosted automatic-mode cold cache publication proof
* fix: bound remote generation setup and honor OpenCode option terminators
Count execution-host profile resolution inside the existing request deadline, cancel its waiter promptly, and pass only the remaining time to the child. Shared bounded profile probes keep their existing cache lifetime; the SSH transport margin is unchanged.
Read OpenCode output format from active final-argv options before -- so literal prompt arguments cannot select the JSON finalizer.
Fresh exact-source controls reproduce nine failures before; 139 related checks pass after, including primary/fallback delays, deadline boundaries, cancellation, parser metadata, and SSH lanes. Node typecheck and strict changed-file lint pass.
* Continue Antigravity IDE and 2.0 history in new CLI conversations (#24692)
* feat(antigravity): bridge IDE history into new CLI conversations
* fix(antigravity): preserve fresh-launch model and environment for IDE references
* fix(antigravity): forward IDE history opt-in through desktop IPC
* fix(antigravity): rebuild remote IDE reference startup on its host
* fix(antigravity): register IDE continuation action labels
* fix(antigravity): confine IDE references and bound metadata reads
* fix(antigravity): localize IDE continuation badges
* Preserve scanner service cache assertions and refresh Antigravity opening metadata
* Preserve Antigravity opening joins and target folder runtime authority
* fix(opencode): retry timed-out SSH plugin updates (#24666)
Preserve bounded retry behavior and the current-main status-envelope fields.
Original-PR: #24124
Reviewed-source: 103144f9c5
Co-authored-by: Justas Brazauskas <brazauskasjustas@gmail.com>
* fix(opencode): keep Go credentials private and resolve backend keys (#24615)
Preserve the complete credential storage, migration, IPC, Settings and rate-limit refresh change alongside standalone GLM plans, current-main database diagnostics and the reviewed unknown-backend environment correction. Keep native discovery cancellation third and selected environment fourth.
Original-topic-commit: 7903f1cddb
Original-topic-commit: 588117b5cf
Original-topic-commit: dedd4f8c86
Original-topic-commit: 6645dae104
Original-topic-commit: a25b80c02c1af7830b0e6a65e72d965b3ad98276
Restacked-from: a25b80c02c1af7830b0e6a65e72d965b3ad98276
Restacked-onto: b032867021
Co-authored-by: kespineira <kespineira@users.noreply.github.com>
Co-authored-by: kevimux <kevimux@users.noreply.github.com>
Reported-by: pullfrog
Reviewed-full-source: 849fe093073f4c1606bd65d79a0c725d955d0d1f
Native-helper-source: 80dbe23237
Reviewed-full-current-source: 36acb57d44adb3d378c0289c8c15f7da0fda214c
* fix(opencode): enforce deadline through executable startup
Keep synchronous Windows PATH resolution and child startup within the existing request budget.
---------
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
Co-authored-by: Justas Brazauskas <brazauskasjustas@gmail.com>
* fix(accounts): recover interrupted profile removal on startup
Scan private removal backups before quarantined cleanup, reuse the existing state schema to validate the exact UUID, and preserve registered or unmarked profiles. Mark account rm as destructive using the existing typo recovery policy. Retain the full original account component and public author ancestry.
Account-PR: 24636
Original-Account-Source: cf71ae4cb6
Reviewed-Public-Parent: 6abaf736e3
Frozen-Integrated-Main: 53f9ea7839
Private source only; no ref or publication.
* fix(persistence): preserve projectGroupOrder on restart for flat folder-scan groups
Retain saved project ranks when the project remains in its folder-scan group.
Co-authored-by: lurunzi <lurunzi@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(runtime): reclaim temp files orphaned by interrupted mobile store writes
Reuse the existing stale temporary-file cleanup policy when each mobile store opens.
Co-authored-by: LDH1103 <ldh517525@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(terminal): show actionable copy for missing folder workspace paths
Show the existing folder-path failure with concrete recovery instructions.
Co-authored-by: Wayn_Liu <wayntingliu@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* docs: reference local orca.yaml and .worktreeinclude
Document the existing workspace configuration, include/copy rules and sharing behavior.
Co-authored-by: Neil <neil@stably.ai>
* fix(orchestration): allow model selection for OMP workers
Pass an explicitly requested model through the existing OMP worker launch catalog.
Co-authored-by: Neil <neil@stably.ai>
* fix(cli): preserve primitive success results in JSON output
Check for an object before inspecting screenshot fields so primitive success results remain printable.
Related: https://github.com/stablyai/orca/pull/14735
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: VXNCXNX <VXNCXNX@users.noreply.github.com>
* Show loaded file changes before deleting a workspace
Expose up to ten already loaded changed paths without altering deletion counts, hydration or authorization.
Related: https://github.com/stablyai/orca/pull/22778
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Michiel de Gooijer <mdgooijer@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(keybindings): record and match Option+digit shortcuts on macOS
Use the physical digit key for explicitly assigned Mac Option+digit shortcuts while retaining modifier checks.
Co-authored-by: marcuslannister <marcus@lannister.cc>
Co-authored-by: Neil <neil@stably.ai>
* fix(editor): don't throw closing a stale active file
Recover stale active-file selection within the owning workspace before choosing a surviving editor.
Related: https://github.com/stablyai/orca/pull/24715
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: Neil <neil@stably.ai>
* Fix Project Settings targeting for multiple local checkouts
Carry the selected checkout identity into the existing project Settings action.
Co-authored-by: fsmeier <1506919+fsmeier@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Neil <neil@stably.ai>
* feat(documents): open CSV and TSV files from the OS
Extend existing OS document associations and delivery to CSV/TSV, preserving restoration and authorization.
Co-authored-by: Neil <neil@stably.ai>
* feat(cli): link GitHub and GitLab items on worktree create and set
Expose and validate the existing workspace link fields, preserving omitted values and provider identity checks.
Related: https://github.com/stablyai/orca/pull/22609
Co-authored-by: marco song <marco.song@mvlchain.io>
Co-authored-by: SongMarco <20613630+SongMarco@users.noreply.github.com>
Co-authored-by: Luca Critelli <lucacri@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* feat(sidebar): include folder workspaces in keyboard navigation
Use the rendered sidebar row order and host identity when cycling through folder and Git workspaces.
Related: https://github.com/stablyai/orca/pull/10555
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: JeongUk Park <jeongph.dev@gmail.com>
* fix(build): turn off MSBuild file tracking for Windows native rebuilds
Default Windows native rebuilds to TrackFileAccess=false while preserving explicit caller preferences.
Co-authored-by: B1nh M1nh <43268322+b1nhm1nh@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Neil <neil@stably.ai>
* Fix Ctrl+M in Linux and Windows terminals
Keep the Minimize menu item but disable its accelerator registration on Linux and Windows.
Co-authored-by: Zhichang Yu <yuzhichang@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Ahmed Nagy <ahmednagy25t@gmail.com>
Related contribution: https://github.com/stablyai/orca/pull/24143
* fix(startup): read a nushell login PATH from $env.PATH
Use Nushell login command syntax and preserve the actual login PATH value.
Related: https://github.com/stablyai/orca/pull/22677
Co-authored-by: Kh05ifr4nD <meandSSH0219@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(cursor): run local hooks through sh for non-POSIX login shells
Keep POSIX hook syntax inside one quoted sh command so the login shell can invoke it safely.
Related: https://github.com/stablyai/orca/pull/22663
Co-authored-by: Kh05ifr4nD <meandSSH0219@gmail.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Neil <neil@stably.ai>
* Fix word wrap for both panes in side-by-side diffs
Forward wrapping to both diff panes through the existing editor option path and clean up listeners.
Co-authored-by: Wooseong Kim <innocarpe@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(monaco): highlight Svelte block closers inside markup
Return Svelte block closers to the existing markup tokenizer state.
Co-authored-by: Wooseong Kim <innocarpe@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(repos): keep active clone dialog open on outside clicks
Ignore accidental outside dismissal only while cloning; Escape, Close and Back still cancel.
Related: https://github.com/stablyai/orca/pull/24581, https://github.com/stablyai/orca/pull/23430
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Nawapat Buakoet <nawapat.b@covest.finance>
* fix(editor): keep chat visible after closing Markdown tabs
Count remaining unified chat tabs before clearing workspace selection during editor close.
Related: https://github.com/stablyai/orca/pull/24273, https://github.com/stablyai/orca/pull/23760
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Wooseong Kim <innocarpe@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* Honor typed starting numbers in chat ordered lists
Forward the existing parsed ordered-list start attribute to both chat Markdown renderers.
Related: https://github.com/stablyai/orca/pull/19765
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Frederic Barthelemy <git@fbartho.com>
* Normalize base source once while checking changed-code diagnostics (#24557)
Reuse the exact existing moved-code matching algorithm with run-local lazy normalization of immutable base source blocks across diagnostics.
* Support real OpenCode sessions in native Chat (#24647)
* Use bounded OpenCode context for vault session continuation
OpenCode database and synthetic row paths are not text transcripts. Use the
vault preview or captured pane context, preserving actual transcript paths
containing a hash and supporting both OpenCode lanes and Windows paths.
Adapted the intent of #11859 and extended it to actual installed v2 vault rows.
Co-authored-by: mrcha033 <mrcha033@users.noreply.github.com>
* Read real OpenCode sessions in terminal-backed native Chat
Reuse the bounded AI Vault SQLite worker for v1 and v2 session pages and live updates. Keep terminal input as the real execution path and pace OpenCode Stop through its two-Escape interrupt.
Co-authored-by: xodmd45-ctrl <xodmd45-ctrl@users.noreply.github.com>
* fix(opencode): publish approval cards for permission requests
* Send OpenCode native approval through its Enter selector
* Resolve mobile Chat readability for folder workspaces
* Bound OpenCode part batches and preserve v2 image attachments
* Prefer live migrated OpenCode sessions over legacy copies
* Consolidate mobile Chat eligibility test imports
* Consolidate OpenCode SQLite protocol type imports
* Update native chat settings contract for both OpenCode agents
* fix(native-chat): reconcile bounded OpenCode transcript reads
* fix(native-chat): dispatch OpenCode questions safely
* fix(native-chat): keep native discovery and transcript windows current
* feat(accounts): link standalone GLM Coding Plans (#24618)
* feat(accounts): link standalone GLM Coding Plans
Adapt the reviewed GLM accounts contribution to current main, retain Antigravity behavior, guard late credential results, expose storage protection, and redact quota errors.
Co-authored-by: Luchong <lu740528977@gmail.com>
* fix(accounts): retain GLM credential results during quota refresh
* fix(accounts): make GLM credential editing desktop-only
* fix(accounts): mirror the host GLM site in paired clients
* fix(accounts): report unknown GLM host details and split web settings tests
Apply the independently reviewed Accounts correction from697284a without the v2 adapter commits. Preserve the saved-key store and serialized write behavior.
* fix(zcode): ship required GLM account translation entries
* chore: record GLM reconciliation hook validation
* chore: validate installed GLM commit hooks
* test: complete GLM account fixtures and web API inventory
---------
Co-authored-by: Luchong <lu740528977@gmail.com>
* fix(native-chat): route transcript requests through shared SQLite worker
* fix(ci): prevent concurrent pnpm refresh during mobile typechecks (#24776)
* fix(ci): run mobile typechecks without concurrent dependency refresh
* test(ci): check effective Linux E2E package list
* test(ci): preserve the mobile production compiler barrier
---------
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
* test(terminal): restore the live fish fixture prerequisites (#24947)
A restored pane waits for the initial status replay before subscribing to
PTY output. This fixture never settled that replay, so fish printed its
mode-2031 arm before the renderer connected. Its PTY API also omitted the
reset-input listener required by the serializer, aborting attachment.
Settle and dispose the existing startup-snapshot registration and provide
the same reset-listener mock used by the other PTY tests. The real fish
child-stdin assertions and timeouts remain unchanged. No production change.
* fix(shortcuts): defer TUI editing chords in terminal-first mode (#24640)
Restack the original focused change onto current main, preserving every owned source and test blob and the merged CI contract and journal cleanup fixes.
Original-commit: f707cde14a
fix(shortcuts): defer TUI editing chords in terminal-first mode
Original-commit: 0c6348e49e
docs(shortcuts): describe deferred preview terminal chords
Original-commit: be62c1b6c6
Align worktree history shortcut metadata with terminal conflict policy
Restacked-from: be62c1b6c6
Restacked-onto: f7b1f9d8be
* Register supervised Qoder China and Qwen Code (#24616)
* Add Qoder session history and search with real CLI coverage
* Allow the real Qoder marker file to end with a newline
* Keep Qoder tool output out of history previews and search
* Keep Qoder search pages readable by older clients
* Verify persisted Qoder history after a real generated and resumed task
* Negotiate Qoder filters before searching an older execution host
* Combine search client imports for the CI plugin gate
* Keep the relay search oracle aligned with legacy agent filtering
* Register supervised Qoder China and Qwen lifecycle integration
* Cover Qoder China mobile assets and mixed-host resume gates
* Verify Qoder provider tags against the older released wire parser
* Verify China and Qwen keep independent Windows hook scripts
* Verify Qoder registrations against the installed older Windows release
* test(qoder): align search capability contracts and pin old-host fencing
* fix(qoder): rank exact picker identities and command aliases first
* test(qoder): preserve the regional CLI shared icon expectation
Keep the full bundled-asset and no-remote-image checks, with an explicit
shared-logo basename for Qoder China. The map also works with older
catalog type unions.
* fix(qoder): align China catalog entry with fallback order
---------
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
* Use the measured pnpm lookup policy automatically in hosted root CI (#24951)
* Select lookup automatically for the measured hosted root-install profile
* Record hosted automatic-mode cold cache publication proof
* fix(native-chat): keep OpenCode history usable at read limits
Continue past failed database probes while preserving discovery cancellation.
Verify rows displaced by a capped tail before deciding whether to replace history.
Represent oversized v1/v2 rows with the existing omission text and stable cursors.
* Continue Antigravity IDE and 2.0 history in new CLI conversations (#24692)
* feat(antigravity): bridge IDE history into new CLI conversations
* fix(antigravity): preserve fresh-launch model and environment for IDE references
* fix(antigravity): forward IDE history opt-in through desktop IPC
* fix(antigravity): rebuild remote IDE reference startup on its host
* fix(antigravity): register IDE continuation action labels
* fix(antigravity): confine IDE references and bound metadata reads
* fix(antigravity): localize IDE continuation badges
* Preserve scanner service cache assertions and refresh Antigravity opening metadata
* Preserve Antigravity opening joins and target folder runtime authority
* fix(opencode): retry timed-out SSH plugin updates (#24666)
Preserve bounded retry behavior and the current-main status-envelope fields.
Original-PR: #24124
Reviewed-source: 103144f9c5
Co-authored-by: Justas Brazauskas <brazauskasjustas@gmail.com>
* fix(opencode): keep Go credentials private and resolve backend keys (#24615)
Preserve the complete credential storage, migration, IPC, Settings and rate-limit refresh change alongside standalone GLM plans, current-main database diagnostics and the reviewed unknown-backend environment correction. Keep native discovery cancellation third and selected environment fourth.
Original-topic-commit: 7903f1cddb
Original-topic-commit: 588117b5cf
Original-topic-commit: dedd4f8c86
Original-topic-commit: 6645dae104
Original-topic-commit: a25b80c02c1af7830b0e6a65e72d965b3ad98276
Restacked-from: a25b80c02c1af7830b0e6a65e72d965b3ad98276
Restacked-onto: b032867021
Co-authored-by: kespineira <kespineira@users.noreply.github.com>
Co-authored-by: kevimux <kevimux@users.noreply.github.com>
Reported-by: pullfrog
Reviewed-full-source: 849fe093073f4c1606bd65d79a0c725d955d0d1f
Native-helper-source: 80dbe23237
Reviewed-full-current-source: 36acb57d44adb3d378c0289c8c15f7da0fda214c
---------
Co-authored-by: mrcha033 <mrcha033@users.noreply.github.com>
Co-authored-by: xodmd45-ctrl <xodmd45-ctrl@users.noreply.github.com>
Co-authored-by: Luchong <lu740528977@gmail.com>
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
Co-authored-by: Justas Brazauskas <brazauskasjustas@gmail.com>
* fix(accounts): protect registered UUID case variants during removal recovery
Compare validated account UUID identities without changing stored IDs or filesystem paths. Protect registered originals and quarantines, including explicit retry variants, and accept equivalent UUID spelling in a valid removal backup. Retain the complete account component and published contributor ancestry.
Private source only; no index, ref, or public mutation.
* Drain removal fixture jobs before resetting and deleting their records (#24977)
* Reuse measured Electron preparation for current Terminal Perf refs (#24968)
* feat(jcode): add Jcode as a supported TUI agent with managed hooks
Ports PR #10521 onto current main: agent catalog, managed hook service,
agent-status listener, session resume, AI Vault parser, per-pane daemon
isolation, and Source Control AI support.
Co-authored-by: Neil <neil@stably.ai>
* feat(jcode): evidence-backed status pipeline (turn_start, live pre_tool, questions)
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): register vault fixture, dynamic model discovery, script refresher
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): keep finished-turn detail so completion notifications fire
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): show the prompt in agent rows instead of jcode's repainting title
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): pre-warm the per-pane daemon so a cold runtime dir cannot time out
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): stop the OSC color skip from crashing every pane connect
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* refactor(jcode): fold three reverse-scan copies into one, reuse the shared hook POST
The jcode journal reader, the Claude transcript reader and the Command Code
transcript reader each carried their own copy of the same reverse chunked line
scan; they now share one tested helper. jcode's managed hook script drops its
hand-rolled curl for buildPosixAgentHookPostCommand, which also gains it the
raw-JSON transport and the --noproxy guard the bespoke copy was missing.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): narrow dynamic reads with predicates, pin the Windows hook shape
CI's anti-slop audit rejects Reflect.get: parse dynamic input into a named type
instead. Adds Windows script-shape tests too, since Windows is the platform this
change could not be exercised on directly.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* test(jcode): measure the gate against the synchronous path, not the clock
The absolute 1s bound was the flake CI shard 3/8 hit: it is tight enough to catch
a synchronous gate on an idle laptop and too tight on a loaded runner. Measuring
the same POST both ways on the same machine makes the claim a ratio, which is what
the test is actually about. Mutation-checked: a synchronous gate reads 6120ms
against a 1517ms observer.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* test(jcode): drop the wall-clock gate assertion
The bound was machine-speed sensitive and flaked on CI shard 3/8 at 1525ms. The
structural assertions (detached gate branch, foreground observer branch) and the
no-hang stdin drain cover the same contract without a timer.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): address CodeRabbit review on #22539
- Windows posted no payload at all: the shared builder reads `payload@-` from
stdin, which the gate has already drained and observer hooks never receive, so
every Windows event was dropped. Write the env var to a temp file and pipe it.
- The daemon pre-warm never fired for daemon-host spawns, which is the default
local path; it now runs there too, and from the final env so the daemon gets the
hook port and token.
- A failed runtime-dir mkdir took down every local terminal, jcode or not.
- removeJcodeManagedHooks matched the raw line, so a user hook whose comment
mentioned the managed script was deleted; matching on Windows never worked.
- A managed entry left by a copied home or a platform switch is now repointed
instead of being reported as user-owned forever.
- The OSC colour skip only checked launchAgent, so a command- or telemetry-named
jcode pane still leaked the reply into its composer.
- `['hooks']` and a commented scalar are recognised, instead of appending a
second [hooks] table that makes jcode reject the whole config.
- A failed tool's error is marked as tool output rather than agent prose.
- The vault keeps a session's stored name, counts its tokens, and skips
background_task and [Scheduled task] turns.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): read the journal once per turn, key prompts by byte offset
The jcode prompt reader ran a synchronous bounded file scan plus a JSON
parse on every hook event. jcode blocks on pre_tool, so a turn that ran
four tools charged the user eight scans of latency it did not need — the
prompt cannot change inside a turn.
- Cache the journal read per pane, refreshed on the turn boundary that
can change it. The cache holds the whole evidence record, since
hasExplicitUserPrompt needs the transcript-evidence flag and not just
the text, and it joins the existing pane-scoped lifecycle (close,
rename, reset) rather than living in a module singleton.
- Key a journal prompt by its absolute byte offset instead of its
region-local line index. The backward scan windows the file from EOF,
so appending shifted every boundary and reminted the key for a prompt
that never moved; a repeated turn_end then slipped past the same-hash
dedupe as a second done event with duplicate telemetry.
- Pin the platform in the daemon pre-warm tests. The pre-warm is a no-op
off POSIX, so the dedupe and retry cases would have passed vacuously
on a Windows runner.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): quote the managed hook path, drop tools from patch prompts
Three fixes, all on paths this PR could not exercise locally.
The managed hook command was stored as a bare path. jcode tokenizes that
string shell-style before exec'ing it directly (parse_hook_command,
crates/jcode-terminal-launch/src/lib.rs): unquoted whitespace splits, and
every unquoted backslash is consumed as an escape. So on Windows
`C:\Users\me\.orca\agent-hooks\jcode-hook.cmd` reached exec as
`C:Usersme.orcaagent-hooksjcode-hook.cmd` and no hook fired at all, and a
POSIX home with a space split into two arguments. Store the path
single-quoted (verbatim, backslashes included), falling back to double
quotes for a path containing a single quote. Existing bare entries are
already repointed by the stale-key path, and getStatus accepts both forms
so the repair is not reported as a user-owned hook. The quoting helper was
previously dead code that only tests called; the three production sites
now use it. isJcodeManagedCommand also normalizes separators, since a
`/`-only needle never matched a Windows entry.
Commit-message generation feeds a staged patch to `jcode run` as the
prompt — attacker-influenced text — while jcode's default profile exposes
shell, read, write, and MCP. Pass `--tool-profile none`, which resolves to
an empty allowed-tool set in jcode's config (base_allowed_tools), matching
the read-only posture claude (plan) and codex (read-only) already take.
docs/reference/jcode-hook-events.md was never actually in this PR: the
repo ignores docs/** and tracks reference docs by allow-list only, so the
captured-payload evidence four source comments point at was silently
dropped. Allow-list it.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* test(jcode): pin the tool-profile flag in the generation plan too
The argv assertion lives in two places; --tool-profile none only landed in
one, so the plan test still expected the unrestricted argv.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(text-generation): refuse an oversized argv prompt on Linux too
The pre-spawn size guard only ran on Windows. Linux caps a single argv
entry at MAX_ARG_STRLEN (32 pages, 128 KiB on a 4-KiB-page host) and
execve fails with E2BIG past it, so an agent that delivers the whole
prompt as one argument — jcode, and the other argv-delivery agents —
failed on a large staged diff with an error the user could not act on.
The cap is per-argument and in bytes, which is why it is not the Windows
line budget: 40k chars trips Windows and is nowhere near the Linux limit,
so folding them together would have refused prompts Linux runs fine.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): register in main's remote-installer guard, drop our duplicate
Rebasing onto 853 commits of main surfaced two things the earlier branch
had hidden.
main already owns a guard for the issue-#7253 bug class
(`remote-hook-service-registry-coverage.test.ts`). This branch had added a
second, near-identical one — a parallel implementation of a test that
already existed, which is what AGENTS.md's reuse rule is about. Deleted
ours and registered jcode in main's, which is the one that has kept pace
with every agent added since.
Also fixes a missing separator in the mobile icon map. `pnpm tc` does not
cover `mobile/`, so only the session-route closure suite caught it.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): decompose the four files jcode pushed over max-lines
Adding an agent tipped four modules past their line budget. AGENTS.md
forbids a `max-lines` disable or a per-file bump, so each is split on a
real seam rather than silenced:
- agent-catalog.tsx keeps `AgentIcon`, which 70+ files import, and the
rows move out. The rows alone exceed the 300-line budget a `.ts` file
gets, so they follow the primary/secondary split this repo already uses
for commit-message agent specs.
- getAgentResumeArgv -> agent-resume-argv.ts, re-exported so the 18 call
sites keep one import path.
- isDiscoverableSessionFile/pathSegments -> session-file-discovery.ts.
- remoteCodexSources -> remote-session-scanner-codex-sources.ts; Codex is
the one remote agent with two CODEX_HOME roots.
Also:
- Records the readiness-census baseline jcode now needs. main added that
gate while this branch was out; the fixture is the recorded 74-case
matrix, not a hand-written one.
- Restores two entries a rebase resolution silently dropped from
config/tsconfig.cli.json (gitlab/project-ref-parser,
startup/shell-path-probe). Nothing to do with jcode; losing them was a
conflict-resolution mistake on this branch.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* feat(native-chat): open structured chats on the paired Orca server that owns the workspace (#24205)
* fix(native-chat): a host admits structured sessions by client capability, not its own chat setting
A host's experimentalStructuredNativeChat decided whether any paired client could reach
agentSession.* at all, and whether session.tabs.* showed it structured tabs. That setting is the
host user's own launch preference: whether a new agent opens as a chat or a terminal is decided by
whoever launches it. Using it as admission control meant a client whose own preference was
"structured chat" was refused on a host whose preference was "terminal", and chats opened while
the setting was on were withheld from mobile once it was turned off.
The gate now asks one thing: did the client advertise agent-session.structured.v1 (in-process
callers negotiate nothing and are always admitted). Tab projection and restore follow the same
rule. With the setting no longer gating anything, the separate cleanup gate (close, cancel,
unsubscribe, release), which existed only so those kept working after the setting was switched
off, is identical to the main gate and is folded into it. The settings listener that republished
tabs when the setting changed is removed, since projection no longer depends on it.
The host setting still picks the default for launches that start on the host itself
(agent.launch from mobile, orchestration worker-start).
* fix(native-chat): the desktop declares structured chat support to paired hosts
The desktop renderer advertised agent-session.structured.v1 (and the Claude, turn-item and
background-task capabilities that go with it) to its own main process but not to a paired Orca
server. The server therefore refused every agentSession.* call from the desktop and stripped
structured chat tabs out of the tab list it published to it, so a structured chat running on a
paired server never appeared on the desktop, even though the renderer already mirrors a host's
agent-session tabs and drives each one against the server that owns its workspace.
The same renderer reads structured chats on either host, so the remote Electron list now carries
the same structured-session capabilities as the local one, and the capability test pins that
nothing is advertised only locally.
* feat(native-chat): open structured chats on the paired server that owns the workspace
With the structured-chat default on, an agent launched in a workspace that lives on a paired Orca
server always opened as a terminal (or the terminal-backed chat view). Three things kept it off
the structured path: the launch check refused every host but this machine, a remote workspace was
handed to the host-published terminal path before the structured route was even considered, and
the structured launch pipeline sent create and every follow-up call to this machine's runtime.
A workspace's owning runtime is fixed, so the pipeline now derives it from the workspace instead
of assuming this machine (structured-agent-session-owner.ts, the same derivation the chat pane
already uses to read a session). The launch intent carries that target; the pre-create support
check, create, the publication check and fence read, the launch prompt send, held option picks,
the "focus this chat" marker, the placeholder tab's host, and tab close/purge all use it.
The launch check now accepts a paired server and asks that server's own capabilities (read from
the status the client already cached for it) rather than this machine's. An SSH workspace stays
terminal-backed: no Orca runtime runs there. The host still answers createSupport before
anything is created, so an older server that refuses shows the failure in the chat tab.
A chat the user closed before its create landed is now also retired on the paired server when it
publishes, as the local sync already does. Orchestration workers placed on another runtime are
unchanged: federation creates terminal agents only.
* fix(native-chat): negotiate client-chosen launch mode so released phones and old servers keep terminals
Hosts advertise agent-session.structured.client-launch-mode.v1: they admit
structured sessions by client capability alone. A remote client that does
not advertise it (phones released before agent.launch) asks createSupport
to pick the launch mode, so the host keeps answering that with its own
setting, exactly as before. Cleanup methods keep their own named gate so a
future admission condition cannot make close or cancel refusable.
* refactor(runtime): keep the Electron client capability list in its own module
protocol-version.ts is at its line budget; the list is what the desktop
advertises to paired hosts, not the host's own contract.
* fix(native-chat): the desktop declares it picks each launch mode itself
Paired hosts and the desktop's own main process then answer createSupport
by the workspace rather than by their own chat setting.
* fix(native-chat): pin each structured chat to the host it was launched on
- Route: a paired server opens a chat only when it advertises the
client-chosen launch mode; an older server keeps its terminal. Its
capabilities come from the store's host status, not the compatibility
cache that is empty after boot or reconnect.
- A launch command override is this machine's: the route applies it only
locally, and a host's createSupport refuses on its own override.
- The owning host is resolved once, from the same value the route used,
and carried on the launch intent, its persisted record (legacy records
load as local), the provisional tab and every mirrored chat tab. Close,
purge, retry and reload read it instead of re-deriving it from a
worktree id two hosts can share; an owner that cannot be named refuses.
- Cancellation tombstones record their host: only that host's
authoritative inventory retires one, restored cleanup closes it there,
and a paired host's tombstone expires after 30 days if it never answers.
- A paired server's frame settles launches it published, as the local
inventory already does for this machine.
* fix(native-chat): a paired server that declines a chat opens its terminal instead
createSupport only reads, so both of its non-answers are settled before
anything is created:
- A paired server that answers it cannot run the chat (a WSL repo, a
Claude account mismatch, its own launch command override) closes the
chat tab and opens the terminal the route would have chosen, with a
notice saying why. This machine's own decline stays a failed chat.
- A host that could not be asked closes the chat tab and leaves one
failure toast, instead of a lingering "could not confirm" chat.
* test(native-chat): a provisional chat carries its launch's host and hands pre-create failures on
* chore(native-chat): justify the two type assertions this change's lines touch
* test(native-chat): state why each staged test fixture is cast
* fix(native-chat): chats that already exist keep showing whatever the chat setting says
The structured chat setting decides only what new agents open as. With it
off, this machine's structured chats used to be hidden while the host,
which no longer reads the setting, still reported them to the workspace
activation gate, so a workspace holding only a chat opened empty. The
local chat mirror and its startup restore now run whatever the setting
says, the continue-after-restart offer follows the chats that exist, and
the setting's copy says it applies to new agents.
* fix(native-chat): the browser client keeps its host terminal on paired servers
A browser client whose own preferences turn structured chat on took the
structured route for every paired-server workspace, but its handshake
never says it reads structured sessions, so the server refused the chat
and the user got a failed chat tab where a host terminal used to open.
The route for a paired host now also asks what this client advertises to
it: the desktop's list does, the browser client's does not. Its handshake
list is now a named constant the route reads, so the two cannot drift.
The chat setting's copy now says it runs on paired Orca servers too;
WSL and SSH hosts still use terminal chat.
* fix(native-chat): a retried launch a paired server declines opens its terminal too
A launch restored after a reload settles only through its Retry, so a
declining paired server left a failed chat there while a first launch got
the server's terminal and a notice. The chat's Retry now hands the same
pre-create failures to the same replacement, carrying the prompt the
launch had staged.
* refactor(native-chat): a paired host's cancelled-chat record ends on its 30-day TTL
The paired census re-read a host's whole inventory after every
authoritative frame to retire tombstones, and a tombstone restored after
a reload needed a second such frame, so in practice it retired nothing.
A tombstone guards a random session id and is inert once stale; the chat
is already closed on its host whenever a frame shows it. The census, its
trigger in the mirror layer and its cleanup are removed; the owner-scoped
tombstones, close-on-sight, the TTL and publication marking from frames
stay.
* fix(native-chat): a chat's pane and status read from the host recorded on its tab
The chat pane and its sidebar status still derived the host from the
workspace id, which two hosts can share; a paired chat in a non-active
same-id workspace was read from this machine. Both now read the owner
stamped on the tab, as close, purge and publication already do.
* fix(native-chat): "Resume in chat" follows the terminal resume's host rule
Agent Session History offered "Resume in chat" for a conversation
recorded on this machine into a paired server's workspace, where its
transcript does not exist. A chat now resumes a conversation only on the
host that recorded it, as the terminal resume does, and that host is the
one asked whether it can resume history.
* fix(native-chat): the chat setting says older paired servers keep terminal chat
* test(native-chat): pin that a host advertises the client-chosen launch mode
* fix(native-chat): mirror this machine's chats only where it holds them
Round 1 ran the local chat mirror for everyone so existing chats show
whatever the setting says. That gave every desktop a permanent
session-tabs listener, which turns on the runtime's phone replication
paths, plus two full session-tab censuses at startup, and made the
browser client mirror its remote host a second time.
The runtime now says whether it holds structured chats: its structured
host is built only when saved chats were restored at startup or a client
created one here, and it announces the moment one is built. The mirror,
the startup restore and the continue-after-restart offer run only when
the setting launches chats or the host holds some, and never in the
browser client. A chat a paired client creates here with the setting off
still appears at once. The chat behaviour settings show wherever chats
exist, and the setting's copy says it picks what new agents open as. The
toggle-off teardown this made dead is removed.
* test(native-chat): route a paired-server launch over the capability lists both sides really advertise
* test(native-chat): record install listeners without a cast
* fix(native-chat): a paired server admits a chat before any of it exists here
The desktop opened a paired server's chat tab, launch record, queued
prompt and focus intent before asking the server, so a "no" needed a
replacement that undid and redid all of it, and every piece it missed
was a bug: the workspace deselected, the caller told "failed" while a
terminal ran its prompt, the caller's arguments and other queued prompts
lost, and a create whose reply was lost treated as never sent.
A paired launch now asks the server first and commits nothing until it
answers. Admitted opens the chat as before. Declined runs the caller's
own launch as the server's terminal, with the existing notice (a resume
fails instead, having no terminal equivalent). Unreachable opens nothing
and names the server in one toast. The new-tab launcher reports the
host's surface for paired workspaces, as it did before paired chats,
with the prompt delivery of whichever surface got the prompt. The
replacement and its error classes are gone, and the probe inside a
launch is back to its old meaning: a "no" is a failed chat with Retry,
and no answer leaves "Could not confirm" with Retry and the prompt kept,
here as on this machine.
* fix(native-chat): mirror this machine's chats only once it holds one, not once its host is built
Session history, resume preparation, terminal resume commands and replay-safe phone launches all
build the structured host for users who never had a chat, which turned on the chat mirror and the
structured-only settings rows until the next restart. The signal is now derived from the host's
records (or a records file still owed its import) and pushed when the first chat is restored or
created. A throwing listener no longer fails the install that fired it.
* fix(native-chat): a fork's reveal never seeds a terminal beside the surface the launcher opens
Forking into a paired-server workspace revealed it as if nothing would open there, so the reveal
created a blank host terminal beside the forked chat (and beside a forked agent terminal on main).
The launcher always opens the fork's surface itself, so the reveal now says so for every surface,
as the fix-checks launch already does.
* fix(native-chat): a declined direct launch keeps the caller's CLI args; an unreachable resume toasts once
When a paired server declines a "Fix checks" chat in a new workspace, the terminal that opens
instead now carries the recipe's saved CLI arguments, launch platform and launch source, as the
terminal route did. "Resume in chat" to a server that cannot be reached showed the admission's
"Could not reach" toast and the vault's generic one; the admission marks its failure notified and
the vault adds nothing.
* fix(native-chat): a declined background create opens its terminal without switching workspaces
Since #23974 a worktree create the user moved away from must not pull them onto the new
workspace. When a paired server declined that create's chat, the fallback terminal opened as a new
agent tab, whose host create selects the workspace. The create now opens its own agent terminal the
way main's background branch does: in place from the request's startup plan (so its CLI args carry),
without selecting the workspace. A create the user is still watching keeps the new-tab fallback.
* test(native-chat): name the launch's host in main's new outbox fence test
Main's new staging-failure test calls settleStructuredAgentLaunchPrompt without the target this PR
made required; it is a local launch, as in the sibling tests.
* fix(native-chat): a paired server's new chat shows no model until the server reports the one it started
A chat on a paired server starts with the server's saved model and options, but the picker showed
this desktop's saved selection (or the catalog default) until the server reported a model, and a
pick made in that window was remembered on the server under that guessed model. A paired launch
now carries no desktop seed, and until the server reports its model the picker names no model and
takes no picks. Local chats are unchanged.
* test(native-chat): seed the paired repo without a cast
The repo literal already satisfies Repo, so the changed-lines cast gate has nothing to excuse.
* feat(native-chat): createSupport reports the saved selection a new chat on this host starts with
A chat on a paired server starts with the server's saved model and options, which the desktop could
not read, so its picker showed a guess. createSupport's answer, which the desktop already waits for
before a paired launch, now also carries that seed as a new optional field (older clients ignore it).
Create and createSupport read it through one resolver so they cannot drift.
* fix(native-chat): a paired server's new chat shows the selection the server will start it with
The paired server now names its saved model and options in the admission answer the desktop
already waits for. That seed goes into the launch intent and its persisted record, so the picker
shows the server's model at once, stays pickable like a local chat, and remembers picks on the
server under that model; a reload shows the same. The locked picker remains only for a server too
old to name a seed.
Also moves host admission and launch-outcome tracking into their own modules: the latest main
merge left structured-agent-session-launch.ts over the max-lines limit.
* test(native-chat): expect the launch intent's new seed argument in exact-call assertions
* refactor(protocol): move the Electron remote client capability list into its own module
Merging main left protocol-version.ts one line over the max-lines limit on this branch. The list of
capabilities the desktop advertises to a paired host moves, unchanged, into
electron-remote-runtime-client-capabilities.ts, the module the next PR in the stack already uses
for it; importers point there.
* fix(native-chat): a paired chat with no saved server model is pickable; Retry shows the server's current seed
A server whose user never saved a chat model sends no seed, and the desktop showed a locked,
model-only picker for it, although that is the common case: no server that can admit a paired chat
predates the seed field. Such a chat now behaves like a local chat with no saved model: the CLI
default, pickable. The lock and its snapshot helper are gone.
Retry kept the first admission's seed while the create probe, which already runs on every attempt,
reported the server's current one and dropped it. The probe's seed now replaces a paired launch's
seed and the picker's, so a retried chat shows what its create will run.
* test(cross-version): stub the launch seed resolver createSupport now reads
* test(protocol): pin the desktop capability divergence against what a paired server receives
Every paired transport sends the shared remote base plus the Electron list, so the
divergence test now compares that union with the renderer's local list instead of
the declared Electron list. A capability added only to the shared base can no
longer slip past it. The two base-only capabilities it surfaced are recorded:
skills.install-result.v2 has no local caller; the authoritative-inventory label is
read by the local tabs sync but dropped by main, and is marked unsettled.
The turn-item and both background-task-stop capabilities were already sent through
the shared base, so the Electron list no longer repeats them. The wire set is
unchanged; this PR's real change on the wire is structured.v1, the Claude
structured capability and the client launch-mode capability.
* fix(native-chat): the desktop tells its own host it picks each launch mode, so retrying an existing chat works with the setting off
* docs(native-chat): name the real exit for the released-phone createSupport rule
* fix(native-chat): the route reads the capabilities a paired host actually receives
The renderer decided whether a paired host would admit a chat from the desktop's Electron list, but
every desktop transport sends that list plus the shared remote base. They agreed only because the
route's checks happened to sit in both. The route input is now built with the same
remoteRuntimeClientCapabilities the transports use (the browser client already sends its list as is),
and a test pins each against the real handshake.
* test(cross-version): a released client still gets the host-setting createSupport answer; a launch-mode client gets supported plus the seed
* test(native-chat): let main's child-records test resolve each chat's owner
Main's new test mocks worktree-runtime-owner with only the runtime environment id, but the status
projection in this PR also resolves each structured chat's owner from the worktree. The mock keeps
the module's real exports and overrides only what the test pins.
* fix(deps): take #24204's lockfile that the merge reverted
* test(native-chat): let the Codex child-approval e2e unit test resolve each chat's owner
Its worktree-runtime-owner mock exported only the runtime environment id, but this PR's status
projection also resolves each structured chat's owner from the worktree. The mock now keeps the
module's real exports and overrides only that id, as structured-child-records-switch does.
* test(native-chat): move the close-race launch cases into their own file
Merging main added launch tests on both sides and took structured-agent-session-launch.test.ts past
the 800-line limit. The three cases where a tab close races a launch move to
structured-agent-session-launch-close-race.test.ts, with the same setup the other split launch
suites copy.
* feat(cli): add --unread/--read to orca worktree set
Pass validated read/unread flags through the existing workspace metadata update.
Contributor attribution correction: the earlier Ctrl+M squash message placed its related link after the co-author footer, so GitHub did not recognize all contributors. Preserve that credit here for https://github.com/stablyai/orca/pull/24100 and its related contribution https://github.com/stablyai/orca/pull/24143 .
Related CLI link validation retained from https://github.com/stablyai/orca/pull/20036 .
Co-authored-by: Raz Shlomo <12373339+razshlomo@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Zhichang Yu <yuzhichang@gmail.com>
Co-authored-by: Ahmed Nagy <ahmednagy25t@gmail.com>
* feat(cli): configure external worktree visibility per repo
Pass show/hide/inherit to the existing repository visibility update.
Related: https://github.com/stablyai/orca/pull/23025
Preserves CLI link validation from https://github.com/stablyai/orca/pull/20036 and read controls from https://github.com/stablyai/orca/pull/22714 .
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: KAPUIST <thsxornjs12@gmail.com>
* perf(git): reuse queries and stop canceled catalog scans (#24923)
* perf(git): reuse queries and stop canceled catalog scans
* refactor(git): check relay filesystem error codes
* Allow cold node-pty setup on Windows ARM CI
Keep a ten-minute bound only for node-pty rebuilt on Windows ARM CI hosts; all other node-gyp calls retain five minutes. Cold setup in two completed ARM jobs left compilation less than a minute.
* test(e2e): require exact Linux package tokens (#24995)
* Allow cold node-pty setup on Windows ARM CI (#24997)
Keep a ten-minute bound only for node-pty rebuilt on Windows ARM CI hosts; all other node-gyp calls retain five minutes. Cold setup in two completed ARM jobs left compilation less than a minute.
* fix(jcode): harden Windows hooks and negotiate remote history (#24998)
Redirect the managed Windows payload file into curl instead of starting
pipeline shells, register native Windows delivery coverage, and document
Jcode v0.89.0+ as the upstream launcher requirement for invisible hooks.
Negotiate Jcode history in both directions with mixed-version Orca hosts,
preserving supported search filters and old-client response compatibility.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
Co-authored-by: JianJia2018 <39438074+JianJia2018@users.noreply.github.com>
* fix(git): preserve SSH review context and worktree ownership (#24945)
* fix(git): preserve SSH arguments and guard background writes
* test(terminal): settle fish fixture startup readiness
* fix(git): preserve bare UNC SSH paths
* fix(git): unescape shell operators in Windows SSH paths
* test(runtime): settle removal writes before fixture cleanup
* test(git): skip optional OpenSSH probe when unavailable
* perf(git): skip equal-tip reads and bound relay discovery
* test(processes): ratchet the removed relay Git spawn
* test(shells): wait for initial zsh output before sending input
* Fix SSH review context and recover Git maintenance cleanup safely
* Keep relay Git compatibility fixtures outside shared client projects
* Fence superseded maintenance and preserve mixed-version session search
* test: model Git child termination and search catalogs
* test: retain catalog authority over history flags
* fix(opencode): keep shared-session TUI activity with its owning pane (#24962)
* fix(opencode): bind legacy attach status to its TUI pane
Adapt the official v1 TUI API to the existing pane reporter and learn structural session identity only after execution-host admission. Preserve known creators, gate new reports to capable hosts, and register the legacy TUI entry without changing user settings or overlay targets.
Credit werlang for the same-directory investigation in #22838 and #22847; no source from that PR is copied.
(cherry picked from commit d9eeb028a23a66b78b08538a5ebf6052499ec24f)
* fix(opencode): fence legacy TUI viewers before creator attribution
Use the execution host's existing admission policy on the physical TUI
pane and launch token before rewriting to a session creator or mutating
the launch-token cache. Preserve live viewers, frozen server stamps and
OpenCode 2 behavior.
Add HTTP regressions for retired panes, closed tabs, replaced tokens and
relay retirement, plus current-launch, alias and compatibility controls.
Follow-up to #22838 and @werlang's original #22847; preserves the original
intentional creator attribution without copying that PR's source.
* fix(opencode): preserve current plugin exports in legacy TUI wrapper
Keep the current server, setup, id and other default export entries when
adding the legacy TUI entry. Cover their preservation and seed the ACL
fixtures with the separate installed TUI/config bytes.
This completes the current-main port of the legacy identity replacement
for issue #22838, carrying @werlang's original PR #22847 contribution.
* fix(opencode): retry timed-out SSH plugin updates (#24666)
Preserve bounded retry behavior and the current-main status-envelope fields.
Original-PR: #24124
Reviewed-source: 103144f9c5
Co-authored-by: Justas Brazauskas <brazauskasjustas@gmail.com>
* fix(opencode): keep Go credentials private and resolve backend keys (#24615)
Preserve the complete credential storage, migration, IPC, Settings and rate-limit refresh change alongside standalone GLM plans, current-main database diagnostics and the reviewed unknown-backend environment correction. Keep native discovery cancellation third and selected environment fourth.
Original-topic-commit: 7903f1cddb
Original-topic-commit: 588117b5cf
Original-topic-commit: dedd4f8c86
Original-topic-commit: 6645dae104
Original-topic-commit: a25b80c02c1af7830b0e6a65e72d965b3ad98276
Restacked-from: a25b80c02c1af7830b0e6a65e72d965b3ad98276
Restacked-onto: b032867021
Co-authored-by: kespineira <kespineira@users.noreply.github.com>
Co-authored-by: kevimux <kevimux@users.noreply.github.com>
Reported-by: pullfrog
Reviewed-full-source: 849fe093073f4c1606bd65d79a0c725d955d0d1f
Native-helper-source: 80dbe23237
Reviewed-full-current-source: 36acb57d44adb3d378c0289c8c15f7da0fda214c
* Skip store notifications when refreshed work items are unchanged (#24530)
An identical forced refresh previously returned an empty Zustand patch, creating a new root state and notifying every subscriber. It now returns the current state after renewing the existing freshness timestamp. Real row or metadata changes still publish once.
* Read project activity once while sorting the sidebar (#24549)
Reuse the existing pure recent rank once per actual entry tuple within each synchronous sort; discard the lookup after sorting.
* Skip impossible link and tag matches in docs search results (#24646)
* Skip impossible HTML matches when formatting docs search results
After the existing link-removal phase, skip the unchanged HTML regex only when the current string has no closing delimiter.
* Skip impossible link and tag matches in docs search results
Skip the five unchanged link regexes when their protected-code input lacks ]; skip the unchanged HTML regex when its post-link input lacks >.
* Decode complete transcript lines without copying their bytes (#24546)
Decode each single owned Buffer synchronously; preserve concatenation for multipart records and remove a redundant tail array copy.
* Clear unused speech worker timers after errors and exit (#24924)
Call the existing idle timer cleanup when private current-worker state is cleared; keep current-worker guards, transcript/error order, deliberate warmth and stop deadline unchanged.
* Reuse documentation search excerpts during result navigation (#24574)
Memoize the existing pure excerpt renderer by complete raw text for the current result array, retaining original rendering as a miss fallback.
* Append subagent handoffs without recopying their message list (#24672)
Append each message reference to its private call-local scope list; retain the final merge and existing ordering.
* Cancel plugin retry timers when their fetch is replaced (#24700)
Reuse the existing timer clear block immediately after the fetch generation advances; preserve live retries and all state publications.
* Avoid rebuilding visible file paths for single-row selection (#24743)
Skip the visible path array only for keyboard replacement selection; the existing replacement branch never reads that order.
* Reuse prepared reads while searching saved agent conversations (#24535)
Use schema-complete session columns to enable the existing bounded statement cache without caching result rows.
* Reuse discovered mobile chunks during static traversal (#24774)
Iterate the existing live visited Set in FIFO discovery order instead of maintaining a second shifted queue; both actual callers own untouched esbuild JSON metadata.
* Skip unused image-size calculations in mobile web browser requests (#24941)
Use the existing mobile density budget only for mobile view; pass the existing constant to the existing assembler for web/default mode, where that argument is discarded. No cache, policy, request or native path changes.
* Skip diff analysis after a result has been discarded (#24711)
Move the unchanged pure render-limit calculation after both existing generation and section-token rejection fences.
* Skip late browser grab toasts after their surface closes (#24727)
Reuse the existing mounted-owner ref at the toast presentation entry while preserving all admitted extraction/screenshot and clipboard completion.
* Classify native chat waiting messages once (#24771)
Use the existing native filter traversal to populate a fresh waiting array, keeping the original predicate, policy, projection and output ordering while classifying each actual private slot once.
* Skip discarded Kanban pruning for an empty selection (#24782)
Guard only the existing open pruning branch when both its private selected Set and nullable anchor are empty; retain all original pre-guard projections and nonempty/anchor-only/public-helper behavior.
* Index discovered test files once during shard selection validation (#24532)
Verified selection plans previously validated each selected filename with a scan of all discovered files. One per-call Set now handles membership checks. Exact source/discovery checks, nonempty selection, fallback to all tests, shard balancing and manifests remain intact.
* Clear the watchdog benchmark heartbeat when CPU sampling fails (#24802)
Move the existing initial CPU observation and histogram enable inside the existing try/finally so their failures clear the sampler's owned heartbeat, preserving successful operation order and original errors.
* Reuse the parsed notification when leasing a push delivery (#24645)
* Reuse the parsed notification when leasing a push delivery
Pass the notification already parsed for the dismissal check into the existing private delivery builder.
* Reuse the parsed notification when leasing a push delivery
Pass the notification already parsed for the dismissal check into the existing private delivery builder.
* Update README downloads badge
* Let retired-cache GC observations settle across the existing six-turn budget (#24967)
* Collect test-selection evidence when full unit tests fail (#24955)
* Collect advisory unit-selection evidence from failed full runs
* Trigger checks after retargeting the evidence fix to main
* Check each project repository once while filtering mobile cards (#24540)
Reuse exact raw source/slug matching decisions within one project filter call, preserving the matcher, membership, negative matches, output order and identity.
* Skip renderer callbacks after their request has ended (#24631)
Reuse the existing pending-request identity check before starting its deferred renderer callback.
* Decode single-piece saved-session tails without copying them (#24632)
Decode the existing owned Buffer directly when the EOF tail has one piece; preserve the existing concatenation for multiple pieces.
* Avoid irrelevant scroll-action searches in Mac snapshots (#24731)
Check the current pure action name before searching for vertical scroll actions in the same immutable array.
* Stop copying every retained browser page before attachment (#24553)
Find the first page/generation match directly in the existing insertion-ordered Map instead of copying all page references into an array.
* Reuse event byte sizes when relay watchers notify several clients (#24558)
Store the existing grouped numeric byte-size array on the already call-local watcher batch sizing object, for reuse by later chunking clients.
* Stop building unused import candidates in the test planner (#24611)
Replace the existing ordered extension map/find with a loop that forms each candidate only when its existing lookup is reached.
* docs(terminal): explain local macOS/Linux shell startup files
Document the actual login/interactive shell startup-file order and existing shell setup behavior.
Co-authored-by: brynnclaw <261708852+brynnclaw@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(editor): recognize Ruby task and configuration files
Add Ruby task/configuration filenames to the existing generated language associations.
Co-authored-by: ggbdpq <ggbdpq@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* Speed up large Markdown Find and render oversized tables (#24948)
* Speed up large Markdown Find and render oversized tables
* Poll table preview geometry outside hidden renderers
* Preserve Markdown navigation through refresh and tab restoration
* Confirm Markdown restoration when refreshed content is ready
* Trace table refresh positions and update Unicode search reference
* Recognize queued measurement scrolls before restoring Markdown anchors
* Rebuild Markdown Find ranges after renderer components change
* fix(source-control): generate clean OpenCode messages locally and over SSH (#24613)
* fix(source-control): generate clean OpenCode answers on local and SSH hosts
Restack the original focused change onto current main, preserving every owned source and test blob and the merged CI contract and journal cleanup fixes.
Original-commit: 64bb15e3f2
fix(source-control): generate clean OpenCode answers on local and SSH hosts
Use configured models and JSON answer/error events, preserve run-first arguments, and handle the precise v2 variant rejection. Hydrate SSH execution-host PATH through the existing bounded login environment resolver before direct spawning.
Credits: andy-murr (PR #5197 SSH environment intent) and coelho-doti (PR #13065 argument-order intent).
Original-commit: 1b60ec5d11
fix(source-control): retry inline OpenCode model and variant options
Original-commit: 98fdecc7a3
Preserve OpenCode named errors without a data message
Restacked-from: 98fdecc7a3
Restacked-onto: f7b1f9d8be
* fix(ci): prevent concurrent pnpm refresh during mobile typechecks (#24776)
* fix(ci): run mobile typechecks without concurrent dependency refresh
* test(ci): check effective Linux E2E package list
* test(ci): preserve the mobile production compiler barrier
---------
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
* test(terminal): restore the live fish fixture prerequisites (#24947)
A restored pane waits for the initial status replay before subscribing to
PTY output. This fixture never settled that replay, so fish printed its
mode-2031 arm before the renderer connected. Its PTY API also omitted the
reset-input listener required by the serializer, aborting attachment.
Settle and dispose the existing startup-snapshot registration and provide
the same reset-listener mock used by the other PTY tests. The real fish
child-stdin assertions and timeouts remain unchanged. No production change.
* fix(shortcuts): defer TUI editing chords in terminal-first mode (#24640)
Restack the original focused change onto current main, preserving every owned source and test blob and the merged CI contract and journal cleanup fixes.
Original-commit: f707cde14a
fix(shortcuts): defer TUI editing chords in terminal-first mode
Original-commit: 0c6348e49e
docs(shortcuts): describe deferred preview terminal chords
Original-commit: be62c1b6c6
Align worktree history shortcut metadata with terminal conflict policy
Restacked-from: be62c1b6c6
Restacked-onto: f7b1f9d8be
* Register supervised Qoder China and Qwen Code (#24616)
* Add Qoder session history and search with real CLI coverage
* Allow the real Qoder marker file to end with a newline
* Keep Qoder tool output out of history previews and search
* Keep Qoder search pages readable by older clients
* Verify persisted Qoder history after a real generated and resumed task
* Negotiate Qoder filters before searching an older execution host
* Combine search client imports for the CI plugin gate
* Keep the relay search oracle aligned with legacy agent filtering
* Register supervised Qoder China and Qwen lifecycle integration
* Cover Qoder China mobile assets and mixed-host resume gates
* Verify Qoder provider tags against the older released wire parser
* Verify China and Qwen keep independent Windows hook scripts
* Verify Qoder registrations against the installed older Windows release
* test(qoder): align search capability contracts and pin old-host fencing
* fix(qoder): rank exact picker identities and command aliases first
* test(qoder): preserve the regional CLI shared icon expectation
Keep the full bundled-asset and no-remote-image checks, with an explicit
shared-logo basename for Qoder China. The map also works with older
catalog type unions.
* fix(qoder): align China catalog entry with fallback order
---------
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
* Use the measured pnpm lookup policy automatically in hosted root CI (#24951)
* Select lookup automatically for the measured hosted root-install profile
* Record hosted automatic-mode cold cache publication proof
* fix: bound remote generation setup and honor OpenCode option terminators
Count execution-host profile resolution inside the existing request deadline, cancel its waiter promptly, and pass only the remaining time to the child. Shared bounded profile probes keep their existing cache lifetime; the SSH transport margin is unchanged.
Read OpenCode output format from active final-argv options before -- so literal prompt arguments cannot select the JSON finalizer.
Fresh exact-source controls reproduce nine failures before; 139 related checks pass after, including primary/fallback delays, deadline boundaries, cancellation, parser metadata, and SSH lanes. Node typecheck and strict changed-file lint pass.
* Continue Antigravity IDE and 2.0 history in new CLI conversations (#24692)
* feat(antigravity): bridge IDE history into new CLI conversations
* fix(antigravity): preserve fresh-launch model and environment for IDE references
* fix(antigravity): forward IDE history opt-in through desktop IPC
* fix(antigravity): rebuild remote IDE reference startup on its host
* fix(antigravity): register IDE continuation action labels
* fix(antigravity): confine IDE references and bound metadata reads
* fix(antigravity): localize IDE continuation badges
* Preserve scanner service cache assertions and refresh Antigravity opening metadata
* Preserve Antigravity opening joins and target folder runtime authority
* fix(opencode): retry timed-out SSH plugin updates (#24666)
Preserve bounded retry behavior and the current-main status-envelope fields.
Original-PR: #24124
Reviewed-source: 103144f9c5
Co-authored-by: Justas Brazauskas <brazauskasjustas@gmail.com>
* fix(opencode): keep Go credentials private and resolve backend keys (#24615)
Preserve the complete credential storage, migration, IPC, Settings and rate-limit refresh change alongside standalone GLM plans, current-main database diagnostics and the reviewed unknown-backend environment correction. Keep native discovery cancellation third and selected environment fourth.
Original-topic-commit: 7903f1cddb
Original-topic-commit: 588117b5cf
Original-topic-commit: dedd4f8c86
Original-topic-commit: 6645dae104
Original-topic-commit: a25b80c02c1af7830b0e6a65e72d965b3ad98276
Restacked-from: a25b80c02c1af7830b0e6a65e72d965b3ad98276
Restacked-onto: b032867021
Co-authored-by: kespineira <kespineira@users.noreply.github.com>
Co-authored-by: kevimux <kevimux@users.noreply.github.com>
Reported-by: pullfrog
Reviewed-full-source: 849fe093073f4c1606bd65d79a0c725d955d0d1f
Native-helper-source: 80dbe23237
Reviewed-full-current-source: 36acb57d44adb3d378c0289c8c15f7da0fda214c
* fix(opencode): enforce deadline through executable startup
Keep synchronous Windows PATH resolution and child startup within the existing request budget.
---------
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
Co-authored-by: Justas Brazauskas <brazauskasjustas@gmail.com>
* fix(persistence): preserve projectGroupOrder on restart for flat folder-scan groups
Retain saved project ranks when the project remains in its folder-scan group.
Co-authored-by: lurunzi <lurunzi@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(runtime): reclaim temp files orphaned by interrupted mobile store writes
Reuse the existing stale temporary-file cleanup policy when each mobile store opens.
Co-authored-by: LDH1103 <ldh517525@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(terminal): show actionable copy for missing folder workspace paths
Show the existing folder-path failure with concrete recovery instructions.
Co-authored-by: Wayn_Liu <wayntingliu@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* docs: reference local orca.yaml and .worktreeinclude
Document the existing workspace configuration, include/copy rules and sharing behavior.
Co-authored-by: Neil <neil@stably.ai>
* fix(orchestration): allow model selection for OMP workers
Pass an explicitly requested model through the existing OMP worker launch catalog.
Co-authored-by: Neil <neil@stably.ai>
* fix(cli): preserve primitive success results in JSON output
Check for an object before inspecting screenshot fields so primitive success results remain printable.
Related: https://github.com/stablyai/orca/pull/14735
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: VXNCXNX <VXNCXNX@users.noreply.github.com>
* Show loaded file changes before deleting a workspace
Expose up to ten already loaded changed paths without altering deletion counts, hydration or authorization.
Related: https://github.com/stablyai/orca/pull/22778
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Michiel de Gooijer <mdgooijer@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(keybindings): record and match Option+digit shortcuts on macOS
Use the physical digit key for explicitly assigned Mac Option+digit shortcuts while retaining modifier checks.
Co-authored-by: marcuslannister <marcus@lannister.cc>
Co-authored-by: Neil <neil@stably.ai>
* fix(editor): don't throw closing a stale active file
Recover stale active-file selection within the owning workspace before choosing a surviving editor.
Related: https://github.com/stablyai/orca/pull/24715
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: Neil <neil@stably.ai>
* Fix Project Settings targeting for multiple local checkouts
Carry the selected checkout identity into the existing project Settings action.
Co-authored-by: fsmeier <1506919+fsmeier@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Neil <neil@stably.ai>
* feat(documents): open CSV and TSV files from the OS
Extend existing OS document associations and delivery to CSV/TSV, preserving restoration and authorization.
Co-authored-by: Neil <neil@stably.ai>
* feat(cli): link GitHub and GitLab items on worktree create and set
Expose and validate the existing workspace link fields, preserving omitted values and provider identity checks.
Related: https://github.com/stablyai/orca/pull/22609
Co-authored-by: marco song <marco.song@mvlchain.io>
Co-authored-by: SongMarco <20613630+SongMarco@users.noreply.github.com>
Co-authored-by: Luca Critelli <lucacri@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* feat(sidebar): include folder workspaces in keyboard navigation
Use the rendered sidebar row order and host identity when cycling through folder and Git workspaces.
Related: https://github.com/stablyai/orca/pull/10555
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: JeongUk Park <jeongph.dev@gmail.com>
* fix(build): turn off MSBuild file tracking for Windows native rebuilds
Default Windows native rebuilds to TrackFileAccess=false while preserving explicit caller preferences.
Co-authored-by: B1nh M1nh <43268322+b1nhm1nh@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Neil <neil@stably.ai>
* Fix Ctrl+M in Linux and Windows terminals
Keep the Minimize menu item but disable its accelerator registration on Linux and Windows.
Co-authored-by: Zhichang Yu <yuzhichang@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Ahmed Nagy <ahmednagy25t@gmail.com>
Related contribution: https://github.com/stablyai/orca/pull/24143
* fix(startup): read a nushell login PATH from $env.PATH
Use Nushell login command syntax and preserve the actual login PATH value.
Related: https://github.com/stablyai/orca/pull/22677
Co-authored-by: Kh05ifr4nD <meandSSH0219@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(cursor): run local hooks through sh for non-POSIX login shells
Keep POSIX hook syntax inside one quoted sh command so the login shell can invoke it safely.
Related: https://github.com/stablyai/orca/pull/22663
Co-authored-by: Kh05ifr4nD <meandSSH0219@gmail.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Neil <neil@stably.ai>
* Fix word wrap for both panes in side-by-side diffs
Forward wrapping to both diff panes through the existing editor option path and clean up listeners.
Co-authored-by: Wooseong Kim <innocarpe@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(monaco): highlight Svelte block closers inside markup
Return Svelte block closers to the existing markup tokenizer state.
Co-authored-by: Wooseong Kim <innocarpe@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Neil <neil@stably.ai>
* fix(repos): keep active clone dialog open on outside clicks
Ignore accidental outside dismissal only while cloning; Escape, Close and Back still cancel.
Related: https://github.com/stablyai/orca/pull/24581, https://github.com/stablyai/orca/pull/23430
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Nawapat Buakoet <nawapat.b@covest.finance>
* fix(editor): keep chat visible after closing Markdown tabs
Count remaining unified chat tabs before clearing workspace selection during editor close.
Related: https://github.com/stablyai/orca/pull/24273, https://github.com/stablyai/orca/pull/23760
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Wooseong Kim <innocarpe@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* Honor typed starting numbers in chat ordered lists
Forward the existing parsed ordered-list start attribute to both chat Markdown renderers.
Related: https://github.com/stablyai/orca/pull/19765
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Frederic Barthelemy <git@fbartho.com>
* Normalize base source once while checking changed-code diagnostics (#24557)
Reuse the exact existing moved-code matching algorithm with run-local lazy normalization of immutable base source blocks across diagnostics.
* Support real OpenCode sessions in native Chat (#24647)
* Use bounded OpenCode context for vault session continuation
OpenCode database and synthetic row paths are not text transcripts. Use the
vault preview or captured pane context, preserving actual transcript paths
containing a hash and supporting both OpenCode lanes and Windows paths.
Adapted the intent of #11859 and extended it to actual installed v2 vault rows.
Co-authored-by: mrcha033 <mrcha033@users.noreply.github.com>
* Read real OpenCode sessions in terminal-backed native Chat
Reuse the bounded AI Vault SQLite worker for v1 and v2 session pages and live updates. Keep terminal input as the real execution path and pace OpenCode Stop through its two-Escape interrupt.
Co-authored-by: xodmd45-ctrl <xodmd45-ctrl@users.noreply.github.com>
* fix(opencode): publish approval cards for permission requests
* Send OpenCode native approval through its Enter selector
* Resolve mobile Chat readability for folder workspaces
* Bound OpenCode part batches and preserve v2 image attachments
* Prefer live migrated OpenCode sessions over legacy copies
* Consolidate mobile Chat eligibility test imports
* Consolidate OpenCode SQLite protocol type imports
* Update native chat settings contract for both OpenCode agents
* fix(native-chat): reconcile bounded OpenCode transcript reads
* fix(native-chat): dispatch OpenCode questions safely
* fix(native-chat): keep native discovery and transcript windows current
* feat(accounts): link standalone GLM Coding Plans (#24618)
* feat(accounts): link standalone GLM Coding Plans
Adapt the reviewed GLM accounts contribution to current main, retain Antigravity behavior, guard late credential results, expose storage protection, and redact quota errors.
Co-authored-by: Luchong <lu740528977@gmail.com>
* fix(accounts): retain GLM credential results during quota refresh
* fix(accounts): make GLM credential editing desktop-only
* fix(accounts): mirror the host GLM site in paired clients
* fix(accounts): report unknown GLM host details and split web settings tests
Apply the independently reviewed Accounts correction from697284a without the v2 adapter commits. Preserve the saved-key store and serialized write behavior.
* fix(zcode): ship required GLM account translation entries
* chore: record GLM reconciliation hook validation
* chore: validate installed GLM commit hooks
* test: complete GLM account fixtures and web API inventory
---------
Co-authored-by: Luchong <lu740528977@gmail.com>
* fix(native-chat): route transcript requests through shared SQLite worker
* fix(ci): prevent concurrent pnpm refresh during mobile typechecks (#24776)
* fix(ci): run mobile typechecks without concurrent dependency refresh
* test(ci): check effective Linux E2E package list
* test(ci): preserve the mobile production compiler barrier
---------
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
* test(terminal): restore the live fish fixture prerequisites (#24947)
A restored pane waits for the initial status replay before subscribing to
PTY output. This fixture never settled that replay, so fish printed its
mode-2031 arm before the renderer connected. Its PTY API also omitted the
reset-input listener required by the serializer, aborting attachment.
Settle and dispose the existing startup-snapshot registration and provide
the same reset-listener mock used by the other PTY tests. The real fish
child-stdin assertions and timeouts remain unchanged. No production change.
* fix(shortcuts): defer TUI editing chords in terminal-first mode (#24640)
Restack the original focused change onto current main, preserving every owned source and test blob and the merged CI contract and journal cleanup fixes.
Original-commit: f707cde14a
fix(shortcuts): defer TUI editing chords in terminal-first mode
Original-commit: 0c6348e49e
docs(shortcuts): describe deferred preview terminal chords
Original-commit: be62c1b6c6
Align worktree history shortcut metadata with terminal conflict policy
Restacked-from: be62c1b6c6
Restacked-onto: f7b1f9d8be
* Register supervised Qoder China and Qwen Code (#24616)
* Add Qoder session history and search with real CLI coverage
* Allow the real Qoder marker file to end with a newline
* Keep Qoder tool output out of history previews and search
* Keep Qoder search pages readable by older clients
* Verify persisted Qoder history after a real generated and resumed task
* Negotiate Qoder filters before searching an older execution host
* Combine search client imports for the CI plugin gate
* Keep the relay search oracle aligned with legacy agent filtering
* Register supervised Qoder China and Qwen lifecycle integration
* Cover Qoder China mobile assets and mixed-host resume gates
* Verify Qoder provider tags against the older released wire parser
* Verify China and Qwen keep independent Windows hook scripts
* Verify Qoder registrations against the installed older Windows release
* test(qoder): align search capability contracts and pin old-host fencing
* fix(qoder): rank exact picker identities and command aliases first
* test(qoder): preserve the regional CLI shared icon expectation
Keep the full bundled-asset and no-remote-image checks, with an explicit
shared-logo basename for Qoder China. The map also works with older
catalog type unions.
* fix(qoder): align China catalog entry with fallback order
---------
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
* Use the measured pnpm lookup policy automatically in hosted root CI (#24951)
* Select lookup automatically for the measured hosted root-install profile
* Record hosted automatic-mode cold cache publication proof
* fix(native-chat): keep OpenCode history usable at read limits
Continue past failed database probes while preserving discovery cancellation.
Verify rows displaced by a capped tail before deciding whether to replace history.
Represent oversized v1/v2 rows with the existing omission text and stable cursors.
* Continue Antigravity IDE and 2.0 history in new CLI conversations (#24692)
* feat(antigravity): bridge IDE history into new CLI conversations
* fix(antigravity): preserve fresh-launch model and environment for IDE references
* fix(antigravity): forward IDE history opt-in through desktop IPC
* fix(antigravity): rebuild remote IDE reference startup on its host
* fix(antigravity): register IDE continuation action labels
* fix(antigravity): confine IDE references and bound metadata reads
* fix(antigravity): localize IDE continuation badges
* Preserve scanner service cache assertions and refresh Antigravity opening metadata
* Preserve Antigravity opening joins and target folder runtime authority
* fix(opencode): retry timed-out SSH plugin updates (#24666)
Preserve bounded retry behavior and the current-main status-envelope fields.
Original-PR: #24124
Reviewed-source: 103144f9c5
Co-authored-by: Justas Brazauskas <brazauskasjustas@gmail.com>
* fix(opencode): keep Go credentials private and resolve backend keys (#24615)
Preserve the complete credential storage, migration, IPC, Settings and rate-limit refresh change alongside standalone GLM plans, current-main database diagnostics and the reviewed unknown-backend environment correction. Keep native discovery cancellation third and selected environment fourth.
Original-topic-commit: 7903f1cddb
Original-topic-commit: 588117b5cf
Original-topic-commit: dedd4f8c86
Original-topic-commit: 6645dae104
Original-topic-commit: a25b80c02c1af7830b0e6a65e72d965b3ad98276
Restacked-from: a25b80c02c1af7830b0e6a65e72d965b3ad98276
Restacked-onto: b032867021
Co-authored-by: kespineira <kespineira@users.noreply.github.com>
Co-authored-by: kevimux <kevimux@users.noreply.github.com>
Reported-by: pullfrog
Reviewed-full-source: 849fe093073f4c1606bd65d79a0c725d955d0d1f
Native-helper-source: 80dbe23237
Reviewed-full-current-source: 36acb57d44adb3d378c0289c8c15f7da0fda214c
---------
Co-authored-by: mrcha033 <mrcha033@users.noreply.github.com>
Co-authored-by: xodmd45-ctrl <xodmd45-ctrl@users.noreply.github.com>
Co-authored-by: Luchong <lu740528977@gmail.com>
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
Co-authored-by: Justas Brazauskas <brazauskasjustas@gmail.com>
* Drain removal fixture jobs before resetting and deleting their records (#24977)
* Reuse measured Electron preparation for current Terminal Perf refs (#24968)
* feat(jcode): add Jcode as a supported TUI agent with managed hooks
Ports PR #10521 onto current main: agent catalog, managed hook service,
agent-status listener, session resume, AI Vault parser, per-pane daemon
isolation, and Source Control AI support.
Co-authored-by: Neil <neil@stably.ai>
* feat(jcode): evidence-backed status pipeline (turn_start, live pre_tool, questions)
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): register vault fixture, dynamic model discovery, script refresher
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): keep finished-turn detail so completion notifications fire
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): show the prompt in agent rows instead of jcode's repainting title
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): pre-warm the per-pane daemon so a cold runtime dir cannot time out
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): stop the OSC color skip from crashing every pane connect
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* refactor(jcode): fold three reverse-scan copies into one, reuse the shared hook POST
The jcode journal reader, the Claude transcript reader and the Command Code
transcript reader each carried their own copy of the same reverse chunked line
scan; they now share one tested helper. jcode's managed hook script drops its
hand-rolled curl for buildPosixAgentHookPostCommand, which also gains it the
raw-JSON transport and the --noproxy guard the bespoke copy was missing.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): narrow dynamic reads with predicates, pin the Windows hook shape
CI's anti-slop audit rejects Reflect.get: parse dynamic input into a named type
instead. Adds Windows script-shape tests too, since Windows is the platform this
change could not be exercised on directly.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* test(jcode): measure the gate against the synchronous path, not the clock
The absolute 1s bound was the flake CI shard 3/8 hit: it is tight enough to catch
a synchronous gate on an idle laptop and too tight on a loaded runner. Measuring
the same POST both ways on the same machine makes the claim a ratio, which is what
the test is actually about. Mutation-checked: a synchronous gate reads 6120ms
against a 1517ms observer.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* test(jcode): drop the wall-clock gate assertion
The bound was machine-speed sensitive and flaked on CI shard 3/8 at 1525ms. The
structural assertions (detached gate branch, foreground observer branch) and the
no-hang stdin drain cover the same contract without a timer.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): address CodeRabbit review on #22539
- Windows posted no payload at all: the shared builder reads `payload@-` from
stdin, which the gate has already drained and observer hooks never receive, so
every Windows event was dropped. Write the env var to a temp file and pipe it.
- The daemon pre-warm never fired for daemon-host spawns, which is the default
local path; it now runs there too, and from the final env so the daemon gets the
hook port and token.
- A failed runtime-dir mkdir took down every local terminal, jcode or not.
- removeJcodeManagedHooks matched the raw line, so a user hook whose comment
mentioned the managed script was deleted; matching on Windows never worked.
- A managed entry left by a copied home or a platform switch is now repointed
instead of being reported as user-owned forever.
- The OSC colour skip only checked launchAgent, so a command- or telemetry-named
jcode pane still leaked the reply into its composer.
- `['hooks']` and a commented scalar are recognised, instead of appending a
second [hooks] table that makes jcode reject the whole config.
- A failed tool's error is marked as tool output rather than agent prose.
- The vault keeps a session's stored name, counts its tokens, and skips
background_task and [Scheduled task] turns.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): read the journal once per turn, key prompts by byte offset
The jcode prompt reader ran a synchronous bounded file scan plus a JSON
parse on every hook event. jcode blocks on pre_tool, so a turn that ran
four tools charged the user eight scans of latency it did not need — the
prompt cannot change inside a turn.
- Cache the journal read per pane, refreshed on the turn boundary that
can change it. The cache holds the whole evidence record, since
hasExplicitUserPrompt needs the transcript-evidence flag and not just
the text, and it joins the existing pane-scoped lifecycle (close,
rename, reset) rather than living in a module singleton.
- Key a journal prompt by its absolute byte offset instead of its
region-local line index. The backward scan windows the file from EOF,
so appending shifted every boundary and reminted the key for a prompt
that never moved; a repeated turn_end then slipped past the same-hash
dedupe as a second done event with duplicate telemetry.
- Pin the platform in the daemon pre-warm tests. The pre-warm is a no-op
off POSIX, so the dedupe and retry cases would have passed vacuously
on a Windows runner.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): quote the managed hook path, drop tools from patch prompts
Three fixes, all on paths this PR could not exercise locally.
The managed hook command was stored as a bare path. jcode tokenizes that
string shell-style before exec'ing it directly (parse_hook_command,
crates/jcode-terminal-launch/src/lib.rs): unquoted whitespace splits, and
every unquoted backslash is consumed as an escape. So on Windows
`C:\Users\me\.orca\agent-hooks\jcode-hook.cmd` reached exec as
`C:Usersme.orcaagent-hooksjcode-hook.cmd` and no hook fired at all, and a
POSIX home with a space split into two arguments. Store the path
single-quoted (verbatim, backslashes included), falling back to double
quotes for a path containing a single quote. Existing bare entries are
already repointed by the stale-key path, and getStatus accepts both forms
so the repair is not reported as a user-owned hook. The quoting helper was
previously dead code that only tests called; the three production sites
now use it. isJcodeManagedCommand also normalizes separators, since a
`/`-only needle never matched a Windows entry.
Commit-message generation feeds a staged patch to `jcode run` as the
prompt — attacker-influenced text — while jcode's default profile exposes
shell, read, write, and MCP. Pass `--tool-profile none`, which resolves to
an empty allowed-tool set in jcode's config (base_allowed_tools), matching
the read-only posture claude (plan) and codex (read-only) already take.
docs/reference/jcode-hook-events.md was never actually in this PR: the
repo ignores docs/** and tracks reference docs by allow-list only, so the
captured-payload evidence four source comments point at was silently
dropped. Allow-list it.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* test(jcode): pin the tool-profile flag in the generation plan too
The argv assertion lives in two places; --tool-profile none only landed in
one, so the plan test still expected the unrestricted argv.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(text-generation): refuse an oversized argv prompt on Linux too
The pre-spawn size guard only ran on Windows. Linux caps a single argv
entry at MAX_ARG_STRLEN (32 pages, 128 KiB on a 4-KiB-page host) and
execve fails with E2BIG past it, so an agent that delivers the whole
prompt as one argument — jcode, and the other argv-delivery agents —
failed on a large staged diff with an error the user could not act on.
The cap is per-argument and in bytes, which is why it is not the Windows
line budget: 40k chars trips Windows and is nowhere near the Linux limit,
so folding them together would have refused prompts Linux runs fine.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): register in main's remote-installer guard, drop our duplicate
Rebasing onto 853 commits of main surfaced two things the earlier branch
had hidden.
main already owns a guard for the issue-#7253 bug class
(`remote-hook-service-registry-coverage.test.ts`). This branch had added a
second, near-identical one — a parallel implementation of a test that
already existed, which is what AGENTS.md's reuse rule is about. Deleted
ours and registered jcode in main's, which is the one that has kept pace
with every agent added since.
Also fixes a missing separator in the mobile icon map. `pnpm tc` does not
cover `mobile/`, so only the session-route closure suite caught it.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* fix(jcode): decompose the four files jcode pushed over max-lines
Adding an agent tipped four modules past their line budget. AGENTS.md
forbids a `max-lines` disable or a per-file bump, so each is split on a
real seam rather than silenced:
- agent-catalog.tsx keeps `AgentIcon`, which 70+ files import, and the
rows move out. The rows alone exceed the 300-line budget a `.ts` file
gets, so they follow the primary/secondary split this repo already uses
for commit-message agent specs.
- getAgentResumeArgv -> agent-resume-argv.ts, re-exported so the 18 call
sites keep one import path.
- isDiscoverableSessionFile/pathSegments -> session-file-discovery.ts.
- remoteCodexSources -> remote-session-scanner-codex-sources.ts; Codex is
the one remote agent with two CODEX_HOME roots.
Also:
- Records the readiness-census baseline jcode now needs. main added that
gate while this branch was out; the fixture is the recorded 74-case
matrix, not a hand-written one.
- Restores two entries a rebase resolution silently dropped from
config/tsconfig.cli.json (gitlab/project-ref-parser,
startup/shell-path-probe). Nothing to do with jcode; losing them was a
conflict-resolution mistake on this branch.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
* feat(native-chat): open structured chats on the paired Orca server that owns the workspace (#24205)
* fix(native-chat): a host admits structured sessions by client capability, not its own chat setting
A host's experimentalStructuredNativeChat decided whether any paired client could reach
agentSession.* at all, and whether session.tabs.* showed it structured tabs. That setting is the
host user's own launch preference: whether a new agent opens as a chat or a terminal is decided by
whoever launches it. Using it as admission control meant a client whose own preference was
"structured chat" was refused on a host whose preference was "terminal", and chats opened while
the setting was on were withheld from mobile once it was turned off.
The gate now asks one thing: did the client advertise agent-session.structured.v1 (in-process
callers negotiate nothing and are always admitted). Tab projection and restore follow the same
rule. With the setting no longer gating anything, the separate cleanup gate (close, cancel,
unsubscribe, release), which existed only so those kept working after the setting was switched
off, is identical to the main gate and is folded into it. The settings listener that republished
tabs when the setting changed is removed, since projection no longer depends on it.
The host setting still picks the default for launches that start on the host itself
(agent.launch from mobile, orchestration worker-start).
* fix(native-chat): the desktop declares structured chat support to paired hosts
The desktop renderer advertised agent-session.structured.v1 (and the Claude, turn-item and
background-task capabilities that go with it) to its own main process but not to a paired Orca
server. The server therefore refused every agentSession.* call from the desktop and stripped
structured chat tabs out of the tab list it published to it, so a structured chat running on a
paired server never appeared on the desktop, even though the renderer already mirrors a host's
agent-session tabs and drives each one against the server that owns its workspace.
The same renderer reads structured chats on either host, so the remote Electron list now carries
the same structured-session capabilities as the local one, and the capability test pins that
nothing is advertised only locally.
* feat(native-chat): open structured chats on the paired server that owns the workspace
With the structured-chat default on, an agent launched in a workspace that lives on a paired Orca
server always opened as a terminal (or the terminal-backed chat view). Three things kept it off
the structured path: the launch check refused every host but this machine, a remote workspace was
handed to the host-published terminal path before the structured route was even considered, and
the structured launch pipeline sent create and every follow-up call to this machine's runtime.
A workspace's owning runtime is fixed, so the pipeline now derives it from the workspace instead
of assuming this machine (structured-agent-session-owner.ts, the same derivation the chat pane
already uses to read a session). The launch intent carries that target; the pre-create support
check, create, the publication check and fence read, the launch prompt send, held option picks,
the "focus this chat" marker, the placeholder tab's host, and tab close/purge all use it.
The launch check now accepts a paired server and asks that server's own capabilities (read from
the status the client already cached for it) rather than this machine's. An SSH workspace stays
terminal-backed: no Orca runtime runs there. The host still answers createSupport before
anything is created, so an older server that refuses shows the failure in the chat tab.
A chat the user closed before its create landed is now also retired on the paired server when it
publishes, as the local sync already does. Orchestration workers placed on another runtime are
unchanged: federation creates terminal agents only.
* fix(native-chat): negotiate client-chosen launch mode so released phones and old servers keep terminals
Hosts advertise agent-session.structured.client-launch-mode.v1: they admit
structured sessions by client capability alone. A remote client that does
not advertise it (phones released before agent.launch) asks createSupport
to pick the launch mode, so the host keeps answering that with its own
setting, exactly as before. Cleanup methods keep their own named gate so a
future admission condition cannot make close or cancel refusable.
* refactor(runtime): keep the Electron client capability list in its own module
protocol-version.ts is at its line budget; the list is what the desktop
advertises to paired hosts, not the host's own contract.
* fix(native-chat): the desktop declares it picks each launch mode itself
Paired hosts and the desktop's own main process then answer createSupport
by the workspace rather than by their own chat setting.
* fix(native-chat): pin each structured chat to the host it was launched on
- Route: a paired server opens a chat only when it advertises the
client-chosen launch mode; an older server keeps its terminal. Its
capabilities come from the store's host status, not the compatibility
cache that is empty after boot or reconnect.
- A launch command override is this machine's: the route applies it only
locally, and a host's createSupport refuses on its own override.
- The owning host is resolved once, from the same value the route used,
and carried on the launch intent, its persisted record (legacy records
load as local), the provisional tab and every mirrored chat tab. Close,
purge, retry and reload read it instead of re-deriving it from a
worktree id two hosts can share; an owner that cannot be named refuses.
- Cancellation tombstones record their host: only that host's
authoritative inventory retires one, restored cleanup closes it there,
and a paired host's tombstone expires after 30 days if it never answers.
- A paired server's frame settles launches it published, as the local
inventory already does for this machine.
* fix(native-chat): a paired server that declines a chat opens its terminal instead
createSupport only reads, so both of its non-answers are settled before
anything is created:
- A paired server that answers it cannot run the chat (a WSL repo, a
Claude account mismatch, its own launch command override) closes the
chat tab and opens the terminal the route would have chosen, with a
notice saying why. This machine's own decline stays a failed chat.
- A host that could not be asked closes the chat tab and leaves one
failure toast, instead of a lingering "could not confirm" chat.
* test(native-chat): a provisional chat carries its launch's host and hands pre-create failures on
* chore(native-chat): justify the two type assertions this change's lines touch
* test(native-chat): state why each staged test fixture is cast
* fix(native-chat): chats that already exist keep showing whatever the chat setting says
The structured chat setting decides only what new agents open as. With it
off, this machine's structured chats used to be hidden while the host,
which no longer reads the setting, still reported them to the workspace
activation gate, so a workspace holding only a chat opened empty. The
local chat mirror and its startup restore now run whatever the setting
says, the continue-after-restart offer follows the chats that exist, and
the setting's copy says it applies to new agents.
* fix(native-chat): the browser client keeps its host terminal on paired servers
A browser client whose own preferences turn structured chat on took the
structured route for every paired-server workspace, but its handshake
never says it reads structured sessions, so the server refused the chat
and the user got a failed chat tab where a host terminal used to open.
The route for a paired host now also asks what this client advertises to
it: the desktop's list does, the browser client's does not. Its handshake
list is now a named constant the route reads, so the two cannot drift.
The chat setting's copy now says it runs on paired Orca servers too;
WSL and SSH hosts still use terminal chat.
* fix(native-chat): a retried launch a paired server declines opens its terminal too
A launch restored after a reload settles only through its Retry, so a
declining paired server left a failed chat there while a first launch got
the server's terminal and a notice. The chat's Retry now hands the same
pre-create failures to the same replacement, carrying the prompt the
launch had staged.
* refactor(native-chat): a paired host's cancelled-chat record ends on its 30-day TTL
The paired census re-read a host's whole inventory after every
authoritative frame to retire tombstones, and a tombstone restored after
a reload needed a second such frame, so in practice it retired nothing.
A tombstone guards a random session id and is inert once stale; the chat
is already closed on its host whenever a frame shows it. The census, its
trigger in the mirror layer and its cleanup are removed; the owner-scoped
tombstones, close-on-sight, the TTL and publication marking from frames
stay.
* fix(native-chat): a chat's pane and status read from the host recorded on its tab
The chat pane and its sidebar status still derived the host from the
workspace id, which two hosts can share; a paired chat in a non-active
same-id workspace was read from this machine. Both now read the owner
stamped on the tab, as close, purge and publication already do.
* fix(native-chat): "Resume in chat" follows the terminal resume's host rule
Agent Session History offered "Resume in chat" for a conversation
recorded on this machine into a paired server's workspace, where its
transcript does not exist. A chat now resumes a conversation only on the
host that recorded it, as the terminal resume does, and that host is the
one asked whether it can resume history.
* fix(native-chat): the chat setting says older paired servers keep terminal chat
* test(native-chat): pin that a host advertises the client-chosen launch mode
* fix(native-chat): mirror this machine's chats only where it holds them
Round 1 ran the local chat mirror for everyone so existing chats show
whatever the setting says. That gave every desktop a permanent
session-tabs listener, which turns on the runtime's phone replication
paths, plus two full session-tab censuses at startup, and made the
browser client mirror its remote host a second time.
The runtime now says whether it holds structured chats: its structured
host is built only when saved chats were restored at startup or a client
created one here, and it announces the moment one is built. The mirror,
the startup restore and the continue-after-restart offer run only when
the setting launches chats or the host holds some, and never in the
browser client. A chat a paired client creates here with the setting off
still appears at once. The chat behaviour settings show wherever chats
exist, and the setting's copy says it picks what new agents open as. The
toggle-off teardown this made dead is removed.
* test(native-chat): route a paired-server launch over the capability lists both sides really advertise
* test(native-chat): record install listeners without a cast
* fix(native-chat): a paired server admits a chat before any of it exists here
The desktop opened a paired server's chat tab, launch record, queued
prompt and focus intent before asking the server, so a "no" needed a
replacement that undid and redid all of it, and every piece it missed
was a bug: the workspace deselected, the caller told "failed" while a
terminal ran its prompt, the caller's arguments and other queued prompts
lost, and a create whose reply was lost treated as never sent.
A paired launch now asks the server first and commits nothing until it
answers. Admitted opens the chat as before. Declined runs the caller's
own launch as the server's terminal, with the existing notice (a resume
fails instead, having no terminal equivalent). Unreachable opens nothing
and names the server in one toast. The new-tab launcher reports the
host's surface for paired workspaces, as it did before paired chats,
with the prompt delivery of whichever surface got the prompt. The
replacement and its error classes are gone, and the probe inside a
launch is back to its old meaning: a "no" is a failed chat with Retry,
and no answer leaves "Could not confirm" with Retry and the prompt kept,
here as on this machine.
* fix(native-chat): mirror this machine's chats only once it holds one, not once its host is built
Session history, resume preparation, terminal resume commands and replay-safe phone launches all
build the structured host for users who never had a chat, which turned on the chat mirror and the
structured-only settings rows until the next restart. The signal is now derived from the host's
records (or a records file still owed its import) and pushed when the first chat is restored or
created. A throwing listener no longer fails the install that fired it.
* fix(native-chat): a fork's reveal never seeds a terminal beside the surface the launcher opens
Forking into a paired-server workspace revealed it as if nothing would open there, so the reveal
created a blank host terminal beside the forked chat (and beside a forked agent terminal on main).
The launcher always opens the fork's surface itself, so the reveal now says so for every surface,
as the fix-checks launch already does.
* fix(native-chat): a declined direct launch keeps the caller's CLI args; an unreachable resume toasts once
When a paired server declines a "Fix checks" chat in a new workspace, the terminal that opens
instead now carries the recipe's saved CLI arguments, launch platform and launch source, as the
terminal route did. "Resume in chat" to a server that cannot be reached showed the admission's
"Could not reach" toast and the vault's generic one; the admission marks its failure notified and
the vault adds nothing.
* fix(native-chat): a declined background create opens its terminal without switching workspaces
Since #23974 a worktree create the user moved away from must not pull them onto the new
workspace. When a paired server declined that create's chat, the fallback terminal opened as a new
agent tab, whose host create selects the workspace. The create now opens its own agent terminal the
way main's background branch does: in place from the request's startup plan (so its CLI args carry),
without selecting the workspace. A create the user is still watching keeps the new-tab fallback.
* test(native-chat): name the launch's host in main's new outbox fence test
Main's new staging-failure test calls settleStructuredAgentLaunchPrompt without the target this PR
made required; it is a local launch, as in the sibling tests.
* fix(native-chat): a paired server's new chat shows no model until the server reports the one it started
A chat on a paired server starts with the server's saved model and options, but the picker showed
this desktop's saved selection (or the catalog default) until the server reported a model, and a
pick made in that window was remembered on the server under that guessed model. A paired launch
now carries no desktop seed, and until the server reports its model the picker names no model and
takes no picks. Local chats are unchanged.
* test(native-chat): seed the paired repo without a cast
The repo literal already satisfies Repo, so the changed-lines cast gate has nothing to excuse.
* feat(native-chat): createSupport reports the saved selection a new chat on this host starts with
A chat on a paired server starts with the server's saved model and options, which the desktop could
not read, so its picker showed a guess. createSupport's answer, which the desktop already waits for
before a paired launch, now also carries that seed as a new optional field (older clients ignore it).
Create and createSupport read it through one resolver so they cannot drift.
* fix(native-chat): a paired server's new chat shows the selection the server will start it with
The paired server now names its saved model and options in the admission answer the desktop
already waits for. That seed goes into the launch intent and its persisted record, so the picker
shows the server's model at once, stays pickable like a local chat, and remembers picks on the
server under that model; a reload shows the same. The locked picker remains only for a server too
old to name a seed.
Also moves host admission and launch-outcome tracking into their own modules: the latest main
merge left structured-agent-session-launch.ts over the max-lines limit.
* test(native-chat): expect the launch intent's new seed argument in exact-call assertions
* refactor(protocol): move the Electron remote client capability list into its own module
Merging main left protocol-version.ts one line over the max-lines limit on this branch. The list of
capabilities the desktop advertises to a paired host moves, unchanged, into
electron-remote-runtime-client-capabilities.ts, the module the next PR in the stack already uses
for it; importers point there.
* fix(native-chat): a paired chat with no saved server model is pickable; Retry shows the server's current seed
A server whose user never saved a chat model sends no seed, and the desktop showed a locked,
model-only picker for it, although that is the common case: no server that can admit a paired chat
predates the seed field. Such a chat now behaves like a local chat with no saved model: the CLI
default, pickable. The lock and its snapshot helper are gone.
Retry kept the first admission's seed while the create probe, which already runs on every attempt,
reported the server's current one and dropped it. The probe's seed now replaces a paired launch's
seed and the picker's, so a retried chat shows what its create will run.
* test(cross-version): stub the launch seed resolver createSupport now reads
* test(protocol): pin the desktop capability divergence against what a paired server receives
Every paired transport sends the shared remote base plus the Electron list, so the
divergence test now compares that union with the renderer's local list instead of
the declared Electron list. A capability added only to the shared base can no
longer slip past it. The two base-only capabilities it surfaced are recorded:
skills.install-result.v2 has no local caller; the authoritative-inventory label is
read by the local tabs sync but dropped by main, and is marked unsettled.
The turn-item and both background-task-stop capabilities were already sent through
the shared base, so the Electron list no longer repeats them. The wire set is
unchanged; this PR's real change on the wire is structured.v1, the Claude
structured capability and the client launch-mode capability.
* fix(native-chat): the desktop tells its own host it picks each launch mode, so retrying an existing chat works with the setting off
* docs(native-chat): name the real exit for the released-phone createSupport rule
* fix(native-chat): the route reads the capabilities a paired host actually receives
The renderer decided whether a paired host would admit a chat from the desktop's Electron list, but
every desktop transport sends that list plus the shared remote base. They agreed only because the
route's checks happened to sit in both. The route input is now built with the same
remoteRuntimeClientCapabilities the transports use (the browser client already sends its list as is),
and a test pins each against the real handshake.
* test(cross-version): a released client still gets the host-setting createSupport answer; a launch-mode client gets supported plus the seed
* test(native-chat): let main's child-records test resolve each chat's owner
Main's new test mocks worktree-runtime-owner with only the runtime environment id, but the status
projection in this PR also resolves each structured chat's owner from the worktree. The mock keeps
the module's real exports and overrides only what the test pins.
* fix(deps): take #24204's lockfile that the merge reverted
* test(native-chat): let the Codex child-approval e2e unit test resolve each chat's owner
Its worktree-runtime-owner mock exported only the runtime environment id, but this PR's status
projection also resolves each structured chat's owner from the worktree. The mock now keeps the
module's real exports and overrides only that id, as structured-child-records-switch does.
* test(native-chat): move the close-race launch cases into their own file
Merging main added launch tests on both sides and took structured-agent-session-launch.test.ts past
the 800-line limit. The three cases where a tab close races a launch move to
structured-agent-session-launch-close-race.test.ts, with the same setup the other split launch
suites copy.
* feat(cli): add --unread/--read to orca worktree set
Pass validated read/unread flags through the existing workspace metadata update.
Contributor attribution correction: the earlier Ctrl+M squash message placed its related link after the co-author footer, so GitHub did not recognize all contributors. Preserve that credit here for https://github.com/stablyai/orca/pull/24100 and its related contribution https://github.com/stablyai/orca/pull/24143 .
Related CLI link validation retained from https://github.com/stablyai/orca/pull/20036 .
Co-authored-by: Raz Shlomo <1237333…
* fix(ci): wait for the fragment navigation policy report (#25017)
---------
Co-authored-by: Justas Brazauskas <brazauskasjustas@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Brynn Bendixen <brynnbendixen@gmail.com>
Co-authored-by: brynnclaw <261708852+brynnclaw@users.noreply.github.com>
Co-authored-by: 小七 <ggbdpq@gmail.com>
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
Co-authored-by: lurunzi <lurunzi@gmail.com>
Co-authored-by: Dongho Lee <126236161+LDH1103@users.noreply.github.com>
Co-authored-by: LDH1103 <ldh517525@gmail.com>
Co-authored-by: Wayn_Liu <115852642+Waynting@users.noreply.github.com>
Co-authored-by: Wayn_Liu <wayntingliu@gmail.com>
Co-authored-by: VXNCXNX <VXNCXNX@users.noreply.github.com>
Co-authored-by: Michiel de Gooijer <mdgooijer@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Marcus <70784566+marcuslannister@users.noreply.github.com>
Co-authored-by: marcuslannister <marcus@lannister.cc>
Co-authored-by: OrcaWin <alpha-eng@stably.ai>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: Florian Schwarzmeier <1506919+fsmeier@users.noreply.github.com>
Co-authored-by: SongMarco <snubflow@gmail.com>
Co-authored-by: marco song <marco.song@mvlchain.io>
Co-authored-by: SongMarco <20613630+SongMarco@users.noreply.github.com>
Co-authored-by: Luca Critelli <lucacri@gmail.com>
Co-authored-by: JeongUk Park <jeongph.dev@gmail.com>
Co-authored-by: B1nh M1nh <43268322+b1nhm1nh@users.noreply.github.com>
Co-authored-by: Zhichang Yu <yuzhichang@gmail.com>
Co-authored-by: Kh05ifr4nD <meandSSH0219@gmail.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Wooseong Kim <2222333+innocarpe@users.noreply.github.com>
Co-authored-by: Wooseong Kim <innocarpe@gmail.com>
Co-authored-by: Nawapat Buakoet <nawapat.b@covest.finance>
Co-authored-by: Frederic Barthelemy <git@fbartho.com>
Co-authored-by: mrcha033 <mrcha033@users.noreply.github.com>
Co-authored-by: xodmd45-ctrl <xodmd45-ctrl@users.noreply.github.com>
Co-authored-by: Luchong <lu740528977@gmail.com>
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
Co-authored-by: razshlomo <razshlomo@users.noreply.github.com>
Co-authored-by: Raz Shlomo <12373339+razshlomo@users.noreply.github.com>
Co-authored-by: Ahmed Nagy <ahmednagy25t@gmail.com>
Co-authored-by: KAPUIST <thsxornjs12@gmail.com>
Co-authored-by: JianJia2018 <39438074+JianJia2018@users.noreply.github.com>
Redirect the managed Windows payload file into curl instead of starting
pipeline shells, register native Windows delivery coverage, and document
Jcode v0.89.0+ as the upstream launcher requirement for invisible hooks.
Negotiate Jcode history in both directions with mixed-version Orca hosts,
preserving supported search filters and old-client response compatibility.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
Co-authored-by: JianJia2018 <39438074+JianJia2018@users.noreply.github.com>
Three fixes, all on paths this PR could not exercise locally.
The managed hook command was stored as a bare path. jcode tokenizes that
string shell-style before exec'ing it directly (parse_hook_command,
crates/jcode-terminal-launch/src/lib.rs): unquoted whitespace splits, and
every unquoted backslash is consumed as an escape. So on Windows
`C:\Users\me\.orca\agent-hooks\jcode-hook.cmd` reached exec as
`C:Usersme.orcaagent-hooksjcode-hook.cmd` and no hook fired at all, and a
POSIX home with a space split into two arguments. Store the path
single-quoted (verbatim, backslashes included), falling back to double
quotes for a path containing a single quote. Existing bare entries are
already repointed by the stale-key path, and getStatus accepts both forms
so the repair is not reported as a user-owned hook. The quoting helper was
previously dead code that only tests called; the three production sites
now use it. isJcodeManagedCommand also normalizes separators, since a
`/`-only needle never matched a Windows entry.
Commit-message generation feeds a staged patch to `jcode run` as the
prompt — attacker-influenced text — while jcode's default profile exposes
shell, read, write, and MCP. Pass `--tool-profile none`, which resolves to
an empty allowed-tool set in jcode's config (base_allowed_tools), matching
the read-only posture claude (plan) and codex (read-only) already take.
docs/reference/jcode-hook-events.md was never actually in this PR: the
repo ignores docs/** and tracks reference docs by allow-list only, so the
captured-payload evidence four source comments point at was silently
dropped. Allow-list it.
Co-authored-by: czzczz <chanzrz_zbf@foxmail.com>
Add Ruby task/configuration filenames to the existing generated language associations.
Co-authored-by: ggbdpq <ggbdpq@gmail.com>
Co-authored-by: Neil <neil@stably.ai>
* Let measured cache producers keep stores without downloading them
* Check that restore-only callers do not publish a producer path
* Enable the measured producer mode and record hosted comparisons
* ci: defer headless dependency installation until graph analysis is needed
* docs: align headless CI rollout with platform and cache policy
* test: isolate headless detector output from the parent CI step
A headless orca serve has no window to write the Codex ready title, so a
Codex tui-idle wait settled only after three quiet seconds. Rule files gain
profile.hooks: "turn-end": a fresh hook done settles the wait, while a
working or permission row leaves the decision to the rules, so a Codex
whose Esc posts no event (before its Interrupt hook) cannot hang the wait.
Codex moves from identity-only to turn-end.
* Let scheduled CI warmers wait and measure WebRTC startup
* Measure a smaller daemon shutdown fixture image
* Counterbalance WebRTC startup and verify retained fixture files
* Record CI fixture measurements and remove temporary pilots
* Clarify fixture build dependency cleanup evidence
* Make coalesced snapshot fixture delivery deterministic
* test: type the PTY write delay observer
* ci: avoid unrelated headless server qualification
* ci: skip headless detection for ineligible draft PRs
* ci: preserve cross-host qualification and skip supplied prerequisites
* ci: include Windows server cache validation in change detection
* fix(worktrees): remove repeated scans and keep prepared checkouts fresh
* fix(worktrees): reclaim unlocked fallback preparations safely
* refactor(worktrees): simplify creation ownership and idle maintenance
* fix(git): keep ref maintenance armed after an index-only pass
An idle attempt that found the pack index due but refs still cooling down
returned without rescheduling, so loose refs from the arming fetch waited
for the next write instead of the ref cooldown.
* fix(codex): install Codex's Interrupt hook so an Esc-cancelled turn settles
Codex 0.150+ fires an Interrupt hook when the user presses Esc on an
approval prompt or mid-tool, and nothing else. Orca did not install it, so
the pane stayed blocked/working until the next prompt.
- Add Interrupt to the managed Codex events and label maps, written with
Codex's 3s cap (a larger value triggers a startup clamp warning).
- Hash the timeout Codex hashes (Interrupt is clamped to [1,3], default 1)
so self-computed trust matches Codex; pinned against a real 0.159.3 hash.
- Map a root Interrupt to the existing cancelled-turn record
(markCodexLeadTurnInterrupted), keeping child work in the fold; a
child-scoped Interrupt is ignored. Relayed rows take the same path.
* test(runtime): add a readiness census pinning every tui-idle verdict
Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.
Refs STA-9098
* test(runtime): pin the census quiet probes to literal windows
A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.
Refs STA-9098
* test(runtime): say which census probe writes runtime state
Refs STA-9098
* refactor(codex): let the hook builder own Codex's per-event timeout
The managed hook's timeout is now Codex's own normalization of the shared
budget, and every installer derives its trust entry from the hook it wrote,
so no installer repeats the Interrupt special case.
Claude-Session: codex-interrupt-hook review
* refactor(codex): route Interrupt through the Stop lead update with an outcome
Interrupt now writes the lead record through the same setCodexMainAgentTurnState
call as Stop, so markCodexLeadTurnInterrupted keeps its original signature.
Drops the child-scoped Interrupt guard: Codex never runs Interrupt hooks for
subagents and its input schema has no agent_id.
Claude-Session: codex-interrupt-hook review
* test(runtime): observe the census through settled panes and caller-visible waits
- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.
* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix
Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.
* test(runtime): read the census baseline field without Reflect.get
The anti-slop lint rejects Reflect.get on parsed input.
* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files
Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.
The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes
Refs STA-9098
* fix(runtime): refuse rule patterns that repeat an optional or alternating group
The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.
* refactor(runtime): give agent state rules and text anchors one when/answer shape
Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.
- Cursor's prompt is two anchors answering working and idle; the one-off
workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.
* docs: point the readiness evidence docs at the agent state rule files
* refactor(runtime): read Codex, Claude, OpenCode, Pi, OMP and Gemini readiness from rule files
The rule engine gains the regions and answers these agents need, as closed-list entries:
- rule regions `title` (the classified title status) and `text` (one of the file's idle text
anchors, settled), and a `predicate` form of the screen region for named engine scans;
- `withoutClock: skip` for strong quiet rules a clockless pane must not believe;
- anchors (renamed from textAnchors) gain a `title` region, and `live` and `hold` answers;
- `profile.screenSource` (trusted grid or live screen), and an `unknown-pane` file for panes
with no known agent.
Codex's header, composer and provisional-startup checks become named predicates referenced
from codex.json; its ready header, header and startup hold become shared text anchors. Native
idle title markers become shared title anchors; name-only title handling becomes each agent's
idle-title rule. The agent-specific branches in terminal-wait-detection.ts and
tui-idle-evidence.ts are deleted, and the "later live prompt cancels a blocker" rule now reads
only rule-file anchors (plus Muse, which moves in part b2).
No behaviour change: the readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the rule engine's title, text and predicate regions and the bundled anchors
Refs STA-9098
* fix(runtime): reject a rule file that repeats an anchor or rule id
A text rule names its anchor by id, so a repeated id let a file pass validation and then throw
while compiling. Also states that engineVersion bumps once a version ships; version 1 is still
being defined.
* refactor(runtime): fold the working anchor answer into live
The engine treated an anchor's working and live answers identically: both mark a live prompt
that cancels an earlier blocker and settles nothing. Cursor's busy prompt now answers live, so
anchors have one non-settling prompt answer.
Refs STA-9098
* refactor(runtime): read the shared π title anchor from pi.json alone
Pi and OMP paint the same `π - <session>` rest title, and title anchors apply to every pane,
so one copy covers both.
Refs STA-9098
* refactor(runtime): key every rule file and read the trusted screen from screenSource alone
readsTrustedScreen no longer also asks for a screen rule (every trusted file has one, and the
schema requires screenSource where it matters), so rule-less files need no filter. A rule's
match is a plain boolean, and compileTitleAnchors is module-private.
Refs STA-9098
* test(runtime): pin that a clocked Codex pane takes no other agent's ready text
No test failed when holdsReadyTextToQuiet was removed; this one does.
Refs STA-9098
* fix(agent-hooks): keep an OMP approval wait until omp resolves it
omp posts tool_execution_start a few milliseconds after
tool_approval_requested, while its Approve/Deny select still holds the
human. Both mapped onto the pane row, so the working event overwrote the
blocked one and the pane read as busy for the whole prompt.
A working event now leaves an OMP approval wait in place; only
tool_approval_resolved or a new turn ends it. An ask row is unchanged:
its own tool_execution_end ends it. The test replays the order a live
omp 17 run posted for a denied bash call.
Refs STA-9100
* refactor(runtime): select the fresh hook row on any of a terminal's handles or pane keys
selectFreshExplicitAgentStatus matched one handle and one pane key and
returned only the mapped status. The row selection now takes sets of
handles and pane keys, an optional received-at floor, and returns the
row itself, so a reader can see the main agent's own state. The old
function keeps its signature and result on top of it.
Refs STA-9100
* feat(runtime): let tui-idle read hook state for agents whose hooks cover the whole turn
tui-idle read no hook state. Hook state reached readiness only through
the `<Agent> ready` titles the window writes, so a headless `orca serve`
never saw it (#16095), and Codex settled only once its screen had been
quiet for three seconds.
Rule files gain `profile.hooks: "authoritative" | "identity-only"`,
defaulting to identity-only. Codex (with its Interrupt hook), OpenCode,
OpenCode 2, Pi and OMP are authoritative. For them a fresh hook-store
row decides ahead of every other lane:
- the main agent's turn decides (`mainAgent.state` when published), so a
subagent's Stop does not end the lead turn: done settles strong,
working holds, a permission wait never settles;
- the tail's blocked text goes through the existing permission arbiter
with the turn as its explicit status, so a denied prompt's dialog left
in the tail no longer blocks a turn the hook says ended;
- the row joins on every pane key and terminal handle the PTY owns.
No row, a stale, restored or other agent's row, a session-start done,
and a row from before a PTY respawn all fall back to today's lanes. That
keeps startup on the screen and text rules: Codex posts SessionStart
only with the first prompt. Claude, Cursor, Gemini and the rest stay
identity-only.
The readiness census has no hook server, so its frames are unchanged.
Refs STA-9100
* docs(agent-status): record readiness as a reader of the hook store
Refs STA-9100
* fix(runtime): ignore a hook done older than the latest input Orca wrote
A finished turn leaves a fresh `done` row. A caller that sends the next
prompt and waits at once could settle on it before the new turn's first
hook arrives, so the wait returned while the agent was starting work.
Orca's own input writes (terminal send, agent prompts, mailbox pointers)
now stamp a per-PTY input clock, and the hook lane reads no `done`
received before it; the pane falls back to the screen and text rules
until the agent reports again. A `working` row is unaffected.
Refs STA-9100
* docs(agent-status): note the input floor on the hook lane's done
Refs STA-9100
* fix(runtime): take the hook lane's input floor from the PTY run's input record
The hook lane ignored a done older than Orca's latest write to the pane, kept in a
new per-PTY map stamped by a wrapper threaded through four write sites. The PTY
run register already sits on both write funnels, so it now records the last
input (launch writes included, terminal replies not) and the lane reads it.
Keys the user types now count too, which closes the restart-in-the-same-shell
gap: typing `codex` to relaunch no longer lets the previous process's done read
ready while the new one boots.
The respawn floor moves from the shared row join into the lane, beside the
input floor; the freshest row predates a floor exactly when every row does.
* test(runtime): drop runtime hook-lane cases the unit suite already proves
Working over a ready title, a permission wait, and an identity-only agent are
decided inside evaluateTuiIdle and covered there; the runtime suite keeps the
wiring: the join, both floors, Pi's own OSC 133 markers and the arbiter.
* fix(runtime): record a PTY's last input even when main adopted it without a spawn commit
A materialized pane re-adopted by the renderer returns before the spawn-commit
site, so it had no run record and its input never moved the hook lane's floor.
The last input now lives beside the run records: any PTY's input counts, and a
new process's commit still clears it.
* fix(runtime): keep a running process's input time when main reattaches or adopts it
A reattach or adoption commit without an incarnation id cleared the PTY's
last-input time, so a prompt sent just before an SSH adoption was forgotten
and the hook lane could accept the previous turn's done as ready. Only a new
process (or a reattach naming a different incarnation) now starts clean; the
first-input fact follows the same rule.
* docs(runtime): say why a lead turn that ended reads ready while a subagent runs
* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail
A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.
* refactor(runtime): state Codex's provisional startup and title anchors as plain rules
The provisional-startup hold becomes a lastOf anchor with an all/none test, so
its TypeScript scan goes. Title anchors drop their status field (every caller
already gates on an idle title), and withoutClock keeps only the value a rule
can set.
* fix(runtime): leave Codex readiness to its title and screen rules
Codex before its Interrupt hook posts nothing for an Esc mid-turn, so its hook
row stays working and a hook-authoritative tui-idle wait hangs until the row
goes stale. Current Codex already settles fast through its ready title.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* Reuse serializer oracle cells and isolate native cache policy
* Preserve native cache post-save paths and record hosted oracle gain
* Record native cache reuse and separate cancel-test startup budget
* test(runtime): add a readiness census pinning every tui-idle verdict
Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.
Refs STA-9098
* test(runtime): pin the census quiet probes to literal windows
A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.
Refs STA-9098
* test(runtime): say which census probe writes runtime state
Refs STA-9098
* test(runtime): observe the census through settled panes and caller-visible waits
- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.
* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix
Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.
* test(runtime): read the census baseline field without Reflect.get
The anti-slop lint rejects Reflect.get on parsed input.
* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files
Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.
The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes
Refs STA-9098
* fix(runtime): refuse rule patterns that repeat an optional or alternating group
The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.
* refactor(runtime): give agent state rules and text anchors one when/answer shape
Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.
- Cursor's prompt is two anchors answering working and idle; the one-off
workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.
* docs: point the readiness evidence docs at the agent state rule files
* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail
A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.
Clients stop an orcad that advertises health.stopRequests through its slot-local request file and keep SIGTERM for older builds. Decommission runs through the activation journal and fence: it refuses while the terminal census is live or uncounted, stops the instance with an instance-bound managed request, cancels a stop orcad never acted on, and deactivates the record only on proven exit. orcad gains --cancel-managed-stop and an exclusive per-transaction decision file so a cancel can never race a dispatched stop. POSIX-only and inert: no production caller.
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
orcad stops through slot-local and instance-bound request files, so a reused PID is never signalled. A managed stop is proven by its completion command and recorded as a receipt. Optional daemon retirement is best effort: an idle daemon retires, while a busy or unverifiable one stays up with its admission fence released. Runtime teardown runs in reverse order and keeps the instance lock and profile admission when any writer fails to stop. Browser discovery no longer delays readiness. Legacy worker recovery and watcher children are drained before the final flush. Headless terminal close no longer waits on a renderer tab that does not exist. No production deployment.
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
* Reduce repeated PR setup and transcript timing waits; add hosted comparisons
* Align parallelism contract with Node-only external rebuild toolchain
* Record hosted coverage and launch package, store, and cancellation comparisons
* Apply hosted Windows setup savings and remove measured test waits
* Keep measured PR package gains and remove completed comparison jobs
* Report measured test counts with precise units
* refactor(native-chat): give the structured chat host one required logger
The structured chat runtime took an optional onError callback that the
desktop never passed, so a late dispatch settlement, an unanswered-dispatch
release, a journal event-sink write and a provider lifecycle delivery that
failed were dropped with no trace. Other host failures went to scattered
console.warn calls, which reach nothing in a packaged desktop build.
The runtime and host now take one required logger (warn/error with a scope
and fields). The production logger writes each entry as a failed span to
<userData>/logs/main.trace.ndjson, which the diagnostic bundle collects, and
to the console (stderr under a supervised headless host). The runtime and the
host wrap it so a logger that throws never fails what it reports, and the
install refuses without one. Sites that deliberately kept a recovery-capsule
error out of the log still log no error object.
* refactor(native-chat): hand the chat host's collaborators the logger, and give orcad its trace file
The delivery loop, idle sweep, queued-message drain, lease renewer, event
sink, conversation map and provider start/exit settlement each took an
internal error callback that the host mapped onto the logger. They now take
the logger itself and log under their own scope. The event sink keeps one
onFailed hook, which decides whether to stop the provider, not whether to
report. The dead-generation settlement returns its failure so each caller
logs it under its own scope.
orcad now installs the desktop's local trace sink under its own data root, so
a headless host's chat failures reach <data-root>/logs/main.trace.ndjson as
well as stderr.
Also passes the logger in the test fixtures the first commit missed, which
tc:node caught.
* fix(native-chat): keep repeated chat failures from flooding the trace file, and record their causes
- The production structured-chat logger writes a repeated failure (same level, scope, session,
message and error text) once per 5 minutes, carrying how many repeats it swallowed; the
tracked set is capped at 256.
- Trace entries now carry the error's code (and SQLite errcode) and up to three causes by name and
message.
- A chat read whose conversation will not open is logged through the host's logger
(open-for-read), and so are the runtime's chat-tab bookkeeping failures that already hold the
host.
- orcad writes its own orcad.trace.ndjson, closes it after every quit handler, and flushes it on
process exit; a trace file that cannot be opened leaves tracing off instead of stopping the app
or orcad.
- Tests: the desktop wiring test proves the logger reaches the trace sink, and the privacy tests
read every level the logger received.
* fix(native-chat): log a created chat's tab-publication and launch-prompt failures through the host's logger
* fix(native-chat): key a repeated chat failure on everything its entry writes
The repeat suppression keyed on the message and the error's text, so two refusals with the same
code but different causes, a plain error and a refusal of one code, or two object-valued errors
shared a key and the second was swallowed for five minutes. The key is now the entry's whole
written content (fields, code, errcode, refusal reason, cause chain, a stable rendering of a
non-error value) plus the error's name and message; a refusal's reason is also written.
* test(native-chat): pin that an error's name keeps two repeated failures apart
* test(native-chat): build the refusal in the repeat-key test as the wire does
* fix(native-chat): read an error's code and a refusal's reason by narrowing, not Reflect.get
* fix(ssh): launch the Windows relay outside sshd's job without WMI
Win32-OpenSSH kills a session's job on close but allows breakaway. relay.js
gains a one-shot launcher mode that starts the detached relay with
CREATE_BREAKAWAY_FROM_JOB through the staged process-tree addon, so a standard
user no longer needs a WMI Remote Enable grant. WMI stays as the fallback for a
relay without the addon, and a refusal there is named. The Windows SSH-host
lanes drop their WMI grant and assert the breakaway route and adoption.
* fix(ssh): find runtime holds without WMI on a standard-user Windows host
The store GC read held runtimes through Get-CimInstance Win32_Process, which
WMI refuses to a standard user's SSH logon, so the pass kept every runtime.
On a refusal it now reads this account's own process image paths through
Get-Process.
* build(relay): ship the Windows relay launcher addon in every desktop package
macOS and Linux packages carried Windows relays without windows-process-tree.node,
so a legacy-runtime relay they uploaded to a Windows SSH host could not launch
outside sshd's job and fell back to WMI, which a standard user is refused.
A reusable Windows job now compiles the x64 and arm64 addons once and uploads
them; release-cut, release-mac-build, and the hourly/daily/adhoc mac builds
download them before build:release and require both arches. Staging now rejects
a binary with the wrong PE machine, the ReadProcessMemory import, or no
spawnOutsideJob export, so a stale pre-launcher build cannot ship.
* ci(ssh): run the Windows SSH-host lanes when the relay process-tree build scripts change
The staging and gyp-rebuild scripts decide which windows-process-tree addon the
relay ships, so a change to either must re-prove the Windows host cells.
* test(ci): find the mac orcad-template download by artifact name
The release mac job now also downloads the relay Windows process-tree addons, so
the first download-artifact step is no longer the template's.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
* fix(ssh): collect the pinned-Node runtime store on Windows hosts
Windows SSH hosts now run runtime-store GC instead of skipping it: one
PowerShell inventory reads .runtime-ref-node-<sha> and .runtime-node refs from
every version dir, and one Get-CimInstance Win32_Process query filtered on an
image path under runtimes\ adds process holds (never by image name; a failed
query keeps everything). Stale upload stages are swept with the same rule as
POSIX. Promotion and the post-upload hold check now take the store lock on
Windows too, and the lock's own commands run unwrapped there.
Windows relay version-dir liveness now honours .relay-pid (design D5): a live
PID answers ALIVE before any pipe is touched, a dead one (ESRCH) plus refusing
pipes is exited, anything else is unverifiable. The runtime probe adopts a
pinned node.exe an earlier vault reader left without a .verified marker after
running it.
* fix(ssh): Windows stage fencing and vault runtime go through the verified node.exe
Upload-stage file identity on Windows no longer compiles an Add-Type P/Invoke
helper when the relay runs on Orca's verified pinned node.exe: the stage
commands run a fixed fs.lstatSync(..., {bigint:true}) script through it. It
prints the legacy helper's vol:high:low lowercase hex, and identity files are
compared after normalising hex spelling, so old and new clients recover each
other's stages. Host-Node relays keep the legacy helper; the choice is
documented in windows-edr-posture.md.
The Windows OpenCode vault reader now installs the pinned runtime through
ensureRemoteOrcadNodeRuntime (official zip, host-side extraction, .verified,
store lock) instead of uploading a client-extracted node.exe, and the relay dir
gains a .runtime-ref-node-<sha> so store GC keeps the runtime the vault uses.
* test(ssh): run the Windows stage-identity and store-GC tests on the Windows lane
The legacy/node.exe identity compatibility test and the Win32_Process hold path
were gated to win32 but no CI lane ran them. Add both files to the Windows
package lane and a real running-node.exe hold test.
* test(ssh): tear down Windows-lane temp trees through removeTreeSync
* test(ssh): grant the store lock to the Windows OpenCode runtime setup test
The Windows promote now runs under runtimes/.store-lock, so the mocked host
must answer the lock's CreateNew step.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
* build(orcad): build server glibc slots on glibc 2.28 and add the glibc 2.17 compat slot
Design D6: the default linux-{x64,arm64}-glibc node-pty slots now build in
manylinux_2_28 (digest-pinned) and are gated at glibc 2.28 / GLIBCXX_3.4.25
through a floor profile on verify-linux-glibc-floor.cjs; the desktop keeps
its Ubuntu 20.04 (2.31) default.
Adds the opt-in linux-x64-glibc217 compat target: NODE_RUNTIME_COMPAT_ASSETS
pins the unofficial glibc-217 Node (update/check pin scripts cover it, outside
SERVER_TARGETS), and a new CI lane builds the compat slot in manylinux2014
with static libstdc++, gates it at glibc 2.17 with no shared C++ runtime in
DT_NEEDED, and smokes it under the glibc-217 Node.
* refactor(node-runtime-pin): route compat lookups through isCompatServerTarget; keep the glibc doc's slot-name paragraph intact
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
* ci(daemon): gate PRs on daemon protocol crossing from the newest release
Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.
* feat(persistence): run profile backups in the worker whenever its entry is bundled
* refactor(orcad): make profile and native preflight runtime-neutral
The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.
* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check
Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.
ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.
* test(persistence): skip plain-Node backup selection tests in the Bun profile suite
* fix(runtime): reject a pinned archive that belongs to another target
* ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol
D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change,
so one PR must not do both. The launcher file list lives in the check script; the
allow-runtime-launcher-protocol-bump label overrides it.
* feat(orcad): select pinned-Node slots by a .runtime-node marker
D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through
.runtime-node instead of .build-target, so Bun-era clients read it as a legacy
slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet.
* fix(runtime): load the Node pin without the typeless-module warning
check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from
its own module, so it no longer loads the update script's build graph.
* fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout
* feat(orcad): 8-slot node-pty prebuilds against the pinned Node headers at N-API 8
- build-orcad-prebuilds.mjs adds win32-x64/arm64 (conpty.node, the vendored
conpty.dll/OpenConsole.exe, upstream's N-API conpty_console_list.node), compiles
in a scratch copy against the hash-verified pinned headers (node.lib pinned per
Windows arch) with NAPI_VERSION=8, rejects post-8 node_api_* imports, and writes
a schema 2 manifest with per-file sha256, N-API level and the glibc need.
- --require-slots [slots] verifies files against hashes; --smoke loads the slot
under the pinned Node and spawns a PTY; --print-slot names the host slot.
- The slot installer gates on N-API, libc, arch, glibc and file hashes instead of
the exact NODE_MODULE_VERSION, and installs nested files (conpty/).
- bun-profile-tests.yml builds, verifies and smokes each runner's slot.
* fix(orcad): scope node-pty's glibc .symver pins to glibc on musl prebuild slots
musl's unversioned libc cannot satisfy openpty@GLIBC_* references at link
time, so the Alpine slot compile would fail. Pin the staged pty.cc guard to
__GLIBC__ and assert both musl transforms against the installed patch.
* feat(orcad): run orcad on the pinned Node instead of Bun
A packaged orcad slot now references the pinned Node 24.21.0 by its
executableSha256 (`.runtime-node`, `.server-target`) instead of carrying
bun-runtime, and ships node-pty from the slot's prebuild, only its own
ripgrep, and no Windows Bun PTY gate. The runtime lives beside the slots
at runtimes/node-<sha>/bin/node (node.exe on Windows, upstream name).
- build:orcad (build-orcad-node.mjs) builds the host slot's prebuild when
missing and places the pinned runtime; the template is schema 3 with
per-target files.
- handoffToBundledOrcad() resolves the slot's runtime reference and checks
process.versions.node against the pin; a host Node >= 18 still hands off.
Startup preflight keys on running as that runtime; callers expect 'node'.
- orcad and its daemon use node-pty (ConPTY + windows-pty-job on Windows);
the Bun PTY sources, gate entry and canUseBunPty branches are removed.
- SSH deploy uploads the official archive once per pin, extracts and
hash-checks it on the host, and self-tests it before publishing. Bun
slots stay launchable for rollback; Node slots never use host Node.
- The runtime materializer is generic over pinned assets; the Bun wrapper
remains only for the OpenCode vault reader (design Phase 2).
- Cross-runtime test: a profile DB written by Bun 1.4.2 (WAL left by
SIGKILL) opens and backs up under the pinned Node, and the reverse.
No daemon PROTOCOL_VERSION change (design D7.1 R3).
* docs(ci): name the headless lanes after the pinned Node, drop Bun shard timings
Design D10: ci-demand-rollout.md and ci-runner-efficiency.md follow the
bun-profile-tests.yml -> node-server-tests.yml rename; shard timings drop the
deleted Bun PTY tests and follow the renamed ones.
* chore(ci): count the runtime archive download as a runtime launcher path
* fix(orcad): pin the macOS C++ standard for node-pty prebuilds
The official Node headers' config.gypi sets clang: 0, so common.gypi skips its
gnu++20 xcode_settings and Apple clang 15 (macos-14 runners) compiles
node-addon-api as C++98.
* fix(orcad): resolve the preflight's slot through realpath, as the handoff does
A symlinked orcad.js handed off to its real slot's pinned Node, but the
startup and profile preflights read the symlink's directory, found no
runtime marker there, and silently skipped the readiness check.
* refactor(ssh): drop materializeCachedNodeRuntime, which nothing calls
Deploys upload the verified official archive (design D5); no client path
needs an extracted Node executable cached by digest.
* test(orcad): gate the Bun-to-Node upgrade and Node-to-Bun rollback with live terminals
Design D7.1 R1/R3/R4 and D7.2. The last Bun orcad and this checkout's Node
slot are installed side by side under ~/.orca-remote, launched and stopped
with the client's own deploy commands, and share one data root. Each
direction proves the incoming orcad adopts the outgoing runtime's daemon
(same PID, same shell, output continues), opens its profile database and
backs it up with its own shipped worker, and that GC keeps the slot the
live daemon was forked from.
The node-server Linux lanes provide Bun 1.4.2 and build that Bun orcad
from main, and run with --cross-runtime. --artifact and --cross-runtime
now make their tests fail on a missing input instead of skipping.
* ci(node-server): pin node:24.21.0-alpine by its multi-arch index digest
* test(ssh): name the runtime archive fixture after its role
* test(node-server): load node-pty from the packaged slot in artifact runs
The node-server lane installs dependencies without building node-pty, and
Linux has no upstream prebuild, so the real-PTY failed-I/O teardown test
(picked up by the pty-subprocess selector) could not load pty.node. In
--artifact runs, alias node-pty to out/orcad's shipped slot so the test
exercises the addon orcad actually runs under the pinned Node.
* fix(orcad): let the Windows profile preflight exit after its PTY probe
On Windows, node-pty keeps the conout worker thread and pseudoconsole alive
until kill(), even after the shell exits. The PTY health probe never killed a
cleanly exited probe, so the packaged preflight printed its readiness line
and then hung until the build's 30s timeout, reported with an empty stderr.
- The probe kills its PTY on Windows after exit and uses the bundled ConPTY
the daemon spawns with.
- The preflight exits once stdout is flushed; its owner reads to EOF.
- Preflight failures now report code, signal, timeout, stdout and stderr.
* test(node-server): load the slot's node-pty in the real-PTY test, not by alias
A vite alias redirected only ESM imports of node-pty; windows-pty-job and
local-pty-utils resolve it through require, so Windows loaded two conpty.node
copies and the Git Bash job-membership proof read an empty job. The failed-I/O
teardown test now loads node-pty through a fixture that picks the packaged slot
in artifact lanes.
The pty-subprocess selector was a prefix that also pulled in its POSIX-host
sibling unit tests, which pr.yml runs and which were never qualified on
Windows. Select the directory plus the two sibling files that belong here.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
* fix(runtime): decide Antigravity readiness from the live screen
agy paints its composer with cursor addressing, so the line-folded wait
text misses the 1.2.14 accept-edits and plan composers and an ended turn,
while the grid keeps the bare `>` caret painted mid-turn and behind the
/model picker. Read the screen's bottom rows instead: rule, caret, rule,
`? for shortcuts`. A clocked pane is held to quiescence (tier 1b) because
the submit repaint reads ready for a moment; a clockless restored pane
settles from the screen alone. When a trustworthy screen exists it
decides, so the name-only title lane no longer settles an open picker.
Retires the visible-read probe's Antigravity branch: the probe now runs
the shared screen rule for any screen-ruled agent without an output
clock, and keeps its generic empty-pane read for everyone else.
Adds twelve agy 1.2.14 recordings and a replay suite shared by
screen-ruled agents. STA-8741.
* fix(runtime): decide Cline readiness from the live screen
Cline paints its composer box with cursor addressing on the alternate
screen, so no text rule saw it and worker-start timed out at
agent_readiness (#23268). Read the box off the grid: rule, an empty
composer with one of the captured placeholders, rule, the Plan/Act row
and the auto-approve row, with no braille spinner above it.
A streaming reply repaints the same box once its spinner has scrolled
away, so Cline is tier 1b only: a clocked pane waits for quiet and a
clockless one never settles from the screen. The screen now decides for
a Cline pane, which shuts the quiet-process lane that would have settled
its unworded tool-approval prompt and the Cline Desktop promo.
readLiveTerminalScreenLines now returns raw rows: the read projection
blanks a composer it takes for a draft, and it takes Cline's
placeholder for one, so a typed draft and an empty composer looked the
same.
Adds nine cline 3.0.66 recordings (macOS) and the 3.0.65 Windows capture
from #23269. STA-8741.
* fix(runtime): decide Prime Agent readiness from the live screen
Prime redraws its composer on the alternate screen, so the text tail
never showed a settled prompt and tui-idle timed out (#22153). Read the
grid instead: a bare `>` directly over the `<- manage` footer, with no
braille status row (`Writing - 6s`) above it. The footer and caret alone
stay painted for a whole turn.
Replayed chunk by chunk, Prime erases that status row before redrawing
it, and on first launch paints the idle composer just before the
trace-sharing question covers it. Both keep repainting, so a clocked
pane is held to quiescence (tier 1b); a clockless restored pane settles
from the screen alone.
Adds nine prime-agent 0.9.8 recordings (isolated HOME, OpenRouter) and
the two 0.9.5 captures from #22154. STA-8741.
* refactor(runtime): drop Cline-only readiness branches
Cline now follows the same pattern as Antigravity and Prime: a screen
rule plus table entries.
- Drop MID_TURN_COMPOSER_AGENTS. onPtyData stamps lastOutputAt on every
chunk, so a re-attached streaming pane has an output clock from its
first byte; the exception only guarded a pane that printed nothing
since attach. A clockless Cline pane now settles from its screen like
the other two.
- Drop the 'ready-body' rest-signal entries for all three agents. The
rest signal is read only by quietForegroundLane, and a readable screen
already shuts that lane and the title lane (isReadinessDecidedByScreen),
so the entries only removed the quiet-process fallback for a pane with
no trustworthy grid. The census now checks that screen-shut instead.
- Drop the Cline rule's auto-approve row check; no recorded verdict
depends on it.
Kept: raw rows from readLiveTerminalScreenLines. Every frame of every
codex-* and qoder-* capture at 120x40, 80x24 and 100x32 gives the same
isKnownReadyPromptBody (with and without a clock) and
isQuietReadyScreenBody verdict through both readers.
Serializer known-failures for the new captures are pre-existing
serializer behaviour, not this branch: row-0 cells restore with a
true-colour background where the source has the default (the DSH
class), and Prime's cursor restores at column 119 instead of the pending
wrap at 120 (the qoder class). STA-8741.
* fix(runtime): trust a screen rule only on the PTY's own grid
Review findings on the screen-ruled readiness (STA-8741):
- A grid out of step with the PTY garbles cursor-addressed chrome, and a
model resize does not make the TUI repaint. readLiveTerminalScreenLines
now returns null unless the emulator's grid matches the PTY's reported
size and was never reflowed without a repaint (a re-attach that learned
the real size late), so the pre-existing lanes decide there instead of
timing out.
- The visible-read probe reads the draft-blanking projection, which
turns Cline's `❯ Ask anything...` into a bare `❯`. It now restores the
blanked composer row before the rule reads it; `terminal read --screen`
output is unchanged.
- The quiet lane no longer ORs the text rules over a trustworthy screen
that refused; without one, tier 1 already ran them. No recorded
verdict changes.
Tests: ready recordings on a mismatched and on a reflowed grid settle
through the old lanes; the restored-pane probe runs every ready
recording through the real projection; the rest-signal census checks
the lane verdict with and without a screen.
* test(runtime): trim STA-8741 recordings to the screens they prove
* refactor(runtime): one screen verdict for every screen-ruled lane
readScreenRuledReady, readScreenRuledQuietReady and isReadinessDecidedByScreen
each re-derived the same thing: the agent's rule applied to a trustworthy live
screen. They collapse into readScreenRuledVerdict (true / false / null), which
tier 1, tier 1b and the lane gate read.
This also makes a refusal final in tier 1: a clockless pane whose trustworthy
screen refused fell through to the text rules, so retained ready text could
settle over an open picker (Greptile review). The quiet tier already refused
there; now both do.
The tier-1b agent set derives the screen-ruled agents from the rule table
instead of listing them again, and the lane test that repeated the census
case is dropped.
* refactor(runtime): let the visible-read probe read its own output clock
The probe's clock was captured at start and threaded through the wait
dependencies as a one-off parameter. The probe now reads it from the live
record when its screen read returns, which is also the fresher answer.
* fix(runtime): trust a reflowed grid again once a PTY resize repaints it
The reattach-reflow flag was never cleared, so a pane stayed on the old lanes
for the rest of its life even after a real resize made the TUI repaint
(Greptile review). The record now keeps the reflowed grid, and a PTY resize
off that grid clears it; an echo of the same size sends no SIGWINCH and keeps
it.
Tests: the reflow case in every screen-ruled suite now includes a same-size
echo, and an Antigravity recording only the screen reads ready settles after a
resize and repaint.
* refactor(runtime): keep screen-rule trust and raw rows to screen-ruled agents
Two shared changes reached agents this PR does not target: the live
screen reader returned raw rows, and it refused a grid that did not
match the PTY. Both now live in readScreenRuledLines, which only the
screen-ruled agents read (screenReader picks it from the rule table);
readLiveTerminalScreenLines is main's again. The probe keeps main's
Antigravity-banner trigger, so a Codex or unknown pane is probed exactly
as before.
Proof: the non-screen-ruled suites give identical pass sets on this
branch and its base (1,781 tests), and replaying every other recording
frame by frame through the readiness and blocked verdicts, for its
agent and for an unknown pane, gives identical results (93 pairs). A new
test keeps a Codex pane reading its screen when the PTY reports another
grid; it fails if the trust check moves back into the shared reader.
* feat(native-chat): publish the host's child records to the status summary and the chat strip
The status summary and the background-task channel now read a session's child
records from the host's canonical store, through the status sink its row
landed in, and derive the legacy task and subagent shapes from the same views.
The parent row folds its child-work liveness from those records at ingest,
not from the summary's task list. The adapters no longer push their task DTO
to clients: the onBackgroundTasksChanged path is gone, and a child-work ingest
is what republishes both the summary and the strip. Finished children stay
listed until the session's own next turn starts. A reader that predates child
views never receives a roster whose rows are all settled.
* feat(native-chat): the chat strip reads the host's child records with its parent's verdict
The strip's roster now renders from the child views its channel carries, and
passes the verdict the session's own status row gives its children, built the
way the sidebar builds it (the row's freshness and the status feed's
observation). So one child reads the same in the strip and the sidebar, live,
after the transport drops, and once the row goes stale. A roster of finished
children stays shown until the next turn but no longer animates the monitoring
indicator or blocks conversation commands.
Tests: an end-to-end run on a host with no renderer (a real hook server as the
status sink) shows the summary and the strip channel carrying the same records
at every step, the parent row folded from them, retention, and an older
client's task list holding live work only; a wired renderer test shows both
surfaces agree when live, lost and stale.
* test(native-chat): a newer host's view degrades, an older reader keeps its live roster, a finished roster holds nothing open
- The view decoder ignores unknown keys, degrades unknown kinds, states,
outcomes and memberships, and drops only rows it cannot identify.
- At the RPC boundary a reader that predates child views gets no strip for a
roster of finished children and never the views themselves; a stop-only
reader keeps rows whose host offers no targeted stop.
- The strip shows finished children without reading them as live work.
- The row keeps its child list's identity when a summary repeats it.
* refactor(native-chat): the summary's task list is the legacy projection's live rows, unfiltered
* test(native-chat): type the switch tests' mocks instead of asserting them
* test: remote clients advertise reading child views
* docs(agent-status): the structured row folds the store's child records
* fix(agent-status): keep the view reader's header from reading as a value import to the renderer boundary
The renderer node-builtin boundary test scans raw text, so a header comment
that said "imports" ahead of the import block turned the type-only import of
agent-status-child-work into a value edge that reaches node:crypto.
* refactor(native-chat): the status summary's broadcast equality gets its own module
The status feed crossed the file-size limit once the summary gained the main agent's turn
outcome beside the child views. Which summary changes reach every session list now lives in
structured-agent-session-status-summary-equality.ts.
* fix(native-chat): command admission reads the strip's child records
A conversation command was refused on the provider tracker's own roster
while the strip read the host's child records, so a drift between the two
rule sets could refuse /clear with a stop instruction the strip had no
button for. Admission now reads the same records through the same read as
the strip, uses the strip's liveness fold, and asks for a stop only when
the strip renders one. The adapter contract no longer exposes the tracker
roster, so no host decision can read it.
Also records when the legacy child shapes die, every earlier death of a
settled child, and the display-precision invariant behind the summary's
clock tolerance.
* refactor(native-chat): command admission takes only what it reads of a turn
* fix(native-chat): the session list drops a session's children when the store does
A session's end no longer removes its child records: a child still running
settles with an outcome nobody reported, and a finished one stays listed.
Records now leave only at the session's own next turn, at the cap on settled
records, or when the host lets go of the session and its row leaves the store.
The summary kept after the host lets go used to strip its children on close,
a rule of its own. It now re-reads them from the store when the row leaves,
through the same read every live summary uses, so the session list and the
chat strip list the same children at each step, including a forget with no
close. Closing only revokes ownership, as before the child records existed.
* test(native-chat): write the Codex frame script's parent row out step by step
Once every surface reads the child records, the provider tracker's roster is
no oracle: it and the records read the same child executions, so agreeing
with it cannot catch a defect in either. Each frame now states the child
liveness and the parent row it must fold to.
* fix(native-chat): the idle sweep and the restart snapshot read the host's child records
The idle sweep (keep an agent running while its subagents or commands run) and
the restart-resume snapshot (what a chat was doing when Orca stopped it) both
read the provider tracker's roster through the adapter interface, which no
longer carries it. Both now take the host's one child-record read, the same one
the status summary, the chat strip and command admission use.
The snapshot's working test also folded that roster through the shared fold's
old `backgroundTasks` input, which the fold no longer reads, so a settled lead
whose subagent was still running would have been offered nothing. It now hands
the fold the records.
* test(native-chat): the child-record tests follow the merged command lifecycle
A command is a live child record from its start and is removed, not settled,
when it stops, whatever Codex tagged it. The end-to-end switch now shows the
child's `npm test` as a live row beside its dev server, and both are gone once
they exit; only the finished subagent stays listed until the next turn. Letting
go of the session is its tab closing, since a closed conversation whose tab
remains keeps its row.
Command admission's finished row is a subagent, the one kind that settles, and
the failed-verdict row test admits its live subagent as a host record, the only
thing the row folds.
* refactor(native-chat): the status feed's journal projection cache gets its own module
The status feed crossed the file-size limit once the child records joined the
agent-start signal and the completion feed's status read. The per-journal
projection, cached per commit, now lives in
structured-agent-session-status-journal-projection.ts.
* test(native-chat): the admission test's compaction resolves with a real outcome
Main's compaction result is a tagged outcome; the host-level admission test
resolved its mock compaction with an empty object.
* feat(native-chat): the sidebar lists running subagents; the strip, running then the newest finished
The host keeps every child record; what each surface lists is picked from them on every read, so
nothing is stored twice. The status summary, which every session list reads, now carries only
running children (and a finished one whose shell still runs, which reads monitoring): a finished
or failed subagent leaves the sidebar and stays in the chat's strip. The strip lists every running
child, then the newest finished ones, 100 rows in all; more than 100 running all show.
This matches common practice: sidebars show live subagents, and finished ones stay in the chat's
panel, newest first. No wire field is added. An older client reads fewer rows: its legacy task
lists were already live-only in the summary, and the strip's settled tasks come from the same
bounded roster.
* fix(sidebar): one rule for what the worktree sidebar lists: running children, from every source
`worktreeSidebarListsChild` is the one definition: a child that runs, counting a finished one
whose own shell still runs (it reads monitoring). The sidebar's row builder applies it to every
child source it reads, a terminal agent's hook roster and a chat session's records alike, and the
host's status summary applies the same predicate, so the sidebar's payload stays small. The chat's
strip keeps finished children, newest first.
A terminal agent's hook roster already drops a child on its own stop, so nothing changes there:
a teammate between turns and a child gone quiet still run, and still show. The selection module
moves to `agent-child-work-listing.ts`, since it now covers every source, not only chat sessions.
* test(native-chat): the switch test passes the startup child key main's status bar takes
* fix(native-chat): a finished child stays until the user's next send, not a turn Claude opens on its own
Claude wakes the agent on its own when a background task ends, and that wake is a
new root turn. Keying retention on the newest root turn retired every finished
child about two seconds after a background agent or shell finished, so its
outcome never showed in the strip.
Retention now keys on the user's newest send the provider accepted (a message, a
steer or a command), with the journal epoch so a rewind still retires. A turn the
provider opens itself and a subagent's turn carry no send. Replayed captured wake
orders through the real adapter, hook server and status feed.
* fix(native-chat): a background Stop reaches the tasks the child records show
The strip draws a row's Stop, and /clear, /compact and rewind wait for background
work, from the host's child records, but the Claude adapter still resolved which
tasks a Stop reached from its own tracker's roster, and refused to stop at all
once that roster was empty. A task the records kept live after the roster dropped
it showed a Stop that sent nothing and blocked those commands until the chat tab
closed.
The host now resolves the provider ids a Stop sends from the records (the same
per-row rule the strip and admission use; every such row for stop-all), and the
adapter stops exactly those, with no tracker guard. An acknowledged stop ends the
record: a running task sends its own stopped frame first, and the CLI answers
success with no frame for a task it no longer knows. A refused stop leaves the
record live. No production code reads the tracker's roster any more.
* fix(native-chat): one rule for a finished child that still owns live work, at any depth
The listing kept a finished child whose work ran through any depth of ownership,
but retention at the user's next turn protected only the direct owner, so a
finished agent whose finished subagent still ran a shell was removed and that
subagent jumped to the top level. Both now read settledOwnersOfLiveWork.
* fix(native-chat): an older client sees a Codex child's shell as it did before views
Clients that predate child views read a flat task roster derived from the views.
It listed a Codex child agent's shell as an extra row beside the running agent,
then as a bare command once the agent finished. The derivation now hides a
running agent's commands and names a finished agent's as "<agent> — <command>",
as the Codex tracker did; the label rule moves to a shared module both use.
* fix(native-chat): the chat decodes a roster's child rows once, as the frame arrives
The client reducer compared raw wire rows, so a row shaped by a newer host could
throw there, and the strip decoded a new array on every render, which defeated its
grouped-rows memo while a turn streamed. Rows are now decoded where the frame
enters the reducer, an unchanged roster keeps its identity, and the strip's parent
context is rebuilt only when one of its values changes.
* fix(native-chat): the strip channel forgets a closed conversation's roster
It kept the last roster fingerprint of every conversation for the host's lifetime.
The conversation map now tells observers when one leaves it. Also corrects the
summary's children comment: it carries running children only.
* docs(native-chat): rewrap the retention comment
* fix(native-chat): a task's own ending replaces a Stop's, and a child finished after the user wrote stays
Two lifecycle gaps from the round-1 fixes.
A Stop acknowledged ahead of the task's own ending relabelled it. The SDK hands
Orca a control answer as soon as it reads it and queues other frames, so when a
task finished just as the user pressed Stop, the acknowledgement arrived before
the task's completion the CLI wrote first, and the task read "Stopped" with its
result lost. An acknowledged Stop now ends a record provisionally (outcome basis
`stop-acknowledged`); the task's own terminal frame replaces it, and nothing
replaces an ending the task reported itself.
A finished child still vanished with no new action from the user when the send
the provider took was written before the child finished: a steer Claude takes at
its next boundary, or a queued draft handed over at the end of the turn. The
user's next turn now carries when they acted (the send's written time, or the
draft's queued time, both on the host clock), and only children that finished at
or before it retire; a later one stays until the user's next send. A rewind
still retires every finished child.
Also: a Stop-all keeps stopping the remaining tasks after one request fails, then
reports the failure.
* fix(native-chat): the strip keeps one empty list for a roster that omits one
A roster with only running or only finished rows made a new empty array on every
render, so the strip regrouped its rows each time. One shared empty list keeps
its memo.
* fix(native-chat): a strip row whose owner the 100-row budget cut renders under the main agent
The budget can keep a finished child and cut its finished owner. The child still
named that owner, so it rendered nowhere. The selection now clears an owner it
did not keep, as the view contract says for an owner outside the projection.
* fix(native-chat): a stop-all that times out stops asking, and "no children" is sent once
A stop-all kept asking after a request timed out, so a Claude CLI that stopped
answering control requests cost one full deadline per task, while the chat's
sends, Stop and /compact waited behind it. A timeout now ends the loop; other
request failures still let the remaining tasks be stopped.
The chat strip channel never remembered that it had sent "no children", so
every change in a chat with none re-sent that frame to each subscriber. It now
remembers it, and forgets only when the conversation closes.
Also renames agent-child-work-stop.ts to agent-child-work-stop-targets.ts, which
says what it answers: the provider ids a background Stop reaches.
* fix(native-chat): the status feed reads a journal snapshot with no submissions, and e2e tests use the child-work reader
CI on dbd2439cd4 was red in three places:
- first-work-branch-rename and the agentSession.subscribeStatus RPC test feed the status feed a
journal whose snapshot lists items only. The projection reads the user's newest accepted send
from `snapshot.submissions`; it now tolerates their absence, as the status projection beside
it already did.
- the cross-version downgrade test still passed `backgroundTasks` to the teardown's working
marker, which now takes `childWork`.
- an e2e unit test still gave the status feed the removed `readBackgroundTasks` dependency
(harmless at run time, a type error in the tests/ project).
* fix(native-chat): the chat strip lists running children only, by the sidebar's rule, and hides when none runs
A finished subagent's result is already in the transcript ("Ran N subagents ·
completed"), so the strip is for work that runs. It now lists exactly what the
sidebar lists, by one predicate (a running child, or a finished one whose own
shell still runs, which reads monitoring), and the host sends no roster once none
runs, so the strip hides.
Gone with it: the 100-row budget and the running-then-newest-finished
selection, the re-homing of a child whose owner the budget cut, and the RPC
gate's rule for a roster of finished rows only (no such roster exists now).
Older clients still get their derived task list, running work only.
Finished records still stay in the host's store until the user's next accepted
message: they refuse a late frame of their run, let a task's own ending replace
an acknowledged Stop's, and keep a running shell's owner. Dating that retention
by when the user wrote the message only kept finished rows visible longer, so it
is removed.
* fix(native-chat): the strip shows running work only from any host, and hides after a released session's last child
- A new app paired with an older host no longer shows that host's finished task
rows: the strip lists running work only, whatever host sent it, and hides when
an older host's roster has only finished rows left.
- A test for the path that hides the strip when a session's last running child
settles after the provider let go of the session (Claude's release path): the
channel sends `null` though no provider answers for the session any more.
- A test comment still described the strip keeping finished children.
* refactor(native-chat): remove the unused terminal handoff
No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.
* fix(native-chat): never let the pre-stop snapshot hold a chat's stop
Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(native-chat): drop helpers only the terminal handoff called
`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.
Co-Authored-By: Claude <noreply@anthropic.com>
* docs(native-chat): stop citing the removed handoff in lifecycle comments
Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): type the stalled snapshot drain without a cast
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): pin that a start dead before proving owes no settlement
The removed restart handoff test pinned this branch; nothing else did.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(native-chat): keep the owner-status read behind an in-flight attach
The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(terminal): remove the agent-session PTY write gate
The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(native-chat): drop the transcript helpers only the handoff called
appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(native-chat): stop calling a starting chat "mid-handoff"
A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): type the stand-in roster decoder without a cast
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(codex): name the pinned rollout lookup for what it does
With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.
* refactor(native-chat): type the owner-status reply as the host sends it
The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.
* refactor(native-chat): normalize terminal-handoff lease values once at decode
Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.
The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:
- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
`conflicted`, the claim every build probes but never stops. A plain native
owner would be stopped by restart recovery, here and in older builds.
Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.
The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.
* refactor(native-chat): stop threading the owner kind through a reservation
A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.
* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else
Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.
* fix(native-chat): name a chat write by its target, not the owner generation
A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.
Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.
Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.
* fix(native-chat): every journal append reaches the chats that are open
A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.
A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.
* test(native-chat): an epoch replacement reaches the open chat
* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map
* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite
The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.
* test(worktree-activation): restore the OMP surfaced-agent resume test
The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.
* perf(native-chat): a publish behind a delivered commit reads nothing
Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.
* test(native-chat): state why the teardown test's fake journal is safe to cast
* docs(native-chat): say mutation admission checks only the writer lease
* docs(native-chat): drop the send rebase from comments that still described it
* fix(native-chat): a message is accepted, then delivered
A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".
A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.
Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.
A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.
Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.
* fix(native-chat): settle queued messages only for the child that ended
A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.
A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.
The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.
* fix(native-chat): an adoption that fails to import keeps the conversation open
The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.
* perf(native-chat): the recovering open reads the journal once
Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.
* fix(native-chat): an attach that fails after indexing its child leaves no child behind
A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.
* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer
The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.
A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.
* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down
The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.
* fix(native-chat): a message rejected while its chat was closed reads as not sent
A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.
* test(orchestration): name why the readiness settlement fakes are cast
* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent
* docs(native-chat): drop the fence from the admission the send effects run behind
* docs(native-chat): give the fence move on release the reason that still holds
* docs(native-chat): stop citing a write fence check in launch and mailbox comments
Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.
* refactor(native-chat): the provider child is its own record
A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.
- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.
* fix(native-chat): the delivery loop alone settles a message its start or child failed
A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.
- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
reads how it ended: a Stop continues; anything else writes one failure row and rejects every
queued message with the same words, then stops. A child still starting whose start the adapter
says did not land fails the same way. The exit, eviction and the settlement retry only settle
the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
closed, with or without a child, and a start the loop already has in flight is waited for so the
child it produces is stopped rather than left behind.
* refactor(native-chat): a stopped child ends on the one reading of its stop
The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.
* feat(native-chat): the host says it accepts a send before any agent has it
The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.
* refactor(native-chat): an attach never opens a journal of its own
The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.
* fix(native-chat): a moved fence resends nothing on a host that accepts first
The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.
The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.
* refactor(native-chat): a child's end says whether the user or the host stopped it
The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.
* fix(native-chat): a chat whose only work is a queued message is not offered for resume
A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.
* test(native-chat): type the queued-message fixtures in the resume-offer tests
* fix(native-chat): a start that dies while a message waits on it is that message's failed start
Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.
* fix(native-chat): a request that failed reads as failed
A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.
The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".
* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now
* test(native-chat): a verdict change republishes the mobile status projection
* refactor(native-chat): the store's retention trigger keeps its flag compare
A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.
* test(native-chat): a user message the provider journaled keeps its session listed
* test(native-chat): pin what a failed start settles, and what a resume offer names
A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.
* test(native-chat): the failed-start pins fail on what the message became, not on a timeout
* fix(native-chat): a late provider-session update keeps a failed recovery record failed
A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.
* test(orchestration): the preamble's host stub is typed, not cast
The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.
* test(native-chat): the terminal-bell check asserts the renamed verdict field
The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.
* fix(native-chat): a failed turn ranks like a completion for attention
Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.
The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.
* fix(native-chat): a failed main agent reads failed while its subagents still work
The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.
Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.
worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.
* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it
The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.
* docs(native-chat): the status-store listing rule names provider-journaled user messages
* fix(native-chat): a refused send notifies failed through the completion feed
The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.
* fix(native-chat): every copy of a row carries the main agent's own status
History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.
- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
rebuilding one; the sync key and history equality compare it.
* test(native-chat): pin the worktree ps verdict across host and phone versions
Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.
* test(mobile): name the parity table's row for its role
* fix(native-chat): a request that settles while the user is asked something notifies once
The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.
The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.
* fix(native-chat): the completion says when the user is being asked
A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.
The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.
* fix(worktree-status): a departed agent's failure yields to live work on the worktree card
A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.
* docs(agent-status): a departed agent's failure ranks below live work on the worktree card
* fix(native-chat): a view never restarts a chat whose last start failed
A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.
* test(native-chat): start the child the loop waits on with an attach, not a second view
A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.
* fix(native-chat): settle a gone generation's turn wherever a conversation opens
A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.
* test(native-chat): prove the next child's start settles the turn an earlier child left
The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.
* test(native-chat): count a failed start's rows by row, not by text
Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.
* test(cross-version): load the phone row readers without mobile's toolchain
Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.
The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.
* test(cross-version): keep the checkout path-guard message and justify the copy import's cast
* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget
* test(native-chat): pin the open's and the send's start and row counts, however the view binds
Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.
* fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm
When the provider gave no verdict, the structured status projection now derives
one from the newest turn's lifecycle: interrupted -> interruption, unverifiable ->
unconfirmed. Nothing new is journaled, the completion feed stays provider-only, and
the legacy interrupted flag stays a user stop only. Every verdict reader handles
both arms explicitly.
* fix(native-chat): settle a gone generation's turn at every open but an acquisition's
The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.
* fix(native-chat): a folded turn a crash cut off reads Interrupted after N
The settled-turn timing now carries the turn's verdict, derived by the same
agentTurnVerdict the status row uses. A turn that ended interrupted with no
provider verdict heads its fold 'Interrupted after N'; a user's stop keeps
'Worked for N'.
* test(native-chat): hold the create's start open until the views bind
The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.
* fix(native-chat): the user's close of a chat records the turn it cuts short as their cancellation
The expected-close settle writes outcome cancellation when the user aimed the
stop at this chat: a Stop while the agent starts, the chat's tab closed (the
agentSession.close RPC, or session.tabs.close with a user reason), or /clear.
A quit, an idle eviction, a worktree teardown or an orchestration stop leaves the
turn with no verdict, so it still reads Interrupted.
* test(native-chat): a Claude turn a newer send superseded reads Interrupted
The supersede fires for any send Orca dispatched, the user's or another agent's,
and nothing at that site records the sender, so the turn keeps no verdict and
folds as Interrupted after N.
* fix(native-chat): the user's close records cancellation on the turn the provider settled on its way out
The Codex and Claude adapters settle their open turn as interrupted, with no
verdict, while the host stops them. The expected-close settle then found no
running turn, so a user's close of a mid-turn chat read Interrupted. The stop
now reads the running turns before it reaches the provider and records the
user's cancellation on each one it cut short, unless the provider gave a
verdict of its own. The close-verdict test's adapter now settles its turn on
close the way the real adapters do.
* test(native-chat): update the close and settled-turn expectations for the host-observed verdict
agentSession.close now passes the user's word to the host, and a settled
interrupted turn with no provider verdict carries `interruption`. Also merge a
duplicate import the code-quality gate rejects.
* refactor(native-chat): drop the composer's second error formatter
After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.
* test(native-chat): pin the reason on a message rejected while its chat was closed
The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.
* fix(native-chat): the user's close cancels a turn whose start landed as the provider stopped
The close read which turns were running before the stop. A turn whose start was
still in flight (a send echo not yet journaled) was absent from that read, so the
provider's verdict-less settle on the way out left it Interrupted. The close now
reads which turns were already over instead, and records the user's cancellation
on every other turn the stop left running or interrupted with no verdict. A turn
cut off earlier, or finished during the stop, keeps its end.
* fix(native-chat): a send the provider never received after a restart has no verdict
Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.
* refactor(native-chat): the adapter settles the turn a stop cuts with the stop's typed cause
The host hands its stop's cause to the adapter's close. Each adapter settles its own
open turn on 'ended' through one mapping, turnVerdictForChildEnd, and the host's
dead-generation fallback uses the same mapping for any turn no adapter settled. A
user's close or stop of this chat is their cancellation; a quit, eviction, teardown
or an exit the adapter saw first is news.
Deletes the snapshot-and-diff reconstruction (endedTurnItemIds,
userStoppedTurnRevisions, settleUserStoppedTurns) and the requestedByUser flag.
host.close now takes a required cause.
* test(native-chat): a user's close drops the chat's status row like an eviction
* test(native-chat): expect the eviction cause on the host closes of idle release, worker stop and worker discard
The stop's typed cause now travels into host.close and the adapter's close, so these
three non-user closes assert the 'evict' they pass.
* refactor(native-chat): every stop names its cause, so none defaults to the user's cancellation
stopStructuredAgentSessionAgentUnderSerialize defaulted its ending to 'user-stop', which now
settles the cut turn as the user's cancellation. Every caller already passes a cause; the
parameter is now required, and a type-level test fails to compile if the default returns.
* test(native-chat): pin who a chat's session.tabs.close speaks for, older clients' reasonless close included
The mapping lived inline in a type-unchecked file, and only the explicit user reason had a test: an older client's reasonless close, or a lifecycle echo read as the user's, stayed green. It is now one exhaustive, type-checked function with a case per reason.
* style(mobile): draw the unconfirmed dot in the theme's status amber, not an inline hex
* docs(agent-status): the main agent's outcome also carries the host-observed end, interruption or unconfirmed
* fix(native-chat): a chat the user closed while its agent started is not a failed start
A still-starting child the user's close cut counted as a failed start, since only 'user-stop' was excluded: a start-failure row, and queued messages rejected as a provider failure. Whether an ending fails its start is now one exhaustive switch, shared by the failed-start read and the delivery loop's handover: a user's stop or close never does; an exit, a failed attach and the host's own stops still do.
* fix(native-chat): a chat the user closed closes its queued messages, and starts no agent for them
After 876b6989f1 a user's close of a still-starting chat went on like a Stop, so when the close did not complete the delivery loop started a new agent for the message queued behind it. A child's end now has three dispositions, not a failed-start boolean: a user's Stop lets the queue go on, the user's close closes what was queued before it, and any other end fails it. The close is the one a completed close does (the provider-closed rejection, no verdict), applied at the top of each delivery step and ordered against the close so a later send still goes on.
* fix(activity): a crash-cut turn draws the interrupted glyph; only a user's Stop keeps the done check
The Activity page drew every Interrupted row with the done check, which #2569 chose for a user's Stop. With a crash now reading Interrupted, that put a green check on a turn nobody asked to stop. The row's glyph is now an exhaustive switch over the verdict: a cancellation keeps the done check, and an interruption draws the existing interrupted dot. An unconfirmed end already drew its own glyph.
* fix(activity): the Interrupted group header draws the done check only when every row is a user's Stop
A user's Stop and a crash share the Interrupted group, and its header took its newest row's glyph, so a Stop newer than a crash put a green check over the crash. The header is now folded over the group's rows: the done check only when every row draws it, the interrupted dot otherwise.
* fix(native-chat): a failed close of what the user closed starts no agent for it
Rejecting the messages a user's close left queued swallowed a journal write failure, so the delivery step went on to start an agent for a message in a chat the user closed. The rejection now reports whether it landed, and a step whose rejection failed stops instead; the next wake re-derives and retries it. Also pins that the ordering against the close holds only within its epoch, since a later epoch's sequences restart.
* test(native-chat): the idle sweep's stop is an eviction, so its close carries that cause
Main's idle sweep now stops an idle agent through the conversation lifetime, which this branch gives the 'evict' cause; its expectations name it.
* fix(native-chat): a retried stop keeps the cause of the stop it finishes
A user's Stop or close whose wind-down failed after the child was proven gone was finished by the idle sweep as an eviction, so the turn it cut read Interrupted. The owed wind-down now carries its stop's cause, and a retry with no child settles with it.
* test(native-chat): the idle sweep's close of a retrying Claude chat carries the eviction cause
Main's new test expected the adapter close with the session id alone; every stop now names its cause, and the idle sweep's is 'evict'.
* fix(status): a user's Stop marks done on the tab and sidebar; red Interrupted is only a turn cut short by something else
The tab, the worktree card and the sidebar rows drew a Stop with the same red dot as a crash. The
verdict mark now maps a cancellation to done, still saying "Interrupted by user" in the row text,
and the mobile mirror follows. The Activity page keeps grouping a Stop under Interrupted with the
done check, as before.
* test(cross-version): a new phone reads a user's Stop as done; an old phone still draws it interrupted
* fix(native-chat): a user's Stop inside a live Claude chat reads as their cancellation
Stopping a running Claude turn interrupts it and keeps the session, so the turn's end comes from
the CLI's result frame. Claude CLIs before 2.1.91 send that frame with no terminal_reason, and later
ones may still omit it, so the user's own Stop was recorded as a failure with an error row.
Orca now records the stop on the open turn when it sends the interrupt. An error result for that
turn reads as the user's cancellation whatever reason the CLI gives. The stop belongs to that one
turn, so it cannot reach the next, and it is withdrawn when the CLI refuses the interrupt.
* docs(agent-status): a user's stop marks done; name the tab close cause by its type
The reference still said a stop marks a row interrupted and ranks between live work and an
unconfirmed end. A cancellation now marks done, and only a turn cut short by something else ranks
as interrupted. The runtime's tab close restated the close cause's union; it now uses the type.
* test(native-chat): a proven crash reads as an interruption on the status feed and in the chat
A crash the relaunch proves now settles its turn interrupted, and the status feed works the verdict
out from that record, so the restart test expects interruption for a proven crash and unconfirmed
for one it cannot prove, never a cancellation. A chat read before the proof lands reports
unconfirmed, then interruption and a folded "Interrupted after 27s" once the proof revises it.
* fix(status): a user's Stop reads Interrupted, and a turn anything else cut short reads Failed
The verdict mark now maps a cancellation, the user's own Stop, to interrupted, and an interruption,
a turn cut short by a crash or a killed agent, to failed, the same as a failure, which outranks live
subagent work. An unconfirmed end is unchanged. This applies to every agent, in a terminal or a chat,
on the tab, the sidebar rows and worktree card, the dashboard row, Cmd+J and the phone. A Stop is
not news, so the rollups rank it below an unconfirmed end, and notifications word an interruption
"failed". Recording is unchanged.
* fix(activity): group a user's Stop under Interrupted and a crash with failures
A user's Stop draws the interrupted glyph and sits alone in Interrupted, and a turn anything else cut
short sits in Failed, titled "Agent failed". Every row in a status group now draws the group's own
glyph, so the header is the group's status and the rule that folded a Stop's done check into the
header is gone. Interrupted ranks below an unconfirmed end, as in the sidebar.
* fix(status): draw a user's Stop in the muted tone, not the fault red
The interrupted dot, which now means only a user's Stop, draws in the muted foreground token on the
agent rows, the sidebar card and the phone. Red stays for a failure or a turn cut short by anything
else, and green for a finish.
* fix(native-chat): fold a stopped turn as "Interrupted after N" and a failed or crash-cut one as "Failed after N"
The settled turn header now follows the verdict mark: a user's Stop reads "Interrupted after N", and
a failure or a turn anything else cut short reads "Failed after N", under the new key
components.native-chat.status.failedAfter in all six catalogs and the boot catalog. Desktop and
phone share the one description, so they agree.
* docs(agent-status): describe the Interrupted and Failed marks
The reference and the phone's turn bar still described a user's stop as done and a crash as
interrupted. A fault now reads failed, a user's stop reads interrupted in the muted tone, and the
rollups rank an unconfirmed end above a stop.
* test(status): a crash the relaunch recovers marks failed
The recovery test still expected a recovered interruption to mark interrupted; it now marks failed,
as a failure does. Formatting only elsewhere.
* fix(native-chat): record a turn a newer request replaced as superseded, and show it Interrupted
A Claude turn that a newer send replaced before its result arrived was recorded as interrupted with
no verdict, which reads as a turn cut short by something else, now "Failed". It is now recorded with
its own outcome, `superseded`, where the replacement is detected. That outcome names no sender, so a
dispatch from another agent is never recorded as the user's Stop, and it sets no legacy flag.
Every reader handles it in an exhaustive switch: it draws the muted Interrupted mark with the plain
text "Interrupted", folds as "Interrupted after N", and attention demotes it with a Stop, through
the renamed agentTurnEndedOnRequest. Older builds read an arm they do not know as no verdict, which
is what this turn carried before, so their rows keep reading done; the cross-version suites pin an
older desktop's journal and status readers and an older phone.
* refactor(status): name the attention predicate for a turn ended on purpose
agentTurnEndedOnRequest becomes agentTurnEndedOnPurpose: a user's Stop or a newer request's
replacement, never a fault. The Claude turn-end comment no longer says a replaced turn carries no
verdict.
* test(native-chat): a turn cut off by a restart or by quitting Orca reads Failed after N
On main a restart-cut turn shows the done tick. Pin the chat's turn bar and
the tab's mark for both cuts, through the recovery settlement and the quit's
child-end mapping, and pin the quit's turn bar through the host's own quit.
* style(native-chat): format the superseded turn-bar expectations
* fix(claude): a Stop that names no turn is the user's stop of the open turn
The chat's Stop button names no turn. Claude's conversation Stop recorded the
user's stop only for a named turn, so an older CLI's error result after that
Stop read Failed. It now records it on the open turn through the same intent,
dropped when Claude refuses the interrupt and never carried to the next turn.
* fix(native-chat): keep the attach context's publishStatus required
The lifetime context type makes publishStatus optional, so the attach context
that spreads it no longer satisfied its own type once its duplicate
publishStatus went. The host's lifetime context is now inferred, and checked
with satisfies, so the spread carries the member it always sets.
* fix(activity): rank a user's Stop below live work in the status grouping
The Activity page's status grouping put the Interrupted group (a user's Stop, or a
turn a newer request replaced) above Working and Monitoring, so a Stop still sorted
like news there while the sidebar, worktree card and Cmd+J rank it below live work.
It now follows live work and stays above Done; Failed and Couldn't confirm keep
their places above live work.
* docs(agent-status): say which turn outcomes the journal records and which are derived
The journal now records superseded as well as the provider's verdict and a stop;
interruption and unconfirmed are derived on read. The resume row no longer claims
interrupted renders red.
* test(native-chat): the idle sweep's held-send rest closes with the evict cause, like its siblings
---------
Co-authored-by: Claude <noreply@anthropic.com>
* docs(security): add the antivirus clearance path for future releases
Every AV false positive here has been handled one vendor and one shipped
version at a time. Document the programs that clear future releases instead --
signer and product enrollment rather than per-build sample submission -- and add
a script that reports an RC's current detection state by hash, so a verdict is
found before users meet it in an issue report.
Hash lookup only by default; --upload transmits the artifact and stays manual.
* fix(windows): replace the managed CLI launcher with a native one
resources\bin\orca.exe was a csc-compiled MSIL assembly: a small, freshly
compiled .NET image in a user-writable directory that mutates environment
variables and proxies a child process. That is the shape .NET dropper
heuristics are trained on, and every verdict against it named the family --
MSILHeracles from two vendors, Wacatac!ml from a third. Signing the file does
not change its shape, so signing never cleared it.
Rebuild it in Rust. Same resolution, same environment contract, same argv
passthrough that keeps newline-bearing orchestration bodies intact (#8374), and
the child still inherits our environment block rather than an explicit map, so
a block carrying both PATH and Path survives (#12046). The PE now carries
publisher, version, icon and an asInvoker manifest from build.rs.
Refs #23383
* ci(windows): install the Rust toolchain before building the CLI launcher
The hosted runners happen to ship cargo, but a real Windows dev box does not --
verified on our own Windows QA host, where cargo and rustc were both absent.
Relying on the image means a future image change fails deep inside
electron-builder's native hook instead of at an obvious step.