mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 16:02:29 +00:00
4e64fa9940f5a4a9bfb7b021a3c405fd76643ca5
12825
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4e64fa9940 |
fix(native-chat): start a Codex chat without waiting on the background model probe (#23831)
* fix(native-chat): start a Codex chat without waiting on the background model probe Opening a Codex chat kicks a session-less model-catalog probe (a throwaway read-only `codex app-server`) for the account, and the new chat's own options read then joined that probe's in-flight refresh through the catalog store's single-flight. When the probe's Codex hung, chat start waited out the probe's 15s deadline and failed with "codex app-server session exceeded 15000ms", even though the chat's own app-server was up. Single-flight now joins only a refresh by the same kind of lister. A live session lists over the connection it already holds; the probe still records its own failure in the store, and the picker keeps serving whichever listing succeeded. * fix(native-chat): never let a chat's model listing join another chat's listing Single-flight in the model catalog store was split by lister kind, so a new chat's acquire-time options read no longer joined the session-less probe, but it still joined any other live chat's in-flight listing for the same account. When that other chat's Codex was wedged or closing, the new chat waited out that chat's request timeout and failed to start with its error. Key in-flight listings by the lister's identity instead: a live session by its own connection, the probe by the probe itself. A lister still joins its own in-flight listing, and shouldRefresh still holds back a probe while any listing for the account is in flight. * test(native-chat): pin that the catalog store releases an account once every listing settles shouldRefresh reads 'any listing in flight for this account' from whether the per-account in-flight map exists, so that map must be dropped exactly when its last listing settles. Nothing covered this: removing the cleanup, or dropping the map on the first settle, passed every catalog test. A leftover map would stop every later background refresh and probe for the account. * fix(native-chat): name the catalog lister type instead of a bare object A live session is keyed by its per-spawn catalog handle (minted in the same acquire as its connection), the probe by itself. * test(codex): pin that a chat's own concurrent option reads share one listing * fix(native-chat): keep model picker current across parallel listings |
||
|
|
020cebeff6 |
Add standalone Agent Client Protocol client layer (#24990)
* Add standalone ACP protocol client and session runtime * Protect ACP transport teardown from late stream errors * Retire incoming ACP request ids before publishing responses * Narrow ACP configuration requests and transport message types * Remove redundant ACP request handler return unions * Keep ACP waits caller-owned and preserve protocol extensions * Preserve open ACP decisions through prompt completion * Generate open ACP enums and check the generated schema offline A newer or vendor enum value (tool kind, tool status, option kind, stop reason) no longer fails the whole message: generated enums accept the known literals plus any other string, typed so callers can still narrow on the known ones. The generated header now records the pinned input digests, the generator digest and a body hash, so `verify:acp-protocol` catches a stale or hand-edited file without network access; it runs in lint and the PR workflow. * Land the ACP runtime contract the agent adapters use - Deliver notifications other than session/update through onExtensionNotification, in arrival order with session updates. - Accept _meta on prompt, setMode, setModel, setConfigOption and cancel. - cancel() always sends session/cancel once the session runs, since the agent can be in a turn it began itself; only a successful send is shared, so a failed write is retried. - Cancel aborts each open agent request's signal and lets its handler send its own answer; -32800 only when the handler rejects. - Permission requests validate only the session, tool call id and options; unreadable fields are dropped with a diagnostic, and any answer Orca cannot send is `cancelled` instead of a JSON-RPC error. Agent-started turns may ask; whether to show it is the caller's decision. - AcpAgentError marks the agent's own errors; AcpInvalidResponseError keeps the raw answer and validation issues for answers Orca could not read. - Lines over the size limit are classified by prefix (shared with the Codex reader): the owed request fails, an oversized agent request is answered with an error, and an unattributable response closes the connection. * Answer every agent request after an ACP cancel A cancel that lands before a permission handler starts now still runs the permission path, so the agent gets the `cancelled` outcome rather than a request-cancelled error. A handler that ignores the abort no longer leaves the agent waiting: once the abort has run through, any request still unanswered gets request-cancelled. Handlers that answer on abort keep their own reply. Also renames a lint-rejected helper parameter, replaces a Reflect.apply in a test, and stops the permission diagnostic from firing with an empty list. * Let each ACP request handler own its answer after a cancel Removes the next-event-loop-turn fallback that answered request-cancelled for any handler still silent after a cancel. It raced answers that were still being saved (an approval mid-journal-write reached the agent as an error) and made the outcome depend on event-loop timing. The handler that owns an agent request now always sends its answer, or throws for request-cancelled; a request it never answers ends when the connection closes. A permission whose handler had not started still answers `cancelled`. * Register the ACP schema verify step in the PR preflight phase test * feat(acp): a steer's cancel asks once and never ends the agent The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle, then close the connection, which ends the agent. A steer used it too, so a slow agent lost its process just because the person added a message. requestSteerCancel() now sends session/cancel once per prompt, cancels the agent's open requests and answers later permissions cancelled, and never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel() stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel paths move into acp-prompt-cancel.ts over one cancel channel. * fix(acp): a repeated steer shares the cancel in flight; say what the caller owns Per review: a second steer before the first write lands returns that write instead of resolving early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that fails instead must not take the steer until the caller rebuilds the session; the Stop's says a prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives the runtime a handler that would allow: the open permission's signal aborts and the late one never reaches it. * test(ratchet): require src/main/acp now that this PR lands it |
||
|
|
635e7f53cb |
feat(github-projects): render Board project views as a kanban with drag-and-drop (#19074)
* Add board layout support for GitHub project views Board-layout views render as a kanban. Columns come from the view's verticalGroupByFields (the host retries without the selection on older GHES schemas and the renderer falls back to the Status field), with one column per single-select option in option order — empty ones included — plus a trailing no-value column whose drop clears the field. Card drops reuse the table's field mutation path, committed from a document-level capture listener because the preload's native-drop bridge stops drop events before React's root ever sees them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(github-projects): harden board capability probe and drop lifecycle Review follow-up. The verticalGroupByFields capability probe now matches parsed GraphQL error messages instead of substring-scanning the whole response body — partial-error responses echo the field name as a data key on healthy schemas, so one SAML/FORBIDDEN partial error could permanently degrade github.com boards for the session. Covered by new project-view-config tests per the capability-cache testing contract. Also: only single-select/iteration vertical fields shape columns (a drifted field kind no longer yields a clear-on-drop no-value column), optimistic patches resolve the column field through the board/group config, drag cleanup moved to a document-level dragend listener, the column dot follows dark mode via the chip CSS variables, the supported- layout allowlist is a single shared predicate, and the board's edit handler is stable so its drop listener stops re-registering per render. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(github-projects): serialize board edits and verify rendered drops * fix(github-projects): preserve refresh baselines and order edits across views * fix(github-projects): accept source settings projection in cache scope * fix(github-projects): distinguish view switches from refreshed field baselines * test(github-projects): type board IPC recordings and verify clear request --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
ca4e239861 | Remove low-value test inventories and duplicate fuzz oracles (#25791) | ||
|
|
c0b07d3a71 |
Admit one provider event's journal writes as one queued operation, decided when it runs (#25141)
* Move the turn message ordinals and the turn-row revision to the neutral timeline folder Pure moves so a shared timeline assembler can use them: Codex's message ordinal counter becomes ProviderTurnMessageOrdinals and Claude's turn-row revision becomes the provider-neutral agent-journal turn-row revision. Only names and import paths change. * Admit one provider event's writes as one transition, and let rows be found again after a restart - A sink transition is admitted whole or not at all; its steps run back to back at their turn in the journal's write queue, and each resolver reads the fold with every earlier write landed. A resolver may also say where the row belongs (turn scope, provider reference), and the writer always hears how the transition landed. A resolved lifecycle batch chooses its settlement mutations from the fold at execution. - New optional row field providerItemRef: the provider's own reference for the item a row is, written only where the row's identity cannot spell it (Codex keys messages by their place in the turn and renumbers its item ids on resume). Set by the creating write, kept by revisions, indexed by the journal fold, never read by clients. A downgrade test shows an older host and client render such rows unchanged. - Provider timeline identity schemes (shared legacy arm, Codex) and the join index that resolves a provider item to its row from memory or the fold: ordinals and request incarnations are read back from the rows, so a restart or an evicted entry finds the original row instead of placing a new one. * Recover message ordinals from the journal's highest place, and forget joins read from a replaced epoch A fresh join index continued a turn's messages at the first free place, so a journal holding only a later ordinal (an imported or removed earlier row) had its sequence back-filled. The place is now one past the highest ordinal any row or echoed send holds there, read through a pure scheme reader. The join caches also drop what they read when the journal's epoch is replaced. * Spell the subagent thread's message slot without spreading an identity union * Keep the journal store under its line limit after the main merge * Drop the provider item reference, join index and identity schemes from the transition PR Nothing in production reaches the state they defended (an assembler that lost its memory while its child keeps streaming the same turn), and the stored Codex id was positional. The legacy identity scheme moves to the assembler PR with its first caller; the Codex scheme and any persisted reference wait for Codex to move onto the assembler. The Codex ordinal counter goes back to codex/, since no neutral code imports it. * Write a resolved settlement in one transaction through enqueueRows A settlement too large for one row now commits all its rows or none, through the journal's existing all-or-nothing write, instead of a row-by-row writer. Every row is built before any commits, so a settlement naming one item twice is refused before anything is written. * Drop the transition's landing report; keep the turn-row write fire-and-forget Nothing reads which steps of a transition wrote. A failed step fails the sink, leaving the steps before it written; the header says so, and tests cover it plus a settlement whose second row fails inside the transaction. writeAgentJournalTurnRow returns nothing again, as on main. * Run a transition's steps as a prefix; drop the paced flag and resolved options A failed step no longer lets the steps after it write: each step checks, at its own turn in the journal's queue, whether the write handed over just ahead of it completed, using the queue's count of completed write bodies (a promise would report the failure only after the next step ran). The sink fails only once every step has had its turn. The item step's `paced` size bypass and the resolver's replacement `options` are removed; nothing planned uses them. * Run a transition's steps in one queued write that loops over them The steps of one event now share one turn in the journal's write queue: a loop writes each in its own transaction through the row writer's synchronous writeRows (split out of enqueueRows) and stops at the first throw. Prefix semantics and "nothing lands between the steps" now hold by construction, so the completed-write counter on the queue, the step gate and the allSettled barrier are gone; the queue is back to main's bytes. * Use current provider handles in transition tests |
||
|
|
dc8b3934b2 |
Add chat appearance settings: text size, code size, width (#25657)
* feat(native-chat): add chat-scoped color tokens * feat(native-chat): soften transcript and composer appearance * fix(native-chat): refine code spacing and faint text styling * fix(native-chat): wrap prose links at word boundaries * test(native-chat): refresh background task strip snapshots * Add native chat appearance settings and card scaffold * Persist chat text and code sizes with appearance width controls * Fix chat appearance typing, shortcut routing, and scoped typography * Use chat call identities in typography regression fixture * Keep markdown metadata typography independent of chat text size * Preserve terminal undo chords and stabilize chat resize estimates * Show only primary shortcuts in chat appearance settings * fix(native-chat): preserve appearance edits and configured zoom shortcuts * Remove unused chat appearance summary translation * Preserve command classification when applying chat code size |
||
|
|
6c693edf40 |
Show native chat tool calls as plain sentences (#25654)
* feat(native-chat): add chat-scoped color tokens * feat(native-chat): soften transcript and composer appearance * fix(native-chat): refine code spacing and faint text styling * fix(native-chat): wrap prose links at word boundaries * test(native-chat): refresh background task strip snapshots * feat(native-chat): show tool calls as plain sentences * fix(native-chat): make tool sentences reflect call state * fix(native-chat): clarify failed commands and subagent sentences * test(native-chat): exercise command disclosure with real results * fix(native-chat): preserve command inputs and localize failure rows * fix(native-chat): retain complete padded command input * fix(native-chat): read named tool input fields safely * Refresh native chat tool rows when the UI language changes |
||
|
|
8d2e9264bc |
fix(native-chat): ignore the terminal launch command when starting a native chat (#25720)
A custom launch command in settings made structured native chat unavailable: the renderer route, the host launch-mode decision, and the host's create-support all treated it as terminal-only, so a user with the chat default on silently got a terminal. The command names the CLI binary and now applies to terminal launches only; native chat ignores it. A start directory outside the workspace still forces a terminal. The internal route field and blocker are renamed to that concept; the wire reason `tui_launch_command` is kept and its receipt text now names the start directory. |
||
|
|
54ded3bc18 |
fix(relay): commit the cell counter in one round trip; cells boot without the database (#25765)
* fix(relay): commit the cell counter in one round trip; cells boot without the DB Step 2 (option B) cell image: - One-round-trip counter commit at acquireActivity, releaseActivity and activateControl: the final counter UPDATE and COMMIT go as one simple-query message. Server errors mean COMMIT never ran (retry as today; 22012 = no row, rolled back and disambiguated outside the transaction); a lost connection is never retried. - Cells skip the schema apply and region backfill, so they listen while the database is down and turn ready on their first successful query. - G13: rehome target connection headroom folded into the existing NOWAIT UPDATE, excluding the host's own reservation by key. - fixLevel on every runtime metrics line, plus declared (not applied) cell fix-level metrics and alert. - Per-desktop drain disconnect-gap measurement from existing log lines. - Census test that fails on floating database promises; fixes two shutdown sites. Lock-wait sample keeps the combined role. * fix(relay): make the outdated-image alert creatable: one PromQL condition, 1 h lookback, fixed floor A PromQL condition must be the only condition in its policy, and alerts on log-based metrics may look back at most 25 h. Replace the 6-day/7-day design with relay_cell_min_fix_level (tfvars, raised by a targeted apply after each wave) and one query: a serving cell below the floor or reporting no level, sustained 6 h. Drops the separate without-level metric. * fix(relay): review fixes: gap-script ordering, wider promise census, fused-path guard, row-busy as scheduled - Drain gap script: sort closes by time (gcloud exports newest first) and refuse an invalid drain start. - Census: any floating promise in relay src, including callback-discarded and never-read ones, with a reviewed never-rejects list. - Test the fused counter commit through the store the server builds, so a wrapper that stops forwarding commitWithFinal fails CI. - Same-cap shadow gate: a row-busy refusal (the host's own release still holds its row) is a scheduled 503, like an own early retry. No client change. * test(relay): judge drain redials by no host refused twice, not a refusal count The row-busy count tracks how many releases are still in flight at the dial (80 of 180 every run at 1 s, against a bar of 90). What matters is that the release has finished by the next dial: assert no host is refused twice, keep the time-to-placed p95 bound. * fix(relay): cap row-busy as scheduled at the drain-return admissions; bound the gap script's window Shadow gate: a row-busy refusal of a drained host follows its drain-return lane admission, so per minute only that many (plus a rounding margin of 2) are scheduled; the rest stay non-drain, so row contention the drain does not explain still fails the budget. Gap script: --drain-ended-at excludes the new container's closes after the roll; later grants still close a gap. |
||
|
|
fff8718c76 |
Improve Quick Open matching, file locations and recent history (#25371)
* Match Quick Open queries across path terms and identifier separators * Honor ignored-file and symlink preferences on file inventory hosts * Open pasted file locations and remember successful workspace file visits * Verify bounded directory listing preferences across execution routes * Preserve symlink preferences on legacy directory fallback * Preserve bounded Quick Open terms and separator alternatives * Revoke stale Quick Open selections and preserve literal file locations * Validate recent Quick Open candidates independently on their host * Use complete old-host inventory fixtures for Quick Open compatibility * Keep Quick Open validation current across palette and workspace changes * Route recent candidates through relay file listing dispatch * Record Quick Open hook renders as test snapshots * Localize Quick Open eligibility errors and search aliases * Validate file inventory with bundled search and runtime preferences * test: use directory junction fixtures on Windows * fix: negotiate Quick Open policies with nested SSH hosts * fix: refresh cached linked folders after target changes * Retry stale expanded links after reads settle and targets recover * Scope stale refresh request lifetime to the visible workspace * test: seed linked-folder fixture before watcher startup * Keep Quick Open responsive and compatible with older hosts * Exercise directory discovery with unprivileged Windows junctions |
||
|
|
c8ff7e8875 |
Stop repeated recovery of finished workers (#25679)
* Settle completed worker assignments and bound recovery retries * Reset queued explicit recovery before starting its retry budget * Use current provider identity in the inherited journal fixture * Keep late worker recovery automatic without repeating persistence work * Repair inherited validation fixtures after updating main * Keep the released parser available to compatibility tests * Limit recovery queries and background process inspection to pending work * Verify historical repair safety and live SSH process retention * Preserve the same provider import as main in the merge result * Reduce historical repair and duplicate workspace recovery work * Align startup prompt fixture with upstream Windows coverage * Keep healthy worker repair indexed and read-only * Align startup prompt tests with current main * Keep pending terminal releases progressing during recovery retries |
||
|
|
e2ccf52dc3 |
fix(remote-runtime): a paired terminal accepts input again after its host app relaunches (#25736)
* test(remote-runtime): reproduce dead input after a paired host relaunch A relaunched desktop host accepts RPC before its renderer publishes a window graph. During that gap session.tabs.list answers with an unpublished empty graph (publicationEpoch "none", snapshotVersion 0, tabs []). The client's reconnect inventory treats that as removal and retires the pane (retireRemoteTerminalId(-1)): the pane goes to "ended", the reconnect overlay disappears, and keystrokes never reach the surviving daemon PTY. Both tests are red on main by design; they are the repro for the fix. - e2e: holds the relaunched host's runtime:syncWindowGraph so the client's reconnect deterministically meets the unpublished host (4/4 red). - unit: transport-level repro of the same retirement (red in ~2s). - restart helper gains a beforeFirstWindow hook; LaunchOptions moves to its own module to stay under the max-lines budget. * fix(remote-runtime): a paired terminal survives its host app relaunching After the host app quit and relaunched (daemon still running), a paired client's terminal cleared its reconnect overlay and then ignored all input. The relaunched host answers session.tabs.list before its renderer publishes, with a synthesized empty frame. A paired client sees that frame through its navigation projection as epoch "none:client-navigation", which the shared "does this frame answer for the worktree" check did not recognise, so the pane read it as removal and retired itself. - hostSnapshotAffirmsWorktreeContents treats the client-projected placeholder as no answer at any version (the projection adds its navigation revision). - The reconnect inventory keeps polling on such a frame instead of retiring; a published frame lacking the surface still retires the pane. - Pushed frames that affirm nothing no longer report the surface absent. - When the bounded inventory wait ends without evidence, the first published snapshot carrying the same handle now reattaches the pane, rather than waiting out the ~3 minute auto-recovery deadline. Client-only: the host already labels the frame, old hosts send the same frame, and the wire is unchanged. The e2e helper drops its type assertions. * refactor(remote-runtime): simplify the host-relaunch reconnect fix - One epoch rule: the placeholder epoch, bare or with the shared client-navigation suffix, is no answer. Hosts only ever send it at version 0, so the version check is dropped. - Unpublished frames are dropped where pushed snapshots enter, instead of a nullable per-subscriber update. - A published same-handle snapshot fires the parked retry through retryNow(), which now also fires a retry parked while still 'recovering' (nothing is in flight then). This replaces state-reading wiring in the listener and lets online/resume fire it too. - The Windows e2e runs the fixture through PowerShell's call operator. * test(remote-runtime): prove the e2e hold beat the relaunched host's first publish The spec now asks the held host for its tab list and requires the unpublished placeholder before releasing, so a hold installed too late fails instead of passing against unfixed code. The helper keeps its gate in a typed global instead of Reflect lookups. * fix(remote-runtime): keep a fenced handle's parked retry revivable A retry parked for a handle that needs a replacement was consumed by an early external trigger and then refused by the epoch's snapshot-wait guard, leaving nothing to revive the pane. An external trigger now takes over that snapshot wait as a fresh attempt, and a republished fenced handle no longer fires the retry. Also refreshes comments that still paired the placeholder epoch with version 0. * fix(remote-runtime): a host answers an empty worktree for real once it has published The host sent its unpublished placeholder for any worktree it had no entry for, including one its window had never opened, long after startup. With the placeholder now read as "ask me later", a paired client opening such a worktree never got its first terminal. The host now sends the placeholder only until its graph first publishes. Afterwards a worktree with no tabs gets a real empty answer, and a client that was told "ask me later" during startup is sent that answer once on publication. Against an older host the client withholds the automatic first terminal for such a worktree, which the user can still create. A host push now fires only a parked retry, never cutting a scheduled backoff short. * chore(reliability-gates): reference the empty-worktree tests in gate commands * test(remote-runtime): find the never-opened worktree by repo id on every platform |
||
|
|
661372d514 |
Cancel abandoned file searches and prevent stale results (#25370)
* Bound relay git-grep records and clean up capacity failures Credits @OrcaWin for the original bounded-record proposal. * Bound filesystem listing and transfer metadata at the execution host Apply mobile limits before transport, list Markdown through its semantic producer, and retain complete directory results only within explicit capacity budgets. Stream SFTP directory packets and reject oversized transfer plans before reporting success; preserve narrow older-peer fallbacks and validate streamed response retention. * Release inactive Markdown candidates and support folder scopes Keep one current document snapshot per consumer and attach completion candidates to the actual editor model lifetime. Resolve folder workspace roots through existing workspace identities so their previews and completions receive the same authorized listing as worktrees. * Preserve full runtime inventories under aggregate byte budgets Leave unqualified inventories complete beyond 20,000 files. Charge retained paths and serialized bytes at local and relay producers, reject oversized full inventories explicitly, and bound old-peer response reassembly; keep caller-specific mobile and explicit result limits. * Bound SFTP directory handle cleanup waits Send CLOSE even after cancellation, stop waiting after five seconds without an acknowledgement, and ignore late callbacks. Preserve the original capacity or cancellation error and close handles returned after an aborted OPENDIR. * Preserve registered root spelling in Markdown document paths * Reject incomplete SFTP directory cleanup and final cancellation * Bound pending response bytes before stream ownership arrives * Preserve bounded directory reads on legacy SSH hosts * Allow bounded per-chunk padding at response byte boundaries * Bound response payload retention after stream metadata * Bound complete legacy Quick Open inventories and cache retention by bytes * Localize Markdown document listing fallback * Exercise relay search decoding with real byte streams * Preserve concrete relay stream fixture types * test: compare distinct bundled ripgrep platforms on Windows * Preserve Markdown discovery across Windows and legacy SSH hosts * Make WSL process termination available to the bounds layer * Cancel abandoned file searches and preserve search keyboard navigation * Invalidate searches when workspace roots or owners change * Cover root replacement and legacy search cancellation together * Exercise search cancellation through the current desktop access seam * Keep search ownership and cancellation contracts current * Share local text-search execution and close canceled subscriptions |
||
|
|
f59194d481 |
Bound file-search inventories and release abandoned Markdown data (#25369)
* Bound relay git-grep records and clean up capacity failures Credits @OrcaWin for the original bounded-record proposal. * Bound filesystem listing and transfer metadata at the execution host Apply mobile limits before transport, list Markdown through its semantic producer, and retain complete directory results only within explicit capacity budgets. Stream SFTP directory packets and reject oversized transfer plans before reporting success; preserve narrow older-peer fallbacks and validate streamed response retention. * Release inactive Markdown candidates and support folder scopes Keep one current document snapshot per consumer and attach completion candidates to the actual editor model lifetime. Resolve folder workspace roots through existing workspace identities so their previews and completions receive the same authorized listing as worktrees. * Preserve full runtime inventories under aggregate byte budgets Leave unqualified inventories complete beyond 20,000 files. Charge retained paths and serialized bytes at local and relay producers, reject oversized full inventories explicitly, and bound old-peer response reassembly; keep caller-specific mobile and explicit result limits. * Bound SFTP directory handle cleanup waits Send CLOSE even after cancellation, stop waiting after five seconds without an acknowledgement, and ignore late callbacks. Preserve the original capacity or cancellation error and close handles returned after an aborted OPENDIR. * Preserve registered root spelling in Markdown document paths * Reject incomplete SFTP directory cleanup and final cancellation * Bound pending response bytes before stream ownership arrives * Preserve bounded directory reads on legacy SSH hosts * Allow bounded per-chunk padding at response byte boundaries * Bound response payload retention after stream metadata * Bound complete legacy Quick Open inventories and cache retention by bytes * Localize Markdown document listing fallback * Exercise relay search decoding with real byte streams * Preserve concrete relay stream fixture types * test: compare distinct bundled ripgrep platforms on Windows * Preserve Markdown discovery across Windows and legacy SSH hosts * Make WSL process termination available to the bounds layer |
||
|
|
b58f8197dc |
Soften native chat colors and code surfaces (#25651)
* feat(native-chat): add chat-scoped color tokens * feat(native-chat): soften transcript and composer appearance * fix(native-chat): refine code spacing and faint text styling * fix(native-chat): wrap prose links at word boundaries * test(native-chat): refresh background task strip snapshots |
||
|
|
bfd1e9c579 |
fix(codex): keep working status when Escape closes search or permissions (#25769)
* fix(codex): preserve working status when Escape dismisses a view * test(codex): cover navigation Escape during terminal exit cleanup |
||
|
|
3ee3a41b6f |
fix(markdown): strengthen dark table grid lines (#25653)
Strengthen only dark-mode table cell borders in Markdown Preview and the rich editor by mixing foreground into the existing border token. Keep light mode and table geometry unchanged. Fixes #24982. Consolidates #24984, #25050, and #25653. Co-authored-by: Paramon <andrii.paramonov@gmail.com> Co-authored-by: kana001-bit <288527232+kana001-bit@users.noreply.github.com> Co-authored-by: Neil <neil@stably.ai> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
5f7c8483e9 |
test(cross-version): load releases that import a package main dropped (#25731)
* test(cross-version): stand in for packages a release declares but main dropped The newest release (v1.4.221) imports @streamparser/json, which main removed. The harness runs release source against the current install, and the release's RPC dispatcher imports its whole method table, so every wire suite failed at import time. Extraction now gives each package the release declares and the current tree no longer declares a stand-in in the checkout's own node_modules that loads but throws, naming the package, on any use. Co-Authored-By: Claude <noreply@anthropic.com> * test(cross-version): decide stand-ins by package resolution, not manifest diff A package gets a stand-in only when release source imports it at runtime, the current tree does not declare it, and it does not resolve from the checkout. A package still reachable from the checkout keeps the real install, and one the current tree declares but cannot resolve stays a loud failure. Type-only imports, comment prose and scripts in strings are ignored. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
3285f8214c |
Share the managed process lifecycle for structured providers (#25204)
* fix(native-chat): a child's root exit is reported even during its close, and bookkeeping after it never reads as unproven - Both connections report the root process's exit once, with `expected` set when a close had begun. A close that came back unproven and whose root exits later is finished by the adapter, and its end reaches the host like any other. - A Claude close whose resume-point write fails after the exit was proven, and a Codex close whose terminal row is refused, now end the session and report the failure, instead of keeping a dead child indexed as if its exit were unproven. - A Codex close whose forced tree kill can't prove the descendants gone but saw the root exit reports the descendants and counts the root exit. - Every child exit with an identity, expected or not, is forwarded to the host. * fix(native-chat): the exit ends the child's record; an unfinished stop is the child's own close, which everyone joins - The host keeps no stored "stop still owed" record any more. A stop begins the child's close (`child.close`), which lives on the child and ends with it. A second Stop, the idle reaper, quit, a send and an option/answer/goal/rewind all join that close instead of retrying a separate obligation. - A caller waits on the close only as long as the step deadline; the close itself is never abandoned. A proof that lands after every caller stopped waiting reaches the host as the adapter's report of that exit, which ends the record through the same handler. - Once the exit is proven, draining, settling, the lease release and the adapter's acknowledgement are each attempted and reported on failure; none keeps the child on record. A start, and the handle's close, write a release that failed from this host's proof of that exit, so a failed write never refuses a send. - A start that meets a close still unverifiable is refused with `previousExitUnverifiable`, so the queued message is rejected with a send-again reason; nothing is held and nothing starts beside the old process. - The idle sweep goes back to idle reaping only. - Removes #24333's retry entry points, the wait row and its hold rule, the ask/failure cursors on the stored record, and the stop's own wake. Tests replace the #24333 unproven-stop test: a send joining an unproven close and an in-flight one, a late proof past the caller's bound, a root exiting after its close gave up, a proven exit whose resume-point write and lease release both failed, an unverifiable close rejecting the send and refusing an option change, a surviving descendant, quit and the idle reaper; and Codex's unverifiable, late-exit and joined-close cases. * fix(native-chat): a message refused because the old process's exit is unverifiable says so, and to send again The start failure for a refusal with reason `previousExitUnverifiable` reads "Orca couldn't confirm Claude's previous process ended. Send your message to try again." instead of "Claude couldn't restart." The status-row kind and the refusal reason stay in the shared lists for rows and hosts that still carry them; the catalogs keep one sentence for both. * fix(native-chat): a close's verdict is the root's exit alone, and what follows it is logged - A Claude close resolves as soon as the root's exit is proven: the session ends and its `ended` report goes out then. Saving the resume point runs afterwards and a failure is logged, so a slow or hung write never reads as an unproven exit or keeps a dead child on record. - A root that exits after its close came back unproven finishes that close through the same path as any close, so the session's child work is published as ended (background tasks and subagents no longer stay shown running for a dead agent), and a failure there is logged. - Codex logs a refused final row, and reports a root exit whose forced tree kill could not prove the rest of the tree gone the way Claude does, so the host logs it and blocks nothing. - Both adapters take the host's logger for this bookkeeping. * fix(native-chat): one handler ends every child's exit, and a join waits on the adapter's own close - One exit handler (`structured-agent-session-child-exit`) ends a child's record for an exit expected or not. `expected` only changes what the chat is told: the stop's cause, its end at the stop's ask, the settlement id, and no crash outcome row. The lease release keeps the exit's evidence; the handoff guard, lifecycle barrier, sink release and adapter acknowledgement apply to both. A Claude journal-sink failure ends in the same step as its stop, as Orca's own fault. - Joining a close is asking the adapter, whose close is memoized while it runs and bounded by its own kill escalation; the host keeps no attempt of its own and no 10 s caller bound. An ask after a close came back unproven runs the stop again. - A close's end is stamped where its stop was asked for (a repeated ask moves it), so the closed chat and failed start checks order a message accepted meanwhile after it. - A start refused because the old exit is unverifiable rejects what was queued in the same step. - The end of a close the host asked for no longer waits on the cross-session recovery chain. - The kill no longer waits for the stop event's write; the journal writes rows in order. * fix(native-chat): an exit's lease release lands whatever the length of its reason A crash's reason can carry kilobytes of the provider's stderr, and a lease whose death detail is over 512 characters fails the store's own check. The exit handler cut it, but the release a start or the chat handle's close re-derives did not, so after a crash whose own release failed every message was refused as not resumable until restart. The record's builder now cuts the detail to the record's bound, so no writer can hand it one too long. * fix(claude): a proven close waits at most 2 s for the output it already wrote Once the root's exit is proven, the close still waited for the SDK's output reader to end. Something outside the process tree that holds the output open would keep that close, and every send, Stop and quit joining it, waiting with no bound. The wait is now bounded; past it the close resolves as proven and the open output is logged. * fix(codex): an exit reported inside Orca's close keeps the reason Orca closed it for The connection reports the app-server's exit inside the close that ends it, so that report ended every Codex close and replaced the close's own reason (for example, a provider frame that could not be recorded) with the connection's stderr text in the ended record and the lease's exit evidence. The session now records Orca's close with its reason, and the exit it ends keeps that reason. The test connection reports its exit inside close the way the real one does. * fix(native-chat): quit stops delivery before it drains exit recovery Every exit now wakes delivery, and teardown drained exit recovery before it stopped delivery, so an exit settled in that window could start a fresh agent that teardown then killed. Teardown stops delivery first; queued messages wait for the next launch. * docs(native-chat): the unverifiable-exit refusal no longer names a caller's wait The caller's bounded wait was removed; the comment describes the close as it is now. * fix(native-chat): a stop whose kill did not take is logged, and the next ask kills again When a close's kill leaves the agent's root running, the host now logs it. Tests pin what a later ask does: each connection runs its whole stop again (Codex sends SIGKILL a second time), refuses input meanwhile, and proves the exit once the kill takes. * fix(native-chat): a start refused over the old process says Orca couldn't stop it The host reaches an unverifiable verdict only after its own kill left the agent's root running, on the machine that runs the agent, so the sentence now says that: "Orca couldn't stop {agent}'s previous process." The refusal reason, failure kind and wire shapes are unchanged. The host test also checks the failed kill is logged. * docs(native-chat): an unverifiable close verdict is a root that survived the kill The host's close runs where the agent runs, so lost contact never yields this verdict; the comment no longer says it does. * fix(native-chat): a kill that did not take is reported once, by whoever met it The log added at the close fired beside a Stop's own failure report for the same event. A stop still reports it through its failure; a send or option change refused over it now logs it at the refusal, the only place it is otherwise invisible. * test(native-chat): a second Stop joins a close the first could not prove and retries its kill * fix(native-chat): say a start refused beside an unstopped process plainly The rejection now reads "Couldn't stop {{agent}} from before. Send your message again to try once more." This kind has its own send-again step; every other failure keeps "Send your message to try again." * Move provider process supervision and stream reading out of Codex * Preserve teardown behavior with checked mock types after move * Apply provider launch environment and caller teardown labels * Share managed provider launch, exit observation, and close * Preserve synchronous stdin closure before the grace wait * Keep the stacked Claude adapter within the module size limit * Handle nullable provider stdin and processless fixture identity * Preserve close-time cleanup diagnostics after a late provider exit * Separate provider root and descendant exit observations * Give provider child env one owner and gate Codex contract on the shared reader resolveProviderChildEnv is now the only place that overlays and strips a provider's environment; the spawn spec and the request-scoped Codex session both call it. supervisedPosixLaunch only accepts a launch without env fields, so an override can no longer be silently ignored there. Edits to the shared stream reader or the env rule now run the real-binary Codex contract job. * Give every provider one close result, stderr tail and root-only default - A close reports the root and the descendants with the one verdict vocabulary (live / unverifiable / exited); `tree` is null when the close made no descendant observation, and the fallback teardown's outcome is written into it instead of a side flag a caller could miss. - The managed process drains stderr and keeps the 8 KiB tail, so no provider can forget to drain the pipe. - Root-only providers take the default close policy and completion rule; only Claude overrides them. The one already-exited guard lives in the managed close and is recorded when no earlier close ran. - A failed spawn is never read as an observed root exit. * Report only observed descendant exits from the fallback teardown The fallback teardown answered "accepted", and the close turned that into `tree: 'exited'`, though a Windows tree kill's outcome is never read and an unreadable process table observes nothing. It now returns what it observed: `exited` only when the captured descendants were verified gone (or the kernel reported the group empty), `live` / `unverifiable` as verification found them, and null when it signalled without observing. Codex's process-tree diagnostic fires in exactly the cases it did before. * Give two test fixtures' casts a SAFETY rationale for the changed-code gate * Claim no observation from an empty process group ESRCH from the dedicated-group signal says only that the group is empty: a descendant that left it, or a root that never led it, may still be running. It now reports no observation instead of `exited`. The unreadable-table comment says why that case stays "no observation" for root-only providers. * fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts On main every close of the chat writes its own Stop and settle. Here a later close joins the first and writes no row, and the first's settle closed when its kill failed, so a turn that opened in between and was cut by the next close read as failed. A person's close joining a person's close whose Stop opened a settle now reopens that settle until its attempt is done. Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and the next reads as the person's cancellation (each fails without its half of the fix). * test(native-chat): build the Stop-opened-turn test's identity with the opaque handle The test (#25056) landed before the opaque provider handle (#24991), so main still built the old {kind, threadId} handle. * test(ratchet): require src/main/provider-process now that it has landed * test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test Main's #25706 made the same fix as this branch at a different line; the merge kept both imports. |
||
|
|
3a03441580 |
refactor(native-chat): structured agents declare their capabilities instead of shared code naming Claude and Codex (#25076)
* refactor(native-chat): keep the provider resume handle opaque to shared code
Shared structured-chat code parsed each provider's resume handle: Claude's
session id and branch leaf, Codex's thread id, through a 'claude' | 'codex'
union every new agent had to widen. The in-memory handle is now
{ transport, agent, nativeId, providerData? }: shared readers use nativeId,
lease and handle-chain checks compare transport and agent, and only the
Claude adapter reads its leaf (providerData).
Stored and wire forms are unchanged for Claude and Codex. One encoding
module writes their typed shapes and decodes both those and the neutral
shape a new transport uses, which an older build refuses as unreadable
rather than reading as Codex. Key and root strings, which fork seeds,
superseded creations and resume offers persist, stay byte-identical.
The journal's own handle type becomes the journal-row and attach-wire
encoding of the same handle, and the journal identity carries the
neutral handle (null before the provider proves one).
No user-visible change.
* fix(native-chat): derive journal-row provider handles from the journal identity
The journal row converter now takes the identity every caller already holds,
so a row's handle has one obvious constructor. Tests that wrote the in-memory
handle straight into journal rows now build it through that converter, and the
processless Claude fixture names a not-yet-proved handle as null.
* fix(native-chat): refuse a stored provider handle written in both forms
A typed Claude or Codex handle that also carries the neutral form's
transport, agent, native id or provider data named two identities; it
was read as Claude or Codex and the next write dropped the other one.
Such a row now stays unreadable and is set aside untouched.
* refactor(native-chat): route structured agents through registered definitions
The structured-chat shared layer named Claude and Codex in a closed router
type, in option reads, in frame classification, and in renderer label and
validation branches, so adding an agent meant editing shared types. Each
agent now has a definition (capabilities, resting option rules) owned by its
own module; the router is a registry of definition + adapter pairs built at
runtime setup, and shared code reads the definition instead of the name.
No behavior change for Claude or Codex; no wire or stored shape change.
* refactor(native-chat): name the structured agent list once in host types
The host types, conversation commands, provider-session ownership and the
runtime's create path spelled the structured agent list inline; they now use
the one shared type the record and wire already validate against.
* fix(native-chat): narrow the record before reading its agent's option rules
* refactor(native-chat): make the router's registrations the only agent definition lookup
The at-rest option reads looked definitions up in a second, separately populated map, so a
registered definition could route and declare capabilities while the static map decided which
options a chat at rest accepted. The router now answers `definition(agent)` from its
registrations, the at-rest readers take it through the host's adapter, and the built-in
definitions are listed only where the runtime composes the router.
* test(native-chat): use opaque handle in queued rejection fixture
* test(native-chat): share one Codex journal identity in the integration suite
Main grew the suite to the 800-line limit; the opaque-handle import pushed it
over. The two tests built the same identity inline.
* refactor(agent-session): name the handle's adapter state resumeCursor
Rename the neutral provider handle's providerData to resumeCursor before any
row persists the neutral form: it is an adapter-owned resume position (Claude's
transcript leaf), never identity. Claude/Codex stored and wire bytes are
unchanged; their typed shapes never carried the field.
State the stored-form contract (a handle's field set is closed; later per-link
data goes on the chain link, which every build preserves) and pin it with a
record round-trip test. Document that transport records the id space the
native id was minted in, which can differ from the agent's current transport.
* refactor(agent-session): one required agent registry; declarations admit what they claim
A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.
/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.
Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).
* refactor(agent-session): the router applies the declared rewind itself
The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.
* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record
* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way
The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.
* test(ratchet): require src/main/provider-process now that it has landed
* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test
Main's #25706 and this branch both added the import at different lines; the merge kept both.
|
||
|
|
db078b6549 |
fix(claude): API-key Claude users are told they're not signed in (#25163)
* fix(claude): open native chat for API-key users instead of saying "not signed in" Claude reports tokenSource "none" beside apiKeySource "ANTHROPIC_API_KEY" when it runs on an API key (environment or settings env). The startup check read only tokenSource, so it refused those chats as signed out. Refuse only when Claude reports no token source and no API-key source. Co-Authored-By: Claude <noreply@anthropic.com> * test(claude): cover a Console /login key at chat start --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
7f2541373a |
Stage long agent launch lines where they are typed, so they arrive whole (step 2 of 7) (#23962)
* feat(agent-launch): host-side prompt delivery for agent.launch The host's agent.launch typed any launch prompt into the shell as part of the launch command. A long or multi-line prompt then ran line by line in the shell, and an agent that never showed readiness or crashed at startup had nothing guarding where its text went. agent.launch now carries a prompt on the typed line only when the line stays one line, control-free and at most 512 bytes; otherwise the agent starts clean and the host pastes the prompt once the agent's own ready signal fires (bracketed paste plus its composer marker or a quiet render, read only after the shell's last hand-off, never while the pane's own shell is proven in front), with main's draft-paste bytes and an Enter 50 ms later. Orchestration worker starts wait on tui-idle as before. A replay-safe launch admits and claims its ledger row in one write, Qwen Code gets a second Enter, the desktop and phone share one launch-refusal classifier, and hosts advertise agent.launch.prompt-carry.v1. Split out of #23748, which moves the desktop source-control buttons onto this path. * fix(agent-launch): keep a short-lined multi-line prompt on a local zsh launch line, as main did #24257 moved every multi-line or over-512-byte prompt off the typed launch line and pasted it after readiness. The phone's AI buttons and review notes, whose multi-line prompts main typed whole into zsh, then reached Claude 0.5-3 s later and their RPC reply waited for the paste. The host now names the shell a local macOS or Linux line is typed into, the way the spawn picks it, and a multi-line prompt rides a zsh line when every line is at most 512 bytes and the whole line at most 8 KB. A real-zsh test types such a line through Orca's own ready barrier and startup write, including when a slow user config makes the write land early. Elsewhere the measured unsafe cases keep the paste: bash 3.2 runs multi-line lines piecemeal, fish drops an early multi-line write, and any shell loses a line over 1 KB written early. * test(agent-launch): keep the real-zsh launch-line test out of the Windows lane's gate scan The Windows lane registration check read `const ZSH_PATH = process.platform === 'win32'` (the head of a multi-line ternary) as a Windows-true flag, so `describe.skipIf(!ZSH_PATH)` looked like a Windows-only suite. The file is POSIX-only; the zsh lookup is now a function. * refactor(protocol): move the agent.launch capabilities into their own module Main's protocol-version.ts sits at the 300-line cap, so the prompt-carry capability pushed it over. The four agent.launch capabilities and their doc move to agent-launch-runtime-capability.ts, re-exported by name and spread into RUNTIME_CAPABILITIES at the same position; the advertised lists and every export are unchanged. * refactor(protocol): import the agent.launch capabilities from their own module `export *` from protocol-version left the four names undefined under the mobile recording loader, which resolves a relative import through a Proxy with no own keys, so 37 phone recordings lost agent.launchReplay. Importers now name agent-launch-runtime-capability directly; protocol-version only spreads its list. * refactor(agent-launch): drop the unshipped viewMode field and trusted local caller id Both were inert in step 1 and existed only for step 2. agent.launch will become a public plugin API, so every wire field is permanent once shipped; a top-level viewMode reads as "choose terminal vs chat", which the host decides. Step 2 introduces placement and view intent under a placement object instead. * fix(agent-launch): read Codex's provisional startup from the rule files' hold anchor Main (#24375) moved Codex's provisional-header check into codex.json's provisional_startup hold anchor and deleted codex-terminal-readiness.ts, so the launch readiness hold now asks showsHoldAnchor, as main's own settled check does. * fix(agent-launch): hold rule-file name titles to quiet for a launch, and census the zsh fixture Main (#24375) answers a name-only title from each agent's rule file ahead of the sustained-title lane, so gemini.json's name_title settled a launch readiness wait on the shell's auto-title while Gemini was still booting. A launch now asks quiet of every weak idle verdict, as that lane did. Main's readiness census requires a recorder for every runtime fixture; the zsh prompt recording is a non-agent control. Gemini's synthetic baseline is regenerated for this PR's stated change: a bare gemini title is no longer its rest mark, so name-only rows settle weak, and a fresh working or blocked status is no longer overridden. * fix(launch): stage long or multi-line launch lines so they arrive whole A launch line that is long or has several lines could lose lines or wedge a shell that is still starting, because it was typed into the terminal before the shell could read it whole. The terminal host (local daemon, and the SSH relay) now writes such a line to a script file and types one short line that runs it, for every POSIX shell; a shell that cannot read Orca's quoting runs Orca's own agent lines through /bin/sh. If the script cannot be written, the terminal says so. The SSH background launch path no longer types the line from the renderer; the relay stages it like any other terminal. Bumps the terminal daemon protocol to 40 (main took 39); a v39 daemon keeps its sessions and gets no staged line. Restacked onto #24257 from the version reviewed on top of #23748; carries none of #23748's changes. * fix(launch): run a staged launch line as its own job in fish and ksh fish and ksh run the commands of a sourced file in the shell's own process group, so a staged agent was not a job: Ctrl-Z was dropped (a typed line stops), and the shell-in-front check saw the shell's group in the foreground while the agent ran. fish now runs the script with `eval (string collect < '<path>')`, which runs it as if typed and leaves the shell's job-control mode alone. ksh and mksh run a long Orca-built line through `/bin/sh '<path>'`, a real job; a long command the user wrote is typed as before. The real-shell test now asserts the agent's process group is its own and is the terminal's foreground group. The staging rule also reuses hasControlByte and the 512-byte budget from startup-line-prompt-carry so the carry and staging limits cannot drift. * fix(agent-launch): paste a launch prompt only when the launched agent is proven in front A launch pasted its prompt unless a shell was proven in the terminal's foreground, so any read that could not prove one let the prompt through. After an agent exited at startup, its shell turned bracketed paste on at the next prompt, readiness fired on it, and the prompt was typed into the shell: - macOS: a pane runs its shell under login, so the process-group fence's root was never the shell's group and never proved it; the cached foreground name could also still name the exited process. - Windows Git Bash and WSL: the shell-alone-in-its-job check never answers. Now one fresh read of the terminal's foreground decides: agent, shell or unknown. Only 'agent' lets a write through (paste, Enter, second Enter, reused panes too); 'shell' still drops a ready signal. A Windows host never proves the agent, so there the launch line carries the prompt at any size, as on main. * test(agent-launch): cover the Windows QA stub, a grok override that exits at once * fix(agent-launch): keep the local socket alive while a prompted launch waits for its agent A launch with a prompt now waits up to 60 s for the terminal agent to be ready before it writes the prompt, and reports not-delivered when the agent never is. The local runtime socket closes a connection idle for 30 s unless the request is a long poll, so a launch whose agent exited at startup lost its reply and the caller saw 'runtime closed the connection' instead of not-delivered. Classify a prompted agent.launch and agent.launchReplay as a long poll, as orchestration.workerStart already is for the same wait. * refactor(agent-launch): narrow the launch params by 'in' instead of a cast * fix(agent-launch): find a launched agent behind a wrapper that leads its process group A tcsh or nu launch line runs the agent from /bin/sh '<script>', and a wrapper script that does not exec its agent does the same: the wrapper leads the terminal's foreground process group and the agent is a member of it. The fresh foreground read names the group's leader, sh, so a prompted launch was refused or pasted late (M4Air tcsh: 2 of 4 not delivered, 2 pasted ~9 s late). Before that read, take the host's process-group observation as positive proof when it names the launched agent among the foreground group's members and is younger than a ready signal's quiet window. It never proves a shell. * fix(launch): type a plain Orca agent line as is in tcsh and nu In tcsh and nu, every Orca-built agent line ran through `/bin/sh '<script>'`, so the agent was a child of sh in sh's process group. When the prompt was pasted after the agent was ready, the paste guard read sh in front and, on a loaded Mac, reported the prompt not delivered (4/4 live runs in tcsh). Now only a line that needs it goes through `/bin/sh`: one over 512 bytes, with a control byte, or with a character those shells would not read literally (`!`, backslash, `"`, `$`, backtick). A plain line such as `claude '--dangerously-skip-permissions'` is typed as on main, so the agent is the shell's own job. * fix(agent-launch): judge the foreground-group proof by when its capture began, not how long ps took The age the host stamps on a process-group observation runs from the start of its whole-machine ps, so on a loaded Mac a capture begun after the read was asked for still read as older than 1 s and the proof was dropped. Count an observation whose capture began after the read was asked for, less the window a shared capture is reused across. * test(agent-launch): keep the crash-guard live test out of the Windows lane's gate scan The Windows-lane registration scan read the const assigned from a platform check as a Windows-only gate, though the suite runs everywhere but Windows; find zsh in a function instead, as the real-zsh typed-line test does. Under load the fresh foreground scan can fail to answer, which lets the shell's prompt settle readiness (2 of 4 paired runs). The guard still refuses that write, so assert the refused write, the property that must always hold. * perf(agent-launch): read a local pane's foreground from its own terminal, not the whole process table The foreground read that gates every launch paste ran the daemon's inspectProcess capture and then a fresh scan, each a whole-machine ps; the fresh one also waits for any capture already running before it starts its own. Measured here at load 5: 1.2 s a read (M4Air QA: 3.4-5.0 s, and worker starts 17.6-32 s against main's 9-12 s at load 25-84). On a local macOS or Linux host, take the pane's root pid from the provider's session inventory and run one ps limited to that pane's terminal. Its foreground process group decides: the launched agent or any non-shell member is the agent (a wrapper that did not exec its agent leads the group), a group of shells alone is the shell. Same pane, same verdict: 2.7 ms a read. SSH hosts keep the relay's observation and name. * test(mobile): re-measure the web app's script sweep after agent.launch's capabilities moved out of protocol-version The mobile web bundle check failed at 124 assets against a ceiling of 123. Main already sat exactly on that ceiling: its sweep table read 69 scripts at 16 routes while the tree builds 73, the whole margin of 4. This branch imports the agent.launch capabilities from their own module, so protocol-version is no longer pulled into the root layout and four other routes. That moves which routes share which modules, and the Qoder capability module, imported by protocol-version and the AI-vault resume path, no longer shares an importer set with anything, so it gets a chunk of its own: 74 scripts. The fence says to re-derive the bound rather than raise it, so the sweep is re-measured on this head (every prefix of the sorted route list). The worst route now adds 10 scripts (session), not 9, which moves the pinned shell crossing from 32 to 30 routes; main re-measured on its own lands on the same crossing. * fix(agent-launch): a worker's brief needs its agent found in front, and Grok's start answers on its composer A paired-server worker start whose agent exited at startup typed its brief into the server's shell, which ran it: the idle wait can settle on a shell back at its prompt, and the brief was written with no foreground read. Both worker-start paths now check before each brief write, as a launch prompt is checked: on a host that can find the agent in front it must be there; on one that cannot (Windows) a shell proven in front still refuses, and anything else writes as before. A Grok worker start waited ~10 s more than main: its only rest signal is its bare name, which a launch holds to quiet output, and Grok animates its logo for ten seconds after its composer glyph. A worker start for an agent whose rest signal is its bare name and whose composer draws a marker (Grok, DSH, mimo-code) now also answers on that marker, whichever comes first. * test(startup-staging): expect the CR that submits a typed launch line since #23672 |
||
|
|
ccc0bf70e4 | Pause Pullfrog reviews while the CI runner queue recovers (#25760) | ||
|
|
059e81a106 |
chore(i18n): use 智能体 for Chinese Agent copy (#25767)
Simplified Chinese rendered the Agent concept as 代理, which collides with 代理 = proxy. Standardize on 智能体 for Agent (and 子智能体 for subagent), while keeping 代理 for proxy senses: HTTP/network proxy, SSH Proxy Command, reverse proxy, and browser user agent. - Converted 131 zh catalog values (incl. 子代理 -> 子智能体); 26 proxy / user-agent values left as 代理. - locale-phrase-fixes.mjs: 客服人员/代理商/座席 -> 智能体, 代理 -> 智能体 when the English names an agent (guard excludes "user agent"); removed the old 智能体 -> 代理 rule so the pipeline no longer reverts it. - Updated value/key/search/macos-tcc overrides to 智能体; proxy keyword and proxy override entries unchanged. - Updated the two policy tests that pinned the old 代理 output. Gates: catalog verify, coverage --check, extraction, runtime-catalog, and the locale vitest suites (296 tests) pass. |
||
|
|
9b7c936929 |
fix(dock): clear stale workspace unread counts (#24877)
* fix(dock): count only the unread the app shows Tab markers outlive the workspace unread flag, so the badge could sit at 1 with no sidebar dot to find or clear. Fixes #23363 * fix(dock): follow the sidebar's host filter in the unread count Counting folder workspaces and per-host rows by flag put rows on the Dock that a host-scoped sidebar does not show. * fix(dock): preserve owned folder alerts without expanding unread counts * fix(dock): honor the existing other-device worktree filter --------- Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> |
||
|
|
c7d74b1160 |
feat(relay): give Asia cell c34 a promotion wave so it can become a general cell (#25757)
* feat(relay): give Asia cell c34 a promotion wave so it can become a general cell c34 launched on 2026-10-05 as a migration-only spare with no promotion path. This adds it to the Asia admission promotion waves, the workflow's promote and canary cases, and the canary evidence map, so the reviewed Asia workflow can promote it with the same five-minute canary c30 and c31 ran. The same-cap migration-only list is deliberately unchanged: a same-cap job reads a cell's class from that list, and c34 must be rolled to the director's image as a migration-only cell before promotion can run. The list moves after promotion, in its own change. Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041 * docs(relay): scope the c34 same-cap pause to the window after promotion Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041 * docs(relay): rewrap the c34 paragraph Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041 |
||
|
|
0e9e273b91 |
fix(relay): deploy driver builds the checked commit, prompts on their own line, summarises progress (#25755)
* fix(relay): publish the reviewed commit before main can move, and quiet the deploy driver The publish workflow builds main's head at dispatch. The driver checked main at preflight but dispatched the publish about four minutes later, after the inspects and the typed phrase, so a busy main stopped the first real deploy. It now dispatches the publish seconds after the check, before anything else, and a build of a moved main stops with the --commit/--publish-run command that deploys it once reviewed. Typed prompts end in a newline, and runs are summarised (status changes plus every 5 min) instead of streaming gh run watch. * fix(relay): a main-moved stop prints only the command that reuses the build The generic re-run line named the reviewed commit without --publish-run, which would only build the moved main again. |
||
|
|
420447e063 |
fix(agent-status): retire a removed worktree's hook-status rows (#23881)
* fix(agent-status): retire a removed worktree's hook-status rows Rows whose terminal was never reattached had no teardown path, so they outlived the worktree in last-status.json for up to 7 days. Fixes #23068 * fix(agent-status): skip panes another owner has since reclaimed * fix(agent-status): clear only the removed owner's claims on a shared pane * fix(agent-status): retire hook-status rows a host scan proved removed `worktrees:forgetRemovedForExecutionHost` is the only path that ever retires an off-host WorktreeMeta row: gcStaleWorktreeMeta skips any row whose repo or hostId is not local. It already prunes the cleanup and space-analysis snapshots for a worktree a remote scan proved gone, but left that worktree's hook-status rows behind in `last-status.json` — the same stranding this branch fixes for the in-Orca delete, reached through the other trigger. A scan the host answered is positive evidence of removal rather than loss of contact, so it is the host evidence `ssh-execution-boundary.md` requires, and it publishes no verdict. The drop is scoped to the scanned host's id, so a same-id worktree on another connection keeps its rows, and a runtime host falls through the method's own early return because a paired server owns its own store. Also names the shared-pane condition in `dropStatusEntriesForRemovedWorktree`, which was the one place the two-owner logic was hard to read. * fix(agent-status): stop a removed worktree's row returning on a shared pane The mixed-owner branch deleted the removed worktree's row but left the pane unfenced, so the stale row came straight back. `getAgentStatusDisposition` returns `accept` for an unfenced pane, and the removed worktree's agent can still post a late turn on a pane it shares with another owner — I reproduced the `working` row being rewritten to memory and to `last-status.json`, which is the symptom #23068 is about. A launch-token fence cannot close this: `ingestTerminalStatus` calls the disposition gate with no event, so the token check never runs and an OSC report carries no token to check. The pane fence the sole-owner path already uses does suppress it, and it lifts on PTY reattach or a new agent's turn. Fencing the pane would otherwise discard the surviving owner's claims, which is what the branch existed to protect, so both of its records are captured first and restored after: its persisted authority commitment (its resume identity) and its current authority observation. The observation also fixes a second defect on this path — `deleteStatusEntry` drops it whatever `preserveAuthority` says, so the surviving owner silently lost `current_runtime` attestation until its next hook event. Both are covered by tests that fail against the previous commit. * fix(agent-status): decide a reused pane by its occupant, and retire rows for local scan-proved removals A pane has one terminal, so its newest row names who occupies it. A saved commitment naming a different owner is what an earlier occupant left behind, and serialization already drops it while the row disagrees. Removal now retires a pane the removed worktree occupies, and on a pane another owner occupies it clears only the removed worktree's outlived commitment. This replaces the capture-and-restore handling of two-owner panes. The local authoritative-scan prune is the local twin of the SSH scan forget: it drops a worktree's metadata once the scan proves it gone, and now retires its status rows too. * fix(agent-status): preserve foreign authority during worktree removal * fix(agent-status): preserve terminal connection validation * fix(agent-status): keep ordinary OSC admission unchanged * refactor(agent-status): colocate the remote envelope type * fix(agent-status): revoke removed startup authority evidence * fix(agent-status): retain foreign startup claims during removal --------- Co-authored-by: Neil <neil@stably.ai> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> |
||
|
|
6012570996 |
feat(files): reveal Source Control files and terminal file links in the file manager (#24502)
* feat(reveal): reveal changed files and terminal file links Reveal only existed on file explorer rows. Every reveal menu now shares one action, so a missing file or a remote workspace behaves the same in all of them. Fixes #24003 * fix(reveal): address review findings on missing links and SSH files A hovered hyperlink to a missing file saved "missing" in the cache plain path links trust, hiding the path once the file existed. A tab holding an SSH file from outside its workspace still offered reveal. * fix(reveal): drop reveal from check-details tabs Their path is a synthetic id, so reveal could only fail. * fix(reveal): use scoped file access for directory checks * fix(reveal): retire superseded Spanish menu labels * test: preserve explicit file access and untranslated-locale contracts * refactor(terminal): reveal a file link from its click popover, not the right-click menu The terminal "Reveal in Finder" action now lives in the file-link popover (the one with "Open file" and "Open with default app"), as a row with no click shortcut. It goes through the shared reveal, so a folder link (including a macOS .app or .xcodeproj bundle) is selected in its parent and never opened. The row is omitted, not disabled, when the file is not on this machine: an SSH or runtime-owned file, a pane whose shell runs on another host, or while a remote runtime is focused. Workspace-root links keep their own "Open in Finder"/"Open folder" row and get no reveal row. Removes the right-click path: the context-menu item, the hovered-link tracking and its OSC 8 / path-existence probe refactors, the directory stat (and its user-named file-access grant, which only served the open-a-folder behaviour), and the right-click e2e case. Terminal files the PR touched are back to main apart from the popover row. Co-authored-by: Kelvin Amoaba <97001695+AmoabaKelvin@users.noreply.github.com> * test(terminal): type the reveal settings fixture without assertion * test(startup): account for encoded Windows prompt lines * test(compatibility): retain the released journal parser dependency * test(journal): retain one canonical provider-handle import * fix(terminal): give the popover's Reveal in Finder row the reveal icon Every other Reveal in Finder item shows the external-link icon; the link popover draws it for rows that open outside Orca, so mark the row that way. --------- Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
a259616f7c |
feat(pi): show Pi and OMP background children as rows (#23201)
* feat(pi): show Pi and OMP background children as rows Refs #22868 * fix(pi): derive the extension's roster cap from the host's The generated extension re-typed the 32-row limit as its own literal, so a change to AGENT_STATUS_MAX_SUBAGENTS would leave the extension posting rosters the host silently truncates on arrival. Interpolate the host constant instead, and cover the cap with a test that posts past it. * test(pi): pin the Pi cancel verdict beside a live child row Publishing a subagent roster for Pi puts its panes behind the same child-work guard Claude and Codex sit behind, so Ctrl+C beside a live child no longer settles a stopped row. The extension reports no main-agent state, so Orca cannot tell a cancelled turn from Ctrl+C at the idle prompt of a lead children alone hold open, and keeps the live row. Both halves of that are now pinned, as is the resume-placeholder guard on the model_select arm of the same code path. * fix(pi): end a run with its session and keep children across reload Pi re-runs the extension on a fresh event bus for /new, resume, fork and /reload, so the roster kept on that bus was lost and nothing ended the run the old session's children held open. * fix(pi): preserve pending completion and OMP pane ownership * test(pi): use the hook owner context type * fix(pi): retain completion timers and delivery across session boundaries * test(omp): preserve unsent completion at session boundaries * test(journal): retain provider handle import across main integration * test(startup): account for encoded Windows prompt lines * fix(pi): serialize status delivery across module reloads * test(compatibility): retain the released journal parser dependency --------- Co-authored-by: Neil <neil@stably.ai> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> |
||
|
|
1b52be6255 |
fix(claude-accounts): preserve shared MCP OAuth credentials across an account switch (#21931)
* fix(claude-accounts): preserve shared MCP OAuth credentials across an account switch Orca's account switch writes the target managed account's own stored credential verbatim to the global "Claude Code-credentials" Keychain item. That credential never carries mcpOAuth/mcpOAuthClientConfig/mcpXaaIdp/mcpXaaIdpConfig/pluginSecrets (they are machine-shared MCP connector state, not per-account), so every switch silently drops any MCP connections the previous session had. Merge the live credential's copy of these shared fields into the target credential right before the Keychain write, live-wins (absence included), mirroring how the third-party claude-swap tool already treats this exact shared Keychain item. Fixes #16098 * fix(claude-accounts): skip reformatting shared credential merge when nothing changed mergeSharedClaudeCredentialFields always re-serialized via JSON.stringify, even when the shared-key set was identical on both sides. The reformatted-but-semantically-equal JSON (e.g. missing the original trailing newline) then read as an external Claude Code refresh to the read-back byte-equality check in runtime-auth-sync, causing every subsequent sync to wrongly adopt it into managed credential storage. Return the original string when the merge changes nothing. * fix(claude-accounts): validate OAuth shape, preserve key order, and degrade Keychain read failures Address CodeRabbit/pullfrog follow-up review on the shared-credential merge: - parseCredentialObject accepted claudeAiOauth as null or a primitive; the merge would then serialize the malformed target instead of passing it through unchanged. Validate it is a non-null object (hasClaudeOauthObject) before merging. - The no-op guard compared whole-object JSON.stringify output, so an existing shared key interleaved among target-only keys got moved to the end of the rebuilt object even when its value did not change, producing a formatting-only rewrite. Compare per shared key and keep the target's own key order via object spread instead. - runtime-auth-sync read the live Keychain credential for the merge with the raw (non-best-effort) function, so a transient Keychain read error aborted the whole account switch instead of just skipping the merge. Use the inherited readAggregateClaudeKeychainCredentialsBestEffort, matching every other Keychain read in this subsystem. - Fixed a dead citation URL (scaryghost/claude-swap 404s; the real repo is realiti4/claude-swap, verified to contain the cited SHARED_CREDENTIAL_KEYS / merge_shared_credential_fields). - Added docstrings to every function touched by this change. Tests: 5 new cases (null/primitive claudeAiOauth, interleaved shared key no-op and update, Keychain read failure does not abort the switch). * fix(claude-accounts): preserve connector state across all runtime credential writes Keep live MCP grants in both runtime files and keychain items, excluding shared secrets from managed account capture and refresh read-back. Preserve rotations and revocations through default restoration and reject unreadable live state before destructive writes. * fix(claude-accounts): retain grants from all live stores on startup Before Orca has a last-written baseline, a missing shared field in one store cannot prove revocation. Carry present fields from the other live stores, prioritizing a proved newer refresh candidate when available. * fix(claude-accounts): reconcile connector updates without guessing token freshness * test(preflight): freeze clock for Windows PATH probe assertion --------- Co-authored-by: tomarai85 <tomarai85@users.noreply.github.com> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
d73efccc7d | Reduce localization audit and relay setup work in CI (#25665) | ||
|
|
08a970a3f1 |
fix(native-chat): show the message rail from the first user message (#25707)
* fix(native-chat): show the message rail from the first user message The rail on the right of native chat stayed hidden until a conversation had three user messages, so short chats had no rail at all. Show it whenever there is at least one user message (still hidden in panes too narrow for it). Co-Authored-By: Claude <noreply@anthropic.com> * fix(test): remove duplicate journal fixture handle import --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
a47e0f5657 |
fix(feedback): optimize oversized screenshots without losing detail (#23664)
* fix(feedback): shrink oversized screenshots instead of refusing them
Fixes #22776
* fix(feedback): release each shrink canvas and skip the PNG ladder for a JPEG
Two defects in the attach-time shrink ladder:
- Every step allocated a fresh full-size canvas and left it for the garbage
collector, so a 32-megapixel source could keep six of them alive at once.
Past the renderer's canvas memory budget Chromium hands back a blank canvas,
which would upload as a blank screenshot with nothing to notice it. Each
step's canvas is now released once toBlob has answered.
- A JPEG source ran two full-size PNG encodes first. A PNG of decoded JPEG
noise is many times larger than the source it has to undercut, so those
steps could only ever burn time and a large buffer before the first JPEG
attempt. A lossy source now starts at the JPEG steps.
* fix(feedback): keep the image read queue usable after a settle callback throws
The queue tail is what the next batch awaits. A throw in either settle callback
left it rejected, so every later attach short-circuited into "Could not read the
attached images" without reading anything — and the rejection went unhandled.
* feat(feedback): say when an attachment is still being prepared
Shrinking a full-screen capture takes long enough that the gap between picking
the file and its thumbnail appearing reads as a dropped attachment: the only
other signal was the Send button quietly staying disabled. The attachment hint
now reads "Preparing attachments…" while a batch is in flight.
* fix(feedback): report an undecodable oversized image as invalid, not too large
* fix(feedback): detect an animated PNG by walking its chunks, not a 64 KB text scan
* docs(feedback): stop claiming toBlob encodes off the renderer thread
* refactor(feedback): seed the image read queue ref directly
* test(feedback): cover the real decode, downscale and release path of the shrink
* perf(feedback): skip image batches still queued when the dialog unmounts
* fix(feedback): show the preparing hint only once a read outlasts a short delay
* test(feedback): pin the animated-PNG walk on corrupt lengths and a refused image's budget
* fix(feedback): blame the shared budget, not the file size, when a shrinkable image has no room
* fix(feedback): cap each shrunk screenshot at half the attachment budget
A shrink kept the largest step that fit everything left, so the first
full-screen screenshot took a 3.8 MB PNG, the second fell to a barely
readable JPEG and the third was refused. Re-encodes now target at most
2 MB; an image that already fits is still attached untouched. The
"larger than 4.0 MB" refusal now applies whenever the capped target an
empty budget would give was the one that refused it.
* test(feedback): pin the half-budget shrink cap independently of its constant
* fix(feedback): reserve a shrinking screenshot's capped size, not its file size, for the paste gate
A 6 MB screenshot still shrinking reserved 6 MB, more than the whole budget,
so a second paste during the shrink fell through to the textarea even though
it attached. Each pending file now reserves what the reader can commit under
the same fit and shrink-target rules. Also corrects the minimum shrink target
comment to the measured smallest step.
* docs(feedback): correct the canvas-release and shrink-queue test comments; mock toast.info in the dialog test
* Revert "refactor(feedback): seed the image read queue ref directly"
This reverts commit
|
||
|
|
8127d94782 |
test(agent-launch): keep the carry control on a prompt that still cannot ride the line after #23672 (#25734)
#23672 quotes a multi-line PowerShell argument as one physical line, so the control's short multi-line prompt now rides the line on a Windows host that can prove the agent, and the unit shard went red. The control now uses a multi-line prompt over the typed budget, which still pastes, and a new test pins that a short multi-line prompt rides the PowerShell line as one physical line. |
||
|
|
4ccfc166d5 | fix(sidebar): preserve host filters across stale UI updates (#25737) | ||
|
|
9756daeced |
chore(i18n): translate 76 new keys to es/fr/ja/ko/zh (#25719)
Delta since last scan (
|
||
|
|
ea6c9d6f63 | fix(test): restore the codexProviderHandle import the #25078 squash removed again (#25728) | ||
|
|
dd39dcba5f |
feat(orchestration): a native chat gets the orchestration pointer a CLI agent gets, through the same send as your messages (#25078)
* feat(orchestration): agent mail to a busy chat waits in the chat's own queue
A mail notice for a busy structured chat used to wait in the orchestration lane's
own invisible "until the chat is free" gate. It now goes through the queue a
person's message uses: sendAgentTurn(..., { delivery: 'queue', source }) makes
the host hold it as a draft card the person sees and can Steer or delete, and
the queue sends it when the turn ends.
- The lane's busy gate (turnRunning / awaitingHuman) is gone; the queue's own
hold decides. Its parking now only waits for its own send or card to settle.
- A queued card's hand-off (sent under a fresh id) is matched by
queuedMessageId, so it stamps the mail once and never sends a second notice.
- A card from before an Orca restart is still the lane's: no second card.
- A card the person deleted counts as handled for that batch; newer mail
notifies again. No stored flag: read off the card and delivered_at.
- `orchestration check` waits for the lane to withdraw a card whose mail it
read, so a stale notice is never sent; more mail replaces the card.
- Each card records who queued it in queued_messages.source_json (versioned,
schema-validated): the person, or Orca for agents, with every distinct
sender as an orchestration party plus host id, and the mailbox, dispatch,
run and message ids. The restart pause holds only the person's cards.
- sendAgentTurn answers with a snapshot, not the journal's live submission.
* test(orchestration): type the mail fixture as a pointer batch message
* fix(orchestration): the queue judges an agent's mail notice as it sends it
A mail notice queued in a busy chat could go out stale or twice: it was
kept true from outside the queue, by orchestration check withdrawing it,
which missed other readers, raced an attempt in flight, lost the card
when /clear carried it (a second notice), and moved it to the back when
new mail replaced it. An agent's card behind a person's paused card never
sent after a restart, so an unattended coordinator stalled.
- The drain asks the card's sender, in its own serialized step, whether
it still stands: send, restate (count and senders of the mail owed
now, written onto the card in the hand-off's transaction), or withdraw
as the host. A failing judge sends as written. Send-now restates but
never withdraws. Orchestration registers the judge on the host.
- check no longer touches any queue (back to main); onMailRead,
notifyOrchestrationMailRead and the lane's withdrawals are gone.
- One unsettled card per agent message, enforced at insert; new mail is
counted when the card sends, never replaces or moves it.
- The lane finds its card by what it is (a notice for the mailbox in the
session the mailbox reaches now); a decline is a withdrawal a person's
operation stamped. /clear moves agent cards as the host. Pointer rows
last only while a direct send is in flight and end with the session.
- Pauses (restart, Stop, /clear) hold only a person's cards, and a held
card never traps an agent's card behind it, keyed on the message kind.
- The source is named for the message (agent-session-message-source),
one payload shape per message kind, no relative host id.
* fix(orchestration): Orca's mail notice waits in the queue out of sight
The notice card read as something the person typed and came and went on
its own, the transient-state message the queue should not show. It is
hidden until B can label who sent it.
- The host leaves mail-notice cards out of the published queue, so the
list, its count, Steer/Delete/Edit, steer-newest and the paused header
never see them, on every client. Keyed on the message kind: D6's task
card will be shown, labelled, in B. They still count toward the queue
limit (temporary, until B).
- A refusal no longer returns a hidden card for the person to act on:
the host withdraws it, and the lane points again only after a turn ran
since, the rule it already had for a refused notice.
- Send-now no longer restates a card: no one can reach a hidden one.
- The stored sender drops the pane key, a mailbox credential.
- The gate covers an abandoned worker: a mailbox that reaches no session
withdraws its notice (tested).
* fix(orchestration): a hidden mail notice never waits on the person
Hiding the notice left three paths where it waited on someone who
cannot see it.
- A failed hand-off write put a send_failed hold on it, which only a
person's Send or Delete clears: that mailbox's notices stopped for
good. The host now withdraws a hidden card instead and hands its
mailbox back to the lane through the ordinary redrive, since an idle
chat gets no other edge. A write that keeps failing is retried once
per edge.
- A person's Stop while the agent started on the notice put it back to
waiting, where no pause holds it, so it was sent again at once. The
host now withdraws it; the lane points again after a turn ran or when
newer mail makes it a different notice, as on main. A restart still
puts it back.
- A notice queued before the person's message went first. One order
rule now: the person's cards go in order and never past one waiting
on them; a card they cannot see never delays one they can, and goes
only when none of theirs may.
The drain moves to its own module (structured-agent-session-queued-
drain.ts). Comments that still described the notice as visible are
corrected.
* fix(orchestration): an accepted notice hand-off stamps its mail; a judge that cannot look decides nothing
- When the agent opened its mail in the notice's own turn (the normal
flow), an open batch left nothing owed, so the lane never marked the
mail the accepted hand-off carried as delivered, unlike an accepted
direct send. The lane now stamps exactly the ids a handed-off notice
carried whenever undelivered mail remains, owed or not. A later
conversation is no longer told again about mail already opened.
- At quit the host registry is cleared before teardown, so the drain's
judge could not resolve the mailbox and withdrew the notice. A judge
that cannot read its inputs now defers: the card waits for a step
that can look. A mailbox that resolves to another session or none is
still withdrawn.
- The judge, the owed-batch selection and the notice body move to
structured-mail-notice.ts. A failing hand-back of a dropped card is
logged.
* refactor(orchestration): a chat receives the agent's mail itself, queued like a person's message
A structured chat used to get a derived "You have N orchestration messages,
run check" notice, which could go stale while it waited in the chat's queue,
so the queue judged, restated or withdrew it as it sent. It now gets the mail
itself as the turn, through sendAgentTurn with delivery 'queue', the path a
person's message takes.
- One turn per mailbox batch: the mailbox's unread mail at delivery, in mail
order, each message led by "[message from <sender>]", its type, subject,
body, payload and reply hint (what check prints). Mail arriving while that
card is still in the queue waits for the next card; a card is never edited.
- An idle chat takes it at once; a busy one queues it as an ordinary card the
person sees and can Steer or delete. Mail is marked read when the chat
accepts the turn, so check does not return it again; a deleted card leaves
its mail unread for check, and it is not pushed again.
- The card stores who it is from (source_json): every distinct sender, and
each message's id, run and sender. Kind 'mail-notice' becomes 'mail'.
- Removed: the send-time judge (QueuedAgentCardVerdict), structured-mail-
notice, structured-pointer-notice-cards, queued-message-restatement, the
hidden-card rules, the agent-card exemptions from the restart, Stop and
/clear pauses, the dropped-card hand-back, and the drain split. An agent's
card now follows the same pauses as a person's.
- Terminal agents are unchanged: they keep the typed pointer and check.
* fix(orchestration): a chat's check skips mail already queued to it; restarts hold only the person's cards
- When a structured chat runs `orchestration check` (consuming, --peek or
--wait), mail an agent's card in that chat's own queue still carries
(any card not deleted) is left out: it is on its way as a turn, so the
agent does not read it twice. Derived from the queued rows at read time;
a deleted card no longer carries it, so it comes back. Terminal callers
and every other read are unchanged. New batches pass the exclusion to
getOrCreateMailboxDelivery; peeks filter it.
- A restart's queue pause now holds only the person's cards: an unattended
coordinator's agent card sends after an app restart without waiting for a
Resume, since the mailbox is the record and reading is marked. Stop and
/clear still hold agent cards. One rule, in queuePauseHolding. An agent
card queued behind the person's held card still waits behind it: the
queue never reorders.
- The duplicated direct-mailbox snapshot routing in check-run and
check-worker becomes one helper, which keeps check-worker under its line
limit.
* fix(orchestration): the mailbox is the only record of read mail; a chat's check takes its waiting cards
A card held an exclusive claim on its mail that nothing reconciled with the
mailbox: queueing it stamped the mail delivered, and a chat's check hid every
card that was not deleted. So a chat's own check --wait in one turn could not
see a result its waiting card held, a hand-off that ended in doubt or came back
"Not sent" stranded its mail, a check racing the card's build read the mail
twice, and a card left behind by run-use still sent and marked it read.
What a card or send holds is now derived each time from the chat's queue and
its sends: an accepted send marks its mail read; a waiting card or an
unanswered send holds its mail; everything else unread is pushed again or
returned by check. A card the person deleted is recorded as not to be pushed,
at the delete. The lane withdraws returned, partly read and moved cards. A
chat's consuming check withdraws the waiting cards holding what it reads, one
at a time with the lane. The sender is stored on the sent submission too
(host-only), so read state survives the card row's prune. A refusal before
anything started ends its operation; a restart pause is raised only by the
person's cards; an unreadable agent source stays an agent's.
* refactor(orchestration): a chat gets the pointer a terminal agent gets, sent through the chat's own queue
A structured chat now receives exactly what a terminal agent is typed: "You have N orchestration
messages ... run `orca orchestration check`" (formatMessagePointer, same CLI name). It goes
through the shared sendAgentTurn with delivery 'queue', the composer's queue-if-active: an idle
chat gets it at once, a busy chat's queue holds it as a normal card and sends it when the turn
ends. The lane's idle gate is gone; the send decides, as for the person's message.
The lane queues no second pointer while one of its cards still waits, read from the queued rows;
mail arriving meanwhile, or after the card is sent, is pointed again as the terminal lane points
new unread mail. A queued pointer counts as delivered, as an accepted one does. `check` and mail
read state are main's: only `check` reads mail.
The card records who it is from (queued_messages.source_json, kind 'agent', message
'mail-notice', with its senders) for the next PR to render. A restart's pause is raised by and
holds only the person's cards; Stop and /clear still hold an agent's.
Removed from the earlier designs: the message-as-turn batching, the check exclusion and taking,
the derived claims, lock and reconciliation, the sender on submissions, and their tests.
* fix(orchestration): a chat's pointer follows a card that leaves its queue unsent
The lane holds new mail while a pointer card waits, and retried it only on the chat's next
status change. A card that leaves the queue without a turn after it (the person deletes it
while Stop holds it, or its hand-off comes back) made none, so that mail sat unpointed until
something unrelated happened. The shared queue wiring now tells its host when an agent's card
stops waiting, read only when the queue changed, and the runtime redrives that chat's parked
mail. The runtime forwards the new host dep like the others.
Tests: that case end to end, and /clear carrying an agent card keeps who it is from. Agent card
bodies in tests are the pointer text; the restart pause's header says why an agent card may send.
* refactor(orchestration): no special handling for agent notices in the chat queue
A structured chat's orchestration notice is now sent through the chat's own send
like any message: an idle chat takes it as a turn, a busy chat's queue holds it
and sends it at turn end, under the same Stop, restart and /clear pauses as the
person's cards. Native chat no longer branches on who a card is from; the card
only records it.
- Restore main's queue pause logic (no restart exemption for agent cards).
- Remove the queue watch that redrove mail when an agent card left the queue,
and the host's queued-row read the lane used for it.
- The lane reads nothing of the chat's queue. It sends no second notice while
mail it already pointed at is unread, read off the mailbox alone.
* fix(orchestration): point new mail like a terminal, even while earlier pointed mail is unread
Drops the native-chat-only rule that held back a notice while the agent had not yet read mail it
was already pointed at. Also fixes main's stop-note test, which still built a provider handle in
the shape #24991 replaced.
* test(native-chat): keep the queued-message rig fixture under the line limit
* fix(test): drop main's duplicate codexProviderHandle import (same line as #25713)
|
||
|
|
41f1103c30 | fix(test): restore the codexProviderHandle import in the submission-positions test (#25722) | ||
|
|
9323443c83 |
fix(test): drop the duplicate codexProviderHandle import in the submission-positions test (#25713)
Two fixes for the same provider-handle change (#22682 and #25709) each added the import, so main fails tc:node with TS2300 (duplicate identifier). |
||
|
|
135c92e01f |
fix(native-chat): an unsent draft survives a reload or quit (#24905)
* fix(native-chat): an unsent draft survives a reload or quit The composer's draft text, its editor document and settled image refs lived only in module memory, so a reload, quit or crash lost whatever was typed, including text a Stop had just given back. They are now saved per pane scope in localStorage: typing is debounced and flushed when the window hides or closes, text given back to the composer is saved at once, an emptied or sent draft is removed at once, and only the newest drafts within the composer caches' existing bound are kept. * fix(native-chat): save draft images outside the state updater * fix(native-chat): one store owns each unsent draft The draft's text and its image chips were saved by two writers that each rebuilt the stored record from what storage held, so a draft that once went over the size cap, or one refused write, lost its images on the next save. One store now holds each composer scope's whole draft in memory and writes storage from it: typing is coalesced, returned text, chips and clears land at once, a refused write stays pending for the next flush, and all drafts together keep to a budget so the outbox can still save a send. Closing a tab drops its panes' drafts. * test(native-chat): returned text is stored before its images arrive * fix(native-chat): keep returned text and drop launch-seed copies across reloads A message given back from the outbox could be larger than one draft was allowed to store, so the saved draft was deleted and the returned text was gone after a reload. One draft may now use the whole draft budget, which holds the largest message a send accepts. An adopted launch link is also parked in the agent's input line. Restoring it after a reload, with no seed left to replace that copy, typed it twice, so the untouched copy is shown but not saved. A composer reused for another pane now shows that pane's images, so adding one never saves the previous pane's over them. * test(native-chat): a plain 260k given-back message stays saved * fix(native-chat): saved drafts drop pasted images and die with their worktree A pasted image lives in a temp file Orca wrote, and the permission to show it is held in memory only, so after a restart a restored pasted chip showed a generic icon, and once the temp file was cleaned it looked fine but the send failed after the composer cleared. Pasted images now stay in the draft for the current run only; images the user attached from their own files are still saved. Removing a worktree now deletes its terminal tabs' drafts, as closing the tab does. A paste that finished late in a composer that had already been replaced wrote that composer's old images back over a newer send. Dropping a placeholder chip no longer writes the saved draft. The pane-change reload in the images hook is removed: the composer is keyed by pane, so a pane change always mounts a new one. * fix(native-chat): the composer renders the draft store instead of keeping a copy A composer that had been replaced, for example by a question prompt, could still finish an attach or a send later and write its own old copy of the draft over the store, so a sent message or image came back. The composer now reads text and images from the draft store and changes the draft as it is now: an added image or file reference joins what is there, and a send that settles after a round trip, or after its composer was replaced, clears the draft only if it is still what was sent. Only chips still being written and an IME composition in progress stay local to the composer. Closing a tab now deletes drafts by the exact tab id: a second chat for one session has an id that extends the first one's, so a prefix match deleted its drafts too. * fix(native-chat): keep images pasted during a command and on shown composers An image pasted while /compact or /clear was still on its way to the host vanished when the command was accepted: the clear that followed also wiped the chip still being saved. A command that settles now clears only the draft it sent, and an image pasted meanwhile stays. A composer on screen could lose its pasted images, and its unsaved launch text, once 128 other drafts pushed its record out of memory, because those are never stored. A draft that a composer is showing is no longer evicted. * fix(native-chat): a draft being written is never evicted from memory A write to a pane while every other cached draft was on screen could evict the record it had just written, losing it. Also adds the test that a send settling after its composer unmounted clears the saved draft, through the real send hook. * fix(native-chat): restored images Orca couldn't keep come back to attach again A pasted screenshot in an unsent draft vanished after a restart with no notice, and a restored image whose file had been deleted looked fine until the send failed. Both now come back as a placeholder chip that names the image and says it wasn't saved with the draft. Send stays disabled, and the send button says to attach the image again or remove it. Attaching a file with the same name takes the placeholder's place. Two browser tabs of the web client on one chat now follow each other: when one sends or edits the draft, the other drops its copy and shows the new one, unless it has an edit of its own not yet saved. * test(native-chat): no send while an image waits to be attached again * fix(native-chat): pasted images in a draft come back after a restart A pasted screenshot lived in the OS temp folder with a read permission held in memory, so a draft could only bring it back as a placeholder. Local pastes are now written to Orca's own native-chat-pastes folder under the app's user data, and a draft keeps them as real images. On restore the composer asks the main process to re-grant each one; main grants a read only when the file's real path, symlinks and junctions resolved, is inside that folder, and reports anything else as not kept, which becomes the placeholder. Pastes older than 30 days are deleted at startup, far past the host's 24 hour window for a resent message; the sweep never follows a link and never blocks startup. SSH pastes and older temp-folder pastes still come back as placeholders. * test(native-chat): the paste sweep test ages the link itself * fix(native-chat): type the paste-folder path helper with PlatformPath * fix(native-chat): type the paste-folder path helper from path.posix * fix(native-chat): placeholder chips say what happened and only a user attach replaces one A placeholder now shows the image's name and a short hint on the chip itself, "Attach again" for a file and "Not kept" for a pasted image, with the longer explanation in its tooltip. The copy no longer says the image wasn't saved: it says it couldn't be brought back, and for a pasted image it asks only for removal, since a new paste can't match it. The send button now says to remove the image to send. Only an image the user attaches (picking, dropping or pasting) takes the place of a placeholder with its file name. An image Stop gives back is added beside it, so a different file with the same name no longer hides the reminder. * fix(native-chat): only composer pastes use the paste folder, and restore grants only real paths Every local clipboard image save had moved into the paste folder, whose macOS path has a space, so a screenshot pasted into a terminal or sent to a terminal-backed agent no longer attached. Only a native-chat composer paste goes there now, through an optional field on the existing save call; terminal, editor and phone pastes stay in the system temp folder as before. An image path sent to an agent's input is escaped the way a dropped image is. The restore re-grant now grants only the file's real path, and only when both the real path and the path as written are inside the paste folder, which must not itself be a link. The sweep deletes only Orca's paste files and skips a linked folder. * fix(native-chat): a restored paste is readable by the path its draft stored Restore granted only a paste's real path, but the composer reads the preview by the path its draft stored, so with user data reached through a link the restored chip showed a generic icon. The stored spelling is granted too; it is already proven to sit inside the paste folder. A test also pins that a composer paste asks for the paste folder. * fix(native-chat): grant a paste's stored spelling only when it names the same file A grant also covers its path's own real path, so a stored spelling that is itself a link inside the paste folder could have granted an outside file while the path as restored reached a real paste. The stored spelling is now granted only when it resolves to the same file; otherwise only the real path is, and the preview falls back to the generic icon. * test(native-chat): check the outside file by its real path too * fix(native-chat): a structured chat's draft belongs to its conversation Drafts of a structured chat were saved under the pane (tab id plus a hash of the session), so the follow-up that keys them by conversation would have left every draft saved by this build invisible after an upgrade, never shown and never deleted. They are now keyed by the conversation (`agent-session:<sessionId>`) from the start: every composer showing the conversation shares one draft, Stop and a queued card's Edit give text back to it, closing the tab keeps it, and removing the worktree deletes it. Terminal-backed chats keep the pane key and lose their draft when the tab is closed. A send now leaves whatever was added since it was sent, text typed or composed and images attached meanwhile, and clears only what was sent; a draft replaced in the meantime is left alone. Lifted from #25207 ( |
||
|
|
9f561703b7 |
fix(i18n): fix garbled Japanese name suggestion and reversed ja/zh quotes (#25255)
* fix(i18n): use opening quotes in ja/zh concatenated labels The Browser Use / mobile emulator prompt examples and the workspace-name row wrap text in separately translated opening and closing quotes, and ja/zh mapped the opening one to the closing mark, rendering 」text」 and ”text”. * fix(i18n): make the Japanese workspace-name row read in Japanese order The row concatenates Use / quote / name / quote / as workspace name in English order, so ja rendered 使用「name」ワークスペース名として. Swap the two word fragments so it reads ワークスペース名として「name」を使用 without touching the component; a full-sentence key (#9294) would also need the emphasized name split out of the translated string. * fix(i18n): give the grab sheet's accessible-name quotes separate keys Both quotes around the accessible name shared one key, so locales with distinct opening and closing marks could only render one of them twice (ja: 」name」, zh: ”name”). A separate opening key lets them render 「name」 and “name”. |
||
|
|
d3943c6d81 |
feat(editor): highlight Quarto and R Markdown files (#17772)
Highlight Quarto and R Markdown documents using the bundled Markdown, YAML, R, Python, and JavaScript tokenizers. Preserve guarded startup registration and bound inherited inline embed recursion. Thanks to Yoshihiko Kunisato (@ykunisato) for the implementation and regression tests. Fixes #17771 Co-authored-by: Yoshihiko Kunisato <12838333+ykunisato@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c6827fe02a |
test: drop the duplicate codexProviderHandle import two main fixes both added (#25717)
* test(ratchet): require src/main/provider-process now that it has landed * test: drop the duplicate codexProviderHandle import two main fixes both added |
||
|
|
b6c5f3f123 |
fix(native-chat): a sent image no longer flickers to its "Pasted image" placeholder while the agent works (#25664)
* fix(native-chat): a sent image no longer flickers to its "Pasted image" placeholder while the agent works The thumbnail's blob URL lease and its release effect were keyed on the runtime-context object, which the owner hook rebuilds with equal values on unrelated store updates. Each rebuild revoked the URL the <img> was showing. Key both on the image cache key instead. * test(native-chat): cover each half of the equal-context image fix; drop releaseLocalImageSrc Hook tests: an equal runtime-context rebuild keeps the pin and URL; an SSH reconnect or worktree change re-reads. Transcript test: an equal rebuild leaves an off-screen image's cached entry alone. releaseLocalImageSrc had no production caller left; tests now release by cache key, and the rich markdown test releases the key the editor actually leases. |
||
|
|
13d94acd45 |
feat(browser): add an eraser to screenshot markup (#24801)
* feat(browser): add an eraser to screenshot markup Click or drag across marks to remove them whole, one undo step per drag. Fixes #23258 * fix(browser): drop the in-flight markup gesture on undo and redo Undo mid-erase left marks hidden against the restored drawing, so release removed them and cleared redo. * fix(browser): settle a markup gesture whose release was lost A gesture now blocks any new press so a second finger cannot steer it. If its own release never arrived, that block also swallowed the next press from the same pointer, so the next drag continued the stale gesture: a stroke jumped from its old end, and an erase swept from the old point, removing marks along a line the user never touched. A pointer cannot press twice without releasing, so a press from the gesture's own pointer now settles the old gesture as its release would have, then starts the new one. * fix(browser): an eraser click takes only the top mark, a drag takes all A click on overlapping marks now erases only the one drawn last, matching what a user expects from pointing at a mark. Once the pointer really moves the press becomes a drag, whose first sweep starts at the press point, so every mark under it and along the path is erased. A move that reports the press point again does not turn a click into a drag. * fix(browser): undo mid-gesture takes back only the gesture Undo pressed while a stroke or an erase is still held now drops just that gesture, as the newest step; the committed marks are untouched, so the next Undo takes back the last of them. Redo and Clear drop the gesture, then apply as before. With the gesture gone, the rest of that drag neither draws nor erases until the next press. Undo is enabled while a gesture is in flight, since it now has something to take back. * fix(browser): keep an eraser tap a click within a 4px slop A touch or pen tap often reports a move a fraction of a pixel from the press point. The eraser treated any move to a different point as a drag, so such a tap erased every mark under the finger instead of only the top one. The erase gesture now records its press point and a pressed/dragging phase: it stays a click until the pointer travels 4 CSS px, then sweeps from the press point as a drag. * fix(browser): undo over an empty erase undoes the last mark Pressing Undo while an eraser was held over empty space only dropped the erase, which had hidden nothing, so Undo looked dead and the Undo button was enabled for a step with no visible effect. Undo now treats a held gesture as a step only when releasing it would change the markup (a stroke being drawn, or an erase that hides at least one mark). Otherwise Undo takes back the last committed mark and ends the erase; with nothing to undo it leaves the erase alone. The Undo button is enabled exactly when pressing it would change something. * fix(browser): redo with nothing to redo keeps a held markup gesture The Redo button is disabled when there is nothing to redo, but the Cmd/Ctrl+Shift+Z shortcut is not. Pressed while a stroke or erase was held down, it replaced the document with itself and dropped the gesture, so the stroke or erase silently stopped. Redo now leaves the state untouched when there is nothing to redo. Clear still always drops the gesture. * fix(browser): discard a markup gesture whose pointer was cancelled pointercancel was wired to the release handler, so when the system took a touch or pen pointer away mid-gesture (an OS gesture or palm rejection), the stroke was saved and an erase was applied as if the user had lifted deliberately. pointercancel now discards the in-flight gesture owned by that pointer: the stroke is not saved and the marks an erase was hiding come back, with no undo step. Lost pointer capture and a press that reveals a missed release still commit, since both stand in for a release that did happen. * test(browser): pin that losing the pointer commits a markup gesture The canvas commits a stroke or erase when it loses pointer capture without a release, so a gesture can never be left open. No test covered that wiring: removing it, or routing it to the cancel handler, left every test green. The new test ends a stroke with only the lost capture and checks that Undo then leaves it to redo, which only a committed stroke does. * fix(browser): do not save a markup shape pressed without dragging A Rectangle, Ellipse or Arrow press released without dragging paints nothing, but it was saved as a zero-size mark. With the eraser, a click takes only the topmost mark, so that invisible mark made the next eraser click on a visible mark at the same spot appear to do nothing. It also added an undo step with no visible effect, and while such a press was held Undo was enabled but only dropped the invisible shape. A shape with no size is now not saved, and a held unmoved shape press is not something Undo takes back. Pen and Highlighter taps still save the dot they draw; text is unchanged. * test(browser): drop a cancelled-stroke test the overlay test already covers MarkupOverlay.test.tsx drives the same press, move, pointercancel and following lost capture through the real canvas wiring and asserts Undo stays disabled, which already proves nothing was committed or held. * docs(browser): note the erase drag re-rasterizes the markup layer The committed layer now renders the visible marks, so an erase drag re-rasterizes it once per newly hit mark; the comment said never. --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
b01df814e6 | test(ratchet): require src/main/provider-process now that it has landed (#25710) | ||
|
|
e43a9f729d |
fix(test): use the current provider handle shape in the submission-positions test (#25709)
#25073 built the test's journal identity with the retired { kind, threadId } provider handle, which the node typecheck rejects (TS2353). Same fix #25706 made in the stop-note test. |
||
|
|
1787585b97 |
fix(daemon): bump the terminal daemon protocol to v40 (#25291)
* fix(daemon): bump the terminal daemon protocol to v40 Daemon-side changes landed under v39 after v1.4.220 shipped v39: the #25130 and #24636 shell-wrapper changes and a wider agent list (jcode, qoder-cn, dsb; resume claims for qoder-cn, qwen-code, cursor and jcode). Without a bump, an update keeps new tabs on the running v39 daemon, which misses the wrapper fixes and rejects those resume claims with agent_session_identity_required. * fix(local-builds): record daemon protocol v40 in the local build compatibility contract |