mirror of
https://github.com/stablyai/orca.git
synced 2026-10-06 16:02:25 +00:00
4bab736f9097317ad6c4bcb90f99d00a4bc4e31c
11645
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4bab736f90 |
fix(native-chat): dock the task strip on the goal tab (#22530)
* fix(native-chat): dock the task strip on the goal tab The background-task strip was the full width of the message box, so it did not share an edge with the narrower goal tab underneath it. When a goal is showing, the strip now uses that tab's width and keeps a square bottom, and the goal tab's top stays square so the two sit on each other. * fix(native-chat): derive the task strip and goal tab seam from adjacency The strip and goal tab were each told by the chat session whether the other was showing, through two flags that had to match what actually rendered. The session also re-derived the goal tab's own visibility rule to compute one of them. Now each bar styles its side of the seam from the DOM: the strip takes the goal tab's width and drops its bottom corners and shadow when the goal tab is its next sibling, and the goal tab drops its top border and corners when the strip comes right before it. The flags and the duplicated goal visibility check are gone, so the seam cannot disagree with what renders, and anything placed between the two bars falls back to the separate look. |
||
|
|
63866c1e27 |
fix(mobile-web): page inputs lose the browser focus ring and hairlines draw one device pixel (OTA phase C follow-up) (#22569)
* fix(mobile-web): drop the UA focus ring from page text inputs Chromium rings every focused text field (:focus-visible); no native TextInput paints one. A zero-specificity rule in its own inline block beside the Expo root reset removes it for every page input; buttons keep the browser's ring. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile-web): draw page hairlines one device pixel thick react-native-web pins StyleSheet.hairlineWidth to 1 CSS px, three device pixels on a 480 dpi phone; native draws one. A build shim replaces that one assignment with React Native's own formula (roundToNearestPixel(0.4), else 1/ratio), so every page hairline matches native without touching components. A rendered check at a real device scale (Playwright's emulated scale floors borders to CSS px, which no phone does) measures both parity fixes. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile-web): apply the hairline shim to react-native-web's CommonJS build The page's dependencies require react-native, so esbuild resolves every importer to react-native-web's dist/cjs build, which the previous filter did not match: the shipped bundle still assigned hairlineWidth=1. The filter now matches both builds, the rendered check requires the package the way the page does, and a builder test reads every hairlineWidth assignment in the bundle. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile-web): draw page hairlines at a width WebKit paints too 1/ratio is exactly one device pixel, and WebKit floors that to 0 and paints nothing (0.3333px at a scale of 3), so the iOS shell would have lost every hairline. The shim now uses native's device-pixel count plus half a pixel; both engines floor a border to whole device pixels, so each paints what React Native paints at ratios 1, 2, 3, 3.5 and 4, measured per engine. The rendered parity check now runs in WebKit at a device scale of 3 as well as Chromium. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile-web): check the parity style in the existing root-reset build Drops a second full build that read one HTML string, plus two assertions that tested the constant against itself. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile-web): round the page hairline up to the 1/64 px layout step Half a pixel over native's count kept borders at one device pixel but made the separators drawn as `height: StyleSheet.hairlineWidth` straddle two rows at about half of all offsets. Both engines lay out in 1/64 CSS px, and WebKit stores an exact 1/3 as 21/64 and paints nothing, so the width is now native's device-pixel count over the ratio, rounded up to the next 1/64 (22/64 at 3). The rendered check adds a separator at a 10.1 px offset in Chromium and WebKit. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
5b6a857e41 |
fix(mobile): route the bottom drawer's keyboard through the platform seam (OTA phase C follow-up) (#22556)
* fix(mobile): route the bottom drawer's keyboard through the platform seam Fill-mode sheets called Keyboard.metrics() directly, which react-native-web does not implement, so opening one on the page threw and the shell re-downloaded the workspace. The drawer now reads useSoftKeyboard, whose native half seeds from metrics() and carries the event duration, and whose web half answers from the window (duration 0). The fill/content-sized seed rule and resolveBottomDrawerKeyboardInset are unchanged. A census keeps Keyboard.metrics/addListener inside the seam plus the tab-sheet hide wait. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(config): retire the drawer's exemption from the page keyboard census The bottom drawer now reads the keyboard seam, so no module in the source-control or review closures names react-native-web's Keyboard stub. The census also flags Keyboard.metrics, which the stub lacks. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): give the drawer an imperative keyboard pair from the seam The seam now exports subscribeSoftKeyboard and currentSoftKeyboardHeight beside its hooks. The drawer's effect is back to its original shape with only its Keyboard calls swapped for the pair, and useSoftKeyboard is back to {height, visible} with no metrics() seed. Seeding every consumer opened an iOS window between willHide and didHide where metrics() still reads open. The web pair answers from visualViewport, so it stays silent inside the shell and lifts sheets in a plain mobile browser. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): start the web keyboard subscription from the current strip A keyboard already covering the page when subscribeSoftKeyboard attached never produced onHide when it closed, so the occlusion hook and a seeded fill sheet stayed lifted. Outside the shell only. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
069dc8a1d8 |
feat(agent-launch): let a caller reserve the chat session, and start terminal launches with the session picks (#22523)
* feat(agent-launch): let a caller reserve the chat session and carry session picks to a terminal launch * fix(agent-launch): keep a caller-minted session id named for its agent, and mint the fallback the same way * test(mobile): model the older host from the launch fields, not the refined schema * docs(agent-launch): describe the reserved session id as conversation identity, not placement The caller mints the session id so it knows which conversation it started; tab placement is not keyed on it. Also puts the terminal surface's doc comment back on createTerminalSurface. * docs(agent-launch): say a terminal launch reads the session picks on the wire contract The `sessionOptions` field doc still said a terminal launch ignores them, which this branch changed. * fix(agent-launch): check a reserved session id's token after the agent name, not the whole id A hyphenated agent name failed the one-token check, so any session id for such an agent was refused at the wire, while every other agent without a chat has its id ignored on the terminal. |
||
|
|
12a7ffdc64 |
test(runtime): stop worker-recovery retries from scanning inside later tests (#22567)
The legacy worker terminal recovery retry re-arms itself on a 1s..30s backoff for as long as a dispatch stays deferred. Nothing in the runtime suite ever resolves one, so a single test armed a loop that kept re-running recovery -- and the worktree scans it issues -- through the shared `listWorktrees` stub for the rest of the file's run, landing inside whichever test was executing when the timer fired. That is what made `lineage-and-scan-cache-part-03`'s per-repo TTL test see 4 scans instead of 3 only under CI load. Give the controller a `cancelAllRetries`, track controllers while they have a timer armed, and cancel them from the shared runtime test lifecycle reset. Measured over the whole 1299-test file with an afterEach probe: 25 stray post-test `listWorktrees` calls before, 0 after. |
||
|
|
6847390c0a |
fix(ai-vault-search): own buffered transcript rows before they pin the parent (#22563)
Capped search rows are slices of the transcript. Own each row with the shared copier before buffering it, and keep the heap regression from #22378. The sidebar test harness now includes the worktree listing map the toolchain banner reads, so that rerender no longer crashes. Co-authored-by: fancivez <fancivez@gmail.com> |
||
|
|
7f203ad5aa |
fix(browser): only pixel-capturing commands wait for the page to be drawn (#22528)
* WIP * WIP2 * fix(browser): only pixel-capturing commands wait for the page to be drawn Every targeted browser command used to take an automation-visibility lease, which waits for two desktop-window animation frames (capped at 2 s). When the desktop window is minimized or throttled, that wait always hits the cap, so each phone tap or agent command stalled for up to 2 s (STA-8024). Input, page JavaScript, layout and accessibility snapshots all work on a hidden page; only pixel capture needs it drawn. The lease is now opt-in via `needsPaint` on the two commands that can capture pixels through this path (`exec` passthrough and `pdf`); `ensureVisible` is removed. Screenshots keep managing their own lease. (Commits |
||
|
|
60bd1dfdea |
feat(native-chat): one shell-environment setting for every structured chat (#22387)
* feat(native-chat): one shell-environment setting for every structured chat Structured Codex chats started from the login-shell environment, while structured Claude chats started from Orca's own process environment, so a variable exported in .zshrc reached one and not the other. Both now start from the same base, chosen by a new setting: - on (default): the whole login-shell environment, as a terminal gets - off: Orca's environment plus PATH, locale, SSH_AUTH_SOCK, and the variable names the user lists The setting is re-read each time a chat starts or resumes. It is shown only when Chat UI, the Chat UI default view, and structured native chat are all on. Terminal-backed chat is unchanged. * fix(native-chat): normalize the shell-environment settings when a profile loads A hand-edited settings file could store the variable list as something other than an array, and the structured runtime called `.filter` on it per launch, so a malformed value failed every structured chat create and resume, and the settings pane render. Normalize both keys where the profile loads, the same way the other array settings are, through one shared normalizer the runtime policy also uses. Also pin that an uncommitted name draft survives an unrelated settings re-render. * fix(native-chat): keep the pinned account as the only source of a structured chat's Claude home The session record owns which Claude home a structured chat uses, and the acquisition pin (claudeConfigDirEnvPatch) is the only emitter of CLAUDE_CONFIG_DIR, compared against what the child would otherwise inherit. With the login-shell snapshot as the inherited base, a CLAUDE_CONFIG_DIR exported only in a shell rc flipped that comparison and produced an explicit pin to the CLI default home, which moves the CLI off its default Keychain item. Drop the inherited CLAUDE_CONFIG_DIR in the Claude launch resolver before the pin runs, as Codex already does for an inherited CODEX_HOME. A configured per-agent overlay still passes through, since the record already honors it. * fix(native-chat): drop Orca's own CLAUDE_CONFIG_DIR from a structured Claude child too The process spawner merges Orca's process env under the launch env, so a CLAUDE_CONFIG_DIR exported to Orca itself reached the child around the launch resolver's drop and unseen by the account pin. One helper now strips it from both inherited bases. Also declare the two shell-environment settings on the runtime store contract and add the six new strings to every locale catalog. * feat(native-chat): add shell variables one at a time with a removable list * fix(native-chat): return focus to the name input after removing a shell variable * fix(native-chat): use a neutral placeholder for the shell variable input The empty input showed a grey HTTPS_PROXY as its placeholder, which reads as a saved value, especially right after that exact entry is removed from the list. Use "Variable name" instead, in every locale catalog. |
||
|
|
9cdbc0c128 |
fix(mobile): keep the shell's window insets out of the page WebView (OTA phase C follow-up) (#22549)
* fix(mobile-web): stop the page declaring viewport-fit=cover The shell already pads the WebView out of the status and navigation bars. With viewport-fit=cover, Android's edge-to-edge WebView still reports the window's bar insets through env(safe-area-inset-*), which react-native-safe-area-context on web reads, so every page-side SafeAreaView padded a full bar a second time. Without it env() reads 0 and the shell's pad is the only one. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(mobile): keep the shell's window insets out of the page WebView WebView M144+ forwards the window's systemBars and displayCutout insets to CSS env(safe-area-inset-*) for every WebView, and Chromium applies them regardless of viewport-fit. The shell already pads the WebView out of both bars, so every page-side SafeAreaView (expo-router's DefaultNavigator and the session header) padded a bar a second time. M139+ likewise resizes the visual viewport for ime(), which the shell has already done by shortening the WebView. The WebView now sees those three types zeroed, per Android's "zeroing" approach (not CONSUMED, so later changes still reach it). A listener replaces the WebView's own onApplyWindowInsets, so the zeroed set is passed back into it. iOS needs nothing: the WKWebView uses contentInsetAdjustmentBehavior = .never inside the padded shell and reports zero insets. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * docs(mobile-web): say why the page declares no viewport-fit The earlier comment claimed dropping viewport-fit=cover makes env() read 0 on Android; Chromium's WebView applies the safe area regardless of viewport-fit. The page simply never asks to extend under the bars, and the shell owns the safe area. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * refactor(mobile): keep the page inset zeroing private to the shell view The transformation has no honest JVM test (the builder runs as SDK 0 there and drops every inset type), so it moves into MobileWebShellView.kt as private members instead of standing alone. The listener comment now covers both the P-R listener and the S+ onApplyWindowInsets path it replaces, and the page document's comment says the env() zeroing is Android's; on iOS the padded WKWebView reports none. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
f1eb1913a6 |
fix(opencode2): block the pane on every session-owned form (#22548)
#22399 admitted an OpenCode 2 form.created as a pane blocker only when metadata.kind === "question". On v2.0.15 that is an allow-list on a field with no contract: packages/schema/src/form.ts declares Metadata as an open Schema.Record and metadata itself as optional, and the public POST /api/session/:sessionID/form endpoint lets any client raise a real blocking form on a real session with no metadata. Orca dropped those, so the pane painted no blocker while OpenCode waited forever. Invert the default. Every form whose owner is a real session blocks; only a form owned by the "global" MCP-elicitation sentinel is dropped, because that owner is not a session and never goes idle, so its blocker could not be retired. That also restores websearch.provider as a blocker: it carries the real context.sessionID, session idle retires it, and while it is pending the agent is genuinely stalled on the user. Resolution is unchanged: clearAttentionForResolution keys on the exact form id plus source session, so a resolution for a dropped form matches nothing and cannot retire a live blocker. |
||
|
|
1866a796fd |
fix: explain Xcode-blocked Git once in the sidebar and rescan on return (#22552)
* fix: explain Xcode-blocked Git once in the sidebar and rescan on return * refactor: simplify Xcode toolchain banner to one classifier and a stateless rescan * refactor: fold banner copy into one table and cover SSH/runtime gating |
||
|
|
90f0c5c8ae |
test(file-search): pin request-key listings against the real runtime shape (#22312)
* test(file-search): pin request-key listings against the real runtime shape * test(file-search): pin intermediate renders and late remote answers Strengthen the stale-answer guard to assert every intermediate render reads as loading (null), add a late-answer drop case, make the tab-entry loading pin non-vacuous, and correct the e2e comment for local listings. * test(file-search): keep only the non-duplicate runtime-listing pins Drop the remote projection cases already covered at the hook level, collapse the classifier integration to the loading pin, drop the local-only rapid-edit e2e, and fix the brittle README absent-file assertion that failed CI. |
||
|
|
845db9e5e2 |
fix(native-chat): underline only file links a click can act on (#22370)
* fix(native-chat): underline only file links a click can act on A chat message could underline a bare file name such as `deck.md` that resolved nowhere, and clicking it did nothing, so it read as a broken link. - Inline code and quoted text become file links only when they name a path (contain a `/` or `\`), matching plain prose; a bare file name stays plain code. - Every file link click now answers: it opens, or says the file was not found, that the host could not be checked, or that the path could not be resolved. - Explicit links like [x](README.md:5) route as files, and linked text keeps `#`, `?` and `%XX` literally instead of re-parsing them as URL syntax. * fix(native-chat): wrap the parsed file location so file URIs in chat text still open Linkified prose, quoted text and inline code wrapped their display text, which the literal wrapped-href route no longer URL-parses, so file:///... resolved as a relative path under the worktree. Wrap pathText[:line[:col]] from the parsed link instead. |
||
|
|
19e8bf319d |
fix(ai-vault-search): detach buffered transcript rows (#22545)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: Andre Ambrósio <56239028+sirambrosio@users.noreply.github.com> |
||
|
|
519bde81df |
feat(ipynb): render notebooks like a notebook, with seamless click-to-edit cells (#22519)
* feat(ipynb): parse ANSI SGR sequences in notebook output Tracebacks and stream output carry terminal colour codes that the notebook printed raw. Splits text into styled runs (16/256/truecolor, bold, italic, underline) and drops non-SGR escapes via the shared stripper. * feat(ipynb): render notebooks like a notebook, not a form - Prose inherits the app UI font instead of a bare terminal font name, which Chromium could not resolve and fell back to Times (removes the now-dead resolveEditorFontFamilyOrInherit). - Markdown cells render by default; double-click or Enter edits them in the same Monaco surface code cells use. Code cells activate on press so a collapsing neighbour cannot swallow the click, which also retires the root pointer-capture deactivation (Monaco blur already covers it). - Code sits on its own tinted surface, the active cell gets a ring, and the editor sizes to its content so activating a cell no longer jumps. - The always-on 8-button toolbar and native select become a hover/focus toolbar (move, delete, and a menu for insert and cell type); the Jupyter [n] prompt turns into the run button. - Outputs show only the richest MIME representation, HTML renders in a script-less, no-network sandboxed frame sized to its content, and ANSI colours use the default terminal palettes. Adopts the MarkdownPreviewBody reuse, richest-MIME selection, extra raster MIME ranks, and CSP-sandboxed auto-height HTML frame from #18542. Co-authored-by: maxidiazbattan <maxidiazbattan@gmail.com> * chore(ipynb): drop the nbformat label and BETA badge from the notebook header parseIpynb already rejects notebooks without a v4 cells array, so the label carried no actionable information; the parsed nbformat field goes with it. Removes those catalog keys and the stale lowercase code/markdown ones. * fix(ipynb): keep cells pixel-stable when they switch to editing The preview and the live editor disagreed on four things, measured over CDP: - Font: the excerpt painted the bare "SF Mono" name (or --font-mono via its row class), while Monaco appended its own fallbacks and landed on Menlo. Both now use resolveEditorFontStack, the editor font plus the terminal fallback chain. - Line height: 20px rows vs Monaco's 21px. Both read CODE_EXCERPT_LAYOUT. - Gutter: a 48px line-number column plus 12px inset vs Monaco's 25px gutter. Notebook cells drop line numbers (the Jupyter and VS Code notebook default) and Monaco's decorations lane is the same 12px inset. - Rows: colorized blank lines collapsed to 0px, and a trailing newline had no preview row. Rows are fixed-height and a trailing newline opens an empty last line, matching the Monaco model. The [n] prompt and run icon now share one grid cell, so the hover swap keeps the label's box and centre. The commented-line tint moves to a theme token. * feat(ipynb): VS Code-style run gutter above a fixed execution count Replaces the in-place [n]/play swap from |
||
|
|
1116c54230 |
fix(relay): accept production cells past c29 in the regional rehome operator (#22518)
The selector membership check capped cell ids at c29, so enable failed closed with "selector membership is invalid" once c30 went general. Accept c1-c99 with no leading zero. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
a77f87ea16 |
feat(sidebar): copy workspace name from context menu (#22338)
* feat(sidebar): copy workspace name from context menu Add Copy Name directly below Copy Path in the workspace context menu. It copies the name through resolveWorktreeDisplayName, the renderer mirror of main's mergeWorktree fallback (custom name, then branch, then folder), so the copied text is what `name:` worktree selectors resolve against and a cleared custom name no longer copies `undefined`. The context-menu model now spreads the command handlers instead of listing each one twice, which keeps it under the file-length limit. Fixes #21980 Linear: STA-8068 * refactor(sidebar): rename Copy Name to Copy Worktree Name - Clarify that this copies the worktree display name, not the path - Update all locale strings and i18n keys - Rename test file to match * test(sidebar): cover folder workspace copy worktree name --------- Co-authored-by: Seongho Bae <me@seonghobae.me> |
||
|
|
641a7f36d9 |
fix(native-chat): keep one live tool-run header from a call's start to the turn's end (#22432)
* fix(native-chat): keep one live tool-run header from a call's start to the turn's end
The collapsed tool run's header was two elements, one for "a call is running"
and one for "nothing is", chosen call by call. Every call start and end
remounted it, the count disappeared while a call ran and came back one
higher, and a call that finished inside a frame still bought the whole swap.
That is the 42→43 flicker in the report.
The header is now one element whose live state belongs to the turn, not to
any call: it stays live from the run's first call until the agent moves past
it (prose, a further run, or the turn's end), and settles in place. While
live the sentence speaks in the present tense and counts the call in flight
("Running 3 commands"), with the latest call's command beside it as a muted
preview; once settled it reads as before ("Ran 3 commands ✓"). The category
glyph is the run's in both states, and the completion mark only appears once
settled, so nothing pops between calls.
Which run is live is derived where the transcript is sliced into rows: the
last row that speaks or acts is the trailing one. A reasoning aside after it
leaves it live; an answer or a further run settles it.
Present-tense forms for the ten sentence categories are added to the shared
copy and the English catalog. The transcript-file lane, which renders with
the structured activity UI off, is unchanged.
* fix(native-chat): settle a run blocked on the reader, keep it live past an approval
- A run whose question is awaiting the reader's answer no longer pulses
"Reading 1 file" while the agent is blocked; it falls back to its calls.
- An approval's receipt no longer moves past the run above it, so the call
it just approved reads as running while it runs.
- The header button is the live region, so the count is announced too.
- Drop the unused live option and record from the shared English sentence;
nothing renders it yet.
* fix(native-chat): stop the settled run's check from fading in on every mount
Windowing remounts settled rows as the reader scrolls, and a restored transcript
mounts them all at once, so the fade replayed where nothing had changed. Also
pin that the live header counts the next call on the same element.
|
||
|
|
563dd5487f |
feat(native-chat): show a Codex chat's goal above the composer, and set it from goal mode (#22377)
* feat(native-chat): show a Codex chat's goal above the composer and set it from goal mode Structured Codex chat now treats the thread goal as session state: a banner above the composer shows the current goal (pursuing / paused) with clear, pause/resume and expand; /goal enters a goal mode whose send calls thread/goal/set; the objective is journaled as a user message marked as sent as a goal. The banner is derived from the journaled goal rows, which Codex's resume snapshot refreshes, so a reopened or adopted chat shows its goal. Fixes STA-8159 * fix(native-chat): replace a recorded goal by clearing first, and recover a lost goal-change response - A set while the journal records a goal (any status) clears it before setting, so the new goal starts with its own time and token counters instead of rewriting the old goal's objective in place. - The threadGoal plan answers an unknown outcome from the goal the journal records and reruns otherwise, so one request timeout no longer refuses every later Clear/Pause/Resume as unknown for the mounted session. - The goal-mode chip says "Exit goal mode"; "Clear goal" stays the banner's action on the provider goal. - A typed bare /goal on Enter enters goal mode, the same as picking it. - The renderer reads the goal off the tail of its ordered snapshot; the host's unordered map keeps the by-sequence reader. - Drop the composer's duplicate in-flight guard; the goal controller already serializes changes. - Pin that a counter-only revision reaches a subscriber's live page under its original sequence. * fix(native-chat): keep a bare /goal inside goal mode as the entrance, and pin goal delivery and serialization - A bare `/goal` submitted while already in goal mode re-enters the mode instead of setting a goal whose objective is the literal text "/goal". - The counter-only revision pin now drives the host's own event sink bound to a real journal, so it goes red when the publish after a lifecycle transition is dropped; the previous fake sink never published. - Pin that a set which threw after journaling its objective puts that objective back exactly once when the ledger reruns the same operation id. - Cover the goal controller hook: absent without host support, the loaded window wins over the host's answer, a second change while one is unsettled answers false without a request, and a refused change frees the next one. * fix(native-chat): resume a blocked or usage-limited goal, and keep goal-mode drafts honest - The goal bar offers Resume on a blocked or usage-limited goal, which the provider resumes exactly as it resumes a paused one; a goal whose token budget is spent still offers only Clear. The rule lives beside the other goal facts in shared code so every reader answers it the same way. - A `/goal <text>` typed inside goal mode sets the objective `<text>`, as it does outside goal mode, instead of a goal whose objective is the literal command. - Setting a goal is a host round trip; a draft edited while it was in flight is no longer wiped when the goal lands, matching every other host command. - Pin that a lost status-change response is read as applied only when the recorded goal is in that status, that a cleared row in the loaded window outranks the host's earlier answer, and that the PTY lane is untouched. * fix(native-chat): keep the load-older anchor on the loaded window when a live revision lands below it A live revision of a row keeps that row's original sequence. When the row is older than the client's loaded window, the shared reducer merged it in and it became the load-older anchor, so paging `before` it skipped every row between. A goal's counter-only revisions during a long goal turn reach any client that attached after the goal row left its window, so a reopened chat lost rows on scroll-back. The reducer now admits live rows only at or above the window's oldest row while older rows remain on the host; the journal keeps the revision and the page reader serves it once the window reaches the row. With nothing older on the host the window is the whole journal, so a row below the head is admitted as before. Also drain accepted provider events before a goal set reads the journal to decide whether it replaces a recorded goal. |
||
|
|
a375936c04 |
feat(agent-launch): let a caller reserve the pane its terminal launch creates (#22291)
* feat(agent-launch): let a caller reserve the pane its terminal launch creates * fix(agent-launch): refuse a launch whose reserved pane is already live * fix(agent-launch): refuse a live reserved pane before it is revealed The live-pane refusal used to fire in the executor, after createTerminal had already issued a handle, published the mobile snapshot and revealed the tab. The reveal re-registered a fresh launch config over the running agent's. agent.launch now passes requireFreshPane with a reserved pane, and createTerminal throws AgentLaunchPaneAlreadyLiveError as soon as spawn reports it attached to a live pane. That is before any handle, snapshot or reveal. The spawn reattach itself is the one terminal.create already uses, so the live PTY is never killed, and the stable-pane create claim is still released in finally. The isReattach plumbing added to the launch factory for the old check is gone. A replay-safe launch refused this way on an existing workspace now records a failed ledger row, the same way a name collision does. Before, the row stayed claimed, so every retry got agent_session_operation_unknown. agent.launchReplay passes the code through. On create-worktree the workspace already exists when the terminal is refused, so the row stays unknown. The code is added to the runtime passthrough list so callers can branch on it. The pane key is now in the replay fingerprint, deliberately. It is not placement: group, anchor and focus still stay out of the request and out of the ledger. It is identity. It is written into the pane's PTY environment and names the tab the caller has placed. A retry that reserved a different pane is therefore a different request. Replaying the first answer would return a key the new reservation can never find. This matches terminal.createAgentSession, which also fingerprints its tab and leaf ids. The key is only folded in when present, so every existing digest is unchanged, and a test pins that. The wire schema now refuses a tab id the runtime would not adopt as sent: one with surrounding whitespace, which the runtime trims, and one longer than 512 characters, which the spawn reservation does not key on. It reuses the tab-id schema that Placement uses. * test(agent-launch): pin that a refused live pane issues no handle The refusal test named handle issuance but only asserted the reveal, so a throw moved to just before the reveal would still pass. Assert no terminal is registered, with the attach test as the positive control. |
||
|
|
dac82f61bc | Update README downloads badge | ||
|
|
8d6fec597b |
Optimize cloud-verify workflow to scan HEAD instead of all history (#22457)
* fix(cloud-verify): scan HEAD instead of all history - Gitleaks now verifies only the checked-out revision - Reduces scan scope and improves verification workflow performance * ci(cloud-verify): clarify that Gitleaks scans HEAD-reachable history |
||
|
|
942d993f1f |
fix(editor): map Salesforce Apex extensions to the apex language id (#14287)
Fixes #22049 Co-authored-by: Nurdaulet Bolat <204565446+nvimq@users.noreply.github.com> |
||
|
|
ae3d380b23 |
fix(relay): lock only the target cell row, last and NOWAIT, in the rehome commit (#22449)
* fix(relay): lock only the target cell row, last and NOWAIT, in the rehome commit The idle-rehome commit runs on the source cell. From an Asia cell each statement is a cross-region round trip, and the transaction locked every relay_cells row plus every runtime, capability and safety row before about twenty more statements, so each Asia-source rehome held the whole fleet's cell rows for ~3.6 s and every reconnect, renewal and placement queued or timed out behind it. The commit now reads the cell inventory and the runtime, capability and safety tables unlocked, keeps the control and worker rows locked (now NOWAIT), and takes one cell lock: the target row, in a single statement that locks it NOWAIT, re-checks enabled, general admission and capacity, and reserves the units, issued as the last statement before COMMIT. A target that changed admission, filled up, or is locked by another writer rolls the whole commit back and answers deferred (candidate-ineligible). The hold is sampled under a site label, so cellInventoryHoldMsMax still sees rehome holds and rehomeTargetRowHoldMsMax reports them apart. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(relay): fail closed on the rehome target-row lock clause The target-row statement now carries FOR UPDATE ... NOWAIT unless the dialect is explicitly SQLite, so a wrapper that omits the optional dialect can no longer run the reservation unlocked. Test wrappers and the fault injection entry forward the dialect they wrap. The latency test also probes the admission and region tables at every round trip; only the target's admission row may be locked, and only before COMMIT. The runbook notes that an Asia-sourced commit holds the rehome control row for about 6.5 s, so a pause that fails once is retried. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
a2a78ab335 |
feat(relay): alert on relay cell table lock convoys (#22446)
* feat(relay): alert on relay cell table lock convoys Adds a log-based metric and alert for cell-inventory lock holds of at least 1,000 ms, and a Cloud SQL log metric and alert for relay-only lock timeout cancels at 20 or more per minute. NOWAIT refusals are excluded: background sweeps produce about 160 per minute even with rehome paused. Replayed over 2026-09-20 14:00 to 2026-09-22 15:00 UTC: the hold filter matches all 93 asia-east2 rehome holds plus 9 director holds, and every one of the 88 cancel burst minutes overlaps an asia-east2 hold. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): page only on cell lock holds; director holds stay visible Director holds of 1-2.5 s recur several times a day with rehoming paused, and pausing rehome does not stop them. The paging hold policy now selects role=cell samples only; a separate policy with no notification channel keeps director holds visible. The burst documentation no longer claims no burst happens while paused, and the runbook points a burst with no cell hold at the director policy. Replayed cell-only: 93 of 93 asia-east2 holds, 0 director holds over 2026-09-20 14:00 to 2026-09-22 15:00 UTC; 0 from then to 2026-09-23 07:30. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
1043dc5e1d |
fix(relay): stop rehoming hosts off Asia cells until the lock fix lands (#22443)
The source cell runs the rehome commit. An Asia source pays a cross-ocean round trip per statement while holding relay_cells row locks every cell needs, which convoys the fleet database. Selection now drops source cells outside the director's region before building the decision window, so the incumbent_region filter shrinks while Asia cells stay valid targets. The preview counts the same hosts as source-outside-director-region and the poll summary reports skippedOffRegionSourceCells. Temporary stopgap. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
51c3434851 |
chore(relay): treat Asia cell c30 as a general cell now that it is promoted (#22439)
C30 was promoted to general on 2026-09-23 (selector generation 286). The same-cap wave now rolls it as a general cell instead of handing it back isolated, and the shadow gate reads its pool beside C27-C29. Follow-up to #22375. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
293c2508fc |
test(mobile): move the session closure pin past the structured tool-line module (#22430)
#22349 added `src/shared/structured-agent-session-tool-call-block.ts`, which the projection and live turn the session route already reaches import. The PR was src/shared-only, so its CI never ran the closure suite; main's pin stayed at 4215 while the closure measures 4216. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
17ffbf3b31 |
fix(runtime): answer startup terminal queries for background-created terminals (#22384)
* fix(runtime): answer startup terminal queries for background-created terminals A runtime-created terminal (orchestration worker-start, `orca terminal create`) has no renderer pane until the user opens its tab. Main only answers terminal queries for PTYs the renderer marked hidden, so on a fresh app sitting on the landing screen nothing answered the agent's startup cursor-position query. Muse waits ~2s per unanswered CPR and then exits 0 with no output, which surfaced as `agent_readiness: timeout`. Background runtime spawns now carry initiallyHidden, mirroring the renderer's hidden-at-spawn path: fresh daemon sessions are marked before byte zero, the committed id is marked and paced as backgrounded, and the mark is released on failure, reattach, or adoption. A pane that later mounts visible unmarks and restores from the model snapshot as before. * fix(runtime): keep background PTYs hidden and paced across renderer reloads A runtime background spawn has no renderer pane to report visibility or re-mark it hidden, so it synced as foregrounded (no backpressure thinning) and a reload/crash gate reset cleared its hidden mark, leaving startup queries unanswered. Track runtime-owned hidden marks: they survive renderer-scoped resets, count as known-hidden for backgrounded pacing until a visible report, and are released by a renderer unmark or PTY teardown. * fix(runtime): don't re-hide a background PTY whose view mounted visible during spawn |
||
|
|
b864a1c775 |
test(mobile): repin the recording corpus to main's tip after #22381 (#22407)
#22381 pinned its own branch commit, which the squash left off main; the corpus now pins main at
|
||
|
|
0b16a31e6e |
fix(runtime): budget explicit terminal close for the daemon's immediate-kill verdict (#22385)
* fix(runtime): budget explicit terminal close for the daemon's immediate-kill verdict Explicit terminal close (worker-release, worker-stop, `orca terminal close`) gave the daemon kill RPC a fixed 2s deadline. The daemon's immediate kill captures descendants, SIGTERMs them with a 2.5s verification window, then waits up to 8s for the root's physical exit. An agent that runs exit hooks after SIGTERM (Muse: ~3s) outlived main's 2s timeout, so close reported the PTY unverifiable and worker-release returned release_unknown even though the daemon confirmed the exit ~200ms later. Derive the close budget from the daemon's own immediate-kill reply budget (now in an import-free module) plus 2s for the post-kill inventory check. A process that exits within the daemon's budget is released; a wedged process or unreachable host still times out as unverifiable. * fix(runtime): budget the force-kill retry and exercise an expired close deadline * test(runtime): drop tautological deadline-expiry assertion |
||
|
|
eb18eaf2b6 |
feat(usage): add Muse Code local usage provider (#22379)
* feat(usage): add Muse Code local usage provider Scan Muse session logs (including subagent logs, which hold usage the parent log does not) for model_completed token events and surface them as a fourth local usage provider: shared scan worker, persisted per-file cache reused by mtime/size, cross-log dedupe, Stats tab, and Usage Overview integration. Muse logs carry no price, so the provider reports tokens only. * fix(usage): name Muse in Stats & Usage copy; skip partial-cost warning when nothing is priced * fix(usage): surface unreadable Muse sessions root; name Muse in remaining Stats & Usage copy * fix(usage): count distinct same-content Muse records within one log |
||
|
|
9fef7a0f04 |
fix(cloud): gate the Asia canary on its own cell's SQL failures, not the directors' (#22405)
* fix(cloud): gate the Asia canary on its own cell's SQL failures, not the directors' The production canary summed sqlFailuresDelta over every director and the canary cell and required zero. Directors log a steady baseline of relay_cells NOWAIT and lock-timeout refusals unrelated to the canary cell, so a C30 canary failed most attempts on that noise. The canary now requires zero SQL failures from the canary cell's own metrics and records the director sum as directorSqlFailures without gating on it. Directors keep every other rule (unavailable regions, fallbacks, pool waiting, transient waiter and wait-time bounds). Staging keeps the combined zero rule and its evidence shape unchanged. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * fix(cloud): gate the Asia canary's pool bounds on its own cell too Directors also show a steady pool-wait baseline (waiting above zero and waits over 50 ms in about 6 of every 60 minutes), so a five-minute canary still failed about half the time on director pool pressure unrelated to the canary cell. With gateDirectorDatabase off, the production canary now applies databasePoolWaitingMax, databasePoolWaitersMax and databasePoolWaitMsMax to the canary cell's metrics only and records the director values under director-prefixed names. Directors still gate Asia selections, region fallbacks and unavailable regions. The staging path keeps its combined values, key order and validation order. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
8757e40063 |
fix(native-chat): keep a structured agent's tool line between tool calls (#22349)
* fix(native-chat): keep a structured agent's tool line between tool calls A structured session's status named a tool only while the call was still running, so the sidebar's tool line went blank whenever the agent was thinking or writing between calls. Terminal agents keep naming the finished tool until the next one starts, and clear it after a failure. The structured status projection now does the same: a running call wins, otherwise the turn's newest root call if it completed. * fix(native-chat): bound the structured tool line by the running turn, not the user row A send made while a turn is running writes its user row into the journal straight away, and the turn keeps going. Stopping the scan at that row blanked the tool line while a tool was still running. The scan now runs to the turn record and names a call only when that record is still running, so a turn that already ended never lends its last tool to a pending follow-up. This lookup was the running-only lookup's only production caller, so it replaces that lookup instead of sitting beside it. * fix(native-chat): keep naming a structured agent's failed tool until the next one Clearing the tool line after a failed call brought the blank gap back for much of a turn: Codex marks any nonzero exit as failed, so a search with no match or a red test run is enough. The failure already shows on the tool's own row in the transcript. The running turn's newest running call still wins; otherwise its newest root call is named whatever it settled to. * fix(native-chat): name a structured Codex edit on the tool line as the chat draws it Once a Codex edit's changes exist, its apply_patch call becomes a diff row, which the status lookup skipped, so the row named the command before the edit. The chat's tool-call block for a journal row now comes from one shared builder, and the status lookup reads the same definition: a diff is named as Diff with its path, and counts as settled since it carries no lifecycle. * docs(native-chat): describe the structured tool field as running-or-latest The status summary's toolName/toolInput now name the running turn's latest tool between calls, not only a running one. Update the wire type and status bridge comments that still said "the running tool". |
||
|
|
996f9cc306 |
feat(mobile-web-bundle): gzipped 384 KiB ranges over a capability-negotiated mobileWeb.bundle.range (OTA phase C follow-up) (#22381)
* feat(mobile-web): serve gzipped 384 KiB bundle ranges behind a capability Adds mobileWeb.bundle.range with its own strict params and result, so shipped chunk readers see no reply change. The host gzips each range at level 6 and sends identity when gzip does not shrink it, sharing the chunk method's read-slot budget and per-asset verification. status.get advertises mobileWeb.bundle.range.v1 beside mobileWeb.bundle.v1, and the method is allowlisted for paired phones. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * feat(mobile): read the bundle range capability and range replies Adds the range reply reader and operation, and picks range or chunk from the status.get capabilities the connection already proved, so an older desktop keeps being paged in chunks with no probe round trip. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * perf(mobile): keep four bundle chunk reads in flight across the whole manifest The fetch ran one worker per asset and paged inside an asset sequentially, so the largest script's 71 chunks were 71 serial round trips while the other readers idled. One window of four chunk reads now covers every (asset, offset) on the host's chunk grid, largest asset first. A read_limited refusal narrows the window and retries the read; eof is still read from the reply. Synthetic manifest (one 71-chunk asset, five small): 72 round trips -> 19. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * chore(mobile): add fflate 0.8.2 for gzip bundle ranges Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * feat(mobile): decode gzip bundle ranges into a bounded buffer Inflates each range into a buffer one byte past its window, so a gzip bomb costs at most that allocation and an overlong body is visible. A corrupt, truncated or unknown-encoding body refuses as range-undecodable; a body of the wrong decoded length refuses as range-length-mismatch. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * feat(mobile): pass the bundle read method from the session to the fetch The download reads the capabilities of the gates the reducer decided under and hands the fetch range or chunk. The fetch does not act on it yet; the range read lands on the pipelined window. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile): re-record the bundle-fetch family under pipelined reads Baseline moves to |
||
|
|
49121c32d8 |
docs(wechat): point community QR code at group 10 (#22403)
Group 9 is full; swap the README QR code and copy (all locales) to the new group 10 invite. |
||
|
|
51d3cafc4f |
feat(editor): add a setting to turn off preview tabs (#22398)
Single-clicking a file in the Explorer, or following a link in Markdown source, opens it as a preview tab that the next preview open replaces. There was no way to turn that off, so browsing files kept swapping one tab. Adds `editorPreviewTabsEnabled` (General -> Navigation, on by default). A caller's `preview` flag is now an intent that `resolveEditorPreviewIntent` resolves against the setting, covering every open path - files, diffs, history diffs, conflicts - in one place. Preview-ness is derived rather than reconciled: readers treat a tab as a preview only when the stored flag and the setting agree, so a flag left over from a saved session, another window, or a host switch is inert while previews are off. Nothing rewrites stored flags when the setting changes, so no settings-landing path has to remember to clean up. Fixes #22397 |
||
|
|
52a1e2875b |
feat(orchestration): accept Muse model and effort for supervised workers (#22383)
* feat(orchestration): accept Muse model and effort for supervised workers `worker-start --agent muse` already launched, but `--model` was refused because Muse had no session-option catalog. Add one that maps worker preferences to `muse --model <id>` and `--reasoning-effort <level>`; it seeds no models, so native-chat surfaces show no picker. opencode stays without `--model`: the opencode 2 TUI (now shipped as `opencode`) rejects the flag, so the refusal now tells callers to rely on the agent's own config. Help, skill guide, and docs list valid `--agent` ids and the agents that accept `--model`. Refs #19823 * test(mobile): repin session route closure for the Muse option catalog |
||
|
|
0afc66ebd3 |
fix(opencode2): only treat the question tool's form as a pane blocker (#22399)
* fix(opencode2): only treat the question tool's form as a pane blocker OpenCode 2 has one form primitive and several producers, and Orca's setup bridge mapped every `form.created` to `question.asked` — the un-evictable "the pane owner must answer this" blocker. Against opencode v2.0.12 only `metadata.kind === "question"` is the agent's question tool; `websearch.provider` is a provider picker and `mcp-elicitation` is an MCP server prompt raised on sessionID "global", which is not a session and so can never be retired by that session going idle. Admit only the question kind, remember the admitted form ids, and drop `form.replied`/`form.cancelled` for forms that were never admitted so an ignored form's resolution cannot retire a live blocker. Evidence (live v2.0.12 capture, real TUI in a PTY against `opencode serve`) in docs/bug-reproductions/opencode2-form-created-kinds. That capture also shows the reported Subagents/Shell/Terminals dock and the agent picker emit no server event at all, so they were never the `form.created` source. Refs #22371 * refactor(opencode2): drop the unreachable form-resolution guard Review was right that the admitted-form-id set defended against nothing. `clearAttentionForResolution` builds the exact key [factoryID, "AskUserQuestion", form.id, sourceSessionID] and returns null on a miss, with no session-wide fallback, and form ids are unique — so a resolution for a form Orca ignored already matches no live blocker. The guard's comment claimed a collision the key structure rules out, which is worse than no comment. Removes the set, its FIFO eviction helper, and the claim; the kind check on form.created is the whole fix. The end-to-end test stays: it pins the behavior that an MCP form raised and cancelled leaves a live question blocker standing, which is the property worth holding regardless of how it is achieved. |
||
|
|
ee1c522070 |
fix(opencode-usage): read OpenCode 2 session_v2 token totals (#22391)
* fix(opencode-usage): read OpenCode 2 session_v2 token totals
OpenCode 2 copies every v1 `session` row into a new `session_v2` table and
then writes only there, so every OpenCode 2 session was invisible to the
usage scanner, which only knew `session`. Reading both tables unfiltered
would double-count the migrated rows, so the scan now builds one
deduplicated session relation where `session_v2` owns any id it shares with
the legacy table, and joins the message-level fallbacks against it too.
Bumps the usage cache schema version so existing caches rescan.
Measured against the real local opencode.db (OpenCode v2.0.12, 234 legacy
`session` rows / 252 `session_v2` rows, 18 of them v2-only):
migrated db before: 206 sessions, 367,419,745 tokens, $79.4704
after: 219 sessions, 374,305,364 tokens, $79.4817
fresh v2 db before: 0 sessions, 0 tokens, $0
(legacy table after: 219 sessions, 374,305,364 tokens, $79.4817
emptied)
Fixes #15841
* fix(opencode-usage): resolve a migrated session to its fuller row
Review found two ways the fixed `session_v2`-wins precedence loses usage.
A `session_v2` without the token columns scores 0 through the source's
fallback literals, but still excluded the legacy row for every shared id, so
a token-bearing legacy `session` paired with a token-less `session_v2`
returned nothing at all — worse than before this branch. And upstream's
importer recomputes v2 totals from decoded messages, so a session whose
messages fail to decode lands *below* its frozen legacy copy, which the
"legacy is frozen, v2 is newer" rationale did not account for.
Both collapse into one rule: per id, the row with the greater token total
wins, and a tie goes to the higher-priority table. A faithful migration is a
tie, so it still resolves to `session_v2`; a v2 row that lost data no longer
erases the legacy record. This also makes `hasSessionUsageColumns`'s `some`
correct rather than merely lenient — a table without the columns can never
outrank a sibling that has them.
Measured on the real local opencode.db: zero shared ids have a legacy row
that beats its `session_v2` counterpart, so the rule is a no-op on healthy
data and only engages on the degraded shapes above.
Both new tests fail on the previous commit and pass here.
|
||
|
|
98a6a5325c |
test(mobile): repin the recording corpus to main's tip after #22376 (#22394)
#22376 pinned its own branch commit, which the squash left off main; the corpus now pins main at
|
||
|
|
483fa0aca2 |
fix(cloud): compare the Asia topology budget gate against the measured 500-connection default (#22386)
The topology workflow's Cloud SQL gate carried a hard-coded 400 for the instance's tier default while the consumer contract records the value measured on the live instance (SHOW max_connections = 500, 2026-09-16, #21163). The gate compares the two and the first production plan run (35815654836) failed silently on that mismatch before Terraform ran. The verified default now lives beside the tier and version it is verified for, as VERIFIED_DEFAULT_MAX_CONNECTIONS, so the contract and the workflow are two independent records of the same measurement and the gate keeps its cross-check. The test pins the new source and forbids a bare literal. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb |
||
|
|
1dcbd4e65d |
test(opencode): pin the installed OpenCode plugin to a v2-loadable default export (#22389)
OpenCode 2 installs under the plain `opencode` executable name and its plugin loader requires the default export to carry `setup` (or `effect`) alongside an `id`; the v1 loader requires `server`. The emitted plugin source already carries both, but nothing asserted it on the file Orca actually installs into OpenCode's config directory — the exact surface that regressed in #22234. Load the installed file as an ES module and assert its default export satisfies both loaders. Reverting getOpenCodePluginSource() to the v1-only options makes this test fail with "expected undefined to be type of 'function'". |
||
|
|
0c2514e7e0 |
perf(mobile): keep four bundle chunk reads in flight across the whole manifest (OTA phase C follow-up) (#22376)
* perf(mobile): keep four bundle chunk reads in flight across the whole manifest The fetch ran one worker per asset and paged inside an asset sequentially, so the largest script's 71 chunks were 71 serial round trips while the other readers idled. One window of four chunk reads now covers every (asset, offset) on the host's chunk grid, largest asset first. A read_limited refusal narrows the window and retries the read; eof is still read from the reply. Synthetic manifest (one 71-chunk asset, five small): 72 round trips -> 19. Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb * test(mobile): re-record the bundle-fetch family under pipelined reads Baseline moves to |
||
|
|
bf3f95245c |
feat(relay): declare Asia cell c30 at the c27 shape (#22375)
* feat(relay): declare Asia cell c30 at the c27 shape Adds production-gce-c30 in asia-east2-a at the reviewed Asia shape (6,000 request units, 3,000/60 connection limits, 16-connection pool, disabled) and the rehome trust the other Asia cells carry. Every Asia enumeration now knows C30. The topology, admission, and director tools treat it as its own reviewed wave so its plan and registration never touch the live launch cells. C30 promotion requires C27 general and fresh staging evidence. The topology and director validators now pin the committed production pool of 16 instead of the stale 10, which had made the topology workflow reject the committed launch cells. * fix(relay): plan C30 at live images and prove it with its own canary The shared URL map pulls every cell into the C30 topology plan, so the workflow now plans each non-target cell at the image its live template serves, and the validator names any change to a cell outside the wave. C30 promotion runs the same five-minute production canary and automatic rollback C27 used, with the load report proving the canary control was placed on C30, instead of relying on staging evidence. C30 leaves the shadow gate's fleet pool list until it serves, rollback rejects mixed partial sets, and a budget test pins the mixed-Asia-pool refusal. * fix(relay): pin C30 to the production director's live image digest C30 promotion requires the director and C30 to report one digest, so C30 takes the director's sha256:4158d8a2 (read 2026-09-22). C27-C29 keep their committed lines; every Asia check compares only the cells named in a run. * fix(relay): read the committed cell map from a plan, not console terraform console evaluates every output against state, and the Relay deployments output indexes each cell's MIG, so it fails with Invalid index while C30 is declared but not created. Read the map from a no-refresh, unlocked plan over the same targets instead, and refuse empty overlay input. * fix(relay): keep console readers working and C30 migration-only until promotion relay_gce_cell_deployments indexed each cell's MIG, backend, and template, so once C30 is declared but not applied every production terraform console reader printed a warning to stdout and broke its jq parse. Wrap those six lookups in try(..., null). Same-cap listed C30 as general, so a rollback dispatch on a migration-only C30 would restore it with activate and skip its canary. List it with the migration-only cells until the promotion follow-up moves it. |
||
|
|
9af6a3d798 |
fix(cli): report a denied runtime connection instead of a dead Orca (#22341)
* fix(cli): report a denied runtime connection instead of a dead Orca Inside Codex's macOS Seatbelt sandbox, connect() on the runtime socket fails with EPERM. The CLI dropped the errno and reported "Could not connect ... Restart Orca", appended "Orca is not running. Run 'orca open' first.", and `orca status` answered ok:true with `starting` (its pid probe also gets EPERM). An agent following that advice restarts a healthy app, which cannot help. EPERM/EACCES on the metadata read or the socket/pipe connect now fails with a CLI-local `runtime_access_denied` error: ok:false, non-zero exit, operation/systemCode/processState:"unverifiable"/retryable:false and nextSteps that say to re-run with escalated permissions and not restart. CODEX_SANDBOX only picks the wording. `orca open` stops before launching. The status pid probe is unchanged: a refused or missing socket proves the caller reached the endpoint, so a later EPERM probe is another uid and keeps #20098's `starting`. Missing, refused, stale-pid and timeout paths are unchanged. Adapted from the diagnosis and tests in #20487 (and #19605, #13583). Co-authored-by: lifeodyssey <zhenjiazhou0127@outlook.com> * docs(skills): tell agents runtime_access_denied means escalate, not restart The shared CLI-resolution block told every bundled skill to run `orca open` when a command says Orca is not running. Add the counterpart for the new access-denied code so sandboxed agents re-run with escalated permissions instead of launching or restarting Orca. Regenerated stubs and manifest. * refactor(cli): classify only a denied runtime connect, with a leaner error A denied metadata read was never observed under a sandbox, and it turned an unreadable user-data path (the Linux launch contract's root-owned HOME) into runtime_access_denied instead of "Orca is not running". Keep metadata reads as on main and classify only the socket/pipe connect. One helper now maps a socket errno to the error or null; the error data keeps only systemCode and nextSteps. Tests drop cases already pinned by status.test.ts. * fix(cli): give not-running advice when a denied socket belongs to a dead Orca A crashed Orca leaves its metadata and socket file behind, and a sandbox denies the connect with EPERM before the CLI can see ECONNREFUSED. The sandbox still reports ESRCH for a gone pid, so a denied connect now probes the metadata pid and falls through to the ordinary unavailable path when the pid is proven gone. isProcessRunning moves to its own module so transport and status share it. * refactor(cli): inline the runtime_access_denied code like other CLI error codes --------- Co-authored-by: lifeodyssey <zhenjiazhou0127@outlook.com> |
||
|
|
83dd047fd9 |
fix(explorer): find files by name in large local workspaces (#22369)
* fix(explorer): search local workspaces by file name across every file The Explorer name filter only searched remote workspaces directly; local workspaces still filtered the first 20,001 listed files, so files beyond that cap never matched in large repos. Local name queries now rank the whole workspace on the host, and fall back to an uncapped git listing when ripgrep is not installed. * fix(explorer): filter capped local listings on the host with the Explorer word rule Replaces the Quick Open fuzzy top-32 routing, which dropped multi-word matches and capped visible results. The Explorer keeps its instant renderer-side filter; only when the local listing hits its cap does it re-list on the host with the same word rule applied before the cap. * fix(explorer): keep capped matches when host name filtering fails - Fall back to the capped listing (and stop re-listing) if the host scan fails - Keep primary matches when the ignored-file pass fails during a filtered scan - Key host scans on normalized filter words; reset capped state per filter session - Bound nameFilter size at the IPC boundary; drop the double readdir walk * fix(explorer): match name filters without locale-dependent lowercasing * fix(explorer): avoid render-time ref writes in the host name filter fallback |
||
|
|
ebed0964a2 |
feat(agents): add first-class Muse Code harness (#22216)
* feat(agents): add first-class Muse Code harness Add Muse as a supervised Orca agent across desktop, mobile, session history, source control, local hooks, SSH, WSL, and native Windows. Preserve user settings, support Muse 1.3 hook environment allowlists, and recognize versioned foreground processes. Include question, waiting, completion, resume, and readiness coverage. Co-authored-by: homesh-dev <300847526+homesh-dev@users.noreply.github.com> Co-authored-by: jeffhuen <32542276+jeffhuen@users.noreply.github.com> Co-authored-by: John Cusack <johncusackccm@gmail.com> Co-authored-by: Adrien De oliveira <75085839+adriendeoliveira@users.noreply.github.com> * test(agents): cover Muse remote hook registration * test(agents): cover Muse hook and source-control contracts * test(agents): exclude Muse hook metadata from script mode check * test(agents): keep Muse skill picker coverage stable * test(ai-vault): include Muse in every-agent fixture * test(mobile): repin Muse agent icon closure * fix(muse): detect questions and approvals from structured Muse signals Muse 1.3 fires no hook for request_user_input, so a pending question left the pane "working". Its internal reminder subagents also post hooks with their own session ids (even after Stop), which surfaced "tool failed" rows and flipped finished panes back to working. - Read pending questions from Muse's session log (user_input_prompt_requested/settled) via the existing transcript poll, now generalized from Codex subagents to Muse on main and relay. - Drop child-session hooks (SubagentStart ids, or turn_id === session_id). - Treat Notification permission_prompt as the approval wait; PermissionRequest also fires for auto-approved calls, so it only caches the approval card. - Ignore Notification copy as the prompt; poll replays are not new prompts or turn boundaries. - Allowlist USERPROFILE so Windows cmd AutoRun doesn't fail every hook. * perf(muse): parse only question events from the session log Most Muse session-log lines are large model/tool records. Filter raw lines by the user_input_prompt_ marker before JSON.parse via an optional readJsonlCursor line filter. * fix(muse): unwrap batched log records and scope questions to the live turn Review follow-ups: question events inside retained_frame batches were skipped, and a question left open by a crash or interrupt stayed pending for the pane's life. Share the history scanner's retained_frame unwrapper, and only report a pending question whose run_id matches the hook turn_id. * refactor(muse): drop type assertion in retained_frame unwrap * fix(agent-hooks): satisfy exhaustive-switch lint in transcript poll policy --------- Co-authored-by: Adrien De oliveira <75085839+adriendeoliveira@users.noreply.github.com> |
||
|
|
7c46a69049 |
feat(telemetry): report the macOS daemon's code identity on adoption and folder-denial events (#22171)
* feat(daemon): import the macOS process code-identity probe from PR #21826 Takes `daemon-mac-code-identity.ts` and its test verbatim from David Bebawy's community PR #21826 (stablyai/orca). The probe asks Security.framework, via `codesign --display --verbose=1 +<pid>`, where a live process's code lives on disk — the question Node cannot answer, and the one that decides whether tccd can still resolve a running daemon's code identity after an app update. Imported unchanged here so the adaptation that follows is reviewable as a diff against the author's original. Co-authored-by: David Bebawy <david.ayad2@gmail.com> * feat(telemetry): report the daemon pid's macOS code identity on the two adoption events Community PR #21826 argues that macOS terminal daemons lose Documents/Desktop/ Downloads access after an update because the daemon's own executable is unlinked — Squirrel parks the outgoing bundle under a ShipIt staging directory and later deletes it — so tccd can no longer map the daemon pid to on-disk code. Today's `spawner_path_class` and `tcc_attribution` read the binary that forked the daemon, which an in-place update deletes and recreates, so neither can see that state. This adds the detector as a measurement only. `code_identity` rides on `daemon_adopted` and `daemon_pty_cwd_denied`, the two events that already describe an adopted daemon, so denied daemons can be cross-tabbed against healthy ones. Nothing reads the verdict: no replacement, no notice, no UI. The probe is David Bebawy's, narrowed from a path-carrying union to the closed enum the wire allows, and memoised per pid so one codesign spawn answers for a whole daemon generation. Off macOS, or with no pid, it reports `probe-failed`, which keeps both schemas strict and non-optional. Co-authored-by: David Bebawy <david.ayad2@gmail.com> * fix(telemetry): read the daemon's code identity fresh on every adoption event The probe memoised its verdict per pid and never expired it, so `daemon_pty_cwd_denied` reported whatever the probe saw at adoption rather than what was true at the denial. That breaks the measurement in both directions: a transient codesign failure during startup pinned `probe-failed` for the rest of the run, and the `parked` to `unresolvable` transition became invisible. Squirrel leaves the parked bundle in place until the next update, which can be days, so a daemon adopted as `parked` and denied as `unresolvable` is the exact crossover this study exists to catch, and the cache hid it. Now every ask runs its own codesign. Only concurrent asks about the same pid share a probe, and that entry is cleared as soon as it settles, so nothing survives to be reported later. Both events are rare enough that one spawn each is not worth a cache. * fix(telemetry): drop the dead existence check from the code-identity probe The classifier stat'd the path codesign displayed and called a missing one unresolvable. That path is unreachable: once the executable is unlinked, `codesign --display` prints no `Executable=` line at all and exits 1 with "No such file or directory", which the fallback below already classifies as unresolvable. Verified directly on Darwin 25.5 against a signed binary deleted out from under a running pid. All the branch actually covered was the window between codesign reading the path and this process stat'ing it, and it paid for that with a synchronous stat on the main thread. * fix(telemetry): never classify a timed-out codesign probe as a verdict `runProcess` kills the child at the deadline and reports `timedOut`, but the runner type dropped that field, so a codesign killed mid-display could still have printed an `Executable=` line and been read as `resolved` or `parked`. A half-written display proves nothing about where the daemon's code lives. The runner result now carries `timedOut`, and a timed-out probe returns `probe-failed` before the output is looked at. * docs(telemetry): state what each code-identity verdict actually asserts A reviewer read `resolved` as a claim that the executable sits inside the installed app and asked for that to be validated. It is not that claim, and we are not making it: proving containment needs the pid record's spawner path, and deciding anything from where the code lives is #21826's proposed behaviour rather than this measurement. The enum doc now spells out all four verdicts in the terms the probe can actually support, and says plainly why `resolved` stops at "exists and is not parked". A matching note sits beside the parked-path pattern. * docs(telemetry): stop asserting how long a parked bundle survives The probe's rationale claimed Squirrel keeps the parked bundle "until the next update". A reviewer claimed the opposite, that it is deleted at the end of the same install. Neither holds up against this Mac's ShipIt log: the install moves the outgoing bundle to a TMPDIR ShipIt directory and logs no removal of it at all, and the one "Couldn't remove owned bundle" line names the incoming download staging copy, not the parked one. Every parked bundle from the last two days is nevertheless gone now. So the rationale in the probe doc, the enum doc, and the reprobe test comment now assert only what is established: the outgoing bundle is moved aside at install and disappears later on a schedule we have not pinned down. That is already enough to justify the design, since one pid's verdict can change within an app run, which is exactly why every ask reads fresh. * feat(telemetry): report readable TCC-gated spawns as the code-identity control `daemon_pty_cwd_denied` gives code_identity's hit rate on denials, but a readable spawn emitted nothing, so an `unresolvable` adoption with no denial could not be told apart from a user who never opened a terminal in Documents, Desktop, or Downloads. The false-positive rate that gates #21826's auto-replacement was unmeasurable. `daemon_pty_cwd_readable` now fires when a daemon reads a TCC-gated cwd, once per daemon and folder class per app run, with the same origin properties as the denial event. The read-out becomes a 2x2 of code_identity against readable/denied on protected-folder spawns. Fire-and-forget on the spawn path like the denial emit, and no app-side directory read. * refactor(telemetry): one emitter and schema for both cwd verdicts, no dedupe state The once-per-daemon dedupe on `daemon_pty_cwd_readable` was keyed before the probe ran, so a daemon first seen readable while `parked` never reported again once it turned `unresolvable` — the one cell that would count most against #21826. It also counted per daemon while denials count per spawn, so the 2x2 mixed units. Readable now reports every spawn, like denied, and both events share one emitter (`trackDaemonPtyCwdVerdict`) and one schema. The TCC-folder gate lives in the verdict branch. The origin fields are one shape spread into both schemas. The codesign probe calls `runProcess` directly and tests mock it, replacing a test-only runner parameter. The repeated "never cached" rationale is now said once. * fix(telemetry): rename the shared origin schema fields for the anti-slop gate no-shape-in-symbol-names rejects daemonOriginShape; the fields are event props. --------- Co-authored-by: David Bebawy <david.ayad2@gmail.com> |
||
|
|
1b85be67d8 |
feat(native-chat): notify on every settled structured turn (#22105)
* feat(native-chat): notify on every settled structured turn
A structured chat that finished while you were elsewhere lit the sidebar
but never raised an OS notification, and a notification that did arrive
for one could not open the chat it came from.
Unread and delivery now come out of the single resolveAgentAttention
decision the terminal lane already uses: the structured dispatcher calls
applyAgentAttention instead of applyAgentAttentionUnread, so the same
policy that decides what to light also decides what to deliver, through
the same sound and blocked-permission tail.
Every settled turn notifies, as the CLI lane does. Success says
"finished"; failure and cancellation say "stopped" through the shipped
agentInterrupted flag rather than a second vocabulary. A turn whose
outcome the host never stated stays unknown and lights nothing.
The host now dedupes mobile fan-out by event identity (scope, session,
turn) beside the existing per-workspace burst cooldown, so a completion
two windows both saw reaches the phone once while each window still
decides its own banner. Clicking a structured notification reveals the
chat tab: its pane key's leaf is synthetic, so focusTerminal would hunt
a split-layout leaf that does not exist.
* fix(notifications): spend each mobile gate only when it actually notifies
Two review findings on the structured-chat notification lane, both real.
The mobile event gate consumed its reservation before the per-workspace
burst cooldown ran. Two chats in one workspace share that cooldown key,
so the second chat's completion could burn its event key and then lose
the cooldown to the first chat — never announced, yet permanently marked
as announced, so a later window dispatching it could no longer reach the
phone. The gate now peeks first and records the event at dispatch, which
also keeps a known duplicate from burning the cooldown slot.
A notification id is minted from the status row's stateStartedAt, and the
row re-projects that field as the turn settles: the working episode's
start moves into stateHistory and the settled start takes its place. A
banner raised in the window before that re-projection therefore carried
an id acknowledgement never rebuilt, leaving it on screen for good.
Acknowledgement now collects ids for the row's left episodes too — the
same episodes the unread check beside it already scanned, so the two
halves finally read the same turns. Lane-neutral: the terminal lane
mints its ids the same way and had the same gap.
* fix(notifications): drop the mobile event gate and reveal chats in folder workspaces
The per-event mobile dedupe defended against one completion being
dispatched by several Orca windows. Only one renderer mounts the
structured attention bridge, the completion feed is live-only with no
replay, and any in-process duplicate lands inside the existing 5s
per-workspace burst cooldown, which already collapses mobile and
desktop alike. The gate never acted on a real sequence, so the wire
field, the shared ledger and its tests go; mobile delivery is back to
main's behavior.
A folder workspace id ("folder:<id>") has no "repoId::" prefix, so the
click binding was skipped and clicking a chat notification there did
nothing. The chat route selects its workspace itself through
ui:focusEditorTab, so it now binds without a repoId; the terminal
route is unchanged.
* fix(notifications): retire the banner ids actually dispatched, not ids rebuilt from a moved row
A banner's id is minted from the status row's stateStartedAt at dispatch, and that field moves
afterwards: a completion can outrun the settled re-projection, and a settled structured row is
re-stamped with no history entry by any later journal row (a cancel appends a status note after
the turn settles). Rebuilding ids from the row's episodes at acknowledgement missed the second
case and fanned out up to 21 mobile dismissals per pane for ids never raised.
The shared delivery tail now records each dispatched id per subject; acknowledgement retires
those plus the current-row rebuild it always had. The acknowledgement collector is back to
main's single-field form.
* refactor(notifications): retire announced notifications by subject in main
Main now records, per pane, the ids it actually announced (a desktop banner
shown or a phone alert sent) and an acknowledgement passes the acknowledged
pane keys so main retires all of them. This replaces the renderer-side record
of dispatched ids: main is where the announcement happens, so it records only
real announcements, including phone alerts whose desktop banner focus
suppressed. The id rebuilt from the current row stays as the fallback after
a restart empties the in-memory record.
|