* Fall back to a merge when a divergent pull has no reconciliation strateg
- Git 2.27+ refuses `git pull` on divergent branches unless pull.rebase or
pull.ff is configured. Retry with `--no-rebase` (Git's historical default)
so pulls succeed out of the box on fresh hosts.
- Skip the fallback whenever the caller already specified a reconciliation
strategy (e.g. --ff-only, --rebase) so explicit policies still fail as
expected on divergence.
- Applied identically in the local git pull path and the relay/SSH git
handler so both surfaces behave the same way.
* Refactor divergent-pull merge fallback into shared helper
Extracts the retry-as-merge logic (duplicated between local git and
relay SSH pull paths) into `runPullWithDivergenceFallback` in
git-remote-error.ts, so both callers share one implementation and
test coverage.
* Prevent index churn from refreshing worktrees
* Cover IPC contract in worktree reliability gate
* Refresh background worktree heads without re-entering structural fanout
External commits, amends, and soft resets in non-active worktrees now reach
store rows through spawn-free Git metadata reads diffed in the watcher's
existing debounce, emitted only on real head moves. HEAD reflog appends become
status-only triggers, config.worktree becomes structural for sparse-flag
freshness, and the non-darwin poller gains a periodic ungated index re-stat
so in-place rewrites on coarse-mtime filesystems cannot be missed forever.
* Reject unsafe symref paths and validate object ids in the head reader
Ref content comes from repo files an attacker can craft. Backslash segments
traverse on Windows where join treats them as separators, and colons are
forbidden in Git ref names; both now fail isSafeRefName before any path is
built. Resolved values are additionally emitted only when they match a hex
SHA-1/SHA-256 object id, so no file content can leak through the identity
event even in principle.
---------
Co-authored-by: Brennan Benson <brennanbenson@Brennans-MacBook-Pro.local>
Co-authored-by: Brennan Benson <>
* fix(agent-status): surface Claude tool failures
* Fix compact sidebar hiding tool-failure errors behind stale tool name
Extract a shared clearActiveToolFieldsUpdate() helper and apply it to
Cursor's postToolUseFailure, Copilot's PostToolUseFailure/ErrorOccurred,
and Grok's post_tool_use_failure events, matching the existing Claude
behavior so the failure message surfaces instead of the last tool name.
* fix(agents): bundle agent icons instead of loading them from Google's favicon service (#8451)
Agents without a hand-authored SVG glyph loaded their icon live from
Google's favicon service (www.google.com/s2/favicons). That service is
unreachable in some regions (e.g. mainland China) and offline, so ~23
agent icons rendered as broken images on the agent settings page, the
terminal title bar, and the status bar.
Bundle each favicon as a build-time asset under resources/agent-icons/
and render it via a new agent id -> URL map (agent-favicon-assets.ts).
The remote favicon service now only serves as a last-resort fallback for
any future agent that lacks a bundled icon. Follows the same pattern as
#7373, which bundled the OpenCode mark.
* fix(agents): bundle mobile agent icons too; drop dead omp faviconDomain (#8451)
Mobile had the same offline/region bug: MobileAgentIcon rendered every
non-glyph agent from Google's favicon service. It actually affected more
agents than desktop, since mobile lacks hand-authored glyphs for
Copilot, OpenCode, Kilocode, Droid, and OpenClaude — all fell through to
the favicon path.
Bundle the 28 favicon-path icons under mobile/assets/agent-icons/ and
render them via a Metro static require() map (mobile-agent-icon-assets.ts).
A node-env invariant test asserts every favicon-path agent ships a
bundled PNG and is wired into the map.
Also remove omp's vestigial faviconDomain from the desktop catalog — omp
renders the hand-authored OmpIcon glyph, so the favicon fallback was
never reachable.
* refactor(agents): share one set of bundled agent icons between desktop and mobile
Desktop and mobile each shipped their own copy of the favicon PNGs (23 +
28, with 23 byte-identical duplicates). Consolidate them into a single
source of truth at src/shared/agent-icons/, reachable by both bundlers:
- Desktop (Vite) imports them via `?url`.
- Mobile (Metro) requires them; Metro already watches src/shared via
metro.config.js sharedRoot, so no config change is needed.
The two per-platform maps stay separate because the import syntax differs
(`?url` string vs `require()` asset ref), but they now point at the same
files. Verified with a real `expo export`: Metro bundles all 28 shared
icons from src/shared/agent-icons.
* fix(agent-status): label Cursor by identity, not a bare "cursor" token
The worktree card, status bar, and mobile all derive an agent label from the
terminal title via getAgentLabel / resolveTerminalTitleAgentType. Both matched
Cursor with `titleHasAgentName(title, 'cursor')`, a whole-token match. But
`cursor` is ordinary editor vocabulary, so a Claude/Codex tab working on Orca's
own code (title like `⠋ preserve cursor visibility across replays`) got
mislabeled as Cursor. The generic braille-spinner Claude fallback even had a
`!lower.includes('cursor')` guard that then dropped the title to no label at
all.
Gate Cursor on its closed identity title set (`isCursorAgentTitle`) instead —
the same predicate @cursor orchestration routing uses. A real cursor-agent
terminal still resolves as Cursor across working/idle/permission; a non-Cursor
tab that merely mentions a text cursor reverts to its true agent. Relax the
braille guard to the same predicate so those titles land on Claude, not null.
Makes display consistent with routing (the follow-up flagged in #8436).
* refactor(agent-status): address review on Cursor identity labeling
- Trim the four Cursor `// Why:` comments in both parallel resolvers
(agent-title-identity.ts, terminal-title-agent-type.ts) to AGENTS.md's
one-to-two-line rule; use identical wording so future drift is visible.
- Add a direct isClaudeAgent assertion in terminal-title-agent-type.test.ts
pinning that file's parallel copy (previously only covered transitively),
plus Cursor Agent / "Cursor - action required" activity-facet assertions.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* feat(diff): add F7/Shift+F7 keyboard navigation for diff changes
Stacks on the Previous/Next change buttons (#6668) to add keyboard
navigation for single-file diffs, matching VS Code / JetBrains diff review.
- Register editor.nextChange (F7) / editor.previousChange (Shift+F7) in the
keybinding registry (Editors group) so they show in Settings and stay
rebindable.
- Teach the keybinding normalizer function keys (F1-F24) and make them
first-class in the bare-key safety model (safe standalone or with Shift,
opt-in per action) - F7 was previously unbindable.
- Install a capture-phase listener from DiffNavigationProvider so keyboard
and the existing header buttons share one goToDiff path; works on
read-only and editable single-file diffs.
- Translate the Previous/Next change strings for es/ja/ko/zh.
Refs #6215
* test(diff): cover change navigation shortcuts
* fix(diff): use shortcut chips in navigation tooltips
---------
Co-authored-by: Brennan Benson <brennanbenson@Brennans-MacBook-Pro.local>
* fix(pr-comments): let users mark comment authors as bots for the Humans/Bots filter
Some review bots post from regular user accounts that defeat both provider
bot metadata and login heuristics, so their comments were misclassified as
human. Adds a persisted prBotAuthorOverrides setting with a "Mark author as
bot" comment action, applied consistently across desktop and mobile.
Fixes#7597
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(pr-comments): address review feedback on bot-author overrides
- Cap sanitized prBotAuthorOverrides at 500 entries so malformed payloads
can't bloat GlobalSettings or slow comment classification
- Reuse the shared normalizePRCommentAuthorLogin in isBotPRComment on
desktop and mobile instead of duplicating the normalization inline
- Pass botAuthorOverrides from CommentRow to CommentMoreMenu instead of
re-subscribing per menu instance
- Re-fetch mobile bot-author overrides alongside each PR refetch so they
don't stay a stale one-shot snapshot for the whole session
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* fix(pr-comments): harden bot author override sync
* fix(pr-comments): bound and recover override updates
* fix(pr-comments): merge overrides from canonical settings
* fix(pr-comments): make bot override updates atomic
* fix(pr-comments): surface rejected bot overrides
* fix(i18n): translate bot override warning
* fix(i18n): translate bot author actions
---------
Co-authored-by: Dzmitry Bachko <dbachko@users.noreply.github.com>
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Brennan Benson <brennanbenson@Brennans-MacBook-Pro.local>
* Surface the Pi CLI's real error when branch auto-naming and commit generation fail
Pi failures (missing provider credentials, HTTP 4xx/5xx, connection errors)
previously collapsed to the generic 'Pi CLI command failed with code 1.'
because extractAgentErrorMessage only recognized Error:-prefixed lines.
Add two stderr-only extraction passes for pi's failure formats and narrow
the unix-path redactor so pi's /login remedy token survives redaction.
Fixes the OP variant of STA-1492 (#7808).
* Replace per-CLI failure parsing with sanitized raw output excerpts
Every agent CLI formats errors differently, and the parsing passes only
ever covered the vendors someone had already debugged (the pi passes
fixed pi and nothing else). Show the output itself instead: a positional
excerpt (first two non-empty lines plus the last one) of stderr, falling
back to stdout when stderr is silent, path-redacted and capped as
before. Exit-0 runs with empty stdout now stay 'returned an empty
<result>' instead of misreporting a command failure. The sanitizer also
strips Cf characters (bidi overrides) now that provider-controlled
bodies flow through verbatim, and OSC sequences are stripped alongside
CSI.
* Keep the full CLI output of failed branch-name generation for on-demand viewing
The persisted rename-failed badge carries only a sanitized excerpt (it
syncs to paired clients), so the complete diagnosis was previously
buried in the main-process log. Failed generations now capture their
full stdout/stderr (bounded to 64 KiB per stream, head and tail) into a
main-memory-only store keyed by worktree — never persisted or synced.
The rename-failed dialog fetches it on demand and shows it in place of
the excerpt, ANSI/control/bidi-stripped; after a restart or on paired
web clients it falls back to the excerpt. Renderer-bound generation
results explicitly drop the capture so IPC payload shapes are
unchanged.
* Cover the rename-failed dialog's full-output fetch and excerpt fallback
* Split the folder-workspace title rename flow into its own module
first-work-branch-rename.ts sat at the max-lines ceiling; the folder
title flow is a self-contained concept and main's concurrent additions
to this file pushed the CI merge ref over the limit.
* Harden the failure-output capture and dialog against review findings
- Flatten bounded captures so V8 sliced strings no longer retain the
multi-megabyte parent stream in the capture map (128 MiB -> ~4 MiB in
a 32-entry probe).
- Require an OSC terminator and stop its char class at newlines so an
unterminated/boundary-truncated sequence can no longer swallow the
omission marker and diagnostic tail.
- Exclude stdout from the persisted branch-name failure detail (it can
echo the prompt into synced metadata); the full local-only capture
still keeps it for the dialog.
- Refetch and reset the dialog's full output when the persisted error
changes so an open dialog never shows or copies a stale run.
- Report signal-terminated generators (null exit code) as 'was
terminated before exiting' instead of 'code null'.
* Redact JSON-escaped Windows drive paths without breaking scheme URLs
Provider JSON bodies double backslashes (C:\\Users\\name), which slipped past the drive-letter redaction into the persisted, client-synced failure excerpt. Allow repeated backslashes only — a URL's :// must stay single so remedy links survive.
* Fix branch rename failure IPC re-registration
---------
Co-authored-by: Brennan Benson <brennanbenson@Brennans-MacBook-Pro.local>
* fix(orchestration): complete worker tasks and improve coordinator UX
* Fix orchestration lifecycle sender resolution and peek/check compat hand
- Lifecycle sends (worker_done/heartbeat) now use ORCA_TERMINAL_HANDLE
verbatim, skipping the liveness probe and pane remint that could
block delivery during restarts or mismatch stale-runtime assignee
handles.
- --peek now round-trips as {peek:true, unread:false} so older runtimes
that strip unknown params degrade to non-destructive "all" instead of
mark-read, with client-side filtering to restore peek semantics and a
clear error when --peek --wait can't be honored.
- Reject combined read-mode flags (--unread/--peek/--all) before calling
the runtime.
- Distinguish suppressed (already-consumed) lifecycle messages from
ignored ones so send doesn't wake --wait waiters for stale heartbeats.
- Fix task summary truncation to avoid splitting UTF-16 surrogate pairs
and to not misreport whitespace normalization as truncation.
* Add shared helper to abbreviate orchestration task specs for brief listi
- Normalizes whitespace and caps spec length at 160 chars, flagging
truncation separately from whitespace-only changes
- Truncates on UTF-16 code point boundaries to avoid splitting
surrogate pairs and emitting malformed strings
* Add pane-key identity to worker_done/heartbeat reconciliation and server
- Records the sender's pane key on messages and dispatch contexts so
worker_done/heartbeat ownership can be verified by the remint-stable
pane leaf instead of the terminal handle, which is reissued across
restarts.
- Rejects lifecycle messages from a genuinely foreign pane while still
tolerating handle remints, tab break-outs, and older CLIs that lack
pane identity.
- Moves task-spec abbreviation server-side (orchestration.taskList
--brief) so full specs no longer cross SSH/relay transports, with a
client-side fallback for older runtimes; consolidates the shared
abbreviation helper under src/shared.
- Adds a stderr warning when a pre-peek runtime's --peek response hits
the 100-row cap, since older unread messages may be missing.
* Isolate ORCA_PANE_KEY in CLI test beforeEach to fix leaked senderPaneKey
Co-authored-by: Orca <help@stably.ai>
* Fix pane-key remint bypassing dispatch mutual-exclusion lock
- Dispatch locking only matched on assignee_handle, so a reminted
terminal handle (tab break-out) could open a second concurrent
dispatch on the same pane.
- Add leaf-UUID-based pane key comparison (parsePaneKey) as a
secondary lock, falling back to exact handle match for legacy
rows without pane keys.
* Update orchestration skill docs for lifecycle authority and CLI flag add
- Clarify that dispatch lifecycle is tied to taskId+dispatchId verified against
the dispatched pane, not the terminal handle, since handles can be reminted
after restart
- Document new `check --peek`/`--all` and `task-list --brief` flags, with
fallback guidance for older CLIs that reject them
- Note that a valid worker_done auto-completes the task/dispatch, so workers
shouldn't also call task-update manually
---------
Co-authored-by: Orca <help@stably.ai>
Persist and safely restore the active top-level view on startup. Unknown, removed, legacy, or unavailable views fall back to the terminal, while cross-window UI sync cannot navigate the current window.\n\nCloses #8264
* fix(settings): make WSL skill commands pasteable (#7795)
* Fix WSL skill commands so PowerShell 7 pastes match PowerShell 5.1 argv
- Encode the WSL login-shell script as base64 and decode/eval it inside
the sh -c invocation, avoiding raw nested quotes at the paste boundary
- Scope $PSNativeCommandArgumentPassing = 'Legacy' to the invocation so
PS 5.1 and PS 7 both hand wsl.exe the same escaped argv
- Extract powershell-native-argument.ts as the shared quoting module and
reuse it from ssh-remote-powershell.ts
* test(runtime): stub getRepo in mobile-tab startup cwd test
Main's #7892 made listMobileSessionTabs validate selectors via
this.store?.getRepo; the mock store here only defined
getWorkspaceSession, so the merged CI build threw 'getRepo is not a
function'. Return null (wt-1 is a worktree id, not a repo).
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Co-authored-by: Orca <help@stably.ai>
* Fix Codex WSL project trust conflating case-distinct Linux paths
normalizeCodexProjectPathForLookup lowercased the entire Windows/UNC
path, including the case-sensitive Linux portion under \\wsl$\<distro>.
Two distinct WSL project dirs (.../Repo vs .../repo) collapsed onto one
trust key, and the mirror dedupe (codex-config-mirror) inherited it.
Preserve case for the Linux path after the case-insensitive
\\wsl$\<distro> / \\wsl.localhost\<distro> share prefix; true Windows
drive letters and normal UNC shares still case-fold as before.
Co-authored-by: Orca <help@stably.ai>
* Apply the same WSL case-fold fix to hook-trust key lookup
normalizeHookTrustKeyForLookup had the identical latent bug: on a win32
host it lowercased the whole Windows-shaped path, folding the
case-sensitive Linux tail of a \\wsl$\<distro> UNC hook path — contrary
to its own comment that WSL sources stay case-sensitive.
Extract foldWindowsCaseInsensitivePath (shared with the project-path
normalizer): fold only the case-insensitive drive/\\wsl$ share prefix,
preserve the Linux tail. Existing suffix (event:group:handler) is already
lowercase, so behavior is unchanged there.
Co-authored-by: Orca <help@stably.ai>
* Fix Codex WSL trust revocation and hook-key case folding
- Case-drifted or share-spelling-varied WSL revocations were being
ignored on merge, letting stale "trusted" entries survive; fold
revocation lookups fully so matching errs toward revoked.
- Hook-trust key folding was gated on process.platform === 'win32',
so WSL/SSH-remote hook sources weren't folded when Orca itself ran
on macOS/Linux; fold by path shape instead, via the new shared
foldWslUncPathCaseInsensitiveParts helper (also covers /mnt drvfs
automounts and wsl$/wsl.localhost share aliasing).
* Fix Codex WSL trust key case-folding regressions
- Don't fold case-variant `/MNT` dirs as if they were the drvfs
automount; only literal lowercase `/mnt/<drive>` folds.
- Preserve an exact-cased trusted project entry in ~/.codex during
runtime merge instead of letting a loosely-matched, case-drifted
revocation clobber a user's re-granted trust on every mirror pass.
- Minor cleanup: inline foldWindowsCaseInsensitivePath and hoist the
repeated normalizeHookTrustKeyForLookup call in findTrustBlockRanges.
* Refactor case-fold WSL trust test to assert real serializer output
Extract path variables and assert against escapeTomlString(incomingPath) instead of a hardcoded escaped string literal, so the fixture can't silently drift from the actual TOML header serialization.
---------
Co-authored-by: Orca <help@stably.ai>
* fix: work-item naming, usage % rounding, and terminal/delete copy
Address prod-release-scan P2s:
- #8238: recognize Bitbucket Server (/projects|users/.../repos/.../pull-requests/N)
and Azure DevOps (/_git/REPO/pullrequest/N) PR URL shapes in
work-item-reference, alongside Bitbucket Cloud; graceful fallback preserved.
- #7574: getDisplayedUsagePercentage now rounds the used value before taking the
`remaining` complement, so the compact status bar (raw usedPercent) and tooltip
(pre-rounded clampUsedPercent) can no longer disagree by 1% at a .5 fraction.
clampUsedPercent moves to the shared module as the single rounding source.
- #7459: mixed remote+local batch delete confirm no longer claims the whole
batch is a permanent "remote host" delete — it now states remote items are
permanent while local items move to the Trash/Recycle Bin.
- #8322: right-click-to-paste settings copy is platform-aware — "Control-click"
on macOS, "Ctrl+right-click" on Windows/Linux — matching the ctrlKey gate.
Localization catalog synced (also picks up pre-existing UsagePercentageDisplayChangeNotice drift).
* Fix NaN% usage bar and label for non-finite provider values
Non-finite usedPercent inputs (NaN/Infinity) propagated through Math.round/min/max into the CSS bar width (`NaN%`) and displayed copy. clampUsedPercent now short-circuits to 0 in that case, with a test covering the divergence from getDisplayedUsagePercentage for the 'remaining' case.
Pi awaits its extension event handlers, so an awaited loopback status
post that stalls (Orca restarting / receiver unavailable) blocked the
running Pi turn and disconnected it. Make post() fire-and-forget with a
latest-only pending slot drained by a single active request and a 1s
AbortController timeout, so a stalled receiver can never hold the turn.
Also stop treating session_shutdown as turn completion: Pi emits it on
reload/new/resume/fork while the PTY stays alive, so only agent_end
proves done (real exit is cleared by PTY teardown). Split the generated
handler registrations into agent-status-handler-source.ts.
Co-authored-by: Orca <help@stably.ai>
* docs: design Grok orchestration group
* docs: plan Grok orchestration group implementation
* fix: add Grok orchestration group
* test(orchestration): accept Windows skill newlines
* Fix @grok orchestration group matching and remove stale planning docs
- Reuse the shared buildAgentNameRe matcher in groups.ts instead of a
divergent local regex, so orchestration groups honor the same
Windows launcher-suffix rule (grok.exe/.cmd/.bat/.ps1) as the rest
of Orca's agent-title detection.
- Add test coverage for real Grok OSC title shapes (spinner-collapsed,
session titles) and Windows launcher-suffix titles.
- Delete the now-completed design and implementation-plan docs for
the Grok orchestration group work.
---------
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Include fallbackGitHubPR alongside linkedGitHubPR/linkedGitLabMR when
determining whether a hosted review link resolves to a push target.
Worktrees without persisted linkedPR metadata (e.g. child worktrees)
were incorrectly blocked with "target unavailable" despite having a
real matching upstream, since their PR was only known via the queue
fallback. Also splits hasPositiveHostedReviewNumberLink to build on
the resolvable subset so the two helpers can't drift.
* Fix F3: use known-tag predicate for explicit-turn decisions
isHarnessInjectedUserTurnText matches any multi-word kebab tag, which is
safe for non-destructive UI classification but wrong for state-transition
decisions: a real prompt starting with a custom <my-element> paste looked
like machinery in resolvePrompt/hasExplicitUserPrompt. Combined with the
post-interrupt working-suppression in agent-hooks/server.ts, that left the
agent visibly done after Ctrl+C. Switch both explicit-turn callsites to the
known-tag isKnownHarnessInjectedUserTurnText predicate; known harness tags
and Grok <user_query> behave unchanged.
Co-authored-by: Orca <help@stably.ai>
* Retire broad harness-tag matcher; unify on known-tag predicate
The broad isHarnessInjectedUserTurnText had one remaining consumer: session
title selection in session-scanner-primary-parsers.ts. Switch it to the
known-tag predicate too — a real first turn that pastes a custom <my-element>
now titles the session instead of being demoted to the meta (fallback) title.
All observed first-turn machinery (system-reminder, caveat, command-name,
task-notification, …) is already in the known list, so the only behavior
change is that unknown, uncatalogued kebab tags stay user turns. Delete the
now-unused broad predicate and fold its coverage into the known-predicate
tests.
Co-authored-by: Orca <help@stably.ai>
* Fix bare <channel> tag being misclassified as harness machinery
Only the attributed `<channel source=…>` form is emitted by the harness;
a bare `<channel>` is legitimate user-pasted RSS/XML content. Remove
'channel' from the known-tag set (which matched any <channel ...>) and
rely solely on the existing '<channel source=' prefix rule.
---------
Co-authored-by: Orca <help@stably.ai>
- gh file-fetch failures (rate limit, auth, unresolved remote) previously
returned an empty array, which the Files tab rendered as "No files
changed." — indistinguishable from a real empty PR
- getPRFiles now returns null on failure; work-item-details surfaces this
as filesUnavailable so GitHubItemDialog and PullRequestPage can show a
retry action instead of a misleading empty state
* feat(status-bar): notify upgraded users usage meters show % used
Show a one-time status-bar callout when upgraded profiles still use the
new percent-used default. Brand-new profiles and users who already chose
remaining stay quiet; dismissing or changing the setting is permanent.
* Add settings deep-link to expand Appearance's Window accordion for one-s
- Replaces the searchQuery-based redirect (fragile, flashed filter UI) with a
dedicated appearanceAccordionDeepLink store field that force-opens the
correct accordion and scrolls to the target row
- Rebuilds the status-bar change notice as a plain elevated card instead of
a Popover, since PopoverContent's glass/backdrop-filter defaults fought
the opaque callout styling and needed heavy overrides
- Simplifies the light/dark card CSS tokens accordingly
* Reposition status-bar usage-change notice via fixed-position portal
Portal the one-shot callout to document.body with fixed positioning
anchored via getBoundingClientRect, since the status-bar's overflow-hidden
flex ancestors clipped or mispositioned the previous in-tree absolute
layout.
* Refine status-bar usage notice styling and test coverage
- Replace hand-tuned light/dark card colors with existing design
tokens (--popover, --border) and the documented floating elevation,
so the callout stays in sync with the design system instead of
duplicating its own palette
- Add tests covering dismiss via X button, "Got it", and Escape
- Scope the foreground-process confirm assertion to the pane's ptyId
so an unrelated pane's delayed confirm can't cause a false failure
* fix(naming): lead workspace and tab names with the work-item identifier
Auto-generated workspace names, branch-rename display names, and tab
titles now lead with the referenced PR/MR/issue/ticket — e.g.
`PR 1033 - Review` instead of a paraphrase like "Review community pr
1094" that buried or dropped the number. Identifiers are the highest-
signal, most searchable token, so leading with them makes the sidebar
and tabs scannable and disambiguates same-verb work items.
- New shared `work-item-reference.ts` extracts the identifier from the
raw prompt: URLs are validated by path structure (owner/repo/pull/N,
GitLab's `/-/` marker) so GitHub Enterprise / self-hosted GitLab still
resolve while stray `/pull/<n>` paths (CDN, docs) do not; a ticket-
prefix denylist keeps `SHA-256` / `UTF-8` / `ISO-8601` from being read
as Jira/Linear keys.
- Reconciles the existing create-from-work-item naming (was action-first
`Review PR 1033`) with the first-work auto-rename onto one identifier-
first format via a shared `formatIdentifierFirst`, so the two paths
can't drift.
- Fixes a pre-existing tab-title bug where markdown punctuation was
stripped before URLs, splitting a GitLab `merge_requests` URL at its
underscore and leaking "requests" into the title.
* fix(naming): keep emphasis-wrapped URLs intact in generated names
The URL/markdown strip reorder left the tab-title URL strip anchored on
\b, which fails when a URL is wrapped in markdown emphasis (_...pull/5_)
and leaked URL fragments into the title. Drop the \b anchor and trim
trailing markdown emphasis (*_~) in the URL identifier parser, keeping
interior underscores (merge_requests) intact. Adds regression tests.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Co-authored-by: Orca <help@stably.ai>
Add a persisted Appearance setting that switches provider usage labels between percent used and percent remaining while keeping meter fill consumption-based.
Cover desktop and web persistence, current providers, settings search, localized copy, and regression tests.
Co-authored-by: gatsby74 <166927047+gatsby74@users.noreply.github.com>
* fix(terminal): let remote desktop viewers own the shared PTY width
A remote (relay/shared-control) desktop viewer resizes the host source PTY to
its own width, but was never registered as a width owner. So the host's own fit
cascade (window resize, split drag, tab reveal, "+"-new-tab re-render) freely
resized the viewed PTY back to the host-local width with no signal to the
viewer. The viewer kept its narrower grid while the host streamed wider
alt-screen frames -> cell-layout garble ("porridge") until a manual resize.
The mobile "presence lock" already solves this shape for phones by suppressing
the host's pty:resize while the phone drives. This does the same for remote
desktop viewers WITHOUT joining the mobile driver state machine (a viewer needs
only resize suppression, not input lock / phone-fit / driver banners), and it
routes every PTY geometry change through the existing enqueueLayout/applyLayout
serialization path rather than resizing the PTY ad hoc.
Design:
- Registry `remoteDesktopViewers: Map<ptyId, Map<subscriptionKey, viewport>>`,
keyed per SUBSCRIPTION so duplicate streams of one client cannot release each
other. isPtyResizeDrivenRemotely() = mobile driver OR any viewer present;
pty:resize (host fit cascade) bails on it. INPUT is never locked (shared
control: host and viewers can both type), unlike mobile.
- New internal layout target { kind: 'remote-desktop' } in applyLayout: clears
terminalFitOverrides like 'desktop', resizes only when dims changed, emits no
fit-override notifications, and does not call onExternalPtyResize (so a
viewer's width never pollutes desktop restore state). enqueueLayout remains
the sole serialized writer of PTY dimensions.
- Smallest client wins: applyRemoteDesktopLayout() sizes the PTY to the SMALLEST
attached viewer, so viewers with different screen sizes never overflow the
narrowest grid.
- No snapshot/replay race: a viewport that arrives while the initial scrollback
snapshot is being serialized is buffered and applied as ONE serialized resize
before serialization, so a mid-repaint alt screen is never baked into the
snapshot.
- Reclaim on detach: when the last viewer leaves, applyRemoteDesktopLayout
resizes the PTY back to the host's own width via enqueueLayout({kind:'desktop'}),
so the host reflows to full width without a manual window jiggle. The host
width is captured at the pty:resize suppression point (the host's own resize
attempts) — a source a remote viewport never pollutes.
- Mobile coexistence: a phone outranks a viewer; when the phone leaves, the
mobile-release paths call applyRemoteDesktopLayout so a surviving viewer keeps
the PTY, else the host reclaims.
Verified live (two local desktop instances): correct render on connect, no
garble on viewer-window resize (synchronous), and host reclaims full width on
disconnect. Unit tests cover the registry lifecycle, smallest-client-wins,
reclaim-via-layout, and mobile coexistence.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Codex <noreply@openai.com>
* fix(terminal): address remote-desktop PTY width review findings
Resolves four CodeRabbit findings on the remote-desktop viewer width
ownership change, plus a round-trip race surfaced while fixing the leak:
- Leak: the one-shot terminal.updateViewport RPC has no disconnect hook,
so it must never create a width floor (nothing releases it, pinning the
host at a stale width). Add refreshRemoteDesktopViewer, which only
refreshes floors the client already owns via its stream subscription
(matched by clientId), mirroring the mobile updateMobileViewport
no-op-without-subscription invariant. The one-shot handler now refreshes
instead of registering.
- Round-trip replay: with the one-shot fallback now refresh-only, a resize
landing during the subscribe round-trip (connected, stream not yet
current) was dropped. The transport now replays the latest viewport over
the stream once it becomes current.
- Reclaim: keep the host reclaim target unless the reclaim resize actually
landed (result.ok), so a failed reclaim can retry against true host
geometry instead of a stale remote width.
- Snapshot drain: drain pendingRemoteDesktopViewport in the
sendRequestedSnapshot finally block; a viewer resize parked during a
SnapshotRequest buffering window was otherwise dropped until the next
resize.
- Key scope: scope the multiplex width-floor key by connectionId so two
connections reusing the same client-local streamId can't overwrite or
release each other's floor.
Tests: refresh-never-creates-floor, reclaim-target-retained-on-failed-resize,
SnapshotRequest parked-resize drain, connection-scoped key assertions, and a
round-trip replay race test (verified to fail without the flush).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(terminal): arbitrate active desktop viewport ownership
Co-authored-by: Orca <help@stably.ai>
* test(terminal): expect viewport claim capability
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): preserve input across ownership teardown
Co-authored-by: Orca <help@stably.ai>
* test(terminal): keep subscribe buffer coverage within lint limit
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: shady <shady2k@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Orca <help@stably.ai>
* fix(automations): deliver prompt to Hermes TUI via process-ready signal
Hermes's prompt_toolkit TUI never emits the DECSET 2004 bracketed-paste
handshake that the default draft-paste readiness waiter gates on, so the
automation prompt was silently dropped and the agent sat idle (terminal
opens, nothing entered). Add a 'process-ready' DraftPasteReadySignal that
arms the quiet-window on first PTY output and pastes once the TUI settles,
and assign it to the hermes agent config.
Adds a Hermes unit test covering the no-handshake path.
Design for the follow-up PR/issue review loop lives in
docs/design/pr-issue-review-loop.md (not committed; gitignored).
* fix(automations): skip broken process-name fallback for process-ready
CodeRabbit: the process-ready fallback consulted waitForAgentReady, which
compares the foreground process basename via isExpectedAgentProcess. Wrapped
interpreter launches surface as 'python3 .../hermes', so the check can never
confirm readiness and only drops the paste. process-ready readiness is the
PTY-quiet window (handled by waitForAgentDraftInputReady); the process-name
fallback is dead for this signal. Skip it and return false on timeout instead.
* test(automations): assert single paste for Hermes quiet-window path
Guard against a duplicate paste in the no-handshake Hermes path by
asserting sendRuntimePtyInputVerified is called exactly once. Addresses
the CodeRabbit nitpick on PR #7862.
* fix(automations): skip broken process-name fallback on pty-bound paste path
pasteDraftToAgentPtyWhenReady (quick-create/work-item route) still ran the
python3-vs-hermes process-name fallback that can never match for
process-ready agents, burning ~1s and dropping the paste. Mirror the
process-ready guard from pasteDraftWhenAgentReady and surface onTimeout.
Adds regression tests for both the quiet-window happy path and the
fallback-skip on the pty-bound path.
Co-authored-by: Orca <help@stably.ai>
* fix(automations): defer submit Enter until the Hermes TUI is interactive
Live-testing against Hermes v0.18.2 showed the fixed 50ms post-paste Enter
is swallowed: the node ui-tui takes 15s+ to boot, the paste fires ~1.5s in
(process-ready quiet window), and the cooked-mode line discipline turns the
early \r into \n, which the editor treats as newline-insert. The prompt
parked in the input box and the automation never executed.
For process-ready agents, defer the Enter until the TUI signals
interactivity (DECSET 2004 enable) or echoes the pasted content marker-free
(legacy prompt_toolkit), with a 120s best-effort cap. PTY input is FIFO, so
the deferred Enter always lands after the buffered paste text. The paste
echo of the content itself (raw or caret-notation markers adjacent) is
rejected so a cooked-mode echo can't release the Enter early.
Verified end-to-end in the dev app: automation prompt pasted, deferred
Enter released on tui-ready, Hermes submitted and ran the turn.
Co-authored-by: Orca <help@stably.ai>
* fix(automations): use Hermes native startup query
Co-authored-by: Orca <help@stably.ai>
* fix(automations): preserve quoted Hermes queries on Windows
* fix(agents): preserve invalid argument error message
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Brandon Bennett <brandonbennett@macbookair.myfiosgateway.com>
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Jinwoo Hong <73622457+Jinwoo-H@users.noreply.github.com>
* Fix mobile terminal query reply authority
* fix(terminal): harden mobile query reply handoffs
* fix(terminal): exclude passive mobile query responders
* fix(terminal): gate mobile query replies on host capability
Older hosts strip terminal.send's inputKind (zod drops unknown keys), so a
forwarded xterm reply would land as ordinary floor-taking shell input. Hosts
now advertise terminal.query-reply-input.v1 via status.get and mobile drops
replies unless the host advertises it (pre-fix behavior). Also documents the
bounded desktop-to-mobile handoff double-reply residual.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): advance snapshot seq across recovery snapshots
The pending-overflow recovery loop trims buffered output against
recovery.seq while query replay and boundary strips kept using the
initial snapshot seq. Unreachable under today's control flow (no await
separates the initial-overflow consume from the loop), but the stale
seq would silently drop covered query replies if that ordering ever
changes. Track the seq that actually covered the buffered chunks.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* fix(win): resume quoted cmd.exe startup commands via stdin, not /K
Resuming an AI Vault session into a cmd.exe tab on Windows failed with
"'...' is not recognized as an internal or external command" and never
ran the resume. The queued command for a cmd live shell is the
self-contained `cmd /d /s /c "cd /d ""cwd"" && claude ""--resume"" ""id"""`
form, which is correct when typed into cmd's interactive parser (its ""
doubling is cmd's convention). But the local PTY provider embedded it in
the `/K` launch argument, where node-pty's C-runtime argv escaping emits
backslash-escaped quotes (\") that cmd.exe cannot parse — the command
arrived as `\"cd /d \"\"cwd\"\" && ...` and was rejected wholesale.
Unlike PowerShell's -EncodedCommand, cmd.exe has no robust argv-quoting
path, so any startup command containing a double quote now falls back to
stdin delivery, where cmd's interactive parser handles the "" doubling
correctly (verified end-to-end against a real ConPTY via node-pty).
Quote-free commands keep the `/K` fast path.
* fix(win): preserve cmd resume command contracts
* fix(win): match copied resume commands to shell
---------
Co-authored-by: Jinwoo Hong <73622457+Jinwoo-H@users.noreply.github.com>
* feat(agent-status): show Claude subagent child rows and gate premature done
A Claude pane that spawned background subagents/teammates showed a green
done check the moment the lead's turn ended, even while a background
review loop was still running. Orca now tracks the pane's live children
from Claude hook events and:
- keeps the pane 'working' while at least one child is working (Stop is
gated; Claude wakes the lead when a child finishes, so the pane
resolves to done on the follow-up Stop with an empty roster)
- renders the children as indented child rows under the pane's sidebar
row (name/type + working/idle dot), reusing the existing lineage UI
Tracking is lifecycle-primary: SubagentStart/SubagentStop/TeammateIdle
(newly registered hooks) plus child-origin tool events (they carry
agent_id) own the roster. Stop's background_tasks is folded only where
unambiguous — verified live on Claude Code 2.1.207 that teammates report
status "running" while idle-alive and their task ids never match
lifecycle agent_ids, so the list cannot decide teammate working-ness.
Child-origin events no longer overwrite the lead's tool/prompt caches
(a live AskUserQuestion card survives child churn); a child's own
PermissionRequest records waitingAgentId so only that child's progress
or death clears the wait. The interrupted flag survives the gated
window, inferred interrupts sync the lead record and refuse while a
child works, and hydration reseeds the roster after a restart.
* fix(agent-status): drop identity icon on subagent child rows
The child's agentType carries its NAME (e.g. "pr-reviewer"), which is not
an iconable agent and rendered the unknown "?" glyph. Nesting under the
parent row already conveys identity.
* fix(agent-status): restore displaced lead state and reconcile phantom subagents
Four review findings from the adversarial pass on the subagent child-row
feature:
- Stash the lead state a child-induced wait displaces
(ClaudeLeadTurnState.stateBeforeWait) and restore it when the wait
clears, instead of inventing 'working' — a lead that had already
stopped left the pane spinning forever after the roster drained,
since the done-gate only ever downgrades done → working.
- Tag snapshot-seeded and background_tasks-recreated roster entries
(backgroundTasksAuthoritative) and demote them when a PRESENT
background_tasks list omits their id. A phantom child seeded before a
restart could otherwise gate the pane 'working' indefinitely in teams
sessions, whose task list is never empty. Live activity clears the
tag so lifecycle-tracked teammates keep their state.
- Match teammate ids with a hyphen-free suffix after `a<name>-` so
TeammateIdle for "lane" cannot idle "lane-hooks"'s rows or clear its
pending permission wait.
- Route turn-boundary events (Stop/StopFailure/UserPromptSubmit) that
carry a KNOWN child agent_id through the child-driven re-emit instead
of adopting them as lead state, and tie the prompt-cache new-turn
reset to lead-origin events so child refreshes can't blank the
prompt label.
---------
Co-authored-by: Brennan Benson <brennanbenson@Brennans-MacBook-Pro.local>
* Add read-only View Log tabs so AI Vault agent sessions open inside Orca'
- Adds a `readOnly` flag on OpenFile that hard-blocks edits, autosave, dirty
state, rename, and drafts, and persists/restores it safely across sessions
- Wires AI Vault's View Log/Open Log actions to open logs as permanent
read-only local tabs instead of shelling out to the OS, gated to
local, single-file, non-synthetic session paths
- Registers a dedicated `jsonl` Monaco language (JSON-style coloring without
whole-document JSON validation) and forces read-only tabs to render as raw
source, bypassing markdown/mermaid/csv/notebook viewers
* Add live-tail streaming for local AI Vault View Log tabs
- Extend the fs:readFile snapshot path with byte-stable file identity so a
read-only tab can resume appending exactly where the snapshot left off.
- Add a ranged local log tail reader plus IPC (read/start/stop watch) that
streams only newly appended bytes, detects truncation/rotation, and
cleans up watchers on tab close or renderer destruction.
- Add a renderer-side UTF-8/line-boundary decoder and useLocalLogTail hook
that appends completed lines into the existing Monaco model, falling
back to a full reload on reset/rotation, keeping snapshot behavior
unchanged when liveTail is not opted in.
- Thread the new `liveTail` flag through OpenFile/PersistedOpenFile,
workspace session persistence/restore, and the AI Vault "View Log" open
path so live tail survives restarts and stays read-only-safe.
* Remove stale planning brief for the live-tail feature
The AI-VAULT-VIEW-LOG-LIVE-TAIL.md pick-up brief is no longer needed now that the live-tail streaming work has landed (c7fdef7b1).
* fix(terminal): send CSI-u Shift+Enter to kitty TUIs (droid) on Windows (#7620)
On Windows, Shift+Enter was always sent as the Alt+Enter byte ESC+CR (added in
#2418 for Codex, which reads win32-input-mode and ignores CSI-u). droid speaks
the kitty keyboard protocol, parses CSI-u directly, and treats ESC+CR as a plain
Enter — so Shift+Enter SUBMITTED the message instead of inserting a newline.
droid works in other terminals (Windows Terminal, Warp) because those honor
win32-input-mode / kitty; Orca (xterm.js) withholds kitty from local Windows
ConPTY panes and emits neither.
Make the Windows Shift+Enter byte pane-aware: latch whether a pane's program
advertised the kitty keyboard protocol (query CSI ? u, push CSI > .. u, or set
CSI = .. u) and send CSI-u (\x1b[13;2u) to those panes, keeping the
Codex-compatible ESC+CR for win32-input-mode-only TUIs. Non-Windows is unchanged
(always CSI-u).
Verified end-to-end against the real droid and Codex CLIs through the actual
production functions: droid now newlines, Codex still newlines.
* fix(terminal): route Windows Shift+Enter safely for Droid
---------
Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
* Cap automationRuns retention to stop unbounded state file growth
`automationRuns` was the only unbounded durable collection in
`orca-data.json`, and the whole blob is re-serialized and rewritten on
every save. On a machine running four `* * * * *` automations it had
grown to 11,184 rows / 21 MB of a 28.5 MB file — all of them
`skipped_precheck` no-ops — so each synchronous `flush()` blocked the
Electron main thread for 190-210 ms and macOS filed 24 `disk writes`
diagnostic reports against Orca (12-30 MB/s sustained, 549 GB/session).
Prune to the newest 100 runs per automation, on load and on append. This
mirrors the retention that `pruneLocalTerminalScrollbackBuffers` and
`pruneWorkspaceSessionBrowserHistory` already apply to their
collections; `automationRuns` was simply missed.
The load-path prune marks state dirty, so an oversized file heals on
first load. Without that flag the shrink lives only in memory: the sole
load-time save trigger is `normalized.changed || loadNeedsSave ||
adaptedProjectGroups`, and `normalized` covers pane identity only. A
user who took the documented workaround (`automations edit --disabled`)
fires no runs, so nothing would ever rewrite the file.
Measured against the affected 28.5 MB file: flush() 190 ms -> 9 ms,
bytes written per save 28.5 MB -> 1.3 MB.
Pruning breaks the old `runNumber` derivation, which counted retained
runs, so every run after the cap would have been titled "run 101".
Carry the ordinal on the run itself and derive the next number from the
highest survivor. Legacy rows are numbered from the highest number their
automation already carries, not from their append position: a downgrade
to a pre-`runNumber` build appends unnumbered runs after pruned
survivors numbered 101+, and a position would reissue one of those,
giving two runs the same title.
Fixes#8118
* Never evict in-flight automation runs from retention
A dispatched run's completion can land hours later (renderer round-trip or
headless completion promise); pruning it makes updateAutomationRun throw
'Automation run not found.'. Only final-status runs are evictable now, with
the final-status predicate shared between retention and the service.
Co-authored-by: Orca <help@stably.ai>
* Skip the usage write when retention evicted the run mid-collection
markDispatchResult finalizes a run, awaits usage collection, then writes
usage by id. The run is final during that await, so a concurrent
create-time prune can evict it and the write threw 'Automation run not
found.' — in the headless path that cascaded into an unhandled rejection.
Co-authored-by: Orca <help@stably.ai>
* Pin backfill-before-prune ordering with true legacy fixture rows
The heal-on-load fixture rows carried runNumber, so swapping backfill and
prune passed every test while renumbering real legacy survivors 1..100 and
re-minting colliding titles. Seed unnumbered rows and assert healed ordinals.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: Orca <help@stably.ai>
Previously a failed resize during reclaimTerminalForDesktop left the
override/lock in place (#7588 semantics), but for an explicit "take
back all terminals" gesture that could strand banners on background
panes whose resize can't converge. Now the take-back unconditionally
releases the driver and clears any held fit-override via a new
releaseDesktopTakeBack helper, while auto-restore and phone-initiated
paths keep the original keep-lock-on-failure behavior.
Also extract the local mapWithConcurrency helper out of
workspace-cleanup.ts into a shared, tested src/shared/map-with-concurrency.ts
and use it to bound concurrent desktop-fit reclaims in
terminal-fit-restore.ts.
* fix(linear): guard mixed-version RPC filtering
* fix(linear): surface filter capability failures correctly
Prevent capability checks from pinning to rejected compatibility cache
entries, and rethrow typed attribute-filter unsupported errors from the
Linear store so TaskPage can show an upgrade message instead of an empty
filtered list.
* fix(runtime): refresh cached capability verdicts
* test(linear): mock isLinearIssueAttributeFilterUnsupportedError
Prevents the invalidation slice test from failing after the runtime
client gained this export, which was otherwise undefined in the mock.
* Fix cold-cache capability probes firing duplicate status.get calls
Coalesce concurrent status.get requests for the same environment by
publishing the in-flight probe to the compatibility cache before
awaiting it, so parallel capability checks share one RPC call. On
failure, drop the cache entry immediately since this probe always
re-fetches and must not leave a stale cached verdict.
Multiplexer/session wrappers prefix pane titles as "prefix | pane-title", pushing the Pi/OMP identity after " | ". Inspect each " | " suffix and prefer a re-ownable compatible identity over the wrapper text, so wrapped OMP/Pi titles normalize to the owner instead of flickering. Includes guard tests for braille-inner labels and no-identity wrappers.
* fix: keep dispatch task preview after single-line status normalization
Agent-status prompts fold multi-KB dispatch preambles to ~200 single-line
chars, which previously kept only lifecycle boilerplate and dropped the
TASK body. Compact dispatch prompts to preserve task id + body, extract
that as the UI fallback before orchestration labels arrive, and prefer
richer labels once they match the live dispatch.
* fix: harden dispatch status prompt compaction
* fix: share dispatch task-marker parsing with UI preview helpers
Round-2 adversarial review found getOrcaDispatchTaskPreview still used
naive indexOf, so raw multi-line preambles with base-drift subjects
mentioning === TASK === could surface the wrong fallback label. Export
the standalone-line marker finder and use it on both compaction and UI
preview paths, with regression coverage for raw and single-line forms.
* fix(agent-status): filter harness-injected turns by tag shape, not a literal list
Harness machinery turns (task-notification, agent-message, bash-input/stdout,
system-reminder, …) were leaking into prompt-derived UI as if the user typed
them — hook-status labels, native-chat bubbles, and AI Vault session titles.
The literal tag blocklist went stale twice in three months because the harness
adds tags faster than we chase them.
Generalize the text classifier to match the harness's tag *shape*: any prompt
beginning with a lowercase multi-word kebab tag (<task-notification>,
<agent-message …>, <bash-input>, …) is machinery. Underscore tags (<user_query>)
are deliberately excluded — Grok wraps real typed prompts in them. Explicit
prefixes remain only for the single-word tag (<channel source=) and the prose
wrappers the harness uses.
Add a structural drop at the native-chat decoder (user turns marked
isMeta/isSynthetic/isCompactSummary, unless they carry tool results) and gate
AI Vault title seeding on the shared classifier, since task-notifications carry
no isMeta.
* fix(native-chat): avoid hiding user markup
---------
Co-authored-by: Brennan Benson <brennanbenson@Brennans-MacBook-Pro.local>