Commit Graph
11496 Commits
Author SHA1 Message Date
Jinwoo Hong a868b090e0 fix: connect mobile emulator in folder workspaces (#14009) 2026-08-12 13:02:23 -07:00
Brennan Benson 1136503c6a fix(mobile): use 'unsupported' for the no-review-creation test premise (#14091) 2026-08-12 12:16:03 -07:00
OrcaWin 09ec516ae5 fix(editor): index WSL watcher aliases per batch (#14015) 2026-08-12 03:24:31 -07:00
github-actions[bot] ad6f8011e2 Update README downloads badge 2026-08-12 09:33:52 +00:00
OrcaWin e5cd5bf542 fix: decouple Floating Workspace shell selection (#13995) 2026-08-12 02:09:57 -07:00
Jinwoo Hong 70ad65fc94 Fix paired remote HTML browser ownership and focus (#13876)
* Fix paired remote HTML browser preview

* Register paired HTML preview reliability gate

* Scope browser worktree selection guard

* Harden paired HTML preview ordering

* Simplify paired HTML preview ownership

* Record final paired HTML preview evidence

* fix(remote): simplify browser preview ownership

* test(remote): refresh HTML preview evidence

* fix(remote): reconcile browser create ordering

* fix(remote): guard browser focus through reconcile

* fix(remote): bound delayed browser focus

* fix(remote): guard browser focus context

* test(remote): record final HTML preview evidence

* fix(remote): keep HTML previews in background

* test(remote): cover browser focus modes

* test(remote): record focused browser contracts
2026-08-12 02:02:10 -07:00
Jinwoo Hong 346e59c879 fix(renderer): keep commit tooltip lines intact (#14000) 2026-08-12 01:10:47 -07:00
Brennan Benson 36d45af062 feat(browser): guide Google sign-in after cookie imports (STA-3811) (#13666) 2026-08-12 00:53:02 -07:00
Neil 9c091cf77e chore(perf): add renderer agent-status benchmark harness (#13905)
Splits run-idle-cpu-benchmark.mjs into a scale fixture, an in-page timing
probe, and process sampling, and records a measured origin/main baseline so
the agent-status batching slice has an auditable before.

The agent-status write workload is not included: it needs setAgentStatuses,
so it lands with the store slice.
2026-08-12 00:37:48 -07:00
Brennan Benson 2249330acf fix(browser): bulk clear cookies during native import (#13966)
* fix(browser): bulk clear cookies during native import

* fix(browser): separate Google cookie import warning

* fix(browser): report encrypted Google cookie exclusions

* fix(browser): skip key lookup for excluded cookies

* test(browser): cover excluded-cookie key bypass

* perf(browser): keep plaintext key scan cheap
2026-08-12 00:33:50 -07:00
NeilandOrca a90a18d43d perf(renderer): anchor spinner phase from animationstart only (#13987)
Co-authored-by: Orca <help@stably.ai>
2026-08-12 00:30:14 -07:00
Jinjing a81224614a fix(agents): detect a live OpenCode pane from its native OC | session title (#13957)
* fix(agents): detect a live OpenCode pane from its native OC | session title

OpenCode publishes `OC | <session>` as its OSC title, which carries no
agent-name token. detectAgentStatusFromTitle gates status on a whole-token
name match, so it returned null and every status consumer read a live
OpenCode pane as a plain shell: no "Send notes to" entry, no title-derived
sidebar row, and a title that the runtime's agent-presence check scored as
neutral. Identity already resolved (getAgentLabel returns OpenCode); only
activity was missing.

Treat the native marker as a live idle agent, placed after the spinner and
glyph checks so the decorated frames pinned by #8940 keep their status, and
accept it in the send-readiness gate the way Claude's U+2733 prefix is
accepted -- only a running OpenCode TUI ever publishes it.

* fix(agents): require spaced `OC | ` marker for native OpenCode detection

Unspaced pipes like `OC|Build` match other tools and would mistakenly
route non-OpenCode panes as send targets. Enforce literal ` | ` as
OpenCode emits it. Also clarify that only spinner decorations carry
working status, not keywords in the session summary.

* fix(agents): require OpenCode foreground process to validate native mark

OpenCode's native `OC | ` title marker now requires an active OpenCode process
to authorize agent sends, preventing false detection when the marker is left on
shell prompts. Extends wrapper prefix matching (ssh, tmux, etc.) and spinner
glyph support. Adds permission-prompt blocking signals for guarded writes.
2026-08-12 00:07:59 -07:00
Jinwoo Hong fd2afc16c8 Fix paired web Add Project folder browsing (#13885) 2026-08-11 23:25:50 -07:00
Jinwoo Hong c45c806220 fix(orchestration): close federated-read benchmark socket (#13939) 2026-08-11 23:19:07 -07:00
Jinwoo Hong 7319d59a10 Make worker completion and cleanup authoritative (#13927)
* fix: require authoritative worker completion verdicts

* Harden federated settlement replay

* fix(orchestration): reconcile dead retained workers

* docs: record SSH worker release coverage

* test(e2e): exercise worker settlement and release CLI

* docs: register combined orchestration CLI oracle

* test(orchestration): pin pre-ack attachment state
2026-08-11 23:06:12 -07:00
Brennan BensonandOrcaWin 4b9fb46646 fix(agent-hooks): stop rewriting typed Prime commands (#13906)
* fix(agent-hooks): preserve typed Prime commands

* test(wsl): enforce Prime bridge exclusion

---------

Co-authored-by: OrcaWin <alpha-eng@stably.ai>
2026-08-11 22:34:54 -07:00
Neil 9e18e1456b fix(workspace-emoji): keep the ':' shortcode suggestions scrolled to the highlight (#13969) 2026-08-11 22:18:26 -07:00
Neil 3d4c968e9b perf(renderer): drop redundant worktree card cache-TTL subscription (#13903)
promptCacheTtlMs was a third store subscription for a field the card
already reads via foundation.settings; derive it instead.

Does NOT bundle the 26 foundation selectors as originally planned.
Measured on this store with React (150 cards, 400 unrelated writes,
26 keys): 26 separate useAppStore hooks 100.7ms; one useShallow bundle
402.5ms (4.0x worse); one bundle with a non-allocating keyed compare
141.5ms (1.4x worse). zustand v5 shallow() rebuilds two Maps from
Object.entries per compare, and a bundle must still build the whole
projection every notification, so collapsing cheap selectors only
trades away useSyncExternalStore dispatches that are cheaper than the
object it allocates. Listener count would have improved 26 -> 1 while
the hot path got slower.
2026-08-11 22:13:39 -07:00
NeilandOrca 971d9548b5 docs(agent-status): make design references self-contained (#13902)
docs/design/agent-status-over-ssh.md was cited from ~10 source files but
does not exist in the repo. Replace each pointer with the invariant the
code actually relies on so the knowledge survives without the doc.

Renderer-side citations (useIpcEvents.ts, agent-status-types.ts) are left
for the concurrent batching change that owns those files.

Co-authored-by: Orca <help@stably.ai>
2026-08-11 22:13:35 -07:00
NeilandOrca 0079623f26 Fix reveal active workspace needing multiple clicks (#13926)
Co-authored-by: Orca <help@stably.ai>
2026-08-11 22:10:44 -07:00
Jinjing 5bee7b5ce9 P1 STA 3887 design Preview Kitty IME (#13940)
* fix(terminal): carry kitty flags through Preview snapshots and pair rele

Preview was omitting the live kitty mirror from the IME bridge and dropping kitty flags from snapshots, so every commit was evaluated at flags 0. A TUI that negotiated bit-3 (report_all_keys_as_escape_codes) would receive the legacy raw text it declined.

Now the snapshot carries proven kitty flags beside their sequence boundary, the forwarder reads flags once per commit, and bit-1 (report_event_types) commits are paired with exactly one release regardless of keyup/insertText ordering. Snapshot authorities expose only the active screen's proven flags, so an old host's absent field stays unknown rather than downgraded to a manufactured zero.

* fix(terminal): sync kitty flags and IME releases across snapshots

* trim wordinesss

* fix(terminal): settle owed IME release before fresh same-key press

When a keyup is lost and the same key is pressed again, settle the stale
record's owed release instead of discarding it — this maintains correct
IME state during recovery. Also refine Kitty flag propagation to only
carry proven baselines across snapshots, and tighten related comments.

* fix(terminal): gate kitty flags on sequence boundaries

- Remote snapshots only include flags when seq is present
- Daemon uses parsed flags value when defined
- Ensures correct flag ordering in snapshot replay
2026-08-11 22:00:44 -07:00
NeilandOrca 5ea7df1a5b fix(terminal): make DECSET 2031 subscriptions silent (#13904)
fish arms `CSI ?2031h` before painting each prompt and withdraws it when it
hands the tty to a child — a ~1ms window. Orca answered that subscribe with
`CSI ?997;Nn` across a 1-3ms renderer hop, so the reply landed after the
withdrawal and was read as stdin by the next child, corrupting `brew`/`npx`
`[y/N]` prompts.

The reply is not stale by Orca's own view when written (measured
staleReplies: 0), so no suppress-the-stale-reply scheme can close this — the
information needed to suppress does not exist yet. Nothing asked for the reply
either. The Contour spec says a terminal "should only send out the DSR when the
palette has been updated"; Ghostty (Termio.zig:729 — force=true reachable only
from the ?996n DSR), iTerm2 (VT100Terminal.m:995 — flag only) and xterm.js
(InputHandler.ts:2035 — flag only) all emit nothing on the DECSET. So stop
entering the race: record the subscription, answer nothing.

Of 17 real programs measured under a pty, only fish, tmux, claude and opencode
subscribe; none block on a reply, and answering produces one redundant palette
re-query and zero rendering difference. tmux is the only one that sends `?996n`,
which Orca still answers.

- Subscribes are record-only at all four emitters (live scan, hidden-gate fact,
  parked byte watcher, parked responder — the last is deleted, it only replied).
- `?996n` answers, the subscription registry, and the theme-flip push are
  unchanged. `paneLastThemeMode` is still seeded at subscribe so the next
  appearance re-apply is not read as a flip.
- Replay grammar carries `?2031l` alongside `?2031h`, so a late-attaching remote
  client no longer registers a subscription the TUI already retired.

Also closes fish-integration gaps found alongside: `unset` (which fish lacks)
becomes `set -e` on paths parsed by the client's login shell, `config.fish` is
parsed for agent-home detection, and bracketed-paste startup delivery is made
consistent across local/daemon/relay.

Regression test drives real fish 4.7.1 under node-pty and asserts on what the
child process reads; it fails against pre-fix code with the exact payload from
the issue. CI installs fish 4 and fails loudly rather than skipping.

Closes #9993

Co-authored-by: Orca <help@stably.ai>
2026-08-11 21:16:36 -07:00
Brennan Benson 137e724119 fix(agent-title): treat Claude Code quarter-circle spinners as working (#13889) (#13925)
* fix(agent-title): treat Claude Code quarter-circle spinners as working

Claude Code 2.1.228 swapped its busy OSC title spinner from braille
(U+2800-U+28FF) to quarter circles (U+25D0/U+25D1). Orca recognized a busy
Claude title only by braille codepoints, so the new frames matched nothing.

The summary-bearing busy frame ("<glyph> Say hi in one word") carries no
"claude" name token, so it resolved to no-status. The tracker's "idle or
permission followed by no-status means the agent exited" rule then fired
mid-turn, confirmPtyAgentExit confirmed it, and the chat surface routed
exitChat -- kicking the tab to the terminal view on every message.

Widen the accepted glyph set via a shared containsAgentSpinnerGlyph helper.
Agent-specific braille frame shapes (Grok, Pi, synthetic Cursor) stay pinned
to their own glyph set.

Fixes #13889

* fix(agent-title): satisfy static analysis and trim scope
2026-08-11 21:13:27 -07:00
948d85cabc fix(terminal): rank cold-park recency by activation order, not random UUID (#13881)
* fix(terminal): rank cold-park recency by activation order, not random UUID

The keep-warm exemption (#8262) broke equal hiddenSinceMs with
id.localeCompare over UUIDv4 tab ids, so which tab stays warm was a coin
flip. Ties are routine, not exotic: use-terminal-tab-cold-parking takes one
Date.now() per effect pass (#7214) and stamps every tab first seen hidden in
that pass with it, so switching away from a worktree ties all of them. The
tab that wins the flip is skipped by selectIdsBeyondHotRetain and never parks.

Ranks by an explicit activation sequence recorded on the hidden->visible edge.
Unit test covers the tie; red before, green after.

Does NOT fix terminal-hidden-view-parking.spec.ts:467, which still fails with
the same 'did not park' error on CI. Shipping this on its own merit.

Co-authored-by: Orca <help@stably.ai>

* fix(terminal): preserve focused tab across parking ties

---------

Co-authored-by: Orca <help@stably.ai>
Co-authored-by: E2E Test <e2e@test.local>
2026-08-11 21:07:57 -07:00
Neil e5971365e4 Update pull_request_template.md 2026-08-11 20:22:17 -07:00
Jinwoo HongandE2E Test 7a58f72814 Fix remote-runtime terminal pane split authority (#13867)
* fix(runtime): preserve remote terminal pane splits

* test(runtime): tighten split authority guards

* fix(runtime): revalidate transient split sources

---------

Co-authored-by: E2E Test <e2e@test.local>
2026-08-11 20:09:13 -07:00
Jinwoo Hong b51ef400ec fix(quick-commands): flatten the settings list and give the command editor room (#13922) 2026-08-11 19:53:22 -07:00
NeilandOrca 64aec94cb2 fix(sidebar): give draft reviews their own glyph instead of a red PR icon (#13919)
* fix(sidebar): give draft reviews their own glyph instead of a red PR icon

The worktree card review icon used one PR glyph for every state and
tinted it by check status, so a draft with failing checks rendered as a
red PR icon that read as closed. Shape now carries review state
(draft/closed/merged), and check tone applies only to open reviews so
the glyph agrees with its tooltip. Same fix for the source-control
header icon, where draft and closed differed only by a muted tone.

Closes #13088

Co-authored-by: Orca <help@stably.ai>

* test(sidebar): pin stateless and GitLab-closed review icon behaviour

Documents why a stateless row keeps its check tone (folder cards render
one while a linked review loads or its details fail) and covers the
closed-MR glyph override, which previously read as already-merged.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-11 19:26:53 -07:00
Neil 444638d96b Revise pull request template for clarity and updates
Updated the pull request template to simplify language and clarify sections. Added new sections for AI disclosure and testing instructions.
2026-08-11 18:40:27 -07:00
Neil ebcb2b6a60 docs: expand PR template with ELI5, before/after screenshots, and X handle
Make the pull request template clearer for contributors: plain-language
ELI5, what/why, mandatory before/after UI proof, testing checklist, AI
disclosure, and an Author X field. Align CONTRIBUTING with the template.
2026-08-11 18:29:01 -07:00
Neil 991a3fe963 chore(lint): update oxlint to 1.77 and enable no-op cleanup rules (#13901)
Enable eleven oxlint rules that simplify code without changing behavior, and fix
every existing violation. Each candidate was gated on measured cost rather than
assumption, so rules that regressed runtime performance or type checking were
dropped instead of suppressed.

typescript/no-redundant-type-constituents is the largest addition: 113 sites, no
autofix. Dead constituents are deleted. Where the redundant literal existed to
document intent (`string | 'all'`), it is preserved as `(string & {})`, which
keeps the autocomplete hint the original code was reaching for instead of
flattening it away. The rule also caught a broken import —
remote-shared-control-retirement-probe.ts pulled RuntimeStatus from
src/shared/types, which does not export it, so the type silently degraded to
`any`; no tsconfig covers that file, so tsc never saw it.

oxlint stays at 1.77.0 rather than 1.78.0 because .npmrc sets
minimum-release-age=4320 and 1.78.0 is younger than that window.

Rules evaluated and rejected, with what disqualified each:
- prefer-string-raw: String.raw is a runtime call, not a literal (184x slower)
- prefer-string-replace-all: 26% slower
- text-encoding-identifier-case: ~5% slower, reproducible
- prefer-spread: [...str] is 110% slower than split('') and differs on surrogates
- no-implicit-coercion: `!!x` narrows types and `Boolean(x)` does not (22 tsc errors)
- prefer-arrow-callback: arrows are not constructible, breaking `new` on mocks
- object-shorthand: rewrites source text asserted by a tracked reliability gate
- switch-case-braces: pushes ten files past max-lines, which cannot be suppressed
- no-useless-switch-case: drops `case undefined:` that switch-exhaustiveness-check needs
- arrow-body-style: 115 violations have no fix, and it breaks max-lines
- newline-after-import: false-positives on the leading-semicolon ASI idiom

electron-vite-output-contract asserted on the literal
Object.prototype.hasOwnProperty.call text; retarget it to Object.hasOwn, which
rejects inherited keys identically.
2026-08-11 18:19:43 -07:00
JinjingandOrca 4c2a10d157 Fix smart sort ranking of done agents by completion time (#13899)
* Fix smart sort ranking of done agents by completion time

Completed entries stayed in the Done sort class indefinitely when
same-state writes refreshed updatedAt without moving stateStartedAt.
Introduce agentEntryCompletionAt() to use actual completion time for
both age display and sort eligibility, ensuring consistent aging
regardless of hook updates.

* Fix smart sort ranking of done agents by completion time

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-11 18:11:34 -07:00
Jinjing 96c17b44d5 Resolve threads as primary acknowledgement, not replies (#13894)
Collapsed threads visibly close on the host — this is the
acknowledgement. An extra reply only adds noise. Snapshot the
resolution target at launch to prevent silent drops if the panel
navigates before agent delivery. Only reply for comments without
resolution endpoints.
2026-08-11 17:38:15 -07:00
Jinjing 675400e318 fix(automations): give list table more breathing room (#13898)
Increase row height and the gap above the table so the Automations list feels less dense.
2026-08-11 17:33:42 -07:00
Neil 0824351fe0 fix(bitbucket): commit credentials atomically and stop blaming the token for network faults (#13887)
* fix(bitbucket): commit credentials atomically and stop blaming the token for network faults

STA-3941 (P0): a credential edit could destroy the last working pair.
`writeFileSync` replaces in place, so a failed or interrupted write truncated
the previous secret, and the secret and metadata files were published
independently — a crash between them left a new secret paired with the old
email, unusable on restart.

- Write credential files through a temp + fsync + rename, so a reader sees
  either the old bytes or the complete new ones, never a truncated file.
- Carry authMode/email/baseUrl inside the encrypted envelope and treat the
  plaintext metadata as display-only. A torn write can now only leave a stale
  displayed account, never an unusable credential. Envelopes written before
  this change still resolve through the metadata fallback.

STA-3944 (P2): timeouts, DNS failures, 5xx and unparseable bodies collapsed to
the same miss as a 401, so Orca told users their credentials were invalid when
the host was simply unreachable — sending them to regenerate a working token.
`/user` now reports rejected vs unreachable, connect explains which happened,
and an unreachable host no longer renders as "Auth failed".

Adds failure-injection coverage for a partial secret write, an interrupt
between publishing the secret and the metadata, legacy envelopes, and the
transport-vs-auth split.

* fix(bitbucket): loop short writes and keep eligibility out of the keychain

Review of the atomic writer found it ignored writeSync's return value. write(2)
may return a short count, so a partial buffer could be fsynced and renamed into
place — the same truncation the change exists to prevent, moved one step later.
Write now loops to completion, with a test that fails without it.

Also stops isBitbucketReviewCreationAuthenticated force-decrypting the stored
secret. Create-PR eligibility is evaluated proactively for the sidebar, so that
popped an OS keychain prompt just from opening a worktree. Presence is enough;
creation itself still fails closed on an unusable credential.
2026-08-11 17:29:22 -07:00
JinjingandOrca 0c13e51097 feat(pr-comments): sort grouped PR comment sections newest-first (#13893)
Grouped/triage sections now read newest-first so recent discussion surfaces
at the top, while Timeline keeps its oldest-first history order.

- rename sortPRCommentGroupsForTimeline to sortPRCommentGroupsByRecency and
  add an order parameter
- rank threads by their latest activity under newest-first so a fresh reply
  refreshes an old thread
- break timestamp ties by numeric comment id in the sort direction, so a
  GitHub review batch sharing one timestamp orders correctly
- sink groups with an unparseable createdAt to the bottom in both orders

Co-authored-by: Orca <help@stably.ai>
2026-08-11 17:26:48 -07:00
Jinjing 55beeebd25 Allow directly search in google (#13863)
* Allow direct search with configured search engines

Users can now search directly from the tab creation menu using their
configured search engine. Forced search mode (`?` prefix) skips file and tab
matching for guaranteed search. Refactored tab-entry operations into focused
modules for clarity: forced-search parsing, network-safe selection, keyboard
focus, copy strings, empty options, and props types. Search routes through
the same workspace browser tab opening mechanism used for URLs, with safe
title and query presentation that doesn't retain sensitive details.

* Allow direct search with configured search engines

- Permit search and URL navigation while file index loads; require explicit
  selection only when needed, not automatic opening
- Block malformed IPv6 addresses in bracket notation to prevent misclassification
- Support Kagi private-session links via searchUrlOptions
- Improve error handling with accessibility: show error messages in status region,
  disable input during submission, display loading state
- Fall back to local browser tab creation when remote creation fails instead of
  throwing; avoids remote availability blocking local search/navigation
- Add error translations for all supported locales (es, ja, ko, zh)

* Allow direct search with configured search engines

- Extract tab create entry lifecycle to key-driven component remounting, replacing conditional state reset with useEffect cleanup
- Consolidate network tab entry classification and request building into reusable helpers, eliminating duplicate logic
- Simplify owner resolution by inlining logic directly into openWorkspaceBrowserTab
- Replace custom surrogate-pair handling with native String.toWellFormed() for search queries
- Disable explicit URL classification to prioritize search-engine queries over raw URLs

* Allow direct search from quick-open tab bar entry

- Single-token queries keep file matches ranked above search (quick-open intent)
- Multi-word phrases promote search to top, since they cannot be file paths
- Arm network actions once file index fails or text is unambiguous search
- Cache prepared file index to avoid re-processing per keystroke
- Generate specific tab titles (e.g. 'example.com/docs') instead of generic 'Open URL'
- Surface opening workspace when launching browser tabs remotely

* Use readOnly instead of disabled for pending search input

Maintain keyboard focus during submission so arrow/Escape navigation continues to work. Use aria-busy to indicate loading state accessibly. Also fixes button hover styling when disabled and cleans up error message handling in the classifier.

* Treat bare searches as prompts; refine path and IP classification

- Bare search queries (e.g., "?") no longer display as error rows
- Path prefixes with existing matches are no longer blocked mid-keystroke
- Private IPv4 addresses now use http, public addresses use https
- Ambiguous inputs with non-numeric ports fall through to search instead of blocking
- Improve diagnostics by logging failure reasons in openFailure
2026-08-11 17:22:15 -07:00
Brennan Benson 23cbe6dfe2 fix(agent-hooks): guard copilot and kimi Windows hooks before they own stdin (#13379)
Same hang class #11568 closed for .cmd: copilot-hook.ps1 ran
[Console]::In.ReadToEnd() and kimi-hook.sh captured stdin before checking the
Orca env, so a user-wide hook fired outside an Orca pane blocked forever when
the caller abandoned the pipe — one stranded powershell.exe or bash.exe per
hook event.

Move the env guard (after the endpoint refresh, which can supply PORT/TOKEN)
above the read on the Windows-local variants. The payload is discarded on the
missing-env path anyway; the broken writer is the trade #11568 established.
POSIX variants keep capture-first — those callers close stdin, and exiting
mid-write there surfaces as EPIPE the agent can see (#8110). Kimi's remote
install now asks for the posix variant explicitly.

Verified on a real Windows host: pre-fix copilot with an abandoning caller
never exits; post-fix both exit 0 immediately, and the env-present copilot
path still consumes stdin and delivers the POST (paneKey + payload received
by a local listener).
2026-08-11 17:06:55 -07:00
Jinwoo HongandE2E Test 9f0bf39b04 fix(github): fail closed when stack metadata is unavailable (#13866)
* fix(github): fail closed on unavailable stack metadata

* fix(github): validate REST pull request response shape

* fix(github): allow omitted stack metadata

* test(github): cover enterprise stack probe failure

* test(github): cover null ordinary stack metadata

---------

Co-authored-by: E2E Test <e2e@test.local>
2026-08-11 16:42:56 -07:00
Brennan Benson 686f5dca1a fix(browser): exclude Google cookies from imports - direct sign-in is the only Google path (STA-3811) (#13670)
* fix(browser): exclude Google cookies from imports (STA-3811)

Imports never write and never remove a google.com-family cookie, on any
path. Signing in directly inside Orca is the only Google session that
survives, so the live jar always beats anything an import could plant.

* fix(browser): roll back selective cookie clears
2026-08-11 16:39:17 -07:00
Brennan Benson f19ff5be68 fix(browser): keep one identity across hosts during Google sign-in (STA-3811) (#13667)
While the auth document is on screen the WebContents UA is Firefox, so its
cross-host subresource/XHR requests (gstatic, play.google.com, the sign-in
challenge endpoints) reached the header layer carrying the Firefox UA yet still
bearing Chromium client hints, which the else-branch rewrote to Chrome. That
paired a Firefox UA with Chrome client hints on every non-auth Google host — a
sharper cross-host identity tell than either signal alone, and a plausible
cause of the password-submit challenge greying out and stalling.

Strip client hints on any request already carrying the Firefox auth UA so the
UA and hint surfaces tell one Firefox story for the whole flow. Gated on the
same googleAuthOverride flag as the auth-host switch, so imported-native
profiles are unaffected and the clean-Chrome default for non-Google sites
(Cloudflare) is untouched.

Extends tests/tools/google-signin-ua-probe.cjs with app-current/app-fixed
modes that mirror the shipped code and log per-request identity; on the real
accounts.google.com load they show 18 firefox-ua-with-chrome-hints cross-host
mismatches before and 0 after.
2026-08-11 16:37:53 -07:00
NeilandOrca ff727c3caf fix(settings): read voice settings from a ref, not a stale closure (#13883)
* fix(settings): read voice settings from a ref, not a stale closure

updateVoiceSettings wrote from a closure captured at mount, so an async
continuation could clobber a newer value. Uses the latest-ref pattern already
in Settings.tsx:405-407.

Found while investigating voice-microphone-selection.spec.ts:107 and confirmed
NOT to be its cause — that failure has no established root cause. The spec
gains a hydration barrier and toBeEnabled assertions here, which improve its
failure message but do not fix it.

Co-authored-by: Orca <help@stably.ai>

* test(e2e): drop the voice spec changes from this PR

The hydration poll was a no-op: prepareVoiceSettings awaits updateSettings
inside the same page.evaluate, so the store is already committed before the
poll samples it, in all three tests (the reload precedes it). Worse, its
comment asserted that the Select is disabled until hydration — the CI
accessibility snapshot at the moment of failure shows the switch checked and
the combobox enabled, so that theory is refuted, and a wrong 'why' comment is
worse than none.

The toBeEnabled assertions go with it: the control was enabled when it hung,
so they would not have caught the real failure either.

Leaves this PR as what it actually is — a product bug fix. The spec's failure
stays open and unexplained.

Co-authored-by: Orca <help@stably.ai>

* test(settings): type the clear-key mock as its API contract

vi.fn(() => clearing) returned Promise<void> where the preload contract is
Promise<{ configured: boolean }>. Vitest does not typecheck, so this passed
locally while tsc -p tsconfig.tc.web.json failed — it would have gone red in
CI's required typecheck job.

Co-authored-by: Orca <help@stably.ai>

* fix(settings): write the voice-settings ref in an effect, not during render

React Doctor's 'Ref mutated during render' rule is blocking in CI's static
analysis job, and it is right: React can replay or discard render work, so a
render-time ref write can leak from UI that never commits.

Moving it to an effect keeps the behavior this fix needs — every writer here
(key-status probe, save-key, clear-key) runs post-commit, so the ref is current
by the time any of them reads it. Test still discriminates: red without the
fix, green with.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-11 16:30:42 -07:00
Jinwoo Hong 09c8597fb7 perf(orchestration): route federated reads over shared control (#13814) 2026-08-11 16:17:17 -07:00
Brennan Benson e77e1fe850 fix(claude): guard cold-restore resume selectors (#13868)
* fix(claude): guard cold-restore resume selectors

Persisted Claude default args or a custom command can carry their own
--resume/-r/--continue/-c selectors (a bare picker default or a stale id).
Cold restore appended the authoritative --resume <id> after them, typing a
command with competing selectors into the restored pane (#12982).

buildAgentResumeStartupPlan now routes Claude through a selector guard that
tokenizes the base with the existing startup tokenizer, strips selectors in
option position only (value-taking options keep dash-leading values), and
appends exactly one authoritative selector, inserting before Claude's own
-- terminator when present. Splicing is span-based so untouched bytes stay
verbatim, wrapper commands are left alone, and any tokenization failure
falls back to the previous append-only behavior. Launch paths, other
agents, persistence, and the wire are unchanged.

* fix(claude): harden resume selector guard against false matches

Round-1 review findings: locate the claude executable by command position
(index 0, after a wrapper --, or behind NAME=value assignments) so an
argument merely ending in /claude can never be mistaken for it; stop
matching the joined -r<id> form, which was ambiguous with dash-leading
option values and forced an unmaintainable arity table (now deleted).
Ambiguous shapes degrade to the pre-guard append-only behavior.

* fix(claude): fail resume guard open on chained shell syntax

Round-2 review findings: an unquoted operator or newline after the claude
token means the base chains other commands, and splicing across that
boundary handed the selector to the wrong command — detect it and fall
back to plain appending. Also recognize claude behind PowerShell's & call
operator, decouple the test oracle from the implementation's selector
predicate, add Windows tokenizer span tests, and rename the module after
its public API.

* fix(claude): flag bare shell operators inside the tokenizers

Round-3 review findings: the guard's operator scan compared raw source to
token value, so one quote or escape anywhere in a token hid a shell-active
operator outside the quotes and the splice crossed a live command boundary,
losing the resume entirely. Both tokenizers now flag tokens carrying an
unquoted, unescaped operator byte (or a word-leading # comment on
posix/powershell) on their spans, where quote state actually lives, and the
guard fails open on that flag. Also strengthens the redirect fail-open test
to carry a stale selector, re-tokenizes each raw span in the shell span
tests, and documents agent-resume-argv-drop as codex-only.

* fix(claude): flag expansions and clamp separator backoff

Round-4 review findings: unquoted multi-token expansions (backtick, $(, ${)
split across whitespace, so removing only the recognized selector token left
a broken construct tail — both tokenizers now raise the span flag (renamed
bareShellSyntax) for those openers, on cmd also for operators between
single quotes, which cmd does not treat as quoting. The separator backoff is
clamped to the previous token's span end so a token ending in an escaped
space can no longer donate its escape to the appended selector.

* fix(claude): treat cmd single-quoted regions as unmodelable

Round-5 review finding: cmd.exe has no single-quote syntax, so the Windows
tokenizer's grouping of a single-quoted region diverges from what cmd
parses — literal argv like 'claude ...--resume... old' was being read as a
real selector and stripped, and a literal '--' as claude's terminator. Flag
any cmd single-quoted token as bareShellSyntax so the guard fails open.

* fix(claude): flag quoted expansions and scope assignment prefixes

Round-6 review findings: the span flag was only evaluated in the unquoted
branch, so an expansion opener inside double quotes went unflagged — and
inside $(…)/backticks a nested quote re-opens a context this tokenizer
does not model, so the splice could cut mid-construct (syntax error, or a
silently mutated substitution body). Both tokenizers now flag those, and
the flag is renamed divergesFromShell to say what it means. Restrict the
NAME=value command-position prefix to posix, where that syntax exists.
Drops two branches proven dead.

* fix(claude): model shell-literal escapes and scan the whole base

Round-7 review findings: (1) the divergence scan started after the claude
token, so an expansion opened in a prefix — $(x; npx -- claude --resume s) —
had its closer spliced away, producing a base bash cannot parse; it now
covers every token including the executable, exempting only PowerShell's
leading call operator. (2) posix drops a double-quoted backslash the shell
keeps literal, and the Windows escape branch ran inside quoted regions where
cmd/PowerShell keep the escape byte literal — both now flagged, so a literal
can never be misread as a selector. (3) an unquoted line continuation hid a
selector inside a token and skipped the newline gap check.

Also removes a third provably dead branch and collapses the cut floor into
the cut itself.

* fix(claude): flag escapes the tokenizer models but the shell removes

Round-8 review findings, all one family — escapes whose token value hides
a selector the shell would see: a double-quoted line continuation (bash
deletes both bytes), posix $'…'/$"…" quoting, a windows escaped newline,
and a trailing unpaired escape. The last one was previously written off as
pre-fix-identical, but once stripping happens the dangling escape swallows
the separator and no exact --resume reaches claude at all — strictly worse
than appending, so it must fail open. Also folds the three gap predicates
into one scan.

* fix(claude): stop over-flagging a literal dollar sign

Round-9 review findings from both lanes: inside double quotes only $( and
${ open an expansion — $' and $" are literal there — and a trailing $
was flagged unconditionally because JS ''.includes('') is true. Both made
the guard fail open on modelable bases, leaving the stale selector to
compete, so #12982 went unfixed for them. Separately, cmd strips ^ before
the child re-splits on the bare whitespace, so an escaped separator hides
two real arguments and must fail open rather than drop one.

* fix(claude): fail open on cmd caret-quotes and bare PowerShell syntax

Round-10 review findings, both Windows-only (a bash oracle cannot reach
them): cmd strips a caret before a quote and the child's parser then reads
a bare quote delimiter, so the tokenizer's word boundaries stop matching
argv — one case turned a working resume into no resume at all, another let
a stale selector survive the splice. And bare (…)/{…} are live PowerShell
syntax in argument position, so splicing through them emitted unbalanced
output that PowerShell cannot parse.

* fix(claude): fail open on the PowerShell stop-parsing token

Round-11 review finding: after a bare --%, PowerShell passes the rest of
the line to the child literally, so the guard stripped a real selector and
then appended quoting that arrives as literal bytes — claude ends up with
no exact --resume at all, worse than leaving the stale one. Quoted "--%"
and cmd, where the token is ordinary, still splice.

* fix(claude): model cmd backslash-escaped quotes

Round-11 review finding: an odd run of backslashes before a quote makes it
a literal byte to the child's CommandLineToArgvW parser, not a delimiter,
so the tokenizer's word boundaries stopped matching argv. Orca manufactures
that pattern itself — quoteStartupArg wraps every token in quotes without
escaping a trailing backslash — so a pasted Windows path was enough to move
the selector into a desynced region and leave claude with no resume flag.
Also replaces a caret test case that was byte-identical before and after
its own fix, and merges two stacked comment blocks.

* fix(claude): fail open on PowerShell double-quoted escape sequences

Round-12 finding: PowerShell expands backtick escapes only inside double
quotes, so a sequence there produces a token value argv never sees — the
guard could strip "-`r" plus the argument after it. Also narrows the
stop-parsing comment: a quoted --% can engage stop-parsing before a
parameter token, where the base is already mangled either way.

* fix(claude): flag PowerShell escape sequences in bare arguments too

Round-13 finding: the previous commit gated on quote === '"', but
PowerShell's tokenizer calls Backtick() from ScanGenericToken, so it
expands these sequences in unquoted arguments as well — bare -`r really
is a control character, not -r. The guard read it as a selector and
dropped it plus the argument after it. Widening to all PowerShell
contexts measures 0 under-flag and 0 over-flag across the full printable
matrix; the backtick-escaped-space idiom still splices. Also swaps a test
case that was byte-identical with and without its own fix.

* fix(claude): drop a token-leading PowerShell backtick before whitespace

Round-14 observations, all pre-existing and measured: PowerShell drops a
token-leading backtick together with the whitespace after it, emitting no
token, so the tokenizer's extra token shifted the locator; and a backtick
before a bare CR is a line continuation too. Flagging both takes the
lane's 329k-base sweep from 87 bad to 0 with no new failures and the
must-splice list byte-unchanged. Also corrects a comment that no longer
listed every PowerShell divergence.

* docs(claude): correct the bare-CR rationale in the tokenizer comment

Round-15 verified against a real PowerShell 7.6.4 engine: a backtick
before a bare CR is not a line continuation there — pwsh keeps the CR in
the token. The flag stays because 5.1 is unverified and failing open costs
nothing, but the comment now says that rather than claiming continuation.
2026-08-11 16:16:24 -07:00
OrcaWinandBrennan Benson 234caecb82 fix(windows): refresh PATH ordering for newly installed tools (#13545)
* fix(windows): refresh persisted PATH ordering

New terminals now preserve Orca-injected entries while adopting current machine and user PATH precedence, so newly installed tools are not shadowed by stale aliases. Invalidate the cached registry snapshot on Windows setting changes and guard in-flight refreshes from restoring stale cache data.

* fix(windows): retry invalidated PATH reads

* fix(windows): reuse current PATH refresh

* fix(windows): preserve per-hive PATH fallback state

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-11 16:10:15 -07:00
Jinjing 648d17372b feat(sidebar): add jump-to-top button for hard upward scrolling (#13864)
* feat(sidebar): add jump-to-top button for hard upward scrolling

Detect intentional hard scroll-up gestures (wheel or scrollbar drag) and
offer a one-click jump-to-top affordance. Auto-hide after idle to avoid
persistent visual clutter. Addresses the common case of fast navigation
through long worktree lists ranked by agent activity.

* test(sidebar): improve scroll-to-top detection for active gestures

Only detect velocity from active scrollbar drag or touch, not
programmatic scrolls. Return focus to list after jump-to-top. Add
comprehensive hook tests with gesture simulation. Improve cumulative
down-delta tracking for dismissal.

* test(sidebar): add scroll-to-top gesture detection tests

Add comprehensive test coverage for the hard-upward-scroll detection hook,
verifying idle timer behavior, gesture suppression, scrollability checks,
and cleanup on unmount. Extract the post-jump suppression window into a
named constant for maintainability.
2026-08-11 15:56:07 -07:00
Jinwoo Hong 077f5a11cd feat(github): create stacked pull requests (#13750)
Adds GitHub stacked pull request creation: a contextual "Stack this PR above #N" option that appears only when the selected base branch has an open PR, plus the main-process stack preflight and registration.

Also reworks the create-review composer for cohesion: shadcn Checkbox and Label primitives, base label above a full-width searchable combobox with attached results, keyboard navigation, and a unified field skin, spacing and typography scale.

Verified end to end against real GitHub: extending an existing stack and creating a new one.
2026-08-11 15:48:57 -07:00
Neilanddevatnull 63271a5933 feat(bitbucket): connect Bitbucket from Settings and create pull requests (#5832)
* feat(bitbucket): connect Bitbucket from Settings with encrypted credential storage

Bitbucket Cloud was the only review provider with no in-app auth: GitHub and
GitLab delegate to the gh/glab CLIs, but Bitbucket has no comparable
first-party CLI, so the only option was ORCA_BITBUCKET_* env vars plus a
restart (discussion #5364).

Adds a Connect/Edit/Disconnect flow on the Bitbucket integration card,
modeled on Linear and Jira:

- Credentials are verified against /user before they are persisted, so a
  dead token is rejected inline instead of silently stored.
- The secret is encrypted with safeStorage (0600 plaintext fallback when no
  OS keyring); non-secret metadata lives in a separate plaintext file so
  status reads render the connected account without decrypting. Opening
  Settings therefore never triggers a keychain prompt.
- Env vars keep precedence over stored credentials, so existing headless and
  SSH setups are unaffected. Env-managed connections hide Disconnect.
- connect/disconnect reset the preflight cache, so no relaunch is needed.

The Bitbucket card moves to its own file to stay under the tsx max-lines cap.

* feat(bitbucket): support creating pull requests from Orca

Bitbucket was the only configured provider whose Create button reported
"This repository provider does not support creating a pull request from
Orca" — supportsReviewCreation was false and the forge provider had no
createReview, so even a correctly authenticated setup was blocked.

Adds createBitbucketPullRequest against POST /repositories/{ws}/{repo}/
pullrequests, using the same env-first / stored-credential resolution as PR
lookups (extracted into resolve-auth.ts so both share one path).

Bitbucket Cloud has no draft pull requests, so a draft request is rejected
with a clear message rather than silently publishing a live PR.

* fix(bitbucket): hide the draft toggle where drafts do not exist, plus review fixes

Bitbucket Cloud has no draft pull requests, so the composer no longer offers
the toggle for it and forces the flag off at submit — better than failing
after the user has filled the form in.

Review fixes:
- writeFileSync's `mode` only applies when it creates the file, so rewriting
  a credential kept whatever permissions it already had. chmod after every
  write, for the secret and the metadata.
- An explicit ORCA_BITBUCKET_API_BASE_URL now wins over a stored base URL.
  Env precedence is per-setting, not all-or-nothing.
- Enter in the credentials dialog only submits from a text field, so it no
  longer hijacks Cancel and the docs link.
- Replace the chmod-based delete-failure test with a mocked unlinkSync: file
  modes are not portable to Windows and elevated runners unlink anyway.

* fix(bitbucket): stop a merged pull request from blocking the branch's next one

Reported on #5832: with a merged PR on a branch, Create reported "Pull
request already exists" and offered no way forward.

The branch lookup queries every PR state and returns the most recently
updated one, so a merged PR came back as the branch's current review and
eligibility blocked on it. Bitbucket only discarded such a match on the repo
default branch (#9171), while GitHub already drops any merged PR it matched
by branch alone — "a merged PR without an explicit link is just a historical
branch match, not implicit review context".

Applies that rule to Bitbucket. An explicitly linked review still resolves
through the linked-number fallback, so merging a PR Orca knows about keeps
showing it.

* fix(bitbucket): add bitbucket to the shared review-creation provider list

Reported on #5832: on a Bitbucket repo with no existing PR, Create still
said "This repository provider does not support creating a pull request
from Orca", even after the forge provider gained createReview.

There are two capability lists. Enabling supportsReviewCreation on the forge
provider was necessary but not sufficient — the blocker and the whole
renderer read the separate shared list, which never included bitbucket.

Adds it, gives Bitbucket its own provider name so review copy stops saying
"GitHub", and asserts the two lists agree so they cannot drift apart again.

* fix(bitbucket): persist pull request links after creation

* fix(bitbucket): fetch linked pull requests by number first

* fix(i18n): use generated Bitbucket integration keys

* test(bitbucket): cover forge creation delegation

* fix(bitbucket): fall back when linked pull request is stale

* docs(bitbucket): explain notFoundIsNull and fix a garbled permissions comment

notFoundIsNull arrived without the rationale its sibling flag carries, and
reads as a bare `true` at the only call site that opts in.

* fix(bitbucket): address review findings before merge

Two of these made the feature unusable in real setups:

- Create PR checked GitHub authentication for Bitbucket. isProviderAuthenticated
  fell through to isGitHubAuthenticated, which was unreachable while Bitbucket
  could not create reviews at all. Anyone with Bitbucket connected but no
  `gh auth login` got auth_required with no way forward.
- The draft flag was only gated in ChecksPanel, not the two SourceControl call
  sites. With "create as draft" saved as a default, the composer hides the
  toggle for Bitbucket, so the flag could not be cleared and creation failed
  every time. Bitbucket now ignores draft instead of rejecting it.

Also:
- Blocked-create copy said "GitHub is not authenticated. Run gh auth login" on
  Bitbucket repos, in both the main-process and renderer paths.
- A decryption failure resolved to an anonymous config and queried anyway; a
  private repo answers 404, which reads as "no pull request" and offers Create
  for a branch that already has one. Requests now fail closed.
- Hiding non-open implicit branch matches was too broad: a declined PR became
  permanently invisible off the default branch. Scoped to merged, restoring the
  default-branch rule (#9171) for the rest.
- A failed disconnect rejected unhandled and the card silently re-rendered as
  connected; a partial delete left the secret live in memory for the session.
- The credentials dialog refused to open on a remote runtime, so a local repo
  could never store a credential. Now only the storage note changes, matching
  the Jira dialog.

---------

Co-authored-by: devatnull <59279509+devatnull@users.noreply.github.com>
2026-08-11 15:32:13 -07:00
OrcaWin e6eec11b9f fix(wsl): preserve single-letter POSIX terminal cwd (#13859) 2026-08-11 15:29:33 -07:00
Brennan Benson 46756b386a fix(terminal): retain WebGL for recently hidden worktrees (#13372)
* fix(terminal): retain WebGL for recently hidden worktrees

Suspending a hidden worktree disposes every pane's WebGL addon, so every
switch-back presents DOM-renderer frames (unfloored ~5% wider advance, and
the channel through which any poisoned cell metrics reach the screen) until
reattach completes — 0.5-3s on loaded sessions. Keep the live addons for the
most recently hidden worktrees instead: LRU capped at 6 contexts (Chromium
allows ~16/page, visible worktrees use 1-4), least-recent evicted to dispose
exactly as before. Reveal then has no renderer swap and nothing to flash.
Window-occlusion callers pass no retention context and are unchanged.

* fix(terminal): pause retained hidden cursor work

* fix(terminal): preserve retained WebGL through deferred rebuilds
2026-08-11 14:43:35 -07:00