Splits run-idle-cpu-benchmark.mjs into a scale fixture, an in-page timing
probe, and process sampling, and records a measured origin/main baseline so
the agent-status batching slice has an auditable before.
The agent-status write workload is not included: it needs setAgentStatuses,
so it lands with the store slice.
* fix(agents): detect a live OpenCode pane from its native OC | session title
OpenCode publishes `OC | <session>` as its OSC title, which carries no
agent-name token. detectAgentStatusFromTitle gates status on a whole-token
name match, so it returned null and every status consumer read a live
OpenCode pane as a plain shell: no "Send notes to" entry, no title-derived
sidebar row, and a title that the runtime's agent-presence check scored as
neutral. Identity already resolved (getAgentLabel returns OpenCode); only
activity was missing.
Treat the native marker as a live idle agent, placed after the spinner and
glyph checks so the decorated frames pinned by #8940 keep their status, and
accept it in the send-readiness gate the way Claude's U+2733 prefix is
accepted -- only a running OpenCode TUI ever publishes it.
* fix(agents): require spaced `OC | ` marker for native OpenCode detection
Unspaced pipes like `OC|Build` match other tools and would mistakenly
route non-OpenCode panes as send targets. Enforce literal ` | ` as
OpenCode emits it. Also clarify that only spinner decorations carry
working status, not keywords in the session summary.
* fix(agents): require OpenCode foreground process to validate native mark
OpenCode's native `OC | ` title marker now requires an active OpenCode process
to authorize agent sends, preventing false detection when the marker is left on
shell prompts. Extends wrapper prefix matching (ssh, tmux, etc.) and spinner
glyph support. Adds permission-prompt blocking signals for guarded writes.
promptCacheTtlMs was a third store subscription for a field the card
already reads via foundation.settings; derive it instead.
Does NOT bundle the 26 foundation selectors as originally planned.
Measured on this store with React (150 cards, 400 unrelated writes,
26 keys): 26 separate useAppStore hooks 100.7ms; one useShallow bundle
402.5ms (4.0x worse); one bundle with a non-allocating keyed compare
141.5ms (1.4x worse). zustand v5 shallow() rebuilds two Maps from
Object.entries per compare, and a bundle must still build the whole
projection every notification, so collapsing cheap selectors only
trades away useSyncExternalStore dispatches that are cheaper than the
object it allocates. Listener count would have improved 26 -> 1 while
the hot path got slower.
docs/design/agent-status-over-ssh.md was cited from ~10 source files but
does not exist in the repo. Replace each pointer with the invariant the
code actually relies on so the knowledge survives without the doc.
Renderer-side citations (useIpcEvents.ts, agent-status-types.ts) are left
for the concurrent batching change that owns those files.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): carry kitty flags through Preview snapshots and pair rele
Preview was omitting the live kitty mirror from the IME bridge and dropping kitty flags from snapshots, so every commit was evaluated at flags 0. A TUI that negotiated bit-3 (report_all_keys_as_escape_codes) would receive the legacy raw text it declined.
Now the snapshot carries proven kitty flags beside their sequence boundary, the forwarder reads flags once per commit, and bit-1 (report_event_types) commits are paired with exactly one release regardless of keyup/insertText ordering. Snapshot authorities expose only the active screen's proven flags, so an old host's absent field stays unknown rather than downgraded to a manufactured zero.
* fix(terminal): sync kitty flags and IME releases across snapshots
* trim wordinesss
* fix(terminal): settle owed IME release before fresh same-key press
When a keyup is lost and the same key is pressed again, settle the stale
record's owed release instead of discarding it — this maintains correct
IME state during recovery. Also refine Kitty flag propagation to only
carry proven baselines across snapshots, and tighten related comments.
* fix(terminal): gate kitty flags on sequence boundaries
- Remote snapshots only include flags when seq is present
- Daemon uses parsed flags value when defined
- Ensures correct flag ordering in snapshot replay
fish arms `CSI ?2031h` before painting each prompt and withdraws it when it
hands the tty to a child — a ~1ms window. Orca answered that subscribe with
`CSI ?997;Nn` across a 1-3ms renderer hop, so the reply landed after the
withdrawal and was read as stdin by the next child, corrupting `brew`/`npx`
`[y/N]` prompts.
The reply is not stale by Orca's own view when written (measured
staleReplies: 0), so no suppress-the-stale-reply scheme can close this — the
information needed to suppress does not exist yet. Nothing asked for the reply
either. The Contour spec says a terminal "should only send out the DSR when the
palette has been updated"; Ghostty (Termio.zig:729 — force=true reachable only
from the ?996n DSR), iTerm2 (VT100Terminal.m:995 — flag only) and xterm.js
(InputHandler.ts:2035 — flag only) all emit nothing on the DECSET. So stop
entering the race: record the subscription, answer nothing.
Of 17 real programs measured under a pty, only fish, tmux, claude and opencode
subscribe; none block on a reply, and answering produces one redundant palette
re-query and zero rendering difference. tmux is the only one that sends `?996n`,
which Orca still answers.
- Subscribes are record-only at all four emitters (live scan, hidden-gate fact,
parked byte watcher, parked responder — the last is deleted, it only replied).
- `?996n` answers, the subscription registry, and the theme-flip push are
unchanged. `paneLastThemeMode` is still seeded at subscribe so the next
appearance re-apply is not read as a flip.
- Replay grammar carries `?2031l` alongside `?2031h`, so a late-attaching remote
client no longer registers a subscription the TUI already retired.
Also closes fish-integration gaps found alongside: `unset` (which fish lacks)
becomes `set -e` on paths parsed by the client's login shell, `config.fish` is
parsed for agent-home detection, and bracketed-paste startup delivery is made
consistent across local/daemon/relay.
Regression test drives real fish 4.7.1 under node-pty and asserts on what the
child process reads; it fails against pre-fix code with the exact payload from
the issue. CI installs fish 4 and fails loudly rather than skipping.
Closes#9993
Co-authored-by: Orca <help@stably.ai>
* fix(agent-title): treat Claude Code quarter-circle spinners as working
Claude Code 2.1.228 swapped its busy OSC title spinner from braille
(U+2800-U+28FF) to quarter circles (U+25D0/U+25D1). Orca recognized a busy
Claude title only by braille codepoints, so the new frames matched nothing.
The summary-bearing busy frame ("<glyph> Say hi in one word") carries no
"claude" name token, so it resolved to no-status. The tracker's "idle or
permission followed by no-status means the agent exited" rule then fired
mid-turn, confirmPtyAgentExit confirmed it, and the chat surface routed
exitChat -- kicking the tab to the terminal view on every message.
Widen the accepted glyph set via a shared containsAgentSpinnerGlyph helper.
Agent-specific braille frame shapes (Grok, Pi, synthetic Cursor) stay pinned
to their own glyph set.
Fixes#13889
* fix(agent-title): satisfy static analysis and trim scope
* fix(terminal): rank cold-park recency by activation order, not random UUID
The keep-warm exemption (#8262) broke equal hiddenSinceMs with
id.localeCompare over UUIDv4 tab ids, so which tab stays warm was a coin
flip. Ties are routine, not exotic: use-terminal-tab-cold-parking takes one
Date.now() per effect pass (#7214) and stamps every tab first seen hidden in
that pass with it, so switching away from a worktree ties all of them. The
tab that wins the flip is skipped by selectIdsBeyondHotRetain and never parks.
Ranks by an explicit activation sequence recorded on the hidden->visible edge.
Unit test covers the tie; red before, green after.
Does NOT fix terminal-hidden-view-parking.spec.ts:467, which still fails with
the same 'did not park' error on CI. Shipping this on its own merit.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): preserve focused tab across parking ties
---------
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: E2E Test <e2e@test.local>
* fix(sidebar): give draft reviews their own glyph instead of a red PR icon
The worktree card review icon used one PR glyph for every state and
tinted it by check status, so a draft with failing checks rendered as a
red PR icon that read as closed. Shape now carries review state
(draft/closed/merged), and check tone applies only to open reviews so
the glyph agrees with its tooltip. Same fix for the source-control
header icon, where draft and closed differed only by a muted tone.
Closes#13088
Co-authored-by: Orca <help@stably.ai>
* test(sidebar): pin stateless and GitLab-closed review icon behaviour
Documents why a stateless row keeps its check tone (folder cards render
one while a linked review loads or its details fail) and covers the
closed-MR glyph override, which previously read as already-merged.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
Make the pull request template clearer for contributors: plain-language
ELI5, what/why, mandatory before/after UI proof, testing checklist, AI
disclosure, and an Author X field. Align CONTRIBUTING with the template.
Enable eleven oxlint rules that simplify code without changing behavior, and fix
every existing violation. Each candidate was gated on measured cost rather than
assumption, so rules that regressed runtime performance or type checking were
dropped instead of suppressed.
typescript/no-redundant-type-constituents is the largest addition: 113 sites, no
autofix. Dead constituents are deleted. Where the redundant literal existed to
document intent (`string | 'all'`), it is preserved as `(string & {})`, which
keeps the autocomplete hint the original code was reaching for instead of
flattening it away. The rule also caught a broken import —
remote-shared-control-retirement-probe.ts pulled RuntimeStatus from
src/shared/types, which does not export it, so the type silently degraded to
`any`; no tsconfig covers that file, so tsc never saw it.
oxlint stays at 1.77.0 rather than 1.78.0 because .npmrc sets
minimum-release-age=4320 and 1.78.0 is younger than that window.
Rules evaluated and rejected, with what disqualified each:
- prefer-string-raw: String.raw is a runtime call, not a literal (184x slower)
- prefer-string-replace-all: 26% slower
- text-encoding-identifier-case: ~5% slower, reproducible
- prefer-spread: [...str] is 110% slower than split('') and differs on surrogates
- no-implicit-coercion: `!!x` narrows types and `Boolean(x)` does not (22 tsc errors)
- prefer-arrow-callback: arrows are not constructible, breaking `new` on mocks
- object-shorthand: rewrites source text asserted by a tracked reliability gate
- switch-case-braces: pushes ten files past max-lines, which cannot be suppressed
- no-useless-switch-case: drops `case undefined:` that switch-exhaustiveness-check needs
- arrow-body-style: 115 violations have no fix, and it breaks max-lines
- newline-after-import: false-positives on the leading-semicolon ASI idiom
electron-vite-output-contract asserted on the literal
Object.prototype.hasOwnProperty.call text; retarget it to Object.hasOwn, which
rejects inherited keys identically.
* Fix smart sort ranking of done agents by completion time
Completed entries stayed in the Done sort class indefinitely when
same-state writes refreshed updatedAt without moving stateStartedAt.
Introduce agentEntryCompletionAt() to use actual completion time for
both age display and sort eligibility, ensuring consistent aging
regardless of hook updates.
* Fix smart sort ranking of done agents by completion time
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
Collapsed threads visibly close on the host — this is the
acknowledgement. An extra reply only adds noise. Snapshot the
resolution target at launch to prevent silent drops if the panel
navigates before agent delivery. Only reply for comments without
resolution endpoints.
* fix(bitbucket): commit credentials atomically and stop blaming the token for network faults
STA-3941 (P0): a credential edit could destroy the last working pair.
`writeFileSync` replaces in place, so a failed or interrupted write truncated
the previous secret, and the secret and metadata files were published
independently — a crash between them left a new secret paired with the old
email, unusable on restart.
- Write credential files through a temp + fsync + rename, so a reader sees
either the old bytes or the complete new ones, never a truncated file.
- Carry authMode/email/baseUrl inside the encrypted envelope and treat the
plaintext metadata as display-only. A torn write can now only leave a stale
displayed account, never an unusable credential. Envelopes written before
this change still resolve through the metadata fallback.
STA-3944 (P2): timeouts, DNS failures, 5xx and unparseable bodies collapsed to
the same miss as a 401, so Orca told users their credentials were invalid when
the host was simply unreachable — sending them to regenerate a working token.
`/user` now reports rejected vs unreachable, connect explains which happened,
and an unreachable host no longer renders as "Auth failed".
Adds failure-injection coverage for a partial secret write, an interrupt
between publishing the secret and the metadata, legacy envelopes, and the
transport-vs-auth split.
* fix(bitbucket): loop short writes and keep eligibility out of the keychain
Review of the atomic writer found it ignored writeSync's return value. write(2)
may return a short count, so a partial buffer could be fsynced and renamed into
place — the same truncation the change exists to prevent, moved one step later.
Write now loops to completion, with a test that fails without it.
Also stops isBitbucketReviewCreationAuthenticated force-decrypting the stored
secret. Create-PR eligibility is evaluated proactively for the sidebar, so that
popped an OS keychain prompt just from opening a worktree. Presence is enough;
creation itself still fails closed on an unusable credential.
Grouped/triage sections now read newest-first so recent discussion surfaces
at the top, while Timeline keeps its oldest-first history order.
- rename sortPRCommentGroupsForTimeline to sortPRCommentGroupsByRecency and
add an order parameter
- rank threads by their latest activity under newest-first so a fresh reply
refreshes an old thread
- break timestamp ties by numeric comment id in the sort direction, so a
GitHub review batch sharing one timestamp orders correctly
- sink groups with an unparseable createdAt to the bottom in both orders
Co-authored-by: Orca <help@stably.ai>
* Allow direct search with configured search engines
Users can now search directly from the tab creation menu using their
configured search engine. Forced search mode (`?` prefix) skips file and tab
matching for guaranteed search. Refactored tab-entry operations into focused
modules for clarity: forced-search parsing, network-safe selection, keyboard
focus, copy strings, empty options, and props types. Search routes through
the same workspace browser tab opening mechanism used for URLs, with safe
title and query presentation that doesn't retain sensitive details.
* Allow direct search with configured search engines
- Permit search and URL navigation while file index loads; require explicit
selection only when needed, not automatic opening
- Block malformed IPv6 addresses in bracket notation to prevent misclassification
- Support Kagi private-session links via searchUrlOptions
- Improve error handling with accessibility: show error messages in status region,
disable input during submission, display loading state
- Fall back to local browser tab creation when remote creation fails instead of
throwing; avoids remote availability blocking local search/navigation
- Add error translations for all supported locales (es, ja, ko, zh)
* Allow direct search with configured search engines
- Extract tab create entry lifecycle to key-driven component remounting, replacing conditional state reset with useEffect cleanup
- Consolidate network tab entry classification and request building into reusable helpers, eliminating duplicate logic
- Simplify owner resolution by inlining logic directly into openWorkspaceBrowserTab
- Replace custom surrogate-pair handling with native String.toWellFormed() for search queries
- Disable explicit URL classification to prioritize search-engine queries over raw URLs
* Allow direct search from quick-open tab bar entry
- Single-token queries keep file matches ranked above search (quick-open intent)
- Multi-word phrases promote search to top, since they cannot be file paths
- Arm network actions once file index fails or text is unambiguous search
- Cache prepared file index to avoid re-processing per keystroke
- Generate specific tab titles (e.g. 'example.com/docs') instead of generic 'Open URL'
- Surface opening workspace when launching browser tabs remotely
* Use readOnly instead of disabled for pending search input
Maintain keyboard focus during submission so arrow/Escape navigation continues to work. Use aria-busy to indicate loading state accessibly. Also fixes button hover styling when disabled and cleans up error message handling in the classifier.
* Treat bare searches as prompts; refine path and IP classification
- Bare search queries (e.g., "?") no longer display as error rows
- Path prefixes with existing matches are no longer blocked mid-keystroke
- Private IPv4 addresses now use http, public addresses use https
- Ambiguous inputs with non-numeric ports fall through to search instead of blocking
- Improve diagnostics by logging failure reasons in openFailure
Same hang class #11568 closed for .cmd: copilot-hook.ps1 ran
[Console]::In.ReadToEnd() and kimi-hook.sh captured stdin before checking the
Orca env, so a user-wide hook fired outside an Orca pane blocked forever when
the caller abandoned the pipe — one stranded powershell.exe or bash.exe per
hook event.
Move the env guard (after the endpoint refresh, which can supply PORT/TOKEN)
above the read on the Windows-local variants. The payload is discarded on the
missing-env path anyway; the broken writer is the trade #11568 established.
POSIX variants keep capture-first — those callers close stdin, and exiting
mid-write there surfaces as EPIPE the agent can see (#8110). Kimi's remote
install now asks for the posix variant explicitly.
Verified on a real Windows host: pre-fix copilot with an abandoning caller
never exits; post-fix both exit 0 immediately, and the env-present copilot
path still consumes stdin and delivers the POST (paneKey + payload received
by a local listener).
* fix(browser): exclude Google cookies from imports (STA-3811)
Imports never write and never remove a google.com-family cookie, on any
path. Signing in directly inside Orca is the only Google session that
survives, so the live jar always beats anything an import could plant.
* fix(browser): roll back selective cookie clears
While the auth document is on screen the WebContents UA is Firefox, so its
cross-host subresource/XHR requests (gstatic, play.google.com, the sign-in
challenge endpoints) reached the header layer carrying the Firefox UA yet still
bearing Chromium client hints, which the else-branch rewrote to Chrome. That
paired a Firefox UA with Chrome client hints on every non-auth Google host — a
sharper cross-host identity tell than either signal alone, and a plausible
cause of the password-submit challenge greying out and stalling.
Strip client hints on any request already carrying the Firefox auth UA so the
UA and hint surfaces tell one Firefox story for the whole flow. Gated on the
same googleAuthOverride flag as the auth-host switch, so imported-native
profiles are unaffected and the clean-Chrome default for non-Google sites
(Cloudflare) is untouched.
Extends tests/tools/google-signin-ua-probe.cjs with app-current/app-fixed
modes that mirror the shipped code and log per-request identity; on the real
accounts.google.com load they show 18 firefox-ua-with-chrome-hints cross-host
mismatches before and 0 after.
* fix(settings): read voice settings from a ref, not a stale closure
updateVoiceSettings wrote from a closure captured at mount, so an async
continuation could clobber a newer value. Uses the latest-ref pattern already
in Settings.tsx:405-407.
Found while investigating voice-microphone-selection.spec.ts:107 and confirmed
NOT to be its cause — that failure has no established root cause. The spec
gains a hydration barrier and toBeEnabled assertions here, which improve its
failure message but do not fix it.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): drop the voice spec changes from this PR
The hydration poll was a no-op: prepareVoiceSettings awaits updateSettings
inside the same page.evaluate, so the store is already committed before the
poll samples it, in all three tests (the reload precedes it). Worse, its
comment asserted that the Select is disabled until hydration — the CI
accessibility snapshot at the moment of failure shows the switch checked and
the combobox enabled, so that theory is refuted, and a wrong 'why' comment is
worse than none.
The toBeEnabled assertions go with it: the control was enabled when it hung,
so they would not have caught the real failure either.
Leaves this PR as what it actually is — a product bug fix. The spec's failure
stays open and unexplained.
Co-authored-by: Orca <help@stably.ai>
* test(settings): type the clear-key mock as its API contract
vi.fn(() => clearing) returned Promise<void> where the preload contract is
Promise<{ configured: boolean }>. Vitest does not typecheck, so this passed
locally while tsc -p tsconfig.tc.web.json failed — it would have gone red in
CI's required typecheck job.
Co-authored-by: Orca <help@stably.ai>
* fix(settings): write the voice-settings ref in an effect, not during render
React Doctor's 'Ref mutated during render' rule is blocking in CI's static
analysis job, and it is right: React can replay or discard render work, so a
render-time ref write can leak from UI that never commits.
Moving it to an effect keeps the behavior this fix needs — every writer here
(key-status probe, save-key, clear-key) runs post-commit, so the ref is current
by the time any of them reads it. Test still discriminates: red without the
fix, green with.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* fix(claude): guard cold-restore resume selectors
Persisted Claude default args or a custom command can carry their own
--resume/-r/--continue/-c selectors (a bare picker default or a stale id).
Cold restore appended the authoritative --resume <id> after them, typing a
command with competing selectors into the restored pane (#12982).
buildAgentResumeStartupPlan now routes Claude through a selector guard that
tokenizes the base with the existing startup tokenizer, strips selectors in
option position only (value-taking options keep dash-leading values), and
appends exactly one authoritative selector, inserting before Claude's own
-- terminator when present. Splicing is span-based so untouched bytes stay
verbatim, wrapper commands are left alone, and any tokenization failure
falls back to the previous append-only behavior. Launch paths, other
agents, persistence, and the wire are unchanged.
* fix(claude): harden resume selector guard against false matches
Round-1 review findings: locate the claude executable by command position
(index 0, after a wrapper --, or behind NAME=value assignments) so an
argument merely ending in /claude can never be mistaken for it; stop
matching the joined -r<id> form, which was ambiguous with dash-leading
option values and forced an unmaintainable arity table (now deleted).
Ambiguous shapes degrade to the pre-guard append-only behavior.
* fix(claude): fail resume guard open on chained shell syntax
Round-2 review findings: an unquoted operator or newline after the claude
token means the base chains other commands, and splicing across that
boundary handed the selector to the wrong command — detect it and fall
back to plain appending. Also recognize claude behind PowerShell's & call
operator, decouple the test oracle from the implementation's selector
predicate, add Windows tokenizer span tests, and rename the module after
its public API.
* fix(claude): flag bare shell operators inside the tokenizers
Round-3 review findings: the guard's operator scan compared raw source to
token value, so one quote or escape anywhere in a token hid a shell-active
operator outside the quotes and the splice crossed a live command boundary,
losing the resume entirely. Both tokenizers now flag tokens carrying an
unquoted, unescaped operator byte (or a word-leading # comment on
posix/powershell) on their spans, where quote state actually lives, and the
guard fails open on that flag. Also strengthens the redirect fail-open test
to carry a stale selector, re-tokenizes each raw span in the shell span
tests, and documents agent-resume-argv-drop as codex-only.
* fix(claude): flag expansions and clamp separator backoff
Round-4 review findings: unquoted multi-token expansions (backtick, $(, ${)
split across whitespace, so removing only the recognized selector token left
a broken construct tail — both tokenizers now raise the span flag (renamed
bareShellSyntax) for those openers, on cmd also for operators between
single quotes, which cmd does not treat as quoting. The separator backoff is
clamped to the previous token's span end so a token ending in an escaped
space can no longer donate its escape to the appended selector.
* fix(claude): treat cmd single-quoted regions as unmodelable
Round-5 review finding: cmd.exe has no single-quote syntax, so the Windows
tokenizer's grouping of a single-quoted region diverges from what cmd
parses — literal argv like 'claude ...--resume... old' was being read as a
real selector and stripped, and a literal '--' as claude's terminator. Flag
any cmd single-quoted token as bareShellSyntax so the guard fails open.
* fix(claude): flag quoted expansions and scope assignment prefixes
Round-6 review findings: the span flag was only evaluated in the unquoted
branch, so an expansion opener inside double quotes went unflagged — and
inside $(…)/backticks a nested quote re-opens a context this tokenizer
does not model, so the splice could cut mid-construct (syntax error, or a
silently mutated substitution body). Both tokenizers now flag those, and
the flag is renamed divergesFromShell to say what it means. Restrict the
NAME=value command-position prefix to posix, where that syntax exists.
Drops two branches proven dead.
* fix(claude): model shell-literal escapes and scan the whole base
Round-7 review findings: (1) the divergence scan started after the claude
token, so an expansion opened in a prefix — $(x; npx -- claude --resume s) —
had its closer spliced away, producing a base bash cannot parse; it now
covers every token including the executable, exempting only PowerShell's
leading call operator. (2) posix drops a double-quoted backslash the shell
keeps literal, and the Windows escape branch ran inside quoted regions where
cmd/PowerShell keep the escape byte literal — both now flagged, so a literal
can never be misread as a selector. (3) an unquoted line continuation hid a
selector inside a token and skipped the newline gap check.
Also removes a third provably dead branch and collapses the cut floor into
the cut itself.
* fix(claude): flag escapes the tokenizer models but the shell removes
Round-8 review findings, all one family — escapes whose token value hides
a selector the shell would see: a double-quoted line continuation (bash
deletes both bytes), posix $'…'/$"…" quoting, a windows escaped newline,
and a trailing unpaired escape. The last one was previously written off as
pre-fix-identical, but once stripping happens the dangling escape swallows
the separator and no exact --resume reaches claude at all — strictly worse
than appending, so it must fail open. Also folds the three gap predicates
into one scan.
* fix(claude): stop over-flagging a literal dollar sign
Round-9 review findings from both lanes: inside double quotes only $( and
${ open an expansion — $' and $" are literal there — and a trailing $
was flagged unconditionally because JS ''.includes('') is true. Both made
the guard fail open on modelable bases, leaving the stale selector to
compete, so #12982 went unfixed for them. Separately, cmd strips ^ before
the child re-splits on the bare whitespace, so an escaped separator hides
two real arguments and must fail open rather than drop one.
* fix(claude): fail open on cmd caret-quotes and bare PowerShell syntax
Round-10 review findings, both Windows-only (a bash oracle cannot reach
them): cmd strips a caret before a quote and the child's parser then reads
a bare quote delimiter, so the tokenizer's word boundaries stop matching
argv — one case turned a working resume into no resume at all, another let
a stale selector survive the splice. And bare (…)/{…} are live PowerShell
syntax in argument position, so splicing through them emitted unbalanced
output that PowerShell cannot parse.
* fix(claude): fail open on the PowerShell stop-parsing token
Round-11 review finding: after a bare --%, PowerShell passes the rest of
the line to the child literally, so the guard stripped a real selector and
then appended quoting that arrives as literal bytes — claude ends up with
no exact --resume at all, worse than leaving the stale one. Quoted "--%"
and cmd, where the token is ordinary, still splice.
* fix(claude): model cmd backslash-escaped quotes
Round-11 review finding: an odd run of backslashes before a quote makes it
a literal byte to the child's CommandLineToArgvW parser, not a delimiter,
so the tokenizer's word boundaries stopped matching argv. Orca manufactures
that pattern itself — quoteStartupArg wraps every token in quotes without
escaping a trailing backslash — so a pasted Windows path was enough to move
the selector into a desynced region and leave claude with no resume flag.
Also replaces a caret test case that was byte-identical before and after
its own fix, and merges two stacked comment blocks.
* fix(claude): fail open on PowerShell double-quoted escape sequences
Round-12 finding: PowerShell expands backtick escapes only inside double
quotes, so a sequence there produces a token value argv never sees — the
guard could strip "-`r" plus the argument after it. Also narrows the
stop-parsing comment: a quoted --% can engage stop-parsing before a
parameter token, where the base is already mangled either way.
* fix(claude): flag PowerShell escape sequences in bare arguments too
Round-13 finding: the previous commit gated on quote === '"', but
PowerShell's tokenizer calls Backtick() from ScanGenericToken, so it
expands these sequences in unquoted arguments as well — bare -`r really
is a control character, not -r. The guard read it as a selector and
dropped it plus the argument after it. Widening to all PowerShell
contexts measures 0 under-flag and 0 over-flag across the full printable
matrix; the backtick-escaped-space idiom still splices. Also swaps a test
case that was byte-identical with and without its own fix.
* fix(claude): drop a token-leading PowerShell backtick before whitespace
Round-14 observations, all pre-existing and measured: PowerShell drops a
token-leading backtick together with the whitespace after it, emitting no
token, so the tokenizer's extra token shifted the locator; and a backtick
before a bare CR is a line continuation too. Flagging both takes the
lane's 329k-base sweep from 87 bad to 0 with no new failures and the
must-splice list byte-unchanged. Also corrects a comment that no longer
listed every PowerShell divergence.
* docs(claude): correct the bare-CR rationale in the tokenizer comment
Round-15 verified against a real PowerShell 7.6.4 engine: a backtick
before a bare CR is not a line continuation there — pwsh keeps the CR in
the token. The flag stays because 5.1 is unverified and failing open costs
nothing, but the comment now says that rather than claiming continuation.
* fix(windows): refresh persisted PATH ordering
New terminals now preserve Orca-injected entries while adopting current machine and user PATH precedence, so newly installed tools are not shadowed by stale aliases. Invalidate the cached registry snapshot on Windows setting changes and guard in-flight refreshes from restoring stale cache data.
* fix(windows): retry invalidated PATH reads
* fix(windows): reuse current PATH refresh
* fix(windows): preserve per-hive PATH fallback state
---------
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
* feat(sidebar): add jump-to-top button for hard upward scrolling
Detect intentional hard scroll-up gestures (wheel or scrollbar drag) and
offer a one-click jump-to-top affordance. Auto-hide after idle to avoid
persistent visual clutter. Addresses the common case of fast navigation
through long worktree lists ranked by agent activity.
* test(sidebar): improve scroll-to-top detection for active gestures
Only detect velocity from active scrollbar drag or touch, not
programmatic scrolls. Return focus to list after jump-to-top. Add
comprehensive hook tests with gesture simulation. Improve cumulative
down-delta tracking for dismissal.
* test(sidebar): add scroll-to-top gesture detection tests
Add comprehensive test coverage for the hard-upward-scroll detection hook,
verifying idle timer behavior, gesture suppression, scrollability checks,
and cleanup on unmount. Extract the post-jump suppression window into a
named constant for maintainability.
Adds GitHub stacked pull request creation: a contextual "Stack this PR above #N" option that appears only when the selected base branch has an open PR, plus the main-process stack preflight and registration.
Also reworks the create-review composer for cohesion: shadcn Checkbox and Label primitives, base label above a full-width searchable combobox with attached results, keyboard navigation, and a unified field skin, spacing and typography scale.
Verified end to end against real GitHub: extending an existing stack and creating a new one.
* feat(bitbucket): connect Bitbucket from Settings with encrypted credential storage
Bitbucket Cloud was the only review provider with no in-app auth: GitHub and
GitLab delegate to the gh/glab CLIs, but Bitbucket has no comparable
first-party CLI, so the only option was ORCA_BITBUCKET_* env vars plus a
restart (discussion #5364).
Adds a Connect/Edit/Disconnect flow on the Bitbucket integration card,
modeled on Linear and Jira:
- Credentials are verified against /user before they are persisted, so a
dead token is rejected inline instead of silently stored.
- The secret is encrypted with safeStorage (0600 plaintext fallback when no
OS keyring); non-secret metadata lives in a separate plaintext file so
status reads render the connected account without decrypting. Opening
Settings therefore never triggers a keychain prompt.
- Env vars keep precedence over stored credentials, so existing headless and
SSH setups are unaffected. Env-managed connections hide Disconnect.
- connect/disconnect reset the preflight cache, so no relaunch is needed.
The Bitbucket card moves to its own file to stay under the tsx max-lines cap.
* feat(bitbucket): support creating pull requests from Orca
Bitbucket was the only configured provider whose Create button reported
"This repository provider does not support creating a pull request from
Orca" — supportsReviewCreation was false and the forge provider had no
createReview, so even a correctly authenticated setup was blocked.
Adds createBitbucketPullRequest against POST /repositories/{ws}/{repo}/
pullrequests, using the same env-first / stored-credential resolution as PR
lookups (extracted into resolve-auth.ts so both share one path).
Bitbucket Cloud has no draft pull requests, so a draft request is rejected
with a clear message rather than silently publishing a live PR.
* fix(bitbucket): hide the draft toggle where drafts do not exist, plus review fixes
Bitbucket Cloud has no draft pull requests, so the composer no longer offers
the toggle for it and forces the flag off at submit — better than failing
after the user has filled the form in.
Review fixes:
- writeFileSync's `mode` only applies when it creates the file, so rewriting
a credential kept whatever permissions it already had. chmod after every
write, for the secret and the metadata.
- An explicit ORCA_BITBUCKET_API_BASE_URL now wins over a stored base URL.
Env precedence is per-setting, not all-or-nothing.
- Enter in the credentials dialog only submits from a text field, so it no
longer hijacks Cancel and the docs link.
- Replace the chmod-based delete-failure test with a mocked unlinkSync: file
modes are not portable to Windows and elevated runners unlink anyway.
* fix(bitbucket): stop a merged pull request from blocking the branch's next one
Reported on #5832: with a merged PR on a branch, Create reported "Pull
request already exists" and offered no way forward.
The branch lookup queries every PR state and returns the most recently
updated one, so a merged PR came back as the branch's current review and
eligibility blocked on it. Bitbucket only discarded such a match on the repo
default branch (#9171), while GitHub already drops any merged PR it matched
by branch alone — "a merged PR without an explicit link is just a historical
branch match, not implicit review context".
Applies that rule to Bitbucket. An explicitly linked review still resolves
through the linked-number fallback, so merging a PR Orca knows about keeps
showing it.
* fix(bitbucket): add bitbucket to the shared review-creation provider list
Reported on #5832: on a Bitbucket repo with no existing PR, Create still
said "This repository provider does not support creating a pull request
from Orca", even after the forge provider gained createReview.
There are two capability lists. Enabling supportsReviewCreation on the forge
provider was necessary but not sufficient — the blocker and the whole
renderer read the separate shared list, which never included bitbucket.
Adds it, gives Bitbucket its own provider name so review copy stops saying
"GitHub", and asserts the two lists agree so they cannot drift apart again.
* fix(bitbucket): persist pull request links after creation
* fix(bitbucket): fetch linked pull requests by number first
* fix(i18n): use generated Bitbucket integration keys
* test(bitbucket): cover forge creation delegation
* fix(bitbucket): fall back when linked pull request is stale
* docs(bitbucket): explain notFoundIsNull and fix a garbled permissions comment
notFoundIsNull arrived without the rationale its sibling flag carries, and
reads as a bare `true` at the only call site that opts in.
* fix(bitbucket): address review findings before merge
Two of these made the feature unusable in real setups:
- Create PR checked GitHub authentication for Bitbucket. isProviderAuthenticated
fell through to isGitHubAuthenticated, which was unreachable while Bitbucket
could not create reviews at all. Anyone with Bitbucket connected but no
`gh auth login` got auth_required with no way forward.
- The draft flag was only gated in ChecksPanel, not the two SourceControl call
sites. With "create as draft" saved as a default, the composer hides the
toggle for Bitbucket, so the flag could not be cleared and creation failed
every time. Bitbucket now ignores draft instead of rejecting it.
Also:
- Blocked-create copy said "GitHub is not authenticated. Run gh auth login" on
Bitbucket repos, in both the main-process and renderer paths.
- A decryption failure resolved to an anonymous config and queried anyway; a
private repo answers 404, which reads as "no pull request" and offers Create
for a branch that already has one. Requests now fail closed.
- Hiding non-open implicit branch matches was too broad: a declined PR became
permanently invisible off the default branch. Scoped to merged, restoring the
default-branch rule (#9171) for the rest.
- A failed disconnect rejected unhandled and the card silently re-rendered as
connected; a partial delete left the secret live in memory for the session.
- The credentials dialog refused to open on a remote runtime, so a local repo
could never store a credential. Now only the storage note changes, matching
the Jira dialog.
---------
Co-authored-by: devatnull <59279509+devatnull@users.noreply.github.com>
* fix(terminal): retain WebGL for recently hidden worktrees
Suspending a hidden worktree disposes every pane's WebGL addon, so every
switch-back presents DOM-renderer frames (unfloored ~5% wider advance, and
the channel through which any poisoned cell metrics reach the screen) until
reattach completes — 0.5-3s on loaded sessions. Keep the live addons for the
most recently hidden worktrees instead: LRU capped at 6 contexts (Chromium
allows ~16/page, visible worktrees use 1-4), least-recent evicted to dispose
exactly as before. Reveal then has no renderer swap and nothing to flash.
Window-occlusion callers pass no retention context and are unchanged.
* fix(terminal): pause retained hidden cursor work
* fix(terminal): preserve retained WebGL through deferred rebuilds