Commit Graph
2191 Commits
Author SHA1 Message Date
Brennan Benson 86b878cfd6 fix(mobile): parse classified PR lookup outcomes (#12659) 2026-08-05 11:31:11 -07:00
Brennan Benson 5d2ad3597a fix(native-chat): add direct Codex model selection (#12657)
* fix(native-chat): select Codex models directly

* fix(native-chat): confirm agent exits before switching views

* fix(runtime): handle unavailable foreground probes
2026-08-05 11:27:12 -07:00
Brennan Benson de64337c26 fix(worktree-watcher): refresh status after external pushes (#12361)
* fix(worktree-watcher): surface external push -u through the git-common watch

An external-shell 'git push -u' writes only the common .git/config (plus
refs/remotes/<remote>/<branch>), both invisible to the git-common event
filter, so the Checks panel stayed on 'No upstream configured' until the
renderer safety poll. Classify the common config and remote-tracking refs
as status-tier signals, poll config alongside the other primary-checkout
metadata files, and keep FETCH_HEAD/reflog/ref-lock churn ignored.

* fix(worktree-watcher): refresh after subsequent pushes
2026-08-05 11:26:13 -07:00
6942871194 feat(editor): add Markdown table structure controls (#11985)
* feat(editor): add markdown table structure controls

* fix(editor): scope table context actions to cells

* Replace table toolbar with context-aware overlay controls

Replace the fixed table toolbar with context-sensitive overlay controls that position themselves around the active table, adding support for direct row/column insertion and full table deletion. This approach is less intrusive and supports click-targeted actions via coordinate-based cell resolution. Enhance structural safety by preventing header removal and ensuring tables never collapse below a single cell, deleting instead when the final row or column is removed. Harden the context-menu query with a 120ms timeout to keep the native menu responsive even if the renderer hangs.

* Replace markdown table context query with IPC coordination

Capture table cell targets on pointerdown and report via IPC channel
instead of executing JavaScript on context-menu events. Eliminates
120ms query timeout and unavailability race conditions. Header cells
now disable incompatible row-level actions.

* fix(editor): make table column rebalancing atomic with insertion

- Refactor rebalanceAddedColumn to mutate the caller's transaction, grouping
  insertion and rebalance into a single undo step
- Add validation for cached cell positions that may outlive the document
- Fix table detection to use isInTable() instead of isActive('table')
- Correct z-index layering to respect menu stacking context
- Fix cleanup of stale animation frames and pending pointer state

---------

Co-authored-by: rainL <WYK15@users.noreply.github.com>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
2026-08-05 11:10:05 -07:00
Jinjing 9fb4dbe8eb fix(ssh): handle rejected PTY deliveries with targeted recovery (#12746)
Add targeted recovery for rejected PTY source frames instead of
terminating the relay channel. Classify rejection reasons (malformed,
generation mismatch, range invalid) and attempt recovery based on the
rejection type. Implement admission control at publication time to ensure
frames aren't delivered after ownership changes. Bound recovery attempts
and retry with backoff to prevent exhaustion. Diagnose and log rejection
reasons to aid debugging.
2026-08-05 10:30:04 -07:00
Jinwoo Hong eea0bb64db fix(ssh): make PTY owner admission explicit and non-destructive (#12673)
An owner-capable `pty.openClient` had two failure modes that presented as something else.

If the relay still held an owner record but the request carried no matching resume proof, admission fell through to a SUBSCRIBER grant — a success-shaped response the client cannot use, which it then rejected as "did not grant an authenticated PTY session owner". And if the relay had forgotten the record the client named, admission threw a stale-recovery error, which the client answered by deleting its own recovery row — `clientInstanceId` included — and reopening. Two round trips, and the identity that lets it resume that target at all went with the deletion.

Now every owner grant carries a required `resumed` flag, a forgotten record mints a fresh claim in one round trip, a held claim returns one of three coded refusals, duplicate opens on one connection are rejected even when identical, and an attached-holder refusal becomes a typed error routed through the terminal-relay-error callback instead of feeding redeploy backoff a link that is working fine.

Independent review caught two regressions in the first attempt, both now fixed and both with tests that fail without them:

**A backpressure teardown could take a live owner's session.** The safety argument was that a record only becomes `disconnected` from an observed peer close — but two of the six paths there are capacity paths, where the relay destroys the client's socket itself because its lane queue filled. That is the signature of a client that is ALIVE but not draining fast enough. Demonstrated: the real owner is torn down for backpressure, a rival is granted ownership 270ms into a nominal 30s grace, and the owner's later reconnect with a valid resume proof is refused permanently, backoff cleared, no retry. Closes now carry a cause (`peer-closed` | `local`, defaulting to `local`, which only ever widens a grace), and the floor applies only to closes the transport actually observed on the peer's side. Capacity teardowns, decode faults and sink failures keep the default.

**A client's own zombie connection blocked it permanently.** Only `SshRelaySession` ever requests owner, and every endpoint-credential client shares one principal — so in a normal single-app deployment an `active` incumbent refusing you is almost always your own half-open connection the relay never saw close. That was refused as terminal, where main recovered on bounded backoff once keepalive noticed. The refusal already held both client identities; a match is now a distinct transient refusal that falls through to relay-lost backoff, restoring that recovery. A genuinely different client is still blocked.

Also: each retry deadline now starts when its own phase begins, instead of both being computed at entry where a slow first phase could leave the second with zero attempts.

Fixes STA-3365.
2026-08-05 02:56:07 -07:00
Jinwoo Hong d15939c5fd fix(terminal): recover rejected paired-runtime input (STA-2830) (#12675)
With a desktop client paired to a remote Orca runtime, terminal panes could report connected, writable, and `terminal.send` returning accepted — yet keystrokes never reached the agent. No error, no banner, no recovery; input silently vanished.

The ticket was really two bugs. The attach half was already fixed by #12589 (subscriber-driven daemon attach), confirmed by reproducing against current main. This fixes the remaining half: a write the host refuses had no way to tell anyone.

A capability-negotiated `WriteUnavailable` opcode carries that refusal back to the client, where it feeds the pane's pre-existing recovery hook. Capability gating matters because decoders reject unknown opcodes on desktop — and, worse, silently drop them on mobile — so the signal is negotiated in the subscribe handshake. Verified per direction: an old host strips the unknown Subscribe key, an old client omits it so the host never emits, and capability cannot be inherited across resubscribe.

Independent review then found the signal was being delivered and discarded: recovery demanded an authoritative liveness answer, and `pty:hasPty` had no `remote:` guard, so a paired pane's id fell through to the LOCAL provider, which returned false, and recovery bailed before remounting. Every test stopped at the transport boundary, so all of them passed while the pane stayed just as stuck. `pty:kill` already had exactly that guard.

The fix makes main answer LESS rather than claim more: `pty:hasPty` now returns unknown for a `remote:` id instead of a fabricated false, because main cannot speak for another host's PTY. The remount is then authorized by positive evidence — the process that owns the PTY stating it refused this specific write over a live negotiated connection — not by inference from silence. Local and app-SSH ids keep the probe, where a false genuinely means the shell died. Nothing is destroyed on this path; the remount rebuilds the renderer over the session it already had.

An end-to-end test now carries a rejected write from the host through to an actual remount, which no prior test did. A surviving mutant was also killed: the legacy-binary capability gate could previously be deleted with nothing turning red.

The reliability gate stays experimental — live paired journeys and mixed installed-release evidence remain uncollected. Fixes STA-2830.
2026-08-05 01:41:27 -07:00
Jinwoo Hong a766ee4bcd fix(runtime): refuse to silently wake a deliberately slept pane (STA-3465) (#12672)
`activateMobileSessionTab` gated only on `publicTab.status !== 'ready'`. A deliberately slept pane publishes as `pending-handle` indefinitely — indistinguishable at that call site from a pane awaiting reconnect — so the reconnect probe added by #11542 respawned it with a re-resolved agent launch, waking something the user had deliberately put to sleep.

The first attempt refused activation for any pane with a `worktree-sleep` record, applied to every path. Independent review found that broke the documented wake gesture: opening the tab IS how those panes are meant to cold-restore (`wake-sleeping-agents-in-background.ts`: "Those panes cold-restore --resume when their own tab is opened"). A mobile tap sends the byte-identical call the reproduction test used, and in three of four topologies no wake clears the record first — so the tap became a permanent no-op with no feedback.

This carries intent explicitly instead of inferring it. A new shared `TabActivationIntent` ('user' | 'automatic') rides the existing ActivateTab schema as an optional additive field; `isAutomaticTabActivation` returns true only for an explicit 'automatic', so an absent value is permissive BY CONSTRUCTION in one place — an older client that does not send it keeps today's behavior rather than silently losing its wake gesture. The field is required on the mobile helper's params, so no call site can be added without declaring who asked.

Every user path (mobile tab switches, paired tab clicks, shortcuts, palette, the pane's own open) is labelled 'user'. The only automatic sender in the codebase is `waitForResubscribeHostSessionHandle`, the #11542 reconnect probe.

Verified per topology: user activation materializes a parked pane under headless serve, a paired runtime client, a completed agent with restoreOnTabOpenOnly, and a running agent whose wake cleared the record. The automatic probe is refused without retiring the surface, and #11542's reconnect tests stay green.

Also fixes a test fixture that made a real bug untestable: the store stub ignored the host id, so mutating the partition lookup to 'local' left the suite green. Correcting it exposed three existing SSH reattach tests that had been relying on that looseness — their workspace session sat in the local partition while their repo was SSH-hosted, a store production would never read. Production was always right; the tests described an impossible world.

Fixes STA-3465.
2026-08-05 01:28:07 -07:00
Jinwoo Hong 2ec36a95c4 test(runtime): pin transport error code/message classification agreement (#12676)
Since #12667, a present error code short-circuits classification: a code that is genuinely transient but missing from `RECOVERABLE_CODES` classifies as FATAL. That is the shape that dead-ends terminal panes — #12650 fixed exactly that for a different error, where a transient failure misclassified as fatal unmounted the Reconnect banner and left recreating the session as the only escape.

Today the code and fragment sets agree. Nothing prevented a future code from being added without a matching entry, and the failure would have been silent.

This pins that agreement: for every reachable transport error, a code that classifies fatal must not carry a message that would have classified recoverable. 57 coded pairs plus 8 code-less ones, derived by invoking the producers where possible so a reworded message updates the corpus instead of leaving a stale copy silently passing. The failure message names the offending fragment and says what to do about it.

Enumeration turned up producers beyond the obvious ones — notably the Tailscale-hinted variants, where `runtime-environment-transport-routing.ts` mutates the message on an already-coded error before it crosses IPC, making those distinct corpus members.

Also documented (not asserted, because it is unreachable today): `runtime_rpc_queue_overloaded` is absent from both host passthrough allowlists, so if it ever crossed `mapRuntimeError` it would flatten to `runtime_error` while keeping its "queue is full" message — precisely the dangerous shape. The queue pool is never instantiated on the server dispatcher, so it cannot happen now.

The known exception is pinned rather than silently exempted: a dedicated test records WHY the guard cannot see `remote_runtime_busy` (fatal by code, matching no fragment, so the two sides have nothing to disagree about). If someone rewords a busy message into connection wording, that test fails and points at STA-3479.

Proven non-vacuous by four separate injections. The only production change is two `const` to `export const`.
2026-08-05 01:01:19 -07:00
Jinwoo Hong 4e370062a8 fix(remote-runtime): make hidden-output recovery reason-driven instead of timer-guessed (#12655)
When a remote terminal tab is hidden, the host stops sending its output and discards what it queued, so on reveal the only way to recover the missed output is to ask the host to serialize its buffer. That reply was ambiguous — one empty answer covered several unrelated situations — so the client inferred "output is lost" from elapsed time, using budgets sized for local IPC. Over a network that guess was routinely wrong: users saw "[Orca skipped hidden terminal output because main recovery was unavailable.]" on a healthy pane and got a permanent scrollback gap, worst exactly when an agent was streaming heavily and there was the most to lose.

The key insight is that there is no provable-absence case at all. A pane with genuinely no retained output returns a SUCCESSFUL snapshot with empty data, because the host serialized fine and found nothing. The real defect was the host sending an untagged empty reply when no serializer answered — reporting an unprovable failure as proven emptiness.

The host now states why a snapshot is unavailable and the client acts on that reason: an empty snapshot is success; retry-worthy retries and then gives up honestly; permanently-unavailable banners immediately with no waiting; and a host too old to say latches that pane to the pre-existing timer heuristic. Local panes are unchanged. The self-heal repaint no longer yanks the viewport of a user scrolled back reading — it waits for the terminal to return to following output.

Retries are bounded by COUNTING REPORTED OUTCOMES, never elapsed time. Independent review found that the single budget also charged attempts for causes returned locally, where the host was never asked — meaning a re-arming resync could exhaust it and banner on a perfectly healthy host, a residual instance of this very bug. Host answers and local gates now have separate budgets; local gates send zero frames, so retrying them cannot pressure the host.

Review also found a duplicate-banner path where a repaint timer armed before a permanent answer survived the abandon; the clear is scoped to the branch that banners, since the retry loop deliberately arms that timer.

Wire change is additive: an optional field on an existing frame, dropped on the success path, so old clients see an unchanged frame. STA-3476 tracks replacing the legacy-host detection (currently inferred from an absent field) with a positive capability signal. Closes STA-3457.
2026-08-05 00:51:39 -07:00
NeilandOrca aca5e8b5b1 fix(terminal): keep a live TUI's mouse modes across daemon reattach (#12461)
Co-authored-by: Orca <help@stably.ai>
2026-08-05 00:51:33 -07:00
Jinwoo Hong 15ef69a814 refactor(runtime): preserve transport error codes over IPC (#12667)
Errors thrown across Electron's `ipcMain.handle` lose their structured error code — only the message survives. So the renderer classified transport failures by matching substrings of English message text. That is how a queue-overload rejection escaped classification during a remote outage and surfaced as a raw error wall: the code was stripped in transit and its message fragment was not in the recoverable list.

This converts `RemoteRuntimeClientError` and `RuntimeRpcCallQueueOverloadError` rejections from `runtimeEnvironments:call` into the existing structured `{ok:false, error:{code,message}}` response, which the preload already passes through unchanged and `unwrapRuntimeRpcResult` already reconstructs with the code intact. Classification now treats a present code as authoritative and consults message fragments only when there is no code.

The fragment list is deliberately RETAINED as a backstop, not deleted: untyped main-handler rejections, subscription-start failures, and older code-less paths still rely on it.

Proven real rather than cosmetic: a test-only patch applied to unmodified main fails (4 failed / 66 passed) because the code does not survive the boundary today, and passes on this branch.

Independent review specifically chased the risk that a present-but-unrecognized code would now short-circuit to fatal where a message fragment previously rescued it — the shape that dead-ends a pane. It enumerated all 34 reachable code/message pairs and confirmed no pair flips recoverable to fatal, that the newly-serialized code set is closed and client-local, and that host-forwarded codes preserve recoverable classification by design. A differential harness over that corpus was verified non-vacuous by injecting the bad shape.

Nothing crosses the paired-runtime wire: desktop main -> IPC -> preload -> renderer only, reusing an existing response shape, no new fields or opcodes.

The connection-level offline state with a single reconnect affordance remains as STA-3456 follow-up work.
2026-08-05 00:44:05 -07:00
Jinwoo Hong 1667b77f0b fix(remote-runtime): keep a remote outage from flooding the error surface and dead-ending the pane (#12650)
When a remote runtime went unreachable (laptop sleep, Tailscale drop), the UI filled with dozens of repeated timeout errors until it was nearly unusable, and the affected terminal then accepted no input after connectivity returned — leaving "close the session and resume it in a new one" as the only escape.

Four causes, three of which were still live:
- Errors accumulated into one ever-growing surface with no de-duplication or cap.
- Queue-overload rejections lose their structured error code crossing the IPC boundary, so they were never classified as recoverable and surfaced raw.
- A transient failure misclassified as fatal called `recovery.cancel()`, setting the pane to an idle phase — which unmounts the Reconnect banner and makes manual retry, online and resume triggers all no-ops. A true dead end, and the reason recreating the session was the only way out.
- Dismissing an error cleared the surface but not the dedup memory, so an identical fatal error recurring in the same outage was suppressed forever while the pane looked healthy; dedup also compared single lines, so multi-line errors never matched and stacked without bound.

The ordinary reconnect loop was already fixed in v1.4.150/160 — bounded backoff, a Reconnect banner and auto-recovery already ship. This fixes what remained.

Note the fix routes fatal resubscribe failures back through the shared terminal error handler: bypassing it had silently dropped stale-handle re-resolution, terminal-gone retirement, SSH-expired recovery and oversized-snapshot suppression — a stuck-pane regression inside the stuck-pane fix, caught in review and covered by 6 dedicated tests.

Verified: reproductions red on main before the fix; after rebasing onto #11542, reverting the dead-end fix still turns its test red. Follow-up STA-3456 tracks preserving typed error codes across the IPC boundary so classification stops matching message text.
2026-08-04 23:21:31 -07:00
Jinjing eebaf47df0 Add 'Has Workspace' mode to show Linear issues linked to local worktrees (#12632)
* feat(linear): add 'Has Workspace' mode to show issues linked to local wo

Enable users to view and open existing workspaces attached to Linear issues
instead of accidentally starting duplicates. Includes shared worktree attachment
labeling for consistent UX across GitHub and Linear surfaces.

* fix(linear): apply search filter in 'in-orca' mode to prevent drops

- Apply search filter in 'in-orca' mode even without active context label to prevent
  team filters from silently hiding linked tickets (no "Fetch more" recovery path)
- Add aria-label to workspace-open button for accessibility
- Update tooltip from "local worktree" to "Orca workspace"
- Reorganize i18n: move workspace.open from lib.linear to components.issue
- Expand test coverage for workspace start and activation scenarios

* fix(linear): avoid mutating in-orca linked refs during render

React Doctor fails static analysis when refs are written during render.
Keep the latest linked refs in an effect so the in-orca loader can still
read them without re-running on identity-only worktree churn.
2026-08-04 22:12:00 -07:00
Brennan Benson c736031773 Fix setup-gated agent startup on long worktree paths (#12623)
* fix(worktrees): preserve gated agent startup on long paths

* fix(wsl): forward sequenced agent startup env
2026-08-04 22:01:35 -07:00
Jinjing fe72eeb75c Add linked issue guidance and ELI5 sections to PR generation prompts (#12613)
* Add linked issue guidance and ELI5 sections to PR generation prompts

Include linked GitHub issues in PR descriptions with Fixes/Refs guidance, and require ELI5 Problem and Solution sections before implementation details. Tests verify linked issue substitution and prompt structure enforcement.

* Include linked issue details in PR description generation

- Fetch the linked GitHub/GitLab issue title and body so generated PRs reference real issue context instead of just a number
- Use provider-specific reference syntax (Fixes/Refs, Closes/Related to, AB#) and label the issue by the active provider
- Feed issue title and description into the generation prompt while treating them as untrusted context, never as instructions
- Fall back to a cached work-item title when the provider lookup fails, and skip cross-provider issue attachment
2026-08-04 20:05:30 -07:00
Brennan Benson c511e51442 fix(mobile): label native-chat tool rows with a clean, expandable input summary (STA-3333) (#12498)
* fix(mobile): label tool rows with a clean summary, expand full input (STA-3333)

Mobile tool rows showed the raw input JSON (`{"file_path":…}`) as the row
label, and the expanded detail just repeated that same truncated string.

- `describeToolInput` labels a row with the target file path, else the
  primary argument (command/cmd/query/pattern/url/description), else the
  bounded JSON preview.
- Codex delivers tool arguments as a JSON string; normalize those into the
  object shape the helpers already understand, so labels, file links,
  run summaries and the expanded detail all work for Codex calls too.
- The expanded detail now renders the fully formatted input, capped at
  MAX_TOOL_RESULT_CHARS like desktop's tool detail (and like the result
  body), and a structured input makes the row expandable.

* fix(mobile): name search rows by their term and keep the filename in path labels (STA-3333)

Review follow-ups to the tool-row summary, all in the shared helper:

- A Grep/Glob row labelled itself with the directory it scanned and dropped
  the pattern entirely, because `toolFilePath` treats `path` as a file target.
  That path is a scan root, so it also rendered a tap-to-open link that asked
  the app to open a folder. `toolFilePath` now ignores the generic `path` key
  for search-shaped input, which lets the pattern win the label and drops the
  bogus link; an explicit `file_path` still wins.

- An overlong path was truncated from the head, cutting off the basename —
  the one part that tells two rows apart. Trim from the front instead, so
  the label reads `…/session/MobileNativeChatMessage.tsx`.

- The primary-argument chain used `??`, so a present-but-blank key selected
  itself and swallowed the keys ranked after it, dropping the label all the
  way back to raw JSON. Take the first key that actually yields a label.

Refs STA-3333.

* fix(mobile): don't offer an expander whose detail repeats the row (STA-3333)

An empty tool input formats back to the row label verbatim, so `{}` and `[]`
advertised an expander and then re-showed the label — the same repeat-the-JSON
problem this change set out to remove. Gate `isStructuredToolInput` on the
collection actually having contents; the lazy detail path is untouched.

Also pins the overlong-path test to the path itself: asserting only length<=80
plus a `…` passed just as well with path labelling deleted.

* fix(mobile): gate the tool detail panel on having detail (STA-3333)

The Tools toggle opens every row at once, bypassing the row's tap guard,
so a row with nothing to expand rendered its own label again underneath
itself — and the tap that would dismiss it is a no-op. Matches desktop.

* fix(mobile): keep a blank tool argument out of the run header (STA-3333)

Skipping a present-but-blank primary key let `briefToolArg` fall through
to the raw JSON preview, so a run header read `Bash {"command":""}` where
it used to read `Bash`. Also state the search-path trade-off honestly:
suppressing the link costs a file-scoped search its tap target.

* fix(mobile): only treat a blank primary key as a missing argument (STA-3333)

The previous guard tested key presence, so a populated but non-string
argument — a mixed argv like ['kill','-9',pid], or a structured query —
dropped out of the run header instead of falling back to the preview.

* test(mobile): pin the tool-row chevron to the detail panel (STA-3333)

The panel gate was covered but the chevron beside it was not: swapping
`showDetail` back to `expanded` on the icon alone left all 909 mobile
tests green, so the affordance lie this branch fixes could return
unnoticed — a down-chevron over no panel, on a row whose tap is guarded
off.

Asserts both icon counts on the fixture that test already renders. The
two halves now die for distinct reasons: the panel gate on the duplicate
label text, the chevron on the icon count.

* test(shared): pin the blank-search-key guard in the tool label (STA-3333)

Dropping `.trim()` from summarizePrimaryToolArg left all 32 tests green,
yet it leaks through isSearchToolInput: a whitespace-only `query` starts
counting as a search term, which suppresses `path`. One character takes
the row's label, its tap-to-open link and its run-header argument at
once, and puts the raw JSON label back — the bug this branch removes.

Asserts all three outputs on that shape. Kills only that mutant; the
isSearchToolInput mutant still dies on the existing search test.

* fix(native-chat): share tool input display semantics (STA-3333)

Build the tool row label, file target, detail eligibility and bounded detail from one normalized input model. Mobile no longer reparses JSON-string input across independent helpers or repeats an already-complete plain label, and desktop now uses the same clean row summary instead of retaining raw JSON.\n\nKeep full detail formatting lazy for collapsed rows and share the 4000-character detail cap across both renderers. Tests pin desktop adoption, mobile disclosure parity, one-pass JSON parsing and the shared bound.
2026-08-04 19:34:59 -07:00
Brennan Benson c3ddc0d5df fix(mobile): keep native chat ask dismissals tab-scoped and gated (STA-3333) (#12497)
* fix(mobile): keep native chat ask dismissals tab-scoped and gated

Dismissal state lived in the chat view subtree, which unmounts on a
chat<->terminal toggle, so an answered ask card came back on return. It
also had no tab scope and no waiting/blocked gate.

- move dismissal into the controller, keyed per session tab
- gate ask cards on waiting/blocked like the permission path already is,
  and retire a dismissal off the ungated detected prompt so a working/done
  status can't be mistaken for the prompt clearing
- ignore a dismissal that settles after its prompt cleared or was replaced

Refs STA-3333.

* fix(mobile): keep an ask dismissal through the transcript re-subscribe

A view toggle or tab switch re-subscribes the native-chat transcript, and
useMobileNativeChatSession withholds `messages` until that read settles. A
transcript-derived ask therefore reads as null while the chat surface is
already visible, so the reset effect took it as "the agent moved on" and
retired a live dismissal — the answered card came back, which is the bug
the off-chat guard was meant to close.

Treat an unobserved null as unobserved: `observing` now also requires the
read to have settled. A prompt that is already detected stays observable on
its own, so a status-derived ask still registers on first paint and an
answer taken during that first load is still accepted.

* fix(mobile): keep the transcript-derived ask outside the paused gate

A hook row idle past AGENT_STATUS_STALE_AFTER_MS (30m) projects to `done`
with no interactivePrompt, so the transcript fallback is the only source
left for a still-pending question. Gating it behind waiting/blocked made
that question unanswerable from mobile. Only the sticky status payload
needs the gate; `extractPendingAsk` clears itself on the tool result.

Also pins the load-window clause in the ask-observability guard, which
was behaviourally load-bearing but killed no test.

* fix(mobile): treat a never-read transcript as unobserved, not as "no ask"

The ask-observability guard only excused `transcriptLoading`, which is true
for an in-flight read alone. useMobileNativeChatSession also withholds
`messages` when the client is gone ('idle') or the tab has not reported a
provider session yet ('waiting-session') — both leave the flag false over an
empty list that was never read. The derived prompt then read as null, the
reset effect took that as "the agent moved on", and a live dismissal was
retired; when the read landed with the question still pending the answered
card came back — the resurfacing bug this guard exists to close.

Gate on the read having actually settled instead. 'error' still counts: it
keeps the last successful read in `messages`, so a prompt that clears under
it is real evidence, unlike a list that was never populated.

Also locks three guards that killed no test: the sticky-status suppression
of the transcript fallback (which is what makes the new paused gate hold in
the post-answer window), the reset effect's identity bail-out, and showAsk's
empty-prompt case. The transcript stand-in now derives `transcriptLoading`
from `status` the way the real hook couples them, so these tests can only
express states the session hook can reach.

Refs STA-3333.

* test(mobile): pin the ask dismissal's tab scope and ungated retirement input

Both wirings were unpinned: swapping `scopeKey` to a constant or feeding the
gated `ask` in as `detectedAsk` left the whole mobile suite green.

* fix(mobile): require a landed read before an errored transcript retires a dismissal

`status === 'error'` was treated as settled on the claim that an error keeps
the last successful read in `messages`. That only holds for an error that lands
on top of an earlier read. The host forwards an initial-drain failure as an
error frame carrying an EMPTY list (transcript-watch-error.test.ts), the mobile
frame applier checks `frame.error` before the messages array so those rows are
discarded, and the session hook's error path never calls `setMessages` — so a
first-read error leaves `messages` at the `[]` the identity-change effect wrote.

That frame is also not terminal: the watcher keeps `initialDrain` true and a
real snapshot follows once the read recovers. So a re-subscribe whose first
read errors made the never-populated list read as "no ask", retired the live
dismissal, and the recovered snapshot brought the answered card back over the
composer — the exact resurfacing this guard exists to close, and most likely on
remote/SSH transcript reads.

Require rows for the error case. Rows can only be present once a read landed,
so the predicate is never wrong in the resurfacing direction; it only declines
to retire a dismissal when the transcript was never observed.

Also drop the dismiss hook's `detectedAsk = ask` default and make both prompts
required. That default silently fed the gated prompt in as the detected one,
which is the pre-fix behavior: a paused-out card would read as "prompt gone"
and retire the dismissal. tsc now enforces the ungated payload at every call
site instead of leaving a trap for the next caller.

* fix(mobile): scope the ask dismissal to the provider session, not the tab

A restart, /clear, or resume swaps the provider session inside one tab. The
next session's first question is often byte-identical, so a tab-keyed dismissal
hid the live card and left the turn blocked with nothing to act on.

* chore: restore upstream formatting
2026-08-04 19:34:45 -07:00
Brennan Benson 38a892c980 feat(mobile): native-chat model/session-option picker + shared slash catalog (STA-3332) (#12366)
* feat(mobile): native-chat model/session-option picker + shared slash catalog (STA-3332)

Piece A — shared slash catalog + send classification:
- Mobile composer now serves getVerifiedNativeChatCommands from the shared
  catalog (agent-aware, with description rows) instead of a hardcoded
  provider-agnostic list that advertised commands Claude does not have.
- classifyNativeChatSend moves to src/shared/native-chat-slash-commands.ts
  (renderer re-exports keep desktop import paths stable); mobile's send seam
  now gates optimistic echoes on it, so slash sends no longer create a
  'Queued' bubble that no transcript echo can ever retire, and the
  ack-lost hold only arms for chat sends.

Piece B — mobile model/session-option pickers:
- New per-tab session-option tracking (state/commands/labels modules) ported
  from the desktop live flow, reading the shared agent-session-option
  catalog for Claude AND Codex.
- Composer pill row (model + options) opening an inline choice card in the
  proven Ask-card pattern; applies use catalog modelApply semantics
  (/model <value> via the existing send path), Codex-style agent-picker
  entries dispatch the picker command and flip the tab to the terminal view.
- Current model seeds from the hook-reported provider model when derivable;
  typed /model-style commands update tracked state (recordOutgoingCommand
  parity); dispatched values render as sent-not-confirmed.

* fix(mobile): keep session option sends scoped

* fix(mobile): synchronize native chat refs after commit

* refactor: share native chat session option logic

* fix(mobile): keep the live tab's session-option record from eviction

`getScopedRecord` returned an existing record without re-inserting it, so the
per-tab record map evicted by insertion order rather than recency. A long-lived
active tab is the oldest key, so crossing the 32-scope cap silently dropped its
tracked model and reset the pill to "Model". Desktop's scope cache does
delete-then-set for exactly this reason.

Also moves the shared session-option tests to src/shared so the root suite runs
them (they only exercised src/shared logic the Electron renderer consumes, but
sat under mobile/ where only mobile's vitest project sees them), and restores
two "why" comments dropped while extracting the shared modules.

* fix(mobile): stop a stale session-start report reverting a model pick

Re-entering a chat tab re-delivers the same `agentStatus.model`, and the
reported-model effect re-applied it unconditionally — so picking a model, moving
to another tab, and coming back reverted the pill to the model the agent reported
at session start, which cannot have observed the `/model` sent after it. The
status stream reconnecting had the same effect.

A report is now only treated as evidence when the matched catalog id CHANGES for
that scope; a genuinely new report still supersedes a local pick. Mobile has no
screen read to confirm a switch against, so the repeat is all we can key off.

* fix(mobile): close four session-option picker defects found in review

D1 — a picker apply could interleave with a composer send. The composer already
blocks a text send while an apply is dispatching, but not the reverse: the host
spaces a send's body and its Enter ~500ms apart, so an apply tapped inside that
window was submitted as part of the user's prompt, and the pill then claimed a
model change that never ran as a command. The pickers render inside the composer,
so they now take its in-flight state directly — the same guard, mirrored.

D2 — an option was filed under the wrong model. `setTrackedSessionOption` resolves
the owning model when it commits, not when the command was built, and the report
effect mutates the same record off-queue. A report landing mid-dispatch therefore
recorded `/effort low` against the model it switched TO. Ports desktop's
supersession guard, which skips the commit when the baseline moved.

D3 — a command template's prefix also matches prose that starts with it, so
"/model is a weird word" tracked that prose as the current model, rendered it as
the pill label, and matched no catalog model, dropping every per-model option.
Parsed values are now canonicalized against the catalog; a typed value containing
whitespace is treated as a prompt rather than a command.

Perf — `/` on a Codex tab returned all 45 commands into a non-virtualized
ScrollView showing ~5, re-reconciled on every streaming tick above the transcript.
Capped at 12.

Also splits the row primitives out of MobileNativeChatSessionOptionPickers.tsx,
which the D1 guard pushed to 402 effective lines against a 400 cap.

* refactor: share the session-option display ordering

CATEGORY_ORDER and the non-model sort were byte-identical in
NativeChatSessionOptionPickers.tsx and mobile's labels module — pure logic with
no i18n in it, so there was no reason for two copies that can drift. Both now
call sortNativeChatSessionOptions from the shared snapshot module.

* refactor(mobile): align model picker layout

* style(mobile): round native chat composer

* fix(mobile): inset rounded chat composer
2026-08-04 19:34:17 -07:00
Brennan Benson 0ce108d935 fix(browser): add native-UA session profiles (#12608)
* fix(browser): add native-UA session profiles

* test(browser): add Google sign-in UA probe

* fix(browser): preserve native profile UA identity
2026-08-04 19:07:23 -07:00
Brennan Benson 8c65dd5094 perf(runtime): keep PowerShell ACL work and a second auth off the remote command path (#12451)
* perf(runtime): keep PowerShell ACL work and a second auth off the remote command path

Two costs sat on the remote authentication path on Windows:

- The E2EE handshake persisted `lastSeenAt` inline, and every secure-file write
  spawns PowerShell synchronously twice to reapply the registry ACL, so the
  client's `e2ee_authenticated` waited on both spawns.
- Every remote CLI command except `status.get` opened a second full WebSocket
  connection just to re-read status for the protocol-compat check, doubling the
  authentications per command.

The first sighting of a device still persists inline (rotation drops entries
disk says were never scanned); later refreshes update memory now and coalesce
onto one deferred write. The compat verdict is saved against the runtime's
per-launch `runtimeId`, so a restarted or upgraded runtime retires it.

* fix(runtime): preserve compatibility on one remote auth

* fix(runtime): flush registry after transport shutdown
2026-08-04 17:04:51 -07:00
Brennan Benson a528b689a9 Prevent Command Code output from hijacking agent icons (#12573)
* fix terminal agent icon ownership

* fix terminal output ownership gaps
2026-08-04 16:06:54 -07:00
Brennan Benson 9deee5ad2f perf(worktrees): delete worktree directories after the removal returns (#12416)
* perf(worktrees): delete worktree directories after the removal returns

`git worktree remove` deleted the whole checkout inline, so the remove IPC held the
watcher/PTY gate for the entire recursive delete (prod traces: worktree.remove.git_remove
p50 8-14s, p90 29s, max 34.7s). Local removals now rename the checkout into a hidden
sibling trash root, clear Git's registration for the missing path, and delete the moved
tree in the background. Renames that cannot run (WSL, other volume, Windows open handles)
fall back to the previous in-place removal unchanged.

* test(worktrees): keep no empty trash root when the rename cannot run

* fix(worktrees): harden deferred trash cleanup

* fix(worktrees): keep WSL trash on its owning host
2026-08-04 15:53:11 -07:00
Brennan Benson 40ea4ece1a Track Claude models from the installed CLI per host (STA-3330) (#12369)
* feat(native-chat): track Claude models from the installed CLI per host (STA-3330)

The Claude seed no longer pins version labels to aliases that resolve
differently across CLI versions, and the catalog now defines listModels
backed by a one-shot list_models control request over --print stream-json.
Hosts whose CLI predates the request answer with a control error and keep
the seed. Discovery also feeds Source Control AI via the commit-message
spec, and the /model echo detector matches resolved model names.

* fix(native-chat): preserve discovered Claude capabilities

* fix(native-chat): tolerate malformed Claude model entries

* fix(native-chat): discover models in folder workspaces

* fix(native-chat): trust discovered Claude capabilities

* fix(native-chat): remove Claude model fallbacks

* fix(native-chat): keep the Claude model picker rendered

The Claude picker rendered nothing until the per-host `list_models` probe
returned, so it popped in ~1s after mount and never appeared at all when
the probe failed — an old CLI without `list_models`, no `claude` on PATH,
or an older remote runtime whose response omits `catalogOrigin`.

Restore the version-neutral family seed as the starting list; discovery
still replaces it wholesale on success, so a host with a real catalog
never shows an obsolete hardcoded row.

Separately, the tracked model could fall outside the active list: the
terminal header scrape yields family ids (`opus`) while a current CLI
lists `opus[1m]` and no plain `opus`. That blanked the picker trigger and
dropped the model's effort and fast-mode controls. Reconcile the tracked
id into the active list once, so the snapshot, the appliers, and typed
command recording all see a labelled, operable row for it.
2026-08-04 15:47:29 -07:00
Brennan Benson 9ee359550b fix(mobile): make native-chat file links and path citations tappable (STA-3331) (#12364)
* fix(mobile): make native-chat file links and path citations tappable (STA-3331)

- Linkify POSIX absolute paths in chat prose (leading-/ regex alternative;
  URL guard now keys off the char before the matched slash)
- Parse agent-style path:line(:col) citations in prose, code spans, and the
  open flow; line/column ride into the mobile file preview route
- Route non-web markdown hrefs (file: URIs, relative/absolute paths) to the
  file opener instead of silently dropping them; unknown schemes stay dead
- Resolve chat paths against the worktree root, not the terminal's live cwd
- Reuse the terminal tap-to-open flow for chat taps (haptic, preview route,
  tab activation with retries) via a shared identity-stable hook, and toast
  on misses instead of silent no-ops
- Keep snake_case paths whole (intraword underscores are literal text),
  scan bold/italic/strike spans for paths, split trailing punctuation off
  autolinks, and let taps land while the composer keyboard is up

* fix(mobile): harden chat file tap handling

* refactor(chat): share native chat href routing

* fix(mobile): detect files directly under path roots

* fix(mobile): keep inline tokens and dunder paths intact around emphasis

Review follow-ups on the chat file-link work:

- A rejected intraword `_` token left the scan index past its closing
  underscore, so every inline token between two snake_case words was
  swallowed and rendered as literal source — including markdown links,
  which became untappable. Rescan from just past the opening delimiter.
- Treat a path separator as an intraword flank so `src/__init__.py` and
  `a/__tests__/x.ts` stay whole; previously they rendered as bold plus a
  remnant that the new absolute-root pattern turned into a tap on `/x.ts`.
- Bound the `:line(:col)` tail so `src/app.ts:1e3` and `:80%` no longer
  parse a line number, while a cited range still opens its first line.
- Route chat tap failures through the composer banner (toast fallback):
  chat taps happen with the keyboard up, which covers the toast.
- Drop the tap-handler mirror's dep list; the call site rebuilds its
  accessors every render, so it could never skip on a route that
  rerenders per keystroke.

* Revert "fix(mobile): keep inline tokens and dunder paths intact around emphasis"

This reverts commit 308bfaf22b.

* fix(mobile): preserve chat file-link parsing and feedback
2026-08-04 15:46:08 -07:00
847c8c852d fix(agent-status): correlate manual Claude compact hooks (#12332)
Co-authored-by: gatsby74 <166927047+gatsby74@users.noreply.github.com>
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-04 15:03:57 -07:00
Jinjing 999e3a3a6d feat(sidebar): link Linear issues from Edit Worktree Details (#12380)
* feat(sidebar): link Linear issues from Edit Worktree Details

The Issue field only accepted GitHub numbers, so a workspace tracking a
Linear issue had no way to say so from the dialog — the link could only be
set at creation time or through `orca worktree set --linear-issue`.

Replaces the field with one provider-aware row: a chip suffix inside the
input selects GitHub or Linear, and pasting a URL flips the chip to match.
A bare key never steers the provider — Linear and Jira issue keys are
byte-identical in shape, so shape alone cannot decide one.

One issue per workspace. A changed field displaces the other provider's
slot and the row names what Save is about to unlink. GitLab and Jira links
are left alone: the row cannot display them, and nothing else in the UI
could restore one it dropped.

- Folder workspaces read-only (their link is creation-time only)
- Remote runtimes assert the capability before writing or clearing, since
  `worktree.set` parses in strip mode and would silently drop the keys
- `updateWorktreeMeta` now reports failure so the dialog can stay open
  instead of closing over a save that refetch reverted
- Parses are length-bounded — `matchGitHubItemPath` strips trailing
  slashes with an unanchored regex that is quadratic on a large paste

* fix(sidebar): respect one-issue-per-workspace rule conditionally

Only clear displaced issue links when they actually existed, preventing
unnecessary Linear keys in GitHub-only workspaces. Skip comment updates
when unchanged to avoid workspace reordering. Add accessibility to
displacement messages and improve folder workspace error handling.

* fix(sidebar): resolve workspace ambiguity and improve Linear issue linki

The same workspace ID can exist under multiple hosts — the owner index reports
this as ambiguous rather than guessing. Dialog callers now pass their repoId so
lookups are unambiguous. Linear identifiers without an org key are resolved
across all workspaces (not just the active organization). Added race-condition
protection for async issue lookups and better change detection to avoid clearing
work-item titles when re-saving an identifier in different spelling.
2026-08-04 13:42:25 -07:00
5Hyeons eed74724ac fix(i18n): localize automation contextual tour (#12270)
The shared Automation tour copy was rendered without passing through
translate(), and the overlay surface hardcoded its default Next and Done
labels. Copy is keyed off the step id rather than its position, so inserting a
step ahead of them cannot shift the text onto the wrong step.

Co-authored-by: 5Hyeons <ohs2251@naver.com>
2026-08-04 03:56:33 -07:00
Evgenii 439a8c46cf fix(plugins): let language packs translate plugin chrome (#12455)
protectedTranslation refused every language-pack key under
auto.components.settings.plugin*, which caught 104 keys that carry no trust
meaning — section titles, empty states, Refresh, Add path. A 35-path exact
allowlist opens those while consent, provenance, and every *Failed string stay
protected; anything new stays protected until it is added deliberately.

PluginsSettingsSection.experimental is held back from the contributed
allowlist: the "Experimental" chip is a trust badge, which the module's own
boundary comment places out of scope.

Co-authored-by: Evgenii <kumiro@me.com>
2026-08-04 03:56:33 -07:00
Neil e1071f59e9 Say why an update install failed instead of stalling for three minutes (#12224)
Surfaces the real install-failure cause instead of letting a failed elevation stall silently, and keeps the reconnect wait inside its total budget by recomputing the remaining time after each awaited RPC.

Relates to #11906 — this fixes the observability half. The functional half (a .deb/.rpm host cannot elevate and can never self-update) is unchanged, so the issue stays open.
2026-08-04 02:02:53 -07:00
Jinjing 00867f06e2 fix(ssh): handle owner displacement and graceful shutdown (#12367)
* fix(ssh): handle owner displacement and graceful shutdown

SSH connections can reconnect with valid session proof after network loss or
device sleep. When the incumbent owner is still half-open, allow the
reconnecting client to displace it outright rather than wait for socket closure
— a window that may never close. Retain displaced deliveries for the new owner
to rotate. During app shutdown, drain SSH sessions without terminating recovery
operations, and retry pending owner grants in case a replacement commits mid-drain.

* fix(ssh): handle owner displacement and graceful shutdown

Make QuitTeardownStartGate a shared singleton so SSH connects use the
same shutdown fence as the main quit path. Track test-connection probes
to ensure they complete before final teardown. Guard owner displacement
to prevent stale owners from clearing recovery state claimed by newer
owners.

* fix(ssh): fix flaky test sync and add error code safety check

Test was using tick-based Promise.resolve() loops which don't guarantee
the async operation has started. Replace with signal-based synchronization
that waits for the actual lease flush. Also add nullish-coalescing to
error code check to prevent crashes if error is null or undefined.

* fix(ssh): fence reset transport opens during shutdown

* fix(ssh): keep recovery leases stable across reconnects

* fix(ssh): close transports owned by cancelled connect attempts

When a connect is cancelled after its transport has opened, that cancelled
attempt still owns the transport and must close it — otherwise it leaks. Add
disconnectConnection() to close by identity (not by target ID) so a cancelled
attempt closes only the transport it minted, without tearing down its
replacement's live transport. Track priorConnection to detect whether this
attempt opened a new transport or reused an existing one, and close only on
abandonment if this attempt owns the session.

* fix(ssh): fence old owner proofs and close superseded transports

When an owner reconnects with a new proof while an old one is still live,
the old proof is now fenced with SUPERSEDED_ERROR instead of retrying
indefinitely. The relay also closes stale transports to signal that their
recovery generation has been overtaken by a newer one.

This ensures overlapping reconnect scenarios complete with the newest proof
rather than getting blocked by stale recovery attempts.

* fix(relay): re-pin stdin/stdout fds after closing to prevent recycling

When the relay closes stdin/stdout to signal EOF to the SSH peer, the OS
can recycle those fds (0 and 1) for new sockets or files. If Node still
treats process.stdin/stdout as those numbers, subsequent operations
corrupt socket clients and trigger shutdown errors. Re-pin the fds by
opening /dev/null to keep them occupied and prevent recycling.
2026-08-04 01:37:25 -07:00
NeilandOrca b86880a9c1 Keep finished and interrupted agent sessions resumable after sleeping a workspace (#12214)
* fix(agent-sleep): keep finished and interrupted agent sessions resumable after sleeping a workspace

Manual workspace sleep ran a liveness filter over the panes it was about to kill
(isValidManualSleepLiveAgentEntry), then wiped every pre-existing sleeping record
in the worktree. A done, interrupted, typed-into, or >30-min-idle pane therefore
lost its only resume handle and woke as a bare shell. A second filter at wake
discarded any record carrying `interrupted: true`, which also killed interrupted
sessions across an app restart.

Capture now records every resumable pane, normalizing only `updatedAt` and
`interrupted` and preserving the entry's real `state` so a finished pane keeps
its passive record and resumes in place when its tab is opened instead of
spawning a duplicate tab. The wake-side interrupted check is gone. A
legacy-orchestration-worker block is carried onto the replacement record, and
the Pi-compatible promoted checkpoint is no longer overwritten by a re-derived
live record.

Closes #11598

* fix(agent-sleep): keep durable slept records a repeat sleep cannot re-derive

removeSleepingRecordsReplacedByManualWorktreeSleep wiped every record in the
worktree, and only the freshly captured set was merged back. A slept `done`
pane stays passive until its tab is opened, so a second sleep found no live
status row to rebuild its record from and deleted the pane's only
`--resume` handle with nothing written back — the original loss of #11598,
one wake/sleep cycle later.

The wipe now skips a record with no replacement in the new capture set when
it is a durable capture (`origin` `worktree-sleep` or `quit`). Provisional
`live`/legacy checkpoints are still cleared, so an unresumable Pi row does
not survive a sleep it cannot back.

Co-authored-by: Orca <help@stably.ai>

* fix(agent-sleep): keep a slept workspace's finished panes out of the mobile wake fan-out

A manual sleep now records every finished pane, and those passive records fed
wakeSleepingAgentsForWorktreeInBackground step (b), which background-mounts one
tab per passive record. A phone opening a slept 12-tab workspace would have
cold-restored 12 agents at once, undoing the process shedding the sleep was for.

Manual-sleep captures of finished panes carry restoreOnTabOpenOnly; step (b)
skips them and the pane resumes in place when its own tab is opened, which is
what desktop activation already did and what the phone's per-tab mount provides.

Co-authored-by: Orca <help@stably.ai>

* fix(agent-sleep): close the retained-row and shared-claim gaps in slept-session capture

Three narrow holes left by the manual-sleep capture rewrite, all in the same
"a slept session must stay resumable" contract:

- The retained pass captured `retained.entry` verbatim, so a retained row -
  stale by construction, since it exists only after the pane's pty died -
  produced a `working` record that wake then discarded on the >30min staleness
  rule. It now takes the same `updatedAt`/`interrupted` normalization and the
  same `automaticResumeBlockedBy` carry-over as the live pass.
- The retained pass also ran after the `origin: 'live'` promotion loop without
  its guard, so a pane holding both a promoted checkpoint and a retained row
  had the checkpoint (connectionId, transcript identity, active `working`
  class) overwritten by a re-derived passive record.
- The mobile wake filtered `restoreOnTabOpenOnly` records *after*
  canonicalization, so a lazy record sharing a provider-session claim with an
  eligible hibernated record could win the claim, delete the hibernated record
  as a duplicate, and then be skipped - stranding the session with nothing
  mounted.

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-04 01:01:19 -07:00
f6878d660f Show Claude's AskUserQuestion card in desktop Chat when the agent runs on a paired headless server (#12223)
* fix(native-chat): show Claude's AskUserQuestion card when the agent runs on a paired headless host

Three gaps kept the question card off the desktop when the agent ran on a
remote `orca serve` host:

- The `session.tabs` projection reduced HTTP agent-hook rows to identity only,
  hard-coding `state: 'done'` and an empty prompt, so `toolName` and the full
  `interactivePrompt` never left the host. It now publishes the newest fresh
  hook row's status fields, bounded by the same staleness window `agentType`
  uses, excluding `providerSessionOnly` resume rows, and yielding to live
  title evidence unless a question is actually pending.
- Nothing republished `session.tabs` when only a hook row changed, and the
  re-emit carried an unchanged `snapshotVersion` that clients drop on their
  monotonic gate. Material hook transitions and pane/SSH status clears now
  bump the version and schedule a coalesced emit.
- The desktop card resolved only from live status. It now falls back to the
  pending ask in the transcript, matching mobile, so a relay gap can no longer
  leave the composer mounted over a pane parked on a selector.

Closes #11761

Co-authored-by: Orca <help@stably.ai>

* fix(native-chat): date the hook-row recency guard against a real clock

`resolveHookLiveAgentRow` compared a hook `receivedAt` (epoch ms) against
title stamps that are title-observation sequence numbers, so the guard could
never fire — any fresh hook row overrode live title-derived state, and a manual
rename (the one epoch writer) inverted it. Stamp the live OSC title path with
wall-clock ms and compare against that alone.

The regression test fabricated epoch-valued title stamps production never
writes; it now drives the title through `onPtyData`, and a new case pins the
opposite direction (hook row newer than the title wins).

Co-authored-by: Orca <help@stably.ai>

* fix(native-chat): stop an orphaned tool call from pinning a dead question card

extractPendingAsk pairs tool results to calls by a global FIFO (tool_use_id
is dropped at decode time), so one call that never gets a result desyncs the
queue for the rest of the transcript and strands an answered ask as pending.
Real transcripts also hold asks the user escaped and typed past. On desktop
that card replaces the composer, so the pane became unsendable.

Drop in-flight calls at a turn boundary — a user turn or the decoders'
interrupt row — since the turn that owned them is over. Claude's tool-result
turns decode as role 'tool', so normal FIFO resolution is untouched.

Co-authored-by: Orca <help@stably.ai>

* refactor(native-chat): trim the headless AskUserQuestion projection

Reuse rather than restate: the invalidator now takes the shared
`AgentHookEventPayload` instead of a locally redeclared row shape, and the
hook live row is a `Pick<>` of the retained OSC snapshot so one projection
branch consumes either carrier. Fold the immediate/coalesced session-tabs
emit into one method (also drops a redundant re-emit on the
provider-session push). Drop card tests that re-route shared-parser
assertions through React. Isolate pane-status-clear subscribers and prove
the no-republish case by version arithmetic instead of a timed silence.

Co-authored-by: Orca <help@stably.ai>

* test(native-chat): pin the AskUserQuestion card render under real Electron

Why: the 13 parser unit tests pin extraction, but nothing proved a card
actually renders where an inert tool call used to. This spec reproduces the
paired-headless topology from the client side — live status carrying agent
identity and state 'working' but no interactivePrompt/toolName, with the
pending ask present only in the transcript — and fails on main.

Refs #11761

Co-authored-by: Orca <help@stably.ai>

* test(native-chat): drop the unused testInfo parameter

Why: oxlint no-unused-vars fails the lint gate on an unused test parameter.

Co-authored-by: Orca <help@stably.ai>

* test(native-chat): drop leftover proof scaffolding from the ask-card spec

The env-var screenshot label and the fixed 2s settle only existed to make
the pre-fix capture comparable; the card assertion already waits.

Co-authored-by: Orca <help@stably.ai>

* test(runtime): use a truly unresolvable pane key in the hook republish guard

#11203 taught pane lookup to recover a reminted tab id by leaf id, so the old
fixture (new tab id, live leaf id) resolved and bumped the snapshot a second
time once this branch merged with main.

Co-authored-by: Orca <help@stably.ai>

* fix(runtime): refuse a hydrated unconfirmed hook row as live pane status

#12346 landed on main after this branch was cut: a nonterminal row restored from
last-status.json is stamped `restoredUnconfirmed` because its transition may have
fired while no receiver was up, and every freshness gate treats it as never-fresh.
The new headless `live` projection here only checked `receivedAt`, so a restart
inside the 30-minute window would republish the hydrated row — resurrecting the
AskUserQuestion card with no agent left to answer it.

`agentType` still reads those rows: they prove identity, just not liveness.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Neil <nwparker@users.noreply.github.com>
2026-08-04 00:48:58 -07:00
OrcaWinandOrcaWin ed7849eb7b fix(worktrees): stop silently switching existing Windows setup scripts to Git Bash (#12406)
* fix(worktrees): stop silently switching existing Windows setup scripts to Git Bash

#6967 derived the Windows setup-runner shell from `terminalWindowsShell`. On
upgrade, any Windows user whose terminal preference resolved to Git Bash had
their existing `orca.yaml` setup script (and issue command) handed to bash
instead of cmd.exe. Scripts authored against the cmd runner — `copy`, `xcopy`,
`set VAR=value`, `if errorlevel 1`, `%VAR%`, backslash paths — broke with no
migration and no warning, and the failure looked like Orca broke the project.

The conflation is also wrong in the steady state: a terminal preference is
per-user, so two people on the same repo got different interpreters for the
same orca.yaml and no project could write a setup script that worked for all
of its Windows contributors.

The interpreter is now a property of the script, declared the standard way:
a leading `#!` line. Native Windows keeps the historical `.cmd` runner unless
the script declares a POSIX shell, so no existing script changes behavior.
`resolveSetupRunnerShell` keeps its role as the feasibility gate — a bash
runner still requires the terminal to resolve to Git Bash, since the launch
command is typed into that shell and uses MSYS `/c/...` paths.

`buildWindowsRunnerScript` now drops a leading `#!` line rather than `call`ing
it, so a declared-bash script that falls back to cmd (Git Bash missing) fails
on a real setup line instead of aborting on errorlevel at line one.

WSL worktrees, POSIX platforms, and SSH hosts are untouched.

* fix(worktrees): keep the cmd setup runner launchable from a Git Bash pane

Adversarial review of this PR found that pinning the runner format per script
reopened issue #6896 one layer down.

- `WorktreeSetupLaunch.shell` had been redefined to mean "the format the runner
  file was written in". `resolveSetupRunnerCommand` consumes it as "the shell
  that types the launch command", so a Git Bash terminal with a batch setup
  script produced `cmd.exe /c "C:\...\setup-runner.cmd"` typed into a bash pane,
  where MSYS rewrites the `/c` switch into a drive path: cmd opens interactively
  and setup never runs. `shell` is the terminal's family again; the runner file's
  .cmd/.sh extension carries the format, and a batch runner launched from a POSIX
  pane reuses the existing PowerShell ProcessStartInfo launcher.
- The cmd runner dropped a leading `#!` line and ran the rest as batch, so a bash
  script reaching cmd (PowerShell/cmd terminal, or any SSH-to-Windows host) got
  its interpreter-agnostic prefix executed before failing mid-way. It now prints
  why and exits 1 without running anything.
- A `#!` line's option flags were discarded: `#!/usr/bin/env -S bash -euo
  pipefail` lost pipefail because the runner is launched as `bash <path>`. The
  generated posix runner now replays declared flags via `set` and drops the
  duplicate interpreter line.
- Docs cover the per-user setup command in repository hook settings, which goes
  through the same `#!` rule, and describe what the `#!` line does and does not
  select.

Tests: composed launch command for a POSIX pane + cmd runner (hooks, shared
runner command, setup sequencing gate, observed-setup signal), the cmd runner's
shebang refusal, and shebang flag replay. Each fails with the source reverted.

* fix(worktrees): replay only real `set` flags and keep the gate in the pane's shell

Two round-2 review findings:

- `#!/bin/bash -l` replayed `set -l`, which exits 2 and aborted the runner under
  its own `set -e` before a single setup line ran (all platforms). Only the flags
  `set` documents are replayed now; a bare `-o` with no option name is dropped
  instead of dumping the shell-option table.
- The wait-for-setup gate picked its language from the runner file, so a batch
  runner launched from a Git Bash pane got the PowerShell gate while the agent
  startup command was already POSIX-quoted — `Invoke-Expression` cannot parse
  `'\''`. The gate now follows the pane; the runner still launches through the
  ProcessStartInfo launcher, never through bash.

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-03 23:25:50 -07:00
OrcaWinandOrcaWin 141b1f43f6 fix(runtime): keep the listener on loopback for a "This computer only" pairing link (#12405)
* fix(runtime): keep the listener on loopback for a "This computer only" pairing link

The runtime pairing URL handler called ensureNetworkExposure() for every
offer, including one whose advertised address is loopback. Settings ->
"Share this Orca server" offers a "This computer only" radio that pairs
against 127.0.0.1 precisely so nothing is reachable off-host, yet choosing
it rebound the WebSocket listener from 127.0.0.1 to 0.0.0.0 — and the widen
never narrows back, so the runtime stayed exposed to the whole LAN for the
rest of the process after the user picked the option that exists to avoid
exactly that.

Gate the widen on the advertised address: only a non-loopback endpoint (LAN,
Tailscale, custom host) needs a listener reachable off this machine, so those
paths keep widening exactly as STA-2370 intended. A loopback link is already
served by the loopback listener, so it now mints without touching the bind.
Classification reuses the shared pairing-address classifier, which also covers
localhost, ::1 and 127.0.0.0/8 typed into the custom-address field.

Tests: a real OrcaRuntimeRpcServer driven through the IPC handler asserts the
bind host stays 127.0.0.1 after a local link and flips to 0.0.0.0 after a
LAN one, plus handler-level cases for 127.0.0.1 / localhost / ::1.

* fix(runtime): gate the pairing widen on the user's declared reach, not the address shape

Review of #12405 found two ways the loopback fix misbehaved.

1. The guarantee died at the next launch. resolveInitialWebSocketBindHost()
   binds 0.0.0.0 whenever any device has lastSeenAt > 0, and MobileSocketWiring
   stamps that for EVERY authenticated socket — including the local browser
   opening a "This computer only" link. So the runtime was still published on
   every interface, one restart later. Grants now carry the reach they were
   minted for (DeviceEntry.pairingReach, persisted); a this-computer grant no
   longer counts as proof that an off-host client may reconnect. Registries
   written before the field default to network reach, so an already-paired
   phone still finds a wide listener after upgrading. A pending grant that is
   re-advertised for the network widens (never narrows) so its link survives.

2. The widen was gated on the shape of the typed address, which the renderer
   never sent the intent for. A Custom `127.0.0.1:8443` — the documented SSH
   tunnel / reverse proxy field — skipped the widen and produced a dead link,
   while `localhost:8443`, `[::1]:6768` and `ws://127.0.0.1:6768` widened, so
   the same loopback intent was handled three different ways. The renderer now
   sends the declared reach ('this-computer' | 'network') and main gates on it;
   the address is only used as a mismatch guard (a this-computer reach carrying
   an off-host address still widens rather than minting an unreachable link),
   resolved through resolveAdvertisedPairingHostname so every accepted address
   form classifies identically.

Also corrected the ensureNetworkExposure invariant comment: the widen is no
longer confined to the first pairing action, so it can now tear down live
loopback sockets — they reconnect on the reused pinned port.

Tests: reach-form matrix + tunnel/undeclared/mismatch cases in mobile.test.ts,
real-server relaunch bind for both reaches, legacy registry compatibility, the
pending-grant reach upgrade, a live-client port-stability guard, hostname
resolver coverage, and the renderer reach plumbing. Reverting only the source
fails 18 of them.

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-03 23:25:44 -07:00
OrcaWinandOrcaWin 20a2901677 fix(worktree): tell the truth about live PTYs, and offer force for a wedged sweep (#12394)
* fix(worktree): tell the truth about live PTYs, and offer force for a wedged sweep

Two gaps in the #11960 force path:

The delete toast described every unstopped-PTY failure as "could not confirm
every terminal has exited", including the case where verification positively
watched them running. Force Delete proceeds either way, so the user was being
asked to waive a doubt that did not exist while a running agent's uncommitted
work died with it. The live verdict now gets copy that says so.

A sweep that rejects before any per-PTY verdict exists (wedged daemon, dropped
SSH channel) fails with a teardown-timeout message that the force classifier did
not recognise, so no Force Delete button appeared — the exact dead end #11960
set out to remove. That error now carries the shared prefix and classifies.

* fix(worktree): close the sweep-rejection wedge and stop racing the delete

Review of #12394 found the fix covered only half the wedge it named, and
routed users into a force path whose own safety comment was untrue.

1. Only the outer deadline was classifiable. When a provider *rejects* the
   sweep — dropped SSH channel, erroring daemon — settleBeforeDeadline
   rejects with the provider's original error, which carries no marker, so
   classifyWorktreeForceDeleteReason still returned null and no Force Delete
   button rendered. That is the exact case #11960 named. A rejected sweep on
   the destructive path is now reworded through the existing unstopped-PTY
   prefix (provider text preserved, original kept as `cause`), so old and new
   clients alike classify it as 'unstopped-pty'.

2. Force could delete files while a sweep was still running. The deadline
   rejects without cancelling run(), so allSettled resolved with shutdown()
   still in flight — by construction the deadline error can only fire while
   something is in flight. Force then deleted the directory a live PTY still
   held open (EBUSY / half-delete on Windows and WSL). Sweeps are now tracked
   so the forced path waits for the abandoned work, bounded by a 2s grace;
   force never wedges, and when the grace expires the warning says handles may
   outlive the delete instead of implying the sweep finished.

3. The toast test named for the classifier passed the reason in as a literal,
   so it never exercised it. It now derives the reason exactly as the store
   does, and fails against main.

4. Added the missing unstoppedPtyLive key to the English catalog.

5. isProvenLivePtyRemovalError anchored the 'still live:' marker to the detail
   separator, so a worktree path can no longer spell out a live verdict and
   flip the toast to the destructive copy.

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-03 23:25:21 -07:00
e59a319ffe fix(sidebar): keep each project's entry-point workspace visible under "Hide sleeping" (#12257)
"Hide sleeping" swept each project's main workspace out of the sidebar as soon as
it had no live PTY, browser tab or agent — even with "Hide default branch" off.
For a project whose only row is that workspace (a folder workspace, a fresh
clone, a detached-HEAD main), the entire project vanished with no in-place way
back.

Adds a shared `isSleepingSweepExemptWorkspace` predicate keyed on
`isMainWorktree` rather than the branch name, so folder workspaces (no branch),
detached-HEAD mains, and SSH rows whose head/branch are blanked while a provider
is disconnected all stay put. Wired into `computeVisibleWorktreeIds` (sidebar,
Cmd+1-9, workspace board), the jump palette's duplicate inline pass, and mobile's
`filterWorktrees`.

Ships default-on with an escape hatch: a persisted
`alwaysShowDefaultBranchWorkspace` setting surfaced as "Except default branch"
under "Hide sleeping". Explicit "Hide default branch" still wins, since it
filters before the sleeping sweep.

Mobile reads the setting but never writes it back, so a desktop opt-out can't be
clobbered by a filter tap before the ui.get roundtrip lands.

Combines the two PRs open against #8873. #8966's exempt set is a strict subset of
this one, so its production diff was subsumed rather than ported; its jump-palette
render harness and e2e spec were carried over, and are the only such coverage here.

Fixes #8873
Closes #8966

Co-authored-by: Rod Boev <rod.boev@gmail.com>
Co-authored-by: Orca <help@stably.ai>
2026-08-03 22:55:17 -07:00
NeilandOrca 0927b9c156 fix(gitlab): load pipeline job traces in the Checks side panel (#7732) (#12266)
* test(repro): demonstrate #7732 GitLab pipeline job details never load in Checks panel

Co-authored-by: Orca <help@stably.ai>

* fix(gitlab): load pipeline job traces in the Checks side panel (#7732)

Expanding a GitLab pipeline job in the Checks panel always showed
"No inline details are available for this check.": the mapper dropped the
numeric job id, `PRCheckDetail` had nowhere to carry it, and every consumer
called the GitHub check-runs API, which returns null for a GitLab job.

- carry `gitlabJobId` on `PRCheckDetail` and add the `gitlab-job:` branch to
  all three identity ladders (panel rows, editor tabs, fix-prompt keys) so
  same-stage jobs with no web_url stop colliding
- add a runtime-routed trace client so SSH/remote workspaces work, not just
  local IPC, and thread the MR's `projectRef` for fork pipelines
- bound the trace in main via the existing `sliceCheckLogTail` (now shared,
  not GitHub-only) so a multi-megabyte CI log never crosses the 1 MB
  transport frame cap; strip ANSI/section markers up to the CR only, which
  keeps each section's visible header and command echo
- render the excerpt inline instead of "Log tail available in full details."
- feed GitLab traces to "Fix with AI", which previously sent bare check names
- skip the fetch for jobs that cannot have a trace (created/manual/skipped)
  so GitLab's 404 does not replace the benign empty state, and re-arm a
  failed load when the job's state changes since the panel has no retry

Co-authored-by: Orca <help@stably.ai>

* fix(gitlab): treat a missing job log as an empty log, not an error (#7732)

Round-1 review follow-up.

- a job canceled before it started (or whose log was erased/expired) is
  `completed`/`cancelled`, so the panel fetched its trace, GitLab answered 404,
  and `classifyGlabError`'s issue-edit copy ("Issue not found — it may have been
  deleted.") landed verbatim on the auto-expanded check row; main now maps that
  404 to an empty trace so the row keeps its benign empty state
- keep a missing project a real error (GitLab masks unauthorized projects as
  404) and add `classifyJobLogError` so 403/unknown failures stop borrowing
  issue-edit wording on a job-log read
- broaden the empty-log copy in all five catalogs: it now covers erased and
  expired logs, not only jobs that never ran
- e2e: derive the repro screenshot dir from `process.cwd()` (or an env
  override) instead of a hardcoded POSIX path to a throwaway worktree
- bound the raw trace before the ANSI/section passes so a multi-megabyte log
  is not scanned in full on the main-process event loop
- drop the redundant `if (repo)` in `handleFixChecksWithAI` and the now-dead
  "Log tail available in full details." catalog entry

Co-authored-by: Orca <help@stably.ai>

* fix(gitlab): address review — project ref on reload, retry re-arm, IPC timeout

- Carry the MR's GitLab project ref on the check-details tab so reloading a
  fork/cross-project job tab fetches the trace from the pipeline's own project.
- Re-arm the sidebar retry when a details load resolves to null, not only when
  it throws; a detail-less row otherwise never retried after the job moved on.
- Bound the local `gl.jobTrace` IPC call with the same 30s timeout the runtime
  RPC path uses — glab runs without a subprocess timeout in main.
- Document that the trace 404 -> empty-log mapping is deliberately broad.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-03 22:48:37 -07:00
Jihwan KimandOrcaWin 5bd2f59d29 fix(runtime): open files from sibling workspaces (#11369)
* feat(runtime): match files to workspace owners

* fix(runtime): resolve terminal paths through sibling workspaces

* fix(editor): route restored sibling workspace files

* fix remote sibling file ownership routing

* fix(editor): migrate restored sibling file owners

* fix(editor): revalidate restored owner activation

* docs(review): record PR 11369 correction evidence

* fix(editor): reject collision before activation prep

* docs(review): record PR 11369 final correction

* fix(editor): retain projected reconciliation narrowing

* chore(review): keep verification artifacts out of PR

* fix(editor): harden restored owner migration

* fix(runtime): resolve workspace root terminal paths

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-03 21:41:58 -07:00
NeilandOrca a6b14eb04c fix(terminal): reset stale mouse tracking on cold restore (#12101); stop OSC color-reply echo leak in POSIX agent panes (#12112) (#12202)
* fix(terminal): reset stale mouse tracking on cold restore (#12101); stop OSC color-reply echo leak in POSIX agent panes (#12112)

#12101: a force-killed TUI never emits its DECRST reset, so its armed mouse
mode is latched into the on-disk checkpoint and re-derived into the
replacement process's emulator via the cold-restore history seed -- through
both rehydrateSequences and SerializeAddon's own mode trailer. The revived
bare shell then echoed SGR motion reports at the prompt. Seed a
RESET_MOUSE_REPORTING segment after the snapshot (before the torn escape
tail), only when there is real recovered content so the empty-array
"nothing to recover" sentinel survives.

#12112: agent panes arm a main-side PtyStartupIngress that answered opencode's
startup OSC 10/11 queries synchronously inside node-pty's onData, while the
POSIX tty still had ECHO on. The line discipline echoed Orca's own reply back
out as visible text. Echo suppression existed but was gated on windows-conpty.
Add PtyStartupReplyDelivery: POSIX defers the write off the query's turn and
recognizes its own echo anywhere in a span (bounded, non-destructive); ConPTY
keeps its synchronous write; windows-wsl is byte-identical to before.

Fixes #12101

Co-authored-by: Orca <help@stably.ai>

* fix(terminal): read the slave's ECHO bit before answering a color query

The startup color reply was written into a PTY still in cooked mode, so the
line discipline echoed it back as visible junk (#12112). Whether that will
happen is readable state on the slave rather than something to infer from
returning bytes, so the reply now waits until the ECHO bit is observably
clear instead of guessing at echo shapes.

Two echo sources exist and only one is readable. A `quiet` verdict proves
the kernel will not echo, so it retires the caret projection; readline
echoes a master write in software with the tty already raw, so that
projection stays armed on every path. Scoping `quiet` narrowly is the whole
correctness argument here: reading it as "no suppression needed"
reintroduces the bug at a plain shell prompt.

Polling is bounded by a wall-clock budget rather than an attempt count,
because each probe is a subprocess and a multi-pane restore serializes them
on fork. Withholding measures flat at ~210ms from 1 to 100 panes.

Also resets a cold-restored pane's mouse reporting (#12101). The armed mode
is re-derived from the dead process's own persisted bytes through two
channels, so the daemon seeds a reset into recovered history and the
renderer stops trusting a persisted "live agent" signal after a cold
restore. The reset literals move to one shared profile module.

Fixes #12101
Fixes #12112

Co-authored-by: Orca <help@stably.ai>

* test(terminal): pin the cold-restore reset on the spawn-adopted reattach path

A spawn can be answered with an adopted session, which reaches the reattach
handler by a door that skips the restored-session path. Pin that the cold-restore
signal survives it, so #12101's junk cannot come back through it.

Co-authored-by: Orca <help@stably.ai>

* test(terminal): note why the adopted-reattach snapshot leaves the cursor visible

Co-authored-by: Orca <help@stably.ai>

* fix(terminal): harden startup reply delivery

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-03 21:00:57 -07:00
Brennan Benson 49dc113a0f Fix terminal corruption after restored snapshot replay (#12363)
* fix(terminal): preserve restored snapshot fidelity

* test(terminal): align legacy history handoff snapshot expectation

* fix(terminal): keep legacy snapshot panes mounted

* fix(terminal): refresh snapshot capability after startup

* fix(terminal): refresh snapshot capability in degraded startup

* fix(terminal): await snapshot provider authority
2026-08-03 20:00:28 -07:00
Brennan Benson 194e1a8d4d fix(persistence): make the renderer unload checkpoint durably flush before reporting success (#12387)
The sync before-unload checkpoint staged renderer state and then queued
store.flushPendingAsync() fire-and-forget, so reload/restart/update paths
navigated while the staged session, scrollback and UI state were still
only in memory. Quit is covered by the will-quit flush barrier; those
paths were not.

Keep staging synchronous (no sync durable writes), but record the flush
outcome and expose it on app:await-before-unload-checkpoint. Restart,
updater install and lazy-chunk recovery reload now join that write before
navigating and abort the attempt when it fails or outlives a 20s deadline.
2026-08-03 19:18:40 -07:00
OrcaWinandBrennan Benson ce8b778d31 perf(runtime): withhold unchanged mobile snapshots from the graph payload (#12245)
* perf(runtime): withhold unchanged mobile snapshots from the graph payload

Every graph sync structured-cloned all 222 worktree snapshots to main even when
none had changed: 374 KB and ~5 ms per clone, paid twice because Electron clones
on serialize and again on deserialize. That transport cost — not the renderer
rebuild — is the bulk of a publication.

The renderer now sends only the snapshots main has not acknowledged and names
the rest in unchangedMobileSessionWorktrees. Detection is object identity, not a
deep compare: an unchanged worktree already returns its cached snapshot object.
Main seeds nextWorktrees from that list so its prune keeps withheld worktrees
live instead of removing them.

The call itself is unconditional. syncWindowGraph is not a one-way publish — its
return value is the only channel carrying agentOrchestrationByPaneKey to the
renderer, and the handler adopts pre-allocated handles, merges detached leaves,
refreshes writable flags, and drains graph-sync callbacks on every sync. Skipping
it would starve all of that.

Two failure modes are closed explicitly. The memo advances only after main
acknowledges, so a publication that throws is resent in full rather than
silently withheld forever. And a worktree main dropped on its own — worktree
metadata removal — comes back in mobileSessionResyncWorktrees, which also clears
the accepted-revision record so the republish is not rejected as a no-op.

Unchanged republish at 222 worktrees / 787 tabs: 374 KB to 3.4 KB, 5.08 ms to
0.02 ms per clone. One changed worktree: 5.3 KB.

* fix(runtime): resync stale withheld mobile snapshots

* fix(runtime): align accepted mobile snapshot membership

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-03 18:11:30 -07:00
Jinjing 637c7e94c9 Add SSH config host picker to add-host dialog (#12334)
* feat(ssh): add SSH config host picker for add-host form

Users can now click 'Fill from ~/.ssh/config…' to browse available SSH
config hosts in a picker, select one, and have the form automatically
prefill with resolved connection details (hostname, port, username, auth).

Previously, an 'import' button provided bulk sync on this form—confusing
and unhelpful when everything was already synced. That action is now
available as a secondary 'Add all' option in the picker.

* fix(ssh): import filter preservation and label fallback

- Reuse search loader on import completion to preserve active filter inside generation guard
- Fall back to hostname when manual host has no label, not empty string
- Make alias duplicate detection case-insensitive to match config picker behavior
- Validate host availability when restoring project group selection
- Add aria-selected attribute to picker options for accessibility

* fix(ssh): harden config picker import, alias folding, and host targeting

Review findings on the ~/.ssh/config picker + bulk add:

- Guard config-host resolution with a generation counter so a late resolve
  cannot overwrite a later pick or a form the user backed out of; freeze the
  other rows while a pick resolves.
- Stop "Add all N" from re-adopting deleted hosts — it now imports without
  reAdopt, matching the new-host count it advertises. Settings → Import keeps
  the explicit re-adopt path.
- Fold SSH aliases through a shared normalizeSshConfigAlias for import
  ownership, delete tombstones, reclaim, picker search, and the save-time
  duplicate check, which now occupies configHost *and* label like the picker.
- Persist GSSAPIAuthentication only when a parsed Host entry asks for it, not
  when `ssh -G` merely echoes the /etc/ssh system default.
- Fail closed with unavailable/setup-not-found when an explicit
  projectHostSetupId names a non-actionable host instead of silently creating
  the workspace on a sibling host.
- Cache the parsed config for the picker session (refresh on open/retry) so
  filter keystrokes no longer reparse and Include-expand the file, keep the
  filter usable during loads, add a Retry on load errors, explain an empty
  Identity file after a config fill, and drop the always-false aria-selected.

* refactor(ssh): centralize host result limit and extract folder group val

Move SSH_CONFIG_HOST_RESULT_LIMIT to shared types so the renderer's limit message
cannot drift from the host's query limit. Extract findActionableFolderProjectGroup
to avoid repeating the folder-host-availability check across the composer hook.

* fix(ssh): pass -F to ssh -G when HOME differs from passwd home

In E2E tests and sandboxes, isolated HOME can differ from the system
passwd home. OpenSSH resolves the default config via getpwuid (passwd),
while Node's loadUserSshConfig uses os.homedir() (HOME-aware). Pass -F
to explicitly specify the config path when they diverge, so ssh -G and
the picker resolve the same file.

* fix(ssh): verify config host exists before resolving with ssh -G

When a user edits ~/.ssh/config and removes a host, the import picker
should not fall back to ssh -G's echoed response (which treats any alias
as valid). Check the reloaded config file before resolving.

- Force reload config on each resolve to catch user edits post-open
- Reject aliases not in the current config before calling ssh -G
- Add test for deleted alias edge case
- Fix workspace-target fallback to honor explicit host selection

* fix(ssh): let tombstoned aliases be re-picked in the config picker

Allow users to reclaim a deleted SSH host by re-picking it from ~/.ssh/config. Tombstoned aliases now appear in the picker with a "Removed from Orca" badge and remain pickable, but don't count toward "Add all" operations — ensuring passive import never resurrects a deleted alias while still giving the user a recovery path.
2026-08-03 17:32:13 -07:00
Brennan Benson 9ec4907cfb fix(agent-status): restore hydrated nonterminal statuses as unconfirmed (#12346)
* fix(agent-status): restore hydrated nonterminal statuses as unconfirmed

A hook transition that fires while Electron is down has no receiver and is
discarded, so last-status.json can restore a stale 'working' as confirmed
truth for up to the 7-day hydrate TTL. Stamp hydrated nonterminal rows with
restoredUnconfirmed, carry it through both IPC paths, and treat such rows as
never-fresh in the shared and renderer freshness gates so the sidebar,
worktree.ps, and the raw snapshot all present the same degraded semantics.
Terminal states restore as-is; any accepted live event clears the flag; the
flag itself is never persisted. Interrupt/question inference refuses to
fabricate transitions onto unconfirmed rows.

* fix(agent-status): shed unconfirmed marker when the liveness sweep verifies done

The restored-subagent reaper's reconciled entry spread carried
restoredUnconfirmed onto a process-probe-verified 'done', making freshness
gates suppress a legitimate completion. Keep the marker only while the
reconciled state stays nonterminal.

* fix(agent-status): let live evidence replace hydrated rows

* fix(agent-status): keep restored rows degraded

Sort accepted live evidence after hydrated rows even across wall-clock rollback. Let unconfirmed rows own their preserved pane titles without asserting live state, while retaining independently live sibling evidence.

* fix(agent-status): suppress unmapped restored titles

Treat a single runtime title as covered by the single restored hook row while layout identity is unavailable. Preserve ordinary age-stale fallback and mapped sibling-pane evidence.
2026-08-03 17:22:43 -07:00
Brennan BensonandOrcaWin f4b2b782b5 feat(orchestration): coordinator-driven release of settled worker terminals (STA-905) (#12355)
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-03 17:17:26 -07:00
Brennan Benson 13f033f091 chore(daemon): disambiguate audit observations (#12343)
* chore(daemon): disambiguate audit observations

* fix(daemon): reject future audit protocol roles

* fix(telemetry): protect daemon audit observations
2026-08-03 16:28:49 -07:00
JinjingandOrca d7fe9d6bcc fix(ai-vault): support session scanning in SSH worktrees (#11004)
* fix(ai-vault): support session scanning in SSH worktrees

Add relay-native aiVault.listSessions scanning that discovers agent
sessions on SSH hosts. Includes fallback to filesystem crawl for
legacy relays, full cancellation support, result validation, and
scan coalescing to reduce redundant work.

* fix(ai-vault): scan sessions in SSH worktrees with coordinated cancellat

- Extract batching logic to `mapRemoteScanBatches` for reuse and proper cancellation checkpoints
- Move `AiVaultScanCoordinator` from relay to main to handle concurrent same-key requests with individual cancellation signals
- Report scope path truncation consistently across relay and SSH fallback paths
- Gracefully degrade relay handler on unsupported platforms instead of aborting startup
- Refactor issue display to separate blocking errors, scope notices, and skipped transcript counts

* fix(ai-vault): stabilize SSH session scan CI

Swallow async WSL relay stdin EPIPE so the live hook-relay shard no longer
fails after all tests pass. Merge main, resolve scan/relay conflicts, and
align cancellation/host-issue reporting with IPC expectations.

* fix(ai-vault): harden session scan cancellation, relay timeouts, and preemption

Thread the abort signal through every scan and parse path so superseded or
cancelled scans stop promptly instead of parsing every remaining transcript
for a caller that already left.  Replace the fragile message-text relay
timeout check with a typed error code so unrelated errors carrying the
phrase "timed out after" no longer suppress the filesystem fallback.  Fix
scan coordinator preemption so a forced Refresh in one window no longer
re-enters as a spurious cancellation in another.  Add a host-leg cache for
the all-hosts view and cap filesystem concurrency so a single slow remote
home cannot stall the whole merge.

Co-authored-by: Orca <help@stably.ai>

* fix(ai-vault): use stable React keys for scan issue banners

Drop array-index keys so react-doctor/no-array-index-as-key passes.
Uniqueness comes from host, kind, agent, path, and message.

* fix(ai-vault): SSH session scanning with configurable depth limits

Implement depth-aware caching and proper scan boundaries to make SSH session
scanning reliable in worktrees. Users can now select between faster (250
sessions) and comprehensive (unlimited) history scans. The scanner:
- Deduplicates scans across relay, host leg, runtime, and renderer layers
- Reuses larger scans to serve smaller depth requests
- Properly bounds in-scope discovery per-limit
- Fixes timeout enforcement when SSH providers ignore abort signals

* Move sessionLimit ref update to useLayoutEffect

Keep render pure for React Doctor by deferring ref updates to
a layout effect, which still executes before render-dependent
effects that consume the ref.

* fix(adhoc): stamp version prefix from main, not the feature branch

Adhoc builds check out arbitrary refs whose package.json often lags
version bumps (e.g. 1.4.165-rc.0 while main is 1.4.168-rc.1). Hourly
always builds main so it already tracks the product line; adhoc now
resolves the base version from origin/main (or ORCA_ADHOC_BASE_VERSION)
so branch builds share that prefix.

* Revert "fix(adhoc): stamp version prefix from main, not the feature branch"

This reverts commit a26a18eb3fd83f7e7d2db9a6a7c3e02e0f79089a.

* fix(ai-vault): fix scoped backfill and coordinator race conditions

Resolve race where the last waiter leaving could abort an already-settled scan (add `settled` flag). Redesign scoped session backfill to keep searching through newer files until the scope reaches its requested session quota instead of stopping at the candidate limit; out-of-scope files no longer consume the scope budget. Centralize scan limit normalization and fix error classification for cancelled scans using the proper helper instead of checking Error.name. Disambiguate cache keys using JSON and add cancellation check after scope discovery phase.

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-03 16:17:00 -07:00
OrcaWinandOrcaWin e25381cdd3 fix(relay): tolerate cell clock skew in pairing invite expiry validation (#12340)
The cell stamps invite expiry at exactly now+10min from its own clock while
the desktop rejected anything past now+10min from the local clock with zero
tolerance, so any cell clock ahead of the machine by more than network
transit made every Relay pairing code fail with an opaque toast. Same
defect class as the host-proof freshness incident; same 30s leeway.

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-03 13:15:30 -07:00
OrcaWin cd68a8b00c fix: preserve live agent PTYs through graph hydration (#11789) 2026-08-03 11:11:14 -07:00