Commit Graph
2191 Commits
Author SHA1 Message Date
Jinwoo Hong 6397668271 Add manual artifact sharing from HTML and Markdown views (#13369) 2026-08-09 16:11:16 -07:00
plaonnandBrennan Benson e20845eedb fix(cursor): accept BOM-prefixed hook JSON (#12652)
* fix(cursor): accept BOM-prefixed hook JSON

* test(cursor): pin the hook BOM allowance to one leading U+FEFF

Document why the BOM strip exists and cover the narrowness the fix
claims: a double BOM, a whitespace-then-BOM prefix, and a BOM inside
the JSON body are all still rejected.

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-09 15:23:55 -07:00
EvgeniiandBrennan Benson 63af0bdd24 feat(agents): add Prime Agent as a supported TUI agent with session history (#12935)
* feat(agents): add Prime Agent as a supported TUI agent with session history

Wire Prime Intellect's prime-agent CLI (a Pi fork) into the desktop and
mobile agent catalogs following the Trae registration pattern, and into
the Agent Session History browser following the OMP pattern:

- types.ts, tui-agent-config.ts: register 'prime-agent' with argv prompt
  injection behind a `--` separator (its own help documents `--` as
  "treat all following arguments as messages"; without it, prompts
  starting with `help`/`agents`/`-…` dispatch as subcommands or
  flags), plus csi-u Shift+Enter encoding matching the Pi TUI it embeds.
- agent-kind.ts, telemetry-events.ts, agent-status-types.ts,
  agent-type-label.ts, tui-agent-display-names.ts,
  tui-agent-selection.ts, skills-cli-agent-keys.ts: standard per-agent
  registrations.
- agent-headless-command.ts: `-p/--print` one-shot runs share the
  print-mode matcher with Claude/Trae so they are not mistaken for live
  interactive panes.
- agent-process-recognition: the npm shim launches a generic bundled
  cli.js, so only the exact package path is an authoritative identity
  (same as Pi and cursor-agent). The three per-agent regex branches are
  now one table in agent-node-entrypoint-identities.ts — the module was
  at its max-lines budget and a table makes the next agent one entry.
- AI Vault: sessions are Pi's message-graph JSONL under
  ~/.prime/agent/sessions (override PRIME_AGENT_CODING_AGENT_DIR —
  Prime Agent brands Pi's env contract instead of sharing
  PI_CODING_AGENT_DIR); parsed by the shared message-graph parser with
  incremental append-resume; discovered locally, in WSL homes, and over
  remote SSH; resumes by absolute transcript path
  (`prime-agent --resume <path>`) like OMP, with session-id fallback.
- skill-discovery-sources.ts: ~/.prime/agent/skills home source.
- Catalog, i18n (en/es/ja/ko/zh), mobile registries, and a bundled
  64x64 favicon (required by mobile's offline-icon invariant).

Scanner-test fixtures for OMP and Prime Agent move into
session-scanner-test-fixtures.ts and the incremental fixture into its
own module, keeping every touched file inside its max-lines budget
without ratchet bumps.

* fix(ai-vault): map custom Prime Agent roots to their sessions child

PRIME_AGENT_CODING_AGENT_DIR is consumed verbatim by the CLI as its agent
config dir, with transcripts always in <agentDir>/sessions — unlike
PI_CODING_AGENT_DIR's <home>/agent/sessions shape the shared normalizer
models. A custom root with a non-special basename (or a `.prime` leaf)
was therefore scanned as-is instead of its sessions child. Dedicated
normalizePrimeAgentSessionsDir appends `sessions` to every configured
root, taking only an explicit `.../sessions` path as-is; the shared
Pi/OMP normalizer drops the `.prime` widening it no longer needs.

Raised in review on #12935.

* fix(ai-vault): guard degenerate Prime Agent roots and cover the remote source

normalizePrimeAgentSessionsDir stripped a filesystem-root value ('/' or '//')
to '', which then joined into the relative root 'sessions' and would walk the
main-process cwd. session-scanner-roots.ts already carries this guard for the
OMP variant; apply the same fallback here.

The remote SSH source had no test: deleting jsonlSource('prime-agent', ...)
left the suite green, unlike the local path which is pinned by the
AI_VAULT_AGENTS exhaustiveness assertion in session-scanner.test.ts. Add a
case that fixes the .prime/agent/sessions root segments, the .jsonl
extension, and parser routing.

Raised in review on #12935.

* fix(ai-vault): honor Prime Agent's sessions-root env and non-interactive modes

Verified against upstream PrimeIntellect-ai/prime-agent source rather than
inferred from the CLI's help text.

config.ts getSessionsDir() reads PRIME_AGENT_SESSION_DIR (and its legacy
PRIME_AGENT_CODING_AGENT_SESSION_DIR alias) ahead of the agent dir and uses it
verbatim; setting either left the vault silently empty. It also appends
`sessions` to the agent dir unconditionally, with no basename escape hatch, so
PRIME_AGENT_CODING_AGENT_DIR=/data/sessions writes to /data/sessions/sessions
while Orca scanned /data/sessions. getAgentDir() and the session-dir override
both run through expandTildePath, so a `~` value set outside a shell resolves.

cli/args.ts also spells the non-interactive runs `--mode json|rpc|acp|daemon`,
which the shared print-mode matcher does not know, so those panes were counted
as live interactive agents and the paste-submit path would write user text into
a JSON-RPC/ACP stream. Match upstream exactly: only the space-separated form,
since `--mode=json` is not parsed by the CLI and does start the TUI.

Raised in review on #12935.

* fix(ai-vault): keep Prime Agent roots absolute and remote segments posix

Two holes in the previous commit.

The degenerate-root guard only rejected pure-separator values, so a relative
env value still resolved against the main-process cwd:
PRIME_AGENT_CODING_AGENT_DIR='.' scanned '<cwd>/sessions' and, worse,
PRIME_AGENT_SESSION_DIR='.' scanned the cwd itself. Require an absolute path in
both branches and fall back to the default otherwise.

remotePrimeAgentSessionsSegments() built its segments with the local-platform
join, so on a Windows client scanning a posix SSH host it produced
'\.prime\agent\sessions' and split('/') collapsed it to one bogus segment —
remote discovery would have found nothing. Remote roots are posix regardless of
client platform, so keep them literal. Pi and OMP are unaffected: their
normalizer returns a '.../sessions' input unchanged and never joins.

Raised in review on #12935.

* test(ai-vault): pin Windows drive roots to the Prime Agent default fallback

'C:\' and 'C:/' strip to the drive-relative 'C:', which isAbsolute
rejects on every platform — assert they land in the default fallback so
a looser truthiness check can't reintroduce a 'C:sessions' scan root.

Raised in review on #12935.

* test(ai-vault): pin the drive-relative root form and state what the posix runner can assert

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-09 14:50:02 -07:00
Jinjing 383665ebb8 Taskpage pure extract (#13367)
* refactor(task-page): extract pure task-kind, jira, and pagination helpers

Moves seven closed sets of pure helpers out of TaskPage.tsx (13485 -> 13308
lines) into domain-named sibling modules. Function bodies are byte-identical
cut/paste; the only production edits in TaskPage.tsx are the removed blocks
and the new import statements.

- task-page-github-task-kind.ts: isPRFocusedTaskView, normalizeGitHubTaskPreset,
  getGitHubTaskKind, getDefaultPresetForGitHubTaskKind, scopeGitHubTaskSearch
- task-page-jira-create-fields.ts: the Jira create-field visibility, allowed-value,
  and payload builders
- task-page-jira-project-selection.ts: getJiraProjectSelectionKey,
  compareJiraProjectsByDisplayLabel
- task-page-jira-status-tone.ts: getJiraStatusTone
- task-page-pr-delta-summary.ts: formatPRDelta
- task-page-pagination-page-numbers.ts: getPageNumbers
- task-page-string-set-equality.ts: areStringSetsEqual

Each module gets a characterization test suite that pins current behavior,
including the quirks (case-sensitive matching, truthiness-based option payload
fallbacks, allowedValues winning over schema type). Quirks are documented, not
fixed.

github-enterprise-slug-routing-boundary.test.ts anchored its source-text
sections on `function formatPRDelta` and `function getPageNumbers`; both moved,
so the sentinels advance to the next declarations. The bounded sections and
their assertions are unchanged.

No intentional behavior change.

* refactor(task-page): clarify helper logic and consolidate tests

- Add comment explaining quoted-form parsing in GitHub task scope
- Refine test descriptions and characterization comments for clarity
- Add test case for quoted `is:"issue"` form in task scoping
- Consolidate test files as part of pure-function extraction

* improvements

* Remove obvious comment from getPageNumbers

The function name and implementation are self-documenting; the comment restates what the code already expresses clearly.
2026-08-09 14:45:34 -07:00
JinjingandOrca 3ec48a74d5 Gate artifact publishing behind off-by-default capability (#13368)
* fix(artifacts): gate agent artifact publishing behind an off-by-default capability

Public artifact sharing was reachable by any agent through `orca artifacts
share`: the Artifacts settings toggle only controlled sidebar visibility, and
nothing in the main process checked a capability before minting a public URL.

Add `artifactSharingEnabled` (default off) and enforce it in
ArtifactCloudService.share/update — before auth, network, or the share-record
write — so the CLI, relay-forwarded remote CLI, and IPC paths are all denied.
The denial carries a stable `artifact_sharing_disabled` code plus next steps
through the RPC error allowlist, so the CLI prints actionable guidance.

list, unshare, and delete stay ungated: turning publishing off must not strand
already-published links. The capability is absent from the `settings.update`
RPC schema, so an agent cannot grant it to itself — only the desktop UI can.

Co-authored-by: Orca <help@stably.ai>

* fix(artifacts): gate agent artifact publishing behind an off-by-default

Publishing is blocked until enabled in Settings → Artifacts. CLI preflights the capability before reading files to avoid unnecessary uploads. RPC surface rejects capability grants so callers cannot self-grant. UI shows opt-in workflow and recovery path when publishing is off. Web clients mirror the host's setting read-only.

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-09 13:19:08 -07:00
Jinjing d3cb02f6a9 fix(grok): carry reasoning effort for models outside the seed catalog (#13365)
Launch resolves options from the static seed, so a discovered model id had no
options and silently dropped --reasoning-effort. unknownModelOptions keeps the
effort menu for those ids. Leave the multi-host launch gate unchanged.
2026-08-09 12:31:47 -07:00
Jinwoo HongandJinwoo-H 394e4bf1c0 Polish Artifacts management UI (#13356)
* refactor(artifacts): polish artifact management UI

* fix(artifacts): address UI polish review

---------

Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
2026-08-09 12:11:33 -07:00
Neilandbbingz 34f2a62cda fix(pty): stop color-scheme 997 replies from painting cooked prompts (#13309)
Route cooked-echo-risk terminal replies through bounded echo-safe delivery across local, daemon, and SSH relay PTYs. Preserve repeated valid replies, bypass the daemon startup input gate, and keep ordinary input plus latency-critical replies on their existing paths.

Closes #13137

Co-authored-by: bbingz <zzb@gxsmjx.com>
2026-08-09 01:45:15 -07:00
Neil 9f7522dee9 fix(workspace): stabilize smart entry keyboard flow (#13319) 2026-08-09 01:36:32 -07:00
NeilandOrca 970696a008 fix(sidebar): show Cursor rows and stop a stray "claude" title hijacking OpenCode (#12466)
* fix(sidebar): show Cursor rows and stop a stray "claude" title hijacking OpenCode

Two defects in the same title-resolution path.

**#10258** — Cursor's only native OSC title is the literal `cursor agent`, which both title trackers dropped unconditionally. A hookless Cursor pane therefore had neither a status entry nor any title carrying Cursor identity, so the worktree card showed nothing at all.

**#8940** — two owner-blind paths let an incidental `claude` token anywhere in an OpenCode session or task title outrank the pane's known owner, so the tab icon and sidebar row flipped to Claude Code.

#10258: let the literal through exactly once as identity, so a restored or mobile tab keeps its Cursor row instead of vanishing. #8940: require an *identity frame* — after stripping status decoration the title must PRESENT Claude, not merely mention it — before a Claude title may reclaim a pane from its prior identity, and make the sidebar row builder owner-aware.

> These two are in one PR because they share the `ownerAgentType` plumbing through `buildTitleDerivedAgentRow` — split apart, neither half compiles on its own.

Fixes #10258
Fixes #8940

Co-authored-by: Orca <help@stably.ai>

* test(e2e): add recordable proof for sidebar-agent-row-identity

Fails on origin/main, passes on this branch.

Test: sidebar keeps a Cursor pane visible and an OpenCode pane out of Claude Code hands

Co-authored-by: Orca <help@stably.ai>

* fix(terminal): preserve restored Cursor identity

* test(terminal): cover restored Cursor redraw suppression

* refactor(terminal): tighten Cursor identity handling and Claude frame matching

Review follow-ups on the title-resolution path:

- pty-transport dropped a native Cursor literal that main emits whenever a
  non-Cursor title preceded it, re-introducing the #10258 blank row in the
  renderer path. The pre-filter now projects the predecessor the drain will
  actually see, and defers to the drain gate while facts are still queued.
- applyTrackedPtyTitle threaded the cursor flag through 12 sites, including
  ptyRecordChanged bookkeeping the sole caller ignores. Force the status null
  once, and the activity-gated effects fall out unchanged.
- isClaudeIdentityFrameTitle missed a multiplexer-wrapped Claude title
  ("zsh | Claude Code"), costing a genuine Claude pane its identity. Reuse
  the ' | ' segment split that agent-title-owner already had inline.
- Keep title normalization on launchAgent: it only rewrites within an
  identity group (OMP wraps Pi), so a split does not make it wrong, and
  the hook-row path normalizes the same way.
- Drop the tab.ptyId tracker fallback, which read a pty that the pane
  identity check had just rejected.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-09 01:32:34 -07:00
Neil c6ded160b2 Revert the Korean Won to backquote mapping (#13312)
* Revert "fix(terminal): detach the Korean input-source probe when its setting goes off (#13283)"

This reverts commit 5a84dbb564.

* Revert "fix(terminal): stop the Korean gate caching an unknown input source as negative (#13182)"

This reverts commit 434959965a.

* Revert "perf(terminal): gate the Korean input-source probe on its setting (#13181)"

This reverts commit 42fc5375e8.

* Revert "feat(terminal): Korean Won (₩) → backquote key mapping for Korean keyboards (#13104)"

This reverts commit 24003936a5.
2026-08-09 00:44:35 -07:00
Neil 78f434dd85 fix(agents): deliver grok launch drafts on its composer frame (8s → 0.7s) (#13308)
* fix(agents): deliver grok launch drafts on its composer frame

Grok has no --prefill-style flag, so a launch draft (e.g. the issue URL of
a worktree created from a GitHub issue) always goes through Orca's
paste-after-ready path. That path used the default readiness signal: DECSET
2004 plus 1.5s of PTY silence. Grok shimmers its startup logo at ~12fps
until the session opens, so the quiet window never settled and the draft
fell through to the 8s hard timeout before it appeared in the composer.

Gate grok on its own composer glyph instead, anchored on the alternate-screen
switch rather than DECSET 2004: the shell that runs the launch command emits
2004 too, and its prompt may itself be the same glyph (starship, pure), so a
Codex-style anchor could paste into the shell. Grok keeps the quiet window
armed as a fallback because it renders differentially and paints the glyph
once, so a late-attaching scanner would otherwise wait out the hard timeout.

Measured against grok 1.0.0 driving the real scanner over a zsh -> grok PTY:
draft delivery moves from 8003ms to 689ms, with the URL landing unsubmitted
in the composer exactly as before.

* fix(agents): keep grok's quiet-window floor on DECSET 2004

The composer-glyph marker is anchored on the alternate-screen switch, but grok
can render inline (`--no-alt-screen`, `--minimal`, `[ui] screen_mode =
"minimal"`), where 1049h never arrives. Anchoring the quiet-window fallback
there too left those launches with no delivery path at all: readiness never
resolved, and the main-process caller drops the draft when it resolves null —
so the issue URL vanished instead of arriving late.

Give the signal two independent anchors: the marker still waits for the
alt-screen switch (so a starship/pure shell prompt can't trip it), while the
quiet window arms off DECSET 2004 exactly as the default signal does. Inline and
legacy-Windows-console launches keep their pre-existing timing; alt-screen
launches keep the fast marker path.

Verified on grok 1.0.0 over a real zsh -> grok PTY: alt-screen delivers at 687ms
via the marker, inline at 1949ms via the quiet window (the default signal
measures 1861ms on the same launch), URL landing unsubmitted in both. Adds a
recorded inline-mode trace fixture so the no-1049h path stays covered.

* fix(agents): revoke grok's alt-screen anchor when the screen is handed back

The composer-glyph anchor latched forever: once \x1b[?1049h had been seen, any
later `❯` counted as grok's composer. Two ways that pastes the launch draft into
the user's shell instead of into grok:

  - grok enters the alternate screen and then dies before painting a composer;
    the shell prompt that follows is `❯` under starship or pure.
  - a pager or editor started from the user's shell rc enters and leaves the
    alternate screen before grok is ever launched, arming the anchor against the
    shell's own prompt.

Track the anchor in stream order instead of as a latch: \x1b[?1049l revokes it,
re-entering re-arms it, and a marker only counts inside a segment where the
anchor is actually held. The chunk is walked segment by segment so ordering
within a single PTY packet is honored, with a 7-char carry — one short of the
escape sequence — so a split sequence rejoins without re-walking scanned output
into a second transition. Signals with no `markerAnchorEnd` (codex, opencode,
the default) keep their existing latch semantics untouched.

Also makes the trace-replay test model the hard timeout: the real waiters settle
at 8s, so a marker landing after that is not a delivery time.
2026-08-09 00:06:55 -07:00
Jinjing 17eefef502 [Tabs] Preserve host routing and reduce search churn (#13114)
* fix tab search host routing and churn

* Fix open-tab search to resolve hosts from worktree when active host unkn

- Use worktree.hostId to resolve execution host instead of defaulting to LOCAL_EXECUTION_HOST_ID
- Correctly populate search results for remote-only worktrees when activeWorkspaceExecutionHostId is null
- Remove automatic focus of terminal tabs after search activation

* Prevent stale tab results when user keeps typing ahead of deferred searc

- useOpenTabSearch now returns {query, results} to track which query the results describe
- Gate tab results on query match so stale results don't appear on user's screen
- Add live region (role=status) for accessibility of tab switch error messages
- Distinguish missing-worktree from missing-page errors in browser page activation
- Improve host resolution to prefer active host when worktree and repo don't specify one

* Re-pin entry to deferred tab results that rank higher

Track whether selection auto-follows the top-ranked result or was
manually positioned. Re-pin entry to tabs when they rank higher,
but preserve manual selection.

* Consolidate browser focus requests and simplify selection state

- Extract requestBrowserFocus to handle queueing + event dispatch atomically
- Simplify omnibox selection tracking with single pinnedOptionId state
- Optimize host resolution in tab search to compute once per query

* Report dead browser workspaces correctly and fold dedupe case by host

Two readiness-checklist fixes for open-tab search:

- Browser page activation checked page/workspace before the worktree, but
  deleting a worktree purges its browser workspaces and pages too, so a dead
  workspace surfaced as "Browser page no longer exists". Check the worktree
  first; routing already maps missing-worktree to the workspace wording.

- Editor-tab/file dedupe compared paths with separator normalization only, so
  a Windows worktree offered both "Switch to tab" and "Open file" for the same
  path in different case. Fold by the worktree path's syntax via the new
  isCaseInsensitiveRuntimeRoot, keeping WSL, POSIX and SSH roots case-sensitive,
  and add NFC so a macOS NFD listing matches an editor's composed path.

* Fix tab deduplication and resolve worktree host collisions

- Only editor tabs should suppress file entries; check contentType instead
  of relying on path being empty for non-editor tabs.
- Add executionHostId to simulator search results to disambiguate when
  the same worktree id exists on multiple execution hosts.
2026-08-08 17:46:11 -07:00
Hyein Cho 24003936a5 feat(terminal): Korean Won (₩) → backquote key mapping for Korean keyboards (#13104)
* feat(terminal): map Korean Won (₩) key to backquote on macOS

Korean keyboard users type markdown code fences and shell backquotes on the key that US layouts reserve for ` — 두벌식 and 세벌식 390 put ₩ there, 세벌식 최종 puts *, so there was no way to type a backquote without switching layouts.

Add a Mac-only terminal setting, "Korean Won (₩) to Backquote (`)", that rewrites the plain backquote-position keystroke to backquote while a Korean input source is active. It sits in Terminal → Advanced, right below the existing JIS Yen (¥) to Backslash (\) mapping.

The rewrite keys on the keystroke position alone — no character or layout-variant knowledge — and follows the live input source through the existing MacNativeTextInputSourceTracker, which refreshes on focus and keyboard activity (Caps Lock / 한영 input switches never blur the window). Modified chords and IME-composed events pass through untouched.

Covered by resolver and input-source tracker unit tests; verified manually in the GUI.

* docs(terminal): clarify Korean Won mapping scope and add docstrings

State in the setting copy that the backquote rewrite applies only while a Korean input source is active, and add JSDoc to the Korean Won resolver exports (addresses CodeRabbit pre-merge docstring coverage and copy-clarity findings).
2026-08-08 02:38:49 -07:00
Jinwoo Hong c991bb27d3 Add account-backed artifact sharing (#13012) 2026-08-07 23:02:29 -07:00
Neilandgatsby74 f968583e95 fix(remote): stop a reachable Orca server with a closed workspace window from reading Ready (#12477)
A remote Orca server whose workspace window is closed keeps answering status RPC, so Settings > Available Hosts showed "Ready" and the status bar showed "Connected" while every graph-backed operation failed. Adds the shared predicate `isRuntimeWorkspaceWindowClosed` (`graphStatus !== 'ready' && desktopWindowStatus === 'openable'`) and one host-health derivation with a new `workspace-window-closed` state, consumed by both surfaces. Hosts that omit `desktopWindowStatus` are unaffected, so the connected-host count and overall dot do not regress.

Fixes #12350

Co-authored-by: gatsby74 <gatsby74@users.noreply.github.com>
2026-08-07 21:11:31 -07:00
ce20a109da Persist the Linear issue list view and per-workspace filters (#12710)
* Persist the Linear issue list view and per-workspace filters

Layout, grouping, ordering, columns, and attribute filters survive a restart.
Facet ids are workspace-scoped, so filters are kept per Linear workspace and the
active filter is *derived* from the selected workspace rather than reset by an
effect on switch — no ordering race can apply workspace A's facets to B, and an
unresolved or cross-workspace selection reads as unfiltered without erasing
anything.

A single shared catalog backs the renderer state, `TaskResumeState`, and the
strict `ui.set` schema, so a new view option cannot leave paired web/mobile/relay
clients rejecting the whole payload. Persisted values are normalized as untrusted
input: a corrupt preference or a single bad workspace entry is dropped without
taking the rest of the resume state with it.

Deriving the filter also removed the guard that used to make three neighbouring
behaviours safe, so they are re-scoped here:

- The primary-team facet reset now fires only on an in-workspace team change.
  A workspace switch also changes the primary team, and clearing there wiped the
  filter that had just been restored for the workspace being switched *to*.
- The list-read force check no longer fires on the session's first read, so a
  restored filter serves warm cache instead of forcing a network round trip
  behind a blocking spinner on every cold start.
- The filter dropdown derives "no single workspace" from `workspaceId` alone.
  With an unresolved workspace it previously rendered the statically populated
  priority section, whose clicks now have nowhere to be stored.

* Harden Linear view persistence against the failures review surfaced

Five issues, each found by a reviewer and reproduced before fixing:

- The filter dropdown's prune effect only ran when the user opened the popover,
  because the filter was always empty at startup. Restoration makes it run on
  mount, where `availableTeams` may still be the issue-scraped fallback rather
  than the real fetch. Metadata complete for a *partial* team set passes every
  R12 guard, so it pruned facets belonging to teams it simply hadn't seen — and
  the write persisted, deleting them permanently. Gated on `teamsSettled`.

- `canonicalize` dedupes but enforces none of the transport bounds; only the
  throwing parser does. So `serialize` could emit a 101-label filter that the
  strict `ui.set` schema rejects, which drops the WHOLE taskResumeState — github,
  jira and linear query included — on every subsequent write, since the renderer
  resends the merged object each time. Added `boundLinearIssueAttributeFilter`
  and a round-trip test built from serializer output rather than a literal, which
  is the only kind that can catch renderer/schema drift.

- `linearIssueView` now carries `.catch(undefined)`: value tolerance stops at the
  top level, so any future instance of the above is a cosmetic reset of the view
  instead of silent loss of every other resume field.

- A workspace switch forced an uncached list read in both directions. The switch
  is a later observation, so the null-baseline fix didn't cover it; the cache is
  already workspace-keyed, making the force pure cost.

- Recency for the 20-workspace cap came from object key order, which is wrong
  twice: re-filtering an existing workspace left it at the head (first evicted,
  though just used), and an array-index-like key enumerates first regardless of
  insertion, so a write could evict the very entry it added. Recency is now an
  explicit ordered key list.

Also adds the nested parity assertion — the top-level one compares only
TaskResumeState's own keys, so a field added to LinearIssueViewResumeState stayed
invisible to it, which is exactly what `.strict()` rejects.

The wiring test was blind: deleting the hydration guard outright left all four
assertions green. The gate is now `shouldPersistLinearIssueView`, unit-tested
directly, and the file is renamed to the repo's `*-boundary.test.ts` convention
with an assertion that fails on that mutation.

* Log discarded Linear views and fix empty-filter serialization

- Schema now logs when linearIssueView is discarded, making validation failures visible
- Fixed serialization: filters that become empty after bounding are now omitted
- Added AssertNoExtraKeys type check for bidirectional schema/type parity
- Refactored view option catalogs to use canonical constants, preventing UI/schema drift

* Remove workspace persistence limits and LRU eviction

Stop capping persisted Linear workspace filters at 20 and evicting
least-recently-used workspaces. Simplify persistence to store all
workspace filters, gate persistence only on resume state application,
and remove tests that pinned implementation details. Users can now
persist filters for all their workspaces without arbitrary limits.

* add test for linear persistence

* Improve Linear filter test clarity and fix e2e overlay dismissal for CI

- Convert parameterized filter-pruning test to sequential assertions
- Fix dismissOverlayChrome to toggle overlay triggers instead of
  force-clicking inert page elements in headless CI

* Prevent TaskPage from stealing Escape from Radix menus

- Add check to detect open Radix dropdown menus and popovers; return
  early from Escape handler to respect their capture-phase ownership
- Update overlay dismissal in e2e tests to use keyboard.press('Escape'),
  now that TaskPage no longer interferes

* The capture-phase Escape guard in TaskPage bailed out for open dropdown menus and popovers, but an open Radix Select matches none of those selectors: the shared SelectContent wrapper (src/renderer/src/components/ui/select.tsx:60) renders data-slot="select-content" and Radix gives its content role="listbox", not role="menu". So with a select open, the window-level capture handler ran first, called preventDefault() and closeTaskPage() — closing the whole task page instead of just the select. Added [data-slot="select-content"] to the guard, as suggested. I did not add [role="listbox"]; the reviewer explicitly notes it's too broad, and the data-slot selector covers every select rendered through the shared wrapper.

---------

Co-authored-by: m4air <m4air@MacBook-Air.localdomain>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
2026-08-07 20:41:59 -07:00
OrcaWinandOrcaWin c3939ebf0e fix(mobile): allow reachable Hyper-V pairing addresses (#13107)
* fix(mobile): allow reachable Hyper-V pairing addresses

* fix(mobile): keep host-local Hyper-V addresses filtered

* fix(mobile): preserve explicit address on empty refresh

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-07 20:37:47 -07:00
Jinwoo HongandJinwoo-H 2f30eb9af5 fix(ai-vault): block deletion of live sessions (#13108)
* fix(ai-vault): block deletion of live sessions

* fix(ai-vault): retain external session authority

---------

Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
2026-08-07 20:36:23 -07:00
BingZ 523feda462 fix(commit-message): use Kimi --prompt instead of Claude --print (#11674)
* fix(commit-message): use Kimi --prompt instead of Claude --print

kimi-code rejects --print (suggesting --prompt). Deliver the generation
prompt as the --prompt argv value so branch auto-rename and commit
message generation work when Kimi is the selected agent.

Fixes #11669

* test(commit-message): cover Kimi argument defaults
2026-08-07 20:19:45 -07:00
NeilandOrca 940f2ff1e4 fix(terminal): quote agent resume for the tab's real Windows shell (cmd.exe) (#12476)
* fix(terminal): quote agent resume commands for the tab's real Windows shell

Cold restore and sleeping-agent resume built their launch line without the
host shell family, so win32 fell back to PowerShell argv quoting. On cmd.exe
tabs those quotes arrived literally and agent CLIs rejected the resume argv
and permission flags after a reboot ("unexpected argument ''<uuid>'' found").

Both call sites now share resolveAgentResumeLaunchTarget, which resolves the
launch platform and the live shell family together via
resolveLocalWindowsAgentStartupShell, honoring a per-tab shell override for
cold restore and leaving SSH / remote-runtime / WSL workspaces on their own
default quoting.

Fixes #12320

Co-authored-by: Orca <help@stably.ai>

* test(shared): cover cmd.exe resume quoting at the plan layer

Adopted from #12321 by @CountClaw.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-07 17:57:15 -07:00
JinjingandOrca 094d6821ef feat(native-chat): add model and effort pickers for grok (#12780)
* feat(native-chat): add model and effort pickers for grok

Grok had no session-option catalog, so the native chat composer showed no
pills and every launch ran the CLI's own defaults with no way to change them.

Adds a `GROK_SESSION_OPTION_CATALOG` (model via `-m`/`/model`, reasoning
effort via `--reasoning-effort`/`/effort`) and the discovery plumbing behind
it. Grok's selectable ids depend on the signed-in account and on `[model.*]`
config, so the seed carries only `grok-4.5` and a runtime `grok models` probe
supplies the rest as authoritative — a retired id must be droppable, since
launching one is a fatal exit rather than a warning.

Because `grok models` publishes `Default model:` and marks the row
`(default)`, the picker can name the model a fresh session is actually
running: `defaultModelIsCliDefault` plus an untracked record means no `-m`
was ever emitted, so the CLI is on its own default. That default scopes the
effort row but is never written to persisted settings — that field is what
authorizes `-m` on every later launch, and adopting a model the user never
picked would pin today's default forever, fatally so on an account without
it. `grok --help` publishes no default for `--reasoning-effort`, so the
effort value stays unnamed until something sets it.

Known gap: that refusal to persist is also a limit. An option set while on
the CLI default is dispatched and honored in-session, but reaches no later
launch — it persists under the default's id with `model` left unset, and
both `resolveNativeChatSessionOptionDefaults` and
`resolveAgentSessionOptionLaunch` bail without that key. Picking a model
explicitly persists normally. Closing this means teaching both to resolve
options from the default model while still refusing to emit `-m`, which is
the launch-args path and wants its own review.

Known gap: the picker infers "no `-m` was emitted" from its own in-memory
record, so a model reaching argv from outside it — the user's own
`agentDefaultArgs`, or a renderer reload that drops the record while the
flagged PTY lives on — leaves the pill claiming the CLI default while
another model runs. No wrong model is persisted.

Extracts `hasFlag` and `labelFromModelId`, and splits the model-probe spec
out of the commit-message registry so discovery no longer implies an agent
can write commit messages.

Co-authored-by: Orca <help@stably.ai>

* docs(native-chat): note the invariant keeping modelIsCliDefault agent-safe

The flag is computed without checking the catalog, so it reads as unsafe for
the four agents with no CLI default. It is safe only because `persist` bails
unless `modelId` is truthy, which for those agents implies a tracked model.
Widening that guard would silently change persistence for every agent.

Co-authored-by: Orca <help@stably.ai>

* fix: retire persisted models on mount and handle -- terminator

- When a pane mounts after model discovery has already settled, it now checks
  the cache and retires persisted models that are no longer available.
- CLI flag detection now respects the `--` option terminator, treating
  everything after it as positional arguments rather than flags.

* Fix: persist grok session options under probe-confirmed defaults

Options set under the CLI default were silently lost on restart.
Distinguish seed guesses from probe-confirmed defaults by renaming
`modelIsCliDefault` to `modelIsUnverifiedDefault`. Once confirmed,
adopt the default as a persisted flag so options survive restarts.

* fix(native-chat): close the retired-model fatal-launch paths from counsel review

Counsel report C1/C2 (High), C3, P1, C4:
- Untrack a session model an authoritative discovery dropped and gate every
  persist path, so option writes can never re-adopt a retired id (C1).
- Resolve launch defaults through the enrichment cache: a persisted model
  missing from every settled probe no longer becomes a fatal `-m` (C2).
- Serialize retirement and picks on one settings write queue that re-reads
  live state at apply time (C3).
- Stabilize onSwitchToTerminal so the session-option surface is not rebuilt
  every TerminalPane render (P1), and cap the enrichment host map (C4).

Co-authored-by: Orca <help@stably.ai>

* Store agent in enrichment entry and extract token utilities

Refactor enrichment to store the agent field directly instead of
parsing it from a composite key, and extract CLI flag token filtering
into a shared utility. Use a dedicated function for tracked model ID
lookup. Improves code reuse and reduces parsing overhead.

* Rename modelIsUnverifiedDefault to adoptModelAsLaunchDefault

Move the model adoption gate into the core session-options module, where probe confirmation and discovered-model status are known. This ensures adoption decisions are gate-checked before persisting to avoid fatal launch flags, and simplifies the picker surface by moving the logic to where it belongs.

* Keep model probe evidence by agent, not host

Store probed model IDs in agent-keyed cache independent of host cache, so
evidence persists across host eviction. Prevents retired models from being
treated as valid when host cache entries are evicted.

* Store agent in enrichment entries instead of separate proof-evidence map

Model probe evidence is now tied to enrichment entries rather than maintained in a separate per-agent map, eliminating the need for eviction logic that could disconnect proof from entries.

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-07 17:48:27 -07:00
NeilandOrca f8786d5224 fix(agent-history): fold trailing-slash, NFD/NFC, and project-fallback folder group keys (#12458)
Co-authored-by: Orca <help@stably.ai>
2026-08-07 15:06:37 -07:00
Pongsakorn PaetrakulandJinwoo-H 6aafb1d318 fix(gitlab): include bridge/child pipeline jobs in Checks (#12863)
Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
2026-08-07 14:11:31 -07:00
a77002c42b feat(ai-vault): delete a provider session from the AI Vault list (#10249)
* feat(ai-vault): validate session-delete targets for single-file providers

Add the pure judgement layer for deleting an Agent Session History entry.
`validateAiVaultSessionDeleteTarget` decides whether a session may be removed:
the agent must be one of the nine providers where a single file is the whole
session (gemini, copilot, cursor, hermes, devin, openclaw, droid, pi, omp),
the host must be local, and the renderer-supplied path must resolve inside
that agent's own session roots and match its discovery predicate.

To keep the delete roots from drifting from the scanner's own roots, the
WSL-expansion helper moves to session-scanner-root-dirs.ts and the OpenClaw
root derivation + session predicate become shared helpers that
discoverOpenClawFiles itself consumes.

The result is path-only and never touches the filesystem; a returned
`allowed: true` still requires an lstat/realpath re-check in the executor
(S-2) before removal, documented as a caller contract on the result type.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XDLggjSAjDnaWi3Y8U622i

* feat(ai-vault): move a validated session transcript to the trash

Add the filesystem executor behind session deletion. It calls the S-1 path
validator, then performs the fs-side guards that validator documented it
could not: lstat().isFile() rejects a directory or symlink, and realpath is
re-fed through the validator so a regular file reached through a symlinked
parent that escapes the agent's roots is rejected too. Only then is the file
moved to the OS trash via shell.trashItem, with ENOENT treated as success so
a delete racing an external removal stays idempotent.

WSL UNC paths (no Recycle Bin) are delegated to tryDeleteWslUncPath before the
Windows-local fs guards, mirroring fs:deletePath. Any non-ENOENT error is
returned as a failure result rather than thrown, since IPC payloads are
untyped at runtime.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XDLggjSAjDnaWi3Y8U622i

* feat(ai-vault): delete-session IPC handler, preload bridge, cache invalidation

Wire the S-2 delete executor to an IPC endpoint and expose it on the preload
bridge. The renderer calls aiVault:deleteSession with { agent, filePath,
executionHostId }; the handler fetches WSL homes, delegates to the executor
(which re-validates and trashes), and on a real delete invalidates the caches
that could otherwise keep serving the deleted session.

Cache invalidation is generation-guarded: a scan already in flight when the
delete lands carries an older generation and must not write its pre-delete
result back into the cache. Without this, an in-flight scan resolving just
after the delete would resurrect the deleted session for the 15s TTL — and
force-refreshing the panel only masks it for the desktop, not for the paired
mobile client or runtime RPC that share the same cache module. Both the shared
local-scope cache and the desktop multi-host cache carry the guard, with
regression tests for the in-flight race.

The delete result type moves to shared/ai-vault-types.ts so the renderer can
import the same contract the executor returns. To keep ai-vault.ts within the
max-lines budget after adding the delete wiring, two cohesive pieces are
extracted to their own files: the delete orchestration (ai-vault-delete.ts)
and listAiVaultSubagentSessions (ai-vault-subagent-list.ts). The latter is the
only handler with no dependency on this module's private cache state, so it is
the one piece that moves verbatim without threading state through a seam.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XDLggjSAjDnaWi3Y8U622i

* feat(ai-vault): renderer judgement for whether Delete is offered

Add the renderer counterpart to the main-side delete validator: given a
session, decide whether the row menu shows Delete enabled, or disabled with a
reason a tooltip can render. It reuses the shared deletable-agent set and
unsupported-reason map so the two sides can never disagree about which agents
are deletable, and reuses the existing local-host / synthetic-path renderer
helpers.

This is intentionally not a security boundary — it validates neither the path
root nor the file predicate. Those are the main process's untrusted-input
defense; the renderer only picks the affordance, and the main side re-checks
on delete regardless.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XDLggjSAjDnaWi3Y8U622i

* docs(ai-vault): correct deletability parity claim; test multi-reason agent

The renderer deletability check runs host -> synthetic -> agent, while the
main validator runs agent -> host -> synthetic. The two layers agree only on
deletable-or-not (renderer-false is a subset of main-false), not on the reason
code a doubly-failing session carries. Document that explicitly instead of
implying the orders match, and add the antigravity case (two reason codes) so
the agentReasonCodes array shape is actually exercised.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XDLggjSAjDnaWi3Y8U622i

* feat(ai-vault): add Delete to the session row menu with a confirmation dialog

Wire the delete affordance into AI Vault. Both the dropdown and the context
menu gain a destructive Delete item; a session that can't be completely deleted
(remote host, synthetic OpenCode-SQLite path, or a directory/registry-backed
agent) shows the item disabled with a reason surfaced both as a tooltip and as
an aria-label so keyboard and screen-reader users learn why. Confirming opens a
dialog that names the session and states it will no longer be resumable from
the provider's own CLI, then calls the delete IPC and force-refreshes the list
for immediate feedback (the main side has already invalidated its caches).

The confirmation copy says the session "will be deleted" rather than "moved to
the trash": on Windows a WSL session is deleted with rm inside the distro (no
Recycle Bin), so promising recoverability would be a lie on that platform.

Deletability is computed once per row and shared by both menus so they can
never disagree. New pure logic — the reason-to-tooltip mapping (including the
multi-reason join) and the delete action hook's deleted/rejected/failed
branches — is covered by unit tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XDLggjSAjDnaWi3Y8U622i

* fix(ai-vault): state that Delete is unavailable without naming the cause

The disabled Delete item explained a provider's storage layout to the user
("Claude sessions can't be deleted here: stores sessions as a folder, not a
single file"). That is Orca's problem, not the reader's — the tooltip now says
which sessions are affected and stops there. The non-local-host string stays as
it was: it states scope, not a cause, and tells the user what would work.

The reason-code plumbing existed only to compose that tooltip, so
AI_VAULT_UNSUPPORTED_DELETE_REASONS, AiVaultUnsupportedDeleteReasonCode, and the
renderer result's agentReasonCodes field go with it. Why each agent is excluded
moves into the comment above AI_VAULT_DELETABLE_AGENTS, where a reader looking
up the deletable set will find it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGuzChimmQ1dYX2raecrH7

* feat(ai-vault): delete claude, rovo, and grok sessions by their directory

These three were excluded only because the delete unit was one file. Their
sessions are directories — claude keeps Task subagent transcripts in a sibling
`<uuid>/subagents/`, rovo and grok keep everything under `<sessionId>/` — and
nothing in them is shared with another session, so a directory-aware delete is
still a complete delete. Supported goes from 9 agents to 12; the four that
remain (antigravity, kimi, codex, opencode) are blocked by a registry or a
SQLite row, which no delete unit fixes.

Validation now returns an ordered removal plan instead of a single path. Each
removal carries the kind it must be on disk and the roots its realpath must
stay inside, so the executor's guard is the same shape for a file and for a
directory. Companions come first and the transcript last: the transcript is
what puts the row on screen, so a part-way failure leaves the row to retry
from rather than dropping it and stranding the rest on disk.

Claude's `session-env/<uuid>/` goes with the transcript — it holds that
session's generated shell exports and nothing else. Its sibling
`file-history/<uuid>/` deliberately does not: it is the rewind buffer holding
earlier versions of the user's own files, and retiring a session is no reason
to take away the only copy that can restore them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGuzChimmQ1dYX2raecrH7

* fix(ai-vault): remove a claude session's own directory, not just its subagents

Deleting a claude session trashed `<uuid>/subagents/` and left `<uuid>/` behind
as an empty directory — one per deleted session, accumulating under every
project. The directory is named after the transcript, so it belongs to that
session as a whole; take it rather than the one subdirectory inside it. Still
derived from the scanner's own subagents path, so the two cannot drift.

Reaching the parent means a degenerate stem now matters: `..jsonl` passes the
extension check and its stem is `.`, which would resolve the session directory
to the project directory holding every session. Reject it instead.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGuzChimmQ1dYX2raecrH7

* fix(ai-vault): keep a session row collapsed when a menu action is chosen

Radix portals the row's dropdown and context menus out of its DOM, but React
still bubbles their clicks back through the component tree, so every menu
selection also hit the row's own click handler and expanded it. The trigger
button already stopped propagation, which is why opening the menu looked fine
and only choosing an item misbehaved.

It shows worst on Delete: the row expands behind the confirm dialog, so
cancelling leaves the list rearranged under a dialog the user just backed out
of. Toggle details only for clicks that land in the row's own subtree — that
covers the context menu and any future portalled surface, not just this one.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGuzChimmQ1dYX2raecrH7

* fix(ai-vault): harden the delete-confirmation flow against IPC rejection and mid-delete dismissal

Two robustness gaps flagged in review:

- handleConfirmDelete only branched on result.outcome. The main handler
  resolves with a 'failed'/'rejected' outcome rather than throwing, but the
  IPC invoke itself can still reject on a transport/serialization error, and
  the caller fires it with `void`. That reject would surface as an unhandled
  rejection with no toast. Catch it and show the same generic failure toast.

- handleDialogOpenChange cleared sessionPendingDelete on every open=false.
  The Cancel button is disabled mid-delete, but Radix still fires its
  Escape/outside-click/X close, which could dismiss an in-flight delete out
  from under itself. Ignore close requests while deletingSession is true.

Both covered by regression tests (verified failing without the fix).

* fix(ai-vault): route WSL UNC directory removals through the WSL rm branch

Directory-shaped deletes (claude's subagents/session-env dirs, rovo/grok's
session dir) gated the WSL branch on kind === 'file', so on Windows a session
under a WSL distro home fell through to shell.trashItem — which can't trash a
WSL-volume item (no Recycle Bin) and throws, or worse is silently stranded when
the 9P filesystem's unreliable lstat false-reports ENOENT and the executor
treats that as success. Single-file deletes predate the directory kinds, so the
file-only gate was correct until directory removals were added.

tryDeleteWslUncPath already supports recursive removal; pass recursive for
directory removals so they take the same WSL rm path as files instead of
shell.trashItem. Covered by two regression tests (file: non-recursive,
directory: recursive), verified failing without the fix.

Also drops the internal ledger-ID references (D-*, S-*) from comments in these
two files; they pointed at a private design doc a reader can't see.

* docs(ai-vault): drop internal design-ledger IDs from shipped comments

Comments across the session-delete feature cited decision/slice IDs (D-1..D-7,
S-1..S-5) from a private design document. Those references are meaningless to
anyone reading the code without that doc, so remove the IDs while keeping the
reasoning each comment carried. No behavior change.

* test(ai-vault): e2e-cover the real on-disk session delete

The unit tests mock lstat/realpath/trashItem, so nothing proved the whole IPC
path actually removes files. This spec seeds sessions into the E2E harness's
isolated HOME and deletes them through window.api.aiVault.deleteSession:

- a single-file session (gemini): the transcript is gone from disk and drops
  out of the list.
- a directory-shaped session (claude): the transcript, the <uuid>/ session
  directory (subagents included, no empty shell left), and the session-env
  companion are all gone, while the file-history rewind buffer is preserved.

Verified failing when the executor's removal is stubbed out. Runs on Linux CI.

* fix(ai-vault): address review findings on the session-delete flow

Three points raised in review:

- Disable Delete for a still-running session. resolveAiVaultSessionDeletability
  now gates on liveState (working/blocked/waiting) last — an otherwise-deletable
  session that is mid-run shows "wait for it to finish" instead of an enabled
  Delete, so trashing a live agent's transcript can't drop writes it is still
  appending. Unsupported/remote sessions keep their permanent reason.

- Realpath the roots, not just the target, in the executor's escape check. The
  roots were only resolve()'d (text), so a session under a symlinked root
  (~/.claude -> /Volumes/…) was falsely rejected; realpath each root (falling
  back to its text form when it can't be resolved) before the membership check.

- Invalidate the parse cache with the raw filePath, not resolve(filePath). The
  cache is keyed by the exact path the scanner discovered, so resolve() could
  normalise it away from the stored key and miss. Drops the now-unused import.

Also moves AiVaultDeleteSessionArgs/Result out of ai-vault-types.ts (which the
upstream merge pushed over the max-lines limit) into the ai-vault-session-deletion
domain module they belong to, and updates importers.

Regression tests added for the live gate, the symlinked-root accept, and the
reason string; verified failing without each fix.

* fix(ai-vault): type the deleteSession preload bridge as its real result

The bridge declared Promise<unknown> while AiVaultApi.deleteSession promises
AiVaultDeleteSessionResult, so the preload object leaned on the api-types
declaration to stay honest instead of being checked against it.

Co-authored-by: Orca <help@stably.ai>

* refactor(ai-vault): tighten the session-delete code to house style

Comments across the delete flow explained HOW alongside WHY and ran to a dozen
lines; they now carry only the non-obvious reasoning. The excluded-agent
rationale, the caller contract on the validator, and the file-history carve-out
are kept — those are knowledge, not narration.

Also removes three duplications the feature introduced:
- AiVaultSessionDeleteExecutionResult was an alias for AiVaultDeleteSessionResult
  whose comment pointed at a module the type no longer lives in.
- The synthetic-path predicate existed twice under near-identical names; the
  renderer now re-exports the shared one it already had a sibling import of.
- The delete-failure toast was written out verbatim in both the rejected and
  the thrown branch.

Co-authored-by: Orca <help@stably.ai>

* refactor(ai-vault): use a design-system dialog width and a stable row selector

The confirm dialog pinned an arbitrary sm:max-w-[440px]; every other dialog in
the right sidebar uses a scale token, and md (448px) covers the role.

The row-expand test selected the row by [draggable="true"], which stopped
naming the row when draggable moved to the title element upstream. It still
passed by bubbling, so the comment was the only thing wrong — now it selects
the title deliberately and says why the query is first-match (Radix's asChild
trigger repeats the subtree, so screen.get* sees duplicates).

Also types the e2e delete helper as AiVaultDeleteSessionResult instead of a
hand-written { outcome: string }, now that the preload bridge returns it.

Co-authored-by: Orca <help@stably.ai>

* refactor(ai-vault): consolidate agent sources and use system dialog

Discovery and deletion now share the same agent source definitions, eliminating the risk of them drifting apart. A single `AI_VAULT_AGENT_SOURCES` table declares each agent's root directories, file extensions, and acceptance predicates. Replaced the custom delete confirmation dialog with the system dialog, simplifying the delete action hook and removing boilerplate state management.

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Co-authored-by: Orca <help@stably.ai>
2026-08-07 11:47:26 -07:00
Jinjing 46b9d3b13a Break out test and generated lines in branch line total (#13057)
* rm comments

* reduce comment
2026-08-07 09:55:54 -07:00
NeilandOrca 02a1251c2d fix(native-chat): classify diff lines whose content begins with -- or ++ (#12459)
* fix(native-chat): stop diff colouring from misreading -- / ++ content lines as file headers

diffFromText skipped every line starting with --- / +++ as a file header, so a
deleted SQL/Lua '-- comment' (git emits '---<content>') or an added '++flag' fell
through to gray context with its marker still attached — and when it was the only
change, the two-marker gate dropped the coloured diff entirely.

Detect real headers structurally instead: an adjacent '--- <old>' / '+++ <new>'
pair outside any hunk. A hunk header or 'diff --git' line now also proves the text
is a diff, so a genuine single-line change renders while prose keeps the guard.

Co-authored-by: Orca <help@stably.ai>

* test(native-chat): adopt #12335 diff-collision vectors and add mobile parity

Pulls in @YuriNachos's test vectors from #12335 (header-less --- deletion, an
adjacent --x/++y content pair, mobile re-export parity) and adds the spaced
-- / ++ pair inside a hunk, which the pair-only rule in that PR misreads.

Co-authored-by: Orca <help@stably.ai>

* fix(native-chat): keep bare --- / +++ rules out of the diff marker count

Dropping the `---`/`+++` prefix exclusions made a bare `---` — a Markdown
thematic break or YAML document separator — classify as a deletion. Tool
results routinely carry those, so `---\na: 1\n---\nb: 2` went from correctly
rejected to rendering as a red diff.

A bare rule is never a file header (those need a path after the marker) and is
only content inside a hunk, so treat it as meta when outside one.

Fold the separate `isStructuredDiff` scan into the same pre-pass and skip
non-marker lines early, so the added guard costs no extra traversal: 5.1 -> 4.3
us per 120-line prose result, diff path unchanged.

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-07 03:13:14 -07:00
9e4e6ddae5 feat(native-chat): render omp transcripts (#11523)
* feat(native-chat): render omp transcripts

omp already ships as a first-class launchable agent with session_id resume, but
its transcripts had no decoder, so native chat could not render it — the agent
runs and the conversation stays a raw terminal. This adds the decoder and wires
it through the same path Claude, Codex and Grok use.

omp writes one envelope per line, `{ type, id, parentId, timestamp, … }`, where
conversation turns are `type: 'message'` and the rest is session bookkeeping.
Reasoning arrives as a `thinking` content block inside the assistant turn, so
the mapping follows Claude rather than Codex: thinking becomes a text block on
an assistant message, where Codex and Grok emit a separate reasoning role only
because their transcripts carry dedicated reasoning records.

  - toolCall -> tool-call, arguments passed through as the object omp writes
  - toolResult -> tool role, isError preserved
  - developer -> system, matching the Codex non-user/non-assistant fallback
  - blob-handle images drop, as the Claude mapper drops an image record with
    neither path nor url
  - bookkeeping and unrecognized types skip rather than throw

Session files are `<ISO timestamp>_<session id>.jsonl` under a per-cwd directory,
so the resolver matches the id as a base-name suffix the way Codex rollout files
are matched, and honors OMP_CODING_AGENT_DIR through normalizeAgentSessionsDir
so it stays consistent with the AI Vault scanner.

omp records no interruption or abort event, so unlike Claude and Codex there is
no NATIVE_CHAT_INTERRUPTED_STATUS_TEXT path.

Verified against 94,603 lines of real omp transcripts across four sessions:
50,546 records decoded, zero malformed, zero thrown.

* fix(native-chat): complete omp record coverage and gate remote transcripts

Review fixes on the omp transcript decoder.

omp writes several record types with no `content` field, so they decoded
to zero blocks and disappeared from the chat view entirely:

- `bashExecution` / `pythonExecution`: TUI `!command` runs, now a tool turn
- `fileMention`: `@path` attachments, listed by path (never `files[].content`,
  which is an auto-read dump)
- `custom_message` and legacy `custom` / `hookMessage` rows, gated on
  `display` the way omp's own renderer gates them

Also:

- `stopReason: 'aborted'` turns now surface as the interrupted row, matching
  the Claude and Codex decoders. An abort carrying partial content keeps it.
- A cancelled command cell now reads as errored. Every omp cancel path emits
  `exitCode: undefined`, which JSON drops, so an `exitCode !== 0` check read a
  cancelled run as a clean success.
- omp joins Grok in requiring a locally readable transcript. Its hook reports
  no transcript path, so under Model-A SSH the chat view opened against a disk
  this process cannot read and never loaded. Applies on mobile too, which
  shares the same allowlist.
- The session-file walk prunes omp's per-session subagent artifact
  directories, matching the AI Vault scanner. It was returning a subagent
  transcript instead of the parent session, and cost a full recursive readdir
  on every resolve.

* style(native-chat): apply oxfmt to the omp review fixes

Mobile CI gates `oxfmt --check`; the two root files were unformatted too,
just ungated there. Line wrapping only, no behavior change.

---------

Co-authored-by: plotarmordev <299844489+plotarmordev@users.noreply.github.com>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-07 00:58:29 -07:00
Jinwoo Hong 2b42de1f52 fix(orchestration): wake coordinators with mail pointers (#12988)
Wake idle Run coordinators with durable orchestration mail pointers while keeping message payloads in the store until check consumes them. Preserve waiter, Cursor, restart, real Codex title, and PTY replacement behavior.\n\nPart of #12953.
2026-08-06 22:50:09 -07:00
Jinwoo HongandJinwoo-H c9485fdded fix(computer): fence macOS HID coordinate clicks (#12981)
Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
2026-08-06 21:18:52 -07:00
4b2603420e fix(terminal): protect local ConPTY Ctrl+Enter without breaking TUIs (#12462)
* fix(terminal): gate Ctrl+Enter CSI-u on a negotiated kitty pane

Ctrl+Enter emitted \x1b[13;5u unconditionally, so a pane that never
negotiated the kitty keyboard protocol (local Windows ConPTY, plain
shell) printed the escape verbatim into the prompt. Mirror the
Shift+Enter guard and fall back to the legacy CR every emulator sends
for this chord. Keeps the intercept, so IME commit ordering and the
single-send dedupe still apply.

Fixes #12329

Co-authored-by: Orca <help@stably.ai>

* test(e2e): negotiate kitty via PTY output in the Ctrl+Enter spec

The Ctrl+Enter gate reads the PTY-output kitty tracker, which
enableKittyKeyboardReporting never feeds (it writes straight into
xterm's parser), so the spec pressed the chord on a pane the policy
still saw as un-negotiated and got the CR fallback. Negotiate from the
application side like the neighbouring Shift+Enter spec, and reset the
flags afterwards for the serial suite.

Co-authored-by: Orca <help@stably.ai>

* fix(terminal): preserve trusted Ctrl+Enter routing

* fix(terminal): scope IME redispatch ownership

* fix(terminal): reject conflicting Ctrl+Enter evidence

---------

Co-authored-by: Orca <help@stably.ai>
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-06 20:07:55 -07:00
Jinjing cb960408f2 fix(mobile): never auto-advertise virtual bridge addresses for pairing (#12962)
* fix(mobile): never auto-advertise virtual bridge addresses for pairing

Container/VM bridges stay manually pickable, but automatic defaults skip them so
QR codes do not race an unreachable direct path. Relay pairs without a local
address; LAN-only and runtime pairing fail closed on bridge-only hosts.

* fix(mobile): never auto-advertise virtual bridge addresses for pairing

- Set endpoint to null when no direct address is advertised, so the QR
  doesn't show an unreachable address to the scanning phone
- Distinguish "No address selected" (bridge exists but not advertised)
  from "No interfaces found" (genuinely nothing to pick)
- Add tests for NetworkInterfacePicker placeholder behavior
2026-08-06 19:17:43 -07:00
Wooseong KimandJinwoo-H e68831f32c fix(github-project): index fork upstream slugs for project row matching (#12822)
* fix(github-project): index fork upstream slugs for project row matching

Project cards often reference the public upstream repo while the open
clone's origin is a personal fork. Map the parent slug to the same Repo
so selected-repo filters no longer hide every board row.

Preserves origin-based getRepoSlug identity for non-project callers.

Fixes #12647

* fix(github-project): match project rows against fork upstream slugs

Resolve the referenced call to a nonexistent `resolveRepoUpstreamSlug` and
match the persisted `repo.upstream` parent instead of issuing an extra
`github.repoUpstream` RPC per repo on every index build — that lookup shells
out to `gh repo view` for non-forks, so it would have gated the Projects tab
on N network calls. `repo.upstream` is already resolved at repo-add time and
backfilled at startup, so the fix costs no IPC.

Origin matches take precedence over upstream ones so an open clone of the
upstream repo itself is never made ambiguous by someone's fork of it.

Also covers the two surfaces the origin-only match broke alongside the desktop
table: mobile's project row matcher and the store-slice row-mutation routing.

* fix(github-project): scope fork upstream matching by host and selection

Round-1 review fixes on top of the upstream-slug index:

- Apply origin-over-upstream precedence among *selected* repos instead of
  globally. An open-but-unselected clone of the upstream repo was shadowing the
  selected fork, so #12647 still reproduced for anyone holding both — and repo
  selection collapses to one repo per project key, which is exactly that case.
- Scope a fork's upstream identity key to the fork's own origin host.
  Persistence strips upstream.host, so GHES forks never matched their own rows
  and a GHES fork's parent could bind a same-named github.com row.

* fix(github-project): skip the fork alias when its own origin is unresolved

Round-2 review fix. `githubHostFromIdentityKey` cannot tell "origin resolved to
github.com" from "origin did not resolve" — both yield no host. A GHES fork
whose slug resolution had failed (auth lapse, unreachable runtime) therefore
landed in the github.com namespace, so an unrelated public Project row matched
it and Start work opened the wrong clone on the wrong server.

Require a resolved origin before indexing the upstream alias: it is the only
host evidence there is, and a repo with an unresolved origin was already absent
from the origin index, so nothing is lost that origin matching had.

* fix(repos): persist the fork upstream host instead of dropping it

`sanitizeRepoUpstream` kept only `{owner, repo}`, so a fork's parent lost the
server it lives on every time the record round-tripped through disk.

That forced the Project row matcher to re-infer the host from `origin`. The
inference is right for an API-resolved fork parent — `getRepoUpstream` stamps
`origin.host` there precisely because "a fork parent lives on the same server as
the fork". It is wrong for the other branch: a local `upstream` remote carries
its own host, so a github.com clone with a GHES `upstream` remote was indexed
into the github.com namespace, where an unrelated same-owner/name public repo
could claim it and Start work would open the wrong clone.

Keeping the host removes the guess. Absent stays absent, so records written
before this hydrate unchanged and the origin-derived fallback still covers them.
Also fixes the avatar for rehydrated GHES forks, which resolved against
github.com for the same reason.

* docs(github-project): correct upstream host fallback comment

Persistence now keeps non-empty upstream.host; originIdentityKey remains
the host fallback for older records without one (CodeRabbit nit).

* fix(github-project): own slug-index retry timer cleanup

Move the failure-retry setTimeout into its own effect so cleanup always
clears it. Scheduling from the async buildIndex then-handler failed the
react-doctor effect-needs-cleanup gate in static analysis.

* test(github-project): guard the slug-index retry timer, fix the mobile twin comment

Two follow-ups on 52298d82 and 2f89c20d:

- Cover the retry timer both ways: a failed resolution still re-resolves after
  the TTL and recovers the match, and the pending timer is gone after unmount.
  The second fails if the timer moves back into the async then-handler, so the
  property is guarded by more than the lint rule.
- The mobile matcher's comment made the same stale "persistence strips
  upstream.host" claim that 2f89c20d fixed on the renderer side.

* test(github-project): unmount slug-index hooks so React cannot flush after teardown

CI shard `tests node 24 6/16` failed with 10 unhandled
`ReferenceError: window is not defined` traced to this file. The tests mounted
hooks without unmounting, so React scheduler work flushed after the DOM
environment was disposed. All assertions passed; the shard failed on the
unhandled errors alone.

`cleanup()` after each test unmounts the trees. Does not reproduce locally in
isolation — it needs CI's worker pooling and file ordering.

---------

Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
2026-08-06 18:43:28 -07:00
Jinjing 8ddf575fe6 Revert "Remove source control group order preference (#12785)" (#12955)
This reverts commit ae1ed5e886.
2026-08-06 16:33:42 -07:00
Brennan Benson b9e9811924 fix(permissions): use standard macOS Local Network request flow (STA-3505) (#12833)
* fix(permissions): surface macOS silent Local Network denial with diagnostic and workaround (STA-3505)

On macOS 27 beta, NECP silently denies Orca's whole process tree Local
Network access: no prompt fires, the app never appears in System
Settings, and terminal child processes fail with EHOSTUNREACH. The
Settings trigger swallowed the probe's socket error and reported
'unknown' + a 'Permission request sent' toast, indistinguishable from
success.

Classify the mDNS probe outcome (EHOSTUNREACH/EHOSTDOWN -> denied,
clean send -> granted, else unknown), remember the verdict for the
status chip, and render an inline diagnostic with the documented
NECP re-evaluation workaround when denial is detected.

* fix(permissions): avoid false Local Network grants

* fix(permissions): use standard Local Network request flow

* feat(permissions): add local network connection test

* fix(permissions): nest local network connection test

* fix(permissions): collapse connection test by default

* fix(permissions): emphasize connection test action

* fix(permissions): restore outlined connection action
2026-08-06 14:10:26 -07:00
Brennan Benson 18b27ac11c fix(terminal): show AI Vault conversation names in tabs (#12778)
* fix(terminal): use provider-native session titles

* refactor(terminal): source session names from AI Vault

* fix(tabs): harden AI Vault title sync
2026-08-06 13:29:06 -07:00
Jinjing 1251e5530f Show SSH worktrees immediately via persisted metadata (#12799)
* Retire SSH worktree metadata an authoritative scan proved gone

The metadata fallback's protection against resurrecting externally deleted
worktrees lived only in renderer module state, so it died on every reload
while the SSH WorktreeMeta it guarded against persists forever
(gcStaleWorktreeMeta exempts any repo with a connectionId, because a local
existsSync cannot probe a remote path). Repro: `git worktree remove` on the
SSH host, let the authoritative scan purge the row, restart — the startup
fetch runs before SSH connects and the fallback re-lists the deleted
worktree as a ghost row.

Chose option (a), deleting the stale persisted meta in main, over persisting
the removal memory: the metadata is the thing that outlives the worktree, and
Orca's own removals already delete it (removeWorktreeMetadataAndTransientState),
so external removals now converge on the same end state instead of accumulating
a second, parallel tombstone list that would itself need eviction. The
in-session memory stays for the window before the async delete lands.

New `worktrees:forgetRemovedForExecutionHost` only accepts SSH hosts, requires
an exact repo owner, skips metas owned by another host, and refuses folder
repos — a folder workspace's meta IS the workspace record (gcStaleWorktreeMeta
skips those keys for the same reason) and no remote scan can retire one. The
renderer only calls it from the authoritative-removal path, so a mere
disconnect never deletes anything.

Also:
- hoist resetAuthoritativelyRemovedWorktreeMemoryForTests into a top-level
  beforeEach; removeWorktree writes that memory too, so suppression could leak
  across describes and silently hide a row.
- cover the requireAuthoritative gate that skips the fallback, which had no test.
- replace the raw NUL byte committed inside the coalesce-key template literal
  with a \0 escape; it made the file scan as binary to grep/ripgrep.

* test(worktrees): verify non-authoritative fallback skips removal

The non-authoritative fallback must not trigger worktree cleanup when it observes an absence — only an authoritative scan should. Tighten the expectation to ensure cleanup happens exactly once, when new data arrives after the connection state changes.
2026-08-06 00:33:12 -07:00
NeilandOrca a7ffb244e4 perf(terminal): bound the reattach payload agent-signal scan (#12681)
hasCursorAgentReattachPayloadScreenSignal built a char-by-char copy of the
entire reattach payload so it could read the last header plus 5000 chars. On a
2MB daemon snapshot that cost 17.5ms of synchronous renderer main-thread work —
~75% of what xterm then spends parsing the same bytes — and the miss case paid
it in full for a result that is always false.

Two changes, both matching existing in-tree precedent: bound the scan to a
256KB tail (as the kitty tracker already bounds its own scan), and strip via
the shared precompiled CSI_SEQUENCE_PATTERN instead of a hand-rolled loop,
which is also faster in V8 because it copies spans rather than building a rope
per character.

  2MB snapshot, header hit   17.5ms -> 0.80ms  (22x)
  2MB snapshot, miss          8.7ms -> 0.52ms  (17x)
  200KB snapshot, header hit  1.5ms -> 0.62ms  (2.4x)

config/scripts/terminal-reattach-payload-scan-benchmark.mjs reproduces this and
asserts every candidate agrees with the baseline before timing it. It also
records a negative result: porting the daemon mouse mirror's includes()
pre-filter to the kitty tracker makes reattach slower, because snapshots always
contain the introducer.

Adds guards for the two behaviours a future shortcut would silently break: a
CSI-split header must still match, and a header behind the tail bound must not.
Also byte-pins POST_REPLAY_REATTACH_RESET_KEEP_MOUSE, which shipped unpinned.

Co-authored-by: Orca <help@stably.ai>
2026-08-06 00:23:57 -07:00
Brennan Benson a2d438db17 fix(daemon): detect severed macOS TCC attribution behind terminal automation denials (STA-3491) (#12848)
* fix(daemon): detect severed macOS TCC attribution and surface daemon-restart remedy (STA-3491)

macOS pins the detached PTY daemon's TCC responsible process to the app
binary that forked it. Once that binary is deleted (packaged updates
replace the bundle), Accessibility/Automation grants on Orca silently
stop covering every daemon-hosted terminal: osascript/System Events
fails with -25211 no matter what the user grants.

- record spawnerExecPath in the daemon pid file at fork
- adoption checks it: severed + 0 live sessions -> replace the daemon
  (reason severed_tcc_attribution); live sessions are preserved
- Settings (Developer Permissions + Manage Sessions) show a visible
  banner pointing at Manage Sessions -> Restart while severed

* fix(daemon): harden TCC attribution recovery
2026-08-05 23:54:30 -07:00
Brennan Benson a30e3b9f61 feat(dashboard): add experimental agent map view (#12168)
* feat(dashboard): add experimental agent map view

* fix(dashboard): harden agent map behavior

* fix(dashboard): harden agent map recovery

* fix(dashboard): close map selection on view change

* fix(agent-map): center sparse layouts

* fix(agent-map): align completion and workspace actions

* Polish agent map interactions and repo labels

* feat(agent-map): add worktree lineage and project actions

* fix(agent-map): use marker for unread agents

* fix(agent-map): compact orchestrated families

* fix(dashboard): harden agent map actions and layout

* fix(agent-map): bound layout work and preserve interactions

* fix(agent-map): move unread marker to ring top-right

* fix(agent-map): seat unread marker on the ring's top-left edge

* feat(agent-map): restore the agent launcher and declutter map labels

Three gaps in the experimental Agent Map:

- The "start a new agent" picker was split onto a preserved branch during the
  08-02 rebase (47829cb226) and never re-landed. Restores that commit and its
  pop-out IPC, keyed on the raw worktree id rather than the map identity.
- Workspace labels draw at a fixed screen size with no collision handling, so a
  zoomed-out map stacked dozens of names on each other. Adds a declutter pass
  that seats project names first, then workspace names by attention, then
  project counts in whatever room is left.
- The pop-out had no workspace right-click at all: its renderer has no store, so
  the shared sidebar menu cannot mount there. Adds a snapshot-driven menu with
  the launcher and Sleep, relayed to the main renderer.

* refactor(agent-map): fold the map's filter rail into the shared toolbar filter

The rail duplicated the toolbar's project filter and cost the canvas 14rem of
width on the surface that needs it most. Agent states move into the toolbar's
Filter dropdown (map view only — the board's columns already separate them) and
count toward its badge; project filtering falls back to the toolbar's own. Show
all is the dropdown's Clear all, and Fit already lives in the viewport controls.

* fix(agent-map): isolate map work from main renderer

* perf(agent-map): stream status updates to popout

* fix(i18n): add agent map catalog entries

* feat(agent-map): glow working entities

* fix(agent-map): prioritize attention ring status

* fix(agent-map): distinguish subagent connectors
2026-08-05 22:57:16 -07:00
Wooseong KimandJinwoo-H 12472b8a63 fix(skills): match official skill files despite local sidecars (#12812)
* fix(skills): match official skill files despite local sidecars

Scope known-snapshot matching to manifest-listed files so agent-written
sidecars (e.g. agents/openai.yaml) no longer mark a package unrecognized
and block updates when official bytes still match.

Preserves fail-closed detection when a listed file's content drifts.

Fixes #12694

* fix(skills): scope lock trust and convergence to official files too

Sidecar tolerance stopped at the snapshot match, leaving three disk-vs-official
comparisons still judging the whole folder.

The lock-comparable hash covered every observed file, so a clean update beside
agents/openai.yaml reported as failed and read 'may be modified'. It is now
carried both whole and scoped to the current bundle's paths, and either may
satisfy the lock: the sidecar case only ever matches scoped, while an upstream
revision that ADDS a file only ever matches whole, so publishing one alone
would trade this bug for #11220.

Convergence re-derived the disk revision from that same whole-folder digest,
which no revision matches once a sidecar lands, retiring the stuck-lock gate
and arming an update the command provably cannot perform; it now honours the
revision observation already resolved.

Subset matching also let an older revision launder drift on a file the current
bundle lists, since that revision does not list it and so read it as a
neighbour. Identity now keys tolerance on what the current bundle owns.

---------

Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
2026-08-05 22:43:22 -07:00
Brennan Benson cd8c66551a fix(agent-hooks): resumed Claude Code session gets its sidebar agent row at SessionStart (STA-3386) (#12859)
* fix(agent-hooks): give resumed Claude sessions a sidebar row at SessionStart (STA-3386)

Claude's hook set never registered SessionStart and normalizeClaudeEvent
dropped it at ingest, so a resumed session that idled produced zero hook
traffic and earned no sidebar agent row until the first prompt.

- Register SessionStart in CLAUDE_EVENTS (local + remote installs).
- Map lead SessionStart (startup/resume/clear) to an idle 'done' row,
  resetting stale roster/task/cron/tool/prompt state like the Codex path;
  compact restarts and child-attributed SessionStart stay dropped.
- Thread hookEventName through the agent-status IPC payload so the
  completion coordinator can tell a session connect from a turn result;
  a SessionStart 'done' no longer raises agent-task-complete.

* fix(agent-hooks): mark SessionStart rows as session boundaries, not completions (STA-3386)

Review follow-up: represent the idle connect as a first-class
sessionBoundary flag on the status payload instead of gating one
renderer consumer on hookEventName.

- sessionBoundary rides AgentStatusPayload/AgentStatusEntry (done-only,
  clamped like interrupted); drops the hookEventName IPC threading.
- Completion-reactive consumers ignore session boundaries: the
  completion coordinator (task-complete notifications), automation
  dispatch observers (a connecting agent no longer completes the run
  and closes its tab), activity unread counts, and the dashboard
  finished timestamp; the status slice keeps boundaries out of
  stateHistory and preserves the flag across done->done repaints.
- SessionStart sources are allowlisted (startup/resume/clear) so
  compact restarts or unknown sources fail closed mid-turn.
- A live SessionStart now un-retires a reusable pane like a fresh
  prompt, so resume-in-reused-pane earns its row too.

* fix(agent-hooks): keep session-boundary dones out of teardown and completion history (STA-3386)

Review round 2:
- A boundary done no longer deletes the pane's launch-config registry
  entry, so a resumed idle TUI keeps its registered-launch-agent
  identity evidence.
- A boundary landing on a REAL done pushes that completion into
  stateHistory so the finished timestamp and unread badge survive a
  resume//clear right after a finish.
- The done->done flag carry yields to turn evidence (assistant message
  or changed prompt) so a genuine completion can never be suppressed.
- Star-nag value-moment observer and the server's OSC-equivalence
  dedupe now discriminate the flag.

* fix(agent-hooks): keep a displaced completion unread in the sidebar badge (STA-3386)

Review round 3: sidebar-badge mode counts only the live entry, so a
session boundary landing on an unacknowledged completion silently
dropped the sidebar badge while the agent-events count kept it. Count
the displaced completion from history for boundary rows, and pin the
behavior with countActivityUnread tests.

* fix(agent-hooks): prevent SessionStart completion side effects (STA-3386)

* fix(agent-hooks): preserve SessionStart through renderer IPC (STA-3386)
2026-08-05 22:06:36 -07:00
Jinwoo Hong b0ba51831c Add per-worker model and effort overrides (#12851) 2026-08-05 21:17:45 -07:00
Jinwoo HongandJinwoo-H 211b2d1a35 fix(runtime): require consecutive missed probes before reaping a paired socket (#12790)
The paired-runtime WS heartbeat terminated a client after a single unanswered
15s ping. One missed pong is UNKNOWN, not proof the peer is gone: a cellular or
Tailscale blackhole, or a stalled TCP retransmit, routinely swallows one pong
from a peer that is still there. Users on flaky paths saw constant drops, each
costing a full redial plus E2EE re-handshake and subscription replay.

Reap now needs MISSED_PROBE_LIMIT (3) consecutive unanswered probes, counted per
socket rather than timed. Any proof of life -- pong or any inbound frame -- clears
the count, as does a resume from a server-loop pause, since a gap the client was
never given a chance to answer must not top up its budget. Missed sweeps still
re-probe, so a recovered path proves itself on the next tick.

Three matches the liveness budgets already in the product: the web client gives
45s (25s idle + 20s probe grace) and the relay control gives 75s. The paired
transport's single miss was the outlier.

Also gives the web client's redial the one-sided jitter the shared-control path
already had, so a fleet dropped by one shared blip does not re-dial in lockstep;
the helper is extracted to src/shared/reconnect-jitter.ts and shared by both.

STA-3320, #12327

Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
2026-08-05 18:59:20 -07:00
Brennan Benson 79896cb9a6 fix(native chat): retain transcript while reconnecting (STA-3333) (#12495)
* fix(mobile): keep the cached transcript visible while reconnecting

A manual retry closes the client and opens a fresh one, so the chat session
hook saw a new client under an unchanged identity, dropped its settled read,
and handed out an empty list — the transcript collapsed to a full-screen
spinner until the swapped client's snapshot landed.

Hold the last settled list per identity (captured post-commit) and keep
rendering it while the re-read is in flight. `transcriptLoading` still gates
consumers that decide from an empty transcript, so the launch-draft seed is
unaffected. The held list is keyed by a new `sourceIdentity` (host/workspace)
in addition to agent/session/transcript, so it can never serve another
source's messages.

Refs STA-3333.

* test(mobile): assert the whole reconnect window, not just its first frame

The re-subscribe lands a commit after the first render of the swap, so a
regression that cleared the held list there left frame 0 green and still
blanked the transcript. Verified: clearing the cache in the subscribe
cleanup now fails this test, where before only the view-toggle test caught it.

* fix(mobile): don't derive a tappable ask card from the held transcript

The cache this PR adds keeps the previous list rendered while a swapped
client re-reads. useMobileNativeChatPrompts was the one consumer reading
`messages` without honouring `transcriptLoading`, so an ask answered on
the terminal resurrected as a live, tappable card during that window.

Gating on `transcriptLoading` is exactly base behaviour: `setRead` only
ever stores 'ready'/'error', so status==='loading' implied an empty list
before this PR. The live `askFromStatus` path is untouched.

* chore: keep merge formatting scoped
2026-08-05 18:06:36 -07:00
Brennan Benson 6df8997c3c fix(file-explorer): sort numbered file names naturally across every listing surface (#11576)
* fix(file-explorer): sort numbered file names naturally

The File Explorer compared names with bare localeCompare, so numbered
files listed 100, 200 before 99. Hoist the numeric collator Source
Control file rows already use (#10850) into src/shared and apply it to
the local and runtime directory listings, the name-filtered view, and
Source Control directory nodes, which were inconsistent with the file
rows one line below (#11426).

* fix(file-explorer): natural sort on SSH funnels, relay, and pickers

Adversarial-review round 1 rework:
- Both readDir funnels short-circuited to the SSH filesystem provider
  before the patched sort, so SSH workspaces kept lexicographic order;
  re-sort locally after the provider returns (the remote relay may be an
  older build), and fix the relay's own comparator for relay-native
  consumers.
- sortDirEntries (shared, unit-tested) owns the directories-first +
  natural-order listing contract used by every funnel.
- compareFileNames breaks numeric-collation ties ('2' vs '02') by code
  units so sibling order stays total instead of readdir order, and pins
  the collator locale to 'en' so every host produces one order.
- The SSH folder browser and runtime server dir picker now match the
  Explorer they browse into.
- Ordering pinned by tests at the relay, source-control tree, and shared
  helper.

* fix(mobile): natural sort in the mobile file explorer

Mobile re-sorted host readDir results with bare localeCompare, undoing
the host funnel's natural order (round-2 review). Reuse the shared
comparator and pin the order in the mobile suite.

* fix(file-explorer): natural sort at the renderer choke point and remaining ties

Round-3 review: the remote-runtime RPC and paired-web routes return the
host's order verbatim, so re-sort in readFileExplorerDirectory where
every desktop route converges; pin the SSH funnel with a handler-level
test; and route Source Control path compares through compareFileNames so
numeric-collation ties share one total order with the Explorer.

* docs(file-name-sort): state the real perf baseline in the hoist comment

* refactor(source-control): drop the dead collator export; pin the test oracle locale

* fix(file-listings): cover remaining natural-sort surfaces
2026-08-05 17:11:28 -07:00
Jinjing 2ff2a1b268 Display SSH worktrees immediately using persisted metadata (#12646)
* Display SSH worktrees immediately using persisted metadata

Users can now see known worktrees for SSH hosts without waiting for the
provider connection to establish. Worktrees are fetched from local metadata
and displayed as non-authoritative, then merged without replacing richer
live data once the provider becomes available.

* Show SSH folder workspaces immediately via persisted metadata

Add safeguards for metadata fallback: track authoritatively removed
worktrees per host to prevent resurrection, position new rows within
the host block to avoid jumping on authoritative scan arrival, and
preserve co-owner detection status during merge. Coalesce concurrent
metadata fetches to dedupe overlapping queries.
2026-08-05 16:52:49 -07:00
Jinjing ae1ed5e886 Remove source control group order preference (#12785)
* Reorder source control to show staged changes first by default

Stages are closest to the commit action and most relevant to the
commit workflow. Merges untracked files into Changes visually while
preserving their Git area. Removes the untracked-first preset and
includes migration logic for existing user settings.

* Drop source control group order user preference

Remove the sourceControlGroupOrder setting and related UI, migrations, and persistence logic. The source control view now always displays sections in the order: staged changes, unstaged changes, untracked files.

* Reorder source control to show changes before staged

Aligns with the edit-stage-commit workflow by showing unstaged
changes (active edits) before staged changes (queued for commit).
2026-08-05 15:29:46 -07:00
Jinjing debf4affe7 Display total lines of code change in branch header (#12771)
* Add branch line total chip to source control header

Display the total lines added and removed across a branch from its fork point, measured via `git diff <mergeBase>`. Only computed when the chip is visible (request gate on merge base OID), with 500ms soft deadline to protect status latency and 15s hard timeout. Deduplicated across concurrent pollers and cached alongside line stats. Omitted on failure — always shows exact or nothing, never a partial estimate. Updates throughout the stack: native git status, relay, renderer store/API, and UI components.

* Pin branch line total to app locale

Format line counts using the app's configured locale instead of the system
locale, ensuring consistent cross-platform display and test reliability.

* test: wait for coalescer joins instead of fixed sleep

Hold the diff until the second status pass actually takes the
branch-total coalescer lease instead of using a fixed 400ms sleep.
Fixes timing-dependent flakiness on slow machines.
2026-08-05 14:46:00 -07:00
Brennan Benson d4dfc35ac4 fix(mobile): preserve multi-image chat attachments (#12639)
* fix(mobile): preserve multi-image chat attachments

* fix(mobile): use preferred array syntax

* fix(mobile): harden multi-image attachment flow

* fix(mobile): retain first-send image previews
2026-08-05 13:23:29 -07:00