Commit Graph
7437 Commits
Author SHA1 Message Date
Jinjing 77c2cd7356 Add startup delivery diagnostics and success announcements
Terminal sessions now report startup command delivery details (whether written, presence, length, and delivery method) without logging the command text—preventing credential leakage and distinguishing missing commands from lost ones in diagnostics.

Setup scripts now announce completion on both POSIX and Windows before executing the startup command, so healthy setups don't appear stuck in the UI with "Waiting for setup..." as the last visible line.

Diagnostics failures are caught and ignored so they never break session creation.
2026-08-31 22:39:01 -07:00
JahyunBaek e3757a1d30 fix(worktree): ignore orphan remote-tracking refs in branch conflicts (#16699)
Refs #16646

Unify native, WSL, and direct SSH conflict checks behind the execution host, remove the duplicated SSH classifier, and cover orphan/configured remote behavior across both paths.
2026-08-29 14:36:24 -07:00
Neil 7655d20f04 fix(terminal): complete OMP stale cwd recovery (#17154) 2026-08-29 14:34:50 -07:00
JahyunBaekandClaude Opus 5 aa95bdb11a test(shared): stop two suites asserting POSIX separators on Windows (#16511)
Both files describe paths with POSIX literals while their subjects compose
paths through `node:path`, so the assertions only hold where the separator
happens to be `/`.

`node-markdown-document-discovery` keys its fake tree at `/repo/docs` and
`/repo/one`, but `discoverMarkdownRelativePaths` descends with
`join(absoluteDirectoryPath, entry.name)` — `\repo\docs` on win32. The child
lookup misses, `readDirectory` yields nothing, and the walk stops at the root:
`docs/guide.mdx` disappears and the depth-limit case never reaches its limit,
so it resolves `[]` instead of rejecting. Keying the children with `join` walks
the tree the subject actually walks.

`git-fetch-head-lock` expects `cwd: '/tmp/repo'` from a subject that returns
`path.resolve(cwd, 'repo')`, which is `C:\tmp\repo` on win32. Asserting through
`path.resolve` pins the behaviour — that `-C` and `--git-dir` are resolved
against the cwd — rather than the separator of whichever machine runs the suite.

Verified on Windows 11: the two files go from 3 failed / 12 passed to
14 passed / 1 skipped, and the wider `src/shared` run shows no regression.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 14:23:22 -07:00
Jinjing 67fa855b6c Make worktree palette hint rows keyboard-clickable (#17272)
* Make worktree palette hint rows keyboard-clickable

Hint entries like "See more" are now CommandItems that can be navigated with arrow keys and activated with Enter, instead of being non-interactive divs. This allows keyboard-only users to access the expand actions without mouse interaction.

* Make worktree palette "See more" keyboard-navigable

Preserve cursor position when expanding via keyboard: auto-select the first
newly revealed item at the previous index and restore input focus.
2026-08-29 14:15:45 -07:00
Neil 2dfaa676d8 chore: update oxlint and oxfmt (#17150) 2026-08-29 14:13:35 -07:00
Neil 78bba0772a perf(markdown): cache repeated syntax highlighting (#17158) 2026-08-29 13:59:45 -07:00
Neil 0bf5361c92 perf(markdown): update code highlighting incrementally (#17147) 2026-08-29 13:52:16 -07:00
Neil eb00123a81 perf(markdown): skip unmatched list tokenizer scans (#17134) 2026-08-29 13:43:27 -07:00
Neil 7c0c9cee57 perf(markdown): skip closed-search update renders (#17162)
* perf(markdown): skip closed-search update renders

* test(markdown): stress closed-search updates
2026-08-29 13:41:33 -07:00
Jinjing 9940355f6e Improve worktree palette layout with flex-1 titles (#17273)
Adds flex-1 to palette open tab titles so they expand to fill
available space, making better use of the palette's horizontal layout.
2026-08-29 13:33:35 -07:00
Neil fb6c2800ee fix(gpu): capture hardware identity in crash reports (#16973)
* fix(gpu): capture hardware identity in crash reports

* fix(gpu): order bounded crash diagnostics before fallback

* fix(gpu): keep fallback persistence ahead of diagnostics
2026-08-29 02:41:53 -07:00
Junhyeok Chae 314dcba98b feat(i18n): localize Agent Dashboard to Korean (#17121)
- Translate Agent Dashboard column headers (Needs You / Working / Idle), board title, and total count.
- Translate empty-column placeholder, "You" message badge, terminal preview actions, and error-boundary copy.
- Resolves English fallback in the Agent Dashboard (dashboardPopout) under the Korean locale; only "Done" was previously translated.
2026-08-29 02:38:32 -07:00
Neil c5591d0893 fix(gpu): persist the safe-graphics marker before the restart prompt (#16945)
* wip: gpu-startup-recovery

* fix(gpu): offer hardware retry after safe recovery
2026-08-29 02:29:15 -07:00
Neil 9db319dc06 fix(terminal): recover OMP from stale working directories (#17128)
* fix(terminal): recover OMP from stale cwd

* fix(terminal): harden OMP cwd recovery
2026-08-29 02:27:44 -07:00
Neil 8d3e32a2ff Fix setup-provisioned skills missing at agent startup (#17124)
* fix(setup): let repos gate agent startup

* test(setup): update runner call expectations
2026-08-29 01:53:31 -07:00
Neil fec82b9bc8 test: widen updater schedule boundary margin (#17130) 2026-08-29 01:49:02 -07:00
Neil d73d36bc99 test: stabilize process and transcript liveness checks (#17126) 2026-08-29 01:45:20 -07:00
Brennan Benson 6bef6d2727 fix(ci): stop the Linux Electron probe step from starving its own probes (#17071)
* fix(ci): stop the Linux Electron probe step from starving its own probes

The package job's "Test Linux Electron lifecycle boundary" step ran five
Electron probe files under Vitest's default file parallelism, so four full
Electron stacks competed for a 4-vCPU runner. Each probe carries its own
in-process deadline (20s for WebRTC, 25s for H3), and every observed failure
was one of those deadlines expiring: exit code 2, "no result", with every
sibling file in the same run slower than its own green maximum.

Run the step with --no-file-parallelism so each probe owns the runner, and
stop each probe nesting a private `xvfb-run --auto-servernum` X server inside
the step's own xvfb-run: reuse an inherited DISPLAY, and only own one when
there is none (shards, local dev), which leaves those lanes unchanged.

Also set each Docker-SSH E2E step's Playwright output aside before the next
step starts, because Playwright empties test-results/ on every run and only
the last lane's traces survived to the artifact.

* fix(ci): route the persisted-worker probe through the same display resolver
2026-08-29 01:38:17 -07:00
Neil 8dc26e152f perf(editor): stop re-rendering every code block on each keystroke (#17008)
* perf(editor): stop re-rendering every code block on each keystroke

Profiling a 305 KB document in a packaged build showed typing was dominated
by two things that had nothing to do with the text being typed.

Tiptap re-renders a React node view whenever its document *position* changes,
even when the node and its decorations are untouched (@tiptap/react 3.22.5,
ReactNodeView.update). Typing shifts the position of every node after the
caret, so one keystroke in a document with 533 code blocks cost 533 React
renders. That re-render only exists so a component can observe a fresh
getPos(); RichMarkdownCodeBlock never reads it, so it now opts out via an
explicit `update`. getPos() stays correct for later callers — Tiptap updates
its position bookkeeping before calling `update`, and passes getPos as a live
function rather than a captured value.

The language <select> also mounted ~25 <option> elements per code block, for
a dropdown almost nobody opens: 13,858 option elements in that document, more
than a quarter of its DOM. The list now mounts on first interaction
(mousedown/focus, flushed synchronously so the native popup never paints a
stale list); until then a single option renders the same visible label. The
labels themselves were getters that re-translated on every property read, so
one render cost thousands of i18next lookups; they are now resolved once per
locale.

Median keystroke latency, packaged build, M-series:

  305 KB   84 ms -> 59 ms
  600 KB  265 ms -> 201 ms

Verified in the running app that the dropdown still expands to the full list
by mouse and by keyboard, that an unknown fence keeps its verbatim label and
fallback option, that changing the language still applies, and that typing
inside a code block still updates with syntax highlighting intact.

This does not move the size limit: the blocking mount (1.7 s at 300 KB) is
what pins that, and it is unchanged. The constant now records the measured
numbers, including that node-view count drives cost far more than byte size.

* fix(editor): refresh cached code language labels
2026-08-29 01:35:28 -07:00
hwantage 4065d053cf feat(i18n): localize onboarding checklist steps to Korean (#17108)
- Add `feature-wall-setup-checklist-localized-copy.ts` using `createLocalizedCatalog` for dynamic step copy lookup.

- Bind localized step name and description in `FeatureWallSetupChecklist.tsx`.

- Add 16 localization keys in `en.json` and verified Korean translations in `ko.json`.

- Add unit tests in `feature-wall-setup-checklist-localized-copy.test.ts`.

- Resolves English fallback for all 8 onboarding checklist steps under Korean locale.
2026-08-28 23:05:29 -07:00
Brennan Benson fd9125ea8c feat(native-chat): Codex structured native chat restructure (#16729)
* feat(native-chat): port structured Codex sessions from restructure-recovery

Rebuilds the desktop structured native-chat implementation from
brennanb2025/native-chat-restructure-recovery (tip 4e31c08db3) on top of
current main as a single commit, scoped to the local Codex path.

Ported:
- Structured agent-session core: durable record store + single-writer lease,
  canonical journal, agent-session wire host/attach/eviction/subscribers,
  `agentSession.*` RPC surface (registered via ALL_RPC_METHODS; host-side
  mobile allowlist included for wire compat), pty write gate, transcript
  additions, and the Codex app-server adapter/launch resolution.
- Renderer: NativeChatStructuredSession view/composer stack, structured
  launch path with the single-flight guard, local structured session tabs
  sync, activation gate + structured inventory (read-only
  `agentSession.handoffStatus` probe), agent-session tabs in the tab strip,
  AI-vault structured session activation, and the settings pane with the
  parent Experimental Chat UI toggle plus the nested "Use updated structured
  native chat" toggle. New sessions require both flags, agent codex, no
  prompt, and a local non-WSL, non-Windows-host execution host
  (structured-native-chat-availability).
- Fixes 72c013cea6 (verified Codex launch recovery), 8ddbaf5e3d (defer
  native terminal view switching affordances), and 4e31c08db3 (release the
  launch gate after a visibility retry) with their regression tests,
  including the third-launch-after-retry guard case.
- Cross-version agent-session wire test + CI lane, packaging entries
  (proper-lockfile, agent-tooling asar excludes), and the wire-compat doc
  section.

Deliberately not ported: mobile/ changes, the Claude structured runtime
(only the claude-transcript-branch-proof and claude-structured-owner-identity
leaf modules remain, backing the kept TUI-recovery arms), the terminal↔chat
adoption/handoff flow (`agentSession.adoptTerminal`/`requestHandoff`, the
handoff request engine, TUI adoption machinery, orca-runtime adoption
methods), renderer switching affordances and their dead leftovers, the
hook/subagent-status refactor cluster, and unrelated branch changes. The
crash-during-acquisition recovery path (restart handoff adjudication,
restore/reverse re-acquire, lease schema handoff keys) is kept because every
plain direct launch depends on it; a trimmed handoff coordinator exposes
only status/restore/close.

Branch edits that targeted files main has since split (ipc/pty.ts,
worktrees.ts, rpc/methods/terminal.ts, useIpcEvents, pty-connection,
store/slices/terminals.ts, runtime-types, web preload) were re-applied to
the split modules, preserving main's newer logic (Windows CIM fallback,
browser tab close rework, cold-restore resume flow, dispatcher threading).

Known seam: the mobile clipboard image-provenance CONSUMER gate ships
(agentSession.send refuses unproven mobile image refs with
agent_session_image_untrusted) but the producer hunk in
rpc/methods/clipboard.ts stays with the unported mobile cluster, so mobile
image sends into structured chat fail closed until that side ports.

* fix(native-chat): trust only authenticated local image uploads

* fix(build): preserve Windows process-tree patch application

* test(windows): include process creation time in addon fixture

* fix(build): run windows-process-tree node-gyp from the physical package dir

gyp expands the node-addon-api dependency by probing node, whose cwd
resolves to the package's physical directory in the store, so the emitted
target is a store-relative ../../../../node-addon-api@... hop. gyp then
resolves that hop against the rebuild cwd; from the node_modules
symlink/junction it escapes the store and configure fails with
"node_addon_api.gyp not found" (run 32999886072).

Rebuild from realpath(package dir) so both bases agree, matching how the
package manager itself runs native install scripts. The regression test
replays gyp's expansion+resolution against the planned cwd and fails
without the fix.

* fix(native-chat): keep chat tabs visible through terminal closes and empty-worktree launches

Two proven blockers in the native Codex tab contract:

closeTerminalTab pre-empted the canonical unified close. With one terminal
left it deactivated the worktree on a terminal/editor/browser-only check,
blanking a workspace that still held a renderable agent-session tab; with
two or more it pre-picked a successor from terminal entities only,
re-stamping the group active before closeUnifiedTab's MRU/neighbor repair
could land on the chat tab. Successor choice now defers to the unified
contract whenever the terminal has a unified row, and deactivation is
gated on the unified renderable count (matching leaveWorktreeIfEmpty),
with the legacy pre-pick kept only for terminals without a unified row.

A structured session created on an empty worktree was published into the
host's headless group while preserveLocalLayout froze the local layout,
leaving the tab in store but permanently off screen. A preserveLocalLayout
owner now always takes client-owned placement — repairing a rendered
leaf whose group record is missing, or materializing a rendered group on a
truly empty worktree — and applies the client-derived layout repair while
still rejecting host-authored layout.

Regression tests drive the real store through closeTerminalTab (git
worktree and folder workspace) and the real snapshot applier for the
empty-worktree adoption states; all fail without the fixes.

* fix(native-chat): close stale turns and retry rejected sends

* fix(native-chat): retire hosted rows on structured tab activation

* fix(native-chat): preserve rpc defaults across main merge

* chore: format remote wire compatibility guide

* test(native-chat): cover retry after unconfirmed send

* fix(native-chat): reload outbox on session switch

* docs(settings): disclose structured chat platform limits

* fix(native-chat): await Codex launch-home preparation

* fix(codex): align child-process allowlist with async trust bridge

* test(identity): update inventory for tab surface refactor

* fix(windows): preserve process-tree CRLF patch sources

* fix(native-chat): anchor an unmatched chat echo where it was sent (#16117)

* fix(native-chat): anchor an unmatched chat echo where it was sent

The reported symptom was old user messages replaying below every new turn, so the
conversation read as scrambled. The cause was not that the echo failed to match a
transcript row. Claude consumes a mid-turn send through a `queued_command`
attachment and writes no `type:"user"` record for it, so some echoes can never
match, and no amount of matching will change that. The cause was WHERE an
unmatched echo rendered: buildMobileNativeChatTransientData appended every pending
item after the entire transcript, so it re-read below each turn that landed
afterwards.

Render each echo directly after the transcript row it was sent against, using the
baseline the send already captures. An unmatched echo is then at worst a duplicate
in the right position rather than a scrambled one, and it stays visible. Echoes
sharing an anchor keep send order; a send with no baseline, or one whose anchor
folding dropped, still falls back to the tail.

Deliberately NOT fixed by deleting the echo. Inferring from send ordering that an
echo can never match, then removing it, loses the user's own text for a message
the agent did receive, and it cannot fire in the common case anyway - measured
drain groups are 1,017 of size 1 against 55 larger. It also escalates an existing
gap: the count pass has no baseline-tail guard, unlike the glue pass, while
`messages` is a 40-row window that head-trims, resets on reconnect and grows at
the front on loadEarlier, so a false landing there would license deleting a
DIFFERENT outstanding message.

That count-pass gap is real and left for a separate change; anchoring makes its
worst case a duplicate in place rather than a scrambled conversation.

* fix(native-chat): preserve folded echo anchors

* fix(native-chat): preserve forward-folded echo anchors

* fix(native-chat): keep leading folded echoes in place

* fix(workspace-cleanup): show git status for every row (#16690)

* fix(native-chat): refuse structured chat on every Windows execution path

canUseStructuredNativeChat only refused win32 when a project runtime
resolved, so folder-workspace keys (and other keys with no project
runtime) failed open into structured chat on Windows. Fail closed on
win32 unconditionally after the host check, matching the settings copy:
local macOS/Linux only; Windows/WSL/SSH stay on terminal chat.

* fix(native-chat): restore runtime refusals behind the win32 gate

506d375de3 replaced the project-runtime checks with a bare platform test,
so a WSL or repair-required runtime resolution would no longer refuse
structured chat off-win32. Keep the unconditional win32 refusal and
re-run the runtime resolution after it, so the gate does not depend on
the resolver's own platform guard. Tests inject WSL and repair-required
resolutions on darwin/linux and fail against the regressed gate.

* fix structured session journal durability

* fix structured tab active pointer after restart

* fix(native-chat): await optional lease renewal callbacks

* refactor(skills): extract install error messages

* fix(agent-session): harden recovery ownership

* fix(native-chat): retain panes across tab activation

* fix(native-chat): address round-one review findings

* test(native-chat): align integration coverage after main merge

* fix(native-chat): harden round-two reliability

* fix(native-chat): harden round-three reliability

* fix(native-chat): close round-four recovery gaps

* fix(native-chat): separate bounded journal key forms

* fix(native-chat): reset outbox error in render on session switch

The switch effect adjusted error state after the sessionId prop changed,
tripping react-doctor's no-adjust-state-on-prop-change on the changed-code
gate and flashing the old session's banner for a frame. Reset it with the
render-time previous-value guard instead.

* fix(native-chat): invalidate stale outbox settlements

* test(native-chat): restore settled-error session-switch regression

a6e2379bd1 replaced this test with the in-flight settlement race test,
leaving the render-time error reset unpinned: deleting the reset block
still passed the whole native-chat suite. Keep both scenarios pinned;
they are distinct (settled error clears on switch vs stale settlement
invalidated in the commit-to-passive window).

* test(wire): make release checkouts race safe

* test(wire): pin cross-process checkout single-flight and importer specifier contract

* test(wire): harden release checkout lifecycle

* fix(build): drop CR-byte residue from windows-process-tree patch

The two trailing CR bytes on the patch's deletion lines are a proven
no-op: pnpm hashes patches CRLF-normalized (both forms hash to the
lockfile's 946ffb2b) and materializes this package without applying the
patch in either form, so the load-bearing build edits come solely from
applyWindowsProcessTreeBuildFixes() (#16947), which handles both source
EOL forms. Restore byte-identity with main and repin the contract test
to the post-#16947 reality: LF-only patch bytes plus lockfile hash sync.

* fix(native-chat): skip empty startup recovery
2026-08-28 16:45:58 -07:00
Brennan Benson 4cb013c0a9 Never let a non-owning provider answer a PTY presence question false during the daemon swap window (#16953)
* fix(pty): answer unverifiable, not false, for presence questions during the daemon swap window

During cold start the installed local provider is still the plain in-process
LocalPtyProvider until daemon-init swaps in the daemon router. It does not own
restored daemon PTY ids, but pty:hasPty and the runtime controller's sync
hasPty still let it answer — and its "not in my table" false read as an
observed absence: the renderer dead-session reconciler tears panes down on
exactly that false, remount recovery refuses on it, and terminal.list records
an observed absence instead of an unverifiable verdict.

- pty:hasPty now waits for the local-provider startup barrier before choosing
  an answering provider (the same #7742 guard pty:kill uses), so the post-swap
  owner answers.
- hasPtyFromRuntimeController is sync and cannot wait; while the startup
  barrier is unsettled it answers null (unverifiable), and it inherits the
  async probe's remote-handle guard: no locally routed provider may answer
  for a paired runtime handle.
- SSH-owned ids keep answering from their own provider without waiting, and a
  registration without a startup barrier (headless/orcad) keeps the in-process
  provider's false authoritative (#12393).

* test(pty): isolate the remote-handle guard from the swap-window gate

* refactor(pty): arm the swap-window settle watcher once per startup promise

* Gate pty:inspectProcess on the daemon-swap startup barrier

During the cold-start swap window the routed local provider is still the
pre-swap LocalPtyProvider, which does not own restored daemon ids; its
answer about one is fabricated, and today reads as unavailable only
because the inspection funnel happens to consult hasPty before the
provider's own inspection. Completion-sensitive inspection must not ride
on that internal ordering: defer until the swap lands, exactly like
pty:kill (#7742) and pty:hasPty. SSH-owned ids and no-barrier
(headless/orcad sole-owner, #12393) registrations keep answering
immediately. The thrice-repeated barrier idiom is now one helper.
2026-08-28 15:57:57 -07:00
Brennan BensonandBrennan Benson 7abdf037d6 Fix Windows daemon host pruning on unverifiable liveness (#16908)
* Fix Windows daemon host prune liveness contract

* Scope host prune evidence per version

* Drop redundant default cases from exhaustive liveness switches

All three switches consume ProcessLivenessVerdict/ProcessSignalEvidence values
constructed in-process by inspectProcessSignal/inspectProcessLiveness; the union
is never deserialized from a wire, RPC, or persisted record, so the defaults are
genuinely unreachable.

* Cover prune liveness gates and quarantine corrupt pid records

* Make the prune delete-gate fail safe and refuse truncated pid salvage

The prune switch shared #16900's delete-gate shape: an unhandled future
verdict status fell through into rmSync, protected only by the lint
exhaustiveness rule. Deletion is now opted into by a positively matched
'exited' via reclaimUnownedDaemonHostDir; a pinning test feeds an
out-of-contract verdict and asserts the host dir survives.

Pid salvage from corrupt records now requires the digit run to be
terminated by a following non-digit byte. A tear inside the digits leaves
a truncated prefix that is a different pid: probing it either quarantined
a record on an unrelated process's death or, when the prefix collided
with an immortal pid (Windows System pid 4), re-created the permanent
prune veto for that record. Unterminated digits mean the writer died
mid-write, so the record quarantines without consulting any probe.

* fix(daemon): stop an in-flight pid publish from being read as a dead version

publishDaemonPidFile creates the record before writing it (writeFileSync with
flag 'wx'), so a concurrent launch can read a live daemon's record as empty. A
two-process probe observed the empty window on 7 of 2273 reads.

An empty record was not treated as corrupt at all: the parser's legacy
bare-integer fallback coerces it to pid 0 (Number('') === 0) with appVersion
null, so the scan skipped it as a pre-relocation daemon, left its version
unpinned, and the prune reclaimed a running daemon's host image -- the exact
destructive outcome this change exists to prevent, reached without any
'unverifiable' verdict. A pid that is not a positive integer names no process
(process.kill(0, 0) probes the caller's own process group), so it is now a
veto rather than a skip.

Quarantine additionally refuses any record written in the last minute: an
in-flight publish is by definition fresh, while a record left corrupt by a
dead writer ages past the floor and is quarantined on a later launch. Fixed
locally rather than in parseDaemonPidFile, whose null result also drives an
unlink in daemon-stale-kill.

Each gate is pinned by a test that fails individually when it is reverted.

---------

Co-authored-by: Brennan Benson <brennanb2025@users.noreply.github.com>
2026-08-28 15:54:21 -07:00
Brennan Benson 5dc09db2cd fix(ui): label agent state glyphs and swap monitoring to a heartbeat (#16981)
* fix(ui): label agent state glyphs and swap monitoring to a heartbeat

The monitoring glyph read as unlabeled: AgentStateDot set only aria-label,
which renders no hover tooltip, so hovering it showed the row's own title —
the same truncated text already visible. Its row siblings (agent icon, model
chip) both had hover titles, leaving this glyph the odd one out.

Give every state a native title in the shared primitive, so done/working/
blocked/idle gain the same affordance across the sidebar, tab bar, dashboard,
kanban, cmd-J palette and AI Vault at once. Callers can override via a new
optional title prop; AiVaultSessionSubagents drops its now-redundant wrapper.

Native title rather than the Radix tooltip: AgentStateDot renders in two
surfaces with no TooltipProvider above it — the dashboard popout is its own
React root, and AgentMapScene — so Radix would throw there. StatusIndicator
already sets a native title for the same reason.

Also swap lucide Radio for Activity. Radio reads as "broadcasting"; the
heartbeat line reads as "still running", which is what the state means.
Mobile keeps its documented 1:1 parity with the desktop primitive.

Fixes STA-5794

* fix(ui): avoid duplicate agent state tooltips

* fix(ui): preserve disabled agent tooltip reason

* fix(ui): stop the state dot from shadowing a row's disabled reason

The shared AgentStateDot now emits a native title on every state, so at
any call site nested inside an element that already has a title, the
dot's generic state word wins on hover over the more useful ancestor
text. That regressed the sidebar agent row, which carries
`sendTargetDisabledReason ?? rowTitle`: hovering the dot showed
"Working" instead of the actionable send-target reason. Same guard the
review-notes send menu already uses.

Also covers three hunks that shipped untested: the Radix opt-outs in
ActivityPrototypePage and the AI Vault subagent line's dropped wrapper
title both stayed green when reverted, and the suppression test was a
`not.toContain` sweep that passed against the pre-fix tree.

* fix(ui): preserve heartbeat hover tooltips

* fix(ui): preserve lineage drop hit zones

* Use styled tooltips for state indicators

* Update jump palette tooltip assertions

* Limit status tooltips to agents

* Restore agent workspace status tooltips

* Keep status tooltips on agent indicators

* Clarify agent status tooltip ownership

* Restore agent-derived workspace status tooltips
2026-08-28 15:34:47 -07:00
Jinwoo Hong 5c10bf9001 fix(sta-5781): stop cross-client resets of workspace view preferences (#17057) 2026-08-28 15:20:21 -07:00
Brennan Benson ca0a6ec9be Stop the OS keyring probe from gating the first window on Linux (#16912)
* Stop the OS keyring probe from gating the first window on Linux

1.4.190 added an at-rest secret protection report and called it from the
`app.whenReady()` startup path, before the first window is created.
`describeProtectionGap()` asks Electron `safeStorage` whether the OS keyring
is usable, and on Linux that is a blocking D-Bus round trip to
`org.freedesktop.secrets`. A keyring that is present but locked with no unlock
prompter never answers, so the call sits until D-Bus times it out and the app
shows no window for over a minute.

Measured on Ubuntu 24.04 against a Secret Service that accepts the connection
and never replies, time to first window:

  1.4.188                          1.06s   (never contacts the keyring)
  1.4.190                         76.05s
  1.4.190 --password-store=basic   1.06s   (probe bypassed)

Nothing on the startup path consumes the report, so it now waits for the first
window's `ready-to-show`, with a timer fallback because that event can fail to
fire when the GPU cannot present and headless serve has no window at all. Same
build under the same hanging keyring: first window 80.48s -> 5.31s, with the
report still delivered.

STA-5765

* Pin the deferral the keyring-probe test exists to protect

The suite passed with the probe fired on browser-window-created instead of
ready-to-show, with setImmediate dropped, and with the fallback stretched to
10 minutes — every one of which reintroduces the STA-5765 stall. Drain the
queue before asserting and bracket the fallback so those mutations fail.

Also reformats the file to oxfmt.

* Report the keyring gap inline in headless serve

Deferring the probe to the first window is right for the desktop app, but serve
never opens one, so the fallback timer became its only path. That moved the
stall to after `printServeReady`: the runtime advertises itself, a relay or
mobile client pairs, and only then does the main thread freeze on the keyring —
stalling pings and PTY pumps, which a client reads as a dead host.

Serve now reports inline, which is the timing it already had, and blocks before
anything is advertised rather than under a live client.

STA-5765

* test(secrets): pin the once-guard against a late window reveal

The fallback can report first and the window reveal arrive after it; without
the guard that probes the keyring a second time, blocking the main thread just
as the user starts interacting. No existing case covered that order — removing
the guard left all six tests green.

* fix(secrets): keep the deferred protection report non-fatal, and pin the wiring

Deferring the report moved it off `whenReady`'s promise chain. A throw there was
an unhandled rejection the app survives; inside `setImmediate` it is an uncaught
exception, and `installUncaughtPipeErrorGuard` re-throws those fatally — so a
diagnostic the module documents as deliberately not fatal could kill the app.
Wrap the deferred call so it degrades to a warn. Serve keeps the inline posture.

Nothing outside index.ts referenced the scheduler, so reverting the call site,
or flipping `deferUntilFirstWindow`, left the whole suite green — including the
headless-serve regression an earlier review already caught once. Pin the wiring
as source text, following the host-port-bootstrap-wiring idiom with every anchor
bounded, and pin the two module gates that were only jointly covered.

* test(secrets): make the deferral wiring pin resist an inert call site

Round 2 of the review gamed the pin it had just added. Both anchors were bounded
against -1 but not against overshoot, and the marker matched anywhere in the
file — so nesting the call in a block, prefixing it with a guard, or commenting
it out all left three green tests standing over a call that never runs.

Bound the slice length, and anchor the marker to a statement at whenReady's own
indent. Commenting the call out, wrapping it in `if (...) schedule(...)` with or
without a block, and flipping the flag each redden now; previously only the
unused-import typecheck error caught the first.
2026-08-28 15:05:24 -07:00
Jinwoo Hong 6256f3d137 Fix lost session-tab changes during initial census (#17064) 2026-08-28 14:57:48 -07:00
Brennan Benson 4fa3022c47 fix(remote): stop painting a disconnected host as connected (#17050)
* fix(remote): stop painting a disconnected host as connected

A remote host row read "Connected" with a green dot in two states where it
was not connected: a cleanly closed control channel (server restart, host
sleep, network blip leaves lastError null, and the mapping required an error
string before it would say disconnected), and a half-open handshake still in
awaiting_ready/awaiting_authenticated.

lastError/lastClose were also never cleared on a successful reconnect, so a
recovered host kept showing "Connected" beside a stale failure indefinitely.
The SSH lane already clears on success; the shared-control lane did not, which
is why only Remote Server rows showed stale text.

* test(remote): cover stale diagnostics after reconnect
2026-08-28 14:39:20 -07:00
Brennan Benson 41ce4fadd9 Fix GitLab MR management menu in Checks sidebar (#16906)
* fix: add GitLab MR management menu

* fix: restore GitLab menu typecheck

* fix stale GitLab review relink updates

* fix review relink guard lifecycle

* test local owner scope for GitLab relinks

* fix(gitlab): honor linked MR during review lookup

* fix: reuse hosted review cache after relink

* fix: avoid duplicate GitLab detail refresh
2026-08-28 14:29:23 -07:00
Brennan Benson df95f03101 Show ready and close actions for draft reviews (#16889)
* Fix draft review sidebar actions

* Drop unused React import in draft actions test

The automatic JSX runtime makes the default React import dead, and
tsconfig.tc.web.json failed the branch on TS6133.

* Add localization keys for draft review actions

The new Ready for review controls introduced five untranslated keys and
the static analysis job requires them present in en.json.

* Name the draft action for what it does

The button read 'Ready for review', which states a status rather than an
action, directly under a header already showing the PR state. The i18n
key (markReady), the in-flight label ('Marking ready...') and the success
toast ('marked ready for review') all already used the verb.
2026-08-28 12:31:55 -07:00
Brennan Benson 2c86d2a3bd fix(agent-hooks): stop test runs and secondary profiles deleting the user's agent hooks (STA-5679) (#16980)
* fix(agent-hooks): stop startup from deleting another instance's managed hooks (STA-5679)

Startup reconciliation removed the managed agent hooks whenever THIS profile had
the agent-status-hooks off switch set. The hook files it removes are user-global
(~/.claude/settings.json, ~/.cursor/hooks.json), so a second Orca profile with the
switch off deleted the hooks every other running instance depends on.

Cursor is the only agent with no title-derived status fallback: its native title is
deliberately parsed as status-less, so a hookless Cursor pane is floored at 'idle'
rather than showing a spinner. A global hook wipe therefore surfaces as "Cursor
loading status missing from the sidebar" while Claude and Codex still paint status
from their own titles, which is why this reads as a Cursor-only bug. Codex is
unaffected either way because its hooks live in an Orca-owned runtime home.

Honoring the off switch only requires skipping the install; removal stays on the
explicit Settings toggle, which is the user-initiated path that should own it.

Regression from #2778, which restored the destructive startup branch.

* fix(cli-tests): stop the deferral suite deleting the developer's real agent hooks

runtime-client-deferral.test.ts runs the REAL `main()` and feeds it
`agent hooks off`. It mocks only ./runtime/environments and ./runtime-client, so
the production handler ran end to end: updateEnabledOnDisk() wrote its state file
and applyAgentStatusHooksEnabled(false) called removeManagedAgentHooks() against
the developer's OWN ~/.claude/settings.json and ~/.cursor/hooks.json.

A green test run therefore deleted every Orca-managed hook on the machine. Agent
status then stopped reporting until the next Orca restart reinstalled them —
silently, because the hook POSTs still return 204 and Cursor has no title-derived
status fallback at all.

The byte-for-byte equivalence twin already refuses these exact tokens, commented
"MUTATING — writes outside ORCA_USER_DATA_PATH (`agent hooks off` parks the real
~/.claude hooks)". The vitest twin never got that guard.

Stub the hook-controls module rather than dropping the row: `agent hooks off` is
the only case in the table that reads ctx.client, so it carries the
null-vs-undefined coverage the other four cannot. All 23 tests still pass, and a
sandboxed HOME now keeps its hooks (5 -> 5) where it previously lost them (5 -> 0).

* fix(cli-tests): ratchet agent hook deferral safety

* fix(agent-hooks): keep startup reconciliation install-only
2026-08-28 11:33:14 -07:00
Jinjing ee1e002d00 Change artifact link button wording (#17037) 2026-08-28 09:38:10 -07:00
Neil 94e7586665 fix(relay): stop advertising an agent session for a pane whose tab is gone (#12447) (#17012)
`pty.shutdown` was the only signal the relay ever got that a tab had closed,
and it retired nothing: it requested a kill and returned. Every retirement path
in the relay keys on proof of process death, so when the kill did not reap the
pane shell the relay kept forwarding the orphaned agent's hook events as a live
agent pane with no tab, and kept publishing its `agentSessionOwners` from
`pty.listProcesses` with no liveness check at all.

The relay now records the retirement the client stated, drops the pane's cached
agent status with it, refuses to forward or replay posts from a retired pane,
verifies the kill actually landed instead of assuming it, and reaps any PTY
whose pid it can prove is gone before listing it.

No wire change: no new RPC, field or stream opcode.
2026-08-28 04:17:45 -07:00
Neil a907dd2ba2 fix(pty): key buffered pre-attach PTY exits on the incarnation, not a sequence fence (#17010)
* fix(pty): key buffered pre-attach exits on the PTY incarnation, not a clock

A restarted SSH relay renumbers PTYs from pty-1, so a fresh spawn is routinely
handed an id whose previous shell is still emitting a late exit. #16970 stopped
that exit blanking the new tab by dating every buffered record and dropping
anything older than the spawn request. That fence is a clock, so it cannot judge
a stale exit that arrives AFTER the request left — the residual risk #16970
documented.

Thread the incarnation main already puts on the pty:exit payload (and the
pty:spawn reply) through preload to the pre-handler buffer, so a buffered exit
names which lifetime of the id died. An exit disagreeing with the incarnation
now attaching is discarded whenever it arrived.

Only a positive disagreement discards: absence stays "unknown", never a
mismatch, so hosts that predate the field keep #16970's behaviour exactly. The
fence is retained for the two cases with no incarnation to compare — buffered
bytes (pty:data carries none) and unnamed exits.

No wire change: incarnationId was already published on the relay's pty.exit
notification and pty.spawn reply, and already forwarded over the in-process
pty:exit / pty:spawn IPC. Only the preload types and the renderer read it now.

* fix(pty): read the incarnation through the shared guard, not truthiness

A malformed incarnation is evidence of nothing, so it must read as "unknown"
rather than as a value that disagrees with every well-formed one — otherwise a
non-string on the payload would discard the very exits the buffer exists to
deliver. Route both the record and the comparison through the existing
isPtyIncarnationId guard.

* refactor(pty): name the bounded-map helper after what it does

It evicts the oldest entry when the map is full and the id is new; it reserves
nothing. Rename only — no behaviour change.

* fix(pty): key the buffered-exit STORAGE on the incarnation too, not just the check

Review caught a swallowed exit. Keying only the comparison on the incarnation
while the storage stayed one slot per pty id left the two races this buffer
exists for able to cancel each other out:

  1. the freshly spawned shell dies before the pane attaches -> its exit (X) is
     buffered;
  2. the relay flushes the previous owner's exit for the same recycled id (W),
     which OVERWRITES X in the single slot;
  3. the spawn reply names X, so the identity discard drops W -- the only
     record left.

registerExit then finds nothing and the pane binds to a PTY that is dead and
will never be reported dead: a hang instead of the blank tab #16970 fixed.

Store one record per lifetime, capped at 4 per id, so W can never evict X. A
duplicate exit for a lifetime replaces that lifetime's record rather than
crowding out another's; drain still delivers the newest survivor, preserving
the last-write-wins behaviour a single slot always had.

* fix(pty): filter buffered exits by lifetime inside the buffer, not at call sites

Review found the identity was enforced only where connectIpcPty calls the
discard, while preHandlerPtyExit has several other consumers. The severe one is
registerEagerPtyBuffer: both background launchers spawn directly and then drain
whatever is buffered for the returned id, so a relay-recycled id holding the
previous owner's exit tore a freshly launched agent session down seconds after
it started -- no fence, no admitPtyId, no identity check at all.

Move the rule into the buffer: every read goes through
admissiblePreHandlerPtyExits, so a record proven to belong to another lifetime
is unreachable by construction rather than because each caller remembered to
discard first. hasPreHandlerPtyExit/drainPreHandlerPtyExit take the asking
lifetime; registerEagerPtyBuffer and registerExit thread it through, and both
background launchers pass the incarnation their own spawn returned.

A reader that cannot name an incarnation still sees everything, which is the
honest answer -- it holds no evidence to discriminate with. That keeps the
pre-spawn fast path in connectIpcPty behaving exactly as it does today; see the
PR for the consumers this still does not cover.

* chore: drop unrelated formatter churn in reliability-gates.jsonc

Repo-wide oxfmt reindented pre-existing entries in a file this change never
touches. Keep the diff to the PTY incarnation work.
2026-08-28 04:17:36 -07:00
Neil 7bbb8adc61 fix(ssh): replay an undelivered remote PTY stop on the next handshake (#12447 item 1) (#17011)
* fix(ssh): replay an undelivered remote PTY stop on the next handshake

A pty.shutdown that dies on the transport left the remote shell running
forever: kill.ts marked liveness unverifiable and nothing retried.

Record the undelivered stop on the existing durable SshRemotePtyLease and
replay it against the authoritative host on the next handshake to that same
target, fenced by the host-minted PTY incarnation so a replay cannot kill a
later PTY that reused a recycled pty-N id. Retire the record on confirmed
delivery, on the host reporting the PTY absent, and on a bounded TTL.

No wire change: the fence reads incarnationId, already published on
pty.listProcesses. A host that does not publish it degrades to no replay.

* fix(ssh): do not leave a replayable kill order behind a reversible stop

Worktree sleep stops through stopAndWait and marks those stops reversible;
when one does not land the pane stays live and the user keeps using it. An
order recorded there would come back on a later handshake and kill that
terminal. Only killPtyFromRuntimeController — where the client gives the PTY
up for good — records one, and it skips any PTY a reversible stop owns.

* fix(ssh): cover the renderer kill route and harden the replay's evidence

pty:kill is a separate implementation from killPtyFromRuntimeController and
is the one an ordinary tab close reaches, so the record was never written on
the path #12447 describes. Extracted it out of inspect.ts (which was over the
line budget and was not what the file is named for) and wired both branches.

Also:
- finishPtyShutdown no longer retires the order. It runs on paths that asked
  the host and on paths that never did, so retiring there was a contract every
  caller had to know, and the one that forgot silently dropped a kill order.
  Retirement is the replay's, on inventory evidence only.
- A recycled relay id now expires its lease. Declining to kill was only half:
  reattach fences on paneKey/tabId, never incarnation, so an untouched lease
  bound the user's old pane to whatever now holds the id.
- Dropped isPtyAlreadyGoneError from the tombstone path. It matches message
  text a transport failure could wear; every tombstone now traces to a listing.
- TTL is owned by a durable prune that actually deletes, not by a branch that
  was unreachable behind the read filter and only looked tested.
- The replay re-reads the inventory per wave and re-checks the fence next to
  each shutdown, and can never reject into the connect path.
2026-08-28 04:08:15 -07:00
Neil 50c88126eb fix(terminal): damp park-verdict oscillation on the rendered verdict (#15136) (#17009)
The verdict-flip pin was computed from flips on the rendered park verdict but
applied only to the cold-park candidate set, so a loop driven by the
worktree-level park prop or the activation-deferred branch kept remounting a
pane at commit cadence while the pin silenced its own breadcrumb for 60s.

Apply the pin to the rendered verdict via selectParkVerdictPinnedTabIds, and
expire pins for every live tab so damping lapses on its own instead of
re-arming forever.
2026-08-28 04:08:06 -07:00
Neil 65dd06a870 feat(editor): add "Open anyway" for oversized rich Markdown files (#16964) (#16971) 2026-08-28 02:38:51 -07:00
Neil 2b0ee06205 docs(env-recipes): warn that snapshotting a started runtime bakes its identity (#17001)
Snapshotting a VM on which `orca serve` has already run captures the
runtime's user-data dir into the image. Every VM booted from that image
then shares one pairing identity and one agent-session-authority key,
which defeats the per-device token design.

Confirmed by booting two VMs from one such snapshot: both emitted
identical deviceToken and pairedDeviceId.

Adds the rule to the base-snapshot section and repeats it for the
agent-auth layer, which is the likelier place to start the runtime by
hand while smoke-testing. Says to delete the whole user-data dir rather
than a named file list, since that list drifts as Orca adds state.
2026-08-28 02:37:55 -07:00
BingZandNeil caef20fec8 fix(gitlab): paginate TaskPage issues beyond 50 (#13538)
Co-authored-by: Neil <neil@stably.ai>
2026-08-28 02:25:48 -07:00
erishandFuzzwah 83320bff9e fix(pty): stop fish printing a test error when codex is absent (#16893) (#16923)
fish splices an unquoted command substitution that produces zero words out of
the argument list entirely, so `test (type -t codex 2>/dev/null) = file` became
`test = file` whenever codex was not resolvable. fish's test rejects that as
malformed and wrote "test: Missing argument at index 3" to stderr on every pane
and agent launch for fish users without codex installed.

Capture the substitution into a local first, mirroring the bash/zsh variant,
then clear it so the guard leaves no variable behind in the user's session.
Quoting in place is not a fix: fish never performs command substitution inside
double quotes, so the wrapper would silently never be installed.

Wrapper behavior is unchanged for every codex state (file, function, alias,
absent) on fish 3.1.0 and 4.7.1, and no other shell's preflight is touched.

Combines #16923 and #16928.

Co-authored-by: erishforG <erish2150@gmail.com>
Co-authored-by: Fuzzwah <rob.crouch@gmail.com>
2026-08-28 01:57:21 -07:00
Jinwoo Hong 8fa1b3c16c test(release): bound Windows skill budget fixture (#16992) 2026-08-28 01:34:28 -07:00
Jinjing c4b39295c1 style: format codebase (#16935)
* style: format codebase

* style: format codebase

* refactor: extract skill install dialog footer and content

Extract footer and content sections from SkillInstallDialog and
SkillInstallManagementDialog into separate components for improved
maintainability and clarity of component responsibilities.
2026-08-28 00:59:21 -07:00
Jinwoo Hong 59515beb70 fix(release): recover immutable patch validation gates (#16984)
* fix(release): recover immutable patch validation gates

* test(e2e): locate wrapped terminal file links

* test(e2e): keep sibling file links on one terminal row
2026-08-28 00:55:45 -07:00
Brennan Benson 4bb337741c feat(terminal): weight-layer forensics for the bold-collapse bug (STA-4042) (#16868)
* feat(terminal): weight-layer forensics for the bold-collapse bug (STA-4042)

Field instrumentation to name the writer behind regular-text-renders-bold:
- metric-weight-change crumbs at the writePaneMetricOptions funnel
  (prev/next/reason; weights never change in normal operation)
- terminal-weight-parity-mismatch audit on every visibility resume
- sentinel weightProbe capture fields: live options vs atlas captured
  config vs renderer-buffer bold census
- Cmd/Ctrl+Shift+click unconditional capture (no divergence gate, no
  recovery) for states the missing-ink detector cannot see
- patched addon-webgl ctx.font readback probe: detects failed font
  assignments that rasterize glyphs at a stale weight

* fix(terminal): treat canvas weight-700-serializes-as-bold as a match in the atlas font probe

Found by live validation: Chromium's ctx.font getter normalizes numeric 700
to the keyword 'bold', which made every legitimate bold rasterization count
as a failed assignment (124 false positives in one session).

* chore: update patch hash for the font-probe normalization fix

* fix(terminal): bound bold glitch diagnostics

* fix(terminal): cover serialized WebGL probe state

* feat(settings): hidden staff toggle to arm terminal render diagnostics

Replaces the reserved hidden-experimental placeholder slot with a real
switch (Shift-click the Experimental sidebar entry to reveal). It arms
and disarms the render-desync capture sentinel live — no localStorage
incantation, no reload — for the bold-glitch investigation. The passive
probes stay always-on; only the capture gestures are gated.

* fix(settings): make render diagnostics disarm exact

* chore(settings): rename hidden group to 'Hidden experimental settings', drop its description

* feat(settings): unlock hidden experimental group via Option-click on the Experimental page title

Replaces the Shift-click-sidebar unlock with the Updates-header idiom:
Option-click the Experimental page title toggles the hidden group.
Removes the now-unused click-modifier plumbing from the settings sidebar.
2026-08-27 23:36:44 -07:00
Jinwoo Hong fc8c981103 fix(browser-preview): enforce canonical runtime grants (#16975) 2026-08-27 23:27:39 -07:00
Neil 6c58038f4a fix(session): defer session writes suppressed by a direct-SSH apply (#16969)
The debounced session writer cleared its pending changed-field set whenever the
persist gate was shut, after it had already advanced its identity-based
detection baseline. Because detection is `prev[key] !== next[key]`, a field
dropped there could never be re-detected, so a mutation made during an SSH
apply (up to 30s plus a 1s tail, on every connect and reconnect) was lost for
good — including a closed-tab tombstone, which is exactly the state that stops
the host from resurrecting the tab.

Keep the pending set instead, and add a one-shot wake-up from the gate owner so
a deferred write is not stranded when the suppression tail expires with no
store update behind it.
2026-08-27 22:53:02 -07:00
Neil f5fb60b13a fix(terminal): a restarted relay must not hand a new tab a dead PTY's exit (#16970)
A redeployed SSH relay renumbers its PTY ids from pty-1, so a freshly spawned
PTY can be handed an id a dead one used to own. The renderer's pre-handler
buffer is keyed on that id alone, so `registerExit` found the dead PTY's
buffered exit and reported the brand-new shell as `exitedBeforeAttach` — the
pane never bound a PTY and the tab came up blank forever.

Date every buffered record with a monotonic sequence and fence a fresh spawn
against it: state recorded before the spawn request left the renderer belongs
to the id's earlier owner. State recorded after it is kept, so the real
pre-attach race (a shell that dies instantly, or writes before the pane
registers its handler) still works.

Makes ssh-lost-kill-tab-resurrection.spec.ts:178 pass; it failed in SETUP.
2026-08-27 22:36:29 -07:00
Jinjing 39535a8f1f fix(ui): rename orphan task action (#16976) 2026-08-27 22:32:16 -07:00
Neil 91e7cd088a refactor(terminal): collapse the five inline remote-execution-host PTY idioms (#16967)
`isRemoteExecutionHostPtyId(id)` (= paired-runtime PTY or direct-SSH app PTY,
i.e. "this request crosses a link") was named in #16941 but only used at its own
call site. Five inline copies of the same disjunction remained.

Moves the helper up out of `pty-connection/` — three of its six call sites are
parking/retention/watcher modules that have nothing to do with pty-connection —
and replaces all five copies. Behaviour-preserving: both prefix-form
`isRemoteRuntimePtyId` implementations (`paired-parked-terminal-restore` and
`runtime-terminal-inspection`) are `startsWith('remote:')`, and
`parseAppSshPtyId(x)` returns an object or null so the truthy and `!== null`
forms agree. The negated site keeps its `ptyId !== null` guard (De Morgan on
`!A && !B`), and the retention site keeps `!ptyId` so empty ids still bail.
2026-08-27 22:16:56 -07:00