Commit Graph
59 Commits
Author SHA1 Message Date
db5df2f6f6 perf: index selected-host SSH leases by PTY (#19467)
* perf: index selected-host SSH leases by PTY

* perf(ssh): build lease index lazily and cover pty id normalization

* style(ssh): restore repository oxfmt formatting

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-08 19:41:44 -07:00
814d3525d4 perf: lazily index workspace tabs during terminal binding replay (#19466)
* perf: lazily index workspace tabs during terminal binding replay

* test(persistence): cover same-worktree duplicate tab ids in binding replay

* style(persistence): restore repository oxfmt formatting

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-08 19:40:21 -07:00
c69d4d4f8b perf: track pane alias singleton or ambiguity without copying buckets (#19479)
* perf: track pane alias singleton or ambiguity without copying buckets

* test(persistence): pin pane-alias ambiguity cardinality parity

Covers 0/1/2/3/4 rows per tab plus a mixed ordering case, so a regression from has() to a truthy check would resurrect an ambiguous tab and fail.

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-08 19:39:26 -07:00
53852c9ca4 feat(terminal): make the contrast floor user-configurable (#10754) (#18126)
* feat(terminal): make the contrast floor user-configurable (#10754)

The xterm minimumContrastRatio floor was hardcoded (3 on dark backgrounds,
4.5 on light) and applied to every pane with no way out, so TUIs that use
deliberately low contrast were rewritten: Powerline separators drawn in the
neighbouring segment's background became visible seams, and dimmed secondary
text lost its hierarchy.

Adds an optional `terminalMinimumContrastRatio` setting under Settings ->
Terminal -> Rendering. Blank keeps today's automatic, background-luminance
gated floor; 1 disables correction entirely (matching VS Code's documented
`terminal.integrated.minimumContrastRatio` and iTerm2's off-by-default
Minimum Contrast); values are clamped to xterm's 1-21 range.

The floor is resolved in one place, so live panes, the Appearance preview
and the dashboard terminal preview all follow it, and the existing
value-gated write still avoids clearing xterm's contrast cache on no-op
re-applies. The clamp also lives at the persistence boundary that every
writer crosses, so a hand-edited profile or CLI write can never hand xterm
a non-finite option. Mobile mirrors the desktop gate, so the resolved floor
travels with the terminal theme payload as a new optional field; hosts that
omit it leave older and newer clients on the luminance gate.

Fixes #10754.

Co-authored-by: Nyanako <44753291+Nanako0129@users.noreply.github.com>

* fix(terminal): refresh mobile payload fixture and clarify contrast target

* feat(terminal): make contrast controls intent-based with custom tuning

---------

Co-authored-by: Nyanako <44753291+Nanako0129@users.noreply.github.com>
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-08 00:39:33 -07:00
Neil 225a47533d fix: preserve paired host sessions during startup residue cleanup (#18922)
* fix: preserve paired host sessions during startup residue cleanup

* refactor(persistence): tighten the paired-host retention pass

Dedupe the owner-key -> repo-id extraction the retention and seeding
passes both needed, and name the `runtime:*` check instead of repeating
the parse three times.

Reach the session walker directly by exporting
`addWorkspaceSessionWorktreeOwners` rather than fabricating a
`{ workspaceSession }` state slice to get at it.

Correct the docstrings: `runtime:*` also covers a serving host's own
partition, and the "authoritative removal" they promised has no product
caller on a paired client today, so say what the exemption actually
costs.

Add a survived-load assertion to the explicit-removal test, which
otherwise passed against the pre-fix sweep -- the partition was already
empty before the removal ran.

No behavior change beyond the docs and the test assertion.
2026-09-06 18:29:19 -07:00
Neil 149732df6d perf(persistence): stop the session write re-scanning and rebuilding unchanged state (#18739) 2026-09-04 23:39:02 -07:00
Neil fb69f00b65 fix(hosts): resolve a folder workspace's SSH host from the repo's host, not its raw connectionId (#18598)
* fix(hosts): resolve a folder workspace's SSH host from the repo's host, not its raw connectionId

`resolveFolderWorkspaceHost` inferred a workspace's host by reading
`repo.connectionId` directly. SSH ownership has two spellings on a repo row, and
a row carrying only `executionHostId: 'ssh:<target>'` has no `connectionId` to
read — so it counted as a local repo and the workspace resolved `{ kind: 'local' }`.
That is an execute-here answer for a workspace whose files are on an SSH host,
the #11163 class, and it fires on a well-formed row.

Resolve the host first, then read the target off it. Every other row keeps its
existing contribution, including a `runtime:` row's nested SSH target: that
target is not this client's to dial, but narrowing it here would be a second
behaviour change riding on this one. The runtime branch above still answers
`local`, and now says so — `FolderWorkspaceHost` has no runtime variant, and
widening the type is its own change, not an oversight to be silently corrected.

Three smaller items that stand on their own:

- `resolveWorktreeExecutionHost` gains a `malformed` reason distinct from
  `unknown`. `unknown` (nothing carries the id) is a verdict the launch path may
  legitimately dispose of as a plain local folder; `malformed` (the row named a
  host that cannot be parsed) must fail closed. One word for two situations is
  the shape that lost the distinction in #18006. The strict read is private to
  that module: `getRepoExecutionHostId` stays the answer everywhere else, since
  its fall-through to `local` is harmless for the grouping, label and index
  callers that are nearly all of its ~340 call sites.
- `readAllWorktreeMetaForRepo` / `readWorktreeMetaForRepo` replace four
  open-coded copies of the same host-qualified read (the F7/F8 lockstep shape).
- `getExecutionHostLabel` answers 'Unknown host' rather than 'All hosts' for an
  id that names no host. Showing one unroutable row as though it were on every
  host is wrong on its own terms. Plain English like every other label in that
  module, none of which resolve through the renderer's i18n catalog.

* fix(hosts): resolve the host in candidate selection too, not just in resolution

The first pass fixed how a repo row is classified once it reaches
`resolveFolderWorkspaceHost`. The candidate filter decides which rows reach it at
all, and it read `repo.connectionId` raw as well — so an SSH-only row outside the
project-group subtree was dropped before the new logic could see it, and the
execute-here bug survived for the population the fix was for, via a different
path. Found in review by CodeRabbit.

Three repo-row reads had the same root cause, not one:

- the scope-connection filter, comparing a path repo's raw field against the
  workspace/group connection;
- the group-connection set, built from group repos' raw fields;
- that set's membership test against path repos' raw fields.

The last two are one comparison with the mismatch on either side, so resolving
only the path side would have reintroduced it from the other direction.

All three, plus the resolution loop, now go through one `getRepoScopeConnectionId`
helper. Non-SSH hosts still fall back to the raw field, so a `runtime:` row keeps
contributing its nested target exactly as before.

The new tests use a repo matched only by path, outside the subtree — the
population every existing test missed, which is why four passing revert-tests
did not catch this. One of them is labelled as pinning the resolver rather than
the filter: under the old raw read both rows came back connectionless and matched
each other by accident, so it survives a filter revert and must not be counted as
coverage for it.
2026-09-04 01:34:47 -07:00
Neil 0593f4e0ad perf(persistence): stop writing every worktree metadata row twice (#18451)
* perf(persistence): stop writing every worktree metadata row twice

`setWorktreeMetaForHost` assigns one object to both `worktreeMeta` and
`worktreeMetaByIdentity`, so the profile serialized every metadata row twice.
On a measured 3.64 MB install, 1,347 of 1,349 locator rows were byte-identical
to their identity twin.

The serializer now omits a `worktreeMeta` row the identity map can rebuild, and
the load path rebuilds it — reinstating the shared object reference `JSON.parse`
splits in two. A row is only omitted when exactly one alias claims the locator
and that alias names exactly one identity key, so the rebuild is a pure function
of the file with no winner selection to disagree about.

Omission rather than an in-value sentinel: a downgraded build reads a non-object
`worktreeMeta` value as corruption and deletes that locator's lineage companions
with it. An absent key is a shape every build already tolerates, and it falls
back to the untouched identity map.

* fix(persistence): keep the lineage maps when a profile file has no worktreeMeta key

The rebuild returned `parsed.worktreeMeta` untouched when it was not a plain
record, so an absent key became an explicit `worktreeMeta: undefined` that
outranked the defaults spread. `normalizeWorktreeLinkedItemMetadata` reads a
non-object `worktreeMeta` as corruption and wipes that file's
`worktreeLineageById` and `workspaceLineageByChildKey` with it, then marks the
state dirty so the wipe is persisted. Before this branch the spread supplied
`{}` and the lineage survived.

Also pins the two raw-file readers the projection made load-bearing: the
history GC recovering projected ids from the alias keys (a miss deletes shell
history a live workspace is using), and the profile-transfer read rebuilding
the omitted locator rows (a miss transfers workspaces with no metadata).

* perf(persistence): drop the identity twin, not the locator row

Reverses the projection direction: `worktreeMeta` stays complete on disk and
`worktreeMetaByIdentity[K]` is omitted instead, only when the locator row
regenerates K by construction (`wt2:<hostId>:<instanceId>`) and the two rows are
equal. That removes the format marker, the deliberate-absence list, both raw-file
reader patches and every downgrade hazard, because "alias present, identity row
absent, locator derives it" is a shape every shipped build already heals to
exactly this state.

Keeps 88% of the byte win (541 KB vs 613 KB) and the whole heap-sharing win.

* chore(persistence): keep the derivation predicate module-private

* test(persistence): pin that a file with no identity map never gains one

Answers the review ask that the projection's absent-key contract be asserted on the
bytes, not inferred: a profile whose file carries no `worktreeMetaByIdentity` must
still not have one after a load+flush.
2026-09-03 21:30:45 -07:00
Neil c11c6878c1 perf(persistence): stop dead SSH leases pinning metadata, retire unreachable tombstones (#18430)
Two unbounded-growth fixes in the persisted profile, which is re-serialised in
full on every save.

`collectPersistedWorkspaceOwners` registered every SSH lease's worktreeId as a
live persisted owner with no state filter, so a route-retired lease — the
operator-close `terminated` tombstone, or an `expired` row already marked
`supersededBy`/`relayIdRecycled` — pinned its worktree's metadata row
permanently. The prune gate's own doc names that failure: "Rows pinned by a
persisted session are never removable, so the repetition cannot even make
progress." Reuses `sshRemotePtyLeaseAllowsReattach`, the predicate that already
decides which leases still name a route.

`sshRemotePtyLeases` had no pruning path at all: removal happens in three
explicit places, none age- or state-based, so `terminated` rows accumulated
forever (137 rows / 54 KB on the reported profile, ~38/day from one target).
Marking a lease `terminated` scrubs its pane bindings in the same write, so once
no persisted binding names the id the row routes nothing — reattach, pane
recovery, the orphan sweep, `ssh:reset` and `ssh:terminateSessions` all behave
identically on an absent row. Delete it then, gated on that reachability check
because a lease freezes its tabId and the tab-qualified scrub cannot reach a
pane that was detached into a new tab.

`expired` rows are deliberately untouched, superseded ones included:
`sweepOrphanedRelayPtys` reads those ids as its leave-alone list, so dropping one
would authorize stopping a remote shell that supersession left running on
purpose (docs/reference/ssh-execution-boundary.md).
2026-09-03 20:59:53 -07:00
Neil ab32c2c0c5 perf(startup): stop the persistence milestone from timing its own details closure (#18439)
`logPersistenceStartupMilestone` resolved the lazy `details` closure before
reading `performance.now()`, so the 1.6 MB `JSON.stringify` that
`persistence-load-done` uses to report `workspaceSessionBytes` was billed to the
milestone it measures. Snapshot `t` first.

Diagnostics output is unchanged; only the recorded timestamp moves.
2026-09-03 20:57:57 -07:00
Neil e42c60e8a3 fix(ssh): resolve a pane's binding from the target partition, not the stale local copy (#18546)
One SSH pane accumulated one extra reattachable lease per relay restart (2, 3, 4,
5, 6 across five), and every one of them costs a `pty.attach` round trip on every
later connect, forever. Nothing prunes `sshRemotePtyLeases`, so the fan-out only
grows.

`supersedeSiblingLeasesForPane` is fenced on the PTY the pane is durably bound to,
and `durablyBoundPtyIdForPane` read `state.workspaceSession` (local) before
`workspaceSessionsByHostId['ssh:<target>']`. But `persistPtyBinding(binding, hostId)`
updates ONLY the host partition:

  AFTER-PERSIST  local= ssh:t@@pty2:old:1   host= ssh:t@@pty2:new:1

So for the length of a reconnect the local copy still names the predecessor, the
fence resolved to it, supersession took an already-`expired` lease as its winner,
and returned having marked nothing. Both partitions agree again once the renderer
republishes its layout, which is why the settled store looks consistent and hid
this.

Read both partitions as an ordered list, target's own first, and test the fence by
membership rather than by equality with whichever was read first. Pick the winner
preferring a lease this client still has a route to, since the stale partition
names an expired one. Never retire a lease that is both bound and live, so a
partition disagreement can't strand a running remote process.

Superseded predecessors stay `expired` and are never `terminated`: losing a lease
is not evidence the shell died (docs/reference/ssh-execution-boundary.md). A pane
with no binding is skipped rather than pruned, so a genuine orphan stays askable.

Also re-runs supersession from the binding side after each spawn commit's binding
write, so the lease/binding order at a call site no longer decides, and reconciles
every pane for a target immediately before `reattachKnownPtys` reads the set it
feeds to `pty.attach` — that repairs stores which already accumulated these rows.

The guard suite could not catch this: every assertion bound the pane BEFORE
upserting the lease, an order no caller uses. Rewritten to the spawn commits' real
order (lease, then binding, then the binding-side trigger); it fails 8 assertions
without this change. Added a suite that drives the real `persistPtyIpcSpawnCommit`
rather than the store primitives, including the exact stale-partition state written
by production's own binding writer.

Verified on the Docker SSH lane: five `relay.js` SIGKILLs with recovery between
each, reattachable leases flat at one per pane.

Note: this bounds the reattach SET, not the store. `sshRemotePtyLeases` still has
no cap or TTL and rows still accumulate; pruning is left alone deliberately, since
an `expired` row without `supersededBy` is a genuine orphan and must not be dropped
on age.
2026-09-03 16:47:32 -07:00
Neil 4cc0b8de61 perf(hot-paths): delete allocation-only work in sort, explorer, monaco, rpc, snapshots (#18372)
* perf(hot-paths): delete allocation-only work in sort, explorer, monaco, rpc, snapshots

* fix(perf): revert snapshot revision fast-path — same revision can carry a new session

* perf(hot-paths): drop the unproven rpc buffer rewrite, dedupe the equality helpers

- Revert the unix-socket chunk-carry change. Its comment claimed it avoided
  O(n^2) rescans, but chunks is reset to [remainder] every data event, so the
  join plus the tail byteLength is two passes where the old code did one;
  benchmarks showed no win. It also moved consumed-frame bookkeeping out of the
  closure, so a synchronous throw from the handler would re-dispatch frames.
- project-host-compatibility: fold the two byte-identical array comparators
  into one generic arraysEqualByJson.
- smart-attention: drop the leftover byTab.size === 0 branch that returned the
  same value as the line after it.
2026-09-03 03:20:59 -07:00
Neil b8da193b7a fix(ssh): route the remaining expired-lease readers through the reattach predicate (#18378) 2026-09-03 02:00:42 -07:00
Neil df420285b0 perf(persistence): stop rewriting redundant bytes in the profile store (#18317)
Two kinds of byte in orca-data.json were provably redundant. Both are paid on
every debounced save (full re-serialize) and every launch (full re-parse).

1. The renderer's host split handed one global-field template to EVERY host, so
   local's browserUrlHistory/workspaceDocHistory were copied verbatim into each
   non-local partition. That is a write-side regression undoing #18161's
   load-time drop: the load path removed the replicas, the next full snapshot
   write put them back. Non-local slices now get a template without the fields
   the merge only ever reads off 'local', and both the host write and the
   serializer strip the residue.

2. mergeWorktreeMetaForWrite materializes all ten linked* slots plus
   isArchived/isPinned on every metadata row, so a 1,200-workspace store carried
   ~534 KB of "field":null pairs across worktreeMeta and worktreeMetaByIdentity.
   The serializer omits slots still at their default and
   normalizeWorktreeLinkedItemMetadata re-fills them at load, so in-memory state
   is unchanged.

Measured on a fixture sized like the reporting install (10 hosts, 1,200 metadata
rows, 200 history entries): 1,445,276 -> 643,238 bytes per save (-55.5%),
164,250 -> 3,411 bytes structured-cloned per persistWorkspaceSessionByHost
across 9 non-local hosts, launch JSON.parse 2.00 ms -> 1.37 ms.
2026-09-02 23:19:28 -07:00
Neil 720c3299ba fix(ssh): require a host death certificate before recreating a pane, and unstick expired leases (#18013)
* fix(ssh): match an expired lease on where its leaf lives now, not its frozen tab

A lease freezes tabId at write time, but detachTerminalPaneToTab moves a live
pane, so the stored tab is the one the pane LEFT. getRecentExpiredSshLease
required lease.tabId === tabId, which is wrong in both directions: a viewer on a
stale mirror matched under the abandoned coordinates (and resolvePersistedStable
PaneOwner then reads an empty layout for that tab, so adoptStablePane is skipped
entirely and a fresh shell is spawned over a possibly-live one, binding the same
leaf in two tabs), while a viewer using the pane's real coordinates matched
nothing and got terminal_not_recoverable.

Resolve the leaf's current tab the way restoreReattachedPtyRuntime already does
and compare against that, falling back to the frozen tabId only when nothing can
say where the leaf lives. Both workspace partitions are read because SSH spawns
bind into ssh:<target> while reattach binds into local.

* fix(ssh): let a proven reattach take an expired lease back to attached

#17965 authorized reattach from `expired` but the state machine refused the
transition back, so a lease that reattached and proved itself alive stayed
`expired` forever. That silently exempted a demonstrably running remote shell
from `ssh:reset` (skips `expired`), from the SSH_TERMINATE_RECONNECT_REQUIRED
ownership fence in `ssh:terminateSessions` (marks it not-owned), and from the
quit-time `detached` sweep, and made it permanently ineligible to win
supersession so its own successors never retired their predecessors.

Only the id-qualified caller carries per-pty proof: markSshRemotePtyLeases
AttachedAsync is fed the relay's `attachedLeaseIds`, so an unqualified bulk mark
over a whole target still cannot revive `expired`. `terminated` stays absorbing.
Re-entering `attached` drops supersededBy/relayIdRecycled, since route
retirement belongs to the shell that lost the pane and this one just proved it
is not that shell — the same invariant upsertSshRemotePtyLease enforces.

* fix(ssh): make the pane-recovery liveness gate refuse without positive evidence of life

The gate refused only `live` and `unverifiable` and passed on `null` — but the
register is an in-memory Map, so `null` is equally what a fresh app start, a
never-asked host and a certified death look like. Absence of evidence was
reading as authorization to spawn a shell over a possibly-live remote process:
`!pty.connected` is cleared for every PTY a dropped relay owned, and `expired`
only ever says the CLIENT lost its route.

- `exited` is now RETAINED rather than deleted, so the register is three-valued
  in the map as well as in the type. Its one writer is a host-delivered exit
  frame — an exit with a real code, or an explicit `hostExitConfirmed` — which
  records the certificate instead of merely dropping the doubt.
- `recoverTerminalPane` refuses on `live` and `unverifiable`, and deliberately
  does NOT demand a positive `exited`. The only answer that ever reaches this
  gate is a reachable relay reporting no such id, and that is a union: pty.attach
  throws not-found for an unknown id with no liveness check, and a relay restart
  makes every previously minted id unknown (ids carry a per-start
  `ptyIdMintEpoch`). No writer of `exited` co-occurs with a reattachable
  `expired` lease either — a host-delivered exit frame tombstones the lease
  `terminated` — so requiring one would close the gate permanently.
- `handlePtyReattachFailure`'s not-found branch publishes `code: -1` to the
  renderer and does not call `runtime.onPtyExit`. The relay's not-found answer is
  not a death certificate, and #17963's ratchet on the same branch pins that.
- The inventory's `observed === false` hunk keeps dropping doubt rather than
  asserting a death: `pty.listProcesses` returns the relay's CURRENT session map,
  so a restarted relay omits every previously minted id whether or not those
  shells died — the same union, one hop away.

A live or unprovable pane refuses; a disowned one still recovers. No wire change.

The gate's ratchets live in terminal-pane-recovery-liveness-gate.test.ts:
config/vitest.config.ts — the config CI runs — matches only `*.test.ts`, so cases
placed under orca-runtime-tests/*.spec.ts would never execute.

* fix(ssh): gate paired-viewer pane recovery on the narrowed session-gone predicate

isSshSessionGoneError landed on the IPC transport, which never calls
terminal.recoverPane. The one caller that does — recoverExpiredHostPane in the
paired-viewer transport — still triggered on a bare SSH_SESSION_EXPIRED
substring, so the identity-mismatch reply (the relay found a LIVE PTY under that
id owned by another pane, which is evidence of presence) still asked the HUB to
replace the pane, putting a second agent on one transcript. Main already refuses
the respawn on that same reply; this makes the two agree.

A pane whose shell genuinely died is unaffected: plain SSH_SESSION_EXPIRED still
matches. The mismatch reply now surfaces as an error instead of a respawn.

* test(persistence): update the reattach ratchet for expired-lease reclaim

markSshRemotePtyLeasesAttachedAsync is id-qualified, so a named pty that
proved itself alive now returns to attached instead of staying expired.
2026-09-02 22:31:59 -07:00
Neil 08c7152ab6 fix(ssh): compare a lease's relay pty id against the pane's app id (#17969)
`getRecentExpiredSshLease` compared the stored lease ptyId (relay form,
written through `toStoredPtyId` -> `toRelaySshPtyId`) raw against the
runtime's app-form `pty.ptyId`, so `'pty-3' === 'ssh:target@@pty-3'` never
held and `recoverTerminalPane` refused every real SSH pane. Normalize with
the same tolerant helper the binding reader already uses, now shared as
`toComparableRelaySshPtyId`.

Switching the path on is only safe on top of #17957 (respawn gated on the
runtime liveness verdict), #17965 (`expired` no longer withdraws bindings)
and #17966 (supersession and id recycling carry their own marks).
`recoverTerminalPane` additionally refuses a lease those marks disqualify,
so it acts only on an `expired` lease that means "reattach gave up".

The path's outcome is a reattach, not a respawn: `createTerminal` calls
`adoptStablePane` first, which attaches attach-only to the retained
binding and only falls through to a fresh shell once the host itself
answers that the PTY is absent.
2026-09-02 21:48:17 -07:00
Neil f2e95e7860 fix(ssh): separate a superseded lease from an orphan so reattach can tell them apart (#17966)
`expired` was one word doing two unrelated jobs — "a newer lease won this pane"
and "reattach lost contact" — so `reattachKnownPtys` had to exclude all of them.
That kept the 2 -> 19 -> 20 fan-out fixed at the cost of never bulk-reattaching a
genuine orphan; those recovered only through the slower `adoptStablePane` path.

The blocker cited in #17965 does not apply. The STA-3077 note guards
`upsertSshRemotePtyLease`'s match against a RECYCLED `pty-N` after a relay
restart. `supersedeSiblingLeasesForPane` is a different path and already holds
`winner.ptyId` when it expires a predecessor, so recording which lease won needs
no relay-start identity.

`SshRemotePtyLease` gains two optional marks, each meaning exactly one thing:

- `supersededBy` — the winner's stored-form ptyId, written only by supersession.
- `relayIdRecycled` — written only by the pending-stop replay's
  `relay-id-recycled` retirement. That retirement wrote `expired` *purely* to
  keep the lease out of the reattach that runs one step later ("hands the user's
  old pane to whatever process now holds the recycled id"), and the reattach
  fences on paneKey/tabId, never on incarnation. Relaxing the filter without
  this would have silently reopened that hole.

Bulk reattach now skips a lease carrying either mark and re-adopts the rest, via
one shared `sshRemotePtyLeaseAllowsReattach`. `terminated` is untouched.

Recycled-id safety: both marks are dropped whenever the id is re-upserted
`attached`/`detached`, so a relay that renumbered onto a new shell cannot inherit
its predecessor's mark. Supersession also stamps an ALREADY-expired predecessor
for the same pane — same evidence, and it is what bounds the reattach set, since
otherwise every past orphan for that pane would stay reattachable forever. Its
`updatedAt` deliberately stays put: bumping it would make a stale lease look
recent to `getRecentExpiredSshLease`.

Persistence: the lease loader is a strict whitelist, so both fields are named in
`normalizeSshRemotePtyLease` or they would be stripped on every launch. Absence
reads as "orphan", which is the only thing an older build could have meant, and
an older build ignores keys it has never heard of (remote-wire Rule 1).
2026-09-02 21:39:08 -07:00
Neil e4c279fa60 fix(ssh): let an expired lease reattach its orphan instead of stranding it (#17965)
* fix(ssh): let an expired lease permit a reattach instead of unbinding the pane

`expired` never means the remote shell exited. Every writer records that the
CLIENT lost its route — a superseded sibling, a recycled relay id, a
persistPtyBinding refusal made *after* pty.attach proved the shell alive, a
failed reattach indistinguishable from a relay restart, a relay reset whose
kill may not have landed. docs/reference/ssh-execution-boundary.md grades all
of those `unverifiable`.

Three readers treated it as death, and together they made the pane unable to
reach a process that is still running:

- `isRestorablePtyBinding` / `hasRestorableSshRemotePtyLease` refused to replay
  a durable binding a renderer snapshot had omitted.
- `markSshRemotePtyLease(s)` wiped the persisted pane->pty binding, which is
  what makes `resolvePersistedStablePaneOwner` return null, `adoptStablePane`
  give up, and `createTerminal` cold-spawn a replacement. The user's terminal
  comes back empty and the running job is orphaned and invisible.

Only `terminated` now withdraws a binding: it is the operator-close state
(`ssh:terminateSessions`) and the one written after a host-acknowledged stop.

This authorizes a reattach ATTEMPT, never a respawn, so #17957's gates are
untouched and in fact fire less often — where the pane previously went straight
to a fresh spawn it now attaches first. A genuinely dead shell still converges:
`attachStablePaneOwner` retires the binding on `isPtyAlreadyGoneError` (the
relay's own absence answer, not a message match) and falls through to a fresh
spawn, so no pane retries forever.

Supersession keeps its own binding scrub in `supersedeSiblingLeasesForPane`,
where a NEWER lease for the same pane is the evidence — the 2 -> 19 -> 20
reattach fan-out stays fixed.

* test(persistence): split SSH remote PTY binding partition cases into their own file
2026-09-02 21:35:28 -07:00
Jinjing c2fce80289 Fix agent dashboard setting configure (#18245)
* Make agents activity always-on; toggle via bell icon

- Remove optional showAgentsSidebar setting
- Replace sidebar view-toggle with bell-button for activity access
- Agents activity now always accessible in sidebar
- Preserve migration flag for introduction to existing users
- Remove visibility inference utilities

* Simplify sidebar when agents view active: hide workspace options, add to

- Hide workspace options menu and add project button when agents view is
  active, reducing UI clutter in that mode
- Add tooltip to the activity bell button for better discoverability
- Localize sidebar search field text
- Move search and filter toggles to local state in SidebarAgentsList,
  removing unused callbacks from thread list components
- Manage search input focus properly when opening
2026-09-02 13:43:06 -07:00
Jinjing e3de6b2ce8 Add automation runs dashboard with pagination and filtering (#18226)
* Add automation runs dashboard with pagination and filtering

Adds a new Runs view in the Automations page that lets users browse all runs across automations with status/host filtering, search, and pagination support. Includes virtualized table rendering for efficient handling of large run histories and summary cards showing 24h/7d success/failure counts.

* Fix missing dependencies in useCallback hooks and imports

Missing dependencies in useCallback can cause stale closure bugs. This
adds missing state setters to dependency arrays and consolidates type
imports for consistency.

* Use keyset pagination for stable automation runs pages

Pagination now uses createdAt:id boundaries instead of offsets, so new
runs arriving between pages don't shift the window. Maintains backwards
compatibility with legacy offset cursors.

Move pagination to shared module, fix outcome counting for future-dated
runs, and improve hook state tracking on authority re-pairing or target
changes.

* Extract automation run details to top-level page view

Moves run display from detail pane to dedicated page, establishing
three-level navigation (Automations → Runs → Run Details) and simplifying
the detail pane component.

* Fix pagination stability when automation runs share createdAt

- Define a stable total order with createdAt and id tiebreaker to prevent runs tied on createdAt from being dropped when the boundary run is pruned between page requests
- Retain cursor on failed pagination so pages remain retryable
- Update ownerNotice type to AutomationActionNotice

* Extract automations list panel and worktree map logic

Split AutomationsPageSurface into smaller, focused modules for better maintainability and reusability. Move list panel UI rendering to AutomationsPageListPanel component and worktree map selection logic to a standalone utility function.

* Add i18n strings for automation runs dashboard

Adds localized strings for the automation runs dashboard view, including search, filtering by host and status, run counts for 24h/7d windows, and empty state messaging across all supported languages.

* fix missing translation

* fix missing translation
2026-09-02 13:42:14 -07:00
Neil 084dbbc3b3 perf(persistence): build the state file once per save instead of seven times (#18161)
Every debounced save stringified the full persisted state, then ran two
`String.replace` passes per secret sentinel — one for the on-disk payload, one
for the guard hash. Each replace returns a rope the next one has to flatten
before it can search, so three sentinels cost seven flattened copies of a
4.65 MB state (a two-byte V8 string, ~8.9 MB each), and the state was then
UTF-8 encoded twice more: once inside `sha1.update(string)` and again inside
`handle.writeFile(payload, 'utf-8')`.

`applySecretSentinelSubstitutions` walks the state once with a single
alternation regex, encodes each literal run to a Buffer exactly once, and feeds
those same buffers to both the payload and the hash. Measured on the author's
4.65 MB store with three live secret slots: 48.8 MB -> 17.9 MB allocated per
save, 26.6 MB -> 0 of large_object_space churn, and 22.1 -> 15.1 ms (min) /
32.3 -> 16.9 ms (median) for build+hash+encode. Bytes on disk and the guard
hash are proven identical to the previous loop.

Separately, non-local host session partitions carried stale replicas of the
`browserUrlHistory` global — 589,807 bytes, 12.7% of the file — that neither
the split (which writes globals only to 'local') nor the merge (which reads
them only from 'local' unless local has none) can ever reach. The load path now
drops them when the local slice already holds the field. Only the two history
globals are dropped: the rest are read out of every partition by the worktree
ownership sweep or the mobile/runtime projections.
2026-09-02 13:05:16 -07:00
Jinjing 61e010079f New agent dashboard (#18222)
* more obvious toggle

* more obvious toggle

* feat(activity): redesign thread rows and add child agent filtering

- Emphasize task title and last activity in row layout over metadata
- Add child agent toggle; hide orchestration workers by default
- Support collapsible groups and ungrouped view mode
- Improve orchestration worker message handling to surface replies
- Add sidebar search and filter controls for agent activity

* periodic checkin

* feat(activity): add "Clear completed" action and performance improvement

- Add "Clear completed" action for activity threads with undo window; clears completed and interrupted rows from view, persists across restart
- Virtualize activity thread list to render only viewport-bounded rows
- Cache activity thread search text to prevent recomputation on every keystroke
- Cache dashboard bucket counts per-worktree for selective invalidation on unrelated changes
- Use useDeferredValue for activity search filtering to keep input responsive
- Make compact mode the default display for activity threads
- Add activity-cleared-at persisted state tracking (per-pane cutoff timestamps)

* improve style

* minor change

* feat(activity): add persisted host and project filters to agents view

Agents scope filters are deliberately separate from workspace-nav filters so a monitoring surface never inherits workspace context silently. Filters survive restarts and always display an active-filter chips row with hidden count, making filtering visible and reversible.

* Graduate Agents view from experimental, refine activity handling

- Agents Dashboard moves from experimental to standard feature with showAgentsSidebar setting controlling visibility
- Add identity-checked cache eviction (dropPersisted IPC) to prevent newer runs from being evicted when UI clears older status, fixing clear-completed safety
- Extract ActivityThreadHoverCardSummary and ActivityThreadListToolbar components for better organization and reusability
- Implement mark-thread-read as separate action from select with clickable bell icon
- Add hasActivityThreadWorkspace helper for checking workspace availability across hosts (SSH/runtime targets)
- Preserve scope filter array identity during hydration for memo optimization
- Track manually-unread turns in auto-ack to prevent re-acknowledgement
- Clean up activity cleared-at cutoffs on pane retirement
- Remove activity-thread-hover-card max-lines lint override (code refactored below threshold)

* Refactor agent cache identity to use timing fields only

- Simplify AgentStatusCacheIdentity: keep only paneKey, receivedAt, stateStartedAt
- This fixes silent no-ops where renderer-enriched fields diverged from main's cache
- Add worktree-jump-navigation for navigating activity to workspaces
- Add manual mark-unread protection separate from auto-ack
- Optimize activity owner resolution with per-build memoization
- Optimize detected worktree lookup with indexed search

* Remove sticky header, add scroll position persistence

Replace the floating sticky header overlay with scroll position memory via
a ref. This preserves the user's scroll location when switching between
threads or remounting the agents list, improving UX without requiring
React state.

* Implement sticky group headers in activity thread list

Keep group headers visible at the top while scrolling when threads are grouped. Headers stick to the viewport while their section is in view, then unstick as the next header approaches.

* add blue flash

* update settings appearnce

* Extracted activity acknowledgement/clearance actions from the oversized UI slice.
  - Removed dead sidebar search/menu props and the unused search ref.
  - Removed the unnecessary sidebar visibility bitmask.
  - Replaced hardcoded sidebar toggle colors with design-system tokens.
  - Removed duplicate “mark all read / clear completed” controls in the sidebar.
  - Preserved manual-unread state correctly across pane retire, transfer, and drop.
  - Made clear-completed cutoffs monotonic so clock skew cannot resurrect old activity.
  - Fixed blank workspace names in hover cards with the existing fallback helper.
  - Added missing localization entries and stabilized hydrated filter array identity.
  - Updated misleading Agents setting copy to describe both sidebar surfaces.

* add onboarding guide for the new agents panel

* Add activity clearance tracking and synced agent view settings

Agent view filters and presentation settings now sync across paired clients.
Preserves per-pane activity clearance cutoffs in persistent state. Improves
activity thread row accessibility with proper ARIA roles, and preserves
terminal host ownership after pane teardown via retained terminal handle.

* rm html

* Graduate Agents from experimental and improve activity visibility

- Migrate `showAgentsSidebar` setting from legacy experimental flags; default new profiles to the agents sidebar
- Replace scoped-thread filtering with visible-thread filtering so bulk actions (mark all read, clear completed) only affect rendered rows
- Rewrite child agent classification as a set of visible pane keys to fix orphan promotion and parent-cycle handling
- Improve activity cleared-at cutoff lifecycle: preserve on row dismissal (pane may still be live) but clear on pane removal
- Add pagehide flush for pending clear-completed evictions so quit/reload cannot replay cleared activity
- Polish agents sidebar: unread count badge, expand button, onboarding intro for migrated/new users
- Extract shared time-ago formatting to a library module
- Fix scroll restoration to defer until content can contain the saved offset
- Improve stable message hold for compact agent rows using state instead of refs
- Add worktree filter-visibility check to distinguish collapsed-but-unfiltered from filtered-hidden

* Graduate Agents from experimental and improve activity visibility

- Remove the deprecated full-page Agents view; fix settings navigation fallback
- Refactor bulk action bindings and separate mark-all-read from visible threads
- Preserve sidebar collapse state across remounts; fix child-agent badge filtering
- Add safety window for scroll-restore and improve worktree host-qualified filtering

* Graduate Agents from experimental and add manual unread tracking

- Move Agents sidebar from experimental settings to standard feature with intro flow
- Add persistent manual unread turn tracking for activity feed
- Consolidate workspace activation through activateAndRevealWorkspace dispatcher
- Improve sidebar view toggle with radio semantics and arrow-key navigation

* Graduate Agents sidebar and separate dashboard experiment

The Agents tab now has its own `showAgentsSidebar` setting (defaults on) independent from the dashboard popout experiment. Activity unread counting is simplified to count all events uniformly without mode-specific filtering. Dashboard visibility is now controlled solely by `experimentalAgentDashboardPopout`, with its own UI in the Experimental settings pane. Migration path updated: only `experimentalActivity=true` graduates to the sidebar; the dashboard experiment remains separate.

* Add agent-session tab support to activity tracking

Build activity event contexts from structured agent-session tabs and
worktree-attributed status entries. When activating a thread, try
agent-session tab activation before falling back to terminal pane.

* • The workspace sidebar tab is now a static Spaces
  label—no grouping-based “Projects” label or hidden
  width-reservation span.

* Show unread count badge and prioritize attention-needing agent threads

Activity group order now surfaces threads needing attention (blocked,
waiting, interrupted) before working/done so they're never buried. The
Agents tab shows an unread count badge while viewing Spaces, since the
open Agents list already highlights unread rows.

Also improves UX text ("Hide Agents" vs "Maybe later"), accessibility
with proper ARIA labels, and handles edge cases: preserves read state
for retained panes on SSH reconnect and handles deleted worktrees
gracefully in navigation.

* Batch agent-status evictions and optimize activity pane rebuilds

- Add dropPersistedStatusEntries batch API; consolidate evictions into one persist
- Implement fallback timeout in clear-completed for unseen toast callbacks
- Project only activity-relevant tabs; memoize terminal tab derivations
- Stabilize activity virtualizer key to prevent unnecessary item measurements

* Remove unread count badge from Agents sidebar tab

Simplify useActivityUnreadCount by removing the enabled parameter and
conditional logic, as the badge is no longer displayed in the UI.

* Deduplicate activity unread counts across source overlaps

Live pane status is the primary source; retained and migration entries
serve as fallback caches that may briefly overlap it during lifecycle
transitions. Count each pane only once by tracking seen keys, prioritizing
the live status as the canonical source.

Also fix monitoring state display: it's a distinct agent state, not a
tool-running row state, so exclude it from tool preview checks.

* Update activity pane tests to remove unread badge assertions

- Remove ActivityPaneVisibility type and readActivityPaneVisibility() helper
- Update agentsSidebarButton selector to match badge-less state
- Simplify assertions to check pane focus instead of visibility isolation
- Remove test for unread badge acknowledgement flow

* Fix activity pane workspace resolution and localization handling

- Thread defaultHostId through activity operations for correct host resolution
- Add language-aware caching for standalone terminal names with cache invalidation
- Fix scroll restoration bounds calculation for tall viewports
- Add focus management to sidebar radio group keyboard navigation
- Refresh localized sidebar content on language changes
- Preserve activity state across heartbeats to prevent history loss
- Improve host-id strictness in worktree jump navigation

* Preserve activity view when settings fetch fails

A failed window.api.settings.get() leaves settings null, which was
incorrectly treated as opt-out. Add the missing null check so the
activity-view gate only applies when settings are available.

Includes tests for this scenario and related edge cases in keyboard
navigation, worktree jumping, and session state handling.
2026-09-02 11:00:24 -07:00
Neil f9db653e14 perf(worktrees): gate worktree metadata hygiene on evidence, not on every listing (#18034)
* perf(worktrees): gate worktree metadata hygiene on evidence, not on every listing

Dangling `worktreeMeta` pruning rode the detected-worktree listing, a polled read
path. Each pass captured a prune expectation over the repo's whole metadata table
(a JSON.stringify per row) and then stat'd every path-missing candidate. Both are
O(all rows), and most rows are refused anyway — pinned by a persisted session, or
structurally unremovable on this host — so the work repeated forever without
converging, pinning the main process in fs completion callbacks (#17775).

Three changes, no behavior lost:

- Probe only rows a delete could still accept. Session ownership and structural
  removability are pure functions of persisted state, so deciding them before the
  filesystem inverts the cheap and expensive halves. The filter is advisory; the
  authoritative checks are unchanged, so it can only shrink the stat fan-out.
- Extract `isLocallyRemovableWorktreeMetadataRow` so probe-avoidance and the
  delete share one definition of removability.
- Gate the metadata + lineage prune on evidence instead of the listing: a worktree
  lifecycle event, a mutation that can make a row more removable (session-owner
  release, metadata removal, SSH lease release, automation run finishing or
  deletion, repo deregistration), or a git listing that differs from the one the
  last pass ran against. With none of those the pass is a provable repeat and is
  skipped, so a quiescent app does no hygiene work at all.

The gate deliberately ignores metadata writes that only add or update a claim:
the listing path itself stamps metadata, so re-arming on those would restore the
storm. A missed signal leaves a row in place until the next one; nothing is
deleted that would not have been deleted anyway.

* refactor(worktrees): fold repo prune-gate teardown behind one call

Merging both import blocks during the rebase pushed the file past the
300-line budget. The two calls are one intention -- retire this repo's
gate state on a full removal, and re-arm the shared inputs either way --
so name that in the module that owns the gate.
2026-09-01 22:47:29 -07:00
Neil 2f105b23d1 fix(persistence): sweep sleeping-agent-only residue and stop mis-seeding orphans
Review found three holes in the load-time sweep.

`sleepingAgentSessionsByPaneKey` and `terminalSurfaceTombstonesByPaneKey` are
pruned by the worktreeId they name, not by their own key, but
`pruneWorktreeStateForRepo` only collected owner keys from `worktreeMeta` and
`lastVisitedAtByWorktreeId`. An orphan whose only residue was a sleeping agent
therefore survived the sweep and re-seeded it on the next load, so the store
never self-cleared and every launch scheduled another save. Collect owner keys
from those records too, which fixes `removeProject` for the same shape.

`ownerKeyBelongsToRepo` is restored to its original body. Reordering its two
readings was not behavior-preserving as claimed: for a repo named `folder` or
`worktree`, checking the workspace-key reading first flips the result. The
census now uses `ownerKeyWorktreeIds`, which returns both readings, and seeds
only when neither names a live repo -- seeding one reading of a key whose
other reading is live would hand the removal pass a live row to delete.

Seed from `activeWorktreeId`, `activeWorkspaceKey` and
`activeWorktreeIdsOnShutdown`, which are pruned by bespoke rules and so were
reachable by no owner-key loop, and record why
`terminalTopologyRevisionByRepoId` stays excluded.

Refs #17776
2026-09-01 17:20:17 -07:00
Neil a2aea5d0b0 fix(persistence): sweep rows owned by deregistered repo ids at load
Deregistering a project stranded every row it owned. Each pruning path is
gated on the repo still being in `state.repos`, so once an id leaves the
catalogue its metadata, identity aliases, lineage and session rows became
unreachable forever -- and on a paired client they rendered as phantom
worktrees under an "Unknown" project.

Reconcile against the repo catalogue on load instead: any repo id that owns
rows but is absent from `state.repos` has its rows removed through the same
path `removeProject` uses. Host-independent and session-independent, because
an orphan has no owner that could object -- which is also why this reaches a
client's mirror of a remote host's session partition, something no local
removal can do.

Only a full `<repoId>::<path>` locator seeds the orphan set; bare keys can be
folder workspace ids or repo-keyed revisions, and guessing wrong there would
delete live state. `retiredWorktreeNamesByRepo` is deliberately untouched so a
re-added repo cannot reissue a name onto a cwd that still holds a prior
occupant's agent state.

Test fixtures that wrote worktree rows without registering their repo were
relying on orphans surviving a reload; they now register the repo they name.

Refs #17776
2026-09-01 17:20:17 -07:00
NeilandBrennan Benson fbe94ceff6 fix: close readiness gaps found by merged-change audit (#17159)
* fix(ssh): fence stale kills and retired pane replay

* fix(ssh): support cancellable interactive authentication

* fix(ssh): await remote catalog before snapshot adoption

* fix(pty): contain Windows ConPTY input failures

* fix(power): avoid redundant macOS display blocking

* perf(editor): narrow markdown override subscriptions

* fix(quick-open): close directory handles after reads

* refactor(linux): remove unused proc socket scanner

* fix(usage): apply flat Sonnet 4.6 pricing

* ci: prime Node next native test cache

* docs(skills): resolve snapshot cleanup data path

* fix(ssh): recover install locks after host reboot

* test(ssh): recognize boot-aware install locks

* test(ssh): prove previous-boot lock recovery live

* test(wire): pin pre-metadata release coverage

* fix(terminal): preserve remote tab ownership through recovery races

* test(runtime): fence replaced terminal handles in agent guard

* fix(ssh): preserve remote snapshot authority across polls

* fix(pty): contain late ConPTY output EPIPE

* test(pty): register Windows exit watcher before kill

* fix: close SSH and tab readiness race gaps

* fix(tabs): retain headless order and placeholder titles

* fix(build): avoid parallel electron-vite config race

* test(windows): avoid MSYS temp path rewriting

* test(windows): avoid killing exited PTY

* fix(pty): avoid late ConPTY input teardown race

* fix(terminal): sync reconnect error ownership after commit

* fix(runtime): use canonical worktree identity comparison

* test(ssh): assert complete cold-hydration baseline

* test(windows): invoke quoted retention fixture via PowerShell

* test(windows): read ConPTY grid through mode con

* fix(terminal): publish PTY replacements atomically

* fix(terminal): infer stale identity on reattach

* fix(terminal): fence stale pane PTY callbacks

* fix(terminal): fence stale pane binds after rebind

* fix(terminal): reject stale pane transport callbacks

* fix(terminal): fence mirrored reattach spawn callbacks

* fix(terminal): replace stale pane PTYs on remount

* fix(ci): size the Windows launcher-compile test budget from measurement

`native-smoke (windows-latest)` fails ~4.5% of runs on
`preserves a multiline argument through the compiled remote launcher`
with "Test timed out in 15000ms" — on unrelated PRs, for reasons that
have nothing to do with them. Across 176 sampled attempts it is the only
red that job produced, and it hit seven different PRs in two days:
#16900, #16904, #16915, #16955 (twice), #16979, #17014, #17085.

The test is six process creations: powershell.exe forks csc.exe, then
the freshly compiled orca.exe forks node.exe, twice. Hosted Windows
runners periodically slow process creation down, and this test amplifies
that far harder than anything else in the job. Comparing the 80 attempts
where it ran under 3s against the 12 where it ran over 12s, its own
median goes 2198ms -> 15917ms (7.2x) while the same file's
powershell-only test moves 556 -> 686ms (1.2x), the cmd.exe and Git Bash
process tests in the neighbouring file move 1.4x, and the other 35 files
put together move 1.5x.

Measured across those 176 attempts: 1881ms to 35438ms, p50 4264ms,
correlation +0.881 with the job's total Vitest duration. 8 of 176 (4.5%)
exceeded the 15s cap; 2 of 176 (1.1%) also exceeded the shared 30s
testTimeout, so deleting the override and inheriting the config is not
enough on its own. 60s clears all 176 with 1.7x headroom on the worst.

This is slow, not hung. Every body here is synchronous spawnSync, so
Vitest cannot interrupt one — the timer fires only after the body
returns and the reported duration is real elapsed time. That is why a
failure reads `× ... 22464ms` under `Test timed out in 15000ms`. The
work finished; the stopwatch was short. Seven reruns at one identical
head measured 2053 / 4680 / 5551 / 8732 / 13506 / 14868 / 21937ms — the
last of those would have been red on code that had not changed.

The 15s came from #8897, which raised this test off Vitest's built-in 5s
default because the job then ran bare `pnpm vitest run`. #8909 landed
3h27m later and pointed the job at config/vitest.config.ts, which is the
real fix for that. The constant stayed behind and has been the binding
budget ever since.

* fix(terminal): fence stale remount reattach ownership

* fix(terminal): reconcile mounted pane identity after replacement

* fix(terminal): fence stale reattach fallback ownership

* fix(terminal): fence deferred SSH reattach ownership

* fix(terminal): fence stale split pane ownership callbacks

* fix(terminal): keep stale spawns from consuming startup

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-31 08:17:40 -07:00
OrcaWinandm4air b1f5d2dd2a fix(agents): preserve manual mode for newly added defaults (#17515)
* fix(agents): preserve manual mode for newly added defaults

* test: handle optional migrated settings fields

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-08-30 18:38:25 -07:00
Brennan BensonandMerge Sim f2e9ba453c fix(agent-hooks): route reminted pane keys to canonical identity (STA-3993) (#15714)
* fix(agent-hooks): route reminted pane keys to canonical identity (STA-3993)

Spawn was stripping $$<base32>:L$$ ORCA_PANE_KEY values (and the launch
token) instead of rewriting them to the metadata-proven tab:leaf key, so
OMP hooks never entered last-status.json and sleeping rows stayed working.

Alias that exact remint form onto the canonical pane so later posts still
route, and keep unmatched tokens from stamping another pane.

* fix(agent-hooks): keep reminted pane-key aliases first-pane-wins

Remint tokens have no embedded tab identity, so a later spawn that reused
the same $$ token with a different tab/leaf was overwriting the alias and
routing leftover hook posts onto the new pane. Refuse destination changes
for that form while still allowing same-pane pty id updates.

* fix(agent-hooks): keep pane alias limit import valid after refactor

* fix(agent-hooks): bound pane alias destination keys

* fix(ssh): keep pane identity env stripped when hooks disabled

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-30 17:54:13 -07:00
Brennan BensonandMerge Sim c3aceacc7b Fix PR unlink for auto-detected reviews (#16898)
* fix: make PR unlink hide auto-detected reviews

* Type the empty-content test double against the real model

The literal narrowed suppressedGitHubPR to number and typed the callback
as Mock, so neither direction was comparable and tsconfig.tc.web.json
failed on TS2352. Keeping the 'as' cast preserves checking of the fields
the double does supply.

* Add localization keys for the unlinked checks-panel state

The unlinked title, relink action, and the remote-runtime upgrade notice
introduced untranslated keys that static analysis requires in en.json.

* Advertise PR suppression capability in the transport test

The client capability list is pinned by websocket-transport.test.ts, and
adding WORKTREE_GITHUB_PR_SUPPRESSION left the expected list stale.

* Fix stale PR suppression in Checks

* fix: harden PR unlink suppression state

* refactor: extract PR unlink state handling

* fix: show PR relink recovery in source control

* fix: add unlinked PR localization

* Clarify workspace-scoped PR unlinking

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-30 12:24:51 -07:00
Chen 07df4bf0be fix(orchestration): recover stripped task deps 2026-08-29 23:07:36 -07:00
Neil 7ae916cebd perf(worktrees): batch-prune stale local metadata (#17278) 2026-08-29 16:06:15 -07:00
Neil 80a0595449 perf(startup): stop rescanning persisted pane tabs (#17179)
* perf(startup): index persisted pane tabs once

* perf(startup): lazily resolve persisted pane tabs
2026-08-29 16:01:23 -07:00
Neil bf884ea866 perf(startup): skip disabled diagnostic serialization (#17176) 2026-08-29 16:01:20 -07:00
Neil 7bbb8adc61 fix(ssh): replay an undelivered remote PTY stop on the next handshake (#12447 item 1) (#17011)
* fix(ssh): replay an undelivered remote PTY stop on the next handshake

A pty.shutdown that dies on the transport left the remote shell running
forever: kill.ts marked liveness unverifiable and nothing retried.

Record the undelivered stop on the existing durable SshRemotePtyLease and
replay it against the authoritative host on the next handshake to that same
target, fenced by the host-minted PTY incarnation so a replay cannot kill a
later PTY that reused a recycled pty-N id. Retire the record on confirmed
delivery, on the host reporting the PTY absent, and on a bounded TTL.

No wire change: the fence reads incarnationId, already published on
pty.listProcesses. A host that does not publish it degrades to no replay.

* fix(ssh): do not leave a replayable kill order behind a reversible stop

Worktree sleep stops through stopAndWait and marks those stops reversible;
when one does not land the pane stays live and the user keeps using it. An
order recorded there would come back on a later handshake and kill that
terminal. Only killPtyFromRuntimeController — where the client gives the PTY
up for good — records one, and it skips any PTY a reversible stop owns.

* fix(ssh): cover the renderer kill route and harden the replay's evidence

pty:kill is a separate implementation from killPtyFromRuntimeController and
is the one an ordinary tab close reaches, so the record was never written on
the path #12447 describes. Extracted it out of inspect.ts (which was over the
line budget and was not what the file is named for) and wired both branches.

Also:
- finishPtyShutdown no longer retires the order. It runs on paths that asked
  the host and on paths that never did, so retiring there was a contract every
  caller had to know, and the one that forgot silently dropped a kill order.
  Retirement is the replay's, on inventory evidence only.
- A recycled relay id now expires its lease. Declining to kill was only half:
  reattach fences on paneKey/tabId, never incarnation, so an untouched lease
  bound the user's old pane to whatever now holds the id.
- Dropped isPtyAlreadyGoneError from the tombstone path. It matches message
  text a transport failure could wear; every tombstone now traces to a listing.
- TTL is owned by a durable prune that actually deletes, not by a branch that
  was unreachable behind the read filter and only looked tested.
- The replay re-reads the inventory per wave and re-checks the fence next to
  each shutdown, and can never reject into the connect path.
2026-08-28 04:08:15 -07:00
Jinwoo Hong cb848647e5 fix(browser-preview): require explicit preview capabilities (STA-5758) (#16921)
* fix(browser-preview): require explicit preview capabilities (STA-5758)

Scope document reads to approved directories, confirm external links before opening them, revoke grants with tab lifecycle, and keep document-preview session state rollback-safe across mixed client/runtime versions.

* Harden document preview lifecycle and permissions

* Document preview DNS prefetch residual

* Make preview E2E guest focus explicit

* fix(browser-preview): entry-file-only authority for root-level docs, contained chip layout, re-issued gate paths (STA-5758)

A grant whose document directory is its own request base — a doc at the
workspace root, or outside any workspace — now reads nothing but the entry
file until the reader approves a directory, at both the lexical and the
canonical containment pass. The DNS-prefetch residual can only beacon what
the page can read, and a root-level document could previously read the
whole worktree silently.

The identity chip's host badge overflowed the chip's layout box under
squeeze (Linux CI): every row member can now shrink and truncate, verified
by a width sweep in isolated Chromium down to ~120px chips.

The Allow banner says what it grants: 'Allow folder', reading files in the
named directory, for the life of the preview.

The reliability-gate manifest command, testFiles entry, assertion refs and
dated evidence naming the deleted doc-preview-external-link-bridge.test.ts
are re-issued at doc-preview-external-link-confirmation.test.ts with a
fresh 189/189 run; the focus-gate assertion text follows the shipped gate.

* fix(browser-preview): hide the chip identity row below 24rem instead of clipping it, ellipsize the host badge, catalog the new i18n keys (STA-5758)

CI's preview pane leaves the chip ~40px: no truncation shows anything
there, so the Workspace-file label and host badge now hide whole below a
24rem container threshold sized so that visible implies contained. The
badge text gains an inner text box — text directly inside the flex pill
clipped both ends with no ellipsis. The e2e geometry oracle asserts
containment when the row shows and the threshold when it does not.

verify:localization-catalog: the hardening's new preview keys (and the
renamed allowDirectory) join en.json via sync:localization-catalog.

* feat(browser-preview): batch blocked folders into one access decision (STA-5758)

Sequential per-folder banners trained the allow reflex without adding
judgment — a reader cannot weigh assets/ against data/. The banner now
accumulates every folder a load surfaces, names them (three, then a
count, full list in the title), and grants exactly that set with one
Allow-N-folders click and one reload. Dismiss fences the whole named
set. The map lives behind a ref with a version tick so a dismissal
fences an offer landing in the same event batch.
2026-08-28 00:42:07 -04:00
Jinjing 52ade074a9 Display host on automation details and dialog (#16823)
* Show automation host in details and support moving between hosts

- Rename AutomationCreateDestinationField to AutomationDestinationField to
  reflect dual use in create and edit modes
- Add host display to automation detail view, showing storage authority
- Allow editing automations to move them to different hosts within same
  authority; project list filters to available projects on chosen host
- Update copy from create-only terminology to mode-agnostic wording

* Display automation host and support cross-authority moves

Users can now move automations to different storage authorities. The
destination picker shows all available hosts, and selecting a new one
displays a warning about the move. The save creates the automation on
the destination and deletes it from the source; if deletion fails,
both copies remain and the user is notified.

* Remove cross-authority move support for automations

An automation's authority (the Orca instance that stores and schedules it)
cannot change; edits now only offer hosts within the same authority and
move logic is removed entirely. This simplifies the destination picker and
removes move-specific UI messaging.

* Fix undefined selectedRowKey in automation host recovery

Replace references to the undefined selectedRowKey variable with
selectedRow?.key to properly access the row's key when recovering
automation runs across hosts.

* Support moving automations across execution authorities

Allows users to move automations between different authorities (desktop ↔ runtime environments) during editing. A save to a different authority creates the automation on the destination and deletes the original with its run history. Includes clear messaging about the move operation, proper handling of workspace id resets, and graceful error handling when deletion fails. Supports destination-aware project and worktree fetching.

* Reuse creationKey across move retries when schedule changes

When retrying a failed automation move, dtstart is minted fresh each
attempt, changing the payload. Previously, operationKey included the
full payload, so retries would mint new creationKeys and risk duplicate
automations on the destination if the initial create failed in transport.
Now key only by the move (source + destination) to ensure stable
creationKey across retries.

Also fix workspace auto-selection to use authority-scoped worktrees
instead of the merged cache, preventing unwanted restoration of
source-host workspaces after switching authorities.

* Rename `note` to `moveWarning` for automation host moves

Clarifies that the field specifically warns when an automation would move to another host, replacing the plain storage line.
2026-08-27 20:10:17 -07:00
Neil a7b1da9e14 fix(ssh): a reattach may bind a pane but must never create one (#16751)
`reattachKnownPtys` treats every non-terminated lease as live and calls
`persistPtyBinding`, which had no way to say "bind only". Two of its branches
then rebuild UI the user is not asking for:

- `pty-binding-persistence.ts:133-143` -- `if (args.incarnationId)`
  unconditionally deletes the pane's close tombstone.
- `:145-160` -- on a tab-not-found it mints one via
  `createMinimalPersistedTerminalTab`. The in-code comment names its only
  intended caller: "pty:spawn can beat the debounced writer." Spawn. Reattach
  took the same branch.

This is the mechanism canceled ticket STA-4268 described: "Leases have no pane
incarnation and upsert only by target/PTY. Every nonterminal lease is reattached;
frozen coordinates are passed to persistPtyBinding, which creates and flushes
missing tabs and layout leaves." It was fixed in #13326, reverted by #14361,
re-applied by #14384, and reverted again by #14395 (opened and merged nine
seconds apart, empty commit body). `grep -rn "mayCreate" src/` returns nothing on
main -- the mechanism is genuinely out of the tree.

Four parts:

1. `mayCreate` (default true). When false, one pre-mutation check mirrors all
   four creating branches and returns `false` without mutating, so a refusal
   leaves nothing half-written.

2. The authority gate -- the part both prior attempts lacked, and probably why
   both were reverted. `mayCreate: false` alone refuses in two situations: "the
   user closed it" AND "the renderer has not published its layout yet". The
   second is routine on disconnect->reconnect and is almost certainly the #14361
   tab-loss mechanism. The fence was not wrong; it was UNCONDITIONED. It is now
   passed only when `hasHostAuthoritativeTerminalMembership()` says the persisted
   membership speaks for this worktree, reusing the function already guarding the
   same question at `orca-runtime.ts:8608`. Losing a tab is worse than keeping a
   duplicate, so an unauthoritative session still gets the creating write.

   Authority is read from `local` because that is the partition the write lands
   in -- it is local's absence being interpreted. But a pane the `ssh:<target>`
   partition still holds is not gone, so it keeps its creating write; refusing
   there would strand a live pane behind a binding reattach can no longer reach.
   (SSH spawns bind into `ssh:<target>` while this reattach binds into `local` --
   GH #12721/#12723, STA-3980. This does not fix that split; it refuses to judge
   from one side of it.)

3. `findTerminalTabIdForLeaf` -- bind resolves the tab from the live layout
   instead of the lease's frozen `tabId`. Only the leaf half of a pane key is
   remint-stable: `detachTerminalPaneToTab` moves a live pane into a new tab, so
   a stored tabId names the tab the pane left. Identity vs location.

4. Pane-keyed supersession -- retires siblings on `(targetId, worktreeId,
   leafId)` to `expired`, guarded by the durable binding, which reads `local`
   then `ssh:<target>` so it is correct whichever partition the binding landed
   in. `upsertSshRemotePtyLease` matched `(targetId, ptyId)` alone
   (`ssh-pty-lease-operations.ts:34-36`), so a new relay pty id on reattach minted
   a SECOND lease instead of updating the first, leaving the predecessor
   non-terminated with nothing to retire it.

On refusal the lease goes `expired`, never `terminated` -- `expired` records that
this shell has no surface to reach it through; `terminated` would assert an exit
nothing here observed, which `docs/reference/ssh-execution-boundary.md` forbids.
The remote process is left running.

Deliberately NOT done:
- A collision guard for `upsertSshRemotePtyLease`. Built, tested, and REMOVED --
  its own test passed with the guard disabled, i.e. vacuous. Telling "same lease"
  from "recycled id on a different shell" needs a relay-start identity, which
  would be a twelfth per-tab identity concept; the codebase already carries
  eleven (784 refs) that SSH-v3 Phase 3 deletes. Left as an in-code NOTE. This
  handles lease DIVERGENCE, not COLLISION.
- A port of #13324. Its own authors deleted its load fold and reverted its
  local-only reader in #13326 ("a headless-owned pane still gets a vote before
  its lease is retired"); porting it ships a state-destroying migration they
  removed.
- `bindPaneShell` from #13325. Its purpose is making `isSupersededPtyId` live,
  and that fence does not exist in main. It would have been a refactor plus a
  silently-ignored `mayCreate` -- TS drops excess props through spreads
  (verified), which is why `mayCreate` is passed as a conditional spread here.

Honesty about scope: this does NOT close the daily-tab-growth report. The
deterministic e2e repro (`ssh-lost-kill-tab-resurrection.spec.ts`, later in this
stack) is byte-identical before and after, and instrumentation shows why -- in
that scenario `restoreReattachedPtyRuntime` is never called at all
(`CREATING TAB` 9 hits, `reattach gate` 0, `BYPASS` 0). The fence is on a path
that bug does not take. It is a real, separately-provable defect; it is not the
headline fix, and must not be claimed as one.

Evidence, A/B on this tree. Disabling the authority gate (`mayCreate = true`) and
the supersession call by hand: **8 of 12 fail**, including "does not resurrect a
tab whose closing pty.kill failed with a transport error" and "holds the live
lease count flat across ten reconnects of one pane"
(`[ 'pty-0', 'pty-1', 'pty-2', …(7) ]` vs `[ 'pty-9' ]`). The 4 that pass both
ways are the over-refusal tripwires, which is the point of having them. Restored:
12 passed; 2,439 passed across `src/main/ssh`, `src/main/persistence` and
`src/main/runtime/workspace-session`.
2026-08-27 19:45:49 -07:00
Neil 7ee8b5e1a6 Refactor lower max-lines modules (#16760) 2026-08-27 16:10:51 -07:00
Neil b241a68ae4 Fix worktree identity collisions across hosts (#16691)
* fix(workspaces): add collision-safe worktree identity

* fix(workspaces): read worktree metadata per host and repair ambiguous identities

The canonical identity store landed write-only: getWorktreeMetaForHost had no
production callers while setWorktreeMetaForHost kept the legacy projection only
for the first known owner, so a second host's edits persisted and were never
read back. Wire the listing paths through host-qualified reads.

An ambiguous alias was also unrecoverable — reads returned undefined and writes
threw forever, and the throw escaped the detected-worktree loop, emptying the
whole repo's sidebar. Fail open onto the most recently active instance instead.

- collapse ambiguous aliases deterministically and persist the repair
- reclaim identity rows in the metadata GC so they cannot outlive their locator
  or resurrect onto a worktree recreated at the same path
- drop every host's rows when a locator is removed outright, not just the owner's
- honour an explicit instanceId so the stale-lineage rotation guard still works
- scope a rename to the moving host; other hosts keep their own locator
- prefer the project host setup matching the repo's own execution host, so a
  repoId registered on two hosts no longer stamps the wrong one durably
- reject an unencoded `|` in a host id, the invariant the alias delimiter needs
- drop the never-populated hostGeneration from the canonical key

* fix(workspaces): close remaining identity review gaps

* fix(workspaces): close remaining review gaps

* fix(workspaces): address review and CI regressions

* test(workspaces): update host-qualified metadata expectations

* fix(workspaces): preserve ambiguous identity records

* fix(workspaces): snapshot metadata during listing

* test(workspaces): mirror listing metadata snapshot in windows fixture

* fix(workspaces): preserve identity routing for metadata writes

* fix(workspaces): scope stale metadata cleanup by host

* fix(workspaces): rekey identities on SSH readoption

* fix(workspaces): fail closed for ambiguous board ids

* perf(workspaces): snapshot metadata across catalog listing

* fix(workspaces): retain neighboring manual order updates

* test(workspaces): cover ambiguous board id index

* fix(persistence): harden host-qualified worktree metadata

* refactor(shared): split project host setup lookup

* refactor(workspaces): simplify host-qualified metadata
2026-08-27 15:08:40 -07:00
Jinjing 9635e6822f Normalize project catalog rows defensively (#16826)
* wip: normalize untrusted project catalog rows at load boundaries

* fix(catalog): make ProjectHostSetup field types true at the ingest boundary

Crash 3bcc5be3: a setup row whose repoId arrived null reached Settings'
projectByRepoId memo and threw on .trim(). The type said `string`; persisted
JSON and remote hosts on other versions can disagree.

Normalize project/setup rows where untrusted data enters typed code — the
persisted-state load (marking dirty, which is the migration), profile
transfer reads, the repo-derived projection, and the renderer's IPC/RPC
ingest and adoption steps — instead of re-guarding each consumer. Also
covers `setup.path`, whose identical crash is on the sidebar render path.

Coercion only: never drops a row, never adds or removes an optional key, and
returns input references when a row already conforms, so selector and useMemo
identity is unchanged.

* refactor(catalog): drop the `as` casts the normalizer introduced

A change whose thesis is "stop the declared types from lying" should not use
`as` to paper over types.

The four source casts all came from the row normalizers returning `readonly`
arrays into mutably-owned fields. Take and return mutable arrays instead, and
copy at the one caller that holds a readonly projection — where identity is
not load-bearing, unlike the persistence path, whose dirty check compares it.

Tests built deliberately malformed rows by casting a literal. Build a valid
row and `Reflect.set` the bad value onto it, which says outright that the
fixture violates its type; parse the non-array case from JSON, which is how
it actually arrives. Fixtures that were merely incomplete needed no cast at
all — `Repo` requires only the five fields they already had.

Also narrow normalizeLoadedProjectCatalog to the two fields it reads.

* refactor(profiles): validate untrusted profile JSON instead of asserting it

Removes the last introduced cast and the unsound pre-existing ones in the
file this change already touches.

The test cast is gone because narrowing normalizeLoadedProjectCatalog to the
two fields it reads made `{}` assignable on its own.

arrayOrEmpty and recordOrEmpty checked the shape and then asserted the
element type, which is the same "declared type is a promise the data does not
keep" problem this change exists to fix. Array.isArray already narrows on its
own, and a generic isRecord narrows the value part, so both assertions delete
outright. JSON.parse returns any, so annotating the binding beats asserting
its result.

The two remaining `as const` in project-host-setup-actions are untouched and
deliberate: a literal assertion narrows a type rather than overriding it.

* Make project catalog normalizers handle null rows defensively

Instead of crashing when corrupt or null catalog rows are encountered,
the normalizers now gracefully repair them with default values. Refactored
helper functions for clarity and added type predicates to improve type
narrowing.
2026-08-27 12:35:03 -07:00
Brennan Benson 8a07bbd8cf fix(orchestration): enforce nested worker depth instead of an accidental fence (#16668)
* fix(orchestration): enforce nested worker depth instead of an accidental fence

Orca documented that "dispatched workers cannot spawn their own sub-workers
(worker-start is coordinator-fenced)". No such check existed. What existed was a
single Run-binding check in the workerStart RPC: a worker's terminal is not bound
to a Run, so worker-start happened to fail. The rule was emergent, asserted by no
test, and written in no doc — and it leaked. A worker could run-create its own
Run, task-create, and worker-start: now bound, the check passed.

Replace it with a real, configurable depth cap.

Depth is derived from the caller's own active Dispatch rather than from Run
binding, which is what dissolves the run-create bypass: creating a Run does not
stop you being a worker. Enforcement lives in a single dispatch-row writer that
owns all three INSERTs that mint a live worker — the generic claim, the supervised
worker-start path (including every retry), and the remote attachment. Two of those
were missed by earlier drafts of this change, so `creator` and `maxDepth` are
required parameters: a new spawn path cannot compile without deciding, and a
boundary test refuses the SQL anywhere else.

Schema v30 adds depth to dispatch_contexts and remote_dispatch_attachments,
NOT NULL DEFAULT 1 and backfilled to 1 so an unstamped or pre-upgrade row fails
closed rather than reading as a root coordinator. The attachment pane indexes
widen to the five states in which a remote worker may still be running:
loss of contact is not evidence of process death, so an unverifiable worker still
counts as a nesting parent.

Also adds the caller-evidence assertion that workerStart was the only Run-scoped
verb to skip, so a declared --from cannot name another terminal's pane and inherit
its depth.

Default is 1, so behaviour is unchanged unless the new setting is raised. Two
limitations are deliberate and documented rather than papered over: this is a
guardrail and not a security boundary, since a caller whose launch evidence is
unverifiable (any ordinary restored terminal) can declare another handle; and it
is enforced at supervised dispatch creation, so a settled worker whose process is
still alive counts as a root again.

* fix(orchestration): share caller resolution and pin worker gaps

* refactor(orchestration): make the caller resolver's pane contract explicit

Overloads so requireStablePane callers get a non-null string instead of casting,
and rename the attestation opt-out to say what it means: the caller asserts it
itself. A flag called assertEvidence:false reads as "attestation optional",
which is the hole this helper exists to close.

* fix(orchestration): propagate dispatch depth to federated workers

* chore(cli): refresh bundled orchestration guide
2026-08-26 13:22:09 -07:00
Jinjing cda2280d63 Show all automations (#16532)
* Add all-host automations with scoped ownership and multi-authority suppo

Enable automations to run on multiple hosts (SSH targets and local) with
owner-fenced mutations, scoped list queries per host, and conflict
resolution. Introduces desktop and runtime authorities as distinct
automation storage owners, with per-host caching, invalidation, and
retry scheduling on the renderer. Captures registration generations for
SSH hosts to survive re-adoption. Adds CLI support for destination
selection and conflict recovery.

* Filter automation create projects by destination host

Only offer projects available on the selected destination, preventing
the mismatches that would fail at submit time. Auto-adjust the project
selection if it becomes unavailable when the destination changes.

* Add runtime storage authority support for automations

- Support both runtime and desktop as automation storage authorities
- Make owner preconditions optional for legacy-client compatibility
- Cache automation list projections to improve performance
- Add per-row repo/worktree resolution for cross-authority collisions
- Extend automation.list RPC to always include owner metadata

* Replace child_process.execFile with runProcess for external automations

- Migrate external-manager to use cross-platform runProcess wrapper per child-process safety policy
- Abstract electron app/ipcMain APIs in orca-runtime via environment accessors
- Install fake app environment in automation tests for consistent setup
- Reorganize imports to use specific module paths (ssh-target-registry, agent-detection, browser-error)
- Remove external-manager from child-process import allowlists (no longer violates direct import)

* Unify desktop automation CRUD onto the local runtime RPC surface

The desktop authority now speaks the same automation.* RPC contract as
remote runtimes, via callRuntimeRpc({kind:'local'}) -> runtime:call ->
the shared RpcDispatcher. The automations:list/listRuns/create/update/
delete/runNow IPC arms, their preload members, and every renderer
desktop-vs-runtime transport fork are retired; the runtime methods are
the single implementation of scoped lists, owner fencing, and change
publication for both transports (mobile clients already exercised them).

The desktop probe scheduler's priority lease survives the move as an
AutomationService hook the IPC registration installs and the runtime
methods take, so Orca's own automation traffic still parks queued
external-manager probes.

External-manager scope arms and dispatch-loop plumbing stay on IPC by
design; automation change events keep their existing channels (renderer
ingestion already converges them by authority).

* Remove automation ghost SSH tombstone scanning

This functionality for synthesizing tombstones for automation-referenced SSH
targets is no longer needed as part of the automation system refactoring.

* Refuse orphan automations at dispatch time, not migration time

Remove migration-time disabling of orphan automations and the `enabledDecidedBy` field. Dispatch now refuses orphans at runtime instead, simplifying state management and UI. Orphans are left unstamped and enabled; dispatch refuses to run them via `resolveAutomationRunTarget`.

* Show all automations in flat table with unified filter menu

- Replace host picker component with comprehensive Filters menu supporting status, last run, agent, and host filters
- Flatten automation list layout to single table instead of host-grouped sections
- Add Host column to display execution host for each automation
- Display active filters as removable pills below toolbar
- Delete unused AutomationHostPicker* components

* Add automation owner fencing and destination validation

- New AUTOMATION_OWNER_FENCING_RUNTIME_CAPABILITY for owner preconditions; legacy clients get owner metadata snapshotted at RPC boundary for compatibility
- Editor captures and revalidates automation destination before save, preventing silent retargeting if SSH infrastructure changes mid-edit
- SSH target types now isolate renderer-authored fields; generation is server-owned and stripped by IPC handlers

* Route automation recovery actions to the origin host

When an automation action fails due to owner fencing, recovery verbs
("Update server", "Reconnect") must run on the host where the refusal
originated: the row's captured owner for row operations, or the
destination the create dialog captured, not the list's filtered host.

* Remove external manager scope limitation notices

Consolidate create destination eligibility checks with a unified predicate
and fix the bug where desktop repo IDs could be sent to runtime hosts where
they cannot resolve.

* Persist only store-derived automation contexts, not client-perspective o

Store contexts must never be based on client-provided runContext or sourceContext
values—clients speak a different perspective (e.g., 'runtime:<id>' for host IDs
they assign), and persisting those makes the store projection orphan automations
it actually owns. Derived contexts now take precedence in create and update paths,
with explicit null still honored to clear a value. Tests verify this by simulating
drift after storage and confirming that moves re-derive while toggles preserve.
2026-08-26 09:50:12 -07:00
Jinwoo HongandJinwoo-H a9781a4118 STA-4150: client-hosted remote browser (consolidated) (#15448)
Co-authored-by: Jinwoo-H <jinwoo@stably.ai>
2026-08-25 15:36:51 -07:00
Neil 9bcaf09869 refactor loading store into cohesive domains (#16192) 2026-08-24 20:45:53 -07:00
Jinjing aa32871a61 Improve cmd j ranking with recency (#16281)
* Track container-only tokens and tab focus for cmd+j ranking

Previously ranked by whether any container-only matches existed (boolean);
now counts tokens matching only containers for finer-grained ranking. Tab
focus recency is now tracked explicitly so recent refocuses rank above
stale worktree activity. Preserves worktree grouping by input order while
applying focused-group MRU within each block.

* fix(cmd-j): preserve duplicate recent tab occurrences

* fix(cmd-j): preserve host scope during worktree purge

* fix(cmd-j): scope repo purge for exact-id host twins

* fix(cmd-j): scope ssh visit recency to local to survive restarts

Boot hydration loads only local + runtime:* partitions, so routing
ssh-qualified recency to ssh partitions strands it across restarts.

- Keep ssh-qualified visit timestamps in local partition
- Route runtime-qualified keys to their partition
- Remove groupId from recent tab occurrence base (unstable on regroup)
- Collapse bare and host-qualified timestamps, preserving max
- Simplify repo pruning host-match logic
- Add robustness: optional chaining, helper function

* Scope focused tab recency by worktree to fix Cmd+J ranking

Tab ids can be duplicated across worktrees; scoping recency keys to per-worktree prevents one worktree's MRU position from overwriting another's in Cmd+J. Scope worktree order blocks to (hostId, worktreeId) to keep same-id worktrees on different hosts separate.

Also fix recency preservation during partial identity migrations and prune orphaned host keys on removal.
2026-08-24 11:23:16 -07:00
Neil 03fcfdfb92 feat(orcad): boot the Orca runtime on plain Node (#15968)
* refactor(host): resolve the app root through the port in fork-reachable modules

`parcel-watcher-entry-path.ts` and `session-scanner-service-entry-path.ts` read the
app root via `require('electron').app` inside a try/catch that already returns null
when Electron is absent. They were therefore correct under plain Node at runtime and
only failed the *static* text check — which is real, not pedantic: the comment in
`ports/port-scan-command-client.ts:19` records that the plain-node-entry-guard fails
on that literal text, try/catch or not.

`hasAppEnvironment() ? getAppEnvironment() : null` gives the identical "no app root
here" answer without the text. That restores `hasAppEnvironment`, which an earlier
commit in this stack deleted as unused — it now has the caller it was waiting for.

Ratchet baseline 27 → 25.

Verified: 74 files / 458 tests; `pnpm typecheck` clean; `oxlint` clean.

* feat(orcad): boot the Orca runtime on plain Node

Closes the last two Electron couplings and makes `orcad` a working artifact:
a 4.43 MB Node bundle that boots, pairs, registers a repo, creates a real git
worktree and round-trips a PTY — with zero `require("electron")`.

Ratchet 2 -> 0, so `config/runtime-electron-baseline.txt` is now empty and its
test asserts exactly that: any reachable electron import is a regression.

- speech: inject the service factories, so importing ModelManager for its type
  no longer drags Electron's streaming net.request into the graph
- filesystem-watcher: add a WorktreeWatcherRemoval port. Every entry in those
  maps arrives through an ipcMain handler carrying a renderer sender, so a host
  with no renderer has nothing to close, restore or forget — the inert default
  is what the desktop code does against empty maps, not a stub hiding work
- user-data-path / profile-storage-paths: resolve userData through
  AppEnvironment. These surfaced only once orcad pulled the store in

Both host ports now anchor to a realm-global symbol. `vi.resetModules()` gives
the re-imported graph a fresh module copy, so a binding installed before the
reset silently read back as uninstalled.

The acceptance smoke drives both hosts through one code path (`--target
orcad|electron`) and seeds its own git repo, so it is hermetic and asserts the
same contract of each. Wired into PR CI.

* test(smoke): remove the seeded workspace container, not just the worktree

* test(smoke): surface the server's stderr when it dies before ready

* fix(smoke): build node-pty for Node before booting orcad in CI

* fix(smoke): drive the CLI built from this checkout, not one on PATH

* docs(ratchet): say the baseline must stay empty, not merely shrink

* build(orcad): externalize only the native modules actually in the graph
2026-08-22 21:47:46 -07:00
Jinwoo Hong da6b9d8065 fix(terminal): stop orphaning live agent terminals across host restarts and graph syncs (#15644) 2026-08-21 17:11:17 -07:00
Brennan BensonandQA 64de8dd637 fix(workspaces): delete on the confirmed host, and make both hosts' rows selectable (STA-4343) (#15013)
* fix(workspaces): host-qualified workspace deletion (STA-4343, STA-4448)

Squashed integration of PR #14606 + the codex review-loop output, replayed
onto current main. Granular history preserved on brennanb2025/sta-4343-review-full.

Fixes the regression from #13413: a workspace id is repoId::path with no host
component, so the same repo at the same path on two hosts published one id for
two workspaces, and deletion routed by that id landed on whichever host routing
preferred - usually the ACTIVE one, not the row the user confirmed.

- removeWorktree takes a REQUIRED host-qualified WorktreeRemovalTarget; omitting
  the host is a type error. All destructive callers migrated.
- Projections dedup on (host, id), so two hosts render as two selectable rows
  while the createWorktree/fetchWorktrees race duplicate still collapses.
- Ephemeral VM cleanup is host-scoped. It matched on bare workspaceId, so the
  host-scoped delete path destroyed the SURVIVING host's VM and its unpushed
  filesystem - a leak fix that had become data destruction.
- Selection, keyboard routing, lineage grouping and Space rows carry host
  identity end to end; fixing the executor dedupe alone would have turned
  one-row intent into deleting both hosts.

Files split to stay under max-lines rather than raising any cap.

* refactor: split files that crossed max-lines

The review-loop commits used --no-verify, so the pre-commit hook never
enforced the caps. Extracted cohesive units rather than raising any limit:
renderer teardown, delete-with-toast, pinned-group rows, host-scope helpers,
workspace-kind predicates, filter actions, kanban drag selection, the
renderer removal result type, and the native-chat persistence tests.

* refactor(workspaces): extract cleanup deletion-phase selector

Clears the last max-lines violation and the import-type side effect the
changed-code gate flagged.

* refactor(sidebar): track the delete-dialog extraction modules

* fix(workspaces): preserve host identity across remaining surfaces

* fix(sidebar): re-carry host through the rewritten palette result model

#15170 replaced PaletteSearchResult while this PR was open. Re-applied the
host qualification on top of the new model instead of taking either side:
results carry worktreeHostId again, and the board filter keys its matched
set on host identity rather than the bare id.

Known gap, documented in the board test rather than deleted: searchWorktrees
resolves evidence through a `documents` map keyed by BARE worktree id, so two
same-id host rows collapse before this code sees them. Closing that belongs
with the palette work.

* test(cmd-j): pin the palette collision gap instead of asserting the old model

The palette collision test asserted two host-qualified rows, which #15170's
rewrite made unreachable: item ids are bare again and worktreeMap is id-keyed.

Rewritten to assert what holds — activation always names a host — and to pin
the defect it exposes: two same-id rows render on ONE command value, so React
sees duplicate keys and a click on the first row activates the second row's
host. That reproduces on main, so it is pre-existing, not from this PR. Pinned
rather than deleted so fixing it must update this test.

---------

Co-authored-by: QA <qa@local>
2026-08-17 15:57:07 -07:00
Jinjing 7c79a0f9e3 fix(persistence): harden persistence edge cases (#15171)
* fix persistence edge cases

* Persist original folderPath value without trimming

The guard validates that the trimmed path is non-empty; persist the original input value that passed validation rather than a transformed version.

* Fix cross-host pane conflicts and persistence edge cases

Prevent ambiguous routing when tab IDs are shared across host partitions by skipping alias registration for colliding tabs. Ensure repaired null lineage maps are marked as changed so they're re-saved on reload. Use execution host instead of connectionId for git username enrichment to handle runtime repos correctly.
2026-08-17 15:25:25 -07:00
Brennan Benson a7f1653415 fix(worktrees): keep retirement tombstones across project and SSH target re-add (#14917)
* fix(worktrees): keep retirement tombstones across project and SSH target re-add

Generated workspace names are retired so a name is never reissued onto a cwd
that still holds another workspace's Claude/Codex history. Two re-add paths
lost that record.

STA-4449 (local): retirement was stored only under `repo.id`. Removing a
project deletes that row and re-adding the same path mints a new id, so the
new repo starts with an empty registry. The on-disk backfill normally
re-seeds local repos, but it cannot recover a name whose only surviving
evidence is a Codex rollout JSONL — those are deliberately not scanned — so a
name spent under Codex with its workspace directory gone came back.

STA-4491 (SSH): the second, path-derived copy embedded the SSH target row id.
Row ids are minted fresh on every re-add, so `ssh:ssh-old:...` became
`ssh:ssh-new:...` and `reassignSshTargetId` migrated other carrier state but
not the retirement namespaces.

Keying the store on the namespace instead of `repo.id` was rejected in
`nestWorkspaces`, `worktreeBasePath` and `repo.path`, so a settings toggle
would orphan every retirement at once. `repo.id` therefore stays primary and
the path-derived namespace stays a mirror — a settings toggle loses the
mirror but keeps the repo row, a re-add loses the repo row but keeps the
mirror, and reads union both.

- Mirror local repos into the namespace too, not just remote ones.
- Key the namespace's host half on the SSH endpoint (host+port+username), the
  thing that actually decides which filesystem a path lands on, instead of the
  target row id. Reads also accept the pre-identity key so an upgrade keeps
  tombstones it already wrote, and `reassignSshTargetId` re-keys the rest.
- Cap the namespace map, which by design outlives the repos that wrote it and
  so has nothing to prune it per repo.

Endpoint identity is extracted from ssh-target-readoption.ts, which already
compared these fields for exactly the same reason, so re-adoption and
retirement cannot drift apart.

* fix(worktrees): copy shared SSH endpoint retirements instead of moving them

An endpoint identity is not owned by the target row that rotates: nothing
dedupes SSH targets by host|port|username, so a second live target can still
resolve to the same host. Moving the bucket stripped that target's tombstones
and reissued a path whose agent history is still on disk.

Row-id identities stay a move — reassignment leaves nothing pointing at them.

* fix(worktrees): carry retirement mirror across in-place SSH endpoint edits

Config sync rewrites host/port/username on the existing target row and a
runtime-owned target takes a fresh address from every provision, both keeping
the row id. No re-adoption runs, so nothing carried the endpoint-keyed mirror
across and it stranded — strictly worse than the pre-change key, which was the
row id and was invariant under these edits.

Also bound the map after a migration: a retained source bucket grows it, so the
cap has to be applied there too, and compare registries by membership rather
than size so an uncompacted destination cannot trade a folded name for a new
one and read as unchanged.

* fix(worktrees): skip retirement migration for on-demand runtime targets

An on-demand VM is discarded between provisions, so its fresh address reaches an
empty filesystem and a reissued name collides with nothing. Migrating there
would spend names against history that no longer exists, and because each
provision mints another address it would add a namespace bucket per run,
evicting the real tombstones of local and ordinary SSH repos.

* fix(worktrees): stop the namespace cap evicting what a migration just wrote

Two defects with one root cause. Assigning to an existing key leaves it in its
original insertion slot, so a merged destination kept the oldest position and the
trim deleted the bucket it had just enriched. Retained source buckets are older
than the destinations a copy appends, so at the cap the trim removed exactly the
sources the copy existed to keep — silently turning it back into a move. The trim
now exempts the keys the migration wrote or deliberately kept.

Also stop on-demand runtime workspaces writing namespace mirrors at all: each
provision reaches a discarded filesystem under a fresh address, so the entry can
never be read back and only spends a capped slot that a local or SSH project
needs. The repo-id row still records the name for the live session.

* fix(worktrees): re-insert migrated namespaces so the cap cannot undo a migration

Exempting keys from the trim protected them for that one call and no other. A
merged destination keeps its original insertion slot, so it sat at the front of
the eviction queue and the next unrelated retirement write dropped it — losing
both the migrated name and the name the destination already held, on a host that
had just been re-added.

Re-insert what the migration writes instead, the same discipline the ordinary
writer already follows, so insertion order reflects use. That also removes the
exemption, which could otherwise leave the map stuck at twice the cap until one
later write evicted the whole excess at once.

Corrects the runtime-gate comment as well: the mirror is unreadable after the
next provision, not immediately, so a remove/re-add inside one provision is a
real if narrow loss.

* fix(worktrees): refresh a retained namespace source even when its merge adds nothing

Replacing the trim exemption with re-insertion narrowed the protection: the
exemption covered every retained source, the re-insertion only covered sources
whose merge actually wrote. A copy whose destination already held the same names
was then neither re-inserted nor exempt, so the migration's own trim evicted the
shared source bucket ahead of hundreds of untouched ones — losing the tombstones
of a live sibling target still on that endpoint, which is what copying exists to
prevent.

A move's destination gets the same treatment: deleting the source makes it the
only remaining copy, so it has been used. Both are order-only and deliberately do
not set the changed flag, keeping an import that moved nothing from scheduling a
save.
2026-08-17 12:23:50 -07:00